Statistics • Software • Applied AI

I build tools that make scientific and technical work more reliable, interpretable, and useful.

My work sits at the intersection of data science, production software, and scientific decision-making. I’ve built statistical pipelines, internal lab-automation tools, and AI-assisted systems that support real-world teams.

Headshot of Lillian Tatka

About

I’m a data scientist who works at the intersection of statistics, software, and decision-making. My work at a seed-stage biotech company spans model research and development on one end and production automation and tooling on the other.

I completed a PhD in Bioengineering Data Science, where I learned to translate ambiguous problems into concrete solutions: production tools, metrics, evaluation frameworks, and actionable decisions.

I’m also drawn to the places where technical work meets business judgment. I’ve supported decisions in quantitative finance and venture capital (due diligence for a VC firm, pricing models at an investment bank). I hold a graduate certificate in technology entrepreneurship from UW, have worked at two early-stage startups, and coached at an entrepreneurship program for high schoolers.

Resume

Selected experience

  • Built a production proteomics hit-calling pipeline from scratch and maintained it for over a year.
  • Owned a lab-automation tool that reduced pick-list generation from hours to about two minutes.
  • Led technical work on a prospective validation effort for a proprietary compound-target prediction model.

Core skills

  • Python, statsmodels, polars, Streamlit
  • Experimental design, hypothesis testing, FDR correction
  • ML evaluation, LLM tool-use systems, MCP server development

Classic resume snapshot

Data Scientist / Technical Builder Current role

Developing statistical, software, and applied AI systems for scientific decision-making in biotech.

Selected highlights Key experience
  • Built and maintained a production proteomics hit-calling pipeline using Python and statsmodels.
  • Created an internal lab-automation tool that reduced workflow time from hours to minutes.
  • Led technical work on model evaluation and validation for drug-discovery applications.
Education PhD

PhD, Bioengineering Data Science. Graduate certificate in technology entrepreneurship, University of Washington.

Projects

Statistical hit-calling pipeline

Built a production statistical engine for proteomics screening from scratch, evaluating multiple modeling approaches and implementing fixes for p-value calibration and variance stability.

Liquid-handler pick-list generator

Replaced a manual, hours-long lab workflow with a backend-plus-Streamlit tool that reduced generation time to around two minutes and made the process self-serve for wet-lab staff.

ML model evaluation for hit-to-lead prioritization

Designed a rigorous retrospective study to assess whether an internal deep learning model added value over random selection, and used the results to prevent adoption of an unvalidated capability.

Conversational data access tool

Architected an MCP-based internal tool that lets users query a large data corpus through natural-language chat, with an evaluation harness designed to validate tool-use behavior before release.

Publications & research

My PhD work reflects early technical training in computational biology and scientific modeling, while my current focus is on applied statistics, software, and real-world decision support.

Selected first-author work

Contact

I’m most easily reached on LinkedIn.