The Virtual Cell Is an Experiment Engine
The new claim is learned biological state: models trained across perturbations, omics, structures, and images that can be queried like an experiment engine.
How did we get here? Scientific computing already modeled parts of biology; AI is trying to learn the shared state underneath them.
The field is moving from one-off predictions to experiment loops.
The older pattern was a model for one assay, one modality, or one target. The newer ambition is a learned cell state that can be perturbed, read out, and used to plan the next test.
Apply a drug, knockout, knockdown, overexpression, or combination.
Ask for expression, structure, regulation, morphology, or other measurements.
Use the prediction to prioritize the next wet-lab experiment.
Decision-making is the anchor.
The clear way to talk about a virtual cell is as an experiment loop. The model proposes or evaluates a perturbation, produces readouts, and helps decide what evidence should be collected next.
The difference now is learned biological state.
AI virtual cells are being built from large biological data across measurements and scales. The model learns representations from data instead of only executing hand-written equations.
Scientific computing is the baseline.
Computational modeling is already part of biology. The new claim is that large neural models can learn a transferable representation of cellular state from growing experimental data.
A world model keeps state between interventions.
GenBio describes AIDO Cell as a world model: a unified latent cell state that can be updated by an intervention and decoded into several predicted readouts.
Illustrative interaction: the intervention updates one model state, then decoders produce several kinds of output.
One latent state supports several readouts.
AIDO Cell is described as a learned latent state. An intervention updates that state, decoders produce the requested readouts, and the control layer preserves the state across consecutive interventions.
Drug discovery is the pressure point.
The useful direction is experiment triage and molecular design: test many candidates in silico, then spend lab time on the ones most worth checking.
GenBio says roughly one compound reaches the clinic for every 10,000 entering the drug-development pipeline.
In a K-562 imatinib case, GenBio says several candidates recovered this share of imatinib's differentially expressed genes.
Triage is the business case.
A useful model can reduce the number of unpromising wet-lab paths or prioritize better ones earlier. That changes cost, time, and scientific attention.
The test is transfer to new biology.
Benchmark papers warn that models can look strong under familiar splits and lose performance under new cell contexts, new perturbations, or cross-dataset evaluation.
Does it generalize beyond the cells it already knows?
Does it handle a new intervention?
Does performance hold across labs and platforms?
Does the prediction survive a real experiment?
Validation is the story.
The field is building models, data infrastructure, and benchmarks at the same time. The honest question is whether the systems transfer to new biological contexts and change wet-lab decisions.