Drug discovery has a human biology problem
Despite ongoing advancements in genomics, artificial intelligence, and precision medicine, clinical trial success rates have remained low, with approximately 90-95% of drug candidates failing during clinical development. This high attrition rate, rather than a lack of computational or financial resources, indicates a translational gap.
Preclinical models currently utilised in candidate selection and progression frequently fail to recapitulate human physiology accurately. While in vivo animal models and immortalised cell lines have historically supported early-stage research, their biological constraints can limit their predictive validity. Optimising clinical success rates and reducing development expenditure requires the integration of physiological models derived directly from human biology at the outset of the discovery pipeline.
The translational gap: Why traditional models limit predictive accuracy
This divergence between preclinical efficacy and clinical outcomes is largely driven by biological complexity. Complex conditions such as Alzheimer’s disease, Parkinson’s disease, neuropsychiatric conditions, and various rare genetic disorders possess distinct human etiologies that are difficult to replicate in non-human systems.
Consequently, evaluation in animal models offers only an approximation of human disease pathology and drug interactions, rather than a precise replication. This discrepancy causes substantial financial asset loss and extended development timelines when candidates with positive preclinical profiles fail to demonstrate efficacy or acceptable safety profiles in human subjects. Ultimately, late-stage attrition in Phase II and III clinical trials drives the rising cost of drug development and delays the availability of new therapeutic interventions for target populations.
Mitigating these risks requires preclinical models that more accurately recapitulate human physiology prior to clinical trials. Human induced pluripotent stem cell (iPSC)-derived models provide a viable methodology to address this requirement. By producing specific, functional human cell types from a renewable, scalable source, researchers can evaluate disease mechanisms and therapeutic responses within a relevant human biological context before initiating clinical trials.
The reproducibility challenge: Ensuring consistency in preclinical data
Transitioning to human-centric biology requires addressing data consistency alongside physiological relevance. For human cell models to effectively validate drug discovery pipelines, they must yield reproducible data. Preclinical research often encounters significant experimental variance, whether that is between lots, between labs, or between sites. Traditional immortalised cell lines frequently lack normal physiological function and are subject to genetic drift, while primary human cells isolated from separate donors exhibit substantial genetic and phenotypic variability. This baseline divergence, while useful at the macro scale to represent population diversity, introduces confounding variables, making it difficult to determine whether an observed effect is a true biological response or a model-induced artefact.
As drug discovery becomes increasingly collaborative, from pharmaceutical companies, contract research organisations (CROs), and academic laboratories, standardisation is essential. Utilising uniform, scalable human cell models minimises unnecessary experimental noise. This technical consistency ensures that datasets generated across different global sites can be accurately compared, providing a reliable foundation for data-driven decisions throughout the development pipeline.
Computational models depend entirely on biological data quality
Artificial intelligence systems are increasingly deployed across drug discovery pipelines, ranging from target and biomarker identification to predicting the behaviour of chemical compounds. While these computational advancements offer significant utility, machine learning models remain entirely dependent on the quality and fidelity of the biological data used to train them. Inconsistent or low-quality datasets inherently limit algorithmic performance, as training models on data that fails to accurately reflect human biology risks embedding those exact limitations into every predictive output.
To realise the full potential of AI in drug discovery, the industry requires LLMs to be trained on precisely controlled biology. In order to learn the language, these systems must be fed clean, reproducible, matched healthy and disease-state human datasets. Standardised human cell models serve as a critical mechanism for generating these inputs, producing data that accurately reflects human physiology and yields consistent experimental results. Rather than being competing approaches, AI and human cell models function symbiotically; analytical algorithms require high-quality, large-scale biological datasets to refine their predictive accuracy, while advanced biological models become substantially more valuable when paired with computational tools capable of analysing complex data.
Implications for the future of drug discovery
To improve the translation of therapeutics from the laboratory to the clinic, the pharmaceutical industry must address three interconnected challenges. These comprise: enhancing clinical translation through relevant human models, eliminating experimental noise via strict data reproducibility, and generating high-quality human datasets to power next-generation, AI-driven drug discovery. The broader adoption of human cell models offers a scalable solution to each of these hurdles.
Achieving this scale, however, requires navigating an unprecedented shift in production volume. The industry currently relies on ~20 million animal models and trillions of primary and immortal cell lines annually. Transitioning to models that more closely resemble human biology necessitates advanced manufacturing technologies capable of delivering billions of human cells, all while maintaining uncompromising quality and remaining compatible with standard research and development budgets.
The transition toward these models is driven not only by the ethical considerations surrounding animal research, but by the operational necessity of utilising systems that accurately mirror human biology, interface effectively with in silico drug discovery workflows, and can be sourced at the vast scale required. The current opportunity lies in optimising existing discovery methodologies by generating more predictive, reproducible human data. By enabling better-informed decisions earlier in the development lifecycle, the industry can increase confidence in preclinical research, mitigate late-stage failure rates, and accelerate the delivery of safe, effective therapies to patients.
About the author
Emma Pepperell, CEO, joined bit.bio in August 2024. She spent 10 years at abcam gaining extensive experience of a fast-growing products business, across Product Management, Marketing, Portfolio & Product Development, and Sales functions, with prior experience in life sciences product distribution. Pepperell has a DPhil from the University of Oxford, where she studied the regulation of stem cell engraftment.
