Three decades of Indian drug discovery: analysis of 1,095 proprietary drug pipeline trajectories
Background: Limitations of a Linear Generic Replication Paradigm and the Transcriptome-Proteome Data Bottleneck in Indian New Chemical Entity (NCE) R&D
The Indian pharmaceutical industry has been entrenched in a generic reverse engineering-centric manufacturing model for half a century since the 1970 patent law revision. ANDA pipelines optimized for Hatch-Waxman bioequivalence (BE) testing have formed the industry's backbone, but this static framework lacks the computational infrastructure for exploring the target protein-ligand binding free energy landscape essential for de novo NCE discovery, ADMET tensor synchronization in the hit-to-lead optimization phase, and single-cell transcriptome noise correction for cellular heterogeneity. In particular, India's unique pharmacogenomic heterogeneity—inter-ethnic variations in CYP2D6 and CYP2C19 allele frequencies—has acted as a baseline confounding variable in global Phase III bridging studies, leading to false-positive efficacy signals or systematic distortions in the dose-response curve. While Sun Pharma, Dr. Reddy's, and Biocon launched NCE programs in the early 2000s, the translational success rate has been significantly lower than the global average of 6.2%, attributable to a structural blind spot: the absence of an in silico virtual screening-wet lab closed-loop. The meta-dataset of 1,095 projects—195 companies, 462 ongoing—captured in this review is the first to panoramically reveal the quantitative dimensions of this bottleneck.
Findings: Activation of Multi-Modal Molecular Architectures and Empirical Demonstration of Pipeline Scale-Independent Tensor Synchronization
The most disruptive finding revealed in this analysis is that the pace of modality diversification in the Indian new drug ecosystem has crossed an exponential inflection point since 2015. While small molecule compounds (811) still dominate, with 684 (84.3%) being NCEs, a multi-modal tensor of monoclonal antibodies/antibody-drug conjugates (ADCs) (114), CAR-T/NK cell therapies (40), gene silencing (siRNA/ASO) (24), and genome editing (CRISPR) (7) is simultaneously being activated. This represents a topological expansion of the molecular diversity landscape, encompassing Bharat Biotech's DNA vaccine platform, Biocon Biologics' biosimilar-to-innovative antibody transition, and ImmunoACT's NexCAR19 (India's first approved CAR-T). In particular, the 22 ADCs align with the global ADC market CAGR of 29.4% (2024-2030, Grand View Research estimate), with the Aurigene-Daiichi Sankyo co-developed pipeline being representative. In terms of clinical phase distribution—Phase I (98), Phase II (63), Phase III (19), approved (21)—the Phase I→II transition rate of approximately 64.3% exceeds the global oncology average of 52%, but the sharp bottleneck in Phase II→III transition (30.2%) suggests an immature biomarker-based patient stratification computational module and adaptive trial design infrastructure. The integration rate of in silico proactive screening in the 276 preclinical projects and the adoption rate of AI-based de novo generative chemistry in the 618 early discovery projects will function as rate-limiting constants in resolving this bottleneck in the future.
Molecular Target Orchestration and Establishment of a Precision Stratification Model for Reversible Pipeline Homeostasis
Reconstructing the omics matrix of 1,095 projects into a four-dimensional tensor of company size, modality, therapeutic area, and development stage reveals the precision stratification landscape of the Indian pipeline. Among the 128 active companies, the top 20 large pharmaceutical companies account for 78.9% of all Phase III projects, while the remaining startup ecosystem—Bugworks (antibacterial), Achilles Therapeutics partner Medigene India (cell therapy)—forms a long-tail distribution concentrated in early discovery and preclinical stages. This asymmetry is both a capital-intensive late-stage clinical bottleneck and, simultaneously, a reversible buffer mechanism in which the Department of Biotechnology (DBT)'s BIRAC funding and the Startup India policy absorb early-stage risk. The 56 prodrug/salt/formulation innovations and 53 repositioning projects represent a strategy of re-orchestrating the pharmacokinetics (PK) parameters—Cmax, AUC, t½—of existing approved molecules through rate-limiting constant up/down-clamping, with Zydus Lifesciences' Saroglitazar (the world's first PPARα/γ dual agonist, an India-approved NCE) being the archetype of this model. Among the 95 gene therapies, the dominant proportion of CAR-T/NK (40) has exponentially increased the computational demand for precision stratification of the HLA polymorphism matrix of Indian T-cell immune phenotypes and the interaction between the tumor microenvironment (TME) immune suppression axis (PD-L1/TIM-3/LAG-3).
Outlook: Establishment of Programmable Drug Governance Standards and Activation of Next-Generation IND Digital Infrastructure
The Indian new drug R&D governance is now at a point where it must declare a complete reset from a static, post-hoc, generic replication system to a fully AI-driven, multi-dimensional tensor-based programmable discovery infrastructure. The 1,095-molecule panorama captured in this review is not merely a catalog but an assertion that the AlphaFold3 structure prediction + Schrödinger FEP+ free energy perturbation + single-cell RNA-seq-based target validation triple computational axis must be seamlessly integrated across all stages of the Indian pipeline. Jubilant Pharmova's AI platform RadiusTM, the Exscientia-India partnership, and the BioNEST incubator network are leading nodes in this transition. When multinational pharmaceutical companies' Indian R&D centers—AstraZeneca Bengaluru, Novartis Hyderabad NIBR—link local genetic polymorphism data as a correction factor in high-throughput screening, the zero-tolerance computational moat for inter-batch drug kinetic deviation will be completed. The CDSCO's (Central Drugs Standard Control Organization) 2023 revised new drug clinical trial rules and ICMR's (Indian Council of Medical Research) biobank guidelines, when converged into a digital companion diagnostics (CDx) standard, will redefine the Indian new drug pipeline as a master asset capable of achieving a disruptive reduction in IND approval timeline—compressed by more than 40% compared to the global average of 12-18 months. India is transitioning from being the world's pharmacy to being the world's discovery engine.
A comprehensive review and analysis of drug discovery efforts at Indian companies between the mid-1990s and 2025 reveals 1095 ongoing (462) or past (633) projects and molecules under investigation at 195 major pharmaceutical, biotechnology, and start-up companies, of which 128 are currently actively pursuing research. They consist mainly of small molecules (811), followed by novel biologics (189) and gene therapies (95), and cover all stages from early discovery (618), preclinical (276), and clinical development phases (Phase 1: 98; Phase 2: 63; Phase 3: 19) up to approved drugs and treatments (21). Small molecules are dominated by new chemical entities (684), followed by prodrugs, salts, and formulations (56), repurposed drugs (53), and others (18). Biologics consist largely of monoclonal antibodies or fragments (92), antibody-drug conjugates (22), various (fusion) proteins (40), enzymes (10), and others (25). Gene therapies use gene silencing (24), gene transfer (21), genome editing (7), modulation of mRNA splicing (3), and genetically modified cells based on chimeric antigen receptor technology using T and NK cells (40). Tracking companies and projects over time illustrates the dynamics and increasing diversification of drug discovery activities in India.
This study's 1,095-molecule panorama mapping transcends theoretical Indian pharmaceutical industry history and directly activates the global finished drug supply chain and the next-generation precision bio-business line.
First, by immediately identifying the rate-limiting step in the Phase II→III transition bottleneck with an AI-based biomarker scanning algorithm, it eliminates the temporal noise of adaptive trial design failures at the source and safeguards computational moats in patient stratification precision.
At the same time, by linking the meta-data of 1,095 projects into open-source ChEMBL, PubChem, and DrugBank databases, it realizes a companion diagnostics (CDx) panel interface that virtually simulates inter-ethnic CYP polymorphism confounding variables and real-time reverse-calculates the effective docking concentration of ADC payloads during clinical trial design.
Furthermore, when multinational corporations conduct large-scale Phase III clinical trials for next-generation CAR-T/NK cell therapies, linking HLA allele frequency and TME immune suppression axis expression levels as correction factors will zero out inter-batch cell potency deviations and maximize the probability of obtaining regulatory approval and cGMP commercial launch permits from global regulatory agencies, functioning as a backbone infrastructure.