Artificial Intelligence in Drug Discovery: A Critical Review of Machine and Deep Learning Applications Across Therapeutic Domains
核心洞察
AI and deep learning have emerged as powerful decision-support tools that accelerate early-stage drug discovery through improved virtual screening, target identification, and candidate prioritization across antibiotics, anticancer agents, antibodies, and small molecules.
Deep learning models including CNNs, GNNs, and generative adversarial networks enable de novo molecular design, protein–ligand binding affinity prediction, and drug repurposing, yet their clinical translation remains constrained by dataset bias and limited external validation.
AI-driven approaches have contributed to the identification of novel antibiotics against multidrug-resistant pathogens and facilitated antibody engineering through epitope prediction and developability assessment, though experimental confirmation remains essential.
The integration of artificial intelligence (AI) into pharmaceutical research has fundamentally reshaped early-stage drug discovery, offering unprecedented capabilities in virtual screening, molecular design, and candidate prioritization. A comprehensive critical narrative review published in Drug Design, Development and Therapy examines how AI methodologies—spanning machine learning (ML) and deep learning (DL)—are being applied across four major therapeutic domains: antibiotics, anticancer agents, antibodies, and small-molecule drugs. The review underscores that while AI has significantly accelerated selected stages of the drug discovery pipeline, its current utility is firmly positioned as an advanced decision-support system rather than an autonomous discovery tool.
The review, which analyzed studies published between 2000 and 2026 from databases including PubMed, Web of Science, and Google Scholar, provides a critical assessment of representative AI applications rather than an exhaustive quantitative analysis. It emphasizes that AI-based approaches enable direct learning from data, allowing for improved capture of complex biological interactions and the ability to handle larger and more heterogeneous datasets compared with conventional computational methods such as classical QSAR models.
AI in Antibiotic Discovery: Addressing Antimicrobial Resistance
Between 2014 and 2019, only 14 new antibiotics were approved—a number considered limited given the increasing burden of antimicrobial resistance, which has been associated with substantial global mortality. This landscape has motivated the exploration of AI-based approaches as supportive tools in antibacterial discovery. Machine learning techniques, including support vector machine (SVM) models and random forest (RF) approaches, have been applied to structure–activity relationship analysis, mechanism-of-action prediction, and potency estimation. In several studies, RF-based approaches evaluated compound activity using minimum inhibitory concentration (MIC) data and bacterial fingerprint information.
A particularly notable development involved a deep learning-guided screening approach that successfully identified a novel structural class of antibiotics specifically targeting the multidrug-resistant pathogen Acinetobacter baumannii (搜索). In another prominent example, a DL framework identified novel antibacterial compounds with reported activity against multidrug-resistant organisms, including Mycobacterium tuberculosis (搜索) and Acinetobacter baumannii, with a proposed mechanism involving disruption of bacterial metabolic processes and iron sequestration. However, the review cautions that these findings were primarily demonstrated in preclinical and in vitro settings, and their broader clinical applicability requires further investigation.
AI in Anticancer Drug Discovery
In oncology, AI-based methods have increasingly been explored to support target identification, genomic data interpretation, and compound prioritization. Graph-based neural network models such as EEG-DTI and DTI-HETA were developed to represent drug–target interactions as heterogeneous graphs and improve interaction prediction using convolutional and attention mechanisms. These models demonstrated improved predictive performance compared with selected baseline models, though results remain dependent on dataset composition and internal validation frameworks.
Network-based methods such as GPSnet utilize patient-specific gene expression data and drug–target interaction networks to identify repositioning candidates. In one reported case, ouabain was identified as a potential candidate in lung adenocarcinoma through network inference and preclinical validation. The SIEVE-Score was developed to improve protein–ligand interaction scoring using ML-based fingerprint representations, while frameworks including MTRD and graph-based models such as DR-HGNN and GDRnet integrate drug–protein–disease interaction networks to predict drug–disease associations.
AI in Antibody Discovery and Engineering
Antibodies remain commonly used therapeutic agents in chronic inflammatory diseases, cancer, and autoimmune disorders. AI and ML-based techniques have been explored to rank clinical candidate antibodies according to predicted immunogenic features and to guide engineering efforts aimed at reducing theoretical T-cell epitope content. In the case of emicizumab, epitope modification strategies were applied to reduce predicted T-cell activation, though the review emphasizes that a reduction in in silico epitope burden does not directly equate to reduced clinical anti-drug antibody (ADA) incidence, as immunogenicity is multifactorial.
Deep learning approaches such as convolutional neural networks (CNNs) were applied during the SARS-CoV-2 pandemic to analyze viral genomic sequences and distinguish variants. Structural and graph-based neural network models, including REFINEGNN, have been applied to antibody structure refinement and CDR conformational modeling. Language model-based approaches such as ProtGPT and AntiBERTy have been utilized for antibody sequence generation, expanding sequence diversity and providing high-throughput design capabilities, though sequence generation quality remains dependent on training data representation.
AI in Small-Molecule Design
Since approximately 2015, AI-assisted approaches have increasingly been incorporated into small-molecule discovery workflows. Examples such as DSP-1181, EXS21546, and DSP-0038 entered early clinical phases following AI-assisted optimization processes, though the review notes these compounds resulted from integrated computational modeling combined with medicinal chemistry and experimental validation rather than being discovered solely by AI.
Platforms such as AtomNet apply deep learning to predict protein–ligand binding probability from structural features. In one study, AtomNet screened millions of compounds and prioritized candidates for synthesis and biological evaluation. Deep learning-based generative models, including recurrent neural networks (RNN), variational autoencoders (VAE), and generative adversarial networks (GAN), have been applied to propose new molecular structures using string-based or graph-based representations. Reinforcement learning approaches including Mol-CycleGAN apply transformation-based learning to modify molecular structures toward predefined objectives.
Persistent Challenges and Limitations
Despite promising integration across therapeutic areas, the review identifies several persistent challenges. Dataset heterogeneity and bias, dependence on high-quality annotated metadata, risk of overfitting, and limited external validation remain significant barriers. The gap between computational prediction and experimental verification is particularly pronounced—while AI models can rapidly generate structurally novel and theoretically active compounds, confirming biological efficacy, synthetic accessibility, and clinical safety remains entirely dependent on conventional experimental workflows.
The review further notes that in generative AI systems, technical issues such as mode collapse may arise, where the generator produces limited structural diversity by repeatedly sampling a narrow subset of outputs. Additionally, the "black box" nature of complex deep learning models raises interpretability concerns, especially when predictions about complex biological interactions must be explained mechanistically.
Future Directions
Looking ahead, the review highlights the potential transition toward autonomous "self-driving" laboratories where AI designs, synthesizes, and tests compounds in closed-loop systems. However, the integration of explainable AI (XAI) and continuous synergy between computational predictions and wet-lab validation will be essential. The review concludes that AI is not expected to replace pharmaceutical scientists but rather to empower them, shifting focus from routine screening to strategic innovation and personalized medicine. Ultimately, the successful progression of AI in pharmaceutical research relies on a combined approach where computational power is continuously guided and validated by human expertise.
