Aureka Releases OpenDDE, an Open-Source AI Drug Discovery Engine with 655M Parameters for Antibody-Antigen Co-Folding
核心洞察
Aureka (搜索) has launched OpenDDE (搜索), an open-source all-atom biomolecular foundation model designed as a structural reasoning core for AI-driven drug discovery, released under the Apache-2.0 license.
The model achieved 51.0% success on PXMeter-AB, 70.0% on FoldBench-AB, and 66.4% on the 2026ARK-AB benchmark under top-ranked selection, with oracle selection rates rising to 65.9%, 81.9%, and 80.1%.
OpenDDE (搜索) contains approximately 655 million trainable parameters and required roughly 414,000 GPU-hours for training, signaling that biomolecular AI is entering a scaling regime analogous to large language models.
Aureka (搜索), an AI TechBio company headquartered in Laguna Hills, California and Shanghai, announced on July 6, 2026 the release of the Open Drug Discovery Engine (OpenDDE (搜索)), an open-source, all-atom biomolecular foundation model designed to serve as the structural reasoning core of next-generation drug discovery systems. The release marks a significant step toward making frontier biomolecular modeling accessible to researchers, startups, academic laboratories, and multinational corporations.
OpenDDE (搜索) uses biomolecular co-folding as the entry point to model interactions across proteins, nucleic acids, small-molecule ligands, and other biomolecular components. Rather than treating structure prediction as an isolated endpoint, the platform is designed as a shared structural reasoning layer for sequence–structure–function modeling, enabling complex structure prediction today while laying the foundation for de novo design, affinity estimation, structure-conditioned optimization, and closed-loop discovery workflows.
Three Technical Contributions
Aureka (搜索) highlighted three core technical innovations underpinning OpenDDE (搜索). First, the model introduces atomic latent reasoning over biomolecular tokens, refining representations of local geometry, chemical context, and cross-molecular interfaces before all-atom structure generation. Second, the folding-centered architecture is designed to be extensible, with a unified framework that currently focuses on complex structure prediction but is built to support de novo molecular design, affinity prediction, and other structure-conditioned modules. Third, Aureka has conducted systematic studies of scaling laws and data distillation, examining scaling directions along model-parameter, data, inference, and training axes to identify practical routes for continued improvement in biomolecular foundation models.
Benchmark Performance in Antibody-Antigen Co-Folding
In its technical report, Aureka (搜索) demonstrated OpenDDE (搜索)'s strong antibody-antigen co-folding performance across three benchmarks. Under top-ranked selection, OpenDDE reached 51.0% success on PXMeter-AB, 70.0% on FoldBench-AB, and 66.4% on the newly curated 2026ARK-AB benchmark. Under oracle selection, the corresponding success rates rose to 65.9%, 81.9%, and 80.1%, indicating strong latent sampling capacity and a clear opportunity for further gains through confidence calibration and candidate ranking.
These results carry particular relevance for therapeutic discovery because antibody-antigen interfaces are inherently difficult, flexible, and chemically diverse. Aureka (搜索) reported that OpenDDE (搜索) improves not only low-threshold recovery but also medium- and high-quality DockQ regimes, suggesting stronger modeling of binding geometry rather than merely producing marginally acceptable complexes. Across in silico benchmarks, OpenDDE showed competitive co-folding performance and narrowed the gap with reported IsoDDE-level results.
Entering the Scaling Era for Biomolecular AI
OpenDDE (搜索) contains approximately 655 million trainable parameters and required approximately 414,000 GPU-hours for training—equivalent to roughly 54 years on a single computing unit. This scale reflects what Aureka (搜索) describes as a broader shift in AI for biology: the frontier is no longer only an algorithm problem, but an infrastructure problem requiring compute, data pipelines, engineering, evaluation, and long-running training windows.
Aureka (搜索)'s analysis identified clear scaling trends for biomolecular foundation models, suggesting that larger effective training corpora, larger models, more capable inference-time sampling, and post-training improvements can systematically translate into stronger biological reasoning and structure prediction. The company characterized this as a signal that biomolecular AI is beginning to enter a scaling regime analogous to the one that transformed large language models.
Open Release and TechBio Infrastructure
OpenDDE (搜索) is being released with training code, inference pipelines, checkpoints, and benchmarks under the Apache-2.0 license, with the goal of enabling independent validation, community-driven extension, and global collaboration. The release is available on GitHub, Hugging Face, and a dedicated project website.
Beyond the computational foundation, Aureka (搜索) is pairing OpenDDE (搜索) with a high-throughput automated wet-lab platform to build a dry-wet closed-loop discovery system. By integrating autonomous antibody-design agents with high-throughput single-cell functional screening and automated yeast evolution, Aureka aims to create a high-throughput, high-content experimental data flywheel for functional antibody discovery. This platform allows AI agents to propose candidates, test them through automated experimental workflows, absorb functional and phenotypic feedback, and iteratively improve their design strategies.
Together, Aureka (搜索)'s TechBio infrastructures are designed to support the next generation of antibody discovery across complex modalities such as epitope-specific antibodies, multispecific antibodies, internalizing antibodies, and pH-switch antibodies, with the long-term goal of developing differentiated First-in-Class and Best-in-Class therapeutic pipelines.
Aureka (搜索) noted that downstream capabilities such as molecular design, affinity prediction, conformational ensemble modeling, active learning, experimental feedback, and clinical translation require further research, validation, and development, and the company does not guarantee the discovery, approval, or commercialization of any specific therapeutic candidate.
