First Complete Diploid Human Genome Ushers in New Era of Personalized Medicine
核心洞察
An international team led by Johns Hopkins University (搜索) has assembled the first complete, telomere-to-telomere phased diploid human genome from a single cell line, capturing both parental chromosome sets.
The new genome reveals approximately 15% more genetic information than previous benchmarks, adding over 900 million DNA letters including regions linked to cancer and neurological disorders.
Sequencing costs have dropped from roughly $5 billion in 2003 to approximately $5,000 today, making routine clinical use of complete personal genomes increasingly realistic.
A Johns Hopkins University (搜索)-led international research team has achieved a landmark in genomics: the first complete, telomere-to-telomere phased diploid human genome assembled from a single cell line. Published as part of a 12-paper collection in Cell and Cell Genomics, the work captures both distinct sets of chromosomes inherited from each parent—a feat that had eluded scientists since the Human Genome Project's initial completion in 2000.
The effort was led by Adam Phillippy, a genomicist at Johns Hopkins University (搜索) and the US National Institutes of Health National Human Genome Research Institute, through the Telomere-to-Telomere (T2T) Consortium. Key collaborators included the National Human Genome Research Institute and the National Institute of Standards and Technology (搜索) (NIST).
From a single blended genome to two distinct parental copies
Phillippy's group previously released the CHM13 genome, which was based on an unusual cell type where both chromosomes in a pair are nearly identical. But humans have two distinct sets of chromosomes—one from each genetic parent. Previous assembly methods, Phillippy explained, would "smash everything together," making the individual assembly of each chromosome in a pair the real challenge.
The new genome is based on the HG002 cell line, the standard human genome reference material used by the NIST Genome in a Bottle consortium. Researchers employed long-read PacBio and Oxford Nanopore sequencing, along with Hi-C sequencing, which helps determine where DNA sequences are in proximity to each other.
"We had the first draft of it actually completed around the same time that CHM13 was finished, maybe even as early as kind of the fall of 2022, but we really spent the next number of years trying to make it as perfect and complete as possible so that it could be used as a benchmark," Phillippy said.
Unlocking 900 million previously hidden DNA letters
The completed genome reveals approximately 15% more genetic information than previous benchmarks, adding more than 900 million DNA letters. These newly accessible regions include sequences associated with cancer, neurological disorders, and other diseases. Computational biologist Steven Salzberg and his team at Johns Hopkins identified genes across each chromosome copy, while Michael Schatz's laboratory helped validate the accuracy of the genome assemblies.
"Our complete reconstruction of the genetic makeup of the first reference material genome was the culmination of years of advances in technologies and analysis methods," said NIST scientist and co-senior author Justin Zook. "The achievement gives technology developers the standard they need to measure and improve accuracy across the most complex regions of the human genome."
Toward routine clinical genomics
The breakthrough carries significant implications for precision medicine. Today, genetic testing often fails to identify the cause of rare diseases because portions of a patient's genome remain difficult to analyze—particularly highly repetitive regions such as those containing genes associated with spinal muscular atrophy (搜索). By creating complete personal genomes, researchers believe physicians will be better equipped to diagnose inherited conditions and identify disease-causing genetic variants that previously went undetected.
"The ultimate vision is that, within say 5 years, this is something that we could routinely do in the clinic for a rare-disease patient," Phillippy said. The cost of sequencing for such an assembly is only about $10,000, he added.
The dramatic decline in sequencing costs makes widespread adoption increasingly realistic. The original Human Genome Project, completed in 2003, cost roughly $5 billion in today's dollars. Comparable—and far more complete—genome sequencing can now be performed for approximately $5,000, representing a million-fold reduction in cost.
Beyond rare disease diagnosis, researchers expect complete personal genomes to improve prediction of an individual's risk for cancer, heart disease, immune disorders, and neuropsychiatric conditions. The comprehensive datasets are also expected to strengthen artificial intelligence models that analyze genetic variation and support more precise medical decision-making.
A benchmark and a pipeline for the future
In addition to creating a benchmark reference, the team has developed an analysis pipeline that should make sequencing and assembling an individual's complete genome much easier. "It will soon become commonplace to sequence an individual's entire genome," Phillippy said. "Complete, personalized genomes are now possible for anyone."
The same technology has also been used to generate complete or near-complete genomes for other vertebrate species, including the common marmoset, zebra finch, macaques, horses, donkeys, giraffes, rats, and voles, with companion studies published alongside the human genome research.
Phillippy acknowledged that the new human genome is not 100% perfect—roughly 100 or more suspicious regions remain, including ribosomal DNA containing tandem arrays of nearly identical, several-kilobase-long units. Still, he argued that for clinical use, resolving those final missing pieces is likely unnecessary. "I think we've basically reached perfection when it comes to a single genome," he said.
