MOSAIC: Self-Assembling DNA Fragments Could Overcome the Bottleneck in Writing Long, Complex Sequences
核心洞察
Wu et al. introduce MOSAIC (搜索) (Molecular Self-Assembly Induced Cloning), a method that couples programmable self-assembly of overlapping oligos with cellular DNA nick repair machinery to generate long DNA sequences.
MOSAIC (搜索) assembles fragments from 300 bp up to 9.9 kb in a one-pot reaction, with correctly aligned products generated even under extreme homology down to a 1-bp difference.
The approach outperforms polymerase cycling assembly (PCA) by correcting mismatching events through reversible strand displacement, avoiding irreversible mispriming errors.
A new method for writing long, complex DNA sequences could help clear a persistent bottleneck in biotechnology. Writing DNA from scratch remains a limiting step for the field, especially for synthesizing long sequences and libraries, even as reading and editing DNA have been transformed over the past two decades by advances in sequencing, gene editing and computation. This month in Nature Biotechnology, Wu et al. present Molecular Self-Assembly Induced Cloning (MOSAIC (搜索)), an approach that couples programmable self-assembly of oligos with cellular DNA nick repair machinery to generate long DNA sequences.
How MOSAIC Works
In MOSAIC (搜索), overlapping oligos self-assemble in a one-pot reaction in vitro, and DNA nick repair machinery in vivo seals nicks to create functional plasmid constructs, even with difficult-to-assemble sequences. To facilitate recombination, two nicking endonuclease cutting sites are digested at the vector ends, generating short, non-continuous strands. Simultaneously, the gene fragment ends are extended using longer oligonucleotides, which readily generate two overhangs. These overhangs on the assembled fragment can then displace the shorter strands and complement the vector, forming a circular double-stranded DNA. The product is transformed into E. coli (搜索), where the nicks are repaired in vivo.
The assembly design is flexible. In one configuration, offsets are positioned at the 5′ end (with segmentation starting at base +30), resulting in 5′ overhangs. If desired, the offset can alternatively be configured at the 3′ end, which correspondingly shifts the overhangs to the 3′ termini.
Assembly Performance Across Fragment Sizes
The researchers tested assembly of fragments ranging from 300 bp up to 3,000 bp, loading fragments of increasing sizes from 300 bp to 3,000 bp in 300 bp increments. Even the 3,000 bp assembly product displayed a distinct band on the gel, indicating successful assembly across this size range. The maximum assembled fragment size reached 9.9 kb, though it yielded only an indistinct band on the gel.
To determine the upper limit of the recombination method, the team conducted a series of experiments. When directly recombined with the vector, an assembled fragment containing 40 nicks yielded viable clones; given an oligo length of 60 nt, 40 nicks correspond to a 1,200 bp fragment. When T4 ligase was applied to the ligation process, up to a 3 kb fragment yielded viable clones with the correct sequence verified by Sanger sequencing.
The authors note that analysis of the length distribution of human coding sequence (CDS) revealed that the majority of human genes range between 600 bp and 2,000 bp, and choosing 3,000 bp as a cut-off point ensures coverage of more than 88% of all human genes.
Correcting Mismatches Through Reversible Strand Displacement
A key advantage of MOSAIC (搜索) over existing methods lies in how it handles mismatching events. When a large number of oligos for multiple genes are mixed, oligos of high homology are inevitable, and mismatching events can lead to mispriming that becomes irreversible in polymerase cycling assembly (PCA). In MOSAIC, the same mismatching events can be corrected by strand displacement because hybridization reactions are reversible.
The researchers demonstrated this using two sets of oligos designed with a 30-bp homologous region containing a short segment (1-, 3- or 5-bp) of different sequences. In the MOSAIC (搜索) method, only correctly aligned products were properly generated with extreme homology down to a 1-bp difference. In contrast, PCA led to both correctly aligned and misaligned products for all samples with 1-, 3- or 5-bp differences in the homologous region.
MOSAIC (搜索) also addresses kinetic traps that arise from oligos with high GC-content and strong secondary structures. Such mismatching events can block desired base pairing and related enzymatic extension in PCA, whereas in MOSAIC these events become kinetically unfavorable at high temperature, with the desired base pairing becoming the favorable interaction.
Building Plasmids and Combinatorial Libraries
The method extends beyond single fragments to full plasmid assembly. In a two-step strategy, the component oligos required for one full-length plasmid were divided into two pools, each self-assembling to form a 600 bp half-plasmid; equimolar amounts of the two half-plasmids were then mixed to assemble a 1.2 kb full-length mini-plasmid. Assembly efficiency was also demonstrated in yeast for a 3.3 kb plasmid, where clones from selective plates were randomly picked for PCR analysis using five specific primer pairs, and enzymatic digestion and Sanger sequencing confirmed proper assembly.
MOSAIC (搜索) was further applied to combinatorial library construction. In one library targeting selected sites, the theoretical combinatorial library size reached 43 million variants. Next-generation sequencing (NGS) revealed an even distribution of library variants, with approximately 1 million of roughly 11 million sequencing reads containing amino acid sequences belonging to the predicted library, and almost all individual variants appearing fewer than three reads. Screening of a small E. coli (搜索) library (approximately 800 clones) identified newly engineered fluorescent proteins with brightness enhancements of up to 13.6-fold compared with controls, as well as PETase variants with improved enzymatic performance.
Implications for DNA Writing and Therapeutics
MOSAIC (搜索) joins a growing set of approaches aimed at overcoming the limitations of phosphoramidite chemistry, the industry standard that can generate DNA oligo sequences up to 200 nucleotides in length but whose workflow has remained essentially unchanged since the early 1980s. Earlier this year, a group introduced a three-way junction technique called Sidewinder, which contains a third strand carrying information about fragment assembly order. Both methods advance DNA assembly in length and speed, but cannot correct synthesis errors in the original oligo sequences, which are still generated chemically.
Efficient and reliable DNA writing technologies would have major opportunities in the therapeutic space. The best commercial opportunity is probably therapeutic DNA manufacturing — complete replacement of phosphoramidite synthesis for DNA used in CRISPR therapeutics, gene therapy, CAR-T engineering and more. Artificial intelligence is driving protein engineering, but to express a designed protein, a DNA sequence needs to exist; faster and cheaper DNA writing accelerates the experimental loop for AI protein design. Cell therapies are becoming more complicated, involving multiple edits, synthetic promoters, safety switches and large knock-ins, and long, accurate DNA writing expands this design space.
As AI continues to accelerate the pace of experimental science, the need for DNA sequences will grow. Many laboratories can computationally design thousands of promising DNA constructs in hours, but physically building and testing them is the slowest and most expensive step. Reliability will be the key: when a sequence or plasmid is ordered from a commercial vendor, it needs to be correct. While companies report high accuracy rates for longer sequences, they still take too long to generate DNA products, on the order of weeks. The field needs to develop ways to make synthesis faster while also retaining accuracy.
