Four molecular letters carry the genetic instructions used by life on Earth. UC San Diego researchers have now shown that a common cellular enzyme can read four additional, synthetic letters in the same strand and produce a corresponding RNA sequence. The finding, published Wednesday in Nature Communications, gives scientists a detailed view of how an eight-letter genetic alphabet can operate during transcription, the step that copies information from DNA into RNA.
The expanded system adds two artificial base pairs, called P:Z and B:S, to the natural pairs formed by adenine, thymine, cytosine and guanine. Researchers sometimes call the resulting set Hachimoji DNA, using the Japanese words for eight and letter. The central question was not simply whether the synthetic bases could sit in a DNA helix. They also had to be recognized with enough speed and selectivity by the molecular machinery that reads a genetic template.
A team led by Dong Wang, a professor at UC San Diego's Skaggs School of Pharmacy and Pharmaceutical Sciences, tested the system with RNA polymerase from E. coli bacteria. RNA polymerase moves along DNA and selects nucleotide building blocks for the RNA copy. Across all 32 combinations of template letters and incoming nucleotides examined in the study, the enzyme efficiently favored the matching partners for both natural and synthetic bases.
Biochemical measurements showed that the P:Z pair was processed at roughly half the rate of the natural G:C pair under the experimental conditions. That difference matters because an expanded alphabet is useful only if copying is both workable and accurate. The team also detected a tendency for the synthetic Z base to pair incorrectly with natural guanine. Replacing one chemical group on Z produced an analogue called Z* and reduced that mismatch in the assays.
To see why the enzyme accepted the unfamiliar letters, the researchers used cryo-electron microscopy, a method that reconstructs molecular structures from many images of frozen samples.
Their structures reached resolutions between 2.42 and 2.75 angstroms. At that scale, the synthetic pairs occupied geometry resembling the familiar Watson-Crick arrangement, while parts of the polymerase moved into positions associated with choosing and adding the next correct RNA nucleotide.
Those images connect the reaction rates to a physical mechanism. They indicate that RNA polymerase does not need an entirely new architecture to handle the extra bases; the synthetic partners fit within an active site already tuned to recognize shape and chemical contacts. The study's forced single-substrate mismatch tests, however, are designed to reveal error pathways and may overstate how often those errors would occur when all competing nucleotides are present together.
An eight-letter alphabet could eventually give synthetic biologists more combinations for storing information or building molecules with functions unavailable to natural DNA and RNA. This experiment does not establish that such systems are ready for use in cells, medicine or manufacturing. It was a controlled biochemical and structural analysis of one transcription enzyme, and the modified chemistry would still have to work reliably through other stages of a biological system.
The immediate advance is narrower and more concrete: the researchers can describe how a natural enzyme distinguishes an expanded set of genetic characters, where it slows and where a mismatch emerges. Z* also provides a testable route to improve one weak point. Future work can ask whether those gains persist in more complete systems, where replication, repair, RNA processing and cellular survival impose constraints that a purified transcription experiment does not.