Most DNA Storage Codes Throw Away Good Sequences. This One Keeps Nearly All of Them

The case for DNA as an archive has been the same for a decade. It is denser than anything engineered, it keeps for centuries in the cold and the dark, and it draws no power while it sits there. The catch was never the molecule. It is that a great many of the sequences you might want to write into it come back wrong, and the standard way of dealing with that throws out far more of them than it needs to.
A nucleotide has four possible letters, so it carries two bits at most. Some patterns, though, are trouble. A long run of the same letter confuses the machines that read DNA back, and a strand weighted too heavily toward G and C, or too far away from them, behaves badly when it is copied. Existing codes avoid those hazards by ruling out whole families of sequences at once. The rules are blunt: for every error-prone sequence one of them excludes, it excludes many that would have worked. Each of those is capacity the code can never reach, and the shortfall has to be made up with more DNA.
In Nature Communications on Aug. 27, Jiasen Li, Haibing Guan and colleagues at Shanghai Jiao Tong University describe a codec built to stop that waste. They call it Siyuan Code, after the university's own byword, from a motto about drinking water and thinking of its source. The idea at its center is a one-to-one map: every string of bits corresponds to exactly one acceptable DNA sequence, and nearly every acceptable sequence stands for some string of bits. Almost nothing usable is left on the floor.
Getting there means dodging the bad patterns precisely instead of in bulk. The team's answer is a data structure that catalogs the error-prone patterns, so the encoder can steer around each one as it goes rather than banning its whole neighborhood in advance. That lets the code work under strict biochemical constraints while adding very little redundancy, two demands that normally pull against each other.
In the experiments they report, the code reached 1.66 bits per nucleotide and a capacity of 41 exabytes per gram of DNA. The second figure is the one to read carefully, because it already includes the cost of redundancy: it counts the six physical copies of each strand that the files were recovered from, all of them intact. Copy number is where the money is in DNA storage; writing the molecule is the expensive step, and every extra copy is another one to write.
Encoding and decoding ran at roughly one megabyte a second on commodity hardware, which the authors compare to an early USB flash drive. That is the speed of the arithmetic, a file turned into letters and back, and not of the chemistry that would put those letters into a molecule, which is a separate and much slower problem.
The authors describe the code as achieving theoretically optimal density, and the scope of that phrase matters. It is optimal for the sequences their own constraints allow, not for DNA in general: at the full two bits per nucleotide, a gram of single-stranded DNA would come to several hundred exabytes, and physical densities in the same range as this one have already been reported by other groups. The claim is that the biochemical rules a real system has to obey now cost almost nothing beyond what they must.
The online version is an accelerated preview: peer review is finished, but copy-editing is still underway. The paper is open access, with the source data, reporting summary and transparent peer review file posted alongside it, allowing the results to be checked directly.
Whether a code this efficient becomes a working archive depends on things no codec can fix, mostly the cost of writing DNA at all. What a codec decides is how much of that expensive material ends up carrying data rather than being ruled out on suspicion, and on the numbers in this paper, nearly all of it does.
Sources
- Peer-reviewedNature Communications
