Sponsored
πŸš€ Future

DNA as a Storage Medium: 1 Gram Could Hold All of Humanity’s Data

refill of liquid on tubes

In 2017, Yaniv Erlich and Dina Zielinski, working between Columbia University and the New York Genome Center, encoded a computer operating system, a movie, a computer virus, and an Amazon gift card into strands of synthetic DNA, then read every bit back without a single error. The technique had a name that sounded almost cute for something this consequential: DNA Fountain. The number that came out of the paper was the one that stuck. One gram of DNA, storing information at a density of 215 petabytes. For scale, 215 petabytes is roughly what you’d need to hold every photo, video, and document produced by a mid-sized country in a year, sitting in something smaller than a raindrop.

I keep coming back to that number not because it’s a novelty statistic, but because it exposes a strange truth about how we’ve been thinking about data storage for seventy years. We built an entire civilization’s memory out of magnetized rust and etched silicon, materials that were never good at this job. Then it turned out the best storage medium ever discovered was sitting inside every living cell the whole time, doing something else entirely.

The Molecule Was Never Designed to Hold Your Tax Returns

DNA stores information the way a hard drive does, as a sequence of discrete states. A hard drive uses magnetic polarity, up or down, one bit at a time. DNA uses four nucleotide bases, adenine, thymine, guanine, cytosine, which gives you two bits per position instead of one. That doubling sounds modest until you consider the scale at which it happens. A single nucleotide is about two nanometers long. A modern hard drive needs roughly a few thousand atoms to reliably store one bit, with error-correction overhead and physical spacing eating into density every generation. DNA needs a handful of atoms per base, and the bases pack into a double helix that is already the most space-efficient information structure evolution ever produced, because evolution needed to fit three billion base pairs of instructions into a nucleus you can’t see without a microscope.

The molecule wasn’t designed for archives. It just happens to be better at the job than anything we’ve engineered.

George Church’s lab at Harvard demonstrated the concept in 2012, encoding his own book, “Regenesis,” roughly 700 kilobytes of text and images, into DNA synthesized on a chip. Nick Goldman and Ewan Birney at the European Bioinformatics Institute followed in 2013 with an encoding scheme robust enough to survive the errors that synthesis and sequencing introduce, storing Shakespeare’s sonnets, an MP3 of Martin Luther King’s “I Have a Dream” speech, and the original Watson and Crick DNA paper, a nod that felt less like a flourish and more like an acknowledgment of debt. Erlich and Zielinski’s 2017 paper pushed the density argument to its limit by proving that DNA could be written near its theoretical information ceiling, roughly 1.8 bits per nucleotide once you account for the redundancy needed to correct errors.

Ancient Genomes Are the Best Argument for the Archive Case

The density number gets the headlines, but longevity is the more interesting claim, because it’s the one backed by evidence that predates the technology entirely. In 2013, a team led by Eske Willerslev at the University of Copenhagen sequenced DNA recovered from a horse bone frozen in Canadian permafrost for roughly 700,000 years. That is not a demonstration project. That is DNA surviving, in readable condition, for a span of time longer than our species has existed, without anyone designing it to do so. Cold, dry, low-oxygen conditions are enough. No refrigerated data center, no tape rotation schedule, no format migration required every decade to dodge obsolescence.

Compare that to the actual lifespan of the media we use now. A hard drive is optimistically rated for five to ten years of reliable operation. LTO magnetic tape, the backbone of enterprise archival storage today, is rated for roughly 30 years under good conditions, and even then it requires periodic rewriting onto newer tape generations because the drives that read old formats stop being manufactured. I’ve spoken with archivists who describe format migration not as a technical chore but as a permanent tax on memory, a bill that comes due every decade forever. DNA doesn’t ask for that. Under cold, dark, dry storage, credible estimates put its readable lifespan in the thousands of years, and the horse genome suggests that under the right conditions, that number could be off by orders of magnitude in the conservative direction.

Every other storage medium we’ve built needs to be rewritten to survive. DNA just needs to be left alone.

The Bottleneck Is Chemistry, Not Physics

Here is where the enthusiasm needs a hard correction, because the popular version of this story tends to skip the part where it’s still expensive and slow to do any of this. Writing DNA means synthesizing it base by base using phosphoramidite chemistry that was developed for research-scale oligonucleotide production, not industrial data storage. It is precise. It is also slow and costly at the volumes data storage would require, and error-prone enough that every encoding scheme has to build in redundancy the way Goldman’s team did in 2013 and Erlich’s DNA Fountain refined in 2017. Reading the data back requires sequencing, which has gotten dramatically cheaper since the Human Genome Project era but still isn’t instant, and isn’t free.

Random access is the other real problem, and it’s the one that gets glossed over most often. A hard drive can jump to any file in milliseconds. DNA, sitting as a pool of mixed strands in a tube, has no inherent index. You have to use PCR amplification with targeted primers to fish out the specific strand you want, which works, researchers at Microsoft and the University of Washington, including Karin Strauss and Luis Ceze, proved random access retrieval from a DNA pool back in 2016, but it’s nowhere close to the access speed of spinning or solid-state media. Companies including Twist Bioscience have partnered with Microsoft and the University of Washington to push synthesis costs down, and startups like Molecular Assemblies and Ansa Biotechnologies are trying to replace phosphoramidite chemistry with enzymatic synthesis, which promises to be faster and cleaner. None of it has yet reached the point where DNA storage competes with tape on cost per gigabyte for anything you’d actually want to retrieve on a normal timescale.

That’s the honest shape of where things stand. DNA is not going to replace the SSD in your laptop. It is a cold-archive technology, built for data you almost never touch but can never afford to lose: national archives, genomic biobanks, satellite imagery, the digital sediment of a civilization that we suspect, correctly, we will want in five hundred years even if no living person can say why yet.

Storing Data Inside Something Alive Raises a Different Kind of Question

In 2017, Seth Shipman, working with George Church’s lab at Harvard, used CRISPR to encode a short digital movie, five frames of a galloping horse, into the genome of a population of living E. coli bacteria, and then recovered the movie by sequencing the bacterial descendants. The bacteria replicated, carrying the encoded file forward into new generations the same way they carry any other gene. That’s a different proposition than a vial of synthesized, inert DNA sitting in a freezer. It is data that reproduces itself, inherits mutation the way any genome does, and exists inside something that is, by any reasonable definition, alive.

I don’t think that distinction is a technical footnote. Once you can write arbitrary information into a genome that replicates on its own, you’ve built a storage medium that doesn’t need a data center, a power grid, or a maintenance contract to persist. It needs an environment where the organism can survive, and it will keep copying your file indefinitely, with a mutation rate you’ll need to correct for, as a side effect of just staying alive. That is either the most elegant archival strategy ever devised or a governance problem nobody quite knows how to name yet, because it blurs a line we’ve relied on being solid: the line between a record of life and a form of it.

Once information can reproduce on its own, the difference between an archive and an organism gets harder to draw.

The biosecurity conversation around synthetic biology has mostly focused on gain-of-function research and DNA synthesis screening for pathogen sequences, and rightly so. But encoding arbitrary digital payloads into replicating organisms is its own category of risk and its own category of possibility, largely unaddressed by existing biosafety frameworks, which were written for organisms, not for organisms carrying someone else’s Wikipedia dump. Nobody has proposed storing meaningful civilizational archives in live bacterial populations at scale. But the Shipman and Church demonstration proved the mechanism works, and mechanisms that work tend to get used eventually, usually before the governance catches up.

What This Technology Is Actually For

Here’s a comparison that makes the tradeoffs concrete rather than abstract.

Medium Density Realistic Lifespan Access Speed
Hard disk drive Roughly 1 terabit per square inch 5 to 10 years Milliseconds
LTO magnetic tape Tens of terabytes per cartridge Roughly 30 years Seconds to minutes, sequential
Synthesized DNA Up to 215 petabytes per gram Centuries to millennia, cold and dry Hours, via PCR and sequencing
Living-genome storage Bound by genome size and copy number Indefinite, self-replicating Days, via culture and sequencing

Read that table honestly and the conclusion isn’t that DNA wins. It’s that we now have four genuinely different tools for four genuinely different jobs, and for the first time in the history of information storage, one of those tools is a molecule that predates every institution we’ve ever built to protect information, including writing itself.

The people who should be paying the closest attention to this aren’t technologists chasing density records. They’re the ones running the Internet Archive, national libraries, and genomic biobanks who already know that the biggest threat to long-term memory has never been a lack of storage capacity. It’s been the quiet failure of every medium we’ve trusted with permanence, from clay tablets that survived by accident to hard drives that die on a schedule we can now practically forecast. DNA doesn’t solve the cost problem yet, and it may not for another decade. But it is the first storage medium in human history built by a process, evolution, that was optimizing for the one variable we’ve always struggled with most: not how much you can write, but how long it survives after everyone who wrote it is gone.

Credit: Louis Reed on Unsplash

archival storage technologybiotechnologydigital archivingDNA data storageDNA Fountain algorithmmolecular data storagemolecular storageSynthetic Biologysynthetic DNA synthesis
Facebook
Twitter
LinkedIn
Stay charged
The electric pulse of discovery, in your inbox.

One weekly email. The most fascinating stories at the intersection of biology, electricity, and the future. No noise.