The human reference genome is still evolving and improving.
The original human genome, when issued to great fanfare in 2000, was an incomplete draft, covering only 92% of the sequence. It caught the most important parts, surely- most of the protein-coding genes- but sequencing technology was not up to the task of doing a complete job. To understand why, imagine the genome as a jigsaw puzzle. The number of pieces corresponds to the sequencing technology- the number of consecutive nucleotides (nt) that can read from the genome at one go. The dominant technology then (and still now) is short read sequencing, which reads about 200 nt at a time, from a randomly sheared up pile of DNA from the genome. Perhaps 300 or 400 if lucky. This means that the 3 billion nucleotide genome is chunked into ~15 million pieces, and that our one-foot jigsaw puzzle has 4,000 pieces per side, each one about 0.08 mm on a side. That is very small. Imagine trying to put together a puzzle like that! Computers are helpful, but they can only do so much.
And like a really hard jigsaw puzzle, our genome has large areas of uniform color- those sky-blue regions and large clouds. Those are the repetitive sequences, which, all together, make up about half of the genome. These are mostly defunct transposons, but also pure repetitive DNA that makes up structural features like centromeres and telomeres. For instance, the sequence GAATGn (where "n" stands for any nucleotide), goes on repetitively for 28 million nucleotides on chromosome 9. That is not an easy puzzle to solve, using tiny sequence reads, no matter how much computer power you have.
| A map of some of the repetitive areas of the genome, focusing on centromeres. Panel A gives a color code for different kinds of repetitive sequence. Panel C gives compositions of the centromere of each chromosome, which are highly variable and different in length as well. Centromeres are underlain by piles of genomic junk, really, with a few key functional sequences. |
So the original project did what it could and released the draft, which was both a huge accomplishment and a large boon to biology ever since. This draft has been patched up several times with improvements. In 2022, finally, a complete genome emerged, which advertised itself as end-to-end, or "T2T", for telomere to telomere. This genome still has a gap on the Y chromosome, which is a degenerate mess. But otherwise, it fills in all gaps, brings in a hundred more protein-coding genes, two thousand other genes, and fixes innumerable other sequencing errors. What got us to this point?
The new technology was long-read sequencing, which is an amazing advance based on threading a single DNA molecule through a tiny pore (a nanopore, which is actually a bacterial protein). You can read the electrical resistance as each base passes through. Each base is ever-so-slightly different electrically, making it possible to figure out the sequence, step-by-step. These methods have been refined to the point that they can read 100,000 nucleotides routinely, and even attain four million nucleotides at a single go. That is revolutionary for closing difficult, repetitive genome gaps.
| Reading sequence, sequentially, as it passes through a tiny pore. |
This long-read nanopore sequencing also benefits from using relatively unprocessed, native molecules (not extensively amplified/copied) helps to reduce errors, though this method has its own error problems, mostly due to the fact that taking electrical readings off the pore is extremely sensitive and tricky. It is also hard to scale up, mostly due to the slow process of threading DNA through those pores, so long-read sequencing has not challenged the huge bulk sequencers for lowest-cost or greatest scale. But when you have to finish a sequence, one that is problematic, with low complexity, long-read sequencing is perfect.
Lastly, what is all that junk DNA doing in our genomes? Mostly, it is honest-to-goodness junk, just carried along out of inertia and lack of selective attention (though there is the story of piRNAs). Transposons can occasionally re-activate and create new mutations, but there are way too many of them to get rid of. They also can gain significant functions, giving rise to new, functional genes and regulatory controls over other genes. So they are a sort of a junk heap / pick-n-pull that can be very useful from time to time. The satellite sequences in the image above are concentrated in structural areas like the centromere, where they do not have classical "gene" type functions, but facilitate equally important processes like genome division and distribution to new cells.
Unfortunately, as we learned from the draft genome, it is not a holy grail full of secrets that will, of themselves, solve our medical mysteries and confer immortality. Its use is as an armature on which we can hang the accumulating knowledge of biology. Each nucleotide in the genome has its own story, of how it got there, what processes it plays in, and how it can affect our health and happiness. The system built upon the sequence is far more involved and interesting than just the sequence.
- More on taxes and inequality.
- Who is Mr. Kremlev?