Sunday, September 27, 2026

An Alternative Vision of China, Past and Future

An appreciation of Chen Jiongming.

China has had a tumultuous and sad history over the last couple of centuries. There was humiliation at the hands of colonial powers, then partial occupation by Japan, followed by Communist revolution, civil war, famine, cultural revolution, and finally its re-adoption of capitalism and rise to immense commercial power, under an autocratic political system. One ray of hope was the first republic, founded in 1912 as the Qing dynasty was overthrown. This was inspired by and initially run by Sun Yat Sen, though he soon bowed out in favor of the warlord Yuan Shikai. Sadly, this republic was just a brief phase in a larger warlord period where competing armies all over China tried to out-maneuver the others. Ultimately, Sun Yat Sen's "nationalist" movement came back to power after his death in 1926 under the leadership of Chiang Kai-shek, who turned out to be little more than another warlord after jettisoning the major principle of democracy from the "Nationalist" platform. 

Well, one leader during this time who paid a bit more attention to the political foundations of China's future was Chen Jiongming. He was part of the great student movement of progressive ideas in the early 1900's, exemplified by the May Fourth movement, the massive student-led revulsion at the sell-outs of China's interests by the Versailles treaty ending World War 1. He was elected to the Guangdong provincial assembly in 1909, and after one abortive attempt at revolution in service to the coming republic, was elected commander of and army based in Hong Kong and aimed to taking over Guangdong province. This he finally did successfully in late 1911. He eventually became the local governor, where he led a brilliant administration devoted to restoring the peace, building the economy, and putting progressive ideals into practice, such as women's political rights. 

After several reverses, particularly after Yuan Shikai clarified his intention to become emperor, Chen was put at the head of another army, which gained control over southern Fujian, next door to Guangdong, in 1918. The next four years would see him take over Guangdong again, and again put in place an efficient and progressive administration. Chen's main principle during this time was federalism, the idea that it would be best to nurture democracy and modernization in single provinces or small areas like Guangdong, rather than devoting all efforts, (as Sun Yat Sen was doing), to take over all of China. The warlord competition was one that seemed to corrupt its participants. When winning is the only goal, other principles get sacrificed, if there were any there to begin with. Few of the warlords had strong political ideals or any kind of progressive basis at all, other than winning. Already Yuan Shikai had shown himself to be afflicted by that virus, and Chiang Kai-shek would take the Nationalists more or less down the same road, after they had indeed conquered all of China and then faced the Japanese and Communists in succession.

Chen Jiongming

Chen was clearly going in a different direction, hoping that successful government and modernization in one area could lead more peacefully and gradually to the similar liberation of other areas. Rather than sacrificing all for power, he wanted to put good government first and stay on the military defensive while tending to an area that is perhaps the premier commercial area of China, which would later lead the way to the PRC's curious blend of capitalism and communism, and whose delta hosts Hong Kong, Macau, and Guangzhou (aka Canton). 

The problem in China, in this and other epochs, is that there are few natural barriers that, in the parlance of biology, foster speciation in geographic isolation. It is too easy to invade other areas, and very difficult to differentiate, grow, and succeed before one's riches are plundered by someone else. The zero-sum logic of warlordism is rampant, and all the people can do is to wait for one of them to win, and hope the winner can be turned to more or less benevolent directions. Chiang Kai-shek and the Nationalists were far from benevolent or competent when their turn came, and their system collapsed in the face of a true peasant revolution (finally, an understandable and broadly appealing political basis, however fleeting) led by Mao Zedong. But that regime turned out to be far, far from benevolent in its turn, and only much later turned its attention to making life better for the Chinese, who are now trapped in an economically growing autocracy, whose leaders are petrified that the people will someday turn on them and recover their natural rights.

It is interesting to note that the Tiananmen square incident happened just as Taiwan was, after a long authoritarian phase of its own, establishing truly functional democratic institutions. Deng Xiaoping, however, crushed the protesters and nascent liberalization in the PRC. Two paths were diverging. Each avoided chaos, but which was the better one?

Perhaps Sun Yat Sen was right to dismiss federalism as a pipe dream in China, given the long history and the political situation at the time. Perhaps it was either eat or be eaten, on this vast and bloody stage. But it is fascinating to consider that an alternate road was in the offing- one that promised far less bloodshed, far less suffering, and a far more civilized outcome.


Saturday, September 19, 2026

Towards a More Perfect Genome

The human reference genome is still evolving and improving. 

The original human genome, when issued to great fanfare in 2000, was an incomplete draft, covering only 92% of the sequence. It caught the most important parts, surely- most of the protein-coding genes- but sequencing technology was not up to the task of doing a complete job. To understand why, imagine the genome as a jigsaw puzzle. The number of pieces corresponds to the sequencing technology- the number of consecutive nucleotides (nt) that can read from the genome at one go. The dominant technology then (and still now) is short read sequencing, which reads about 200 nt at a time, from a randomly sheared up pile of DNA from the genome. Perhaps 300 or 400 if lucky. This means that the 3 billion nucleotide genome is chunked into ~15 million pieces, and that our one-foot jigsaw puzzle has 4,000 pieces per side, each one about 0.08 mm on a side. That is very small. Imagine trying to put together a puzzle like that! Computers are helpful, but they can only do so much.

And like a really hard jigsaw puzzle, our genome has large areas of uniform color- those sky-blue regions and large clouds. Those are the repetitive sequences, which, all together, make up about half of the genome. These are mostly defunct transposons, but also pure repetitive DNA that makes up structural features like centromeres and telomeres. For instance, the sequence GAATGn (where "n" stands for any nucleotide), goes on repetitively for 28 million nucleotides on chromosome 9. That is not an easy puzzle to solve, using tiny sequence reads, no matter how much computer power you have.

A map of some of the repetitive areas of the genome, focusing on centromeres. Panel A gives a color code for different kinds of repetitive sequence. Panel C gives compositions of the centromere of each chromosome, which are highly variable and different in length as well. Centromeres are underlain by piles of genomic junk, really, with a few key functional sequences.

So the original project did what it could and released the draft, which was both a huge accomplishment and a large boon to biology ever since. This draft has been patched up several times with improvements. In 2022, finally, a complete genome emerged, which advertised itself as end-to-end, or "T2T", for telomere to telomere. This genome still has a gap on the Y chromosome, which is a degenerate mess. But otherwise, it fills in all gaps, brings in a hundred more protein-coding genes, two thousand other genes, and fixes innumerable other sequencing errors. What got us to this point?

The new technology was long-read sequencing, which is an amazing advance based on threading a single DNA molecule through a tiny pore (a nanopore, which is actually a bacterial protein). You can read the electrical resistance as each base passes through. Each base is ever-so-slightly different electrically, making it possible to figure out the sequence, step-by-step. These methods have been refined to the point that they can read 100,000 nucleotides routinely, and even attain four million nucleotides at a single go. That is revolutionary for closing difficult, repetitive genome gaps.

Reading sequence, sequentially, as it passes through a tiny pore.

This long-read nanopore sequencing also benefits from using relatively unprocessed, native molecules (not extensively amplified/copied) helps to reduce errors, though this method has its own error problems, mostly due to the fact that taking electrical readings off the pore is extremely sensitive and tricky. It is also hard to scale up, mostly due to the slow process of threading DNA through those pores, so long-read sequencing has not challenged the huge bulk sequencers for lowest-cost or greatest scale. But when you have to finish a sequence, one that is problematic, with low complexity, long-read sequencing is perfect. 

Lastly, what is all that junk DNA doing in our genomes? Mostly, it is honest-to-goodness junk, just carried along out of inertia and lack of selective attention (though there is the story of piRNAs). Transposons can occasionally re-activate and create new mutations, but there are way too many of them to get rid of. They also can gain significant functions, giving rise to new, functional genes and regulatory controls over other genes. So they are a sort of a junk heap / pick-n-pull that can be very useful from time to time. The satellite sequences in the image above are concentrated in structural areas like the centromere, where they do not have classical "gene" type functions, but facilitate equally important processes like genome division and distribution to new cells. 

Unfortunately, as we learned from the draft genome, it is not a holy grail full of secrets that will, of themselves, solve our medical mysteries and confer immortality. Its use is as an armature on which we can hang the accumulating knowledge of biology. Each nucleotide in the genome has its own story, of how it got there, what processes it plays in, and how it can affect our health and happiness. The system built upon the sequence is far more involved and interesting than just the sequence.


Saturday, September 12, 2026

Origins of The Trumpocene: Something in the Air?

I have a theory about our predicament.

How did we get here? With a madman as president, a cabinet of incompetants and sycophants, a supine legislature barely able to keep the lights on, and a dominant block of voters that see nothing wrong with it? This president has, for no discernable reason, decided to wage a trade war against basically the entire world, with special emphasis on our closest friend, Canada. This president launched an idiotic war against Iran that has destroyed our military and diplomatic status, and isn't doing much for our economy either. This president has launched vendettas against anyone with brains, such as colleges, medical researchers, military, diplomatic and intelligence professionals. And this president has made past corruption scandals look like penny ante with shameless shakedowns, pump-and-dumps, and quid-pro-quos reaching into the billions.

It's not good, and stands at the end of a political tradition with far higher accomplishments. We used to expect complete sentences, coherent thoughts, and reasoned policies. All that is now out the window. The civil war era and the revolutionary era were both characterized by excellent, sometimes great writing, lengthy and reasoned debates, and careful argument. Just think of the Federalist papers, the Lincoln-Douglas debates, and other works of Abraham Lincoln. Even into the end of the 20th century, we accomplished great things, and had elevated goals, such as fostering a world-leading research enterprise, and addressing poverty. 

Today, politics revolves around zingers and memes, and we can hardly hold things together while our madman in chief burns down one institution after another. But therein lies a clue- what is in the air? What is making us crazy? As we pass the September 11 anniversary, we recall, inter alia, the devastating effects of pollution around ground zero. Among its many health effects are cognitive impairment and dementia. There is a large literature on the cognitive impairments caused by particulate pollution, particularly in the young. That is a clue.

Long ago, someone had the bright idea to add lead to gasoline. Lead was already known as a neurological poison, but cars were king, and already exquisitely dangerous, so it hardly seemed significant to add an entirely different kind of risk to the fuel mix. Lead was never added to diesel fuel, as diesel is supposed to auto-combust from engine compression. But gasoline is supposed to, instead, be ignited by the spark plug, not compression. An "octane" rating was even devised to make higher-lead gasoline more attractive at the pump. 

Leaded gasoline proceeded to pollute our landscapes, increasingly as car use rose during the 20th century. Urban areas were particularly hard-hit, though agricultural areas can also be highly polluted. There is a well-supported theory connecting this pollution with rises in crime in the US, and later declines as lead was removed from gasoline and gradually the environment. The phaseout of lead was gradual, so if we take its peak in the mid-70's to 80's, then the generation exposed to the most lead in childhood would have been at prime criminal age in the mid-90's. That is one reason why crime has been dropping since that time. But prime criminal age is only one stage of life. Our political system is run by older adults, and the prime voting age is more like fifties and sixties. Thus a peak of criminality in the mid 90's due to the cognitive ravages of lead would correspond to a peak of political influence right about now. 

This is not to say that our president is suffering from lead poisoning. He has other, and far more serious issues with psychopathy and Hitler/Putin worship. (Though he is remarkably touchy about not being very bright.) But the political generation that elected him and forms his base- our political generation- may well be afflicted, explaining to some degree why our political environment is so degraded right now. We are not entirely up to scratch. 

But it isn't just about lead. Particulate pollution is pervasive, from all sorts of fossil fuel burning. We used to think nothing of cooking with natural gas, but it is a major source of indoor pollution. Coal burning spewed vast amounts of mercury, another neurotoxin, among many other pollutants. The bottom line is that the clean energy transition can not come soon enough, if not for our planet, then for our political sanity. 


Saturday, September 5, 2026

Institutions of Truth

Society is built up around institutions, which come in several varieties.

We gather together to do more than we can possibly so alone, into families, tribes, and nations. As human capabilities increased through time, and our populations increased, ever-more diverse and specialized groups assembled to deal with the incredibly complex opportunities at the technological and social frontier. Putting astronauts on the moon required a vast organization with all sorts of skills, but especially technical skills. Such an institution runs on truth. Managers need to get truthful statements from the financial bean counters, the astronauts need to have confidence in the testing and engineering operations, the engineers need to count on the vendors and technicians for quality products and work. 

As a scientist, I often think of my own community as the epitome of truth-seeking and moral righteousness. But all human communities are truthful in their own way, though rarely so explicitly and doggedly. The legal system, for example, is predicated on a search for truth. Granted, it uses an adversarial system that relies on each side bending the truth to its interests, which is an ethic and practice that can easily get out of hand. Even politics is a sort of truth-finding system, not in its speechifying and rhetoric, but in gauging the feelings of an electorate- how it views its problems and who it thinks might solve them, assuming that voting is truly secret and civically engaged. It has been fascinating to see the lies of our current president, so meaningful to his base, functioning as permission to be their worst selves- racist, greedy, incontinent, misogynist, hateful, ignorant.

The news media is an explicitly truth-seeking and publishing enterprise, but can be compromised, both by its market incentives and by political pressures and interests. Markets and business are likewise truth-seeking, not in that they always tell the truth, but that they are predicated on finding a ground truth of what makes money- what is the most efficient and useful thing to be done at this moment. And connected with this is advertising, apparently one of the best ways to make money, and the financial foundation of the news media. Advertising is a curious phenomenon. Its practitioners are paid to tell lies, but only little lies, that skirt the boundaries of fraud. Their art is more typically to insinuate their lies rather than state them outright. That pretty women will flock to the owner of this watch or that car, or that betting at this casino will make you both handsome and rich. 

The arts dispense with explicit truths entirely, and tell human truths though imagined stories. The mark of great art is that it cuts to the heart of the human condition using all its artifice and archetypes. The Matrix is a great example of a wildly creative and frankly insane story that poses the central moral dilemma- can we take the truth, and can we- do we even want to- overcome pleasant illusions to get there?

At the bottom of it all is religion, which offers spiritual purpose and community, founded on pure fantasy. It used to be, in pagan times, that religions were not taken so literally and seriously. Each cult coexisted with the others, and none claimed to be more than an art of the archetypes whereby people could get in touch with a sense of binding social and philosophical purpose. A civic pleasure and ceremonial. But then came Christianity, child of the collision between Judaism and Egyptian and Greek philosophies. Christianity claimed a true historical story, even though it was fabricated more or less out of whole cloth. It claimed absolute truth, certain life after death, and salvation for something... apparently, the human condition was not good enough and had to be redeemed, by blood and death. It certainly was dramatic, but truth? Only on some campy (and masochistic) emotional level.

This institution has had incredibly pervasive effects on our view of truth, and it has taken hundreds of years to extricate ourselves from its presumptions and bizarre claims. Even now there is a rump movement of biblical literalists who are wreaking evil on our politics and society. All in the name of a hot mess of fabrications and primitive morals. We need to take a closer look at the role that truth plays in our day-to-day lives and society, in our moral lives and imaginations. 


Saturday, August 29, 2026

Every Cell Has Its Antenna

Some cilia are for motion, others are for sensing.

The classic image of eukaryotic cilia is of a paramecium covered in its motile fur. These are cilia, and they are composed of an elegant bundle of microtubules, with cross-bridging and dynein / kinesin motors. In our bodies, we have motile cilia in only a few places, like the airway epithelium, the ventricles of the brain, the female reproductive tract, the flagellum of sperm cells, and the ear, both in the middle ear epithelium and the stereocilia of the cochlea. Each of these have very specific roles, generally to move fluids around the cell, or the cell through fluids. 

Classic differential interference contrast image of a paramecium.

Cross sectional structures of a motile cilum (left) and primary, non-motile cilium (right). This 9+0 arrangement further degrades and reduces as one travels outwards to the distal end of primary cilia.

But there is another kind of cilium, called the primary cilium, (or sensory cilium), because virtually all cells have one and only one of them. These are less elegant, and generally immobile. They seem to be a relic of our free-living cellular past, but are hardly without function. There is a class of obscure diseases called ciliopathies that trace their origin to defects in primary cilia and have related and overlapping characteristics. For example, they often feature the mirror reversal of internal organs, which originates from (a lack of) directional fluid flow caused by these cilia (in their only known motile role) in very early embryogenesis. Another common feature is obesity, thought to be due to lack of critical sensing of satiety by the hypothalamus through the primary cilia of its neurons. 

All cilia contain microtubules, and also contain traveling globs called IFT, or intra-flagellar transport complexes. These move up and down the cilium, hitched to motor proteins, carrying the various proteins that make up the cilium, after they were synthesized in the endoplasmic reticulum and sent through the golgi sorting system, into vesicles addressed to the base of the cilium. So, the cilium has its own long-standing growth and internal transport system. 

Cargoes to maintain the cilium are ferried up and down the cilium in rafts.


Some of the specialized proteins of primary cilia are sensory, such as PC1, or polycystin 1, which is a sensor of fluid flow and generates critical signals, such as to mTOR, which is a global regulator of cell energy management and proliferation. The lack of signaling when there are ciliary defects causes over-stimulation of mTOR, which leads to over-proliferation in certain places, such as the kidney, causing polycystic kidney disease, another aspect of many ciliopathies.

Immunofluorescence image of the length-influencing kinase CDKL5 (bottom) and merged with another protein that is known to mark cilia (top). Note how there is only one primary cilium on this mouse fibroblast cell.

The primary cilium is so tiny and sparse that it has been difficult to study physically, despite its functional importance. It is shorter, thinner, and much less organized than motile cilia. But recent papers have shown it in both electron microscopy and fluorescence microscopy.


Electron micrographs of one primary cilium. Note how the well-organized structure at the base (left) gets thinner towards the distal end. The middle panel tracks which microtubules from the base peter out along the cilium length. Below, note the thick raft-like structures of the intraflagellar transport system.

These images show that the microtubule structure is pretty solid at the base, but then progressively degrades along the ciliary length. So, the structure is not much to write home about. It is the cilium's role as a signaling hub, probably of very ancient origin, which is its main importance in most cells. For example, a recent paper delves into CDKL5 deficiency disorder, or CDD. As the authors state, "... patients with CDD variably present with cranial facial and hand anomalies, have significant gastrointestinal dysfunction and sleep disorders along with recurrent pneumonia and respiratory disease. The affected gene is X-linked, and most patients are females. CDKL5 is a protein kinase that is conserved across ciliated organisms including the green alga Chlamydomonas reinhardtii, where CDKL5 localizes to flagella and its loss causes the flagella to be abnormally long." These genetics suggest that males with only one copy of the affected gene are dead, (one X), while females are heterozygous, and still have one good copy. 

The CDKL5 protein is located in both motile and primary cilia, and appears to restrict their length/growth, as its most obvious phenotype. It is concentrated at the base, and may restrict the activities of the intraflagellar transport system, while being transported by it as well. It is a protein kinase that likely regulates some aspect of the IFT by phosphorylating one or more IFT proteins. The authors find that CDKL5 is inhibited by another protein kinase, CDK20, which is tied to the cell cycle and cell proliferation. CDK20 is an oncogene upregulated in many cancers, thus making one more connection between cilia and the wider mechanisms of cellular growth control.

So, a tiny organelle common to all our cells is a bit like the radio antenna on our cars.. a bit antiquated and curious to look at, but still critical for some functions. It is hard to believe that its functions could not be re-engineered to happen on the plasma membrane, which has, if needed, other localization mechanisms like lipid rafts and caveolae, but design is not part of the equation here, rather, making do with and elaborating on whatever has worked in the past is always the way forward.


  • Whence the debt?
  • The organization of visual processing and abstraction may be vectorial, but not be sequential.
  • Some problems with beef.
  • What is happening to our institutions of truth?
  • Bill Mitchell on why bond sales by our government are unnecessary.

Saturday, August 22, 2026

Curious Computation

A biologist and former employee of IBM visits the Computer History Museum, in Mountain View, California.

Walking into the Computer History Museum, it seems unassuming enough- a slightly clean aesthetic, some open space, an early Waymo. But it is one of the most intense religious shrines I have been to. If Silicon Valley is a religion, one of industry, hope, and salvation, this is its interpretive and devotional reliquary. It honors the efforts of countless acolytes who have put their hearts into the technical organization of knowledge, from punch cards to SQL to data centers. And it honors its prophets, above all Thomas Watson and Steve Jobs. While the Silicon Valley religion has fallen on tumultuous times now that its powers have grown so intrusive, even reviled, and fallen into decidedly lesser hands, its eschatology has never been closer, as its greatest monastic centers are working towards "general" artificial intelligence, humanoid robots, and (ironically) the end of work. 

In the last century, IBM (with Thomas Watson as its benevolent dictator) was the leader of creating and handling mechanized information. Punch cards were a wonder of their day, allowing for efficient storage and flexible tabulation of huge amounts of data. Businesses quickly became dependent on them for modeling (and forecasting) their inventories, employees, payrolls, and financial streams. There was a tinge of omniscience and prophecy in the powers these machines could bring. But IBM was slow to get into electronics, slow to get into relational databases, and slow to get into small/personal computers. Particularly heartbreaking to me was the exhibit on SQL. Researchers at IBM invented the relational database and its SQL language. But IBM didn't see its advantage over its existing, less flexible data search/organization methods, and let Larry Ellison steal the idea and set up Oracle corporation, now as large as IBM itself.

The IBM system 360 was a revolution in style as well as computation.

The museum has a special room devoted to Apple and Steve Jobs, the quintessential prophet of Silicon Valley and its positive message of hope through technology. He has died and gone to heaven, but who has not stared into his handiwork? Who can say that he is not living still among us? Seeing the small beginnings and flickering flame of Apple's early challenges, before it became what one may now say is the official religion of technologically advanced society, well ... it was stirring, to say the least. 


Is that a halo?

The collection of the museum is very rich, dealing as it does with the relatively recent past (other than the abacus collection!). Indeed, I have a minor version of it in my basement. It is remarkable how diverse the flow of ideas was in the early computer age, with computers made in Japan, Germany, along analog as well as digital lines, by countless designs. There was a ferment that is characteristic of nascent technological fields (and, not incidentally, of biological radiations and innovations) where new possibilities are interpreted in various ways, and different prophets try out contrasting visions of the future, some regimented (like IBM), some more free flowing (like the home-brew computer club).

But the scope of the museum is also a bit confining, focused as it is on the human-built apparatuses of data processing. Information and computation is a larger subject, and as a biologist, I have to say that the finest computers at the Computer History Museum were walking around, not in the display cases. We are computers, and darn good ones. The fact that AI companies with their Sauronic datacenters can not accomplish what one person can with one brain, powered by little more than a burrito, means that some fundamental aspect of intelligence is missing in what is so blithely called AI. Computers have far surpassed what humans can do in data storage, and in mathematical computation. But in other areas our intelligence remains untouched. Technology will undoubtedly get there, but I don't think it will get there in the current dispensation, just by "scaling". 

So, for the moment, biology still wins. Indeed, information is so fundamental that biology has from the start created information systems. The simplest bacterium senses its environment and can tell whether there is food, or danger. The Cambrian explosion featured eyes- one of the most powerful weapons in the battle of life. Humans are naked in part because of our advanced sensing and cognitive abilities. We don't need shells, scales, or even fur. We can see far away, into the future, and around corners. We can organize flexibly to defeat any opponent on the planet, to the point that we now pine for those monsters we have lost, and are still losing. Information is power.

Not only that, but biology is doubly about information, since life has not only persistently created computational machinery, it is a computational machine, with that four-bit code at its heart. While software developers have a cycle of [ write - test - fix ], biology uses reproduction to get around the fixing part. Good programs are replicated/manifested, bad ones die. The cycle is [ reproduce - test ]. That is it, with ineluctable variation thrown into the "reproduce" part of the process. Simplicity itself, but profoundly transformative to the Earth, not to mention creating the greatest computers currently in existence. Of course, some physicists will say the universe is a big computer as well. It is computers all the way down!

The Silicon Valley religion is, as all religions are, about ourselves, really, and the image of ourselves that we see and strive for. That is why the quest for artificial intelligence is so haunting. Originally, that vision was progressive- one that achieves the deep yearning of the defeat of all enemies, of poverty, of conflict, of limitation and bounded-ness. That achieves freedom. Does Amazon help us achieve freedom? Does Netflix? Yes, a little, and so the works and wonders keep coming, even though they may come out of decidedly unfree channels. And some of the leaders have, sadly, gone in a different direction, preaching freedom as libertarianism, coarseness, even racism.

So there is a dark side to this religion, as the powers of all this technology conjure up greed and conflict, surveillance and control. IBM's biggest customer was the government, after all, and it happily sold to Nazi Germany. Religions, especially ones that promise power, can turn bad, and they have to consitute an ongoing conversation that comes back to people, not technology, to determine what the future should look like. The Computer History Museum houses an immensely positive story, about people who helped us overcome age-old limitations and achieve incredible dreams. But look in the mirror- we are the real subjects and agents of this story.


Saturday, August 15, 2026

Sensing Unfolded Proteins

Do cells take joy in tidying? Absolutely! But how can they tell where the messes are?

Any household knows that without cleaning and disposal, nothing else is possible. Things pile up, messes accumulate, everything comes to a standstill. Cells are the same, needing a constant flow of new materials in, trash out, recycling, and cleanups. One of the major kinds of mess in cells is unfolded proteins, which when they accumulate can cause aggregations and diseases like Huntington's, other dementias, atherosclerosis, and inflammation generally. While there are a lot of mechanisms (chaperones, chaperonin cage complexes, proteasomes) dedicated to proper folding, sometimes they are not enough, and excess junk in the form of unfolded proteins builds up. One major response this is the unfolded protein response, or UPR, which is centered at the endoplasmic reticulum (ER).

Why the ER? This is where countless ribosomes dock to feed in nascent proteins destined for secretion, the plasma membrane, and many locations inside the cell other than the cytoplasm. So it is a key organelle for protein synthesis and distribution, and thus for sensing how cellular proteins are doing. This sensing, the UPR, starts with a sensor protein, called IRE1 (also called ERN1), which sits on the inside face of the endoplasmic membrane. IRE1 binds to unfolded proteins inside the ER, and when the concentration rises high enough, it gathers into condensates of its own that trigger a variety of responses that include, in the extreme, cell death. Yes, lack of timely tidying can have serious consequences!

IRE1 is multifunctional, containing a kinase that activates AMPK and JNK, proteins which influence metabolism and inflammation. It also contains an RNAse enzymatic activity, which activates a whole other program by altering the mRNA of XBP1, which then becomes active and produces a transcription regulator that starts building more ER, more chaperones, and more protein folding capacity generally. The RNAse also digests other mRNAs indiscriminately, lightening the load of translation and thus incoming nascent proteins. So IRE1 is a key homeostatic regulator that adjusts the size of the ER to cellular needs, and informs the rest of the cell how things are going in the ER. 

One structure of IRE1, luminal domain only. This is a dimer, with halves shown symmetrically across the middle. How does this bind unfolded proteins?

But how does IRE1 sense unfolded proteins? A recent paper went through a modeling study to show how IRE1 operates. Sitting in the ER membrane, the enzymatic, signaling parts of the protein are outside, while the sensor is inside, called the core luminal domain (cLD). The authors demonstrate that this domain binds unfolded proteins directly, using them as bridging elements to bring multiple IRE1 proteins together and form the glob that nudges aside the inhibitory protein BiP and promotes the phosphorylation activity and thus activation, of IRE1. BiP also binds to unfolded proteins itself, so forms another sensor that indirectly activates IRE1 under stress conditions. 

Firstly, IRE1 is a dimer, and the luminal domain is where it dimerizes. The structure (just of the luminal domain, above) shows a beta ribbon core and peripheral unstructured areas and alpha helices. After running the crystal structure (shown) though very brief molecular dynamics simulations, the unstructured areas flop around dramatically, while the core and dimeric structure remain stable. Intriguingly, the beta stranded center of the dimer resembles (between the blue helices, above) the peptide binding cleft of MHC molecules- those that present random antigens to T-cell receptors of the immune system. Yet the authors show that unfolded proteins bind differently to IRE1. They flop across the whole central region, attracted by negative charge as well as the open hydrophobic surface. The next image shows molecular dynamics before/after shots of various unfolded peptides known to bind IRE1 in vitro, showing how after starting straight across the central region, they each find different comfort zones, mostly over this central region, but some off to a side. 

Collection of IRE1 structures (gray) with unstructured peptides of various kinds bound to the luminal face. t=0 is the starting point for the molecular dynamics simulation, which went for one microsecond. End points are below. In each case, the unfolded peptide stayed bound, and settled into some comfortable position, based on the hydrophobic and negatively charged landscape of the IRE1 face. 

The intriguing thing about this binding is that the structure of IRE1 itself is unaffected. The dimer interface remains stable and open, leading to the hypothesis that these peptides form bridges by becoming the meat of a two-dimer IRE1 sandwich. And it is this that then displaces the inhibitory protein BiP and leads to condensation and activation, by bringing together multiple IRE1 kinase domains (on the outside of the ER), and activating their kinase/signaling functions, turning on the various downstream pathways.

A helpful model from another lab, describing how, once the unfolded proteins are bound inside the ER (by mechanisms unknown at the time), the outside domains would gather, trans-phosphorylate, and activate downstream pathways. 


  • Those flock cameras.
  • As if charging for previews of presidential posts wasn't corrupt enough.

Saturday, August 8, 2026

Our Loopy Way of Learning

Different ways of learning go through different parts of the brain.

Humans are champion learners. Other animals are smart, but we are smarter, spending more time in childhood soaking up the mysteries of the world around us, and storing them in a larger and better brain. Learning is our calling card, allowing us to adapt to any environment and defeat any foe. Indeed, we have overwhelmed the earth's biosphere, and need to exercise some deeper and longer vision by pulling back from our successes in mastering every possible resource.

But how the brain does all this is something we have yet to fully learn about. Our sense organs bring in tons of information, but as we have learned in computer science and AI, it takes a lot more to create usable information than just amassing data. How do the various parts of our brains work together to create actionable models of the world out of that data? An innovative theory was elaborated a quarter century ago by Kenji Doya, who has apparently gone on to a research sideline on soccer(!)

The problem began with, in part- what does the cerebellum do? From lesion data, it has long been clear that this small part of our brain has a big role in motor accuracy and coordination. But cerebellar neurons project all over the brain, not just to motor areas, and it gradually became clear that the cerebellum affects those other areas as well, helping us learn and manage many cognitive tasks. But structurally, the cerebellum is a peculiar organ, highly repetitive, with relatively linear and parallel processing, a bit like a GPU. It is clearly specialized for something- some kind of fine tuning, but what?

Doya's theory about the distinct learning styles contributed by different parts of the brain.

Doya proposed a general theory about learning in the brain, which splits learning up into three types. One type is supervised learning, which is what the cerebellum facilitates. We grab something, we recognize an error in eye-hand coordination, and we learn to grasp better. There is a loop involved, both physically and logically, by which some task is evaluated, and error signal sent back to the source, and a small adjustment is made for the next iteration. Humans have great hand coordination, partly thanks to our advanced cerebellar-cortical connections. 

The second form of learning is reward-based learning. This prototypically features dopamine, and is centered in the basal ganglia, where emotions are processed, but which is also a major switchboard for connections upwards to the cortex and downwards to the brainstem, and is also critical for motion control. Parkinson's and Huntington's diseases both affect this region. Doya proposed in broad terms that this area doesn't do the kind of error correction that the cerebellum does, but rather learns by reward- whether the action's goal was reached. If the goal was reached, a spurt of dopamine is issued, which reinforces whatever connections helped that action happen.

Interestingly, there is a progressive temporal aspect to this learning. As an action is learned, the reward happens sooner, given the predictive capability of our learning / cognitive systems. The whole point, after all, is to figure out as early as possible what will be happening in the future. If a bell is associated with food, at first the food prompts the reward. But as learning happens, and the bell is reliably associated with the later food, and the bell starts to set off the reward all by itself. Gradually, the sight of the person getting about to ring the bell, or the expected time of day, sets off the reward... it is cast progressively back to whatever the perceived cause is, going back in the learned causal chain. Eventually, as everything becomes rote and the novelty wears off, reward lessens, and learning is complete. Mealtime is here, as it is every day... boring.

Lastly, the third type of learning is unsupervised learning. This is the job of the cortex, and is the most interesting and subtle form of learning. The world has patterns, and the cortex's job is to figure out what those patterns are. No one outside tells it what is right or wrong, what is or is not real. It has to figure all this out from the stream of input, including its own blundering actions and experiments. Unsupervised learning is thought to be well suited to the basic mechanism of the brain, its Hebbian circuitry where activity reinforces any active connection. There are a lot of mathematical methods that have been devised to approximate such processes, like principal component analysis. Data can be binned / categorized in various ways to simplify and organize reality, and eventually we have ... language and concepts, labels and abstract understanding. This categorization of the world is immensely powerful, if always inaccurate and approximate. 

Doya's model of the cerebellum. Purkinje cells at the center are also conceptually at the center, as the only cells that are plastic in this scheme, able to change their connection weights due to input from the other layers, which in turn get input from other brain areas including the cortex and thalamus. This has been likened to a perceptron in computer science terms.

Doya thus created an influential heuristic that divided up learning styles / mechanisms by anatomy in the brain. Obviously, all these mechanisms work closely with each other. The cerebellum doesn't know anything without conceptual categorizations coming from the cortex, which provide the foundation of error signals. Similarly, whether goals have or haven't been reached has to be informed by conceptual data from the cortex (or food signals, or other sensations from elsewhere in the body). Learning is a whole-brain activity, and while it is prodigious and effortless early in life, it continues to be essential throughout, lest we be overwhelmed by the fresh challenges the world keeps presenting to us.


Saturday, August 1, 2026

Ohnologs and Paralogs: The Wages of Gene Duplication

Two whole genome duplications lie at the root of vertebrate evolution.

Another week, another story about the power of gene duplication in evolution. While rare during normal reproduction, gene duplication happens pretty frequently over longer time scales. How else would we get a thousand olfactory receptor genes, all similar to each other? But other accidents can occur as well, like whole chromosome duplications (such as what leads to Down syndrome), and whole genome duplication, when the cell division process stops early, but otherwise proceeds with cells remaining viable with double the genomes as before. This is common in plants. Corn is tetraploid, wheat is hexaploid, and strawberries are octaploid. 

But among animals, whole genome duplication is less common. Two decades ago, however, two researchers working from the newly sequenced human genome revealed that there were two such duplication events at the beginning of vertebrate evolution, explaining some oddities and also perhaps the speed and power of subsequent evolution. The basic evidence is the genome sequence, which is full of related genes. Genes that do the same thing in various species, and are lineally related, are called orthologs. That is relatively simple, per the Darwinian tree of descent of all life. Genes within one organism / one genome that are similar to each other due to ancient duplication events are called paralogs. All those olfactory receptors are paralogs, for instance. Lastly, genes that are paralogs stemming from a whole-genome duplication event are called ohnologs, in honor of Susumu Ohno, who led the field of molecular evolution in recognizing the importance of gene duplication, and speculated about whole genome duplication well before it was discovered.

A classic example of this evidence is the hox cluster, a linear sequence of genes that have been extensively studied in flies as providing an important set of regulatory controls over the linear body plan. They lie in the middle of the developmental cascade, downstream of egg and body polarity genes, but upstream of specific appendage and tissue expression programs. They encode DNA-binding (homeobox) transcription regulators, and their position in the genome is co-linear with the body parts they activate because there is a progressive chromatin opening process by which this whole locus becomes activated. Well, flies have one hox locus, but vertebrates have four. What happened?

Hox loci across evolution. Where flies and primitive chordates have one hox cluster, vertebrates have four. How did that happen?

Obviously, once researchers lined everything up, it became pretty clear that there were two massive duplications along the way, creating four hox loci in the genome, after which quite a few of the duplicated genes fell away. After a duplication event, gene survival is a race between neo-functionalization (which leads to preservation by selection) and deleterious mutation, degradation, and disposal. Enough of these ohnologs survived to help fuel the substantially greater complexity of the vertebrate body plan, now including intricate wrists and hands, and ever more involved head structures.

Similar findings were made all over the human genome. The original paper has a graph that shows that, across the genome, most paralogs exist in families of four. If genome duplication were not the applicable hypothesis, then one would expect a smooth asymptotic curve downwards from one member (implicit) to two, then three, and fewer from there outwards, since single gene duplications would each be independent statistical events. But no, there is a peak at four, indicating that some process yoked together many, many genes into parallel duplication events, twice in succession. 

A graph of count of paralogs vs their frequency by count, in the human genome. What should in principle have been a smoothly declining curve from 1 or 2 turned out to have a weird peak at the number four. This was important evidence that many of our paralogs arose through some common event, such as a pair of ancient whole genome duplications. 

OK, so far, so old hat. A bunch of more recent papers flesh out this story a bit, showing how the vertebrate duplications were timed, and how they affected various types of genes. miRNAs, for instance, turn out to be ancient genetic elements, and were duplicated along with everything else. miRNAs that originate from (and were preserved from) the whole genome duplications tend to be more conserved, have more targets than other miRNAs, are more highly expressed, and have higher rates of targeting RNAs of transcription regulatory proteins that likewise originated from whole-genome duplication. The traces of this history are thus interestingly preserved.

Another paper traced the effects of the vertebrate ohnologs on brain development. Using the pre-vertebrate amphioxus as an out-group for comparison, they find that the ohnologs that date from the whole duplications are more highly associated with brain cell types and their developmental programs than are other gene duplications- either ones dating from the same time as the whole genome duplications, or since.

"Compared to their closest invertebrate relatives—tunicates and amphioxus—vertebrate brains are highly regionalized and complex."

"Ohnologues were enriched in development, cell-fate commitment, signaling and neurotransmitter transport. By contrast, SSD [small-scale duplication] paralogues were enriched for immune response and sensory perception in all species, a result that matched previous reports."

Lastly, a paper focusing on intermediate genomes, part of a plethora of genome sequencing that has happened since the original analysis, reveals exactly what happened at this time, about 500 million years ago. From a molecular perspective, hagfish are in the same group as lampreys- both parasitic eels that lack real jaws but arise from the vertebrate lineage. They are cyclostomes, while we are gnathostomes. After the first whole genome duplication about 530 million years ago in the common stem lineage, the gnathostomes and the cyclostomes split and each experienced their own, separate genome duplications roughly 490 million years ago. This can be concluded from the differing gene collections that survived from each respective duplication, and their sequences relative to the cyclostome/gnathostome split. Indeed, lampreys and hagfish have six hox clusters instead of four, indicating that a triplication event happened, instead of duplication event, between their split from gnathostomes and the split between lamprey and hagfish. 

A phylogenetic tree with dates, locating the genome duplications deep in the vertebrate stem lineages. 1R is the first genome duplications, common to all vertebrates. 2R is the duplication in the lineage leading to jawed vertebrates, while CR is the apparent triplication that happened in the lineage leading to cyclostomes. Apologies for the antiquated human being. 

All this genome renovation didn't benefit cyclostomes the way it did gnathostomes, however. Only one group went on to globe-straddling glory, indicating that while genome duplications may be helpful for evolutionary / developmental innovation, they certainly aren't determinative. They are grist for an evolutionary process that remains largely shrouded in mystery- the context and environments of the time, the competitive biosphere, and the internal molecular environment. There is no reason to think that all this is unaccountable by natural processes, but that doesn't mean we have the information to reconstruct it in detail. We should thus be deeply grateful that the scientific community can dig up even this much of our history.


  • But how can we take more money away from workers?
  • The smearing of Anthony Fauci. And of basic reason.
  • Reward and addiction.
  • Still stupid ... the Discovery Institute and "information".

Saturday, July 25, 2026

Revealing the Origin of Eukaryotes

Tracing the breadcrumbs of deep phylogeny through gene duplications. 

Last week, we discussed the jamboree of mutation that is the MHC locus in animals. One part of that story was the persistent duplication and pseudogene-i-zation (which is to say, the birth and death) of genes in that locus through primate evolution, another consequence of the never-ending arms race with our pathogens. Gene duplication has much deeper significance, however, as it has given rise to new genes throughout biological evolution, as it is merely an accident of DNA (or RNA) replication, thus operationally rather frequent. Significant transitions, such as the advent of eukaryotes, and the advent of plants and of angiosperms, all feature rapid gene duplication, even whole-genome duplication, which provide grist for new functions, refinement of old functions, and speciation.

A recent paper did a deep dive into the origin of eukaryotes, perhaps the watershed event in the history of life, second only to the origin of life itself on the early Earth. It is very difficult to reconstruct what happened during these extremely early events, which these authors put as long as three billion years ago. There has been so much mutational churn in our molecules, and no fossil record to speak of, that, again like the origin of life itself, we do not have a great deal to go on. Thankfully, the discovery of the kingdom of archaea, and its especially eukaryotic-like class, the Asgard archaea, has provided a much better platform for speculating about what the originating organisms might have been like.

Eukaryotes differ from bacteria and archaea by numerous characteristics, implying extensive development/evolution before the last common ancestor from which we can trace all the extant descendants. These include the eponymous nucleus, an internal system of membranes and membrane-bound organelles, a complex cytoskeleton of both actin and microtubules, a mitochondrion, meiosis, a microtubule spindle-driven division system, and many other, more obscure molecular differences. The lack of intermediate forms is certainly frustrating from a scientific perspective, making it difficult to go through the catalog of life today to identify the steps involved in this transition. It also implies that the process ended with an evolutionary bang that in essence caused the final iteration to wipe out all the preceding forms, a bit like humans vs the many other Australopithecines and their descendants. 

These differences are enormous and momentous, making eukaryotic cells far larger, energetically capable, and complex than their antecedents. Another key characteristic is a prevalence of gene duplication and specialization, leading to much larger genomes along with all the other complexity. It has long been theorized that it was the union of the proto-mitochondrion, which was a bacterium, with the archaeal founding cell, that set the whole process off, particularly providing the massive amount of energy required for all this rise in complexity. 

However, the current authors claim that this definitively is not the case- that the mitochondrial symbiosis was a rather late event, coincident with the oxygenation of the atmosphere about two billion years ago. They argue that, since gene duplication and specialization is in any case a clear characteristic of eukaryotes, using molecular clocks to time these many gene duplications can give us relatively specific and data-rich insight into the epochs at which each of the processes that they participate in first developed. 

The authors present estimated times of origin for key duplications involved in RNA synthesis. The RNA polymerases I, II, and III were duplicated from one polymerase in archaea, and are composed of numerous gene products, as well as ancillary regulatory proteins. While some subunits are still shared, the ones that duplicated did so in the range of 2.3 to 2.8 billion years ago, before the symbiosis with the mitochondrial precursor, timed at "mFECA", or the mitochondrial first eukaryotic common ancestor, about 2.2 billion years ago. "nFECA" is nuclear first eukaryotic common ancestor, while "LECA" is last eukaryotic common ancestor. The axis at the top is time before present, in giga-years ago, or Ga.

For example, eukaryotes all have three RNA polymerases, while bacteria and archaea only have one. These RNA polymerases specialize on mRNA synthesis (RNA pol II), ribosomal RNA synthesis (RNA pol I), and tRNA, 5S RNA and U6 snRNA (RNA pol III). Why this specialization exists remains a bit mysterious, though once it started there was no going back. In any case, it happened and is totally characteristic of eukaryotes, so these duplications (which involve the duplication of numerous genes encoding proteins of the polymerase complex itself and its regulators and way-finding helpers) can be used to date eukaryogenesis, using a very carefully calibrated and species-sampled molecular clock. And what they find is that all of these duplications, given as ranges in the graphs above, with green dots at the likeliest time points, all lie prior to the point of "mFECA", which is the mitochondrial first eukaryotic common ancestor, at roughly 2.2 billion years ago. Tabulated over studies of almost a hundred other genes that are similarly duplicated and characteristic of eukaryotes, the authors place the "nFECA", or nuclear eukaryotic common ancestor way back at almost 3 billion years ago. 

That is a very long time! It is a very long time ago that the characteristics of eukaryotes started to come together, and it is also a very long time (well over a billion years) over which that development spanned before culminating in the last common ancestor of all extant eukaryotes, which the authors put at about 1.7 billion years ago. And note that that still comes a billion years before the Cambrian explosion of animal life. We are talking about really deep time here. 

If mitochondrial symbiosis came later, (causing another bout of genomic change as hundreds of bacterial genes moved from the symbiont to the nucleus), just as the atmosphere was getting oxygenated, what was going on in the preceding billion years? One can speculate that these evolving proto-eukaryotes were the top predators of their world, a bit like eukaryotic protists in microbial environments today. The competition would as always have been intense, but there were no higher life forms (i.e. animals or protists) to worry about. Size was not a problem, but rather an advantage. The focus was on quality and effectiveness in finding, disabling, and digesting prey. Thus there was little downside to complexity, unlike the constraints on prey species, which had to optimize for rapid growth and reproduction under both nutrient and predation constraints. New and costly systems for protein secretion, for prey digestion, and intelligent, responsive regulation of all cell processes might have been consistently advantageous.

I take this work as being definitive, at least in terms of settling the debate about early vs late symbiosis. It benefits from a clear theoretical basis and new, voluminous sequence data from appropriately related species in its molecular comparisons. I recommend it highly, as I can only scratch the surface of its analysis here. The advent of mitochondria was then a later step change in complexity, but brought incredible metabolic productivity that, at least in oxygen-rich environments, would have definitively superseded all the prior forms of the proto-eukaryote, and (unfortunately for us) erased that record of evolution. A later innovation, placed between mFECA and the last eukaryotic common ancestor "LECA" was sexual meiosis, another fascinating and characteristic innovation of eukaryotes. These authors thus provide a true revolution in our understanding of eukaryogenesis, extending it across well over a billion years.

One more sample from the author's presentation of molecular duplications that help time eukaryogenesis. Here, DNA repair proteins are emphasized, including some involved in meiosis (light blue dots, see legend). Some of these duplications involve genes from the mitochondrial endosymbiont, (MSH4), and some (such as EME1/2 and MUS81) come from older repair processes, but later specialized for meiosis. 


Saturday, July 18, 2026

MHC Through Evolution: Breaking All the Rules

The immunologically critical MHC gene cluster plays by its own rules through the arms race of life.

Last week, I discussed the very general landscape of variation in the human genome, specifically the tradeoff between prevalence in the population and effect size. Given that the vast majority of variants are deleterious, those that affect our phenotypic traits are more heavily selected against the greater effect they have. The result is that at a gross level, most genes and most traits share the same general distribution of lots of variants (alleles) with minor effects, and far fewer with large effects. And those with minor effects also turn out to be tangential for biologists, rarely informative about the nature of the traits they (sort-of) affect.

This week, another paper and another view of evolution, though the eyes of one the more critical genes of the immune system, the major histocompatibility complex, or MHC group of genes. While our adaptive immune system has developed the extraordinary and powerful ability (though semi-controlled DNA recombination of the antibody and TCR genes) to recognize practically any antigen, foreign or domestic, that system requires stringent controls. One of those controls on T cells, which carry the antigen-recognizing TCR receptor, is that it can only see antigens that are "presented" on MHC molecules. MHC proteins have a surface cleft that gets loaded with and holds small peptides (8 to 12 amino acids long) that are cleaved from other proteins, either from pathogens or from the cell itself. The MHC+peptide complex then sits on the surface of the cell, announcing either that 1) I am healthy, full of normal cell proteins, going about their business, nevermind, or 2) I have some other proteins inside, either from a viral infection (class I MHC) or from some bacterium I have just phagocytosed to deal with an infection (class II MHC). In the second case, T cells carrying the TCR receptor lock onto the MHC+peptide complex, and start up the process of killing that cell. 

MHC proteins (beige, pink) present an antigen (red) from one cell, and dock to a T-cell which recognizes the antigen+MHC complex by shape, using the TCR receptor. The additional CD8 or CD4 receptors help to verify the proper binding. While class I MHC are present on all cells, class II MHC are on phagocytosing cells like dendritic cells and macrophages that commonly ingest, reprocess, and re-display bits of encountered pathogens.  


Given that the TCR gene recombination process is unbiassed and produces a galaxy of random binding specificities, how do these cells distinguish self from non-self antigens? This is a deep question that is not fully resolved. But one major mechanism is thymic selection, which is what gives T cells their name. Special cells in the thymus display a wide range of self-antigens, and T cells, which are obliged to pass through the thymus during their maturation, are induced to commit suicide if they react to any of them. A paper from 2018 fascinatingly discussed how it is possible to create a T cell population that knows the "language" of foreign vs domestic after what is known to be a rather haphazard selection process, which displays only a partial range of self-antigens, and leaves quite a few self-reactive T cells around.

At any rate, the MHC proteins do not benefit from hyper-variation provided by genetic recombination. Yet it turns out that variation is beneficial here as well. The way foreign antigen peptides nestle in the MHC groove can be varied by mutations in the MHC molecule, providing a rich field of variation in antigen recognition and thus disease resistance. So, our MHC genes have not just a few alleles in the population, not just a few dozen, but over six thousand alleles. For each individual MHC protein, each person has only two, but over the population, there myriads with different properties. MHC was first recognized for its role in self vs non-self recognition and transplant rejection, (thus the "compatibility" in its name), and it quickly became evident that people vary tremendously in their MHC complement. And this variation plays a big role in keeping us (and all other animals) going as populations, in the face of pathogens that evolve a lot faster than we do. 

A recent paper provided a phylogenetic history of MHC molecules in monkeys, covering the last sixty million years of evolution in our lineage. It is a festival of gene birth, death, and duplication, quite apart from the smaller mutations that are constantly accumulating and cycling through the population. The MHC region carries about 200 related genes, most of which have minor roles, and only six of which (three MHC class I, and three MHC class II) follow the high-mutation pattern because they encode the main antigen presenting proteins. These genes are subject to, quite obviously, unique selective forces. 

How the MHC gene cluster looks, when aligned and identified by gene, over the primates. Note the deep divergence between the new- and old-world primates. Genes A, B, and C are the major MHC class I genes, which vary the most over this time. Note also how some of these genes have gone through extensive duplication in some old-world monkey lineages.

The first force is balancing selection. As soon as one allele becomes common, pathogens evolve to evade its presentation skills, rendering it less effective and less desirable. This results in a population full of minor variants. Indeed, for any individual person, having two MHC molecules that are the same would be bad. Being heterozygous at these genes is highly advantageous, thus enforcing both the retention of minor alleles, and an observed behavior in mating to favor partners with different MHC complements. Apparently, our MHC makeup is reflected in our personal aroma! 

A second force, conversely, is the retention of ancient alleles. It turns out that, across the old-world monkeys, many MHC alleles are preserved and cluster more closely in sequence comparisons with each other than they do with other alleles in the same species. That is, despite the general speed of MHC evolution and constant accumulation of new alleles, old alleles are preserved in all monkey populations as well, due to their distinct capabilities, under balancing selection. This is part of what makes population bottlenecks so damaging to near-extinction species. They lose critically valuable genetic resources (in the form of rare MHC alleles) that represent millions of years of accumulated variation. 

Incidentally, the trees shown here again reinforce the history of primate evolution, with new world monkeys splitting off from the old-world monkeys quite early on and developing a very distinct set of MHC molecules.

So, while virtually every other gene in the genome is being relentlessly optimized, sticking to its knitting, doing one thing and being beaten down whenever any mutation steers it from its optimized path, the MHC genes follow quite a different path, at least in portions of their sequence that provide variation in antigen binding and presentation. These genes revel in endless diversity, throw off pseudogenes at a high rate, wink out of existence and come back in other forms. Natural selection is the motor in each case, but meets the challenge of survival in different ways.