Showing posts with label deep time. Show all posts
Showing posts with label deep time. Show all posts

Saturday, July 25, 2026

Revealing the Origin of Eukaryotes

Tracing the breadcrumbs of deep phylogeny through gene duplications. 

Last week, we discussed the jamboree of mutation that is the MHC locus in animals. One part of that story was the persistent duplication and pseudogene-i-zation (which is to say, the birth and death) of genes in that locus through primate evolution, another consequence of the never-ending arms race with our pathogens. Gene duplication has much deeper significance, however, as it has given rise to new genes throughout biological evolution, as it is merely an accident of DNA (or RNA) replication, thus operationally rather frequent. Significant transitions, such as the advent of eukaryotes, and the advent of plants and of angiosperms, all feature rapid gene duplication, even whole-genome duplication, which provide grist for new functions, refinement of old functions, and speciation.

A recent paper did a deep dive into the origin of eukaryotes, perhaps the watershed event in the history of life, second only to the origin of life itself on the early Earth. It is very difficult to reconstruct what happened during these extremely early events, which these authors put as long as three billion years ago. There has been so much mutational churn in our molecules, and no fossil record to speak of, that, again like the origin of life itself, we do not have a great deal to go on. Thankfully, the discovery of the kingdom of archaea, and its especially eukaryotic-like class, the Asgard archaea, has provided a much better platform for speculating about what the originating organisms might have been like.

Eukaryotes differ from bacteria and archaea by numerous characteristics, implying extensive development/evolution before the last common ancestor from which we can trace all the extant descendants. These include the eponymous nucleus, an internal system of membranes and membrane-bound organelles, a complex cytoskeleton of both actin and microtubules, a mitochondrion, meiosis, a microtubule spindle-driven division system, and many other, more obscure molecular differences. The lack of intermediate forms is certainly frustrating from a scientific perspective, making it difficult to go through the catalog of life today to identify the steps involved in this transition. It also implies that the process ended with an evolutionary bang that in essence caused the final iteration to wipe out all the preceding forms, a bit like humans vs the many other Australopithecines and their descendants. 

These differences are enormous and momentous, making eukaryotic cells far larger, energetically capable, and complex than their antecedents. Another key characteristic is a prevalence of gene duplication and specialization, leading to much larger genomes along with all the other complexity. It has long been theorized that it was the union of the proto-mitochondrion, which was a bacterium, with the archaeal founding cell, that set the whole process off, particularly providing the massive amount of energy required for all this rise in complexity. 

However, the current authors claim that this definitively is not the case- that the mitochondrial symbiosis was a rather late event, coincident with the oxygenation of the atmosphere about two billion years ago. They argue that, since gene duplication and specialization is in any case a clear characteristic of eukaryotes, using molecular clocks to time these many gene duplications can give us relatively specific and data-rich insight into the epochs at which each of the processes that they participate in first developed. 

The authors present estimated times of origin for key duplications involved in RNA synthesis. The RNA polymerases I, II, and III were duplicated from one polymerase in archaea, and are composed of numerous gene products, as well as ancillary regulatory proteins. While some subunits are still shared, the ones that duplicated did so in the range of 2.3 to 2.8 billion years ago, before the symbiosis with the mitochondrial precursor, timed at "mFECA", or the mitochondrial first eukaryotic common ancestor, about 2.2 billion years ago. "nFECA" is nuclear first eukaryotic common ancestor, while "LECA" is last eukaryotic common ancestor. The axis at the top is time before present, in giga-years ago, or Ga.

For example, eukaryotes all have three RNA polymerases, while bacteria and archaea only have one. These RNA polymerases specialize on mRNA synthesis (RNA pol II), ribosomal RNA synthesis (RNA pol I), and tRNA, 5S RNA and U6 snRNA (RNA pol III). Why this specialization exists remains a bit mysterious, though once it started there was no going back. In any case, it happened and is totally characteristic of eukaryotes, so these duplications (which involve the duplication of numerous genes encoding proteins of the polymerase complex itself and its regulators and way-finding helpers) can be used to date eukaryogenesis, using a very carefully calibrated and species-sampled molecular clock. And what they find is that all of these duplications, given as ranges in the graphs above, with green dots at the likeliest time points, all lie prior to the point of "mFECA", which is the mitochondrial first eukaryotic common ancestor, at roughly 2.2 billion years ago. Tabulated over studies of almost a hundred other genes that are similarly duplicated and characteristic of eukaryotes, the authors place the "nFECA", or nuclear eukaryotic common ancestor way back at almost 3 billion years ago. 

That is a very long time! It is a very long time ago that the characteristics of eukaryotes started to come together, and it is also a very long time (well over a billion years) over which that development spanned before culminating in the last common ancestor of all extant eukaryotes, which the authors put at about 1.7 billion years ago. And note that that still comes a billion years before the Cambrian explosion of animal life. We are talking about really deep time here. 

If mitochondrial symbiosis came later, (causing another bout of genomic change as hundreds of bacterial genes moved from the symbiont to the nucleus), just as the atmosphere was getting oxygenated, what was going on in the preceding billion years? One can speculate that these evolving proto-eukaryotes were the top predators of their world, a bit like eukaryotic protists in microbial environments today. The competition would as always have been intense, but there were no higher life forms (i.e. animals or protists) to worry about. Size was not a problem, but rather an advantage. The focus was on quality and effectiveness in finding, disabling, and digesting prey. Thus there was little downside to complexity, unlike the constraints on prey species, which had to optimize for rapid growth and reproduction under both nutrient and predation constraints. New and costly systems for protein secretion, for prey digestion, and intelligent, responsive regulation of all cell processes might have been consistently advantageous.

I take this work as being definitive, at least in terms of settling the debate about early vs late symbiosis. It benefits from a clear theoretical basis and new, voluminous sequence data from appropriately related species in its molecular comparisons. I recommend it highly, as I can only scratch the surface of its analysis here. The advent of mitochondria was then a later step change in complexity, but brought incredible metabolic productivity that, at least in oxygen-rich environments, would have definitively superseded all the prior forms of the proto-eukaryote, and (unfortunately for us) erased that record of evolution. A later innovation, placed between mFECA and the last eukaryotic common ancestor "LECA" was sexual meiosis, another fascinating and characteristic innovation of eukaryotes. These authors thus provide a true revolution in our understanding of eukaryogenesis, extending it across well over a billion years.

One more sample from the author's presentation of molecular duplications that help time eukaryogenesis. Here, DNA repair proteins are emphasized, including some involved in meiosis (light blue dots, see legend). Some of these duplications involve genes from the mitochondrial endosymbiont, (MSH4), and some (such as EME1/2 and MUS81) come from older repair processes, but later specialized for meiosis. 


Sunday, July 12, 2026

How Do Complex Human Traits Add Up?

Notes on the genetics, traits, and evolution.

Everything about us is a trait. Not everything about our traits is genetic, though. The conundrum of nature vs nurture, of genes vs environment, and the structure and meaning of genetics goes to the heart of biology. A few traits, like eye color, are simple enough. But they are the exception, by far. Body mass index is influenced by practically every gene we have. And autism has, by this point, hundreds of contributing genes. Both traits are highly heritable, in the sense that inheritance/genes are the dominant influence, vs environment (as seen in twins) when most conditions are equal. But environment can easily be dominant over BMI when conditions change and starvation sets in. 

A puzzle that came out of the early human genome studies was how unhelpful it was to do genome-wide association studies (GWAS) to approach some of these questions- that is, studies of what variants in the population at large contribute to particular traits, especially to serious diseases. Study after study was done, and disappointment mounted that what were found were genes with minor effects, in tangential biological processes. This was supposed to be the holy grail- the payoff for sequencing the human genome- and what came up was dud after dud.

A recent paper plows over this ground again with a new mathematical synthesis of genetic structure of human traits, genetic alleles, and selection. What it finds is sort of obvious, but there are some intriguing observations along the way. It is critical to note at the outset that evolution as Darwin understood and described, and natural selection in particular, is absolutely at the heart of this or any contemporary analysis of genetics. While Darwin's understanding was revolutionary and broad, subsequent decades of work have brought these concepts to a very concrete, operational, indeed mathematical, level. 

Let's start with the concept of allele frequency, also called minor allele frequency, of MAF. Over any individual genome, there will be millions of "variants", which are coding differences from the reference genome. Which does not have any special status... it is just the genome of some guy from Buffalo. Variants (or alleles) are bases in the genome that differ from the reference. With three billion positions and four possible bases per position, that means that there are nine billion possible variants. How common is a particular variant in a population? That is its allele frequency. For GWAS and related studies, the threshold is commonly set at variants seen in the population at over one percent frequency, while minor alleles are seen under that frequency. A variant that causes some devastating disease is typically one that sprang up recently, and is heavily selected against. That is why it must have an exceedingly low allele frequency. On the other hand, a variant may have no discernable effect at all, not being selected for or against, thus just drifts along in the genome, not subject to natural selection. Such alleles may, over long periods of time, drift to higher or lower frequency by random chance.

However, the focus of GWAS association studies are variants between these extremes. These are variants that have some effect on a trait (or may be physically close to others that do, thus get "carried" along over time). At the same time, they are also common in the population, at least common enough to be discernable in an association study. Such a study needs some statistical correlation between the occurrence of the variant, and the occurrence of the trait. That means that the variant can not just appear once, but must appear many times over a large population. At the same time, a study that focuses on a trait like, say, high blood pressure, will be seeing variants that are, by definition, deleterious. That means that any allele with a large effect will be subject to strong selection, and driven out of the population. Only alleles with more modest effects will be able to survive at all, and even then, at low frequencies. So a GWAS focuses on medium-to-low prevalence variants, hunting for alleles on the loose in a large population that have typically modest effects on a given disease or other trait. Such alleles will have typically survived for tens or even hundreds of thousands of years, so they will have some complex relationship to natural selection. 

In contrast to all this is the family study, which focuses on a dramatic variant that causes some terrible disease. Such studies have been remarkably productive, because they deal with extremely rare, high-effect variants, which are as a rule very informative about the genesis of that disease. Such variants, as mentioned above, would be heavily selected against, thus disappear rapidly. But mutation is always happening, so all sorts of mutations arise in large enough populations. Assuming that, as biologists, we are interested in the core ten or fewer genes that most influence a given trait or condition, the hundreds or thousands of significant, but low-effect variants that come out of GWAS are almost by definition guaranteed to be tangential and minor. It turns out that most traits are complex, in the sense of being influenced in various minor ways by hundreds of genes.

So, what is the typical genetic structure of complex traits? That is- what this paper set out to answer. "Structure" in this case means ... what is the normal distribution of selective target / effect size versus frequency/prevalence in the population of variants that, in combination, add up to a complex trait? The assumption (as discussed above) is that the core armature of such traits does not have variants at all, due to strong selection, while the available variants in the population all have minor effects that in sum form the genetic variation seen in the trait in the population. 

While other researchers have attempted to fit the variation distribution of complex traits to typical formulas like the normal distribution, these authors found that a natural selection-informed approach gave a clear and simple result. All traits follow the same general scaling, with only two parameters- the mutational target size of the trait (that is, the proportion of the genome capable of appearing as relevant variants), and the effect that each site has on the given trait, termed (very poorly) the site's "heritability". It is important to note that every site in the genome is equally and fully heritable. The term refers to the trait, and the size of the effect from variations of that site on that trait. Summed over all sites in the genome and all variations in the population, this heritability ultimately equates to the overall variation of the trait that is genetically caused.

A comparison of two traits, and how they might look in a genetic variation study. In blue is a trait skewed towards small effect variants, with weak selection and consequently variants with higher frequency. In red is a different trait that partakes more from stronger effect variants. On the whole, this kind of difference is not common among complex traits that arise from thousands of loci. MAF = minor allele frequency; Z-score is the score in a GWAS study indicating how statistically significant the variant's effect on the trait is. Note how lower Z-score correlates with more variants at the higher MAF frequencies. At the same time, the GWAS method overall has some skew to higher MAF frequencies, since only those provide sufficient statistical power to get any results at all. Log(s) is the strength of selection; L is the genetic target size for the whole trait, and h*2 is the heritability, or proportion of the trait effect due to the causal variant.

The model they come up with accounts for the selective effect of trait effects (the larger the effect of the variation on traits, the lower its frequency in the population). It also accounts for the fact that variants that affect one trait often affect other traits as well, so the selective effect needs to be considered over all affected traits, most of which are probably unknown, but can be inferred. And conversely, most traits are composed of contributions from many genes and their alleles, sometimes thousands- they are genetically complex traits. 

The model the researchers come up with can normalize among the huge population of variants that affect a single trait, in this case blood pressure. Left shows the effect sizes of individual variants, and right shows the scaled (normalized) version from the paper's model, showing that all these variants follow the same overall rule / logic, using the custom parameters of h*2 and L- trait heritability and mutational target size. All this is to say that the lower effect variants (skewed to left) are assumed to have higher selection coefficients.

The researchers go on to show that various traits do look different under this analysis. Some differ mostly by target size, accounting for more or fewer variants, but having a similar spread of effect sizes over the population. Others differ by the scale of effects that each variant contributes, thus skewing toward higher or lower allele frequencies overall. Interestingly, they add an analysis of the age of these low-frequency, low-effect variants that are the grist for GWAS, finding that they are on average 137,000 years old. That compares with an average age of 600,000 years for variants that are neutral, thus would not come up in GWAS or be under selection. This is fascinating in its implications both for human population genetics in general, and for the fact that most human variation- even that under modest selection- predates the divergence between African and non-African populations. 


Saturday, June 20, 2026

From Icebox to Hothouse, and Back Again

Better modeling, by including the biosphere, retrodicts more of Earth's dynamic climate history.

Climate change, while ignored by the current administration, is not ignoring us. The Earth is warming well past where it has been for millions of years. But before that? While the planet has generally had stable climates, they have varied substantially through time, and have gone through occasional catastrophes. There was a little ice age, in the middle of the last millennium, thought to have been caused in part by the depopulation of the Americas due to European diseases. The ensuing regrowth of forests covering the Americas drew down CO2 from the atmosphere and cooled the climate. But more to the point, there have been far more severe episodes, both of heat (the end-Permian extinction event) and cold (the Sturtian glaciation of the Precambrian). All of these arise from CO2 levels, as CO2 is the master controller of heat in the atmosphere, thanks to the greenhouse effect. (As it is on Venus as well.) 

For example, the end-Permian extinction is thought to have been caused by unusual volcanism in what is now Siberia. Over a mere 100,000 years, this poured an estimated 26,000 petagrams of CO2 into the atmosphere, causing its concentration to shoot up to about 2500 ppm (parts per million) and temperatures to shoot up as well, killing off 90% of all species. What we are doing now is much faster, though admittedly in early days. We are pouring roughly 11 petagrams of CO2 into the atmosphere yearly, which has raised CO2 concentrations from a preindustrial 280 ppm to 427 ppm today. It would take us another one to two thousand years to cause a 90% extinction event!

A bedrock of our climate thermostat is the silicate cycle. Since the vast majority of carbon on earth is locked up in rocks, (carbonates of silicon, magnesium, and calcium), not in the biosphere, it is rocks that have a dominant effect. Volcanoes belch out CO2 in huge amounts. That CO2 slowly eats away at rocks that are exposed, re-forming carbonate compounds that are weathered off and back into the ocean. Where these compounds (with those built by shelled animals of all kinds) are gradually deposited on the sea floor and subducted back into the Earth's crust. Some of those carbonates are reduced at depth and brought forth again by volcanic activity. The more CO2 there is in the atmosphere, and the warmer it is, the more weathering happens and thus the faster CO2 levels are brought back down. That is the elegant thermostat that has kept Earth at mostly mild temperatures through its long history. 

However, this is a slow thermostat, taking hundreds of thousands of years to equilibrate. Unusual events, like an asteroid impact, prodigious volcanism, or the advent of human ingenuity, can make a mess of things way faster than the silicate cycle can deal with in its slow, grinding way. Many subtler influences can also come into play, like cycles in the tilt of the Earth towards the Sun, or continental arrangements that lead to particular patterns of ocean circulation, can create variations such as ice ages. A recent paper brought out peculiar influences from the biosphere that can also affect, and even destabilize, the thermostat on longer time horizons

The oceans are responsible for roughly half of photosynthetic productivity, and they are also where the carbonate minerals get buried. So how they react to changes in the atmosphere are very influential in the whole cycle. These authors ran half-million-year simulations of climate perturbations while including not only the silicate cycle, but also reactions by the biosphere and especially the phosphorous cycle, which has a strong influence on biological productivity. It turns out that when the atmosphere has lower levels of oxygen than we do today, (as was the case during the Precambrian epoch), high CO2 levels cause long-term rises in biosphere productivity and also in phosphate recycling out of the ocean floor. The extra phosphate increases biological productivity even more, and thus causes CO2 drawdown to persist past where the silicate cycle would level out for the long term. The result can be a rebounding ice age after a hot phase. 

Model results over 500,000 years, showing rebound from an injection of high CO2 at year 10,000. A shows concentrations of CO2 over time, B shows O2 concentration, and C shows sea ice, which goes to zero at first, the rebounds sharply, especially under the blue condition of 0.6 times current oxygen concentration in the atmosphere. It shows how exquisitely sensitive the climate is to CO2.

These models make some sense of the Precambrian climate cycles, which had a few dramatic swings that went through so-called snowball Earth phases where the entire surface of the planet seems to have iced over. The silicate cycle naturally came to the rescue eventually, spewing enough CO2 from volcanoes to overcome the snow / albedo effects of all the ice and cause a rebound hot phase. Between the rising oxygen levels and the extreme climatic swings, the stage was somehow set for the rise of animal life, leading the so-called Cambrian explosion, though there was a fair amount of simpler precursor animal live in the Precambrian.dd

https://www.science.org/doi/10.1126/science.adh7730

A schematic of the proposed cycle, with CO2 coming in from vulcanism (red) and being disposed of by various means, first and foremost the silicate cycle (blue). OC = organic carbon, P = phosphorous/phosphate, OCpetro = organic carbon weathered out of sediments, coal, limestone, and other geologic formations. Thus, the brown color shows this paper's additions to the classical silicate cycle.

While it is just a modeling paper, models are what we think and do in science. It is nice to have laboratory confirmation for areas of science (like molecular biology) that permit it, but historical sciences, especially those pertaining to whole planets as systems, have to be more forensic and speculative. This new model is a refinement on the basic silicate cycle, and thus seems a strong improvement on what has heretofore been a science of more or less back-of-the-envelope estimation. And judging from this new model, the authors propose that the next ice age is not being put off indefinitely by our profligate emissions, but rather that organic burial feedbacks will bring it closer (than 400k years away) with additional overcooling thereafter!


  • Medicine is toast. "MIRA outperformed physicians in diagnostic accuracy and made guideline-concordant, medication-safe and appropriate admission decisions."
  • A death sentence for US science.
  • Revaluing trash.
  • Apparently, fungi in the ocean are a thing.

Saturday, March 28, 2026

Death and Resurrection ... Of a Gene

The SLAMF9 gene became non-functional in the human lineage, and then later was re-activated. Why?

Biology is amazingly intricate, but it is often also needlessly complex- evidence for the haphazard, if eventually pointed, mechanisms of the evolutionary process. We will take up the discussion of "junk" DNA again next week, but molecular biology is full of redundant and excessive processes, which should certainly be mystifying from a "design" perspective. At the frontier of natural selection are neutral and near-neutral genetic elements, which change over time due to chance, lacking selection pressure towards conservation. Pseudogenes (of which we have about 20,000- almost as many as functional genes) are one form of neutral element. They are typically remnants of functional genes that have been duplicated and inactivated by mutation. They are a lively area of genome annotation because it is hard to be sure that they are really dead. Despite what looks like an inactivating mutation, they typically still produce RNA transcripts, and may produce partial or alternative proteins as well. The literature is full of experiments finding products and activities from genes annotated elsewhere as pseudogenes. And what looks like a pseudogene from one sample might just be an allele, the same gene being whole and active in other people.

So, it is hard to know what any particular genetic region is doing without a lot of evolutionary, functional, and even population analysis. A recent paper looked deeply at one gene- a gene that seems to have flipped back and forth between functional and non-functional states in the human lineage. It is a rare example of a gene coming back from what is usually a one-way trip into mutational oblivion, once its function- and thus selective pressure for conservation- have disappeared.

SLAMF9 is one of a family (signaling lymphocyte activation molecule family) of surface receptors that occur in many cells of the immune system, help activate responses in these cells, and also recognize some viruses and bacteria. They bind to each other and to other components of the immune system, creating complex signaling networks. Genes involved in our immune systems are commonly subject to rapid evolution, the arms race against our many pathogens being relentless. Sometimes that takes the form of gene inactivation, if a particular receptor, for instance, has been turned against us by a pathogen that uses it for binding and cell entry. 

This week's authors were facing a conundrum. They were studying SLAMF9, and found the mouse version easy to clone and express in the lab. But the human version ... that was another story, frustratingly impossible to express in usable amounts. When they looked at the protein sequence, they were in for a big surprise:

At the front end of SLAMF9, there is very strong conservation across mammals... except when it comes to humans! The signal peptide is what directs this protein to be inserted into the plasma membrane, and is cleaved off the mature protein. In red is highlighted the region starkly different in humans, which naturally affects (not in a good way) the signal cleavage process. "a" and "b" point to important domains of the cytoplasmic side of the final protein, which are just barely preserved/conserved in the human form.

This alignment among various mammalian versions (orthologs) of SLAMF9 shows that they are all pretty much the same... except for the human version. All the way from mouse to chimpanzee nothing has changed at the front end of this protein. That is amazing in itself, showing very strong conservation. But then after our lineage split from chimpanzees, something weird. A small segment at the front of this protein is totally different. This area is important because it carries the cleavage site of the signal sequence. The signal sequence directs the protein to be sent to the membrane (as this is a trans-membrane receptor), and this cleavage site is bad, explaining why the author's attempt to express this protein went so poorly. It might be enough for modest expression in the natural setting, but not enough for their investigations.

At the DNA level, it is clear that what happened to the protein was a double frame shift in translation, out of frame at the front, then recovered frame at the second mutation. The mutations must have been independent events, but the order of their occurrence is not known. The first intron trails off to the left, while the coding sequence tails off to the right.

When they looked at the DNA sequence, the reason for this change in the protein sequence became clearer. There was a frame shift, with only small changes in the DNA sequence that led to the bigger change in the protein sequence. On the left, there is a shift in the splice site at the end of the first intron (splice acceptor). This shifts the mRNA product by four bases (vs the start site of translation), creating a frame shift in translation, as portrayed in the amino acid codes given. On the right, there is a one nucleotide deletion, causing another frame shift that brings the translation back into the normal frame. 

They sampled all the available archeological samples from the human lineage- Neanderthals and Denisovans, and each were the same as the current human sequence. So, whatever happened did so between the split from chimpanzees and the advent of these available homo species. And what happened were two distinct events- the second frame shift and the first frame shift are independent genetic mutations. 

Which happened first? That is uncertain, but the authors show that the right-most frame shift (called g.621delT) did not influence the change in the splice site. The splice site change was caused by a series of about six mutations within the first intron, (not shown), which shifted the pattern of mRNA self-hybridization that helps direct splice site selection. So it is likely that the splice site change happened first, essentially killing the gene. And then the downstream frameshift happened later on to rescue it in a partial, not very well-expressed way. However, either mutation could have happened first to functionally kill off this gene, and then further mutation(s) to recover its function. In any case, both events happened within this roughly six-million-year time span that generated our immediate lineage, becoming firmly fixed as the only version of this gene now in our collective genome.

What might cause these events? It all goes back to the function of SLAMF9. As shown above, it is very highly conserved. But, being part of the immune system and the interface we show to pathogens, it is also on the front line of the bio-warfare arms race. As humans started ranging far beyond their original habitats, they doubtless encountered many new pathogens. It seems likely that killing off this gene might have resolved one such fight, at least for a little while, perhaps by removing a pathogen entry point. But later on, it became beneficial to recover it, which is to say that new mutations that restored its function even a little bit were evidently selected for, and spread in the population. There was a race at this point between the accumulation of more (now neutral) mutations that would have permanently inactivated this gene, and the advent of that one special mutation that could save it. The overall conservation of SLAMF9 argues that saving it must have conferred significant benefits.


Saturday, February 14, 2026

We Have Rocks in Our Heads ... And Everywhere Else, Too

On the evolution and role of iron-sulfur complexes.

Some of the more persuasive ideas about the origin of life have it beginning in the rocks of hydrothermal vents. Here was a place with plenty of energy, interesting chemistry, and proto-cellular structures available to host it. Some kind of metabolism would by this theory have come first, followed by other critical elements like membranes and RNA coding/catalysis. This early earth lacked oxygen, so iron was easily available, not prone to oxidation as now. Thus life at this early time used many minerals in its metabolic processes, as they were available for free. Now, on today's earth, they are not so free, and we have complex processes to acquire and manage them. One of the major minerals we use is the iron-sulfur complex, (similar to pyrite), which comes in a variety of forms and is used by innumerable enzymes in our cells. 

The more common iron-sulfur complexes, with sulfur in yellow, iron in orange.


The principle virtue of the iron-sulfur complex is its redox flexibility. With the relatively electronically "soft" sulfur, iron forms semi-covalent-style bonds, while being able to absorb or give up an electron safely, without destroying nearby chemicals as iron alone typically does. Depending on the structure and liganding, the voltage potential of such complexes can be tuned all over the (reduction potential) map, from -600 to +400 mV. Many other cofactors and metals are used in redox reactions, but iron-sulfur is the most common by far.

Reduction potentials (ability to take up an electron, given an electrical push) of various iron-sulfur complexes.

Researchers had assumed that, given the abundance of these elements, iron-sulfur complexes were essentially freely acquired until the great oxidation event, about two to three billion years ago, when free oxygen started rising and free iron (and sulfur) disappeared, salted away into vast geological deposits. Life faced a dilemma- how to reliably construct minerals that were now getting scarce. The most common solution was a three enzyme system in mitochondria that 1) strips a sulfur from the amino acid cysteine, a convenient source inside cells, 2) scaffolds the construction of the iron-sulfur complex, with iron coming from carrier proteins such as frataxin, and 3) employs several carrier proteins to transfer the resulting complexes to enzymes that need them. 

But a recent paper described work that alters this story, finding archaeal microbes that live anaerobically and make do with only the second of these enzymes. A deep phylogenetic analysis shows that the (#2) assembly/scaffold enzymes are the core of this process, and have existed since the last common ancestor of all life. So they are incredibly ancient, and it turns out to that iron-sulfur complexes can not just be gobbled up from the environment, at least not by any reasonably advanced life form. Rather, these complexes need to be built and managed under the care of an enzyme.

The presented structures of the dimer of SmsB (orange) and SmsC (blue) that dimerize again to make up a full iron-sulfur scaffolding and production enzyme in the archaean Methanocaldococcus jannaschii. Note the reaction scheme where ATP comes in and evicts the iron-sulfur cluster. On right is shown how ATP fits into the structure, and how it nudges the iron-sulfur binding area (blue vs green tracing).

A recent paper from this group extended their analysis to the structure of the assembly/scaffold enzyme. They find that, though it is a symmetrical dimer of a complex of two proteins, it only deals with one iron-sulfur complex at at time. It also binds and cleaves ATP. But ATP seems to have more of an inhibitory role than one that stimulates assembly directly. The authors suggest that high levels of ATP signal that less iron-sulfur complex is needed to sustain the core electron transport chains of metabolism, making this ATP inhibition an allosteric feedback control mechanism in these archaeal cells. I might add, however, that ATP binding may well also have a role in extricating the assembled iron-sulfur cluster from the enzyme, as that complex is quite well coordinated, and could use a push to pop out into the waiting arms of target enzymes.

"These ancestral systems were kept in archaea whereas they went through stepwise complexification in bacteria to incorporate additional functions for higher Fe-S cluster synthesis efficiency leading to SUF, ISC and NIF." - That is, the three-component systems present in eukaryotes, which come in three types.

In the author's structure, the iron-sulfur complex, liganded by three cysteines within the SmsC protein. But note how, facing the viewer, the complex is quite exposed, ready to be taken up by some other enzyme that has a nice empty spot for it.

Additionally, these archaea, with this simple one-step iron cluster formation pathway, get their sulfur not from cysteine, but from ambient elemental sulfur. Which is possible, as they live only in anaerobic environments, such as deep sea hydrothermal vents. So they represent a primitive condition for the whole system as may have occurred in the last common ancestor of all life. This ancestor is located at the split between bacteria and archaea, so was a fully fledged and advanced cell, far beyond the earlier glimmers of abiogenesis, the iron sulfur world, and the RNA world.


Saturday, January 24, 2026

Jonathan Singer and the Cranky Book

An eminent scientist at the end of his career writes out his thoughts and preoccupations.

Jonathan Singer was a famous scientist at my graduate school. I did not interact with him, but he played a role in attracting me to the program, as I was interested in biological membranes at the time. Singer himself studied with Linus Pauling, and they were the first to identify a human mutation in a specific gene as a cause for a specific disease- sickle cell disease. After further notable work in electron microscopy, he reached a career triumph by developing, in 1972, the fluid mosaic model of biological membranes. This revolutionized and clarified the field, showing that cells are bounded by something incredibly simple- a bilayer of phospholipids that naturally order themselves into a remarkably stable sheet, (a bubble, one might say), all organized by their charged headgroups and hydrophobic fatty tails. This model also showed that proteins would be swimming around freely in this membrane, and could be integrated in various ways, ether lightly attached on one side, or spanning it completely, thereby enabling complex channel and transporter functions. The model implied the typical length of a protein alpha helix that, by virtue of its hydrophobic side chains, would naturally be able to do this spanning function- a prediction that was spot-on. He could have easily won a Nobel for this work.

I was intrigued when I learned recently that Singer had written a book near the end of his career. It is just the kind of thing that a retired professor loves to do in the sunset of his career, sharing the wisdom and staving off the darkness by taking a stab at the book biz. And Singer's is a classic of the form- highly personal, a bit stilted, and ultimately meandering. I will review some of its high points, and then take a stab of my own at knitting together some of the interesting themes he grapples with.

For at base, Singer turns out to be a spiritual compadre of this blog. He claims to be a rationalist, in a world where, as he has it, no more than 9% of people are rational. Definition? It is the poll question of whether one believes that god created man, rather than the other way around. Singer recognizes that the world around him is crazy, and that the communities he has been a part of have been precious oases amid the general indifference and grasping of the world. But changing it? He is rather fatalistic about that, recognizing that reason is up against overwhelming forces.

His specific themes cover a great deal of biology, and then some more mystical reflections on balance and diversity in biology, and later, in capitalism and politics. He points out that the nature/nurture debate has been settled by twin studies. Nature, which is to say, genetics, is the dominant influence on human characteristics, including a wide variety of psychological traits, including intelligence. Environment and nurture is critical for reaching one's highest potential, and for using it in socially constructive ways, but the limits of that potential are largely set by one's genes. Singer does not, however, draw the inevitable conclusion from these observations, which is that some kind of long-term eugenic approach would be beneficial to our collective future, assuming machines do not replace us forthwith. Biologists know that very small selective coefficients can have big effects, so nothing drastic is needed. But what criteria to use- that is the sticky part. Just as success in the capitalist system hardly signals high moral or personal qualities, nor does incarceration by the justice system always show low ones. It is virtually an insoluble problem, so we muddle along, destined probably for continued cycles of Spenglerian civilizational collapse.

Turning to social affairs, Singer settles on "structural chaos" as his description of how the scientific enterprise works, and how capitalism at large works. With a great deal of waste, and misdirected effort, it nevertheless ends up providing good results- better than those that top-down direction can provide. He seems a sigh a little that "scientific" methods of social organization, such as those in Soviet Russia, were so ineffective, and that the best we can do is to muddle along with the spontaneous entrepreneurship and occasional flashes of innovation that push the process along. Not to mention the "monstrous vulgarity" of advertising, etc. Likewise, democracy is a mess, with most people totally incapable of making the reasoned decisions needed to maintain it. Again, the chaos of democracy is sadly the best we can do, and the duty of rational people, in Singer's view, is to keep alive the flame of intellectual freedom while outside pressures constantly threaten.

Art, and science.

What can we do with this? I think that the unifying thread that Singer was groping for was competition. One can frame competition as a universal principle that shapes the physical, biological, and social worlds. Put two children on a teeter-totter, and you can see how physical forces (e.g. gravitation) compete all the time, subtly producing equilibria that characterize the universe. Chemical equilibria are likewise a product of constant competition, even including the perpetual jostling of phospholipids to find their lowest energy configuration amidst the biological membrane bilayer, which has the side-effect of creating such a stable, yet highly flexible, structure. With Darwin, competition reaches its apotheosis- the endless proliferation, diversification, and selection of organisms. Singer marvels at the fragility of individual life, at the same time that life writ large is so incredibly durable and prolific. Well, the mechanism behind that is competition. And naturally, economics of any free kind, including capitalism and grant-making in science, are based on competition as well- the natural principle that selects which products are useful, which employees are productive, and which technologies are helpful. Waste is part of the process, as diversity amidst excess production is the essential ingredient for subsequent selection. 

And yet.. something is missing. The earth's biosphere would still be a mere bacterial soup if competition were the only principle at work. Bacteria (and their viruses) are the most streamlined competition machines- battlebots of the living world. It took cooperation between a bacterial cell and an archaeal cell to make a revolutionary new entity- the eukaryotic cell. It then took some more cooperation for eukaryotic cells to band together into bodies, making plants and animals. And among animals, cooperation in modest amounts provides for reproduction, family structure, flock structures, and even complex insect societies. It is with humans that cooperation and competition reach their most complex heights, for we are able to regulate ourselves, rationally. We make rules. 

Without rules, human society is anarchic mayhem- a trumpian, dystopian and corrupt nightmare. With them, it (ideally) balances competition with cooperation to harness the benefits of each. Our devotion to sports can be seen as a form of rule worship, and explicit management of the competitive landscape. Can there be too many rules? Absolutely, there are dangers on both sides. Take China as an example. In the last half-century, it revamped its system of rules to lower the instability of political competition, harness the power of economic competition, and completely transform its society. 

The most characteristic and powerful human institution may be the legislature, which is our ongoing effort to make rational rules regulating how the incredibly powerful motive force of competition shapes our lives. Our rules, in the US, were authored, at the outset, by the founders, who were- drumroll please- rationalists. To read the Federalist Papers is to see exquisite reasoning drawing on wide historical precedent, and particularly on the inspirations of the rationalist enlightenment, to formulate a new set of rules mediating between cooperation and competition. Not only were they more fair than the old rules, but they were designed for perpetual improvement and adjustment. The founding was, at base, a rationlist moment, when characters like Franklin, Hamilton, Madison, and Jefferson- deists at best and rationalists through and through, led the new country into a hopeful, constitutional future. At the current moment, two hundred and fifty years on, as our institutions are being wantonly destroyed and anything resembling reason, civility, and truth is under particularly vengeful attack, we should appreciate and own that heritage, which informs a true patriotism against the forces of darkness.


Saturday, December 13, 2025

Mutations That Make Us Human

The ongoing quest to make biologic sense of genomic regions that differentiate us from other apes.

Some people are still, at this late date, taken aback by the fact that we are animals, biologically hardly more than cousins to fellow apes like the chimpanzee, and descendants through billions of years of other life forms far more humble. It has taken a lot of suffering and drama to get to where we are today. But what are those specific genetic endowments that make us different from the other apes? That, like much of genetics and genetic variation, is a tough question to answer.

At the DNA level, we are roughly one percent different from chimpanzees. A recent sequencing of great apes provided a gross overview of these differences. There are inversions, and larger changes in junk DNA that can look like bigger differences, but these have little biological importance, and are not counted in the sequence difference. A difference of one percent is really quite large. For a three gigabyte genome, that works out to 30 million differences. That is plenty of room for big things to happen.

Gross alignment of one chromosome between the great apes. [HSA- human, PTR- chimpanzee, PPA- bonobo, GGO- gorilla, PPY- orangutan (Borneo), PAB- orangutan (Sumatra)]. Fully aligned regions (not showing smaller single nucleotide differences) are shown in blue. Large inversions of DNA order are shown in yellow. Other junk DNA gains and losses are shown in red, pink, purple. One large-scale jump of a DNA segment is show in green. One can see that there has been significant rearrangement of genomes along the way, even as most of this chromosome (and others as well) are easly alignable and traceable through the evolutionary tree.


But most of those differences are totally unimportant. Mutations happen all the time, and most have no effect, since most positions (particularly the most variable ones) in our DNA are junk, like transposons, heterochromatin, telomeres, centromeres, introns, intergenic space, etc. Even in protein-coding genes, a third of the positions are "synonymous", with no effect on the coded amino acid, and even when an amino acid is changed, that protein's function is frequently unaffected. The next biggest group of mutations have bad effects, and are selected against. These make up the tragic pool of genetic syndromes and diseases, from mild to severe. Only a tiny proportion of mutations will have been beneficial at any point in this story. But those mutations have tremendous power. They can drag along their local DNA regions as they are positively selected, and gain "fixation" in the genome, which is to say, they are sufficiently beneficial to their hosts that they outcompete all others, with the ultimate result that mutation becomes universal in the population- the new standard. This process happens in parallel, across all positions of the genome, all at the same time. So a process that seems painfully slow can actually add up to quite a bit of change over evolutionary time, as we see.

So the hunt was on to find "human accelerated regions" (HAR), which are parts of our genome that were conserved in other apes, but suddenly changed on the way to humans. There roughly three thousand such regions, but figuring out what they might be doing is quite difficult, and there is a long tail from strong to weak effects. There are two general rationales for their occurrence. First, selection was lost over a genomic region, if that function became unimportant. That would allow faster mutation and divergence from the progenitors. Or second, some novel beneficial mutation happened there, bringing it under positive selection and to fixation. Some recent work found, interestingly, that clusters of mutations in HAR segments often have countervailing effects, with one major mutation causing one change, and a few other mutations (vs the ancestral sequence) causing opposite changes, in a process hypothesized to amount to evolutionary fine tuning. 

A second property of HARs is that they are overwhelmingly not in coding regions of the genome, but in regulatory areas. They constitute fine tuning adjustments of timing and amount of gene regulation, not so much changes in the proteins produced. That is, our evolution was more about subtle changes in management of processes than of the processes themselves. A recent paper delved in detail into HAR5, one of the strongest such regions, (that is, strongest prior conservation, compared with changes in human sequence), which lies in the regulatory regions upstream of Frizzled8 (FZD8). FZD8 is a cell surface receptor, which receives signals from a class of signaling molecules called WNT (wingless and int). These molecules were originally discovered in flies, where they signal body development programs, allowing cells to know where they are and when they are in the developmental program, in relation to cells next door, and then to grow or migrate as needed. They have central roles in embryonic development, in organ development, and also in cancer, where their function is misused.

For our story, the WNT/FZD8 circuit is important in fetal brain development. Our brains undergo massive cell division and migration during fetal development, and clearly this is one of the most momentous and interesting differences between ourselves and all other animals. The current authors made mutations in mice that reproduce some of the HAR5 sequences, and investigated their effects. 

Two mouse brains at three months of age, one with the human version of the HAR5 region. Hard to see here, but the latter brain is ~7% bigger.

The authors claim that these brains, one with native mouse sequence, and the other with the human sequences from HAR5, have about a seven percent difference in mass. Thus the HAR5 region, all by itself, explains about one fourteenth of the gross difference in brain size between us and chimpanzees. 

HAR5 is a 619 base-pair region with only four sequence differences between ourselves and chimpanzees. It lies 300,000 bases upstream of FZD8, in a vast region of over a million base pairs with no genes. While this region contains many regulatory elements, (generally called enhancers or enhancer modules, only some of which are mapped), it is at the same time an example of junk DNA, where most of the individual positions in this vast sea of DNA are likely of little significance. The multifarious regulation by all these modules is of course important because this receptor participates in so many different developmental programs, and has doubtless been fine-tuned over the millennia not just for brain development, but for every location and time point where it is needed.

Location of the FZD8 gene, in the standard view of the genome at NIH. I have added an arrow that points to the tiny (in relative terms) FZD8 coding region (green), and a star at the location of HAR5, far upstream among a multitude of enhancer sequences. One can see that this upstream region is a vast area (of roughly 1.5 million bases) with no other genes in sight, providing space for extremely complicated and detailed regulation, little of which is as yet characterized.

Diving into the HAR5 functions in more detail, the authors show that it directly increases FZD8 gene expression, (about 2 fold, in very rough terms), while deleting the region from mice strongly decreases expression in mice. Of the four individual base changes in the HAR5 region, two have strong (additive) effects increasing FZD8 expression, while the other two have weaker, but still activating, effects. Thus, no compensatory regulation here.. it is full speed ahead at HAR5 for bigger brain size. Additionally, a variant in human populations that is responsible for autism spectrum disorders also resides in this region, and the authors show that this change decreases FZD8 expression about 20%. Small numbers, sure, but for a process that directs cell division over many cycles in early brain development, this kind of difference can have profound effects.


The HAR5 region causes increased transcription of FZD8, in mice, compared to the native version and a deletion.

The HAR5 region causes increased cell proliferation in embryonic day 14.5 brain areas, stained for neural markers.

"This reveals Hs-HARE5 modifies radial glial progenitor behavior, with increased self-renewal at early developmental stages followed by expanded neurogenic potential. ... Using these orthogonal strategies we show four human-specific variants in HARE5 drive increased enhancer activity which promotes progenitor proliferation. These findings illustrate how small changes in regulatory DNA can directly impact critical signaling pathways and brain development."

So there you have it. The nuts and bolts of evolution, from the molecular to the cellular, the organ, and then the organismal, levels. Humans do not just have bigger brains, but better brains, and countless other subtle differences all over the body. Each of these is directed by genetic differences, as the combined inheritance of the last six million years since our divergence versus chimpanzees. Only with the modern molecular tools can we see Darwin's vision come into concrete focus, as particular, even quantum, changes in the code, and thus biology, of humanity. There is a great deal left to decipher, but the answers are all in there, waiting.


Saturday, September 6, 2025

How to Capture Solar Energy

Charge separation is handled totally differently by silicon solar cells and by photosynthetic organisms.

Everyone comes around sooner or later to the most abundant and renewable form of energy, which is the sun. The current administration may try to block the future, but solar power is the best power right now and will continue to gain on other sources. Likewise, life started by using some sort of geological energy, or pre-existing carbon compounds, but inevitably found that tapping the vast powers streaming in from the sun was the way to really take over the earth. But how does one tap solar energy? It is harder than it looks, since it so easily turns into heat and lost energy. Some kind of separation and control are required, to isolate the power (that is to say, the electron that was excited by the photon of light), and harness it to do useful work.

Silicon solar cells and photosynthesis represent two ways of doing this, and are fundamentally, even diametrically, different solutions to this problem. So I thought it would be interesting to compare them in detail. Silicon is a semiconductor, torn between trapping its valence electrons in silicon atoms, or distributing them around in a conduction band, as in metals. With elemental doping, silicon can be manipulated to bias these properties, and that is the basis of the solar cell.

Schematic of a silicon solar cell. A static voltage exists across the N-type to P-type boundary, sweeping electrons freed by the photoelectric effect (light) up to the conducting electrode layer.


Solar cells have one side doped to N status, and the bulk set to P doping status. While the bulk material is neutral on both sides, at the boundary, a static charge scheme is set up where electrons are attracted into the P-side, and removed from the N-side. This static voltage has very important effects on electrons that are excited by incoming light and freed from their silicon atoms. These high energy electrons enter the conduction band of the material, and can migrate. Due to the prevailing field, they get swept towards the N side, and thus are separated and can be siphoned off with wires. The current thus set up can exert a pressure of about 0.6 volt. That is not much, nor is it equivalent to the 2 to 3 electron volts received from each visible photon. So a great deal of energy is lost as heat.

Solar cells do not care about capturing each energized electron in detail. Their purpose is to harvest a bulk electrical voltage + current with which to do some work in our electrical grids. Photosynthesis takes an entirely different approach, however. This may be mostly for historical and technical reasons, but also because part of its purpose is to do chemical work with the captured electrons. Biology tends to take a highly controlling approach to chemistry, using precise shapes, functional groups, and electrical environments to guide reactions to exact ends. While some of the power of photosynthesis goes toward pumping protons out of the membrane, setting up a gradient later used to make ATP, about half is used for other things like splitting water to replace lost electrons, and making reducing chemicals like NADPH.

A portion of a poster about the core processes of photosynthesis. It provides a highly accurate portrayal of the two photosystems and their transactions with electrons and protons.

In plants, photosynthesis is a chain of processes focused around two main complexes, photosystems I and II, and all occurring within membranes- the thylakoid membranes of the chloroplast. Confusingly, photosystem II comes first, accepting light, splitting water, pumping some protons, and sending out a pair of electrons on mobile plastoquinones, which eventually find their way to photosystem I, which jacks up their energy again using another quantum of light, to produce NADPH. 

Photosystem II is full of chlorophyll pigments, which are what get excited by visible photons. But most of them are "antenna" chlorophylls, passing the excitation along to a pair of centrally located chlorophylls. Note that the light energy is at this point passed as a molecular excitation, not as a free electron. This passage may happen by Förster resonance energy transfer, but is so fast and efficient that stronger Redfield coupling may be involved as well. Charge separation only happens at the reaction center, where an excited electron is popped out to a chain of recipients. The chlorophylls are organized so that the pair at the reaction center have a slightly lower energy of excitation, thus serve as a funnel for excitation energy from the antenna system. These transfers are extremely rapid, on the picosecond time scale.

It is interesting to note tangentially that only red light energy is used. Chlorophylls have two excitation states, excited by red light (680 nm = 1.82 eV) and blue light (400-450 nm, 2.76 eV) (note the absence of green absorbance). The significant extra energy from blue light is wasted, radiated away to let it (the excited electron) relax to the lower excitation state, which is then passed though the antenna complex as though it had come from red light. 

Charge separation is managed precisely at the photosystem II reaction center through a series of pigments of graded energy capacity, sending the excited electron first to a neighboring chlorophyll, then to a pheophytin, then to a pair of iron-coordinated quinones, which then pass two electrons to a plastoquinone that is released to the local membrane, to float off to the cytochrome b6f complex. In photosystem II, another two photons of light are separately used to power the splitting of one water molecule, (giving two electrons and pumping two protons). So the whole process, just within photosystem II, yields, per four light quanta, four protons pumped from one side of the membrane to the other. Since the ATP sythetase uses about three protons per ATP, this nets just over one ATP per four photons. 

Some of the energetics of photosystem II. The orientations and structures of the reaction center paired chlorophylls (Pd1, Pd2), the neighboring chlorophyll (Chl), and then the pheophytin (Ph) and quinones (Qa, Qb) are shown in the inset. Energy of the excited electron is sacrifice gradually to accomplish the charge separation and channeling, down to the final quinone pairing, after which the electrons are released to a plastoquinone and send to another complex in the chain.

So the principles of silicon and biological solar cells are totally different in detail, though each gives rise to a delocalized field, one of electrons flowing with a low potential, and the other of protons used later for ATP generation. Each energy system must have a way to pop off an excited electron in a controlled, useful way that prevents it from recombining with the positive ion it came from. That is why there is such an ornate conduction pathway in photosystem II to carry that electron away. Overall, points go to the silicon cell for elegance and simplicity, and we in our climate crisis are the beneficiaries, if we care to use it. 

But the photosynthetic enzymes are far, far older. A recent paper pointed out that no only are photosystems II and I clearly cousins of each other, but it is likely that, contrary to the consensus heretofore, photosystem II is the original version, at least of the various photosystems that currently exist. All the other photosystems (including those in bacteria that lack oxygen stripping ability) carry traces of the oxygen evolving center. It makes sense that getting electrons is a fundamental part of the whole process, even though that chemistry is quite challenging. 

That in turn raises a big question- if oxygen evolving photosystems are primitive (originating very roughly with the last common ancestor of all life, about four billion years ago) then why was earth's atmosphere oxygenated only from two billion years ago onward? It had been assumed that this turn in Earth history marked the evolution of photosystem II. The authors point out additionally that there is also evidence for the respiratory use of oxygen from these extremely early times as well, despite the lack of free oxygen. Quite perplexing, (and the authors decline to speculate), but one gets the distinct sense that possibly life, while surprisingly complex and advanced from early times, was not operating at the scale it does today. For example, colonization of land had to await the buildup of sufficient oxygen in the atmosphere to provide a protective ozone layer against UV light. It may have taken the advent of eukaryotes, including cyanobacterial-harnessing plants, to raise overall biological productivity sufficiently to overcome the vast reductive capacity of the early earth. On the other hand, speculation about the evolution of early life based on sequence comparisons (as these authors do) is notoriously prone to artifacts, since what evolves at vanishingly slow rates today (such as the photosystem core proteins) must have originally evolved at quite a rapid clip to attain the functions now so well conserved. We simply can not project ancient ages (at the four billion year time scales) from current rates of change.