Showing posts with label article review. Show all posts
Showing posts with label article review. Show all posts

Saturday, July 25, 2026

Revealing the Origin of Eukaryotes

Tracing the breadcrumbs of deep phylogeny through gene duplications. 

Last week, we discussed the jamboree of mutation that is the MHC locus in animals. One part of that story was the persistent duplication and pseudogene-i-zation (which is to say, the birth and death) of genes in that locus through primate evolution, another consequence of the never-ending arms race with our pathogens. Gene duplication has much deeper significance, however, as it has given rise to new genes throughout biological evolution, as it is merely an accident of DNA (or RNA) replication, thus operationally rather frequent. Significant transitions, such as the advent of eukaryotes, and the advent of plants and of angiosperms, all feature rapid gene duplication, even whole-genome duplication, which provide grist for new functions, refinement of old functions, and speciation.

A recent paper did a deep dive into the origin of eukaryotes, perhaps the watershed event in the history of life, second only to the origin of life itself on the early Earth. It is very difficult to reconstruct what happened during these extremely early events, which these authors put as long as three billion years ago. There has been so much mutational churn in our molecules, and no fossil record to speak of, that, again like the origin of life itself, we do not have a great deal to go on. Thankfully, the discovery of the kingdom of archaea, and its especially eukaryotic-like class, the Asgard archaea, has provided a much better platform for speculating about what the originating organisms might have been like.

Eukaryotes differ from bacteria and archaea by numerous characteristics, implying extensive development/evolution before the last common ancestor from which we can trace all the extant descendants. These include the eponymous nucleus, an internal system of membranes and membrane-bound organelles, a complex cytoskeleton of both actin and microtubules, a mitochondrion, meiosis, a microtubule spindle-driven division system, and many other, more obscure molecular differences. The lack of intermediate forms is certainly frustrating from a scientific perspective, making it difficult to go through the catalog of life today to identify the steps involved in this transition. It also implies that the process ended with an evolutionary bang that in essence caused the final iteration to wipe out all the preceding forms, a bit like humans vs the many other Australopithecines and their descendants. 

These differences are enormous and momentous, making eukaryotic cells far larger, energetically capable, and complex than their antecedents. Another key characteristic is a prevalence of gene duplication and specialization, leading to much larger genomes along with all the other complexity. It has long been theorized that it was the union of the proto-mitochondrion, which was a bacterium, with the archaeal founding cell, that set the whole process off, particularly providing the massive amount of energy required for all this rise in complexity. 

However, the current authors claim that this definitively is not the case- that the mitochondrial symbiosis was a rather late event, coincident with the oxygenation of the atmosphere about two billion years ago. They argue that, since gene duplication and specialization is in any case a clear characteristic of eukaryotes, using molecular clocks to time these many gene duplications can give us relatively specific and data-rich insight into the epochs at which each of the processes that they participate in first developed. 

The authors present estimated times of origin for key duplications involved in RNA synthesis. The RNA polymerases I, II, and III were duplicated from one polymerase in archaea, and are composed of numerous gene products, as well as ancillary regulatory proteins. While some subunits are still shared, the ones that duplicated did so in the range of 2.3 to 2.8 billion years ago, before the symbiosis with the mitochondrial precursor, timed at "mFECA", or the mitochondrial first eukaryotic common ancestor, about 2.2 billion years ago. "nFECA" is nuclear first eukaryotic common ancestor, while "LECA" is last eukaryotic common ancestor. The axis at the top is time before present, in giga-years ago, or Ga.

For example, eukaryotes all have three RNA polymerases, while bacteria and archaea only have one. These RNA polymerases specialize on mRNA synthesis (RNA pol II), ribosomal RNA synthesis (RNA pol I), and tRNA, 5S RNA and U6 snRNA (RNA pol III). Why this specialization exists remains a bit mysterious, though once it started there was no going back. In any case, it happened and is totally characteristic of eukaryotes, so these duplications (which involve the duplication of numerous genes encoding proteins of the polymerase complex itself and its regulators and way-finding helpers) can be used to date eukaryogenesis, using a very carefully calibrated and species-sampled molecular clock. And what they find is that all of these duplications, given as ranges in the graphs above, with green dots at the likeliest time points, all lie prior to the point of "mFECA", which is the mitochondrial first eukaryotic common ancestor, at roughly 2.2 billion years ago. Tabulated over studies of almost a hundred other genes that are similarly duplicated and characteristic of eukaryotes, the authors place the "nFECA", or nuclear eukaryotic common ancestor way back at almost 3 billion years ago. 

That is a very long time! It is a very long time ago that the characteristics of eukaryotes started to come together, and it is also a very long time (well over a billion years) over which that development spanned before culminating in the last common ancestor of all extant eukaryotes, which the authors put at about 1.7 billion years ago. And note that that still comes a billion years before the Cambrian explosion of animal life. We are talking about really deep time here. 

If mitochondrial symbiosis came later, (causing another bout of genomic change as hundreds of bacterial genes moved from the symbiont to the nucleus), just as the atmosphere was getting oxygenated, what was going on in the preceding billion years? One can speculate that these evolving proto-eukaryotes were the top predators of their world, a bit like eukaryotic protists in microbial environments today. The competition would as always have been intense, but there were no higher life forms (i.e. animals or protists) to worry about. Size was not a problem, but rather an advantage. The focus was on quality and effectiveness in finding, disabling, and digesting prey. Thus there was little downside to complexity, unlike the constraints on prey species, which had to optimize for rapid growth and reproduction under both nutrient and predation constraints. New and costly systems for protein secretion, for prey digestion, and intelligent, responsive regulation of all cell processes might have been consistently advantageous.

I take this work as being definitive, at least in terms of settling the debate about early vs late symbiosis. It benefits from a clear theoretical basis and new, voluminous sequence data from appropriately related species in its molecular comparisons. I recommend it highly, as I can only scratch the surface of its analysis here. The advent of mitochondria was then a later step change in complexity, but brought incredible metabolic productivity that, at least in oxygen-rich environments, would have definitively superseded all the prior forms of the proto-eukaryote, and (unfortunately for us) erased that record of evolution. A later innovation, placed between mFECA and the last eukaryotic common ancestor "LECA" was sexual meiosis, another fascinating and characteristic innovation of eukaryotes. These authors thus provide a true revolution in our understanding of eukaryogenesis, extending it across well over a billion years.

One more sample from the author's presentation of molecular duplications that help time eukaryogenesis. Here, DNA repair proteins are emphasized, including some involved in meiosis (light blue dots, see legend). Some of these duplications involve genes from the mitochondrial endosymbiont, (MSH4), and some (such as EME1/2 and MUS81) come from older repair processes, but later specialized for meiosis. 


Saturday, July 18, 2026

MHC Through Evolution: Breaking All the Rules

The immunologically critical MHC gene cluster plays by its own rules through the arms race of life.

Last week, I discussed the very general landscape of variation in the human genome, specifically the tradeoff between prevalence in the population and effect size. Given that the vast majority of variants are deleterious, those that affect our phenotypic traits are more heavily selected against the greater effect they have. The result is that at a gross level, most genes and most traits share the same general distribution of lots of variants (alleles) with minor effects, and far fewer with large effects. And those with minor effects also turn out to be tangential for biologists, rarely informative about the nature of the traits they (sort-of) affect.

This week, another paper and another view of evolution, though the eyes of one the more critical genes of the immune system, the major histocompatibility complex, or MHC group of genes. While our adaptive immune system has developed the extraordinary and powerful ability (though semi-controlled DNA recombination of the antibody and TCR genes) to recognize practically any antigen, foreign or domestic, that system requires stringent controls. One of those controls on T cells, which carry the antigen-recognizing TCR receptor, is that it can only see antigens that are "presented" on MHC molecules. MHC proteins have a surface cleft that gets loaded with and holds small peptides (8 to 12 amino acids long) that are cleaved from other proteins, either from pathogens or from the cell itself. The MHC+peptide complex then sits on the surface of the cell, announcing either that 1) I am healthy, full of normal cell proteins, going about their business, nevermind, or 2) I have some other proteins inside, either from a viral infection (class I MHC) or from some bacterium I have just phagocytosed to deal with an infection (class II MHC). In the second case, T cells carrying the TCR receptor lock onto the MHC+peptide complex, and start up the process of killing that cell. 

MHC proteins (beige, pink) present an antigen (red) from one cell, and dock to a T-cell which recognizes the antigen+MHC complex by shape, using the TCR receptor. The additional CD8 or CD4 receptors help to verify the proper binding. While class I MHC are present on all cells, class II MHC are on phagocytosing cells like dendritic cells and macrophages that commonly ingest, reprocess, and re-display bits of encountered pathogens.  


Given that the TCR gene recombination process is unbiassed and produces a galaxy of random binding specificities, how do these cells distinguish self from non-self antigens? This is a deep question that is not fully resolved. But one major mechanism is thymic selection, which is what gives T cells their name. Special cells in the thymus display a wide range of self-antigens, and T cells, which are obliged to pass through the thymus during their maturation, are induced to commit suicide if they react to any of them. A paper from 2018 fascinatingly discussed how it is possible to create a T cell population that knows the "language" of foreign vs domestic after what is known to be a rather haphazard selection process, which displays only a partial range of self-antigens, and leaves quite a few self-reactive T cells around.

At any rate, the MHC proteins do not benefit from hyper-variation provided by genetic recombination. Yet it turns out that variation is beneficial here as well. The way foreign antigen peptides nestle in the MHC groove can be varied by mutations in the MHC molecule, providing a rich field of variation in antigen recognition and thus disease resistance. So, our MHC genes have not just a few alleles in the population, not just a few dozen, but over six thousand alleles. For each individual MHC protein, each person has only two, but over the population, there myriads with different properties. MHC was first recognized for its role in self vs non-self recognition and transplant rejection, (thus the "compatibility" in its name), and it quickly became evident that people vary tremendously in their MHC complement. And this variation plays a big role in keeping us (and all other animals) going as populations, in the face of pathogens that evolve a lot faster than we do. 

A recent paper provided a phylogenetic history of MHC molecules in monkeys, covering the last sixty million years of evolution in our lineage. It is a festival of gene birth, death, and duplication, quite apart from the smaller mutations that are constantly accumulating and cycling through the population. The MHC region carries about 200 related genes, most of which have minor roles, and only six of which (three MHC class I, and three MHC class II) follow the high-mutation pattern because they encode the main antigen presenting proteins. These genes are subject to, quite obviously, unique selective forces. 

How the MHC gene cluster looks, when aligned and identified by gene, over the primates. Note the deep divergence between the new- and old-world primates. Genes A, B, and C are the major MHC class I genes, which vary the most over this time. Note also how some of these genes have gone through extensive duplication in some old-world monkey lineages.

The first force is balancing selection. As soon as one allele becomes common, pathogens evolve to evade its presentation skills, rendering it less effective and less desirable. This results in a population full of minor variants. Indeed, for any individual person, having two MHC molecules that are the same would be bad. Being heterozygous at these genes is highly advantageous, thus enforcing both the retention of minor alleles, and an observed behavior in mating to favor partners with different MHC complements. Apparently, our MHC makeup is reflected in our personal aroma! 

A second force, conversely, is the retention of ancient alleles. It turns out that, across the old-world monkeys, many MHC alleles are preserved and cluster more closely in sequence comparisons with each other than they do with other alleles in the same species. That is, despite the general speed of MHC evolution and constant accumulation of new alleles, old alleles are preserved in all monkey populations as well, due to their distinct capabilities, under balancing selection. This is part of what makes population bottlenecks so damaging to near-extinction species. They lose critically valuable genetic resources (in the form of rare MHC alleles) that represent millions of years of accumulated variation. 

Incidentally, the trees shown here again reinforce the history of primate evolution, with new world monkeys splitting off from the old-world monkeys quite early on and developing a very distinct set of MHC molecules.

So, while virtually every other gene in the genome is being relentlessly optimized, sticking to its knitting, doing one thing and being beaten down whenever any mutation steers it from its optimized path, the MHC genes follow quite a different path, at least in portions of their sequence that provide variation in antigen binding and presentation. These genes revel in endless diversity, throw off pseudogenes at a high rate, wink out of existence and come back in other forms. Natural selection is the motor in each case, but meets the challenge of survival in different ways.


Sunday, July 12, 2026

How Do Complex Human Traits Add Up?

Notes on the genetics, traits, and evolution.

Everything about us is a trait. Not everything about our traits is genetic, though. The conundrum of nature vs nurture, of genes vs environment, and the structure and meaning of genetics goes to the heart of biology. A few traits, like eye color, are simple enough. But they are the exception, by far. Body mass index is influenced by practically every gene we have. And autism has, by this point, hundreds of contributing genes. Both traits are highly heritable, in the sense that inheritance/genes are the dominant influence, vs environment (as seen in twins) when most conditions are equal. But environment can easily be dominant over BMI when conditions change and starvation sets in. 

A puzzle that came out of the early human genome studies was how unhelpful it was to do genome-wide association studies (GWAS) to approach some of these questions- that is, studies of what variants in the population at large contribute to particular traits, especially to serious diseases. Study after study was done, and disappointment mounted that what were found were genes with minor effects, in tangential biological processes. This was supposed to be the holy grail- the payoff for sequencing the human genome- and what came up was dud after dud.

A recent paper plows over this ground again with a new mathematical synthesis of genetic structure of human traits, genetic alleles, and selection. What it finds is sort of obvious, but there are some intriguing observations along the way. It is critical to note at the outset that evolution as Darwin understood and described, and natural selection in particular, is absolutely at the heart of this or any contemporary analysis of genetics. While Darwin's understanding was revolutionary and broad, subsequent decades of work have brought these concepts to a very concrete, operational, indeed mathematical, level. 

Let's start with the concept of allele frequency, also called minor allele frequency, of MAF. Over any individual genome, there will be millions of "variants", which are coding differences from the reference genome. Which does not have any special status... it is just the genome of some guy from Buffalo. Variants (or alleles) are bases in the genome that differ from the reference. With three billion positions and four possible bases per position, that means that there are nine billion possible variants. How common is a particular variant in a population? That is its allele frequency. For GWAS and related studies, the threshold is commonly set at variants seen in the population at over one percent frequency, while minor alleles are seen under that frequency. A variant that causes some devastating disease is typically one that sprang up recently, and is heavily selected against. That is why it must have an exceedingly low allele frequency. On the other hand, a variant may have no discernable effect at all, not being selected for or against, thus just drifts along in the genome, not subject to natural selection. Such alleles may, over long periods of time, drift to higher or lower frequency by random chance.

However, the focus of GWAS association studies are variants between these extremes. These are variants that have some effect on a trait (or may be physically close to others that do, thus get "carried" along over time). At the same time, they are also common in the population, at least common enough to be discernable in an association study. Such a study needs some statistical correlation between the occurrence of the variant, and the occurrence of the trait. That means that the variant can not just appear once, but must appear many times over a large population. At the same time, a study that focuses on a trait like, say, high blood pressure, will be seeing variants that are, by definition, deleterious. That means that any allele with a large effect will be subject to strong selection, and driven out of the population. Only alleles with more modest effects will be able to survive at all, and even then, at low frequencies. So a GWAS focuses on medium-to-low prevalence variants, hunting for alleles on the loose in a large population that have typically modest effects on a given disease or other trait. Such alleles will have typically survived for tens or even hundreds of thousands of years, so they will have some complex relationship to natural selection. 

In contrast to all this is the family study, which focuses on a dramatic variant that causes some terrible disease. Such studies have been remarkably productive, because they deal with extremely rare, high-effect variants, which are as a rule very informative about the genesis of that disease. Such variants, as mentioned above, would be heavily selected against, thus disappear rapidly. But mutation is always happening, so all sorts of mutations arise in large enough populations. Assuming that, as biologists, we are interested in the core ten or fewer genes that most influence a given trait or condition, the hundreds or thousands of significant, but low-effect variants that come out of GWAS are almost by definition guaranteed to be tangential and minor. It turns out that most traits are complex, in the sense of being influenced in various minor ways by hundreds of genes.

So, what is the typical genetic structure of complex traits? That is- what this paper set out to answer. "Structure" in this case means ... what is the normal distribution of selective target / effect size versus frequency/prevalence in the population of variants that, in combination, add up to a complex trait? The assumption (as discussed above) is that the core armature of such traits does not have variants at all, due to strong selection, while the available variants in the population all have minor effects that in sum form the genetic variation seen in the trait in the population. 

While other researchers have attempted to fit the variation distribution of complex traits to typical formulas like the normal distribution, these authors found that a natural selection-informed approach gave a clear and simple result. All traits follow the same general scaling, with only two parameters- the mutational target size of the trait (that is, the proportion of the genome capable of appearing as relevant variants), and the effect that each site has on the given trait, termed (very poorly) the site's "heritability". It is important to note that every site in the genome is equally and fully heritable. The term refers to the trait, and the size of the effect from variations of that site on that trait. Summed over all sites in the genome and all variations in the population, this heritability ultimately equates to the overall variation of the trait that is genetically caused.

A comparison of two traits, and how they might look in a genetic variation study. In blue is a trait skewed towards small effect variants, with weak selection and consequently variants with higher frequency. In red is a different trait that partakes more from stronger effect variants. On the whole, this kind of difference is not common among complex traits that arise from thousands of loci. MAF = minor allele frequency; Z-score is the score in a GWAS study indicating how statistically significant the variant's effect on the trait is. Note how lower Z-score correlates with more variants at the higher MAF frequencies. At the same time, the GWAS method overall has some skew to higher MAF frequencies, since only those provide sufficient statistical power to get any results at all. Log(s) is the strength of selection; L is the genetic target size for the whole trait, and h*2 is the heritability, or proportion of the trait effect due to the causal variant.

The model they come up with accounts for the selective effect of trait effects (the larger the effect of the variation on traits, the lower its frequency in the population). It also accounts for the fact that variants that affect one trait often affect other traits as well, so the selective effect needs to be considered over all affected traits, most of which are probably unknown, but can be inferred. And conversely, most traits are composed of contributions from many genes and their alleles, sometimes thousands- they are genetically complex traits. 

The model the researchers come up with can normalize among the huge population of variants that affect a single trait, in this case blood pressure. Left shows the effect sizes of individual variants, and right shows the scaled (normalized) version from the paper's model, showing that all these variants follow the same overall rule / logic, using the custom parameters of h*2 and L- trait heritability and mutational target size. All this is to say that the lower effect variants (skewed to left) are assumed to have higher selection coefficients.

The researchers go on to show that various traits do look different under this analysis. Some differ mostly by target size, accounting for more or fewer variants, but having a similar spread of effect sizes over the population. Others differ by the scale of effects that each variant contributes, thus skewing toward higher or lower allele frequencies overall. Interestingly, they add an analysis of the age of these low-frequency, low-effect variants that are the grist for GWAS, finding that they are on average 137,000 years old. That compares with an average age of 600,000 years for variants that are neutral, thus would not come up in GWAS or be under selection. This is fascinating in its implications both for human population genetics in general, and for the fact that most human variation- even that under modest selection- predates the divergence between African and non-African populations. 


Saturday, July 4, 2026

Performing Search, as a Transcription Regulator

Billions of years have created some weird tricks in DNA search.

Search is all around us, as we increasingly rely on search engines to find everything we need on the internet, want to watch, or want to buy. Search looks into databases, which hold the sought-after information. All our accounts, all the domain names, all the products... everything is held in databases of one kind or another, and those databases are indexed in clever ways to provide virtually instant pointers from the question we ask to the answer held online. AI merely puts a linguistic gloss on this, and most people are still encountering AI first as a feature of search, such as the top of current Google search results.

Well, our genomes are databases as well- rich and ancient storehouses of jewels that encode the body and its doings. How does search work there, and what is search even for? At any moment, each cell of the body has certain needs, stresses, and goals, as expressed in its DNA programming. The tools available are proteins and RNAs, which carry out the cell's functions. The needs may arise from signals coming from previously expressed receptors, say, for insulin, which may trigger and tell the cell to take up glucose from the blood. The receptor turns on a kinase, which may turn on another kinase, which turns on a transcription regulator, which goes into the nucleus and ... does a search. This regulator is searching for places (specific sequences) in the genomic DNA where it can bind, after which it helps to turn on (or off) the nearby gene, executing the desired function / tool. 

General introduction to transcription regulators (or "factors") and their role in gene activation and the whole process of gene expression.

Obviously a very different kind of search than what Google does on our behalf across documents, but there are similarities. Internet search depends on patterns, matching the user's input with the vast corpus of the internet also held as text symbols. Transcription regulators match patterns, in this case patterns of DNA that they like to bind, which may occur only once in the genome, or occur tens of thousands of times. The pattern here is a complementary physical/electrochemical shape, rather than an abstract same-symbol match. The genome is, to a protein, truly vast. Our three billion-base genome is forty million times larger than an average regulatory protein of, say, fifty kilodaltons (kDa). Search is also, here, a difficult problem, which researchers have been wondering about for decades. Several recent papers discuss different aspects of the problem and shed some modern light on it.

We have roughly 1600 transcription regulators in our genomes, so there is something going on all the time. DNA is always being queried. And what it replies with is RNA- a transcript issued/copied from a gene, which either goes off to instruct creation of a protein, or is itself functional in some way. So, how do proteins bind to DNA, executing their search? It was transformative when the first atomic structures of such proteins were solved. They were clearly complementary with their DNA targets, with nicely positioned positive charges to mate with the backbone of the DNA and amino acid fingers reaching into the helix to feel the shapes of the nucleotides they wanted to bind. All very neat, and paradigmatic for bacteria whose genomes are quite small. But there is more to the story. Binding sites in human genes tend to be quite short- five to seven bases. That really isn't enough to be very specific, across a vast genome. Eukaryotes have developed several weird tricks, as it were, to encourage efficient search over much larger genomes and at the same time increase precision while maintaining evolvability and flexibility.

Eukaryotes have nucleosomes, chromatin, and packaging. The DNA is not just splayed out randomly, but wound up on protein spools. One would think that this would impair search by regulators. But paradoxically, there is a fine balance between hunting around on a given piece of DNA for a preferred site (one-dimensional search, 1D), and jumping off, letting go, and trying somewhere else (by diffusion; three-dimensional search, 3D). The compaction of genomic DNA into nucleosomes that wind up most of the DNA while leaving linking DNA in between free appears to provide a nice balance of landing spots that allow searching regulators to jump very long distances (in linear DNA terms) while not going very far in absolute terms. Regulators vary in how aggressively they can plow through nucleosomes to try out their internal DNA sites, but many (called pioneer factors) can do so.

Secondly, transcription regulators cooperate with other proteins to create longer, more complex DNA sites for precise gene identification and higher binding affinity. As biologists have characterized the enhancers and promoters of important developmental genes, they have found that DNA binding sites occur in bunches, and have much weaker effects when broken down and separated. Sometimes there is direct side-to-side cooperation between two regulators that bind the DNA. At other times, they combine with other non-DNA-binding proteins to create complexes at such sites. The DNA recognition sequences of these combinatorial sites can be changed significantly, even beyond (our) recognition, by the addition of cooperative proteins. This is something that makes prediction of where a given regulator binds particularly perilous. 

Thirdly, many regulators contain not only DNA binding domains, but also extra disordered domains that facilitate DNA search and binding. This has been a recent realization that accounts for some of the speed and flexibility of regulator search and DNA interaction in eukaryotes. The stable crystal structures of paradigmatic bacterial regulators are not the whole story, and indeed are insufficient to explain what is happening in the much larger setting of our own cells. The authors note that eighty percent of human gene regulators have large disordered domains, (called IDRs, for intrinsically disordered region), upwards of 500 amino acids long. These never showed up in crystal structures, naturally. Being disordered, they are also poorly conserved. So, they have been difficult to study. 

Comparison of binding by one regulator, MSN2, which has a large IDR, to its genomic sites. At top is its native binding pattern, across a whole genome. At bottom are mapped its core motif occurrences on that DNA. Second from top is the MSN2 protein mutated to contain only its core motif-binding domain, and third from top is the MSN2 protein mutated to remove that domain and retain everything else. Note how different the patterns of binding by each of these proteins are, though how each approximates to some degree the wild-type pattern.

In related work, researchers have divided up such proteins into the core binding site part and the IDR part. They find that both parts work partially, directing binding to some of the native sites around the genome. In fact, the IDR part does a more statistically accurate job than the core DNA binding motif. This is fascinating, showing that in eukaryotes, a new search mechanism arose, supplementing discrete and precise binding with a floppy / fuzzy code in the IDR and its binding sites. It turns out that regions of hundreds of bases around core target sites (which in one case amount to only the motif AGGGG) are preferentially bound by the respective IDR protein domain, with multiple weak interactions that remain structurally uncharacterized. In fact, neither the protein structures responsible, nor the DNA sites they bind are known yet, though deletion studies through IDR domains show that binding is distributed throughout.

Relationship between IDR binding site size, and the ratio of 1D vs 3D search time, by simulation. The bottom axis is size of the IDR binding region, the Y axis is time taken for search. Time spent in total (yellow) goes down to minimum at an optimum between 1D search that is slowed by longer IDR-binding regions, while 3D search is strongly accelerated by longer IDR-binding regions.

The combination of core binding and loosely unstructured binding in one regulatory / search protein provides powerful benefits. In dimensionality terms, if the effective landing site is expanded from five to five hundred bases, then the time required for 3D search through the space of the nucleus is dramatically shortened. Secondly, loose binding by the IDR then promotes an "octopus"-like 1D search along the local DNA, resulting in efficient settling on the core binding site to get ultimately precise positioning. The ultimate affinity of the regulator with the local DNA is also enhanced compared to what it could manage over a five base pair site. The researchers conclude that with these domains, the search problem is, in net terms, reduced by one dimension, from 3D to 2D. The surrounding areas of DNA that have marginal affinity for the IDR domain are called "antenna" regions, and the author's simulations show how they alter search behavior.


Schematic explanation of the current work, describing how IDR domains help to speed up the transition from 3D search through space, to 1D search across the DNA. And then also to facilitate 1D search by preventing full detachment from the DNA while the core binding motif continues to search by diffusion for its binding site (yellow).

For computers and databases, search is a huge problem that has led to technical innovation, as well as large drains on resources. Every search engine combs the internet, gobbling up all available information, creating indexes, and updating them constantly in order to give us the instant access we want. This infrastructure has been raised to a new level by AI, which transforms search into a new form, combining it with language translation and prediction methods that allow a search for corkscrew to bring back results for wine. Whether it understands anything is unlikely, but the desire to upgrade search from a simply determinative process to one that is more fuzzy and richly interpretive, and thus more useful, is not a new phenomenon.


Saturday, June 20, 2026

From Icebox to Hothouse, and Back Again

Better modeling, by including the biosphere, retrodicts more of Earth's dynamic climate history.

Climate change, while ignored by the current administration, is not ignoring us. The Earth is warming well past where it has been for millions of years. But before that? While the planet has generally had stable climates, they have varied substantially through time, and have gone through occasional catastrophes. There was a little ice age, in the middle of the last millennium, thought to have been caused in part by the depopulation of the Americas due to European diseases. The ensuing regrowth of forests covering the Americas drew down CO2 from the atmosphere and cooled the climate. But more to the point, there have been far more severe episodes, both of heat (the end-Permian extinction event) and cold (the Sturtian glaciation of the Precambrian). All of these arise from CO2 levels, as CO2 is the master controller of heat in the atmosphere, thanks to the greenhouse effect. (As it is on Venus as well.) 

For example, the end-Permian extinction is thought to have been caused by unusual volcanism in what is now Siberia. Over a mere 100,000 years, this poured an estimated 26,000 petagrams of CO2 into the atmosphere, causing its concentration to shoot up to about 2500 ppm (parts per million) and temperatures to shoot up as well, killing off 90% of all species. What we are doing now is much faster, though admittedly in early days. We are pouring roughly 11 petagrams of CO2 into the atmosphere yearly, which has raised CO2 concentrations from a preindustrial 280 ppm to 427 ppm today. It would take us another one to two thousand years to cause a 90% extinction event!

A bedrock of our climate thermostat is the silicate cycle. Since the vast majority of carbon on earth is locked up in rocks, (carbonates of silicon, magnesium, and calcium), not in the biosphere, it is rocks that have a dominant effect. Volcanoes belch out CO2 in huge amounts. That CO2 slowly eats away at rocks that are exposed, re-forming carbonate compounds that are weathered off and back into the ocean. Where these compounds (with those built by shelled animals of all kinds) are gradually deposited on the sea floor and subducted back into the Earth's crust. Some of those carbonates are reduced at depth and brought forth again by volcanic activity. The more CO2 there is in the atmosphere, and the warmer it is, the more weathering happens and thus the faster CO2 levels are brought back down. That is the elegant thermostat that has kept Earth at mostly mild temperatures through its long history. 

However, this is a slow thermostat, taking hundreds of thousands of years to equilibrate. Unusual events, like an asteroid impact, prodigious volcanism, or the advent of human ingenuity, can make a mess of things way faster than the silicate cycle can deal with in its slow, grinding way. Many subtler influences can also come into play, like cycles in the tilt of the Earth towards the Sun, or continental arrangements that lead to particular patterns of ocean circulation, can create variations such as ice ages. A recent paper brought out peculiar influences from the biosphere that can also affect, and even destabilize, the thermostat on longer time horizons

The oceans are responsible for roughly half of photosynthetic productivity, and they are also where the carbonate minerals get buried. So how they react to changes in the atmosphere are very influential in the whole cycle. These authors ran half-million-year simulations of climate perturbations while including not only the silicate cycle, but also reactions by the biosphere and especially the phosphorous cycle, which has a strong influence on biological productivity. It turns out that when the atmosphere has lower levels of oxygen than we do today, (as was the case during the Precambrian epoch), high CO2 levels cause long-term rises in biosphere productivity and also in phosphate recycling out of the ocean floor. The extra phosphate increases biological productivity even more, and thus causes CO2 drawdown to persist past where the silicate cycle would level out for the long term. The result can be a rebounding ice age after a hot phase. 

Model results over 500,000 years, showing rebound from an injection of high CO2 at year 10,000. A shows concentrations of CO2 over time, B shows O2 concentration, and C shows sea ice, which goes to zero at first, the rebounds sharply, especially under the blue condition of 0.6 times current oxygen concentration in the atmosphere. It shows how exquisitely sensitive the climate is to CO2.

These models make some sense of the Precambrian climate cycles, which had a few dramatic swings that went through so-called snowball Earth phases where the entire surface of the planet seems to have iced over. The silicate cycle naturally came to the rescue eventually, spewing enough CO2 from volcanoes to overcome the snow / albedo effects of all the ice and cause a rebound hot phase. Between the rising oxygen levels and the extreme climatic swings, the stage was somehow set for the rise of animal life, leading the so-called Cambrian explosion, though there was a fair amount of simpler precursor animal live in the Precambrian.dd

https://www.science.org/doi/10.1126/science.adh7730

A schematic of the proposed cycle, with CO2 coming in from vulcanism (red) and being disposed of by various means, first and foremost the silicate cycle (blue). OC = organic carbon, P = phosphorous/phosphate, OCpetro = organic carbon weathered out of sediments, coal, limestone, and other geologic formations. Thus, the brown color shows this paper's additions to the classical silicate cycle.

While it is just a modeling paper, models are what we think and do in science. It is nice to have laboratory confirmation for areas of science (like molecular biology) that permit it, but historical sciences, especially those pertaining to whole planets as systems, have to be more forensic and speculative. This new model is a refinement on the basic silicate cycle, and thus seems a strong improvement on what has heretofore been a science of more or less back-of-the-envelope estimation. And judging from this new model, the authors propose that the next ice age is not being put off indefinitely by our profligate emissions, but rather that organic burial feedbacks will bring it closer (than 400k years away) with additional overcooling thereafter!


  • Medicine is toast. "MIRA outperformed physicians in diagnostic accuracy and made guideline-concordant, medication-safe and appropriate admission decisions."
  • A death sentence for US science.
  • Revaluing trash.
  • Apparently, fungi in the ocean are a thing.

Sunday, June 7, 2026

Strides in Cancer Treatment

A new paper shows that CART therapies can be unleashed against solid tumors.

We are finally in the payoff period in the decades-long war on cancer. Slowly, painfully, precision approaches are being developed to treat specific molecular lesions in ways that are superior to the old blunderbuss kill-everything approaches. At first, these treatments had only marginal effects, at astounding costs. But increasingly, the effects are lengthening and cures are in sight in some forms of cancer. One unexpected area of revolutionary progress has been immunotherapies, which in various ways help our immune systems attack cancers. It turns out that many cancers have tricks to hide from the immune system, and once those tricky dampening molecules are circumvented, dramatic reductions are possible. One paper recently described an anticancer vaccine made up of a witch's brew of targeting molecules, cancer antigens, and adjuvants, that achieves strong anti-melanoma action.

Another one of these immunotherapies is CART, or chimeric antigen receptor T-cell therapy. T cells have a receptor repertoire, just as B-cells do, which target things to be attacked- foreign pathogens, diseased states, etc. at molecules called antigens. One problem in cancer is that the cells are, originally at least, our own, so they mostly evade immune detection by having few "foreign" antigens. But there are nevertheless some antigens, comprised of normal molecules that are out of place (such as DNA found outside the cell) and "neoantigens" that are proteins expressed from the mutations in cancer cells. Additionally, as mentioned above, cancer cells express additional molecules (PD-L1) that can dampen even the immune response that does get generated by these few cancer antigens. So, the chimeric part of CART is taking the patient's own T-cells and engineering some of them to express new anti-antigen receptors that are relevant to the patient's cancer. Perhaps there is a mutant fusion protein that the cancer depends on. Perhaps the cancer displays an unusual surface molecule. Perhaps the tables need to be turned and PD-L1 targeted. There are many possible targets. 

CART therapies have, to date, been mostly directed at blood tumors. Solid tumors have extra protection in their micro-environments, and have not been good targets, though they necessarily have blood supplies and thus exposure to systemic T-cells. A recent paper blows open this field by revealing a magic molecule that plays a very significant role in the structure of solid tumors- the urokinase receptor. The urokinase plasminogen activator receptor (uPAR) is heavily expressed on senescent cells and many solid tumors, but rarely expressed elsewhere. Indeed, its expression correlates with tumor aggressiveness. Plasminogen is a protease that is sort of a cleanup crew for the circulatory system and body generally. It breaks up blood clots, and digests follicle tissues allowing ovulation. It encourages wound healing and discourages fibrosis- the buildup of scar tissue. However, in the cancer setting, the same activity seems to encourage fibrosis in a sort of constant wound healing state. Reviews in this field are rather confused about the direction of action. But one thing is clear- uPAR has myriad signaling activities relevant to tissue repair and immune activation that are not all dependent on the uPA (plasminogen activator) and plasmin activation system. Indeed, it is expressed not just in cancers, but in many other fibrotic settings.

A wide array of proteins are assessed here for their expression in a cancer tissue sample. uPAR is in red at the upper left. The matrix on the right shows the correlation of expression in a wide variety of cell types and tissues, like cancer-associated fibroblasts (CAFs), monocytes/macrophages (Mo/Mac), and with the protein fibroblast activation protein alpha (FAP).

The authors sought to target CART cells against uPAR, principally as a targeting device, since this marks many solid tumors and correlates with metastasis and rapid cancer progression, in addition to inflammation and fibrosis. While only tested in mice, the results were remarkable. 

"CAR T cells targeting the D2-D3 domain of uPAR display broad antitumor activity in xenograft, syngeneic, and patient-derived models, including in adjuvant and combination settings, supporting the concept that targeting conserved malignant cell states can enable therapeutic strategies that transcend tumor type. ... our prior work shows that uPAR CAR T cells targeting senescent cells remodel fibrotic tissues, and, as shown herein, this remodeling is associated with CAR T cell infiltration and cytotoxic activity. Similarly, parallel work demonstrates that uPAR CAR T cells exhibit potent efficacy in glioblastoma models and can co-target supportive stromal cells."

This is to say that these CART cells target not only tumor cells, but the surrounding solid tissues (stromal cells) that they rely on. That is the key to defeating solid tumors. It also indicates that other autoimmune and fibrotic conditions may be addressable with this therapy as well. 


Treatment effects from the CART therapy in mice, against several tumors. The red graphs are controls, and the blue graphs are treatments. Top is the tumor volume over time, while at bottom is survival of the mice over time.  Lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), high-grade serous ovarian carcinoma (HGSOC), and pancreatic ductal adenocarcinoma (PDAC).

The results of treatment of xenografted human ovarian tumors into susceptible mice, at 3 weeks, bottom. On the left is the control, while the other two sets were treated with CART cells against uPAR.


The authors note that relapses were seen occasionally, but that in these cases, the uPAR target was still highly expressed. That suggests firstly that it is difficult for tumors of these targeted types to do without uPAR, and secondly that something else went wrong with the tailored CART therapy, other than that its target went away. Perhaps future work can enhance its penetration or activity. The researchers also strained to make their model systems as human-relevant as possible, using cancer tissue transplanted (xenografted) from human cell lines, human CART cells, and mice with transplanted immune systems from humans. This work is thus not only a scientific breakthrough of the highest order, but is a technical tour de force as well. It also ends up with a variety of patent declarations and commercial ties, indicating that this breakthrough is being fully milked by its inventors and commercialized at breakneck speed.

One major problem with this mode of therapy is that CART cells require a great deal of engineering. First, antibodies against uPAR were developed in mice or other species. Then the genes from those immune systems were recovered from those mice, to get the precisely recombined gene that expressed the antibody with highest binding activity against uPAR. Then that gene, hooked up to new transmembrane and intracellular domains, (specially selected to activate the T cell they will be put into), was introduced into a transformation vector and put into the T cells collected from the diseased mice. In humans, this treatment routinely runs a half a million dollars. It is incredibly ornate, and one expects that gene therapy will someday allow the patient's T cells to be directly modified in the body, without all the collection and laboratory work, (which takes months), given a high-quality gene encoding the antibody fragment that is generally applicable- not tailored to a specific patient.


  • The example of Spain.
  • AI does insurance... as you would expect!
  • AI is not what it is cracked up to be. And way more expensive than it has to be.
  • Japan is surprisingly willing to keep importing fossil fuels, despite exchange rate degradation.
  • Renewables and batteries have stabilized California's grid and made electricity cheaper.
  • Map of where electricity has gotten more expensive in the last year.
  • How other animals deal with inequality.
  • What economic warfare looks like.
  • Can we survive the internet?
  • Electric cars are good for everyone.

Sunday, April 26, 2026

The History and Future of a Single Mutation

The CCR5delta32 confers resistance to HIV. Where did it come from?

We are edging into an age of precision medicine, where the causes of our maladies will be known in molecular detail, allowing treatments that address them at the root. Given the parlous state of medicine today, in the midst of financial breakdown and a continued mediocre level of basic diagnosis, it is hard to believe this is a corner we can turn. But vaccines have long been in this category, of addressing the precise pathogenic causes of disease, and oncology is fitfully getting there, given advances in DNA sequencing and in treatments based on specific mutations.

HIV is also a beneficiary of this approach, since the discovery of its pathogen led directly to a variety of effective (if not yet permanent) treatments. A researcher in China created gene-edited humans with a specific mutation that will render them resistant to HIV. The mutation he chose for this work is called CCR5delta32, and it does not naturally exist in Chinese populations. 

But it does exist in European populations, at a roughly 10% rate in single copy. When present in two copies, it provides complete immunity to HIV, while if present in one copy, it slows infection substantially. A recent paper rooted through the available ancient and present genomes to figure out where this mutation came from. 

CCR5 is a cell surface receptor for cytokines 3, 4, and 5. These are all pro-inflammatory cytokines, and they interact with multiple receptors. Here, as in so many other respects, the immune system is riven with redundancy, so that it can grapple with as many contingencies as possible. Cytokines are signaling molecules for the immune system, which is an unusual organ, being dispersed all over the body with numerous cell types all patrolling around, and communicating with each other by long- and short-range chemical messages. It turns out that the major form of HIV uses the CCR5 protein to get into our immune cells, explaining why CCR5delta32, which is totally non-functional, has such a dramatic effect on HIV susceptibility. 

While people carrying CCR5delta32 are generally fine, this defect does confer a variety of subtle changes to their susceptibility to other infectious diseases and cancers. That explains why this mutation has settled at its low level in the European populations, probably balancing the occasional benefit against a specifically CCR5-seeking pathogen against its natural functions that form the basis of its existence in the first place as a part of immune system that is conserved in all mammals. The Chinese gene-editing researcher came under withering criticism not only for breaching the generally agreed moratorium on human germline gene editing, but also because the net effect of this mutation is, on the whole, negative, raising risks of numerous diseases, despite its beneficial effect on HIV. 

The authors run several models and populations in an attempt to time the origin of the CCR5delta32 mutation, and portray its positive selection over the ensuing millenia.  CHG- Caucasus hunter-gatherer; EHG- Eastern hunter-gatherer; WHG- Western hunter-gatherer; ANA- Anatolian Neolithic ancestry. The bottom axis is time, and the Y axis is the frequency of the mutation in these populations. "Modern DAF" refers to the inclusion of the data set of current (not ancient) population frequencies, (top), which the authors claim leads to continued rates of selection (last 2,000 years) that are artifactual.

So where did it come from? The new authors gather up a large variety of population samples from around the world, and from ancient humans, back to about ten thousand years ago. They find the first instance of the mutation in one sample at 5.8 thousand years ago. After that, its frequency rises dramatically, up to about two thousand years ago, when it levels off. They conclude that this mutation originated about seven to nine thousand years ago, in the steppes of Eastern Europe / Western Asia, and was under strong positive selection at first, spreading to the current frequency of about 10% of the population / alleles. All occurrences on other continents can be accounted by the spread from this source.

Does this mean that HIV was prevalent long, long before the current pandemic? Hardly. The authors can not say anything about it, but one theory would be that some other disease had a similar profile. It certainly was not the Black Death, as the authors show that this mutation had no change in frequency over that gruesome pandemic. Another hypothesis is that general reduction in inflammatory response might be beneficial in some settings, as has been found for Covid-19, though here again, this mutation does not have any known positive or negative net effect on Covid-19 susceptibility or course. 

It is amazing that we have enough sequences of ancient DNA to be able to reconstruct this kind of thing- to be able to trace where and when some influential mutation occurred, and how it traveled and spread. It is a tour-de-force of bio-archeological reconstruction.


  • When you escape reality, and morality.
  • Some environmental benefits are flowing from the current war.
  • We may be at peak oil, courtesy of the US.

Sunday, April 19, 2026

The Death of Boredom and the Future of Politics

Can politics work without a civic sphere?

How can we have a loneliness epidemic when we are connected like never before? It is a problem that perplexed Robert Putnam in "Bowling Alone". He put it mostly down to TV, internet, and the growth of passive and isolated forms of entertainment generally. When you read between the lines of history of any time before about one hundred years ago, you realize that people were, before the modern age, bored out of their minds. Who plays cards? Who puts on operas, or runs numbers, or goes bowling? Who needs an Easter pageant, or a three-to-four-hour baseball game? Only people with nothing better to do. If you wanted music, you had to make it. If you wanted conversation, you had to share it. Human society was built on simple quid pro quos- social rewards and resolution of boredom and isolation for personal participation.

But that deal has broken down dramatically in the modern age. We have a thousand channels, talk radio, recorded music. With AI, we are getting personal chatbots and bespoke romantasy partners. Sports have slid tectonically from participation to spectation. Boredom is a thing of the past, though if you do want to play cards, plenty of computers are willing to take a hand.

An interesting article in the New Republic knit this together very nicely with the problems we are having in politics. In the US, political engagement is increasingly shallow, leaving the field to extremists who can still call up foot soldiers to storm the ramparts. What happened to the Occupy movement? For all its inherent logic and flash organization, it fizzled into nothing because it gave little thought to its own institutionalization (indeed, was allergic to organization) and durable engagement, all the while railing against the overwhelming organization and deep pockets of the entrenched systems of capitalism. The Left is notoriously inable to herd itself into an effective, organized force. While capitalism is naturally organized and institutionalized by virtue of naked self-interest and corporate structures, civic groups grow out of far more disparate, and evanescent, motivations. Unions have been an attempt to organize around a countervailing, while still self-interested logic, which inherently limits their reach and coherence. The true civic sphere, however, is threadbare.

Political parties have similarly shallow roots. In California, the governor's race has 61 candidates, and little control by the party establishment, particularly by the Democratic establishment that supposedly runs the state. Like other non-profits, parties ask little of their adherents, other than possibly a monetary contribution, and wouldn't dream of holding truly social events that could deepen civic engagement. Expectations of civic engagement have hit rock bottom, mostly because people have tuned out across the civic spectrum. The testimonial dinner is a relic. The ice cream social is unheard of. Service organizations like Rotary and Elks are fossils, unions are on life support. Events and organizations that previously kept people entertained and involved in a civic way are scarce. These traditions both trained people for common action, and led to the kind of networking and contact that fed political consciousness and activity. They also helped to vet people directly for office holding (see the recent Swalwell case). 

Bernie Sanders can draw a crowd, but do those crowds go out, organize, and persist?

Republicans have found a partial solution to these problems by ginning up endless outrage through their propaganda outlets, predominantly talk radio and hate TV. While motivating, the results have, naturally, been intellectually disastrous and have us teetering on the edge of fascism. Democrats, as the more level-headed and progressive temperament, have not used the same tools effectively, and shouldn't. What should they do? Well, the field for civic engagement is pretty wide open. For example, one could imagine a tax on political advertisements, say 10%, which is collected by the government / FEC, and sent to counties or municipalities for civic engagement purposes, either election-related or not. This would create a fund for local talks, events, civic education, and the like that would, in theory, complement the advertising that is increasingly vacuous and meretricious. 

Another approach is direct action, where Democrats could use some of their energy and resources to build civic engagement, outside of straight campaigns. Just as the Republicans have harnessed ancillary issues like abortion and tax cuts that energized specific segments of their base, Democrats have to be a bit more canny about asking for more engagement and offering more involvement. Climate change is a great example, where a wide spectrum of individual action (trash pickups, solar panel installation, water quality testing) could be integrated into civic engagement that builds party alignment and ultimately, institutional strength. All great religions know that the more you ask, the more you get, and the deeper the commitment of followers. Additionally, the left already has a bewildering array of non-profits, whose efforts would ideally be more closely integrated with the Democratic umbrella to generate more organizational power- synergy or leverage, in business-speak.

On the other hand, how could civic disengagement be accommodated rather than fought? One approach might be to enhance the vetting and exposure of candidates by having nominating conventions at the local level. Even though California has an open primary, and thus does not grant each party automatic spots on each ticket, the parties should not shy away from selecting, testing, and promoting candidates. This should not be a central commitee operation hidden in the dark, the province of interested apparatchiks, but open forums that promote philosophies as well as people.

We are in a tough position, trying to keep politics alive in a world where its underpinnings- of civic engagement, communal organization and leadership, and simple conviviality- are fading in a deluge of individualized enjoyments. Political parties are at the forefront of this change, and need to think very deeply about how to keep themselves relevant and effective.


Saturday, April 11, 2026

Pumping Calcium

An ornate ion pump manages rapid outflow of calcium.

In the beginning, the egg cell experienced a wave of calcium release, triggered by union with a sperm cell. This blocked other sperm from entering, and prepared the egg to become a zygote and embark on embryogenesis. It is but one example of the pervasive role of calcium signaling among animals. Another is the muscle activation cycle, which relies on calcium release from the specialized sarcoplasmic reticulum (in response to a nerve activation) to get the cell as a whole contracting. Generally, calcium is kept very low in the cytoplasm, and high in the endoplasmic reticulum and outside the cell. Thus, channels gated by electrical activation or other signals can cause rapid cytoplasmic calcium spikes and signal widely within a cell. 

On the flip side, there have to be pumps that keep the cytoplasmic concentration low, and a recent paper elucidates the structure of one such pump that is remarkably fast, while also closely regulated. It is an impressive machine. PMCA2 is an ATP-using calcium pump that sits in the plasma membrane and carries out what is called the Post-Albers cycle. This is a flip-switch mechanism for pumping ions, where ATP drives conformational switches alternately exposing ion binding sites to each side of the membrane. When the pore is open to the cytoplasm, there is no competition from higher concentrations outside, so the active site can bind one internal calcium, given a high-affinity site. Then, after the conformational switch, the pore is exposed to the outside, and at the same time the site is reconfigured to be lower-affinity, releasing the calcium ion into a high concentration environment. Neurons especially use calcium signaling extensively to operate synapses and regulate growth and development. Their rapid and frequent signaling requires a pump that has especially high capacity. PMCA2 operates at a maximal rate of several thousands of Ca2+ ions pumped per second.

Cartoon of the Post-Albers cycle, which is shared by a large family of active ATP-using pumps that transfer ions against their chemical concentration gradient. M is the main transmembrane domain of the pump, where the ions traverse the membrane. The N, P, and A domains are regulatory, especially binding and cleaving ATP  at an interface between the N, P, and A domain. The cycle links power steps (1,2) with conformational changes that carefully gate the pumping process.

And that is not all. Since calcium has a charge of 2+ and this pump does not intend to alter charge across the membrane, the pump simultaneously has binding sites for counter-ions (generally two OH-) that are transferred in the opposite direction from the calcium. Not only that, but every pump of this kind requires regulation of various kinds. PMAC2 is activated by phosphatidyl inositol 4,5 bisphosphate (PIP2), which is another important signaling molecule generated by specific PI kinases in response to activation of G-protein coupled receptors or protein kinase C, which may respond to external signals. In very general terms, these tend to be pro-growth or stress-induced pathways. These regulatory processes can tune the overall rate of recovery from rapid Ca2+ signaling events, by adjusting the level and activity of pumps like PMAC2. 

ATP binds at the N/P/A domain interface, and its hydrolysis (and loss of ADP) generates extensive shape changes, including into the transmembrane M domain. At the very bottom, the calcium ion is shown in green, bound inside the M domain pumping channel. The motions here are subtle, but enough to dramatically reshape the calcium channel.

The authors, using various substrate variants and other tricks, were able to develop structures of PMAC2 in several steps of the pumping cycle, using cryo-electron microscopy. The ATPase site in the N domain (red) is far from the channel that conducts the calcium ion (brown, far bottom). They show extensive shape changes from binding or losing the ATP molecule, though they mostly concern the intracellular domains (red, blue, yellow). The effects on the transmembrane pore domain are rather subtle, shown on right. The authors claim that, compared to other pumps of this large family, the structural changes are significantly less, suggesting that evolution for speed has caused the mechanism to become more efficient, with less wasted motion per conduction event, at least in the channel region itself.

Relation of the PIP2 binding domain (orange/red stick figures) to the calcium core binding site. PIP2 appears to be essential for rapid pump operation. At bottom is shown some schematics of the gating provided by PIP2 in bound and unbound states, especially via the D873 side chain (negatively charged aspartic acid).


They also find that the activating molecule PIP2 is neatly parked right next to the main calcium binding and conduction region, and is more or less essential for enzyme activity. In the graph above (e), they show that several single mutations made in the calcium binding high affinity site, for example switching the negatively charged D873 for the positively charged K (lysine), kills ion pumping activity. Mutation of the PIP2 binding pocket (KKQ->TLL, around position 347) likewise kills enzyme activity.

Relation of the counter-ion channel (red dots) with the calcium channel. Both are essential parts of the mechanism. Closeups with the coordinating protein side chains shown on the right.

The whole mechanism is alluded to in the last figure, where the central calcium binding site is shown, with the general direction of calcium pumping. The counter-ion transport area is shown nearby as a flurry of red dots (standing for water molecules, which at this scale are interchangeable with OH ions). Specific single mutations in either area, either changing negatively charged E412 to positively charged lysine at the calcium binding pocket, or changing polar S877 in the water/hydroxy binding area to the bulky and hydrophobic F (phenylalanine), each kill pumping activity (graph). 

While it would be ideal to have a more dynamic representation of what is going on, the new structures give tremendous detail, including the associated ATP, PIP2, calcium, and water molecules. The mutations also nail down several functional points. Obviously a rather intricate and well-oiled machine that keeps its bit of cellular calcium homeostasis on an even keel. It is hard to believe that the sum of thousands of machines like this one is life, but the deeper we look the more true that appears to be.