Showing posts with label genetics. Show all posts
Showing posts with label genetics. Show all posts

Saturday, August 1, 2026

Ohnologs and Paralogs: The Wages of Gene Duplication

Two whole genome duplications lie at the root of vertebrate evolution.

Another week, another story about the power of gene duplication in evolution. While rare during normal reproduction, gene duplication happens pretty frequently over longer time scales. How else would we get a thousand olfactory receptor genes, all similar to each other? But other accidents can occur as well, like whole chromosome duplications (such as what leads to Down syndrome), and whole genome duplication, when the cell division process stops early, but otherwise proceeds with cells remaining viable with double the genomes as before. This is common in plants. Corn is tetraploid, wheat is hexaploid, and strawberries are octaploid. 

But among animals, whole genome duplication is less common. Two decades ago, however, two researchers working from the newly sequenced human genome revealed that there were two such duplication events at the beginning of vertebrate evolution, explaining some oddities and also perhaps the speed and power of subsequent evolution. The basic evidence is the genome sequence, which is full of related genes. Genes that do the same thing in various species, and are lineally related, are called orthologs. That is relatively simple, per the Darwinian tree of descent of all life. Genes within one organism / one genome that are similar to each other due to ancient duplication events are called paralogs. All those olfactory receptors are paralogs, for instance. Lastly, genes that are paralogs stemming from a whole-genome duplication event are called ohnologs, in honor of Susumu Ohno, who led the field of molecular evolution in recognizing the importance of gene duplication, and speculated about whole genome duplication well before it was discovered.

A classic example of this evidence is the hox cluster, a linear sequence of genes that have been extensively studied in flies as providing an important set of regulatory controls over the linear body plan. They lie in the middle of the developmental cascade, downstream of egg and body polarity genes, but upstream of specific appendage and tissue expression programs. They encode DNA-binding (homeobox) transcription regulators, and their position in the genome is co-linear with the body parts they activate because there is a progressive chromatin opening process by which this whole locus becomes activated. Well, flies have one hox locus, but vertebrates have four. What happened?

Hox loci across evolution. Where flies and primitive chordates have one hox cluster, vertebrates have four. How did that happen?

Obviously, once researchers lined everything up, it became pretty clear that there were two massive duplications along the way, creating four hox loci in the genome, after which quite a few of the duplicated genes fell away. After a duplication event, gene survival is a race between neo-functionalization (which leads to preservation by selection) and deleterious mutation, degradation, and disposal. Enough of these ohnologs survived to help fuel the substantially greater complexity of the vertebrate body plan, now including intricate wrists and hands, and ever more involved head structures.

Similar findings were made all over the human genome. The original paper has a graph that shows that, across the genome, most paralogs exist in families of four. If genome duplication were not the applicable hypothesis, then one would expect a smooth asymptotic curve downwards from one member (implicit) to two, then three, and fewer from there outwards, since single gene duplications would each be independent statistical events. But no, there is a peak at four, indicating that some process yoked together many, many genes into parallel duplication events, twice in succession. 

A graph of count of paralogs vs their frequency by count, in the human genome. What should in principle have been a smoothly declining curve from 1 or 2 turned out to have a weird peak at the number four. This was important evidence that many of our paralogs arose through some common event, such as a pair of ancient whole genome duplications. 

OK, so far, so old hat. A bunch of more recent papers flesh out this story a bit, showing how the vertebrate duplications were timed, and how they affected various types of genes. miRNAs, for instance, turn out to be ancient genetic elements, and were duplicated along with everything else. miRNAs that originate from (and were preserved from) the whole genome duplications tend to be more conserved, have more targets than other miRNAs, are more highly expressed, and have higher rates of targeting RNAs of transcription regulatory proteins that likewise originated from whole-genome duplication. The traces of this history are thus interestingly preserved.

Another paper traced the effects of the vertebrate ohnologs on brain development. Using the pre-vertebrate amphioxus as an out-group for comparison, they find that the ohnologs that date from the whole duplications are more highly associated with brain cell types and their developmental programs than are other gene duplications- either ones dating from the same time as the whole genome duplications, or since.

"Compared to their closest invertebrate relatives—tunicates and amphioxus—vertebrate brains are highly regionalized and complex."

"Ohnologues were enriched in development, cell-fate commitment, signaling and neurotransmitter transport. By contrast, SSD [small-scale duplication] paralogues were enriched for immune response and sensory perception in all species, a result that matched previous reports."

Lastly, a paper focusing on intermediate genomes, part of a plethora of genome sequencing that has happened since the original analysis, reveals exactly what happened at this time, about 500 million years ago. From a molecular perspective, hagfish are in the same group as lampreys- both parasitic eels that lack real jaws but arise from the vertebrate lineage. They are cyclostomes, while we are gnathostomes. After the first whole genome duplication about 530 million years ago in the common stem lineage, the gnathostomes and the cyclostomes split and each experienced their own, separate genome duplications roughly 490 million years ago. This can be concluded from the differing gene collections that survived from each respective duplication, and their sequences relative to the cyclostome/gnathostome split. Indeed, lampreys and hagfish have six hox clusters instead of four, indicating that a triplication event happened, instead of duplication event, between their split from gnathostomes and the split between lamprey and hagfish. 

A phylogenetic tree with dates, locating the genome duplications deep in the vertebrate stem lineages. 1R is the first genome duplications, common to all vertebrates. 2R is the duplication in the lineage leading to jawed vertebrates, while CR is the apparent triplication that happened in the lineage leading to cyclostomes. Apologies for the antiquated human being. 

All this genome renovation didn't benefit cyclostomes the way it did gnathostomes, however. Only one group went on to globe-straddling glory, indicating that while genome duplications may be helpful for evolutionary / developmental innovation, they certainly aren't determinative. They are grist for an evolutionary process that remains largely shrouded in mystery- the context and environments of the time, the competitive biosphere, and the internal molecular environment. There is no reason to think that all this is unaccountable by natural processes, but that doesn't mean we have the information to reconstruct it in detail. We should thus be deeply grateful that the scientific community can dig up even this much of our history.


  • But how can we take more money away from workers?
  • The smearing of Anthony Fauci. And of basic reason.
  • Reward and addiction.
  • Still stupid ... the Discovery Institute and "information".

Saturday, July 18, 2026

MHC Through Evolution: Breaking All the Rules

The immunologically critical MHC gene cluster plays by its own rules through the arms race of life.

Last week, I discussed the very general landscape of variation in the human genome, specifically the tradeoff between prevalence in the population and effect size. Given that the vast majority of variants are deleterious, those that affect our phenotypic traits are more heavily selected against the greater effect they have. The result is that at a gross level, most genes and most traits share the same general distribution of lots of variants (alleles) with minor effects, and far fewer with large effects. And those with minor effects also turn out to be tangential for biologists, rarely informative about the nature of the traits they (sort-of) affect.

This week, another paper and another view of evolution, though the eyes of one the more critical genes of the immune system, the major histocompatibility complex, or MHC group of genes. While our adaptive immune system has developed the extraordinary and powerful ability (though semi-controlled DNA recombination of the antibody and TCR genes) to recognize practically any antigen, foreign or domestic, that system requires stringent controls. One of those controls on T cells, which carry the antigen-recognizing TCR receptor, is that it can only see antigens that are "presented" on MHC molecules. MHC proteins have a surface cleft that gets loaded with and holds small peptides (8 to 12 amino acids long) that are cleaved from other proteins, either from pathogens or from the cell itself. The MHC+peptide complex then sits on the surface of the cell, announcing either that 1) I am healthy, full of normal cell proteins, going about their business, nevermind, or 2) I have some other proteins inside, either from a viral infection (class I MHC) or from some bacterium I have just phagocytosed to deal with an infection (class II MHC). In the second case, T cells carrying the TCR receptor lock onto the MHC+peptide complex, and start up the process of killing that cell. 

MHC proteins (beige, pink) present an antigen (red) from one cell, and dock to a T-cell which recognizes the antigen+MHC complex by shape, using the TCR receptor. The additional CD8 or CD4 receptors help to verify the proper binding. While class I MHC are present on all cells, class II MHC are on phagocytosing cells like dendritic cells and macrophages that commonly ingest, reprocess, and re-display bits of encountered pathogens.  


Given that the TCR gene recombination process is unbiassed and produces a galaxy of random binding specificities, how do these cells distinguish self from non-self antigens? This is a deep question that is not fully resolved. But one major mechanism is thymic selection, which is what gives T cells their name. Special cells in the thymus display a wide range of self-antigens, and T cells, which are obliged to pass through the thymus during their maturation, are induced to commit suicide if they react to any of them. A paper from 2018 fascinatingly discussed how it is possible to create a T cell population that knows the "language" of foreign vs domestic after what is known to be a rather haphazard selection process, which displays only a partial range of self-antigens, and leaves quite a few self-reactive T cells around.

At any rate, the MHC proteins do not benefit from hyper-variation provided by genetic recombination. Yet it turns out that variation is beneficial here as well. The way foreign antigen peptides nestle in the MHC groove can be varied by mutations in the MHC molecule, providing a rich field of variation in antigen recognition and thus disease resistance. So, our MHC genes have not just a few alleles in the population, not just a few dozen, but over six thousand alleles. For each individual MHC protein, each person has only two, but over the population, there myriads with different properties. MHC was first recognized for its role in self vs non-self recognition and transplant rejection, (thus the "compatibility" in its name), and it quickly became evident that people vary tremendously in their MHC complement. And this variation plays a big role in keeping us (and all other animals) going as populations, in the face of pathogens that evolve a lot faster than we do. 

A recent paper provided a phylogenetic history of MHC molecules in monkeys, covering the last sixty million years of evolution in our lineage. It is a festival of gene birth, death, and duplication, quite apart from the smaller mutations that are constantly accumulating and cycling through the population. The MHC region carries about 200 related genes, most of which have minor roles, and only six of which (three MHC class I, and three MHC class II) follow the high-mutation pattern because they encode the main antigen presenting proteins. These genes are subject to, quite obviously, unique selective forces. 

How the MHC gene cluster looks, when aligned and identified by gene, over the primates. Note the deep divergence between the new- and old-world primates. Genes A, B, and C are the major MHC class I genes, which vary the most over this time. Note also how some of these genes have gone through extensive duplication in some old-world monkey lineages.

The first force is balancing selection. As soon as one allele becomes common, pathogens evolve to evade its presentation skills, rendering it less effective and less desirable. This results in a population full of minor variants. Indeed, for any individual person, having two MHC molecules that are the same would be bad. Being heterozygous at these genes is highly advantageous, thus enforcing both the retention of minor alleles, and an observed behavior in mating to favor partners with different MHC complements. Apparently, our MHC makeup is reflected in our personal aroma! 

A second force, conversely, is the retention of ancient alleles. It turns out that, across the old-world monkeys, many MHC alleles are preserved and cluster more closely in sequence comparisons with each other than they do with other alleles in the same species. That is, despite the general speed of MHC evolution and constant accumulation of new alleles, old alleles are preserved in all monkey populations as well, due to their distinct capabilities, under balancing selection. This is part of what makes population bottlenecks so damaging to near-extinction species. They lose critically valuable genetic resources (in the form of rare MHC alleles) that represent millions of years of accumulated variation. 

Incidentally, the trees shown here again reinforce the history of primate evolution, with new world monkeys splitting off from the old-world monkeys quite early on and developing a very distinct set of MHC molecules.

So, while virtually every other gene in the genome is being relentlessly optimized, sticking to its knitting, doing one thing and being beaten down whenever any mutation steers it from its optimized path, the MHC genes follow quite a different path, at least in portions of their sequence that provide variation in antigen binding and presentation. These genes revel in endless diversity, throw off pseudogenes at a high rate, wink out of existence and come back in other forms. Natural selection is the motor in each case, but meets the challenge of survival in different ways.


Sunday, July 12, 2026

How Do Complex Human Traits Add Up?

Notes on the genetics, traits, and evolution.

Everything about us is a trait. Not everything about our traits is genetic, though. The conundrum of nature vs nurture, of genes vs environment, and the structure and meaning of genetics goes to the heart of biology. A few traits, like eye color, are simple enough. But they are the exception, by far. Body mass index is influenced by practically every gene we have. And autism has, by this point, hundreds of contributing genes. Both traits are highly heritable, in the sense that inheritance/genes are the dominant influence, vs environment (as seen in twins) when most conditions are equal. But environment can easily be dominant over BMI when conditions change and starvation sets in. 

A puzzle that came out of the early human genome studies was how unhelpful it was to do genome-wide association studies (GWAS) to approach some of these questions- that is, studies of what variants in the population at large contribute to particular traits, especially to serious diseases. Study after study was done, and disappointment mounted that what were found were genes with minor effects, in tangential biological processes. This was supposed to be the holy grail- the payoff for sequencing the human genome- and what came up was dud after dud.

A recent paper plows over this ground again with a new mathematical synthesis of genetic structure of human traits, genetic alleles, and selection. What it finds is sort of obvious, but there are some intriguing observations along the way. It is critical to note at the outset that evolution as Darwin understood and described, and natural selection in particular, is absolutely at the heart of this or any contemporary analysis of genetics. While Darwin's understanding was revolutionary and broad, subsequent decades of work have brought these concepts to a very concrete, operational, indeed mathematical, level. 

Let's start with the concept of allele frequency, also called minor allele frequency, of MAF. Over any individual genome, there will be millions of "variants", which are coding differences from the reference genome. Which does not have any special status... it is just the genome of some guy from Buffalo. Variants (or alleles) are bases in the genome that differ from the reference. With three billion positions and four possible bases per position, that means that there are nine billion possible variants. How common is a particular variant in a population? That is its allele frequency. For GWAS and related studies, the threshold is commonly set at variants seen in the population at over one percent frequency, while minor alleles are seen under that frequency. A variant that causes some devastating disease is typically one that sprang up recently, and is heavily selected against. That is why it must have an exceedingly low allele frequency. On the other hand, a variant may have no discernable effect at all, not being selected for or against, thus just drifts along in the genome, not subject to natural selection. Such alleles may, over long periods of time, drift to higher or lower frequency by random chance.

However, the focus of GWAS association studies are variants between these extremes. These are variants that have some effect on a trait (or may be physically close to others that do, thus get "carried" along over time). At the same time, they are also common in the population, at least common enough to be discernable in an association study. Such a study needs some statistical correlation between the occurrence of the variant, and the occurrence of the trait. That means that the variant can not just appear once, but must appear many times over a large population. At the same time, a study that focuses on a trait like, say, high blood pressure, will be seeing variants that are, by definition, deleterious. That means that any allele with a large effect will be subject to strong selection, and driven out of the population. Only alleles with more modest effects will be able to survive at all, and even then, at low frequencies. So a GWAS focuses on medium-to-low prevalence variants, hunting for alleles on the loose in a large population that have typically modest effects on a given disease or other trait. Such alleles will have typically survived for tens or even hundreds of thousands of years, so they will have some complex relationship to natural selection. 

In contrast to all this is the family study, which focuses on a dramatic variant that causes some terrible disease. Such studies have been remarkably productive, because they deal with extremely rare, high-effect variants, which are as a rule very informative about the genesis of that disease. Such variants, as mentioned above, would be heavily selected against, thus disappear rapidly. But mutation is always happening, so all sorts of mutations arise in large enough populations. Assuming that, as biologists, we are interested in the core ten or fewer genes that most influence a given trait or condition, the hundreds or thousands of significant, but low-effect variants that come out of GWAS are almost by definition guaranteed to be tangential and minor. It turns out that most traits are complex, in the sense of being influenced in various minor ways by hundreds of genes.

So, what is the typical genetic structure of complex traits? That is- what this paper set out to answer. "Structure" in this case means ... what is the normal distribution of selective target / effect size versus frequency/prevalence in the population of variants that, in combination, add up to a complex trait? The assumption (as discussed above) is that the core armature of such traits does not have variants at all, due to strong selection, while the available variants in the population all have minor effects that in sum form the genetic variation seen in the trait in the population. 

While other researchers have attempted to fit the variation distribution of complex traits to typical formulas like the normal distribution, these authors found that a natural selection-informed approach gave a clear and simple result. All traits follow the same general scaling, with only two parameters- the mutational target size of the trait (that is, the proportion of the genome capable of appearing as relevant variants), and the effect that each site has on the given trait, termed (very poorly) the site's "heritability". It is important to note that every site in the genome is equally and fully heritable. The term refers to the trait, and the size of the effect from variations of that site on that trait. Summed over all sites in the genome and all variations in the population, this heritability ultimately equates to the overall variation of the trait that is genetically caused.

A comparison of two traits, and how they might look in a genetic variation study. In blue is a trait skewed towards small effect variants, with weak selection and consequently variants with higher frequency. In red is a different trait that partakes more from stronger effect variants. On the whole, this kind of difference is not common among complex traits that arise from thousands of loci. MAF = minor allele frequency; Z-score is the score in a GWAS study indicating how statistically significant the variant's effect on the trait is. Note how lower Z-score correlates with more variants at the higher MAF frequencies. At the same time, the GWAS method overall has some skew to higher MAF frequencies, since only those provide sufficient statistical power to get any results at all. Log(s) is the strength of selection; L is the genetic target size for the whole trait, and h*2 is the heritability, or proportion of the trait effect due to the causal variant.

The model they come up with accounts for the selective effect of trait effects (the larger the effect of the variation on traits, the lower its frequency in the population). It also accounts for the fact that variants that affect one trait often affect other traits as well, so the selective effect needs to be considered over all affected traits, most of which are probably unknown, but can be inferred. And conversely, most traits are composed of contributions from many genes and their alleles, sometimes thousands- they are genetically complex traits. 

The model the researchers come up with can normalize among the huge population of variants that affect a single trait, in this case blood pressure. Left shows the effect sizes of individual variants, and right shows the scaled (normalized) version from the paper's model, showing that all these variants follow the same overall rule / logic, using the custom parameters of h*2 and L- trait heritability and mutational target size. All this is to say that the lower effect variants (skewed to left) are assumed to have higher selection coefficients.

The researchers go on to show that various traits do look different under this analysis. Some differ mostly by target size, accounting for more or fewer variants, but having a similar spread of effect sizes over the population. Others differ by the scale of effects that each variant contributes, thus skewing toward higher or lower allele frequencies overall. Interestingly, they add an analysis of the age of these low-frequency, low-effect variants that are the grist for GWAS, finding that they are on average 137,000 years old. That compares with an average age of 600,000 years for variants that are neutral, thus would not come up in GWAS or be under selection. This is fascinating in its implications both for human population genetics in general, and for the fact that most human variation- even that under modest selection- predates the divergence between African and non-African populations. 


Sunday, April 26, 2026

The History and Future of a Single Mutation

The CCR5delta32 confers resistance to HIV. Where did it come from?

We are edging into an age of precision medicine, where the causes of our maladies will be known in molecular detail, allowing treatments that address them at the root. Given the parlous state of medicine today, in the midst of financial breakdown and a continued mediocre level of basic diagnosis, it is hard to believe this is a corner we can turn. But vaccines have long been in this category, of addressing the precise pathogenic causes of disease, and oncology is fitfully getting there, given advances in DNA sequencing and in treatments based on specific mutations.

HIV is also a beneficiary of this approach, since the discovery of its pathogen led directly to a variety of effective (if not yet permanent) treatments. A researcher in China created gene-edited humans with a specific mutation that will render them resistant to HIV. The mutation he chose for this work is called CCR5delta32, and it does not naturally exist in Chinese populations. 

But it does exist in European populations, at a roughly 10% rate in single copy. When present in two copies, it provides complete immunity to HIV, while if present in one copy, it slows infection substantially. A recent paper rooted through the available ancient and present genomes to figure out where this mutation came from. 

CCR5 is a cell surface receptor for cytokines 3, 4, and 5. These are all pro-inflammatory cytokines, and they interact with multiple receptors. Here, as in so many other respects, the immune system is riven with redundancy, so that it can grapple with as many contingencies as possible. Cytokines are signaling molecules for the immune system, which is an unusual organ, being dispersed all over the body with numerous cell types all patrolling around, and communicating with each other by long- and short-range chemical messages. It turns out that the major form of HIV uses the CCR5 protein to get into our immune cells, explaining why CCR5delta32, which is totally non-functional, has such a dramatic effect on HIV susceptibility. 

While people carrying CCR5delta32 are generally fine, this defect does confer a variety of subtle changes to their susceptibility to other infectious diseases and cancers. That explains why this mutation has settled at its low level in the European populations, probably balancing the occasional benefit against a specifically CCR5-seeking pathogen against its natural functions that form the basis of its existence in the first place as a part of immune system that is conserved in all mammals. The Chinese gene-editing researcher came under withering criticism not only for breaching the generally agreed moratorium on human germline gene editing, but also because the net effect of this mutation is, on the whole, negative, raising risks of numerous diseases, despite its beneficial effect on HIV. 

The authors run several models and populations in an attempt to time the origin of the CCR5delta32 mutation, and portray its positive selection over the ensuing millenia.  CHG- Caucasus hunter-gatherer; EHG- Eastern hunter-gatherer; WHG- Western hunter-gatherer; ANA- Anatolian Neolithic ancestry. The bottom axis is time, and the Y axis is the frequency of the mutation in these populations. "Modern DAF" refers to the inclusion of the data set of current (not ancient) population frequencies, (top), which the authors claim leads to continued rates of selection (last 2,000 years) that are artifactual.

So where did it come from? The new authors gather up a large variety of population samples from around the world, and from ancient humans, back to about ten thousand years ago. They find the first instance of the mutation in one sample at 5.8 thousand years ago. After that, its frequency rises dramatically, up to about two thousand years ago, when it levels off. They conclude that this mutation originated about seven to nine thousand years ago, in the steppes of Eastern Europe / Western Asia, and was under strong positive selection at first, spreading to the current frequency of about 10% of the population / alleles. All occurrences on other continents can be accounted by the spread from this source.

Does this mean that HIV was prevalent long, long before the current pandemic? Hardly. The authors can not say anything about it, but one theory would be that some other disease had a similar profile. It certainly was not the Black Death, as the authors show that this mutation had no change in frequency over that gruesome pandemic. Another hypothesis is that general reduction in inflammatory response might be beneficial in some settings, as has been found for Covid-19, though here again, this mutation does not have any known positive or negative net effect on Covid-19 susceptibility or course. 

It is amazing that we have enough sequences of ancient DNA to be able to reconstruct this kind of thing- to be able to trace where and when some influential mutation occurred, and how it traveled and spread. It is a tour-de-force of bio-archeological reconstruction.


  • When you escape reality, and morality.
  • Some environmental benefits are flowing from the current war.
  • We may be at peak oil, courtesy of the US.

Saturday, April 4, 2026

Not Every Transcript is Golden

 Reflections on junk DNA, and junk transcripts.

Some time ago, a large project in molecular biology determined that most regions of the genome are transcribed. The authors and most observers took this to mean that most regions are functional, quite in contrast to the reigning theory up to that point, that our genomes host a smattering of genes floating in a sea of "junk" DNA. That theory was based on the now-ancient observations of reannealing curves for bulk DNA from humans and other species which found that most of our DNA re-anneals very quickly, due to the fact that it is repetitive. Most of our genomes (60%) are taken up with LINE repeats, SINE repeats, old retro-transposons, stray duplications, and other repetitive material that, at a first glance, seems like junk. There has been a battle ever since, between proponents of junk DNA and those who see function around every corner. As we learn more about the genome, many more functions have indeed come to light, like distant enhancers and regulatory RNAs of many flavors. But overall, there still seems to be a lot of junk. 

A recent paper took an oblique shot at this field, looking at the profusion of alternative gene transcripts, which can number into the hundreds for a single gene. (This was also reviewed.) These are generally called isoforms, and arise due to variable ways one gene's RNA products can be initiated, terminated, and spliced. So not only are most regions of the genome transcribed in some form, actively transcribed regions can be transcribed and processed in myriad ways to lead to different RNA products. Here again, there has been an analogous argument, about whether every such isoform has a function, or whether isoforms might arise from more or less sporadic processes, often as unintended and non-functional sparks coming out of the machinery. The importance of isoforms is very well documented in many cases, so the possibility of function, sometimes highly conserved, is not in question. Only the importance of every last variation in combinatorial collections of isoforms that can number into the hundreds.

Here is an image from the first page (of about six pages) of RNA transcripts coming off the notorious BRCA1 gene, which is intensely studied for its role in breast cancer. Each line is a distinct mRNA transcript. Each darker bar is an exon, which are separated by introns. The darker colored exons are in the protein coding region, while the lighter exons signify the untranslated upstream and downstream ends. I count about 315 transcripts described for this genetic locus. The idea that each of these has some evolutionarily constrained and important function is, on the face of it, absurd.

The authors took an interesting evolutionary approach, reasoning that species with larger population sizes experience more stringent purifying selection, and thus should (in theory) show tighter control over stray genomic products such as isoforms, if most transcript isoforms are neutral (or even deleterious) accidents, rather than intentional and functional forms. Thankfully, animals come in a wide range of population sizes, from insects to crocodiles and primates; very large to very small. While population size is hard to calculate, several convenient proxies are known, like lifespan, body size, etc. When they totted everything up, they saw clear correlations between these proxies and the number of alternative RNA products per gene- also termed transcript diversity. They sliced up the data by organ where the RNA was expressed, and by the source of the RNA variation- either different initiation, different termination, different splicing. In all cases the trend was the same. In species with larger population sizes, the diversity of transcripts was lower, agreeing with their hypothesis that when greater selecive force is available, the slop from the transcription and transcript processing machinery declines.

The authors draw correlations between alternative splicing (AI) diversity in an organism's cells and its population size. 

The authors additionally note that there is a similar relationship between alternative splice site usage and expression level of a gene. That is, the higher the gene expression, the less likely that minor splice sites are used, indicating that here again, higher selective pressure helps to clear out non-functional off-products of the transcription apparatus.

The correlations found here are only that- correlations. While significant, they are not terribly strong, let alone stark. So it is evident that our gene expression machinery has a lot of play in it, and this falls on a spectrum from deleterious to critically functional. It is, after all, machinery, not divine. It is also grist for evolution itself- it is useful to have some slop so that there is always some diversity in the gearing to accommodate new selective pressures. But the idea that just because a distinct transcript exists, it is biologically functional, or that, similarly, because a genomic region is transcribed, it is a "gene" rather than junk DNA.. that does not hold water. Every nucleotide in the genome has its own unique selective constraints, and for many of them, that constraint is zero.


  • The world order, and our position in it, is crumbling.
  • Whence Hungary?
  • Another AI tax, as if gobbling up power wasn't bad enough.
  • Mindless.

Saturday, March 28, 2026

Death and Resurrection ... Of a Gene

The SLAMF9 gene became non-functional in the human lineage, and then later was re-activated. Why?

Biology is amazingly intricate, but it is often also needlessly complex- evidence for the haphazard, if eventually pointed, mechanisms of the evolutionary process. We will take up the discussion of "junk" DNA again next week, but molecular biology is full of redundant and excessive processes, which should certainly be mystifying from a "design" perspective. At the frontier of natural selection are neutral and near-neutral genetic elements, which change over time due to chance, lacking selection pressure towards conservation. Pseudogenes (of which we have about 20,000- almost as many as functional genes) are one form of neutral element. They are typically remnants of functional genes that have been duplicated and inactivated by mutation. They are a lively area of genome annotation because it is hard to be sure that they are really dead. Despite what looks like an inactivating mutation, they typically still produce RNA transcripts, and may produce partial or alternative proteins as well. The literature is full of experiments finding products and activities from genes annotated elsewhere as pseudogenes. And what looks like a pseudogene from one sample might just be an allele, the same gene being whole and active in other people.

So, it is hard to know what any particular genetic region is doing without a lot of evolutionary, functional, and even population analysis. A recent paper looked deeply at one gene- a gene that seems to have flipped back and forth between functional and non-functional states in the human lineage. It is a rare example of a gene coming back from what is usually a one-way trip into mutational oblivion, once its function- and thus selective pressure for conservation- have disappeared.

SLAMF9 is one of a family (signaling lymphocyte activation molecule family) of surface receptors that occur in many cells of the immune system, help activate responses in these cells, and also recognize some viruses and bacteria. They bind to each other and to other components of the immune system, creating complex signaling networks. Genes involved in our immune systems are commonly subject to rapid evolution, the arms race against our many pathogens being relentless. Sometimes that takes the form of gene inactivation, if a particular receptor, for instance, has been turned against us by a pathogen that uses it for binding and cell entry. 

This week's authors were facing a conundrum. They were studying SLAMF9, and found the mouse version easy to clone and express in the lab. But the human version ... that was another story, frustratingly impossible to express in usable amounts. When they looked at the protein sequence, they were in for a big surprise:

At the front end of SLAMF9, there is very strong conservation across mammals... except when it comes to humans! The signal peptide is what directs this protein to be inserted into the plasma membrane, and is cleaved off the mature protein. In red is highlighted the region starkly different in humans, which naturally affects (not in a good way) the signal cleavage process. "a" and "b" point to important domains of the cytoplasmic side of the final protein, which are just barely preserved/conserved in the human form.

This alignment among various mammalian versions (orthologs) of SLAMF9 shows that they are all pretty much the same... except for the human version. All the way from mouse to chimpanzee nothing has changed at the front end of this protein. That is amazing in itself, showing very strong conservation. But then after our lineage split from chimpanzees, something weird. A small segment at the front of this protein is totally different. This area is important because it carries the cleavage site of the signal sequence. The signal sequence directs the protein to be sent to the membrane (as this is a trans-membrane receptor), and this cleavage site is bad, explaining why the author's attempt to express this protein went so poorly. It might be enough for modest expression in the natural setting, but not enough for their investigations.

At the DNA level, it is clear that what happened to the protein was a double frame shift in translation, out of frame at the front, then recovered frame at the second mutation. The mutations must have been independent events, but the order of their occurrence is not known. The first intron trails off to the left, while the coding sequence tails off to the right.

When they looked at the DNA sequence, the reason for this change in the protein sequence became clearer. There was a frame shift, with only small changes in the DNA sequence that led to the bigger change in the protein sequence. On the left, there is a shift in the splice site at the end of the first intron (splice acceptor). This shifts the mRNA product by four bases (vs the start site of translation), creating a frame shift in translation, as portrayed in the amino acid codes given. On the right, there is a one nucleotide deletion, causing another frame shift that brings the translation back into the normal frame. 

They sampled all the available archeological samples from the human lineage- Neanderthals and Denisovans, and each were the same as the current human sequence. So, whatever happened did so between the split from chimpanzees and the advent of these available homo species. And what happened were two distinct events- the second frame shift and the first frame shift are independent genetic mutations. 

Which happened first? That is uncertain, but the authors show that the right-most frame shift (called g.621delT) did not influence the change in the splice site. The splice site change was caused by a series of about six mutations within the first intron, (not shown), which shifted the pattern of mRNA self-hybridization that helps direct splice site selection. So it is likely that the splice site change happened first, essentially killing the gene. And then the downstream frameshift happened later on to rescue it in a partial, not very well-expressed way. However, either mutation could have happened first to functionally kill off this gene, and then further mutation(s) to recover its function. In any case, both events happened within this roughly six-million-year time span that generated our immediate lineage, becoming firmly fixed as the only version of this gene now in our collective genome.

What might cause these events? It all goes back to the function of SLAMF9. As shown above, it is very highly conserved. But, being part of the immune system and the interface we show to pathogens, it is also on the front line of the bio-warfare arms race. As humans started ranging far beyond their original habitats, they doubtless encountered many new pathogens. It seems likely that killing off this gene might have resolved one such fight, at least for a little while, perhaps by removing a pathogen entry point. But later on, it became beneficial to recover it, which is to say that new mutations that restored its function even a little bit were evidently selected for, and spread in the population. There was a race at this point between the accumulation of more (now neutral) mutations that would have permanently inactivated this gene, and the advent of that one special mutation that could save it. The overall conservation of SLAMF9 argues that saving it must have conferred significant benefits.


Saturday, December 13, 2025

Mutations That Make Us Human

The ongoing quest to make biologic sense of genomic regions that differentiate us from other apes.

Some people are still, at this late date, taken aback by the fact that we are animals, biologically hardly more than cousins to fellow apes like the chimpanzee, and descendants through billions of years of other life forms far more humble. It has taken a lot of suffering and drama to get to where we are today. But what are those specific genetic endowments that make us different from the other apes? That, like much of genetics and genetic variation, is a tough question to answer.

At the DNA level, we are roughly one percent different from chimpanzees. A recent sequencing of great apes provided a gross overview of these differences. There are inversions, and larger changes in junk DNA that can look like bigger differences, but these have little biological importance, and are not counted in the sequence difference. A difference of one percent is really quite large. For a three gigabyte genome, that works out to 30 million differences. That is plenty of room for big things to happen.

Gross alignment of one chromosome between the great apes. [HSA- human, PTR- chimpanzee, PPA- bonobo, GGO- gorilla, PPY- orangutan (Borneo), PAB- orangutan (Sumatra)]. Fully aligned regions (not showing smaller single nucleotide differences) are shown in blue. Large inversions of DNA order are shown in yellow. Other junk DNA gains and losses are shown in red, pink, purple. One large-scale jump of a DNA segment is show in green. One can see that there has been significant rearrangement of genomes along the way, even as most of this chromosome (and others as well) are easly alignable and traceable through the evolutionary tree.


But most of those differences are totally unimportant. Mutations happen all the time, and most have no effect, since most positions (particularly the most variable ones) in our DNA are junk, like transposons, heterochromatin, telomeres, centromeres, introns, intergenic space, etc. Even in protein-coding genes, a third of the positions are "synonymous", with no effect on the coded amino acid, and even when an amino acid is changed, that protein's function is frequently unaffected. The next biggest group of mutations have bad effects, and are selected against. These make up the tragic pool of genetic syndromes and diseases, from mild to severe. Only a tiny proportion of mutations will have been beneficial at any point in this story. But those mutations have tremendous power. They can drag along their local DNA regions as they are positively selected, and gain "fixation" in the genome, which is to say, they are sufficiently beneficial to their hosts that they outcompete all others, with the ultimate result that mutation becomes universal in the population- the new standard. This process happens in parallel, across all positions of the genome, all at the same time. So a process that seems painfully slow can actually add up to quite a bit of change over evolutionary time, as we see.

So the hunt was on to find "human accelerated regions" (HAR), which are parts of our genome that were conserved in other apes, but suddenly changed on the way to humans. There roughly three thousand such regions, but figuring out what they might be doing is quite difficult, and there is a long tail from strong to weak effects. There are two general rationales for their occurrence. First, selection was lost over a genomic region, if that function became unimportant. That would allow faster mutation and divergence from the progenitors. Or second, some novel beneficial mutation happened there, bringing it under positive selection and to fixation. Some recent work found, interestingly, that clusters of mutations in HAR segments often have countervailing effects, with one major mutation causing one change, and a few other mutations (vs the ancestral sequence) causing opposite changes, in a process hypothesized to amount to evolutionary fine tuning. 

A second property of HARs is that they are overwhelmingly not in coding regions of the genome, but in regulatory areas. They constitute fine tuning adjustments of timing and amount of gene regulation, not so much changes in the proteins produced. That is, our evolution was more about subtle changes in management of processes than of the processes themselves. A recent paper delved in detail into HAR5, one of the strongest such regions, (that is, strongest prior conservation, compared with changes in human sequence), which lies in the regulatory regions upstream of Frizzled8 (FZD8). FZD8 is a cell surface receptor, which receives signals from a class of signaling molecules called WNT (wingless and int). These molecules were originally discovered in flies, where they signal body development programs, allowing cells to know where they are and when they are in the developmental program, in relation to cells next door, and then to grow or migrate as needed. They have central roles in embryonic development, in organ development, and also in cancer, where their function is misused.

For our story, the WNT/FZD8 circuit is important in fetal brain development. Our brains undergo massive cell division and migration during fetal development, and clearly this is one of the most momentous and interesting differences between ourselves and all other animals. The current authors made mutations in mice that reproduce some of the HAR5 sequences, and investigated their effects. 

Two mouse brains at three months of age, one with the human version of the HAR5 region. Hard to see here, but the latter brain is ~7% bigger.

The authors claim that these brains, one with native mouse sequence, and the other with the human sequences from HAR5, have about a seven percent difference in mass. Thus the HAR5 region, all by itself, explains about one fourteenth of the gross difference in brain size between us and chimpanzees. 

HAR5 is a 619 base-pair region with only four sequence differences between ourselves and chimpanzees. It lies 300,000 bases upstream of FZD8, in a vast region of over a million base pairs with no genes. While this region contains many regulatory elements, (generally called enhancers or enhancer modules, only some of which are mapped), it is at the same time an example of junk DNA, where most of the individual positions in this vast sea of DNA are likely of little significance. The multifarious regulation by all these modules is of course important because this receptor participates in so many different developmental programs, and has doubtless been fine-tuned over the millennia not just for brain development, but for every location and time point where it is needed.

Location of the FZD8 gene, in the standard view of the genome at NIH. I have added an arrow that points to the tiny (in relative terms) FZD8 coding region (green), and a star at the location of HAR5, far upstream among a multitude of enhancer sequences. One can see that this upstream region is a vast area (of roughly 1.5 million bases) with no other genes in sight, providing space for extremely complicated and detailed regulation, little of which is as yet characterized.

Diving into the HAR5 functions in more detail, the authors show that it directly increases FZD8 gene expression, (about 2 fold, in very rough terms), while deleting the region from mice strongly decreases expression in mice. Of the four individual base changes in the HAR5 region, two have strong (additive) effects increasing FZD8 expression, while the other two have weaker, but still activating, effects. Thus, no compensatory regulation here.. it is full speed ahead at HAR5 for bigger brain size. Additionally, a variant in human populations that is responsible for autism spectrum disorders also resides in this region, and the authors show that this change decreases FZD8 expression about 20%. Small numbers, sure, but for a process that directs cell division over many cycles in early brain development, this kind of difference can have profound effects.


The HAR5 region causes increased transcription of FZD8, in mice, compared to the native version and a deletion.

The HAR5 region causes increased cell proliferation in embryonic day 14.5 brain areas, stained for neural markers.

"This reveals Hs-HARE5 modifies radial glial progenitor behavior, with increased self-renewal at early developmental stages followed by expanded neurogenic potential. ... Using these orthogonal strategies we show four human-specific variants in HARE5 drive increased enhancer activity which promotes progenitor proliferation. These findings illustrate how small changes in regulatory DNA can directly impact critical signaling pathways and brain development."

So there you have it. The nuts and bolts of evolution, from the molecular to the cellular, the organ, and then the organismal, levels. Humans do not just have bigger brains, but better brains, and countless other subtle differences all over the body. Each of these is directed by genetic differences, as the combined inheritance of the last six million years since our divergence versus chimpanzees. Only with the modern molecular tools can we see Darwin's vision come into concrete focus, as particular, even quantum, changes in the code, and thus biology, of humanity. There is a great deal left to decipher, but the answers are all in there, waiting.


Saturday, November 22, 2025

Ground Truth for Genetic Mutations

Saturation mutagenasis shows that our estimates of the functional effect of uncharacterized mutations are not so great.

Human genomes can now be sequenced for less than $1,000. This technological revolution has enabled a large expansion of genetic testing, used for cancer tissue diagnosis and tracking, and for genetic syndrome analysis both of embryos before birth and affected people after birth. But just because a base among the 3 billion of the genome is different from the "reference" genome, that does not mean it is bad. Judging whether a variant (the modern, more neutral term for mutation) is bad takes a lot of educated guesswork.

A recent paper described a deep dive into one gene, where the authors created and characterized the functional consequence of every possible coding variant. Then they evaluated how well our current rules of thumb and prediction programs for variant analysis compare with what they found. It was a mediocre performance. The gene is CDKN2A, one of our more curious oddities. This is an important tumor suppressor gene that inhibits cell cycle progression and promotes DNA repair- it is often mutated in cancers. But it encodes not one, but two entirely different proteins, by virtue of a complex mRNA splicing pattern that uses distinct exons in some coding portions, and parts of one sequence in two different frames, to encode these two proteins, called p16 and p14. 

One gene, two proteins. CDKN2A has a splicing pattern (mRNA exons shown as boxes at top, with pink segments leading to the p14 product, and the blue segments leading the p16 product) that generates two entirely different proteins from one gene. Each product has tumor suppressing effects, though via distinct mechanisms.

Regardless of the complex splicing and protein coding characteristics, the authors generated all possible variants in every possible coded amino acid (156 amino acids in all, as both produced proteins are relatively short). Since the primary roles of these proteins are in cell cycle and proliferation control, it was possible to assay function by their effect when expressed in cultured pancreatic cells. A deleterious effect on the protein was revealed as, paradoxically, increased growth of these cells. They found that about 600 of the 3,000 different variants in their catalog had such an effect, or 20%.

This is an expected rate of effect, on the whole. Most positions in proteins are not that important, and can be substituted by several similar amino acids. For a typical enzyme, for instance, the active site may be made up of a few amino acids in a particular orientation, and the rest of the protein is there to fold into the required shape to form that active site. Similar folding can be facilitated by numerous amino acids at most positions, as has been richly documented in evolutionary studies of closely-related proteins. These p16 and p14 proteins interact with a few partners, so they need to maintain those key interfacial surfaces to be fully functional. Additionally, the assay these researchers ran, of a few generations of growth, is far less sensitive than a long-term true evolutionary setting, which can sift out very small effects on a protein, so they were setting a relatively high bar for seeing a deleterious effect. They did a selective replication of their own study, and found a reproducibility rate of about 80%, which is not great, frankly.

"Of variants identified in patients with cancer and previously reported to be functionally deleterious in published literature and/or reported in ClinVar as pathogenic or likely pathogenic (benchmark pathogenic variants), 27 of 32 (84.4%) were functionally deleterious in our assay"

"Of 156 synonymous variants and six missense variants previously reported to be functionally neutral in published literature and/or reported in ClinVar as benign or likely benign (benchmark benign variants), all were characterized as functionally neutral in our assay "

"Of 31 VUSs previously reported to be functionally deleterious, 28 (90.3%) were functionally deleterious and 3 (9.7%) were of indeterminate function in our assay."

"Similarly, of 18 VUSs previously reported to be functionally neutral, 16 (88.9%) were functionally neutral and 2 (11.1%) were of indeterminate function in our assay"

Here we get to the key issues. Variants are generally classified as benign, pathogenic/deleterious, or "variant of unknown/uncertain significance". The latter are particularly vexing to clinical geneticists. The whole point of sequencing a patient's tumor or genomic DNA is to find causal variants that can illuminate their condition, and possibly direct treatment. Seeing lots of "VUS" in the report leaves everyone in the dark. The authors pulled in all the common prediction programs that are officially sanctioned by the ACMG- Americal College of Medical Genetics, which is the foremost guide to clinical genetics, including the functional prediction of otherwise uncharacterized sequence variants. There are seven such programs, including one driven by AI, AlphaMissense that is related to the Nobel prize-winning AlphaFold. 

These programs strain to classify uncharacterized mutations as "likely pathogenic", "likely benign", or, if unable to make a conclusion, VUS/indeterminate. They rely on many kinds of data, like amino acid similarity, protein structure, evolutionary conservation, and known effects in proteins of related structure. They can be extensively validated against known mutations, and against new experimental work as it comes out, so we have a pretty good idea of how they perform. Thus they are trusted to some extent to provide clinical judgements, in the absence of better data. 

Each of seven programs (on bottom) gives estimations of variant effect over the same pool of mutations generated in this paper. This was a weird way to present simple data, but each bar contains the functional results the authors developed in their own data (numbers at the bottom, in parentheses, vertical). The bars were then colored with the rate of deleterious (black) vs benign (white) prediction from the program. The ideal case would be total black for the first bar in each set of three (deleterious) and total white in the third bar in each set (benign). The overall lineup/accuracy of all program predictions vs the author data was then overlaid by a red bar (right axis). The PrimateAI program was specially derived from comparison of homologous genes from primates only, yielding a high-quality dataset about the importance of each coded amino acid. However, it only gave estimates for 906 out of the whole set of 2964 variants. On the other hand, cruder programs like PolyPhen-2 gave less than 40% accuracy, which is quite disappointing for clinical use.

As shown above, the algorithms gave highly variable results, from under 40% accurate to over 80%. It is pretty clear that some of the lesser programs should be phased out. Of programs that fielded all the variants, the best were AlphaMissense and VEST, which each achieved about 70% accuracy. This is still not great. The issue is that, if a whole genome sequence is run for a patient with an obscure disease or syndrome, and variants vs the reference sequence are seen in several hundred genes, then a gene like CDKN2A could easily be pulled into the list of pathogenic (and possibly causal) variants, or be left out, on very shaky evidence. That is why even small increments in accuracy are critically important in this field. Genetic testing is a classic needle-in-a-haystack problem- a quest to find the one mutation (out of millions) that is driving a patient's cancer, or a child's inherited syndrome.

Still outstanding is the issue of non-coding variants. Genes are not just affected by mutations in their protein coding regions (indeed many important genes do not code for proteins at all), but by regulatory regions nearby and far. This is a huge area of mutation effects that are not really algorithmically accessible yet. As a prediction problem, it is far more difficult than predicting effects on a coded protein. It will requiring modeling of the entire gene expression apparatus, much of which remains shrouded in mystery.


Saturday, July 5, 2025

Water Sensing by WNKs

WNK kinases sense osmotic condition as well as chloride concentration to keep us hydrated.

"Water, water, everywhere, nor any drop to drink." This line from Coleridge evokes the horror of thirst on the ghost ship, as its crew can not drink salt water. Other species can, but ocean water is too strong for us, roughly four times as salty as our blood. Nevertheless, our bodies have exquisite mechanisms to manage salt concentrations, with each cell managing its own traffic, and the kidneys managing most electrolytes in the blood. It is a very difficult task that has led to clever evolutionary solutions like counter-current exchange across the nephron loops, and stark differences in those nephron cell membranes, over water or salt permeability, to maximize use of passive ion gradients. But at the heart of the system, one has to know what is going on- one has to monitor all of the electrolyte levels and overall osmotic stress.

One such monitoring thermostat for chemical balances turns out to be the WNK kinases- a family of four proteins in humans that control (by phosphorylating them) a secondary set of regulators, which in turn control many salt transporters, such as SLC12A2 and SLC12A4. These latter are passive, though regulated, co-transporters that allow chloride across the membrane when combined with a matching cation like sodium or potassium. The cations drive the process, because they are normally kept (pumped) to strong gradients across cell membranes, with high sodium outside, and high potassium inside. Thus when these co-transporters are turned on (or off), they use the cation gradients to control the chloride level in the cell, in either direction, depending on the particular transporter involved. Since the sodium and potassium levels are held at relatively static, pumped levels, it is the chloride level that helps control the overall osmotic pressure in a finely tuned way. 

A few of the ionic transactions done in the kidney.


The WNK kinases were discovered genetically, in families that showed hypertension and raised levels of chloride and potassium in the blood. These syndromes mirrored complementary syndromes caused by mutations in SLC12A2, the Na/Cl co-transporter, indicating the WNK kinases inhibit SLC12A2. It turns out that WNK, which are named for an unusual catalytic site (with no lysine [K]) are sensors for both chloride, which inhibit them, and for osmotic pressure, which activates them. They are expressed in different locations and have slightly different activities, (and control many more transporters and processes than discussed here), but I will treat them interchangeably here. The logic of all this is that, if osmotic pressure is low, that means that internal salt levels are low, and chloride needs to be let into the cell, by activating the cation/chloride co-transporters. Likewise, if chloride levels inside the cell are high, the WNK kinase needs to be inhibited, reducing chloride influx. 

A recent paper (and prior work from the same lab) discussed structures of the WNK regulators that explain some of this behavior. WNK kinases are dimers at rest, and in that state mutually inhibit their auto-phosphorylation. It is separation and auto-phosphorylation that turns them on, after which they can then phosphorylate their target proteins, such as the secondary kinases STK39 and OSR1. The authors had previously found a chloride binding site right at the active site of the enzyme that promotes dimerization. In the current paper, they reveal a couple of clusters of water molecules which similarly affect the dimerization, and thus activity, of the enzyme.

Location of the inhibitory chloride (green) binding site in WNK1. This is right in the heart of the protein, near the active kinase site and dimerization interface with the other WNK1 partner.

While X-ray crystal structures rarely show or care much about water molecules, (they are extremely small and hard to track), here, those waters were hypothesized to be important, since WNK kinases are responsive to osmotic pressure. One way to test this is to add PEG400 to the reaction. This is a polymer (400 molecular weight) that is water-like and inert, but large in a way that crowds out water molecules from the solution. At 15% or 25% of the volume, PEG400 displaces a lot of water, lowers the water activity of a solution, and thus increases the osmotic pressure- that is its tendency to draw water in from outside. Plants use osmotic pressure as turgor pressure, and our cells, not having cells walls, need to always be at an osmotic pressure similar to the outside, lest they swell up, or conversely shrink away. Anyhow, WKN kinases can be switched from an inactive to active state just by adding PEG400- a sure sign that they are sensors for osmotic pressure.


Water network (blue dots) within the WNK1 kinase protein. Most of the protein is colored teal, while the active site kinase area is red, and a tiny amount of the dimer partner is colored green. When this crystal is osmotically challenged, the water network collapses from 14 waters to 5, changing the structure and promoting dissociation of the dimer. In B is show a sequence alignment over a wide evolutionary range where the amino acids that coordinate the water network (yellow) are clearly very well conserved, thus quite important.

Above is shown a closeup of the WNK1 protein, showing in teal the main backbone, including the catalytic loop. In red is the activation loop of the kinase, and in green is a little bit from the other WNK1 protein in the dimer pair. The chloride, if bound, would be located right at top center, at K375. Shown in blue are a series of fourteen water molecules that make up one so-called water network. Another smaller one was found at the interface between the two WNK1 proteins. The key finding was that, if crystalized with PEG400, this water network collapsed to only five water molecules, thereby changing the structure of the protein significantly and accounting for the dissolution of the dimer. 

Superposition of WNK1 with PEG400 (purple) and activated vs WNK1 without, in an inactive state (teal). Most of the blue waters would be gone in the purple state as well. This shows the significant structural transition, particularly in the helixes above the active site, which induce (in the purple state) dissociation of the dimer, auto-phosphorylation, and activation.

Thus there is a delicate network of water molecules tentatively held together within this protein that is highly sensitive to the ambient water activity (aka osmotic pressure). This dynamic network provides the mechanism by which the WNK proteins sense and transmit the signal that the cell requires a change in ionic flows. Generally the point is to restore homeostatic balance, but in the kidney these kinases are also used to control flows for the benefit of the organism as a whole, by regulating different transporters in different parts of the same cell- either on the blood side, or the urine side.