Saturday, July 25, 2026

Revealing the Origin of Eukaryotes

Tracing the breadcrumbs of deep phylogeny through gene duplications. 

Last week, we discussed the jamboree of mutation that is the MHC locus in animals. One part of that story was the persistent duplication and pseudogene-i-zation (which is to say, the birth and death) of genes in that locus through primate evolution, another consequence of the never-ending arms race with our pathogens. Gene duplication has much deeper significance, however, as it has given rise to new genes throughout biological evolution, as it is merely an accident of DNA (or RNA) replication, thus operationally rather frequent. Significant transitions, such as the advent of eukaryotes, and the advent of plants and of angiosperms, all feature rapid gene duplication, even whole-genome duplication, which provide grist for new functions, refinement of old functions, and speciation.

A recent paper did a deep dive into the origin of eukaryotes, perhaps the watershed event in the history of life, second only to the origin of life itself on the early Earth. It is very difficult to reconstruct what happened during these extremely early events, which these authors put as long as three billion years ago. There has been so much mutational churn in our molecules, and no fossil record to speak of, that, again like the origin of life itself, we do not have a great deal to go on. Thankfully, the discovery of the kingdom of archaea, and its especially eukaryotic-like class, the Asgard archaea, has provided a much better platform for speculating about what the originating organisms might have been like.

Eukaryotes differ from bacteria and archaea by numerous characteristics, implying extensive development/evolution before the last common ancestor from which we can trace all the extant descendants. These include the eponymous nucleus, an internal system of membranes and membrane-bound organelles, a complex cytoskeleton of both actin and microtubules, a mitochondrion, meiosis, a microtubule spindle-driven division system, and many other, more obscure molecular differences. The lack of intermediate forms is certainly frustrating from a scientific perspective, making it difficult to go through the catalog of life today to identify the steps involved in this transition. It also implies that the process ended with an evolutionary bang that in essence caused the final iteration to wipe out all the preceeding forms, a bit like humans vs the many other Australopithecines and their descendents. 

These differences are enormous and momentous, making eukaryotic cells far larger, energetically capable, and complex than their antecedents. Another key characteristic is a prevalence of gene duplication and specialization, leading to much larger genomes along with all the other complexity. It has long been theorized that the union of the proto-mitochondrion, which was a bacterium, with the archaeal founding cell that set the whole process off, particularly providing the massive amount of energy required for all this rise in complexity. However the current authors claim definitively that this is not the case, that the mitochondrial symbiosis was a rather late event, coincident with the oxygenation of the atmosphere about two billion years ago (GYA). The authors argue that, since gene duplication and specialization is in any case a clear characteristic of eukaryotes, using molecular clocks to time these duplications can give us the epochs at which the processes they participate in first developed. 

The authors present estimated times of origin for key duplications involved in RNA synthesis. The RNA polymerases I, II, and III were duplicated from one polymerase in archaea, and are composed of numerous gene products, as well as ancillary regulatory proteins. While some subunits are still shared, the ones that duplicated did so in the range of 2.3 to 2.8 billion years ago, before the symbiosis with the mitochondrial precursor, timed at "mFECA", or the mitochondrial first eukaryotic common ancestor, about 2.2 billion years ago. "nFECA" is nuclear first eukaryotic common ancestor, while "LECA" is last eukaryotic common ancestor.

For example, eukaryotes all have three RNA polymerases, while bacteria and archaea only have one. These RNA polymerases specialize on mRNA synthesis (RNA pol II), ribosomal RNA synthesis (RNA pol I), and tRNA, 5S RNA and U6 snRNA (RNA pol III). Why this specialization exists remains a bit mysterious, though once it started there was no going back. In any case, it happened and is totally characteristic of eukaryotes, so these duplications (which involve the duplication of numerous genes encoding proteins of the polymerase complex itself and its regulators and way-finding helpers) can be used to date eukaryogenesis, using a very carefully calibrated and species-sampled molecular clock. And what they find is that all of these duplications, given as ranges in the diagram below, with green dots at the likeliest time point, all lie prior to the point of "mFECA", which is the mitochondrial first common ancestor, at roughly 2.2 billion years ago. Tabulated over studies of almost a hundred other genes that are similarly duplicated and characteristic of eukaryotes, the authors place the "nFECA", or nuclear eukaryotic common ancestor at almost 3 billion years ago. 

That is a very long time! It is a very long time ago that the characteristics of eukaryotes started to come together, and that is also a very long time (well over a billion years) over which that development spanned before culminating in the last common ancestor of all extant eukaryotes, which the authors put at about 1.7 billion years ago. And note that that still comes a billion years before the cambrian explosion of animal life. We are talking about really deep time here. 

If mitochondrial symbiosis came later, (causing another bout of genomic change as hundreds of bacterial genes moved from the symbiont to the nucleus), just as the atmosphere was getting oxygenated, what was going on in the preceding billion years? One can speculate that these evolving proto-eukaryotes were the top predators of their world, a bit like eukaryotic protists in microbial environments today. The competition would as always have been intense, but there were no higher life forms (i.e. animals or protists) to worry about. Size was not a problem, but rather an advantage. The focus was on quality and effectiveness in finding, disabling, and digesting prey. Thus there was little downside to complexity, unlike the constraints on prey species, which had to optimize for rapid growth and reproduction under both nutrient and predation constraints. New and costly systems for protein secretion, for prey digestion, and intelligent, responsive regulation of all cell processes might have been consistently advantageous.

I take this work as being definitive, at least in terms of setting the debate about early vs late symbiosis. It benefits from a clear theoretical basis and new, vouminous data from appropriately related species in its molecular comparisons. I recommend it highly, as I can only touch on its analysis. The advent of mitochondria was then a later step change in complexity, but brought incredible metabolic productivity that, at least in oxygen-rich environments, would have definitively superceded all the prior forms of the proto-eukaryote, and (unfortunately for us) erased that record of evolution. Another later innovation, between mFECA and the last eukaryotic common ancestor "LECA" was meiosis

One more sample from the author's presentation of molecular duplications that help time eukaryogenesis. Here, DNA repair proteins are emphasized, including some involved in meiosis (light blue dots, see legend). Some of these duplications involve genes from the mitochondrial endosymbiont, (MSH4), and some (such as EME1/2 and MUS81) come from older repair processes, but later specialized for meiosis. 



Saturday, July 18, 2026

MHC Through Evolution: Breaking All the Rules

The immunologically critical MHC gene cluster plays by its own rules through the arms race of life.

Last week, I discussed the very general landscape of variation in the human genome, specifically the tradeoff between prevalence in the population and effect size. Given that the vast majority of variants are deleterious, those that affect our phenotypic traits are more heavily selected against the greater effect they have. The result is that at a gross level, most genes and most traits share the same general distribution of lots of variants (alleles) with minor effects, and far fewer with large effects. And those with minor effects also turn out to be tangential for biologists, rarely informative about the nature of the traits they (sort-of) affect.

This week, another paper and another view of evolution, though the eyes of one the more critical genes of the immune system, the major histocompatibility complex, or MHC group of genes. While our adaptive immune system has developed the extraordinary and powerful ability (though semi-controlled DNA recombination of the antibody and TCR genes) to recognize practically any antigen, foreign or domestic, that system requires stringent controls. One of those controls on T cells, which carry the antigen-recognizing TCR receptor, is that it can only see antigens that are "presented" on MHC molecules. MHC proteins have a surface cleft that gets loaded with and holds small peptides (8 to 12 amino acids long) that are cleaved from other proteins, either from pathogens or from the cell itself. The MHC+peptide complex then sits on the surface of the cell, announcing either that 1) I am healthy, full of normal cell proteins, going about their business, nevermind, or 2) I have some other proteins inside, either from a viral infection (class I MHC) or from some bacterium I have just phagocytosed to deal with an infection (class II MHC). In the second case, T cells carrying the TCR receptor lock onto the MHC+peptide complex, and start up the process of killing that cell. 

MHC proteins (beige, pink) present an antigen (red) from one cell, and dock to a T-cell which recognizes the antigen+MHC complex by shape, using the TCR receptor. The additional CD8 or CD4 receptors help to verify the proper binding. While class I MHC are present on all cells, class II MHC are on phagocytosing cells like dendritic cells and macrophages that commonly ingest, reprocess, and re-display bits of encountered pathogens.  


Given that the TCR gene recombination process is unbiassed and produces a galaxy of random binding specificities, how do these cells distinguish self from non-self antigens? This is a deep question that is not fully resolved. But one major mechanism is thymic selection, which is what gives T cells their name. Special cells in the thymus display a wide range of self-antigens, and T cells, which are obliged to pass through the thymus during their maturation, are induced to commit suicide if they react to any of them. A paper from 2018 fascinatingly discussed how it is possible to create a T cell population that knows the "language" of foreign vs domestic after what is known to be a rather haphazard selection process, which displays only a partial range of self-antigens, and leaves quite a few self-reactive T cells around.

At any rate, the MHC proteins do not benefit from hyper-variation provided by genetic recombination. Yet it turns out that variation is beneficial here as well. The way foreign antigen peptides nestle in the MHC groove can be varied by mutations in the MHC molecule, providing a rich field of variation in antigen recognition and thus disease resistance. So, our MHC genes have not just a few alleles in the population, not just a few dozen, but over six thousand alleles. For each individual MHC protein, each person has only two, but over the population, there myriads with different properties. MHC was first recognized for its role in self vs non-self recognition and transplant rejection, (thus the "compatibility" in its name), and it quickly became evident that people vary tremendously in their MHC complement. And this variation plays a big role in keeping us (and all other animals) going as populations, in the face of pathogens that evolve a lot faster than we do. 

A recent paper provided a phylogenetic history of MHC molecules in monkeys, covering the last sixty million years of evolution in our lineage. It is a festival of gene birth, death, and duplication, quite apart from the smaller mutations that are constantly accumulating and cycling through the population. The MHC region carries about 200 related genes, most of which have minor roles, and only six of which (three MHC class I, and three MHC class II) follow the high-mutation pattern because they encode the main antigen presenting proteins. These genes are subject to, quite obviously, unique selective forces. 

How the MHC gene cluster looks, when aligned and identified by gene, over the primates. Note the deep divergence between the new- and old-world primates. Genes A, B, and C are the major MHC class I genes, which vary the most over this time. Note also how some of these genes have gone through extensive duplication in some old-world monkey lineages.

The first force is balancing selection. As soon as one allele becomes common, pathogens evolve to evade its presentation skills, rendering it less effective and less desirable. This results in a population full of minor variants. Indeed, for any individual person, having two MHC molecules that are the same would be bad. Being heterozygous at these genes is highly advantageous, thus enforcing both the retention of minor alleles, and an observed behavior in mating to favor partners with different MHC complements. Apparently, our MHC makeup is reflected in our personal aroma! 

A second force, conversely, is the retention of ancient alleles. It turns out that, across the old-world monkeys, many MHC alleles are preserved and cluster more closely in sequence comparisons with each other than they do with other alleles in the same species. That is, despite the general speed of MHC evolution and constant accumulation of new alleles, old alleles are preserved in all monkey populations as well, due to their distinct capabilities, under balancing selection. This is part of what makes population bottlenecks so damaging to near-extinction species. They lose critically valuable genetic resources (in the form of rare MHC alleles) that represent millions of years of accumulated variation. 

Incidentally, the trees shown here again reinforce the history of primate evolution, with new world monkeys splitting off from the old-world monkeys quite early on and developing a very distinct set of MHC molecules.

So, while virtually every other gene in the genome is being relentlessly optimized, sticking to its knitting, doing one thing and being beaten down whenever any mutation steers it from its optimized path, the MHC genes follow quite a different path, at least in portions of their sequence that provide variation in antigen binding and presentation. These genes revel in endless diversity, throw off pseudogenes at a high rate, wink out of existence and come back in other forms. Natural selection is the motor in each case, but meets the challenge of survival in different ways.


Sunday, July 12, 2026

How Do Complex Human Traits Add Up?

Notes on the genetics, traits, and evolution.

Everything about us is a trait. Not everything about our traits is genetic, though. The conundrum of nature vs nurture, of genes vs environment, and the structure and meaning of genetics goes to the heart of biology. A few traits, like eye color, are simple enough. But they are the exception, by far. Body mass index is influenced by practically every gene we have. And autism has, by this point, hundreds of contributing genes. Both traits are highly heritable, in the sense that inheritance/genes are the dominant influence, vs environment (as seen in twins) when most conditions are equal. But environment can easily be dominant over BMI when conditions change and starvation sets in. 

A puzzle that came out of the early human genome studies was how unhelpful it was to do genome-wide association studies (GWAS) to approach some of these questions- that is, studies of what variants in the population at large contribute to particular traits, especially to serious diseases. Study after study was done, and disappointment mounted that what were found were genes with minor effects, in tangential biological processes. This was supposed to be the holy grail- the payoff for sequencing the human genome- and what came up was dud after dud.

A recent paper plows over this ground again with a new mathematical synthesis of genetic structure of human traits, genetic alleles, and selection. What it finds is sort of obvious, but there are some intriguing observations along the way. It is critical to note at the outset that evolution as Darwin understood and described, and natural selection in particular, is absolutely at the heart of this or any contemporary analysis of genetics. While Darwin's understanding was revolutionary and broad, subsequent decades of work have brought these concepts to a very concrete, operational, indeed mathematical, level. 

Let's start with the concept of allele frequency, also called minor allele frequency, of MAF. Over any individual genome, there will be millions of "variants", which are coding differences from the reference genome. Which does not have any special status... it is just the genome of some guy from Buffalo. Variants (or alleles) are bases in the genome that differ from the reference. With three billion positions and four possible bases per position, that means that there are nine billion possible variants. How common is a particular variant in a population? That is its allele frequency. For GWAS and related studies, the threshold is commonly set at variants seen in the population at over one percent frequency, while minor alleles are seen under that frequency. A variant that causes some devastating disease is typically one that sprang up recently, and is heavily selected against. That is why it must have an exceedingly low allele frequency. On the other hand, a variant may have no discernable effect at all, not being selected for or against, thus just drifts along in the genome, not subject to natural selection. Such alleles may, over long periods of time, drift to higher or lower frequency by random chance.

However, the focus of GWAS association studies are variants between these extremes. These are variants that have some effect on a trait (or may be physically close to others that do, thus get "carried" along over time). At the same time, they are also common in the population, at least common enough to be discernable in an association study. Such a study needs some statistical correlation between the occurrence of the variant, and the occurrence of the trait. That means that the variant can not just appear once, but must appear many times over a large population. At the same time, a study that focuses on a trait like, say, high blood pressure, will be seeing variants that are, by definition, deleterious. That means that any allele with a large effect will be subject to strong selection, and driven out of the population. Only alleles with more modest effects will be able to survive at all, and even then, at low frequencies. So a GWAS focuses on medium-to-low prevalence variants, hunting for alleles on the loose in a large population that have typically modest effects on a given disease or other trait. Such alleles will have typically survived for tens or even hundreds of thousands of years, so they will have some complex relationship to natural selection. 

In contrast to all this is the family study, which focuses on a dramatic variant that causes some terrible disease. Such studies have been remarkably productive, because they deal with extremely rare, high-effect variants, which are as a rule very informative about the genesis of that disease. Such variants, as mentioned above, would be heavily selected against, thus disappear rapidly. But mutation is always happening, so all sorts of mutations arise in large enough populations. Assuming that, as biologists, we are interested in the core ten or fewer genes that most influence a given trait or condition, the hundreds or thousands of significant, but low-effect variants that come out of GWAS are almost by definition guaranteed to be tangential and minor. It turns out that most traits are complex, in the sense of being influenced in various minor ways by hundreds of genes.

So, what is the typical genetic structure of complex traits? That is- what this paper set out to answer. "Structure" in this case means ... what is the normal distribution of selective target / effect size versus frequency/prevalence in the population of variants that, in combination, add up to a complex trait? The assumption (as discussed above) is that the core armature of such traits does not have variants at all, due to strong selection, while the available variants in the population all have minor effects that in sum form the genetic variation seen in the trait in the population. 

While other researchers have attempted to fit the variation distribution of complex traits to typical formulas like the normal distribution, these authors found that a natural selection-informed approach gave a clear and simple result. All traits follow the same general scaling, with only two parameters- the mutational target size of the trait (that is, the proportion of the genome capable of appearing as relevant variants), and the effect that each site has on the given trait, termed (very poorly) the site's "heritability". It is important to note that every site in the genome is equally and fully heritable. The term refers to the trait, and the size of the effect from variations of that site on that trait. Summed over all sites in the genome and all variations in the population, this heritability ultimately equates to the overall variation of the trait that is genetically caused.

A comparison of two traits, and how they might look in a genetic variation study. In blue is a trait skewed towards small effect variants, with weak selection and consequently variants with higher frequency. In red is a different trait that partakes more from stronger effect variants. On the whole, this kind of difference is not common among complex traits that arise from thousands of loci. MAF = minor allele frequency; Z-score is the score in a GWAS study indicating how statistically significant the variant's effect on the trait is. Note how lower Z-score correlates with more variants at the higher MAF frequencies. At the same time, the GWAS method overall has some skew to higher MAF frequencies, since only those provide sufficient statistical power to get any results at all. Log(s) is the strength of selection; L is the genetic target size for the whole trait, and h*2 is the heritability, or proportion of the trait effect due to the causal variant.

The model they come up with accounts for the selective effect of trait effects (the larger the effect of the variation on traits, the lower its frequency in the population). It also accounts for the fact that variants that affect one trait often affect other traits as well, so the selective effect needs to be considered over all affected traits, most of which are probably unknown, but can be inferred. And conversely, most traits are composed of contributions from many genes and their alleles, sometimes thousands- they are genetically complex traits. 

The model the researchers come up with can normalize among the huge population of variants that affect a single trait, in this case blood pressure. Left shows the effect sizes of individual variants, and right shows the scaled (normalized) version from the paper's model, showing that all these variants follow the same overall rule / logic, using the custom parameters of h*2 and L- trait heritability and mutational target size. All this is to say that the lower effect variants (skewed to left) are assumed to have higher selection coefficients.

The researchers go on to show that various traits do look different under this analysis. Some differ mostly by target size, accounting for more or fewer variants, but having a similar spread of effect sizes over the population. Others differ by the scale of effects that each variant contributes, thus skewing toward higher or lower allele frequencies overall. Interestingly, they add an analysis of the age of these low-frequency, low-effect variants that are the grist for GWAS, finding that they are on average 137,000 years old. That compares with an average age of 600,000 years for variants that are neutral, thus would not come up in GWAS or be under selection. This is fascinating in its implications both for human population genetics in general, and for the fact that most human variation- even that under modest selection- predates the divergence between African and non-African populations. 


Saturday, July 4, 2026

Performing Search, as a Transcription Regulator

Billions of years have created some weird tricks in DNA search.

Search is all around us, as we increasingly rely on search engines to find everything we need on the internet, want to watch, or want to buy. Search looks into databases, which hold the sought-after information. All our accounts, all the domain names, all the products... everything is held in databases of one kind or another, and those databases are indexed in clever ways to provide virtually instant pointers from the question we ask to the answer held online. AI merely puts a linguistic gloss on this, and most people are still encountering AI first as a feature of search, such as the top of current Google search results.

Well, our genomes are databases as well- rich and ancient storehouses of jewels that encode the body and its doings. How does search work there, and what is search even for? At any moment, each cell of the body has certain needs, stresses, and goals, as expressed in its DNA programming. The tools available are proteins and RNAs, which carry out the cell's functions. The needs may arise from signals coming from previously expressed receptors, say, for insulin, which may trigger and tell the cell to take up glucose from the blood. The receptor turns on a kinase, which may turn on another kinase, which turns on a transcription regulator, which goes into the nucleus and ... does a search. This regulator is searching for places (specific sequences) in the genomic DNA where it can bind, after which it helps to turn on (or off) the nearby gene, executing the desired function / tool. 

General introduction to transcription regulators (or "factors") and their role in gene activation and the whole process of gene expression.

Obviously a very different kind of search than what Google does on our behalf across documents, but there are similarities. Internet search depends on patterns, matching the user's input with the vast corpus of the internet also held as text symbols. Transcription regulators match patterns, in this case patterns of DNA that they like to bind, which may occur only once in the genome, or occur tens of thousands of times. The pattern here is a complementary physical/electrochemical shape, rather than an abstract same-symbol match. The genome is, to a protein, truly vast. Our three billion-base genome is forty million times larger than an average regulatory protein of, say, fifty kilodaltons (kDa). Search is also, here, a difficult problem, which researchers have been wondering about for decades. Several recent papers discuss different aspects of the problem and shed some modern light on it.

We have roughly 1600 transcription regulators in our genomes, so there is something going on all the time. DNA is always being queried. And what it replies with is RNA- a transcript issued/copied from a gene, which either goes off to instruct creation of a protein, or is itself functional in some way. So, how do proteins bind to DNA, executing their search? It was transformative when the first atomic structures of such proteins were solved. They were clearly complementary with their DNA targets, with nicely positioned positive charges to mate with the backbone of the DNA and amino acid fingers reaching into the helix to feel the shapes of the nucleotides they wanted to bind. All very neat, and paradigmatic for bacteria whose genomes are quite small. But there is more to the story. Binding sites in human genes tend to be quite short- five to seven bases. That really isn't enough to be very specific, across a vast genome. Eukaryotes have developed several weird tricks, as it were, to encourage efficient search over much larger genomes and at the same time increase precision while maintaining evolvability and flexibility.

Eukaryotes have nucleosomes, chromatin, and packaging. The DNA is not just splayed out randomly, but wound up on protein spools. One would think that this would impair search by regulators. But paradoxically, there is a fine balance between hunting around on a given piece of DNA for a preferred site (one-dimensional search, 1D), and jumping off, letting go, and trying somewhere else (by diffusion; three-dimensional search, 3D). The compaction of genomic DNA into nucleosomes that wind up most of the DNA while leaving linking DNA in between free appears to provide a nice balance of landing spots that allow searching regulators to jump very long distances (in linear DNA terms) while not going very far in absolute terms. Regulators vary in how aggressively they can plow through nucleosomes to try out their internal DNA sites, but many (called pioneer factors) can do so.

Secondly, transcription regulators cooperate with other proteins to create longer, more complex DNA sites for precise gene identification and higher binding affinity. As biologists have characterized the enhancers and promoters of important developmental genes, they have found that DNA binding sites occur in bunches, and have much weaker effects when broken down and separated. Sometimes there is direct side-to-side cooperation between two regulators that bind the DNA. At other times, they combine with other non-DNA-binding proteins to create complexes at such sites. The DNA recognition sequences of these combinatorial sites can be changed significantly, even beyond (our) recognition, by the addition of cooperative proteins. This is something that makes prediction of where a given regulator binds particularly perilous. 

Thirdly, many regulators contain not only DNA binding domains, but also extra disordered domains that facilitate DNA search and binding. This has been a recent realization that accounts for some of the speed and flexibility of regulator search and DNA interaction in eukaryotes. The stable crystal structures of paradigmatic bacterial regulators are not the whole story, and indeed are insufficient to explain what is happening in the much larger setting of our own cells. The authors note that eighty percent of human gene regulators have large disordered domains, (called IDRs, for intrinsically disordered region), upwards of 500 amino acids long. These never showed up in crystal structures, naturally. Being disordered, they are also poorly conserved. So, they have been difficult to study. 

Comparison of binding by one regulator, MSN2, which has a large IDR, to its genomic sites. At top is its native binding pattern, across a whole genome. At bottom are mapped its core motif occurrences on that DNA. Second from top is the MSN2 protein mutated to contain only its core motif-binding domain, and third from top is the MSN2 protein mutated to remove that domain and retain everything else. Note how different the patterns of binding by each of these proteins are, though how each approximates to some degree the wild-type pattern.

In related work, researchers have divided up such proteins into the core binding site part and the IDR part. They find that both parts work partially, directing binding to some of the native sites around the genome. In fact, the IDR part does a more statistically accurate job than the core DNA binding motif. This is fascinating, showing that in eukaryotes, a new search mechanism arose, supplementing discrete and precise binding with a floppy / fuzzy code in the IDR and its binding sites. It turns out that regions of hundreds of bases around core target sites (which in one case amount to only the motif AGGGG) are preferentially bound by the respective IDR protein domain, with multiple weak interactions that remain structurally uncharacterized. In fact, neither the protein structures responsible, nor the DNA sites they bind are known yet, though deletion studies through IDR domains show that binding is distributed throughout.

Relationship between IDR binding site size, and the ratio of 1D vs 3D search time, by simulation. The bottom axis is size of the IDR binding region, the Y axis is time taken for search. Time spent in total (yellow) goes down to minimum at an optimum between 1D search that is slowed by longer IDR-binding regions, while 3D search is strongly accelerated by longer IDR-binding regions.

The combination of core binding and loosely unstructured binding in one regulatory / search protein provides powerful benefits. In dimensionality terms, if the effective landing site is expanded from five to five hundred bases, then the time required for 3D search through the space of the nucleus is dramatically shortened. Secondly, loose binding by the IDR then promotes an "octopus"-like 1D search along the local DNA, resulting in efficient settling on the core binding site to get ultimately precise positioning. The ultimate affinity of the regulator with the local DNA is also enhanced compared to what it could manage over a five base pair site. The researchers conclude that with these domains, the search problem is, in net terms, reduced by one dimension, from 3D to 2D. The surrounding areas of DNA that have marginal affinity for the IDR domain are called "antenna" regions, and the author's simulations show how they alter search behavior.


Schematic explanation of the current work, describing how IDR domains help to speed up the transition from 3D search through space, to 1D search across the DNA. And then also to facilitate 1D search by preventing full detachment from the DNA while the core binding motif continues to search by diffusion for its binding site (yellow).

For computers and databases, search is a huge problem that has led to technical innovation, as well as large drains on resources. Every search engine combs the internet, gobbling up all available information, creating indexes, and updating them constantly in order to give us the instant access we want. This infrastructure has been raised to a new level by AI, which transforms search into a new form, combining it with language translation and prediction methods that allow a search for corkscrew to bring back results for wine. Whether it understands anything is unlikely, but the desire to upgrade search from a simply determinative process to one that is more fuzzy and richly interpretive, and thus more useful, is not a new phenomenon.