Saturday, August 15, 2026

Sensing Unfolded Proteins

Do cells take joy in tidying? Absolutely! But how can they tell where the messes are?

Any household knows that without cleaning and disposal, nothing else is possible. Things pile up, messes accumulate, everything comes to a standstill. Cells are the same, needing a constant flow of new materials in, trash out, recycling, and cleanups. One of the major kinds of mess in cells is unfolded proteins, which when they accumulate can cause aggregations and diseases like Huntington's, other dementias, atherosclerosis, and inflammation generally. While there are a lot of mechanisms (chaperones, chaperonin cage complexes, proteasomes) dedicated to proper folding, sometimes they are not enough, and excess junk in the form of unfolded proteins builds up. One major response this is the unfolded protein response, or UPR, which is centered at the endoplasmic reticulum (ER).

Why the ER? This is where countless ribosomes dock to feed in nascent proteins destined for secretion, the plasma membrane, and many locations inside the cell other than the cytoplasm. So it is a key organelle for protein synthesis and distribution, and thus for sensing how cellular proteins are doing. This sensing, the UPR, starts with a sensor protein, called IRE1 (also called ERN1), which sits on the inside face of the endoplasmic membrane. IRE1 binds to unfolded proteins inside the ER, and when the concentration rises high enough, it gathers into condensates of its own that trigger a variety of responses that include, in the extreme, cell death. Yes, lack of timely tidying can have serious consequences!

IRE1 is multifunctional, containing a kinase that activates AMPK and JNK, proteins which influence metabolism and inflammation. It also contains an RNAse enzymatic activity, which activates a whole other program by altering the mRNA of XBP1, which then becomes active and produces a transcription regulator that starts building more ER, more chaperones, and more protein folding capacity generally. The RNAse also digests other mRNAs indiscriminately, lightening the load of translation and thus incoming nascent proteins. So IRE1 is a key homeostatic regulator that adjusts the size of the ER to cellular needs, and informs the rest of the cell how things are going in the ER. 

One structure of IRE1, luminal domain only. This is a dimer, with halves shown symmetrically across the middle. How does this bind unfolded proteins?

But how does IRE1 sense unfolded proteins? A recent paper went through a modeling study to show how IRE1 operates. Sitting in the ER membrane, the enzymatic, signaling parts of the protein are outside, while the sensor is inside, called the core luminal domain (cLD). The authors demonstrate that this domain binds unfolded proteins directly, using them as bridging elements to bring multiple IRE1 proteins together and form the glob that nudges aside the inhibitory protein BiP and promotes the phosphorylation activity and thus activation, of IRE1. BiP also binds to unfolded proteins itself, so forms another sensor that indirectly activates IRE1 under stress conditions. 

Firstly, IRE1 is a dimer, and the luminal domain is where it dimerizes. The structure (just of the luminal domain, above) shows a beta ribbon core and peripheral unstructured areas and alpha helices. After running the crystal structure (shown) though very brief molecular dynamics simulations, the unstructured areas flop around dramatically, while the core and dimeric structure remain stable. Intriguingly, the beta stranded center of the dimer resembles (between the blue helices, above) the peptide binding cleft of MHC molecules- those that present random antigens to T-cell receptors of the immune system. Yet the authors show that unfolded proteins bind differently to IRE1. They flop across the whole central region, attracted by negative charge as well as the open hydrophobic surface. The next image shows molecular dynamics before/after shots of various unfolded peptides known to bind IRE1 in vitro, showing how after starting straight across the central region, they each find different comfort zones, mostly over this central region, but some off to a side. 

Collection of IRE1 structures (gray) with unstructured peptides of various kinds bound to the luminal face. t=0 is the starting point for the molecular dynamics simulation, which went for one microsecond. End points are below. In each case, the unfolded peptide stayed bound, and settled into some comfortable position, based on the hydrophobic and negatively charged landscape of the IRE1 face. 

The intriguing thing about this binding is that the structure of IRE1 itself is unaffected. The dimer interface remains stable and open, leading to the hypothesis that these peptides form bridges by becoming the meat of a two-dimer IRE1 sandwich. And it is this that then displaces the inhibitory protein BiP and leads to condensation and activation, by bringing together multiple IRE1 kinase domains (on the outside of the ER), and activating their kinase/signaling functions, turning on the various downstream pathways.

A helpful model from another lab, describing how, once the unfolded proteins are bound inside the ER (by mechanisms unknown at the time), the outside domains would gather, trans-phosphorylate, and activate downstream pathways. 


  • Those flock cameras.
  • As if charging for previews of presidential posts wasn't corrupt enough.

Saturday, August 8, 2026

Our Loopy Way of Learning

Different ways of learning go through different parts of the brain.

Humans are champion learners. Other animals are smart, but we are smarter, spending more time in childhood soaking up the mysteries of the world around us, and storing them in a larger and better brain. Learning is our calling card, allowing us to adapt to any environment and defeat any foe. Indeed, we have overwhelmed the earth's biosphere, and need to exercise some deeper and longer vision by pulling back from our successes in mastering every possible resource.

But how the brain does all this is something we have yet to fully learn about. Our sense organs bring in tons of information, but as we have learned in computer science and AI, it takes a lot more to create usable information than just amassing data. How do the various parts of our brains work together to create actionable models of the world out of that data? An innovative theory was elaborated a quarter century ago by Kenji Doya, who has apparently gone on to a research sideline on soccer(!)

The problem began with, in part- what does the cerebellum do? From lesion data, it has long been clear that this small part of our brain has a big role in motor accuracy and coordination. But cerebellar neurons project all over the brain, not just to motor areas, and it gradually became clear that the cerebellum affects those other areas as well, helping us learn and manage many cognitive tasks. But structurally, the cerebellum is a peculiar organ, highly repetitive, with relatively linear and parallel processing, a bit like a GPU. It is clearly specialized for something- some kind of fine tuning, but what?

Doya's theory about the distinct learning styles contributed by different parts of the brain.

Doya proposed a general theory about learning in the brain, which splits learning up into three types. One type is supervised learning, which is what the cerebellum facilitates. We grab something, we recognize an error in eye-hand coordination, and we learn to grasp better. There is a loop involved, both physically and logically, by which some task is evaluated, and error signal sent back to the source, and a small adjustment is made for the next iteration. Humans have great hand coordination, partly thanks to our advanced cerebellar-cortical connections. 

The second form of learning is reward-based learning. This prototypically features dopamine, and is centered in the basal ganglia, where emotions are processed, but which is also a major switchboard for connections upwards to the cortex and downwards to the brainstem, and is also critical for motion control. Parkinson's and Huntington's diseases both affect this region. Doya proposed in broad terms that this area doesn't do the kind of error correction that the cerebellum does, but rather learns by reward- whether the action's goal was reached. If the goal was reached, a spurt of dopamine is issued, which reinforces whatever connections helped that action happen.

Interestingly, there is a progressive temporal aspect to this learning. As an action is learned, the reward happens sooner, given the predictive capability of our learning / cognitive systems. The whole point, after all, is to figure out as early as possible what will be happening in the future. If a bell is associated with food, at first the food prompts the reward. But as learning happens, and the bell is reliably associated with the later food, and the bell starts to set off the reward all by itself. Gradually, the sight of the person getting about to ring the bell, or the expected time of day, sets off the reward... it is cast progressively back to whatever the perceived cause is, going back in the learned causal chain. Eventually, as everything becomes rote and the novelty wears off, reward lessens, and learning is complete. Mealtime is here, as it is every day... boring.

Lastly, the third type of learning is unsupervised learning. This is the job of the cortex, and is the most interesting and subtle form of learning. The world has patterns, and the cortex's job is to figure out what those patterns are. No one outside tells it what is right or wrong, what is or is not real. It has to figure all this out from the stream of input, including its own blundering actions and experiments. Unsupervised learning is thought to be well suited to the basic mechanism of the brain, its Hebbian circuitry where activity reinforces any active connection. There are a lot of mathematical methods that have been devised to approximate such processes, like principal component analysis. Data can be binned / categorized in various ways to simplify and organize reality, and eventually we have ... language and concepts, labels and abstract understanding. This categorization of the world is immensely powerful, if always inaccurate and approximate. 

Doya's model of the cerebellum. Purkinje cells at the center are also conceptually at the center, as the only cells that are plastic in this scheme, able to change their connection weights due to input from the other layers, which in turn get input from other brain areas including the cortex and thalamus. This has been likened to a perceptron in computer science terms.

Doya thus created an influential heuristic that divided up learning styles / mechanisms by anatomy in the brain. Obviously, all these mechanisms work closely with each other. The cerebellum doesn't know anything without conceptual categorizations coming from the cortex, which provide the foundation of error signals. Similarly, whether goals have or haven't been reached has to be informed by conceptual data from the cortex (or food signals, or other sensations from elsewhere in the body). Learning is a whole-brain activity, and while it is prodigious and effortless early in life, it continues to be essential throughout, lest we be overwhelmed by the fresh challenges the world keeps presenting to us.


Saturday, August 1, 2026

Ohnologs and Paralogs: The Wages of Gene Duplication

Two whole genome duplications lie at the root of vertebrate evolution.

Another week, another story about the power of gene duplication in evolution. While rare during normal reproduction, gene duplication happens pretty frequently over longer time scales. How else would we get a thousand olfactory receptor genes, all similar to each other? But other accidents can occur as well, like whole chromosome duplications (such as what leads to Down syndrome), and whole genome duplication, when the cell division process stops early, but otherwise proceeds with cells remaining viable with double the genomes as before. This is common in plants. Corn is tetraploid, wheat is hexaploid, and strawberries are octaploid. 

But among animals, whole genome duplication is less common. Two decades ago, however, two researchers working from the newly sequenced human genome revealed that there were two such duplication events at the beginning of vertebrate evolution, explaining some oddities and also perhaps the speed and power of subsequent evolution. The basic evidence is the genome sequence, which is full of related genes. Genes that do the same thing in various species, and are lineally related, are called orthologs. That is relatively simple, per the Darwinian tree of descent of all life. Genes within one organism / one genome that are similar to each other due to ancient duplication events are called paralogs. All those olfactory receptors are paralogs, for instance. Lastly, genes that are paralogs stemming from a whole-genome duplication event are called ohnologs, in honor of Susumu Ohno, who led the field of molecular evolution in recognizing the importance of gene duplication, and speculated about whole genome duplication well before it was discovered.

A classic example of this evidence is the hox cluster, a linear sequence of genes that have been extensively studied in flies as providing an important set of regulatory controls over the linear body plan. They lie in the middle of the developmental cascade, downstream of egg and body polarity genes, but upstream of specific appendage and tissue expression programs. They encode DNA-binding (homeobox) transcription regulators, and their position in the genome is co-linear with the body parts they activate because there is a progressive chromatin opening process by which this whole locus becomes activated. Well, flies have one hox locus, but vertebrates have four. What happened?

Hox loci across evolution. Where flies and primitive chordates have one hox cluster, vertebrates have four. How did that happen?

Obviously, once researchers lined everything up, it became pretty clear that there were two massive duplications along the way, creating four hox loci in the genome, after which quite a few of the duplicated genes fell away. After a duplication event, gene survival is a race between neo-functionalization (which leads to preservation by selection) and deleterious mutation, degradation, and disposal. Enough of these ohnologs survived to help fuel the substantially greater complexity of the vertebrate body plan, now including intricate wrists and hands, and ever more involved head structures.

Similar findings were made all over the human genome. The original paper has a graph that shows that, across the genome, most paralogs exist in families of four. If genome duplication were not the applicable hypothesis, then one would expect a smooth asymptotic curve downwards from one member (implicit) to two, then three, and fewer from there outwards, since single gene duplications would each be independent statistical events. But no, there is a peak at four, indicating that some process yoked together many, many genes into parallel duplication events, twice in succession. 

A graph of count of paralogs vs their frequency by count, in the human genome. What should in principle have been a smoothly declining curve from 1 or 2 turned out to have a weird peak at the number four. This was important evidence that many of our paralogs arose through some common event, such as a pair of ancient whole genome duplications. 

OK, so far, so old hat. A bunch of more recent papers flesh out this story a bit, showing how the vertebrate duplications were timed, and how they affected various types of genes. miRNAs, for instance, turn out to be ancient genetic elements, and were duplicated along with everything else. miRNAs that originate from (and were preserved from) the whole genome duplications tend to be more conserved, have more targets than other miRNAs, are more highly expressed, and have higher rates of targeting RNAs of transcription regulatory proteins that likewise originated from whole-genome duplication. The traces of this history are thus interestingly preserved.

Another paper traced the effects of the vertebrate ohnologs on brain development. Using the pre-vertebrate amphioxus as an out-group for comparison, they find that the ohnologs that date from the whole duplications are more highly associated with brain cell types and their developmental programs than are other gene duplications- either ones dating from the same time as the whole genome duplications, or since.

"Compared to their closest invertebrate relatives—tunicates and amphioxus—vertebrate brains are highly regionalized and complex."

"Ohnologues were enriched in development, cell-fate commitment, signaling and neurotransmitter transport. By contrast, SSD [small-scale duplication] paralogues were enriched for immune response and sensory perception in all species, a result that matched previous reports."

Lastly, a paper focusing on intermediate genomes, part of a plethora of genome sequencing that has happened since the original analysis, reveals exactly what happened at this time, about 500 million years ago. From a molecular perspective, hagfish are in the same group as lampreys- both parasitic eels that lack real jaws but arise from the vertebrate lineage. They are cyclostomes, while we are gnathostomes. After the first whole genome duplication about 530 million years ago in the common stem lineage, the gnathostomes and the cyclostomes split and each experienced their own, separate genome duplications roughly 490 million years ago. This can be concluded from the differing gene collections that survived from each respective duplication, and their sequences relative to the cyclostome/gnathostome split. Indeed, lampreys and hagfish have six hox clusters instead of four, indicating that a triplication event happened, instead of duplication event, between their split from gnathostomes and the split between lamprey and hagfish. 

A phylogenetic tree with dates, locating the genome duplications deep in the vertebrate stem lineages. 1R is the first genome duplications, common to all vertebrates. 2R is the duplication in the lineage leading to jawed vertebrates, while CR is the apparent triplication that happened in the lineage leading to cyclostomes. Apologies for the antiquated human being. 

All this genome renovation didn't benefit cyclostomes the way it did gnathostomes, however. Only one group went on to globe-straddling glory, indicating that while genome duplications may be helpful for evolutionary / developmental innovation, they certainly aren't determinative. They are grist for an evolutionary process that remains largely shrouded in mystery- the context and environments of the time, the competitive biosphere, and the internal molecular environment. There is no reason to think that all this is unaccountable by natural processes, but that doesn't mean we have the information to reconstruct it in detail. We should thus be deeply grateful that the scientific community can dig up even this much of our history.


  • But how can we take more money away from workers?
  • The smearing of Anthony Fauci. And of basic reason.
  • Reward and addiction.
  • Still stupid ... the Discovery Institute and "information".