{"id":78142,"date":"2026-08-27T21:16:23","date_gmt":"2026-08-28T01:16:23","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=78142"},"modified":"2026-08-27T21:16:23","modified_gmt":"2026-08-28T01:16:23","slug":"pottsmpnn-protein-design-mit-78142","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/pottsmpnn-protein-design-mit-78142\/","title":{"rendered":"PottsMPNN Expands Protein Design Beyond Natural Sequences"},"content":{"rendered":"<p>Protein design has long been constrained by a paradox: the same three-dimensional structure can arise from countless amino acid sequences, yet most computational models for designing new proteins have been judged by how closely they can reproduce the single sequence that nature happened to select. A new machine-learning framework called <a href=\"https:\/\/github.com\/keatinglab\/PottsMPNN\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">PottsMPNN<\/a>, developed in the Department of Biology at MIT, directly challenges that metric. By incorporating the physical principles that govern protein stability and folding, PottsMPNN expands the universe of viable protein sequences far beyond those found in nature, opening the door to completely novel proteins that could be tailored for therapeutic, industrial, and research applications.<\/p>\n<h2>The Fundamental Problem with Current Protein Design Metrics<\/h2>\n<p>A protein\u2019s function is determined by its structure, and that structure \u2014 the way a protein folds \u2014 is determined by its sequence of amino acids. Many methods for designing novel proteins, including those intended to bind to disease-causing molecules in human cells, follow a two-step process: first, a target structure is defined, and then a machine-learning framework generates a repertoire of sequences predicted to adopt that structure. In nature, many different amino acid sequences can fold into the same architecture, and conversely, a single sequence can potentially adopt multiple conformations depending on flexibility or functional triggers. The challenge for artificial intelligence, therefore, is to recognize that there are many potentially useful answers \u2014 that a vast number of sequences can achieve the same fold.<\/p>\n<p>\u201cFor years, the field has measured success by asking whether a model can reproduce the protein sequence that evolution happened to select \u2014 our work shows that this isn\u2019t the best metric for protein design,\u201d says <a href=\"https:\/\/biology.mit.edu\/profile\/amy-e-keating\/\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Amy E. Keating<\/a>, Department of Biology head, Jay A. Stein (1968) Professor of Biology, professor of biological engineering, and senior author of a paper recently published in <em><a href=\"https:\/\/www.pnas.org\/\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">PNAS<\/a><\/em>. The new framework, PottsMPNN, addresses this gap by modeling the sequence-energy landscape \u2014 the relationship between the identity of each amino acid and the stability <a href=\"https:\/\/overcentral.com\/en\/servant-of-the-lake-achievement-guide\/\" title=\"Servant Of The Lake Unlocks Every Achievement\" data-iacss-internal=\"1\">of the<\/a> protein \u2014 more accurately than existing methods.<\/p>\n<h2>What Is PottsMPNN and How Does It Differ from Existing Models?<\/h2>\n<p>PottsMPNN is a machine-learning framework that incorporates pairwise interactions between amino acids and introduces controlled noise during training to prevent the model from merely mimicking natural sequences. The name \u201cPotts\u201d refers to a statistical model (the Potts model) that captures pairwise couplings, a concept borrowed from physics to describe interactions in systems with many components. In proteins, each position can be occupied by one of 20 possible amino acids, and the energetic favorability of a given pair at two interacting positions strongly influences overall stability. By explicitly modeling these pairwise distributions, PottsMPNN gains a deeper understanding of which sequence combinations are physically plausible, beyond what is encoded in natural sequences.<\/p>\n<p>Perhaps the most widely used protein sequence design model today was released in 2022. For a field that moves as fast as machine learning in biology, that model has not been surpassed. Graduate student and lead author Foster Birnbaum says, \u201cWe\u2019ve been trying to understand why that is, and what it is about that model that makes it so useful.\u201d The MIT team sought to build on its strengths while addressing its limitations \u2014 particularly its tendency to overfit to native sequences. PottsMPNN uses noise (adding variations to protein structures during training) to decrease that tendency, increasing the diversity of structures for which it can generate sequences. It also incorporates sets of evolutionarily related sequences during training to teach the model how different sequences can adopt the same folded structure.<\/p>\n<h2>A Featured Snippet Answer: How Does PottsMPNN Improve Protein Sequence Design?<\/h2>\n<p>PottsMPNN improves protein sequence design by modeling the physical interactions between all possible amino acid pairs at every position in a protein, rather than relying solely on patterns found in natural sequences. It uses a pairwise Potts distribution to capture these interactions, introduces structural noise during training to avoid overfitting to natural sequences, and includes evolutionarily related sequences to teach the model that many different sequences can adopt the same fold. This combination allows the framework to generate sequences that are more likely to fold into the desired structure and to predict the effects of mutations on protein stability more accurately, even for completely novel, designed proteins that have no native counterpart.<\/p>\n<h2>Beyond the Noise: The Technical Innovations Behind PottsMPNN<\/h2>\n<p>The development of PottsMPNN involved three key technical advances. First, Birnbaum focused on the strategic application of \u201cnoise.\u201d Adding variations to a protein structure during training decreases the model\u2019s tendency to overly mimic native sequences, thereby increasing the diversity of structures for which it can generate viable sequences. This is crucial for designing proteins with folds that have never been observed in nature.<\/p>\n<p>Second, the model uses a pairwise distribution to capture interactions between amino acids. The ability to account for the physical interactions between all 20 possible sequence options at a pair of positions in the protein is a key reason PottsMPNN more accurately models the sequence-energy landscape than other methods. Many existing models simplify these interactions, but the MIT team demonstrated that explicitly modeling pairwise couplings yields better predictions of stability and mutational effects.<\/p>\n<p>Third, Birnbaum and colleagues introduced sets of evolutionarily related sequences into the <a href=\"https:\/\/overcentral.com\/en\/mit-study-reveals-ai-art-lacks-traceable-training-data\/\" title=\"MIT Study Reveals AI Art Lacks Traceable Training Data\" data-iacss-internal=\"1\">training data<\/a>. This step teaches the model how different sequences can adopt the same folded structure. Birnbaum acknowledges that incorporating evolutionary information is, in some ways, still a reliance on natural sequences, but PottsMPNN succeeded in demonstrating that as the model depends less and less on native sequences, structural compatibility and energy prediction \u2014 including for novel proteins \u2014 improve.<\/p>\n<h2>Practical Implications for Protein Engineering and Drug Design<\/h2>\n<p>Adding PottsMPNN to a protein design pipeline will allow researchers to design structurally feasible proteins with sequences that do not resemble those of any native protein. \u201cIf we\u2019re thinking about a completely novel, designed structure, there would be no native sequence to compare it to,\u201d Birnbaum explains. \u201cWhat we actually care about is how likely the generated sequences are to fold into the desired structures, how well the model understands the sequence-energy landscape, and how well it can predict the effect of mutations on the stability of the protein.\u201d<\/p>\n<p>This capability is especially valuable for designing proteins that bind to specific disease-causing molecules, such as those involved in cancer or viral infections. Traditional drug discovery often relies on screening large libraries of small molecules or antibodies, but computationally designed proteins that can be tailored for high affinity and specificity offer a powerful alternative. PottsMPNN\u2019s improved understanding of the energetic consequences of mutations also means that researchers can more reliably predict how a designed protein will behave when introduced into a cell or organism.<\/p>\n<h2>Protein Design in the Age of AI: Strategic Significance<\/h2>\n<p>\u201cOnce we can design any protein we want, that enables us to do a potentially scary amount of biological engineering,\u201d Birnbaum says. \u201cIt\u2019s a difficult task, but I\u2019m really optimistic about this century\u2019s progress in biology.\u201d The potential applications range from creating new enzymes for industrial catalysis (e.g., breaking down plastic waste) to designing protein-based sensors for environmental monitoring, and even engineering therapeutic proteins that can target previously undruggable targets. The ability to generate sequences that are not constrained by evolutionary history means that designers can explore regions of sequence space that nature never visited, potentially yielding properties that are impossible for natural proteins.<\/p>\n<p>Birnbaum hopes that PottsMPNN can be further improved and fine-tuned for specific tasks, which has in the past led to better predictions \u2014 for example, on the outcome or consequence of a particular mutation. The MIT team\u2019s approach is modular: the core framework can be adapted with additional training data or specialized loss functions to prioritize certain features, such as binding affinity or thermal stability.<\/p>\n<h2>The Broader Context: Replacing Outdated Benchmarks in Computational Biology<\/h2>\n<p>The field of protein design has been dominated by a single metric for more than two decades: the ability to recover the wild-type sequence for a given structure. This metric, while intuitive, ignores the fundamental biological reality that evolution has sampled only a tiny fraction of all possible sequences. By reframing the goal as generating sequences that are thermodynamically feasible and structurally compatible, rather than sequence-similar to natural proteins, PottsMPNN aligns computational design with actual physical principles. Keating emphasizes, \u201cOur methods move the field toward designing useful new-to-nature proteins for diverse applications while providing a stronger foundation for future advances.\u201d<\/p>\n<p>This reframing has implications far beyond the MIT lab. As the cost of DNA synthesis continues to drop and high-throughput screening technologies improve, the bottleneck in protein engineering is increasingly the quality of the design. Models like PottsMPNN that generate a wider array of viable sequences provide more raw material for experimental validation, increasing the likelihood that a designed protein actually works as intended. That, in turn, accelerates the cycle of design-build-test-learn that underpins modern biotechnology.<\/p>\n<h2>Technical Details: Pairwise Potts Distributions and the Sequence-Energy Landscape<\/h2>\n<p>To understand why PottsMPNN outperforms earlier models, it helps to consider the sequence-energy landscape. Each amino acid at a given position interacts with others through hydrogen bonds, electrostatic interactions, van der Waals forces, and hydrophobic effects. The total energy of a protein is a function of these pairwise and higher-order interactions. Most machine-learning models approximate this energy using single-position preferences (e.g., from a position-specific scoring matrix) combined with simple neural networks. PottsMPNN, by contrast, explicitly computes pairwise energies using a Potts model \u2014 a statistical framework that dates back to the 1920s in physics but has found renewed relevance in computational biology for analyzing coevolutionary signals in protein families.<\/p>\n<p>The Potts model assigns a weight to every possible pair of amino acids at two positions. During training, the model learns these weights from crystallographic structures and from mutation stability data. The resulting energy function can then be used to score candidate sequences or to guide the generative process. Birnbaum and colleagues showed that this approach leads to significantly better predictions of mutant stability compared to models that ignore pairwise couplings. Furthermore, because the energy function is grounded in physics, it generalizes better to structures that are far from any training example, which is essential for truly novel designs.<\/p>\n<h2>Challenges and Future Directions<\/h2>\n<p>Despite its successes, PottsMPNN is not a complete solution to protein design. Birnbaum notes that the model could be further improved by incorporating dynamics \u2014 proteins are not static structures, and flexibility often plays a critical role in function. Additionally, while PottsMPNN captures pairwise interactions, it <a href=\"https:\/\/overcentral.com\/en\/ai-search-moves-cognitive-load-does-not-remove-it\/\" title=\"AI Search Moves Cognitive Load, Does Not Remove It\" data-iacss-internal=\"1\">does not<\/a> yet account for higher-order (three-body or more) effects that can be important in some contexts. Extending the framework to handle these would require more data and more sophisticated training algorithms.<\/p>\n<p>Another challenge is the need for accurate structural input. PottsMPNN takes a protein backbone structure as input; if the target structure is poorly predicted or incomplete, the sequences generated may not be feasible. Integrating PottsMPNN with state-of-the-art structure prediction models (such as AlphaFold2 or RoseTTAFold) could create an end-to-end pipeline that begins with a functional requirement and ends with a validated sequence ready for synthesis.<\/p>\n<p>The MIT team has made the code for PottsMPNN openly available, encouraging other researchers to test, modify, and improve upon it. This open-science approach aligns with the rapid pace of development in the field, where the best models often emerge from collaborative iterations rather than isolated efforts.<\/p>\n<h2>Looking to the Horizon: A New Paradigm for Protein Design<\/h2>\n<p>The publication in <em>PNAS<\/em> marks an important step, but the real test of PottsMPNN will come in the laboratory. Already, the team is planning experiments to validate sequences designed by the model for a range of target structures, including some that have never been observed in nature. If these early validations succeed, the framework could become a standard tool in the protein engineer\u2019s toolkit, complementing existing models by providing a physically principled alternative that prioritizes stability and diversity over sequence recovery.<\/p>\n<p>In a field where the ultimate goal is to design proteins with properties that evolution never produced, the ability to escape the gravitational pull of natural sequences is liberating. PottsMPNN does not simply generate more sequences; it generates better ones \u2014 sequences that are energetically feasible, structurally compatible, and entirely new. For researchers developing protein-based drugs, biosensors, or industrial enzymes, that is exactly the kind of tool needed to transform promise into practice.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Protein design has long been constrained by a paradox: the same three-dimensional structure can arise from countless amino acid sequences, yet most computational models for designing new proteins have been judged by how closely they can reproduce the single sequence that nature happened to select. A new machine-learning framework called PottsMPNN, developed in the Department [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":78156,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/pub-4d4fc17555de4152be07eaf2a416a31e.r2.dev\/en\/ocie_1787879805526.jpg","fifu_image_alt":"PottsMPNN Expands Protein Design Beyond Natural Sequences","footnotes":""},"categories":[31],"tags":[],"class_list":["post-78142","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology"],"fifu_image_url":"https:\/\/pub-4d4fc17555de4152be07eaf2a416a31e.r2.dev\/en\/ocie_1787879805526.jpg","fifu_image_alt":"PottsMPNN Expands Protein Design Beyond Natural Sequences","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/78142","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=78142"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/78142\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/78156"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=78142"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=78142"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=78142"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}