Protein Plasticity and Evolution

Protein Plasticity and Evolution

Proteins are the most intriguing biomolecules of life, driving essential functions. The Toth-Petroczy lab is fascinated by how the protein diversity observed in nature is generated and distributed across protein families. We aim to understand protein diversity at multiple levels: from individual protein sequences, to their collective organization into biomolecular condensates, and to erroneous protein production through transcriptional and translational errors that generate phenotypic mutations.

The lab is currently moving from MPI-CBG to CASUS:
=> MPI-CBG website

Our research therefore focuses on three areas that extend beyond the classical emphasis on structured proteins, membrane-bound organelles, and genetic mutations. Specifically, we investigate the sequence–function relationships of

  1. intrinsically disordered regions (IDRs)
  2. biomolecular condensates and other higher-order protein assemblies, and
  3. phenotypic mutations and their roles in evolution and disease.

We address these questions through an interdisciplinary combination of computation and experiment. We develop machine learning and mechanistic algorithms that exploit evolutionary and structural information, curate and generate ground-truth datasets for systematic analyses, perform computational experiments to uncover principles of protein evolution, and validate predictions through targeted experiments and high-throughput screening.

In the long term, we aim to integrate AI models with experimental proteomics to build a unified framework for understanding protein diversity, dynamics and function across species and cell states.

Prof. Dr. Agnes Toth-Petroczy

CASUS Research Team Leader &
Professor of Systems Biology of Protein Evolution, Biotechnology Center (BIOTEC) at the Center for Molecular and Cellular Bioengineering (CMCB), Dresden University of Technology

Contact

Center for Advanced Systems Understanding

Conrad-Schiedt-Straße 20

D-02826 Görlitz

Identifying functional motifs in disordered regions ̬

Increasing insights into how sequence motifs in intrinsically disordered regions (IDRs) provide functions underscore the need for systematic motif detection. Contrary to structured regions where motifs can be readily identified from sequence alignments, the rapid evolution of IDRs limits the usage of alignment-based tools in reliably detecting motifs within.

Here, we developed SHARK-capture (10.1002/pro.70091), an alignment-free motif detection tool designed for difficult-to-align regions. SHARK-capture innovates on word-based methods by flexibly incorporating amino acid physicochemistry to assess motif similarity without requiring rigid definitions of equivalency groups and offers consistently strong performance in a systematic benchmark, with superior residue-level performance.

Chi Fung Willis Chow, Swantje Lenz, Maxim Scheremetjew, Soumyadeep Ghosh, Doris Richter, Ceciel Jegers, Alexander von Appen, Simon Alberti, Agnes Toth-Petroczy. SHARK-capture identifies functional motifs in intrinsically disordered protein regions. Protein Sci. 2025 Apr;34(4):e70091. doi: 10.1002/pro.70091. PMID: 40100159; PMCID: PMC11917139.

Evolution of Instrinsically Disordered Regions (IDRs) ̬

Most of our understanding about the molecular history of the cell is biased towards well-conserved ordered protein domains. Intrinsically disordered protein regions (IDRs) are largely unexplored due to their low sequence complexity and low conservation. Yet they have pivotal roles in the cell including the formation of biomolecular condensates.

We have designed an alignment-free algorithm to assess homology between un-ulignable IDR sequences (SHARK-dive, 10.1073/pnas.2401622121) and developed a method to identify conserved motifs within a set of homologous IDR sequences (SHARK-capture). Based on evolutionary information, we can classify IDR families using our new tools, and propose functional sites that we are validating experimentally.

Chi Fung Willis Chow, Soumyadeep Ghosh, Anna Hadarovich, and Agnes Toth-Petroczy. SHARK enables sensitive detection of evolutionary homologs and functional analogs in unalignable and disordered sequences. Proc Natl Acad Sci U S A. 2024 Oct 15;121(42):e2401622121. doi: 10.1073/pnas.2401622121. Epub 2024 Oct 9. PMID: 39383002.

Biomolecular Condensates ̬

To integrate interdisciplinary scientific knowledge about the function and composition of biomolecular condensates, we developed the CrowDsourcing COndensate Database and Encyclopedia (CD-CODE.org). CD-CODE is a community-editable platform, which includes a database of biomolecular condensates based on the literature, an encyclopedia of relevant scientific terms, and a crowdsourcing web application.

Ksenia Kuznetsova, Maxim Scheremetjew, Jialin Yin, HongKee Moon, Diego A Vargas, Anna Hadarovich, Natasha Lewis, Carsten Hoege, Chi Fung Willis Chow, David Kuster, Jik Nijssen, Alberto Hernandez-Armendariz, Jonathan C Savage, Yu Wei, Silja Zedlitz, Hari Raj Singh, Soumyadeep Ghosh, Allysa P Kemraj, Lena Hersemann, Anthony A Hyman, Diana M Mitrea, Agnes Toth-Petroczy. CD-CODE 2.0: an enhanced condensate knowledgebase integrating pathobiology, condensate modulating drugs, and host–pathogen interactions. Nucleic Acids Research, 2025, gkaf1104. Open access

Behind the paper blog post: Condensing information on condensates

MPI-CBG NewsCD-CODE – Database and Encyclopedia for membraneless liquid droplets

Proteins involved in condensates in cells (PICNIC) ̬

To enable systematic detection of condensate proteins, we developed an algorithm to recognize proteins involved in in vivo biomolecular condensates regardless of the mechanism of condensate formation. Using a curated dataset of CD-CODE, we trained a machine learning classifier based on sequence- and structure-based features to predict if a protein is part of a condensate. Our model, PICNIC (Proteins Involved in CoNdensates In Cells) enables systematic detection of condensate proteins across organisms.

We show that protein sequence and structure encode the condensate forming function. PICNIC accurately predicts which proteins are involved in condensates, as 18 out of 24 proteins tested experimentally indeed formed high-confidence in cellulo condensates. We infer that ~40% of the proteomes of organisms across the tree of life form condensates with no apparent correlation with disorder content and organismal complexity.

Precomputed scores for 14 organisms are available at picnic.cd-code.org.

Hadarovich A, Singh HR, Ghosh S, Scheremetjew M, Rostam N, Hyman AA, Toth-Petroczy A. PICNIC accurately predicts condensate-forming proteins regardless of their structural disorder across organisms. Nat Commun. 2024 Dec 11;15(1):10668. doi: 10.1038/s41467-024-55089-x. PMID: 39663388.

Clinical Variant Assessment ̬

Assessing the phenotypic effects of single nucleotide variants remains a challenge even in genes with well-established clinical impact. We have developed DeMAG (Deciphering Mutations in Actionable Genes, demag.org) that reaches unprecedented accuracy in the classification of missense variants in clinically actionable genes. Our approach makes use of protein 3D structures and epistatic residue interactions to derive new predictive features, as well as abundant clinical diagnostic data in these genes to improve performance.

Luppino, Federica, Ivan A. Adzhubei, Christopher A. Cassa, and Agnes Toth-Petroczy. 2023. “DeMAG Predicts the Effects of Variants in Clinically Actionable Genes by Integrating Structural and Evolutionary Epistatic Features.” Nature Communications 2023, 14 (1): 1–14. Open Access

MPI-CBG  News articleNew tool facilitates clinical interpretation of genetic information.

Phenotypic Mutations ̬

Transcriptional and translational errors generate diverse sequences, most dysfunctional, some harboring novel functions (Romero et. al, Prot Sci. 2022).

To detect and quantify phenotypic mutations proteome-wide, we combine theoretical modeling, machine learning and experiments. We developed a theoretical model and mass spectrometry pipeline to study amino acid misincoroprations at proteome scale (Landerer et. al, Mol Bio Evol 2024).

Using a library of reporters for STOP codon miscoding in E. coli, we found that under stress conditions, stop codon miscoding may occur with a rate as high as 80%, depending on the nucleotide context (Romero et. al, Nat. Commu. 2024). Further, our analysis of selected reporters by mass spectrometry and RNA-seq showed that not only translation but also transcription errors contribute to stop codon miscoding. The RNA polymerase is more likely to misincorporate a nucleotide at premature stop codons. 

We are also developing algorithms for predicting frameshifts by using proteomics data. We aim at discovering novel frameshift and STOP codon read-through variants, and we will explore the evolutionary potential of slippage sites and test whether they could lead to promiscuous protein functions.

Selected publications

Ksenia Kuznetsova, Maxim Scheremetjew, Jialin Yin, HongKee Moon, Diego A Vargas, Anna Hadarovich, Natasha Steffi Lewis, Carsten Hoege, Chi Fung Willis Chow, David Kuster, Jik Nijssen, Alberto Hernandez-Armendariz, Jonathan C Savage, Yu Wei, Silja Zedlitz, Hari Raj Singh, Soumyadeep Ghosh, Allysa P Kemraj, Lena Hersemann, Anthony Hyman, Diana M Mitrea, Agnes Toth-Petroczy (2026) . NAT. Nucleic Acids Res.

CD-CODE 2.0 (https://cd-code.org) is an enhanced web application and database of condensates, expanding the utility of version 1.0 for biomedical research. New features include data on nucleic acids condensate components, infectious condensates, condensate modulating drugs, and disease-linked condensates. Enhanced search functions, programmatic access, and relational architecture enable interconnectivity across major biomedical databases (e.g. condensates, proteins, chemistry, and disease), fostering systems-level insights and accelerating hypothesis generation and therapeutic discovery through the lens of condensate biology.

Anna Hadarovich, Maxim Scheremetjew, Hari Raj Singh, HongKee Moon, Lena Hersemann, Agnes Toth-Petroczy (2026). Bioinformatics.

Biomolecular condensates have been implicated in key cellular processes such as gene regulation, stress response, and signaling, and dysregulation of condensates has been linked to neurodegeneration and other diseases. Computational algorithms that predict protein condensation can aid systematic characterization of biomolecular condensates at the proteome scale. However, many experimental labs may lack the computational background or resources to run sophisticated prediction tools locally.

Team members

_

Dr. Anna Hadarovich

Postdoctoral Researcher

Alumni

Jonathan Berthold

Dr. Jonathan Berthold

Trina De

Dr. Trina De

_

Chanwoong Hwang

Rui Li

Dr. Rui Li

Ashkan Mokarian

Ashkan Mokarian

_

Henrique Romão