publications
Reverse chronological order. (*) equal contribution.
See also Google Scholar and ORCID.
2026
- MLSTSteering sequence generation in protein language models through Iterative Lookback Monte Carlo SamplingFrancesco Calvanese, Gianluca Lombardi, Martin Weigt*, and Jorge FERNANDEZ-DE-COSSIO-DIAZ*Machine Learning: Science and Technology, 2026Spotlight at the GenBio Workshop, ICML 2026
Protein language models (pLMs) leverage large-scale evolutionary data to generate novel sequences, but steering generation toward desired physicochemical properties without sacrificing diversity remains a major challenge. Existing approaches often induce severe diversity loss or require computationally expensive retraining. We introduce Iterative Lookback Monte Carlo (ILMC), a training-free inference-time sampling strategy that interleaves autoregressive elongation with Metropolis–Hastings refinement to approximate sampling from a maximum-entropy target distribution balancing generative quality and steering objectives. We show theoretically that this target distribution is entropy-maximizing under fixed generative quality and steering constraints, and empirically that ILMC produces more diverse samples than standard autoregressive baselines at matched generative quality. Using simple steering potentials, ILMC improves desired molecular properties, yielding higher predicted melting temperatures than both unsteered autoregressive sampling and compute-matched rejection sampling at matched generative quality. ILMC naturally applies to classifier-guided steering, where it outperforms purely autoregressive guidance in diversity while maintaining comparable enrichment of target properties. We validate ILMC on family-specific pLMs and on the multi-family model ProGen3.
@article{calvanese2026steering, title = {Steering sequence generation in protein language models through Iterative Lookback Monte Carlo Sampling}, author = {Calvanese, Francesco and Lombardi, Gianluca and Weigt, Martin and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge}, journal = {Machine Learning: Science and Technology}, volume = {7}, number = {5}, pages = {055003}, year = {2026}, doi = {10.1088/2632-2153/ae923a}, note = {Spotlight at the GenBio Workshop, ICML 2026}, } - PNASCross-individual translation of spontaneous zebrafish brain activity through a shared latent representationMattéo Dommanget-Kott, Jorge FERNANDEZ-DE-COSSIO-DIAZ, Guillaume Faye-Bédrin, Georges Debrégeas, and Volker BormuthProceedings of the National Academy of Sciences, 2026
Spontaneous activity is a hallmark of brain function, reflecting the underlying circuit organization. Identifying conserved structure across individuals in this self-sustained activity has remained a longstanding challenge, especially in vertebrates where one-to-one neuron correspondence is inaccessible. Here, we introduce latent-aligned Restricted Boltzmann Machines (LaRBMs), an unsupervised generative approach that uncovers a common representational space from cell-resolved whole-brain recordings in larval zebrafish. This latent space consists of spatially localized coactivation motifs, or cell assemblies, that generalize across animals and form interpretable building blocks of population-wide activity. LaRBMs enable bidirectional mapping of instantaneous whole-brain activity patterns between individuals: Activity patterns from one fish can be encoded into the latent space and decoded into another. The translated patterns are assigned high probability by the recipient model and retain the original spatial organization. These results show that spontaneous activity in the vertebrate brain is highly stereotyped at the level of functional cell assemblies and can be reliably captured through a common latent code. Because it provides an interpretable and quantitative framework for functional cross-individual alignment, LaRBM paves the way for comparative phenotyping of brain activity across developmental, genetic, and pathological variation.
@article{dommangetkott2026cross, title = {Cross-individual translation of spontaneous zebrafish brain activity through a shared latent representation}, author = {Dommanget-Kott, Mattéo and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Faye-Bédrin, Guillaume and Debrégeas, Georges and Bormuth, Volker}, journal = {Proceedings of the National Academy of Sciences}, volume = {123}, number = {20}, pages = {e2529064123}, year = {2026}, doi = {10.1073/pnas.2529064123}, } - PLOS CBLinking brain and behavior states in Zebrafish Larvae locomotion using hidden Markov modelsMattéo Dommanget-Kott*, Jorge FERNANDEZ-DE-COSSIO-DIAZ*, Monica Coraggioso, Volker Bormuth, Rémi Monasson, Georges Debrégeas, and 1 more authorPLOS Computational Biology, 2026
Understanding how collective neuronal activity in the brain orchestrates behavior is a central question in integrative neuroscience. Addressing this question requires models that can offer a unified interpretation of multimodal data. In this study, we jointly examine video-recordings of zebrafish larvae freely exploring their environment and calcium imaging of the Anterior Rhombencephalic Turning Region (ARTR) circuit, which is known to control swimming orientation, recorded in vivo under tethered conditions. We show that both behavioral and neural data can be accurately modeled using a Hidden Markov Model (HMM) with three hidden states. In the context of behavior, the hidden states correspond to leftward, rightward, and forward swimming. The HMM robustly captures the key statistical features of the swimming motion, including bout-type persistence and its dependence on bath temperature, while also revealing inter-individual phenotypic variability. For neural data, the three states are found to correspond to left- and right-lateral activation of the ARTR circuit, known to govern the selection of left vs. right reorientation, and a balanced state, which likely corresponds to the behavioral forward state. To further unify the two analyses, we exploit the generative nature of the HMM, using neural sequences to generate synthetic swimming trajectories, whose statistical properties are similar to the behavioral data. Overall, this work demonstrates how state-space models can be used to link neuronal and behavioral data, providing insights into the mechanisms of self-generated action.
@article{dommangetkott2026linking, title = {Linking brain and behavior states in Zebrafish Larvae locomotion using hidden Markov models}, author = {Dommanget-Kott, Mattéo and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Coraggioso, Monica and Bormuth, Volker and Monasson, Rémi and Debrégeas, Georges and Cocco, Simona}, journal = {PLOS Computational Biology}, volume = {22}, number = {1}, pages = {e1013762}, year = {2026}, doi = {10.1371/journal.pcbi.1013762}, } - eLifeDesign and experimental characterization of specificity-switching mutational paths of WW domainsAhmed Rehan, Eugenio Mauri, Jorge FERNANDEZ-DE-COSSIO-DIAZ, Pierre-Guillaume Brun, Remi Monasson, Marco Ribezzi-Crivellari, and 1 more authoreLife (Reviewed Preprint), 2026
Specific interactions between proteins and other biomolecules are ubiquitous in cellular processes. How specificity is encoded in the protein sequence and can be modified through a minimal set of concerted mutations is a complex issue. In this work, we focus on the WW protein domain, whose variants specifically bind to different classes of proline-rich peptides. Combining unsupervised learning of homologous WW sequence data with Restricted Boltzmann Machines (RBM) and path-sampling methods, we design mutational paths of putative WW domains interpolating between two natural WW domains with either distinct or similar specificities. Sequences along the designed paths are then experimentally validated with high-throughput in-vitro binding assays against 3 peptides of different classes. The vast majority (93%) of intermediate sequences along the designed paths are responsive to the initial or/and final peptides. On the contrary, domains along scrambled paths, in which the same mutations are introduced in random order are not functional, emphasizing how successful design crucially depends on the ability to model epistatic interactions. Interestingly, switch in specificity between classes I and IV whose representative peptides bind to different pockets on the WW domain appears to be smooth, with intermediates displaying some level of binding cross-reactivity with all tested peptides. We finally show that the RBM paths share a high identity with internal nodes obtained from ancestral sequence reconstruction based on the seed WW domains.
@article{rehan2026design, journal = {eLife (Reviewed Preprint)}, title = {Design and experimental characterization of specificity-switching mutational paths of WW domains}, author = {Rehan, Ahmed and Mauri, Eugenio and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Brun, Pierre-Guillaume and Monasson, Remi and Ribezzi-Crivellari, Marco and Cocco, Simona}, year = {2026}, doi = {10.7554/elife.110491}, } - arXivReplica Theory of Spherical Boltzmann Machine EnsemblesThomas Tulinski, Jorge FERNANDEZ-DE-COSSIO-DIAZ, Simona Cocco, and Rémi MonassonarXiv preprint, 2026
Training in machine learning generally consists in finding one model, whose parameters minimize a data-dependent loss. Yet, empirical work shows that ensemble learning, an approach in which multiple models are sampled, can improve performance. Here, we provide an analytical framework to understand these observations in the case of Boltzmann machines, exploiting a duality between ensemble learning and large deviations of the free energy in spin-glass models. Replica calculations allow us to fully solve the case of spherical Boltzmann machine ensembles, and clarify when ensemble learning improves over standard loss minimization, in particular for nearly finite-dimensional data. Our framework can also be applied to complex data distributions, in agreement with numerical simulations on deep networks.
@article{thomastulinski2026replica, title = {Replica Theory of Spherical Boltzmann Machine Ensembles}, author = {Tulinski, Thomas and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Cocco, Simona and Monasson, Rémi}, journal = {arXiv preprint}, year = {2026}, } - NeurIPSSpherical Boltzmann machines: a solvable theory of learning and generation in energy-based modelsThomas Tulinski, Simona Cocco, Rémi Monasson, and Jorge FERNANDEZ-DE-COSSIO-DIAZIn Advances in Neural Information Processing Systems (NeurIPS), 2026
Energy-based models (EBMs) are flexible generative architectures inspired by statistical physics, but their learning and generative properties remain poorly understood. Here, we analyze a solvable EBM in the high-dimensional limit: the spherical Boltzmann machine (SBM). Combining tools from random matrix theory and dynamical mean-field theory, we: solve exact equations describing the training dynamics of the SBM; compute the Bayesian evidence, which acts as a partition function in parameter space and encodes global properties of the trained model; and uncover cascades of phase transitions that occur both during training and as a function of hyperparameters, related to successive alignment and condensation of the top modes of the coupling matrix to the data. We connect these transitions to sampling-time generative phenomena in a teacher-student scenario, including: sampling temperature tuning, double descent as a function of regularization strength, tempered posterior effects, and out-of-equilibrium effects during training that induce biases in the trained model. We provide numerical evidence demonstrating that all these phenomena appear in standard generative architectures, beyond the SBM.
@inproceedings{thomastulinski2026spherical, title = {Spherical Boltzmann machines: a solvable theory of learning and generation in energy-based models}, author = {Tulinski, Thomas and Cocco, Simona and Monasson, Rémi and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS)}, year = {2026}, }
2025
- Nat. Commun.Designing molecular RNA switches with Restricted Boltzmann machinesJorge FERNANDEZ-DE-COSSIO-DIAZ*, Pierre Hardouin*, Francois-Xavier Moutier, Andrea Di Gioacchino, Bertrand Marchand, Yann Ponty, and 3 more authorsNature Communications, 2025
Riboswitches are structured allosteric RNA molecules that change conformation upon metabolite binding, triggering a regulatory response. Here we focus on the de novo design of riboswitch-like aptamers, the core part of the riboswitch undergoing structural changes. We use Restricted Boltzmann machines (RBM) to learn generative models from homologous sequence data. We first verify, on four different riboswitch families, that RBM-generated sequences correctly capture the conservation, covariation and diversity of natural aptamers. The RBM model is then used to design new SAM-I riboswitch aptamers. To experimentally validate the properties of the structural switch in designed molecules, we resort to chemical probing (SHAPE and DMS), and develop a tailored analysis pipeline adequate for high-throughput tests of diverse sequences. We probe a total of 476 RBM-designed and 201 natural sequences. Designed molecules with high RBM scores, with 20% to 40% divergence from any natural sequence, display ≈ 30% success rate of responding to SAM with a structural switch similar to their natural counterparts. We show how the capability of the designed molecules to switch conformation is connected to fine energetic features of their structural components.
@article{fernandezdecossiodiaz2025designing, title = {Designing molecular RNA switches with Restricted Boltzmann machines}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Hardouin, Pierre and Lyonnet du Moutier, Francois-Xavier and Di Gioacchino, Andrea and Marchand, Bertrand and Ponty, Yann and Sargueil, Bruno and Monasson, Rémi and Cocco, Simona}, journal = {Nature Communications}, volume = {16}, number = {1}, pages = {11223}, year = {2025}, doi = {10.1038/s41467-025-66265-y}, } - arXivA High-Order Cumulant Extension of Quasi-Linkage EquilibriumKai S. Shimagaki, Jorge FERNANDEZ-DE-COSSIO-DIAZ, Mauro Pastore, Rémi Monasson, Simona Cocco, and John P. BartonarXiv preprint, 2025
A central question in evolutionary biology is how to quantitatively understand the dynamics of genetically diverse populations. Modeling the genotype distribution is challenging, as it ultimately requires tracking all correlations (or cumulants) among alleles at different loci. The quasi-linkage equilibrium (QLE) approximation simplifies this by assuming that correlations between alleles at different loci are weak – i.e., low linkage disequilibrium – allowing their dynamics to be modeled perturbatively. However, QLE breaks down under strong selection, significant epistatic interactions, or weak recombination. We extend the multilocus QLE framework to allow cumulants up to order K to evolve dynamically, while higher-order cumulants (>K) are assumed to equilibrate rapidly. This extended QLE (exQLE) framework yields a general equation of motion for cumulants up to order K, which parallels the standard QLE dynamics (recovered when K = 1). In this formulation, cumulant dynamics are driven by the gradient of average fitness, mediated by a geometrically interpretable matrix that stems from competition among genotypes. Our analysis shows that the exQLE with K=2 accurately captures cumulant dynamics even when the fitness function includes higher-order (e.g., third- or fourth-order) epistatic interactions, capabilities that standard QLE lacks. We also applied the exQLE framework to infer fitness parameters from temporal sequence data. Overall, exQLE provides a systematic and interpretable approximation scheme, leveraging analytical cumulant dynamics and reducing complexity by progressively truncating higher-order cumulants.
@article{kaisshimagaki2025high, title = {A High-Order Cumulant Extension of Quasi-Linkage Equilibrium}, author = {Shimagaki, Kai S. and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Pastore, Mauro and Monasson, Rémi and Cocco, Simona and Barton, John P.}, journal = {arXiv preprint}, year = {2025}, }
2024
- BMC Bioinf.Unsupervised modeling of mutational landscapes of adeno-associated viruses viabilityMatteo De Leonardis, Jorge FERNANDEZ-DE-COSSIO-DIAZ, Guido Uguzzoni, and Andrea PagnaniBMC Bioinformatics, 2024
Adeno-associated viruses 2 (AAV2) are minute viruses renowned for their capacity to infect human cells and akin organisms. They have recently emerged as prominent candidates in the field of gene therapy, primarily attributed to their inherent non-pathogenic nature in humans and the safety associated with their manipulation. The efficacy of AAV2 as gene therapy vectors hinges on their ability to infiltrate host cells, a phenomenon reliant on their competence to construct a capsid capable of breaching the nucleus of the target cell. To enhance their infection potential, researchers have extensively scrutinized various combinatorial libraries by introducing mutations into the capsid, aiming to boost their effectiveness. The emergence of high-throughput experimental techniques, like deep mutational scanning (DMS), has made it feasible to experimentally assess the fitness of these libraries for their intended purpose. Notably, machine learning is starting to demonstrate its potential in addressing predictions within the mutational landscape from sequence data. In this context, we introduce a biophysically-inspired model designed to predict the viability of genetic variants in DMS experiments. This model is tailored to a specific segment of the CAP region within AAV2’s capsid protein. To evaluate its effectiveness, we conduct model training with diverse datasets, each tailored to explore different aspects of the mutational landscape influenced by the selection process. Our assessment of the biophysical model centers on two primary objectives: (i) providing quantitative forecasts for the log-selectivity of variants and (ii) deploying it as a binary classifier to categorize sequences into viable and non-viable classes.
@article{deleonardis2024unsupervised, title = {Unsupervised modeling of mutational landscapes of adeno-associated viruses viability}, author = {De Leonardis, Matteo and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Uguzzoni, Guido and Pagnani, Andrea}, journal = {BMC Bioinformatics}, volume = {25}, number = {1}, pages = {229}, year = {2024}, doi = {10.1186/s12859-024-05823-5}, } - ICLRAccelerated Sampling with Stacked Restricted Boltzmann MachinesJorge FERNANDEZ-DE-COSSIO-DIAZ*, Clément Roussel*, Simona Cocco, and Rémi MonassonIn The Twelfth International Conference on Learning Representations (ICLR), 2024
Sampling complex distributions is an important but difficult objective in various fields, including physics, chemistry, and statistics. An improvement of standard Monte Carlo (MC) methods, intensively used in particular in the context of disordered systems, is Parallel Tempering, also called replica exchange MC, in which a sequence of MC Markov chains at decreasing temperatures are run in parallel and can swap their configurations. In this work we apply the ideas of parallel tempering in the context of restricted Boltzmann machines (RBM), a paradigm of unsupervised architectures, capable to learn complex, multimodal distributions. Inspired by Deep Tempering, an approach introduced for deep belief networks, we show how to learn on top of the first RBM a stack of nested RBMs, using the representations of a RBM as ‘data’ for the next one along the stack. In our Stacked Tempering approach the hidden configurations of a machine can be exchanged with the visible configurations of the next one in the stack. Replica exchanges between the different RBMs is facilitated by the increasingly clustered representations learnt by deeper RBMs, allowing for fast transitions between the different modes of the data distribution. Analytical calculations of mixing times in a simplified theoretical setting shed light on why Stacked Tempering works, and how hyperparameters, such as the aspect ratios of the RBMs and weight regularization should be chosen. We illustrate the efficiency of the Stacked Tempering method with respect to standard and replica exchange MC on several datasets: MNIST, in-silico Lattice Proteins, and the 2D-Ising model.
@inproceedings{fernandezdecossiodiaz2024accelerated, title = {Accelerated Sampling with Stacked Restricted Boltzmann Machines}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Roussel, Clément and Cocco, Simona and Monasson, Rémi}, booktitle = {The Twelfth International Conference on Learning Representations (ICLR)}, year = {2024}, } - Book chapterGenerative Modeling of RNA Sequence Families with Restricted Boltzmann MachinesJorge FERNANDEZ-DE-COSSIO-DIAZIn RNA Design, 2024
In this chapter, we discuss the potential application of Restricted Boltzmann machines (RBM) to model sequence families of structured RNA molecules. RBMs are a simple two-layer machine learning model able to capture intricate sequence dependencies induced by secondary and tertiary structure, as well as mechanisms of structural flexibility, resulting in a model that can be successfully used for the design of allosteric RNA such as riboswitches. They have recently been experimentally validated as generative models for the SAM-I riboswitch aptamer domain sequence family. We introduce RBM mathematically and practically, providing self-contained code examples to download the necessary training sequence data, train the RBM, and sample novel sequences. We present in detail the implementation of algorithms necessary to use RBMs, focusing on applications in biological sequence modeling.
@incollection{fernandezdecossiodiaz2024generative, title = {Generative Modeling of RNA Sequence Families with Restricted Boltzmann Machines}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge}, booktitle = {RNA Design}, publisher = {Springer US}, pages = {163--175}, year = {2024}, doi = {10.1007/978-1-0716-4079-1_11}, } - PLOS CBInference and design of antibody specificity: From experiments to models and backJorge FERNANDEZ-DE-COSSIO-DIAZ*, Guido Uguzzoni*, Kévin Ricard, Francesca Anselmi, Clément Nizak, Andrea Pagnani, and 1 more authorPLOS Computational Biology, 2024
Exquisite binding specificity is essential for many protein functions but is difficult to engineer. Many biotechnological or biomedical applications require the discrimination of very similar ligands, which poses the challenge of designing protein sequences with highly specific binding profiles. Experimental methods for generating specific binders rely on in vitro selection, which is limited in terms of library size and control over specificity profiles. Additional control was recently demonstrated through high-throughput sequencing and downstream computational analysis. Here we follow such an approach to demonstrate the design of specific antibodies beyond those probed experimentally. We do so in a context where very similar epitopes need to be discriminated, and where these epitopes cannot be experimentally dissociated from other epitopes present in the selection. Our approach involves the identification of different binding modes, each associated with a particular ligand against which the antibodies are either selected or not. Using data from phage display experiments, we show that the model successfully disentangles these modes, even when they are associated with chemically very similar ligands. Additionally, we demonstrate and validate experimentally the computational design of antibodies with customized specificity profiles, either with specific high affinity for a particular target ligand, or with cross-specificity for multiple target ligands. Overall, our results showcase the potential of leveraging a biophysical model learned from selections against multiple ligands to design proteins with tailored specificity, with applications to protein engineering extending beyond the design of antibodies.
@article{fernandezdecossiodiaz2024inference, title = {Inference and design of antibody specificity: From experiments to models and back}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Uguzzoni, Guido and Ricard, Kévin and Anselmi, Francesca and Nizak, Clément and Pagnani, Andrea and Rivoire, Olivier}, journal = {PLOS Computational Biology}, volume = {20}, number = {10}, pages = {e1012522}, year = {2024}, doi = {10.1371/journal.pcbi.1012522}, } - PreprintReplica symmetry breaking and clustering phase transitions in undersampled restricted Boltzmann machinesJorge FERNANDEZ-DE-COSSIO-DIAZ, Thomas Tulinski, Simona Cocco, and Rémi MonassonHAL preprint, 2024
Restricted Boltzmann machines (RBMs) are among the simplest unsupervised models implementing data/representation duality. The learning curves of RBMs trained on structured data are nevertheless difficult to characterize analytically, in part due to the presence of a partition function that depends on the trainable parameters. In this work, we present the exact solution of RBMs trained on structured data in the undersampled regime. The solution involves gradual symmetry breaking among the hidden units for decreasing regularization strength, as they specialize to finer-level details of the data. Hidden units form extensive blocks with identical weight parameters. Trained RBMs with different block sizes are separated by large barriers in the posterior distribution of the weights, which makes the optimal block size inaccessible during training with local gradient descent.
@article{fernandezdecossiodiaz2024replica, title = {Replica symmetry breaking and clustering phase transitions in undersampled restricted Boltzmann machines}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Tulinski, Thomas and Cocco, Simona and Monasson, Rémi}, journal = {HAL preprint}, year = {2024}, } - PLOS CBInference of annealed protein fitness landscapes with AnnealDCALuca Sesta, Andrea Pagnani, Jorge FERNANDEZ-DE-COSSIO-DIAZ*, and Guido Uguzzoni*PLOS Computational Biology, 2024
The design of proteins with specific tasks is a major challenge in molecular biology with important diagnostic and therapeutic applications. High-throughput screening methods have been developed to systematically evaluate protein activity, but only a small fraction of possible protein variants can be tested using these techniques. Computational models that explore the sequence space in-silico to identify the fittest molecules for a given function are needed to overcome this limitation. In this article, we propose AnnealDCA, a machine-learning framework to learn the protein fitness landscape from sequencing data derived from a broad range of experiments that use selection and sequencing to quantify protein activity. We demonstrate the effectiveness of our method by applying it to antibody Rep-Seq data of immunized mice and screening experiments, assessing the quality of the fitness landscape reconstructions. Our method can be applied to several experimental cases where a population of protein variants undergoes various rounds of selection and sequencing, without relying on the computation of variants enrichment ratios, and thus can be used even in cases of disjoint sequence samples.
@article{sesta2024inference, title = {Inference of annealed protein fitness landscapes with AnnealDCA}, author = {Sesta, Luca and Pagnani, Andrea and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Uguzzoni, Guido}, journal = {PLOS Computational Biology}, volume = {20}, number = {2}, pages = {e1011812}, year = {2024}, doi = {10.1371/journal.pcbi.1011812}, }
2023
- eLifeA transfer-learning approach to predict antigen immunogenicity and T-cell receptor specificityBarbara Bravi, Andrea Di Gioacchino*, Jorge FERNANDEZ-DE-COSSIO-DIAZ*, Aleksandra M Walczak, Thierry Mora, Simona Cocco, and 1 more authoreLife, 2023
Antigen immunogenicity and the specificity of binding of T-cell receptors to antigens are key properties underlying effective immune responses. Here we propose diffRBM, an approach based on transfer learning and Restricted Boltzmann Machines, to build sequence-based predictive models of these properties. DiffRBM is designed to learn the distinctive patterns in amino-acid composition that, on the one hand, underlie the antigen’s probability of triggering a response, and on the other hand the T-cell receptor’s ability to bind to a given antigen. We show that the patterns learnt by diffRBM allow us to predict putative contact sites of the antigen-receptor complex. We also discriminate immunogenic and non-immunogenic antigens, antigen-specific and generic receptors, reaching performances that compare favorably to existing sequence-based predictors of antigen immunogenicity and T-cell receptor specificity.
@article{bravi2023transfer, title = {A transfer-learning approach to predict antigen immunogenicity and T-cell receptor specificity}, author = {Bravi, Barbara and Di Gioacchino, Andrea and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Walczak, Aleksandra M and Mora, Thierry and Cocco, Simona and Monasson, Rémi}, journal = {eLife}, volume = {12}, pages = {e85126}, year = {2023}, doi = {10.7554/elife.85126}, } - Data BriefHEK293 producing the extracellular domain HER1: Full datasets of continuous fermentation process and metabolites analysisLisandra Calzadilla, Erick Hernández, Julio Dustet, Jorge FERNANDEZ-DE-COSSIO-DIAZ, Kalet León, Matthias Pietzke, and 3 more authorsData in Brief, 2023
The data for provide evidences of the multi steady state of the human cell line HEK 293 was obtained from 2 L bioreactor continuous culture. A HEK 293 cell line transfected to produce soluble HER1 receptor was used. The bioreactor was operated at three different dilution rates in sequential manner. Daily samples of culture broth were collected, a total of 85 samples were processed. Viable cell concentration and culture viability was addressing by trypan blue exclusion method using a hemocytometer. Heterologous HER1 supernatant concentration was quantified by a specific ELISA and the metabolites by mass spectrometry coupled to a liquid chromatography. The primary data were collected in excel files, where it was calculated the kinetic and other variables by using mass balance and mathematical principles. It was compared the steady states behavior each other’s to find out the existence of steady states’ multiplicity, taking into account the stationary phase with respect to the cell density (which means its coefficient of variation is less than 20%). From the metabolic measurements by using Liquid Chromatography coupled to mass spectrometry (LC-MS), it was also built the data matrix with the specific rates of the 76 metabolites obtained. The data were processed and analyzed, using multivariate data analysis (MVDA) to reduce the complexity and to find the main patterns present in the data. We describe also the full data of the metabolites not only for steady states but also in the time evolution, which could help others in terms of modeling and deep understanding of HEK293 metabolism, especially under different culture conditions.
@article{calzadilla2023hek, title = {HEK293 producing the extracellular domain HER1: Full datasets of continuous fermentation process and metabolites analysis}, author = {Calzadilla, Lisandra and Hernández, Erick and Dustet, Julio and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and León, Kalet and Pietzke, Matthias and Vazquez, Alexei and Mulet, Roberto and Boggiano, Tammy}, journal = {Data in Brief}, volume = {50}, pages = {109604}, year = {2023}, doi = {10.1016/j.dib.2023.109604}, } - Biochem. Eng. J.Multiple steady states and metabolic switches in continuous cultures of HEK293: Experimental evidences and metabolomicsLisandra Calzadilla, Erick Hernández, Julio Dustet, Jorge FERNANDEZ-DE-COSSIO-DIAZ, Kalet León, Matthias Pietzke, and 3 more authorsBiochemical Engineering Journal, 2023
The optimization of the production process in the biotechnology industry needs a deep understanding of the metabolic patterns developed in the bioreactors. In particular, the possibility to induce changes between different metabolic states in these cultures has opened a new path to reach this optimization. The goal is to drift the culture toward a metabolic state of maximum productivity. In this work we experimentally explore and analyze this path using a HEK293 cell line, transfected to produce the HER1 glycoprotein, cultured in continuous mode. We first show that this cell culture exhibits steady state multiplicity, i.e., different cell densities and protein concentrations for the same experimental parameters. We also demonstrate that the switch between these steady states can be triggered manipulating the dilution rate in the bioreactor. Furthermore, we present an extensive metabolic characterization of the steady states measuring metabolic concentrations through Liquid Chromatographic-Mass Spectrometry (LC-MS). The data obtained is processed using Principal Component Analysis (PCA), unveiling a correspondence between the culture multiplicity and the existence of distinctive metabolic states. Our results support the idea that different stationary states, although obtained for the same experimental parameters, are consistent with major metabolic readjustments in the culture. Finally, the comparison of the different steady states demonstrates that, in this case, the state with reduced lactate production benefits volumetric productivity.
@article{calzadilla2023multiple, title = {Multiple steady states and metabolic switches in continuous cultures of HEK293: Experimental evidences and metabolomics}, author = {Calzadilla, Lisandra and Hernández, Erick and Dustet, Julio and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and León, Kalet and Pietzke, Matthias and Vazquez, Alexei and Mulet, Roberto and Boggiano, Tammy}, journal = {Biochemical Engineering Journal}, volume = {198}, pages = {109010}, year = {2023}, doi = {10.1016/j.bej.2023.109010}, } - PRXDisentangling Representations in Restricted Boltzmann Machines without AdversariesJorge FERNANDEZ-DE-COSSIO-DIAZ, Simona Cocco, and Rémi MonassonPhysical Review X, 2023
A goal of unsupervised machine learning is to build representations of complex high-dimensional data, with simple relations to their properties. Such disentangled representations make easier to interpret the significant latent factors of variation in the data, as well as to generate new data with desirable features. Methods for disentangling representations often rely on an adversarial scheme, in which representations are tuned to avoid discriminators from being able to reconstruct information about the data properties (labels). Unfortunately adversarial training is generally difficult to implement in practice. Here we propose a simple, effective way of disentangling representations without any need to train adversarial discriminators, and apply our approach to Restricted Boltzmann Machines (RBM), one of the simplest representation-based generative models. Our approach relies on the introduction of adequate constraints on the weights during training, which allows us to concentrate information about labels on a small subset of latent variables. The effectiveness of the approach is illustrated with four examples: the CelebA dataset of facial images, the two-dimensional Ising model, the MNIST dataset of handwritten digits, and the taxonomy of protein families. In addition, we show how our framework allows for analytically computing the cost, in terms of log-likelihood of the data, associated to the disentanglement of their representations.
@article{fernandezdecossiodiaz2023disentangling, title = {Disentangling Representations in Restricted Boltzmann Machines without Adversaries}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Cocco, Simona and Monasson, Rémi}, journal = {Physical Review X}, volume = {13}, number = {2}, pages = {021003}, year = {2023}, doi = {10.1103/physrevx.13.021003}, } - ImmunoInform.Benchmarking solutions to the T-cell receptor epitope prediction problem: IMMREP22 workshop reportPieter Meysman, Justin Barton, Barbara Bravi, Liel Cohen-Lavi, Vadim Karnaukhov, Elias Lilleskov, and 13 more authorsImmunoInformatics, 2023
Many different solutions to predicting the cognate epitope target of a T-cell receptor (TCR) have been proposed. However several questions on the advantages and disadvantages of these different approaches remain unresolved, as most methods have only been evaluated within the context of their initial publications and data sets. Here, we report the findings of the first public TCR-epitope prediction benchmark performed on 23 prediction models in the context of the ImmRep 2022 TCR-epitope specificity workshop. This benchmark revealed that the use of paired-chain alpha-beta, as well as CDR1/2 or V/J information, when available, improves classification obtained with CDR3 data, independent of the underlying approach. In addition, we found that straight-forward distance-based approaches can achieve a respectable performance when compared to more complex machine-learning models. Finally, we highlight the need for a truly independent follow-up benchmark and provide recommendations for the design of such a next benchmark.
@article{meysman2023benchmarking, title = {Benchmarking solutions to the T-cell receptor epitope prediction problem: IMMREP22 workshop report}, author = {Meysman, Pieter and Barton, Justin and Bravi, Barbara and Cohen-Lavi, Liel and Karnaukhov, Vadim and Lilleskov, Elias and Montemurro, Alessandro and Nielsen, Morten and Mora, Thierry and Pereira, Paul and Postovskaya, Anna and Rodríguez Martínez, María and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Vujkovic, Alexandra and Walczak, Aleksandra M and Weber, Anna and Yin, Rose and Eugster, Anne and Sharma, Virag}, journal = {ImmunoInformatics}, volume = {9}, pages = {100024}, year = {2023}, doi = {10.1016/j.immuno.2023.100024}, }
2022
- iScienceInference of metabolic fluxes in nutrient-limited continuous cultures: A Maximum Entropy approach with the minimum informationJosé Antonio Pereiro-Morejón, Jorge FERNANDEZ-DE-COSSIO-DIAZ, and Roberto MuletiScience, 2022
We propose a new scheme to infer the metabolic fluxes of cell cultures in a chemostat. Our approach is based on the Maximum Entropy Principle and exploits the understanding of the chemostat dynamics and its connection with the actual metabolism of cells. We show that, in continuous cultures with limiting nutrients, the inference can be done with \it limited information about the culture: the dilution rate of the chemostat, the concentration in the feed media of the limiting nutrient and the cell concentration at steady state. Also, we remark that our technique provides information, not only about the mean values of the fluxes in the culture, but also its heterogeneity. We first present these results studying a computational model of a chemostat. Having control of this model we can test precisely the quality of the inference, and also unveil the mechanisms behind the success of our approach. Then, we apply our method to E. coli experimental data from the literature and show that it outperforms alternative formulations that rest on a Flux Balance Analysis framework.
@article{pereiromorejon2022inference, title = {Inference of metabolic fluxes in nutrient-limited continuous cultures: A Maximum Entropy approach with the minimum information}, author = {Pereiro-Morejón, José Antonio and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Mulet, Roberto}, journal = {iScience}, volume = {25}, number = {12}, pages = {105450}, year = {2022}, doi = {10.1016/j.isci.2022.105450}, }
2021
- Biotech. Bioeng.In-silico media optimization for continuous cultures using genome scale metabolic networks: The case of CHO-K1Bárbara A. Pérez-Fernández, Jorge FERNANDEZ-DE-COSSIO-DIAZ, Tammy Boggiano, Kalet León, and Roberto MuletBiotechnology and Bioengineering, 2021
The cell culture is the central piece of a biotechnological industrial process. It includes upstream (e.g. media preparation, fixed costs, etc.) and downstream steps (e.g. product purification, waste disposal, etc.). In the continuous mode of cell culture, a constant flow of fresh media replaces culture fluid until the system reaches a steady state. This steady state is the standard operation mode which, under very general conditions, is a function of the ratio between the cell density and the dilution rate and depends on the media supplied to the culture. To optimize the production process it is widely accepted that the concentration of the metabolites in this media should be carefully tuned. A poor media may not provide enough nutrients to the culture, while a media too rich in nutrients may be a waste of resources because, either the cells do not use all of the available nutrients, or worse, they over-consume them producing toxic byproducts. In this study, we show how an in-silico study of a genome scale metabolic network coupled to the dynamics of a chemostat could guide the strategy to optimize the media to be used in a continuous process. Given a known media we model the concentrations of the cells in a chemostat as a function of the dilution rate. Then, we cast the problem of optimizing the production process within a linear programming framework in which the goal is to minimize the cost of the media keeping fixed the cell concentration for a given dilution rate in the chemostat. We evaluate our results in two metabolic models: first a simplified model of mammalian cell metabolism, and then in a realistic genome-scale metabolic network of mammalian cells, the Chinese hamster ovary cell line. We explore the latter in more detail given specific meaning to the predictions of the concentrations of several metabolites.
@article{perezfernandez2021silico, title = {In-silico media optimization for continuous cultures using genome scale metabolic networks: The case of CHO-K1}, author = {Pérez-Fernández, Bárbara A. and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Boggiano, Tammy and León, Kalet and Mulet, Roberto}, journal = {Biotechnology and Bioengineering}, volume = {118}, number = {5}, pages = {1884--1897}, year = {2021}, doi = {10.1002/bit.27704}, } - IJMSAMaLa: Analysis of Directed Evolution Experiments via Annealed Mutational Approximated LandscapeLuca Sesta, Guido Uguzzoni, Jorge FERNANDEZ-DE-COSSIO-DIAZ, and Andrea PagnaniInternational Journal of Molecular Sciences, 2021
We present Annealed Mutational approximated Landscape (AMaLa), a new method to infer fitness landscapes from Directed Evolution experiments sequencing data. Such experiments typically start from a single wild-type sequence, which undergoes Darwinian in vitro evolution via multiple rounds of mutation and selection for a target phenotype. In the last years, Directed Evolution is emerging as a powerful instrument to probe fitness landscapes under controlled experimental conditions and as a relevant testing ground to develop accurate statistical models and inference algorithms (thanks to high-throughput screening and sequencing). Fitness landscape modeling either uses the enrichment of variants abundances as input, thus requiring the observation of the same variants at different rounds or assuming the last sequenced round as being sampled from an equilibrium distribution. AMaLa aims at effectively leveraging the information encoded in the whole time evolution. To do so, while assuming statistical sampling independence between sequenced rounds, the possible trajectories in sequence space are gauged with a time-dependent statistical weight consisting of two contributions: (i) an energy term accounting for the selection process and (ii) a generalized Jukes–Cantor model for the purely mutational step. This simple scheme enables accurately describing the Directed Evolution dynamics and inferring a fitness landscape that correctly reproduces the measures of the phenotype under selection (e.g., antibiotic drug resistance), notably outperforming widely used inference strategies. In addition, we assess the reliability of AMaLa by showing how the inferred statistical model could be used to predict relevant structural properties of the wild-type sequence.
@article{sesta2021amala, title = {AMaLa: Analysis of Directed Evolution Experiments via Annealed Mutational Approximated Landscape}, author = {Sesta, Luca and Uguzzoni, Guido and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Pagnani, Andrea}, journal = {International Journal of Molecular Sciences}, volume = {22}, number = {20}, pages = {10908}, year = {2021}, doi = {10.3390/ijms222010908}, }
2020
- Sci. Rep.A self-consistent probabilistic formulation for inference of interactionsJorge Fernandez-de-Cossio, Jorge FERNANDEZ-DE-COSSIO-DIAZ, and Yasser Perera-NegrinScientific Reports, 2020
Large molecular interaction networks are nowadays assembled in biomedical researches along with important technological advances. Diverse interaction measures, for which input solely consisting of the incidence of causal-factors, with the corresponding outcome of an inquired effect, are formulated without an obvious mathematical unity. Consequently, conceptual and practical ambivalences arise. We identify here a probabilistic requirement consistent with that input, and find, by the rules of probability theory, that it leads to a model multiplicative in the complement of the effect. Important practical properties are revealed along these theoretical derivations, that has not been noticed before.
@article{fernandezdecossio2020self, title = {A self-consistent probabilistic formulation for inference of interactions}, author = {Fernandez-de-Cossio, Jorge and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Perera-Negrin, Yasser}, journal = {Scientific Reports}, volume = {10}, number = {1}, pages = {21435}, year = {2020}, doi = {10.1038/s41598-020-78496-8}, } - PREStatistical mechanics of interacting metabolic networksJorge FERNANDEZ-DE-COSSIO-DIAZ and Roberto MuletPhysical Review E, 2020
We cast the metabolism of interacting cells within a statistical mechanics framework considering both, the actual phenotypic capacities of each cell and its interaction with its neighbors. Reaction fluxes will be the components of high-dimensional spin vectors, whose values will be constrained by the stochiometry and the energy requirements of the metabolism. Within this picture, finding the phenotypic states of the population turns out to be equivalent to searching for the equilibrium states of a disordered spin model. We provide a general solution of this problem for arbitrary metabolic networks and interactions. We apply this solution to a simplified model of metabolism and to a complex metabolic network, the central core of the \emphE. coli, and demonstrate that the combination of selective pressure and interactions define a complex phenotypic space. Cells may specialize in producing or consuming metabolites complementing each other at the population level and this is described by an equilibrium phase space with multiple minima, like in a spin-glass model.
@article{fernandezdecossiodiaz2020statistical, title = {Statistical mechanics of interacting metabolic networks}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Mulet, Roberto}, journal = {Physical Review E}, volume = {101}, number = {4}, pages = {042401}, year = {2020}, doi = {10.1103/physreve.101.042401}, } - MBEUnsupervised Inference of Protein Fitness Landscape from Deep Mutational ScanJorge FERNANDEZ-DE-COSSIO-DIAZ, Guido Uguzzoni, and Andrea PagnaniMolecular Biology and Evolution, 2020
The recent technological advances underlying the screening of large combinatorial libraries in high-throughput mutational scans deepen our understanding of adaptive protein evolution and boost its applications in protein design. Nevertheless, the large number of possible genotypes requires suitable computational methods for data analysis, the prediction of mutational effects, and the generation of optimized sequences. We describe a computational method that, trained on sequencing samples from multiple rounds of a screening experiment, provides a model of the genotype–fitness relationship. We tested the method on five large-scale mutational scans, yielding accurate predictions of the mutational effects on fitness. The inferred fitness landscape is robust to experimental and sampling noise and exhibits high generalization power in terms of broader sequence space exploration and higher fitness variant predictions. We investigate the role of epistasis and show that the inferred model provides structural information about the 3D contacts in the molecular fold.
@article{fernandezdecossiodiaz2020unsupervised, title = {Unsupervised Inference of Protein Fitness Landscape from Deep Mutational Scan}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Uguzzoni, Guido and Pagnani, Andrea}, journal = {Molecular Biology and Evolution}, volume = {38}, number = {1}, pages = {318--328}, year = {2020}, doi = {10.1093/molbev/msaa204}, } - Cell Death Dis.Formate induces a metabolic switch in nucleotide and energy metabolismKristell Oizel, Jacqueline Tait-Mulder, Jorge FERNANDEZ-DE-COSSIO-DIAZ, Matthias Pietzke, Holly Brunton, Sergio Lilla, and 9 more authorsCell Death & Disease, 2020
Formate is a precursor for the de novo synthesis of purine and deoxythymidine nucleotides. Formate also interacts with energy metabolism by promoting the synthesis of adenine nucleotides. Here we use theoretical modelling together with metabolomics analysis to investigate the link between formate, nucleotide and energy metabolism. We uncover that endogenous or exogenous formate induces a metabolic switch from low to high adenine nucleotide levels, increasing the rate of glycolysis and repressing the AMPK activity. Formate also induces an increase in the pyrimidine precursor orotate and the urea cycle intermediate argininosuccinate, in agreement with the ATP-dependent activities of carbamoyl-phosphate and argininosuccinate synthetase. In vivo data for mouse and human cancers confirms the association between increased formate production, nucleotide and energy metabolism. Finally, the in vitro observations are recapitulated in mice following and intraperitoneal injection of formate. We conclude that formate is a potent regulator of purine, pyrimidine and energy metabolism.
@article{oizel2020formate, title = {Formate induces a metabolic switch in nucleotide and energy metabolism}, author = {Oizel, Kristell and Tait-Mulder, Jacqueline and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Pietzke, Matthias and Brunton, Holly and Lilla, Sergio and Dhayade, Sandeep and Athineos, Dimitri and Blanco, Giovanny Rodriguez and Sumpton, David and Mackay, Gillian M. and Blyth, Karen and Zanivan, Sara R. and Meiser, Johannes and Vazquez, Alexei}, journal = {Cell Death & Disease}, volume = {11}, number = {5}, pages = {310}, year = {2020}, doi = {10.1038/s41419-020-2523-z}, }
2019
- Sci. Rep.Cell population heterogeneity driven by stochastic partition and growth optimalityJorge FERNANDEZ-DE-COSSIO-DIAZ, Roberto Mulet, and Alexei VazquezScientific Reports, 2019
A fundamental question in biology is how cell populations evolve into different subtypes based on homogeneous processes at the single cell level. Here we show that population bimodality can emerge even when biological processes are homogenous at the cell level and the environment is kept constant. Our model is based on the stochastic partitioning of a cell component with an optimal copy number. We show that the existence of unimodal or bimodal distributions depends on the variance of partition errors and the growth rate tolerance around the optimal copy number. In particular, our theory provides a consistent explanation for the maintenance of aneuploid states in a population. The proposed model can also be relevant for other cell components such as mitochondria and plasmids, whose abundances affect the growth rate and are subject to stochastic partition at cell division.
@article{fernandezdecossiodiaz2019cell, title = {Cell population heterogeneity driven by stochastic partition and growth optimality}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Mulet, Roberto and Vazquez, Alexei}, journal = {Scientific Reports}, volume = {9}, number = {1}, pages = {9406}, year = {2019}, doi = {10.1038/s41598-019-45882-w}, } - PLOS CBMaximum entropy and population heterogeneity in continuous cell culturesJorge FERNANDEZ-DE-COSSIO-DIAZ and Roberto MuletPLOS Computational Biology, 2019
Continuous cultures of mammalian cells are complex systems displaying hallmark phenomena of nonlinear dynamics, such as multi-stability, hysteresis, as well as sharp transitions between different metabolic states. In this context mathematical models may suggest control strategies to steer the system towards desired states. Although even clonal populations are known to exhibit cell-to-cell variability, most of the currently studied models assume that the population is homogeneous. To overcome this limitation, we use the maximum entropy principle to model the phenotypic distribution of cells in a chemostat as a function of the dilution rate. We consider the coupling between cell metabolism and extracellular variables describing the state of the bioreactor and take into account the impact of toxic byproduct accumulation on cell viability. We present a formal solution for the stationary state of the chemostat and show how to apply it in two examples. First, a simplified model of cell metabolism where the exact solution is tractable, and then a genome-scale metabolic network of the Chinese hamster ovary (CHO) cell line. Along the way we discuss several consequences of heterogeneity, such as: qualitative changes in the dynamical landscape of the system, increasing concentrations of byproducts that vanish in the homogeneous case, and larger population sizes.
@article{fernandezdecossiodiaz2019maximum, title = {Maximum entropy and population heterogeneity in continuous cell cultures}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Mulet, Roberto}, journal = {PLOS Computational Biology}, volume = {15}, number = {2}, pages = {e1006823}, year = {2019}, doi = {10.1371/journal.pcbi.1006823}, } - Sci. Rep.Directed evolution of super-secreted variants from phage-displayed human Interleukin-2Gertrudis Rojas, Tania Carmenate, Julio Felipe Santo-Tomás, Pedro A. Valiente, Marlies Becker, Annia Pérez-Riverón, and 6 more authorsScientific Reports, 2019
Selection from a phage display library derived from human Interleukin-2 (IL-2) yielded mutated variants with greatly enhanced display levels of the functional cytokine on filamentous phages. Introduction of a single amino acid replacement selected that way (K35E) increased the secretion levels of IL-2-containing fusion proteins from human transfected host cells up to 20-fold. Super-secreted (K35E) IL-2/Fc is biologically active in vitro and in vivo , has anti-tumor activity and exhibits a remarkable reduction in its aggregation propensity- the major manufacturability issue limiting IL-2 usefulness up to now. Improvement of secretion was also shown for a panel of IL-2-engineered variants with altered receptor binding properties, including a selective agonist and a super agonist that kept their unique properties. Our findings will improve developability of the growing family of IL-2-derived immunotherapeutic agents and could have a broader impact on the engineering of structurally related four-alpha-helix bundle cytokines.
@article{rojas2019directed, title = {Directed evolution of super-secreted variants from phage-displayed human Interleukin-2}, author = {Rojas, Gertrudis and Carmenate, Tania and Santo-Tomás, Julio Felipe and Valiente, Pedro A. and Becker, Marlies and Pérez-Riverón, Annia and Tundidor, Yaima and Ortiz, Yaquelín and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Graça, Luis and Dübel, Stefan and León, Kalet}, journal = {Scientific Reports}, volume = {9}, number = {1}, pages = {800}, year = {2019}, doi = {10.1038/s41598-018-37280-5}, }
2018
- Sci. Rep.A physical model of cell metabolismJorge FERNANDEZ-DE-COSSIO-DIAZ and Alexei VazquezScientific Reports, 2018
Cell metabolism is characterized by three fundamental energy demands: to sustain cell maintenance, to trigger aerobic fermentation and to achieve maximum metabolic rate. The transition to aerobic fermentation and the maximum metabolic rate are currently understood based on enzymatic cost constraints. Yet, we are lacking a theory explaining the maintenance energy demand. Here we report a physical model of cell metabolism that explains the origin of these three energy scales. Our key hypothesis is that the maintenance energy demand is rooted on the energy expended by molecular motors to fluidize the cytoplasm and counteract molecular crowding. Using this model and independent parameter estimates we make predictions for the three energy scales that are in quantitative agreement with experimental values. The model also recapitulates the dependencies of cell growth with extracellular osmolarity and temperature. This theory brings together biophysics and cell biology in a tractable model that can be applied to understand key principles of cell metabolism.
@article{fernandezdecossiodiaz2018physical, title = {A physical model of cell metabolism}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Vazquez, Alexei}, journal = {Scientific Reports}, volume = {8}, number = {1}, pages = {8349}, year = {2018}, doi = {10.1038/s41598-018-26724-7}, }
2017
- PLOS CBCharacterizing steady states of genome-scale metabolic networks in continuous cell culturesJorge FERNANDEZ-DE-COSSIO-DIAZ, Kalet Leon, and Roberto MuletPLOS Computational Biology, 2017
We present a model for continuous cell culture coupling intra-cellular metabolism to extracellular variables describing the state of the bioreactor, taking into account the growth capacity of the cell and the impact of toxic byproduct accumulation. We provide a method to determine the steady states of this system that is tractable for metabolic networks of arbitrary complexity. We demonstrate our approach in a toy model first, and then in a genome-scale metabolic network of the Chinese hamster ovary cell line, obtaining results that are in qualitative agreement with experimental observations. More importantly, we derive a number of consequences from the model that are independent of parameter values. First, that the ratio between cell density and dilution rate is an ideal control parameter to fix a steady state with desired metabolic properties invariant across perfusion systems. This conclusion is robust even in the presence of multi-stability, which is explained in our model by the negative feedback loop on cell growth due to toxic byproduct accumulation. Moreover, a complex landscape of steady states in continuous cell culture emerges from our simulations, including multiple metabolic switches, which also explain why cell-line and media benchmarks carried out in batch culture cannot be extrapolated to perfusion. On the other hand, we predict invariance laws between continuous cell cultures with different parameters. A practical consequence is that the chemostat is an ideal experimental model for large-scale high-density perfusion cultures, where the complex landscape of metabolic transitions is faithfully reproduced. Thus, in order to actually reflect the expected behavior in perfusion, performance benchmarks of cell-lines and culture media should be carried out in a chemostat.
@article{fernandezdecossiodiaz2017characterizing, title = {Characterizing steady states of genome-scale metabolic networks in continuous cell cultures}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Leon, Kalet and Mulet, Roberto}, journal = {PLOS Computational Biology}, volume = {13}, number = {11}, pages = {e1005835}, year = {2017}, doi = {10.1371/journal.pcbi.1005835}, } - Sci. Rep.Limits of aerobic metabolism in cancer cellsJorge FERNANDEZ-DE-COSSIO-DIAZ and Alexei VazquezScientific Reports, 2017
Cancer cells exhibit high rates of glycolysis and glutaminolysis. Glycolysis can provide energy and glutaminolysis can provide carbon for anaplerosis and reductive carboxylation to citrate. However, all these metabolic requirements could be in principle satisfied from glucose. Here we investigate why cancer cells do not satisfy their metabolic demands using aerobic biosynthesis from glucose. Based on the typical composition of a mammalian cell we quantify the energy demand and the OxPhos burden of cell biosynthesis from glucose. Our calculation demonstrates that aerobic growth from glucose is feasible up to a minimum doubling time that is proportional to the OxPhos burden and inversely proportional to the mitochondria OxPhos capacity. To grow faster cancer cells must activate aerobic glycolysis for energy generation and uncouple NADH generation from biosynthesis. To uncouple biosynthesis from NADH generation cancer cells can synthesize lipids from carbon sources that do not produce NADH in their catabolism, including acetate and the amino acids glutamate, glutamine, phenylalanine and tyrosine. Finally, we show that cancer cell lines have an OxPhos capacity that is insufficient to support aerobic biosynthesis from glucose. We conclude that selection for high rate of biosynthesis implies a selection for aerobic glycolysis and uncoupling biosynthesis from NADH generation.
@article{fernandezdecossiodiaz2017limits, title = {Limits of aerobic metabolism in cancer cells}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Vazquez, Alexei}, journal = {Scientific Reports}, volume = {7}, number = {1}, pages = {13488}, year = {2017}, doi = {10.1038/s41598-017-14071-y}, } - Sci. Rep.Microenvironmental cooperation promotes early spread and bistability of a Warburg-like phenotypeJorge FERNANDEZ-DE-COSSIO-DIAZ, Andrea De Martino, and Roberto MuletScientific Reports, 2017
We introduce an in silico model for the initial spread of an aberrant phenotype with Warburg-like overflow metabolism within a healthy homeostatic tissue in contact with a nutrient reservoir (the blood), aimed at characterizing the role of the microenvironment for aberrant growth. Accounting for cellular metabolic activity, competition for nutrients, spatial diffusion and their feedbacks on aberrant replication and death rates, we obtain a phase portrait where distinct asymptotic whole-tissue states are found upon varying the tissue-blood turnover rate and the level of blood-borne primary nutrient. Over a broad range of parameters, the spreading dynamics is bistable as random fluctuations can impact the final state of the tissue. Such a behaviour turns out to be linked to the re-cycling of overflow products by non-aberrant cells. Quantitative insight on the overall emerging picture is provided by a spatially homogeneous version of the model.
@article{fernandezdecossiodiaz2017microenvironmental, title = {Microenvironmental cooperation promotes early spread and bistability of a Warburg-like phenotype}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and De Martino, Andrea and Mulet, Roberto}, journal = {Scientific Reports}, volume = {7}, number = {1}, pages = {3103}, year = {2017}, doi = {10.1038/s41598-017-03342-3}, } - arXivMissing and spurious interaction in additive, multiplicative and odds ratio modelsJorge Fernandez-de-Cossio, Jorge FERNANDEZ-DE-COSSIO-DIAZ, Toshifumi Takao, and Yasser PereraarXiv preprint, 2017
Additive, multiplicative, and odd ratio neutral models for interactions are for long advocated and controversial in epidemiology. We show here that these commonly advocated models are biased, leading to spurious interactions, and missing true interactions.
@article{jorgefernandezdecossio2017missing, title = {Missing and spurious interaction in additive, multiplicative and odds ratio models}, author = {Fernandez-de-Cossio, Jorge and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Takao, Toshifumi and Perera, Yasser}, journal = {arXiv preprint}, year = {2017}, }
2016
- JSTATFast inference of ill-posed problems within a convex spaceJorge FERNANDEZ-DE-COSSIO-DIAZ and R MuletJournal of Statistical Mechanics: Theory and Experiment, 2016
In multiple scientific and technological applications we face the problem of having low dimensional data to be justified by a linear model defined in a high dimensional parameter space. The difference in dimensionality makes the problem ill-defined: the model is consistent with the data for many values of its parameters. The objective is to find the probability distribution of parameter values consistent with the data, a problem that can be cast as the exploration of a high dimensional convex polytope. In this work we introduce a novel algorithm to solve this problem efficiently. It provides results that are statistically indistinguishable from currently used numerical techniques while its running time scales linearly with the system size. We show that the algorithm performs robustly in many abstract and practical applications. As working examples we simulate the effects of restricting reaction fluxes on the space of feasible phenotypes of a genome scale Escherichia coli metabolic network and infer the traffic flow between origin and destination nodes in a real communication network.
@article{fernandezdecossiodiaz2016fast, title = {Fast inference of ill-posed problems within a convex space}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Mulet, R}, journal = {Journal of Statistical Mechanics: Theory and Experiment}, volume = {2016}, number = {7}, pages = {073207}, year = {2016}, doi = {10.1088/1742-5468/2016/07/073207}, }
2014
- EJM B/FluidsFree wave modes in elliptic cylindrical containersM. Oliva-Leyva, Jorge FERNANDEZ-DE-COSSIO-DIAZ, and C. Trallero-GinerEuropean Journal of Mechanics - B/Fluids, 2014
The linear theory of unforced surface gravity–capillary waves in cylindrical containers with an elliptical cross-section is studied in detail. General solutions for the velocity potential and the free surface amplitude are given in terms of Mathieu functions. Our numerical results show the dependence of the natural frequencies on the fluid properties and the eccentricity e of the container cross-section. The well-known case of a circular tank for e = 0 is retrieved and remarkable crossings of the mode frequencies for certain values of e are found. The frequency shift and the wall damping ratio due to viscous dissipation in the Stokes boundary layers are evaluated numerically. The effect of the viscous dissipation in the bulk, the wall damping ratio, is estimated.
@article{olivaleyva2014free, title = {Free wave modes in elliptic cylindrical containers}, author = {Oliva-Leyva, M. and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Trallero-Giner, C.}, journal = {European Journal of Mechanics - B/Fluids}, volume = {43}, pages = {185--190}, year = {2014}, doi = {10.1016/j.euromechflu.2013.09.003}, }
2013
- PRLOptimally Designed Quantum Transport across Disordered NetworksMattia Walschaers, Jorge FERNANDEZ-DE-COSSIO-DIAZ, Roberto Mulet, and Andreas BuchleitnerPhysical Review Letters, 2013
We establish a general mechanism for highly efficient quantum transport through finite, disordered 3D networks. It relies on the interplay of disorder with centro-symmetry and a dominant doublet spectral structure, and can be controlled by proper tuning of only coarse-grained quantities. Photosynthetic light harvesting complexes are discussed as potential biological incarnations of this design principle.
@article{walschaers2013optimally, title = {Optimally Designed Quantum Transport across Disordered Networks}, author = {Walschaers, Mattia and {FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Mulet, Roberto and Buchleitner, Andreas}, journal = {Physical Review Letters}, volume = {111}, number = {18}, pages = {180601}, year = {2013}, doi = {10.1103/physrevlett.111.180601}, }
2012
- Anal. Chem.Computation of Isotopic Peak Center-Mass Distribution by Fourier TransformJorge FERNANDEZ-DE-COSSIO-DIAZ and Jorge Fernandez-de-CossioAnalytical Chemistry, 2012
We derive a new efficient algorithm for the computation of the isotopic peak center-mass distribution of a molecule. With the use of Fourier transform techniques, the algorithm accurately computes the total abundance and average mass of all the isotopic species with the same number of nucleons. We evaluate the performance of the method with 10 benchmark proteins and other molecules; results are compared with BRAIN, a recently reported polynomial method. The new algorithm is comparable to BRAIN in accuracy and superior in terms of speed and memory, particularly for large molecules. An implementation of the algorithm is available for download.
@article{fernandezdecossiodiaz2012computation, title = {Computation of Isotopic Peak Center-Mass Distribution by Fourier Transform}, author = {{FERNANDEZ-DE-COSSIO-DIAZ}, Jorge and Fernandez-de-Cossio, Jorge}, journal = {Analytical Chemistry}, volume = {84}, number = {16}, pages = {7052--7056}, year = {2012}, doi = {10.1021/ac301296a}, }
software
- RestrictedBoltzmannMachines.jl: Restricted Boltzmann machines in Julia.
More on GitHub.