The motivation: simulation as infrastructure for a bio-industrial age
 

We are pursuing the challenge of a new industrial age in which sustainability is not a constraint bolted onto chemical manufacturing but the consequence of switching its basis — from petrochemistry and high-temperature catalysis to biological processes that run at ambient conditions on renewable inputs. Enzymes replace heavy-metal catalysts; engineered cells replace reactors; grown and self-assembled materials replace energy-intensive fabrication; biologics and precision therapeutics replace blunt small molecules. This transition is currently paced by empiricism: directed evolution, screening, and formulation are slow, costly, and largely unpredictive. A general-purpose biomolecular simulator that could predict rather than screen would be the enabling infrastructure of that age — the equivalent of computational modelling for aerospace or semiconductors. But the simulation stack we have was built to answer structural biology’s question — “what is this molecule and how does it bind” — and the bio-industrial age asks a different one: what does this molecule do, in context, over time, and can we make it and trust it. 

The cross-cutting challenge 

Almost every high-value gap shares one root: today’s tools assume dilute conditions, a single equilibrium structure, fixed chemistry, and reachable timescales. The bio-industrial questions violate all four. The grand challenge is to reorient simulation from binding in a box to consequence in a body, a cell, or a material, over time — non-equilibrium, kinetic, reactive, and multiscale by default. 

Therapeutics 

The dominant challenge is not affinity — the comp-chem industry already pursues that — but molecular fate. Specifically: predicting how the body chemically transforms a drug (metabolism), into what, and how fast, at a throughput that informs design. This is reactive enzyme chemistry (cytochrome P450 oxidation), historically blocked by the cost of per-compound QM/MM. Reaction-specific machine-learned potentials are the credible route, because they move the generalisation burden from infinite drug-space to the small, finite repertoire of metabolic reaction chemistries. The residual, under-discussed wall is the trustworthiness of the reference quantum chemistry for iron-oxo catalysis. Alongside this sit kinetics/residence time (over equilibrium affinity) and proteome-scale selectivity as design-time safety. 

Industrial biotechnology 

The bottleneck is manufacturability, not function: whether an engineered enzyme or protein will express, fold in vivo, avoid aggregation, and work in the crowded, non-equilibrium cellular context rather than a dilute simulation box. A related grand challenge is predictive, high-throughput enzyme catalysis — ranking mutations by their effect on catalytic rate (a transition-state problem) before wet-lab screening. This is pursued but remains unsolved at the accuracy and throughput design requires, and cracking it would displace the largest empirical cost in biocatalysis. 

Biomaterials 

Here the material is its processing history: properties emerge from non-equilibrium assembly, from evolving chemistry (crosslinking, self-healing, degradation, where network topology is the dependent variable), and from the mesoscale between atoms and continuum. The under-solved challenges are simulating assembly-and-processing itself, reactive materials whose bonds form and break during the simulation, and failure and lifetime under load — not just the equilibrium native state. Coarse-graining and mesoscale methods have real communities, but a principled, chemistry-preserving bridge across scales remains out of reach. 

The delivery approach 

A simulator this broad cannot be built as a monolith, and should not be. The tractable path is component by component: each module tackles a specific, well-posed question that sits inside one of the grand challenges — a reaction-specific potential for metabolic oxidation, an in-cell crowding environment, an assembly-and-processing engine, a mutation-to-rate estimator — built and validated on its own terms before being composed with the others. This mirrors how the challenges themselves decompose: they are not one problem but a family of physics, and each yields to a focused effort where a general assault would stall. 

Each component should be released as professionally maintained open-source software — not as a code dump accompanying a paper, but as durable, documented, tested infrastructure with real maintenance behind it. This matters for three reasons. First, hardening: a component only becomes trustworthy when a community can stress it against problems its authors never imagined, and open, maintained code is the only mechanism that generates that adversarial testing at scale. Second, composability: openly specified, well-engineered modules can be built on and combined by others, so the simulator grows through contribution rather than solely through one organisation’s effort. Third, trust: bio-industrial and clinical decisions cannot rest on black boxes, and open methods are auditable in a way proprietary tools are not. 

Written by Julien Michel and proofed by Claude.