All posts

Science

SyntheMol-RL designs synthecin, an antibiotic candidate that curbed MRSA wound infections in mice

A reinforcement learning model from McMaster and Stanford picked drug candidates from 46 billion easy-to-make molecules. Of 79 synthesized, 13 were potent against S. aureus.

HackHoster Team · · 10 min read

Scanning electron micrograph of clusters of round, golden-colored MRSA bacteria on a dark green background
Photo: Janice Haney Carr, CDC, public domain

At a glance

  • SyntheMol-RL, published in Molecular Systems Biology on April 23, 2026, searches 46 billion molecules that commercial vendors can build on demand.
  • Of 196 compounds synthesized and tested, 13 of the 79 from the two reinforcement learning models stopped S. aureus growth at 8 micrograms per millilitre or less.
  • The graph-neural-network version hit 11 of 38, while the older tree-search version and 44 random compounds produced no hits at all.
  • Mice with MRSA-infected wounds treated with 2% synthecin cream carried about 100 times fewer bacteria after 24 hours than vehicle-treated controls.
  • The code is on GitHub under an MIT licence, and the authors say its mechanism of action is still unknown.

Researchers at McMaster University and Stanford published SyntheMol-RL in Molecular Systems Biology on April 23. It is a generative model that designs small-molecule drug candidates, with one important restriction: every molecule it proposes can be built from commercially available building blocks using known chemical reactions. Pointed at Staphylococcus aureus, it proposed compounds from which the team synthesized 79. Thirteen were potent against the bacterium in lab tests, and one of them, which the team named synthecin, held back a drug-resistant MRSA infection in a mouse wound model when applied as a cream.

The paper's first author is Kyle Swanson of Stanford's computer science department, with James Zou at Stanford and Jonathan Stokes at McMaster as corresponding authors. McMaster names graduate student Gary Liu as the lead developer of the model and Denise Catacutan as the lead on the wet-lab work.

A modern glass and stone medical building at McMaster University on a sunny day, with young trees in front
A McMaster University medical school building in Hamilton, Ontario. SyntheMol-RL came out of Jonathan Stokes's lab at McMaster. Photo: Captain108 / Wikimedia Commons, CC BY-SA 4.0

Why a new antibiotic is hard to find

Staphylococcus aureus is a common bacterium. The US Centers for Disease Control and Prevention says roughly one in three people carry staph on their skin or in their nose without trouble. The methicillin-resistant form, MRSA, resists several antibiotics at once, spreads through contact with infected people or wounds, and most often shows up as a skin infection, though it can reach the lungs, the bloodstream and surgical sites.

Inside Precision Medicine, reporting on the study, frames the stakes: drug-resistant bacteria were linked to close to 5 million deaths worldwide in 2019, a figure projected to double by 2050 if no new antibiotics arrive.

The discovery pipeline has not kept pace. Alexander Fleming's 1928 observation of penicillin opened a few decades in which new classes arrived regularly, and then the well ran dry. Most antibiotics in use descend from scaffolds found in the mid-20th century, and commercial incentives are weak: a drug taken for a week, and held in reserve to slow resistance, does not earn what a chronic-disease drug earns.

Alexander Fleming seated at a cluttered laboratory bench, examining a culture dish
Alexander Fleming in his laboratory at St Mary's Hospital, London, during the Second World War, 14 years after his discovery of penicillin. Photo: Ministry of Information Photo Division photographer / Wikimedia Commons, public domain

Machine learning entered the field in February 2020, when Jonathan Stokes, then at MIT, and colleagues published a deep-learning screen in Cell. They trained a neural network on 2,335 molecules tested against E. coli, then applied it to libraries of more than 107 million compounds. The standout was halicin, a molecule from the Drug Repurposing Hub that is structurally unlike conventional antibiotics, kills a broad range of pathogens, and treated Clostridioides difficile and pan-resistant Acinetobacter baumannii infections in mice. From 23 predictions tested out of the ZINC15 database, the model turned up eight more antibacterial compounds unlike known drugs.

Chemical structure diagram of halicin, showing two linked rings with amino and nitro groups
Halicin, identified in 2020 by a deep-learning screen of existing molecules. It is structurally unlike conventional antibiotics. Diagram: Meodipt / Wikimedia Commons, public domain

That approach screens molecules that already exist. Generating new ones is harder, because most generative chemistry models produce structures freely and many of the results turn out to be difficult or impossible to make.

Only design what a chemist can actually build

SyntheMol takes the opposite approach. It assembles molecules from building blocks and reaction templates taken from two make-on-demand catalogs, Enamine REAL Space and WuXi GalaXi. According to the paper, the 2022 version of Enamine REAL covers about 31 billion molecules from 139,517 building blocks and 169 reactions, of which the team used the 13 most common, while WuXi GalaXi covers about 16 billion molecules from roughly 15,000 building blocks and 36 reactions. Together that is about 46 billion molecules. McMaster rounds the ingredients to roughly 150,000 building blocks and 50 reactions. Anything SyntheMol suggests can be ordered.

Make-on-demand space, defined. Vendors such as Enamine and WuXi publish catalogs not of molecules in stock but of molecules they are confident they can synthesize on request, enumerated by combining stocked building blocks with reliable reactions. Searching that space means every hit comes with a recipe.

The first version, published in Nature Machine Intelligence in March 2024, searched this space with Monte Carlo tree search and was used against Acinetobacter baumannii. That work covered about 30 billion molecules built from around 132,000 fragments and 13 reactions. The team synthesized 58 of the generated molecules and found six structurally novel ones with antibacterial activity. James Zou said at the time that the model generates the recipe for each new molecule as well as the molecule itself, which matters because chemists otherwise have no route to an AI-designed structure.

The new version replaces the tree search with reinforcement learning. A learned value function predicts how good the molecules built from a given building block are likely to be, and the search samples building blocks in proportion to that value rather than always taking the best one, which keeps some exploration alive while exploiting promising regions. Each run does 10,000 rollouts and yields about 10,000 unique molecules.

It can also optimize several properties at once. Each property gets its own value model, and their scores are combined with weights that adjust over time to steer the search toward regions where molecules meet every threshold. In this study the targets were antibacterial activity against S. aureus and water solubility. McMaster notes that the two often pull in opposite directions, which is why the team built solubility into generation itself rather than filtering for it afterwards. Liu describes that as the main change from the 2024 model. Stokes put the broader point plainly to McMaster: bleach is antibacterial, and so is fire, but neither ticks the other boxes.

Hoover Tower and the sandstone buildings of Stanford University's main quad in evening light
Stanford University, where first author Kyle Swanson and co-senior author James Zou are based. The wet-lab work was done at McMaster. Photo: King of Hearts / Wikimedia Commons, CC BY-SA 3.0

How the model was trained and filtered

The team tried two property predictors inside the search. Chemprop-RDKit is a graph neural network whose 300-dimensional learned representation is concatenated with 200 computed chemical features. MLP-RDKit is a faster feed-forward network that uses only those 200 features.

For antibacterial activity, the predictors trained on 10,658 molecules screened against S. aureus at 50 micromolar, of which 1,137 (10.7%) counted as active. For solubility, they used 9,982 molecules from AqSolDB. Under ten-fold cross-validation, the Chemprop-RDKit activity model reached an ROC-AUC of 0.875 and the solubility model an R-squared of 0.822. Training each model took under 80 minutes on eight CPUs and one GPU, which is modest by modern standards.

Generation is only the first step. The paper describes a filtering cascade applied to the output: keep molecules whose predicted activity clears 0.5 and whose predicted solubility clears a threshold; drop anything too similar to the training set or to known antibiotics in ChEMBL; thin the remainder so no two survivors are closely similar to each other; take the top 150 by predicted activity per method; then keep the 50 with the lowest predicted clinical toxicity from the ADMET-AI tool.

A researcher in a biosafety cabinet using a multichannel pipette over a 96-well plate
High-throughput liquid handling with a multichannel pipette and a 96-well plate, the format used for the activity screening behind the model's training data. Photo: Siduduziwe Nxumalo / Wikimedia Commons, CC BY-SA 4.0

In total the team ordered 250 compounds. Enamine successfully made 177 of 224 (79%) and WuXi 17 of 24 (71%), giving 196 unique compounds that were tested.

The scale of the search is worth comparing with the alternatives. McMaster notes that a physical screen in a lab can realistically cover about a million molecules, several orders of magnitude below the 46 billion the model searches. The paper also ran a conventional virtual screen as a control, scoring 21 million existing molecules with the same predictors, which took about seven days of compute.

What the lab results show

A compound counted as potent if it stopped bacterial growth at 8 micrograms per millilitre or less. The comparison across generation methods is the heart of the paper.

MethodPotent hitsHit rate
Reinforcement learning with Chemprop11 of 3829%
Reinforcement learning with the faster MLP2 of 415%
Monte Carlo tree search (the 2024 approach)0 of 380%
Conventional virtual screen2 of 375%
Randomly chosen compounds0 of 440%

The paper reports the advantage of the Chemprop-based model over each other group as statistically significant after correcting for multiple comparisons. The spread between the two reinforcement learning variants is striking: same search, same chemical space, different property predictor, and a six-fold difference in hit rate.

After a detailed literature search, seven of the 13 hits were judged structurally novel. The other six overlapped with salicylanilides, a known antibacterial class.

Chemical structure diagram of salicylanilide, two rings joined by an amide group with a hydroxyl substituent
Salicylanilide, a long-known antibacterial scaffold. Six of the 13 hits overlapped with this class; the other seven survived a manual novelty check. Diagram: Edgar181 / Wikimedia Commons, public domain

Synthecin was the one taken into animals. The paper describes it as distinct from the salicylanilides, sharing some chemical groups with them but built around a different, more flexible core. It held its potency against USA300, a community-acquired MRSA strain, and against four vancomycin-intermediate S. aureus isolates from the CDC's resistance panel. It is narrow-spectrum: the paper reports no notable activity against the ESKAPE panel organisms such as E. coli, Pseudomonas aeruginosa and Klebsiella pneumoniae. In broth it is bacteriostatic, holding bacteria in check rather than killing them outright.

A simple illustration of a grey laboratory mouse
Five mice per group were used in the wound infection experiment. Illustration: Ryan Kissinger, NIAID, public domain

In the mouse experiment, C57BL/6N mice were given MRSA-infected skin wounds and treated with a cream containing 2% synthecin, or with the vehicle alone, five times over the first 20 hours. Each group had five animals. At 24 hours, vehicle-treated wounds carried about 6.4 billion bacteria per gram of tissue, while synthecin-treated wounds carried about 51 million, roughly the level present when treatment started and more than a hundred times lower than the controls. The paper reports the difference as statistically significant and notes that treated wounds showed no notable tissue inflammation, while control wounds did.

Caveats, and the authors' own criticism

This is an early candidate, not a drug. Synthecin has been tested only as a topical treatment, in one small animal experiment, against one species. The team has not yet worked out how it inhibits bacteria, a step Stokes says is key to judging its safety, and McMaster says mechanism-of-action work is under way. Toxicity entered the pipeline only as a computational prediction used to choose which compounds to make, which is not the same as testing it.

Key caveat. The potency threshold here is a laboratory measurement of growth inhibition, and the animal work is a five-mouse topical model over 24 hours. Neither says anything yet about whether the compound is safe or useful in people.

Narrow spectrum cuts both ways. A compound that acts on S. aureus and leaves E. coli and Pseudomonas alone spares a patient's other bacteria, which is often desirable, but it also means a diagnosis has to come first and the market is smaller. The same is true of being bacteriostatic rather than bactericidal in broth: holding an infection in check is useful on a wound, less so in a bloodstream infection.

The authors are candid about the model's limits. Only part of what it generates turns out to be active, which they attribute to the accuracy of the property predictors, and they write that this shows the need to improve those models so that high-scoring compounds translate into real hits. They also report that the settings controlling property weighting and exploration are highly sensitive and inconsistent between models, and that generated molecules tend to cluster around shared building blocks rather than spreading across the space. The search strongly favored Enamine compounds, which the paper links to Enamine's much larger set of building blocks.

There is one number that makes the synthesizability constraint concrete. When the team compared against GFlowNet and REINVENT 4, two generative models that design molecules freely, the vendors would only offer limited synthesis of those molecules, at up to 62 times the cost and up to 5.7 times the synthesis time of the SyntheMol-RL compounds.

What builders can take from it

For developers, the idea that travels is constraining generation to things that can actually be executed, the same instinct behind grammar-constrained decoding in language models. Generating freely and filtering afterwards wastes most of the search; generating inside a space where every point is reachable does not.

The code is open source on GitHub in the SyntheMol repository under an MIT license, and its 2.0.0 release added the reinforcement learning version. It installs with pip, supports Chemprop, Chemprop-RDKit, MLP-RDKit and random-forest predictors, and the repository says it runs on a standard laptop with 16 GB of memory, with GPUs useful mainly for faster model training.

A few things to check if you want to work with it or build something similar:

  • Match the predictor to the task. The 11-of-38 versus 2-of-41 gap in this study came down to the property model, not the search.
  • Budget for the filter cascade. Novelty, diversity and predicted toxicity cut a generated set down hard before anything gets ordered.
  • Mind the catalog bias. The search followed the larger building-block set. If you use two sources, check whether results reflect chemistry or inventory.
  • Hold out a test set you trust. The paper's random-compound arm, with zero hits from 44 molecules, is what makes the headline rate meaningful.
  • Count the whole loop. Ordering 250 compounds yielded 196 that could be made and tested. Synthesis success, not just generation, sets the pace of a campaign like this.

Stokes told Inside Precision Medicine that the tool was built to be disease agnostic, and could generate candidates for diabetes or cancer just as easily. The framework is described as compatible with any property predictor and any combinatorial chemical space, so the substitution is a matter of swapping the scoring model and the catalog.

What to watch

McMaster says the lab expects a more robust version of SyntheMol later this year, and the paper was selected for the June issue of Molecular Systems Biology. The nearest-term scientific question is synthecin's mechanism, which the team is working on now and which gates any assessment of its safety. Beyond that, the open questions are whether the seven structurally novel hits hold up under broader testing, and whether the same pipeline produces comparable hit rates when pointed at a different target.

Sources