Science
SyntheMol-RL designs synthecin, an antibiotic candidate that curbed MRSA wound infections in mice
A reinforcement learning model from McMaster and Stanford picked drug candidates from 46 billion easy-to-make molecules. Of 79 synthesized, 13 were potent against S. aureus.
HackHoster Team · · 10 min read

At a glance
- SyntheMol-RL, published in Molecular Systems Biology on April 23, 2026, searches 46 billion molecules that commercial vendors can build on demand.
- Of 196 compounds synthesized and tested, 13 of the 79 from the two reinforcement learning models stopped S. aureus growth at 8 micrograms per millilitre or less.
- The graph-neural-network version hit 11 of 38, while the older tree-search version and 44 random compounds produced no hits at all.
- Mice with MRSA-infected wounds treated with 2% synthecin cream carried about 100 times fewer bacteria after 24 hours than vehicle-treated controls.
- The code is on GitHub under an MIT licence, and the authors say its mechanism of action is still unknown.
Researchers at McMaster University and Stanford published SyntheMol-RL in Molecular Systems Biology on April 23. It is a generative model that designs small-molecule drug candidates, with one important restriction: every molecule it proposes can be built from commercially available building blocks using known chemical reactions. Pointed at Staphylococcus aureus, it proposed compounds from which the team synthesized 79. Thirteen were potent against the bacterium in lab tests, and one of them, which the team named synthecin, held back a drug-resistant MRSA infection in a mouse wound model when applied as a cream.
The paper's first author is Kyle Swanson of Stanford's computer science department, with James Zou at Stanford and Jonathan Stokes at McMaster as corresponding authors. McMaster names graduate student Gary Liu as the lead developer of the model and Denise Catacutan as the lead on the wet-lab work.

Why a new antibiotic is hard to find
Staphylococcus aureus is a common bacterium. The US Centers for Disease Control and Prevention says roughly one in three people carry staph on their skin or in their nose without trouble. The methicillin-resistant form, MRSA, resists several antibiotics at once, spreads through contact with infected people or wounds, and most often shows up as a skin infection, though it can reach the lungs, the bloodstream and surgical sites.
Inside Precision Medicine, reporting on the study, frames the stakes: drug-resistant bacteria were linked to close to 5 million deaths worldwide in 2019, a figure projected to double by 2050 if no new antibiotics arrive.
The discovery pipeline has not kept pace. Alexander Fleming's 1928 observation of penicillin opened a few decades in which new classes arrived regularly, and then the well ran dry. Most antibiotics in use descend from scaffolds found in the mid-20th century, and commercial incentives are weak: a drug taken for a week, and held in reserve to slow resistance, does not earn what a chronic-disease drug earns.

Machine learning entered the field in February 2020, when Jonathan Stokes, then at MIT, and colleagues published a deep-learning screen in Cell. They trained a neural network on 2,335 molecules tested against E. coli, then applied it to libraries of more than 107 million compounds. The standout was halicin, a molecule from the Drug Repurposing Hub that is structurally unlike conventional antibiotics, kills a broad range of pathogens, and treated Clostridioides difficile and pan-resistant Acinetobacter baumannii infections in mice. From 23 predictions tested out of the ZINC15 database, the model turned up eight more antibacterial compounds unlike known drugs.

That approach screens molecules that already exist. Generating new ones is harder, because most generative chemistry models produce structures freely and many of the results turn out to be difficult or impossible to make.
Only design what a chemist can actually build
SyntheMol takes the opposite approach. It assembles molecules from building blocks and reaction templates taken from two make-on-demand catalogs, Enamine REAL Space and WuXi GalaXi. According to the paper, the 2022 version of Enamine REAL covers about 31 billion molecules from 139,517 building blocks and 169 reactions, of which the team used the 13 most common, while WuXi GalaXi covers about 16 billion molecules from roughly 15,000 building blocks and 36 reactions. Together that is about 46 billion molecules. McMaster rounds the ingredients to roughly 150,000 building blocks and 50 reactions. Anything SyntheMol suggests can be ordered.
Make-on-demand space, defined. Vendors such as Enamine and WuXi publish catalogs not of molecules in stock but of molecules they are confident they can synthesize on request, enumerated by combining stocked building blocks with reliable reactions. Searching that space means every hit comes with a recipe.
The first version, published in Nature Machine Intelligence in March 2024, searched this space with Monte Carlo tree search and was used against Acinetobacter baumannii. That work covered about 30 billion molecules built from around 132,000 fragments and 13 reactions. The team synthesized 58 of the generated molecules and found six structurally novel ones with antibacterial activity. James Zou said at the time that the model generates the recipe for each new molecule as well as the molecule itself, which matters because chemists otherwise have no route to an AI-designed structure.
The new version replaces the tree search with reinforcement learning. A learned value function predicts how good the molecules built from a given building block are likely to be, and the search samples building blocks in proportion to that value rather than always taking the best one, which keeps some exploration alive while exploiting promising regions. Each run does 10,000 rollouts and yields about 10,000 unique molecules.
It can also optimize several properties at once. Each property gets its own value model, and their scores are combined with weights that adjust over time to steer the search toward regions where molecules meet every threshold. In this study the targets were antibacterial activity against S. aureus and water solubility. McMaster notes that the two often pull in opposite directions, which is why the team built solubility into generation itself rather than filtering for it afterwards. Liu describes that as the main change from the 2024 model. Stokes put the broader point plainly to McMaster: bleach is antibacterial, and so is fire, but neither ticks the other boxes.

How the model was trained and filtered
The team tried two property predictors inside the search. Chemprop-RDKit is a graph neural network whose 300-dimensional learned representation is concatenated with 200 computed chemical features. MLP-RDKit is a faster feed-forward network that uses only those 200 features.
For antibacterial activity, the predictors trained on 10,658 molecules screened against S. aureus at 50 micromolar, of which 1,137 (10.7%) counted as active. For solubility, they used 9,982 molecules from AqSolDB. Under ten-fold cross-validation, the Chemprop-RDKit activity model reached an ROC-AUC of 0.875 and the solubility model an R-squared of 0.822. Training each model took under 80 minutes on eight CPUs and one GPU, which is modest by modern standards.
Generation is only the first step. The paper describes a filtering cascade applied to the output: keep molecules whose predicted activity clears 0.5 and whose predicted solubility clears a threshold; drop anything too similar to the training set or to known antibiotics in ChEMBL; thin the remainder so no two survivors are closely similar to each other; take the top 150 by predicted activity per method; then keep the 50 with the lowest predicted clinical toxicity from the ADMET-AI tool.

In total the team ordered 250 compounds. Enamine successfully made 177 of 224 (79%) and WuXi 17 of 24 (71%), giving 196 unique compounds that were tested.
The scale of the search is worth comparing with the alternatives. McMaster notes that a physical screen in a lab can realistically cover about a million molecules, several orders of magnitude below the 46 billion the model searches. The paper also ran a conventional virtual screen as a control, scoring 21 million existing molecules with the same predictors, which took about seven days of compute.
What the lab results show
A compound counted as potent if it stopped bacterial growth at 8 micrograms per millilitre or less. The comparison across generation methods is the heart of the paper.
| Method | Potent hits | Hit rate |
|---|---|---|
| Reinforcement learning with Chemprop | 11 of 38 | 29% |
| Reinforcement learning with the faster MLP | 2 of 41 | 5% |
| Monte Carlo tree search (the 2024 approach) | 0 of 38 | 0% |
| Conventional virtual screen | 2 of 37 | 5% |
| Randomly chosen compounds | 0 of 44 | 0% |
The paper reports the advantage of the Chemprop-based model over each other group as statistically significant after correcting for multiple comparisons. The spread between the two reinforcement learning variants is striking: same search, same chemical space, different property predictor, and a six-fold difference in hit rate.
After a detailed literature search, seven of the 13 hits were judged structurally novel. The other six overlapped with salicylanilides, a known antibacterial class.

Synthecin was the one taken into animals. The paper describes it as distinct from the salicylanilides, sharing some chemical groups with them but built around a different, more flexible core. It held its potency against USA300, a community-acquired MRSA strain, and against four vancomycin-intermediate S. aureus isolates from the CDC's resistance panel. It is narrow-spectrum: the paper reports no notable activity against the ESKAPE panel organisms such as E. coli, Pseudomonas aeruginosa and Klebsiella pneumoniae. In broth it is bacteriostatic, holding bacteria in check rather than killing them outright.

In the mouse experiment, C57BL/6N mice were given MRSA-infected skin wounds and treated with a cream containing 2% synthecin, or with the vehicle alone, five times over the first 20 hours. Each group had five animals. At 24 hours, vehicle-treated wounds carried about 6.4 billion bacteria per gram of tissue, while synthecin-treated wounds carried about 51 million, roughly the level present when treatment started and more than a hundred times lower than the controls. The paper reports the difference as statistically significant and notes that treated wounds showed no notable tissue inflammation, while control wounds did.
Caveats, and the authors' own criticism
This is an early candidate, not a drug. Synthecin has been tested only as a topical treatment, in one small animal experiment, against one species. The team has not yet worked out how it inhibits bacteria, a step Stokes says is key to judging its safety, and McMaster says mechanism-of-action work is under way. Toxicity entered the pipeline only as a computational prediction used to choose which compounds to make, which is not the same as testing it.
Key caveat. The potency threshold here is a laboratory measurement of growth inhibition, and the animal work is a five-mouse topical model over 24 hours. Neither says anything yet about whether the compound is safe or useful in people.
Narrow spectrum cuts both ways. A compound that acts on S. aureus and leaves E. coli and Pseudomonas alone spares a patient's other bacteria, which is often desirable, but it also means a diagnosis has to come first and the market is smaller. The same is true of being bacteriostatic rather than bactericidal in broth: holding an infection in check is useful on a wound, less so in a bloodstream infection.
The authors are candid about the model's limits. Only part of what it generates turns out to be active, which they attribute to the accuracy of the property predictors, and they write that this shows the need to improve those models so that high-scoring compounds translate into real hits. They also report that the settings controlling property weighting and exploration are highly sensitive and inconsistent between models, and that generated molecules tend to cluster around shared building blocks rather than spreading across the space. The search strongly favored Enamine compounds, which the paper links to Enamine's much larger set of building blocks.
There is one number that makes the synthesizability constraint concrete. When the team compared against GFlowNet and REINVENT 4, two generative models that design molecules freely, the vendors would only offer limited synthesis of those molecules, at up to 62 times the cost and up to 5.7 times the synthesis time of the SyntheMol-RL compounds.
What builders can take from it
For developers, the idea that travels is constraining generation to things that can actually be executed, the same instinct behind grammar-constrained decoding in language models. Generating freely and filtering afterwards wastes most of the search; generating inside a space where every point is reachable does not.
The code is open source on GitHub in the SyntheMol repository under an MIT license, and its 2.0.0 release added the reinforcement learning version. It installs with pip, supports Chemprop, Chemprop-RDKit, MLP-RDKit and random-forest predictors, and the repository says it runs on a standard laptop with 16 GB of memory, with GPUs useful mainly for faster model training.
A few things to check if you want to work with it or build something similar:
- Match the predictor to the task. The 11-of-38 versus 2-of-41 gap in this study came down to the property model, not the search.
- Budget for the filter cascade. Novelty, diversity and predicted toxicity cut a generated set down hard before anything gets ordered.
- Mind the catalog bias. The search followed the larger building-block set. If you use two sources, check whether results reflect chemistry or inventory.
- Hold out a test set you trust. The paper's random-compound arm, with zero hits from 44 molecules, is what makes the headline rate meaningful.
- Count the whole loop. Ordering 250 compounds yielded 196 that could be made and tested. Synthesis success, not just generation, sets the pace of a campaign like this.
Stokes told Inside Precision Medicine that the tool was built to be disease agnostic, and could generate candidates for diabetes or cancer just as easily. The framework is described as compatible with any property predictor and any combinatorial chemical space, so the substitution is a matter of swapping the scoring model and the catalog.
What to watch
McMaster says the lab expects a more robust version of SyntheMol later this year, and the paper was selected for the June issue of Molecular Systems Biology. The nearest-term scientific question is synthecin's mechanism, which the team is working on now and which gates any assessment of its safety. Beyond that, the open questions are whether the seven structurally novel hits hold up under broader testing, and whether the same pipeline produces comparable hit rates when pointed at a different target.
Sources
- SyntheMol-RL, a flexible reinforcement learning framework for designing easily synthesizable antibiotics (Molecular Systems Biology)
- McMaster-built AI model speeds up drug discovery, designs new antibiotic (McMaster News)
- AI tool creates designer antibiotics (Inside Precision Medicine)
- SyntheMol source code (GitHub)
- Researchers invent artificial intelligence model to design new superbug-fighting antibiotics (ScienceDaily, 2024)
- A deep learning approach to antibiotic discovery (Cell, 2020)
- About MRSA (US Centers for Disease Control and Prevention)
More from the blog

Science ·
Anthropic's wet lab reports a CRISPR-like enzyme family surfaced by Claude agents
About 950 Claude agent sessions mined sequence data for reverse transcriptases and flagged ART, a phage system with CRISPR-like repeats. Humans ran the experiments, and its function is unknown.
11 min read

Science ·
Six proteomic aging clocks read younger in rentosertib's lung fibrosis trial
A Nature Biotechnology study ran six protein-based aging clocks on serum from Insilico's IPF drug trial. The signal is consistent but small, early and hard to separate from the lung disease.
11 min read

Science ·
MIT and Oak Ridge's CrysVCD builds charge balance into AI crystal generation
A small transformer writes charge-balanced formulas before a diffusion model builds the crystal. Tuned for stability, 85% of outputs were predicted metastable. The code is open.
10 min read