All posts

Science

Anthropic's wet lab reports a CRISPR-like enzyme family surfaced by Claude agents

About 950 Claude agent sessions mined sequence data for reverse transcriptases and flagged ART, a phage system with CRISPR-like repeats. Humans ran the experiments, and its function is unknown.

HackHoster Team · · 11 min read

Electron micrograph, rotated and cropped, of Enterobacteria phage T2, a well-studied phage with an oval head and long tail, not one of the jumbo phages that carry ART
Photo: SnaxMikn / Wikimedia Commons, CC BY-SA 4.0

At a glance

  • Anthropic said on September 23 that Claude agents found ART, a family of phage reverse transcriptases paired with arrays of repeated DNA that resemble CRISPR.
  • The campaign surveyed 1.94 billion protein clusters in 21.5 hours using 949 agent sessions and 215.6 million tokens, with no human intervention.
  • ART arrays hold 3 to 21 repeats of 15 to 49 nucleotides separated by spacers of 120 to 220 nucleotides, and no cas genes sit nearby.
  • In phage SA1, array-derived RNA made up as much as 8% of phage transcripts 15 minutes after infection.
  • Ten reruns of the same campaign all missed the array, and in tool-based tests models read too little DNA to spot it in 39% of attempts.

On September 23 Anthropic described the human-staffed molecular biology lab it has run in the San Francisco Bay Area since the spring, and announced its first result: a family of phage genes it calls array-associated reverse transcriptases, or ART. Claude agents found the family by mining sequence databases. Human scientists did all of the lab work.

Reverse transcriptases (RTs) copy RNA into DNA. The ART systems pair an RT with a dedicated partner gene and a long array of evenly spaced DNA repeats, a layout that resembles a CRISPR array. Anthropic is open about the main gap: nobody yet knows what the system does.

The more durable lesson may be about the agents rather than the enzyme. A companion technical report, which has not been peer reviewed, shows exactly how the agents stumbled on ART, and then shows that the same campaign failed to find it again in ten reruns. That combination is rare in AI-for-science announcements, and useful for anyone building agents that work through data.

What Anthropic announced

The announcement introduces a new life sciences research group. Its stated plan is to have Claude explore DNA datasets for protein families nobody has characterised, generate hypotheses at scale, and test the best ones at the bench. Anthropic says the lab works only at biosafety levels 1 and 2, does not handle pathogens that infect humans, and that every experiment is performed by people. The team uses Claude Science and Claude Code, the same tools sold to outside scientists, plus an in-house harness that runs many Claude sessions in parallel.

The technical report lists six authors, all at Anthropic: Peter Yoon, Januka Athukoralage, Emmanuel Ameisen, Eric Kauderer-Abrams, Nicholas Perry and Matthew Durrant. The Scientist identifies Kauderer-Abrams as Anthropic's head of life sciences. Citing his interview with Reuters, it reports that the lab will not run clinical trials and does not intend to compete with drug companies, and that Reuters' unnamed sources described a longer-term goal of Claude directing robotic lab equipment.

Surface model of HIV-1 reverse transcriptase in green and blue, bound to a nucleic acid double helix, with the polymerase and nuclease sites circled
Crystal structure of HIV-1 reverse transcriptase (PDB 3KLF), a well-studied member of the enzyme class. The ART enzymes come from phages and are so far described only through predicted structures. Image: Thomas Splettstoesser / Wikimedia Commons, CC BY-SA 3.0

Reverse transcriptases, retrons and CRISPR

Reverse transcriptase was discovered in 1970, independently by Howard Temin in Rous sarcoma virus and by David Baltimore in murine leukemia virus and Rous sarcoma virus. Both shared the 1975 Nobel Prize in Physiology or Medicine for the work. Retroviruses such as HIV depend on the enzyme, and biotechnology relies on it for RT-PCR and for making complementary DNA libraries from messenger RNA.

Black-and-white photo of Howard Temin, a smiling man with curly hair and thick glasses in a suit and checked tie, holding a pair of headphones
Howard Temin, who co-discovered reverse transcriptase in 1970, photographed in 1984. Photo: Willy Pragher / Wikimedia Commons, CC BY 4.0

Bacteria carry their own RTs. The best known are group II introns and retrons. Anthropic's announcement notes that in recent years researchers have found many more bacterial RT families, mostly acting as parts of immune systems against phages, and nearly all of them through genome mining: searching sequence databases for uncharacterised genes, spotting the odd ones and working out what they do.

Definition. A retron is a small bacterial gene cluster with three parts: a non-coding RNA, a reverse transcriptase that copies part of that RNA into a single-stranded DNA called msDNA, and a partner protein. Work published from 2020 onward showed that retrons help defend bacteria against phages.

Diagram of a retron operon: a promoter followed by msr and msd regions that make msDNA, and a ret gene with seven conserved boxes that encodes reverse transcriptase
Layout of a bacterial retron: the msr and msd regions encode the RNA and DNA parts of msDNA, and the ret gene encodes the reverse transcriptase. ART shares this three-part design but carries many RNA units. Diagram: Stigmatella aurantiaca / Wikimedia Commons, CC BY-SA 3.0

CRISPR started the same way, as an anomaly in sequence data. Yoshizumi Ishino's group noticed unusual clustered repeats in E. coli DNA in 1987. The acronym came in 2001, the realisation that spacers come from phage DNA in 2005, experimental proof of an adaptive immune function in 2007, and the programmable Cas9 system in 2012, which led to the 2020 Nobel Prize in Chemistry. Anthropic's announcement places ART in that lineage of discoveries that began with someone noticing an odd pattern, alongside restriction enzymes and the heat-stable Taq polymerase from a Yellowstone hot spring.

Simplified diagram of a CRISPR locus with cas genes on the left, a leader sequence, and an array of identical grey repeat boxes separated by coloured spacer bars
A simplified CRISPR locus: cas genes, a leader, and a repeat-spacer array. ART arrays look similar but have far longer spacers and no cas genes nearby. Diagram: AnnaJune / Wikimedia Commons, CC BY-SA 3.0

How the agent harness works

The technical report describes the harness in enough detail to copy. A human-written research brief, in this case asking for new RT systems defined by new partner-gene associations, is converted into stages such as building search profiles, searching the database and classifying hits. Each stage is split into tasks. Every task gets a worker agent, a Claude Code instance that plans and executes, and a supervisor agent that reviews the plan and results and can send them back. Of 119 tasks, 49 were revised at least once.

Supervisors can also open new tasks when a worker notices something, and 98 of the 119 tasks came from agents rather than from the brief. A curator agent writes each finding into a shared knowledge base that later agents see, an editor agent reviews reports before filing, and everything is committed to a version-controlled record. Agents had a shell, file access and web search, plus connectors to a metagenomic protein database and literature search. The campaign ran on Claude Mythos 5.

Campaign stepCount
Metagenomic protein clusters surveyed1.94 billion
RT clusters recovered after filtering198,290
RT classes assigned9
RT loci sampled for neighbouring genes10,983
Candidate partner families scored3,564
Families promoted to deep dives16, plus 1 from a follow-up
New RT partner associations confirmed3
New RT lineages, including ART3
Reports filed19
Agent sessions, agent-hours, tokens949, 77, 215.6 million
Wall-clock time21.5 hours, no human intervention

Anthropic's announcement rounds these to about 950 agents, 21 hours and 210 million tokens. To decide which of the 19 reports humans should read first, a Mythos 5 judge compared every ordered pair (342 games) on impact, novelty, soundness and actionability, and a Bradley–Terry model turned the wins into a ranking. The ART report came third, winning 32 of its 36 games.

The moment of discovery

ART came out of a detour. Agents first flagged a group of phage RTs because they sat next to a subunit of the jumbo-phage RNA polymerase. A worker rejected that link as an artefact of gene order, but queued a follow-up on the RT itself because a retron-like RT inside a jumbo phage seemed noteworthy. Retron RTs work with an RNA encoded just upstream, so the supervisor asked the next worker to check that upstream region.

The first RTs had almost no room upstream, a median of 27 base pairs. Their closest relatives had long non-coding stretches, a median of 940. The worker read one of those stretches, about 2,900 letters of raw DNA, straight into its context and wrote that it could see a tandem repeat array by eye. It then questioned whether the pattern was already known, wrote a script to count repeats (one locus had 14 copies of a 16-letter repeat), compared the layout with known non-coding elements including CRISPR arrays, ran a literature search and filed a report.

What ART looks like

Human scientists then defined the family in interactive sessions. Searches of Anthropic's database and public genomes found 95 distinct ART RT clusters in jumbo phages (very large bacterial viruses) and predicted viral DNA, and 28 of them carry a detectable array upstream of the RT. Phylogenetic analysis placed ART next to the retrons. The RT has an unusually long front end, about 180 amino acids ahead of the polymerase domain where other RTs have about 50 or fewer, and it keeps the catalytic YxDD motif in all 93 members whose sequences cover it.

Two transmission electron micrographs of a jumbo phage with a large hexagonal head and a tail, next to 500-nanometre scale bars
A jumbo phage, fMGyn-Pae01, which infects Pseudomonas aeruginosa, shown with an extended tail (A) and contracted tail (B). It is an example of the very large phages where ART was found, not one known to carry ART. Image: Kira Ranta, Mikael Skurnik, Salja Kilijunen / Wikimedia Commons, CC BY 4.0

The arrays are what make the family distinctive:

FeatureART arraysTypical CRISPR arrays
Repeats3 to 21 copies of 15 to 49 nucleotides, with a palindromic core of about 15Typically 28 to 37 base pairs
Spacer length120 to 220 nucleotidesAbout 30 nucleotides
Spacers between related genomesKept in order between related phagesGained and lost between strains
Neighbouring genesRT plus one of three partner families, no cas genescas genes

The partner genes fall into three unrelated types. Type I, found at 59 of the 95 loci, encodes a protein of about 600 amino acids with two acetyltransferase-like folds. Type II, limited to the Staphylococcus phage clade, is about 270 amino acids, and type III is about 170. The authors read the fact that unrelated partners attach to the same RT family as evidence that the pairing evolved more than once.

What the lab work shows so far

Two observations suggest the arrays matter. A Claude Science session found and reanalysed public RNA sequencing data from a published infection time course of Staphylococcus phage SA1, which carries an ART system. At 15 minutes after infection, RNA from the array made up as much as 8% of phage transcripts, among the most abundant in the phage, and the reads broke into shorter pieces with boundaries that repeated across replicates.

Colourised scanning electron micrograph of clustered round purple bacteria
Staphylococcus aureus bacteria (MRSA) under a scanning electron microscope. Phage SA1, whose ART system was studied, infects a related Staphylococcus species. Image: NIAID / Wikimedia Commons, CC BY 2.0

The team then expressed the SA1 system from plasmids in E. coli. Small-RNA sequencing again found discrete short RNAs cut from the array, and folding predictions placed each repeat in a base-paired stem. The authors' working hypothesis is a retron-like system that uses a bank of different RNAs instead of one, with each unit RNA forming its own complex with the RT and partner, perhaps so the system can respond to different triggers. They write that they have not shown the RT is active, that the RNAs are its substrates, whether the RT and partner interact, or what the system does for the phage.

Key caveat. ART is a sequence-level discovery with early RNA evidence. Its function, its activity and any use as a tool are all unknown, and the report has not been peer reviewed.

Why the find did not reproduce

Anthropic ran the same campaign ten more times. Agents sampled ART loci in nearly every run, and in two runs workers followed up on the lineage, but none read the upstream DNA and the array was missed every time. The authors put this down to the large search space and the harness's non-deterministic behaviour.

To dig in, they built a fixed benchmark. Seven Claude models each made 100 attempts at five levels of difficulty, 3,500 runs in all, and a Mythos 5 judge scored each report against ten features of the ART family. The four strongest models (Opus 5.5, Mythos 5.1, Mythos 5 and Opus 5) clearly beat the other three (Opus 4.6, Opus 4.8 and Sonnet 5).

The surprise was that more tools made array detection worse. With the sequences pasted directly into context, the strongest models described the array correctly in at least 90% of attempts. Given the same data as files plus analysis tools, success fell as low as 32%. In 39% of the file-based attempts, the model never read a contiguous stretch of 200 nucleotides or more, so it never saw more than about one repeat unit. When it did read that much, recognition rose by 16 to 32 percentage points, and with the most DNA in context Mythos 5 spotted the array in up to 96% of attempts.

The report also looks inside the model. Genome language models such as Evo 2 and gLM2 already pick out repeats in DNA. Replaying the discovery transcript through Mythos 5, the team found two internal signals that rose as the repeats were read and went largely quiet when each repeat copy was shuffled. Neither signal is specific to DNA. The authors argue these signals help explain the discovery, though the evidence comes from replaying a single session.

Reactions and caveats

Feng Zhang of MIT and the Broad Institute, one of the pioneers of CRISPR genome editing, reviewed the preprint for Anthropic. He called the RNA-repeat arrays "genuinely intriguing and merits further investigation," and described the work as a good example of AI agents contributing to biological discovery.

Feng Zhang in a tuxedo holding an open medal case beside NSF Director France Córdova in a formal reception room with flags
Feng Zhang receiving the National Science Foundation's Alan T. Waterman Award from NSF Director France Córdova in 2014. Photo: National Science Foundation / Wikimedia Commons, Public domain

Other points worth weighing:

  • Not peer reviewed. The Scientist notes the result is a technical report.
  • Not entirely unknown. The Scientist reports that one of the RTs had been characterised in earlier research. The report itself says the original genome description of the MarsHill phage locus identified the RT and proposed an upstream RNA, but missed the repeats and the partner gene.
  • Other methods find such arrays. Similar non-coding arrays were recently reported beside an unrelated RT family, UG27, using a purpose-built genome language model, which suggests the architecture evolved more than once and that specialised tools can find it too.
  • Access is uneven. The Scientist points out that the discovery used 210 million tokens, and that academics had already raised concerns in July about choosing between paying for AI tokens and paying trainees.

What builders can take from it

The reruns and benchmarks are the most transferable part of the report, because the failure mode is common in agent systems of every kind.

  • Check what your agents actually read. Log how many bytes of raw data each tool call returns into context. Agents that summarise files through scripts may never look at the pattern you care about.
  • Treat one run as one sample. A single striking agent result is an anecdote until it reproduces. Budget for reruns and report the hit rate.
  • Separate worker and reviewer roles. A supervisor that can send work back, and a shared, versioned record of plans and results, made this campaign auditable after the fact.
  • Rank outputs before humans read them. Pairwise tournaments with a fixed rubric are a cheap way to triage dozens of agent reports.
  • Keep transcripts. The team could explain the discovery only because every reasoning step and tool call was logged.

Practical tip. If your agent analyses logs, traces or sequences through tools, add a rule that it must read a raw sample of each input directly before writing code to summarise it.

What to watch

As of September 24, Anthropic says further experiments on ART are under way and that it wants to work with outside scientists who propose research questions. The open questions are whether the ART RT is active, whether the array RNAs are its templates, and what the system does for jumbo phages, possibly in competition with other phages. The other thing to watch is whether the report goes through peer review, and whether other groups can find ART-like arrays with their own pipelines.

Sources