The search campaign
Anthropic’s preprint reports a 21.5-hour computational campaign in which 949 Claude Code sessions, organised as workers and supervisors, searched a clustered database of about 1.9 billion protein sequences for unusual reverse-transcriptase systems. The run used about 215.6 million tokens and produced 19 final reports. The computational campaign ran without human intervention, while researchers performed the biological follow-up afterwards.
The workflow split 119 tasks across agents, sampled more than 200,000 representative reverse-transcriptase clusters, and narrowed 3,564 protein families to 17 candidates after prevalence filters. One report drew attention to a reverse transcriptase found next to a tandem repeat array and a conserved accessory gene, a combination the authors call array-associated reverse transcriptases, or ARTs.
What the agents surfaced
The reverse transcriptase itself was already represented in databases, and related enzymes had been described previously. The proposed novelty is the broader genomic arrangement: the enzyme, an upstream repeat array and a neighbouring partner gene, with many examples found in jumbo phages.
The authors analysed published RNA-sequencing data from phage-infected cells and then expressed an ART locus on plasmids in Escherichia coli. They report discrete short RNAs from the repeat array in both settings. That supports expression of the array as separate RNA units, but it does not establish what the reverse transcriptase does with those RNAs or whether the partner protein is required for a biological effect.
The rerun problem
The same discovery campaign was then run ten more times. According to the preprint, nearly every completed rerun sampled ART loci and two investigated the lineage as follow-up, yet none read the upstream DNA where the repeat array sits. The defining array was therefore missed in all ten reruns.
A fixed-input benchmark separated recognition from search. When models were directly shown ART loci, the most capable systems recognised the array at least 90% of the time. With command-line tools, performance fell as low as 32%; 39% of tool attempts never read a contiguous DNA segment of at least 200 nucleotides. The authors conclude that the discovery bottleneck was often navigation and tool use rather than recognising the pattern once it was visible.
What remains unknown
ART’s biological function is still unknown. The paper proposes a retron-like model in which several distinct RNAs could interact with one reverse transcriptase and partner protein, but this remains a hypothesis. The study does not demonstrate programmable gene editing, CRISPR activity or a confirmed antiviral mechanism.
The work is a preprint from Anthropic researchers and has not been peer reviewed or independently reproduced. It provides a concrete example of an agent-generated hypothesis reaching wet-lab follow-up, while the failed reruns show that the autonomous search path is highly non-deterministic. Independent genomic replication and functional experiments are still needed to establish the system’s novelty and biological role.