On September 23rd, Anthropic disclosed an early-stage biological experiment: Claude independently completed data retrieval, data analysis, and candidate screening under the condition that scientists only provided high-level objectives, identifying an enzyme system related to clustered DNA repeat sequences. Since the morphology of these repeat sequences resembles CRISPR, the news was easily packaged as " AI discovers a new CRISPR ". However, the official statement did not conclude this way. More accurately, the team discovered a set of candidate enzymes that had not been described by the system before, and further laboratory verification is still needed to confirm their functions, mechanisms, and biological significance.
This work comes from the life science research program launched by Anthropic this spring. The question the program aims to answer is not complicated: can a general model string together literature reading, database queries, code analysis, and hypothesis generation into a truly useful scientific research process, rather than just making paper summaries more coherent? The results obtained this time are among the first set of public cases, indicating that the model is already able to narrow down the search space for open-ended problems, but it is still far from crossing the threshold of experimental verification.
The real innovation lies in the model organizing a long and messy search process on its own.
In bioinformatics research, it is often not a case of a lack of data, but rather an abundance of it. Researchers need to cross-check between genomics, protein structures, species distributions, and existing papers to determine whether a certain pattern is merely coincidental or whether it represents a clue worth pursuing with actual experiments. The role played by Claude in this task is more akin to that of an tireless computational researcher: it first breaks down the problem, then selects the appropriate tools, writes analysis code, checks intermediate results, and continues to follow up on any anomalies that arise.
Anthropic emphasizes that what scientists provide is a high-level direction, rather than step-by-step operational instructions. This detail is crucial. If every step were designed by humans, the model would merely be a faster script generator; it is only when the model can decide for itself what to look at, what to compare, and what to discard that it begins to affect the scientific research process itself. However, "autonomy" does not equate to unmonitored operation. The research questions, scope of data, available tools, and final judgments are still the responsibility of human teams, and the candidates identified by the model must still undergo independent review.
Such tasks also expose the most difficult aspect of evaluating AI's scientific research. For chatbots, answering a question correctly can be scored using a standard answer; however, in open-ended research, there is no ready-made answer, and success often depends on whether a direction that has not been previously emphasized can be found and can be tested through experiments. It is of no value for a model to generate hundreds of thousands of seemingly reasonable hypotheses; only by narrowing down the candidates to a few that the laboratory is willing to invest time in can real savings in research costs be achieved.
Therefore, the most important indicator in this case is not how much code the model has written, but rather the "verifiability" of the clues. Whether the candidate sequences can be reproduced in independent datasets, whether the relevant enzymes have the activity predicted, and whether repetitive structures are involved in defense, regulation, or other mechanisms, all these need to be answered by subsequent experiments. If any one of these steps fails, it could turn what seems like a compelling computational result into an ordinary false positive.
From candidate clues to biological conclusions, there is an entire chain of evidence in between.
The reason CRISPR has shifted its focus to biotechnology is that researchers have not only observed special repetitive sequences but have also gradually proven their immune functions, molecular mechanisms, and programmability. Anthropic The system disclosed this time is currently at a more advanced stage. Describing "similar morphology" as "same function" or referring to "candidate discoveries" as "new gene editing tools" would be an exaggeration of the existing evidence.
The next steps in reality include running the analysis again, eliminating database annotation errors, checking for species and sample biases, and then designing experiments, combining or splitting them. Even if experiments confirm the activity of a molecule, it is still necessary to determine whether it is stable, selective, and capable of functioning in different cellular environments. Molecular systems that are to be used in medical or industrial applications must also undergo safety, manufacturability, and regulatory assessments. Today's announcement did not provide these conclusions.
But being cautious doesn’t mean that this result is unimportant. One of the most costly aspects of basic research is deciding which experiment to conduct first within a vast range of possibilities. If a general model can consistently provide high-quality, reproducible candidates, it could improve the utilization rate of experimental equipment and researchers. AI will not replace pipettes, nor can it replace the scientific judgment of abnormal results, but it may change the way preparations are done before experiments begin.
For research institutions, a more reasonable approach now is "model proposal, pipeline review, experimental validation." The search logs of the models, code versions, and data sources should be fully preserved; key conclusions need to be replicated using another method; negative results should also be recorded to prevent teams from only presenting successful cases. Only by making the failure rate public can the outside world determine whether this represents a replicable research capability or merely a lucky coincidence.
Anthropic still needs to address scalability issues: whether a single case can be transferred to other protein families, how dependent the model is on clues that already exist in the training data but are not prominent, and whether different models or hints can lead to similar candidates. If these tests are successful, AI's value in life sciences will shift from 'helping researchers write papers' to 'helping researchers decide what to do next.' For now, the most appropriate assessment is that it represents an early result worth serious verification, rather than a completed biological breakthrough.











