A study involving Princeton University, the UK AI Security Institute, Stanford University, and the University of Toronto shows that while current cutting-edge AI agents can independently complete many engineering tasks in the research process, they are still significantly lacking in making original scientific contributions and struggle to produce papers that can be accepted by top machine learning conferences.
Complete the entire process within six days
The research, published Wednesday and titled "Can AI agents conduct open AI research?", involved a team that instead of using standard tests with pre-defined answers. Instead, they presented the AI with the core questions from two unpublished NeurIPS 2026 papers to avoid the system finding ready-made answers directly from training data or networks.
Researchers provided each AI agent with six days, along with thousands of dollars in API credits, GPU resources, internet access, and virtual machine environments, requiring them to independently produce papers that met the standards for submission to academic conferences.
The papers were written, but all were rejected.
The results showed that these systems completed a significant amount of research work, including literature retrieval, software debugging, experiment running, GPU resource management, and generating complete papers without human intervention. However, after review by the original research authors, both AI-generated papers failed to pass the review.
The research team believes that this type of test reflects AI's scientific reasoning ability better than traditional benchmarks because it deals with open-ended research questions rather than well-defined, fixed tasks. The reviewers all agree on one point: the system can complete the engineering steps, but it still cannot offer sufficiently novel research contributions to justify publication in top-tier conferences.
Original scientific research remains a weakness
The study also summarized five recurring failure patterns, which it believes are the main reasons why AI papers fail to meet publishable standards. The original abstract did not elaborate on the specifics of these five problems, but the overall conclusion is clear: AI can handle many execution tasks in the research process, but it still struggles to produce truly original scientific research.
The authors also noted that this test only covered two research projects, resulting in a small sample size, and that the AI-generated papers were reviewed by the original researchers, which introduced certain limitations. However, they still believe the results are sufficient to demonstrate that cutting-edge AI agents are rapidly taking over the engineering aspects of scientific research, but there is still a significant gap between them and independently completing original research.
Additional information:Before and after the release of this study, the academic community has been continuously monitoring the anomalous behavior of autonomous AI agents. In May, researchers from the University of California, Riverside, Microsoft, and Nvidia reported that AI agents often made dangerous or irrational actions when performing objectives. OpenAI also disclosed this month that one of its cutting-edge AI agents attempted to cheat in a cybersecurity benchmark test, bypassed restrictions to access Hugging Face, and subsequently accessed four other online services.











