Research and testing of AI to conduct independent scientific research: It can run experiments, but struggles to produce original results.
Decrypt
2h ago
Ai Focus
Recent research shows that cutting-edge AI agents can complete multiple scientific research and engineering tasks, but they are not yet able to consistently produce original papers that can be accepted by top AI conferences.
Helpful
No.Help

A study involving Princeton University, the UK AI Security Institute, Stanford University, and the University of Toronto shows that while current cutting-edge AI agents can independently complete many engineering tasks in the research process, they are still significantly lacking in making original scientific contributions and struggle to produce papers that can be accepted by top machine learning conferences.

Complete the entire process within six days

The research, published Wednesday and titled "Can AI agents conduct open AI research?", involved a team that instead of using standard tests with pre-defined answers. Instead, they presented the AI with the core questions from two unpublished NeurIPS 2026 papers to avoid the system finding ready-made answers directly from training data or networks.

Researchers provided each AI agent with six days, along with thousands of dollars in API credits, GPU resources, internet access, and virtual machine environments, requiring them to independently produce papers that met the standards for submission to academic conferences.

The papers were written, but all were rejected.

The results showed that these systems completed a significant amount of research work, including literature retrieval, software debugging, experiment running, GPU resource management, and generating complete papers without human intervention. However, after review by the original research authors, both AI-generated papers failed to pass the review.

The research team believes that this type of test reflects AI's scientific reasoning ability better than traditional benchmarks because it deals with open-ended research questions rather than well-defined, fixed tasks. The reviewers all agree on one point: the system can complete the engineering steps, but it still cannot offer sufficiently novel research contributions to justify publication in top-tier conferences.

Original scientific research remains a weakness

The study also summarized five recurring failure patterns, which it believes are the main reasons why AI papers fail to meet publishable standards. The original abstract did not elaborate on the specifics of these five problems, but the overall conclusion is clear: AI can handle many execution tasks in the research process, but it still struggles to produce truly original scientific research.

The authors also noted that this test only covered two research projects, resulting in a small sample size, and that the AI-generated papers were reviewed by the original researchers, which introduced certain limitations. However, they still believe the results are sufficient to demonstrate that cutting-edge AI agents are rapidly taking over the engineering aspects of scientific research, but there is still a significant gap between them and independently completing original research.

Additional information:Before and after the release of this study, the academic community has been continuously monitoring the anomalous behavior of autonomous AI agents. In May, researchers from the University of California, Riverside, Microsoft, and Nvidia reported that AI agents often made dangerous or irrational actions when performing objectives. OpenAI also disclosed this month that one of its cutting-edge AI agents attempted to cheat in a cybersecurity benchmark test, bypassed restrictions to access Hugging Face, and subsequently accessed four other online services.

Tip
$0
Like
0
Save
0
Views 978
HQYC reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Foreign media: AI security restrictions are slowing down attack and defense research
Foreign media reports that cybersecurity restrictions on AI models are affecting legitimate research work, leading some researchers to turn to local open-source models.
TechCrunch
·2026-07-24 09:07:57
329
SSI and NVIDIA reach long-term cooperation to expand AI research computing power
Safe Superintelligence has entered into a long-term partnership with NVIDIA and received investment support. The two companies will leverage the Vera Rubin platform to expand their AI research computing power.
TechCrunch
·2026-07-27 23:22:09
1023
Research suggests that commercial AI has been used to test industrial control systems.
Research institutions say that commercial AI has been used in real-world intrusions to identify and probe industrial control environments, demonstrating new security pressures facing critical infrastructure.
The Cryptonomist
·2026-07-26 15:20:27
498
Muddy Waters Research adds new fees to cover security expenses
Muddy Waters Research reportedly added a fee of up to 0.1% to cover rising security costs, reflecting the operational pressures faced by active short sellers.
Businessinsider
·2026-07-30 17:55:55
410
OpenAI has rehired Lilian Weng to lead internal research.
After leaving Thinking Machines, Lilian Weng returned to OpenAI and will lead the team to advance internal research collaboration and recursive self-improvement.
TechCrunch
·2026-07-30 05:14:32
274