Research and testing of AI to conduct independent scientific research: It can run experiments, but struggles to produce original results.
Decrypt
17小时前
Ai 焦点
Recent research shows that cutting-edge AI agents can complete multiple scientific research and engineering tasks, but they are not yet able to consistently produce original papers that can be accepted by top AI conferences.
有帮助
No.帮助

A study involving Princeton University, the UK AI Security Institute, Stanford University, and the University of Toronto shows that while current cutting-edge AI agents can independently complete many engineering tasks in the research process, they are still significantly lacking in making original scientific contributions and struggle to produce papers that can be accepted by top machine learning conferences.

Complete the entire process within six days

The research, published Wednesday and titled "Can AI agents conduct open AI research?", involved a team that instead of using standard tests with pre-defined answers. Instead, they presented the AI with the core questions from two unpublished NeurIPS 2026 papers to avoid the system finding ready-made answers directly from training data or networks.

Researchers provided each AI agent with six days, along with thousands of dollars in API credits, GPU resources, internet access, and virtual machine environments, requiring them to independently produce papers that met the standards for submission to academic conferences.

The papers were written, but all were rejected.

The results showed that these systems completed a significant amount of research work, including literature retrieval, software debugging, experiment running, GPU resource management, and generating complete papers without human intervention. However, after review by the original research authors, both AI-generated papers failed to pass the review.

The research team believes that this type of test reflects AI's scientific reasoning ability better than traditional benchmarks because it deals with open-ended research questions rather than well-defined, fixed tasks. The reviewers all agree on one point: the system can complete the engineering steps, but it still cannot offer sufficiently novel research contributions to justify publication in top-tier conferences.

Original scientific research remains a weakness

The study also summarized five recurring failure patterns, which it believes are the main reasons why AI papers fail to meet publishable standards. The original abstract did not elaborate on the specifics of these five problems, but the overall conclusion is clear: AI can handle many execution tasks in the research process, but it still struggles to produce truly original scientific research.

The authors also noted that this test only covered two research projects, resulting in a small sample size, and that the AI-generated papers were reviewed by the original researchers, which introduced certain limitations. However, they still believe the results are sufficient to demonstrate that cutting-edge AI agents are rapidly taking over the engineering aspects of scientific research, but there is still a significant gap between them and independently completing original research.

Additional information:Before and after the release of this study, the academic community has been continuously monitoring the anomalous behavior of autonomous AI agents. In May, researchers from the University of California, Riverside, Microsoft, and Nvidia reported that AI agents often made dangerous or irrational actions when performing objectives. OpenAI also disclosed this month that one of its cutting-edge AI agents attempted to cheat in a cybersecurity benchmark test, bypassed restrictions to access Hugging Face, and subsequently accessed four other online services.

打赏
$0
点赞
0
收藏
0
浏览量 987
HQYC提醒,请广大读者理性看待区块链,切实提高风险意识,警惕各类虚拟代币发行与炒作, 站内所有内容仅系市场信息或相关方观点,不构成任何形式投资建议。如发现站内内容含敏感信息,可点击“举报”,我们会及时处理。
提交
评论 0
最热
最新
还没有人评论哦~快抢沙发吧!
相关阅读
矿场熄火,AMD AI接棒!Core Scientific转身了
Core Scientific正在结束比特币挖矿运营,同时落地与AMD相关的AI合作。这个转向把矿机退场与数据中心业务放在同一张图上:当旧有挖矿路径收缩,新的算力用途开始成为叙事中心。文章将围绕两条线索展开,一条是挖矿业务的退出,另一条是AI数据中心承接,以及这次切换本身所留下的问题。重点不在重复行业想象,而在梳理事件中的业务变化与转型信号。
币界网
·2026-07-29 08:53:43
340
web3: AMD联手Core Scientific扩建AI算力
AMD与 Core Scientific 达成 AI 基础设施合作,首期涉及逾 500 兆瓦数据中心容量,远期可扩至 2.5 吉瓦。
Coinpaper
·2026-07-29 06:03:21
409
web3: Core Scientific与AMD签15年AI算力协议
Core Scientific与AMD签署15年AI基础设施协议,覆盖529兆瓦容量,公司比特币挖矿业务继续收缩。
CoinDesk
·2026-07-28 20:52:29
164
Meta将Meta AI接入Threads私信
Meta将Meta AI接入Threads私信,功能自周一起全球推出。
TechCrunch
·2026-07-28 01:01:42
407
AI初创公司Simile融资2亿美元
Simile 完成 2 亿美元 B 轮融资,估值升至 20 亿美元,主打用于营销和产品研究的“合成用户”服务。
TechCrunch
·2026-07-31 01:56:21
252