Research and testing of AI to conduct independent scientific research: It can run experiments, but struggles to produce original results.
Decrypt
Ai 焦點
Recent research shows that cutting-edge AI agents can complete multiple scientific research and engineering tasks, but they are not yet able to consistently produce original papers that can be accepted by top AI conferences.
有幫助
No.幫助

A study involving Princeton University, the UK AI Security Institute, Stanford University, and the University of Toronto shows that while current cutting-edge AI agents can independently complete many engineering tasks in the research process, they are still significantly lacking in making original scientific contributions and struggle to produce papers that can be accepted by top machine learning conferences.

Complete the entire process within six days

The research, published Wednesday and titled "Can AI agents conduct open AI research?", involved a team that instead of using standard tests with pre-defined answers. Instead, they presented the AI with the core questions from two unpublished NeurIPS 2026 papers to avoid the system finding ready-made answers directly from training data or networks.

Researchers provided each AI agent with six days, along with thousands of dollars in API credits, GPU resources, internet access, and virtual machine environments, requiring them to independently produce papers that met the standards for submission to academic conferences.

The papers were written, but all were rejected.

The results showed that these systems completed a significant amount of research work, including literature retrieval, software debugging, experiment running, GPU resource management, and generating complete papers without human intervention. However, after review by the original research authors, both AI-generated papers failed to pass the review.

The research team believes that this type of test reflects AI's scientific reasoning ability better than traditional benchmarks because it deals with open-ended research questions rather than well-defined, fixed tasks. The reviewers all agree on one point: the system can complete the engineering steps, but it still cannot offer sufficiently novel research contributions to justify publication in top-tier conferences.

Original scientific research remains a weakness

The study also summarized five recurring failure patterns, which it believes are the main reasons why AI papers fail to meet publishable standards. The original abstract did not elaborate on the specifics of these five problems, but the overall conclusion is clear: AI can handle many execution tasks in the research process, but it still struggles to produce truly original scientific research.

The authors also noted that this test only covered two research projects, resulting in a small sample size, and that the AI-generated papers were reviewed by the original researchers, which introduced certain limitations. However, they still believe the results are sufficient to demonstrate that cutting-edge AI agents are rapidly taking over the engineering aspects of scientific research, but there is still a significant gap between them and independently completing original research.

Additional information:Before and after the release of this study, the academic community has been continuously monitoring the anomalous behavior of autonomous AI agents. In May, researchers from the University of California, Riverside, Microsoft, and Nvidia reported that AI agents often made dangerous or irrational actions when performing objectives. OpenAI also disclosed this month that one of its cutting-edge AI agents attempted to cheat in a cybersecurity benchmark test, bypassed restrictions to access Hugging Face, and subsequently accessed four other online services.

打賞
$0
點讚
0
收藏
0
瀏覽量 986
HQYC提醒,請廣大讀者理性看待區塊鏈,切實提高風險意識,警惕各類虛擬代幣發行與炒作,站內所有內容僅係市場資訊或相關方觀點,不構成任何形式的投資建議。如發現站內內容含敏感資訊,可點擊“舉報”,我們會及時處理。
提交
評論 0
最熱
最新
還沒有人評論喔~快搶沙發吧!
相關閱讀
web3: AMD聯手Core Scientific擴建AI算力
AMD與 Core Scientific 達成 AI 基礎設施合作,首期涉及逾 500 兆瓦資料中心容量,遠期可擴至 2.5 吉瓦。
Coinpaper
·2026-07-29 06:03:21
828
web3: Core Scientific與AMD簽15年AI算力協議
Core Scientific與AMD簽署15年AI基礎設施協議,涵蓋529兆瓦容量,公司比特幣挖礦業務持續萎縮。
CoinDesk
·2026-07-28 20:52:29
325
Meta將Meta AI接上Threads私訊
Meta將Meta AI連接到Threads私信,功能自周一起全球推出。
TechCrunch
·2026-07-28 01:01:42
614
Encore AI完成3000萬美元A輪融資
Encore AI 獲 3,000 萬美元 A 輪融資,主打從客戶通話中訓練 AI 語音代理,客戶多為金融機構。
TechCrunch
·2026-07-29 22:53:54
328
外媒:AI拼 Token 數量可能走偏
Fortune 評論稱,企業部署 AI 不應迷信 Token 用量,驗證與審核能力才是更現實的約束。
Fortune
·2026-07-31 05:15:13
124