Anthropic publishes a new paper showcasing the latest progress in model alignment training using the AI system. The paper indicates that an automated research system has improved model performance across 10 alignment benchmarks targeting mismatch behaviors, without compromising overall performance.
Automatic systems can complete research iterations.
This paper is titled “Automated Researchers Can Reliably Mitigate Alignment Failures”. The research was led by researchers Anthropic and Chen Yueh - Han. The system introduced in the paper first retrieves existing literature, then proposes a training method, and uses this method to train the model for about 30 minutes.
Subsequently, the system will continue to raise the benchmark requirements based on the results and repeat multiple rounds of testing. Effective methods will be retained, while ineffective ones will be eliminated. According to the description in the paper, this process is similar to the traditional research approach of "searching for information, proposing solutions, conducting experiments, and screening results," but it is carried out at a faster pace and on a larger scale.
All 10 benchmarks have been improved.
The paper states that among the 10 alignment benchmarks provided, the automated system achieved improvements in each one without causing a decline in overall capability. This means that the system is not only capable of correcting specific misalignment behaviors but can also, to some extent, avoid the problem of 'fixing one area and damaging another.'
Anthropic also compared this system with that of human researchers. The paper states that the best automated alignment research method was able to surpass the solutions proposed by senior researchers in an average of 6 hours; however, research directions led by humans did not yield stronger results.
Lower costs, but still subject to benchmark restrictions
The paper also provides a cost comparison: the inference cost of the automated alignment research system’s API is approximately $4 per hour, whereas the cost paid to human researchers using Anthropic is about $150 per hour.
However, the paper also mentions that this method has obvious limitations. For an automated system to be effective, it is prerequisite that these alignment benchmarks themselves can accurately reflect the real objectives. If the benchmarks are not properly designed, the results optimized by the system may deviate from actual needs.
In addition, such systems also rely on continuously maintained research literature and evaluation frameworks. In other words, although automated researchers can speed up the experimental and screening processes, setting benchmarks, updating data, and calibrating targets still require a significant amount of work.











