Does Reinforcement Learning Lead to Increasingly Single Answers? New Research from Meta Attempts to Allow Models to Maintain Diversity “According to Target Proportions”
CoinMeta
09-25 10:02
Ai Focus
On September 24th, Meta AI published a paper MaD-RL that addresses a less obvious but very practical issue in large model reinforcement learning: training typically only rewards high scores for a single output, which may cause the model to gradually concentrate its probabilities on one solution or expression. For mathematical problems, converging to a highly successful answer might be advantageous; however, in tasks involving synthetic data, strategy exploration, or those that require satisfying population distribution constraints, overly singular outputs can result in the loss of viable options.
Helpful
No.Help

Meta AI published a paper on September 24th, MaD-RL, focusing on a less obvious but very practical issue in large model reinforcement learning: training usually only rewards high scores for a single output, which may cause the model to gradually concentrate its probabilities on one solution or expression. For math problems, converging to an answer with a high success rate might be beneficial; however, in tasks involving synthetic data, strategy exploration, and those that require satisfying population distribution constraints, overly singular outputs can result in the loss of viable options.

The research team proposed the Distribution Matching framework, which does not merely aim to maximize the average reward, but rather to make the distribution of a certain potential category attribute in the model's output approach a pre-set target. The so-called “target distribution” can refer to the proportions of different problem-solving strategies, code structures, or other discernible categories. The paper demonstrates the effectiveness of this method through mathematical reasoning and programming experiments; however, these are still research findings and do not imply that all general models have automatically solved the issue of pattern collapse, nor does it mean that a single set of proportions can be used to control all open-ended texts.

GRPO may push the probability towards a single mode, and adjusting the temperature does not always solve the problem.

Commonly used verifiable rewards in post-training of large models: points are awarded for correct answers, and no points are given for incorrect ones, followed by using reinforcement learning to increase the probability of high-score outputs. Methods like GRPO compare a set of candidate outputs for the same question and then reinforce the samples that perform relatively better. The mechanism is simple and effective, but when multiple answers are correct, training may favor the first dominant pattern, gradually weakening other viable paths.

The paper describes this phenomenon as a concentration of the output distribution. Increasing the sampling temperature can enhance the randomness at the Token level, and entropy regularization can also prevent the probability from becoming too sharp too quickly, but it does not guarantee that the model will achieve the desired proportion in terms of semantic categories. Although there may be many surface changes in the text, the underlying strategies could still be the same; conversely, forcibly pursuing uniformity may not necessarily align with business objectives. Some tasks require a 70% conservative strategy and a 30% exploratory strategy, rather than an equal amount for each category.

MaD-RL transforms the goal into a distribution distance problem. Researchers first define a potential attribute that can classify the output, then compare the current distribution of the model with the target distribution, and use the differences to construct rewards. The paper discusses different degrees of dispersion such as L2, KL, and Jensen-Shannon; each measure has a different sensitivity to rare categories, zero probabilities, and degrees of deviation. The value of the method lies not in a particular formula itself, but in the expansion of the training goal from "whether this answer is good or not" to "whether many answers combined meet the requirements."

This also brings new dependencies: categories must be reliably identifiable. If a classifier misclassifies different strategies as the same, or if the target attribute itself is biased, then distribution matching will only make the errors more stable. Mathematical and programming tasks can easily benefit from validators and structural rules; however, open dialogue, creative writing, and social attributes are more difficult to assign clear, undisputed labels to.

Distributable control does not equate to content fairness; who sets the goals remains the core issue.

The paper mentions applications such as synthetic data and fairness constraints. When the model generates data for downstream training, if a large number of samples repeat the same template, it is difficult to cover real-world variations no matter how large the data scale is; distribution matching can force the retention of multiple types. However, 'diversity' is merely a statistical attribute and does not automatically guarantee that the information is correct, fully representative, or free from harmful content. Both quality and distribution thresholds must be present simultaneously.

In fairness-related scenarios, the risks are higher. Setting different output proportions for different groups may help to correct obvious imbalances, but it may also simplify complex social issues into rigid quotas. The target proportions should come from specific policies, data, and discussions with stakeholders, rather than being arbitrarily determined by training engineers. Models may also learn to cater to classifiers, seemingly meeting the proportions while in reality continuing to replicate biases.

For developers, MaD-RL is more suitable to be viewed as a training tool rather than an out-of-the-box product feature. Before use, it is necessary to clarify the attributes, target distribution, classification accuracy rate, and failure cost; after training, it is also essential to verify on unseen tasks to confirm that the model has not traded off accuracy for other metrics. The mathematical and programming experiments presented in the paper provide evidence of feasibility, but cross-domain generalization still requires more independent reproductions.

This research also reminds the industry to re-examine the definition of “best answer.” Many evaluations only consider the highest score or average accuracy rate, and are unable to determine whether a model relies on a single approach. When AI is used for solution generation, scientific research hypotheses, and code candidates, what users truly need may be a set of different options that all meet quality standards. Therefore, evaluations should also record success rates, distribution coverage, repetition, and calibration errors simultaneously.

After reinforcement learning, training is moving from single-reward approaches to multi-objective constraints. Correctness, safety, cost, and diversity may pull in opposite directions; extreme optimization of any one metric can harm other aspects. The contribution made by MaD-RL is to explicitly incorporate semantic distributions into the optimization objectives and to compare the effects of different distance metrics. While it does not solve the governance issue of 'who defines the ideal distribution', it transforms this problem from a complaint after training into a parameter that must be clearly defined before training begins.

The most cautious judgment at this stage is that the Meta research team has demonstrated that in several verifiable tasks, reinforcement learning can directly match the target output distribution, and they have pointed out that common GRPO training methods may compress diversity. Whether this approach can work stably in larger models, more complex open tasks, and with long-term training still requires further experimentation; however, it has already provided a testable technical path for the idea that "models should not only get the answers right but also retain a reasonable range of choices."

Tip
$0
Like
0
Save
0
Views 139
HQYC reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
XRP Ledger A key indicator surges by 535% within two days, what is the impact on prices?
U.Today reports that online activities on XRP Ledger have significantly heated up, with the number of newly created accounts increasing by 535.2% to 2,900 within two days, the number of active accounts rising by 218.1% to 15,300, and the number of successful transactions reaching 1.7 million. The volume of payments and the value of transfers also increased substantially. However, the number of active users decreased by 70.6% to 140,000, indicating that the growth may have been concentrated among a smaller but more active user group.
U.Today
·2026-09-28 05:51:21
2
Investigation on Sunday: Who Should Be Responsible When an Autonomous Vehicle Has an Accident?
Electrek launched a poll around the question of "Who should be responsible when a taxi driven by AI gets out of control?" The results showed that nearly 90% of respondents attributed the responsibility to the developers of the autonomous driving system, while only about 3.5% believed that the person sitting in the driver's seat should bear the responsibility.
The Block
·2026-09-28 05:00:28
13
Anthropic CEO to Have Dinner with President Trump
Anthropic CEO Dalio Amodei has frequently become the focus this weekend. Axios was the first to report that he would have dinner with U.S. President Donald Trump at the White House on Sunday night, and TechCrunch subsequently confirmed this arrangement by citing a source familiar with his itinerary. This will be their first face-to-face meeting, as both parties have recently held differing views on AI security issues.
TechCrunch
·2026-09-28 04:51:25
13
After mortgage rates rise above 7%, considering buying stocks instead of a house? The S&P 500 has outperformed the housing market by far over the past decade
From an investment perspective, in recent years, the returns on housing in the United States have significantly lagged behind those of the stock market, and mortgage rates rising above 7% may further widen this gap. The article cites economists who argue that buying a home and investing are two different things; at the same time, many regions in the U.S. still have a buyer's market, with sellers making more concessions and offering more discounts.
Fortune
·2026-09-28 04:31:23
11
Traders expect XRP to be most likely to rise to $1.70 this month
As XRP continues its strong momentum in the strongest quarter of the year, the latest market data predicted by Kalshi indicates that platform traders believe XRP is most likely to reach $1.70 this month. According to the report, at the price of $1.53 at the time of writing, XRP needs to rise by more than 11% in the next three days to achieve this target.
U.Today
·2026-09-28 04:22:13
15
View More