Random Experiment with a Thousand Students: ChatGPT Makes Answers More Professional, Critical Thinking Training Makes Ideas More Unique
CoinMeta
1h ago
Ai Focus
On August 27th, OpenAI and researchers from Bocconi University in Italy announced a randomized experiment aimed at answering one of the most difficult questions in the education field: how do students' assignment quality, reasoning abilities, and originality change after they are given generative AI. Over 1,000 first-year undergraduate students from Bocconi University participated in the experiment, where they were required to come up with real marketing suggestions for the university's souvenir shop. The researchers randomly divided them into four groups based on class time periods: those who used GPT-4o, those who received causal reasoning training, those who used both, and those who did not use either.
Helpful
No.Help

On August 27th, researchers from Bocconi University in Italy announced a randomized experiment aimed at answering one of the most difficult questions in the field of education: how do students' assignment quality, reasoning abilities, and creativity change after they are exposed to generative AI. Over 1,000 first-year undergraduate students from Bocconi University participated in the experiment, where they were required to come up with real marketing suggestions for the university's souvenir shop. The researchers randomly divided them into four groups based on their class periods: those who used GPT-4o, those who received causal reasoning training, those who used both, and those who did neither.

The experimental results are not simply “AI good” or “AI bad”. Students who received ChatGPT improved by nearly 1 point on a five-point manual grading scale on average; their answers contained more ideas, were more logically clear, and were also closer to expert recommendations. Students who underwent causal reasoning training did not show significant improvement on traditional grading scales, but they proposed broader and more diverse ideas, and were better able to explain why the solutions were effective and under what conditions they might fail. Students who received both AI and reasoning training retained the advantages of both high-quality answers and diversity of ideas.

This study provides causal evidence from a specific classroom task, and it is not proof that “AI has comprehensively improved students’ abilities.” The sample consists of first-year students from the same business school, and the task involved a marketing case study using the GPT-4o model; whether these conclusions can be generalized to mathematics proofs, history writing, programming, medicine, or elementary and secondary school classrooms requires new experiments. The study also involved the participation of OpenAI Economic Research; readers should also review the paper’s methodology, data, and potential conflicts of interest.

AI has improved the quality of finished products, but it has also exposed the blind spots of traditional scoring systems.

The content submitted by students is evaluated by trained human graders using a five-point scale, which focuses on whether the suggestions effectively contribute to enhancing the store's visibility and usage rate. The ChatGPT group saw an improvement of nearly 1 point, indicating that the model can help beginners quickly acquire a structured approach, proper expression, and common professional frameworks. Students are not merely required to copy and present what is given; they still need to decide how to ask questions, filter suggestions, and combine them into a final answer. However, the AI significantly lowered the barrier for generating something that “looks like work by an expert.”

The problem is that traditional scoring systems tend to reward clear, complete, and conventional answers more, but not necessarily uniqueness. The causal reasoning training group was able to come up with solutions that were even less similar to those of their peers, yet these differences were not fully reflected in the overall scores. It was only through automated text analysis that researchers were able to observe the various changes between the number of ideas, the degree of difference, causal explanations, and the similarity to expert texts.

If schools continue to focus only on the final product, it will become increasingly difficult to distinguish what students have truly understood in the AI era. A well-expressed solution may stem from the student's in-depth reasoning, or it may mainly come from the structure provided by a model. On the contrary, an answer that is not perfect but proposes a new approach may contain a higher level of independent judgment. Evaluation of design needs to expand beyond just whether the final product meets standards to include issue definition, evidence selection, hypothesis testing, process documentation, and oral defense.

Causal reasoning training itself is not aimed at AI. Students learn to connect causes and effects through games, cases, questions, and feedback, and think about why a plan may succeed or fail. It does not increase traditional test scores, but it does enhance reasoning skills and the diversity of ideas. This reminds schools that critical thinking is not just an abstract slogan; it can be broken down into specific actions that can be practiced, and it also requires specialized evaluation indicators to be observed.

The most reasonable classroom strategy is not to ban or indulge, but to combine the two abilities.

In the experiment, students who received both ChatGPT and causal reasoning training had scores and numbers of ideas that were similar to those of the group that only used AI, while their idea diversity was closer to that of the group that only received reasoning training; they also demonstrated stronger logical coherence and were more willing to seek explanations and question hypotheses. This indicates that there is no zero-sum relationship between the AI tool and cognitive training. The model can help students overcome barriers in knowledge and expression, while the training helps them avoid all converging to the same set of perfect answers.

Classroom implementation can be divided into several stages based on this approach. Students first independently write down questions, hypotheses, and preliminary plans, then use AI to expand or refute their views; subsequently, they indicate which parts come from the model and explain the reasons for adopting or rejecting them. Finally, their understanding is verified through oral discussions or by applying the knowledge to new situations. Evaluation takes into account the quality of the final product, originality, the causal chain, the reliability of evidence, and the reflection process. In this way, AI improves efficiency, but it cannot replace human judgment.

Teachers also need to control the boundaries of tasks and data. Research tasks do not involve highly sensitive personal information, but real classrooms may contain student records, health, or family-related materials. Schools need to use managed accounts, limit the data that can be uploaded, and clarify who will review model errors. For minors, it is also necessary to handle parental consent, age requirements, and exit mechanisms.

Replicating studies is also very important. Subsequent experiments should involve changing schools, disciplines, models, and the difficulty of tasks to compare whether short-term high scores can translate into independent performance a few weeks later. It is also necessary to record the actual hints given to students, the process of their modifications, and factual errors. If students perform better only when the AI is available, but do not develop the ability to transfer this knowledge to other situations, then the educational value and the dependence on the tool need to be evaluated separately. On the contrary, if structured use helps students gradually learn to ask questions and examine causal chains, it indicates that the classroom design has produced sustained learning effects.

The most important conclusion of this experiment is not that ChatGPT made the students “smarter,” but rather that different teaching interventions improved various aspects. AI led to more professional and complete answers; causal reasoning training broadened their thinking and enabled them to better explain boundaries. If schools only reward polished and answer, they will overlook creativity; if they only prohibit AI, they will also forego tools that help beginners approach professional frameworks. What truly needs to be reworked are homework and grading systems, so that both tool capabilities and independent reasoning skills become observable and verifiable learning outcomes.

Tip
$0
Like
0
Save
0
Views 19
HQYC reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
New York's second home tax comes into effect, wealthy homeowners seek ways to avoid it
New York's high-value second-home tax enters implementation phase; wealthy homeowners seek exemption routes, but lawyers say there is little room for evasion under the regulations.
Businessinsider
·2026-08-28 08:26:59
23
web3 : Android Starting from version 17, accessing website domain names will be hidden by default.
Google will integrate ECH within Android 17, and ECH GREASE will be enabled by default to enhance domain name privacy protection during website visits.
The Cryptonomist
·2026-08-28 07:15:09
36
web3: Foreign media: AMM exempting arbitrage from fees can increase pool earnings
Research indicates that if AMM eliminates fees for internal arbitrage and uses hooks to rebalance within the same transaction, it may reduce MEV losses and increase the retained value of the pool.
The Cryptonomist
·2026-08-28 07:15:06
36
web3: Sparrow Wallet releases security update, with multiple fixes coming from AI review
Sparrow Wallet releases version 2.5.4, developers say most of the fixes come from AI auxiliary code review; updates involve transaction verification, Tor privacy, and hardware wallet security.
Coinpaper
·2026-08-28 05:52:01
43
View More