Give students ChatGPT and their assignment scores rise. Teach them a causal framework and their ideas spread out. In a new classroom experiment, those were different gains produced by different interventions.
A preregistered ChatGPT critical-thinking study at Bocconi University assigned 1,053 first-year students to one of four conditions: causal-reasoning training, GPT-4o access, both, or neither. Students then completed a 45-minute university-merchandising recommendation in no more than 180 words.
ChatGPT access improved the main rubric score. Causal training changed the range of ideas students produced. The experiment does not show that one short intervention created durable learning, broad transfer, or better independent reasoning outside the task.
The trial separated tool access from thinking training
The 2-by-2 design is the study’s strongest feature. Instead of comparing “AI” with “no AI” and guessing why outcomes changed, the researchers independently varied access to GPT-4o and a short causal-reasoning lesson.
| Group | GPT-4o access | Causal training | What the comparison isolates |
|---|---|---|---|
| Control | No | No | Baseline assignment performance |
| Tool only | Yes | No | Effect of GPT access |
| Training only | No | Yes | Effect of causal framework |
| Combined | Yes | Yes | Whether the interventions reinforce each other |
The students came from 13 classes. That is a meaningful classroom sample, but it remains one university, one course context, and one short business recommendation.
GPT access added 0.862 points to the main score
With the study’s controls, GPT access increased the main 1-to-5 evaluation score by 0.862 points. The estimated control score was 2.09. Relative to that baseline, the reported increase is about 41%, calculated as 0.862 divided by 2.09.
That percentage is a way to understand the scale inside this assignment, not a general productivity claim. The score measured the submitted recommendation under the study rubric. It did not measure retention a month later, performance without ChatGPT, or transfer to another subject.
The practical reading is straightforward: students with access to a strong language model produced work that graders rated higher on this documented task.
Causal training produced the diversity gain
The causal-reasoning lesson increased diversity within each student’s solution by roughly half a standard deviation. It was also the only intervention with a clear positive effect on idea diversity across students.
That finding deserves attention because AI-assisted assignments can converge on fluent, familiar answers. The model may improve structure and completeness while nudging many students toward the same obvious recommendation.
A thinking framework can force students to consider mechanisms, alternatives, and downstream effects before they ask for prose. The tool then helps express an analysis that already has more branches.
The study used both human and model-based grading
Trained human graders scored the main business-quality outcome. The causal-reasoning measures were scored by GPT-5 across 31,590 API evaluations.
That split should remain visible. Human grading supports the headline assignment-score result. Model-based grading made the larger causal-analysis evaluation practical, but it can inherit the judge model’s preferences and blind spots.
- Publish the scoring rubric and prompts.
- Check agreement on a human-scored subset.
- Test whether another judge model changes the ranking.
- Blind graders to treatment where possible.
- Keep the main human outcome separate from automated diagnostic scores.
One polished answer is not a learning measure
The assignment lasted 45 minutes and allowed 180 words. That is suitable for testing immediate work quality under controlled conditions. It is too narrow to answer whether students learned a reusable reasoning method.
A stronger follow-up would remove GPT access and ask students to solve a new problem days or weeks later. It would measure whether they can explain the causal chain, identify an omitted variable, revise after counterevidence, and transfer the method to a different domain.
The study also does not tell us whether weaker and stronger students benefited equally, whether students followed similar prompting strategies, or how much of the score increase came from editing rather than new analysis. Those questions need process traces, subgroup analysis, and a delayed assessment. They should not be filled in with assumptions simply because the average effect was large.
Our guide to AI use in college and student judgment makes the same distinction: the finished answer can improve while the student’s independent judgment remains unknown.
A classroom design that uses both gains
- Give students the reasoning framework before opening the model.
- Require an initial causal map or hypothesis written without ChatGPT.
- Use the model to challenge assumptions, find alternatives, or improve expression.
- Ask students to mark which ideas came from the model and which they rejected.
- Finish with a short independent explanation or transfer question.
This design treats ChatGPT as an amplifier rather than the source of the whole analysis. It also gives the teacher evidence about process, not just a final paragraph.
The existing OpenAI education plugins can help produce materials and practice. The Bocconi result suggests those tools should be paired with explicit reasoning instruction if schools care about variety and independent thought.
My verdict: better work and better thinking are separate targets
The trial gives educators a more useful answer than “ChatGPT helps” or “ChatGPT harms learning.” On this task, GPT access improved the graded product. Causal training broadened the ideas.
Schools should design for both outcomes and measure them separately. Let the model help students produce stronger work, but include a reasoning step the model cannot silently complete and a later task that tests what the student can carry forward alone.
Read the study
- Read the preregistered classroom-study paper.
- Review Bocconi University’s summary.
- Compare OpenAI’s account of the experiment.
Checked August 29, 2026. The 41% comparison is Musthave.ai’s calculation from the reported 0.862-point effect and estimated 2.09 control score. It applies to this study’s rubric and task, not general learning.