Whether teaching a cognitive skill remains valuable in this era of AI is a crucial question in education and for firms. This column reports on an experiment in which around 1,000 undergraduates in Italy received either (1) causal-reasoning training, (2) access to ChatGPT, (3) both, or (4) neither. While students with ChatGPT scored better in an assignment to write a 180-word recommendation on a real-world problem, only those who had received causal-reasoning training produced unique ideas. The findings highlight the benefits of combining training in first-principles thinking with large language models.
Our ability to solve problems is changing with the evolution of large language models (LLMs), but we do not know how LLMs combine with cognitive skills. Garicano and Rayo (2026) show that this question is urgent. Their theory predicts that junior employees with a knowledge background above a certain threshold complement AI, benefit from early-career training, and their careers progress as they did in the pre-AI world. In contrast, juniors with knowledge below the threshold are replaced by AI and their career progression falls behind. Therefore, according to their theory, the knowledge background of juniors has long-term implications.
Recent evidence confirms the urgency of the question. Surveys put the share of undergraduates who use generative AI in their studies at 80% or more across rich countries (The Economist 2026). Evidence about the implications cuts both ways. In the workplace, giving people ChatGPT makes them faster and raises the quality of what they produce, with the largest gains for the least experienced (Noy and Zhang 2023, Brynjolfsson et al. 2025). In education, the same tools can hollow out learning. A study of 27,000 Chinese secondary school pupils finds that those who used AI for homework raised their homework scores and cut the time they spent on it, but scored 20% lower in exams (Strömberg et al. 2026). Studies in universities show the same tension. After ChatGPT arrived, grades rose in AI-compatible courses at an Israeli university, especially for weaker students, and the grade distribution compressed, so that grades now tell employers less about skills than before (Hausman et al. 2025).
What this evidence cannot tell us is whether teaching a way of thinking still pays once the tool is available. Existing studies vary access to AI, but they measure the human skill rather than varying it. To know whether a cognitive skill is worth teaching in the age of LLMs, one has to teach it to some and not to others, give the tool to some and not to others, and look at all four combinations. This is what we did (Asirvatham et al. 2026).
We ran a pre-registered randomised controlled trial with 1,053 first-year undergraduates at Bocconi University in Milan, all enrolled in the same introductory management course across 13 classes. Classes were randomly assigned to one of four conditions: training in causal reasoning (‘first-principles thinking’), access to ChatGPT, both, or neither. We chose causal reasoning because it is a fundamental and teachable form of thinking, and because earlier randomised trials show that it improves managerial decisions (Camuffo et al. 2020).
The training took the form of a game of 12 questions with feedback, which taught students to build explicit chains of cause and effect, to state the conditions under which their claims would fail, and to explain the mechanisms through which a cause produces its effect. Students in the other conditions played a placebo game with the same questions but no feedback. Those with ChatGPT access used it through the university’s institutional subscription, in a separate browser tab.
All students had 45 minutes to write a 180-word recommendation on a real problem: how to increase alumni awareness and usage of the university’s merchandising shop, a standard marketing brief (Bendle et al. 2016).
We measured eight outcomes in three domains.
Figure 1 summarises the results as differences from the control group. Overall, students who received both treatments did better than the control group on every outcome, never did worse than students with a single treatment, and on three outcomes (coherent logic, mechanisms, and the number of ideas) the two treatments reinforced each other.
Figure 1 Effects of causal-reasoning training and ChatGPT access on eight outcomes, differences from the control group
The two treatments work on different outcomes. GPT is what raised performance: students with ChatGPT wrote recommendations that the evaluators rated higher and that came closer to the experts’ solutions. Causal reasoning training did not improve performance, and adding it to GPT did not raise performance further. Both treatments made the texts more coherent and richer in ideas, GPT more so, and the two together more than either alone.
Deeper thinking, however, is another matter. Only the students who underwent training in first-principles thinking stated the conditions under which their recommendations would fail, explained the mechanisms behind them, and produced ideas that differed from those of their classmates. GPT alone did none of this, and adding GPT to the training took nothing away. In short, GPT expands the volume of ideas and raises the evaluations; trained reasoning changes the character of the ideas.
Our results highlight the benefits of combining training in first-principles thinking with LLMs. It may seem easier to ban LLMs, or to replace first-principles thinking with them. However, according to our study, nurturing first-principles thinking in the AI world can raise the background of juniors above the Garicano-Rayo bar. The message is for universities, the education system more generally, and firms.
Additionally, GPT delivers coherent, well-evaluated answers that resemble what experts would write. But counterfactual thinking, attention to mechanisms, and the diversity of ideas – which is where originality lives – remain the prerogative of trained human reasoning. LLMs do not erode these capacities, but they do not supply them either. If anything, AI agents may need special training to provide them.
Finally, evaluations are often themselves the measure of performance. Students’ careers depend on evaluations by professors, employees’ salaries and careers depend on evaluations by their managers, and even entrepreneurs, and the value of firms, depend on evaluations by investors. To make large-scale evaluations feasible, and perceived as fair, it is natural to adopt standard criteria. Similarly, investors may not know how to evaluate novel ideas and therefore prize standard solutions.
Our study shows that the unintended consequence of standardised evaluations is to discourage diversity of ideas and first-principles thinking. While not all the ideas produced this way are good ideas, most good ideas come from these processes. In a world with AI, first-principles thinking must be nurtured and evaluations have to follow suit. Otherwise, they may provide younger generations with incentives to stay below, rather than above, the Garicano-Rayo bar.
Source : VOXeu
Taiwan has not passed on most of the rising costs to consumers, and it heavily…
Only 13% of companies were on track with their AI initiatives, as regulatory hurdles and…
The European Securities and Markets Authority said on Wednesday that European regulators should be given…
Oil prices fell 1% on Thursday as recovering crude exports from the Gulf and a…
China's trade surplus is widening, its manufactured exports are surging, and the political backlash is…
Fed's Williams says there's time to parse data before raising rates again Federal Reserve Bank…