The OECD came out with numbers that should make any operations leader sit up. In its review of artificial intelligence in education, it found that students who leaned on generative AI scored lower on tests of critical thinking. The tool built to make them faster had made them weaker. Then came the detail that matters most for business. When those same students received training in how to critically assess the machine's output, the penalty vanished. They performed no worse than peers who had never touched AI at all. The tool had not softened their judgment. Their relationship to the tool had.
This is a small lab, but it is a useful lab. Schools exist to build judgment. That is partly what a university does, and it is exactly why the OECD study is a warning light rather than a curiosity. If handing AI the thinking can dull the critical assessment of people who spend their days exercising that muscle, the same thing is happening in the office, quieter and slower, in the rooms that train analysts and managers and decision-makers.
What actually happens to the judgment
When a model hands you a finished answer, the natural reaction is relief. You do not have to think about it. You read it, you nod, you send it along. Cognitive scientists call this offloading, and there is nothing wrong with offloading heavy lifting. But the brain is not wired to use a tool without it changing how the brain works. Every time you accept a machine's conclusion without testing it, you skip the exact exercise that keeps your own evaluation sharp. The muscle is not used, so it weakens.
The danger is subtle because the model is usually right. Most of the time the output is good enough that nobody notices the gap. You save the hour, you meet the deadline, and the person next to you is doing the same. Nobody has to own the fact that none of you spent the effort to check.
The thing that reversed the trend
The OECD data did not simply show a decline. It showed a decline that could be cured, and the cure was specific. The students were not trained to prompt better or read faster. They were trained to treat the machine's answer as a claim, not a verdict. To figure out whether the reasoning held, to look for what the model had missed, and to override it when it did. In other words, they were taught to judge the tool before trusting it.
The result is the important part. After the training, those students scored no worse than the people who never used AI. The intervention did not merely prevent harm. It restored judgment to a working standard. The people were as capable as the baseline of humans who had done the work with nothing but their own head. That tells you the skill was never gone, only dormant.
Why this matters for the office
Most companies have not thought hard about what judgment-heavy work actually means once it can be assisted. A strategy analyst reviewing a competitor's filing. A buyer weighing a supplier quote. A marketer reviewing copy. An engineer reading a proposed fix. In each case the human's value is the ability to assess what the machine produces, to spot the half-truth, to ask the question the tool did not raise. That assessment is a skill, and skills decay when they are not practiced.
If a company wires AI into those jobs so that the model produces the first draft and the human simply approves it, the company is asking every employee to quietly give up the one thing that justifies the salary. The fade will be gradual, so it will not be caught. Then comes the moment the model is wrong in a way that costs money, and the human who has been outsourcing their judgment has nothing to fall back on.
What to build instead
The fix is not to pull AI out of judgment-heavy roles. It is to redesign the role so the assessment stays human, and to teach the assessment as a craft. That means keeping the person responsible for evaluating the output, not just requesting it. It means training employees to check the machine's reasoning, to test its assumptions, and to find what it left out, as a regular skill rather than an afterthought. It means making verification a required step in the workflow, not a personal preference. And it means measuring how well employees catch the machine's errors, because that is the capability you are trying to preserve.
Some of this is already the discipline of senior work. A good manager does not simply sign what a report produces. A good editor does not print the first line a copywriter submits. The insight is that AI now makes those jobs easier to do badly, so the training must be deliberate and explicit, or the whole floor will drift toward approval without thought.
So what happens to employee judgment, really
It depends entirely on how the company designs the work. If you hand AI the thinking and let the employee merely collect the output, the judgment will fade, quietly and permanently, because the brain only sharpens what it is asked to use. That is the OECD's first lesson. If you instead build the workflow so the human stays the person who assesses, verifies, and decides, the fade stops. The OECD's second lesson, and the one that matters more, is that this is not just a hope. Training people to judge the machine's work restores their own judgment to a full working standard. The tool does not have to take the thinking. The outcome depends on which choice the company makes, and most have not made it yet.