Perhaps the safest way to use AI is to learn how to disagree with it

AI can improve visible performance faster than the judgement supporting it. Emerging research suggests the effect depends partly on how much thinking people continue to do for themselves.

Share this article:

Boardroom Coach Part

Psychology and workplace research is beginning to raise a more difficult question about AI: whether visible performance can improve faster than the judgement supporting it.

That distinction matters because AI can make work look better very quickly. The harder question is whether the person producing it is still exercising enough independent judgement to challenge, reconstruct and defend the conclusion.

What’s Happening

Research is beginning to distinguish between AI that assists thinking and AI that progressively removes the need to think through the task.

A 2025 Microsoft Research study surveyed 319 knowledge workers and collected 936 first-hand examples of people using generative AI in their work. Higher confidence in GenAI was associated with less reported critical thinking, while greater confidence in people's own ability to perform the task was associated with more critical engagement. The researchers also found that critical thinking changed shape around AI, shifting towards information verification, response integration and maintaining oversight of the task. The study is observational, so the findings show associations rather than causation. (Lee et al., 2025)

Experimental evidence raises a related question about skill formation. Anthropic ran a randomised controlled study involving 52 mostly junior software engineers who were learning to use an unfamiliar Python library. Participants using AI assistance scored an average of 50% on a subsequent knowledge assessment, compared with 67% for those who worked without AI. The AI-assisted group completed the original task around two minutes faster on average, but the productivity difference was not statistically significant. (Shen and Tamkin, 2026)

The pattern within the Anthropic study is particularly relevant. Participants who largely delegated coding or debugging to AI tended to perform poorly on the subsequent assessment. Those who used AI to ask conceptual questions, request explanations or test their understanding tended to perform better. The researchers are careful not to claim that these interaction styles caused the differences in learning, but the pattern points towards a distinction between using AI to complete work and using it while remaining cognitively engaged. (Shen and Tamkin, 2026)

A 2026 paper in Trends in Cognitive Sciences places this in the wider history of cognitive offloading. The authors conclude that handing cognitive activity to AI can impede skill acquisition and contribute to skill decay, while emphasising that the risk depends considerably on what is being offloaded and how the technology is used. (Cash et al., 2026)

The American Psychological Association's July 2026 review reaches a similarly measured conclusion. It notes evidence that heavy reliance on generative AI can reduce critical-thinking activity and job-specific skill development, while structured use can preserve more human engagement. It also stresses that many questions about longer-term cognitive consequences remain unresolved. (Abrams, 2026)

The evidence does not establish that challenging AI preserves leadership judgement.

It supports a narrower and more useful proposition.

The cognitive consequences of AI appear to depend partly on how much verification, evaluation, independent problem-solving and human oversight remain in the work.

"Psychology and workplace research is beginning to raise a more difficult question about AI: whether visible performance can improve faster than the judgement supporting it.

That distinction matters because AI can make work look better very quickly. The harder question is whether the person producing it is still exercising enough independent judgement to challenge, reconstruct and defend the conclusion."

What’s Being Said

Most organisational discussion about AI still concentrates on adoption and productivity.

How quickly are people using it? How much time is being saved? Which workflows can be automated? How good are the prompts becoming? How much more output can people produce?

Those questions are commercially understandable.

They are also easier to measure than what may be happening underneath the output.

A document can improve while the person producing it practises less of the reasoning required to create one independently.

A junior employee can produce work that looks considerably more experienced than the judgement behind it.

An executive can review a sophisticated recommendation without ever having developed an independent position against which to test it.

That creates a difficult possibility.

Visible capability may improve faster than underlying judgement.

The exposure may therefore remain hidden until the situation no longer fits the pattern the technology has already seen.

"Visible capability may improve faster than underlying judgement."

What I’ve Noticed

Perhaps the safest way to use AI isn’t to teach people how to get better answers from it.

It may be to teach them how to disagree with it intelligently.

Much of the training around generative AI understandably focuses on extraction. Write a better prompt. Give the system more context. Specify the format. Refine the answer. Automate the routine work.

All useful.

Professional judgement develops rather differently.

It develops when something refuses to fit.

You form a view before everyone else has one. Evidence conflicts with experience. A forecast looks technically defensible but commercially unlikely. A recommendation sounds persuasive, yet one assumption keeps bothering you.

You stay with the discomfort long enough to understand why.

Those moments are inefficient.

They are also where judgement accumulates.

AI is exceptionally good at removing cognitive friction. That is part of its value.

The more demanding question is which friction wastes time and which friction develops capability.

That distinction may matter more than adoption rates.

"Perhaps the safest way to use AI isn’t to teach people how to get better answers from it.

It may be to teach them how to disagree with it intelligently."

What This Means

The organisational exposure may initially resemble progress.

Papers arrive faster. Analysis becomes cleaner. Meetings are better prepared. People process greater volumes of information. Junior employees appear capable of operating further beyond their experience.

Executives can move between issues with analytical support that would previously have required considerably more time.

None of those outcomes should cause concern on their own.

The tension appears when capability with the system develops faster than capability to evaluate the system.

People become experienced partly by encountering consequences. They make imperfect judgements. They discover where assumptions fail. They learn which questions should have been asked earlier. They recognise patterns because they have previously watched those patterns unfold badly.

Much of that development occurs through doing difficult work.

If AI progressively completes the cognitively difficult parts of the task, organisations may eventually find that the quality of visible output has developed faster than the judgement supporting it.

That becomes particularly important in succession and promotion.

A future senior leader may arrive with years of excellent AI-assisted output, increasingly sophisticated analysis and apparent exposure to complex decisions.

The question will be how much of the underlying judgement was actually exercised.

For experienced executives, there is a different pressure.

Years of accumulated judgement make AI enormously useful because there is something substantial against which to test its output.

Pressure can change how consistently that judgement is exercised.

A late paper arrives. The diary is full. The recommendation is coherent. The assumptions appear reasonable. AI has already synthesised the evidence.

Review can quietly become acceptance without anyone consciously deciding that it should.

The twenty minutes saved may have contained the doubt that would otherwise have changed the decision.

The developmental challenge around AI may therefore be less about protecting people from technology and more about protecting the moments in which judgement is still required.

"Pressure Test"

Discover where your organisatin is leaking critical thinking

Pressure-test this:

If AI disappeared tomorrow from one important decision process in your organisation, would the people responsible for that decision still possess enough independent understanding to reconstruct the reasoning, challenge its assumptions and defend the conclusion?

If that produces hesitation or disagreement, responsibility and understanding may have begun moving apart.

The first senior move is unlikely to involve reducing AI adoption. It is to identify the moments of judgement that are commercially too important to remove from the human part of the process.

The productivity case for AI is becoming easier to see.

The judgement question will take longer to surface because the early signs may look like improved performance.

The organisations worth watching will be those that learn which cognitive friction can safely disappear and which must remain because capability, challenge and accountability still depend on it.

To pressure test this with me. 
Book a complimentary Discovery Call
Back to blog

Read more blog posts

05 Feb 2026

Boardrooms: The Eye of the Storm

Why boardrooms often feel calm while organisations struggle, how culture adapts to protect leadership, and what happens when dissent is quietly punished instead of surfaced.

Read article
01 Jul 2026

Compliance Is Not Governance: Why Organisations Confuse Evidence with Judgement

Compliance proves obligation. Governance directs judgement. When organisations confuse the two, they risk creating false assurance: evidence without effect and structure without control.

Read article
17 Jul 2025

Why Your Team Isn’t Learning Fast Enough

(And What To Do About It)

You can deliver fast. But can you learn fast enough to keep delivering the right things?

Most teams confuse movement with momentum — speed with sense-making.

Read article

Get in touch

We want to hear from you

If you're ready to break bias, decode decisions and unlock success, we're here to help. Let's get your transformation journey started!