Coaching exists to build other people's judgement. It's the mechanism, more than any course or deck, that turns a manager's feedback into a rep's habit, or a leader's intent into a team's behaviour. Which raises an odd, slightly recursive question: if a coach's job is to develop someone else's judgement, who develops the coach's? And now that AI can hold a coaching conversation almost as well as an early-career human coach, does that answer change?
I went looking for actual evidence rather than opinion here, because this is exactly the kind of question people answer from vibes. What I found was more interesting, and more specific, than "AI is coming for coaching."
Can AI actually coach?
Better than most people assume, and worse than a good coach, at the same time. A review of the evidence by Tomas Chamorro-Premuzic in Forbes found AI coaching performs at roughly the level of an early-career certified human coach on formal coaching standards, and can competently handle up to 90% of routine coaching interactions: goal structuring, tracking progress, prompting reflection, being available at 11pm when a human coach obviously isn't. In text, AI empathy is sometimes rated higher than a human's, because it's patient, consistent, and never has an off day.
But there's a catch worth sitting with. The same research describes what Chamorro-Premuzic calls an empathy paradox: when people know they're talking to AI, they report feeling less understood, even when the actual words are identical to a human coach's. The content of the response isn't what's missing. What's missing is the fact that someone was actually there, at some cost to themselves, choosing to pay attention.
That's the same shape of problem I wrote about in an earlier piece on AI, herd mentality and the doorman fallacy: a role has a visible task (ask good questions, track goals, hold you accountable) and an invisible one (someone genuinely caring what happens to you). Automate the visible task well enough and you can fool the measurement. You can't fool the feeling, at least not yet, and the research above suggests people notice the difference even when they can't articulate why.
So: replace, complement, or noise?
The honest answer, based on where the evidence actually points, is "it depends which layer you're coaching."
- Entry level: mostly complement, sometimes genuine access. Most people never got coached at all. AI coaching turns "no coaching" into "some coaching" for a huge number of people who couldn't previously get it, at close to zero marginal cost. That's not noise; that's real value that didn't exist before.
- Mid-level, specific skills: strong complement. For discrete, well-defined skills (delegation, feedback delivery, running a meeting) where a coach applies a consistent framework, AI is genuinely good, and consistent in a way tired humans aren't.
- Executive level, judgement and identity work: still overwhelmingly human. Pinnacle's evidence-based framework for CHROs found roughly 70% of leaders still prefer personalised human guidance for the development that actually matters at that level: navigating ambiguity, office politics, identity, high-stakes trust. That tracks with the judgement article I wrote earlier: the mechanical part of a job compresses fast, and what's left, uncomfortably exposed, is judgement, which is exactly what senior coaching is for.
HR Executive has started describing this as a three-tier coaching stack: AI for structure and volume, human coaches for the emotionally complex and high-stakes work, and a deliberate handoff between the two rather than a straight substitution. That maps almost exactly onto the Know → Practice → Apply → Reinforce framework I've used for enablement generally: AI is genuinely strong at Know and Practice (drilling, structure, always-available repetition). Reinforce, the step that actually changes behaviour, is still where a human has to show up.
The part that worries me more
Here's where this connects back to herd mentality specifically, not just the doorman fallacy in general. A meaningful amount of the current AI-coaching rollout in organisations isn't happening because someone diagnosed a real coaching gap. It's happening because a competitor announced an AI coaching platform, and not having one started to feel like the risk. That's the exact pattern from the earlier piece: adopt the visible thing everyone else is adopting, skip the diagnosis of what the actual coaching gap was, and call it transformation.
Applying Copy → Diagnose → Redesign → Reinforce to coaching AI
- Copy (honestly): most organisations buying AI coaching tools right now are matching a competitor's announcement, not fixing a diagnosed development gap. Fine as a starting point; not fine as the whole plan.
- Diagnose: which parts of coaching in your organisation are actually structure-and-volume problems (not enough coaching happens) versus judgement-and-trust problems (the coaching that happens isn't good enough)? AI fixes the first. It doesn't touch the second.
- Redesign: don't bolt an AI coaching tool onto an unchanged manager population and call it done. Redesign what human coaching time is actually spent on once the routine, repetitive part is handled elsewhere.
- Reinforce: the follow-through that makes coaching stick (real accountability, a real relationship, real stakes) still needs a human on the other end who'd notice if you quietly stopped trying.
Who coaches the coach, then?
This is where it gets genuinely interesting, because the coaching profession is already living this question, not just theorising about it. Coaches have long used supervision, sometimes called meta-coaching, as their own version of coaching: a reflective relationship with a more experienced practitioner whose job isn't to direct the coach's work but to help them see their own patterns more clearly.
AI is starting to enter that space too, and the way it's being positioned is telling. Reflecta, a reflection tool built specifically for coaches, describes itself in its own materials this way: "Reflecta does not replace supervision. It deepens it." Its own pitch draws the same line this whole piece keeps drawing: "AI can optimise, humans can care. AI can calculate, humans can feel significance."
Even the people building AI tools for coaches, who have every commercial incentive to oversell what the tool can do, are positioning it as a mirror between sessions, not a substitute for a human supervisor. That's a stronger signal than any market-size projection: the closer you get to the people who actually understand what coaching is for, the more careful they get about the word "replace."
The human question
So here's the question I keep landing on, and it's not really about AI. If a coach's real job is to build someone else's judgement, and supervision exists because judgement doesn't develop itself, then: who is watching you closely enough to notice if your own judgement quietly stopped developing? Not your calendar of 1:1s. Not a dashboard. A person who'd actually notice.
AI can ask you that question at 2am and never get tired of your excuses. It genuinely can't be the answer to it.
---
Sources and further reading:
- Tomas Chamorro-Premuzic, "Does AI Coaching Work? What The Evidence Shows", Forbes: the empathy paradox, and the "roughly early-career human coach" performance benchmark.
- Pinnacle, "Can AI Replace a Human Executive Coach?": the 70% leader preference for human guidance figure.
- HR Executive, "The new coaching stack: a 3-tier model for AI and human coaching": the tiered AI/human handoff model.
- Reflecta, an AI reflection tool for coaches: the "does not replace supervision, it deepens it" framing.
- Institute of Coaching, "Reflective Practice and Supervision for Coaches": background on coaching supervision as a professional norm, not a remediation programme.
- This article was researched and drafted with Claude (Anthropic), consistent with how this site discloses AI use elsewhere.