I sit across the table from CEOs who have spent six or seven figures on leadership development, and when I ask what it delivered, most of them go quiet. The answer usually isn't disappointing, though sometimes it is. There simply isn't an answer at all. Nobody wrote one down. Most companies can't measure leadership development ROI, and the reason has nothing to do with how slippery leadership is as a concept. It has everything to do with what they decided to track in the first place.
This isn't a hunch about how these programmes usually go. McKinsey researchers Pierre Gurdjian, Thomas Halbeisen and Kevin Lane looked closely at why leadership programmes fail, and one finding sits underneath most of the others: companies, in their words, "pay lip service to the importance of developing leadership skills" while carrying nothing that could prove what the spend actually bought them. Say that back to yourself slowly. The organisations spending the most on leadership are frequently the ones with the least idea whether it's working.
I've sat in enough of these rooms to recognise the pattern before the CFO even finishes the question. Someone presents a slide with an attendance number and a smiley-face average. Everyone nods. Nobody asks what the programme actually changed, because nobody built the means to answer that question, and asking it out loud would only expose the gap. It's a quiet failure, not a loud one. The invoice gets paid, the calendar clears, and the organisation moves on believing something useful happened, because it wants to believe that, not because anyone checked.
The Discipline You Need to Measure Leadership Development ROI
The distinction that gets missed matters more than it sounds like it should. It's not that leadership ROI shows up somewhere down the track and businesses simply haven't waited long enough to see it. It's that most businesses never built the instrument that would read it, whenever it arrived. You cannot measure something you never defined, never baselined and never checked twice. That is a measurement problem, not a patience problem, and treating it as the second lets everyone off the hook for fixing the first.
Ask a typical L&D leader how last year's flagship leadership programme performed and you'll get a satisfaction score. Ninety-one percent said the workshop was valuable. The venue was good. The facilitator was engaging. That tells you whether people had a pleasant few days out of the office. It does not tell you whether a single leader in that room now makes better calls under pressure, delegates work they used to hoard, or handles a difficult conversation any differently six months later. Those are two entirely different questions, and only one of them gets asked.
What a Real Measurement System Actually Requires
A workshop is not a system. A system has inputs, a defined output and a way of checking whether the output arrived. Most leadership development spend skips straight from input (the programme) to a feeling (the satisfaction score), with no defined output and no check in between. Building the missing instrument is not complicated, but it does require four things most programmes never bother to put in place.
- A defined behaviour, not a theme: Name the specific decision or behaviour this programme is meant to change, before it starts. Not 'better leadership'. A named shift specific enough that two observers watching the same meeting would agree whether it happened.
- A baseline before the room: Measure where each leader actually stands on that named behaviour before the programme starts. Without a starting point, nothing you observe afterward means anything at all.
- A checkpoint after the room empties: Come back at ninety days, then again at six months. Not to ask whether people enjoyed it. To check whether the named behaviour actually moved, in the room where it counts.
- An owner who isn't the facilitator: The person accountable for the result should not be the person who ran the workshop. Otherwise the only number anyone has an incentive to report is the one that flatters the programme.
Why Satisfaction Surveys Are the Wrong Instrument
Run a satisfaction survey and you'll learn how people felt leaving the room. That's a real thing to know, and it isn't nothing. But it measures the wrong layer. It measures the classroom, not the job. McKinsey's research puts a number on why that gap matters so much: adults typically hold onto only around 10 percent of what they hear in a classroom lecture, while learning by doing lifts that figure to nearly two thirds. A satisfaction score taken at the end of a lecture-heavy workshop is measuring the layer of learning least likely to survive the walk back to the desk.
That single statistic should change how you read every glowing feedback form that ever crosses your desk. A high satisfaction score after a classroom-style session tells you the content landed well in the room. It says almost nothing about what survives contact with a Monday morning full of real decisions. If the programme was mostly lecture, most of what was taught is already gone before the behaviour it was meant to change ever gets tested.
This is where engagement and impact get swapped for one another without anyone noticing, and it happens so smoothly that most executives never catch the substitution. Engagement is easy to see: full seats, good energy, warm applause at the close. Impact is harder, because it only shows up later, in a different room, under different pressure, and nobody sent anyone to go and look. A programme can score highly on the first and deliver almost nothing on the second, and the only reason that goes unnoticed is that the second was never measured to begin with.
- Whether the room was comfortable and the schedule respected people's time
- Whether the facilitator was likeable and easy to follow
- Whether the content felt relevant in the moment it was delivered
- Whether people would recommend the session to a colleague
Every one of those four questions is worth asking. But not one of them is the question a CFO actually needs answered before approving a leadership development budget. What follows is what a satisfaction score cannot tell you, no matter how high it climbs.
- Whether a leader delegates differently three months later
- Whether a difficult conversation gets handled with more skill and less avoidance
- Whether the specific behaviour the business needed to change actually changed
- Whether the money was well spent, against any defined standard
What Measuring Leadership Development Actually Looks Like
Building the missing instrument is a sequence, not a single decision. It has to happen in order, because each step depends on the one before it. Skip the first step and the last three are just theatre dressed up as rigour.
- Name the decision, not the topic: Before you book a single session, write down the specific decision or behaviour this programme needs to change, in language specific enough that two observers watching the same meeting would agree whether it happened.
- Baseline it properly: Use a real assessment, not a gut feel from HR, to establish where each leader is starting from on that specific behaviour. This is the number everything else gets measured against, and it has to exist before the programme starts, not after.
- Build the checkpoint into the calendar, not the wish list: Put a 90-day and a six-month review into the programme's actual budget and calendar at the outset, the same way you'd schedule a project milestone. A checkpoint that isn't scheduled is a checkpoint that will not happen.
- Report the number that matters to the person who signed the cheque: The CEO or CFO who approved the spend should see the behaviour-change number, not the satisfaction score. If the only report that reaches them is a happiness average, you have told them nothing about return.
10%: retained from classroom lectures: Adults typically retain only around 10 percent of what they hear in a classroom lecture, even after basic training, per McKinsey's 2014 research on leadership development.
~2/3: retained from learning by doing: The same research found retention rises to nearly two thirds when adults learn by doing, which is exactly the layer most satisfaction surveys never touch.
The Real Cost of Not Measuring
This is where the cost actually sits, and it is bigger than the wasted training budget. A leadership function nobody can measure is a leadership function nobody can defend. When the next downturn arrives and someone finally has to justify the line item, "people said they enjoyed it" will not survive that conversation. I've written before about what actually causes leadership development programmes to fail, and a missing measurement discipline sits near the top of that list every time I've gone looking for the real reason.
Boards are starting to ask the question directly, even where they didn't used to. A board that once approved a leadership development line item without much scrutiny will now ask what it bought, in the same breath it asks about marketing spend or a technology rollout. An L&D leader who can only answer with a satisfaction average is not just under-prepared for that meeting. They are describing a function the business has no real way to trust, at exactly the moment trust is what's being tested.
The same discipline applies to individual coaching spend, not just group programmes. When I've walked through the real numbers behind executive coaching ROI, the organisations that could actually answer the question had done the unglamorous work upfront: a baseline, a named behaviour, a follow-up that wasn't optional. The ones that couldn't answer it had skipped straight to the workshop and hoped the number would show up on its own.
If you want the fuller playbook for tracking capability at organisation scale rather than programme by programme, I've laid that out separately in how to measure leadership capability across an organisation. It's the difference between running one honest experiment and building a standing system that keeps reporting the truth long after this year's programme is finished. That's also the real dividing line between a leadership programme and what I'd call evidence-based leadership development: one runs on hope, the other runs on a number you checked twice.
It's also why the assessment work inside Capability AI starts with a baseline before a single session gets booked, and why leadership assessments exist as a distinct, repeatable instrument in their own right, rather than a one-off survey attached to the end of a workshop.
The Question to Ask Before the Next Programme
Before you approve the next leadership development spend, ask one question first: what will we measure, and against what baseline? If nobody in the room can answer that in a single sentence, you are not funding leadership development. You are funding a pleasant couple of days out of the office and calling it strategy.
Fix that, and everything downstream gets easier. You stop guessing whether the programme worked. You stop defending a budget with a feeling instead of a number. And the next time someone at the top of the business asks what leadership development actually bought the company, there's a real answer waiting instead of a long silence. The only leadership investment worth defending is one you started measuring before you ever needed to defend it.
What to connect before publication
The useful test for why most companies can't measure leadership development roi (and what that costs you) is whether a leader can make a better decision in the work that already exists. Start with one live case, write down the judgement that was used and compare it with the result that followed. Then ask what the team would do when the same pressure appears next month. That question keeps the article close to practice rather than turning it into another list of principles. It also gives a sponsor something concrete to challenge: the owner, the evidence, the trade-off and the point at which the plan should change.
A second check is transfer. If the idea in this article depends on one senior person, a special workshop or a project team standing beside the work, it has not yet become an organisational capability. Try the decision with a new manager or an unfamiliar case. Note what remains clear, what needs coaching and what still relies on personal memory. That small test gives why most companies can't measure leadership development roi (and what that costs you) a boundary. It says what the approach can support now, what must be tested next and which claim should not be made until the evidence is stronger.
Related decisions to test
This topic sits beside three decisions that often get separated in practice. Read the team collaboration and effectiveness framework. The comparison is useful because it keeps the argument in this article specific: which decision is changing, who owns it and what evidence would show that the change has travelled into everyday work? Use the linked pieces as contrasts, not as a substitute for the judgement required here.
