Skip to main content
What Is the Human-AI Decision Ladder?

What Is the Human-AI Decision Ladder?

I built this tool because executive teams kept discovering, mid-conversation, that a decision had drifted somewhere nobody remembered approving. The ladder makes that drift visible before it becomes a problem.

By · Published

The Human-AI Decision Ladder is a genuinely practical framework for mapping exactly where a given category of decision actually sits between fully human-made and fully autonomous AI action, and comparing that against where the organisation actually, deliberately chose it to sit. Most executive teams, when they map their real decisions honestly against the ladder for the first time, find at least a few sitting one or two rungs higher than anyone consciously approved.

The Five Rungs of the Human-AI Decision Ladder

I use five rungs specifically because fewer collapses genuinely different accountability postures into one label, and more starts producing distinctions too fine to be practically useful in a real conversation with an executive team who need to act on the answer, not just admire its precision. At the bottom rung, a human makes the decision entirely unassisted, AI plays no role at all. One rung up, AI informs the human, surfacing information or analysis, but the human still does the actual reasoning and makes the call independently. The middle rung is where AI recommends a specific course of action and a human has to actively approve it before anything happens, a meaningfully different posture from simply being informed, because the recommendation carries an implicit pull toward agreement that pure information doesn't. Near the very top of the ladder, AI acts first and a human reviews the action only afterward, which reverses the sequence entirely, the decision has already happened by the time a person sees it. At the very top rung, AI acts with no human review at all, anywhere in the loop.

Why the Middle Rungs Are the Dangerous Ones: The bottom rung is obviously safe and the top rung is obviously bold, which is exactly why both get scrutinised. The real risk lives in the middle, where a decision has quietly moved from 'AI informs' to 'AI recommends' to 'AI acts, human reviews', each individual step reasonable on its own, the cumulative position never actually chosen by anyone at all.

Why Decisions Drift Upward Without Anyone Deciding

This is genuinely worth dwelling on for a moment, because it explains why the drift so rarely gets caught by anyone actively looking for it, even conscientious leaders. It's rarely a conscious choice. It happens because each individual step upward feels like a small, sensible efficiency gain. A team that's used to being informed by AI starts, understandably, giving more weight to its recommendations once the recommendations have proven reliable a few times in a row. A team that's used to approving AI recommendations starts approving them faster, then almost automatically, then only spot-checking a fraction of them. None of those individual moves looks reckless in the moment. The cumulative effect, over months, is a decision category sitting two or three rungs higher than anyone would have approved if asked directly, in a single sitting, whether that's actually where it should be.

I'd add a genuinely important distinction here that matters more than it first appears, and that most teams miss entirely the first time through: the ladder position isn't fixed permanently once chosen. A decision category can legitimately move up over time as trust is genuinely earned through track record, and it can also legitimately need to move back down if the context changes, new regulatory scrutiny, a near-miss that revealed a gap, a change in who's actually doing the work day to day. The ladder isn't a one-time exercise. It's a position that deserves periodic, deliberate re-evaluation, the same way you'd revisit any other structural decision as circumstances change around it.

How to Use the Ladder With Your Own Leadership Team

The exercise itself is genuinely straightforward to run, takes a single focused working session rather than a lengthy multi-week project, and consistently produces useful, sometimes genuinely uncomfortable, results worth sitting with afterward. List every meaningful decision category AI currently touches in the business. For each one, ask two separate questions rather than one: where does this decision actually sit on the ladder right now, based on what genuinely happens in practice, and where would this team deliberately choose to place it if starting from a blank page today. The gap between those two answers, not the ladder position itself, is the useful output of the exercise. A wide gap on a low-stakes decision is mostly interesting. A wide gap on a high-stakes decision is the finding that should change how the business operates before its next review cycle, not at some point in the future.

  1. List every decision category AI genuinely touches — Be specific rather than general. 'Hiring decisions' is too broad. 'Initial resume screening' and 'final offer approval' belong on the ladder separately.
  2. Rate where it actually sits, not where you assume it sits — Ask the people closest to the decision, not just the leader who originally approved the AI tool. Actual practice and assumed practice often diverge.
  3. Rate where it should deliberately sit — Answer this as if designing the process fresh today, without anchoring on wherever it happens to currently be.
  4. Prioritise gaps by stakes, not by size — A large gap on a genuinely low-stakes decision matters less than a smaller gap on a decision with real consequences if it goes wrong.

I'd flag one very specific trap that catches almost every single team the first time they run this exercise, because I've genuinely watched it happen consistently enough to name it directly here: rating the ladder position based on policy rather than practice. A business will often confidently state that a decision sits at 'AI recommends, human approves' because that's what the original process document says, while the people actually doing the work describe something closer to 'AI acts, human reviews occasionally', because the approval step has quietly become a formality nobody genuinely interrogates before clicking through it. The honest rating has to come from what actually happens on a Tuesday afternoon, not from what the process documentation claims should happen. Those two answers diverge more often than most leadership teams expect going in.

What to Do Once You've Found a Gap

Once a genuine gap has shown up, resist the urge to fix everything in that same meeting. Pick the two or three gaps with the highest actual stakes attached, close those deliberately and properly, and schedule the rest for a follow-up rather than redesigning every decision category's position in one sitting. Rushed fixes applied to every gap at once tend to be shallow and don't survive contact with the next busy quarter, whereas a smaller number of properly closed gaps, with a named owner and a real review cadence attached, tend to actually hold.

It's worth being clear about this before running the exercise, because it changes the tone of the conversation entirely. Closing a gap doesn't always mean pulling the decision back down the ladder. Sometimes the honest answer is that the decision has earned its current position, the higher rung is genuinely appropriate given the track record and the actual stakes involved, and the useful outcome of the exercise is simply making that position deliberate and documented rather than accidental. Other times the gap reveals a decision that drifted somewhere it genuinely shouldn't be, and the fix is pulling it back down a rung, reinstating a human approval step that had quietly stopped happening in practice, or assigning explicit review where review had become a formality nobody actually performed carefully any more.

A genuinely useful discipline once a gap has been closed is explicitly naming who owns re-checking it, and on what schedule, rather than treating the fix as permanent. A decision pulled back down a rung today can just as easily drift back up over the following year through the exact same accumulation of small, reasonable steps that caused the original drift. Building a recurring review into the calendar, even briefly, quarterly for high-stakes decision categories, is what prevents the exercise from becoming a one-time correction that quietly unwinds itself over the following months without anyone noticing until the next crisis forces another look.

The Distinction That Actually Matters

I built this specific framework, and named it a ladder deliberately rather than a scale or a spectrum, because a ladder implies you climb it one rung at a time, one deliberate step after another, and each rung is a distinct, nameable place you can actually stand on, rather than a blurry point somewhere along a continuous line that's hard to pin down or defend when someone asks you to justify it. That specificity matters in practice. Asking an executive team where a decision sits on a vague spectrum from human to automated tends to produce vague, hedged answers. Asking which of five distinct rungs it occupies right now produces a much sharper, more honest conversation, because there's nowhere comfortable to hide between two clearly defined positions.

The Human-AI Decision Ladder isn't a tool for slowing AI down or for pushing every decision toward full autonomy either. It's a tool for making sure the position a decision actually occupies is the position someone deliberately chose, rather than the position it drifted into through a sequence of individually reasonable steps nobody stopped to add up. Most organisations discover the drift only after something's gone wrong. The ladder is how you find it before that happens, while the fix is still a straightforward conversation rather than a genuine crisis response.