The Human-AI Partnership Framework

Where human judgment belongs

Understanding the HITL Maturity Scale

Why human involvement is measured on a five-point scale — and why the goal isn't the top of it

If you've followed this framework for a while, you've likely seen the maturity levels — M1 through M5 — on the dashboard without a clear explanation of what they mean or why the scale is shaped the way it is. That gap in understanding is common, and it's worth closing directly, because the scale is doing more work than it looks like at a glance.

Start Here: Design, Not Safety

Most people's first association with "human-in-the-loop" is safety — keeping a person in the chain to catch AI errors before they cause harm. That's correct, but it's not the whole idea, and it's not what this scale measures.

HITL Design vs. HITL Safety

HITL Safety asks: where must humans stay involved to prevent harm? HITL Design asks: where should humans stay involved to produce the best outcome? The design question is broader, more strategic, and it's the one this framework is built around.

Framed as a design question, human involvement isn't something that just happens to a process as AI tools get adopted — it's something an organization can decide, deliberately, activity by activity. The maturity scale is how that decision gets measured and tracked over time.

The Three Layers Every Process Sits In

Before the scale makes sense, it helps to see the mental model underneath it. Every piece of enterprise work belongs in one of three layers:

Human Layer
Judgment & Oversight
Approvals, interpretation, escalations, ethical decisions — the work that requires wisdom, context, and accountability.
Automation Layer
Execution
Routing, matching, scheduling, orchestration — rules-based workflows that are consistent, reliable, and repeatable.
AI Layer
Prediction & Generation
Drafting, summarizing, pattern detection, forecasting — fast, scalable, tireless processing and generation.

The most common failure in human-AI operating model design is work sitting in the wrong layer — humans doing Transaction work that automation should handle, or AI generating outputs for Decision work where human judgment and accountability are non-negotiable. The five pattern types below are what determine which layer an activity actually belongs in. Full detail in the HITL Design Methodology →

The Five Patterns, and Two Numbers Per Pattern

Every activity in the framework's process hierarchy gets classified into one of five patterns, based on what the work fundamentally requires — not what tools happen to be in use today. Each pattern carries two reference numbers: a typical current-state baseline, and where that pattern's design intent target sits.

PatternRequiresTypical baseline → design intent
DecisionJudgment, authority, accountability — approvals, selections, sign-offs that can't be delegated without real risk.90% → 75%
KnowledgeExpert interpretation and synthesis — AI can assist with research and pattern detection; humans lead the interpretation.80% → 60%
DocumentDrafting, summarizing, structured content creation — AI can produce first drafts; humans review, refine, and approve.65% → 35%
TransactionRules-based, repeatable processing and routing — high automation potential; humans handle exceptions and quality oversight.45% → 30%
ExceptionNon-standard events and edge cases — AI flags and routes; humans investigate and decide.92% → 75%
The Design Intent Number Is Not a Target to Minimize

A 75% design intent for Decision work doesn't mean 25% of decisions should be made by AI. It means AI should be handling roughly a quarter of the supporting analysis, preparation, and documentation — while human judgment still owns the actual decision. The goal is optimal balance, not maximum automation.

From Two Numbers to a Five-Point Scale

The two reference numbers above — baseline and design intent — are the anchor points. The five-level scale (M1–M5) is how the framework tracks the real, ongoing journey between and beyond those two points for every individual activity, so progress is a measured trajectory rather than a single before/after snapshot.

LevelNameWhat it means
M1UndesignedThe Current H% baseline — human involvement exists because that's how the process happened to evolve, not because anyone decided it should sit there.
M2EmergingAutomation has started to reduce human involvement, but piecemeal and reactively — individual tools bolted on, not a deliberate redesign.
M3 ★Design IntentThe Design Intent H% target — the point where human involvement is deliberately engineered, grounded in the activity's pattern type. This is the framework's actual target state, not a waypoint to M5.
M4OptimizedThe M3 design intent, refined and tuned — friction removed, workflows smoothed — without changing who's accountable for what.
M5LeadingThe technical ceiling: as much automation as the pattern type genuinely allows. For Decision- and Exception-pattern work, that ceiling stays close to M3, because it's bounded by what shouldn't be delegated, not by what's technically possible.

M1 is the baseline; M3 is the design intent target. M2 is the honest, often messy middle — automation arriving unevenly, before anyone's redesigned the work on purpose. M4 and M5 exist because design intent isn't the technical ceiling — once a process is deliberately designed at M3, it can still be refined (M4) and pushed toward its real limit (M5). For Transaction- and Document-pattern work, that limit is genuinely close to full automation. For Decision- and Exception-pattern work, it isn't — and that difference is the whole point of the next section.

How the Human Role Is Guarded at Design Intent

This is the mechanism that makes M3 more than a label, and it's the piece most engagement and questions on this topic seem to miss: the pattern classification isn't a description added after the fact — it's the reason each activity's numbers land where they do, and it's what actively protects certain work from compressing toward M5 the way Transaction or Document work does.

PatternHow guarded, in practice
DecisionMost heavily guarded. Real examples from this framework's built roles: approving a job requisition, selecting a candidate, negotiating an offer, and establishing a succession plan all show gaps as shallow as 12 points — far tighter than the 90%→75% reference. The design intent barely moves these activities, because the judgment itself is the point.
ExceptionEqually guarded, for a related reason: an Exception-pattern activity is, by definition, the case a standard workflow couldn't resolve. Automating the exception-handling would just move the exception up a level, not remove it. These activities carry the highest baselines found in the framework — up to 95% — and stay high through M5.
KnowledgePartially guarded. Genuinely compresses with AI assistance (synthesis and drafting speed up), but the interpretive judgment underneath doesn't fully delegate — gaps land close to the 80%→60% reference.
DocumentCompresses the most of any pattern with real human accountability still attached — gaps of 30+ points are common, since generating a first draft is exactly what generative AI does well. A human still reviews and approves; the drafting labor largely moves.
TransactionCompresses fastest and furthest. Least guarded by design, because it's rules-based by definition — there's no judgment being protected, just execution.

The practical effect: within the same process, at the same overall maturity level, two activities can have wildly different human-involvement trajectories — not because one was designed more carefully, but because the pattern itself defines what "designed well" looks like. A shallow gap on a Decision or Exception activity isn't unfinished work. It's the design intent succeeding — protecting exactly the judgment that shouldn't be automated away.

Reading the Numbers

When you see a Δ M1→M3 figure anywhere in this framework, it's worth reading against the pattern behind it, not just the headline number:

A Note on Where the Numbers Come From

The baseline and design intent percentages by pattern are informed estimates — developed through research and reasoning, not yet validated against large-scale organizational data. Where this framework has done real per-activity assessment (as in PCF 7.0, built out in full), the actual figures for a given activity can differ from the pattern's generic reference numbers above; the reference table is a starting point for a pattern, not a substitute for assessing the real activity in front of you.

← The full HITL Design Methodology
The complete methodology this page draws from — the HITL Stack, the five pattern types, the ratio methodology, a full worked example, and the governance model.

Where This Shows Up

This same logic runs through everything built on this framework: the Maturity Table's per-activity M1–M5 rows, the RACI Table's Accountable/Responsible assignments, every Value Stream Profile's AI Capability Map, and — on the People side — every Role Transition Profile's account of which parts of a role's work are genuinely reinvented versus which stay anchored. Same five patterns, same design-intent logic, queried from a different angle each time.