Framework · AI autonomy & interface design

Designing for variable trust

AI is right most of the time, wrong some of the time, and occasionally not there at all. Most interfaces are built for the first situation only.

This framework places work on a spectrum of AI autonomy, then derives the interface from that placement.

Design for both: when the AI contributes, and when it does not.

The seven-rung autonomy ladder, Manual to Autonomous with each rung's verb, beside an illustration of a winding path climbing toward a glowing marker, with a figure standing at its base looking up the path.
1 · The problem

Autonomy is not a setting

Imagine a car with several modes of automation.

On a clear, well-mapped highway, the car can lead while the driver supervises. In construction, it should reveal what is uncertain and return more control. If its map becomes stale—or the system goes offline—the steering wheel cannot disappear.

Most AI products are designed as if autonomy were one setting for the whole product. Teams choose how much AI to use, then design a screen around that choice. That produces three failures:

  • • People asked to check work they cannot evaluate.
  • • Finished output handed over with no way to look behind it.
  • • Work that stops the moment the AI does, halfway through, with a deadline still running.

The mistake is answering "how much AI?" once for the whole journey. The useful unit is each decision within it. And because AI may be right, wrong, or absent, every decision needs a range: a default when AI contributes fully, an escape when its output becomes doubtful, and a manual floor when AI is not there.

This framework places each decision on that range, then derives the interface required at every position. Five questions get you there.

2 · Question

What is the unit?

Not a product. Not a workflow. Something smaller.

On one trip, lane keeping, merging, and responding to an unmarked detour may deserve different levels of automation. The trip is only the container; each decision is the unit being placed.

Split the journey into flows. A new flow starts when any of these changes.

  • • The job. Model, configure, monitor are different work.
  • • The artifact's state. Proposal, signed, configured, live.
  • • The trigger. Someone starting work and an event arriving are different flows even when the screens match. This is the split teams miss.
  • • The actor. Reviewer to specialist, contributor to manager. Miss this and you place a flow at a rung no one person occupies.

Split flows into steps, and mark the decisions inside them. A step is an action. A decision is a step where a reasonable person could have done otherwise. Uploading a file is not a decision. Accepting a proposed change is.

Placement attaches to decisions. Flows and steps are containers. A multi-part document is not one decision. Each part someone accepts or rejects is one, and two of them in the same step can sit at different rungs.

A vertical tree: one journey branches into flows, one flow into four steps, and one step into two decisions — the decision is the unit you place.

3 · Question

Who leads the decision?

The spectrum is a climb. Start at Manual. Every rung above has a gate below it. Stop at the first that closes.

Gates two and three are one question at two scales. The distance between them is the authorship line. Below it the person composes and AI supplies. Above it AI proposes and the person adjudicates.

One decision can hold several placements. A decision recurs across many cases, and those cases are not alike. Carve them into slices by attribute, and each slice climbs to its own height. Run the walk per slice.

Two questions sit outside the walk.

Can the recipient tell it is wrong?

If not, validation moves to someone who can. That sets who stands on the rung, not which rung. It caps only through capacity: no qualified reviewer at volume means the rung is unreachable.

What if it never gets done at all?

Every gate asks how high you may climb. This asks how far down you must still be able to work. The higher you place, the more that floor costs: a rarely used path breaks quietly and is found broken at the moment it is needed.

The walk: a solid arrow forward for the gate question, a dashed arrow back for what closes it.

The answer is a band, not a point.

  • • Clear road, current map — Default: full AI contribution.
  • • Construction or stale map — Escape: reduce AI authority and support investigation.
  • • System unavailable — Floor: the person can complete the work without AI.

The driving mode changes with the road segment. In the framework, placement changes with the decision.

The band measures how much AI contribution you can count on. It has three positions, and something specific sends you to each.

• Default. The placement. Full contribution, and where decisions land.

• Escape. Below the default, as far down as the doubt goes. Reached when the output is wrong, or built on stale reference data. Stale is the one that fools every confidence threshold, because confidence stays high while the answer does not.

• Floor. No AI at all. Reached when the model is absent: budget exhausted, rate limited, timed out. Nothing to doubt, because there is nothing there.

Design all three. Most teams design the first and assume the other two.

Autonomy is earned. Placement moves up on track record. Drift or an incident pulls it back down.

The band: full-contribution edge, low-contribution escape, and the manual floor.

4 · Question

What does the person do?

The band gives you three positions. Each one puts the person in a different job, and the rung alone does not tell you which.

The unit of attention changes at gate four. Below it the person works one case at a time. Above it, a population. That is a different job, and usually a different person, which is why the actor test splits flows in the first place.

Triage is not a primitive. It is Verify on the flagged plus Navigate across the unflagged. Teams build the queue and skip the population view, and the rung is incomplete without both.

Govern runs alongside every rung. Someone owns the threshold, the sampling rate, the placement. That happens whether the work sits at Assess or at Autonomous, usually by someone outside the task flow.

The verb inherits the band. The escape is the affordance to drop down the verbs. Usually one: Verify falls to Assess, Navigate falls to Verify on a slice, Triage falls to Navigate. Deeper doubt drops further.

When the model is absent there is nothing to assess, so the verb lands on Author. That is the floor, and it has to be somewhere a person can work.

Verb and unit by rung: one case from Manual to Review everything, many cases at review in aggregate and review exceptions, the system itself at Autonomous.

5 · Question

What may the interface render?

Four types. Each applies to a region, which is usually a whole screen and occasionally part of one.

Unaided. No AI content.

Slotted. The designer set the layout. AI fills named slots.

Assembled. AI arranges certified components into designer-defined patterns. Composition varies, vocabulary does not.

Generated. AI produces the region freely. No grammar, so nothing to certify. Excluded wherever the output carries consequence.

Test: can you screenshot the empty state? If the wireframe exists before you know the case, it is Slotted. If it depends on the case, it is Assembled. A repeating row with a fixed template is still Slotted. Varying the count is not composing.

The interface pattern scale, Unaided at the bottom to Generated at the top, with the certification limit before Generated.

Autonomy and pattern are one decision read twice. The rung sets what the person does, and what the person does picks the pattern. Author by hand and it is Unaided. Let the AI inform or produce parts — you assess or compose — and it is Slotted or Assembled, depending on whether the AI fills a slot or arranges a region. The moment your job becomes oversight — verify, navigate, triage, govern — it collapses to Slotted: one legible instrument, the same every time.

Generated never lands on a rung. It is the explore mode, off the commit path.

Each autonomy rung maps to the interface pattern its job needs: Unaided for authoring, Slotted or Assembled for making with AI, Slotted all the way up for oversight.

Bringing it together

One ladder, three surfaces, per decision

Each band position renders as a type, but the type is not fixed by the position. On one decision the default is Slotted and the escape is Assembled. On another, the task is ambiguous enough that even the default is Assembled and the escape stays Assembled below it. The pattern each position takes is a call about the task, not a rule you apply once.

That is the whole point. Nothing here fixes the mapping for you: you read the task, and the task tells you how much structure the default can afford and how far the escape has to fall. The only line that holds is the floor. A floor that still needs the model is not a floor, so the bottom of every band is Unaided — an AI-powered fallback is not a fallback.

Explore dynamically, commit structurally. A decision of record lands on a Slotted region. It may sit inside an assembled surface. It may not be a sentence in generated prose with a button under it.

A flow gives you one ladder per decision. Five decisions, five bands, each at its own height with its own three surfaces. That is the build plan, and it is why placement is worth doing before anything is designed.

One flow, several steps: a default, an escape, and a floor across the steps, with each point marked by the pattern its rung lands on — so the same line takes different patterns as it rises.

6 · Question

Which patterns?

Seven rungs, each with a default, an escape, and the specific ways it fails.

At every rung

Availability

The AI stops at step 3 of 5. The deadline does not move. What does the person see?

Availability: five steps in a row, a break between step 3 and step 4. Steps 1 and 2 done, step 3 unfinished at the break, steps 4 and 5 continue on the unaided floor. Three Slotted patterns mark where they act: forecast before step 1, handoff at the break, unaided floor across steps 4 and 5.

Three patterns, all Slotted, needed at every placement.

• Completion forecast. At step 1, say whether there is capacity for all five. If not, offer the choice up front: run the steps that matter most, or stay manual throughout. Budget exhaustion is arithmetic, not a surprise, so it does not belong in an error state.

• Handoff artifact. At the break, state what is finished and verified, what is unverified, and how long is left.

• Unaided floor. The remaining steps finish without AI, on the pre-AI path, kept in working order.

Anti-pattern. A generic failure message. It discards the partial work, says nothing about what was done, and leaves the clock running.

Test. Kill AI access at step 3 of 5. Can the work still finish on time?

Manual

Author

No AI in the region. Two patterns still matter.

• Unaided-surface label. Where some regions are AI-composed, the absence of a marker is ambiguous, not reassuring.

• On-ramp. Manual is often a position in a band, not a permanent state.

AI informs

Assess

The person produces the work. AI surfaces references. The artifact is untouched.

Default
  • • Peripheral evidence rail. Beside the work surface, never in it.
  • • Base-rate panel. Frequency with a denominator. "34%, n=22" not a badge.
  • • Anomaly marking without ranking. Flag the unusual. Do not sort the queue. Sorting is allocation, and allocation is part of the work.
  • • Scoped corpus controls. What it looked at precedes what it found.
Escape
  • • Retrieval with citations as the payload
  • • Contextual inline query, read-only by construction
  • • What-if probe
Anti-patterns & test

Anti-patterns. A synthesized answer with no sources. A pre-ranked worklist. A percentage with no denominator.

Test. Close the rail. Is the artifact unchanged?

Evidence rail
Base rate: 34%, n=22
Anomaly marks
AI produces components

Compose

AI produces finished parts. The person assembles them.

Default
  • • Structured slot-filling. Designer owns the layout, AI owns the contents, person owns whether each slot stands.
  • • Candidate tray. Parallel options, not one answer. Choosing keeps authorship human. One candidate makes the person an editor, which is the next rung.
  • • Seam marking. Authorship visible per region, persisting after the person edits an AI part.
  • • Draft mode. An explicitly marked not-yet-the-artifact space.
Escape
  • • Targeted regeneration
  • • Component provenance
  • • Rationale for a candidate not taken
Anti-patterns & test

Anti-patterns. An invisible seam. The quietest failure on the spectrum: you have crossed into supervision without installing any of what supervision requires.

Test. Can the person tell which parts they wrote?

Candidate tray
Author-written
AI-authored (seam)
Review everything

Verify

AI authored it. The person checks each item.

Default
  • • Delta as the surface. Show what changed. Rendering the whole output and calling it review manufactures rubber stamps.
  • • Assertion-level provenance. Each assertion linked to source, with version and date. Collapsed by default, expanded on low confidence.
  • • Structured rationale. An enumerable list of what drove the call. Prose reads as persuasion.
  • • Explicit verification act. Commit is a gesture, not a side effect of navigating away.
  • • Forcing function. Record the person's read before the recommendation is actionable. Explanations do not prevent over-reliance. Interrupting the shortcut does.
  • • Review ergonomics. Keyboard disposition. Rejection carries a reason. At 100+ items a day this is load-bearing.
Escape
  • • Source document
  • • Full reasoning trace
  • • Prior similar decisions and how they resolved
Anti-patterns & test

Anti-patterns. Bulk approve. A confidence badge as the only signal. A diff larger than the time available. The documented case: physicians signing denials at 1.2 seconds each, with roughly 90% of appealed denials later reversed.

Test. Is override rate non-trivial, and does time-on-item have a real tail?

Delta (what changed)
Assertion-level provenance
Forcing function: record your read first
Verify
Reject
Review in aggregate

Navigate

The person reads distributions, then chooses what to inspect.

Default
  • • Distribution-first landing. Every item present, shape readable without reading items.
  • • Control totals that tie out. A whole-population check that costs one glance.
  • • Person-chosen decomposition axis. If the product picks the axis, it picked the finding.
  • • Stratified sample with a stated interval. The evidence mechanism for this rung, equivalent to override rate one rung left.
  • • Simulation before commit. Run against real volume, show the outcome distribution, then decide.
Escape
  • • Brushing and drill-through to the rows behind any bar
  • • Baseline comparison against a known-good prior run
  • • Per-item review of a chosen slice
Anti-patterns & test

Anti-patterns. Landing on a pre-sorted top ten, which is an exception queue in costume. A single headline number, which is unfalsifiable and therefore not review.

Test. Does the decomposition axis actually get changed? If nobody touches it, the rung has collapsed.

Note. The public AI UX pattern libraries have almost nothing here. These come from business intelligence, statistical audit, and observability practice.

Distribution
Decomposition axis (chosen)
Stratified sample
Review exceptions

Triage

The system flags. The rest executes.

Default
  • • Queue with typed dispositions. Approve, reject, escalate and edit are four acts with four records.
  • • Visible threshold. An object with an owner, a version, a history. Burying it means nobody can audit the thing that decides what gets audited.
  • • Circuit breaker. Trips on override-rate, volume, or cost anomaly. With no per-item check, population tripwires are the remaining net.
  • • Shadow audit. A random sample of what nothing flagged, reviewed anyway. This is the Navigate half of the composite verb, and the instrument that measures whether the gate into this rung is still open.
  • • Rejection reason capture. Reasons feed threshold tuning.
Escape
  • • Full case inspection
  • • Population context: what the queue did not flag, and why
  • • Threshold change history
Anti-patterns & test

Anti-patterns. Queue volume beyond reviewer capacity, which is a documented attack surface and not only an operational problem. A system defining its own exceptions with no independent check is governing itself.

Test. Is the shadow audit reviewed at rate? Do rejections carry reasons?

Exception queue
Visible threshold
Circuit breaker
Shadow audit sample
Autonomous

Govern

Autonomy is already common, including in regulated work. Payment matching, fraud scoring, duplicate detection, and small-value refunds all run without per-item review, and nobody considers that reckless.

Confidence does not license the rung. Confidence is a property of one item. Autonomy is a policy over a population that includes items nobody has seen. 99.7% on $50 transactions and on $180,000 transactions is the same number and a different decision. High confidence on stale data is exactly what a confident wrong answer looks like.

Five entry conditions. All of them. None is confidence.

  • • Bounded worst day
  • • Institution-side detectability
  • • Reversible inside the detection window
  • • Recipient not carrying an error burden they cannot bear
  • • Governance present

Place a slice, not a whole decision. Under a value threshold, cause in a named set, counterparty in a named list, reference data current, no specialist judgment required. That slice runs autonomously and every other slice of the same decision stays where it was.

Default (Slotted)
  • • Drift monitoring against the accuracy that earned the placement
  • • Slice definition as an owned, versioned object
  • • Capability disclosure
  • • Circuit breaker and pause
Escape (Assembled)
  • • Incident replay
  • • Slice change history and the evidence behind each widening
Anti-pattern, demotion path & test

Anti-pattern. Slice creep. The slice performs, the return case says widen it, and it widens until it holds cases it was never validated on. It is under permanent pressure to expand.

The demotion path. If autonomy is earned, it can be un-earned. Who pulls it back, on what evidence, what happens to work in flight, and how downstream people learn their supervision burden just changed. Under-designed across the field.

Test. When the slice last widened, what evidence supported it and who signed?

7 · Guardrails

Five checks catch a false placement

01

Does the interaction match the placement?

A line-by-line screen is not aggregate review. A pre-filtered queue is not free navigation. One-click approval is not verification. If the interaction does not match, the placement is false.

02

Can the person drop down the ladder?

No escape means the placement is a single setting and the band is decorative.

03

Is the bottom of the band operable with no AI at all?

Kill AI access at step 3 of 5. If the work cannot finish by the deadline, the band has no floor. If a step must be manually completable and cannot be, the placement is too high.

04

Would you know it went wrong?

If detection runs only through someone else's complaint, the placement cannot move up no matter how good the model gets.

05

Who validates, and can they?

If the recipient cannot, validation moves institution-side, and that is a staffing constraint before it is an interface one.

8 · Applied & close

Where it came from

This came out of designing AI features in healthcare revenue cycle, where a wrong output is a denied claim or a compliance exposure, and where the person reviewing it is often not the person who can tell whether it is right.

The framework says the same thing three times at three scales. Regions, not whole screens. Decisions, not steps. Slices, not workflows. The unit is always smaller than it looks, and that is where the design work is.

It is domain-neutral, and it is applied end to end in the two case studies linked below.

Contracts Intelligence applies both halves of the analogy: first make the map accurate, then assign driving authority decision by decision.

Sources to credit

Public libraries organize on a different axis. Shape of AI sorts by interaction moment. Microsoft's HAX Toolkit by lifecycle stage. Google's People + AI Guidebook by design concern. None sorts by who holds authorship. Their governance patterns land in this framework's supervised zone; their input and tuning patterns land in the assisted zone.

The nearest academic ancestor is the Parasuraman, Sheridan and Wickens model of levels of automation, which varies automation across four stages of cognition. This varies it by consequence and ambiguity of the task instead.