Case study · AI tooling

Better on average. Still not a release decision.

Holdout, a release gate for AI features. Designed end to end, from the review that finds what an average hides to the launch page.

Scope
Flow, identity, product, launch
Platforms
Web and iOS
Built on
Own system on React Aria
Deliverable
Figma, FigJam and recorded motion
00Overview

40 bad answers, hidden in a better score.

v43 scored 93.4 against 91.2. The average went up while 40 refund conversations got worse, and the person who owns refunds never saw them.

2,000 conversations. 40 of them got worse.

Returns and refunds · 31 break the refund rule

9

Steps to a customer, down from 9

4

Tool, down from four

2,000

Conversations a person reads, not 2,000

0

Screens mapped, web and mobile

01The release as it stood

Four people, four days, one average.

Walked as it really happened. A coral pin wherever a regression slipped through.

Swimlane of how Juno's prompt change reached customers today, with five leaks marked
Drag to explore →
  1. 1The average hides the category
  2. 2The decision is a Slack reply
  3. 3Random checks miss rare failures
  4. 4The policy owner never sees it
  5. 5All or nothing, no way back
02Where it got hard

Five places where two good requirements collided.

None of the hard parts were visual. Each one was decided before a screen was drawn.

01

Iteration speed vs review rigor

Review the difference, not a sample. Only changed conversations reach a person.

02

One headline score vs six categories

A total that cannot hide a loss. Every category has a floor.

03

Model judges vs human judgment

Rules block, judges rank, people decide.

04

Ship the fix vs ship it safely

Release by category. Returns stays on v42.

05

Who owns the ship button

The categories a change touches pick its approvers.

03The release, rebuilt

One afternoon, and the policy owner is in the loop.

Eight steps, one tool. Every fix answers a coral pin.

Swimlane of the same change through Holdout, with Holdout as its own lane
Drag to explore →
  1. 1The score is split by category
  2. 2The decision lives in the product
  3. 3People read the changed, not the random
  4. 4The policy owner is routed in
  5. 5Release by category, gradually

A release is a state, not a button

Nobody approves their own change

04Identity

A release gate that looks like one.

Graphite and bone, lime for Holdout itself, and colour only where a result got better or worse.

An arch and a dot.

The gate, and the one conversation resting on its threshold.

05Explorations

A mark has to survive 16 pixels.

Five of the twenty candidates, each tested where a mark actually lives. Pick one.

E · Doorway

App icon, 96 to 16 pixels

Browser tab

Holdout · Release review×
holdout.dev/changes/v43

Lockup

holdout

E · Doorway

An arch, and one conversation resting on its threshold. The dot is nearly as wide as the doorway, so it is still a gate at 16 pixels.

06Colour and type

Colour is a verdict, never decoration.

Sky means better, coral means worse, lime means Holdout. The badges speak in sentences, not in glowing mono.

07Map and roles

44 screens, mapped before a pixel was drawn.

Drag the boards. Every surface has an owner, and separation of duties is a role, not a setting.

Screen map: 38 web screens in 8 areas and 6 mobile screens
Drag to explore →
Five tensions on a dial, with the decision and its cost
Drag to explore →
08The product

One change, from test to canary.

Every number matches across screens: 2,000 conversations, 212 better, 40 worse, 1,748 the same.

Returns is below its floor, the rest can ship
40 conversations, the most confident mistakes first
The one reply that changed, side by side
Five categories to 5% of traffic
Guards watch live traffic for 24 hours

Close up

Close up of the verdict: Returns is below its floor, next to the score, 212 better, 40 worse and 1,748 the same
Close up of the v43 reply with the refund promise underlined
Close up of the policy rules: refund_rule with 31 fails

Approve from a phone. Roll back from anywhere.

Releases need the right approvals. Rolling back never does, and the log records who did it.

The support lead reads the one reply that changed

Five categories to 5%, Returns stays on v42

The release, screen by screen

Canary monitor
Review queue
Release dialog
Conversation diff
Release review
09Motion

Recorded frame by frame, not animated by hand.

The key moments were rebuilt from the Figma screens and recorded at 60 fps, so the motion and the designs never disagree.

0

Recordings, web and mobile

0

Frames a second, every one real

0

Screens designed in full

0

Conversations in every number

10Launch page

A launch page built from the product.

Real screens, the dot field, and no borrowed logos. Keep scrolling.

holdout.dev
Holdout landing page
11What we test next

Three assumptions that need evidence.

On the designs, with support leads and AI engineers, with tasks rather than opinions.

Assumption 1

The review load

Can a support lead keep up with 40 conversations on a busy release day?

Assumption 2

The category floors

Do teams set floors honestly, or lower them the first time a release is held?

Assumption 3

The 24 hours

Is 5% of traffic for a day enough to catch a failure that shows up once in 2,000?

Is an average hiding something in your product?

We find the decision your screen should be making, and hand it back designed in five days.