Design is judged by what moved, not what shipped

Every design review I sat in for the first decade of my career ended the same way. Two smart people disagreed about a screen, the more senior one won, and everybody went back to work not knowing whether the product had got better. We shipped a lot. We could not say what any of it did.

When I took over a design org inside a payments business, I decided the team would be judged by what moved, not by what shipped. This is what that took, and what it changed.

The one sentence that started it

Use metrics count actions. Outcome metrics verify that the action left the user better off.

That distinction sounds small. It reorganised everything. A conversion rate counts how many people finished a flow. It cannot tell a correct completion from a mistaken one. We had a real case of this: an identity verification flow was redesigned, conversion went up, and the team celebrated. Within a month the rejections and cancellations had climbed too, because people were completing the flow with the wrong documents. Conversion had counted every one of those as a success.

Breaking the funnel into finer steps does not fix that. A finer funnel still measures use. A page can convert perfectly into a decision the customer regrets. The fix is not more granularity. It is a change of role: the number that judges the work has to be derived from the customer's life, written down before the build, and the conversion rate becomes something you watch rather than something you are graded on.

The same problem hides in every metric a design team is usually handed. A satisfaction score measures what someone says at the moment of asking, about the whole relationship, including the pricing and the mood they are in. A click rate cannot distinguish understood and proceeded from confused and clicked anyway. Task success looks like a UX metric, but as normally used it only counts completions. It only becomes an outcome metric when the definition includes correctness and independence: completed correctly, unaided, and needing no help afterwards.

We kept every one of those numbers. We stopped letting any of them be the judge.

Writing an outcome that can fail

The tool we replaced them with is a sentence. Before any significant piece of work starts, the team writes one outcome statement to a fixed template:

If we do a great job on this, we will improve a named person's life by an observable change, bounded in time and context, even when a hard situation occurs. So that specific miseries stop happening.

Every clause is doing a job.

The named person is not a persona. It is a real customer the team has met, in the field or in a conversation, whose evidence is compressed into one name. If you cannot picture him, you have not done the research yet.

The observable change is the part most drafts get wrong. Understands the page easily is not observable. Confirms in under a minute, on his phone, that every rupee he collected today has either arrived or has a stated reason and date is observable. Four questions test each draft: could I stand next to him and verify it, could it fail, is it bounded, and does it belong to his life rather than to our product.

The even when clause names the hard situation the experience must survive. This is where reputations are made, because the situation it names is exactly the one a team under deadline would quietly exclude. One alert per event works fine in the demo. The customer with five events in one evening gets spammed and switches notifications off. If the team would be relieved to descope a situation, that is the edge case. Write it in.

A statement written this way can be wrong. That is the point. A goal that cannot fail is not a goal, it is a wish.

Three kinds of number, and one party

From that sentence, every metric follows. There are three kinds.

Success metrics are the finish line, defined at the start: the precise, observable moment the outcome is achieved. Progress metrics track the climb towards it, as a trajectory published on a rhythm, as milestones, and as firsts, the human-scale moments that prove new ground before any percentage has moved. Problem-value metrics price what it costs to leave the problem unsolved, in money, every month, so that priority is set by cost rather than by who argues loudest.

The threshold for success is the part people push back on. Why 97 percent and not another number? The answer is floor, ceiling and clock. The floor is the instrumented baseline today. The ceiling is the irreducible failure rate, the share of cases no design could prevent. The clock is whether the climb is achievable in the time available with the levers you actually control. A threshold published without that derivation is a number. With it, it is an argument, and arguments can be examined.

And every success metric gets a party moment: one concrete sentence stating what we will celebrate. We celebrate when a small merchant closes his day in under a minute and the support tickets from people like him have halved. If the team cannot write that sentence, the metric is not defined yet. It is also how leadership prefers to read progress: as outcomes delivered, not percentages moved.

When the business overrides the experience

This is the objection I hear most: what happens when a decision above the team is final and it makes the experience worse? A pricing detail disclosed in the contract rather than on the screen. A feature the sales team promised.

The framework has an answer, in four steps, in order. Translate: write the customer's outcome under the business decision, because most conflicts dissolve here, and the better path for the customer is usually also the cheaper route to the business number. Price: if the decision can only be met by making the customer's life worse, estimate what the harm will cost in abandonment, support contacts and lost trust, and present two routes to the same business number with their costs. Log: if the decision stands, record it at decision time with its estimated cost and the guardrail metrics that will detect the harm. Guard: pre-register those detectors, so the result is settled later by published numbers rather than relitigated opinion.

Neither the business outcome nor the customer outcome outranks the other. The business sets the destination. The customer outcome is the causal route. When they conflict, the decision is priced in writing and made with the price visible. Once a team can do this, it stops being a service desk. It is questioning business outcomes with evidence, and that is when its influence starts to grow.

Exposure, or why decision-makers must see customers

None of the above works if the people making decisions have never watched a customer. Organisations naturally build insulation between deciders and users. Exposure is how you break it.

The rule I hold the team to is small and non-negotiable: twenty minutes a week, every week, watching a real customer use the product, a prototype, a competitor, or doing the job with no product at all. The last one is the highest-value gap you will ever find, watching someone reconcile with a bank SMS and a notebook. One session fades in about six weeks, so it has to repeat, and the invitation has to stay open to product, engineering and leadership. A decision-maker who has stood next to the customer argues differently in the next review.

What changed

Design reviews got shorter and quieter. Decisions are checked against a pre-agreed outcome and its threshold rather than against opinion, so the loudest voice stopped mattering. Design's contribution became traceable, because every outcome links to a business goal, and its effect on revenue, adoption and support cost can be shown rather than asserted. Effort moved to where it pays, because unsolved problems carry a price. Results stayed honest, because thresholds agreed before the results exist cannot be quietly changed after them, in either direction. And the organisation got faster, because teams stopped shipping outputs that moved no outcome and stopped relitigating decisions the data had already settled.

There is a cost. It takes discipline to write a sentence that can fail, and courage to publish a number before you know whether you will hit it. Some of ours did not move for months while the foundations were built, and we reported milestones and firsts instead of percentages and said so. A team does not owe a delta every cycle. It owes honesty about the ceiling.

The mechanisms around it

Measurement is one of six mechanisms that let a design team deliver outcomes rather than screens. The others are unglamorous and I have written about some of them before: one agreed path from request to launch with quality gates on the drafting table rather than the construction site, a quarterly intake so work is planned rather than ad hoc, a weekly critique so the quality bar is decided together rather than per project, a biweekly note to the whole organisation so nobody has to ask what design is doing, and written expectations for every level so people know what growth looks like. Plan, execute, communicate. Each one turned something unpredictable into something predictable.

But the measurement is the one that changed how the team is treated. When you can say what moved, you get invited to the conversation about what should move next. That is the whole difference between a design team that is consulted and one that is in the room.