LearnProduct Judgment

The Launch Call

Chapter 5 · 17 min read

On this page

Start here

The beta has been running six weeks. Launch is in six days. Marketing has been briefed, the blog post is drafted, and two people have rearranged their holidays around the date.

The dashboard says: 40 users enrolled, 62% completed the new flow, feedback mixed but mostly warm, one power user unhappy and vocal about it.

Is 62% good?

You do not know, and — this is the part that matters — you cannot find out from this data. With 40 users, a true rate of 50% and a true rate of 75% both produce something around 62% often enough that the number cannot tell them apart. The dashboard is not giving you a weak signal. It is giving you a number that looks like a signal.

Which puts you somewhere specific and uncomfortable. This is not a case where more analysis helps, and it is not a case where the data is bad — the beta ran fine, the number is correctly computed. It is a case where the test as designed could not have answered the question, and you are finding that out six days before launch.

So the question is not "ship or delay." It is: what do you do when the information you needed was never going to arrive, and how do you avoid arriving here again?

Core concepts

A test that cannot fail is not a test

Read the plan the beta came from. It probably said something close to this: six weeks, 200 users, we'll learn and iterate. Success: gather qualitative and quantitative insight into how users experience the new flow.

Every word of that is reasonable and the whole thing is unfalsifiable. There is no result it could have produced that anyone would call failure. "We gathered insight" is true whatever happens, including if nothing happens.

A test with no losing outcome cannot produce information. That is not a stylistic complaint — it is what the word test means. If both branches lead to the same action, you did not run a test, you built something with extra steps and a research vocabulary attached.

How to use it. Before running anything, write down: what result would make us stop? Put it in the document, before the data exists. "If fewer than 20% of the pilot group complete the flow, we stop" is a test. "We'll evaluate the results" is not.

Then check three things about the threshold:

  1. Is there a baseline? A threshold with nothing to compare against gets reinterpreted after the fact. What is the current rate?
  2. Can the sample distinguish the outcomes? If the threshold is 20% and the sample cannot tell 15% from 25%, the threshold is decorative.
  3. Did the comparison get written down first? Stating it afterwards is unfalsifiable by construction.

How to spot it. Success described in verbs about learning — learn, iterate, gather, understand, validate. Those are what a plan says when nobody wants to name a number.

One thing that has changed. "Just build it and see" is now cheap enough to be superficially rational, and for small bets it genuinely often is. The distinction that matters is whether the learning is cheap, not whether the building is. A week-long build that produces no interpretable signal is still a wasted week, and now you also own a codebase.

What would change my mind

There is a question that costs nothing to ask and is almost never in the document: what would I have to see to reverse this?

Committing in advance to a disconfirming observation is the only cheap defence against interpreting every result as support — and every result will be interpreted as support, by competent honest people, because that is how reading works when you already have a conclusion.

How to use it. Write the answer down, in the document, before the data. It does two things. It makes the recommendation falsifiable, and it tells your team what evidence to go and find, which is usually more valuable than the discipline itself.

How to spot the gap. A confident recommendation with no named disconfirming observation. That combination is the reliable signature of a conclusion that preceded its analysis — not a lie, just an order of operations that got reversed somewhere and never got noticed.

Which risk are you shipping against

Chapter one's four risks come back here under time pressure, and the pressure changes how they behave.

Value, usability, feasibility, business viability. A launch decision is a decision to stop testing some of them. That is legitimate — you cannot test everything — but it should be chosen rather than defaulted into.

How to use it. Name which risks the beta actually addressed, and say the rest out loud. "This de-risked usability; value is untested and we are shipping on the assumption that the enterprise segment behaves like the beta group" is a sentence that lets a reader agree or disagree with something real.

Under time pressure, feasibility gets over-reported, because it is the one that resolves cleanly and produces confident sentences. Business viability gets under-reported, because its costs — support load, brand, legal, margin — arrive after the launch and belong to other teams.

How to spot it. A launch memo with a risks table that lists only delivery risks. Documentation slipping, volume spiking, support briefed. Every row is about whether the launch happens, none about whether it should.

One thing that has changed. When shipping is cheap, more decisions become reversible, and reversibility genuinely should change how much evidence you demand — demanding the same rigour for a reversible bet as an irreversible one is its own error. What does not become reversible: the promise made to customers, the support surface created, and the thing your positioning now claims.

What "ship a subset" actually buys

The middle option in almost every launch call is to ship to some but not all. It is usually right, and it is usually right for a reason nobody states.

Shipping to a subset buys you a second sample, under real conditions, at a size you choose — but only if you decide in advance what that sample is supposed to resolve, and what you will do with each answer.

Shipping to a subset with no threshold and no decision date is not a cautious launch. It is a full launch with a smaller denominator and a delay attached, and it feels responsible, which is why it is the most common way this decision gets made badly.

How to use it. If you ship a subset, write three things: which segment and why that one, what number decides, and the date the decision gets made. Absent the third, the subset becomes permanent by drift.

Back to the 62%

The honest read: the beta established that the flow works mechanically and that some people are fine with it. It did not establish the completion rate, because it could not.

That leaves three defensible calls, and — this is the point — they are defensible. Ship, with the untested risk named. Delay, if the cost of being wrong is high enough to buy a real sample. Ship to a subset, with a threshold and a date.

What is not defensible is treating 62% as though it meant something. The failure here is not the choice. It is the plan, six weeks ago, that produced a number nobody could act on.

Worked examples

The threshold that arrived too late

A team runs a two-week experiment on a new pricing page. Result: conversion 3.4% against 3.1% for the old page.

In the review, the argument is about whether 3.4% versus 3.1% is meaningful at their traffic. It is not — the difference is well inside the noise for that sample — but the argument runs for half an hour anyway, because both sides are arguing about interpretation and there is nothing to appeal to.

Someone eventually asks what the plan said the threshold was. There was no plan with a threshold; the experiment was set up in an afternoon.

Here is the trap: the honest answer now is "we don't know," but nobody can say it without appearing to advocate for the old page. Because there is no pre-registered comparison, "inconclusive" reads as a position rather than a finding, and the person saying it is treated as the opposition.

That is the real cost of a missing threshold, and it is a political cost rather than an analytical one. It is why the threshold has to be written before, when nobody has anything at stake in it.

What they do next is right: they define the threshold and the required sample for the rerun, and they compute up front that the rerun needs five weeks at current traffic. That number is itself a finding — it tells them this page is not worth testing at this traffic level, which is more useful than either result.

The pilot that could only succeed

A six-week onboarding pilot: 200 users, described as "learn and iterate," success defined as gathering insight. Approved by everyone.

A reviewer redlines it before it starts, and the redline is short.

What is the current completion rate? Nobody knows. First fix: measure the baseline for one week before the pilot begins. Without it there is nothing to compare against, and any post-hoc comparison will use whatever prior period makes the result look right.

What result would make us stop? Silence. Second fix: write it down. The team lands on: if completion does not improve by at least 8 points over baseline, the approach is wrong and we revert rather than iterate.

Can 200 users distinguish an 8-point improvement? This is the question that usually gets skipped. Roughly: yes, if the baseline is near 50% — an 8-point move is detectable at that size. So the threshold is honest.

Who decides, and when? Third fix: a named person and a date, before the result exists, because a threshold with no decision date gets extended.

Total cost of the redline: one week of baseline measurement and about an hour of argument. What it buys is a pilot that can return "no" — which is the only thing that makes the six weeks worth spending.

The reason this is hard has nothing to do with knowing the technique. Committing to a kill threshold means sometimes having to kill your own project in public, and leaving it open preserves the option to be right. Everyone in the room understands this, which is why nobody proposes the threshold.

The launch that was right to ship and wrong to claim

A company ships a feature to all customers after a beta with 30 accounts. The right call, plausibly: the feature is additive, reversible, and the beta surfaced no serious problems.

The launch memo says the beta "validated strong demand."

That sentence is the error, and it is worth separating from the decision, which was fine. Thirty accounts that opted into a beta are the most enthusiastic users of the product — a sample selected for enthusiasm cannot establish demand in the general population, and everyone involved knows this if asked directly.

The damage is not immediate. It arrives one quarter later, when adoption comes in at a fraction of the beta rate and the team treats it as a mystery to be investigated. Two weeks go into diagnosing a discrepancy that was created by the sentence in the memo, not by the product.

The version that costs nothing and prevents all of it: "the beta established that the feature works and that enthusiastic users like it. Demand in the general base is untested. We are shipping because it is cheap and reversible, not because it is validated."

Same decision. Same launch date. A team that is not confused in three months.

Case studies

The beta whose data cannot distinguish the outcomes

An invented composite, built from the pattern rather than any real company.

The situation. A project-management product for construction firms. The team has rebuilt the flow where a site manager creates their first project — the point where new accounts have historically stalled.

Six-week beta. 40 accounts, opted in. 62% completed the new flow within a week of signup. Historical baseline for the old flow is "roughly half," measured on a different definition eighteen months ago and never re-derived.

GA is in six days. Marketing is briefed, a partner announcement is coordinated, two engineers have holidays booked on the assumption this ships.

Qualitative feedback: mostly positive. One power user, who runs eleven concurrent projects, is loudly unhappy — the new flow assumes a single project and makes his setup slower.

What the data can and cannot say. With 40 users, the 95% interval around 62% runs roughly from 46% to 76%. The old flow's "roughly half" sits inside that interval. So the beta cannot establish that the new flow is better than the old one. It also cannot establish that it is worse.

And the baseline is worse than useless: measured on a different definition means the comparison is void even before the sample size problem, which is chapter three's problem arriving inside chapter five's.

What the beta did establish: the flow works mechanically at small scale, and there is a segment — multi-project users — for whom it is a regression. That second finding is real, is not a sample-size question, and is the most solid thing in the data.

Three defensible resolutions.

Ship on the date. The argument: the flow is not worse, the mechanical validation is real, the delay costs a coordinated partner announcement and six days of engineering idle, and the change is reversible. The memo must say the completion improvement is unestablished and that shipping is a bet on it being at least neutral. Add the multi-project regression as a known issue with an owner and a date, because that one is not a sample-size question. The cost: if the new flow is worse, it is worse for every new account until somebody notices, and nobody is measuring in a way that would notice quickly.

Delay four weeks. The argument: this is the first-run experience for every future customer, which makes it high-leverage and expensive to get wrong. Four weeks at current signup volume gets to roughly 160 accounts, enough to distinguish a 10-point difference. The cost: the partner announcement is real and moving it has a relationship cost. Six days of engineering idle becomes four weeks of context-switching. And "we delayed for data" is a precedent that gets invoked later by whoever wants to block something. The question that decides it: what does being wrong for a quarter actually cost? If new-account volume is 40 a month, the answer is small and delay is over-cautious. If it is 400, delay is obviously right. Nobody in the room has said that number out loud, and it is the number the decision turns on.

Ship to a subset. Route single-project accounts to the new flow, keep multi-project on the old one. This is not a compromise — it directly addresses the one solid finding in the data, and it turns GA into the larger sample the beta could not provide. Requirements, or it becomes the worst option: the segment split has to be defined by a rule, not a judgment call; the threshold and the decision date go in the document now; someone owns the date. The cost: two flows to maintain, which is real engineering debt, and a permanent temptation to leave it that way because it is working.

What the case is actually about. All three are defensible and the choice depends on facts not in the memo — new-account volume, reversibility cost, whether the partner announcement can move. A strong response says which fact it needs and what it would decide under each value, rather than picking one and defending it.

The deeper failure is upstream. The beta was designed six weeks ago with no baseline, no threshold, and a sample chosen by who volunteered. Every option now is worse than the option that existed then, which was to spend one week measuring the baseline properly and to compute the sample size before starting.

Where a good decision could still go wrong. The subset split becomes permanent because it works and nobody wants to own the migration. The four-week delay produces 160 accounts that are still self-selected, so the sample is bigger and biased in exactly the same direction. And the "known issue" for multi-project users gets an owner and no date, which is how known issues become permanent.

What great operators do

Design the test so it can fail. A test with no losing outcome cannot produce information. Real operators state the kill condition before running it, because stating it afterwards is unfalsifiable. Tell: "if fewer than 20% of the pilot group complete the flow, we stop," written before the pilot. Its absence means the result will be read favourably whatever it is.

Ask "what would change my mind?" and answer it in writing. Committing in advance to a disconfirming observation is the only cheap defence against reading every result as support. It also tells the team what evidence to go and find. Tell: the document names the observation that would reverse the decision. Its absence beside a confident recommendation is the signature of a conclusion that preceded its analysis.

Name which of the four risks you actually tested. Most proposals conflate "we scoped it" with "we validated it." Naming the risk addressed exposes the ones that weren't. Tell: strong artifacts say "this de-risks value; usability and business viability are untested." Weak ones present an engineering estimate as validation.

Common failure patterns

Unfalsifiable pilot

What it looks like. A test with no stated failure condition, or thresholds set after the results arrive. Often paired with a sample too small to distinguish any outcome from noise, and a plan to "learn and iterate" regardless.

Why smart people do it. Pre-committing to a kill threshold means sometimes having to kill your own project in public. Leaving it open preserves the option to be right, and nobody in the approval meeting has an incentive to close it.

The correction. Write the kill condition before the test and put it in the document. If nobody will agree to a threshold, the organisation is not running a test; it is building with extra steps.

Fluency mistaken for rigour

What it looks like. A well-structured, confidently-worded launch memo — appropriate headings, a risks table, a phased plan — in which no claim is sourced and no number is traceable. It reads like the good version of this document.

Why smart people do it. Format was a real signal for decades. Producing a polished strategy document used to require the thinking that justified it. That correlation broke; the reading habit didn't.

The correction. Read for traceability rather than structure. Pick the three load-bearing claims and ask where each came from. Treat "it reads well" as carrying no information about whether the analysis happened.

Make the call

Two checkpoints, and they test opposite ends of the chapter — one is the decision under pressure, the other is the plan that created the pressure.

Ship it on forty users? is the situation from Start here, handed over with six days to launch and marketing already briefed. You commit to a call — ship, delay, or ship to a subset — and then defend it in a short justification. All three are defensible; the score lives entirely in the reasoning, and specifically in whether you notice that the sample cannot answer the question it is being asked, and that the baseline it is compared against does not survive one look.

The pilot with no way to fail is the plan that produced it: six weeks, 200 users, no threshold, no baseline, no kill condition, approved by three teams. You redline it so it could actually return a negative result. The trap is that adding a threshold is the obvious move and is not sufficient — without a baseline and a pre-registered comparison, the threshold gets reinterpreted in the readout.

Honest difficulty note. Both are 2s. Do them in that order: making the call badly is the experience that makes the redline feel urgent rather than procedural.