Evidence, Not Encouragement
Chapter 2 · 14 min read
On this page
Start here
Twelve interviews. Three weeks of scheduling, an hour each, notes written up the same evening. The summary doc runs to four pages and its conclusion is one line: strong validation — recommend we build.
Read the highlights. "This is exactly the kind of thing we've been looking for." "Honestly, if you had this today I'd use it." "You're solving a real problem." Nine of twelve rated the concept 4 or 5 out of 5.
The team is proud of this, and they should be. Twelve strangers gave up an hour each. That is genuinely hard to arrange and the notes are careful.
Now try to answer one question from the document: what does any of these twelve people currently do about this problem? Not what they said about your idea — what they do on a Tuesday, right now, before you existed. What they spend on it. What spreadsheet they built. Who they email.
The document does not say. Not because the team hid it, but because nobody asked, and nobody asked because every question was about the idea.
So the question this chapter answers is not "did we do enough research." Three weeks is plenty. It is: what kind of thing did these twelve conversations produce, and is it the kind that can support a decision?
Core concepts
Evidence and encouragement
Sit in on any customer conversation where the interviewer describes their idea, and watch the other person's face. What comes back is almost always warm. Not because they are lying — because a polite adult, told about something a hopeful person built, produces encouragement. It is what the situation calls for.
A customer conversation produces one of two things: facts about what the person has already done, or opinions about what they might do. Only the first is evidence.
Compliments, hypothetical enthusiasm, and "I'd definitely use that" are social output. They are what a polite person emits when you describe an idea, and they correlate with nothing — not with future purchase, not with future usage, not even with the person's own honest private assessment.
How to use it. Interrogate any discovery write-up — the summary of research done before building — for three things:
- Past behaviour, not predicted behaviour. Does it report what happened, or what someone says would happen?
- Specifics, not generalities. Dates, amounts, tools currently paid for, workarounds already built. "It's frustrating" is a generality. "She rebuilds the report by hand every Monday morning, takes about two hours" is a specific.
- Something real given up. Time, an introduction, money, a scheduled follow-up with a colleague. Or only words.
If a discovery summary contains no facts and no commitments, it validated nothing — regardless of how many conversations it reports. Twelve of nothing is nothing.
How to spot it. Read only the quotes. If they are about your product, the conversation drifted to you and away from the person's life. If they are about spreadsheets, contractors, current spend, and things that broke last month, the conversation stayed where it belonged.
One thing that has changed. Discovery synthesis is now routinely machine-summarised, and summarisation strips exactly the texture that distinguishes evidence from encouragement. A model summarising twelve interviews will happily render "several users expressed strong interest" out of twelve polite deflections, and the sentence will be accurate as a summary and useless as evidence. The move is to demand the underlying quotes and behaviours, not the synthesis.
The chain from outcome to experiment
A team has decided to build in-app messaging. Ask why and you get a fluent answer: users want better communication. Ask what else was considered and the room goes quiet, because in-app messaging was not chosen from a set. It arrived, and the reason was assembled afterwards.
There is a structure that makes this visible. It connects four levels:
- the outcome you want to move (a behaviour change, with a number),
- the opportunities that could move it (customer needs, pains, desires — the things people actually experience),
- the solutions that could address each opportunity,
- the experiments that test each solution.
Its value is not tidiness. It is that it makes the reasoning chain from "what we want to change" to "what we're building" explicit, and therefore checkable.
How to use it. Trace any proposal in two directions. Upward: which opportunity does this serve, and which outcome does that opportunity move? A solution that cannot be traced up is unmoored — it may still be a good idea, but nobody can say why. Sideways: what other solutions to that same opportunity were considered? A tree with exactly one branch usually means the solution came first and the opportunity was reverse-engineered to justify it.
How to spot it. Ask the sideways question out loud. The tell is not a bad answer, it is a surprised one.
One thing that has changed. Generating the alternative branches is now nearly free, which removes the last excuse for a single-option proposal — and floods the tree with plausible options nobody has any evidence for. The scarce work moved from generating branches to pruning them against real opportunity evidence.
The riskiest assumption
Every plan rests on things that must be true. Most plans never list them, which means nobody can tell which one, if false, takes the whole thing down.
The discipline: before building, enumerate the assumptions the solution depends on. Plot each on two axes — how much it matters if we're wrong, and how little we know. Test the one in the top-right corner first, with the smallest test that could actually change your mind.
How to use it. Write the list. For each item, mark it evidenced or assumed. Then find the one that is both load-bearing and unevidenced: that is this week's work. Then apply the second test, which is the one people skip — could the proposed test actually return "no"? A test with no failing outcome is theatre.
How to spot it. Look at the plan's risks section. If every entry is an execution risk — timeline may slip, the vendor may be late — and none is a belief risk, the plan has not been examined. It has been scheduled.
One thing that has changed. Cheap building makes "just build it and see" superficially rational, and for genuinely small bets it now often is. The distinction that matters is whether the learning is cheap, not whether the building is. A week-long build that produces no interpretable signal is still a wasted week, and now you also have a codebase to maintain.
Back to the twelve interviews
The twelve conversations produced encouragement. That is not a small failure to be fixed with more interviews — and this is the part most teams get wrong. Doing twelve more of the same conversations produces twelve more compliments, because the questions were forward-looking, and forward-looking questions have only one kind of answer.
The fix is upstream of volume. It is in what gets asked.
Worked examples
The conversation that changed shape
Sana is building software for independent physiotherapy clinics. Her first eight interviews all went well and taught her nothing, which she only realised when she tried to write the summary and found she had no sentence to write.
For the ninth, she changed one thing: she did not describe the product at all.
She opened with "walk me through the last time a patient no-showed." What came back was fifteen minutes of detail she could not have invented. The receptionist keeps a paper list because the practice-management system's reminder feature sends at a fixed time and half the patients work shifts. Someone rebuilt the recall list in a spreadsheet two years ago and it is still the thing the clinic actually runs on. They pay £40 a month for a text-message service they use for exactly one purpose, having abandoned the other four.
None of that is about Sana's idea. All of it is evidence — a paid-for workaround, a rebuilt tool, an abandoned feature set. It is a description of a problem expensive enough that someone already spent money and effort on it.
She left with one more thing: the receptionist offered to send her the spreadsheet. That is a commitment. It cost something to give, which is what makes it worth more than the eight enthusiastic hours before it.
The uncomfortable part: interviews one through eight were not badly executed. They were well-run conversations about the wrong subject.
What the summary removed
Marcus runs product at a company selling inventory tools to small retailers. His team ran fourteen discovery calls, recorded and transcribed them, and had a model produce the synthesis. It is a good document — themes, frequencies, representative quotes, a clear recommendation.
Theme two reads: "Users find the reorder workflow confusing (mentioned in 9/14 interviews)."
Marcus asks for the nine underlying passages. This takes twenty minutes and changes the finding entirely.
Six of the nine are not about the workflow being confusing. They are about the same thing: the reorder screen shows warehouse stock, but four of those six retailers hold stock in the shop and the back room and count them separately, so the number on screen is never the number they need. They are not confused. They are looking at a figure that is wrong for their business, and they have each built a private correction for it.
The other three are genuinely about confusion, and they are new users.
The synthesis was not inaccurate. "Confusing" is a fair one-word summary of nine people expressing difficulty. But the compression merged two different problems with two different fixes and one different audience, and the recommendation that followed — a workflow redesign — would have solved the smaller one.
The general form: summarisation removes specificity, and specificity is the thing that made the evidence evidence. The step from fourteen interviews to "users want X" is where the information was lost. It used to be a person doing that step; now it happens automatically and faster, which means it happens more often and gets questioned less.
The survey that answered a different question
A team wants to know whether to add a compliance-reporting module. They run a survey: 340 responses, 71% say they would find it valuable, 44% say they would pay extra for it.
The number is real. The survey was competently run. And it cannot support the decision, for a reason that has nothing to do with sample size.
Every respondent was asked a question about the future — would you value, would you pay — and answers to those questions are cheap to give and cost nothing to be wrong about. Nobody in the sample has yet had this problem cost them anything in a way that this module would have prevented.
The convert-it-backwards test: what would the same question look like pointed at the past? In the last twelve months, has a compliance requirement caused you work? What did you do about it? What did that cost you? Do you pay anyone for it today? Those questions have expensive answers, and expensive answers are informative.
Run that version and the sample shrinks — most respondents have no story, which is itself the finding. What remains is a smaller number of people who describe a real recurring cost, and they are the ones whose willingness to pay would mean something.
Surveys are hypothesis generators. They are very good at telling you what to go and investigate. They are not validation, and 340 of a bad question is not better than 12 of a good one.
What great operators do
Ask about their past, never about your idea. People are reliable reporters of what they did and unreliable forecasters of what they will do, especially about a product whose creator is in the room. Past behaviour is data; predicted behaviour is politeness. Tell: present — "walk me through the last time this happened." Absent — "would you use something that…", "how much would you pay for…".
Treat a compliment as a signal that the conversation failed. A compliment means the conversation drifted to your idea and away from their life. Strong operators visibly deflect praise and steer back to specifics. Tell: a write-up whose highlights are enthusiasm quotes has no findings. One whose highlights are workflows, spreadsheets, and existing spend has findings.
Look for what they already pay for or hack around. Existing spend and existing workarounds are revealed preference — someone already paid a cost to solve this. It is the strongest cheap evidence that a problem is real and ranked. Tell: the artifact names current tools, spreadsheets, contractors, or manual processes and what they cost. Absent: the problem described only in adjectives.
Push for a commitment, not a next meeting. Interest that survives a request for time, money, reputation, or an introduction is real; interest that evaporates was never there. Advancing the relationship is the test. Tell: "they agreed to pilot it with their team next month" beats "they were very interested" by an enormous margin. "Let's reconnect in Q3" is a polite no.
Demand the raw material behind any synthesis. The move from twelve interviews to "users want faster onboarding" is where the information was lost, and it now happens automatically. Tell: the write-up links to or quotes actual conversations and behaviours. A findings document with no traceable underlying observation is an opinion citing itself.
Common failure patterns
Compliment laundering
What it looks like. A discovery or research doc whose evidence is enthusiasm: "users loved the concept," "strong positive signal from 9 of 12 interviews." No past behaviour, no existing spend, no commitments.
Why smart people do it. The conversations genuinely happened and genuinely felt good. Effort was expended, so it feels like evidence was gathered. And nobody enjoys re-reading twelve interviews in order to discover they learned nothing.
The correction. Re-derive the finding from behaviour only, discarding every opinion. If nothing survives, the honest write-up is "we have not yet validated demand" — which is itself a finding, and a cheap one compared to finding out after the build.
Hypothetical validation
What it looks like. Demand established by asking about the future: survey results on willingness to pay, "would you use this," feature-request rankings collected from people who have never had the problem cost them anything.
Why smart people do it. It scales, it produces charts, and it arrives fast. Real behavioural evidence is slow and small-n, which looks weaker in a deck — a sample of nine feels embarrassing next to 340, even when the nine are the only ones who know anything.
The correction. Convert every forward-looking question into a backward-looking one and see how much of the case survives. Treat surveys as hypothesis generators, never as validation.
Premature automation
What it looks like. Automating onboarding, support, or research before anyone understands the process, justified by efficiency. Increasingly: replacing customer conversations with generated summaries at exactly the point where the team was about to learn something.
Why smart people do it. Manual work feels like a failure of engineering, and automation is satisfying to build. It also removes the emotionally uncomfortable part of early-stage work — talking to people who don't want your product yet.
The correction. Automate what is understood and repetitive; protect direct contact longest. Ask what signal the automation removes, not only what labour it saves. The temptation has inverted in the last few years: it is now easier to automate customer contact than to do it, so teams skip the manual phase by default and buy fluent throughput at the cost of the only signal that mattered.
Make the call
Twelve interviews, zero findings is the checkpoint: the discovery synthesis from Start here, written out in full — enthusiasm quotes, a themes section, a strong-validation conclusion, a recommendation to build. Your task is to say what the research actually established.
It exercises the whole chapter, and the trap is specific. The obvious critique is "the sample is too small" or "do more interviews," and both are wrong. Twelve is plenty. Read to the end of the document, where the interview guide is: the questions were forward-looking, so more of them produces more of the same. Saying that to a team proud of three weeks of real work is the socially expensive part, and doing it without being dismissive is most of the skill.
Honest difficulty note. This is a 1, the gentlest scenario in the library, and it is the right place to start if this is your first. The compliment quotes are easy to find. What separates a good answer from an adequate one is what you recommend next.