Rules
Part of Problem discovery: steps, examples and decisions for 2027
Problem discovery metrics that show whether progress is real
Problem discovery metrics that can fall rather than only rise: surprise rate, cold sample share and workarounds found, and why no threshold is offered.
Discovery resists measurement, and the usual response is to measure the wrong thing: conversations held, hours spent, notes taken. Those rise whatever happens, which is exactly why they get reported.
The counts below are different. Each one can fall. Each says something about the quality of what you are collecting rather than the quantity, and every one can be produced from a log you should be keeping anyway.
No target numbers appear here, and any source offering you one is guessing. A useful threshold depends on how varied your segment is and how expensive your next commitment is, and nobody who has seen neither can hand you a figure.
What to take away
- Count things that can fall. Conversations held only ever rises, so it measures effort and nothing else.
- The most informative count is the share of conversations that changed your written statement, and it should be falling before you stop.
- Ratios from small samples are noise wearing a percentage sign. Write the fraction, not just the rate.
Counts worth keeping
| What to count | How to get it | What a low number means | What it does not mean |
|---|---|---|---|
| Unprompted problem mentions | Conversations where they described it before you named it | You are explaining the problem into existence | That the problem is not real |
| Workarounds found | Distinct effortful things people already do about it | You are hearing complaints, not costs | That nobody has the problem |
| Existing spend observed | People paying, hiring or subscribing to reduce it | No budget line exists yet | That one could not exist |
| Buyer contact | Conversations with someone who could authorize a purchase | You have not left the sympathetic layer | That the enthusiasm is fake |
| Cold sample share | Fraction reached through a route you could repeat | Your evidence is about your network | That the audience is unreachable |
| Surprise rate | Conversations that changed the written statement | Either saturation, or you stopped listening | That you are finished |
| Trigger recall | People who could name the event that made them look | You do not know when anyone buys | That purchases are random |
| Statement revisions | Dated changes to the problem statement | Nothing has been allowed to contradict you | That the statement is correct |
Existing spend is the row that carries the most weight, and it is the evidence the SBA's guide to market research and competitive analysis treats as the difference between a market and a category.
Keep them as counts first, with the denominator visible. Twelve out of thirty and two out of five are the same percentage, and only one of them is worth acting on.
The two that do most of the work
Surprise rate. For each conversation, record whether it changed anything in the written statement. Early in a segment this is high. It falls as you learn, and when several in a row change nothing, you have saturated that segment.
The trap is reading saturation as market knowledge. You have saturated the people you talked to. If your cold sample share is low, the flatness means you exhausted your own network, which arrives much earlier and looks identical from the inside.
Cold sample share. Split every conversation by how you got it: personal route, or a route anyone could use. Warm conversations teach you about the problem. Only cold ones teach you anything about whether the audience is reachable, and reach is what most plans assume without checking. This is the count that becomes a real question in market validation.
Numbers that look like progress
- Conversations held. Rises with effort and is unrelated to learning. Useful only as a denominator.
- Positive responses. Measures how agreeable your questions were. High here with low workarounds found means the questions are leading.
- Waiting list size. Measures the appeal of a promise. It decays: a name from six months ago is not a person waiting.
- Hours spent. Reliably higher for people avoiding the uncomfortable part of the work.
- Survey completion rate. Tells you the survey was short. Nothing about the answers, which are shaped by wording and order to a degree set out at length in the guidance on writing survey questions.
None of these are lies. They are upstream of anything you can decide with, and they are the ones that feel best to report.
Reading them in pairs
Single counts mislead here. The pairs are informative.
- High positive responses with few workarounds: your questions are doing the work, not the problem. Rewrite them toward past events, as in customer interviews.
- High surprise rate with a low cold sample share: you are learning, but only about people like you.
- Flat surprise rate with no buyer contact: you saturated the users and never met the buyer, which is the most common way discovery ends too early.
- Existing spend observed with no trigger recall: the money is there and you do not yet know when it moves. That is a timing problem, and it decides your channel.
Why no thresholds
A figure like twenty interviews is a way of not saying the real condition, which is that new conversations have stopped changing the statement. In a narrow, uniform segment that arrives quickly. In a varied one it may not arrive at all, and the honest response is to split the segment rather than keep counting.
The same applies to any share you might be offered for how many people should report the problem. It depends entirely on how the sample was drawn, which is the part nobody can see from outside. Produce your own numbers, write the denominator and the sampling route beside them, and compare later rounds against your own earlier ones rather than against a figure from somewhere else.
Turning counts into a decision
Before a round, write down which count you expect to move and what value would make you change course. After the round, compare, and write the comparison down whether or not you liked it.
That is the whole apparatus. It works because the threshold gets set while you are still neutral, which is the same reason the claim by claim loop insists on naming the failing observation before the check runs.
Common questions
What is the smallest set worth tracking?
Three: cold sample share, workarounds found, and whether the statement changed. They fit in a spreadsheet with one row per conversation and answer most of what matters.
Do I need a tool for this?
No. It is a spreadsheet, and it stays one for longer than people expect. The record matters more than what holds it, as the discovery log in problem discovery sets out.
Can I compare my numbers with someone else's?
Not usefully. Their sampling route, segment and question wording differ and are usually undisclosed, so the comparison is between two things measured differently. Compare your round three against your round one.
When do these stop being the right measurements?
The moment you ask someone for money. From then on the evidence is behavior rather than description, and the counts that matter are about acceptance and refusal: see offer testing.
