In 2006, Greg Guest and two colleagues coded sixty interviews from a West African health study and noticed the new themes stopped coming around interview twelve. The number stuck. A 2022 review put the range at nine to seventeen. Twenty years of textbooks and procurement templates settled it: qual saturates early, keep it small.
We wanted to see what that settlement costs. So we ran a real brief, an airline segmentation called Traveler Under Pressure: fixed script, screener, ratings battery, 495 AI-moderated interviews with US travelers. Then the part nobody budgets for. We coded every transcript. All of them, 7,686 verified spans, and suddenly the corpus is an answer key.
Because once you know what all 495 said, you can deal out two hundred smaller studies at every size and grade each one: what would this study have seen?
What each sample size buys
● each dot below the axis: one of the 20 traveler findings, at the n where a study is 95% likely to meet it. The pencil deepens as the study grows n=12 → 495. Hover any dot for the finding.
Draw the study, read both sides
Every cell is one of the 495 interviews (the 33 dashed cells failed the usability bar and can’t be drawn). Pick a sample size and draw. You get two sides of the same study: the front stage, the report a team would actually ship, proto-personas and themes and quotes at full confidence. And the backstage, what only the full corpus can see: the denominators, the sizing error, the segmentation agreement, the voice your draw never met. It opens at n=10, the study at its most confident and most wrong.
one cell per interview · hover for who · filled cells are in the current draw · fill color = the draw’s n on the pencil scale
Draw at 10 a few times and watch the front stage stay perfectly confident while the backstage swings. Then climb to 145 and feel everything settle, except the tail, and except the personas. That’s the ladder, in your hands.
“I'm not worried about the time or the money, I just need to know... don't keep telling me it's 20 minutes away when you don't have a part or you don't have the airplane”
He wasn’t asking for much. Just to know. Ninety-five of the 462 were asking with him, and it took all four hundred ninety-five interviews to count them.
Next corpus, new constants. Let’s see.
For the auditors. Skippable, checkable: every session scored, the full instrument verbatim, and who we talked to.
Every interview, scored
We scored the moderator too: question loops, false rejections of valid answers, echo-misquotes, seven distinct ways of mishandling a 1-to-7 scale. Every failure class is a fixable behavior, which makes the 2.8 a roadmap, not a confession. And still, 93% of sessions produced usable data, partly because participants ran the repairs themselves.
One cell per coded interview (495). Cell color is participant signal; the corner wedge is the AI moderator's score for that session; dashed outlines are the 33 sessions excluded from analysis. Study level: mean IQS 3.67 (Signal 3.79 · Moderator 2.8 · Usability 3.96). Hover any cell for its decomposition.
What was asked
The pre-registered engine design, then the instrument’s guide, verbatim.
Who we talked to
Age
Screener price identity (Q7)
Airline type (Q5)
Flights in past 2 years
Trip type
Top states
Built with AI + analyzed with human care. The Observatory x Dialogue AI, July 2026.