Dialogue AI × The Observatory · a study of studies

The Twelfth Interview

Qualitative research decided long ago that saturation arrives early. We coded 495 interviews to find out what arrives after.

the story ↓what n buysdraw your own study
May–July 2026 · 495 AI-moderated interviews · 7,686 verified codes · 852 verified quotes

In 2006, Greg Guest and two colleagues coded sixty interviews from a West African health study and noticed the new themes stopped coming around interview twelve. The number stuck. A 2022 review put the range at nine to seventeen. Twenty years of textbooks and procurement templates settled it: qual saturates early, keep it small.

We wanted to see what that settlement costs. So we ran a real brief, an airline segmentation called Traveler Under Pressure: fixed script, screener, ratings battery, 495 AI-moderated interviews with US travelers. Then the part nobody budgets for. We coded every transcript. All of them, 7,686 verified spans, and suddenly the corpus is an answer key.

Because once you know what all 495 said, you can deal out two hundred smaller studies at every size and grade each one: what would this study have seen?

The key findings

What each sample size buys

25%50%75%100% 0% 102550100145300495 sample size drawn (log scale) · mean of 200 seeded draws per size share of the corpus’s 108 themes themes discovered (any of 108) rare themes captured (the 53 under 5%) S15 · Fare class is contextual · n95=2 (88.7% carry it)S04 · Delegation is a gated contract · n95=5 (52.2% carry it)S01 · Control loss survives remedy · n95=6 (42.2% carry it)S11 · Fee fairness has several grammars · n95=6 (40.9% carry it)S17 · Captivity retains customers · n95=7 (39.0% carry it)S06 · Mode preference is a state, not a trait · n95=11 (24.7% carry it)S02 · The void is its own injury · n95=14 (20.6% carry it)S08 · Remedy adequacy is the loyalty hinge · n95=15 (19.0% carry it)S07 · The burden of options, mechanized · n95=27 (10.8% carry it)S03 · “Gives a damn” modulates everything · n95=35 (8.2% carry it)S13 · Convergence math and the Spirit memes · n95=85 (3.5% carry it)S20 · Era fingerprints, incl. rebooking fraud · n95=114 (2.6% carry it)S19 · What actually worked · n95=172 (1.7% carry it)S16 · Disruption costs are bodily · n95=197 (1.5% carry it)S09 · Pseudo-remedy frictions · n95=230 (1.3% carry it)S10 · Four ways claims get abandoned · n95=276 (1.1% carry it)S18 · The premium that didn’t protect · n95=276 (1.1% carry it)S05 · Folk models of airline incentives · n95=345 (0.9% carry it)S14 · Spreadsheet-grade tradeoff machinery · n95=345 (0.9% carry it)S12 · Price reads as an integrity signal · n95=460 (0.6% carry it)

● each dot below the axis: one of the 20 traveler findings, at the n where a study is 95% likely to meet it. The pencil deepens as the study grows n=12 → 495. Hover any dot for the finding.

n≈12
The headlineControl loss outlives the fix; the orthodoxy earns its twelve.
≈55% of themes6 of 20 findings
n≈25
Existence, answeredFeels finished. It isn’t: found is not sized.
73% of themessizing ±5.6pp
n≈50–85
“Saturation”New themes nearly stop (86% at 50, 94% by 100); a third of the rare tail is still invisible.
86–94% of themestail ⅔ captured
n≈122
Roadmap-grade sizingPrevalence holds to ±2 points; ranking stops being a coin flip.
sizing ±2pp
n≈83–460
The decision tailFraud, bodily costs, folk economics: ten of twenty findings live below 5% prevalence.
53 themes under 5%80% tail at n≈83
>145
Structure: never reachedMean ARI peaks at 0.28 (n=145, 200 draws), still rising; your draws below will scatter around it.
mean ARI 0.28 @145
495
The audit852 verified quotes, every session scored, the moderator graded 2.8 of 5.
852 quotesmoderator 2.8/5
These are this corpus’s numbers: one instrument, one market, one language. Our own replication pool saturated at a different n because its theme space was smaller. The doctrine travels; the constants don’t.
Your turn

Draw the study, read both sides

Every cell is one of the 495 interviews (the 33 dashed cells failed the usability bar and can’t be drawn). Pick a sample size and draw. You get two sides of the same study: the front stage, the report a team would actually ship, proto-personas and themes and quotes at full confidence. And the backstage, what only the full corpus can see: the denominators, the sizing error, the segmentation agreement, the voice your draw never met. It opens at n=10, the study at its most confident and most wrong.

n =

one cell per interview · hover for who · filled cells are in the current draw · fill color = the draw’s n on the pencil scale

Front stage · what ships
Backstage · what’s true

Draw at 10 a few times and watch the front stage stay perfectly confident while the backstage swings. Then climb to 145 and feel everything settle, except the tail, and except the personas. That’s the ladder, in your hands.

“I'm not worried about the time or the money, I just need to know... don't keep telling me it's 20 minutes away when you don't have a part or you don't have the airplane”
P201 65+ New Baltimore, MI · 14:51

He wasn’t asking for much. Just to know. Ninety-five of the 462 were asking with him, and it took all four hundred ninety-five interviews to count them.

Next corpus, new constants. Let’s see.

Appendix · the receipts

For the auditors. Skippable, checkable: every session scored, the full instrument verbatim, and who we talked to.

Every interview, scored

3.8/5
participants’ average signal
2.8/5
the AI moderator, graded honestly

We scored the moderator too: question loops, false rejections of valid answers, echo-misquotes, seven distinct ways of mishandling a 1-to-7 scale. Every failure class is a fixable behavior, which makes the 2.8 a roadmap, not a confession. And still, 93% of sessions produced usable data, partly because participants ran the repairs themselves.

sort:
113
230
109
36
high signal (113) good (230) thin (109) low (36) void (7) excluded (33)
corner wedge = moderator: strong 4–5 functional 3 weak 1–2

One cell per coded interview (495). Cell color is participant signal; the corner wedge is the AI moderator's score for that session; dashed outlines are the 33 sessions excluded from analysis. Study level: mean IQS 3.67 (Signal 3.79 · Moderator 2.8 · Usability 3.96). Hover any cell for its decomposition.

What was asked

The pre-registered engine design, then the instrument’s guide, verbatim.

PRE
The meta-study pre-registrationthe design, committed before the engine ran
Four yield curves: common-code discovery, prevalence stability, rare-theme capture, segment stability (ARI).
200 without-replacement draws per rung, seed 20260713; rungs n = 10 / 25 / 50 / 100 / 145
Exclusion: Data Usability ≤ 2 (33 sessions); analysis set 462
Sensitivities: elicitation-form (A-family dropped), word-count residualization, replication pool (R1)
No LLM anywhere in the analysis loop; plain Python + scikit-learn
Q1
Travel self-imageidentity vs actual behavior
“What kind of flyer are you?”
“What makes you say that?”
“Can you tell me about a recent flight that really shows that side of you?”
“Are there situations where you become a different kind of traveler?”
Q2
How you choose flightsreal tradeoffs (avoidance over preference)
“Think about the last time you booked a flight. How did you decide which one to pick?”
“What did you look at first?”
“What almost made you choose a different option?”
“Did you end up paying more than the cheapest option? Why or why not?”
“What were you trying to avoid when making that decision?”
Q3
Tradeoff momentcertainty vs price
“Imagine two flights: one is cheaper but has a worse schedule or higher chance of hassle; one costs more but feels easier and more predictable. Which would you choose? What would make you switch your answer?”
“How big does the price difference need to be?”
“What kind of “hassle” matters most to you?”
Q4
When things go wrong (core)behavior under real stress
“Tell me about a time something went wrong while flying. What happened?”
“What was the most frustrating part of that experience?”
“What did you want from the airline in that moment?”
“Was it more about time, money, uncertainty, or how you were treated?”
“When did it go from “annoying” to “not okay”?”
“Consistency probe (always asked): Earlier you said [their priority]. In this situation, did that still feel true?”
Q5
Control vs airline responsibilitythe options-vs-delegation axis
“When something goes wrong, what do you prefer: having options to choose from, or having the airline just take care of it for you?”
“When do options feel helpful?”
“When do they feel like extra work?”
“What would make you trust the airline to decide for you?”
Q6
Fairness and breaking pointdignity, fairness, switching threshold
“When does a cheap flight stop feeling like a good deal?”
“Is it about the total cost, or how the costs show up?”
“What kinds of fees feel fair vs unfair?”
“What kind of experience would make you say “I’m never flying them again”?”
RF
Rapid-fire calibrationforced choices + five 1-to-7 ratings
Forced choice: options vs airline decides · cheapest vs predictable · worst problem (time / money / control / not knowing / disrespect) · reason for staying.
1. I am willing to pay more to avoid uncertainty
2. Having too many options during travel makes things more stressful
3. I trust airlines to make good decisions on my behalf
4. Hidden fees bother me more than higher upfront prices
5. Once an airline loses my trust, it is hard to win me back

Who we talked to

495 interviews coded · 462 analyzed · all 50-state US, English

Age

18-24
9.7%
25-34
32.7%
35-44
29.7%
45-54
16.9%
55-64
9.1%
65+
1.9%

Screener price identity (Q7)

“cheapest that works” 46.3% · “pay more to avoid stress” 53.7%

Airline type (Q5)

A mix of both
24.9%
Mostly low-cost airlin
5.4%
Mostly major/full-serv
69.7%

Flights in past 2 years

1 time
6.3%
10+ times
16.9%
2–4 times
40.9%
5–9 times
35.9%

Trip type

A mix of both
35.1%
Mostly business/work t
5.0%
Mostly leisure (vacati
53.5%
Visiting family or oth
6.5%

Top states

CA
14.7%
FL
8.2%
TX
6.1%
NY
5.2%
OH
4.3%
PA
3.9%
GA
3.9%
NC
3.9%

Built with AI + analyzed with human care. The Observatory x Dialogue AI, July 2026.