
A/B Test Development Process: How to Build Tests That Don’t Break Your Site

Meera runs CRO for a D2C footwear brand. Every Monday she opens the same Notion page: 40 unread session recordings, a growing pile of product reviews, a support inbox nobody has time to tag. She calls it “the graveyard.” Her team calls it Tuesday.

She’s not lazy and she’s not understaffed by industry standards. She’s normal. And the numbers back that up:
Debt is the right word. It compounds. Every unread transcript is a decision Meera’s team makes on gut feel instead of evidence.

Read a file
Read a file
Meera’s first move was the one almost every team makes. She pasted 18 product reviews into an AI chat and typed: “Can you summarize what customers are saying?”
Read a file
Read a file
It read well. It also told her nothing she could act on:

This is the trap. A chronological summary feels like research. It isn’t. Real synthesis means ignoring the play-by-play and pulling out patterns that only show up when you look across every response at once, something a recap can’t do by design.
Same 18 reviews. Different prompt. This time Meera attached the file and gave the AI rules, not a request:
Read a file
Read a file
Same data. Same tool. A completely different, checkable output:
Every number here survives the question “show me the actual review.” That’s the entire difference between the two screenshots above: not the model, the instructions.
Read a file
Read a file
AI owns exactly one box on this list: stage 3. Stage 4 is what makes stage 3 trustworthy. Skip it, and you’ve automated guessing instead of replacing it.

Here’s the part most CRO teams never think to ask: is this a counting question or an understanding question?

Academic researchers have argued this distinction since 1987, calling it Small-q versus Big-Q research. A 2025 paper in the International Journal of Qualitative Methods lays out exactly why it matters for anyone using AI on qualitative data.
Read a file
Read a file
Meera’s rule now: before she opens a chat window, she decides which side of that line her question is on.
The footwear example works in one prompt. It doesn’t scale. By Q3, Meera’s team had 1,800 reviews and a full year of support tickets, too much for one inductive pass to stay consistent, especially with that 36-47% intercoder reliability ceiling from the JMIR data hanging over it.

The fix, borrowed directly from the same academic research: seed-and-scale.
Read a file
Read a file
Use inductive AI-first clustering for small, unfamiliar datasets. Use seed-and-scale once the pile is too big to read end to end. Pick the wrong one and you get themes nobody can explain six weeks later.
Formal inter-rater reliability testing needs a statistician. Most CRO teams don’t have one on call. There’s a shortcut, and it comes straight from the research: run the identical prompt three times, in three separate conversations.

One study ran this at scale, ten independent ChatGPT passes compared against each other, and used the agreement rate as a reliability proxy. You don’t need ten. Three catches most of the wobble, for the price of two extra prompts.

A clustered theme is not a test idea. It becomes one the moment it fits this sentence:
Read a file
Read a file
Can’t fill in the metric blank? The theme isn’t ready. Go find more evidence first.
Read a file
Read a file
CXL’s PXL framework does this for test ideas generally: ten objective yes/no questions, scored and summed, no gut feel allowed. Applied here:
Two themes ship this quarter. Two get benched, on the record, instead of quietly forgotten or forced onto the roadmap because someone felt strongly about them.
Fail two or more: it’s not ready. One more round of triangulation, then it goes on the queue.
Meera’s fix took one afternoon: pick the oldest ignored export, strip anything personally identifying, run it through the pipeline above, three passes, human audit, one hypothesis per surviving theme. That’s the whole first deposit against your own Transcript Debt.
Want a second pair of eyes on the process? OptiPhoenix runs a free mini audit checking whether your last quarter of research would survive this exact checklist. Need the tagging protocol and scoring model built from scratch? Let’s talk growth.
No. AI accelerates the mechanical parts of qualitative analysis, coding, clustering, summarizing at volume, but interpretation, cultural context, and judgment calls still require a human. The JMIR study behind this article found AI matched human coders on only 71% of themes in the best-case scenario (inductive coding), dropping further when AI was constrained to a predefined structure.
It’s the practice of using a large language model to help identify, cluster, and count recurring themes in qualitative data such as reviews, support tickets, or interview transcripts, with a human reviewing and validating the output before it’s used for decisions.
Research puts inductive (AI-discovers-the-themes) coding agreement with human coders at around 71%, and intercoder reliability on matched themes at a fair-to-moderate 36-47%. Accuracy drops further with deductive coding, where AI is told what themes to look for in advance.
Inductive coding lets themes emerge from the data with no preset categories. Deductive coding checks data against a fixed list of themes decided in advance. AI performs measurably better at inductive coding.
Only after removing personally identifying information. Commercial AI tools require sharing whatever is pasted in with a third-party platform, and no formal industry-wide ethical guideline yet governs AI-assisted qualitative research.

Attach the raw data as a file, then instruct: cluster into themes, cite the specific source for every mention, set a minimum mention threshold for “actionable,” and forbid inferring reasons not stated directly.
Three times, in three separate conversations, is a practical minimum.
Small-q is structured and countable; Big-Q is interpretive and meaning-driven. AI is a reliable first-pass coder for Small-q. For Big-Q, it’s a brainstorming partner only.
Once the dataset is too large to read end to end, typically past a few hundred records.
No universal number, but three is a common practical floor, decided before you see results.
Strip or mask identifying fields before anything goes into a general-purpose AI tool, and check your vendor’s actual data retention policy.
