In August 2024 our team ran an experiment. We had a fresh round of client interviews in hand, and we synthesized the same transcripts two ways. One team worked the traditional process: notes broken down into key points, clustered into themes across participants. A second team gave the transcripts to GPT-4 and prompted it toward cross-cutting themes.
At the time, the results were humbling for the machine. The themes it produced were accurate but basic, while the hand synthesis was consistently more nuanced. The AI track added time instead of saving it, because preparing transcripts and correcting misreadings became its own line of work. And when we asked it for something as small as removing filler words from a transcript, it changed what people had actually said. Those were 2024 findings about 2024 tools. We state them as history, not as claims about today's models, which we have not re-tested in the same controlled way.

Two years on, the more interesting question is not the scoreboard. The tools have improved enormously at the mechanics: longer context, cleaner handling of transcripts, fewer of the clumsy errors that filled our 2024 notes. What has not changed is the discipline, and the discipline was always the point.
Three principles from that experiment still govern how we work.
Firsthand observation comes first. Synthesis is not summarization. Much of what an interview teaches lives outside the transcript: cadence, hesitation, the question a participant answers instead of the one you asked. A model working from text alone cannot see any of that, no matter how capable it becomes. So we still listen to the interviews ourselves, and we still form our own read before asking a model for one.
AI supplements the synthesis; it does not start it. Beginning from machine-generated themes constrains your thinking to them. You end up hunting for evidence that fits, which is confirmation bias with better tooling. Run your own synthesis first, then put the model to work where it is genuinely strong: pulling supporting quotes quickly, checking that a theme holds across every participant rather than one vivid interview, surfacing the plain-sight patterns a human eye skips past.
Verify the work. The filler-word episode is dated, but its lesson is not. A model's output reads equally confident whether it is right or wrong, and the errors that matter are the quiet ones. Anything that will carry weight with a client gets checked against the raw material before it travels.
What strikes us now is how closely this tracks something we see across product work generally. The mechanical parts of the job keep getting cheaper: drafting, transcribing, clustering, formatting. What is left is judgment. Knowing which theme matters, hearing what a participant did not say, deciding what the evidence can honestly support. The 2024 experiment read at the time like a verdict on a tool. It reads now like an early sighting of that larger shift. The tools improved at the mechanics, and the work that remains is exactly the work we would want to keep.
