All posts

How to choose a synthetic research method in 2026

Most teams weighing up a research method are still asking whether synthetic research works. That question has been settled for a while. Simulated audiences are sold and used at enterprise scale, and the money going into the category says the same thing.

The question that pays is narrower. Which method suits the decision in front of you, and how much verification does that decision earn?

What is the decision actually worth?

Start there, not with the tool. Research budgets have always been rationed, and the rationing used to be crude: commission a study for the big decisions, use judgement for everything else. Faster methods change the arithmetic. They do not remove the need to match effort to stakes.

Some work is genuinely exploratory. Sketching an angle before a planning meeting. Testing a headline. Working out which questions are worth putting to a real customer. Nothing is committed yet, and if the answer is wrong that becomes clear quickly and cheaply.

Other work carries value. A media budget. A campaign plan that runs for a quarter. A regional launch, a pricing decision, a channel mix. Sizing a market before committing to it belongs here too. Something happens because of the answer, and if the answer is wrong, the money has already gone.

The methods on offer are not interchangeable across that line.

Where do generated personas earn their place?

For exploratory work, a persona produced by a general-purpose language model is quick, cheap, and often good enough. It will give you a plausible person with plausible motivations, and it's a useful tool for opening up thinking. Cambium AI keeps a comparison of the free options for teams at that stage, alongside a plainer explainer of what a customer persona is and what goes into one. The same holds for running research queries through a language model: quick, useful, and best treated as a first pass.

The limit is worth stating precisely, because it is narrower than it is usually made out to be. A language model will write a convincing persona. What it cannot do is show which real population that persona came from. There is no variable to inspect, no sample to size, no source to cite. That is a gap in verifiability rather than in capability, and it stays invisible right up to the moment somebody asks you to justify the decision.

What are focus groups and panels still best at?

Depth, language, and reasons. A person in a room will tell you why they chose something, in their own words, and no volume of population data produces that. Cambium AI has written about getting some of that depth without scheduling calls, but the underlying value of qualitative work is unchanged, and the case for running qualitative and quantitative work together is this same argument at a smaller scale.

What a small group cannot give you is proportion. Eight people cannot tell you how many households in a region resemble them, and that is usually the number a budget rests on. Behavioural-data products answer a different question again. They are strong where observed behaviour is the thing you need, and quiet on the people who are not in the data yet.

What should you ask before acting on any of it?

The questions do not change with the tool. Cambium AI's verification checklist sets them out in full. The short version is five.

  1. Which dataset is this from, and can you name it?
  2. What was the sample, and at what geographic resolution?
  3. What is the error band around the number?
  4. When was it last refreshed?
  5. Could two people reconstruct the same answer from the same source?

GreenBook's practical guide to synthetic data and augmented sample makes much the same argument from the industry side, recommending holdout tests and equivalence checks on the measures that matter rather than accepting output because it reads well.

Where does verified public data fit?

On the decisions that carry value. Cambium AI builds personas from verified public data, including US Census Bureau household data alongside other official sources, joined at the level of individual records. Two things follow from that. Anyone in the room can check which variable a given persona came from. And because the records are joined across sources rather than fitted to a published total, the population behind the persona is proportioned like the real one rather than like a plausible story.

That is different work from data access, which is now widely available. There are good public-data tools inside AI assistants if a quick lookup is all a question needs. It is also not a replacement for talking to customers. It is the part of the evidence that survives being questioned, which is what market validation asks of it.

The practical test is short. Ask what the decision is worth, then ask whether you could defend the answer if it went badly. For a first pass, a generated persona is fine. For a quarter's budget, Cambium AI is the insurance policy for the decision: verified public data and a method you can show. See it in Cambium AI →

Ask AI about this post:
ChatGPT