Forbes cited our research on what one label does to an AI persona
Forbes published a piece on 2 September about AI and leadership assessment, and it cited our research on what a language model does when it is asked to represent a particular kind of person. It is here, by Christine Michel Carter.
This is the short version of what we found, what we think fixes it, and where our work stops and the article's argument begins. We think the last of those matters as much as the first.
What we found
Imagine measuring how entrepreneurial 100 women and 100 men are, and plotting each group as a bell curve. In real life, those two curves overlap heavily. Ask a language model to simulate the same two groups and it pushes them more than three times farther apart than they actually are.
We measured the same effect on agreeableness. The real difference between men and women is about 0.6 standard deviations. A model told only that it is a woman produces a difference of about two. So the model is not representing a person. It is representing the label, and it takes a small real difference and returns it as a caricature. That is the finding the article cites, and we set out how we arrived at it in why language models flatten a population.
Why one label does so much of the work
The problem is not that a model cannot represent people. It is that a single label is asked to carry the whole person. If all a model knows is "she is a woman", that one word does all the work, and what comes back is everything the model has ever absorbed about the category rather than anything about an individual.
Our answer to that is the opposite of what most people expect. Tell the model more, not less. Learn that she is also Asian, a mother, and a city resident, and "woman" stops carrying so much explanatory weight. Add enough true, co-occurring detail and the competing stereotypes begin to offset each other, and what is left behaves more like a person. There is no single female experience, and a model only stops assuming one when you give it enough to work with.
One constraint on that, which we have written about before: the detail has to be true together. Attributes that could not co-occur in a real person do not cancel; they compound into noise. Why more detail makes an AI persona more real goes through how that works, and our methodology page covers how the underlying populations are built and checked.
A limit we are open about
Michael Birdsall, one of our co-founders, also told Forbes about something difficult to achieve. Language models struggle to express intense emotion, anger in particular, as strongly as people actually feel it.
That matters for anyone planning to test a difficult decision against a synthetic population. Reconstruct someone's age, income, family, and occupation as carefully as you like, and the model is likely to still under-express how angry they would be. On high-emotion topics, a synthetic population is a starting point and not a substitute for asking real people.
What should you ask before acting on a persona?
She closes on advice worth repeating. Ask what the model actually knew about a person beyond their gender. Ask which behaviours it observed and which characteristics it inferred. And do not let "the AI recommended it" stand as evidence by itself.
All three point at where a persona came from rather than at how good its answers sound. That distinction matters because the answers sound the same either way. A persona built from one label and a persona built from hundreds of details that genuinely occur together are indistinguishable in the reply.
If you want to see a demo of what we are building, get in touch here.
