A researcher built a synthetic population of 1,000 AI personas, handed each one the Langer Mindfulness Scale, and got back the wrong shape. A deliberately varied population should have produced something close to a bell curve. What came back clustered in the middle and upper part of the scale. His explanation is not psychology but training: reinforcement learning from human feedback rewards curiosity and flexibility, the model absorbs that preference, and a 14-item instrument built for people picked it up.
An AI persona, in the sense used here, is a generative-AI simulation that computationally reproduces possible human thoughts and behavior. A researcher can run one, dozens, hundreds, thousands, or far more. The proposal is to treat them as a methodological tier in their own right, sitting alongside studies on human subjects and hybrid designs in which people and personas both participate. Personas do not replace human subjects, but they can carry a share of the research load.
The hybrid case leaves an unsettled question of bookkeeping: whether the synthetic half of an experiment counts as preparatory work or as a study in its own right. The results worth having, on this account, are the ones that can be set beside data collected from people.
The scale itself is about 15 years old and measures the capacity to perceive things anew. A mindful person does not stop at an automatic response to the world but notices what is happening and draws new distinctions. The instrument has 14 questions, each scored from 1 to 7, summed into a single mindfulness figure. It has been used on individuals and also to look at the trait at the level of organizations and societies.
The author names four uses for personas here. They can pilot a study before it runs on people, since recruiting humans costs money, time and organizational effort. They can test new hypotheses about mindfulness as a theoretical construct quickly enough that a researcher can vary the approach, compare versions and carry only the promising ones forward to human subjects. They can be used to study the scale itself, as an instrument. And the scale can be turned around and used to study the AI, as a readout of its computational behavior.
The warnings that follow are more specific than the proposal. The scale has three dimensions: novelty seeking, novelty producing, and engagement. Build every persona out of a disposition to seek novelty and the scale will duly report a population oriented toward novelty, and a researcher may conclude that personas tend that way in general when the result only describes how they were made. The prompt should also not mention mindfulness at all: ask the personas to be mindful and the model will write mindfulness into their characteristics. Deliberate bias is permissible if the researcher states plainly what was done and how it shaped the outcome.
The run itself was modest. Personas can be created in essentially any large language model with the right prompts, and popular models often ship with ready-made personas built by their developers; for psychological work the author recommends his own taxonomy. Each of the 1,000 characters was asked to complete the scale and answer all the questions from 1 to 7. The raw data went to a file, then a spreadsheet, then a statistical tool. The expected distribution did not appear.
Here is what I think this exercise demonstrates, as against what it sets out to. It is a pitch for AI personas as a psychological instrument, and its single empirical result is the strongest argument against taking one at face value. The warnings are the substance. Every one of them describes the same failure: the output encodes the input. RLHF is that failure in its least tractable form, a prompt nobody in the experiment wrote, applied by raters nobody names, baked in before the researcher opens the model. The proposed fix is a prompt designed to compensate for it, which is another way of saying the researcher now has to guess the size of an effect he cannot see in order to cancel it.
What the write-up does not have is a comparison. The claim is that the synthetic population scored higher than a randomly chosen group of people probably would — probably, with no human data collected here to check it against. Nor does the text say which model produced the personas. RLHF is not one procedure; it differs between labs and between versions of the same model. A distributional result attributed to it, without a model and without a date, is not something a second researcher can reproduce or refute.
The follow-up is where the real test sits. The plan is to work through mean, standard deviation, skewness, kurtosis, correlations and Cronbach's alpha, then give the personas a novelty-detection task and see whether each one's score predicts how it behaves. If it does, the scale is measuring something stable inside the simulation. If it does not, then 1,000 rows of numbers filed under the name of a human trait are a portrait of a training pipeline, and the reason they skew high is the most informative thing about them.