Replacing human subjects with AI surrogates, or digital twins, is on some social scientists’ wish lists. Human subjects are expensive, tire easily and can suffer psychological distress from some studies.
But those wishes may take more time to be realized, a study appearing September 2 in Science Advances suggests. AI twins designed to mimic a given individual’s behavior instead seem to distort their surrogate’s views, creating a “funhouse mirror” effect, the researchers note.
“There’s some promise,” says Olivier Toubia, a computational social scientist at Columbia Business School in New York City. “But [the twins’ performance] was overall a bit disappointing.”
In work reported last year in Marketing Science, Toubia and colleagues recruited more than 2,000 individuals from across the United States. Those respondents answered 500-plus questions about characteristics including age, ethnicity, income, education, religious practices, political preferences, personality traits, spending habits, mathematical abilities and vocabulary skills. Many of the questions come from scales that are used commonly in psychology, economic and business research. Respondents also completed various online tests designed to assess thought patterns and biases.
Because research testing the value of digital twins remains limited, the goal of that project was to create an open-source dataset for others to use, Toubia says. “I think it’s been downloaded like 25,000 times at this point.”
To develop the twins in the newly published study, the researchers fed all the information for each individual into a large language model and prompted the LLM to respond as if it were that person. Across 19 social science experiments, the team evaluated everything from how individuals and their twins respond to people who donate to both Republican and Democratic party candidates to what they “think” about algorithmic hiring.
These digital twins performed better than chance, the team found, but they were wrong on average about a quarter of the time. They performed roughly on par with chatbots that received demographic information alone.
The digital twins did, however, better capture real variation in people’s responses compared with the LLMs that knew only demographic info. For example, one person may rate themselves as a 2 on a scale of self-control while another rates themselves as a 4. The LLM with more limited info might report a 3 for each person, washing out any differences. But the digital twin might report a 3 and 5. Though still wrong, the predictions give a better sense of potential differences across the group.
Toubia and colleagues attribute the digital twins’ poor performance overall to several key distortions: Twins’ responses tended to be more homogenous than people’s responses and often skewed to demographic stereotypes. And their accuracy increased with more affluent and educated participants. They also displayed certain biases, such as expressing more trust in others and showing less concern about technological threats. Compared with their human surrogates, the twins also appeared more rational.
AI researcher and economist Hadi Hosseini of Penn State University has seen a similar effect in his own research into how AI agents make health care decisions when resources are scarce: “LLMs distort human judgment in a very specific direction toward something that is very rational [and] more reasonable than actually what people are.”
But he says there are ways that the digital twins could be improved. The researchers used “a very static set of questions,” he says. Adding more dynamic approaches, such as having a chatbot shadow an individual across the day or converse regularly with a person, as is already common, could make for a better dataset, he says.
Toubia wants to try more complex training methods. Still, he says, there are cases where even digital twins of the type tested in this study would be useful. For instance, sometimes researchers need a detailed response to a question. Tired humans might provide single-sentence answers while indefatigable digital twins churn out essays. Similarly, pretesting an experiment with digital twins could help researchers ensure their design is working before sapping their human respondents’ limited bandwidth.
But Toubia also urges humility. Social scientists tend to think that scales and surveys capture the full range of human experience, he says. But it’s incredibly hard to predict human behavior with a machine. “We need to be realistic in terms of the expectations we have from synthetic data.”
Read the full article here














