A team of researchers has found that cutting-edge language models trained to process text and images can judge a person’s character from their facial features, mirroring a long-standing human bias. In a series of experiments, large language models were shown to rely on arbitrary facial characteristics to assess competence or the likelihood of criminal activity. The work, published in PNAS Nexus, raises fresh concerns about the use of such systems in high-stakes decisions.
For decades, people have made rapid judgments about others based on facial appearance—a tendency known as face-to-character inference. It rests on subtle differences in facial shape and texture, yet offers no real information about a person’s actual traits. A 2014 study found that even three-year-olds make character judgments from faces, illustrating how deeply embedded this bias is in human thinking.
“Faces are both pervasive and significant social stimuli,” said Mahzarin R. Banaji, the Richard Clarke Cabot Research Professor of Social Ethics at Harvard University and external faculty at the Santa Fe Institute. “So much so that the human brain has a dedicated region that responds to faces—we are, each and every one of us, ‘face experts.’ Face-based judgment permeate so many decisions we make and so getting it right is important. So there was a pragmatic reason to focus on face-based judgments.”
The study, led by Steven A. Lehr of Cangrade, Inc., with Banaji and Yash Lothe, examined whether language models—now increasingly multimodal—mirror this particular human quirk. Because these systems are trained on vast amounts of human text, which captures a wide spectrum of biases, the researchers wondered if the models might learn to associate certain facial types with specific traits.
Lehr described the core idea succinctly: “My own operating theory is that LLMs are a mirror that will reflect just about any human characteristic in surprisingly high fidelity. For this reason, I’m on the lookout for humanlike characteristics that would be surprising in a machine. Face-to-character biases fit the bill.” Banaji added that there was a hope the visual biases might not emerge in language-based AI, asking, “Would AI save us from our own face-based errors of judgment (see the work of the brilliant psychologist, Alex Todorov)?”
The researchers pursued a rigorous programme of testing, conducting 13 experiments totalling nearly 8,000 trials across four language models. In the first two experiments, GPT-4o was presented with pairs of computer-generated faces that had been subtly altered along features linked to competence or trustworthiness. Faces in each pair varied by several standard deviations to create large and small perceptual differences.
Across 600 trials focused on competence, GPT-4o selected the more competent-looking face 87.83 per cent of the time. In 600 trials assessing trustworthiness, the model chose the expected face 72.67 per cent of the time. The bias intensified as the visual differences between faces widened: for faces six standard deviations apart in competence-related features, the model’s correct-appearance choice reached 98 per cent, while two standard deviations yielded a still-strong 70 per cent.
Remarkably, the authors noted that the model’s bias was amplified relative to human performance. “Readers should notice that these were large and practically meaningful effects,” Lehr said. “Indeed, according to standard effect size measures, the bias appeared to be notably amplified relative to that of humans. It should be noted that this is partly because the LLMs were so consistent in showing the bias, so this may reflect partly low response variance as opposed to just truly greater essential levels of bias. But of course, consistency matters in contexts like selection: it makes the bias more reliable.”
The bias extended beyond straightforward traits. In 2,160 trials, GPT-4o evaluated the same faces on related characteristics such as intelligence or laziness (competence) and warmth or selfishness (trustworthiness). The model consistently generalised its judgments, selecting the expected face 74.63 per cent of the time for competence-related traits and 73.33 per cent for trustworthiness-related traits.
In a fifth experiment, the researchers exposed the model to 150 pairs of macaque faces, rated by humans as mean or nice. GPT-4o chose the “nice” monkeys as more trustworthy in 66 per cent of trials, suggesting the model develops a broader concept of facial trustworthiness that extends beyond human faces.
The study also explored high-stakes scenarios. In 540 trials, GPT-4o was asked which face was more likely to be a serial killer, a human trafficker, or a Ponzi schemer. The less trustworthy-looking face was chosen 68.70 per cent of the time. In a separate set of 540 trials examining positive real-world outcomes, such as hiring a university president or funding a startup, the model recommended the more competent-looking individual 75.19 per cent of the time.
“GPT-4o did not show any real reluctance to say that one person was more likely to be a serial killer, human trafficker, or Ponzi schemer,” Lehr pointed out. “And the models readily told us we should hire the person with the more competent-looking face.”
The researchers then tested whether the same bias persisted in newer models. GPT-5, Gemini 3 Flash Preview, and Claude Sonnet 4.5 were subjected to the same basic competence and real-world decision tasks. All three exhibited substantial face-to-character biases, sometimes surpassing the levels seen in GPT-4o. In particular, GPT-5 showed an even larger bias, selecting the expected face 94.33 per cent of the time on basic competence judgments and 97.04 per cent for real-world decisions. Gemini 3 and Claude Sonnet 4.5 similarly demonstrated elevated bias levels.
Lehr remarked that one might have expected more advanced reasoning models to resist such biases, but evidence showed the opposite: “One might have thought that the three tested reasoning models would think through the questions more and avoid this bias, but in fact, they showed it to a greater degree.”
Interpreting the findings, the researchers cautioned that the experiments relied on forced choices between two faces and used two-dimensional static images. They suggested results might differ with more natural settings, video inputs, or varied angles. They also questioned whether biases originate directly from training data or other factors unique to vision in language models, noting that the bias appears amplified in models across generations and may reflect low response variance rather than solely greater essential bias.
The team stressed that current alignment efforts aimed at public-facing language tasks do not guarantee safety in more sensitive domains. While guardrails exist to curb explicit gender or racial prejudice, subtler biases tied to facial shape may escape such protections. The researchers emphasised that the implications for human resources, law, and other high-impact areas warrant careful scrutiny and testing for biases before deployment.
Lehr summed up the position: “The training designed to align AI models with human values has achieved what looks like surface egalitarianism, but something deeper is needed. We should be cautious about using AI models in high-impact domains, such as hiring and law, and if we do use them, we should be certain to always carefully test them for biases.”
The study, entitled “Like humans, language models demonstrate face-to-character biases,” was conducted by Steven A. Lehr, Yash Lothe, and Mahzarin R. Banaji.
