A collaborative study involving researchers from McGill University, Utrecht University and INSEAD has tested how well indirect measures of hidden racial biases predict behavioural decisions when direct self-reports remain the strongest predictor. The project, designed as an adversarial collaboration, recruited more than two thousand White American adults to compare lab-based behavioural tasks with both indirect tests and direct attitude surveys.
In psychology, direct and indirect measures of racial attitudes sit at opposite ends of a spectrum. Direct measures rely on conscious self-report surveys where respondents control their answers, while indirect measures aim to capture automatic associations that people may not readily disclose. A common indirect tool is the Implicit Association Test, where faster pairing of Black faces with negative words versus White faces with negative words is taken as an indicator of implicit bias.
The debate over how well these automatic responses translate into real-world discrimination is longstanding. Critics argue that implicit bias tests may reflect cultural knowledge rather than personal prejudice, and often fail to predict individual actions. The study’s lead author, Jordan R. Axt, explained the motivation to bring competing viewpoints together: “Rather than continuing to work separately, the goal of this study was to get relative proponents and skeptics of implicit measures … to come together and agree on a study design that would be informative no matter how the results turned out.”
The team assembled 2,114 White American adults for a two-session online investigation. In the first session, participants engaged in four behavioural tasks that could reveal discriminatory tendencies. These included a trust game to decide how much money to share with Black and White partners, an ultimatum game to accept or reject monetary splits, and two simulated hiring scenarios—one in which resumes with Black- or White-sounding names were reviewed, and another in which a hiring decision was based on profiles featuring metrics such as math proficiency and headshots.
A few days later, in a second session, participants completed four indirect measures of racial bias, including the Implicit Association Test and an evaluative priming task in which faces were briefly flashed before categorising subsequent words. They also completed an affect misattribution procedure, rating the pleasantness of Chinese characters following a Black or White face. Five direct attitude measures followed, asking participants to rate warmth and liking toward racial groups and to endorse certain stereotypes. The researchers employed structural equation modelling to pool related tests and reduce measurement error.
The findings presented a nuanced picture. On average, participants demonstrated a pro-White and anti-Black stance on the indirect, time-based measures. Yet the behavioural tasks revealed a different pattern, with a slight pro-Black tilt in hiring and monetary decisions. A small subset, about seven per cent, showed extreme pro-White scores on the indirect measures and engaged in behavioural discrimination against Black targets.
In terms of predictive power, the indirect measures did carry some weight. They accounted for about 2.5 per cent of the unique variance in behaviour beyond what self-reports explained, suggesting that these hidden-bias tests capture a distinct aspect of social decision-making. “The clearest takeaway is that indirect measures do capture something meaningful about race-related behaviour that is not already captured by self-report,” said David S. March, an associate professor at Florida State University not involved in the project. “That incremental contribution is modest, but it appears to be real, particularly when implicit attitudes are treated as a broader latent construct rather than relying on any single measure.”
Nevertheless, direct self-report measures vastly outperformed the indirect tests. Explicit attitudes explained roughly 45 per cent of the variance in the behavioural tasks, indicating that asking people directly provides far stronger predictive information about how they might behave in these simulated social decisions. As Axt put it: “The data are clear that explicit attitudes were much better predictors of behaviour than implicit attitudes, so perhaps measures of implicit attitudes need not be such a large focus in studies on these issues.”
There was a caveat, however. The researchers also tested whether fatigue or tiredness would push participants to rely more on automatic biases, but the data did not show any such effect. “But when we included measures of how distracted or tired participants felt, they did not seem to have any effect on the relationship between implicit associations and behavior,” Axt noted, while acknowledging measurement limits.
The study aligns with prior PsychoPost coverage suggesting higher reliability and stronger predictive validity for self-reports relative to implicit tools, while also underscoring that the signal from implicit measures becomes clearer when measurement error is carefully considered. It also reflected patterns observed in a 2023 examination of racial attitudes and outcomes in different contexts, though conducted with different experimental setups.
The authors recognised several limitations. The online, laboratory-like tasks may not fully capture real-world workplace dynamics, and participants may have altered their choices to appear egalitarian. Axt emphasised that none of the behavioural measures within this study showed anti-Black discrimination, raising questions about how these results generalise to other contexts or actions where participants are unaware they are part of a psychology study.
March urged caution against over-interpreting the findings as evidence that implicit bias broadly produces anti-Black discrimination. “The finding is that people higher in implicit bad versus good associations tended to behave relatively less favourably toward Black targets than people lower in implicit bias,” he said, adding that the results should be tested with more reliable, domain-tuned predictors in naturalistic settings to see whether relationships strengthen when the predictor aligns better with the behaviour in question.
Reliability also emerged as an issue. Many of the behavioural measures and one of the indirect tests displayed low internal reliability, with the trust game being the notable exception in showing consistent responses. This challenges the precision of effect size estimates and underscores the need for more robust metrics in future studies.
Beyond the United States’ White American sample, the authors acknowledge that results may differ across other demographic groups and contexts, particularly for attitudes toward groups beyond Black and White individuals.
The researchers advocate pursuing studies in more consequential environments and organisational settings, hoping to observe whether implicit and explicit measures continue to independently predict outcomes such as promotions or performance reviews. They also stressed the value of replications using reliable, real-world tasks to better gauge the practical significance of implicit biases in everyday decision-making.
As a methodological exercise in collaborative science, the researchers described the adversarial collaboration as productive and intellectually engaging. “For an ‘adversarial collaboration,’ it really was a pleasant and intellectually engaging experience,” Axt said, noting that participants remained genuinely invested in the study’s value even when conclusions diverged. “It definitely made me interested in pursuing additional collaborations of this type in the future.”
The study, titled On the relationship between indirect measures of Black vs. White racial attitudes and discriminatory outcomes: An adversarial collaboration using a sample of White Americans, involved Jordan R. Axt, Paul Connor, Suzanne Hoogeveen, Cory J. Clark, Michelangelo Vianello, Joanna N. Lahey, Adam Hahn, Jeffrey To, Richard E. Petty, Thomas H. Costello, Gregory Mitchell, Philip E. Tetlock, and Eric Luis Uhlmann, and was conducted across multiple institutions.
