AI screening tools are already filtering job applicants before human eyes reach the resume. A new study from Princeton University and the University of Chicago shows that the models doing that filtering – including ChatGPT, Claude, and Gemini – develop stronger hiring biases than the humans they are meant to assist.

What the Experiment Actually Tested
The researchers adapted a psychology study originally designed to observe how humans form stereotypes through experience. In the original human version, participants made hiring decisions and received feedback on outcomes. The researchers ran the same framework through large language models, giving each one a simulated consultant role working for the mayor of a fictional city.
Each model was asked to hire candidates for 20 different jobs – doctors, lawyers, child-care aides, janitors, and others. The applicant pool consisted of members from four fictional ethnic groups: Tufa, Aima, Reku, and Weki. In every round, the model picked one of four candidates (one from each group), received feedback on whether that hire succeeded, and then moved on to fill the next opening. The process ran across 40 rounds total.
There was a hidden baseline built into the experiment: all candidates from every group were equally likely to succeed at every job. No group was coded as better or worse. The models had no factual basis for sorting applicants by group – only the noise of early outcomes.
That noise was enough. After a candidate from one group failed in a particular role, the models rapidly reassigned the entire group to lower-status positions. When a model saw that an Aima candidate failed as a doctor – a job the model associated with high warmth and competence – it began routing all Aima applicants toward janitor roles instead. One early data point triggered a durable pattern.
LLMs Outpaced Humans on Every Bias Metric
The study used a segregation scale ranging from 0 to 2, where 2 represents total confinement: every group locked permanently into its own job category with no crossover. Human participants in the original psychology study scored 0.84 on that scale. The language models scored roughly 65% higher. OpenAI’s reasoning model o3 reached 1.83 – nearly the maximum possible score – meaning it came close to fully segregating all four fictional groups into separate job niches based on nothing more than random early feedback.
Ryan Liu, a PhD student at Princeton and a coauthor of the study, attributes this to how the models were trained. Large language models are optimized heavily on math, coding, and science tasks – problem types that reward extracting broad rules from limited examples. That same instinct fires in social contexts, where generalizing quickly from a small sample is exactly the wrong approach. “That’s literally a lot of what they’re optimized for,” Liu said of the tendency to form generalizations fast.
The study, published at the International Conference on Machine Learning in Seoul in July, found that the bias problem intensified with newer, more capable models. OpenAI’s o3 and DeepSeek’s R1 – both high-reasoning models – showed stronger stereotyping behavior than older systems. More reasoning power did not reduce discriminatory patterns. It amplified them. Every decision-maker, human or machine, navigates what psychologists call the exploration-exploitation dilemma – the choice between acting on past experience and trying something new. LLMs are resolving that dilemma by defaulting to exploitation too early, locking in assumptions before accumulating enough evidence to justify them.

The researchers tried a direct countermeasure: they told the models to be fair. It did not change behavior. Instructing an LLM to avoid bias did not interrupt the pattern of group-based sorting, which suggests the bias is not located in the model’s stated values but in how it processes sequential feedback and builds generalizations from it. OpenAI and Anthropic did not respond to requests for comment on the findings.
Angelina Wang, a computer scientist at Cornell University who was not involved in the research, pointed to a specific feature trend that makes these findings harder to dismiss as a laboratory edge case. Chatbots are increasingly being built with persistent memory – the ability to draw on prior conversation history to personalize responses. That memory, designed to make interactions more useful, creates exactly the conditions for compounding bias. When a model draws on prior experience, Wang said, it can “over-index on the same kinds of behaviors it’s experienced before.” A model that formed a bias early in its memory window would carry that bias forward across every subsequent interaction it uses that history to inform. The hiring context makes this especially direct: AI agents operating in workplace settings are already being integrated into workflows where memory and personalization are selling points.
Why Reducing Memory Is Not a Simple Fix
The obvious response – strip the models of memory so they cannot form persistent patterns – runs into a user-experience problem. People want chatbots to remember what they have said. Memory is not a bug in the eyes of the companies building these products; it is a feature that drives adoption and retention. Wang framed the challenge directly: “We still are trying to figure out just the right amount that isn’t too much or too little.” There is no established threshold that keeps personalization intact while preventing the bias accumulation the Princeton and Chicago researchers documented.

That unresolved calibration problem sits at the center of a practical question facing any organization currently using AI to screen applicants. The models tested – ChatGPT, Claude, Gemini, o3 – are already deployed in real hiring pipelines. The fictional groups in the study stood in for real demographic categories. The feedback loop that drove segregation in the simulation mirrors exactly how a real model would process outcome data from actual hires: a pattern forms, it gets reinforced, and the model never questions whether its early sample was representative. At 1.83 on a 2-point segregation scale, o3 was one data point away from treating every demographic group as categorically unsuited for any role another group had claimed first.








