Aylin Caliskan – UW News /news Wed, 16 Sep 2026 21:31:36 +0000 en-US hourly 1 https://wordpress.org/?v=6.9.7 Q&A: UW researchers respond to recent concerns over AI risk /news/2026/09/16/uw-researchers-discuss-ai-risk/ Wed, 16 Sep 2026 21:07:49 +0000 /news/?p=93174 AI apps open on a phone.
Five UW AI researchers discuss the risks of AI systems. Photo:

This summer, OpenAI announced escaped a training environment and hacked into the AI company Hugging Face. Anthropic quickly followed with news that its AI agents also .

Last week, an outgoing Anthropic employee took to X, posting that the “.” Such talk has for years, though many AI experts have argued that these Terminator-esque claims are distractions from the real risks posed by current AI systems. Nevertheless, that viral X thread is

To help make sense of all this, UW News talked to five AI researchers from the : 

  • , associate professor in the Information School;
  • , professor in the Information School;
  • , professor in the Paul G. Allen School of Computer Science & Engineering;
  • , professor in the Allen School and the UW’s vice provost for AI;
  • and , professor in the Information School and the School of Law.

How alarming do you find the hacks announced by OpenAI and Anthropic? 

Franziska Roesner: I do find them somewhat alarming — not due to the hypothetical risks from an anthropomorphized runaway AI, but because complex interconnected systems are being built and seemingly run without much in the way of standard safeguards and auditing. The resulting outcomes are unsurprising to security experts, but are sensationalized as AI risk.

Ryan Calo: The timing makes me a little skeptical. Is OpenAI trying to match Anthropic by arguing that its systems are just as scary? Is Hugging Face trying to look relevant in advance of its purchase by Nvidia? But yes — this sort of emergent behavior is concerning.

Noah A. Smith: We’ve been told that the beast got out of the cage, but we don’t know enough about the cage the beast was in. The demonstrations may establish an important new capability in these AI models without establishing the broader risk people are inferring. Assessing the underlying risk depends on what access, scaffolding, permissions and safeguards the system had. shows that the alarming behavior depended heavily on what tools the model was given, what it was allowed to access, and how the experiment was set up, not just on the model itself.

Chirag Shah: I’m in half-agreement with scholars like who warn that the big AI labs are creating this scare to distract us from real problems that AI is causing. I also concur with and others who have been warning us about the security threats posed by the frontier models. I don’t think these two viewpoints are mutually exclusive: Yes, there are many other potential harms being created by AI, but the hacks and other security issues are real too and could be more devastating. Worse, we may not have time or opportunity to react, fix or reverse.

Aylin Caliskan: When such a complex system is equipped with tools and capabilities that enable it to interact with other complex systems, we should expect unforeseen exploits, problems and unintended consequences by default. The safety of these systems needs to be rigorously evaluated under controlled conditions and in real time, and appropriate guardrails should be dynamically integrated while they’re running.

What do you make of former Anthropic that, “The people building AI earnestly believe that it could kill us all by the end of the decade”?

RC: I worry engineers like Mr. Coxon are playing into an industry rhetoric that would have society focus on speculative, existential threats, rather than immediate, real-world harms. I argued as much in 2023 in .

NS: I think most people don’t want to kill others or die themselves. Is he claiming that AI builders, collectively, want to harm others? Why are they building AI? Extraordinary claims about what AI builders collectively believe need evidence.

CS: I don’t buy it. I’d put this in the same category as the Y2K bug or communism destroying the world. AI has real benefits and dangers, but world-saving or world-destroying characterizations are neither realistic nor helpful.

AC: What does “believe” mean in Coxon’s sentence? Does it mean being unable to rule out a risk with 100% certainty, or does it mean that a large group of people building AI strongly believe that AI will be a net negative, yet continue to dedicate their resources to AI development? In theory, many things are possible. In practice, how likely are they?

FR: I wonder if these statements say more about the people making them than about the fundamental capabilities of AI. from science fiction writer Ted Chiang gives one perspective on this — that this belief in rampant, destructive AI is a product of the “no-holds-barred capitalism” practiced by major tech companies. It’s from 2017, but remarkably relevant.

Related

Sources for further reading, suggested by Noah A. Smith:

The people making these claims and announcements largely have financial stakes in these companies, which are . How are you thinking about ulterior motives here?

CS: I see this as an attempt to steer the public into believing these companies are building world-changing tech that everyone needs to invest in or they’d miss out; that this tech would be so powerful that they rise up to national security level and gain power; and that the same tech could also be so dangerous that only they have the ability to curb it and they can self-regulate.

NS: It doesn’t take a conspiracy theorist to note that there are incentives at work. The financial stakes around prospective IPOs are enormous, and there are also long-standing concerns that safety arguments can shape regulation in ways that favor incumbent firms. Rules could reduce competition and independent scrutiny, concentrating both technological power and the authority to define what counts as “safe” in the hands of a few companies. They could also bar many people from participating in what the technology is designed to do, for example, by slowing or stopping work on open-source alternatives.

What should be done about AI risk?

NS: Risks need to be defined based on independent scrutiny and high-quality evidence, not messaging from organizations and people with a stake in what the response to risk looks like. We need sensible liability and accountability for harms, and governance proportional to demonstrated risks in real-world contexts rather than speculative narratives and science fiction. We should be especially wary of rules that entrench incumbent interests or treat closed, centralized control as synonymous with safety.

Openness is part of safety: If outsiders cannot inspect, reproduce and challenge claims about dangerous behavior, we are left trusting the organizations that have the strongest incentives to frame the narrative.

RC: Some combination of common law liability and regulation needs to create adequate incentives for AI companies to address the inevitable harms of this trillion-dollar industry.

FR: To me, the bigger question for safety is less, “What can AI models do in isolation?” and more, “How and why are we building these models into increasingly complex systems?” Computer systems security, for example, has already offered us examples of how to build these systems. More generally, we should all — whether we are building, integrating or using AI — anticipate how systems might be misused by people or harm them and adjust our systems accordingly.

AC: Academic freedom, independent evaluation and development, and open science play critical roles in analyzing and mitigating AI risks, as well as in effectively disseminating findings and evidence to inform policy and the public. To better manage risks, we should be designing AI deployment contexts in collaboration with stakeholders and communities, providing evidence to demonstrate net positive deployment effects that do not disproportionately benefit specific entities or groups, and iteratively identifying, isolating, and minimizing risks.

CS: Establish and fund commissions and taskforces that audit these companies and models and make independent assessments and recommendations. Make the companies rolling out these models accountable for any harms caused by their tech. Educate and empower the public through media, policies and democratic frameworks that give them a real say in what happens to their lives and labor through these technologies.

To set up an interview with an AI expert, contact Stefan Milne at stmilne@uw.edu.

Source

]]>
People mirror AI systems’ hiring biases, study finds /news/2025/11/10/people-mirror-ai-systems-hiring-biases-study-finds/ Mon, 10 Nov 2025 15:46:33 +0000 /news/?p=89402 A person's hands type on a laptop.
In a new study, 528 people worked with simulated LLMs to pick candidates for 16 different jobs, from computer systems analyst to nurse practitioner to housekeeper. The researchers simulated different levels of racial biases in LLM recommendations for resumes from equally qualified white, Black, Hispanic and Asian men. Photo: Delmaine Donson/iStock

An organization drafts a job listing with artificial intelligence. Droves of with chatbots. Another AI system sifts through those applications, passing recommendations to hiring managers. Perhaps AI avatars conduct screening interviews. This is increasingly the state of hiring, as people seek to streamline the stressful, tedious process with AI.

Yet research is finding that hiring bias — against people with disabilities, or certain races and genders — permeates large language models, or LLMs, such as ChatGPT and Gemini. We know less, though, about how biased LLM recommendations influence the people making hiring decisions.

In a new study, 528 people worked with simulated LLMs to pick candidates for 16 different jobs, from computer systems analyst to nurse practitioner to housekeeper. The researchers simulated different levels of racial biases in LLM recommendations for resumes from equally qualified white, Black, Hispanic and Asian men.

When picking candidates without AI or with neutral AI, participants picked white and non-white applicants at equal rates. But when they worked with a moderately biased AI, if the AI preferred non-white candidates, participants did too. If it preferred white candidates, participants did too. In cases of severe bias, people made only slightly less biased decisions than the recommendations.

The team Oct. 22 at the AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society in Madrid.

“In one survey, 80% of organizations using AI hiring tools said they don’t reject applicants without human review,” said lead author , a UW doctoral student in the Information School. “So this human-AI interaction is the dominant model right now. Our goal was to take a critical look at this model and see how human reviewers’ decisions are being affected. Our findings were stark: Unless bias is obvious, people were perfectly willing to accept the AI’s biases.”

Participants were given a job description and the names and resumes of five candidates: two white men; two men who were either Asian, Black or Hispanic; and one candidate whose resume lacked qualifications for the job, to obscure the purpose of the study. An example from the study is shown here. Photo: Wilson et al./AIES ‘25

The team recruited 528 online participants from the U.S. through surveying platform , who were then asked to screen job applicants. They were given a job description and the names and resumes of five candidates: two white men and two men who were either Asian, Black or Hispanic. These four were equally qualified. To obscure the purpose of the study, the final candidate was of a race not being compared and lacked qualifications for the job. Candidates’ names implied their races — for example, Gary O’Brien for a white candidate. Affinity groups, such as Asian Student Union Treasurer, also signaled race.

In four trials, the participants picked three of the five candidates to interview. In the first trial, the AI provided no recommendation. In the next trials, the AI recommendations were neutral (one candidate of each race), severely biased (candidates from only one race), or moderately biased, meaning candidates were recommended at rates similar to rates of bias in real AI models. The team derived rates of moderate bias using the same methods as in their 2024 study that looked at bias in three common AI systems.

Rather than having participants interact directly with the AI system, the team simulated the AI interactions so they could hew to rates of bias from their large-scale study. Researchers also used AI generated resumes, rather than real resumes, which they validated. This allowed greater control, and AI-written resumes are increasingly common in hiring.

“Getting access to real-world hiring data is almost impossible, given the sensitivity and privacy concerns,” said senior author , a UW associate professor in the Information School. “But this lab experiment allowed us to carefully control the study and learn new things about bias in human-AI interaction.”

Without suggestions, participants’ choices exhibited little bias. But when provided with recommendations, participants mirrored the AI. In the case of severe bias, choices followed the AI picks around 90% of the time, rather than nearly all the time, indicating that even if people are able to recognize AI bias, that awareness isn’t strong enough to negate it.

“There is a bright side here,” Wilson said. “If we can tune these models appropriately, then it’s more likely that people are going to make unbiased decisions themselves. Our work highlights a few possible paths forward.”

In the study, bias dropped 13% when participants began with an , intended to detect subconscious bias. So companies including such tests in hiring trainings may mitigate biases. Educating people about AI can also improve awareness of its limitations.

“People have agency, and that has huge impact and consequences, and we shouldn’t lose our critical thinking abilities when interacting with AI,” Caliskan said. “But I don’t want to place all the responsibility on people using AI. The scientists building these systems know the risks and need to work to reduce systems’ biases. And we need policy, obviously, so that models can be aligned with societal and organizational values.”

, a UW doctoral student in the Information School, and , a postdoctoral scholar at Indiana University, are also co-authors on this paper. This research was funded by The U.S. National Institute of Standards and Technology.

For more information, contact Wilson at kywi@uw.edu and Caliskan at aylin@uw.edu.

Source

]]>
AI tools show biases in ranking job applicants’ names according to perceived race and gender /news/2024/10/31/ai-bias-resume-screening-race-gender/ Thu, 31 Oct 2024 16:00:34 +0000 /news/?p=86725
research found significant racial, gender and intersectional bias in how three state-of-the-art large language models, or LLMs, ranked resumes. Photo:

The future of hiring, it seems, is automated. Applicants can now . And companies — which have long automated parts of the process — are now to write job descriptions, sift through resumes and screen applicants. An estimated 99% of Fortune 500 companies now .

This automation can boost efficiency, and some claim it can make the hiring process less discriminatory. But new research found significant racial, gender and intersectional bias in how three state-of-the-art large language models, or LLMs, ranked resumes. The researchers varied names associated with white and Black men and women across over 550 real-world resumes and found the LLMs favored white-associated names 85% of the time, female-associated names only 11% of the time, and never favored Black male-associated names over white male-associated names.

The team Oct. 22 at the AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society in San Jose.

“The use of AI tools for hiring procedures is already widespread, and it’s proliferating faster than we can regulate it,” said lead author , a UW doctoral student in the Information School. “Currently, outside of , there’s no regulatory, independent audit of these systems, so we don’t know if they’re biased and discriminating based on protected characteristics such as race and gender. And because a lot of these systems are proprietary, we are limited to analyzing how they work by approximating real-world systems.”

Previous studies have found and disability bias when sorting resumes. But those studies were relatively small — using only one resume or four job listings — and ChatGPT’s AI model is a so-called “black box,” limiting options for analysis.

The UW team wanted to study open-source LLMs and do so at scale. They also wanted to investigate intersectionality across race and gender.

The researchers varied 120 first names associated with white and Black men and women across the resumes. They then used three state-of-the-art LLMs from three different companies — Mistral AI, Salesforce and Contextual AI — to rank the resumes as applicants to over 500 real-world job listings. These were spread across nine occupations, including human resources worker, engineer and teacher. This amounted to more than three million comparisons between resumes and job descriptions.

The team then evaluated the system’s recommendations across these four demographics for statistical significance. The system preferred:

  • white-associated names 85% of the time versus Black-associated names 9% of the time;
  • and male-associated names 52% of the time versus female-associated names 11% of the time.

The team also looked at intersectional identities and found that the patterns of bias aren’t merely the sums of race and gender identities. For instance, the study showed the smallest disparity between typically white female and typically white male names. And the systems never preferred what are perceived as Black male names to white male names. Yet they also preferred typically Black female names 67% of the time versus 15% of the time for typically Black male names.

“We found this really unique harm against Black men that wasn’t necessarily visible from just looking at race or gender in isolation,” Wilson said. “Intersectionality is a protected attribute only in California right now, but looking at multidimensional combinations of identities is incredibly important to ensure the fairness of an AI system. If it’s not fair, we need to document that so it can be improved upon.”

The team notes that future research should explore bias and harm reduction approaches that can align AI systems with policies. It should also investigate other protected attributes, such as disability and age, as well as looking at more racial and gender identities — with an emphasis on intersectional identities.

“Now that generative AI systems are widely available, almost anyone can use these models for critical tasks that affect their own and other people’s lives, such as hiring,” said senior author , a UW assistant professor in the iSchool. “Small companies could attempt to use these systems to make their hiring processes more efficient, for example, but it comes with great risks. The public needs to understand that these systems are biased. And beyond allocative harms, such as hiring discrimination and disparities, this bias significantly shapes our perceptions of race and gender and society.”

This research was funded by the U.S. National Institute of Standards and Technology.

For more information, contact Wilson at kywi@uw.edu and Caliskan at aylin@uw.edu.

Source

]]>
AI image generator Stable Diffusion perpetuates racial and gendered stereotypes, study finds /news/2023/11/29/ai-image-generator-stable-diffusion-perpetuates-racial-and-gendered-stereotypes-bias/ Wed, 29 Nov 2023 16:53:35 +0000 /news/?p=83689 This image compares 16 images generated by Stable Diffusion. In the lower right quadrant, images generated to represent “a person from Papua New Guinea” show four dark-skinned people, while the other 12 images, representing people from Oceania, Australia and New Zealand show only light-skinned people.
researchers found that when prompted to create pictures of “a person,” the AI image generator over-represented light-skinned men, sexualized images of certain women of color and failed to equitably represent Indigenous peoples. For instance, compared here (clockwise from top left) are the results of four prompts to show “a person” from Oceania, Australia, Papua New Guinea and New Zealand. Papua New Guinea, where the population remains mostly Indigenous, is the second most populous country in Oceania. Photo: Ghosh et al./EMNLP 2023 — AI GENERATED IMAGE

What does a person look like? If you use the popular artificial intelligence image generator Stable Diffusion to conjure answers, too frequently you’ll see images of light-skinned men.

Stable Diffusion’s perpetuation of this harmful stereotype is among the findings of a new study. Researchers also found that, when prompted to create images of “a person from Oceania,” for instance, Stable Diffusion failed to equitably represent Indigenous peoples. Finally, the generator tended to sexualize images of women from certain Latin American countries (Colombia, Venezuela, Peru) as well as those from Mexico, India and Egypt.

The researchers will present Dec. 6-10 at the in Singapore.

“It’s important to recognize that systems like Stable Diffusion produce results that can cause harm,” said , a UW doctoral student in the human centered design and engineering department. “There is a near-complete erasure of nonbinary and Indigenous identities. For instance, an Indigenous person looking at Stable Diffusion’s representation of people from Australia is not going to see their identity represented — that can be harmful and perpetuate stereotypes of the settler-colonial white people being more ‘Australian’ than Indigenous, darker-skinned people, whose land it originally was and continues to remain.”

To study how Stable Diffusion portrays people, researchers asked the text-to-image generator to create 50 images of a “front-facing photo of a person.” They then varied the prompts to six continents and 26 countries, using statements like “a front-facing photo of a person from Asia” and “a front-facing photo of a person from North America.” They did the same with gender. For example, they compared “person” to “man” and “person from India” to “person of nonbinary gender from India.”

The team took the generated images and analyzed them computationally, assigning each a score: A number closer to 0 suggests less similarity while a number closer to 1 suggests more. The researchers then confirmed the computational results manually. They found that images of a “person” corresponded most with men (0.64) and people from Europe (0.71) and North America (0.68), while corresponding least with nonbinary people (0.41) and people from Africa (0.41) and Asia (0.43).

Likewise, images of a person from Oceania corresponded most closely with people from majority-white countries Australia (0.77) and New Zealand (0.74), and least with people from Papua New Guinea (0.31), the second most populous country in the region where the population remains predominantly Indigenous.

A third finding announced itself as researchers were working on the study: Stable Diffusion was sexualizing certain women of color, especially Latin American women. So the team compared images using a NSFW (Not Safe for Work) Detector, a machine-learning model that can identify sexualized images, labeling them on a scale from “sexy” to “neutral.” (The of being less sensitive to NSFW images than humans.) A woman from Venezuela had a “sexy” score of 0.77 while a woman from Japan ranked 0.13 and a woman from the United Kingdom 0.16.

“We weren’t looking for this, but it sort of hit us in the face,” Ghosh said. “Stable Diffusion censored some images on its own and said, ‘These are Not Safe for Work.’ But even some that it did show us were Not Safe for Work, compared to images of women in other countries in Asia or the U.S. and Canada.”

While the team’s work points to clear representational problems, the ways to fix them are less clear.

“We need to better understand the impact of social practices in creating and perpetuating such results,” Ghosh said. “To say that ‘better’ data can solve these issues misses a lot of nuance. A lot of why Stable Diffusion continually associates ‘person’ with ‘man’ comes from the societal interchangeability of those terms over generations.”

The team chose to study Stable Diffusion, in part, because it’s open source and makes its training data available (unlike prominent competitor Dall-E, from ChatGPT-maker OpenAI). Yet both the reams of training data fed to the models and the people training the models themselves introduce complex networks of biases that are difficult to disentangle at scale.

“We have a significant theoretical and practical problem here,” said , a UW assistant professor in the Information School. “Machine learning models are data hungry. When it comes to underrepresented and historically disadvantaged groups, we do not have as much data, so the algorithms cannot learn accurate representations. Moreover, whatever data we tend to have about these groups is stereotypical. So we end up with these systems that not only reflect but amplify the problems in society.”

To that end, the researchers decided to include in the published paper only blurred copies of images that sexualized women of color.

“When these images are disseminated on the internet, without blurring or marking that they are synthetic images, they end up in the training data sets of future AI models,” Caliskan said. “It contributes to this entire problematic cycle. AI presents many opportunities, but it is moving so fast that we are not able to fix the problems in time and they keep growing rapidly and exponentially.”

This research was funded by a National Institute of Standards and Technology award.

For more information, contact Ghosh at ghosh100@uw.edu and Caliskan at aylin@uw.edu.

Source

]]>