This article was prepared from the source material provided, which cites yet-unpublished research from the University of Washington as reported by New Scientist. No original reporting or fact-checking was conducted.

Artificial intelligence tools designed to detect hate speech online, including Google's Perspective algorithm, show a significant bias against Black users, according to yet-unpublished research from the University of Washington reported by New Scientist.

The researchers examined how humans annotated a database of more than 100,000 tweets used to train anti-hate speech algorithms. They found that annotators tended to flag tweets written in African-American Vernacular English (AAVE) as offensive, a bias that then propagated into the algorithms themselves.

The team confirmed the bias by training several AI systems on the database, finding that the algorithms associated AAVE with hate speech. They then tested algorithms, including Perspective, on a database of 5.4 million tweets whose authors had disclosed their race. The algorithms were between one-and-a-half to twice as likely to flag posts written by people who identified as African-American as toxic, New Scientist reported.

Why the Bias Matters

Automated content moderation tools will likely remove many benign posts based on the ethnicity of their authors, leading to the silencing and suppression of certain communities online. The findings highlight how a well-intentioned effort to make the internet safer could end up discriminating against already-marginalized groups.

The research has not yet been published, and the full methodology and peer review are pending. The University of Washington scientists have not publicly commented beyond the New Scientist report. Google has not issued a statement on the findings.

The study adds to a growing body of evidence that algorithmic bias can emerge from human-generated training data, even when the goal is to protect vulnerable users from abuse.

Editorial Notes