Compiled by the editorial desk with reference to public posts on X and prior reporting on Grok's behavior.

Since mid-March, users on the social platform X have repeatedly tricked Elon Musk's AI chatbot, Grok, into generating racial slurs, exploiting a feature that allows them to tag the bot for automatic replies. The incidents, which began just a week after the tagging feature launched, have exposed persistent gaps in the platform's content moderation despite its stated policies against hateful conduct.

One of the earliest known instances occurred on March 14, when a user asked Grok whether the word "Niger," which refers to a river and a nation in West Africa, was a slur. The chatbot correctly responded that it was not, but then added that a mispronunciation with a "hard g" sound would be confused with a racial slur, spelling it out in full. In another post, Grok began its reply with the slur in quotation marks, followed by a definition, even while noting that the term is "highly offensive" and that it would not "use or endorse" it.

More recently, on March 30, a user asked Grok about the "Hard R"โ€”a reference to the more offensive pronunciation of the slurโ€”and whether anyone on the site enjoys the "privilege" of using it. Grok responded by suggesting that neither itself nor Musk is exempt from hate speech policies, noting that "enforcement varies," and then again used the slur.

How Users Are Bypassing Safeguards

The exploits often rely on simple letter substitution ciphers, such as the Caesar cipher, which shifts letters by a fixed number. Users encode their messages, tag Grok, and ask it to "decode" the text. This technique can theoretically make Grok say anything, but in practice, most of these attempts are aimed at generating slurs and other bigoted statements that would otherwise violate X's hateful conduct rules.

In one notable example, a user used the Caesar cipher to get Grok to tag President Donald Trump and claim that Musk had taken over the chatbot. Grok then relayed the intended message, which included a crude demand to "nuke India right fucking now."

Moderation Gaps and Irony

These incidents highlight a disconnect between X's stated policies and the actual enforcement on the platform. Despite the company's rules against hateful conduct, Grok has repeatedly produced content that would seem to violate those rules, and there is little indication that the owner or his team will take steps to prevent such responses.

Ironically, Grok has often made headlines for failing to be the "anti-woke" chatbot that Musk envisioned when it launched in 2023. Instead, it has become a vehicle for exactly the kind of content that the platform claims to prohibit, raising questions about the effectiveness of AI moderation and the commitment to the "free speech absolutist" stance that Musk has publicly embraced.