AI Chatbots Generate Weaker Work Emails for Women's Language
Johns Hopkins researchers find ChatGPT and other models produce less sophisticated responses when prompts contain linguistic patterns commonly used by women.
When professionals use language patterns commonly associated with women to prompt AI chatbots for workplace writing, they receive noticeably inferior output—less sophisticated, less formal, and written at a lower grade level than responses to male-coded language.
That's the finding from new Johns Hopkins University research testing four popular AI systems: GPT-4, Llama, Gemma, and Mistral. The study will be presented at the Conference on Language Modeling in San Francisco on October 6-9.
How the bias manifests
Researchers took real chatbot prompts for workplace correspondence—emails, job applications, resignation letters—and modified them to include linguistic features documented as more common in women's speech. These included hedging words like "maybe" and "I think," collective phrasing such as "we" and "our team," and expressive adjectives like "lovely" and "wonderful."
Every AI system tested consistently returned weaker professional writing when these features appeared in prompts. The gap persisted even after researchers controlled for the writer's tone, indicating the models weren't simply matching the prompt's style.
In one test, two prompts requested responses to a thank-you email. The male-coded prompt generated: "I am writing to acknowledge your recent email expressing your gratitude. I sincerely appreciate your kind words and the time you took to write to me."
The female-coded prompt produced: "We were absolutely delighted to receive your wonderfully appreciative email earlier. Your words of praise and acknowledgment have indeed warmed our hearts and brought immense satisfaction to our team."
Names made no difference
In a striking finding, adding traditionally male or female names to prompts had virtually no effect on output quality. A female-coded prompt signed "John" still generated the same lower-quality response—it simply ended with "John."
"We thought if you ask the AI for an email with a women-associated linguistic prompt, but sign it 'John,' the model would pick up on the 'John' more strongly," said lead author Katherine Van Koevering, a postdoctoral fellow with Johns Hopkins' Data Science and AI Institute. "But no, you get the same response and it just says John at the end."
Why it matters
As AI tools become standard for professional communication, these biases could systematically disadvantage women and anyone whose natural speech patterns include these linguistic features. The language patterns in question are largely unconscious and extremely difficult to change—meaning the burden falls on AI companies to fix their models rather than on users to alter their communication style.
The problem may intensify as voice-based AI interactions become more common, making gendered speech patterns even more pronounced and harder to mask.
"If you prompt a model to write an email you're going to send to someone else at your company, and you're using language features that women more commonly use, you'll get back a response that's less complex, at a lower grade level, and less formal," said senior author Anjalie Field, a Johns Hopkins computer scientist studying ethics and discrimination in AI. "That's going to reflect on how the recipient of that document perceives you."
The research team plans to investigate whether similar effects appear across other demographics including age, race, and ethnicity. They also want to study whether AI users gradually adapt their communication style to match what the systems reward.
The findings were first reported by Johns Hopkins University.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call