GPT-4, Llama, Gemma and Mistral Respond Differently to “Male” and “Female” Prompt Styles

Researchers at Johns Hopkins University have found that large language models can shape professional texts differently depending on the user’s linguistic style. Prompts containing features that appear statistically more often in women’s speech produced shorter, less complex and less formal responses.
The work is titled It’s How You Ask: Gender-Associated Linguistic Bias in LLMs. The authors tested GPT-4, Llama, Gemma and Mistral on three types of work texts – emails, job applications and resignation letters.
Importantly, the researchers did not simply divide prompts into those written by men and those written by women. They used linguistic features that are statistically associated with different genders in American English.
The “female” versions included hedging expressions such as “possibly” and “I think,” collective phrasing like “we” and “our team,” and more emotional adjectives. The opposite versions used features that appear more often in men’s speech.
The difference showed up in all four models
The result held consistently across all the systems tested. Prompts with female linguistic markers produced responses with lower formality, shorter length and less complex language structure.
Prompts with features more often linked to male speech, by contrast, produced longer and more formal texts.
The authors separately tested an obvious explanation – whether the models were simply mimicking the user’s own style. But the effect persisted even after statistically accounting for the complexity of the original prompt and the carryover of its linguistic features into the response.
At the same time, simpler or less formal text is not necessarily objectively worse in every situation. The problem the researchers point to arises specifically in professional contexts, where writing style can shape how an employer or colleagues perceive the author.
Johns Hopkins University notes that this effect could potentially affect not only women but anyone whose manner of speech carries the same linguistic features.
A male or female name changed almost nothing
In another experiment, identical prompts were signed with traditionally male and female names.
This produced virtually no noticeable effect. The way a prompt was phrased influenced the result far more than a direct gender signal in the name.
The researchers consider this especially important, since many features of everyday speech are used unconsciously and are shaped by culture and social environment. So the advice to simply “write prompts differently” does not fully solve the problem.
The authors stress that the work does not prove deliberate “discrimination against women” by ChatGPT or other models. It demonstrates a narrower effect: LLMs can pick up on indirect gender-associated linguistic cues and systematically change the characteristics of the professional text they generate.
The preprint was published in August 2026. The work has been accepted to the Conference on Language Modeling (COLM 2026), which will take place October 6–9 in San Francisco. The authors name checking similar effects for age, race and ethnicity as the next direction.