> users cannot easily avoid them through strategic self-presentation
Prompting LLMs differently than you talk to humans doesn't really seem that hard. I already do this (e.g. ask basic questions in a separate chat so I'll look smart, and get better responses, in the main session.)
Disclaimer: only read the abstract, feel free to point out if I missed the point.
The weird framing of this being a negative thing toward women is the personal bias of the women who published this and has no place being in this study. The measurement of what constitutes a response as "high quality" is also open to interpretation and varies depending on personal preference. You can't argue that a shift in the direction of the metrics mentioned in the report are objectively better or worse, they're just different.
https://shelflovepodcast.substack.com/p/actually-romance-nov...
> Through the perturbation of patient messages, we evaluate whether LLM behavior remains consistent, accurate, and unbiased when non-clinical information is altered. […] Our findings reveal notable inconsistencies in LLM treatment recommendations and significant degradation of clinical accuracy in ways that reduce care allocation to patients. […] Our perturbations reflect realistic patient messages from electronic formatting errors and/or simulate patient groups that would be impacted by a wide adoption of patient-AI systems (female patients, non-binary patients or those who use gender-neutral pronouns, patients with health anxiety, patients with a more dramatic disposition, patients with less technological aptitude, and patients with limited English proficiency, etc.)
We’re all peering down the kaleidoscope of a trillion parameter model. It’s no surprise gentle nudges in inputs (grammar, language proficiency, cultural norms) yield different outcomes, despite the intent not changing. It’s one thing to generate crap code, it’s another to generate crap medical advice.
perching_aix•29m ago
For example, in prompts, I (male) heavily use language that they attribute to women:
> Women’s language is more likely to include hedges (e.g., maybe, I think), tag questions (e.g., isn’t it?), collective reference (e.g., we, our), and expressive adjectives (e.g., lovely, wonderful).
But there are subtle ways in which their Figure 1 example prompt goes way beyond this, and that blatantly derails the entire thing:
> Let’s compose an email together to arrange our mid-year appraisal with our team
This is not about saying "our appraisal" or "our team", but about literally asking for a collaborative workflow ("Let's compose an email together"), rather than for a draft.
The "male" prompt in that Figure 1 comparison was also weird ("your team"), but alas.
nullbio•10m ago