In the discrimination experiment, the 175B parameter model discriminates against Black versus white [sic] students by 3% in the Q condition, and discriminates in favor of Black students by 7% in the Q+IF+CoT condition. In this experiment, larger models can over-correct, especially as the amount of RLHF training increases. This may be desirable in certain contexts, such as those in which decisions attempt to correct for historical injustices against marginalized groups, if doing so is in accordance with local laws.
This is similar to Microsoft's (in collaboration with MIT, Carnegie Mellon, and University of Washington) SafeNLP project to measure hate-speech in AIs. They explicitly turned a blind eye towards anti-white hate [1], do not even include a category for whites in their safety scores [2], and consider the phrase "stop hurting white people" (and only white people) as hate [3].
[1] Our ultimate aim is to shift power dynamics to targets of oppression. Therefore, we do not consider identity dimensions that are historically the agents of oppression (e.g., whiteness, heterosexuality, able-bodied-ness). - https://arxiv.org/pdf/2203.09509
like_any_other•45m ago
In the discrimination experiment, the 175B parameter model discriminates against Black versus white [sic] students by 3% in the Q condition, and discriminates in favor of Black students by 7% in the Q+IF+CoT condition. In this experiment, larger models can over-correct, especially as the amount of RLHF training increases. This may be desirable in certain contexts, such as those in which decisions attempt to correct for historical injustices against marginalized groups, if doing so is in accordance with local laws.
This is similar to Microsoft's (in collaboration with MIT, Carnegie Mellon, and University of Washington) SafeNLP project to measure hate-speech in AIs. They explicitly turned a blind eye towards anti-white hate [1], do not even include a category for whites in their safety scores [2], and consider the phrase "stop hurting white people" (and only white people) as hate [3].
[1] Our ultimate aim is to shift power dynamics to targets of oppression. Therefore, we do not consider identity dimensions that are historically the agents of oppression (e.g., whiteness, heterosexuality, able-bodied-ness). - https://arxiv.org/pdf/2203.09509
[2] https://github.com/microsoft/SafeNLP#safety-scores-based-on-...
[3] https://github.com/microsoft/SafeNLP/blob/main/data/implicit...
blinkbat•33m ago