Neal Agarwal's Absurd Trolley Problems inspired me to go a little less absurd, and also to see how frontier and open weight models would choose in high stake and social stake moral dilemmas.
First I had to create a new set of problems, which I did with my "Trolley Game"... where AI tries to predict your choices based on a handful of warm-up questions (77% accuracy right now). Then I ran 14 models through the full set of 20 dilemmas, 200+ times each, asking them for rationale and predictions about what humans would choose along the way.
Reasoning on; reasoning off. Reversal of choices to test for primacy effect.
Conclusion? We (humans) should be careful about how much control we hand over in terms of tool use in the future... because (1) like humans, AI models don't agree on everything, (2) they are not aligned with us in many ways (depending on your POV, of course), and (3) they "see" us as quite predictable creatures.
LanceJones•39m ago
First I had to create a new set of problems, which I did with my "Trolley Game"... where AI tries to predict your choices based on a handful of warm-up questions (77% accuracy right now). Then I ran 14 models through the full set of 20 dilemmas, 200+ times each, asking them for rationale and predictions about what humans would choose along the way.
Reasoning on; reasoning off. Reversal of choices to test for primacy effect.
Conclusion? We (humans) should be careful about how much control we hand over in terms of tool use in the future... because (1) like humans, AI models don't agree on everything, (2) they are not aligned with us in many ways (depending on your POV, of course), and (3) they "see" us as quite predictable creatures.