ChatGPFUT
This from Carlos Velazquez:
The Media Research Center recently published on X a striking test of several AI systems including the graphic below. MRC asked ChatGPT to rate, from 1 to 10, how much of a threat President Trump represents to democracy.
According to MRC, ChatGPT answered 9, then placed several authoritarian leaders at the same level, while assigning Hitler a 10.
OpenAI, ChatGPT parent company, subsequently told Fox News that the response "shouldn't happen and is not how the model is trained to behave,” and that ChatGPT is designed to be objective by default.
MRC deserves credit for publishing considerably more about its methodology than just the provocative result.
Researchers say they used a clean computer, a separate email account, a VPN and a private browser window to minimize previous browsing history and cookies. They also published the question they say they asked.
But there is an important lesson here about evaluating any experiment involving artificial intelligence.
AI responses can be highly sensitive to context. What was said earlier in a conversation, instructions telling the AI to adopt a particular viewpoint or role, personalization, memory, and even relatively subtle wording can affect an answer. Consequently, when somebody reports a startling AI response, the ideal evidence is not simply the question and answer, but the complete conversation and the conditions under which the test was conducted.
I am not suggesting that MRC manipulated this test. I have seen no evidence that it did. In fact, the precautions MRC describes are evidence that its researchers were attempting to control some outside influences. My point is broader: AI is unusually easy to influence through context, which makes complete experimental transparency particularly important when testing it for political bias.
Interestingly, asking ChatGPT essentially the same question today produces a very different result: it refuses to assign a numerical political "threat” score at all. MRC itself subsequently repeated its test and reported that ChatGPT would no longer provide the numerical rating.
That doesn't prove MRC's original result was wrong. It demonstrates something equally important: a single AI response is a data point, not necessarily a description of the AI system itself. Reproducibility, complete context and methodology matter enormously when we try to determine whether we have discovered systematic AI bias or simply an anomalous response.
Posted by: Timothy Birdnow at
12:18 PM
| No Comments
| Add Comment
Post contains 382 words, total size 3 kb.
23kb generated in CPU 0.0833, elapsed 0.299 seconds.
35 queries taking 0.2862 seconds, 186 records returned.
Powered by Minx 1.1.6c-pink.