Fun, though as hinted at the end, the point of LLM "truth" probes is to measure the model's internal judgment of truthfulness. There's no reason this judgment, even if measured with 100% accuracy, couldn't be mistaken or logically inconsistent.
Vecr•13m ago
Someone at MIRI must know how to solve this. Good luck getting them to tell you how!
aesthesia•31m ago