Research

#Mechanistic-Interpretability

Why does an LLM sometimes conclude that a correct answer it just produced is incorrect?

Why does an LLM sometimes conclude that a correct answer it just produced is incorrect?