The Battleship Potemkin
June 28, 2025 · AI
Another chapter in the “LLM cannot reason saga”

Today, I read this post on mastodon from my PhD advisor Diomidis Spinellis, regarding a paper called,
Potemkin Understanding in Large Language Models
which introduces the phenomenon of “Potemkin understanding” in large language models (I will explain later on what this actually is). After my previous post on The Strawberry Problem and the recent paper from Apple regarding LLM reasoning capability collapse, I decided to take a closer look :).
So, the paper introduces and investigates the phenomenon of “Potemkin understanding” in large language models. The authors define this as the illusion of conceptual understanding, where an LLM can pass benchmark tests but misunderstands concepts in ways no human would. The research argues that benchmarks designed for humans are only valid for LLMs if the models misunderstand concepts in a manner similar to humans. When this is not the case, an LLM can correctly answer “keystone” questions — those that would typically confirm understanding in a human — while lacking a true grasp of the concept.
The paper presents two procedures to quantify the prevalence of Potemkin understanding:
A new benchmark dataset: This dataset spans 32 concepts across three domains: literary techniques, game theory, and psychological biases. The authors test seven different LLMs on their ability to first define a concept (the “keystone” task) and then apply it in classification, generation, and editing tasks. The “Potemkin rate” is the measure of how often a model fails at application tasks after correctly defining the concept.
An automated evaluation procedure: This general method provides a lower-bound estimate of Potemkin understanding. It works by first prompting an LLM to answer a question correctly. Then, it asks the same LLM to generate related questions and answer them. Finally, it prompts the LLM again to grade its own generated answer. Disagreement between the generated answer and the final grade indicates a Potemkin instance. This procedure also measures “incoherence,” where an LLM’s own evaluation of its output is inconsistent.
In order to visually understand it, the authors have some very nice set of illustrations:

The key findings of the study are:
Potemkin understanding is widespread: The study finds high rates of Potemkin understanding across all tested models, tasks, and domains. While models correctly defined concepts 94.2% of the time, their performance dropped significantly when required to apply these concepts.
Models exhibit internal incoherence: The research reveals that these failures are not just due to incorrect understanding but also to internal incoherence in the models’ representation of concepts. The automated procedure identified high rates of incoherence across models.
Current benchmarks may be invalid for LLMs: The high prevalence of Potemkin understanding suggests that standard benchmarks, often composed of keystone questions designed for humans, are not reliable measures of true conceptual understanding in LLMs. Success on these benchmarks does not guarantee that the model will perform well on other related tasks.
Epilogue
Ιn conclusion, the paper argues that Potemkin understanding is a significant and overlooked failure mode in LLMs. The authors propose that identifying these instances is crucial for developing more reliable models and for re-evaluating the validity of current benchmarking practices. They suggest that future work should focus on developing methods to detect and mitigate Potemkin rates in model training and evaluation.
So, it seems that in terms of reasoning, there are numerous papers and approaches that exhibit limitations in certain capabilities of the LLMs and this current stream of technology. But things are moving on, very large investments are in place and the game is still on. We will see what happens.