GSM8K and the Limitations of LLMs
October 15, 2024 · AI

Recently, I bumped into the following paper by Apple engineers, titled:
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Note: The paper can be found here.
Initially, I thought that it was, yet another LLM-cannot-reason paper, which actually is, but I liked very much the approach of the authors.
In short, they take the GSM8K benchmark and test two variations a test them in many “edge” LLM models. The results show that for these two, all the models have inferior performance, thus they cannot really reason, making the statement that the results lie heavily on the prompts that are used. Take also into consideration that the author play fair and they do not try cheap tricks on LLMs.
With these two variations they introduce a new benchmark, named GSM-Symbolic. So, let’s take a quick look at the results.
Template Generation
Their variation included problem generation through templates e.g. they switched names on variables and values. For example,

On the left is the initial problem form and on the right the template, which replaces formulates several variations of the problem. With this little change, all models dropped in accuracy.

The GPT-4o had the smallest drop, around 0.3% and Mistral 9.2% (!). All models dropped a little, but lets say that some of the flagship models have not significant drops. But on the other hand, that 2.2% of o1-preview, really hurt my feelings.
The No-Op Variation & Symbolic Difficulty
The second variation was by introducing several statements inside the problem definition that were valid, but with no meaning at all, ones that a human would not take into consideration to solve the problem. For example:

You can see a sample problem that has that irrelevant statement and the answers from o1-mini and Llama3–8B.
The authors also provided problems with increased symbolic difficulty:

These problems are variations that require more advanced reasoning to solve them.
The results were dissapointing:

The accuracy drop was 17.5% for the o1-preview, which was the best model, but still it is a slaughter.
Takeaways
It seems that a lot of the accuracy results you can find on the internet regarding top models are there for marketing reasons. Of course, these models are very capable of doing lots of things, but right now, only with human supervision, as assistants.
We have a lot way to go still, these models get better with each iteration and we will see how it ends.
If you liked this article, see also: