Workshop · 2012
Comparative Language Fuzz Testing: Programming Languages vs. Fat Fingers
Diomidis Spinellis, Vassilios Karakoidas, Panos Louridas
Ten programming languages, the same programs, and a fuzzer that makes small typo-like changes to the source — a way of measuring how much a language's design catches before a mistake becomes a wrong answer.
- Published in
- Proceedings of the Workshop on Evaluation and Usability of Programming Languages and Tools (PLATEAU) at the ACM Onward! and SPLASH Conferences, 2012
- Citations
- 5 on Google Scholar, read 5 September 2026 — 21 of 30 by count
- Cite as
- SKL12
The idea
Arguments about type systems usually run on anecdote. This work turns one of them into a measurement: take programs that solve the same tasks in ten languages, perturb their source the way a slipped keystroke would, and count what happens. A language that rejects the mutated program at compile time has protected the programmer; one that runs it and quietly returns the wrong answer has not.
The method
- A corpus from Rosetta Code: fourteen diverse tasks, each implemented in all ten languages — C, C++, C#, Haskell, Java, JavaScript, Perl, PHP, Python and Ruby.
- A mutation-based fuzzer applying five kinds of change: identifier substitution, integer perturbation, random character substitution, similar-token substitution and random-token substitution.
- Each mutated program compiled and run, and the outcome classified — rejected, crashed, or ran and produced wrong output.
What it found
Languages with weak type systems are significantly likelier than languages enforcing strong typing to let a fuzzed program compile, run, and produce an erroneous result. The broader claim is methodological: comparative fuzz testing is a usable way to evaluate programming language designs empirically.
The harness is on this site as the fuzzing harness.
Written from the authors' HTML copy of the paper; the published version is behind the DOI above and is not hosted here.