All publications

Workshop · 2012

Comparative Language Fuzz Testing: Programming Languages vs. Fat Fingers

Diomidis Spinellis, Vassilios Karakoidas, Panos Louridas

Ten programming languages, the same programs, and a fuzzer that makes small typo-like changes to the source — a way of measuring how much a language's design catches before a mistake becomes a wrong answer.

Published in
Proceedings of the Workshop on Evaluation and Usability of Programming Languages and Tools (PLATEAU) at the ACM Onward! and SPLASH Conferences, 2012
Citations
5 on Google Scholar, read 5 September 2026 — 21 of 30 by count
Cite as
SKL12

The idea

Arguments about type systems usually run on anecdote. This work turns one of them into a measurement: take programs that solve the same tasks in ten languages, perturb their source the way a slipped keystroke would, and count what happens. A language that rejects the mutated program at compile time has protected the programmer; one that runs it and quietly returns the wrong answer has not.

The method

  • A corpus from Rosetta Code: fourteen diverse tasks, each implemented in all ten languages — C, C++, C#, Haskell, Java, JavaScript, Perl, PHP, Python and Ruby.
  • A mutation-based fuzzer applying five kinds of change: identifier substitution, integer perturbation, random character substitution, similar-token substitution and random-token substitution.
  • Each mutated program compiled and run, and the outcome classified — rejected, crashed, or ran and produced wrong output.

What it found

Languages with weak type systems are significantly likelier than languages enforcing strong typing to let a fuzzed program compile, run, and produce an erroneous result. The broader claim is methodological: comparative fuzz testing is a usable way to evaluate programming language designs empirically.

The harness is on this site as the fuzzing harness.

Written from the authors' HTML copy of the paper; the published version is behind the DOI above and is not hosted here.