Conference · 2014
The Bug Catalog of the Maven Ecosystem
Dimitris Mitropoulos, Vassilios Karakoidas, Panos Louridas, Georgios Gousios, Diomidis Spinellis
A dataset of FindBugs results for every version of every Java project in Maven Central — 115,214 JARs, about 265 GB of input — published so that bug-related questions about the ecosystem can be asked without repeating the analysis.
- Published in
- MSR 2014, The 11th Working Conference on Mining Software Repositories, 2014
- Citations
- 52 on Google Scholar, read 5 September 2026 — 2 of 30 by count
- Cite as
- MKLGS14
The idea
A software ecosystem is a collection of projects that evolve together, with interdependencies and many versions each. Running a static analyser over all of it — not the latest release of each project, but every version — produces something no single project's data can: a record of how defects appear and persist across an entire ecosystem's history.
How it was built
- A January 2012 snapshot of Maven Central, cross-checked later that year against the live repository so that versions released in the meantime were not missing. Projects in other JVM languages (Scala, Groovy, Clojure) were filtered out, since FindBugs analyses Java bytecode.
- Distributed processing: tasks on a RabbitMQ queue, twenty-five Python workers checking them out, running FindBugs and storing the results. The tooling is 14 Python scripts, 1,256 lines.
- For every JAR, the metrics FindBugs reports plus metadata — the JAR's size, its dependencies and more.
What is in it
17,505 projects and 115,214 versions: a median of 3 versions per project, a mean of 6.58, and a maximum of 338. The paper works through example analyses to show what the data supports.
The dataset has a page here — the Maven bug catalog. Its security-focused sibling is the vulnerability dataset, and the study built on it is Dismal Code.
Written from the paper itself — the PDF linked above, which this site hosts.