← All blog posts

The static analysis toolbox in 2026

September 2, 2026 · Tools & Systems

A dense dependency graph drawn across four layers separated by dashed lines. On the left, one large dark node with edges converging on it from every layer. On the right, three highlighted nodes joined by arrows into a closed loop whose last edge crosses a layer boundary upward.
Three things a dependency graph will show you: a component that everything depends on (left), a cycle whose last edge climbs back up through a layer (right), and layer lines that no tool can confirm anyone actually intended.

I spent a good part of my PhD years on static analysis and software metrics. Back then, running a detector over a large Java system was something you kicked off before going home (and hoped had finished by the morning, because quite often it had run out of heap around 3am). In 2026 the same analysis is a pipeline step that finishes before the coffee is ready.

So we won, more or less. The tools are fast, most of them are free, and you can run four of them on every pull request without anyone noticing the bill.

And yet almost every team I talk to tells me the same story; thousands of findings, nobody reads them, the gate is set to "warn" and everybody moved on with their lives.

I do not think the tools are the problem. I think we bought them without deciding what question we wanted answered.

Three questions, not one

"Static analysis" is one phrase covering three different jobs. From the outside they look the same (a tool reads your code, a tool prints a list) and they do not substitute for each other at all.

  • Is this code wrong? A null dereference, a file that is never closed, an unsanitised string reaching a query, a race. The answer is local and binary; a finding is a defect or a false positive, and usually you can tell in a minute.
  • Is this class badly designed? Nothing is wrong here. It compiles, the tests pass, and the class is 900 lines that spend their time touching everybody else's data. It is not broken, it is expensive (you pay every time you change it).
  • Does this dependency belong where it is? Also nothing wrong. It is a perfectly legal import. It just runs from the domain layer straight into persistence, across a line somebody drew on a whiteboard two years ago and that exists nowhere in the code.

Why does the distinction matter, if you are going to run the tools anyway? Because the three do not correlate, and this is not my opinion. Fontana and colleagues checked it properly in 2019, on a decent sample, and architectural smells turn out to be essentially independent from code smells.

In plain words; you can clean up every God Class in the system and still have all of your cycles. The dashboard improves. The system stays exactly as hard to change as it was.

So one tool will not do it. But (and this is the part everybody skips) you do not need eleven either. One per question, and know which question you are in when a finding shows up.

What is actually in the box

Is it wrong. Start with the boring answer, which is the type checker. mypy or pyright for Python, strict mode in TypeScript, Error Prone for Java, staticcheck for Go, Clippy for Rust.

These are the cheapest to run, the least noisy, and the ones your engineers already trust (which matters far more than people admit; a tool nobody believes is a tool nobody runs).

Then the pattern engines. Semgrep is the one I would learn if I only learned one, because a rule there is a piece of code with holes in it, rather than an AST visitor you have to write in Java. There is a community fork, OpenGrep, after the 2025 licence change, so decide which one you are on before you write two hundred rules.

At the heavy end sit CodeQL, which turns your codebase into a database you write queries against, and Infer, which reasons across procedures and finds things nothing else finds. Both are excellent. Both are a project, not an afternoon.

Is it badly designed. PMD, SpotBugs, the maintainability half of Sonar. The good idea here is older than all three and still under-used; Marinescu called it a detection strategy. One metric is too fine-grained to name a design problem, so you compose several until the conjunction means something.

PMD's God Class rule is exactly that: weighted methods per class over 47, more than five accesses to foreign data, cohesion under a third. A class that does a lot, constantly reaches into other objects, and whose own methods share almost no state. That is a sentence a person can act on, which was always the point.

About that 47. It is not a law of nature, it is a percentile from a Java corpus that is roughly twenty years old (so, from before lambdas, streams and records, which is to say before half of what your code is made of).

Thresholds were the known weakness of this approach in 2006 and they still are, which is why two detectors will cheerfully disagree about the same class. These are not verdicts. They are a shortlist for a human, and that is all they were ever meant to be.

Does the dependency belong. This is the question most teams have nothing at all for. ArchUnit for Java, import-linter for Python, dependency-cruiser for JavaScript, Konsist for Kotlin; you write the rule as a test and it fails in CI like any other test.

On the smell side there is Arcan, Designite, jQAssistant, which read the component graph and name the cycles, the god components, the unstable dependencies.

One thing to be clear about here, because the vendors are not. A tool can reverse-engineer your dependency graph every night, and that is genuinely useful, but the graph is a measurement, not an architecture. Nothing in it knows what you meant.

The allow-list, the layer definition, the "this package may not import that one" is a decision, and a human has to write it down. A beautiful generated map that nobody ever approved is a picture of what happened to you, not of what you chose.

Wiring it up

Four things that, from experience, separate a toolbox from a wall of noise:

  • Emit SARIF and forget about dashboards. The quiet win of the last few years. Every serious analyser now speaks one output format, so five tools land in one place and you can diff two runs. Before that, each tool arrived with a portal nobody logged into.
  • Gate on the diff, not the repository. Baseline what exists today, accept it, fail only on what is new. On a codebase older than five years this is the only version of adoption that works (the alternative is 12,000 open issues and a team that has learned to ignore the gate).
  • One tool per question. Two analysers with overlapping rules is how you get four findings for one problem, and an engineer who stops reading them.
  • Keep the architecture rules in the repository, as tests. Not in a modelling tool, not in a wiki diagram from 2023. When somebody breaks a layer it should read as a failing test with a name, not a bot comment linking to a portal.

And then the agents happened

All of the above was true in 2016 as well. What changed is the ratio.

Code now arrives a lot faster than anybody can read it, and an agent optimises for the thing it can see, which is the test going green. The test does not see an illegal import. It does not see a five-line block cloned for the fourth time, and it certainly does not see a new cycle between two components.

The obvious answer is to make the reviewer work harder. It does not scale, because reading was always the slow part and it has not got faster since 1998 (I have tried; the eyes are the same eyes).

The better answer is to move the analyser to the other side. Hand the checks to the agent and let it fix its own violations before a human opens the diff; type checker, one pattern engine, the architecture test file. If it cannot get past those three, there is nothing for me to review yet.

That also changes what the rules are. An allow-list used to be documentation that quietly drifted. Now it is a specification that something actually reads, on every single run.

Just watch that it does not freeze, because a layer definition nobody is allowed to change is a layer definition that somebody eventually deletes at 2am :)

Epilogue

ISO/IEC 25010 still names the target honestly enough; modularity, reusability, analysability, modifiability, testability. We have far more machinery for those five words than we have a theory of them, and I do not see that gap closing. We can measure a hundred things about a system and still not say, in a way I would defend at a viva, what makes it easy to change.

So my advice is boring, and I am at peace with that. Decide which of the three questions you are asking. Pick one tool for it, gate it on the diff, and put it in front of the agent rather than behind it.

And keep the number of analysers below the number of engineers who will actually read their output. That ratio, far more than any threshold, decides whether any of this survives the quarter.

Reading

  • R. Marinescu. Detection strategies: metrics-based rules for detecting design flaws. ICSM 2004. (ICSME most influential paper, 2014.)
  • M. Lanza and R. Marinescu. Object-Oriented Metrics in Practice. Springer, 2006.
  • J. Garcia, D. Popescu, G. Edwards and N. Medvidovic. Toward a catalogue of architectural bad smells. QoSA 2009.
  • F. Arcelli Fontana et al. Are architectural smells independent from code smells? An empirical study. Journal of Systems and Software, 2019.
  • ISO/IEC 25010:2023, product quality model, maintainability.
  • Tools mentioned: Error Prone, staticcheck, Clippy, mypy, pyright, Semgrep / OpenGrep, CodeQL, Infer, PMD, SpotBugs, Sonar, ArchUnit, import-linter, dependency-cruiser, Konsist, Arcan, Designite, jQAssistant.