SDriver/XPath, seventeen years on: the same overhead
October 4, 2026 · Projects & Releases
For years my software page has had an entry called SDriver/XPath sitting in the Deprecated section, tagged "abandonware", with no page of its own (lately, a link to its paper). That was a fair description. The code was last built in 2009, with Ant, for Java 1.6, against a native MD5 library that shipped binaries for PowerPC Macs (yes, really).
This weekend it became a public repository again, building on JDK 17 with no dependencies at all. As with the applets last month, I did not write the new code; Claude Code did, in one session, working from the 2009 tree and the paper. What I wanted to know was what it costs today.
What it was
In 2009 Dimitris Mitropoulos, Diomidis Spinellis and I published Fortifying Applications Against XPath Injection Attacks at MCIS (the Mediterranean Conference on Information Systems). Dimitris had already built SDriver, a JDBC driver that refuses SQL it does not recognise, and the paper carried the same idea over to XPath.
The mechanism is simple. Every query gets an identifier, made of two parts; the query with its literals and numbers taken out, and the chain of methods that called it. You run the application in a training mode, every identifier goes into a registry, and after that anything the registry does not know is refused. The values do not matter. The shape of the query and the place it comes from do.
The nice part (and the reason I still like it) is that it plugs in through the standard JAXP factory lookup. You ask for org.sdriver.xpath.SecureXPathFactory by name, and the rest of the application stays as it is.
The same overhead
The paper measured one million calls to XPath.compile() on a Core 2 Duo; 21,311 ms for plain JAXP against 48,761 ms through SDriver/XPath, 128% slower. The repository now has a benchmark that repeats that measurement, so I ran it on my M4 Max with OpenJDK 26, three times for each build.
JAXP takes about 1,580 ms now, thirteen times faster. The 2009 code costs about 178% on top of that. The rebuilt code costs between 121% and 129%. So the machine got an order of magnitude faster and the ratio ended up almost exactly where the paper left it. Each check is about two microseconds, and most of it goes on walking the stack, which is the very thing that makes the identifier location-specific. You do not get one without the other.
Epilogue
I should say clearly what I would tell anyone today. If you bind user input as XPath variables, through an XPathVariableResolver, it never becomes part of the query text, and that is the actual fix. SDriver/XPath is a safety net for code that still builds queries by gluing strings together. There is plenty of that around, and there will be for a long time.
The weak spot is the same one we wrote in the paper's conclusions. The registry is only as good as the training run. Rename a method and every query below it needs training again. So the real question was never the two microseconds; it is who keeps the registry current after every release. In 2009 the answer was "the development and testing teams", which is a polite way of saying nobody in particular.
The code is at bkarak/sdriver-xpath, BSD licensed, and it has a page here now. What changed in each release (including 2.0.1, which fixed three bugs found in 2.0.0 the same day) is in its changelog. It is out of the Deprecated section, at least for a while :)