FIRE/J
A regular expression engine for Java that compiles each expression into its own JVM class, so matching runs as generated bytecode rather than an interpreted automaton.
FIRE/J — Fast Implementation of Regular Expressions for Java — is a regular expression engine for the Java (external link, opens in a new tab) platform, written entirely in Java. Where the usual engine walks an automaton it has built in memory, FIRE/J turns every regular expression into a class of its own: the automaton is compiled to JVM bytecode at run time, loaded, and then the match is a straight run through generated code. It follows the POSIX extended syntax, has pluggable parsers and code generator back-ends, caches what it compiles, and is distributed under the Apache License, Version 2.0 (external link, opens in a new tab).
The idea comes from treating a regular expression as a small domain-specific language and compiling it into the host language instead of interpreting it — the source-to-source transformation pattern. The result is a DFA engine that, in the paper's benchmarks, beat every Java engine it was compared with and the C libraries too, while staying portable across machines.
What is what in the diagram
The figure below is the engine's architecture. A regular expression goes in at the top, a compiled class comes out at the bottom, and five components sit in between.
- Regular Expression Programming Interface — the public API, kept compatible with the regular expression API of the Java 5 SDK so that existing code can switch engines. In 0.71 that is
RegexFactory.createRegex(pattern)giving back aRegexwhoserun(text, offset)returns how far the match runs from the offset, or −1. - Preprocessor — rewrites macros into plain syntax before anything is parsed.
\p{Alpha}becomes[a-zA-Z], or, with the locale set to Greek,[α-ωΑ-Ω]; this is also where internationalisation happens. It can work at the syntax level too, turningab+ainto the equivalentabb*a. - Compiler — builds a deterministic finite automaton (DFA) from the simplified expression. FIRE/J does not construct the automaton itself; it uses the parser and DFA builder of Anders Møller's dk.brics.automaton (external link, opens in a new tab) library, which is why the two engines accept exactly the same expressions.
- Code Generation Back-End and its Writer — translate the DFA into code. Two back-ends ship: an experimental one that writes Java source and calls the Java compiler, and the production one that writes JVM bytecode directly. The writer is the stage that assembles the class file and hands it to the class loader.
- Regular Expression Caching Engine — generating and loading a class is the expensive part, so compiled expressions are kept and handed back on the next request for the same pattern, the way Perl's
/oswitch and .NET'sCompileToAssemblywork. It caches the output of both the compiler and the back-end, and it is a singleton — one per virtual machine — so every application in a shared JVM, a servlet container for instance, benefits from what any of them compiled. - Buffer Subsystem — the one standard way of reading the input text, integrated tightly with the generated class for speed. It is optional: a module developer can override it with a buffering strategy of their own.
Why the generated code is faster
A state machine in a language without goto is a loop around a big switch: read a character, jump to the case for the current state, compute the next state, go round again. That extra hop — from a state number to the code for the state — is paid on every character. Java has no goto, but the JVM does, so the bytecode back-end lays every state out as its own block that reads a character, looks up the next state and jumps straight to it. No dispatch, no method calls: helpers like check_state and isTerminalState are inlined into the class.
The Java-source back-end exists to measure exactly that. It emits the switch version, so the gap between the two back-ends is the cost of the indirection alone. On the two hand-crafted expressions below it is a factor of two to four.
Performance
The paper benchmarks FIRE/J against the Java 5 SDK engine, dk.brics.automaton, the GNU and Henry Spencer C libraries, Mono's .NET engine and Perl 5.8, on a 2.8 GHz Pentium 4 running Linux 2.6. Each engine ran in its own process after a warm-up, with FIRE/J's cache switched off. Two hand-crafted expressions — an IPv4 address matcher against a 15-character address, and an Apache access-log line matcher against a 193-character entry — give the clearest picture:
| Engine | IPv4 (µs) | Apache log (µs) | Compile (µs) |
|---|---|---|---|
| FIRE/J, bytecode | 0.1 | 1.2 | 115,690 |
| FIRE/J, Java source | 0.2 | 4.9 | 260,904 |
| dk.brics.automaton | 0.3 | 4.6 | 43,992 |
| Java 5 SDK | 3.5 | 33.8 | 62 |
| Perl 5.8 | 3.8 | 6.0 | 0.9 |
| GNU regex (C) | 9.2 | 67.4 | 160 |
| Mono (C#) | 14.5 | 39.0 | 5.3 |
| Henry Spencer (C) | 118.4 | 7,094.8 | 62 |
The larger experiment ran 849 real-world expressions collected from regexlib.com, narrowed to the 629 that every engine accepted, and plotted matching time against Ehrenfeucht and Zeiger's size and length measures of each expression. FIRE/J's bytecode came out fastest throughout, with dk.brics.automaton second — the only other DFA — and the native C libraries, all NFA-based and written decades earlier, well behind both.
The price is the compile step: turning the DFA into a class costs about 72 ms for bytecode and 217 ms for Java source, which makes FIRE/J the slowest engine to compile by a wide margin. The break-even against Perl, the fastest compiler, is around 24,000 matches of the Apache-log expression. Below that the cache is what makes the engine worth using; above it, or in a long-running server, the compile cost disappears.
The engine in 2026
The engine was modernised in September 2026: Java 21 and Maven, ASM in place of Jasmin, a JUnit suite of a thousand cases, and two corpora of real-world expressions — the 2007 regexlib dump the paper used and a slice of regex101.com. Three things changed in how it works; the source has been public since 9 September 2026, at bkarak/firej-oss (external link, opens in a new tab).
- The dialect is Perl's, as far as a DFA can take it. The preprocessor translates what people actually write —
\d \w \s, character escapes,(?:…)and named groups, lazy quantifiers,.excluding newline, a leading^and trailing$— into the automaton library's syntax, and rejects at compile time what no DFA can express: word boundaries, backreferences, lookaround, inline flags. Before, several of those were silently misread; the letter t stood in for a tab. - The generated code is jump-threaded again, one block per state jumping straight to the next, as the paper describes — the 2007 template had done it with
jsr, which modern verifiers reject — and each state with more than two character ranges now resolves its character through a class map: one table lookup and one jump table, instead of a compare-and-branch pair per range. That last step is what a table-driven DFA does, and without it dk.brics.automaton was ahead on real-world expressions. - Capturing groups are recovered after the DFA has fixed the match span, by walking the expression over that slice, so the generated code stays group-agnostic. Compile time is still the price: a few hundred microseconds per expression against the JDK's few.
The four back-ends — class-map threading, plain threading, the paper's switch loop and an interpreter — are all kept, so the effect of each step can be measured. Rerunning the paper's IPv4 measurement on an Apple M4 Max with OpenJDK 26, one million matches after warm-up:
| Engine | 2008 (µs per match) | 2026 (ns per match) |
|---|---|---|
| FIRE/J, threaded with class map | — | 6 |
| FIRE/J, threaded, compare chains (2008: bytecode) | 0.1 | 8 |
| FIRE/J, switch loop (2008: Java source) | 0.2 | 17 |
| FIRE/J, interpreter | — | 58 |
| dk.brics.automaton | 0.3 | 19 |
| Java SDK | 3.5 | 184 |
The ratios have held up better than the absolute numbers. Threading beats the switch loop by two to one, as in the paper; against the JDK the lead is 30 to 1 where the paper had 35, the JDK's own engine having been rewritten several times since Java 5. On an expression with wide character classes, an email matcher, the class map is what keeps FIRE/J ahead of the automaton library: 6 ns per match against 10, where the compare chains alone manage 7.
The paper's benchmark, rerun with ten engines
A standalone benchmark now follows the paper's section 6 as written — one process per engine, warm-up, compile time and amortised match time, break-even, the size and length of every expression — and runs FIRE/J beside the JDK, dk.brics.automaton, RE2/J, Joni (Oniguruma on the JVM), Perl and Python. The expressions are the paper's two, the 2007 regexlib dump the paper drew on, and 3,311 from regex101.com; only rows every engine compiles are compared, 823 and 3,200 of them. Medians of three repetitions on the same Apple M4 Max, nanoseconds per match:
| Engine | IPv4 | Apache log | regexlib 2007 | regex101 |
|---|---|---|---|---|
| FIRE/J, class map | 8.5 | 118 | 6.1 | 5.6 |
| FIRE/J, compare chains | 10.8 | 210 | 6.4 | 6.0 |
| FIRE/J, switch loop | 19.8 | 466 | 10.3 | 7.4 |
| FIRE/J, interpreter | 59.3 | 1,410 | 22.1 | 12.9 |
| dk.brics.automaton | 13.1 | 440 | 8.1 | 5.9 |
| java.util.regex | 196 | 1,080 | 64.1 | 40.2 |
| RE2/J | 506 | 4,210 | 207 | 264 |
| Joni | 393 | 1,090 | 168 | 121 |
| Perl 5.34 | 524 | 688 | 231 | 172 |
| Python 3.14 | 360 | 588 | 104 | 78.6 |
FIRE/J's class-map code is the fastest participant in every column: 1.5 times faster than the automaton library on IPv4, 3.7 times on the Apache log line, 1.3 times over the 2007 dump, and level on regex101, whose sample inputs mostly fail within a few characters. The class map is worth a quarter over the compare chains on the corpora and nearly twice on the long expression. Seeing that took three rounds of work on the harness rather than the engine: the first standings had the automaton library two to five times ahead on the corpora, and every one of those numbers came from the JIT — a class per pattern means a compilation per pattern, thousands of them queued behind each other, so a fixed warm-up was timing FIRE/J's profiled first-tier code. The harness now warms each pattern up until its pace is steady. Compile time remains the price the paper described: hundreds of microseconds against the JDK's half a microsecond, a break-even of 3,000 to 6,500 matches. The backtracking engines show their tail: Joni, Perl and Python each exceed a 60-second timeout on a few regex101 rows, while the JDK never does but still spends 16 ms on its slowest.
Where it stops
- It is a DFA engine. Matching is text-directed and always returns the longest leftmost match, so backreferences, lookaround and the rest of the Perl extensions are not available. The 0.71 syntax is POSIX extended; the 2026 engine reads Perl's and refuses those constructs at compile time.
- It accepts what dk.brics.automaton accepts. Of the 849 regexlib expressions, 665 compiled — 78%, against 95% for the SDK engine. That was a deliberate trade: reusing the automaton library isolates the gain from code generation alone.
- Big automata make big classes. With many states the generated class grows large enough to miss the instruction cache, and the Java-source back-end already falls behind dk.brics.automaton on the Apache-log expression for that reason. The bytecode back-end has the same ceiling, further out. The 2026 engine meets a harder one first: HotSpot never compiles a method whose bytecode passes 8,000 bytes, so a threaded walk that would exceed that is split into several methods that hand the state number between them, and large automata keep running compiled code. The switch form is not split: past that size it runs in the JVM's interpreter, and past the JVM's 64 KiB method cap it falls back to FIRE/J's own.
- Compilation is slow, as above. The paper's own advice is to use FIRE/J server-side or in anything that matches intensively — log analysers being the example — and not for one-off matches.
Using it
The 0.71 API is two calls. Example.java in the downloads runs five expressions this way, the benchmark's IPv4 matcher among them. The shortest of the five:
import org.firej.Regex;
import org.firej.RegexFactory;
Regex r = RegexFactory.createRegex("[0-9]{3}-[0-9]{2}-[0-9]{4}");
String v = "333-22-4444";
int end = r.run(v, 0); // length matched from the offset, -1 on failure
boolean full = (end == v.length());The zip includes its dependencies — the automaton library and the bytecode assembler — and the Javadocs are a separate download. The benchmark utility that produced the 2008 numbers above has a page of its own.
The 2026 engine keeps RegexFactory as a facade and adds a java.util.regex-shaped one, Pattern and Matcher, with Regex.compile for the one-call case. The same expression, and a search:
import org.firej.Regex;
import org.firej.Pattern;
boolean ok = Regex.compile("[0-9]{3}-[0-9]{2}-[0-9]{4}").matches("333-22-4444");
var m = Pattern.compile("[0-9]+").matcher("ab12cd34");
while (m.find()) {
System.out.println(m.group()); // 12, then 34
}Firej.builder() picks the parser, the code generator and the cache, which is how the four back-ends in the tables above are chosen. A compiled template may be shared between threads; a matcher may not.
Source
The 2026 engine has been public since 9 September 2026 at bkarak/firej-oss (external link, opens in a new tab), under the Apache License 2.0. It is a Maven project for JDK 21 or later: mvn verify builds it and runs the suite, which the repository's CI runs on Java 21 and 25. It holds the engine alone — the preprocessor, the four back-ends, the capture recovery, the caches, the JUnit suite and the two corpora it is checked against. The benchmark that produced the 2026 tables above is a separate program and is not published. The 0.71 release stays available as the zip below.
Publications
Vassilios Karakoidas and Diomidis Spinellis, FIRE/J — Optimizing Regular Expression Searches with Generative Programming, Software: Practice and Experience, 38(6):557–573, 2008. PDF. Every 2008 figure on this page comes from it; the 2026 figures are from a rerun of its IPv4 measurement.
Vassilios Karakoidas, Compiling regular expressions into Java bytecodes, MSc dissertation, Athens University of Economics and Business, 2004 (in Greek). PDF. The work FIRE/J grew out of.
Related projects
- SARE — Simple API for Regular Expressions (external link, opens in a new tab)
- dk.brics.automaton (external link, opens in a new tab), the DFA library FIRE/J builds on
The FIRE/J logo was designed by George Zouganelis (link no longer available). The last packaged release is 0.71; the current engine is in the repository above. Disclaimer: FIRE/J is not related to the FIRE++ regular expressions engine, nor to Bruce Watson's FIRE toolkit for C++.
Architecture
Downloads
- fire-0.71.zipFIRE/J binary zip, includes dependencies
- fire-docs.tar.gzJavadocs
- Example.javaExample code
- fire-benchmark.zipThe benchmark utility — ~250Kb, includes binaries for all supported engines
These are the original files, kept as they were published. Most are decades old and are here as a record rather than as working software.