JSLoCCount
A single-jar line counter for a whole directory tree: source and comment lines for 356 text formats, file counts for everything else, written as CSV.
JSLoCCount takes a directory and says what is in it: source and comment lines for every text format it knows, and a file count for everything else, from images to compiled classes. One scan, one account of the whole tree, written as CSV so the numbers go straight into whatever you compare them with.
It dates from August 2012 and the PhD years (the CKJM extension on this site is from the same period, and the same obsession with software metrics), and it was a single jar with no dependencies then as it is now. What changed in 2026 is everything behind the jar: the counting was rewritten, the table of file types grew from 83 to 458, and the tool gained a test suite. This page describes version 2.1.0, released on 11 September 2026, and every figure on it was taken from a run of that release.
Getting it
The jar is attached to each release on GitHub. It needs a Java 21 runtime and nothing else: no installer, no configuration file, no dependency to fetch.
curl -LO https://github.com/bkarak/jsloccount/releases/latest/download/jsloccount.jar java -jar jsloccount.jar <directory>
Building from a checkout is ant release, and ant test runs the suite. The tests are plain Java as well, so there is nothing to fetch for those either.
Running it
java -jar jsloccount.jar [options] <directory>
-o, --output <dir> write the reports into <dir> (default: the working directory)
-n, --name <name> base name for the report files (default: the scanned directory)
--stdout write one combined report to standard output instead of files
-x, --exclude <name> skip files and directories called <name>; repeatable
--include-hidden scan hidden files and directories too
-q, --quiet suppress progress messages
--list-languages list every recognized file type and exit
-h, --help show this help and exit
-V, --version show the version and exitTwo reports are written, <name>-filestats.csv and <name>-sizestats.csv, named after the scanned directory unless -n says otherwise. Both spellings, --opt value and --opt=value, are accepted, and -- ends option processing. Progress and warnings go to standard error, which is what lets --stdout be piped:
java -jar jsloccount.jar --stdout --exclude node_modules ~/src/project | column -s, -t
The exit status is 0 on success, 1 when the directory cannot be scanned and 2 for a usage error. A bare invocation is a usage error rather than a silent success, so an empty variable in a script cannot pass for a clean run.
What it reports
Run over its own tree after ant release (from a neighbouring directory, or the CSVs it writes land in the tree and the next scan counts them), --stdout gives one table:
Resource Type,File Count,Total File Count,Source Lines of Code,Comments Lines of Code Java,28,56,2300,704 Java Compiled Class File,23,56,, Markdown,2,56,793,0 ANT Build File,1,56,45,3 JAR,1,56,,
Each row carries two file counts, the files of that type and the files in the whole tree, so it reads as a share. The line counts are empty rather than zero for the binary types, which are counted by file alone. A file of a type it does not recognise is left out of the table and reported on standard error instead, as a count per suffix; on this tree that is the one file with no extension, the jar manifest.
Without --stdout the same scan writes the two CSVs, the file statistics for every type and the size metrics for the text types:
Resource Type,File Count,Total File Count Java,28,56 Java Compiled Class File,23,56 Markdown,2,56 ANT Build File,1,56 JAR,1,56
Resource Type,Source Lines of Code,Comments Lines of Code Java,2300,704 Markdown,793,0 ANT Build File,45,3
How it counts
The lines are physical. A line is counted once as source, once as comment, or once as each, so int c; /* inline */ int d; is one source line and one comment line, and a blank line is neither. Most of the work is in not being fooled:
- String literals are skipped, not searched. A marker inside a literal opens no comment, so a line holding
"http://example.com"is code, and a quote inside a comment opens no literal. Java text blocks, Python docstrings and JavaScript templates carry across lines the way block comments do. - A block comment stays open across blank lines until its closing marker, and code after the close on the same line counts as source.
- Nesting is per marker, not per language. Rust, Swift, Scala, Kotlin, Dart, Haskell, OCaml, F# and Julia nest; D nests
/+ +/but not/* */; C, C++, C#, Java, JavaScript and Go deliberately do not, since a nesting/* */there would swallow everything after/* a /* b */. - Position counts where the language says so. Fixed-form Fortran's
Copens a comment in column one and nowhere else. Vim Script's"is a comment at the start of a line and, elsewhere, a string if it closes on the line and a comment if it does not. - The awkward spellings are spelled out. Lua's long brackets,
--[[through--[===[, in comments and in strings; the block comment written with one delimiter used twice, as Smalltalk, CoffeeScript and Algol 68 do; the*line comment of COBOL, SAS and ABAP, which opens a comment only at the start of a line because anywhere else it is a multiplication.
Every rule has a hand-written case in the test suite, in the language the rule exists for, with the counts worked out by reading the sample rather than by running the counter.
What it recognises
458 file types in ten groups. The first four groups are the text types, 356 of them, and are counted for source and comment lines; the rest are binary and counted by file, so they appear in the file statistics only. (Two exceptions: SVG is XML and PostScript carries a % marker, so both are line-counted despite their group.) --list-languages prints the whole list from the tool itself.
| Group | Types |
|---|---|
| Programming languages | 226 |
| Markup, data and configuration | 75 |
| Build and project files | 33 |
| Documentation and text | 20 |
| Images, fonts and design assets | 23 |
| Audio and video | 15 |
| Archives and packages | 26 |
| Office documents | 11 |
| Databases, datasets and models | 14 |
| Compiled artifacts | 15 |
The text half is kept in step with scc (external link, opens in a new tab): all 366 languages in its table resolve to a type here, and the comment markers and string quotes were taken from its languages.json rather than from memory. The binary half is there on purpose. A tree that is half translation catalogs or model checkpoints should say so, and a counter that discards whatever it cannot line-count reports it as half "other".
Where one suffix belongs to two languages a winner was picked, and the loser is counted under the winner's name: .m is Objective-C rather than MATLAB, .ts TypeScript rather than a Qt translation, .v Verilog rather than Coq, .pl Perl rather than Prolog (which keeps .pro). The numeric man-page sections are left out deliberately, since .3 would also claim every libfoo.so.3 in a tree. The suffixes it did not recognise are named after every scan, so the next gap is visible rather than buried in "other".
Tests
ant test runs 9,257 checks over eight suites, with no framework behind them, and any failure fails the build. This is the closing of a run against 2.1.0:
resource table 4,718 checks, 0 failures invariants every file type must keep detection 2,085 checks, 0 failures every extension resolves to its own type languages 1,780 checks, 0 failures one counting case per text type counting rules 88 checks, 0 failures the documented rules, sample by sample project statistics 15 checks, 0 failures the directory walk and what it totals options 51 checks, 0 failures the command line grammar and its refusals csv 48 checks, 0 failures RFC 4180 escaping, by round trip stream output 472 checks, 0 failures the combined report behind --stdout OK — 9,257 checks over 8 suites, 0 failures
The languages suite is generated: for each of the 356 text types a sample is assembled from that type's own markers, so every language has a counting case and a new type gets one for free. A generated case cannot tell whether the declared markers are the right ones for the real language (scc's table is the oracle for that), which is why the counting rules keep hand-written cases as well, and why the detection suite resolves all 785 declared suffixes back to the type that declares them, the check that catches one type shadowing another.
Limits
Some of what it gets wrong, known and left alone for now:
- Whole-filename entries match by suffix, so
mybuild.xmlis an Ant build file. - Lua long brackets past level three,
[====[and up, are not recognised. - Vim Script's
"is resolved by heuristic. Against Neovim's runtime it lands within about 2% of an independent estimate, and it misreads a comment that follows an unterminated single-quoted string. - Forth's
\is a word in the language rather than a prefix, so a backslash anywhere opens a comment. - Where a suffix collision was settled, the loser is mislabelled. Only Coq miscounts, its
(* *)comments reading as code under Verilog's markers.
History
The first commit is from 19 August 2012. By 2024 the tool knew 83 file types, wrote its two CSVs and had had the odd language added (Swift and Go in 2017), and that is where it stood when the 2024 post Software Complexity Metrics meet LLMs mentioned it. Two releases in 2026 are most of what this page describes.
- 2.0.0, August 2026. Modernised for Java 21, and three defects that predated it fixed: the two CSVs had their columns crossed (file counts under "Source Lines of Code" and the other way round), a block comment counted as code beyond its opening line, and a marker inside a string literal opened a comment. Nested comments, RFC 4180 escaping and the command line above arrived with it, and coverage went from 83 types to 192.
- 2.1.0, September 2026. Coverage from 192 types to 458, and the first test suite. Neither the interface nor the report format changed, and a type that was counted before is counted the same way now, with two exceptions: PostScript and
.vhdfiles, both binary before, are counted as text.
Source
- bkarak/jsloccount (external link, opens in a new tab)
- Releases, with the jar attached to each (external link, opens in a new tab)
Requires Java 21 or newer, BSD-licensed. Of the pages here this is one of the few describing software still being changed, so where the repository and this page differ, the repository is current.