PEP 8 and the standard library: half the violations of 2008, and the one rule it gave up on
September 28, 2026 · Programming
Back in November I wrote that the Python standard library ignores PEP 8 a little more with every release, and I had a table to prove it; 4,479 violations in Python 3.0.1, 16,560 in 3.14.0, and a steady climb in between. Last week I went back to that script, because I wanted to turn it into something more serious, and the table fell apart in an afternoon.
The numbers were right. I can reproduce them today, to the last digit. What they measured was not what I thought they measured, and the conclusion I drew from them ("the quality is falling") was wrong. Time to do it properly, and to ask a better question at the end.
What the script actually measured
The November script was a hundred lines; download every 3.x tarball from python.org, run ruff check --select E,W on its Lib/ directory, keep the line that says Found N errors. Four things were hiding behind that one flag, and I knew none of them.
First, the rule set. ruff 0.14.4 implements 67 of pycodestyle's rules, but 42 of them sit behind a preview flag; every indentation rule (E1), every whitespace rule (E2) and every blank-line rule (E3), which is most of what PEP 8 is about. Without --preview, "E,W" means 25 rules, and it means it silently. Second, the columns. ruff's default line length is 88 (the number black made popular) and PEP 8 says 79. Third, the denominator; Lib/ grew from 357 thousand lines in 3.0 to 960 thousand in 3.14, and I was comparing raw counts. Nobody would compare the number of bugs in two programs without asking how big they are, and I did exactly that.
The fourth one is the good story. ruff looks for a configuration file next to the code it lints, and CPython added a project-wide .ruff.toml on May 1, 2025 (backported to the 3.13 branch in August). It ships in the tarballs from 3.13.8 on, and it says line-length = 79. So from 3.13.8 my script measured at 79 columns, and for every release before it at 88. E501, line too long, went from 3,718 in 3.13.0 to 12,690 in 3.13.9, with every other rule moving by less than a hundred. That is the entire "jump" at the end of the November table; not the library changing, but the ruler changing under it, because the library had started shipping its own ruler.
Ouch. I had learned this lesson before, in my PhD days, mining Java projects; the tool's defaults are part of the experiment, and if you let them move, your result moves with them. Apparently some lessons have to be learned twice.
The Process, take two
The new setup is not clever, it is just careful. The instrument is pycodestyle 2.15.0 at 79 columns with its defaults, because pycodestyle is the reference implementation of the checkable part of PEP 8 (if I remember correctly it started life as pep8.py, by Johann C. Rocholl, and was renamed in 2016 because Guido did not want a tool called pep8 to look official). ruff runs next to it as a cross-check, with every configuration file ignored, all 67 rules on, 79 columns, and the target version set to the release being measured;
$ python -m pycodestyle --max-line-length=79 <every .py under Lib/> $ ruff check --isolated --preview --select E,W --line-length 79 --target-version py314 <the same files>
Both tools get the same explicit list of files, so nothing is excluded by one and counted by the other. Syntax errors and unreadable files are counted on the side and never added to the style count. Everything is reported per thousand lines, and every file goes in one of three bins; the library proper, the tests (Lib/test and the smaller test directories inside packages), and generated files, which is 78 files in 3.14 (the charmap codecs, pydoc_data/topics.py, keyword.py, token.py and friends).
The series is one point per minor release, 3.0 to 3.14. In November I averaged every patch release into its minor, which sounds thorough and is not; a patch release only backports fixes, so the 26 patch releases of 3.9 are 26 copies of the same point, and the 3.13 average mixed eight releases measured at 88 columns with two measured at 79. This time every release python.org offers went through the harness, 264 of them, and a patch release stays within a few points of its minor's first (2.7, which lived for ten years, drifted by twenty); the whole thing takes about half an hour on my M4 Max, most of it downloading. The tarballs are pinned by SHA-256, the tools by version, and the protocol, the harness and every number below are in the repository.
The Results
| Release | Lines | Library proper | Tests | Generated |
|---|---|---|---|---|
| 3.0 (2008) | 356,513 | 87.6 | 99.1 | 805.7 |
| 3.3.0 | 517,181 | 75.4 | 81.1 | 810.2 |
| 3.6.0 | 664,121 | 58.6 | 67.0 | 753.7 |
| 3.9.0 | 781,304 | 51.5 | 60.3 | 738.6 |
| 3.12.0 | 856,398 | 50.5 | 59.4 | 704.0 |
| 3.14.0 (2025) | 959,894 | 48.1 | 56.0 | 645.0 |
The library proper went from 87.6 to 48.1 diagnostics per thousand lines, and it fell at every release but one (3.11 to 3.12, up by 0.2). The tests went from 99.1 to 56.0. ruff agrees with pycodestyle within 2.5% on every release, which, after last time, is the only reason I trust the table at all.
I also went back before 3.0, since pycodestyle can tokenise Python 2 even if ruff cannot. The library of 2.0.1 ran at 261 per thousand lines, half of it the tabs that PEP 8 had just asked everyone to stop using, and 2.7 at 129. The library the guide was written for was five times as far from it as today's.
Then the families. Whitespace (E2) fell from 43.1 to 16.0 per thousand lines in the library proper, blank lines (E3) from 24.0 to 13.8, statements (E7, the "two statements on one line" kind) from 6.0 to 2.4. W605, the invalid escape sequence, went from 159 to zero, and it did so at 3.6, which is when Python itself started warning about it; one warning in the interpreter did in a single release what fifteen years of the style guide had not. And one family went the other way. Line length (E5) doubled, from 5.5 to 10.9 per thousand lines, and E501 is the only rule in the top twenty with more instances in 3.14 than in 3.0 (2,951 against 886).
Where are the violations, then? In 3.14 the library you actually import holds 17% of them, on 28% of the lines. The test suite holds 47%, on 67%. And 78 generated files, 4.5% of the lines, hold 36%; the charmap codecs are tables with a two-space inline comment on every single line, 645 diagnostics per thousand, and in November every one of them counted as "the standard library".
Ok, 48 per thousand lines is still a lot.
Every twentieth line of the library has something pycodestyle would underline; 2,457 top-level definitions with one blank line above them instead of two, 942 inline comments with a single space before the hash, 718 spaces before a colon. The zealots from the first post would have a field day :). But it is half of what it was, and the direction is the opposite of what I wrote.
And it is almost entirely layout. The rules about what the code does (comparisons to None, bare excepts, lambdas assigned to names, a variable called l) come to 308 instances in the whole library, 1.1 per thousand lines, down from 5 in 2.0. Ninety-eight percent of what pycodestyle underlines is a question of where the whitespace goes.
Who wrote the violations
PEP 8 says two things about itself that matter here. It was written for the standard library (that is its first paragraph), and it tells you not to reformat old code just to follow it. So whether a module from 1992 follows a guide from 2001 says nothing about whether anybody cares. Whether the code written after the guide follows it does.
For that you need history, so I cloned CPython (900 MB) and ran git blame on every line of every file under Lib/ at the 3.14.0 tag; 1,830 files, 959,894 lines, two minutes on twelve threads. Every line gets the date of the commit that last edited it, and every diagnostic goes to the cohort of the line it sits on. I cut the cohorts at four events; the day PEP 8 was created (July 5, 2001), the release of Python 3.0 (December 3, 2008), the move to GitHub pull requests (February 2017) and the first ruff hook in CPython's pre-commit configuration (September 2023).
| Last edited | Share of lines | All rules | Without line length | E501 alone |
|---|---|---|---|---|
| before PEP 8 | 9% | 49.3 | 47.3 | 1.9 |
| PEP 8 to 3.0 | 23% | 70.2 | 66.0 | 4.0 |
| 3.0 to GitHub | 31% | 36.9 | 30.3 | 6.5 |
| GitHub to ruff | 22% | 40.8 | 24.2 | 16.5 |
| ruff in pre-commit | 15% | 47.4 | 19.5 | 27.9 |
Apart from line length, code from every era is cleaner than the last; 66 per thousand lines, then 30, then 24, then 19.5. Whitespace goes from 36 to 6 across the same four eras, blank lines from 21 to 7, and the small rules (a space after a comma, spaces around an operator, the shape of a comment) fall eight- to twenty-fold. Line length goes the other way; 2 per thousand lines in the code nobody has touched since before PEP 8, then 4, 6.5, 16.5 and 28. The newest code has the highest total since the Python 2 days, and it is entirely long lines.
Two more things from that table. The worst cohort is Python 2's; the code last edited between PEP 8 and 3.0 carries 70 diagnostics per thousand lines, worse than the lines nobody has touched since the nineties (49). That era is the big packages of the early 2000s together with every line the 2-to-3 conversion rewrote, which is the caveat of the method; blame dates the last edit of a line, not its birth, so a line from 1995 that the conversion touched in 2007 counts as 2007. For a style question I think that is the right date (the last editor was the last one who could have fixed it), but it flatters the nineties and it loads the 2000s. The other thing is that the library is old. 63% of its lines were last edited before CPython moved to GitHub, 9% before PEP 8 existed. 2024 alone touched 7% of it, more than any year before.
Six more measurements
A number on its own is not a result, and 48 per thousand lines is a number on its own. So I ran the same setup, same tools, same flags, on five other code bases; Django, pip, NumPy, requests and black itself. At 79 columns they come out at 27, 26, 18, 41 and 52 per thousand lines of their own source, against CPython's 48. Take line length out and the four projects that let black do their layout sit between 0.2 and 4, NumPy at 9, CPython at 37. Measured by everything except the column limit, the standard library is an order of magnitude further from PEP 8 than a project with a formatter in its pre-commit. Which is, I suppose, the whole point of having one.
The column limit itself deserved a closer look, because "line too long" says nothing about how long. In the newest cohort 2.8% of the lines pass 79, but the median of those lines is 85 characters and nine in ten are under 99. There is no unwritten 88 or 99 in the standard library; there is a 79 that people miss by a handful of characters, because nothing tells them. And PEP 8 has a second limit that nobody measures, 72 columns for docstrings and comments, which pycodestyle checks only when asked. Asked, it finds 10,676 such lines in the library proper, three and a half times the number of code lines over 79. The most broken rule in PEP 8 is the one no tool enables by default.
Naming was my original observation in November (the camelCase next to the snake_case), and pycodestyle does not check names at all, so ruff's port of pep8-naming did it. The library runs at 8 naming violations per thousand lines, down from 12.8 in 3.0. Of the 2,185 in 3.14, 883 are function names that are not lowercase, and they live in xml (317 of them; the W3C DOM's own names), unittest (150), idlelib, logging and multiprocessing; the mixedCase interfaces that PEP 8 explicitly allows "where that's already the prevailing style". What I had noticed was the guide's own exception, applied.
Then the clause about not reformatting old code, which I had been treating as an excuse. It is a policy, and the library follows it. From the point where 3.8 branched off main (June 2019) to 3.14, 74% of the library's lines are unchanged, and so are 70% of the lines carrying a violation; where the line survived, the violation survived 96.5% of the time. Nobody cleans up in passing. The lines that do get touched are the long ones (37% of them changed or removed in six years) rather than the ones with an extra blank line above them (21%).
The violations are not spread evenly either. mimetypes runs at 287 per thousand lines and the alias table of encodings at 213, both because their colons are aligned in columns; locale runs at 171 because every entry of its alias table carries a comment one space too close. All three are lookup tables, the same shape as the generated codecs. asyncio, the largest clean package, runs at 9 on fifteen thousand lines.
And the eras hold up, for what it is worth. The ordering survives a bootstrap over files and moving every boundary a year either way; the two oldest eras carry wide intervals, the three newest do not. Dating each line by its content instead of its last edit (whitespace changes ignored, moved code kept with its origin) shifts a few percent of lines into older cohorts and changes no density by more than four.
What survives from November
The observation survives. The standard library does not follow its own style guide; not in 2008 and not now, and nobody makes it. CPython's pre-commit runs ruff on the tests, the docs and the tools, never on Lib/ itself, and the project-wide configuration that broke my script says 79 columns and then selects no rule that would check them. The list of sixteen rules survives too, more or less, although it was a list of the rules ruff happened to have stable, not a list of what the library violates.
What does not survive is the trend. Measured with one ruler, per line, and without the test suite and the codec tables, the library is twice as conformant as it was in 3.0, and every generation of its contributors has been tidier than the last. The November table was a picture of ruff's configuration, and I signed it as a picture of Python.
Conclusions
Why does line length go the other way, when everything else converges? I think because it is the one rule whose cost is paid by the writer, every time, while the benefit goes to a reader with an 80-column terminal, who is increasingly nobody. Nobody enforces it in CPython, so nobody pays for it, and the code drifts to whatever the editor window shows. PEP 8 itself allows 99 columns for teams that agree on it. The library did not even do that; its long lines stop at 85, not 99, which is the signature of a limit nobody enforces, not of a limit somebody raised.
The guide has a section heading for this situation, and it has had it for as long as I remember; A Foolish Consistency is the Hobgoblin of Little Minds. In November I read the numbers as the library ignoring its guide. It seems the library had read that section more carefully than I had.
The November post stays where it is; being wrong in public is part of the job, and the correction should be public too. The protocol, the harness and the numbers are in the repository, and 3.15 is due any day now; its second release candidate already went through the same harness and lands at 48.3 per thousand lines, so the line seems to know where it is going.