Aggressive Regression
· 17 min read
testing elixir programming opinion ai llm hacking
There is a particular species of engineer who, upon being asked whether the codebase has regression tests, answers “we have ninety-four percent coverage.” This is roughly equivalent to answering the question “is your house insured against fire?” with “I own eleven smoke detectors, and I polish them weekly.” The two statements inhabit the same general neighbourhood of concern, but they are not even distantly related, and only one of them will matter on the night the kitchen goes up.
Let me state the thesis up front, in the hope that it survives the journey to the comment section intact:
Each regression test is a test, but not each test is a regression test.
This is not a clever aphorism. It is a statement about sets, and it is the sort of statement that most teams nod along to and then immediately violate for the next four years.
The Taxonomy Nobody Bothered to Learn
A test is any piece of code that fails when something is wrong. A regression test is a piece of code that fails when something that used to work stops working. The second category is a strict subset of the first, and it is a far smaller subset than your coverage badge implies.
The distinction lives entirely in the word “used.” A regression test is a memorial. It commemorates a specific moment in history when the software was observably correct, and it stands guard over that moment for as long as the project draws breath. It is, in the most literal sense, institutional memory encoded as an executable artifact—which makes it the only kind of memory your team has that does not resign, retire, or forget everything over a long weekend.
Most of what we write is not that. Most of what we write is a developer reassuring himself, at two in the afternoon, that the function he just typed does what he just thought it should do while typing it. That is a perfectly respectable activity. It is also a tautology wearing a lab coat.
Unit Tests Are Mostly Mirrors
Here is the uncomfortable thing about the average unit test: it asserts that the implementation does what the implementation does. The author wrote the function, then wrote a test derived from the function he had just written, and then expressed surprise when both agreed. The two of them were never going to disagree. They were conceived in the same five minutes, by the same brain, operating under the same misconception. Even a mediocre developer like myself is able to wrote a simple function without a glitch.
A unit test typically reaches inside. It knows the module’s private helpers, the shape of its intermediate accumulator, the exact order in which it walks a list. It is intimate with details that no caller will ever observe, and it is therefore brittle in exactly the places where you most want freedom to move. Refactor the accumulator from a list to a map, change nothing whatsoever about the observable behaviour, and watch forty tests go red while the software remains flawless. That is not a safety net. That is a tripwire you installed in your own hallway.
And the whole edifice rests on the author’s imagination. The unit test checks the three inputs he thought of. The bug, with the unerring instinct of a cat finding the one guest in the room who is allergic, lives in the fourth.
This is why a single property-based test is worth a bushel of hand-typed examples. Instead of enumerating what you happened to think of, you declare the law the function must obey, and then let the machine hunt for the counterexample with a diligence no human possesses at four in the afternoon on a Friday.
Consider the usual hand-rolled approach:
test "encodes and decodes" do
assert "abc" == decode(encode("abc"))
assert "" == decode(encode(""))
assert "hello world" == decode(encode("hello world"))
end
Three assertions, all of them chosen by a human being who already knew how the encoder worked. Compare:
property "decode/1 is the left inverse of encode/1" do
check all input <- StreamData.string(:printable) do
assert input == input |> encode() |> decode()
end
end
One statement, infinitely many assertions, and—this is the part people underestimate—an automatic shrinker that will hand you the minimal failing input rather than the 4 KiB blob that happened to trip it. The first version tests your memory. The second tests your software.
The law, incidentally, is the valuable artifact here.
decode(encode(x)) == x
is a sentence about the system that remains true across three rewrites, two language versions, and the departure of everyone who originally wrote it. Your
assert "abc" == ...
is true about one string.
Round-trips, idempotency, commutativity, invariants, monotonicity, order-independence. If you cannot name a single law your module obeys, that is not an argument against property-based testing. That is a code review finding about the module.
Elixir Will Not Let You Test Private Functions, and This Is a Feature
Every few months, someone arrives in the forums with a question phrased as a bug report: “How do I test a
defp
?” The replies are patient and the asker is rarely satisfied, because the honest answer sounds like obstruction:
you don’t, and you shouldn’t want to.
A private function is, by construction, not part of your contract. It is scaffolding. It exists because you got tired of a forty-line function and cut it in half, or because you needed the same three lines twice. Nobody outside the module can call it, nobody outside the module can observe it, and nobody outside the module can be harmed by its replacement. It has no behaviour in the sense that matters—it has only implementation, and implementation is precisely the thing we have just agreed not to test.
The language is not being coy. It is telling you something: if you feel a burning need to test a private function in isolation, you have discovered a design pressure, not a tooling gap. One of three things is true.
It should be public. The thing you are trying to pin down is actually part of the module’s contract, and you have been hiding your own API from yourself out of some misplaced sense of tidiness. Make it public, document it, test it as the public thing it always was.
It should be its own module. The private helper has grown a personality: its own vocabulary, its own invariants, its own reasons to change. That is a module trying to be born. Deliver it. Now it has a public interface, and testing it is no longer a philosophical crisis.
It should not be tested at all. It is three lines of plumbing between two public functions, both of which are tested, and both of which would scream if the plumbing broke. Leave it alone. Every test you write is a line of code you must maintain forever; spending that budget on a private two-liner is the testing equivalent of insuring your shoelaces.
Naturally, someone always discovers that
Module.function
can be reached through the back door with
apply/3
and a little dishonesty about module attributes, or that one can simply
@compile :export_all
and get on with their life. Yes. You can also remove the guard rail from the balcony because it was obstructing your view. The guard rail was not a design flaw.
There is a second, subtler gift hidden here. Because the language makes implementation untestable, it quietly pushes the lazy developer—and we are all the lazy developer on Thursday afternoons—toward testing behaviour, which is the only thing worth testing. Elixir has encoded good taste into a compiler restriction. Few languages have the nerve.
Behaviour Is a Virtue, Implementation Is a Vice
Here is the mechanical reason behavioural tests catch regressions and implementation tests do not.
A regression happens when someone changes the implementation and accidentally changes the behaviour. That is the entire anatomy of the bug. Now ask yourself what happens to each kind of test under that event.
The implementation test was written against the old implementation. The developer changed the implementation, so the test broke—but it broke because the implementation changed, not because anything is wrong. The developer, who is reasonable and in a hurry and has seen this test cry wolf nine times this month, updates the test to match the new implementation. The test goes green. The bug ships. The test participated enthusiastically in its own defeat.
The behavioural test was written against the contract, which nobody intended to change. It breaks, and there is no plausible way to “update” it, because updating it would mean openly declaring that the software now does something different. That is a conversation, not a quick fix. The bug does not ship.
This is why a proper regression test has a very particular pedigree. It is not written during feature development by a cheerful developer. It is written in the aftermath of an incident, by someone slightly grim, and it reads like a deposition:
describe "issue #1472: timezone-aware expiry" do
test "a token minted at 23:59 UTC is not considered expired at 00:01 local" do
# Reported 2026-04-02 by a customer in UTC+2 whose entire
# team was logged out overnight, every night, for nine days.
token = mint_token(at: ~U[2026-04-02 23:59:00Z], ttl: {2, :hour})
assert {:ok, _claims} = verify(token, now: ~U[2026-04-03 00:01:00Z])
end
end
Note what the test knows and what it does not. It knows the symptom, the date, and the victim. It does not know whether expiry is computed with
DateTime.diff/3
, a cached monotonic offset, or a sundial. You may rewrite the entire expiry subsystem tomorrow; this test will follow you there without complaint, and the day it goes red, something real has broken.
Note also the comment. A regression test without the story of its birth is a landmine for the next person, who will eventually look at a strange assertion, fail to imagine why anyone would write it, and delete it. Every regression test should carry its own obituary of the bug it buried. The cheapest documentation in your repository is the kind that fails the build when it becomes false.
The rule, stated without hedging: every bug fix is accompanied by a test that fails before the fix and passes after it. If you cannot write a failing test, you have not understood the bug—you have merely disturbed it, and it will be back.
And Then the Machines Arrived
Everything above has been true since roughly 1975. What changed is the volume and the provenance of the code flowing into the repository.
Generated code is not worse than human code in the way the pessimists claim. It is worse in a far more specific and dangerous way: it is
locally plausible and globally unanchored
. A model producing a refactor has no memory of the incident in April. It has never been paged at three in the morning. It does not know that the retry loop has a weird
+1
in it because a vendor’s API counts from one, nor that the normalisation step strips a particular Unicode character because of a bug in a client shipped four years ago that six hundred customers still run. All of that knowledge lives in exactly two places: in the heads of people who will eventually leave, and in your regression suite.
When a human engineer stumbles across a strange-looking line, he hesitates. He runs
git blame
. He asks in the channel. A model does not hesitate—it cannot, because hesitation requires knowing what you do not know, and a statistical amplifier is constitutionally incapable of that particular form of humility. It sees an oddity, recognises it as deviation from the idiom it was trained on, and tidies it away with the serene confidence of a cleaner throwing out an artist’s installation because it looked like a pile of rubbish.
Your regression suite is the only thing in the building that can say no to that. Not the linter, which thinks the tidied version is objectively nicer. Not the type checker, which agrees the types line up. Not the reviewer, who is looking at a diff of four hundred lines with a meeting in six minutes. Only the test that was written by someone who was there in April.
Which inverts the economics in a way worth stating plainly. The cost of writing code has collapsed. The cost of verifying code has not moved an inch, because verification is where the domain knowledge lives, and the domain knowledge was never in the keystrokes. So the ratio shifts. In 2015, a team producing ten thousand lines a year could afford a casual relationship with its test suite. A team now producing ten thousand lines a month cannot—not because the code is worse, but because no human being is reading all of it any more, and the suite has quietly been promoted from a safety net to the primary reader.
Two corollaries, both of which people hate.
Never let the model write the regression test from the fixed code. The whole point of the test is that it was derived from the symptom , independently of the implementation. Hand the model a patched function and ask for tests, and you will get an exquisitely thorough description of the patch—a mirror again, just a faster and better-spelled one. Write the failing test first, from the bug report, in your own words. Then let the machine fix the code until the test goes green. That direction is enormously productive. The other direction is theatre.
Be very suspicious of a green suite after a large generated refactor.
Either your suite is genuinely behavioural, in which case congratulations and go to lunch; or your suite is thin enough that nothing in it was ever going to notice. The way to tell the two apart is to break something on purpose and check that the suite screams. If you want this done systematically rather than by intuition, that is what mutation testing is for: it is a test suite for your test suite, and the first run is always a humbling experience. Try
muex
btw, it shines.
Further Specimens from the Wild
Since the invitation was open, some additional patterns that live on the correct side of the line—and some impostors that do not.
Characterisation tests for code nobody understands. You inherit forty thousand lines with no tests and no author. You cannot write behavioural tests, because nobody knows the intended behaviour. So you write down the actual behaviour—feed it real inputs, record the real outputs, assert they stay that way. The tests are not asserting correctness; they are asserting stability , which is the only thing you can honestly assert about a system you do not understand. This is regression testing in its purest form, and it is the standard entry point for making a legacy system safe to touch. Mark these tests clearly, because some of them are faithfully memorialising bugs, and one day somebody will fix one and be baffled by the red.
Doctests as behavioural contracts.
Elixir’s
doctest
is a quietly radical idea: the example in your documentation
is
the test. Documentation lies constantly, in every project, in every language—except here, where a lie fails the build. And because documentation examples are written from the caller’s perspective by definition, they are behavioural by construction. You cannot document a private function’s intermediate accumulator in a way a user would find useful, so you do not, so you do not test it. The medium enforces the message.
Contract tests at service boundaries. The regression most likely to ruin your quarter is not in your code at all; it is in the thing on the other end of the wire, which changed a field from a string to an integer on a Tuesday without telling anyone. A contract test pinned against a recorded, versioned interaction catches this on your CI run rather than in production. Behaviour at the seam, which is exactly where behaviour is hardest to reason about and most expensive to get wrong.
Mox
and explicit contracts, as opposed to mocks as a noun.
Mocking a private collaborator so you can assert it was called twice is implementation testing with extra ceremony and a worse error message. Defining a behaviour, implementing it twice—once for real, once for tests—and asserting on the
interaction contract
is behavioural testing with a swappable backend. The first makes refactoring harder; the second makes it possible. Same library, opposite outcomes, and the difference is entirely in whether the thing you mocked is a contract or an accident.
The flaky test is a regression test for your concurrency, and you are ignoring it.
Every team has the one test that fails on CI roughly once a fortnight, and every team has a
@tag :skip
or a retry wrapper over it. That test is not broken. That test is the only member of your suite that has noticed the race condition, and you have gagged it. Nondeterminism in tests is nearly always a true positive about nondeterminism in the system.
Time, locale, and encoding: the three horsemen.
A startling proportion of all regressions ever recorded come from a date crossing a boundary, a locale formatting a number differently, or a string containing a character the author did not personally know existed. These are behavioural by nature and almost never covered by hand-written unit tests, because the author’s imagination is bounded by the author’s timezone, locale, and alphabet. Property-based generators do not share those limitations, which is why
StreamData
finds the emoji in the username field within about nine runs.
Golden files, used correctly.
Snapshot testing has a bad reputation, entirely earned by teams who run “update all snapshots” as a reflex and commit the result unread. A golden file reviewed with the same seriousness as source code is an excellent regression test for anything with a complex output shape—rendered templates, generated schemas, serialised payloads. A golden file updated by muscle memory is a
git add .
with extra steps.
Performance budgets.
A function that returns the correct answer in nine seconds, having previously taken forty milliseconds, has regressed—the behaviour of a system includes how long it takes to behave. A crude assertion on an upper bound, or a benchmark gated in CI, catches the accidental
Enum.map
over a database query that someone slipped into a loop. Nobody writes these until the first outage, which is a pity, because they are cheap.
The test you delete. Last and most heretical: a test that has never failed for a good reason in three years is not protecting you. It is a line item in your maintenance budget, a drag on your CI time, and a small daily tax on everyone’s attention. Regression suites should grow by accretion from real incidents, not by a coverage quota imposed from above. Coverage measures which lines were executed, not which behaviours were verified—a suite that calls every function and asserts nothing scores a perfect hundred, as does one that asserts everything and understands nothing.
The Short Version
Tests are cheap and plentiful; regression tests are expensive and rare, because each one costs an actual outage to acquire. Treat them accordingly. Write them from the bug report rather than from the fix. Aim them at the contract rather than the plumbing. Let the compiler stop you from testing private functions, and thank it quietly for the discipline. Prefer the law to the example, and let the generator find the counterexample you were never going to think of.
And now that an increasing share of your codebase is arriving from a source with no memory of the last five years, understand what your suite has become. It is not a quality gate any more. It is the only institutional memory you have that does not sleep, does not leave, and cannot be talked out of its position by a confident paragraph of plausible reasoning.
Be aggressive about it. The regressions certainly will be.
Previously, on the subject of not deceiving yourself with your own tooling: