Field notes · 15 Sep 2026 · AI code generation
AI writes the code. You review every line.
That plan has a math problem.
The new division of labor sounds clean: the model produces, the human reads, the pipeline ships. But generation is cheap and getting cheaper, while reading is expensive and capped by biology. A pipeline where the cheap side emits and the expensive side inspects every line is a queue with a person standing in it. This is a field note about the exit: what a deterministic floor removes from the reading pile, what stays probabilistic, and why the merge should stay human even when nothing else does.
The math that breaks first
Every organization that adopts code generation discovers the same arithmetic. If the model doubles the code entering a workday, the diff surface doubles and review load doubles with it. Humans do not scale to match. Attention is a fixed budget, and correctness review by eye degrades fast: past the first screens of a diff, a reviewer stops finding bugs and starts finding confirmation that the code looks fine. The bug at line 870 ships with a signature on it.
The industry's answer was a second model: let AI review AI. On paper it closes the loop. In practice, teams that run model review and count honestly report that only about a third of the comments are worth acting on; the rest is noise, style preference, or flat wrong. Each call costs orders of magnitude more than a pattern match. And the noise is not free. A reviewer that cries wolf gets muted, and muting is where the one real finding goes to die.
Give the cheap problems to the cheap layer
Look at what model review is actually asked to catch. A hardcoded API
key. A dynamic eval on user input. A curl
piped into a shell. A diff touching a private credentials file. None
of that needs judgment. All of it needs exact matching. Running a
probabilistic, expensive layer as the first gate while the
deterministic layer idles is inverted engineering: the costly tool
does the cheap work badly, and the cheap tool does nothing at all.
Sober's preflight is the cheap tool, placed where it costs least: a
pre-commit and pre-push hook, pattern-exact, offline, milliseconds,
before the diff exists as a commit. It catches secrets, private-file
leaks, unsafe eval and exec calls,
curl | sh, prompt injection, and the AI-slop
shapes that agent workflows produce. It either matches or it does
not. It cannot hallucinate a missed secret; it has no off day at
line 870; the same diff produces the same verdict every time it
runs. For this class of defects, zero variance is the entire
feature.
| Layer | What it covers | Cost | Failure mode |
|---|---|---|---|
| Deterministic floor (preflight) | Secrets, private-file leaks, unsafe eval/exec calls, curl | sh, prompt injection, slop shapes |
Milliseconds, offline, no model, no tokens | Blind to semantics. Within its class: none. |
| Model review (advisory) | Intent vs. diff, logic, practice | Seconds to minutes, per call | Confident misses; a third of comments actionable |
| Human merge | Stakes, accountability, product judgment | Human attention, capped | Rubber-stamping under load |
The harness is the product
Here is the open secret of the AI coding boom: the tools everyone calls AI are mostly ordinary code wrapped around a model. The agent harness you are typing into is itself hundreds of thousands of lines of classical engineering: sandboxing, diffing, retries, permissions, parsers, prompts. The industry selling "AI writes your software" is selling deterministic scaffolding around a stochastic core. Accept that, and the real design question becomes placement. Where in the pipeline do the deterministic parts go?
Put them after the model, at review time, and you pay model prices
to generate a defect, then pay again to catch and explain it. Put
them at the commit boundary, and the class dies before it exists.
Same rules, same engine, different slot in the pipeline; the second
slot is cheaper by the entire cost of generating and reviewing that
defect. Sober puts the floor at the boundary. One command:
sober hooks install all.
The fast loop belongs to the agent
A quiet shift is underway in teams that measure their pipelines: models implement analyzer findings remarkably reliably. Hand a model a precise finding and it fixes the code; hand it a vague review comment an hour later and it renegotiates. The implication is uncomfortable for the "developer as permanent reviewer" plan. The best reviewer of machine-generated code includes the machine that generated it, in the same session, with the same context.
That loop only works if feedback is fast. A hook that runs in
milliseconds turns findings into the agent's native cycle: emit,
check, fix, continue. A review queue that answers in an hour forces
a context reload you pay for twice. Sober feeds both consumers:
preflight at the boundary, and every run, finding, and verdict lands
signed in .sober/store.sqlite, queryable by a person or
by the agent itself. Evidence as infrastructure, not a vendor
dashboard.
Laws for pipelines that generate code
Stated as doctrine, because we run on them:
"Generation is cheap. Verification is not. A pipeline that leaves verification last pays full price twice."
"A reviewer who cannot keep up does not ask for help; they start approving."
"The merge is the last cheap control point. Models advise; humans own."
What the deterministic floor buys
Zero variance
A pattern matches or it does not. No hallucinated passes, no fatigue drift, no bad day at line 870. For the machine-checkable class, "the gate blinked" is structurally impossible.
The reading pile shrinks by construction
The floor spends zero reviewer attention, and every secret, unsafe-exec, and slop shape it kills is one less diff to read. The pile was the problem.
The agent reads its own evidence
.sober/store.sqlite is local, signed, queryable,
exportable, deletable. Your agent can act on its own run history;
you can audit both of you. The diff never leaves the machine by
default.
Quiet when clean
The engine says Sober clean and stops talking. Alert
floods train humans to ignore alerts; a gate that is silent when
clean keeps its signal when it is not.
What the floor cannot do, and we will not pretend
No semantics
A pattern match cannot tell a broken contract from a working one. Cross-file meaning stays with model review and humans. Sober's model review is single-pass today; the floor narrows the pile, it does not understand the code.
No product judgment
Only your team knows whether a change matches intent. The floor and the model narrow the reading; the decision needs context no engine has.
No liability transfer
The model vendor will not carry your outage, and "the model approved it" will not survive an auditor or a court. The person who ships owns the outcome. That is why Sober's binary has no auto-merge path: never auto-merges, never closes, never deletes is a source-level guarantee, not a policy setting.
Reading, demoted to its proper size
Ask the headline question about a pipeline with no gates and the answer is yes: you will read late, tired, and badly, and the queue will win. Add the deterministic floor and the arithmetic changes shape. The machine-checkable class never reaches review. The semantic class gets advisory review with a known noise rate. The human reads the escalation instead of the emission.
This is more authority for the developer, not less. Reading every emitted line is servitude to the diff. Reading the escalation while holding the only key that can merge, with a floor underneath that never blinks: that is the job. The developer of the near future is not the pipeline's proofreader. They are its final authority.
What this means for your team
-
Install the floor first.
sober hooks install all. Kill the machine-checkable class at the commit boundary, in milliseconds, before it exists as a commit. - Keep merge authority human. If your pipeline ships an auto-approve flag, deadline pressure will find it. Sober refuses the permission, not the pace.
- Give the model the fast loop. Wire your agent to preflight findings and the evidence store, and let it fix model mistakes in-session. The same session still has the context; next sprint does not.
- Budget attention like the resource it is. Read what the gates escalated and trust what they cleared. That discipline is what the reading pile never let you have.