Lessons on Shifting Left with AI Development
How I keep quality up while AI writes most of the code for Read Master (opens in new tab), an AI-powered reading comprehension and retention app — and what happened when the quality process itself became the problem. A companion to Lessons on Parallel Sessions in Claude Code.
How this was written: Claude drafted this article, and I reviewed, edited, and fact-checked it. The incidents and numbers are real and mine.
For seven weeks, my app stored content that was supposed to be encrypted at rest in plaintext.
There was an encryption helper. It worked. The code path that mattered didn't call it. Some session (AI or me, it doesn't matter) wrote a database insert by hand instead of going through the helper, and nothing anywhere objected. Every test passed. Every review looked fine. Seven weeks.
The fix took an afternoon. The interesting part is what I did after the fix: I wrote a lint rule that makes it an error to write that table outside the helper, watched the rule catch the bad pattern, and merged both in the same pull request. That class of bug is now extinct in my codebase: the commit that reintroduces it cannot land.
That's what "shifting left" means in practice, and I've come to believe it's the only quality strategy that survives AI-speed development. This article is the playbook: how findings become guards, how to know your guards actually work, and the part nobody warns you about, which is what to do when the guard collection itself starts eating your velocity.
Why AI changes the review math
"Shift left" is old advice (older than most of the tools (opens in new tab)): catch problems earlier, where they're cheaper. In CI rather than in production. At push time rather than in CI, at commit rather than at push, in the editor rather than at commit — and cheapest of all, make the mistake impossible to write.
AI development changes the math in two ways.
First, volume. When code gets produced at 5–10x your old pace, review becomes the bottleneck, and any recurring mistake recurs fast. A human team member who gets review feedback remembers it next week. An AI session doesn't; tomorrow's session is a fresh instance that never saw today's review. Feedback given in review evaporates; feedback encoded as a check compounds.
Second, and this took me longer to see: AI is better than humans at complying with mechanical gates. A lint error with a clear message gets fixed correctly, immediately, without ego. So encoding your standards as machine-checkable rules pays double with AI. It's the only feedback that persists, and it's the feedback AI handles best.
So the rule I run the whole project on: when a review or audit finding gets fixed, the fix isn't done until that finding's whole class is caught earlier, or made impossible. Fixing the instance is necessary. It is not sufficient.
Lesson 1: Every finding gets a guard, at the earliest rung you can reach
When something slips through, I pick the highest rung on this ladder that fits:
- A custom lint rule. The workhorse. My repo has 125 of them now, each born from a real incident: "every database query must declare which fields it selects," "every background job must declare a concurrency queue," "no raw request-body parsing outside the validation helper."
- A type or schema that makes the mistake unrepresentable. The compiler as reviewer. If the bad state can't be expressed, nobody has to catch it.
- A test that fails on the bad pattern.
- A scaffold default. Bake the right thing into the generator, so new code starts correct instead of getting corrected.
- A written rule. Last resort, for genuine judgment calls.
The ordering is by who does the catching. Rungs 1–4 are machines. Rung 5 is future-you, at 11pm, skimming.
The ladder is one axis: what kind of check catches the mistake. The second axis is where it runs, and the same check can run at several stations. The editor gives you a squiggle at save time. A pre-commit hook runs fast, on staged files only. A pre-push hook runs the scoped test suite. CI runs everything, unscoped, as the backstop that doesn't care how anyone's laptop is configured. Then come required merge checks, and deploy gates at the end; my deploys only ship on green CI, and one post-merge workflow even audits, after the fact, that merges followed policy. The rule of thumb: run each guard at the earliest station where it's fast enough not to annoy, and keep a later station running it too. Every stage exists because something occasionally slips past the one before it.
Try it: next time a code review catches something for the second time, stop and ask which rung could have caught it. If it happened twice, it's a class, not an accident.
Make it automatic: custom lint rules are less work than they sound. ESLint's custom-rule API (opens in new tab) is a visitor over the syntax tree, and your existing rules become templates for new ones; mine mostly start as a copy of the nearest neighbor. I also keep a slash command that takes a finding description and scaffolds the rule plus its test. The AI writes most of the rule; the incident tells it exactly what to match.
Lesson 2: A helper without a guard is a suggestion
The plaintext incident generalizes, and it's the hardest-won lesson here: introducing "the one correct way to do X" does nothing unless something rejects the other ways.
The plaintext story wasn't a one-off; it's a shape. The same repo later lost user-visible text from all six of its export formats, because a hand-rolled database query dropped a field the query-fields helper would have included. Later still, an inline cache key drifted from the key-builder helper, and the price was double-fetching plus a cache that wouldn't invalidate. Three subsystems, one pattern: the helper exists, nothing enforces it, the next producer hand-rolls its own version, and the divergence is silent. A reviewer can't catch this class, because you can't diff a missing call against N correct ones.
So the rule is now structural: a new canonical helper ships with its call-site guard in the same PR. The lint rule usually writes itself, because you already know exactly the shape you're replacing.
Make it automatic: this is a convention until you make it a checklist item and a scaffold behavior. My PR template asks "does this introduce a canonical helper? where's its guard?", and the failure stories above are documented next to the rules they produced, so the next session sees why the pattern exists.
Lesson 3: Watch every guard fail before you trust it
There's an embarrassing counterpart to a growing guard collection: some of your guards don't work, and you can't tell, because a guard that never fires looks exactly like a codebase with no violations.
I know the ways this happens because I've shipped most of them. I shipped a check script that pattern-matched the things it recognized and silently skipped whatever it couldn't parse; it took four rounds of AI audit review to land on the principle that a guard must require what must be true, not approve what it happens to understand. Fail closed. Another time I "confirmed" a new lint rule by grepping for violations instead of running the linter; the grep was subtly wrong, the rule had a bug, and both said everything was fine. Verify a rule with the rule. And for months, files that lived in no package escaped every gate entirely, because all the checks were dispatched per-package. When a check runs "everywhere," ask what's outside "everywhere."
The discipline that catches all of these: when you add a guard, feed it a violation and watch it fail. Red first, then green. Test-driven development, applied to the guards themselves. A recent PR of mine took this seriously enough to test the boundaries — the size cap at exactly its limit and one byte over, all 17 protected path classes individually, both ways git renders a binary diff — and that PR's adversarial review still found a gap in one of the glob patterns. Guards are code. Code has bugs.
Make it automatic: every custom rule ships with its own test file, and those tests run in CI like any other test. An untested guard is a hope.
Lesson 4: Keep a ledger — but measure the right thing
Every finding in my repo gets a one-line ledger entry: what it was, where it's now caught. The ledger's job is to make the trend visible. Shift-left is working if findings trend down, or at least if the same class never appears twice.
Honest admission: my ledger got the first part wrong for months. Raw finding counts (138 one month, 94 the next, 210 the month after) turned out to conflate how hard we audited with how much was wrong. The number that actually matters is different: is this finding a NEW class, or a repeat of a class that already had a guard? A new class means the system is learning. A repeat means a guard failed — the one signal worth alarming on. My counts buried it.
There's a subtler measurement trap underneath: a working guard erases its own evidence. The lint rule fires on your machine, you fix the line, you commit, and no trace remains that the rule ever caught anything. Your most valuable guards look permanently idle. Worth knowing before you conclude a guard is dead weight.
Make it automatic: the ledger is generated from per-PR fragment files (no merge conflicts, no hand-editing), and classifying new-class vs. repeat is a field on the fragment. Cheap at write time, impossible to reconstruct later.
Lesson 5: The counterweight — when shift-left becomes the bloat
Everything above compounds. That's the point. It's also the problem.
This August I measured what the accumulated system actually cost, across my last 40 merged PRs. The visible pipeline was fast: CI in 3–7 minutes, merge in under an hour. The cost was hiding earlier, in-session. A 3-line fix paid the same ceremony as a 1,200-line feature: full multi-model audit with 12-minute-per-leg timeouts, disposition write-ups, attestation, context updates. One PR went seven rounds with a nondeterministic AI reviewer that surfaced newly-worded findings each round. The guard inventory (125 lint rules, ~54 check scripts, 20 CI workflows) only ever grew; nothing retired anything. Roughly a third of my recent PRs were process upkeep rather than product.
The trap has a shape specific to AI workflows, and it's the sentence I'd tattoo on the practice: a human quietly right-sizes process for a small change; an AI session dutifully executes the maximal documented path every time. You built the ceremony for the risky change; the AI performs it for the typo. Right-sizing has to be documented and mechanized too, or it simply never happens.
What fixing it looks like, without weakening anything:
- Tier the ceremony to the diff, and let the tooling decide. Small changes get a light audit, but eligibility is mechanical, not a judgment call made in a hurry: diff under a size cap, no binary changes (git's text diff lies about binaries; a 52-icon commit "measured" 14.6 KB), no file on a risk path (server code, schema, migrations, CI, the guard scripts themselves), fresh base. Ineligible? The runner refuses, names the reason, and you take the full path.
- Close the laundering routes. A light pass must be distinguishable from a full pass, and stale artifacts from an earlier full run can't be allowed to dress up a light one. Fast paths that can impersonate slow paths become the default within a week.
- Put soft caps on unbounded loops, not hard ones. The reviewer-round counter warns loudly at the limit instead of blocking, because the empirical record showed a real finding surfacing in round four. Cap the churn, keep the escape hatch.
- Optimize when failures are discovered and how many nondeterministic rounds run, never whether things are checked. The deterministic gates (types, lint, tests) stay untouched. They're cheap, and they don't argue.
- Schedule the gardening. Guard retirement and consolidation don't happen ambiently; they need their own recurring slot, like dependency updates.
A detail I enjoy: the PR that introduced the fast path took the slow path, because it modified the audit tooling itself — a risk path under its own new rules.
Make it automatic: the tier floor lives in the audit runner as an exported, tested constant, not in prose. Prose is where the AI's judgment goes to agree with whoever wrote last.
Lesson 6: Close the loop — let the process propose its own guards
The flywheel, end to end: incidents become findings, findings become ledger entries, entries become guards, guards get watched failing, and the ceremony gets tiered so the whole thing stays affordable. Three more loops make it self-feeding.
The first runs at the end of a session. After anything gnarly (a debugging saga, a tricky integration), a short reflective pass routes each lesson to exactly one home: a guardable class becomes a new rule, a reusable procedure becomes a skill file the AI auto-loads next time, a subsystem gotcha goes into that package's context doc. The worst outcome for a hard-won lesson is "it stayed in the chat log."
The second runs on a schedule. A weekly cloud job clusters the last few weeks of review comments and reverts, looks for recurring classes, and proposes guards, deduplicated against the ledger so it never re-derives one twice. Report-only; a human approves. The AI reviewing the AI's patterns to suggest what should become un-writable still feels like living in the future.
The third points outward. Those two learn from my own incidents; a small skill I call glean learns from everyone else's. Point it at a URL — a postmortem, a best-practices post, someone's lessons-learned — and it reads the page, checks what my repo already does, and reports only the genuinely applicable deltas, each mapped to where it would land: a rule, a hook, a doc. Most articles yield nothing, which is the point. The filter is "not already done here," so whatever survives is worth a PR. (If you're reading this with an AI assistant handy, this article is glean-able too.)
Try it: you don't need the cloud job to start. End your next debugging session by asking your AI assistant: "what class of mistake was this, and what's the earliest check that would have caught it?" Then make it write that check.
The cheat sheet
- Fixing the instance is necessary, never sufficient. Kill the class.
- Ladder, top rung first: lint rule → unrepresentable type → test → scaffold default → prose.
- Review feedback evaporates; encoded checks compound. AI obeys machines better than memos.
- A canonical helper without a call-site guard is a suggestion. Same PR, both.
- Watch every new guard fail once. An untested guard is a hope.
- Track new-class vs. repeat, not raw finding counts. A repeat means a guard failed.
- A working guard erases its own evidence. Don't confuse idle with dead.
- Tier ceremony to risk, mechanically: an AI will execute your maximal documented process every time, even for a typo.
- Soft caps on nondeterministic review loops; never weaken deterministic gates.
- Schedule guard gardening, or the inventory only grows.
- Harvest inbound too: point the AI at other people's postmortems and keep only what your repo doesn't already do.
Why I care about this
I'm building Read Master (opens in new tab) alone: an AI-powered reading comprehension and retention app, with pre-reading guides that prime you before a book, assessments that check what stuck, and spaced-repetition flashcards so it stays stuck, all wrapped around an accessibility-first reader. Solo plus AI is a real team size now, but only if quality doesn't decay with speed. Shift-left is how I sleep at night, and every rule in that repo is a scar with a test attached.
Read Master launches publicly in September; the waitlist is open (opens in new tab). And if you've built guard systems around AI development — especially if you've hit the bloat wall — I want to hear what you retired, not just what you added.
Claude drafted this article; I reviewed, edited, and fact-checked it. The draft itself was checked by some of the lint rules it describes.