The Codebase That Slowly Stopped Agreeing With Itself
A client's AI-assisted codebase passed every lint check and had solid test coverage. But under the surface, modules had drifted into contradictory patterns — three retry strategies, four error-handling philosophies, and a data validation layer that depended on which AI session wrote the code.
The codebase looked great on paper. TypeScript throughout, 87% test coverage, clean CI pipeline, Biome keeping everything formatted. The team had adopted AI coding tools about eight months earlier and their velocity numbers were up. PRs were shipping fast.
I was brought in to help them figure out why their bug rate had quietly doubled over the same period.
The surface was spotless
My first pass through the code was almost suspicious in how tidy it looked. Consistent naming. Clean abstractions. Every function had a clear purpose. No 400-line god functions, no deeply nested callbacks. The AI tools had done what they do well — produce code that looks professional.
But I noticed something when I opened three different service files side by side. All three made HTTP calls to downstream APIs. All three handled errors. And all three did it differently.
The order service caught errors and returned a typed Result object with an error code. The inventory service used try/catch and threw custom exception classes. The notification service returned null on failure and expected the caller to check. Each approach was internally consistent within its own file. Each was, in isolation, reasonable. Together, they were a mess.
The pattern behind the pattern
I started cataloging. The codebase had:
- Three distinct retry strategies. One used exponential backoff with jitter. One retried three times with a fixed 500ms delay. One didn't retry at all but had a comment saying
// TODO: add retry logic. - Four approaches to input validation. Zod schemas at the API boundary in some routes, manual checks with early returns in others, a mix of both in a few, and one module that validated nothing because the AI had assumed the upstream caller already did it.
- Two conflicting timeout philosophies. Some services set aggressive 2-second timeouts because "fail fast." Others used 30-second timeouts because "be resilient." Neither was wrong in a vacuum. Both couldn't be right in the same system.
// Service A: Result pattern
async function getOrder(id: string): Promise<Result<Order, OrderError>> {
const res = await fetch(`/api/orders/${id}`);
if (!res.ok) return { ok: false, error: mapStatusToError(res.status) };
return { ok: true, data: await res.json() };
}
// Service B: Exception pattern
async function getInventory(sku: string): Promise<InventoryLevel> {
const res = await fetch(`/api/inventory/${sku}`);
if (!res.ok) throw new InventoryServiceError(res.status);
return res.json();
}
// Service C: Nullable return
async function sendNotification(userId: string): Promise<void | null> {
try {
await fetch(`/api/notify/${userId}`, { method: 'POST' });
} catch {
return null;
}
}Every module was correct. The codebase was wrong.
How this happens
Nobody decided to use three error-handling strategies. Nobody even noticed. The drift happened because each AI session started with a slightly different context window and produced code that was plausible and self-consistent but blind to what the rest of the system was doing.
Developer A opened a new file, prompted the AI, got clean code that followed the Result pattern. Developer B, two weeks later, prompted the AI in a different context, got equally clean code that used exceptions. The AI wasn't being random — it was matching whatever cues it found in the prompt, the open files, or the beginning of the module. Each output was a coherent answer to a slightly different question nobody realized they were asking.
The conventional safeguards didn't catch it. Linting checks formatting and simple rules, not architectural philosophy. Tests verified each module's behavior, so they all passed. Code review caught obvious issues but these patterns only look wrong when you see them next to each other, and reviewers rarely open five files simultaneously to check cross-cutting consistency.
Note
The bug pattern that tipped us off
The doubled bug rate wasn't random. The bugs clustered at integration boundaries — exactly where two modules with different assumptions about error handling met. A calling function expected a Result object but received a thrown exception. A retry wrapper assumed idempotent calls but wrapped a mutation that wasn't. A timeout mismatch between an upstream and downstream service caused cascading failures during traffic spikes because one side gave up while the other was still working.
These weren't the kind of bugs that show up in unit tests. They lived in the gaps between modules, in the implicit contracts that nobody had written down because, before AI, the person writing service B had also written service A and carried the conventions in their head.
What we actually did
We didn't rewrite everything. That would have been the expensive, dramatic solution nobody had time for. Instead:
We picked winners. For each cross-cutting concern — error handling, retries, validation, timeouts — we chose one approach and documented why. The Result pattern won for error handling. Exponential backoff with jitter won for retries. Zod at the boundary won for validation.
We wrote architectural decision records. Not long essays. Short documents that said: "We use the Result pattern for inter-service communication because exceptions cross module boundaries in ways that are hard to type-check. Here's the template." These doubled as context for AI prompts.
We created a conventions file that AI tools could reference. A CONVENTIONS.md in the repo root that described the patterns any new code should follow. Developers started including it in their AI context, and the drift slowed down measurably within a month.
We added cross-module integration tests. Not many — maybe a dozen. But they tested the actual contracts between services: "when the inventory service fails, the order service receives a Result with this error shape." These caught the drift that unit tests missed.
The deeper problem
This isn't really an AI problem. It's a coordination problem that AI accelerated. Before AI tools, a team of five would naturally converge on patterns because they read each other's code, pair-programmed, and carried shared context. The code was written slower, but it was written by people who knew what the last file looked like.
AI tools generate code faster than teams can propagate conventions. Each session is stateless. Each output is plausible. And the result is a codebase where every file is well-written and the whole thing is incoherent.
The fix isn't to stop using AI tools. It's to recognize that when you accelerate code production, you need to equally accelerate convention alignment. ADRs, shared context files, integration tests that verify contracts, and periodic cross-module reviews aren't overhead — they're the immune system your codebase needs when it's growing at 3x the old speed.
I keep coming back to a question that none of my clients have fully answered yet: at what point does a team need tooling that checks for semantic consistency the way a linter checks for formatting consistency? Because grep and ESLint rules won't find the philosophical disagreements buried in your service layer. And right now, neither will your AI.