A language model writes correct code. That is not sarcasm — it really does. The problem is that "correct" and "still running in production a year later" are two different things, and the gap between them does not live in the syntax.

Below are ten cases from our own projects. Each one really happened, each looks sensible in the code, and each would pass a review by someone who has not yet watched the same thing break.

1. The cache that ate the disk

A catch-all route behind catalog filters cached every URL permutation a crawler fed it. The cache grew to 38 GB across 1,395,316 files, the server disk hit one hundred percent and took five services down: one application’s API for six hours, the automation engine, analytics, a PostgreSQL database and three systemd units.

The code was correct. The problem appears after weeks, under bot traffic, in the container’s writable layer — invisible in a code review and invisible in tests. The fix was a targeted find -delete from inside the container, with no downtime: rm -rf on the whole directory would have killed the route, because build artefacts sit next to the cache.

2. A file name that was a template

In the same cache sat a file literally named ${encodeURIComponent(e.slug)}.html. An uninterpolated template string that became a real URL.

The syntax is fine, TypeScript does not complain, the site builds. We only saw it while going through the cache after the outage above.

3. A recording the model never heard

An .m4a recording from a phone went straight to the transcription model. The model reads files through libsndfile, which does not know AAC — so the note was produced without the content of the recording.

The code "works": no exception, just an empty result. And a model will happily write a text "about" a file it never read. The fix: ffmpeg in the image, re-encoding to 16 kHz mono, segmenting, and an explicit ban on guessing content in the prompt.

4. The same scan page twice

Attachments were matched by file name. A phone sends every photo as image.jpg, so two pages of a scan produced two copies of the first page.

It looks reasonable: find(a => a.name === p.name). The bug only shows up on real phone data. The fix: match by index, plus [i/N] numbering in the prompt.

5. A shop in the Gulf of Guinea

The content panel stores a missing coordinate as 0, and lat: 0, lng: 0 is a point in the Atlantic off the African coast. The shop pin landed in the ocean.

A typeof === "number" check passes, because zero is a valid number. What you need is a deliberate zero filter in the data layer — and nobody writes one before they have seen a pin float in the ocean.

6. A 57 kB saving that cost 11 points

A monospace font pulled out of preload to "save 57 kB": a badge wrapped onto two lines in the fallback font and the whole page shifted when the real font arrived. CLS 0.22, minus 11 Lighthouse points.

The gain was calculated, the cost was not. Without measurement it sounds like an obvious optimisation. The change went back, and we took the saving another way — a custom font subset.

7. Text that reads well and lies

A content model running at 97.6 percent fidelity twisted readable parts of the source: "Canada" became a biblical place name, a relationship between two people was reversed, song lyrics were attributed to scripture.

The output sounds good and reads smoothly. The error is only visible against the source. It surfaced in a benchmark: two independent runs, counted critical facts, manual verification in the transcript.

8. A tutorial describing an older version

Since version 2.0 the automation engine no longer exposes environment variables in the code node. Every tutorial online shows the older version, so the model reproduces what it learned.

The only thing that helps: testing against the live instance before building the automation.

9. Archiving that deletes history

Archiving a task by moving it to another bucket resets its completion date. The API accepts the request and returns 200, so nothing signals a problem.

The "what got done" history disappears quietly. We archive with a label, not with a bucket.

10. A partial POST that clears fields

In the same API a partial POST clears the fields that are missing from the request body. You send {"priority": 4} and lose the description.

REST without a schema allows it, and the client sees nothing wrong. Caught during the first integration and written into the tool’s canon so it would not come back.

What follows from this

None of these is a syntax error. They are errors of meaning, of data and of operation — specific to a library version, to the hosting, and to real data. A language model has no knowledge of them, because it has no access to your production or your disk.

They are caught by a developer who knows where to look, and who is still around after launch.

That is exactly the difference worth paying for: not writing the code, which AI genuinely does faster, but knowing what to check before that code starts living a life of its own.