ts/Tanmay’s place
Writing · July 2026

Regeneration is a rebase.

What building a merge engine for generated code taught me about trusting git history.

Fern generates SDKs from API specs. That is the whole product in one sentence. A customer describes their API, and we produce idiomatic client libraries in TypeScript, Python, Java, Go, and a few other languages. The generated code is good, but it is never the last word. Customers add helper methods. They patch a type the generator got wrong, or add the retry policy their infra team insists on. Then the API changes, the SDK regenerates, and every one of those edits is gone.

For years the official answer was a file called .fernignore. List a file there and the generator will never touch it again. It works, in the way that unplugging a smoke alarm works. You now own the whole file forever, you stop receiving generated improvements to it, and the list only grows. One customer had 66 files in .fernignore to protect a single type patch. When I saw that number I stopped filing it under annoyances.

The insight

The fix came from noticing what regeneration actually is. When new generated code replaces old generated code, and a human has edited the old code in the meantime, you are not looking at an overwrite problem. You are looking at a rebase. Git solved this decades ago with the 3-way merge, and the mapping is exact.

BASE   = the tree the generator produced last time
OURS   = the tree the generator produced today
THEIRS = what the human made of the old tree

Run diff3 over each touched file and most edits reapply cleanly, because generators tend to change different lines than humans do. This is Replay, the system I built at Fern. On every regeneration it detects the commits a customer made since the last generated commit, stores each one as a patch in a lockfile, and replays them on top of the fresh output.

Conflicts belong in editors, not pull requests

The first real design decision was what to do when the merge does conflict. The obvious move is to commit conflict markers and let the customer sort it out. We never do that. A regeneration lands as a pull request, and a pull request full of <<<<<<< is radioactive. Nobody merges it, it goes stale, and now the customer is behind on SDK updates too. Instead the conflicting patch is marked unresolved in the lockfile, the PR stays clean and mergeable, and the customer runs fern replay resolve locally, where their editor shows an ordinary merge conflict. An open conflict carries forward across regenerations until someone resolves it, so nothing is dropped while it waits.

The boundary problem

All of that was the easy half. The hard half was a question that sounds trivial. Where does the last generation end and the customer's work begin?

The first version answered it the obvious way. When we generate, record the commit SHA. Next time, diff from that SHA. I trusted that pointer, and git taught me not to, four separate times.

A squash merge creates a new commit and abandons the original, so the recorded SHA still exists but no branch can reach it. A force push rewrites it away entirely. Garbage collection eventually deletes whatever nothing references. And a shallow clone, which is what every CI system uses, never fetched the old commits in the first place.

Each of these produced its own strange bug report. The most memorable came from the squash merge case. When the recorded SHA was unreachable, the detector fell back to scanning commits, and on one large repository it walked straight past the beginning of Fern's involvement and captured years of unrelated history as customer edits. The result was a lockfile holding 135 patches, 101 of them anchored to commits that no longer existed, and a pull request whose description exceeded GitHub's 65,536 character limit for PR bodies. I did not know that limit existed. I do now.

Derive, don't store

The lesson underneath all four failures is the same. Any state you store about git history is a cache, and it will eventually disagree with the history itself. Squash merges, force pushes, GC, and shallow clones are just how people use git.

So the rewrite stopped storing the answer and started deriving it. On every run, Replay walks git log --first-parent from HEAD and classifies each commit with a narrow predicate that recognizes generated commits. The first match is the boundary. There is nothing to go stale because nothing is remembered between runs. That one change deleted about three thousand lines of recovery code whose only job had been to repair the stored pointer after git moved underneath it. Deleting code that defends a bad assumption is the best feeling in this job.

Two holes worth describing

Deriving the boundary opened its own set of holes, and two changed how I think about git.

First, --first-parent follows the mainline of the branch, which is what you want, until a customer pushes a fix directly onto a generated pull request and merges it with a merge commit. Their fix now lives on the second parent side of history, and a mainline walk never sees it. The detector had to switch to examining every non-merge commit individually, so each human commit becomes its own attributed patch regardless of which side of a merge it arrived on.

Second, a patch stored months ago can reference a base tree that no longer exists on the machine doing the merge, because of GC or a shallow clone. You still have the patch itself though, and a unified diff contains more information than people give it credit for. Context lines plus removed lines reconstruct what BASE looked like around the edit. Context lines plus added lines reconstruct THEIRS. You can rebuild enough of both sides from the diff alone to run a real 3-way merge against the new output. I called this ghost reconstruction.

You cannot mock git

A mocked git returns what you told it to return, and every bug above came from git doing something I had not imagined. So the engine's test suite, more than 800 cases now, runs against real repositories. It squash merges, force pushes, truncates clones with --depth=1, garbage collects, and then asserts that no customer edit was lost. There are dedicated testbed repositories whose entire purpose is reproducing the exact history shapes we first met in the wild. The rule now is that when git surprises you, the surprise becomes a permanent test before the fix merges.

Where this generalizes

Replay is generally available and running for companies like ElevenLabs and Auth0, and the engine is public on npm as @fern-api/replay. The launch post covers the product side.

The engineering lesson travels further than SDKs. If your system remembers a fact about an external source of truth, whether that is a SHA or a cache key, you have created a second copy of reality, and the two will eventually diverge. When you can afford to derive the fact each time, derive it. Git is fast enough. So is almost everything else.

I work at Standard Template Labs. Before that, I worked at Fern (now part of Postman). I built SDK generators and agent tooling there. More about me or play something instead.

TANMAY’S PLACE / VISITOR

Your memory card

Six souvenirs are hidden around the site. You’ll find them by looking around.

0 / 6 souvenirs collected

Find all six for a small surprise back at home.

Save it or share it. It’s made in your browser, and nothing gets uploaded.

Your card stays in this browser. No account needed. Clearing browser data clears your card.