The atlas demo toggling a lead scoring card from 'New score: 50 (was 35)' to 'Lead score: 50'
The same score, two audiences. Flip the toggle in the atlas.

Ask an agent to change a scoring model in your dashboard and you’ll find “New score: 50 (was 35)”. Ask it to rename a function and you get scoreV2 sitting next to the legacy_weights it didn’t dare to delete. Ask for a docs edit and the page says “now supports Parquet”. None of that is what you asked for, but the agent loves to prove it did the work, inside the product, even though no user ever saw the old version (or no user has ever used the product at all…).

I’ve started calling it history leakage. The artifact telling you how it changed when the user only needs to know the current state. To catch it, ask your agent: Would a capable implementer, given only the current requirements, write this line? So I wrote a /dehistorize skill. A cleanup pass for artifacts that already shipped with the leak. Finding leaks is the easy half. The hard part is not deleting the changelog so most of the skill is about what to keep: changelogs, release notes, ADRs, audit logs, and any feature that compares data over time rather than over revisions. A “(was 35)” that refers to last week’s lead score is a feature. The same string pointing at what the old code computed is a leak.

The principle surprisingly common in humans too. You might know it as ‘anchoring’ or ‘Einstellung’. So I also built a small page on where the evidence comes from. Luchins (1942) and Tversky and Kahneman (1974) to anchoring in LLM-as-a-Judge systems in 2026.

Read the history leakage atlas →