Prior state enters context
The request contains the old design, the change, and the desired outcome.
History leakage
An artifact exposes knowledge of how it changed, although its audience needs only the current state.
History leakage is a working term from my own engineering playbooks, not an established research label. It is written up as the keep-history-out-of-the-artifact principle, with dehistorize as the cleanup pass.
The request contains the old design, the change, and the desired outcome.
The model predicts from every visible token. Prior values and names remain available.
Repeating the before and after is an easy way for the agent to show its work.
The process appears where the user expected only the finished state.
02 / See it happen
Switch views. Both versions can be factually true. Only one belongs in the product.
New score: 50 (was 35)
Updated with the improved scoring model.
History leaked
The requester cares that the score moved from 35 to 50. The sales rep opening this account cares that the score is 50.
Correct destination: Put the comparison in the reply, commit, release note, or audit log.
03 / Research map
History leakage is a local engineering term. The evidence beneath it comes from several distinct lines of work. Open a card to see the mechanism, study, finding, limit, and design consequence together.
The short version. Human studies explain how prior values, familiar procedures, examples, and private knowledge constrain later judgment. LLM studies show comparable sensitivity to preceding code, numerical hints, and prior scores. The connection to product-copy residue is a reasoned application, not a result any paper tested directly.
A familiar procedure becomes a mental set. It can block a simpler method or persist even when it stops working.
Does learning one successful procedure make people keep using it when a simpler method exists, or when the learned method fails?
Participants solved water-jar problems. Repeated training problems favored one multi-step formula. Later problems allowed a direct solution, and one could not be solved by the learned formula.
Many trained participants persisted with the familiar formula. Control participants, who did not receive the training set, found the direct method more often.
The task is a constrained laboratory puzzle from 1942. It demonstrates a human mental set, not an LLM mechanism or software-design behavior.
The old implementation can become a procedure to preserve instead of a state to replace. That is fixation applied to revision work.
An initial value becomes a reference point. Later estimates move away from it, but often not far enough.
Which shortcuts do people use for uncertain judgments, and when do those shortcuts create predictable errors?
The paper synthesizes experiments on representativeness, availability, and adjustment from an anchor. In the best-known anchoring task, a random starting value shifted estimates of the share of African countries in the United Nations.
People often make insufficient adjustments from an initial value. The estimate stays closer to the anchor than the evidence warrants.
The paper concerns human judgment. Later work debates whether all anchoring comes from insufficient adjustment or whether anchors also make consistent information easier to retrieve.
The previous score, name, or layout can pull the output toward itself because it is already present, even when the current requirement does not need it.
People who know more struggle to reconstruct the perspective of someone who lacks that information.
Can better-informed people reproduce the judgments of people who lack their private information?
Participants predicted uninformed estimates of corporate earnings. The researchers compared individual judgments with trading markets that rewarded accurate anticipation of the uninformed group.
Informed participants could not fully ignore what they knew. Market forces reduced the measured bias by about half, but did not remove it.
The study tests information asymmetry in experimental economics, not communication design or language models.
The requester and the artifact user do not share the same context. The artifact fails when the builder writes as if they do.
Seeing an example narrows later design work. Incidental details can survive alongside useful ones.
Does seeing an example design constrain later conceptual design, including the transfer of weak or flawed features?
Engineering students solved design problems after some participants saw example solutions. The examples included details that did not satisfy the brief.
Participants who saw the examples repeated more of their details and flaws. The examples narrowed the range of designs they produced.
The experiments involve human engineering design and a small set of conceptual tasks. Fixation strength varies with expertise, task, and how examples are presented.
Showing the old artifact gives the agent more than requirements. It also gives the agent names, components, and layouts to repeat in the replacement.
Irrelevant preceding code acts as an anchor. Generated code borrows its structure even when the requested function points elsewhere.
Can cognitive-bias experiments help researchers find predictable qualitative failures in open-ended model output?
The authors transformed HumanEval prompts for Codex and CodeGen. They prepended irrelevant functions or related but wrong anchor functions, then measured functional accuracy and copied anchor elements.
Irrelevant framing lines appeared in 81% of Codex outputs and 70.7% of CodeGen outputs, compared with 4.5% and 0% on original prompts. Functional accuracy fell by 22.3 to 30.5 points for Codex. Related wrong code also pulled generated solutions toward its structure.
The models are older code generators. The transformed prompts are synthetic stress tests. The work shows predictable sensitivity, not a complete account of internal cause.
This paper gives the closest code-level evidence. Irrelevant prompt history can enter the generated artifact even when it conflicts with the function signature.
Low and high hints become reference points for later model estimates. The study also asks whether prompting can neutralize them.
How do factual, expert, and irrelevant numerical hints affect GPT-3.5, GPT-4, and GPT-4o?
The study uses 62 everyday prediction questions. Each question receives a control hint and either a low or high anchor. The authors collect 30 answers per treatment and compare distributions with t-tests.
Most response shifts matched the anchor direction. Expert anchors produced especially consistent effects. Chain-of-thought, principle generation, explicit ignore instructions, and reflection did not reduce the effect on the expert-anchor subset.
The dataset is small, numeric, and adapted from a prediction game. The analysis relies on repeated samples from three GPT model families and does not establish broad causal mechanisms.
Telling a model to ignore the anchor is a weak safeguard. Remove irrelevant prior information or supply independent evidence before generation.
Semantic and numerical anchors influence answers. Activation patching probes where that influence appears inside one open model.
How common is LLM anchoring, where does it act inside a model, and which interventions reduce it?
SynAnchors contains 60 semantic questions and 40 numerical questions. The study tests 11 models from 0.5B parameters through GPT-4o and reasoning models. It also uses activation patching on Llama 3.1 8B.
Every tested model shows some anchoring. The reported total occurrence ranges from 22% for Qwen3 thinking mode to 60.5% for Qwen 2.5 0.5B. The activation-patching probe locates the strongest anchor influence before the middle layers. Tested mitigations reduce some results but do not eliminate the pattern.
The dataset is generated and human-curated. The mechanism study covers one open model and ten selected questions. The result supports a shallow processing hypothesis, not a universal circuit.
Independent criteria and explicit reasoning can help, but the effect remains. A clean review pass still matters.
A prior score included as metadata changes the next judgment, even when the text being judged stays fixed.
Does prior-evaluation metadata compromise an LLM judge's independence when it scores a revised answer?
Eight models scored 20 fixed texts under no metadata, revision-only metadata, and metadata with a prior score. The study attempted 192,000 evaluations and obtained 185,271 valid responses. It also tests 441 human-labeled industry samples.
Seven of eight judges shifted toward the prior-score condition. On industry data, anchored metadata cut error correction by 47.9%. It changed 10.18% of baseline-correct cases to the assigned wrong label. Chain-of-thought and a warning did not give a general numerical mitigation.
The main benchmark has one screened answer for each of 20 tasks. The full metadata condition changes more than the score field, and the token-level probes cover selected model-task pairs.
This study most closely matches the lead-score example. A previous score placed in context can contaminate the next evaluation. Excluding irrelevant metadata remains the safest control.
Words such as new, now, and currently depend on when the reader encounters them. Their reference point decays while the product state remains.
Describe how the product works. Avoid time-bound terms in product and reference documentation unless a date or version supplies the missing reference point.
Timeless prose reduces maintenance and does not assume that the reader knows an earlier product version.
Release notes, migration guides, and dated announcements exist to describe change. Before-and-after language belongs there.
This is practitioner guidance, not an experiment. It gives the closest established writing rule for keeping construction history out of current-state documentation.
The applied rule matches the history-leakage test. Product surfaces describe the state. Change records describe the delta.
04 / Recent field notes
These posts describe operational responses to overloaded context. They add current vocabulary and practice to the research map, but they are observations and recommendations rather than controlled evidence.
Accumulated distractions and dead ends can poison a long conversation. The working name is "context rot."
The important point is not only that models forget material near a context limit. Stale attempts remain available and continue to compete with the current task. History leakage is one visible symptom. The agent starts treating abandoned work as live design material.
For long-context writing, aggressively clean the source material and compress it into annotated notes.
This treats context selection as part of authorship. Compression should retain decisions, evidence, and unresolved questions while removing discarded phrasings and obsolete alternatives. A summary that repeats every detour in shorter form preserves the same anchor problem.
Subagents are useful partly because each one starts with fresh context instead of inheriting the full working history.
Fresh context is an isolation boundary. A narrowly briefed reviewer can inspect the finished artifact without sharing the builder's attachment to its earlier form. That makes separate review useful even when parallelism is not the goal.
Infinite context is not an unqualified benefit. Old information already leaks into current answers and makes the work tiring.
More memory increases recall and interference at the same time. The design question is therefore not how to retain everything. It is which state the next step needs, and which history should remain in a log rather than in the active prompt.
When a chat becomes stubborn and attached, compact the work into a Markdown handoff and start a clean conversation.
A handoff creates a deliberate reset. The new conversation receives the goal, current state, constraints, and evidence without replaying every failed attempt. The quality of the handoff matters because a chronological summary can smuggle the same fixation into the fresh thread.
Design separate context layers for humans and agents. Do not assume that the same history belongs in both.
The requester, builder, reviewer, and artifact user need different information. Treating one transcript as the source for all four collapses those audiences. History leakage occurs when process context crosses that boundary and becomes product content.
The summaries and interpretations are paraphrases from an authenticated X session on 9 October 2026. Profile images are locally cached copies of the posters' public avatars. Follow each source link for the original wording and surrounding thread.
05 / Detect and prevent
Prompt reminders help, but several studies above show that "ignore the anchor" is weak. Remove irrelevant input, then inspect the finished artifact.
Separate the audiences. The requester needs the delta; the product user needs the state.
Write the current requirement by itself. Remove old values, abandoned names, and rejected designs when they are not needed.
Draft from the current state. Describe what the artifact is and does, not how it arrived there.
Search for residue such as new, now, was, previously, V2, and legacy.
Run the from-scratch test. Ask whether the line would exist without access to the edit history.
Move valid before-and-after detail to the reply, commit, pull request, changelog, migration guide, or audit log.
A match is a review prompt, not a verdict. Release notes and migration guides are supposed to describe change.
Describe the finished artifact in the artifact. Put the change history in the reply, commit, release note, migration guide, or audit log.
Back to the top