Home Recent Research People Papers Contact

DriftWell: Measuring Semantic Drift in Recursive LLM Generation.

Recursive generation

Meaning moves when AI talks to itself.

Each generated answer becomes the next input. Small semantic changes accumulate, producing measurable displacement from the original meaning.

51 states per chain 1,248 chains 63,648 observations 4 frontier models
Finding 1 · content vulnerability

The fragile content was not slang.

Formal concepts drifted farthest. The result points toward contested definitions and competing legitimate framings as a source of semantic instability.

0.411 formal 0.346 jargon 0.335 slang 0.329 technical
Finding 2 · retrieval grounding

Perfect retrieval is not a universal anchor.

Oracle grounding reduced semantic drift, but the effect varied almost five-fold by content type. Even perfect retrieval helps unevenly, and the endpoint remains content-dependent.

31% slang 12.6% formal 9.4% technical 8% jargon
Finding 3 · attractor dynamics

It looks like drift. It resolves into basins.

Recursive generation was predominantly dissipative rather than chaotic. Trajectories converged toward eight identifiable semantic attractors.

8 basins 69% neutral 30% dissipative 1% chaotic
The grounding paradox

Closer to the source can still mean more switching.

Grounding reduced endpoint drift but increased basin switching. Translational stability and dynamical stability are not the same thing.

22.4% ungrounded rate 28.5% grounded rate +27.5% more switching closer ≠ quieter
Implications for agentic systems

Every handoff is another chance to re-encode meaning.

Planning, memory, summarization, critique, and delegation repeatedly transform state. Long-horizon reliability therefore depends on monitoring how meaning moves, not just whether an answer sounds coherent.

agent memory content summaries task delegation critique loops
Research takeaway

Semantic integrity must be measured over time.

The central takeaway from DriftWell is that semantic drift in recursive LLM generation is not an occasional failure mode but a persistent property of the process. Across models and conditions, meaning moved away from its starting point, yet it did not wander without structure; trajectories repeatedly settled into a finite set of semantic attractor basins. Retrieval grounding reduced endpoint drift, but unevenly across content types, and in some cases increased movement between basins, showing that being closer to the source does not necessarily mean being more semantically stable. For recursive and agentic AI systems, the implication is that semantic integrity cannot be assumed from model quality or retrieval alone; it must be measured over time as meaning is repeatedly re-encoded, transferred, and reused.