Biuron

Language agents that plan over long horizons without drift

R. Adeyemi · S. Park · Reasoning Group

arXiv:2606.11902

Agents built on language models are fluent but fragile: coherence decays as plans lengthen, and small errors compound. We study drift as a measurable quantity and identify its sources in the interaction between memory, planning, and tool use.

Our training regime introduces explicit checkpoints where the agent verifies its own state against the goal, recovers from detected errors, and defers when uncertain. On internal evaluations spanning research and engineering tasks, agents hold coherent plans past two hundred steps.

We argue that reliability, not fluency, is the binding constraint on useful agents, and that it is best addressed as a first-class research problem rather than a prompting detail.

This is a plain-language summary. The full manuscript and supplementary material are available on request.

Collaborate

Want to build on this?

We work with researchers and institutions extending our published work into the real world.