Biuron

Program · AGT·02

Reasoning & Assistants

We build agents that plan, reason, and act, turning intent into reliable work across research, engineering, and everyday operations. The hard problem is not fluency but reliability over long horizons: staying coherent across hundreds of steps without drifting from the goal.

ProgramAGT·02
DisciplineLanguage · Planning · Tool use
Publications2
StatusActive

What we work on

A program is a small set of hard, coupled problems. These are the ones this group is working on now.

01Long-horizon planning without drift
02Verifiable tool use and self-correction
03Memory that persists across sessions
04Calibrated uncertainty and deferral
01

Reliability is the research problem

A model that is right ninety percent of the time is not ninety percent of a reliable assistant, because errors compound across a plan. We study the mechanisms of drift and design systems that notice when they are wrong, recover, and know when to ask rather than guess.

02

Reasoning you can inspect

An assistant that acts in the world must expose its reasoning as structured, checkable steps rather than a black box. We build agents whose plans can be audited, replayed, and constrained, so trust is earned through transparency rather than assumed.

03

Grounded in tools, not just text

Real work happens through tools: code, search, data, and instruments. Our agents are trained to use them deliberately, verify the results, and compose them into workflows that hold up outside the demo.

Signals from the program

91%
success rate on our internal agent evals
200+
steps held on a single coherent plan
4.2s
median assistant response time

Selected writing from this program

← Previous program · FIN·01Financial IntelligenceNext program · BIO·03Living Systems Models

Collaborate

Work on this with us.

We hire researchers and engineers who want to push one of these programs forward, and partners who want to put the results to work.