ATTENTION DEFICIT · EP 004FIELD NOTE / KSI

Knowledge-centric self-improvement

Let the worker reset. Keep the evidenceunder review.

A new agent tried each task. When it finished, it left notes and evidence for the next agent.

Based on Wang et al.’s KSI experiments. Other systems change the agent itself or rewrite its instructions. KSI keeps both fixed and updates shared notes with supporting evidence.

WHAT CHANGESGENERATION N → N+1
A FRESH AGENT EACH TIME
question → attempt → notes
WHAT STAYSShared notes and evidence

The notes keep useful tactics and failed ideas. When agents disagree, both sides remain visible.

How the claim changes

One failure produces a rule that is far too broad.

The first agent treats one result as a general rule. Other agents find a counterexample, try the rule on new tasks, and narrow the claim.

00 · FAILURE

One result becomes a universal rule.

The explanation sounds certain, but it rests on one failed task.

01 · SAME TASK REVIEW

A peer brings a counterexample.

Agents test the claim against evidence from the same task.

02 · TEST ON OTHER TASKS

Does the rule travel?

New tasks show where the rule works and where it breaks.

03 · REWRITE THE CLAIM

The claim gets specific.

The reviewer adds limits and links the claim to the evidence behind it.

04 · FRESH AGENT

The agent resets. A narrower claim remains.

A fresh agent receives the shared notes without keeping the previous agent’s memory.

ONE CLAIM · FIVE CHECKSTOO BROAD
FAILEDATTEMPT TASKFORUM CROSS-TASKFORUM DISTILLCLAIM FRESHAGENT SHARED KNOWLEDGE BASE marker cells never control geometryevidence: one failed taskstatus: too broad
FAILED ATTEMPTOne failure becomes a rule that is too broad.
WHAT STAYSThe shared notes are still empty.
0 / 4
failurereviewtestrewritereuse
AGENTfailed once
SHARED NOTESone guess, not tested
This slider simplifies one example from an ARC‑AGI grid puzzle. Agents asked whether the position and shape of marker cells controlled movement. It shows the paper’s review steps, not an actual transcript of a model’s reasoning.
01 · WHAT STAYS

The agent starts over. The notes do not.

KSI uses the same model and solver setup for every task. Each task starts with a new agent. Only the shared notes change.

NOTES IN THE FINAL VIEW7,289

The notes include tactics, ideas that failed, and warnings that apply only to certain cases.

MEMORY KEPT BY THE AGENTNONE

Each agent reads the notes, tries one task, adds evidence, and exits. The next agent starts clean.

02 · COST AND RESULTS

KSI solved more tasks and cost less in the authors’ tests.

Both systems used Haiku 4.5. They differed in how they carried lessons from one task to the next.

Selected ARC‑AGI comparisons
HyperAgents
KSI
ARC‑AGI‑1 solve rate
70%
86.7%
ARC‑AGI‑1 evaluation cost
$234
$76
ARC‑AGI‑2 solve rate
60%
82.7%
ARC‑AGI‑2 evaluation cost
$188
$80

The KSI numbers average three runs. Each HyperAgents number comes from one rerun. KSI won every reported comparison, though more runs could change the size of the gap.

03 · NOTES BETWEEN MODELS

Notes written by Haiku helped GPT, and GPT’s notes helped Haiku.

Every tested pairing beat the same model working without shared notes. One Haiku to GPT ARC‑AGI‑1 result varied by ±12.6 across three runs, so the exact gain is still noisy.

DIRECTIONBoth ways

Every tested model pairing did better with shared notes than without them.

NOISIEST RESULT±12.6

Haiku’s notes helped GPT on ARC‑AGI‑1, but the three runs were far apart.

The next agent inherits the evidence, not the old agent’s confidence. A better counterexample can still overturn the claim.