One result becomes a universal rule.
The explanation sounds certain, but it rests on one failed task.
Knowledge-centric self-improvement
A new agent tried each task. When it finished, it left notes and evidence for the next agent.
Based on Wang et al.’s KSI experiments. Other systems change the agent itself or rewrite its instructions. KSI keeps both fixed and updates shared notes with supporting evidence.
The notes keep useful tactics and failed ideas. When agents disagree, both sides remain visible.
How the claim changes
The first agent treats one result as a general rule. Other agents find a counterexample, try the rule on new tasks, and narrow the claim.
The explanation sounds certain, but it rests on one failed task.
Agents test the claim against evidence from the same task.
New tasks show where the rule works and where it breaks.
The reviewer adds limits and links the claim to the evidence behind it.
A fresh agent receives the shared notes without keeping the previous agent’s memory.
KSI uses the same model and solver setup for every task. Each task starts with a new agent. Only the shared notes change.
The notes include tactics, ideas that failed, and warnings that apply only to certain cases.
Each agent reads the notes, tries one task, adds evidence, and exits. The next agent starts clean.
Both systems used Haiku 4.5. They differed in how they carried lessons from one task to the next.
The KSI numbers average three runs. Each HyperAgents number comes from one rerun. KSI won every reported comparison, though more runs could change the size of the gap.
Every tested pairing beat the same model working without shared notes. One Haiku to GPT ARC‑AGI‑1 result varied by ±12.6 across three runs, so the exact gain is still noisy.
Every tested model pairing did better with shared notes than without them.
Haiku’s notes helped GPT on ARC‑AGI‑1, but the three runs were far apart.
The next agent inherits the evidence, not the old agent’s confidence. A better counterexample can still overturn the claim.