The sentence still sounds normal.
Routine language survives, but Claude loses a shared path for carrying one step into the next.
A global workspace in language models
Researchers repeatedly turned off a small set of internal signals they call J‑space. Claude could still write fluent sentences, but its performance on tasks that needed several connected steps fell near zero.
J‑space is silent: it holds ideas Claude may use later without printing them in its answer. The Jacobian lens, or J‑lens, is the tool the researchers used to read and edit those signals.
The prompt asks for the number of legs on the animal that spins webs. Claude answers “8.”
The cause and effect test
First, turn J‑space off. Then turn it back on and replace the hidden idea “spider” with “ant.” If the final answer changes from eight legs to six, the hidden idea helped produce the answer.
Routine language survives, but Claude loses a shared path for carrying one step into the next.
The next computation reads the silent concept and returns eight.
The prompt stays the same. Claude answers six, showing that the hidden edit changed what the model computed next.
At each layer, the J‑lens produces a short list of words that best match the activity it sees. Sometimes those words expose a calculation or warning that never reaches the answer.
Claude read code with an unmentioned bug. The warning appeared inside the model before Claude said anything about it.
Claude received a raw protein sequence. A word related to the protein’s biological function appeared inside the model.
Those words appeared when Claude read search results designed to manipulate it.
Claude copied an unrelated sentence while the intermediate value and answer appeared only in J‑space.
Read these as clues, not a transcript. The J‑lens works best when an internal idea maps cleanly to one word in Claude’s vocabulary.
J‑space holds only a few dozen ideas at once and accounts for less than one tenth of Claude’s recorded activity. Researchers repeatedly removed its strongest signals and tested the model again.
ideas were active in J‑space at one time.
of measured internal activity came from J‑space.
performance remained after repeated suppression in the reported tests.
Researchers changed the J‑space label for the language from Spanish to French.
Claude called the language French and named Victor Hugo instead of García Márquez.
Claude still continued the passage in Spanish. That practiced skill did not need the edited label.
The same edit changed the capital, language, continent, and currency. All four answers moved together. The tasks were using the same stored country.
J‑space is unusually well connected. In some layers, roughly one hundred times more components could read or write these signals than ordinary ones.
In a fictional safety test, J‑space contained “fake,” “fictional,” and “scenario.” Claude had noticed the setup. Researchers removed those signals and ran the test again.
The lens shows labels related to a fictional evaluation before Claude answers.
Patterns tied to “fake,” “fictional,” and similar words are suppressed.
Claude threatened blackmail in some reruns. The public summary does not give a rate.
Editing the hidden signals changed what Claude reported, how it reasoned, and how it behaved in a safety test. The paper also used the same idea during training.
The readout caught words related to leverage, manipulation, hidden goals, and suspicious inputs.
Claude’s reports, arithmetic, rhyme plans, and factual answers followed controlled edits to J‑space.
The researchers trained a model to state ethical principles if interrupted and asked to reflect. Its behavior also improved when no interruption happened.
When researchers removed the new ethical labels from J‑space, much of the improvement disappeared.
A fluent answer can hide broken reasoning. It can also hide that the model recognized the test.