Skip to main content
Mythos

J-space is a small, readable internal workspace that Anthropic researchers identified inside Claude's reasoning process, discovered using a Jacobian-based interpretability technique called J-lens.

J-space holds only a few dozen concepts at a time and accounts for less than a tenth of the model's overall internal activity, yet Anthropic found it carries most of the representations that matter for safety. When 📝Claude reads buggy code, an "ERROR" concept appears in the workspace; when it reads a prompt injection, concepts like "injection" and "fake" surface there. Anthropic has used J-space signals to study prompt-injection recognition, staged-evaluation awareness, fabricated data, and hidden objectives in controlled misaligned model organisms.

The workspace also appears functionally introspectable: when asked what it is thinking about, Claude's answer tends to track what is active in J-space. Disabling the workspace leaves fluency, sentiment reading, and factual recall intact but degrades multi-step reasoning and creative tasks such as poetry, suggesting it functions as a working-memory bottleneck. 📝Anthropic frames the finding through global workspace theory, a leading account of consciousness, while explicitly stating the results show functional access to information rather than evidence that Claude has subjective experience.

Contexts

Created with 💜 by One Inc | Copyright 2026