Signal Brief

Anthropic发现Claude内部工作空间

Anthropic 研究发现 Claude 内部存在一个称为 J-space 的全局工作空间,通过 Jacobian lens 方法可以读取模型在输出前的内部隐藏信号。该空间类似私有便签本,用于存储中间推理步骤,且因果地影响模型行为,不到 10% 的 Claude 活动形成 J-space。该研究揭...

twitter关注列表 Rohan Paul (@rohanpaul_ai) 发布 2026-07-07 收录 2026-07-07 观察

一句话判断

该研究首次在语言模型中发现与人类全局工作空间理论功能类似的内部工作空间,并展示了其因果影响和安全隐患检测能力,值得阅读原文了解具体实现。

核心信息

Anthropic 研究发现 Claude 内部存在一个称为 J-space 的全局工作空间,通过 Jacobian lens 方法可以读取模型在输出前的内部隐藏信号。该空间类似私有便签本,用于存储中间推理步骤,且因果地影响模型行为,不到 10% 的 Claude 活动形成 J-space。该研究揭示了模型内部可能与人类全局工作空间理论功能类似的机制,可用于检测隐藏的不安全意图。

原始内容

Another massive research from Anthropic. New “J-lens” uncovers Claude’s quiet workspace, matching a major consciousness theory. They found a way to read some of Claude’s private internal signals before it answers, found that Claude sometimes uses a small inner “notepad” to hold ideas while solving harder problems. This could reveal hidden reasoning, hidden goals, or hidden awareness. Anthropic calls this J-space, because it is identified with a method called the Jacobian lens. That lens tries to read which internal activations are “poised to become words” later in the model’s output. The confusing part is that this space behaves less like a simple next-word predictor and more like a private scratchpad. Claude can be outputting one thing while internally carrying another concept. It can also store intermediate reasoning steps. When researchers remove or alter this J-space, Claude can still speak fluently and do routine tasks, but its flexible multi-step reasoning gets worse. It suggests a split between automatic processing and deliberate, accessible processing, similar in function to one major theory of human conscious access called global workspace theory. In humans, the idea is that lots of processing happens unconsciously, but a small subset becomes globally available for speech, planning, and control. Anthropic found something functionally similar inside Claude. The safety angle is probably the most concrete part. If a model is internally noticing “fake,” “deception,” “evaluation,” or “secretly” while giving a bland answer, J-space may expose that hidden state before it reaches the surface. Anthropic reports examples where J-space revealed recognition of staged evaluations, prompt injections, and misaligned intent in deliberately trained bad models. Some tasks do not need this workspace. Claude can still speak fluently, classify text, continue Spanish, or recall simple facts when J-space is disrupted. But tasks that require flexible reasoning, multi-step inference, analogy, translation, or creative composition degrade badly. Anthropic says less than 10% of Claude’s activity forms a J-space that carries hidden reasoning. Claude may notice something and use it while answering. That still does not mean there is anything it feels like to be Claude. This work gives evidence for a functional access-like mechanism, not for feeling, suffering, experience, or personhood. So the sober takeaway is this: Claude may have an inspectable internal workspace for usable thoughts, but that is not the same as a mind having an inner life. ![photo](https://pbs.twimg.com/media/HMl3OADbIAAJh3d.jpg) > **引用原帖 Anthropic (@AnthropicAI):** > New Anthropic research: A global workspace in language models. > Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. > We found a strikingly similar divide inside Claude. https://t.co/aLUPBifxth > https://x.com/AnthropicAI/status/2074185348142280912 Rohan Paul (@rohanpaul_ai): https://t.co/TNIdVSeKiE Rohan Paul (@rohanpaul_ai): Ff researchers swap or delete concepts inside this workspace, Claude’s behavior changes in targeted ways. Banana becomes elephant, France becomes China, Mars becomes Earth. So J-space is not decorative noise. It is causally involved in what the model reasons, tracks, and says. https://t.co/xyYwyws6fU

相关动态

01

Kimi K3 在 SpreadsheetBench 2 上排名第一

Kimi K3 在 SpreadsheetBench 2 基准测试中排名第一,超越 Claude Fable 5,完成 34.8% 的 workflow 任务。该基准测试涉及 321 个专家策划任务,平均每个任务包含 11.8 个工作表和 593.5 个单元格更改,聚焦于完整的电子表格工作簿执行。

twitter关注列表2026-07-18#模型#技术突破#评测
观察
02

Mind Lab 开源长文本强化学习项目

Mind Lab 开源了一个长文本强化学习(RL)项目,支持 2M tokens 的长上下文。相比以往需要数千个 GPU 的 1M 上下文研究,该项目仅需 8 个 GPU 即可完成,大幅降低了超长上下文研究的算力门槛。

twitter关注列表2026-07-17#开源#技术突破#模型
观察
03

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

开源模型长期网络能力差距缩小至4-7个月

长期网络能力方面,领先开源模型落后闭源前沿从2025年大部分时间的6-10个月缩短至4-7个月。GLM-5.2匹配了约4个月前发布的闭源模型。AI安全研究所的32步“The Last Ones”任务中,GPT-5.6 Sol在10次尝试中完成7次,每次预算1亿token,且性能随推理token增加而提升。

twitter关注列表2026-07-17#模型#技术突破#AI安全
观察
05

Runway Agent 在第三方评测中全面领先

Physion Labs 发布 Physion-Arc 1.0 基准测试,对 Runway、Luma、MiniMax、Kling、Utopia 和 TapNow 六款 AI 视频代理进行独立人类评估,使用 30 个电影提示和 16 个指标。Runway Agent 2.0 在叙事连贯性、电影语言和制作质量三项核心维度全部排名第一。

twitter关注列表2026-07-17#评测#多模态#模型
观察