Anthropic’s research team published A global workspace in language models in July 2026, saying it had identified an observable, intervenable and causally effective region of neural activity inside Claude through a tool called J-lens. The region was labeled J-Space. MarsBit later carried a long-form commentary, originally from the WeChat account Xinzhiyuan and signed by ASI Apocalypse, that used the paper to reopen a broader argument over how large language model interpretability should be approached.
J-Space is presented as a window into internal model states
The article says the finding drew broad interest because it appeared to expose something like a model’s internal monologue during reasoning. In that framing, interpretability research would no longer stop at explaining visible model behavior from the outside. It would move toward real-time observation of internal states.
J-Space is discussed through the lens of the global workspace theory from cognitive neuroscience. The article says that framework treats language-model reasoning as comparable, at least functionally, to information processing associated with human consciousness. On that basis, it describes the work as a methodological and epistemological step forward, and as a new monitoring dimension for AI safety.
Still, the piece argues that the approach remains fundamentally internalist. It treats the central interpretability question as understanding what is happening inside the model, much as a neuroscientist might use fMRI to scan a human brain. In this case, the analogue is scanning a language model’s neural activity with J-lens.
The author argues that this assumes the answer to interpretability is hidden inside the model’s body. Whether a model’s output can be understood, the article says, depends not only on whether its internal states can be observed, but also on how those states relate to states of affairs in the world, to semantic norms and to the cognitive framework of the user. Looking only at neural activity to understand model speech, it argues, is like trying to understand what a person says only by watching electrical activity in the brain. Neural correlations may be visible, but meaning itself remains out of reach.
The article argues that the J-Space path has three built-in limits
The first limit, in the author’s account, is that the model itself becomes both the starting point and the endpoint of explanation. The second lies in the theory transfer. J-Space borrows from global workspace theory, but in doing so, the article says, a subtle category mistake appears: functional isomorphism is treated as if it were epistemic equivalence.
Models do not have subjective experience, the article says. Activation patterns in J-Space are mathematical products of computation, not mental states in any meaningful sense. A model may look similar to a cognitive structure at the level of function, but that does not grant it the same epistemic standing as consciousness.

The third limit is framed as an epistemological one. The article says J-Space is ultimately an engineering-driven project that narrows interpretability into observability and intervention. In a wider tradition of epistemology, explanation does more than that. It places phenomena under broader lawful structures, provides reasons and grounds, and addresses the justification of decisions.
By that measure, the article argues, J-Space may tell researchers what the model is thinking, but not why it thinks in that way, what reasons it relies on, or in what sense those reasons count as good reasons. Those answers are not contained in neural activity patterns themselves.
The broader diagnosis is that J-Space, and much of interpretability work centered on neural networks, treats the model as the only object of explanation. The question begins with the model and ends there too.
From model internals to an ontology of information
The article then sets out a different route. It proposes moving the question of interpretability away from the inside of the model and toward the information the model processes. In its wording, that means shifting from an internalist path modeled on neuroscience to an epistemological path grounded in an ontology of information.
The basis for that move is simple: a large language model is fundamentally an information processor, and both its input and output are text. The meaning of text, which is what people actually need explained, does not reside in neuron activation values. It resides in the relations between symbols and the world, between symbols and knowledge, and between symbols and human practice.
The article uses a basic statement as an example. If a model answers that Paris is the capital of France, the thing that needs explaining is not just which internal region activated. It is also the knowledge system within which that statement is valid, the grounds on which it stands, the reliability and legitimacy of those grounds, and the relation between the output and established human geographic knowledge. None of that, the article says, can be resolved by scanning neural activity.

From there, the author argues that the core question should move from how the model thinks to what kind of information it has processed and what ontological status that information has. The object of explanation should expand from the model itself to the wider information ecology in which the model is embedded, including:
- the structure of training data,
- the way knowledge is represented,
- the flow of information during reasoning, and
- the mapping between output and external knowledge systems.
In that account, J-Space contributes by making it possible to inspect what happens inside a model. But the article says its internalism, its reliance on functional analogy and its engineering-driven narrowing of explanation together create a threefold epistemic limit. Progress, it argues, will require a systematic examination of the information the model processes: where it comes from, how it is structured, how it is represented, how it moves and how it relates to knowledge outside the model.
Kant’s categories are introduced as a philosophical basis for intelligibility
The piece then turns to philosophy and invokes Kant’s Critique of Pure Reason. It begins with an old question: how do humans understand the world at all? Kant’s answer, as the article summarizes it, is that the mind does not passively receive external stimuli. It comes equipped with twelve pure concepts of understanding, or twelve categories, that function as formal conditions for cognition.
Those categories are derived from twelve forms of logical judgment and grouped into quantity, quality, relation and modality. Quantity concerns how much. Quality concerns what something is like. Relation concerns how things are connected. Modality concerns modes of existence.
The article treats Kant’s theory as an ontological commitment about intelligibility. Only what can be placed under those categorical frameworks can become an object of knowledge. The thing-in-itself that lies outside the framework remains unknowable. Applied to AI interpretability, the author says, what is truly explainable is not the physical activation of neurons but the process through which information is categorized and structured into intelligible knowledge.
In that distinction, neural activation belongs to the level of the thing-in-itself, while the meaning of model output belongs to the level of appearances. Meaning can be understood and judged only when it is placed inside a cognitive structure.

That is why the article calls ontology the key to AI interpretability. It gives two reasons:
- At the analytical level, ontology offers a conceptual framework for describing the structured form of the information a model processes. One can ask whether a statement contains an attribution of substance and accident, a judgment of causality, or a modal commitment.
- At the normative level, ontology supplies a standard of evaluation. If the model’s internal representation forms structured patterns corresponding to ontology, its output has a basis for being understood. If it cannot be mapped onto ontology, the article says, then even fluent output remains epistemically uninterpretable.
The article also stresses that using Kant’s categories as a philosophical key does not mean the model must literally possess those categories. Kant’s categories are conditions of cognition for a subject; a model faces a functional implementation problem. A model may distinguish substance, causality or modality through very different neural computation paths while remaining functionally equivalent. Interpretability, in this view, does not require transparency down to every weight. It requires showing that the structures formed at the information-processing level can be mapped onto the categorical framework humans use to understand the world.
Ontology engineering is described as the bridge from theory to systems
After the philosophical section, the article moves into engineering. Ontology tells us what an intelligible structure should look like, it says, but that answer does not by itself become an operational technical system. Without ontology engineering, ontology remains a conceptual exercise suspended in abstraction.
The article defines ontology engineering as the practical field that instantiates philosophical categories as computable, maintainable and traceable technical entities. In AI interpretability, ontology tells us what kind of knowledge structure to ask for. Ontology engineering builds that structure across models, data and systems.
Large language models, the piece says, have given ontology engineering fresh momentum while also creating new engineering problems. Traditional ontology construction depends heavily on domain experts, takes time, costs a great deal and struggles to keep pace with knowledge updates and shifting domains.
LLMs, because they can extract semantic patterns and knowledge relations from massive text corpora, are changing that practice. The article names class definition, relation extraction and attribute construction as core ontology-learning tasks, and says language models can perform large-scale structured extraction with much higher efficiency than manual work.

It goes further, arguing that language models show semantic sensitivity in identifying hierarchical relations, synonym relations and associative relations between concepts. That shifts ontology building from expert handcrafting toward human-machine collaboration and then toward automatically generative construction. The significance, the article says, is not limited to better efficiency. It also gives ontology construction wider scalability and broader domain coverage, opening ontology support to more vertical scenarios and faster-changing knowledge areas rather than only a few critical fields.
How ontology engineering can work back on the model
The article does not frame ontology engineering merely as something helped by large models. It also presents it as a means to strengthen them. Large language models are powerful, it says, but their reasoning process is opaque, their output is hard to verify, and their behavior leans heavily on statistical regularities in training data. Together, those traits create a core obstacle to interpretability.
Within that setup, ontology is given several engineering roles:
- as a provider of structured knowledge, supplying the model with a validated domain knowledge base,
- as a framework for reasoning checks, applying consistency constraints and logical calibration to model output, and
- as an anchoring structure for explanation, mapping each step of model reasoning onto clearly defined classes, attributes and relations.
Once a model’s output can be traced back to ontology entries it depends on, explanation no longer relies on guesses about hidden neural states. It rests on tracing the knowledge structure itself. The article presents this as the engineering basis for moving from seeing through the black box to displaying knowledge structure. The former faces severe technical barriers. The latter, it argues, is a designable, optimizable and verifiable engineering problem.
That is where the article introduces the idea of an AI-friendly ontology framework. Traditional ontologies were built for description-logic reasoners, with syntax, axioms and inference mechanisms optimized for deterministic symbolic deduction. The arrival of large language models changes the consumer of ontology and the scenarios in which it is used.
As a result, the author says, ontology design principles need to shift. Ontology should narrow its responsibility and focus on clearly defining objects, relations, behaviors and rules within a domain, thereby giving the model the semantic skeleton it needs for reasoning. The detailed reasoning process, including how rules are selected, combined and applied, should be left to the language model’s own generalization capacity.

That division of labor brings clear engineering gains, according to the article. Ontology does not need to chase logical completeness and sink into highly complex axiomatization. It can instead prioritize simplicity and maintainability while supplying stable semantic coordinates for model output.
Within that framework, ontology construction should also be optimized for LLM-facing interfaces. The article says class definitions and relation descriptions should be easy for models to understand and use; structured knowledge should be easy to retrieve and cite; and constraint rules should make output validation easier for the model to perform. In that form, ontology is neither a symbolic engine replacing model reasoning nor a static background reference. It becomes explanatory infrastructure embedded in the reasoning chain and available for real-time use and traceability.
Explaining the model or explaining the impact
In its final section, the article returns to the future direction of interpretability. It says the discussion begins with J-Space, moves through Kant’s twelve categories as philosophical grounding, and ends with practical integration between large language models and ontology engineering. The thread running through all of it is that the interpretability crisis of large language models does not stem only from the invisibility of internal mechanisms. It also comes from a long habit of treating explanation as if it were the same thing as transparency.
To illustrate that limit, the article cites Stanislaw Lem’s Solaris and its gelatinous ocean that covers an entire planet, reads human memory and materializes it. The author treats it as an ultimate metaphor for the AI black box: something that can process vast amounts of information and generate results beyond human expectation, while its underlying logic remains unreadable to humans. It is neither benevolent nor malicious. It simply follows laws humans cannot penetrate. In the article’s reading, the more pessimistic point is that the ocean ultimately rejects every human attempt to tame or understand it, hinting at an objective boundary to cognition itself.
That image is used to make a narrower claim. Even if researchers can observe what a model is thinking, they may still fail to understand why it is thinking that way. The real difficulty of interpretability may lie not only in limited tools but in a framework that has become too narrow.
For that reason, the article argues that a practical path beyond the current interpretability impasse should not be confined to opening the black box. It should give equal, or even greater, weight to observing, understanding and controlling model outputs and their real-world effects.

Ontology engineering is presented as the practical framework for that shift. By building AI-friendly semantic skeletons that can be called by the model and traced afterward, the article says, model reasoning can be anchored in clearly defined knowledge structures. That gives the classes, attributes and relations underlying output a formal descriptive basis and a traceable verification path.
When every model statement can be mapped onto an ontological concept framework, explanation is no longer an autopsy of neural network weights. It becomes a display of knowledge structure. When the grounds for a model’s output can be traced and checked at the ontology level, control no longer means forcefully intervening in hidden activations. It means governing the pathways through which information moves.
The article closes by arguing that this shift turns interpretability from an almost impossible technical task into a governance objective that engineering can keep approaching over time. The aim is not complete transparency in the model itself, but making the model’s effects in the real world understandable, traceable and accountable.
In the final paragraph, the piece says Tongfudun has been working in the ontology-engineering and interpretability framework described above, and that its core product, LegionSpace, is built on the same technical ideas. The product is described as an enterprise AI infrastructure with ontology at its core, designed to place the information handled by models and the knowledge they rely on into formal ontology engineering so that each step of reasoning and decision-making is anchored in interpretable knowledge structures. Its stated vision is to make ontology a common language between AI and human understanding, and to turn interpretability into an engineered governance reality.
The article was originally published by the WeChat account Xinzhiyuan and credited to ASI Apocalypse.

