Kaiming He’s team has introduced VISTA, described in the paper VISTA: A Visual Harness for Reasoning in an Interactive World, as a visual-native harness for multimodal models. The framework is meant to let existing models observe an environment directly, preserve raw frames, and revisit, zoom in on, or inspect details during reasoning without training a new base model from scratch.
The central idea is to avoid reducing an environment too early into text or code. VISTA keeps the original visual information available, so a model can bring previously seen frames back into context when a current question requires them. In the paper’s framing, visual experience becomes reusable context rather than a one-time input.
On evaluation, the paper says Claude Opus 5.0 paired with VISTA completed all 25 public ARC-AGI-3 games, reached a relative human action-efficiency score of 100, and used 57.4% fewer game actions than the first-play human baseline. GPT-5.6 Sol also completed all 25 public games and scored 99.
The team also tested the method in browser games, mazes, and connection puzzles, and identified embodied tasks in environments closer to the physical world as a next research direction.
What problem VISTA is trying to solve
From ResNet and Mask R-CNN to MAE, visual perception and representation learning have been a recurring line in He’s research. VISTA shifts that focus to a different question: how a model can accumulate, preserve, and use visual experience during ongoing interaction.
The paper’s answer is a harness. In November last year, Anthropic used long-horizon coding tasks to explain how a harness can help a model continue work across multiple context windows. In that setup, task lists track unfinished items, while progress files and code preserve completed work. When a new context window opens, the model can read those records and recover what has already been done and what still needs to happen.
That text-and-code style of harness has helped complex tasks be broken into steps, checked, and resumed across context windows. The paper argues that it has also contributed, to some extent, to the recent jump in agent capabilities.
But once a model enters a real visual environment, text memory alone is not enough. It has to keep track of object positions and orientations, and it also has to understand what changed in the environment before and after each action. A detail that looks unimportant when first seen may become critical later, and those visual details are hard to capture fully in advance with written notes.
In benchmarks such as ARC-AGI-3, one existing approach is to convert the screen into a text grid made of numbers, then have the model write programs that simulate the environment’s rules and test action plans. The paper notes that ARC-AGI-3 is an interactive visual reasoning benchmark that does not provide rules or goals ahead of time; the agent has to infer how to solve each game through observation and action.
As environments become more complex, though, object appearance, spatial relations, and dynamic changes become harder to translate completely into text and code. As interaction histories grow longer, early frames are also pushed out of context or replaced by summaries.
That leads to the paper’s core question: how can an agent re-examine visual information it saw earlier, at the moment reasoning requires it, without additional training of the base model?

The three core components of the visual-native harness
VISTA answers that by storing every frame returned by the environment outside the context window, including intermediate animation frames generated during actions, and retrieving them only when needed. In ARC-AGI-3, that lets the model compare frames from different moments and zoom in on local regions to inspect orientation markers on small blocks.
The paper calls this mechanism explicit attention over interaction history. Text notes capture the model’s current understanding, while raw frames preserve the evidence it may need to reassess that understanding later. For that to work, the system has to keep the original images and provide tools for retrieval and inspection.
VISTA is built around three components: visual observation, lossless visual memory, and active visual inspection.
Visual observation
Visual observation is the part that presents the environment to the model. In ARC-AGI-3 experiments, VISTA upsamples the official 64×64 screen to a 512×512 PNG image while preserving object color, appearance, and spatial relations, so the model can inspect the environment directly.
Lossless visual memory
Lossless visual memory stores what the model has seen. After each action, VISTA saves all frames returned by the environment, including intermediate animation frames, and indexes them by turn number and frame number. That means the model does not have to decide in advance what is worth remembering. It can preserve the full visual record first and use it later if needed.

Active visual inspection
Active visual inspection gives the model a way to revisit history on demand. Through an inspect tool, it can request frames from a specific earlier turn, crop or zoom into a region, and check details. If it wants to know what exactly changed after an action, it can retrieve multiple frames at once and compare before-and-after states.
The paper describes this as giving the model a visual archive it can consult at any time. It sees the current environment, but it can also go back to earlier observations and use them as evidence for the next move.
How the framework operates during interaction
Under VISTA’s execution flow, each turn starts with the model observing the current screen and available actions. It combines that with prior experience to judge the current state of the environment and form a hypothesis about the next move.
If the available information is not enough, it can call tools to revisit historical frames, zoom into local regions, or even read specific pixel values to gather more evidence. Once it has enough evidence, it does not act immediately. It first predicts what the action should change. After the action is executed, it compares the actual result with that prediction, checks whether its judgment was correct, and updates its understanding of the game’s rules.
The paper presents this as an iterative loop in which the model explores the environment while building experience over time.
To keep long tasks coherent, VISTA also uses two text notes:

- GUIDE.md, which stores reusable rules and experience across levels.
- WORKING.md, which records the current level’s state, progress, and next steps.
When the context window approaches its limit, the model prepares a handoff summary and continues in a new context window. The text notes, action history, and visual archive are all preserved.
One design choice stands out. VISTA keeps the full visual history, but it does not dump every image into the model’s context. In the full setup, after each action the framework shows the model only the final frame by default. Intermediate frames from the action sequence are stored in the visual archive and retrieved only when the model asks for them. That preserves complete visual information without letting historical images consume context space continuously.
Just as important, the framework does not train a new model for this. Environment understanding, action planning, and reasoning are still handled by off-the-shelf multimodal models. The harness manages tool calls, stores visual history, and returns evidence when the model requests it. In other words, VISTA changes how a model gets, stores, and uses visual information rather than changing the model itself.
Results beyond ARC-AGI-3
Beyond the ARC-AGI-3 results cited at the start, the team evaluated VISTA on three more benchmarks: GameWorld, AI GameStore, and BabyVision, covering browser games, mazes, and connection tasks.
Using the same GPT-5.6 Sol model, VISTA raised the success rate on GameWorld’s 170 tasks from 40.0% to 63.3% compared with the official base framework.

On AI GameStore’s 10 games, the aggregate score increased from 47.3 to 140.3, with the human median normalized to 100.
On 39 maze and connection problems selected from BabyVision, accuracy improved from 41.0% to 63.2%. The paper notes that these tasks use only static images, yet the model still benefited from zooming into local regions and checking pixels.
Those results suggest that preserving raw frames and letting a model revisit them on demand can extend what current multimodal models can do. The same VISTA framework, with limited adaptation, was used across different games and puzzles. In the conclusion, the paper points to embodied tasks in more physically grounded environments as the next place to test the approach: settings where a model must preserve, review, and use its own visual experience before deciding what to do next.
Authors and affiliations
The paper lists three co-first authors: Qiushi Han, Keya Hu, and Linlu Qiu. The other two authors are Cathy Wu and Kaiming He.
Qiushi Han, also referred to as Josh Han, is a PhD student at the MIT Operations Research Center and is advised by Cathy Wu.

Keya Hu is a PhD student in MIT’s Department of Electrical Engineering and Computer Science, jointly advised by Kaiming He and Jacob Andreas. The article says she graduated from Shanghai Jiao Tong University’s ACM class and works at the intersection of language and vision, with an interest in building agents that are more data-efficient and generalize better.
Linlu Qiu is a PhD student in MIT’s Department of Electrical Engineering and Computer Science and the Computer Science and Artificial Intelligence Laboratory, advised by Yoon Kim and Jacob Andreas. Her research covers natural language processing and machine learning, and she has previously worked at Google Research and Meta FAIR.
Cathy Wu is an associate professor in MIT’s Department of Civil and Environmental Engineering and the Institute for Data, Systems, and Society. Her research focuses on using machine learning and reinforcement learning to improve complex systems such as transportation. She earned her bachelor’s and master of engineering degrees at MIT and later completed her PhD at the University of California, Berkeley.
Kaiming He is a tenured associate professor in MIT’s Department of Electrical Engineering and Computer Science and a principal author of ResNet, Mask R-CNN, and MAE. He received his undergraduate degree from Tsinghua University and his PhD from the Chinese University of Hong Kong, worked at Microsoft Research Asia and FAIR, and joined MIT in 2024.
The reference link provided in the source is https://arxiv.org/pdf/2610.02200. The original Chinese article was published by the WeChat account Quantum Position and credited to henry.

