At the intersection of AI and crypto, researcher Dwarkesh Patel has proposed a paradigm shift in AI training: moving from pre-deployment training to continuous post-deployment learning. He argues that while reinforcement learning from verifiable rewards (RLVR) can effectively optimize model performance in controlled test environments, it fails to adapt to the dynamic and unpredictable nature of real-world tasks.
Real-World Learning: Stepping Out of the Sandbox
Patel emphasizes that AI training must evolve from "closed sandboxes" to "open worlds." Models should continuously accumulate experience from user interactions, system feedback, on-chain operations, and other real-world scenarios, feeding that experience back into their own weights. This concept is described as "malleability" — the ability of a model to remain receptive to new experience even after deployment, rather than being frozen in its training state.
OPSD and "Dreaming": Efficient Mechanisms for Continuous Learning
To enable real-world learning, Patel proposes two core technical paths: on-policy self-distillation (OPSD) and "dreaming" simulations. OPSD allows the model to distill new data using its own current policy while avoiding catastrophic forgetting. "Dreaming" involves generating high-value trajectories in simulated environments, allowing the model to trial-and-error repeatedly and accelerate experience consolidation. Together, these mechanisms could enable AI to self-iterate without reliance on human annotation.
Implications for Crypto: A Training Revolution for On-Chain Agents
In the crypto domain, AI-driven agents are already emerging in DeFi, DAO governance, and NFT trading, but most training still relies on offline RLVR or human feedback. If Patel's OPSD and real-world learning paradigm materialize, future on-chain agents could learn and optimize strategies directly from real transactions, governance votes, and cross-chain interactions, dramatically improving adaptability and autonomy. Moreover, crypto infrastructure — on-chain data, oracles, decentralized computation — naturally provides a rich environment and verification mechanism for such continuous learning.
Currently, Patel's vision remains theoretical, but his critique of the prevailing training paradigm has sparked deeper discussion in the crypto community about "evolvable agents." If real-world learning and OPSD can be combined with decentralized verification networks, they could give rise to next-generation, self-evolving on-chain AI systems.

