Google DeepMind, the University of Maryland, and the University of Virginia have introduced "Dream-RSI," a framework designed to achieve recursive self-improvement (RSI) in AI agents by optimizing their exploration policies offline. Traditionally, coding agents rely on hardcoded or fixed exploration policies to decide whether to exploit a promising path or explore new alternatives. Dream-RSI bypasses this limitation by caching the history of all previous attempts—including code, scores, and execution failures—on a disk. This cached data acts as a zero-cost "replay simulator," allowing the agent to "dream" by testing thousands of alternative exploration policies against historical data without rerunning the underlying LLM or evaluator. Once the optimal policy is identified, it is deployed on the next live run. In benchmarks across eight algorithmic and mathematical tasks, Dream-RSI powered by Gemini-3.1-Pro successfully generated a Lasso solver that outperformed Python's scikit-learn library in just 317 attempts, compared to 550 for static policies and over 51,000 for previous state-of-the-art methods. However, critics note this is not true recursive self-improvement in the classical sense, as the underlying model's weights remain static; instead, it represents an advanced, highly efficient search optimization technique.
Dream-RSI optimizes AI exploration policies by utilizing cached historical attempt logs as a zero-cost replay simulator. The framework allows AI agents to test thousands of search strategies offline without rerunning the underlying large language model.
In testing, Dream-RSI wrote a Lasso solver that surpassed Python's scikit-learn library in approximately 317 attempts. The system relies on a detailed prompt instructing the agent to analyze all past attempts and avoid repetitive minor code tweaks.
Critics argue that Dream-RSI is not true recursive self-improvement because the underlying model's weights and capabilities remain unchanged. Recent major AI mathematical breakthroughs have relied on static models wrapped in custom orchestration harnesses rather than self-improving weights.
Chapter guide
Worth noting
- The video contains a paid sponsorship segment for Blacksmith and its product CodeSmith.
- While the paper claims to achieve recursive self-improvement, the video notes that the underlying model's weights remain static, meaning the AI itself does not actually become more intelligent.