Загрузка...

Recursive Language Models: LLMs That Call Themselves

Recursive Language Models (RLMs) are an inference-time scaffold around a base LM that treats an arbitrarily long prompt as a variable in a persistent Python REPL instead of feeding it into the network. The model writes code to peek, decompose, and recursively call itself over slices of the prompt, with only constant-size metadata ever entering its context window. Three design choices make it expressive: a symbolic handle to the prompt, output assembled as a REPL variable, and symbolic recursion inside arbitrarily large loops. The result: inputs beyond 10M tokens, double-digit gains over compaction, CodeAct-with-sub-calls, and Claude Code on four long-context tasks at comparable cost, plus a small Qwen3-8B post-trained as the first natively recursive model. The one-sentence takeaway: move the prompt out of the window and into the environment, then let the LM recurse over it programmatically.

This video breaks down the fascinating "Recursive Language Models" paper, which fundamentally rethinks how Large Language Models (LLMs) handle long contexts. Traditionally, as prompts grow longer, they are fed entirely into the model's context window. This approach scales poorly, as attention mechanisms become computationally expensive, and even models with massive windows struggle to effectively reason over millions of tokens. RLMs propose a radical alternative: stop putting the prompt into the model. Instead, the prompt is bound to a variable in an external environment (a Python REPL).

The LLM acts as an agent that interacts with this environment. It generates code to examine chunks of the prompt, process them, and crucially, it can write code that recursively invokes itself on these smaller chunks. By doing so, the model's actual context window only ever contains small, manageable pieces of metadata and code snippets, never the full prompt. We explore the three core design pillars of RLMs. First, the use of symbolic handles allows the model to reference vast amounts of data without ingesting it. Second, the outputs are constructed programmatically within the REPL, enabling the generation of outputs that exceed the model's output length limits.

Third, and most importantly, we examine the mechanism of symbolic recursion within loops. This allows the model to perform complex, iterative reasoning over data structures of arbitrary size. We review the impressive empirical results: RLMs easily handle inputs exceeding 10 million tokens and demonstrate significant performance improvements on long-context tasks compared to existing methods like simple compaction or standard agentic workflows. The video concludes with a look at their post-training process, where a smaller Qwen model is fine-tuned to natively understand and execute these recursive strategies, proving the viability of this architectural paradigm.

Видео Recursive Language Models: LLMs That Call Themselves канала Latent AI
Яндекс.Метрика
Все заметки Новая заметка Страницу в заметки
Страницу в закладки Мои закладки
На информационно-развлекательном портале SALDA.WS применяются cookie-файлы. Нажимая кнопку Принять, вы подтверждаете свое согласие на их использование.
О CookiesНапомнить позжеПринять