arXiv preprint 2026Long-context processing
Coding Agents are Effective Long-Context Processors
Off-the-shelf coding agents can process long contexts as files. With only a path and a question, Codex and Claude Code search, slice and script through the text, and exceed the best published scores on four of five long-context benchmarks.
- Result
- The best coding-agent configuration exceeds the best published score on four of five benchmarks, e.g. BrowseComp-Plus accuracy 88.50 vs 80.00; LongBench is the exception (62.50 vs 63.30).
- New vs prior
- RAG and ReAct agents reach the text through fixed retrieval tools, and Recursive Language Models prompt a model to call itself over snippets; here an unmodified coding agent explores the text as files with its own commands and scripts.
- Limitation
- Scores come from 200-question samples without reported variance, and three of the five best published scores were measured on full test sets.
Long contexts are easy to access but hard to process
Frontier models now accept millions of tokens, and drop-in long-context models can outperform retrieval-augmented generation. The paper argues that this scales access rather than processing: accuracy falls as inputs grow, and attention leaves little trace of which parts of the context produced an answer. Retrieval pipelines make access explicit, but a fixed retrieval stage cannot easily let one finding shape the next query, which multi-hop questions need.
The paper starts from an observation about coding agents: they are trained on repositories with long files and nested directories, and they already search, slice and transform text with programs. In the paper’s Figure 2, asked for the last spell Vax’ildan casts in each episode of a 385K-token Oolong-Real transcript, the agent wrote a script that split the transcript into episodes, kept lines spoken by or mentioning the character, and matched a list of spell keywords. Many episodes came back “Unknown”, so it printed lines from the failing episodes, found phrasings such as “uses Lay on Hands”, added 12 domain-specific spells and new patterns, and re-ran the extraction.
The question is whether this transfers from one case to long-context tasks in general. If a corpus or a long document is laid out as files, can an off-the-shelf coding agent, with no task-specific training and only a minimal prompt, outperform long-context models, retrieval pipelines and search agents?
Closest prior work
Standard RAG retrieves through a fixed, shallow pipeline that struggles with multi-hop questions, and code-executing agents such as CodeAct reason with code but, according to Zhang et al. (2025a), handle long-context tasks poorly. Agentic RAG systems reformulate queries iteratively but are typically fine-tuned or trained with reinforcement learning for web search or open-domain QA. The closest, concurrent work is Recursive Language Models (Zhang et al., 2025a), which also treats long text as an external environment, explored from a Python REPL through recursive LLM sub-calls under a specialized system prompt. This paper instead hands the text to an unmodified coding agent with no task-specific prompting and lets it use file-system tools such as grep and sed and scripts of its own.
Large gains where search or code helps; parity where reading suffices
88.50 vs 80.00
BrowseComp-Plus accuracy (%) · Codex without a retriever vs best published
The 100K-document corpus holds 750M tokens. The published 80.00 was measured on the full set; on the paper’s 200-question sample, a GPT-5 ReAct search agent scores 72.50 and RAG 65.00 (Table 1).
88.50 → 78.50
BrowseComp-Plus accuracy (%) · Codex without vs with a BM25 retriever
Across the nine Codex retriever comparisons in Table 1, a retriever lowers the score eight times and ties once. With a retriever, native search commands per query fall from 14.92 to 9.84 (BM25) or 8.33 (Gemini embeddings; Table 4).
Measured · Tables 1 and 5 · Computed · differences and relative changes
What does the 17.3% average compare?
The abstract reports that coding agents outperform the published state of the art by 17.3% on average, a mean of per-benchmark relative changes. Choose a coding-agent row and a reference row from Table 1: each benchmark shows both scores, their difference, the relative change and the cost per query from Table 5, and the summary recomputes the average.
- Mean relative change
- +17.3%
- Median
- +10.6%
- Agent higher
- 4 of 5
- BrowseComp-PlusCodex (No Retriever)88.50Best Published80.00*Difference+8.50 ppRelative+10.6%Cost per query$0.703 vs –
- Oolong-SynCodex (No Retriever)71.75Best Published64.38Difference+7.37 pointsRelative+11.4%Cost per query$0.194 vs –
- Oolong-RealClaude Code + BM2537.46Best Published24.09Difference+13.37 pointsRelative+55.5%Cost per query$0.380 vs –
- LongBenchClaude Code + BM2562.50Best Published63.30*Difference−0.80 ppRelative−1.3%Cost per query$0.319 vs –
- NQCodex (No Retriever)56.00Best Published50.90*Difference+5.10 ppRelative+10.0%Cost per query$0.111 vs –
Selected comparison
Best per benchmark vs Best Published: higher on 4 of 5 benchmarks, lower on LongBench. Mean relative change +17.3%, matching the abstract; the largest term is Oolong-Real at +55.5% (+13.37 points). Rounded, the five terms are Figure 1’s labels (+11%, +11%, +56%, −1%, +10%). Best per benchmark takes Codex (No Retriever) on BrowseComp-Plus, Oolong-Syn and NQ; Claude Code + BM25 on Oolong-Real and LongBench. 3 of the 5 reference scores (*) were measured on the full test set, not on the paper’s 200-question sample. Table 5 lists no cost for best published results.
Coding agent (paper’s method)Reference row* best published score measured on the full test set
Measured: scores from Table 1 (accuracy for BrowseComp-Plus and LongBench, exact match for NQ, Oolong’s score; every row except Best Published was run on the paper’s random 200-question sample per benchmark) and average cost per query in US dollars from Table 5. Best Published scores marked * were measured on the full test set and are given in the paper for reference; Table 1 cites them to openJiuwen (2025), Zhang et al. (2025a, the Recursive Language Models paper, which used Oolong-Synthetic’s trec_coarse subset), Singh et al. (2025, the GPT-5 system card), Comanici et al. (2025, Gemini 2.5) and Hui et al. (2025, Interact-RAG). Table 1 prints 64.38 for both RLM and Best Published on Oolong-Syn. Context lengths are as printed in Table 1. Computed here: differences, relative changes (agent − reference) / reference, their mean and median over the benchmarks where both rows have a score, and the best or strongest row per benchmark.
Table 1 compares all methods on the same random 200 questions per benchmark. Codex without a retriever beats every re-run baseline on all five benchmarks, but the margins differ widely: 16.00 pp over the ReAct agent on BrowseComp-Plus (88.50 vs 72.50) and 10.66 points over RLM on Oolong-Real (33.73 vs 23.07), against 0.50 pp over full-context GPT-5 on LongBench (61.50 vs 61.00) and 0.67 pp over RLM on NQ (56.00 vs 55.33). The abstract’s 17.3% instead averages relative changes over the best published scores, three of them measured on full test sets, and its largest term is Oolong-Real, where Claude Code with BM25 scores 37.46 against 24.09.
The file-structure ablation tests the role of the corpus layout directly. On a 100-question BrowseComp-Plus subset, one file per document instead of a single JSON dictionary raises accuracy from 83.0 to 89.0 without a retriever and from 86.0 to 90.0 with Gemini embeddings, and makes no difference with BM25 (82.0 for both; Table 2, below). Command counts suggest why: with folders, the agent runs sed 4.48 times per query instead of 0.61 and rg 18.46 times instead of 31.05 (Table 3), reading selected line ranges instead of repeatedly scanning the whole corpus.
Retrieval tools rarely help the agent. In Table 1, adding a retriever to Codex lowers its score in eight of nine comparisons and ties once, on LongBench with Gemini embeddings; on BrowseComp-Plus, BM25 costs 10.00 pp. With a retriever available, the agent issues fewer native search commands per query (14.92 without, 9.84 with BM25, 8.33 with Gemini embeddings; Table 4). The paper hypothesizes that the retriever becomes the default discovery tool and displaces broader exploration, so ranking errors hide relevant context, and leaves the mechanism to future work.
The agent’s behavior also changes with the task (Figure 5): BrowseComp-Plus accounts for most of the text read, the two Oolong benchmarks for nearly all generated code, and LongBench draws little tool use, with accuracy close to full-context GPT-5 (61.50 vs 61.00). The extra work has a price. Codex without a retriever costs $0.703 per BrowseComp-Plus query, against $0.237 for the ReAct agent and $0.111 for RAG, but on Oolong-Syn it costs $0.194 against $1.421 for full-context GPT-5 (Table 5).
Table 2. Accuracy (%) of Codex with GPT-5 on a 100-question subset of BrowseComp-Plus, with the corpus stored as one text file per document in a folder or as a single JSON dictionary keyed by document id; higher is better. Change (folder minus single file) is computed here. The subset is smaller than Table 1’s 200-question sample, so the no-retriever folder score (89.0) differs from Table 1 (88.50). Source ↗
| Retriever | Folder structure | Single file | Change |
|---|---|---|---|
| No Retriever | 89.0 | 83.0 | +6.0 pp |
| Gemini Emb. | 90.0 | 86.0 | +4.0 pp |
| BM25 | 82.0 | 82.0 | 0.0 pp |
Scope. Each benchmark is a single random sample of 200 questions (100 in the file-structure ablation), and no run-to-run variance is reported. Three of the five best published scores were measured on full test sets, and Claude Code was run only on Oolong-Real and LongBench. The folder-versus-single-file ablation changes only the corpus layout; that the benefit comes from familiarity acquired in code training is a hypothesis the paper does not test directly. Table 1 lists the NQ corpus as 3T tokens, while the authors’ code README gives three billion. Read the study ↗
Read the abstract
Large Language Models (LLMs) have demonstrated remarkable progress in scaling to access massive contexts. However, the access is via the latent and uninterpretable attention mechanisms, and LLMs fail to effective process long context, exhibiting significant performance degradation as context length increases. In this work, we study whether long-context processing can be externalized from latent attention into explicit, executable interactions, by allowing coding agents to organize text in file systems and manipulate it using its native tools. We evaluate off-the-shelf frontier coding agents as the general interface for tasks that require processing long contexts, including long-context reasoning, retrieval-augmented generation, and open-domain question answering with large-scale corpus contains up to three trillion tokens. Across multiple benchmarks, these agents outperform published state-of-the-art by 17.3% on average. We attribute this efficacy to two key factors: native tool proficiency, which enables agents to leverage executable code and terminal commands rather than passive semantic queries, and file system familiarity, which allows them to navigate massive text corpora as directory structures. These findings suggest that delegating long-context processing to coding agents offers an effective alternative to semantic search or context window scaling, opening new directions for long-context processing in LLMs.
Full paper ↗Citation BibTeX
@misc{cao2026codingagentseffectivelongcontext,
title={Coding Agents are Effective Long-Context Processors},
author={Weili Cao and Xunjian Yin and Bhuwan Dhingra and Shuyan Zhou},
year={2026},
eprint={2603.20432},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.20432}
}