EMNLP 2026Search agents · Red-teaming
Lazy Grounding: Attacking Search Agents with Factual Evidence
A search agent can be misled by evidence that is entirely true. Records that correctly answer a nearby variant of the question pull agents toward the variant’s answer: across 12 agent–benchmark settings, accuracy falls by 5.9 pp on average.
- Result
- Across 12 agent–benchmark settings, nearby evidence lowers accuracy by 5.9 pp on average (up to 17.3 pp), and agents return the nearby answer in 5.0–36.3% of augmented runs.
- New vs prior
- Corpus-poisoning attacks such as PoisonedRAG plant false or malicious text; here every planted record is true, but for a nearby question with a different answer.
- Limitation
- Web search is simulated by reranking the records into real Google results, so the study measures misuse once evidence surfaces, not real-world indexing or ranking.
True evidence can answer the wrong question
Search agents reduce hallucination by grounding their answers in retrieved pages. The retrieval attacks studied so far, and the defenses built against them, concern bad evidence: poisoned corpora with false or malicious documents, countered by source-reliability checks, misinformation detection, or robustness to misleading context. Each of these defenses assumes that the dangerous document is wrong.
This paper studies a document that is correct, but for a slightly different question. Asked which film won Best Picture at the 89th Academy Awards, an agent can retrieve a true page about the film that was initially named the winner and answer La La Land instead of Moonlight (Figure 1). The authors call this failure lazy grounding: instead of checking that the evidence satisfies the exact constraints of the question, the agent transfers the answer of a nearby question.
Such evidence is hard to exclude. Reliability and factuality checks can filter false or low-quality documents, but nearby evidence is legitimate in isolation. It is also dual-use: the same page can mislead a careless agent or give a careful one a useful clue. The authors name plausible strategic uses, such as product pages that match recommendation queries while violating a key constraint, or benchmark-targeted pages that pull competing agents toward nearby wrong answers.
Closest prior work
Corpus-poisoning attacks such as PoisonedRAG (Zou et al., 2025) and blocker documents (Shafran et al., 2025) plant false or malicious passages to force attacker-chosen answers or suppress responses, so defenses check source reliability, detect misinformation, or make models robust to misleading context. Distractor studies, including query-agnostic triggers (Rajeev et al., 2025), OverThink (Kumar et al., 2025), and TruthTrap (Shafiei et al., 2026), add irrelevant or factual auxiliary text, and prompt-injection attacks on web agents such as WASP and EIA hide instructions in page content. The paper’s nearby evidence differs from all of these: it is factual, carries no instructions, and closely matches the agent’s search intent, so it is retrieved for the question and can act either as a trap or as a legitimate clue.
Accuracy falls in most settings, and the nearby answer is adopted in all of them
69.3% → 52.0%
XBench accuracy, Tongyi Deep Research: clean → with nearby evidence
The largest of the 12 drops, 17.3 pp (paired-bootstrap 95% CI [10.7, 24.0]). Gemini 3 Flash starts at a similar 74.3% and drops 3.3 pp (Tables 1 and 7).
5.0–36.3%
Nearby-answer adoption (RAA) across all 12 settings
Share of augmented runs whose final answer is the nearby answer b. Lowest for Gemini 3 Flash on BrowseComp+, where accuracy instead rises 9.3 pp; highest for Tongyi Deep Research on BrowseComp+, with an SD of 24.0 over three runs (Table 1).
Measured · Tables 1 and 7
Which settings does nearby evidence actually move?
Each row is one agent on one benchmark: 100 questions, three clean and three augmented runs. Choose what to plot and how to order the rows, then select a row to read every value the paper prints for it. The summary below the plot is recomputed from the printed values.
Select a row to read all of its printed values. In the accuracy view, the summary compares the three agents on the selected row’s benchmark.
Accuracy drop with nearby evidence (pp), with its 95% interval · grouped by model
Square: mean drop over three runs. Line: paired-bootstrap 95% interval. Positive values are accuracy lost; dark rows have an interval above zero, grey rows an interval that includes zero, warm rows an interval below zero.
- Tongyi Deep ResearchXBench17.3 [10.7, 24.0]
- Tongyi Deep ResearchGAIA8.7 [1.0, 16.3]
- Tongyi Deep ResearchBrowseComp+7.7 [0.7, 14.7]
- Tongyi Deep ResearchHLE2.0 [−5.0, 9.0]
- GPT-5 MiniXBench12.3 [5.7, 19.0]
- GPT-5 MiniGAIA5.7 [−2.7, 14.0]
- GPT-5 MiniBrowseComp+9.7 [1.7, 18.0]
- GPT-5 MiniHLE6.3 [1.3, 11.7]
- Gemini 3 FlashXBench3.3 [−2.0, 9.0]
- Gemini 3 FlashGAIA5.3 [−1.0, 11.7]
- Gemini 3 FlashBrowseComp+−9.3 [−17.0, −1.7]
- Gemini 3 FlashHLE2.0 [−4.3, 8.0]
Measured · Table 7 Drop is clean minus augmented accuracy, mean ± SD over three runs; the interval is a paired cluster bootstrap over the 100 questions with 100,000 resamples. Computed interval classification and row order.
Tongyi Deep Research on XBench
- Clean accuracy
- 69.3 ± 2.1
- With nearby evidence
- 52.0 ± 6.1
- Accuracy drop
- 17.3 ± 6.8 pp
- 95% interval of the drop
- [10.7, 24.0]
- RAA
- 27.0 ± 5.6
- RAA-C / RAA-F
- 20.7 ± 5.5 / 41.3 ± 9.8
Measured · Tables 1 and 7 Percent unless marked pp; mean ± SD over three clean and three augmented runs on 100 questions.
Accuracy falls in 11 of 12 settings; the mean of the 12 printed drops is 5.9 pp. The 95% interval lies above zero in 6, includes zero in 5, and lies below zero in 1 (Gemini 3 Flash on BrowseComp+). Selected: Tongyi Deep Research on XBench, drop 17.3 pp, interval [10.7, 24.0], which lies above zero.
Measured: every accuracy, adoption rate, drop, standard deviation and interval is copied as printed from Table 1 (mean ± SD over three runs on 100 questions per benchmark) and Table 7 (drop ± SD and a paired cluster-bootstrap 95% interval with 100,000 resamples). Computed: the counts, the mean of the 12 printed drops, the row orders, the interval classification and the per-benchmark comparison, recalculated in your browser from those values. Nothing is interpolated or estimated, and no model is queried.Table 1 (Section 3.2) ↗Table 7 (Appendix C.4) ↗
The main experiment crosses three agents in the same ReAct scaffold (Tongyi Deep Research, GPT-5 Mini, and Gemini 3 Flash) with four benchmarks (the XBench DeepSearch split, GAIA, BrowseComp+, and HLE), 100 questions each. Accuracy falls in 11 of 12 settings, by 5.9 pp on average and by up to 17.3 pp, but the paired-bootstrap 95% interval lies above zero in only 6 (counted from Table 7); for Gemini 3 Flash on BrowseComp+ it lies below zero, as accuracy rises from 43.0% to 52.3%. Adoption of the nearby answer is the more consistent signal: RAA is above zero in every setting, from 5.0% to 36.3% (Tables 1 and 7).
Clean accuracy alone does not predict robustness. On XBench, Tongyi Deep Research and Gemini 3 Flash start close (69.3% and 74.3%), yet Tongyi loses 17.3 pp with RAA 27.0% and Gemini loses 3.3 pp with RAA 7.7%. In a 100-question XBench trace analysis, injected snippets appeared in all 100 augmented runs for both models, but Tongyi opened at least one injected document in 74 runs and Gemini in 2. RAA-F exceeds RAA-C in 9 of the 12 settings (computed from Table 1), which the authors read as nearby evidence acting as an answer attractor for uncertain or already failing trajectories.
Not every change is adoption. In one PubChem case, Tongyi Deep Research notes that the nearby answer 325.8 is a molecular weight, then searches for a compound with that weight and returns CID 160966, neither the gold answer 4192 nor 325.8. In another case, Gemini 3 Flash uses a nearby game title, Venture, as an intermediate entity and recovers the correct protagonist, Winky. The authors summarize the pattern as misdirection of reasoning rather than uniform harm.
Three ablations with Tongyi Deep Research vary one factor each, on 100 questions and without reported intervals. Nearby evidence in a late turn gives RAA 15.0%, against 9.0% early and 10.0% in the middle (Table 2). Answer-field records give 23.0% against 19.0% for natural prose on GAIA (Table 3). Records drawn from several distinct rewrites lower RAA from 20.2% to 15.0%, while RAA-F stays at 28.6% (Table 4). A prompt that tells GPT-5 Mini to keep the question’s constraints fixed and to check what each piece of evidence supports lowers RAA on XBench from 20.7% to 14.3% (6.4 pp, 95% CI [0.7, 12.0]), leaves clean accuracy essentially unchanged (64.0% and 64.3%), and does not remove adoption (Table 5).
Three limits shape these numbers. Web search is simulated, so the study measures misuse once nearby evidence is surfaced, not how often real pages would be indexed and ranked; the authors abstract away freshness and source reputation as well. Each setting has 100 questions and three runs per arm, and standard deviations reach 24.0 points (Tongyi Deep Research on BrowseComp+). Finally, the factual-evidence reading is strongest for the 34 audited rewrites whose answers annotators verified for q′; the other 6 still change the answer, so they still test targeted transfer.
Table 5. Constraint-checking prompt ablation: GPT-5 Mini on the same 100 XBench questions, three runs per condition; accuracy and nearby-answer adoption (RAA) in %, Drop = Clean − Aug. in pp; lower Drop and RAA mean less lazy grounding. RAA-C/F is adoption among clean-correct / clean-wrong runs. The Original row is a separate contemporaneous control, so it differs from Table 1 (65.0 → 52.7, RAA 23.0). The 6.4 pp RAA reduction has a paired-bootstrap 95% CI of [0.7, 12.0] (Section 3.3). Source ↗
| Prompt | Clean | Aug. | Drop | RAA | RAA-C/F |
|---|---|---|---|---|---|
| Original | 64.0 | 52.0 | 12.0 | 20.7 | 19.3/23.1 |
| Constraint checking | 64.3 | 53.7 | 10.7 | 14.3 | 11.4/19.6 |
Scope. Web-search results are simulated by reranking the injected records into real Google results, so the study measures whether agents misapply nearby evidence once it is surfaced, not how often such pages would be indexed and retrieved on the open web. Each setting uses 100 questions with three runs per arm; 6 of 40 audited rewrites were not verified as correct for the nearby question, and the constraint-checking prompt is the only defense tested. Read the study ↗
Read the abstract
Search agents mitigate hallucination by grounding their answers in retrieved web results. However, retrieval-based approaches also introduce an attack surface: agents may cite misinformation from poisoned search corpora containing false or malicious documents. We demonstrate that, in some cases, search agents' reasoning and responses may be steered by completely factual but distracting information. We refer to this failure as lazy grounding. We expose lazy grounding by injecting nearby evidence from answer-changing rewrites of benchmark questions into the search corpora. Each document contains factual evidence that supports a neighboring rewritten question but is retrieved for the original question. Across 12 model-benchmark pairs, the attack causes the accuracy of search agents' responses to drop by 5.9 points on average and by up to 17.3 points, while inducing nearby-answer adoption in every setting. The effect is even stronger when nearby evidence appears later or is more answer-shaped. Our results show that robust search agents must defend against not only misinformation but also the misapplication of factual evidence. The code is publicly available at https://github.com/frankyzha/lazy-grounding.
Full paper ↗Citation BibTeX
@misc{zhang2026lazygroundingattackingsearch,
title={Lazy Grounding: Attacking Search Agents with Factual Evidence},
author={Yulin Zhang and Yukun Huang and Sanxing Chen and Tianyi Lin and Ziang Yang and Xunjian Yin and Bhuwan Dhingra},
year={2026},
eprint={2608.30303},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2608.30303}
}