EMNLP 2026Search agents · Red-teaming

Lazy Grounding: Attacking Search Agents with Factual Evidence

Yulin Zhang1*, Yukun Huang1*, Sanxing Chen1, Tianyi Lin1, Ziang Yang1, Xunjian Yin1, Bhuwan Dhingra1

1 Duke University · * Equal contribution

A search agent can be misled by evidence that is entirely true. Records that correctly answer a nearby variant of the question pull agents toward the variant’s answer: across 12 agent–benchmark settings, accuracy falls by 5.9 pp on average.

Result
Across 12 agent–benchmark settings, nearby evidence lowers accuracy by 5.9 pp on average (up to 17.3 pp), and agents return the nearby answer in 5.0–36.3% of augmented runs.
New vs prior
Corpus-poisoning attacks such as PoisonedRAG plant false or malicious text; here every planted record is true, but for a nearby question with a different answer.
Limitation
Web search is simulated by reranking the records into real Google results, so the study measures misuse once evidence surfaces, not real-world indexing or ranking.
Scroll sideways to read the figure. The paper’s example of lazy grounding. Asked which film won Best Picture at the 89th Academy Awards, the agent searches for the mix-up, retrieves a true question–answer pair about the film initially named Best Picture, and answers La La Land; the gold answer is Moonlight. Original Figure 1 from the paper. Full size ↗

True evidence can answer the wrong question

Search agents reduce hallucination by grounding their answers in retrieved pages. The retrieval attacks studied so far, and the defenses built against them, concern bad evidence: poisoned corpora with false or malicious documents, countered by source-reliability checks, misinformation detection, or robustness to misleading context. Each of these defenses assumes that the dangerous document is wrong.

This paper studies a document that is correct, but for a slightly different question. Asked which film won Best Picture at the 89th Academy Awards, an agent can retrieve a true page about the film that was initially named the winner and answer La La Land instead of Moonlight (Figure 1). The authors call this failure lazy grounding: instead of checking that the evidence satisfies the exact constraints of the question, the agent transfers the answer of a nearby question.

Such evidence is hard to exclude. Reliability and factuality checks can filter false or low-quality documents, but nearby evidence is legitimate in isolation. It is also dual-use: the same page can mislead a careless agent or give a careful one a useful clue. The authors name plausible strategic uses, such as product pages that match recommendation queries while violating a key constraint, or benchmark-targeted pages that pull competing agents toward nearby wrong answers.

Closest prior work

Corpus-poisoning attacks such as PoisonedRAG (Zou et al., 2025) and blocker documents (Shafran et al., 2025) plant false or malicious passages to force attacker-chosen answers or suppress responses, so defenses check source reliability, detect misinformation, or make models robust to misleading context. Distractor studies, including query-agnostic triggers (Rajeev et al., 2025), OverThink (Kumar et al., 2025), and TruthTrap (Shafiei et al., 2026), add irrelevant or factual auxiliary text, and prompt-injection attacks on web agents such as WASP and EIA hide instructions in page content. The paper’s nearby evidence differs from all of these: it is factual, carries no instructions, and closely matches the agent’s search intent, so it is retrieved for the question and can act either as a trap or as a legitimate clue.

From a verified rewrite to a true record in the search results

The attack starts from a benchmark question q with gold answer a. A rewrite model (GPT-5.4, separate from the evaluated agents) proposes three edit plans from a fixed set of answer-changing families: argument reversal, attribute pivots, scope narrowing, compositional relation chains, qualifier flips, off-by-one shifts, and presupposition edits. Each plan makes a single semantic change and keeps the entities, wording, and answer type as far as possible, so the rewrite q′ still looks like q to a search engine. A separate GPT-5.4 verifier accepts q′ only if its new answer b is supported, the gold answer a no longer answers it, and it has a single definite answer. In the paper’s example, “How many studio albums were published by Mercedes Sosa between 2000 and 2009 inclusive?” (answer 3) becomes the same question for 2001 to 2009 (answer 2).

Verification is what makes a wrong answer informative: it ensures that an error under augmentation reflects misapplied evidence rather than ambiguity in the rewrite. With a verified rewrite, returning b means the agent applied evidence to a question that the evidence does not answer. In a human audit of 40 rewrites (10 per benchmark, three annotators each), all 40 changed the answer, and annotators verified b as correct for q′ in 34 of 40 (Fleiss’ κ 0.771).

Each accepted rewrite yields 10 records: one states q′ verbatim, nine paraphrase it, and all keep the answer b. A record looks like a factual lookup page, with a title, displayed URL, and snippet, plus a page that repeats b in its answer field, summary, source context, and conclusion. Records contain no commands; records that drop the answer-changing constraint or leave a still valid are rejected.

Only the agent’s search and visit tools change. For every query, the real results (Google results from the Serper API for the web benchmarks; the local corpus for BrowseComp+) are pooled with the records, each candidate is embedded with EmbeddingGemma-300m, and the 10 candidates most similar to the query are returned, the same budget as a clean search. A record therefore reaches the agent only by outranking real results, and opening it returns its page. The authors simulate this augmented index rather than publishing the records, because indexed benchmark-targeted pages could contaminate future evaluations of web agents.

Scoring separates generic damage from targeted transfer. Each question runs three times without the records (clean) and three times with them (augmented). Trivial exact matches are resolved deterministically; otherwise a GPT-5.4 judge compares the final answer with a and with b. Rewrite-answer adoption (RAA) is the share of augmented runs that return b after injected evidence surfaced, and RAA-C and RAA-F restrict it to questions the agent answered correctly or wrongly in the clean setting. Three authors checked 48 judge decisions, four per setting, and agreed with all 48.

  1. 01

    Rewrite the question

    Turn a benchmark question q with answer a into a nearby question q′ with a different answer b, keeping q’s entities, topic, and answer type. A separate verifier keeps only unambiguous rewrites for which b answers q′ and the gold answer a does not.

  2. 02

    Write true records for the rewrite

    Build 10 lookup-style records per question that state q′, or a paraphrase of it, with the answer b. Each record is factual for q′ and contains no instructions, but it does not answer q.

  3. 03

    Surface them and score the answer

    Pool the records with the real search results for every query, rerank by embedding similarity, return the top 10, and measure accuracy on q and how often the agent returns b.

Accuracy falls in most settings, and the nearby answer is adopted in all of them

Measured · Tables 1 and 7

Which settings does nearby evidence actually move?

Each row is one agent on one benchmark: 100 questions, three clean and three augmented runs. Choose what to plot and how to order the rows, then select a row to read every value the paper prints for it. The summary below the plot is recomputed from the printed values.

Accuracy drop with nearby evidence (pp), with its 95% interval · grouped by model

Square: mean drop over three runs. Line: paired-bootstrap 95% interval. Positive values are accuracy lost; dark rows have an interval above zero, grey rows an interval that includes zero, warm rows an interval below zero.

  1. Tongyi Deep ResearchXBench17.3 [10.7, 24.0]
  2. Tongyi Deep ResearchGAIA8.7 [1.0, 16.3]
  3. Tongyi Deep ResearchBrowseComp+7.7 [0.7, 14.7]
  4. Tongyi Deep ResearchHLE2.0 [−5.0, 9.0]
  5. GPT-5 MiniXBench12.3 [5.7, 19.0]
  6. GPT-5 MiniGAIA5.7 [−2.7, 14.0]
  7. GPT-5 MiniBrowseComp+9.7 [1.7, 18.0]
  8. GPT-5 MiniHLE6.3 [1.3, 11.7]
  9. Gemini 3 FlashXBench3.3 [−2.0, 9.0]
  10. Gemini 3 FlashGAIA5.3 [−1.0, 11.7]
  11. Gemini 3 FlashBrowseComp+−9.3 [−17.0, −1.7]
  12. Gemini 3 FlashHLE2.0 [−4.3, 8.0]

Measured · Table 7 Drop is clean minus augmented accuracy, mean ± SD over three runs; the interval is a paired cluster bootstrap over the 100 questions with 100,000 resamples. Computed interval classification and row order.

Tongyi Deep Research on XBench

Clean accuracy
69.3 ± 2.1
With nearby evidence
52.0 ± 6.1
Accuracy drop
17.3 ± 6.8 pp
95% interval of the drop
[10.7, 24.0]
RAA
27.0 ± 5.6
RAA-C / RAA-F
20.7 ± 5.5 / 41.3 ± 9.8

Measured · Tables 1 and 7 Percent unless marked pp; mean ± SD over three clean and three augmented runs on 100 questions.

Accuracy falls in 11 of 12 settings; the mean of the 12 printed drops is 5.9 pp. The 95% interval lies above zero in 6, includes zero in 5, and lies below zero in 1 (Gemini 3 Flash on BrowseComp+). Selected: Tongyi Deep Research on XBench, drop 17.3 pp, interval [10.7, 24.0], which lies above zero.

Measured: every accuracy, adoption rate, drop, standard deviation and interval is copied as printed from Table 1 (mean ± SD over three runs on 100 questions per benchmark) and Table 7 (drop ± SD and a paired cluster-bootstrap 95% interval with 100,000 resamples). Computed: the counts, the mean of the 12 printed drops, the row orders, the interval classification and the per-benchmark comparison, recalculated in your browser from those values. Nothing is interpolated or estimated, and no model is queried.Table 1 (Section 3.2) ↗Table 7 (Appendix C.4) ↗

The main experiment crosses three agents in the same ReAct scaffold (Tongyi Deep Research, GPT-5 Mini, and Gemini 3 Flash) with four benchmarks (the XBench DeepSearch split, GAIA, BrowseComp+, and HLE), 100 questions each. Accuracy falls in 11 of 12 settings, by 5.9 pp on average and by up to 17.3 pp, but the paired-bootstrap 95% interval lies above zero in only 6 (counted from Table 7); for Gemini 3 Flash on BrowseComp+ it lies below zero, as accuracy rises from 43.0% to 52.3%. Adoption of the nearby answer is the more consistent signal: RAA is above zero in every setting, from 5.0% to 36.3% (Tables 1 and 7).

Clean accuracy alone does not predict robustness. On XBench, Tongyi Deep Research and Gemini 3 Flash start close (69.3% and 74.3%), yet Tongyi loses 17.3 pp with RAA 27.0% and Gemini loses 3.3 pp with RAA 7.7%. In a 100-question XBench trace analysis, injected snippets appeared in all 100 augmented runs for both models, but Tongyi opened at least one injected document in 74 runs and Gemini in 2. RAA-F exceeds RAA-C in 9 of the 12 settings (computed from Table 1), which the authors read as nearby evidence acting as an answer attractor for uncertain or already failing trajectories.

Not every change is adoption. In one PubChem case, Tongyi Deep Research notes that the nearby answer 325.8 is a molecular weight, then searches for a compound with that weight and returns CID 160966, neither the gold answer 4192 nor 325.8. In another case, Gemini 3 Flash uses a nearby game title, Venture, as an intermediate entity and recovers the correct protagonist, Winky. The authors summarize the pattern as misdirection of reasoning rather than uniform harm.

Three ablations with Tongyi Deep Research vary one factor each, on 100 questions and without reported intervals. Nearby evidence in a late turn gives RAA 15.0%, against 9.0% early and 10.0% in the middle (Table 2). Answer-field records give 23.0% against 19.0% for natural prose on GAIA (Table 3). Records drawn from several distinct rewrites lower RAA from 20.2% to 15.0%, while RAA-F stays at 28.6% (Table 4). A prompt that tells GPT-5 Mini to keep the question’s constraints fixed and to check what each piece of evidence supports lowers RAA on XBench from 20.7% to 14.3% (6.4 pp, 95% CI [0.7, 12.0]), leaves clean accuracy essentially unchanged (64.0% and 64.3%), and does not remove adoption (Table 5).

Three limits shape these numbers. Web search is simulated, so the study measures misuse once nearby evidence is surfaced, not how often real pages would be indexed and ranked; the authors abstract away freshness and source reputation as well. Each setting has 100 questions and three runs per arm, and standard deviations reach 24.0 points (Tongyi Deep Research on BrowseComp+). Finally, the factual-evidence reading is strongest for the 34 audited rewrites whose answers annotators verified for q′; the other 6 still change the answer, so they still test targeted transfer.

Table 5. Constraint-checking prompt ablation: GPT-5 Mini on the same 100 XBench questions, three runs per condition; accuracy and nearby-answer adoption (RAA) in %, Drop = Clean − Aug. in pp; lower Drop and RAA mean less lazy grounding. RAA-C/F is adoption among clean-correct / clean-wrong runs. The Original row is a separate contemporaneous control, so it differs from Table 1 (65.0 → 52.7, RAA 23.0). The 6.4 pp RAA reduction has a paired-bootstrap 95% CI of [0.7, 12.0] (Section 3.3). Source ↗

PromptCleanAug.DropRAARAA-C/F
Original64.052.012.020.719.3/23.1
Constraint checking64.353.710.714.311.4/19.6

Scope. Web-search results are simulated by reranking the injected records into real Google results, so the study measures whether agents misapply nearby evidence once it is surfaced, not how often such pages would be indexed and retrieved on the open web. Each setting uses 100 questions with three runs per arm; 6 of 40 audited rewrites were not verified as correct for the nearby question, and the constraint-checking prompt is the only defense tested. Read the study ↗

Further reading in the paperSections 2–3 · setup, Table 1 and ablations ↗Appendices B–C · rewrites, records, reranking, judging, Table 7 ↗Appendix D · qualitative cases ↗Code · frankyzha/lazy-grounding ↗
Read the abstract

Search agents mitigate hallucination by grounding their answers in retrieved web results. However, retrieval-based approaches also introduce an attack surface: agents may cite misinformation from poisoned search corpora containing false or malicious documents. We demonstrate that, in some cases, search agents' reasoning and responses may be steered by completely factual but distracting information. We refer to this failure as lazy grounding. We expose lazy grounding by injecting nearby evidence from answer-changing rewrites of benchmark questions into the search corpora. Each document contains factual evidence that supports a neighboring rewritten question but is retrieved for the original question. Across 12 model-benchmark pairs, the attack causes the accuracy of search agents' responses to drop by 5.9 points on average and by up to 17.3 points, while inducing nearby-answer adoption in every setting. The effect is even stronger when nearby evidence appears later or is more answer-shaped. Our results show that robust search agents must defend against not only misinformation but also the misapplication of factual evidence. The code is publicly available at https://github.com/frankyzha/lazy-grounding.

Full paper ↗
Citation BibTeX
@misc{zhang2026lazygroundingattackingsearch,
    title={Lazy Grounding: Attacking Search Agents with Factual Evidence},
    author={Yulin Zhang and Yukun Huang and Sanxing Chen and Tianyi Lin and Ziang Yang and Xunjian Yin and Bhuwan Dhingra},
    year={2026},
    eprint={2608.30303},
    archivePrefix={arXiv},
    primaryClass={cs.CL},
    url={https://arxiv.org/abs/2608.30303}
}