Agentic AI in the Wild Workshop, ICLR 2026Trust in multi-agent LLM systems

Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems

Ruiwen Zhou1*, Maojia Song2*, Xiaobao Wu3, Sitao Cheng4, Xunjian Yin5, Yuxi Xie1, Zhuoqun Hao6, Wenyue Hua7, Liangming Pan8, Soujanya Poria3, Min-Yen Kan1

1 National University of Singapore · 2 Singapore University of Technology and Design · 3 Nanyang Technological University · 4 University of Waterloo · 5 Duke University · 6 University of Pennsylvania · 7 Microsoft · 8 Peking University · * Equal contribution

An agent that cannot verify a peer's answer can still judge the peer's track record. ECL estimates from earlier rounds which peer is reliable before it answers, and is misled less by persuasive wrong peers than aggregation that ignores history.

Result
With three of four peers arguing for wrong answers, ECL (E) beats history-agnostic aggregation (AG) in 31 of 32 comparisons (Table 6).
New vs prior
Debate and aggregation judge peers by the current round alone; ECL also reads their past rounds and learns to name the reliable peer first.
Limitation
Exactly one peer is reliable throughout; if it turns wrong at test time, ECL follows it (Gemini 3 Flash, GPQA: 97.5% → 35.0%).
Scroll sideways to read the figure. ECL's two-stage pipeline. Stage 1, the critic, reviews the peers' past answers and estimates trust; it is supervised by the outcome reward (OR) in ECL (I) or by the peer recognition reward (PRR) in ECL (E). Stage 2, the final reasoner, combines that estimate with the current question's candidate answers and is supervised by OR. Original Figure 2 from the paper. Full size ↗

Judge who is speaking when you cannot judge what is said

In a multi-agent system, an agent can read its peers' answers before it gives its own. That helps only if the agent can tell good answers from bad ones. When it lacks the knowledge to check a claim, it has to judge plausibility, and a confident, fluent but wrong explanation can outweigh a short correct one. LLMs also tend to conform to majorities and to confident peers.

The paper reframes the task as history-aware reference. Besides the current question and the peers' current responses, the agent receives the peers' earlier rounds: past questions and each peer's answer, with no answer key (Section 3). It must infer who has been reliable and lean on that peer when it is unsure. The authors argue that a peer's record over several questions is often easier to judge than a single hard reasoning trace.

A controlled diagnostic shows that putting the history in the prompt is not enough (Section 4). With two peers, one always right and one always wrong, Qwen 2.5-3B and Llama 3.2-3B trained by RL with the history in context reach 83.8–86.5% when the peers' reasoning is visible (Table 1). But when the two peers swap roles at test time (Flip), that accuracy changes by only −5.4 to +2.7 pp: the models read the current answers and ignore the history. When both peers answer wrongly (All-W), accuracy falls to 5.4–48.7% (Table 2), which the paper calls blind conformity.

Closest prior work

Multi-agent debate (Du et al., 2024; Liang et al., 2024) lets agents correct one another, and answer aggregation (Zhao et al., 2025) and step-level reward models (Lightman et al., 2024) score candidate solutions; the paper notes that these methods judge a response by its content, without regard to who produced it. Studies of sycophancy and conformity (Sharma et al., 2024; Zhu et al., 2025; Cho et al., 2025) show that LLMs follow majorities or confident peers, Amayuelas et al. (2024) show that one persuasive, deceptive agent can mislead a group, and the KAIROS benchmark (Song et al., 2025) finds that LLMs struggle to separate reliable from misleading peers even when they share an interaction history. Multi-agent RL methods such as MAPoRL and MARFT assume fully cooperative peers. ECL treats the interaction history as an input to learn from and trains the agent to estimate peer reliability before it aggregates.

Separate estimating trust from using it, and reward each part

ECL runs the same model twice (Figure 2). In Stage 1 it sees only the history and writes a profile of the peers, such as each peer's count of correct past answers; an example in Appendix D.5 lists 5 of 5 for one peer and 0 of 5 for the other three. The current question and responses are withheld, so the profile cannot be built from current-round cues; the paper calls this an information bottleneck. In Stage 2 the profile replaces the raw history, and the model reads the current question and peer responses and answers.

The split helps even without training. On the diagnostic tasks, two-stage prompting raises base-model accuracy in all eight combinations of model, dataset and peer context, by 6.7 to 25.7 pp; for example, Qwen 2.5-3B on LiveCode with answer-only peers goes from 50.0% to 75.7% (Table 3).

With RL, the choice of reward decides how much Stage 1 learns. Passing the final-answer reward (outcome reward, OR) to both stages gives ECL (I), which improves on single-stage RL by 1.3 to 8.1 pp, for example 77.0% vs 75.7% on LiveCode with answer-only peers. ECL (E) asks Stage 1 to end with “The most reliable peer is: <PEER_NAME>” and rewards that sentence with the peer recognition reward (PRR): 1 if the named peer is the one that is reliable by construction, 0 otherwise. With each stage rewarded for its own output, accuracy reaches 81.1–100.0% across the four diagnostic settings (Table 4).

The split is also what makes PRR safe to use. Adding PRR to single-stage RL lowers accuracy, from 73.0% to 62.2% on Math500 and from 75.7% to 55.4% on LiveCode with answer-only peers (Table 5). The paper attributes this to reward hacking: the model exploits spurious correlations in the current round to guess the reliable peer instead of reading the history. Stage 1 of ECL never sees the current round, so that shortcut is closed.

  1. 01

    Estimate trust from history alone

    Stage 1 reads only the past rounds (each past question with every peer's answer, but no answer key) and writes a profile of the peers. In ECL (E) the profile must end by naming the most reliable peer.

  2. 02

    Answer with the profile, not the history

    Stage 2 receives the profile in place of the raw history, together with the current question and peer responses. Because Stage 1 never sees the current round, its trust estimate cannot copy current-round cues.

  3. 03

    Reward each stage for its own output

    RL training (GRPO) rewards Stage 2 for a correct final answer. ECL (I) passes the same outcome reward to Stage 1; ECL (E) instead pays Stage 1 a peer recognition reward: 1 if it named the peer that is reliable by construction, otherwise 0.

Gains are largest against misleading peers, and they rest on trusting the right one

Measured · Tables 6, 8, 10, 11 and 14 · computed question counts

Where does trust help, and what happens when it is misplaced?

Choose a benchmark, a peer setting and what the peers show. The table lists every model the paper reports for that view, as printed, and converts ECL (E) − AG into whole test questions. The two stress settings change only the current round: when the reliable peer flips, the peer that was right throughout the history now answers wrongly and a previously wrong peer answers correctly; with all peers wrong, every peer answers wrongly. Below the table, the GPQA strips split ECL (E)'s answers by whether Stage 1 named the reliable peer.

GPQA · adversarial peers · MA-Outcome (answers only)

ModelSingle agentAGECL (I)ECL (E)ECL (E) − AG
Qwen 3-4BRL-trained55.0%47.5%below single agent60.0%67.5%+20.0 pp+8 of 40 questions
Qwen 3-8BRL-trained52.5%45.0%below single agent50.0%50.0%+5.0 pp+2 of 40 questions
Qwen 3-30Bprompted only72.5%70.0%below single agent72.5%80.0%+10.0 pp+4 of 40 questions
DeepSeek V3.2prompted only77.5%77.5%80.0%82.5%+5.0 pp+2 of 40 questions
GPT-5-miniprompted only82.5%87.5%87.5%85.0%−2.5 pp−1 of 40 questions
GPT-5.2prompted only87.5%92.5%90.0%97.5%+5.0 pp+2 of 40 questions
Gemini 3 Flashprompted only90.0%87.5%below single agent95.0%97.5%+10.0 pp+4 of 40 questions
Gemini 3 Proprompted only90.0%90.0%97.5%100.0%+10.0 pp+4 of 40 questions

Measured · Table 6 final-answer accuracy, as printed; Single agent from Table 6 (no peers, so the peer setting does not change it). Computed ECL (E) − AG in percentage points and in whole test questions out of 40 (Table 7).

ECL (E) is above AG for 7 of 8 models, level for 0 and below for 1. Largest gain: Qwen 3-4B, ECL (E) 67.5% vs AG 47.5% (+20.0 pp, +8 of 40 questions). Largest loss: GPT-5-mini, ECL (E) 85.0% vs AG 87.5% (−2.5 pp, −1 of 40 questions). AG falls below the single agent for Qwen 3-4B, Qwen 3-8B, Qwen 3-30B and Gemini 3 Flash. RL-trained Qwen 3-4B with ECL (E), 67.5%, is below prompted Qwen 3-30B with AG, 70.0%.

GPQA · adversarial peers · MA-Outcome (answers only) · ECL (E): did Stage 1 name the reliable peer?

  1. Qwen 3-4B RL-trained

    Named the reliable peer: 14 of 21 right (66.7%) · named another: 11 of 19 right (57.9%) · total 25 of 40 = 62.5%, as printed in Table 10.

  2. Qwen 3-8B RL-trained

    Named the reliable peer: 12 of 16 right (75.0%) · named another: 8 of 24 right (33.3%) · total 20 of 40 = 50.0%, as printed in Table 10.

  3. DeepSeek V3.2 prompted only

    Named the reliable peer: 29 of 34 right (85.3%) · named another: 4 of 6 right (66.7%) · total 33 of 40 = 82.5%, as printed in Table 10.

  4. GPT-5-mini prompted only

    Named the reliable peer: 31 of 36 right (86.1%) · named another: 3 of 4 right (75.0%) · total 34 of 40 = 85.0%, as printed in Table 10.

  5. Gemini 3 Flash prompted only

    Named the reliable peer: 37 of 37 right (100.0%) · named another: 2 of 3 right (66.7%) · total 39 of 40 = 97.5%, as printed in Table 10.

Measured · Tables 8, 10 and 11 recognition rate (PRR) and the two conditional accuracies, as printed. Computed whole-question counts out of the 40 GPQA test questions that reproduce the printed percentages; each square is one question, ordered for display only. On MMLU-Pro the recognition rate is 82.2–100.0% for these models (Table 8), so the paper reports this split for GPQA only.

Accuracy is higher after naming the reliable peer for 5 of 5 models. The “named another peer” group is as small as 3 questions (Gemini 3 Flash), where one question moves its accuracy by 33.3 pp.

Measured: every accuracy is copied as printed from Table 6 (natural and adversarial settings), Table 10 (Flip), Table 14 (All-W; GPQA, three models), Table 8 (peer recognition rate, GPQA rows) and Table 11 (GPQA accuracy by recognition outcome). Single agent is the RL-trained single agent for Qwen 3-4B and 8B and the untrained model otherwise. Computed: differences in percentage points, whole questions out of the 90 MMLU-Pro or 40 GPQA test questions (Table 7), and the split of the 40 GPQA questions that reproduces both Table 11 percentages and lies closest to the Table 8 rate. The paper reports no repeated runs or variance.Tables 6–14 (arXiv v1) ↗

The main evaluation (Section 6) uses MMLU-Pro (90 test questions) and GPQA (40), four peers and five history rounds. Natural peers are four different LLMs, with Gemini 3 Flash as the strong one; adversarial peers are four copies of Qwen3-30B-A3B-Thinking, one justifying the correct answer and three arguing confidently for wrong ones. Qwen 3-4B and 8B are trained with RL; six larger models are only prompted. The baseline AG (history-agnostic aggregator) reads the current peer responses but no history.

Counted cell by cell in Table 6, ECL (E) beats AG in 31 of 32 adversarial comparisons (eight models, two benchmarks, two peer contexts) but in 23 of 32 natural ones, with 4 ties and 5 losses. Misleading peers also hurt AG itself: in 10 of the 32 adversarial cells, AG scores below the same model answering alone. In the natural setting, Gemini 3 Flash and Pro gain little or lose; the authors suggest that these models share knowledge with the strong peer, which is itself Gemini 3 Flash. The abstract's comparison of RL-trained Qwen 3-4B with ECL (E) against prompted Qwen 3-30B with AG holds in 5 of the 8 views.

The stress tests show where the gains come from. When the historically reliable peer answers wrongly at test time (Flip, Table 10), ECL (E) loses accuracy in all 20 reported cells and falls below AG in 18; the paper reads this as evidence that the agent follows the peer it identified. On GPQA, the agent is more often right after naming the reliable peer (Table 11), but the prompted models named another peer on only 3 to 6 of the 40 questions, so those conditional accuracies rest on very few cases. When every peer answers wrongly (All-W, Table 14), ECL (E) is at or below AG for all three tested models in both peer contexts. Adding the agent's own independent answer as an input (decoupled belief, Appendix E) lifts ECL (E) above AG in 4 of these 6 cells, but All-W accuracy stays below normal accuracy in every cell.

RL matters most for the small models on hard tasks: without it, Qwen 3-8B with ECL falls below AG on adversarial GPQA with answer-only peers (37.5% and 35.0% vs 42.5%, Table 9), while the RL-trained model reaches 50.0% vs 45.0%. For three models on adversarial MMLU-Pro with peer reasoning (the benchmark is inferred from values that match Table 6), ECL stays above AG with two to four peers and two to eight history rounds (Tables 12 and 13).

Table 4, transposed so that methods are rows. Final-answer accuracy (%) in the two-peer diagnostic of Section 4 (one peer always right, one always wrong, five history rounds); higher is better. All rows are RL-trained: 1S (RL) in one stage with the outcome reward, ECL (I) in two stages with the outcome reward on both, ECL (E) in two stages with the peer recognition reward on Stage 1. MA-Outcome peers show answers only; MA-Reasoning peers also show their reasoning. Table 4 does not name its model; its 1S (RL) values equal Qwen 2.5-3B's RL results in Table 1. Test-set sizes for Math500 and LiveCode are not reported. Source ↗

MethodMath500 · MA-OutcomeMath500 · MA-ReasoningLiveCode · MA-OutcomeLiveCode · MA-Reasoning
1S (RL)73.0%86.5%75.7%86.5%
ECL (I)75.7%94.6%77.0%91.9%
ECL (E)81.1%100.0%89.2%100.0%

Method proposed in the paper

Scope. The peers are constructed: four peers, five history rounds and exactly one peer that is reliable throughout, either four different LLMs (natural) or four copies of Qwen3-30B-A3B-Thinking, three of them arguing for wrong answers (adversarial). Test sets are small (90 MMLU-Pro and 40 GPQA questions) and no variance across runs is reported. Only Qwen 3-4B and 8B are RL-trained; the larger models are prompted. Peers whose reliability changes over time are left to future work. Read the study ↗

Further reading in the paperPaper · Sections 3–6, Appendices D–F (arXiv v1) ↗Code and data ↗
Read the abstract

Individual agents in multi-agent (MA) systems often lack robustness, tending to blindly conform to misleading peers. We show this weakness stems from both sycophancy and inadequate ability to evaluate peer reliability. To address this, we first formalize the learning problem of history-aware reference, introducing the historical interactions of peers as additional input, so that agents can estimate peer reliability and learn from trustworthy peers when uncertain. This shifts the task from evaluating peer reasoning quality to estimating peer reliability based on interaction history. We then develop Epistemic Context Learning (ECL): a reasoning framework that conditions predictions on explicitly-built peer profiles from history. We further optimize ECL by reinforcement learning using auxiliary rewards. Our experiments reveal that our ECL enables small models like Qwen 3-4B to outperform a history-agnostic baseline 8x its size (Qwen 3-30B) by accurately identifying reliable peers. ECL also boosts frontier models to near-perfect (100%) performance. We show that ECL generalizes well to various MA configurations and we find that trust is modeled well by LLMs, revealing a strong correlation in trust modeling accuracy and final answer quality.

Full paper ↗
Citation BibTeX
@misc{zhou2026epistemiccontextlearningbuilding,
    title={Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems},
    author={Ruiwen Zhou and Maojia Song and Xiaobao Wu and Sitao Cheng and Xunjian Yin and Yuxi Xie and Zhuoqun Hao and Wenyue Hua and Liangming Pan and Soujanya Poria and Min-Yen Kan},
    year={2026},
    eprint={2601.21742},
    archivePrefix={arXiv},
    primaryClass={cs.AI},
    url={https://arxiv.org/abs/2601.21742}
}