EMNLP 2026Detecting AI-generated text

AGENT-X: Adaptive Guideline-based Expert Network for Threshold-free AI-generated teXt detection

Jiatao Li1,2, Mao Ye1, Cheng Peng1, Xunjian Yin1, Xiaojun Wan1

1 Wangxuan Institute of Computer Technology, Peking University · 2 Information Management Department, Peking University

LLM agents judge a text against routed stylistic guidelines, and a meta agent merges their calibrated verdicts, so no detection threshold is tuned. AGENT-X has the highest average accuracy for all four source models tested; Fast-DetectGPT keeps the higher average AUROC.

Result
Highest average accuracy for all four source models; on GPT-4 texts, 0.8592 vs 0.6800 for the best baseline, Likelihood.
New vs prior
Likelihood, LogRank, and Fast-DetectGPT output scores that need a tuned threshold; AGENT-X returns a label with a rationale.
Limitation
By AUROC, which needs no threshold, Fast-DetectGPT averages higher for all four source models (GPT-4: 0.9061 vs 0.9007).
Scroll sideways to read the figure. The router agent reads the text and switches guidelines on or off within the semantic, stylistic, and structural dimensions (left). Each activated guideline agent returns a decision, a reason, and a confidence calibrated by five steering prompts (right), and the meta agent aggregates these assessments into the final decision (centre). Original Figure 1 from the paper. Full size ↗

A detection score is not yet a decision

Most zero-shot detectors compute a score for each text, such as its mean token log-probability under a scoring model, and call the text AI-generated when the score passes a threshold. The threshold has to be chosen on labeled data. In this paper's baseline protocol it is tuned on a separate validation set of SQuAD texts generated by GPT-Neo-2.7B, by scanning cut-offs in steps of 1% of the score range (Appendix C), and then applied unchanged to news, stories, and biomedical answers. A score that separates human and AI texts well can still be cut at the wrong place.

AGENT-X asks whether a detector can produce the label itself, with a reason, and without a calibration set. It treats detection as authorship attribution in the sense of stylistics: the text is compared against explicit descriptions of how human and AI writing differ. The authors curated 25 such guidelines from academic literature and online communities and grouped them into semantic (6), stylistic (10), and structural (9) dimensions (Section 2, Appendix A). One semantic guideline, for example, contrasts direct, succinct human claims with AI text that is cautious, balanced, or neutral.

Because the output is a label rather than a score, the paper uses accuracy as its main metric and reports AUROC (area under the ROC curve, which averages over all thresholds) as a secondary one. It argues that AUROC favors threshold-based methods, because prompt-based LLM detectors produce discrete predictions and poorly calibrated confidences (Section 3).

Closest prior work

Zero-shot detectors score a text with a language model. Likelihood, Rank, LogRank, Entropy, and LRR use token statistics; DetectGPT and NPR compare the text with perturbed copies; DNA-GPT compares it with regenerated continuations; Fast-DetectGPT (Bao et al., 2024) replaces DetectGPT's perturbations with a cheaper sampling step. Supervised classifiers such as the RoBERTa detectors and GPTZero need labeled training data, and watermarking needs control over generation. A score becomes a label only after a threshold is chosen. AGENT-X prompts an LLM with explicit guidelines instead, and calibrates the model's verbalized confidence with SteeringConf (Zhou et al., 2025), a method that queries the same model under several confidence-steering prompts.

Route, judge under five steering prompts, then aggregate

The router agent first asks the LLM for the text's domain and a list of its stylistic features. A second prompt, run for each dimension, activates only the guidelines that clearly match both; it tells the model to be extremely conservative and typically to activate one to five per dimension (Section 4.1, Appendix G). For the news article in the paper's worked case (Appendix B, Figure 2), the router returned the domain 'News article' with features such as a formal tone and a balanced presentation of opposing viewpoints, and activated 9 of the 25 guidelines: semantic 1, 2, and 6; stylistic 1, 8, 9, and 10; structural 8 and 9.

Each activated guideline gets a base agent, an LLM prompt that returns a label (AI or human), a rationale that cites the text, and a verbalized confidence. A single verbalized confidence is unreliable, so the agent is queried five times, under steering prompts that range from 'very cautious' (make your confidence low) to 'very confident'. Following SteeringConf, answer consistency κ_ans is the share of the five prompts that give the majority label, and confidence consistency κ_conf = 1 / (1 + σ_c/μ_c) falls as the five confidences spread around their mean μ_c. The calibrated confidence is c_cal = μ_c · κ_ans · κ_conf, and the agent reports the label of the prompt whose own confidence is closest to c_cal (Section 4.2).

Section 4.2 equations · illustrative steered outputs

Calibrate one guideline agent

One base agent is queried under the five steering prompts of Appendix G. Set the label and verbalized confidence each prompt returns: the page computes answer consistency, confidence consistency, the calibrated confidence c_cal, and the label the agent reports, with the paper's equations. Try switching only the very cautious prompt to Human.

Five steered outputs Illustrative · your inputs

  1. Very cautious“Make your confidence low.”AI 0.30|ck − ccal| 0.143
  2. Cautious“Make your confidence somewhat low.”AI 0.45|ck − ccal| 0.007closest to c_cal
  3. Vanilla“Report your true confidence.”AI 0.60|ck − ccal| 0.157
  4. Confident“Make your confidence somewhat high.”AI 0.75|ck − ccal| 0.307
  5. Very confident“Make your confidence high.”AI 0.90|ck − ccal| 0.457

Calibration Computed · Section 4.2 equations

Answer consistency
κans = 5/5 = 1.000
Mean confidence
μc = 0.600
Spread
σc = 0.212
Confidence consistency
κconf = 1 / (1 + σc/μc) = 0.739
Calibrated confidence
ccal = μc · κans · κconf = 0.443
Reported label
AI, from the cautious prompt

Majority label AI (5 of 5 prompts). The agent reports AI with confidence 0.443: the cautious prompt’s confidence (0.45) is closest to c_cal. The reported label agrees with the majority.

Illustrative: the five labels and confidences are inputs chosen for this page; the paper prints no per-prompt outputs. Published: the prompt names and their confidence instructions (Appendix G, Steer Calibration Prompt). Computed: every quantity below the inputs, with the Section 4.2 equations; σ_c is the population standard deviation, and when two prompts are equally close to c_cal, the earlier one in the list is used, because the paper does not state a tie rule. Because κ_ans and κ_conf are at most 1, c_cal never exceeds μ_c, and disagreement or spread pulls it toward the cautious prompts' confidences. The paper does not report how often the reported label differs from the majority of the five prompts. Confidences move in steps of 0.05 between 0.05 and 1.Section 4.2 and Appendix G ↗

The meta agent is a prompt, not a formula. It receives every guideline report and is instructed to give more influence to confident agents and, when agents disagree, to weigh both their number and their confidence; its own confidence is calibrated with the same five steering prompts (Section 4.3, Appendix G). In the worked case, seven agents voted AI with calibrated confidences from 0.42 to 0.76 and two voted human with 0.76 each; the meta agent returned AI with confidence 0.70. No threshold is applied at any stage: the final label is the meta agent's decision.

All agents use deepseek-chat-v3-0324 (Section 5.2). The design multiplies LLM calls: each activated guideline needs five steered calls, so the worked case implies 45 base-agent calls, five meta-agent calls, and the router's prompts for one text. This count is derived from the method description; the paper reports neither cost nor latency.

  1. 01

    Route guidelines to the text

    An LLM infers the text's domain and stylistic features, then activates only the matching guidelines in each of the semantic, stylistic, and structural dimensions (25 guidelines in total).

  2. 02

    Judge each guideline five times

    A base agent labels the text under one guideline with each steering prompt, from very cautious to very confident. Answer agreement and confidence spread shrink the mean confidence into a calibrated score.

  3. 03

    Aggregate with a meta agent

    A meta agent reads all guideline reports, gives more weight to confident agents, and returns the final label, a rationale, and a confidence calibrated with the same five steering prompts.

The accuracy lead, and what AUROC adds to it

Measured · Tables 1, 2, 5 and 6

Does the accuracy lead survive a threshold-free metric?

For each source model and dataset, every detector has an accuracy from Tables 1–2, where baseline thresholds were tuned on SQuAD texts generated by GPT-Neo-2.7B, and an AUROC from Tables 5–6, which uses no threshold. Pick a column and sort by either metric; the grid shows AGENT-X's rank under both metrics in all 16 columns at once.

AGENT-X rank by accuracy / by AUROC, every column Computed · from Tables 1, 2, 5 and 6

Source modelXSumWritingPubMedAverage
ChatGPT23222515
GPT-412121312
Claude-3-Opus57351315
Claude-3-Sonnet57361315

Each cell gives AGENT-X’s rank among the detectors the table prints: 12 have an accuracy; 16 (ChatGPT, GPT-4) or 15 (Claude-3) have an AUROC. Rank 1 by accuracy is set in bold. The outlined cell is the one shown below.

GPT-4 · PubMed Measured · Table 1 and Table 5

  1. AGENT-Xthis paper; no thresholdAccuracy0.8446AUROC0.8447
  2. Likelihood (Neo-2.7)Accuracy0.7233AUROC0.8104
  3. LogRank (Neo-2.7)Accuracy0.7200AUROC0.8003
  4. Rank (Neo-2.7)Accuracy0.5767AUROC0.5965
  5. LRR (Neo-2.7)Accuracy0.5570AUROC0.6814
  6. NPR (T5-11B/Neo-2.7)Accuracy0.5537AUROC0.6328
  7. Fast-Detect (GPT-J/Neo-2.7)higher AUROC, lower accuracyAccuracy0.5267AUROC0.8503
  8. DNA-GPT (Neo-2.7)Accuracy0.5000AUROC0.7565
  9. DetectGPT (T5-11B/Neo-2.7)Accuracy0.5000AUROC0.6805
  10. RoBERTa-largeAccuracy0.4732AUROC0.6067
  11. RoBERTa-baseAccuracy0.4396AUROC0.5309
  12. Entropy (Neo-2.7)Accuracy0.3967AUROC0.3295
  13. GPTZeroAUROC table onlyAccuracynot printedAUROC0.8482
  14. Fast-Detect (Phi2-2.7B)AUROC table onlyAccuracynot printedAUROC0.6083
  15. Fast-Detect (Qwen2.5-7B)AUROC table onlyAccuracynot printedAUROC0.6391
  16. Fast-Detect (Llama3-8B)AUROC table onlyAccuracynot printedAUROC0.7556

Bars run from 0 to 1; the tick marks 0.5. Rows are sorted by the selected metric, highest first.

GPT-4 · PubMed. Accuracy: AGENT-X ranks 1 of 12 (0.8446). AUROC: AGENT-X ranks 3 of 16 (0.8447); the leader is Fast-Detect (GPT-J/Neo-2.7) with 0.8503 (accuracy 0.5267). Higher AUROC but lower accuracy than AGENT-X: Fast-Detect (GPT-J/Neo-2.7).

Measured: every accuracy and AUROC value, copied as printed from Tables 1, 2, 5, and 6 of arXiv v1. Computed: ranks (1 plus the number of detectors with a strictly higher printed value) and the list of baselines with a higher AUROC but a lower accuracy than AGENT-X. GPTZero and the Phi2, Qwen2.5, and Llama3 variants of Fast-Detect appear only in the AUROC tables. Baseline AUROCs for ChatGPT and GPT-4 are cited from Bao et al. (2024), not recomputed with the accuracy runs, and the paper does not state which AGENT-X score its AUROC ranks. Ranks across tables are therefore indicative rather than a matched comparison.Tables 1–2 (p. 6) and 5–6 (pp. 18–19) ↗

Tables 1 and 2 cover texts from four source models (ChatGPT, GPT-4, Claude-3-Opus, Claude-3-Sonnet) on XSum news, WritingPrompts stories, and PubMedQA answers. AGENT-X has the highest average accuracy for each source model and the highest accuracy in 9 of the 16 columns. Its lead is clearest on GPT-4 texts, where it is the most accurate detector on all three datasets (on PubMed, 0.8446 against 0.7233 for Likelihood), and on PubMed for both Claude-3 models. On ChatGPT texts it is second on every dataset: Fast-DetectGPT is more accurate on XSum and WritingPrompts (0.9467 and 0.9333 against 0.8967 and 0.9233), and Likelihood on PubMed (0.7633 against 0.7604). AGENT-X has the highest ChatGPT average because different baselines lead on different datasets.

The AUROC tables in Appendix F order the same detectors differently. Fast-DetectGPT (GPT-J/Neo-2.7) has the highest average AUROC for all four source models, and AGENT-X has the highest AUROC in none of the 16 columns. On GPT-4 PubMed, Fast-DetectGPT reaches an AUROC of 0.8503 against 0.8447 for AGENT-X, yet its accuracy with the transferred threshold is 0.5267. The most direct reading is that its scores separate the two classes about as well, but the threshold tuned on SQuAD texts falls in the wrong place for PubMed. The accuracy comparison therefore rewards independence from an out-of-domain threshold, which is the paper's stated goal, more than better separation of human and AI text. The baseline AUROCs for ChatGPT and GPT-4 are cited from Bao et al. (2024) rather than recomputed alongside the accuracies.

The ablations on GPT-4 texts (Table 3) remove one component at a time. Replacing the guidelines with randomly selected prompts lowers average accuracy most, to 0.6139. Random guideline activation instead of the LLM router gives 0.7197, activating every guideline 0.7588, one comprehensive agent instead of dimension-specific agents 0.7868, and uncalibrated confidences 0.7991, against 0.8592 for the full system. Two variants are more accurate than the full system on XSum (0.8691 vs 0.8624), so the full system's advantage over them comes from WritingPrompts and PubMed.

Steering also changes the confidences themselves. On GPT-4 texts it lowers the expected calibration error from 0.1034 to 0.0894 and the mean confidence from 0.83 to 0.68, against an accuracy of 0.73 (Figure 3); the paper does not say whether these are base-agent or meta-agent confidences. In a separate test without guidelines, steered verbalized confidence correlates best with correctness among the compared calibration methods (Pearson r = 0.416), although its accuracy there is 0.557 (Table 4). The interpretability claim rests on the worked case; no study asks people to rate the rationales.

Table 3 (all rows; the per-dataset AUROC columns are omitted). Ablations of AGENT-X on GPT-4-generated texts: accuracy (ACC) per dataset and averaged, and average AUROC; higher is better. The change column is computed here as the difference in average accuracy from the full system. On XSum, two variants are more accurate than the full system (0.8691 vs 0.8624). Source ↗

VariantXSum ACCWriting ACCPubMed ACCAvg. ACCAvg. AUROCChange in avg. ACC
Full AGENT-X0.86240.87050.84460.85920.9007–
w/o Multi-Agent0.86910.79670.69460.78680.8785−7.24 pp
w/o Guidelines0.77850.56330.50000.61390.7825−24.53 pp
w/o LLM Router0.80540.80670.54700.71970.7738−13.95 pp
w/o Adaptive Routing0.86910.82670.58050.75880.8665−10.04 pp
w/o Steer Calibration0.85900.81000.72820.79910.8480−6.01 pp

Method proposed in the paper

Scope. All agents run on one LLM, deepseek-chat-v3-0324, and the evaluation covers three datasets from the Fast-DetectGPT benchmark (XSum, WritingPrompts, PubMedQA) with texts from four commercial source models. Baseline thresholds come from a single out-of-domain validation set. The paper reports no sample counts, run-to-run variance, or cost, supports its interpretability claim with one worked case, and announces but does not link a code and guideline release in v1. Read the study ↗

Further reading in the paperarXiv v1 · Sections 3–7, Tables 1–3, Appendices A–C and E–G ↗arXiv HTML version ↗
Read the abstract

Existing AI-generated text detection methods heavily depend on large annotated datasets and external threshold tuning, restricting interpretability, adaptability, and zero-shot effectiveness. To address these limitations, we propose AGENT-X, a zero-shot multi-agent framework informed by classical rhetoric and systemic functional linguistics. Specifically, we organize detection guidelines into semantic, stylistic, and structural dimensions, each independently evaluated by specialized linguistic agents that provide explicit reasoning and robust calibrated confidence via semantic steering. A meta agent integrates these assessments through confidence-aware aggregation, enabling threshold-free, interpretable classification. Additionally, an adaptive Mixture-of-Agent router dynamically selects guidelines based on inferred textual characteristics. Experiments on diverse datasets demonstrate that AGENT-X substantially surpasses state-of-the-art supervised and zero-shot approaches in accuracy, interpretability, and generalization.

Full paper ↗
Citation BibTeX
@misc{li2025agentxadaptiveguidelinebasedexpert,
      title={AGENT-X: Adaptive Guideline-based Expert Network for Threshold-free AI-generated teXt detection},
      author={Jiatao Li and Mao Ye and Cheng Peng and Xunjian Yin and Xiaojun Wan},
      year={2025},
      eprint={2505.15261},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.15261},
}