EMNLP 2026Detecting AI-generated text
AGENT-X: Adaptive Guideline-based Expert Network for Threshold-free AI-generated teXt detection
LLM agents judge a text against routed stylistic guidelines, and a meta agent merges their calibrated verdicts, so no detection threshold is tuned. AGENT-X has the highest average accuracy for all four source models tested; Fast-DetectGPT keeps the higher average AUROC.
- Result
- Highest average accuracy for all four source models; on GPT-4 texts, 0.8592 vs 0.6800 for the best baseline, Likelihood.
- New vs prior
- Likelihood, LogRank, and Fast-DetectGPT output scores that need a tuned threshold; AGENT-X returns a label with a rationale.
- Limitation
- By AUROC, which needs no threshold, Fast-DetectGPT averages higher for all four source models (GPT-4: 0.9061 vs 0.9007).
A detection score is not yet a decision
Most zero-shot detectors compute a score for each text, such as its mean token log-probability under a scoring model, and call the text AI-generated when the score passes a threshold. The threshold has to be chosen on labeled data. In this paper's baseline protocol it is tuned on a separate validation set of SQuAD texts generated by GPT-Neo-2.7B, by scanning cut-offs in steps of 1% of the score range (Appendix C), and then applied unchanged to news, stories, and biomedical answers. A score that separates human and AI texts well can still be cut at the wrong place.
AGENT-X asks whether a detector can produce the label itself, with a reason, and without a calibration set. It treats detection as authorship attribution in the sense of stylistics: the text is compared against explicit descriptions of how human and AI writing differ. The authors curated 25 such guidelines from academic literature and online communities and grouped them into semantic (6), stylistic (10), and structural (9) dimensions (Section 2, Appendix A). One semantic guideline, for example, contrasts direct, succinct human claims with AI text that is cautious, balanced, or neutral.
Because the output is a label rather than a score, the paper uses accuracy as its main metric and reports AUROC (area under the ROC curve, which averages over all thresholds) as a secondary one. It argues that AUROC favors threshold-based methods, because prompt-based LLM detectors produce discrete predictions and poorly calibrated confidences (Section 3).
Closest prior work
Zero-shot detectors score a text with a language model. Likelihood, Rank, LogRank, Entropy, and LRR use token statistics; DetectGPT and NPR compare the text with perturbed copies; DNA-GPT compares it with regenerated continuations; Fast-DetectGPT (Bao et al., 2024) replaces DetectGPT's perturbations with a cheaper sampling step. Supervised classifiers such as the RoBERTa detectors and GPTZero need labeled training data, and watermarking needs control over generation. A score becomes a label only after a threshold is chosen. AGENT-X prompts an LLM with explicit guidelines instead, and calibrates the model's verbalized confidence with SteeringConf (Zhou et al., 2025), a method that queries the same model under several confidence-steering prompts.
The accuracy lead, and what AUROC adds to it
0.8446 vs 0.5267
GPT-4 · PubMed accuracy: AGENT-X vs Fast-DetectGPT
Fast-DetectGPT's threshold was tuned on SQuAD texts from GPT-Neo-2.7B. By AUROC, which needs no threshold, it scores slightly higher than AGENT-X (0.8503 vs 0.8447; Tables 1 and 5).
0.8592 vs 0.6800
GPT-4 average accuracy: AGENT-X vs the best baseline, Likelihood
AGENT-X has the highest average accuracy for all four source models; its smallest margin is on Claude-3-Sonnet (0.8000 vs 0.7856 for Likelihood, +1.44 pp; Tables 1 and 2).
−24.53 pp
GPT-4 average accuracy without the guidelines
Replacing the guidelines with randomly selected prompts lowers accuracy from 0.8592 to 0.6139, the largest drop among five ablations; uncalibrated confidences cost 6.01 pp (Table 3).
Measured · Tables 1, 2, 5 and 6
Does the accuracy lead survive a threshold-free metric?
For each source model and dataset, every detector has an accuracy from Tables 1–2, where baseline thresholds were tuned on SQuAD texts generated by GPT-Neo-2.7B, and an AUROC from Tables 5–6, which uses no threshold. Pick a column and sort by either metric; the grid shows AGENT-X's rank under both metrics in all 16 columns at once.
AGENT-X rank by accuracy / by AUROC, every column Computed · from Tables 1, 2, 5 and 6
| Source model | XSum | Writing | PubMed | Average |
|---|---|---|---|---|
| ChatGPT | 23 | 22 | 25 | 15 |
| GPT-4 | 12 | 12 | 13 | 12 |
| Claude-3-Opus | 57 | 35 | 13 | 15 |
| Claude-3-Sonnet | 57 | 36 | 13 | 15 |
Each cell gives AGENT-X’s rank among the detectors the table prints: 12 have an accuracy; 16 (ChatGPT, GPT-4) or 15 (Claude-3) have an AUROC. Rank 1 by accuracy is set in bold. The outlined cell is the one shown below.
GPT-4 · PubMed Measured · Table 1 and Table 5
Bars run from 0 to 1; the tick marks 0.5. Rows are sorted by the selected metric, highest first.
GPT-4 · PubMed. Accuracy: AGENT-X ranks 1 of 12 (0.8446). AUROC: AGENT-X ranks 3 of 16 (0.8447); the leader is Fast-Detect (GPT-J/Neo-2.7) with 0.8503 (accuracy 0.5267). Higher AUROC but lower accuracy than AGENT-X: Fast-Detect (GPT-J/Neo-2.7).
Measured: every accuracy and AUROC value, copied as printed from Tables 1, 2, 5, and 6 of arXiv v1. Computed: ranks (1 plus the number of detectors with a strictly higher printed value) and the list of baselines with a higher AUROC but a lower accuracy than AGENT-X. GPTZero and the Phi2, Qwen2.5, and Llama3 variants of Fast-Detect appear only in the AUROC tables. Baseline AUROCs for ChatGPT and GPT-4 are cited from Bao et al. (2024), not recomputed with the accuracy runs, and the paper does not state which AGENT-X score its AUROC ranks. Ranks across tables are therefore indicative rather than a matched comparison.Tables 1–2 (p. 6) and 5–6 (pp. 18–19) ↗
Tables 1 and 2 cover texts from four source models (ChatGPT, GPT-4, Claude-3-Opus, Claude-3-Sonnet) on XSum news, WritingPrompts stories, and PubMedQA answers. AGENT-X has the highest average accuracy for each source model and the highest accuracy in 9 of the 16 columns. Its lead is clearest on GPT-4 texts, where it is the most accurate detector on all three datasets (on PubMed, 0.8446 against 0.7233 for Likelihood), and on PubMed for both Claude-3 models. On ChatGPT texts it is second on every dataset: Fast-DetectGPT is more accurate on XSum and WritingPrompts (0.9467 and 0.9333 against 0.8967 and 0.9233), and Likelihood on PubMed (0.7633 against 0.7604). AGENT-X has the highest ChatGPT average because different baselines lead on different datasets.
The AUROC tables in Appendix F order the same detectors differently. Fast-DetectGPT (GPT-J/Neo-2.7) has the highest average AUROC for all four source models, and AGENT-X has the highest AUROC in none of the 16 columns. On GPT-4 PubMed, Fast-DetectGPT reaches an AUROC of 0.8503 against 0.8447 for AGENT-X, yet its accuracy with the transferred threshold is 0.5267. The most direct reading is that its scores separate the two classes about as well, but the threshold tuned on SQuAD texts falls in the wrong place for PubMed. The accuracy comparison therefore rewards independence from an out-of-domain threshold, which is the paper's stated goal, more than better separation of human and AI text. The baseline AUROCs for ChatGPT and GPT-4 are cited from Bao et al. (2024) rather than recomputed alongside the accuracies.
The ablations on GPT-4 texts (Table 3) remove one component at a time. Replacing the guidelines with randomly selected prompts lowers average accuracy most, to 0.6139. Random guideline activation instead of the LLM router gives 0.7197, activating every guideline 0.7588, one comprehensive agent instead of dimension-specific agents 0.7868, and uncalibrated confidences 0.7991, against 0.8592 for the full system. Two variants are more accurate than the full system on XSum (0.8691 vs 0.8624), so the full system's advantage over them comes from WritingPrompts and PubMed.
Steering also changes the confidences themselves. On GPT-4 texts it lowers the expected calibration error from 0.1034 to 0.0894 and the mean confidence from 0.83 to 0.68, against an accuracy of 0.73 (Figure 3); the paper does not say whether these are base-agent or meta-agent confidences. In a separate test without guidelines, steered verbalized confidence correlates best with correctness among the compared calibration methods (Pearson r = 0.416), although its accuracy there is 0.557 (Table 4). The interpretability claim rests on the worked case; no study asks people to rate the rationales.
Table 3 (all rows; the per-dataset AUROC columns are omitted). Ablations of AGENT-X on GPT-4-generated texts: accuracy (ACC) per dataset and averaged, and average AUROC; higher is better. The change column is computed here as the difference in average accuracy from the full system. On XSum, two variants are more accurate than the full system (0.8691 vs 0.8624). Source ↗
| Variant | XSum ACC | Writing ACC | PubMed ACC | Avg. ACC | Avg. AUROC | Change in avg. ACC |
|---|---|---|---|---|---|---|
| Full AGENT-X | 0.8624 | 0.8705 | 0.8446 | 0.8592 | 0.9007 | – |
| w/o Multi-Agent | 0.8691 | 0.7967 | 0.6946 | 0.7868 | 0.8785 | −7.24 pp |
| w/o Guidelines | 0.7785 | 0.5633 | 0.5000 | 0.6139 | 0.7825 | −24.53 pp |
| w/o LLM Router | 0.8054 | 0.8067 | 0.5470 | 0.7197 | 0.7738 | −13.95 pp |
| w/o Adaptive Routing | 0.8691 | 0.8267 | 0.5805 | 0.7588 | 0.8665 | −10.04 pp |
| w/o Steer Calibration | 0.8590 | 0.8100 | 0.7282 | 0.7991 | 0.8480 | −6.01 pp |
Method proposed in the paper
Scope. All agents run on one LLM, deepseek-chat-v3-0324, and the evaluation covers three datasets from the Fast-DetectGPT benchmark (XSum, WritingPrompts, PubMedQA) with texts from four commercial source models. Baseline thresholds come from a single out-of-domain validation set. The paper reports no sample counts, run-to-run variance, or cost, supports its interpretability claim with one worked case, and announces but does not link a code and guideline release in v1. Read the study ↗
Read the abstract
Existing AI-generated text detection methods heavily depend on large annotated datasets and external threshold tuning, restricting interpretability, adaptability, and zero-shot effectiveness. To address these limitations, we propose AGENT-X, a zero-shot multi-agent framework informed by classical rhetoric and systemic functional linguistics. Specifically, we organize detection guidelines into semantic, stylistic, and structural dimensions, each independently evaluated by specialized linguistic agents that provide explicit reasoning and robust calibrated confidence via semantic steering. A meta agent integrates these assessments through confidence-aware aggregation, enabling threshold-free, interpretable classification. Additionally, an adaptive Mixture-of-Agent router dynamically selects guidelines based on inferred textual characteristics. Experiments on diverse datasets demonstrate that AGENT-X substantially surpasses state-of-the-art supervised and zero-shot approaches in accuracy, interpretability, and generalization.
Full paper ↗Citation BibTeX
@misc{li2025agentxadaptiveguidelinebasedexpert,
title={AGENT-X: Adaptive Guideline-based Expert Network for Threshold-free AI-generated teXt detection},
author={Jiatao Li and Mao Ye and Cheng Peng and Xunjian Yin and Xiaojun Wan},
year={2025},
eprint={2505.15261},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2505.15261},
}