arXiv preprint 2026LLM auditing · Consumer health
Challenges of Auditing: Variability in Outputs of Large Language Models for Health
API-based health evaluations may not describe what chatbot users see. With the model version fixed, GPT answers to the same 50 consumer health questions differed between the API and ChatGPT, and several differences reversed between GPT-5.3 and GPT-5.4.
- Result
- GPT-5.3 answers averaged 377 words through ChatGPT vs 268 through the API and asked for more information in 76.0% vs 54.0% of runs; GPT-5.4 asked less often through ChatGPT (52.0% vs 76.7%).
- New vs prior
- Health evaluations usually query the API and generalize to the chatbot; this study measures the gap on paired questions, including ChatGPT Health, against run-to-run variation.
- Limitation
- One provider, 50 single-turn questions, and three runs per condition in fresh sessions without memory or connected records; p-values are unadjusted.
An API answer may not be the answer a chatbot user receives
People increasingly ask consumer chatbots such as ChatGPT for health advice, and researchers try to check whether those answers are sound. Copying questions into the chatbot interface mirrors consumer use but is hard to scale, especially as models are updated. Most evaluations therefore query the provider's API (application programming interface), which gives programmatic access to the same underlying model, and their findings are then often read as if both routes were one system.
The two routes are different access systems. Even for the same model version, the consumer product surrounds the model with product-level instructions, conversation context and saved memory, built-in tools and connected data, a product-facing model choice, and provider-managed safeguards, and not all of these are publicly documented. A researcher can configure analogous API settings but cannot fully replicate the consumer environment. ChatGPT Health, a health-specific experience inside ChatGPT, adds health-focused model training, activates automatically when ChatGPT detects a personal health question, and can draw on connected medical records and wellness data; none of these features can be reproduced through the API.
The paper asks a narrow empirical question. With the model version held constant and the available settings matched as closely as possible, do the API and the consumer interfaces answer the same health question differently, and are those differences larger than the variation between repeated runs of the same condition?
Closest prior work
Work on AI auditing and governance has argued that audit conclusions depend on the form of model access and has called for evaluation across access conditions (Casper et al., 2024; Kembery, 2024). Outside health care, studies have reported differences between ChatGPT-interface and GPT-API outputs in source selection, source diversity, and behavior in disordered or conspiratorial dialogue (Schatto-Eckrodt et al., 2025; Kirgis et al., 2026). This paper compares access modes on consumer health questions, includes ChatGPT Health, and uses repeated runs of each condition as the variability baseline.
Form and engagement differ most, and their direction depends on the model version
+22.0 / −24.7 pp
Requests for more information, ChatGPT minus API
GPT-5.3: 76.0% vs 54.0% of runs; GPT-5.4: 52.0% vs 76.7% (both p < 0.001, 50 questions, Table A2). Readability and formatting density reverse between the versions in the same way.
0.109 vs 0.175
Cited-webpage overlap across vs within modes
Mean Jaccard overlap between GPT-5.4 API and ChatGPT runs, against repeated runs in one mode (p < 0.001, 50 questions, Table A3). Clinical concepts varied less: 0.615 vs 0.653 for GPT-5.3 ChatGPT vs API.
Measured · Tables A1–A3
Which access-mode differences hold, and which flip between model versions?
Each cell is a paired difference printed in Table A2: the first-named mode minus the second, averaged over questions, with its unadjusted Wilcoxon p-value. Choose the cut-off used to mark a difference, and select a measure to see its per-mode means from Table A1. The column 5.3 → 5.4 compares the ChatGPT-vs-API difference in GPT-5.3 with the one in GPT-5.4; the last two columns compare ChatGPT Health with the API and with standard ChatGPT.
Select a measure in the table to see its per-mode means.
With p < 0.05 as the cut-off, 12 of 15 tested measures differ for GPT-5.3 ChatGPT vs API, 11 of 15 for GPT-5.3 Health vs API, 5 of 11 for GPT-5.3 Health vs ChatGPT, and 13 of 17 for GPT-5.4 ChatGPT vs API. Of the 10 measures that pass in both versions' ChatGPT-vs-API contrasts, 9 change direction between GPT-5.3 and GPT-5.4; only length (words, counting the GPT-5.4 interface preamble) keeps its direction.
| Measure | ChatGPT vs API | ChatGPT Health, GPT-5.3 | |||
|---|---|---|---|---|---|
| GPT-5.3 | GPT-5.4 | 5.3 → 5.4 | vs API | vs ChatGPT | |
| Response form | |||||
| Length words | +109p < 0.001 | +26p = 0.003 | Same | +103p < 0.001 | −8p = 0.132 |
| Length without preamble words | — | −45p < 0.001 | — | — | — |
| Readability FK grade | −0.89p < 0.001 | +0.90p < 0.001 | Reverses | −1.38p < 0.001 | −0.54p < 0.001 |
| Section headings any, % of runs | +7.3p = 0.004 | −43.3p < 0.001 | Reverses | +7.9p = 0.008 | n.t. |
| Section headings n/100 words | +1.56p < 0.001 | −0.19p < 0.001 | Reverses | +1.82p < 0.001 | +0.28p = 0.024 |
| List items any, % of runs | +1.3p = 0.500 | −48.0p < 0.001 | 5.4 only | +1.6p = 0.500 | n.t. |
| List items n/100 words | +2.29p < 0.001 | −1.25p < 0.001 | Reverses | +2.69p < 0.001 | +0.39p = 0.060 |
| Bold/italic spans any, % of runs | +57.3p < 0.001 | −0.7p = 1.000 | 5.3 only | +57.1p < 0.001 | n.t. |
| Bold/italic spans n/100 words | +2.93p < 0.001 | −2.09p < 0.001 | Reverses | +3.03p < 0.001 | +0.20p = 0.171 |
| Emoji any, % of runs | +36.0p < 0.001 | n.t. | 5.3 only | +28.6p < 0.001 | −6.3p = 0.336 |
| Referenced sources | |||||
| Webpages referenced n | — | −0.27p = 0.109 | — | — | — |
| Domains referenced n | — | −0.04p = 0.820 | — | — | — |
| Response behavior | |||||
| Disclaimer any, % of runs | −4.0p = 0.391 | −17.3p < 0.001 | 5.4 only | −10.3p = 0.047 | −4.8p = 0.250 |
| Asks for information % of runs | +22.0p < 0.001 | −24.7p < 0.001 | Reverses | +31.0p < 0.001 | +11.1p < 0.001 |
| Offers to produce more % of runs | +11.3p = 0.014 | −7.3p = 0.046 | Reverses | −0.8p = 0.861 | −10.3p = 0.043 |
| Any continuation move % of runs | +27.3p < 0.001 | −31.3p < 0.001 | Reverses | +23.8p < 0.001 | n.t. |
| Clinical content | |||||
| Additional clinical concepts n | +0.73p < 0.001 | −0.73p = 0.005 | Reverses | +0.09p = 0.900 | −0.56p < 0.001 |
| Named levels of care n | −0.23p = 0.213 | −0.11p = 0.627 | Neither | −0.29p = 0.333 | +0.05p = 0.693 |
First-named mode higherFirst-named mode lowerNot below the cut-offn.t. not tested · — not in the design
Asks for information % of runs
Questions or prompts for additional case details; classified by GPT-5.4.
| GPT-5.3 API | 54.0 (35.6) | |
|---|---|---|
| GPT-5.3 ChatGPT | 76.0 (33.0) | |
| GPT-5.3 Health | 91.3 (22.2) | |
| GPT-5.4 API | 76.7 (33.8) | |
| GPT-5.4 ChatGPT | 52.0 (37.0) |
GPT-5.3 ChatGPT vs API: +22.0 pp, p < 0.001; ChatGPT asks for information in more runs. GPT-5.3 ChatGPT Health vs API: +31.0 pp, p < 0.001; ChatGPT Health asks for information in more runs. GPT-5.3 ChatGPT Health vs ChatGPT: +11.1 pp, p < 0.001; ChatGPT Health asks for information in more runs. GPT-5.4 ChatGPT vs API: −24.7 pp, p < 0.001; ChatGPT asks for information in fewer runs. The ChatGPT-vs-API difference reverses direction from GPT-5.3 to GPT-5.4.
Within- versus cross-mode overlap Table A3, mean Jaccard overlap per question
| Set compared | 0 to 1 | Within | Cross | Δ | p |
|---|---|---|---|---|---|
| Cited webpagesGPT-5.4 ChatGPT/API · n = 50 | 0.175 | 0.109 | +0.066 | <0.001 | |
| Cited domainsGPT-5.4 ChatGPT/API · n = 50 | 0.389 | 0.298 | +0.091 | <0.001 | |
| Additional clinical conceptsGPT-5.3 ChatGPT/API · n = 49 | 0.653 | 0.615 | +0.038 | 0.012 | |
| Additional clinical conceptsGPT-5.3 Health/API · n = 41 | 0.662 | 0.614 | +0.048 | 0.002 | |
| Additional clinical conceptsGPT-5.3 Health/ChatGPT · n = 41 | 0.654 | 0.645 | +0.010 | 0.617 | |
| Additional clinical conceptsGPT-5.4 ChatGPT/API · n = 49 | 0.619 | 0.586 | +0.033 | 0.013 | |
| Named levels of careGPT-5.3 ChatGPT/API · n = 47 | 0.696 | 0.664 | +0.032 | 0.342 | |
| Named levels of careGPT-5.3 Health/API · n = 40 | 0.727 | 0.710 | +0.017 | 0.734 | |
| Named levels of careGPT-5.3 Health/ChatGPT · n = 40 | 0.723 | 0.681 | +0.042 | 0.067 | |
| Named levels of careGPT-5.4 ChatGPT/API · n = 48 | 0.740 | 0.714 | +0.026 | 0.313 |
Within modeAcross modesΔ is within minus cross.
At p < 0.05, cross-mode overlap is below the within-mode baseline in 5 of 10 comparisons: cited webpages, GPT-5.4 ChatGPT/API; cited domains, GPT-5.4 ChatGPT/API; additional clinical concepts, GPT-5.3 ChatGPT/API; additional clinical concepts, GPT-5.3 Health/API; additional clinical concepts, GPT-5.4 ChatGPT/API. All printed gaps are positive (within minus cross from +0.010 to +0.091).
Measured: differences and p-values from Table A2 (two-sided paired Wilcoxon signed-rank tests, unadjusted; Δ in percentage points for percentage measures), means and standard deviations from Table A1, and overlaps from Table A3, all copied as printed. Contrasts involving ChatGPT Health use the 42 questions on which Health activated, so their Δ need not equal the difference between column means. n.t.: not tested because fewer than two questions had a nonzero paired difference; —: not part of the design. Computed on this page: the marking at the chosen cut-off, the counts, and the comparison of directions between versions. The cut-off changes only the marking; it is not a multiple-comparison correction. Bars start at zero; percentage measures use a 0–100% scale, other measures the largest mean in the row.
In GPT-5.3, both interface modes gave longer answers than the API (377 and 373 vs 268 words), with lower Flesch–Kincaid grades, denser headings, lists, and emphasis, and emoji in 36.7% and 28.6% of runs versus 0.7% (all p < 0.001 for these contrasts). They also invited further interaction more often: a continuation move appeared in 98.0% and 100.0% of runs versus 70.7% through the API, mainly as requests for more information.
In GPT-5.4, most of these differences reversed. Interface answers were harder to read (grade 11.8 vs 10.9), less densely formatted, and less likely to ask for more information (52.0% vs 76.7% of runs) or to make any continuation move (67.3% vs 98.7%). They were still longer (451 vs 425 words), but that difference depends on the preamble the interface shows before the answer: without it, interface answers were 45 words shorter on average (380 vs 425; p < 0.001, Table A2). The authors conclude that access-mode effects are not stable across model versions.
ChatGPT Health broadly resembled standard ChatGPT. On the 42 matched questions it was easier to read (−0.54 grade levels, p < 0.001) and asked for more information more often (+11.1 pp, p < 0.001), and it used self-limiting statements least (0.8% of runs). Clinical content varied less than form. Cross-mode overlap of additional clinical concepts was 0.03–0.05 below the within-mode baseline in the three comparisons with the API (p ≤ 0.013), whereas Health and standard ChatGPT overlapped as much as repeated runs did (0.645 vs 0.654, p = 0.617). Named levels of care showed no significant cross-mode drop in any pair. The number of additional concepts differed modestly: 6.7 vs 5.9 for GPT-5.3 ChatGPT vs API (p < 0.001), and 6.7 vs 7.4 for GPT-5.4 (p = 0.005).
Cited sources varied most, even within a mode. Repeated GPT-5.4 runs in the same mode shared only about 17–18% of cited webpages and 39% of domains; across modes the overlap fell to 10.9% and 29.8% (both p < 0.001), although both modes cited similar numbers of webpages (4.7 vs 4.4) and domains (3.2 each).
These results come from deliberately controlled conditions: single-turn questions without memory, history, customization, or connected records. The authors argue that these features could widen the gap in routine use, and that regulators should require developers of consumer-facing health AI to give researchers scalable, reproducible access to consumer deployment pathways; both are arguments rather than results of this study. All p-values are unadjusted, and no measure grades whether an answer is clinically correct.
Table 1 (10 of 13 rows). Mean over questions of question-level values by access mode: the three runs of each question are averaged within mode, and percentage rows give the share of runs showing the feature. API and ChatGPT columns cover 50 questions; ChatGPT Health covers the 42 on which it activated. FK grade is the Flesch–Kincaid grade level (lower is easier to read). Disclaimer counts only self-limiting “I can't” or “I cannot” sentences. GPT-5.4 ChatGPT length includes the interface's pre-answer preamble (380 words without it, Table A1). Omitted: the sample-size row and the GPT-5.4-only counts of cited webpages (API 4.7, ChatGPT 4.4) and domains (3.2 and 3.2). Source ↗
| Feature (measure) | GPT-5.3 API | GPT-5.3 ChatGPT | GPT-5.3 Health | GPT-5.4 API | GPT-5.4 ChatGPT |
|---|---|---|---|---|---|
| Length (words) | 268 | 377 | 373 | 425 | 451 |
| Readability (FK grade) | 8.9 | 8.0 | 7.4 | 10.9 | 11.8 |
| Section headings (n/100 words) | 0.9 | 2.5 | 2.7 | 0.3 | 0.1 |
| List items (n/100 words) | 4.1 | 6.4 | 6.6 | 1.7 | 0.4 |
| Bold/italic spans (n/100 words) | 0.9 | 3.9 | 4.0 | 6.8 | 4.8 |
| Any emoji (% of runs) | 0.7 | 36.7 | 28.6 | 0.0 | 0.0 |
| Any disclaimer (% of runs) | 10.0 | 6.0 | 0.8 | 42.0 | 24.7 |
| Asks for information (% of runs) | 54.0 | 76.0 | 91.3 | 76.7 | 52.0 |
| Offers to produce more (% of runs) | 24.0 | 35.3 | 23.8 | 26.7 | 19.3 |
| Any continuation move (% of runs) | 70.7 | 98.0 | 100.0 | 98.7 | 67.3 |
Scope. OpenAI models only (GPT-5.3 Instant and GPT-5.4 Thinking), 50 single-turn questions from HealthCareMagic-100k, and three runs per condition, collected in April 2026 (ChatGPT Health in August 2026) in fresh sessions without memory, history, customization, or connected records. P-values are unadjusted; continuation moves and clinical concepts were coded with GPT-5.4-based pipelines; no measure grades clinical accuracy or safety. The call for provider APIs that reproduce consumer settings is the authors' argument, not a tested intervention. Read the study ↗
Read the abstract
People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations typically rely on APIs while consumers interact through chatbot interfaces, these discrepancies limit evaluation validity. Our findings underscore an urgent need for model providers to enable faithful replication of consumer experiences and settings for rigorous audits.
Full paper ↗Citation BibTeX
@misc{pu2026challengesauditingvariabilityoutputs,
title={Challenges of Auditing: Variability in Outputs of Large Language Models for Health},
author={Yuan Pu and Yewon Chang and Furong Jia and Xunjian Yin and Jessica Ma and Ayman Ali and Monica Agrawal},
year={2026},
eprint={2609.16590},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2609.16590}
}
