arXiv preprint 2026LLM auditing · Consumer health

Challenges of Auditing: Variability in Outputs of Large Language Models for Health

Yuan Pu1, Yewon Chang1, Furong Jia1, Xunjian Yin1, Jessica Ma2,3, Ayman Ali4, Monica Agrawal1,5

1 Duke University, Department of Computer Science · 2 Duke University, Department of Medicine · 3 Durham VA Health System, Geriatrics and Extended Care · 4 Duke University, Department of Surgery · 5 Duke University, Department of Biostatistics and Bioinformatics

API-based health evaluations may not describe what chatbot users see. With the model version fixed, GPT answers to the same 50 consumer health questions differed between the API and ChatGPT, and several differences reversed between GPT-5.3 and GPT-5.4.

Result
GPT-5.3 answers averaged 377 words through ChatGPT vs 268 through the API and asked for more information in 76.0% vs 54.0% of runs; GPT-5.4 asked less often through ChatGPT (52.0% vs 76.7%).
New vs prior
Health evaluations usually query the API and generalize to the chatbot; this study measures the gap on paired questions, including ChatGPT Health, against run-to-run variation.
Limitation
One provider, 50 single-turn questions, and three runs per condition in fresh sessions without memory or connected records; p-values are unadjusted.
Scroll sideways to read the figure. Original Figure 2: one consumer health question answered by GPT-5.3 through the API (left) and ChatGPT Health (right), with the authors' annotations. The interface answer is longer and more formatted (312 vs 203 words), names more specific conditions, gives different alternative causes and different reasons to seek urgent care, and ends with follow-up questions. The authors note that the differences in specificity and clinical suggestions were also observed between repeated runs in the same access mode. Full size ↗

An API answer may not be the answer a chatbot user receives

People increasingly ask consumer chatbots such as ChatGPT for health advice, and researchers try to check whether those answers are sound. Copying questions into the chatbot interface mirrors consumer use but is hard to scale, especially as models are updated. Most evaluations therefore query the provider's API (application programming interface), which gives programmatic access to the same underlying model, and their findings are then often read as if both routes were one system.

The two routes are different access systems. Even for the same model version, the consumer product surrounds the model with product-level instructions, conversation context and saved memory, built-in tools and connected data, a product-facing model choice, and provider-managed safeguards, and not all of these are publicly documented. A researcher can configure analogous API settings but cannot fully replicate the consumer environment. ChatGPT Health, a health-specific experience inside ChatGPT, adds health-focused model training, activates automatically when ChatGPT detects a personal health question, and can draw on connected medical records and wellness data; none of these features can be reproduced through the API.

The paper asks a narrow empirical question. With the model version held constant and the available settings matched as closely as possible, do the API and the consumer interfaces answer the same health question differently, and are those differences larger than the variation between repeated runs of the same condition?

Closest prior work

Work on AI auditing and governance has argued that audit conclusions depend on the form of model access and has called for evaluation across access conditions (Casper et al., 2024; Kembery, 2024). Outside health care, studies have reported differences between ChatGPT-interface and GPT-API outputs in source selection, source diversity, and behavior in disordered or conspiratorial dialogue (Schatto-Eckrodt et al., 2025; Kirgis et al., 2026). This paper compares access modes on consumer health questions, includes ChatGPT Health, and uses repeated runs of each condition as the variability baseline.

Hold the question and the model fixed, then vary only the access mode

The authors sampled 50 patient-authored questions from HealthCareMagic-100k, a corpus of online medical questions, and excluded questions addressed to a doctor so that the sample resembles what people type into a general chatbot. GPT-5.3 Instant, then the default ChatGPT model for all account holders, was queried through the API, standard ChatGPT, and ChatGPT Health. GPT-5.4 Thinking, available to paying subscribers, was queried through the API and ChatGPT with extended thinking and web search enabled. Figure 1 shows which parts of each route a researcher controls.

Original Figure 1: columns for the GPT API, ChatGPT, and ChatGPT Health compared row by row: access point, user input, system instructions, context and memory, tools and data, model version, and control and safeguards.
Scroll sideways to read the figure. Original Figure 1, a conceptual schematic; the authors state that it does not depict OpenAI's internal architecture. All three routes share provider-managed model-level instructions and model behavior and safeguards (grey). Through the API, a researcher sets the instructions, supplied context, tools, and model snapshot (solid outline). ChatGPT adds provider-managed product-level instructions and product-exposed settings (dashed): custom instructions, the current conversation, saved memory, product tools and connected data, and a product-facing model choice. ChatGPT Health adds health-specific data integration and health-specific model behavior and safeguards. Full size ↗

Each question was submitted verbatim as the only input, three times per condition: 600 responses in the primary comparison (April 1–22, 2026) and 126 ChatGPT Health responses (August 4–5, 2026). Interface runs started new conversations in Temporary Chat, which keeps them out of saved memory. Health offered no Temporary Chat, so each Health conversation was deleted after collection, and no records were connected. API runs used an empty system prompt, gpt-5.3-chat-latest without reasoning or tools, gpt-5.4 (snapshot gpt-5.4-2026-03-05) with high reasoning effort and web search, and default settings otherwise. Health activated for 42 of the 50 questions; the other eight concerned another person's health and were excluded from the Health analyses.

Each response was coded on four dimensions. Form: length in words and Flesch–Kincaid grade after removing Markdown, the frequency of headings, list items, and bold or italic spans, and emoji use. Behavior: self-limiting statements (sentences with “I can't” or “I cannot”) and continuation moves, that is, requests for more information and offers to produce more, classified by GPT-5.4. Sources: the webpages and domains that GPT-5.4 cited. Clinical content: additional clinical concepts not raised in the question, extracted by a GPT-5.4 pipeline with clinician-reviewed prompts and a question-specific codebook, and the levels of care named, from self-care to calling 911.

The question is the unit of analysis. For each question and condition, the three runs are averaged, or turned into the share of runs showing a yes/no feature, and access modes are compared within question with paired Wilcoxon signed-rank tests, separately for each model version. Cited sources and clinical concepts are sets, so they are compared with Jaccard overlap: shared items divided by distinct items. Because repeated runs of one condition already differ, the cross-mode overlap of a question is tested against the mean of its two within-mode overlaps. The explorer below applies this procedure to illustrative citation lists.

  1. 01

    Fix the question and the model

    Fifty patient-authored questions from HealthCareMagic-100k, each submitted verbatim as the only input, in fresh sessions without memory, history, customization, or connected records.

  2. 02

    Vary only the access mode

    GPT-5.3 Instant through the API, ChatGPT, and ChatGPT Health; GPT-5.4 Thinking with web search through the API and ChatGPT. Three runs per question and condition, 726 responses in total.

  3. 03

    Test against run-to-run variation

    Paired Wilcoxon signed-rank tests within question; for cited sources and clinical concepts, cross-mode Jaccard overlap against the overlap among repeated runs in one mode.

Section 5.5 procedure · illustrative citation lists

When does low overlap point to the access mode?

Each row is one run of the same question and each column a source that a run may cite. The lists are illustrative, not study data. Pick a preset or switch single citations: the page computes the Jaccard overlap of every pair of runs, averages the three pairs inside each mode and the nine pairs across modes, and compares the cross-mode value with the mean of the two within-mode values, as the paper does for each question. Compare the second and third presets: which one has the higher cross-mode overlap, and which one shows an access-mode gap?

Sources cited by each run (source labels 1–8)
Run12345678n
API 1● cited● cited● cited● cited· not cited· not cited· not cited· not cited4
API 2● cited● cited● cited· not cited● cited· not cited· not cited· not cited4
API 3● cited· not cited· not cited● cited● cited● cited· not cited· not cited4
ChatGPT 1· not cited· not cited● cited● cited· not cited● cited● cited· not cited4
ChatGPT 2· not cited· not cited● cited· not cited· not cited● cited● cited● cited4
ChatGPT 3· not cited· not cited· not cited● cited● cited· not cited● cited● cited4
Jaccard overlap of each pair of runs
API 1API 2API 3ChatGPT 1ChatGPT 2
API 20.60
API 30.330.33
ChatGPT 10.330.140.33
ChatGPT 20.140.140.140.60
ChatGPT 30.140.140.330.330.33

Within a modeAcross modesShading grows with overlap.

Within API mean of 3 pairs
0.422
Within ChatGPT mean of 3 pairs
0.422
Within-mode baseline mean of the two
0.422
Cross-mode mean of 9 pairs
0.206
Within minus cross tested across questions
+0.216

Within-mode baseline 0.422, cross-mode 0.206: within minus cross is +0.216. Runs share fewer sources across access modes than between repeated runs in one mode. A gap of this sign, consistent across questions, is what the paper's paired test detects.

Measured counterpart (Table A3, GPT-5.4, 50 questions): cited webpages overlap 0.175 within a mode and 0.109 across modes; cited domains 0.389 and 0.298 (both p < 0.001).

Illustrative: the six citation lists, the eight source labels, and the presets. Computed: the Jaccard overlap (shared ÷ distinct sources) of every pair of runs, the within-mode means over three pairs, the cross-mode mean over nine pairs, and within minus cross, following Section 5.5. A pair of runs that cite nothing has no overlap value and is excluded, as in the paper. The paper computes these values for every question and tests the gap across questions with a paired Wilcoxon signed-rank test; one question cannot be tested.

Form and engagement differ most, and their direction depends on the model version

Measured · Tables A1–A3

Which access-mode differences hold, and which flip between model versions?

Each cell is a paired difference printed in Table A2: the first-named mode minus the second, averaged over questions, with its unadjusted Wilcoxon p-value. Choose the cut-off used to mark a difference, and select a measure to see its per-mode means from Table A1. The column 5.3 → 5.4 compares the ChatGPT-vs-API difference in GPT-5.3 with the one in GPT-5.4; the last two columns compare ChatGPT Health with the API and with standard ChatGPT.

With p < 0.05 as the cut-off, 12 of 15 tested measures differ for GPT-5.3 ChatGPT vs API, 11 of 15 for GPT-5.3 Health vs API, 5 of 11 for GPT-5.3 Health vs ChatGPT, and 13 of 17 for GPT-5.4 ChatGPT vs API. Of the 10 measures that pass in both versions' ChatGPT-vs-API contrasts, 9 change direction between GPT-5.3 and GPT-5.4; only length (words, counting the GPT-5.4 interface preamble) keeps its direction.

MeasureChatGPT vs APIChatGPT Health, GPT-5.3
GPT-5.3GPT-5.45.3 → 5.4vs APIvs ChatGPT
Response form
Length words+109p < 0.001+26p = 0.003Same+103p < 0.001−8p = 0.132
Length without preamble words—−45p < 0.001———
Readability FK grade−0.89p < 0.001+0.90p < 0.001Reverses−1.38p < 0.001−0.54p < 0.001
Section headings any, % of runs+7.3p = 0.004−43.3p < 0.001Reverses+7.9p = 0.008n.t.
Section headings n/100 words+1.56p < 0.001−0.19p < 0.001Reverses+1.82p < 0.001+0.28p = 0.024
List items any, % of runs+1.3p = 0.500−48.0p < 0.0015.4 only+1.6p = 0.500n.t.
List items n/100 words+2.29p < 0.001−1.25p < 0.001Reverses+2.69p < 0.001+0.39p = 0.060
Bold/italic spans any, % of runs+57.3p < 0.001−0.7p = 1.0005.3 only+57.1p < 0.001n.t.
Bold/italic spans n/100 words+2.93p < 0.001−2.09p < 0.001Reverses+3.03p < 0.001+0.20p = 0.171
Emoji any, % of runs+36.0p < 0.001n.t.5.3 only+28.6p < 0.001−6.3p = 0.336
Referenced sources
Webpages referenced n—−0.27p = 0.109———
Domains referenced n—−0.04p = 0.820———
Response behavior
Disclaimer any, % of runs−4.0p = 0.391−17.3p < 0.0015.4 only−10.3p = 0.047−4.8p = 0.250
Asks for information % of runs+22.0p < 0.001−24.7p < 0.001Reverses+31.0p < 0.001+11.1p < 0.001
Offers to produce more % of runs+11.3p = 0.014−7.3p = 0.046Reverses−0.8p = 0.861−10.3p = 0.043
Any continuation move % of runs+27.3p < 0.001−31.3p < 0.001Reverses+23.8p < 0.001n.t.
Clinical content
Additional clinical concepts n+0.73p < 0.001−0.73p = 0.005Reverses+0.09p = 0.900−0.56p < 0.001
Named levels of care n−0.23p = 0.213−0.11p = 0.627Neither−0.29p = 0.333+0.05p = 0.693

First-named mode higherFirst-named mode lowerNot below the cut-offn.t. not tested · — not in the design

Asks for information % of runs

Questions or prompts for additional case details; classified by GPT-5.4.

Mean (SD) over questions, Table A1. API and ChatGPT: 50 questions; Health: 42.
GPT-5.3 API54.0 (35.6)
GPT-5.3 ChatGPT76.0 (33.0)
GPT-5.3 Health91.3 (22.2)
GPT-5.4 API76.7 (33.8)
GPT-5.4 ChatGPT52.0 (37.0)

GPT-5.3 ChatGPT vs API: +22.0 pp, p < 0.001; ChatGPT asks for information in more runs. GPT-5.3 ChatGPT Health vs API: +31.0 pp, p < 0.001; ChatGPT Health asks for information in more runs. GPT-5.3 ChatGPT Health vs ChatGPT: +11.1 pp, p < 0.001; ChatGPT Health asks for information in more runs. GPT-5.4 ChatGPT vs API: −24.7 pp, p < 0.001; ChatGPT asks for information in fewer runs. The ChatGPT-vs-API difference reverses direction from GPT-5.3 to GPT-5.4.

Within- versus cross-mode overlap Table A3, mean Jaccard overlap per question

Set compared 0 to 1WithinCross Δp
Cited webpagesGPT-5.4 ChatGPT/API · n = 500.1750.109+0.066<0.001
Cited domainsGPT-5.4 ChatGPT/API · n = 500.3890.298+0.091<0.001
Additional clinical conceptsGPT-5.3 ChatGPT/API · n = 490.6530.615+0.0380.012
Additional clinical conceptsGPT-5.3 Health/API · n = 410.6620.614+0.0480.002
Additional clinical conceptsGPT-5.3 Health/ChatGPT · n = 410.6540.645+0.0100.617
Additional clinical conceptsGPT-5.4 ChatGPT/API · n = 490.6190.586+0.0330.013
Named levels of careGPT-5.3 ChatGPT/API · n = 470.6960.664+0.0320.342
Named levels of careGPT-5.3 Health/API · n = 400.7270.710+0.0170.734
Named levels of careGPT-5.3 Health/ChatGPT · n = 400.7230.681+0.0420.067
Named levels of careGPT-5.4 ChatGPT/API · n = 480.7400.714+0.0260.313

Within modeAcross modesΔ is within minus cross.

At p < 0.05, cross-mode overlap is below the within-mode baseline in 5 of 10 comparisons: cited webpages, GPT-5.4 ChatGPT/API; cited domains, GPT-5.4 ChatGPT/API; additional clinical concepts, GPT-5.3 ChatGPT/API; additional clinical concepts, GPT-5.3 Health/API; additional clinical concepts, GPT-5.4 ChatGPT/API. All printed gaps are positive (within minus cross from +0.010 to +0.091).

Measured: differences and p-values from Table A2 (two-sided paired Wilcoxon signed-rank tests, unadjusted; Δ in percentage points for percentage measures), means and standard deviations from Table A1, and overlaps from Table A3, all copied as printed. Contrasts involving ChatGPT Health use the 42 questions on which Health activated, so their Δ need not equal the difference between column means. n.t.: not tested because fewer than two questions had a nonzero paired difference; —: not part of the design. Computed on this page: the marking at the chosen cut-off, the counts, and the comparison of directions between versions. The cut-off changes only the marking; it is not a multiple-comparison correction. Bars start at zero; percentage measures use a 0–100% scale, other measures the largest mean in the row.

In GPT-5.3, both interface modes gave longer answers than the API (377 and 373 vs 268 words), with lower Flesch–Kincaid grades, denser headings, lists, and emphasis, and emoji in 36.7% and 28.6% of runs versus 0.7% (all p < 0.001 for these contrasts). They also invited further interaction more often: a continuation move appeared in 98.0% and 100.0% of runs versus 70.7% through the API, mainly as requests for more information.

In GPT-5.4, most of these differences reversed. Interface answers were harder to read (grade 11.8 vs 10.9), less densely formatted, and less likely to ask for more information (52.0% vs 76.7% of runs) or to make any continuation move (67.3% vs 98.7%). They were still longer (451 vs 425 words), but that difference depends on the preamble the interface shows before the answer: without it, interface answers were 45 words shorter on average (380 vs 425; p < 0.001, Table A2). The authors conclude that access-mode effects are not stable across model versions.

ChatGPT Health broadly resembled standard ChatGPT. On the 42 matched questions it was easier to read (−0.54 grade levels, p < 0.001) and asked for more information more often (+11.1 pp, p < 0.001), and it used self-limiting statements least (0.8% of runs). Clinical content varied less than form. Cross-mode overlap of additional clinical concepts was 0.03–0.05 below the within-mode baseline in the three comparisons with the API (p ≤ 0.013), whereas Health and standard ChatGPT overlapped as much as repeated runs did (0.645 vs 0.654, p = 0.617). Named levels of care showed no significant cross-mode drop in any pair. The number of additional concepts differed modestly: 6.7 vs 5.9 for GPT-5.3 ChatGPT vs API (p < 0.001), and 6.7 vs 7.4 for GPT-5.4 (p = 0.005).

Cited sources varied most, even within a mode. Repeated GPT-5.4 runs in the same mode shared only about 17–18% of cited webpages and 39% of domains; across modes the overlap fell to 10.9% and 29.8% (both p < 0.001), although both modes cited similar numbers of webpages (4.7 vs 4.4) and domains (3.2 each).

These results come from deliberately controlled conditions: single-turn questions without memory, history, customization, or connected records. The authors argue that these features could widen the gap in routine use, and that regulators should require developers of consumer-facing health AI to give researchers scalable, reproducible access to consumer deployment pathways; both are arguments rather than results of this study. All p-values are unadjusted, and no measure grades whether an answer is clinically correct.

Table 1 (10 of 13 rows). Mean over questions of question-level values by access mode: the three runs of each question are averaged within mode, and percentage rows give the share of runs showing the feature. API and ChatGPT columns cover 50 questions; ChatGPT Health covers the 42 on which it activated. FK grade is the Flesch–Kincaid grade level (lower is easier to read). Disclaimer counts only self-limiting “I can't” or “I cannot” sentences. GPT-5.4 ChatGPT length includes the interface's pre-answer preamble (380 words without it, Table A1). Omitted: the sample-size row and the GPT-5.4-only counts of cited webpages (API 4.7, ChatGPT 4.4) and domains (3.2 and 3.2). Source ↗

Feature (measure)GPT-5.3 APIGPT-5.3 ChatGPTGPT-5.3 HealthGPT-5.4 APIGPT-5.4 ChatGPT
Length (words)268377373425451
Readability (FK grade)8.98.07.410.911.8
Section headings (n/100 words)0.92.52.70.30.1
List items (n/100 words)4.16.46.61.70.4
Bold/italic spans (n/100 words)0.93.94.06.84.8
Any emoji (% of runs)0.736.728.60.00.0
Any disclaimer (% of runs)10.06.00.842.024.7
Asks for information (% of runs)54.076.091.376.752.0
Offers to produce more (% of runs)24.035.323.826.719.3
Any continuation move (% of runs)70.798.0100.098.767.3

Scope. OpenAI models only (GPT-5.3 Instant and GPT-5.4 Thinking), 50 single-turn questions from HealthCareMagic-100k, and three runs per condition, collected in April 2026 (ChatGPT Health in August 2026) in fresh sessions without memory, history, customization, or connected records. P-values are unadjusted; continuation moves and clinical concepts were coded with GPT-5.4-based pipelines; no measure grades clinical accuracy or safety. The call for provider APIs that reproduce consumer settings is the authors' argument, not a tested intervention. Read the study ↗

Further reading in the paperPaper · Sections 3 and 5, Figures 1–2, Table 1, Appendix Tables A1–A3 ↗Study repository (code and data announced) ↗
Read the abstract

People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations typically rely on APIs while consumers interact through chatbot interfaces, these discrepancies limit evaluation validity. Our findings underscore an urgent need for model providers to enable faithful replication of consumer experiences and settings for rigorous audits.

Full paper ↗
Citation BibTeX
@misc{pu2026challengesauditingvariabilityoutputs,
    title={Challenges of Auditing: Variability in Outputs of Large Language Models for Health},
    author={Yuan Pu and Yewon Chang and Furong Jia and Xunjian Yin and Jessica Ma and Ayman Ali and Monica Agrawal},
    year={2026},
    eprint={2609.16590},
    archivePrefix={arXiv},
    primaryClass={cs.CL},
    url={https://arxiv.org/abs/2609.16590}
}