arXiv preprint 2023Summarization evaluation · LLM evaluators
Human-like Summarization Evaluation with ChatGPT
Given the instructions that human annotators follow, ChatGPT completed four summarization evaluation protocols. In this early-2023 study it agreed with human judgments better than common automatic metrics on SummEval but not on Newsroom, and its system-level agreement shifted sharply with the prompt.
- Result
- ChatGPT beats all 11 automatic metrics in every SummEval correlation column but ranks third, behind two BARTScore variants, in every Newsroom column.
- New vs prior
- Earlier metrics score n-gram or embedding overlap (ROUGE, BERTScore), BART likelihood (BARTScore), or sentence factuality (FactCC, DAE); this study prompts ChatGPT with four human evaluation protocols instead.
- Limitation
- One model snapshot (gpt-3.5-turbo-0301), one prompt per protocol, and no significance tests; a longer prompt moved system-level consistency from 0.833 to −0.149.
Can ChatGPT take the place of a human annotator?
Summaries are usually scored automatically with ROUGE, which counts n-gram overlap with a reference summary, or with model-based metrics such as BERTScore and BARTScore. The paper notes that surface word matching misses much of what makes a summary good, that factual accuracy is hard to judge without the source document, and that even the newer metrics remain unsatisfactory in performance, usability, and interpretability. Human evaluation is the reference standard, but it is expensive, slow, and hard to reproduce.
Human annotators work from written instructions and examples. Because instruction-tuned language models can follow instructions and learn from examples in context, the authors ask whether ChatGPT can take the annotator's place: read the same instructions, return the same kind of judgment, and agree with people. They call this human-like automatic evaluation. Its appeal is flexibility: the reply can carry a score, a choice between summaries, a label, or an explanation, so one model can imitate several different human protocols.
Closest prior work
Common summarization metrics compare a summary with a reference by n-gram overlap (ROUGE) or contextual embeddings (BERTScore, MoverScore), score how likely BART is to generate it from the source or reference, or the reference from it (BARTScore), or label each sentence as factual or not (FactCC, DAE). Concurrent studies also used LLMs as evaluators: Kocmi and Federmann (2023) for translation quality, Wang et al. (2023) on three NLG meta-evaluation datasets, Ji et al. (2023) for ranking generated content, Luo et al. (2023) for factual consistency in summarization, and Liu et al. (2023) with ChatGPT, GPT-4, and chain-of-thought. This study concentrates on summarization: it recasts four human summarization protocols, including Pyramid and pairwise comparison, as prompts, and compares ChatGPT with both the metrics and expert annotators.
Ahead on SummEval, behind on Newsroom, and sensitive to the prompt
12 of 12
SummEval columns where ChatGPT beats all 11 metrics
Spearman ρ of gpt-3.5-turbo-0301 with human Likert scores for four dimensions at three levels; its narrowest lead is system-level consistency, 0.833 against 0.800 for BARTScore_s_h (Table 1). No significance tests are reported.
3rd of 12
ChatGPT's rank in every Newsroom column
BARTScore_s_h and BARTScore_cnn_s_h are higher in all 12 columns; at sample-level coherence they reach 0.679 and 0.653 against ChatGPT's 0.484 (Table 2).
+12.13 pp
Accuracy lead on QAGS_XSUM factuality
ChatGPT 0.7573 against DAE 0.6360 on binary sentence factuality. Its other accuracy leads are small: +0.29 pp on QAGS_CNN, +1.32 pp on REALSumm SCUs, and +0.27 pp over BARTScore_h_r on TLDR pairwise comparison (Tables 3–5).
−0.149
SummEval system-level consistency with a longer prompt, from 0.833
Adding the dimension definitions and step instructions (Figure 5) reverses ChatGPT's system-level ρ, while its sample-level ρ falls only from 0.435 to 0.338; invalid replies rise from 0 to 1.06% (Tables 6 and 9).
Measured · Tables 1–5
Where does ChatGPT lead the automatic metrics?
Tables 1–5 compare ChatGPT with automatic metrics in 28 columns: Spearman ρ with human Likert scores for four dimensions at three levels on SummEval and Newsroom, and accuracy against human labels on TLDR, REALSumm, QAGS_CNN, and QAGS_XSUM. Pick a dataset, dimension, and correlation level: the bars rank every row the paper prints for that column, and the grid shows ChatGPT minus the strongest other row in all 28 columns at once.
SummEval and Newsroom print Spearman ρ for four dimensions at three levels; the other datasets print one accuracy each. Select a cell of the grid to open that column.
ChatGPT minus the strongest other row, all 28 columns
| Sample | System | Dataset | |
|---|---|---|---|
| Consistency | Consistency, sample level+0.068 | Consistency, system level+0.033 | Consistency, dataset level+0.091 |
| Relevance | Relevance, sample level+0.077 | Relevance, system level+0.136 | Relevance, dataset level+0.051 |
| Fluency | Fluency, sample level+0.070 | Fluency, system level+0.143 | Fluency, dataset level+0.125 |
| Coherence | Coherence, sample level+0.113 | Coherence, system level+0.132 | Coherence, dataset level+0.149 |
| Sample | System | Dataset | |
|---|---|---|---|
| Coherence | Coherence, sample level−0.195 | Coherence, system level−0.143 | Coherence, dataset level−0.180 |
| Fluency | Fluency, sample level−0.190 | Fluency, system level−0.357 | Fluency, dataset level−0.144 |
| Informativeness | Informativeness, sample level−0.125 | Informativeness, system level−0.214 | Informativeness, dataset level−0.137 |
| Relevance | Relevance, sample level−0.080 | Relevance, system level−0.179 | Relevance, dataset level−0.067 |
- TLDRTable 3 · Pairwise comparisonAccuracy+0.27 pp
- REALSummTable 4 · PyramidAccuracy+1.32 pp
- QAGS_CNNTable 5 · Binary factuality evaluationAccuracy+0.29 pp
- QAGS_XSUMTable 5 · Binary factuality evaluationAccuracy+12.13 pp
Computed from the printed values: correlation margins in Spearman ρ, accuracy margins in percentage points (pp). Warm cells are columns where a metric is ahead of ChatGPT: 12 of 28.
SummEval · Consistency · sample level · Spearman ρ
ChatGPTStrongest other row
Measured · Table 1 Each summary is rated 1–5 on four dimensions in one prompt (Figure 1); the score is Spearman ρ with the human ratings.
The prompt called this dimension faithfulness, because it gave no definitions (footnote 5); Table 1 keeps SummEval’s term, consistency.
SummEval · Consistency · sample level (Table 1). ChatGPT 0.435 ranks 1 of 12. The strongest other row is BARTScore_cnn_s_h at 0.367, so ChatGPT leads by 0.068. Across all 28 columns, ChatGPT is ahead of the strongest other row in 16 and behind it in 12, all of them on Newsroom.
Measured: every value is copied as printed from Tables 1–5 of arXiv v1 (gpt-3.5-turbo-0301, temperature 0). Computed: the ranking, ChatGPT's rank, and its margin over the strongest other row; correlation margins are differences in ρ and accuracy margins are percentage points. Derived: the TLDR chance level of 0.5. The paper reports no significance tests or confidence intervals and does not define the three correlation levels. Row names are printed as in the paper; Section 2.1 describes BARTScore's generation directions (from the source or the reference to the summary, or from the summary to the reference) but does not explain the suffixes.Tables 1–5 · Section 3.5 ↗
With the base prompt, ChatGPT beats all eleven automatic metrics in every SummEval column (Table 1); its narrowest lead is at system-level consistency, 0.833 against 0.800 for BARTScore_s_h, and BARTScore_cnn_s_h is the strongest metric in the other eleven columns. On Newsroom (Table 2) ChatGPT ranks third in every column, behind BARTScore_s_h and BARTScore_cnn_s_h. The accuracy protocols show smaller margins: +0.27 pp over the best metric on TLDR, +1.32 pp over DAE on REALSumm, +0.29 pp on QAGS_CNN, and +12.13 pp on QAGS_XSUM, the one large lead. The paper reports no significance tests or confidence intervals for these differences.
Prompt changes move system-level correlation much more than the other two levels (Table 6). With definitions and step instructions, system-level ρ falls from 0.833 to −0.149 for consistency and from 0.901 to −0.079 for relevance, while sample- and dataset-level values fall by at most 0.123; the system prompt alone lowers system-level consistency to 0.007. Definitions alone raise sample- and dataset-level ρ for consistency, relevance, and coherence but lower all three fluency values, which the paper summarizes as a modest improvement in a few cases. An inference that the paper does not test: system-level correlation is usually computed over per-system averages and therefore rests on far fewer points, and on Newsroom four to six metrics print identical system-level values in each dimension; both point to this level as the least stable.
Expert annotators stay ahead. Each annotator's correlation with the three-annotator average exceeds ChatGPT's in 35 of the 36 annotator-column pairs; the exception is system-level fluency, where ChatGPT's 0.889 exceeds Annotator_1's 0.843. Because each annotator is part of the average they are compared with, the human figures are likely somewhat inflated. The paper's two practical arguments rest on estimates rather than experiments: fixing the decoding temperature should make the judgments reproducible, and the cost is lower. For SummEval it computes 0.002 × 1600 = 3.2 USD at about 1,000 tokens per summary, against 12 × 5 = 60 USD for one annotator assumed to need 5 hours, and puts human evaluation at 10 to 20 times the cost.
ChatGPT often adds an explanation even when the prompt does not ask for one. In the Table 7 example both the base prompt and +def give low faithfulness or consistency scores and point to a factual error in the summary, but the paper notes that the correctness of such explanations was not tested. The explanations also show how ChatGPT reads the dimensions: without definitions it explains fluency and coherence together, and its fluency and coherence scores correlate at 0.960 at dataset level; with definitions the explanations separate and the correlation drops to 0.843. Invalid replies, such as refusing to evaluate, writing a new summary, or continuing the given one, stay at about 1% or less: 0 with the base prompt and 0.0106 with +def+ins (Table 9).
Table 6 (sample- and system-level columns for consistency and fluency; the dataset-level columns and the relevance and coherence dimensions are omitted). Spearman ρ on SummEval, higher is better; gpt-3.5-turbo-0301 at temperature 0. +def adds the dimension definitions, +ins the step instructions (Figure 5), and +sys_prompt sets the system prompt “You are a human annotator that rates the quality of summaries.” Each annotator row correlates one expert’s scores with the average of all three experts, which includes that expert. The omitted columns follow the same pattern: with +def+ins, system-level ρ falls to −0.079 for relevance and 0.338 for coherence, while every dataset-level value of the three variants stays within 0.123 of the base prompt. Source ↗
| Evaluator | Consistency sample | Consistency system | Fluency sample | Fluency system |
|---|---|---|---|---|
| ChatGPT | 0.435 | 0.833 | 0.419 | 0.889 |
| ChatGPT+def | 0.471 | 0.786 | 0.347 | 0.606 |
| ChatGPT+def+ins | 0.338 | −0.149 | 0.349 | 0.016 |
| ChatGPT+sys_prompt | 0.414 | 0.007 | 0.390 | 0.149 |
| Annotator_0 | 0.843 | 0.990 | 0.740 | 0.960 |
| Annotator_1 | 0.813 | 0.965 | 0.847 | 0.843 |
| Annotator_2 | 0.712 | 0.973 | 0.613 | 0.923 |
Scope. All results come from one ChatGPT snapshot, gpt-3.5-turbo-0301, queried in early 2023 (arXiv v1, 5 April 2023, the only version) with temperature 0 and one prompt per protocol; later models, snapshots, and prompts may behave differently. The study covers five summarization datasets, reports no significance tests or confidence intervals, and does not define its sample-, system-, and dataset-level correlations. Its reproducibility argument is not tested by repeated runs, and its cost comparison (about 3.2 USD in API fees against 60 USD for SummEval) assumes 5 annotator hours at 12 USD per hour rather than measuring human effort. Read the study ↗
Read the abstract
Evaluating text summarization is a challenging problem, and existing evaluation metrics are far from satisfactory. In this study, we explored ChatGPT's ability to perform human-like summarization evaluation using four human evaluation methods on five datasets. We found that ChatGPT was able to complete annotations relatively smoothly using Likert scale scoring, pairwise comparison, Pyramid, and binary factuality evaluation. Additionally, it outperformed commonly used automatic evaluation metrics on some datasets. Furthermore, we discussed the impact of different prompts, compared its performance with that of human evaluation, and analyzed the generated explanations and invalid responses.
Full paper ↗Citation BibTeX
@misc{gao2023humanlikesummarizationevaluationchatgpt,
title={Human-like Summarization Evaluation with ChatGPT},
author={Mingqi Gao and Jie Ruan and Renliang Sun and Xunjian Yin and Shiping Yang and Xiaojun Wan},
year={2023},
eprint={2304.02554},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2304.02554}
}