arXiv preprint 2023Summarization evaluation · LLM evaluators

Human-like Summarization Evaluation with ChatGPT

Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, Xiaojun Wan

Wangxuan Institute of Computer Technology, Peking University

Given the instructions that human annotators follow, ChatGPT completed four summarization evaluation protocols. In this early-2023 study it agreed with human judgments better than common automatic metrics on SummEval but not on Newsroom, and its system-level agreement shifted sharply with the prompt.

Result
ChatGPT beats all 11 automatic metrics in every SummEval correlation column but ranks third, behind two BARTScore variants, in every Newsroom column.
New vs prior
Earlier metrics score n-gram or embedding overlap (ROUGE, BERTScore), BART likelihood (BARTScore), or sentence factuality (FactCC, DAE); this study prompts ChatGPT with four human evaluation protocols instead.
Limitation
One model snapshot (gpt-3.5-turbo-0301), one prompt per protocol, and no significance tests; a longer prompt moved system-level consistency from 0.833 to −0.149.
Scroll sideways to read the figure. The paper's four zero-shot prompt templates, one per human evaluation protocol: Likert scale scoring (Figure 1; SummEval and Newsroom), pairwise comparison (Figure 2; TLDR), Pyramid (Figure 3; REALSumm, up to 16 SCUs), and binary factuality evaluation (Figure 4; QAGS). Text in blue braces is filled in for each example. Original Figures 1–4 from the paper, arranged in a 2×2 grid. Full size ↗

Can ChatGPT take the place of a human annotator?

Summaries are usually scored automatically with ROUGE, which counts n-gram overlap with a reference summary, or with model-based metrics such as BERTScore and BARTScore. The paper notes that surface word matching misses much of what makes a summary good, that factual accuracy is hard to judge without the source document, and that even the newer metrics remain unsatisfactory in performance, usability, and interpretability. Human evaluation is the reference standard, but it is expensive, slow, and hard to reproduce.

Human annotators work from written instructions and examples. Because instruction-tuned language models can follow instructions and learn from examples in context, the authors ask whether ChatGPT can take the annotator's place: read the same instructions, return the same kind of judgment, and agree with people. They call this human-like automatic evaluation. Its appeal is flexibility: the reply can carry a score, a choice between summaries, a label, or an explanation, so one model can imitate several different human protocols.

Closest prior work

Common summarization metrics compare a summary with a reference by n-gram overlap (ROUGE) or contextual embeddings (BERTScore, MoverScore), score how likely BART is to generate it from the source or reference, or the reference from it (BARTScore), or label each sentence as factual or not (FactCC, DAE). Concurrent studies also used LLMs as evaluators: Kocmi and Federmann (2023) for translation quality, Wang et al. (2023) on three NLG meta-evaluation datasets, Ji et al. (2023) for ranking generated content, Luo et al. (2023) for factual consistency in summarization, and Liu et al. (2023) with ChatGPT, GPT-4, and chain-of-thought. This study concentrates on summarization: it recasts four human summarization protocols, including Pyramid and pairwise comparison, as prompts, and compares ChatGPT with both the metrics and expert annotators.

Four annotator protocols become four prompts

Each prompt stays as close as possible to the original human instructions (Figures 1–4). Likert scale scoring asks for 1–5 ratings on four dimensions in a single prompt: relevance, faithfulness, fluency, and coherence for SummEval, and relevance, informativeness, fluency, and coherence for Newsroom. SummEval's consistency dimension is called faithfulness because the prompt gives no definitions. Pairwise comparison (TLDR) asks which of two summaries is better and requests the answer "Summary 0" or "Summary 1" without an explanation. Pyramid (REALSumm) lists the semantic content units (SCUs) extracted from the reference summary, up to 16, and asks for Yes or No on whether each can be inferred from the summary. Binary factuality evaluation (QAGS) asks whether one sentence of a generated summary is supported by the article.

Every query goes to the ChatGPT API model gpt-3.5-turbo-0301 with temperature 0, to reduce randomness, and max_tokens 256. Simple rules extract the score or label from the reply, and a reply without one is marked invalid. Likert scores are compared with the human scores by Spearman correlation at three levels, sample, system, and dataset, which the paper names but does not define; the other three protocols are scored by accuracy against the human labels. As baselines, ROUGE-1/2/L, BERTScore, MoverScore, and six BARTScore variants are scored on the Likert and pairwise datasets, and the classifiers FactCC and DAE, which split a summary into sentences with NLTK and label each one as factual or not, on REALSumm and QAGS.

Section 4 then varies the SummEval prompt in three ways (Table 6). +def adds the definitions of the four dimensions from the original human evaluation; +def+ins adds those definitions together with step-by-step instructions (Figure 5, which opens with "Imagine you are a human annotator now."); +sys_prompt sets the API system prompt to "You are a human annotator that rates the quality of summaries." The same table gives a human reference point: the correlation of each of the three expert annotators with the average of all three.

  1. 01

    Reuse human protocols as prompts

    Turn four human evaluation methods (Likert scale scoring, pairwise comparison, Pyramid, and binary factuality evaluation) into zero-shot prompts that stay as close as possible to the original annotator instructions (Figures 1–4).

  2. 02

    Query ChatGPT and parse its replies

    Send each example to gpt-3.5-turbo-0301 with temperature 0 and max_tokens 256, extract the score or label with simple rules, and mark replies without one as invalid.

  3. 03

    Compare with humans and metrics

    Correlate Likert scores with human scores (Spearman ρ at sample, system, and dataset level) and measure accuracy for the other protocols; score ROUGE, BERTScore, MoverScore, BARTScore, FactCC, and DAE against the same human judgments.

Ahead on SummEval, behind on Newsroom, and sensitive to the prompt

Measured · Tables 1–5

Where does ChatGPT lead the automatic metrics?

Tables 1–5 compare ChatGPT with automatic metrics in 28 columns: Spearman ρ with human Likert scores for four dimensions at three levels on SummEval and Newsroom, and accuracy against human labels on TLDR, REALSumm, QAGS_CNN, and QAGS_XSUM. Pick a dataset, dimension, and correlation level: the bars rank every row the paper prints for that column, and the grid shows ChatGPT minus the strongest other row in all 28 columns at once.

ChatGPT minus the strongest other row, all 28 columns

SummEval · Table 1 · Spearman ρ
SampleSystemDataset
ConsistencyConsistency, sample level+0.068Consistency, system level+0.033Consistency, dataset level+0.091
RelevanceRelevance, sample level+0.077Relevance, system level+0.136Relevance, dataset level+0.051
FluencyFluency, sample level+0.070Fluency, system level+0.143Fluency, dataset level+0.125
CoherenceCoherence, sample level+0.113Coherence, system level+0.132Coherence, dataset level+0.149
Newsroom · Table 2 · Spearman ρ
SampleSystemDataset
CoherenceCoherence, sample level−0.195Coherence, system level−0.143Coherence, dataset level−0.180
FluencyFluency, sample level−0.190Fluency, system level−0.357Fluency, dataset level−0.144
InformativenessInformativeness, sample level−0.125Informativeness, system level−0.214Informativeness, dataset level−0.137
RelevanceRelevance, sample level−0.080Relevance, system level−0.179Relevance, dataset level−0.067
  • TLDRTable 3 · Pairwise comparisonAccuracy+0.27 pp
  • REALSummTable 4 · PyramidAccuracy+1.32 pp
  • QAGS_CNNTable 5 · Binary factuality evaluationAccuracy+0.29 pp
  • QAGS_XSUMTable 5 · Binary factuality evaluationAccuracy+12.13 pp

Computed from the printed values: correlation margins in Spearman ρ, accuracy margins in percentage points (pp). Warm cells are columns where a metric is ahead of ChatGPT: 12 of 28.

SummEval · Consistency · sample level · Spearman ρ

  1. ChatGPTgpt-3.5-turbo-0301, zero-shot prompt0.435
  2. BARTScore_cnn_s_hBART generation likelihood0.367
  3. BARTScore_s_hBART generation likelihood0.299
  4. ROUGE-2n-gram overlap0.179
  5. BARTScore_cnn_h_rBART generation likelihood0.171
  6. ROUGE-1n-gram overlap0.153
  7. MoverScoreembedding similarity0.151
  8. ROUGE-Ln-gram overlap0.111
  9. BERTScoreembedding similarity0.105
  10. BARTScore_h_rBART generation likelihood0.097
  11. BARTScore_cnn_r_hBART generation likelihood0.001
  12. BARTScore_r_hBART generation likelihood−0.075

ChatGPTStrongest other row

Measured · Table 1 Each summary is rated 1–5 on four dimensions in one prompt (Figure 1); the score is Spearman ρ with the human ratings.

The prompt called this dimension faithfulness, because it gave no definitions (footnote 5); Table 1 keeps SummEval’s term, consistency.

SummEval · Consistency · sample level (Table 1). ChatGPT 0.435 ranks 1 of 12. The strongest other row is BARTScore_cnn_s_h at 0.367, so ChatGPT leads by 0.068. Across all 28 columns, ChatGPT is ahead of the strongest other row in 16 and behind it in 12, all of them on Newsroom.

Measured: every value is copied as printed from Tables 1–5 of arXiv v1 (gpt-3.5-turbo-0301, temperature 0). Computed: the ranking, ChatGPT's rank, and its margin over the strongest other row; correlation margins are differences in ρ and accuracy margins are percentage points. Derived: the TLDR chance level of 0.5. The paper reports no significance tests or confidence intervals and does not define the three correlation levels. Row names are printed as in the paper; Section 2.1 describes BARTScore's generation directions (from the source or the reference to the summary, or from the summary to the reference) but does not explain the suffixes.Tables 1–5 · Section 3.5 ↗

With the base prompt, ChatGPT beats all eleven automatic metrics in every SummEval column (Table 1); its narrowest lead is at system-level consistency, 0.833 against 0.800 for BARTScore_s_h, and BARTScore_cnn_s_h is the strongest metric in the other eleven columns. On Newsroom (Table 2) ChatGPT ranks third in every column, behind BARTScore_s_h and BARTScore_cnn_s_h. The accuracy protocols show smaller margins: +0.27 pp over the best metric on TLDR, +1.32 pp over DAE on REALSumm, +0.29 pp on QAGS_CNN, and +12.13 pp on QAGS_XSUM, the one large lead. The paper reports no significance tests or confidence intervals for these differences.

Prompt changes move system-level correlation much more than the other two levels (Table 6). With definitions and step instructions, system-level ρ falls from 0.833 to −0.149 for consistency and from 0.901 to −0.079 for relevance, while sample- and dataset-level values fall by at most 0.123; the system prompt alone lowers system-level consistency to 0.007. Definitions alone raise sample- and dataset-level ρ for consistency, relevance, and coherence but lower all three fluency values, which the paper summarizes as a modest improvement in a few cases. An inference that the paper does not test: system-level correlation is usually computed over per-system averages and therefore rests on far fewer points, and on Newsroom four to six metrics print identical system-level values in each dimension; both point to this level as the least stable.

Expert annotators stay ahead. Each annotator's correlation with the three-annotator average exceeds ChatGPT's in 35 of the 36 annotator-column pairs; the exception is system-level fluency, where ChatGPT's 0.889 exceeds Annotator_1's 0.843. Because each annotator is part of the average they are compared with, the human figures are likely somewhat inflated. The paper's two practical arguments rest on estimates rather than experiments: fixing the decoding temperature should make the judgments reproducible, and the cost is lower. For SummEval it computes 0.002 × 1600 = 3.2 USD at about 1,000 tokens per summary, against 12 × 5 = 60 USD for one annotator assumed to need 5 hours, and puts human evaluation at 10 to 20 times the cost.

ChatGPT often adds an explanation even when the prompt does not ask for one. In the Table 7 example both the base prompt and +def give low faithfulness or consistency scores and point to a factual error in the summary, but the paper notes that the correctness of such explanations was not tested. The explanations also show how ChatGPT reads the dimensions: without definitions it explains fluency and coherence together, and its fluency and coherence scores correlate at 0.960 at dataset level; with definitions the explanations separate and the correlation drops to 0.843. Invalid replies, such as refusing to evaluate, writing a new summary, or continuing the given one, stay at about 1% or less: 0 with the base prompt and 0.0106 with +def+ins (Table 9).

Table 6 (sample- and system-level columns for consistency and fluency; the dataset-level columns and the relevance and coherence dimensions are omitted). Spearman ρ on SummEval, higher is better; gpt-3.5-turbo-0301 at temperature 0. +def adds the dimension definitions, +ins the step instructions (Figure 5), and +sys_prompt sets the system prompt “You are a human annotator that rates the quality of summaries.” Each annotator row correlates one expert’s scores with the average of all three experts, which includes that expert. The omitted columns follow the same pattern: with +def+ins, system-level ρ falls to −0.079 for relevance and 0.338 for coherence, while every dataset-level value of the three variants stays within 0.123 of the base prompt. Source ↗

EvaluatorConsistency sampleConsistency systemFluency sampleFluency system
ChatGPT0.4350.8330.4190.889
ChatGPT+def0.4710.7860.3470.606
ChatGPT+def+ins0.338−0.1490.3490.016
ChatGPT+sys_prompt0.4140.0070.3900.149
Annotator_00.8430.9900.7400.960
Annotator_10.8130.9650.8470.843
Annotator_20.7120.9730.6130.923

Scope. All results come from one ChatGPT snapshot, gpt-3.5-turbo-0301, queried in early 2023 (arXiv v1, 5 April 2023, the only version) with temperature 0 and one prompt per protocol; later models, snapshots, and prompts may behave differently. The study covers five summarization datasets, reports no significance tests or confidence intervals, and does not define its sample-, system-, and dataset-level correlations. Its reproducibility argument is not tested by repeated runs, and its cost comparison (about 3.2 USD in API fees against 60 USD for SummEval) assumes 5 annotator hours at 12 USD per hour rather than measuring human effort. Read the study ↗

Further reading in the paperPaper · Section 3, Experiments (arXiv v1 HTML) ↗Section 4 · prompts, human comparison, explanations ↗PDF · arXiv v1 ↗
Read the abstract

Evaluating text summarization is a challenging problem, and existing evaluation metrics are far from satisfactory. In this study, we explored ChatGPT's ability to perform human-like summarization evaluation using four human evaluation methods on five datasets. We found that ChatGPT was able to complete annotations relatively smoothly using Likert scale scoring, pairwise comparison, Pyramid, and binary factuality evaluation. Additionally, it outperformed commonly used automatic evaluation metrics on some datasets. Furthermore, we discussed the impact of different prompts, compared its performance with that of human evaluation, and analyzed the generated explanations and invalid responses.

Full paper ↗
Citation BibTeX
@misc{gao2023humanlikesummarizationevaluationchatgpt,
    title={Human-like Summarization Evaluation with ChatGPT},
    author={Mingqi Gao and Jie Ruan and Renliang Sun and Xunjian Yin and Shiping Yang and Xiaojun Wan},
    year={2023},
    eprint={2304.02554},
    archivePrefix={arXiv},
    primaryClass={cs.CL},
    url={https://arxiv.org/abs/2304.02554}
}