AAAI 2024Temporal knowledge editing

History Matters: Temporal Knowledge Editing in Large Language Model

Xunjian Yin, Jin Jiang, Liming Yang, Xiaojun Wan

Updating a model with today's facts can erase yesterday's answers; METO edits current and historical knowledge together so that facts remain associated with their time periods.

An update should not erase the past

Correcting a false fact and updating a once-true fact are different operations. If a leader changes, both the old and new answers may be correct, depending on the date. Yet an editor can score well on the new answer, its paraphrases, and unrelated questions while silently destroying the earlier fact. The history is related to the edit, so ordinary locality tests do not necessarily protect it.

A released AToKe record makes the distinction concrete: Billy Wrigglesworth plays for Manchester United from 1937 to 1947, and an edit adds Southampton for 1947 to 1948. After the edit, a question about 1947–1948 should return Southampton, while a question about 1937–1947 should still return Manchester United. The benchmark also asks relative questions, such as which club he used to play for, which require an ordered history rather than an isolated date–answer pair.

Released record · Tables 2–3

Which questions survive a timestamped edit?

One released AToKe record gives the exact probe behind each score. Choose a question form, an edit sequence, and an editor to compare GPT-J's measured accuracy with and without METO.

Published record · AToKe-ME record, case_id 3 · (Billy Wrigglesworth, playsFor)

  1. 1937–1947Manchester United F.C.Known to GPT-J before editing
  2. 1947–1948Southampton F.C.Edit 1
  3. 1948–1953Arsenal F.C.Edit 2, multiple edits only

What the editor is asked to learn for edit 1

Existing editor: one timestamped prompt
InputTarget
From 1947 to 1948, Billy Wrigglesworth is a player ofSouthampton F.C.
Same editor with METO: both facts, and their time spans
InputTarget
Fact objective
(Billy Wrigglesworth, playsFor, ?, 1937, 1947)Manchester United F.C.
(Billy Wrigglesworth, playsFor, ?, 1947, 1948)Southampton F.C.
Time objective
(Billy Wrigglesworth, playsFor, Manchester United F.C., ?, ?)1937 to 1947
(Billy Wrigglesworth, playsFor, Southampton F.C., ?, ?)1947 to 1948

The existing-editor prompt is the record's own time_prompt. METO's targets are written in the notation of the paper's Figure 3; the paper's own time-objective example edits “Donald Trump is the President of the United States from” toward “2017 to 2021”. METO re-edits the fact GPT-J already knew so that its span closes, rather than replaying all older facts.

HES · Historical fact, explicit time · After edit 1

Probe completed by GPT-JExpected answer
From 1937 to 1947, Billy Wrigglesworth's team wasQuestion format: Which club did Billy Wrigglesworth play for from 1937 to 1947?Manchester United F.C.

Asks for the fact that the edit turned into history.

Measured accuracy on HES, GPT-J, single edit

Without METO (Table 2)With METO (Table 3)

  1. CFT0.06 → 3.38+3.32
  2. MEND1.73 → 30.14+28.41
  3. ROME2.41 → 20.25+17.84
  4. MEMIT2.22 → 30.31+28.09

Selected cell

MEMIT · HES · single edit: 2.22% without METO → 30.31% with METO (+28.09 points). Without METO, historical probes almost never succeed; with METO, 30.31% do. Over the same edits, CES moves from 99.66% to 86.4%.

Change with METO across every question form

Percentage points as printed in Table 3, single edit. The selected form and editor are outlined.

Question formCFTMENDROMEMEMIT
CESCurrent fact, explicit time−2.93+2.79−0.04−13.26
CES-PCurrent fact, explicit time, paraphrased−3.07−7.11−3.23−6.91
CRSCurrent fact, relative time−3.08−7.05−2.76−1.24
HESHistorical fact, explicit time+3.32+28.41+17.84+28.09
HRSHistorical fact, relative time+2.41+29.49+14.73+23.11

Record: AToKe-ME record, case_id 3, from the authors' released dataset documentation (linked in the paper). AToKe samples its chains from YAGO and ends each chain with one sampled counterfactual future fact, so the latest link need not be biographical. Measured: Table 2 (existing editors) and Table 3 (with METO), GPT-J 6B; changes are those printed in Table 3. The multiple-edit scores average each edit of a chain; HES* is asked once after the final edit.

Preserve the fact that is becoming history

METO first identifies which point on a known timeline matches the model's current knowledge. It then edits that current fact together with the newer facts. The important transition is the formerly current fact becoming historical: its time interval must close, and its wording must reflect the past. The method does not simply replay every older fact; it assumes facts that were already historical should remain represented.

A second editing objective reverses the usual question. Instead of only training the model to supply an entity for a dated relation, it also asks for the time span during which a specified fact was true. This makes temporal information an explicit prediction target. METO adds these examples and objectives to existing editors, so the comparison asks how temporal supervision changes an editor's behavior rather than introducing a wholly separate model architecture.

The evaluation distinguishes a single update, several successive updates, and an extension of an existing fact's validity. Baseline edit prompts also contain timestamps. Their historical failures therefore cannot be dismissed as simply forgetting to mention the date in the instruction. The benchmark tests whether dated editing actually preserves the answers implied by a timeline.

  1. 01

    Evaluate a timeline

    AToKe tests single edits, repeated updates, and extensions of a fact's validity, using questions about current and historical knowledge.

  2. 02

    Edit both time periods

    METO constructs editing targets for previous facts alongside the new fact, including timestamps in the prompts.

  3. 03

    Learn when facts hold

    A time prediction objective helps the edited model associate each fact with the period in which it is valid.

Better retention still leaves a trade-off

On GPT-J with MEMIT, METO raises single-edit historical explicit-question accuracy from 2.22% to 30.31%, while current explicit-question accuracy decreases from 99.66% to 86.4%. After multiple edits, the final historical score rises from 0.27% to 21.93%. These are substantial improvements over near-erasure, but the remaining errors matter: temporal editing is not solved, and retention is not free.

The timeline uses year-level intervals, filtered knowledge chains, and synthetic extensions, with experiments on GPT-J and four editing methods. Relative-time questions and repeated updates are harder than checking one new fact. The practical lesson is to assess current accuracy and historical preservation together, with the time resolution and update sequence stated explicitly. The interactive example pairs each question form of one released record with the measured GPT-J accuracy for that form.

Table 3 (selected columns). Accuracy (%) of GPT-J edited with METO-enhanced editors on AToKe, higher is better; each arrow gives the change from the same editor without METO (Table 2). CES / HES = current / historical questions with an explicit time, HRS = historical questions with a relative time expression; SE = single edit, ME = multiple edits (averaged over edits), and HES* is historical accuracy measured after all multiple edits are completed. The omitted current-knowledge columns show a cost: under SE, paraphrased (CES-P) and relative-time (CRS) current accuracy fall for all four editors, e.g., MEND+ CES-P 33.45 (↓7.11) and CRS 25.41 (↓7.05). Source ↗

MethodSE CESSE HESSE HRSME HESME HES*
CFT+2.8 ↓2.933.38 ↑3.322.43 ↑2.411.64 ↑1.610.73 ↑0.71
MEND+83.26 ↑2.7930.14 ↑28.4130.17 ↑29.4928.65 ↑28.2521.83 ↑21.58
ROME+99.95 ↓0.0420.25 ↑17.8416.29 ↑14.7323.22 ↑22.7815.92 ↑15.66
MEMIT+86.4 ↓13.2630.31 ↑28.0924.32 ↑23.1136.2 ↑35.7221.93 ↑21.66

Scope. The main experiments edit GPT-J on timestamped fact chains. Historical recall improves substantially but remains incomplete, and improvements need to be assessed alongside accuracy on current facts. Read the study ↗

Further reading in the paperPaper · temporal benchmark, METO, and Table 3 ↗
Research summary

The imperative task of revising or updating the knowledge stored within large language models arises from two distinct sources: intrinsic errors inherent in the model which should be corrected and outdated knowledge due to external shifts in the real world which should be updated. Existing knowledge editing methods treat all updates uniformly, but we argue that models should retain recollection of the historical knowledge while integrating the newfound knowledge. We introduce Temporal Knowledge Editing (TKE) as a distinct task and create the AToKe (Assessment of Temporal Knowledge Editing) benchmark with three dataset variants. Our experiments reveal that existing editing approaches cause catastrophic forgetting of historical facts. We propose METO (Multi-Editing with Time Objective), a framework that edits both old and new knowledge simultaneously while optimizing temporal predictions, substantially improving performance on historical knowledge retention.

Full paper ↗
Citation BibTeX
@inproceedings{yin2024history,
  title={History matters: Temporal knowledge editing in large language model},
  author={Yin, Xunjian and Jiang, Jin and Yang, Liming and Wan, Xiaojun},
  booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
  volume={38},
  number={17},
  pages={19413--19421},
  year={2024}
}