Life sciences · Preprint
arXiv · September 4, 2026
Early or partial results. Treat as a signal, not a conclusion.
LexFlip is a newly released diagnostic dataset of 373 minimal legal text perturbations designed to test whether semantic metrics preserve legal meaning while holding surface form constant. The benchmark reveals that standard embedding and BERTScore metrics allocate only 0.022–0.039 of their response range to legal force reversals that preserve 93% of tokens, whereas bidirectional NLI allocates 0.670. This dissociation diagnostic exposes a systematic failure mode in current meaning-preservation metrics for legal text.
Diagnostic benchmark validation study with controlled text perturbations. Quebec statutory French legal clauses; semantic metrics and language models (embedding-based, BERTScore, NLI); human raters as judges.. Intervention: LexFlip dataset: minimal perturbations reversing legal force while preserving surface form.. Compared with: Baseline metrics (embedding, BERTScore, NLI); human ceiling; length feature.. n = 373.
LexFlip comprises 373 minimal perturbations of Quebec statutory French reversing legal force while preserving 0.93 token overlap Seven embedding and BERTScore metrics allocate only 0.022 to 0.039 of their identical-to-unrelated range to legal force reversals Bidirectional NLI allocates 0.670 of its range, the only metric family the identical-pair check would typically disqualify
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a methodological validation study presenting a diagnostic dataset and benchmark (LexFlip) to test whether existing metrics preserve legal meaning; it demonstrates a gap in current metrics but does not evaluate a clinical, regulatory, or deployed intervention.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Does a simplified legal clause still say what the original said? The checks in current use cannot establish that it does: requiring an identical pair to score highest and an unrelated pair lowest moves lexical overlap and legal force together, so any monotone function of token overlap satisfies both. Our remedy is a dissociation, an item holding surface form fixed while legal force moves. We release LexFlip, 373 minimal perturbations of Quebec statutory French that reverse legal force while preserving 0.93 of the tokens, with a harness scoring metrics, regressors and prompted judges alike. The seven embedding and BERTScore metrics we test spend only 0.022 to 0.039 of their identical-to-unrelated range on such an edit, against 0.670 for bidirectional NLI, the one family the identical-pair check would disqualify. On FrJudge, against a measured human ceiling of r=0.597, a bare length feature outscores every semantic metric and has the lowest margin we measure.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.