Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
MedDeID is a locally deployed de-identification framework combining real and synthetic training data for clinical notes. The system achieves high character-level detection rates (96.1–99.7%) on small annotated benchmarks, with synthetic-only training showing comparable or superior performance on some tasks, but clinical deployment safety and generalizability remain unvalidated.
Technical validation study on annotated benchmarks. Clinical notes from a Dutch hospital (including an independently adjudicated 300-note benchmark), primary-care settings (100 notes), and external English synthetic datasets; no patient-level characteristics reported.. Intervention: MedDeID framework: compact transformer model trained on real or synthetic clinical notes, with integrated pseudonymisation and evaluation.. Compared with: Hospital-trained model versus synthetic-only model; synthetic-trained versus hospital-trained on primary-care notes.. Dutch hospital and primary-care setting; English instantiation tested on external synthetic benchmarks (location not specified)..
Hospital-trained compact transformer detected 98.9% of identifying text on independent 300-note Dutch benchmark while redacting only 0.24% of non-identifier text Synthetic-only model detected 96.1% of identifiers on the same benchmark Synthetic-trained model achieved 90.3% recall on 100 primary-care notes versus 87.0% for hospital-trained model
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
This framework offers potential value for institutions requiring on-premises de-identification without external data transfer, but performance on real clinical English notes remains unvalidated and the clinical safety and compliance implications of synthetic-trained models are not addressed.
A proof-of-concept framework for clinical text de-identification with promising performance metrics on limited datasets, but lacking independent validation, clinical outcome data, and peer review.
As stated by the source record.
Quoted from the source exactly as published.
This framework offers potential value for institutions requiring on-premises de-identification without external data transfer, but performance on real clinical English notes remains unvalidated and the clinical safety and compliance implications of synthetic-trained models are not addressed.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Clinical notes contain personally identifiable information (PII), restricting reuse for research and medical AI, especially when data cannot leave an institution. We developed MedDeID, an on-premises framework combining in-house annotation and synthetic-note generation with model training, inference, pseudonymisation and evaluation. On an independently annotated, adjudicated 300-note Dutch hospital benchmark, a hospital-trained compact transformer detected 98.9% of identifying text while redacting 0.24% of text outside annotated identifiers; a synthetic-only counterpart detected 96.1%. On 100 primary-care notes, the synthetic-trained model achieved higher recall than the hospital-trained model (90.3% versus 87.0%) and greater robustness to identifier-format perturbations. An English instantiation trained without real text detected 99.7% and 98.9% of annotated identifier characters on two external synthetic benchmarks. These results demonstrate transfer of the workflow to another language, but not clinical English performance. MedDeID provides a route to locally governed de-identification using real or synthetic training data.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.