Life sciences · Preprint
arXiv · September 3, 2026
Raises a question worth testing. It does not answer one.
This preprint describes Targeted Active Search (TAS), a black-box attack that recovers forgotten prompts from machine learning models that have undergone unlearning. The attack achieves 100% entity recovery and reconstructs up to 95% of forgotten prompts using 99.7% fewer queries than naive methods, revealing a potential vulnerability in current unlearning defenses.
Controlled experimental attack study on unlearned machine learning models. Three unlearned large language models subjected to refusal-aligned unlearning (NPO, DPO, LUNAR methods) on three datasets.. Intervention: Targeted Active Search (TAS) attack using template-entity queries to recover forgotten prompts. Compared with: Naive probing baseline for query efficiency comparison.
TAS recovers forgotten entities with 100% accuracy across tested conditions TAS reconstructs up to 95% of forgotten prompts from unlearned models Attack uses up to 99.7% fewer queries than naive baseline probing
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a preprint demonstrating a novel attack methodology on machine learning systems; it raises security questions about unlearning robustness but does not evaluate clinical, regulatory, or patient-relevant outcomes.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Recent unlearning methods (e.g. NPO, DPO, LUNAR) make use of refusal alignment to suppress forgotten data. However, it has been shown that refusal responses might leave traces of unlearning, and recent attacks have been able to successfully recover some of the unlearned knowledge. In this paper, we uncover a new vulnerability. Existing attacks typically assume that the forgotten prompts are already known to the adversary and focus on recovering their answers. However, we show that the forgotten prompts themselves can be extracted by using the retained data and black-box access to the model. Our attack, Targeted Active Search (TAS), first identifies the forgotten entities by constructing canonical templates and entity pool, and selectively querying the model using the most informative template-entity pair under a limited query budget. Once the entities are identified, TAS instantiates prompt templates with those entities to probe the unlearned model and reconstruct the forgotten prompts. Experiments across three unlearning methods with three datasets and three LLMs shows that TAS recovers the forgotten entity with $100\%$ accuracy and reconstructs up to $95\%$ of forgotten prompts, all while using up to $99.7\%$ fewer queries than naive probing.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.