Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint introduces Temporal Deepfake Localisation in Multi-Speaker Conversations (TDLMC) as a formal problem and proposes a zero-shot, training-free pipeline that wraps existing binary deepfake detectors to pinpoint synthetic speech segments. On constructed ASVspoof 5 multi-speaker conversations, the system achieves temporal intersection-over-union 0.90 and temporal detection rate 0.95, with false-alarm rates under 6% on genuine speech and below 2% on genuine multi-speaker dialogue (AMI). The work establishes the first baseline and benchmark for this task but remains unreviewed and relies on synthetic test data.
Preprint. Constructed multi-speaker conversations with synthetic speech injection created from ASVspoof 5; real multi-speaker conversation recordings from AMI corpus.. Intervention: Zero-shot temporal deepfake localisation pipeline: five-stage training-free system combining frozen binary detector with two-threshold hysteresis finite-state-machine decoder.. Compared with: Trained localiser under identical pipeline; three frozen detectors evaluated under one decoder with held-out calibration..
On 180 constructed ASVspoof 5 multi-speaker conversations, temporal intersection-over-union 0.90, temporal detection rate 0.95, MS-DCF 0.26 False-alarm rate on genuine speech below 6%, falling under 2% on genuine real multi-speaker dialogue (AMI) Trained localiser improves temporal IoU by only approximately 0.04 over zero-shot pipeline, bounding the cost of forgoing supervision
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A novel methodological contribution demonstrating proof-of-concept for an unsolved problem, with constructed test data and no peer review, suitable for establishing a baseline rather than clinical or operational deployment.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Voice-cloning fraud increasingly relies on surgical injection: a genuine conversation in which only one or two sentences are replaced by synthetic speech. Utterance-level deepfake detectors emit a single real/fake label per clip and cannot report where the synthetic speech lies. We formalise this as Temporal Deepfake Localisation in Multi-Speaker Conversations (TDLMC), show that equal error rate and min-DCF are ill-posed once a file contains both classes, and propose temporal metrics for this regime. Our contribution is a training-free five-stage pipeline that wraps a frozen binary detector and adds segment-level output with no retraining, using a two-threshold hysteresis finitestate-machine decoder to turn noisy window scores into coherent intervals. On 180 constructed multi-speaker conversations from ASVspoof 5, the system attains temporal intersection-over-union 0.90, temporal detection rate 0.95, and MS-DCF 0.26 with a strong backbone, and its false-alarm rate on genuine speech is below 6%, falling under 2% on genuine real multi-speaker dialogue (AMI). Under an identical pipeline, a trained localiser improves temporal IoU by only about 0.04, bounding the cost of forgoing supervision. Evaluated across three frozen detectors under one decoder whose constants are selected on a held-out calibration split, and with a controlled analysis attributing the residual false-alarm rate to a backbone domain gap rather than to the decoder, this provides the first zero-shot baseline and a reusable benchmark for TDLMC.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.