Life sciences · Preprint
arXiv · September 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
SIRF is a foundation model designed to internalize content risk policies via continued pretraining, evaluated against a policy-matched baseline using controlled comparison. The preprint reports that SIRF-8B-SFT achieves 71.3% Black Recall at P95 precision (a +15.1 percentage-point gain over baseline) and shows deployment benefits; however, the work is unrefereed and lacks independent validation, peer review, or detailed statistical analysis.
Controlled same-source comparison of two model variants. Content moderation system evaluation; no human participant population described.. Intervention: SIRF-8B-SFT: foundation model with internalized platform policies via continued pretraining, policy synthesis via EntiGraph, MAGA rewriting, and account-level chain-of-thought, deployed at verdict-only interface.. Compared with: Qwen3-8B-SFT baseline with identical policy injection and output format but without policy-grounded continued pretraining..
SIRF-8B-SFT reaches 71.3% Black Recall@P95 precision, +15.1pp over baseline using ~70M continued pretraining tokens Deployed as tree-model adjudication layer recovers 20% more mis-penalized samples Transfers to freezing scenario with ~70% relative mis-penalization reduction
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a preprint describing a novel machine learning system for content moderation with controlled comparisons and deployment metrics, but it lacks peer review, clinical or health-related endpoints, and independent validation.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
For industrial content risk control, the real deployment constraint is not average accuracy but how much risk can be auto-handled under high precision and second-level latency. We present SIRF (Spec-Internalized Risk Foundation Model), which internalizes a platform's complex policies, synthesized without additional human annotation via EntiGraph, MAGA rewriting and account-level chain-of-thought (CoT), into the weights via continued pretraining (CPT), so rules are applied at high precision under an ultra-low-latency, verdict-only deployment. A controlled same-source comparison (Qwen3-8B-SFT vs. SIRF-8B-SFT, identical policy injection and verdict-only output form, differing only in policy-grounded CPT) attributes the gain to internalization: SIRF-8B-SFT reaches 71.3% Black Recall@P95, +15.1pp over the baseline, using only ~70M CPT tokens without harming general ability, and among included, logprob-available models under this interface it matches or exceeds far larger systems. SIRF is deployed as a tree-model adjudication layer (20% more mis-penalized samples recovered) and transfers to a freezing scenario at low cost (~70% relative mis-penalization reduction).
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.