SEP 8, 2026 · PREPRINT
Suan: Rectifying Direct Preference Safety Alignment in Large Language Models
arXiv
This is an unrefereed arXiv preprint describing a novel algorithmic approach to LLM safety; it has not undergone peer review and reports no clinical outcomes or validated benchmarks suitable for practice guidance.
Study details
InterventionSuan: a gradient-level preference optimizat…
ComparatorExisting preference optimization methods fo…