AUG 14, 2026 · PREPRINT
More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It
arXiv
This is an unreviewed preprint describing a methodological problem in language model sampling and proposing a fix, demonstrated empirically across benchmarks but without peer review or clinical/real-world validation.
Reported
Accuracy drop from Power Samplingup to 18.5 percentage points