SEP 4, 2026 · PREPRINT
A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR
arXiv
This is an unrefereed preprint describing a machine learning method development study on a held-out test set of 438 examples, with no peer review, no clinical validation, and no comparison to established baselines or human performance.
Reported
P3 (reasoning depth/explainabilit…50.68%
P3 (reasoning depth/explainabilit…72.20%
Hybrid P1 (answer correctness) af…55.94%