SEP 9, 2026 · PREPRINT
PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling
arXiv
This is an unrefereed technical preprint describing a software optimization method for on-device LLM inference; it reports comparative performance metrics but has not undergone peer review.
Reported
Speedupup to 23.1%
Energy consumption reductionup to 52.4%