March 25, 2026 · papers

WISP accepted at ACM SIGMETRICS

Speculative decoding wastes work whenever the draft model guesses wrong, and at the edge that waste competes with the very requests it was meant to accelerate. WISP suppresses both through dynamic drafting and SLO-aware batching. The paper appears in Proceedings of the ACM on Measurement and Analysis of Computing Systems, led by Xiangchen Li and Jiakun Fan with collaborators at Queen's University Belfast, University College Dublin and Virginia Tech.

← All news