Speculative decoding paper shows drafting tokens ahead can speed up LLM inference
Leviathan, Kalman and Matias showed a small draft model verified by the large model can cut inference latency 2-3x with identical outputs.
Google DeepMindIdeas & essays