
Liquid AI ships draft models that cut function-calling latency by 57%
DSpark checkpoints add speculative decoding to three LFM2.5 models. Liquid AI reports up to 3.18x more throughput on a GPU and 2.87x on device, with quality unchanged, and llama.cpp and SGLang support upstream on day one.
Hugging Faceverified
