
Perplexity open-sources Lily for faster local AI inference
Perplexity has open-sourced Lily, a local inference engine built for hybrid compute in Perplexity Computer and specialized for Qwen3.6-35B-A3B on Apple silicon.
Unlike general-purpose MLX-LM, Lily handles prompt prefill and token-by-token decode as separate workloads, mapping Qwen operations to Apple silicon's compute and memory architecture.
In benchmarks on an M5 Max MacBook Pro, Perplexity reported average throughput gains over MLX-LM of 1.23x for prefill and 1.35x for decode, with effectively unchanged output quality.
No comments so far, maybe you want to be first?
Gu


