Perplexity open-sources Lily for faster local AI inference

Perplexity open-sources Lily for faster local AI inference

Perplexity has open-sourced Lily, a local inference engine built for hybrid compute in Perplexity Computer and specialized for Qwen3.6-35B-A3B on Apple silicon.

Unlike general-purpose MLX-LM, Lily handles prompt prefill and token-by-token decode as separate workloads, mapping Qwen operations to Apple silicon's compute and memory architecture.

In benchmarks on an M5 Max MacBook Pro, Perplexity reported average throughput gains over MLX-LM of 1.23x for prefill and 1.35x for decode, with effectively unchanged output quality.

by Fla

Add as a preferred source on Google
  • ...

An agentic AI assistant that uses your computer to complete tasks for you, from booking travel to filling out forms.

No comments so far, maybe you want to be first?
Gu