Perplexity has launched Hybrid Compute for macOS 15, allowing Pro and Max subscribers to split AI tasks between cloud models and local hardware to protect sensitive data. Alongside this, Perplexity open sourced Lily, a specialized Rust and Metal inference engine optimized for the Qwen3.6-35B-A3B model on Apple silicon. By bypassing PyTorch and MLX, Lily achieves 1.23x faster prefill and 1.35x faster decode speeds than Apple's default MLX-LM stack on an M5 Max.
Lily open source release
- ▪The standalone demo of Lily is publicly available in the pplx-garden GitHub repository, offering greedy text generation through an OpenAI-compatible HTTP API.
- ▪Perplexity open sourced Lily, a single-process local inference engine written in Rust and Metal designed for Qwen3.6-35B-A3B on Apple silicon.
Hybrid Compute launch
- ▪Hybrid Compute is designed to keep sensitive data secure on a user's local machine and reduce inference costs by offloading tasks from expensive cloud models.
- ▪Perplexity launched Hybrid Compute, a feature for Perplexity Computer that allows users to split tasks between cloud-based frontier models and a local LLM running on their Mac.
- ▪Hybrid Compute is available to Pro and Max subscribers on Apple Silicon Macs running macOS 15 or higher, requiring a minimum of 24 GB of unified memory with 32 GB recommended.
Metal kernel architecture
- ▪Lily's architecture is specialized for Qwen3.6-35B-A3B, which mixes 10 full-attention layers using grouped-query attention with 30 Gated DeltaNet layers.
- ▪Lily executes model operations using hand-written Metal kernels and a Rust driver layer, bypassing PyTorch and MLX entirely in its execution path.
Prefill decode optimization
- ▪Lily improves decode performance by writing the selected token directly into the GPU-resident input slot of the next step, eliminating a per-token CPU round trip.
- ▪Lily optimizes prefill performance by fusing 4-bit dequantization into the grouped GEMM, which raised end-to-end prefill by 77.4% at a 512-token prompt in Perplexity's ablation.
MLX-LM benchmark comparison
- ▪Lily achieved an average of 4,156 prefill tokens/s and 170.0 decode tokens/s compared to MLX-LM's 3,388 prefill and 126.4 decode tokens/s on a 40-core M5 Max.
- ▪In a teacher-forced check across 192 positions, Lily's perplexity was 0.04% higher than MLX-LM, matching the same top-ranked token 96.35% of the time.
Story comments
Loading comments…