AI
Perplexity Open Sourced Its Mac Inference Engine. It Reads 24 Times Faster Than It Writes.
Perplexity turned on hybrid compute in its Mac app on 1 September, then open sourced Lily, the engine underneath it, on 2 September. Lily averages 4,156 prefill tokens per second and 170 decode tokens per second on an M5 Max. That 24x gap is the real design constraint, and it explains exactly which half of the work Perplexity kept on the device.