Friday, 4 September, 2026
Perplexity Shows How Software Optimisation Can Bring Large AI Models On-Device
By TechShots Studio

Perplexity has demonstrated that a 35-billion-parameter AI model can run on consumer hardware by optimising software rather than upgrading hardware. Its custom Lily inference engine uses 4-bit quantisation and memory optimisations to run Qwen3.6-35B-A3B on Apple Silicon. On an M5 Max Mac, Lily reportedly delivered up to 1.32x faster response generation than Apple’s MLX-LM framework.
Read full story at Business Standard