Perplexity Hybrid Compute Routes Sensitive Data to Local AI
New feature splits workloads between cloud frontier models and on-device LLMs to protect confidential information and reduce inference costs.
Perplexity has introduced Hybrid Compute, a new capability that intelligently divides AI workloads between cloud-based frontier models and local language models running directly on user hardware. The system aims to keep sensitive information on-device while leveraging powerful cloud AI for non-confidential portions of tasks.
The feature, available through Perplexity's Mac application, automatically scans files and data for sensitive content using a newly trained privacy classifier. When detected, the system routes that information to local models rather than sending it to cloud servers. Users review these decisions before processing begins and can adjust which models handle each portion of their work.
How the split processing works
According to Jon Staff, who oversees Perplexity's Mac products, the system integrates directly into the application. When users upload files or input information, the privacy classifier evaluates content and flags what should remain local. Users then confirm these suggestions and select their preferred models for both local and cloud processing.
For on-device computation, Perplexity currently offers Gemma E4B and two versions of Qwen's 35-billion parameter 3.6 model, one of which Perplexity post-trained internally. The company plans to expand local model options over time. Cloud-side processing can leverage frontier models including Opus 5 and GPT-5.6 Sol.
The installation process requires no terminal access—Perplexity's application manages local model deployment automatically. During operation, users see real-time visualization of CPU, GPU, and memory usage alongside token consumption metrics. Importantly, tokens generated by local models incur no charges.
Target use cases and tradeoffs
Perplexity highlights legal work as a primary application, where attorneys could compare client cases against public case law while keeping confidential client data entirely on their machines. The company also positions Hybrid Compute as a cost-optimization tool, allowing users to reserve expensive frontier model inference for tasks that genuinely require maximum capability.
Staff acknowledged that fully cloud-based frontier models will generally produce superior outputs for raw artifact creation. However, he emphasized that many users prioritize data privacy or cost efficiency over absolute performance. The system gives users control over this tradeoff based on their specific needs for each task.
Why it matters
Hybrid Compute represents a practical response to enterprise AI adoption barriers. Many organizations hesitate to route proprietary or regulated data through cloud AI services, even as they recognize the technology's potential. By automatically identifying and isolating sensitive information for local processing, Perplexity addresses compliance concerns while maintaining access to cutting-edge models for appropriate workloads. The approach could accelerate AI deployment in regulated industries like legal, healthcare, and finance where data residency requirements have slowed adoption.
The feature currently requires Apple Silicon Macs running macOS 15 with at least 32GB of unified memory. Access is limited to Pro, Max, and enterprise subscribers. These details were first reported by Engadget.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

