The Verge · September 2, 2026 · 1Cifer
Perplexity splits AI work between device and cloud: the hybrid scheme saves money and protects data
Perplexity demonstrated a hybrid compute scheme: part of the AI work runs locally on the user's device, and the system calls cloud models only for complex tasks. The benefits come in threes — lower latency, a smaller cloud bill, and less data leaving the device.
It answers a growing tension in the industry: cloud models are powerful, but every request costs money and means sending data out. Local models are weaker, yet free to run and private. Hybrid takes the best of both worlds, and nearly every major player is now moving down this path.
For companies in Kazakhstan, where data sensitivity is high — banking secrecy, customers' personal data, commercial terms — the hybrid approach is worth keeping in mind when choosing AI solutions. The right question for a vendor: which data is processed where, and what exactly reaches external models. At 1Cifer we answer it with a protected-perimeter architecture: the agent works with company data inside the company's own space rather than pouring it into a common pot.


