Perplexity and Nvidia Push AI Agents Onto the Device

Perplexity is pushing AI agents away from the cloud and directly onto local hardware.
Working with $NVDA, the company is launching Portable Computer, designed to run AI workloads locally with near-zero marginal inference cost.
The idea is a hybrid AI architecture.
Smaller local models handle routine tasks directly on the device.
More powerful frontier models are only called when the task actually requires them.
That could significantly reduce cloud token usage and inference costs while also keeping more sensitive user data on-device.
The economics are important.
Today, every additional cloud inference request carries compute and infrastructure costs.
Local AI changes that equation because once the hardware is purchased, additional inference can become extremely cheap.
For Nvidia, this expands the AI opportunity beyond massive data centers.
The future may not be cloud AI versus local AI.
It may be an intelligent combination of both.

Perplexity and Nvidia Push AI Agents Onto the Device