Dell Deskside AI cuts token costs 87% by moving agents on-premise
As agentic workflows consume up to a million tokens per task, Dell's workstation-based approach shifts compute from metered cloud to owned hardware.

The token cost explosion
Agentic AI systems don't just answer questions—they plan, execute, revise, and loop until tasks complete. Each step consumes tokens, the small text units that AI models process at roughly three-quarters of a word each. While a simple chatbot exchange burns a few hundred tokens, a basic agent uses up to 15,000 per task. Complex multi-agent systems can consume 200,000 to over a million tokens for a single workflow.
Stanford researchers found that agentic coding uses approximately 1,000 times more tokens than standard code chat. Goldman Sachs projects total enterprise token consumption will multiply 24-fold by 2030, reaching 120 quadrillion tokens monthly. Token prices have dropped roughly 80% between mid-2023 and early 2026, but enterprise AI spending climbed 320% in the same period as companies deployed more agents across more workflows.
The result: lower unit costs met dramatically higher volume, and total bills climbed.
Why it matters
Token consumption transforms AI spending from predictable software licenses to variable compute costs that scale with usage. Each agentic workflow resends accumulated context with every step, multiplying costs in ways that make infrastructure placement a financial decision, not just a technical one. Organizations that treat token strategy as purely an IT concern will find themselves managing runaway budgets instead of controlled deployments.
Running agents where you own the hardware
Bharat Patel, a solution architect at Dell Technologies Customer Solution Center, frames the shift simply: where you run AI matters as much as which AI you run. His recommendation is to handle routine agentic tasks on hardware the organization already owns, reducing the marginal cost of each token to essentially the price of electricity. Reserve metered cloud access for the largest frontier models when specialized capability justifies the expense.
Dell Deskside Agentic AI, launched in May, executes this approach by running production-ready agents on Dell workstations equipped with the Nvidia NemoClaw open-source stack. The system handles validated workflows for coding, research, and private assistants on models ranging from compact 30-billion-parameter versions to trillion-parameter frontier models.
Agents operate inside an OpenShell environment that enforces policy-based privacy and security rules while logging all actions. When a prototype requires larger infrastructure, it can migrate to Dell PowerEdge servers in the data center without architectural changes. The system targets organizations that cannot send work to public cloud—engineers keeping source code in-house or researchers handling pre-publication data under privacy regulations.
Analysis by Signal65 and Futurum, commissioned by Dell, found the approach delivers up to 87% savings on token spending over two years compared to public cloud APIs, with break-even occurring in as few as three months.
Patel distills his guidance to six words: start local, govern early, scale smart. The cost structure of agentic AI becomes a design decision made before the first agent ships. Organizations that determine workload placement and establish governance frameworks early transform token bills from uncontrolled expenses into managed line items.
These details were first reported by Business Insider.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call