Free U.S. Shipping Up to 10 LBS | Hassle-free Return Policy | 24/7 Customer Support

Buy now pay later with Shop Pay!

NVIDIA Vera CPU: The First Chip Built for AI Agents

NVIDIA Vera CPU system being hand-delivered to an AI lab

Orange Hardwares |

NVIDIA has moved its Vera CPU from announcement to real customer hardware. On May 15, NVIDIA Vice President of Hyperscale and High-Performance Computing Ian Buck personally delivered the first Vera systems to three leading AI labs in the San Francisco Bay Area.

Anthropic, OpenAI, and SpaceXAI each received a unit that day, and Oracle Cloud Infrastructure followed on the next business day at its Santa Clara facility.

The rollout marks the moment NVIDIA's agentic CPU strategy left the lab and entered production environments at some of the industry's most closely watched companies.

Why NVIDIA Built a CPU for Agents

Jensen Huang introduced Vera at GTC San Jose in March, describing it as NVIDIA's next major standalone business line. The chip targets a problem that has grown alongside the rise of AI agents: GPUs handle the heavy model computation, but agents generate an enormous amount of surrounding work that lands on the CPU instead.

Every sandboxed tool call, every orchestration step, every retrieval operation across a long context window pulls CPU cycles that traditional server chips were not designed to prioritize.

Buck framed the shift plainly, noting that agentic AI is creating a new CPU moment inside the AI factory as models move from simply answering questions to actively completing tasks.

Vera was built around that demand rather than around the core density priorities that shaped earlier server processors.

The specifications back up that focus:

  • 88 custom NVIDIA-designed Olympus cores
  • 1.2 terabytes per second of memory bandwidth
  • 50 percent faster per-core performance than its predecessor under sustained load

NVIDIA says that combination keeps agentic workloads moving efficiently across the entire AI factory, which translates into faster responses for end users running concurrent, real-time tasks.

Four Deliveries with Four Different Reactions

The rollout read almost like a cross-town tour of the AI industry's biggest names, each stop shaped by the priorities of the company receiving the hardware.

1. Anthropic, San Francisco

The first handoff took place at Anthropic's SoMa offices, where James Bradbury, the company's head of compute, accepted the delivery.

Buck walked Bradbury through the server using a bare Vera motherboard as a guide. Bradbury said scaling compute remains an important accelerant for model growth and called Vera a promising addition to the ecosystem for agentic workloads.

2. OpenAI, Mission Bay

The second stop moved outdoors, onto a balcony at OpenAI's headquarters. Sachin Katti, OpenAI's head of compute infrastructure, thanked Buck for bringing the system over in person. Buck pulled a screwdriver from his pocket and opened the chassis to show the internal layout.

3. SpaceXAI, Palo Alto

The final delivery of the day went to SpaceXAI, where Elon Musk reviewed the system's interior and asked detailed questions about core layout, memory architecture, and cooling.

SpaceXAI is evaluating Vera for reinforcement learning workloads and the agent-based simulation pipelines that support its training stack.

4. Oracle Cloud Infrastructure, Santa Clara

The following Monday, a team from OCI toured an unboxed Vera system inside the Oracle AI Customer Excellence Center.

Karan Batta, who leads product management at OCI, and Gary Miller, the company's chief customer and partner success officer, joined the walkthrough while an NVIDIA GPU rack processed customer workloads nearby.

What Vera Solves for Cloud Providers

Buck explained the underlying technical problem during the OCI stop. When an AI model receives a complex question, the answer is rarely ready to output immediately.

The model often needs to generate code, such as a Python script, to work through the problem before it can respond.

That step is CPU-intensive, and it is a major reason demand for capable CPUs has grown so quickly alongside agentic AI adoption.

OCI is treating that demand as a long-term infrastructure priority. Batta and Miller pointed to a few specific reasons why:

  • OCI plans to deploy hundreds of thousands of Vera CPUs starting in 2026, driven by the sustained performance agentic AI requires at scale
  • Batta described Vera's architecture as purpose built for high throughput reasoning workloads
  • The platform offers the efficiency, density and footprint OCI needs for its next generation of enterprise AI services
  • OCI is the first cloud provider to deploy Vera at hyperscale, giving its enterprise customers access to production grade agentic AI infrastructure at a scale competitors cannot yet match

Miller said the OCI team is looking forward to seeing how customers react to the new hardware as they test and validate their own agentic workloads on it.

Vera at a Glance

Category

Details

What it is

NVIDIA's first custom CPU, designed specifically for agentic AI

What it handles

Orchestration, tool calling, reinforcement learning workloads, data analytics, agent sandboxing, long context state management

Who it's for

AI labs, cloud providers, and enterprises running agentic AI at scale

Core specs

88 custom Olympus cores, 1.2 TB/s memory bandwidth, 50 percent faster per core under full load

How Vera Fits Into NVIDIA's Broader Platform

Vera is one piece of NVIDIA's wider hardware codesign strategy. For buyers exploring NVIDIA products across this ecosystem, the platform also includes

  • The Rubin GPU
  • The BlueField-4 DPU
  • Spectrum-X networking
  • The MGX rack architecture

Beyond running as a standalone CPU system, Vera also serves as the host processor inside the Vera Rubin NVL72 platform, where it connects to a pair of Rubin GPUs through NVIDIA's second-generation NVLink-C2C interconnect.

In that paired configuration, Vera and Rubin share a unified memory architecture designed to keep GPU utilization high. Vera's cores and interconnect handle the orchestration, control, and data movement needed to keep GPUs fed with work, and NVIDIA says this approach runs at roughly twice the energy efficiency of traditional infrastructure setups.

As Vera Rubin systems reach hyperscalers and cloud providers, enterprise buyers are already asking what this generational shift means for their existing fleets.

For a closer look at how the new platform is affecting H100 and A100 prices, including current depreciation trends and secondary market data, see our full pricing breakdown.

Recommended: Why GPU Prices Are Surging in 2026 (And What to Buy Instead)

Reading the Delivery as a Market Signal

The choice to hand-deliver Vera to four companies in person, rather than simply announce general availability, says something about how NVIDIA wants this launch understood. Each stop paired the hardware with a direct conversation between Buck and the executive responsible for that company's compute strategy.

It also reinforced a point NVIDIA has been making since Huang's GTC keynote in March: agentic AI needs dedicated CPU capacity alongside GPU compute. With OCI alone planning hundreds of thousands of units starting this year, that message appears to be landing with the market's largest buyers.

What Comes Next

With Vera now shipping to production customers rather than sitting in demonstration labs, the real test begins. Anthropic, OpenAI, SpaceXAI and OCI will each put the chip through its own workload patterns over the coming months, from reinforcement learning pipelines to enterprise-scale reasoning tasks. OCI's plan to deploy hundreds of thousands of units starting in 2026 suggests the industry views agentic CPU capacity as an urgent priority.

The delivery marks a clear signal from NVIDIA about where it sees the next phase of AI infrastructure heading. As agents take on more autonomous, multi-step tasks, the CPU work behind the scenes becomes just as important as the GPU compute that gets most of the public attention. Vera is NVIDIA's answer to that shift, and its first customers are already putting it to work.

Frequently Asked Questions

Q: What makes Vera different from a standard server CPU?

A: Vera is engineered around agentic workloads specifically. Its 88 Olympus cores and 1.2 TB/s memory bandwidth target orchestration and tool calling rather than raw batch processing alone.

Q: Can Vera run without a GPU?

A: Yes. NVIDIA ships Vera as a standalone CPU system, in addition to pairing it with Rubin GPUs inside the NVL72 platform.

Q: Why does agentic AI need more CPU power?

A: Agents generate code, manage long context windows and coordinate tool calls between steps. That work scales directly with agent volume.

Q: Who has Vera hardware right now?

A: Anthropic, OpenAI, SpaceXAI and Oracle Cloud Infrastructure all received systems in May 2026, with OCI planning a hyperscale rollout.

Leave a comment

Please note: comments must be approved before they are published.

Don’t Leave Yet, Wait!

Request a free quote now for exclusive pricing or bulk discounts. Save big before you leave!

By providing a telephone number and submitting this form you are consenting to be contacted by SMS text message. Message & data rates may apply. You can reply STOP to opt-out of further messaging.

Don't miss out
Need Assistance?

Request a quote for exclusive pricing or bulk orders.

By providing a telephone number and submitting this form you are consenting to be contacted by SMS text message. Message & data rates may apply. You can reply STOP to opt-out of further messaging.