Quick Facts
- Nvidia is developing a new AI inference chip expected to debut at GTC 2026 conference March 16-19
- The chip integrates technology from Nvidia’s $20 billion Groq acquisition in December 2024
- OpenAI will be among the earliest adopters with access to 3 gigawatts of dedicated inference capacity
Nvidia Corp. is developing a dedicated inference processor that will launch at its annual GTC developer conference later this month, according to a Wall Street Journal report. The new chip integrates technology from the company’s $20 billion acquisition of Groq Inc. in December.
The chip addresses growing demand for specialized inference processors as AI companies seek more efficient alternatives to traditional GPUs for running trained AI models in production. The new processor uses Groq’s language processing unit architecture, which delivers significantly lower energy consumption for inference tasks.
OpenAI has received early access to the new chip and will become one of its first major customers. The partnership stems from Nvidia’s recent $30 billion investment in OpenAI’s $110 billion funding round announced last week.
Groq Technology Integration
Nvidia’s acquisition of Groq represents the company’s largest purchase ever, surpassing its $7 billion Mellanox deal in 2019. The $20 billion transaction was structured as a nonexclusive technology license and included hiring Groq’s founding CEO Jonathan Ross and President Sunny Madra.
Groq’s chips use a novel architecture optimized for inference workloads. The technology enables much lower energy usage compared to traditional GPUs, addressing cost concerns that have driven companies to seek alternatives to Nvidia’s processors.
CEO Jensen Huang told employees the deal will expand Nvidia’s capabilities. “We plan to integrate Groq’s low-latency processors into the NVIDIA AI factory architecture, extending the platform to serve an even broader range of AI inference and real-time workloads,” Huang wrote in an internal email.
Market Shift to Specialized Chips
The inference chip reflects broader industry trends as AI workloads shift from training to production deployment. A November 2025 Futurum Group survey found GPUs accounted for 58% of data center compute spending in 2025. However, specialized processors are expected to lead growth in 2026 at 22%, outpacing GPUs at 19%.
Nvidia faces increasing competition in the inference market from Google’s TPUs, Amazon’s Trainium2 chips, and startups like Cerebras Systems. Amazon claims its Trainium2 delivers 30-40% better price-performance than Nvidia’s H100-based instances.
Under its investment agreement with OpenAI, Nvidia has secured commitments for 3 gigawatts of dedicated inference capacity and 2 gigawatts of training on its upcoming Vera Rubin systems. This capacity adds to OpenAI’s existing infrastructure across Microsoft Azure, Oracle Cloud, and CoreWeave.
The chip launch at GTC 2026 will provide more details on technical specifications and availability. Nvidia had $60.6 billion in cash and short-term investments as of October, giving it substantial resources to compete in the evolving AI chip market.
Read more: Report: Nvidia is working on a top secret AI inference chip that could debut next month
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
