Quick Facts
- AWS will deploy Cerebras’ WSE-3 chips with 900,000 cores through Amazon Bedrock starting in Q3 2026
- The disaggregated architecture combining WSE-3 and AWS Trainium promises 5x faster AI inference speeds
- Cerebras’ chip delivers 21 petabytes per second memory bandwidth, 2,600 times faster than Nvidia’s Blackwell B200
Amazon Web Services announced a partnership with Cerebras Systems to bring the chip maker’s wafer-scale WSE-3 processors to AWS data centers. The collaboration will combine Cerebras’ massive AI chip with AWS’s own Trainium processors to create a specialized inference solution.
The WSE-3 chip contains 900,000 cores and 44 gigabytes of on-chip SRAM, manufactured using a 5nm TSMC process. The processor delivers 21 petabytes per second of memory bandwidth, approximately 2,600 times faster than Nvidia’s flagship Blackwell B200 chip.
AWS will deploy these processors through a disaggregated architecture where Trainium chips handle prefill work and compute the KV cache, then send results to the WSE via Amazon’s high-speed EFA interconnect. The Cerebras chip focuses exclusively on decode operations, generating thousands of output tokens per second.
“Inference is where AI delivers real value to customers, but speed remains a critical bottleneck for demanding workloads like real-time coding assistance and interactive applications,” said David Brown, Vice President of Compute & ML Services at AWS.
Cerebras CEO Andrew Feldman said the partnership will “bring the fastest inference to a global customer base” through AWS’s existing environment. The company recently raised $1 billion at a $23 billion valuation in February 2026 and plans an IPO in Q2 2026.
The solution targets AI coding applications where agents generate approximately 15 times more tokens per query than conversational chat. Recent benchmarks show the Cerebras CS-3 system running Llama 3.1 70B at 2,100 tokens per second per user, eight times faster than Nvidia’s H200.
AWS customers can expect a phased rollout starting in US-East (N. Virginia) and US-West (Oregon) regions by Q3 2026. The partnership represents AWS’s strategic move to challenge Nvidia’s 90% market share in AI inference chips through specialized hardware optimization.
Deloitte projects inference will account for two-thirds of all AI compute spending by 2026, up from one-third in 2023. This growing market creates urgent demand for faster inference solutions across the industry.
Read more: AWS will bring Cerebras’ wafer-size WSE-3 chip to its cloud platform
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
