Quick Facts
- Laguna S 2.1 is a 118-billion-parameter Mixture-of-Experts model that activates only 8 billion parameters per token, trained in under nine weeks on 30 trillion tokens.
- The model scores 78.5% on SWE-Bench Multilingual and 70.2% on Terminal-Bench 2.1, topping all open, disclosed-size models on both benchmarks.
- No Western lab had released an open-weight model in this parameter class for 11 months before this launch.
San Francisco AI lab Poolside released Laguna S 2.1 on July 21, a 118-billion-parameter open-weight foundation model built for agentic coding tasks. The model tops published leaderboards among open, disclosed-size models on several key benchmarks, despite activating just 8 billion parameters per token.
The model uses a Mixture-of-Experts architecture, activating roughly 6.8% of its parameters on any given token. It supports a context window of up to 1 million tokens in both thinking and no-thinking modes, and is small enough to run on a single DGX Spark.
Benchmark Results
Laguna S 2.1 scores 78.5% on SWE-Bench Multilingual, placing first on Poolside’s published leaderboard. On Terminal-Bench 2.1 with thinking enabled, it scores 70.2%, again leading all open, disclosed-size models. On DeepSWE v1.1, it scores 40.4% against DeepSeek-V4-Pro-Max’s 9.0%, with roughly one-sixth the active parameters.
Closed frontier models, including Claude Fable 5 and Kimi K3, still lead on several benchmarks. But Laguna S 2.1 holds its own against much larger open systems, including DeepSeek-V4-Pro-Max, NVIDIA’s Nemotron 3 Ultra, and Thinking Machines’ Inkling.
Training and Infrastructure
Pre-training began May 22, 2026, on 4,096 NVIDIA H200 GPUs. The full training run completed in under nine weeks on 30 trillion tokens. Poolside says this is the first model where reinforcement learning ran in FP8 precision.
Poolside publishes weights in BF16, FP8, INT4, and NVFP4 formats, along with official GGUF and MLX conversions. The company says it ships new models on roughly a five-week cadence through its internal Model Factory platform, which automates architecture ablations and reinforcement learning from code execution.
The Open-Weight Gap
No Western lab had released an open-weight model in the 118-billion-parameter class for 11 months before this launch, since gpt-oss-120b arrived in August 2025. In that window, Chinese labs including DeepSeek, Alibaba’s Qwen, and Moonshot’s Kimi dominated the open-weight category.
Co-CEO Jason Warner framed the release in direct terms. “The West needs open-weight models it can trust, run, and build on,” he said. “Laguna S 2.1 is our answer. It is a model that enterprises and governments can put into production today, on their own hardware, at a cost that makes agentic coding practical at scale.”
Co-founder and Co-CEO Eiso Kant argued that open models must match or beat closed equivalents to win adoption. “Users simply want the best intelligence for the task at hand,” Kant said. He added that the open ecosystem “will not win by being the best in its own category.”
Business Implications
Because Laguna S 2.1 runs on a single DGX Spark, engineering teams can move high-volume agentic coding work off metered APIs and onto hardware they control. For companies running large numbers of coding agents, that shift cuts per-token costs and removes third-party data exposure.
Poolside spent most of its three-year existence selling coding models to governments and defense agencies. The open release of Laguna S 2.1 marks a shift toward the broader enterprise market, arriving as questions about who supplies the West’s open-weight AI have moved from research circles into boardrooms and Washington.
Read more: Poolside drops Laguna S 2.1, an open-weight coding model that beats rivals 10x its size
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
