Laguna S 2.1 Review: 6 Hours of Failed AI Coding Tests

A Single Prompt, Six Hours of Frustration

Not every model launch lives up to the hype. A developer found that out the hard way while testing Poolside’s newly released Laguna S 2.1, a 118 billion parameter coding model pitched as a capable agent for software engineering. His review on YouTube walks through six painful hours of broken tool calls, empty UI outputs, and zero successful runs on a prompt other models had nailed in one shot.

The test centred on a familiar challenge: build a 3D globe dashboard. He’d already done this exact task with Tencent HY3 and Kimi K3, and both delivered a working, visually rich result right away. Laguna S 2.1 never got close.

The Model’s Big Promises

Poolside markets Laguna S 2.1 as a coding agent first, not a general chatbot. It runs on a mixture-of-experts design, activating just 8 billion of its 118 billion parameters per token. The idea is to pair large-model knowledge with small-model speed. On Poolside’s own benchmark table, it sits just behind Tencent HY3 on the terminal coding task, scoring 70.2 versus 71.7. Close enough to call it near parity.

A 250,000 token context window and native tool-calling support round out the pitch. And as one early reviewer pointed out, the permissive license and free access on OpenRouter make it tempting for developers who want to steer clear of proprietary lock-in.

The 3D Globe Dashboard That Never Rendered

The tester started on Poolside’s own chat interface at chat.poolside.ai. He entered the globe dashboard prompt and got back a near-blank page. No 3D globe. No dashboard layout. No visualisation, no data, just a broken shell.

The prompt itself wasn’t the problem. The same instruction had worked flawlessly on HY3, producing an interactive globe with real-time stats. With Laguna S 2.1, the model either misread what was being asked or didn’t have the tool-calling precision to pull the app together.

Two Agent Environments, Zero Success

Thinking the issue might be specific to the chat interface, he switched to two agent environments that had integrated the model: Hermes and Open Code. Both gave free access to Laguna S 2.1.

In Hermes, the model started responding but kept dropping the connection with a “reconnecting” message, even though the network was fine. Tool calls never finished. Open Code showed the same pattern. He tried prompting the model to self-correct, but after multiple attempts over three hours, the output stayed broken. “Nothing worked,” he said.

These aren’t small edge-case glitches. A coding agent that can’t get through a multi-step frontend task, whether from its own chat or through established agent wrappers, points to a deeper gap in integration or capability.

How Other Models Handled the Same Task

This wasn’t an unreasonable ask. He’d run the exact same 3D globe dashboard prompt on Tencent HY3 and Kimi K3 before. Both produced a complete, interactive dashboard on the first try. Those results are up on his channel for a direct comparison. The prompt wasn’t the issue. Laguna S 2.1 just couldn’t deliver.

Questioning the Benchmark Numbers

What’s hardest to square away, in his view, is the gap between the published benchmarks and what actually happened. A one-point difference on a leaderboard doesn’t mean much if the model keeps failing at a practical task that its supposed rival finishes in seconds. His verdict, plainly stated: “I think it is a flop model.”

He also questions why Poolside put Gemini 1.5 in its comparison table at all, since it’s clearly in a different weight class. That choice, on top of his own null results, leaves him doubting whether the reported scores reflect real, repeatable performance.

Mixed Signals From the Community

Not every test has gone badly. A quick demo by AI With Yousaf, using the Kilocode extension, showed Laguna S 2.1 correctly listing a directory’s contents and reading a simple HTML file. Kacper Rutkiewicz’s in-depth review got a working 2048 game, a transcript-to-deliverables tool, and a self-correcting REST API out of it, paired with the PI agent and Hermes. He did flag some limits too: no vision support, a “thinking-mode token tax,” and a ceiling on how far it can go with truly frontier-class problems.

None of that erases the six-hour failure on the globe prompt. But it does suggest Laguna S 2.1 can handle simpler, well-scoped tasks while stumbling on complex UI builds that need precise tool orchestration. For a model marketed as a coding agent, that kind of inconsistency is a warning sign.

What the Architecture Couldn’t Save

The mixture-of-experts setup, activating only 8B parameters per token, is supposed to deliver efficiency without giving up depth. Here, that efficiency didn’t translate into reliable tool use or multi-step planning. The repeated connection drops, empty outputs, and failed self-correction all point to an agent that isn’t ready for real, multi-file development work.

Real-World Tests Define a Model’s Worth

A benchmark score is a snapshot. A developer’s six-hour battle is a verdict. Until Laguna S 2.1 can reliably handle the same prompts its competitors clear on the first try, the skepticism will stick around. The community’s mixed early results back up a simple point: no spec sheet or leaderboard ranking replaces hands-on testing with real tasks.

Frequently Asked Questions

What exactly is Laguna S 2.1?

It’s a 118 billion parameter mixture-of-experts coding model from Poolside AI. It activates 8 billion parameters per token and supports up to 250,000 tokens of context. It’s released under a permissive license and can be used for commercial projects.

Why did the model fail the 3D globe dashboard test?

The developer tested it on three platforms: the official chat, Hermes, and Open Code. He got either empty pages or broken tool calls each time. The exact cause isn’t clear, but the model’s tool-calling and self-correction abilities seem to fall short on complex, front-end-heavy tasks.

Is Laguna S 2.1 worth using for other coding projects?

Some users have had success with simpler tasks, like basic file operations or game generation. But if your project needs precise UI rendering, multi-step agent workflows, or reliable tool use, you might run into the same problems documented in this test.

SAVE WHILE SHOPPING 1 - How to make money from home online part time jobs

Recommended For You