Laguna S 2.1

Poolside · open MoE · weights public · OpenMDW-1.1 · 2026-07-21


118B total / 8B active MoE · 1M ctx · agentic coding. Hub: poolside/Laguna-S-2.1. Blog: introducing Laguna S 2.1. Models: poolside.ai/models.

piecechoice
body48 layers · hidden 3072 · vocab 100352
FFNL0 dense (I=12288) · L1–47 MoE
MoE256 routed · top-10 · 1 shared add · scale 2.5
routersigmoid + aux-free bias (not softmax)
attn12 global · 36 SWA (win 512) · ≈1:3 · GQA 8 KV
headsglobal 48Q · SWA 72Q · head_dim 128
PEYaRN on global · RoPE on SWA
gateper-head softplus before o_proj · QK-norm
thinkingoff / max (default max) · preserved reasoning_content
text → 48× (GQA + QK-norm + RoPE/YaRN + softplus gate)
  → SWA/global 1:3
  → L0 dense · L1–47 MoE sigmoid top-10/256 + shared add
  → text (+ thinking / tools)

Benches (2026-07-21, pool): TB2.1 70.2 (no-think 60.4) · SWE-Multi 78.5 · SWE-Pro 59.4 · DeepSWE 40.4 · Atlas 46.2 · Toolathlon 49.7. Trajectories: trajectories.poolside.ai.

Train: 30T · 4096×H200 from 2026-05-22 · FP8 RL · multi-harness rollouts · ~409k envs. Serve: vLLM/SGLang/TRT-LLM/Ollama · DFlash draft · OpenRouter free 256K.

limits: harness overfitting · nested JSON tool args · overthinking · off/max only · text-only

Chat: <think> · tools <tool_call>/arg_key · eos ids [2, 24] (〈|EOS|〉 + </assistant>) · parsers poolside_v1 · DFlash draft n_spec=15.

package: studies/models/laguna-s-2.1 · compare · code truths: sigmoid router · shared add · softplus attn gate

← home