118B total / 8B active MoE · 1M ctx · agentic coding. Hub: poolside/Laguna-S-2.1. Blog: introducing Laguna S 2.1. Models: poolside.ai/models.
| piece | choice |
|---|---|
| body | 48 layers · hidden 3072 · vocab 100352 |
| FFN | L0 dense (I=12288) · L1–47 MoE |
| MoE | 256 routed · top-10 · 1 shared add · scale 2.5 |
| router | sigmoid + aux-free bias (not softmax) |
| attn | 12 global · 36 SWA (win 512) · ≈1:3 · GQA 8 KV |
| heads | global 48Q · SWA 72Q · head_dim 128 |
| PE | YaRN on global · RoPE on SWA |
| gate | per-head softplus before o_proj · QK-norm |
| thinking | off / max (default max) · preserved reasoning_content |
text → 48× (GQA + QK-norm + RoPE/YaRN + softplus gate) → SWA/global 1:3 → L0 dense · L1–47 MoE sigmoid top-10/256 + shared add → text (+ thinking / tools)
Benches (2026-07-21, pool): TB2.1 70.2 (no-think 60.4) · SWE-Multi 78.5 · SWE-Pro 59.4 · DeepSWE 40.4 · Atlas 46.2 · Toolathlon 49.7. Trajectories: trajectories.poolside.ai.
Train: 30T · 4096×H200 from 2026-05-22 · FP8 RL · multi-harness rollouts · ~409k envs. Serve: vLLM/SGLang/TRT-LLM/Ollama · DFlash draft · OpenRouter free 256K.
Chat: <think> · tools <tool_call>/arg_key ·
eos ids [2, 24] (〈|EOS|〉 + </assistant>) ·
parsers poolside_v1 · DFlash draft n_spec=15.