Beam benchmarks
Scores reported by Reflection AI in the Beam launch announcement. Independent reproductions and head-to-head comparisons will be added once the weights are public.
Reflection AI states that Beam scores comparably to GLM-5.2 while using 3–4× less inference compute.
Coding & agentic
| Benchmark | Score | |
|---|
| SWE-Bench Verified | 80.9 | |
| Terminal-Bench v2.1 | 80.1 | |
| SWE-Bench Multilingual | 78.0 | |
Reasoning
| Benchmark | Score | |
|---|
| AIME 2026 | 97.8 | |
| GPQA Diamond | 90.5 | |
| SciCode | 49.7 | |
General & long context
| Benchmark | Score | |
|---|
| AA-LCR | 79.3 | |
| LongBench v2 | 65.5 | |
| IFBench | 79.7 | |
Source: https://reflection.ai/blog/introducing-beam