BeamModel

Beam benchmarks

Scores reported by Reflection AI in the Beam launch announcement. Independent reproductions and head-to-head comparisons will be added once the weights are public.

Reflection AI states that Beam scores comparably to GLM-5.2 while using 3–4× less inference compute.

Coding & agentic

BenchmarkScore
SWE-Bench Verified80.9
Terminal-Bench v2.180.1
SWE-Bench Multilingual78.0

Reasoning

BenchmarkScore
AIME 202697.8
GPQA Diamond90.5
SciCode49.7

General & long context

BenchmarkScore
AA-LCR79.3
LongBench v265.5
IFBench79.7

Source: https://reflection.ai/blog/introducing-beam