Deploying Beam
The weights are not public yet. The commands below follow each engine’s standard usage and will be tested and updated on release day. Replace the placeholder repo id with the official one.
Before you start
At FP8, the weights alone take about 500 GB, so plan for a multi-GPU node such as 8× H200 or 8× B200. For quantized GGUF builds on Apple Silicon or CPU, see the VRAM calculator.
1. Download the weights
pip install -U "huggingface_hub[cli]"
huggingface-cli download <beam-repo-id> --local-dir ./beam2. vLLM (recommended for serving)
pip install -U vllm
vllm serve ./beam \
--tensor-parallel-size 8 \
--max-model-len 131072 \
--port 80003. SGLang
pip install -U "sglang[all]"
python -m sglang.launch_server \
--model-path ./beam \
--tp 8 \
--port 80004. llama.cpp (GGUF, consumer hardware)
Community GGUF quantizations usually appear within days of a release. Large models are split into several files, so point -m at the first shard.
llama-server \
-m ./beam-gguf/<first-shard>.gguf \
-c 32768 \
-ngl 99 \
--port 80005. Test the OpenAI-compatible endpoint
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "./beam", "messages": [{"role": "user", "content": "Hello"}]}'