Files
llama.cpp/scripts/rx580-bench/README.md
T
LumpiastyandClaude Opus 5 f9a5c231ed scripts: add the RX 580 benchmark harness
Encodes the production config, fixed corpus slices, repeat/median discipline
and the noise floor, so the measurement method does not have to be
rediscovered each time. Runs the corpus prefill test through llama-server and
the llama-bench sweep as a controlled cross-check, with interleaved A/B.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PZz44SLQvTXMyWGio6t9DZ
2026-09-10 17:06:01 +02:00

201 lines
7.9 KiB
Markdown

# rx580-bench
Prompt-processing benchmark harness for the RX 580 (Polaris) serving pod.
It exists so that no future session has to rediscover how to measure this
machine. It encodes the production config, fixed corpus slices, the repeat
and median discipline, and the noise floor.
## What it measures
Two things, deliberately:
1. **Corpus prompt processing** (the number we actually care about). Real
Polish prose from `pan-tadeusz.txt` pushed through `llama-server`'s
`/completion` endpoint with `n_predict: 1`, `cache_prompt: false`. This
mimics agentic tool-result parsing: a big blob of real text arriving cold.
Reported as `timings.prompt_per_second` at prompt lengths of about 4k, 16k
and 32k tokens.
2. **`llama-bench` sweep** (the controlled cross-check). Synthetic random
tokens, `pp2048`, `pp8192`, `tg32`. This is what prior work on this box
recorded, so it is the continuity metric.
They do not measure the same thing and they do not agree. `llama-bench`
pp2048 is a cold depth-0 batch; the corpus numbers include the cost of a
growing KV cache, so they fall off with prompt length. Both are useful.
## How to run it
One command from the workstation:
```
./scripts/rx580-bench/run.sh --build /root/arms/new-clean/build --label new-clean
```
Interleaved A/B, which is the only valid way to compare two builds:
```
./scripts/rx580-bench/run.sh \
--build /root/arms/new-clean/build --label new-clean \
--build-b /root/arms/new-fork/build --label-b new-fork \
--rounds 3
```
Summarize whatever has accumulated:
```
./scripts/rx580-bench/run.sh --summarize
```
`run.sh` copies the harness onto the pod and execs it there. The wrapper is
thin on purpose: all the methodology lives in the pod-side scripts, so it can
also be driven directly on the pod if kubectl is inconvenient.
The run is detached on the pod and then tailed, so Ctrl-C on the wrapper does
not orphan a half-finished run. `bench.sh` restarts `llama-swap` from an EXIT
trap on success, failure and interrupt alike.
### Pod access
```
export KUBECONFIG=/home/user/Projects/klaster/talos/generated/kubeconfig
kubectl -n llama exec -i deploy/supervisord -- sh -c '<command>'
```
The pod has `curl` and `python3` but **no `jq`**; the harness uses python3 for
all JSON.
## Files
| file | side | what |
|---|---|---|
| `run.sh` | workstation | one-command wrapper; installs and execs the rest |
| `bench.sh` | pod | orchestrator: stops llama-swap, waits for idle GPU, runs both harnesses, restarts llama-swap via trap |
| `ppbench.py` | pod | the corpus prompt-processing harness |
| `lbsweep.sh` | pod | the llama-bench sweep |
| `summarize.py` | pod | turns `results.txt` into a median/min-max table with delta significance |
Results append to `/root/bench/results.txt` as parseable `RESULT` /
`SUMMARY` / `LBRESULT` lines. Server logs land in `/root/bench/server-*.log`.
## The config it assumes
Production serving config, matching `/root/config.yaml`:
```
-t 6 -ngl 99 --n-cpu-moe 40 -b 2048 -ub 2048 -fa 1 --no-mmap
--ctx-size 40960 --no-warmup
```
Model:
```
/root/.cache/huggingface/hub/models--unsloth--Qwen3.6-35B-A3B-GGUF/snapshots/a483e9e6cbd595906af30beda3187c2663a1118c/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf
```
Qwen3.6-35B-A3B, 20.8 GiB, 40 layers, 256 experts of which 8 are used.
Machine: AMD RX 580 8 GB (Polaris/GCN, RADV, no fp16 accel), PCIe 3.0 x16,
12 CPU cores, 62 GB RAM.
The `llama-bench` equivalent of that config is:
```
-t 6 -ngl 99 -ncmoe 40 -b 2048 -ub 2048 -fa 1 -mmp 0 -p 2048,8192 -n 32 -r 3
```
## The corpus slices are constants, not recalibrated
`pan-tadeusz.txt` is 447334 chars / 482907 bytes of Polish text. Against this
model's tokenizer it runs **2.6702 chars/token**, and `/completion` reports
`prompt_n` identical to `/tokenize` (BOS offset 0).
The slices are `corpus[0:NCHARS]` and are baked into `ppbench.py`:
| target tokens | chars | utf-8 bytes |
|---|---|---|
| 4096 | 10776 | 11608 |
| 16384 | 43858 | 47175 |
| 32768 | 87165 | 93843 |
These hit the target `prompt_n` exactly. Every build is therefore measured on
byte-identical input. `ppbench.py` warns loudly if the reported `prompt_n`
ever stops matching the target, which is the signal that the table is stale.
Only recalibrate (`ppbench.py <build> <label> --calibrate`) if the model file
or the corpus changes.
## Reading the numbers
* `prompt_n` is the token count actually prefilled. It must equal the target.
* `prompt_ms` is wall-clock prefill time.
* `tps` / `prompt_per_second` is `prompt_n / prompt_ms * 1000`. Higher better.
* `pp2048` / `pp8192` are llama-bench prefill throughput at depth 0.
* `tg32` is token generation at DEFAULT depth (n_ctx 32). **It is NOT a valid
decode metric on this box.** At that size decode is launch/barrier bound, not
compute bound, and measures a regime serving never reaches: it read -36%
between two upstream commits that are at parity in real use (the cause was
the view-alias fix in ggml_vk_graph_optimize). For decode use **tg256 with an
explicit depth**, e.g. `-p 0 -n 256 -d 2048`. tg is also NOT monotonic in
n_ctx here - on one binary n_ctx 256 measured slower than both 32 and 544 -
so only ever compare same-depth cells.
## Noise floor: deltas under about 5 percent are unproven
This is the single most important thing in this directory.
`llama-bench`'s within-run error bars understate **cross-invocation** variance
by roughly **14x**. The same config measured `276.7 +/- 0.9` in one invocation
and `296.4 +/- 0.6` in another, 7 percent apart, with error bars that claimed
0.3 percent precision.
Therefore:
* Any comparison you want to draw a conclusion from must live **inside a
single `llama-bench` invocation**, using comma-separated sweeps.
* Two different *builds* cannot share an invocation, so run them
**interleaved A/B/A/B/A/B across at least 3 rounds** and compare medians.
That is what `--build-b` does.
* State explicitly that any cross-build delta under about 5 percent is
unproven. `summarize.py` labels them `unproven` for you.
## Traps
* **`-tb` is not a `llama-bench` flag.** It is a `llama-server` flag. Passing
it makes `llama-bench` print usage and exit silently, mid-sweep, which looks
exactly like a run that produced nothing. Always check the row count.
* **A ubatch larger than the prompt never fills.** With `-ub 2048`, prompt
lengths must be multiples of 2048 or the sweep measures nothing meaningful.
* **Nothing else may touch the GPU.** `bench.sh` stops `llama-swap` and
refuses to start if any `llama-*` process is still alive.
* **Never leave the pod not serving.** `bench.sh` restarts `llama-swap` from
an EXIT trap. If you bypass the script, restart it by hand:
`supervisorctl -c /root/supervisord.conf start llama-swap`.
* **Do not run `llama-cli` detached with closed stdin** on this box; it loops
forever. Use the server harness or `llama-bench -d` for depth timing.
* `--no-mmap` is deprecated upstream in favour of `--load-mode` / `-lm none`,
but still accepted and still what production passes.
## Known-good reference values
If a fresh run disagrees with these by much more than the noise floor,
something is wrong (something else on the GPU, a thermal problem, a bad
build) before you believe you found a speedup.
Model Qwen3.6-35B-A3B-UD-Q4_K_XL, production config, current upstream
(`f3f1a8f27`) and the deployed fork, measured 2026-09-08/09:
| metric | clean upstream | fork |
|---|---|---|
| corpus pp @ 4k | about 240 t/s | about 255 t/s |
| corpus pp @ 16k | about 206 t/s | about 225 t/s |
| corpus pp @ 32k | about 171 t/s | about 195 t/s |
| llama-bench pp2048 | see results.txt | about 285 t/s |
| llama-bench tg32 | | not a valid metric, see above |
Older reference points for the same model/config: `pp2048` depth-0 about
285 t/s. The historical "tg32 about 19 t/s" figure is retained only as trivia -
do not treat it as a target, and see the tg32 warning above.
Model load with `--no-mmap` takes about 26 s warm, longer cold.
A full 3-round A/B run of both harnesses takes roughly 100 minutes;
add about 16 minutes per extra arm per round.