Benchmarks
Measured results, treated like data rather than opinion. Every benchmark on this site records the full hardware and software configuration it was run with — pinned versions, exact commands, sample counts — so you can reproduce it or judge whether it applies to your machine.
How to read a benchmark page here
- The tested environment block states the hardware, OS, drivers, and exact software versions. If your stack differs, the numbers are a reference point, not a prediction.
- Interpretations are scoped to the models, quantizations, and workloads actually run — and say so. A ranking on one model is not a claim about every model.
- Where available, raw output files are linked alongside the tables, so anyone can recompute the claims from the source data.
- Pages are pinned to the versions tested, not to "current" software — upstream moves faster than test runs do.
All benchmarks
- RTX PRO 5000 Blackwell — llama.cpp Benchmarks
Prompt processing and token generation speeds for popular GGUF models on dual RTX PRO 5000 Blackwell, with full reproduction setup.
- PCIe 5.0 x8/x8 Isn't Enough: Testing GPU P2P Bandwidth on Z890
Two CPU-attached Gen5 x8 GPUs that move only 5.3 GB/s GPU-to-GPU in one direction and 2.0 GB/s in the other — measured with a custom probe and NVIDIA's p2pBandwidthLatencyTest, and why PCIe generation and width alone mislead.
- ROCm vs Vulkan on AMD Strix Halo: llama.cpp Benchmarks
Measured llama.cpp prompt-processing and token-generation rates for ROCm/HIP vs Vulkan backends on Ryzen AI Max+ / Radeon 8060S — dense and MoE models, short and large context, plus MTP speculative decoding on a realistic coding workload.