<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>How Quant</title><description>Local AI Benchmarks, Hardware &amp; Inference Guides</description><link>https://howquant.dev/</link><language>en-us</language><item><title>Check Your GPU&apos;s Real PCIe Bandwidth on Linux</title><link>https://howquant.dev/guides/check-gpu-pcie-bandwidth/</link><guid isPermaLink="true">https://howquant.dev/guides/check-gpu-pcie-bandwidth/</guid><description>See what PCIe link width and generation your GPU actually negotiated (lspci vs nvidia-smi), what the numbers should be, and when a downgrade actually matters.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate></item><item><title>RTX PRO 5000 Blackwell — llama.cpp Benchmarks</title><link>https://howquant.dev/benchmarks/rtx-pro-5000-llama-cpp/</link><guid isPermaLink="true">https://howquant.dev/benchmarks/rtx-pro-5000-llama-cpp/</guid><description>Two RTX PRO 5000 Blackwell cards (72 + 48 GB) running dense and MoE models on llama.cpp — prompt processing at pp512–pp8192 plus token generation, with the full pinned configuration.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate></item><item><title>PCIe 5.0 x8/x8 Isn&apos;t Enough: Testing GPU P2P Bandwidth on Z890</title><link>https://howquant.dev/benchmarks/z890-gpu-p2p-bandwidth/</link><guid isPermaLink="true">https://howquant.dev/benchmarks/z890-gpu-p2p-bandwidth/</guid><description>Two CPU-attached Gen5 x8 GPUs that move only 5.3 GB/s GPU-to-GPU in one direction and 2.0 GB/s in the other — measured with a custom probe and NVIDIA&apos;s p2pBandwidthLatencyTest, and why PCIe generation and width alone mislead.</description><pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Run MiniMax H3 Locally with Diffusers</title><link>https://howquant.dev/guides/run-minimax-h3-locally/</link><guid isPermaLink="true">https://howquant.dev/guides/run-minimax-h3-locally/</guid><description>MiniMax H3 generates video and audio together and needs two ~62 GB BF16 components. Run it on one high-memory GPU or split it across two 48 GB cards — Modular Diffusers, no ComfyUI.</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Video Resolution Calculator</title><link>https://howquant.dev/tools/video-resolution-calculator/</link><guid isPermaLink="true">https://howquant.dev/tools/video-resolution-calculator/</guid><description>Find video and image generation sizes divisible by your model&apos;s required multiple (8, 16, 32 or 64) for a target aspect ratio — or check whether a resolution you already have divides cleanly.</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate></item><item><title>llama.cpp Common Options</title><link>https://howquant.dev/reference/llama-cpp-options/</link><guid isPermaLink="true">https://howquant.dev/reference/llama-cpp-options/</guid><description>The llama.cpp flags you will actually use — GPU offloading, context and KV cache, performance, and sampling — grouped and one line each.</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Install ROCm 7.14 on AMD Strix Halo / Ryzen AI Max+</title><link>https://howquant.dev/guides/install-rocm-7-14-strix-halo/</link><guid isPermaLink="true">https://howquant.dev/guides/install-rocm-7-14-strix-halo/</guid><description>Install and verify ROCm 7.14 with gfx1151 support on a Strix Halo (Ryzen AI Max+) laptop — repository, the amdrocm-core-dev package set, permissions, and rocminfo/HIP/llama.cpp verification.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Run llama.cpp on AMD Strix Halo: ROCm and Vulkan Setup</title><link>https://howquant.dev/guides/llama-cpp-strix-halo/</link><guid isPermaLink="true">https://howquant.dev/guides/llama-cpp-strix-halo/</guid><description>Practical llama.cpp setup for Strix Halo (Ryzen AI Max+ / Radeon 8060S) — both backends (ROCm/HIP and Vulkan), which to pick, build commands, device selection, memory and large-context notes.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Recommended llama.cpp Setup for Strix Halo</title><link>https://howquant.dev/guides/llama-cpp-strix-halo-recommended-setup/</link><guid isPermaLink="true">https://howquant.dev/guides/llama-cpp-strix-halo-recommended-setup/</guid><description>The starting point we recommend on AMD Strix Halo (Ryzen AI Max+ / Radeon 8060S) for llama.cpp: ROCm as the default backend, Vulkan as the fallback, and when MTP speculative decoding is worth enabling.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>RTX PRO 5000 Blackwell eGPU on Linux with Sonnet 850T5</title><link>https://howquant.dev/guides/sonnet-850t5-rtx-pro-5000-linux-egpu/</link><guid isPermaLink="true">https://howquant.dev/guides/sonnet-850t5-rtx-pro-5000-linux-egpu/</guid><description>Get an NVIDIA RTX PRO 5000 Blackwell working as a Thunderbolt 5 eGPU under Linux — the proven boot/hot-plug sequence, the Sonnet 850T5 PCIe topology, small-BAR behavior, every workaround and its sourcing, BIOS outcomes.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Set Up Vulkan Compute on AMD Strix Halo / Radeon 8060S</title><link>https://howquant.dev/guides/vulkan-compute-strix-halo/</link><guid isPermaLink="true">https://howquant.dev/guides/vulkan-compute-strix-halo/</guid><description>Get Vulkan compute working on Strix Halo with the Mesa RADV driver — Kisak PPA setup, exact packages, vulkaninfo verification, and a llama.cpp Vulkan build.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>GPU Model Fit Calculator</title><link>https://howquant.dev/tools/gpu-model-fit/</link><guid isPermaLink="true">https://howquant.dev/tools/gpu-model-fit/</guid><description>Calculate whether an LLM fits in GPU memory. Uses exact GGUF artifact sizes for preset models, parameter-count estimates with provenance labels otherwise, and includes measured GTT usage from the How Quant test machine.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>ROCm vs Vulkan on AMD Strix Halo: llama.cpp Benchmarks</title><link>https://howquant.dev/benchmarks/strix-halo-rocm-vs-vulkan/</link><guid isPermaLink="true">https://howquant.dev/benchmarks/strix-halo-rocm-vs-vulkan/</guid><description>Measured llama.cpp prompt-processing and token-generation rates for ROCm/HIP vs Vulkan backends on Ryzen AI Max+ / Radeon 8060S — dense and MoE models, short and large context, plus MTP speculative decoding on a realistic coding workload.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Linux eGPU Startup Script Reference (Sonnet 850T5)</title><link>https://howquant.dev/reference/linux-egpu-startup-script-sonnet-850t5/</link><guid isPermaLink="true">https://howquant.dev/reference/linux-egpu-startup-script-sonnet-850t5/</guid><description>The recovered egpu.sh for an RTX PRO 5000 Blackwell in a Sonnet 850T5 under Linux, block by block — runtime power control, the Link Control 2 poke, persistence mode, power cap, audio unbind — with what each step does and how its rationale is sourced.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>AMD Strix Halo Local AI Compatibility</title><link>https://howquant.dev/reference/strix-halo-local-ai-compatibility/</link><guid isPermaLink="true">https://howquant.dev/reference/strix-halo-local-ai-compatibility/</guid><description>Which local-AI runtimes work on Strix Halo (Ryzen AI Max+ / Radeon 8060S) — llama.cpp on ROCm and Vulkan tested at pinned versions, PyTorch tracked but untested, with backend and status per combination.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Understanding RAM, GTT, and GPU Memory on AMD Strix Halo</title><link>https://howquant.dev/reference/strix-halo-memory-gtt/</link><guid isPermaLink="true">https://howquant.dev/reference/strix-halo-memory-gtt/</guid><description>Why tools report different memory sizes on Strix Halo — unified memory architecture, GTT aperture vs usable RAM, what the kernel, ROCm, Vulkan and llama.cpp each report, with measured numbers.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Build llama.cpp with CUDA for NVIDIA Blackwell</title><link>https://howquant.dev/guides/build-llama-cpp-blackwell/</link><guid isPermaLink="true">https://howquant.dev/guides/build-llama-cpp-blackwell/</guid><description>Compile llama.cpp for Blackwell GPUs (sm_120) with a current CUDA toolkit — correct architecture flags, build commands, and fixes for the errors you will actually hit.</description><pubDate>Sat, 22 Aug 2026 00:00:00 GMT</pubDate></item></channel></rss>