This is the Windows fork of Gufo: a native Windows port for AMD Strix
Halo (Ryzen AI Max+ 395 with Radeon 8060S, gfx1151), maintained on this
repository's feature/window-native branch. Upstream Gufo targets Linux;
the original README — models, benchmarks, philosophy and the Linux build —
follows below the separator.
Getting started: download a ready-made build from the
Releases page, unzip the
gufo-<version>-windows-gfx1151-<commit>.zip asset (for example
gufo-0.8.0-windows-gfx1151-1b2da4a.zip) and run gufo.exe from the
extracted folder — no installation needed. Every update to the port
publishes its own release, so older builds stay available. A recent AMD
graphics driver is required.
Next, download a model. The hf command ships with Hugging Face's
Python package (pip install -U huggingface_hub). This fetches the
Qwen3.8 Flash-Next UD-Q4_K_XL GGUF (four shards), the shared MTP predictor
and the vision projector into the default Hugging Face cache
(.cache\huggingface\hub in your user profile); see the
model guide for other quants
and the full file list:
hf download unsloth/Qwen3.8-Flash-Next-GGUF `
--revision 38bb39ee97821de2c9009abb7e93950eec396e66 `
--include "UD-Q4_K_XL/*" "MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf" "mmproj-BF16.gguf"Then serve it with MTP speculative decoding at the full 256K context for
two concurrent sessions. Run PowerShell from your user profile folder so
the .cache paths below resolve; the loader finds the remaining shards
and the MTP predictor automatically:
.\Downloads\gufo-0.8.0-windows-gfx1151-1b2da4a\gufo.exe serve llm --model .\.cache\huggingface\hub\models--unsloth--Qwen3.8-Flash-Next-GGUF\snapshots\38bb39ee97821de2c9009abb7e93950eec396e66\UD-Q4_K_XL\Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf --speculative mtp `
--served-model-name qwen3.8-flash-next `
--context 262144 --sessions 2 `
--cache-ram-bytes 8589934592 `
--cache-disk .\.cache\gufo-qwen4exp `
--cache-disk-bytes 21474836480 `
--cache-disk-staging-bytes 8589934592 `
--temperature 1 --top-p 0.95 --top-k 20 --min-p 0 `
--think on --reasoning-effort xhigh --preserve-thinking on `
--repeat-penalty 1 --frequency-penalty 0 --presence-penalty 0 `
--host 0.0.0.0 --port 11434 --log-progress --kv-cache q8_0-q4_kThe --kv-cache option shrinks the attention KV cache so deep 256K
sessions fit in memory: q8_0 stores both planes as Q8_0 blocks (about
half the bytes), and q8_0-q4_k stores Q8_0 keys with Q4_K super-block
values (9,984 bytes per token vs 24,576 for F16, −59%). Both are lossy
and off by default; f16 (the default) keeps full precision. Snapshots
and disk-cache entries are not shared across cache modes, so toggling
the option re-prefills existing conversations once.
When the server is up, point OpenAI-compatible clients at
http://localhost:11434/v1 (see the API contract).
You can also build it yourself on Windows with MSVC, the TheRock HIP SDK and vcpkg.
Gufo is a vertical local inference engine specifically built and optimized for the AMD Strix Halo hardware:
Ryzen AI MAX+ 395 systems with Radeon 8060S (gfx1151), up to 128 GiB of unified memory.
Contributions are welcome!
See the changelog and GitHub Releases for user-facing changes and release history.
Tip
There are two Gufo variants that haven't been merged yet: A Windows port and Gufo RDMA for Dual Strix Halo. Check them out!
All model documentation lives under docs/models:
| Model | Inference modes | Hugging Face weights | Benchmarks | Quality |
|---|---|---|---|---|
| Qwen3.8 27B | Q4/Q8, images, AR, DFlash2 | Unsloth Q4_K_XL / Q8_K_XL · DFlash2 Q4_K_M | Q4: 656.33 tok/s pp; up to 70.56 tok/s tg single user and 123.00 aggregated tok/s on 8 concurrent requests with DFlash2 · Benchmarks | Quality |
| Qwen3.8 Flash-Next | Q4, images, AR, MTP | Unsloth Q4_K_XL · MTP Q8_0 | 1,700.52 tok/s pp; up to 60.39 tok/s tg single user and 162.98 aggregated tok/s on 8 concurrent requests with MTP · Benchmarks | Quality |
| DeepSeek V4 Flash | AR, DSpark | antirez Flash 0731 IQ2XXS · DSpark | 484.62 tok/s pp; up to 26.62 tok/s tg single user and 54.74 aggregated tok/s on 8 concurrent requests with DSpark · Benchmarks | Quality |
| Qwen3-ASR 1.7B | Speech recognition | BF16 | 15.27× realtime · Benchmarks | Quality |
| Qwen3-TTS 1.7B | Speech synthesis and voice cloning | BF16 CustomVoice / VoiceDesign / Base | Up to 2.54× realtime; 201 ms to first audio (CustomVoice) · Benchmarks | Quality |
| Qwen-Image-2.1 | BF16 image generation and editing | Complete pipeline | In progress · Benchmarks | Quality |
| MiniMax H3 | BF16 text to video/audio | FL2VA pipeline | In progress · Benchmarks | Quality |
Peak measured workloads; text pp is autoregressive (AR), while tg uses the named speculative mode. Peaks include repetitive output; aggregate tg sums individual request decode rates. Qwen27B's single-user peak uses the short-prompt C1 workload. Audio excludes loading. Each model guide lists the required files and complete benchmark settings.
- Contributions are welcome! We need the help of Strix Halo community to keep improving gufo!
- We would like this to be the one-stop shop for Strix Halo Local AI enthusiasts: batteries included for text, audio, image, and video models.
- Build and optimize specifically for the Strix Halo 128 GiB hardware. Smaller memory configurations should still work and preserve the speed benefits for models that can fit on memory.
- Support only the best available models for their size that can run on this hardware: less code to maintain, more focused optimization and testing work.
- Preserve quality when optimizing. Each model's quality report records independent numerical checks, execution consistency and unresolved gaps. Don't reuse kernels across different models to limit blast radius of a code change.
- Treat concurrent requests, cancellation and conversation caching as first-class workloads.
- Keep production dependencies small and development tools separate.
hf download unsloth/Qwen3.8-27B-GGUF \
Qwen3.8-27B-UD-Q8_K_XL.gguf \
--revision 4ca720788d1e01f1bff70c033e0d0028fd02e502 \
--repo-type model \
--local-dir models/Qwen3.8-27B-GGUF
hf download z-lab/Qwen3.8-27B-DFlash2-GGUF \
Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
--revision 2d9571f8ce46e151f61c6499c99dee6079e1d610 \
--repo-type model \
--local-dir models/Qwen3.8-27B-DFlash2-GGUF
podman pull ghcr.io/gufo-org/toolboxes/gufo-runtime:latest
podman run --rm \
--userns=keep-id:uid=1000,gid=1000 \
--device /dev/kfd \
--device /dev/dri \
--group-add keep-groups \
--ulimit memlock=-1 \
-p 8080:8080 \
-v ./models:/models:ro \
ghcr.io/gufo-org/toolboxes/gufo-runtime:latest \
gufo serve --host 0.0.0.0 --port 8080 llm \
--model /models/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q8_K_XL.gguf \
--speculative dflash2 \
--dflash-model /models/Qwen3.8-27B-DFlash2-GGUF/Qwen3.8-27B-DFlash2-Q4_K_M.ggufRootless Podman needs crun for --group-add keep-groups. Your host user must
have read/write access to /dev/kfd and /dev/dri/renderD*, usually through the
render and video groups; log out and back in after changing membership.
Container groups named video/render do not preserve host supplementary
groups. Check id, ls -l /dev/kfd /dev/dri/renderD*, and
podman info --format '{{.Host.OCIRuntime.Name}}' if ROCm reports no device.
See Podman's rootless group-access guidance.
On Fedora or another SELinux-enforcing host, GPU enumeration can succeed while
SELinux blocks mapping /dev/kfd, causing ROCr to report a misleading
“Memory critical” error. Check the host audit log:
sudo ausearch -m avc -ts recent | grep -E '/dev/kfd|hsa_device_t'If it shows a denied map for the container, Podman documents this fix:
sudo setsebool -P container_use_devices trueThis persistently allows containers to access device labels for devices passed into them; it affects all containers on that host. Review that policy scope before enabling it. See Podman's device documentation and the SELinux container policy.
Then, from another terminal, ask it something through the OpenAI-compatible API:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen3.8-27B-UD-Q8_K_XL",
"messages": [{"role": "user", "content": "Say something"}]
}'For Open WebUI, VS Code, OpenAI SDK and Responses-API clients (for example
Codex), set the API base URL to http://localhost:8080/v1. Chat Completions and
Responses support text, images, function tools, structured output and streaming.
See the API contract.
The text server uses the model's native context by default and generates until
EOS or the context is full. --context N sets context capacity per session;
--max-tokens N sets a default response limit that clients can override.
Reasoning tokens count toward that response limit.
- Qwen with Codex: Codex sends developer messages mid-conversation, after compaction or a settings change. Qwen's template accepts only one leading system turn, so Gufo moves them there. The request that introduces one is prefilled again; later requests reuse the cache. See the API contract.
Linux x86-64 on AMD Strix Halo (gfx1151) is the supported target. CMake owns
one production configuration for both Nix and ordinary Linux builds. Tests,
profilers, tuning executables and Python reference runners are not installed
with the production package. No build.sh wrapper is needed.
nix build
./result/bin/gufo diagnose
./result/bin/gufo serve llm --model /path/to/model.ggufflake.lock pins the dependencies. nix develop adds profiling,
model-download and independent evaluation tools; these are not runtime
requirements. Optional benchmark baselines are selected separately with
nix shell .#ds4-reference, .#llama-cpp-reference or
.#llama-cpp-mtp-reference; see benchmarking.
See testing for the small hosted CI suite and
explicit local quality checks.
Install a C++20 compiler, CMake 3.21+, Ninja, pkg-config and the following development libraries. The currently qualified toolchain is GCC 15.3 and ROCm 7.2.3. Attention and audio convolution kernels are compiled directly from HIP. Python, Triton/AOTriton, Composable Kernel and MIOpen are not production build or runtime requirements.
| Dependency | Used for |
|---|---|
| ROCm HIP compiler/runtime, hipBLAS, hipBLASLt, rocBLAS | GPU execution and matrix multiplication |
| hipCUB, rocPRIM, rocWMMA headers | Compiled GPU kernels |
| ICU, libcurl, OpenSSL, libpng, libjpeg, libwebp | Tokenization, HTTPS, hashing and images |
| FFmpeg and ffprobe | Video/audio output; invoked as separate executables |
Install ROCm using AMD's Linux instructions.
Use the development packages for the libraries above. ROCm normally installs
under /opt/rocm.
For example, on Debian/Ubuntu the ordinary system libraries are:
sudo apt install build-essential cmake ninja-build pkg-config \
libicu-dev libcurl4-openssl-dev libssl-dev libpng-dev libjpeg-dev libwebp-dev ffmpeg
# ROCm libraries from the table, named as AMD's repository ships them.
sudo apt install hipblas-dev hipblaslt-dev rocblas-dev \
hipcub-dev rocprim-dev rocwmma-dev
cmake --preset release -DCMAKE_INSTALL_PREFIX="$HOME/.local"
cmake --build --preset release --parallel 4
./build/release/gufo diagnose
./build/release/gufo serve llm --model /path/to/model.ggufConfiguring fails at find_package(hipblas) when those ROCm packages are
missing. Other distributions name them -devel instead of -dev.
For nonstandard installations, pass ordinary CMake paths, for example
cmake --preset release -DCMAKE_PREFIX_PATH="/opt/rocm".
If compiler discovery picks a system Clang, also pass
-DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++.
Use cmake --install build/release to install Gufo,
its runtime data and license notices. The GPU driver must allow your user to
access /dev/kfd and /dev/dri; model weights are acquired separately.
The same source, compiler flags and install rules serve both builds. Nix pins the complete toolchain for reproducible comparisons; changing the compiler or math libraries requires the affected model's quality checks.
Gufo's original code is MIT licensed. Adapted code and dependencies
retain their own notices in NOTICE, THIRD_PARTY_NOTICES.md
and licenses/, installed under share/licenses/gufo. Model weights are not
bundled and retain their publishers' terms.
The initial design is informed by the following open source projects:
- llama.cpp for compact model serving, GGUF, and CPU/GPU correctness paths.
- LaurentZuijdwijk/llama.cpp, Nathanw1014/strix-halo-llamacpp, and gaetan-puleo/llama-cpp-strix-halo for Strix Halo optimization inspiration.
- vLLM for continuous batching and paged request scheduling.
- hipEngine for torch-free HIP execution, and native speculative-cycle work.
- ds4 for DeepSeek V4 Flash, MoE scheduling, and DSpark.
- audio.cpp for audio models for tts and asr tasks.
