3 ms·
FYI: nothing seems to be able to run this (easily) yet. llama.cpp, vllm etc I couldn't get working because of no support in the mainline version.
by martinald 1mo ago
FYI: nothing seems to be able to run this (easily) yet. llama.cpp, vllm etc I couldn't get working because of no support in the mainline version.
- kzrdude 1mo agoThey are giving pointers to how to run it now using for example https://recipes.vllm.ai/Qwen/Qwen3.8-Flash-Next https://recipes.vllm.ai/Qwen/Qwen3.8-Flash-Next (and an especially provided vllm release).
- martinald 1mo agoah I don't think that page was up when I checked, it 404ed!
- a_humean 1mo agoProbably going to take a 1-3 days for support to land in llama.cpp and vllm.
- flounder3 1mo agoRelevant PR: https://github.com/ggml-org/llama.cpp/pull/27742 https://github.com/ggml-org/llama.cpp/pull/27742 This branch works now: https://github.com/unslothai/llama.cpp/tree/qwen4exp/qwen3.8-flash-next https://github.com/unslothai/llama.cpp/tree/qwen4exp/qwen3.8... cmake -B build -DGGML_CUDA=ON or cmake -B build -DGGML_METAL=ON then cmake --build build --config Release -j --target llama-server llama-cli