Qwen 3.8-27B: Benchmarks, Versions, and Serving Strategies
Discover the power of Qwen 3.8-27B! Learn which model version to use and how to serve it fast for local AI. Benchmarks, speed tests, and optimization tips.
The recent release of Qwen 3.8 Max, with its 2.4 trillion parameters, highlights the rapid advancements in large language models. However, its substantial size limits local deployment for most users. The more accessible Qwen 3.8-27B, released on Friday, offers a compelling alternative, balancing impressive capabilities with practical usability for local AI applications. This article examines the performance of Qwen 3.8-27B, compares it to previous models, and explores optimal deployment strategies for achieving efficient local execution.
Performance Benchmarks and Comparisons
Qwen 3.8-27B demonstrates a significant improvement over its predecessor, Qwen 3.6-27B, across various benchmarks. This advancement is particularly notable when compared to Meta's Muse Glimmer model. While Muse Glimmer, trained for broader generalization, remains a viable option for specific use cases, Qwen 3.8-27B consistently outperforms it in benchmark tests.
Recent intelligence index scores from Artificial Analysis reveal Qwen 3.8-27B achieving a score of 52. This places it competitively close to models like GLM 5.2 (53) and DeepSeek V4 Pro, and substantially ahead of Qwen 3.6-27B and other open models requiring significant hardware investment. Furthermore, in agentic index tests, Qwen 3.8-27B surpasses GLM 5.2 and some GPT 5.6 models, offering remarkable performance for a locally runnable model with decent token speeds.
The model also shows enhanced vision performance, with capabilities in computer use and browser interaction benchmarks exceeding even Opus 4.6 Max in some instances. This suggests a strong focus on multimodal functionalities in the Qwen 3.8 series.
Model Variants and Quantization
The Qwen 3.8 family includes several variants, notably the 2.4 trillion parameter models and the 27 billion parameter models. For the 27B version, Qwen provides both a bfloat16 (full resolution) and an FP8 version. Beyond these official releases, the open-source community has developed further optimizations.
Unsloth has released 4-bit quantized versions utilizing NVFP4 quantization, which can significantly reduce memory requirements. However, compatibility with specific GPU hardware is a consideration for these advanced quantizations.
Additionally, numerous fine-tuned versions are emerging. These include models trained on specialized datasets for improved results and "uncensored" versions, achieved through fine-tuning or obliteration techniques. For users on macOS, the MLX community has developed multiple Qwen 3.8 versions, including bfloat16, 8-bit, and 4-bit variants, as well as MLX-specific quantizations.
Optimizing Reasoning and Token Usage
A critical factor in achieving optimal performance with Qwen 3.8-27B is the configuration of its reasoning capabilities, which directly impacts token consumption. Testing reveals distinct differences in output and token usage across various reasoning settings:
- No Thinking: This setting bypasses reasoning entirely, leading to rapid generation but potentially less sophisticated output. For instance, in an HTML website generation task, this mode produced a functional website with minimal thinking tokens.
- Low Thinking: This setting introduces reasoning, resulting in more refined outputs. In the HTML example, it generated a more elaborate website using approximately 512 thinking tokens.
- Medium Thinking: This setting appears to offer a balance, though in some specific tasks, it may consume fewer tokens than the "low" setting. For the HTML task, it produced a comparable output to the "low" setting, with variations in token usage observed across multiple runs.
- X-High Thinking: This setting engages extensive reasoning, leading to highly detailed, albeit token-intensive, outputs. In the HTML example, this mode consumed over 17,500 tokens, exceeding a 32K token limit and preventing completion. The model engages in extensive self-dialogue, significantly increasing processing time. This behavior is consistent across FP8, 16-bit, and Unsloth quantized versions.
Similar patterns were observed in an SVG Pelican generation test. "X-High" reasoning used approximately 11,000 tokens for a decent output. "Medium" reasoning reduced this to under 1,000 tokens while still producing a good result. "Low" reasoning used more tokens than "medium" and missed some details, while "No Thinking" resulted in a rudimentary output.
For a red dragon SVG test, "X-High" reasoning consumed 21,000 tokens out of 35,000 total, producing a passable but not exceptional result. "Medium" reasoning, using 2,000 tokens, showed some regression on the dragon compared to the bicycle. "Low" reasoning used 420 tokens. The "sweet spot" for this user appears to be "medium" reasoning, balancing quality and efficiency.
Inference Libraries and Speed Optimization
The choice of inference library and configuration significantly impacts token generation speed. Testing was conducted on a Dell T2 Pro Max with an Nvidia RTX Pro 6000 GPU, featuring 96GB of VRAM.
- bfloat16 Version: Running the original bfloat16 version yielded approximately 30 tokens per second (t/s) for pure generation without speculative decoding.
- Speculative Decoding: Enabling speculative decoding, particularly with Multi-Token Prediction (MTP) set to 3, provided a substantial speed increase.
- FP8 Version: Using Qwen's FP8 version with speculative decoding achieved speeds of 80-120 t/s, with minimal perceived quality difference from the bfloat16 version.
- Obliterated Versions: While offering expanded capabilities, these versions were found to frequently enter repetitive loops, making them less practical for general use at this time.
- Unsloth Model: The Unsloth quantized model delivered good token rates, reaching around 120 t/s with MTP.
- SGLang with NVFP4: Utilizing SGLang's serving framework, specifically configured for NVFP4 quantization and DSpark speculative decoding, yielded the highest speeds. This setup, running within a Docker container using SGLang's provided image, achieved speeds of up to 206 t/s. With their specific model weights, average speeds of 173 t/s were recorded, and for shorter sequences, speeds between 200-220 t/s were observed. Even with extensive token usage (150-170 t/s average for 36,000 tokens), this configuration proved exceptionally fast. The full 262K context window was utilized effectively.
For users with lower VRAM GPUs, Llama CPP is recommended as an alternative inference library.
Conclusion and Future Outlook
Qwen 3.8-27B represents a significant leap forward for locally deployable AI models. Achieving optimal performance requires careful consideration of model quantization, reasoning token configuration, and the selection of an appropriate inference library. SGLang, when paired with specific NVFP4 quantizations, currently offers the fastest inference speeds.
The ongoing development of fine-tuned versions, such as potential "ThinkingCap" adaptations that optimize token usage for intelligence, and advancements in fused kernels for faster serving, suggest that Qwen 3.8-27B's capabilities will continue to expand. For users seeking to run powerful AI models on prosumer hardware, Qwen 3.8-27B is currently the model to beat. Experimentation with different configurations and inference libraries is crucial to identify the best setup for individual use cases.
Introduction to Qwen 3.8-27B
Introduces the Qwen 3.8-27B model, highlighting its impressive capabilities while acknowledging the challenge of its 2.4 trillion parameter predecessor. It sets the stage for discussing which version of the 27B model is best for local use and how to serve it efficiently.
- Qwen released Qwen 3.8 Max (2.4 trillion parameters), which is powerful but difficult to run locally.
- The Qwen 3.8-27B model, released on Friday, is the focus for local deployment.
- The video will cover benchmarks, model versions, and serving strategies for Qwen 3.8-27B.
Performance Benchmarks and Comparisons
Compares Qwen 3.8-27B to previous models like Qwen 3.6-27B and Meta's Glimmer 30B, noting that Qwen 3.8-27B significantly outperforms them in benchmarks. It also touches on the potential rush behind Meta's Glimmer release.
- Qwen 3.6-27B was popular for local coding and agents.
- Meta's Glimmer 30B was benchmarked against Qwen 3.6-27B.
- Qwen 3.8-27B substantially outperforms Meta's Muse Glimmer and Qwen 3.6.
- Meta may have rushed Glimmer's release knowing Qwen 3.8 was coming.
Exploring Qwen 3.8-27B Model Variants
Discusses the various versions of Qwen 3.8-27B available, including different parameter counts, quantizations (bfloat16, FP8, 4-bit), fine-tunes, and versions for specific hardware like Macs (MLX). The importance of choosing the right version and configuration for local systems is emphasized.
- Qwen 3.8 family includes 2.4 trillion parameter models and 27B models.
- Available versions include bfloat16 (full resolution) and FP8.
- 4-bit versions (e.g., Unsloth's NVFP4) are available but require compatible GPUs.
- Numerous fine-tuned and uncensored versions exist.
- MLX versions are available for Mac users.
The Role of Reasoning Tokens
Analyzes the impact of 'thinking' or reasoning token settings on model output quality and token consumption. Demonstrates how different levels of thinking (none, low, medium, x-high) affect the generation of websites and SVGs, showing that 'medium' often strikes a good balance.
- Reasoning token settings significantly impact output quality and token usage.
- No thinking results in basic, often uninspired output.
- Low and medium thinking produce better results with moderate token usage.
- X-high thinking consumes a vast number of tokens, leading to excessive detail or incomplete generations.
- The optimal thinking level varies by task; 'medium' is often a sweet spot.
Optimizing Serving Speed and Inference Libraries
Evaluates the serving speed of different Qwen 3.8-27B versions using libraries like vLLM and SGLang. Compares bfloat16, FP8, and Unsloth quantizations, ultimately finding SGLang with a specific NVFP4 model to provide the fastest token generation speeds (up to 200+ tokens/sec).
- Basic bfloat16 version achieved ~30 tokens/sec without speculative decoding.
- Speculative decoding (MTP=3) significantly improves speed.
- FP8 version with speculative decoding reached 80-120 tokens/sec.
- Unsloth quantization achieved ~120 tokens/sec.
- SGLang with a specific NVFP4 model and DSpark speculative decoding achieved the highest speeds, averaging 150-173 tokens/sec, with peaks over 200 tokens/sec.
- Abliterated versions were found to get stuck in loops.
- vLLM and SGLang were tested; SGLang was the winner for high-speed serving.
