The Bosgame M5 mini PC demonstrates significantly faster AI model inference performance when running Ubuntu compared to Windows 11. Testing with the Qwen3.8 27B Q8 model revealed that Ubuntu, utilizing the Vulkan backend, completed long-prompt responses in 25.27 seconds, a 33% improvement over the 37.94 seconds required by Windows 11. For short prompts, Ubuntu achieved a first-text latency of 0.3 seconds, compared to 2.6 seconds on Windows. While the ASUS GX10 remains the faster overall system, the M5’s performance on Ubuntu is notably more efficient than on Windows. Within the Ubuntu environment, the Vulkan backend consistently outperformed the HIP/ROCm backend across prompt processing, token generation rates, and total response time. The system maintained stable thermal performance during testing, with CPU and GPU peaks reaching 84°C and 85°C respectively at 28°C ambient temperatures, well below the 95°C safety threshold. The configuration utilized 96 GB of the system's 128 GB shared memory for graphics, ensuring sufficient overhead for large models without resorting to memory swapping.
Ubuntu reduces long-prompt inference time by 33% compared to Windows 11 on the Bosgame M5. The Vulkan backend provides faster inference speeds than HIP/ROCm on the Ubuntu platform.
Short-prompt latency is significantly lower on Ubuntu, with first-text output appearing in 0.3 seconds. The Bosgame M5 maintains safe operating temperatures under heavy AI workloads at 28°C ambient.
Allocating 96 GB of shared memory to the GPU allows the system to run large models without swapping.
Chapter guide
Worth noting
- The testing methodology is limited to a single specific AI model (Qwen3.8 27B Q8) and may not reflect performance across all AI applications.
- The results are based on a specific BIOS memory allocation (96 GB for graphics) which may not be optimal for all user scenarios.