Qwen3.6 27B-MTP
Q4_K_M·27B params·GGUF
reasoning
intelligence: see on Artificial Analysis →
checkpoint:
unsloth/Qwen3.6-27B-MTP-GGUF:Q4_K_MAll runs (43)
| Hardware | Backend | Mode | Shape | Conc. | Gen tok/s ↓ | Prefill tok/s | TTFT | TPOT (ms) | Prompt tok | Out tok | Total | VRAM Δ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | MTP n=2 | codegen | 1 | 63.7 | 195.4 | 335ms | 0.1 | 62 | 1000 | 15.69s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | MTP n=2 | agent | 1 | 61.7 | 1157.5 | 525ms | 0.1 | 599 | 500 | 8.11s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | MTP n=2 | chat | 1 | 59.5 | 119.7 | 259ms | 0.1 | 30 | 100 | 1.68s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | MTP n=3 | codegen | 1 | 59.2 | 175.8 | 372ms | 0.1 | 62 | 1000 | 16.88s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | MTP n=3 | chat | 1 | 58.7 | 123.8 | 259ms | 0.1 | 30 | 100 | 1.70s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | MTP n=3 | agent | 1 | 57.8 | 1208.4 | 537ms | 0.1 | 599 | 500 | 8.66s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | MTP n=2 | rag | 1 | 55.9 | 1186.1 | 879ms | 0.1 | 842 | 200 | 3.58s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | MTP n=3 | rag | 1 | 54.6 | 1096.3 | 934ms | 0.1 | 842 | 200 | 3.66s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | baseline | codegen | 1 | 40.4 | 188.9 | 355ms | 23.7 | 62 | 1000 | 24.76s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | baseline | agent | 1 | 39.3 | 1196.7 | 505ms | 23.8 | 599 | 500 | 12.73s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | baseline | chat | 1 | 38.7 | 127.0 | 238ms | 23.4 | 30 | 100 | 2.58s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | baseline | rag | 1 | 35.6 | 1303.0 | 738ms | 23.6 | 842 | 200 | 5.62s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | mtp-3-pl-200w | chat | 1 | 34.2 | 109.6 | 283ms | 0.1 | 30 | 100 | 2.92s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | mtp-2-pl-200w | chat | 1 | 32.0 | 113.4 | 271ms | 0.1 | 30 | 100 | 3.12s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | mtp-2-pl-200w | codegen | 1 | 31.8 | 185.9 | 377ms | 0.1 | 62 | 1000 | 31.41s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | mtp-3-pl-200w | rag | 1 | 31.2 | 856.1 | 1.05s | 0.1 | 842 | 200 | 6.40s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | mtp-3-pl-200w | agent | 1 | 31.2 | 964.4 | 621ms | 0.1 | 599 | 500 | 16.03s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | mtp-3-pl-200w | codegen | 1 | 31.1 | 182.4 | 384ms | 0.1 | 62 | 1000 | 32.18s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | mtp-2-pl-200w | agent | 1 | 30.4 | 972.1 | 616ms | 0.0 | 599 | 500 | 16.44s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | mtp-2-pl-200w | rag | 1 | 29.8 | 860.6 | 1.05s | 0.0 | 842 | 200 | 6.71s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | MTP n=2 | agent | 4 | 25.3 | 57.3 | 12.42s | 0.1 | 599 | 500 | 20.39s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | MTP n=3 | agent | 4 | 24.0 | 59.7 | 13.43s | 0.1 | 599 | 500 | 21.54s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | baseline-pl-200w | codegen | 1 | 21.5 | 191.7 | 345ms | 45.9 | 62 | 1000 | 46.44s | 0.000 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | MTP n=3 | chat | 1 | 21.2 | 78.2 | 386ms | 3.0 | 30 | 100 | 4.72s | 0.006 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | baseline-pl-200w | chat | 1 | 21.1 | 112.9 | 266ms | 43.9 | 30 | 100 | 4.74s | 0.000 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | baseline-pl-200w | agent | 1 | 21.0 | 984.2 | 609ms | 46.2 | 599 | 500 | 23.77s | 0.000 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | MTP n=3 | codegen | 1 | 20.6 | 133.5 | 475ms | 0.1 | 62 | 1000 | 48.47s | 0.041 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | MTP n=3 | agent | 1 | 20.4 | 1984.4 | 302ms | 0.0 | 599 | 500 | 24.51s | 0.022 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | MTP n=2 | codegen | 1 | 20.0 | 136.7 | 460ms | 0.1 | 62 | 1000 | 49.96s | 0.044 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | baseline-pl-200w | rag | 1 | 20.0 | 1035.7 | 923ms | 44.9 | 842 | 200 | 10.01s | 0.000 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | MTP n=3 | rag | 1 | 19.9 | 426.2 | 1.69s | 0.0 | 842 | 200 | 10.05s | 0.011 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | MTP n=2 | chat | 1 | 19.7 | 81.6 | 375ms | 0.0 | 30 | 100 | 5.07s | 0.006 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | MTP n=2 | agent | 1 | 19.4 | 1949.4 | 315ms | 0.0 | 599 | 500 | 25.73s | 0.024 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | MTP n=2 | rag | 1 | 18.8 | 432.5 | 1.68s | 0.0 | 842 | 200 | 10.66s | 0.012 GiB |
GeForce RTX 3090 · 24 GiB450 Wdrv 590 | llama.cpp cuda-4f13cb7 (cuda) | baseline | agent | 4 | 16.2 | 35.9 | 19.73s | 23.8 | 599 | 500 | 32.01s | 0.000 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | baseline | codegen | 1 | 12.0 | 146.7 | 428ms | 82.7 | 62 | 1000 | 83.17s | 0.016 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | baseline | agent | 1 | 12.0 | 2090.1 | 287ms | 82.8 | 599 | 500 | 41.69s | 0.009 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | baseline | chat | 1 | 11.7 | 87.3 | 345ms | 82.4 | 30 | 100 | 8.52s | 0.003 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | baseline | rag | 1 | 11.1 | 470.6 | 1.51s | 82.8 | 842 | 200 | 18.01s | 0.006 GiB |
GeForce RTX 3090 · 24 GiB200 Wdrv 590 | llama.cpp 4f13cb7-mtp (cuda) | baseline-pl-200w | agent | 4 | 9.8 | — | 4.04s | 89.5 | — | 341 | 34.74s | 0.040 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | MTP n=3 | agent | 4 | 8.4 | 18.8 | 36.60s | 0.1 | 599 | 500 | 61.49s | 0.090 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | MTP n=2 | agent | 4 | 8.2 | 21.9 | 38.32s | 0.1 | 599 | 500 | 63.14s | 0.097 GiB |
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified | llama.cpp 4f13cb7-mtp (rocm) | baseline | agent | 4 | 5.0 | 12.6 | 62.69s | 82.8 | 599 | 500 | 104.08s | 0.037 GiB |
Environment
GeForce RTX 3090 · 24 GiB
cpuAMD EPYC 7302P 16-Core Processor
gpuNVIDIA GeForce RTX 3090
archNVIDIA
vram24 GiB (system 64.0 GiB)
power200 W / 450 W max(44% cap)
backendllama.cpp 4f13cb7-mtp (cuda)
serverlemonade unknown
osUbuntu 24.04 LTS
kernel6.17.13-7-pve
driver590.48.01
python3.12.3
containerizedtrue
runs/cell5
warmups2
endpoint/v1/chat/completions
streamingtrue
GeForce RTX 3090 · 24 GiB
cpuAMD EPYC 7302P 16-Core Processor
gpuNVIDIA GeForce RTX 3090
archNVIDIA
vram24 GiB (system 64.0 GiB)
power450 W / 450 W max
pcieGen 4 x16 / Gen 4 x16 max
clocksgfx 1980/2100 MHz · mem 9501 MHz
temp38°C idle · 83°C peak
peak draw436 W
backendllama.cpp cuda-4f13cb7 (cuda)
serverlemonade unknown
osUbuntu 24.04 LTS
kernel6.17.13-7-pve
driverNVIDIA 590.48.01 + CUDA 13.1
libc2.39
python3.12.3
containerizedtrue
llama.cppversion: 18 (4f13cb7) built with GNU 13.3.0 for Linux x86_64
build flagsGGML_CUDA=ON CMAKE_BUILD_TYPE=Release
runs/cell5
warmups2
endpoint/v1/chat/completions
streamingtrue
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)
cpuAMD RYZEN AI MAX+ 395 w/ Radeon 8060S
gpuAMD Radeon 8060S
archStrix Halo (gfx1151)
vram96 GiB (system 31.1 GiB, unified)
backendllama.cpp 4f13cb7-mtp (rocm)
serverlemonade unknown
osUbuntu 24.04 LTS
kernel7.0.2-2-pve
python3.12.3
containerizedtrue
runs/cell5
warmups2
endpoint/v1/chat/completions
streamingtrue