Skip to content

Qwen3.6 27B-MTP

Q8_0·27B params·GGUF
reasoning
checkpoint: unsloth/Qwen3.6-27B-MTP-GGUF:Q8_0

All runs (43)

legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=3codegen1
57.1
57.1192.4—383ms0.1—62100017.50s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=3chat1
55.9
55.9111.5—287ms0.1—301001.79s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=3agent1
55.0
55.01189.9—503ms0.1—5995009.10s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=3rag1
53.8
53.81285.1—857ms0.1—8422003.72s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=2codegen1
53.2
53.2205.8—317ms0.1—62100018.79s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=2agent1
51.2
51.21210.6—495ms0.1—5995009.77s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-3-pl-200wcodegen1
50.6
50.6193.1—380ms0.1—62100019.75s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-3-pl-200wchat1
50.5
50.5118.2—265ms0.1—301001.98s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-3-pl-200wagent1
49.9
49.9961.4—623ms0.1—59950010.01s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=2chat1
49.6
49.6120.8—258ms0.1—301002.02s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-2-pl-200wcodegen1
48.7
48.7190.1—390ms0.1—62100020.51s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=2rag1
47.8
47.81293.5—854ms0.1—8422004.18s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-2-pl-200wchat1
47.4
47.4115.1—275ms0.0—301002.11s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-2-pl-200wagent1
47.2
47.2927.1—646ms0.1—59950010.59s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-3-pl-200wrag1
45.1
45.1891.0—1.05s0.1—8422004.44s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-2-pl-200wrag1
42.5
42.5929.2—1.13s0.1—8422004.71s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)baselinechat1
28.1
25.7131.2—236ms35.6—301003.89s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)baselinecodegen1
28.0
27.0236.9—328ms35.7—62100037.04s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)baselineagent1
28.0
26.21160.3—516ms35.7—59950019.12s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)baselinerag1
28.0
24.61088.0—810ms35.7—8422008.13s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)baselineagent4
28.0
11.126.3—28.39s35.7—59950046.78s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)baseline-pl-200wchat1
27.1
25.1126.0—238ms37.0—301003.98s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)baseline-pl-200wcodegen1
26.9
26.1198.6—337ms37.2—62100038.31s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)baseline-pl-200wrag1
26.9
23.51243.4—911ms37.2—8422008.50s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)baseline-pl-200wagent1
26.8
25.3957.8—625ms37.4—59950019.73s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=3agent4
24.3
24.345.7—13.08s0.1—59950021.22s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB each450 W × 2 maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=2agent4
22.3
22.343.0—14.11s0.1—59950022.99s0.000 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=3agent1
18.7
18.71808.0—331ms0.0—59950026.81s0.023 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 590
llama.cpp 4f13cb7-mtp (cuda)baseline-pl-200wagent4
18.2
15.3——3.11s55.0——34122.31s0.000 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=3chat1
18.1
18.160.1—509ms0.0—301005.51s0.007 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=3codegen1
17.4
17.4127.7—486ms0.1—62100057.37s0.041 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=3rag1
17.1
17.1419.2—1.87s0.0—84220011.71s0.011 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=2agent1
16.1
16.12070.4—292ms0.0—59950031.11s0.025 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=2codegen1
15.7
15.7132.6—484ms0.1—62100063.67s0.044 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=2chat1
15.7
15.763.3—501ms0.0—301006.37s0.008 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=2rag1
15.4
15.4438.1—1.68s0.0—84220013.01s0.011 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=3agent4
7.9
7.919.1—39.55s0.0—59950065.22s0.088 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)baselinechat1
7.7
7.466.4—455ms129.5—3010013.44s0.003 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)baselinecodegen1
7.7
7.7142.8—434ms129.7—621000130.50s0.017 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)baselineagent1
7.7
7.72096.0—286ms129.7—59950065.13s0.010 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)baselineagent4
7.7
3.28.1—97.82s129.8—599500162.65s0.039 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)baselinerag1
7.7
7.3476.3—1.53s129.8—84220027.40s0.006 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=2agent4
6.8
6.815.8—44.75s0.0—59950075.79s0.096 GiB

Environment

2× GeForce RTX 3090 · 24 GiB each
cpuAMD EPYC 7302P 16-Core Processor
gpuNVIDIA GeForce RTX 3090 × 2
archNVIDIA
vram48 GiB (system 64.0 GiB)
power200 W × 2 / 450 W × 2 max(44% cap)
hardware probes
copy 42% of theoryFP16 peak 65.4 TFcopy/math flat across caps
384-bit9751 MHz82 SM/CU
Microbenchmarks for memory copy and tensor math; raw-engine decode and API workload rows measure model-serving speed.
captheorycopyfp16bf16
200 W936 GB/s391 GB/s65.4 TF65.4 TF
300 W936 GB/s391 GB/s65.4 TF65.3 TF
450 W936 GB/s391 GB/s65.4 TF65.4 TF
compute: 8.6
backendllama.cpp 4f13cb7-mtp (cuda)
osUbuntu 24.04 LTS
kernel6.17.13-7-pve
driver590.48.01
python3.12.3
runs/cell5
warmups2
endpoint/v1/chat/completions
streamingtrue
2× GeForce RTX 3090 · 24 GiB each
cpuAMD EPYC 7302P 16-Core Processor
gpuNVIDIA GeForce RTX 3090 × 2
archNVIDIA
vram48 GiB (system 64.0 GiB)
power450 W × 2 / 450 W × 2 max
pcieGen 4 x16 / Gen 4 x16 max
clocksgfx 1800/2100 MHz · mem 9501 MHz
temp60°C idle · 69°C peak
peak draw294 W
backendllama.cpp cuda-4f13cb7 (cuda)
osUbuntu 24.04 LTS
kernel6.17.13-7-pve
driverNVIDIA 590.48.01 + CUDA 13.1
libc2.39
python3.12.3
llama.cppversion: 18 (4f13cb7) built with GNU 13.3.0 for Linux x86_64
build flagsGGML_CUDA=ON CMAKE_BUILD_TYPE=Release
runs/cell5
warmups2
endpoint/v1/chat/completions
streamingtrue
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)
cpuAMD RYZEN AI MAX+ 395 w/ Radeon 8060S
gpuAMD Radeon 8060S
archStrix Halo (gfx1151)
vram96 GiB (system 31.1 GiB, unified)
hardware probes
copy 41% of theoryFP16 peak 30.3 TF
256-bit8000 MHz20 SM/CU
Microbenchmarks for memory copy and tensor math; raw-engine decode and API workload rows measure model-serving speed.
captheorycopyfp16bf16
fixed256 GB/s106 GB/s30.3 TF-
compute: 11.5
backendllama.cpp 4f13cb7-mtp (rocm)
osUbuntu 24.04 LTS
kernel7.0.2-2-pve
python3.12.3
runs/cell5
warmups2
endpoint/v1/chat/completions
streamingtrue