Skip to content

Qwen3.6 27B-MTP

Q4_K_M·27B params·GGUF
reasoning
checkpoint: unsloth/Qwen3.6-27B-MTP-GGUF:Q4_K_M

All runs (127)

legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=2codegen1
63.7
63.7195.4—335ms0.1—62100015.69s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=2agent1
61.7
61.71157.5—525ms0.1—5995008.11s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=2chat1
59.5
59.5119.7—259ms0.1—301001.68s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=3codegen1
59.2
59.2175.8—372ms0.1—62100016.88s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=3chat1
58.7
58.7123.8—259ms0.1—301001.70s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=3agent1
57.8
57.81208.4—537ms0.1—5995008.66s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=2rag1
55.9
55.91186.1—879ms0.1—8422003.58s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=3rag1
54.6
54.61096.3—934ms0.1—8422003.66s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx1k_answer1
50.2
50.2656.8—1.48s0.1—9695009.96s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx1k_answer1
47.5
47.5671.1—1.45s0.1—96950010.54s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)baselinechat1
42.7
38.7127.0—238ms23.4—301002.58s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)baselinerag1
42.3
35.61303.0—738ms23.6—8422005.62s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)baselinecodegen1
42.2
40.4188.9—355ms23.7—62100024.76s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)baselineagent1
42.1
39.31196.7—505ms23.8—59950012.73s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)baselineagent4
41.9
16.235.9—19.73s23.8—59950032.01s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx1k_probe1
40.1
6.3870.2—1.09s24.9—95181.27s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx1k_answer1
39.5
35.2824.9—1.18s25.3—96950014.21s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx4k_answer1
39.1
30.81238.8—3.05s25.5—377850016.22s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx4k_answer1
38.0
38.0827.9—4.57s0.1—377850013.15s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx16k_probe1
37.8
0.71350.7—11.21s26.5—15136811.40s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx16k_answer1
37.6
20.21344.6—11.27s26.6—1515450024.81s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx4k_answer1
36.7
36.7823.7—4.59s0.1—377850013.62s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx4k_probe1
36.5
2.41256.4—3.05s27.4—383583.29s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx32k_probe1
36.1
0.31269.9—23.85s27.7—30287824.05s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx32k_answer1
35.7
13.11269.6—23.87s28.0—3030550038.03s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-3-pl-200wchat1
34.2
34.2109.6—283ms0.1—301002.92s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx64k_answer1
32.8
7.11107.2—54.68s30.5—6060150070.34s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-2-pl-200wchat1
32.0
32.0113.4—271ms0.1—301003.12s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx64k_probe1
32.0
0.11098.8—55.14s31.2—60585855.37s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-2-pl-200wcodegen1
31.8
31.8185.9—377ms0.1—62100031.41s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-3-pl-200wrag1
31.2
31.2856.1—1.05s0.1—8422006.40s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-3-pl-200wagent1
31.2
31.2964.4—621ms0.1—59950016.03s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-3-pl-200wcodegen1
31.1
31.1182.4—384ms0.1—62100032.18s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-2-pl-200wagent1
30.4
30.4972.1—616ms0.0—59950016.44s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx100k_probe1
29.9
0.1939.3—100.77s33.4—946458101.01s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)mtp-2-pl-200wrag1
29.8
29.8860.6—1.05s0.0—8422006.71s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx100k_answer1
29.1
4.2937.4—100.91s34.4—94590500118.33s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx128k_answer1
27.4
3.1852.8—142.02s36.5—121117500160.72s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)baselinectx128k_probe1
27.2
0.1856.6—141.38s36.8—1210998141.66s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=2agent4
25.3
25.357.3—12.42s0.1—59950020.39s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiB450 W maxdrv 590
llama.cpp cuda-4f13cb7 (cuda)MTP n=3agent4
24.0
24.059.7—13.43s0.1—59950021.54s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)baseline-pl-200wchat1
22.8
21.1112.9—266ms43.9—301004.74s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)baseline-pl-200wrag1
22.3
20.01035.7—923ms44.9—84220010.01s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)baseline-pl-200wcodegen1
21.8
21.5191.7—345ms45.9—62100046.44s0.000 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)baseline-pl-200wagent1
21.6
21.0984.2—609ms46.2—59950023.77s0.000 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx1k_answer1
21.3
21.3286.6—3.38s0.0—96950023.50s0.013 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=3chat1
21.2
21.278.2—386ms3.0—301004.72s0.006 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=3codegen1
20.6
20.6133.5—475ms0.1—62100048.47s0.041 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=3agent1
20.4
20.41984.4—302ms0.0—59950024.51s0.022 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=2codegen1
20.0
20.0136.7—460ms0.1—62100049.96s0.044 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=3rag1
19.9
19.9426.2—1.69s0.0—84220010.05s0.011 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=2chat1
19.7
19.781.6—375ms0.0—301005.07s0.006 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=2agent1
19.4
19.41949.4—315ms0.0—59950025.73s0.024 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx1k_answer1
19.3
19.3287.4—3.37s0.0—96950025.91s0.013 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=2rag1
18.8
18.8432.5—1.68s0.0—84220010.66s0.012 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx16k_answer1
17.2
17.2791.3—19.15s0.1—1515450029.01s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx16k_answer1
17.2
17.2790.1—19.18s0.1—1515450029.05s0.000 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx4k_answer1
15.3
15.3303.8—12.44s0.0—377850032.64s0.012 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx4k_answer1
14.1
14.1302.7—12.48s0.0—377850035.35s0.014 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)baselinechat1
12.1
11.787.3—345ms82.4—301008.52s0.003 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)baselinecodegen1
12.1
12.0146.7—428ms82.7—62100083.17s0.016 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)baselineagent1
12.1
12.02090.1—287ms82.8—59950041.69s0.009 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx1k_probe1
12.1
2.3336.1—2.83s82.8—95183.42s0.002 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)baselineagent4
12.1
5.012.6—62.69s82.8—599500104.08s0.037 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx1k_answer1
12.1
11.3319.1—3.04s82.8—96950044.42s0.005 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)baselinerag1
12.1
11.1470.6—1.51s82.8—84220018.01s0.006 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx4k_answer1
11.9
9.4339.2—11.14s83.7—377850053.00s0.006 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx4k_probe1
11.7
0.7340.5—11.26s85.8—3835811.89s0.004 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx16k_answer1
11.5
5.4311.1—48.72s86.7—1515450092.05s0.006 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx16k_probe1
11.5
0.2312.4—48.44s87.2—15136849.07s0.002 GiB
legacystack comparable
GeForce RTX 3090 · 24 GiBcap 200 Wdrv 590
llama.cpp 4f13cb7-mtp (cuda)baseline-pl-200wagent4
11.2
9.8——4.04s89.5——34134.74s0.040 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx32k_answer1
11.0
3.2276.3—109.69s90.9—30305500155.13s0.005 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx32k_probe1
11.0
0.1276.3—109.63s91.1—302878110.28s0.001 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx64k_answer1
10.1
1.6224.9—269.20s99.3—60601500318.82s0.005 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx64k_probe1
10.0
0.0224.8—269.46s99.5—605858270.17s0.002 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx32k_answer1
9.9
9.9737.1—41.12s0.1—3030550050.70s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx32k_answer1
9.8
9.8739.3—40.99s0.1—3030550051.28s0.000 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx100k_answer1
9.2
0.9185.8—509.02s108.6—94590500563.37s0.006 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx100k_probe1
9.2
0.0185.8—509.52s108.8—946458510.31s0.002 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx128k_answer1
8.6
0.6163.8—739.48s115.9—121117500797.46s0.005 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)baselinectx128k_probe1
8.6
0.0163.6—740.26s115.9—1210998741.09s0.001 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=3agent4
8.4
8.418.8—36.60s0.1—59950061.49s0.090 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unified
llama.cpp 4f13cb7-mtp (rocm)MTP n=2agent4
8.2
8.221.9—38.32s0.1—59950063.14s0.097 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx16k_answer1
6.4
6.4280.1—54.12s0.0—1515450077.75s0.013 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx16k_answer1
6.3
6.3280.2—54.09s0.0—1515450079.44s0.014 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx1k_probe1
5.4
5.4693.4—1.37s0.1—95181.49s0.010 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx1k_probe1
5.3
5.3694.6—1.37s0.1—95181.50s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx64k_answer1
4.7
4.7631.4—95.98s0.1—60601500105.45s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx64k_answer1
4.7
4.7632.7—95.67s0.1—60601500105.88s0.000 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx32k_answer1
3.4
3.4250.0—121.24s0.0—30305500147.14s0.013 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx32k_answer1
3.4
3.4249.5—121.48s0.1—30305500148.56s0.015 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx100k_answer1
2.7
2.7536.0—176.47s0.1—94590500187.99s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx100k_answer1
2.7
2.7535.0—176.82s0.1—94590500188.56s0.000 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx1k_probe1
2.3
2.3299.4—3.18s0.0—95183.50s0.003 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx1k_probe1
2.3
2.3299.5—3.17s0.0—95183.55s0.003 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx128k_answer1
1.9
1.9478.8—252.98s0.1—121117500263.24s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx128k_answer1
1.9
1.9476.8—254.00s0.1—121117500265.22s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx4k_probe1
1.7
1.7833.5—4.60s0.1—383584.72s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx4k_probe1
1.7
1.7821.7—4.67s0.1—383584.81s0.000 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx64k_answer1
1.5
1.5204.8—295.96s0.0—60601500322.74s0.010 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx64k_answer1
1.5
1.5204.7—296.05s0.0—60601500326.27s0.009 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx100k_answer1
0.9
0.9170.1—556.16s0.0—94590500585.80s0.010 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx100k_answer1
0.8
0.8169.4—558.28s0.0—94590500590.24s0.010 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx4k_probe1
0.6
0.6302.9—12.66s0.0—3835812.98s0.002 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx4k_probe1
0.6
0.6304.1—12.61s3.6—3835813.02s0.001 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx128k_answer1
0.6
0.6150.2—806.44s0.0—121117500837.93s0.010 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx128k_answer1
0.6
0.6149.5—810.10s0.0—121117500844.87s0.010 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx16k_probe1
0.4
0.4795.0—19.04s0.1—15136819.16s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx16k_probe1
0.4
0.4793.2—19.08s0.0—15136819.24s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx32k_probe1
0.2
0.2739.3—40.97s0.1—30287841.10s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx32k_probe1
0.2
0.2738.5—41.01s0.0—30287841.17s0.000 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx16k_probe1
0.1
0.1281.4—53.80s0.1—15136854.24s0.002 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx16k_probe1
0.1
0.1279.9—54.08s0.0—15136854.42s0.003 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx64k_probe1
0.1
0.1635.1—95.39s0.1—60585895.58s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx64k_probe1
0.1
0.1633.4—95.65s0.0—60585895.82s0.000 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx32k_probe1
0.1
0.1250.1—121.13s0.0—302878121.58s0.003 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx32k_probe1
0.1
0.1249.8—121.26s0.0—302878121.61s0.002 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx100k_probe1
0.0
0.0537.7—176.03s0.1—946458176.22s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx100k_probe1
0.0
0.0536.4—176.45s0.0—946458176.66s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=2ctx128k_probe1
0.0
0.0477.9—253.42s0.1—1210998253.61s0.000 GiB
legacystack comparable
2× GeForce RTX 3090 · 24 GiB eachcap 200 W × 2drv 595
llama.cpp direct (cuda)MTP n=3ctx128k_probe1
0.0
0.0479.0—252.82s0.0—1210998252.96s0.000 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx64k_probe1
0.0
0.0204.8—295.80s0.0—605858296.33s0.003 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx64k_probe1
0.0
0.0204.8—295.90s0.0—605858296.29s0.003 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx100k_probe1
0.0
0.0169.7—557.77s0.0—946458558.32s0.002 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx100k_probe1
0.0
0.0169.9—557.01s0.0—946458557.61s0.003 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=2ctx128k_probe1
0.0
0.0149.9—807.88s0.0—1210998808.42s0.003 GiB
legacystack comparable
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)unifieddrv 7
llama.cpp direct (rocm)MTP n=3ctx128k_probe1
0.0
0.0150.2—806.45s0.0—1210998806.93s0.002 GiB

Environment

GeForce RTX 3090 · 24 GiB
cpuAMD EPYC 7302P 16-Core Processor
gpuNVIDIA GeForce RTX 3090
archNVIDIA
vram24 GiB (system 64.0 GiB)
power200 W / 450 W max(44% cap)
hardware probes
copy 42% of theoryFP16 peak 65.4 TFcopy/math flat across caps
384-bit9751 MHz82 SM/CU
Microbenchmarks for memory copy and tensor math; raw-engine decode and API workload rows measure model-serving speed.
captheorycopyfp16bf16
200 W936 GB/s391 GB/s65.4 TF65.4 TF
300 W936 GB/s391 GB/s65.4 TF65.3 TF
450 W936 GB/s391 GB/s65.4 TF65.4 TF
compute: 8.6
backendllama.cpp 4f13cb7-mtp (cuda)
osUbuntu 24.04 LTS
kernel6.17.13-7-pve
driver590.48.01
python3.12.3
runs/cell5
warmups2
endpoint/v1/chat/completions
streamingtrue
GeForce RTX 3090 · 24 GiB
cpuAMD EPYC 7302P 16-Core Processor
gpuNVIDIA GeForce RTX 3090
archNVIDIA
vram24 GiB (system 64.0 GiB)
power450 W / 450 W max
pcieGen 4 x16 / Gen 4 x16 max
clocksgfx 1980/2100 MHz · mem 9501 MHz
temp38°C idle · 83°C peak
peak draw436 W
backendllama.cpp cuda-4f13cb7 (cuda)
osUbuntu 24.04 LTS
kernel6.17.13-7-pve
driverNVIDIA 590.48.01 + CUDA 13.1
libc2.39
python3.12.3
llama.cppversion: 18 (4f13cb7) built with GNU 13.3.0 for Linux x86_64
build flagsGGML_CUDA=ON CMAKE_BUILD_TYPE=Release
runs/cell5
warmups2
endpoint/v1/chat/completions
streamingtrue
2× GeForce RTX 3090 · 24 GiB each
cpuAMD EPYC 7302P 16-Core Processor
gpuNVIDIA GeForce RTX 3090 × 2
archNVIDIA
vram48 GiB (system 64.0 GiB)
power200 W × 2 / 450 W × 2 max(44% cap)
pcieGen 4 x16 / Gen 4 x16 max
clocksgfx 1800/2100 MHz · mem 9501 MHz
temp41°C idle · 53°C peak
peak draw195 W
backendllama.cpp direct (cuda)
osUbuntu 24.04 LTS
driverNVIDIA 595.71.05 + CUDA 13.2
libc2.39
python3.12.3
llama.cppversion: 18 (4f13cb7) built with GNU 13.3.0 for Linux x86_64
build flagsGGML_CUDA=ON CMAKE_BUILD_TYPE=Release
runs/cell3
warmups1
endpoint/v1/chat/completions
streamingtrue
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)
cpuAMD RYZEN AI MAX+ 395 w/ Radeon 8060S
gpuAMD Radeon 8060S
archStrix Halo (gfx1151)
vram96 GiB (system 31.1 GiB, unified)
pcieGen 4 x16 / Gen 4 x16 max
clocksgfx 1158 MHz · mem 1000 MHz
temp47°C idle · 77°C peak
peak draw103 W
hardware probes
copy 41% of theoryFP16 peak 30.3 TF
256-bit8000 MHz20 SM/CU
Microbenchmarks for memory copy and tensor math; raw-engine decode and API workload rows measure model-serving speed.
captheorycopyfp16bf16
fixed256 GB/s106 GB/s30.3 TF-
compute: 11.5
backendllama.cpp direct (rocm)
osUbuntu 24.04 LTS
driverROCm 7.2.3
libc2.39
python3.12.3
llama.cppversion: 1 (4f13cb7) built with Clang 22.0.0 for Linux x86_64
build flagsGGML_HIP=ON AMDGPU_TARGETS=gfx1151 CMAKE_BUILD_TYPE=Release
runs/cell3
warmups1
endpoint/v1/chat/completions
streamingtrue
Strix Halo · Radeon 8060S · 128 GiB unified (96 GiB VRAM)
cpuAMD RYZEN AI MAX+ 395 w/ Radeon 8060S
gpuAMD Radeon 8060S
archStrix Halo (gfx1151)
vram96 GiB (system 31.1 GiB, unified)
backendllama.cpp 4f13cb7-mtp (rocm)
osUbuntu 24.04 LTS
kernel7.0.2-2-pve
python3.12.3
runs/cell5
warmups2
endpoint/v1/chat/completions
streamingtrue