Qwen/Qwen2.5 32B-Instruct
unknown·32B params·unknown
intelligence: see on Artificial Analysis →
checkpoint:
Qwen/Qwen2.5-32B-Instruct-AWQcommit:
5c7cb76a268fweights 18.00 GiB
All runs (5)
| Hardware | Backend | Shape | Conc. | Gen tok/s ↓ | TTFT | TPOT (ms) | Out tok | Total | VRAM Δ |
|---|---|---|---|---|---|---|---|---|---|
| GeForce RTX 3090 · 24 GiB | vLLM 0.21.0 (cuda) | codegen | 1 | 19.3 | 189ms | 51.4 | 625 | 32.42s | 0.000 GiB |
| GeForce RTX 3090 · 24 GiB | vLLM 0.21.0 (cuda) | chat | 1 | 19.2 | 112ms | 51.3 | 87 | 4.20s | 0.000 GiB |
| GeForce RTX 3090 · 24 GiB | vLLM 0.21.0 (cuda) | agent | 1 | 19.2 | 119ms | 51.5 | 301 | 15.70s | 0.000 GiB |
| GeForce RTX 3090 · 24 GiB | vLLM 0.21.0 (cuda) | rag | 1 | 18.8 | 114ms | 52.6 | 55 | 3.41s | 0.000 GiB |
| GeForce RTX 3090 · 24 GiB | vLLM 0.21.0 (cuda) | agent | 4 | 18.7 | 155ms | 53.4 | 301 | 16.11s | 0.000 GiB |
Environment
GeForce RTX 3090 · 24 GiB
cpuAMD EPYC 7302P 16-Core Processor
gpuNVIDIA GeForce RTX 3090
archNVIDIA
vram24 GiB (system 64.0 GiB)
power200 W / 450 W max(44% cap)
backendvLLM 0.21.0 (cuda)
serverlemonade unknown
osUbuntu 24.04 LTS
kernel6.17.13-7-pve
driver590.48.01
python3.12.3
containerizedtrue
runs/cell5
warmups2
endpoint/v1/chat/completions
streamingtrue