One containerised dashboard per GPU host, each reading its own
/proc, driver and inference servers. The table below is gpumon’s own telemetry,
pushed by every host at 5 s and read live at render time; the wall underneath is live.
| host | gpu | peak · bandwidth | vram used / total | power | thermal | util · 3 min trend | engines | models answering | state |
|---|---|---|---|---|---|---|---|---|---|
| spark-1 unified memory · most models at once |
GB10 Grace BlackwellUNIFIED | 1 PFLOP fp4273 GB/s | 63.0 / 121.7 G52% | 15 Wno cap | 51° | 0% | 1 | qwen3-embedding:8b qwen-vl freelawproject/modernbert-embed-base_finetune_512 qwen3.6-35b-a3b |
LIVE |
| spark-2 unified memory · 80B MoE under vLLM |
GB10 Grace BlackwellUNIFIED | 1 PFLOP fp4273 GB/s | 4.1 / 121.7 G3% | 5 Wno cap | 45° | 0% | 1 | qwen3.8-flash-next |
LIVE |
| nvidia-one 24.0 GB VRAM 24 GB · dense 27B · 200 W |
GeForce RTX 3090 | 71 TF fp16936 GB/s | 22.3 / 24.0 G93% | 9 Wcap n/r | 43° | 0% | 1 | qwen3.8-27b-exl3 |
LIVE |
| nvidia-two 12.0 GB VRAM 12 GB · dense 12B |
GeForce RTX 3080 Ti | 68 TF fp16912 GB/s | 8.1 / 12.0 G67% | 5 Wcap n/r | 34° | 0% | 1 | gemma4-12b |
LIVE |
| nvidia-three 20.0 GB VRAM 20 GB · dense 27B |
RTX 4000 SFF Ada | 77 TF fp16280 GB/s | 11.4 / 20.0 G57% | 6 Wcap n/r | 38° | 0% | — | freelawproject/modernbert-embed-base_finetune_512 qwen3-reranker-0.6b |
LIVE |
| nvidia-four 20.0 GB VRAM 20 GB · same MoE as spark-1 · 50 W cap |
RTX 4000 SFF Ada | 77 TF fp16280 GB/s | 17.2 / 20.0 G86% | 6 Wcap n/r | 41° | 0% | — | qwen3.6-35b-a3b |
LIVE |
| Z14K000APLL/A (2022) Mac Studio · Mac13,2 unified memory · 64-core GPU · native, not Docker |
Apple M1 UltraUNIFIED | 21 TF fp32800 GB/s | 85.9 / 128.0 G67% | not measured | Nominal | 0% | — | qwen3.8-27b-heretic thomson-1.0-small qwen3-coder-30b-a3b |
LIVE |
| Z14V00173LL/A (2021) MacBook Pro · MacBookPro18,2 unified memory · 32-core GPU · native, not Docker |
Apple M1 MaxUNIFIED | 10 TF fp32400 GB/s | unknown | 0 Wno cap | — | 0%no trend | 0 | none answering | LIVE |
Keys inside any dashboard —
a all panels · p perf · h layers · m experts ·
b memory pipeline · v compare models · Tab/1–9 focus a model ·
r rescan · s settings · P palette · q quit (reload to restart)
On this page — Shift+P switches the whole page and every tile: light / dark / solarized ·
1/2/3 tiles per row. A readout in the top-right names every key you press.
Rendered 2026-10-06 19:48:14Z · hub/render-hub.ts ·
UNIFIED marks a part with no framebuffer counters: total is system RAM and used is the sum of
nvidia-smi --query-compute-apps, which is what the patched dashboard shows too.
Font Awesome Pro Thin is not bundled on the build machine, so this page ships without glyphs
rather than loading them from a CDN.