federation · gpu hosts

llm-visuals — fleet

One containerised dashboard per GPU host, each reading its own /proc, driver and inference servers. The table below is gpumon’s own telemetry, pushed by every host at 5 s and read live at render time; the wall underneath is live.

hostgpupeak · bandwidthvram used / totalpowerthermal util · 3 min trendenginesmodels answeringstate
spark-1
unified memory · most models at once
GB10 Grace BlackwellUNIFIED 1 PFLOP fp4273 GB/s 63.0 / 121.7 G52% 15 Wno cap 51° 0% 1 qwen3-embedding:8b qwen-vl freelawproject/modernbert-embed-base_finetune_512 qwen3.6-35b-a3b LIVE
spark-2
unified memory · 80B MoE under vLLM
GB10 Grace BlackwellUNIFIED 1 PFLOP fp4273 GB/s 4.1 / 121.7 G3% 5 Wno cap 45° 0% 1 qwen3.8-flash-next LIVE
nvidia-one
24.0 GB VRAM
24 GB · dense 27B · 200 W
GeForce RTX 3090 71 TF fp16936 GB/s 22.3 / 24.0 G93% 9 Wcap n/r 43° 0% 1 qwen3.8-27b-exl3 LIVE
nvidia-two
12.0 GB VRAM
12 GB · dense 12B
GeForce RTX 3080 Ti 68 TF fp16912 GB/s 8.1 / 12.0 G67% 5 Wcap n/r 34° 0% 1 gemma4-12b LIVE
nvidia-three
20.0 GB VRAM
20 GB · dense 27B
RTX 4000 SFF Ada 77 TF fp16280 GB/s 11.4 / 20.0 G57% 6 Wcap n/r 38° 0% — freelawproject/modernbert-embed-base_finetune_512 qwen3-reranker-0.6b LIVE
nvidia-four
20.0 GB VRAM
20 GB · same MoE as spark-1 · 50 W cap
RTX 4000 SFF Ada 77 TF fp16280 GB/s 17.2 / 20.0 G86% 6 Wcap n/r 41° 0% — qwen3.6-35b-a3b LIVE
Z14K000APLL/A (2022)
Mac Studio · Mac13,2
unified memory · 64-core GPU · native, not Docker
Apple M1 UltraUNIFIED 21 TF fp32800 GB/s 85.9 / 128.0 G67% not measured Nominal 0% — qwen3.8-27b-heretic thomson-1.0-small qwen3-coder-30b-a3b LIVE
Z14V00173LL/A (2021)
MacBook Pro · MacBookPro18,2
unified memory · 32-core GPU · native, not Docker
Apple M1 MaxUNIFIED 10 TF fp32400 GB/s unknown 0 Wno cap — 0%no trend 0 none answering LIVE

Keys inside any dashboard — a all panels · p perf · h layers · m experts · b memory pipeline · v compare models · Tab/1–9 focus a model · r rescan · s settings · P palette · q quit (reload to restart)
On this page — Shift+P switches the whole page and every tile: light / dark / solarized · 1/2/3 tiles per row. A readout in the top-right names every key you press.

tiles per row − + 1 — full width 2 3 drag a tile by its caption to reorder
spark-1 GB10 Grace Blackwell
spark-2 GB10 Grace Blackwell
nvidia-one GeForce RTX 3090
nvidia-two GeForce RTX 3080 Ti
nvidia-three RTX 4000 SFF Ada
nvidia-four RTX 4000 SFF Ada
Z14K000APLL/A Apple M1 Ultra
Z14V00173LL/A Apple M1 Max

Rendered 2026-10-06 19:48:14Z · hub/render-hub.ts · UNIFIED marks a part with no framebuffer counters: total is system RAM and used is the sum of nvidia-smi --query-compute-apps, which is what the patched dashboard shows too. Font Awesome Pro Thin is not bundled on the build machine, so this page ships without glyphs rather than loading them from a CDN.