llm-visuals
Per-host terminal dashboards for every GPU box in the fleet, plus a fleet hub page that
tabulates all eight. Lives in services/llm-viz/; served at
https://gpu.atsignhandle.xyz/llm-viz/.
Two halves, deployed independently:
services/llm-viz/agent/ |
services/llm-viz/hub/ |
|
|---|---|---|
| what | one container per GPU host, six ttyd dashboards |
the fleet table at /llm-viz/ |
| runs on | each of the eight hosts | spark-1 (nginx :7690) + swarm (:7680) |
| refresh | live, in the terminal | rendered on a timer |
Six ports per host
Each agent starts six ttyd processes — three writable, three public read-only — one per
palette:
| port | access | palette | flags |
|---|---|---|---|
| 7681 / 7682 / 7683 | writable, loopback + tailnet | light / dark / solarized | -W |
| 7684 / 7685 / 7686 | public, read-only | light / dark / solarized | -b /llm-viz/<machine>/<palette> |
The -b flag is load-bearing. It makes ttyd emit absolute token and WebSocket URLs
that already carry the public path. That is why nothing in the chain strips a prefix: put a
prefix-stripping proxy in front of these and every one of the 24 tiles breaks, because the
URLs ttyd minted no longer match the path it is reached on. The dashboard's nginx.conf
carries the same warning at the /llm-viz/ note, where the decision not to proxy is
recorded alongside its cost: on the LAN origin 192.168.1.211 Traefik is not in front of
the dashboard, so /llm-viz/ falls through the SPA try_files rule and quietly serves the
dashboard shell instead of the hub. /seaweed has the same shape. The link is correct on
gpu.atsignhandle.xyz and only there.
Routing chain
sunnypants tunnel
→ spark-1 Traefik (~/Docker/traefik/data/config.yml)
/llm-viz/ → hub nginx on spark-1, :7690 (bind mount)
/llm-viz/<machine>/<p>/ → http://<lan-ip>:7684|7685|7686
The writable trio is never routed publicly. Reach it over the tailnet:
https://<dns>.chihuahua-aeolian.ts.net:7681/. Note dns is not always the hostname —
mepmbp2022 registers as mepmbp2022-6.
Two address schemes, one renderer
render-hub.ts emits two different pages and shipping either to the other place is the
failure the split exists to prevent:
| command | output | links | |
|---|---|---|---|
| tailnet | bun render-hub.ts |
dist/index.html |
https://<dns>.chihuahua-aeolian.ts.net:<port>/ |
| public | PUBLIC=1 bun render-hub.ts |
dist/public.html |
/llm-viz/<machine>/<palette>/ |
timer/llmvis-hub-refresh.sh guards both directions before deploying: it aborts if the
public build contains a tailnet hostname, and aborts if the LAN build does not. A crossed
pair is silently wrong otherwise — both pages render, both look fine, neither works.
Telemetry source
The hub reads gpumon's own writer. It does not collect anything itself.
GET /api/gpu/latest → per-host util, mem, power, temp, running_model
GET /api/gpu/series-all?bucket=5&since=<ts> → 180 s sparkline, ~36 points per host
GET /api/local-hosts → litellm endpoint registry (engine count)
Base URL is GPUMON_API, default http://192.168.1.211:2290.
This replaced an eight-way SSH fan-out that ran nvidia-smi/ioreg remotely and sampled
utilisation 20 × 0.35 s per sweep. That fan-out was a third copy of work gpumon already
does — services/collector SSH-polls the same eight hosts and skips any that self-report,
and services/gpu-reporter pushes from the hosts themselves every 5 s. It was also the
worst of the three, because the hub renders on a timer: its numbers could be a quarter of
an hour old while the writer held the same reading at 3 s. A sweep is now three HTTP GETs
and zero SSH sessions, so the page stamp matches wall-clock time.
Fields this source does not carry
Stated rather than guessed, because a plausible wrong number is worse than a blank:
power.limit— no column in the writer. Unified parts (GB10, Apple) genuinely expose no cap and renderno cap; discrete parts rendercap n/r, which says the cap exists and this source does not report it.- engine count —
/api/local-hostsis litellm's routing table, so a host serving through its own ollama registers nothing there. Zero-with-models renders as—, not0. - Apple power —
apple-gpu-statsneeds thepowermetricssudoers drop-in fromagent/macos/install-macos-telemetry.sh. Without it the host reports a sub-watt floor; that renders as0.12 W · unreliablerather than rounding to a confident0 W. - Apple temperature — macOS exposes no die temperature to userspace.
temp_con a Mac isapple-gpu-statsencodingOSThermalPressureLevelas a pseudo-celsius (30/60/80/90/95) to keep its CSVnvidia-smishaped. The hub decodes it toNominal/Moderate/Heavy/Trapping/Sleepingand leaves temperature empty, so nothing can render it as heat.
Rendering and checks
cd services/llm-viz/hub
bun render-hub.ts # dist/index.html (tailnet)
PUBLIC=1 bun render-hub.ts # dist/public.html (public)
bun badge-lint.ts # square-badge contract
bun design-lint.ts # house design system
tokens.ts is a symlink to the shared house design system; the renderer adds no stylesheet
and every page is correct with JavaScript off.
When probing the 24 public tiles, probe serially. Parallel TLS handshakes through the
tunnel produce phantom 000 responses that read as an outage on healthy hosts — that false
signal has cost a debugging session already.