gpumon · integration API

broker · writer · ingress · ocr-api

llm-visuals

Per-host terminal dashboards for every GPU box in the fleet, plus a fleet hub page that tabulates all eight. Lives in services/llm-viz/; served at https://gpu.atsignhandle.xyz/llm-viz/.

Two halves, deployed independently:

services/llm-viz/agent/ services/llm-viz/hub/
what one container per GPU host, six ttyd dashboards the fleet table at /llm-viz/
runs on each of the eight hosts spark-1 (nginx :7690) + swarm (:7680)
refresh live, in the terminal rendered on a timer

Six ports per host

Each agent starts six ttyd processes — three writable, three public read-only — one per palette:

port access palette flags
7681 / 7682 / 7683 writable, loopback + tailnet light / dark / solarized -W
7684 / 7685 / 7686 public, read-only light / dark / solarized -b /llm-viz/<machine>/<palette>

The -b flag is load-bearing. It makes ttyd emit absolute token and WebSocket URLs that already carry the public path. That is why nothing in the chain strips a prefix: put a prefix-stripping proxy in front of these and every one of the 24 tiles breaks, because the URLs ttyd minted no longer match the path it is reached on. The dashboard's nginx.conf carries the same warning at the /llm-viz/ note, where the decision not to proxy is recorded alongside its cost: on the LAN origin 192.168.1.211 Traefik is not in front of the dashboard, so /llm-viz/ falls through the SPA try_files rule and quietly serves the dashboard shell instead of the hub. /seaweed has the same shape. The link is correct on gpu.atsignhandle.xyz and only there.

Routing chain

sunnypants tunnel
  → spark-1 Traefik            (~/Docker/traefik/data/config.yml)
      /llm-viz/                → hub nginx on spark-1, :7690 (bind mount)
      /llm-viz/<machine>/<p>/  → http://<lan-ip>:7684|7685|7686

The writable trio is never routed publicly. Reach it over the tailnet: https://<dns>.chihuahua-aeolian.ts.net:7681/. Note dns is not always the hostname — mepmbp2022 registers as mepmbp2022-6.

Two address schemes, one renderer

render-hub.ts emits two different pages and shipping either to the other place is the failure the split exists to prevent:

command output links
tailnet bun render-hub.ts dist/index.html https://<dns>.chihuahua-aeolian.ts.net:<port>/
public PUBLIC=1 bun render-hub.ts dist/public.html /llm-viz/<machine>/<palette>/

timer/llmvis-hub-refresh.sh guards both directions before deploying: it aborts if the public build contains a tailnet hostname, and aborts if the LAN build does not. A crossed pair is silently wrong otherwise — both pages render, both look fine, neither works.

Telemetry source

The hub reads gpumon's own writer. It does not collect anything itself.

GET /api/gpu/latest                          → per-host util, mem, power, temp, running_model
GET /api/gpu/series-all?bucket=5&since=<ts>  → 180 s sparkline, ~36 points per host
GET /api/local-hosts                         → litellm endpoint registry (engine count)

Base URL is GPUMON_API, default http://192.168.1.211:2290.

This replaced an eight-way SSH fan-out that ran nvidia-smi/ioreg remotely and sampled utilisation 20 × 0.35 s per sweep. That fan-out was a third copy of work gpumon already does — services/collector SSH-polls the same eight hosts and skips any that self-report, and services/gpu-reporter pushes from the hosts themselves every 5 s. It was also the worst of the three, because the hub renders on a timer: its numbers could be a quarter of an hour old while the writer held the same reading at 3 s. A sweep is now three HTTP GETs and zero SSH sessions, so the page stamp matches wall-clock time.

Fields this source does not carry

Stated rather than guessed, because a plausible wrong number is worse than a blank:

  • power.limit — no column in the writer. Unified parts (GB10, Apple) genuinely expose no cap and render no cap; discrete parts render cap n/r, which says the cap exists and this source does not report it.
  • engine count — /api/local-hosts is litellm's routing table, so a host serving through its own ollama registers nothing there. Zero-with-models renders as —, not 0.
  • Apple power — apple-gpu-stats needs the powermetrics sudoers drop-in from agent/macos/install-macos-telemetry.sh. Without it the host reports a sub-watt floor; that renders as 0.12 W · unreliable rather than rounding to a confident 0 W.
  • Apple temperature — macOS exposes no die temperature to userspace. temp_c on a Mac is apple-gpu-stats encoding OSThermalPressureLevel as a pseudo-celsius (30/60/80/90/95) to keep its CSV nvidia-smi shaped. The hub decodes it to Nominal/Moderate/Heavy/ Trapping/Sleeping and leaves temperature empty, so nothing can render it as heat.

Rendering and checks

cd services/llm-viz/hub
bun render-hub.ts              # dist/index.html  (tailnet)
PUBLIC=1 bun render-hub.ts     # dist/public.html (public)
bun badge-lint.ts              # square-badge contract
bun design-lint.ts             # house design system

tokens.ts is a symlink to the shared house design system; the renderer adds no stylesheet and every page is correct with JavaScript off.

When probing the 24 public tiles, probe serially. Parallel TLS handshakes through the tunnel produce phantom 000 responses that read as an outage on healthy hosts — that false signal has cost a debugging session already.