dooner.tech
Inference / Live view

VLLM operations

Live, readable telemetry from every vLLM-like Docker workload discovered on AI01. Pick a container to inspect its runtime details and watch the log flow by subsystem.

connecting
HOST AI01
LAST SAMPLE waiting
log refresh every 3s
State
---
waiting for Docker inventory
Model
---
served model
Image
---
Docker image
Log lines
0
bounded tail
Latest observed telemetry

Performance

Prefill / prompt
not reported
latest prompt throughput
Decode / generation
not reported
latest generation throughput
GPU KV cache
not reported
current cache utilization
Prefix cache hit
not reported
internal and external hit rate
MTP / speculative
not reported
depth and acceptance
LMCache tokens
not reported
hit, computed, and load
CKV / cache transfer
not reported
CKV or transfer activity
Requests
not reported
running and waiting
Values are the newest matching measurements in the fetched log tail.
Container runtime

Docker details

Container

Name
---

Model and runtime

no model settings yet
Delineated log stream

Subsystem flow

LMCache / KV transfer

0 lines
No matching lines yet.

Prefill / prompt

0 lines
No matching lines yet.

Decode / engine

0 lines
No matching lines yet.

Requests / serving

0 lines
No matching lines yet.

Other / startup / warnings

0 lines
No matching lines yet.