← Back to blog

Giving back: observability that answers a real question

October 3, 2026

Observability and automation were part of my work at Magic Leap, and SRE work at Chewy reinforced why telemetry needs to help explain a running system. A dashboard is most useful when it answers a question you actually have.

The LGTM Kubernetes lab gives readers a small way to follow that evidence. Loki stores logs, Tempo stores traces, Mimir stores metrics, and Grafana lets you query and explore them. Alloy is the collector that routes the configured signals to those backends.

Follow one signal end to end

Start with one known log, metric, or instrumented request. Confirm it is emitted, collected, stored, and queryable. When one of those stages fails, an empty dashboard can look deceptively like a quiet application.

Logs do not automatically create traces. Applications need suitable instrumentation, context propagation, and an exporter. This example receives application OTLP spans, and its configured log coverage is deliberately limited. Expanding coverage requires matching discovery configuration and permissions rather than granting a collector access to everything.

Read the lab's limits

The chart uses single-replica components and persistent filesystem storage. Those choices make the system inspectable, but they allow upgrade downtime and do not provide a highly available observability service. The cluster needs a suitable StorageClass; a persistent volume is not a backup.

The backends are private, unauthenticated lab services behind a namespace network policy, with plaintext connections inside that restricted network. The policy needs a CNI that enforces it. Production or shared environments require a different review of authentication, TLS, tenant isolation, capacity, and storage.

Retention also does not guarantee a disk cannot fill. Observe the observability stack itself: storage use, collector queues, and component health are part of the system you are operating. Alert evaluation and successful notification delivery are separate checks.

Use it to understand a workload

The repository's render workflow lets you inspect manifests before installation. Choose the GitHub or GitLab delivery path you intend to own, and use its native runner and protected configuration. The GitOps guide connects that choice to the rest of the lab.

For my direction toward AI infrastructure, I want the same discipline around an inference workload: know what the signal measures, what coverage is missing, and what evidence supports the conclusion. This lab is a place to practice that discipline with understandable components.