Observability Across Cloud and Datacenter Boundaries
October 3, 2026
Observability Across Cloud and Datacenter Boundaries
As my work has grown from network engineering into platform and SRE, observability has become a way to connect the layers. A device, Kubernetes workload, provisioning workflow, and application can each report something useful. The challenge is making those signals answer an operational question without losing their ownership or history.
I am interested in that problem across both datacenter and cloud infrastructure. It is also central to the AI infrastructure roles I am pursuing: a model endpoint is another service with dependencies, resource constraints, and failures that need to be explained. Useful visibility should help an engineer decide what to inspect next.
A metrics tenancy milestone
On June 15, 2026, I committed a platform Mimir tenancy cutover. Headerless platform writes were directed into a dedicated platform tenant, while Grafana's federated read path could query both the new tenant and the previous historical tenant.
The important detail is that historical data remained readable in place. I did not need to describe a data migration that had not happened. The read configuration preserved continuity across the change, while new writes followed the platform tenancy decision.
The scope was equally important. This was a platform metrics change. It did not establish that Loki and Tempo had received the same tenancy migration. Those systems were unchanged in that commit. Calling the result a complete observability-stack conversion would have hidden the actual boundary of the work.
That is a lesson I keep returning to: observability configuration is part of the system's behavior, and its claims should be as precise as application claims. A dashboard that displays data is useful evidence, but I also need to understand which tenant and source supplied it.
Connecting the architecture to the workflow
CloudSpinUp is another part of my platform work. Its repository describes a self-service hybrid cloud architecture: a Next.js customer portal connects to an API gateway and microservices, Temporal coordinates provisioning workflows, NATS carries events, and PostgreSQL and Redis support the services.
The platform brings infrastructure integrations into that workflow, including Proxmox, public-cloud providers, inventory and addressing, and billing. Provider MCPs sit beneath the application boundary to adapt provider inventory and lifecycle operations. This is the architecture I am building around; it should not be read as a claim that every provider or catalog offer has completed production qualification.
Observability needs to follow those boundaries. An order accepted by the portal is different from a workflow finishing, a resource being provisioned, or a service becoming usable. When a step fails, the useful record connects the user-facing request to the component responsible for the next action.
The work I want to keep doing
These projects bring together the areas I enjoy: networking, cloud infrastructure, delivery automation, operational diagnosis, and software that turns infrastructure into a service. They also show why I am drawn to forward deployed engineering. Understanding the environment and the intended outcome is part of making a product work.
I want to keep building platforms where the evidence is clear enough to support the next decision. For me, good observability is practical: it helps explain a failure, protects the boundaries of a change, and gives engineers a better basis for recovery.