When logs live in one tool, metrics in another, and traces nowhere, every incident starts with a scavenger hunt. I designed a central observability platform so everyone starts from the same place: Grafana on top, Loki for logs, Mimir for metrics, Tempo for traces, and OpenTelemetry as the way data gets in.
Instrument once
OpenTelemetry is the most important choice in the stack, because it separates applications from backends. Services send data over OTLP to a collector, and the collector decides where it goes. If a backend changes later, application code doesn't have to. Here's a trimmed-down collector pipeline:
receivers:
otlp:
protocols: { grpc: {}, http: {} }
processors:
batch: {}
exporters:
otlp/tempo: { endpoint: tempo:4317, tls: { insecure: true } }
prometheusremotewrite: { endpoint: http://mimir:9009/api/v1/push }
otlphttp/loki: { endpoint: http://loki:3100/otlp }
service:
pipelines:
traces: { receivers: [otlp], processors: [batch], exporters: [otlp/tempo] }
metrics: { receivers: [otlp], processors: [batch], exporters: [prometheusremotewrite] }
logs: { receivers: [otlp], processors: [batch], exporters: [otlphttp/loki] }Connect the signals
When every log line carries a trace ID, you can jump from an error in Loki to the full request in Tempo, and from there to the metrics of the service that was slow. Those links are what make the separate backends feel like one system.
Watch your labels
Loki and Mimir are cheap to run when labels have low cardinality and expensive when they don't. User IDs, request IDs, and full URLs belong in the log line or the trace, not in labels. That one rule does more for cost and query speed than most tuning.
Dashboards answer questions
A dashboard should answer a question someone actually asks during an incident: is it us or a dependency, when did it start, which version is running? Dashboards that nobody opens are noise. Start with a few service-level views and add more when a real incident shows what's missing.
