Observability¶
You add four things to your application. Everything else is already running.
| To get | You add |
|---|---|
| Logs | Nothing: write to standard output |
| Traces | Two environment variables |
| Metrics | A /metrics endpoint and x-metrics |
| Availability checks | x-status-check on a route |
Logs¶
Write one JSON object per line. Add trace_id to jump from a line to its trace.
{"level": "info", "message": "request served", "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736"}
Traces¶
compose.production.yaml
services:
web:
environment:
OTEL_SERVICE_NAME: my-app
OTEL_EXPORTER_OTLP_ENDPOINT: http://alloy.o11y.svc.cluster.local:4318 # (1)!
- The collector of the platform. It exists on the cluster, so declare it in the file of that environment.
Metrics¶
compose.yaml
services:
web:
image: my-app:local
expose:
- "9090" # (1)!
x-metrics:
port: 9090 # (2)!
path: /metrics
interval: 30s
- The port where your application serves metrics.
- It must be a port declared in
exposeorports.
Generated on every release. You never write these files or see them.
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
labels:
com.docker.compose.project: my-app
com.docker.compose.service: web
release: monitoring # (1)!
name: web
namespace: my-app
spec:
endpoints:
- interval: 30s # (2)!
path: /metrics
port: p-9090-tcp # (3)!
selector:
matchLabels:
com.docker.compose.project: my-app
com.docker.compose.service: web
- Makes the collector of the platform pick it up.
- Your
intervalandpath. - Your
port, as the service exposes it.
Only the collector may reach that port.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
labels:
com.docker.compose.project: my-app
com.docker.compose.service: web
name: web-metrics
namespace: my-app
spec:
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: o11y
podSelector:
matchLabels:
app.kubernetes.io/name: prometheus # (1)!
ports:
- port: 9090
protocol: TCP
podSelector:
matchLabels:
com.docker.compose.project: my-app
com.docker.compose.service: web
policyTypes:
- Ingress
- The collector of the platform, and nobody else.
| Field | Required | Default |
|---|---|---|
port |
Yes | |
path |
No | /metrics |
interval |
No | 30s |
The platform collects the endpoint at that interval. It does not instrument your application: the endpoint is yours.
Availability checks¶
A check belongs to a route: it asks the public address, from outside, like a user would.
compose.production.yaml
services:
web:
x-ingress:
routes:
- hostname: app.example.com
port: 8000
x-status-check:
name: My app # (1)!
group: my-app
url: https://app.example.com/healthz # (2)!
interval: 1m
conditions:
- "[STATUS] == 200" # (3)!
- How the check appears in Gatus.
- The address the check calls.
- What a healthy answer looks like, in the syntax of Gatus.
Generated on every release. You never write this file or see it. The check travels with the route of your domain.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: web-28059829b1
namespace: my-app
labels:
com.docker.compose.project: my-app
com.docker.compose.service: web
annotations:
gatus.home-operations.com/endpoint: 'name: My app
group: my-app
url: https://app.example.com/healthz
interval: 1m
conditions:
- ''[STATUS] == 200''
' # (1)!
spec:
hostnames:
- app.example.com
parentRefs:
- name: envoy-external
namespace: network
rules:
- backendRefs:
- name: web
port: 8000
matches:
- path:
type: PathPrefix
value: /
- Your
x-status-check, as Gatus reads it.
| Field | Required | Default |
|---|---|---|
name |
Yes | |
group |
Yes | |
url |
Yes | |
interval |
No | 1m |
conditions |
No | ["[STATUS] == 200"] |
Where to look¶
| Question | Open |
|---|---|
| Is it up? | Gatus |
| What did it log? | Grafana → Explore → Loki |
| Why was this request slow? | Grafana → Explore → Tempo |
| How is it trending? | Grafana → Dashboards |
Logs of your application, in Loki
{namespace="my-app-staging"}