Field notes
Observability decisions from real production work
I write about the decisions behind reliable telemetry: what I inspect, what I change and why it matters during an incident.
By Rafael Fahreev, Observability ExpertPRODUCTION / INCIDENT 042LIVE
Gateway99.98%
Checkout3.8s
Payments APIOWNER?
Incident response7 min read
Your Monitoring Works. Why Does Incident Response Still Start With Guessing?
Dashboards can be green, alerts can be firing, and responders can still have no reliable path from user impact to the failing component. This is how I find the missing links.
22 July 2026Read article →
OTEL / PIPELINEHEALTHY
01RECEIVEOTLP
02PROCESSbatch · memory
03EXPORTqueue · retry
QUEUE38 / 100
OpenTelemetry8 min read
How I Take the OpenTelemetry Collector From Proof of Concept to Production
A Collector that exports a demo trace is not yet a production pipeline. I check topology, backpressure, state, routing and the Collector's own telemetry before rollout.
22 July 2026Read article →
ON-CALL / DECISION PATHOWNED
SIGNALUser error rate
SEVERITYPage
OWNERCheckout SRE
ACTIONProtect checkout
BURN RATE12.4×
Alerting and SLOs7 min read
I Do Not Fix Alert Fatigue by Tuning Thresholds
Threshold tuning can reduce noise for a week. I fix the operating model behind paging: user impact, ownership, severity, action and a deliberate SLO policy.
22 July 2026Read article →
