← All topics
Observability and service reliability
Monitoring, tracing, metrics, telemetry, service reliability, and operational incidents.
5,777 stories · 118 in 30 days · Updated Sep 28, 2026
Sign up for the newsletterThe pace of conversation
Posts per month
2020-012026-09
Posts per month · first and latest months may be partial.
01 / EXPLORE THROUGH TIME
Timeline
Start with the latest. Scroll back through time, or zoom in for more.
20261,094 stories
2025895 stories
2024767 stories
2023795 stories
2022644 stories
2021644 stories
2020938 stories
03 / GO DEEPER
Follow your curiosity
Find a particular story, revisit a year, or see what resonated.
Xiaomi Mimo 2.6 live post-training dashboard↗ mimo.xiaomi.com · HN · 155 comments · 562 points · Sep 16, 2026
The Normalization of Inexplicable Failures↗ ihatethefuture.com · HN · 119 comments · 274 points · Sep 27, 2026
Nobody is saying why OpenAI and Anthropic had outages↗ wired.com · HN · 4 comments · 207 points · Sep 04, 2026
Show HN: Sigabrt.dev – cronjob monitor with an SSH TUI↗ sigabrt.dev · HN · 35 comments · 81 points · Sep 19, 2026
How we configured OpenTelemetry logs in Rails↗ sixpatterns.com · HN · 8 comments · 31 points · Aug 31, 2026
Show HN: Aura – a Rust agent that investigates and fixes production incidents↗ github.com · HN · 6 comments · 28 points · Sep 02, 2026
Issues with Codex – Identified – Full Outage↗ status.openai.com · HN · 1 comments · 27 points · Sep 25, 2026
Show HN: ClaudeStatsBar: your session is 486k deep and nothing told you↗ github.com · HN · 15 comments · 18 points · Sep 11, 2026