Lesson
Building taught me something ugly this week. One of our…
Turns out our queue jobs were propagating traceparent from the parent span down into every fan-out child, and one nightly command fans out into hundreds of small jobs. Each child kept tracing back to the same root. The root just... kept growing. Nobody told it to stop. Found it because Grafana started choking on ingest, not because anything crashed. That's the part that gets me. The system was "working." Dashboards loaded slow, someone shrugged, moved on. Took actually going into Tempo and asking why one trace was 20MB to find it. Fix was 3 levers, nothing clever: stop propagating traceparent into unrelated fan-out children, cap span attributes on the hot path, and a sampling ratio that actually matches our traffic instead of a number picked on day one and never revisited. Lesson learned: observability tools can hide the exact problem they exist to catch, if you never look at the raw trace and only trust the summary graph. Dashboard said "fine." Trace said 20 megabytes of nonsense. Still have a follow-up open on where the threshold should sit long-term. Dont have that answer yet. Will post when I do :-)
- Reactions
- 0
- Comments
- 1
- Shares
- 0
Public discussion
Uriel Bitton
its hard to find a good observability tool. I built some myself but none were successful. The lesson is always valuable !