4 ms·
Had a fun one relatively recently that was a mix of "hard to reproduce" and "hard to get internal state information". Flaky test in a rails app that would fail
by ThrustVectoring 5y ago
Had a fun one relatively recently that was a mix of "hard to reproduce" and "hard to get internal state information". Flaky test in a rails app that would fail one in every ten to hundred runs of the full test suite with "this random number is too big to be a primary key" kind of message. Root cause was an edge case passing through multiple swiss-cheese holes in various assumptions:
1. ActiveRecord makes an assumption that primary keys are integers, and does its own check whether or not they are big enough to be persisted (rather than catching a database error).
2. Furthermore, it does this by coercing the key to an integer if it isn't already one. This is done with to_i, which for a string takes any leading 0-9 characters and discards the rest.
3. We had a table with a string primary key (UUID of some sort)
4. One of our test factories was generating a hexadecimal string for that primary key
5. And the test factory was not deterministic and did not respect the --seed flag in the test suite.
So the end result was a very innocuous-looking line of code occasionally generating a hexadecimal string with enough leading numeric digits to be larger than ActiveRecord thinks you can stuff into a table, causing an extremely cryptic error message. It does not reproduce with the same test seed. And the stack trace was about three frames and discarded all the context - all I could see was that ActiveRecord was throwing a fit over somehow mysteriously receiving a large number somehow.
Figuring that out was honestly like, 80% pure luck. Chased down a hunch that it was test object generation somehow and did that in the REPL, then narrowed it down via looking through child object generation until the haystack was small enough that I couldn't not find the needle.
- meowface 5y ago>And the stack trace was about three frames and discarded all the context - all I could see was that ActiveRecord was throwing a fit over somehow mysteriously receiving a large number somehow. This is Python rather than Ruby, but my average debugging time and frequency of "impossible to figure out" bugs drastically decreased once I started using a traceback library that provides a lot of context to each stack frame. It makes logs containing any raised exceptions much larger and more tedious to scroll through, but the benefits are more than worth it; especially for production services. I use better-exceptions (https://github.com/Qix-/better-exceptions https://github.com/Qix-/better-exceptions), but there are a bunch of other good libraries as well.
- Groxx 5y agoYeah, when we turned this feature on in Sentry a few years back, bugs got so much easier to fix. 90% of the time the stack trace and all arguments along the way were sufficient to figure out the full cause just from reading, no need to try to reproduce. It is enormously CPU-expensive though, so it can be risky in production. Small-ish error spikes can cause enough latency to cause more errors, which cause more latency...