6 ms·
How We Found a Missing Scala Class
- koube 8y agoFYI this domain is blocked by default for uBlock users.
- kazinator 8y agoWild-assed kazinator guess: probably be a substring/regex match on "analytics". Must be a tracking domain!
- manigandham 8y agoIt's on the Pete Lowe Adserver list directly, since Heap is in fact a tracking system and domain: https://pgl.yoyo.org/as/serverlist.php?showintro=0;hostformat=hosts https://pgl.yoyo.org/as/serverlist.php?showintro=0;hostforma...
- pharrington 8y agoWhile Heap Analytics indeed is tracking software, clicking the "temporarily unblock" button allowed me to read a pretty good detective story.
- draw_down 8y agoSucks for ublock users I guess.
- Inetgate 8y agoYa, I see. And Google and Wayback machine have not yet cached the issued url content.
- emilfihlman 8y agoArchive.org is getting a 403 forbidden nginx. Archive.is worked: http://archive.is/kKcRV http://archive.is/kKcRV
- drob 8y agoHeap CTO here – would love to answer any questions you have. This was my first exposure to btrace, which a super useful swiss army knife for JVM debugging. That made this a worthwhile adventure for sure.
- lixtra 8y agoHow much time did your team spend on the debugging effort?
- drob 8y agoIirc Ivan (post author) spent a few days tracking this down. There were some other debugging dead ends that we omitted in this writeup. One red herring was that the issue appeared to happen during the US morning, so there was some time-of-day component, and we thought it might be a system load issue. The fix turned out to be fairly involved too – on the order of a week I think. Ivan works from Bulgaria so sadly he is asleep right now.
- lostmyoldone 8y agoIsn't that a bit odd, your stack trace there in the article? Usually the stack trace when a NoClassDefFoundError is thrown contains a "Cause" clearly showing the name of the class loader that was supposed to know about the class in question, and which then rather obviously failed to load it. If it actually isn't in the real trace, the exception/error logging is probably a bit wonky. I have seen logging procedures that don't traverse/print the entire traceback chain, but only prints the message and the immediate stack. But this is unfortunately a rather terrible idea. Quite often it will exclude the actual cause from the printed trace while simulataneously retaining the error message, not rarely leading to liberal amounts of confusion. The pattern with exceptions being thrown with a cause is not uncommon in the JDK, so making sure to log causes are important. In general, although it's probably obvious, I would like to mention that having code loaded by class loaders with different lifetimes interact is rife with "interesting" issues. Wherever possible I would recommend serializing messages over any boundaries where class loaders have different lifetimes, as it both prevents all of the strangest causes of errors, and can also lead to a cleaner design. An exception would be if prohibitively expensive from a performance perspective, of course.
- nambit 8y agoWhy doesn't java just spit out a classLoaderClosed error?
- th3iedkid 8y agoYes, believe that was the bad part from JMV implementation!
- SuspiciousSwan 8y agoIt sounds like you have a lot of operation issues due to the technologies that you used. I mean, at least you aren't doing your backend in node, but running an actor system on top of an actor system is going to be brutal to properly analyze once you actually have scale. What sort of process do you have for picking trendy technologies vs tested ones, and how much do you talk to people who have built large scale systems before implementing things like scala?
- jjjensen90 8y agoYeah... Reading this, it smacked of a possible combination of poor tool choice and over-engineering (which I've been guilty of plenty). I built a video processing/workflow application in Scala with Akka a few years ago and debugging that was hard enough, eventually it was refactored to a simpler Kotlin/Spring application... Actor systems are great for certain use cases but you can really hurt the transparency of your app if you aren't careful. I can't imagine maintaining the OP's application at scale for this use case, but maybe they have someone smarter than me!
- rozap 8y agoCounterpoint: debugging erlang systems in production is a cakewalk. The tracing and introspection tools that come bundled in OTP make tracking problems down really easy. It's really hard to go back to systems that don't have erlang level visibility, so much so that it's kind of a crutch sometimes. This is an ecosystem problem and not something inherent in a program using an actor abstraction.
- jjjensen90 8y agoAh yes, that is a great point. Erlang/OTP were designed to be used as actor systems, whereas the actor implementations in Scala and other JVM languages/frameworks are at least one level of abstraction above that. Definitely agreed on Erlang/OTP having a wonderful set of tools for debugging/visibility, but I still stand by my assessment that OPs problems are from over-engineering (and secondarily from the ecosystem).
- userbinator 8y agoNoClassDefFoundError? But it’s right there! Although in this case the cause was very different, it reminds me of an old "trap for young players" with loading shared libraries dynamically --- the library itself can exist and be readable and executable, and yet attempting to load it fails with a "file not found" error. This happens when one of its dependencies, directly or indirectly, is missing.
- sk5t 8y agoAh! This was one of my very, very least favorite things about developing win32 DLLs, way back in the day.
- Lazare 8y agoI ran across another similar-yet-very-different example of this once in a completely different language. In my case, we added a new class, it worked fine on dev, then failed on staging (and would have failed in production if we'd let it go that far). This was confusing, because we were using Vagrant to ensure our dev and staging environments were identical. What could be going on? Well, our linux VMs were being hosted in OS X, with shared folders for the code, and by default OS X volumes are not case sensitive. Meanwhile the linux staging and production environments were using actual linux filesystems, which were case sensitive. So someone added a new class MyFancyClass, then tried to import it as MyFancyclass, and it worked great in dev since the underlying FS of the host OS could find the file, then failed on staging. A fun debugging ride, and a good reminder that 1) having dev and staging the same is really important and 2) that might be harder than you think. :)
- djsumdog 8y agoThe moment the article mentioned "Fat jar" I knew that'd be the problem. I don't recommend using any type of fat jar plugin (like OneJar) or even Google Guice for that matter. Custom class loaders are a nightmare. Thanks to Docker containers, you should never really need a far jar again. Just find a decent Docker packager for your build system (sbt, gradle, etc.) and it can plop all your dependencies in there in a nice, isolated container that uses the standard class loader.
- Quekid5 8y agoFat jars are evil, but there's no need for Docker: Just collect the dependency jars in a lib/ (or whatever) folder and explicitly give them on the classpath when running the 'java' executable. We use the sbt 'pack' plugin where I work and it works a treat. (Docker has its place, but it's massive overkill just to avoid fat jars.)
- thecatspaw 8y agoWhat is problematic about Fat jars? The problem seemed to be Flink's implementation to unload the FatJar's classloader when erroring. This would have happened with slim jars as well, wouldnt it? I also dont see how docker relates exactly, you can have hundreds of library jars in a classpath with standard classloaders, no docker required
- Quekid5 8y agoI'm not sure about whether this particular problem has anything to do with fat jars, but there a couple of really big annoyances with fat jars which have bitten me occasionally: * If you flatten the classpath completely when building the fat jar then you need to somehow be able reconcile duplicate classpath entries[1]. (Duplicate class path entries are perfectly within spec as far as I can tell -- at least as long as they are from distinct jars. Not sure if they're allowed in a single jar.) * If you use nested jars (+ custom classloader perhaps) then it becomes impossible to refer to classpath entry resources in nested jars in a standard way (via java.net.URL, that is). Sometimes you can use foo.jar!bar.jar/blah, but even with that non-standard syntax I've encountered at least one case where it was impossible to refer to a doubly-nested classpath resource (don't ask). The lack of standard support for nesting jars seems like an oversight in the class loading/resource API, but there it is. There's probably more, but that's at least a couple of the ones that come to mind. [1] For example the plugin data file used by log4j 2.x where you have to have custom merge logic to handle that specific (binary!) file. You might blame this on log4j, but as a practical matter it's hard to avoid it. This is a pretty rare scenario, but any custom build logic or special-case plugins can be a huge pain for maintenance.
- fgheorghe 8y agoHow do you lose a class in a programming language?!
- mb720 8y agoThe title is misleading. The class was there all along, the NoClassDefFoundError was thrown because the class loader was closed when trying to load a class.
- saati 8y agoBut why can you even do that?
- rad_gruchalski 8y agoIt says right there in the article.
- GrumpyNl 8y agoI keep running in this type of problems all the time with our developers. Please keep it simple. Take a step back and ask yourself, do i need all this stuff, is this the best approach. Often they just blindly accept all the external libs. For me as an old school guy, i don't trust all those dependencies at all.