5 ms·
Scalding at Etsy
- morgante 13y agoTotally read this as scaling at Etsy. Still interesting though. And somewhat accurate.
- krick 13y agoIt's nothing, I read this as "Soldering is easy" and then spent about 5 seconds trying to understand why the material is something different from what I expected. Probably should sleep more.
- Edmond 13y agoIs anyone aware of psychological research into why some people are prone to this type of error? This is a problem I have and often find bemusing but sometimes alarming. At least a few times a week I encounter text that I initially grossly misread by automatically inserting my own words instead of what was actually written. I am inclined to think it may have something to do with doing a lot complicated programming. Basically after years of programming and having to deal with detail, you survive by being good at applying triage to problems, ie knowing what to ignore and what to pay attention to. Perhaps misreading text is your brain attempting to apply this type of triage.
- stormbrew 13y agoI don't think it's likely to have anything to do with 'some people' or any kinds of tasks you perform. It's probably more to do with word shapes [1], which are probably at least part of how you read. [1] http://en.wikipedia.org/wiki/Bouma http://en.wikipedia.org/wiki/Bouma
- Stratoscope 13y agoWe are all in this together. I thought Etsy was test-buying some of the kitchen items listed there and finding them painfully unfit for use. (No, really. I'm not trying to be clever. That honestly was my first thought.)
- StefanKarpinski 13y agoI was the one at Adtuitive who chose Cascading.JRuby. It was a pre-existing DSL on top of Cascading. It needed some work, but it was a pretty nice concise way to generate Cascading jobs. However, with half a dozen to a dozen people hacking features into it over time at Etsy, and with no real design coordination, things got pretty out of hand. There were a couple of fundamental problems: 1. Type system mismatch. Ruby is dynamic and not type checked; Java is static and type checked; Cascading is somewhere in between and unfortunately doesn't seem to use Java's type system as well as it could. 2. User-defined functions in strings. For some reason Cascading lets you write user-defined functions as strings and compiles them dynamically during job execution. This was the way to write user-defined code in Cascading.JRuby. Cascading requires a compilation step, yet since you're writing Ruby code, you get get none of the benefits of static type checking. It was standard to discover a type issue only after kicking off a job on, oh, 10 EC2 machines, only to have it fail because of a type mismatch. And user code embedded in strings would regularly fail to compile – which you again wouldn't discover until after your job was running. Each of these were bad individually, together, they were a fucking nightmare. The interaction between the code in strings and the type system was the worst of all possible worlds. No type checking, yet incredibly brittle, finicky and incomprehensible type errors at run time. I will never forget when one of my friends at Etsy was learning Cascading.JRuby and he couldn't get a type cast to work. I happened to know what would work: a triple cast. You had to cast the value to the type you wanted, not once, not twice, but THREE times. Scalding fixes both problems since it is statically typed and lets you write real user-defined functions instead of stuffing them in strings. To me, the main moral is never, ever design an API that involves writing code in strings. It's just bound to be a disaster. Also, if you're going to have a compilation step anyway, you might as well get some static checking for obvious problems like type errors as part of the bargain.
- mcfunley 13y agoOne other unforced error with c.j was having almost every function take a dictionary of args instead of a sane argument list. That left "grep the codebase" and "read the entire function definition and god help you if it passes on the argument dict" as the two horrible options for figuring out how to even call most of the functions. But we probably could have erected a blast shield around that mess if only we could have written functions and aggregators in ruby.
- bkirwi 13y agoI'd like to get in a plug for my personal favourite Scala / Hadoop productivity framework, Scoobi.[0] If you're already familiar with Scala, Scoobi has the more familiar and intuitive API, and it leverages the type system to provide stronger guarantees before the code is even run. (For example, if Cascading / Scalding don't know how to serialize your data you'll get a runtime error; in Scoobi, it complains at compile-time. This is really useful when your compile / deploy / run / error loop may take several minutes...) I've also run into a few bugs in Scalding, while I've found Scoobi to be much more solid. OTOH, Scalding has the edge in terms of community size and ecosystem. I haven't found this to be a big issue -- shimming an existing Hadoop input / output to work with either project is quite simple -- but YMMV. [0] https://github.com/NICTA/scoobi https://github.com/NICTA/scoobi
- dallasmarlow 13y agoI'm really excited to see scoobi listed here as it's something I have really grown to love over the last couple of years. I find that it's very flexible and end up using it for jobs that do more unconventional mapred things.
- avibryant 13y agoYeah, for those same reasons I much prefer using Scalding's typed API [1], which feels very similar to Scoobi. The tuple API shown in these slides is great for places like Etsy that already have a large investment in Cascading, but otherwise you're better off getting the added type safety and similarity to the standard Scala API. [1] https://github.com/twitter/scalding/wiki/Type-safe-api-reference https://github.com/twitter/scalding/wiki/Type-safe-api-refer...
- bkirwi 13y agoI'm familiar with the typed API, but it still doesn't quite bring me as far as Scoobi does. I recognize your name from the Scalding code, so I'll say that this is meant as helpful criticism and not a complaint.* - There a couple types for datasets in the Scalding API: TypedPipe, and KeyedList and subclasses. Scoobi subsumes both of these under DList; thanks to the usual Scala wizardry, this has all the methods to operate on key-value pairs without loss of typesafety. This isn't a huge deal, but it removes the tiny pains of constantly converting back and forth between the two. - Scoobi's other abstraction, DObject, represents a single value. These are usually created by aggregations or as a way to expose the distributed cache, and have all the operations you'd expect when joining them together or with full datasets. You can emulate this in Cascading / Scalding, but it's a bit less explicit and more error-prone. - There's no equivalent to the compile-time check for serialization in Scalding, AFAICT. - Scoobi has less opinions about the job runner itself... there are some helpers for setting up the job, but all features are available as a library. For some reason, I found the two harder to separate in Scalding? - IIRC, Scalding did job setup by mutating a Cascading object that was available implicitly in the Job. In Scoobi, you build up an immutable datastructure describing the computation and hand that to the compiler. This suits my sense of aesthetics better, I suppose... * Also, thanks to you guys for Algebird! That's a really fantastic little project, and I use it all the time.