13 ms·
Introducing Drake, a kind of ‘make for data’
- abraininavat 14y agoWhy Clojure?
- hvs 14y agoProbably because that's one of the languages that they use internally. http://www.factual.com/jobs/oTR1Vfwq/Software-Engineer---Palo-Alto http://www.factual.com/jobs/oTR1Vfwq/Software-Engineer---Pal... Another answer would be, "Why not?"
- aboytsov 14y agoWe love Clojure. Lisp is an extremely powerful language, and Clojure brings all this to the practical JVM world. And Lisp is quite good in operating on lists and graphs, which is a big part of Drake.
- pencilcode 14y agoout of curiosity, why did you go the clojure route instead of the scala route? From what i understand, scala has more libraries available, including ai and nlp libraries but maybe my impression is not correct?
- aboytsov 14y agoIt's hard to compare Clojure and Scala. Scala is a multi-paradigm programming language with strong OOP support and functional support. It's arguably more verbose than Clojure but looks much more similar to Java. Clojure is a Lisp. Lisp stands aside all other programming languages, first of all, because it supports syntactic abstraction (a.k.a. "code is data"). Hardcode addicts (I'm not one of them) say there are only two programming languages - Lisp and non-Lisp. Here's a good comparison of Scala and Clojure: http://stackoverflow.com/questions/1314732/scala-vs-groovy-vs-clojure http://stackoverflow.com/questions/1314732/scala-vs-groovy-v... When we made the decision to switch to Clojure, several things affected it, in no particular order: - we had some people who were already very proficient in Lisp - we liked how expressive and compact it was - Lisp is considered to possess immense expressive power (see http://www.paulgraham.com/lisp.html http://www.paulgraham.com/lisp.html) - we were enamoured by Cascalog (http://nathanmarz.com/blog/introducing-cascalog-a-clojure-based-query-language-for-hado.html http://nathanmarz.com/blog/introducing-cascalog-a-clojure-ba...), and it's written in and for Clojure. This one payed off very well. - Lisp has a reputation of being great at manipulating data: lists, graphs, etc. Here's a good answer from one of our engineers: http://www.quora.com/Clojure/Why-would-someone-learn-Clojure http://www.quora.com/Clojure/Why-would-someone-learn-Clojure As for libraries, both Clojure and Scala are JVM-based, and Clojure has a very good syntax for Java interop, so all Java libraries are available to us. But, of course, Clojure community also spits out libraries like crazy, for example, take a look at this marvel which we use in Drake for parsing: https://github.com/joshua-choi/fnparse https://github.com/joshua-choi/fnparse.
- dirtyvagabond 14y agoanother way to look at this is: we didn't say to ourselves, "we have to be on the JVM, so which JVM language should we use"? instead we asked, "what would be a great language for this project?" and arrived at Clojure based on the above.
- pencilcode 14y agoThanks for your feedback. I've been playing around with both languages, and was leaning towards scala since it seemed more likely i could use it professionally, even though i liked clojure a bit more, sortta like the lisp like syntax.
- SilasX 14y agoSome see a Lisp and ask "why"? I see a non-Lisp that could be replaced with a Lisp and ask "why not?"
- jonathanjaeger 14y agoAm I the only one who immediately thought of Drake the rapper? He's pretty famous, not sure if this was considered during the naming process. Even if it's not a legal problem, it's an SEO/social media problem.
- wickeand000 14y agoAlthough I don't agree that the name "Drake" is an issue, I do find it interesting that an even more apt name for an application of this type might be "Usher"!
- jonathanjaeger 14y agoHa touché
- sehugg 14y agoA drake is a male duck. They were pretty famous back in the day.
- prospero 14y agoAlso a privateer and a mythical beast. I think there's sufficient prior art on this one.
- jonathanjaeger 14y agoTrue, but I wouldn't call my product 'Queen', 'Cream', 'Journey', or another noun that could be confused with someone or something famous. This distracts from the conversation of the product, so perhaps I shouldn't have brought it up.
- yb66 14y agoI think Queen, cream and journey all also have prior uses that outdo the use by the bands ;)
- logn 14y ago
- jboggan 14y agoI really wish that I had a tool like this back in grad school. I was doing bioinformatics work and merging, chopping, and processing various datasets over many months. When a new version of the underlying data came out it was not an easy task to go back and re-process it through dozens of steps in Perl and R. Having a tool like this would have made it a single command to do so and also ensured repeatability and transparency in my data, something which is often sorely lacking in an academic setting. I am one of the data engineers at Factual and though I didn't have a role in creating it I definitely enjoy using it on a day to day basis. You begin to see the utility of it when you have a dozen people working up and down a data pipeline and need to coordinate as product specs evolve or schemas change. I also really like the tagging features - you can add specific tags to different steps in the build and run different "flavors" of your workflow depending upon what is needed. For example, you might build a workflow that collects, cleans, filters, and performs calculations on data from all over the world - but you might also want alternative versions of the build that only work on specific regions or smaller debug datasets. Tags make that really simple to do, even when many steps are shared by the different versions or the dependencies are complicated.
- xaa 14y agoAs a fellow bioinformatician I can agree that this looks quite useful. Although (since you mention R), I wonder why there's no love for R in Drake, given that R is perhaps the quintessential data processing language.
- dirtyvagabond 14y agoThere is love for R in Drake! As of about an hour ago: https://github.com/Factual/drake/commit/f63dd2630ca3e5e4a6a6baa4296d62dcd078690e https://github.com/Factual/drake/commit/f63dd2630ca3e5e4a6a6...
- ori_b 14y agoIt looks like all of the drakefiles could be replaced pretty trivially with Makefiles. Replacing '<-' with ':', ';' with '#', and '$INPUT', '$OUTPUT' with '$<' and '$@', and inserting shell invocations of the Python interpreter looks like it would do the job. The major differences I see are: - Inline support for Python et al. - Confirming the steps that will be taken. - HDFS support. Are there any other big differences?
- dirtyvagabond 14y agoMake was a major inspiration for us, and so Drake definitely has similarities to Make. The differences you list were non-trivial to us in usefulness, but of course YMMV. Also, there are a lot of (possibly) interesting future features described in the spec. https://docs.google.com/document/d/1bF-OKNLIG10v_lMes_m4yyaJtAaJKtdK0Jizvi_MNsg/edit# https://docs.google.com/document/d/1bF-OKNLIG10v_lMes_m4yyaJ...
- JoshTriplett 14y agoMake can support Python, or any other language you'd like. Just set ONESHELL to avoid splitting commands by line, and then set SHELL to your preferred language interpreter. Make will then hand that interpreter the entire body of commands to rebuild a target.
- aboytsov 14y agoDrake supports "protocol" abstraction, which is much more than just specifying an interpreter. Python is a trivial protocol, not much more complicated than shell. There are slightly more complicated protocols, for example, "eval", which runs the first line as a shell command before putting everything else in $CMDS environment variable. There could be protocols for running an HBase query, a Pig query, Cascalog query, or an SQL query. Some of these things could involve building a JAR file and giving it to Hadoop binary. Currently only a handful of protocols is implemented, but more are described in the spec.
- aboytsov 14y agoThe example in the blogpost is understandably trivial, and it can be implemented in almost any Make-like system. The concept of Make is not unique. Everything that has dependencies and executes steps is similar to Make in concept. Drake is no exception, and it can be replaced with Make, but no more so than Rake, Ant or Maven can be replaced by Make. That is, if it's trivial - yes. Just a bit more complicated - no. Some things are merely painful to implement with Make, some are just impossible: - multiple outputs - no-input and no-output steps - HDFS support - Hadoop's partial files support (part-?????) - forced execution of any subbranch, up or down the tree or any individual targets (crucial for debugging and development) - target exclusions - protocol abstraction - inline Python is just one example - tags - branching - methods These are just what's implemented already. Other things are planned such as: - automated data versioning (backup and revert) - parallelization - real-time status console - retries, email notifications - etc. Requirements for building executables and working with large, complicated and expensive data workflows are quite visible different, and the most important thing about Drake is that it provides the platform for convenient features (such as versioning or email notifications) to be implemented. And once they are, every data workflow can take advantage of them. I guess, if Make was really, really extendable, we could have considered it as a platform for all this. But it's not, and hacking all of that into Make's source code in C would be, I'm sure, a much greater pain than writing Drake. Artem.
- moonboots 14y agoDjb redo[1], a make alternative, feels like a good fit for these type of data manipulation and dependency representations. Below is a port of the first example. The build script is just shell, so you can do stuff like embed python with a heredoc. One bit of syntactic sugar is that redo assumes stdout is the desired contents of the generated file, so you don't need to explicitly pipe to an OUTPUT variable. #!/bin/sh case $1 in contracts.csv) curl http://www.ferc.gov/docs-filing/eqr/soft-tools/sample-csv/contract.txt ;; evergreens.csv) redo-ifchange contracts.csv grep Evergreen contracts.csv ;; report.txt) input=evergreens.csv redo-ifchange $input python2 <<-EOF linecount = len(file("$input").readlines()) print("File $input has {0} lines.\n".format(linecount)) EOF ;; esac [1] https://github.com/apenwarr/redo https://github.com/apenwarr/redo
- aboytsov 14y agoPlease see my response to Make comparison: http://news.ycombinator.com/item?id=5111527 http://news.ycombinator.com/item?id=5111527 I suspect most of the points I made would be applicable to redo as well, if not more so. Trivial things don't require Drake. Heck, they often times don't require Make as well - just put it in a linear shell script if the steps are not too expensive. It's when things are getting complicated you need something like Drake.
- moonboots 14y agoRedo lacks features baked into Drake, especially the Hadoop integration, but I believe it would be easier to incorporate custom functionality into redo versus hacking Make or writing a custom build system. I haven't used Drake, so I would be interested in a small but complicated Drake script which tackles an intractable problem in Make. I don't claim redo can provide a cleaner solution than a purpose-built system, but I think it will be unexpectedly simple.
- aboytsov 14y agoThe most crucial thing that Make lacks is multiple outputs and precise control over execution. When you're debugging/developing a large and expensive workflow, you absolutely must have the ability to say things like: - run only this step, I'm debugging it - I've changed implementation of this step, re-build it and everything that depends on it - build everything except this branch, it's expensive and I don't need to rebuild it that often (example: model training) Other examples of intractable problems in Make would be timestamped dependency resolution between local and HDFS files. If Make can't look at HDFS, it can't say if the step needs to be built or not. I don't think you can fix it with external commands. But generally, search for intractable problems is a futile one. Remember, everything you can code in Java, you can code in a Turing machine. :)
- Xion 14y agoThere seems to be few differences between Drake and just rolling out Makefiles for data processing, but I definitely see this project has potential. Distributed processing over AWS/Compute Engine/etc. clusters would be one nice thing to have, as a kind of simpler alternative to Hadoop. I really like the inline, multi-language scripting though.
- aboytsov 14y agoThanks! We feel that in practice, there's quite a lot of differences between Drake and most Make-like systems. See this response for details: http://news.ycombinator.com/item?id=5111527 http://news.ycombinator.com/item?id=5111527
- madMilo 14y agoReminds me of Makeflow: A Portable Abstraction for Data Intensive Computing on Clusters, Clouds, and Grids, Workshop on Scalable Workflow Enactment Engines and Technologies (SWEET) at ACM SIGMOD, May, 2012. https://www3.nd.edu/~ccl/software/makeflow/ https://www3.nd.edu/~ccl/software/makeflow/
- aboytsov 14y agoNice. Surprisingly, we weren't aware of Makeflow and kinda missed it completely. On the first look, it seems like Drake is quite a bit more feature-rich than Makeflow. Please see the designdoc and/or the tutorial video for details.
- jcromartie 14y agoI like the idea that the tasks can be implemented in any language, but I feel like this has limitations compared to something like Rake, where the step definition is code, too. What this means is that in Rake I am not just limited to defining new task bodies, but new ways of defining tasks themselves. I see that Drake is implemented in Clojure, so I'd imagine you understand the value of homoiconicity and extensible languages. So I wonder why you didn't just use Clojure all the way through?
- aboytsov 14y agoThis is a great question. Our approach to this is described here: http://www.youtube.com/watch?feature=player_detailpage&v=BUgxmvpuKAs#t=2393s http://www.youtube.com/watch?feature=player_detailpage&v... In short, we don't feel like it's an either or question. We want to have Drake as a command-line frontend to the core functionality, but we would love to see/have other frontends developed as well. Currently, there's no Clojure DSL for Drake, but I think it'd be totally awesome. The reason we started from command-line is because our workflows are heterogenous, and we also didn't want to limit Drake to developers and associate it with coding. Clojure can be quite a big learning curve if you only need it to specify steps and link them together through file dependencies. We had an important design goal in mind: Drake should be as simple as writing a shell script. If it's not, our experience shows that most workflow start as trivial shell-scripts with one or two steps, and by the time it grows into something unmanageable, it's kinda too late. :) On a related note, Drake supports Clojure code inlining for manipulation of the parse tree. It's not an equivalent, just a somewhat related feature. It allows you to modify the steps, dependencies, and anything else in the parse tree directly from Clojure.
- jboggan 14y agoI'm glad the step definitions are not in Clojure or a unified programming language. It makes it much easier to pull in data specialists, product managers, and other non-engineers to help build and maintain a data workflow while leaving them the autonomy to run and troubleshoot the steps of the build specific to their skillsets.
- madhadron 14y agoI wrote a workflow processing system (http://github.com/madhadron/bein http://github.com/madhadron/bein) that's still running around the bioinformatics community in southern Switzerland, and came to the conclusion that something like make isn't actually what you want. Unfortunately, what you want varies with the task at hand. The relevant parameters are: - The complexity of your analysis. - How fixed your pipeline is over time. - The size of a data set. - How many data sets you are running the analysis on. - How long the analysis takes to run. If you are only doing one or two tasks, then you barely need a management tool, though if your data is huge, you probably want memoization of those steps. If your pipeline changes continuously, as it does for a scientist mucking around with new data, then you need executions of code to be objects in their own right, just like code. Make-like systems are ideal when: - Your analysis consists of tens of steps. - You have only a couple of data sets that you're running a given analysis on. - The analysis takes minutes to hours, so you need memoization. Another Swiss project, openBIS, is ideal for big analyses that are very fixed, but will be run on large numbers of data sets. It's very regimented and provides lots of tools for curating data inputs and outputs. The system I wrote was meant for day to day analysis where the analysis would change with every run, was only being run on a few data sets, and the analysis tool minutes to hours to run. Having written it and had a few years to think about it, there are things I would do very differently today (notably, make executions much more first class than they are, starting with an omniscient debugger integrated with memoization, which is effectively an execution browser). So bravo for this project for making a tool that fits their needs beautifully. More people need to do this. Tools to handle the logistics of data analysis are not one size fits all, and the habits we have inherited are often not what we really want.
- aboytsov 14y agoThank you very much. We're really looking forward to other people using this tool. You raise some interesting points (for example, a frequently changing code), which we ran into as well. Our current approach to it is not as fundamental, and basically includes ability to force re-build any target and everything down the tree and methods, and you can also add your binaries as a step's dependency. I'm sure as we and other people use the tool, we'll have better ideas. For example, Drake could automatically sense that the step's definition has changed and offer to rebuild or dismiss. Other points you raised are also definitely worth thinking about.
- circa 14y agoWhen you run it. It tells you, "you're the fuckin' best, you da fuckin' best."
- gojomo 14y agoI could imagine a bash shell that helps create drake files, by remembering in a richer history structure all files read/modified by subprocesses. (A degenerate drake file, one line per 'step', would almost be a 1:1 representation of this richer history... though you then might want to coalesce and reorder atomic steps to represent the real shape of your workflow and dependencies.)
- aaronjg 14y agoI've spent a lot of time working with pipelining software, first for my last job doing bioinformatics research, and now for handling analytics workflows at Custora. We ultimately decided to write our own (which we are considering open sourcing, email me if you are interested in learning more). The initial system that I used was pretty similar to Paul Butler's technique, with a whole bunch of hacks to inform Make as to the status of various MySQL tables, and to allow jobs to be parallelized across the cluster. At Custora, we needed a system specifically designed for running our various machine learning algorithms. We are always making improvements to our models, and we need to be able to do versioning to see how the improvements change our final predictions about customer behavior, and how these stack up to reality. So in addition to versioning code, and rerunning analysis when the code is out of date we also need to keep track of different major versions of the code, and figure out exactly what needs to be recomputed. We did a survey of a number of different workflow management systems such as JUG, Taverna, and Kepler. We ended up finding a reasonable model in an old configuration management program called VESTA. We took the concepts from VESTA and wrote a system in Ruby and R to handle all of our workflow needs. The general concepts are pretty similar to to Drake, but it is specialized for our ruby and R modeling. Some more useful links for those interested: JUG https://github.com/luispedro/jug https://github.com/luispedro/jug Taverna http://www.taverna.org.uk/ http://www.taverna.org.uk/ Kepler https://kepler-project.org/ https://kepler-project.org/ VESTA http://vesta.sourceforge.net/ http://vesta.sourceforge.net/
- jeffdavis 14y agoCool project. I expected to be underwhelmed, but when I saw the dependency stuff, I was impressed. Maybe it should include a hook so that it can detect dataset changes automatically by running a separate command (or did I miss it?). With a bit of creativity, I think there may be a lot of applications here.
- aboytsov 14y agoThis is an awesome idea. Currently Drake only supports timestamped and forced evaluations, but it would be great to have an evaluation abstraction where you could provide your own implementation of whether a target's changed and/or whether a target is to be considered fresher/younger than another target. Timestamped would compare modification times, forced would return true, and it could be extended indefinitely. If you're serious about it, please submit a feature request (https://github.com/Factual/drake/issues https://github.com/Factual/drake/issues), and describe more specifically what you would like to be able to do in your case. Thank you for a great thought. Artem.
- swalsh 14y agoWhoa, this is the first time i'm hearing of "Factual" but playing around i'm impressed! There was a side project I had a while ago, which i eventually gave up because I couldn't source some data. These guys found it!
- danpalmer 14y agoWith an empty workflow, this is the result of `drake --version`. $ time drake --version Drake Version 0.1.0 Target not found: ... drake --version 5.42s user 0.18s system 188% cpu 2.969 total For short scripts that you should be running in the shell, this is really bad. I expect basic make commands on small projects to be effectively instant. Compilation might take a bit longer, but 5.4s to print the version points to a 5s overhead on all executions. I'm guessing this is due to the JVM overhead, so that pretty much says this project isn't suited to the JVM. The JVM is great for long running processes, and applications where the overhead is a very small percentage of the total running time, but if it takes 5s longer than `make` to print it's version, that's really not a good sign. This is a fantastic idea, and I will definitely be using it. But this overhead needs fixing.
- aboytsov 14y agoHey, thanks for trying out our tool! First of all, --version shouldn't try to run any targets. This seems like a bug. Thanks. Yes, you guessed correctly - this is the JVM startup time. I just hate JVM for that. We experimented with Nailgun and Drip to eliminate it - Nailgun is problematic because it uses a shared JVM for all runs, and it can get quite hairy sometimes. In the long run, Nailgun is almost certainly not an answer, since it assumes things we have no control over (i.e. Clojure runtime) don't do destructive tear down. Drip is a bit more promising, but we didn't succeed running Drake under it (simpler things worked fine though). So, we're still looking into it, and we're looking for other ideas, too. In the meantime, you could run Drake under REPL: (-main "...") The only problem is that Drake calls System/exit but we can add a flag ("--repl") that would prevent it from doing so, and you'll stay in REPL. Thoughts? P.S. JVM is unfortunate but Clojure is a fantastic language for something like Drake.
- danpalmer 14y agoThanks for the detailed and well explained reply. I have limited experience with Clojure, but it does seem to be a good match to this sort of task due to it's structure. However the JVM seems to be a real drawback to me. Perhaps with something like Scheme or Lisp you might get a similar program structure, and be able to compile to faster binaries? The REPL is a solution, but as many developers are using tools like make with many other tools in the shell, running a REPL like that would prevent them from using other things efficiently. Ultimately I think the overhead time needs to be removed. If it takes far longer than something like make, that's not necessarily an issue. The key point is making it fast from the user's perspective. As long as it runs in a fraction of a second, I can't see much of a difference between 0.1s and 0.0001s, so I don't think that sort of difference really matters, it's when it gets over 1s that it becomes an issue. Running something like Nailgun in the background may be a good solution, I don't have any experience with it. But if it requires starting a daemon in the background, that could get in the way of using the tool in a normal way. I don't really know what the best solution to this problem is. I'm not sure Clojure is the best tool for the job.
- daemon13 14y agoArtem, the approach you guys are using is really EXCELLENT! I think that a bit of a disconnect here may be because some OPs might be used to 'compiling' code versus 'compiling' data angle that you are using. This is especially evident by make dependencies discussion with lars512. To give a simple specific example: I have a dataset of say 5000-50000 SKUs that are aggregated across 9-12 dimensions. My final report/analysis uses 3 scenarios. Now one sub-set of one scenario has changed [that's the raw input] - of course running 'data compilation' by using data that changed and ONLY what depends on it is the most effective&efficient approach. Just my 2 financial cents...
- aboytsov 14y agoThank you very much for your kind words and support, and we certainly are looking forward to your feedback, feature requests and bug reports, as well as your code contributions, should you so desire. We built this based on our own pain points with a larger audience in mind. We hope we got some things right, because the success of any tool is defined by its users. So, if you like it, let's build a thriving community together! Artem.
- roolio_ 14y agoKudos for your work! Do you plan to integrate Amazon S3 the same way you did for hdfs?
- aboytsov 14y agoThank you. Why not? We would love to see it, but we're also not actively using Amazon S3 at the moment. But we would be more than happy to review code contributions. First of all, you can file a feature request: https://github.com/Factual/drake/issues https://github.com/Factual/drake/issues Adding a new filesystem to Drake's source is very easy. You just create a filesystem object that implements a bunch of methods for: listing directory, removing file, renaming file and getting file's timestamps, and then put it along with the corresponding prefix in the filesystem map. That's pretty much it. Assuming there's client JAR for Amazon S3, written either in Clojure or in Java, it should be quite simple to do. Artem.
- fnbr 14y agoPerhaps I am the only one having issues here, but I cannot seem to get drake to run. Is there anything that is supposed to be done after building the uberjar? Further, I don't understand how I'm supposed to alter my path to be able to run drake by simply entering 'drake'- would it be possible to get some help? (I'm sorry if this is really obvious)
- aboytsov 14y agoThe project's README file (https://github.com/Factual/drake https://github.com/Factual/drake - scroll down) contains building and running instructions, as well as how to create a simple script to run Drake which you can put on your PATH.
- fnbr 14y agoAh, sorry, I should have been more clear. I've actually gone through the readme a few times, to no avail. I'll triple-check it though.
- aboytsov 14y agoRead the "Installation" section, there's "A nicer way to run Drake" subsection. But I would advise to read the whole "Installation" section carefully.
- fnbr 14y agoI actually did that, several times. My mistake was that I didn't realize I was supposed to have Drake.jar in the same folder as the workflow that I was trying to execute (I'd keep getting the error 'Unable to access jarfile drake.jar'). Naive error, I suppose. However, I'm still having trouble executing the 'A nicer way to run Drake' instructions. I created a file named 'drake' on my path, and inserted the given text. However, I keep getting the error 'Exception in thread "main" java.lang.NoClassDefFoundError: drake/core' Was I supposed to alter the script in any way? I just naively copy/pasted.