7 ms·
Bonobo – A data processing toolkit for Python 3.5+
- IanCal 9y agoIt'd be good to see some comparisons, why this and not one of the other currently available systems? Why should I use this over, for example, Luigi? What scale is this intended for? Is it intended to nearly solve a simple problem over my 20TB of data on S3? Big complex graphs? Or more for transitioning a small local report system that's currently in three excel files into a tested python script?
- ah- 9y agoFrom looking at their examples and interfaces, it's clearly for simple, small scale processing.
- rdorgueil 9y agoIt's indeed intended for «small data», by opposition to «big data». I know, that does not say much, but I basically wanted to handle small flux of data without having to install the "big weapons". I'm preparing explanation pages for a lot of the questions I got, including comparisons, volumes of data, where it is good and where it is not ... All that will be well ready before 1.0, but for now, we're at 0.2 ... Thanks for all the hackerlove, though!
- rdorgueil 9y agoWith the ancestor of bonobo, I was processing 5M lines of data in around 1 hour, including extraction, joins, api calls and a few loads. That should give a first info about the size target.
- spangry 9y agoI haven't tried this yet, but am praying that it delivers even half of what it promises. For whatever reason I just can't get my head around pandas, despite multiple attempts. If this also turns out to be inscrutable I may be forced to conclude that I'm stupid...
- unixhero 9y agoTried these video series? That guy explains it really nicely and is really bright. https://pythonprogramming.net/search/?q=pandas https://pythonprogramming.net/search/?q=pandas
- spangry 9y agoThanks for this! It's funny in a way: I'm trying to learn the basics, but don't have a clear idea of what the basics actually are. This looks like it could be just the ticket. Cheers!
- gjreda 9y agoIn case you're looking for more, this tutorial series hit the front page of HN a few years ago: http://www.gregreda.com/2013/10/26/intro-to-pandas-data-structures/ http://www.gregreda.com/2013/10/26/intro-to-pandas-data-stru... Modern pandas is a bit more idiomatic now though: https://tomaugspurger.github.io/modern-1.html https://tomaugspurger.github.io/modern-1.html
- fnord123 9y agoAre you trying to learn pandas just to learn pandas or do you have a motivating example?
- maxerickson 9y agoWhat have you wanted to use it for? Pandas is basically an R data frame for Python. A sloppy description of that is a text mode spreadsheet. The description of Bonobo doesn't immediately invite the comparison to Pandas, to me anyway.
- rdorgueil 9y agoYou're very right, as I'm using both pandas and bonobo for different reasons. Mostly, when I want a quasi-mathematical look over a dataset, pandas is my tool of choice. For all those data pipeline things that reasonably fit on one computer, I do use bonobo.
- rkda 9y agoAll those references to monkeys hurt my head. Bonobos are not monkeys. If they wanted to name it after monkeys, they should've called it Capuchin or something.
- rdorgueil 9y agoNoted, sorry for that. I'll get more infos about bonobos.
- nn3 9y agoThe picture looks more like a Gorilla than a Bonobo too
- rdorgueil 9y agoCurrently realizing that we only have one word in french for both ape and monkeys ...
- e5an 9y agoCame here to post just that. It's called 'Bonobo', there's a picture of a gorilla, and the page keeps saying 'monkey'- as petty as it sounds, you're probably losing potential users to zoological nerdrage.
- rdorgueil 9y ago
- payne92 9y agoI'm trying to figure out if this is "all hat, no cattle". There seems to be a lot of "framework" here, without much core functionality. Stated more precisely: if I'm stitching together things that process data and 'yield' results, why can't I just do that in pure Python? What does this framework add?
- rdorgueil 9y agoShort answer : parralel execution.
- ColanR 9y ago"parallel" :)
- cicero 9y agoRemember, the double l's make parallel lines.
- rs86 9y agoMultiprocessing or muilthreading? Why don't you market it as parallel coroutines processing? That gets me interested. Because there are dozens of frameworks with overgeneral descriptions.
- rdorgueil 9y agoToday, as a default, multithreading. But that's an implementation detail. Actually, Bonobo does not support coroutines (as in asyncio coroutines) so it would be a lie to market it this way. The plan though is to allow to use coroutines/futures in the future, for specific reasons (like long running/blocking operations where keeping output order tied to input order is of no importance). Still, there is a lot on the roadmap before this becomes a priority. I note that I still have a lot of work explaining in simple terms what is actually bonobo, without falling in the trap of "overgeneral description".
- rjurney 9y ago
- maxerickson 9y agoThere's a syntax error in mutate_my_dict_like_crazy at http://docs.bonobo-project.org/en/0.2/guide/purity.html http://docs.bonobo-project.org/en/0.2/guide/purity.html. Seems the documentation is still quite WIP.
- vittore 9y agoInteresting, right now we are using PETL[1] that we used to do with SSIS, bonobo for some reason reminds me of Bubbles library. [1]: https://petl.readthedocs.io/en/latest/ https://petl.readthedocs.io/en/latest/
- ctippett 9y ago+1 for petl, I'm using it right now on a project that deals with a lot of tabular data and it's been a huge time saver.
- zfrenchee 9y agoWho is behind this?
- jredwards 9y agoThe government.
- rdorgueil 9y agoMe (as an individual), and a few great people that helped me along the way. Not commercially endorsed, or supported.
- lookACamel 9y agoHow does this compare to Dask, Luigi or Airflow?
- rdorgueil 9y agoAs soon as I can, I'll include comparison pages to the documentation, trying to keep it as objective as possible. I can't seriously answer this question in depth here, but it is planned, so at least experts from other systems can also jump in and complement/correct my understanding of each systems. I used a bunch of them, but I'm in no mean expert user of each so making it collaborative sound like a better idea than just giving my point of view.
- rcarmo 9y agoHmm. So what would be the advantage over Dask, which lets me scale out over a cluster?
- ziikutv 9y agoAh, interesting. The example on the 'on-boarding' page reminds me of what we used to do at work. We used itertools chains to write producers and consumer to create 'Chain' objects that process data exactly as the bonobo.Graph. Can't wait to try this.
- throwaway_374 9y agoSo how is this different from Airflow, other than Windows compatibility and a lack of dashboard?
- glial 9y agoThat's my question too. I've come to heavily rely on Airflow. As an Apache project now it's becoming mature. From what I can tell browsing the site, Bonobo looks like it's designed to do data processing within the framework. Airflow insists that it's really a task coordinator/scheduler...however, tasks can be Python function calls. So it seems like Bonobo is a specific use case, where Airflow is the more general case (tasks can be SQL queries, bash commands, etc).
- srean 9y agoSweet! generator based utilities for ETL. I think this really a good use of generators and coroutines. Reminds me of stackless based https://bitbucket.org/diji/pypes/src https://bitbucket.org/diji/pypes/src (backing video http://pyvideo.org/pycon-us-2011/pycon-2011--large-scale-data-conditioning--amp--p.html http://pyvideo.org/pycon-us-2011/pycon-2011--large-scale-dat... )
- goodside 9y agoI'm annoyed that I bothered to read the tutorial to this. The TLDR: "Write some generators or functions, put them in a list, and Bonobo will call them all for you in order. Look at the example files for more." The example files are all basic string transformations. The docs are mostly blank pages and missing sections. What little is written has more jokes and conversational tics than information. What does this even do? There's mention of DAGs and different execution strategies if you really dig through the docs, but is that it? If so, why would you use this instead of joblib or some other established parallelism lib?
- ben_jones 9y agoYeah there seems to be a lot of marketing but I found a concise definition on the author's personal website: "extract transform load for python 3.5+". It could be noted that some of the earlier commit messages include "more marketing".
- rjurney 9y agoYeah, why does anyone need something to run some functions in order for them? I can do that, thanks. If it ran them on some... say 'big data' platform, that would be something. As is, this does not deserve to be front page. This is vaporware.
- rdorgueil 9y agoBonobo runs each functions in the pipeline in parallel and make the fifo queues plumbing and thread pool management completely transparent. The TLDR would then be "Write some generators or functions, link them in a graph, and call them in order on each line of data as soon as the previous transformation node output is ready.". For example if you have a database cursor that yields each line of a query as its output, it starts to run the next step(s) in the graph as soon as the first result is ready (yet not stop yielding from database until the graph is done for the current row). I did not find it easy to do with the libraries I tried. The docs clearly lacks completion to say the least, and would need an example with a big dataset, one with long individual operations and one with a non linear graph, so it's more obvious that, of course, it's not made to process strings to uppercase twice in a row. Stay tuned, I'm very happy HN brought it to homepage, did not really think it could happen at this stage though and I understand you. But that's a good thing for the project to move forward.
- teilo 9y agoBefore there was Pandas I wrote a website (using Django) to transform grid data into a denormalized CSV file. In other words, a reverse pivot. Basically it converts multiple header rows and header columns as separate fields for each intersecting value. I've written this basic routine several times over in my career (once in Access VBA!) for different reasons. The current version of it is used to convert a store/item/quantity grid into per-store pick/pack slips. Pandas has a built-in function that can de-pivot a table. I'm not sure it can handle my use case, however, with multiple header rows. Mine also has extra goodies like populating blank row or column values with the previous value in the row-column, among other bizarre features written to grapple with the inconsistent ways our clients make their distro spreadsheets. Trying to break them of their reliance on Excel for this type of planning has proven futile. I'll have to spend some time with Bonobo and Pandas before I take on refactoring our grid tool. It needs a refactor mostly because I'm the only one who understands it. The new data munging libraries would surely simplify some very gnarly logic, and make it accessible to other developers should I get hit by a bus or leave the company.
- madenine 9y agoSeems like sklearn pipelines with a more generalized use case + additional helpful features for ETL. Very interested
- jwilk 9y agoThis name is already taken. https://wiki.gnome.org/Attic/Bonobo https://wiki.gnome.org/Attic/Bonobo
- erwinvaneyk 9y agoSo, what is the advantage of using this over existing workflow management systems, such as Airflow, Azkaban and Luigi?