4 ms·
I don't have very much background in ML or distributed systems, so forgive my naive questions... > After all, most ML in industry today seems to be lightweight
by int3 4y ago
I don't have very much background in ML or distributed systems, so forgive my naive questions...
> After all, most ML in industry today seems to be lightweight models applied to heavily engineered features
I assume "lightweight models" are those that don't have too many parameters, and "heavily engineered features" mean that the data fed into the model has undergone significant pre-processing via potentially complicated UDFs -- hence the motivation for the project. Is that right?
> Quokka is an open-source push-based vectorized query engine ... it is meant to be much more performant than blocking-shuffle based alternatives like SparkSQL
Does anyone have pointers to what push-based vs blocking-shuffle engines are? Any good papers?
> It should work on local machine no problem (and should be a lot faster than Pandas!)
So I understand why Quokka is faster than Spark, but I'm a bit uncertain as to why the author is also making a comparison with Pandas on a single machine. Is it because the streaming pipeline design means that Quokka can better take advantage of multiple cores?
- marsupialtail_2 4y agoThat's right. My background is mostly in quantitative finance, where we would use models like linear regression on expert-engineered features based on market data, instead of throwing a deep neural network at raw price data like what some people might imagine. For push vs. pull, I'd recommend: https://news.ycombinator.com/item?id=27006476 https://news.ycombinator.com/item?id=27006476. On single machine, you really should just use Polars. Quokka is faster than Pandas because it can take advantage of multiple cores, but so can Polars -- and it is likely to be faster.
- int3 4y agoThanks for the answers! And the push vs pull link is a great explanation indeed :)