10 ms·
> Benchmarking cutting-edge graph-processing algorithms running on 128-core clusters against a single-threaded 2014 Macbook Pro. The laptop consistently wins, s
by CapmCrackaWaka 4y ago
> Benchmarking cutting-edge graph-processing algorithms running on 128-core clusters against a single-threaded 2014 Macbook Pro. The laptop consistently wins, sometimes by an order of magnitude.
LOL, this hits close to home. My company had a modeling specific VM set up to run our predictive modeling pipelines. Typical pipeline is about 50,000 to 5 million rows of training data. At best, using an expensive VM, we managed to get 2x training speed from lightgbm on the VM vs my personal work laptop. We tried GPU boxes, hyper threaded machines, you name it. At the end of the day, we decided to let our data scientists just run models locally.
- christophilus 4y agoI found the same thing when doing video transcoding. The VPSs were all woefully underpowered. Netcup bare metal (root servers) ended up getting pretty close and were by far the best bang for the buck of anything I found.
- bilekas 4y agoCurious what the setup of VPS' was and why you would expect better than real hardware, video transcoding is quite a beast from what I remember and I just can't imagine there's a VPS solution that expects to keep up
- otterley 4y agoThe Intel Xeon processors that cloud providers typically use don't have the Intel Quick Sync core that provides hardware A/V encoding/decoding on typical desktop/laptop CPU SKUs. So the software has to fall back to CPU-based codecs, which are much slower. AWS EC2 has a VT1 instance family that enables high-speed A/V encoding via a Xilinx media accelerator card.
- thom 4y agoThere appear to be slightly weird commercial reasons behind this, because gaming GPUs have great CUDA performance but NVIDIA won’t let you put them in a datacentre. So buying your data scientists gaming laptops (RGB and all) generally works out faster for any reasonable price point. That said, a dedicated server with a decent Xeon and MKL set up correctly generally outperforms CPU-bound stuff.
- CapmCrackaWaka 4y agoI think it really depends on your data size. All the benchmarks I can find are on massive datasets, with tens of millions of rows or thousands of columns. I’m sure there are significant performance gains in these situations. Our data just wasn’t big enough.
- makeset 4y agoWith larger data it really depends on the algorithm. If you must iterate over more than a few GB at a time, GPU memory capacity and bus speeds become prohibitive, while a dead-simple implementation on a single CPU with 100+ cores and TBs of RAM goes brrr.
- lostmsu 4y agoCPU RAM is generally much slower. 8 channels of DDR4-3200 only provide 200GB/s bandwidth. RTX 3090 has 936GB/s. So even 4 socket Xeon won't catch up.
- lmz 4y agoThat would only be an advantage if you had to do multiple passes over the data, otherwise the data would still go through the CPU RAM before getting loaded onto the GPU, no?
- lostmsu 4y agoDefinitely the case in the state of the art stuff like neural networks. Many if not most other algorithms are iterative. Hell, even sorting is.
- VWWHFSfQ 4y agoHaha! Back in ~2014 or so my company was spending nearly $30,000/month on an EC2 "compute-optimized" cluster to transcode live video streams to multiple renditions. One of our engineers said hey, why don't we try to colo some real hardware? We did a test with a single bare-metal 8-core Xeon server and it completely destroyed the performance of the EC2 "compute-optimized" cluster! After that we colo'd 4 big Xeon servers for about $1,600/month total. Looking back on it now it's just so insane...$30,000? no way.
- jamal-kumar 4y agoI don't use AWS for a damn thing because of exorbitant costs. I just don't get why people think that it's necessary other than that they're the types to get drawn into marketing hype. There's just so many better things for your company to be spending the money on.
- redredrobot 4y agoThe argument is - engineers are expensive so why pay for the expertise to setup and run machines? There's just so many better things for your company to be spending the money on.
- jamal-kumar 4y agoWhat if my company is pretty much all engineers? We don't let pencil pushers with MBAs anywhere near what we're doing, and it's going great. I know this isn't the most usual configuration but if undervaluing my skills and trying to bottom dollar on them is going to be their rules, then I'm just going to do my own thing, and they're just going to have to scrape the bottom of the barrel for talent. I hope the zeitgeist changes any time soon. God knows how many unicorns have been sacrificed with that kind of paradigm which could be successful companies by now.
- nemothekid 4y ago>What if my company is pretty much all engineers? I don't know how that changes the equation. No one is undervaluing your skills; it's would you rather spend your time driving to a colo center to replace a RAID array or working on $product. With AWS you are outsourcing an IT team, not just processors and how you approach pricing should reflect that.
- ip26 4y agoThat seems like a problem of matching the workload to the hardware- graph computing isn’t embarrassingly parallel in the regular sense.
- nostrademons 4y agoMy rule-of-thumb is that if you have less than a terabyte of data, you're better off processing it locally, and even that is pretty conservative. Big data is for when you have problem sets that simply will not fit on a single machine. With 4TB hard drives going for about $60, a lot of problems are better solved by simple algorithms in efficient programming languages on a single box. There are some data sets where you really do need big-data tools, but it's for when you have petabyte-scale data, not megabyte/gigabyte-scale data.
- atty 4y agoAlso depends on the complexity of the algorithm (specifically thinking of large neural networks). We have a model that requires 8 A100s for training due to the size of the activations. No way to replicate that on a local machine and have it train successfully in any reasonable time frame. However the complexity of the algorithm many times scales with the size of the dataset, either the full corpus or the size of individual examples.
- bobbylarrybobby 4y agohttps://adamdrake.com/command-line-tools-can-be-235x-faster-than-your-hadoop-cluster.html https://adamdrake.com/command-line-tools-can-be-235x-faster-...
- jamal-kumar 4y agoOh yeah I love simply avoiding memory allocation at all costs and keeping things to the processor cache and streaming APIs. awk/sed is fantastic for this if you're just working with CSV data, but I've done it in my own custom code processing hundreds of gigabytes of JSON in seconds as well. I think data scientists just aren't really hugely concerned with programming optimizations or bottlnecks or whatever. Most of them are just intermediate-level python programmers, and that's completely fine until they think they need a hadoop cluster for whatever they're doing and the costs start piling up.
- siboehm 4y agoI built this decision tree (LightGBM) compiler last summer: https://github.com/siboehm/lleaves https://github.com/siboehm/lleaves It get's you ~10x speedups for batch predictions, more if your model is big. It's not complicated, it ended up being <1K lines of Python code. I heard a couple of stories like yours, where people had multi-node spark clusters running LightGBM, and it always amused me because by if you compiled the trees instead you could get rid of the whole cluster.
- CapmCrackaWaka 4y agoWow, very interesting, thanks for this. Daily batch predictions is all we do. I’m the maintainer of miceforest[1], do you think this would integrate well into the package at a brief glance? I’m always looking for ways to make this package faster. [1] https://github.com/AnotherSamWilson/miceforest https://github.com/AnotherSamWilson/miceforest
- siboehm 4y agoI had a brief look at your package, and my impression was that it's only changing model training. If this is correct then the format of the model.txt (calling `lgbm.save(model, "model.txt")`) is the same as regular lightgbm. This would mean you can use my library for inference.