3 ms·
We recently ran a large clustering job over billions of records (and a few TBs of data on spinning disks) on a single machine with minimal command line tooling
by throwaway556179 6y ago
We recently ran a large clustering job over billions of records (and a few TBs of data on spinning disks) on a single machine with minimal command line tooling in a few hours. Not really optimized yet. People forget how fast modern hardware is and overestimate how much useful data they have (or need).
I think I should start a company around minimalistic data tooling or the like - the amount of waste seems large across the industry.
I saw the de-skilling a few year ago, where a guy stiched together a compete application from a couple of SaaS APIs. Cool, but it somehow does not impress me.
- paulryanrogers 6y agoBare metal can be incredibly fast, if you can get access for a reasonable price. Virtualization is becoming a continuum but the overhead is always there.
- Proven 6y agowhat?? virtualization has negligible overhead.
- MrPowers 6y agoGreat point, r5.metal instances have 96 CPUs and 768 GB of RAM. Lots of "big data problems" can actually be solved with a single big EC2 instance. Cluster computing should always be avoided when a single node will do.
- throwaway556179 6y agoWe had something like 24G of RAM, but we have 500GB RAM machines as well (we own the hardware) and the job would have been even more of a breeze there.
- bastawhiz 6y ago> People forget how fast modern hardware is and overestimate how much useful data they have (or need). Where your model breaks down is when you have folks throughout your engineering org who have data needs but don't have folks to spin up bespoke pipelines like this and spend time optimizing them. You have data, you want to write something SQL-ish or have some nice APIs, and you want to get results dumped into a predictable place. When you have mixed workloads, mixed data sources, and those jobs are being tweaked and changed frequently, the actual underlying compute cost is hardly the issue. Getting the data, chewing on it [fast enough] without having to spend much time optimizing, getting it to its destination, and making that happen regularly and reliably is where the value is.