6 ms·
Presumably the "order" you mention is a primary key to another table, likely one that references the individual items that make up that order, so the data will
by tomwheeler 4y ago
Presumably the "order" you mention is a primary key to another table, likely one that references the individual items that make up that order, so the data will be much larger than you estimate.
It will grow larger still if you include web logs from your e-commerce site and event data from your mobile app so that you can correlate these orders with items that customers considered but ultimately didn't buy. How will your laptop and SSD perform when you then build a user-item matrix to generate product recommendations for each of those 1.2 billion customers?
While plenty of organizations unnecessarily use Big Data tools to store and analyze relatively small amounts of data, there are plenty of customers with enough data to require them. I've seen plenty of them firsthand.
- ilyt 4y agoThat's still well within 1U server with some RAM and bunch of NVMes reach
- 0xB31B1B 4y agoThere are functionally less than 1000 organizations that currently require distributed compute for data analysis. You can get off the shelf AWS units with 1000 cores, terabytes of ram and storage, etc. The cost of compute has decreased faster than the amount of data we have to store and process. What we used to do with spark jobs we can do with python on a single box.
- doug_durham 4y agoCitations please? That's a pretty bold statement to make in the face of observed reality.
- fho 4y agohttps://yourdatafitsinram.net/ https://yourdatafitsinram.net/
- threeseed 4y agoThis is such a lazy response. I/O performance is just one of many characteristics that impact performance and from experience the one you least need to worry about. RAID 0 across multiple high-end NVME drives with OS file caching is going to be more than fast enough for most use cases. The issue is running out of CPU performance and being able to seamlessly scale up/down compute with live running workloads.
- beckingz 4y agoA large computer is radically CPU overprovisioned for most workloads.
- threeseed 4y agoBut we aren't talking about most workloads.
- fho 4y agoBut ... we are ... basically by definition. Vanishing little projects actually need cloud scale infrastructure. And, to address your previous statement: one beefy server is actually pretty scalable. Soft threads spin up in microseconds to serve incoming requests, communication between threads is blazing fast, caching is simpler on one machine, etc. You don't even have to worry to much about scaling, the CPU just throttles itself when there is no load. And every once in a while you just upgrade to the next gen beefy machine.
- beckingz 4y agoEven if this is off by two orders of magnitude and it's only 100,000 companies that need distributed compute, that means that almost all companies just need a single large computer. Looking at the distribution of companies by employee count and assuming that data scales with employee count (dangerous assumption, but probably true enough on average), that means that companies don't need distributed compute until they get several hundred employees. [0] [0] https://www.statista.com/statistics/487741/number-of-firms-in-the-us-by-employment-size/ https://www.statista.com/statistics/487741/number-of-firms-i...
- pocket_cheese 4y agoThis is not true. Any column store database (bigquery, Redshift, snowflake) implements distributed compute behind the scenes. When an analyst/business intelligence people have a query return in 3 seconds instead of 15 seconds, it's actually huge. Not just in aggregate amount of time saved, but in creating a quick feedback loop in testing hypothesizes. This is especially true considering that most analyst type people look at data as aggregates across some dimension (e.g. sales per month , unique visitors per region, etc...) These types of questions are orders of magnitude faster with a distributed backend.
- glogla 4y agoYup. I was just playing with some data from our manufacturing system, about 30 GB. I pulled the data to my laptop (very expensive Apple one) and while it fits on my disk just fine, it took about 15 minutes to download. I imported it to ClickHouse which took a while due to figuring out whatever compression and LowCardinality() and so on. I ran a query and it took ClickHouse about 15 seconds. DuckDB pointed to the parquet files on my SSD took 19 seconds to do the same. Our big data tool took 2 seconds, while working with data directly in cloud storage. Now of course this is entirely unfair - the big data thingie has over twenty times more CPUs than my laptop, and cloud storage is also quite fast when accessed from many machines at once. If I ran ClickHouse or DuckDB on 100 CPU machine with terabyte of RAM it might have still turned out faster. But this experiment (I was thinking of using some of the new fancy tech to serve interactive applications with less latency) made me realize that big data is still a thing. This was a sample - one building from one site, which we have quite a few of.
- ryguyrg 4y agoI'd love to understand the shape of this data and some of the types of queries you're performing. It would be very helpful as we build our product here at motherduck. I have no doubt that there are situations where the cloud will be faster, especially when provisioned for max usage [which many companies do not]. However, there are a lot of these situations even where the local machine can supplement the cloud resources [think re decisions a query planner can make]. Feel free to reach out at ryan at motherduck if you want to chat more.
- threeseed 4y agoLet's assume your completely made-up 1000 organisations claim is true. Right now I work for one of them: a global investment bank. Within that organisation we have at least 100+ Spark clusters across the organisation doing distributed compute. And at least in our teams we have tight SLAs where a simple Python script simply can't deliver the results quick enough. Those jobs underpins 10s of billions of dollars in revenue and so for us money is not important, performance is. So 1000 x 100 = 100,000 teams, all of whom I speak for, disagree with you.
- 0xB31B1B 4y agoDisagree with what? I never said _you_ are a dummy for using distributed compute. There are many good applications for distributed compute. I used spark and flink at a big tech job. The stack worked well for some things, and for others it was a hammer looking for a nail. What you do not see is that for every team that you work with and consider a peer group to you, there are 100 teams that really do not need distributed compute, because they have an org wide infra budget of <3M dollars and a total addressable data lake of less than 1TB, but they are implementing very expensive distributed compute solutions recommended from either a Deloitte consultant or a very junior engineer. Should an IB with an infra budget in the 100M+ infra budget zone use distributed compute solutions, absolutely. There just aren't that many of these orgs.
- crabbone 4y ago> You can get off the shelf AWS units with 1000 cores, terabytes of ram and storage, etc. Hold your horses... the beefiest servers that are in production today, unless you count custom-made stuff go to somewhere between 128 and 256 cores per board. These are hugely expensive. Also, I don't know if you can rent those from Amazon. Typical, affordable servers range between 4..16 cores. Doesn't matter if you buy them yourself, or you rent them from Amazon. It's much cheaper to command a fleet of affordable servers than to deal with a high-end few. This is both because the one-time price of buying is quite different and because with smaller individual servers you have a fighting chance to scale your application with demand. Especially this is true in case of Amazon as you could theoretically buy spot instances and by so doing you'd share the (financial) load with other Amazon's customers. Now... storage. Well, you see, in Amazon you can get very expensive storage that's guaranteed to be "directly" attached to the CPU you rent, the so-called ephemeral storage. This is the storage that's included with the VM image you use. It's very hard to get a lot of it. I couldn't find the numbers for Amazon, instead, I know that Azure tops out at 2 TB. In principle, this kind of storage cannot exceed a single disk, so, think Amazon probably offers the same 2 TB, maybe 4. But, again, it's cheaper to have a bunch of EBS's attached... but then you'll have to have more of them as the latency will suffer, and in order to compensate for that you would try to increase throughput, perhaps. Also, think that, in practice, you'd want to have a RAID, probably RAID5, and this means you need upwards from 3 disks. Also, if you are using something like a relational database, you'd most likely want to put the OS on a single device, the database data on a RAID and database journal on a yet another device, and, probably, you'd want that device to be something like persistent memory / optane / something from higher-tier disks with dedicated power supply. And all this is not due to size, but due to different contingencies you need to have in order to prevent huge data loss... Now, add to this backups and snapshots, perhaps replication in 2-3 different geographical areas if you are running an international business... and that's quite a bill to foot. There are similar problems with memory, since there can only be so many legs on memory bus and only so many pieces of memory you can attach to a single CPU, and if you also want a lot of storage, then, similarly, there can be only so many individual storage devices attached and so on. Bottom line... even to reproduce the performance of your laptop in the cloud you would probably end up with some distributed solution, and you would still struggle with latency.
- deltarholamda 4y agoDon't forget the cool JS library you included to track mouse movements so you can optimize your UI to make sure Important Money Making Things are easily clickable. That's 8.4 hojillion megabytes per second right there.