5 ms·
Agreed, I am amazed at how often I hear of a few million rows being considered a large table.
by Guest42 5y ago
Agreed, I am amazed at how often I hear of a few million rows being considered a large table.
- PTOB 5y agoIt's not the rows that kill you. It's the columns.
- Guest42 5y agoTrue, and improper indexing.
- at_compile_time 5y agoWhy is that?
- bigiain 5y agoRows with an integer id and a few integer or bit flag columns or maybe a short short fixed width char column or two and a few well though out indexes, is a way different beast compared to something with dozens of big varchar() or blob columns (or worse still, json or xml text columns with un-indexable key/values inside that you need to parse out before you can run selects on them...)
- at_compile_time 5y agoAh, I think I understand now. It's not the amount of data (rows), it's how that data is structured (columns). I thought you were saying that number of columns was more important than number of rows, which clashed with my naive intuition.
- zozbot234 5y agoModern RDBMS's can support JSON and XML as native data types, i.e. you absolutely can index within them and they have facilities to speed up the parsing.
- bigiain 5y agoYeah. But not if you've just shoved them into text columns.
- onemoresoop 5y agoLets not kid ourselves, a few million rows is not small either and if not done properly can quickly become a system's bottleneck. But it's not humongous either. Depending on ones perspective it could be seen as medium or large but certainly not small.
- nomel 5y agoI think it's safe to call anything that fits within the memory of a mid range notebook computer "small". I guess it depends on how wide those tables are.
- bigiain 5y agoThe data scientist I've worked with that I most respect once told me "It's not 'big data' if it fits on your laptop". Anything that fits in ram, at least in the context of "a table" in a RDBM probably counts as 'micro data'...
- bigiain 5y agoAnybody who has a "bottleneck" caused by a table with a few million rows is functionally incompetent. From any half-modern database's perspective that's "tiny". Any reasonable production machine would be able to fit (at least the important parts of) a million row table in memory. Even if this is your hobby project running on a RaspberryPi Zero, you'd have at least the index in ram. Either you don't care about that query's runtime (a perfectly valid approach if your use case is "this creates a report after COB on Friday that needs to have completed before 9am Monday morning"), or you need someone who knows what they're doing take over from you (also a valid approach if this is in your tech demo or POC, and you're the Technical Founder who's now out of their depth and needs a proper engineer to build a scalable product for you).
- hnthrowaway0315 5y agoI mean Excel can handle a few million if use Power Pivot.
- pdpi 5y agoA million 1k-wide rows is still just 1GB. You can go double that width and triple the rows, and you still comfortably fit in memory on a raspberry pi. I don’t know about “big” or “small”, but the threshold for “done properly” has to be really low for that to become a bottleneck.
- runnerup 5y agoThis thought has a very narrow perspective. During the space program, a few million rows was very large. Today, you have better tools so a few million rows is small to you. But someone else who works on Fugaku has tools that you don't have and will find that your "large" amount of data is small to them -- they also get to use the electrical equivalent of 20,000 homes to process that data. Most people today don't have the computing tools that you do. Yes, a consumer laptop running Numpy can process it quickly. But they have Excel, not Numpy...and Excel cannot process millions of rows. So in the context of the tools they have, it is a large table.
- hnthrowaway0315 5y agoIMHO when discussing "Big Data" you have to go up PB (in total) and some large numbers (in daily) to reach "Big Data". If the data can be stored inside a commercial Xeon box then it's not even "Medium Data".
- dylan604 5y agoto me, Big Data is more of a concept than just a description of the amount of data. So ontop of the storage, it's also the analysis and usage of that data. So Big Data services could be just as useful to a small company with a mere 100GB of data once they learn how to squeeze the juice from the berries they've collected.
- funcDropShadow 5y agoBut, if it is 100 GB of data, Big Data tools are the wrong tools, because they are inefficient at that size. Most Big Data come with lot's of restricitions and complications which are fine if they are the only alternative. But don't ever think about using Cassandra instead of Postgres if you have less than 10 TiB. Even over 10 TiB it is not clear when to switch to something more scalable.
- dylan604 5y agoAgain, this comment is focused on the size of the data on storage and how much data there is. To me, this is just a single aspect of "Big Data". To me, "Big Data" is also what information can be gleaned from that data that is being stored whether it be Postgres, NoSQL, Cassandra, etc. That part of it is only relevant in conversations about "how much data" is available to the "Big Data" processes. People paying for "Big Data" are excited about "how much data" they have other than a bragging point. What they are paying for is the information that can be garned for having that data.