27 ms·
Big data is dead
- andreygrehov 4y agoWith all the LLM craziness, this is just the beginning. How else are they going to train all those models? I'm not an expert, just imho.
- fijiaarone 4y agoSomewhere along the line people were tricked into thinking that logging was data, and that to we needed to turn up every trace log to 11 on every production system. Logs are where data goes to die.
- pelatimtt 4y agoAgree. And thing I noticed is that tools like #apache spark have become the de-facto standard for any data engineer work even when data size does not require it. Result is that many jobs are much harder to mantain and often slower (due to all the shuffling) than running on a single node.
- dbjt_baki 4y agoWell then if businesses do not require data, then the "AI world" might need some. So changing practice to be a machine learning engineer might not seem too bad.
- singularity2001 4y agoBig Data lives on in LLMs.
- meindnoch 4y agoGood riddance.
- guardiangod 4y agoThere is literally a post on front page on ChatGPT, and Microsoft and Google are preparing to duke it out starting in the _next 2 days_ over big-data generated 'chat' result. Big data was never going to be useful to even medium size enterprises, unless anyone can get public access to PBs of data, but that doesn't mean big data is dead. ChatGPT is literally changing how school will test their students, for a start. Maybe what the author is trying to say is 'small-scale big data is dead, but big data chugs on.'
- eppp 4y agoI kind of doubt they trained chatgpt on petabytes of application logs and web server logs. Is keeping all of this crap even useful for more than a small amount of time at this scale? Actual good information will always be useful, most of this "big data" seems to be the equivalent of recording background static.
- miguelazo 4y ago>ChatGPT is literally changing how school will test their students, for a start. Sure, instead of schools checking for plagiarism from other students' papers using turnitin.com, they'll check for plagiarism using ChatGPT tools that scan for known output from their industrial-scale amalgamation of plagiarized materials. Big whoop.
- humanizersequel 4y agoAll math homework through the high school level is now as simple as figuring out how to describe it to ChatGPT (or maybe ChatGPT 2.0 for particularly tricky examples). Paper-writing is now a matter of figuring out how to rephrase LLM output in your own words to get around any watermarking or pattern detection.
- eganist 4y agoWolfram alpha has been around for math cheats (and people like me who just needed a more visual representation to learn) for a while now. Including proof of work.
- LarryMullins 4y agoYears before wolfram alpha, we had TI-89s with computer algebra systems for cheating your way through highschool math.
- eganist 4y agoOh yeah, it's why most of my classes restricted us to TI-83s. The TI-89 was restricted in schools to basically calc and above, and the TI-92 was just banned. Lol
- alexpetralia 4y agoI am writing an essay series on this topic: last-mile analytics and how an abundance of data must be ultimately converted into (measurably correct) action. If anyone wants to follow along, the series is here! https://alexpetralia.com/2023/01/19/working-with-data-from-start-to-finish/ https://alexpetralia.com/2023/01/19/working-with-data-from-s...
- imachine1980_ 4y agoThis sounds like "sane planning, sensible tomorrow." Book for Al gore
- blakeburch 4y agoThat looks like a huge undertaking, but kudos for taking the time. I'll be following along. Totally agree that all data should be tied to the business value that it's driving. Unfortunately, I've found that many data teams focus more on making the data clean and available. They never drive the conversation about what actions are being taken with the data. That leads to them being treated as cost centers. Wrote a similar post about my perspective on it - https://bytesdataaction.substack.com/p/transform-your-data-team-into-a-performance https://bytesdataaction.substack.com/p/transform-your-data-t... I'd love to chat about the space more with you if you're interested! Email in bio.
- danuker 4y agoI agree with many of the points here. My cheap no-name old laptop SSD writes with 170MB/s. A customer has a name, address, email and order. Let's say 200 bytes for each. That means I can write 844000 new customers per second, far outside my personal marketing reach. My disk is 240GB, which means I can store data for 1.2 billion customers. It'll take a while until I become that successful.
- tomwheeler 4y agoPresumably the "order" you mention is a primary key to another table, likely one that references the individual items that make up that order, so the data will be much larger than you estimate. It will grow larger still if you include web logs from your e-commerce site and event data from your mobile app so that you can correlate these orders with items that customers considered but ultimately didn't buy. How will your laptop and SSD perform when you then build a user-item matrix to generate product recommendations for each of those 1.2 billion customers? While plenty of organizations unnecessarily use Big Data tools to store and analyze relatively small amounts of data, there are plenty of customers with enough data to require them. I've seen plenty of them firsthand.
- ilyt 4y agoThat's still well within 1U server with some RAM and bunch of NVMes reach
- 0xB31B1B 4y agoThere are functionally less than 1000 organizations that currently require distributed compute for data analysis. You can get off the shelf AWS units with 1000 cores, terabytes of ram and storage, etc. The cost of compute has decreased faster than the amount of data we have to store and process. What we used to do with spark jobs we can do with python on a single box.
- doug_durham 4y agoCitations please? That's a pretty bold statement to make in the face of observed reality.
- ThereIsNoWorry 4y agoBig Data is dead? Seems well and alive to me. If you're not a big company with big customers, it never affected you to begin with.
- dig1 4y agoBig Data is far from dead. On the contrary, people (on most daily projects) are more mindful now wrt all Big Data liabilities and benefits (infrastructure cost vs. what you get from it) thanks to the experience of the failed ones. But many analytics companies are thriving. Also, using BigQuery as a metric of how Big Data is used is, IMHO, wrong. Real analytics companies usually have custom solutions because BigQuery is too expensive for any serious usage unless you are Google.
- miguelazo 4y agoOn to the next hype theme(AI)!
- sgt101 4y agoLooks at 15 hr Spark job (running since this morning) Sighs...
- lucidguppy 4y agoSome of mongo's leveling off is the adoption of good jsonb columns in postgres. mongo's got sharding out of the box - which is nice - but you have to get your key right or it will suck. Also no one should want to host a mongo db - unless that's your business.
- threeseed 4y agoMongoDB grew revenue 52.8% in the previous financial year [1]. And if there is any levelling off it's going to be because of the move towards cloud managed options e.g. Snowflake, DocumentDB rather than because PostgreSQL decided to add JSONB support. [1] https://www.macrotrends.net/stocks/charts/MDB/mongodb/revenue https://www.macrotrends.net/stocks/charts/MDB/mongodb/revenu...
- nicklaf 4y ago"Shards are the secret ingredient in the webscale sauce": https://www.youtube.com/watch?v=b2F-DItXtZs https://www.youtube.com/watch?v=b2F-DItXtZs
- fdgsdfogijq 4y ago"For more than a decade now, the fact that people have a hard time gaining actionable insights from their data has been blamed on its size." The real issue is that business people usually ignore what the data says. Wading through data takes a huge amount of thought, which is in short supply. Data Scientists are commonly disregarded by VPs in large corporations, despite the claims about being "data driven". Most corporate decision making is highly political, the needs of/whats best for the business is just one parameter in a complex equation.
- revolvingocelot 4y agoIt's absolutely this. "Decision-based evidence-making" is what I've seen it called.
- capableweb 4y agoIs that what's happening at Amazon as well? As they seemingly is loosing more and more track of the "Customer Obsession" schtick.
- aintgonnatakeit 4y agoThey are encouraging their customers to have a bias toward action. Away from that asshole Bezos.
- fuzzylightbulb 4y ago"customer obsession" was always at the mercy of the real obsession: "making money hand over fist". The former will ALWAYS lose out to the latter given enough cycles.
- ralph84 4y agoBig Data got replaced by Big Parameters.
- hgsgm 4y agoParameters come from data.
- rvieira 4y agoWhat about IoT?
- glogla 4y agoI agree with a lot of the sentiments of the MotherDuck people, but boy are they loud and proud for someone who never delivered anything more than blogposts and vague promise to somehow exploit the MIT licensed DuckDB. Meanwhile for example boilingdata.com seems to have already done that - by using AWS Lambda + DuckDB as distributed compute engine which I can't decide if its awesome, deranged or both.
- mytherin 4y agoWe (the DuckDB team) are very happy working together with MotherDuck in a close partnership [1]. [1] https://duckdblabs.com/news/2022/11/15/motherduck-partnership.html https://duckdblabs.com/news/2022/11/15/motherduck-partnershi...
- travisgriggs 4y agoI've made anecdotal observations similiar to this over the last 10 years. I work in AgTech. A big push for a while here has been "more and more more data". Sensor-the-heck out of your farm, and We'll Tell You Things(tm). Most of what we as an industry are able to tell growers is stuff they already know or suspect. There is the occasional suprise or "Aha" moment where some correlation becomes apparent, but the thing about these is that once they've been observed and understood, the value of ongoing observation drops rapidly. A great example of this is soil moisture sensors. Every farmer that puts these in goes geek-crazy for the first year or so. It's so cool to see charts that illustrate the effect of their irrigation efforts. They may even learn a little and make some adjustments. But once those adjustments and knowledge have been applied, it's not like they really need this ongoing telementry as much anymore. They'll check periodically (maybe) to continue to validate their new assumptions, but 3 years later, the probes are often forgotten and left to rot, or reduced in count.
- barathr 4y agoClassic paper on soil moisture sensors (from 2010!) -- the title says it all: "Mate, we don't need a chip to tell us the soil's dry" https://doi.org/10.1145/1858171.1858211 https://doi.org/10.1145/1858171.1858211
- chudi 4y agoMost of the time this story is true, but think this way, the person that was using the system was an expert on the subject. If you can replace the expert with a person just looking at a graph from time to time to know if you have to irrigate the soils it's a different thing. Most of the data or ML tools show us something that the client as an expert already knows, but the true power of this tools is to give them to a non expert user and have roughly the same level of proficiency
- ladyattis 4y agoI think there's a problem at the heart of the matter, specifically the idea that the act of measurement is in itself powerful when in point of fact that this isn't universally the case. As the old adage goes: "garbage in, garbage out." Even more troubling, there is a physical limit to our ability to model what we measure. Take the retina, it has around a million light receptors and even if you assumed they only have two valid states then you're left with around 10^300,000 bits of information to process, so good luck with that. Same thing applies to whatever firms are measuring and what they think is conveying relevant information as they'll have similarly exponential increases if they don't filter out the vast majority of irrelevant data points and states.
- juujian 4y ago> Most data is rarely queried Right on point. In the past I have been obsessed with big data, looking for insights. Then I realized that a medium-sized specific data set is always better than a gargantuan general big data monster. There is so many applications in my field where only outliers matter anyways, and everything is very "centralized" to a few relevant observations. So the only thing about big data is that you maybe throw away 99.9% of the data right away and then you have some observations that you actually care about. There is soooo much data out there that is just noise, and so little that I actually care about. And that's why I still end up hand collecting stuff every now and then.
- donretag 4y agoMy personal definition of Big Data has always been when you gather/store data without having a planned use for it. Do we need this data? Don't know, let's just store it for now. The article does allude to this definition when it states that "Most data is rarely queried". We have become data hoarders. Technology has made it easy (and relatively cheap) to store data, but the ideas of what to do with this data have not scaled in comparison.
- ankrgyl 4y agoI love DuckDB and am cheering for MotherDuck, but I think bragging about how fast you can query small data is really no different than bragging about big data. In reality, big data's success is not about data volume. It's about enabling people to effectively collaborate on data and share a single source of truth. I don't know much about MotherDuck's plans, but I hope they're focused on making it as easy to collaborate on "small data" as Snowflake/etc. have made it to collaborate on "big data".
- andix 4y agoI see it all the time: people develop applications that will never ever get a database size of over 100GB and are using big data databases or distributed cloud databases. Often queries only hit a small subset of the date (one customer, one user). So you could easily fit everything into one SQL database. Using any of the traditional SQL databases takes away a lot of complications. You can do transactions, you can query whatever you want, … And if the database may get up to 1TB, still no problem with SQL. If exceed that, you may need a professional OPs team for your database and a few giant servers, but they should easily be able to go up to 10 TB, offload some queries to secondary servers, …
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- tootie 4y agoI think a lot of data tech has come full circle is now mostly just relational databases. Our org is invested in redshift which lets us mostly pay as we go. The DB itself is just a Postgres facade on scalable storage with some native connectors to file stores and third-parties. After rolling over our stack like three times, we're now just dumping tons of raw data into staging tables, then creating views on top of them. It's 97% raw SQL with a smattering of python for clunky extractions. And we're now true believers in ELT vs ETL.
- threeseed 4y agoRedshift with S3 storage is no different to Spark SQL with S3 storage. Both are distributed compute. Except that Spark allows you to mix/match code with SQL.
- primax 4y agoI think a key driver of this is not having to use SQL. I like DynamoDB and EdgeDB because I can use a more modern and reasonable language to interact with the database.
- bfrog 4y agoThis reminds me of a great blog post by Frank McSherry (Materialize, timely dataflow, etc) talking about how using the right tools on a laptop could beat out a bunch of these JVM distributed querying tools because... data locality basically. https://github.com/frankmcsherry/blog/blob/master/posts/2015-02-04.md https://github.com/frankmcsherry/blog/blob/master/posts/2015...
- cubefox 4y agoThis is a bit ironic given that generative AI models like GPT-3 and Dall-E only work because they were trained on very large datasets.
- fredliu 4y agoThe title might be hyperbole (intentionally), but the observations are more or less in line with what I experienced through a few the Big Data initiatives over the years under different enterprise environments (although I have reservation about the one 1%er comment). To me, Big Data was never about how "big" the data was, but more about the tools/system/practice needed to overcome the limitation of the previous generation. From that perspective, yes, the "monolith" may be having a "coming back" for now due to the improvement of underlying single node performance. But I do think Data size will keep growing, everything needed to make Big Data work would still be there when the pendulum swings back where a single node can't handle it anymore.
- Agingcoder 4y agoI remember the big data craze. People had very little data and low quality at that so they had a data problem before they had a big data one!
- mejakethomas 4y agoYes! This!!! Volume != Quality
- spaintech 4y agoNot that big data is dead, more like real time data is coming to life, but you need the old stuff around to make a buck or two… Well, that my view. LLMs are transformer model technique are making data more relevant than ever. If you are a business, well you are in for a “now real” digital transformation. Making data the centerpiece of your business business could mean that your effectiveness of business process could increase several order of magnitudes. Funny thing is, you will not use some else’s model, unless you are building a ChatBox to infer, but you will need to build your own model and be trained in your own business process to be successful. Consider a bank, here is my prediction of expected outcomes: Enhanced Customer Experience: The system can act as a virtual banking assistant, providing customers with instant access to their account information, real-time transactions, and balance updates. The system can also answer customer inquiries and provide relevant information, improving the overall customer experience. Improved Fraud Detection: The system can monitor the bank's financial transactions in real-time and identify any potential fraud, helping the bank reduce its exposure to financial losses. Automated Loan Processing: The system can analyze loan applications, credit scores, and other relevant data to approve or reject loan applications in real-time, reducing the time and effort required for manual loan processing. Personalized Marketing: The system can analyze customer behavior, transaction history, and demographic information to provide personalized marketing and cross-selling opportunities, increasing the bank's revenue and customer loyalty. Real-Time Insights: The system can provide real-time insights into the bank's financial performance, customer behavior, and market trends, enabling the bank to make informed decisions and respond to market changes quickly. What is interesting to me is, this is just the beginning of what could be…
- mr_tristan 4y agoYeah, I've noticed more applications just need to focus on making sense of raw information really quickly, but usually don't need an archive to make decisions. There are lots of interesting things that can happen with "big streaming" than necessarily "big data". Like, cybersecurity is evolving to monitoring and reacting what everyone's machine is doing in the last 15 minutes, instead of having a huge database of hashes you trust. But not a ton of things really utilize what happened, say, 10 years ago on people's machines. There's definitely some things that can use massive archives of old data, but I have found far, far fewer things that would benefit from it, and often that comes with some very big maintenance hassles. Most of the time, you can just set data retention to 30 days and be done.
- idlewords 4y agoPretty funny to see this when every other headline on this site is about how large language models are about to revolutionize dentistry, beekeeping, etc.
- moooo99 4y agoI feel like big data has rarely lived in most organizations. My own experience working in large orgs largely supports the point that collected data is rarely queried. But this is rarely due to a lack of interest, it is mostly because a) nobody really has a great overview over what even is collected b) even if you know/assume something is collected, you usually have no idea where c) if you find the data, there is a decent chance that it is in some sort of weird format that requires a ton of processing to be usable. This has been - to varying extends - my own experience working in large organizations that don't have tech as their core business. Although there are some successful data analysis project, the potential of the collected data remains largely underutilized.
- posharma 4y agoWe're going to reach a point where we might say the same thing about large language models. Fine tuned LMs (based off of their large parents) are going to be the bread and butter.
- luckydata 4y agoIt's kinda weird to read this. The whole argument is "we didn't have databases that could handle the sizes and use cases emerging, we worked on the problem for 20 years and now it's no biggie". Mission accomplished more than big data is dead IMHO.
- pier25 4y ago> Are you in the big data one percent? Exactly, and I'd go further. Are you in the perf/scale/data one percent? So many people worry about scaling when in reality 99% of web apps will never reach above 100reqs/s. I've been in web dev for 20+ years. Only once when working for a big international corporate client I had to worry about traffic spikes. And that was just for one of their multiple web apps.
- articsputnik 4y agoI love DuckDB's simplicity and think it will solve many problems. Still, transitioning from a local single file DB to concurrent updates and serving it online will be different. I'm curious about what MotherDuck will come up with to solve DuckDB at scale. I love use cases like the Rill Data (https://youtube.com/watch?v=XvP2-dJ4nVM https://youtube.com/watch?v=XvP2-dJ4nVM), where you can suddenly run analytics with a single cmd line prompt and see your data just instantly visualized. Such use cases are only possible because of the "small" data approach that DuckDB tries.
- low_tech_punk 4y agoLong live Big Model, I guess? Instead of independent data warehouses, we are now moving towards a few centralized companies using supercomputer in physical data centers. The "winner takes all" effect will only increase as the trend goes on.
- AaronBBrown 4y agoThe truth is that most "big data" problems aren't big and can often be solved with awk and xargs.
- carlineng 4y agoMotherDuck has been making the rounds with a big funding announcement [1], and a lot of posts like this one. As a life-long data industry person, I agree with nearly all of what Jordan and Ryan are saying. It all tracks with my personal experience on both the customer and vendor side of "Big Data". That being said, what's the product? The website says "Commercializing DuckDB", but that doesn't give much of an idea of what they're offering. DuckDB is already super easy to use out of the box, so what's their value-add? It's still a super young company, so I'm sure all that is being figured out as we speak, but if any MotherDuckers are on here, I'd love to hear more about the actual thing that you're building. [1]: https://techcrunch.com/2022/11/15/motherduck-secures-investment-from-andreessen-horowitz-to-commercialize-duckdb/ https://techcrunch.com/2022/11/15/motherduck-secures-investm...
- danielmarkbruce 4y agoDeliberately speculating so someone will correct it: I'd guess they'll make a bunch of enterprise tools to do things like: enable access and synch the data in a way which complies with various policy, encrypt/tokenize/hide certain columns etc, monitor queries, ensure data is encrypted at rest, stuff like that. Assuming the above it true: I'll bet the reason they aren't so loud about exactly what they are doing is they want to get a head start on it. In theory anyone can build this stuff around DuckDB. From a marketing perspective the clever thing to do would be drive up usage of DuckDB while they build out all this functionality and then the minute corporates start seeing problems with their people using it (compliance etc), they have the solutions.
- itamarst 4y agoThis is an excellent summary, but it glosses over part of the problem (perhaps because the author has an obvious, and often quite good solution, namely DuckDB). The implicit problem is that even if the dataset fits in memory, the software processing that data often uses more RAM than the machine has. And unlike using too much CPU, which just slows you down, using too much memory means your process is either dead or so slow it may as well be. It's _really easy_ to use way too much memory with e.g. Pandas. And there's three ways to approach this: * As mentioned in the article, throw more money at the problem with cloud VMs. This gets expensive at scale, and can be a pain, and (unless you pursue the next two solutions) is in some sense a workaround. * Better data processing tools: Use a smart enough tool that it can use efficient query planning and streaming algorithms to limit data usage. There's DuckDB, obviously, and Polars; here's a writeup I did showing how Polars uses much less memory than Pandas for the same query: https://pythonspeed.com/articles/polars-memory-pandas/ https://pythonspeed.com/articles/polars-memory-pandas/ * Better visibility/observability: Make it easier to actually see where memory usage is coming from, so that the problems can be fixed. It's often very difficult to get good visibility here, partially because the tooling for performance and memory is often biased towards web apps, that have different requirements than data processing. In particular, the bottleneck is _peak_ memory, which requires a particular kind of memory profiling. In the Python world, relevant memory profilers are pretty new. The most popular open source one at this point is Memray (https://bloomberg.github.io/memray/ https://bloomberg.github.io/memray/), but I also maintain Fil (https://pythonspeed.com/fil/ https://pythonspeed.com/fil/). Both can give you visibility into sources of memory usage that was previous painfully difficult to get. On the commercial side, I'm working on https://sciagraph.com https://sciagraph.com, which does memory and also performance profiling for Python data processing applications, and is designed to support running in development but also in production.
- edpichler 4y agoI believe we are living in the "emotional era", so data has being ignored and 'feelings' come first when making decisions or creating processes. This is happening not only in companies but in our current society in general.
- mordechai9000 4y agoPerhaps I'm somewhat cynical, but I believe this is a feature of the human condition, not an attribute of our age in particular. Reason and analysis are tools that are used to justify what we already believe.
- maxfurman 4y agoAgreed! The so-called "Age of Reason" was the anomaly, and probably not that much more reasonable than our own time.
- tootie 4y agoI think there's absolutely a place for this. I often of the old Henry Ford quote about people wanting faster horses. Data and analytics are great for optimization, but sometimes you need to trust your gut and give people something they didn't ask for to have a breakthrough.
- jacobsenscott 4y agonosql is dead, client side SPAs are dead. Nice to see the complexity pendulum swinging back to the correct side again. Curious what the merchants of complexity will reach for next. Are applets going to be the new hot thing?
- Flatcircle 4y agoSeems like just yesterday, every business magazine's cover story was about "big data." Wonder what the next batch of business buzz words will be?
- anon223345 4y agoLong live big data!
- therealbilly 4y agoI think server hardware solved the big data issue. The stuff we have now can blitz through data in the blink of an eye. For national governments like our own, mainframes still have a place. For me personally, I don't even talk about big data anymore.
- revskill 4y agoMain goal of Big Data as i see is to profile performance and metrics. Number of user registration, number of converted users,...
- datan3rd 4y agoDetailed web event telemetry is where I have seen the "biggest" data, not application-generated data. Orders, customers, products will always be within reasonable limits. Generating 100s of events (and their associated properties) for every single page/app view to track impressions, clicks, scrolls, page-quality measurements can get you to billions of rows and TBs of data pretty quickly for a moderately popular site. Convincing technical leaders to delete old, unused data has been difficult; convincing product owners to instrument fewer events is even harder.
- kthejoker2 4y agoSo the argument is you can do everything with an OLAP Database because we shrunk "Big Data" back inside RAM? K, good luck!
- H8crilA 4y agoBig data starts somewhere around a petabyte, maybe a bit lower than that. That's when you need some serious, dedicated systems. But as always everyone wants to (pretend to) do what the big players do.
- zzzeek 4y ago> The most surprising thing that I learned was that most of the people using “Big Query” don’t really have Big Data. wow, ya think? Must have been eye opening to see all those customers with a few million rows thinking they had "Big data" huh?
- wizwit999 4y agoPerhaps this is true for business data (though I'm skeptical of the claims), but, for example, for security data, this isn't true at all. Collecting cloud, identity, SaaS, and network logs/data can easily exceed hundreds of terabytes. A big reason why we're building Matano as a data lake for security. It seems an odd pitch in general to say, hey my product specifically performs poorly on large datasets.
- CobrastanJorji 4y agoOn the contrary, identifying what your product is explicitly not aiming to do is extremely helpful. "Big" adds a lot of complexity and pain, most people don't do that, our product avoids the complexity and pain and is the best choice for most people. Seems like a good, simple pitch, and all it requires is the humility to say that your solution isn't the best for some use cases.
- simonw 4y agoSounds like you're in the "Big Data One-Percenter" category described at the very bottom of the article.
- CobrastanJorji 4y agoTableau's "Medium Data" April Fools Day ad from several years ago still rings amazingly true.
- hinkley 4y agoIt's always kind of amazed me how closely Big Data was followed by the KonMari method and it really seems like the nerds were not paying attention to that at all. Or just happy to take a paycheck from people who weren't paying attention. Hoarding is not a winning strategy.
- KaiserPro 4y agobig data isn't big anymore. 1) 10 years ago, having access to 300tb of data that could sustain 10gigabytes/s of throughput would require something like two racks of disks with some SSD cache and junk. 2) people thought hadoop was a good idea 3) People assumed that everything could be solved with map:reduce 3) machine learning was much less of a thing. 4) people realised that postgres does virtually everything that mongo claimed it could. 5) people realised that cassandra was a very expensive way to make a write only database. I gave a talk about using big data, and basically at the time the best definition I could come up with was "anything that's too big to reasonably fit in one computer. so think 4, 60 disk direct attached SAS boxes". Most of the time people were chasing the stuff for the CV, rather than actually stopping to think if it was a good idea. (think k8s two years ago, chatGPT now, chat bots in 2020). Most buisnesses just wanted metrics, and instead of building metrics into the app, they decided to boil the ocean by parsing unstructured logs. Not surprisingly it turned to shit pretty quick. Nowadays people are much better at building metrics generation directly into apps, so its much easier to easily plot and correlate stuff.
- swyx 4y agowhat is your current explanation for why hadoop turned out NOT to be a good idea and everything couldn't be solved with map:reduce?
- nerpderp82 4y agoBig Data was whatever someone couldn't handle in a spreadsheet or on their laptop using R. This paper is 8 years old and it was somewhat obvious then. Scalability! But at what COST? https://www.usenix.org/system/files/conference/hotos15/hotos15-paper-mcsherry.pdf https://www.usenix.org/system/files/conference/hotos15/hotos... A big single machine can handle 98% of peoples data reduction needs. This has always been true. Just because your laptop only has 16GB doesn't mean you need a Hadoop (or Spark, or Snowflake) cluster. And it was always in the best interest of the BD vendors and Cloud vendors to say, "collect it all" and analyze on/or using our platform. The future of data analysis is doing it at the point of use and incorporating it into your system directly. Your actionable insights should be ON your grafana dashboard seconds after the event occurred.
- mywittyname 4y agoYou can do a petabytes of analysis with regular old BigQuery just as easily as you can analyze megabytes of data. This solves the scalability issue for a lot of companies, IMHO.
- nerpderp82 4y agoI agree, BQ is a gem on GCP. You pay for storage (or not, you can use federated queries) and don't pay anything when you aren't using it. The ability to dynamically scale reservations is pretty nice as well.
- angry_moose 4y agoMy experience with "Big Data" is it was something that couldn't be handled in a spreadsheet or on their laptop using R because it was so inefficiently coded. I got sucked into "weekly key metric takes over 14 hours to run on our multi-node kubernetes cluster" a while back. I'm not sure how many nodes it actually used, nor did I really care. Digging into it, the python code ingested about ~50GB of various files, made well over a dozen copies of everything, leaving the whole thing extremely memory starved. I replaced almost all of the program with some "grep | sed | awk | sed | grep" abomination that stripped about 98% of the unnecessary info first and it ran in under 2 minutes on my laptop. I probably should have tightened it up more but I was more than happy to wash my hands of the whole thing by that point. Instead of improving the code, they just kept tossing more compute at it. Still heard all kinds of grumbling about os.system('grep | sed | awk | sed | grep') not being "pythonic" and "bad practice"; but not enough that they actually bothered to fix it.
- heisenbit 4y agoSampling has proven extremely useful. Pi can be approximated with it as were nuclear bombs designed using statistical methods. Flame graphs based on stack samples are used to optimize servers. Government does planning with it. Management does its thing by wandering around. It usually does not take many data points for an actionable insight and most actions then will invalidate small details in old data anyhow. Better to start every round with fresh eyes.
- lern_too_spel 4y agoPeople don't want to deal with having to rearchitect when their workload does not fit on a single instance. Yes, optimize for the small data case, but if you build a product that can handle only the small data case, you have a tough sell.
- jl6 4y agoTo add to the “the real issue is…” pile: Most orgs collect the data that is easy to collect, and they are extremely lucky if that happens to be the data that enables the insights they desire. When the data they really need looks too hard to get, the org tries to compensate by collecting more of the easy stuff, and hoping that if blood can’t be squeezed out of a stone, maybe it can be squeezed out of 100bn stones.
- CommieBobDole 4y agoIt's not dead, it's just entered the plateau of productivity. Where people use it for whatever it's useful for and don't try to solve every problem with it just because it's the cool new thing.
- poorman 4y agoThis entire post reads like "you probably don't actually have big data". What do these blockchains do that have to keep data around forever, with high throughput, and need to expose it quickly do? Are you saying they should delete parts of data in the chain? Seriously, I've spent my career working on big data systems, and while the answer is sometimes "yes you need to delete your data", I don't think that's going to always work.
- PeterisP 4y agoAnd what about these blockchains? The full history of Bitcoin blockchain is less than 500gb, so for any analysis just getting a machine with a terabyte of RAM is both simpler and cheaper (once you include dev+ops time) than doing any horizontal scaling across multiple machines with "Big Data" approaches. "You probably don't actually have big data" is a very valid point, not that many organizations do - most businesses haven't generated enough actionable data in their lifetime to need more than a single beefy machine without ever deleting data.
- poorman 4y agoBitcoin is notoriously slow. I don't think it's a good example of a high-throughput system. There are chains out there with 100x the number of transactions per second than that of Bitcoin. https://realtps.net/ https://realtps.net/
- blipvert 4y agoListen to “Reason”
- mikepk 4y agoWe need to re-think how to make data _useful_. The fact that the value hasn't materialized after decades of attempts, billions of dollars, and lots of tools and technology points to the fact that our core assumptions and patterns are wrong. This post doesn't go far enough. It challenges the assumption that everyone's data is "big data" or that every company's data will eventually grow to be big data. I agree that "big data" was the wrong model. We also need to challenge that all data should be stored in one place (warehouse, lake, lakehouse). We need to challenge that one tool can be used for every data need. We need to challenge how we build systems both from a technology and people standpoint. We need to embrace that the problems and needs of companies _are always changing_. We are living with conceptual inertia. Many of our patterns are an evolution from the 70's and 80's and the first relational databases. It's time to rethink how we "do data" from first principles.
- blakeburch 4y agoThe problem is that no tool alone can make data useful. It requires human ingenuity to come up with a theory, gather the required data, then test and verify the theory. We've gotten to a point where the first and last step get skipped. Business leaders see other companies doing interesting things with data, so the answer must be "gather all the data"! Internal teams end up focused on gathering the data without the context of how it might be used. We need to train data teams to not focus on the data as the product. Instead, they should be responsible for driving business actions. Gathering and cleaning the data should just a byproduct of that activity.
- nemo44x 4y agoWhy would I use DuckDB instead of Clickhouse or similar? Is it just because I want to have the database embedded in my app and not connect to a server?
- tylerhannan 4y agoOne great reason to use DuckDB was when ClickHouse took up too much memory on Parquet files. https://github.com/ClickHouse/ClickHouse/issues/45741#issuecomment-1419990419 https://github.com/ClickHouse/ClickHouse/issues/45741#issuec... helps with that though. Also, clickhouse-local exists https://clickhouse.com/blog/extracting-converting-querying-local-files-with-sql-clickhouse-local https://clickhouse.com/blog/extracting-converting-querying-l... as a thing. But, yes, when I think of DuckDB...I think embedded use cases...i'm also not a power user. I also think of this very much as a 'horses for courses' or 'different strokes, different folks' sort of scenario. There is, naturally, overlap because 'analytical data.' But also, there is naturally overlap with R and this giant scary mess of data-munging PERL code I maintain for a side project. The DuckDB team, the MotherDuck team, the ClickHouse team...we all want your experience interacting with data to be amazing. In some scenarios, ClickHouse is better. In some scenarios, DuckDB. I'm biased (as I work for ClickHouse in DevRel), but I <3 ClickHouse. Try both. Pick the one that is best for you. Then...you know...tell the other(s) why so that we all can get better at what we do.
- nemo44x 4y agoThanks but I’m looking for specific use cases. Like I get SQLite. And I get Clickhouse. But I just don’t get why I’d use DuckDB specifically. I’m sure it’s awesome and super useful but I have a gap in my understanding.
- morelisp 4y agoTo the extent "Big Data" originally and is still often claimed to mean "data beyond what fits on a single [process/RAM/disk/etc]", it's always been strange to me how much it's identified with analytics pipelines doing largely trivial transformations producing ultra-expensive "BI" pablum. Yes, thank goodness that part is dead. But meanwhile - we've still got more actual data than ever to store, and ever-tighter deadlines on finding and delivering it. If we can get back to that and let the PySpark bootcampers fade away, maybe things can get a little better for once. In other words: Even when querying giant tables, you rarely end up needing to process very much data. Modern analytical databases can do column projection to read only a subset of fields, and partition pruning to read only a narrow date range. They can often go even further with segment elimination to exploit locality in the data via clustering or automatic micro partitioning. Other tricks like computing over compressed data, projection, and predicate pushdown are ways that you can do less IO at query time. And less IO turns into less computation that needs to be done, which turns into lower costs and latency. Big data is "dead" because data engineers (the programming ones, not the analysts-in-all-but-title) spent a ton of effort building DBs with new techniques that scale better than before, with other storage patterns than before. Someone still has to write and maintain those! And it would be even better if those tools and techniques could escape the half dozen major data cloud companies and be more directly accessible to the average small team.
- cmrdporcupine 4y agoFrom about 2008/2009/2010 or so on there was perhaps an over-emphasis on specialized tools for the mass acquisition of streams of data. Maybe in large part due to the explosion of $$ in ad-tech. Some people had legitimately insane click/impression streams -- I worked at a couple companies like that. Development of DBs based on LSM trees or other write-specialized storage structures became important. Existing relational databases weren't particularly well built for this stuff. This was part of, but not the whole story with the whole NoSQL thing. People were willing to go completely denormalized in order to gain some advantage or ability here. It helped that much of the data looked at was of perhaps little structural complexity. In the meantime SSD storage took off, so the IOPS from a stock drive have skyrocketed, business domains for large data sets have broadened beyond click/impression streams, and the challenge now is not "can I store all this data" it's "WTH do I do with it?" Regardless of quantity of data, structuring and analysis and querying of said data remains paramount. The challenge for anybody working with data is to represent and extract knowledge. I remain convinced that logic -- first order logic and its offshoot in the relational model -- remains the best tool for reasoning about knowledge. Codd's prognostications on data from the 1970s are still profound. I think we're in a space now where we can turn our attention to knowledge management, not just accumulating streams of unstructured data. The challenge in a business is to discover and capture the rules and relationship in data. SQL is an existing but poor tool for this, based on some of the concepts in the relational model but tossing them together in a relatively uncomposable and awkward way (though it remains better than the dogs breakfast of "NoSQL" alternatives that were tossed together for a while there.) My employer is working in this space, I think they have a really good product: https://relational.ai/ https://relational.ai/
- winterismute 4y agoThe database was the key technology in the 2001-2011 decade: it allowed companies to store massive amount of data in an organized way, so that they could provide basic functionality (search, monitoring) to users. Statistical learning is being the key "technology" of 2011-today: it allowed companies, which had stored massive amount of data, to feedback predictions to users. I think AR/Computer Graphics will be the key technology of the next decade: it will allow users to interact directly and seamlessly with the insights produced by ML systems, and possibly feed-back information.
- deleted 4y ago[deleted]
- hugesniff 4y ago"Very often when a data warehousing customer moves from an environment where they didn’t have separation of storage and compute into one where they do have it, their storage usage grows tremendously..." Can someone explain why this is the case? Is it due to more replications or maintaining more indices?
- mejakethomas 4y agoSo what I'm hearing is it's not the size of your data that matters, it's how you use it?
- alluro2 4y agoI'm quite surprised with data sizes mentioned in the article, and wondering if I'm missing something...We are a very small 2yo company, handling route optimization and delivery management / field service. Even with our very small number of customers, their relatively small sizes (e.g. number of "tasks" per day), being very early in development in terms of data that we collect - our database containing just customer data for 2 years is ~100GB. Which I previously considered small, and if we collected useful user metrics, had more elaborate analytics, location tracking history etc, I would expect it to be at least 3x. We don't use any "BigData" products yet, as there wasn't any need for them, even when we provide full search and relatively nice and rich set of analytics over all the data. Yet, based on the article, we're way above most of the companies relying heavily on such tools. Confusing.
- taftster 4y agoThis posting was great. Highly recommended reading through. It gets really good when the author hits "Data is a Liability". > An alternate definition of Big Data is “when the cost of keeping data around is less than the cost of figuring out what to throw away.” This is exactly it. It's way too hard to go through and make decisions about what to throw away. In many respects, companies are the ultimate hoarders and can't fathom throwing any data way, Just In Case. Really appreciated the post overall. Very insightful. As an anecdote to this article, when business folks have come up to me and asked about storing their data in a Big Data facility, I have never found the justification to recommend it. Like, if your data can fit into RAM, what exactly are we talking about Big Data for?
- disqard 4y ago> if your data can fit into RAM, what exactly are we talking about Big Data for? That's a fantastic point, and I keep mentioning the COST paper to anyone who cares: https://www.usenix.org/system/files/conference/hotos15/hotos15-paper-mcsherry.pdf https://www.usenix.org/system/files/conference/hotos15/hotos...
- DriverDaily 4y agoCOVID was proof The Right Data is better than Big Data. All those data sources to measure how many sick people we have and it turns out we just need one: Wastewater.
- bombcar 4y agoOr another way to look at it - if we make a data lake that collects everyone’s shit we might find something useful in it!
- CliffStoll 4y agoIn a larger sense, it's a challenge to throw away stuff, just as it's difficult to trim big data. As I reach retirement, our attic, bookshelves, and cabinets must be trimmed -- and each item requires attention and a decision. Some things in the attic are obvious liabilities (what to do with a mercury barometer? A radium dial pocket watch? Old electronics?) Disposing of other stuff requires time, insight, and a sense of the future (should we keep those fingerpainted scribbles from when the kids were 3? How about those cheesy trophies from chess club? Computer books from the 1970's? Betamax home movies? Record albums?)
- LeanderK 4y agoWho has ever believed those claims? There's a common saying "garbage in, garbage out" about what happens with all those fancy models if the data quality is not high. That's really independent from dataset-size. There's no magic insight you get because your dataset is bigger. You need a quality analyst to handle your data, irrelevant of its size. Also, who thought their company would cease to function because surely they will hit google-scale dataset-sizes in the near future? Impossible for most except the biggest of the biggest
- zX41ZdbW 4y agoMy presentation from FOSDEM 2023 is very sympathetic to the "Big data is dead" statement: https://www.youtube.com/watch?v=JlcI2Vfz_uk https://www.youtube.com/watch?v=JlcI2Vfz_uk It is about using modern tools (ClickHouse) for data engineering without the fluff - when you can take whatever dataset or data stream and make what you need without the need for complex infrastructure. Nevertheless, the statement "big data is dead" is short-sighted, and I don't entirely follow this opinion. For example, here is one of ClickHouse's use-case: > Main cluster is 110PB nvme storage, 100k+ cpu cores, 800TB ram. The uncompressed data size on the main cluster is 1EB. And when you have this sort of data for realtime processing, no other technology can help you.
- gesman 4y agoCustomer pays data analytics vendor to tackle bunch of their [low quality, big size] data. If you have no tangible capabilities to do above, asking customer "ARE YOU IN THE BIG DATA ONE PERCENT?" will be the quickest way out of the door.
- blakeburch 4y agoGreat post and really resonates with my experience. Good to have some confirmation that most organizations aren't using their large swaths of data. Although I don't think most organizations are blaming lack of actionable insights on the data size. It's the lack of prioritizing data usage over data accessibility. We need to be teaching data people business levers and teaching business people data levers. Data should be a byproduct of an actionable idea that you want to execute. It shouldn't exist until you have that experiment in mind.
- ryadh 4y agoWhile I get that they're sometimes useful to trigger debate, I don't really subscribe to very bold statements. We are drowning in data, it's all around us. Information overload is real. Data enables most of our daily digital experiences, from operational data to insights in the form of user facing analytics. Data systems are the backbone of the digital life. It's is an ocean and it's all about the vessel you pick to navigate it. I don't believe that the vessel should dictates the size of the ocean, it's simply constrained by it's capabilities. The trick is to pick the right vessel for the job, whether you want to go fast, go far or fish for insights (ok, I need to stop pushing on this metaphor ) This visionary paper from Michael Stonebreaker (2005) predicted it quite accurately and I think is still relevant: https://cs.brown.edu/~ugur/fits_all.pdf https://cs.brown.edu/~ugur/fits_all.pdf Databases come in various flavours and the "trends" are simply a reflection of what the current era needs Disclaimer: I work at ClickHouse
- spopejoy 4y agoI guess the article title is a "bold statement" but maybe the biggest insight in there is that people don't think hard enough about throwing old data away, and it hurts them. This is a liferaft for drowning in data and is more "bold" organizationally, as it actually takes a certain kind of courage to realize you should just throw stuff away instead of succumb to the false comfort that "hey you never know when you might need it". Weirdly there's a similar thing that can happen to codebases, specifically unit tests and test fixtures that outlive any of their original programmers, nobody understands what's actually being tested and before each release lose days/weeks hammering to "fix the test". The only solution is to throw it away, but good luck getting most teams to ever do that, because of the false comfort they get -- even though that fixture is now just testing itself and not protecting you from any actual bugs. I mean how often does Netflix need to look a viewing habits from 2015? Summarize and throw it away.
- crazygringo 4y agoI am baffled by this comment. Throwing out unit tests? If you make a change and it fails a test, then you fix the bug or fix the test. I can't even imagine in what universe it's a good idea to throw away a test if it covers code in use. In what universe are unit tests "false comfort"? And if "nobody understands what's actually being tested" then you've got huge problems with your development practices. Similarly, viewing habits from 2015 are tremendously important. There may be a show they're releasing soon that is most similar to a title released in 2015, and those stats will provide the best model. "Summarize" requires knowing how data will be used in the future, but will likely throw away what you need. Not to mention how useful and profitable vast quantities of data are for ML training. Storing data is incredibly cheap. I'm actually curious where this desire to throw away old data comes from? I've literally never encountered it before, and it flies in the face of everything I've ever learned. The only context I know it from is data retention policies, but that's solely to limit legal liability.
- freedude 4y ago"Among customers who were using the service heavily, the median data storage size was much less than 100 GB" Eye-opening. Especially when combined with a recent quote from Satya Nadella, "First, as we saw customers accelerate their digital spend during the pandemic, we’re now seeing them optimize their digital spend to do more with less." Conclusion: SaaS is easy to drop off in downturns. Just as easy as it is to buy initially.
- zmmmmm 4y agoTo be honest, I slightly disagree about data size. I think the big data is there to be had, the real story is that data science itself has not panned out to provide the business value that people asserted would come from it. Data volumes haven't risen more because in the end, it turns out most of the things businesses need to know are easily ascertainable from much smaller data and their ability to action even these smaller very obvious things is already saturated. It doesn't help that we've shifted into a climate where hoarding data comes with a huge regulatory and compliance price tag, not to mention risk. But if the value was there we would do it, so this is not the primary driver.
- papito 4y agoFirst they came for the sacred microservices, now they are after Bid Data. What. Is. Happening. Don't get me wrong, I love it. It's about time people got off these stupid and shockingly expensive bandwagons.
- cmollis 4y agowe regularly run audits on over 12 years of customer order histories. This requires scanning of about 40TB of data and growing. They used to jump through hoops on the Oracle cluster just to get data out for one customer. We pushed all of the order history into s3 parquet using Spark and I can query this in about 20 seconds using Spark or Presto. It's now streamed through kafka and Spark structured streaming so it's up to date in about 3 minutes. The click-bait-y title notwithstanding, I get that not all data is 'big' and duckdb (and datafusion, polars, etc) is probably great for certain use-cases but what I work on every day can't be done on a single machine.
- greggyb 4y agoI mean, you can almost fit all that data on one SSD. Micron's latest are 30TB. Commodity servers are available with 24 NVMe drive bays. At 7 GB/s read, across 24 drives, you can scan 40 TB in 240 seconds. Such a server could readily fit 384 threads with dual EPYC and would be available with enough RAM to keep more than 10% of that data in cache. Your workload absolutely fits the definition of "fits on one machine". I am not suggesting here that you should put it on one machine. You definitely could, though.
- college_physics 4y agoNot dead, just complying with the Gartner cycle for hypes. There is probably a rational, well thought out classification of different types of data bigness, as in CERN-big, Google-big, MegaBank-big, down to wordpress-log big and on the basis of that one would probably find that different designs are indispensable, address different pain points and cannot really "die". Hype has a more erratic lifecycle than real needs
- ThomPete 4y agoOn the contrary. Now that AI is here, big data is going to be more alive than ever.
- sixdimensional 4y agoIt is amusing that in 2005, "VLDB" (precursor term to "big data") was defined in Wikipedia to be "larger than 1TB".. after reading through the post and the author's experience.. it would appear that this was not actually a completely terrible estimate, although there are larger and smaller: https://en.wikipedia.org/w/index.php?title=Very_large_database&oldid=20738417 https://en.wikipedia.org/w/index.php?title=Very_large_databa... The current version of that article states: "There is no absolute amount of data that can be cited. For example, one cannot say that any database with more than 1 TB of data is considered a VLDB. This absolute amount of data has varied over time as computer processing, storage and backup methods have become better able to handle larger amounts of data.[5] That said, VLDB issues may start to appear when 1 TB is approached,[8][9] and are more than likely to have appeared as 30 TB or so is exceeded.[10]" https://en.wikipedia.org/wiki/Very_large_database https://en.wikipedia.org/wiki/Very_large_database
- anonymousDan 4y agoI think the relevant phrase here is 'Selling your book'
- vonnik 4y agoAgree with this post. Big data was vendor generated hype that convinced many engineers to confuse the size of their dataset with their, ahem, shoe size. They didn’t do their employers any favors.
- xnx 4y agoSimilar to the "we must have microservices so that we can scale" fad a lot of people thought they had big data even though their records easily fit on a single machine.
- siliconc0w 4y agoBig data is dead because executives are rewarded when decisions are reactionary and politically savvy, data doesn't enter the picture.
- wowJustwow 4y agoInteresting take from a Googler. Big data hype never felt to me like anything more than a hype campaign to help big tech research ML/AI. Larry even rambled as much: https://arstechnica.com/information-technology/2013/05/larry-page-wants-you-to-stop-worrying-and-let-him-fix-the-world/ https://arstechnica.com/information-technology/2013/05/larry... It appears not all Googlers got the memo. Everyone else is in the way of him solving big problems! Not like such work could not be distributed among technologists and researchers around the globe via the internet. Help Google do it! I am leaning into “Deep Work” going forward; will slowly iterate on my own model creation and collaborate with like minded folks. I’m fucking done with intentionally empowering billionaire minority who convinced an ignorant political gerontocracy that minority is capable of magic. Anyone prattling on with common tropes of “longtermism”; nation state nutters, religious, technocrats; are appealing to non-existent authority they see a a magical future for us! Give them your money to insure it arrives! They have zero ability to insure such outcomes and a lot of upside to making people believe such today.
- noobermin 4y agoThe less than a terabyte datasets being common had me awestruck. I, singular post-doctoral scientist noobermin[0], have processed terabytes of data at a time on HPC systems. Sure, a lot of it was garbage and I had to wade through it, but no one paid me millions to do it, I just did it to publish the papers. Sure, I needed the system which cost someone a lot of money, I suppose. But, I considered myself a small fry compared to some of the things others did on the system, particularly, hyrdrodynamics modellers. Moreover, I know I can probably process 100GB datasets on my own home PC, which isn't too impressive, it would just take longer (say a day or so instead of a hour or a few minutes). And this is with idk, python scripts using MPI. Yes, MPI because I'm a computational scientist and that's what HPC systems use, nothing fancy and likely the "legacy systems" he railed against in his pitches, but it worked. I'm just awestruck, I could tell anyone that "large data" isn't really a bottleneck, but making sense of it is the very difficult part. My mentors kept pushing me to mention the sheer size of the datasets I process in talks because it sounds impressive, and I do do so, but I always knew it didn't matter because the interpretation and analysis is the hard part, not just the "sheer size." [0] not going to use my real name
- wongarsu 4y ago> I'm just awestruck, I could tell anyone that "large data" isn't really a bottleneck, but making sense of it is the very difficult part Especially now that 1TB datasets fit in memory on off-the-shelf servers and 100GB fits in memory on consumer hardware. You need a lot of data to run into real technical challenges that can't be solved by throwing a couple hundreed dollars a month (amortized cost) at hardware. And often you can get by with much, much less than even that.
- hackernewds 4y agoNot to mention if you're physically storing it (not in the cloud) 1TB can fit on your pinky finger
- krab 4y agoThat's a different category of big data. I worked for a big pharma and they were building their big data department with Spark and friends. I was quite surprised that their biggest dataset had something like 200 GB. At the same time, though, there was a lot of DNA sequencing data, we were designing CRISPR probes etc. But Spark and Hadoop aren't really that helpful in this area, so the Big Data team wasn't involved in those.
- sortalongo 4y ago> Customer data sizes followed a power-law distribution. The largest customer had double the storage of the next largest customer, the next largest customer had half of that, etc I’m no statistician, but I’m like 99% sure that’s an exponential, not a power law There’s a world of difference. The point of an exponential is that you can ignore big things. The point of a power law is that you can’t.
- lelanthran 4y ago>> Customer data sizes followed a power-law distribution. The largest customer had double the storage of the next largest customer, the next largest customer had half of that, etc > I’m no statistician, but I’m like 99% sure that’s an exponential, not a power law I'm no expert either, but it seems correct. The power law distribution has each X value in an X/Y series decreasing by a specific factor, like this: https://en.wikipedia.org/wiki/Power_law#/media/File:Long_tail.svg https://en.wikipedia.org/wiki/Power_law#/media/File:Long_tai... The exponential has each X value increasing by a specific factor, like this: https://en.wikipedia.org/wiki/Exponential_function#/media/File:Exp.svg https://en.wikipedia.org/wiki/Exponential_function#/media/Fi...
- deleted 4y ago[deleted]
- cbreynoldson 4y agoLong live Big Meaning
- alentred 4y agoAnother problem with "BigData": hiring and the tendency of the ecosystem to "sustain" itself (like any system). As a company hires traditional BigData Architects, Developers, Data Scientists, Engineers, etc. it will naturally have a tendency to choose the traditional BigData technology and solutions like BigQuery, Spark, storing everything in HDFS, etc. A trick I saw is companies hiring experienced jack-of-all-trades back-end engineers into Data teams. A lot of things get migrated from Spark to Postgres, from Kafka to REST API calls, and keep working fine and become generally more responsive. I'm on the same page as the author here: traditional BigData tech has its place and its uses, but before choosing it companies (CTOs, architects) should carefully consider if it is necessary, especially considering the cost of it and the risk of locking themselves down in a very specialized domain.
- 5tefan 4y agoI quite often say: If you need KPIs you're too far removed from how the company actually conducts business.
- qikInNdOutReply 4y agoCongratulations on the birth of little data, to the proud dad Big Data are in order. Well, obviously, the realization that management is largely emotion driven and little data driven, is a prelude for the CEO-AI yet in the makings. Of course this still got a face to it. A CEO who speaks and talks, as the voice commands, but does not do the part that even humans who think they are good at it, are bad at, decision making. The ground truth is there ("Cooperate history") going back to the merchants of sumeria. Lets learn that lesson, pack it into a decission tree, and wrap that bundle with Chat GPT smooth talking.
- xiaodai 4y agoMedium data is where it's at! diskframe.com and polars and arrows are good enough for most use cases!
- anktor 4y ago"90% of queries processed less than 100 MB of data. [in big query]" I think there is a problem when someone with such proclaimed knowledge of the sector gets to this, and similar, pieces of data, and does not attribute it to pricing. Could it be queries are short because bigquery pricing for analysis, as confusing as this models are, is based on amount of data?[0] Because the other line of reasoning is that a big chunk of that 90% of professionals being paid to do their jobs, do NOT take into account pricing of the tool and are using it for small data, instead of thinking that people are using the best tool with the lowest price, because there's plenty of options to process and analyse data right now in the cloud. On the "business have low amount of data", that matches my experience as well. At first I thought I was simply dealing with smaller sized companies, but it's a trend of doing big data projects for data that'd fit a pendrive. [0] https://cloud.google.com/bigquery/pricing#analysis_pricing_models https://cloud.google.com/bigquery/pricing#analysis_pricing_m...
- ammar_x 4y agoWell, we have less than 2 TB of data, and although we are running MySQL on a large instance with ~120 GB of RAM, it's extremely slow when dealing with big tables (like a 25 GB table) and that's why we need "big data" tools like BigQuery.
- diceduckmonk 4y agoUnlike quantum which cracks computationally complex algorithms, BigData was just about costs. SSDs we’re limited in capacity and still expensive. Parallelizing work with MapReduce allowed using cheap fault-prone commodity hardware and disks. If you’re dealing with terabytes rather than petabytes of data, you probably don’t need BigData
- twwittr1 4y agoThis is what most of the companies are doing.
- harish_dash 4y agoConfirmation bias exists almost everywhere. Confirmation bias especially among senior management is highly dangerous as decisions are based on not on data and facts, rather they are based on anecdotes, hunch/feelings, with high probability of going wrong. This is precisely where data scientists play a significant role, by providing recommendations and presenting facts based on hard data and mathematical models, in order to ensure that senior management decisions are based on facts/data, and not on anecdotes and hunches. Furthermore, a data driven organisation must have a supporting culture, where data driven decisions are given precedence, and data scientists (data messengers) must be empowered to present facts as is, no matter whether these facts are aligned or not with the basic assumptions and biases held by the senior management team. Creating such a supporting organization culture is extremely important but definite not easy. Culture is one of the factors that makes a difference between success or failure in a data driven organisation.