17 ms·
When I was hiring data scientists for a previous job, my favorite tricky question was "what stack/architecture would you build" with the somewhat detailed requi
by kmarc 2y ago
When I was hiring data scientists for a previous job, my favorite tricky question was "what stack/architecture would you build" with the somewhat detailed requirements of "6 TiB of data" in sight. I was careful not to require overly complicated sums, I simply said it's MAX 6TiB
I patiently listened to all the big query hadoop habla-blabla, even asked questions about the financials (hardware/software/license BOM) and many of them came up with astonishing tens of thousands of dollars yearly.
The winner of course was the guy who understood that 6TiB is what 6 of us in the room could store on our smart phones, or a $199 enterprise HDD (or three of them for redundancy), and it could be loaded (multiple times) to memory as CSV and simply run awk scripts on it.
I am prone to the same fallacy: when I learn how to use a hammer, everything looks like a nail. Yet, not understanding the scale of "real" big data was a no-go in my eyes when hiring.
- wslh 2y agoIn my context 99% of the problem is the ETL, nothing to do with complex technology. I see people stuck when they need to get this from different sources in different technologies and/or APIs.
- geraldwhen 2y agoI ask a similar question on screens. Almost no one gives a good answer. They describe elaborate architectures for data that fits in memory, handily.
- mcny 2y agoI think that’s the way we were taught in college / grad school. If the premise of the class is relational databases, the professor says, for the purpose of this course, assume the data does not fit in memory. Additionally, assume that some normalization is necessary and a hard requirement. Problem is most students don’t listen to the first part “for the purpose of this course”. The professor does not elaborate because that is beyond the scope of the course.
- acomjean 2y agoI took a Hadoop class. We learned hadoop and were told by the instructor we probably wouldn’t’t need it, and learned some other Java processing techniques (streams etc)
- kmarc 2y agoFWIW if they were juniors, I would've continued the interview and direct them with further questions, and observer their flow of thinking to decide if they are good candidates to pursue further. But no, this particular person had been working professionally for decades (in fact, he was much older than me).
- geraldwhen 2y agoYeah. I don’t even bother asking juniors this. At that level I expect that training will be part of the job, so it’s not a useful screener.
- Joel_Mckay 2y agoPeople can always find excuses to boot candidates. I would just back-track from a shipped product date, and try to guess who we needed to get there... given the scope of requirements. Generally, process people from a commercially "institutionalized" role are useless for solving unknown challenges. They will leave something like an SAP, C#, or MatLab steaming pile right in the middle of the IT ecosystem. One could check out Aerospike rather than try to write their own version (the dynamic scaling capabilities are very economical once setup right.) Best of luck, =3
- boppo1 2y agoYou have 6 TiB of ram?
- ninkendo 2y agoYou don’t need that much ram to use mmap(2)
- marginalia_nu 2y agoTo be fair, mmap doesn't put your data in RAM, it presents it as though it was in RAM and has the OS deal with whether or not it actually is.
- ninkendo 2y agoRight, which is why you can mmap way more data than you have ram, and treat it as though you do have that much ram. It’ll be slower, perhaps by a lot, but most “big data” stuff is already so god damned slow that mmap probably still beats it, while being immeasurably simpler and cheaper.
- marginalia_nu 2y agoReally depends on the shape of the data. mmap can be suboptimal in many cases. For CSV it flat out doesn't matter what you do since the format is so inefficient and needs to be read start to finish, but something like parquet probably benefits from explicit read syscalls, since it's block based and highly structured, where you can predict the read patterns much better than the kernel can.
- cess11 2y agoThe "(multiple times)" part probably means batching or streaming. But yeah, they might have that much RAM. At a rather small company I was at we had a third of it in the virtualisation cluster. I routinely put customer databases in the hundreds of gigabytes into RAM to do bug triage and fixing.
- sfilipco 2y agoI agree that keeping data local is great and should be the first option when possible. It works great on 10GB or even 100GB, but after that starts to matter what you optimize for because you start seeing execution bottlenecks. To mitigate these bottlenecks you get fancy hardware (e.g oracle appliance) or you scale out (and get TCO/performance gains from separating storage and compute - which is how Snowflake sold 3x cheaper compared to appliances when they came out). I believe that Trino on HDFS would be able to finish faster than awk on 6 enterprise disks for 6TB data. In conclusion I would say that we should keep data local if possible but 6TB is getting into the realm where Big Data tech starts to be useful if you do it a lot.
- nottorp 2y ago> I agree that keeping data local is great and should be the first option when possible. It works great on 10GB or even 100GB, but after that starts to matter what you optimize for because you start seeing execution bottlenecks. The point of the article is 99.99% of businesses never pass even the 10 Gb point though.
- sfilipco 2y agoI agree with the theme of the article. My reply was to parent comment which has a 6 TB working set.
- hectormalot 2y agoI wouldn't underestimate how much a modern machine with a bunch of RAM and SSDs can do vs HDFS. This post[1] is now 10 years old and has find + awk running an analysis in 12 seconds (at speed roughly equal to his hard drive) vs Hadoop taking 26 minutes. I've had similar experiences with much bigger datasets at work (think years of per-second manufacturing data across 10ks of sensors). I get that that post is only on 3.5GB, but, consumer SSDs are now much faster at 7.5GB/s vs 270MB/s HDD back when the article was written. Even with only mildly optimised solutions, people are churning through the 1 billion rows (±12GB) challenge in seconds as well. And, if you have the data in memory (not impossible) your bottlenecks won't even be reading speed. [1]: https://adamdrake.com/command-line-tools-can-be-235x-faster-than-your-hadoop-cluster.html https://adamdrake.com/command-line-tools-can-be-235x-faster-...
- pdimitar 2y agoBlows my mind. I am a backend programmer and a semi-decent sysadmin and I would have immediately told you: "make a ZFS or BCacheFS pool with 20-30% redundancy bits and just go wild with CLI programs, I know dozens that work on CSV and XML, what's the problem?". And I am not a specialized data scientist. But with time I am wondering if such a thing even exists... being a good backender / sysadmin and knowing a lot of CLI tools has always seemed to do the job for me just fine (though granted I never actually managed a data lake, so I am likely over-simplifying it).
- WesolyKubeczek 2y ago> just go wild with CLI programs, I know dozens that work on CSV and XML ...or put it into SQLite for extra blazing fastness! No kidding.
- pdimitar 2y agoThat's included in CLI tools. Also duckdb and clickhouse-local are amazing.
- WesolyKubeczek 2y agoI need to learn more about the latter for some log processing...
- fijiaarone 2y agoLog files aren’t data. That’s your first problem. But that’s the only thing that most people have that generates more bytes than can fit on screen in a single spreadsheet.
- thfuran 2y agoOf course they are. They just aren't always structured nicely.
- 2y ago
- mattbillenstein 2y agoI can appreciate the vertical scaling solution, but to be honest, this is the wrong solution for almost all use cases - consumers of the data don't want awk, and even if they did, spooling over 6TB for every kinda of query without partitioning or column storage is gonna be slow on a single cpu - always. I've generally liked BigQuery for this type of stuff - the console interface is good enough for ad-hoc stuff, you can connect a plethora of other tooling to it (Metabase, Tableau, etc). And if partitioned correctly, it shouldn't be too expensive - add in rollup tables if that becomes a problem.
- __alexs 2y agoA moderately powerful desktop processor has memory bandwidth of over 50TB/s so yeah it'll take a couple of minutes sure.
- fijiaarone 2y agoThe slow part of using awk is waiting for the disk to spin over the magnetic head. And most laptops have 4 CPU cores these days, and a multiprocess operating system, so you don’t have to wait for random access on a spinning plate to find every bit in order, you can simply have multiple awk commands running in parallel. Awk is most certainly a better user interface than whatever custom BrandQL you have to use in a textarea in a browser served from localhost:randomport
- Androider 2y ago> The slow part of using awk is waiting for the disk to spin over the magnetic head. If we're talking about 6 TB of data: - You can upgrade to 8 TB of storage on a 16-inch MacBook Pro for $2,200, and the lowest spec has 12 CPU cores. With up to 400 GB/s of memory bandwidth, it's truly a case of "your big data problem easily fits on my laptop". - Contemporary motherboards have 4 to 5 M.2 slots, so you could today build a 12 TB RAID 5 setup of 4 TB Samsung 990 PRO NVMe drives for ~ 4 x $326 = $1,304. Probably in a year or two there will be 8 TB NVMe's readily available. Flash memory is cheap in 2024!
- _ugfj 2y agohttps://x.com/garybernhardt/status/600783770925420546 https://x.com/garybernhardt/status/600783770925420546 (Gary Bernhardt of WAT fame): > Consulting service: you bring your big data problems to me, I say "your data set fits in RAM", you pay me $10,000 for saving you $500,000. This is from 2015...
- RandomCitizen12 2y agohttps://yourdatafitsinram.net/ https://yourdatafitsinram.net/
- crowcroft 2y agoI wonder if it's fair to revise this to 'your data set fits on NVME drives' these days. Astonishing how fast and how much storage you can get these days.
- fbdab103 2y agoYou can always check available ram: https://yourdatafitsinram.net/ https://yourdatafitsinram.net/
- xethos 2y agoBased on a very brief search: Samsung's fastest NVME drives [0] could maybe keep up with the slowest DDR2 [1]. DDR5 is several orders of magnitude faster than both [2]. Maybe in a decade you can hit 2008 speeds, but I wouldn't consider updating the phrase before then (and probably not after, either). [0] https://www.tomshardware.com/reviews/samsung-980-m2-nvme-ssd-review https://www.tomshardware.com/reviews/samsung-980-m2-nvme-ssd... [1] https://www.tomshardware.com/reviews/ram-speed-tests,1807-3.html https://www.tomshardware.com/reviews/ram-speed-tests,1807-3.... [2] https://en.wikipedia.org/wiki/DDR5_SDRAM https://en.wikipedia.org/wiki/DDR5_SDRAM
- dralley 2y agoThe statement was "fits on", not "matches the speed of".
- 2y ago
- marginalia_nu 2y agoProblem is possibly that most people with that sort of hands-on intuition for data don't see themselves as data scientists and wouldn't apply for such a position. It's a specialist role, and most people with the skills you seek are generalists.
- deepsquirrelnet 2y agoYeah it’s not really what you should be hiring a data scientist to do. I’m of the opinion that if you don’t have a data engineer, you probably don’t need a data scientist. And not knowing who you need for a job causes a lot of confusion in interviews.
- the_real_cher 2y agoHow would six terabytes fit into memory? It seems like it would get a lot of swap thrashing if you had multiple processes operating on disorganized data. I'm not really a data scientist and I've never worked on data that size so I'm probably wrong.
- coldtea 2y ago>How would six terabytes fit into memory? What device do you have in mind? I've seen places use 2TB RAM servers, and that was years ago, and it isn't even that expensive (can get those for about $5K or so). Currently HP allows "up to 48 DIMM slots which support up to 6 TB for 2933 MT/s DDR4 HPE SmartMemory". Close enough to fit the OS, the userland, and 6 TiB of data with some light compression. >It seems like it would get a lot of swap thrashing if you had multiple processes operating on disorganized data. Why would you have "disorganized data"? Or "multiple processes" for that matter? The OP mentions processing the data with something as simple as awk scripts.
- fijiaarone 2y ago“How would six terabytes fit into memory?” A better question would be: Why would anyone stream 6 terabytes of data over the internet? In 2010 the answer was: because we can’t fit that much data in a single computer, and we can’t get accounting or security to approve a $10k purchase order to build a local cluster, so we need to pay Amazon the same amount every month to give our ever expanding DevOps team something to do with all their billable hours. That may not be the case anymore, but our devops team is bigger than ever, and they still need something to do with their time.
- the_real_cher 2y agoWell yeah streaming to the cloud to work around budget issues is a while nother convo haha.
- Terr_ 2y agoI'm having flashbacks to some new outside-hire CEO making flim-flam about capex-vs-opex in order to justify sending business towards a contracting firm they happened to know.
- pyronik19 2y ago[dead]
- rr808 2y agoIf you look at the article the data space is more commonly 10GB which matches my experience. For these sizes definitely simple tools are enough.
- deleted 2y ago[deleted]
- randomtoast 2y agoNow, you have to consider the cost it takes for you whole team to learn how to use AWK instead of SQL. Then you do these TCO calculations and revert back to the BigQuery solution.
- deleted 2y ago[deleted]
- tomrod 2y agoAbout $20/month for chatgpt or similar copilot, which really they should reach for independently anyhow.
- randomtoast 2y agoAnd since the data scientist cannot verify the very complex AWK output that should be 100% compatible with his SQL query, he relies on the GPT output for business-critical analysis.
- tomrod 2y agoOnly if your testing frameworks are inadequate. But I belive you could be missing or mistaken on how code generation successfully integrates into a developer and data scientist's work flow. Why not take a few days to get familiar with AWK, a skill which will last a lifetime? Like SQL, it really isn't so bad.
- randomtoast 2y agoIt is easier to write complex queries in SQL instead of AWK. I know both AWK and SQL, and I find SQL much easier for complex data analysis, including JOINS, subqueries, window functions, etc. Of course, your mileage may vary, but I think most data scientists will be much more comfortable with SQL.
- elicksaur 2y agoMany people have noted how when using LLMs for things like this, the person’s ultimate knowledge of the topic is less than it would’ve otherwise been. This effect then forces the person to be reliant on the LLM for answering all questions, and they’ll be less capable of figuring out more complex issues in the topic. $20/mth is a siren’s call to introduce such a dependency to critical systems.
- bee_rider 2y agoThere’d still have to be some further questions, right? I guess if you store it on the interview group’s cellphones you’ll have to plan on what to do if somebody leaves or the interview room is hit by a meteor, if you plan to store it in ram on a server you’ll need some plan for power outages.
- apwell23 2y agoWhat kind of business just has a static set of 6TiB data that people are loading on their laptops. You tricked candidates with your nonsensical scenario. Hate smartass interviewers like this that are trying some gotcha to feel smug about themselves. Most candidates don't feel comfortable telling ppl 'just load on your laptops' even if they think thats sensible. They want to present a 'professional solution', esp when you tricked them with the word 'stack'. which is how most of them prbly perceived your trick question. This comment is so infuriating to me. Why be assholes to each other when world is already full of them.
- tomrod 2y agoI disagree with your take. Your surly rejoinder aside, the parent commenter identifies an area where senior level knowledge and process appropriately assess a problem. Not every job interview is satisfying checklist of prior experience or training, but rather assessing how well that skillset will fit the needed domain. In my view, it's an appropriate question.
- apwell23 2y agoWhat did you gather as 'needed domain' from that comment. 'needed domain' is often implicit, its not a blank slate. candidates assume all sorts of 'needed domain' even before the interview starts, if i am interviewing at bank I wouldn't suggest 'load it on your laptops' as my 'stack'. OP even mentioned that it his favorite 'tricky question' . It would def trick me because they used the word 'stack' which has specific meaning in the industry. There are even websites dedicated to 'stack's https://stackshare.io/instacart/instacart https://stackshare.io/instacart/instacart
- yxwvut 2y agoWell put. Whoever asked this question is undoubtedly a nightmare to work with. Your data is the engine that drives your business and its margin improvements, so why hamstring yourself with a 'clever' cost saving but ultimately unwieldy solution that makes it harder to draw insight (or build models/pipelines) from? Penny wise and pound foolish, plus a dash of NIH syndrome. When you're the only company doing something a particular way (and you're not Amazon-scale), you're probably not as clever as you think.
- throwaway_20357 2y agoIt depends on what you want to do with the data. It can be easier to just stick nicely-compressed columnar Parquets in S3 (and run arbitrarily complex SQL on them using Athena or Presto) than to try to achieve the same with shell-scripting on CSVs.
- fock 2y agohow exactly is this solution easier than putting the very Parquet files on a classic filesystem. Why does the easy solution require an amazon-subscription?
- pdinny 2y agoThis is adjacent to "why would I need EC2 when I can serve from my laptop?" In terms of maturity of solution and diversity of downstream applications you'll go much further with BigQuery/Athena (at comically low cost) for this amount of data than some cobbled together "local" solution. I thoroughly agree with the author but the comments in this thread are an indication of people who haven't actually had to do meaningful ongoing work with modest amounts of data if they're suggesting just storing it as plain text on personal devices. I'm not advocating for complicated or expensive solutions here, BigQuery and Athena are very low complexity compared to any of the Hadoop et-al tooling (yes Athena is Trino is in the family, but it is managed and dirt cheap).
- filleokus 2y agoI think I've written about it here before, but I imported ≈1 TB of logs into DuckDB (which compressed it to fit in RAM of my laptop) and was done with my analysis before the data science team had even ingested everything into their spark cluster. (On the other hand, I wouldn't really want the average business analyst walking around with all our customer data on their laptops all the time. And by the time you have a proper ACL system with audit logs and some nice way to share analyses that updates in real time as new data is ingested, the Big Data Solution™ probably have a lower TCO...)
- marcosdumay 2y ago> And by the time you have ... the Big Data Solution™ probably have a lower TCO... I doubt it. The common Big Data Solutions manage to have a very high TCO, where the least relevant share is spent on hardware and software. Most of its cost comes from reliability engineering and UI issues (because managing that "proper ACL" that doesn't fit your business is a hell of a problem that nobody will get right).
- yourapostasy 2y ago> ...managing that "proper ACL" that doesn't fit your business is a hell of a problem that nobody will get right... I'm not sure there is a way to get this right unless there is a programmatic integration into the org chart, and ability to describe and parse in a declarative language the organizational rules of who has access to what, when, under what auth, etc. It has otherwise been for me an exercise in watching massive amounts of toil manually interpreting between the SOT of the org chart and all the other applications mediated by many manual approval policies and procedures. And at every client I've posed this to, I've always been denied that programmatic access for integration. A lot of sites try to avoid this by designing ACL's around certain activity or data domains because those are more stable than organizations, but this breaks down when you get to the fine-grained levels of the ACL's so we get capped benefits from this approach. I'd love to hear how others solve this in large (10K+ staff) organizations that frequently change around teams.
- riku_iki 2y ago
- thunky 2y ago> requirements of "6 TiB of data" How could anyone answer this without knowing how the data is to be used (query patterns, concurrent readers, writes/updates, latency, etc)? Awk may be right for some scenarios, but without specifics it can't be a correct answer.
- marginalia_nu 2y agoThose are very appropriate follow up questions I think. If someone tasks you to deal with 6 TiB of data, it is very appropriate to ask enough questions until you can provide a good solution, far better than to assume the questions are unknowable and blindly architect for all use cases.
- kbolino 2y agoEven if a 6 terabyte CSV file does fit in RAM, the only thing you should do with it is convert it to another format (even if that's just the in-memory representation of some program). CSV stops working well at billions of records. There is no way to find an arbitrary record because records are lines and lines are not fixed-size. You can sort it one way and use binary search to find something in it in semi-reasonable time but re-sorting it a different way will take hours. You also can't insert into it while preserving the sort without rewriting half the file on average. You don't need Hadoop for 6 TB but, assuming this is live data that changes and needs regular analysis, you do need something that actually works at that size.
- rcxdude 2y agoYeah, but it very well not be data that needs random access or live insertion. A lot of data is basically just one big chunk of time-series that just needs some number crunching run over the whole lot every half a year.
- 7thaccount 2y agoI am a big fan of these simplistic solutions. In my own area, it was incredibly frustrating as what we needed was a database with a smaller subset of the most recent information from our main long-term storage database for back end users to do important one-off analysis with. This should've been fairly cheap, but of course the IT director architect guy wanted to pad his resume and turn it all into multi-million project with 100 bells and whistles that nobody wanted.
- palata 2y agoOne thing that may have an impact on the answers: you are hiring them, so I assume they are passing a technical interview. So they expect that you want to check their understanding of the technical stack. I would not conclude that they over-engineer everything they do from such an answer, but rather just that they got tricked in this very artificial situation where you are in a dominant position and ask trick questions. I was recently in a technical interview with an interviewer roughly my age and my experience, and I messed up. That's the game, I get it. But the interviewer got judgemental towards my (admittedly bad) answers. I am absolutely certain that were the roles inverted, I could choose a topic I know better than him and get him in a similarly bad position. But in this case, he was in the dominant position and he chose to make me feel bad. My point, I guess, is this: when you are the interviewer, be extra careful not to abuse your dominant position, because it is probably counter-productive for your company (and it is just not nice for the human being in front of you).
- emptyfile 2y ago[dead]
- ufo 2y agoFrom the point of view of the interviewee, it's impossible to guess if they expect you to answer "no need for big data" or if they expect you to answer "the company is aiming for exponential growth so disregard the 6TB limit and architect for scalability"
- kmarc 2y agoFWIW, it's a 2.5 second extra to say "Although you don't need big data, but if you insist, ..." and gimme the hadoop answer.
- whamlastxmas 2y agoIs this like interviewing for a chef position for a fancy restaurant and when asked how to perfectly cook a steak, you preface it with “well you can either go to McDonald’s and get a burger, or…” It may not be reasonable to suggest that in a role that traditionally uses big data tools
- mrtimo 2y ago.parquet files are completely underrated, many people still do not know about the format! .parquet preserves data types (unlike CSV) They are 10x smaller than CSV. So 600GB instead of 6TB. They are 50x faster to read than CSV They are an "open standard" from Apache Foundation Of course, you can't peek inside them as easily as you can a CSV. But, the tradeoffs are worth it! Please promote the use of .parquet files! Make .parquet files available for download everywhere .csv is available!
- sph 2y agoThird consecutive time in 86 days that you mention .parquet files. I am out of my element here, but it's a bit weird
- fifilura 2y agoFWIW I am the same. I tend to recommend BigQuery and AWS/Athena in various posts. Many times paired with Parquet. But it is because it makes a lot of things much simpler, and that a lot of people have not realized that. Tooling is moving fast in this space, it is not 2004 anymore. His arguments are still valid and 86 days is a pretty long time.
- ok_computer 2y agoSometimes when people discover or extensively use something they are eager to share in contexts they think are relevant. There is an issue when those contexts become too broad. 3 times across 3 months is hardly astroturfing for big parquet territory.
- mrtimo 2y agoI've downloaded many csv files that were mal-formatted (extra commas or tabs etc.), or had dates in non-standard formats. Parquet format probably would not have had these issues!
- swyx 2y agono need to be so suspicious when its an open standard not even linked to a startup?
- EdwardDiego 2y agoIf you were hiring me for a data engineering role and asked me how to store and query 6 TiB, I'd say you don't need my skills, you've probably got a Postgres person already.
- hotstickyballs 2y agoAnd how many data scientists are familiar with using awk scripts? If you’re the only one then you’ll have failed at scaling the data science team.
- jrm4 2y agoThis feels representative of so many of our problems in tech, overengineering, over-"producting," over-proprietary-ing, etc. Deep centralization at the expense of simplicity and true redundancy; like renting a laser cutter when you need a boxcutter, a pair of scissors, and the occasional toenail clipper.
- deleted 2y ago[deleted]
- rgrieselhuber 2y agoThis is a great test / question. More generally, it tests knowledge with basic linux tooling and mindset as well as experience level with data sizes. 6TiB really isn't that much data these days, depending on context and storage format, etc. of course.
- deepsquirrelnet 2y agoIt could be a great question if you clarify the goals. As it stands it’s “here’s a problem, but secretly I have hidden constraints in my head you must guess correctly”. The OPs desired solution could have been found from probably some of those other candidates if asked “here is the challenge, solve in most McGuyver way possible”. Because if you change the second part, the correct answer changes. “Here is a challenge, solve in the most accurate, verifiable way possible” “Here is a challenge, solve in a way that enables collaboration” “Here is a challenge, 6TiB but always changing” ^ These are data science questions much more than the question he was asking. The answer in this case is that you’re not actually looking for a data scientist.
- 6510 2y agoI dont know anything but when doing that I always end up next Thursday having the same with 4TB and the next with 17 at which point I regret picking a solution that fit so exactly.
- wg0 2y agoI have lived through the hype of Big data it was a time of HDFS+HTable I guess and Hapoop etc. One can't go wrong with DuckDB+SQLite+Open/Elasticsearch either with 6 to 8 even 10 TB of data. [0]. https://duckdb.org/ https://duckdb.org/
- michaelcampbell 2y agoMy smartphone cannot store 1TiB. <shrug>
- dfgdfg34545456 2y agoThe problem with your question is that they are there to show off their knowledge. I failed a tech interview once, question was build a web page/back end/db that allows people to order let's say widgets, that will scale huge. I went the simpleton answer route, all you need is Rails, a redis cache and an AWS provisioned relational DB, solve the big problems later if you get there sort of thing. Turns out they wanted to hear all about microservices and sharding.
- rqtwteye 2y agoPlenty of people get offended if you tell them that their data isn’t really “big data”. A few years ago I had a discussion with one of my directors about a system IT had built for us with Hadoop, API gateways, multiple developers and hundreds of thousands of yearly cost. I told him that at our scale (now and any foreseeable future) I could easily run the whole thing on a USB drive attached to his laptop and a few python scripts. He looked really annoyed and I was never involved again with this project. I think it’s part of the BS cycle that’s prevalent in companies. You can’t admit that you are doing something simple.
- noisy_boy 2y agoIn most non-tech companies, it comes down to the motive of the manager and in most cases it is expansion of reporting line and grabbing as much budget as possible. Using "simple" solutions runs counter to this central motivation.
- disqard 2y agoThis is also true of tech companies. Witness how the "GenAI" hammer is being used right now at MS, Google, Meta, etc.
- eloisant 2y ago- the manager wants expansion - the developers want to get experience in a fancy stack to build up their resume Everyone benefits from the collective hallucination
- boh 2y agoThat's the tech sector in a nutshell. Very few innovations actually matter to non-tech companies. Most companies could survive on Windows 98 software.
- sien 2y agoThe flipside of this is that people at some places are fully aware of this and get very suspicious of consultants offering to handle their 'big data' challenges with loads of highly proprietary stuff. Data organisation can be more of an issue, but the general issue with this is often a lack of internal discipline on the data owners to carefully manage their data. But that isn't nearly as attractive for management. Putting 'brought in new cloud vendor / technology' looks better than 'improved data organisation' on a CV even if the new vendor was a waste of money.
- lizknope 2y agoI'm on some reddit tech forums and people will say "I need help storing a huge amount of data!" and people start offering replies for servers that store petabytes. My question is always "How much data do you actually have?" Many times you they reply with 500GB or 2TB. I tell that that isn't much data when you can get 1TB micro SD card the size of a fingernail or a 24TB hard drive. My feeling is that if you really need to store petabytes of data that you aren't going to ask how to do it on reddit. If you need to store petabytes you will have an IT team and substantial budget and vendors that can figure it out.
- KronisLV 2y ago> The winner of course was the guy who understood that 6TiB is what 6 of us in the room could store on our smart phones, or a $199 enterprise HDD (or three of them for redundancy), and it could be loaded (multiple times) to memory as CSV and simply run awk scripts on it. If it's not a very write heavy workload but you'd still want to be able to look things up, wouldn't something like SQLite be a good choice, up to 281 TB: https://www.sqlite.org/limits.html https://www.sqlite.org/limits.html It even has basic JSON support, if you're up against some freeform JSON and not all of your data neatly fits into a schema: https://sqlite.org/json1.html https://sqlite.org/json1.html A step up from that would be PostgreSQL running in a container: giving you the support for all sorts of workloads, more advanced extensions for pretty much anything you might ever want to do, from geospatial data with PostGIS, to something like pgvector, timescaledb etc., while still having a plethora of drivers and still not making your drown in complexity and having no issues with a few dozen/hundred TB of data. Either of those would be something that most people on the market know, neither will make anyone want to pull their hair out and they'll give you the benefit of both quick data writes/retrieval, as well as querying. Not that everything needs or can even work with a relational database, but it's still an okay tool to reach for past trivial file storage needs. Plus, you have to build a bit less of whatever functionality you might need around the data you store, in addition to there even being nice options for transparent compression.
- hipadev23 2y agoHuh? How are you proposing loading a 6TB CSV into memory multiple times? And then processing with awk, which generally streams one a line at a time. Obviously we can get boxes with multiple terabytes of RAM for $50-200/hr on-demand but nobody is doing that and then also using awk. They’re loading the data into clickhouse or duckdb (at which point the ram requirement is probably 64-128GB) I feel like this is an anecdotal story that has mixed up sizes and tools for dramatic effect.
- lelanthran 2y ago> How are you proposing loading a 6TB CSV into memory multiple times? And then processing with awk, which generally streams one a line at a time. Ramdisk would work.
- assmanreturns 2y agoAwk doesn't load things into memory. It processes one line at a time. So memory usage is basically zero. That said awk isn't that fast. I mean your looking at "query" times in the range of at least 30 minutes if not more. Awk is imo a poor solution. I use awk all the time and I would never use it for something like this. Why not just use postgres. Its a lot more powerful, easy to setup and you get SQL which is extremely powerful. Normally I might even go with sqllite but for me 6TB is too much for sqllite.
- dahart 2y agoWait, how would you split 6 TiB across 6 phones, how would you handle the queries? How long will the data live, do you need to handle schema changes, and how? And what is the cost of a machine with 15 or 20 TiB of RAM (you said it fits in memory multiple times, right?) - isn’t the drive cost irrelevant here? How many requests per second did you specify? Isn’t that possibly way more important than data size? Awk on 6 TiB, even in memory, isn’t very fast. You might need some indexing, which suddenly pushes your memory requirement above 6 TiB, no? Do you need migrations or backups or redundancy? Those could increase your data size by multiples. I’d expect a question that specified a small data size to be asking me to estimate the real data size, which could easily be 100 TiB or more.
- torginus 2y agoIt's astonishing how shit the cloud is compared to boring-ass pedestrian technology. For example, just logging stuff into a large text file is so much easier, performant and searchable that using AWS CloudWatch, presumably written by some of the smartest programmers who ever lived. On another note I was once asked to create a big data-ish object DB, and me, knowing nothing about the domain, and a bit of benchmarking, decided to just use zstd-compressed json streams with a separate index in an sql table. I'm sure any professional would recoil at it in horror, but it could do literally gigabytes/sec retrieval or deserialization on consumer grade hardware.
- jandrewrogers 2y agoAs a point of reference, I routinely do fast-twitch analytics on tens of TB on a single, fractional VM. Getting the data in is essentially wire speed. You won't do that on Spark or similar but in the analytics world people consistently underestimate what their hardware is capable of by something like two orders of magnitude. That said, most open source tools have terrible performance and efficiency on large, fast hardware. This contributes to the intuition that you need to throw hardware at the problem even for relatively small problems. In 2024, "big data" doesn't really start until you are in the petabyte range.
- jodrellblank 2y ago> "most open source tools have terrible performance and efficiency on large, fast hardware." What do you use?
- deleted 2y ago[deleted]
- buremba 2y agoI can’t really think of a product with the requirement of max 6TiB data. If the data is big as TiB, most products have 100x TiB rather than a few ones.
- citizenpaul 2y agoThe funny thing is that is exactly the place I want to work at. I've only found one company so far and the owner sold during the pandemic. So far my experience is that amount of companies/people that want what you describe is incredibly low. I wrote a comment on here the other day that some place I was trying to do work for was using $11k USD a month on a BigQuery DB that had 375MB of source data. My advice was basically you need to hire a data scientist that knows what they are doing. They were not interested and would rather just band-aid the situation for a "cheap" employee. Despite the fact their GCP bill could pay for a skilled employee. As I've seen it for the last year job hunting most places don't want good people. They want replaceable people.
- itronitron 2y ago>> "6 TiB of data" is not somewhat detailed requirements, as it depends quite a bit on the nature of the data.
- tonetegeatinst 2y agoI'm not even in data science, but I am a slight data hoarder. And heck even I'd just say throw that data on a drive and have a backup in the cloud and on a cold hard drive.
- SkipperCat 2y agoThat makes total sense if you're archiving the data, but what happens when you want to have 10,000 people have access to read/update the data concurrently. Then you start to need some fairly complex solutions.
- kmarc 2y agoThis thread blew up a lot, and some unfriendly commenters made many assumptions about this innocent story. You didn't, and indeed you have a point (missing specification of expected queries), so I expand it as a response here. Among the MANY requirements I shared with the candidate, only one was the 6TiB. Another one was that it was going to be serving as part of the backend of an internal banking knowledge base, with at maximum 100 request a day (definitely not 10k people using it). To all the upset data infrastructure wizards here: calm down. It was a banking startup, with an experimental project, and we needed the sober thinker generalist, who can deliver solutions to real *small scale* problems, and not the one who was the winner on the buzzword bingo. HTH.
- SkipperCat 2y agoThanks for the follow up. I've always felt any questions is good for an interview if it starts a conversation. Your thread did just that so I'd consider it a success!
- citizen_friend 2y agoThis load is well handled by a Postgres instance and 15-25k thrown at hardware.
- paulddraper 2y agoStoring 6TB is easy. Processing and querying it is trickier.
- TeamDman 2y agoWould probably try https://github.com/pola-rs/polars https://github.com/pola-rs/polars and go from there lol
- xLaszlo 2y ago6TB - Snowflake Why? That's the boring solution. If you don't have a use case, what kind of queries you would run then opt for maximum flexibility with the minimum setup of a managed solution. If cost is prohibitive on the long run, you can figure out a more tailored solution based on the revealed preferences. Fiddling with CSVs is the DWH version of the legendary "Dropbox HN commenter".
- nostrademons 2y agoI would've said "Pandas with Parquet files". If you're hiring a DS it's implied that you want to do some sort of aggregate or summary statistics, which is exactly what Pandas is good for, while awk + shell scripts would require a lot of clumsy number munging. And Parquet is an order of magnitude more storage efficient than CSV, and will let you query very quickly.
- atomicnumber3 2y agoIt's really hard because I've failed interviews by pitching "ok we start with postgres, and when that starts to fall over we throw more hardware at it, then when that fails we throw read replicas in, then we IPO, then we can spend all our money and time doing distributed system stuff". Whereas the "right answer" (I had a man on the inside) was to describe some wild tall and wide event based distributed system. For some nominal request volume that was nowhere near the limits of postgres. And they didn't even care if you solved the actual hard distributed system problems that would arise like distributed transactions etc. Anyway, I said I failed the interview, really they failed my filter because if they want me to ignore pragmaticism and blindly regurgitate a YouTube video on "system design" FAANG interview prep, then I don't want to work there anyway.
- metadat 2y agoCan you get a single machine with more than 6TiB of memory these days? That's quite a bit..
- 1vuio0pswjnm7 2y ago"... or a $199 enterprise HDD" External or internal? Any examples? "... it could be loaded (multimple times) to memory" All 6TiB at once, or loaded in chunks?
- gen220 2y agohttps://diskprices.com https://diskprices.com yields https://www.amazon.com/dp/B0C363Y5BQ https://www.amazon.com/dp/B0C363Y5BQ, fwiw. (16TB for $129.99 at time of writing)
- 1vuio0pswjnm7 2y agoThat 1st site is great. Many thanks.
- zeroq 2y agoReminds me an old story from Steve Yegge: I gave him the "find the phone numbers in 50,000 html files" question, and he decided to write a huge program with an ad-hoc state machine. When I asked how long it would take to write the program, he said he'd have to hit N files, with M lines per file, so... I interrupted him: no, WRITE. How long to WRITE the program? Oh. He estimated it at 5 days of work. At this point I was 50% ready to throw him out.
- mewpmewp2 2y agoOn the other hand if salaries are at 300k then 10k compared to that is not a huge cost. If a scalable tool can make you even 10 percent more effective it would be worth 30k.
- fatnoah 2y agoSomewhat ironically, I'm pretty sure I failed a system design interview at a big tech last year for not drastically over building to solve a problem that probably had far less data (movie showtimes, and IIRC there are less than 50k screens in the US).