11 ms·
I've done a handful of interviews recently where the 'scaling' problem involves something that comfortably fits on one machine. The funniest one was ingesting
by jesse__ 9mo ago
I've done a handful of interviews recently where the 'scaling' problem involves something that comfortably fits on one machine. The funniest one was ingesting something like 1gb of json per day. I explained, from first principals, how it fits, and received feedback along the lines of "our engineers agreed with your technical assessment, but that's not the answer we wanted, so we're going to pass". I've had this experience a good handful of times.
I think a lot of people don't realize machines come with TBs of RAM and hundreds of physical cores. One machine is fucking huge these days.
- coliveira 9mo agoYes, but then how are these people going to justify the money they're spending on cloud systems?... They need to find only reasons to maintain their "investment", otherwise they could be held as incompetent when their solution is proven to be ineffective. So, they have to show that it was a unanimous technical decision to do whatever they wanted in the first place.
- kevmo314 9mo agoThe wildest part is they’ll take those massive machines, shard them into tiny Kubernetes pods, and then engineer something that “scales horizontally” with the number of pods.
- jesse__ 9mo agoYeah man, you're running on a multitasking OS. Just let the scheduler do the thing.
- mystraline 9mo agoIts all fun and games, until the control plane gets killed by the OOMkiller. Naturally, that detaches all your containers. And theres no seamless reattach for control plane restart.
- dgxyz 9mo agoOr your CNI implementation is made of rolled up turds and you lose a node or two from the cluster control plane every day. (Large EKS cluster)
- zacmps 9mo agoUntil you need to schedule GPUs or other heterogenous compute...
- jesse__ 9mo agoAre you saying that running your application in a pile of containers somehow helps that problem ..? It's the same problem as CPU scheduling, we just don't have good schedulers yet.. Lots of people are working on it though
- zacmps 9mo agoNot really? At the moment it's done by some user-land job scheduler. That could be something container based like k8s, something in-process like ray, or a workload manager like slurm.
- dgxyz 9mo agoYeah this. As I explain many times to people, processes are the only virtualisation you need if you aren’t running a fucked up pile of shit. The problem we have is fucked up piles of shit not that we don’t have kubernetes and don’t have containers.
- jesse__ 9mo agoHahhah, yuuuup. I can maybe make a case for running in containers if you need some specific security properties but .. mostly I think the proliferation of 'fucked up piles of shit' is the problem.
- jchw 9mo agoContainers are just processes plus some namespacing, nothing really stops you from running very huge tasks on Kubernetes nodes. I think the argument for containers and Kubernetes is pretty good owing to their operational advantages (OCI images for distributing software, distributed cron jobs in Kubernetes, observability tools like Falco, and so forth). So I totally understand why people preemptively choose Kubernetes before they are scaling to the point where having a distributed scheduler is strictly necessary. Hadoop, on the other hand, you're definitely paying a large upfront cost for scalability you very much might not need.
- dgxyz 9mo agoTime to market and operational costs are much higher on kubernetes and containers from many years of actual experience. This is both in production and in development. It’s usually a bad engineering decision. If you’re doing a lift and shift, it’s definitely bad. If you’re starting greenfield it makes sense to pick technology stacks that don’t incur this crap. It only makes sense if you’re managing large amounts of large siloed bits of kit. I’ve not seen this other than at unnamed big tech companies. 99.9% of people are just burning money for a fashion show where everyone is wearing clown suits because someone said clown suits are good.
- bartread 9mo agoThanks. You’ve reassured me that I’m not going mad when I look at our project repo and seriously consider binning the Dockerfile and deploying direct to Ubuntu. The project is a Ruby on Rails app that talks to PostreSQL and a handful of third party services. It just seems unnecessary to include the complexity of containers.
- ahartmetz 9mo agoI think my brain hurts
- andai 9mo agoI had to re-read this a few times. I am sad now.
- cyberpunk 9mo agoTo be fair each of those pods can have dedicated, separate external storage volumes which may actually help and it’s def easier than maintaining 200 iscsi or more whatever targets yourself
- jayd16 9mo agoI mean, a large part of the point is that you can run on separate physical machines, too.
- pnt12 9mo agoThis is especially aggravating when the os inside the container and the language runtimes are much heavier than the process itself. I've seen arguments for nano services (I wouldn't even call them micros services), that completely ignored that part. Split a small service in n tiny services, such that you have 10(os, runtime, 0.5) rather than 2(os, runtime, x).
- SpaceNugget 9mo agoThere is no os inside the container. That's a big part of the reason containerization is so popular as a replacement for heavier alternatives like full virtualization. I get that it's a bit confusing with base image names like "ubuntu" and "fedora", but that doesn't mean that there is a nested copy of ubuntu/fedora running for every container.
- yieldcrv 9mo ago“there’s no wrong answer, we just want to see how you think” gaslighting in tech needs to be studied by the EEOC, Department of Labor, FTC, SEC, and Delaware Chancery Court to name a few let’s see how they think and turn this into a paid interview
- ahartmetz 9mo agoEvery one of these cores is really fast, too!
- jesse__ 9mo agoyeah man, computers are completely bananacakes
- yndoendo 9mo agoI recently had to parse 500MB to 2GB daily log files into analytical information for sales. Quick and dirty, the application would of needed 64GB RAM and work laptop only has 48GB RAM. After taking time cleaning it up, it was using under 1GB of RAM and worked faster by only retaining records in RAM if need be between each day. It is not about what you are doing, it is always about how you do it. This was the same with doing OCR analysis of assembly and production manuals. Quick and dirty, it would of took over 24 hours of processing time, after moving to semaphores with parallelization it took less than two hours to process all the information.
- maest 9mo ago> It is not about what you are doing, it is always about how you do it. It saddens me to see how the LinkedIn slop style is expanding to other platforms
- bauerd 9mo agoIn interviews just give them what they are looking for. Don't overthink it. Interviews have gotten so stupidly standardized as the industry at large copied the same Big Tech DSA/System Design/Behavioral process. And therefore interview processes have long been decoupled from the business reality most companies face. Just shard the database and don't forget the API Gateway
- jesse__ 9mo agoMeh .. I've played that game; it doesn't work out well for anyone involved. I optimize my answers for the companies I want to work for, and get rejected by the ones I don't. The hardest part of that strategy is coming to terms with the idea that I constantly get rejected by people that I think are mostly <derogatory_words_here>, but I've developed thick skin over the years. I'd much rather spend a year unemployed (and do a ton of painful interviews) and find a company who's values align with mine, than work for a year on a team I disagree with constantly and quit out of frustration.
- bauerd 9mo agoThe company's values may align to yours, even though they reject you. It's because the interview process doesn't need to have anything to do with their real-world process. Their engineers probe you for the same "best practices" that they themselves were constantly probed for in their own interviews. Interviewing is its very own skill that doesn't necessarily translate into real-life performance.
- jesse__ 9mo agoI agree with your observation. My issue is (from experience) it's really hard to tell from the outside if a teams' values align with mine. Many teams talk the talk, but don't walk the walk, as the saying goes. It's just easier to not participate than it is to guess, and be wrong. I also believe that running a broken interview process actively selects for qualities you actually don't want, so it's much more likely that teams conducting those interviews aren't teams I want to work on. Edit: As credence for my claims, the best team I've ever worked on was a team I did 90%+ of the hiring for, and we didn't do any of the 'typical' interview bullshit most companies do. What we did instead was sit people down and have deep technical conversations about systems they'd worked on in the past. The candidate would explain, in as much detail as they could muster, a system they'd worked on in the past, down to the lowest level details. Usually, they would talk to us for at least 20-30 minutes, then, we (the interviewers) would pose questions, usually starting with the form 'if we changed X, what effect would it have'. Doing interviews in this style make a few things immediately obvious: 1. Did the candidate have a deep, systemic understanding of the system they worked on? 2. Does the candidate have a good mental model for evaluating change in the system? That's how I conduct interviews, and unsurprisingly, when I get interviewed like that, my success rate is 100%. I don't think I've ever done an interview like that which did not result in an offer. Anyways, there's some rambling and unsolicited opinions for you :)
- dehrmann 9mo ago> but that's not the answer we wanted You could have learned this if you were better about collecting requirements. You can tell the interviewer "I'd do it like this for this size data, but I'd do it like this for 100x data. Which size should I design this for?" If they're looking for one direction and you ask which one, interviewers will tell you.
- jesse__ 9mo agoI've done that too and, in my experience, people that ask a scaling question that fits on a single machine don't have the capacity to have that nuanced conversation. I usually try to help the interviewer adjust the scale to something that actually requires many machines, but they usually don't get it. Said another way, how do you have a meaningful conversation about scaling with a person who thinks their application is huge, but in reality only requires a tiny fraction of a single machine? Sometimes, there's such a massive gulf between perception and reality that the only thing to do is chuckle and move on.
- esafak 9mo agoThe burden of wisdom.
- badgersnake 9mo agoThis kind of bad interview is rife. It’s often more a case of guess what the interviewer thinks than come up with a good solution.
- colechristensen 9mo agoYeah I had this problem at a couple of times in startup interviews where the interviewer asked a question I happened to have expertise in and then disagreed with my answer and clearly they didn't know all that much about it. It's ok, they did me a favor. It may or may not be related that the places that this happened were always very ethnically monotone with narrow age ranges (nothing against any particular ethnic group, they were all different ethnic monotones)
- jesse__ 9mo agoHah yeah, that's a funny one, being able to run circles around the interviewer.
- jitl 9mo ago1gb of json u can do in one parse ¯\_(ツ)_/¯ big batches are fast
- ytoawwhra92 9mo ago> I explained, from first principals, how it fits, and received feedback along the lines of "our engineers agreed with your technical assessment, but that's not the answer we wanted, so we're going to pass". I've had this experience a good handful of times. Probably a better outcome than being hired onto a team where everyone know you're technically correct but they ignore your suggestions for some mysterious (to you) reason.
- jesse__ 9mo agoOh, absolutely.
- franciscop 9mo agoI have a funny story I need to tell some day about how I could get a 4GB JSON loaded purely in the browser at some insane speed, by reading the bytes, identifying the "\n" then making a lookup table. It started low stakes but ended up becoming a multi-million internal project (in man-hours) that virtually everyone on the company used. It's the kind of project that if started "big" from the beginning, I'd bet anything it wouldn't have gotten so far. Edit: I did try JSON.parse() first, which I expected to fail and it did fail BUT it's important that you try anyway.
- mr_toad 9mo agoCurious about which browser and hardware. In my experience browsers often choke on 0.5GB strings, or decide to kill the tab/proccess.
- franciscop 9mo agoYes, but I didn't read the full file, I kept the File reference and read the bytes in pages of 10MB IIRC to find all of the line break offsets. Then used those to slice and only read the relevant parts.
- sharadov 9mo agoYes, yes but how are we going to get HA with one machine.. Fuck off ..you're 10 person startup with an MVP and no revenue stream needs customers first..
- winrid 9mo agoI've actually worked on distributed systems that were so broken, I created a script to connect to prod and just create the report from my laptop. My manager offered to buy me a second laptop for running the report since it was easier than getting approval from the architects to get rid of the distributed report system (it only created that one report).
- anshumankmr 9mo agoThough I do not know the situation AT the firm you were interviewing in, if there is some unexpected increase in data volume OR say a job fails on certain days or you need to do some sort of historical data load (>= 6 months of 1 gig of data per day), the solution for running it on a single VM might not scale. BUT again, interviews are partially about problem solving, partially about checking compliance at least for IC roles (IN my anecdotal experience). That being said yeah I too have done some similar stuff where some data engineering jobs could be run on a single VM but some jobs really did need spark, so the team decision was to fit the smaller square peg into a larger square peg and call it a da.In fact, I had spent time refactoring one particular pivotal job to run as an API deployed on our "macrolith" and integrated with our Airflow but it was rejected, so I stopped caring about engineering hygiene.
- ahoka 9mo ago“6 months of 1 gig of data per day” Then you would need an enormous 2TB storage device. \s
- wongarsu 9mo agoIf we are talking about cloud VMs: sure, their cpu performance is atrocious and io can be horrible. This won't scale to infinity But if there's the option to run this on a fairly modest dedicated machine, I'd be comfortable that any reasonable solution for pure ingest could scale to five orders of magnitude more data, and still about four orders of magnitude if we need to look at historical data. Of course you could scale well beyond that, but at that point it would be actual work
- johndough 9mo ago(>= 6 months of 1 gig of data per day) You can parse JSON at several GB/s: https://github.com/simdjson/simdjson https://github.com/simdjson/simdjson And you could scale that by one or two orders of magnitude with thread-based parallelism on recent AMD Epyc or Intel Xeon CPUs. So parsing alone should not pose a problem (maybe even sub-second for 6 months of data). We would need a more precise problem statement to judge whether horizontal scaling is needed.
- ahoka 9mo agoThey wanted to see if you would be on board with their embezzlement scheme.