6 ms·
What Dan got right: Brooks clearly missed the boat when he said, "Well how many MIPS can one use fruitfully?". I'm not sure where the fault lies on that one. A
by tl 6y ago
What Dan got right:
Brooks clearly missed the boat when he said, "Well how many MIPS can one use fruitfully?". I'm not sure where the fault lies on that one. A recent article [1] bemoans the slowness of modern software.
What Dan got wrong:
Both problems discussed were created by computers. A large collection of machines generating logs is possible to analyze, but that possibility comes from massive investment in tools designed to parse, extract and report on linear and structure data. Brooks covers this in "Buy vs. Build", although his example is closer to Excel than ggplot.
Also, if you are dashing off a python script in a single wall-clock minute, you either write insane amounts of Python for purposes similar to this or are way smarter than the average engineer. In Advent of Code (where Python is popular), the fastest solution to the easiest puzzle (Day 1) this year was 0:07:11. [2]
[1] https://news.ycombinator.com/item?id=25386098 https://news.ycombinator.com/item?id=25386098
[2] https://adventofcode.com/2020/leaderboard/day/1 https://adventofcode.com/2020/leaderboard/day/1
- Olreich 6y agoI just tracked my time for doing that Advent of Code day 1. It took 10 minutes from start to finish. Across the two parts of the question, I spent approximately 2 minutes building the scripts, 1 minute determining an algorithm, and 7 minutes reading the prompt. I suspect similar breakdowns for the 7:11 you quote. If Dan excludes the prompt reading and solution determination because he's seen this sort of problem before, 1 minute using tools he's comfortable with doesn't seem unreasonable. Considering all it is is a list of hosts, a loop, a try/except block, and a scp command, using much more than a couple of minutes writing would be surprising to me.
- Jtsummers 6y agoThere was a 6 minute delay in getting the inputs for day one for many people. The fastest part one solution was 35 seconds. And the fastest part two (after finishing part one) was also about 30 seconds. If there hadn’t been a server issue I suspect many of the top 100 would’ve finished both parts in under two minutes
- nemo1618 6y agoIndeed, just look at previous years (Day 1 is always pretty simple): 2015: 00:05:38 (part 1) 00:10:55 (part 2) 2016: 00:03:57 (part 1) 00:07:01 (part 2) 2017: 00:00:57 (part 1) 00:01:16 (part 2) !! 2018: 00:00:26 (part 1) 00:01:48 (part 2) !! 2019: 00:00:23 (part 1) 00:01:39 (part 2) !!
- crispyambulance 6y agoI think Dan has braggadoccio'd his time estimates, or his task is somewhat different from what he describes. I mean, the guy talks fast, like really fast, so I suppose he's quick but mere minutes for something like this doesn't seem realistic unless it's extremely routine. Instead, it seems ad-hoc-ish and exploratory to me, it seems like something that needs to be considered and planned out rather than done between 2 sips of coffee. (I am considering his whole task here, not just the scp'ing of files). He's talking about log files from a few hundred THOUSAND servers that results in several terabytes of data that have to be parsed. He doesn't say exactly what he's looking for, but the point is he's trying to answer some questions about performance for more than a few different applications. Are these simple questions, or involved ones which spawn other questions? We don't know, but even if they're easy questions, there's many applications involved and many servers. Right off the bat, for something THAT BIG, I think it's reasonable to figure out what you're going to do with a sample of logs before downloading "home depot" onto your hard-drive. So this is definitely a multi-pass kind of job: start with a survey, then try a bigger chunk, if everything's OK do the rest. Next, I think it's advisable to consider factors about the servers themselves: the application versions, whether or not the applications were running (and why not), the hardware, the role of the server, whether or not the server was up (and why not). Is this metadata about each server available (can you trust it?) or is it something that has to be queried each time on each server? Dan says this supposed to be a couple of years of data, has each server been through upgrades? when? Is that relevant? We don't know any of these, but they would have to at least be considered for someone doing this task. After the data is parsed there's slicing and dicing to do for the purpose of graphs. That takes lotsa of time-- I am assuming he's not just talking about extracting one figure for each application and plotting it. For someone that is all set-up and on top of things, this seems like something that is a day's work and easily more, not counting follow-up work and validation to further investigate the additional questions that would inevitably (in my opinion) be raised on such a big dataset.
- swiftcoder 6y agoI think you underestimate the value of pipelining here. You could spend time narrowing down the set of logs to download... but in the time it takes to figure that out, you might as well just download them all. Having "home depot" locally available for analysis is never a bad thing, plus you may be racing against time re log rotation, etc. > For someone that is all set-up and on top of things, this seems like something that is a day's work and easily more, not counting follow-up work and validation to further investigate the additional questions that would inevitably (in my opinion) be raised on such a big dataset. In the middle of a SEV, you don't have a day to perform this kind of analysis. 15 minutes till the SLA clock starts ticking and customers are owed refunds.
- deleted 6y ago[deleted]
- adwn 6y ago> What Dan got wrong: Both problems discussed were created by computers. A large collection of machines generating logs is possible to analyze, but that possibility comes from massive investment in tools designed to parse, extract and report on linear and structure data. I don't understand your argument. In what way does that disprove Dan's point?
- taeric 6y agoIt isn't one user's work. This would be a literal technical debt of the modern user. Edit: I meant this as a question. Used wrong punctuation and don't want to ninja edit.
- tl 6y agoAnalysis of data coming from computers (logs coming from servers in this case) is guaranteed to have a structure that makes it easier to process, compared to data coming from more organic sources (reports collected by people, measurements from instruments, etc...) which are closer to the domain Brooks would have dealt with. When dealing with these problems today, the greatest challenge isn't writing a script; it's deciding whether data points that don't fit the model invalidate the data or the model. Moreover, we've spent five decades systemizing the former. Doing the latter is more challenging than ever and fraught with controversy.
- adwn 6y ago> [...] which are closer to the domain Brooks would have dealt with. If anything, that confirms Dan's argument that we're solving problems today that Brooks couldn't even imagine back in the 80s.
- earthboundkid 6y agoReading the article had me thinking: “computers are good at automating… computers.” The quantity of logs he’s describing could only be created by having a machine take readings and dump them somewhere. The cause of the problem had to be a computer in the first place. For human data sources, even back in the 60s we had machines fast enough to tally the census, balance checkbooks, or book flights. So yeah, Dan had a problem of unimaginable scope in the 80s but also the problem is kind of fake? “Help, my computers are spitting out too many numbers and I can’t graph them all!” I feel like Dan was on to something here but also something in the article didn’t quite line up.
- loeg 6y agoFor leaderboarders, and especially the easier challenges, total AoC solve time is largely prompt comprehension. If you already know what you want to do, and that thing is straightforward, it is not unimaginable to bang out a short python script in a minute.
- deleted 6y ago[deleted]