9 ms·
> mostly here on HN Which may or may not be an accurate depiction (as we read personal accounts and thoughts of the commenters) of a quite marginal subset of r
by haylem 10y ago
> mostly here on HN
Which may or may not be an accurate depiction (as we read personal accounts and thoughts of the commenters) of a quite marginal subset of real-life IT-professionals.
I wouldn't worry too much about what's being said or not said on HN. There are great ideas and topics to be covered here for sure, but they're sprinkled on top of a giant cake made with 1-part self-loathing, 2-parts day-dreaming, and 1-part regular huff-and-puffing.
It makes for good entertainment and procrastination.
> No idea what they mean, this books sounds great to me
That being said, not knowing Big-O while doing CS or IT work seems worrying. Sure, it's not absolutely necessary for most of the grunt work. But you should definitely have the same understanding of performance issues without knowing the fancy notation and terminology. Big-O is just a notation and a formalization of these concepts, and it helps with communication. I'd say it's still better to know it.
So, indeed the book probably doesn't hurt.
- teekert 10y agoI'm just a biologist that switched to Python because Excel and Origin weren't dealing very well with my ever increasing pile of data (Typical data: Every row is cell in a Tissue sample, every column is a quantified parameter (size, marker intensity, ...) of that cell, typically I deal with 10s to 100s of tissues samples) Pandas is great, I spend my time turning DataFrames into histograms, scatter plots and ROC curves in Jupyter Notebooks. I have the feeling knowing Big-O is not very relevant. Still, learning new languages, new words and new abstractions is almost guaranteed to influence ones way of working and thinking at some level. So, indeed indeed the book probably doesn't hurt ;) Edit: Just glanced over the link in Practicality's comment about big-O and sure enough I think it may actually be useful as my ever increasing pile of data increases even further! I have to admit; as the parameters increase I find myself doing over night calculations more and more.
- emodendroket 10y agoAsymptotic complexity comes up pretty frequently in bioinformatics contexts because the volumes of data can be huge.
- vonmoltke 10y agoHow does it change what you do, though? I did signal processing with massive data streams rather than bioinformatics, but I assume the situation is similar. The algorithms are what they are. They are complex mathematical equations or transformations that need to be run on data and are often optimized without being able to change their asymptotic complexity.
- emodendroket 10y agoSkiena's Algorithm Design Manual mentions him being brought in as an algorithmic consultant to modify some genetics analysis software so that it'd actually finish but I don't really remember the details or know enough about the field to give you plausible examples.
- vonmoltke 10y agoI can see that; I did a lot of similar work with signal processing algorithms. None of what I did affected asymptotic complexity at all, though. The asymptotic complexity was tied to the algorithms chosen, and changing those was an issue of trading computational performance for system performance.
- deleted 10y ago[deleted]
- emodendroket 10y agoI pulled it up and found it; the problem involved modeling genome sequences as strings and finding possible substrings.
- Jtsummers 10y agoEDIT: Misstated the big-O, in this particular case (should've found my coworkers actual code). Both are O(m x n), one just has a large constant. Here's a pattern I've noticed with code written for processing a data file by a lot of people (python-esque, using a function (match) that's "left as an exercise for the reader" to implement): def search(filename, value): with open(filename, "r") as f: for line in f: if match(value,line): print(line) # we don't care about not matching def main(): for v in [search1, search2, search3, ...]: search("data.dat", v) What happened is that one time they needed that search function, and so they made search and it worked well. They realized they could run that same search function repeatedly, and for small data files and few searches it was quick enough. But the performance is O(m x n) [EDIT: originally wrote O(m x n)], where m is the number of lines, n is the number of search values. [EDIT: wrong: a second m because it takes a time proportional to the size of the file to read the file.] The data file is read every time something is searched. If you've got an SSD, it's not really noticeable. If you've got a spinning disk, it becomes a problem. If you're hitting network storage, you're downloading that file n times. The main issue being that each read (each iteration of the inner for) hits the hard drive, network, or similar. A simple performance hack is to move the read into main, put the whole thing into one list of lines and pass that list to search instead of the filename (modifying search appropriately): def search(data, value): for line in data: if match(value,line): print(line) # we don't care about not matching def main(): with open("data.dat", "r") as f: data = f.read().splitlines() for v in [search1, search2, search3, ...]: search(data, v) It's still O(m x n) [EDIT: It's now O(m x n). We've removed one of the m factors because we do the read once, and never again.] For very large files and very large search parameter lists, this will still take a long time, but it's much faster than the previous version when you're dealing with large files. EDIT: Shortest code I can think of to get the actual worst case that I've had a few coworkers pull off: def search(filename, value): with open(filename, "r") as f: data = f.read().splitlines() for line in data: if match(value,line): print(line) # we don't care about not matching def main(): for v in [search1, search2, search3, ...]: search("data.dat", v) With, of course, other code in between because as vonmoltke points out, the above has clear problems. My point was about the structure of the bad pattern, not the specific implementation of it.
- lafay 10y agoThis is the best description of HN I've ever read: > There are great ideas and topics to be covered here for sure, but they're sprinkled on top of a giant cake made with 1-part self-loathing, 2-parts day-dreaming, and 1-part regular huff-and-puffing.
- dineshp2 10y ago> I wouldn't worry too much about what's being said or not said on HN. There are great ideas and topics to be covered here for sure, but they're sprinkled on top of a giant cake made with 1-part self-loathing, 2-parts day-dreaming, and 1-part regular huff-and-puffing. I don't understand your argument regarding why you would not pay much attention to what is being said on HN. Could you explain?
- blowski 10y agoProbably because there is a noisy minority who voice dogmatic opinions without understanding the constraints of the problem at hand. They are lilliputians, spouting wonderful ideas that collapse in the face of deadlines and budgets. For those of us that have to live in reality, they can be very annoying and disheartening, so it's essential to take their opinions with a fistful of salt.
- haylem 10y agoThat, yes. But also because a significant amount of even the day's top popular topics have in the end very little relevance and impact to most businesses. Which doesn't mean it's not interesting, though.