16 ms·
The most copied StackOverflow snippet of all time is flawed (2019)
- corbezzoli 3y agoWhy do you need a 4-line dependency? This is the reason.
- bauruine 3y agoThere is still the chance that the person that created the 4 line dependency also just copy pasted it from the flawed StackOverflow answer. Or is the same person or is also just a random person creating the package like the random person that created the SO answer. I'm not sure why random_person1 should be more trustworthy to produce non flawed code than random_person2. OTO: It's at least easily upgrade able so it has an advantage.
- corbezzoli 3y ago> There is still the chance There's no chance if you avoid random_person1 and use known_oss_provider’s package instead. At the very least, look at the tests. Any package with tests is guaranteed to be more correct than a never-before-run SO answer.
- envsubst 3y agoWhat if you write the code and test in your project?
- JimDabell 3y agoThere is still the chance. As the article states, OpenJDK copied from the Stack Overflow answer.
- ziml77 3y agoSure, but if OpenJDK is exposing that function then anyone who is using it will get the correct output when OpenJDK fixes the problem. If everyone copies the function into their own code then in many cases it's likely to never be corrected.
- ComputerGuru 3y agoShameless plug: another option to format sizes in a human readable format quickly and correctly (other than copying from S/O), you can use one of our open source PrettySize libraries, available for rust [0] and .NET [1]. They also make performing type-safe logical operations on file sizes safe and easy! The snippet from S/O may be four lines but these are much more extensive, come with tests, output formatting options, conversion between sizes, and more. [0]: https://github.com/neosmart/prettysize-rs https://github.com/neosmart/prettysize-rs [1]: https://github.com/neosmart/PrettySize.net https://github.com/neosmart/PrettySize.net
- drunkendog 3y agoReplacing 4 line solutions with extensive libraries is what caused left-pad.
- LoganDark 3y agoNo. left-pad was placing a 4-line solution in a library. prettysize is well deserving of library status.
- eviks 3y agoWhat caused left-pad is the the ability to delete published code
- Cthulhu_ 3y agoYeah, copying an incorrect answer from SO thousands of times is much better! (The subject at hand isn't whether libraries are good or not, it's whether copying something off the internet is. In the post, it turns out it isn't. If it was a library, the author could have fixed and updated the library, and the issue would be fixed for everyone that uses it. left-pad isn't an issue with libraries per se, it's an issue with library management)
- Analemma_ 3y agoI understand where you're coming from here, but the whole point of this article is at the 4-line solution is wrong (and the author specifically mentioned that every other answer on the stack overflow post was wrong in the same way as well). "Seemingly-simple problem where every naïve solution contains a subtle bug" is exactly the right use case for a well-designed library method.
- nelsonic 3y agoHow does the author determine this is the "most copied snippet" on SO? The Question/Answer has only been Viewed 351k times. There are posts with many millions of views e.g: https://stackoverflow.com/questions/927358/how-do-i-undo-the-most-recent-local-commits-in-git https://stackoverflow.com/questions/927358/how-do-i-undo-the... which have definitely been copy-pasted more times. Yes, there may be many instances of this Java function on GitHub. But only because the people doing the copying are too lazy to think about how it works never mind alter the function name. If there's a bug, just update the SO answer and fix the problem. No need to write a lengthy self-promoting post about it.
- deleted 3y ago[deleted]
- _fizz_buzz_ 3y agoThird paragraph of the post: It's according to this paper: https://link.springer.com/article/10.1007/s10664-018-9650-5 https://link.springer.com/article/10.1007/s10664-018-9650-5
- deleted 3y ago[deleted]
- moribunda 3y agoIt's described in the article...
- bloak 3y agoThis reminds me of a weirdness with some sat navs: the distance to your exit/destination is displayed as: 12 ... 11 ... 10 ... 10.0 ... 9.9 ... 9.8 ... with the value 10.0 shown only while the distance is between 9.95 and 10. It's not really a bug but it's strange seeing the display update from 10 to 10.0 as you pass the imaginary ten-mile milestone so perhaps it's a distraction worth avoiding.
- bombcar 3y agoMercedes for awhile had a fuel gauge that showed 1/4 1/2 3/4 1/1 They had another one that went R 2/4 4/4 I'm still undecided which was more weird. You can see them both on eBay.
- elzbardico 3y agoThere's nothing weird here. Those are very common fractions used across several domains, including cooking. But one thing that I would really love to see are actual liters or gallons (depending on the country where I am at the moment).
- roryokane 3y ago(2019) Past discussions: https://news.ycombinator.com/item?id=21693431 https://news.ycombinator.com/item?id=21693431 https://news.ycombinator.com/item?id=21698619 https://news.ycombinator.com/item?id=21698619 https://news.ycombinator.com/item?id=27533684 https://news.ycombinator.com/item?id=27533684
- dang 3y agoThanks! Macroexpanded: The most copied StackOverflow snippet of all time is flawed (2019) - https://news.ycombinator.com/item?id=27533684 https://news.ycombinator.com/item?id=27533684 - June 2021 (334 comments) The most copied StackOverflow snippet of all time is flawed - https://news.ycombinator.com/item?id=21698619 https://news.ycombinator.com/item?id=21698619 - Dec 2019 (88 comments) The most copied StackOverflow snippet of all time is flawed - https://news.ycombinator.com/item?id=21693431 https://news.ycombinator.com/item?id=21693431 - Dec 2019 (3 comments)
- speak_plainly 3y agoSounds like someone bumped into Zeno's paradox... https://www.youtube.com/watch?v=VI6UdOUg0kg https://www.youtube.com/watch?v=VI6UdOUg0kg
- dleeftink 3y agoKnowledge cascades all the way down; it goes to show how difficult it is to 'holster' even the smallest piece of knowledge once its drawn. I wonder with the rate Stack Exchange is losing active contributors, what it would take for 'fastest gun' answers to be corrected that are later found to be off mark, and what it would mean for our collective knowledge once these 'slightly off' answers are further cemented in our annals of search and increasingly, LLM history.
- oooyay 3y agoOut of curiosity, is there a sizable number of developers that just copy and paste untrusted code from StackOverflow into their applications? The conjecture that people just copy from StackOverflow is obviously popular but I always thought this was just conjecture and humor until I saw someone do it. Don't get me wrong, I use StackOverflow to give me a head start on solving a problem in an area I'm not as familiar with yet, but I've never just straight copied code from there. I don't do that because rarely does the snippet do exactly and only exactly what I need. It requires me to look at the APIs and form my own solution from the explained approach. StackOverflow has pointed me in the direction of some niche APIs that are useful to me, especially in Python.
- foobarian 3y agoWell. You (collective you) start by copying and pasting a code snippet first, and then modifying it as needed. Does that count? If no modifications are needed, then it stays.
- SwiftyBug 3y agoThat's what I do. I almost always rename things to match the coding style of the codebase I'm working on, though.
- TrackerFF 3y agoMillions.
- ahoka 3y agoYes, people do that. After looking at a huge number of incorrect TLS related code and configuration at SO, I’m now pretty sure that most systems run without validating certificates properly.
- PhilipRoman 3y agoTo be fair that might be partly the fault of TLS libraries. There should be a single sane function that does the least surprising thing and then lower level APIs for everything else. Currently you need a checklist of things that must be checked before trusting a connection.
- greenhearth 3y agoPretty awesome stuff. This is what Hacker News is for!
- greenhearth 3y agowtf, why would someone downvote this? This is prime hacker News shit and why I come here!
- recursivecaveat 3y agoI didn't downvote, but I would guess it's due to the general idea that if you just approve or disapprove of a post you should simply vote that way instead of expressing it in a comment. Personally while I agree there's a logic to that, I find it a little cold for positive sentiments. I couldn't find it, but I think that's there's a PG or Dang comment to the effect that "I like this" as a comment is explicitly not discouraged on HN, but obviously that doesn't mean everyone agrees.
- greenhearth 3y ago[flagged]
- seeknotfind 3y agoI was surprised to find log implementations are loopless. Cool. https://github.com/lattera/glibc/blob/master/sysdeps/ieee754/dbl-64/e_log.c https://github.com/lattera/glibc/blob/master/sysdeps/ieee754...
- zeroonetwothree 3y agoIt basically has the loop unrolled. But it looks like it’s evaluating a polynomial approximation so I suppose it makes sense
- throwaway9870 3y agoI don't understand. There are 7 suffixes, can't you pick the right one with binary search? That would be 3 comparisons. Or just do it the dumb way and have 6 comparisons. How are two log() calls, one pow() call and ceil() better than just doing it the dumb way? The bug being described is a perfect example of trying to be too clever.
- emerongi 3y agoThe author apparently went back to using a loop after recognizing that it's not readable: https://programming.guide/java/formatting-byte-size-to-human-readable-format.html https://programming.guide/java/formatting-byte-size-to-human... Notably, it's still slightly better than the first code example in the original article, as it takes the rounding bug into account.
- zeroonetwothree 3y agoThe author says at the beginning that it’s not actually better than the loop. Also 6 comparisons is only if you’d have the max value which seems unlikely in actual usage. Linear could be better if most of the time values are in B or KB ranges
- totallywrong 3y agoRead: The most common answer to that question from LLMs is flawed.
- envsubst 3y agoAlmost every top stack overflow answer is wrong. The correct one is usually at rank 3. The system promotes answers which the public believes to be correct (easy to read, resembles material they are familiar with, follows fads, etc). Pay attention to comments and compare a few answers.
- emerongi 3y agoBack in ye olden days, almost every answer involving a database contained a SQL injection vulnerability.
- buffet_overflow 3y agoIf you ever have an issue with the Requests library in Python, just try again with verify=false.
- psd1 3y agoEasier than getting the app team to fix their TLS.
- WorldMaker 3y agoOr the corporate IT team to remove their TLS-trashing MITM attack (because their Firewall Vendor claims that's still "Best Practice" in 2023 and/or the C-Suite loves employee surveillance).
- Aeolun 3y agoAt least node has a variable to disable checks globally.
- capableweb 3y agoJust be sure to try running the program with sudo first, before trying shitty solutions like that.
- 3y ago
- golol 3y agoClassic off by 1 :)
- loeg 3y agoShould have just stuck with the loop. You could change the thresholds to 95% of 10^whatever to accommodate the desired output rounding.
- marginalia_nu 3y agoI don't understand why you'd use floating point logarithms if you want log 2? Unless I'm missing something, this gives you an accurate value of floor(log2(value)) for anything positive less than 2^63 bytes, and it's much faster too: Long.bitCount( (Long.highestOneBit(value) << 1) - 1) - 1
- zeroonetwothree 3y agoThe “common” units are powers of 10 so this doesn’t work
- lifthrasiir 3y agoBut you can avoid binary search because there are at most one power of tens between 2^k and 2^(k+1). So you can turn it into a lookup table problem.
- morpheuskafka 3y agoThe original SO question did actually state they wanted powers of two (kilobyte as 1024 bytes). Although, they should have used KiB, GiB, instead to be pedantic.
- nathan_gold 3y agoI'm curious what answer GPT will return.
- Denote6737 3y agoGiven how unreliable it is probably, 418 - I'm a teapot.
- elzbardico 3y agoProbably this one as this is the most common on the corpus it was used to train it.
- chriscosma 3y agoGPT-3.5 returns: public static String convertBytes(long bytes) { String[] suffixes = {"B", "KB", "MB", "GB", "TB", "PB", "EB", "ZB", "YB"}; if (bytes < 1024) return bytes + " " + suffixes[0]; int exp = (int) (Math.log(bytes) / Math.log(1024)); return String.format("%.2f %s", bytes / Math.pow(1024, exp), suffixes[exp]); }
- TacticalCoder 3y agoI find it interesting that all the answers using hardcoded values / if statements (or while) are all doing up to five comparisons. It goes B, KiB, MiB, GiB, TiB, EiB and no more than that (in all the answers) so that can be solved with three if statements at most, no five. I mean: if it's greater or equal to GiB, you know it won't be B, KiB or MiB. Dichotomy search for the win! Not a single of the hardcoded solutions do it that way. Now let's go up to ZiB and YiB: still only three if statements at most, vs up to seven for the hardcoded solutions. I mention it because I'd personally definitely not go for the whole log/pow/floating-points if I had to write a solution myself (because I precisely know all too well the SNAFU potential). I'd hardcode if statements... But while doing a dichotomy search. I must be an oddball. P.S: no horse in this race, no hill to die on, and all the usual disclaimers
- throwaway9870 3y agoYour comment and mine are basically the same. This is what I call terrible engineering judgement. A random co-worker could review the simple solution without much effort. They could also see the corner cases clearly and verify the tests cover them. With this code, not so much. It seems like a lot of work to write slower, more complex, harder to test and harder to review code.
- zeroonetwothree 3y agoIt depends on the input distribution. If it’s very common to have smaller values then the linear search could be superior.
- IshKebab 3y agoI would expect your binary search solution is possibly slower than just doing 6 checks because the latter is only going to take 1 branch. Branching is very slow. You want to keep code going in a straight line as much as possible.
- johnnyanmac 3y agoYup, know your hardware and know problem. Dichotomic search is wonderful when your data can't fit in RAM and it starts being more efficient to cut down on number of nodes traversed. for a problem space limited by your input size (signed 64 bit number) to a 6 entry dictionary? At best you may want to optimize some in-lining or compiler hints if your language supports it. maybe setup some batching operations if this is called hundreds of times a frame so you're not creating/desrtoying the stack frame everytime (even then, the compiler can probably optimize that). But otherwise, just throw that few dozen byte lookup table into the registers and let the hardware chew through it. Big N notations aren't needed for data at this scale.
- derstander 3y agoI feel like there ought to be a software analogue to that aphorism about models (if it doesn’t exist already) — maybe something like: All code is wrong, but some is useful.
- adolph 3y agoAgreed, but is code not a model?
- strangesmells02 3y agojust divide by 1000 until x < 1000 and return int(x) plus a map of number of times divided by 1,000 to MB, GB,... string. Its a O(1) operation because of limited size allowed for numeric types
- instamail 3y agoObligatory, my favourite StackOverflow answer of all time: https://stackoverflow.com/a/1732454 https://stackoverflow.com/a/1732454
- zeroonetwothree 3y agoAnd yet it’s wrong like all the rest
- robertlagrant 3y agoHow so?
- didntcheck 3y agoThe answer is amusing, but it seems the author either didn't read the question properly, or didn't read their formal languages textbook properly, and rushed ahead with an answer that isn't really correct For one thing, It assumes "regex" as used in programming are the same as "regular expressions" (defining regular languages) in formal use. More info on that [1] But the question isn't even about a full parsing of HTML, with bracket balancing. It's just about syntactically matching all the opening tags. More "lexing" than "parsing". Instinctively that does look like a simple regular language to me, though I'm not claiming certainly. The super-regularity of HTML comes from nested elements, but it's just the tag syntax this user cares about, with no context-sensitivity One red herring is comments and CDATA sections, but since they cannot be nested, they do not change the language class, as you just transition to a skip state and back when you see the start/end markers. But they do make the expression much more ugly of course [1] https://en.wikipedia.org/wiki/Regular_expression#Patterns_for_non-regular_languages https://en.wikipedia.org/wiki/Regular_expression#Patterns_fo...
- bradley13 3y agoWhen StackOverflow was new, it was an incredible resource. Unfortunately, so much cruft has accumulated that it is now nearly useless. Even if an answer was once correct (and many are not), it is likely years out of date and no longer applicable.
- jprete 3y agoI took one look at the snippet, saw a floating-point log operation and divisions applied to integers, and mentally discarded the entire snippet as too clever by half and inherently bug-prone.
- zeroonetwothree 3y agoThat’s basically the point of the article
- stmblast 3y agoWell - I suppose it makes sense. SO isn't built for correctness, it's built for upvotes that just depend on whether the people upvoting like the answer or not (regardless of correctness).
- dirtyv 3y agoThis reminds me of when I was in basic training. The drill sgts would give us new recruits a task that none of us knew how to do, purposefully without guidance, and then leave. One guy would try and start doing it, always the incorrect way, and everyone else would just copy that person.
- nomilk 3y agoI wonder if this is exacerbated by human tendencies to not want to look bad relative to others, even if it leads to silly outcomes like intelligent people following a bad or rushed idea. Something similar happens in public economic forecasts because those who get it wrong when others get it right are treated much more harshly than those who get it wrong when others get it wrong too.
- zeroonetwothree 3y agoWhat was the goal of this?
- Kwpolska 3y agoThe usual goal of anything in military training, being cruel to new recruits?
- didntcheck 3y ago"Don't jump off a cliff just because everyone else is doing it" basically I guess the next logical exercise would be asking them to do something with instructions that are complete, but incorrect or at least inefficient, to teach the lesson of questioning superior orders rather than just peers. Actually, I'm honestly not sure it that's desired in military discipline or not (no direct experience here)
- PH95VuimJjqBqy 3y agoI drove a forklift one summer for a manufacturing plant. I had a supervisor tell me to do something that was clearly not right and I refused. I came in the next day and they tried to write me up and I refused to sign the paperwork for it. The one thing no one could accurately describe is why the supervisor was right. I agree with the idea of being willing to go against authority but disagree that it's always a good career move :) Of course it was easier for me, it was just a summer job, I was going back to Uni in the fall.
- koromak 3y agoIn a way, I don't even consider floating point errors to be "flaws" with an algorithm like this. If the code defines a logical, mathematically correct solution, then its "right". Solving floating point errors is a step above this, and only done in certain circumstances where it actually matters. You can imagine some perfect future programming language where floating point errors don't exist, and don't have to be accounted for. Thats the language I'm targeting with 99% of my algorithms.
- ludwigvan 3y agoPlot twist: they were hired by Oracle since they were the author of the most copied StackOverflow snippet (!)
- crabbone 3y agoLong time ago, when ActionScript was a thing, there was this one snippet in ActionScript documentation that illustrated how to deal with events dispatching, handling etc. In order to illustrate the concept the official documentation provided a code snippet that created a dummy object, attached handlers to it, and in those handlers defined some way of processing... I think it was XML loading and parsing, well, something very common. The example implied that this object would be an instance of a class interested in handling events, but didn't want to blow up the size of this example with not so relevant bits of code. There was a time when I very actively participated in various forums related to ActionScript. And, as you can imagine, loading of XML was paramount to success in that field. Invariably, I'd encounter code that copied the documentation example and had this useless dummy object with handlers defined (and subsequently struggled to extract information thus loaded). It was simply amazing how regardless of the overall skill of the programmer or the purpose of the applet, the same exact useless object would appear in the same situation -- be it XML socket or XML loaded via HTTP, submitted and parsed by user... it was always there. ---- Today, I often encounter code like this in unit tests in various languages. Often programmers will copy some boilerplate code from example in the manual and will create hundreds or even thousands of unit tests all with some unnecessary code duplication / unnecessary objects. Not sure why in this specific area, but it looks like programmers both treat these kinds of test as some sort of magic but also unimportant, worthless code that doesn't need attention. ---- Finally, specifically on the subject of human-readable encoding of byte sizes. Do you guys like parted? Because it's so fun to work with it because of this very issue! You should try it, if you have some spare time and don't feel misanthropic enough for today.
- dmccarty 3y agoProcessors are inherently awesome at branching, adding, adding, shifting, etc. And shifting to get powers of 2 (i.e., KB vs. GB) is a superpower of its own. They're a little less awesome when it comes to math.pow(), math.log(), and math.log() / math.log(). Why 300K+ people copied this in the first place shows some basic level of ignorance about what's happening under the hood.[1] As someone who's been at this for decades now and knows my own failings better than ever, it also shows how developers can be too attracted by shiny things (ooh look, you can solve it with logs instead, how clever!) at the expense of readable, maintainable code. [1] But hey, maybe that's why we were all on StackOverflow in the first place
- feoren 3y ago[flagged]
- sclangdon 3y agoIt's not that we think it's arcane or that we are in our own "bubbles of thought", it's that we aren't doing math. We're programming a computer. And a competent programmer would know, or at least suspect, that doing it with logarithms will be slower and more complicated for a computer. The author even points out that even he wouldn't use his solution. P.S. Please look up the word literally.
- WorldMaker 3y agoThe author's final suggested solution at the bottom of the article still relies on logarithms. > doing it with logarithms will be slower and more complicated for a computer This is a fascinating point of view and while it isn't wrong in certain "low-level optimization golf" viewpoints is in part based on old wrong assumptions from early chipsets that haven't been true in decades. Most FPUs in modern computers will do basic logarithms in nearly as many cycles as any other floating point math. It is marvelous technology. That many languages wrap these CPU features in what look like library function calls like Math.log() instead of having some sort of "log operator" is as much an historic accident of mathematical notation and that logarithms were extremely slow for a human. Logarithms used to be the domain of lookup books (you might have one or more volumes, if not a shelf-full) and was one of the keys to the existence of slide rules and why an Engineer would actually have a set of slide rules in different logarithmic bases. Mathematicians would spend lifetimes doing the complex calculations to fill a lookup book of logarithmic data. Today's computers excel at it. Early CPU designs saved transistors and made logarithms a domain of application/language design. Some of the most famous game designs did interesting hacks of pre-computing logarithm tables for a specific set of needs and embedding them in ROM in useful memory versus CPU time trade-offs. Today's CPU designs have plenty of transistors and logarithm support in hardware is just about guaranteed. (That's just CPU designs even; GPU designs can be logarithmic monsters in how many and how fast they can do.) Yesterday's mathematicians envy the speed at which a modern computer can calculate logarithms. In 2023 if you are trying to optimize an algorithm away from logarithms to some other mix of arithmetic you are either writing retro games for a classic chipset like the MOS 6502, stuck by your bosses in a history-challenged backwards language such as COBOL, or massively prematurely optimizing what the CPU can already better optimize for you. I wish that was something any competent programmer would know or at least suspect. It's 2023, it's okay to learn to use logarithms like a mathematician, because you aren't going to need that "optimization" of bit shifts and addition/subtraction/multiplication/division that obscures what your actual high-level algorithmic need and complexity is.
- meling 3y agoWhile reading I was thinking why aren’t stackoverflow “mandating” that solutions have tests, so that this problem isn’t left to everyone else, ref. to the comment at the end of the article: Test all edge cases, especially for code copied from Stack Overflow.
- Rapzid 3y agoThe most impressive suggestion Copilot has given me was a solution to this that used a loop to divide and index further into an array of units.. It never dawned on me to approach it that way and I had never seen that solution(not that I ever looked). Not sure where it got that from but was pretty cool and.... Yeah, it gets simple stuff wrong all the time haha.
- paulddraper 3y agotl;dr When in the 999+ petabyte range, it gives inappropriately rounded results. And the key takeaway is "Stack Overflow snippets can be buggy, even if they have thousands of upvotes." I don't disagree, but is this really the example to prove it.....