10 ms·
Java's URL.equals() Performs DNS Resolution
- ceejayoz 7y ago> Note: The defined behavior for equals is known to be inconsistent with virtual hosting in HTTP. That seems... less than useful.
- strictnein 7y ago> "Two hosts are considered equivalent if both host names can be resolved into the same IP addresses" Uhmm... yikes. Why are they resolving anything? A URL is a string, this should just be doing a string comparison. All of the parts of the URL object are strings or ints, so at worst they should just be comparing all of those individually, not resolving domains and comparing IP addresses. That makes no sense at all.
- boofgod 7y agoIt's possible to have URLs that are not equal as strings but resolve to the same thing so this actually makes perfect sense. If you really want to test if URLs are equal as strings, I'm sure you could still do it that way.
- Someone1234 7y ago> It's possible to have URLs that are not equal as strings but resolve to the same thing so this actually makes perfect sense. No, that doesn't make "perfect sense." Two different URLs are different, they cannot be equal if they're different, and a library called "URL" shouldn't use voodoo-magic to say that two different URLs are equal when they're not. If it was e.g. Net.DNS.Equal(Uri, Uri), perhaps. But even then it is ambiguous as DNS has multiple record types, not just "A" records. So is it pulling all records (A, AAAA, MX, etc) and comparing them all? Or just arbitrary comparing one? But as it stands, it is a URL library that is ignoring the URL part of the URL and using resolution to decide equality instead. It is nonsensical. And even if they did resolve identically they may not be treated identically by routing or endpoints.
- deleted 7y ago[deleted]
- jerf 7y agoNot only is it possible to have two hostnames that resolve to the same IP address, hostname -> IP address is not unique either: $ dig amazon.com ... ;; ANSWER SECTION: amazon.com. 60 IN A 205.251.242.103 amazon.com. 60 IN A 176.32.103.205 amazon.com. 60 IN A 176.32.98.166 Something resolving "amazon.com" is welcome to resolve it to any of those three addresses. Code that tries to "resolve" these IPs is at the very least coming perilously close to deciding that amazon.com != amazon.com, nondeterministically, if the underlying resolution call changes its behavior. The further obvious change of trying to compare the whole record just gets worse; what is the answer if domain 1 resolves to IPs (V, X, Y) and domain 2 resolves to (X, Y, Z)? Oh, and let's not forget, DNS can be different depending on your geographical location, so in the US two domains may happen to resolve to the same IP but in Europe they may be different. What's the use of a URL equality operator that changes behavior based on where the user is? Whatever the use of such a thing may be (I mean, yeah, I get the internet contrarian impulse, yeah, I can construct some bizarre situation in which it is useful), it is certainly less useful than a simpler operator. Basically, the entire idea is just fundamentally flawed and shouldn't be used. I make no claim to have exhaustively enumerated all the ways in which this is a bad idea, merely added to the pile, and demonstrated sufficient evidence for my claim that it shouldn't be used. Edit: They know, see yzmtf2008: https://news.ycombinator.com/item?id=21766138 https://news.ycombinator.com/item?id=21766138 and consider my post here a lightly educational post on further reasons why it's a bad idea. If you only take one thing away, remember, a function that takes a hostname and yields an IP doesn't have very many useful properties beyond just that it yields an IP or failure. You can count on very little else about that IP. It should be treated as opaque and not compared, stored (other than logging), etc.
- jefftk 7y agoUnlike virtual hosting, which wasn't possible when java.net.URL was written, this kind of DNS behavior had been around for years: https://tools.ietf.org/html/rfc1794 https://tools.ietf.org/html/rfc1794
- lmkg 7y agoI mean, I can see host resolution equivalence being a useful piece of functionality. I could even see it being built-in to a std URL library. Just, y'know... not as the default equality operator. And yes, I know that overriding .equals() does not change the behavior of == in Java. But I would still consider .equals() to be the "default" equality operator for objects.
- BenjiWiebe 7y agoThanks to virtual hosting (which is extremely common) host resolution is actively harmful to URL comparison.
- jefftk 7y ago> Thanks to virtual hosting (which is extremely common) It's extremely common now, but it wasn't possible yet when they made this design decision.
- pdw 7y agoI'm pretty sure the specification of this class predates name-based virtual hosting. After all, that only appeared in 1997, or perhaps late 1996, and it took a couple of years to be standardized. The URL class existed in Java in 1995 :) (Edit: And really, virtual hosts are an extremely weird HTTP feature. What other internet protocol cares about what domain name you used to establish a connection?)
- merb 7y ago> (Edit: And really, virtual hosts are an extremely weird HTTP feature. What other internet protocol cares about what domain name you used to establish a connection?) kerberos? (btw. this is older than 1995) or smtp (little bit different than the http version, tough)
- tootie 7y agoSounds like a developer with too much time on their hands decided to add some humongous side effects.
- kevin_thibedeau 7y agoProbably a PHB who wanted to "put the . in .com" with network enhanced string processing.
- whalesalad 7y agoIt makes a lot of sense if you pull your head out of the code and consider a URL is a universal resource locator. I don’t agree with what this is doing... but you can’t say it makes no sense at all.
- strictnein 7y agoWith this definition two different resources could be said to have the equivalent universal resource locators. foo.bar/baz would be said to be equivalent to bar.wack/baz if they shared the same IP address. If there is some sense to be made from that, it eludes me.
- lkschubert8 7y agoI think the argument is the domain name is an abstraction over ip address and ip address is the definitive "resource".
- kstrauser 7y agoIt definitely doesn't make sense, because two unrelated hosts can share the same IP. If you tried to fetch http://a.b.c.d/some/path http://a.b.c.d/some/path, then the server will fail because it doesn't know which host you mean. In this case, Java has thrown away the critical distinction between the two.
- sjwright 7y agoAnd it especially doesn’t make sense because resolving to the same IP on the server right now doesn’t guarantee it will resolve the same on the client, or even on the same server in 15 minutes time.
- desc 7y ago* A URL is a name for a pointer. * A URL is a parameter to a function returning a resource. * A URL is a string, conforming to a certain format. Neither of these specifies that the resource it points to is always the same resource, nor that it's always possible to resolve it. That's by design; hostnames change. The Web is not permanent and that's why we have eg. HTTP 404. Network topologies also change. However, the default equals() comparison between two objects is supposed to compare those two objects, not the current topology of the Internet. There is no way this behaviour ever made sense, nor any way it ever could make sense. It's moronic, through and through, for any language or library, to implement default equality in this way. If you want to implement a comparer which does stuff like this and accepts a hostname resolver as a dependency, great. But there is simply no excuse for this kind of stupidity in a default dependency-less implementation.
- tyre 7y agoThat's an interesting way to think about URLs. If both `url`.com` and `url2.com` resolve to the same IP then they are functionally the same. It's like comparing two strings whose variable names are different but that resolve to the same space in memory. It also means that if one URL changes, then they could be equal sometimes and not others.
- kohtatsu 7y agoThey are not functionally the same. Full stop. It sends a different Host HTTP header; servers send different responses based on that header ~100% of the time.
- ses1984 7y agoNo, they aren't functionally the same for all protocols, because of things like SNI.
- duskwuff 7y agoOr, even before that: because of things like HTTP 1.1 virtual hosting.
- sli 7y ago> Note: The defined behavior for equals is known to be inconsistent with virtual hosting in HTTP. They even seem to know it makes no sense.
- Gaelan 7y agoPresumably, at the point they realized that, something already depended on this behavior.
- kuschku 7y agoAt the point someone invented virtual hosting, sun already depended on this behavior ;)
- Twirrim 7y agoThis feature was introduced in 1995. Before virtual hosting was a thing. (HTTP/1.1 wasn't formalised until 1999)
- eliaspro 7y ago...and what if my resolver uses round-robin and returns a different IP on each request?
- duskwuff 7y agoOr if the computer running the JVM doesn't have a reliable network connection? Does the behavior of URL.equals() change if your network connection is down? ... and what happens when you try to compare a URL which contains a nonexistent hostname, like http://asdfghjkl.example.com/ http://asdfghjkl.example.com/ ? Does that compare as equal to all other URLs with unresolvable hostnames?
- wolfgang42 7y ago> if either host name can't be resolved, the host names must be equal without regard to case; or both host names equal to null.
- cotillion 7y agoI wonder if anything is using URL.equals to check for 'safe' urls. Some creative use of DNS TTL could probably make things interesting.
- adrusi 7y agoBuy a domain name, point it at Enterprise Business Inc.'s server, submit a link using your domain to a form the operate which does a if (internalURLs.contains(submittedURL)) { check. Then change your DNS records to point to some other server, once your domain is in their database and they assume it to have already been validated as internal.
- kstrauser 7y agoExploiters, start your engines. This seems absolutely ripe for abuse, as now there's a nice string you can search for in GitHub to see where that's used as a security feature. I'm imagining things like "if URL("http://example.com/some/path" http://example.com/some/path").equals(URL(checkedUrl)) { return AllowEditRights }", and checkedUrl = "http://wiki.example.com/some/path" http://wiki.example.com/some/path" or similar.
- toyg 7y ago"now"? This has been the case for more than 20 years...
- ivan_gammel 7y agoThere's one thing in Javadoc that says it all: `@since JDK 1.0` Java is as good as Windows in the sense, that its standard library and set of APIs is very stable and supports a lot of legacy software. I doubt there's a real need for URL class in the new code by now, given that URI class was introduced in JDK 1.4 almost 18 years ago. There's plenty of dependencies though, so URL will probably stay in the core library forever, but URI represents a superset for URLs, has reasonable implementation of equals/hashCode and is sufficient for majority of the uses.
- toyg 7y agoIndeed, and tbh I'm somewhat surprised this is front-page in 2019. I guess younglings are (re-)learning java and unhearting its umpteen WTFs...
- kelnos 7y agoEh, I dunno, I've been an on-and-off Java person since ~2000 (with consistent JVM development over the last 8 years or so), and I only learned about this URL gotcha within the last couple years. I think https://imgs.xkcd.com/comics/ten_thousand_2x.png https://imgs.xkcd.com/comics/ten_thousand_2x.png applies here, just in a more restricted programmer-y sense..
- vbezhenar 7y ago> I doubt there's a real need for URL class in the new code by now URL class is used to establish actual connection to a resource.
- dionian 7y agoI'm a seasoned java dev (since before 1.4) and i didnt realize i was supposed to use URI, i did know of it, but I'm always confused what to use and wind up using URL. But, just like Vector remains in use in Swing, URL remains in use in the core java library (looking at you, classloader.getresource), so it's easy for me to have made that mistake. then again i dont use it often. mostly come in contact with it when using ClassLoader API.,
- boring_twenties 7y agoWhat about http://foo/ http://foo/ and http://foo:80/ http://foo:80/? Are those different URLs?
- beart 7y agoYes, they just resolve to the same thing
- iso-8859-1 7y agoPort numbers do not resolve. And your 'thing' is ambiguous.
- spc476 7y agoRFC-3986, section 6.2.3: For example, because the "http" scheme makes use of an authority component, has a default port of "80", and defines an empty path to be equivalent to "/", the following four URIs are equivalent: http://example.com http://example.com/ http://example.com:/ http://example.com:80/
- patrickthebold 7y ago> January 2005
- spc476 7y agoAnd? It has not been obsoleted, and it's marked as STD0066.
- kelnos 7y agoOne thing I find interesting is that it doesn't canonicalize the path portion at all. E.g.: scala> new URL("http://localhost/foo/bar/baz").equals(new URL("http://localhost/foo/../foo/bar/baz")) res0: Boolean = false It's just funny to me that they went to the effort to do a full network request to see if the hostnames resolve to the same IP, but didn't bother to normalize paths.
- jacinabox 7y agoStill this sort of expansionary philosophy in library design does give some cause for optimism. In general library authors want to make their specs efficiently implementable, so when an method has this sort of semantics in an official spec it makes one suspect that a breakthrough in AI/algorithms research is right around the corner.
- smarks 7y agoThis might have something to do with the way that name resolution worked inside of Sun at the time. If there was a host, say doppio.eng.sun.com it would be referred to simply as "doppio" from within the engineering ("eng.sun.com") domain, or possibly as "doppio.eng" from other domains inside of Sun. It was fairly rare to use FQDNs inside of Sun to refer to other hosts inside of Sun. Thus, the following URLs all referred to the same resource: http://doppio/foo.html http://doppio.eng/foo.html http://doppio.eng.sun.com/foo.html It's a plausible point of view that URL.equals() should report true for any two of the above URLs. (That doesn't mean that I think it was a good idea, though.)
- cryptonector 7y agoBut then again, an HTTP server is allowed to provide different semantics for an URLs depending on the authority portion of those URLs. So, no, this is not remotely Ok at all.
- kelnos 7y agoAgreed, but consider that the internet, and state of HTTP hosting, was a very different place in 1995. Was the Host header even a thing back then? (And if it was, was it widely deployed?) This is clearly wrong in hindsight, but I could see why it was naively designed that way in the first place. Regardless, given all Java's warts, this doesn't even make my top 10.
- cryptonector 7y agoSure, but since no one in their right mind wants this behavior now, it's best to rip it out and fix the spec to not require it.
- zmzrr 7y agoThis is like when `strings' was found to be vulnerable to code execution because it was parsing ELF files.
- divyekapoor 7y agoAnd even with the DNS resolution, they'll get it wrong.
- lmkg 7y agoFrom the linked documentation: > Note: The defined behavior for equals is known to be inconsistent with virtual hosting in HTTP. So, yes.
- yzmtf2008 7y agoTL;DR: use URI instead. URL is behaving this way due to backwards compatibility. --- To quote a JDK bug ticket: https://bugs.java.com/bugdatabase/view_bug.do?bug_id=4434494 https://bugs.java.com/bugdatabase/view_bug.do?bug_id=4434494 > We are well aware of the problem with URL.equals and URL.hashCode. The cause of the problem is due to the existing spec and implementation, where we will try to compare the two URLs by resolving the host IP addresses, instead of just doing a string comparison. Because hashCode has to maintain certain relationships with equals, namely if two objects are equal, they should have the same hashCode, the implementation of hashCode also tries to resolve the host in the URL into an IP address. As a result, we are facing problems with http virtual hosting, as described in the Description part, and performance hit due to DNS name resolutions. > Unfortunately, changing the behavior now would break backward compatibility in a serious way, plus Java Security mechanism depends on it in some parts of the implementation. We can't change it now. > However, to address URI parsing in general, we introduced a new class called URI in Merlin (jdk1.4). People are encouraged to use URI for parsing and URI comparison, and leave URL class for accessing the URI itself, getting at the protocol handler, interacting with the protocol etc. So, at present, we don't plan on changing the URL.equals/hashCode behavior and we will leave the bug open until Tiger, when we re-investigate our options.
- tuwtuwtuwtuw 7y agoOut of curiosity do you get a compiler warning when you use URL.equals?
- ivan_gammel 7y agoNo and probably it will never happen in javac. Some IDEs may try to do static code analysis to identify possible uses of URL.equals (e.g. Map<? extends URL, ?>), but I doubt anyone would invest in this feature.
- matsemann 7y agoLooks like IntelliJ has it: - equals() or hashCode() called on java.net.URL object - Map or Set may contain java.net.URL objects https://www.jetbrains.com/help/idea/list-of-java-inspections.html https://www.jetbrains.com/help/idea/list-of-java-inspections...
- ajkjk 7y agoJava has a bunch of relics like this: holdovers from the more zealous OOP days of the past, when updating objects was supposed to update the database, URLs were the concept of domain names instead of the data of a physical URL, etc. I think most people have away from this mindset because it leads to code being hard to reason about.
- Someone1234 7y agoJava could really benefit from more pruning/obsoleting. There's far too many "gotcha" pieces that nobody uses because better alternatives have come along, or ideas have been lost to time. Instead of leaving them in as traps for noobies, they should mark them bad or outright remove them.
- ivan_gammel 7y agoThere's plenty of software worth somewhere between 10^11-10^12$ which is still working because all those legacy APIs still exist. Java is not <put the name of any other platform in permanent beta stage>. It prospered for 20 years and it will stay alive for another half of century, because it is conservative enough.
- yashap 7y ago“Cleanup” vs “backwards compatibility” is a really tough problem for language designers/maintainers. Breaking backwards compatibility leads to a fractured community (think Python 2/3), as well as resentment from all the users who want stability, and don’t want to waste time on painful upgrades. But over time, not breaking backwards compatibility leads to all sorts of weird warts, adding lots of overhead to learning the language, and using it safely. Ultimately I think it’s a somewhat natural part of language evolution that they accumulate cruft over time, because the costs of breaking backwards compatibility are too great. And after a few decades or so, new languages come out that are (currently) much cleaner, and they take over until the cycle repeats. i.e. “cleaner Java” will not be a new version of Java, but a new language solving similar problems (Go?).
- 7y ago
- nostromo 7y agoAnyone who has used Java knows that the URL class is a total mess. Just avoid it entirely.
- badrequest 7y agoLove to use a language where you just kind of have to glean from convention which standard libraries are worthless.
- scarejunba 7y agoIsn't that true for practically any old language. Certainly true for C, C++, Java. Most of them carry the weight of early decisions.
- theandrewbailey 7y agoIt's unavoidable when a language/library ecosystem has been around for over 20 years.
- Polyisoprene 7y agoI’m sure there are some languages without legacy in their stdlib, but I still haven’t seen one.
- josefx 7y agoFrom the URL documentation: > The recommended way to manage the encoding and decoding of URLs is to use URI, and to convert between these two classes using toURI() and URI.toURL(). and at that point you might as well just use URI. The resource resolution functionality is basically the only thing URL offers over URI.
- Someone1234 7y agoWhy is URL.equals(url) doing DNS Resolution comparison under the hood? That seems wildly outside the scope of what the URL library should be doing. It might be a useful feature, but there's no rational reason it should be here. Some quick googling suggests people are using java.net.URI instead to bypass this poor design.
- peeters 7y agoWell the "why are they doing DNS resolution" comes down to "because the contract requires it": > Two URL objects are equal if they have the same protocol, reference equivalent hosts, have the same port number on the host, and the same file and fragment of the file. > Two hosts are considered equivalent if both host names can be resolved into the same IP addresses; else if either host name can't be resolved, the host names must be equal without regard to case; or both host names equal to null. > Since hosts comparison requires name resolution, this operation is a blocking operation. So it's not an implementation bug, it's a requirements bug. Now why on Earth they thought this was a reasonable requirement for an equals() method, that's a fair question.
- shellac 7y agoThis is very old news now, but still worth repeating. Use URI
- TimTheTinker 7y agoI can see why a method like this could be helpful. The URL specification RFC has an incredibly flexible specification for what constitutes a URL. The hostname/address section can be particularly hairy - to the point that it likely becomes impractical/infeasible to compare hostnames for equality. Comparison via DNS resolution is likely a simple, desirable solution to a real problem. That being said, URL.equals() is a terribly opaque and non-obvious method signature for performing DNS-based comparisons of hostnames. It lacks any indication that calling it involves network IO.
- tuwtuwtuwtuw 7y agoThis is only a desirable solution if you ignore basic internet stuff such as round robin resolution, non-infinite-TTL, SNI and Host headers. Even the authors realized it was a bad idea...
- keymone 7y agoI can’t see when can it be helpful. Resolving an ip shows you only a subset of dns information, basing equality on this behavior makes zero sense. It should have been deprecated and removed long time ago. Backward compatibility like this is cancer.
- KoenDG 7y agoIntellij's code analysis(and other tools as well) warn about this. Particularly they recommend not sticking URL objects in Set or Map structures, due to the inherent equality check.
- oldgeek 7y agoOk kids. URL was deprecated in favor of URI before some of you were born. Easy to be a harsh judge now but back when java was first developed the concept of a "design pattern" did not even exist. In fact the development design patterns was initially driven primarily by people figuring out good ways of doing things in java.
- dukoid 7y agoIs URL really deprecated -- I thought URL is still used to actually open the connection?
- layer8 7y agoIt’s not deprecated. As you say, it is needed to actually open a connection.
- oldgeek 7y agoI don't mean @deprecated. Just that the recommendation has been to use URI instead of URL wherever possible since URI was introduced. URL is not needed any longer with the new HttpClient stuff in java 11.
- JoeAltmaier 7y agoRevisionist history? The seminal work in Design Patterns (coined the term) came out in 1995 Design Patterns (Elements of Reusable Object-Oriented Software), Erich Gamma, Richard Helm, Ralph Johnson, John Vlissides, 1995] Same year Java was invented. Hard to imagine Java inspired it.
- ivan_gammel 7y ago1994 and the book contained examples of code on C++
- oldgeek 7y agoOk, java was influential in a ton of subsequent design patterns. For a long time java was pretty much the test-bed for these ideas. And the other half of the point remains: modern design patterns were not part of the standard toolbox at the time. To be clear, I am not defending the design of the URL class. Just that you need to judge it in context.
- deleted 7y ago[deleted]
- layoutIfNeeded 7y agoThis is similar to how NSURL on macOS/iOS accesses the filesystem in its constructor: https://developer.apple.com/documentation/foundation/nsurl/1410301-init https://developer.apple.com/documentation/foundation/nsurl/1... “ This method assumes that path is a directory if it ends with a slash. If path does not end with a slash, the method examines the file system to determine if path is a file or a directory.” People often overlook this and then wonder why their app stutters randomly when the UI thread gets blocked on this NSURL ctor :^)
- anfilt 7y agoWow, that seems like a bad idea. who thought that was a good idea let alone expected behavior.
- thfuran 7y agoIf you think that's crazy, check this out: at java.net.Inet6AddressImpl.lookupAllHostAddr(Native Method) at java.net.InetAddress$2.lookupAllHostAddr(InetAddress.java:928) at java.net.InetAddress.getAddressesFromNameService(InetAddress.java:1323) at java.net.InetAddress.getLocalHost(InetAddress.java:1500) - locked <0x00000000800a1578> (a java.lang.Object) at sun.font.FcFontConfiguration.getFcInfoFile(FcFontConfiguration.java:352) at sun.font.FcFontConfiguration.readFcInfo(FcFontConfiguration.java:425) at sun.font.FcFontConfiguration.init(FcFontConfiguration.java:94) - locked <0x00000000d5af3c58> (a sun.font.FcFontConfiguration) at sun.font.FcFontConfiguration.<init>(FcFontConfiguration.java:76) at sun.awt.X11FontManager.createFontConfiguration(X11FontManager.java:768) at sun.font.SunFontManager$2.run(SunFontManager.java:431) at java.security.AccessController.doPrivileged(Native Method) at sun.font.SunFontManager.<init>(SunFontManager.java:376) at sun.awt.FcFontManager.<init>(FcFontManager.java:35) at sun.awt.X11FontManager.<init>(X11FontManager.java:57) at sun.reflect.NativeConstructorAccessorImpl.newInstance0(Native Method) at sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62) at sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45) at java.lang.reflect.Constructor.newInstance(Constructor.java:423) at java.lang.Class.newInstance(Class.java:442) at sun.font.FontManagerFactory$1.run(FontManagerFactory.java:83) at java.security.AccessController.doPrivileged(Native Method) at sun.font.FontManagerFactory.getInstance(FontManagerFactory.java:74) - locked <0x00000000d5abbb00> (a java.lang.Class for sun.font.FontManagerFactory) at sun.font.SunFontManager.getInstance(SunFontManager.java:250) at sun.font.FontDesignMetrics.getMetrics(FontDesignMetrics.java:264) at sun.swing.SwingUtilities2.getFontMetrics(SwingUtilities2.java:1113) at javax.swing.JComponent.getFontMetrics(JComponent.java:1626) at javax.swing.plaf.basic.BasicLabelUI.getPreferredSize(BasicLabelUI.java:227) at javax.swing.JComponent.getPreferredSize(JComponent.java:1662)
- howaboutyes 7y agoPanel: What size pants do you require? Button: Note sure. Hold on, let me call my mother about that.
- khc 7y agoSubmitter here. Someone asked me how often this is actually a problem given there's already URI class. Unfortunately github doesn't do global exact match in searching, but randomly clicking around I found this in a minute: https://github.com/biomics/icef/blob/d69f9be9b1f773598b47de5736f16c0daaffccea/coreASIM/org.coreasim.eclipse/src/org/coreasim/eclipse/util/IconManager.java#L23 https://github.com/biomics/icef/blob/d69f9be9b1f773598b47de5... Query I used: https://github.com/search?q=%22Map%3CURL%2C%3E%22+language%3AJava+language%3AJava&type=Code https://github.com/search?q=%22Map%3CURL%2C%3E%22+language%3...
- khc 7y agoI've filed a few issues so far: https://github.com/apache/dubbo/issues/5462 https://github.com/apache/dubbo/issues/5462 https://github.com/DoctorD1501/JAVMovieScraper/issues/310 https://github.com/DoctorD1501/JAVMovieScraper/issues/310 https://github.com/MovingBlocks/Terasology https://github.com/MovingBlocks/Terasology focusing on repos that actually seem to be used. Someone should write a tool to scan this
- AndrewStephens 7y agoThree of the truisms in my varied programming career are: * DNS will fail even if it is implemented correctly on your clients' network * It probably isn't implemented correctly on your clients' network * Your software is probably doing more DNS queries than you think This seems like a particularly unfortunate example but things like this are not uncommon. Doing any kind of RPC, even just to a server on your local machine? Half the time your library is doing DNS queries under the hood for no good reason. Or performing reverse DNS queries just to display a hostname in a log file. Its easy to accidentally trigger a lookup. If it happens in a loop, a query that should take .2 seconds now takes 30 minutes.
- JaDogg 7y agoThere's even a warning about this in effictive Java. Glad I read that.
- burtonator 7y agoStory time... My company Datastreamer has been around for a decade. We provide crawl data (usually a massive amount of crawl data, north of 300GB per day) to our customers. ... so we have a LOT of real-world experience pushing data to customers in production over long time periods. Here's what we've learned. Networking libraries around HTTP are and have been fundamentally broken for a long time and they're broken in pathological ways that you don't realize until years later in production. DNS caching is a good one. A lot of systems do infinite DNS caching. Java, until at least Java 8, does infinite DNS caching. Some do infinite HTTP timeout. Timeouts are awesome. You should use a timeout. Without a timeout if the network breaks your code just locks up. Some libs provide no API to change TCP buffer sizes (which you have to do at the kernel level). So about 5 years ago we took a harsh stance. NO CUSTOM CLIENTS. We have a streaming firehose client that we implemented from the ground up to do everything properly. The API is literally that we just stream JSON files to disk. It's a docker container now so not too hard to deploy. Your job is to just to listen to the disk and wait for new files to be written. We do a move from a tmpdir to the final dir so the entire file is written and you don't have to worry about partial reads. About 80% of our customers love it. The other 20% of customers seem to initially hate it and we have to explain to their CTO or senior architect that, no, you DO NOT want to implement this from the ground up. What happens is that it works immediately, but then 18 months in it will break pathologically and everyone running it has moved on or it's in some datacenter that no one has access too. This causes us to break our SLAs and means we have upset customers. This decision by far was one of the best decisions I've ever made and has really helped our growth and stability over time. It's really really really nice to keep customers for 5-10 years. They're happy and you get steady checks and predictable growth.
- oneepic 7y agoA long time ago I remember seeing those timeout parameters when I took Java class, and saying, "ehhh it's probably fine to leave it as null". Also the docs at the time seemed to agree that an infinite timeout was totally fine.
- macintux 7y agoSomeone in the Erlang community, might have been Joe himself, said (paraphrased) about timeouts: When I see an infinite timeout, I suggest the developers change it to 30 years. They inevitably respond "That's crazy!" and prove my point.
- soulclap 7y agoJava: a fractal of bad design.
- gaul 7y agoURL's behavior is unfortunate and users should prefer URI introduced in Java 1.4 in 2002. error-prone warns about dangerous uses of URL: https://errorprone.info/bugpattern/URLEqualsHashCode https://errorprone.info/bugpattern/URLEqualsHashCode
- mrkeen 7y agoThis has bitten me. I inserted two URLs into a set. Then the set contained one item, so I tried to debug the set.
- peeters 7y agoBeen there. I remember fixing a performance bug ages ago where we had URL (or maybe some other address class) used as a key in a HashMap (which seemed like a perfectly valid thing to do with a value class). We were doing literally thousands of DNS lookups for what seemed like the most trivial algorithms. equals() and hashCode() are probably one of the weakest points of Java. While it seems like an obvious candidate for a contract, the issue has always been that one person ends up defining equality for everyone, when often different usecases will warrant different definitions of equality. Are objects equal if they have the same identity? If they have the same data? If they resolve to the same thing? It's easy enough to leave them unimplemented but the issue then is a lack of standard library support for providing custom hash and equality functions for Maps and Sets.
- kjeetgill 7y agoI agree! I get that it's easy to critique in hindsight but it's clear that they eventually solved this sort of thing well via the Comparable/Comparator relationship. It's less an issue so I get why it's been backlogged, but it's one that I'd love to see worked out.
- ben509 7y agoWhat's frustrating in hindsight is they knew the URL class was obviously terrible, but never made a serious effort to stop people using it. I'm fine with not breaking old jars, but they could have made newer javac's by default die with a notice "use -legacy to build this". Same thing with Date, Vector, all the broken thread semantics, you name it. These classes should have been relegated to a handful of ancient jars, they shouldn't keep popping up in modern Java because newbies have to learn a whole host of classes they aren't supposed to use.
- kjeetgill 7y agoI sorta agree when in comes to Vector and some of the Thread stuff, but URL is only broken in gethashcode() and equals(). The rest of the class is perfectly fine. It's a great simple curl that takes you pretty far before you reach for a real http client lib.
- jrochkind1 7y agoWell THAT's clearly a legacy inherited mistake. The semantics don't even seem desirable in 2019, let alone the performance characteristics.
- skissane 7y agojava.net.URL needs @Deprecated(forRemoval=true)