8 ms·
Impossible Java
- bmc7505 10y agoI was surprised to learn that in Kotlin, it is possible to disambiguate overloaded functions based only on their return type. I had no idea the JVM even supports such semantics. [1] [1]: http://stackoverflow.com/q/42916801/1772342 http://stackoverflow.com/q/42916801/1772342
- mbel 10y agoAccording you your link JVM does not support this feature. It's implemented by Kotlin itself. I guess it's probably some kind of name-mangling scheme.
- bmc7505 10y agoNote that the first example compiles without any further changes, producing valid JVM bytecode.
- masklinn 10y agoAccording to the link Java/Javac does not support the feature. The JVM does, `bar(foo: List<String>): String` can be compiled to `bar(Ljava/util/List;)Ljava/lang/String;` and `bar(foo: List<Int>): Int` to `bar(Ljava/util/list;)Ljava/lang/Integer;` or somesuch, there is no ambiguity at the bytecode level. In fact, the documentation for Class#getMethod specifically outlines this issue[0]: > Note that there may be more than one matching method in a class because while the Java language forbids a class to declare multiple methods with the same signature but different return types, the Java virtual machine does not. [0] http://docs.oracle.com/javase/8/docs/api/java/lang/Class.html#getMethod-java.lang.String-java.lang.Class...- http://docs.oracle.com/javase/8/docs/api/java/lang/Class.htm...
- mbel 10y agoThe article describes Dalvik which is different from JVM, so it's not really a proof. But you are actually right. It looks that JVM bytecode includes full function signature in function invocation: https://www.ibm.com/developerworks/library/it-haggar_bytecode/ https://www.ibm.com/developerworks/library/it-haggar_bytecod... (if I'm reading examples correctly).
- barrkel 10y agoIt could hardly be otherwise, or the JVM would need to do overload resolution at runtime, which would not be a pleasant experience for anyone.
- mbel 10y agoWell, it could just use function name and leave it to javac/kotlin/other JVM front end to handle overloads and generate unique names. Although that wouldn't work nice with refection and other run-time features I guess.
- masklinn 10y agoIt would also screw up interop. You'd basically have the issue everyone has when trying to call C++.
- masklinn 10y ago> The article describes Dalvik which is different from JVM, so it's not really a proof. My comment describes the actual JVM, hence having linked to official Java/JVM documentation.
- mbel 10y agoSorry, I misread it.
- matharmin 10y agoThis is about Dalvik bytecode format, but the same applies to standard Java bytebode files. Practically any obfuscated Java code will have this, which makes reverse engineering much more difficult without the tools to handle it.
- HighlandSpring 10y agoWhy is it not as simple as mapping the method call instructions to returntype_methodname format? Am I missing something?
- Flowdalic 10y agoWhat if the call site does not use the returned value?
- tokenizerrr 10y agoIt's still present in the bytecode (since the VM has to know which method to call).
- masklinn 10y agoThe bytecode encodes that information, the method identifier includes name, parameter types and return type.
- barahilia 10y agoYes. Because of overloading there may be 2 functions with the same name and the same return type but different arguments. So you need to include them too.
- tokenizerrr 10y agoNo. Java itself handles the differing arguments.
- MichaelGG 10y ago
- deleted 10y ago[deleted]
- _old_dude_ 10y agoEven the Java compiler uses that trick, by example with a bridge method. public static void main(String[] args) { class Fun implements Supplier<String> { public String get() { return null; } } Arrays.stream(Fun.class.getMethods()) .filter(m -> m.getDeclaringClass() == Fun.class) .forEach(System.out::println); }
- kuschku 10y agowhy not use .getDeclaredMethods()?
- _old_dude_ 10y agono real reason :)
- DorothySim 10y agoThe same thing exists in .NET IL where you can overload methods based only on return values (among other interesting things like modopt/modreq [0] etc.). [0]: http://stackoverflow.com/a/5294456 http://stackoverflow.com/a/5294456
- 0x0 10y agoAnother oddity to think about in java source code: In regular Java, all objects extend java.lang.Object, including java.lang.Class. So how do you bootstrap building java.lang.Object from source?
- pjmlp 10y agoYou don't. This is part of the bootstraping process of a programming language. Usually such special types are built manually in the compiler data structures, or make use of special primitives, like native methods on Java's case.
- seanmcdirmid 10y agojava.lang.Object is just part of the VM. Same thing with native methods.
- 0x0 10y agoWhile many of the methods in java.lang.Object have a native implementation, there are still other methods that do sport a pure java implementation: http://hg.openjdk.java.net/jdk7/jdk7/jdk/file/tip/src/share/classes/java/lang/Object.java http://hg.openjdk.java.net/jdk7/jdk7/jdk/file/tip/src/share/...
- seanmcdirmid 10y agoSure, but...if you look at the JVMS, there are a lot of special cases concerning java.lang.Object to help with bootstrapping.
- Gaelan 10y ago> Any compiler of sound mind and memory will issue an error GHC would beg to differ.
- mattnewton 10y agoThat was my first thought, that this is being presented as a fundamental tenant of compilers and really it's a design choice that a whole family of languages, some of them that people even use, didnt make.
- masklinn 10y agoAlso rustc, though it doesn't quite have overloading in the Java sense: https://is.gd/s0oF79 https://is.gd/s0oF79
- cjensen 10y agoAda also allows return type overloading. A complaint against Ada is that it is hard for the compiler to figure out the right overloading in a complex statement with many nested function calls. My compilers prof took the time to show how to do it, and why it doesn't take much time. Too bad C's automatic conversions prevent this from being used in C++.
- tenkeyless 10y agoThis is a well known problem in the field of decompilers and disassemblers. It figures that the pseudo-code it outputs for generic compilers is pretty good, but when encountered with a man-made assembly or byte-code, they go places.
- mnarayan01 10y agoIt seems like this is simply a disassembler error (albeit an understandable one). Am I missing something? Edit: Based on the responses below, I guess the point is that the disassembler can't generate Java code that will "naively" (wrong word, but I can't think of a better one) generate the same output. Notable (I assume) in that name munging would be problematic outside the current compilation unit.
- shawnz 10y agoWhat did the disassembler do wrong? It just happens that there is no valid Java code which could produce that (valid) bytecode. What should it have outputted instead?
- tokenizerrr 10y agoOutput the method names as a_void and a_string instead of just a?
- shawnz 10y agoThen it wouldn't be true to the real structure of the program (although it would be true to the real behaviour of the program). Why is that less wrong?
- tokenizerrr 10y ago> Why is that less wrong? It compiles.
- drdrey 10y agoBut may not run.
- bluejekyll 10y agoIt's obviously not impossible to create legit Java code during disassembly. This is like an arms race; go ahead and dissemble my code, but I've made sure that you now need a better disassembler to produce valid code. Eventually that disassembled will come along, and then more tricks will be used by the obfuscators...
- nebulous1 10y agotl;dr Java needs to be able to decide which overloaded method implementation to use at compile time, meaning you can't differentiate between method implementations based on return type alone. However, when compiled to bytecode, each method and method call includes the return type as part of the method identification, so Java bytecode is actually capable of differentiating between implementations with the same name based on return type alone. This fact is used by bytecode obfuscators, and can lead to bugs in decompiled code if the decompiler doesn't account for it.
- masklinn 10y ago> Java needs to be able to decide which overloaded method implementation to use at compile time, meaning you can't differentiate between method implementations based on return type alone. That's not correct. Rust has no issue statically dispatching based on the return type. Java-the-language does not allow it so it does not have to deal with calls which don't use the return value e.g. int getSomething() String getSomething() If the return value is not used, you have to explicitly disambiguate this call somehow, and Java provides no way to do so.
- bradleyjg 10y agoEven if you have a language that forces you to assign the return value, what does the compiler do in a situation like: Integer getSomething() String getSomething() when the calling site is: Object something = obj.getSomething();
- anonymousDan 10y agoI believe the technical term for the correspondence between what it is possible to compile and what is valid bytecode/machine code is fully abstract compilation. It's an interesting concept with many interesting implications (e.g. for security). In the past at least there were various examples of Java programs that were illegal in the language but nonetheless could be created directly as bytecode and would be loaded by the JVM. This obviously becomes a security problem if your program loads bytecode dynamically and makes assumptions about its capabilities at the language level as opposed to the bytecode level.
- drdrey 10y agoIf you load untrusted code dynamically, it strikes me as wrong to assume anything about its capabilities. Even more so "at the language level". Untrusted code can do anything unless you sandbox it.
- anonymousDan 10y agoYou're right but I guess it's perhaps easy to overlook, e.g. that if you decompile/disassemble a valid bytecode program it might give you a program that is not valid at the source level. Some interesting examples and discussion here: http://lambda-the-ultimate.org/node/5364 http://lambda-the-ultimate.org/node/5364
- teraflop 10y agoFun fact: in Java versions 5 and 6, it was actually possible to write valid Java source code with overloaded return types! The trick is that generic types in Java are subject to type erasure at runtime. Due to an oversight, it was possible to declare a class with methods like: String foo(List<X> l); double foo(List<Y> l); which would be erased to: String foo(List l); double foo(List l); At runtime, even though the type information for your List was no longer available, the compiler would be able to locate the correct method using the method signature stored in the caller's bytecode, giving the appearance of return-type dispatch. Technically this violates the Java language spec, and javac 7 was updated to be stricter and prevent this sort of code: http://bugs.java.com/bugdatabase/view_bug.do?bug_id=6182950 http://bugs.java.com/bugdatabase/view_bug.do?bug_id=6182950 I guess it's ultimately not that mysterious, but I first encountered it "in the wild", and ended up scratching my head for a while before I figured out why our code suddenly stopped compiling when we upgraded the JDK.
- jxi 10y agoInteresting. Then, why isn't the bytecode always stored with the caller, that way you can have return-type dispatch?
- deleted 10y ago[deleted]
- aardvark179 10y agoThe bytecode for method invocation does store the type of the method (which includes the return type), and it is quite valid for a class file to contain multiple methods that are only distinguishable by their return type. You can't normally do this visibly in Java, but I think it's used for bridge methods and a few other things by the compiler.
- josefx 10y agoThe Java bytecode stores it, however the Java language normally does not expose it. Having the compiler select the correct overload based on return type would most likely add a lot of complexity to the language itself ( have you seen the current overload resolution rules? ) without much benefit. I think the compiler actually has to work around that when you narrow the return type of an overriden function: A method Object Base::get() overriden with Integer Child::get() will result in an additional compiler generated Object Child::get() in the bytecode.
- mooman219 10y agoNot only is this possible, as stated by the article, this level of overloading is incredibly useful for in a number of areas and has been used for a long time. One use case is API backwards compatibility. If your API wants to change the return type of a function, say from int to double, but also wants to maintain binary backwards compatibility, you can do that. See OverMapped [ https://github.com/Wolvereness/OverMapped https://github.com/Wolvereness/OverMapped ]. Obfuscation is another area and ProGuard employs this to make decompliling more difficult iirc.
- ipsum2 10y agonever mind.
- usmannk 10y agoI could see his comment just fine. What's the evidence of an HN shadowban?
- sctb 10y agoWe detached this subthread from https://news.ycombinator.com/item?id=13978371 https://news.ycombinator.com/item?id=13978371 and marked it off-topic.
- Tideflat 10y agoI, for one, can see what he posts and they aren't hidden in any way.
- reitanqild 10y agoI call clickbait: Author admits this is not valid Java. It is not even compilable. If I read correctly it is just artifacts from a partially sucessful decompile. Interesting and this discussion is interesting but this is not and have never been valid Java.
- cremp 10y agoIt should be noted that the JVM allows for things that javac does not. This is one of those things, and describes it accurately, albeit in the android java environment, however it applies to the standard JVM too. One of the more fun things I've done is use the ASM and BCEL libraries to make (and unmake) these kinds of manipulations (manipulations that javac won't let you do.)