6 ms·
The gains seem to not have been high enough to sustain that project. Nowadays CPUs plan, fuse and reorder so much of micro-code that lower-level languages can s
by BenoitP 3y ago
The gains seem to not have been high enough to sustain that project. Nowadays CPUs plan, fuse and reorder so much of micro-code that lower-level languages can sort of be considered virtual as well.
But Java and similar languages extract more freedom-of-operation from the programmer to the runtime: no memory address shenanigans, richer types, and to some extent immutability and sealed chunks of code. All these could be picked up and turned into more performance by the hardware; with some help from the compiler. Sort of like SQL being a 4th-gen language, letting the runtime collect statistics and chose the best course of execution (if you squint at it in the dark with colored glasses)
More recent work about this is to be found on the RISC-V J extension [1], still to be formalized and picked up by the industry. Three features could help dynamic languages:
* Pointer masking: you can fit a lot in the unused higher bits of an address. Some GCs use them to annotate memory (refered-to/visited/unvisited/etc.), but you have to mask them. A hardware assisted mask could help a lot.
* Memory tagging: Helps with security, helps with bounds-checking
* More control over instruction caches
It is sort of stale at the moment, and if you track down the people working on it they've been reassigned to the AI-accelerator craze. But it's going to come back, as Moore's law continues to end and Java's TCO will again be at the top of the bean-counter's stack.
[1] https://github.com/riscv/riscv-j-extension https://github.com/riscv/riscv-j-extension
- lionkor 3y ago> as Moore's law continues to end more like Wirths law proving itself still
- pjmlp 3y agoAs free beer AOT compilers for Java are commonly available, and as shown on Android since version 5, I doubt special opcodes will matter again. Ironically when one dives into computer archeology, old Assembly languages are occasionally referred as bytecodes, the reason being that in CISC designs with microcoded CPUs they were already seen that way by hardware teams.
- BenoitP 3y agoI'm still not decided on AOT vs JIT being the endgame. In theory JIT should be higher performance, because it benefits from statistics taken at actual runtime. Given a smart enough compiler. But as a piece of code matures and gets more stable, the envelope of executions is better known and programmers can encode that at compile-time. That's the tradeoff taken by Rust: ask for more proofs from the programmers, and Rust is continuing to pick up speed. That's also what the Leyden project / condensers [1] is about, if I understand correctly. Pick up proofs and guarantees as early as possible and transform the program. For example by constant-propagating a configuration file taken up during build-time. Something I've pondered over the years: a programmer's job is not to produce code. It is to produce proofs and guarantees (yet another digression/rant: generating code was never a problem. Before LLMs we could copy-paste code from StackOverflow just fine) In the end it's only about marginal improvements though. These could be superseded by changes of paradigm like RAM getting some compute capabilities; or programs being split into a myriad of specialized instructions. For example filters, rules and parsing going inside the network card; SQL projections and filters going into the SSD controller; or matrix-multiplication going into integrated GPU/TPU/etc just like now. [1] https://openjdk.org/projects/leyden/notes/03-toward-condensers https://openjdk.org/projects/leyden/notes/03-toward-condense...
- pjmlp 3y agoThe best solution isn't AOT vs JIT, rather JIT and AOT, having both available as standard part of the tooling. Android has learnt to have both, and thanks to PGO being shared across devices via Play Store, the AOT/JIT outcome reaches the ideal optimum for a specific application. Azul and IBM have similar approaches on their JVMs with a cluster based JIT, and JIT caches as AOT alternative. Also stuff like GPGPU is a mix of AOT and JIT, and is doing quite alright. I am not so confident with LLMs, when they get good enough programmers will be left out of the loop, and will have to contend to similar roles as when doing no-code SaaS configs or some form of architects. A few programmers will remain as the LLMs high priests.
- BenoitP 3y ago
- vextea 3y agoRemember when for a while Azul tried to sell custom CPUs to support features in their JVM (e.g. some garbage collector features that required hardware interrupts and some other extra instructions). Although they dropped it pretty quickly in favor of just working on software https://www.cpushack.com/2016/05/21/azul-systems-vega-3-54-cups-of-coffee/ https://www.cpushack.com/2016/05/21/azul-systems-vega-3-54-c...
- sillywalk 3y agoIBM's Z14 (and later I assume) supported Guarded Storage Facility for 'pauseless Java Garbage collection.'
- toast0 3y ago> Pointer masking: you can fit a lot in the unused higher bits of an address. Some GCs use them to annotate memory (refered-to/visited/unvisited/etc.), but you have to mask them. A hardware assisted mask could help a lot. If you're building hardware masking, it should be viable for low bits too. If you define all your objects to be n-byte aligned, it frees up low bits for things too, and might not be an imposition, things like to be aligned.
- ithkuil 3y agoThe sparc ISA had tagged arithmetic instructions so that you could tag integers using LSBs and ignore them
- pjc50 3y agoOne of the few elements left like this is the ARM Javascript instruction: https://news.ycombinator.com/item?id=24808207 https://news.ycombinator.com/item?id=24808207
- funcDropShadow 3y agoThe Java ecosystems initially started with optimizing Java compilers. That setup could benefit from direct hardware support for Java bytecode. Later, it was discovered that it is more beneficial to remove the optimization from javac in order to provide more context to the JIT compiler. Which enables better optimizations from JIT compilers. By directly running Java bytecode, you would loose so many optimizations done by Hotspot, that it is hard to get on par just by interpreting bytecode in hardware. The story may be different for restricted JVMs that don't have a sophisticated JIT.
- miohtama 3y agoThe current (largest) end-user Java ecosystem is in practice Android and it ahead-of-time compiling ART. Java itself got very good. Though Oracle was blocked to leech money, or have return for their investment, depending on the viewpoint.
- kaba0 3y agoI don’t really get your last point - java’s improvements are due to Oracle, not despite it. They have a terrible name, but they have been excellent stewards of the platform.
- miohtama 3y agoAndroid ART would unlikely to exist if Oracle would have been enforcing licensing requirements as they wished. ART runs on devices for 1B+ users and is more relevant for the world population as Oracle. Although we can speculate likely Android would have switched to something else if Oracle were to win in the court. Ironically Android more realised Java’s original light client vision “Write once, run everywhere” if you consider “everywhere” as all around the world, by every human, with various device architectures.
- kaba0 3y agoI wouldn’t be so quick to dismiss the huge deal of internet services running OpenJDK. Like, AWS itself, Apple’s backends, a huge part of Google’s infrastructure, the whole of Alibaba that is responsible for some crazy amount of transactions, just to mention a few.