6 ms·
> Due to limitations in the underlying programming model, TornadoVM doesn’t support objects (except for trivial cases), recursion, dynamic memory allocation or
by Raphael_Amiard 7y ago
> Due to limitations in the underlying programming model, TornadoVM doesn’t support objects (except for trivial cases), recursion, dynamic memory allocation or exceptions.
So basically Java syntax for some kind of restricted C/CUDA dialect. How can you even say you're running Java if you don't have objects or dynamic allocation? Everytime the promise of a general purpose programming language running on GPUs is made, this is actually what is delivered, eg. marketing fluff, not compilers actually getting smarter in any fashion
And when you think about how a GPU works, it completely makes sense. A high level language for a GPU will not look like Java
- alcidesfonseca 7y agoBack in 2010 I worked on a user-library and preprocessor that did the same thing [1]. Had exactly the same restrictions. Classes are supported if they act as dumb structs. Recursion is now supported in CUDA via dynamic parallelism (or faking the stack) but it is not performant at all. The ideal language for GPU programming is closer to Julia or a MATLAB-like language where arrays and matrices are first-class. [1] https://github.com/AEminium/AeminiumGPU https://github.com/AEminium/AeminiumGPU
- eternalban 7y ago> The ideal language .. That pulled up "Chapel" in my head and a couple of clicks later the net pulled this up: "PGAS (Partitioned Global Address Space) programming models were originally designed to facilitate productive parallel programming at both the intra-node and inter-node levels in homogeneous parallel machines. However, there is a growing need to support accelerators, especially GPU accelerators, in heterogeneous nodes in a cluster. Among high-level PGAS programming languages, Chapel is well suited for this task due to its use of locales and domains to help abstract away low-level details of data and compute mappings for different compute nodes, as well as for different processing units (CPU vs. GPU) within a node. In this paper, we address some of the key limitations of past approaches on mapping Chapel on to GPUs ..." https://pldi19.sigplan.org/details/CHIUW-2019-papers/4/GPUIterator-Bridging-the-Gap-between-Chapel-and-GPU-Platforms https://pldi19.sigplan.org/details/CHIUW-2019-papers/4/GPUIt...
- gauravphoenix 7y ago>And when you think about how a GPU works, it completely makes sense. A high level language for a GPU will not look like Java Can someone explain this further? I am genuinely interested in learning what makes Java etc not suitable for GPUs.
- AnimalMuppet 7y agoOther way around - GPUs are not suitable for Java. Java is a dynamically-allocated, garbage-collected language. So far as I know, GPU code doesn't have the ability to dynamically allocate memory. (I mean, I guess you could, but you'd probably need one memory pool per parallel computational unit. Otherwise, you'd need a global lock on the global memory pool.) Java is a general programming language. GPUs are not general-purpose processors. That's off the top of my head. There may be other reasons as well.
- pjmlp 7y agoCUDA can dynamically allocate, more recente versions even support unified shared memory blocks with the host processor. It is also possible to do native memory allocation in Java. Imposing restrictions on a Java subset is hardly any different of imposing restrictions on C and C++, which apparently is ok to do.
- whizzter 7y agoThe big difference that makes Java cumbersome is that you always have object references (ie pointers), thus an array of objects is really an array of pointers to objects that are in another memory location. C/C++ structs(classes) on the other hand represents a memory layout so an array of "objects" would be laid out as a single contigious block of memory. (and this works well for a GPU) There is/was a propsal for Java for value types but it hasn't yet been approved. https://openjdk.java.net/jeps/169 https://openjdk.java.net/jeps/169 (As a side note, C# classes are used as refernces like in Java but they also have "struct" in C# that behaves like C structs, and this is what the Unity boost compiler leverages and could be used for a GPU variant with less restrictions)
- untog 7y agoThe way I sometimes think about stuff like this is that it's useful in reverse: you're going to have to write optimised code that doesn't really resemble Java. But if you're in a situation where there is no GPU available, that same code will execute fine on the JVM.
- thu2111 7y agoYeah, that's TornadoVM's big thing. You just write normal code and develop it locally as you would when running on the CPU (most friendly dev environment). And then it just figures out what HW is available and uses it. This fits really nicely with scaling up on the cloud. If you have a parallel workload that you need to auto-scale very quickly, you could migrate from small CPU to big CPU, to AVX512, to GPUs to FPGAs, all automatically. Maybe it won't produce code that can beat hand-tuned Verilog but it'll do it a lot faster and cheaper. Re: no objects. That doesn't actually mean no objects. Recall this is based on Graal. Graal is really, really good at scalar replacing objects automatically, i.e. converting allocations into local variables, even when they're being e.g. allocated in a method and then returned. This is how Truffle gets such good performance: it just recursively inlines everything into one giant method and then wires up all the allocation sites to use sites. It turns out you can really use a lot of objects this way, yet at runtime they all disappear. Given it's based on the same compiler infrastructure I'd expect that to also be true of TornadoVM; perhaps a TornadoVM developer can chime in. Probably you can do method calls and working with wrapper types without hitting the limitation, or at least, it can be added without too much effort (I guess the hard part is flattening the types into arrays as value types, which may require a more sophisticated optimisation pass).
- kotselidis 7y agoCorrect, since we employ Graal for parts of the compilation, we also enjoy the benefits of inlining and escape analysis. If an object is replaced by scalars then we can use it on a GPU. Otherwise (if a real object allocation is required) then the compiler will bail out and execution will continue normally on the CPU (with compiled code from Graal or C2).
- rbanffy 7y ago
- lumost 7y agoI've never understood why modern GPUs don't address the programmability restrictions with a minimal CPU core on die. While one wouldn't choose a GPU for programs best written dynamically or recursively, having to block on the CPU for simple statements seems like an addressable bottleneck.
- kllrnohj 7y ago> I've never understood why modern GPUs don't address the programmability restrictions with a minimal CPU core on die. They basically do. There's a minimal "core" that does things like branching and that core manages a bunch of threads that all execute the same instruction in parallel. But there's a lot of those minimal cores still, so transistors spent on it need to be worth it. Spending transistors to make a nicer programming model means fewer transistors spent going faster. Alternatively the design you're looking for is Intel's Xeon Phi, which died. Intel's Xe looks maybe more traditional GPU in architecture but TBD, it might be less restrictive in this regard. Maybe.
- gridlockd 7y agoGPU programs run fast exactly because programs are written with these limitations in mind. GPUs trade stuff like branch prediction for data parallelism. You can't just have both on the same die and achieve the same throughput. Using dynamic memory liberally, having lots of indirect branches and non-local memory access - that's pretty bad for performance on the CPU as well, even though it is optimized for that. It's just that people don't notice because they're used to programming that way.
- jfim 7y ago> How can you even say you're running Java if you don't have objects or dynamic allocation? It's not any different from, say, JavaCard, which doesn't even have java.lang.String or garbage collection. > And when you think about how a GPU works, it completely makes sense. A high level language for a GPU will not look like Java It can still have Java syntax, which can also leverage existing IDEs and other libraries (eg. for unit testing).
- moron4hire 7y agoI mean, without all that stuff, is it even still Java? Is it even something more than just a few text-replacements on identically-mapped syntax elements?