4 ms·
I do not feel like writing an entire essay on language integration right now - most likely I will open an issue 6-12 months from now on Futhark's Github repo ju
by abstractcontrol 9y ago
I do not feel like writing an entire essay on language integration right now - most likely I will open an issue 6-12 months from now on Futhark's Github repo just to talk about it, but let me start with agreeing everything what DannyBee said and adding a few thoughts of my own.
1) Let me just say that language integration is a very serious issue. It is not just between completely different languages, but also between modules written in the same language. Once you start using algebraic datatypes to emulate language features the main language lacks, you essentially step into a dynamic sublanguage and have to deal with the friction caused by crossing module boundaries. This friction is both on the programmer's side – he has to deal with writing boilerplate for crossing the boundaries, and on the computer's side which has to do marshaling which is terrible for performance and just nasty.
2) This friction in the case of Futhark will be magnified manifold as it is a completely different language. Since Futhark is not a general purpose language, a realistic use case for it to be called indirectly directly by other languages who will generate code for it. This is generally how high-speed anything is used today. There is a collection of highly optimized, assembly written routines (such as BLAS) bundled into a library and they are called from very slow high-level languages such as Python.
Futhark today is not fit for such a purpose.
* It would be difficult to partition the program written for it into separate pieces. For very simple programs, at a minimum it will generate 2k lines of code (in the C backend).
* This is compounded by the fact that it does not link to the aforementioned optimized libraries, but generates all the code internally. You could then imagine using Futhark intermittently – calling those fast libraries in the main language and using Futhark for the rest since writing code in Futhark is much more convenient compared to C, but then you would need to partition the program and will immediately run into the code bloat issues in the first bullet point.
* Futhark is a very high level language and takes on all the responsibility for managing memory on itself. That even further clashes with the idea of partitioning the program. Memory allocations are extremely slow on the GPU, and in addition to that, they block the whole device meaning they are not asynchronous.
* A minor point of friction compared to the above is that Futhark supports OpenCL which has minor market share instead of Cuda.
3) Based on the above, I question the current integration strategy by Futhark of making backend for different languages. It currently has a C and a Python backend, and an OpenCL CPU backend, and F# is planned, and you can imagine many different backends…
What might be worth trying instead would be to make Futhark an embedded interpreted language. This is not as crazy as it seems – it would define a natural API point for other languages to access it and the interpreter would be responsible for managing GPU memory. It would be a much better model than disposing everything once the program stops running and would allow for efficient intermingling of multiple Futhark programs that could reside in memory. Right now, that sort of thing would be very high friction.
I do not have much advice on how this could be accomplished and would no doubt require much design work, but I am going to try something like that in my own language at some point. I had this crucial insight when I was trying to use it from the language it was written in and realized that it is actually very difficult.
Scala, Clojure and F# in particular had the master stroke of latching themselves to already established ecosystems. Languages targeting the GPU cannot use that strategy directly and will need to be more inventive.
- Athas 9y agoThanks for your response! Do you think some of the pain would be alleviated if Futhark had a simple FFI that would allow calling hand-written primitives when they exist? (With the property that these would be black boxes and not fused with anything else.) There will still be some friction, but it would be lessened. There are a few things I should correct: first, memory allocations on GPUs are not unusually slow (although copies from CPU to GPU are). Futhark presently stores all memory on the GPU at all times, so CPU<->GPU traffic is very low (this has other problems, however, but it seems the newest GPU hardware has features that can be used to solve this more elegantly). I've also started cooling a bit on the idea of making a lot of language-specific backends. While convenient, they are not really scalable. What will likely happen is that we will focus on improving the C backend as a target for the FFIs of other languages (since most languages provide convenient ways of calling C libraries). Possibly we'll also keep a few other strategic backends around, like the Python backend for demonstration, and possibly Java and C# backends as they can be targeted by huge (non-C) ecosystems. Using Futhark as an embedded language is an interesting idea, but I'm not sure it would solve the problems you bring up. First, the compilation technique needed to obtain good performance leads invariably to fairly slow compile times, which makes it a bit more awkward to re-compile on startup. Second, I don't see why an interpreter would be any better at managing memory than the current Futhark runtime system. If used as a library, a Futhark program does not dispose everything once it stops running, but only once the library is unloaded. For example, if a Futhark function returns an array, that array still lives on the GPU, and if you use it as an argument to another Futhark function, there will have been no traffic (except bookkeeping stuff) between CPU and GPU.
- abstractcontrol 9y ago> Do you think some of the pain would be alleviated if Futhark had a simple FFI that would allow calling hand-written primitives when they exist? Yeah, definitely. You should go a bit further and allow embedding C code for performance oriented users. I do not foresee using it, but somebody is going to need it eventually I guarantee it. I can envision some better alternatives than C, such as the language I am working on, but right now it is still incomplete and won't be for some time. > first, memory allocations on GPUs are not unusually slow The time it takes to allocate a chunk is linear in its size which is quite slow. I am not sure why that is, but maybe it is faster using OpenCL instead of Cuda? What were your timings for allocating and disposing raw memory plotted against size? > If used as a library, a Futhark program does not dispose everything once it stops running, but only once the library is unloaded. That is interesting. Can Futhark be used as a library apart from Haskell (in which it is written)? > For example, if a Futhark function returns an array, that array still lives on the GPU, and if you use it as an argument to another Futhark function, there will have been no traffic (except bookkeeping stuff) between CPU and GPU. But still, Futhark will probably not be able to optimize away all the intermediates. And more to the point, some programs like neural nets do in fact accumulate intermediates by necessity. If a particular Futhark program is run multiple times, the memory in those intermediates should be held in a pool. > Second, I don't see why an interpreter would be any better at managing memory than the current Futhark runtime system. It could potentially allow memory pool to be shared amongst multiple Futhark programs. The idea is not to turn Futhark into an interpreted language per se, it would still be a compiled language, but to instead add an extra layer that would allow easier communication with other languages. I do not have a concrete vision of how this should be done. It goes back to what you mentioned about using Futhark as a library. I am expecting a negative answer that it can be used as a library from anything other than Haskell, but if you were to go more in the direction I am suggesting, instead of making backends for C#, Java and such what you could do is make something that will allow Futhark to be used as a library. Thinking about to some of the C examples that I have seen, I do not think users will appreciate having massively bloated code files needed to compile the stuff dumped into their projects folders by Futhark. Not to mention, C# will need to be compiled to C which will result in more temporary files. Futhark as an embedded language could take responsibility for managing all of that.