25 ms·
I want a good parallel language [video]
- ChadNauseam 11mo agoRaph and I also talked about this subject here: https://www.popovit.ch/interviews/raph-levien-simd https://www.popovit.ch/interviews/raph-levien-simd The discussion covers things at a relatively basic level as we wanted it to be accessible to a wide audience. So we explain SIMD vs SIMT, predication, multiversioning, and some more. Raph is a super nice guy and a pleasure to talk to. I'm glad we have people like him around!
- SamInTheShell 11mo agoWent in thinking "Have you heard of Go?"... but this turned out to be about GPU computing.
- nasretdinov 11mo agoWell, they said "good" :). Go we already have, that is correct. P.S. I'm joking, I do love Go, even though it's by no means a perfect language to write parallel applications with
- SamInTheShell 11mo agoHaha, yeah. Idk any other language where I practically get a free parallelization+concurrency sandwich. It's kept me coming back to Go for a decade now, despite them using a signal that prevents using it for system level libraries. They literally broke my libnss-go package years ago when they selected the signal to use to control the concurrency portion of the runtime.
- pdimitar 11mo agoElixir. Parallelism is trivial and front-and-center. And no it's not a niche language. Don't listen to the army of Python technicians.
- SamInTheShell 11mo agoThere seems to be about the same level of effort with Elixir as there is for a language like Kotlin. The free sandwich I'm referring to with Go is the ability to just do `go funcnamehere()` and that's running concurrently and in parallel. If I need coordination of those goroutines, I can still do that with any number of locking patterns. It's extremely convenient, making the trade off of having a runtime baked in worth it imo.
- pdimitar 11mo agoWell, almost exactly the same goes for Elixir. You have to declare a supervisor first but from then on it's trivial, almost on the level of `go doThisThing()`. That's why I made Elixir my main language. A lot of the tech sphere stubbornly pretends we don't live in a world with multicore CPUs, even to this day.
- deleted 11mo ago[deleted]
- fifilura 11mo agoSQL. It is a joke, but an SQL engine can be massively parallel. You just don't know it, it just gives you what you want. And in many ways the operations resembles what you do for example in CUDA. CUDA backend for DuckDB or Trino would be one of my go-to projects if i was laid off.
- drivebyhooting 11mo agoMy issue with SQL is lack of composability and difficulty of debugging intermediate results.
- asadm 11mo agois it a language problem though? it's just lack of tooling.
- theLiminator 11mo agoThe dataframe paradigm (a good example being polars) is another good alternative that's more composable (imo).
- fifilura 11mo agoIt is true. I still hate it. I think because it always offers 10 different ways to do the same thing. So it is just too much to remember.
- mamcx 11mo agoYes, SQL is poor. What could be good is relational + array model. I have some ideas on https://tablam.org https://tablam.org, and building not just the language but the optimizer in tandem I think will be very nice.
- oembar4 11mo agoThe programming style reminds me of the old days of clipper and xbase family, even ABAP. I like the syntax.
- v9v 11mo agoThere were a few languages designed specifically for parallel computing spurred by DARPA's High Productivity Computing Systems project. While Fortress is dead, Chapel is still being developed.
- zokier 11mo agoiirc those were oriented more towards large HPC clusters rather than computation on single node?
- Jtsummers 11mo agoChapel, at least, aims for both. You can write loops that it will try to compile to use SIMD instructions, or even for the GPU: https://chapel-lang.org/docs/technotes/gpu.html https://chapel-lang.org/docs/technotes/gpu.html
- convolvatron 11mo agothe distinction matters less and less. Inside the GPU there is already plenty of locality to exploit (catches, schedulers, warps). nvlink is a switch memory access network, so that already gets you some fairly large machines with multiple kinds of locality. throwing infiniband or IP on top is really structurally more of the same. Chapel definitely can target a single GPU.
- jandrewrogers 11mo agoThose languages were not effective in practice. The kind of loop parallelism that most people focus on is the least interesting and effective kind outside of niche domains. The value was low. Hardware architectures like Tera MTA were much more capable but almost no one could write effective code for them even though the language was vanilla C++ with a couple extra features. Then we learned how to write similar software architecture on standard CPUs. The same problem of people being bad at it remained. The common thread in all of this is people. Humans as a group are terrible at reasoning about non-trivial parallelism. The tools almost don't matter. Reasoning effectively about parallelism involves manipulating a space that is quite evidently beyond most human cognitive abilities to reason about. Parallelism was never about the language. Most people can't build the necessary mental model in any language.
- cubefox 11mo agoUnfortunately his microphone did not cooperate.
- abejfehr 11mo agoBend comes to mind as an attempt at this: https://github.com/HigherOrderCO/Bend https://github.com/HigherOrderCO/Bend Disclaimer: I did not watch the video yet
- dandanua 11mo agoI think a good parallel language will be the one that takes your code written with tasks and channels, understands its logic, rewrites and compiles it in the most efficient way. I don't feel that I have to write something harder than that as a pity human.
- convolvatron 11mo agomapping from channels to SIMD seems kind of intractable, its a kind of lifting that involves looking across the producers and the consumers. going the other direction, making channel runtimes run SIMD, is trivial
- Munksgaard 11mo agoInteresting talk. He mentions Futhark a few times, but fails to point out that his ideal way of programming is almost 1:1 how it would be done in Futhark. His example is: sequence .map(|x: T0| ...: T1) .scan(|a: T1, b: T1| ...: T1) .filter(|x: T1| ...: bool) .flat_map(|x: T1| ...: sequence<T2>) .collect() It would be written in Futhark something like this: sequence |> map (\x -> ...) |> scan (\x y -> ...) |> filter (\x -> ...) |> map (\x -> ...) |> flatten
- Munksgaard 11mo agoAlso, while not exactly the algorithm Raph is looking for, here is a bracket matching function (from Pareas, which he also mentions in the talk) in Futhark: https://github.com/Snektron/pareas/blob/master/src/compiler/parser/bracket_matching.fut https://github.com/Snektron/pareas/blob/master/src/compiler/... I haven't studied it in depth, but it's pretty readable.
- pythomancer 11mo ago(author here) check_brackets_bt is actually exactly the algorithm that Raph mentions
- Munksgaard 11mo agoThanks for clarifying! It would indeed be interesting to see a comparison between similar implementations in other languages, both in terms of readability and performance. I feel like the readability can hardly get much better than what you wrote, but I don't know!
- raphlinus 11mo agoRight. This is the binary tree version of the algorithm, and is nice and concise, very readable. What would take it to the next level for me is the version in the stack monoid paper, which chunks things up into workgroups. I haven't done benchmarks against the Pareas version (unfortunately it's not that easy), but I would expect the workgroup optimized version to be quite a bit faster.
- MangoToupe 11mo agoprolog?
- swatson741 11mo agoSo he wants a good parallel language? What's the issue? I haven't had problems with concurrency, multiplexing, and promises. They've solved all the parallelism tasks I've needed to do.
- pbronez 11mo agoThe audio is weirdly messed up
- raphlinus 11mo agoYes, sorry about that. We had tech issues, and did the best we could with the audio that was captured.
- awaymazdacx5 11mo agoLower-level programming language, which is either object-oriented like python or after compilation a real-time system transposition would assemble the microarchitecture to an x86 chip.
- pancakeguy 11mo agoWhat about burla.dev ? Or basically a generic nestable `remote_parallel_map` for python functions over lists of objects. I haven't had a chance to fully watch the video yet / I understand it focuses on lower levels of abstraction / GPU programming. But I'd love to know how this fit's into what the speaker is looking for / what it's missing (other than obviously it not being a way to program GPU's) (also full disclosure I am a co-founder).
- RobotToaster 11mo agoVHDL?
- raphlinus 11mo agoI almost mentioned it in the talk, as an example of a language that's deployed very successfully and expresses parallelism at scale. Ultimately I didn't, as the core of what I'm talking about is control over dynamic allocation and scheduling, and that's not the strength of VHDL.
- TJSomething 11mo agoIt seems like there are two sides to this problem, both of which are hard and go hand in hand. There is the HCI problem of having abstractions are rich enough to handle problems like parsing and scheduling on the GPU. Then you need a sufficiently smart compiler problem of lowering these problems to the GPU. But of course, there's a limit to how smart a compiler can be, which loops back to your abstraction design. Overall, it seems to be a really interesting problem!
- AllegedAlec 11mo agoctrl-f Erlang Nothing yet? Damn...
- rramadass 11mo agoYeah, i too was looking for Erlang. The thing i would really like to see is some research on how to run the Erlang concurrency model on a GPU.
- deleted 11mo ago[deleted]
- jerf 11mo agoThere's no need for research. The answer is simple: You can't run Erlang concurrency on a GPU. GPUs fundamentally get their advantage by running the same operations on a huge set of cores across different data. They aren't just Platonically faster than CPUs, they're faster than CPUs on very, very specific tasks. Out of the context of those tasks, they are in fact massively, massively slower. Some of the operations Erlang does, GPUs don't even want to do at all, including basic things like pattern matching. GPUs do not want that sort of code at all. "Erlang" is being over specific here. No conventional CPU language makes sense on a GPU at all.
- rramadass 11mo agoNot what i meant (these are superficialities). Erlang is a concurrency-oriented language though its concurrency architecture (multicore/node/cluster/etc.) is different from that modeled by GPUs (Vectorized/SIMD/SIMT/etc.) Since share-nothing Processes (so-called Actor model) are at the heart of the Erlang Run Time System(ERTS)/BEAM it is easy to imagine a "group of Erlang processes" being mapped directly to a "group of threads in a warp on a GPU". Of course the Erlang scheduler being different (it is reduction based and not time sliced) one would need to rethink some fundamental design decisions but that should not be too out-of-the-way since the system as a whole is built for concurrency support. The other problem would be memory transfers between CPU and GPU (while still preserving immutability) but this is a more general one. You can call out to CUDA/OpenCL/etc. from Erlang through its C interface (Kevin Smith did a presentation years ago) but i have seen no new research since then. However, there has been some new things in Elixir land notably "Nx" (Numerical Elixir) and "GPotion" (a DSL for GPU programming in Elixir). But note that none of the above is aimed at modifying the Erlang language/runtime concurrency model itself to map to GPU models which is what i would very much like to see.
- nailer 11mo agoWas trying to remember where I recognised this name, Raph Levien is the Ghostscript and Advogato creator and helped legalize crypto https://en.wikipedia.org/wiki/Raph_Levien https://en.wikipedia.org/wiki/Raph_Levien
- zhethoughts 11mo agohttps://docs.ray.io/en/latest/ https://docs.ray.io/en/latest/
- coffeeaddict1 11mo agoI was expecting the author to at least mention Halide https://halide-lang.org/ https://halide-lang.org/.