8 ms·
Blaze: A High Performance C++ Math Library
- stargrazer 2y agoIs this in represented here for posterity? Last news item is 15.8.2020. There are recent commits for compiler compatibility testing (feb 2024). What is of import here?
- sevagh 2y agoPeople can post whatever they want on HN. It's a neat library, why not post it? Previous post (by the same submitter): https://news.ycombinator.com/item?id=34407106 https://news.ycombinator.com/item?id=34407106
- dannyz 2y agoIt seems like every large project these days has coalesced around Eigen, what are some of the advantages that Blaze has over Eigen?
- queuebert 2y agoOr cuBLAS. In practice, if I'm going through the trouble to rewrite math in C++, I'd rather just make GPU kernels.
- VHRanger 2y agoI mean, that only works for a small subset of workloads where the data movement patterns fit, the bandwidth is more important than the latency, etc. The reality is that almost all workloads aren't anywhere near saturating the AVX instruction max bandwidth on a CPU since Haswell.
- queuebert 2y agoDepends on whether you measure workloads as "jobs" or "flops". If "flops", I would hazard that the bulk of computing on the planet right now is happening on GPUs.
- chrsig 2y agoI'm by no means an expert in the topic, but to share my take anyway: It seems to me like there's just diminishing returns in SIMD approaches. If you're going to organize your data well for SIMD use then it's not a far reach to make it work well on a gpu, which will keep getting more cores. I imagine we'll get to a point where CPUs are actually just pretty dumb drivers for issuing gpu commands.
- gdiamos 2y agoAs someone who worked on CUDA 15 years ago - it’s amazing to me that someone on the internet posted this statement. Did GPUs win?
- thrtythreeforty 2y agoYes and no. The compute density and memory bandwidth is unmatched. But the programming model is markedly worse, even for something like CUDA: you inherently have to think about parallelism, how to organize data, write your kernels in a special language, deal with wacky toolchains, and still get to deal with the CPU and operating system. There is great power in the convenience of "with open('foo') as f:". Most workloads are still stitching together I/O bound APIs, not doing memory-bound or CPU-bound compute.
- gdiamos 2y agoCUDA was always harder to program - even if you could get better perf It took a long time to find something that really took advantage of it, but we did eventually. CUDA enabled deep learning which enabled LLMs . That's history. What surprised me about the statement was that it implied that the model of python driving optimized GPU kernels was broader than deep learning. That was the original vision of CUDA - most of the computational work being done by massively parallel cores
- chrsig 2y agoI don't think that there's a "win" here. It's just sort of which way you tilt your head, how much space do you have to cram a ton of cores connected to a really wide memory bus and how close can you get the storage while keeping everything from catching on fire, no? ("just sort of" is going to have to skip leg day because of the herculean lift it just did) It's a fairly fractal pattern in distributing computing. Move the high throughput heavy computation bits away from the low latency responsive bits ("low latency" here is relative to the total computation). Use an event loop for the reactive bits. Eventually someone will invert the event loop to use coroutines so everything looks synchronous (Go, anyone? python's gevent?). After it seems to me that the only real question is if takes too long or costs too much to move the data to the storage location the heavy computation hardware uses. There's really not much of a conceptual difference between airflow driving snowflake and c++ running on a cpu driving cuda kernels. It takes a certain scale to make going from a OLTP database to an OLAP database worth it, just like it takes a certain scale to make a GPU worth it over simd instructions on the local processor.
- Const-me 2y ago> almost all workloads aren't anywhere near saturating the AVX instruction max bandwidth on a CPU since Haswell That’s true, but GPUs aren’t only good at FLOPs, the memory bandwidth in them is also an order of magnitude faster than system memory. In my previous computer, the numbers were 484 GB/second for 1080 Ti, and 50 GB/second for DDR4 system memory. In my current one, they are 672 GB/second for 4070 Ti super, and 74 GB/second for DDR5 system memory.
- oispakaljaa 2y agoAccording to the provided benchmarks [1], it seems to be quite a bit faster. [1] https://bitbucket.org/blaze-lib/blaze/wiki/Benchmarks https://bitbucket.org/blaze-lib/blaze/wiki/Benchmarks
- dannyz 2y agoThese benchmarks look to be ~8 years old, and don't really agree with benchmarks done by other sources (https://romanpoya.medium.com/a-look-at-the-performance-of-expression-templates-in-c-eigen-vs-blaze-vs-fastor-vs-armadillo-vs-2474ed38d982 https://romanpoya.medium.com/a-look-at-the-performance-of-ex..., https://eigen.tuxfamily.org/index.php?title=Benchmark https://eigen.tuxfamily.org/index.php?title=Benchmark) In general I would be skeptical about any benchmark that claims to beat MKL significantly on standard operations
- adgjlsfhk1 2y agobeating MKL for <100x100 is pretty doable. the BLAS framework has a decent amount of inherent overhead, so just exposing a better API (e.g. one that specifies the array types and sizes well) makes it pretty easy to improve things. For big sizes though, MKL is incredibly good.
- Lockal 2y agoIf you are talking about non-small matrix multiplication in MKL, is now in opensource as a part of oneDNN. It literally has exactly the same code, as in MKL (you can see this by inspecting constants or doing high-precision benchmarks). For small matmul there is libxsmm. It may take tremendous efforts make something faster than oneDNN and libxsmm, as jit-based approach of https://github.com/oneapi-src/oneDNN/blob/main/src/gpu/jit/gemm/gen_gemm_kernel_generator.cpp https://github.com/oneapi-src/oneDNN/blob/main/src/gpu/jit/g... is too flexible: if someone finds a better sequence, oneDNN can reuse it without major change of design. But MKL is not limited to matmul, I understand it...
- VHRanger 2y agoCompile times for one. Eigen uses C++ templates to do most things, which explodes compile times.
- planede 2y agoAFAIK blaze is also somewhat heavy on templates, but maybe it uses more modern metaprogramming techniques.
- a_t48 2y agoCompile times and binary sizes :(
- touisteur 2y agoAaaand debug times. And profiling. I'd forgotten the joys of debugging/tracing heavily templated code before I jumped back into Eigen. Not that MKL was easier to debug but nowadays most of oneapi is open-source, at least the parts I use?
- 1over137 2y agoIs Eigen still alive? There's been no release in 3 years, and no news about it: https://gitlab.com/libeigen/eigen/-/issues/2699 https://gitlab.com/libeigen/eigen/-/issues/2699
- sevagh 2y agoThe master branch is active and people use Eigen today. The Discord has maintainers that are still active. Not sure how it could be considered "dead"?
- infamouscow 2y agoThe rise of frontend developers over the last 5 years learned everything must be new. That a math library of all things could be complete is several orders of thinking beyond their ability. I'm sure the gut reaction is to downvote this for the embarrassing criticism, but in all seriousness, this is the right answer.
- klaussilveira 2y agoWhat? You mean I don't need to refactor and break API every 6 months?
- sevagh 2y agoI realize asking for a new 4.0 release is fair (and the GitLab issue does have a highly upvoted request for a release). But you can't just call things "dead" for no reason, it's in poor taste. It's feature-complete, not dead!
- stanleykm 2y agoSure code can be “feature complete” but the reality is the rest of the world changes, so there will be more and more friction for your users over time. For example someone in the issue mentions they need to use mainline to use eigen with cuda now.
- flemishgun 2y agoI'm surprised people think this, there is also the widely-used Armadillo linear algebra library. In my opinion it has a much nicer syntax. https://arma.sourceforge.net/ https://arma.sourceforge.net/
- UncleOxidant 2y agoHow's the performance? EDIT: also being on Sourceforge is kind of a hinderance to discovery these days. I wonder why they chose to be on there instead of github?
- stagger87 2y agoIt's slower but maybe the target audience is different? Armadillo prioritizes MATLAB like syntax. I use armadillo as a stepping stone between MATLAB prototypes and a hand rolled C++ solution, and in many scenarios it can get you a long ways down the road.
- flemishgun 2y agoTough to say something as blanket as "it's slower"... there are lots of operations in any linear algebra library. It's not a direct comparison with other C++ linear algebra libraries, but hard to say Armadillo is slow based on benchmarks like this: https://conradsanderson.id.au/pdfs/sanderson_curtin_armadillo_pasc_2017.pdf https://conradsanderson.id.au/pdfs/sanderson_curtin_armadill...
- touisteur 2y agoOn this exact sequence, is there a LLM of choice that is really performant in this translation task? To armadillo, Eigen, Blaze or even numpy? I have had very little success with most of the open self-hosted ones, even with my 4xA40 setup, as they either don't know the c++ libraries or generate very good-looking numpy stuff, full of horrors, simple and very very subtle bugs... Looking for the same thing from any linear algebra library or language to cuda BTW (yes, calls to cu-blas/solver/sparse/tlass/dnn are OK), I haven't found one model able to write cuda code properly - not even kernels themselves but at least chaining library calls. Probably doesn't exist (invoking Cunningham's Law).
- klaussilveira 2y agoIf you want something similar, but for games: https://github.com/EricLengyel/Terathon-Math-Library https://github.com/EricLengyel/Terathon-Math-Library
- OnionBlender 2y agoWhat is the advantage over glm? The geometric algebra stuff?
- Arelius 2y agoAnother good PGA library https://github.com/jeremyong/Klein https://github.com/jeremyong/Klein
- Solvency 2y agoout of curiosity, when and/or how often do these high-performance math libraries get folded into game physics engines? Like would Blaze offer any sort of advantage if you were to develop a new 3d soft/hard body physics engine?
- cyber_kinetist 2y agoFor typical game physics engines... not that much. Math libraries like Eigen or Blaze use lots of template metaprogramming techniques under the hood that can help when you're doing large batched matrix multiplications (since it can remove temporary allocations at compile-time and can also fuse operations efficiently, as well as applying various SIMD optimizations), but it doesn't really help when you need lots of small operations (with mat3 / mat4 / vec3 / quat / etc.). Typically game physics engines tend to use iterative algorithms for their solvers (Gauss-Seidel, PBD, etc...) instead of batched "matrix"-oriented ones, so you'll get less benefits out of Eigen / Blaze compared to what you typically see in deep learning / scientific computing workloads. The codebases I've seen in many game physics engines seem to all roll their own minimal math libraries for these stuff, or even just use SIMD (SSE / AVX) intrinsics directly. Examples: PhysX (https://github.com/NVIDIA-Omniverse/PhysX https://github.com/NVIDIA-Omniverse/PhysX), Box2D (https://github.com/erincatto/box2d https://github.com/erincatto/box2d), Bullet (https://github.com/bulletphysics/bullet3 https://github.com/bulletphysics/bullet3)...
- floor_ 2y agoI don't know if I would call a math library that uses templates so liberally "high performance". High performance also includes compile time in my opinion.
- murderfs 2y agoYour opinion is wrong.
- oivey 2y agoYeah. Avoiding templates almost certainly leads to losing run time performance. The compile time is a drop in the bucket.
- anonfordays 2y agoI get the template hate, they take a while to wrap your head around and can create cryptic bugs. Nonetheless they can be extremely powerful and enable performance and reduced complexity by being a bit complex upfront.
- 392 2y agoAre there any benchmarks to show it would be noticeably faster to compile with a non-templated design?