Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cburdick13
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
cburdick13
3y ago
The sample is really designed to show the simplicity of the syntax, and the performance is just a side effect. Where you'll see a bigger performance difference with numPy/cuPy is when kernel fusion happens where MatX is typically
2.
▲
by
cburdick13
3y ago
Hi, if you don't mind opening an issue asking for this we can run these and put in the readme.
3.
▲
by
cburdick13
3y ago
Hi, I looked into this and it seems that it is indeed using >= Depends: cuda-libraries-11-8 (>= 11.8.0), cuda-drivers (>= 520.61.05)
4.
▲
by
cburdick13
3y ago
Hi, yes, the original development was started for radar users who did not know CUDA but needed to write in c++. Many of our examples and code are radar related for that reason.
5.
▲
by
cburdick13
3y ago
Hi, I addressed the comparisons in other comments, but in general this is for c++ users and not Python. It's more of a comparison to numPy/cuPy, and we do have a table showing the comparison in the docs. We don't support auto
6.
▲
by
cburdick13
3y ago
We have benchmarks in the benchmarks directory, but these are for things like convolution, matrix multiples, etc. It's not for running a traditional benchmark set like resnet. Like most benchmarks it really depends on what you want to
7.
▲
by
cburdick13
3y ago
Hi, those libraries don't have a GPU counterpart, and for matrix multiplication we only support GPU right now.
8.
▲
by
cburdick13
3y ago
I have to admit I'm only tangentially familiar with gnuradio, but matx should be able to integrate with any C++17 codebase. We have several examples of integration with our streaming sensor pipeline called holoscan. see this radar pipe
9.
▲
by
cburdick13
3y ago
That's a fair point. I'll look into it.
10.
▲
by
cburdick13
3y ago
Hi, what specifically are you looking to benchmark on the K80? Users are free to contribute and we've had many external PRs. Contribution guide is here: https://github.com/NVIDIA/MatX/blob/main/CONTR
11.
▲
by
cburdick13
3y ago
If you use the run files rather than the package manager files it's always installed in completely separate folders inside /use/local/cuda. What you're describing with an "all" install can somewhat be acco
12.
▲
by
cburdick13
3y ago
I agree we should update it given that it's also a very old comparison (2 years now). The cuPy comparison should be fair since it was the same GPU. I addressed it more here: https://news.ycombinator.com/item?id=37760120
13.
▲
by
cburdick13
3y ago
Hi, you can choose not to install the driver with a CUDA install, and to download the driver separately.
14.
▲
by
cburdick13
3y ago
The main difference is Jax is for python primarily, while MatX is c++. This might seem like a poor answer, but in many domains (quasi-real time, signal processing, etc) the language is important to give certainty guarantees on performance.
15.
▲
by
cburdick13
3y ago
The main difference is the GPU part. This is a large difference because the same lazy evaluated template type can be run on the CPU or GPU through what we call an executor. On the CPU it's likely very similar to how xtensor is already
16.
▲
by
cburdick13
3y ago
It was mentioned elsewhere, but we support down to cuda 11.4, which supports down to the Maxwell architecture (nearly 10 years old now).
17.
▲
by
cburdick13
3y ago
Good point, and agreed the landing page is a bit sensational. I mentioned it elsewhere but between MatX and cuPy we see a 3-4x performance difference on average. The gap tends to widen with more complex workflows where compile-time kernel f
18.
▲
by
cburdick13
3y ago
We typically support whatever the underlying library supports. For int8 the support would come from cuBLASLt currently. I don't believe that or Cutlass supports mixed precision inputs, but I can check.
19.
▲
by
cburdick13
3y ago
We've tried our best to match python as well as we can, or falling back to matlab-style if Python doesn't have it. Many of our unit tests are verified against python, so the conversion is typically very easy. The one thing that py
20.
▲
by
cburdick13
3y ago
Hi, what matrix sizes and types are you working with and how many batches? In general it sounds be similar to eigen, but with GPU support. We have several svd methods for different scenarios, so if you give us the info above we can let you
21.
▲
by
cburdick13
3y ago
That's right, we've tested down to pascal, but this should work on Kepler too since CUDA and the underlying libraries support it.
22.
▲
by
cburdick13
3y ago
Hi, besides an (subjectively) easier syntax, the performance should be higher compared to libtorch. Every operator expression (think of it as an arithmetic expression) is evaluated at compile-time and is often fused into a single GPU kernel
23.
▲
by
cburdick13
3y ago
Hi, having feature parity with cuPy is a daunting task, especially for a C++ library. At this point we feel we have a good foundation for all kinds of basic and advanced tensor manipulations, and have a growing number of library-based funct
24.
▲
by
cburdick13
3y ago
Hi, MatX currently has partial support for CPUs too. Please see this comment: https://news.ycombinator.com/item?id=37758635
25.
▲
by
cburdick13
3y ago
Currently we support both CUDA and CPU to some extent. CPU is done through standard C++ (and soon stdpar). Obviously standard C++ is problematic since it doesn't include everything we support (FFTs, matrix multiplies, etc). One option
26.
▲
by
cburdick13
3y ago
Hi all, I'm one of the maintainers of MatX. I didn't expect it to hit HN this soon, but happy to answer any questions.