Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
DTolm
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
AMD MI300X and Nvidia H100 benchmarking: VkFFT, cuFFT and rocFFT comparison
(old.reddit.com)
1 points
by
DTolm
2y ago
|
0 comments
2.
▲
by
DTolm
3y ago
Hello, I am the author of VkFFT, Tolmachev Dmitrii. I remember VkFFT got a lot of initial traction thanks to Hacker News three years ago. Back then VkFFT was a simple collection of pre-made shaders for powers of two FFTs. Nowadays it is bas
3.
▲
VkFFT – A Performant, Cross-Platform and Open-Source GPU FFT Library
(ieeexplore.ieee.org)
2 points
by
DTolm
4y ago
|
1 comments
4.
▲
by
DTolm
4y ago
The white paper of VkFFT is out. It can be interesting to people who want to know more about performant GPU algorithms in HPC and how VkFFT is designed. VkFFT is an efficient GPU-accelerated multidimensional Fast Fourier Transform library f
5.
▲
VkResample – real-time Vulkan FFT upscaling
(github.com)
3 points
by
DTolm
6y ago
|
0 comments
6.
▲
VkResample – real-time Vulkan FFT upscaling
(github.com)
2 points
by
DTolm
6y ago
|
0 comments
7.
▲
by
DTolm
6y ago
In VkFFT FFT is done as a part of the so called compute queue. There is no need to transfer data through PCI to show it on screen - there is a graphics queue that has this as its main purpose. You can use compute queue results in it and do
8.
▲
by
DTolm
6y ago
Big 1D FFTs also take a lot of memory by themselves (i.e. 2^28 takes 2GB just to store complex data). Multiple smaller batches can be used in ML applications for example for big kernel convolutions. All learning can actually be done without
9.
▲
by
DTolm
6y ago
There are surely many different ways to get the job done. VkFFT.h file by itself desn't do any computations btw - it is more like a configurator that launches shaders (sth similar to kernels in CUDA). Having only one header also makes
10.
▲
by
DTolm
6y ago
1k FFT size in single precision is 1024 x 2 x sizeof(float) = 8KB. If we don't think that it won't utilize full GPU (not even one compute unit) and assume that it scales similarly to big systems then: 1)165GB/s is an algorith
11.
▲
by
DTolm
6y ago
FFT is an extremely bandwidth limited problem, so if most time is taken by one upload by both algorithms, the overall time will be similar. More in-depth analysis of how VkFFT and cuFFT scales with memory clocks and bandwidth can be found h
12.
▲
by
DTolm
6y ago
The library only includes vkFFT.h file (in C) and a set of shaders (C-like language compiled to SPIR-V). Vulkan_FFT.cpp is only an example that shows how VkFFT can be used. It also contains the benchmark in it, but it is not a part of the l
13.
▲
by
DTolm
6y ago
Actually, it is still best to aim at zero transfers between GPU and CPU during the execution. The GPU is limited by VRAM-chip bandwidth which is much bigger than the PCI-E bandwidth. And it should not be affected by SAM.
14.
▲
by
DTolm
6y ago
Yes, this is indeed something I would like to add in the future. While adding different radix kernels support for small prime factors is not that hard, writing efficient scheduler is a much more challenging task (each sequence, even for pow
15.
▲
by
DTolm
6y ago
It is a great open-source license for library projects. For example, Eigen uses it: https://eigen.tuxfamily.org/index.php?title=News:Relicensing... !
16.
▲
by
DTolm
6y ago
Hello! Since the last post VkFFT has experienced a number of huge improvements and optimizations. Namely: -It now supports sequences up to 2^32 in all dimensions (algorithmically, in reality limited to allocatable memory size, switch to 64-
17.
▲
VkFFT – Vulkan Fast Fourier Transform library
(github.com)
123 points
by
DTolm
6y ago
|
50 comments
18.
▲
RTX 3080 and Radeon VII benchmark results in VkFFT against cuFFT
(github.com)
2 points
by
DTolm
6y ago
|
0 comments
19.
▲
VkFFT – Vulkan Fast Fourier Transform Library Now MPL 2.0
(github.com)
2 points
by
DTolm
6y ago
|
0 comments
20.
▲
by
DTolm
6y ago
Hello, I am the creator of VkFFT library( https://github.com/DTolm/VkFFT ). I was recently invited to be a speaker at the EUROfusion Webinar #5 about Vulkan Compute. I would like to share the recording of it with the com
21.
▲
Vulkan Compute: sample transposition code and webinar recording
(github.com)
2 points
by
DTolm
6y ago
|
1 comments
22.
▲
by
DTolm
6y ago
If you happen to need any assistance in refining VkFFT for your use case, feel free to contact me.
23.
▲
by
DTolm
6y ago
Well, most likely I won't be able to help explaining the fluctuation easily then, as you have spent a lot of time on it already. It would be cool to try VkFFT in this usage scenario at some pont in the future though - it also can do 1D
24.
▲
by
DTolm
6y ago
I have some of these routines like Reduce and Scan in my other project https://github.com/DTolm/spirit . It also has implementations of linear algebra solvers like CG, VP, Runge-Kutta and some others. These routines hav
25.
▲
by
DTolm
6y ago
Small FFTs like 2048 only utilize one SM and the way they are given to the GPU may produce some fluctuations. It also depends on the way your code works. Synchronizations are also more impactful in this case. Do you launch a big grid that c
26.
▲
by
DTolm
6y ago
He is not wrong, convolutions between an image and a small kernel can be done faster by direct multiplication than by padding the kernel and performing FFT + iFFT. This is what tensor cores are aiming to do really fast. However, doing a con
27.
▲
by
DTolm
6y ago
I use the averaged data of 1000 merged launches and then average the end result over a number of runs. Merging FFT calls is actually the way how I use VkFFT in Vulkan Spirit (with some other shaders between), so this benchmark is fairly clo
28.
▲
by
DTolm
6y ago
GPU is a very consistent device, so the purpose of such big sample sizes and multiple launches with averaging is to reduce all the deviations almost to zero. The error is <1% in this case and showing it on the plot will not really change
29.
▲
by
DTolm
6y ago
The FFT and iFFT are performed consecutively up to 1000 times and then each run is done 5 more times. The total result is averaged both for VkFFT and cuFFT and stays roughly the same between launches. The minor performance gains (5-20%) are
30.
▲
by
DTolm
6y ago
I have used VkFFT to create GPU version of a magnetic simulation software Spirit ( https://github.com/DTolm/spirit ). Except for FFT it also has a lot of general linear algebra routines, like efficient GPU reduce/sc
More ›