29 ms·
VkFFT – Vulkan Fast Fourier Transform Library
- phkahler 6y agoIsn't LGPL 2.1 is an odd license for something like this? Does it produce a library?
- microcolonel 6y ago> Does it produce a library? It is a library.
- bialpio 6y agoA _header-only_ library. Not sure how LGPL works for those - not much to avoid linking against... Throw it in your own .dll / .so and use that in your closed-source projects? Standard disclosure: IANAL.
- loa_in_ 6y agoParagraph 5 of the LGPL version 2.1 states: A program that contains no derivative of any portion of the Library, but is designed to work with the Library by being compiled or linked with it, is called a "work that uses the Library". Such a work, in isolation, is not a derivative work of the Library, and therefore falls outside the scope of this License.
- kbumsik 6y agoI am confused with that paragraph. So if I statically compiled a project with Qt without any modification, does it fall outside the scope of LGPL as well?
- rjeli 6y agoIIRC it’s unclear if static linking LGPL is okay, so most people steer clear
- cycloptic 6y agoStatic linking is okay and is allowed as a derivative work, but you need to ship object files. https://www.gnu.org/licenses/gpl-faq.en.html#LGPLStaticVsDynamic https://www.gnu.org/licenses/gpl-faq.en.html#LGPLStaticVsDyn... >(1) If you statically link against an LGPLed library, you must also provide your application in an object (not necessarily source) format, so that a user has the opportunity to modify the library and relink the application. This is covered in sections 6 and 6a of the LGPL 2.1.
- kbumsik 6y agoSo then the OP is a LGPL header-only library, which is always statically compiled in every source. Do we need to do the same thing, and if yes, how we provide an application to allow recompiling?
- cycloptic 6y agoIf you can't use the object files to rebuild another version of the program with a modifier version of the library then no that probably wouldn't be compliant. In that case you would probably have to modify the library so it's not inlining every function. (Disclaimer: IANAL)
- DTolm 6y agoHello, I am the author of VkFFT. When I made VkFFT I wanted license to be like this: your project doesn't have to be open-source, but please share your modifications to VkFFT. I think about switching it to MPL 2.0, is this one better for everybody?
- cycloptic 6y agoThat would probably do what you want because it applies on a per-file basis and not to the entire program, but make sure you read the MPL FAQ and the fine print before making such a change: https://www.mozilla.org/en-US/MPL/2.0/FAQ/ https://www.mozilla.org/en-US/MPL/2.0/FAQ/
- 6y ago
- deleted 6y ago[deleted]
- pabs3 6y agoThe program seems like it would contain inlined code from the header though right?
- formerly_proven 6y agoUh, no, it's not. The shaders are clearly part of the work, so you need to make sure that the shaders are "dynamically linked"; i.e. can be replaced by the end user with their own version in order to comply with the terms of the LGPL.
- zdw 6y agoIf I were a hiring person at AMD or Intel, I'd shortlist this guy for a job, as they need help competing against the headstart CUDA has in the GPU-base compute space.
- jjeaff 6y agoYa, but the important question is can they invert a binary tree on a whiteboard?
- mangamadaiyan 6y ago... or are leetcode-proficient, these days.
- umvi 6y agoJust get a clear whiteboard, draw the binary tree, then flip the whiteboard 180 around the vertical axis so you are now looking through the back of the whiteboard.
- repsilat 6y agoHonestly, it's an O(1) solution depending on choice of data structure, possibly constant-factor more efficient through your program depending on language, and if you're programming on a whiteboard it might also arguably be the idiomatic way to do it in that context. Limited in that you can only have one tree per whiteboard, though.
- qppo 6y ago"Write an FFT" is the DSP engineer interview question that's analogous to tree traversal algorithm whiteboarding. The hard part is remembering how a butterfly computation works, and you'll almost never need to implement it.
- TomVDB 6y agoOne should hope that the non-CUDA GPU compute library ecosystem has already advanced beyond being able to calculate FFTs!
- oxxoxoxooo 6y agoWhat is "Native zero padding to model open systems"? And how come it is "up to 2x faster than simply padding input array with zeros"?
- gct 6y agoSo you can pad your input array with zeros, but the algorithm doesn't know that it's padded, and will just compute with those zeros like any other value. If you could tell it that they were zeros it could take advantage of x*0=0 and x+0=x to significantly reduce computation. That's what I think that is.
- DTolm 6y agoThat is almost the correct answer. To go even further, there are sequences that are completely full of zeros in the padded case of multidimensional FFTs and we can omit their FFTs entirely.
- oxxoxoxooo 6y agoThank you for the reply! Could you be more specific? In the case of 1D FFT, the right half (possibly zero-padded) of the signal is completely mixed up with the left half after the first pass [of breath-first FFT]. If the right half was all zeros, would it still be twice as fast in 1D case? Do you have any pointers to literature which discusses this?
- DTolm 6y agoNo, the 1D case will mostly save on the fact that it transfers 2x times less data from the vram to the chip. The up to 2x times increase in performance was mainly related to 2D and 3D cases, where only 1/4 or 1/8 of the data is nonzero. In 2D, when doing 1D FFTs over x-axis, we omit sequences after Ny/2 because we know they are full of 0 and thus their result will be 0. So we do 0.5Ny x-axis ffts and full Nx y-axis ffts. For a square system this will mean a drop from 2N to 1.5N sequences. In 3D the drop will be even bigger, from 3N^2 to (1/4+1/2+1)=1.75N^2 sequences (almost 2x).
- Jhsto 6y agoI think this guy will have no problem getting hired. Being conscious enough to push code online works so much better than the CV preparation courses. You know you're on the right path when you are asked to play up your CV abstract than to downplay it. Personally, I would have a hard time hiring anyone without a Github account and less so working in a place where nobody has one.
- ncmncm 6y agoTo me a Gitlab account, instead, would signify superior judgment.
- adamnemecek 6y agoNot if you want your work to be discovered.
- ncmncm 6y agoOnly the best will discover it.
- solipsism 6y agoGetting downvoted, but this is no more arbitrary, myopic, and unfair to the applicant than the parent.
- deleted 6y ago[deleted]
- ncmncm 6y agoThe Microsoft Defense Force has been activated. Despite it, my statement remains true: I do, in fact, adjudge candidates the more favorably for a Gitlab account than a Github account. It demonstrates conscious choice in a knee-jerk world. (Microsoft doesn't need your assistance, boiz.)
- 6y ago
- p1mrx 6y agoHow does using Vulkan for computation fit into the OpenCL/CUDA landscape? Is CUDA's proprietary nature doing meaningful harm, and does Vulkan help?
- Jhsto 6y agoYou can run OpenCL kernels on Vulkan at least in theory: SPIR-V supports OpenCL memory model. CUDA might be machine translatable if you can compile into LLVM target (clang seems to have experimental support developed outside of Nvidia) which you then retarget into SPIR-V using a cross-compiler. The LLVM to SPIR-V cross-compiler however is limited in its translation for the time being. In general, Vulkan is a thing which commands the GPU, but is not opinionated on what the language used to represent the kernel is as long as it compiles to SPIR-V. SPIR-V in itself is like parallel LLVM IR. If you look into the project source, the shaders are in GLSL which have been pre-compiled using a cross-compiler into SPIR-V. The C file you find on the project root constitutes as the loader program for the SPIR-V files. Futhark project did some initial benchmarks on translating OpenCL to Vulkan. The results were mainly slowdowns. You can read about it in here: https://futhark-lang.org/student-projects/steffen-msc-project.pdf https://futhark-lang.org/student-projects/steffen-msc-projec...
- jgavris 6y agoWe run OpenCL on top of Vulkan in a production application on Android, thanks to a project from Google / Codeplay and other contributors https://github.com/google/clspv https://github.com/google/clspv. SPIR-V can't represent all of OpenCL, but maybe enough for most people's use cases.
- pjmlp 6y agoBadly, OctaneRender had moved away from Vulkan into CUDA, because they found out that Vulkan compute wasn't at the level that they wanted. https://home.otoy.com/octane2020-rndr-released/ https://home.otoy.com/octane2020-rndr-released/ "OTOY | GTC 2020: Real-Time Raytracing, Holographic Displays, Light Field Media and RNDR Network" https://www.youtube.com/watch?v=Qfy6CTaSHcc https://www.youtube.com/watch?v=Qfy6CTaSHcc
- meisel 6y agoWarning: LGPL license
- ncmncm 6y ago... which, being a header-only library, happens to place no restrictions or requirements of any kind on the calling program.
- detaro 6y agoI don't think it's that easy? LGPLv3 has an explicit carve-out for headers which makes that scenario easy, but this is 2.1...
- loa_in_ 6y agoParagraph 5 of the LGPL version 2.1 states: A program that contains no derivative of any portion of the Library, but is designed to work with the Library by being compiled or linked with it, is called a "work that uses the Library". Such a work, in isolation, is not a derivative work of the Library, and therefore falls outside the scope of this License.
- meisel 6y agoIn that case, if header-only is outside the scope of the license, it begs the question why they would pick that license in the first place. But anyways, it doesn't seem clear from the passage how headers fit in, considering that these headers are not just APIs, they contain the implementation themselves.
- kbumsik 6y agoIn this case I interpret LGPL as "Please don't maintain your own fork (with bugfixs) in your company internally, contribute to my repo directly", which make sense.
- 6y ago
- Mizza 6y agoI'm very eager to see GPU acceleration make its way into audio production, which is all still heavily CPU bound. A Free GPUFFT implementation will certainly help! Great work.
- adamnemecek 6y agoIt's not gonna happen, audio is much less throughput intensive but a lot more latency sensitive.
- singhrac 6y agoI've heard credible claims that GPUs these days (esp. TPUs) have lower latency for big models than CPUs. I haven't really investigated, but I could see it happening if you give the TPU a huge L1 cache or something.
- someguydave 6y agoPerhaps for large calculations? Otherwise the PCI transfer delay would be a big latency hit?
- adamnemecek 6y agoYeah until TPUs can directly communicate with the sound card, it sounds slow.
- codetrotter 6y agoI would think a GPU might help if you have a lot of audio channels and a lot of effects on each channel. But even if that is not the case, machine learning is making its way into music production tools more and more. No doubt a beefy GPU will be useful to a lot of music production professionals in the future at least, as the tools they are using begin to leverage ML more and more.
- colejohnson66 6y agoCould it be possible to “prerender” the audio on the GPU when it’s not being worked on (say, a track not being edited)? Then just play that track if it’s not edited before the user hits play?
- person_of_color 6y agoThis guy will get a foot in but still have to do a gotcha interview loop
- rektide 6y agomay someday please someone help dethrone the underlord of AI & rise us up
- slavik81 6y agoWhat are the common applications for these sorts of GPU-accelerated FFTs? We mostly just solved problems analytically in undergrad, and the little bit of naive coding we did seemed pretty fast. I feel like this must be used for problems I would have learned about in grad school, if I had continued in electrical engineering.
- HelloNurse 6y agoThe same as any FFT, but accelerated; with the tradeoff that the cost of moving data from and to the GPU needs to be amortized. It's also a good proof of concept for other kinds of GPU computations.
- Reelin 6y agoLikely any HPC application that has an FFT somewhere in its pipeline and is otherwise amenable to being run on a GPU. Fluid flow, heat transfer, and other such physical phenomena that you might want to simulate. Phase correlation in image processing is another example. (https://en.wikipedia.org/wiki/Phase_correlation https://en.wikipedia.org/wiki/Phase_correlation) MD simulations rely on FFT but I'm not sure how much is typically (or can be) done on the GPU. For example, NAMD employs cuFFT on the GPU in some cases. (https://aip.scitation.org/doi/10.1063/5.0014475 https://aip.scitation.org/doi/10.1063/5.0014475)
- amelius 6y agoMachine learning uses CNNs, which are directly based on FFTs.
- deleted 6y ago[deleted]
- fluffything 6y ago> Support for big FFT dimension sizes. Current limits: C2C - (2^24, 2^15, 2^15), What about bigger than big? > 2^29 or so ? Are these sizes for double precision ?
- DTolm 6y agoCurrently, I hit the limit of maximum workgroups amount for one submit dispatch (this is why y and z axis are lower than x one for now). It can be removed by adding multiple dispatches to the code, which I will do in one of the next updates. To go past 2^24 I need to polish the four stage FFT algorithm to allow for >2 data transfers, which I have implemented, but not yet tested. There will also be a single precision limit in this range, as the twiddle factors values will be close to 1e-8 which will be close to a machine error.
- querez 6y ago"VkFFT aims to provide community with an open-source alternative to Nvidia's cuFFT library, while achieving better performance." There are no error bars on the graphs, so it's very hard to judge if the minor differences are significant. I work in research, so probably I'm peculiar about this point, but: I'd expect better from anyone who's taken basic statistics. But from a quick look, it seems like the performance is pretty much just "on par". It would also be nice to know how performance is on other hardware. I'm assuming it's tuned to nvidida GPUs (or maybe even the specific GPU mentioned). But how does this perform on Intel or AMD hardware? How does it compare to `rocFFT` or Intel's own implementation?
- DTolm 6y agoThe FFT and iFFT are performed consecutively up to 1000 times and then each run is done 5 more times. The total result is averaged both for VkFFT and cuFFT and stays roughly the same between launches. The minor performance gains (5-20%) are noticeable. If you have a better testing technique, I am open to the suggestions. I have tested VkFFT on Intel UHD620 GPU and the performance scaled on the same rate as most benchmarks do. There are a couple of parameters that can be modified for different GPUs (like the amount of memory coalesced, which is 32bits on Nvidia GPUs after Pascal and is 64bits for Intel). I have no access to an AMD machine, otherwise I would have refined the lauch configuration parameters for it too. I have not tested other libraries than cuFFT yet.
- querez 6y agoThanks for the further clarification! If you ran this several times, you could calculate standard deviations or confidence intervals. It would be nice if you could report one such measure, so it's clearer that the differences are not just some random fluctuations. E.g. you could include them as error bars in your plots. You could also run a statistical test (in this case, a t-test is very easy to do) and report the p-value. Those are the things I'd expect my students to do if they'd have to do something like this for a report or a project, because it's the only way for people to judge if differences show clear signal or are just random fluctuations due to measurement noise. Also: I should've said this in my first post already, which in hindsight might sound too negative: I think this is a cool project and you did a great job! I just thought this might improve the presentation of your results a bit.
- bobowzki 6y agoI wonder if this works on the raspberry pi with the new Vulkan drivers.
- Lichtso 6y agoVery cool! Seems a bit more feature complete than my take on the problem: https://github.com/Lichtso/VulkanFFT https://github.com/Lichtso/VulkanFFT Still, to beat CUDA with Vulkan a lot is still missing: Scan, Reduce, Sort, Aggregate, Partition, Select, Binning, etc.
- DTolm 6y agoI have some of these routines like Reduce and Scan in my other project https://github.com/DTolm/spirit https://github.com/DTolm/spirit. It also has implementations of linear algebra solvers like CG, VP, Runge-Kutta and some others. These routines have to be inlined in users shaders in some way to have a good performance. Releasing them as a standalone library will require some thinking due to the fact that some routines have multiple shader dispatches.