7 ms·
Beamforming in PulseAudio
- tacos 10y agoVersion 9.0 of the product and the guy has time to draw pictures but can't be bothered to implement an FIR filter. (Bonus: incorrectly defines an FIR filter.) Admits his solution is bad, might be buggy. Finds a better replacement, then ships his anyway. What is it about audio that attracts this curious level of "engineering"?
- arunarunarun 10y ago1. I was pretty clear about not being a DSP-person 2. The module provides infrastructure for other beamformer implementations to be plugged in, my implementation is intended to be a trivial test 3. module-beamformer did not ship (the link points to a branch in my git repository)
- jjbiotech 10y ago> What is it about audio that attracts this curious level of "engineering"? What is it about HN that attracts haters of open source POCs? Why don't you contribute and implement an FIR filter? You probably won't, because it's easier to bash someone's work rather than do the work yourself.
- tacos 10y agoI believe you spelled "open source POSs" incorrectly. Also if you want an FIR filter go here and press the button. It provides the source code for you, too. This may be the most trivial algorithm in all of signal processing, which is why it's so frustrating to see somebody completely whiff on it and a bunch of HN amateurs upvote and defend it. http://t-filter.engineerjs.com/ http://t-filter.engineerjs.com/
- dang 10y agoI believe that you know a lot about signal processing, but please don't express it in a snarky, dismissive way. If you can't post civilly and substantively, please don't post at all. https://news.ycombinator.com/newswelcome.html https://news.ycombinator.com/newswelcome.html https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- hatsunearu 10y agoPOC? Why is this a race thing?
- justinsaccount 10y agoproof of concept.
- deleted 10y ago[deleted]
- CamperBob2 10y agoFor what it's worth, I had exactly the same thought. My dang-compliant version of the same sentiment would go something like, "If you're going to write posts for public consumption, please consider limiting the topic to things you actually know about. You never know who might read and/or learn from what you write."
- JorgeGT 10y agoIf anyone is looking to experiment with beamforming algorithms, here is a free dataset from MIT with a 1020-microphone array: http://groups.csail.mit.edu/cag/mic-array/ http://groups.csail.mit.edu/cag/mic-array/
- anonbanker 10y agoStill nothing that isn't already possible in JACK. over ethernet or Wifi.
- Bromskloss 10y agoHow do you mean?
- natermer 10y agoJack is a audio daemon for Linux that performs a similar function to Pulseaudio, but is intended for a wholly different audience. Jack is intended as a low-latency audio and midi router for audio production. It routes audio data between applications and whatever analog/digital/midi I/O you have plugged into it that Linux supports reasonably well. In other words it's for build a audio workstation. So you could, provided you had the skills, quite easily create a audio processing pipeline that will do any type of 'beamforming' you want and more. Pulseaudio on the other hand is designed as a general purpose desktop and mobile audio daemon that is designed to make it simple to manage typical hardware you find in a typical PC. This sort of stuff is designed to make it nice to have a webcam for simple blog or talking to your mom, not necessarily the best choice if you want to run a music studio from your laptop.
- anonbanker 10y agoPulseaudio is still trying to catch up to features JACK implemented years ago. One of which is a beamforming implementation that could easily be done in any post-processing application plugged into the JACK chain (Trivial to do in Ardour or even FL Studio under WineASIO (Wine for Realtime music applications). This has been done in live recording environments using a couple of mic'ed-up laptops over 802.11n/ac for a while now; producing a properly-sync'ed master recording of the performance takes less than 15 minutes after the show is complete to finish.
- mmastrac 10y agoKudos to the author. It takes a lot to write about something in a way that makes it accessible to folks not skilled in the art. He's written some other interesting posts on PulseAudio features too, but this is more well-written and approachable: https://arunraghavan.net/2016/05/improvements-to-pulseaudios-echo-cancellation/ https://arunraghavan.net/2016/05/improvements-to-pulseaudios... https://arunraghavan.net/2011/08/hello-hello-hello/ https://arunraghavan.net/2011/08/hello-hello-hello/
- zimbatm 10y agoWouldn't it be possible to assume that by default, the user will be placed between both microphones? It would be less than ideal but just keeping the sound that hits both mics at the same time could be a nice default improvement.
- arunarunarun 10y agoThat is actually the default with the webrtc beamformer (the sample recordings point straight forwards which effectively works out to being between the two microphones).
- mgraczyk 10y agoYes this is the default. The beamformer is steerable and you'll notice that a significant portion of the code is responsible for steering.
- mgraczyk 10y agoGreat to see this work being talked about. I work on the media signal processing team at Google. My team built this beamformer before I started, but I'm happy to see it being used in PulseAudio. The paper hasn't been released, but the nonlinear beamformer code is open source. You can find it in WebRTC. https://chromium.googlesource.com/external/webrtc/+/master/webrtc/modules/audio_processing/beamformer/nonlinear_beamformer.cc https://chromium.googlesource.com/external/webrtc/+/master/w...
- arunarunarun 10y agoKudos to your team -- having access to the fantastic work you guys have done (especially the AudioProcessing module that we use in PulseAudio) has been very useful indeed.
- tacos 10y agoRepost as a semi-useful thread below didn't meet humor standards and people who aren't logged into HN should see it, too. If you need an FIR filter, click here and push the button. Generates the code too. http://t-filter.engineerjs.com/ http://t-filter.engineerjs.com/ Also if you don't know what you're talking about, kindly refrain from wandering into it in the middle of an article that might otherwise be useful. "An Intro To Beamforming" is a hell of a lot stronger if it doesn't have several flaming errors about basic DSP processing in the middle of it. Those sorts of errors may cause experts discovering you for the first time to avoid your project, not devote time to fixing it.
- amluto 10y agoSad that the thread got flagged into oblivion. > There is a fair amount of academic work describing methods to perform filtering on a sample to provide a fractional delay. One common way is to apply an FIR filter. However, to keep things simple, the method I chose was the Thiran approximation — the literature suggests that it performs the task reasonably well, and has the advantage of not having to spend a whole lot of CPU cycles first transforming to the frequency domain (which an FIR filter requires). I'm not a real DSP expert by any stretch of the imagination, but applying an FIR filter does not require transforming to the frequency domain. To the contrary: a basic FIR filter just means that the ith output sample is a[0]x[i] + a[1]x[i-1] + a[2]x[i-2] + ... + a[N-1]x[i-N+1] where x is the input, a is the filter coefficients, and N+1 is the (finite) length of the filter. (If the input is all zeros except that x[0]=1, then the input is an impulse and a is the output, i.e. the impulse response. Hence the name: Finite Impulse Response.) Some very long FIR filters are more efficient to apply by Fourier transforming the input, but that's almost certainly not the case here. It's worth noting that FIR filters vectorize very nicely. (My FIR description isn't quite 100% accurate. I described only the discrete-time causal case. If you drop the causality requirement (which is fine but can be awkward in real-time processing) then you add negative indices to a. If you switch to continuous time, you end up with a convolution instead of a sum of products.)
- arunarunarun 10y agoThanks for the explanation. I've corrected the post to not assert that the FFT is necessary.
- chipsy 10y agoHave some DSP resources: Richard G. Lyons, Understanding Digital Signal Processing [0] Gareth Loy, Musimathics: The Mathematical Foundations of Music (volume 2) [1] r8brain-free-src (high quality sample rate conversion algorithms) [2] KVR's DSP forum, frequented by actual pro audio developers [3] [0] https://www.amazon.com/Understanding-Digital-Signal-Processing-3rd/dp/0137027419?ie=UTF8&tag=stackoverfl08-20 https://www.amazon.com/Understanding-Digital-Signal-Processi... [1] https://www.amazon.com/gp/product/026251656X/ref=pd_cp_0_1?ie=UTF8&refRID=589F25CMK0PJ0P3B1WRW https://www.amazon.com/gp/product/026251656X/ref=pd_cp_0_1?i... [2] https://github.com/avaneev/r8brain-free-src https://github.com/avaneev/r8brain-free-src [3] https://www.kvraudio.com/forum/viewforum.php?f=33 https://www.kvraudio.com/forum/viewforum.php?f=33
- Shish2k 10y agoI wonder if this could be auto-calibrated, like prompt the user with "sit in a quiet place and say '1,2,3'" then brute-force the audio offsets to get the highest peak signal; and from then on you could have the mic focus follow the user's head as they move it by constantly trying slightly different offsets and jumping to a new offset if one sounds stronger?
- microcolonel 10y agoBetter yet, you could listen for breathing sounds if they're not speaking, and use those to find speakers. ;-)
- throwaway7767 10y agoIf this becomes popular, I could see laptop manufacturers putting information about the microphone position relative to the screen/camera somewhere (in the ACPI tables?). Then there would be less need for calibration. Of course, we'll have the same problems as all the other information in there, that some cheap equipment will have bogus values as someone just copied the ACPI tables from the previous hardware without verifying the information...
- brandmeyer 10y agoThe reason the author's beamformer doesn't work very well may have nothing whatsoever to do with his implementation. A short-baseline microphone array won't be able to have good selectivity in the human hearing range no matter what you do. All synthetic aperture systems are fundamentally limited by the wavelength of the thing being sampled. You need an aperture many wavelengths wide in order to get a significant effect at any given frequency. The tiny two-mic device shown in the picture will only be an effective aperture in the near-ultrasonic range. That's why commercially available phased array microphone systems that actually work are typically one to several meters wide.
- throwaway7767 10y agoI understood his article such that he used the same laptop to test both his own beamforming algo as well as the google one. Since the google one seems to work very well, it seems likely that the limit is not hardware. I'm actually very surprised that it works as well as it does, based on the sample recordings.
- brandmeyer 10y agoAny claim that someone beat the diffraction limit should be viewed with the same level of skepticism as faster-than-light particles, engines that don't conserve momentum, perpetual motion machines, etc.