3 ms·
I'm very intrigued by AI note takers, but I'm absolutely unwilling to expose me or my clients to this exact problem. The solution (theoretically) is a purely l
by wkirby 2mo ago
I'm very intrigued by AI note takers, but I'm absolutely unwilling to expose me or my clients to this exact problem.
The solution (theoretically) is a purely local note taker, but I haven't found one that's any good. Tried meetily and others in the same vein, including briefly rolling my own. The breakdown in the pipeline seems to be reliable local diarization and speaker identification; even if the transcription is good, when speakers aren't accurately identified and speech isn't well grouped, there's no rescuing it in the summary step.
- cyberge99 2mo agoDrafts.app is hideous but it has great routing capability and a dictation feature. I use it to capture what my thoughts and route based on content. I have a button that routes to an internal voice agent named KiKo. Ideas get routed to Things or todoist. Issues get routed to github, etc. It’s one universal surface for note capture. But man is it ugly.
- wkirby 2mo agoMy ideal use case is to pipe audio from both my microphone and capture system audio so things like our weekly team standup or my 1:1s with my devs can all have reliable, decent notes without taking me out of the flow of the conversation. I think clearly the _leader_ in the space is granola, but I'm just not going to use a cloud provider for this. Drafts have anything like that?
- properbrew 2mo agoI'm definitely biased as the developer, but maybe try https://whistle-enterprise.com https://whistle-enterprise.com and see how it works for you. It's a hard problem I've been working away on for a while now. It's far from perfect but every step brings it a bit closer.
- wkirby 2mo agoLiterally starting my monday weekly standup now, I'll run it and see what's up. Thanks!
- properbrew 2mo agoThank you! Feel free to drop me and email (info in bio or on the website) if you run into any issues or have any suggestions.
- WhrRTheBaboons 2mo agoSeems interesting. What would you say are the biggest missing points currently or things you want to get working/improved but couldn't yet?
- properbrew 2mo agoThank you, great question! Hard one to answer, thought about it a lot and it's going to be the diarisation of more than 5+ speakers per audio stream (your microphone + system audio for a max of 10). I actually spent a lot of time that went completely nowhere trying to fine tune my own diarisation model, it was fun to a degree but painful to see my output end up worse than what I currently had after days of work. Having a bot join the call would be such an easy way of diarising, the call software has already done it for you, but feels like a bit of a cop out. Two more improvements, audio quality improvements which is currently in the works and close to release and a new document generation model. I'm currently using a custom fine tuned Phi-4 (released December 2024!) model, that's _so old_ in the grand scheme of LLMs, I just haven't had time to benchmark and properly test some new models whilst this currently does a good job as it is. There has to be some gains here, but who knows!
- artemisart 2mo agoInterested also, which text-to-speech model do you use? For diarisation Granola uses a chrome extension instead of a bot if that can give you ideas.
- properbrew 2mo agoI think this is just to capture the audio to ship off to their servers and do the crunching. You mentioned the word "extension" and thank you so much, I've been thinking about how to do integrations as it's something a few users have mentioned, but keep the whole "completely offline" angle. I could build standalone extensions that integrate with it if it's something a user wants. So damn obvious in hindsight! Ahh yea as for the models: Speech to text - Nvidia Parakeet TDT 0.6b V3 Diarisation - Nvidia Marblenet for the speech detection, TitaNet-Large for the embeddings and then using NeMo multi-scale to do clustering around them
- halfcat 2mo ago> I'm very intrigued by AI note takers, but I'm absolutely unwilling to expose me or my clients to this exact problem Unfortunately it’s mostly not up to you. It’s a weakest-link problem. It doesn’t matter if you don’t use a note taker AI, if even one person on the call uses one. Their tool doesn’t notify you and usually the person doesn’t either. It also has the reverse impact to the person using the note taker, where people say less around them. Same as if I'm talking to someone with Meta glasses. I wonder if the people who use these tools know the people they meet with speak less during their meetings, and then all of the participants have a post-meeting call without them to say what they really thought.
- wkirby 2mo ago> It doesn’t matter if you don’t use a note taker AI, if even one person on the call uses one Yeah, but I'm unwilling to be that person.
- eterm 2mo ago> The breakdown in the pipeline seems to be reliable local diarization Yep, diarization just hasn't been well solved yet. As soon as it has, the quality in note-takers, meeting transcripts, etc, will sky-rocket across the board.
- tommat32 2mo agoThat’s a very fair concern. I completely understand why you’d want to avoid exposing yourself or your clients to that risk. I also agree that diarization and speaker identification are probably one of the hardest parts to get right if the speakers aren’t separated correctly, even a great summary won’t fix it. We’re looking into more privacy preserving approaches for MindNote (a multimodal AI Notetaker), including local/on-device processing, so this is really useful feedback. Thanks for sharing your experience!
- deleted 2mo ago[deleted]
- ljokr 2mo ago[dead]