7 ms·
Show HN: Mikey – No bot meeting notetaker for Windows
- dmantis 2y agoLooks cool. Is it possible to use a local model (like whisper) to avoid leaking conversations to the cloud-based AI?
- hotrod46 2y agoThat’s what’s planned next :)
- peterhorvath01 2y ago[dead]
- hotrod46 2y agohi, ive added this, lmk what you think
- alkonaut 2y agoSomething I find annoying with automatic transcriptions and summaries, like the one built into Teams, is that they lack the context necessary to properly interpret what's being said. Example if I have a meeting discussing products, abbreviations or systems with "internal" names then it can't discern them or statistically rejects them, replacing them with its best guess for a dictionary word instead. So say we have a long call involving frequent mentions about a measure called pNet pronounced in the meeting "Peenet". Then you end up with a transcription of a bunch of guys having a discussion about penises. Hilarious, the first few times. OK always hilarious, but not so useful. Being able to set the system prompt for these transcriptions would be very useful. Like "You are a friendly bot transcribing meetings at a software company. Some common terms and abbreviations you'll encounter are...".
- deleted 2y ago[deleted]
- jvanderbot 2y agoThis should be trivially solveable with a glossary as context, as you suggest. I bet the above repo would love a PR, too!
- sesm 2y agoBut the error happens in 'audio to text' part, so text prompt won't solve it. The way to fix it is probably fine-tuning the underlying audio to text model.
- alkonaut 2y agoDoing audio-to-text requires having a statistical model for what word or phrase a piece of sound is most likely to be. Without context, you can't do better than ranking the most likely candidates where a common word is more likely than an uncommon one. Having a task-specific dictionary at that point would help. One could also imagine doing it at the summary step where the AI could simply be asked to do phonetic analysis. "Here is a transcription of a meeting. Here is a list of terms/names/participants etc. Given the transcription, the meeting context/topics and assuming the transcriptor has made errors, replace similarly sounding words and terms with more likely ones from the context"
- ukuina 2y agoWhisper accepts a system prompt.
- _joel 2y agoMy favourite was Kubernetes in our meeting being referred to as Cuban Eighties. ⎈
- sys_64738 2y agoPerhaps these will be flagged for the CIA or DEA to investigate due to illegal importation of Cubans from the enemy!
- thih9 2y agoAnecdotally, if you have an accent and want to reference Maltese Falcon[1], your voice recognition software may understand it as “Maltese f* off”. [1]: https://en.m.wikipedia.org/wiki/The_Maltese_Falcon_(1941_film) https://en.m.wikipedia.org/wiki/The_Maltese_Falcon_(1941_fil...
- collinmcnulty 2y agoGong has such a feature. It’ll even expand out acronyms the first time they show up in the transcript.
- dartos 2y ago[flagged]
- deleted 2y ago[deleted]
- oersted 2y agoThere's still a surprising lack of good video call recording services that can be controlled programmatically, unlike the end-to-end SaaS apps like Read.ai or Otter.ai. The only open-source one I could find is Amurex, which looks promising. But it only supports Google Meet for now, it does it a bit differently with a Chrome extension, and it is generally rather immature, but I do wish them the best. The only API services available are Recall.ai and MeetingBaaS, they both support the big three (Google Meet, Microsoft Teams and Zoom), but they are rather expensive at $0.5 - $1 per hour. The Calendar Syncing feature is also locked behind enterprise tiers with additional monthly fees in the hundreds, and it is rather important real-world use.
- deleted 2y ago[deleted]
- jtswole 2y agoHey there The creator of Amurex here. Thank you for the kind words :D More platform support is coming very soon ;) (read next week) > The only API services available are Recall.ai and MeetingBaaS, they both support the big three (Google Meet, Microsoft Teams and Zoom), but they are rather expensive at $0.5 - $1 per hour. seems like someone has told you our internal roadmap xD but I am glad to see we are on the right track to solve the problem :D
- oersted 2y agoYou are doing great work, and I do think making it open-source is a smart strategic choice. There's still so much potential for building AI intelligence products on top of video call recordings, and right now you are offering the only practical foundation to build such systems on. I've been keeping a close eye because $1/h is unsustainable for what we are building, and there's no good reason why it should cost so much. It's manageable for early traction, but soon we'll need to consider either to build all those integrations ourselves or to build on top of Amurex. We might be contributing soon. I did see in GitHub that Teams support was almost done, exciting! Do you plan to continue with the browser extension model, or are you also looking for solutions to record meetings that happen in the Teams/Zoom native client? I think this is why most companies do it by creating a bot that joins the meeting, it's also great free advertising for them. Of course it's a bit awkward for the user, but it's becoming a normal thing, and ethically it's better to be explicit about the fact you are recording.
- sirjaz 2y agoLooks awesome, love that it is a local native app
- ForHackernews 2y ago>transcribing it using the Groq API It's not really local: it sends all the audio to some cloud AI API.
- troyvit 2y agoI'm not familiar with Groq, but it looks like: https://sdk.vercel.ai/providers/ai-sdk-providers/groq https://sdk.vercel.ai/providers/ai-sdk-providers/groq Some open models support it. It seems in theory that you could use your own cloud AI then right?
- hotrod46 2y agothats true, plan is to update to transcribe locally next
- hotrod46 2y agoive fixed it now, it now runs whisper locally to transcribe
- ttul 2y agoHas anyone done this on the Mac? I hate sending audio to Otter; it creeps me out.
- deleted 2y ago[deleted]
- doug_life 2y agohttps://speechpulse.com https://speechpulse.com does fully local audio transcription. The UI and settings are not the most intuitive, but it works fairly well and they are making constant updates.
- simplemindedbot 2y agoSpellar.ai does a great job. There’s others out there for Mac but I like Spellar’s calendar integration. Interestingly, their initial raison d’être was to help with English pronunciation and speaking speed, giving you real time feedback. They’ve downplayed this in recent releases, but the functionality is still there. Though, I’m a native English speaker and it always flagged me as pronouncing words incorrectly even though I’ve got little regional accent (I’ve been told this by others, not just my opinion. I had a speech therapist as a mother, hence little accent)
- simplemindedbot 2y agoAs an additional note, Spellar does let you bring your own Open AI key but does not allow for purely local processing. You’ve still got to send the audio out for transcription and interpretation. Also, I have no affiliation with Spellar, just a user.
- mpdaugherty 2y agoWe do this at quillmeetings.com - the audio stays on your device and is transcribed by whisper. We also do speaker splitting and recognition with a combination of models. If you share or sync notes/meetings they are e2e encrypted. FYI, the transcript-only product is free forever (it's local, so why not?), but generating AI notes, interpreting screenshots if you enable that, etc. are in the Pro plan and do require using a cloud API.
- 2y ago
- bbor 2y agoWhat does “no bot” mean? I don’t see any elaboration, tho maybe I’m just blind!
- maccard 2y agoNot affiliated, but I'd guess it doesn't have a "bot" account join the zoom/meets call
- hotrod46 2y agoThe other meeting note takers usually have a bot join the meet to take notes, that seemed a bit strange to me.
- simplemindedbot 2y agoThere’s not a “bot” that needs to attend the meeting and show up in the list of attendees thus giving away the recording of the call. Otter.ai, for instance, shows up as “Otter” (or another name) on a Zoom call when it is recording and taking notes.
- Cheer2171 2y agoOh, so it is for more "seamlessly" helping people commit the crime of wiretapping in two-party consent jurisdictions, like California? If you don't like people knowing you are recording them, you probably have a consent issue.
- stevenAthompson 2y agoYou could have said this exact same thing without it sounding like a personal attack, but you chose to be unkind instead. I wonder why?
- Cheer2171 2y agoBecause crime is bad and I don't have to be nice to those who support criminals doing crimes. If your marketing differentiator vs all the AI recording bot products is that with your product, you can record people without them knowing you are recording them... then your business model is literally to facilitate crime in many jurisdictions, including California. Let me be clear: if you have a bot capturing audio in a call you have with someone in California, and you do not tell that person you are recording them, then you have committed a felony, even if you are not in California. And what is it about you that makes you so allergic to me calling this out? I wonder why.... See, I can do that too. How does that feel? We having a good conversation here?
- mijoharas 2y agoI was looking into something like this for linux recently. Didn't find anything obviously simple (considered hooking up whisper.cpp and a bit of audio magic to make it at least transcribe, but it firstly seemed like a fair bit of a pain and secondly I couldn't think of a nice way to do speaker detection.)
- utrack 2y agohttps://github.com/m-bain/whisperX https://github.com/m-bain/whisperX looks promising - I'm hacking away on an always-on transcriber for my notes for later search&recall. It has support for diarization (the speaker detection you're looking for). I'm currently hacking away on a mix of https://github.com/speaches-ai/speaches https://github.com/speaches-ai/speaches + https://github.com/ufal/whisper_streaming https://github.com/ufal/whisper_streaming though - mostly because my laptop doesn't have a decent GPU, I stream the audio to a home server instead. But overall it's pretty simple to do after you wrangle the Python dependencies - all you need is a sink for the text files (for example, create a new file for every Teams meeting, but that's another story...)
- ewuhic 2y agoSo which are you "hacking away on" in the end?
- mijoharas 2y agoAny good solutions for capturing the audio streams and piping them where they're needed? (I.e both microphone and speakers. I was wondering if I needed to mess with pulseaudio and/or jack (I mean pipewire under the hood, but I think those APIs sit on top and might be clearer))
- mijoharas 2y agoNever mind, played around a little, and pulseaudio's cli API makes it easy enough to sling some loopback/virtual devices around that you can then read from easily enough.
- rs186 2y agoMicrosoft Teams already provides similar built-in features, along with translation, and I have to say it is one of the rare AI tools from Microsoft that makes sense and actually works -- I had good experience using it for reviewing meetings in non English language. It's not hard to imagine that this will be a standard feature of all mainstream video conference software. Wonder what is the place for these tools.
- darknavi 2y agoI've thoroughly enjoyed not having to anoint a "note taker" in my meetings in the last few months.
- deleted 2y ago[deleted]
- m348e912 2y agoI don't think this tool can do what native AI transcription integrations can do, track who is speaking. Is there any novel way of addressing that gap?
- mpdaugherty 2y agoWe did a lot of work at https://www.quillmeetings.com https://www.quillmeetings.com to build a diarization & speaker recognition pipeline that works locally on mac and windows. Basically, we can create embeddings of parts of the audio, like you might create embeddings for text for a RAG system, and cluster them (simplifying a lot of details from the "last 80%" that has taken a lot of effort to get working...) The speaker recognition can't be as perfect as listening to each stream separately like Zoom itself can do, but it also learns your contacts over time and can recognize voices for ad-hoc in-person meetings, etc. which I've found really magical since we launched it.
- prollyjethi 2y agonot open source :/
- jtswole 2y agoAh yes, a locally-run, mostly-accurate speaker recognition pipeline that isn't open source. Love to see cool features locked away while the rest of us plebs make do with whatever scraps the OSS world has managed to build. But hey, at least it kind of works, so you can enjoy your slightly-wrong diarization in private. Truly the future of meetings.
- someonehere 2y agoI’m using Granola for macOS and it’s limited to that platform. Hoping this is a good windows alternative. Wondering if anyone out there has an OSS macOS client similar to this one so I can ditch payware.
- deleted 2y ago[deleted]
- peterhorvath01 2y ago[dead]
- peterhorvath01 2y ago[dead]
- lukeluc 2y agoI'm not sure if you have any interest in porting this to Mac, but in case you do, here's some native Swift code that might help. It was built by me and a friend originally for Electron, but the repo should act as a general template. It's completely open source, and if you (or anyone) need any license modifications for any reason, just reach out: https://github.com/O4FDev/electron-system-audio-recorder/blob/main/src/swift/Recorder.swift https://github.com/O4FDev/electron-system-audio-recorder/blo...
- hotrod46 2y agothats cool, ill look into it