4 ms·
Are there easy ways to implement something like that (without the phone part) with gpt4all, or other open alternatives? It would be very interesting to have his
by ensocode 2y ago
Are there easy ways to implement something like that (without the phone part) with gpt4all, or other open alternatives? It would be very interesting to have history as well and the assistant should be able to use the history for future answers.
- CuriouslyC 2y agoUse a local AI setup, SillyTavern does some context management and pseudo memory for you.
- creesch 2y agoAs far as what you need the building blocks are relatively straightforward. Though I wouldn't go as far as saying it is easy if you want everything to be easy if you want it to be open, as that effectively means self hosting. Effectively, what you are looking at is three things: 1. Voice transcription. 2. An LLM using the transcription as input with a custom system prompt. 3. Text to speech to return the text output as speech. For the first point you could look into OpenAI's whisper. Whisper is open source so you can run it locally in theory, although I don't quite know what the requirements and performance is there. The second point can be anything, really. If we keep it to being open, Ollama with a capable enough model can fit the bill. The third TTS bit I don't quite know to be honest. I have no clue what the quality is of open TTS speech engines, but I am pretty sure there are a bunch out there. Running everything locally probably does require a pretty beefy GPU to be able to run it all in a competent manner. Once you have access to all three, making something like a coach should be fairly easy. Though, fairly does a lot of heavy lifting there. It entirely depends on your coding capabilities and front-end knowledge. One thing you need to decide on is how portable this needs to be. Mostly because of security and authentication if you want to access all of this from a phone outside your local network. Assuming you want to be able to access it everywhere, I'd maybe go for a single endpoint API that accepts audio files as input and also gives audio as a response. Then on the backend you can access the various components. You need to create a UI for voice that facilitates a voice call. Having a continuously open microphone is a bit tricky because then you need to figure out stuff like detecting pauses to send the audio off. So you could for a press to talk method to simplify things. You need to figure out the prompt. Also, how much history you want to keep. This depends on the context window the LLM you are using is capable of. If you just want the LLM to respond, then this is actually easy enough. The conversation you can then simply store for reading it back later. You then need to pass the output to the TTS engine you are using and send that back to your front-end for the response. Doing all of this with the OpenAI API is also possible and would simplify a lot of things. Although, I'd probably still go for my own API and handle accessing the various OpenAI endpoints in the backend. Personally, I have been thinking about it lately. Though without the TTS aspect back, as I am not looking for a coach but have been brainstorming about an implementation that looks at my ramblings and then makes notes, reminders in my calendar and other automated actions. Basically something like your regular Google assistent/Siri/Alexa but without the vendor lock in and tailored at my own tools. So far I have kept it at brainstorming, though. But it does mean I did look into the tech stack you might need, so maybe it is helpful to you :)