4 ms·
LLaVA, Whisper and a few bash scripts should be able to do it. I don't know how helpful the model is with screenshots though. 1. Download LLaVA from https://gi
by dave1010uk 3y ago
LLaVA, Whisper and a few bash scripts should be able to do it. I don't know how helpful the model is with screenshots though.
1. Download LLaVA from https://github.com/Mozilla-Ocho/llamafile https://github.com/Mozilla-Ocho/llamafile
2. Run Whisper locally for speech to text
3. Save screenshots and send to the model, with a script like https://til.dave.engineer/openai/gpt-4-vision/ https://til.dave.engineer/openai/gpt-4-vision/