3 ms·
I've actually coded something like this. The trick is to split it in a smart way. What I did is the following: 1. Download the video 2. Split it in max length
by v4dok 3y ago
I've actually coded something like this. The trick is to split it in a smart way. What I did is the following:
1. Download the video
2. Split it in max length for whisper
3. Used whisper for transcription on the chunks
4. Used gpt for cleaning the transcription
5. Unioned the cleaned transcribed chunks
5. Used the langchain "refine" function of the summary chain
It is quite slow and expensive because the refine summary is doing a lot of calls, but the result is amazing. And you don't need a vector DB for that because the summary is serial.
For Q&A tho, an embeddings DB is a must