2 ms·
I've tried using LLMs to restore and clean up raw unpunctuated transcripts, but they tend to hallucinate new words. And chunking is an issue since long transcri
by ldenoue 2y ago
I've tried using LLMs to restore and clean up raw unpunctuated transcripts, but they tend to hallucinate new words. And chunking is an issue since long transcripts need to be split according to the LLM input context size. But where do we split if we don't have the punctuation yet?
In Scribe I chose instead to restore punctuation marks using a token-classifier (here a DistilBert running in your browser) https://www.appblit.com/scribe https://www.appblit.com/scribe
Laurent