7 ms·
Show HN: Semantic search for video
Hello HN
Over the New Year's break, I created semanticvideosearch.com. This can search any video based on meaning and context. I would love to get your feedback on it. What should I change and what can be improved?
The preprocessed videos can be search very quickly, while the youtube video links take some time (yt videos also have a upper duration limit due to compute issues). I intend to add search based on the frames of the video soon.
I would love to know your thoughts on the demo and any suggestions for improvements.
Thanks!
PS: the inspiration to create this was to get the 2 mins of content from youtube videos with 18 other mins of fluff.
- AlexanderTheGr8 4y agoConsidering that such a huge amount of information exists in videos compared to texts, I think that we are missing out on all that information from our current search engines. I hope that this can be used to bridge that gap.
- andre-z 4y ago[flagged]
- AlexanderTheGr8 4y agoI am calculating vector embeddings on 60 sec intervals with 30 sec overlap window. So overall, I only get 20-30 embeddings for a video. So a linear lookup works fine.
- andre-z 4y agoGot it. However, depends on the number of videos in the end, right? :)
- AlexanderTheGr8 4y agoI am semantically searching within 1 video only.
- andre-z 4y agoHm, a search across all videos or at least a series of videos could be a valuable feature.
- gabrielgrant 4y agoto be clear: you're calculating embeddings based on the timestamped transcript, yes? or are you doing something with the actual video content? what are you using for the embeddings?
- throwawaybutwhy 4y agoIt is generally advisable to announce potential conflicts of interest while touting a particular solution in every other comment.
- AlexanderTheGr8 4y agoThe github repo has > 3k stars so does the person still have an incentive to plug their solution everywhere?
- andre-z 4y agoAgree, should mention that I'm affiliated.
- entrepy123 4y ago0. Excellent general space, much needed. I have scripted something related out for my own needs (still need to automate even more), but a web services that accepts web URLs is much appreciated and could help others. 0b. Compare/contrast with https://freesubtitles.ai https://freesubtitles.ai from https://news.ycombinator.com/item?id=33663486 https://news.ycombinator.com/item?id=33663486 and https://github.com/mayeaux/generate-subtitles https://github.com/mayeaux/generate-subtitles The OP site looks down right now, but it seems that semantic search in this context might mean something like object detection, whereas I am very interested in semantic search as in concept, idea, topic, linguistic search. (The simplest form of this would be direct transcript search, but if that can be transformed into additional useful semantic models to be searched, that would be gravy.) That said, a constructive critique out of love follows, which may or may not fit with your goals: 1. The 10 minute length limitation makes this useless for most buried spoken content, because, you know, the longer videos have the more buried spoken content. Most detailed presentations are 15-60+ minutes where content would want to be searched IMO. A 3 hour limit would make more sense, but ideally none. 2. The heading says "any video", but the form specifies YouTube link. This needs to support Bitchute, Odysee, and Rumble, at least, because much interesting spoken content (which is totally reasonable content highly suitable for semantic search, if not to say, critically important, BTW) has been repeatedly banned from YouTube. (At least YouTube gives auto-transcripts if enabled for the video, that can be downloaded with tools. This web service actually most needs to exist for video platforms OTHER THAN YouTube, which have not implemented built-in machine-transcription options yet.) 3. The site needs SSL, or else the traffic (video links, titles, search terms) is communicated in plain text and plainly visible to all looking. Other platforms, I'm sure you can figure out. However, I don't know how you will make this support larger videos affordably, which would be a requirement for it to be really useful IMO. MOST IMPORTANT IDEA: I think you could support searching GROUPS of videos. So I want to provide URLs of 500 videos, and get their transcript indexed and searchable (with reasonable search capability--verbatim and fuzzy and regex, for example). And I want to share the link to upload more to that group, or to search just that group, with friends. But I do NOT want to simultaneously search the 50,000 videos that other people have listed on your service. (Extension 1: Another reasonable way to approach searching COLLECTIONS of videos that might satisfy 90% of the use case would be to say: let me input a list of CHANNELS TO SEARCH--YouTube, Bitchute, Odysee, Rumble channel URLs--and show me the text search results for all videos indexed that were retrieved from those CHANNELs.) I might build custom "LISTS OF CHANNELS" to search, and share those collections with friends interested in searching the same content, instead of necessarily curating granular "lists of videos" to search. Though, both might be useful. (Extension 2: Similarly, accepting input of CHANNELS TO INDEX (YouTube, Bitchute, Odysee, Rumble channel URLs) would be helpful. That way, I don't have to wait each n=? days and add a link to your site manually.)
- funfunfunction 4y agoHey this is really cool! I built something similar for https://addcontext.xyz https://addcontext.xyz over the break. We just launched a self-service version that allows anyone to enter a YouTube playlist link and create a mini search-engine from the content. I've seen a few other people working on similar stuff in the wild. Very curious to see where this goes.
- AlexanderTheGr8 4y agoVery cool work. It supports lots of videos/playlists. I presume you are using FAISS or something for vector lookups?
- funfunfunction 4y agoWe’re using Pinecone. Great experience so far.
- zikohh 4y ago> No interface is running right now
- AlexanderTheGr8 4y agosorry about that. Should work now. Was facing some compute issues.
- deleted 4y ago[deleted]
- monkeydust 4y agoNice. Is this a whisper, gpt3, embeddings OpenAI mashup?
- AlexanderTheGr8 4y agoNo gpt3. Embeddings alone can be used for semantic search.
- marginalia_nu 4y agoIt's a bit curious Google isn't already doing this, given they have both unfettered access to a metric shit-ton of video material (through youtube), a lot of computational power, as well as a lot of expertise with regards to search.
- AlexanderTheGr8 4y agoI think that google thinks that it's not worth it. They could believe that videos don't have as much relevant information as text. Or if videos do have more information, it's not worth it to try to spend so much compute for it. Plus large scale retrieval of video information is pretty challenging (even for them). Facebook had to create a new framework (FAISS) to solve this problem. So I presume it must be pretty challenging. But on a short scale, they already do this. For some queries, in the search results, they show a youtube link which automatically goes to the relevant part of the video.
- hbarka 4y agoVery cool. Any tips if the search criteria can be extended with Boolean operators?
- transitivebs 4y agoI built a similar open source app recently that you can use with any YouTube channel / playlist. The demo uses the All-In Podcast. https://github.com/transitive-bullshit/yt-semantic-search https://github.com/transitive-bullshit/yt-semantic-search
- AlexanderTheGr8 4y agoCool project!
- xatalytic 4y agoCurious if you can share more about the stack. From another comment it sounds like you're using Whisper to generate the text from audio. My default out of the box way to approach this would be something straightforward like a BERT-alike encoder to embed each target sentence in a FAISS index (hell, podcasts aren't long -- it could be brute force lookup, I suppose) or similar, with the same encoder running on the queries. Something I've been playing with is Flan-T5 (https://huggingface.co/docs/transformers/model_doc/flan-t5 https://huggingface.co/docs/transformers/model_doc/flan-t5), which has really strong out of the box question answering capabilities. I could see chunking in larger blocks and using the blocks as a context passage and the query as a question-oriented prompt. I've run some fine-tuning experiments with this setup for text generation (e.g. write me a summary of Huberman's key takes on dopamine) and find that the Flan-T5 model forgets a lot of its other capabilities when subject to fine tuning. In any event, understand if you're not inclined to share, but love talking shop on this stuff.
- AlexanderTheGr8 4y agoI love to talk about this stuff as well. My stack is video -> extract audio -> whisper for transcript -> break it into segments -> create embddings for each segment -> get query from user -> get query embeddings -> compare and show the best results There are a couple other tricks I use such as an overlap window for segments and a little post-processing for better results from the comparisons; but overall, this is the gist of it. An issue with this approach is that ques-ans doesn't work as well as I'd like bec question and answer don't necessarily have similar embeddings ("What's the dog doing?", "He's sleeping" can be 2 completely independent sentences). So I would love to investigate more into Flan-T5 for this. I am "AlexanderTheGreat#9743" on Discord if you want to discuss more.
- transitivebs 4y agoThe biggest question I have after building something similar is: what's the best way to break up transcripts into segments? You want the segments to be long enough to extract useful semantic info, but you don't want them to be too long either.
- tremarley 4y agoExcellent job. This idea was on my 'build one day' list aswel
- AlexanderTheGr8 4y agoLol same. So I decided to finally make it over new years.
- jimmySixDOF 4y agoI just stumbled across Reduct.Video who have something similar released as a service for $25/20hrs per month. They have some text to timestamp based clip editor and collaboration stuff thrown in but search is the main driver and I imagine they use a similar pipeline.
- rahuldan 4y agoHey really cool project!! I also built a semantic search engine for codebases using Openai's embedding and FAISS https://github.com/rahuldan/codesearch https://github.com/rahuldan/codesearch