14 ms·
Show HN: YouTube Full Text Search – Search all of a channel from the commandline
yt-fts is a simple python script that uses yt-dlp to scrape all of a youtube channels subtitles and load them into an sqlite database that is searchable from the command line. It allows you to query a channel for specific key word or phrase and will generate time stamped youtube urls to the video containing the keyword.
- theptrk 3y agoNice work highlighting that "Life In The Big City" classic from the Ben Avery days
- foderking 3y agonice
- derefr 3y agoI love that a third party is stepping up here, but it's incomprehensible to me that Google doesn't do this themselves. They're a search company, and they own YouTube. The YouTube data — including the subtitle files — is already sitting there on their servers; they don't have to scrape it, they just have to index it. What are they even doing? Fun thing to try: do a Google search with "site:youtube.com" in it. You get basically nothing, no matter what keywords you use. It seems that Google actually entirely ignores/excludes YouTube from their regular HTML indexing, and instead only relies on the YouTube backend to actively push content into (a special, separate part of) the search index. Which gets you "results from YouTube" and "video search" — but doesn't get you the ability to search youtube videos pages qua web pages. (Consider: you can find a post in a Reddit comment thread on Google. Can you find a post in a YouTube video comments section on Google?) Heck, when I first heard about YouTube's autogenerated captions, my first thought was "oh, so this is Google building deep indexing of video through audio transcription, because they can't trust externally-provided subtitles, right?" But it's been 10 years, and I couldn't have been more wrong.
- mrazomor 3y agoProbably because YouTube =! Google Search, while YouTube is still a subset of Google. So, going an extra mile for YouTube and not for others might put Google Search in anti competition issues. Then again, I also find it absurd. YouTube is one of the most valuable parts of the Internet. And its lack of searchability is criminal. At least the YT search itself should make up for it. It's shame it doesn't.
- derefr 3y agoGoogle doesn't necessarily have to do anything special for YouTube, though. Google could "just" index YouTube videos as if they were any other web pages, in a standard way. It would then be YouTube's job, to make the data inside those video pages legible to Google's indexer. Where Google could enable this, by pushing for web standards to increase machine-legibility of video in HTML — e.g. standardized ARIA-accessible captions sources for the <video> element, etc. If they got it set up such that in theory any web spider could come along and index a YouTube video — then there would be no anti-trust reason that Google couldn't just directly ingest the subtitle files off their own servers; it'd just be a bandwidth-saving optimization over the scraping process that they could otherwise do.
- userbinator 3y agoe.g. standardized ARIA-accessible captions sources for the <video> element, etc. YouTube could literally be a minimal web forum with a video tag in the first post of each thread, but likely due to DRM and related motivations, they instead wrap everything in thick layers of obfuscated JS. Comments were easily indexable too, before that was also removed: https://news.ycombinator.com/item?id=11053204 https://news.ycombinator.com/item?id=11053204 For a while there were various shady-looking sites that seemed to scrape YouTube video pages (including comments) and I could sometimes find them through Google (then going back to YouTube for the original video), but within the past few years those have unfortunately also either been delisted/censored from search results or died out.
- stingraycharles 3y ago
- ivan2kh 3y agoNext step is to prettify subtitles into sentences using one of LLMs.
- dpflan 3y agoNice. Combine this with an "ascii-art" the video converter in the terminal? There are some existing tools, a brief search yields this UNIX StackExchange discussion: https://unix.stackexchange.com/questions/160212/watch-youtube-videos-in-terminal/160221#160221 https://unix.stackexchange.com/questions/160212/watch-youtub...
- gorgoiler 3y agoVery nice! FYI: sqlite ships with a full text search engine featuring a Boolean query language, highlight(), snippet() and scoring: https://www.sqlite.org/fts5.html https://www.sqlite.org/fts5.html I’ve not used it with enough content to know how much faster it is than LIKE ‘%my query%’ but it should be a lot quicker. (Also, in most cases you don’t need to create an id column — every table has one already in the form of rowid.)
- Boltgolt 3y agoIn what cases is it unwise to rely on rowid over a id field?
- gorgoiler 3y agoI don’t think there are any. They are one and the same — if you create an integer primary key named id it is aliased to rowid: https://www.sqlite.org/rowidtable.html https://www.sqlite.org/rowidtable.html
- Scaevolus 3y agoThis is safer-- relying on implicit rowids can break things if you use a plaintext database dump (those don't have rowids), and having a column "id integer primary key" is clearer in the schema.
- eternauta3k 3y agoI think the integer autoincrement primary key is more explicit / less mysterious than the implicit rowid. Even if most of us have run into that explanation in the sqlite manual.
- lennxa 3y agoNot sure if this includes fuzzy search, but having it will make this much more usable.
- lessname 3y agoWhat I really would like to see on youtube is a full text search on video content, at least for videos with subtitles.
- warangal 3y agoFor what it is worth, we work on a tool[0] to index all local videos and images and later allowing query just using natural language. It is based on CLIP which has been trained on image-text pairs, but seems to work great for videos after applying some naive heuristics. [0] https://github.com/ramanlabs-in/hachi https://github.com/ramanlabs-in/hachi
- aabbcc1241 3y agoYouglish [1] is a website that allow you to search video with timestamp by transcript text [1] https://youglish.com/ https://youglish.com/
- jimmySixDOF 3y agoI stumbled across a ShowHN that did not get to the front page but seems to fit here: https://clipbase.xyz/ https://clipbase.xyz/
- ggerganov 3y agoFor videos without subtitles one could chain Whisper to auto-generate transcripts, though that would require downloading the audio and processing it
- cced 3y agoThis is exactly what I’ve built. Nothing fancy just ytdlp + whisper + ripgrep + fzf and I’ve got a pretty interesting way to ctrl+f my YT history.
- gukoff 3y agoMind to share? I'd like to try this out
- zdwolfe 3y agoNot OP, but I too wrote something nearly identical with whisper so I could creep on old EthosLab videos. Here's the gist: from pytube import Channel import whisper channel_yt = Channel(channel_url) video_yt = channel_yt.videos[0] video_yt_stream = video_yt.streams.filter(mime_type="video/mp4").filter(res="720p").first() video_yt_video_file_path = video_yt_stream.download() audio = whisper.load_audio(video_yt_video_file_path) model = whisper.load_model("tiny") transcript = model.transcribe(audio)
- monkeydust 3y agoNice work. You could encode the text, load this into a vector database and allow semantic search.
- Reflecticon 3y agoSomething like this? https://dexa.ai https://dexa.ai
- monkeydust 3y agoYes in theory although they are pretty expensive. I am doing something like this at work as I wanted to unlock the wealth of information we have in our tutorials, webinars etc.
- pyinstallwoes 3y agohttps://weaviate.io/ https://weaviate.io/ Looks interesting. I was just reading about it.
- rahimnathwani 3y agoIf you're getting started with Weaviate, these two are probably what you need: 1. Wizard to create a docker-compose file: https://weaviate.io/developers/weaviate/installation/docker-compose https://weaviate.io/developers/weaviate/installation/docker-... (e.g. choose the embedding model) 2. Sample notebook showing how to index items using the python library: https://github.com/weaviate-tutorials/vector-provision-options/blob/main/tutorial.ipynb https://github.com/weaviate-tutorials/vector-provision-optio...
- pknerd 3y agoPardon my ignorance as I have not worked on Vector DBs yet, could you come up with an example how it'd be different than a full text search?
- bigzyg33k 3y agoHere's a (kinda) ELI5: you would use a language model to create "embeddings" of the text, which you can think of as a set of numbers representing the "meaning" of a set of characters. These numbers can be plotted as points in a space, and embeddings of things with similar meanings are plotted close to each other. So things like "exam preparation" would have embeddings close to things like "top study tips". Say you have created embeddings for a large corpus of text (in this case all youtube captions) once. If you create embeddings for a user query, you can search for embeddings close to it, and these will be "semantically" similar to the query. The advantage is that unlike traditional full-text search, the user doesn't need a query that includes words present in the text.
- BoBoiBoy 3y ago[flagged]
- rmdes 3y agoWondering if this could be expanded to also search for comments
- pknerd 3y agoLoved the idea, yet to try it out. I would definitely download all videos of Lex and run some text analyzer/text cloud generator to learn about things being discussed
- xurukefi 3y ago$ python yt_fts.py download 'https://www.youtube.com/@ycombinator/videos' [...] File "/app/yt_fts.py", line 176, in get_channel_id channel_id = re.search('channelId":"(.{24})"', html).group(1) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ AttributeError: 'NoneType' object has no attribute 'group' $ python yt_fts.py download 'https://www.youtube.com/@ycombinator/videos' --channel-id UCj089h5WsDdh1q8t54K3ZCw [...] File "/app/yt_fts.py", line 191, in get_channel_name data = json.loads(script.string) ^^^^^^^^^^^^^ AttributeError: 'NoneType' object has no attribute 'string' works great
- LaputanMachine 3y agoAs a workaround you can manually set the channel_name in line 82
- lfconsult 3y agoThis tool itself does works great. This behavior is due to the YouTube cookies consent page. I opened an issue about this specific issue: https://github.com/NotJoeMartinez/yt-fts/issues/1 https://github.com/NotJoeMartinez/yt-fts/issues/1 Glad if you want to help.
- lfconsult 3y agoFixed.
- smcleod 3y agoReminds me how much I dislike python error messages, not as much as java but still so much noise to signal by default.
- nomilk 3y ago> yt-fts is a simple python script that uses yt-dlp to scrape all of a youtube channels subtitles and load them into an sqlite database that is searchable from the command line. Critically, this is per channel. I wonder if we can optionally configure this to share the downloaded transcripts to a central repository so eventually a good proportion of youtube's transcripts could be downloaded as one big text file.
- hnarn 3y ago> share the downloaded transcripts to a central repository Sure, are you willing to host it and handle the absolutely inevitable legal issues?
- deleted 3y ago[deleted]
- moritonal 3y agoI wish IPFS was better. It'd be an obvious solution to this. Content hash the YouTube ID and then distribute hosting.
- prometheon1 3y agoThere is something similar to a central repository at https://filmot.com/ https://filmot.com/
- liampulles 3y agoPut the subs into a vector db instead and enable semantic search. :)
- lopatin 3y agoThis will come in handy. I’ve always wanted to count how many times Lex Fridman has referred to something as a beautiful dance.
- noman-land 3y agoI want to do a word count on the word "love".
- notjoemartinez 3y ago> yt-fts search "love" --channel "Lex Fridman" | grep "love" | wc -l > 7060
- lopatin 3y agoLegend
- noman-land 3y agoThis comment blew my mind.
- noman-land 3y agoUpdate, I couldn't get this to work. It returns 0 for me. Running it without the grep etc says channel not found.
- noman-land 3y agoUpdate, I had to download all the video subtitles and then query the sqlite tables directly. Below are the top 10 from just the podcast. https://www.youtube.com/playlist?list=PLrAXtmErZgOdP_8GztsuKi9nrraNbKKp4 https://www.youtube.com/playlist?list=PLrAXtmErZgOdP_8GztsuK... First column is "love"s per episode. Total "love" count in 376 episodes = 7,614 Average "love"s per episode = 20.25 ----------------- 107 | Sarma Melngailis: Bad Vegan, Fraud, Pris | iZjby1LkTWQ 98 | Andrew Huberman: Focus, Stress, Relation | lvh3g7eszVQ 92 | Bishop Robert Barron: Christianity and t | WgytXF0SPh0 80 | David Buss: Sex, Dating, Relationships, | sndW9hzX-wA 79 | Duncan Trussell: Comedy, Sentient Robots | jdIyNMkusLE 76 | Rana el Kaliouby: Emotion AI, Social Rob | 36_rM7wpN5A 75 | Edward Frenkel: Reality is a Paradox - M | Osh0-J3T2nY 75 | Todd Howard: Skyrim, Elder Scrolls 6, Fa | H9AAnV59ddE 74 | Travis Oliphant: NumPy, SciPy, Anaconda, | gFEE3w7F0ww 74 | Kelsi Sheren: War, Artillery, PTSD, and | PbN3HzKkW4M ----------------- SELECT count(s.video_id) AS love_count, substr(v.video_title, 1, 40), s.video_id FROM Subtitles s, Videos v WHERE s.video_id = v.video_id AND s.video_id IN ( SELECT v.video_id FROM Videos v WHERE v.video_title LIKE "%Podcast%" AND v.video_title NOT LIKE "%Podcast Clips%" ) AND s.text LIKE "%love%" GROUP BY s.video_id ORDER BY love_count DESC LIMIT 10
- _aaed 3y agoPerhaps I'm wrong but how is this full text search? It's just using the LIKE operator
- mitesh07 3y agoThank You for share with us, Looking good to me
- expertentipp 3y agoI think I'll start to use exclusively CLI tools for discovering and downloading of YT content. The entire experience which starts from typing "youtube.com" in the address bar and pressing enter is obnoxiously unbearable.
- hackernewds 3y agoDownloading? Do you also not believe creators should be compensated for their content?
- galleywest200 3y agoI already block advertisements on the web, so I see none when on a Desktop web version of YouTube. But I do not use Sponsor Block so those creators still get to show me their ExpressVPN ads or whatever the flavor is today. Also I use Patreon.
- smcleod 3y agoI think they should be compensated, but I don't think Google should be.
- lawrencehook 3y agoself-promo, but you might find my extension helpful. https://lawrencehook.com/rys/ https://lawrencehook.com/rys/
- deleted 3y ago[deleted]
- DyslexicAtheist 3y agopython3 yt_fts.py download https://www.youtube.com/@PerspectiveArts/videos --channel-id UCUCN8V_pO0xOFKLL4XG1tshnw Downloading channel Saving vtt files to /tmp/user/1000/tmpm4xoskpo Traceback (most recent call last): File "/home/user/src/yt-fts/yt_fts.py", line 273, in <module> cli() File "/home/user/src/yt-fts/.env/lib/python3.11/site-packages/click/core.py", line 1130, in __call__ return self.main(*args, \*kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/user/src/yt-fts/.env/lib/python3.11/site-packages/click/core.py", line 1055, in main rv = self.invoke(ctx) ^^^^^^^^^^^^^^^^ File "/home/user/src/yt-fts/.env/lib/python3.11/site-packages/click/core.py", line 1657, in invoke return _process_result(sub_ctx.command.invoke(sub_ctx)) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/user/src/yt-fts/.env/lib/python3.11/site-packages/click/core.py", line 1404, in invoke return ctx.invoke(self.callback, \*ctx.params) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/user/src/yt-fts/.env/lib/python3.11/site-packages/click/core.py", line 760, in invoke return __callback(*args, \*kwargs) ^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/user/src/yt-fts/yt_fts.py", line 31, in download download_channel(channel_id) File "/home/user/src/yt-fts/yt_fts.py", line 82, in download_channel channel_name = get_channel_name(channel_id) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/user/src/yt-fts/yt_fts.py", line 191, in get_channel_name data = json.loads(script.string) ^^^^^^^^^^^^^ AttributeError: 'NoneType' object has no attribute 'string'
- vrglvrglvrgl 3y ago[dead]
- harlanji 3y agoI hate being the poo pooer who says that subtitles are available via the API and wish the tool went that route. I'm all for stuff being archivable with tools like youtube-dl, but I much prefer to see tools like this use the API despite its quotas because it goes beyond archiving a copy for reference. Tools that (ab)use scraping only justify anti-scraping efforts that journalists and the like use and escalate that arms race. I think one could still scrape a channel or two per day within API usage limits [1]--50 units per list, 200 units per download; quota 10,000 units per day. [1]: https://developers.google.com/youtube/v3/docs/captions/download https://developers.google.com/youtube/v3/docs/captions/downl...
- tonto 3y agoagree that using the API is likely the nicer route. you can also apply for a quota increase, I recently applied for youtube API quota increase to 100,000 units and it was approved for my app (https://cmdcolin.github.io/ytshuffle/ https://cmdcolin.github.io/ytshuffle/) I was concerned they wouldn't like that the app downloads so much data but it was approved without much question, they just wanted terms of service prominently displayed to end users
- progman32 3y agoAnother avenue for this is to use Tube Archivist, which takes the approach of locally mirroring videos and serving them up in a web interface, complete with comments, subtitles, and an index containing all the above. Definitely overkill if you just want to do a couple text searches though.
- teelelbrit 3y agoWhy not fold this into an LLM interface?
- almog 3y agoThis reminded me that Google Podcast search function clearly uses some kind of transcription index for its search. However, I wish they included at least partial matches as part of the search results (similar to how Google Books search works).
- dredmorbius 3y agoDoes this rely on an API key? There have been earlier tools which permitted command-line / terminal access to Youtube, one of the best being mps-youtube: <https://github.com/mps-youtube/mps-youtube https://github.com/mps-youtube/mps-youtube> That permitted searches for terms and channels (though not subtitle text within channel AFAIK), for music specifically, and compiliation of either temporary or saved playlists, with the option to play through a full selection of videos. It also offered either full-video or audio-only playback. Google killed it by throttling API-key access. Discussed previously: <https://news.ycombinator.com/item?id=32919545 https://news.ycombinator.com/item?id=32919545> <https://news.ycombinator.com/item?id=28571421 https://news.ycombinator.com/item?id=28571421>
- philsnow 3y agoIt seems to just use yt-dlp, which supports fetching subtitles.
- dredmorbius 3y agoThanks.
- bsnnkv 3y agoThis is great! I had made something like this for my own use, but it was way more complicated. It took a user account, downloaded all the liked videos, ran it through some model and vectorized it, and then you could use natural language search to describe a scene in your history of liked videos and it would show you the timestamp and the thumbnail (with the link to start watching from the timestamp). I ended up taking it offline because I didn't use it much, it was expensive and I couldn't see a path to monetization. This solution being posted however, is really elegant, in particular because it is very resource-efficient.
- bob_theslob646 3y agoIs it as effective as what you did? Did you need an API key for yours to work? What was your cost per video?
- simonw 3y agoIt looks like you're running searches using LIKE: https://github.com/NotJoeMartinez/yt-fts/blob/050981c0519a966593d2e2405bf2116997a00909/db_scripts.py#L81 https://github.com/NotJoeMartinez/yt-fts/blob/050981c0519a96... SQLite has a really power full-text search mechanism built in - FTS5. It can handle things like stemming and stop words and relevance ranking. My sqlite-utils Python library includes helper methods for setting that up: https://sqlite-utils.datasette.io/en/stable/python-api.html#full-text-search https://sqlite-utils.datasette.io/en/stable/python-api.html#...
- notjoemartinez 3y agoThank you! I was able to integrate this into the project[1]. I'm also looking into using your openai-to-sqlite[2] library for semantic search. [1]https://github.com/NotJoeMartinez/yt-fts/pull/25 https://github.com/NotJoeMartinez/yt-fts/pull/25 [2]https://github.com/simonw/openai-to-sqlite https://github.com/simonw/openai-to-sqlite
- lfconsult 3y agoYou're right. Thanks for sharing the link to your full-text search helper, really neat.
- vram22 3y agoApropos of the post: sometime back I had wondered if it would be possible to search videos for content by text keywords, but not to match text occurring in the video title, chapter titles or comments. Instead, if an app or library could somehow search for the given keywords by matching them with the spoken words in the video. That would be a potentially cool and useful feature. I realize this may be technically impossible [1] or very difficult, but thought of mentioning the idea here. [1] On further thought, speech recognition (as seen in mobile phones at least) has progressed to quite a good level (speaking as a layman for this part), so maybe the idea is not wholly infeasible. If an app or lib could somehow "internally" play the video, and speech-recognize the spoken words into text, then the problem would reduce to normal text pattern matching. Putting the idea out here to see what devs think of it.
- vram22 3y ago>somehow "internally" play the video, Analogous to how headless browsers are used for automated testing or scraping of web apps.
- bob_theslob646 3y ago>yt-fts is a simple python script that uses yt-dlp to scrape all of a youtube channels subtitles and load them into an sqlite database that is searchable from the command line "Scrape all of YouTube subtitles" So if that data is not good...then how is this useful. The captions in YouTube have been pretty bad. Are these the same thing? I'll test it out.
- theo_champion 3y agoI have actually been working on a full blow full-text search engine for youtube by transcripts as a web application: https://clipbase.xyz https://clipbase.xyz I'd love to know what y'all think!
- harryvederci 3y agoNice. Possible (?) improvement: provide a bit of context with each clip. For example: 2 seconds before the searched bit, and 2 seconds after.