19 ms·
Yark: Advanced and easy YouTube archiver now stable
- Owez 4y agoI've been working on polishing my YouTube archiver project for the last while and I've finally released a solid version of it, it has an offline web viewer/visualiser for archived channels and it's managed using a cli. Most importantly, its easy to use :)
- theandrewbailey 4y agoCan you edit your submission to add a "Show HN:" before the title? Like these: https://news.ycombinator.com/show https://news.ycombinator.com/show
- Owez 4y agoWill do
- bityard 4y agoI have a cron job that uses yt-dlp to download just the audio tracks of any videos that I have saved to a public playlist, can this be a replacement for that?
- na4ma4 4y agoI like it, it's much better than what I've used previously. I made a docker container to run it (https://github.com/na4ma4/docker-yark https://github.com/na4ma4/docker-yark), when I get time I'll do a PR if you're interested so it isn't a separate project. (I'll also fix it so the host is a command line argument not just changing the binding from 127.0.0.1 to 0.0.0.0)
- ekianjo 4y agoIs there a longer documentation anywhere? It's not clear from the README if you add a whole channel when you create an archive, or if you can add videos one by one to archive them?
- googlryas 4y agoGreat project! I have a playlist that I use to keep track of videos that my young kid likes to watch, but wanted to get away from yt because of the ads, related videos, comments, autoplaying next vids, etc. So I just followed the simple instructions and bam, I have all the videos on my computer now with a UI that is even easier to use for him. Very nice, thank you!
- alexvoda 4y agoFrom the readme, it is not clear to me what is meant by metadata. Will this archive subtitles? Will it archive comments? If is can, can comments be updated? Also, from the readme it looks like all metadata is kept centralized instead of each video having its own metadata file. Are containers other than mp4 supported?
- newsclues 4y agoWhy not use https://youtube-dl.org https://youtube-dl.org ?
- hbn 4y agoytdl is one of its dependencies You clearly didn't even skim the readme to see what this does
- naavis 4y agoTo be fair, the readme does not mention ytdl.
- blowski 4y agoFrom the HN guidelines: > Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
- hbn 4y agoThere's a difference between not reading an entire article or missing reading a paragraph versus not even clicking the link or skimming a couple screenshots to figure out what you're commenting on.
- alwayslikethis 4y agoyoutube-dl basically became unusuable for a while now. It gets only tens of KB/s on my 1gbps connection. yt-dlp is a more maintained alternative, I think. It seems to get great speeds the last time I checked.
- iforgotpassword 4y agoDid you write your own YouTube scraper, which would be quite a task, or is this based on something like ytdl? Might be worth mentioning in the readme.
- amarshall 4y agoIndeed it appears to use yt-dlp https://github.com/Owez/yark/blob/676074ee3d9e379d15e52ffe2e0fa5436b2a31fb/pyproject.toml#L21 https://github.com/Owez/yark/blob/676074ee3d9e379d15e52ffe2e...
- chii 4y ago> Might be worth mentioning in the readme. i mean, it's not really important to a user which ever library it uses to scrape youtube - i suppose it's important if you want to contribute/develop it.
- toomuchtodo 4y agoIf you want to provide an option to upload artifacts to the Internet Archive, you could crib off of https://github.com/bibanon/tubeup https://github.com/bibanon/tubeup It too relies on yt-dlp for extraction. Importantly, pay close attention to what artifacts are uploaded to an item created, and what metadata is set as part of the upload process.
- 2kwatts 4y agoThis is really slick. I was figuring it was just a simple wrapper for yt-dlp which scraped some additional things (comments, views, etc) but you went above and beyond with the web interface. Nice job!
- prometheus76 4y agoyt-dlp and a batch file that runs via Task Scheduler has been doing this for me for a couple of years now. I also grab the captions and throw that into a database so that I can search transcripts for a clip that I can remember but can't remember which video it's in. It was a fun weekend project.
- 2OEH8eoCRo0 4y agoHow do you deal with file numbering? I prefer the file prefixed with a number that indicates "air date". 01 being the first uploaded video. The default is by index and the top of the channel or playlist is number 01 which is the most recent.
- prometheus76 4y agoI just use the publish date in the format of YYYY-MM-DD at the beginning of the filename so that they sort properly.
- NegativeLatency 4y agoI run mine with cron and it puts files in a special folder for plex: https://github.com/nburns/utilities/blob/master/youtube.fish https://github.com/nburns/utilities/blob/master/youtube.fish Pulls from my watch later playlist which is quite handy
- jamessb 4y agoIt looks like this depends on a "./add-video.py" script that isn't in the repository.
- seligman99 4y agoLong ago I had my podcast downloader keep all files it downloads and recently I've been using OpenAI's Whisper to go through and create transcripts of the 8000 or so hours of data I have downloaded over the years. It's very cool to be able to search through and remind myself of something I heard once. Not exactly life changing, but still, nice to be able to quickly drill down and find audio for something when a curiosity strikes me.
- eats_indigo 4y agoBit of a noob question here. What's an archiver for? Is it a library for things you've watched and want to store outside of youtube? Or is this for storing content you've created / managing your own portfolio of content?
- pessimizer 4y agoThere are coded hints in the link, like: > Yark lets you continuously archive all videos and metadata for YouTube channels. You can also view your archive as a seamless offline website
- eats_indigo 4y agoSnarky and not answering the question, well done.
- iforgotpassword 4y agoPersonally, for me it's archiving. In case I want to go back to it. Videos just keep disappearing from YouTube because channels get deleted by YouTube, by their owners, videos get copystriked, geoblocked, privated, and so on. As I'm lazy as f### I didn't create anything as sophisticated as OP, but a simple 10 line PHP Script on my home server that just pretends being Kodi enough to fool yatse (android remote for Kodi). So every time I watch a video on YouTube (on my phone) that I want to keep I tap "share" and then "play on Kodi", my php script gets the video url from the post data and launches youtube-dl. It sucks because I never get feedback if it worked and when it's finished, but I log all the URLs and at some point in the future I'll eventually add a cronjob that checks the list and sends reports and whatnot. Some day.
- eats_indigo 4y agoGreat, thanks for the insight!
- cocacola1 4y agoI think it’s the latter. I’ve no issue with most things being one-and-done. But some channels have phenomenal content that I’d like to keep for the long term. Something might happen to their channel that makes it difficult to get, so I’ll regularly update my downloads with new videos, pictures, etc. This applies to ripping, too. Funimation removed Drifters years ago, but I’ll always have a copy of it because I ripped it. Of course, I need to store it so it still costs money. But I can be content that I have the content.
- amelius 4y agoHow cool would it be if everyone had IPFS running in their browser, and everyone dedicated some time to filling it with a backup of the internet, including YouTube.
- KMnO4 4y agoI did some napkin math. If 1 billion people each backed up 10gb, we’d almost have enough to store a copy of YouTube with zero data redundancy. Google is massive.
- CamelCaseName 4y agoIt's unbelievable that YouTube was as free as it was for as long as it was. We got too great a deal for so long that we many people can't see things any other way.
- yesco 4y agoWas probably easier when the videos had a time limit and didn't support 4K (or 1080p even).
- judge2020 4y agoFor reference, 1080p was in 2009: https://blog.youtube/news-and-events/1080p-hd-comes-to-youtube/ https://blog.youtube/news-and-events/1080p-hd-comes-to-youtu... while ads came out much before then: https://blog.youtube/news-and-events/partner-program-expands/ https://blog.youtube/news-and-events/partner-program-expands...
- wintermutestwin 4y agoIt is not "free" at all if you are paying with your data. My data privacy is worth way more than the cost of streaming some video with crap discovery.
- rchaud 4y agoIt's unbelievable that Wikipedia is free and survives on donations. Youtube sells ads and is subsidized by one of the biggest ad companies in the world that happens to have a lot of cheap cloud storage available.
- causality0 4y agoDoes this have the ability to bet set to "grab highest available resolution" instead of specifying one? A lot of the material I'd like to archive has material from well before and after Youtube started supporting HD resolutions.
- 2OEH8eoCRo0 4y agoWhat if the highest available resolution does not have an audio stream?
- TonyTrapp 4y agoOnly some formats which I guess are to be used with older browsers contain both video and audio. In general, these days video and audio are delivered through separate streams on YouTube.
- NegativeLatency 4y agograb the audio stream from something else and stitch them together with ffmpeg (like youtube-dl and others do)
- rollcat 4y agoNormally in modern adaptive streaming, every video variant is muxed into a separate stream without audio, and different audio variants are muxed into their own individual streams.
- jxramos 4y agowow, I wonder if that's why it always feels so frequent an experience of mine where the audio and video feel subtly out of sync with one another. It's very minute but detectable. Feels like that experience has increased in the last 6 months or so.
- crazygringo 4y agoYouTube has been that way (separate streams) for a long time, definitely not anything new in the last 6 months. And they reassemble to be indistinguishable from the original combined stream, so that's not going to be the cause. There are plenty of causes of delayed audio, however. Bluetooth is a big one, if your device and software aren't properly compensating for the Bluetooth transmission delay.
- swyx 4y agodoes anyone have recs on how to run this on a continuous basis in the cloud? this obviously will take a lot more storage than like a normal heroku setup (not that I would use heroku). should i use Railway or Render or is that overkill compared to something else? gasp can i run it as a github action???
- pablo24602 4y agoI'd add some more flags/options for downloading specific videos or updating the library of downloaded content. I don't really want to download all the videos from one specific channel- instead I want to download the last 10 videos, for example.
- Owez 4y agoYou can do yark refresh [name] --videos=10
- nomilk 4y agoOccasionally my personal/literature/academic/tech notes link to a youtube source, but when clicking on it I find the video has vanished with no way of knowing what it was or even what it was called (it's sometimes impossible to track down an identical/replacement source). I lost many valuable references that way. Wayback machine solves for webpages, but nothing I'm aware of (short of youtube-dl-ing the video yourself and storing it somewhere openly, probably at risk of various infringements) solves this. Quite a lot of hassle for something rather simple. It would be great to be able to immortalise them on a per-video basis, so if it's important enough, we can be sure that references made to the content will still be there in the future when needed.
- 0cf8612b2e1e 4y agoI do not see how that becomes possible without the Internet archive effectively mirroring a large percentage of YouTube. I recall at one point, IA wanted to archive just the video metadata and realized even that would be technically challenging.
- RockRobotRock 4y agoIt would be nice to prioritize videos which are deemed at "high risk" of being deleted, with bayes statistics or machine learning or something like that.
- judge2020 4y agoIA does seems to archive YT video content, at least last time I tried to watch a deleted but popular video.
- sodality2 4y agoOnly if a user chooses to submit the URL - I think parent comment is referring to an organized attempt by IA to archive a significant portion of YT's videos.
- svnpenn 4y agoAny video you care about, you need to make it your own responsibility to backup the metadata and/or streams. If you're lucky you can internet search the video ID to get the metadata, even after deletion.
- salutonmundo 4y ago[flagged]
- HEHENE 4y agoI'm a big fan of the historical information that Yark shows. Arrimus 3D recently replacing a large chunk of their 3D modeling tutorials with religious content was a pretty big lightbulb moment for me that so much of the content I rely on - not just for the initial learning of a new skill, but as a continual reference when I forget something - is so fragile. I immediately bought a NAS and began backing up everything that I gleam even the tiniest bit of learning from using a similar project, TubeArchivist[0]. Projects like this are really important for maintaining all of the great knowledge on the web. [0] https://github.com/tubearchivist/tubearchivist https://github.com/tubearchivist/tubearchivist
- bakkoting 4y agoSeconding tubearchivist. One of the killer features IMO is the browser extension [0], which adds a button on every video to send that video to the server to archive. [0] https://github.com/tubearchivist/browser-extension https://github.com/tubearchivist/browser-extension
- jxramos 4y agothat's quite a pivot. Can you link to an example of a typical before and after video of that content creator? I'm curious what the connection is.
- dotancohen 4y agoYou might consider making torrents of those reference videos. It will help other people find and use them, and provide robustness to your collection.
- wahnfrieden 4y agoAnyone know if Apple bans this kind of lib from use in the app store?
- pvg 4y agoThere have been iOS video players with such functionality built-in whose authors have had to remove the functionality at Apple's request.
- wahnfrieden 4y agoThank you. I have a couple like that and didn’t know how confident to be with replicating It seems Apple is fine with adblocking by default for web content, but inconsistently as with youtube. Hard to predict what's risky to invest dev time into
- pvg 4y agoI'm not an expert on Apple's rules but my understanding is the download thing is circumvention of the terms under which Youtube licenses/re-licenses content so it's treated fairly strictly.
- wahnfrieden 4y agoThen why do they allow other ad-blocking products or browsers that block ads by default? I don't understand the consistency of their position except that YouTube has a powerful lobby
- alwayslikethis 4y agoApparently they don't like when your app does something in a way that breaks ToS or EULA, even if they are not worth the pixels they are written on. Apple being the trillion dollar company it is, is naturally averse to the legal risk posed by having a ToS-breaking app on their store. This is why having a single company controlling your entire device is bad -- they do what's best for them, rather than whatever you might happen to want.
- j1elo 4y agoAfrer reading the description, this project seems to be solely focused on downloading all of a specific channel's videos. I've been taking my first steps at having a home server, and one of the things I'd love to do with it is having an archive of the videos that I have saved in my private playlists on YouTube. In my mind, the service would periodically check all my playlists, compare with what exists locally, and download any missing video. Maybe even with a nice web UI so it's easier to visually configure and use. Does such a service already exist so I can self-host it?
- aquova 4y agoI haven't used it personally, but Tube Archivist might be what you're looking for. https://www.tubearchivist.com/ https://www.tubearchivist.com/
- erinnh 4y agoLook at tubearchivist. It can do what you want. You can subscribe to playlists, as well as automatically update and download videos.
- Fang_ 4y agoNice web ui aside, if I'm not mistaken youtube-dl already supports this kind of usage. You can `youtube-dl --download-archive archive.txt https://youtu.be/your-playlist https://youtu.be/your-playlist` and it'll keep track in the archive.txt of everything it's already downloaded. Supplement with authentication options as necessary, set up a cronjob, done.
- itsyaboi 4y agoIt's even simpler than that, just give youtube-dl the channel name and it will download all videos skipping any that already exists in your current directory.
- _0ffh 4y agoI think youtube-dl as well as yt-dlp can both download playlists. You can create a script to download all your playlist and make it a cron job. Videos that do already exist in the target folder will be skipped automatically.
- superkuh 4y agoAny tool that tries to use youtube in a non-standard way that is stable is soon obsolete.
- walrus01 4y agoDoes this use youtube-dl / yt-dlp underneath for the retrieval of each video's URL and highest-quality video/audio format, and merge with ffmpeg?
- iKlsR 4y agoThis is pretty cool, I did a similar personal project that I describe here https://news.ycombinator.com/item?id=28480790 https://news.ycombinator.com/item?id=28480790. The only historical thing I log however is if a video was removed/reuploaded.
- 6510 4y agoAn idea I had some time ago: BitTorrent magnet uri's or hashes are suppose to be made from torrents but they really just point at a torrent. One could make torrents for each video, take the youtube url v param and make a hash from that and point it ("erroneously") at the torrent. That way, provided it was downloaded before, anyone who has the url can obtain the video. The idea needs one more trick to validate the download. I suppose one could compare a chunk of yt to the same piece downloaded over BitTorrent but perhaps there are better ideas to be had. Eventually, with tit for tad one could swap one chunk of one video with a chunk from a different video on the same channel.
- qwerty456127 4y agoBy the way, does anybody know a way to include the YouTube time stamps as chapters in a downloaded video file?
- xk3 4y agoIn yt-dlp is --embed-metadata (and --embed-chapters if you only want chapters)
- RichardVeems 4y agoHere's a similar project we've been working on, it's running 24/7 https://github.com/VeemsHQ/yt-channels-archive https://github.com/VeemsHQ/yt-channels-archive This just provides the latest HQ of the videos + thumbs & metadata, no historic information such as changes in video titles.
- cactusplant7374 4y agoI have a youtube archiver script I'm using right now that I pulled from a thread on the data hoarder subreddit. My main issue is that I need the downloader to remove emojis from the filename because I sometimes sync them to Dropbox. Can this project do that?
- xk3 4y agoyt-dlp has a `--restrict-filenames` flag
- Mikescher 4y agoA while ago I did something similar. I'm already downloading various playlists via yt-dlp and wrote a web interface [1] to view/play/search them. My biggest annoyance at the time was importing my existing videos (and converting them to a streamable format, generating thumbnails and hover-previews etc). Do you have any plans of allowing to import existing yt-dlp folder (in the standard layout with a bunch of mkv files, the info.json, the subtitles etc). Because my current archive contains a lot of already-deleted videos :( [1] https://github.com/Mikescher/youtube-dl-viewer https://github.com/Mikescher/youtube-dl-viewer
- Owez 4y agoI do eventually, might be in v1.3 next month or sadly v1.4 depending on how much I've got on my plate :)
- mwest 4y agoIf you're interested in this kind of thing, you may also want to check out: - the Distributed YouTube Archive Discord: https://discord.com/invite/PQqks7eSKc https://discord.com/invite/PQqks7eSKc - ArchiveTeam also do a significant amount of YT archiving: https://wiki.archiveteam.org/index.php/YouTube https://wiki.archiveteam.org/index.php/YouTube - a similar, but private effort: https://reddit.com/r/Archivists/comments/5uvfpw/youtube_archive_still_running_already_40k_videos/ https://reddit.com/r/Archivists/comments/5uvfpw/youtube_arch...
- xk3 4y agoI also archive many playlists with some code I wrote but I don't use a GUI. https://github.com/chapmanjacobd/library/blob/f778e22bf80c58f9689f88b3c1719225ec7e626b/xklb/tube_backend.py#L279 https://github.com/chapmanjacobd/library/blob/f778e22bf80c58... My focus is on error handling and trying to differentiate between unrecoverable errors and recoverable ones (try different proxy) but there's still a lot of work to be done. Also look into https://github.com/swolegoal/squid-dl https://github.com/swolegoal/squid-dl
- mwest 4y agoWow, I love the "daily tabs" concept. I'll install this and give it a go. Thanks! "The use-case of tabs are websites that you know are going to change: subreddits, games, or tools that you want to use for a few minutes daily, weekly, monthly, quarterly, or yearly."