14 ms·
YouTube-dl has an interpreter for a subset of JavaScript in 870 lines of Python
- Too 4y agoThey must have been inspired by this PyCon presentation, where David Beazley live codes a fully working webassembly interpreter, in under one hour. https://youtu.be/VUT386_GKI8 https://youtu.be/VUT386_GKI8
- homarp 4y agothe tests for it: https://github.com/ytdl-org/youtube-dl/blob/master/test/test_jsinterp.py https://github.com/ytdl-org/youtube-dl/blob/master/test/test...
- M30 4y agoHow should a programming noob interpret this? Be impressed at what was achieved here? Be concerned about security implications using the tool? Something else entirely?
- smcl 4y agoAll of the above, really.
- lolinder 4y agoIt's an extremely tiny subset of JS—as an example, the only object that can be instantiated is Date. Anything other than "Date" after "new" throws an exception. It's definitely neat, but not especially useful outside of the confines of its current application, and the security concerns of such a tiny subset will be minimal.
- petters 4y ago> Anything other than "Date" after "new" throws an exception It's even very sensitive to white space.
- tenebrisalietum 4y ago> How should a programming noob interpret this? The browser is client-facing and everything there is possible to reverse engineer and figure out. So if you design a web-based application, and are depending on client-side Javascript for any security or distribution enforcement, it can be helpful, but can ultimately be unwound and cracked even if obfuscated, etc. > Be impressed at what was achieved here? Yes. Try to download a YouTube video with out it or an online service which is probably using it internally.
- Supermancho 4y agoYoutube-dl is impressive. This particular hack is not.
- pwdisswordfish9 4y agoyoutube-dl as a whole is not particularly impressive either. It’s a big pile of unresolved technical debt, of hacks-upon-hacks and quick-and-dirty temporary solutions just like this one staying there for years.
- Test0129 4y ago> How should a programming noob interpret this? Usually in a virtual machine.
- rkangel 4y agoThis is the compiler writer equivalent of parsing HTML with regex: It is technically wrong - it isn't a sufficiently rich and powerful approach to handle all JS (HTML) that you might throw at it. It'll work for a while until it eventually barfs when you least expect it. EXCEPT that if the inputs you are giving it come from some understood source(s) that aren't likely to change, then a simpler approach to the "all singing all dancing" correct may be appropriate and justified. E.g. because it might be easier to write, easier to maintain and/or less attack surface etc.
- pwdisswordfish9 4y ago> some understood source(s) that aren't likely to change Does that apply to YouTube? Or any of the other hundreds of supported sites?
- rkangel 4y agoPresumably because it gets tested with those sites and the JS doesn't change that much it can be fixed or adjusted as required.
- bjt2n3904 4y agoThe goal of youtube-dl is to download a video off of YouTube for offline storage. This isn't something YouTube particularly enjoys. They would rather you keep coming back -- every visit is more ad revenue for them. If you have an offline copy, you don't need to visit YouTube anymore. YouTube has an incentive, therefore, to make it more difficult to download (or "scrape") their content. I'm not particularly sure of the specific details, but apparently YouTube has added JavaScript (a programming language that executes in the browser) as a hurdle to jump over. A simple python script doesn't have enough brains to execute JavaScript, only enough to realize that it exists. (Clearly, youtube-dl is sophistication enough to have jumped over it.) These are the conclusions I come to, having written software for about a decade. 1) Once you give information to someone, be it text, pictures, sound, or video -- they will do whatever they want with it, and you have no control. Oh, yes -- it may be illegal. Maybe unethical. But the fact of the matter is you do not have control over information once it leaves your hands. 2) Adding hurdles to make it harder to access the information does little to stop someone who is dedicated to accessing it. 3) Implementing a subset of JavaScript in such an elegant and tiny manner is quite impressive. How you interpret these facts depends on your worldviews. If you are a media and content creator, you will view these facts differently than a politician, and a teenager. As an engineer and amateur philosopher, I certainly support the rights of content creators to be paid for their work. And yet, I fear that more and more, content creators want to lease me a right to listen their music, instead of own a copy of it. I used to own CDs, DVDs, movies, and books. What happens if Amazon or YouTube decides to not serve me anymore? Anything I've "purchased" from them, I lose access to. Further more, if I create a song, I used to be able to burn copies of CDs and distribute it on the street corners. Now, you have to sign up to stream on Spotify. This is a double edged sword -- I get a wide audience, but Spotify will do whatever they want with me. This troubles me.
- chlorion 4y agoThe "interpreter" in the youtube-dl source is probably safe from a security standpoint. yt-dlp seems to support running javascript in a full javascript interpreter/headless browser called phantomjs though. Running javascript in a full interpreter like this is a lot more scary from a security standpoint. I am not sure whether phantomjs sandboxes the javascript evaluation from the rest of the system, and if it does, whether the sandbox actually works properly at all. It looks like the project is not being maintained which is another bad sign. Big projects with lots of manpower behind them such as chromium have trouble keeping javascript evaluation safe, so I would really suggest not trusting phantomjs on untrusted input.
- Tao3300 4y agoIn the face of weird shit like this, I give you the permission to go with your gut.
- anony23 4y agoWhat purpose does it serve?
- deleted 4y ago[deleted]
- throwaway0984 4y agoIIRC it's used to extract/generate the signatures needed for YouTube media URLs
- oynqr 4y agoYou need to run some obscured JS to get decent download speeds from Youtube. Something along the lines of PoW.
- db48x 4y agoIt’s not like proof of work at all. It’s just a challenge and response; youtube includes a random number in the webpage for each video, and expects to see a request parameter with a particular value calculated from that random number when you request the video. If you don’t do the arithmetic it throttles you to 50kb/s. Since the calculation of the response is done in JS, and they occasionally change the formula, some download programs are moving towards running the JS rather than trying to keep up with the changes. It’s really just bullshit to make people’s lives harder.
- xg15 4y agoNext step will probably be moving the calculation to webassembly or requiring the script to fetch the result via websocket or webrtc...
- mistrial9 4y ago.. pirate determination is a thing to behold, as is crazed-repetitive digital grabs.. Its not a fair or accurate characterization to dismiss it as "making people's lives harder" .. it is remarkable that the Debian distros now include ytdl; lets do what is reasonable to make it continue
- rcarmo 4y agoAwesome. Even if it's likely incomplete, it might come in really handy for some scraping I need to do...
- haunter 4y agoThe same in yt-dlp https://github.com/yt-dlp/yt-dlp/blob/master/yt_dlp/jsinterp.py https://github.com/yt-dlp/yt-dlp/blob/master/yt_dlp/jsinterp... Interesting to see the diffcheck between the two https://www.diffchecker.com/8EJGN27K https://www.diffchecker.com/8EJGN27K
- cheschire 4y agoIs yt-dlp's implementation being better the reason why I have fewer throttling issues than with youtube-dl?
- deleted 4y ago[deleted]
- LeoPanthera 4y agoMaybe this isn't true anymore, but for a while they would hit different APIs. yt-dlp was using the Android YouTube API because it had no throttling.
- sylware 4y agoNowadays "javascript" refers to the scriptable, grotesquely and absurdely complex and massive web engines, aka google financed blink and geeko, then apple financed webkit, that with their SDK. The currently obfuscated javascript media players will try to break yt-dlp by leveraging the complexity and size of those scripted web engines. They will make them out of reach to small teamns or individuals and it is even "better", it will force ppl to use apple or google web engine, killing any attempt to provide a real alternative. A standalone javascript interpreter is actually some work, but seems to stay in the "reasonable" realm: look at quickjs from M. Bellard and friends (the guy who created qemu, ffmpeg, tinycc, etc): plain and simple C (no need of a c++ compiler), doing the job more that well enough. That's why noscript/basic (x)html is so much important.
- dtx1 4y ago> but seems to stay in the "reasonable" realm > M. Bellard and friends Chose one, that dude is a wizard wielding c like a brain surgeon wields a scalpel.
- olliej 4y agoYeah I agree with almost all of this - the massive size and complexity of commercial engines makes it seem like JS the language must also be complex. I also agree with the idea that these sites will probably be able to/want to create JS that breaks these small/lightweight engines requiring constant work :-/ This final point I disagree with entirely. You can't point to Bellard doing something as evidence that it's reasonable. This is a guy that wrote a program that generated a TV signal via a VGA card. :D
- randyrand 4y agoChrome and Safari both have open source JS engines…
- userbinator 4y agoThat's beside the point. Open-source is not useful to the smaller players if it is too complex to comprehend and constantly churned.
- olliej 4y agoThis is super cool. Some of the stuff is kind of questionable to me in the sense that I could believe you could probably make some kind of sufficiently wonky JS that this would do the "wrong" thing. But it's super cool that they are able to do this as I think it shows that claims of JS complexity based on the size of JS engines is overlooking just how much of that size/complexity comes from the "make it fast" drive vs. what the language requires. Here you have a <1000LoC implementation of the core of the JS language, removed from things like regex engines, GCs, etc. Mad props to them for even attempting it as well - it simply would not have ever occurred to me to say "let's just write a small JS engine" and I would have spent stupid amounts of time attempting to use JSC* from python instead. [* JSC appears to be the only JS engine with a pure C API, and the API and ABI are stable so on iOS/macOS at least you can just use the system one which reduces binary size+build annoyance. The downside is that C is terrible, and C++ (differently terrible? :D) APIs make for much more pleasant interfaces to the VM - constructors+destructors mean that you get automatic lifetime management so handles to objects aren't miserable, you can have templates that allow your API to provide handles that have real type information. JSC only has JSValueRef and JSObjectRef, and as a JSObjectRef is a JSValueRef it's actually just a typedef to const JSValueRef :D OTOH other hand I do thing JSC's partially conservative GC is better for stack/temporary variables is superior to Handles for the most part, but it's also absolutely necessary to have an API that isn't absolutely wretched. The real problem with JSC's API is that it has not got any love for many many many .... many years so it doesn't have any way to handle or interact with many modern features without some kludgy wrappers where you push your API objects into JS and have the JS code wrap them up. The API objects are also super slow, as they basically get treated as "oh ffs" objects that obey no rules. I really do wish it would get updated to something more pleasant and really usable.]
- esprehn 4y agoThis doesn't actually implement any of the JS language though, it just reuses all of python's semantics and hard coded a tiny list of ex. String methods I also assume you mean mainstream JS engine, but Duktape, JerryScript and QuickJS are all C APIs. They probably could have used ex. https://github.com/PetterS/quickjs https://github.com/PetterS/quickjs instead of the hacks in the OP linked file.
- lolinder 4y agoTo be clear, this is an extremely tiny subset of JS. It looks like they only implemented the features needed to run a very specific function. For example, the only symbol allowed after "new" is "Date", everything else throws an exception. It's still fun that it's there, but it's not as big a deal as it sounds from the tweet.
- krab 4y agoIt will only grow - as new scripts will need to be interpreted, new features will be added.
- lolinder 4y agoI would be horrified if this grew much further. It's perfectly fine for its current scope, but the architecture would not scale at all to a full interpreter without essentially starting from scratch.
- kelnos 4y agoYeah, at some point you have to question if it's worth spending time maintaining a quirky, error-prone, ever-growing mini-JS interpreter, or just adding a dependency on v8 or node or something. And then you don't have to worry about supporting new scripts, as they'll just always work.
- jchw 4y agoIf you were going to use a C library, the most logical is QuickJS since it has Python bindings, is small, perfectly fast enough for the kind of needs yt-dl has, and it has excellent coverage of the standard and passes conformance tests. That said I think a decent Python-native JS interpreter isn't that bad of an idea, it definitely needs a separate project and a more sophisticated architecture but it's an attainable goal.
- d0mine 4y agoBeing pure Python has an advantage where it can be run e.g., in Pythonista 3 for iOS (which allows to implement various functions such as: download just audio, send link to be viewed on a separate device).
- deleted 4y ago[deleted]
- esprehn 4y agoThis isn't really JS, it's a purpose built evaluator that's only for evaluating a particular script on YouTube, assuming a huge list of things are true about how YouTube JS is written. Ex. Its got a hard coded list of methods for String, and it doesn't respect prototypes. It only supports creating Date instances, and won't work if you override the global Date. It parses with regexes and implements all operators with python's operator module (which is the wrong type semantics) etc. Nearly none of the semantics of JS are implemented. It's sort of the sandwich categorization problem: If I write a C# "interpreter" in perl thats only 200 lines and just handles string.Join, string.Concat and Console.WriteLine, and it doesn't actually try to implement C# syntax or semantics at all and just uses perl semantics for those operations is it actually C#? :P I say "not a sandwich".
- tra3 4y agoIt’s quacks like a duck at midnight, but it’s actually a frog?
- Test0129 4y agoThis really isn't fair. Just because it doesn't faithfully implement whatever standard Javascript is on doesn't mean it isn't an interpreter. All an interpreter is is something that executes a script directly rather than requiring compilation. It is a defacto interpreter for a subset of javascript. Nothing more, nothing less. The title could be more clear, however.
- baobabKoodaa 4y agoThere's a huge difference between an interpreter for "JavaScript" and an interpreter for a "subset of JavaScript".
- Test0129 4y agoMaking a pedantic argument on what constitutes an interpreter is silly. The title is bad. It is an interpreter. I'll continue to eat downvotes on this because of the pedantry of HN.
- kristopolous 4y agoTo understand why, I have a far simpler tool that focuses on a subset of sites (adult content video aggregators) https://github.com/kristopolous/tube-get https://github.com/kristopolous/tube-get It too deals with this problem but does so in a way that'd be easy to maliciously sabotage Look right about here https://github.com/kristopolous/tube-get/blob/master/tube-get.py#L111 https://github.com/kristopolous/tube-get/blob/master/tube-ge... As to why this program exists, this was originally written between about 2010-2015 or so technically predates the yt-* ecosystem. The tool still works fine and it's not a strict subset of yt-dlp or YouTube-dl because being a different approach, although it's overall site coverage is smaller, I've had it be a "second try" system when yt-* fails and it comes up with success maybe about half the time
- pabs3 4y agoWould you mind switching to subprocess with shell=False? os.popen is obsolete and insecure because it passes the command through the shell. PS: I found it quite easy to contribute to yt-dlp and the reviewers are ultra-helpful and kind, you might want to migrate all of your extractors there.
- kristopolous 4y ago1. It's ancient code but sure 2. They're fundamentally not compatible approaches. This is worthless to them
- jraph 4y agoI do wonder why YouTube does not try harder to make it difficult to do this computation meant to prove you are a legit YouTube web client. Providing an easy-to-find, simple JS function interpretable with 900 lines of Python is like they don't try at all. They might as well do nothing. Or is their goal just to make youtube-dl not 100% reliable? Or to be able to say "look, you are running our code in a way we did not intend, you can't do this because you are breaking the EULA"?
- Arnavion 4y agoThey do make it harder from time to time. In fact yt-dlp's interpreter has been broken for a month or so now and the devs finally gave up and told users to just install PhantomJS (which itself hasn't been updated since 2016 and probably has bugs / vulns of its own, but whatever). https://github.com/yt-dlp/yt-dlp/issues/4635#issuecomment-1235595110 https://github.com/yt-dlp/yt-dlp/issues/4635#issuecomment-12...
- whywhywhywhy 4y agoI mean if this is the direction it’s heading it makes more sense to port yt-dlp to node. It’s already dependent on a scripting language, it may as well be the one YouTube speaks.
- Cthulhu_ 4y agoI'm guessing the amount of people using it is low enough to not bother with mitigation. Then again, there's a LOT of YT videos that take clips from other videos (which in most cases falls under fair use), which I can imagine would use this tool.
- zuminator 4y agoI'd guess that their efforts to make it harder are limited by the fact that they want YouTube to be able to play on thousands of different low powered set top boxes and cheap phones. So whatever obfuscated code they use has to be simple enough to be run and periodically updated by all these different devices, and that same simplicity makes it emulable.
- deleted 4y ago[deleted]
- lewisl9029 4y agoAnother really cool JS dialect I recently learned about is njs from the nginx team: https://github.com/nginx/njs https://github.com/nginx/njs This video goes into some of the design and tradeoffs: https://www.youtube.com/watch?v=Jc_L6UffFOs https://www.youtube.com/watch?v=Jc_L6UffFOs TL;DW: they optimized for fast creation/destruction of low-footprint VMs with no JIT or garbage collection.
- Uptrenda 4y agoAnyone who has ever pulled a website from a script knows the pain that is Javascript. Normally you want to just get some text and work out the API actions but a lot of sites use horribly obfuscated Javascript -- either because that's what modern web development is (lolz) -- or because its part of their 'security.' That means if you want to write browser-based bots properly -- you ought to use a browser. There are special browsers that run 'headlessly' or are designed mostly for bot use. Like https://www.selenium.dev/ https://www.selenium.dev/ which plugs into a few different 'browser engines.' But now you have another problem. Your simple script goes from being small, simple, self-contained, and elegant gem, to requiring a full browser, specialized drivers, and/or daemons running just to work. If you're using something like Python you just frankly don't have very good packaging. So it's hard to string together all that into a solution and have it magically work for everyone. What YouTube-dl have done is good engineering. Even though it's not a full JS interpreter: they've kept their software lean, self-contained, and easier to use.
- eurasiantiger 4y agoJust npm install puppeteer.
- lolinder 4y agoPuppeteer is cool, but it's exactly what OP is warning against: it's a full browser that is downloaded and run through npm. It's remarkably well packaged, but still far more error prone than a simple HTTP request, and far more likely to break on its own just with the passage of time.
- eurasiantiger 4y agoYes, but: ”Your simple script goes from being small, simple, self-contained, and elegant gem, to requiring a full browser, specialized drivers, and/or daemons running just to work” Complex problems cannot be solved by simple scripts, but they can be abstracted away to vendor libraries when/if they are well maintained, such as in this case. While it can break with time, at least someone else fixes it for you.
- mdaniel 4y agoI was expecting this to be about Duktape <https://github.com/svaarala/duktape https://github.com/svaarala/duktape>, but heh, for sure no. I'd bet $1 there's no way youtube-dl would switch, but I wonder if yt-dlp would?
- delusional 4y agoCan we stop the trend of linking to tweets that just contain another link to the content? what's the point? Wouldn't this be 10x better if it was a link directly to the github?
- kelnos 4y agoI was thinking the same thing; link to the file on Github, with the same title text as is there now, and it saves me an extra click. And any time I don't have to visit Twitter, I consider that a win.
- naikrovek 4y ago
- derangedHorse 4y agoI like the Twitter linking since it's almost like the OP is giving credit to where they found the information from.
- plaguepilled 4y agoAgreed. If you only know this from someone else's observation, you should link the observation.
- mkl 4y agoThat is against HN guidelines: "Please submit the original source. If a post reports on something found on another site, submit the latter." - https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- Tao3300 4y agoCiting the guidelines is against the guidelines, if not by the letter, in spirit. It's boring and it lacks curiosity. It assumes too much about the sharer. "Can we stop this trend" is a dog whistle for the "HN is getting worse" complaint. Instead we could be considering if we're meant to read the Twitter conversation as well, or sharing a laugh about the link in the tweet author's bio. Or maybe the sharer didn't feel comfortable enough seeming like they made the claim but still wanted to share it because it's kind of cool. AFAIK there's no junior HN mod of the year award.
- tonetheman 4y agoIf this got much bigger I would switch it to quickjs
- aeyes 4y agoThey just don't want to use any external dependencies... There is also an AES implementation: https://github.com/ytdl-org/youtube-dl/blob/master/youtube_dl/aes.py https://github.com/ytdl-org/youtube-dl/blob/master/youtube_d...
- Tao3300 4y agoGreenspun's Tenth Rule: > Any sufficiently complicated C or Fortran program contains an ad hoc, informally-specified, bug-ridden, slow implementation of half of Common Lisp. [1] And here we have a complicated Python program with a partial JS implementation in it. [1] https://en.wikipedia.org/wiki/Greenspun's_tenth_rule https://en.wikipedia.org/wiki/Greenspun's_tenth_rule
- atan2 4y agoThis seems to be a pretty small subset of JavaScript, but I personally love small projects like this for educational purposes. Removing the noise and keeping things minimal helps my brain reason about things. Earlier this year I enrolled in an online class called "Building a Programming Language" taught by Roberto Ierusalimschy (creator of Lua) and Gustavo Pezzi (creator of pikuma.com). We created a toy language interpreter/VM and the final code was around of 1,800 lines of Lua code. Keeping things as simple (and sometimes naive) as possible was definitely the right choice for me to really wrap my head around the basic theory and connect the dots. Thanks for the link.