10 ms·
This kind of thing always makes me nervous, because you end with a mix of methods where you can (supposedly) pass arbitrary user input to them and they'll safel
by entuno 7mo ago
This kind of thing always makes me nervous, because you end with a mix of methods where you can (supposedly) pass arbitrary user input to them and they'll safely handle it, and methods where you can't do that without introducing vulnerabilities - but it's not at all clear which is which from the names. Ideally you design that in from the state, so any dangerous functions are very clearly dangerous from the name. But you can't easily do that down the line.
I'm also rather sceptical of things that "sanitise" HTML, both because there's a long history of them having holes, and because it's not immediately clear what that means, and what exactly is considered "safe".
- voxic11 7mo agoThe idea is you wouldn't mix innerHTML and setHTML, you would eliminate all usage of innerHTML and use the new setHTMLUnsafe if you needed the old functionality.
- post-it 7mo ago> you would eliminate all usage of innerHTML The mythical refactor where all deprecated code is replaced with modern code. I'm not sure it has ever happened. I don't have an alternative of course, adding new methods while keeping the old ones is the only way to edit an append-only standard like the web.
- noduerme 7mo agoFinally, a good use case for AI.
- josefx 7mo agoWouldn't AI be trained on data using innerHTML?
- Aachen 7mo agoMy experience is that they somehow print quite modern code despite things like ES6 being too new to be standard knowledge even for me and I'm not even middle-aged yet Maybe the last 10 years saw so much more modern code than the last cumulative 40+ years of coding and so modern code is statistically more likely to be output? Or maybe they assign higher weights to more recent commits/sources during training? Not sure but it seems to be good at picking this up. And you can always feed the info into its context window until then
- skeeter2020 7mo agoThis is not my experience. Claude has been happily generating code over the past week that is full of implicit any and using code that's been deprecated for at least 2 years. >> Maybe the last 10 years saw so much more modern code than the last cumulative 40+ years of coding and so modern code is statistically more likely to be output? The rate of change has made defining "modern" even more difficult and the timeframe brief, plus all that new code is based on old code, so it's more like a leaning tower than some sort of solid foundation.
- SahAssar 7mo agoES6 is 11 years old. It's not that new.
- Aachen 7mo agoHence the example of how long it takes non-LLMs to pick that up, whereas LLMs seem to get it despite there being loads of old code out there See also my reply to the sibling comment with the same remark https://news.ycombinator.com/item?id=47151211 https://news.ycombinator.com/item?id=47151211 My mistake for saying 10 instead of 11 years btw, but I don't think it changes the point
- chrisweekly 7mo ago> "ES6 being too new to be standard knowledge" Huh? It's been a decade.
- Aachen 7mo agoYeah, using a kilowatt GPU for string replacement is going to be the killer feature. I probably shouldn't even be joking, people are using it like this already
- charcircuit 7mo agoWhen the condition for when you want to replace is hard to properly specify, AI shines for such find and replaces.
- Aachen 7mo agoThis one is literally matching "innerHTML = X" and setting "setHTML(X)" instead. Not some complex data format transformation But I can see what you mean, even if then it would still be better for it to print the code that does what you want (uses a few Wh) than doing the actual transformation itself (prone to mistakes, injection attacks, and uses however many tokens your input data is)
- charcircuit 7mo agoThat can break the site if you do the find and replace blindly. The goal here is to do the refactor without breaking the site.
- lelanthran 7mo ago> When the condition for when you want to replace is hard to properly specify, AI shines for such find and replaces. And, in your opinion, this is one of those cases?
- charcircuit 7mo agoIt is because the new API purposefully blocks things the old API did not.
- littlestymaar 7mo agoThis ship has sailed unfortunately, no later than yesterday I've seen coworkers redact a screenshot using chatGTP.
- Vinnl 7mo agoI kinda like the way JS evolved into a modern language, where essentially ~everyone uses a linter that e.g. prevents the use of `var`. Sure, it's technically still in the language, but it's almost never used anymore. (Assuming transpilers have stopped outputting it, which I'm not confident about.)
- thunderfork 7mo agoDepending on the transpiler and mode of operation, `var` is sometimes emitted. For example, esbuild will emit var when targeting ESM, for performance and minification reasons. Because ESM has its own inherent scope barrier, this is fine, but it won't apply the same optimizations when targeting (e.g.) IIFE, because it's not fine in that context. https://github.com/evanw/esbuild/issues/1301 https://github.com/evanw/esbuild/issues/1301
- delaminator 7mo agofor some values of "everyone" and "never".
- yawaramin 7mo agoActually... https://github.com/microsoft/TypeScript/issues/52924 https://github.com/microsoft/TypeScript/issues/52924
- Vinnl 7mo agoAh yeah, I remember that. General point still stands: in terms of the lived experience of developers, `var` is essentially deprecated.
- plorkyeran 7mo agoI touch JS that uses var heavily on a daily basis and I would be incredibly surprised to find out that I am alone in that.
- Vinnl 7mo agoThat is indeed why I added qualifiers to "everyone" and "never".
- thenewnewguy 7mo agoIf you want to adopt this in your project, you can add a linter that explicitly bans innerHTML (and then go fix the issues it finds). Obviously Mozilla cannot magically fix the code of every website on the web but the tools exist for _your_ website.
- bulbar 7mo agoIt for sure happens for drop in replacements.
- littlestymaar 7mo agoNobody's talking about old code here. Having an alternative to innerHTML means you can ban it from new code through linting.
- croes 7mo agoIf I need the old functionality why not stick to innerHTML?
- orf 7mo agobecause the "unsafe" suffix conveys information to the reader, whereas `innherHTML` does not?
- goatlover 7mo agoAny potential reader should be familiar with innerHTML.
- kennywinker 7mo agoRight. Like how any potential reader is familiar with the risks of sql injection which is why nothing has ever been hacked that way. Or how any potential driver is familiar with seat belts which is why everybody wears them and nobody’s been thrown from a car since they were invented.
- deleted 7mo ago[deleted]
- orf 7mo agoyes, and bugs shouldn't exist because everyone should be familiar with everything.
- croes 7mo agoBut if some are marked unsafe and others are not it gives a false sense of security if something is not marked unsafe.
- orf 7mo agoSo we shouldn’t mark anything as unsafe then? And give no indication whatsoever? The issue isn’t that the word “safe” doesn’t appear in safe variants, it’s that “unsafe” makes your intentions clear: “I know this is unsafe, but it’s fine because of X and Y”.
- reddalo 7mo agoYou can't rename an existing method. It would break compatibility with existing websites.
- extraduder_ire 7mo agoI looked up setHTMLUnsafe on MDN, and it looks like its been in every notable browser since last year. Good idea to ship that one first, when it's easier to implement and is going to be the unsafe fallback going forward.
- onion2k 7mo agoI looked up setHTMLUnsafe on MDN, and it looks like its been in every notable browser since last year. Oddly though, the Sanitizer API that it's built on doesn't appear to be in Safari. https://developer.mozilla.org/en-US/docs/Web/API/Sanitizer https://developer.mozilla.org/en-US/docs/Web/API/Sanitizer
- DoctorOW 7mo agoThey do link the default configuration for "safe": https://wicg.github.io/sanitizer-api/#built-in-safe-default-configuration https://wicg.github.io/sanitizer-api/#built-in-safe-default-... But I agree, my default approach has usually been to only use innerText if it has untrusted content: So if their demo is this: container.SetHTML(`<h1>Hello, {name}</h1>`); Mine would be: let greetingHeader = container.CreateElement("h1"); greetingHeader.innerText = `Hello, {name}`;
- itishappy 7mo agoWhat if I wanted an <h2>? Edit: I don't mean this flippantly. If I want to render, say, my blog entry on your site, will I need to select every markup element from a dropdown list of custom elements that only accept text a la Wordpress?
- post-it 7mo agorealSetSafeHTML()
- jncraton 7mo agoYou are right that the concept of "safe" is nebulous, but the goal here is specifically to be XSS-safe [1]. Elements or properties that could allow scripts to execute are removed. This functionality lives in the user agent and prevents adding unsafe elements to the DOM itself, so it should be easier to get correct than a string-to-string sanitizer. The logic of "is the element currently being added to the DOM a <script>" is fundamentally easier to get right than "does this HTML string include a script tag". [1] https://developer.mozilla.org/en-US/docs/Web/API/Element/setHTML https://developer.mozilla.org/en-US/docs/Web/API/Element/set...
- entuno 7mo agoIt's certainly an improvement over people trying to homebrew their own sanitisers. But that distinction of being XSS-safe is a potentially subtle one, and could end up being dangerous if people don't carefully consider whether XSS-safe is good enough when they're handling arbitrary users input like that.
- intrasight 7mo agoAlso has made me nervous for years that there's been no schema against which one can validate HTML. "You want to validate? Paste your URL into the online validation tool."
- Dylan16807 7mo agoThis help? https://github.com/validator/validator https://github.com/validator/validator But for html snippets you can pretty much just check that tags follow a couple simple rules between <> and that they're closed or not closed correctly.
- intrasight 7mo agoThat app does look helpful!
- snowhale 7mo ago[dead]
- pornel 7mo agoBTW, HTML allows inline SVG with an XML-flavored syntax that interprets <script/> and <title> differently. It's a goldmine for sanitizer escapes. There are completely bonkers syntax switching and error recovery rules that interact with parsing modes (there's even an edge case where a particular attribute value switches between HTML and XML-ish parsing rules). Don't even try to allow inline <svg> from untrusted sources! (and then you still must sanitise any svg files you host)
- kccqzy 7mo agoIf you just serve SVGs through <img> tag it’ll be much safer. I never understood the appeal of inline <svg> anyways.
- rwj 7mo agoInline reduces round trips.
- toast0 7mo agoYou can use img with a data url?
- lenkite 7mo agoInline SVG is stylable with CSS styles in the same HTML page.
- runarberg 7mo agoAlso animatible with the same context (Animation API, etc.) as the parent page, so different SVGs can influence each other’s animations.
- cxr 7mo ago
- noduerme 7mo agoSome sanitization is better than none? If you're relying on the browser to handle it for you, you're already in a lot of trouble.
- deleted 7mo ago[deleted]
- jaffathecake 7mo agofwiw, if you serve your page with: Content-Security-Policy: require-trusted-types-for 'script' …then it blocks you from passing regular strings to the methods that don't sanitize.
- Cthulhu_ 7mo agoIdeally you should be able to set a global property somewhere (as a web developer) that disallows outdated APIs like `innerHTML`, but with the Big Caveat that your website will not work on browsers older than X. But maybe there's web standards for that already, backup content if a browser is considered outdated.
- afavour 7mo agoI like the idea of that. But I imagine linting rules are a much more immediate answer in a lot of projects.
- staticassertion 7mo agoDoesn't using TrustedTypes basically do that? I'm not really web-y, someone please correct me if I'm off.
- madeofpalk 7mo agoYup, this is basically what TrustedTypes is for!
- cxr 7mo agoIt's not an "outdated API". It's still good for what it was always meant for: parsing trusted, application-generated markup and atomically inserting it into the content tree as a replacement for a given element's existing children. > set a global property somewhere (as a web developer) that disallows[…] `innerHTML` Object.defineProperty(Element.prototype, "innerHTML", { set: (() => { throw Error("No!") }) }); (Not that you should actually do this—anyone who has to resort to it in their codebase has deeper problems.)
- onion2k 7mo agoit's not at all clear which is which from the names There's setHTML and setHTMLUnsafe. That seems about as clear as you can get.
- hahn-kev 7mo agoBut you can use InnerHTML to set HTML and that's not safe.
- onion2k 7mo agoAt this point that API has been around for decades and is probably impossible to deprecate without breaking fairly large amounts of the web. The only option is to introduce a new and better API, and maybe eventually have the browser throw out console warnings if a page still uses the old innerHTML API. I doubt any browser vendor will be gung ho enough to actually remove it for a very long time.
- entuno 7mo agoIf that'd been the design from the start, then sure. But it's not at all obvious that setHTML is safe with arbitrary user input (for a given value of "safe") and innerHTML is dangerous.
- HWR_14 7mo agoThat's why I only allow user input of alphanumeric ascii characters. No need to worry about sanitation then, and you can just remove all the characters that don't match. (It's a joke, but it is also 100% XSS, SQL injection, etc. safe and future proof)
- thaumasiotes 7mo ago> I'm also rather sceptical of things that "sanitise" HTML, both because there's a long history of them having holes, and because it's not immediately clear what that means, and what exactly is considered "safe". What is safe depends on where the sanitized HTML is going, on what you're doing with it. It isn't possible to "sanitize HTML" after collecting it so that, when you use it in the future, it will be safe. "Safe" is defined by the use. But it is possible to sanitize it before using it, when you know what the use will be.
- cxr 7mo ago> it's not at all clear which is which from the names. Ideally you design that in from the [start] It was, and there is: setting elementNode.textContent is safe for untrusted inputs, and setting elementNode.innerHTML is unsafe for untrusted inputs. The former will escape everything, and the latter won't escape anything. You are right that these "sanitizers" are fundamentally confused: > "HTML sanitization" is never going to be solved because it's not solvable.¶ There's no getting around knowing whether or any arbitrary string is legitimate markup from a trusted source or some untrusted input that needs to be treated like text. This is a hard requirement. <https://news.ycombinator.com/item?id=46222923 https://news.ycombinator.com/item?id=46222923> The Web platform folks who are responsible for getting fundamental APIs standardized and implemented natively are in a position to know better, and they should know better. This API should not have made it past proposal stage and should not have been added to browsers.
- Dylan16807 7mo ago> There's no getting around knowing whether or any arbitrary string is legitimate markup from a trusted source or some untrusted input that needs to be treated like text. This is a hard requirement. It is not a hard requirement that untrusted input is "treated like text". And this API lets you customize exactly what tags/attributes are allowed in the untrusted input. That's way better than telling everyone to write their own; it's not trivial.
- cxr 7mo ago[flagged]
- Dylan16807 7mo agoI don't see how I differed from what you said? You divided strings going into HTML into two categories, where one category uses textContent and the other category uses innerHTML. My point is to disagree with those categories, not whatever subtle thing you're taking issue with.
- 7mo ago
- duxup 7mo agoYeah someone tells me something has been made “safe” is nice but unless I know exactly what that means … it’s easy to say safe by someone who doesn’t have to deal with it when the bad corner case happens. Oh and it’s safe… in this browser… not that one, so this idea of safety is kinda dead to me for now.