4 ms·
If I were making a cookbook on HTML/JS I would probably leave out this: https://phuoc.ng/collection/html-dom/sanitize-html-strings/#eliminating-the-script-tags
by bobmaxup 3y ago
If I were making a cookbook on HTML/JS I would probably leave out this:
https://phuoc.ng/collection/html-dom/sanitize-html-strings/#eliminating-the-script-tags https://phuoc.ng/collection/html-dom/sanitize-html-strings/#...
- panzi 3y ago`value.startsWith('javascript:')` and there you have a vulnerability. There can be arbitrary white-space before the URL, so you'd need to do `value.trim().startsWith('javascript:')` instead. However, I much prefer white listing instead, i.e. only allowing `http:`, `https:` and maybe `mailto:`, `ftp:`, `sftp:` and such. Maybe allow relative URLs starting with `/`. Maybe. Though that would mean to correctly handle all attributes that actually can be URLs. Again, I'd just white list a few tags plus a few of their attributes.
- zeroCalories 3y agoWhy? This is a common need. Is there an issue with the implementation or another reason why you would want to use a library?
- deleted 3y ago[deleted]
- chrismorgan 3y agoI’m not sure quite why you’re against removing script tags, but honestly that entire article is poor, riddled with disastrously bad advice: • “Using regular expressions”: it suggests that this approach is acceptable within its limits. It’s not at all. As a simple example, the expression shown is trivially bypassed by "<script>…</script >". This is why, unlike the post claims claims, using regular expressions for cleaning HTML is not a common approach. • (“Eliminating the script tags”: I want to grumble about using `[...scriptElements].forEach((s) => s.remove())` instead of `for (const s of scriptElements) { s.remove(); }` or even `Array.prototype.forEach.call(scriptElements, (s) => s.remove())`. Creating an array from that HTMLCollection is just unnecessary and a bad habit.) • “Removing event handlers”: `value.startsWith('javascript:') || value.startsWith('data:text/html')` is inadequate. Tricks like capitalising and adding whitespace (which the browser will subsequently normalise) in order to bypass such poor checks have been common for decades. • “Retrieving the sanitized HTML”: you are now vulnerable to mXSS attacks, which undo all your effort. • “Elements and attributes to remove from the DOM tree”: this proposes a blacklist approach and mentions a few examples of things that should be removed. Each example misses adjacent but equally-important things that should be removed. You will not get acceptable filtering if you start from this approach. • “Simplifying HTML sanitization with external libraries”: this is pitched merely as easier, faster and cheaper, rather than as the only way to have any confidence in the result. • “Conclusion”: as I hope I’ve shown, “The DOMParser API is one tool you can use to get the job done right.” is not an acceptable position. Really, the article could be significantly improved by presenting it as what a common developer might think, and then scribbling all over the problematic things with these explanations of why they’re so bad, and ending with the conclusion “so: just use the DOMPurify library; consider nothing else acceptable”. (There have at times been a couple of other libraries of acceptable quality, but as far as I’m concerned, DOMPurify has long been the one that everyone should use. I note also that this article is talking about client-side filtration. I’m not familiar with the state of the art in server-side HTML sanitisation, where you probably don’t have an actual DOM; this is also a reasonable place to wish to do filtering, but the remaining active mXSS vectors might pose a challenge. I’d want to research carefully before doing anything.) I look forward to the Sanitizer API <https://wicg.github.io/sanitizer-api/ https://wicg.github.io/sanitizer-api/> being completed and deployed, so that DOMPurify can become just a fallback library for older browsers.
- bobmaxup 3y ago> I’m not sure quite why you’re against removing script tags My bad, I left a fragment in the URL.