17 ms·
Fun Facts on Producing Minimal HTML
- ShaneMcGowan 6y agoThis is gross, please don't tell people this stuff, pleeeeease
- buzzerbetrayed 6y agoAgreed. I gave up after they suggested using <pre> to force your <a> tags on to separate lines. Scary stuff. There are many ways to do this, and using <pre> should not be one of them.
- oefrha 6y agoThe funny thing is correctly using <br> actually produces shorter HTML than the grossly unsemantic (and certainly doesn’t render nicely) <pre> suggestion, at least in the given example.
- schwartzworld 6y agoa { display: block; } was that so hard?
- boznz 6y agoIt's never going to look nice but I have a couple of embedded devices with a few K of memory from 15 years ago still serving pages and still working fine in google and Firefox. Nit-picking, but rather than saying "99% of browsers" (and I am not sure there are 100 browsers out there to get that particular stat) it would be best to just mention the ones it doesn't work on.
- vpzom 6y agoI think there are some differences in rendering if you omit the DOCTYPE Also, why?
- sjwright 6y agoBecause your browser will enable a compatibility shim to improve rendering of web pages authored 15+ years ago. If you omit DOCTYPE, your HTML is assumed to be very old.
- thinkloop 6y ago> <meta name=viewport content="width=device-width, initial-scale=1"> > This snippet is copy and pasted a whole lot around the web. Most people don't explain how it actually functions though. "width" sets the initial width to the mobile's physical display width in 100% pixels. I still don't get this - if the browser is the width of the device, why doesn't the site flow to 100% of the browser?
- SahAssar 6y agoIt's because mobile browsers have to deal with sites built only for desktops. So if they don't get this tag they wil emulate a desktop resolution but on a mobile screen. The tag sets the width to the devices actual width (although pixel scaling can change the display resolution) and the initial scale sets the scale to not imitate a desktop.
- giantrobot 6y agoThe original iPhone's width in portrait orientation was 320px, the viewport was 980px. When a page loaded it was rendered to a viewport 980px wide, like a browser window 980px wide. Without setting a viewport setting Safari would render a page as if it was a 980px side browser window. The viewport meta tag let's you adjust this virtual window size. You can set it to whatever value you want but "device-width" will just use the screen's current width according to its current orientation. You can add it to a page's header and in many cases it'll look way better on mobile even without specific CSS targeting mobile. The default viewport size of the original iPhone was in place to let it use the "normal" web. At the iPhone's introduction this was a big deal because other smartphone browsers worked best with simple "mobile" layouts. Safari on the iPhone rendered web pages as they looked on the desktop. Making the viewport sized to the display usually makes it so a user doesn't need to zoom in to read text or horizontally scroll to read wonder content. This behavior makes modern mobile browsers render pages more like old mobile browsers where the viewport was usually the portrait screen width (320px typically).
- masklinn 6y agoUsefully though inconveniently for my purposes https://shachaf.net/w/b-trees https://shachaf.net/w/b-trees had this exact issue when it first made the rounds, despite the simplicity of the page it was rather difficult to read on most mobile devices as it did not have a viewport meta, so you’d get a pretty tiny font and it would not reflow when zoomed to a readable level.
- btrettel 6y agoPractically speaking these optimizations won't make much of a difference, but I still find them interesting. I have been keeping a list of similar optimizations. Here are some not in the linked article: - Use relative URLs when possible, i.e., /page.html, no need to specify the protocol. - Use shorthand CSS properties like font, background, margin, border, padding, and list. - Use lowercase tags as they compress better. See: https://encode.su/threads/1889-gzthermal-pseudo-thermal-view-of-Gzip-Deflate-compression-efficiency https://encode.su/threads/1889-gzthermal-pseudo-thermal-view... - Use shorthand hex colors. These optimizations are harmless compared against some of the ones recommended in the linked article...
- cdubzzz 6y agoWeird. I have always incorrectly referred to URLs beginning with a “/“ as absolute. Oops.
- johncmouser 6y agoI thought that "/" was absolute and "../../something/else" was relative.
- btrettel 6y agoIn my notes I had this called "site relative" following this Stack Exchange post: https://webmasters.stackexchange.com/a/71376/12374 https://webmasters.stackexchange.com/a/71376/12374 But you're right; relative by itself would refer to relative to the current location. Error on my part in not being specific enough. In retrospect "site relative" is not good terminology.
- deaddodo 6y agoIt's URL relative. Example, for "http://mysite.com/dir/page1.html" http://mysite.com/dir/page1.html": * Relative: "../other_dir/page2.html" * URL relative: "/other_dir/page2.html" * Absolute: "http://mysite.com/other_dir/page2.html"
- divbzero 6y ago> Use relative URLs when possible Protocol relative URLs can also be used for different domains as long as protocol remains the same, which is often the case with https: so widespread, e.g.: <script src=//code.jquery.com/jquery-3.5.1.slim.min.js integrity="sha256-4+XzXVhsDmqanXGHaHvgh1gMQKX40OUvDEBTu8JcmNs=" crossorigin=anonymous> </script>
- mindctrl-org 6y ago> You don't need to close your tags. Not always true. I’ve run into numerous issues caused by the lack of closing tags, and just did earlier this week.
- duxup 6y agoYeah so many code formatters and such tell me to close the tag .. I'm gonna close it... must be a reason.
- lazyjones 6y agoThe reason is that the code formatters are broken and don't support the full html5 specification.
- johncmouser 6y agothey will be nested right. so <div> <p> <h1> <h2> without closing tabs create the DOM tree div->p->h1->h2 if you were actually developing production code and misplaced, let's say, <p>'s closing tag, then that would mess up the rest of your tree (from your perspective -- the computer doesnt care)
- naniwaduni 6y agoThis actually produces the DOM equivalent to <div> <p> </p><h1> </h1><h2> </h2></div>. Many of the rules for unclosed tags are more there so that browsers can agree on what to do with garbage first, and for you to rely on only incidentally! They defer to historical practice before common sense! In order to predict this reliably, you essentially need to have the list of content categories[1] memorized (or look them up). Not all of them are ... necessarily intuitive. [1]: https://developer.mozilla.org/en-US/docs/Web/Guide/HTML/Content_categories https://developer.mozilla.org/en-US/docs/Web/Guide/HTML/Cont...
- ufo 6y agoIs there a way to get warnings for HTML that looks valid with matching start and end tags but doesn't actually parse the way it is written? I get the impression that we end up needing to memorize those content categories even if we plan to only generate html with all the start and end tags. For example, <p>A<p>B</p>C</p> looks like two nested <p> but it is parsed as 3<p> next to each other: <p>A</p><p>B</p>C<p></p>.
- SahAssar 6y ago> Text outside of tags is acceptable in modern browsers. Please don't. If it's plaintext serve it as such and if it's html serve it and format it as such. > <!DOCTYPE html> is required by the HTML spec Yes, it is and do it. Disregard everything the author wrote after that on this point. > <html>, <head>, <body> are not required by modern browsers. Agreed, they should be left behind. This is a valid HTML doc: <!DOCTYPE html> <title>test</title> <p>test doc here</p> > You don't need to close your tags. Some tags need closing, some don't. This is documented in the standard. Follow it, don't just freestyle it. > If you do not define <meta charset="utf-8">, then most browsers will default to ASCII or Windows-1252. So if you are confidently in the ASCII range, skip the declaration. Please don't. Set the charset in your headers (and it should be utf-8 unless you have a very good reason). > Using a preformatted block (<pre>) of links This is just bad. You have a list of links but you don't use the exact element created for creating lists?
- traes 6y ago> Please don't. Set the charset in your headers (and it should be utf-8 unless you have a very good reason). Legitimate question: why? If I'm not planning on using non-ASCII characters, why bother?
- rglullis 6y agoThere is a huge distance between what you are planning to do and what happens when reality shows up.
- dragonwriter 6y ago> Legitimate question: why? If I'm not planning on using non-ASCII characters, why bother? Because not all character sets are ASCII compatible, and you don't know that your user's default is, even though most browsers' defaults if not customized are.
- thristian 6y agoBecause if you don't specify a character set, the browser will be forced to guess, and even with plain ASCII input the guess is not always reliable: https://en.wikipedia.org/wiki/Bush_hid_the_facts https://en.wikipedia.org/wiki/Bush_hid_the_facts
- lhorie 6y agoThese are certainly fun (in a for-teh-lulz sort of way), but production grade html minifiers actually use techniques like omitting quotes too. Many of the other techniques are highly questionable, so again, as a rule of thumb, if an html minifier doesn't do it, you probably shouldn't either Also worth mentioning, the ultimate minimalism "hack" is to simply serve a txt or md file w/ Content-Type: text/plain
- johncmouser 6y agonot sure how right this is but this is on twitter https://twitter.com/hncynic/status/1258916263562219520 https://twitter.com/hncynic/status/1258916263562219520
- MR4D 6y agoI wonder why we don’t have Markdown browsers. Seems like that would help a ton.
- Minor49er 6y agoApparently Markdown has a text/markdown mime type: https://stackoverflow.com/a/25812177 https://stackoverflow.com/a/25812177 It would be simple to have a browser or plugin detect and render these, assuming they don't already.
- w0mbat 6y agoEx-browser dev here. Please don't ship invalid HTML like this. Web standards specify how to render valid HTML in a standard way, and browsers have become better and better at that over the years. Invalid HTML destroys all that. Each browser will have to guess how to repair and fill in the blanks on malformed incomplete tag soup, and there is no one right way to do that. Browser X will make different guesses from Browser Y, and the next version of each will be different too. Please just write actual HTML that is valid and your website will render much more consistently and reliably across shifting platforms, browsers and versions. It is not the size of HTML that slows the web down anyway.
- Zarel 6y agoAs of HTML5, web standards specify how to render invalid HTML as well, precisely to avoid this problem: https://html.spec.whatwg.org/multipage/parsing.html#parse-errors https://html.spec.whatwg.org/multipage/parsing.html#parse-er... In addition, most of the suggestions in the link (leaving off <html> and <head>, leaving off quotes for attributes, not closing <p>) are valid HTML in the first place.
- lhorie 6y agoTo add to this, one of suggestions in the article was to drop the doctype. While that puts the browser in tag soup parsing mode rather than html5, in practice browsers don't actually start diverging much in behavior until you get to truly messed up markup, like malformed tags. So even the more egregious suggestions here are probably still fine, provided that you have equally egregious CSS hacks to get around ancient box model quirks
- oefrha 6y agoTrue, and the suggestions in the article are pretty tame. But let’s set aside the article and talk about invalid HTML for a minute. I write scrapers a lot (not the irresponsible kind and never for monetary gains) and invalid HTML, while technically parseable and to spec, are often a pain in the ass. You have to bring in a full blown HTML5 parser, and they could be way slower (e.g. lxml.html vs html5lib). Depending on your language of choice there might not even be a to-spec HTML5 parser available. So, just close your damn tags (except self-closing ones), and close them in order, it’s not hard, the size increase is minimal, it will help with your own sanity and people will thank you for it.
- emilfihlman 6y agoIt seems people are entirely missing the point of the exercise.
- arkitaip 6y agoThis is HN; missing the point is the point.
- johncmouser 6y agoAh, maybe a better name would have helped: HTML Code Golf - How to make really small HTML that doesn't break Firefox or Chrome, currently at least
- pmiller2 6y agoExcept that it isn't HTML if it doesn't follow the standard.
- johncmouser 6y agookay, quasi-html
- yoloClin 6y agoHTML Code Golf - How to make really small _extremely fragile_ HTML that doesn't break Firefox or Chrome, currently at least FTFY
- pmiller2 6y ago
- syrrim 6y agoCouple more tips for those who /really/ want to slim down their html: - opening <a> tags close the previous <a> tag... but without an href they do nothing. Use them as if they were closing <a>s to save on slashes - <select>...<select> is completely equivalent in html to <select>...</select>. Save more slashes this way - formatting elements won't be closed automatically, they rear their head again like so many hydras until your burn the wounds with end tags. But! there's a way around this: consider '<div><b><b><b><b></div></b></b></b>X'; that's funny... why isn't the X bold? it turns out that only three identical (down to attributes) formatting tags are remembered. Use this to your advantage when nesting identical formatting tags inside themselves. - since we're saving on slashes: <table> (while in a row, ie <tr><table>) closes the last table, and opens a new one. You might think: "I don't want to start another table so soon!" fear not! the browser will move anything you put in a table, above the start of the table, right up until you start putting cells in it. This also avoids having to open the table later, you can start <tr>ing right away. - you saved characters by dropping the doctype... but was it worth it? only with a valid doctype declaration will you <p> tags be closed automatically when you open tables. Just 4 closing </p> tags will make you wish you included that doctype. still think it's worth it?
- johncmouser 6y agowait, for which tags does this apply? because won't <p>one<p>two<p>three<p>four create four <p> elements nested together? does this only work with <a> tags? your third point about the triple-identical (reminds me about TCP ACK and Re-transmit haha) is pretty nuts though
- dragonwriter 6y ago> won't “<p>one<p>two<p>three<p>four” create four <p> elements nested together? No, because “A p element's end tag may be omitted if the p element is immediately followed by an address, article, aside, blockquote, details, div, dl, fieldset, figcaption, figure, footer, form, h1, h2, h3, h4, h5, h6, header, hgroup, hr, main, menu, nav, ol, p, pre, section, table, or ul element, or if there is no more content in the parent element and the parent element is an HTML element that is not an a, audio, del, ins, map, noscript, or video element, or an autonomous custom element.” https://html.spec.whatwg.org/multipage/syntax.html#syntax-tag-omission https://html.spec.whatwg.org/multipage/syntax.html#syntax-ta...
- panic 6y agoAnother helpful meta tag is <meta charset=utf-8> Without this, your document may be interpreted using an implementation-defined ASCII-like encoding (e.g., windows-1252 for English-speaking locales) if served without a Content-Type.
- chrismorgan 6y ago<meta name=viewport content="width=device-width, initial-scale=1"> The `, initial-scale=1` has been unnecessary for a few years now (sorry, not searching for the citation now, hopefully you can find it if you’re interested), so it’s slimmer to use this instead: <meta name=viewport content="width=device-width"> If you drop the quotes, it’ll parse the same way (and parsing is well-defined in HTML now, so you can be confident all browsers will handle all parsing the same) but be nominally non-conformant. Up to you how much you care about non-conformance, but I avoid writing non-conformant documents, though I do regularly hand-write minimal HTML (omitting html/head/body, skipping unnecessary closing tags, unquoting attribute values, &c.) ---- <!DOCTYPE html> I strongly recommend against omitting this, because removing it throws you into quirks mode. Also I recommend spelling it `<!doctype html>`, because that will regularly save a byte or two in gzipping due to the much greater frequency of lowercase letters. ---- <meta charset=utf-8> I strongly recommend keeping this; it’s a safe bet that your text editor is working in UTF-8, so specifying the charset thus ensures that if later on you insert some non-ASCII (e.g. pasting a quote that includes curly quotes) it will work properly. You can also save one more byte by spelling this `<meta charset=utf8>`. There’s fun history around that in the encoding spec, where that used to not be a valid value, but based on observing people spelling it that way sometimes they added it. So it’s now valid, but particularly old browsers might not like it. ---- Using <pre> for line breaks? Please no. Just don’t do this. The side-effects are awful. Use <br> if line breaks is all you want. ---- > Both single-quotes and double-quotes are valid for tag parameters. This is useful for producing valid HTML output in programs without resorting to escaped double-quote. I like to do minimal encoding. Within a quoted attribute value, the only characters that need to be escaped are & and the particular quote used, so for an attribute that includes a large blob of JSON I like to use single quotes, so that I only need to escape single quotes and ampersands: <a data-json='{"department":"R&D"}'>…</a> This yields a smaller and more human-readable result, which is also nice. ---- > You don't need to close your tags. Well, there are three cases to consider here: 1. Self-closing tags like <meta>, which don’t have a closing tag (so it’s actively wrong to include one). 2. Tags for which the end-tag is optional, depending on what follows it, e.g. <p> doesn’t need </p> if it’s followed by various elements such as another <p>. 3. Non-conformant documents where the well-defined parsing behaviour just happens to produce what you want, despite what you’ve written being probably nonsensical and something where a human wouldn’t be sure what you meant.
- _bxg1 6y agoPlease please please don't write article text in preformatted blocks. It makes it impossible to read on mobile. Even reader mode doesn't work, because it respects preformatting.
- treeman79 6y agoReader view in iPhone is broken on this page. My vision is poor, and reader view is an amazing help.
- unicornporn 6y agoFirefox mobile on Android works perfect in reader mode. However, you shouldn't need to switch browser to read the content.
- divbzero 6y agoFrom my own notes on minimal HTML5… Closing tags are optional for the following tags: html head body p dt dd li option thead th tbody tr td tfoot colgroup The following tags are self-closing and should not have closing tags: meta img input hr br Attributes can be left unquoted if the following characters don’t appear in the attribute value: - Single quote (') - Double quote (") - Space ( ) - Equal sign (=) - Greater-than sign (>)
- recursive 6y agoYou can even have an unquoted attribute value with a space, if you replace it with a plus sign.
- hannob 6y agoA less well known tip if you want to micro-optimize html: Use protocol-relative external links. If your own site is https only (which it should be) then <a href="//example.org/">example</a> is the same as <a href="https://example.org/">example</a> https://example.org/">example</a>
- jraph 6y agoI'm conflicted on omitting tags to make pages lighter. I'm sensitive to making things lighteight but I really do like the XHTML parser catching dumb errors that would be silent bugs in HTML. I also like the readability of a document where all closing tags are here. Some tricks can arguably make the code more readable (ommitting head and body) but some tricks require effort to understand if you are not used to them. We write code for human beings first. Maybe an HTML minimizer could be used if one wants to save bytes?
- eska 6y agoI use XHTML in my own static site generator, together with an external XML minifier library, then validate the output at build time. In my tests XHTML had a significant parsing advantage over HTML, and I didn't need to do any questionable stuff like in the suggestions here.