8 ms·
* PDFs are self-contained and offlineable HTML can easily be offline-able. Base64 your images or use SVG, put your CSS in the HTML page, remove all 2-way data
by monkeynotes 5y ago
* PDFs are self-contained and offlineable
HTML can easily be offline-able. Base64 your images or use SVG, put your CSS in the HTML page, remove all 2-way data interaction, basically reduce HTML to the same performance as PDF and allow it to be downloaded.
* PDFs are files
HTML is files
* PDFs are decentralised
This should be "PDFs can be decentralised". PDFs aren't inherently any more decentralised than any other kind of file, including HTML.
The store is the thing that becomes decentralised, not the content.
* PDFs are page-oriented
HTML can be page-oriented. Simply build your website with pagination. PDFs can also be abused to have hugely long pages. Bad UX can be encapsulated in any medium.
* PDFs used to be large (bla bla bla Javascript weighs a lot)
Nope, PDFs are still objectively larger than the equivalent HTML. PDFs don't have any dynamic interaction, rip all that out and produce the HTML of yesteryear and your HTML will be tiny in comparison to the PDF.
Edit: I'm sorry, the more I think about this the dumber I feel. The web is useful because it's 2-way. I am excited by the web because I can interact with other people. I come to hacker news to engage with thinkers, not to just read a published article from one single author. I want to read ad-hoc opinions and user submitted content. PDF web, really?
- playpause 5y agoThese all seem like technical quibbles that miss the point.
- wlesieutre 5y agoUnless I'm on a paper-sized tablet I would definitely rather have an offline HTML file than a PDF. Nobody likes to pan back and forth on lines of text to read something.
- Robotbeat 5y agoI had the exact opposite reaction. I’m reading this on an iPhone SE2020, and I MUCH appreciate reading this in pdf form. I didn’t have to pan back and forth or even put the phone in landscape orientation. This is one of the smallest smartphones you can still buy, and the experience of PDF is WAY better than the user-hostile auto-flow text forced down mobile users’ throats. I was skeptical at first, but I think the author made the point fantastically well.
- nemetroid 5y agoI’m using a 2016 iPhone SE, and it’s largely unreadable without being very up close.
- cunthorpe 5y agoWhat. Your browser has a zoom functionality that lets you make the text smaller, essentially replicating the PDF site above. Only the opposite of what you say is correct: I can’t read that PDF’s text without turning my phone into landscape and picking up my glasses.
- wlesieutre 5y agoTo get equally small text on my desktop I have to turn the font size all the way down to 7. God forbid you have readers with less than stellar eyesight. I get what they're going for but the PDF is not exactly an accessible reading experience.
- apotheon 5y agoEPUB would beat the shit out of PDF for that. (EPUB is basically a subset of HTML with client-oriented context.)
- pseingatl 5y agoPDF is size-agnostic. There's nothing to stop you from creating documents the size of a phone screen.
- wlesieutre 5y agoI’m commenting here as a user reading a PDF. The fact that someone else could have laid it out differently doesn’t change the fixed layout of the PDF that I’m trying to read. There’s a reason responsive design has been a big deal for the last 10+ years and I don’t think the benefits of PDF are worth throwing it out.
- JohnFen 5y agoAs someone who really detests responsive design, the lack of it in a PDF strikes me as a feature, not a bug.
- monkeynotes 5y agoThe guy outlines his whole case based on those exact points which are, as you have observed, technical quibbles and not a basis for abandoning HTML. Under the hood it seems apparent to me that the real premise is an emotional one, not a technical one. The internet is plastic not because of HTML, but because of money and people. When you have teens driving content it's going to feel plastic. When Walmart uses the internet to sell you crap it's gonna be plastic. Gossip / social platforms are trash, no matter the medium. It could be argued that TV is an incredible learning platform ruined by HD. Back in the standard definition days we had proper news, documentaries that were substantial, and no reality TV. We need to go back to black and white standard definition. Sorry, but the PDF web is not a solution to societal rot.
- adolph 5y agoIs the medium the message? Does style have substance? Is form also a function?
- leetcrew 5y agoI'm not exactly sure what point you're trying to make here, but I don't think two different formats for encoding formatted text with images constitute different "mediums".
- megameter 5y agoOf course they are, and we run into it constantly in computing. You can encode text with images as a bitmap, as vector graphics, as symbolic content that references bitmaps or vectors, as an algorithm that procedurally generates any of the above... While you can produce identical outputs from the different methods, it's not hair-splitting to say that the authoring process and hence the nature of the medium to shape expression is affected by choosing one. When you opt towards maximizing generality your production cycle can grow without bound because everything is possible by layering different media, even if all of it is unnecessary. That's how you end up with creative projects that take multiple years to decades to accomplish.
- 5y ago
- jedimastert 5y agoThis statement could be for both the comment you're replying to and the original article.
- quietbritishjim 5y ago> These all seem like technical quibbles that miss the point. If these all "miss the point", what is the point? It seems to me that the article's point is that PDF as a format has attributes that satisfy the author's goal, whereas HTML does not. The parent comment says that HTML does have those attributes after all (if you choose to use HTML that way). That is very directly addressing the article's point, as I understand it.
- JohnFen 5y agoPerhaps I misunderstood, but I believe the author's point was to highlight what a steaming mess the modern web is. The PDF aspect strikes me as illustrating a point, not a seriously proposed solution.
- gunapologist99 5y agoagreed. and, ancient HTML can still be easily read by modern browsers, so that's not exactly a special attribute of PDF either.
- noduerme 5y agoHonestly, if you're going to put out a manifesto as a PDF, at least take some time "layouting" your design. The one advantage of that format is that you control the aspect ratio. Every font is permissible, everything is absolutely positioned. Using a generator to create it is cringey. Show the art that's possible. Really sell the format. FWIW I deliver PDFs daily as an art director; not ideal, but they work in most cases. There's certainly nothing rebellious or non-commercial about them.
- hyperpape 5y agoSaying HTML can be offlineable is like saying C can be provably terminating. There's a subset of programs where that's true, but it's not inherent to the form. A PDF is inherently self-contained, standard web technologies are not. When you open the page and it's a PDF, it gives you certain guarantees, when you open it and it's HTML, you have to have to do further investigation.
- monkeynotes 5y agoI don't buy that the problem with the web is that HTML is not inherently offlineable. HTML may not be inherently offlineable but it can be. PDF isn't inherently a web friendly format, but it can be. There really isn't any good argument for PDFing the web.
- pajko 5y agoPrint the page to PDF.
- tablespoon 5y ago> Print the page to PDF. Even that usually sucks nowadays, because web developers don't care anymore. Probably 75% of the time before I do that, I have to go into the dev console to delete overlay elements that obscure content and garbage that will waste 10 pages (e.g. grossly oversized images, related article recommendations, etc.). There was a time when most websites had a print view that gave you a simplified html page that worked well, but I think most of those are gone now. Now it's all some print "media-type" CSS that no one ever put the time in to do properly or keep up to date.
- EugeneOZ 5y ago> A PDF is inherently self-contained, standard web technologies are not What technologies exactly? You can have absolutely everything you need inside the HTML. You can inline css, js, svg and images. What technologies you can’t inline?
- 5y ago
- Koshkin 5y agoAll true. Incidentally, I do not see pagination as necessary or in most cases even desirable; rather, I see it as a vestige of the printing technology, while the need for printing has shrunk dramatically over the past 20 years.
- Frost1x 5y ago>PDFs don't have any dynamic interaction... Just a caveat to that statement, you can literally do interactive and dynamic 3D graphics rendering in PDFs: https://helpx.adobe.com/acrobat/using/enable-3d-content-pdf.html https://helpx.adobe.com/acrobat/using/enable-3d-content-pdf.... You can also embed JS in PDFs: https://helpx.adobe.com/acrobat/using/applying-actions-scripts-pdfs.html https://helpx.adobe.com/acrobat/using/applying-actions-scrip...
- dathinab 5y agoYes, and many of this things are "in general" not well supported by anything but adobe PDF. Even most simple interactive things can easily not work correctly even in more widely spread PDF readers. IMHO PDF is in many ways worse then HTML, it's just that this ways are less commonly used, but if you start a PDF instead of HTML trend it's just a matter of time until this "not so compatible" aspects of PDF become widely used by some people.
- monkeynotes 5y agoJS in a PDF? You can do that in HTML, why not use the tools you already have that work together by design? This guy is arguing that removing JS is what makes the web better. Having published, static, paper-like content is the way forward.
- Frost1x 5y agoJust caveating a technical statement I knew wasn't quite true, not making any sort of assessment either way. As someone who has had to extract data from large sets of PDFs and modern web presentation formats, I'm not a fan of either, really. Even verifying that a visibly presented string exists in a PDF document programmatically can be a non-trivial task, as with a given website as well. That to me says a lot.
- chalst 5y agomonkeynotes seems to take the line that technical defects in claims others make fatally undermines their case, but technical defects in his/her arguments are irrelevancies. For what it's worth, the same objection occured to me. The use of scripting I've seen in PDFs has been use-supporting and consistent with their book-like feel.
- rexreed 5y agoAlso - how are PDFs exactly "discoverable"? I have petabytes of PDFs and making them easily "discoverable" for any mass use, such as analytics, search, or data analysis is a massive pain. I'd rather have them in a non-PDF format.
- relaxing 5y agoThe author calling for new content to be authored as PDF, which can easily be made discoverable. I’m guessing your data set is made of scans with poor or no OCR.
- rexreed 5y agoNot a single researcher or data analyst I know of would prefer "discoverable" content to be in PDF format, regardless of just how awesome the OCR is (which it often isn't, especially for tabular data). Even for all-text, non-tabular documents, OCR does not provide the metadata needed to make sense of the documents. Why PDF is claimed to have superior "discoverability" in the OP essay is a mystery to me. For the sake of "discoverability", PDF is definitely not the way to go.
- relaxing 5y agoThe essay claimed > PDFs are discoverable. Search engines index them as easily as any other format. What you’re taking about has nothing to do with that.
- marcosdumay 5y ago> PDFs don't have any dynamic interaction Oh, you are set for a world of surprises. Nearly every single one bad, but running our current web over PDFs is well within the specs.
- pajko 5y agoThis. Also who hates the huge double margins? The slow rendering? The unnatural break-up of text? Meaningless headers and footers? And the whole page-based layout? PDF is not meant for the web. Period.
- goodpoint 5y agoYou seem to miss the point of the post: ---- Call to action Publish in static file formats Date and hash your work Stop spying on your users ---- All this cannot be GUARANTEED by HTML/pdf/epub and requires active cooperation from the author. This is bad.
- ChrisMarshallNY 5y ago> Simply build your website with pagination. My experience is that browsers are terrible with CSS pagination support in their display and printing directly. The only place it seems to actually work is...saving as a PDF...
- tablespoon 5y ago>> * PDFs are self-contained and offlineable > HTML can easily be offline-able. Base64 your images or use SVG, put your CSS in the HTML page, remove all 2-way data interaction, basically reduce HTML to the same performance as PDF and allow it to be downloaded. You're missing the point. Even a relatively computer-illiterate person can easily save a PDF to my hard drive, and it's significantly more difficult with HTML. At a minimum you're probably going to get an HTML file with a sidecar directory (or I believe a sometimes browser-specific archive, it's been a long time since I tried since it works so poorly), and even that may not have the content you want to due to dynamic sites.
- stzups 5y ago>> it's significantly more difficult with HTML Right Click > Save as Try it with this page!
- tablespoon 5y ago> Right Click > Save as > Try it with this page! Say hello to your new sidecar directory (or broken CSS/images/God knows what else)! I tried to save an NY Times article, and it 1) needed JS to display anything, 2) even with the sidecar stuff was broken, 3) it was so plastered with ads and other junk I thought it was incomplete (it wasn't, I just had to scroll waaay down past something that looked like a footer and some voids after that). If you save a PDF, you get that exact PDF on your hard drive, and when you open it (even in 10 years) it will look exactly the same as it did on the site. With PDF WYSIWYS: What you see is what you save.
- trey-jones 5y agoThis is of course the point of the article - that the web is a giant steaming pile of shit for the most part, plagued by JS and external resource requirements, all of which contribute to massive total page size. I'll preface by saying I have some expertise in HTML, but none in PDF (the format). The point of most commenters who suggest that HTML is still a better alternative than PDF (I agree), are assuming that if this is an important issue to you, that you would craft your page in a simpler style compared to most of what we see on the web, making Print to PDF or Save As... more viable. > PDFs and a PDF tool ecosystem exist today. No need for another ghost town GitHub repo with a promising README and v0.1 in progress. This is news to me. I'm not sure that I buy it. PDFs have always been a pain in the ass to work with in my opinion. Maybe there are tools, but in my experience they aren't very good. In general, we know that HTML is going to be much more compact (and compressible!) than PDF and that's the biggest advantage I see on a web where bandwidth still matters. Another downside shows itself by trying to copy and pasting the above quote: PDF formatting seems to be weird.
- majkinetor 5y agoPDF - does not reflow, major suck - is binary format, another major suck So no thx, PDF is outdated tech, while HTML and friends are just abused.
- anigbrowl 5y agoWhat I like best about pdf files is that I can just give them to someone and be almost certain that any questions will be about the content rather than the format of the file.
- chowderman 5y ago> HTML can easily be offline-able. Base64 your images or use SVG, put your CSS in the HTML page, remove all 2-way data interaction, basically reduce HTML to the same performance as PDF and allow it to be downloaded. I built a tool for this exact purpose[0] since the HTML specification and modern browsers have a lot of nice features for creating and reading documents compared to PDF (reflow and responsive page scaling, accessibility, easily sharable, a lot of styling options that are easy to use, ability for the user to easily modify the document or change the style, integration with existing web technologies, etc.). In general I would rather read an HTML document than the PDF document since I like to modify the styling in various ways (dark theme extensions in the browser for example) which may be hard to do with a PDF, but its more of a personal preference. Some people will prefer that the document adjusts to the screen size of the device (many HTML pages), and others will prefer the exact same or similar rendering regardless of the screen size (PDF). Either way, kind of a fun idea making a website using just PDFs. Not the most practical choice, but fun none-the-less. [0] https://github.com/chowderman/hyperfiler https://github.com/chowderman/hyperfiler
- baybal2 5y agoHTML used to be a very nice format at the age of xhtml 1.1, very formally specified, and a tie with DOM was assured by vert strictly standardised DOM v3. And ACID3 was giving you a pixel for pixel repeatability during rendering. HTML+JS today... now it's effectively a standard in name only, and Chrome is the new IE6. The standard is now "what has worked in the last stable release" Now go to http://acid3.acidtests.org/ http://acid3.acidtests.org/ and see how the latest stable Chrome release can't render a decade old CSS testcase.
- kemitche 5y agoPDFs are also horrible to view on mobile, as the text doesn't reflow.
- LeifCarrotson 5y agoWhen you find a page - inherently a document-oriented term - like an article, blog post, how-to, or project writeup that's interesting or useful, and you want to make sure it's available to you later, what do you do? Do you save the HTML, CSS, and Javascript, and hope that it works offline? I used to use the "Save page as..." tool back in the early 2000s, but it's become less and less useful, with too many dysfunctional disappointments. No, I cut out some junk I don't need with the Printliminator [1] bookmarklet, then I do a *print-to-PDF.* This gives me a file. I can save the file, back it up to my NAS, search for it later, keep it with other files from a project where it was useful, and otherwise hang onto it. This is so common, in fact, that it's gone from being an obscure thing you could do with a Postscript-to-PDF converter or (before the adware/Ask toolbar scandal) the installing the CutePDF virtual printer. Modern OSes bundle a PDF printer, and print dialogs understand that you want to "Save as PDF". Google Docs and Office 365 editors allow downloading a document as a PDF. I totally agree that a dynamic, interactive page or a comment section is not compatible with this model of usage. There's a lot of consumption of endless feeds, and a lot of one-time video views that also don't make sense to save as offline files. However, the web for creators, where people write articles that are worth hanging onto, has a definite place for PDFs. [1]: http://css-tricks.github.io/The-Printliminator/ http://css-tricks.github.io/The-Printliminator/
- derefr 5y ago> When you find a page [...] and you want to make sure it's available to you later, what do you do? Instead of doing a bad and lossy job of archiving the page myself, I notify† our friendly neighbourhood archivists at the Internet Archive of the page; and they then do the best, most lossless job of preserving the page that they're able, given their cumulative experience. † http://blog.archive.org/2017/01/25/see-something-save-something/ http://blog.archive.org/2017/01/25/see-something-save-someth... As a side-benefit, they also then take care of keeping the archive they've made around and available online in perpetuity, with no additional marginal effort on my part. The same can't be said for something in my own "private collection."
- deleted 5y ago[deleted]
- supperburg 5y agoThis reminds me of the guy who said drop box was stupid because he could set up an ftp server. It’s the exact same argument. People understand PDFs, they are extremely common in the academic and business world as “digital paper” standalone documents. Hypothetically, anything in memory can be made into a file but in this scenario what matters is the practical goal of people actually using these files. I think it makes sense for the web to be made up of discreet primitives not only so that the web can be browsed in an intuitive and frictionless way but also because it lends itself to being backed up and easily re-hosted.
- camgunz 5y agoYou got nerd sniped by the HTML vs. PDF format thing and missed the entire point of TA: > Isn’t it a good thing that we enjoy rapid progress? To the extent that we get to enjoy things like YouTube and sandspiel, yes! But to the extent that we want the internet to be a place where we can work and live and think and communicate free of malware, surveillance, dark patterns and the insidious influence of advertising, the answer is, empirically, sadly, no. The web has become ad-corrupted hand-in-hand with growth in technological capability, and the symbiotic relationship between web and browser means they feed on each others’ churn. Ads demand new sources of novelty to put themselves on, so the web expands continually, the specs grow in complexity, the browsers grow in sophistication, the barrier to entry grows ever higher, the vast cost of it all demands more ad revenue to fund it... and thus the perpetual motion machine is complete.
- cxr 5y agoThe author does identify a problem, and so you want to focus on that. That's fine. There is the issue of triviality, however. The problem described is widely felt, and also widely discussed. We already know this stuff to be a problem. For the piece to be worthwhile, then, it should do something that is not present in the other instances where the topic has been raised. It should articulate (or at the very least exhibit, without necessarily articulating) a solution for us. It doesn't. A bad remedy to a genuine problem does not yield a solved problem.
- slashdot2008 5y agoThe author brings a solution, it is to publish documents in PDF instead of HTML.
- grishka 5y agoPDFs aren't really meant to be read off a screen, they're much better suited for stuff that's meant to be printed out. And you can have a single self-contained file with a webpage, it's called a "web archive", with .mhtml extension.
- Tomte 5y ago> Base64 your images […], put your CSS in the HTML page Is there a tool that does those two things (or at least the first one) and that can be used by non-programmers (command line use is fine, a Python library would not be)?
- gildas 5y agoYou can use SingleFile for this, see https://github.com/gildas-lormeau/SingleFile/ https://github.com/gildas-lormeau/SingleFile/
- 1vuio0pswjnm7 5y ago"I come to hacker news to engage with thinkers, not just read a published article from a single author." And how many websites today are anything like HN, in terms of relative simplicity, e.g., no images^1, 3rd party requests or ads, only a tiny bit of (gratuitous)^2 JS. 1. I do not particpate in the voting scheme but I could vote from the command line if I wanted to. I use a text-only browser so the grey, fading text gimmick is irrelevant. I see all comments and treat them according to the thinking not the voting. 2. If we exclude the .ico and a .gif There seems to be a double-standard, for lack of a better term, where many HN commenters and voters appear to work for companies that make websites with tracking and ads and various gimmicks targeted at "non-thinkers" which are nothing at all like HN. Whatever these commenters and voters see and appreciate in HN they are not working to bring it to the rest of the web. I seriously doubt they comment and vote on HN out of fear of so-called "power users" or a belief that the HN type of simplicity could become more popular and threaten their jobs that depend on surveillance, online ads and a non-thinking audience of "powerless" users. Rather, a more rational explanation might be that they see some value in a website that shows no ads and generally uses no gimmicks; that's something to think about. "PDF web" may not make sense to many folks who have invested heavily in JS and Big Tech web browsers, but Postscript is arguably more elegant than Javascript. "Thinkers" usually like FORTH. https://en.m.wikipedia.org/wiki/Display_PostScript https://en.m.wikipedia.org/wiki/Display_PostScript The tracking section mentions the Abe Vigoda status page. http://www.abevigoda.com/ http://www.abevigoda.com/
- novok 5y agoSounds a lot like epub.
- anigbrowl 5y agoHTML can easily be offline-able. Sure - if the publisher cares. From the user's standpoint, the safe assumption is that they don't. Of course PDF is No Good for many contexts, but for any sort of long-form document that is primarily meant to be read, it's so often better. Also, if something is available in pdf, I can be moderately sure that someone else took the time to make sure it would be formatted correctly and print out OK.* If it only exists in HTML it's more of a roulette wheel experience. * Unless some graphic designer thought 'gee this report would look so cool if the cover pages were black or some other highly saturated block of solid color.'
- stjohnswarts 5y agoso because someone chooses to publish their website in an open format that they prefer "it's dumb" because they don't agree with you.