20 ms·
I've said this before, but it bears repeating: Moby Dick is 1.2mb uncompressed in plain-text. That's lower than the "average" news website by quite a bit--I ju
by cheezymoogle 8y ago
I've said this before, but it bears repeating:
Moby Dick is 1.2mb uncompressed in plain-text. That's lower than the "average" news website by quite a bit--I just loaded the New York Times front page. It was 6.6mb. that's more than 5 copies of Moby Dick, solely for a gateway to the actual content that I want. A secondary reload was only 5mb.
I then opened a random article. The article itself was about 1,400 words long, but the page was 5.9mb. That's about 4kb per word without including the gateway (which is required if you're not using social media). Including the gateway, that's about 8kb per word, which is actually about the size of the actual content of the article itself.
So all told, to read just one article from the New York Times, I had to download the equivalent of ten copies of Moby Dick. That's about 4,600 pages. That's approaching the entirety of George R.R. Martin's A Song of Ice and Fire, without appendices.
If I check the NY Times just 4 times a day and read three articles each time, I'm downloading 100mb worth of stuff (83 Moby-Dicks) to read 72kb worth of plaintext.
Even ignoring first-principles ecological conservatism, that's just insanely inefficient and wasteful, regardless of how inexpensive bandwidth and computing power are in the west.
EDIT: I wrote a longer write-up on this a while ago on a personal blog, but don't want it to be hugged to death:
http://txti.es/theneedforplaintext http://txti.es/theneedforplaintext
- soared 8y agoI don't think thats a meaningful comparison. Moby Dick is a book, written by 1 guy and maybe an editor or two. NYT employs 1,300 people. When you read a book all you get is the text. NYT has text, images, related articles, analytics, etc. Moby Dick doesn't have to know what pages you read. NYT needs to know how long you spent, on which articles, etc. They need data to produce the product and you can only achieve that with javascript tracking pixels (Server logs aren't good enough). If Moby Dick was being rewritten and optimized every single day it would be a few mb. Its not, so you can't compare the two. Yes NYT should be lighter, no your comparison is not meaningful. A better comparison would by Moby Dick to the physical NYT newspaper.
- zajd 8y ago> They need data to produce the product [citation needed] There's no reason they need to use Javascript to track user behavior down to "how long have they read this article".
- throwawaymath 8y ago> NYT needs to know how long you spent, on which articles, etc. They need data to produce the product and you can only achieve that with javascript tracking pixels (Server logs aren't good enough). No they don't. They really don't need to know any of that. They don't even get a pass on tracking because they're providing a free whatever - I pay for a subscription to the NYT. The business, or a meaningfully substantial core of it, is viable without tracking. It would be nice if the things I pay for didn't start stuffing their content with bullshit. What and who do I have to pay to get single second page loads? It's not a given that advertising has to be so bloated and privacy-invasive. Various podcasts and blogs (like Daring Fireball) plug the same ad to their entire audience each post/episode for set periods of time. If you're going to cry about needing advertising then take your geographic and demographic based targeting. But no war of attrition will get me to concede you need user-by-user tracking. You want me to pay for your content? Fine, I like it well enough. You want to present ads as well? Okay sure, the writing and perspectives are worth that too I suppose. But in addition to all of this you want to track my behavior and correlate it to my online activity that has nothing to do with your content? No, that's ridiculous.
- koolba 8y agoInterestingly if you pay them, and thus are logged in when you view an article, then they can better track you. In contrast if you never sign up, disable JS, and periodically clear your cookies, then the entire site works fine and none of the third party trackers work. At best they can link your browser user agent and IP to a hit on the server side.
- deleted 8y ago[deleted]
- closeparen 8y agoNYT needs to produce and recommend content that people find engaging to continue earning their subscription dollars. The idea that tracking is purely or primarily there to support a business model of selling user data is a strawman invented by self-righteous HNers. You need to know what parts of your product are effective to make it competitive in today’s marketplace.
- paulsutter 8y agoYou’re right about the problem: web pages tend to scale with the size of the organization serving them, not the size of the content. But this is the failure, not a defense. It’s a big problem on mobile and the reason I read HN comments before the article. > NYT employs 1,300 people
- mvdwoord 8y agoThey need data to produce the product and you can only achieve that with javascript tracking pixels (Server logs aren't good enough). I disagree. They need journalists, and they need to find some way to monetize. Your argument implies there is no other way than to add user tracking. Sure, images take up space, but I refuse to believe the current way of the web is the only viable option.
- cheezymoogle 8y agoA random archive of the New York Times frontpage in 2005 is 300kb. Articles were probably comparable in size. Are you honestly saying that the landscape of the internet and/or the staffing needs of the NY Times has changed so drastically that they actually needed a 22x increase in size to deliver fundamentally text-based reporting?
- kazagistar 8y agoI mean, if most of that is a few images, then those images could just be bigger today for nicer screens and faster internet. Not that that is the case.
- goatlover 8y ago> NYT needs to know how long you spent, on which articles, etc. They need data to produce the product and you can only achieve that with javascript tracking pixels (Server logs aren't good enough). This just seems like such an abuse of what the web was meant to be. I can imagine the horror people in the 90s would have experienced if they new what JS was going to be used for when perusing news sites. Sometimes I wonder if it would have been better keeping the web as a document platform without any scripting, and creating a separate one for apps. Anyway, an alternative model news sites could use is to let users choose which content they want to pay for. That's a way to track which content users prefer.
- jstarfish 8y ago> I can imagine the horror people in the 90s would have experienced if they new what JS was going to be used for We understood the horror no less than we do now. Javascript in the 90s gave us infinite pop-ups, pop-unders, evasive controls, drive-by downloads and otherwise hijacked your browser and/or computer. There's a reason Proxomitron and other content blockers hit the scene by the early 2000s-- the need to shut that shit off was clear.
- mahranch 8y agoYeah, people under 33 seem to have a romanticized view of the internet. They believed there was no ads and it was flush with the kinds of content we enjoy today. Nope. Content existed but it was scarce/thin. Many of the internet users just stayed on AOL/Prodigy/Compuserve and never left to explore the WWW side of things. Those service providers were essentially national level BBSs. There was no youtube, wikipedia, itunes or reddit. No instagram, twitter or google earth. The internet was basically geocities where most webpages were fan pages or pages/forums about niche interests. I think people want to believe that because they believe that if ads were to disappear off the internet tomorrow, nothing would change. They don't realize that ads subsidize the content they consume, whether it's a youtube video they're watching or a reddit thread, ads are paying for that content. Nothing is free.
- 8y ago
- kingbirdy 8y agoThe New York Times existed for 145 years from its founding in 1851 to the creation of its website in 1996, and it got by just fine without tracking pixels in all those years.
- freediver 8y agoThis is a fallacy. Humans got by just fine without smartphones, Internet, electricity for thousands of years. While you could do just fine without those things today it is impractical. Times change (pun intended).
- jaredklewis 8y ago> NYT needs to know how long you spent, on which articles, etc. They need data to produce the product and you can only achieve that with javascript tracking pixels (Server logs aren't good enough). Nonsense. I subscribe to the NYT so that I can read the news. Nothing about that necessitates tracking which users read which articles. If the NYT uses page view data for anything other than statistics for their advertising partners, it's a shame. I don't want the NYT to tailor write their articles to maximize page views, time spent, or any other vanity statistic; if I felt like reading rage bait fed to me by an algorithm personally customized for all my rage buttons, there is plenty of that elsewhere. NYT's differentiating factor is that they are one of the few businesses left that pays people to conduct actual journalism. If they give up on that, then I imagine their customers will just go to buzzfeed or wherever.
- lazyasciiart 8y agoSo, how does the NYT style section fit into your idea of 'actual journalism'?
- jaredklewis 8y agoIt doesn't. Not every word printed in the NYT is journalism, but it doesn't change the fact that they are one of the few websites that have any journalism at all. If the NYT cut their paper down to just the style section, horoscopes, and other garbage, they would be just another Buzzfeed and are probably not equipped to compete. On the other hand, Info Wars, Mother Jones and friends offer publications with basically no journalism at all. That's the space the NYT, WSJ, Miami Herland, Chicago Tribune, and so on fill. They do Pulitzer prize worthy reporting. If these papers become run-of-the-mill click farms, I'm sure silicon valley will run them out of business, as rage bait is not really their core competency.
- ryandrake 8y ago> Moby Dick is a book, written by 1 guy and maybe an editor or two. NYT employs 1,300 people. Totally irrelevant. Why should the number of employees in the company have any bearing on the size or cost of the product? Ford has 5x as many employees as Tesla. Should their cars be 5x as big or 5x more expensive? > NYT needs to know how long you spent, on which articles, etc. They need data to produce the product They may want this but they don’t need it. They successfully produced their product in the past without it. > If Moby Dick was being rewritten and optimized every single day it would be a few mb. Irrelevant and likely false. If anything, books and other text media tend to get smaller after subsequent editing and revising. > A better comparison would by Moby Dick to the physical NYT newspaper. Comparing a digital text product (Moby Dick) with a digital text product (a NYT article) is as close as it gets.
- dwild 8y ago> Totally irrelevant. Why should the number of employees in the company have any bearing on the size or cost of the product? Ford has 5x as many employees as Tesla. Should their cars be 5x as big or 5x more expensive? If the cost or the size wasn't a constraint, for sure Ford would build a car 5x as big or 5x as expensive. The website size isn't a constraint here, if it was, they would works on it and make it smaller. It's only a constraint for highly technical people here. Currently at my job I'm optimizing some queries that takes way too long. It has been like that for years but we hit a wall recently, our SQL Server can't take it anymore. I always found it stupid that it took so long to optimize it... but at the end of the day, the clients just didn't care that it took 3 seconds to load the page. I could be working on more features right now, something that the client actually care about. What makes the number of employees relevant to the size? Well if you were the only one building that website, you would know everything about it right? You would always use the exact same component, reuse everything you can, you already know every single part of the code. Add a second employee, now you don't know exactly what he does, you do know some of it but some time you forget and you may duplicate something or do it badly or whatever. At one point, something is just too big to be understood by a any single employee and you get code badly reused, stuff that serve no direct purpose too but make maintenance easier, etc... You never decrease the size simply because it's never worth it to but each and every single one of the employee add stuff to it.
- LolNoGenerics 8y ago> NYT needs to know how long you spent, on which articles, etc. They need data to produce the product and you can only achieve that with javascript tracking pixels (Server logs aren't good enough). That's analytics. The marketing stuff that is causing the bloat isn't doing that much in comparison. Trackers are often coded very wasteful and are redundant by nature. You can have easily dozens to hundreds of them all doing the same stuff just with different APIs. It is insane and out of control and has absolutely nothing to do with gathering insights about your app and improving. It is pure external 3rd party marketing.
- kdl20kxkrk 8y ago> NYT employs 1,300 people. See 2 in the article. A sample: “As Graeber observed in his essay and book, bullshit jobs tend to spawn other bullshit jobs for which the sole function is a dependence on the existence of more senior bullshit jobs:“ I work for one of the vacation rentals No reason private owners couldn’t be doing the work over email. But certain fetishized models of doing, in this case cloud and web apps, get the focus. It’s all for eyeballs and buy in at scale to justify the bullshit. “Look everyone is watching us talk up this shit! Better keep justifying it, bringing them into our flock!” It’s turned us all into corporate sycophants. Religious conviction isn’t limited to belief in sky wizards Anything sufficiently magical to the layman will instill blind allegiance And despite all the smart people here, life as is seems magical and there’s a lot of blind buy-in
- Joeri 8y agoNewspapers did just fine for centuries without tracking. The business is viable without tracking.
- deleted 8y ago[deleted]
- _louisr_ 8y agoThis is both an appeal to people's universal appreciation of efficiency, and a weak denunciation of the modern web. Your argument is 1) that a website's value is the number of words on the page, and 2) that raw text is the highest value data that can be transmitted over the internet, and 3) that inefficiency and wastefulness of bits is a bad thing First off, you need to defend your first two assumptions. Don't websites do a lot more than display text? Does HTML/markup not have a magnitude more value than raw text? And how exactly is being inefficient with something that is abundant a tautologically bad thing?
- cheezymoogle 8y ago1. A website's value is the amount of information that it provides to the end user. I would argue that the entirety of the works of Shakespeare or the 1911 Brittanica Encyclopedia provides more information than a high-definition picture of Donald Trump grimacing at EU leaders or an autoplaying ad for Doritos Locos Tacos NEW AT TACO BELL. As far as supplemental uses, at the end of the day, people are using plaintext to communicate with other people. Images and videos are secondary. If that ever changes, then society is already doomed as literacy is fundamental to the maintenance of technology. 2. Raw text is the highest value data/information return that can be transmitted over the internet. There's a reason that Morse Code and APRS are still around: they're reliable, appropriate tech, and require little to no middlemen outside of the transceivers themselves. 3. If wantonly (namely, for no enduringly good reason) increasing the amount of entropy in the universe isn't tautologically bad to you, then I really doubt that any argument would sway you to the contrary. Concerning the relative value of HTML and CSS, yes, you could argue that UX matters in that department, but even the most bloated static HTML/CSS page is going to pale dramatically in comparison to the size of what's considered acceptable throughput today.
- fyrabanks 8y agoIn response to #2, there is also a reason that photojournalism exists--unless you believe in a world where everyone imagines what important figures and historic events look like based solely on textual descriptions. This is completely ignoring the fact that I, nor most of society, would never attempt to receive the day's news via Morse Code.
- pc86 8y agoComparing the raw text of a fiction novel to the code of a website is a pretty asinine comparison, honestly.
- Pete_D 8y agoMaybe you'd prefer comparing the code of a website to the amount of useful content on the website, which OP also did. Taking "I'm downloading 100mb worth of stuff (83 Moby-Dicks) to read 72kb worth of plaintext" at face value, we could also say that 0.072% of the data transferred is useful, or, equivalently, that 99.928% of it is crap.
- pc86 8y agoYou don't have to download the typography of a physical book but it still plays a huge role in the readability and enjoyment of it. So I guess the typography of websites is "crap" because it has to be downloaded? It's a ridiculous apples-and-hammers comparison thinly veiled as an intelligent critique.
- cheezymoogle 8y agoI can put raw text on a kindle and read it. In fact, I frequently do. Comparing digital text to digital text is not apples-and-hammers.
- chickenfries 8y agoExactly, it’s not the NYTs fault that plain text compresses well compared to jpgs. Also, is a world where he NYT is subscriber only really preferable?
- cheezymoogle 8y agoThis is a textbook definition of a false dichotomy. There are other distribution models for digital news services. There are other methods for transmitting digital content. It's not an either/or situation.
- yathern 8y agoI just loaded up a nytimes[1] article too - and only weighed in at 1.0MB. For a 1000 word article. Subsequent reloads dropped it to ~1000KB. I don't think that's too bad, considering there are images in there as well. Now of course, I'm running an ad blocker. I assume the remaining MB that you noticed had come from advertising sources. In which case, bloat isn't the issue, ads are. [1] - https://www.nytimes.com/2018/07/31/us/politics/facebook-political-campaign-midterms.html https://www.nytimes.com/2018/07/31/us/politics/facebook-poli...
- cheezymoogle 8y agoAre you also running noscript? I'm running a DNS sinkhole and still get 3mb on reloads.
- iforgotpassword 8y agoThat distinction is nonsense. The ads are part of the page and are no more or less bloat then the rest of the useless junk that gets embedded. It's deliberately put there by the NY times, they don't end up there by accident.
- rhizome 8y agoweighed in at 1.0MB. For a 1000 word article. Subsequent reloads dropped it to ~1000KB. You really can't beat savings like that.
- welly 8y agoThat'll knock dollars... no, cents... no half-cents off his ISP bill!
- zanny 8y agoI must say, despite what I imagine is a good bit of traffic the instant load times on your site were a joy to behold. I never get to experience that kind of speed on the "modern" web. Even HN loads orders of magnitude slower than that.
- cheezymoogle 8y agoNot my website, but please send along thanks and awareness to @thebarrytone on Twitter!
- ChuckMcM 8y agoI like this rant, you should go the next step: All you need to 'fix' this is a fast loading news website that gets enough paid subscribers to earn enough margin from subscriptions that you can pay for a news staff, an office, and various overheads. That is a longish way of saying that 99.9% of the overhead in any modern web site can be traced almost entirely to the mechanisms by which that web site is attempting to extract value from you for visiting/reading. If people would visit with a 56K modem and deal with a 3 - 10 second page load, then that is the bar. And any spare bandwidth you might have is available for the web site to exploit in some way to generate revenue. The more bandwidth between you and them, the more ways they can come up with to exploit that bandwidth for additional surveillance, ads, or analytics that will get them more money. When you are the customer, which to say it is your purchasing of a subscription or articles is the only revenue the site needs in order to survive, then the things that retain you as a customer have the highest priority (like fast page load times, minimal bandwidth usage). But when you are a data cow, a random bit of insight into a picture much bigger than you can comprehend, a pixel in a much larger tapestry, or an action droplet in a much larger river of action. Well then there isn't really any incentive to make your life better, as long as the machine we have milking you for data can get even a couple of molecules more of that precious data milk without scaring you out of the barn. Well we'll build right up to that limit.
- TeMPOraL 8y agoHence the rant mentioning Molochian economy, under which we operate. And that reference explains in depth why this is a very hard problem. I like the rant too. Except maybe the bit about sending content at the speed of humans - I for one would like to take lightweight, bullshit-free content as fast as it can be sent, to pipe it to further processing on my end, in the never-ending quest to automate things in my life.
- cheezymoogle 8y agoI mean, if you're just reducing the content even further, just request that they make the reduction possible server-side and everybody wins.
- gooseus 8y agoI can't agree more with the points you make, I've spent a decent amount of time and effort reducing the overhead of my blog, for example - https://goose.us/thoughts/on-the-purpose-of-life/ https://goose.us/thoughts/on-the-purpose-of-life/ That page includes images and "embedded" youtube videos, but loads 454kb of 3,187 words with 10 requests in 450-600ms - if anyone has any suggestions on how to reduce that further, I'd love to hear. Going to bookmark that txti.es service for the future, definitely seems useful for publishing simple content without needing to fit it into a blog theme.
- overcast 8y agoI'd start with the 100+KB PNG images. None of those pictures are complex, they can be compressed much more.
- gooseus 8y agoI put them all through ImageOptim, so not sure if those PNGs can compress anymore... I agree on those complexity graphs since they're actually scanned from Sean Carroll's book and edited in Pixelmator - if I was better with graphics I could probably recreate them as SVG or in an editor and make it only a couple kb. Same for the MinutePhysics screen grab, though I couldn't bring myself to attempt a poor recreation of their work and couldn't get rid of the weird pink gradient background which probably prevents better compression. The Youtube cover images are still being pulled in from youtube - I debated downloading/compressing/hosting them myself, but figured I'd gain more from downloading them from the separate domain since I still think there is a default limit of connections the browser will open to a single domain at the same time.
- overcast 8y agoThat's because PNG is lossless, use JPG, and you'll dramatically cut down the space.
- irrational 8y agoHave you tried opening the image(s) in Photoshop and try outputting them at different resolutions and compression settings?
- chickenfries 8y agoYou want to compare plain text to a newspaper, which even in paper form, people expect to contain pictures and advertisements.
- krrrh 8y agoI support your effort to make Moby-Dicks the football-field-like unit of measurement for text-focused data. It’s close enough to the 1.44 MB floppy disk to handle easy mental conversion of historical rants, and half of the people reading this have probably never held one of those. I still remember downloading a text version of a 0.9 Moby-Dick book from some FTP site and carrying it around on a floppy so I could read it on whatever computer was handy. That aside, the most shocking part of your analysis is how inefficient the nytimes was at caching resources for your reload.
- userbinator 8y agoFor a rather more technical comparison, 4,600 pages is more than the size of Intel's x86/64 Software Developer's Manual, which is ~3-4k pages.
- Gravityloss 8y agoI don't know how many people here read usenet or were on old mailing lists. You could have removed a some hard edges in usability and everything would have been connected to phones with monochrome text displays back in the nineties already. They didn't even provide decent email experience. But instead we got the technology developing through some ringtone stuff advertised on TV. I guess it's something that you can instantly show to your friends.
- cheezymoogle 8y agoThat's one of the wildest things! We had the technological capacity to run Unix systems with several hundred users simultaneously and access at the speed of thought with 26kbps modems back in 1992, complete with instant messaging and personal directories! What happened?!
- TeMPOraL 8y agoAnother wild thing like this is what you'll notice when you read up on Lisp machines. We had development environments in the 70s/80s that would seem magical today.
- dri_ft 8y agoIt's heartbreaking.
- Gravityloss 8y agoShows we are not limited by technology but somehow get distracted by other things.
- scroot 8y agoThe "other things" are the short-termism and appealing to the lowest common denominator that go with the pursuit of profit before anything else.
- 8y ago
- skepticmoron 8y agoYour math does not add up. It is 1 * (6.6 + 3 * 5.9) + 3 * (5 + 3 * 5.9) == 92.4. Unless you are speaking about approximations obviously.
- adamsea 8y agoIMHO you are confusing data with information with knowledge. And mixing mediums. You can't compare a novel - the plainest of plain-text mediums, with the front page online of a major news organization in 2018 - of course it will be interactive content, its an entirely different medium, a different market, and different sets of user expectations and competition. https://www.quora.com/What%E2%80%99s-the-difference-between-data-information-and-knowledge-in-machine-systems-Or-is-there-a-difference https://www.quora.com/What%E2%80%99s-the-difference-between-... DATA: a "given" or a fact; number; picture represents something in real world raw materials in production of information INFORMATION: Data that have meaning in context Data related Data after manipulation KNOWLEDGE: familiarity, awareness and understanding of someone or something acquired through experience or learning it is a concept mainly for humans unlike data and information.
- liquidwax 8y agoI first realised how heavy these pages are when I disabled javascript. Things load in the blink of an eye. \Most\ pages work and the web remains largely usable.
- cjohansson 8y agoYes this is my experience as well, JavaScript is often the key antagonist. Unfortunately many websites require JavaScript to function
- freediver 8y agobbc.co.uk will load just fine without JS and actually be more enjoyable (IMO) than JS version. cnn.com fails miserably without JS.
- dri_ft 8y agoMy pet comparisons for everything being too big nowadays are Mario 64 (8mb!), Super Mario World (512kb!), and Super Mario (32kb!!).
- kevin_thibedeau 8y agoNoScript cuts the bullshit down to 1.35 MB with all scripts blocked and it's still readable. I can barely tolerate the web without it.
- kodablah 8y agoThis dovetails into an idea I had [0]. Basically just client side scrape the web as it's used and deliver people this plain text and simple forms. It would have a maintained set of definitions and potentially even logic to put a better "front" on all this bullshit. It's like reverse ad block where you only whitelist some content instead of blacklisting it. You could argue sites will get good at fighting it, but if used enough by the common user, they'd just alienate them (e.g. my scrape/front for Google search makes it clear which results the app has a friendly scrape/front for). 0 - https://github.com/cretz/software-ideas/issues/82 https://github.com/cretz/software-ideas/issues/82
- cheezymoogle 8y agoI've often toyed with the idea of using multi-user systems over SSH running Gopher and Lynx to achieve something like this. In the process, it would also decentralize communities and establish digital equivalents of coffee shops (i.e. places to work in public and meet strangers)--basically SDF, but deployable on Raspberry Pis with more modern userland toys (i.e. software actually designed to be multi-user on the same system). [1] https://en.wikipedia.org/wiki/Gopher_(protocol) https://en.wikipedia.org/wiki/Gopher_(protocol) [2] https://en.wikipedia.org/wiki/SDF_Public_Access_Unix_System https://en.wikipedia.org/wiki/SDF_Public_Access_Unix_System
- kodablah 8y agoMy main reason for client side is to skirt legal troubles that can result from running a web-filtering proxy (not whether it's legal or not, but whether you will be in legal fights). Either way, needs to be as transparent as possible and as usable by the less-tech-savvy as possible. But that's really all it is, a web server (or an app, or an extension, or a combo) that serves you up the web looking like Craigslist. Would require strongly curated set of "fronts"/"recipes".
- Pete_D 8y agoSounds a bit like tedunangst's miniwebproxy[0]. I've been wondering about writing either something like it or a youtube-dl-like "article-dl" for my own use, but haven't quite been annoyed enough into doing it yet. [0] https://www.tedunangst.com/flak/post/miniwebproxy https://www.tedunangst.com/flak/post/miniwebproxy (self-signed cert)
- aalleavitch 8y agoMoby Dick doesn’t have any pictures or video. Audiovisual media has value.
- BLKNSLVR 8y agoAudiovisual media of value has value. Audiovisual media that I choose to donate my bandwidth towards downloading may have value. Audiovisual media in and of itself has no value. Audiovisual media that automatically loads, thus slowing down the loading of everything else, has negative value.
- conanbatt 8y ago> Moby Dick is 1.2mb uncompressed in plain-text How many times did you read moby dick online.
- lifeformed 8y agoThe 100mb pays the bills for the other 72kb of text.
- textmode 8y ago"I just loaded the New York Times front page. It was 6.6mb." ftp -4o 1.htm https://www.nytimes.com du -h 1.htm 206K For the author, 206K somehow grew to 6.6M. Could it have anything to do with the browser he is using? Does it automatically load resources specified by someone other than the user, without any user input? Above I specified www.nytimes.com. I did not specify any other sources. I got what I wanted: text/html. It came from the domain I specified. (I can use a client that does not do redirects.) But what if I used a popular web browser to download the front page? What would I get then? Maybe I would get more than just text, more than 206K and perhaps more from sources I did not specify. If the user wants application/json instead of text/html, NYTimes has feeds for each section: curl https://static01.nyt.com/services/json/sectionfronts/$1/index.jsonp where $1 is the section name, e.g., "world". The user can use the json to create the html page she wants, client-side. Or she can let a browser javascript engine use it to construct the page that someone else wants, probably constructing it in a way that benefits advertisers.
- smadge 8y agoI don’t think there is anything wrong with user agents downloading resources (like images and stylesheets) linked to by an html document. It is the providers, not the user agents, who have violated the trust of users by including unnecessary scripts, fonts, spyware, advertisements, etc.
- textmode 8y ago"I don't think there is anything wrong with user agents downloading resources (like images and stylesheets) linked to by an html document." Neither do I. For some websites, this is both necessary and appropriate. However, in cases where the user does not want/need these resources, or where she does not trust the provider, I do not think there is anything wrong with not downloading images, stylesheets, unnecessary scripts, fonts, spyware, advertisements, etc.
- userbinator 8y agobut don't want it to be hugged to death: Incidentally, this is also another reason for keeping pages small --- bandwidth costs. I remember when free hosts with quite miniscule monthly bandwidth and disk space allotments were the norm, and kept my pages on those as small as possible.
- freyr 8y agoBad comparison, Moby Dick never tracked your activity across the web and sold your data to advertisers.
- krzrak 8y ago> mb I think you didn't mean milibits, but megabytes (MB). 1.2 mb = 1.5e-10 MB.
- PurpleRamen 8y ago> to read just one article from the New York Times, I had to download the equivalent of ten copies of Moby Dick. But how many Mona Lisas are this?
- norepicycle 8y ago> As my father was a news and politics junkie (as well as a collector of email addresses) I feel like there's something I'm missing here: something like getting his hands on clever usernames?
- JoshuaRLi 8y agoWell said.