3 ms·
Hi, we're trying to lower the requests:pageview ratio in general, but for what it's worth this article essentially: - ignores the vast majority of "image servi
by tristan9 5y ago
Hi, we're trying to lower the requests:pageview ratio in general, but for what it's worth this article essentially:
- ignores the vast majority of "image serving" (most is handled by DDG and our custom CDN)
- the JS fragments thankfully should load only on first visit and then get aggressively cached by DDG/your browser
One of the pain points is that there are a lot of settings for users to decide what they should or shouldn't see (content rating, original language of origin, search tags, etc) and some are already specifically denormarlized (when querying chapter entities, ES indices for those contain some manga-level properties to avoid needing to dereference that first too) -- however this also makes caching substantially less efficient in many places, alas
Thanks!
- jiggawatts 5y agoHi, I'm a performance tuning expert, and this thread piqued my interest. The first thing that I noticed is that even with caching enabled, you're loading "too much data". After loading the main page and then clicking one of the tiles, there are several JSON API calls. Here's an example, 195 kB transferred (528 kB size): https://api.mangadex.org/manga/bbaa17c4-0f36-4bbb-9861-34fc8fdf20fc/feed?limit=96&includes[]=scanlation_group&includes[]=user&order[volume]=desc&order[chapter]=desc&offset=0&contentRating[]=safe&contentRating[]=suggestive&contentRating[]=erotica&contentRating[]=pornographic https://api.mangadex.org/manga/bbaa17c4-0f36-4bbb-9861-34fc8... Oof. Half a megabyte of JSON! Ignore the network traffic for a moment, because GZIP does wonders. The real problem is that generating that much JSON is very "heavy" on servers. Lots and lots of small object allocations, which gives the garbage collector a ton of work to do. It's also expensive to decode on the browser for similar reasons. On my computer, this took a whopping 455ms to transfer, nearly half a second. That results in a noticeable latency hit to the site. In my consulting gig I always give developers the same advice: "Displaying 1 kilobyte of data should take roughly 1 kilobyte of traffic". In other words, there's isn't 500 KB of text anywhere on that page! A quick cut & paste shows about 8 KB of user-visible text in the final HTML rendering. That's a 1:60 ratio of content-to-data, which is very poor. I bet that behind the scenes, this took a heck of a lot more back-end network traffic and in-memory processing to generate. Probably tens to hundreds of megabytes of internal traffic, all up. This is one of the core reasons most sites have difficulty scaling, because for every kilobyte of content output to the screen, they're powering through megabytes or even gigabytes of data behind the scenes. Can this API query be cut down to match what's displayed on the screen? Can it be cached for all users? Can it be cached precompressed? Etc...
- tristan9 5y ago> The real problem is that generating that much JSON is very "heavy" on servers. Lots and lots of small object allocations, which gives the garbage collector a ton of work to do. It's also expensive to decode on the browser for similar reasons. For what it's worth, this isn't generated live but a mix of existing entity documents Most of it is page filenames which indeed could be made optional and fetched only by the reader, but that'd be us actively nulling them out in the returned entity, since they are there in the ES documents for the chapters (a manga feed like this being a list of chapters)
- jiggawatts 5y agoYou're basically dumping down a database to the web browser, including all of the internal metadata that's likely irrelevant to rendering the HTML. For example, user role memberships: { "id": "c80b68c5-09ae-4a50-a447-df7c5a4a6d01", "type": "user", "attributes": { "username": "kinshiki", "roles": [ "ROLE_MEMBER", "ROLE_GROUP_MEMBER", "ROLE_POWER_UPLOADER" ], "version": 1 } } Also record timestamp dates like created/changed, along with contact details that may be revealing sensitive info: "attributes": { "name": "SENPAI TEAM", "locked": true, "website": "https:\/\/discord.gg\/84e3j9b", "ircServer": null, "ircChannel": null, "discord": "84e3j9b", "contactEmail": "senpai.info@gmail.com", "description": null, "official": false, "verified": false, "createdAt": "2021-04-19T21:45:59+00:00", "updatedAt": "2021-04-19T21:45:59+00:00", "version": 1 } But let's just go back to your response: > Most of it is page filenames which indeed could be made optional Do that! If you strip them out, the 529 kB document shrinks to 280 kB, which hardly seems worth the hassle, but when gzipped, this is a miniscule 13 kB! This is because those strings are hashes, which significantly reduces their compressibility compared to general JSON, which usually compresses very well. It's basic stuff like this that can make a website absolutely fly. Avoid giving computers unnecessary, mandatory work: https://blog.jooq.org/many-sql-performance-problems-stem-from-unnecessary-mandatory-work/ https://blog.jooq.org/many-sql-performance-problems-stem-fro...
- deleted 5y ago[deleted]
- the8472 5y agoOne issue I see is that flipping back and forth between chapters reloads images from different URLs which means they're uncachable. I guess that's somehow related to the mangadex@home thing, but if the URLs were generated in a more deterministic manner (keyed on some client ID + the chapter being loaded) then the browser could avoid redundant traffic.
- tristan9 5y agoThat's very close to how MD@H works, but it also has a time component and tokens are not generated by our main backends, so it'd require a separate internal http call per chapter
- the8472 5y agoAnother thing. For each page that's being loaded there's a report being sent. Instead this could be aggregated (e.g. once a second) and then processed as a batch on the server side which should be faster. And if your JS assets are hashed then you can add cache-control: immutable so that a browser doesn't have to reload them when the user F5s.
- maxk42 5y ago> the JS fragments thankfully should load only on first visit and then get aggressively cached by DDG/your browser According to Alexa you have a 46.4% bounce rate. [1] When 46% of your users aren't coming back, how does 31 round-trips to your server for 100% of first-page visitors save anyone time or bandwidth? Your pageviews per visitor is 6.8, meaning the 53.6% that stick around view an average of 11.8 pages each. Even if there are zero subsequent js requests on other pages (clicking a random page I see 8) you would be generating 31 requests up-front to save 10.8 subsequent requests for about half of your users. (And again - in any scenario where the number of js fragments transferred on subsequent requests >= 1 even this benefit goes out the window.) How does that save you or your users bandwidth, server load, or other overhead? The scale is not quite linear, but generally speaking, if you get your number of requests down from > 100 to < 5, you'll be able to handle around 20x the traffic with the same number of web-facing servers. Or alternatively the same amount of traffic with around 1 / 20th the servers. Would that have a material effect on your costs? [1] https://www.alexa.com/siteinfo/mangadex.org https://www.alexa.com/siteinfo/mangadex.org
- tristan9 5y agoDefinitely needs optimising for user experience indeed! However the serving of this JS has nearly no cost to us (as they are cached at the edge by DDoS-Guard and the frontend is otherwise entirely static on our end)
- rowanG077 5y agoIt does have a cost it's just hidden. The cost is that it increases your bounce rate because of bad UX.