4 ms·
It's something of a UX question too. One could design a page that visually loaded, then jumped to a redirect after 10 seconds on the page. But who would? The
by ethbro 6y ago
It's something of a UX question too.
One could design a page that visually loaded, then jumped to a redirect after 10 seconds on the page. But who would?
The primary approach is always event-based, because most pages do that sanely.
If not... the best approach I've found is looking for sentinel elements.
Essentially, something that only matches once the website is de facto loaded (regardless of events). Sometimes it's a "search results found" bit of text, sometimes a first element. But more or less, "How do I (a human) know when the page is ready?"
- luckylion 6y ago> But who would? Affiliate networks (the shadier they are, the more likely they will), because they are weird and load third party tracking beacons in transitional pages and want to make extra sure that the beacons (who also can redirect multiple times) have been loaded. To add to the fun, they're also adding random new tracking domains (to avoid being blocked, I assume), so you can't even say whether you expect some domain to be transitional or final to increase your confidence in what you measure. You're right though, looking for elements is a pretty good way if you know the page you're checking. If you're going in blind, you can still look for things they probably have (e.g. <nav>, <header>, <section> etc), but I haven't found any that are reliably on a "real" page and reliably not on a redirect page.
- ethbro 6y agoThat's a use case I haven't encountered, nor considered! Most of my work is making known transitions (e.g. page1 to page2) work reliably, so I have the benefit of knowing the landing page structure. If you're crawling pathological, client-side redirect chains, maybe do pattern-matching scans on loaded code for the full set of redirect methods? There's only so many, and includes / doesn't-include seems a fair way to bucket pages.
- luckylion 6y agoYeah, we had been doing that initially and found that there are lots of imaginative ways to use e.g. refresh meta-tags that browsers do accept but we did not (e.g. somebody might say content="0.0; url=https://example.com/"* https://example.com/"*) and more and more networks and agencies switching to JS-based redirects lead to a headless browser being easier in the end, despite dealing with these specific issues. A simple self.location.href = ...* is still doable (-ish, because I've seen conditional changes that were essentially if(false)... to disable a redirect, which we obviously didn't consider when pattern matching), but once they include e.g. jQuery (and some do on a simple redirect page) it got far too complicated.