4 ms·
Happy to add this. Do you happen to know if this affects systems doing the scraping for social previews, which are already parsing meta tags?
by statico 5y ago
Happy to add this. Do you happen to know if this affects systems doing the scraping for social previews, which are already parsing meta tags?
- nhoughto 5y agoYou could probably just do both? Once a redirect has fired, off you go..?
- Klonoar 5y agoCurious if there's reasoning against just checking the user agent and sending back the HTML with OG tags for them, but then straight up HTTP redirecting in all other cases? I feel like any reputable agent calling for open graph stuff is going to be labeled correctly, though I'm open to being wrong.
- spiffytech 5y agoIt'll depend on the system doing the scraping. If it's just fetching the HTML with curl or similar, using the meta tag shouldn't be any different. I'm not familiar with HTML parsers that follow meta redirects. But if the scraper is using Puppeteer, either approach will result in the redirect being followed and the Open Graph data being ignored. For example, Microlink follows the JS redirect and seems to ignore your Open Graph data: https://microlink.io/meta?url=https%3A%2F%2Fnews.ycombinator1.com%2Fitem%3Fid%3D30181167 https://microlink.io/meta?url=https%3A%2F%2Fnews.ycombinator... Scraped data at this time: https://i.imgur.com/pRLIt0S.png https://i.imgur.com/pRLIt0S.png This is normally the behavior you want from a scraper: if someone pastes a bit.ly / t.co link you want the link preview to show data on the other side of the redirect. You could try to figure out how to distinguish between a scraper and a human operator, but that'll be tough, since people work very hard to make Puppeteer look 100% legit.
- statico 5y agoOK, meta tag added!
- andjd 5y agoThis is probably and edge case you don't need to worry about. For the most part, you can rely on the user agent for a scraper that cares about these meta tags to identify itself as a bot (e.g. twitterbot). Since meta tags are machine, not user, facing, I don't think there's much of any benefit in trying to serve the meta tags to a bot that is intentionally trying to emulate a human.
- hyperhopper 5y agoWhile this is how web scrapers work, this post is targeted at url unfurlers, which have different goals and almost never follow redirects. Facebook, slack, etc, have documentation publicly available on how their unfurlers work: most effectively just curl the html then parse meta tags
- junon 5y agoThey follow redirects but only authoritative ones from the server. Not meta tag redirects. Just to be clear.
- hyperhopper 5y agoHence the "almost"