6 ms·
I have an idea for a project (I call it 'cannon' for now)[0] which would 'canonicalize' URLs and extract semantic information from it, ideally just by looking a
by karlicoss 5y ago
I have an idea for a project (I call it 'cannon' for now)[0] which would 'canonicalize' URLs and extract semantic information from it, ideally just by looking at the URL, without doing any extra requests. For example, a tweet URL usually encodes the tweet author and tweet ID; and by extracting such entities one could determine 'relations' between URLs. I'm using a simple prototype in Promnesia [1], a browser extension aiming to make the web browsing history more useful and aid knowledge management.
This effort is really ought to be shared, it's potentially a lot of manual work, and could benefit many projects. ClearURLs seems like one of the most promising existing projects doing similar stuff; have been meaning to approach the devs, feels like it's something we could cooperate on. Although ClearURL has a somewhat narrower scope, but still I feel like there is a potential to share.
[0] https://beepb00p.xyz/exobrain/projects/cannon.html https://beepb00p.xyz/exobrain/projects/cannon.html
[1] https://github.com/karlicoss/promnesia#readme https://github.com/karlicoss/promnesia#readme
- ForHackernews 5y agoHow would this be possible? Different sites have different ideas of what counts as a "canonical" URL: example.com/page1 vs example.org?page=2
- karlicoss 5y agoYep, that's kind of the main problem :) Hence the need for some manual curation. (e.g. ClearURLs seems to do it here https://github.com/ClearURLs/Rules/blob/master/data.json https://github.com/ClearURLs/Rules/blob/master/data.json) For 80% of sites just throwing away the query parameters work, for the rest sadly it's necessary to do more sophisticated normalizing. I'm also thinking that it might be possible by some simple machine learning, by looking at the corpus of existing URLs. E.g. if a human looks at a corpus of different URLs they would more or less guess what is useful, and what's tracking garbage, so perhaps it's possible to automate it with a high accuracy? Then, I also feel if it's paired with some UI to allow the user to 'fix' the algorithm for entity extraction (e.g. by pointing at the 'relevant' parts of the URL), it would already be good enough for the user -- they would fix the sites that are worst offenders for them. Then these fixes could be optionally contributed back and merged to the upstream 'rules database'.
- StavrosK 5y agoThis is a nit, but why "cannon" and not "canon"? Presumably it's not firing projectiles at things.
- karlicoss 5y agoJust the first pun I came up with :)