4 ms·
Here's a startup idea that I'd been holding close to my heart, but now that my teammate's working for a BigCorp, earning beyond her dreams, it's time for it to
by benten10 11y ago
Here's a startup idea that I'd been holding close to my heart, but now that my teammate's working for a BigCorp, earning beyond her dreams, it's time for it to go public. Steal this idea!
Link rot is only going to get bad from now. It's real, and it's awful.
It's not even random websites. NYT will post a link to a popular Youtube video, two days later the video gets pulled, and one week into the article, the links are already stale.
As Wiki gets more 'reputable', newspapers will posts links to it. Wiki being what it is, the reference to the page will go stale.
So here's an idea for a service: hash all the outgoing/inside pages on a website. If they change, either 1) give options for users to review, 2) update the link 3) delete the reference. If the original website 404's, change the link to the archive page. If a video link is outdated, provide tools to search for similar videos to link, from inside the dashboard. For Wikipedia, automatically link to a certain version of history of a page, not the general page. This is different from the idea posted below because it integrates with existing publishing systems, so it'd be more b2b, and one would start cashflow right away.
For links to social media, use a 'photo snippet' tool that looks to check if the link is valid, and goes to the image version of the link if validity is dead.
I'm certain people would pay GOOD MONEY for this service. I know I would if I were running a publication.
Give me some spare change if it succeeds. : )
- sosuke 11y agoThe ultimate version being complete capture of the Internet at large with hour to hour snapshots and history perhaps using deltas between versions. Git Internet! I've personally started saving sites I want to keep using the print-to-PDF feature. Bookmarks aren't enough when you really care to save the data.
- mmebane 11y agoI use the Firefox version of Zotero[1] to archive links. It works quite well, and is more searchable than PDFs. [1]: https://www.zotero.org/ https://www.zotero.org/
- wodenokoto 11y agoYou can save a lot of space by using a "reader-view" version of the webpage.
- stanleydrew 11y agoSome friends of mine built https://preserve.io https://preserve.io to automate that. It's basically bookmarks-plus-print-to-pdf-as-a-service.
- TazeTSchnitzel 11y ago> As Wiki gets more 'reputable', newspapers will posts links to it. Wiki being what it is, the reference to the page will go stale. Unless the page was non-notable and got deleted, there'll be a redirect.
- thyrsus 11y agoI found myself so enraged at what Wikipedia privileged editors considered non-notable that I haven't had the energy to contribute in the past few years. I'm not talking about my 2nd cousin's garage band: see the talk page on reStructuredText (which managed to survive).
- TazeTSchnitzel 11y agoThe great challenge is proving something is notable. If contested it can be an uphill struggle, because journalists don't talk about obscure things much.
- Animats 11y agoThat's the whole point of Wikipedia. Otherwise it would be like PR Newswire.
- cooper12 11y agoI don't understand why you were enraged. The editor who spoke with you made it very clear what was needed to prove notability (reliable secondary sources) and politely explained things to you while linking to official policy. I'd hardly call them privileged. It'd be another story if they said something like "I don't want this article on my encyclopedia"; he even recommended using the IBM DeveloperWorks link. The tag that was on the article is merely a signal to editors that the article needs better sourcing. To be deleted, it'd need to be nominated, where it would need to be proven that the topic is non-notable, so the system actually works in favor of keeping articles that are actually notable. It might seem like the rules are arbitrary or wielded by wikilawyers but if you read them they're not that bad. I don't agree with all of them, but "when in Rome..."
- toomuchtodo 11y agoYou've just described IPFS: https://ipfs.io/ https://ipfs.io/
- thesteamboat 11y agoThanks for the link! That's a pretty interesting project I was previously unaware of.
- toomuchtodo 11y agoYou're welcome!
- btrask 11y agoYou can't simply hash the whole page, because for most sites the hash will change constantly due to any varying content (e.g. rotating ads, new comments, random messages, A/B testing...). I wrote up a proposal for using a CSS rule to select only the "relevant" portion of a page to hash.[1] [1] https://bentrask.com/?q=hash://sha256/2c9e53858b5312564a2b8f7dc9be20df0c0193c5eadff9f2ff6d490923647a37 https://bentrask.com/?q=hash://sha256/2c9e53858b5312564a2b8f...
- bhartzer 11y agoI hate to say it, but you've just described what a good SEO does. The good SEOs (and SEO firms) will analyze and take care of your internal links, your internal link structure, and make sure the site isn't linking out to outdated/old links. If you want to do this yourself, there's several crawlers out there that you can find the data yourself. Many publications have been reluctant to link out to sites over the years because of this very thing--link rot.
- jsutton 11y agoGood SEO still doesn't fix the root of the problem--lost content. If you write an article about a video, and that video gets pulled, there goes your article.