3 ms·
This is why I think scraping rights is so valuable. As the article says: > Best case scenario is if the service is local-first in the first place. However, th
by arciini 7y ago
This is why I think scraping rights is so valuable.
As the article says:
> Best case scenario is if the service is local-first in the first place. However, this may be a long way ahead and there are certain technical difficulties associated with such designs.
> I'm suggesting a data mirror app, that merely runs in background on client side and continuously/regularly sucks in and synchronizes backend data to the latest state.
Here are a few premises:
1. It's a fact that only a small portion of users care heavily about centralizing and truly owning their data
2. As such, it's reasonable for companies to not focus on exporting data. That's not how they get value out of their data
3. That being said, companies should at least not punish that small group of users for taking matters into their own hands
Scraping is our solution to this problem, and the least companies can do is allow well-behaved (rate-limited) scraping.
- choward 7y agoI have thought about these issues a lot; especially lately. With regards to being able to scrape your data or just get your data back from 3rd party's, I think that's a losing battle. You need to be in control of your data before it gets to them. Web sites and APIs are constantly changing and sometimes just disappear. This idea of polling for changes seems very brittle and would never be up to date. What I picture is an a program that you use to store your own microblogs, blogs, contacts, comments, etc. and then you publish to whoever from that app via their API or crawling. Imagine you just created a new microblog entry. You can now either post to your Twitter, Mastodon, etc. accounts with the click of a button. You would have to poll for replies though and it would be up to you to store them if you wished (you probably want to if you are storing your replies). As an added benefit you could see the replies in one place instead of bouncing between two sites. The point is, when you create the data, it's yours first. Then if you want to, you can post it other places. Tools like this are abundant for businesses, but we don't seem to build tools for actual people anymore.
- shaki-dora 7y ago> With regards to being able to scrape your data or just get your data back from 3rd party's, I think that's a losing battle. Both Google and Facebook have tools that allow you to export your data in both (usable) JSON and HTML.
- choward 7y agoThat's fine but how often do you export? How do you merge it with other exported data? Are you exporting your entire history every time? What about when Google/Facebook breaks their API or get rid of it?
- sroussey 7y agoThat is why we are creating PrivicyPal -- to set a schedule to download and merge data from many sources and monitor over time. Will be adding beta testers soon. See https://www.privicy.com/privicypal/about https://www.privicy.com/privicypal/about
- karlicoss 7y agoYep, exactly that! As in, you can do GDPR export or Google Takeout now and then, but then you request it, enter your password/etc; in few days you will get a link to the archive. You can put a reminder to do it now and then, but then your export stales, it's just so frustrating. It's almost ok as a means of backup, but it's hard to use this data in a meaningful day.
- fuzzybeard 7y agoThe lack of an export feature is rather common among smaller services. What Google and Facebook are capable of doesn't scale down to other service providers.
- jeena 7y agoSounds like you're inventing POSSE https://indieweb.org/POSSE https://indieweb.org/POSSE That's how I do it, see the original https://jeena.net/photos/524 https://jeena.net/photos/524 the mastodon copy https://toot.jeena.net/@jeena/103214370709720207 https://toot.jeena.net/@jeena/103214370709720207 and the Twitter copy https://twitter.com/jeena/status/1199954031134887936 https://twitter.com/jeena/status/1199954031134887936 Facebook on the other hand removed that API so I stopped crossposting there: https://github.com/snarfed/bridgy/issues/817 https://github.com/snarfed/bridgy/issues/817
- choward 7y agoThe problem I see with this as with most other projects linked in this thread is that they aren't actually "solutions". They're just ideas. You need programming and/or sysadmin experience to even begin using these approaches. What I'm thinking of is more of a solution that people can just start using without any friction.
- pcr910303 7y agoOkay, copy-pasting over my other top-level comment as I believe that solutions other to scraping exists: In particular, scraping means that the fundamental data source is the company. It shouldn't be like that, why should the company own the data I have produced? > The Solid Project [0] (AFAIK which is led by Tim Berners-Lee) is made to tackle this exact problem. > It is about defining a standard/protocol to store personal data in a 'pod' and give minimal/granual access of my pod to web apps. > Apps are separated to pods, and it allows the decoupling of data & functionality. > You can switch(upgrade) from text message to instant messengers without losing any chat history for example. > It also has an advantage that it prevents lock-in, since one can move their data around trivially. > Looks like the OP would greatly benefit from this. > [0] https://solidproject.org https://solidproject.org