5 ms·
Human Web – Data Collection Without Privacy Side-Effects
- lamkhanhsong 7y agoThis is the 3rd post from the advent. day 1: Why the world needs more search engines https://www.0x65.dev/blog/2019-12-01/the-world-needs-cliqz-the-world-needs-more-search-engines.html https://www.0x65.dev/blog/2019-12-01/the-world-needs-cliqz-t... day 2: Is data collection evil https://www.0x65.dev/blog/2019-12-02/is-data-collection-evil.html https://www.0x65.dev/blog/2019-12-02/is-data-collection-evil...
- deleted 7y ago[deleted]
- chrmod 7y agoAsking HN: point another company describing why and how it does data collection at this extent.
- dessant 7y agoThere is an even better solution: do not collect any data, especially when your customers are paying for your product. If your business model involves subsidizing product prices with the prospect of tracking users and collecting personal data, consider releasing a version of your product that cost more, but does not engage in any data harvesting.
- chrmod 7y agoCheck the post from previous day titled “Is Data Collection Evil?” https://www.0x65.dev/blog/2019-12-02/is-data-collection-evil.html https://www.0x65.dev/blog/2019-12-02/is-data-collection-evil... By now it should be obvious that no privacy oriented company (or in fact any other company) should treat user data as product.
- gfodor 7y agoFor a given product requirement, you have a set of potential reifications which include their data impact. There's not a zero or one dichotomy as you suggest. You can cut requirements, but all the ones that remain will have varying reliance upon data collection. The goal should be to minimize that impact to its least possible surface area while still delivering the same user value. It's especially toxic to pre-emptively collect data for unknown, or conjectured later requirements. YAGNI applies.
- dessant 7y agoOf course, some products require data collection to function, and I agree that in such cases data collection should be kept at a minimum. I was referring to products such as Windows. The consumer versions of Windows do not have UI controls for stopping telemetry entirely, nor is there a paid extension offered by Microsoft that disables it. There is clearly a demand for a Windows version that offers complete privacy, yet Microsoft is not interested in meeting that demand.
- netankit 7y ago(Disclaimer: I work at Cliqz) Exactly! It's a high time for everyone to have an open discussion and have some strict implementation around what kind of data collection is considered fair? Search being a complex application as discussed in our previous articles on our blog (https://0x65.dev/ https://0x65.dev/) requires data and by that I mean real "Quality" data. When dealing with large datasets, circumventing through all the noise is probably a harder problem than search itself. This is where Human Web shines for us. Consider this realistic scenario, for a user query, we have to navigate through an index of billions of webpages and come up with at-least 10 relevant results (which fill up page one). On top of that, make sure that the result on first position (top result) is most Relevant or Vital (a very broad and subjective metric) - All of this under fraction of a second. Add to this, the complexity involved dealing with an ever evolving and exploding content on the web, where we have different languages and region locales to deal with. Through our blog articles, we want to re-emphasize that building an "Independent" search engine is very very complex and a challenging task, specially when we incorporate constraints encountered at our scale and the fact that we want to really abide by the principles of "Privacy by Design". As with most we too could assume that "X" data point is readily available, let's collect it! But In reality, we always circle back. Literally, ask ourselves ... Does collecting X datapoint violate privacy for our user? If the answer to this is ever a "Yes", the X datapoint is dropped immediately and we DO NOT COLLECT IT! For us, features which rely on that datapoint X are just not deemed worthy!!! If the industry starts to replicate this practice, it would be a great win for Cliqz.
- dessant 7y agoI'm wary of companies that insist so much on collecting user data, especially when it's done at such a granular level. I keep wondering what did Mozilla have to gain from associating themselves with Cliqz. https://en.wikipedia.org/wiki/Cliqz#Integration_with_Firefox https://en.wikipedia.org/wiki/Cliqz#Integration_with_Firefox
- tschakkaMarc 7y agoHi – thanks for the feedback. I’m Marc (disclaimer, I work at Cliqz). The goal of the collaboration was to jointly build a better, more private search engine. Don’t forget that every major browser today sends every keystroke of the Omnibar to either Google or Bing. No privacy mechanism in place (and don’t even get me started on all the tracker madness). We wanted to jointly replace that. Cliqz is building a new search from ground up with privacy by design. In the end the collaboration didn’t work out (for many reasons, lack of privacy was not one of them though). Firefox changed back to Google as search provider. Coming back to your point: There are many services that can’t be built without data. Search is one of them, without data you will have a very bad search engine, impossible to compete. We explain this in detail here: https://0x65.dev/blog/2019-12-02/is-data-collection-evil.html https://0x65.dev/blog/2019-12-02/is-data-collection-evil.htm... . We took maximum scrutiny, and this article about Human Web is exactly there to explain how we collect data that is needed, without the side effect of collecting personal data. We are so transparent about this, because we want the scrutiny. Our business does not depend on collecting personal data or actually any data. But our product needs a lot of data. Denying anyone to collect data – even if they are as open, transparent, and without any interest in personal information – just means you support those that are the incumbents and have no interest in privacy.
- deleted 7y ago[deleted]
- dessant 7y agoIs there a detailed writeup on the Human Web proxy network, specifically on the data transmission? It would be interesting to see how does it prevent Cliqz and proxy server operators from learning the user's IP address. Was Tor evaluated as an alternative for data transmission?