4 ms·
Will these systems be prone to false data? I.e. if an organization fighting for free speech would create software generating false personas who create false me
by ChaosDegenerate 7y ago
Will these systems be prone to false data?
I.e. if an organization fighting for free speech would create software generating false personas who create false messages, and other events, would they know any better?
For example, it wouldn't be too hard to generate hundreds if not thousands of fake Facebook, Instragram, Airbnb, Coinbase, Uber, Gmail, Amazon, etc. accounts doing "stupid things". Like ordering stuff and canecelling orders right away. Like generating fake emails en masse with "trigger words" in millions a day. Like creating fake posts on Facebook warning of fake incidents like "15% of Uber drivers have serious mental issues".
The whole thing could be scripted and run on bots worldwide in millions. After some time, serious chunk of internet traffic would be fake. How would these companies know any better?
- tuukkah 7y agoData from Google Pay and Apple Pay must be extremely valuable as it links your user account to actual physical purchases - expensive to fake those en masse.
- Rebelgecko 7y agoWhen I went to check out the new Apple TV+, the fine print mentioned that my usage of the service was contingent on my "device trust score", which depends on various metrics like how often I use my device, how many emails I send, etc. Considering I was trying to sign up in a web browser, it was a surprise that Apple may be tracking these metrics for me. I wonder if people that don't own iDevices will get lower resolution streams due to a lower social credit—excuse me, "device trust"— score
- Canada 7y agoThat wouldn't matter. These companies would still have real data for actual people. And it's not easy to create tons of fake accounts for services like Coinbase because they have strict identity verification. Any account that doesn't complete it obviously isn't really a real account.
- Majromax 7y ago> Any account that doesn't complete it obviously isn't really a real account. Without very careful system design, that itself could be used to corrupt data. Suppose a thousand obviously fake accounts are made and use the phrase "giraffe-eater" (to coin something random-looking) in messages. If the system is designed as a fancy nonlinear regressor, then it doesn't have to obey causality -- observing P("giraffe-eater"|fake) = 100% will increase the modeled prediction of P(fake|"giraffe-eater"). Worse yet, if the system is in fact being built in a non-interpretable way (throwing everything into a random forest or deep neural network), then it will be impossible to prove anything about how the system makes its judgement. Ultimately, these models are at risk from anti-discrimination statutes if (for example) they can key on use of dialect or other protected-ground-correlates.
- mr__y 7y agoIf all this data would be more or less random (which is kind of implied by "run on bots worldwide in millions") this would merely add some random noise to large datasets. For this to create any real problems the noise generated by those bots would have to be somehow biased, of course in an incorrect way. This implies then, that for such idea to be even remotely effective one would need at least a peak in aggregated data to know real trends and program bots to blur the image generating fake trends.
- ChaosDegenerate 7y agoThank you for your reply. Granted that Google, Microsoft, Facebook, and all the others collect data on us. And this is used by the Government security agencies (i.e. NSA/CIA as described by Snowden) to spy on us -- as we all know now the agencies have backdoors to these Corporations enabled with stuff like 'secret rooms' (AT&T) common. In which case this is used to detect terrorism threats. Why state-operators with enormous funding, like let's say Iran or North Korea or Russia or China -- won't generate systems like this to make NSA/CIA job harder (if anything). If I were Iran state security institution planning, targeting and executing terrorist attacks against the US and their allies -- I would be very interested to do things like millions of email accounts sending millions of emails daily with trigger words like 'bomb', 'attack', 'death to america', etc, etc. So my real communication would be a drop in the ocean of similar-looking noise. What is stopping these actors from doing just that? I have wondered about this for long time and I'm just curious. If the whole 'terrorism threat' thing is real, we would see counter-actions like this done by terrorism sponsoring organizations or states. No?
- jlokier 7y agoCredit reference agencies are already prone to false data, and that's without anyone trying to game them. When I checked this year, 2 out of the 3 major CRAs in my country (Equifax, Experian) had: - An incorrect flag saying I was not registered to vote, and a note saying this significantly affected my credit rating. - A bogus address that didn't correspond to any physical location. Nor did it correspond with any address used by any business, that I know of. (If it does, they won't be able to mail me!) - Three credit application searches (hard searches) in 3 days, for applications I didn't make. (I complained to the relevant company, who agreed they were added due to software errors on their side, and "resolved" the complaint by agreeing to remove them; but in the end they didn't remove them, so my history has unremovable entries for applications I've never made.) - An account with the largest telecoms provider (BT) in the country that I didn't have (I'd left them 2 years earlier, account fully closed and settled). - Fictitious monthly entries on the above account showing new amounts being added each month, of seemingly random amounts (no obvious pattern), and flagged as severe, overdue, late payer etc. Not a good look on a credit record, and entirely fictitious. (Fixing this proved arduous, and I ended up having to use three companies in a daisy-chain of each one passing along a formal complaint to the next. I later learned from BT customer support that BT does this to other former customers without their knowledge as well, so for ethical reasons if I can muster the energy I'll be complaining about this to the government regulator) I could say so much more about the complaints process, terrible customer service in every conceivable way from Equifax specifically, and more, but it would be rather off-topic. Remarkably, just complaining about the above caused the errors to be acknowledged and correct data found magically by the companies involved, without me needing to provide replacement data. It's as if the companies involved had all the data they needed already, they just aren't using it until a customer finds out and complains.
- therealx 7y agoSo common. Ive had all of these happen, too. I didn't know not being registered to vote could hurt you, was that in the US? I assume not due to BT being mentioned. AT&T is currently doing the same thing BT did to you. I closed my account and shipped back hardware a year ago. Since then, they have been reporting either a bill with a random number from $100-3,000 or 0, and being either on time or late or none. It's maddening.