3 ms·
>And where the training data for your AI advice comes from? It seems like the unstated assumption in that question assumes that the world totally depends on th
by jasode 1y ago
>And where the training data for your AI advice comes from?
It seems like the unstated assumption in that question assumes that the world totally depends on the information from small independent blogs like this thread's article. I.e. all other information sources would be derivatives of the independent blogs.
There are many other sources of organic info to feed AI training. Examples:
- transcripts of Youtube videos. E.g. somebody (maybe a travel agent or a well-traveled vacationer) records a video giving advice and uploads it to Youtube. Google auto-transcribes the audio and feeds the text to the training algorithm.
- AI assistants used by normal people to "analyze/summarize information" can feed that same data to the AI cloud. E.g. a travel agent types out an email giving advice to a customer. That customer then submits that same email content (or the AI autoscans the customer's email inbox) to enable the customer to ask the AI assistant, "Is this travel advice good? Is there anything this travel agent overlooked?"
Of course, the travel advisor would want to limit his "proprietary and valuable travel knowledge" to only his direct clients in that private email but they stop the customer from exposing it to AI assistants.
The common theme is that AI engines can insert themselves in between many types of communication between people. Those are the scenarios where you can think creatively about where all the new training data will come from. If AI assistants are used as mediators in private communication, information (including "travel advice") can "leak out" into the public. Independent blogs are a good source -- but they're not the only source.
- Aldipower 1y agoThanks, that wants me to use AI even less. Both examples has nothing to do with an open and free internet. Meaning I cannot trust AI at all. All those examples of data source here, also in the other replies, using mainly highly biased sources. Wikipedia (biased by a small group of mods), YouTube filtered by Google itself, pasting customer travel advice email heavily violates GDPR, social forums also funneled. If we loose organic sites, we loose freedom. Fair enough, organic sites does mean the information there is correct, but still it is open and free, so organic sites can be treated as Gaussian distributed.
- jasode 1y ago>, pasting customer travel advice email heavily violates GDPR Not seeing how pasting text with no personal identity information would violate GDPR. E.g. someone sends an email saying "For your career prospects, I think you should learn Rust instead of COBOL." Copy&pasting that into AI or an AI scanning that sentence with no identity information isn't going to violate GDPR. There's no personal data to violate. (If the AI companies deliberately want to ignore privacy laws and want to secretly attach personal data to that "Rust/COBOL" sentence, then yes, that violates GDPR.) EDIT reply to : >or the AI autoscans the customer's email inbox That auto-scan scenario still doesn't require the AI to save the personal identifiers attached to text fragments. Many ways to do that without violating GDPR. Consider how today's global spam filters "auto scan" customers' incoming emails to automatically categorize some of them for the customers' "Junk folder" without any intervention or violation of GDPR.
- Aldipower 1y ago> or the AI autoscans the customer's email inbox