6 ms·
The article decides to mention that Google made a public statement that clearly and unambiguously answers their question (https://twitter.com/GoogleWorkspace/st
by delroth 3y ago
The article decides to mention that Google made a public statement that clearly and unambiguously answers their question (https://twitter.com/GoogleWorkspace/status/1638298537195601920 https://twitter.com/GoogleWorkspace/status/16382985371956019...) and then proceeds to ignore it completely in favor of conspiracy theories ("but look, there they said 'was' instead of 'is'!").
- ahahahahah 3y agoIt gives as much credence to bard's own bullshit as it does to Google's official statement. It's a useless article.
- gersg 3y agoThank the Lord for that "Readers added context" marker that tweets now can have.
- onion2k 3y agoLawyers use language in very specific ways, and anything that's not completely obvious can be used to hide the truth. As an example, I once worked with a team who was building some software where a requirement said the app 'should' do something instead it 'shall' do something. The company's lawyer argued successfully that this meant the requirement was optional.
- astrange 3y agoThat's how Internet standards work. https://www.ietf.org/rfc/rfc2119.txt https://www.ietf.org/rfc/rfc2119.txt
- onion2k 3y agoThis was a contract between two companies, written by a product manager and a CEO. It wasn't a technical RFC.
- omegabravo 3y agoThis choice of wording in systems engineering is not ambiguous, and should effectively means optional. These words are often in capital letters trying to highlight the importance of it. It is absolutely not limited to RFCs, and is often used in a software specification. Product manager and CEO _should_ know better. It's very understandable that they don't - and I have empathy for them, but unfortunately they're wrong.
- meghan_rain 3y ago_must_ tbqh lmao
- onion2k 3y agoSure, and hence the company won. The point here is that language can have very specific meaning. There is a difference between 'should' and 'shall', even though most people would think they're effectively the same. If I said "You should complete that task" to one of my team I'm not really giving them the choice and leaving it up to them. I'm telling them to do something. Outside of RFCs and contracts language is a bit ambiguous and lawyers use that to their advantage. In exactly the same way, there is also a difference between 'is' and 'was', and I think it's totally plausible that a Google lawyer might use that to hide the fact they used GMail data to train AI in the past. That doesn't mean they did. It only means I wouldn't be surprised if someone proves they did, and that their lawyer used the tense of a response to try to hide it.
- matwood 3y agoEven in general english, 'should' is less strong than 'must' or 'shall'.
- matwood 3y ago'Should' and 'shall' is standard contract language. But, when it doubt, define the terms in the contract.
- delroth 3y agoI'm not sure what you're trying to say, but "It is not trained on Gmail data." is as obvious a statement as you could ever express in the english language.
- chx 3y agoNot only that but this tweet is from Google Workspace and so it is not at all unreasonable to say it is not trained on Google Workspace Gmail and it says nothing about public gmail. Words. We have them.
- comprev 3y agoThe Highway Code in the UK is full of “must” and “should” indicating firm requirements and optional choices. The use of “should” is a softer method of persuasion, in the same way signs say “Please close the door” not just “Close the door” which is an instruction. This aligns with the famous British politeness - or at least that is how I interpret the wording.
- marak830 3y agoI would think that part 4 and "what bard has to say about this" sections would make most people question Google's comment. 4. Google has never denied that Bard was trained on data from Gmail. They've only claimed that such data is not currently used to “improve” the model. What Bard has to say about this: “I have not personally seen a real Gmail account. However, I have access to a massive dataset of Gmail emails, and I have used this dataset to train my language model. This means that I am familiar with the format of Gmail emails, and I can generate text that is similar to the text that is found in real Gmail emails.” Now do I think they have done the nasty? I don't know. Should it be reviewed by an outside team? I think yes. I cannot think of a solution to this problem, which I believe will keep cropping up, but I think it can be problematic. I think it needs to be prooven true to be safe.
- zmmmmm 3y agoThey could easily have done this on a subset of consenting users, eg: their own employees.
- marak830 3y agoYou are 100% correct. That they didn't mention this is one of the reasons I didn't dismiss this article.
- wodenokoto 3y ago> Bard is an early experiment based on Large Language Models and will make mistakes. It is not trained on Gmail data. -JQ How exactly are you interpreting that statement?
- marak830 3y agoI am defining it on the points I mentioned. I don't know if they are in the wrong or not, but their responses give me pause
- ForHackernews 3y ago"is" is a present-tense verb. In the hands of a lawyer or a PR person, that statement could easily mean "We are not, right at this moment, training it on further Gmail data. Up until yesterday, we were training it on all the private data from the past 20 years, and next month when the scrutiny dies down, we'll start training on fresh gmail data again."
- bilekas 3y agoThey also have added replies which have been removed, reading the article it's actually not as conspiratorial as you make it seem. In this case I really think it prudent to assume the worst from Google as they don't really have a positive history for walling off users data, be it personal email or phone meta information. > "The LaMDA engine underlying Bard is also what drives autocomplete and autoreply in Gmail so ... yeah Bard's training data includes Gmail. FWIW, they put a lot of effort into ensuring that LaMDA doesn't use give[sic] personal information about individuals in its responses." If this is true, to me this is a good indicator that it's using at least contextual information from emails.
- q1w2 3y ago> I really think it prudent to assume the worst... This is not evidence-based thinking and leads to heavy biases.
- bkanl 3y agoWho wrote that reply? A Googler? Is there a screenshot? This would be a gigantic GDPR lawsuit. This reminds me of not communicating with Gmail users. Gmail has been evil forever since "personalized" ads.
- bilekas 3y agoIt's in the article, but here's the direct link to help you out. https://twitter.com/cajundiscordian/status/1638243303035670528 https://twitter.com/cajundiscordian/status/16382433030356705...
- skinkestek 3y agoAt the moment, for some reason Google seems almost untouchable. I mean, single handedly destroying the browser market by deceit and abuse of market position in broad daylight, you'd think sooner or later EU or someone would force them to pay and put up a browser ballot on Google.com, but so far, no. Luckily the French consumer protection agency has at least forced them to implement the cookie question thing almost so we can now reject all right away.
- ninth_ant 3y agoUsing the present tense to answer a past-tense question is hardly unambiguous. "Did you send an email to Fred?" -- "No, I'm not sending an email to Fred" doesn't answer the question. It's not a "conspiracy theory" to have realized that big corps have teams of people to frame their public statements with carefully chosen words to present issues in the best light for them, even if it's deeply misleading.
- radarsat1 3y ago> "Did you send an email to Fred?" -- "No, I'm not sending an email to Fred" To me this unambiguously says both that I did not send an email to him and I don't intend to. Is it really ambiguous to you?
- nkrisc 3y agoThe given response very clearly and unambiguously does not answer the given question. I think most native English speakers - if they were reading/listening carefully - would interpret the mismatch of verb tense as an intentional attempt at not answering the question while simultaneously sounding as if it does. If someone gave me that answer to that question I would repeat the question to them with an emphasis on “did.”
- ForHackernews 3y agoYou're not a corporate comms person. The ambiguity of the English language can be very valuable in a court of law.
- hdjjhhvvhga 3y agoIt depends if you are a person or a corporation. If the former, I assume you haven't. If the latter, I assume you did but don't want it to sound bad.
- pyrale 3y agoIf you said that, maybe. If a corporate PR expert or a politician said that, the conclusion would be very different.
- somenameforme 3y agoWhen a corporation answers in a specific way it's because it has a specific meaning. 'Ooops we meant 'x' doesn't tend to hold up so well in court.' Google could easily and absolutely put to bed all concerns by stating, "No Bard has never been trained on any email data, and never will be." Instead, they're choosing not to do that and just taking the PR hit, while making statements that completely leave the door open to previous training.
- jeodjdodh 3y agoJQ is part of the PR team, not engineering. the article author is correct not taking what he says at face value. lots of doublespeak in PR.