3 ms·
> "Do whatever you want, just make sure you conceal the results and impede progress and understanding." This is not a fair characterization of what's going on
by kajecounterhack 4y ago
> "Do whatever you want, just make sure you conceal the results and impede progress and understanding."
This is not a fair characterization of what's going on here. Google spent a ton of money on researchers & training infra (it's wildly expensive even just hardware-wise) to train these models. It's not different from other proprietary technologies -- they don't owe the public anything here. Providing the research findings + methodology in a paper without the implementation & data is a _tradeoff_ as a participant in the field. If someone else implements the model with their money and uses it for nefarious purposes, that's more acceptable than if they directly use Google's _already known to be flawed_ models.
> I'm curious what ethical reasons you think require that new technology only be used in secret and without oversight by trillion dollar companies. This is supposed to be AI safety?
If I make a chair and I know it's not always safe to sit on, maybe I should not sell that chair. We can talk about this proof-of-concept chair as a research subject, but if you go to build one and use it to prank someone, that's on you.
That's all that's going on here. If the model could be used to generate CSAI, maybe Google doesn't want to be part of that.
> Google is developing image and video generation models and equivalent versions will be open source by the year's end I expect. These models aren't especially dangerous.
Maybe that's the disconnect -- you don't think generative models are dangerous, but they can be, and Google would know because they have entire teams dedicated to AI fairness & safety researching this topic.
It's also not trivial to reproduce these models. Given the cost to simply train even if you had the source data, any organization releasing these models has to have a bit of money and skill. The onus will always be on the team building these models to think about what their ethics are and how they want to proceed knowing there may be negative externalities.
> Yes, people will use them to be racist or mean, same as they use their phones or computers or books or whatever to be those things.
Tools empowering large-scale inauthenticity & disinformation are not comparable to individuals making comments.
- ALittleLight 4y agoGoogle uses research, published models, and data that was freely shared with them and iterates on it, making use of their vast budgets and hardware, to develop new models. Then, Google uses those models internally and doesn't share the models. This is a violation of academic norms under the pretense of "safety". As I characterized previously Google is able to do whatever they want, conceal their results, and impede progress and understanding because they aren't sharing their results. You say this isn't a "fair characterization" but it is exactly what is happening - which part is wrong? You say that Google doesn't "owe the public anything" and that may, or may not, be true from a legal standpoint, but obviously, from a norms, ethical, and moral standpoint Google does have a massive obligation to the public that they are breeching. Google uses the public's data to train, public research, and publicly shared models to iterate on. Then, after building on the shoulders of giants, Google refuses to share what they have built in contravention of the norms that they benefit from. Regarding your chair metaphor - the "danger" of these models, if there is such, is not that they would hurt the user, like a faulty chair, but that they could be used to hurt others - e.g. a bot army to manipulate public opinion or create fake news. Google isn't building a chair that might break and hurt the user then, but a gun that might hurt others. It's true that guns shouldn't be widely available - not even a die hard libertarian would want a child to have access to a gun, but the entity that sets rules regarding availability is a representative government for the people for whom those rules are being set - not a private company. In other words, if these tools can cause harm they should be regulated by the government, not Google. If the tools are dangerous, that is not an argument that Google should keep them secret.
- kajecounterhack 4y ago> Google uses research, published models, and data that was freely shared with them and iterates on it, making use of their vast budgets and hardware, to develop new models. Then, Google uses those models internally and doesn't share the models. This is a violation of academic norms under the pretense of "safety". Google's not doing this (LLM, generative image model) research on academic datasets freely shared with them. They're doing this research on data they gathered at their expense. This is not a violation of academic norms. Again, Google shares a lot of datasets and models, just not LLMs and generative sets trained on problematic source datasets. > As I characterized previously Google is able to do whatever they want, conceal their results, and impede progress and understanding because they aren't sharing their results. You say this isn't a "fair characterization" but it is exactly what is happening - which part is wrong? Anyone can do research and not share back to the community. Google _does_ share back to the community in the form of papers (and again, very frequently with models and datasets). If you have the money and expertise to implement the papers, more power to you. Every technology company has some secret sauces they don't share with everyone. That Google may have some of those is not a moral failing. > from a norms, ethical, and moral standpoint Google does have a massive obligation to the public that they are breeching. Google uses the public's data to train, public research, and publicly shared models to iterate on From the other end: Google gets user data and has a responsibility to not proliferate that data, no? I wouldn't want them to share a dataset that has my personal data, even if anonymized because there are ways to deanonymize. There are levels to everything, and choosing "I'll release the paper but not the model + data" for some potentially sensitive models seems sane. > Then, after building on the shoulders of giants, Google refuses to share what they have built in contravention of the norms that they benefit from. People are building on the shoulders of Google's research all the time, and plenty of companies are doing similar things to Google and being way less open about their work. I mean, every company that trains a big model on data collected from the public -- are they all required to share their models with everyone? Is Cruise sharing their pedestrian detection model? I don't think what you're suggesting could possibly be the standard. > Regarding your chair metaphor - the "danger" of these models, if there is such, is not that they would hurt the user, like a faulty chair, but that they could be used to hurt others - e.g. a bot army to manipulate public opinion or create fake news. Sure, I was trying not to be hyperbolic and compare LLMs to guns since they have plenty of awesome use cases (whereas guns really don't). A faulty chair that you set out for anyone to use can hurt people other than the chair's creator / people who are aware of the specific risks. But yeah, seems like you now agree these models have the potential to cause great harm. > In other words, if these tools can cause harm they should be regulated by the government, not Google I agree that gov't regulation can be helpful for setting a minimum standard. But I strongly disagree that lack of laws means we should abdicate our own moral responsibilities. If I sell / provide something, I need to be able to sleep at night knowing I didn't make the world worse. Googlers typically try to do this.