3 ms·
This is a great first step. It's a joke that Open AI thinks they can get away with saying they use "both publicly available data (such as internet data) and dat
by numberalltheway 3y ago
This is a great first step. It's a joke that Open AI thinks they can get away with saying they use "both publicly available data (such as internet data) and data licensed from third-party providers" in their Technical Report.
There isn't anything left at that point! With that information they could actually have used anything.
If you're going to pretend to be doing science you should at least be held to some of the standards we typically associate with doing science.
I know the article talks about copyright, but not stating any sources for data is a bad precedent to allow.
- nullc 3y ago> is a bad precedent to allow. Woah. Doing bad science isn't illegal, and making it so would be quite chilling. It's common in many fields to be quite imprecise about data used in the work, and entirely uncommon in many for data to be externally reproducible. Legislation restricting research isn't the right way to improve science and is unlikely to achieve the intended effect for many reasons, including that it's easier and safer to just not touch the impacted area. In some domains this causes whole areas to go unstudied or understudied, e.g. because it runs into IRB and just isn't worth doing... but at least the rules demanding IRB approval are intended to keep people from suffering grave harm and even those are less strong than blanket regulation (they're rules tied to federal funding, not research in the abstract).
- jruohonen 3y agoI agree but things are changing: many publishers already require a disclosure statement about data. I think both the US and the EU are slowly moving to a direction of open data in scientific research. What is this "legislation restricting research"? These companies are not doing "science".
- simion314 3y agoChat GPT is a comerical product not research. I want to know if OpenAI used say GPL or other copyrighted software and then the bastards had the genius idea to put restrictions on the output in their ToS. I want stuff to be fair, if MS/OpenAI can train on GPL then I should also e allowed to train on MS proprietary code or on Disney images and video, it is not fair that big companies can screw the public but the public can't do the same to the big companies. The first step is clearly have the big companies reveal if they used copyrighted stuff.
- jeswin 3y ago> I want to know if OpenAI used say GPL or other copyrighted software and then the bastards had the genius idea to put restrictions on the output in their ToS. This is a bit of a gray area. Are you allowed to read GPL'ed code and use a similar pattern in a closed source project?
- simion314 3y agoI am a human , I am not a machine that inputs all the GPL code on the internet and then outputs similar code with very small differences. I am fine with OpenAI and MS using GPL code as long as the open source community can also train on proprietary code and art of the big companies. What happens now is that some big companies say that is OK to train on any licensed stuff and on the other hand some people are sued because they done it, I want it clarified ASAP. And personally I would not give a shit on the ToS of OpenAI and use their poutput as I like as similar as they did.
- bostik 3y agoIf you're doing any kind of science, recording provenance for your inputs should be table stakes. Bad science isn't about hiding or obscuring the origin of the data, it's about being sloppy, incompetent, or even flat out willfully misinterpreting the results. We have names for what looks like science but is done without documenting - let alone outright falsifying - where the data came from. Hoax. Advertising. Propaganda. Parallel construction. Let's not lump incompetence and malice in the same bucket, please. And if you're unsure of the data provenance, then state that fact.
- nullc 3y agoThe consumers of scholarship are able to look at and determine if its the sort of thing where access to underlying data is important-- and they're free to discount it when it doesn't provide enough. Journals and grant writers are free to set standards for the work they publish or sponsor-- and they should! But no one needs to legislate that publications such as your conclusion-- that science done without documenting its sources is properly called Hoax, Advertising, Propaganda, or Parallel construction-- itself properly document its sources. We can take it for what it is, an opinion-- one no doubt supported by some data but none of us need to see it, and we can evaluate it without calling it propaganda. If you wanted to make your point stronger, I'm sure you'd give us some supporting data (if you could figure out where those views came from...). Though people sometimes pretend otherwise, a lot of research is dressed up informed opinion, put into a formal setting with standardized argument styles so that it can be compared and assessed against other informed opinions. None the less, such work done honestly and diligently advances the human condition. The ways in which it can be best improved are field specific and can only really be judged by the people attempting to use the scholarship. In some cases the data should be published, in others its provenance documented (sometimes publication of the data would be a violation of the law, too!), in others access to source code should be paramount, in yet others the authors biases may be the primary concern, and in some fields all publications should be directly sent to the incinerator. Applying the wrong standards will just make things worse. People are smart, they tend to figure out what works for them and their field over time.
- bostik 3y ago
- rhn_mk1 3y agoDoing bad "science" is not illegal, but maybe should be, considering the replication crisis that is upon us. It diminishes the utility of the work that is being done, and makes it difficult to tell apart actual scientific discoveries from flukes and forgeries.
- duskwuff 3y ago> With that information they could actually have used anything. I mean, at least it excludes "nonpublic data we didn't obtain permission to use". So I guess that's a start? (/s)