3 ms·
Can anyone point to a more nuanced perspective on this? Ideally not only with regards to OpenAI; it is still quite common that research published in some of the
by timkam 6y ago
Can anyone point to a more nuanced perspective on this?
Ideally not only with regards to OpenAI; it is still quite common that research published in some of the "top" venues does not disclose the underlying source code.
Anecdote: I once pushed researchers to publish their code when I was peer-reviewing and it turned out that the code was super slow in comparison to the benchmark algorithms they compared with (comparison only included accuracy etc, not speed), something that I am sure the authors were aware of, but chose to not report in the paper.
- unityByFreedom 6y ago> it is still quite common that research published in some of the "top" venues does not disclose the underlying source code. The whole point of "OpenAI" was to share their innovations openly. They were a non profit. Now they're a "capped profit" and not sharing what's needed to reproduce their results (full code and training data). They acquired talent on a bunch of good will and shifted into a for profit.
- IAmEveryone 6y agoI don't necessarily agree with the following with regards to OpenAI, but it is what I have been told by a few professors and other academics in bioinformatics: Money is terribly tight in academia, and even at hot think tanks like OpenAI. The difference between a working but not very pretty prototype and a polished product you are willing to be associated with may seem like just a bunch of trivial boilerplate such as documentation or tests. It is indeed boring work (at least for scientists), but a lot of it. And if the intention is to publish and support something on an ongoing basis, and for widespread use, you probably need to invest 5x to 10x as much time and money. So in the end most code in academia is hidden because it's hideous, not to hide how they hardcoded all the good jokes GTP-3 comes up with.
- woah 6y agoSuch an odd attitude. There’s no way it takes 5-10x engineering effort to write some tests for components of a one-shot python script, as most scientific code is. It’s maybe the easiest type of software to test. Would this type of attitude fly for anything else? Can you imagine if chemists worked like this? “Yea most chemists don’t like to release their procedures, since taking the trouble of accurately measuring out reagents could be 5-10x the effort and grant money is tight. Plus sometimes they don’t want to go all the way to the cabinet to grab a beaker so they just do the reaction in a coffee mug. The truth is that sometimes the procedure is just too ugly to release. And if they come out of the lab with a new chemical, what does it matter?”
- mennis16 6y agoI mean a lot of biology papers are loose on protocol details to be honest, dunno about Chemistry but for Bio it can often be difficult to exactly replicate a method because of lack of details alone. That said I think this should change given how easy it is these days to share supplemental info online. As far as ugly code I wouldn't be surprised if a lot of academics won't take the time to fix it up, but that doesn't mean it shouldn't be released IMO. Maybe an open source community would pop up that could help clean this code for them.
- albntomat0 6y agoHere's my alternative view: If OpenAI had open sourced GPT-3, there would be an equivalently angry thread about how they were endangering democracy/social order/etc, and not being responsible with the powerful tool they had created (see other threads on HN regarding those working in facial recognition). Both that group, and the one posting here have valid points, and there would be strong, valid critique of their decision no matter which way they chose.
- sudosysgen 6y agoBut we do realize that sooner or later a model even more powerful will be released to the public, right? And as far as protecting democracy, I assure you that the geopolitical enemies of your nation either have similar models, or will have them very soon. There is very significant investment on ML models for text manipulation going on behind closed doors, funded by States.
- albntomat0 6y ago> But we do realize that sooner or later a model even more powerful will be released to the public, right? And at that time, we can equally criticize whomever releases it. > And as far as protecting democracy, I assure you that the geopolitical enemies of your nation either have similar models, or will have them very soon. There is very significant investment on ML models for text manipulation going on behind closed doors, funded by States. True, but it's a good idea to keep a high bar for developing and using it. There's a large difference between the resources of a nation state and various criminal enterprises. Per [0], GPT-3 took $12 million to train. That does not include the people with the relevant skills need to train it, and access to the compute hardware. [0]: https://venturebeat.com/2020/06/01/ai-machine-learning-openai-gpt-3-size-isnt-everything/ https://venturebeat.com/2020/06/01/ai-machine-learning-opena...
- sudosysgen 6y agoI mean, I don't really see the benefit of limiting such a model to such famously reliable and disinterested actors as nation-states and multinational corporations. The kind of criminals that are going to threaten "democracy/social order/etc..." aren't going to be stopped by a price tag of 12 million dollars, or 120 million dollars for that matter. Just a single Mexican cartel would have the resources to pay for that, to say nothing of criminal groups that work in lockstep with states. The proceeds of crime are 900 billion dollars per year. Training GPT-3 is well within the means of dozens of criminal enterprises. But sure, we can praise OpenAI for protecting us from small enterprises and individuals, known to democracy, and instead reserving it for the use of megacorps and states, that are known not to threaten democracy or social order.