6 ms·
This is great news! It feels like open-source is closing the gap with "Open"AI which is really exciting, and the acceleration towards parity is faster than mor
by burcs 3y ago
This is great news!
It feels like open-source is closing the gap with "Open"AI which is really exciting, and the acceleration towards parity is faster than more advancements made on the closed source models. Maybe it's wishful thinking though?
- udev4096 3y agoIs it tho? It's not really open source if they don't give us the information regarding training datasets
- jerpint 3y agoIt definitely is open source even if they don’t disclose all details behind the training
- SOLAR_FIELDS 3y agoThe very definition of what constitutes open source is being called into question in these kinds of discussions about AI. Without the training details and the weights being made fully open it’s hard to really call something truly open, even if it happens to meet some arbitrary definition of “open source”. A good definition of “truly open” is whether the exact same results can be reproduced by someone with no extra information from only what has been made available. If that is not possible, because the reproduction methodology is closed (a common reason, like in this case) then what has been made available is not truly open. We can sit here and technically argue whether or not the subject matter violated some arbitrary “open source” definition but it still doesn’t change the fact that it’s not truly open in spirit
- okaram 3y agoNotice you are creating your own arbitrary definition of 'truly open', which IMHO corresponds more with 'reproducible'. We already have a definition of open source. I don't see any reason to change it.
- losteric 3y agoThe inference runtime software is open, the weights are an opaque binary. Publishing the training data, hyperparameters, process, etc - that would make the whole thing "open source".
- magicalhippo 3y agoThe quake engine is still open source even though it doesn't come with the quake game assets, no? It seems unreasonable to require the training data just to be called open source, given it has similar copyright challenges as game assets. Of course, this wouldn't make the model reproducible. But that's different from open source.
- darkwater 3y agoGood example. And in fact you are calling the "engine" opensource, not the whole Quake game. The 'assets" in most "opensource" AI models are not available.
- EGreg 3y agoImagine if the Telegram client was open source but not the backend. Imagine if Facebook open-sourced their front-end libraries like React but not the back-end. Imagine if Twitter or Google didn’t publish its Algorithm for how they rank things to display to different people. You don’t need to imagine. That’s exactly what’s happening! Would you call them open source because their front end is open source? Could you host your own back end on your choice of computers? No. That’s why I even started https://qbix.com/platform https://qbix.com/platform
- darkwater 3y agoI completely agree with you (and the example you mention are singled out in the "antifeatures" list in F-Droid, to name an example)
- torginus 3y ago
- rolisz 3y agoThen a lot of stuff is not open source. Have you tried reproducing random GitHub repos, especially in machine learning?
- richardw 3y agoSo if someone includes images in their project they need to tell you every brush stroke that led to the final image? All sorts of intangibles end up in open source projects. This isn’t a science experiment that needs replication. They’re not trying to prove how they came up with the image/code/model.
- xnorswap 3y agoThose "Brush Strokes" are effectively the source code. To be considered open source, yes source code needs to be provided along side the binaries (the "image").
- EGreg 3y agoIt’s more like someone giving you an open source front end client, but not giving you a way to host your own backend. Look into Affero GPL. Images are inert static assets. Here we are talking about the back end engine. The fact that neural networks and model weights are non-von-neumann architecture doesn’t negate the fact that they are executable code and not just static assets!
- abriosi 3y agoImagine someone giving you a executable binary without the source code and calling it "open source"
- jyrkesh 3y agoI'm actually mostly in your camp here. But it's complicated with AI. What if someone gave you a binary and the source code, but not a compiler? Maybe not even a language spec? Or what if they gave you a binary and the source code and a fully documented language spec, and both of 'em all the way down to the compiler? BUT it only runs on special proprietary silicon? Or maybe even the silicon is fully documented, but producing that silicon is effectively out of reach to all but F100 companies? It's turtles all the way down...
- krageon 3y agoThere is the binary (the model) and the source (the thing that allows you to recreate the model, the dataset and methodology). Compilers and how art is made quite simply doesn't factor in here, because nobody is talking about the compiler layer. Art isn't even close to what is present. Trying to make this more complicated than it is is playing into companies' hands by troubling the waters around what constitutes open source.
- r3trohack3r 3y agoTo be fair, OpenSource troubled the waters around what constitutes free software. Free(dom Respecting) Software wasn’t just about the source code. https://www.gnu.org/philosophy/open-source-misses-the-point.en.html https://www.gnu.org/philosophy/open-source-misses-the-point....
- DougBTX 3y agoYou can pass in any command line arguments you like, so it must be open source
- otikik 3y agoWell the other day on this very website there were some very opinionated voices stating that Open Source is “exclusively what OSI defines”. I am not on that camp, more like in yours. To me there’s open source and OSI-approved open source. But you will encounter people very set on that other opinion, which I found interesting. Make no mistake, I am super grateful to OSI for their efforts and most of my code out there uses one of their licenses. I just think they are limited by the circumstances. Some things I consider open are not conforming to their licenses and, like here, some things that conform might not be really open.
- m3at 3y agoTo take an other example, would you call a game that has its code and all assets (ex. character sprites) freely available open source? Or would the process that was used to create the assets in the first place also be required to be considered open? The parallel can be made with model weights being static assets delivered in their completed state. (I favor the full process being released especially for scientific reproducibility, but this is an other point)
- pjc50 3y agoThe old Stallman definition used the phrase "preferred form for modification" rather than the more specific "source code". What do you need to effectively modify an AI model?
- kordlessagain 3y agoUsually the datasets, not the source code.
- selcuka 3y agoHow do you define "source", then? By this logic any freely downloadable executable software (a.k.a. freeware) is also open source, even though they don't disclose all details on how to build it.
- mogwire 3y agoSource would be the way the data is produced so that you can replicate it yourself and make changes. If I hand you a beer for free that’s freeware. If I hand you the recipe and instructions to brew the beer that is open source. We muddy the waters too much lately and call “free” to use things “open source”.
- TeMPOraL 3y ago> If I hand you a beer for free that’s freeware. If I hand you the recipe and instructions to brew the beer that is open source. Yeah, but what those "open source" models are is like you handing me a bottle of beer, plus the instructions to make the glass bottle. You're open-sourcing something, just not the part that matters. It's not "open source beer", it's "beer in an open-source bottle". In the same fashion, those models aren't open source - they're closed models inside a tiny open-source inference script.
- imranhou 3y agoPerhaps one more thing that is missing in context is that I'm also getting the right to alter that beer by adding anything I like to it and redistributing it, without knowing its true recipe.
- szundi 3y agoInteresting as the literal source of the result is not open
- EGreg 3y agoPeople need to realize something… The model weights in eg TensorFlow are the source code. It is not a von-Neumann architecture but a gigabyte of model weights is the executable part, no less than a gigabyte of imperative code. Now, the training of the model is akin to the process of writing the code. In classical imperative languages that code may be such spaghetti code that each part would be intertwined with 40 others, so you can’t just modify something easily. So the fact that you can’t modify the code is Freedom 2 or whatever. But at least you have Freedom 0 of hosting the model where You want and not getting charged for it an exorbitant amount or getting cut off, or having the model change out from under you via RLHF for political correctnesss or whatever. OpenAI has not even met Freedom Zero of FSR or OSI’s definition. But others can.
- simonw 3y agoThat doesn't work for me. The model weights aren't source code. They are the binary result of compiling that source code. The source code is the combination of the training data and configuration of model architecture that runs against it. The model architecture could be considered the compiler. If you give me gcc and your C code I can compile the binary myself. If you give me your training data and code that implements your model architecture, I can run those to compile the model weights myself.
- EGreg 3y agoNo, you would need to spend “eye watering amounts of compute” to do it, similar to hiring a lot of developers to produce the code. The compiling of the code to an executable format is a tiny part of that cost.
- simonw 3y agoI still think of millions of dollars of GPU spend crunching away for a month as a compiler. A very slow, very expensive compiler - but it's still taking the source code (the training material and model architecture) and compiling that into a binary executable (the model). Maybe it helps to think about this at a much smaller scale. There are plenty of interesting machine learning models which can be trained on a laptop in a few seconds (or a few minutes). That process feels very much like a compiler - takes less time to compile than a lot of large C++ projects. Running on a GPU cluster for a month is the exact same process, just scaled up. Huge projects like Microsoft Windows take hours to compile and that process often runs on expensive clusters, but it's still considered compilation.
- Gasp0de 3y agoThey compare it to OpenAI's ada model though, which is light-years away from ChatGPT.
- infecto 3y agoDoes that not conflate two different things though? Embedding model != LLM Model ?
- simonw 3y agoDon't confuse the current Ada embedding model the old Ada GPT3 model. It turns out OpenAI have used the name "Ada" for several very different things, purely because they went through a phase of giving everything Ada/Babbage/Curie/DaVinci names because they liked the A/B/C/D thing to indicate which of their models were largest.
- infecto 3y agoWishful thinking? Embeddings to me were never the interesting or bleeding edge thing at OpenAI. Maybe the various ada models at one point reigned supreme but there have been open-source models at the top of the leaderboard for a while and from a cost/performance perspective, often even the Bert models did a really fine job.