3 ms·
You have a weird definition of open source. OS software developers don't release the books they have read or the tools they've used to write code. This is full
by ko27 2y ago
You have a weird definition of open source. OS software developers don't release the books they have read or the tools they've used to write code.
This is fully 100% OSI compliant source code with an approved license (Apache 2.0). You are not entitled to anything more than this.
- Zambyte 2y agoThey don't have a weird definition of open source. I recently outlined a LLM chat that I think clearly outlines this: https://news.ycombinator.com/item?id=40035688 https://news.ycombinator.com/item?id=40035688
- ko27 2y agoA bunch of code was autocompleted or generated by IDEs, are open source developers supposed to release the source code of that IDE to be OSI compliant?
- Zambyte 2y agoIs the IDE a primary input for building the program? Is the IDE a build dependency? Probably not. Certainly not based on the situation you described. The LLM equivalent here would be programmatically generating synthetic input or cleaning input for training. You don't need the tools used to generate or clean the data in order to train the model, and thus they can be propriety in the context of an open source model, so long as the source for the model is open (the training data).
- ko27 2y ago> Is the IDE a primary input for building the program? Is the IDE a build dependency? No, the same way training is not a build dependency for the weights source code. You can literally compile and run them without any training data.
- Zambyte 2y agoTraining data is a build dependency for the weights. You cannot realistically get the same weights without the same training data.
- ko27 2y agoDeveloper's mindset, knowledge and tooling is also a build dependency for any open source code. You can not realistically get the same code without it.
- Zambyte 2y ago> you can not realistically get the same code without it You mean the same source code? Because... I agree. That's why it's important for the source to be open. Both in the context of software and language models.