3 ms·
You can open source dataset without all the details how it was assembled. Models are lossy compressed datasets you can pick up and amend (fine tune / continue
by mirekrusin 1mo ago
You can open source dataset without all the details how it was assembled.
Models are lossy compressed datasets you can pick up and amend (fine tune / continue training / alter) according to license they were released under.
Hy4 is released under OSI approved Apache License 2.0.
- kennywinker 1mo agoParent poster is technically right - open “source” implies the source used to make something is open. The model source is training data and code, not just weights. But the reality is, the weights are a useful artifact that you can use to create derivative works. So, dismissing it as a photoshop binary is as technically wrong as calling it open source.
- Alpha3031 1mo agoIIRC Nvidia claims to release enough data that it should be possible to fully reproduce Nemotron, so even if it's not as good as the current best models, GPT 5.1 or Opus 4.1.was still useful right? I guess it depends on what you wanted to do with them.
- LtWorf 1mo agoSo windows is open source because the binaries are a lossy compression of the original source?
- NitpickLawyer 1mo agoWeights are not binary. A model is created at init time, with random values. After that, it is being modified using data. The key point is that the labs modify the models "as weights". That means that weights are the intended / preferred way of modifying a model. Which, coincidentally, matches the definition of source in Apache 2.0. There is no "higher level" place where editing takes place. It all happens in weight space. Through the license you get the same rights as the lab that created it: view, inspect, run, modify, re-release. That's it. That's the only thing a license can grant you. The rest is semantics, misunderstandings, and FUD. A model released under an open source license is open source. Training data is lab knowhow / IP. Which, historically, has never been required for any open source release.
- frabcus 1mo agoWell, you can't add or alter data in pre-training from just the weights. Which, as I understand it, means you can't fundamentally increase core knowledge or cognitive ability, only what the model likes to do with those. You can only post-train, and you're subject as a result to catastrophic forgetting. To explain simply as far as I can tell (would love to be corrected) the large number of pre-training tokens only works because the documents are randomly ordered. So if you e.g. took a foundation model with open weights, then tried post-training it all the new data since its cut-off period, it would then end up over-trained on that new data, and forget older things.
- mirekrusin 1mo agoAs I live next to EPFL, I'll give you example from them: their Meditron-70B model is adapted to the medical domain from Llama-2-70B through continued pretraining. They took weights of Llama-2-70B and continued training on PubMed, medical guidelines and general data. Weights aren't just executable artifact that's consumed by users. Third parties actually use released parameter state as the editable starting point for further training and produce new foundation models from it.
- frabcus 1mo agoNice example - although it seems it ended up specialised in medical texts, so did indeed ("catastrophically") forget other knowledge?
- mirekrusin 1mo agoNo, it didn't. It lost 69.2% -> 67.8% on MMLU while improving medical performance. If you're trying to argue that loss of ~1.4 points is "catastrophic forgetting" (it's not) then look at later work, ie. Me-LLaMA that clearly demonstrates continued pretraining that improved both general MMLU and medical performance. Not sure why you're fixating on catastrophic forgetting. How do you think model training works? Model training is just a sequence of checkpoints: pretraining produces it, training resumes from last, continued pretraining starts from last, supervised fine tuning starts from last, RL/post-training starts from last - it's just a sequence of checkpoints. There isn't some fundamental distinction where original author continuing training from checkpoint X is training but a third party downloading checkpoint X and continuing training from it suddenly isn't. ie. checkpoint doesn't somehow become a different kind of artifact when it's published.
- petu 1mo agoBefore we worry about source code, Microsoft doesn't grant me rights to modify/redistribute/sell copy of Windows I have.
- mirekrusin 1mo agoYou can't take windows binaries and continue development on them. Model weight release is a snapshot/checkpoint you can take and resume training on new data, producing new model. You don't need original training history to modify it further.
- LtWorf 1mo ago> You can't take windows binaries and continue development on them. It's easier than working on weights.
- mirekrusin 1mo agoReverse engineering binaries is easier than continuing training on weights? You're joking, right? Labs themselves use weight snapshots, that's how you do training, it's normal part of training process.
- villish 1mo agoCountries that aren’t competitive need access to training datasets so that they may train their own similarly capable models and be sure of the inputs. Governments cannot blindly trust open weight models from China and the US.