3 ms·
This is why I asked this question yesterday: "Ask HN: Why don't programming language foundations offer "smol" models?" https://news.ycombinator.com/item?id=45
by xrd 11mo ago
This is why I asked this question yesterday:
"Ask HN: Why don't programming language foundations offer "smol" models?"
https://news.ycombinator.com/item?id=45840078 https://news.ycombinator.com/item?id=45840078
If I could run smol single language models myself, I would not have to worry.
- XzAeRosho 11mo agoThe answer to most convenient solutions is money. There's no money in that.
- deleted 11mo ago[deleted]
- jazzyjackson 11mo agoAnd or, the lower parameter models are straight up less effective than the giants? Why is anyone paying for sonnet and opus if mixtral could do what they do?
- xrd 11mo agoBut, for example, Zig as a language has prominent corporate support. And, Mitchell Hashimoto is incredibly active and a billionaire. It feels like this would be a rational way to expand the usage of a language.
- xvector 11mo agoNo, it's because that's not how training an LLM works.
- trvz 11mo agoHave you even tried Qwen3-Coder-30B-A3B?
- embedding-shape 11mo ago> I wonder why I can't find a model that only does Python and is good only at that I don't think it's that easy. The times I've trained my own tiny models on just one language (programming or otherwise), they tend to get worse results than the models I've trained where I've chucked in all the languages I had at hand, even when testing just for single languages. It seems somewhat intuitive to me that it works like that too, programming in different (mainstream) languages is more similar than it's different (especially when 90% of all the source code is Algol-like), so makes sense there is a lot of cross-learning across languages.
- acedTrex 11mo agobecause a smol model that any of the nonprofits could feasibly afford to train would be useless for actual code generation. Hell, even the huge foundational models are still useless in most scenarios.