3 ms·
The municipality of Rio de Janeiro (via its IT company IplanRIO) released Rio-3.5-Open-397B, presented as a homegrown Qwen3.5 fine-tune that beats comparable op
by unrvl22 4mo ago
The municipality of Rio de Janeiro (via its IT company IplanRIO) released Rio-3.5-Open-397B, presented as a homegrown Qwen3.5 fine-tune that beats comparable open models on benchmarks. The linked issue argues it's actually a weighted merge of ~60% Nex-N2 Pro + ~40% Qwen3.5-397B-A17B - Nex-N2 having been released about a week earlier.
- deleted 4mo ago[deleted]
- clear-octopus 4mo ago[dead]
- Lucasoato 4mo agoSo the problem isn’t in the missing attribution to Qwen, but with the fact that they didn’t mention Nex-N2 Pro right?
- Aurornis 4mo agoThe problem is that they claimed to have made a big achievement with their home grown post training, and they expected to receive a lot of praise for it. Then researchers looked at the weights and there is no post training at all. They are now attributing both models they merged, but their excuse for the lack of post training is to claim they accidentally uploaded the wrong files.
- serial_dev 4mo agoI’d believe they accidentally uploaded the wrong files if they uploaded the correct ones. To state that they accidentally uploaded something else and then not upload the correct version means they probably do not have anything and either hope people forget about this or they are scrambling to have something that is at least close to their original claim.
- evilduck 4mo ago"Oops, we uploaded the wrong files" is the standard deflection every time people like this get caught. Look up "Reflection 70B" drama.
- DonsDiscountGas 4mo agoI didn't know model merging like that was possible. (Obviously possible from a pure software standpoint but I'm surprised it's effective)
- bwhitty 4mo agoAs another poster above linked, it’s been shown to be effective since 2022: https://arxiv.org/abs/2203.05482 https://arxiv.org/abs/2203.05482
- nightpool 4mo agoit works because Nex N2 is also a derivative of the original base Qwen model. If it was two completely unrelated models it wouldn't work.
- hypercube33 4mo agoEven merging models with themselves as shown here in the post how they got to the top of hugging face with two gpus
- baobabKoodaa 4mo agoA few years back these used to be called "Frankenstein models"
- vasco 4mo agoRio better have the best IT infrastructure and software in the world if they are spending time on LLMs. What a waste of tax payer money.
- vitorgrs 4mo agoPiaui state it's also doing a LLM it seems. But indeed it would make more sense if it was a national thing rather than local...