2 ms·
Please source this claim. What 1GB models are capable of has increased generation-on-generation. > For example: you can't make a mice-sized brain as smart as a
by Philpax 3mo ago
Please source this claim. What 1GB models are capable of has increased generation-on-generation.
> For example: you can't make a mice-sized brain as smart as a human brain no matter how hard you try.
Sure. We don't know where the ceiling is for our digital minds, though.
- himata4113 3mo agoThey have not increased in capabilities, they have increased in specialization. If you train a small model in another domain it will begin losing capabilities in the former domain. This is effectively the sigmoid problem. Although I will admit that if we discover a higher information density algorithm that it might change, but not by a substantial amount to where "super intelligence" in 1gb would be possible.
- Philpax 3mo agoOver the last two years, this weight class has doubled its scores and/or saturated several benchmarks in the Qwen lineup alone without loss of generality: https://claude.ai/public/artifacts/9f249169-3623-417e-86cd-771b9562f8ac https://claude.ai/public/artifacts/9f249169-3623-417e-86cd-7... There is undoubtedly a limit somewhere (there is only so much you can pack into a given size) but it's really not particularly clear where that limit is. I don't think it's superintelligence - that much I agree with you - but I think "We already have a 1gb model that is as capable as it will ever be" is strictly false.
- himata4113 3mo agoThe measured entropy of the model remains nearly unchanged though which means we have lost capabilities we have not measured, the model hasn't become "denser" it just became more specialized. It's like comparing two person A and B of similar intelligence where A is smarter and B is a genius at signing, but signing was not on the test so person A won.
- Philpax 3mo agoAssumes facts not in evidence. Please show your working.
- himata4113 3mo agohttps://en.wikipedia.org/wiki/Model_collapse https://en.wikipedia.org/wiki/Model_collapse - you want to use sigmoid 1.0, but the closer you are to 1.0 the higher the chance your model will collapse so you use 0.99-0.98, but those lead to data loss so after n passes all the original data becomes lost so you have a strict data limit there. The rest is just the general reality I am sure you are familiar with: - https://en.wikipedia.org/wiki/Catastrophic_interference https://en.wikipedia.org/wiki/Catastrophic_interference - https://en.wikipedia.org/wiki/Fine-tuning_(deep_learning) https://en.wikipedia.org/wiki/Fine-tuning_(deep_learning) - https://en.wikipedia.org/wiki/Entropy_(information_theory) https://en.wikipedia.org/wiki/Entropy_(information_theory)
- Philpax 3mo agoYou... haven't shown any evidence that we're near collapse. That's what I'm asking you for. Show me some evidence that we are losing capabilities with we have today.
- neuroticnews25 3mo ago> The measured entropy of the model remains nearly unchanged though which means we have lost capabilities Or we've lost random noise, or redundancy not captured by the entropy measurement.