4 ms·
FWIW I run a clinical analysis backend and directly compared Luna with DSv4 Flash 0731 - DS was a bit ahead, and cheaper even considering it used 40% more token
by flyinglizard 1mo ago
FWIW I run a clinical analysis backend and directly compared Luna with DSv4 Flash 0731 - DS was a bit ahead, and cheaper even considering it used 40% more tokens.
The exposure to Deepseek made me question the valuation house of cards built on SOTA providers. There are more companies producing competitive and useful models than there are companies producing jet engines for airliners, and not for the lack of trying. China has been trying to make these engines for decades and so far failed (their flagship C919 airliner is using CFM, American/French), but it has produced at least three competitive model companies within 3 years even though they're handicapped by their hardware.
It's simply not that hard, and diminishing returns will, in fact, diminish.
- andai 1mo agoYou reminded me of this recent post, which I found very interesting: Why jet engines aren't made in China https://news.ycombinator.com/item?id=48740971 https://news.ycombinator.com/item?id=48740971 That's a very interesting point of comparison though. I wonder why that should be the case? Why is it so much harder to make a jet engine than a language model? Maybe with LLMs the iteration times are lower? Or there's more public information about technique? Or are jet engines just intrinsically a harder problem?