4 ms·
So the problem with Chris’ take is “This one for fun project didn’t produce anything particularly interesting.” So outside of the fact that we have magic now t
by conception 7mo ago
So the problem with Chris’ take is “This one for fun project didn’t produce anything particularly interesting.”
So outside of the fact that we have magic now that can just produce “conventional “ compilers. Take it to a Moore’s Law situation. Start 1000 create a compiler projects- have each have a temperature to try new things, experiment, mutate. Collate - find new findings - reiterate- another 1000 runs with some of the novel findings. Assume this is effectively free to do.
The stance that this - which can be done (albeit badly) today and will get better and/or cheaper - won’t produce new directions for software engineering seems entirely naive.
- runarberg 7mo agoMoors law states that the number of transistors in an integrated circuit doubles about every two years. It has nothing to say about the capabilities of statistical models. In fact in statistics we have another law which states that as you increase parameters the more you risk overfitting. And overfitting seems to already be a major problem with state of the art LLM models. When you start overfitting you are pretty much just re-creating stuff which is already in the dataset.
- vidarh 7mo agoIn their example it doesn't matter is this case if the models get better or not. It matters whether inference gets cheaper to the point that we can afford to basically throw huge amounts of tokens at exploring the problem space. Further model improvements would be a bonus, but it's not required for us to get much further.
- warkdarrior 7mo agoModern LLMs showed that overfitting disappears if you add more and more parameters. "Double descent" is well documented, if not well understood.
- runarberg 7mo ago> Modern LLMs showed that overfitting disappears if you add more and more parameters. I have not seen that. In fact this is the first time I hear this claim, and frankly it sounds ludicrous. I don‘t know how modern LLMs are dealing with overfitting but I would guess there is simply a content matching algorithm after the inference, and if there is a copyright match the program does something to alter or block the generation. That is, I suspect the overfitting prevention is algorithmic and not part of the model.