3 ms·
This was nowhere near the top submission. But even if a solo engineer could get a top kernel, you don't think that having thousands of engineers, infinite token
by dzbarsky 2mo ago
This was nowhere near the top submission. But even if a solo engineer could get a top kernel, you don't think that having thousands of engineers, infinite tokens, and stronger models than are available to the public would give the labs a significant edge?
- genxy 2mo agoDoes the edge matter? I know you added significant as your hedge, but once you have feedback, your gain is largely irrelevant. Gain buys you bandwidth, so we are constructing systems run by the most powerful corporations where they are now optimizing for latency, as Archer says, do you want to flash crash civilization? This is how you do it.
- datakan 2mo agoI don't know. That just sounds like throwing money at a problem until it goes away. I'm not convinced that is the correct path forward.
- qeternity 2mo agoYou say this like that isn’t how the vast majority of problems are solved…
- dejavucoder 2mo ago1. labs have lots of inference capacity 2. they will have domain experts working on this so their efficiency is gonna be exponentially more (can direct LLM better, save money, reach same results faster)
- fooblaster 2mo agoYou can't exceed roofline performance on hardware. There is an performance cap you can hit. This recursive self improvement stuff lets you be closer to the pareto frontier, but the idea that it is leading to some exponential growth is a total pipe dream.
- dejavucoder 2mo agofair enough
- DANmode 2mo agoThat’s not what a moat is =] Which I believe was the word intended.
- dejavucoder 2mo agohello author here. yes, it gives labs edge and leads to self-recursive improvement loops. also i was myself able to finish 7th in a later competition with 2-3 other approaches which are variants of the method discussed in this blog. in general, having a harness as thin as possible with some problem specific instructions while controlling for context rot is the key. point i am trying to make is there are a lot of optimisation surface areas possible.
- dejavucoder 2mo agoyou may notice Kimi, GLM have also started telling how their model is able to optimise it's own inference pipeline https://www.kimi.com/blog/kimi-k3 https://www.kimi.com/blog/kimi-k3