5 ms·
With how stochastic the process is it makes it basically unusable for any large scale task. What's the plan? To roll the dice until the answer pops up? That wou
by margorczynski 1y ago
With how stochastic the process is it makes it basically unusable for any large scale task. What's the plan? To roll the dice until the answer pops up? That would be maybe viable if there was a way to automatically evaluate it 100% but with a human in the loop required it becomes untenable.
- diggan 1y ago> What's the plan? Call me old school, but I find the workflow of "divide and conquer" to be as helpful when working with LLMs, as without them. Although what is needed to be considered a "large scale task" varies by LLMs and implementation. Some models/implementations (seemingly Copilot) struggles with even the smallest change, while others breeze through them. Lots of trial and error is needed to find that line for each model/implementation :/
- mjburgess 1y agoThe relevant scale is the number of hard constraints on the solution code, not the size of task as measured by "hours it would take the median programmer to write". So eg., one line of code which needed to handle dozens of hard-constraints on the system (eg., using a specific class, method, with a specific device, specific memory management, etc.) will very rarely be output correctly by an LLM. Likewise "blank-page, vibe coding" can be very fast if "make me X" has only functional/soft-constraints on the code itself. "Gigawatt LLMs" have brute-forced there way to having a statistical system capable of usefully, if not universally, adhreading to one or two hard constraints. I'd imagine the dozen or so common in any existing application is well beyond a Terawatt range of training and inference cost.
- cyanydeez 1y agoKeep in mind that the model of using LLM assumes the underlying dataset converges to production ready code. Thats never been proven, cause we know they scraped sourcs code without attribution.
- nonethewiser 1y agoIts hard for me to think of a small, clearly defined coding problem an LLM cant solve.
- jodrellblank 1y ago"Find a counter example to the Collatz conjecture".
- mrguyorama 1y agoThere are several in the linked post, primarily: "Your code does not compile" and "Your tests fail" If you have to tell an intern that more than once on a single task, there's going to be conversations.
- safety1st 1y agoI mean I guess this isn't very ambitious, but it's a meaningful time saver if I basically just write code in natural language, and then Copilot generates the real code based on that. I don't have to look up syntax details, or what some function somewhere was named, etc. It will perform very accurately this way. It probably makes me 20% more efficient. It doubles my efficiency in a language I'm unfamiliar with. I can't fire half my dev org tomorrow with that approach, I can't really fire anyone, so I guess it would be a big letdown for a lot of execs. Meanwhile though we just keep incrementally shipping more stuff faster at higher quality so I'm happy... This works because it treats the LLM like what it actually is: an exceptionally good if slightly random text transformer.
- eterevsky 1y agoThe plan is to improve AI agents from their current ~intern level to a level of a good engineer.
- ethanol-brain 1y agoSeems like that is taking a very long time, on top of some very grandiose promises being delivered today.
- DrillShopper 1y agoThird AI Winter from overpromise/underdeliver when?
- rsynnott 1y agoThird? It’ll be the tenth or so.
- infecto 1y agoI look back over the past 2-3 years and am pretty amazed with how quick change and progress have been made. The promises are indeed large but the speed of progress has been fast. Not defending the promise but “taking a very long time” does not seem to be an accurate representation.
- owebmaster 1y ago> The promises are indeed large but the speed of progress has been fast And at the same time, absurdly slow? ChatGPT is almost 3 years old and pretty much AI has still no positive economic impact.
- infecto 1y agoSaying “AI has no economic impact” ignores reality. The financials of major players clearly show otherwise—both B2C and B2B applications are already profitable and proven. While APIs are still more experimental, and it’s unclear how much value businesses can ultimately extract from them, to claim there’s no economic impact is willful blindness. AGI may be far off, but companies are already figuring out value from both the consumer side and slowly API.
- rsynnott 1y agoI suspect that the plan is that MS has spent a lot, really a LOT, of money on this nonsense, and there is now significant pressure to put, something, anything, out even if it is worse than useless.
- Traubenfuchs 1y ago> to roll the dice This was discussed here https://news.ycombinator.com/item?id=43988913 https://news.ycombinator.com/item?id=43988913