4 ms·
> I'm pretty sure we will solve this issue. Either already through World Models or another architecture. Please elaborate. How? With which technique? Currently
by germandiago 2mo ago
> I'm pretty sure we will solve this issue. Either already through World Models or another architecture.
Please elaborate. How? With which technique? Currently the only path forward is to feed more data and tweak for specific situations (fitting, basically). How does that help in the general case or in new situations with current tecchnology (LLMs, concretely). Noatter how far you get, this is not a general or reliable solution. It van only simulate more generality or more reliability by training and tweaking. Nothing else. At least, with this paradigm.
This does not mean they will not be useful. What I challenge here is the AGI or singularity. We are far from that.
> I'm now team lead for 10 years and every single year I teach them the same thing over and over and over again.
I have been a lead and an architect also for years at different position. I think you miss how much tacit knowledge and judgement there is inside the brains of each of us that an LLM is not capable of. And if it is, then you have to dumo so much context that it is better to go do it yourself. There is a cost to that also actually. It is not just so "dry and technical" the knowledge. Maybe yes to learn Java patterns or C++ constructors or the like.
But not for "given this situation with all these specifics", which solution would you bet on? Probably the LLM will give you a shitty REST API that is not what u need at all.So u tell the AI. It gives u something else generati g 30-50% of "decorated code". Now it seems to workso you use it. Now you do this every day. Come back in 2 months. You generated a lot of fat.
Now you have a bug. You do not know even where to start. Thisis theprice to payfor speed, as usual: technical debt.
Now you tell me you put three agents to talk and burn 2000 usd in tokens. Great! Is the final solution better than what you would have achieved? Not sure at all.
TBH I am not into agents bc I do not trust a tool sniffing all my code and for copyright concerns. But I saw some and use a prompt with limited access and the best I can take out for my speed + control when coding is tech discussions to decide on it, error catching, test generation, one-off scripts... But never "make an app like this or that". If I ever do that (I did it a couple of times) is for scaffolding and later throw away 70%.
Namely, to see something that runs on screen quickly. But later you need to spend time yourself as usual. Not a bad thing, just that this is not what you deliver and need the work done. Iterations etc.
- Yopolo 2mo ago> Please elaborate. How? With which technique? Reinforcement learning can just solve things even if they are new. It doesn't understand how a tool works? Give it a vm with the tool, a thousand agents and let it discover it automatically. Use the thumbs up/down emoji + chat analysis when a customer is unhappy, feed that to a RL Loop. The AI Researchers though work on World Models, grounding the AI and letting it simulate. It can do the simulation in parallel (unlimited) and choose what is best. > I think you miss how much tacit knowledge and judgement there is inside the brains of each of us that an LLM is not capable of. But thats my problem. Soooo many do not have this even as senior developers. > Now you have a bug. You do not know even where to start. Thisis theprice to payfor speed, as usual: technical debt. Yeah now i just ask the LLM to describe to me the bug. Works very well. > Now you tell me you put three agents to talk and burn 2000 usd in tokens. Great! Is the final solution better than what you would have achieved? Not sure at all. This is the thing. It only needs to make the team 10-30% better to compensate token budget with one work collegue. We have reached this level in my opinion already. Choosing a head count vs. choosing tokens. But it becomes cheaper and easier and better. So you will not just be able to do ith with 3 agents but with 20, 50 or 100. It will be better if your team is an avg team. It will be worse if you have a high profile team, for now. But man our industry has such a weird broad quality spectrum. > TBH I am not into agents bc I do not trust a tool sniffing all my code and for copyright concerns In worst case, my team always do code reviews, I do a code review on an ai instead of a human and adjust the harness or the infos the ai can access. I can actually work on making this workflow better and then i can clone it or spin it up for every single PR. For a human? I have to train them and they might leave. But there are plenty of cases were code doesn't matter. Researchers write a lot of random shitty uggly code as long as it does what it does, it doesn't matter. I have scripts for small tasks, we have microservices which do one thing because it is a tech stack we only need for one use case.
- germandiago 2mo ago> eah now i just ask the LLM to describe to me the bug. Works very well. I think you are confusing giving theories about what a bug might be with certainty. It does help bc it csn accelerste things, but many times I had AIs with challenging bugs throwing a lot of misleading theories to me. For the easier bugs, I was just as capable most of the time. Not every time, so there is some potential time saving there. But also time waste. As for research and fast prototyping you are right: I find it a good tool to explore bc yiu do not need the quality of a final product and researxh is in big part throwaway work. But I was talking about software that needs features, maintenance, etc. This is just not the same thing.