5 ms·
No we don't know your 'points' The Hallucinations are becoming less, significantly by now. It also might be already were it is cheaper for one of the big few
by Yopolo 2mo ago
No we don't know your 'points'
The Hallucinations are becoming less, significantly by now.
It also might be already were it is cheaper for one of the big few companies to spend millions and billions to teach the LLM / creating the training data necessary for an LLM to do something which it is not yet good enough due to the fact, that they sell this capability then to everyone who wants to use this capabilitiy.
We have not seen the end of Reinforcement Learning, which does need a lot less training data but more compute.
I'm 'vibing' on the side a handfull of small things, no LLM trained on particular what i'm asking to do. Its very capable of stringing together enough things so it can clearly follow handwavy things i tell it to do, analyse error messages, analysing screenshots etc. all by itself.
There is not a single real ceilling in sight, we only have clear barriers like compute but constant fast progress.
The field of mathematics went from 'useless' to 'you start better using it' to 'gamechanger' in how fast? 1 year after coding? less?
I want signes that we hit a real problem, instead I get cheaper tokens, Chinese models becoming very good as open models, new model updates from the others, mathematicans now saying how good it is etc.
Only half a year ago I had to babysit an LLM, now i tell it 1-3 sentences and it just goes and does it. And that stuff runs without compile errors etc.
If AI makes us 10% or 20% betteer, which is not that much, this alone will lead to companies reduing their expensive staff by 10-20%, which will has real impact on a job area. Some jobs are already hard to sell like cyber security and basic image tasks.
- germandiago 2mo ago> The Hallucinations are becoming less, significantly by now. Yes? What is the mega-solid technique that is used for it? Armies of people using curated data and reviewing it by hand? That is exactly one of my points: shifting the work elsewhere for specialized tasks. More replicable, improved, but, it scales infinitely and is autonomous? Can you assert that? I am not denying there is some use (a lot of uses!) for this, but this is more nuanced than just: oh, they will replace us. Not at all, that day, with the current technology, is not going to arrive. This is just a systematization, fitting and tweaking of human knowledge by curated data. It is not the one true superintelligence they are selling us. To begin with, they do not have a concept of truth, but of probabilistic truth. Only that poses already a very, very big problem for the path to perfection. > We have not seen the end of Reinforcement Learning, which does need a lot less training data but more compute. Noone said the opposite, but I would like to know at which cost and if it is feasible. We do not have even enough compute power for current technology. > . Its very capable of stringing together enough things so it can clearly follow handwavy things i tell it to do, analyse error messages, analysing screenshots etc. all by itself. I use it every day for these tasks and it works well BECAUSE I review the output and makes me go faster. It finds a lot of things I would have not found and it also hallucinates another handful of them, which confirms my point about AIs not being able to be fully autonomous in any future point in time unless tweaked exactly for the task, and even then, it can still miss judgement a human could have for edge cases. So I am not sure of how bad or good it can be compared to a human but I am pretty sure it cannot be more reliable than an expert in many situations. > Chinese models becoming very good as open models I think they will be better in the long term if they follow this path. Not absolutely better but when mixing with economics and the fact that no frontier model is totally reliable anyway... why pay a lot for something that needs human inspection anyway? > There is not a single real ceilling in sight, we only have clear barriers like compute but constant fast progress. The ceiling is the paradigm itself, as I mentioned above. There is not a single chance with current technology that something could become "generically knowledgeable" and "reliable" both at the same time. If it becomes generically knowledgeable and reliable, it is bc of data fed into it and curated and tweaked by humans. This is not an original idea from myself, there are armies of people doing this every day around the world, you can check. This is where a lot of improvement comes from. Can this be reused? Of course. It is a generic solution? No way. > Only half a year ago I had to babysit an LLM, now i tell it 1-3 sentences and it just goes and does it. And that stuff runs without compile errors etc. Yes, I also do one-off scripts like this and code snippets, even reviews and others. Now go design a full distributed system. Use agents if you want. We come back in six months and compare it to a system that was properly written and tested by humans and we can compare the quality on some grounds: 1. how long it takes to add new features? 2. which ones act more according to spec once added? 3. when adding features, which ones have more bugs? 4. in the face of an error, will the agent delete my whole AWS infra (count the money losses if possible also)? 5. will I understand (or need to understand, but I bet yes) this code at some point in the future? You have to count all that money also, not just I vibe coded something and it seemed to work. With full systems things become super messy. Now add the human factor of requirements and back and forth (iterations can be admittedly faster with AI, especially prototypes, but that comes with other costs also)... Not easy at all.
- Yopolo 2mo ago> shifting the work elsewhere for specialized tasks. More replicable, improved, but, it scales infinitely and is autonomous? Can you assert that? I would say yes and it will scale. It will either happen through central LLM just paying for it and scaling it up to everyone on the planet (literaly) OR by the agentic layer every big business is building into their systems. You needed some human to use your tool optimized for their company? With agentic layer you no longer need this. And if you look at companies like Google, they were pushing this notion for ages already because they saw an adoption problem of more 'complex' tools and trying to make it simpler and easier. Now you can act from the other side too. > We do not have even enough compute power for current technology. Exactly. Right? There is no ceiling if its clear that we don't even have the hrdware. But the hardware is a bottleneck not a ceiling. > The ceiling is the paradigm itself, as I mentioned above. There is not a single chance with current technology that something could become "generically knowledgeable" and "reliable" both at the same time. It doesn't need to be perfect, it only needs to be better than the avg human. And the current LLMs are already better than aat least 1-2 people in my team. > Now add the human factor of requirements and back and forth Yeah for now. Grill me skill made it a lot easier. Harness engineering is also being worked on, agentic layer, ai factories etc. And all of this can be copy and pasted. There is only one harness needed which becomes the expert security reviewer and tomorrow everyone can have it. I'm still discusing progress with LLMs with people and still not everyone is using it or playing around with harnesses or developgn an agentic layer. We still have a lot of work to do to even see how good it will become while it already is really good. People are already borred of AI today and making wrong decisions based on the current level of AI while i think we will see continues progress for years.
- germandiago 2mo ago> I would say yes and it will scale So you mean AI will be useful for any general job without lots of training for those jobs? How about new tasks? Tasks it has not been tweaked for. When I deviated from the average, and not really weird things, when programming, the output was way worse than average stuff. And this is an explicit target of AIs nowadays. I think you are missing a lot of details here, honestly. > Exactly. Right? There is no ceiling if its clear that we don't even have the hrdware. But the hardware is a bottleneck not a ceiling. No, the hardware is a bottleneck, the paradigm as we know it is a ceiling unless you massively and continuously feed this system with average tasks (which is useful). Which is exactly the opposite of what singularity and AGI have been promising. The systems we have now (unless the paradigm changes) will keep doing, essentially, fitting. No concept of truth and limited inference. That inference is based on already existing data, not on future data. In fact, there have been experiments about feeding output back to the input of LLMs and the degradation of the quality is very visible. If they are supposed to be so "intelligent", why it happens? > It doesn't need to be perfect I can agree that for lots of tasks it does not. But for others it is just not a tool good enough. > Yeah for now. Grill me skill made it a lot easier. Harness engineering is also being worked on, agentic layer, ai factories etc. I will not deny there could be progress, but nothing similar to "autonomous", "reliable", "super intelligence" or "singularity" with this paradigm. In fact, often in my experience, this is a waste of tokens for subpar results that shift the technical debt elsewhere. I mean if you try to develop full systems by "vibe-code like" techniques. If you use them judiciously, you can accelerate your workflow, maybe 2x, but not much beyond that if you want to have something worth to be used. Note that here I am talking about the full thing: with testing, quality, maintenance concerns and everything together. If you want to ship a sub-par thing that will go to the rubbish in a couple of months, then yes, you can do that. But that will fail commercially any way. Unless your job is convincing enough people that you can go 10x faster every time, deliver some sub-par thing, and find another customer, which, to me, would equal a scam.
- ubercore 2mo agoHallucinations were low hanging fruit in some ways. As someone working on a large-ish complex-ish distributed system that has to be maintained and support customers, it's still very high value to have Claude in the mix, but the core problem of needing to monitor, advise, course correct, and make sure you don't end up with more code and complexity than you need is, at least in my experience, still roughly the same. The sharp edges are being filed off very rapidly, but the core experience of "make and maintain a large system" isn't advancing nearly as fast, IMO.
- Yopolo 2mo agoI'm waiting for the agentic ai platform layer. We see AI factories going in this direction but there is no real 'the open source ai platform' thingy. It needs connectors to integrate with k8s, hyperscalers etc. it needs to be able to have a basic router, a way of configuring expert agents and interaction options for the human in the loop. There is for sure things we need to build or change, but it def feels like to me that it would immediadly fix a few things today. I'm not disappointed that it doesn't advance as fast as it feels