7 ms·
Question for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs? I lead a team
by LASR 2y ago
Question for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs?
I lead a team exploring cutting edge LLM applications and end-user features. It's my intuition from experience that we have a LONG way to go.
GPT-4o / Claude 3.5 are the go-to models for my team. Every combination of technical investment + LLMs yields a new list of potential applications.
For example, combining a human-moderated knowledge graph with an LLM with RAG allows you to build "expert bots" that understand your business context / your codebase / your specific processes and act almost human-like similar to a coworker in your team.
If you now give it some predictive / simulation capability - eg: simulate the execution of a task or project like creating a github PR code change, and test against an expert bot above for code review, you can have LLMs create reasonable code changes, with automatic review / iteration etc.
Similarly there are many more capabilities that you can ladder on and expose into LLMs to give you increasingly productive outputs from them.
Chasing after model improvements and "GPT-5 will be PHD-level" is moot imo. When did you hire a PHD coworker and they were productive on day-0 ? You need to onboard them with human expertise, and then give them execution space / long-term memories etc to be productive.
Model vendors might struggle to build something more intelligent. But my point is that we already have so much intelligence and we don't know what to do with that. There is a LOT you can do with high-schooler level intelligence at super-human scale.
Take a naive example. 200k context windows are now available. Most people, through ChatGPT, type out maybe 1500 tokens. That's a huge amount of untapped capacity. No human is going to type out 200k of context. Hence why we need RAG, and additional forms of input (eg: simulation outcomes) to fully leverage that.
- amelius 2y agoYes, but literally anybody can do all those things. So while there will be many opportunities for new features (new ways of combining data), there will be few business opportunities.
- Miraste 2y agoHN always says this, and it's always wrong. A technical implementation that's easy, or readily available, does not mean that a successful company can't be built on it. Last year, people were saying "OpenAI doesn't have a moat." 15 years before that, they were saying "Dropbox is just a couple of chron jobs, it'll fail in a few months."
- amelius 2y ago> HN always says this The meaning here is different. What I'm saying is that big companies like OpenAI will always strive to make a generic AI, such that anyone can do basically anything using AI. The big companies therefore will indeed (like you say) have a profitable business, but few others will.
- hartator 2y agoAll of these hacks do sound like we are at that diminishing return point.
- namaria 2y agoIt all just sounds to me like we're back at expert systems. Doesn't bode well...
- ianbutler 2y agoHonest question, how would you expect systems to get external knowledge etc without tools like the OP is suggesting? Action oriented through self exploration? What is your thought for how these systems integrate with the existing world? Why does the OP's suggested mode of integration make you think of those older systems?
- namaria 2y agoThe premise of deep learning is the automated 'absorption' of knowledge. If we're back to curating it by hand and imparting it by writing code manually, how exactly are these systems an improvement on the 80's idea of building expert systems?
- brookst 2y agoHey look, it's Gordon Moore visiting us from 2005! :)
- crystal_revenge 2y agoI don't think we've even started to get the most value out of current gen LLMs. For starters very few people are even looking at sampling which is a major part of the model performance. The theory behind these models so aggressively lags the engineering that I suspect there are many major improvements to be found just by understanding a bit more about what these models are really doing and making re-designs based on that. I highly encourage anyone seriously interested in LLMs to start spending more time in the open model space where you can really take a look inside and play around with the internals. Even if you don't have the resources for model training, I feel personally understanding sampling and other potential tweaks to the model (lots of neat work on uncertainty estimations, manipulating the initial embedding the prompts are assigned, intelligent backtracking, etc). And from a practical side I've started to realize that many people have been holding on of building things waiting for "that next big update", but there a so many small, annoying tasks that can be easily automated.
- dr_dshiv 2y ago> I've started to realize that many people have been holding on of building things waiting for "that next big update" I’ve noticed this too — I’ve been calling it intellectual deflation. By analogy, why spend now when it may be cheaper in a month? Why do the work now, when it will be easier in a month?
- vbezhenar 2y agoWhy optimise software today, when tomorrow Intel will release CPU with 2x performance?
- sdenton4 2y agoCuriously, Moore's law was predictable enough over decades that you could actually plan for the speed of next year's hardware quite reliably. For LLMs, we don't even know how to reliably measure performance, much less plan for expected improvements.
- 2y ago
- msabalau 2y agoThere are all sorts of valuable things to explore and build with what we have already. But understanding how likely it is that we will (or will not) see a new models quickly and dramatically improve on what we have "because scaling" seems valuable context for everyone in ecosystem to make decisions.
- ben_w 2y ago> Question for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs? IMO we've not even exhausted the options for spreadsheets, let alone LLMs. And the reason I'm thinking of spreadsheets is that they, like LLMs, are very hard to win big on even despite the value they bring. Not "no moat" (that gets parroted stochastically in threads like these), but the moat is elsewhere.
- alach11 2y agoMy team and I also develop with these models every day, and I completely agree. If models stall at current levels, it will take 10 (or more) years for us to capture most of the value they offer. There's so much work out there to automate and so many workflows to enhance with these "not quite AGI-level" models. And if peak model performance remains the same but cost continues to drop, that opens up vastly more applications as well.
- alangibson 2y agoI think you're playing a different game than the Sam Altmans of the world. The level of investment and profit they are looking for can only be justified by creating AGI. The > 100 P/E ratios we are already seeing can't be justified by something as quotidian as the exceptionally good productivity tools you're talking about.
- gizajob 2y agoYeah I keep thinking this - how is Nvidia worth $3.5Trillion for making code autocomplete for coders
- drawnwren 2y agoNvidia was not the best example. They get to moon in the case that any AI exponential hits. Most others have less of a wide probability distribution.
- BeefWellington 2y agoYeah they're the shovel sellers of this particular goldrush. Most other businesses trying to actually use LLMs are the riskier ones, including OpenAI, IMO (though OpenAI is perhaps the least risky due to brand recognition).
- hluska 2y agoNowhere near, but the market seems to have priced in that scaling would continue to have a near linear effect on capability. That’s not happening and that’s the issue the article is concerned with.
- HarHarVeryFunny 2y agoSure, there's going to be a lot of automation that can be built using current GPT-4 level LLMs, even if they don't get much better from here. However, this is better thought of as "business logic scripting/automation", not the magic employee-replacing AGI that would be the revolution some people are expecting. Maybe you can now build a slightly less shitty automated telephone response system to piss your customers off with.
- brookst 2y ago> Question for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs? Certainly not. But technology is all about stacks. Each layer strives to improve, right up through UX and business value. The uses for 1µm chips had not been exhausted in 1989 when the 486 shipped in 800nm. 250nm still had tons of unexplored uses when the Pentium 4 shipped on 90nm. Talking about scaling at the the model level is like talking about transistor density for silicon: it's interesting, and relevant, and we should care... but it is not the sole determinent of what use cases can be build and what user value there is.
- senko 2y agoNo. The scaling laws may be dead. Does this mean the end of LLM advances? Absolutely not. There are many different ways to improve LLM capabilities. Everyone was mostly focused on the scaling laws because that worked extremely well (actually surprising most of the researchers). But if you're keeping an eye on the scientific papers coming out about AI, you've seen the astounding amount of research going on with some very good results, that'll probably take at least several months to trickle down to production systems. Thousands of extremely bright people in AI labs all across the world are working on finding the next trick that boosts AI. One random example is test-time compute: just give the AI more time to think. This is basically what O1 does. A recent research paper suggests using it is roughly equivalent to an order of magnitude more parameters, performance wise. (source for the curious: https://lnkd.in/duDST65P https://lnkd.in/duDST65P) Another example that sounds bonkers but apparently works is quantization: reducing the precision of each parameter to 1.58 bits (ie only using values -1, 0, 1). This uses 10x less space for the same parameter count (compared to standard 16-bit format), and since AI operatons are actually memory limited, directly corresponds to 10x decrease in costs: https://lnkd.in/ddvuzaYp https://lnkd.in/ddvuzaYp (Quite apart from improvements like these, we shouldn't forget that not all AIs are LLMs. There's been tremendous advance in AI systems for image, audio and video generation, interpretation and munipulation and they also don't show signs of stopping, and there's possibility that a new or hybrid architecture for the textual AI might be developed). AI winter is a long way off.
- limaoscarjuliet 2y agoScaling laws are not dead. The number of people predicting death of Moore's law doubles every two years. - Jim Keller https://www.youtube.com/live/oIG9ztQw2Gc?si=oaK2zjSBxq2N-zj1&t=451 https://www.youtube.com/live/oIG9ztQw2Gc?si=oaK2zjSBxq2N-zj1...
- nyrikki 2y agoThere are way too many personal definitions of what "Moore's Law" even is to have a discussion without deciding on a shared definition before hand. But Goodhart's law; "When a measure becomes a target, it ceases to be a good measure" Directly applies here, Moore's Law was used to set long term plans at semiconductor companies, and Moore didn't have empirical evidence it was even going to continue. If you say, arbitrarily decide CPU, or worse, single core performance as your measurement, it hasn't held for well over a decade. If you hold minimum feature size without regard to cost, it is still holding. What you want to prove usually dictates what interpretation you make. That said, the scaling law is still unknown, but you can game it as much as you want in similar ways. GPT4 was already hinting at an asymptote on MMLU, but the question is if it is valid for real work etc... Time will tell, but I am seeing far less optimism from my sources, but that is just anecdotal.
- afro88 2y ago> potential applications > if you ... > for example ... Yes there seems to be lots of potential. Yes we can brainstorm things that should work. Yes there is a lot of examples of incredible things in isolation. But it's a little bit like those youtube videos showing amazing basketball shots in 1 try, when in reality lots of failed attempts happened beforehand. Except our users experience the failed attempts (LLM replies that are wrong, even when backed by RAG) and it's incredibly hard to hide those from them. Show me the things you / your team has actually built that has decent retention and metrics concretely proving efficiency improvements. LLMs are so hit and miss from query to query that if your users don't have a sixth sense for a miss vs a hit, there may not be any efficiency improvement. It's a really hard problem with LLM based tools. There is so much hype right now and people showing cherry picked examples.
- jihadjihad 2y ago> Except our users experience the failed attempts (LLM replies that are wrong, even when backed by RAG) and it's incredibly hard to hide those from them. This has been my team's experience (and frustration) as well, and has led us to look at using LLMs for classifying / structuring, but not entrusting an LLM with making a decision based on things like a database schema or business logic. I think the technology and tooling will get there, but the enormous amount of effort spent trying to get the system to "do the right thing" and the nondeterministic nature have really put us into a camp of "let's only allow the LLM to do things we know it is rock-solid at."
- sdesol 2y ago> "let's only allow the LLM to do things we know it is rock-solid at." Even this is insanely hard in my opinion. The one thing that you would assume LLM to excel at is spelling and grammar checking for the English language, but even the top model (GPT-4o) can be insanely stupid/unpredictable at times. Take the following example from my tool: https://app.gitsense.com/?doc=6c9bada92&model=GPT-4o&samples=5 https://app.gitsense.com/?doc=6c9bada92&model=GPT-4o&samples... 5 models are asked if the sentence is correct and GPT-4o got it wrong all 5 times. It keeps complaining that GitHub is spelled like Github, when it isn't. Note, only 2 weeks ago, Claude 3.5 Sonnet did the same thing. I do believe LLM is a game changer, but I'm not convinced it is designed to be public-facing. I see LLM as a power tool for domain experts, and you have to assume whatever it spits out may be wrong, and your process should allow for it. Edit: I should add that I'm convinced that not one single model will rule them all. I believe there will be 4 or 5 models that everybody will use and each will be used to challenge one another for accuracy and confidence.
- whiplash451 2y agoThe main difference between GPT5 and a PhD-level new hire is that the new hire will autonomously go out, deliver and take on harder task with much fewer guidance than GPT5 will ever require. So much of human intelligence is about interacting with peers.
- ben_w 2y agoHuman interaction with peers is also guidance. I don't know how many team meetings PhD students have, but I do know about software development jobs with 15 minute daily standups, and that length meeting at 120 words per minute for 5 days a week, 48 weeks per year of a 3 year PhD is 1.296.000 words.
- eastbound 2y agoI have 3 remote employees whose job is consistently as bad as LLM. That means employees who use LLM are, on average, recognizably bad. Those who are good enough, are also good enough to write the code manually. To the point I wonder whether this HN thread is generated by OpenAI, trying to create buzz around AI.
- ben_w 2y ago1. The person I'm replying to is hypothesising about a future, not yet existent, version, GPT5. Current quality limits don't tell you jack about a hypothetical future, especially one that may not ever happen because money. 2. I'm not commenting on the quality, because they were writing about something that doesn't exist and therefore that's clearly just a given for the discussion. The only thing I was adding is that humans also need guidance, and quite a lot of it — even just a two-week sprint's worth of 15 minute daily stand-up meetings is 18,000 words, which is well beyond the point where I'd have given up prompting an LLM and done the thing myself.
- whiplash451 2y agoThey definitely tell you jack. GPTs have reach their glass ceiling as they’ve sucked all available data and overfit to benchmarks. Their models have tons of use cases, but OpenAI and Anthropic are now in a product/commercial play.
- EGreg 2y agoI want to stuff a transcript of a 3 hour podcast into some LLM API and have it summarize it by: segmenting by topic changes, keeping the timestamps, and then summarizing each segment. I wasn’t able to get it do it with Anthropic or OpenAI chat completion APIs. Can someone explain why? I don’t think the 200K token window actually works, is it looking sequentially or is it really looking at the whole thing at once or something?
- anonzzzies 2y agoThe current models are very powerful and we definitely didn't get most out of them yet. We are getting more and more out of them every week when we release new versions of our toolkits. So if this is it; please make it faster and take less energy. We'll be fine until the next AI spring.
- simonw 2y agoRight. I've been saying for a while that if all LLM development stopped entirely and we were stuck with the models we have right now (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama 3.1/2, Qwen 2.5 etc) we could still get multiple years worth of advances just out of those existing models. There is SO MUCH we haven't figured out about how to use them yet.
- dgfitz 2y agoLLMs use historic data to help create useful current data. It works well sometimes. I find that a human is able to solve a P=NP situation, and an LLM can’t quite yet do that. When they can the game changes.
- niobe 2y ago> There is SO MUCH we haven't figured out about how to use them yet. I mean, it's pretty clear to me they're a potentially great human-machine interface, but trying to make LLMs - in their current fundamental form - a reliable computational tool.. well, at best it's an expensive hack, but it's just not the right tool for the job. I expect the next leap forward will require some orthogonal discovery and lead to a different kind of tool. But perhaps we'll continue to use LLMs as we knownthem now for what they're good at - language.
- simonw 2y agoOne of the biggest challenges in learning how to use and build on LLMs is figuring out how to work productively with a technology that - unlike most computers - is inherently unreliable and non-deterministic. It's possible, but it's not at all obvious and requires a slightly skewed way of looking at them.
- XenophileJKO 2y agoThis really reminds me of a trend years ago to create probabilistic programming constructs. I think it was just a trend way ahead of its time. Typical software engineers tend to be very ill-suited to think in probabilities and how to build reasonably reliable systems around them.
- 23B1 2y agoThe user interface for LLMs is stuck in C:\ That's where I'd focus.
- kenjackson 2y agoVoice for LLMs is surprisingly good. I'd love to see LLMs used in more systems like cars and in-home automation. Whatever cars use today and Alexa in the home simply are much worse than what we get with ChatGPT voice today.
- ericmcer 2y agoI have tried a few AI coding tools and always found them impressive but I don't really need something to autocomplete obvious code cases. Is there an AI tool that can ingest a codebase and locate code based on abstract questions? Like: "I need to invalidate customers who haven't logged in for a month" and it can locate things like relevant DB tables, controllers, services, etc.
- james_marks 2y agoI haven’t seen quite that, but it’s an interesting question; like a semantic search.
- fullstackchris 2y agoCursor (Claude behind the scenes) can do that, however as always, your mileage may vary. I tried building a whole codebase inspector, essentially what you are referring to with Gemini's 2 million token context window but had troubles with their API when the payload got large. Just 500 error with no additional info so...
- disgruntledphd2 2y agoI've played around with Claude and larger docs and it's honestly been a bit of a crapshoot, it feels like only some of the information gets into the prompt as the doc gets larger. They're great for converting PDF tables to more usable formats though.
- spacebanana7 2y agoIt's ugly, but I've had some success with uploading a few files from a project and a sketch of the schema. Then asking for new functionality. ChatGPT and Claude seem to be pretty good at maintaining an implicit understanding of the codebase based on a subset of files.
- yk 2y agoTo a certain extent I think we get a better understanding what llms can do, and my estimation for the next ten years is more like best UI ever rather than llms will replace humanity. Now best UI ever is something that can certainly deliver a lot of value, 80% of all buttons in a car should be replaced by actually good voice control, and I think that is were we are going to see a lot of very interesting applications: Hey washing machine, this is two t-shirts and a jeans. (The washing machine can then figure out it's program by itself, I don't want to memorize the table in the manual.)
- lokimedes 2y agoTo each their own, but I don’t look forward to having my kids yelling, a podcast in my ears and having to explain to my tumbler that wool must be spun at 1000 RPM. Humans have varying preferences when it comes to communication and sensing, making our machine interactions favor the extroverted talkative exhibitionists is really only one modality.
- weweersdfsd 2y agoI think buttons should not be replaced, but rather augmented with voice control. I certainly want to be able to adjust air conditioning or use my washing machine while listening music or having otherwise noisy environemnt.
- machiaweliczny 2y agoLong context is a scam. Claude is best but it’s still gets lost with longer context
- bbor 2y agoI have no data, but I whole-heartedly agree. Well, perhaps not “scam”, but definitely oversold. One of my best undergrad professors taught me the adage “don’t expect a model to do what a human expert cannot”, and I think it’s still a good rule of thumb. Giving someone an entire book to read before answering your question might help, but it would help way, way more to give them a few paragraphs that you know are actually relevant.
- cruffle_duffle 2y agoIn my experience, the reality of long context windows doesn’t live up to the hype. When you’re iterating on something, whether it's code, text, or any document, you end up with multiple versions layered in the context. Every time you revise, those earlier versions stick around, even though only the latest one is the "most correct". What gets pushed out isn’t the last version of the document itself (since it’s FIFO), but the important parts of the conversation—things like the rationale, requirements, or any context the model needs to understand why it’s making changes. So, instead of being helpful, that extra capacity just gets filled with old, repetitive chunks that have to be processed every time, muddying up the output. This isn’t just an issue with code; it happens with any kind of document editing where you’re going back and forth, trying to refine the result. Sometimes I feel the way to "resolve" this is to instead go back and edit some earlier portion of the chat to update it with the "new requirements" that I didn't even know I had until I walked down some rabbit hole. What I end up with is almost like a threaded conversation with the LLM. Like, I sometimes wish these LLM chatbots explicitly treated the conversion as if it were threaded. They do support basically my use case by letting you toggle between different edits to your prompts, but it is pretty limited and you cannot go back and edit things if you do some operations (eg: attach a file). Speaking of context, it's also hard to know what things like ChatGPT add to it's context in the first place. Many of times I'll attach a file or something and discover it didn't "read" the file into it's context. Or I'll watch it fire up a python program it writes that does nothing but echo the file into it's context. I think there is still a lot of untapped potential in strategically manipulating what gets placed into the context window at all. For example only present the LLM with the latest and greatest of a document and not all the previous revisions in the thread.
- bbor 2y agoGreat question. Im very confident in my answer, even though it’s in the minority here: we’re not even close to exhausting the potential. Imagine that our current capabilities are like the Model-T. There remains many improvements to be made upon this passenger transportation product, with RAG being a great common theme among them. People will use chatbots with much more permissive interfaces instead of clicking through menus. But all of that’s just the start, the short term, the maturation of this consumer product; the really scary/exciting part comes when the technology reaches saturation, and opens up new possibilities for itself. In the Model-T metaphor, this is analogous to how highways have (arguably) transformed America beyond anyone’s wildest dreams, changing the course of various historical events (eg WWII industrialization, 60s & 70s white flight, early 2000s housing crisis) so much it’s hard to imagine what the country would look like without them. Now, automobiles are not simply passenger transportation, but the bedrock of our commerce, our military, and probably more — through ubiquity alone they unlocked new forms of themselves. For those doubting my utopian/apocalyptic rhetoric, I implore you to ask yourself one simple question: why are so many experts so worried about AGI? They’ve been leaving in droves from OpenAI, and that’s ultimately what the governance kerfluffle there was. Hinton, a Turing award winner, gave up $$$ to doom-say full time. Why? My hint is that if your answer involves less then a 1000 specialized LLMs per unified system, then you’re not thinking big enough.
- fire_lake 2y ago> Hinton, a Turing award winner, gave up $$$ to doom-say full time This is a hint of something but a weak argument. Smart people are wrong all the time.
- mrandish 2y ago> why are so many experts so worried about AGI? FYI, I find this line of reasoning to be unconvincing both logically and by counter-example ("why are so many experts so worried about the Y2K bug?") Personally, I don't find AI foom or AI doom predictions to be probable but I do think there are more convincing arguments for your position than you're making here.
- bbor 2y ago
- robrenaud 2y ago> For example, combining a human-moderated knowledge graph with an LLM with RAG allows you to build "expert bots" that understand your business context / your codebase / your specific processes and act almost human-like similar to a coworker in your team. I'd love to hear about this. I applied to YC WC 25 with research/insight/an initial researchy prototype built on top of GPT4+finetuning about something along this idea. Less powerful than you describe, but it also works without the human moderated KG.
- bloppe 2y ago> you can have LLMs create reasonable code changes, with automatic review / iteration etc. Nobody who takes code health and sustainability seriously wants to hear this. You absolutely do not want to be in a position where something breaks, but your last 50 commits were all written and reviewed by an LLM. Now you have to go back and review them all with human eyes just to get a handle on how things broke, while customers suffer. At this scale, it's an effort multiplier, not an effort reducer. It's still good for generating little bits of boilerplate, though.
- Aeolun 2y agoIf the last 50 commits were reviewed by an AI and it took that long for an issue to happen I’d immediately mandate all PR’s are reviewed by an AI.
- moogly 2y ago> you can have LLMs create reasonable code changes Could you define "code changes" because I feel that is a very vague accomplishment.
- nonameiguess 2y agoYour hypothesis here is not exclusive of the hypothesis in this article. Name your platform. Linux. C++. The Internet. The x86 processor architecture. We haven't exhausted the options for delivering value on top of those, but that doesn't mean the developers and sellers of those platforms don't try to improve them anyway and might struggle to extract value from application developers who use them.
- purple-leafy 2y agoDoesn’t sound cutting edge at all? Every man and his dog is doing a similar process
- rco8786 2y agoI think there’s a long way to go also. I think people expected that AI would eventually be like a “point and shoot” where you would tell it to go do some complicated task, or sillier yet, take over someone’s entire job. More realistically it’s like a really great sidekick for doing very specific mundane but otherwise non deterministic tasks. I think we’ll start to see AI permeate into nearly every back office job out there, but as a series of tools that help the human work faster. Not as one big brain that replaces the human.
- RayVR 2y agoI am definitely not an expert, nor do I have inside information on the directions of research that these companies are exploring. Yes, existing LLMs are useful. Yes, there are many more things we can do with this tech. However, existing SOTA models are large, expensive to run, still hallucinate, fail simple logic tests, fail to do things a poorly trained human can do on autopilot, etc. The performance of LLMs is extremely variable, and it is hard to anticipate failure. Many potential applications of this technology will not tolerate this level of uncertainty. Worse solutions with predictable and well understood shortcomings will dominate.
- jeswin 2y agoIn my view, an escape hatch if we are truly stuck would be radical speed ups (like Cerebras) in compute time. If we get outputs in milli-seconds instead of seconds and at much lower costs, it would make backtracking viable. This won't allow AGI, but can make a new class of apps possible.
- Lonestar1440 2y agoNo, we have not even scratched the surface of what current-gen LLMs can do for an organization which puts the correct data into them. If indeed the "GPT 5!" Arms race has calmed down, it should help everyone focus on the possible, their own goals, and thus what AI capabilities to deploy. Just as there won't be a "Silver Bullet" next gen model, the point about Correct Data In is also crucial. Nothing is 'free' not even if you pay a vendor or integrator. You, the decision making organization, must dedicate focus to putting data into your new AI systems or not. It will look like the dawn of original IBM, and mechanical data tabulation, in retrospect once we learn how to leverage this pattern to its full potential.
- hamburga 2y agoI think there's a ton to be tapped based on the current state of the art. As a developer, I'm making much more progress using the SOTA (Claude 3.5) as a Socratic interrogator. I'm brainstorming a project, give it my current thoughts, and then ask it to prompt me with good follow-up questions and turn general ideas into a specific, detailed project plan, next steps, open questions, and work log template. Huge productivity boost, but definitely not replacing me as an engineer. I specifically prompt it to not give me solutions, but rather, to just ask good questions. I've also used Claude 3.5 as (more or less) a free arbitrator. Last week, I was in a disagreement with a colleague, who was clearly being disingenuous by offering to do something she later reneged on, and evading questions about follow up. Rather than deal with organizational politics, I sent the transcript to Claude for an unbiased evaluation, and it "objectively" confirmed what had been frustrating me. I think there's a huge opportunity here to use these things to detect and call out obviously antisocial behavior in organizations (my CEO is intrigued, we'll see where it goes). Similarly, in our legal system, as an ultra-low-cost arbitrator or judge for minor disputes (that could of course be appealed to human judges). Seems like the level of reasoning in Claude 3.5 is good enough for that. My mental model is always "low-risk search". https://muldoon.cloud/2023/10/29/ai-commandments.html https://muldoon.cloud/2023/10/29/ai-commandments.html
- soheil 2y agoWe have not exhausted what html can do either. LLMs not getting smarter is orthogonal to its currently unexplored search space.
- Roger-L 2y agoYes, I personally think that training an "all-knowing" artificial intelligence is not as good as training n "experts" in a single field.
- sky2224 2y ago> do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs? I know we absolutely have not, but I think we have reached the limit in terms of the Chatbot experience that ChatGPT is. For some reason the industry keeps trying to force the chatbot interface to do literally everything to the point that we now have inflated roles like "Prompt Engineers". This is to say that people suck at knowing what they want off the rip, and LLMs can't help with that if they're not integrated in technology in such a way where a solid foundation is built to allow the models to generate good output. LLMs and other big data models have incredible potential for things like security, medicine, and the power industry to name a few fields. I mean I was recently talking with a professor about his research in applying deep learning to address growing security concerns in cars on the road. The application is far from reaching the ceiling.
- amw-zero 2y agoWe might not have exhausted their applications, but everything I’ve witnessed them being used for has been extremely disappointing. That is, other than me using them to bounce ideas off of and create small snippets of code.
- zmmmmm 2y ago> combining a human-moderated knowledge graph with an LLM with RAG allows you to build "expert bots" that understand your business context / your codebase / your specific processes and act almost human-like similar to a coworker in your team It's been a while though, we've had great models now for a 18 months plus. Why are we still yet to see these type of applications rolling out on a wide scale? My anecdotal experience is that almost universally, 90-95% type accuracy you get from them is just not good enough. Which is to say, having something be wrong 10% or even 5% of the time is worse than not having at all. At best, you need to implement applications like that in an entirely new paradigm that is designed to extract value without bearing the costs of the risks. It doesn't mean LLMs can't be useful, but they are kind of stuck with applications that inherently mesh with human oversight (like programming etc). And the thing about those is that they don't really scale, because the human oversight has to scale up with whatever the LLM is doing.
- mycall 2y agoWe are just scratching the surface of what LLMs can do. Case in point, ESM3. https://www.biorxiv.org/content/10.1101/2024.07.01.600583v1 https://www.biorxiv.org/content/10.1101/2024.07.01.600583v1
- reissbaker 2y agoBeyond just RAG, I'm fairly bullish on finetuning. For example, Qwen2.5-Coder-32B-Instruct is much better than Qwen2.5-72B-Instruct at coding... Despite simply being a smaller version of the same model, finetuned on code. It's on par with Sonnet 3.5 and 4o on most benchmarks, whereas the simple chat-tuned 72B model is much weaker. And while Qwen2.5-Coder-32B-Instruct is a pretty advanced finetune — it was trained on an extra 5 trillion tokens — even smaller finetunes have done really well. For example, Dracarys-72B, which was a simpler finetune of Qwen2.5-72B using a modified version of DPO on a handmade set of answers to GSM8K, ARC, and HellaSwag, significantly outperforms the base Qwen2.5-72B model on the aider coding benchmarks. There's a lot of intelligence we're leaving on the floor, because everyone is just prompting generic chat-tuned models! If you tune it to do something else, it'll be really good at the something else.
- jalapenos 2y agoWell I have a question for you: do you think this format of AI can actually think? I.e. can it ruminate on the data it's ingested, and rather than returning the response of highest probability, return something original? I think that's the key. If LLMs can't ultimately do that, there's still a lot to be gained from utilising the speed and fluidly scalable resources of computers. But like all the top tech companies know, it's not quantity of bodies in seats that matters but talent, the thing that's going to prevail is raw intelligence. If it can't think better than us, just process data faster and more voluminously but still needing human verification, we're on an asymptotic path.
- malthaus 2y agoit's the equivalent of the "we overestimate the impact of technology in the short-term and underestimate the effect in the long run" quote. everyone is looking at llm scores & strawberry gotchas while ignoring the trillions of market potential in replacing existing systems and (yes) people with the current capabilities. identifying the use cases, finetuning the models and (most importantly) actually rolling this out in existing organizations/processes/systems will be the challenge long before the base models' capabilities will be it is worth working on those issues now and get the ball rolling, switching out your models for future more capable ones will be the easy part later on.
- _Algernon_ 2y agoI have yet to see LLMs provide a positive net value in the first place. They have a long way to go to weigh up for its negative uses in the form of polluting the commons that is the web, propaganda use, etc.
- corimaith 2y agoLooks you independently arrived at the original context that language models existed in as interfaces for deeper knowledge system in chatbots. But the knowledge system here is doing the grunt of the work, and progressing past it's own limitations goes right hack to the pitfalls of the rules based AI winter. That's not a engineering problem, it's a foundational mathematics problems that only a few people are seriously working on.
- raxxorraxor 2y agoThe context is a strict limitation if you work with data analysis or knowledge bases. Embeddings work, but the products we know get left and right mostly do not offer such capabilities at all. In that case most of these products remain decent chat bots. For coding LLMs certainly are helpful, but I prefer local models instead of anything on offer right now. There is just much more potential here.