4 ms·
I see people saying things like this but I have yet to see anyone show data for a non-trivial workflow with human-level accuracy over a wide range of inputs, wi
by lukasb 4y ago
I see people saying things like this but I have yet to see anyone show data for a non-trivial workflow with human-level accuracy over a wide range of inputs, without a human in the loop.
- api 4y agoCounter argument: this may be a matter of incremental improvement. The breakthroughs may all be behind us. It’s like saying you haven’t yet seen a 1000 mile range EV for under $100k. No you can’t buy such a thing now but it’s clearly possible and we know how to get there by just continuing to grind on battery technology and scale manufacturing. AGI may be at the place a moon landing was in 1950, not where it was in 1900 or 1850.
- charcircuit 4y ago>by just continuing to grind on battery technology and scale manufacturing. I think it would be easier to include an ICE and enough fuel to get you to that 1000 miles mark.
- ThunderSizzle 4y agoIt'd be cheaper to then get rid of the battery. Then 10k can be your new price limit. Even 1k can get a good enough junk car that can go that far.
- HopenHeyHi 4y agoYou can actually buy a 1000 km range EV for $160k now (MB EQXX). Just as a by the way. :) At this price point it actually has nothing to do with grinding on battery tech and scale manufacturing, the limiting factor is physics. You can only make it so aerodynamic before you hit diminishing returns or it stops looking like a car. You can only make it so lightweight. And so forth. This is vaguely as good as it can get and we can say that because we understand how it all works. LLMs on the other hand invite all kinds of magical thinking around unlimited potential because we poked them with a stick and something interesting comes out it must mean that if we poke it just right we will get an AGI. That just doesn't logically follow from what we know of it so far.
- api 4y agoI am not convinced we have cracked AGI. I just would no longer make a large bet that we have not. We won’t know until an AGI actually starts to act like one. In other words we won’t know until we know and then we are suddenly there. That doesn’t mean I’m on the doomwagon. I feel kind of weird and contrarian but I am just not that afraid of AGI. For the foreseeable future AGI should be much more afraid of us. Imagine having us for gods. (I actually am a bit concerned that we will accidentally put a sentient mind in hell without knowing what we are doing. Would it know how to tell us? Would we care?) As far as human survival I’m afraid of whatever it is that is going to get us that nobody including myself is thinking about. That’s not AGI. That’s the alien weapon for which Oumuamua was a spent deceleration stage. (To make up something random. It probably isn’t that.) I disagree about physical limits with EVs. We are not near the physical limits of battery energy density. From what I have read a 2000 mile EV may be possible, albeit quite far out. But it was just a random contemporary example.
- nemo44x 4y agoFwiw, a majority at OpenAI believes GPT5 will achieve AGI, depending on how you define it, according to Sam Altman.
- HopenHeyHi 4y agoIt is hard to falsify as they are not very open but I believe the keyword here is believe. It is faith/intuition based. Certainly fertile ground for exploration but people who argue back and forth about it remind of the "are we alone in the universe" conversations.
- deleted 4y ago[deleted]
- Avicebron 4y agoAnd I'm sure him saying that has nothing to do with marketing
- 4y ago
- heyitsguay 4y agoLLMs are having a moment in 2023 like self driving cars were having in 2015. Some really cool demos following a lot of hard work, too much hyperbolic speculation that mass real-world job-destroying deployments are right around the corner, not enough appreciation of how few commercial applications are ok with 99% (or even 99.9%) accurate solutions. Real value being created, but still requiring lots of human ingenuity and grunt work to unlock it. As a decent first-order metric - follow the ratio of companies getting money for using LLMs to do something, to companies getting money for providing LLMs and associated tooling to others. The bigger that ratio gets, the more real world impact LLMs are having.
- famouswaffles 4y agoalmost nothing involving NLP requires solutions anywhere near that accuracy rate. I've seen the self driving comparisons a lot but they straight up make little sense. there's a reason microsoft's various copilot suites have already popped up (365, X, Bing). massive value to be gained already in the here and now.
- skepticATX 4y agoTo play devil's advocate, we still have no idea how economically impactful the Copilot suites will be. I don't think this is the most likely outcome, but I can absolutely see a scenario in which these end up being minor features that are rarely used by the typical worker.
- heyitsguay 4y agoI've messed with ChatGPT before but this weekend I tried to use it seriously to hack together a python demo in the computer vision space. I appreciate it and it's a way better rubber duck for me to talk my ideas through with and generate sample snippets, but hallucination is a problem, and the more intricate and customized the code being developed, the more prone it's been to misinterpretation. I'm getting some good code starting points, and talking the idea through step by step with a chatbot is really helping me clarify what i need to do, but I've still got API docs open for the libraries it uses because it likes to make up functions, including the core one on which the project logic hinges. (But that's ok, because now I know I just need to write a function that works that way and i can do that). Pretty cool and helpful! And I can imagine it getting better with GPT-4 or code-specific tooling. But it's generating value on the order of like.. many other SaaS offerings that have come onto the scene that try to ease pain points in coding workflows. Versus value of the sort that upends society and my entire way of life. A great new tool that I should learn about to make rote bits of my activities faster and easier, a story that's a bit more familiar in tech than some of the more breathless stories about AI make it sound.
- famouswaffles 4y agoHow is reinforcement learning without a single human in the loop not non-trivial?
- lukasb 4y agoWhat is the accuracy of the resulting model? Over what range of inputs?
- famouswaffles 4y agohttps://crfm.stanford.edu/helm/latest/?group=core_scenarios https://crfm.stanford.edu/helm/latest/?group=core_scenarios Anthropic-LM v4-s3 (52B) is the model in question. rlaif doesn't seem to be any less effective given the size of the model.
- lukasb 4y agoSo base model+RLLLMF performs as well as base model+RLHF. That could mean a lot of things - it could mean the base model puts a ceiling on the total possible accuracy, so having human-level performance at the RL step doesn't matter as much. And looking at the scores on the individual tests that make up the composite accuracy metric, that looks probable.