7 ms·
Absolutely correct. We already know this is about self-driving cars. Passing a driver's test was already possible in 2015 or so, but SDCs clearly aren't ready
by thwayunion 4y ago
Absolutely correct.
We already know this is about self-driving cars. Passing a driver's test was already possible in 2015 or so, but SDCs clearly aren't ready for L5 deployment even today.
There are also a lot of excellent examples of failure modes in object detection benchmarks.
Tests, such as driver's tests or standardized exams, are designed for humans. They make a lot of entirely implicit assumptions about failure modes and gaps in knowledge that are uniquely human. Automated systems work differently. They don't fail in the same way that humans fail, and therefore need different benchmarks.
Designing good benchmarks that probe GPT systems for common failure modes and weaknesses is actually quite difficult. Much more difficult than designing or training these systems, IME.
- zer00eyz 4y ago> good benchmarks ... failure modes and weaknesses is actually quite difficult. Much more difficult than designing or training these systems Is it? Based on the restrictions placed on the systems we see today and the way people are breaking it, I would say that some failure modes are known.
- thwayunion 4y agoA good benchmark is not simply a set of unit tests. What you want in a benchmark is a set of things you can use to measure general improvement; doing better should decrease the propensity of a particular failure mode. Doing this in a way that generalizes beyond specific sub-problems, or even specific inputs in the benchmark suite, is difficult. Building a benchmark suite that's large and comprehensive enough that generalization isn't necessary is also a challenge. Think about an analogy to software security. Exploiting a SQL injection attack in insecure code is easy. Coming up with a set of unit tests that ensures an entire black box software system is free of SQL injection attacks is quite a bit more difficult. Red teaming vs blue teaming, except the blue team doesn't get source code in this case. So the security guarantee has to come from unit tests alone, not systematic design decisions. Just like in software security, knowing that you've systematically eliminated a problem is much more difficult than finding one instance of the problem.
- brookst 4y agoI think the hard / unknown part is how you know you’ve identified all of the failure modes that need to be tested. Tests of humans have evolved over a long time and large sample size, and humans may be more similar to each other than LLMs are, so failure modes may be more universal. But very short history, small sample size, and diversity of architecture and training means we really don’t know how to test and measure LLMs. Yes, some failure modes are known, but how many are not?
- zer00eyz 4y ago>. Tests of humans have evolved over a long time and large sample size, and humans may be more similar to each other than LLMs are, so failure modes may be more universal. In reading this the idea that sociopaths and psychopaths pass as "normal" springs to mind. Is what an LLM doing any different than what these people do? https://medium.datadriveninvestor.com/the-best-worst-funniest-most-absurd-etc-chatgpt-responses-9094dda976fb https://medium.datadriveninvestor.com/the-best-worst-funnies... For people language is spoken before it is written... there is a lot of biology in the spoken word (visual and audio queue)... I think without these these sorts of models are going to hit a wall pretty quickly.
- brookst 4y ago> In reading this the idea that sociopaths and psychopaths pass as "normal" springs to mind. > Is what an LLM doing any different than what these people do? I think it's too big of a question to have any meaning. Which sociopaths? Which LLMs? For what differences? It's like asking "is a car any different from an airplane"? Yes, obviously in some ways. No, they are identical in other ways.
- dcolkitt 4y agoI'd also add that the almost all standardized tests are designed for introductory material across millions of people. That kind of information is likely to be highly represented in the training corpus. Whereas most jobs require highly specialized domain knowledge that's probably not well represented in the corpus, and probably too expansive to fit into the context window. Therefore standardized tests are probably "easy mode" for GPT, and we shouldn't over-generalize its performance there to its ability to actually add economic value in actually economically useful jobs. Fine-tuning is maybe a possibility, but its expensive and fragile, and I don't think its likely that every single job is going to get a fine-tuned version of GPT.
- Tostino 4y agoFrom what i've gathered, fine tuning should be used to train the model on a task, such as: "the user asks a question, please provide an answer or follow up with more questions for the user if there are unfamiliar concepts." Fine tuning should not be used to attempt to impart knowledge that didn't exist in the original training set, as it is just the wrong tool for the job. Knowledge graphs and vector similarity search seem like the way forward for building a corpus of information that we can search and include within the context window for the specific question a user is asking without changing the model at all. It can also allow keeping only relevant information within the context window when the user wants to change the immediate task/goal. Edit: You could think of it a little bit like the LLM as an analog to the CPU in a Von Neumann architecture and the external knowledge graph or vector database as RAM/Disk. You don't expect the CPU to be able to hold all the context necessary to complete every task your computer does; it just needs enough to store the complete context of the task it is working on right now.
- fud101 4y ago>From what i've gathered, fine tuning should be used to train the model on a task, such as: "the user asks a question, please provide an answer or follow up with more questions for the user if there are unfamiliar concepts." That isn't what finetuning usually means in this context. It usually means to retrain the model using the existing model as a base to start training.
- Robotbeat 4y agoI tend to think that it would not be particularly hard for current self driving systems to exceed the safety of a teenager right after passing the drivers test.
- jstummbillig 4y ago> Designing good benchmarks that probe GPT systems for common failure modes and weaknesses is actually quite difficult. Much more difficult than designing or training these systems, IME. What do you think is the difficulty?
- thwayunion 4y agoA good benchmark provides a strong quantitative or qualitative signal that a model has a specific capability, or does not have a specific flaw, within a given operating domain. Each part of this difficult -- identifying/characterizing the operating domain, figuring out how the empirically characterize a general abstract capability, figuring out how to empirically characterize a specific type of flaw, and characterizing the degree of confidence that a benchmark result gives within the domain. To say nothing of the actual work of building the benchmark.
- jstummbillig 4y agoSure – but how does this specificially concern GPT like systems? Why not test them for concrete qualifications in the way we test humans, using the tests we already designed to test concrete qualifications in humans?
- sebzim4500 4y agoThe difference is the impact of contaminated datasets. Exam boards tend to reuse questions, either verbatim or slightly modified. This is not such a problem for assessing humans, because it is easier for a human to learn the material than to learn 25 years of prior exams. Clearly that is not the case for current LLMs.
- thwayunion 4y agoAgain, because machines have different failure modes than humans.
- simiones 4y agoTo take a simplistic example, because a human who can provide a long motivated solution to a math problem that you re-use every three years likely understands the math behind it, while an LLM providing the same solution is likely just copying it from the training set and would be fully unable to resolve a similar problem that did not appear in the training set. Lots of exams are designed to prove certain knowledge given safe assumptions of the known limitations of humans, which are completely wrong for machines. The relative difficulty of rote memorization versus having an accurate domain model is perhaps the most obvious one, but there are others. Also, the opposite problem will often exist - if the exam is provided in the wrong format to the AI, we may underestimate its abilities (i.e. a very similar prompt may elicit a significantly better response).
- sebzim4500 4y agoYes, I think that we really don't have a good way of benchmarking these systems. For example, GPT-3.5-turbo apparently beats davinci on every benchmark that OpenAI has, yet anecdotally most people who try to use them both end up strongly preferring davinci despite the much higher cost. Presumably, this is what OpenAI is trying resolve with their 'Evals' project, but based on what I have seen so far it won't help much.
- kolbe 4y agoWe still struggle on benchmarking people.
- Waterluvian 4y agoOn topic of the driver's test analogy: I've known people who have passed the test and still said, "I'm don't yet feel ready to drive during rush hour or in downtown Toronto." And then at some point in the future they then recognize that they are ready and wade into trickier situations. I wonder how self-aware these systems can be? Could ChatGPT be expected to say things like, "I can pass a state bar exam but I'm not ready to be a lawyer because..."
- PaulDavisThe1st 4y agoYour comment has no doubt provided some future aid to a language model's ability to "say" precisely this.
- tsukikage 4y agoThe problem ChatGPT and the other language models currently in the zeitgeist are trying to solve is, "given this sequence of symbols, what is a symbol that is likely to come next, as rated by some random on fiverr.com?" Turns out that this is sufficient to autocomplete things like written tests. Such a system is also absolutely capable of coming up with sentences like "I can pass a state bar exam but I'm not ready to be a lawyer because..." - or, indeed, sentences with the opposite meaning. It would, however, be a mistake to draw any conclusions about the system's actual capabilities and/or modes of failure from the things its outputs mean to the human reader; much the same way that if you have dice with a bunch of words on and you roll "I", "am", "sentient" in that order, this event is not yet evidence for the dice's sentience.
- Waterluvian 4y agoI generally agree. But I remain cautiously skeptical that perhaps our brains are also little more than that. Maybe we have no capacity for that kind of introspection but we demonstrate what looks like it, just because of how sections of our brains light up in relationship to other sections.
- tsukikage 4y agoI don't believe that AI models can become introspective without such a capability either being explicitly designed in (difficult, since we don't really know how our own brains accomplish this feat and we don't have any other examples to crib) or being implicitly trained in (difficult, because the random person on fiverr.com rating a given output during training doesn't really know much of anything about the model's internal state and therefore cannot rate the output based on how introspective it actually is; moreover, extracting information about a model's actual internal state in some manner humans can understand is an active area of research, which is to say we don't really know how to do this, and so we couldn't provide enough feedback to train the ability to introspect even if we were trying to). I have no doubt that both these research areas can be improved on and that eventually either or both problems will be solved. However, the current generation of chatbots is not even trying for this.
- KKKKkkkk1 4y ago> We already know this is about self-driving cars. Passing a driver's test was already possible in 2015 or so, but SDCs clearly aren't ready for L5 deployment even today. Who told you that? Passing a driver's test was not possible in 2015 and it's not possible today. You might pass, but only if there are no awkward interactions with other drivers or bicyclists or pedestrians, no construction zones, and you don't enter areas where your map is out of date. The guy testing you would have to go out of his way to help you pass.
- thwayunion 4y ago>> We already know this is about self-driving cars. Passing a driver's test was already possible in 2015 or so, but SDCs clearly aren't ready for L5 deployment even today. > Who told you that? Passing a driver's test was not possible in 2015 and it's not possible today. You might pass, but only if there are no awkward interactions with other drivers or bicyclists and pedestrians, no construction zones, and you don't enter areas where your map is out of date. My, myself, and I. Driver's exams are de facto geo-fenced around the DMV where you choose to take the exam, and you get to choose from a few DMV locations, and you get to choose the time and day that you take the exam. Having spent some time working on self driving cars, I know that there existed at least one SDC platform in 2015 that was capable of passing the driving exam that I took when I got my driver's license (which involved leaving the parking lot, driving down a 4 lane road, turning into and driving around in a subdivision, taking another couple turns at well-marked intersections, pulling into the parking lot, and parallel parking). It's a low bar; mostly testing that you can follow four different types of road signs, navigate an unprotected left turn, and parallel park. I suppose following the officer's verbal instructions about where to go wasn't part of the SDC platform, but the actual driving part it would've been capable of passing.
- logifail 4y ago> Passing a driver's test was not possible in 2015 and it's not possible today My friend moved from Europe to the USA and took a driver's test in California (been driving in Europe since the 1980s). He tracked the test, he drove a whopping 2 miles (forwards) plus had to reverse about 30 feet. Commented to me afterwards that "signing the form was the hardest bit" and that "a blind person could probably pass it with the help of a guide dog". Passing a driving test isn't a proxy for anyone and anything being a good driver anywhere, but it's a good enough proxy for a human being a reasonable driver in the location where they take the test, which is what society has determined acceptible. Acceptible, for a human! I'm not sure it's useful for us to repeatedly attempting to measure AI's capabilities the same way we measure humans. Turing tests are all very well, but there are only so many fire hydrants I want to have to click on before I'm allowed to log into my hotel chain's loyalty scheme (Hilton, looking at you...)
- fatherzine 4y ago"SDCs clearly aren't ready for L5 deployment" Apologies for the tangent to the OP topic. The metric to watch is 'insurance damage per million miles driven'. At some point SDCs will overperform the human driver pool, possibly by a large margin. Wouldn't that be the point where SDCs are clearly ready for L5? Not even sure if that point is in the past or the future, does anyone -- not named Elon ;) -- have reasonably up-to-date trend charts and willing to share?
- TaylorAlexander 4y agoDamage per mile does not imply L5 readiness. My throttle only cruise control system in my car has never led to an accident, but only because I’m still there to operate the steering and to disable the cruise control at a moments notice. A self driving system that has been proven to be safe with humans diligently monitoring its behavior does not imply that this system can operate just as safely without the human.
- dekhn 4y agothat's exactly what's being tested by waymo in SF and Phoenix- there is no driver.
- TaylorAlexander 4y agoAh fair, but I believe L5 also means “all weather conditions” and probably “all reasonable roads”. No snow in either location and only certain kinds of roads. I wonder how they would handle a snowy single lane dirt road.
- hn_throwaway_99 4y agoGiven human nature, I still think society at large will reject self driving cars if they fail in ways a human never/rarely would, even if they are overall safer. That is, if a self driving car has, on average, fewer accidents than a human driver, but every 100 million miles or whatever it decides to randomly drive into a wall, I don't think people will accept them. Obviously this is a gray area (after all, humans sometimes decide to randomly drive into walls), but cars will need to be pretty far on "the right side of the gray" before they are accepted.
- rileymat2 4y ago> There are also a lot of excellent examples of failure modes in object detection benchmarks. I am curious if there are counter examples with better object detection. As a kid I used to see faces and to some extent still do in the dark. This is a really common thing that the human brain does. https://www.wired.com/story/why-humans-see-faces-everyday-objects/ https://www.wired.com/story/why-humans-see-faces-everyday-ob... https://en.wikipedia.org/wiki/Pareidolia https://en.wikipedia.org/wiki/Pareidolia Part of me wonder if in the face of novel environments that a sufficiently intelligent system needs to make these errors. But AI errors will always be different than human errors like you say.
- alexvoda 4y agoThe very big and dangerous difference is that while SDCs need approval in order to be allowed on the streets, there will be no quality control rules for reliance on LLMs. Corporate incentives to raise KPIs will mean that LLMs will be used and output verification will be superficial.
- SergeAx 4y ago> Passing a driver's test was already possible in 2015 or so I think we can talk about 2005. Check out the DARPA Grand Challenge, it was way harder: https://en.wikipedia.org/wiki/DARPA_Grand_Challenge_(2005) https://en.wikipedia.org/wiki/DARPA_Grand_Challenge_(2005)
- YeGoblynQueenne 4y ago>> Passing a driver's test was already possible in 2015 or so, but SDCs clearly aren't ready for L5 deployment even today. Wait, what are you saying? Passing a driver's test has been possible for much longer than since 2015 for a human. When did a self-driving car pass a driving test? In what jurisdiction? Under what conditions? Who gave it the test? What do you mean?