5 ms·
Can you actually test a component that is, by design, a black box? ("by design" of the current models. Yes, the systems could be re-made traceable. But then it
by Piskvorrr 2y ago
Can you actually test a component that is, by design, a black box? ("by design" of the current models. Yes, the systems could be re-made traceable. But then it would also turn out on how much pirated material the LLM has trained, and hoo boy would the IP lawyers show up. That's the con: "if it's a blackbox, you can't prove whence the corpus, nyah nyah!")
- sensanaty 2y agoOur "testing" involves sending it 50 prompts and seeing if it bullshits too much and revising the prompt until it bullshits to acceptable levels. I suspect that's how the model makers test it as well, just fling crap at the wall and see which wall lets the most crap slide off of it (but never all of it)
- ksaj 2y agoThis is called Fuzzing. And it is a major component of software testing. Especially in infosec, but I digress. The reasoning behind it is exactly the same.
- ksaj 2y agoGood call. A case that is actively in court right now, is based on a point that you can sorta remake Johnny Be Good with Sonu. However, to do so, you have to already break copyright law by feeding it copyrighted lyrics. I did their experiment, but instead of "Go, Johnny, Go, Go Go!" I put in lyrics like "Run, Forrest, Run, Run, Run" (a Gump reference) and described the music style accordingly. It came out exactly like you'd expect. And the case is demonstrating a mirage where if you steal the song's actual lyrics, describe the style, you'll get something akin to what they might have written... except that there's no way to prove then that the model has even heard that original song before, because you gave it the lyrics and the style, and the rest is clearly bound to be similar as a result.
- emseetech 2y agoI brought this up in a previous HN thread, it's what I call "probabilistic UX". https://news.ycombinator.com/item?id=39954719 https://news.ycombinator.com/item?id=39954719 We don't have much experience building automated systems that are non-deterministic. Normally, in computer engineering, if we couldn't predict the results of an operation, we'd consider that operation buggy, broken or at best, flaky. Building consumer user interfaces for black box magic is a whole new ballgame.
- surfingdino 2y ago> We don't have much experience building automated systems that are non-deterministic Our brains evolved to turn chaos into patterns that help us survive. Determinism is good from that point of view and nondeterministic behaviour is discarded as useless. Nature in general operates in a deterministic way, a banana tree doesn't randomly produce kittens. AI and LLMs in particular are nothing but misinformation and IP appropriation systems. No wonder people reject them.
- ImHereToVote 2y agoI bet that if you included an expensive verification step by GPT 4 a lot of errors would go away.
- Piskvorrr 2y agoIt always circles back to "follow the money": - the system could theoretically trace the inputs, but then IP lawyers would eat company alive; pretending the corpus isn't pirated, so it can be cheap. - the system could verify and re-check, but that would require a massive compute increase; spouting bullshit, so it can be cheap - indeed a lot of answers around here go with "let's waste money by throwing it at the system until a human likes the result" - not cheap, but profitable for the provider
- Piskvorrr 2y ago"It sorta kinda works, most of the time, when it doesn't, Retry Reboot Reinstall^W^W, and if it's still broken, you're SOL - and thats by-design, we're not fixing the product." The word you're looking for is enshittification.