6 ms·
> Correctness will go from a binary pass/fail to a probability Excellent point. The pervasiveness of neural nets will require engineers (and really, everyone)
by danrocks 4y ago
> Correctness will go from a binary pass/fail to a probability
Excellent point. The pervasiveness of neural nets will require engineers (and really, everyone) to start thinking more probabilistically and establish acceptability thresholds instead of certainty. It's the way of the future.
- marginalia_nu 4y agoThat seems like a regression in many regards, especially given how much of this seems like a solution looking for a problem it can solve rather than the other way around.
- guhidalg 4y agoAgreed. We may imagine ourselves very smart but in my experience most people do not understand probability and statistics well. Is 90% accuracy good enough? Is 95%? 99%? 99.9%? No matter the answer, you have to tolerate errors. Now your stakeholders have to tolerate errors. Are they going to accept errors just because "Software 2.0" is here and that's what we all have to live with? Nope.
- staunton 4y agoAll software is riddled with errors and for most purposes that's fine. Any developer, stakeholder, whatever, who thinks otherwise is living in a parallel universe. It may be possible in the future to have "for all practical purposes flawless" software, which might make sense for select special applications. That would be a new thing though, rather than something we have and could lose due to adopting AI development.
- marginalia_nu 4y agoFormal verification of software has been around for a while though.
- staunton 4y agoWhich has almost never been used to build software anyone ever used... That might be changing very slowly
- rco8786 4y ago> All software is riddled with errors and for most purposes that's fine. This is typically not that true when it comes to correctness. Most software does the correct thing in the eyes of the user, nearly 100% of the time. And when it doesn't, the bug gets fixed and that edge case is corrected for every other user going forward. AI generated software, from what I've seen, has a wide range of errors in correctness (along with all the other errors that you mentioned all software having..which is true). Like it literally just does the wrong thing given what the user is expecting it to do. The path toward iteratively improving and getting it to an acceptable level of correctness for any given application might be there, but so far I have not seen it.
- staunton 4y agoThat's an interesting difference but I'm not convinced it's valid. Perhaps you can explain what you mean in more detail? Let's say the software is good enough if it does the right thing 99.9% of the time it is used. I take it, you're saying that if an AI starts modifying it and only writes correct code in 99.9% of cases (yes, current AI is not even close, but it will improve), that makes it worse because the software might start failing completely. However, if you have proper tests and release management, such obvious flaws will quickly be detected and fixed or rolled back. For most applications that seems pretty much equivalent to what we have now. The other case is software that is completely AI generated and cannot reasonably be modified by humans anymore. In that case, again, you have tests and a sane deployment strategy that mitigates failures to a sufficient degree depending on application. So the only issue is when you start making a completely AI generated software and fail to ever meet requirements, or pass the human written test cases? Users at least will never be impacted by that. Even now, many human-written software projects never get to the stage where they can be used. Is this really a problem, especially if the attempt at AI generation of software is cheap?
- rco8786 4y ago> yes, current AI is not even close This is really my main argument currently... > but it will improve With this being the "if" question. Improve, but improve to a point where we can trust it to do the things with the level of correctness actually required? Unclear so far. GPT3, the state of the art, can't be trusted to answer basic questions correctly (yet). I also don't personally buy into the other notions in these comments that "the future is probabilistic software". I think that's wishful thinking outside of some specific domains and an attempt to bend our actual requirements to meet the capabilities of AI software, rather than the opposite. > pass the human written test cases I'm not super sold on this idea either. It seems reasonably possible that writing the test cases to a level of specification necessary to ensure that correctness we're after is just as much effort as just writing the code. But, time will tell with everything.
- madeforhnyo 4y agoIt might be easier for an engineer to fix a bug by changing some lines of text than readjust a neural network, for the time being at least. As pointed out, there are already formal languages that allow formal verification like B [0] notably for like-critical systems. [0] https://en.wikipedia.org/wiki/B-Method https://en.wikipedia.org/wiki/B-Method
- hwayne 4y agoThere's a difference between errors in UI/reliability/performance/etc and errors in business logic. When there's an error in the business logic, then heads roll.
- duckmysick 4y ago> No matter the answer, you have to tolerate errors. We already do that in manufacturing. Physical parts are imperfect and we design with such variation in mind.
- badloginagain 4y agoThe entire field of Service Reliability is dedicated to finding the exact boundaries of acceptable errors, and defining the response function when those boundaries are crossed. I would assume this actually gels very nicely with neural nets since its constantly optimizing for fitness. Hell in theory you could bake in your SLA/SLIs into the models to self correct? Give the model direct feedback that its unfit?
- davedx 4y agoLess of a regression than you might think. How many software systems (outside of very well unit tested components) really have strong correctness proofs? It’s an almost futile task in modern software engineering with all those distributed systems everywhere
- xp84 4y agoI'd argue that it's a solution to the problem that programmers are expensive and only a fraction can be counted on to produce 100% reliable code anyway. While it's possible for a highly-skilled, highly-professional developer can both write code that will solve a given problem 100% correctly and write tests that will prove that it solves them for the entire domain, in practice most developers fall short on both counts. Every time you interact with a date or phone number field that chastises you for your use or non-use of punctuation, you know this. So, for many use cases, it's possible imperfect programmers will be replaced with neural networks that are 95% accurate, perhaps with a differently-trained one checking the work of the first one.
- __MatrixMan__ 4y agoI've been looking for an opportunity to try out a pattern where the two approaches improve each other: 1. Generate a 95% accurate model 2. Use it to generate test cases 3. Code the thing, with the help of the cases 4. Manually remove cases in the 5% I image steps 1 and 2 being completed by a product owner and 3 and 4 being completed by a software engineer. We're so horrifically bad at communicating a requirement's intent, I wonder what would happen if we tried to use AI to communicate them via their extent instead.
- lifeisstillgood 4y agoThat's probably genius :-) I mean ... can it be done? - build a platform (ie the data we care about and are going to build some workflow over) - have business describe what should happen in english - How does GPT build something that will run? Can it create the infrastructure? does it speak AWS? - then ... OK - I am actually excited by that
- __MatrixMan__ 4y agoI'm not familiar enough with AI workflows to weigh in on how easy or hard it would be. I've just noticed that it's much easier to describe why something is not what you want than describing what you want. So if AI can give us a mockup that's workable enough to skip the first few iterations of "no that's not what I want" and get right to the part where the engineer is asking questions about the edge cases that weren't explicit in the requirements... That's a win. I imagine you'd still want to have it in a box of some sort re: creating infra. Like you give it a very small cluster and probably make the stack decisions "write me a postgres schema for... write a fastapi API for the schema... write a react UI for the API... write me a k8s operator that up/down's the above components... Workshop the idea with other product people... ...and only then involve the engineer like: "make this AI-generated house of cards into a fortress".
- agileAlligator 4y agoIt seems to me that current AI techniques are more about reducing programmer workload than actually doing something that isn't possible with traditional methods.
- islon 4y agoOur software has a 99% chance to calculate you taxes correctly! And only a 1% chance of committing tax fraud.
- kklisura 4y agoNot quite, let me rephrase it... > Our software has a 99% chance to calculate your taxes correctly! And only a 1% chance of failure in which it's your fault and it's you that's committing tax fraud
- gtirloni 4y agoConsidering ChatGPT spits out innacurate information all the time, I think "our software has a 99% chance of you going to jail for tax fraud" is more accurate.
- bryanrasmussen 4y agothe problem really is that ChatGPT writes things based on what is most likely given what it wrote before and what the input is, so what are the possible error scenarios: 1. making a mistake in this part is really common, ChatGPT makes common mistake. 2. you have uncommon situation affecting here, ChatGPT ignores and writes things that cause you to get in trouble, or it writes things that cause you to pay more than you should. Also the longer is goes on writing things the more likely that things it writes does not hang together with the past things it wrote, when a human lies they try to make their lies at least follow a sensible pattern. ChatGPT would be likely to get you flagged for audits because you can't be sure that what it wrote on page 1 jibes with what it writes on page 3.
- alfor 4y agoBetter than most accountants. Family member got ripped by a government audit of what was supposed all fine by the person doing his company taxes.
- sebzim4500 4y agoI wonder what accuracy existing tax calculators have.
- fock 4y agothat doesn't look to bright if we look back towards Covid and vaccines. People are even bad at plain frequentist statistics!
- duderific 4y agoThat assumes people even care what the statistics are. Many people just go with their emotional response based on a political outlook when deciding what course of action to take.
- goatlover 4y agoThat doesn't sound like progress. Maybe for CRUD apps it will be okay. I have a hard time imagining it will be acceptable for financial or critical systems.
- ako 4y agoHow about self driving cars?
- bee_rider 4y agoYou’ll rent a self-driving car from a ride sharing app, they’ll make sure that the expected cost of fines and lawsuits will be significantly less than the expected revenue from providing rides. Although they can help push down the former value by making sure they operate from a favorable jurisdiction.
- rco8786 4y ago> It's the way of the future. For some use cases, sure. We go through painstaking efforts to ensure things like correctness, consistency, and idempotency for a reason though. Most things we want to be deterministic, and when something's not deterministic we freak out and fix it ASAP (including waking people up in the middle of the night to do so)
- fassssst 4y agoYea, aka how engineers working on safety critical or high reliability stuff already have to think. Assume your software has a probability to fail or have bugs or gets hit by bit flips or unreliable hardware. There’s a whole field for dealing with those kind of things that typical web devs haven’t had to worry about as much.
- belter 4y agoThese are not those type of Engineers. Not the ones doing triple computing using three different processors with a summation and voting process, inclusive using different programming languages and compilers, because the compiler can also have bugs. These are the engineers running self driving beta neural nets on public roads...
- williamcotton 4y agoI use a process called sample-and-vote when having LLMs return executable solutions. I imagine we will see this kind of design a lot with probabilistic computing.
- belter 4y agoSounds interesting. Any references?
- williamcotton 4y agoI show an approach here: https://github.com/williamcotton/empirical-philosophy/blob/main/articles/from-prompt-alchemy-to-prompt-engineering-an-introduction-to-analytic-agumentation.md https://github.com/williamcotton/empirical-philosophy/blob/m...
- flangola7 4y agoEngineer brains are just a type of neural network too, which also have a probability of putting out bad code. Space exploration budgets run into the hundreds of millions of dollars each and recruit the brightest people on the planet to design and program software, yet there is a long list of projects that failed due to code errors. Mars probes have failed to reach the planet, orbited too low and burned up, crashed into the surface, or landed successfully but later overwrote critical memory due to a flawed software update. The AI doesn't have to be perfect, but only offer a lower error rate than humans.
- belter 4y agoEngineer brains are not a neural net. It's only our brain lack of knowledge of how the brain works that led us to anthropomorphize Matrices and Weighted Graphs :-) Reading a bit on the Brain will quickly help dismiss those analogies. I suggest these two as good starting points: - Lange Clinical Neurology - 11th Edition - Bradley's Neurology in Clinical Practice, 8th Edition
- Xeoncross 4y ago"We're 98.7% certain the user's payment go to the right account with this change"
- krab 4y agoMore like we're 98.7 % certain that the screen layout looks ok for all device types and languages. The question is how hard is it to fix the problematic cases without turning the whole thing upside down when you find out the original solution is not enough.
- kabdib 4y ago"Good news, analysis of the anti-lock brake system failure shows that you're only sixteen percent dead."
- bryanrasmussen 4y agoFrom the neck up, then?