5 ms·
Those massive gains have yet to considered reliable enough to be considered trusthworthy. Would you consider them trusthworthy in court, where lives are at stak
by turingspiritfly 8y ago
Those massive gains have yet to considered reliable enough to be considered trusthworthy. Would you consider them trusthworthy in court, where lives are at stake? Gains are nice but we are still so far from the essence of AI systems and considering how much resources we are pouring into learning, at this point all of them appear as nothing more than massive fat expensive toys
- laretluval 8y agoHumans are often not considered trustworthy and reliable in court.
- EpicEng 8y agoWhat does that have to do with engineering?
- laretluval 8y agoTrustworthiness in court was advanced as a metric for AI success at which current methods are failing. > Would you consider them trusthworthy in court, where lives are at stake?
- solotronics 8y agoI have the perspective of an informed layman as a programmer who hasn't messed with ML yet. Wouldn't the "safest" solution be a system with multiple algorithms and a consensus mechanism?
- bumby 8y agoI believe some models do precisely that. Random forest ML as an example tallies "votes" on the outcome. I'm not sure how robustly multiple algorithms have been applied to this voting technique, but it would be an interesting read if anyone has information on it.
- notafraudster 8y ago"Vote tallying" is basically taking a mode/weighted mode response rather than a mean/median/trimmed mean/weighted mean response. There are contexts where this is ideal (for example, classification in multiple unordered classes where mode is the only measure of central tendency that is even reasonable); cases where it's superfluous (in 2-class classification the mode is the median is the sign of the mean); and cases where it's bad (in regression with a continuous outcome where the modal prediction has probability exactly 0). So really it depends on the space where you want to use it. Typically in a regression setting with an ensemble learner you're using a kind of weighted mean, where the weights are selected based on cross-validation performance. This is sometimes called a "super learner". See van der Laan, Polley, and Hubbard 2007. Note that this suffers from being similarly awful in terms of theory as a lot of other ML stuff. It is not, for example, the case that two apparently similar datasets or problem domains will produce similar super-learner weights. Which is disturbing, because it's easy to believe that say SVM does better at X and penalized ordered logit at Y, but it's hard to believe that they both do better seemingly at random.
- chartpath 8y agoYeah. I'm coming from the same background, but with some experience orchestrating these systems in production. AFAIK a bunch of the submissions to the various "AGI" awards like Alexa prize use ensembles of models and some way of weighing each one based on a context in order to choose which classifier to trust in a particular scenario. E.g. MILABOT. There is so much more than one monolithic NN since those are easy to saturate in terms of precision and recall with enough training data and features, but are not enough to provide good UX in any complex domain/ontology. So it makes more sense to have many different models trained on each subdomain/taxonomy so that each can be specialized and then combined orthogonally. Then the question becomes "how do we orchestrate them?" Well, there is a lot of research from the 80s and 90s that kinda got left by the wayside due to hype cycles (see the last "AI winter"). My faves are Collagen and Ravenclaw. And there is a lot of literature around topic frame stack modelling, which can be combined with various expert systems or other logics. I am currently using CLIPS (PyKnow) with custom Ravenclaw implementation. I believe b4.ai is doing something similar without the logic/rules engine, and actually applying ML to topic selection as well. My systems are goal oriented so I like to give them a teleology for business reasons, which would not suffice for AGI ambitions. TLDR data science isn't enough on its own. We need engineers to architect things properly to solve problems.
- module0000 8y ago>> Those massive gains have yet to considered reliable enough to be considered trusthworthy. We're using them at Generic Health Insurance Megacorp in production - lots of enterprises are. If you are in the IT industry, it might be useful to spend some lab time with ML. Possibly you have a misconception of ML and/or confuse it with AI.
- crankylinuxuser 8y agoSo, long story short, when groups were talking about governmental death panels, they in actuality were black box AIs that we have no understanding of, yet they make the core decisions and recommendations? Indeed...
- ineedasername 8y agoI think the phrase "governmental death panels" is a rather loaded politically partisan phrase. Any system of insuring folks includes people who decide what treatments will be approved and what won't. And in the realm of universal health care systems where decisions like that must be made (as they're made an any insurance company) Something like the UK's quality-adjusted life year (QALY) is, if imperfect, not a wholly unreasonable concept that doesn't include black box AI.
- module0000 8y agoGeezus, not that type of health insurance. We're an insurance provider over a century old, blue cross and blue something or other. We're not in the business of death panels, and we have similar feelings towards the tin-foil-hat wearing masses that Catholic priests likely have towards people who assume since some priests are pedophiles, they must all be pedophiles. Some insurance companies(in the minority) are driven by greed and have a reputation that reflects that. Not all priests are pedos, and not all(or any that I've heard of) insurance providers are murderous thugs trying to extract every last cent of wealth from their customers. We're a non-profit organization for that matter. If we somehow extracted a ton of wealth, we wouldn't be allowed to keep it anyway. We have some strict(and irritating for profit-centric people) governance... one of the more interesting pieces of governance is called the "85/15 rule", which roughly translated, means that if we take in $100 dollars, the government mandates we use $85 of them to pay your costs, and $15 to pay our staff, light bills, and any other expense we have. If we end up only using $80 to pay your costs, we have to refund the remaining $5 to your group plan. Here's the obvious secret about health insurance that people like to have conspiracy theories about..I can't speak for other institutions in other countries, however...our stance is really simplistic: you can't pay premiums if you are not alive, therefore it is in our mutual interest for you to remain alive. All the conspiracy theories such as "but you don't want that cancer patient in your insurance group plan!" are just that..conspiracy theories. We absolutely do want that person in the group, because then that group's rates go up! The costs for that patient's care are more or less fixed(and known), built on the assumption of a terminal outcome. We're going to pay for it anyway, and try to make that miserable experience as pleasant as possible for everyone involved. That type of service is how you get repeat business, and a good reputation. This likely falls on deaf ears. Feel free to return to the zealous insurance hatred, and I'm going to return to writing code. Not for death panel machines. Promise.
- another-one-off 8y ago> Would you consider them trusthworthy in court, where lives are at stake? Probably. Human intelligence is extremely fallible - based on the statistics the only reason we trust humans to do half the stuff they do is because there is literally no choice. If we held humans to a high objective engineering standard We wouldn't: * Let them drive * Let them present their memories as evidence in a court case * Entrust them with monitoring jobs * Allow them to perform surgical operations Humans are the best we have at those things, but from a "did we secure the best result with the information we had" perspective they are not very reliable. A testable and consistently performing AI with known failure modes might even be able to outperform a human with a higher failure rate (eg, we can reconfigure our road systems if there is just one scenario an AI driver can't handle). Basically, you might be dead on the money that they are not 'trustworthy enough', but lets not lose sight of the fact that even being an order of magnitude from human performance might be enough after costs and engineering benefits get factored in. The weakest link is the stupidest human, and that is quite a low bar.
- perfmode 8y agohttps://www.usenix.org/conference/usenixsecurity18/presentation/mickens https://www.usenix.org/conference/usenixsecurity18/presentat...
- denzil_correa 8y ago> Basically, you might be dead on the money that they are not 'trustworthy enough', but lets not lose sight of the fact that even being an order of magnitude from human performance might be enough after costs and engineering benefits get factored in. Ironically, the thing that is lost in this comment would be "accountability". In case of a human, you can go back / trace decision making criteria and hold someone accountable. In case of an algorithm, everyone washes their hands off. Performance is not the only criteria to make a decision if algorithms are "trustworthy" over humans.
- RA_Fisher 8y agoLinear models are highly interpretable and an operator can be held accountable.