18 ms·
Uncertain<T>
- mackross 1y agoAlways enjoy mattt’s work. Looks like a great library.
- boscillator 1y agoDoes this handle covariance between different variables? For example, the location of the object your measuring your distance to presumably also has some error in it's position, which may be correlated with your position (if, for example, if it comes from another GPS operating at a similar time). Certainly a univarient model in the type system could be useful, but it would be extra powerful (and more correct) if it could handle covariance.
- layer8 1y agoTo properly model quantum mechanics, you’d have to associate a complex-valued wave function with any set of entangled variables you might have.
- evanb 1y agoIf you need to track covariance you might want to play with gvar https://gvar.readthedocs.io/en/latest/ https://gvar.readthedocs.io/en/latest/ in python.
- joerick 1y agoI've been wondering for a while if a program could "learn" covariance somehow. Through real-world usage. Otherwise, it feels to me that it'd be consistently wrong to model the variables as independent. And any program of notable size is gonna be far too big to consider correlations between all the variables. As for how one might do the learning, I don't know yet!
- karelpeeters 1y agoUsing this sampling-based approach you get correct covariance modeling for free. You have to only sample leaf values that are used in multiple places once per evaluation, but it looks like they do just that: https://github.com/mattt/Uncertain/blob/962d4cc802a2b179685d33919cb02588218d063e/Sources/Uncertain/Uncertain.swift#L1509-L1521 https://github.com/mattt/Uncertain/blob/962d4cc802a2b179685d...
- jakubmazanec 1y ago[flagged]
- cobbal 1y agoI don't think inference is part of this at all, frequentist or otherwise. It's not part of the type system, it's just the giry monad as a library.
- frizlab 1y ago> And why does it need to be part of the type system? As presented in the article, it is indeed just a library.
- geocar 1y ago> What if I want Bayesian? Bayes is mentioned on page 46. > And why does it need to be part of the type system? It could be just a library. It is a library that defines a type. It is not a new type system, or an extension to any particularly complicated type system. > Am I missing something? Did you read it? https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/uncertaint-asplos-2014-slides.pdf https://www.microsoft.com/en-us/research/wp-content/uploads/... https://github.com/klipto/Uncertainty/ https://github.com/klipto/Uncertainty/
- jakubmazanec 1y ago> Bayes is mentioned on page 46. Bayes isn't mentioned in the linked article. But thanks for the links.
- geocar 1y agoThat did not surprise me because I did not think the article was about anything but about adapting the dot-net library they linked to on Microsoft's site to swift, and I figured that if I wanted to understand the library and the approach I had better read the links that indicated I might be able to learn from them.
- muxl 1y ago
- AlotOfReading 1y agoA small note, but GPS is only well-approximated by a circular uncertainty in specific conditions, usually open sky and long-time fixes. The full uncertainty model is much more complicated, hence the profusion of ways to measure error. This becomes important in many of the same situations that would lead you to stop treating the fix as a point location in the first place. To give a concrete example, autonomous vehicles will encounter situations where localization uncertainty is dominated by non-circular multipath effects. If you go down this road far enough you eventually end up reinventing particle filters and similar.
- mikepurvis 1y agoVehicle GPS is usually augmented by a lot of additional sensors and assumptions, notably the speedometer, compass, and knowledge the you'll be on one of the roads marked on its map. Not to mention a fast fix because you can assume you haven't changed position since you last powered on.
- monocasa 1y agoAs well as a fast fix because you know what mobile cell or wifi network you're on.
- astrange 1y agoNone of the inputs you mention work against multipath effects in cities, which means car GPS won't know which lane you're in and in a grid system may think you're on the next street over. If you have an HD map you can solve for it using building shapes or by looking at the street with cameras. WiFi seems like it would help, but the locations of the WiFi terminals are themselves based on crowdsourced GPS.
- o11c 1y ago> Not to mention a fast fix because you can assume you haven't changed position since you last powered on. ... until you use a ferry.
- jeffreygoesto 1y agoWell. Some part of the 101 was moved a bunch of feet sideways after construction. Really hard to correct for, the GPS and the map localization were constantly fighting like an old couple... Had to re-map that stretch quickly...
- layer8 1y agoArguably Uncertain should be the default, and you should have to annotate a type as certain T when you are really certain. ;)
- esafak 1y agoA complement to Optional.
- nine_k 1y agoOnly for physical measurements. For things like money, you should be pretty certain, often down to exact fractional cents. It appears that a similar approach is implemented in some modern Fortran libraries.
- XorNot 1y agoMoney has the problem that no matter how clever you are someone will punch all the values into Excel and then complain they don't match. Or specify they're paying X per day, but want hourly itemized billing...but it should definitely come out to X per day (this was one employer which meant I invoiced them with like 8 digits of precision due to how it divided, and they refused to accept a line item for mathematical uncertainty aggregates).
- rictic 1y agoA person might have mistyped a price, a barcode may have been misread, the unit prices might be correct but the quantity could be mistaken. Modeling uncertainty well isn't just about measurement error from sensors. I wonder what it'd look like to propagate this kind of uncertainty around. You might want to check the user's input against a representative distribution to see if it's unusual and, depending on the cost of an error vs the friction of asking, double-check the input.
- bee_rider 1y agoTypos seem like a different type of error from physical tolerances, and one that would be really hard to reason about mathematically.
- munchler 1y agoIs this essentially a programmatic version of fuzzy logic? https://en.wikipedia.org/wiki/Fuzzy_logic https://en.wikipedia.org/wiki/Fuzzy_logic
- esafak 1y agohttps://en.wikipedia.org/wiki/Probabilistic_programming https://en.wikipedia.org/wiki/Probabilistic_programming more like. It is already a thing; see, for example, https://pyro.ai/ https://pyro.ai/
- krukah 1y agoMonads are really undefeated. This particular application feels to me akin to wavefunction evolution? Density matrices as probability monads over Hilbert space, with unitary evolution as bind, measurement/collapse as pure/return. I guess everything just seems to rhyme under a category theory lens.
- valcron1000 1y agoRelevant (2006): https://web.engr.oregonstate.edu/~erwig/pfp/ https://web.engr.oregonstate.edu/~erwig/pfp/
- 8note 1y agofor mechanical engineering drawings to communicate with machinists and the like, we use tolerances eg. 10cm +8mm/-3mm for what the acceptable range is, both bigger and smaller. id expect something like "are we there yet" referencing GPS should understand the direction of the error and what directions of uncertainty are better or worse
- mabster 1y agoSomething that's bugged me about this notation though is that sometimes it means "cannot exceed the bounds" and sometimes it means "only exceeds the bounds 10% of the time"
- taneq 1y agoI don’t think I’ve ever seen mechanical drawings have “90% confidence” dimensions like this. If a part’s too big then it won’t fit, and it’s probably useless.
- kevin_thibedeau 1y agoIf a test procedure is verifying all dimensional accuracy, it can be assumed to be bounding tolerance. If it's a mass production line with less than 100% testing of parts, you'd have to expect that some outliers get by and the tolerance is something like 3-sigma on a Gaussian.
- mabster 1y agoYeah it's probably field specific and I guess Gaussian-based uncertainty would be more about statistical sampling rather than tolerances. I've noticed that if arithmetic is being done on it it's almost certainly Gaussian. I just mean whenever I see uncertainty like this, I don't know what is meant!
- brabel 1y agoIn Mechanical Engineering, tolerances ensure that when you put parts together, they will fit as long as the tolerances were respected. It's not statistical. If the machinist makes a part that's not within the +/- bounds, they throw it away and start again. If you tried to fit multiple parts, all with only statistical respect for tolerances, you would run into trouble almost 100% of the time with just a few pieces.
- cb321 1y agoIf you are in an even more "approximate" mindset (as opposed to propagating by simulation to get real world re-sampled skewed distributions, as often happens in experimental physics labs, or at least their undergraduate courses), there is an error propagation (https://en.wikipedia.org/wiki/Propagation_of_uncertainty https://en.wikipedia.org/wiki/Propagation_of_uncertainty) simplification for "small" errors thing you can do. Then translating "root" errors to "downstream errors" is just simple chain rule calculus stuff. (There is a Nim library for that at https://github.com/SciNim/Measuremancer https://github.com/SciNim/Measuremancer that I use at least every week or two - whenever I'm timing anything.) It usually takes some "finesse" to get your data / measurements into territory where the errors are even small in the first place. So, I think it is probably better to do things like this Uncertain<T> for the kinds of long/fat/heavy tailed and oddly shaped distributions that occur in real world data { IF the expense doesn't get in your way some other way, that is, as per Senior Engineer in the article }.
- black_knight 1y agoThis seems closely related to this classic Functional Pearl: https://web.engr.oregonstate.edu/~erwig/papers/PFP_JFP06.pdf https://web.engr.oregonstate.edu/~erwig/papers/PFP_JFP06.pdf It’s so cool! I always start my introductory course on Haskell with a demo of the Monty Hall problem with the probability monad and using rationals to get the exact probability of winning using the two strategies as a fraction.
- internet_points 1y agoSee also the Haskell library monad-bayes https://monad-bayes.netlify.app/tutorials/ https://monad-bayes.netlify.app/tutorials/ https://www.tweag.io/blog/2019-09-20-monad-bayes-1/ https://www.tweag.io/blog/2019-09-20-monad-bayes-1/
- droideqa 1y agoCould this be implemented in Rust or Clojure? Does Anglican kind of do this?
- j2kun 1y agoThis concept has been done many times in the past, under the name "interval arithmetic." Boost has it [1] as does flint [2] What is really curious is why, after being reinvented so many times, it is not more mainstream. I would love to talk to people who have tried using it in production and then decided it was a bad idea (if they exist). [1]: https://www.boost.org/doc/libs/1_89_0/libs/numeric/interval/doc/interval.htm https://www.boost.org/doc/libs/1_89_0/libs/numeric/interval/... [2]: https://arblib.org/ https://arblib.org/
- Tarean 1y agoInterval arithmetic is only a constant factor slower but may simplify at every step. For every operation over numbers there is a unique most precise equivalent op over intervals, because there's a Galois connection. But just because there is a most precise way to represent a set of numbers as an interval doesn't mean the representation is precise. A computation graph which gets sampled like here is much slower but can be accurate. You don't need an abstract domain which loses precision at every step.
- bee_rider 1y agoIt would have been sort of interesting if we’d gone down the road of often using interval arithmetic. Constant factor slower, but also the operations are independent. So if it was the conventional way of handling non-integer numbers, I guess we’d have hardware acceleration by now to do it in parallel “for free.”
- eru 1y agoYou can probably get the parallelism for interval arithmetic today? Though it would probably require a bit of effort and not be completely free. On the CPU you probably get implicit parallel execution with pipelines and re-ordering etc, and on the GPU you can set up something similar.
- pklausler 1y agoInterval arithmetic makes good intuitive sense when the endpoints of the intervals can be represented exactly. Figuring out how to do that, however, is not obvious.
- nicois 1y agoIs there a risk that this will underemphasise some values when the source of error is not independent? For example, the ROI on financial instruments may be inversely correlated to the risk of losing your job. If you associate errors with each, then combine them in a way which loses this relationship, there will be problems.
- deleted 1y ago[deleted]
- tricky_theclown 1y agoS
- lloydatkinson 1y agoIS there the complete C# available for this? I looked over the original paper and it's just snippets.
- kittoes 1y agohttps://github.com/klipto/uncertainty https://github.com/klipto/uncertainty
- Pxtl 1y ago10 years since commit and no attached documents besides a tiny readme. Pass.
- miffy900 1y agoThis is still some code, as opposed to no code. It does seem to model everything in the research paper. Aside from the original research paper needing to be included in the repo, it definitely does not need anything more than what's already there. It all builds and compiles without errors, only 2 warnings for the library proper and 6 warnings for the test project. Oh and it comes with a unit testing project: 59 tests written that covers about 73% of the library code. Only 2 tests failed. Even having a unit testing library means it beats out like 50% of all repos you see on GitHub.
- kittoes 1y agoBlame Microsoft Research, as the link came directly from them: https://www.microsoft.com/en-us/research/project/uncertainty/ https://www.microsoft.com/en-us/research/project/uncertainty.... I don't think they ever really took the project past the initial paper/presentation.
- naasking 1y agoSometimes things can just be "done", and the paper is pretty good documentation if the implementation is faithful to what is described there.
- contravariant 1y agoI feel like if you're worried about picking the right abstraction then this is almost certainly the wrong one.
- lxe 1y agoI really like that this leans on computing probabilities instead of forcing everything into closed-form math or classical probability exercises. I’ve always found it way more intuitive to simulate, sample, and work directly with distributions. With a computer, it feels much more natural to uh... compute: you just run the process, look at the results, and reason from there.
- keeganpoppen 1y agooh man i had forgotten about this blog from when i orbited the swift ecosystem a bit... it's clearly as great as always! fun post!
- dcsommer 1y agoSeems more proper to call it a `ProbabilityDistribution` type. It's a more general and intuitive way to handle the concept.
- btown 1y agoOnce one understands that a variable (in a programming context) can hold a specification for a variable (in a mathematical context), one opens up incredible doors that are at the foundation of modern AI. When you see y = m * x + b, your recollections of math class may note that you can easily solve for "m" or find a regression for "m" and "b" given various data points. But from a programming perspective, if these are all literal values, all this is is a "render" function. How can you reverse an arbitrary render function? There are various approaches, depending on how Bayesian you want to be, but they boil down to: if your language supports redefining operators based on the types of the variables, and you have your variables contain a full specification of the subgraphs of computations that lead to them... you can create systems that can simultaneously do "forward passes" by rendering the relationships, and "backward passes" where the system can automatically calculate a gradient/derivative and thus allow a training system to "nudge" the likeliest values of variables in the right direction. By sampling these outputs, in a mathematically sound way, you get the weights that form a model. Every layer in a deep neural network is specified in this way. Because of the composability of these operations, systems like PyTorch can compile incredibly optimal instructions for any combination of layers you can think of, just by specifying the forward-pass relationships. So Uncertain<T> is just the tip of the iceberg. I'd recommend that everyone experiment with the idea that a numeric variable might be defined by metadata about its potential values at any given time, and that you can manipulate that metadata as easily as adding `a + b` in your favorite programming language.
- jonahx 1y agoVery interesting. Are there PLs that support this kind of thing at the language level as you are describing?
- btown 1y agohttps://colcarroll.github.io/ppl-api/ https://colcarroll.github.io/ppl-api/ is likely a good starting point to get a taste of examples in Python; some use custom languages, but the success of Python-native frameworks in the LLM world I think has shown that embracing that makes interop and composability more possible at scale. https://news.ycombinator.com/item?id=28941145 https://news.ycombinator.com/item?id=28941145 has some discussion here as well, though it’s a few years old. Pyro and NumPyro seem to be popular at the moment!
- webcoon 1y agoAwesome! This speaks to something, which I've been thinking (and wishing) for a long time. I've already done probabilistic programming in a scientific context (Python) and classical software engineering for web development (TypeScript, Python, Rust), and I've always wondered why I couldn't have the real-world modelling capacity of the former with the static type assurances of the latter. Love that you (and Microsoft) are thinking along the same lines! Do you perhaps know of any Python implementations for this? There are plenty of dynamic stats programming libraries, but none offer typing solutions AFAIK.
- akst 1y agoSomething I've wanted to make was a data type to represent a value that may or may not be known with a level of certainty over a certain distribution (or probability density function), but you could apply various transforms that may or may not have their own level of uncertainty, and you end up with a refined set of probability distributions each observation (or a new set of classifications based on whatever conditionals). With the eventual goal of running various simulations over different randomly generated outcomes based on those probability distributions.
- thekoma 1y agoWe designed a processor microarchitecture [1] at the University of Cambridge, inspired by Uncertain<T> (James Bornholt) and related work. In addition to assuming parametric distributions (e.g., Gaussian, Rayleigh), it lets you load arbitrary sets of samples into registers/memory so program values are carried and propagated as nonparametric distributions through ordinary arithmetic. A spin-off, Signaloid, is taking this technology to market. I'm also researching using this in state estimation (e.g., particle filters). [1]: https://dl.acm.org/doi/10.1145/3466752.3480131 https://dl.acm.org/doi/10.1145/3466752.3480131
- deleted 1y ago[deleted]
- naasking 1y agoReally interesting to see how long ideas take to go mainstream. From my recollection, Oleg and Chung-chieh Shan did this first back in 2009 as a library in OCaml [1,2]. [1] https://groups.google.com/g/fa.caml/c/CbXeoR_Rzrk?pli=1 https://groups.google.com/g/fa.caml/c/CbXeoR_Rzrk?pli=1 [2] https://okmij.org/ftp/kakuritu/ https://okmij.org/ftp/kakuritu/
- captainmuon 1y agoBack when I was studying physics, we frequently had to do calculations with error propagation. I tried to implement something very similar in C++ and in Python, but never finished it. I also thought it would be neat if a spreadsheet program could understand uncertainties, and also units, so you could enter 1m +- 10cm and it would propagate the errors correctly. If you laid out the data with one column for the values and one for the errors, I had a couple of OpenOffice macros that would perform the calculations. Another place where I think this would be neat would be in CAD. Imagine if you are trying to create a model of an existing workpiece or of a room, and your measurements don't exactly add up. It's really frustrating and you have to go back and measure again, and you usually end up idealizing the model and putting in rounder numbers to make it fit, but it is less true to reality. It would be cool if you could put in uncertainties for all lengths and angles, and it would run a solver to minimize the total error.