26 ms·
1. Faster? Really? Those values only live in registers for a short while, is saving one or two cycles all that important? 2. That is an interesting point. Gut
by anonymous 13y ago
1. Faster? Really? Those values only live in registers for a short while, is saving one or two cycles all that important?
2. That is an interesting point. Gut instinct says tmp should end up being f64 and then both expressions would be the same. Equally valid is the interpretation that the user selected explicitly for conversion to f32 of the intermediate value. We all know computers don't perform platonically perfect math with real numbers, there isn't much you can do about it, but I think the choice which results in the most precision should win at the end of the day.
Perhaps in a language which doesn't require specifying types like Rust, the best thing to do is to simply dispense with 32-bit floating point values entirely. I'd be equally happy with compiler flags like -fmore-precision. Or maybe you could have other arithmetic operators like +64 and +80 which cast all their arguments to 64/80-bit floating point numbers and produce a result with that precision. So I'd write
let a: f32 = b (+80) c (/80) d as f32
Given that Rust is a pretty young language and still far from a 1.0 release, do consider if you can provide any kind of syntax to actually make full use of computers' numerical capabilities.
- pcwalton 13y agoRegarding point (1), it starts to matter when autovectorization kicks in. If this expression is in a loop and the compiler can autovectorize (which rustc does, if you have a new enough version), you really want to be able to pack more values into a SIMD register if you can. (This is one point in which this paper is out of date…) As for point (2), well, I think that making the size of "tmp" f64 would interact badly with other features like operator overloading and generics, since you're playing fast and loose with types. Gory details (warning, type theory ahead): "+" is implemented by a trait, Add(RHS,Result) where RHS and Result are concrete types and there is a functional dependency so that Self + RHS → Result. This fundep is necessary because otherwise "let tmp: b + c" wouldn't work, as the compiler doesn't know what the type of "tmp" is (is it f32 or f64? It won't guess.) So you can't simultaneously implement f32 : Add(f32,f32) and f32 : Add(f32,f64). We'd have to introduce a lot more experimental type machinery to make this work (a "floating point functional dependency", I guess?) and I'm not sure there wouldn't be fallout (for example, I can foresee issues with higher-kinded type parameters).
- acqq 13y agoIt's very simple: as soon as the floating value is read, it should go to the 64-bit virtual register unless you specify that register to be 32-bit too. If you want to use 32-bit partial results, you should say that explicitly. Autovectorization can be applied even if the rules are like suggested.