5 ms·
What time-weighted averages are and why you should care
- Animats 5y ago"a fully managed TimescaleDB service" Weighted averaging as a commercial service, all by itself? The SQL solution seems complicated, but this is a task easily performed on time-ordered data. I had to code such a time-weighted average last year, for a game client. The client gets updates for moving objects when they move a small distance or when some dwell time has elapsed. I needed a smoothed velocity from those irregularly spaced samples. The math had to be worked out, but it was just two lines of C++ in the implementation.[1] filtermult = 1.0f - 1.0f / pow(1.0f + 1.0f / fsRegionCrossingSmoothingTime, secs); // filter scale factor mFiltered = val * filtermult + mFiltered * (1.0f - filtermult); // low pass filter [1] https://vcs.firestormviewer.org/phoenix-firestorm/files/tip/indra/newview/fsregioncross.cpp https://vcs.firestormviewer.org/phoenix-firestorm/files/tip/...
- PaulWaldman 5y ago>Weighted averaging as a commercial service, all by itself? No. The full quote is: "If you’d like to get started with the time_weight hyperfunction - and many more - right away, spin up a fully managed TimescaleDB service" They're just indicating you can try this with their hosted TimescaleDB service. Not that this a standalone service.
- KarlKode 5y agoTimescale is a postgres extension that allows you to store time series data in a postgres database. Besides offering a very performant way to store specialized data it also offers additional functions to work with your time series data. Besides the postgres extension there is also a database-as-a-service option where you can get started immediately and try the technique described in the blogpost.
- g8oz 5y agoA data analyst is more likely to know SQL rather than C++.
- Animats 5y agoMy point is that this is easy to do if you just have the data in sequence. The actual computation could be done in Matlab, or Julia, or R, or Python. SQL isn't a good language for doing computations on time-series data, but you can always write out a sorted data list.
- nl 5y agotime window functions exist in SQL now and are great.
- lupire 5y agoThat looks like a non-linear time-decay function, not a time-weighted average. Why are you setting filtermult=1-stuff and then immediately doing 1-filtermult?
- monkeydust 5y agoIn finance referred to as TWAP https://en.m.wikipedia.org/wiki/Time-weighted_average_price https://en.m.wikipedia.org/wiki/Time-weighted_average_price
- donquichotte 5y agoI don't mean to be demeaning, but the concept illustrated in the last figure is known as trapezoidal integration [1], has been well-know for millennia and is easy to implement. I fail to see how the advertised "hyperfunctions" would make the process of analyzing such data any easier by adding web-services and SQL. [1] https://en.wikipedia.org/wiki/Trapezoidal_rule https://en.wikipedia.org/wiki/Trapezoidal_rule
- jcelerier 5y agoisn't trapezoidal rule the thing that gets "rediscovered" regularly ? https://academia.stackexchange.com/questions/9602/rediscovery-of-calculus-in-1994-what-should-have-happened-to-that-paper https://academia.stackexchange.com/questions/9602/rediscover...
- belter 5y agoAnd that paper was then cited 112 times. I really do not know if I should laugh or cry...
- jhgb 5y ago> the concept illustrated in the last figure is known as trapezoidal integration And, as a consequence, the concept they call "time-weighted average" is just called "mean value of a function", at least in my country. Also "hyperfunction" seems to be a terrible name. Unless we're talking about mathematics where the term has very specific meaning, how exactly is it different in programming from any ordinary function? (Possibly an aggregate function in SQL, of course.)
- TeMPOraL 5y agoOh yes, no ends of fun with this. The article mentions "Industrial IoT", but applications of time-weighed averages in industry are much, much older than the term IoT (or "Industry 4.0"). It's the cornerstone of recording data from industrial controllers and sensors, and as a functionality, it's built into most of the relevant industrial tooling (e.g. historians[0]). Support for time-weighed average and other means of retrieving and analyzing such compressed data is included in core industrial protocols like OPC Classic and OPC UA. As the article says, it enables data compression - but the article is mistaken (or simplifying too much) by saying data points are retained only "when the value changes". In typical implementations, data is retained when rate of change changes (i.e. first derivative). As long as the data points fit on the line, you can recover them all from line's end points with linear interpolation. There's some more sophistication involved in this, too: for example, when configuring data archiving, you may set up "deadbands" - basically defining how thick the interpolated line is; data points that fall on this thick line will be considered as if they were perfectly centered, and not recorded. On top of that, most systems also track metadata like "sample quality" or "engineering units used" - there are rules specifying when to record a sample if its metadata changed. All in all, it's an ingenious approach, but it comes with certain consequences: One - what this article is about - when data is stored in such fashion, you need to use time-weighed aggregations (average or other functions) to compensate for some (usually most) of original data not being recorded. Two, this technique assumes a single, continuous process being recorded. Imagine you're using a "smart scale" to record your daily weight. After a month, you give that scale to your partner. Right there you introduce a discontinuity; any data query that overlaps the period between your last stored (not measured, but stored under compression) measurement and your partner's first stored measurement, will return nonsense. These two caveats sound obvious, but I've seen industrial project getting both of these wrong at some point. -- [0] - https://en.wikipedia.org/wiki/Operational_historian https://en.wikipedia.org/wiki/Operational_historian
- djk447 5y agoNB: I'm the author of the post :) That's a nifty way of doing things that I didn't know as much about (definitely left out the historians bit as I didn't want to get too into the weeds) and have sometimes encountered the linear interpolation fit coming from them as well as the LOCF type fit, but didn't realize that's how they were sampling under the hood. That's pretty cool. I was mostly giving that example so people would understand why the LOCF option was available for the function as well, and one of the most common places where I've seen it is in (usually slowly changing) sensors in industrial settings that only record when they change. On the single continuous process, yes, this is one of the most common mistakes people can make when they do this. We find it's best to model that as a change to the relational part of the data and potentially generate a new "id" for the sensor when that sort of discontinuity happens.
- mahathu 5y agoSimpson's rule is a more accurate approach to numerical integration without being more computationally expensive: https://en.wikipedia.org/wiki/Simpson%27s_rule https://en.wikipedia.org/wiki/Simpson%27s_rule Nevertheless, this was a very interesting article! When the topic of a time-weighted average came up, my first intuition was to give each sample y_n a weight of x_n - x_n-1, that is, the distance/time to the sample that came before. Wouldn't that also suffice?
- djk447 5y agoNB: Article author here. Interesting idea with Simpson's rule! Perhaps we'll add another method... The other approach is very similar to the LOCF (last observation carried forward) approach, except that it uses the distance to the next sample instead of distance to the previous. I'm sure there are reasons to do the other way, but for most things that use LOCF they record when the value changes so it makes more sense to carry it forward. And it works pretty well as well. These are all also other integral approximation techniques, and which to use probably depends mostly on use case.
- jordache 5y agoI can't deal with this article.. every paragraph contains a reference to the product these folks are trying to sell..
- dannyz 5y agoTime-weighted average has always seemed like such an odd name to me. I assume that taking a simple average of irregularly spaced points is a very common mistake and so that is likely where this name comes from. But, as another commenter pointed out, it really is just the mean value of a function.