4 ms·
You could certainly come up with something better than this. At the very least for such a comparison to be generalizable, the raw data should be a set of timed
by fwilliams 9y ago
You could certainly come up with something better than this. At the very least for such a comparison to be generalizable, the raw data should be a set of timed routes, app estimates, and traffic/weather conditions. Given this information, you can extract similar routes and do an apples-to-apples comparison. If you want to extract general facts, you can group routes by certain attributes that you care about (e.g. high traffic areas, time of day, etc...) and do an error analysis.
But all that is beyond the point of my original comment, which states that the "benchmark" the author uses likely does not generalize to very well for the following reasons:
* The distribution of all possible routes is large and depends on many variables that I mentioned (traffic, weather, location, etc...). The author's sample of this distribution is tiny and biased towards routes frequently taken by him and his wife. You could probably choose a different 120 routes and compute completely different results.
* Averages can be misleading if you don't know the underlying distribution you are sampling from. As a trivial example, the average value of a set of samples from a Bernoulli distribution with p=0.5 is 0.5 even though all the samples are either 0 or 1. In this case, the average is not a good tool to summarize the data (unless you know it's a Bernoulli distribution). So even if the author had used a million routes, simply taking the average error doesn't say anything meaningful about the error without also understanding the error distribution.
So I think the data and methodology presented in the article are not sufficient to draw any general conclusions about how good the error estimates are in each app.