18 ms·
Forecasting s-curves is hard
- api 6y ago"It may not be surprising that in the exponential growth phase the estimate is very bad, but even in the linear phase (when 40+ points are available) the correct curve has not been found. In fact, it is only once the data starts to level-off that the correct s-curve is found. This is especially unhelpful when you consider that it can be quite hard to tell which part of the curve you on; hindsight is 20-20." Does this perhaps explain why we are so bad at calling the top of economic bubbles and similar phenomena? Maybe there literally is not enough information until the very end. It's not that we're dumb. It's that we can't do the mathematically impossible. I have a sense that this might be a very important and profound principle that might explain a lot of seemingly irrational behaviors.
- wazoox 6y agoMaybe. But there also seems to be a tendency not to admit that any exponential phenomenon in the real world (Moore's Law, economic growth, etc) is always the climbing part of an s-curve. Call it wishful thinking, or optimism, or self-delusion, according to your preferences :)
- ahartmetz 6y agoFor Moore's Law it has always been clear that it was going to end due to the size of atoms, it just slowed down earlier and faster than expected. EUV is at least 10 years late and no elegant solution has been found. It is still very difficult, complicated and expensive. I know somebody who works on EUV sources. The optics are even more difficult than the light sources. For the whole system, you get milliwatts on the wafer for megawatts of electrical power to the EUV source. But it's starting to work just well enough that it makes sense to use it.
- KarlKemp 6y agoI believe our supposed inability to predict economic bubbles (and other financial crises) is mostly just tautological: those crises we can predict never happen because they are "self-falsifying prophesies": If enough people believe the stock or housing market is overheating, they will stop buying or might even speculate against further rising prices. Thus, any bubble that does manage to grow to significant size before bursting is necessarily "unforeseen". (of course because of that thing with the monkeys, the keyboards, and online message boards, the will always be plenty of people that did see it coming, but I would wait to buy their book until they do it a second time). I know it's always trendy to hate on economists (some people have taken that idea all the way to creating cryptocurrencies and reinventing economics along the way). But comparing, say, the 2008 crisis with the 1920 or even the 1970s, I can't shake the feeling that maybe economists have become slightly better over time. The gold standard fandom that was all the rage for a while essentially rests on the idea that interventions by central banks are worse than doing nothing, and the evidence now seems overwhelming that we can do better than this (admittedly low) benchmark.
- willis936 6y agoEconomists are the analyzers, not the players. I’m not under the impression that it’s popular to dislike economists. I disagree that bubbles are unpredictable because the assumption that the market is self regulating is the exact opposite of the market’s actual nature.
- shazzzm 6y agoI think Taleb makes a similar point in his books - if you can forsee an economic crisis, you take steps to avoid it and therefore it doesn't happen. But then everyone asks why you were wrong in the first place, as the crisis didn't happen
- yters 6y agoThere is no such thing as an 'exponential curve' in our finite world. And the derivative of the sigmoid is a bell curveish thing.
- Gravityloss 6y agoPiecewise. A exponential forecast may be a better model for example for policy making on some time period than a linear one would be.
- littlestymaar 6y agoIt's a logistic function, which is different from the “bell curve”. They look pretty much the same, but their mathematical properties are really different so they should not be confused.
- CrazyStat 6y agoThe left tail of a sigmoid looks a hell of a lot like an exponential. The left tail of the derivative of a sigmoid also looks like an exponential.
- tlarkworthy 6y agoYes, thats why exponential curves taper off to an S-curve in the real world (e.g. population growth is exponential at first, then tapers off when hitting food limits). This article is about estimating the more realistic s-curve in the real world. The point it: You cannot get a sense of the end by observing the beginning... even in the super idealized case of the data coming from a perfect s-curve, let alone the real world. Its a great article.
- empath75 6y agoWhen you’re talking about a highly contagious disease it’s more or less a distinction without a difference. It’s exponential until very close to the point where a significant percentage of the vulnerable population is infected. It’s not super interesting to point out that it’s no longer growing exponentially, when half the population has it already.
- 6y ago
- BenoitP 6y agoWell, considering all s-curves are exponentials at the beginning, it naturally flows that you can only make an accurate model after the inflexion point. Since we're on HN, I'll plug a question I've had for a long time: Do IPOs/liquidity events/exits/etc all happen when the right price has been determined; ie when we can see what size the company will be, and it is no longer useful to capitalize it further? When all growth paths have been explored. Is that time the inflexion point? What does the conversation look like with the VCs when they see the inflexion point?
- claudiusd 6y ago> considering all s-curves are exponentials at the beginning Sorry, but I hear a lot of people saying this and it's driving me crazy. S-curves are S-curves from the beginning, not exponentials. It can be useful to use an exponential growth model at the beginning of the curve for short-term forecasting, but these two models will diverge dramatically at the S-curve inflection point. Not that we shouldn't plan as if exponential growth will occur in a crisis like we're in now, but many people I know don't understand these dynamics and it has lead to a lot of undue panic.
- JadeNB 6y agoThank you for this response. All smooth curves are also approximately linear at all points, but that doesn't mean that we can't usefully predict and model an appropriate non-linear fit.
- pbhjpbhj 6y agoNot disagreeing, what you say seems right to me, but it ("All smooth curves ...") also seems like the sort of result that there might be a counterexample to, like a smooth curve that at no scale has a linear approximation. Maybe a fractal with a smoothly varying generator?? For curves of infection rates I don't doubt your verity.
- umanwizard 6y ago
- littlestymaar 6y ago> In other words, data enthusiasts (such as myself) should leave the modelling up to the professionals. This is really important to understand, and unfortunately often overlooked, especially by economists, who often work with good data and solid mathematical background but no prior domain knowledge and no reference to literature of the given field.
- glofish 6y agono, I really don't think so. no epidemiologist actually has seen anything like what we are seeing today, they are not all that well prepared. this disease defies the models, youth is unaffected etc. If they were prepared well then they would not insist on imposing the exact same rules for NYC as for rural NY State. The two places could not be any more different.
- pbhjpbhj 6y agoDefine "like". People haven't seen the exact scenario, but people - epidemiologists and others - have studied the Spanish Flu, and MERS, SARS, H1N1, seasonal flus, etc., which all have useful similarities.
- glofish 6y agonone of these diseases are anything like this, no person under 18 has died of the disease in Italy, 3% or less of cases in population under 20 etc. many people don't show symptoms but are quite infectious etc which disease is like that? none actually.
- MrPatan 6y agoGiven how bad the data about the current pandemic is, why do you think data about previous pandemics was any better?
- frank2 6y agoBecause there has been more time to collect and analyze the data about the previous ones.
- m3kw9 6y agoIs hard because for example to predict case load, at any point on the curve you could have policy changes, surprise increase of case load for various reasons(sudden increase of tests), these are just a few of possible hundreds of high impact events that can impact length and slope at any point
- panarky 6y agoYes, exactly. It's tempting to try to compute coefficients X weeks in, so we can forecast X+26 weeks in the future. But the coefficients aren't fixed. Coefficients change drastically due to public policy, individual actions in response to news and social media, culture of local communities, degree of compliance with public policy, travel between regions with different rates of infection, etc. So you're not fitting a curve to the data, you're modeling dynamic human actions which are not nearly as easy to forecast.
- deleted 6y ago[deleted]
- tlarkworthy 6y agoThis article is about estimating a an s-curve in the real world. The point it: You cannot get a sense of the end by observing the beginning... even in the super idealized case of the data coming from a noisy s-curve. This learning obviously transfers over to the real world, where the data is going to be strictly worse than the idealized case (i.e. it won't be a perfect s-curve). Its a great article, with applications to pandemic forecasting.
- pif 6y ago> However, in my experience “intuition” and “mathematics” can often be hard to reconcile. They can be hard to do, but you must absolutely reconcile them for your effort to be fruitful. Until they are separated, your intuition or your mathematics is wrong, and you can't know which is.
- vajrabum 6y agoLeaving modeling up to the professionals is completely the wrong lesson. Time series forecasting from business, epidemiological or economic data forecasting is hard so you should always take the results with a large grain of salt. The professionals get it wrong too and just like the weather nearby points from the model are more likely to be accurate than distant ones. Smoothing helps. Estimating error helps. Domain knowledge helps. Experience in applying the model helps. There are techniques other than the one demonstrated in the article. Sometimes those help. There are also professionals who create models which are used to support the bias of the modeler or the modelers employer (i.e. cherry picking). Politicians and pundits more often than not take the results of these models and draw conclusions which are unwarranted or at least highly uncertain without mentioning the uncertainty or only giving it lip service.
- glofish 6y agoyes, I do think that there is nothing wrong with fitting the data and asking questions suggesting that only an "expert" can possibly fit the data correctly is the wrong conclusion
- kurthr 6y agoIt's certainly true that forcasting using exponential fits is bound to have huge error bars, but it also seems foolish to "show" that it can't be accurately done early on a curve without noticing that it is fairly easy to fit (and relatively stable) once it becomes non-exponential (e.g. once rate of change is constant). Of course in many cases there are external bounds that are similarly useful, for example when you know the maxima (e.g. full population size). Then the question to the early data is simply whether a fit to an exponential is "better" than successively higher order polynomials. Currently, this forcasting information seems not particularly helpful at all for fitting epidemiological data for Covid19 since nowhere do we have data for a case where strict distancing hasn't been an effect of high death and hospitalization rates well before 50% of the population would provide any effective immunity.
- wenc 6y agoYes. On one hand, I tire of numerous hot takes that start with "I'm not an epidemiologist, but..." which is followed by opinion un-anchored by practical experience and are coupled with a disregard for what has come before (community context). There are also a number of people in adjacent fields who are taking advantage of the situation to advance their careers. On the other hand, I realize most experts carry the baggage of unrecognized assumptions in their toolkit, which need to be challenged (at the first principles level) by people outside the domain who don't carry the same baggage. The ideal situation would be to leverage the expertise of both epidemiologists and non-epidemiologists on a common platform (that is not Twitter) and have them check and build on top each other's work.
- darksaints 6y agoThe reason it is hard is that it is using a stateless model to approximate an inherently stateful and often chaotic process. Take tech adoption for example. Often incorrectly represented as s-curves, they are the result of the inherent cost/benefit of the technology combined with the stateful diffusion process of communication and inherently human resistance to change. And being inherently stateful, you can have chaotic influences in that diffusion process. For example, the idea of microservice architecture had an absolutely massive diffusion jump the moment Amazon sent out that now-famous email mandating adoption. It wasn't linear, it wasn't exponential...it was a discrete step, and a very large one at that. These are everywhere too, because communication doesn't propogate like bacteria grows, it propagates due to extremely non-linear levels of influence. Bill Smith, 45 year old mid-level programmer for a tiny Midwest bank, will never have the tech-adoption influence levels of a Steve Jobs or Alan May or Linus Torvalds. A better option for modeling would be to use Monte Carlo methods or systems methods. Something that acknowledges the inherent statefulness of the process.
- KarlKemp 6y ago> Often incorrectly represented as s-curves, they are the result of the inherent cost/benefit of the technology [...] You're taking from two completely different levels of abstraction here. Tech adoption rather obviously happens in an S-shaped curve, at least sometimes. See the article for examples. And of course that shape, and its exact parameters, are the result of some underlying processes. These two things aren't contradictory. Outside air temperature follows a roughly sinusoid curve. It's the result of the earth turning and therefore alternating between night and day. But it's still sinusoid. And the S-curve does acknowledge state, or it would just be exponential. Yes, there are different approaches to disease modelling, such as agent- and rule-base simulations. Unfortunately, they tend to be really bad because we just don't enough data to satisfactorily simulate societies at the level necessary for this application.
- darksaints 6y agoBut having an s-shaped curve is not the same thing as being a sigmoid function, in the same way that having a bell-shaped density is not the same thing as having a gaussian distribution. There a tons of processes out there that can approximate well with either of those two, while being extremely different in extrapolation. One thing that can immediately disprove mathematical sigmoid modeling: curve symmetry. If the early adoption exponential growth is not exactly the same shape as the late adoption slowdown, then you don't have a sigmoid function. And that's the problem: if you're using a single function to model the result of two (or more) separate processes with distinct mechanics and parameters, you're going to have the same exact pitfalls as you would trying to model a bimodal process with a single probability distribution. Namely that they might fit well with interpolative methods, but completely fail with extrapolative methods (like forecasting!).
- whiw 6y ago>Many of us will have learnt in school that if there are three parameters to be found, you need three data points to define the function. 4 points are needed in my universe.
- bhouston 6y agoI was thinking the same thing. Exponentials are so sensitive to almost nothing at the beginning and then it is too late to react. The only solution to not be late in reacting is to over react. I was correct in my predictions of the disease coming to North America and the stages it would entail but I was wrong about the timeline because so as it happened it happened so fast.
- kurthr 6y agoSimilar situation, and I knew it could happen very fast (3 day doubling periods do that), but without data I still wouldn't have predicted (and can't now) where it would get bad and where it would middle around. What we do now know is that it changes VERY fast so you can't wait for your ICUs to fill up or a lot of people will die... you have to act early. And if it came once this fast... it can again, if our guard goes down (and it will).
- krastanov 6y agoI am bothered the animation does not include confidence intervals or error bars for the fit. The way these confidence intervals would shrink as more data points are available would tell a just as important part of the story.
- chillingeffect 6y agoyes, a better problem definition would in this case reveal that the estimate becomes quite good (in my definition) around the 50% time mark, when the transition from ^x to ^-x takes place. More specifically, I mean that although the stable point is still off by e.g. 25%, the important thing is that the stable time of the curve is well-estimated. you know it's no longer increasing exponentially...
- nmca 6y agoWould be intrigued to see what the story is like with priors on the parameters and a credible interval on the output.
- paulpauper 6y ago>S-curves have only three parameters, and so it is perhaps impressive that they fit a variety of systems so well no it does not. it just so happens that the solution of certain differential equations produces an s-shape. it has nothing to do with having only 3 parameters. Very sophisticated models with many of parameters and conditions can produce this shape.
- ericjang 6y agoThe author was probably talking about the logistic function being parameterized by 3 parameters, x0, L, and k [1]. The point the author is probably making is that if you are fitting a perfect logistic model to data, 3 data points should be sufficient to determine 3 parameters that unambiguously parameterize the curve. A separate set of 3 parameters also parameterizes the SIR compartmental model (https://en.wikipedia.org/wiki/Compartmental_models_in_epidemiology#The_SIR_model https://en.wikipedia.org/wiki/Compartmental_models_in_epidem...), the solution of which also looks like a logistic curve. But this is a model-based (dynamics-based) solution whereas one may be interested in just fitting the logistic model based on the assumption that it's going to be logistic. [1] https://en.wikipedia.org/wiki/Logistic_function https://en.wikipedia.org/wiki/Logistic_function
- clairity 6y agothe article correctly points out that 3 parameters need to be estimated but then jumps directly to modeling those 3 parameters with 3 points and that that will always be wrong. that’s not the right intuition. you could model a logistic curve with just 3 points if the error in those measurements tended toward infinity. the further apart the points are, the less tight the error bars need to be. the problem with real-world modeling/curve-fitting is that measurements are super noisy and the errors in them are significant.
- graycat 6y agoYes, as in the OP, S curves can be challenging to work with, in particular, to use, say, early in the history of smart phones, to make long term projections from the number of smart phones sold each day for each of the last 30 days. But there is some good news: The data used can vary, and in some cases good projections can be easier to make. Can see, e.g., for COVID-19 the recent https://news.ycombinator.com/item?id=22898015 https://news.ycombinator.com/item?id=22898015 https://news.ycombinator.com/item?id=22897967 https://news.ycombinator.com/item?id=22897967 https://news.ycombinator.com/item?id=22900104 https://news.ycombinator.com/item?id=22900104 https://news.ycombinator.com/item?id=22902667 https://news.ycombinator.com/item?id=22902667 In the third one of those we have that the projection from a FedEx case is the solution to the first order ordinary differential equation initial value problem y'(t) = k y(t) (b - y(t)) There for data we used y(0) and b. Then we guessed at k. Had we used values of y for the past month, we could have picked a better, likely fairly good, value for k. Lesson: Fitting an S curve does not have to be terribly bad. The key here is the b: The S curve of the solution is the logistic curve, and it rises to be asymptotic to b from below. Knowing b helps a LOT! When have b, are no longer doing a projection or extrapolation but nearly just an interpolation -- much better. For FedEx, the b was the capacity of the fleet. For COVID-19 the b would be the population needed for herd immunity (from recovering from the virus, from therapeutics that confer immunity, and a vaccine that confers immunity). Knowing b makes the fitting much easier/better. To know b, likely need to look at the real situation, e.g., population of candidate smart phone users, candidate TV set owners, market potential of FedEx (as it was planned at the time), or population needed for herd immunity for the people in some relatively isolated geographic area. Then in TeX source code, the solution is y(t) = { y(0) b e^{bkt} \over y(0) \big ( e^{bkt} - 1 \big ) + b} Can also use a continuous time discrete state space Markov process subordinated to a Poisson process. Here's how that works: Have some states, right, they are discrete. For FedEx, that would be (i) the number of customers talking about the service and (ii) the number of target customers listening. Then the time to the next customer is much like the time to the next click of a Geiger counter, that is, has exponential distribution, that is, is the time of the next arrival in a Poisson arrival process (e.g., the time of the next arrival at the Google Web site). So at this arrival, the process moves to a new state where we have 1 more current customer and 1 less target customer. Then start again to get the next new customer. The Markov assumption is that the past and future of the process are conditionally independent given the present state; so that justifies our getting to the next state using only the current state -- given the current state, for predicting the future, everything before that is irrelevant. What is a Markov process, what satisfies that the Markov assumption, can depend on what we select for the state -- roughly the more we have in the state, the closer we are to Markov. In particular, if we take the whole past history of the process as the state, IIRC every process is Markov. But Markov helps in something like the FedEx application since that state is so simple. We get to use continuous time since the time to the next change of state is from a Poisson process whose arrival times are the continuum -- that is, we don't have to make time discrete although it is true that the history of the process (one sample path) has state changes only at discrete times. So, for state change, and for some positive integer n we have some n possible states, then for i, j = 1, 2, ..., n, we can have some p(i,j) which is the probability of jumping from state i to state j, that is, we have an n x n matrix of transition probabilities. [p(i,j) is the conditional probability of entering state j given that the last state was i.] For two jumps, square that matrix. Now there is a lot of pretty math -- get some limits and eigenvectors of states, etc. Actually fairly generally there is a closed form solution to the process. Alas, often in practice that closed form is useless because the n and the n x n are so large, maybe n^2 in the trillions. E.g., in a problem I solved for war at sea, there were Red weapons, Blue weapons, on each side some number types and some number of weapons of that type. The states were the combinatorial explosion. Then there were the one on one Red-Blue encounters where one died, the other died, both died, or neither died. The time to an encounter was the next arrival of Poisson processes, also Poisson. Well, that was an example where there was a closed form solution but n and n x n were wildly too large for the closed form solution but running off, say, 500 sample paths via Monte-Carlo was easy to program and fast for the computer. So, sure the software reported the average of the 500 sample paths. On a PC today, my software would be done before could get finger off the mouse button or the Enter key. This approach is fairly general. And since what I did included attack submarines, SSBN submarines, anti-submarine destroyer ships, long range airplanes, etc., there should be no difficulty building such a model for COVID-19 that included babies, grade school kids, ..., nursing home residents, people at home, people working nearly alone on farms, .... Back to S curves, IIRC dropping out of the math for the n x n matrix and its powers is an S curve. So, in a broad range of cases, always get an S curve although a different curve depending on, yes, the p(i,j) and the initial state. Uh, when no one is left sick, the Markov process handles that as an absorbing state -- once get there, don't leave. For the n in the billions, the n x n is really a biggie. So, for the submarine problem I did, J. Keilson, Green's Function Methods in Probability Theory. asked "How can you possibly fathom that enormous state space?". That is a good question, and my answer was: "After, say, 5 days, the number of SSBNs left is a random variable. It is bounded. So it has finite variance. So, both the strong and weak laws of large numbers apply. So, run off 500 sample paths, average them, and get the expectation within a gnat's ass nearly all the time. Intuitively, Monte Carlo puts the effort where the action is.". Keilson was offended by "gnat's ass" but liked the math and approved my work for the US Navy. That question and answer are good to keep in mind. There is more in, say, Erhan Çinlar, Introduction to Stochastic Processes, ISBN 0-13-498089-1, Prentice-Hall, Englewood Cliffs, NJ, 1975. For why the arrival times have exponential distribution and why we get a Poisson process, Çinlar has a nice simple, intuitive, useful axiomatic derivation. There is more via the renewal theorem in William Feller, An Introduction to Probability Theory and Its Applications, Second Edition, Volume II, ISBN 0-471-25709-5, John Wiley & Sons, New York, 1971.
- surroundingbox 6y agoI suppose that the s-curve with 3 parameters that the author is talking about is the logistic function. In general, if you consider a differentiable function of three parameters and try to determine and interval for the values of the parameters of that model then the length of that interval is bound by the ratio of the error in the data over the derivative with respect the parameter. For example estimating the parameter k (wikipedia logistic growth rate) with points such that x near x0 = (wikipedia midpoint of the sigmoid) is hard, since the derivative of the function with respect to k at x=x0 is zero. So mathematically this seems to be a well known fact when one try to estimate parameters from datapoints.
- cracker_jacks 6y ago> the length of that interval is bound by the ratio of the error in the data over the derivative with respect the parameter This is interesting! Could you expand on this a bit? Why is the length of the interval bound by the ratio of the data error over the derivative?
- surroundingbox 6y agoThe general case require some work and conditions. But to give a hint, the case of only one parameter is an application of the mean value theorem (1). Suppose a model (y = f(p,x) ) with only one parameter p0 and an exact point (x0,y0) (that is y0=f(p0,x)) and a data point (x0,y1) such that y1-y0=error in the data. And that there is a value p1 of the parameter such that f(p1,x0) = y1, then y1 - y0 = f(p1,x0) - f(p0,x0) = f'(sigma) . (p1-p0), so that p1-p0 = (y1-y0)/f'(sigma) that is (error in the parameter) = (error in the data)/(derivative with respect to the parameter) where sigma is between p0 and p1. The general case is a generalization of this idea using the mean value inequality. (1) https://en.wikipedia.org/wiki/Mean_value_theorem https://en.wikipedia.org/wiki/Mean_value_theorem
- tel 6y agoThis is accounted for in more professional methods which estimate error. During a period of exponential growth, the error 6 weeks out is very sensitive to tiny errors occurring in immediate measurement. It's not so much that fitting exponential or S-curves is hard as much as even a very good fit is likely to have very significant error bars. The hard parts about modeling the current situation involve the poor and limited data. The dynamics of the system are highly variable and unobserved and thus require clever tricks to reduce uncertainty by expert knowledge, clever alternative signals, and leveraging other models. So ultimately, it's just not the forecasting but the response. People rarely handle uncertainty very well and a situation with uncertain exponential growth ends up with exponential uncertainty. Making policy decisions where the predicted outcome spans orders of magnitude is tremendously challenging.
- acqq 6y agoExactly. And I have a very recent example, yesterday somebody responded to one of my replies with: "Cuomo claimed to need 40k ventilators after the lockdown was put in place. He ended up needing only a fraction of that. Obviously, without the lockdown, there would likely be more needed. Of course, modeling has errors, but this is an extremely large error, and one that has direct policy implications." (1) My response was that it wasn't "a large error" but what only 6 more days of exponential growth would have made still insufficient: https://news.ycombinator.com/item?id=22911465 https://news.ycombinator.com/item?id=22911465 Also as an illustration of the "uncertainty" one can see the shaded areas clearly here: https://covid19.healthdata.org/united-states-of-america https://covid19.healthdata.org/united-states-of-america As I write this, the model on that page was last time updated based on the data up to the 15th of April, only 4 days ago, and its best estimate for number of deaths in USA on 19th of April was 37,056. Also as I check the https://www.worldometers.info/coronavirus/country/us/ https://www.worldometers.info/coronavirus/country/us/ it's already 40,423. But the 95% error area for 19th is 39,721-56,108. It's that hard. Also at this moment the projection for 1st of June is 60,262 (34,063-140,106) and we can compare all these values after some later updates, e.g. in 6 and 12 days. 1) careful observers know that there is one more talked about person who is used the same pseudo-argument before.
- 6y ago
- nabla9 6y agoTo predict the s-curve for epidemic, you need to know R0. The upper limit of the curve is 1 - 1/R0. If you use mitigation and effective reproductive number the number is 1 - 1/Rt, where Rt varies with mitigation effort.
- hinkley 6y agoA turning point in my understanding of S Curves came when I encountered an article that showed an S Curve as the cumulative area under a normal distribution. Which is great if your velocity on the project retains a normal distribution over time. But missing requirements, bad risk analysis, and team dynamic changes can make for a long velocity tail. Which means the last 10% of the project accounts for the other 90% of the project time. If however you have managed to do the important work first, you drop less important features and ship 95% of the planned features on time.
- FabHK 6y agoIt’s a common misconception, but the S curve generated by disease, for example, with exponential growth at the beginning, corresponds not to the normal distribution, but the logistic distribution, which has fatter tails than the normal. The bell curve drops to zero with exp(-x^2), that is extremely fast. The logistic distribution (derivative of the S-curve described here) drops to zero with exp(-|x|), that is exponentially, but not as fast as the normal distribution. https://en.wikipedia.org/wiki/Logistic_distribution https://en.wikipedia.org/wiki/Logistic_distribution
- JoelJacobson 6y agoI made a R web-app in Shiny which does curve fitting of a Four Parameter Log Logistic function (which is the S-curve discussed in the article) against the John Hopkins data: https://joelonsql.shinyapps.io/coronalyzer/ https://joelonsql.shinyapps.io/coronalyzer/ https://github.com/joelonsql/coronalyzer https://github.com/joelonsql/coronalyzer "Sweden FHM" is the default country, which is a different data source, it's using data from the Folkhälsomyndigheten FHM (Swedish Public Health Agency), which is adjusted by death date and not reporting date as the John Hopkins data is.
- deleted 6y ago[deleted]
- autokad 6y agoI did the covid19 week1 and week2 kaggle competitions (I think they had 4 of them) below. If you are interested, this is a fun way to play around with the data, and it shows how hard it is. This I tried: - Weibull Gamma distributions, but it was impossible to find good parameters for the distributions without exploding. It would only work if I put in an additional parameter saying 99% of the population wouldnt get it. It would come up with good shapes of growth but the predictions were far too below actual values in the future. - Logistic curves. usually great for countries that already ran up the curve but terrible for ones still in the exponential phase (as the article states). Also kind of useless for countries that didnt even begin their journey up the curve. - light gbm: good for predicting the next day but terrible many days out. It seems other counties curves do not help that much - SARIMAX: really good but later predictions would explode, like showing 4 million deaths in france, etc. I tried to get around these by ensembling them together, but overall I did very poorly at predicting coronvavirus. I still want to get better at this, so if anyone has any good suggestions, please share. Also you can check out what other kagglers have done as well https://www.kaggle.com/c/covid19-global-forecasting-week-1 https://www.kaggle.com/c/covid19-global-forecasting-week-1 https://www.kaggle.com/c/covid19-global-forecasting-week-2 https://www.kaggle.com/c/covid19-global-forecasting-week-2
- LolWolf 6y agoThe problem is just too ill-defined to fit directly with a logistic model or really any general-purpose model (such as gbm, etc). You need better priors, either in the form of good, specified dynamics, or in other terms. You could do a regularized stratified model [0], for example, in the case where you have plenty of data in one country but not in others, or similar cases which incorporate dynamics (log-log models of growth are also pretty good). Overall, I think it's just too hard of a problem as it is very ill-posed—even if we knew the exact dynamics, any small perturbation in the data samples significantly changes the best fit (in other words, the dynamics of the system are 'chaotic'). ----- [0] https://stanford.edu/~boyd/papers/pdf/eigen_strat.pdf https://stanford.edu/~boyd/papers/pdf/eigen_strat.pdf
- earthicus 6y agoI briefly studied mathematical biology, and remember there being some debate about whether tumor growth was more accurately modeled by logistic growth or Gompertz growth [1]. I'd be curious to know whether your fits get better or worse if you replace your logistic-based model with a Gompertz-based model. [1] https://en.wikipedia.org/wiki/Gompertz_function https://en.wikipedia.org/wiki/Gompertz_function
- FabHK 6y agoOne thing to note (from looking at the graph) is that the noise seems to be additive with constant standard deviation (and presumably floored floored such that the sum doesn’t go negative). That means that there is huge relative error initially (we have 10 infections +/- 100), and very little relative error eventually (we have 1000000 infections +/- 100). I assume the forecasts would be better if the error were multiplicative (in other words, with standard deviation proportional to the current value). However, I think the main point stands: the forecasts get much better once one approaches the inflection point.
- YetAnotherNick 6y agoI tried to do some exploration with the coronavirus data to get some idea on the final number. One of the best plot I found that could tell the final number is plotting the percentage of cases in the next week with the cases till now. It is like negative half parabola and the time it meets the x axis will give the final number per country. This is the final plot: https://i.imgur.com/o54t0Ts.png https://i.imgur.com/o54t0Ts.png. You could make a good guess from the data how many people it will affect in the lifetime per country by continuing the same pattern till it reaches x axis.
- anonytrary 6y agoForecasting solutions to a differential equation is hard, especially when there are infinitely many solutions. It is a matter of keeping constants up to date in light of new information. If those constants are wrong, the whole model is essentially useless.
- jkqwzsoo 6y ago> It is a matter of keeping constants up to date in light of new information. I'd argue it's a matter of calculating confidence intervals for all of your fitted parameters and displaying joint confidence intervals if they display a large degree of interdependence/nonlinearity, as well as showing the fitted residuals beneath the main plot so it's easily apparent what issues the fitted function might have, such as high leverage, or overfitting. The prediction made will be from the best-fit parameters, but the contributions of other possible models can be used to infer a distribution of potential outcomes. A parameter sensitivity analysis wouldn't hurt either. Then you'll immediately know if you're looking at (or about to publish) junk.
- LolWolf 6y ago> Forecasting solutions to a differential equation is hard, especially when there are infinitely many solutions. Not sure what you mean by infinitely many solutions, but this is not true in general. A silly example is y'(t) = Ct, where, for most distributions, you would only need a few points to get both C and the initial condition to reasonably high accuracy (~O(sqrt(n))). More complicated examples exist that have much more interesting dynamics, but their general trajectories are just not as sensitive/chaotic (w.r.t. the initial parameters). > If those constants are wrong, the whole model is essentially useless. I think what makes this hard is not if they're wrong, but rather, being even just a tiny bit wrong makes the whole future prediction change drastically. In other words, two possible inferred parameters which are statistically indistinguishable given our current observations will yield incredibly different outcomes under many of these models. > It is a matter of keeping constants up to date in light of new information. Indeed! :)
- mirimir 6y agoI've loved s-curves for decades. Way back in the lab, to analyze ligand-binding assays. And not that long ago, as a litigation consultant. > For technological changes, can the final level-off be reasonably estimated? It helps a lot if it's 0% or 100% :) Given that, you get a decent long-term fit, after about half the time to plateau.
- 6gvONxR4sf7o 6y agoWhat's even harder is that you're usually trying to forecast for a reason other than pure academic curiosity. Like, should I be worried about the coronavirus? Should we stay inside? When can we reopen business? But the decisions people take determine the curve and the predicted curve determines people's decisions. If you decide not to social distance, you're changing the future. For that reason, forecasting is better done under different scenarios. For example, don't tell people approximately X (+/- a lot) will die. Tell them approximately X (+/- a little less) will die if we don't social distance and approximately Y (+/- a little less) will die if 95% of us self-quarantine. It's harder than just fitting a curve, but that's kind of the point. If you want actionable predictions, it is harder.
- sradman 6y agoOne of the only good things to come out of this pandemic is the increased emphasis on S-Curves. Models tend to use the predictive power of exponential curves to estimate the steep part of the S-Curve but these predictions are short term and are best applied to planning scenarios. What is missing from this article is the relationship between S-Curves and Bell Curves. We can use the Rules of Thumb associated with the Normal Distribution to think about peak growth rate and standard deviations. The health data.org curve fitting is a decent Fermi Estimate based on observed data. Models are always wrong but sometimes they are useful and I hope we start to discuss the underlying key assumptions used in each case rather than focusing on their imperfect predictive power.