3 ms·
Even better! Knative will allow you to even scale your containers out-of-band if you know that something BIG will happen. See https://github.com/knative/serving
by markusthoemmes 8y ago
Even better! Knative will allow you to even scale your containers out-of-band if you know that something BIG will happen. See https://github.com/knative/serving/issues/1656 https://github.com/knative/serving/issues/1656 for the discussion.
- pwaai 8y agoThat tells me not much....is it pinging to keep instances warm?
- jacques_chester 8y agoNo. The intention is that the autoscaler will support both floor and ceiling values for its work. So while scale-to-zero is the default, you can make an economic decision that you want scale-to-1, scale-to-10, whatever makes sense for your case. Pinging won't be necessary to artificially create this behaviour. This is an example of nesting reactive control (the autoscaler) with predictive control (the min/max values).
- pwaai 8y agointeresting...by floor and ceiling is that like the minimum and maximum threshold for latency? here's my pain point. I built a serverless REST API with token authentication on Lambda. However, if many people aren't using it all the time it will sleep and then the next sucker who calls the endpoint is stuck with waiting for the serverless instance to wake up. In some cases even getting a token from a simple POST request would take an awful long time. This was a few years ago and I stopped using serverless since then. But now I'm interested in serverless because I've been hearing that the cold startup problem is being reduced. I wonder if in the future developers will be just taking core logics from serverless repository and wiring up the components, sort of like how wordpress does it without the crazy layers of PHP and bloat.
- jacques_chester 8y agoThere are two aspects here: engineering and economic. Our engineering view is that we want startup to be as fast as possible, which is a surprisingly nuanced problem with lots of moving parts that need to collectively do something smart, even before you get to the startup time of your own code. This will show a lot of improvement as we go, but right now it's early days. The economic question is about trading off the risk of hitting a slow start vs the cost of maintaining idle instances. It is impossible for Knative's contributors to solve that problem with a black box solution. What we can do is to provide you with some knobs and dials to express your preferences. Edit: I didn't answer this question -- > interesting...by floor and ceiling is that like the minimum and maximum threshold for latency? Not for now, this would be bounds on what scale the autoscaler can choose. Latency is an example of an output setpoint that an autoscaler could attempt to control, as opposed to a process input. We have in mind to make autoscaling somewhat pluggable because different people want to target different signals.