4 ms·
Well, if you want to "contest most the points in that page." I suggest you do so on substance, rather than by going after the person ? There is a difference be
by phkamp 9y ago
Well, if you want to "contest most the points in that page." I suggest you do so on substance, rather than by going after the person ?
There is a difference between "being ignorant" and "ignoring".
I don't think there is much in modern HiPerf computing I'm ignorant of (hint: Guess who's writing one of the most intense HiPerf scientific code right now), but there sure is a lot of it I'm deliberately ignoring with respect to Varnish.
TLB flushing, instruction caching ? When your CPU is 50+% idle most of the time, that's not really what (should) worry about.
And the "downright difficult job of an OS scheduler" isn't really that hard when all threads a waiting for I/O, wake up for a few microseconds, then go back to waiting for I/O again.
And so on, and so forth...
Again: Your total lack of actual experience running Varnish shows.
You should try it some day, you might actually find it useful :-)
- kev009 9y agoYou keep accusing me of logical fallacy but are using them yourself so I can't have a rational conversation with you. There is little CPU idle when you are doing TCP packet pacing, TLS in core, and edge compute. That's what I do on a 30Tbit/s CDN. I was just hoping you might reconsider the thread per connection but it appears not yet. Anyway congrats on the release and have a good day, sorry if the only takeaway you got from this is to be antagonized.
- phkamp 9y agoThread-per-request seems even more correct to me, given the increased cost of system-calls, now that Spectre and Meltdown fixes are in.
- jeremiep 9y agoI used to do thread-per-request, that has mediocre scaling at best. Even on the JVM this barely scales to a few thousand connections; native threads are heavier than that. I've also done a lot of task-per-request (with each thread's affinity locked to a single hardware thread to avoid ripple effects), which does scale at least an order of magnitude more than threads. I now use fiber-per-request for the best of both worlds: the easy sequential model of threads, yet the simple performance of tasks. I can't understand why someone would defend the thread-per-request model at this point. It wasn't even a good model 10 year ago.
- kev009 9y agoCorrect
- olavgg 9y agoIsn't TCP packet shaping normally done on the network interface these days so you have CPU offloading?
- kev009 9y agos/shaping/pacing/, we're trying to smooth out the bursty nature of sending from high speed links and TSO so a flow is less likely to incur large contiguous buffer drops and tail drops along its path rather than limit the flow's throughput. I have a large fleet of intel NICs on the back half of their life cycle. NIC pacers sound good in theory and I'd like to eventually partial offload but they have limitations in terms of flows and number of pacing rates so will still require a software fallback. On a timerwheel, the system overhead for no offload and a partial offload is not nearly economical as theoretical full offload. And contrary to some misinformation from one of the varnish devs in this thread, taking an SWI/context switch/scheduler overhead (these basic concepts will probably get called "babble" by the dude) comprise much of the overhead.