3 ms·
Forget parsing time, just taking the time takes an awful lot of time. Try measuring a very quick routine (a few nanoseconds) with clock_gettime - you have to do
by necessity 10y ago
Forget parsing time, just taking the time takes an awful lot of time. Try measuring a very quick routine (a few nanoseconds) with clock_gettime - you have to do various calls and take a mean because of resolution issues, and also measure and reduce the overhead of the logging routine. This becomes a real issue because taking a mean of various calls is different than doing just one call (cache, branch prediction, etc). You could add something else inside the loop but then your fast routine becomes a tiny part of what you are measuring and the relative error explodes. It is simply not possible to benchmark routines that take just a few nanoseconds precisely and exactly with userspace routines. There are whitepapers from Intel on how to do this on their processors with a kernel module to disable preemption and IRQs and read the TSC directly with asm, but then you can't have userspace stuff... Benchmarking quick stuff is no fun.
- Beltiras 10y agoHardest thing I ever had to profile had to do with accessing Postgres from Python. I needed the cold-start time. I ended up kinda fuzzing it to find an estimate. Filled a table with proper data and accessed random keys to find an average. Even then some optimization occurred.