4 ms·
And haskell does this. This is way over engineering for performance. Different level. It's the same in other language when you go deep into performance engineer
by UserIsUnused 7y ago
And haskell does this.
This is way over engineering for performance. Different level. It's the same in other language when you go deep into performance engineering, it just depends how deep considering if the algorithm fits better a functional or imperative model.
- mschaef 7y ago> It's the same in other language when you go deep into performance engineering Take a look at the optimized code for the C version of wc: https://opensource.apple.com/source/text_cmds/text_cmds-68/wc/wc.c.auto.html https://opensource.apple.com/source/text_cmds/text_cmds-68/w... There are a couple special cases for a line count and a total character count, but it's essentially a couple of nested loops. I agree with you that going deep into performance engineering brings out a lot of esoteric edge cases, but in this case for the C program it wasn't even necessary to go down that path. Basic concepts in C outran or at least matched something heavily optimized in Haskell. I'm a big fan of higher level languages, and use them almost exclusively myself in my day job, but there's something to be said for basic constructs that just run quickly out of the box. It may be harder to optimize the C ode, but you're also less likely to need to do so.
- cmroanirgo 7y agoDisclaimer: old timer (not sure if you whippersnappers do the same anymore). The few times that I really needed performance out of a particular C/C++ function, I'd enable the compiler's output to generate ASM, and then see what was being generated, and either rework the code (to keep it readable for the next guy), or optimise directly with ASM, and often an iterative mix of the two. Loop unrolling (as in Duff's device) was a common one for array iteration performance improvements. I don't see a lot of "optimized C" in the wc impl. I'd imagine that wc optimised for use in a GPU would be pretty spectacular, even beyond using a few ordinary cores like the Haskell result did... but as you say, it's pretty unnecessary in this case.
- mschaef 7y ago> Disclaimer: old timer (not sure if you whippersnappers do the same anymore). Heh... I don't know that I qualify as an 'old timer', but it's been a while since 'whippersnapper' would've applied to me. :-) > The few times that I really needed performance out of a particular C/C++ function For me, at least, performance has almost never been a problem, particularly when writing in C. > I don't see a lot of "optimized C" in the wc impl. I'd imagine that wc optimised for use in a GPU would be pretty spectacular, Whenever I think of optimized C code, I immediately think of the work done to make Gnu Grep fast. https://lists.freebsd.org/pipermail/freebsd-current/2010-August/019310.html# https://lists.freebsd.org/pipermail/freebsd-current/2010-Aug... A lot of that boils down to picking an algorithm that lets them defer as much work as possible until you know you need to do it. I suppose that's one of lazy evaluation's claims to fame, but I think the level reached in Grep would be difficult to achieve in a more modern functional lannguage.
- cmroanirgo 7y agoI was writing rasterisers, image manipulators and various matrix/vec math operators for 2d and 3d. Every clock cycle counts in some situations.