For decades people have said that compiler's code generation is at par... and it slowly has become more true. I read a lot of disassembly listings and I'm often pleasantly surprised at how clever some of the code is. However, there's also plenty of times where I see plainly wasted instructions as well.
That's not the real reason that assembly is often not worth it, even in situations where it once was. The real culprit is that for a long time CPUs were outpacing memory in speed gains, so a cache miss became more and more expensive relative to the clock speed.
This meant that more and more micro-optimization work has gotten focused on cache behavior. C gives you just as much control over how your data structures are stored as assembly would. Maybe you'll need to define some prefetch() macros depending on your compiler, but that's about it.
There are certainly some remaining cases where you really want to control things at a register-allocation level (encryption, codecs, fancy floating-point things) However, most projects are better off focusing on improving their memory behavior rather than trying to get to that cache miss in 80 instead of 82 cycles.
The other huge shift in performance-oriented computing is, of course, the availability of more and more CPU cores. Again, lots of work to do but assembly doesn't give you any advantage at all.
So even if you're good enough at writing assembly to beat the compiler (and most people aren't) you probably should have been spending your optimization effort on other things.
Comments
For decades people have said that compiler's code generation is at par... and it slowly has become more true. I read a lot of disassembly listings and I'm often pleasantly surprised at how clever some of the code is. However, there's also plenty of times where I see plainly wasted instructions as well.
That's not the real reason that assembly is often not worth it, even in situations where it once was. The real culprit is that for a long time CPUs were outpacing memory in speed gains, so a cache miss became more and more expensive relative to the clock speed.
This meant that more and more micro-optimization work has gotten focused on cache behavior. C gives you just as much control over how your data structures are stored as assembly would. Maybe you'll need to define some prefetch() macros depending on your compiler, but that's about it.
There are certainly some remaining cases where you really want to control things at a register-allocation level (encryption, codecs, fancy floating-point things) However, most projects are better off focusing on improving their memory behavior rather than trying to get to that cache miss in 80 instead of 82 cycles.
The other huge shift in performance-oriented computing is, of course, the availability of more and more CPU cores. Again, lots of work to do but assembly doesn't give you any advantage at all.
So even if you're good enough at writing assembly to beat the compiler (and most people aren't) you probably should have been spending your optimization effort on other things.