4 ms·
This has more to do with the ISA than C I'd assume. C was built as an "easy assembler". Furthermore, the computer doesn't really care about structure (in the st
by ablob 4y ago
This has more to do with the ISA than C I'd assume. C was built as an "easy assembler". Furthermore, the computer doesn't really care about structure (in the structured programming sense).
In this case the implementation of the ISA stores information on each 'ret' depending on some 'bl' that came before. One can imagine that a different optimization technique which actually leads to a speedup exists. Without a branch-predictor, the program that Matt wrote might've been faster.
Imo, this has nothing to do with paradigms, but how control-flow interacts with the system. This interaction is different between different implementations of an architecture.
Code written for Cache-locality and paradigms that work well with it, for example, only became a "winner" after caches were widely implemented. Before that, the cost of sequential array access and random array access was identical.
With caches, a linear search for an element can be faster than a binary search, even though it requires a lot more memory accessess. Thus, the optimal implementation for finding an element in an array is now dependent on size as well. (I.e. after an ordered array reaches a certain size, binary search should become faster on average).
- int_19h 4y agoThe point is that the ISA is assuming that BL and RET opcodes come in pairs - which is an assumption that does reflect what structured code is typically compiled down to. Going by the semantics of the opcodes themselves, there's no reason why RET should be treated differently from a simple indirect jump here.
- masklinn 4y ago> there's no reason why RET should be treated differently from a simple indirect jump here. There’s its very existence. If you’re looking for a general purpose indirect jump, use that.
- tsimionescu 4y agoAs far as I understand, the ISA explicitly intends for BL and RET to come in pairs. The only point of RET is to jump to the address last stored by BL. If you don't need this behavior, there's no reason to use RET - as the article itself shows, B x30 does the exact same job, and doesn't come with the extra assumptions.
- int_19h 4y agoThat's my point - the ISA is designed around the notion that function calls (a structured programming concept!) are a thing that BL will be used for, and not arbitrary jumps where you happen to have some clever use of the link address. And while it's not something that the example code in the article demonstrates, but based on the description of how it confuses the branch detector, it sounds like having multiple BLs without a matching RET for each would also be a problem.
- andrepd 4y agoSubroutines predate structured programming.
- int_19h 4y agoI was using the terminology as OP did, and yes, it is not quite right. But the point as I understood it was that branch predictors optimize around specific "structured" usage patterns of opcodes - in this case, the particular way to use BL/RET to implement function calls typical of C - to the point where any other use is too slow to be practical for anything.
- dustbitying 4y agowhat's the difference between a subroutine an a proper C function?? theyr'e so close to what they're in principle. IMO, C functions are the answer to the problem exposed in the famous quote "goto considered harmful". assembly is all about GOTOs.. C provides a *structure* way to deal with them without getting bored to day (it all becomes way to much to soon) but I think (for reasons that I wish I could get into) that in the end, the goto-based assembly code can go up to multiplication, but then with C and their computer-functions one can go beyond exponentiation. What is there beyond, I can only name by reference but wouldn't say I undesrtand it... which tetration. so riddle me this: why is 2[op]2=4 for all these operations: addition, multiplication, exponentiation, tetration, pentation!? imma go keep being insane. thxbai
- 4y ago
- raverbashing 4y ago> Going by the semantics of the opcodes themselves, This is RISC wishful thinking that makes it sound like BL is not CALL. Maybe in ARM 1 it wasn't, it is now.
- zerohp 4y ago> Going by the semantics of the opcodes themselves, there's no reason why RET should be treated differently from a simple indirect jump here. RET is documented as a subroutine return hint in the official ARM architecture reference manual. That's the only reason it has a distinct opcode from BR. BL is also documented as a subroutine call hint.
- classichasclass 4y agoI agree, because this semantic difference between br x30 and ret doesn't exist in many other RISCs which also mostly run C-like languages. Power just has blr, and MIPS jr $ra, for example. This feels more like a well-intentioned footgun in the ISA.
- shadowofneptune 4y agoThe potential population for that footgun is compiler writers, who should safely fall into 'know what they are doing' territory.