13 ms·
Apple M1 Assembly Language Hello World
- nneonneo 6y agoThe author appears to have conflated "Linux" with the "kernel" - for example, "When calling a Linux service the function number goes in X16 rather than X8.", "X16 - linux function number", "Call linux to output the string", etc. in the macOS assembly code. Linux refers specifically to the operating system kernel that's used on Linux systems; the macOS kernel is called Mach. Technically, the non-negative-numbered system calls on macOS are those derived from BSD; macOS additionally has another (negative-numbered) set of system calls for its Mach microkernel.
- athrun 6y agoYes I think he's getting some of terminology wrong. Saying that "Linux [is] based on Unix" is not really accurate either. But ultimately this is just nitpicking. It's great that he's sharing what he learned.
- derefr 6y ago> Saying that "Linux [is] based on Unix" is not really accurate either. The context of the author's statement was syscall ABIs. And Linux's original (x86) syscall ABI is based on [a snapshot of] the syscall ABI from a Unix (4.2BSD, I think.) Loosely based, mind you — 4.2BSD wasn't targeting a 32-bit architecture, let alone x86, so the registers et al weren't the same. But the syscall numbers match up, and the number and order of registers used have direct parallels. Compare and contrast: • FreeBSD ABI (direct descendant of 4.2BSD): https://github.com/freebsd/freebsd-src/blob/master/sys/kern/syscalls.master https://github.com/freebsd/freebsd-src/blob/master/sys/kern/... • Linux x86 ABI: https://chromium.googlesource.com/chromiumos/docs/+/master/constants/syscalls.md#x86-32_bit https://chromium.googlesource.com/chromiumos/docs/+/master/c... Everything lines up until you get to the OS-proprietary stuff.
- spijdar 6y agoSo, a few things. 4.2BSD was targeting VAX, which is most assuredly a 32 bit machine. Further, the common denominator you're seeing goes back much, much further than BSD itself. Behold! "Version 1" of UNIX, written in just 14 files of PDP-11 assembly. And if you look at u1.s, you'll see the sysent routine which handles syscalls. The numbering should be familiar :-) https://github.com/dspinellis/unix-history-repo/blob/Research-V1-Snapshot-Development/u1.s#L35 https://github.com/dspinellis/unix-history-repo/blob/Researc... If you go to 4.2BSD, you'll see many of these syscalls labeled as "old", e.g. "old pause", "old wait", "old break". https://github.com/dspinellis/unix-history-repo/blob/BSD-4_2-Snapshot-Development/usr/src/sys/sys/syscalls.c#L6 https://github.com/dspinellis/unix-history-repo/blob/BSD-4_2... You can go a little further back to the original PDP7 assembly, but I don't believe the "kernel" actually ran on a separate "process" at all, and the userland simply linked into kernel symbols and called "read", "readdir" etc directly, hence why even later unix documentation tends to call these "routines" or "library calls".
- derefr 6y ago> Further, the common denominator you're seeing goes back much, much further than BSD itself. Interesting, I didn’t know that! There’s definitely an essay to be written here about the evolution of this “Unix system-call convention” over the decades, going into how this table of calls survived each transition and port mostly-unscathed. (Given that we’re throwing actual source syscall tables back and forth as proofs, there’s definitely some narrativizing to be done here, The Old New Thing-style.) I assume the syscalls weren’t retained in ports with the goal of “binary compatibility”, given that these descendant Unix ports were on different architectures that couldn’t literally exec(2) binaries from their ancestor. Guesses: • Toolchain compatibility? • Some shared cross-compiling assembler/linker that nevertheless had hard-coded syscalls? • (The perhaps never-achieved-in-practice goal of) emulation-assisted descendant cross-compatibility, ala z/OS? • The existence of hybrid/transitional minicomputer generations, that had application processors for both the old and new architectures (or application processors that could execute on both ISAs!), such that at least some of the systems being ported to could exec(2) the ancestor’s binaries straight from tape? • Or just the expectation that, despite Unix and C being so intertwined, there were still enough people writing ASM for these machines — using a non-macro assembler — who had developed reflex-memory for the existing syscall numbers, that it would be a bad idea to change the table out from under them? > the userland simply linked into kernel symbols and called "read", "readdir" etc directly So basically, the original Unix was a DOS, rather than a kernel/supervisor. Was that just because the PDP7 didn’t have virtual memory management, or was it a conscious design decision that was later reversed?
- Blikkentrekker 6y agoThe term is so often conflated because it seems as though what is wanted be a technical definition of a concept that is not technical, but social. “Linux” the kernel is technical, but that other thing for which “Linux” is often used has no technical barriers to it, and is entirely a social thing — something is “Linux” when it declares itself part of the tribe, and is so accepted. SteamOS as a marketing strategy always used the phrase “Linux and SteamOS” and was thus successful in moving itself outside of the tribe, for example. There is no technical definition of “an operating system” and there are no technical reasons for what “operating systems" are and aren't “Linux”; it is purely based on social cohesion.
- userbinator 6y agoNothing a little s/Linux/macOS/ won't fix... but I sure hope the book he is promoting is less confusing than that. To his credit, he does say "in MacOS it's X, in Linux it's Y". (I'm a long-time x86 programmer, who has done a very little of ARM but finds it boring since you can't beat a compiler as easily as you can on x86... and yet instinctively felt unease at moving 0 into a register. ;-)
- nneonneo 6y agoYeah, there's a bunch of stuff in AArch64 that takes getting used to - for example, fixed-length instructions (everything is a 32-bit instruction!), tons of sanely-named registers (X0~X30, W0~W30), shorter literals (all literals need to fit in a 32-bit instruction), and so much more. Plus, it's a total break from AArch32 - I'm pretty comfortable with 32-bit ARM code, and yet 64-bit ARM code still looks pretty unusual to me.
- lovelyviking 6y ago> I'm pretty comfortable with 32-bit ARM Could you give a direction to a good place for learning it ? Where is better to start? I wish to write hello world without linker or assembler for linux.
- saagarjha 6y agoTry https://azeria-labs.com/writing-arm-assembly-part-1/ https://azeria-labs.com/writing-arm-assembly-part-1/.
- madog 6y agoI found Azeria Labs [0] to be a pretty good introduction for learning ARM assembly basics. Although the main focus is ARM exploitation, it's still a good primer [0]: https://azeria-labs.com/writing-arm-assembly-part-1/ https://azeria-labs.com/writing-arm-assembly-part-1/
- lovelyviking 6y agoWhat do you mean by “beat the compiler” ? like running code without it nor linker ? Looks like I have a mood to try that now
- bananabreakfast 6y agoSmall nitpick, but the macOS kernel is actually called XNU which is a hybrid of Mach and FreeBSD.
- reaperducer 6y agothe macOS kernel is actually called XNU Since you seem knowledgeable about these things, how does one pronounce "XNU?" Is it "ex-en-you" or "zznoo" or "she-no" or something else?
- glhaynes 6y agoFairly certain I've heard it as "ex-noo". But it could just be that I've said it in my head that way forever!
- joshspankit 6y agoZA-na-doo
- deleted 6y ago[deleted]
- coldtea 6y ago>The author appears to have conflated "Linux" with the "kernel" I don't think he conflated it (in his mind). In the post he's quite clear which is which. He probably just reused an ARM Linux example code (probably from his book), made some changes required for macOS ARM, but forgot to change the name of the kernel.
- raxxorrax 6y agoNot yet given the M1 or MacOS a spin, but the syscall adresses should be available under /sys/syscall.h according to an answer on stack exchange. Cannot find a table on the net and syscall adresses could be subject to changes.
- saagarjha 6y agohttps://opensource.apple.com/source/xnu/xnu-7195.50.7.100.1/bsd/kern/syscalls.master.auto.html https://opensource.apple.com/source/xnu/xnu-7195.50.7.100.1/...
- wtsnz 6y agoSidenote: What a cool personal blog. Started in 2009 and still going strong. Inspirational!
- userbinator 6y agoClass AMSupportURLConnectionDelegate is implemented in both ?? (0x1edb5b8f0) and ?? (0x122dd02b8). One of the two will be used. Which one is undefined. It's both amusing and sad to see even simple command-line tools spewing those warnings which are most commonly found in the system log. I don't have a system to check at the moment, but I seem to recall using xcrun a long time ago and it didn't do that. I guess it is as they say, "beauty is only skin deep"...
- marcan_42 6y agoI see your userland M1 assembly language hello world and raise you a bare-metal M1 kernel assembly language hello world :-) https://github.com/AsahiLinux/m1n1/blob/main/src/start.S https://github.com/AsahiLinux/m1n1/blob/main/src/start.S The first few lines of that will print 'm1n1' to the serial port as it initializes other things and eventually jumps to C code.
- nneonneo 6y agoThis looks basically like assembly code for practically any embedded system - reading/writing hardware registers to interact with UARTs. Nice job (& best of luck with the Asahi Linux project - very exciting!)
- marcan_42 6y agoWell, the Apple M1 is an embedded SoC after all :) Thank you!
- yitchelle 6y agoThat is the most exciting part of the M1 for me. What do you think will be the first embedded application for the M1, assuming Apple allows it to be purchased by 3rd parties? Autonomous driving, or robotics, or other..??
- saagarjha 6y agoApple won't allow it to be purchased by third parties.
- magnusmundus 6y agoNot quite as exciting as your examples, but I guess an evolution of Apple CarPlay might be likely.
- WanderPanda 6y agoYes they need to capture the car UI part as legacy automakers will not be able to build an ecosystem for that.
- megablast 6y agoLooking forward to the updates benchmarks on hello world programs.
- cylinder714 6y agoMany thanks for any assembly-language posts, but can we agree that a FizzBuzz implementation would be more interesting? ;-)
- saagarjha 6y ago#include <sys/syscall.h> modulo: udiv x9, x0, x1 msub x0, x9, x1, x0 ret print_fizz: mov x0, #0 adr x1, fizz mov x2, #4 mov x16, SYS_write svc #0x80 ret print_buzz: mov x0, #0 adr x1, buzz mov x2, #4 mov x16, SYS_write svc #0x80 ret print_number: mov x9, x0 mov x10, #10 udiv x11, x0, x10 mov x0, #0 adr x1, number_table add x1, x1, x11 mov x2, #1 svc #0x80 msub x11, x11, x10, x9 mov x0, #0 adr x1, number_table add x1, x1, x11 mov x2, #1 svc #80 ret print_newline: mov x0, #0 adr x1, newline mov x2, #1 mov x16, SYS_write svc #0x80 ret .globl _main .align 2 _main: mov x19, #0 loop: mov x20, #0 mov x0, x19 mov x1, #3 bl modulo cmp x0, 0 cset x20, eq b.ne not_3 bl print_fizz not_3: mov x0, x19 mov x1, #5 bl modulo cmp x0, 0 cset x21, eq b.ne not_5 bl print_buzz not_5: orr x20, x20, x21 cmp x20, #0 b.ne divisible mov x0, x19 bl print_number divisible: bl print_newline add x19, x19, #1 cmp x19, #100 b.le loop mov x16, #0 svc #0x80 fizz: .ascii "Fizz" buzz: .ascii "Buzz" number_table: .ascii "0123456789" newline: .ascii "\n"
- ffhhj 6y agoNice tutorial! In the 90's got really excited with x86, trying all the DOS/BIOS interrupts and CGA/EGA/VGA/VESA programming. Got a computer lab evacuated with a fake "computer virus" Mcafee AV couldn't detect, a resident program that would display a snake moving on the screen. So much fun. It would be great to get a book on the new M1 and MacOS services, specially if it includes some guide to their Neural Engine and GPU programming.
- rckoepke 6y agoI believe the first attempts at doing something like this were back in July of 2020. https://github.com/below/HelloSilicon https://github.com/below/HelloSilicon : Someone with an early Developer Transition Kit (pre-M1 release) worked to convert the code from "Programming with 64-Bit ARM Assembly Language" to the M1's syntax. Porting additional textbooks to M1's (ARMv8) syntax could help a lot in terms of making assembly accessible to more people. I believe there's a lot of value in learning it on a particularly popular real-world platform like x86 or M1 - where it may directly translate to reverse engineering userspace applications without having to then learn another assembler language as you might if you started with, say, RISC-V. Truthfully I think someone should really add the M1 syntax to the new version of VisUAL: https://github.com/scc416/Visual2 https://github.com/scc416/Visual2 . I've been intending to work on that to go along with porting Bob Plantz' "Introduction to Computer Organization: ARM Assembly Language Using the Raspberry Pi" to ARMv8 for the same educational purpose but haven't quite found the time to really dive in. A tool like VisUAL2 could help a lot of people learn this even if they don't own an M1 themselves. Very tangentially to all of this, I'd like to showcase https://github.com/cornell-brg/pydgin https://github.com/cornell-brg/pydgin , which is a flexible toolkit for simulating ISA's in Python, and was used to help validate the first version of VisUAL during its own development.
- KMag 6y agoI'm currently auditing the Stanford compilers course online. I understand the inertia in introductory courses, and the simplicity of MIPS, but hopefully the course is eventually ported to Aarch64 or RISC-V. Presumably a spim-like simulator for Aarch64 or RISC-V is the biggest missing component. 6.004 was one of my favorite classes at MIT, and the DEC Alpha AXP was a fine architecture to simplify for the pedagogical Beta architecture. However, I'm glad to hear they've moved to RISC-V. I presume they'd been avoiding porting the course to Aarch64 (or a simplified version thereof) due to intellectual property issues.
- remexre 6y agoFor AArch32, not AArch64, but there's https://salmanarif.bitbucket.io/visual/index.html https://salmanarif.bitbucket.io/visual/index.html as a SPIM replacement.
- wk_end 6y agoI low-key think ARM64 might be the nicest ISA around. It reminds me of a modern, RISCier 68k.
- bni 6y agoHow is it possible that a RISC ISA can be nice to program for in Assembly? Wasn't PPC horrible for this for example? Wasn't the point of CISC to make it easier for Assembly programmers (made sense when x86 and 68k was defined)
- socialdemocrat 6y agoI started with 68k on Amiga which made me love assembly. Then got exposed to x86 which I hated with a passion. Keep in mind I was quite young then and x86 had so many oddities and inconsistencies that it made it really hard for a beginner. In Uni when making a compiler I choose to target PPC as I figured anything was better than x86. I kind of regretted it. I was new to RISC and the heavy use of registers threw me off. E.g. using registers for all function arguments except when there where too many. Using a register for return address but then needing to use stack anyway of call stack got too deep. How every address had to be loaded in two steps. On top of that PPC did not offer a smaller and simpler instruction set. It was quite large from what I remember. So getting something akin to feeling like 68k was a dream that got popped. But I would say RISC-V gives me a bit of the 68k feeling as things are kept really small and beginner friendly.
- adrian_b 6y agoMost RISC ISA's, including PPC and ARM are much nicer to program in assembly than the Intel/AMD x86 ISA. The only exception is with RISC ISA's that are far too simple, e.g. the base RISC-V, so they lack many common instructions or addressing modes, forcing the programmer to use sequences of instructions instead of the single instructions of other, more complete, ISA's. Most RISC ISA's are relatively orthogonal, in the sense that you have to memorize only a small list of things, e.g. of kinds of instructions, kinds of registers, kinds of addressing modes, and then you can express any program as a combination of those few elements. On the other hand, for the Intel/AMD x86 ISA, to be able to write an optimized program, you need to know a huge list of things. The original x86 registers are all different having different features, there are not even 2 identical registers, so to find the optimal register allocation is difficult. Even the extra 8 registers added by the AMD 64-bit extension are not identical, but they are divided in 3 or 4 groups with different properties regarding the encoding of the instructions, which result in different program izes. Among the registers added by more recent ISA extensions, especially the SIMD extensions, SSE/AVX, there are finally groups of identical registers that are easier to allocate, but these must be used together with the integer registers, which are needed at least for addressing. It is difficult to determine the total number of different instructions in the x86 ISA, but they are much more than a thousand, many of which are obsolete and which should be avoided. Even when compared to older normal ISA's, which are not so complex like x86, e.g. Motorola 68k or IBM mainframes, many RISC ISA's, including POWER & ARMv8, are still an easier target for assembly programming, when you attempt to write an optimized program. When the performance of the program was not important, assembly programs for some traditional CISC ISA's e.g. Motorola 68k or DEC VAX could be simpler than for RISC ISA's, due to using fewer more complex instructions, but those instructions typically had slower implementations, so an optimized program might have been forced to avoid them. Knowing the very variable timing of the CISC instructions and accounting for that in order to write an optimized program, made optimized assembly programming more difficult than for RISC CPUs.
- high_priest 6y agoDoes this blog have RSS feed?
- banana_giraffe 6y agohttps://smist08.wordpress.com/feed/ https://smist08.wordpress.com/feed/
- pedro1976 6y agoJust in case you are looking for an rss feed, you can always give rss-proxy [0] a try. [0] https://github.com/damoeb/rss-proxy https://github.com/damoeb/rss-proxy
- SecurityLagoon 6y agoThanks for this. Seems like another great service to run on my pi. Getting a little tangential but how well do you find it works in general? I imagine it is only truly reliable with SSG type websites.
- pedro1976 6y agoI think it works fine, it makes a website of choice accessible for your Reader. Since it just maps HTML -> RSS the entry description can be sparse, but then your reader should grab the fulltext of the referenced site. There is basic JavaScript support [0], but this is rather new and not sure how well it works. I tested it with a couple of webapps and craigslist and it worked ok. If you host it on your pi, you can just run the big js-image. [0] https://github.com/damoeb/rss-proxy#javascript-support https://github.com/damoeb/rss-proxy#javascript-support
- mhh__ 6y agoIs mentioning M1 just guaranteed karma now?
- 1_player 6y agoApparently yes, I doubt a generic ARM assembly tutorial that writes "Hello, world" to stdout would get almost 400 points. It's not like M1 is a totally new architecture. Conversely, saying anything bad about M1 or anything related to it, like this article, is a good way to piss people off on this forum.
- andy_threos_io 6y agoAnd it's a total lame, bad assembly code. Lacks any kind of material knowledge or practice.
- Hackbraten 6y agoYour comments in this thread have been belittling and unhelpful. We don’t like that here at all. Please stop.
- andy_threos_io 6y agoThe example program presented is not a good solution, it goes against all the rules and recommendations. Your program (made this way) will NOT WORK at all in the near future, ex when the system call number is changed (it's happened).
- fellellor 6y agoSorry for the noob question, but here it is anyway. What kind of projects and career paths does this kind of programming enable?
- mhh__ 6y agoMy snarky answer to that is that I genuinely think any software engineer should be able to read and write assembly, even if purely to remind them just how far up the stack we usually operate at. If we take "this" to mean low-level programming, it can open doors into anywhere from reverse engineering, to OS development etc (or more boringly, writing toolchains). Some parts of compilers rely heavily on having real problem solving skills at this level - for example most books won't teach you how to integrate ELF into your compiler, let alone DWARF. I am hesitant to say performance optimization because many people have an idea of a CPU just executing one instruction at a time rather than the heavily pipelined and out of order monsters we have today. To become Anger Fog, you must first invent the CPU-niverse. (Agner is a almost mythical figure for many compiler and microarchitecture folks)
- Intermernet 6y agoI'm somewhat delighted that your first mention of Agner got auto-corrected to "Anger Fog" :-) I like to think that Agner would appreciate that!
- mhh__ 6y agoAgner Fog becomes Anger Fog when measures Intel Compilers on AMD...
- 2sk21 6y agoWhen I was a CS PhD student, I spent the summer of 1988 writing assembly code on a parallel machine (Encore Multimax) that was running a variant of Unix. I wrote the runtime system (a virtual machine actually) for a parallel programming language. Haven't used assembly language since then!
- fellellor 6y ago
- nsonha 6y agoisn't this just ARM assembly, is there such a thing called M1 assembly (genuine question)?
- fulafel 6y agoYes. I guess the author is (successfully) going along with the general popular hype that paints it as a new architecture.
- _ph_ 6y agoI am certain, he isn't. The book also has a special chapter for iOS, explaining the changes needed to what is shown in the book to run the code with the iOS operation system and tooling.
- _ph_ 6y agoYes, it is just ARM assembly. But the tooling (Clang vs GCC) and the operation system differ, so the code from the book won't work without changes on a Mac. The book already contains a chapter for using the ARM assembly on iOS, because you have to do the same changes vs. running the sample code on Linux. The linked blog posts lists these changes. They are not large, but significant enough, that someone learning how to code in ARM assembly will be thankful for adjusted example code.
- khanan 6y agoFunny how you can replace MacOS with FreeBSD in this article and it's the same thing. OH WAIT! eyeroll
- astrange 6y agoIt isn't the same thing since macOS doesn't use ELF.
- musicale 6y agoThe macOS kernel is Xnu or Darwin, not Linux. Try this on macOS vs. Linux: uname -a macOS system call numbers and interfaces are derived from (Free)BSD.
- 95014_refugee 6y agoSome, a long time ago. Others come from Mach, and there’s ~20 years worth of independent work and divergence between then and now...
- andy_threos_io 6y agoPrime example of how you NOT write assembly code, no matter what cpu are you using. First thing first, macro assemblers have been around for at least 40+ years. So don't use numbers in your code. use macros or defines. Ex. you can compile with gcc flags -x assembler-with-cpp and you can have nice defines in your code. Second on a decent OS even in assembly link against operating system call library, no matter what. The system call numbers can change. so use symbols. Also don't write the string length in the code, it's total lame. and use .asciz not .ascii and define your function symbols to function gcc: .global my_func_name .type my_func_name, function my_func_name: if you use C preprocessor for assembly compile just make an include like this: #define _FUNCTION(A) .global A ;\ .type A, function EDIT: For the not well informed HN readers: https://developer.apple.com/library/archive/qa/qa1118/_index.html https://developer.apple.com/library/archive/qa/qa1118/_index... "Apple does not support statically linked binaries on Mac OS X. A statically linked binary assumes binary compatibility at the kernel system call interface, and we do not make any guarantees on that front. Rather, we strive to ensure binary compatibility in each dynamically linked system library and framework." So DON'T write direct system call numbers in your code!
- saagarjha 6y agoUsing .ascii is fine; the string is going to a syscall–not libc. And using the standard library in a program like this is kind of beside the point…
- andy_threos_io 6y agonope, every decent OS has system call library you can write like: bl _Open ; whatever syscall name you have and the linker or the OS dynamic linker will get you the proper system call code, with the numbers. BTW even in our small threos.io os we use dynamic linked system calls.
- saagarjha 6y agoI know how dynamic linkers work and am not arguing against using the libc interface in general. I'm saying that this program intentionally doesn't use those interfaces.
- tcmb 6y agoI'm confused, why is this showing a Linux assembly file and then the Makefile and build instructions for macOS?
- mark_l_watson 6y agoThat is so cool. I think that I am going to but his book. It has been years since I wrote assembly code. I will time box an hour or two this weekend to play with this, but just for fun.
- mark_l_watson 6y agoTo late to edit, I meant “buy his book”. Typo.
- deleted 6y ago[deleted]
- _ph_ 6y agoThis is very nice to see. I recently got this book to get into ARM assembly language, with the long-term goal of using it on an M1 Mac, I barely could resist on getting a MacBook Air immediately :). Right now, I am setting up a Raspberry Pi 400 for first experiments. Unfortunately, it still ships with a 32 bit install. While the differences between a M1 Mac and Linux-ARM machine are small, it is very nice to have the examples ported and especially tested on the Mac, as especially during learning a small mistake can be difficult to be noticed.
- bobrippling 6y agoA few extra things to note for enthusiasts: > The MacOS linker/loader doesn’t like doing relocations, so you need to use the ADR rather than LDR instruction to load addresses. You could use ADR in Linux and if you do this it will work in both. I think this is more the assembler - the linker's perfectly happy performing relocations > In MacOS you need to link in the System library even if you don’t make a system call from it or you get a linker error. This sample Hello World program uses software interrupts to make the system calls rather than the API in the System library and so shouldn’t need to link to it. You can get around this by creating a statically linked executable, which requires a bit of wrangling, but is supported (and perhaps handy if you're going on to write a kernel). > In MacOS the default entry point is _main whereas in Linux it is _start. This is changed via a command line argument to the linker. In macOS the default entry point is start (linux is still _start), the C runtime still needs to be setup - the kernel can't jump a program straight to main.
- 95014_refugee 6y ago> I think this is more the assembler - the linker's perfectly happy performing relocations No, this is “all executables are PIE because applying relocations to shared code is stupid and inefficient in the presence of ASLR”.
- 95014_refugee 6y ago> In macOS the default entry point is start (linux is still _start), the C runtime still needs to be setup - the kernel can't jump a program straight to main. macOS has no interest in what the symbol is called; it pulls the initial PC from the LC_MAIN command in the Mach-o header. ld64 (the linker) will by default populate that load command by looking up “_start”, but that’s a separate thing...