6 ms·
Mu: A minimal hobbyist computing stack
- Quequau 7y agoThis dudes stuff has been posted here recently and it's worth it go read what he's said about his work. https://news.ycombinator.com/item?id=21242190 https://news.ycombinator.com/item?id=21242190
- minxomat 7y agoAnnouncement link from that thread: https://lists.tildeverse.org/hyperkitty/list/tildeclub@lists.tildeverse.org/message/Z7EBQ2ZCBIQ7YMA7Q3RUJWWB4LBIFS3M https://lists.tildeverse.org/hyperkitty/list/tildeclub@lists...
- akkartik 7y agoThanks! I took the opportunity to try to clarify those questions in this version.
- eterps 7y agoInteresting project. So far the only desktop OS + compiler I know that can be reasonably understood in its entirety is http://www.projectoberon.com http://www.projectoberon.com
- snvzz 7y agoI'd add TempleOS and its HolyC to that list.
- grawprog 7y agoTempleOS may be understandable, but trying to actually make and do anything with it's a completely different story.
- e12e 7y agoSome interesting parallels to choices made in Pascal in the second level language, too. I'm not convinced that this approach is the best - the first level language is a bit like "literal programming" machine code (not assembler) - but I think a level of simple substitution would help a bit (think m4 as an assembler) - allowing at least mnemonics. The insitence on numbers seem somewhat arbitrary as they still need to be translated from strings (letter sequences) to actual integers[e] - I don't think a table/variable look up would be too onerous. On the whole this feels a bit more convoluted than building a Forth as level one - but no matter what I applaud the effort. [e] assuming program source code is text, not binary? If it's binary I suppose "filtering out" comments might be an approach as opposed to "translating" the source.
- akkartik 7y agoThis is a very acute critique, so it seems like the perfect place to elaborate on why I chose numbers. We tend to think of Assembly programming as using mnemonics, but really the names of instructions in x86 are more than that: there are 14 different kinds of add instructions (https://c9x.me/x86/html/file_module_x86_id_5.html https://c9x.me/x86/html/file_module_x86_id_5.html) If a single name can map to this many distinct cases, it's not a mnemonic anymore, it's a little more than that. When a single word can do so many different things, and when the underlying instruction set has gaps, it becomes hard to create good error messages. Oh you added two immediate values together. You can't do that. But what can you do? Any impedance mismatch between external notation and target format quickly adds to the amount of code needed to generate good, clear error messages. Since my goal was to build SubX in itself, it seemed like a good trade-off to force people to live closer to the target language. That makes the error messages easier to write. There's really only one case to address at any point. (The alternative would be to give each opcode a distinct mnemonic. Does that seem useful? I didn't initially think it was, but I'm open to trying it.)
- e12e 7y agoMaybe I skimmed, but I didn't note anything on this (anything this clear at any rate) in the text. It certainly is a justification for exposing the different op codes. I still think symbolic names for symbolic numbers make sense,though. And I'm not sure a table in one place looses out on clarity. I do think it's good to re-think assembly notation (que gas/masm/nasm/x86/6800 notation wars). I guess I feel at the level of level 1 - there is some symbolic manipulation and meta programming already (eg loops, labels for computed jumps) - and that "dumb" replacement might even turn out to me more clear. Comments go away (replace with nothing), jump-to-labels become jump-to-offset, opcode mnemonic become binary numbers, numbers become binary numbers etc. I confess I wonder at the choice of 32bit as I find amd64 to be both clearer, more sane (better? Calling convention) - and it should be more powerful (eg more registers) - faster and allow addressing more memory (if/when that becomes and issue - say mapping an entiere 6tb drive into contiguous memory?).
- ntoll 7y agoNote to self (as the author of the "Mu" editor, https://codewith.mu/ https://codewith.mu/) - two letter project names are just asking for name collisions. ;-)
- svxml 7y agoAnd also https://www.djcbsoftware.nl/code/mu/mu4e.html https://www.djcbsoftware.nl/code/mu/mu4e.html the mail client for emacs :)
- perfunctory 7y ago> I'd like Mu to be a stack a single person can hold in their head all at once, and modify in radical ways. This reminds me of Alan Kay's STEPS project - an attempt to reinvent personal computer in 20K lines of code [0]. Always happy to see efforts in this direction. [0] http://www.vpri.org/pdf/tr2011004_steps11.pdf http://www.vpri.org/pdf/tr2011004_steps11.pdf
- akkartik 7y agoAuthor here; I just found this thread. Feel free to ask me anything.
- radEd 7y agoWould you recommend this for someone who hasn't yet officially learned operating systems/compilers, but is looking to break into the realm?
- akkartik 7y agoI hope so! It's new so you may find breakage, but I'll try to be very responsive answering questions. Mu doesn't have its own OS at the moment. It runs on Linux and on Soso (https://github.com/ozkl/soso https://github.com/ozkl/soso). I don't know either well, but hopefully I'll at least be a good study partner. And you can definitely ask me questions about the language side and the interface with the OS. Mu doesn't have every feature of modern compilers and OSs, but it should hopefully give you a view of something simple working from end to end that you can then carry with you when making sense of more complex systems.
- radEd 7y ago"Mu doesn't have every feature of modern compilers and OSs, but it should hopefully give you a view of something simple working from end to end that you can then carry with you when making sense of more complex systems." Awesome, that's what I was hoping for: something that will lower the entry bar but still provide useful education for the greater world.
- kragen 7y agoI know you would like to make Mu open source, which is to say, free software. What are the problems standing in your way, and how can we help to solve them?
- bitminer 7y agoWhat's old is new again. Your thoughts on the benefits of language slightly higher level than assembler are interesting. Others have trod this path before. Bliss32 was the systems language for vax/vms back in the 1980s through 2000s. Most of vms was written in it. It was heavily influenced by PL/360 from the days when IBM 360 or 370 mainframes were mostly programmed in assembler. Both languages emitted very predictable assembler from their higher level constructs. They had no optimizations. But their objectives seemed to be focused on readability and predictable output, like your system. The source code for 99% of vms was published on microfiche and distributed with the binary tapes. Worth finding a copy and the accompanying text on vms internals and data structures. An excellent intro to nonTorvalds operating system architecture.
- carapace 7y agoI find this really inspiring! I've got a similar hobby project (that I've been neglecting recently...) but with some different design choices. I don't want to hijack the thread but I thought it might be interesting to compare and contrast. Instead of targeting x86 it uses Prof. Wirth's RISC CPU for Project Oberon. This is a really simple chip, there is a model function for it written in Oberon that is less than a page of code. It's also very well documented, and emulators exist in Java, C, Python, and Javascript, and there is a VHDL (or is it Verilog?) description which people have used to make FPGA-based workstations. For the basis notation I'm using a dialect of Joy, an extremely simple and elegant purely functional, stack-based, concatinative language. It's like a combination of the best parts of Forth and Lisp. https://en.wikipedia.org/wiki/Joy_(programming_language) https://en.wikipedia.org/wiki/Joy_(programming_language) (I'm pretty sure there's no simpler useful language than Joy.) To bridge Joy to the CPU I took the high road: a compiler written in Prolog. It would be possible to use Forth lore to bootstrap from the metal up, but I was gobsmacked by David Warren's "Logic Programming and Compiler Writing" ( https://news.ycombinator.com/item?id=17674859 https://news.ycombinator.com/item?id=17674859 ) and have spent the last year "moulting" from Python to Prolog programmer. Anyhow, the compiler so far is concise but the model it implements for Joy on the CPU is still very crude. Between life, work, and learning more Prolog I haven't had much time to improve it. The thing is, Prolog is a very simple language, the compiler is very simple and small, and Joy is also very very simple and small (especially if it's implemented in Prolog), and the underlying CPU is very simple and small. The tricky bit would be to, say, implement a Prolog interpreter in Joy. But if you do that you've closed the loop for self-hosting: Prolog-in-Joy, running on Joy-on-the-metal, implementing a compiler for Joy to the metal, bootstrapped through Prolog on the host. (It would be simple to convert the Joy compiler in Prolog to a Joy compiler in Joy using the Prolog interpreter in Joy and partial evaluation. FWIW.) For the OS I'm going to build a simple clone of OberonOS but using Joy rather than Oberon language. I made a model of this OS in Python and Tkinter to guide the reimplementation in Joy. This will also serve as a stress test for Joy: can you implement a whole (albeit simple) "OS" in a purely functional, stack-based, concatinative language? Maybe it sucks? To sum up, a simple text-based UI/OS inspired by and loosely imitating OberonOS, targeting a simple but capable 32-bit RISC chip, with Joy as the shell and glue language and Prolog as the under-the-hood powerhouse systems language. And the whole thing should fit in under a hundred pages of code.
- eterps 7y agoakkartik, I see the following 3 syntaxes for 'MOV ESP, EBP': 89/<- ebp 4/r32/esp # copy esp to ebp 89/<- 5/rm32/ebp 4/r32/esp # copy esp to ebp 89/<- 3/mod/direct 5/rm32/ebp 4/r32/esp # copy esp to ebp Only the last one contains all information needed to encode the instruction. And the first one is used in the factorial example. Can you explain why all the 3 syntaxes are in use?
- akkartik 7y agoYeah, the last one is the 'ground truth' that both translators understand. sigils.subx adds some syntax sugar so you can say 89/<- %ebp 4/r32/esp By my self-imposed rules I can't use this sugar until I build it, so there's little code that uses it at the moment. But any future Subx code will use it, such as the implementation of Level 2. You can also see it used in calls.subx, which is another later layer of syntax sugar. Could you point me at where you see your first two versions?
- eterps 7y agoIn http://akkartik.name/post/mu-2019-1 http://akkartik.name/post/mu-2019-1 if you search for the pattern '/89' you'll find the first two versions. The 3rd one is here: https://github.com/akkartik/mu/blob/master/apps/factorial.subx#L23 https://github.com/akkartik/mu/blob/master/apps/factorial.su... Can you explain in a sentence what rules are used to complete the encoding for the variant: > 89/<- %ebp 4/r32/esp
- akkartik 7y agoAh, I see. Yes, I mentioned in the post that I was dropping some magic numbers. Mu is intended to be read in a text editor, and I didn't want to scare people too much while reading the post in a browser. Slightly evil, I know. The concrete rules are illustrated at https://github.com/akkartik/mu/blob/469a1a9ace9b290658beb9d0799b58ce3bc2a4b4/apps/sigils.subx#L9 https://github.com/akkartik/mu/blob/469a1a9ace9b290658beb9d0.... In this case `%<reg>` always translates to `3/mod <reg-code>/rm32`.
- kbob 7y agoThis reminds me very much of 1970s Unix. Unix was also an research exercise into OS and userland minimalism and a new, streamlined language (yes, C) that created a new niche in the "like asm but better" space. In other words, it had nothing in common with today's Linux. (-:
- akkartik 7y agoIndeed, much of Mu's design was motivated by the question, "if we imitated the codesign of OS and language that led to Unix and C, given what we know now, what would we create?" This question doesn't seem to be often asked in recent decades, with languages and OSs evolving separately.