4 ms·
I am also pursuing something similar (complementary?) to this (recently started writing a Rust-based "distributed OS" that manages ML resources in a network of
by softwarewright 12d ago
I am also pursuing something similar (complementary?) to this (recently started writing a Rust-based "distributed OS" that manages ML resources in a network of heterogenous systems: varied cores, system RAM, GPU VRAM, I/O). So my focus is not so much distributed agents, but more distributed inference (and fine tuning) that would benefit agents, distributed or not. I may blog about my work soon.
Your posted link is helpful, thanks.
- kgcgfva 12d agoThanks, softwarewright; would love to hear more about your work. I've been implementing small-model inference in Zug inside WunderOS, but it's not done yet, so nothing public just yet.
- softwarewright 12d agoI'm in a research/prototyping phase with no truly verified results yet, but I have all the hardware I need to do actual testing. I have an OS that boots in a VM, it just isn't verified to have practical value yet. It could be an AI coding agent's fever dream until I use it anger. So far Apple Silicon only, but most of my systems run Linux and most of my GPUs are NVIDIA, so moving development from Mac to Linux soon. It does have tests, demos, docs. Iterating on it daily. https://github.com/sw-ml-study/sw-os-ml https://github.com/sw-ml-study/sw-os-ml I also have been implementing small model inference: https://github.com/sw-ml-study/moe-microscope https://github.com/sw-ml-study/moe-microscope I should clarify that by Apple Silicon I mean it boots Rust no_std on ARM. Does not use GPU yet. Plan is to use Rust without CUDA libraries. This project is more likely to use an NPU on a ARM dev board before it can use an NVIDIA GPU, and might never be able to use Apple GPUs. Goal: run on no-longer-supported by CUDA GPUs.
- kgcgfva 11d agoGood stuff here, much of which resonates with me. FWIW, my notion of an OS for agents isn't as close to the metal yet, ie, my view of agent lifecycle is that they look a lot like WhatsApp or Discord traffic. So, based on that DEEP analysis, I decided to build on BEAM/OTP for the control plane: basically, Elixir for the UX and Gleam for everything else. And the data plane is pure Zig: data sidecar into BEAM/OTP via NIF. So that gives me a certain freedom for deploy: bare metal, containers, VMs, even K8S. And since there are some tools for pickling all that into a single binary, and running BEAM/OTP on u-kernel sorts of things, I can get all the way to the metal in the way that you are. Whether or when that happens remains to be seen, etc. Thanks for sharing! See https://pentad.ai/PLRN https://pentad.ai/PLRN for more about what I'm up to.
- softwarewright 11d agoInteresting link/content and it seems complimentary. One thing that is missing from both of our approaches is the ability re-train (fine-tune) coding models "overnight" so that they can "learn" from the prior day and changes since their training cutoff date. I have found some things I can do to improve my work based on this, thanks. _Pentad idea_ -/- _MLOS relevance_ -/- _Action_ Closed autonomic loops -/- Very high -/- Adopt architecture vocabulary Deterministic replay -/- Very high -/- Strengthen event/replay contract Model minimalism -/- Very high. -/- Extend later to compute-placement ladder Durable vs active population -/- High -/- Define registered vs resident capacity metrics Standing queries -/- High -/- Future policy/watch abstraction Provenance by construction -/- High -/- Record policy decision causality No model/NLP in hot path -/- High -/- State explicitly as invariant
- kgcgfva 10d agoOne thing that falls out of Model Minimalism is adding native WunderOS model hosting, which I've been working on this week, natively in Zig, NIF'd into BEAM/OTP. IMO vertical integration in AI infra is underrated; by adding model serving I can exploit a range of optimizations that 'best of breed'/glue code architecture makes harder. This week's example: native semantic entropy implementation -- following Spanda -- in Zig, such that hallucination detection at K=5 (batch size) is 650us per turn, i.e., in the noise. I've spec'd how to do QLoRA, too, but it's unclear when or if I'll implement it, not least because it's not clear that I should bother given the training data integration issues. I'm glad someone else is thinking about this stuff!