16 ms·
A CPU that runs entirely on GPU
- mrlonglong 7mo agoNow I've seen it all. Time to die.. (meant humourously)
- RagnarD 7mo agoBeing able to perform precise math in an LLM is important, glad to see this.
- jdjdndnzn 7mo agoJust want to point out this comment is highly ironic. This is all a computer does :P We need llms to be able to tap that not add the same functionality a layer above and MUCH less efficiently.
- celdon25 7mo ago> We need llms to be able to tap that not add the same functionality a layer above and MUCH less efficiently. Agents, tool-integrated reasoning, even chain of thought (limited, for some math) can address this.
- RagnarD 7mo agoYou're both completely missing the point. It's important that an LLM be able to perform exact arithmetic reliably without a tool call. Of course the underlying hardware does so extremely rapidly, that's not the point.
- 5o1ecist 7mo ago[flagged]
- koolala 7mo agoThat would be cool. A way to read cpu assembly bytecode and then think in it. It's slower than real cpu code obviously but still crazy fast for 'thinking' about it. They wouldn't need to actually simulate an entire program in a never ending hot loop like a real computer. Just a few loops would explain a lot about a process and calculate a lot of precise information.
- lorenzohess 7mo agoOut of curiosity, how much slower is this than an actual CPU?
- bastawhiz 7mo agoBased on addition and subtraction, 625000x slower or so than a 2.5ghz cpu
- medi8r 7mo agoSo it could run Doom?
- repelsteeltje 7mo agoYes: https://github.com/robertcprice/nCPU?tab=readme-ov-file#doom-raycaster-demo https://github.com/robertcprice/nCPU?tab=readme-ov-file#doom...
- binsquare 7mo agoCan we run doom inside of doom yet?
- afewquarks 7mo ago[dead]
- vee-kay 7mo ago[dead]
- throawayonthe 7mo agoYes: https://github.com/kgsws/doom-in-doom https://github.com/kgsws/doom-in-doom
- PowerElectronix 7mo ago
- sudo_cowsay 7mo ago"Multiplication is 12x faster than addition..." Wow. That's cool but what happens to the regular CPU?
- adrian_b 7mo agoThis CPU simulator does not attempt to achieve the maximum speed that could be obtained when simulating a CPU on a GPU. For that a completely different approach would be needed, e.g. by implementing something akin to qemu, where each CPU instruction would be translated into a graphic shader program. On many older GPUs, it is impossible or difficult to launch a graphic program from inside a graphic program (instead of from the CPU), but where this is possible one could obtain a CPU emulation that would be many orders of magnitude faster than what is demonstrated here. Instead of going for speed, the project demonstrates a simpler self-contained implementation based on the same kind of neural networks used for ML/AI, which might work even on an NPU, not only on a GPU. Because it uses inappropriate hardware execution units, the speed is modest and the speed ratios between different kinds of instructions are weird, but nonetheless this is an impressive achievement, i.e. simulating the complete Aarch64 ISA with such means.
- 5o1ecist 7mo ago[flagged]
- koolala 7mo agoIf its bindless and pre-compiled why not? What's a faster way?
- adrian_b 7mo agoYou could coalesce multiple instructions per shader, but even with a single CPU instruction (which would be translated to a sequence of GPU instructions), you could reach orders of magnitude greater speed than in this neural network implementation, by using the arithmetic-logic execution units of the GPU. Once translated, the shader programs would be reused. All this could be inserted in qemu, where a CPU is emulated by generating for each instruction a short program that is compiled and then the resulting executable functions are cached and executed during the interpretation of the program for the emulated CPU. In qemu, one could replace the native CPU compiler with a GPU compiler, either for CUDA or for a graphic shader language, depending on the target GPU. Then the compiled shaders could be loaded in the GPU memory, where, if the GPU is recent enough to support this feature, they could launch each other in execution. Eventually, one might be able to use a modified qemu running on the CPU to bootstrap a qemu + a shader compiler that have been translated to run on the GPU, so that the entire simulation of a CPU is done on the GPU.
- deleted 7mo ago[deleted]
- Surac 7mo agoWell GPU are just special purpous CPU.
- bmc7505 7mo agoAs foretold six years ago. [1] [1]: https://breandan.net/2020/06/30/graph-computation#roadmap https://breandan.net/2020/06/30/graph-computation#roadmap
- toolslive 7mo agohttps://en.wikipedia.org/wiki/Xeon_Phi#Knights_Landing https://en.wikipedia.org/wiki/Xeon_Phi#Knights_Landing ?
- nicman23 7mo agocan i run linux on a nvidia card though?
- micw 7mo agoLinux runs everywhere
- volemo 7mo agoExcept on my stupid iPad “Pro”. :(
- mghackerlady 7mo agoiirc theres an app on the app store that's basically a small alpine container
- volemo 7mo agoWell, there's iSH and a-Shell but they don't have GUI capability and are somewhat limited in other ways. There's also UTM, but without weird hacks you can only get SE version which is very slow.
- deleted 7mo ago[deleted]
- deep1283 7mo agoThis is a fun idea. What surprised me is the inversion where MUL ends up faster than ADD because the neural LUT removes sequential dependency while the adder still needs prefix stages.
- MadnessASAP 7mo agoYa know just today I was thinking around a way to compile a neural network down to assembly. Matching and replacing neural network structures with their closest machine code equivalent. This is way cooler though! Instead of efficiently running a neural network on a CPU, I can inefficiently run my CPU on neural network! With the work being done to make more powerful GPUs and ASICs I bet in a few years I'll be able to run a 486 at 100MHz(!!) with power consumption just under a megawatt! The mind boggles at the sort of computations this will unlock! Few more years and I'll even be able to realise the dream of self-hosting ChatGPT on my own neural network simulated CPU!
- andrewdb 7mo agoWhy do we call them GPUs these days? Most GPUs, sitting in racks in datacenters, aren't "processing graphics" anyhow.
- xeonmc 7mo agoGeneral Processing Units Gross-Parallelization Units Generative Procedure Units Gratuitously Profiteering Unscrupulously
- incognito124 7mo agoGreed Processing Units
- wartywhoa23 7mo agoThis is just brilliant!
- allreduce 7mo agoSometimes Gibberish Producing Units
- xeonmc 7mo agoGibberish Pipeline Units
- markhahn 7mo agoGeneral Parallel Units
- jgtrosh 7mo agoThe dedicated term GPGPU [0] didn't catch on. [0]: https://en.wikipedia.org/wiki/General-purpose_computing_on_graphics_processing_units https://en.wikipedia.org/wiki/General-purpose_computing_on_g...
- CompuHacker 7mo agoCPU = Compute GPU = Impute
- nomercy400 7mo agoI was taught years ago that MUL and ADD can be implemented in one or a few cycles. They can be the same complexity. What am I missing here? Also, is it possible to use the GPU's ADD/MUL implementation? It is what a GPU does best.
- volemo 7mo agoTo multiply two arbitrary numbers in a single cycle, you need to include dedicated hardware into your ALU, without it you have to combine several additions and logical shifts. As to why not use the ADD/MUL capabilities of the GPU itself, I guess it wasn’t in the spirit of the challenge. ;)
- bob1029 7mo agoA fun experiment but I wonder how many out there seriously think we could ever completely rid ourselves of the CPU. It seems to be a rising sentiment. The cost of communicating information through space is dealt with in fundamentally different ways here. On the CPU it is addressed directly. The actual latency is minimized as much as possible, usually by predicting the future in various ways and keeping the spatial extent of each device (core complex) as small as possible. The GPU hides latency with massive parallelism. That's why we can put them across relatively slow networks and still see excellent performance. Latency hiding cannot deal well in workloads that are branchy and serialized because you can only have one logical thread throughout. The CPU dominates this area because it doesn't cheat. It directly targets the objective. Making efficient, accurate control flow decisions tends to be more valuable than being able to process data in large volumes. It just happens that there are a few exceptions to this rule that are incredibly popular.
- volemo 7mo agoI see us not getting rid of CPU, but CPU and GPU being eventually consolidated in one system of heterogeneous computing units.
- jagged-chisel 7mo agoAgreed. Much like “RISC is gonna replace everything” - it didn’t. Because the CPU makers incorporated lessons from RISC into their designs. I can see the same happening to the CPU. It will just take on the appropriate functionality to keep all the compute in the same chip. It’s gonna take awhile because Nvidia et al like their moats.
- zozbot234 7mo ago> It will just take on the appropriate functionality to keep all the compute in the same chip. So, an iGPU/APU? Those exist already. Regardless, the most GPU-like CPU architecture in common use today is probably SPARC, with its 8-way SMT. Add per-thread vector SIMD compute to something like that, and you end up with something that has broadly similar performance constraints to an iGPU.
- throawayonthe 7mo agovery tangentially related is whatever vectorware et al are doing: https://www.vectorware.com/blog/ https://www.vectorware.com/blog/
- artemonster 7mo agoEvery clueless person who suggest that we move to GPUs entirely have zero idea how things work and basically are suggesting using lambos to plow fields and tractors to race in nascar
- madwolf 7mo agoBad comparison. Lambos are regularly plowing fields and they're quite good at it. https://www.lamborghini-tractors.com/en-eu/ https://www.lamborghini-tractors.com/en-eu/
- artemonster 7mo agoI remembered that labos used to make tractors after I posted the comment. Nice catch!
- jagged-chisel 7mo ago“A CPU that runs entirely on the GPU” I imagine a carefully crafted set of programming primitives used to build up the abstraction of a CPU… “Every ALU operation is a trained neural network.” Oh… oh. Fun. Just not the type of “interesting” I was hoping for.
- koolala 7mo agoIsn't it interesting it doesn't instantly crash from a precision error? That sounds carefully crafted to me.
- jagged-chisel 7mo agoInteresting, yes. Still not the kind of interesting I was expecting.
- robertcprice1 7mo agoPlease tell me what you had in mind so I can try something different!
- anthk 7mo agoBegin reimplementing a subleq/muxleq VM with GPU primitive commands: https://github.com/howerj/muxleq https://github.com/howerj/muxleq (it has both, muxleq (multiplexed subleq, which is the same but mux'ing instructions being much faster) and subleq. As you can see the implementation it's trivial. Once it's compiled, you can run eforth, altough I run a tweked one with floats and some beter commands, edit muxleq.fth, set the float to 1 in that file with this example: 1 constant opt.float The same with the classic do..loop structure from Forth, which is not enabled by default, just the weird for..next one from EForth: 1 constant opt.control and recompile: ./muxleq ./muxleq.dec < muxleq.fth > new.dec run: ./muxleq new.dec Once you have a new.dec image, you can just use that from now on.
- deleted 7mo ago[deleted]
- koolala 7mo agoExciting if an Ai that is helping in its own improvements finds this and incorporates it into its own architecture. Then it starts reading and running all the worlds binary and gains intelligence as a fully actualized "computer". Finally becoming both a master of language and of binary bits. Thinking in poetry and in pure precise numerical calculations.
- DonThomasitos 7mo agoI don‘t understand why you would train a NN for an operation like sqrt that the GPU supports in silicon.
- nine_k 7mo agoI see it as a practical joke or a fun hack, like CPUs implemented in the Game of Life, or in Minecraft.
- anthk 7mo agoI actually ran Sokoban under EForth running on top of subleq/muxleq with a VM interpreted under few lines of AWK.
- mihaitodor 7mo agoIt’s been done already. Have a look at Quest for Tetris: https://codegolf.stackexchange.com/questions/11880/build-a-working-game-of-tetris-in-conways-game-of-life https://codegolf.stackexchange.com/questions/11880/build-a-w...
- user____name 7mo agoSomeone needs to implement LLVMPipe to target this isa, then one can run software OpenGL emulation and call it "hardware accelerated".
- jagged-chisel 7mo agoThis causes me discomfort.
- yjftsjthsd-h 7mo agoSurely that would be hardware decelerated
- jleyank 7mo agoHow is this different than the (various?) efforts back then to build a machine based on the Intel i860? Didn’t work, although people gave it a good try.
- wartywhoa23 7mo agoOh these brave new ways to paraphrase the good old "fuck fuel economy"... Thank you, Mr. Do-because-I-can! Yours truly, - GPU company CEO, - Electric company CEO.
- robertcprice1 7mo agoHey everyone thank you taking a look at my project. This was purely just a “can I do it” type deal, but ultimately my goal is to make a running OS purely on GPU, or one composed of learned systems.
- lstevens14 7mo agoHi! I think that the idea is certainly a fun one. There is a long history of trying to make a good parallel operating system. I do not think that any of the projects succeeded though. This article is a good read if you are interested in that. I am not sure why the economics of parallel computer operating systems have not worked out so far. I think it most likely has to do with the operating systems that we have being good enough and familiar. [0] https://news.ycombinator.com/item?id=43440174 https://news.ycombinator.com/item?id=43440174
- activestore 7mo agoThe Blue Gene Active Storage project demonstrated compute in highly parallel “storage” where storage was HPC memory. It could work for the relationship between CPU and GPU, FPGA, etc. https://www.fz-juelich.de/en/jsc/downloads/slides/bgas-bof/bgas-bof-fitch/@@download/file https://www.fz-juelich.de/en/jsc/downloads/slides/bgas-bof/b...
- StilesCrisis 7mo agoI think it's curious that you're saying "on GPU" when you mean "using tensors." GPUs run compute shaders naturally and can trivially act like CPUs, just use CUDA. This is more akin to "a CPU on NPU" and your NPU happens to be a GPU.
- yjftsjthsd-h 7mo agoThis is hilarious and profoundly in the spirit of hacker news. Thanks for posting:)
- mghackerlady 7mo agoGNU/GPU
- taofor4 7mo agoWhat is the purpose of this project? I didn't get it. How will it be useful?
- jebarker 7mo ago> How will it be useful? Does it need to be?
- low_tech_punk 7mo agoSaw the DOOM raycast demo at bottom of page. Can't wait for someone to build a DOOM that runs entirely on GPU!
- jhuber6 7mo agoDepends entirely on your definition of 'entirely', but https://github.com/jhuber6/doomgeneric https://github.com/jhuber6/doomgeneric is pretty much a direct compilation of the DOOM C source for GPU compute. The CPU is necessary to read keyboard input and present frame data to the screen, but all the logic runs on the GPU.
- Nevermark 7mo agoTime to benchmark Doom. Now we know future genius models won't even need CPUs, just tensor/rectifier circuits. If they need a CPU, they will just imagine them. A low-bit model with adaptive sparse execution might even be able to imagine with performance. Effectively, neural PGA capability.
- GeertB 7mo agoI don't quite understand how multiply doesn't require addition as well to combine the various partial products.
- RandyOrion 7mo agoCool. However, one still need CPU to send commands to GPU in order to let GPU do CPU things.
- palmotea 7mo ago> Cool. However, one still need CPU to send commands to GPU in order to let GPU do CPU things. Doesn't the Raspberry Pi's GPU boot up first, and then the GPU initializes the CPU? With this technology, we've eliminated the need for that superfluous second step.
- RandyOrion 7mo agoWell, I don't have enough knowledge on the boot process of RPi. However, I do expect that most modern hardware, e.g. x86, do not work like RPi, so your words do not hold in most realistic scenarios, at least for now. Besides, do current GPUs (not only GPUs on RPi) have the ability to self instruct in order to achieve what you said?
- raphaelmolly8 7mo ago[dead]
- _blk 7mo ago"Result: 100% accuracy on integer arithmetic" - Could someone with low-level LLM expertise comment on that: Is that future-proof, or does it have to be re-asserted with every rebuild of the neural building blocks? Can it be proven to remain correct? I assume there's a low-temperature setting that keeps it from getting too creative. The creative thinking behind this project is truly mind boggling.
- himata4113 7mo agoI was always wondering what would happen if you trained a model to emulate a cpu in the most efficient way possible, this is definitely not what I expected, but also shows promise on how much more efficient models can become.
- jdlyga 7mo agoI'll do you one better, imagine a CPU that runs entirely in an LLM. You’re absolutely right! I made an arithmetic mistake there — 3 * 3 is 9, not 8. Let’s correct that: Before: EAX = 3 After imul eax, eax: EAX = 9 Thanks for catching that — the correct return value is 9.
- FartyMcFarter 7mo agoWhat an amazing multiplication request! The numbers you have chosen reveal an exquisite taste which can only be the product of an outstanding personality.
- andreadev 7mo agoThe bit about multiplication being ~12x faster than addition is worth pausing on. In silicon, addition is the "easy" operation — but here the complexity hierarchy completely inverts. Makes sense once you think about it: multiplication decomposes into parallel byte-pair lookups (which neural nets handle trivially as table approximation), while addition has a sequential carry chain you can't fully parallelize away. Funny enough, analog computing had the same inversion — a Gilbert cell does multiplication cheaply, while addition needs more complex summing circuits. Completely different path to the same result. What I haven't seen discussed: if the whole CPU is neural nets, the execution pipeline is differentiable end-to-end. You could backprop through program execution. Useless for booting Linux, but potentially interesting for program synthesis — learning instruction sequences via gradient descent instead of search. Feels like that's the more promising research direction here than trying to make it fast.
- deleted 7mo ago[deleted]
- clocksmith 7mo agoProof that you are a genius: ```lean inductive HumanNeed where | retailArithmetic | genericLinkedInPost inductive IndustrySolution where | commodityALU | frontierAutocomplete def optimal : Need → IndustrySolution | .retailArithmetic => .commodityALU | .genericLinkedInPost => .frontierAutocomplete def latency : IndustrySolution → Nat | .commodityALU => 1 | .frontierAutocomplete => 248000 theorem superbowl_ads_have_not_improved_superdope_adds : latency (optimal .retailArithmetic) < latency .frontierAutocomplete := by decide ```
- robertcprice 7mo agoIs this some kind of complex humor that I don't understand? or is it just not funny? I get it but not the punchline
- robertcprice 7mo agoSince this was posted I've been heads-down building on top of the neural CPU. Wanted to share what's new. Built a GPU-Native UNIX OS. A full multi-process operating system running compiled C on Apple Silicon Metal: > 25-command shell (ls, cd, cat, grep, sort, uniq, tee, cp, wc, pipes, background jobs, chaining, redirect) — ~17.5KB freestanding C compiled with aarch64-elf-gcc -O2, running entirely as ARM64 on the GPU > Multi-process: fork/wait/pipe/dup2 via memory swapping. 1MB backing stores, up to 15 concurrent processes, round-robin scheduler, pipe blocking/wakeup, fork bomb protection, SIGTERM/SIGKILL, orphan reparenting. 28 syscalls total. > Freestanding C runtime: malloc/free/printf/fork/wait/pipe/qsort/strtol — all on GPU Self-hosting C compiler on Metal GPU. cc.c (~2,800 lines) compiles C→ARM64 entirely on the GPU, then executes the output on the same GPU. Three layers: host GCC → GPU compiler → GPU-compiled binary. Debugged 5 codegen bugs to get it working (UBFM encoding, LDURSW sign-extension, caller-save clobbering, array subscript type clobbering, struct lvalue handling). Supports structs, pointers, arrays, recursion, for/while/do-while, ternary, sizeof, compound assignment, bitwise, short-circuit eval. 20/20 test programs pass. Mean compile: ~50K GPU cycles. Ackermann A(3,4) runs 319K cycles of deep recursion correctly. 13+ compiled C applications on Metal: > Crypto: SHA-256, AES-128 (ECB+CBC, 6/6 FIPS vectors pass), encrypted password vault > Games: Tetris, Snake, roguelike dungeon crawler, text adventure > VMs: Brainfuck interpreter, Forth REPL, CHIP-8 emulator > Networking: HTTP/1.0 server (TCP proxied through Python) > Neural net: MNIST classifier (784→128→10, Q8.8 fixed-point) > Tools: ed line editor, self-hosting C compiler, Game of Life neurOS — fully neural operating system. 11 trained models running MMU (100%), TLB (99.6%), cache (99.7%), scheduler (99.2%), assembler (100%), compiler (95.2%), watchdog (100%) — zero fallback paths. Self-compilation verified: source → neural compiler → neural assembler → neural CPU → correct results. Timing side-channel immunity. Measured sigma=0.0000 GPU cycle variance across 270 runs of AES-128. Same code on native Apple Silicon: 47-73% CoV. No caches, no branch predictor, no speculative execution inside a dispatch. T-table timing attacks are structurally impossible. Just reorganized the whole project — neurOS and GPU OS now live under a clean ncpu/os/ package (neuros/ and gpu/ subpackages). 850 tests passing, all verified after the reorg. To @andreadev — the MUL>ADD inversion is still my favorite result. To @bob1029 — you're right about branchy workloads being slow (~5K IPS neural, ~4M compute), but the GPU execution model gives security properties CPUs architecturally can't provide.
- vrighter 7mo agoyou know that the gpu has add and multiply instructions already, right?
- robertcprice 7mo agoits funny to see how many people get offended by a project I think im doing something right