3 ms·
What you're describing is a cellular automaton, in the same vein as of Conway's Game of Life. You can do lots of interesting things with those, but it's emphati
by ooterness 2y ago
What you're describing is a cellular automaton, in the same vein as of Conway's Game of Life. You can do lots of interesting things with those, but it's emphatically not where I'd start for a flexible computing platform.
Why not go the extra mile, and make each tile a small CPU? There's a Zachtronics have called TIS-100 with this premise.
https://store.steampowered.com/app/370360/TIS100/ https://store.steampowered.com/app/370360/TIS100/
- mikewarot 2y agoBecause each cell has it's own state, and 64 bits of "program" (16 bits in each of the 4 LUTs), it's unlike the game of life, where the rule is the same for each cell. I looked at a lot of choices for architecture, and wanted to allow data paths to cross without conflict, and the 4 in/4 out choice worked best without going too far. Someone did work out how you could run the game of life on a BitGrid, it's in the Esoteric Languages wiki https://esolangs.org/wiki/Bitgrid https://esolangs.org/wiki/Bitgrid I see it as something like a Turing machine, a bit less abstract, and much, much faster at computing real results. I hope it can democratize access to PetaFLOPS. The question I can't seem to find an answer to is simple... how much power does a 4 bit in/out set of LUTs with a latch take statically? How many femtojoules does it take to switch? If those numbers are good enough, it's entirely possible that really fast compute is on the table of possibilities. If not, it's another Turing machine.
- ooterness 2y agoThe power is easy enough to calculate. Let's take the a Kintex Ultrascale+ from Xilinx as a fairly typical example of a modern FPGA. Relevant documentation is the UltraScale Architecture CLB User Guide [1] and the Xilinx Power Estimator spreadsheet [2]. Each "slice" contains two flip-flops and a lookup table with 6 input bits and 2 output bits. So two slices is enough to implement each cell with room to spare. Let's say you have a 200 x 200 grid = 40k cells. That's 80k LUTs and 160k flip-flops. That's about 29% of the resources on a XCKU9P. If we assume a 100 MHz clock and 25% toggle rate (somewhat arbitrary), that's 4e12 state-changes per second. The spreadsheet indicates that circuit will consume 850 mW, or about 200 fJ per state-change. That said, this is NOT an efficient way to do arithmetic. You'd need N cells to do a fixed-point addition with N-bit arguments, and O(N^2) (give or take) to do a fixed-point multiplication. Floating point requires orders of magnitude more. There's a reason modern FPGAs have dedicated paths for fast addition and hardwired multiplier macros. [1] https://www.xilinx.com/content/dam/xilinx/support/documents/user_guides/ug574-ultrascale-clb.pdf https://www.xilinx.com/content/dam/xilinx/support/documents/... [2] https://www.xilinx.com/products/technology/power/xpe.html https://www.xilinx.com/products/technology/power/xpe.html
- imtringued 2y ago>Why not go the extra mile, and make each tile a small CPU? Xilinx AI Engine and Ryzen AI is exactly that.