3 ms·
I think the problem with such an architecture is that latency isn't a limit on timing closure from a hardware sense now, but you still have to consider it now f
by ChrisKjellqvist 2y ago
I think the problem with such an architecture is that latency isn't a limit on timing closure from a hardware sense now, but you still have to consider it now from the software compilation perspective, and it might severely impact performance.
From what I've thought about this, the problem applies to any high fanout signal. For instance, if you want to implement a multiplexer for two n-bit operands, you'll need the select bit to be in ~n places at once (if you compile the n-bit multiplex into n 1-bit multiplexes). Compiling LUTs to route this select signal to the right places in the grid synchronously with the arrival of the data signals is complex and amounts to a similar sort of problem one faces with hardware compilation (akin to setup timing). In this architecture, you're replacing actual routing resources (wires) with LUT entries. Instead of considering the propagation of a signal down a wire in terms of nanoseconds, you'll be thinking about it in terms of cycles to traverse the grid. Unless the clock rate is absurdly high, signals like this will probably cause a performance problem for your design.
FPGAs/ASICs also have a problem of this flavor but it generally only happens for one signal: the clock. FPGAs address this by not using regular routing resources for the clock, and instead using special, pre-routed nets for clock distribution. I imagine you'd probably need a solution like this to deal with high fanout signals in an efficient way.