3 ms·
Some assorted ideas: * Expose the data dependency graph directly to the processor, rather than forcing the processor to infer it from the instructions. * Anno
by piinbinary 10y ago
Some assorted ideas:
* Expose the data dependency graph directly to the processor, rather than forcing the processor to infer it from the instructions.
* Annotate when data is read-only, to reduce communication between cores (for the sake of avoiding latency, not for bandwidth savings).
* Add a mechanism for much cheaper (if more limited) parallelism where a single core would work on multiple, related thunks in parallel, even if those units of work would be far too small to be worth coordinating with another thread to offload. This would likely be largely implicit, taking advantage of the first feature.
* Instructions for graph traversal. The CPU could (for some uses) order the traversal in a way that improves cache locality, based on how the graph is actually laid out in memory (and prefetch uncached nodes while working on cached ones).
* Something like map & reduce, where you can apply a small (pure) function to a list of data. Again, this would likely be done in parallel.
- qznc 10y ago> data dependency graph directly to the processor What format would you propose? How is it different to instructions? > Annotate when data is read-only What overhead would you save? On the cache coherence protocol level? > a single core would work on multiple, related thunks in parallel VLIW https://en.wikipedia.org/wiki/Very_long_instruction_word https://en.wikipedia.org/wiki/Very_long_instruction_word > instructions for graph traversal Your Intel CPU can already do prefetching. Others have special instructions for it. Hyperthreading does the "work on cached stuff first" part. > like map & reduce Use the GPU.
- yvdriess 10y agoAnswering with my own take on the subject. > What format would you propose? How is it different to instructions? It would still be instructions, but either by reworking what their operands are and introducing a destination (instruction address + operand port). Another option that could maybe work is a second and mirroring instruction stream that contains only the dataflow/dependencies. > What overhead would you save? On the cache coherence protocol level? Current strategy is write-invalidate, I believe. In some heavy-contention situations (e.g. lots of spin-loops on one variable), using a 'write-once' instruction would traffic. Thinking more radically: With an overhaul of the entire virtual-memory system, a write-once instruction becomes potentially very interesting (Jack Dennis had an interesting paper on the subject: http://www.cs.ucy.ac.cy/dfmworkshop/wp-content/uploads/2014/08/DFM2014-9-On-the-Feasibility-of-a-Codelet-Based-Multi-core-Operating-System.pdf http://www.cs.ucy.ac.cy/dfmworkshop/wp-content/uploads/2014/...). > VLIW Only if you mean something like Mill (https://en.wikipedia.org/wiki/Mill_architecture https://en.wikipedia.org/wiki/Mill_architecture), instead of Itanium.
- Symmetry 10y agoA VLIW doesn't "work on multiple, related thunks in parallel" but instead works on multiple instructions from a stream. What the GP is talking about would be something that can parralelize across function calls or loop iterations and there's now way in a conventional VLIW to do that since each loop iteration contains the same instructions.