5 ms·
Related: Generate fully specialized, stand-alone, human-readable C code for AVX-512 neural net inference: https://NN-512.com https://NN-512.com Far outperfor
by 37ef_ced3 4y ago
Related:
Generate fully specialized, stand-alone, human-readable C code for AVX-512 neural net inference:
https://NN-512.com https://NN-512.com
Far outperforms TensorFlow, PyTorch, etc.
Code generator is also stand-alone (no dependencies) and is written in Go.
- touisteur 4y agoI see a lot of references to NN512 in HN and while I'm quite a fan of the idea of a high performance, well crafted neural network code generator, I'm afraid such a library answers to only a very small part of the problem. But, on the NN512 page I don't see how to take an existing trained network (from tensorflow, torch, mxnet, onnx...) to generate the code? I'm not sure whether NN512 does any of the quite efficient layer fusion or layer-op fusions. How does it compare, with openvino, tvm, tensorrt (I'm restricting myself to C/C++ stuff, but comparing to the python stuff would be nice too)? When one says 'far outperforms' other libraries I'd rather like to know all that. And ConvNet is quite old now, many state of the art network architecture are going away from pure cnns and going into rnns, transformers, and even deeper networks where I'm not sure AVX512 really helps (except for batching and even then...). As for 'human readable' that's quite the stretch, if you're not an avx512 intrinsic reader it's quite the sequence of instructions. Is there a comprehensive, recent, answer to all or some of this in the NN512 docs, that I missed? Also it seemed last time that I checked that the author had tried to monetize the project somehow and was quite disheartened and was (iirc?) kind of abandoning the effort? I can understand how the ML/AI inference can be discouraging, and how hard it is to follow the big 2 (tf, torch) even for deep pocket orgs. Can't imagine how much people like FINN/Tensil/Tenstorrent will spend supporting all operators... And new ones keep appearing...
- 37ef_ced3 4y agoNN-512 was monetized by making its author a Senior Architect at the most important company in the machine learning space, a company whose hardware and software you almost certainly use. NN-512 was an EXTREMELY profitable project for its author. Most competent engineers will be able to dump their network weights to a file as described in the headers. Do you understand your network or not? The weights, etc., are literally a C struct in the header file. Use fread() to load the floats. If that's too complicated, NN-512 is not for you. As clearly stated on its brief webpage (which you might care to read before commenting) NN-512 does every kind of fusion you can imagine, fusing convolutions, fusing elementwise operations, fusing batch normalization, etc. See https://nn-512.com/example/28 https://nn-512.com/example/28 for an example. If you're using a classical ConvNet (e.g., a DenseNet-121), as many people are, and you want to do inference on AVX-512 CPUs, nothing will outperform NN-512. And your inference code will be completely stand-alone. No dependencies.
- touisteur 4y agoI'm glad the author (you?) was successful and managed to monetize NN512. I seemed to remember reading some rant about it but it seems I was mistaken, sorry. I didn't mean to be disparaging and I guess that's my fault asking questions or stating doubts on HN but last time I tried, with some not off the shelf models, to use it wasn't 'just' loading structs. There were some ops missing and not every one uses mainstream convnets. I guess I was wrong too and couldn't make some lstms work right aways. As for the comments on 'competent', not caring to read, asking if I know my networks (as I talked about using other inference libraries that I seem to be able to make work), etc. I think that was uncalled for, but since I seem to annoy you, I promise not chip in on NN512 anymore. Thanks.
- 37ef_ced3 4y agoI didn't mean to insult you. My apologies. All I meant was: it's not hard to write your weights to a file as floats and fread() them into a struct. I view this as the simplest possible universal interface. And you're right: NN-512 supports only the classical ConvNet operations. It does that one thing, and it does it well. If you need LSTMs or transformers, or training, etc., another tool is needed. Have a nice day.
- mirker 4y agoYou should cross-reference the author’s alias with the NN512 web-page.
- touisteur 4y agoI should have yes, thanks. I didn't think it would be a problem, I usually don't check before posting, but there am I.