5 ms·
I had the opportunity to work on TinyML, it's a wonderful field! You can do a lot even with very small hardware. For example, it's possible to get real-time c
by cooootce 3y ago
I had the opportunity to work on TinyML, it's a wonderful field!
You can do a lot even with very small hardware.
For example, it's possible to get real-time computer vision system with an esp32-s3 (dual-core XTensa LX7 @ 240 MHz cost like 2$), of course using the methods given in the article (Pruning, Quantization, Knowledge distillation, etc.). The more important thing is to craft the model to fit as much as possible your need.
More than that, it's not that hard to get into, with solution named AutoML that do a lot for you. Checkout tool like Edge impulse [0], NanoEdge AI Studio [1], eIQ® ML [2]
There is a lot of tooling that is more low-level too, like model compiler (TVM or glow) and Tensorflow Lite Micro [3].
It's very likely that TinyML will get a lot more of traction. A lot of hardware companies are starting to provide MCU with NPU to keep consumption as low as possible. Company like NXP with the MCX N94x, Alif semiconductor [4], etc.
At my work we have done an article with a lot of information, it's in French but you can check it out: https://rtone.fr/blog/ia-embarquee/ https://rtone.fr/blog/ia-embarquee/
[0]: https://edgeimpulse.com/ https://edgeimpulse.com/
[1]: https://stm32ai.st.com/nanoedge-ai/ https://stm32ai.st.com/nanoedge-ai/
[2]: https://www.nxp.com/design/design-center/software/eiq-ml-development-environment:EIQ https://www.nxp.com/design/design-center/software/eiq-ml-dev...
[3]: https://www.tensorflow.org/lite/microcontrollers https://www.tensorflow.org/lite/microcontrollers
[4]: https://alifsemi.com/ https://alifsemi.com/
- Archit3ch 3y agoWhat about the Milk-V Duo? 0.5 TOPS INT8 @ $5.
- cooootce 3y agoDidn't know about it but their design decision is really cool (not very clear with the difference between the normal version and the "256 Mo" confusing). The software side doesn't seem very mature with very few help regarding TinyML. But this course seem interesting https://sophon.ai/curriculum/description.html?category_id=48 https://sophon.ai/curriculum/description.html?category_id=48
- mysterydip 3y agoOne thing I've wondered in this space: Let's say for a really basic example I want to identify birds and houses. Is it better to make one large model that does both, or two small(er) models that each does one?
- flyingcircus3 3y agoWhy not three models? One model does basic feature detections, like lines, shapes, etc. A second model that can take the first model's output as its input, and identify birds. A third model can take the first model's output as its input, and identify houses.
- wegfawefgawefg 3y agoThis is a lesson I've watched people, and companies learn for the past 7-8 years. An end to end model will always outperform a sequence of models designed to target specific features. You truncate information when you render the data into output space (the model output vector) from feature space (much richer data inside the model), thats the primary reason why to do transfer learning all layers are frozen, the final layer is chopped off, and then the output of the internal layer is sent into the next model. Not the output itself. Yes you can create a large tree of smaller models, but the performance cieling is still lower. Please don't tell people to do this. Ive seen millions wasted on this. When you train a vision model it will already develop a heirarchy of fundamental point, and line detectors in the first few layers. And they will be particularly well chosen for the domain. It happens automatically. No need to manually put them there.
- DoingIsLearning 3y agoAs someone not in ML but curious about the field this is really interesting. Intuitively indeed it would be natural to aim for some sort of inspectable composition of models. Is there specific tooling to inspect intermediate layers or will they be unintelligible for humans?
- 3y ago
- anigbrowl 3y agoGreat post. surprised and excited to discover Tensorflow models can run on commodity hardware like the ESP32.
- Reviving1514 3y agoI ended up hand rolling a custom micropython module for the S3 to do a proof of concept handwriting detection demo on an ESP32, might be interesting to some. https://luvsheth.com/p/running-a-pytorch-machine-learning https://luvsheth.com/p/running-a-pytorch-machine-learning
- cooootce 3y agoGreat post with very interesting detail, thanks ! Another optimization could be to quantize the model, this transform all compute as int compute and not as floating point compute. You can lose some accuracy, but for any bigger model it's a requirement ! Espressif do a great job on the TinyML part, they have different library for different level of abstraction. You can check https://github.com/espressif/esp-nn https://github.com/espressif/esp-nn that implement all low level layers. It's really optimized and if you use the esp32-s3 it will unlock a lot of performance by using the vector instructions.
- Reviving1514 3y agoYou are right I should definitely be looking into how to run these models as ints as well, especially with the C optimizations to micropython you would see a lot larger performance gains using ints compared to floats. Definitely need to find some time to try it! On the other hand the tinyML library looks great too and if I was going to do this for a product that would likely be the direction I would end up taking just cause it would be more extensible and better supported. Thank you for the links!
- Cacti 3y agoProblems reducible even partially to matrix math are for many practical purposes embarrassing parallel even within a single core. A couple hundred million FLOPS with 1990s SIMD support will let you run nearly all near-SOTA models within, idk, 3s, with most running in 0.1 or 0.01s. That’s pretty fast considering it’s an EP32 and some of these capabilities/models didn’t even exist a year ago. Your expectation was not really wrong, because for most purposes, when discussing a “model” one is really talking about “capabilities”. And capabilities often require many calls to the model. And that capability may be reliant on being refreshed very rapidly… and now your 0.1s is not even slow, it’s almost existentially slow. Re: training. even on the EP32, training is entirely doable, so long as you pretend you are in 2011 solving 2011 problems hahaha
- demondemidi 3y agoI think we know each other. ;)
- Cacti 3y agothank you for the post and good work. can I ask, is the focus primarily on inference? is there anything serious going on with training at the power scale you are talking about?
- cooootce 3y agoThanks ! Yes, the main focus is on inference. It's possible to re-train a simple model at this power scale, but it's often time very small model and not deep-learning. Nanoedge AI studio from STelectronic give you some tool to train the model after deployment on device. It's often time used for predictive maintenance, in order to adapt each ML model at the water pump plugged, for example.