5 ms·
Show HN: Simply Reading Analog Gauges – GPT4, CogVLM Can't
- omeze 3y agoThis is really cool - based on the project page it uses off the shelf models like Yolo and Tesseract plus a ton of synthetic data (which is also freely available). Also solves a real industrial use case!
- icecoffee 3y agoYup, the components can be replaced to your liking and expand the data to your match your environment better. I found it really interesting how, no matter how I prompted foundational models they couldn't read gauges.
- bsenftner 3y agoI'm surprised something like reading gauges is not one of the entry level computer vision problem data sets. It seems to obviously useful.
- rob74 3y agoWell, I hope that for any serious application they would replace the analog gauges with sensors that produce values that can be directly read by a computer system, rather than teaching it to read the analog gauges. So I think the actual real-world usefulness of this is rather limited...
- bsenftner 3y agoI guess you're unaware of the gargantuan number of legacy manufacturing plants littered all over the globe that are trying to retrofit current tech over 40+ year old tech while it is still operating. Think 2nd and 3rd world manufacturing; far larger than you probably imagine.
- _fizz_buzz_ 3y agoAlso 1st world manufacturing.
- Workaccount2 3y agoMy company has a huge legacy business where we are still manufacturing brand new PC boards with intel 8086's on them. We have a huge inventory of 40 year old chips for it, lol. Nobody wants to upgrade when the current system still works and can be supported (however painfully).
- icecoffee 3y agoExactly, I use to think the same thing! Many legacy system, the industrial machines that run factories, infrastructure, etc. all have hundreds of gauges and some of them do get "watched". Cost of digitization via smart gauges just cost too much and planning that goes into replacing them often isn't tenable. So many instrument readings don't get logged - until some machine fails. There are systems out there, but they aren't affordable to make a broad impact - I hope we can progress to change that.
- idiotsecant 3y agoIts very common for plant operators to do 'rounds' and inspect gauges that are not digital because they are not usually terribly important, because the gauge provides a backup to the electronic reading, etc. Put this on a 'spot' robot from Boston dynamics and you have an operator that won't be hurt when there is a steam leak, or whatever. There are definitely use cases.
- Workaccount2 3y agoIt would be even more trivial to just have cameras and a human who can quickly scroll through the different cams to check the gauges. The question is how much you trust this vision system, not whether or not you still need a guy to do rounds.
- idiotsecant 3y agoNo, fixed cameras have a linear maintenance cost and are subject to all the regular problems fixed infrastructure has. A robot has the same cost whether it spends all day walking in circles or doing rounds.
- Workaccount2 3y agoYou would probably have to pay for 80 years of "camera maintenance" before you met the cost of 1 robot.
- icecoffee 3y agoThe cv reading here is really a POC base on off-the-self and will require some work. Fix-camera vs Spot, I can definitely see the pros and cons. So at what price point do you think industry will adopt Spot for doing rounds?
- pona-a 3y agoI recall there was a project integrating utility meters into Home Assistant, where a tiny ESP board with on-device ML used a camera to read the analog gauge, basically the most practical solution given you can't replace or modify your utility meter and only a handful offer a serial port.
- mkl 3y agoDoesn't seem to work me for uploaded images on mobile. Firefox Android tries to take a photo instead of uploading one, and Chrome Android consistently stops and says "Error" 6 seconds after pressing "Submit".
- eurekin 3y agoI was just about to whip up a openCV version for sth very similar! If this works, I might be able to diagnose the eCVT issue I have with my car. EDIT: ah, it's for the other types of gauges. Errors out on mine :) https://imgur.com/a/J4IHP0r https://imgur.com/a/J4IHP0r
- icecoffee 3y agoYea... the model wasn't trained on speedometers. I curious how this might help with your "eCVT issue", can you elaborate?
- eurekin 3y agoIt manifests only at a certain speed range (40 to 55 kmph): acceleration pedal press isn't translated into car acceleration instantly, but with up to 3 second pause, during which the ICE is revving up almost to the max. I wanted to record speed, rpm and acceleration to graph that and show at the repair station. At one visit in a authorised repair station the mechanic couldn't diagnose it (unrelated heavy snowing conditions).
- eurekin 3y agoIf anyone's interested: it works as designed. I just bumped into a inherent flaw of a eCVT, which was exaggarated by bad weather.
- cibyr 3y agoSounds like a job for a cheap Bluetooth OBDII dongle.
- eurekin 3y agoYes, and that's what I'll probably end up doing. It surprised me that there is no off the shelf solution using computer vision and that piqued my interest to refresh the SOTA
- bondarchuk 3y ago
- mitjam 3y agoThere is also the "AI on the Edge Device" project [1] which can read both analog gauges and digital displays and runs on an ESP32. One read-out takes more than a minute but inference runs on the ESP32 with TensorFlow Lite. [1]: https://github.com/jomjol/AI-on-the-edge-device https://github.com/jomjol/AI-on-the-edge-device
- hncomb 3y agoIs this commercially usable? Great work!
- coder543 3y agoI don’t see clear licensing on any of this, including the synthetic dataset.
- amelius 3y ago"The model was built only with synthetic data (e.g. examples). Hence, it probably will not work on significantly different images - give it a try. Let us know, so we can keep improving."
- dotancohen 3y agoThere are even lots of errors on the four example images' readings. I'm on mobile so I won't file the bugs, but this project has a lot of work to do.
- RecycledEle 3y agoThis is not surprising. LLMs trained in Internet data are Internet simulators. How often have you asked the Internet to read an analog gauge for you? Probably never.
- RecycledEle 3y agoMaybe we need an LLM to follow grandpa around his auto shop and learn a few things. Glasses with cameras and mics might help record the data. This could also be a way to train a bot to do grandpa's job.
- devmandan 3y agoI take issues with the approach. Anyone know anyone who wants dial gauges retrofitted? I got a way to do it that seems good enough.
- bearjaws 3y agoWould fine tuning be able to fix this kind of problem in a model like LLaVA? I am not sure if it even has the right foundation data set to even understand the myriad of gauges that exist in the world, so I doubt fine tuning will fix anything.
- icecoffee 3y agoI suspect models probably need "reasoning" ability - very lacking in current models. Perhaps, AlphaGeometry can do better.
- observationist 3y agoAll watches and clocks generated by dall-e (and almost all other generative models) show 10:10, because this is the standard display "pose" in watch ads, so most of the training data shows things in a particular configuration. Synthetic data showing relative gauge positions and a wide variety of contexts anchored to a particular category, like "dial gauge showing 31%" is the solution, just like watch images showing a full range of perspectives and times is the solution to the 10:10 problem. I think the trickiest problem in transformer models is identifying situations like this where there's a mode collapse, but it's one or more degrees of separation from being apparent. This can lead to generation being technically and structurally correct (like a beautifully rendered wristwatch,) but the range of available answers will be limited to whatever the model has collapsed on (i.e. all watches and clocks only show one time, even if you ask it to display 4:59 PM and construct a perfect prompt about something right before work is done for the day.) These are subtle distortions in the world model, but things like the 10:10 problem likely also affect other concepts that overlap with clocks and watches, like analog gauges. Maybe there are patterns in the model that can be associated with these types of issues, and using noise and synthetic data can target and unstick a model?
- icecoffee 3y agoYes, I encountered a similar issues with the first iteration of dataset. The prediction location of the needle tip would be wildly off when the view angle was high. The fix was easy, just generated high viewing angle images for the 2nd dataset.
- pona-a 3y agoCan it be remedied at inference time by decaying the collapsed tokens' probability or is the world model itself defective?
- godelski 3y agoI tried this, with an example image they give[0] and it failed. It shows a value of 1.9, which is where the shadow is pointing but not the actual needle. I am suspect that the training data is good. Many tasks highly depend on labeled datasets and unfortunately these need to be very large (given current ML methods). The major problem is that these labels are generated from cheap or free labor where labelers are also expected to know the nuance and label appropriately. Nearly (if not every) dataset you've seen has major errors in them that matter. Honestly, one of the only ways to combat this is be insistent on viewing benchmarks as guides. There's been countless times I've demonstrated to others how I can perform worse the training dataset but absolutely demolish results when using customer datasets or real world data (i.e. generalization). Remember, it is absolutely possible to overfit your model even if your validation accuracy does not decrease. There are many ways to overfit and evaluation is unfortunately substantially harder than looking at results on a test set. There's a lot of complexity to this and if you want an effective product you should probably have at least one person on your team that cares about statistics and will nitpick stuff (though you can also over-saturate your team with these people too. I say that as one of these people). Edit: Actually I believe every example shows an incorrect reading. [0] https://synanthropic-reading-analog-gauge.hf.space/--replicas/bkh3j/file=/tmp/gradio/195709a4da49f2cbabbffb74cd6ac25b22b3ed36/example0.jpg https://synanthropic-reading-analog-gauge.hf.space/--replica...
- icecoffee 3y agoThough I haven't effort into analyzing or tweaking the training, some of the perditions does look like it points at the shadow instead of the tip. This synthetic dataset is tiny (not even 100k) - but still larger than what I could find. I think there's a lot of room for improvement. What do you think is the most important requirement?
- godelski 3y agoThat's all fair, and I'm a big believer in context for proper evaluation. My issue is really about the presentation. There's a pervasive issue in research that exaggeration compounds as if one work exaggerates its effectiveness that subsequent works have a higher bar to meet to show improvements and the frequent result is additional exaggeration (even in just small amounts). > What do you think is the most important requirement? Data quality and augmentation. I looked at the dataset[0] and the quality is very poor. There is very low diversity and it appears that it is composed of a few images and increased through augmentation. Augmentation should not be handled through dataset generation but dataset processing. Torch will even move bounding box keypoints when cropping or resizing images. Data quality is one of the most important factors and I know it is common to just get labels to Turk or cheap services but you absolutely get what you pay for. [0] Please don't harvest emails and usernames to require downloading. I created a dummy account for this because that's absolutely not acceptable. DO NOT SEND PEOPLE SPAM. DO NOT HARVEST EMAILS AND SELL THOSE EMAILS. I happily flag emails that come to my account through means like this until Google identifies all emails from that sender as spam. DO NOT DO THIS. It will just make people hate you and build a devoted base against you. There is no justification for this. You should absolutely rescind the gate and delete all emails you harvested. Your licensing is already an agreement. Don't be shady.