7 ms·
OpenCV 5 Is Here: The Biggest Leap in Years for Computer Vision
- leoncos 4mo agoWhen I use Codex/Claude to complete a computer vision task, such as extracting assets from an image, OpenCV is their default solution. However, I believe that using YOLO and other methods is outdated. The best solution now is to directly use Nano Banana or other AI image models. A paper has proven that image generation models can perform most CV tasks well. I believe the new OpenCV should become a wrapper for VLM or AI image models.
- TZubiri 4mo agoI am confused, how can functions that output images help with functions that should take images as input?
- taneq 4mo agoThey’re multimodal LLMs trained for image generation. Turns out that if you want to generate images you gotta know what things look like.
- TZubiri 4mo agoThat's not helpful my brother. If you have details share them, if not, don't pretend you are more illuminated than me. Is the image(text) function reversible? Or are they brute force searching a nearest neighbor like word2vec/hash brute forcing.
- sorenjan 4mo agoGoogle recently released their paper "Image Generators are Generalist Vision Learners" about exactly this. They fine tuned Nano Banana pro into what they call Vision Banana which can do segmentation etc. https://arxiv.org/abs/2604.20329 https://arxiv.org/abs/2604.20329
- TZubiri 4mo agovery interesting, it seems that they use image(image,text) functions to process/filter images, effectively generating arbitrary bitmap(image), where bitmap is of the same dimension as image.
- serf 4mo agodo you realize how many edge or unconnected nodes do OpenCV work? some SBC w/ an industrial camera that is doing pick-place or go/no-go operations on a conveyor belt against a singular object type doesn't need a huge image-gen/llm model governing it. I mean have you even considered the kind of performance an opencv function can get w/ just mask-matching? I mean even with a fancy YOLO model these answers get thrown out in 1.5-50ms ; this is just a wholly different time scaling.
- nicolailolansen 4mo agoWhenever you can run a model like Nano Banana or other vision-LLM with the same compute and time performance/restrictions as an OpenCV or YOLO call, you can make that comparison. Until then, I would not call YOLO and OpenCV outdated, it's simply wrong. There's a time and place for big V-LLMs just as there is a time and place for more "traditional" computer vision methods.
- mirsadm 4mo agoThat is a very uninformed view. Real time CV is not going to be doing that anytime soon.
- regularfry 4mo agoI've built hardware with a pi zero 2 + pi cam running a mildly fine-tuned YOLO doing local-only object detection as a USB-OTG device, in a use case where any off-device API calls would have been totally unacceptable, and where the object detection was part of the human interaction loop with a hard ceiling of 300ms on the total interaction time of which the object detection was only one process among many. We're not going to fit Nano Banana or anything like it on a device with 512MB RAM and a GPU old enough to be irrelevant, and again, API calls just aren't on the menu.
- Hendrikto 4mo ago> API calls just aren't on the menu Even if they were an option, your 300ms latency requirement would exclude them anyway.
- kryptiskt 4mo agoIf I want to identify and measure the size of round things in my orange sorter machine, I shouldn't have to resort to an unnecessarily complicated solution just because some AI bros can't understand that not everything needs to be an AI model. Like, the AI model tools already exist, all that would be accomplished if OpenCV pivoted would be to take it away for people who want to do low-level vision programming. It wouldn't add anything useful to the world, just destroy an excellent library.
- wongarsu 4mo agoI can get great results from a YOLO model with 30M to maybe 300M params. To get decent CV from a LLM 8B params is the absolute minimum, closer to 30B for interesting tasks I might be on board about LLMs being the future of OCR (though many would disagree), but for general CV they are very inefficient for very limited benefit
- IanCal 4mo agoThey can however be extremely useful for curating training data. Also things like SAM and the DINO (/grounding dino) models. Also if they are better then you can also have a flow that’s cheap model -> marginal cases go to more complex thing (and a chain of these). The yolo models are really shockingly good for their cost and how well they can work with not much training data as well.
- charcircuit 4mo ago>for very limited benefit Due to how simple they are to work with they will become popular. Compare NLP before and after GPT-3. GPT-3 majorly brought down the complexity and skill needed for doing NLP tasks even if traditional NLP is much much faster. Ultimately ease of development will win out and the industry will work towards optimizing running such LLMs to make it cheap enough to run.
- sebmellen 4mo agoGreat, let me know when those models can run on-server and process/analyze streams of ID images with less than 100ms of latency. You’ll need to make sure you have a massive set of training data including all manner of slightly blurred and slightly distorted ID cards
- _the_inflator 4mo agoExactly, and all on an embedded system with quite restrictive settings and no overclocked Intel lastest generation combined with NVIDIA's 10k graphic cards.
- charcircuit 4mo agoEmbedded systems can make network calls to powerful, GPU equipped servers.
- Chu4eeno 4mo agoThey really shouldn't, though.
- charcircuit 4mo agoIt can offer a ton of user value. There is a whole industry built upon this idea, Internet of Things.
- ceejayoz 4mo agoIoT wasn't not built on "send all the data off to a hosted GenAI". It predated them by quite a few years.
- charcircuit 4mo agoThe GPUs were doing video transcoding instead of GenAI.
- Qhemlomo 4mo ago100.000 pictures take a lot of time with LLMs. Its a lot better, faster, cheaper to use LLMs for initial labeling together with hand finetuning and then training YOLO with this. Training YOLO takes a few hours and is then very fast.
- _the_inflator 4mo ago"When I use..." Dude, in business we think in terms of large numbers, internationally easily in billion times processing images. This wouldn't cut it. Also, do you buy the mega expensive super individually designed shoes from the best shoemaker there is to march along though some dirt or simply stick to gumboots? OpenCV is used behind the scenes for many of the fancy stuff those major AI provider pretend to do. Claude is a huge system and not a LLM anymore.
- hbcondo714 4mo ago> LLMs and VLMs, Running Inside OpenCV…Qwen 2.5, Gemma 3, PaliGemma, and the GPT-2 / GPT-4 family Why these specific models / versions?
- mkl 4mo agoYes, it's weird that they're so old.
- globalnode 4mo agodoes this mean im actually able to try object detection in opencv now? i mean i know basic image processing techniques, and i know "in theory" how ML works but ive never really seen a case where i can just say "heres an image now detect all the apples". theres always 1. find a model that has the knowledge, 2. hook it up to an inference engine, 3. do something useful. i always get stuck at 1.
- fnands 4mo agoThat seems to be the way things are going. Large general models have taken over in NLP, and (outside of embedded/low latency applications) it seems like they are coming for CV next. So you should soon be able to have large generic model that can detect whatever for you. It's already pretty much possible with open-vocabulary detectors like SAM3, where you could just prompt it with "Apple": https://ai.meta.com/research/sam3/ https://ai.meta.com/research/sam3/
- wongarsu 4mo agoYOLO has basically solved that for my use cases for a couple years now. If you want labels that are not in the pretrained labels it's also easy to fine-tune, provided you're willing to label 200 or so images If you need something less restricted to existing labels (say wanting all the red apples, or all cardboard signs) SAM3 is great, as the sibling comment says
- IanCal 4mo ago> provided you're willing to label 200 or so images A quick note to say that this is also a task you can hand to things like gemini.
- dekhn 4mo agoYep- this is what I do. I use a high quality VLM to generate labelled boxes (in my case, around tardigrades in a microscope image), do some light editing to fix the small number of errors, and then train YOLO26 with it. Works great, saved me tens of hours of labelling. It's a bit scary that there is a VLM that works as well as my fine-tuned model (although much slower).
- ftchd 4mo ago> One practical detail is worth knowing. The new engine is CPU-only at the moment, so if you select a non-CPU backend and target (for example CUDA or OpenVINO through setPreferableBackend and setPreferableTarget), you will want the classic engine. So there's room for even better performance!
- wongarsu 4mo agoIt's certainly a choice to make your headline feature a new ONNX engine, feature a bunch of comparisons how it's better than ONNXRuntime, while casually mentioning on the side that the cool new much faster engine is CPU-only Sure, running models on the CPU is very much a thing in computer vision (the benchmarked YOLOv8n has 37M params). But this whole announcement feels more like OpenCV catching up to the modern world, not "The Biggest Leap in Years for Computer Vision" Still great, needing fewer libraries is a good thing, but maybe a bit oversold
- VadimPR 4mo agoThe release post is AI-written with little human oversight and it shows.
- vdfs 4mo agoThe illustrations couldn't be any more generic-ai
- kphorn 4mo agomy code, my commit - ugh
- claytongulick 4mo agoI had to stop reading after: "This is not just another incremental release. OpenCV 5 is a major step forward." If a human can't be bothered to write a piece, I can't be bothered to read it.
- arcanine 4mo agoThey really improved the performance. I tested yolov8 medium segmentation model on intel i7 11th gen cpu. Opencv 4.11 : ~255ms Opencv 5.0.0 : ~185ms with the same code.
- bobmcnamara 4mo agoIntel never really improved their memory controller and busses and it shows.
- oliveiracwb 4mo agoComputer vision was the formative school for many autodidacts. Although I acquired substantial knowledge from articles translated via Power Translator and Babylon (whose outputs closely mirror those of any 2-million-parameter SLM), it was OpenCV that made concepts like convolutions, softmax, minmax, and others finally click for me. I have consistently viewed OpenCV as an intrinsically open, educational, and adaptable library. Any developer can dissect its codebase to extract a specific filter or algorithmic implementation and tailor it to their requirements. It is certainly not cruising at the velocity of trillion-dollar capital. But it holds its altitude. And it will always be there.
- maelito 4mo agoCan it detect the speed of the car without any hand-made measurement ?
- MaxikCZ 4mo agoIn pixels/second? Sure!
- monster_truck 4mo agoDo you know the focal length/AOV of your webcam?
- brk 4mo agoThat would be pretty hard to do with any level of accuracy or external calibration/input.
- deleted 4mo ago[deleted]
- sixothree 4mo agoDo you even have one known reference?
- imJack 4mo ago[dead]
- charankilari 4mo agowow its been ages
- shelled 4mo agoA few years ago I was using OpenCV is a commercial Android SDK (it might still be being used; also because iOS provided almost all of those "needs" ready-made and Android just didn't, neither did Firebase, or Jetpack suites/tools). I was the one who had added it in the SDK. There was a lot I/we could do but as an Android developer (barely any exposure to CV or even C/C++) what I felt we lacked was documentation, a community. We struggled with even shaving off parts that we did not want to ship with our SDK. Speed was such an issue. The problem was someone who just wanted to use the lib (on mobile) a lot of things felt esoteric and out of reach i.e difficult. It didn't have to be.Sadly LLM wasn't at full speed back then, barely useable, not even talked about. Something like this would have been a perfect use case of AI/LLM. A coder, not from the exact/specific field the tool was made in/from, but being able to take full advantage of its capabilities in a nuanced/selective manner.
- plasticeagle 4mo agoThe thing I love about OpenCV is that it remains hands down the best library for simply loading images and video. I've never even used any of its fancy computer vision features, but if I need to load a video file and look at the pixels - which I did need to do recently for an art project - OpenCV does it in about four lines of code.
- Joel_Mckay 4mo agoDone a few projects with OpenCV over the years, and I agree it can be fun. However, it has a few issues: 1. Patented algorithms that are effectively impossible to license in a commercial setting. 2. Permuted API that change how identically named functions behave over versions. 3. Hardware CUDA version coupling deprecating support every major release. 4. Inconsistent and contradictory documentation in the constant subtle permutations. Downstream projects tend to version lock the lib for really practical reasons. 5. A shift away from core C libraries like ImageMagick & V4l, and into C++ abstractions with legacy Swig wrapper libraries in Java or Python. 6. Perpetual-Beta culture means the library will unlikely ever really fully stabilize. It is a fun library, until people actually try to deploy something serious. As users will often simply suggest using an old version release if there is a bug. Everything from Build flags to the API documentation has never fully stabilized. ymmv =3
- harrall 4mo agoAgree with this too. OpenCV is functionality great but its constituent parts are written by many different people who all kind of do things a little differently and it shows. But I can’t really complain because it’s open source and added to by contributors.
- Joel_Mckay 4mo agoOne can... and should report when stuff is broken, or the project becomes worthless to all but one persons passing interest. =3
- Sesse__ 4mo ago
- Magnets 4mo agoThe announcement itself is pure AI slop
- thunky 4mo agoWhat about the post was not up to your standards?
- pzo 4mo agoQuite a good release although not sure why they invest so much time into their ONNX engine. I don't think they have enough stuff and big pockets to compete with ONNXRuntime, CoreAI, ExecuTorch, LiteRT. I'm happy they added option for ONNXRuntime. I wish their cv.dnn was mostly that unified wrapper around many different backends (ONNXRuntime, Executorch, LiteRT, CoreAI) and maybe just some tooling around it (performance metrics tools, model downloads etc). Transformers(.js) approach looks better for me. Wish they also invested more time into better production ready Camera I/O (for mobiles, device/format discovery, manual settings, depthmap support, etc) and better Highgui that could use different backends (skia, webgpu) and on mobiles.
- GreenSalem 4mo agoAI written release post and it shows...
- oceansky 4mo agoI can't say for sure, but there is a suspicious amount of "it's not x, it's y". At least there are no em-dashes.
- _qua 4mo agoThe diagrams definitely look like LLM output as well
- M4v3R 4mo agoThe diagrams were generated with Nano Banana Pro (most probably, or alternatively with ChatGPT Image 2), if you look closely in high contrast areas you'll see artifacts in the background that give it away. I personally don't mind AI generated content when it's properly reviewed, but unfortunately more often than not the author just glances at the result and decides it's good enough. Example: https://opencv.org/wp-content/uploads/2026/06/image-1.jpeg https://opencv.org/wp-content/uploads/2026/06/image-1.jpeg I'm not knowledgable enough to determine whether this diagram is 100% accurate, but some things look off - the arrows in the bottom left seem superficial, some arrows are connected in weird ways, the mini diagram in AttentionLayer block doesn't look right (it has two Softmax icons and one MatMul icon, while the "before" diagram is the opposite).
- bl0b 4mo agoYeah that diagram is all over the place. The arrows on the left branching from the outline of the diagram itself?
- deleted 4mo ago[deleted]
- saberience 4mo agoTested one of the diagrams: "Yes, the digital watermark indicates that most or all of this image was generated or edited using Google AI."
- pimlottc 4mo ago[dead]
- cdogukank 4mo ago[dead]
- xavierforge 4mo ago[flagged]
- boredemployee 4mo agoHow can I learn the practical side of computer vision in 2026? I'm not interested in understanding papers or the math behind it, but rather in how to put a system into production, whether it's object detection, running 20 cameras in parallel on a single computer, like sizing hardware for a specific task, and so on. Any tips?
- yayitswei 4mo agoTry a coding agent for writing and tuning the OpenCV part, and have it explain its choices. That's probably the most practical path to shipping a working system. Speaking from experience: never used OpenCV before, recently vibe coded a tool that makes supercuts of pool videos, trimming each clip from the cue ball's first strike to when the motion stops.
- eastof 4mo agoOne of the great things about OpenCV is how ubiquitous it is, there's a ton samples online and well represented in frontier model training data. I recently vibe-coded an object detector for my own personal photo library so I could separate out my pictures with humans in them. Very approachable with Codex + feeding it a sample from Github.
- bonoboTP 4mo agoBy doing it. Decide on a small project, like tracking your cat, detecting food items in your fridge, then take it step by step. Then do a slightly more ambitious project. Start with something very simple. It also heavily depends on what you already know regarding programming, image processing etc.
- kelvinjps10 4mo agoJust start cooking , python is easy and the bindings are not that hard
- ge96 4mo agoI remember trying to do photo stitching myself (panoramas) then I failed miserably but it's built into opencv ha. I've used quite a bit of OpenCV features eg. laplace variance for an automatic zoom/focusing mechanical lens camera system (steppers) and contour/blob finding for crude color segmentation.
- deleted 4mo ago[deleted]
- owenpalmer 4mo ago> This is not just another incremental release. OpenCV 5 is a major step forward. Am I the only one that finds this sentence very cheesey?
- wiradikusuma 4mo agoCurious how do people usually use OpenCV with CCTV? (Use cases)
- trollbridge 4mo agoGreat to hear. OpenCV was so easy and smooth to set up for doing tasks like generating thumbnails from uploads from arbitrary photo uploads regardless of format (including funky new formats like webp, avif, or heic).
- maxdo 4mo agocurious how many people model killed on battlefields of ukraine and russia.
- noobcoder 4mo ago[flagged]
- hdgvhicv 4mo agoThat page does t say “what is computer vision”
- riazrizvi 4mo agoI guess i'm the only one blown away by this announcement and super excited to get back into image processing.
- johnAthan_ 4mo agojudging by the amount of upvotes, it is unlikely you're the only one.
- wolfgangK 4mo agoOpenCV being in the list of Pyodide modules [0] was the biggest boon for my online teaching experience because remotely dealing with install woes (corporate proxies & cie) was a show stopper for regular Python. I'm hoping that they will package this new version and that it will bring the new neural networks engines goodies to the no-install crowd ! [0] https://pyodide.org/en/latest/usage/packages-in-pyodide.html https://pyodide.org/en/latest/usage/packages-in-pyodide.html
- dadachi 4mo agoSame on mobile. I use Apple's Vision framework on-device to find people in photos for a printing app. Sending users' personal photos to an image-model API is a non-starter on privacy, latancy, and per-photo cost alone. Less flexible than a V-LLM, but for "find the people, give me box" it's instant, free, and works offline.
- Lukevigoss 4mo ago[dead]
- ternaus 4mo agoImage augmentations library Albumentations is heavily based on OpenCV, which allows it to beat torchvision, Kornia, PIL, and other similar libraries. But there is still a huge room for improvement in terms of performance, as for some low level operations StringZilla or Numkong are faster, for some, especially for float32 images, numpy is the best. The most annoying component is that OpenCV is limited to input shapes like (H, W, C), which limits its application to videos and volumes with shapes (X, H, W, C)
- ternaus 4mo agoTo the question about OpenCV peformance to measure JPEG image decoding sequentially and as a part of the PyTorch Dataloader. TL;DR OpenCV is fast, but torchvision is faster. https://arxiv.org/abs/2605.08731 https://arxiv.org/abs/2605.08731
- mattcox12 4mo agoI'm more curious about the hardware acceleration on Snapdragon and ARM. That matters way more for real deployment than any new DNN feature.
- deleted 4mo ago[deleted]