3 ms·
An excellent book for fundamentals. Still haven't found a good textbook that covers the next level, that takes you from a student to competent practitioner. Adv
by krapht 1y ago
An excellent book for fundamentals. Still haven't found a good textbook that covers the next level, that takes you from a student to competent practitioner. Advanced knowledge that I've picked up in this field has been from coworkers, painfully gained experience, and reading Kaggle writeups.
- bonoboTP 1y agoIt gets specialized after that. You need to be more specific about the area you are interested in. Computer vision is a very broad field. For newer topics, there are often no textbooks yet because it takes time to write books and the methods and practices change quite fast, so it takes time to stand the test of time. Your best bet is arXiv and GitHub to learn the latest things. Object detection / segmentation, human pose (2D/3D), 3D human motion tracking and modeling, multi-object tracking, re-identification and metric learning, action recognition, OCR, handwriting, face and biometrics, open-vocabulary recognition, 3D geometry and vision-language-action models, autonomous driving, epipolar geometry, triangulation, SLAM, PnP, bundle adjustment, structure-from-motion, 3D reconstruction (meshes, NeRFs, Gaussian splatting, point clouds), depth/normal/optical flow estimation, 3D scene flow, recovering material properties, inverse rendering, differentiable rendering, camera calibration, sensor fusion, IMUs, LiDAR, birds eye view perception. Generative modeling, text-to-image diffusion, video generation and editing, question answering, un- and self-supervised representation learning (contrastive, masked modeling), semi/weak supervision, few-shot and meta-learning, domain adaptation, continual learning, active learning, synthetic data, test-time augmentation strategies, low-level image processing and computational photography, event cameras, denoising, deblurring, super-resolution, frame-interpolation, dehazing, HDR, color calibration, medical imaging, remote sensing, industrial inspection, edge deployment, quantization, distillation, pruning, architecture search, auto-ML, distributed training, inference systems, evaluation/benchmarking, metric design, explainability etc. You can't put all that into a single generic textbook.
- greenavocado 1y agoPlus photogrammetric scale recovery, rolling-shutter & generic-camera (fisheye, catadioptric) geometry, vanishing-point and Manhattan-world estimation, non-rigid / template-based SfM, reflectance/illumination modelling (photometric stereo, BRDF/BTDF, inverse rendering beyond NeRF), polarisation, hyperspectral, fluorescence, X-ray/CT/microscopy, active structured-light, ToF waveform decoding, coded-aperture lensless imaging, shape-from-defocus, transparency & glass segmentation, layout/affordance/physics prediction, crowd & group activity, hand/eye/gaze performance capture, sign-language, document structure & vectorisation charts, font/writer identification, 2-D/3-D primitive fitting, robust RANSAC variants, photometric corrections (rolling-shutter rectification, radial distortion, HDR glare, hot-pixel mapping), adversarial/corruption robustness, fairness auditing, on-device streaming perception and learned codecs, formal verification for safety-critical vision, plus reproducibility protocols and statistical methods for benchmarks
- thenobsta 1y agoIt's astounding how much there is to this field.