mirror of
https://github.com/opencv/opencv.git
synced 2026-07-25 13:23:02 +04:00
60c705e93b
Feat: add OpenCL SIFT detector and descriptor #29450 Add OpenCL implementation of SIFT detector and descriptor extractor, triggered via T-API when the caller passes a UMat. Targets the features module (5.x directory layout). Implementation: - 4 OpenCL kernels in sift.cl: gaussian_blur_h, gaussian_blur_v, detect_and_orient, and compute_descriptor. - Host-side dispatch in sift.dispatch.cpp using CV_OCL_RUN_ macro, pure additions to the existing CPU path. - T-API entry: detectAndCompute, detect, compute. Gating and fallback: - 64-bit only. 32-bit falls back to CPU via sizeof(void*) > 4 guard. - OpenCL 1.1 or later required. Devices below 1.1 fall back to CPU. - No-OpenCL builds fall back to CPU. Verified with WITH_OPENCL=OFF: 79 tests pass, same 3 pre-existing failures (DISK/AFFINE_FEATURE missing .npy data files). Correctness: - nOctaveLayers > 3 OOB fix: gk_coeffs and gk_radius resized to nOctaveLayers+2 via std::vector (were fixed size 5). - Regression test for nOctaveLayers > 3. - 10 OCL SIFT tests pass, 60 OCL tests pass, 139 features tests pass (3 pre-existing failures unrelated to SIFT). Performance optimizations: - native_exp, native_sqrt, native_recip for hardware approximations (OpenCL 1.1 builtins). - Eliminate intermediate rawDst[128] buffer; normalize in-place from hist[]. - Non-blocking kernel launches with single ocl::finish() before host reads keypoint count. - Custom separable Gaussian blur for init image bypassing T-API GaussianBlur dispatch overhead. - exp2() replaces pow(2.0f, x); scl_octv reused. - Hoist UMat tmp allocation out of pyramid loop. - Interleaved keypoint output buffer (6 arrays to 1) for coalesced writes and fewer copies. - Consolidated descriptor keypoint copies (2N to 2 host-to-device transfers). - std::map replaced with flat vector indexed by pyramid level (O(1) vs O(log n)). - DoG computed on-the-fly in detect kernel via READ_DOG macro, eliminating 45 subtract() calls and dog_pack allocation. - Shared gauss_packs between detect and descriptor, eliminating duplicate copyTo. - Cross-level packing per octave: one descriptor kernel launch per octave instead of per level. Keypoint buffer expanded to 5 floats (added layer index). Levels packed into single UMat. Reduces kernel launch overhead and improves GPU utilization for octaves with few keypoints. - __local memory tiling for Gaussian blur kernels with 16x16 workgroups and cooperative halo loading. Reduces global memory traffic. - Direct uchar descriptor output for CV_8U: kernel writes convert_uchar_sat_rte directly instead of float buffer + host convertTo pass. Eliminates temp allocation and extra kernel launch. Performance result (stitching/s2.jpg resized): - DetectAndCompute: OCL wins at >=960x540 (1.32x at 960x540, 1.80x at 1280x720, 1.87x at 1920x1080). - Compute-only: OCL wins at >=1280x720 (1.60x at 1280x720, 1.78x at 1920x1080). Tests: - modules/features/test/ocl/test_feature2d.cpp with OCL SIFT correctness tests on real images (leuven img1.png, a3.png, s2.jpg). - DescriptorType, Regression_26139, and Batch tests mirror CPU coverage (CV_8U descriptor type, single-keypoint edge case, 6-image detect+compute). - modules/features/perf/opencl/perf_sift.cpp with real image and 5-size sweep. - CPU SIFT_Scaled fixture added to perf_feature2d.cpp for aligned OCL vs CPU comparison. ### Pull Request Readiness Checklist See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request - [x] I agree to contribute to the project under Apache 2 License. - [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV - [x] The PR is proposed to the proper branch - [ ] There is a reference to the original bug report and related work - [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable Patch to opencv_extra has the same branch name. - [x] The feature is well documented and sample code can be built with the project CMake