mirror of
https://github.com/opencv/opencv.git
synced 2026-07-21 19:33:03 +04:00
a61ff0fa81
### Pull Request Readiness Checklist See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request - [x] I agree to contribute to the project under Apache 2 License. - [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV - [x] The PR is proposed to the proper branch - [ ] There is a reference to the original bug report and related work - [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable Patch to opencv_extra has the same branch name. - [ ] The feature is well documented and sample code can be built with the project CMake <!-- Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the PR title. --> ### Summary This PR improves SIMD coverage for masked `cv::accumulate()` on 4-channel images. The existing masked accumulate SIMD implementation has special paths for `cn == 1` and `cn == 3`, while 4-channel images fall back to the general implementation. This PR adds `cn == 4` SIMD paths using OpenCV Universal Intrinsics. Compare link: https://github.com/opencv/opencv/compare/4.x...feitianduowen:opencv:4.x Modified files: - `modules/imgproc/src/accum.simd.hpp` - Add masked `cn == 4` SIMD paths for: - `CV_32FC4 -> CV_32FC4` - `CV_8UC4 -> CV_32FC4` - `modules/video/test/test_accum.cpp` - Add deterministic correctness regression tests for the new masked 4-channel accumulate paths. - `modules/imgproc/perf/perf_accumulate.cpp` - Add performance coverage for masked 4-channel accumulate. imgproc: add SIMD paths for masked 4-channel accumulate #29094 ### Implementation The new SIMD branches are added in: - `modules/imgproc/src/accum.simd.hpp` They handle: - `src`: `CV_32FC4`, `dst`: `CV_32FC4`, `mask`: `CV_8UC1` - `src`: `CV_8UC4`, `dst`: `CV_32FC4`, `mask`: `CV_8UC1` For the `CV_32FC4` path, the implementation uses `v_load_deinterleave`, applies the pixel mask to all 4 channels, accumulates into `dst`, and stores the result with `v_store_interleave`. For the `CV_8UC4 -> CV_32FC4` path, the implementation loads 4 interleaved uchar channels, applies the pixel mask, widens the values to float, accumulates into `dst`, and stores the results back as interleaved `CV_32FC4`. The masked SIMD block is guarded with: ```cpp if (x <= len - cVectorWidth) { ... } ``` This check ensures that SIMD setup and vector operations are only executed when at least one full vector iteration can run. For very small inputs or tail-only cases, the function skips SIMD initialization and lets the existing scalar fallback handle the data through `acc_general_(src, dst, mask, len, cn, x)`. This keeps tail handling unchanged and avoids unnecessary SIMD initialization when the vector loop would be skipped. ### Accuracy tests Added a deterministic regression test in: * `modules/video/test/test_accum.cpp` New test: ```bash Video_Acc.accuracy_32FC4_masked_cn4 Video_Acc.accuracy_8UC4_to_32FC4_masked_cn4 ``` Test commands: ```bash opencv_test_video --gtest_filter=Video_Acc.accuracy_32FC4_masked_cn4 opencv_test_video --gtest_filter=Video_Acc.accuracy_8UC4_to_32FC4_masked_cn4 opencv_test_video --gtest_filter="Video_Acc.*:Video_AccSquared.*:Video_AccProduct.*:Video_RunningAvg.*" ``` The filtered accumulate-related tests passed locally. I added dedicated deterministic tests instead of only extending the existing randomized accumulate tests because this PR only changes the masked `cv::accumulate()` paths for specific 4-channel type combinations. It does not add `cn == 4` SIMD coverage for `accumulateSquare`, `accumulateProduct`, or `accumulateWeighted`. The existing base accumulate tests are shared by several accumulation functions. Extending the common randomized channel selection to include `cn == 4` would also affect tests for functions that are not optimized by this PR. The new deterministic tests directly cover the modified paths: * `cv::accumulate(src, dst, mask)` * `CV_32FC4 -> CV_32FC4` * `CV_8UC4 -> CV_32FC4` * `mask`: `CV_8UC1` The test also covers small sizes and non-vector-multiple sizes, so both the new SIMD path and the existing scalar tail fallback are exercised. ### Performance tests Added performance coverage in: * `modules/imgproc/perf/perf_accumulate.cpp` Perf command: ```bash opencv_perf_imgproc --gtest_filter="*AccumulateMask32FC4*" --perf_min_samples=1000 --perf_force_samples=1000 opencv_perf_imgproc --gtest_filter="*AccumulateMask8UC4To32FC4*" --perf_min_samples=1000 --perf_force_samples=1000 ``` Test environment: * Release build * AVX2 enabled * IPP disabled * OpenCL disabled * Windows MinGW build Performance results: `CV_32FC4 -> CV_32FC4` | Size | Before median | After median | Speedup | | --------- | ------------- | ------------ | ------- | | 1920x1080 | 3.07 ms | 2.74 ms | 1.12x | | 1280x720 | 0.84 ms | 0.73 ms | 1.15x | | 640x480 | 0.17 ms | 0.16 ms | 1.06x | | 320x240 | 0.04 ms | 0.04 ms | ~1.00x | `CV_8UC4 -> CV_32FC4` | Size | Before median | After median | Speedup | | --------- | ------------- | ------------ | ------- | | 1920x1080 | 6.09 ms | 1.75 ms | 3.48x | | 1280x720 | 2.69 ms | 0.82 ms | 3.28x | | 640x480 | 0.96 ms | 0.28 ms | 3.43x | | 320x240 | 0.24 ms | 0.06 ms | 4.00x | | 127x61 | 0.02 ms | 0.01 ms | 2.00x | ### Build and test commands The following commands were run from Windows `cmd.exe`. ```sh cmake -S . -B build_release -G Ninja -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTS=ON -DBUILD_PERF_TESTS=OFF -DBUILD_EXAMPLES=OFF -DWITH_IPP=OFF -DWITH_OPENCL=OFF -DCMAKE_C_FLAGS=-mstackrealign -DCMAKE_CXX_FLAGS=-mstackrealign cmake --build build_release --target opencv_test_video -j 4 build_release\bin\opencv_test_video.exe --gtest_filter=Video_Acc.accuracy_32FC4_masked_cn4 build_release\bin\opencv_test_video.exe --gtest_filter=Video_Acc.accuracy_8UC4_to_32FC4_masked_cn4 build_release\bin\opencv_test_video.exe --gtest_filter=Video_Acc.*:Video_AccSquared.*:Video_AccProduct.*:Video_RunningAvg.* ``` Accuracy test commands passed locally. For the performance comparison, I kept the new perf test in `modules/imgproc/perf/perf_accumulate.cpp` and only reverted `modules/imgproc/src/accum.simd.hpp` to measure the baseline. Then I restored the SIMD patch and measured the optimized version. Baseline measurement: ```sh git diff -- modules/imgproc/src/accum.simd.hpp > accum_cn4_simd.patch git checkout -- modules/imgproc/src/accum.simd.hpp mkdir perf_logs cmake -S . -B build_perf_before -G Ninja -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTS=OFF -DBUILD_PERF_TESTS=ON -DBUILD_EXAMPLES=OFF -DWITH_IPP=OFF -DWITH_OPENCL=OFF -DCMAKE_C_FLAGS=-mstackrealign -DCMAKE_CXX_FLAGS=-mstackrealign > perf_logs\before_cmake_config.txt 2>&1 cmake --build build_perf_before --target opencv_perf_imgproc -j 4 > perf_logs\before_build.txt 2>&1 build_perf_before\bin\opencv_perf_imgproc.exe --gtest_filter=*AccumulateMask32FC4* --perf_min_samples=1000 --perf_force_samples=1000 > perf_logs\before_perf.txt 2>&1 ``` ```sh git diff -- modules/imgproc/src/accum.simd.hpp > accum_cn4_u8_simd.patch git checkout -- modules/imgproc/src/accum.simd.hpp cmake -S . -B build_perf_before_u8 -G Ninja -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTS=OFF -DBUILD_PERF_TESTS=ON -DBUILD_EXAMPLES=OFF -DWITH_IPP=OFF -DWITH_OPENCL=OFF -DCMAKE_C_FLAGS=-mstackrealign -DCMAKE_CXX_FLAGS=-mstackrealign cmake --build build_perf_before_u8 --target opencv_perf_imgproc -j 4 build_perf_before_u8\bin\opencv_perf_imgproc.exe --gtest_filter=*AccumulateMask8UC4To32FC4* --perf_min_samples=1000 --perf_force_samples=1000 > perf_logs\before_perf1.txt 2>&1 ``` Optimized measurement: ```sh git apply accum_cn4_simd.patch cmake -S . -B build_perf_after -G Ninja -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTS=OFF -DBUILD_PERF_TESTS=ON -DBUILD_EXAMPLES=OFF -DWITH_IPP=OFF -DWITH_OPENCL=OFF -DCMAKE_C_FLAGS=-mstackrealign -DCMAKE_CXX_FLAGS=-mstackrealign > perf_logs\after_cmake_config.txt 2>&1 cmake --build build_perf_after --target opencv_perf_imgproc -j 4 > perf_logs\after_build.txt 2>&1 build_perf_after\bin\opencv_perf_imgproc.exe --gtest_filter=*AccumulateMask32FC4* --perf_min_samples=500 --perf_force_samples=500 > perf_logs\after_perf.txt 2>&1 ``` ```sh git apply accum_cn4_u8_simd.patch cmake -S . -B build_perf_after_u8 -G Ninja -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTS=OFF -DBUILD_PERF_TESTS=ON -DBUILD_EXAMPLES=OFF -DWITH_IPP=OFF -DWITH_OPENCL=OFF -DCMAKE_C_FLAGS=-mstackrealign -DCMAKE_CXX_FLAGS=-mstackrealign cmake --build build_perf_after_u8 --target opencv_perf_imgproc -j 4 build_perf_after_u8\bin\opencv_perf_imgproc.exe --gtest_filter=*AccumulateMask8UC4To32FC4* --perf_min_samples=1000 --perf_force_samples=1000 > perf_logs\after_perf1.txt 2>&1 ``` Performance comparison was measured by keeping the new perf tests and only reverting modules/imgproc/src/accum.simd.hpp for the baseline run. `-mstackrealign` was used for the MinGW Windows build to avoid stack-alignment issues with AVX code generation during local testing. ### Notes The implementation keeps the existing scalar fallback path unchanged. SIMD is used only for the newly covered masked `cn == 4` vectorizable part, and any remaining tail elements are still handled by `acc_general_()`. The improvement is most visible on larger images. Small images are dominated by overhead and do not always show meaningful speedup.
113 lines
4.8 KiB
C++
113 lines
4.8 KiB
C++
// This file is part of OpenCV project.
|
|
// It is subject to the license terms in the LICENSE file found in the top-level directory
|
|
// of this distribution and at http://opencv.org/license.html.
|
|
#include "perf_precomp.hpp"
|
|
|
|
namespace opencv_test {
|
|
|
|
typedef Size_MatType Accumulate;
|
|
|
|
#define MAT_TYPES_ACCUMLATE CV_8UC1, CV_16UC1, CV_32FC1
|
|
#define MAT_TYPES_ACCUMLATE_C MAT_TYPES_ACCUMLATE, CV_8UC3, CV_16UC3, CV_32FC3
|
|
#define MAT_TYPES_ACCUMLATE_D MAT_TYPES_ACCUMLATE, CV_64FC1
|
|
#define MAT_TYPES_ACCUMLATE_D_C MAT_TYPES_ACCUMLATE_C, CV_64FC1, CV_64FC1
|
|
|
|
#define PERF_ACCUMULATE_INIT(_FLTC) \
|
|
const Size srcSize = get<0>(GetParam()); \
|
|
const int srcType = get<1>(GetParam()); \
|
|
const int dstType = _FLTC(CV_MAT_CN(srcType)); \
|
|
Mat src1(srcSize, srcType), dst(srcSize, dstType); \
|
|
declare.in(src1, dst, WARMUP_RNG).out(dst);
|
|
|
|
#define PERF_ACCUMULATE_MASK_INIT(_FLTC) \
|
|
PERF_ACCUMULATE_INIT(_FLTC) \
|
|
Mat mask(srcSize, CV_8UC1); \
|
|
declare.in(mask, WARMUP_RNG);
|
|
|
|
#define PERF_TEST_P_ACCUMULATE(_NAME, _TYPES, _INIT, _FUN) \
|
|
PERF_TEST_P(Accumulate, _NAME, \
|
|
testing::Combine( \
|
|
testing::Values(sz1080p, sz720p, szVGA, szQVGA, szODD), \
|
|
testing::Values(_TYPES) \
|
|
) \
|
|
) \
|
|
{ \
|
|
_INIT \
|
|
TEST_CYCLE() _FUN; \
|
|
SANITY_CHECK_NOTHING(); \
|
|
}
|
|
|
|
/////////////////////////////////// Accumulate ///////////////////////////////////
|
|
|
|
PERF_TEST_P_ACCUMULATE(Accumulate, MAT_TYPES_ACCUMLATE,
|
|
PERF_ACCUMULATE_INIT(CV_32FC), accumulate(src1, dst))
|
|
|
|
PERF_TEST_P_ACCUMULATE(AccumulateMask, MAT_TYPES_ACCUMLATE_C,
|
|
PERF_ACCUMULATE_MASK_INIT(CV_32FC), accumulate(src1, dst, mask))
|
|
|
|
PERF_TEST_P_ACCUMULATE(AccumulateMask32FC4, CV_32FC4,
|
|
PERF_ACCUMULATE_MASK_INIT(CV_32FC), accumulate(src1, dst, mask))
|
|
|
|
PERF_TEST_P_ACCUMULATE(AccumulateMask8UC4To32FC4, CV_8UC4,
|
|
PERF_ACCUMULATE_MASK_INIT(CV_32FC), accumulate(src1, dst, mask))
|
|
|
|
PERF_TEST_P_ACCUMULATE(AccumulateDouble, MAT_TYPES_ACCUMLATE_D,
|
|
PERF_ACCUMULATE_INIT(CV_64FC), accumulate(src1, dst))
|
|
|
|
PERF_TEST_P_ACCUMULATE(AccumulateDoubleMask, MAT_TYPES_ACCUMLATE_D_C,
|
|
PERF_ACCUMULATE_MASK_INIT(CV_64FC), accumulate(src1, dst, mask))
|
|
|
|
///////////////////////////// AccumulateSquare ///////////////////////////////////
|
|
|
|
PERF_TEST_P_ACCUMULATE(Square, MAT_TYPES_ACCUMLATE,
|
|
PERF_ACCUMULATE_INIT(CV_32FC), accumulateSquare(src1, dst))
|
|
|
|
PERF_TEST_P_ACCUMULATE(SquareMask, MAT_TYPES_ACCUMLATE_C,
|
|
PERF_ACCUMULATE_MASK_INIT(CV_32FC), accumulateSquare(src1, dst, mask))
|
|
|
|
PERF_TEST_P_ACCUMULATE(SquareDouble, MAT_TYPES_ACCUMLATE_D,
|
|
PERF_ACCUMULATE_INIT(CV_64FC), accumulateSquare(src1, dst))
|
|
|
|
PERF_TEST_P_ACCUMULATE(SquareDoubleMask, MAT_TYPES_ACCUMLATE_D_C,
|
|
PERF_ACCUMULATE_MASK_INIT(CV_64FC), accumulateSquare(src1, dst, mask))
|
|
|
|
///////////////////////////// AccumulateProduct ///////////////////////////////////
|
|
|
|
#define PERF_ACCUMULATE_INIT_2(_FLTC) \
|
|
PERF_ACCUMULATE_INIT(_FLTC) \
|
|
Mat src2(srcSize, srcType); \
|
|
declare.in(src2);
|
|
|
|
#define PERF_ACCUMULATE_MASK_INIT_2(_FLTC) \
|
|
PERF_ACCUMULATE_MASK_INIT(_FLTC) \
|
|
Mat src2(srcSize, srcType); \
|
|
declare.in(src2);
|
|
|
|
PERF_TEST_P_ACCUMULATE(Product, MAT_TYPES_ACCUMLATE,
|
|
PERF_ACCUMULATE_INIT_2(CV_32FC), accumulateProduct(src1, src2, dst))
|
|
|
|
PERF_TEST_P_ACCUMULATE(ProductMask, MAT_TYPES_ACCUMLATE_C,
|
|
PERF_ACCUMULATE_MASK_INIT_2(CV_32FC), accumulateProduct(src1, src2, dst, mask))
|
|
|
|
PERF_TEST_P_ACCUMULATE(ProductDouble, MAT_TYPES_ACCUMLATE_D,
|
|
PERF_ACCUMULATE_INIT_2(CV_64FC), accumulateProduct(src1, src2, dst))
|
|
|
|
PERF_TEST_P_ACCUMULATE(ProductDoubleMask, MAT_TYPES_ACCUMLATE_D_C,
|
|
PERF_ACCUMULATE_MASK_INIT_2(CV_64FC), accumulateProduct(src1, src2, dst, mask))
|
|
|
|
///////////////////////////// AccumulateWeighted ///////////////////////////////////
|
|
|
|
PERF_TEST_P_ACCUMULATE(Weighted, MAT_TYPES_ACCUMLATE,
|
|
PERF_ACCUMULATE_INIT(CV_32FC), accumulateWeighted(src1, dst, 0.123))
|
|
|
|
PERF_TEST_P_ACCUMULATE(WeightedMask, MAT_TYPES_ACCUMLATE_C,
|
|
PERF_ACCUMULATE_MASK_INIT(CV_32FC), accumulateWeighted(src1, dst, 0.123, mask))
|
|
|
|
PERF_TEST_P_ACCUMULATE(WeightedDouble, MAT_TYPES_ACCUMLATE_D,
|
|
PERF_ACCUMULATE_INIT(CV_64FC), accumulateWeighted(src1, dst, 0.123456))
|
|
|
|
PERF_TEST_P_ACCUMULATE(WeightedDoubleMask, MAT_TYPES_ACCUMLATE_D_C,
|
|
PERF_ACCUMULATE_MASK_INIT(CV_64FC), accumulateWeighted(src1, dst, 0.123456, mask))
|
|
|
|
} // namespace
|