1
0
mirror of https://github.com/opencv/opencv.git synced 2026-07-21 19:33:03 +04:00

Merge pull request #29094 from feitianduowen:4.x

### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake

<!-- Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the PR title. -->

### Summary

This PR improves SIMD coverage for masked `cv::accumulate()` on 4-channel images.

The existing masked accumulate SIMD implementation has special paths for `cn == 1` and `cn == 3`, while 4-channel images fall back to the general implementation. This PR adds `cn == 4` SIMD paths using OpenCV Universal Intrinsics.

Compare link:

https://github.com/opencv/opencv/compare/4.x...feitianduowen:opencv:4.x

Modified files:

- `modules/imgproc/src/accum.simd.hpp`
  - Add masked `cn == 4` SIMD paths for:
    - `CV_32FC4 -> CV_32FC4`
    - `CV_8UC4 -> CV_32FC4`
- `modules/video/test/test_accum.cpp`
  - Add deterministic correctness regression tests for the new masked 4-channel accumulate paths.
- `modules/imgproc/perf/perf_accumulate.cpp`
  - Add performance coverage for masked 4-channel accumulate.
imgproc: add SIMD paths for masked 4-channel accumulate #29094
  
### Implementation

The new SIMD branches are added in:

- `modules/imgproc/src/accum.simd.hpp`

They handle:

- `src`: `CV_32FC4`, `dst`: `CV_32FC4`, `mask`: `CV_8UC1`
- `src`: `CV_8UC4`, `dst`: `CV_32FC4`, `mask`: `CV_8UC1`

For the `CV_32FC4` path, the implementation uses `v_load_deinterleave`, applies the pixel mask to all 4 channels, accumulates into `dst`, and stores the result with `v_store_interleave`.

For the `CV_8UC4 -> CV_32FC4` path, the implementation loads 4 interleaved uchar channels, applies the pixel mask, widens the values to float, accumulates into `dst`, and stores the results back as interleaved `CV_32FC4`.

The masked SIMD block is guarded with:

```cpp
if (x <= len - cVectorWidth)
{
    ...
}
```
This check ensures that SIMD setup and vector operations are only executed when at least one full vector iteration can run. For very small inputs or tail-only cases, the function skips SIMD initialization and lets the existing scalar fallback handle the data through `acc_general_(src, dst, mask, len, cn, x)`.

This keeps tail handling unchanged and avoids unnecessary SIMD initialization when the vector loop would be skipped.

### Accuracy tests

Added a deterministic regression test in:

* `modules/video/test/test_accum.cpp`

New test:

```bash
Video_Acc.accuracy_32FC4_masked_cn4
Video_Acc.accuracy_8UC4_to_32FC4_masked_cn4
```

Test commands:

```bash
opencv_test_video --gtest_filter=Video_Acc.accuracy_32FC4_masked_cn4
opencv_test_video --gtest_filter=Video_Acc.accuracy_8UC4_to_32FC4_masked_cn4
opencv_test_video --gtest_filter="Video_Acc.*:Video_AccSquared.*:Video_AccProduct.*:Video_RunningAvg.*"
```
The filtered accumulate-related tests passed locally.

I added dedicated deterministic tests instead of only extending the existing randomized accumulate tests because this PR only changes the masked `cv::accumulate()` paths for specific 4-channel type combinations. It does not add `cn == 4` SIMD coverage for `accumulateSquare`, `accumulateProduct`, or `accumulateWeighted`.

The existing base accumulate tests are shared by several accumulation functions. Extending the common randomized channel selection to include `cn == 4` would also affect tests for functions that are not optimized by this PR. The new deterministic tests directly cover the modified paths:

* `cv::accumulate(src, dst, mask)`
* `CV_32FC4 -> CV_32FC4`
* `CV_8UC4 -> CV_32FC4`
* `mask`: `CV_8UC1`

The test also covers small sizes and non-vector-multiple sizes, so both the new SIMD path and the existing scalar tail fallback are exercised.

### Performance tests

Added performance coverage in:

* `modules/imgproc/perf/perf_accumulate.cpp`

Perf command:

```bash
opencv_perf_imgproc --gtest_filter="*AccumulateMask32FC4*" --perf_min_samples=1000 --perf_force_samples=1000
opencv_perf_imgproc --gtest_filter="*AccumulateMask8UC4To32FC4*" --perf_min_samples=1000 --perf_force_samples=1000
```

Test environment:

* Release build
* AVX2 enabled
* IPP disabled
* OpenCL disabled
* Windows MinGW build

Performance results:

`CV_32FC4 -> CV_32FC4`

| Size      | Before median | After median | Speedup |
| --------- | ------------- | ------------ | ------- |
| 1920x1080 | 3.07 ms       | 2.74 ms      | 1.12x   |
| 1280x720  | 0.84 ms       | 0.73 ms      | 1.15x   |
| 640x480   | 0.17 ms       | 0.16 ms      | 1.06x   |
| 320x240   | 0.04 ms       | 0.04 ms      | ~1.00x  |

`CV_8UC4 -> CV_32FC4`

| Size      | Before median | After median | Speedup |
| --------- | ------------- | ------------ | ------- |
| 1920x1080 | 6.09 ms       | 1.75 ms      | 3.48x   |
| 1280x720  | 2.69 ms       | 0.82 ms      | 3.28x   |
| 640x480   | 0.96 ms       | 0.28 ms      | 3.43x   |
| 320x240   | 0.24 ms       | 0.06 ms      | 4.00x   |
| 127x61    | 0.02 ms       | 0.01 ms      | 2.00x   |

### Build and test commands

The following commands were run from Windows `cmd.exe`.

```sh
cmake -S . -B build_release -G Ninja -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTS=ON -DBUILD_PERF_TESTS=OFF -DBUILD_EXAMPLES=OFF -DWITH_IPP=OFF -DWITH_OPENCL=OFF -DCMAKE_C_FLAGS=-mstackrealign -DCMAKE_CXX_FLAGS=-mstackrealign

cmake --build build_release --target opencv_test_video -j 4

build_release\bin\opencv_test_video.exe --gtest_filter=Video_Acc.accuracy_32FC4_masked_cn4
build_release\bin\opencv_test_video.exe --gtest_filter=Video_Acc.accuracy_8UC4_to_32FC4_masked_cn4
build_release\bin\opencv_test_video.exe --gtest_filter=Video_Acc.*:Video_AccSquared.*:Video_AccProduct.*:Video_RunningAvg.*
```

Accuracy test commands passed locally.

For the performance comparison, I kept the new perf test in `modules/imgproc/perf/perf_accumulate.cpp` and only reverted `modules/imgproc/src/accum.simd.hpp` to measure the baseline. Then I restored the SIMD patch and measured the optimized version.

Baseline measurement:

```sh
git diff -- modules/imgproc/src/accum.simd.hpp > accum_cn4_simd.patch
git checkout -- modules/imgproc/src/accum.simd.hpp

mkdir perf_logs

cmake -S . -B build_perf_before -G Ninja -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTS=OFF -DBUILD_PERF_TESTS=ON -DBUILD_EXAMPLES=OFF -DWITH_IPP=OFF -DWITH_OPENCL=OFF -DCMAKE_C_FLAGS=-mstackrealign -DCMAKE_CXX_FLAGS=-mstackrealign > perf_logs\before_cmake_config.txt 2>&1
cmake --build build_perf_before --target opencv_perf_imgproc -j 4 > perf_logs\before_build.txt 2>&1

build_perf_before\bin\opencv_perf_imgproc.exe --gtest_filter=*AccumulateMask32FC4* --perf_min_samples=1000 --perf_force_samples=1000 > perf_logs\before_perf.txt 2>&1
```

```sh
git diff -- modules/imgproc/src/accum.simd.hpp > accum_cn4_u8_simd.patch
git checkout -- modules/imgproc/src/accum.simd.hpp

cmake -S . -B build_perf_before_u8 -G Ninja -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTS=OFF -DBUILD_PERF_TESTS=ON -DBUILD_EXAMPLES=OFF -DWITH_IPP=OFF -DWITH_OPENCL=OFF -DCMAKE_C_FLAGS=-mstackrealign -DCMAKE_CXX_FLAGS=-mstackrealign
cmake --build build_perf_before_u8 --target opencv_perf_imgproc -j 4
build_perf_before_u8\bin\opencv_perf_imgproc.exe --gtest_filter=*AccumulateMask8UC4To32FC4* --perf_min_samples=1000 --perf_force_samples=1000 > perf_logs\before_perf1.txt 2>&1
```

Optimized measurement:

```sh
git apply accum_cn4_simd.patch

cmake -S . -B build_perf_after -G Ninja -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTS=OFF -DBUILD_PERF_TESTS=ON -DBUILD_EXAMPLES=OFF -DWITH_IPP=OFF -DWITH_OPENCL=OFF -DCMAKE_C_FLAGS=-mstackrealign -DCMAKE_CXX_FLAGS=-mstackrealign > perf_logs\after_cmake_config.txt 2>&1
cmake --build build_perf_after --target opencv_perf_imgproc -j 4 > perf_logs\after_build.txt 2>&1
build_perf_after\bin\opencv_perf_imgproc.exe --gtest_filter=*AccumulateMask32FC4* --perf_min_samples=500 --perf_force_samples=500 > perf_logs\after_perf.txt 2>&1
```

```sh
git apply accum_cn4_u8_simd.patch

cmake -S . -B build_perf_after_u8 -G Ninja -DCMAKE_BUILD_TYPE=Release -DBUILD_TESTS=OFF -DBUILD_PERF_TESTS=ON -DBUILD_EXAMPLES=OFF -DWITH_IPP=OFF -DWITH_OPENCL=OFF -DCMAKE_C_FLAGS=-mstackrealign -DCMAKE_CXX_FLAGS=-mstackrealign
cmake --build build_perf_after_u8 --target opencv_perf_imgproc -j 4
build_perf_after_u8\bin\opencv_perf_imgproc.exe --gtest_filter=*AccumulateMask8UC4To32FC4* --perf_min_samples=1000 --perf_force_samples=1000 > perf_logs\after_perf1.txt 2>&1
```

Performance comparison was measured by keeping the new perf tests and only reverting modules/imgproc/src/accum.simd.hpp for the baseline run.

`-mstackrealign` was used for the MinGW Windows build to avoid stack-alignment issues with AVX code generation during local testing.

### Notes

The implementation keeps the existing scalar fallback path unchanged. SIMD is used only for the newly covered masked `cn == 4` vectorizable part, and any remaining tail elements are still handled by `acc_general_()`.

The improvement is most visible on larger images. Small images are dominated by overhead and do not always show meaningful speedup.
This commit is contained in:
Siyu Wang
2026-05-28 17:44:16 +08:00
committed by GitHub
parent 3bb68212da
commit a61ff0fa81
3 changed files with 382 additions and 101 deletions
+6
View File
@@ -45,6 +45,12 @@ PERF_TEST_P_ACCUMULATE(Accumulate, MAT_TYPES_ACCUMLATE,
PERF_TEST_P_ACCUMULATE(AccumulateMask, MAT_TYPES_ACCUMLATE_C,
PERF_ACCUMULATE_MASK_INIT(CV_32FC), accumulate(src1, dst, mask))
PERF_TEST_P_ACCUMULATE(AccumulateMask32FC4, CV_32FC4,
PERF_ACCUMULATE_MASK_INIT(CV_32FC), accumulate(src1, dst, mask))
PERF_TEST_P_ACCUMULATE(AccumulateMask8UC4To32FC4, CV_8UC4,
PERF_ACCUMULATE_MASK_INIT(CV_32FC), accumulate(src1, dst, mask))
PERF_TEST_P_ACCUMULATE(AccumulateDouble, MAT_TYPES_ACCUMLATE_D,
PERF_ACCUMULATE_INIT(CV_64FC), accumulate(src1, dst))
+265 -101
View File
@@ -317,79 +317,183 @@ void acc_simd_(const uchar* src, float* dst, const uchar* mask, int len, int cn)
}
else
{
v_uint8 v_0 = vx_setall_u8(0);
if (cn == 1)
if (x <= len - cVectorWidth)
{
for ( ; x <= len - cVectorWidth; x += cVectorWidth)
v_uint8 v_0 = vx_setall_u8(0);
if (cn == 1)
{
v_uint8 v_mask = vx_load(mask + x);
v_mask = v_not(v_eq(v_0, v_mask));
v_uint8 v_src = vx_load(src + x);
v_src = v_and(v_src, v_mask);
v_uint16 v_src0, v_src1;
v_expand(v_src, v_src0, v_src1);
for ( ; x <= len - cVectorWidth; x += cVectorWidth)
{
v_uint8 v_mask = vx_load(mask + x);
v_mask = v_ne(v_0, v_mask);
v_uint8 v_src = vx_load(src + x);
v_src = v_and(v_src, v_mask);
v_uint16 v_src0, v_src1;
v_expand(v_src, v_src0, v_src1);
v_uint32 v_src00, v_src01, v_src10, v_src11;
v_expand(v_src0, v_src00, v_src01);
v_expand(v_src1, v_src10, v_src11);
v_uint32 v_src00, v_src01, v_src10, v_src11;
v_expand(v_src0, v_src00, v_src01);
v_expand(v_src1, v_src10, v_src11);
v_store(dst + x, v_add(vx_load(dst + x), v_cvt_f32(v_reinterpret_as_s32(v_src00))));
v_store(dst + x + step, v_add(vx_load(dst + x + step), v_cvt_f32(v_reinterpret_as_s32(v_src01))));
v_store(dst + x + step * 2, v_add(vx_load(dst + x + step * 2), v_cvt_f32(v_reinterpret_as_s32(v_src10))));
v_store(dst + x + step * 3, v_add(vx_load(dst + x + step * 3), v_cvt_f32(v_reinterpret_as_s32(v_src11))));
v_store(dst + x, v_add(vx_load(dst + x), v_cvt_f32(v_reinterpret_as_s32(v_src00))));
v_store(dst + x + step, v_add(vx_load(dst + x + step), v_cvt_f32(v_reinterpret_as_s32(v_src01))));
v_store(dst + x + step * 2, v_add(vx_load(dst + x + step * 2), v_cvt_f32(v_reinterpret_as_s32(v_src10))));
v_store(dst + x + step * 3, v_add(vx_load(dst + x + step * 3), v_cvt_f32(v_reinterpret_as_s32(v_src11))));
}
}
}
else if (cn == 3)
{
for ( ; x <= len - cVectorWidth; x += cVectorWidth)
else if (cn == 3)
{
v_uint8 v_mask = vx_load(mask + x);
v_mask = v_not(v_eq(v_0, v_mask));
v_uint8 v_src0, v_src1, v_src2;
v_load_deinterleave(src + (x * cn), v_src0, v_src1, v_src2);
v_src0 = v_and(v_src0, v_mask);
v_src1 = v_and(v_src1, v_mask);
v_src2 = v_and(v_src2, v_mask);
v_uint16 v_src00, v_src01, v_src10, v_src11, v_src20, v_src21;
v_expand(v_src0, v_src00, v_src01);
v_expand(v_src1, v_src10, v_src11);
v_expand(v_src2, v_src20, v_src21);
for ( ; x <= len - cVectorWidth; x += cVectorWidth)
{
v_uint8 v_mask = vx_load(mask + x);
v_mask = v_ne(v_0, v_mask);
v_uint8 v_src0, v_src1, v_src2;
v_load_deinterleave(src + (x * cn), v_src0, v_src1, v_src2);
v_src0 = v_and(v_src0, v_mask);
v_src1 = v_and(v_src1, v_mask);
v_src2 = v_and(v_src2, v_mask);
v_uint16 v_src00, v_src01, v_src10, v_src11, v_src20, v_src21;
v_expand(v_src0, v_src00, v_src01);
v_expand(v_src1, v_src10, v_src11);
v_expand(v_src2, v_src20, v_src21);
v_uint32 v_src000, v_src001, v_src010, v_src011;
v_uint32 v_src100, v_src101, v_src110, v_src111;
v_uint32 v_src200, v_src201, v_src210, v_src211;
v_expand(v_src00, v_src000, v_src001);
v_expand(v_src01, v_src010, v_src011);
v_expand(v_src10, v_src100, v_src101);
v_expand(v_src11, v_src110, v_src111);
v_expand(v_src20, v_src200, v_src201);
v_expand(v_src21, v_src210, v_src211);
v_uint32 v_src000, v_src001, v_src010, v_src011;
v_uint32 v_src100, v_src101, v_src110, v_src111;
v_uint32 v_src200, v_src201, v_src210, v_src211;
v_expand(v_src00, v_src000, v_src001);
v_expand(v_src01, v_src010, v_src011);
v_expand(v_src10, v_src100, v_src101);
v_expand(v_src11, v_src110, v_src111);
v_expand(v_src20, v_src200, v_src201);
v_expand(v_src21, v_src210, v_src211);
v_float32 v_dst000, v_dst001, v_dst010, v_dst011;
v_float32 v_dst100, v_dst101, v_dst110, v_dst111;
v_float32 v_dst200, v_dst201, v_dst210, v_dst211;
v_load_deinterleave(dst + (x * cn), v_dst000, v_dst100, v_dst200);
v_load_deinterleave(dst + ((x + step) * cn), v_dst001, v_dst101, v_dst201);
v_load_deinterleave(dst + ((x + step * 2) * cn), v_dst010, v_dst110, v_dst210);
v_load_deinterleave(dst + ((x + step * 3) * cn), v_dst011, v_dst111, v_dst211);
v_float32 v_dst000, v_dst001, v_dst010, v_dst011;
v_float32 v_dst100, v_dst101, v_dst110, v_dst111;
v_float32 v_dst200, v_dst201, v_dst210, v_dst211;
v_load_deinterleave(dst + (x * cn), v_dst000, v_dst100, v_dst200);
v_load_deinterleave(dst + ((x + step) * cn), v_dst001, v_dst101, v_dst201);
v_load_deinterleave(dst + ((x + step * 2) * cn), v_dst010, v_dst110, v_dst210);
v_load_deinterleave(dst + ((x + step * 3) * cn), v_dst011, v_dst111, v_dst211);
v_dst000 = v_add(v_dst000, v_cvt_f32(v_reinterpret_as_s32(v_src000)));
v_dst100 = v_add(v_dst100, v_cvt_f32(v_reinterpret_as_s32(v_src100)));
v_dst200 = v_add(v_dst200, v_cvt_f32(v_reinterpret_as_s32(v_src200)));
v_dst001 = v_add(v_dst001, v_cvt_f32(v_reinterpret_as_s32(v_src001)));
v_dst101 = v_add(v_dst101, v_cvt_f32(v_reinterpret_as_s32(v_src101)));
v_dst201 = v_add(v_dst201, v_cvt_f32(v_reinterpret_as_s32(v_src201)));
v_dst010 = v_add(v_dst010, v_cvt_f32(v_reinterpret_as_s32(v_src010)));
v_dst110 = v_add(v_dst110, v_cvt_f32(v_reinterpret_as_s32(v_src110)));
v_dst210 = v_add(v_dst210, v_cvt_f32(v_reinterpret_as_s32(v_src210)));
v_dst011 = v_add(v_dst011, v_cvt_f32(v_reinterpret_as_s32(v_src011)));
v_dst111 = v_add(v_dst111, v_cvt_f32(v_reinterpret_as_s32(v_src111)));
v_dst211 = v_add(v_dst211, v_cvt_f32(v_reinterpret_as_s32(v_src211)));
v_dst000 = v_add(v_dst000, v_cvt_f32(v_reinterpret_as_s32(v_src000)));
v_dst100 = v_add(v_dst100, v_cvt_f32(v_reinterpret_as_s32(v_src100)));
v_dst200 = v_add(v_dst200, v_cvt_f32(v_reinterpret_as_s32(v_src200)));
v_dst001 = v_add(v_dst001, v_cvt_f32(v_reinterpret_as_s32(v_src001)));
v_dst101 = v_add(v_dst101, v_cvt_f32(v_reinterpret_as_s32(v_src101)));
v_dst201 = v_add(v_dst201, v_cvt_f32(v_reinterpret_as_s32(v_src201)));
v_dst010 = v_add(v_dst010, v_cvt_f32(v_reinterpret_as_s32(v_src010)));
v_dst110 = v_add(v_dst110, v_cvt_f32(v_reinterpret_as_s32(v_src110)));
v_dst210 = v_add(v_dst210, v_cvt_f32(v_reinterpret_as_s32(v_src210)));
v_dst011 = v_add(v_dst011, v_cvt_f32(v_reinterpret_as_s32(v_src011)));
v_dst111 = v_add(v_dst111, v_cvt_f32(v_reinterpret_as_s32(v_src111)));
v_dst211 = v_add(v_dst211, v_cvt_f32(v_reinterpret_as_s32(v_src211)));
v_store_interleave(dst + (x * cn), v_dst000, v_dst100, v_dst200);
v_store_interleave(dst + ((x + step) * cn), v_dst001, v_dst101, v_dst201);
v_store_interleave(dst + ((x + step * 2) * cn), v_dst010, v_dst110, v_dst210);
v_store_interleave(dst + ((x + step * 3) * cn), v_dst011, v_dst111, v_dst211);
v_store_interleave(dst + (x * cn), v_dst000, v_dst100, v_dst200);
v_store_interleave(dst + ((x + step) * cn), v_dst001, v_dst101, v_dst201);
v_store_interleave(dst + ((x + step * 2) * cn), v_dst010, v_dst110, v_dst210);
v_store_interleave(dst + ((x + step * 3) * cn), v_dst011, v_dst111, v_dst211);
}
}
else if (cn == 4)
{
v_uint8 v_zero = vx_setzero_u8();
for (; x <= len - cVectorWidth; x += cVectorWidth)
{
v_uint8 v_mask = vx_load(mask + x);
v_mask = v_ne(v_mask, v_zero);
v_uint8 v_src0, v_src1, v_src2, v_src3;
v_load_deinterleave(src + x * cn,
v_src0, v_src1, v_src2, v_src3);
v_src0 = v_and(v_src0, v_mask);
v_src1 = v_and(v_src1, v_mask);
v_src2 = v_and(v_src2, v_mask);
v_src3 = v_and(v_src3, v_mask);
v_uint16 v_src0_u16_0, v_src0_u16_1;
v_uint16 v_src1_u16_0, v_src1_u16_1;
v_uint16 v_src2_u16_0, v_src2_u16_1;
v_uint16 v_src3_u16_0, v_src3_u16_1;
v_expand(v_src0, v_src0_u16_0, v_src0_u16_1);
v_expand(v_src1, v_src1_u16_0, v_src1_u16_1);
v_expand(v_src2, v_src2_u16_0, v_src2_u16_1);
v_expand(v_src3, v_src3_u16_0, v_src3_u16_1);
v_uint32 v_src0_u32_0, v_src0_u32_1, v_src0_u32_2, v_src0_u32_3;
v_uint32 v_src1_u32_0, v_src1_u32_1, v_src1_u32_2, v_src1_u32_3;
v_uint32 v_src2_u32_0, v_src2_u32_1, v_src2_u32_2, v_src2_u32_3;
v_uint32 v_src3_u32_0, v_src3_u32_1, v_src3_u32_2, v_src3_u32_3;
v_expand(v_src0_u16_0, v_src0_u32_0, v_src0_u32_1);
v_expand(v_src0_u16_1, v_src0_u32_2, v_src0_u32_3);
v_expand(v_src1_u16_0, v_src1_u32_0, v_src1_u32_1);
v_expand(v_src1_u16_1, v_src1_u32_2, v_src1_u32_3);
v_expand(v_src2_u16_0, v_src2_u32_0, v_src2_u32_1);
v_expand(v_src2_u16_1, v_src2_u32_2, v_src2_u32_3);
v_expand(v_src3_u16_0, v_src3_u32_0, v_src3_u32_1);
v_expand(v_src3_u16_1, v_src3_u32_2, v_src3_u32_3);
v_float32 v_src0_f0 = v_cvt_f32(v_reinterpret_as_s32(v_src0_u32_0));
v_float32 v_src0_f1 = v_cvt_f32(v_reinterpret_as_s32(v_src0_u32_1));
v_float32 v_src0_f2 = v_cvt_f32(v_reinterpret_as_s32(v_src0_u32_2));
v_float32 v_src0_f3 = v_cvt_f32(v_reinterpret_as_s32(v_src0_u32_3));
v_float32 v_src1_f0 = v_cvt_f32(v_reinterpret_as_s32(v_src1_u32_0));
v_float32 v_src1_f1 = v_cvt_f32(v_reinterpret_as_s32(v_src1_u32_1));
v_float32 v_src1_f2 = v_cvt_f32(v_reinterpret_as_s32(v_src1_u32_2));
v_float32 v_src1_f3 = v_cvt_f32(v_reinterpret_as_s32(v_src1_u32_3));
v_float32 v_src2_f0 = v_cvt_f32(v_reinterpret_as_s32(v_src2_u32_0));
v_float32 v_src2_f1 = v_cvt_f32(v_reinterpret_as_s32(v_src2_u32_1));
v_float32 v_src2_f2 = v_cvt_f32(v_reinterpret_as_s32(v_src2_u32_2));
v_float32 v_src2_f3 = v_cvt_f32(v_reinterpret_as_s32(v_src2_u32_3));
v_float32 v_src3_f0 = v_cvt_f32(v_reinterpret_as_s32(v_src3_u32_0));
v_float32 v_src3_f1 = v_cvt_f32(v_reinterpret_as_s32(v_src3_u32_1));
v_float32 v_src3_f2 = v_cvt_f32(v_reinterpret_as_s32(v_src3_u32_2));
v_float32 v_src3_f3 = v_cvt_f32(v_reinterpret_as_s32(v_src3_u32_3));
v_float32 v_dst0, v_dst1, v_dst2, v_dst3;
v_load_deinterleave(dst + x * cn, v_dst0, v_dst1, v_dst2, v_dst3);
v_store_interleave(dst + x * cn,
v_add(v_dst0, v_src0_f0),
v_add(v_dst1, v_src1_f0),
v_add(v_dst2, v_src2_f0),
v_add(v_dst3, v_src3_f0));
v_load_deinterleave(dst + (x + step) * cn, v_dst0, v_dst1, v_dst2, v_dst3);
v_store_interleave(dst + (x + step) * cn,
v_add(v_dst0, v_src0_f1),
v_add(v_dst1, v_src1_f1),
v_add(v_dst2, v_src2_f1),
v_add(v_dst3, v_src3_f1));
v_load_deinterleave(dst + (x + 2 * step) * cn, v_dst0, v_dst1, v_dst2, v_dst3);
v_store_interleave(dst + (x + 2 * step) * cn,
v_add(v_dst0, v_src0_f2),
v_add(v_dst1, v_src1_f2),
v_add(v_dst2, v_src2_f2),
v_add(v_dst3, v_src3_f2));
v_load_deinterleave(dst + (x + 3 * step) * cn, v_dst0, v_dst1, v_dst2, v_dst3);
v_store_interleave(dst + (x + 3 * step) * cn,
v_add(v_dst0, v_src0_f3),
v_add(v_dst1, v_src1_f3),
v_add(v_dst2, v_src2_f3),
v_add(v_dst3, v_src3_f3));
}
}
}
}
@@ -500,49 +604,109 @@ void acc_simd_(const float* src, float* dst, const uchar* mask, int len, int cn)
}
else
{
v_float32 v_0 = vx_setzero_f32();
if (cn == 1)
if (x <= len - cVectorWidth)
{
for ( ; x <= len - cVectorWidth ; x += cVectorWidth)
v_float32 v_0 = vx_setzero_f32();
if (cn == 1)
{
v_uint16 v_masku16 = vx_load_expand(mask + x);
v_uint32 v_masku320, v_masku321;
v_expand(v_masku16, v_masku320, v_masku321);
v_float32 v_mask0 = v_reinterpret_as_f32(v_not(v_eq(v_masku320, v_reinterpret_as_u32(v_0))));
v_float32 v_mask1 = v_reinterpret_as_f32(v_not(v_eq(v_masku321, v_reinterpret_as_u32(v_0))));
for ( ; x <= len - cVectorWidth ; x += cVectorWidth)
{
v_uint16 v_masku16 = vx_load_expand(mask + x);
v_uint32 v_masku320, v_masku321;
v_expand(v_masku16, v_masku320, v_masku321);
v_float32 v_mask0 = v_reinterpret_as_f32(v_not(v_eq(v_masku320, v_reinterpret_as_u32(v_0))));
v_float32 v_mask1 = v_reinterpret_as_f32(v_not(v_eq(v_masku321, v_reinterpret_as_u32(v_0))));
v_store(dst + x, v_add(vx_load(dst + x), v_and(vx_load(src + x), v_mask0)));
v_store(dst + x + step, v_add(vx_load(dst + x + step), v_and(vx_load(src + x + step), v_mask1)));
v_store(dst + x, v_add(vx_load(dst + x), v_and(vx_load(src + x), v_mask0)));
v_store(dst + x + step, v_add(vx_load(dst + x + step), v_and(vx_load(src + x + step), v_mask1)));
}
}
else if (cn == 3)
{
for ( ; x <= len - cVectorWidth ; x += cVectorWidth)
{
v_uint16 v_masku16 = vx_load_expand(mask + x);
v_uint32 v_masku320, v_masku321;
v_expand(v_masku16, v_masku320, v_masku321);
v_float32 v_mask0 = v_reinterpret_as_f32(v_not(v_eq(v_masku320, v_reinterpret_as_u32(v_0))));
v_float32 v_mask1 = v_reinterpret_as_f32(v_not(v_eq(v_masku321, v_reinterpret_as_u32(v_0))));
v_float32 v_src00, v_src01, v_src10, v_src11, v_src20, v_src21;
v_load_deinterleave(src + x * cn, v_src00, v_src10, v_src20);
v_load_deinterleave(src + (x + step) * cn, v_src01, v_src11, v_src21);
v_src00 = v_and(v_src00, v_mask0);
v_src01 = v_and(v_src01, v_mask1);
v_src10 = v_and(v_src10, v_mask0);
v_src11 = v_and(v_src11, v_mask1);
v_src20 = v_and(v_src20, v_mask0);
v_src21 = v_and(v_src21, v_mask1);
v_float32 v_dst00, v_dst01, v_dst10, v_dst11, v_dst20, v_dst21;
v_load_deinterleave(dst + x * cn, v_dst00, v_dst10, v_dst20);
v_load_deinterleave(dst + (x + step) * cn, v_dst01, v_dst11, v_dst21);
v_store_interleave(dst + x * cn, v_add(v_dst00, v_src00), v_add(v_dst10, v_src10), v_add(v_dst20, v_src20));
v_store_interleave(dst + (x + step) * cn, v_add(v_dst01, v_src01), v_add(v_dst11, v_src11), v_add(v_dst21, v_src21));
}
}
else if (cn == 4)
{
for (; x <= len - cVectorWidth; x += cVectorWidth)
{
v_uint16 v_masku16 = vx_load_expand(mask + x);
v_uint32 v_masku320, v_masku321;
v_expand(v_masku16, v_masku320, v_masku321);
v_float32 v_mask0 = v_reinterpret_as_f32(v_ne(v_masku320, v_reinterpret_as_u32(v_0)));
v_float32 v_mask1 = v_reinterpret_as_f32(v_ne(v_masku321, v_reinterpret_as_u32(v_0)));
v_float32 v_src00, v_src01;
v_float32 v_src10, v_src11;
v_float32 v_src20, v_src21;
v_float32 v_src30, v_src31;
v_load_deinterleave(src + x * cn,
v_src00, v_src10, v_src20, v_src30);
v_load_deinterleave(src + (x + step) * cn,
v_src01, v_src11, v_src21, v_src31);
v_src00 = v_and(v_src00, v_mask0);
v_src10 = v_and(v_src10, v_mask0);
v_src20 = v_and(v_src20, v_mask0);
v_src30 = v_and(v_src30, v_mask0);
v_src01 = v_and(v_src01, v_mask1);
v_src11 = v_and(v_src11, v_mask1);
v_src21 = v_and(v_src21, v_mask1);
v_src31 = v_and(v_src31, v_mask1);
v_float32 v_dst00, v_dst01;
v_float32 v_dst10, v_dst11;
v_float32 v_dst20, v_dst21;
v_float32 v_dst30, v_dst31;
v_load_deinterleave(dst + x * cn, v_dst00, v_dst10, v_dst20, v_dst30);
v_load_deinterleave(dst + (x + step) * cn, v_dst01, v_dst11, v_dst21, v_dst31);
v_store_interleave(dst + x * cn,
v_add(v_dst00, v_src00),
v_add(v_dst10, v_src10),
v_add(v_dst20, v_src20),
v_add(v_dst30, v_src30));
v_store_interleave(dst + (x + step) * cn,
v_add(v_dst01, v_src01),
v_add(v_dst11, v_src11),
v_add(v_dst21, v_src21),
v_add(v_dst31, v_src31));
}
}
}
else if (cn == 3)
{
for ( ; x <= len - cVectorWidth ; x += cVectorWidth)
{
v_uint16 v_masku16 = vx_load_expand(mask + x);
v_uint32 v_masku320, v_masku321;
v_expand(v_masku16, v_masku320, v_masku321);
v_float32 v_mask0 = v_reinterpret_as_f32(v_not(v_eq(v_masku320, v_reinterpret_as_u32(v_0))));
v_float32 v_mask1 = v_reinterpret_as_f32(v_not(v_eq(v_masku321, v_reinterpret_as_u32(v_0))));
v_float32 v_src00, v_src01, v_src10, v_src11, v_src20, v_src21;
v_load_deinterleave(src + x * cn, v_src00, v_src10, v_src20);
v_load_deinterleave(src + (x + step) * cn, v_src01, v_src11, v_src21);
v_src00 = v_and(v_src00, v_mask0);
v_src01 = v_and(v_src01, v_mask1);
v_src10 = v_and(v_src10, v_mask0);
v_src11 = v_and(v_src11, v_mask1);
v_src20 = v_and(v_src20, v_mask0);
v_src21 = v_and(v_src21, v_mask1);
v_float32 v_dst00, v_dst01, v_dst10, v_dst11, v_dst20, v_dst21;
v_load_deinterleave(dst + x * cn, v_dst00, v_dst10, v_dst20);
v_load_deinterleave(dst + (x + step) * cn, v_dst01, v_dst11, v_dst21);
v_store_interleave(dst + x * cn, v_add(v_dst00, v_src00), v_add(v_dst10, v_src10), v_add(v_dst20, v_src20));
v_store_interleave(dst + (x + step) * cn, v_add(v_dst01, v_src01), v_add(v_dst11, v_src11), v_add(v_dst21, v_src21));
}
}
}
#endif // CV_SIMD
acc_general_(src, dst, mask, len, cn, x);
@@ -3106,4 +3270,4 @@ CV_CPU_OPTIMIZATION_NAMESPACE_END
} // namespace cv
///* End of file. */
///* End of file. */
+111
View File
@@ -246,4 +246,115 @@ TEST(Video_AccSquared, accuracy) { CV_SquareAccTest test; test.safe_run(); }
TEST(Video_AccProduct, accuracy) { CV_MultiplyAccTest test; test.safe_run(); }
TEST(Video_RunningAvg, accuracy) { CV_RunningAvgTest test; test.safe_run(); }
typedef testing::TestWithParam<tuple<Size, int, int> > Video_Acc_Cn4;
TEST_P(Video_Acc_Cn4, accuracy)
{
const Size size = get<0>(GetParam());
const int pattern = get<1>(GetParam());
const int srcType = get<2>(GetParam());
RNG& rng = theRNG();
Mat src(size, srcType);
Mat dst(size, CV_32FC4);
Mat mask(size, CV_8UC1);
if (srcType == CV_8UC4)
rng.fill(src, RNG::UNIFORM, Scalar::all(0), Scalar::all(256));
else
rng.fill(src, RNG::UNIFORM, Scalar::all(-10.0), Scalar::all(10.0));
rng.fill(dst, RNG::UNIFORM, Scalar::all(-1000.0), Scalar::all(1000.0));
for (int y = 0; y < mask.rows; ++y)
{
uchar* row = mask.ptr<uchar>(y);
for (int x = 0; x < mask.cols; ++x)
{
switch (pattern)
{
case 0:
row[x] = 0;
break;
case 1:
row[x] = 255;
break;
case 2:
row[x] = ((x + y) % 2) ? 255 : 0;
break;
case 3:
row[x] = ((x * 13 + y * 7) % 5) ? 255 : 0;
break;
default:
row[x] = ((x * 17 + y * 11) % 3) ? 255 : 0;
break;
}
}
}
Mat dstRef = dst.clone();
if (srcType == CV_32FC4)
{
for (int y = 0; y < src.rows; ++y)
{
const Vec4f* srcRow = src.ptr<Vec4f>(y);
Vec4f* dstRefRow = dstRef.ptr<Vec4f>(y);
const uchar* maskRow = mask.ptr<uchar>(y);
for (int x = 0; x < src.cols; ++x)
{
if (maskRow[x])
{
for (int c = 0; c < 4; ++c)
dstRefRow[x][c] += srcRow[x][c];
}
}
}
}
else
{
CV_Assert(srcType == CV_8UC4);
for (int y = 0; y < src.rows; ++y)
{
const Vec4b* srcRow = src.ptr<Vec4b>(y);
Vec4f* dstRefRow = dstRef.ptr<Vec4f>(y);
const uchar* maskRow = mask.ptr<uchar>(y);
for (int x = 0; x < src.cols; ++x)
{
if (maskRow[x])
{
for (int c = 0; c < 4; ++c)
dstRefRow[x][c] += static_cast<float>(srcRow[x][c]);
}
}
}
}
cv::accumulate(src, dst, mask);
const double err = cv::norm(dst, dstRef, NORM_INF);
EXPECT_EQ(0.0, err)
<< "size=" << size
<< ", pattern=" << pattern
<< ", srcType=" << srcType;
}
INSTANTIATE_TEST_CASE_P(Accumulate,
Video_Acc_Cn4,
testing::Combine(
testing::Values(Size(1, 1),
Size(3, 5),
Size(17, 7),
Size(37, 19),
Size(128, 16),
Size(641, 37)),
testing::Values(0, 1, 2, 3, 4),
testing::Values(CV_32FC4, CV_8UC4)));
}} // namespace