Filter2D IPP extraction to HAL for 5.x #29428
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
dnn: silence false-positive CPU target warning for New graph engine #29457
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
* [x] I agree to contribute to the project under Apache 2 License.
* [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
* [x] The PR is proposed to the proper branch (`5.x`)
* [ ] There is a reference to the original bug report and related work
No existing OpenCV issue or pull request directly covers this warning. This PR is the original report and fix.
* [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Not applicable. This patch changes only diagnostic behavior for an existing CPU no-op path.
* [ ] The feature is well documented and sample code can be built with the project CMake
Not applicable. This patch adds no feature or public API.
## Summary
Suppress a misleading warning emitted when OpenCV 5 New DNN graph engine receives its default CPU target request.
`FaceDetectorYN` and `FaceRecognizerSF` call:
```cpp
net.setPreferableTarget(DNN_TARGET_CPU);
```
after loading their networks.
When New graph engine is active, target switching is not supported. OpenCV therefore currently prints:
```text
Targets are not supported by the new graph engine for now
```
However, New graph engine already executes on CPU. The CPU request does not require a target change, is ignored, and inference continues through New graph engine.
The warning therefore suggests a failed configuration or fallback to Classic DNN even though neither occurs.
## Change
Treat `DNN_TARGET_CPU` as a silent no-op when generic New graph engine is active.
```text
New graph engine + CPU target
→ no target change needed
→ return unchanged
→ no warning
New graph engine + non-CPU target
→ target remains unsupported
→ preserve existing warning
→ return unchanged
```
## Behavior before
```text
New graph engine active
→ caller requests CPU
→ request is ignored
→ warning emitted
→ inference continues on New graph engine
```
## Behavior after
```text
New graph engine active
→ caller requests CPU
→ request is ignored
→ no warning
→ inference continues on New graph engine
```
## Unchanged behavior
* New graph engine selection is unchanged.
* Inference execution is unchanged.
* CPU remains the effective target for generic New graph engine.
* Classic DNN behavior is unchanged.
* ONNX Runtime handling is unchanged.
* Non-CPU targets continue to emit the existing warning.
## Motivation
The current diagnostic is a false positive for the default CPU request. It reports unsupported target selection even though CPU is already the active execution target and inference succeeds through New graph engine.
## Testing
* Built OpenCV locally.
* Ran relevant DNN tests.
* Verified New graph engine still loads and executes affected models.
* Verified CPU target requests no longer emit the misleading warning.
* Verified non-CPU target requests retain the existing unsupported-target warning.
Merge pull request #29414 from Prasadayus:box_filter_refactor
Moving IPP functions to HAL for box_filter in Imgproc #29414
**Performance Numbers on Intel(R) Core(TM) i9-11900K:** https://docs.google.com/spreadsheets/d/1puWmOSTtAFwWjPu8J1SpuccPQvxUJU8NlFQs7mAZwL8/edit?usp=sharing
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Refactoring and optimizing cvtcolor in Imgproc #29389
**Key Changes:**
1. Unified three divergent SIMD swap implementations into a single shared `v_swap` helper.
2. Consolidated the repeated `CV_8U/CV_16U/CV_32F` depth-dispatch chains into a single `CvtColorLoopDepth` template.
3. Extracted the duplicated coefficient-selection loops in the YCrCb/YUV constructors into `selectYuvCoeffs`.
4. Collapsed the shared YUV420 store logic duplicated across both decode invokers into `storeYUV420block`.
5. Generalized the inline blue-channel coefficient swaps in `color_lab.cpp` into `swapBlueCoeffsCols`/`swapBlueCoeffsRows`.
6. Added a `v_dotprod`-based SIMD path (`v_RGB2Y`/`v_RGB2UV`) for the previously scalar `RGB8toYUV422Invoker`.
7. Migrated all inline `CV_IPP_CHECK` cvtColor paths out of the dispatch files and into the IPP HAL plugin (`hal/ipp/src/color_ipp.cpp`), replacing each call site with `CALL_HAL`.
**Performance Numbers on Intel(R) Core(TM) i9-11900K**: https://docs.google.com/spreadsheets/d/1pz0aHlTeG4Ao8RT7LipgdNSjhCJz92QfZG3GmNa63hU/edit?usp=sharing
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Extended Attention layer support #29333
Implemented present/past KV support.
Merge with: https://github.com/opencv/opencv_extra/pull/1381
Co-authored by: @Akansha-977
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
dnn: add Winograd F(6,3) RVV implementation for RISC-V (VLEN≥256) #29411
### Summary
Adds a RISC-V Vector (RVV) backend for the Winograd F(6,3) convolution path
in the DNN module, targeting VLEN≥256 processors (tested on SpacemiT K1,
rv64gcv).
### Background
The Winograd F(6,3) path was already gated on `useSIMD128 || useAVX ||
useAVX2 || useNEON` at runtime, and protected by the same set of compile-time
macros. RVV was simply missing from both guards, so it always fell through to
generic convolution even when the hardware supports it.
A second issue: `conv_winograd_f63` was registered with
`ocv_add_dispatched_file`, which skips ISA variants not present in the host
toolchain during cross-compilation. Changing to `ocv_add_dispatched_file_force_all`
forces the `.rvv.cpp` translation unit to be generated unconditionally, matching
how every other RVV-enabled kernel in the DNN module is registered.
### Implementation notes
**Atom width.** `vsetvlmax_e32m1()` returns 8 on VLEN=256, so `winoAtomF32=8`
is chosen, matching the AVX2 atom width. The `impl_accum_F32` and transform
functions are structured identically to the AVX2 path (4 output channels ×
6 input tiles per atom).
**Input/output transform.** AVX2 uses `_mm256_unpacklo/hi_ps` + `permute2f128`
for an in-register 8×8 transpose. RVV has no equivalent cross-lane shuffle at
this width without `vrgather`, which adds index-vector overhead for a
non-bottleneck step. Instead the 8×8 intermediate matrix is transposed
scalar-in-memory between the two `wino_bt8x8_rvv` / `wino_at8x6_rvv` passes.
This keeps the transform code simple and correct; the GEMM in `impl_accum_F32`
dominates runtime.
**VLEN guard.** `getWinofunc_F32` checks `vsetvlmax_e32m1() >= 8` at runtime
and returns an empty functor on narrower implementations, so the code is safe
on VLEN=128 targets without a separate code path.
### Testing
**Correctness:** `ConvolutionWinograd.Accuracy` passes on the K1 board (9 ms).
**Performance:** `opencv_perf_dnn`, SpacemiT K1 (rv64gcv, VLEN=256),
GCC 13, `-O2 -march=rv64gcv`. Reported times are per-iteration medians from
`[ PERFSTAT ]` output (`--perf_min_samples=5`). "Generic" is the same build
with Winograd disabled via `OPENCV_DNN_DISABLE_WINOGRAD=1`.
| Input → Output C | Winograd (ms) | Generic (ms) | Speedup |
|--------------------------|--------------|-------------|---------|
| {1,128,52,52} → 256 | 35.97 | 62.86 | 1.75× |
| {1,512,13,13} → 1024 | 65.18 | 112.50 | 1.73× |
| {1,256,75,75} → 256 | 159.75 | 271.65 | 1.70× |
| {1,64,300,300} → 64 | 174.21 | 295.32 | 1.69× |
| {1,1152,16,16} → 1152 | 176.12 | 245.33 | 1.39× |
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
dnn: implement LpPool ONNX operator #29217
## Summary
Implements the `LpPool` ONNX operator (opset 1–18), which was previously unregistered and caused a parse failure. `LpPool` computes the Lp-norm pooling: `(sum(|x|^p))^(1/p)` over a sliding window.
## Changes
- New `LpPoolLayer` in `modules/dnn/src/layers/lppool_layer.cpp`
- Supports `kernel_shape`, `strides`, `dilations`, `pads`, `auto_pad` (NOTSET/SAME_UPPER), `ceil_mode`, and `p` (default 2)
- SIMD fast paths for p=1 (abs + accumulate) and p=2 (square + accumulate + sqrt); scalar fallback for other values of p
- Registered `LpPool` dispatch entry in both `onnx_importer.cpp` (classic engine) and `onnx_importer2.cpp` (new graph engine)
- Added `LpPoolLayer` declaration to `modules/dnn/include/opencv2/dnn/all_layers.hpp`
- Registered layer class in `modules/dnn/src/init.cpp`
- Re-enabled 8 lppool conformance tests in `test_onnx_conformance.cpp` (previously in parser denylist)
- `test_lppool_2d_same_lower` added to the global conformance denylist — same known SAME_LOWER padding bug that affects `averagepool` and `maxpool`
## Testing
All applicable ONNX conformance tests pass:
| Test | Result |
|------|--------|
| test_lppool_1d_default | PASSED |
| test_lppool_2d_default | PASSED |
| test_lppool_2d_dilations | PASSED |
| test_lppool_2d_pads | PASSED |
| test_lppool_2d_same_lower | SKIPPED (known SAME_LOWER padding bug, consistent with avgpool/maxpool) |
| test_lppool_2d_same_upper | PASSED |
| test_lppool_2d_strides | PASSED |
| test_lppool_3d_default | PASSED |
Tested on: macOS (x86_64/SSE4, Rosetta 2) and Linux x86_64 (AVX2/AVX-512, GCC 13.3.0), Release build
OpenCV version: 5.0.0-pre
## Related Issues
None
---
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
Port of #29377 to 5.x. Wires Log, Erf, Exp, Sin, Cos, Sinh, Cosh, Tan,
Softplus, BNLL, Asinh, Acosh, Atanh into the dispatched activation_kernels
registry, plus a Layer_Activation perf test. 1.4-12.3x across
M4/A76/Threadripper/Xeon, correct to <=5e-7 vs scalar.
Validate onnx tensor payload size in getMatFromTensor (5.x) #29370
5.x port of #29314. The DNN ONNX importer diverged here, so the change is ported manually.
`getMatFromTensor()` sizes the output blob from `TensorProto.dims` and then reads that many elements out of the tensor payload, but the payload is sized independently in the model and never checked against the shape. In 5.x the payload reaches the read through three sources: a typed `*_data` field, `raw_data`, or external-file data routed through `getTensorRAWData()`. An initializer whose dims claim more elements than the payload holds makes the `Mat` copy/convert read past the buffer.
### Before
A FLOAT tensor declaring `[1000000]` backed by a 4-byte `raw_data` reads about 4 MB out of bounds (ASan flags a heap over-read in `cv::Mat::copyTo`). The same mismatch is present in every datatype branch and for both the typed-field and `rawdata` sources, reachable from `readNetFromONNX`/`readNetFromONNXBuffer` for every initializer.
### After
Compute the element count the shape implies once, then check each payload source holds at least that many elements before the read. `getTensorRAWData()` now reports the byte count it returns (covering both `raw_data` and external-file data), and `getMatFromTensor()` validates the typed fields and the raw byte count per datatype. A short payload throws a clear `cv::Exception` instead of over-reading.
### Tradeoffs
The check lives in `getMatFromTensor` because that is the one place the shape and the payload meet, so every initializer and attribute tensor is covered without each caller repeating it. The added cost is one multiply-accumulate over the dims plus a comparison per tensor at load time; well-formed models, where the payload already matches the shape, are unaffected. The element-count accumulation saturates on overflow so an oversized shape can't wrap to a small total and slip past the check.
### Pull Request Readiness Checklist
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
- [ ] The feature is well documented and sample code can be built with the project CMake
fixed Dynamic quantized linear layer error #29386
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
dnn: vectorize fp32 convolution on RISC-V RVV #29357
### Summary
Vectorizes the new DNN engine's fp32 convolution on RISC-V **RVV**. Convolution was the last major block-layout operator still running scalar on RVV; this brings it to parity with the already-vectorized depthwise/pool/batchnorm path.
Single file, **purely additive** (`#elif CV_SIMD_SCALABLE` branches only) — x86/AVX2, ARM/NEON and AArch64 codegen are unchanged.
### Background
- **#28585** introduced block-layout conv with fp32 kernels specialized for AVX2 (`C0=8`) and NEON, scalar `#else` for everything else, and explicitly noted *"we could add the respective kernels later."*
- On RVV that `#else` meant conv ran **scalar at every VLEN** — the C0=8 fixed width approach doesn't transfer to a variable-VLEN ISA.
- **#29304** made the block size track the hardware width (`C0 = vlanes()`), fixing the VLEN≥256 assertion but leaving conv scalar.
- **This PR** supplies the deferred RVV conv vectorization, on top of that `C0=vlanes()` foundation.
### Approach
- Re-enable the `SPAT_BLOCK_SIZE=10` **blocked path** on RVV. Because `K0 == vlanes()`, one `v_float32` accumulator covers a full output-channel block, so 10 accumulators process 10 output positions together. A single K0-wide weight `vx_load` is reused across all 10 — amortizing the per-output-point load that bottlenecks the scalar path (this is what enables multi-core scaling; the scalar/per-FMA-load path is bandwidth-bound).
- Vectorize the **scalar tail** (the `<10`-position remainder) the same way.
- **Vector-length-agnostic:** `C0` is runtime, the kernel is written in terms of `v_float32`/`vlanes()`, so the same binary runs at full width on VLEN 128/256/512/1024 with no recompile. No per-VLEN kernels.
- The six `C0=8` specialized kernels stay `#if !CV_SIMD_SCALABLE`-gated (they can't run at `C0=vlanes()`); on RVV all shapes use the generic kernel.
### Performance
Convolution throughput, blocked (this PR) vs scalar, same machine/engine (VLEN=256, 8× rv64 @ 2.4 GHz), GFLOP/s:
| layer | scalar 1T | this 1T | scalar 8T | this 8T | 8T speedup |
|---|---|---|---|---|---|
| 3×3, 256→256, 64² | 1.13 | 13.23 | 8.28 | 75.62 | **9.1×** |
| 1×1, 256→256, 64² | 1.08 | 11.73 | 8.02 | 71.70 | 8.9× |
| 3×3, 128→128, 128² | 1.10 | 13.02 | 8.45 | 89.14 | **10.6×** |
≈ **12× single-thread, 9–11× at 8 threads**. Convolution dominates CNN inference time, so this is a large end-to-end win on RVV hardware.
### Validation
K3-class RVV board, VLEN=256: `Test_ONNX_layers` **264/264**, `Test_ONNX_conformance` **1641/1641**, **0 assertions**. Output is bit-identical to the scalar path (same accumulation order, just vectorized over `K0`).
Vector-length-agnosticism confirmed by running the VLEN=256 binary unchanged at **VLEN=1024**.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Extended support for Resize and Split layers #29323
Merge with: https://github.com/opencv/opencv_extra/pull/1379
Key changes:
Resize layer:
_antialias_ (linear & cubic) :- PIL-style separable resampling with stretched filter support and edge-clamped, renormalized weights.
_axes_ (incl. reversed [3,2]) :- getOutShape, the scale override, and runtime _tf_crop_and_resize_ ROI parsing now map 2-element sizes/scales/roi by the axes order instead of assuming [2,3].
_keep_aspect_ratio_policy_ (not_larger/not_smaller) :- output size from min/max per-axis scale.
_half_pixel_symmetric_ :- new coordinate-transform mode + importer mapping.
_align_corners_ downsampling :- coordinate scale uses the unfloored scaled length (in−1)/(in·x_scale−1).
Split Layer:
_convertTo empty 1-D Mat:_ the empty-Mat branch collapsed a 1-D [0] to 2-D [1,0] via cv::Size(); now uses allowTransposed like the non-empty path.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
predictOptimalVectorWidth() built its vectorWidths table with 8 entries
(depths CV_8U..CV_16F), but checkOptimalVectorWidth() indexes it by depth.
The 5.x depths CV_16BF, CV_Bool, CV_64U, CV_64S and CV_32U therefore read
past the end of the array; the garbage value can slip past the ckercn <= 0
guard, and the divider normalization loop then shifts the divider to zero,
so "offsets[i] % dividers[i]" raises SIGFPE. The failure is allocation
dependent and shows up as a sequence-dependent crash, e.g. in the OpenCL
Norm tests for CV_32U (cv::norm on a UMat reaches this via ocl_sum, which
calls predictOptimalVectorWidth before its own depth guard bails out).
Size the table to CV_DEPTH_MAX so every depth is in bounds and map the new
fixed-size integer depths to their natural OpenCL vector widths; CV_16BF
has no OpenCL vector type and stays scalar. Add a regression test covering
all depths.
MlasThreadedBufAlloc() and its ThreadedBufHolder guarded the
_aligned_malloc / _aligned_free path on _MSC_VER. Non-MSVC Windows
toolchains (MinGW, clang) therefore fell through to posix_memalign(),
which the Windows CRT does not provide, so the build failed with
"'posix_memalign' was not declared in this scope".
Guard the aligned-allocation path on _WIN32 instead, in both the holder
declaration (mlasi.h) and definition (platform.cpp) so the unique_ptr
deleter type stays consistent across translation units, and include
<malloc.h> on Windows since MinGW only declares the _aligned_* functions
there. The non-Windows posix_memalign / aligned_alloc fallback is
unchanged.
Fixes#29350
Fix ONNXRuntime dll path mismatch #29309
**1. C2664 build error in `net_impl_backend.cpp`**
`EnableProfiling()` expects `const wchar_t*` on Windows (`ORTCHAR_T`), but was passed `const char*`.
Fixed by converting to `std::wstring`, as suggested in #29278
---
**2. Wrong ORT DLL loaded at runtime (`modules/dnn/CMakeLists.txt`)**
With `DOWNLOAD_ONNXRUNTIME=ON`, the DLL glob only searched `bin/` but the downloaded package places DLLs in `lib/`. This left the build tree with no ORT DLL, causing Windows to fall back to the stale `System32\onnxruntime.dll` (1.17.1), crashing against the ORT 1.25.1 API.
Fixed by adding `lib/` as fallback and staging DLLs into the build bin directory at configure time.
Closes : #29278
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake