Fix ONNXRuntime dll path mismatch #29309
**1. C2664 build error in `net_impl_backend.cpp`**
`EnableProfiling()` expects `const wchar_t*` on Windows (`ORTCHAR_T`), but was passed `const char*`.
Fixed by converting to `std::wstring`, as suggested in #29278
---
**2. Wrong ORT DLL loaded at runtime (`modules/dnn/CMakeLists.txt`)**
With `DOWNLOAD_ONNXRUNTIME=ON`, the DLL glob only searched `bin/` but the downloaded package places DLLs in `lib/`. This left the build tree with no ORT DLL, causing Windows to fall back to the stale `System32\onnxruntime.dll` (1.17.1), crashing against the ORT 1.25.1 API.
Fixed by adding `lib/` as fallback and staging DLLs into the build bin directory at configure time.
Closes : #29278
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
dnn: fix new-engine block layout for RVV when VLEN > 128 #29304
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
<!-- Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the PR title. -->
#### Problem
Fixes#28852. The new DNN engine packs activations in a blocked NCHWc layout with a fixed channel-block size `C0 = 8`. The blocked-layout kernels (MaxPool, AveragePool, BatchNorm, depthwise Conv, ConvTranspose, generic Conv) load one block into a SIMD vector and assert: `C0 == nlanes || C0 == nlanes*2 || C0 % (nlanes*4) == 0`
On RISC-V the universal intrinsics are **scalable with LMUL=2**, so `nlanes = (VLEN/32)*2` — **16 at VLEN=256, 64 at VLEN=1024** — which exceeds the fixed `C0=8`. The assertion fails with `-215` and the ONNX pooling/conv conformance tests abort. #29180 worked around it by disabling the RVV SIMD path in those kernels.
#### Fix
Make the block size track the hardware vector width on RVV instead of assuming 8:
- `net_impl.cpp`: `defaultC0 = max(8, VTraits<v_float32>::vlanes())` (16/64 on RVV).
- The fp32 blocked kernels read `C0` at runtime and size scratch buffers by `max_nlanes` (bounded by the compile-time `CV_RVV_MAX_VLEN`, default 1024).
- All changes are guarded by `CV_SIMD_SCALABLE`, so **x86/NEON are unchanged**.
- The int8 quantized kernels are hardwired to `C0=8` (VNNI/NEON packing + per-channel quant) and are scalar on RVV, so quantized graphs are pinned to `C0=8` in `prepareForInference()` (detected via a `CV_8S` arg); only fp32 graphs use the wider block.
#### Validation (native RISC-V)
I have already conducted tests on the Spacemit K3, covering both the X100 (VLEN=256) and the A100 (VLEN=1024).
`opencv_test_dnn`, ENGINE_AUTO (new engine), no assertions:
| VLEN | Test_ONNX_conformance | Test_ONNX_layers (conv/pool/depthwise/quantized/MobileNet_v4) |
|------|-----------------------|--------------------------------------------------------------|
| 256 | 1641 / 1641 | all pass (incl. `Quantized_Convolution`) |
| 1024 | 1641 / 1641 | all pass (incl. `Quantized_Convolution`) |
[GSOC] feat: Add ALIKED feature extractor and LightGlue matcher with DNN integration #28986
## PR Description
### Summary
Integrate ALIKED and LightGlue into OpenCV's `features` module as native`Feature2D` and `DescriptorMatcher` implementations, enabling end-to-end neural feature matching within OpenCV's ecosystem.
---
### What's included
#### New classes
- **`cv::ALIKED`** extends `Feature2D`
- CNN-based keypoint detection
- 128-D descriptor extraction via ONNX Runtime
- **`cv::LightGlueMatcher`** extends `DescriptorMatcher`
- Deep feature matching with spatial context
- Uses keypoints and image sizes during matching
---
#### API design
- Standard OpenCV patterns:
- `detectAndCompute()`
- `match()`
- `knnMatch()`
- Multiple factory methods:
- ONNX model path
- In-memory model buffer
- Pre-loaded `dnn::Net`
- `Params` structs use `CV_EXPORTS_W_SIMPLE`
for Python/Java bindings support
- Optional DNN dependency:
- `HAVE_OPENCV_DNN` guards
- Stub implementations throw `StsNotImplemented`
---
### Files added
| File | Description |
|------|-------------|
| `src/feature2d_aliked.cpp` | ALIKED implementation |
| `src/matchers_lightglue.cpp` | LightGlueMatcher implementation |
| `src/aliked_context.hpp` | Shared internal context struct |
| `test/test_aliked_lightglue.cpp` | Unit tests (9 test cases) |
| `samples/cpp/example_features_aliked_lightglue.cpp` | Demo application |
---
### Files modified
- `CMakeLists.txt`
- Add `opencv_dnn` as optional dependency
- `features.hpp`
- Add ALIKED and LightGlueMatcher declarations
- `precomp.hpp`
- Add DNN include guard
---
### Usage
```cpp
// Feature extraction
Ptr<ALIKED> aliked =
ALIKED::create("aliked-n16rot-top1k-640.onnx");
vector<KeyPoint> kpts;
Mat descs;
aliked->detectAndCompute(image, Mat(), kpts, descs);
// Feature matching
Ptr<LightGlueMatcher> lg =
LightGlueMatcher::create("aliked_lightglue.onnx");
lg->setPairInfo(
kpts1Mat,
kpts2Mat,
img1.size(),
img2.size()
);
vector<DMatch> matches;
lg->match(descs1, descs2, matches);
````
please refer to samples/cpp/example_features_aliked_lightglue.cpp
---
### Test plan
* Build with `BUILD_LIST=features,dnn`
* Build without DNN:
* Verify stubs compile
* Verify `StsNotImplemented` is thrown
* Run:
* `ctest -R Features2d_ALIKED`
* `ctest -R Features2d_LightGlueMatcher`
* Run sample application with:
* Real images
* Real ONNX models
* Verify Python/Java bindings compile and work
---
### Related
Phase 1 of the
"End-to-End AI Feature Extraction and LightGlue Matching Pipeline"
GSoC project.
Designed to be extensible to:
* XFeat
* SuperPoint
* Other neural feature extractors
### test dependency
Depends on the opencv_extra PR adding ALIKED and LightGlue test models:
- [opencv_extra PR](https://github.com/opencv/opencv_extra/pull/1366)
This PR adds the following ONNX models to `download_models.py`:
- `aliked-n16rot-top1k-640.onnx`
- `aliked_lightglue.onnx`
These models are required for the `features2d` tests in the main OpenCV repository to validate the ALIKED and LightGlue feature extraction and matching pipeline.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
[FOLLOW UP] : Documentation optimizations for the new Sphinx structure #29220
### Pull Request Readiness Checklist
This PR serves as a follow-up to the new documentation system introduced in [#29206](https://github.com/opencv/opencv/pull/29206)
Co-authored by: @abhishek-gola @kirtijindal14 @Akansha-977 @Prasadayus @varun-jaiswal17
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Fixed Out-of-Memory issue and added VLM sample #29221
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Added DISK feature extractor support #29073
closes: https://github.com/opencv/opencv/issues/27083
Merge with: https://github.com/opencv/opencv_extra/pull/1368/
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Caffe importer cleanup #28678
Merge with: https://github.com/opencv/opencv_extra/pull/1324
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
- [x] The feature is well documented and sample code can be built with the project CMake
WeChatQR-fix conversion Caffe to ONNX #28746
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Temporarily disabled RVV intrinsics in several layers of the new DNN tested on musebook K1.
Kernels will be re-enabled when we find some good solution for the current 'm2' issue in rvv_scalable intrinsics.
This should fix#28852
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Update default YuNet model to new dynamic inputs #29107
Update the default model in `face_detect.py` and `face_detect.cpp` to
`face_detection_yunet_2026may.onnx`, which has symbolic `height`/`width` input dims.
## Changes
- `samples/dnn/face_detect.py`: update default `--face_detection_model` to `face_detection_yunet_2026may.onnx`
- `samples/dnn/face_detect.cpp`: update default `fd_model` to `face_detection_yunet_2026may.onnx`
Companion PR :
- https://github.com/opencv/opencv_zoo/pull/310
- https://github.com/opencv/opencv_extra/pull/1373
Closes : https://github.com/opencv/opencv/issues/28769
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Moved geometry transformations from imgproc to 3d, future geometry module #29101
The first step of 2d geometry operations migration to the future geometry module.
I created 2d.hpp to isolate the moved functions for now. I propose to create geometry.hpp when the module is renamed and include all things there.
OpenCV contrib: https://github.com/opencv/opencv_contrib/pull/4126
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [ ] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
Added SDPA layer (Scaled Dot Product Attention) #29104
Merge with: https://github.com/opencv/opencv_extra/pull/1374
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
fix MSVC error: use name constexp for alignas array size in INT8 kernel #29156
### Problem
Windows CI fails with MSVC error C2131 in `conv2_int8_kernels.simd.hpp`:
error C2131: expression did not evaluate to a constant
failure was caused by a read of a variable outside its lifetime
see usage of 'this'
Two array declarations inside `parallel_for_` lambdas used `constexpr`
variables defined inside the lambda body as array sizes:
alignas(32) int8_t wbuf[128 * K0]; // K0 defined inside lambda
alignas(32) int32_t sumbuf[SPAT_BLOCK_SIZE * K0]; // both defined inside lambda
### Fix
Move the combined size constants to function scope (before the lambda),
where they are true compile-time constants with no `this` involvement:
constexpr int WBUF_SIZE = 128 * 8; // 128 * K0
constexpr int SUMBUF_SIZE = 8 * 8; // SPAT_BLOCK_SIZE * K0
### Related
Fixes Windows CI failure introduced by #29126.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
<!-- Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the PR title. -->
Attention graph fusion with MLAS FlashAttention #29126
Performance numbers for Owl-v2 model on intel i9:
```
ORT: Average inference time over 10 runs: 1411.55 ms (min 1399.75, max 1438.89)
NEW: Average inference time over 10 runs: 1078 ms (min 1048.04, max 1110.61)
```
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
dnn int8 optimization #29079
all_layers.hpp
- Add float_input flag to Conv2Int8Params and Conv2Int8Layer to let the first conv accept raw FP32 input and quantize internally.
graph_fusion_qdq.cpp :
- Fuse DQ → Sigmoid → QL into SigmoidInt8, Similarly for MAxPool.
- Fuse the input QuantizeLinear node into the first Conv2Int8.
conv2_int8_layer.cpp
- Add quantizeInterleaveBlock()
conv2_int8_kernels.simd.hpp
- Add spatial tiling to both convInt8BlockVNNI and convInt8BlockDepthwise: splits output pixels into tiles so total task count is N × ngroups × Kblk × ntiles, fully utilizing all threads even when the channel count is small.
elementwise_layers.cpp
- Widen CV_Assert to accept CV_8U in addition to CV_8S.
eltwise2_int8_layer.cpp
- Add QLinearMul support: new Mul math path for both signed and unsigned int8.
- Add numpy-style broadcast support so QLinearMul / QLinearAdd with scalar
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Add KV cache with paged attention and prefetch #29127
Closes: https://github.com/opencv/opencv/issues/27159
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
FP32 KV Cache #28840
OpenCV Extra: https://github.com/opencv/opencv_extra/pull/1348
## This PR introduces basic (Paged ) KV-Cache to use on CPU
### Summary:
1. To ensure proper gemm-prepacking,
1.1. The Page Size of Key Cache is currently hardcoded as `FAST_GEMM_F32_NR`(which is 8, 12 or 16 depending on CPU architecture)
1.2. The Page Size of Values Cache is hardcoded as `FAST_GEMM_F32_PACKED_STRIDE_K`
2. there are two phases supported - prefill & generate.
2.1. prefill grows cache by `N` tokens and is allowed **only** for empty cache
2.2. generate grows cache by 1 token.
2.3. **Improtant**: it is currently not allowed to grow non-empty cache by more than one token at a time (thisbehaviour is sufficient for normal LLM querying, but should be extended if we want to implement speculative decoding)
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [ ] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
Dnn pow layer output type fix#29124
### Description
This pull request needs https://github.com/opencv/opencv_extra/pull/1372 to run.
**issue:** when running disk-lightglue `.onnx` model, such error message is thrown:
```bash
Error message OpenCV(5.0.0-pre) /home/tapakya/cpp/opencv_mod/opencv/modules/dnn/src/layers/nary_eltwise_layers.cpp:410: error: (-2:Unspecified error) in function 'virtual void cv::dnn::NaryEltwiseLayerImpl::getTypes(const std::vector<int>&, int, int, std::vector<int>&, std::vector<int>&) const'
> All inputs should have equal types (expected: 'inputs[0] == input'), where
> 'inputs[0]' is 11 (CV_64SC1)
> must be equal to
> 'input' is 5 (CV_32FC1)
```
This is because in `nary_eltwise_layers.cpp`, `getTypes` didn't set the output type of power operation correctly. `int`^`float` should output `float`, whereas `getTypes` set the output type to `int` in this case.
```c++
if (op == OPERATION::POW) {
CV_Assert(inputs.size() == 2);
auto isIntegerType = [](int t) {
return t == CV_8S || t == CV_8U || t == CV_16S || t == CV_16U || t == CV_32S || t == CV_32U || t == CV_64S || t == CV_64U;
};
auto isFloatType = [](int t) {
return t == CV_32F || t == CV_64F || t == CV_16F || t == CV_16BF;
};
int out_type;
const bool baseIsInt = isIntegerType(inputs[0]);
const bool expIsInt = isIntegerType(inputs[1]);
const bool baseIsFloat = isFloatType(inputs[0]);
const bool expIsFloat = isFloatType(inputs[1]);
if ((baseIsInt && expIsInt) || (baseIsFloat && expIsFloat))
{
out_type = (inputs[0] == inputs[1]) ? inputs[0] : CV_32F;
}
else if (baseIsFloat != expIsFloat)
{
out_type = inputs[0];
// if base is int and exp is float, output type should be float
}
```
**fix:** corrected the logic of determining output type
```c++
if ((baseIsInt && expIsInt) || (baseIsFloat && expIsFloat))
{
out_type = (inputs[0] == inputs[1]) ? inputs[0] : CV_32F;
}
else if (baseIsFloat && !expIsFloat)
{
out_type = inputs[0];
}
else if (!baseIsFloat && expIsFloat)
{
out_type = CV_32F;
}
```
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
dnn: skip Conv2Int8 fusion for grouped convs with Kg<8 (#28798) #28920
OpenCV Extra: https://github.com/opencv/opencv_extra/pull/1360
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another
license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
---
Fixes#28798.
The YuNet 2023mar int8 model from opencv_zoo doesn't detect anything on 5.x. With `--score_threshold
0.05` the detector still returns no boxes, so it's not just a confidence drop — the network output is
broken.
After dumping intermediate tensors and comparing against onnxruntime, the `obj_*` branches collapse to
~0 (saturated int8 `-128`), and the `cls_*` branches saturate the other way. Final score is `cls * obj`
so nothing ever crosses threshold.
The cause is in `Conv2Int8`'s VNNI kernel. It processes `K0 = 8` output channels per SIMD iteration and
writes them as one `K0`-wide block to `out + (n*K1 + k1)*planesize` with `k1 = k_base / K0`. When
`ngroups > 1` and `Kg = K/ngroups < K0`, consecutive groups share the same `k1` slot and overwrite each
other — only the last group's result survives. YuNet has six `1x1x3x3` depthwise conv heads (`Kg = 1`),
which is exactly this case.
The unfused `DequantizeLinear → Conv2 → QuantizeLinear` path is fine. The minimal fix here is to skip
the rewrite to `Conv2Int8` when `Kg < 8` so we fall back to the float path. A proper depthwise int8
kernel can be added later as a separate optimization.
Verified with `samples/dnn/face_detect.py` on Lena:
Before:
AssertionError: Cannot find a face in samples/data/lena.jpg
After:
Face 0, top-left coordinates: (204, 187), box width: 149, box height 212, score: 0.89
Cross-check vs onnxruntime: `cls_8` mean `0.673` (was `0.93`, ORT `0.673`), `obj_8` max `0.0039` (was
`0`, ORT `0.0039`).
The regression test is a single 16-group depthwise `QLinearConv` with non-zero `x_zp`, added to
`Quantized_Convolution`. Test data is in opencv_extra on the matching branch (~9 KB total).
<img width="1864" height="1060" alt="image" src="https://github.com/user-attachments/assets/a1b15c64-6b84-42b1-86a4-1d55474cb38c" />
Added MLAS third party module and integrated into GeMM path #28934
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
fix int32 overflow in shape_utils::total() for large tensors #29093
Fixes https://github.com/opencv/opencv/issues/24914
### Problem
When running inference with a ConvTranspose (deconvolution) layer on large inputs
(e.g. 4864×4864 with 30 channels), the DNN module crashes with:
OpenCV net_impl.cpp: error:
(expected: 'total(ints[i]) > 0'), where
'total(ints[i])' is -1455947776
must be greater than
'0' is 0
The root cause is `shape_utils::total()` which returns `int` (32-bit signed).
`ENGINE_CLASSIC` catches this via `CV_CheckGT` and throws.
`ENGINE_NEW` was silently bypassing the check — the overflow in `total()` itself
was never addressed.
### Changes
**`modules/dnn/include/opencv2/dnn/shape_utils.hpp`** — root fix
- Changed return type of both `total()` overloads from `int` to `size_t`
- Changed accumulator from `int elems = 1` to `size_t elems = 1`
**`modules/dnn/src/net_impl.cpp`**
- Updated `CV_CheckGT(total(...), 0)` to `CV_CheckGT(total(...), (size_t)0)`
to match the new return type
**`modules/dnn/src/net_impl2.cpp`**
- Added the same `CV_CheckGT` shape validation that `ENGINE_CLASSIC` has in
`net_impl.cpp:1333-1337` — `ENGINE_NEW` was missing this check entirely
**`modules/dnn/src/legacy_backend.hpp`**
- Removed the now-incorrect `(int)` cast in `CV_CheckEQ` — both sides are
now `size_t`
### Test
Added `Net.ShapeUtils_total_no_int32_overflow` in `modules/dnn/test/test_misc.cpp`:
- The shape [1920 × 1,478,656] is the exact im2col buffer from the bug report.
EXPECT_EQ verifies total() returns the correct size_t value 2,839,019,520.
EXPECT_LT documents that casting it to int wraps to -1,455,947,776 — the
value that caused the original crash.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
<!-- Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the PR title. -->
requires (YOLO26n , RetinaFace model download) : https://github.com/opencv/opencv_extra/pull/1357
added performance tests for 14 modern deep learning models to `modules/dnn/perf/perf_net.cpp.`
**Tests added**
- RT_DETR_L
- RF_DETR
- Grounding_DINO
- BlazeFace
- RetinaFace
- YOLO26n
- YOLO26m_Seg
- SegFormer_B2_Clothes
- Depth_Anything_V2
- SigLIP
- OWLv2
- RAFT
- SAM2_Encoder
- SAM2_Decoder
Models are available at: https://drive.google.com/drive/folders/1lpYAWSLMtxcmfDtYjiw0qwzZX0bCH-i4
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Add missing DataLayout constants in generated bindings for 5.x #29074
- Added DATA_LAYOUT_* constants to missing_consts so they
are manually injected into Core.java (same approach used for CV_8U, FILLED etc.)
- Added DataLayout entry to type_dict so the generator correctly maps
DataLayout
### Tests
modules/dnn/misc/java/test/DnnBlobFromImageWithParamsTest.java:
New test added:
- testDataLayoutConstants: verifies all DATA_LAYOUT_* constants are accessible from Core
Pre-existing tests enabled (were commented out earlier):
- testBlobFromImageWithParamsNHWCScalarScale: verifies blobFromImageWithParams
produces correct output with DATA_LAYOUT_NHWC and per-channel scalar scaling
- testBlobFromImageWithParams4chMultiImage: verifies blobFromImagesWithParams
correctly handles a batch of images with DATA_LAYOUT_NHWC layout
Closes : #27264
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
[Bug Fix] Slice2 empty-range crash in new DNN engine#29029
Closes: https://github.com/opencv/opencv/issues/29023
Merge with: https://github.com/opencv/opencv_extra/pull/1363
## Summary
ENGINE_NEW asserted `CV_Assert(outsz >= 0)` in `slice2_layer.cpp` when an ONNX Slice op had `start > end` with positive step. Per ONNX spec, such slices are valid and
produce an empty-range output (size 0). The hard assertion crashed SwinIR inference on ENGINE_NEW.
## Fix
- Replaced the fatal assertion with a graceful clamp: when `outsz < 0`, set `outsz = 0` and `end = start`.
- Moved `allEnds[axis] = end` to after the clamp so downstream shape inference receives a consistent value (the previous order propagated un-clamped garbage end values
into `normalize_axis()`, causing SIGSEGV).
## Test
Adds `Reproducibility_SwinIR_ONNX` accuracy test in `test_model.cpp`. Test data registered in opencv_extra PR with matching branch name `fix/dnn-slice2-empty-range`.
- Without fix: 3/3 FAIL with `slice2_layer.cpp:124: (-215:Assertion failed) outsz >= 0`
- With fix: 3/3 PASS (CPU, OCL, OCL_FP16)
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
SIMD Kernel speedup for DNN layers #28889
This PR add following speedups for Grounding Dino tiny model.
For Device: Intel(R) Core(TM) i9-14900KS, x86, 32 Cores, ubuntu 24.04,
| Model | `ENGINE_NEW (Before)` | `ENGINE_NEW (After)` | `ENGINE_ORT` |
| :--- | :--- | :--- | :--- |
| **Grounding Dino Tiny** | 3130 ms | 1872.06 ms| 1800.18 ms|
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
Fixed gatherND offset calculation #29084
Bug fix: fixed segmentation fault when running ALIKED onnx model with ENGINE_NEW.
### Description
issue: when running ALIKED onnx model with ENGINE_NEW, segmentation fault happens during the inference.
fix: fixed the offset calculation in gatherND.cpp, so gather does not access out-of-bound memory.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
KleidiCV update to verison 26.03 #28744
KleidiCV release: https://gitlab.arm.com/kleidi/kleidicv/-/releases/26.03
Tuned DNN test threshold as resize linear produces slightly different result with the new KleidiCV version.
### Pull Request Readiness Checklist
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [ ] The PR is proposed to the proper branch
- [ ] There is a reference to the original bug report and related work
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
Patch to opencv_extra has the same branch name.
- [ ] The feature is well documented and sample code can be built with the project CMake
randomnormallike_layer.cpp registered itself via CV_DNN_REGISTER_LAYER_CLASS_STATIC at file scope. Nothing else in the translation unit was referenced from elsewhere, so when OpenCV is built statically (BUILD_SHARED_LIBS=OFF, as in the reporter's MSVC build) the linker drops the object file and the static-init registration never runs. The ONNX importer then sets layerParams.type = "RandomNormalLike" but getLayerInstance has no factory for that type, and parsing fails with "Can't create layer of type RandomNormalLike".
Switch to the standard pattern used by every other layer in the module: declare RandomNormalLikeLayer in all_layers.hpp, expose a static create() factory from the .cpp, and register it explicitly from initializeLayerFactory in init.cpp. Because init.cpp is referenced by the DNN module init path, this pulls in the layer .cpp regardless of build mode and the registration always runs.
The failure was incorrectly attributed to MSVC in the bug report. The bug is build-mode sensitive (static vs shared), not platform sensitive.
Verified locally on linux/gcc-13 by building modules/dnn and
opencv_test_dnn against the patched tree, then running opencv_test_dnn --gtest_filter='Test_ONNX_layers.RandomNormalLike_basic/0:Test_ONNX_layers.RandomNormalLike_complex/0'
Both tests pass; without the patch the same binary reproduces the
"Can't create layer of type RandomNormalLike" error from the report.