1
0
mirror of https://github.com/opencv/opencv.git synced 2026-07-30 15:53:03 +04:00

Merge pull request #29126 from abhishek-gola:flash_attention

Attention graph fusion with MLAS FlashAttention #29126

Performance numbers for Owl-v2 model on intel i9:
```
ORT: Average inference time over 10 runs: 1411.55 ms (min 1399.75, max 1438.89)
NEW: Average inference time over 10 runs: 1078 ms (min 1048.04, max 1110.61)
```
### Pull Request Readiness Checklist

See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request

- [x] I agree to contribute to the project under Apache 2 License.
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
- [x] The PR is proposed to the proper branch
- [x] There is a reference to the original bug report and related work
- [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
      Patch to opencv_extra has the same branch name.
- [x] The feature is well documented and sample code can be built with the project CMake
This commit is contained in:
Abhishek Gola
2026-05-27 12:02:27 +05:30
committed by GitHub
parent 2405556989
commit bf0cf34963
11 changed files with 2049 additions and 41 deletions
+1 -1
View File
@@ -408,7 +408,7 @@ size_t
#if defined(__aarch64__) && defined(__linux__)
typedef size_t(MLASCALL MLAS_SBGEMM_FLOAT_KERNEL)(
const float* A,
const bfloat16_t* B,
const void* B,
float* C,
size_t CountK,
size_t CountM,