mirror of
https://github.com/opencv/opencv.git
synced 2026-07-25 21:33:04 +04:00
d95badefa4
Parallelize DNN layers using chunking #28821 At a high level, I replaced tensor-level parallelism with chunk-level parallelism. Previously, parallel_for_ was dispatched over the number of input or output tensors i.e. one thread handled one whole tensor's copy. The new approach precomputes each tensor's destination offset and per-slice size upfront, then slices the total byte work into fixed 64 KB chunks. The full chunk count is handed to parallel_for_ as a single flat range, and each worker decodes its chunk index back into (tensor, slice, byte_offset) using a prefix-sum table before running a plain memcpy on its piece. A small-size threshold falls back to the sequential path so we don't get threading overhead on small tensors. Performance numbers after these optimizations: For Device: Intel(R) Core(TM) i9-14900KS, x86, 32 Cores, ubuntu 24.04, | Model | `ENGINE_NEW` | `ENGINE_ORT` | | :--- | :--- | :--- | | **YOLOv8n** |10.9 ms| 12.15 ms| | **YOLOv5n** | 8.36 ms| 9.23 ms| | **YOLOX-S** | 23.46 ms| 25.16 ms| For Device: Macbook M1 Air | Model | `ENGINE_NEW` | `ENGINE_ORT` | | :--- | :--- | :--- | | **YOLOv8n** |34.45 ms| 42.52 ms| | **YOLOv5n** | 31.62 ms| 25.52 ms| | **YOLOX-S** | 88.9 ms| 116.7 ms| ### Pull Request Readiness Checklist See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request - [x] I agree to contribute to the project under Apache 2 License. - [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV - [x] The PR is proposed to the proper branch - [x] There is a reference to the original bug report and related work - [x] There is accuracy test, performance test and test data in opencv_extra repository, if applicable Patch to opencv_extra has the same branch name. - [x] The feature is well documented and sample code can be built with the project CMake