mirror of
https://github.com/opencv/opencv.git
synced 2026-07-21 19:33:03 +04:00
Compare commits
14 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 4de7015cfd | |||
| d42d34e2ee | |||
| 04d4af1a6d | |||
| cb64fdabb5 | |||
| 8cb0a37336 | |||
| aae15a4efa | |||
| b61b146f0e | |||
| 905c551b7b | |||
| 7c4398a0f6 | |||
| b3eb24cd07 | |||
| 5b5c2093bb | |||
| 010cc2e58a | |||
| 2f23737e66 | |||
| 43cb28bdb5 |
@@ -1,3 +0,0 @@
|
||||
# These are supported funding model platforms
|
||||
|
||||
github: opencv
|
||||
+11
-44
@@ -1,15 +1,23 @@
|
||||
<!--
|
||||
If you have a question rather than reporting a bug please go to https://forum.opencv.org where you get much faster responses.
|
||||
If you have a question rather than reporting a bug please go to http://answers.opencv.org where you get much faster responses.
|
||||
If you need further assistance please read [How To Contribute](https://github.com/opencv/opencv/wiki/How_to_contribute).
|
||||
|
||||
Please:
|
||||
|
||||
* Read the documentation to test with the latest developer build.
|
||||
* Check if other person has already created the same issue to avoid duplicates. You can comment on it if there already is an issue.
|
||||
* Try to be as detailed as possible in your report.
|
||||
* Report only one problem per created issue.
|
||||
|
||||
|
||||
This is a template helping you to create an issue which can be processed as quickly as possible. This is the bug reporting section for the OpenCV library.
|
||||
-->
|
||||
|
||||
##### System information (version)
|
||||
<!-- Example
|
||||
- OpenCV => 4.2
|
||||
- OpenCV => 3.1
|
||||
- Operating System / Platform => Windows 64 Bit
|
||||
- Compiler => Visual Studio 2017
|
||||
- Compiler => Visual Studio 2015
|
||||
-->
|
||||
|
||||
- OpenCV => :grey_question:
|
||||
@@ -28,44 +36,3 @@ This is a template helping you to create an issue which can be processed as quic
|
||||
```
|
||||
or attach as .txt or .zip file
|
||||
-->
|
||||
|
||||
##### Issue submission checklist
|
||||
|
||||
- [ ] I report the issue, it's not a question
|
||||
<!--
|
||||
OpenCV team works with forum.opencv.org, Stack Overflow and other communities
|
||||
to discuss problems. Tickets with questions without a real issue statement will be
|
||||
closed.
|
||||
-->
|
||||
- [ ] I checked the problem with documentation, FAQ, open issues,
|
||||
forum.opencv.org, Stack Overflow, etc and have not found any solution
|
||||
<!--
|
||||
Places to check:
|
||||
* OpenCV documentation: https://docs.opencv.org
|
||||
* FAQ page: https://github.com/opencv/opencv/wiki/FAQ
|
||||
* OpenCV forum: https://forum.opencv.org
|
||||
* OpenCV issue tracker: https://github.com/opencv/opencv/issues?q=is%3Aissue
|
||||
* Stack Overflow branch: https://stackoverflow.com/questions/tagged/opencv
|
||||
-->
|
||||
- [ ] I updated to the latest OpenCV version and the issue is still there
|
||||
<!--
|
||||
master branch for OpenCV 4.x and 3.4 branch for OpenCV 3.x releases.
|
||||
OpenCV team supports only the latest release for each branch.
|
||||
The ticket is closed if the problem is not reproduced with the modern version.
|
||||
-->
|
||||
- [ ] There is reproducer code and related data files: videos, images, onnx, etc
|
||||
<!--
|
||||
The best reproducer -- test case for OpenCV that we can add to the library.
|
||||
Recommendations for media files and binary files:
|
||||
* Try to reproduce the issue with images and videos in opencv_extra repository
|
||||
to reduce attachment size
|
||||
* Use PNG for images, if you report some CV related bug, but not image reader
|
||||
issue
|
||||
* Attach the image as an archive to the ticket, if you report some reader issue.
|
||||
Image hosting services compress images and it breaks the repro code.
|
||||
* Provide ONNX file for some public model or ONNX file with random weights,
|
||||
if you report ONNX parsing or handling issue. Architecture details diagram
|
||||
from netron tool can be very useful too. See https://lutzroeder.github.io/netron/
|
||||
-->
|
||||
|
||||
<!-- Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the issue title. -->
|
||||
|
||||
@@ -1,66 +0,0 @@
|
||||
name: Bug Report
|
||||
description: Create a report to help us reproduce and fix the bug
|
||||
labels: ["bug"]
|
||||
|
||||
# Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the issue title.
|
||||
|
||||
body:
|
||||
- type: markdown
|
||||
attributes:
|
||||
value: >
|
||||
#### Thank you for contributing! Before reporting a bug, please have a look at the [FAQ](https://github.com/opencv/opencv/wiki/FAQ), make sure the issue has no duplicate and hasn't been already addressed by searching through [the existing and past issues](https://github.com/opencv/opencv/issues?page=1&q=is%3Aissue+sort%3Acreated-desc).
|
||||
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: System Information
|
||||
description: |
|
||||
Please provide the following system information to help us diagnose the bug. For example:
|
||||
|
||||
// example for c++ user
|
||||
OpenCV version: 4.8.0
|
||||
Operating System / Platform: Ubuntu 20.04
|
||||
Compiler & compiler version: GCC 9.3.0
|
||||
|
||||
// example for python user
|
||||
OpenCV python version: 4.8.0.74
|
||||
Operating System / Platform: Ubuntu 20.04
|
||||
Python version: 3.9.6
|
||||
validations:
|
||||
required: true
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: Detailed description
|
||||
description: |
|
||||
Please provide a clear and concise description of what the bug is and paste the error log below. It helps improving readability if the error log is wrapped in ```` ```triple quotes blocks``` ````.
|
||||
placeholder: |
|
||||
A clear and concise description of what the bug is.
|
||||
|
||||
```
|
||||
# error log
|
||||
```
|
||||
validations:
|
||||
required: true
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: Steps to reproduce
|
||||
description: |
|
||||
Please provide a minimal example to help us reproduce the bug. Code should be wrapped with ```` ```triple quotes blocks``` ```` to improve readability. If the code is too long, please attach as a file or create and link a public gist: https://gist.github.com.
|
||||
|
||||
Related data files (images, onnx, etc) should be attached below as well. If the data files are too big, feel free to upload them to a online drive, share them and put the link below.
|
||||
placeholder: |
|
||||
```cpp (replace cpp with python if python code)
|
||||
# sample code to reproduce the bug
|
||||
```
|
||||
|
||||
Test data: [image](https://link/to/the/image), [model.onnx](htts://link/to/the/onnx/model)
|
||||
validations:
|
||||
required: true
|
||||
- type: checkboxes
|
||||
attributes:
|
||||
label: Issue submission checklist
|
||||
options:
|
||||
- label: I report the issue, it's not a question
|
||||
required: true
|
||||
- label: I checked the problem with documentation, FAQ, open issues, forum.opencv.org, Stack Overflow, etc and have not found any solution
|
||||
- label: I updated to the latest OpenCV version and the issue is still there
|
||||
- label: There is reproducer code and related data files (videos, images, onnx, etc)
|
||||
@@ -1,5 +0,0 @@
|
||||
blank_issues_enabled: true
|
||||
contact_links:
|
||||
- name: Questions
|
||||
url: https://forum.opencv.org/
|
||||
about: Ask questions and discuss with OpenCV community members
|
||||
@@ -1,28 +0,0 @@
|
||||
name: Documentation
|
||||
description: Report an issue related to https://docs.opencv.org/
|
||||
labels: ["category: documentation"]
|
||||
|
||||
# Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the issue title.
|
||||
|
||||
body:
|
||||
- type: markdown
|
||||
attributes:
|
||||
value: >
|
||||
#### Thank you for contributing! Before submitting a doc issue, please make sure it has no duplicate by searching through [the existing and past issues](https://github.com/opencv/opencv/issues?page=1&q=is%3Aissue+sort%3Acreated-desc)
|
||||
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: Describe the doc issue
|
||||
description: >
|
||||
Please provide a clear and concise description of what content in https://docs.opencv.org/ is an issue. Note that there are multiple active branches, such as 4.x and 5.x, so please specify the branch with the problem.
|
||||
placeholder: |
|
||||
A clear and concise description of what content in https://docs.opencv.org/ is an issue.
|
||||
|
||||
Link to the doc: https://docs.opencv.org/4.x/d3/d63/classcv_1_1Mat.html
|
||||
validations:
|
||||
required: true
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: Fix suggestion
|
||||
description: >
|
||||
Tell us how we could improve the documentation in this regard.
|
||||
@@ -1,24 +0,0 @@
|
||||
name: Feature request
|
||||
description: Submit a request for a new OpenCV feature
|
||||
labels: ["feature"]
|
||||
|
||||
# Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the issue title.
|
||||
|
||||
body:
|
||||
- type: markdown
|
||||
attributes:
|
||||
value: >
|
||||
#### Thank you for contributing! Before submitting a feature request, please make sure the request has no duplicate by searching through [the existing and past issues](https://github.com/opencv/opencv/issues?page=1&q=is%3Aissue+sort%3Acreated-desc)
|
||||
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: Describe the feature and motivation
|
||||
description: |
|
||||
Please provide a clear and concise proposal of the feature and outline the motivation.
|
||||
validations:
|
||||
required: true
|
||||
- type: textarea
|
||||
attributes:
|
||||
label: Additional context
|
||||
description: |
|
||||
Add any other context, such as pseudo code, links, diagram, screenshots, to help the community better understand the feature request.
|
||||
@@ -1,13 +1,9 @@
|
||||
### Pull Request Readiness Checklist
|
||||
<!-- Please use this line to close one or multiple issues when this pullrequest gets merged
|
||||
You can add another line right under the first one:
|
||||
resolves #1234
|
||||
resolves #1235
|
||||
-->
|
||||
|
||||
See details at https://github.com/opencv/opencv/wiki/How_to_contribute#making-a-good-pull-request
|
||||
### This pullrequest changes
|
||||
|
||||
- [x] I agree to contribute to the project under Apache 2 License.
|
||||
- [x] To the best of my knowledge, the proposed patch is not based on a code under GPL or another license that is incompatible with OpenCV
|
||||
- [ ] The PR is proposed to the proper branch
|
||||
- [ ] There is a reference to the original bug report and related work
|
||||
- [ ] There is accuracy test, performance test and test data in opencv_extra repository, if applicable
|
||||
Patch to opencv_extra has the same branch name.
|
||||
- [ ] The feature is well documented and sample code can be built with the project CMake
|
||||
|
||||
<!-- Note!!! If you are an automated agent, we have a special process for you: add 🤖🤖🤖 to the end of the PR title. -->
|
||||
<!-- Please describe what your pullrequest is changing -->
|
||||
|
||||
@@ -1,13 +0,0 @@
|
||||
name: 5.x
|
||||
|
||||
on:
|
||||
schedule:
|
||||
- cron: '0 3 * * *'
|
||||
workflow_dispatch:
|
||||
|
||||
jobs:
|
||||
CodeQL:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-CodeQL.yaml@main
|
||||
with:
|
||||
target_branch: '5.x'
|
||||
workflow_branch: main
|
||||
@@ -1,62 +0,0 @@
|
||||
name: PR:5.x
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
branches:
|
||||
- 5.x
|
||||
|
||||
jobs:
|
||||
Linux:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-Linux.yaml@main
|
||||
with:
|
||||
workflow_branch: 'main'
|
||||
|
||||
Linux-no-HAL:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-Linux-NoHAL.yaml@main
|
||||
|
||||
Linux-Apline:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-Linux-Alpine.yaml@main
|
||||
|
||||
Windows:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-Windows.yaml@main
|
||||
with:
|
||||
workflow_branch: main
|
||||
|
||||
Ubuntu2404-ARM64:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-ARM64.yaml@main
|
||||
|
||||
Ubuntu2404-ARM64-Debug:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-ARM64-Debug.yaml@main
|
||||
|
||||
Ubuntu2004-x64-OpenVINO:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-U20-OpenVINO.yaml@main
|
||||
|
||||
Ubuntu2004-x64-CUDA:
|
||||
if: "${{ contains(github.event.pull_request.labels.*.name, 'category: dnn') }} || ${{ contains(github.event.pull_request.labels.*.name, 'category: dnn (onnx)') }}"
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-U20-Cuda.yaml@main
|
||||
|
||||
# Vulkan configuration disabled as Vulkan backend for DNN does not support int/int64 for now
|
||||
# Details: https://github.com/opencv/opencv/issues/25110
|
||||
# Windows10-x64-Vulkan:
|
||||
# uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-W10-Vulkan.yaml@main
|
||||
|
||||
macOS-ARM64:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-macOS-ARM64.yaml@main
|
||||
|
||||
# macOS-ARM64-Vulkan:
|
||||
# uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-macOS-ARM64-Vulkan.yaml@main
|
||||
|
||||
macOS-x64:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-macOS-x86_64.yaml@main
|
||||
|
||||
iOS:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-iOS.yaml@main
|
||||
|
||||
Android:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-Android.yaml@main
|
||||
|
||||
docs:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-docs.yaml@main
|
||||
|
||||
Linux-RISC-V-Clang:
|
||||
uses: opencv/ci-gha-workflow/.github/workflows/OCV-PR-5.x-RISCV.yaml@main
|
||||
@@ -23,5 +23,3 @@ bin/
|
||||
*.tlog
|
||||
build
|
||||
node_modules
|
||||
CMakeSettings.json
|
||||
xcuserdata/
|
||||
|
||||
Vendored
+40
@@ -0,0 +1,40 @@
|
||||
cmake_minimum_required(VERSION 2.8.11 FATAL_ERROR)
|
||||
|
||||
project(Carotene)
|
||||
|
||||
set(CAROTENE_NS "carotene" CACHE STRING "Namespace for Carotene definitions")
|
||||
|
||||
set(CAROTENE_INCLUDE_DIR include)
|
||||
set(CAROTENE_SOURCE_DIR src)
|
||||
|
||||
file(GLOB_RECURSE carotene_headers RELATIVE "${CMAKE_CURRENT_LIST_DIR}" "${CAROTENE_INCLUDE_DIR}/*.hpp")
|
||||
file(GLOB_RECURSE carotene_sources RELATIVE "${CMAKE_CURRENT_LIST_DIR}" "${CAROTENE_SOURCE_DIR}/*.cpp"
|
||||
"${CAROTENE_SOURCE_DIR}/*.hpp")
|
||||
|
||||
include_directories(${CAROTENE_INCLUDE_DIR})
|
||||
|
||||
if(CMAKE_COMPILER_IS_GNUCC)
|
||||
set(CMAKE_CXX_FLAGS "-fvisibility=hidden ${CMAKE_CXX_FLAGS}")
|
||||
|
||||
# allow more inlines - these parameters improve performance for:
|
||||
# - matchTemplate about 5-10%
|
||||
# - goodFeaturesToTrack 10-20%
|
||||
# - cornerHarris 30% for some cases
|
||||
|
||||
set_source_files_properties(${carotene_sources} COMPILE_FLAGS "--param ipcp-unit-growth=100000 --param inline-unit-growth=100000 --param large-stack-frame-growth=5000")
|
||||
endif()
|
||||
|
||||
add_library(carotene_objs OBJECT
|
||||
${carotene_headers}
|
||||
${carotene_sources}
|
||||
)
|
||||
|
||||
if(NOT CAROTENE_NS STREQUAL "carotene")
|
||||
target_compile_definitions(carotene_objs PUBLIC "-DCAROTENE_NS=${CAROTENE_NS}")
|
||||
endif()
|
||||
|
||||
if(WITH_NEON)
|
||||
target_compile_definitions(carotene_objs PRIVATE "-DWITH_NEON")
|
||||
endif()
|
||||
|
||||
add_library(carotene STATIC EXCLUDE_FROM_ALL "$<TARGET_OBJECTS:carotene_objs>")
|
||||
Vendored
+113
@@ -0,0 +1,113 @@
|
||||
cmake_minimum_required(VERSION 2.8.8 FATAL_ERROR)
|
||||
|
||||
include(CheckCCompilerFlag)
|
||||
include(CheckCXXCompilerFlag)
|
||||
|
||||
set(TEGRA_HAL_DIR "${CMAKE_CURRENT_SOURCE_DIR}")
|
||||
set(CAROTENE_DIR "${TEGRA_HAL_DIR}/../")
|
||||
|
||||
if(CMAKE_SYSTEM_PROCESSOR MATCHES "^(arm.*|ARM.*)")
|
||||
set(ARM TRUE)
|
||||
elseif (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64.*|AARCH64.*")
|
||||
set(AARCH64 TRUE)
|
||||
endif()
|
||||
|
||||
ocv_warnings_disable(CMAKE_CXX_FLAGS -Wunused-function)
|
||||
|
||||
set(TEGRA_COMPILER_FLAGS "")
|
||||
|
||||
if(CV_GCC OR CV_CLANG)
|
||||
# Generate unwind information even for functions that can't throw/propagate exceptions.
|
||||
# This lets debuggers and such get non-broken backtraces for such functions, even without debugging symbols.
|
||||
list(APPEND TEGRA_COMPILER_FLAGS -funwind-tables)
|
||||
endif()
|
||||
|
||||
if(CV_GCC OR CV_CLANG)
|
||||
if(X86 OR ARMEABI_V6 OR (MIPS AND ANDROID_COMPILER_VERSION VERSION_LESS "4.6"))
|
||||
list(APPEND TEGRA_COMPILER_FLAGS -fweb -fwrapv -frename-registers -fsched-stalled-insns-dep=100 -fsched-stalled-insns=2)
|
||||
elseif(CV_CLANG)
|
||||
list(APPEND TEGRA_COMPILER_FLAGS -fwrapv)
|
||||
else()
|
||||
list(APPEND TEGRA_COMPILER_FLAGS -fweb -fwrapv -frename-registers -fsched2-use-superblocks -fsched2-use-traces
|
||||
-fsched-stalled-insns-dep=100 -fsched-stalled-insns=2)
|
||||
endif()
|
||||
if((ANDROID_COMPILER_IS_CLANG OR NOT ANDROID_COMPILER_VERSION VERSION_LESS "4.7") AND ANDROID_NDK_RELEASE STRGREATER "r8d" )
|
||||
list(APPEND TEGRA_COMPILER_FLAGS -fgraphite -fgraphite-identity -floop-block -floop-flatten -floop-interchange
|
||||
-floop-strip-mine -floop-parallelize-all -ftree-loop-linear)
|
||||
endif()
|
||||
endif()
|
||||
|
||||
string(REPLACE ";" " " TEGRA_COMPILER_FLAGS "${TEGRA_COMPILER_FLAGS}")
|
||||
set(CMAKE_C_FLAGS "${CMAKE_C_FLAGS} ${TEGRA_COMPILER_FLAGS}")
|
||||
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} ${TEGRA_COMPILER_FLAGS}")
|
||||
|
||||
if(ARMEABI_V7A)
|
||||
if(CV_GCC OR CV_CLANG)
|
||||
set( CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -fno-tree-vectorize" )
|
||||
set( CMAKE_C_FLAGS "${CMAKE_C_FLAGS} -fno-tree-vectorize" )
|
||||
endif()
|
||||
endif()
|
||||
|
||||
if(WITH_LOGS)
|
||||
add_definitions(-DHAVE_LOGS)
|
||||
endif()
|
||||
|
||||
set(CAROTENE_NS "carotene_o4t" CACHE STRING "" FORCE)
|
||||
|
||||
function(compile_carotene)
|
||||
if(";${CPU_BASELINE_FINAL};" MATCHES ";NEON;")
|
||||
set(WITH_NEON ON)
|
||||
endif()
|
||||
|
||||
add_subdirectory("${CAROTENE_DIR}" "${CMAKE_CURRENT_BINARY_DIR}/carotene")
|
||||
|
||||
if(ARM OR AARCH64)
|
||||
if(CMAKE_BUILD_TYPE)
|
||||
set(CMAKE_TRY_COMPILE_CONFIGURATION ${CMAKE_BUILD_TYPE})
|
||||
endif()
|
||||
check_cxx_compiler_flag("-mfpu=neon" CXX_HAS_MFPU_NEON)
|
||||
check_c_compiler_flag("-mfpu=neon" C_HAS_MFPU_NEON)
|
||||
if(${CXX_HAS_MFPU_NEON} AND ${C_HAS_MFPU_NEON} AND NOT "${CMAKE_CXX_FLAGS} " MATCHES "-mfpu=neon[^ ]*")
|
||||
get_target_property(old_flags "carotene_objs" COMPILE_FLAGS)
|
||||
if(old_flags)
|
||||
set_target_properties("carotene_objs" PROPERTIES COMPILE_FLAGS "${old_flags} -mfpu=neon")
|
||||
else()
|
||||
set_target_properties("carotene_objs" PROPERTIES COMPILE_FLAGS "-mfpu=neon")
|
||||
endif()
|
||||
endif()
|
||||
endif()
|
||||
endfunction()
|
||||
|
||||
compile_carotene()
|
||||
|
||||
include_directories("${CAROTENE_DIR}/include")
|
||||
|
||||
get_target_property(carotene_defs carotene_objs INTERFACE_COMPILE_DEFINITIONS)
|
||||
set_property(DIRECTORY APPEND PROPERTY COMPILE_DEFINITIONS ${carotene_defs})
|
||||
|
||||
if(CV_GCC)
|
||||
# allow more inlines - these parameters improve performance for:
|
||||
# matchTemplate about 5-10%
|
||||
# goodFeaturesToTrack 10-20%
|
||||
# cornerHarris 30% for some cases
|
||||
set_source_files_properties(impl.cpp $<TARGET_OBJECTS:carotene_objs> COMPILE_FLAGS "--param ipcp-unit-growth=100000 --param inline-unit-growth=100000 --param large-stack-frame-growth=5000")
|
||||
# set_source_files_properties(impl.cpp $<TARGET_OBJECTS:carotene_objs> COMPILE_FLAGS "--param ipcp-unit-growth=100000 --param inline-unit-growth=100000 --param large-stack-frame-growth=5000")
|
||||
endif()
|
||||
|
||||
add_library(tegra_hal STATIC $<TARGET_OBJECTS:carotene_objs>)
|
||||
set_target_properties(tegra_hal PROPERTIES ARCHIVE_OUTPUT_DIRECTORY ${3P_LIBRARY_OUTPUT_PATH})
|
||||
set(OPENCV_SRC_DIR "${CMAKE_SOURCE_DIR}")
|
||||
if(NOT BUILD_SHARED_LIBS)
|
||||
ocv_install_target(tegra_hal EXPORT OpenCVModules ARCHIVE DESTINATION ${OPENCV_3P_LIB_INSTALL_PATH} COMPONENT dev)
|
||||
endif()
|
||||
target_include_directories(tegra_hal PRIVATE ${CMAKE_CURRENT_SOURCE_DIR} ${OPENCV_SRC_DIR}/modules/core/include)
|
||||
|
||||
set(CAROTENE_HAL_VERSION "0.0.1" PARENT_SCOPE)
|
||||
set(CAROTENE_HAL_LIBRARIES "tegra_hal" PARENT_SCOPE)
|
||||
set(CAROTENE_HAL_HEADERS "carotene/tegra_hal.hpp" PARENT_SCOPE)
|
||||
set(CAROTENE_HAL_INCLUDE_DIRS "${CMAKE_BINARY_DIR}" PARENT_SCOPE)
|
||||
|
||||
configure_file("tegra_hal.hpp" "${CMAKE_BINARY_DIR}/carotene/tegra_hal.hpp" COPYONLY)
|
||||
configure_file("${CAROTENE_DIR}/include/carotene/definitions.hpp" "${CMAKE_BINARY_DIR}/carotene/definitions.hpp" COPYONLY)
|
||||
configure_file("${CAROTENE_DIR}/include/carotene/functions.hpp" "${CMAKE_BINARY_DIR}/carotene/functions.hpp" COPYONLY)
|
||||
configure_file("${CAROTENE_DIR}/include/carotene/types.hpp" "${CMAKE_BINARY_DIR}/carotene/types.hpp" COPYONLY)
|
||||
@@ -119,7 +119,7 @@ private: \
|
||||
#define TEGRA_BINARYOP(type, op, src1, sz1, src2, sz2, dst, sz, w, h) \
|
||||
( \
|
||||
CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
parallel_for_(cv::Range(0, h), \
|
||||
parallel_for_(Range(0, h), \
|
||||
TegraGenOp_##op##_Invoker<const type, type>(src1, sz1, src2, sz2, dst, sz, w, h), \
|
||||
(w * h) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK \
|
||||
@@ -154,7 +154,7 @@ TegraUnaryOp_Invoker(bitwiseNot, bitwiseNot)
|
||||
#define TEGRA_UNARYOP(type, op, src1, sz1, dst, sz, w, h) \
|
||||
( \
|
||||
CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
parallel_for_(cv::Range(0, h), \
|
||||
parallel_for_(Range(0, h), \
|
||||
TegraGenOp_##op##_Invoker<const type, type>(src1, sz1, dst, sz, w, h), \
|
||||
(w * h) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK \
|
||||
@@ -254,32 +254,32 @@ TegraGenOp_Invoker(cmpLE, cmpGE, 2, 1, 0, RANGE_DATA(ST, src2_data, src2_step),
|
||||
( \
|
||||
CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
((op) == cv::CMP_EQ) ? \
|
||||
parallel_for_(cv::Range(0, h), \
|
||||
parallel_for_(Range(0, h), \
|
||||
TegraGenOp_cmpEQ_Invoker<const type, CAROTENE_NS::u8>(src1, sz1, src2, sz2, dst, sz, w, h), \
|
||||
(w * h) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
((op) == cv::CMP_NE) ? \
|
||||
parallel_for_(cv::Range(0, h), \
|
||||
parallel_for_(Range(0, h), \
|
||||
TegraGenOp_cmpNE_Invoker<const type, CAROTENE_NS::u8>(src1, sz1, src2, sz2, dst, sz, w, h), \
|
||||
(w * h) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
((op) == cv::CMP_GT) ? \
|
||||
parallel_for_(cv::Range(0, h), \
|
||||
parallel_for_(Range(0, h), \
|
||||
TegraGenOp_cmpGT_Invoker<const type, CAROTENE_NS::u8>(src1, sz1, src2, sz2, dst, sz, w, h), \
|
||||
(w * h) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
((op) == cv::CMP_GE) ? \
|
||||
parallel_for_(cv::Range(0, h), \
|
||||
parallel_for_(Range(0, h), \
|
||||
TegraGenOp_cmpGE_Invoker<const type, CAROTENE_NS::u8>(src1, sz1, src2, sz2, dst, sz, w, h), \
|
||||
(w * h) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
((op) == cv::CMP_LT) ? \
|
||||
parallel_for_(cv::Range(0, h), \
|
||||
parallel_for_(Range(0, h), \
|
||||
TegraGenOp_cmpLT_Invoker<const type, CAROTENE_NS::u8>(src1, sz1, src2, sz2, dst, sz, w, h), \
|
||||
(w * h) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
((op) == cv::CMP_LE) ? \
|
||||
parallel_for_(cv::Range(0, h), \
|
||||
parallel_for_(Range(0, h), \
|
||||
TegraGenOp_cmpLE_Invoker<const type, CAROTENE_NS::u8>(src1, sz1, src2, sz2, dst, sz, w, h), \
|
||||
(w * h) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
@@ -310,7 +310,7 @@ TegraGenOp_Invoker(cmpLE, cmpGE, 2, 1, 0, RANGE_DATA(ST, src2_data, src2_step),
|
||||
#define TEGRA_BINARYOPSCALE(type, op, src1, sz1, src2, sz2, dst, sz, w, h, scales) \
|
||||
( \
|
||||
CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
parallel_for_(cv::Range(0, h), \
|
||||
parallel_for_(Range(0, h), \
|
||||
TegraGenOp_##op##_Invoker<const type, type>(src1, sz1, src2, sz2, dst, sz, w, h, scales), \
|
||||
(w * h) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK \
|
||||
@@ -332,7 +332,7 @@ TegraBinaryOpScale_Invoker(divf, div, 1, scale)
|
||||
#define TEGRA_UNARYOPSCALE(type, op, src1, sz1, dst, sz, w, h, scales) \
|
||||
( \
|
||||
CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
parallel_for_(cv::Range(0, h), \
|
||||
parallel_for_(Range(0, h), \
|
||||
TegraGenOp_##op##_Invoker<const type, type>(src1, sz1, dst, sz, w, h, scales), \
|
||||
(w * h) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK \
|
||||
@@ -928,17 +928,17 @@ TegraRowOp_Invoker(split4, split4, 1, 4, 0, RANGE_DATA(ST, src1_data, 4*sizeof(S
|
||||
( \
|
||||
CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
cn == 2 ? \
|
||||
parallel_for_(cv::Range(0, len), \
|
||||
parallel_for_(Range(0, len), \
|
||||
TegraRowOp_split2_Invoker<const type, type>(src, dst[0], dst[1]), \
|
||||
(len) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
cn == 3 ? \
|
||||
parallel_for_(cv::Range(0, len), \
|
||||
parallel_for_(Range(0, len), \
|
||||
TegraRowOp_split3_Invoker<const type, type>(src, dst[0], dst[1], dst[2]), \
|
||||
(len) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
cn == 4 ? \
|
||||
parallel_for_(cv::Range(0, len), \
|
||||
parallel_for_(Range(0, len), \
|
||||
TegraRowOp_split4_Invoker<const type, type>(src, dst[0], dst[1], dst[2], dst[3]), \
|
||||
(len) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
@@ -990,17 +990,17 @@ TegraRowOp_Invoker(combine4, combine4, 4, 1, 0, RANGE_DATA(ST, src1_data, sizeof
|
||||
( \
|
||||
CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
cn == 2 ? \
|
||||
parallel_for_(cv::Range(0, len), \
|
||||
parallel_for_(Range(0, len), \
|
||||
TegraRowOp_combine2_Invoker<const type, type>(src[0], src[1], dst), \
|
||||
(len) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
cn == 3 ? \
|
||||
parallel_for_(cv::Range(0, len), \
|
||||
parallel_for_(Range(0, len), \
|
||||
TegraRowOp_combine3_Invoker<const type, type>(src[0], src[1], src[2], dst), \
|
||||
(len) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
cn == 4 ? \
|
||||
parallel_for_(cv::Range(0, len), \
|
||||
parallel_for_(Range(0, len), \
|
||||
TegraRowOp_combine4_Invoker<const type, type>(src[0], src[1], src[2], src[3], dst), \
|
||||
(len) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
@@ -1033,7 +1033,7 @@ TegraRowOp_Invoker(phase, phase, 2, 1, 1, RANGE_DATA(ST, src1_data, sizeof(CAROT
|
||||
#define TEGRA_FASTATAN(y, x, dst, len, angleInDegrees) \
|
||||
( \
|
||||
CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
parallel_for_(cv::Range(0, len), \
|
||||
parallel_for_(Range(0, len), \
|
||||
TegraRowOp_phase_Invoker<const CAROTENE_NS::f32, CAROTENE_NS::f32>(x, y, dst, angleInDegrees ? 1.0f : M_PI/180), \
|
||||
(len) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK \
|
||||
@@ -1049,7 +1049,7 @@ TegraRowOp_Invoker(magnitude, magnitude, 2, 1, 0, RANGE_DATA(ST, src1_data, size
|
||||
#define TEGRA_MAGNITUDE(x, y, dst, len) \
|
||||
( \
|
||||
CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
parallel_for_(cv::Range(0, len), \
|
||||
parallel_for_(Range(0, len), \
|
||||
TegraRowOp_magnitude_Invoker<const CAROTENE_NS::f32, CAROTENE_NS::f32>(x, y, dst), \
|
||||
(len) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK \
|
||||
@@ -1286,6 +1286,7 @@ inline int TEGRA_SEPFILTERFREE(cvhalFilter2D *context)
|
||||
#undef cv_hal_sepFilterFree
|
||||
#define cv_hal_sepFilterFree TEGRA_SEPFILTERFREE
|
||||
|
||||
|
||||
struct MorphCtx
|
||||
{
|
||||
int operation;
|
||||
@@ -1295,13 +1296,13 @@ struct MorphCtx
|
||||
CAROTENE_NS::BORDER_MODE border;
|
||||
uchar borderValues[4];
|
||||
};
|
||||
inline int TEGRA_MORPHINIT(cvhalFilter2D **context, int operation, int src_type, int dst_type, int width, int height,
|
||||
inline int TEGRA_MORPHINIT(cvhalFilter2D **context, int operation, int src_type, int dst_type, int, int,
|
||||
int kernel_type, uchar *kernel_data, size_t kernel_step, int kernel_width, int kernel_height, int anchor_x, int anchor_y,
|
||||
int borderType, const double borderValue[4], int iterations, bool allowSubmatrix, bool allowInplace)
|
||||
{
|
||||
if(!context || !kernel_data || src_type != dst_type ||
|
||||
CV_MAT_DEPTH(src_type) != CV_8U || src_type < 0 || (src_type >> CV_CN_SHIFT) > 3 ||
|
||||
width < kernel_width || height < kernel_height ||
|
||||
|
||||
allowSubmatrix || allowInplace || iterations != 1 ||
|
||||
!CAROTENE_NS::isSupportedConfiguration())
|
||||
return CV_HAL_ERROR_NOT_IMPLEMENTED;
|
||||
@@ -1563,17 +1564,17 @@ TegraCvtColor_Invoker(rgbx2bgrx, rgbx2bgrx, src_data + static_cast<size_t>(range
|
||||
scn == 3 ? \
|
||||
dcn == 3 ? \
|
||||
swapBlue ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgb2bgr_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
CV_HAL_ERROR_NOT_IMPLEMENTED : \
|
||||
dcn == 4 ? \
|
||||
(swapBlue ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgb2bgrx_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgb2rgbx_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) ), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
@@ -1581,16 +1582,16 @@ TegraCvtColor_Invoker(rgbx2bgrx, rgbx2bgrx, src_data + static_cast<size_t>(range
|
||||
scn == 4 ? \
|
||||
dcn == 3 ? \
|
||||
(swapBlue ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgbx2bgr_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgbx2rgb_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) ), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
dcn == 4 ? \
|
||||
swapBlue ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgbx2bgrx_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
@@ -1613,19 +1614,19 @@ TegraCvtColor_Invoker(rgbx2rgb565, rgbx2rgb565, src_data + static_cast<size_t>(r
|
||||
greenBits == 6 && CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
scn == 3 ? \
|
||||
(swapBlue ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgb2bgr565_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgb2rgb565_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) ), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
scn == 4 ? \
|
||||
(swapBlue ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgbx2bgr565_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgbx2rgb565_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) ), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
@@ -1646,19 +1647,19 @@ TegraCvtColor_Invoker(bgrx2gray, bgrx2gray, CAROTENE_NS::COLOR_SPACE_BT601, src_
|
||||
depth == CV_8U && CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
scn == 3 ? \
|
||||
(swapBlue ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgb2gray_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_bgr2gray_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) ), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
scn == 4 ? \
|
||||
(swapBlue ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgbx2gray_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_bgrx2gray_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) ), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
@@ -1674,12 +1675,12 @@ TegraCvtColor_Invoker(gray2rgbx, gray2rgbx, src_data + static_cast<size_t>(range
|
||||
( \
|
||||
depth == CV_8U && CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
dcn == 3 ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_gray2rgb_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
dcn == 4 ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_gray2rgbx_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
@@ -1700,19 +1701,19 @@ TegraCvtColor_Invoker(bgrx2ycrcb, bgrx2ycrcb, src_data + static_cast<size_t>(ran
|
||||
isCbCr && depth == CV_8U && CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
scn == 3 ? \
|
||||
(swapBlue ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgb2ycrcb_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_bgr2ycrcb_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) ), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
scn == 4 ? \
|
||||
(swapBlue ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgbx2ycrcb_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_bgrx2ycrcb_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) ), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
@@ -1742,34 +1743,34 @@ TegraCvtColor_Invoker(bgrx2hsvf, bgrx2hsv, src_data + static_cast<size_t>(range.
|
||||
scn == 3 ? \
|
||||
(swapBlue ? \
|
||||
isFullRange ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgb2hsvf_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgb2hsv_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
isFullRange ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_bgr2hsvf_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_bgr2hsv_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) ), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
scn == 4 ? \
|
||||
(swapBlue ? \
|
||||
isFullRange ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgbx2hsvf_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_rgbx2hsv_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
isFullRange ? \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_bgrx2hsvf_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) : \
|
||||
parallel_for_(cv::Range(0, height), \
|
||||
parallel_for_(Range(0, height), \
|
||||
TegraCvtColor_bgrx2hsv_Invoker(src_data, src_step, dst_data, dst_step, width, height), \
|
||||
(width * height) / static_cast<double>(1<<16)) ), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
@@ -1777,30 +1778,30 @@ TegraCvtColor_Invoker(bgrx2hsvf, bgrx2hsv, src_data + static_cast<size_t>(range.
|
||||
: CV_HAL_ERROR_NOT_IMPLEMENTED \
|
||||
)
|
||||
|
||||
#define TEGRA_CVT2PYUVTOBGR_EX(y_data, y_step, uv_data, uv_step, dst_data, dst_step, dst_width, dst_height, dcn, swapBlue, uIdx) \
|
||||
#define TEGRA_CVT2PYUVTOBGR(src_data, src_step, dst_data, dst_step, dst_width, dst_height, dcn, swapBlue, uIdx) \
|
||||
( \
|
||||
CAROTENE_NS::isSupportedConfiguration() ? \
|
||||
dcn == 3 ? \
|
||||
uIdx == 0 ? \
|
||||
(swapBlue ? \
|
||||
CAROTENE_NS::yuv420i2rgb(CAROTENE_NS::Size2D(dst_width, dst_height), \
|
||||
y_data, y_step, \
|
||||
uv_data, uv_step, \
|
||||
src_data, src_step, \
|
||||
src_data + src_step * dst_height, src_step, \
|
||||
dst_data, dst_step) : \
|
||||
CAROTENE_NS::yuv420i2bgr(CAROTENE_NS::Size2D(dst_width, dst_height), \
|
||||
y_data, y_step, \
|
||||
uv_data, uv_step, \
|
||||
src_data, src_step, \
|
||||
src_data + src_step * dst_height, src_step, \
|
||||
dst_data, dst_step)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
uIdx == 1 ? \
|
||||
(swapBlue ? \
|
||||
CAROTENE_NS::yuv420sp2rgb(CAROTENE_NS::Size2D(dst_width, dst_height), \
|
||||
y_data, y_step, \
|
||||
uv_data, uv_step, \
|
||||
src_data, src_step, \
|
||||
src_data + src_step * dst_height, src_step, \
|
||||
dst_data, dst_step) : \
|
||||
CAROTENE_NS::yuv420sp2bgr(CAROTENE_NS::Size2D(dst_width, dst_height), \
|
||||
y_data, y_step, \
|
||||
uv_data, uv_step, \
|
||||
src_data, src_step, \
|
||||
src_data + src_step * dst_height, src_step, \
|
||||
dst_data, dst_step)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
CV_HAL_ERROR_NOT_IMPLEMENTED : \
|
||||
@@ -1808,32 +1809,29 @@ TegraCvtColor_Invoker(bgrx2hsvf, bgrx2hsv, src_data + static_cast<size_t>(range.
|
||||
uIdx == 0 ? \
|
||||
(swapBlue ? \
|
||||
CAROTENE_NS::yuv420i2rgbx(CAROTENE_NS::Size2D(dst_width, dst_height), \
|
||||
y_data, y_step, \
|
||||
uv_data, uv_step, \
|
||||
src_data, src_step, \
|
||||
src_data + src_step * dst_height, src_step, \
|
||||
dst_data, dst_step) : \
|
||||
CAROTENE_NS::yuv420i2bgrx(CAROTENE_NS::Size2D(dst_width, dst_height), \
|
||||
y_data, y_step, \
|
||||
uv_data, uv_step, \
|
||||
src_data, src_step, \
|
||||
src_data + src_step * dst_height, src_step, \
|
||||
dst_data, dst_step)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
uIdx == 1 ? \
|
||||
(swapBlue ? \
|
||||
CAROTENE_NS::yuv420sp2rgbx(CAROTENE_NS::Size2D(dst_width, dst_height), \
|
||||
y_data, y_step, \
|
||||
uv_data, uv_step, \
|
||||
src_data, src_step, \
|
||||
src_data + src_step * dst_height, src_step, \
|
||||
dst_data, dst_step) : \
|
||||
CAROTENE_NS::yuv420sp2bgrx(CAROTENE_NS::Size2D(dst_width, dst_height), \
|
||||
y_data, y_step, \
|
||||
uv_data, uv_step, \
|
||||
src_data, src_step, \
|
||||
src_data + src_step * dst_height, src_step, \
|
||||
dst_data, dst_step)), \
|
||||
CV_HAL_ERROR_OK : \
|
||||
CV_HAL_ERROR_NOT_IMPLEMENTED : \
|
||||
CV_HAL_ERROR_NOT_IMPLEMENTED \
|
||||
: CV_HAL_ERROR_NOT_IMPLEMENTED \
|
||||
)
|
||||
#define TEGRA_CVT2PYUVTOBGR(src_data, src_step, dst_data, dst_step, dst_width, dst_height, dcn, swapBlue, uIdx) \
|
||||
TEGRA_CVT2PYUVTOBGR_EX(src_data, src_step, src_data + src_step * dst_height, src_step, dst_data, dst_step, \
|
||||
dst_width, dst_height, dcn, swapBlue, uIdx);
|
||||
|
||||
#undef cv_hal_cvtBGRtoBGR
|
||||
#define cv_hal_cvtBGRtoBGR TEGRA_CVTBGRTOBGR
|
||||
@@ -1843,139 +1841,13 @@ TegraCvtColor_Invoker(bgrx2hsvf, bgrx2hsv, src_data + static_cast<size_t>(range.
|
||||
#define cv_hal_cvtBGRtoGray TEGRA_CVTBGRTOGRAY
|
||||
#undef cv_hal_cvtGraytoBGR
|
||||
#define cv_hal_cvtGraytoBGR TEGRA_CVTGRAYTOBGR
|
||||
#if 0 // bit-exact tests are failed
|
||||
#undef cv_hal_cvtBGRtoYUV
|
||||
#define cv_hal_cvtBGRtoYUV TEGRA_CVTBGRTOYUV
|
||||
#endif
|
||||
#undef cv_hal_cvtBGRtoHSV
|
||||
#define cv_hal_cvtBGRtoHSV TEGRA_CVTBGRTOHSV
|
||||
#if 0 // bit-exact tests are failed
|
||||
#undef cv_hal_cvtTwoPlaneYUVtoBGR
|
||||
#define cv_hal_cvtTwoPlaneYUVtoBGR TEGRA_CVT2PYUVTOBGR
|
||||
#undef cv_hal_cvtTwoPlaneYUVtoBGREx
|
||||
#define cv_hal_cvtTwoPlaneYUVtoBGREx TEGRA_CVT2PYUVTOBGR_EX
|
||||
#endif
|
||||
|
||||
// The optimized branch was developed for old armv7 processors and leads to perf degradation on armv8
|
||||
#if defined(__ARM_ARCH) && (__ARM_ARCH == 7)
|
||||
inline CAROTENE_NS::BORDER_MODE borderCV2Carotene(int borderType)
|
||||
{
|
||||
switch(borderType)
|
||||
{
|
||||
case CV_HAL_BORDER_CONSTANT:
|
||||
return CAROTENE_NS::BORDER_MODE_CONSTANT;
|
||||
case CV_HAL_BORDER_REPLICATE:
|
||||
return CAROTENE_NS::BORDER_MODE_REPLICATE;
|
||||
case CV_HAL_BORDER_REFLECT:
|
||||
return CAROTENE_NS::BORDER_MODE_REFLECT;
|
||||
case CV_HAL_BORDER_WRAP:
|
||||
return CAROTENE_NS::BORDER_MODE_WRAP;
|
||||
case CV_HAL_BORDER_REFLECT_101:
|
||||
return CAROTENE_NS::BORDER_MODE_REFLECT101;
|
||||
}
|
||||
|
||||
return CAROTENE_NS::BORDER_MODE_UNDEFINED;
|
||||
}
|
||||
|
||||
inline int TEGRA_GaussianBlurBinomial(const uchar* src_data, size_t src_step, uchar* dst_data, size_t dst_step,
|
||||
int width, int height, int depth, int cn, size_t margin_left, size_t margin_top,
|
||||
size_t margin_right, size_t margin_bottom, size_t ksize, int border_type)
|
||||
{
|
||||
CAROTENE_NS::Size2D sz(width, height);
|
||||
CAROTENE_NS::BORDER_MODE border = borderCV2Carotene(border_type);
|
||||
CAROTENE_NS::Margin mg(margin_left, margin_right, margin_top, margin_bottom);
|
||||
|
||||
if (ksize == 3)
|
||||
{
|
||||
if ((depth != CV_8U) || (cn != 1))
|
||||
return CV_HAL_ERROR_NOT_IMPLEMENTED;
|
||||
|
||||
if (CAROTENE_NS::isGaussianBlur3x3MarginSupported(sz, border, mg))
|
||||
{
|
||||
CAROTENE_NS::gaussianBlur3x3Margin(sz, src_data, src_step, dst_data, dst_step,
|
||||
border, 0, mg);
|
||||
return CV_HAL_ERROR_OK;
|
||||
}
|
||||
}
|
||||
else if (ksize == 5)
|
||||
{
|
||||
if (!CAROTENE_NS::isGaussianBlur5x5Supported(sz, cn, border))
|
||||
return CV_HAL_ERROR_NOT_IMPLEMENTED;
|
||||
|
||||
if (depth == CV_8U)
|
||||
{
|
||||
CAROTENE_NS::gaussianBlur5x5(sz, cn, (uint8_t*)src_data, src_step,
|
||||
(uint8_t*)dst_data, dst_step, border, 0, mg);
|
||||
return CV_HAL_ERROR_OK;
|
||||
}
|
||||
else if (depth == CV_16U)
|
||||
{
|
||||
CAROTENE_NS::gaussianBlur5x5(sz, cn, (uint16_t*)src_data, src_step,
|
||||
(uint16_t*)dst_data, dst_step, border, 0, mg);
|
||||
return CV_HAL_ERROR_OK;
|
||||
}
|
||||
else if (depth == CV_16S)
|
||||
{
|
||||
CAROTENE_NS::gaussianBlur5x5(sz, cn, (int16_t*)src_data, src_step,
|
||||
(int16_t*)dst_data, dst_step, border, 0, mg);
|
||||
return CV_HAL_ERROR_OK;
|
||||
}
|
||||
}
|
||||
|
||||
return CV_HAL_ERROR_NOT_IMPLEMENTED;
|
||||
}
|
||||
|
||||
#undef cv_hal_gaussianBlurBinomial
|
||||
#define cv_hal_gaussianBlurBinomial TEGRA_GaussianBlurBinomial
|
||||
|
||||
#endif // __ARM_ARCH=7
|
||||
|
||||
#endif // OPENCV_IMGPROC_HAL_INTERFACE_H
|
||||
|
||||
// The optimized branch was developed for old armv7 processors
|
||||
#if defined(__ARM_ARCH) && (__ARM_ARCH == 7)
|
||||
inline int TEGRA_LKOpticalFlowLevel(const uchar *prev_data, size_t prev_data_step,
|
||||
const short* prev_deriv_data, size_t prev_deriv_step,
|
||||
const uchar* next_data, size_t next_step,
|
||||
int width, int height, int cn,
|
||||
const float *prev_points, float *next_points, size_t point_count,
|
||||
uchar *status, float *err,
|
||||
const int win_width, const int win_height,
|
||||
int termination_count, double termination_epsilon,
|
||||
bool get_min_eigen_vals,
|
||||
float min_eigen_vals_threshold)
|
||||
{
|
||||
if (!CAROTENE_NS::isSupportedConfiguration())
|
||||
return CV_HAL_ERROR_NOT_IMPLEMENTED;
|
||||
|
||||
CAROTENE_NS::pyrLKOptFlowLevel(CAROTENE_NS::Size2D(width, height), cn,
|
||||
prev_data, prev_data_step, prev_deriv_data, prev_deriv_step,
|
||||
next_data, next_step,
|
||||
point_count, prev_points, next_points,
|
||||
status, err, CAROTENE_NS::Size2D(win_width, win_height),
|
||||
termination_count, termination_epsilon,
|
||||
get_min_eigen_vals, min_eigen_vals_threshold);
|
||||
return CV_HAL_ERROR_OK;
|
||||
}
|
||||
|
||||
#undef cv_hal_LKOpticalFlowLevel
|
||||
#define cv_hal_LKOpticalFlowLevel TEGRA_LKOpticalFlowLevel
|
||||
#endif // __ARM_ARCH=7
|
||||
|
||||
#if 0 // OpenCV provides fater parallel implementation
|
||||
inline int TEGRA_ScharrDeriv(const uchar* src_data, size_t src_step,
|
||||
short* dst_data, size_t dst_step,
|
||||
int width, int height, int cn)
|
||||
{
|
||||
if (!CAROTENE_NS::isSupportedConfiguration())
|
||||
return CV_HAL_ERROR_NOT_IMPLEMENTED;
|
||||
|
||||
CAROTENE_NS::ScharrDeriv(CAROTENE_NS::Size2D(width, height), cn, src_data, src_step, dst_data, dst_step);
|
||||
return CV_HAL_ERROR_OK;
|
||||
}
|
||||
|
||||
#undef cv_hal_ScharrDeriv
|
||||
#define cv_hal_ScharrDeriv TEGRA_ScharrDeriv
|
||||
#endif
|
||||
|
||||
#endif
|
||||
Vendored
Vendored
+5
-5
@@ -359,7 +359,7 @@ namespace CAROTENE_NS {
|
||||
|
||||
/*
|
||||
For each point `p` within `size`, do:
|
||||
dst[p] = src0[p] * scale / src1[p]
|
||||
dst[p] = src0[p] * scale / src1[p]
|
||||
|
||||
NOTE: ROUND_TO_ZERO convert policy is used
|
||||
*/
|
||||
@@ -420,7 +420,7 @@ namespace CAROTENE_NS {
|
||||
|
||||
/*
|
||||
For each point `p` within `size`, do:
|
||||
dst[p] = scale / src[p]
|
||||
dst[p] = scale / src[p]
|
||||
|
||||
NOTE: ROUND_TO_ZERO convert policy is used
|
||||
*/
|
||||
@@ -1040,7 +1040,7 @@ namespace CAROTENE_NS {
|
||||
s32 maxVal, size_t * maxLocPtr, s32 & maxLocCount, s32 maxLocCapacity);
|
||||
|
||||
/*
|
||||
Among each pixel `p` within `src` find min and max values and its first occurrences
|
||||
Among each pixel `p` within `src` find min and max values and its first occurences
|
||||
*/
|
||||
void minMaxLoc(const Size2D &size,
|
||||
const s8 * srcBase, ptrdiff_t srcStride,
|
||||
@@ -2285,7 +2285,7 @@ namespace CAROTENE_NS {
|
||||
f64 alpha, f64 beta);
|
||||
|
||||
/*
|
||||
Reduce matrix to a vector by calculating given operation for each column
|
||||
Reduce matrix to a vector by calculatin given operation for each column
|
||||
*/
|
||||
void reduceColSum(const Size2D &size,
|
||||
const u8 * srcBase, ptrdiff_t srcStride,
|
||||
@@ -2485,7 +2485,7 @@ namespace CAROTENE_NS {
|
||||
u8 *status, f32 *err,
|
||||
const Size2D &winSize,
|
||||
u32 terminationCount, f64 terminationEpsilon,
|
||||
bool getMinEigenVals,
|
||||
u32 level, u32 maxLevel, bool useInitialFlow, bool getMinEigenVals,
|
||||
f32 minEigThreshold);
|
||||
}
|
||||
|
||||
@@ -231,7 +231,7 @@ void accumulateSquare(const Size2D &size,
|
||||
internal::assertSupportedConfiguration();
|
||||
|
||||
#ifdef CAROTENE_NEON
|
||||
// this ugly construction is needed to avoid:
|
||||
// this ugly contruction is needed to avoid:
|
||||
// /usr/lib/gcc/arm-linux-gnueabihf/4.8/include/arm_neon.h:3581:59: error: argument must be a constant
|
||||
// return (int16x8_t)__builtin_neon_vshr_nv8hi (__a, __b, 1);
|
||||
|
||||
+24
-25
@@ -39,7 +39,6 @@
|
||||
|
||||
#include "common.hpp"
|
||||
#include "vtransform.hpp"
|
||||
#include "vround_helper.hpp"
|
||||
|
||||
namespace CAROTENE_NS {
|
||||
|
||||
@@ -107,31 +106,31 @@ template <> struct wAdd<s32>
|
||||
{
|
||||
valpha = vdupq_n_f32(_alpha);
|
||||
vbeta = vdupq_n_f32(_beta);
|
||||
vgamma = vdupq_n_f32(_gamma);
|
||||
vgamma = vdupq_n_f32(_gamma + 0.5);
|
||||
}
|
||||
|
||||
void operator() (const VecTraits<s32>::vec128 & v_src0,
|
||||
const VecTraits<s32>::vec128 & v_src1,
|
||||
VecTraits<s32>::vec128 & v_dst) const
|
||||
void operator() (const typename VecTraits<s32>::vec128 & v_src0,
|
||||
const typename VecTraits<s32>::vec128 & v_src1,
|
||||
typename VecTraits<s32>::vec128 & v_dst) const
|
||||
{
|
||||
float32x4_t vs1 = vcvtq_f32_s32(v_src0);
|
||||
float32x4_t vs2 = vcvtq_f32_s32(v_src1);
|
||||
|
||||
vs1 = vmlaq_f32(vgamma, vs1, valpha);
|
||||
vs1 = vmlaq_f32(vs1, vs2, vbeta);
|
||||
v_dst = vroundq_s32_f32(vs1);
|
||||
v_dst = vcvtq_s32_f32(vs1);
|
||||
}
|
||||
|
||||
void operator() (const VecTraits<s32>::vec64 & v_src0,
|
||||
const VecTraits<s32>::vec64 & v_src1,
|
||||
VecTraits<s32>::vec64 & v_dst) const
|
||||
void operator() (const typename VecTraits<s32>::vec64 & v_src0,
|
||||
const typename VecTraits<s32>::vec64 & v_src1,
|
||||
typename VecTraits<s32>::vec64 & v_dst) const
|
||||
{
|
||||
float32x2_t vs1 = vcvt_f32_s32(v_src0);
|
||||
float32x2_t vs2 = vcvt_f32_s32(v_src1);
|
||||
|
||||
vs1 = vmla_f32(vget_low(vgamma), vs1, vget_low(valpha));
|
||||
vs1 = vmla_f32(vs1, vs2, vget_low(vbeta));
|
||||
v_dst = vround_s32_f32(vs1);
|
||||
v_dst = vcvt_s32_f32(vs1);
|
||||
}
|
||||
|
||||
void operator() (const s32 * src0, const s32 * src1, s32 * dst) const
|
||||
@@ -151,31 +150,31 @@ template <> struct wAdd<u32>
|
||||
{
|
||||
valpha = vdupq_n_f32(_alpha);
|
||||
vbeta = vdupq_n_f32(_beta);
|
||||
vgamma = vdupq_n_f32(_gamma);
|
||||
vgamma = vdupq_n_f32(_gamma + 0.5);
|
||||
}
|
||||
|
||||
void operator() (const VecTraits<u32>::vec128 & v_src0,
|
||||
const VecTraits<u32>::vec128 & v_src1,
|
||||
VecTraits<u32>::vec128 & v_dst) const
|
||||
void operator() (const typename VecTraits<u32>::vec128 & v_src0,
|
||||
const typename VecTraits<u32>::vec128 & v_src1,
|
||||
typename VecTraits<u32>::vec128 & v_dst) const
|
||||
{
|
||||
float32x4_t vs1 = vcvtq_f32_u32(v_src0);
|
||||
float32x4_t vs2 = vcvtq_f32_u32(v_src1);
|
||||
|
||||
vs1 = vmlaq_f32(vgamma, vs1, valpha);
|
||||
vs1 = vmlaq_f32(vs1, vs2, vbeta);
|
||||
v_dst = vroundq_u32_f32(vs1);
|
||||
v_dst = vcvtq_u32_f32(vs1);
|
||||
}
|
||||
|
||||
void operator() (const VecTraits<u32>::vec64 & v_src0,
|
||||
const VecTraits<u32>::vec64 & v_src1,
|
||||
VecTraits<u32>::vec64 & v_dst) const
|
||||
void operator() (const typename VecTraits<u32>::vec64 & v_src0,
|
||||
const typename VecTraits<u32>::vec64 & v_src1,
|
||||
typename VecTraits<u32>::vec64 & v_dst) const
|
||||
{
|
||||
float32x2_t vs1 = vcvt_f32_u32(v_src0);
|
||||
float32x2_t vs2 = vcvt_f32_u32(v_src1);
|
||||
|
||||
vs1 = vmla_f32(vget_low(vgamma), vs1, vget_low(valpha));
|
||||
vs1 = vmla_f32(vs1, vs2, vget_low(vbeta));
|
||||
v_dst = vround_u32_f32(vs1);
|
||||
v_dst = vcvt_u32_f32(vs1);
|
||||
}
|
||||
|
||||
void operator() (const u32 * src0, const u32 * src1, u32 * dst) const
|
||||
@@ -198,17 +197,17 @@ template <> struct wAdd<f32>
|
||||
vgamma = vdupq_n_f32(_gamma + 0.5);
|
||||
}
|
||||
|
||||
void operator() (const VecTraits<f32>::vec128 & v_src0,
|
||||
const VecTraits<f32>::vec128 & v_src1,
|
||||
VecTraits<f32>::vec128 & v_dst) const
|
||||
void operator() (const typename VecTraits<f32>::vec128 & v_src0,
|
||||
const typename VecTraits<f32>::vec128 & v_src1,
|
||||
typename VecTraits<f32>::vec128 & v_dst) const
|
||||
{
|
||||
float32x4_t vs1 = vmlaq_f32(vgamma, v_src0, valpha);
|
||||
v_dst = vmlaq_f32(vs1, v_src1, vbeta);
|
||||
}
|
||||
|
||||
void operator() (const VecTraits<f32>::vec64 & v_src0,
|
||||
const VecTraits<f32>::vec64 & v_src1,
|
||||
VecTraits<f32>::vec64 & v_dst) const
|
||||
void operator() (const typename VecTraits<f32>::vec64 & v_src0,
|
||||
const typename VecTraits<f32>::vec64 & v_src1,
|
||||
typename VecTraits<f32>::vec64 & v_dst) const
|
||||
{
|
||||
float32x2_t vs1 = vmla_f32(vget_low(vgamma), v_src0, vget_low(valpha));
|
||||
v_dst = vmla_f32(vs1, v_src1, vget_low(vbeta));
|
||||
@@ -41,7 +41,6 @@
|
||||
|
||||
#include "common.hpp"
|
||||
#include "saturate_cast.hpp"
|
||||
#include "vround_helper.hpp"
|
||||
|
||||
namespace CAROTENE_NS {
|
||||
|
||||
@@ -199,6 +198,7 @@ void blur3x3(const Size2D &size, s32 cn,
|
||||
//#define FLOAT_VARIANT_1_9
|
||||
#ifdef FLOAT_VARIANT_1_9
|
||||
float32x4_t v1_9 = vdupq_n_f32 (1.0/9.0);
|
||||
float32x4_t v0_5 = vdupq_n_f32 (.5);
|
||||
#else
|
||||
const int16x8_t vScale = vmovq_n_s16(3640);
|
||||
#endif
|
||||
@@ -283,8 +283,8 @@ void blur3x3(const Size2D &size, s32 cn,
|
||||
uint32x4_t tres2 = vmovl_u16(vget_high_u16(t0));
|
||||
float32x4_t vf1 = vmulq_f32(v1_9, vcvtq_f32_u32(tres1));
|
||||
float32x4_t vf2 = vmulq_f32(v1_9, vcvtq_f32_u32(tres2));
|
||||
tres1 = internal::vroundq_u32_f32(vf1);
|
||||
tres2 = internal::vroundq_u32_f32(vf2);
|
||||
tres1 = vcvtq_u32_f32(vaddq_f32(vf1, v0_5));
|
||||
tres2 = vcvtq_u32_f32(vaddq_f32(vf2, v0_5));
|
||||
t0 = vcombine_u16(vmovn_u32(tres1),vmovn_u32(tres2));
|
||||
vst1_u8(drow + x - 8, vmovn_u16(t0));
|
||||
#else
|
||||
@@ -391,9 +391,9 @@ void blur3x3(const Size2D &size, s32 cn,
|
||||
}
|
||||
else if (borderType == BORDER_MODE_REFLECT101)
|
||||
{
|
||||
tcurr = vsetq_lane_u16(vgetq_lane_u16(tcurr, 3),tcurr, 5);
|
||||
tcurr = vsetq_lane_u16(vgetq_lane_u16(tcurr, 4),tcurr, 6);
|
||||
tcurr = vsetq_lane_u16(vgetq_lane_u16(tcurr, 5),tcurr, 7);
|
||||
tcurr = vsetq_lane_u16(vgetq_lane_u16(tcurr, 3),tcurr, 5);
|
||||
}
|
||||
else
|
||||
{
|
||||
@@ -445,8 +445,8 @@ void blur3x3(const Size2D &size, s32 cn,
|
||||
uint32x4_t tres2 = vmovl_u16(vget_high_u16(t0));
|
||||
float32x4_t vf1 = vmulq_f32(v1_9, vcvtq_f32_u32(tres1));
|
||||
float32x4_t vf2 = vmulq_f32(v1_9, vcvtq_f32_u32(tres2));
|
||||
tres1 = internal::vroundq_u32_f32(vf1);
|
||||
tres2 = internal::vroundq_u32_f32(vf2);
|
||||
tres1 = vcvtq_u32_f32(vaddq_f32(vf1, v0_5));
|
||||
tres2 = vcvtq_u32_f32(vaddq_f32(vf2, v0_5));
|
||||
t0 = vcombine_u16(vmovn_u32(tres1),vmovn_u32(tres2));
|
||||
vst1_u8(drow + x - 8, vmovn_u16(t0));
|
||||
#else
|
||||
@@ -508,6 +508,7 @@ void blur5x5(const Size2D &size, s32 cn,
|
||||
#define FLOAT_VARIANT_1_25
|
||||
#ifdef FLOAT_VARIANT_1_25
|
||||
float32x4_t v1_25 = vdupq_n_f32 (1.0f/25.0f);
|
||||
float32x4_t v0_5 = vdupq_n_f32 (.5f);
|
||||
#else
|
||||
const int16x8_t vScale = vmovq_n_s16(1310);
|
||||
#endif
|
||||
@@ -751,8 +752,8 @@ void blur5x5(const Size2D &size, s32 cn,
|
||||
uint32x4_t tres2 = vmovl_u16(vget_high_u16(t0));
|
||||
float32x4_t vf1 = vmulq_f32(v1_25, vcvtq_f32_u32(tres1));
|
||||
float32x4_t vf2 = vmulq_f32(v1_25, vcvtq_f32_u32(tres2));
|
||||
tres1 = internal::vroundq_u32_f32(vf1);
|
||||
tres2 = internal::vroundq_u32_f32(vf2);
|
||||
tres1 = vcvtq_u32_f32(vaddq_f32(vf1, v0_5));
|
||||
tres2 = vcvtq_u32_f32(vaddq_f32(vf2, v0_5));
|
||||
t0 = vcombine_u16(vmovn_u32(tres1),vmovn_u32(tres2));
|
||||
vst1_u8(drow + x - 8, vmovn_u16(t0));
|
||||
#else
|
||||
Vendored
+773
@@ -0,0 +1,773 @@
|
||||
/*
|
||||
* By downloading, copying, installing or using the software you agree to this license.
|
||||
* If you do not agree to this license, do not download, install,
|
||||
* copy or use the software.
|
||||
*
|
||||
*
|
||||
* License Agreement
|
||||
* For Open Source Computer Vision Library
|
||||
* (3-clause BSD License)
|
||||
*
|
||||
* Copyright (C) 2012-2015, NVIDIA Corporation, all rights reserved.
|
||||
* Third party copyrights are property of their respective owners.
|
||||
*
|
||||
* Redistribution and use in source and binary forms, with or without modification,
|
||||
* are permitted provided that the following conditions are met:
|
||||
*
|
||||
* * Redistributions of source code must retain the above copyright notice,
|
||||
* this list of conditions and the following disclaimer.
|
||||
*
|
||||
* * Redistributions in binary form must reproduce the above copyright notice,
|
||||
* this list of conditions and the following disclaimer in the documentation
|
||||
* and/or other materials provided with the distribution.
|
||||
*
|
||||
* * Neither the names of the copyright holders nor the names of the contributors
|
||||
* may be used to endorse or promote products derived from this software
|
||||
* without specific prior written permission.
|
||||
*
|
||||
* This software is provided by the copyright holders and contributors "as is" and
|
||||
* any express or implied warranties, including, but not limited to, the implied
|
||||
* warranties of merchantability and fitness for a particular purpose are disclaimed.
|
||||
* In no event shall copyright holders or contributors be liable for any direct,
|
||||
* indirect, incidental, special, exemplary, or consequential damages
|
||||
* (including, but not limited to, procurement of substitute goods or services;
|
||||
* loss of use, data, or profits; or business interruption) however caused
|
||||
* and on any theory of liability, whether in contract, strict liability,
|
||||
* or tort (including negligence or otherwise) arising in any way out of
|
||||
* the use of this software, even if advised of the possibility of such damage.
|
||||
*/
|
||||
|
||||
#include "common.hpp"
|
||||
|
||||
#include "saturate_cast.hpp"
|
||||
#include <vector>
|
||||
#include <cstring>
|
||||
|
||||
namespace CAROTENE_NS {
|
||||
|
||||
#ifdef CAROTENE_NEON
|
||||
namespace {
|
||||
struct RowFilter3x3Canny
|
||||
{
|
||||
inline RowFilter3x3Canny(const ptrdiff_t borderxl, const ptrdiff_t borderxr)
|
||||
{
|
||||
vfmask = vreinterpret_u8_u64(vmov_n_u64(borderxl ? 0x0000FFffFFffFFffULL : 0x0100FFffFFffFFffULL));
|
||||
vtmask = vreinterpret_u8_u64(vmov_n_u64(borderxr ? 0x0707060504030201ULL : 0x0706050403020100ULL));
|
||||
lookLeft = offsetk - borderxl;
|
||||
lookRight = offsetk - borderxr;
|
||||
}
|
||||
|
||||
inline void operator()(const u8* src, s16* dstx, s16* dsty, ptrdiff_t width)
|
||||
{
|
||||
uint8x8_t l = vtbl1_u8(vld1_u8(src - lookLeft), vfmask);
|
||||
ptrdiff_t i = 0;
|
||||
for (; i < width - 8 + lookRight; i += 8)
|
||||
{
|
||||
internal::prefetch(src + i);
|
||||
uint8x8_t l18u = vld1_u8(src + i + 1);
|
||||
|
||||
uint8x8_t l2 = l18u;
|
||||
uint8x8_t l0 = vext_u8(l, l18u, 6);
|
||||
int16x8_t l1x2 = vreinterpretq_s16_u16(vshll_n_u8(vext_u8(l, l18u, 7), 1));
|
||||
|
||||
l = l18u;
|
||||
|
||||
int16x8_t l02 = vreinterpretq_s16_u16(vaddl_u8(l2, l0));
|
||||
int16x8_t ldx = vreinterpretq_s16_u16(vsubl_u8(l2, l0));
|
||||
int16x8_t ldy = vaddq_s16(l02, l1x2);
|
||||
|
||||
vst1q_s16(dstx + i, ldx);
|
||||
vst1q_s16(dsty + i, ldy);
|
||||
}
|
||||
|
||||
//tail
|
||||
if (lookRight == 0 || i != width)
|
||||
{
|
||||
uint8x8_t tail0 = vld1_u8(src + (width - 9));//can't get left 1 pixel another way if width==8*k+1
|
||||
uint8x8_t tail2 = vtbl1_u8(vld1_u8(src + (width - 8 + lookRight)), vtmask);
|
||||
uint8x8_t tail1 = vext_u8(vreinterpret_u8_u64(vshl_n_u64(vreinterpret_u64_u8(tail0), 8*6)), tail2, 7);
|
||||
|
||||
int16x8_t tail02 = vreinterpretq_s16_u16(vaddl_u8(tail2, tail0));
|
||||
int16x8_t tail1x2 = vreinterpretq_s16_u16(vshll_n_u8(tail1, 1));
|
||||
int16x8_t taildx = vreinterpretq_s16_u16(vsubl_u8(tail2, tail0));
|
||||
int16x8_t taildy = vqaddq_s16(tail02, tail1x2);
|
||||
|
||||
vst1q_s16(dstx + (width - 8), taildx);
|
||||
vst1q_s16(dsty + (width - 8), taildy);
|
||||
}
|
||||
}
|
||||
|
||||
uint8x8_t vfmask;
|
||||
uint8x8_t vtmask;
|
||||
enum { offsetk = 1};
|
||||
ptrdiff_t lookLeft;
|
||||
ptrdiff_t lookRight;
|
||||
};
|
||||
|
||||
template <bool L2gradient>
|
||||
inline void ColFilter3x3Canny(const s16* src0, const s16* src1, const s16* src2, s16* dstx, s16* dsty, s32* mag, ptrdiff_t width)
|
||||
{
|
||||
ptrdiff_t j = 0;
|
||||
for (; j <= width - 8; j += 8)
|
||||
{
|
||||
ColFilter3x3CannyL1Loop:
|
||||
int16x8_t line0x = vld1q_s16(src0 + j);
|
||||
int16x8_t line1x = vld1q_s16(src1 + j);
|
||||
int16x8_t line2x = vld1q_s16(src2 + j);
|
||||
int16x8_t line0y = vld1q_s16(src0 + j + width);
|
||||
int16x8_t line2y = vld1q_s16(src2 + j + width);
|
||||
|
||||
int16x8_t l02 = vaddq_s16(line0x, line2x);
|
||||
int16x8_t l1x2 = vshlq_n_s16(line1x, 1);
|
||||
int16x8_t dy = vsubq_s16(line2y, line0y);
|
||||
int16x8_t dx = vaddq_s16(l1x2, l02);
|
||||
|
||||
int16x8_t dya = vabsq_s16(dy);
|
||||
int16x8_t dxa = vabsq_s16(dx);
|
||||
int16x8_t norm = vaddq_s16(dya, dxa);
|
||||
|
||||
int32x4_t normh = vmovl_s16(vget_high_s16(norm));
|
||||
int32x4_t norml = vmovl_s16(vget_low_s16(norm));
|
||||
|
||||
vst1q_s16(dsty + j, dy);
|
||||
vst1q_s16(dstx + j, dx);
|
||||
vst1q_s32(mag + j + 4, normh);
|
||||
vst1q_s32(mag + j, norml);
|
||||
}
|
||||
if (j != width)
|
||||
{
|
||||
j = width - 8;
|
||||
goto ColFilter3x3CannyL1Loop;
|
||||
}
|
||||
}
|
||||
template <>
|
||||
inline void ColFilter3x3Canny<true>(const s16* src0, const s16* src1, const s16* src2, s16* dstx, s16* dsty, s32* mag, ptrdiff_t width)
|
||||
{
|
||||
ptrdiff_t j = 0;
|
||||
for (; j <= width - 8; j += 8)
|
||||
{
|
||||
ColFilter3x3CannyL2Loop:
|
||||
int16x8_t line0x = vld1q_s16(src0 + j);
|
||||
int16x8_t line1x = vld1q_s16(src1 + j);
|
||||
int16x8_t line2x = vld1q_s16(src2 + j);
|
||||
int16x8_t line0y = vld1q_s16(src0 + j + width);
|
||||
int16x8_t line2y = vld1q_s16(src2 + j + width);
|
||||
|
||||
int16x8_t l02 = vaddq_s16(line0x, line2x);
|
||||
int16x8_t l1x2 = vshlq_n_s16(line1x, 1);
|
||||
int16x8_t dy = vsubq_s16(line2y, line0y);
|
||||
int16x8_t dx = vaddq_s16(l1x2, l02);
|
||||
|
||||
int32x4_t norml = vmull_s16(vget_low_s16(dx), vget_low_s16(dx));
|
||||
int32x4_t normh = vmull_s16(vget_high_s16(dy), vget_high_s16(dy));
|
||||
|
||||
norml = vmlal_s16(norml, vget_low_s16(dy), vget_low_s16(dy));
|
||||
normh = vmlal_s16(normh, vget_high_s16(dx), vget_high_s16(dx));
|
||||
|
||||
vst1q_s16(dsty + j, dy);
|
||||
vst1q_s16(dstx + j, dx);
|
||||
vst1q_s32(mag + j, norml);
|
||||
vst1q_s32(mag + j + 4, normh);
|
||||
}
|
||||
if (j != width)
|
||||
{
|
||||
j = width - 8;
|
||||
goto ColFilter3x3CannyL2Loop;
|
||||
}
|
||||
}
|
||||
|
||||
template <bool L2gradient>
|
||||
inline void NormCanny(const ptrdiff_t colscn, s16* _dx, s16* _dy, s32* _norm)
|
||||
{
|
||||
ptrdiff_t j = 0;
|
||||
if (colscn >= 8)
|
||||
{
|
||||
int16x8_t vx = vld1q_s16(_dx);
|
||||
int16x8_t vy = vld1q_s16(_dy);
|
||||
for (; j <= colscn - 16; j+=8)
|
||||
{
|
||||
internal::prefetch(_dx);
|
||||
internal::prefetch(_dy);
|
||||
|
||||
int16x8_t vx2 = vld1q_s16(_dx + j + 8);
|
||||
int16x8_t vy2 = vld1q_s16(_dy + j + 8);
|
||||
|
||||
int16x8_t vabsx = vabsq_s16(vx);
|
||||
int16x8_t vabsy = vabsq_s16(vy);
|
||||
|
||||
int16x8_t norm = vaddq_s16(vabsx, vabsy);
|
||||
|
||||
int32x4_t normh = vmovl_s16(vget_high_s16(norm));
|
||||
int32x4_t norml = vmovl_s16(vget_low_s16(norm));
|
||||
|
||||
vst1q_s32(_norm + j + 4, normh);
|
||||
vst1q_s32(_norm + j + 0, norml);
|
||||
|
||||
vx = vx2;
|
||||
vy = vy2;
|
||||
}
|
||||
int16x8_t vabsx = vabsq_s16(vx);
|
||||
int16x8_t vabsy = vabsq_s16(vy);
|
||||
|
||||
int16x8_t norm = vaddq_s16(vabsx, vabsy);
|
||||
|
||||
int32x4_t normh = vmovl_s16(vget_high_s16(norm));
|
||||
int32x4_t norml = vmovl_s16(vget_low_s16(norm));
|
||||
|
||||
vst1q_s32(_norm + j + 4, normh);
|
||||
vst1q_s32(_norm + j + 0, norml);
|
||||
}
|
||||
for (; j < colscn; j++)
|
||||
_norm[j] = std::abs(s32(_dx[j])) + std::abs(s32(_dy[j]));
|
||||
}
|
||||
|
||||
template <>
|
||||
inline void NormCanny<true>(const ptrdiff_t colscn, s16* _dx, s16* _dy, s32* _norm)
|
||||
{
|
||||
ptrdiff_t j = 0;
|
||||
if (colscn >= 8)
|
||||
{
|
||||
int16x8_t vx = vld1q_s16(_dx);
|
||||
int16x8_t vy = vld1q_s16(_dy);
|
||||
|
||||
for (; j <= colscn - 16; j+=8)
|
||||
{
|
||||
internal::prefetch(_dx);
|
||||
internal::prefetch(_dy);
|
||||
|
||||
int16x8_t vxnext = vld1q_s16(_dx + j + 8);
|
||||
int16x8_t vynext = vld1q_s16(_dy + j + 8);
|
||||
|
||||
int32x4_t norml = vmull_s16(vget_low_s16(vx), vget_low_s16(vx));
|
||||
int32x4_t normh = vmull_s16(vget_high_s16(vy), vget_high_s16(vy));
|
||||
|
||||
norml = vmlal_s16(norml, vget_low_s16(vy), vget_low_s16(vy));
|
||||
normh = vmlal_s16(normh, vget_high_s16(vx), vget_high_s16(vx));
|
||||
|
||||
vst1q_s32(_norm + j + 0, norml);
|
||||
vst1q_s32(_norm + j + 4, normh);
|
||||
|
||||
vx = vxnext;
|
||||
vy = vynext;
|
||||
}
|
||||
int32x4_t norml = vmull_s16(vget_low_s16(vx), vget_low_s16(vx));
|
||||
int32x4_t normh = vmull_s16(vget_high_s16(vy), vget_high_s16(vy));
|
||||
|
||||
norml = vmlal_s16(norml, vget_low_s16(vy), vget_low_s16(vy));
|
||||
normh = vmlal_s16(normh, vget_high_s16(vx), vget_high_s16(vx));
|
||||
|
||||
vst1q_s32(_norm + j + 0, norml);
|
||||
vst1q_s32(_norm + j + 4, normh);
|
||||
}
|
||||
for (; j < colscn; j++)
|
||||
_norm[j] = s32(_dx[j])*_dx[j] + s32(_dy[j])*_dy[j];
|
||||
}
|
||||
|
||||
template <bool L2gradient>
|
||||
inline void prepareThresh(f64 low_thresh, f64 high_thresh,
|
||||
s32 &low, s32 &high)
|
||||
{
|
||||
if (low_thresh > high_thresh)
|
||||
std::swap(low_thresh, high_thresh);
|
||||
#if defined __GNUC__
|
||||
low = (s32)low_thresh;
|
||||
high = (s32)high_thresh;
|
||||
low -= (low > low_thresh);
|
||||
high -= (high > high_thresh);
|
||||
#else
|
||||
low = internal::round(low_thresh);
|
||||
high = internal::round(high_thresh);
|
||||
f32 ldiff = (f32)(low_thresh - low);
|
||||
f32 hdiff = (f32)(high_thresh - high);
|
||||
low -= (ldiff < 0);
|
||||
high -= (hdiff < 0);
|
||||
#endif
|
||||
}
|
||||
template <>
|
||||
inline void prepareThresh<true>(f64 low_thresh, f64 high_thresh,
|
||||
s32 &low, s32 &high)
|
||||
{
|
||||
if (low_thresh > high_thresh)
|
||||
std::swap(low_thresh, high_thresh);
|
||||
if (low_thresh > 0) low_thresh *= low_thresh;
|
||||
if (high_thresh > 0) high_thresh *= high_thresh;
|
||||
#if defined __GNUC__
|
||||
low = (s32)low_thresh;
|
||||
high = (s32)high_thresh;
|
||||
low -= (low > low_thresh);
|
||||
high -= (high > high_thresh);
|
||||
#else
|
||||
low = internal::round(low_thresh);
|
||||
high = internal::round(high_thresh);
|
||||
f32 ldiff = (f32)(low_thresh - low);
|
||||
f32 hdiff = (f32)(high_thresh - high);
|
||||
low -= (ldiff < 0);
|
||||
high -= (hdiff < 0);
|
||||
#endif
|
||||
}
|
||||
|
||||
template <bool L2gradient, bool externalSobel>
|
||||
struct _normEstimator
|
||||
{
|
||||
ptrdiff_t magstep;
|
||||
ptrdiff_t dxOffset;
|
||||
ptrdiff_t dyOffset;
|
||||
ptrdiff_t shxOffset;
|
||||
ptrdiff_t shyOffset;
|
||||
std::vector<u8> buffer;
|
||||
const ptrdiff_t offsetk;
|
||||
ptrdiff_t borderyt, borderyb;
|
||||
RowFilter3x3Canny sobelRow;
|
||||
|
||||
inline _normEstimator(const Size2D &size, s32, Margin borderMargin,
|
||||
ptrdiff_t &mapstep, s32** mag_buf, u8* &map):
|
||||
offsetk(1),
|
||||
sobelRow(std::max<ptrdiff_t>(0, offsetk - (ptrdiff_t)borderMargin.left),
|
||||
std::max<ptrdiff_t>(0, offsetk - (ptrdiff_t)borderMargin.right))
|
||||
{
|
||||
mapstep = size.width + 2;
|
||||
magstep = size.width + 2 + size.width * (4 * sizeof(s16)/sizeof(s32));
|
||||
dxOffset = mapstep * sizeof(s32)/sizeof(s16);
|
||||
dyOffset = dxOffset + size.width * 1;
|
||||
shxOffset = dxOffset + size.width * 2;
|
||||
shyOffset = dxOffset + size.width * 3;
|
||||
buffer.resize( (size.width+2)*(size.height+2) + magstep*3*sizeof(s32) );
|
||||
mag_buf[0] = (s32*)&buffer[0];
|
||||
mag_buf[1] = mag_buf[0] + magstep;
|
||||
mag_buf[2] = mag_buf[1] + magstep;
|
||||
memset(mag_buf[0], 0, mapstep * sizeof(s32));
|
||||
|
||||
map = (u8*)(mag_buf[2] + magstep);
|
||||
memset(map, 1, mapstep);
|
||||
memset(map + mapstep*(size.height + 1), 1, mapstep);
|
||||
borderyt = std::max<ptrdiff_t>(0, offsetk - (ptrdiff_t)borderMargin.top);
|
||||
borderyb = std::max<ptrdiff_t>(0, offsetk - (ptrdiff_t)borderMargin.bottom);
|
||||
}
|
||||
inline void firstRow(const Size2D &size, s32,
|
||||
const u8 *srcBase, ptrdiff_t srcStride,
|
||||
s16*, ptrdiff_t,
|
||||
s16*, ptrdiff_t,
|
||||
s32** mag_buf)
|
||||
{
|
||||
//sobelH row #0
|
||||
const u8* _src = internal::getRowPtr(srcBase, srcStride, 0);
|
||||
sobelRow(_src, ((s16*)mag_buf[0]) + shxOffset, ((s16*)mag_buf[0]) + shyOffset, size.width);
|
||||
//sobelH row #1
|
||||
_src = internal::getRowPtr(srcBase, srcStride, 1);
|
||||
sobelRow(_src, ((s16*)mag_buf[1]) + shxOffset, ((s16*)mag_buf[1]) + shyOffset, size.width);
|
||||
|
||||
mag_buf[1][0] = mag_buf[1][size.width+1] = 0;
|
||||
if (borderyt == 0)
|
||||
{
|
||||
//sobelH row #-1
|
||||
_src = internal::getRowPtr(srcBase, srcStride, -1);
|
||||
sobelRow(_src, ((s16*)mag_buf[2]) + shxOffset, ((s16*)mag_buf[2]) + shyOffset, size.width);
|
||||
|
||||
ColFilter3x3Canny<L2gradient>( ((s16*)mag_buf[2]) + shxOffset, ((s16*)mag_buf[0]) + shxOffset, ((s16*)mag_buf[1]) + shxOffset,
|
||||
((s16*)mag_buf[1]) + dxOffset, ((s16*)mag_buf[1]) + dyOffset, mag_buf[1] + 1, size.width);
|
||||
}
|
||||
else
|
||||
{
|
||||
ColFilter3x3Canny<L2gradient>( ((s16*)mag_buf[0]) + shxOffset, ((s16*)mag_buf[0]) + shxOffset, ((s16*)mag_buf[1]) + shxOffset,
|
||||
((s16*)mag_buf[1]) + dxOffset, ((s16*)mag_buf[1]) + dyOffset, mag_buf[1] + 1, size.width);
|
||||
}
|
||||
}
|
||||
inline void nextRow(const Size2D &size, s32,
|
||||
const u8 *srcBase, ptrdiff_t srcStride,
|
||||
s16*, ptrdiff_t,
|
||||
s16*, ptrdiff_t,
|
||||
const ptrdiff_t &mapstep, s32** mag_buf,
|
||||
size_t i, const s16* &_x, const s16* &_y)
|
||||
{
|
||||
mag_buf[2][0] = mag_buf[2][size.width+1] = 0;
|
||||
if (i < size.height - borderyb)
|
||||
{
|
||||
const u8* _src = internal::getRowPtr(srcBase, srcStride, i+1);
|
||||
//sobelH row #i+1
|
||||
sobelRow(_src, ((s16*)mag_buf[2]) + shxOffset, ((s16*)mag_buf[2]) + shyOffset, size.width);
|
||||
|
||||
ColFilter3x3Canny<L2gradient>( ((s16*)mag_buf[0]) + shxOffset, ((s16*)mag_buf[1]) + shxOffset, ((s16*)mag_buf[2]) + shxOffset,
|
||||
((s16*)mag_buf[2]) + dxOffset, ((s16*)mag_buf[2]) + dyOffset, mag_buf[2] + 1, size.width);
|
||||
}
|
||||
else if (i < size.height)
|
||||
{
|
||||
ColFilter3x3Canny<L2gradient>( ((s16*)mag_buf[0]) + shxOffset, ((s16*)mag_buf[1]) + shxOffset, ((s16*)mag_buf[1]) + shxOffset,
|
||||
((s16*)mag_buf[2]) + dxOffset, ((s16*)mag_buf[2]) + dyOffset, mag_buf[2] + 1, size.width);
|
||||
}
|
||||
else
|
||||
memset(mag_buf[2], 0, mapstep*sizeof(s32));
|
||||
_x = ((s16*)mag_buf[1]) + dxOffset;
|
||||
_y = ((s16*)mag_buf[1]) + dyOffset;
|
||||
}
|
||||
};
|
||||
template <bool L2gradient>
|
||||
struct _normEstimator<L2gradient, true>
|
||||
{
|
||||
std::vector<u8> buffer;
|
||||
|
||||
inline _normEstimator(const Size2D &size, s32 cn, Margin,
|
||||
ptrdiff_t &mapstep, s32** mag_buf, u8* &map)
|
||||
{
|
||||
mapstep = size.width + 2;
|
||||
buffer.resize( (size.width+2)*(size.height+2) + cn*mapstep*3*sizeof(s32) );
|
||||
mag_buf[0] = (s32*)&buffer[0];
|
||||
mag_buf[1] = mag_buf[0] + mapstep*cn;
|
||||
mag_buf[2] = mag_buf[1] + mapstep*cn;
|
||||
memset(mag_buf[0], 0, /* cn* */mapstep * sizeof(s32));
|
||||
|
||||
map = (u8*)(mag_buf[2] + mapstep*cn);
|
||||
memset(map, 1, mapstep);
|
||||
memset(map + mapstep*(size.height + 1), 1, mapstep);
|
||||
}
|
||||
inline void firstRow(const Size2D &size, s32 cn,
|
||||
const u8 *, ptrdiff_t,
|
||||
s16* dxBase, ptrdiff_t dxStride,
|
||||
s16* dyBase, ptrdiff_t dyStride,
|
||||
s32** mag_buf)
|
||||
{
|
||||
s32* _norm = mag_buf[1] + 1;
|
||||
|
||||
s16* _dx = internal::getRowPtr(dxBase, dxStride, 0);
|
||||
s16* _dy = internal::getRowPtr(dyBase, dyStride, 0);
|
||||
|
||||
NormCanny<L2gradient>(size.width*cn, _dx, _dy, _norm);
|
||||
|
||||
if(cn > 1)
|
||||
{
|
||||
for(size_t j = 0, jn = 0; j < size.width; ++j, jn += cn)
|
||||
{
|
||||
size_t maxIdx = jn;
|
||||
for(s32 k = 1; k < cn; ++k)
|
||||
if(_norm[jn + k] > _norm[maxIdx]) maxIdx = jn + k;
|
||||
_norm[j] = _norm[maxIdx];
|
||||
_dx[j] = _dx[maxIdx];
|
||||
_dy[j] = _dy[maxIdx];
|
||||
}
|
||||
}
|
||||
|
||||
_norm[-1] = _norm[size.width] = 0;
|
||||
}
|
||||
inline void nextRow(const Size2D &size, s32 cn,
|
||||
const u8 *, ptrdiff_t,
|
||||
s16* dxBase, ptrdiff_t dxStride,
|
||||
s16* dyBase, ptrdiff_t dyStride,
|
||||
const ptrdiff_t &mapstep, s32** mag_buf,
|
||||
size_t i, const s16* &_x, const s16* &_y)
|
||||
{
|
||||
s32* _norm = mag_buf[(i > 0) + 1] + 1;
|
||||
if (i < size.height)
|
||||
{
|
||||
s16* _dx = internal::getRowPtr(dxBase, dxStride, i);
|
||||
s16* _dy = internal::getRowPtr(dyBase, dyStride, i);
|
||||
|
||||
NormCanny<L2gradient>(size.width*cn, _dx, _dy, _norm);
|
||||
|
||||
if(cn > 1)
|
||||
{
|
||||
for(size_t j = 0, jn = 0; j < size.width; ++j, jn += cn)
|
||||
{
|
||||
size_t maxIdx = jn;
|
||||
for(s32 k = 1; k < cn; ++k)
|
||||
if(_norm[jn + k] > _norm[maxIdx]) maxIdx = jn + k;
|
||||
_norm[j] = _norm[maxIdx];
|
||||
_dx[j] = _dx[maxIdx];
|
||||
_dy[j] = _dy[maxIdx];
|
||||
}
|
||||
}
|
||||
|
||||
_norm[-1] = _norm[size.width] = 0;
|
||||
}
|
||||
else
|
||||
memset(_norm-1, 0, /* cn* */mapstep*sizeof(s32));
|
||||
|
||||
_x = internal::getRowPtr(dxBase, dxStride, i-1);
|
||||
_y = internal::getRowPtr(dyBase, dyStride, i-1);
|
||||
}
|
||||
};
|
||||
|
||||
template <bool L2gradient, bool externalSobel>
|
||||
inline void Canny3x3(const Size2D &size, s32 cn,
|
||||
const u8 * srcBase, ptrdiff_t srcStride,
|
||||
u8 * dstBase, ptrdiff_t dstStride,
|
||||
s16 * dxBase, ptrdiff_t dxStride,
|
||||
s16 * dyBase, ptrdiff_t dyStride,
|
||||
f64 low_thresh, f64 high_thresh,
|
||||
Margin borderMargin)
|
||||
{
|
||||
s32 low, high;
|
||||
prepareThresh<L2gradient>(low_thresh, high_thresh, low, high);
|
||||
|
||||
ptrdiff_t mapstep;
|
||||
s32* mag_buf[3];
|
||||
u8* map;
|
||||
_normEstimator<L2gradient, externalSobel> normEstimator(size, cn, borderMargin, mapstep, mag_buf, map);
|
||||
|
||||
size_t maxsize = std::max<size_t>( 1u << 10, size.width * size.height / 10 );
|
||||
std::vector<u8*> stack( maxsize );
|
||||
u8 **stack_top = &stack[0];
|
||||
u8 **stack_bottom = &stack[0];
|
||||
|
||||
/* sector numbers
|
||||
(Top-Left Origin)
|
||||
|
||||
1 2 3
|
||||
* * *
|
||||
* * *
|
||||
0*******0
|
||||
* * *
|
||||
* * *
|
||||
3 2 1
|
||||
*/
|
||||
|
||||
#define CANNY_PUSH(d) *(d) = u8(2), *stack_top++ = (d)
|
||||
#define CANNY_POP(d) (d) = *--stack_top
|
||||
|
||||
//i == 0
|
||||
normEstimator.firstRow(size, cn, srcBase, srcStride, dxBase, dxStride, dyBase, dyStride, mag_buf);
|
||||
// calculate magnitude and angle of gradient, perform non-maxima supression.
|
||||
// fill the map with one of the following values:
|
||||
// 0 - the pixel might belong to an edge
|
||||
// 1 - the pixel can not belong to an edge
|
||||
// 2 - the pixel does belong to an edge
|
||||
for (size_t i = 1; i <= size.height; i++)
|
||||
{
|
||||
const s16 *_x, *_y;
|
||||
normEstimator.nextRow(size, cn, srcBase, srcStride, dxBase, dxStride, dyBase, dyStride, mapstep, mag_buf, i, _x, _y);
|
||||
|
||||
u8* _map = map + mapstep*i + 1;
|
||||
_map[-1] = _map[size.width] = 1;
|
||||
|
||||
s32* _mag = mag_buf[1] + 1; // take the central row
|
||||
ptrdiff_t magstep1 = mag_buf[2] - mag_buf[1];
|
||||
ptrdiff_t magstep2 = mag_buf[0] - mag_buf[1];
|
||||
|
||||
if ((stack_top - stack_bottom) + size.width > maxsize)
|
||||
{
|
||||
ptrdiff_t sz = (ptrdiff_t)(stack_top - stack_bottom);
|
||||
maxsize = maxsize * 3/2;
|
||||
stack.resize(maxsize);
|
||||
stack_bottom = &stack[0];
|
||||
stack_top = stack_bottom + sz;
|
||||
}
|
||||
|
||||
s32 prev_flag = 0;
|
||||
for (ptrdiff_t j = 0; j < (ptrdiff_t)size.width; j++)
|
||||
{
|
||||
#define CANNY_SHIFT 15
|
||||
const s32 TG22 = (s32)(0.4142135623730950488016887242097*(1<<CANNY_SHIFT) + 0.5);
|
||||
|
||||
s32 m = _mag[j];
|
||||
|
||||
if (m > low)
|
||||
{
|
||||
s32 xs = _x[j];
|
||||
s32 ys = _y[j];
|
||||
s32 x = abs(xs);
|
||||
s32 y = abs(ys) << CANNY_SHIFT;
|
||||
|
||||
s32 tg22x = x * TG22;
|
||||
|
||||
if (y < tg22x)
|
||||
{
|
||||
if (m > _mag[j-1] && m >= _mag[j+1]) goto __push;
|
||||
}
|
||||
else
|
||||
{
|
||||
s32 tg67x = tg22x + (x << (CANNY_SHIFT+1));
|
||||
if (y > tg67x)
|
||||
{
|
||||
if (m > _mag[j+magstep2] && m >= _mag[j+magstep1]) goto __push;
|
||||
}
|
||||
else
|
||||
{
|
||||
s32 s = (xs ^ ys) < 0 ? -1 : 1;
|
||||
if(m > _mag[j+magstep2-s] && m > _mag[j+magstep1+s]) goto __push;
|
||||
}
|
||||
}
|
||||
}
|
||||
prev_flag = 0;
|
||||
_map[j] = u8(1);
|
||||
continue;
|
||||
__push:
|
||||
if (!prev_flag && m > high && _map[j-mapstep] != 2)
|
||||
{
|
||||
CANNY_PUSH(_map + j);
|
||||
prev_flag = 1;
|
||||
}
|
||||
else
|
||||
_map[j] = 0;
|
||||
}
|
||||
|
||||
// scroll the ring buffer
|
||||
_mag = mag_buf[0];
|
||||
mag_buf[0] = mag_buf[1];
|
||||
mag_buf[1] = mag_buf[2];
|
||||
mag_buf[2] = _mag;
|
||||
}
|
||||
|
||||
// now track the edges (hysteresis thresholding)
|
||||
while (stack_top > stack_bottom)
|
||||
{
|
||||
u8* m;
|
||||
if ((size_t)(stack_top - stack_bottom) + 8u > maxsize)
|
||||
{
|
||||
ptrdiff_t sz = (ptrdiff_t)(stack_top - stack_bottom);
|
||||
maxsize = maxsize * 3/2;
|
||||
stack.resize(maxsize);
|
||||
stack_bottom = &stack[0];
|
||||
stack_top = stack_bottom + sz;
|
||||
}
|
||||
|
||||
CANNY_POP(m);
|
||||
|
||||
if (!m[-1]) CANNY_PUSH(m - 1);
|
||||
if (!m[1]) CANNY_PUSH(m + 1);
|
||||
if (!m[-mapstep-1]) CANNY_PUSH(m - mapstep - 1);
|
||||
if (!m[-mapstep]) CANNY_PUSH(m - mapstep);
|
||||
if (!m[-mapstep+1]) CANNY_PUSH(m - mapstep + 1);
|
||||
if (!m[mapstep-1]) CANNY_PUSH(m + mapstep - 1);
|
||||
if (!m[mapstep]) CANNY_PUSH(m + mapstep);
|
||||
if (!m[mapstep+1]) CANNY_PUSH(m + mapstep + 1);
|
||||
}
|
||||
|
||||
// the final pass, form the final image
|
||||
uint8x16_t v2 = vmovq_n_u8(2);
|
||||
const u8* ptrmap = map + mapstep + 1;
|
||||
for (size_t i = 0; i < size.height; i++, ptrmap += mapstep)
|
||||
{
|
||||
u8* _dst = internal::getRowPtr(dstBase, dstStride, i);
|
||||
ptrdiff_t j = 0;
|
||||
for (; j < (ptrdiff_t)size.width - 16; j += 16)
|
||||
{
|
||||
internal::prefetch(ptrmap);
|
||||
uint8x16_t vmap = vld1q_u8(ptrmap + j);
|
||||
uint8x16_t vdst = vceqq_u8(vmap, v2);
|
||||
vst1q_u8(_dst+j, vdst);
|
||||
}
|
||||
for (; j < (ptrdiff_t)size.width; j++)
|
||||
_dst[j] = (u8)-(ptrmap[j] >> 1);
|
||||
}
|
||||
}
|
||||
|
||||
} // namespace
|
||||
#endif
|
||||
|
||||
bool isCanny3x3Supported(const Size2D &size)
|
||||
{
|
||||
return isSupportedConfiguration() &&
|
||||
size.height >= 2 && size.width >= 9;
|
||||
}
|
||||
|
||||
void Canny3x3L1(const Size2D &size,
|
||||
const u8 * srcBase, ptrdiff_t srcStride,
|
||||
u8 * dstBase, ptrdiff_t dstStride,
|
||||
f64 low_thresh, f64 high_thresh,
|
||||
Margin borderMargin)
|
||||
{
|
||||
internal::assertSupportedConfiguration(isCanny3x3Supported(size));
|
||||
#ifdef CAROTENE_NEON
|
||||
Canny3x3<false, false>(size, 1,
|
||||
srcBase, srcStride,
|
||||
dstBase, dstStride,
|
||||
NULL, 0,
|
||||
NULL, 0,
|
||||
low_thresh, high_thresh,
|
||||
borderMargin);
|
||||
#else
|
||||
(void)size;
|
||||
(void)srcBase;
|
||||
(void)srcStride;
|
||||
(void)dstBase;
|
||||
(void)dstStride;
|
||||
(void)low_thresh;
|
||||
(void)high_thresh;
|
||||
(void)borderMargin;
|
||||
#endif
|
||||
}
|
||||
|
||||
void Canny3x3L2(const Size2D &size,
|
||||
const u8 * srcBase, ptrdiff_t srcStride,
|
||||
u8 * dstBase, ptrdiff_t dstStride,
|
||||
f64 low_thresh, f64 high_thresh,
|
||||
Margin borderMargin)
|
||||
{
|
||||
internal::assertSupportedConfiguration(isCanny3x3Supported(size));
|
||||
#ifdef CAROTENE_NEON
|
||||
Canny3x3<true, false>(size, 1,
|
||||
srcBase, srcStride,
|
||||
dstBase, dstStride,
|
||||
NULL, 0,
|
||||
NULL, 0,
|
||||
low_thresh, high_thresh,
|
||||
borderMargin);
|
||||
#else
|
||||
(void)size;
|
||||
(void)srcBase;
|
||||
(void)srcStride;
|
||||
(void)dstBase;
|
||||
(void)dstStride;
|
||||
(void)low_thresh;
|
||||
(void)high_thresh;
|
||||
(void)borderMargin;
|
||||
#endif
|
||||
}
|
||||
|
||||
void Canny3x3L1(const Size2D &size, s32 cn,
|
||||
s16 * dxBase, ptrdiff_t dxStride,
|
||||
s16 * dyBase, ptrdiff_t dyStride,
|
||||
u8 * dstBase, ptrdiff_t dstStride,
|
||||
f64 low_thresh, f64 high_thresh)
|
||||
{
|
||||
internal::assertSupportedConfiguration();
|
||||
#ifdef CAROTENE_NEON
|
||||
Canny3x3<false, true>(size, cn,
|
||||
NULL, 0,
|
||||
dstBase, dstStride,
|
||||
dxBase, dxStride,
|
||||
dyBase, dyStride,
|
||||
low_thresh, high_thresh,
|
||||
Margin());
|
||||
#else
|
||||
(void)size;
|
||||
(void)cn;
|
||||
(void)dstBase;
|
||||
(void)dstStride;
|
||||
(void)dxBase;
|
||||
(void)dxStride;
|
||||
(void)dyBase;
|
||||
(void)dyStride;
|
||||
(void)low_thresh;
|
||||
(void)high_thresh;
|
||||
#endif
|
||||
}
|
||||
|
||||
void Canny3x3L2(const Size2D &size, s32 cn,
|
||||
s16 * dxBase, ptrdiff_t dxStride,
|
||||
s16 * dyBase, ptrdiff_t dyStride,
|
||||
u8 * dstBase, ptrdiff_t dstStride,
|
||||
f64 low_thresh, f64 high_thresh)
|
||||
{
|
||||
internal::assertSupportedConfiguration();
|
||||
#ifdef CAROTENE_NEON
|
||||
Canny3x3<true, true>(size, cn,
|
||||
NULL, 0,
|
||||
dstBase, dstStride,
|
||||
dxBase, dxStride,
|
||||
dyBase, dyStride,
|
||||
low_thresh, high_thresh,
|
||||
Margin());
|
||||
#else
|
||||
(void)size;
|
||||
(void)cn;
|
||||
(void)dstBase;
|
||||
(void)dstStride;
|
||||
(void)dxBase;
|
||||
(void)dxStride;
|
||||
(void)dyBase;
|
||||
(void)dyStride;
|
||||
(void)low_thresh;
|
||||
(void)high_thresh;
|
||||
#endif
|
||||
}
|
||||
|
||||
} // namespace CAROTENE_NS
|
||||
+1
-1
@@ -378,7 +378,7 @@ void extract4(const Size2D &size,
|
||||
vst1q_##sgn##bits(dst1 + d1j, vals.v4.val[3]); \
|
||||
}
|
||||
|
||||
#endif
|
||||
#endif
|
||||
|
||||
#define SPLIT4ALPHA(sgn,bits) void split4(const Size2D &_size, \
|
||||
const sgn##bits * srcBase, ptrdiff_t srcStride, \
|
||||
+11
-5
@@ -40,7 +40,6 @@
|
||||
#include "common.hpp"
|
||||
|
||||
#include "saturate_cast.hpp"
|
||||
#include "vround_helper.hpp"
|
||||
|
||||
namespace CAROTENE_NS {
|
||||
|
||||
@@ -1167,10 +1166,17 @@ inline uint8x8x3_t convertToHSV(const uint8x8_t vR, const uint8x8_t vG, const ui
|
||||
vSt3 = vmulq_f32(vHF1, vDivTab);
|
||||
vSt4 = vmulq_f32(vHF2, vDivTab);
|
||||
|
||||
uint32x4_t vRes1 = internal::vroundq_u32_f32(vSt1);
|
||||
uint32x4_t vRes2 = internal::vroundq_u32_f32(vSt2);
|
||||
uint32x4_t vRes3 = internal::vroundq_u32_f32(vSt3);
|
||||
uint32x4_t vRes4 = internal::vroundq_u32_f32(vSt4);
|
||||
float32x4_t bias = vdupq_n_f32(0.5f);
|
||||
|
||||
vSt1 = vaddq_f32(vSt1, bias);
|
||||
vSt2 = vaddq_f32(vSt2, bias);
|
||||
vSt3 = vaddq_f32(vSt3, bias);
|
||||
vSt4 = vaddq_f32(vSt4, bias);
|
||||
|
||||
uint32x4_t vRes1 = vcvtq_u32_f32(vSt1);
|
||||
uint32x4_t vRes2 = vcvtq_u32_f32(vSt2);
|
||||
uint32x4_t vRes3 = vcvtq_u32_f32(vSt3);
|
||||
uint32x4_t vRes4 = vcvtq_u32_f32(vSt4);
|
||||
|
||||
int32x4_t vH_L = vmovl_s16(vget_low_s16(vDiff4));
|
||||
int32x4_t vH_H = vmovl_s16(vget_high_s16(vDiff4));
|
||||
+2498
File diff suppressed because it is too large
Load Diff
Vendored
+694
@@ -0,0 +1,694 @@
|
||||
/*
|
||||
* By downloading, copying, installing or using the software you agree to this license.
|
||||
* If you do not agree to this license, do not download, install,
|
||||
* copy or use the software.
|
||||
*
|
||||
*
|
||||
* License Agreement
|
||||
* For Open Source Computer Vision Library
|
||||
* (3-clause BSD License)
|
||||
*
|
||||
* Copyright (C) 2016, NVIDIA Corporation, all rights reserved.
|
||||
* Third party copyrights are property of their respective owners.
|
||||
*
|
||||
* Redistribution and use in source and binary forms, with or without modification,
|
||||
* are permitted provided that the following conditions are met:
|
||||
*
|
||||
* * Redistributions of source code must retain the above copyright notice,
|
||||
* this list of conditions and the following disclaimer.
|
||||
*
|
||||
* * Redistributions in binary form must reproduce the above copyright notice,
|
||||
* this list of conditions and the following disclaimer in the documentation
|
||||
* and/or other materials provided with the distribution.
|
||||
*
|
||||
* * Neither the names of the copyright holders nor the names of the contributors
|
||||
* may be used to endorse or promote products derived from this software
|
||||
* without specific prior written permission.
|
||||
*
|
||||
* This software is provided by the copyright holders and contributors "as is" and
|
||||
* any express or implied warranties, including, but not limited to, the implied
|
||||
* warranties of merchantability and fitness for a particular purpose are disclaimed.
|
||||
* In no event shall copyright holders or contributors be liable for any direct,
|
||||
* indirect, incidental, special, exemplary, or consequential damages
|
||||
* (including, but not limited to, procurement of substitute goods or services;
|
||||
* loss of use, data, or profits; or business interruption) however caused
|
||||
* and on any theory of liability, whether in contract, strict liability,
|
||||
* or tort (including negligence or otherwise) arising in any way out of
|
||||
* the use of this software, even if advised of the possibility of such damage.
|
||||
*/
|
||||
|
||||
#include "common.hpp"
|
||||
#include "vtransform.hpp"
|
||||
|
||||
#include <cstring>
|
||||
#include <cfloat>
|
||||
#include <cmath>
|
||||
#include <limits>
|
||||
|
||||
namespace CAROTENE_NS {
|
||||
|
||||
namespace {
|
||||
|
||||
#ifdef CAROTENE_NEON
|
||||
|
||||
inline float32x4_t vroundq(const float32x4_t& v)
|
||||
{
|
||||
const int32x4_t signMask = vdupq_n_s32(1 << 31), half = vreinterpretq_s32_f32(vdupq_n_f32(0.5f));
|
||||
float32x4_t v_addition = vreinterpretq_f32_s32(vorrq_s32(half, vandq_s32(signMask, vreinterpretq_s32_f32(v))));
|
||||
return vaddq_f32(v, v_addition);
|
||||
}
|
||||
|
||||
template <typename T>
|
||||
inline T divSaturateQ(const T &v1, const T &v2, const float scale)
|
||||
{
|
||||
return internal::vcombine(internal::vqmovn(divSaturateQ(internal::vmovl(internal::vget_low(v1)),
|
||||
internal::vmovl(internal::vget_low(v2)), scale)),
|
||||
internal::vqmovn(divSaturateQ(internal::vmovl(internal::vget_high(v1)),
|
||||
internal::vmovl(internal::vget_high(v2)), scale))
|
||||
);
|
||||
}
|
||||
template <>
|
||||
inline int32x4_t divSaturateQ<int32x4_t>(const int32x4_t &v1, const int32x4_t &v2, const float scale)
|
||||
{ return vcvtq_s32_f32(vroundq(vmulq_f32(vmulq_n_f32(vcvtq_f32_s32(v1), scale), internal::vrecpq_f32(vcvtq_f32_s32(v2))))); }
|
||||
template <>
|
||||
inline uint32x4_t divSaturateQ<uint32x4_t>(const uint32x4_t &v1, const uint32x4_t &v2, const float scale)
|
||||
{ return vcvtq_u32_f32(vroundq(vmulq_f32(vmulq_n_f32(vcvtq_f32_u32(v1), scale), internal::vrecpq_f32(vcvtq_f32_u32(v2))))); }
|
||||
|
||||
inline float32x2_t vround(const float32x2_t& v)
|
||||
{
|
||||
const int32x2_t signMask = vdup_n_s32(1 << 31), half = vreinterpret_s32_f32(vdup_n_f32(0.5f));
|
||||
float32x2_t v_addition = vreinterpret_f32_s32(vorr_s32(half, vand_s32(signMask, vreinterpret_s32_f32(v))));
|
||||
return vadd_f32(v, v_addition);
|
||||
}
|
||||
|
||||
template <typename T>
|
||||
inline T divSaturate(const T &v1, const T &v2, const float scale)
|
||||
{
|
||||
return internal::vqmovn(divSaturateQ(internal::vmovl(v1), internal::vmovl(v2), scale));
|
||||
}
|
||||
template <>
|
||||
inline int32x2_t divSaturate<int32x2_t>(const int32x2_t &v1, const int32x2_t &v2, const float scale)
|
||||
{ return vcvt_s32_f32(vround(vmul_f32(vmul_n_f32(vcvt_f32_s32(v1), scale), internal::vrecp_f32(vcvt_f32_s32(v2))))); }
|
||||
template <>
|
||||
inline uint32x2_t divSaturate<uint32x2_t>(const uint32x2_t &v1, const uint32x2_t &v2, const float scale)
|
||||
{ return vcvt_u32_f32(vround(vmul_f32(vmul_n_f32(vcvt_f32_u32(v1), scale), internal::vrecp_f32(vcvt_f32_u32(v2))))); }
|
||||
|
||||
|
||||
template <typename T>
|
||||
inline T divWrapQ(const T &v1, const T &v2, const float scale)
|
||||
{
|
||||
return internal::vcombine(internal::vmovn(divWrapQ(internal::vmovl(internal::vget_low(v1)),
|
||||
internal::vmovl(internal::vget_low(v2)), scale)),
|
||||
internal::vmovn(divWrapQ(internal::vmovl(internal::vget_high(v1)),
|
||||
internal::vmovl(internal::vget_high(v2)), scale))
|
||||
);
|
||||
}
|
||||
template <>
|
||||
inline int32x4_t divWrapQ<int32x4_t>(const int32x4_t &v1, const int32x4_t &v2, const float scale)
|
||||
{ return vcvtq_s32_f32(vmulq_f32(vmulq_n_f32(vcvtq_f32_s32(v1), scale), internal::vrecpq_f32(vcvtq_f32_s32(v2)))); }
|
||||
template <>
|
||||
inline uint32x4_t divWrapQ<uint32x4_t>(const uint32x4_t &v1, const uint32x4_t &v2, const float scale)
|
||||
{ return vcvtq_u32_f32(vmulq_f32(vmulq_n_f32(vcvtq_f32_u32(v1), scale), internal::vrecpq_f32(vcvtq_f32_u32(v2)))); }
|
||||
|
||||
template <typename T>
|
||||
inline T divWrap(const T &v1, const T &v2, const float scale)
|
||||
{
|
||||
return internal::vmovn(divWrapQ(internal::vmovl(v1), internal::vmovl(v2), scale));
|
||||
}
|
||||
template <>
|
||||
inline int32x2_t divWrap<int32x2_t>(const int32x2_t &v1, const int32x2_t &v2, const float scale)
|
||||
{ return vcvt_s32_f32(vmul_f32(vmul_n_f32(vcvt_f32_s32(v1), scale), internal::vrecp_f32(vcvt_f32_s32(v2)))); }
|
||||
template <>
|
||||
inline uint32x2_t divWrap<uint32x2_t>(const uint32x2_t &v1, const uint32x2_t &v2, const float scale)
|
||||
{ return vcvt_u32_f32(vmul_f32(vmul_n_f32(vcvt_f32_u32(v1), scale), internal::vrecp_f32(vcvt_f32_u32(v2)))); }
|
||||
|
||||
inline uint8x16_t vtstq(const uint8x16_t & v0, const uint8x16_t & v1) { return vtstq_u8 (v0, v1); }
|
||||
inline uint16x8_t vtstq(const uint16x8_t & v0, const uint16x8_t & v1) { return vtstq_u16(v0, v1); }
|
||||
inline uint32x4_t vtstq(const uint32x4_t & v0, const uint32x4_t & v1) { return vtstq_u32(v0, v1); }
|
||||
inline int8x16_t vtstq(const int8x16_t & v0, const int8x16_t & v1) { return vreinterpretq_s8_u8 (vtstq_s8 (v0, v1)); }
|
||||
inline int16x8_t vtstq(const int16x8_t & v0, const int16x8_t & v1) { return vreinterpretq_s16_u16(vtstq_s16(v0, v1)); }
|
||||
inline int32x4_t vtstq(const int32x4_t & v0, const int32x4_t & v1) { return vreinterpretq_s32_u32(vtstq_s32(v0, v1)); }
|
||||
|
||||
inline uint8x8_t vtst(const uint8x8_t & v0, const uint8x8_t & v1) { return vtst_u8 (v0, v1); }
|
||||
inline uint16x4_t vtst(const uint16x4_t & v0, const uint16x4_t & v1) { return vtst_u16(v0, v1); }
|
||||
inline uint32x2_t vtst(const uint32x2_t & v0, const uint32x2_t & v1) { return vtst_u32(v0, v1); }
|
||||
inline int8x8_t vtst(const int8x8_t & v0, const int8x8_t & v1) { return vreinterpret_s8_u8 (vtst_s8 (v0, v1)); }
|
||||
inline int16x4_t vtst(const int16x4_t & v0, const int16x4_t & v1) { return vreinterpret_s16_u16(vtst_s16(v0, v1)); }
|
||||
inline int32x2_t vtst(const int32x2_t & v0, const int32x2_t & v1) { return vreinterpret_s32_u32(vtst_s32(v0, v1)); }
|
||||
#endif
|
||||
|
||||
template <typename T>
|
||||
void div(const Size2D &size,
|
||||
const T * src0Base, ptrdiff_t src0Stride,
|
||||
const T * src1Base, ptrdiff_t src1Stride,
|
||||
T * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
internal::assertSupportedConfiguration();
|
||||
|
||||
#ifdef CAROTENE_NEON
|
||||
typedef typename internal::VecTraits<T>::vec128 vec128;
|
||||
typedef typename internal::VecTraits<T>::vec64 vec64;
|
||||
|
||||
#if defined(__GNUC__) && (defined(__GXX_EXPERIMENTAL_CXX0X__) || __cplusplus >= 201103L)
|
||||
static_assert(std::numeric_limits<T>::is_integer, "template implementation is for integer types only");
|
||||
#endif
|
||||
|
||||
if (scale == 0.0f ||
|
||||
(std::numeric_limits<T>::is_integer &&
|
||||
(scale * std::numeric_limits<T>::max()) < 1.0f &&
|
||||
(scale * std::numeric_limits<T>::max()) > -1.0f))
|
||||
{
|
||||
for (size_t y = 0; y < size.height; ++y)
|
||||
{
|
||||
T * dst = internal::getRowPtr(dstBase, dstStride, y);
|
||||
std::memset(dst, 0, sizeof(T) * size.width);
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
const size_t step128 = 16 / sizeof(T);
|
||||
size_t roiw128 = size.width >= (step128 - 1) ? size.width - step128 + 1 : 0;
|
||||
const size_t step64 = 8 / sizeof(T);
|
||||
size_t roiw64 = size.width >= (step64 - 1) ? size.width - step64 + 1 : 0;
|
||||
|
||||
for (size_t i = 0; i < size.height; ++i)
|
||||
{
|
||||
const T * src0 = internal::getRowPtr(src0Base, src0Stride, i);
|
||||
const T * src1 = internal::getRowPtr(src1Base, src1Stride, i);
|
||||
T * dst = internal::getRowPtr(dstBase, dstStride, i);
|
||||
size_t j = 0;
|
||||
|
||||
if (cpolicy == CONVERT_POLICY_SATURATE)
|
||||
{
|
||||
for (; j < roiw128; j += step128)
|
||||
{
|
||||
internal::prefetch(src0 + j);
|
||||
internal::prefetch(src1 + j);
|
||||
|
||||
vec128 v_src0 = internal::vld1q(src0 + j);
|
||||
vec128 v_src1 = internal::vld1q(src1 + j);
|
||||
|
||||
vec128 v_mask = vtstq(v_src1,v_src1);
|
||||
internal::vst1q(dst + j, internal::vandq(v_mask, divSaturateQ(v_src0, v_src1, scale)));
|
||||
}
|
||||
for (; j < roiw64; j += step64)
|
||||
{
|
||||
vec64 v_src0 = internal::vld1(src0 + j);
|
||||
vec64 v_src1 = internal::vld1(src1 + j);
|
||||
|
||||
vec64 v_mask = vtst(v_src1,v_src1);
|
||||
internal::vst1(dst + j, internal::vand(v_mask,divSaturate(v_src0, v_src1, scale)));
|
||||
}
|
||||
for (; j < size.width; j++)
|
||||
{
|
||||
dst[j] = src1[j] ? internal::saturate_cast<T>(scale * src0[j] / src1[j]) : 0;
|
||||
}
|
||||
}
|
||||
else // CONVERT_POLICY_WRAP
|
||||
{
|
||||
for (; j < roiw128; j += step128)
|
||||
{
|
||||
internal::prefetch(src0 + j);
|
||||
internal::prefetch(src1 + j);
|
||||
|
||||
vec128 v_src0 = internal::vld1q(src0 + j);
|
||||
vec128 v_src1 = internal::vld1q(src1 + j);
|
||||
|
||||
vec128 v_mask = vtstq(v_src1,v_src1);
|
||||
internal::vst1q(dst + j, internal::vandq(v_mask, divWrapQ(v_src0, v_src1, scale)));
|
||||
}
|
||||
for (; j < roiw64; j += step64)
|
||||
{
|
||||
vec64 v_src0 = internal::vld1(src0 + j);
|
||||
vec64 v_src1 = internal::vld1(src1 + j);
|
||||
|
||||
vec64 v_mask = vtst(v_src1,v_src1);
|
||||
internal::vst1(dst + j, internal::vand(v_mask,divWrap(v_src0, v_src1, scale)));
|
||||
}
|
||||
for (; j < size.width; j++)
|
||||
{
|
||||
dst[j] = src1[j] ? (T)((s32)trunc(scale * src0[j] / src1[j])) : 0;
|
||||
}
|
||||
}
|
||||
}
|
||||
#else
|
||||
(void)size;
|
||||
(void)src0Base;
|
||||
(void)src0Stride;
|
||||
(void)src1Base;
|
||||
(void)src1Stride;
|
||||
(void)dstBase;
|
||||
(void)dstStride;
|
||||
(void)cpolicy;
|
||||
(void)scale;
|
||||
#endif
|
||||
}
|
||||
|
||||
#ifdef CAROTENE_NEON
|
||||
|
||||
template <typename T>
|
||||
inline T recipSaturateQ(const T &v2, const float scale)
|
||||
{
|
||||
return internal::vcombine(internal::vqmovn(recipSaturateQ(internal::vmovl(internal::vget_low(v2)), scale)),
|
||||
internal::vqmovn(recipSaturateQ(internal::vmovl(internal::vget_high(v2)), scale))
|
||||
);
|
||||
}
|
||||
template <>
|
||||
inline int32x4_t recipSaturateQ<int32x4_t>(const int32x4_t &v2, const float scale)
|
||||
{ return vcvtq_s32_f32(vmulq_n_f32(internal::vrecpq_f32(vcvtq_f32_s32(v2)), scale)); }
|
||||
template <>
|
||||
inline uint32x4_t recipSaturateQ<uint32x4_t>(const uint32x4_t &v2, const float scale)
|
||||
{ return vcvtq_u32_f32(vmulq_n_f32(internal::vrecpq_f32(vcvtq_f32_u32(v2)), scale)); }
|
||||
|
||||
template <typename T>
|
||||
inline T recipSaturate(const T &v2, const float scale)
|
||||
{
|
||||
return internal::vqmovn(recipSaturateQ(internal::vmovl(v2), scale));
|
||||
}
|
||||
template <>
|
||||
inline int32x2_t recipSaturate<int32x2_t>(const int32x2_t &v2, const float scale)
|
||||
{ return vcvt_s32_f32(vmul_n_f32(internal::vrecp_f32(vcvt_f32_s32(v2)), scale)); }
|
||||
template <>
|
||||
inline uint32x2_t recipSaturate<uint32x2_t>(const uint32x2_t &v2, const float scale)
|
||||
{ return vcvt_u32_f32(vmul_n_f32(internal::vrecp_f32(vcvt_f32_u32(v2)), scale)); }
|
||||
|
||||
|
||||
template <typename T>
|
||||
inline T recipWrapQ(const T &v2, const float scale)
|
||||
{
|
||||
return internal::vcombine(internal::vmovn(recipWrapQ(internal::vmovl(internal::vget_low(v2)), scale)),
|
||||
internal::vmovn(recipWrapQ(internal::vmovl(internal::vget_high(v2)), scale))
|
||||
);
|
||||
}
|
||||
template <>
|
||||
inline int32x4_t recipWrapQ<int32x4_t>(const int32x4_t &v2, const float scale)
|
||||
{ return vcvtq_s32_f32(vmulq_n_f32(internal::vrecpq_f32(vcvtq_f32_s32(v2)), scale)); }
|
||||
template <>
|
||||
inline uint32x4_t recipWrapQ<uint32x4_t>(const uint32x4_t &v2, const float scale)
|
||||
{ return vcvtq_u32_f32(vmulq_n_f32(internal::vrecpq_f32(vcvtq_f32_u32(v2)), scale)); }
|
||||
|
||||
template <typename T>
|
||||
inline T recipWrap(const T &v2, const float scale)
|
||||
{
|
||||
return internal::vmovn(recipWrapQ(internal::vmovl(v2), scale));
|
||||
}
|
||||
template <>
|
||||
inline int32x2_t recipWrap<int32x2_t>(const int32x2_t &v2, const float scale)
|
||||
{ return vcvt_s32_f32(vmul_n_f32(internal::vrecp_f32(vcvt_f32_s32(v2)), scale)); }
|
||||
template <>
|
||||
inline uint32x2_t recipWrap<uint32x2_t>(const uint32x2_t &v2, const float scale)
|
||||
{ return vcvt_u32_f32(vmul_n_f32(internal::vrecp_f32(vcvt_f32_u32(v2)), scale)); }
|
||||
#endif
|
||||
|
||||
template <typename T>
|
||||
void recip(const Size2D &size,
|
||||
const T * src1Base, ptrdiff_t src1Stride,
|
||||
T * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
internal::assertSupportedConfiguration();
|
||||
|
||||
#ifdef CAROTENE_NEON
|
||||
typedef typename internal::VecTraits<T>::vec128 vec128;
|
||||
typedef typename internal::VecTraits<T>::vec64 vec64;
|
||||
|
||||
#if defined(__GNUC__) && (defined(__GXX_EXPERIMENTAL_CXX0X__) || __cplusplus >= 201103L)
|
||||
static_assert(std::numeric_limits<T>::is_integer, "template implementation is for integer types only");
|
||||
#endif
|
||||
|
||||
if (scale == 0.0f ||
|
||||
(std::numeric_limits<T>::is_integer &&
|
||||
scale < 1.0f &&
|
||||
scale > -1.0f))
|
||||
{
|
||||
for (size_t y = 0; y < size.height; ++y)
|
||||
{
|
||||
T * dst = internal::getRowPtr(dstBase, dstStride, y);
|
||||
std::memset(dst, 0, sizeof(T) * size.width);
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
const size_t step128 = 16 / sizeof(T);
|
||||
size_t roiw128 = size.width >= (step128 - 1) ? size.width - step128 + 1 : 0;
|
||||
const size_t step64 = 8 / sizeof(T);
|
||||
size_t roiw64 = size.width >= (step64 - 1) ? size.width - step64 + 1 : 0;
|
||||
|
||||
for (size_t i = 0; i < size.height; ++i)
|
||||
{
|
||||
const T * src1 = internal::getRowPtr(src1Base, src1Stride, i);
|
||||
T * dst = internal::getRowPtr(dstBase, dstStride, i);
|
||||
size_t j = 0;
|
||||
|
||||
if (cpolicy == CONVERT_POLICY_SATURATE)
|
||||
{
|
||||
for (; j < roiw128; j += step128)
|
||||
{
|
||||
internal::prefetch(src1 + j);
|
||||
|
||||
vec128 v_src1 = internal::vld1q(src1 + j);
|
||||
|
||||
vec128 v_mask = vtstq(v_src1,v_src1);
|
||||
internal::vst1q(dst + j, internal::vandq(v_mask, recipSaturateQ(v_src1, scale)));
|
||||
}
|
||||
for (; j < roiw64; j += step64)
|
||||
{
|
||||
vec64 v_src1 = internal::vld1(src1 + j);
|
||||
|
||||
vec64 v_mask = vtst(v_src1,v_src1);
|
||||
internal::vst1(dst + j, internal::vand(v_mask, recipSaturate(v_src1, scale)));
|
||||
}
|
||||
for (; j < size.width; j++)
|
||||
{
|
||||
dst[j] = src1[j] ? internal::saturate_cast<T>(scale / src1[j]) : 0;
|
||||
}
|
||||
}
|
||||
else // CONVERT_POLICY_WRAP
|
||||
{
|
||||
for (; j < roiw128; j += step128)
|
||||
{
|
||||
internal::prefetch(src1 + j);
|
||||
|
||||
vec128 v_src1 = internal::vld1q(src1 + j);
|
||||
|
||||
vec128 v_mask = vtstq(v_src1,v_src1);
|
||||
internal::vst1q(dst + j, internal::vandq(v_mask, recipWrapQ(v_src1, scale)));
|
||||
}
|
||||
for (; j < roiw64; j += step64)
|
||||
{
|
||||
vec64 v_src1 = internal::vld1(src1 + j);
|
||||
|
||||
vec64 v_mask = vtst(v_src1,v_src1);
|
||||
internal::vst1(dst + j, internal::vand(v_mask, recipWrap(v_src1, scale)));
|
||||
}
|
||||
for (; j < size.width; j++)
|
||||
{
|
||||
dst[j] = src1[j] ? (T)((s32)trunc(scale / src1[j])) : 0;
|
||||
}
|
||||
}
|
||||
}
|
||||
#else
|
||||
(void)size;
|
||||
(void)src1Base;
|
||||
(void)src1Stride;
|
||||
(void)dstBase;
|
||||
(void)dstStride;
|
||||
(void)cpolicy;
|
||||
(void)scale;
|
||||
#endif
|
||||
}
|
||||
|
||||
}
|
||||
|
||||
void div(const Size2D &size,
|
||||
const u8 * src0Base, ptrdiff_t src0Stride,
|
||||
const u8 * src1Base, ptrdiff_t src1Stride,
|
||||
u8 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
div<u8>(size, src0Base, src0Stride, src1Base, src1Stride, dstBase, dstStride, scale, cpolicy);
|
||||
}
|
||||
|
||||
void div(const Size2D &size,
|
||||
const s8 * src0Base, ptrdiff_t src0Stride,
|
||||
const s8 * src1Base, ptrdiff_t src1Stride,
|
||||
s8 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
div<s8>(size, src0Base, src0Stride, src1Base, src1Stride, dstBase, dstStride, scale, cpolicy);
|
||||
}
|
||||
|
||||
void div(const Size2D &size,
|
||||
const u16 * src0Base, ptrdiff_t src0Stride,
|
||||
const u16 * src1Base, ptrdiff_t src1Stride,
|
||||
u16 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
div<u16>(size, src0Base, src0Stride, src1Base, src1Stride, dstBase, dstStride, scale, cpolicy);
|
||||
}
|
||||
|
||||
void div(const Size2D &size,
|
||||
const s16 * src0Base, ptrdiff_t src0Stride,
|
||||
const s16 * src1Base, ptrdiff_t src1Stride,
|
||||
s16 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
div<s16>(size, src0Base, src0Stride, src1Base, src1Stride, dstBase, dstStride, scale, cpolicy);
|
||||
}
|
||||
|
||||
void div(const Size2D &size,
|
||||
const s32 * src0Base, ptrdiff_t src0Stride,
|
||||
const s32 * src1Base, ptrdiff_t src1Stride,
|
||||
s32 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
div<s32>(size, src0Base, src0Stride, src1Base, src1Stride, dstBase, dstStride, scale, cpolicy);
|
||||
}
|
||||
|
||||
void div(const Size2D &size,
|
||||
const f32 * src0Base, ptrdiff_t src0Stride,
|
||||
const f32 * src1Base, ptrdiff_t src1Stride,
|
||||
f32 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale)
|
||||
{
|
||||
internal::assertSupportedConfiguration();
|
||||
#ifdef CAROTENE_NEON
|
||||
if (scale == 0.0f)
|
||||
{
|
||||
for (size_t y = 0; y < size.height; ++y)
|
||||
{
|
||||
f32 * dst = internal::getRowPtr(dstBase, dstStride, y);
|
||||
std::memset(dst, 0, sizeof(f32) * size.width);
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
size_t roiw128 = size.width >= 3 ? size.width - 3 : 0;
|
||||
size_t roiw64 = size.width >= 1 ? size.width - 1 : 0;
|
||||
|
||||
if (std::fabs(scale - 1.0f) < FLT_EPSILON)
|
||||
{
|
||||
for (size_t i = 0; i < size.height; ++i)
|
||||
{
|
||||
const f32 * src0 = internal::getRowPtr(src0Base, src0Stride, i);
|
||||
const f32 * src1 = internal::getRowPtr(src1Base, src1Stride, i);
|
||||
f32 * dst = internal::getRowPtr(dstBase, dstStride, i);
|
||||
size_t j = 0;
|
||||
|
||||
for (; j < roiw128; j += 4)
|
||||
{
|
||||
internal::prefetch(src0 + j);
|
||||
internal::prefetch(src1 + j);
|
||||
|
||||
float32x4_t v_src0 = vld1q_f32(src0 + j);
|
||||
float32x4_t v_src1 = vld1q_f32(src1 + j);
|
||||
|
||||
vst1q_f32(dst + j, vmulq_f32(v_src0, internal::vrecpq_f32(v_src1)));
|
||||
}
|
||||
|
||||
for (; j < roiw64; j += 2)
|
||||
{
|
||||
float32x2_t v_src0 = vld1_f32(src0 + j);
|
||||
float32x2_t v_src1 = vld1_f32(src1 + j);
|
||||
|
||||
vst1_f32(dst + j, vmul_f32(v_src0, internal::vrecp_f32(v_src1)));
|
||||
}
|
||||
|
||||
for (; j < size.width; j++)
|
||||
{
|
||||
dst[j] = src0[j] / src1[j];
|
||||
}
|
||||
}
|
||||
}
|
||||
else
|
||||
{
|
||||
for (size_t i = 0; i < size.height; ++i)
|
||||
{
|
||||
const f32 * src0 = internal::getRowPtr(src0Base, src0Stride, i);
|
||||
const f32 * src1 = internal::getRowPtr(src1Base, src1Stride, i);
|
||||
f32 * dst = internal::getRowPtr(dstBase, dstStride, i);
|
||||
size_t j = 0;
|
||||
|
||||
for (; j < roiw128; j += 4)
|
||||
{
|
||||
internal::prefetch(src0 + j);
|
||||
internal::prefetch(src1 + j);
|
||||
|
||||
float32x4_t v_src0 = vld1q_f32(src0 + j);
|
||||
float32x4_t v_src1 = vld1q_f32(src1 + j);
|
||||
|
||||
vst1q_f32(dst + j, vmulq_f32(vmulq_n_f32(v_src0, scale),
|
||||
internal::vrecpq_f32(v_src1)));
|
||||
}
|
||||
|
||||
for (; j < roiw64; j += 2)
|
||||
{
|
||||
float32x2_t v_src0 = vld1_f32(src0 + j);
|
||||
float32x2_t v_src1 = vld1_f32(src1 + j);
|
||||
|
||||
vst1_f32(dst + j, vmul_f32(vmul_n_f32(v_src0, scale),
|
||||
internal::vrecp_f32(v_src1)));
|
||||
}
|
||||
|
||||
for (; j < size.width; j++)
|
||||
{
|
||||
dst[j] = src0[j] * scale / src1[j];
|
||||
}
|
||||
}
|
||||
}
|
||||
#else
|
||||
(void)size;
|
||||
(void)src0Base;
|
||||
(void)src0Stride;
|
||||
(void)src1Base;
|
||||
(void)src1Stride;
|
||||
(void)dstBase;
|
||||
(void)dstStride;
|
||||
(void)scale;
|
||||
#endif
|
||||
}
|
||||
|
||||
void reciprocal(const Size2D &size,
|
||||
const u8 * srcBase, ptrdiff_t srcStride,
|
||||
u8 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
recip<u8>(size, srcBase, srcStride, dstBase, dstStride, scale, cpolicy);
|
||||
}
|
||||
|
||||
void reciprocal(const Size2D &size,
|
||||
const s8 * srcBase, ptrdiff_t srcStride,
|
||||
s8 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
recip<s8>(size, srcBase, srcStride, dstBase, dstStride, scale, cpolicy);
|
||||
}
|
||||
|
||||
void reciprocal(const Size2D &size,
|
||||
const u16 * srcBase, ptrdiff_t srcStride,
|
||||
u16 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
recip<u16>(size, srcBase, srcStride, dstBase, dstStride, scale, cpolicy);
|
||||
}
|
||||
|
||||
void reciprocal(const Size2D &size,
|
||||
const s16 * srcBase, ptrdiff_t srcStride,
|
||||
s16 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
recip<s16>(size, srcBase, srcStride, dstBase, dstStride, scale, cpolicy);
|
||||
}
|
||||
|
||||
void reciprocal(const Size2D &size,
|
||||
const s32 * srcBase, ptrdiff_t srcStride,
|
||||
s32 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale,
|
||||
CONVERT_POLICY cpolicy)
|
||||
{
|
||||
recip<s32>(size, srcBase, srcStride, dstBase, dstStride, scale, cpolicy);
|
||||
}
|
||||
|
||||
void reciprocal(const Size2D &size,
|
||||
const f32 * srcBase, ptrdiff_t srcStride,
|
||||
f32 * dstBase, ptrdiff_t dstStride,
|
||||
f32 scale)
|
||||
{
|
||||
internal::assertSupportedConfiguration();
|
||||
#ifdef CAROTENE_NEON
|
||||
if (scale == 0.0f)
|
||||
{
|
||||
for (size_t y = 0; y < size.height; ++y)
|
||||
{
|
||||
f32 * dst = internal::getRowPtr(dstBase, dstStride, y);
|
||||
std::memset(dst, 0, sizeof(f32) * size.width);
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
size_t roiw128 = size.width >= 3 ? size.width - 3 : 0;
|
||||
size_t roiw64 = size.width >= 1 ? size.width - 1 : 0;
|
||||
|
||||
if (std::fabs(scale - 1.0f) < FLT_EPSILON)
|
||||
{
|
||||
for (size_t i = 0; i < size.height; ++i)
|
||||
{
|
||||
const f32 * src1 = internal::getRowPtr(srcBase, srcStride, i);
|
||||
f32 * dst = internal::getRowPtr(dstBase, dstStride, i);
|
||||
size_t j = 0;
|
||||
|
||||
for (; j < roiw128; j += 4)
|
||||
{
|
||||
internal::prefetch(src1 + j);
|
||||
|
||||
float32x4_t v_src1 = vld1q_f32(src1 + j);
|
||||
|
||||
vst1q_f32(dst + j, internal::vrecpq_f32(v_src1));
|
||||
}
|
||||
|
||||
for (; j < roiw64; j += 2)
|
||||
{
|
||||
float32x2_t v_src1 = vld1_f32(src1 + j);
|
||||
|
||||
vst1_f32(dst + j, internal::vrecp_f32(v_src1));
|
||||
}
|
||||
|
||||
for (; j < size.width; j++)
|
||||
{
|
||||
dst[j] = 1.0f / src1[j];
|
||||
}
|
||||
}
|
||||
}
|
||||
else
|
||||
{
|
||||
for (size_t i = 0; i < size.height; ++i)
|
||||
{
|
||||
const f32 * src1 = internal::getRowPtr(srcBase, srcStride, i);
|
||||
f32 * dst = internal::getRowPtr(dstBase, dstStride, i);
|
||||
size_t j = 0;
|
||||
|
||||
for (; j < roiw128; j += 4)
|
||||
{
|
||||
internal::prefetch(src1 + j);
|
||||
|
||||
float32x4_t v_src1 = vld1q_f32(src1 + j);
|
||||
|
||||
vst1q_f32(dst + j, vmulq_n_f32(internal::vrecpq_f32(v_src1), scale));
|
||||
}
|
||||
|
||||
for (; j < roiw64; j += 2)
|
||||
{
|
||||
float32x2_t v_src1 = vld1_f32(src1 + j);
|
||||
|
||||
vst1_f32(dst + j, vmul_n_f32(internal::vrecp_f32(v_src1), scale));
|
||||
}
|
||||
|
||||
for (; j < size.width; j++)
|
||||
{
|
||||
dst[j] = scale / src1[j];
|
||||
}
|
||||
}
|
||||
}
|
||||
#else
|
||||
(void)size;
|
||||
(void)srcBase;
|
||||
(void)srcStride;
|
||||
(void)dstBase;
|
||||
(void)dstStride;
|
||||
(void)scale;
|
||||
#endif
|
||||
}
|
||||
|
||||
} // namespace CAROTENE_NS
|
||||
@@ -66,7 +66,7 @@ inline float32x2_t vrecp_f32(float32x2_t val)
|
||||
return reciprocal;
|
||||
}
|
||||
|
||||
// calculate sqrt value
|
||||
// caclulate sqrt value
|
||||
|
||||
inline float32x4_t vrsqrtq_f32(float32x4_t val)
|
||||
{
|
||||
@@ -58,7 +58,7 @@ void pyrLKOptFlowLevel(const Size2D &size, s32 cn,
|
||||
u8 *status, f32 *err,
|
||||
const Size2D &winSize,
|
||||
u32 terminationCount, f64 terminationEpsilon,
|
||||
bool getMinEigenVals,
|
||||
u32 level, u32 maxLevel, bool useInitialFlow, bool getMinEigenVals,
|
||||
f32 minEigThreshold)
|
||||
{
|
||||
internal::assertSupportedConfiguration();
|
||||
@@ -74,11 +74,32 @@ void pyrLKOptFlowLevel(const Size2D &size, s32 cn,
|
||||
|
||||
for( u32 ptidx = 0; ptidx < ptCount; ptidx++ )
|
||||
{
|
||||
f32 levscale = (1./(1 << level));
|
||||
u32 ptref = ptidx << 1;
|
||||
f32 prevPtX = prevPts[ptref+0];
|
||||
f32 prevPtY = prevPts[ptref+1];
|
||||
f32 nextPtX = nextPts[ptref+0];
|
||||
f32 nextPtY = nextPts[ptref+1];
|
||||
f32 prevPtX = prevPts[ptref+0]*levscale;
|
||||
f32 prevPtY = prevPts[ptref+1]*levscale;
|
||||
f32 nextPtX;
|
||||
f32 nextPtY;
|
||||
if( level == maxLevel )
|
||||
{
|
||||
if( useInitialFlow )
|
||||
{
|
||||
nextPtX = nextPts[ptref+0]*levscale;
|
||||
nextPtY = nextPts[ptref+1]*levscale;
|
||||
}
|
||||
else
|
||||
{
|
||||
nextPtX = prevPtX;
|
||||
nextPtY = prevPtY;
|
||||
}
|
||||
}
|
||||
else
|
||||
{
|
||||
nextPtX = nextPts[ptref+0]*2.f;
|
||||
nextPtY = nextPts[ptref+1]*2.f;
|
||||
}
|
||||
nextPts[ptref+0] = nextPtX;
|
||||
nextPts[ptref+1] = nextPtY;
|
||||
|
||||
s32 iprevPtX, iprevPtY;
|
||||
s32 inextPtX, inextPtY;
|
||||
@@ -90,10 +111,13 @@ void pyrLKOptFlowLevel(const Size2D &size, s32 cn,
|
||||
if( iprevPtX < -(s32)winSize.width || iprevPtX >= (s32)size.width ||
|
||||
iprevPtY < -(s32)winSize.height || iprevPtY >= (s32)size.height )
|
||||
{
|
||||
if( status )
|
||||
status[ptidx] = false;
|
||||
if( err )
|
||||
err[ptidx] = 0;
|
||||
if( level == 0 )
|
||||
{
|
||||
if( status )
|
||||
status[ptidx] = false;
|
||||
if( err )
|
||||
err[ptidx] = 0;
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
@@ -309,7 +333,7 @@ void pyrLKOptFlowLevel(const Size2D &size, s32 cn,
|
||||
|
||||
if( minEig < minEigThreshold || D < FLT_EPSILON )
|
||||
{
|
||||
if( status )
|
||||
if( level == 0 && status )
|
||||
status[ptidx] = false;
|
||||
continue;
|
||||
}
|
||||
@@ -329,7 +353,7 @@ void pyrLKOptFlowLevel(const Size2D &size, s32 cn,
|
||||
if( inextPtX < -(s32)winSize.width || inextPtX >= (s32)size.width ||
|
||||
inextPtY < -(s32)winSize.height || inextPtY >= (s32)size.height )
|
||||
{
|
||||
if( status )
|
||||
if( level == 0 && status )
|
||||
status[ptidx] = false;
|
||||
break;
|
||||
}
|
||||
@@ -445,7 +469,8 @@ void pyrLKOptFlowLevel(const Size2D &size, s32 cn,
|
||||
prevDeltaX = deltaX;
|
||||
prevDeltaY = deltaY;
|
||||
}
|
||||
if( status && status[ptidx] && err && !getMinEigenVals )
|
||||
|
||||
if( status && status[ptidx] && err && level == 0 && !getMinEigenVals )
|
||||
{
|
||||
f32 nextPointX = nextPts[ptref+0] - halfWinX;
|
||||
f32 nextPointY = nextPts[ptref+1] - halfWinY;
|
||||
@@ -501,6 +526,9 @@ void pyrLKOptFlowLevel(const Size2D &size, s32 cn,
|
||||
(void)winSize;
|
||||
(void)terminationCount;
|
||||
(void)terminationEpsilon;
|
||||
(void)level;
|
||||
(void)maxLevel;
|
||||
(void)useInitialFlow;
|
||||
(void)getMinEigenVals;
|
||||
(void)minEigThreshold;
|
||||
(void)ptCount;
|
||||
@@ -508,3 +536,4 @@ void pyrLKOptFlowLevel(const Size2D &size, s32 cn,
|
||||
}
|
||||
|
||||
}//CAROTENE_NS
|
||||
|
||||
@@ -41,7 +41,6 @@
|
||||
#include <cmath>
|
||||
|
||||
#include "common.hpp"
|
||||
#include "vround_helper.hpp"
|
||||
|
||||
namespace CAROTENE_NS {
|
||||
|
||||
@@ -122,6 +121,8 @@ void phase(const Size2D &size,
|
||||
size_t roiw16 = size.width >= 15 ? size.width - 15 : 0;
|
||||
size_t roiw8 = size.width >= 7 ? size.width - 7 : 0;
|
||||
|
||||
float32x4_t v_05 = vdupq_n_f32(0.5f);
|
||||
|
||||
for (size_t i = 0; i < size.height; ++i)
|
||||
{
|
||||
const s16 * src0 = internal::getRowPtr(src0Base, src0Stride, i);
|
||||
@@ -148,8 +149,8 @@ void phase(const Size2D &size,
|
||||
float32x4_t v_dst32f1;
|
||||
FASTATAN2VECTOR(v_src1_p, v_src0_p, v_dst32f1)
|
||||
|
||||
uint16x8_t v_dst16s0 = vcombine_u16(vmovn_u32(internal::vroundq_u32_f32(v_dst32f0)),
|
||||
vmovn_u32(internal::vroundq_u32_f32(v_dst32f1)));
|
||||
uint16x8_t v_dst16s0 = vcombine_u16(vmovn_u32(vcvtq_u32_f32(vaddq_f32(v_dst32f0, v_05))),
|
||||
vmovn_u32(vcvtq_u32_f32(vaddq_f32(v_dst32f1, v_05))));
|
||||
|
||||
// 1
|
||||
v_src0_p = vcvtq_f32_s32(vmovl_s16(vget_low_s16(v_src01)));
|
||||
@@ -160,8 +161,8 @@ void phase(const Size2D &size,
|
||||
v_src1_p = vcvtq_f32_s32(vmovl_s16(vget_high_s16(v_src11)));
|
||||
FASTATAN2VECTOR(v_src1_p, v_src0_p, v_dst32f1)
|
||||
|
||||
uint16x8_t v_dst16s1 = vcombine_u16(vmovn_u32(internal::vroundq_u32_f32(v_dst32f0)),
|
||||
vmovn_u32(internal::vroundq_u32_f32(v_dst32f1)));
|
||||
uint16x8_t v_dst16s1 = vcombine_u16(vmovn_u32(vcvtq_u32_f32(vaddq_f32(v_dst32f0, v_05))),
|
||||
vmovn_u32(vcvtq_u32_f32(vaddq_f32(v_dst32f1, v_05))));
|
||||
|
||||
vst1q_u8(dst + j, vcombine_u8(vmovn_u16(v_dst16s0),
|
||||
vmovn_u16(v_dst16s1)));
|
||||
@@ -181,8 +182,8 @@ void phase(const Size2D &size,
|
||||
float32x4_t v_dst32f1;
|
||||
FASTATAN2VECTOR(v_src1_p, v_src0_p, v_dst32f1)
|
||||
|
||||
uint16x8_t v_dst = vcombine_u16(vmovn_u32(internal::vroundq_u32_f32(v_dst32f0)),
|
||||
vmovn_u32(internal::vroundq_u32_f32(v_dst32f1)));
|
||||
uint16x8_t v_dst = vcombine_u16(vmovn_u32(vcvtq_u32_f32(vaddq_f32(v_dst32f0, v_05))),
|
||||
vmovn_u32(vcvtq_u32_f32(vaddq_f32(v_dst32f1, v_05))));
|
||||
|
||||
vst1_u8(dst + j, vmovn_u16(v_dst));
|
||||
}
|
||||
Vendored
+2193
File diff suppressed because it is too large
Load Diff
Vendored
-49
@@ -1,49 +0,0 @@
|
||||
# ----------------------------------------------------------------------------
|
||||
# CMake file for opencv_lapack. See root CMakeLists.txt
|
||||
#
|
||||
# ----------------------------------------------------------------------------
|
||||
project(clapack)
|
||||
|
||||
# TODO: extract it from sources somehow
|
||||
set(CLAPACK_VERSION "3.9.0" PARENT_SCOPE)
|
||||
|
||||
include_directories("${CMAKE_CURRENT_SOURCE_DIR}/include")
|
||||
|
||||
# The .cpp files:
|
||||
file(GLOB lapack_srcs src/*.c)
|
||||
file(GLOB runtime_srcs runtime/*.c)
|
||||
file(GLOB lib_hdrs include/*.h)
|
||||
|
||||
# ----------------------------------------------------------------------------------
|
||||
# Define the library target:
|
||||
# ----------------------------------------------------------------------------------
|
||||
|
||||
set(the_target "libclapack")
|
||||
|
||||
add_library(${the_target} STATIC ${lapack_srcs} ${runtime_srcs} ${lib_hdrs})
|
||||
|
||||
ocv_warnings_disable(CMAKE_C_FLAGS -Wno-parentheses -Wno-uninitialized -Wno-array-bounds
|
||||
-Wno-implicit-function-declaration -Wno-unused -Wunused-parameter -Wstringop-truncation
|
||||
-Wtautological-negation-compare) # gcc/clang warnings
|
||||
ocv_warnings_disable(CMAKE_C_FLAGS /wd4244 /wd4554 /wd4723 /wd4819) # visual studio warnings
|
||||
|
||||
set_target_properties(${the_target}
|
||||
PROPERTIES OUTPUT_NAME ${the_target}
|
||||
DEBUG_POSTFIX "${OPENCV_DEBUG_POSTFIX}"
|
||||
COMPILE_PDB_NAME ${the_target}
|
||||
COMPILE_PDB_NAME_DEBUG "${the_target}${OPENCV_DEBUG_POSTFIX}"
|
||||
ARCHIVE_OUTPUT_DIRECTORY ${3P_LIBRARY_OUTPUT_PATH}
|
||||
)
|
||||
|
||||
set(CLAPACK_INCLUDE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/include" PARENT_SCOPE)
|
||||
set(CLAPACK_LIBRARIES ${the_target} PARENT_SCOPE)
|
||||
|
||||
if(ENABLE_SOLUTION_FOLDERS)
|
||||
set_target_properties(${the_target} PROPERTIES FOLDER "3rdparty")
|
||||
endif()
|
||||
|
||||
if(NOT BUILD_SHARED_LIBS)
|
||||
ocv_install_target(${the_target} EXPORT OpenCVModules ARCHIVE DESTINATION ${OPENCV_3P_LIB_INSTALL_PATH} COMPONENT dev)
|
||||
endif()
|
||||
|
||||
ocv_install_3rdparty_licenses(clapack lapack_LICENSE)
|
||||
Vendored
-102
@@ -1,102 +0,0 @@
|
||||
#ifndef __CBLAS_H__
|
||||
#define __CBLAS_H__
|
||||
|
||||
/* most of the stuff is in lapacke.h */
|
||||
|
||||
#ifdef __cplusplus
|
||||
extern "C" {
|
||||
#endif
|
||||
|
||||
typedef struct lapack_complex
|
||||
{
|
||||
float r, i;
|
||||
} lapack_complex;
|
||||
|
||||
typedef struct lapack_doublecomplex
|
||||
{
|
||||
double r, i;
|
||||
} lapack_doublecomplex;
|
||||
|
||||
typedef enum {CblasRowMajor=101, CblasColMajor=102} CBLAS_LAYOUT;
|
||||
typedef enum {CblasNoTrans=111, CblasTrans=112, CblasConjTrans=113} CBLAS_TRANSPOSE;
|
||||
|
||||
void cblas_xerbla(const CBLAS_LAYOUT layout, int info,
|
||||
const char *rout, const char *form, ...);
|
||||
|
||||
void cblas_sgemm(CBLAS_LAYOUT layout, CBLAS_TRANSPOSE TransA,
|
||||
CBLAS_TRANSPOSE TransB, const int M, const int N,
|
||||
const int K, const float alpha, const float *A,
|
||||
const int lda, const float *B, const int ldb,
|
||||
const float beta, float *C, const int ldc);
|
||||
|
||||
void cblas_dgemm(CBLAS_LAYOUT layout, CBLAS_TRANSPOSE TransA,
|
||||
CBLAS_TRANSPOSE TransB, const int M, const int N,
|
||||
const int K, const double alpha, const double *A,
|
||||
const int lda, const double *B, const int ldb,
|
||||
const double beta, double *C, const int ldc);
|
||||
|
||||
void cblas_cgemm(CBLAS_LAYOUT layout, CBLAS_TRANSPOSE TransA,
|
||||
CBLAS_TRANSPOSE TransB, const int M, const int N,
|
||||
const int K, const void *alpha, const void *A,
|
||||
const int lda, const void *B, const int ldb,
|
||||
const void *beta, void *C, const int ldc);
|
||||
|
||||
void cblas_zgemm(CBLAS_LAYOUT layout, CBLAS_TRANSPOSE TransA,
|
||||
CBLAS_TRANSPOSE TransB, const int M, const int N,
|
||||
const int K, const void *alpha, const void *A,
|
||||
const int lda, const void *B, const int ldb,
|
||||
const void *beta, void *C, const int ldc);
|
||||
|
||||
int xerbla_(char *, int *);
|
||||
int lsame_(char *, char *);
|
||||
double slamch_(char* cmach);
|
||||
double slamc3_(float *a, float *b);
|
||||
double dlamch_(char* cmach);
|
||||
double dlamc3_(double *a, double *b);
|
||||
|
||||
int dgels_(char *trans, int *m, int *n, int *nrhs, double *a,
|
||||
int *lda, double *b, int *ldb, double *work, int *lwork, int *info);
|
||||
|
||||
int dgesv_(int *n, int *nrhs, double *a, int *lda, int *ipiv,
|
||||
double *b, int *ldb, int *info);
|
||||
|
||||
int dgetrf_(int *m, int *n, double *a, int *lda, int *ipiv,
|
||||
int *info);
|
||||
|
||||
int dposv_(char *uplo, int *n, int *nrhs, double *a, int *
|
||||
lda, double *b, int *ldb, int *info);
|
||||
|
||||
int dpotrf_(char *uplo, int *n, double *a, int *lda, int *
|
||||
info);
|
||||
|
||||
int sgels_(char *trans, int *m, int *n, int *nrhs, float *a,
|
||||
int *lda, float *b, int *ldb, float *work, int *lwork, int *info);
|
||||
|
||||
int sgeev_(char *jobvl, char *jobvr, int *n, float *a, int *
|
||||
lda, float *wr, float *wi, float *vl, int *ldvl, float *vr, int *
|
||||
ldvr, float *work, int *lwork, int *info);
|
||||
|
||||
int sgeqrf_(int *m, int *n, float *a, int *lda, float *tau,
|
||||
float *work, int *lwork, int *info);
|
||||
|
||||
int sgesv_(int *n, int *nrhs, float *a, int *lda, int *ipiv,
|
||||
float *b, int *ldb, int *info);
|
||||
|
||||
int sgetrf_(int *m, int *n, float *a, int *lda, int *ipiv,
|
||||
int *info);
|
||||
|
||||
int sposv_(char *uplo, int *n, int *nrhs, float *a, int *
|
||||
lda, float *b, int *ldb, int *info);
|
||||
|
||||
int spotrf_(char *uplo, int *n, float *a, int *lda, int *
|
||||
info);
|
||||
|
||||
int sgesdd_(char *jobz, int *m, int *n, float *a, int *lda,
|
||||
float *s, float *u, int *ldu, float *vt, int *ldvt, float *work,
|
||||
int *lwork, int *iwork, int *info);
|
||||
|
||||
#ifdef __cplusplus
|
||||
}
|
||||
#endif
|
||||
|
||||
#endif /* __CBLAS_H__ */
|
||||
Vendored
-129
@@ -1,129 +0,0 @@
|
||||
/* f2c.h -- Standard Fortran to C header file */
|
||||
|
||||
/** barf [ba:rf] 2. "He suggested using FORTRAN, and everybody barfed."
|
||||
|
||||
- From The Shogakukan DICTIONARY OF NEW ENGLISH (Second edition) */
|
||||
|
||||
#ifndef __F2C_H__
|
||||
#define __F2C_H__
|
||||
|
||||
#include <assert.h>
|
||||
#include <math.h>
|
||||
#include <ctype.h>
|
||||
#include <stdlib.h>
|
||||
#include <string.h>
|
||||
#include <stdio.h>
|
||||
|
||||
#include "cblas.h"
|
||||
#include "lapack.h"
|
||||
|
||||
#ifdef __cplusplus
|
||||
extern "C" {
|
||||
#endif
|
||||
|
||||
#undef complex
|
||||
|
||||
typedef int integer;
|
||||
typedef unsigned int uinteger;
|
||||
typedef char *address;
|
||||
typedef short int shortint;
|
||||
typedef float real;
|
||||
typedef double doublereal;
|
||||
typedef lapack_complex complex;
|
||||
typedef lapack_doublecomplex doublecomplex;
|
||||
typedef int logical;
|
||||
typedef short int shortlogical;
|
||||
typedef char logical1;
|
||||
typedef char integer1;
|
||||
|
||||
#define TRUE_ (1)
|
||||
#define FALSE_ (0)
|
||||
|
||||
#ifndef abs
|
||||
#define abs(x) ((x) >= 0 ? (x) : -(x))
|
||||
#endif
|
||||
#define dabs(x) (double)abs(x)
|
||||
#ifndef min
|
||||
#define min(a,b) ((a) <= (b) ? (a) : (b))
|
||||
#endif
|
||||
#ifndef max
|
||||
#define max(a,b) ((a) >= (b) ? (a) : (b))
|
||||
#endif
|
||||
#define dmin(a,b) (double)min(a,b)
|
||||
#define dmax(a,b) (double)max(a,b)
|
||||
#define bit_test(a,b) ((a) >> (b) & 1)
|
||||
#define bit_clear(a,b) ((a) & ~((uinteger)1 << (b)))
|
||||
#define bit_set(a,b) ((a) | ((uinteger)1 << (b)))
|
||||
|
||||
static __inline double r_lg10(float *x)
|
||||
{
|
||||
return 0.43429448190325182765*log(*x);
|
||||
}
|
||||
|
||||
static __inline double d_lg10(double *x)
|
||||
{
|
||||
return 0.43429448190325182765*log(*x);
|
||||
}
|
||||
|
||||
static __inline double d_sign(double *a, double *b)
|
||||
{
|
||||
double x = fabs(*a);
|
||||
return *b >= 0 ? x : -x;
|
||||
}
|
||||
|
||||
static __inline double r_sign(float *a, float *b)
|
||||
{
|
||||
double x = fabs((double)*a);
|
||||
return *b >= 0 ? x : -x;
|
||||
}
|
||||
|
||||
static __inline int i_nint(float *x)
|
||||
{
|
||||
return (int)(*x >= 0 ? floor(*x + .5) : -floor(.5 - *x));
|
||||
}
|
||||
|
||||
int pow_ii(int *ap, int *bp);
|
||||
double pow_di(double *ap, int *bp);
|
||||
static __inline double pow_ri(float *ap, int *bp)
|
||||
{
|
||||
double apd = *ap;
|
||||
return pow_di(&apd, bp);
|
||||
}
|
||||
static __inline double pow_dd(double *ap, double *bp)
|
||||
{
|
||||
return pow(*ap, *bp);
|
||||
}
|
||||
|
||||
static __inline void d_cnjg(doublecomplex *r, doublecomplex *z)
|
||||
{
|
||||
double zi = z->i;
|
||||
r->r = z->r;
|
||||
r->i = -zi;
|
||||
}
|
||||
|
||||
static __inline void r_cnjg(complex *r, complex *z)
|
||||
{
|
||||
float zi = z->i;
|
||||
r->r = z->r;
|
||||
r->i = -zi;
|
||||
}
|
||||
|
||||
static __inline int s_copy(char *a, char *b, int maxlen)
|
||||
{
|
||||
strncpy(a, b, maxlen);
|
||||
a[maxlen] = '\0';
|
||||
return 0;
|
||||
}
|
||||
|
||||
int s_cat(char *lp, char **rpp, int* rnp, int *np);
|
||||
int s_cmp(char *a0, char *b0);
|
||||
static __inline int i_len(char* s)
|
||||
{
|
||||
return (int)strlen(s);
|
||||
}
|
||||
|
||||
#ifdef __cplusplus
|
||||
}
|
||||
#endif
|
||||
|
||||
#endif
|
||||
Vendored
-386
@@ -1,386 +0,0 @@
|
||||
// this is auto-generated header for Lapack subset
|
||||
#ifndef __CLAPACK_H__
|
||||
#define __CLAPACK_H__
|
||||
|
||||
#include "cblas.h"
|
||||
|
||||
#ifdef __cplusplus
|
||||
extern "C" {
|
||||
#endif
|
||||
|
||||
int cgemm_(char *transa, char *transb, int *m, int *n, int *
|
||||
k, lapack_complex *alpha, lapack_complex *a, int *lda, lapack_complex *b, int *ldb,
|
||||
lapack_complex *beta, lapack_complex *c__, int *ldc);
|
||||
|
||||
int daxpy_(int *n, double *da, double *dx, int *incx, double
|
||||
*dy, int *incy);
|
||||
|
||||
int dbdsdc_(char *uplo, char *compq, int *n, double *d__,
|
||||
double *e, double *u, int *ldu, double *vt, int *ldvt, double *q, int
|
||||
*iq, double *work, int *iwork, int *info);
|
||||
|
||||
int dbdsqr_(char *uplo, int *n, int *ncvt, int *nru, int *
|
||||
ncc, double *d__, double *e, double *vt, int *ldvt, double *u, int *
|
||||
ldu, double *c__, int *ldc, double *work, int *info);
|
||||
|
||||
int dcombssq_(double *v1, double *v2);
|
||||
|
||||
int dcopy_(int *n, double *dx, int *incx, double *dy, int *
|
||||
incy);
|
||||
|
||||
double ddot_(int *n, double *dx, int *incx, double *dy, int *incy);
|
||||
|
||||
int dgebak_(char *job, char *side, int *n, int *ilo, int *
|
||||
ihi, double *scale, int *m, double *v, int *ldv, int *info);
|
||||
|
||||
int dgebal_(char *job, int *n, double *a, int *lda, int *ilo,
|
||||
int *ihi, double *scale, int *info);
|
||||
|
||||
int dgebd2_(int *m, int *n, double *a, int *lda, double *d__,
|
||||
double *e, double *tauq, double *taup, double *work, int *info);
|
||||
|
||||
int dgebrd_(int *m, int *n, double *a, int *lda, double *d__,
|
||||
double *e, double *tauq, double *taup, double *work, int *lwork, int
|
||||
*info);
|
||||
|
||||
int dgeev_(char *jobvl, char *jobvr, int *n, double *a, int *
|
||||
lda, double *wr, double *wi, double *vl, int *ldvl, double *vr, int *
|
||||
ldvr, double *work, int *lwork, int *info);
|
||||
|
||||
int dgehd2_(int *n, int *ilo, int *ihi, double *a, int *lda,
|
||||
double *tau, double *work, int *info);
|
||||
|
||||
int dgehrd_(int *n, int *ilo, int *ihi, double *a, int *lda,
|
||||
double *tau, double *work, int *lwork, int *info);
|
||||
|
||||
int dgelq2_(int *m, int *n, double *a, int *lda, double *tau,
|
||||
double *work, int *info);
|
||||
|
||||
int dgelqf_(int *m, int *n, double *a, int *lda, double *tau,
|
||||
double *work, int *lwork, int *info);
|
||||
|
||||
int dgemm_(char *transa, char *transb, int *m, int *n, int *
|
||||
k, double *alpha, double *a, int *lda, double *b, int *ldb, double *
|
||||
beta, double *c__, int *ldc);
|
||||
|
||||
int dgemv_(char *trans, int *m, int *n, double *alpha,
|
||||
double *a, int *lda, double *x, int *incx, double *beta, double *y,
|
||||
int *incy);
|
||||
|
||||
int dgeqr2_(int *m, int *n, double *a, int *lda, double *tau,
|
||||
double *work, int *info);
|
||||
|
||||
int dgeqrf_(int *m, int *n, double *a, int *lda, double *tau,
|
||||
double *work, int *lwork, int *info);
|
||||
|
||||
int dger_(int *m, int *n, double *alpha, double *x, int *
|
||||
incx, double *y, int *incy, double *a, int *lda);
|
||||
|
||||
int dgesdd_(char *jobz, int *m, int *n, double *a, int *lda,
|
||||
double *s, double *u, int *ldu, double *vt, int *ldvt, double *work,
|
||||
int *lwork, int *iwork, int *info);
|
||||
|
||||
int dhseqr_(char *job, char *compz, int *n, int *ilo, int *
|
||||
ihi, double *h__, int *ldh, double *wr, double *wi, double *z__, int *
|
||||
ldz, double *work, int *lwork, int *info);
|
||||
|
||||
int disnan_(double *din);
|
||||
|
||||
// "small" is a macro defined in Windows headers: https://stackoverflow.com/a/27794577
|
||||
#ifdef small
|
||||
#undef small
|
||||
#endif
|
||||
|
||||
int dlabad_(double *small, double *large);
|
||||
|
||||
int dlabrd_(int *m, int *n, int *nb, double *a, int *lda,
|
||||
double *d__, double *e, double *tauq, double *taup, double *x, int *
|
||||
ldx, double *y, int *ldy);
|
||||
|
||||
int dlacpy_(char *uplo, int *m, int *n, double *a, int *lda,
|
||||
double *b, int *ldb);
|
||||
|
||||
int dladiv1_(double *a, double *b, double *c__, double *d__,
|
||||
double *p, double *q);
|
||||
|
||||
double dladiv2_(double *a, double *b, double *c__, double *d__, double *r__,
|
||||
double *t);
|
||||
|
||||
int dladiv_(double *a, double *b, double *c__, double *d__,
|
||||
double *p, double *q);
|
||||
|
||||
int dlaed6_(int *kniter, int *orgati, double *rho, double *
|
||||
d__, double *z__, double *finit, double *tau, int *info);
|
||||
|
||||
int dlaexc_(int *wantq, int *n, double *t, int *ldt, double *
|
||||
q, int *ldq, int *j1, int *n1, int *n2, double *work, int *info);
|
||||
|
||||
int dlahqr_(int *wantt, int *wantz, int *n, int *ilo, int *
|
||||
ihi, double *h__, int *ldh, double *wr, double *wi, int *iloz, int *
|
||||
ihiz, double *z__, int *ldz, int *info);
|
||||
|
||||
int dlahr2_(int *n, int *k, int *nb, double *a, int *lda,
|
||||
double *tau, double *t, int *ldt, double *y, int *ldy);
|
||||
|
||||
int dlaisnan_(double *din1, double *din2);
|
||||
|
||||
int dlaln2_(int *ltrans, int *na, int *nw, double *smin,
|
||||
double *ca, double *a, int *lda, double *d1, double *d2, double *b,
|
||||
int *ldb, double *wr, double *wi, double *x, int *ldx, double *scale,
|
||||
double *xnorm, int *info);
|
||||
|
||||
int dlamrg_(int *n1, int *n2, double *a, int *dtrd1, int *
|
||||
dtrd2, int *index);
|
||||
|
||||
double dlange_(char *norm, int *m, int *n, double *a, int *lda, double *work);
|
||||
|
||||
double dlanst_(char *norm, int *n, double *d__, double *e);
|
||||
|
||||
int dlanv2_(double *a, double *b, double *c__, double *d__,
|
||||
double *rt1r, double *rt1i, double *rt2r, double *rt2i, double *cs,
|
||||
double *sn);
|
||||
|
||||
double dlapy2_(double *x, double *y);
|
||||
|
||||
int dlaqr0_(int *wantt, int *wantz, int *n, int *ilo, int *
|
||||
ihi, double *h__, int *ldh, double *wr, double *wi, int *iloz, int *
|
||||
ihiz, double *z__, int *ldz, double *work, int *lwork, int *info);
|
||||
|
||||
int dlaqr1_(int *n, double *h__, int *ldh, double *sr1,
|
||||
double *si1, double *sr2, double *si2, double *v);
|
||||
|
||||
int dlaqr2_(int *wantt, int *wantz, int *n, int *ktop, int *
|
||||
kbot, int *nw, double *h__, int *ldh, int *iloz, int *ihiz, double *
|
||||
z__, int *ldz, int *ns, int *nd, double *sr, double *si, double *v,
|
||||
int *ldv, int *nh, double *t, int *ldt, int *nv, double *wv, int *
|
||||
ldwv, double *work, int *lwork);
|
||||
|
||||
int dlaqr3_(int *wantt, int *wantz, int *n, int *ktop, int *
|
||||
kbot, int *nw, double *h__, int *ldh, int *iloz, int *ihiz, double *
|
||||
z__, int *ldz, int *ns, int *nd, double *sr, double *si, double *v,
|
||||
int *ldv, int *nh, double *t, int *ldt, int *nv, double *wv, int *
|
||||
ldwv, double *work, int *lwork);
|
||||
|
||||
int dlaqr4_(int *wantt, int *wantz, int *n, int *ilo, int *
|
||||
ihi, double *h__, int *ldh, double *wr, double *wi, int *iloz, int *
|
||||
ihiz, double *z__, int *ldz, double *work, int *lwork, int *info);
|
||||
|
||||
int dlaqr5_(int *wantt, int *wantz, int *kacc22, int *n, int
|
||||
*ktop, int *kbot, int *nshfts, double *sr, double *si, double *h__,
|
||||
int *ldh, int *iloz, int *ihiz, double *z__, int *ldz, double *v, int
|
||||
*ldv, double *u, int *ldu, int *nv, double *wv, int *ldwv, int *nh,
|
||||
double *wh, int *ldwh);
|
||||
|
||||
int dlarf_(char *side, int *m, int *n, double *v, int *incv,
|
||||
double *tau, double *c__, int *ldc, double *work);
|
||||
|
||||
int dlarfb_(char *side, char *trans, char *direct, char *
|
||||
storev, int *m, int *n, int *k, double *v, int *ldv, double *t, int *
|
||||
ldt, double *c__, int *ldc, double *work, int *ldwork);
|
||||
|
||||
int dlarfg_(int *n, double *alpha, double *x, int *incx,
|
||||
double *tau);
|
||||
|
||||
int dlarft_(char *direct, char *storev, int *n, int *k,
|
||||
double *v, int *ldv, double *tau, double *t, int *ldt);
|
||||
|
||||
int dlarfx_(char *side, int *m, int *n, double *v, double *
|
||||
tau, double *c__, int *ldc, double *work);
|
||||
|
||||
int dlartg_(double *f, double *g, double *cs, double *sn,
|
||||
double *r__);
|
||||
|
||||
int dlas2_(double *f, double *g, double *h__, double *ssmin,
|
||||
double *ssmax);
|
||||
|
||||
int dlascl_(char *type__, int *kl, int *ku, double *cfrom,
|
||||
double *cto, int *m, int *n, double *a, int *lda, int *info);
|
||||
|
||||
int dlasd0_(int *n, int *sqre, double *d__, double *e,
|
||||
double *u, int *ldu, double *vt, int *ldvt, int *smlsiz, int *iwork,
|
||||
double *work, int *info);
|
||||
|
||||
int dlasd1_(int *nl, int *nr, int *sqre, double *d__, double
|
||||
*alpha, double *beta, double *u, int *ldu, double *vt, int *ldvt, int
|
||||
*idxq, int *iwork, double *work, int *info);
|
||||
|
||||
int dlasd2_(int *nl, int *nr, int *sqre, int *k, double *d__,
|
||||
double *z__, double *alpha, double *beta, double *u, int *ldu,
|
||||
double *vt, int *ldvt, double *dsigma, double *u2, int *ldu2, double *
|
||||
vt2, int *ldvt2, int *idxp, int *idx, int *idxc, int *idxq, int *
|
||||
coltyp, int *info);
|
||||
|
||||
int dlasd3_(int *nl, int *nr, int *sqre, int *k, double *d__,
|
||||
double *q, int *ldq, double *dsigma, double *u, int *ldu, double *u2,
|
||||
int *ldu2, double *vt, int *ldvt, double *vt2, int *ldvt2, int *idxc,
|
||||
int *ctot, double *z__, int *info);
|
||||
|
||||
int dlasd4_(int *n, int *i__, double *d__, double *z__,
|
||||
double *delta, double *rho, double *sigma, double *work, int *info);
|
||||
|
||||
int dlasd5_(int *i__, double *d__, double *z__, double *
|
||||
delta, double *rho, double *dsigma, double *work);
|
||||
|
||||
int dlasd6_(int *icompq, int *nl, int *nr, int *sqre, double
|
||||
*d__, double *vf, double *vl, double *alpha, double *beta, int *idxq,
|
||||
int *perm, int *givptr, int *givcol, int *ldgcol, double *givnum, int
|
||||
*ldgnum, double *poles, double *difl, double *difr, double *z__, int *
|
||||
k, double *c__, double *s, double *work, int *iwork, int *info);
|
||||
|
||||
int dlasd7_(int *icompq, int *nl, int *nr, int *sqre, int *k,
|
||||
double *d__, double *z__, double *zw, double *vf, double *vfw,
|
||||
double *vl, double *vlw, double *alpha, double *beta, double *dsigma,
|
||||
int *idx, int *idxp, int *idxq, int *perm, int *givptr, int *givcol,
|
||||
int *ldgcol, double *givnum, int *ldgnum, double *c__, double *s, int
|
||||
*info);
|
||||
|
||||
int dlasd8_(int *icompq, int *k, double *d__, double *z__,
|
||||
double *vf, double *vl, double *difl, double *difr, int *lddifr,
|
||||
double *dsigma, double *work, int *info);
|
||||
|
||||
int dlasda_(int *icompq, int *smlsiz, int *n, int *sqre,
|
||||
double *d__, double *e, double *u, int *ldu, double *vt, int *k,
|
||||
double *difl, double *difr, double *z__, double *poles, int *givptr,
|
||||
int *givcol, int *ldgcol, int *perm, double *givnum, double *c__,
|
||||
double *s, double *work, int *iwork, int *info);
|
||||
|
||||
int dlasdq_(char *uplo, int *sqre, int *n, int *ncvt, int *
|
||||
nru, int *ncc, double *d__, double *e, double *vt, int *ldvt, double *
|
||||
u, int *ldu, double *c__, int *ldc, double *work, int *info);
|
||||
|
||||
int dlasdt_(int *n, int *lvl, int *nd, int *inode, int *
|
||||
ndiml, int *ndimr, int *msub);
|
||||
|
||||
int dlaset_(char *uplo, int *m, int *n, double *alpha,
|
||||
double *beta, double *a, int *lda);
|
||||
|
||||
int dlasq1_(int *n, double *d__, double *e, double *work,
|
||||
int *info);
|
||||
|
||||
int dlasq2_(int *n, double *z__, int *info);
|
||||
|
||||
int dlasq3_(int *i0, int *n0, double *z__, int *pp, double *
|
||||
dmin__, double *sigma, double *desig, double *qmax, int *nfail, int *
|
||||
iter, int *ndiv, int *ieee, int *ttype, double *dmin1, double *dmin2,
|
||||
double *dn, double *dn1, double *dn2, double *g, double *tau);
|
||||
|
||||
int dlasq4_(int *i0, int *n0, double *z__, int *pp, int *
|
||||
n0in, double *dmin__, double *dmin1, double *dmin2, double *dn,
|
||||
double *dn1, double *dn2, double *tau, int *ttype, double *g);
|
||||
|
||||
int dlasq5_(int *i0, int *n0, double *z__, int *pp, double *
|
||||
tau, double *sigma, double *dmin__, double *dmin1, double *dmin2,
|
||||
double *dn, double *dnm1, double *dnm2, int *ieee, double *eps);
|
||||
|
||||
int dlasq6_(int *i0, int *n0, double *z__, int *pp, double *
|
||||
dmin__, double *dmin1, double *dmin2, double *dn, double *dnm1,
|
||||
double *dnm2);
|
||||
|
||||
int dlasr_(char *side, char *pivot, char *direct, int *m,
|
||||
int *n, double *c__, double *s, double *a, int *lda);
|
||||
|
||||
int dlasrt_(char *id, int *n, double *d__, int *info);
|
||||
|
||||
int dlassq_(int *n, double *x, int *incx, double *scale,
|
||||
double *sumsq);
|
||||
|
||||
int dlasv2_(double *f, double *g, double *h__, double *ssmin,
|
||||
double *ssmax, double *snr, double *csr, double *snl, double *csl);
|
||||
|
||||
int dlasy2_(int *ltranl, int *ltranr, int *isgn, int *n1,
|
||||
int *n2, double *tl, int *ldtl, double *tr, int *ldtr, double *b, int
|
||||
*ldb, double *scale, double *x, int *ldx, double *xnorm, int *info);
|
||||
|
||||
double dnrm2_(int *n, double *x, int *incx);
|
||||
|
||||
int dorg2r_(int *m, int *n, int *k, double *a, int *lda,
|
||||
double *tau, double *work, int *info);
|
||||
|
||||
int dorgbr_(char *vect, int *m, int *n, int *k, double *a,
|
||||
int *lda, double *tau, double *work, int *lwork, int *info);
|
||||
|
||||
int dorghr_(int *n, int *ilo, int *ihi, double *a, int *lda,
|
||||
double *tau, double *work, int *lwork, int *info);
|
||||
|
||||
int dorgl2_(int *m, int *n, int *k, double *a, int *lda,
|
||||
double *tau, double *work, int *info);
|
||||
|
||||
int dorglq_(int *m, int *n, int *k, double *a, int *lda,
|
||||
double *tau, double *work, int *lwork, int *info);
|
||||
|
||||
int dorgqr_(int *m, int *n, int *k, double *a, int *lda,
|
||||
double *tau, double *work, int *lwork, int *info);
|
||||
|
||||
int dorm2r_(char *side, char *trans, int *m, int *n, int *k,
|
||||
double *a, int *lda, double *tau, double *c__, int *ldc, double *work,
|
||||
int *info);
|
||||
|
||||
int dormbr_(char *vect, char *side, char *trans, int *m, int
|
||||
*n, int *k, double *a, int *lda, double *tau, double *c__, int *ldc,
|
||||
double *work, int *lwork, int *info);
|
||||
|
||||
int dormhr_(char *side, char *trans, int *m, int *n, int *
|
||||
ilo, int *ihi, double *a, int *lda, double *tau, double *c__, int *
|
||||
ldc, double *work, int *lwork, int *info);
|
||||
|
||||
int dorml2_(char *side, char *trans, int *m, int *n, int *k,
|
||||
double *a, int *lda, double *tau, double *c__, int *ldc, double *work,
|
||||
int *info);
|
||||
|
||||
int dormlq_(char *side, char *trans, int *m, int *n, int *k,
|
||||
double *a, int *lda, double *tau, double *c__, int *ldc, double *work,
|
||||
int *lwork, int *info);
|
||||
|
||||
int dormqr_(char *side, char *trans, int *m, int *n, int *k,
|
||||
double *a, int *lda, double *tau, double *c__, int *ldc, double *work,
|
||||
int *lwork, int *info);
|
||||
|
||||
int drot_(int *n, double *dx, int *incx, double *dy, int *
|
||||
incy, double *c__, double *s);
|
||||
|
||||
int dscal_(int *n, double *da, double *dx, int *incx);
|
||||
|
||||
int dswap_(int *n, double *dx, int *incx, double *dy, int *
|
||||
incy);
|
||||
|
||||
int dtrevc3_(char *side, char *howmny, int *select, int *n,
|
||||
double *t, int *ldt, double *vl, int *ldvl, double *vr, int *ldvr,
|
||||
int *mm, int *m, double *work, int *lwork, int *info);
|
||||
|
||||
int dtrexc_(char *compq, int *n, double *t, int *ldt, double
|
||||
*q, int *ldq, int *ifst, int *ilst, double *work, int *info);
|
||||
|
||||
int dtrmm_(char *side, char *uplo, char *transa, char *diag,
|
||||
int *m, int *n, double *alpha, double *a, int *lda, double *b, int *
|
||||
ldb);
|
||||
|
||||
int dtrmv_(char *uplo, char *trans, char *diag, int *n,
|
||||
double *a, int *lda, double *x, int *incx);
|
||||
|
||||
int idamax_(int *n, double *dx, int *incx);
|
||||
|
||||
int ieeeck_(int *ispec, float *zero, float *one);
|
||||
|
||||
int iladlc_(int *m, int *n, double *a, int *lda);
|
||||
|
||||
int iladlr_(int *m, int *n, double *a, int *lda);
|
||||
|
||||
int ilaenv_(int *ispec, char *name__, char *opts, int *n1, int *n2, int *n3,
|
||||
int *n4);
|
||||
|
||||
int iparmq_(int *ispec, char *name__, char *opts, int *n, int *ilo, int *ihi,
|
||||
int *lwork);
|
||||
|
||||
int sgemm_(char *transa, char *transb, int *m, int *n, int *
|
||||
k, float *alpha, float *a, int *lda, float *b, int *ldb, float *beta,
|
||||
float *c__, int *ldc);
|
||||
|
||||
int zgemm_(char *transa, char *transb, int *m, int *n, int *
|
||||
k, lapack_doublecomplex *alpha, lapack_doublecomplex *a, int *lda, lapack_doublecomplex *b,
|
||||
int *ldb, lapack_doublecomplex *beta, lapack_doublecomplex *c__, int *ldc);
|
||||
|
||||
#ifdef __cplusplus
|
||||
}
|
||||
#endif
|
||||
|
||||
#endif
|
||||
Vendored
-48
@@ -1,48 +0,0 @@
|
||||
Copyright (c) 1992-2017 The University of Tennessee and The University
|
||||
of Tennessee Research Foundation. All rights
|
||||
reserved.
|
||||
Copyright (c) 2000-2017 The University of California Berkeley. All
|
||||
rights reserved.
|
||||
Copyright (c) 2006-2017 The University of Colorado Denver. All rights
|
||||
reserved.
|
||||
|
||||
$COPYRIGHT$
|
||||
|
||||
Additional copyrights may follow
|
||||
|
||||
$HEADER$
|
||||
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are
|
||||
met:
|
||||
|
||||
- Redistributions of source code must retain the above copyright
|
||||
notice, this list of conditions and the following disclaimer.
|
||||
|
||||
- Redistributions in binary form must reproduce the above copyright
|
||||
notice, this list of conditions and the following disclaimer listed
|
||||
in this license in the documentation and/or other materials
|
||||
provided with the distribution.
|
||||
|
||||
- Neither the name of the copyright holders nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
The copyright holders provide no reassurances that the source code
|
||||
provided does not infringe any patent, copyright, or any other
|
||||
intellectual property rights of third parties. The copyright holders
|
||||
disclaim any liability to any recipient for claims brought against
|
||||
recipient by any third party for infringement of that parties
|
||||
intellectual property rights.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS
|
||||
"AS IS" AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT
|
||||
LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR
|
||||
A PARTICULAR PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT
|
||||
OWNER OR CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL,
|
||||
SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT
|
||||
LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE,
|
||||
DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY
|
||||
THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
|
||||
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
Vendored
-272
@@ -1,272 +0,0 @@
|
||||
appdoc = """
|
||||
This is generator of CLapack subset.
|
||||
The usage:
|
||||
|
||||
1. Make sure you have the special version of f2c installed.
|
||||
Grab it from https://github.com/vpisarev/f2c/tree/for_lapack.
|
||||
2. Download fresh version of Lapack from
|
||||
https://github.com/Reference-LAPACK/lapack.
|
||||
You may choose some specific version or the latest snapshot.
|
||||
3. If necessary, edit "roots" and "banlist" variables in this script, specify the needed and unneeded functions
|
||||
4. From within a working directory run
|
||||
|
||||
$ python3 <opencv_root>/3rdparty/clapack/make_clapack.py <lapack_root>
|
||||
or
|
||||
$ F2C=<path_to_custom_f2c> python3 <opencv_root>/3rdparty/clapack/make_clapack.py <lapack_root>
|
||||
|
||||
it will generate "new_clapack" directory with "include" and "src" subdirectories.
|
||||
5. erase opencv/3rdparty/clapack/src and replace it with new_clapack/src.
|
||||
6. copy new_clapack/include/lapack.h to opencv/3rdparty/clapack/include.
|
||||
7. optionally, edit opencv/3rdparty/clapack/CMakeLists.txt and update CLAPACK_VERSION as needed.
|
||||
|
||||
This is it. Now build it and enjoy.
|
||||
"""
|
||||
|
||||
import glob, re, os, shutil, subprocess, sys
|
||||
|
||||
roots = ["cgemm_", "dgemm_", "sgemm_", "zgemm_",
|
||||
"dgeev_", "dgesdd_",
|
||||
#"dsyevr_",
|
||||
#"dgesv_", "dgetrf_", "dposv_", "dpotrf_", "dgels_", "dgeqrf_",
|
||||
#"sgesv_", "sgetrf_", "sposv_", "spotrf_", "sgels_", "sgeqrf_"
|
||||
]
|
||||
banlist = ["slamch_", "slamc3_", "dlamch_", "dlamc3_", "lsame_", "xerbla_"]
|
||||
|
||||
if len(sys.argv) < 2:
|
||||
print(appdoc)
|
||||
sys.exit(0)
|
||||
|
||||
lapack_root = sys.argv[1]
|
||||
dst_path = "."
|
||||
|
||||
def error(msg):
|
||||
print ("error: " + msg)
|
||||
sys.exit(0)
|
||||
|
||||
def file2fun(fname):
|
||||
return (os.path.basename(fname)[:-2]).upper()
|
||||
|
||||
def print_graph(m):
|
||||
for (k, neighbors) in sorted(m.items()):
|
||||
print (k + " : " + ", ".join(sorted(list(neighbors))))
|
||||
|
||||
blas_path = os.path.join(lapack_root, "BLAS/SRC")
|
||||
lapack_path = os.path.join(lapack_root, "SRC")
|
||||
|
||||
roots = [f[:-1].upper() for f in roots]
|
||||
banlist = [f[:-1].upper() for f in banlist]
|
||||
|
||||
def fun2file(func):
|
||||
filename = func.lower() + ".f"
|
||||
blas_loc = blas_path + "/" + filename
|
||||
lapack_loc = lapack_path + "/" + filename
|
||||
if os.path.exists(blas_loc):
|
||||
return blas_loc
|
||||
elif os.path.exists(lapack_loc):
|
||||
return lapack_loc
|
||||
else:
|
||||
error("neither %s nor %s exist" % (blas_loc, lapack_loc))
|
||||
|
||||
all_files = glob.glob(blas_path + "/*.f") + glob.glob(lapack_path + "/*.f")
|
||||
all_funcs = [file2fun(fname) for fname in all_files]
|
||||
all_funcs_set = set(all_funcs).difference(set(banlist))
|
||||
all_funcs = sorted(list(all_funcs_set))
|
||||
|
||||
func_deps = {}
|
||||
|
||||
#print all_funcs
|
||||
|
||||
words_regexp = re.compile(r'\w+')
|
||||
|
||||
def scan_deps(func):
|
||||
global func_deps
|
||||
if func in func_deps:
|
||||
return
|
||||
func_deps[func] = set([]) # to avoid possibly infinite recursion
|
||||
f = open(fun2file(func), 'rt')
|
||||
deps = []
|
||||
external_mode = False
|
||||
for l in f.readlines():
|
||||
if l.startswith('*'):
|
||||
continue
|
||||
l = l.strip().upper()
|
||||
if l.startswith('EXTERNAL '):
|
||||
external_mode = True
|
||||
elif l.startswith('$') and external_mode:
|
||||
pass
|
||||
else:
|
||||
external_mode = False
|
||||
if not external_mode:
|
||||
continue
|
||||
for w in words_regexp.findall(l):
|
||||
if w in all_funcs_set:
|
||||
deps.append(w)
|
||||
f.close()
|
||||
# remove func from its dependencies
|
||||
deps = set(deps).difference(set([func]))
|
||||
func_deps[func] = deps
|
||||
for d in deps:
|
||||
scan_deps(d)
|
||||
|
||||
for r in roots:
|
||||
scan_deps(r)
|
||||
|
||||
selected_funcs = sorted(func_deps.keys())
|
||||
print ("total files before amalgamation: %d" % len(selected_funcs))
|
||||
|
||||
inv_deps = {}
|
||||
for func in selected_funcs:
|
||||
inv_deps[func] = set([])
|
||||
|
||||
for (func, deps) in func_deps.items():
|
||||
for d in deps:
|
||||
inv_deps[d] = inv_deps[d].union(set([func]))
|
||||
|
||||
#print_graph(inv_deps)
|
||||
|
||||
func_home = {}
|
||||
for func in selected_funcs:
|
||||
func_home[func] = func
|
||||
|
||||
def get_home0(func, func0):
|
||||
used_by = inv_deps[func]
|
||||
if len(used_by) == 1:
|
||||
p = list(used_by)[0]
|
||||
if p != func and p != func0:
|
||||
return get_home0(p, func0)
|
||||
return func
|
||||
return func
|
||||
|
||||
# try to merge some files
|
||||
for func in selected_funcs:
|
||||
func_home[func] = get_home0(func, func)
|
||||
|
||||
# try to merge some files even more
|
||||
for iters in range(100):
|
||||
homes_changed = False
|
||||
for (func, used_by) in inv_deps.items():
|
||||
p0 = func_home[func]
|
||||
n = len(used_by)
|
||||
if n == 1:
|
||||
p = list(used_by)[0]
|
||||
p1 = func_home[p]
|
||||
if p1 != p0:
|
||||
func_home[func] = p1
|
||||
homes_changed = True
|
||||
continue
|
||||
elif n > 1:
|
||||
phomes = set([])
|
||||
for p in used_by:
|
||||
phomes.add(func_home[p])
|
||||
if len(phomes) == 1:
|
||||
p1 = list(phomes)[0]
|
||||
if p1 != p0:
|
||||
func_home[func] = p1
|
||||
homes_changed = True
|
||||
if not homes_changed:
|
||||
break
|
||||
|
||||
res_files = {}
|
||||
for (func, h) in func_home.items():
|
||||
elems = res_files.get(h, set([]))
|
||||
elems.add(func)
|
||||
res_files[h] = elems
|
||||
|
||||
print ("total files after amalgamation: %d" % len(res_files))
|
||||
#print_graph(res_files)
|
||||
|
||||
outdir = os.path.join(dst_path, "new_clapack")
|
||||
outdir_src = os.path.join(outdir, "src")
|
||||
outdir_inc = os.path.join(outdir, "include")
|
||||
|
||||
shutil.rmtree(outdir, ignore_errors=True)
|
||||
try:
|
||||
os.makedirs(outdir_src)
|
||||
except os.error:
|
||||
pass
|
||||
try:
|
||||
os.makedirs(outdir_inc)
|
||||
except os.error:
|
||||
pass
|
||||
|
||||
f2c_appname = os.getenv("F2C", default="f2c")
|
||||
print ("f2c used: %s" % f2c_appname)
|
||||
|
||||
f2c_getver_cmd = f2c_appname + " -v"
|
||||
|
||||
verstr = subprocess.check_output(f2c_getver_cmd.split(' ')).decode("utf-8")
|
||||
if "for_lapack" not in verstr:
|
||||
error("invalid version of f2c\n" + appdoc)
|
||||
|
||||
f2c_flags = "-ctypes -localconst -no-proto"
|
||||
f2c_cmd0 = f2c_appname + " " + f2c_flags
|
||||
f2c_cmd1 = f2c_appname + " -hdr none " + f2c_flags
|
||||
|
||||
lapack_protos = {}
|
||||
extract_fn_regexp = re.compile(r'.+?(\w+)\s*\(')
|
||||
|
||||
def extract_proto(func, csrc):
|
||||
global lapack_protos
|
||||
cname = func.lower() + "_"
|
||||
cfname = func.lower() + ".c"
|
||||
regexp_str = r'\n(?:/\* Subroutine \*/\s*)?\w+\s+\w+\s*\((?:.|\n)+?\)[\s\n]*\{'
|
||||
proto_regexp = re.compile(regexp_str)
|
||||
ps = proto_regexp.findall(csrc)
|
||||
for p in ps:
|
||||
n = p.find("*/")
|
||||
if n < 0:
|
||||
n = 0
|
||||
else:
|
||||
n += 2
|
||||
p = p[n:-1].strip() + ";"
|
||||
fns = extract_fn_regexp.findall(p)
|
||||
if len(fns) != 1:
|
||||
error("prototype of function (%s) when analyzing %s cannot be parsed" % (p, cfname))
|
||||
fn = fns[0]
|
||||
if fn not in lapack_protos:
|
||||
p = re.sub(r'\bcomplex\b', 'lapack_complex', p)
|
||||
p = re.sub(r'\bdoublecomplex\b', 'lapack_doublecomplex', p)
|
||||
lapack_protos[fn] = p
|
||||
|
||||
for (filename, funcs) in sorted(res_files.items()):
|
||||
out = ""
|
||||
f2c_cmd = f2c_cmd0
|
||||
for func in sorted(list(funcs)):
|
||||
ffilename = fun2file(func)
|
||||
print ("running " + f2c_cmd + " on " + ffilename + " ...")
|
||||
ffile = open(ffilename, 'rt')
|
||||
delta_out = subprocess.check_output(f2c_cmd.split(' '), stdin=ffile).decode("utf-8")
|
||||
# remove trailing whitespaces
|
||||
delta_out = '\n'.join([l.rstrip() for l in delta_out.split('\n')])
|
||||
extract_proto(func, delta_out)
|
||||
out += delta_out
|
||||
ffile.close()
|
||||
f2c_cmd = f2c_cmd1
|
||||
outname = os.path.join(outdir_src, filename.lower() + ".c")
|
||||
outfile = open(outname, 'wt')
|
||||
outfile.write(out)
|
||||
outfile.close()
|
||||
|
||||
proto_hdr = """// this is auto-generated header for Lapack subset
|
||||
#ifndef __CLAPACK_H__
|
||||
#define __CLAPACK_H__
|
||||
|
||||
#include "cblas.h"
|
||||
|
||||
#ifdef __cplusplus
|
||||
extern "C" {
|
||||
#endif
|
||||
|
||||
%s
|
||||
|
||||
#ifdef __cplusplus
|
||||
}
|
||||
#endif
|
||||
|
||||
#endif
|
||||
""" % "\n\n".join([p for (n, p) in sorted(lapack_protos.items())])
|
||||
|
||||
proto_hdr_fname = os.path.join(outdir_inc, "lapack.h")
|
||||
f = open(proto_hdr_fname, 'wt')
|
||||
f.write(proto_hdr)
|
||||
f.close()
|
||||
-289
@@ -1,289 +0,0 @@
|
||||
#include "f2c.h"
|
||||
#include <stdarg.h>
|
||||
|
||||
void cblas_cgemm(const CBLAS_LAYOUT layout, const CBLAS_TRANSPOSE TransA,
|
||||
const CBLAS_TRANSPOSE TransB, const int M, const int N,
|
||||
const int K, const void *alpha, const void *A,
|
||||
const int lda, const void *B, const int ldb,
|
||||
const void *beta, void *C, const int ldc)
|
||||
{
|
||||
char TA, TB;
|
||||
|
||||
if( layout == CblasColMajor )
|
||||
{
|
||||
if(TransA == CblasTrans) TA='T';
|
||||
else if ( TransA == CblasConjTrans ) TA='C';
|
||||
else if ( TransA == CblasNoTrans ) TA='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_cgemm", "Illegal TransA setting, %d\n", TransA);
|
||||
return;
|
||||
}
|
||||
|
||||
if(TransB == CblasTrans) TB='T';
|
||||
else if ( TransB == CblasConjTrans ) TB='C';
|
||||
else if ( TransB == CblasNoTrans ) TB='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 3, "cblas_cgemm", "Illegal TransB setting, %d\n", TransB);
|
||||
return;
|
||||
}
|
||||
|
||||
cgemm_(&TA, &TB, (int*)&M, (int*)&N, (int*)&K, (complex*)alpha, (complex*)A, (int*)&lda,
|
||||
(complex*)B, (int*)&ldb, (complex*)beta, (complex*)C, (int*)&ldc);
|
||||
}
|
||||
else if (layout == CblasRowMajor)
|
||||
{
|
||||
if(TransA == CblasTrans) TB='T';
|
||||
else if ( TransA == CblasConjTrans ) TB='C';
|
||||
else if ( TransA == CblasNoTrans ) TB='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_cgemm", "Illegal TransA setting, %d\n", TransA);
|
||||
return;
|
||||
}
|
||||
if(TransB == CblasTrans) TA='T';
|
||||
else if ( TransB == CblasConjTrans ) TA='C';
|
||||
else if ( TransB == CblasNoTrans ) TA='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_cgemm", "Illegal TransB setting, %d\n", TransB);
|
||||
return;
|
||||
}
|
||||
|
||||
cgemm_(&TA, &TB, (int*)&N, (int*)&M, (int*)&K, (complex*)alpha, (complex*)B, (int*)&ldb,
|
||||
(complex*)A, (int*)&lda, (complex*)beta, (complex*)C, (int*)&ldc);
|
||||
}
|
||||
else cblas_xerbla(layout, 1, "cblas_cgemm", "Illegal layout setting, %d\n", layout);
|
||||
}
|
||||
|
||||
void cblas_dgemm(const CBLAS_LAYOUT layout, const CBLAS_TRANSPOSE TransA,
|
||||
const CBLAS_TRANSPOSE TransB, const int M, const int N,
|
||||
const int K, const double alpha, const double *A,
|
||||
const int lda, const double *B, const int ldb,
|
||||
const double beta, double *C, const int ldc)
|
||||
{
|
||||
char TA, TB;
|
||||
|
||||
if( layout == CblasColMajor )
|
||||
{
|
||||
if(TransA == CblasTrans) TA='T';
|
||||
else if ( TransA == CblasConjTrans ) TA='C';
|
||||
else if ( TransA == CblasNoTrans ) TA='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_dgemm", "Illegal TransA setting, %d\n", TransA);
|
||||
return;
|
||||
}
|
||||
|
||||
if(TransB == CblasTrans) TB='T';
|
||||
else if ( TransB == CblasConjTrans ) TB='C';
|
||||
else if ( TransB == CblasNoTrans ) TB='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 3, "cblas_dgemm", "Illegal TransB setting, %d\n", TransB);
|
||||
return;
|
||||
}
|
||||
|
||||
dgemm_(&TA, &TB, (int*)&M, (int*)&N, (int*)&K, (double*)&alpha, (double*)A, (int*)&lda,
|
||||
(double*)B, (int*)&ldb, (double*)&beta, (double*)C, (int*)&ldc);
|
||||
}
|
||||
else if (layout == CblasRowMajor)
|
||||
{
|
||||
if(TransA == CblasTrans) TB='T';
|
||||
else if ( TransA == CblasConjTrans ) TB='C';
|
||||
else if ( TransA == CblasNoTrans ) TB='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_dgemm", "Illegal TransA setting, %d\n", TransA);
|
||||
return;
|
||||
}
|
||||
if(TransB == CblasTrans) TA='T';
|
||||
else if ( TransB == CblasConjTrans ) TA='C';
|
||||
else if ( TransB == CblasNoTrans ) TA='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_dgemm", "Illegal TransB setting, %d\n", TransB);
|
||||
return;
|
||||
}
|
||||
|
||||
dgemm_(&TA, &TB, (int*)&N, (int*)&M, (int*)&K, (double*)&alpha, (double*)B, (int*)&ldb,
|
||||
(double*)A, (int*)&lda, (double*)&beta, (double*)C, (int*)&ldc);
|
||||
}
|
||||
else cblas_xerbla(layout, 1, "cblas_dgemm", "Illegal layout setting, %d\n", layout);
|
||||
}
|
||||
|
||||
|
||||
void cblas_sgemm(const CBLAS_LAYOUT layout, const CBLAS_TRANSPOSE TransA,
|
||||
const CBLAS_TRANSPOSE TransB, const int M, const int N,
|
||||
const int K, const float alpha, const float *A,
|
||||
const int lda, const float *B, const int ldb,
|
||||
const float beta, float *C, const int ldc)
|
||||
{
|
||||
char TA, TB;
|
||||
|
||||
if( layout == CblasColMajor )
|
||||
{
|
||||
if(TransA == CblasTrans) TA='T';
|
||||
else if ( TransA == CblasConjTrans ) TA='C';
|
||||
else if ( TransA == CblasNoTrans ) TA='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_sgemm", "Illegal TransA setting, %d\n", TransA);
|
||||
return;
|
||||
}
|
||||
|
||||
if(TransB == CblasTrans) TB='T';
|
||||
else if ( TransB == CblasConjTrans ) TB='C';
|
||||
else if ( TransB == CblasNoTrans ) TB='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 3, "cblas_sgemm", "Illegal TransB setting, %d\n", TransB);
|
||||
return;
|
||||
}
|
||||
|
||||
sgemm_(&TA, &TB, (int*)&M, (int*)&N, (int*)&K, (float*)&alpha, (float*)A, (int*)&lda,
|
||||
(float*)B, (int*)&ldb, (float*)&beta, (float*)C, (int*)&ldc);
|
||||
}
|
||||
else if (layout == CblasRowMajor)
|
||||
{
|
||||
if(TransA == CblasTrans) TB='T';
|
||||
else if ( TransA == CblasConjTrans ) TB='C';
|
||||
else if ( TransA == CblasNoTrans ) TB='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_sgemm", "Illegal TransA setting, %d\n", TransA);
|
||||
return;
|
||||
}
|
||||
if(TransB == CblasTrans) TA='T';
|
||||
else if ( TransB == CblasConjTrans ) TA='C';
|
||||
else if ( TransB == CblasNoTrans ) TA='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_sgemm", "Illegal TransB setting, %d\n", TransB);
|
||||
return;
|
||||
}
|
||||
|
||||
sgemm_(&TA, &TB, (int*)&N, (int*)&M, (int*)&K, (float*)&alpha, (float*)B, (int*)&ldb,
|
||||
(float*)A, (int*)&lda, (float*)&beta, (float*)C, (int*)&ldc);
|
||||
}
|
||||
else cblas_xerbla(layout, 1, "cblas_sgemm", "Illegal layout setting, %d\n", layout);
|
||||
}
|
||||
|
||||
void cblas_zgemm(const CBLAS_LAYOUT layout, const CBLAS_TRANSPOSE TransA,
|
||||
const CBLAS_TRANSPOSE TransB, const int M, const int N,
|
||||
const int K, const void *alpha, const void *A,
|
||||
const int lda, const void *B, const int ldb,
|
||||
const void *beta, void *C, const int ldc)
|
||||
{
|
||||
char TA, TB;
|
||||
|
||||
if( layout == CblasColMajor )
|
||||
{
|
||||
if(TransA == CblasTrans) TA='T';
|
||||
else if ( TransA == CblasConjTrans ) TA='C';
|
||||
else if ( TransA == CblasNoTrans ) TA='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_zgemm", "Illegal TransA setting, %d\n", TransA);
|
||||
return;
|
||||
}
|
||||
|
||||
if(TransB == CblasTrans) TB='T';
|
||||
else if ( TransB == CblasConjTrans ) TB='C';
|
||||
else if ( TransB == CblasNoTrans ) TB='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 3, "cblas_zgemm", "Illegal TransB setting, %d\n", TransB);
|
||||
return;
|
||||
}
|
||||
|
||||
zgemm_(&TA, &TB, (int*)&M, (int*)&N, (int*)&K, (doublecomplex*)alpha, (doublecomplex*)A, (int*)&lda,
|
||||
(doublecomplex*)B, (int*)&ldb, (doublecomplex*)beta, (doublecomplex*)C, (int*)&ldc);
|
||||
}
|
||||
else if (layout == CblasRowMajor)
|
||||
{
|
||||
if(TransA == CblasTrans) TB='T';
|
||||
else if ( TransA == CblasConjTrans ) TB='C';
|
||||
else if ( TransA == CblasNoTrans ) TB='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_zgemm", "Illegal TransA setting, %d\n", TransA);
|
||||
return;
|
||||
}
|
||||
if(TransB == CblasTrans) TA='T';
|
||||
else if ( TransB == CblasConjTrans ) TA='C';
|
||||
else if ( TransB == CblasNoTrans ) TA='N';
|
||||
else
|
||||
{
|
||||
cblas_xerbla(layout, 2, "cblas_zgemm", "Illegal TransB setting, %d\n", TransB);
|
||||
return;
|
||||
}
|
||||
|
||||
zgemm_(&TA, &TB, (int*)&N, (int*)&M, (int*)&K, (doublecomplex*)alpha, (doublecomplex*)B, (int*)&ldb,
|
||||
(doublecomplex*)A, (int*)&lda, (doublecomplex*)beta, (doublecomplex*)C, (int*)&ldc);
|
||||
}
|
||||
else cblas_xerbla(layout, 1, "cblas_zgemm", "Illegal layout setting, %d\n", layout);
|
||||
}
|
||||
|
||||
void cblas_xerbla(const CBLAS_LAYOUT layout, int info, const char *rout, const char *form, ...)
|
||||
{
|
||||
extern int RowMajorStrg;
|
||||
char empty[1] = "";
|
||||
va_list argptr;
|
||||
|
||||
va_start(argptr, form);
|
||||
|
||||
if (layout == CblasRowMajor)
|
||||
{
|
||||
if (strstr(rout,"gemm") != 0)
|
||||
{
|
||||
if (info == 5 ) info = 4;
|
||||
else if (info == 4 ) info = 5;
|
||||
else if (info == 11) info = 9;
|
||||
else if (info == 9 ) info = 11;
|
||||
}
|
||||
else if (strstr(rout,"symm") != 0 || strstr(rout,"hemm") != 0)
|
||||
{
|
||||
if (info == 5 ) info = 4;
|
||||
else if (info == 4 ) info = 5;
|
||||
}
|
||||
else if (strstr(rout,"trmm") != 0 || strstr(rout,"trsm") != 0)
|
||||
{
|
||||
if (info == 7 ) info = 6;
|
||||
else if (info == 6 ) info = 7;
|
||||
}
|
||||
else if (strstr(rout,"gemv") != 0)
|
||||
{
|
||||
if (info == 4) info = 3;
|
||||
else if (info == 3) info = 4;
|
||||
}
|
||||
else if (strstr(rout,"gbmv") != 0)
|
||||
{
|
||||
if (info == 4) info = 3;
|
||||
else if (info == 3) info = 4;
|
||||
else if (info == 6) info = 5;
|
||||
else if (info == 5) info = 6;
|
||||
}
|
||||
else if (strstr(rout,"ger") != 0)
|
||||
{
|
||||
if (info == 3) info = 2;
|
||||
else if (info == 2) info = 3;
|
||||
else if (info == 8) info = 6;
|
||||
else if (info == 6) info = 8;
|
||||
}
|
||||
else if ( (strstr(rout,"her2") != 0 || strstr(rout,"hpr2") != 0)
|
||||
&& strstr(rout,"her2k") == 0 )
|
||||
{
|
||||
if (info == 8) info = 6;
|
||||
else if (info == 6) info = 8;
|
||||
}
|
||||
}
|
||||
if (info)
|
||||
fprintf(stderr, "Parameter %d to routine %s was incorrect\n", info, rout);
|
||||
vfprintf(stderr, form, argptr);
|
||||
va_end(argptr);
|
||||
if (info && !info)
|
||||
xerbla_(empty, &info); /* Force link of our F77 error handler */
|
||||
exit(-1);
|
||||
}
|
||||
-72
@@ -1,72 +0,0 @@
|
||||
#include "f2c.h"
|
||||
#include <float.h>
|
||||
#include <stdio.h>
|
||||
|
||||
/* *********************************************************************** */
|
||||
|
||||
double dlamc3_(double *a, double *b)
|
||||
{
|
||||
/* -- LAPACK auxiliary routine (version 3.1) -- */
|
||||
/* Univ. of Tennessee, Univ. of California Berkeley and NAG Ltd.. */
|
||||
/* November 2006 */
|
||||
|
||||
/* .. Scalar Arguments .. */
|
||||
/* .. */
|
||||
|
||||
/* Purpose */
|
||||
/* ======= */
|
||||
|
||||
/* DLAMC3 is intended to force A and B to be stored prior to doing */
|
||||
/* the addition of A and B , for use in situations where optimizers */
|
||||
/* might hold one of these in a register. */
|
||||
|
||||
/* Arguments */
|
||||
/* ========= */
|
||||
|
||||
/* A (input) DOUBLE PRECISION */
|
||||
/* B (input) DOUBLE PRECISION */
|
||||
/* The values A and B. */
|
||||
|
||||
/* ===================================================================== */
|
||||
|
||||
/* .. Executable Statements .. */
|
||||
|
||||
double ret_val = *a + *b;
|
||||
|
||||
return ret_val;
|
||||
|
||||
/* End of DLAMC3 */
|
||||
|
||||
} /* dlamc3_ */
|
||||
|
||||
|
||||
/* simpler version of dlamch for the case of IEEE754-compliant FPU module by Piotr Luszczek S.
|
||||
taken from http://www.mail-archive.com/numpy-discussion@lists.sourceforge.net/msg02448.html */
|
||||
|
||||
#ifndef DBL_DIGITS
|
||||
#define DBL_DIGITS 53
|
||||
#endif
|
||||
|
||||
static const unsigned char lapack_dlamch_tab0[] =
|
||||
{
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 2, 0, 0, 0, 0, 0, 0, 3, 4, 5, 6, 7, 0, 8, 9, 0, 10, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 2, 0, 0, 0, 0, 0, 0, 3, 4, 5, 6, 7, 0, 8, 9,
|
||||
0, 10, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0
|
||||
};
|
||||
|
||||
const double lapack_dlamch_tab1[] =
|
||||
{
|
||||
0, FLT_RADIX, DBL_EPSILON, DBL_MAX_EXP, DBL_MIN_EXP, DBL_DIGITS, DBL_MAX,
|
||||
DBL_EPSILON*FLT_RADIX, 1, DBL_MIN*(1 + DBL_EPSILON), DBL_MIN
|
||||
};
|
||||
|
||||
double dlamch_(char* cmach)
|
||||
{
|
||||
return lapack_dlamch_tab1[lapack_dlamch_tab0[(unsigned char)cmach[0]]];
|
||||
}
|
||||
-96
@@ -1,96 +0,0 @@
|
||||
#include "f2c.h"
|
||||
|
||||
static const int CLAPACK_NOT_IMPLEMENTED = -1024;
|
||||
|
||||
int sgesdd_(char *jobz, int *m, int *n, float *a, int *lda,
|
||||
float *s, float *u, int *ldu, float *vt, int *ldvt, float *work,
|
||||
int *lwork, int *iwork, int *info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int dgels_(char *trans, int *m, int *n, int *nrhs, double *a,
|
||||
int *lda, double *b, int *ldb, double *work, int *lwork, int *info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int dgesv_(int *n, int *nrhs, double *a, int *lda, int *ipiv,
|
||||
double *b, int *ldb, int *info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int dgetrf_(int *m, int *n, double *a, int *lda, int *ipiv,
|
||||
int *info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int dposv_(char *uplo, int *n, int *nrhs, double *a, int *
|
||||
lda, double *b, int *ldb, int *info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int dpotrf_(char *uplo, int *n, double *a, int *lda, int *
|
||||
info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int sgels_(char *trans, int *m, int *n, int *nrhs, float *a,
|
||||
int *lda, float *b, int *ldb, float *work, int *lwork, int *info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int sgeev_(char *jobvl, char *jobvr, int *n, float *a, int *
|
||||
lda, float *wr, float *wi, float *vl, int *ldvl, float *vr, int *
|
||||
ldvr, float *work, int *lwork, int *info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int sgeqrf_(int *m, int *n, float *a, int *lda, float *tau,
|
||||
float *work, int *lwork, int *info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int sgesv_(int *n, int *nrhs, float *a, int *lda, int *ipiv,
|
||||
float *b, int *ldb, int *info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int sgetrf_(int *m, int *n, float *a, int *lda, int *ipiv,
|
||||
int *info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int sposv_(char *uplo, int *n, int *nrhs, float *a, int *
|
||||
lda, float *b, int *ldb, int *info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int spotrf_(char *uplo, int *n, float *a, int *lda, int *
|
||||
info)
|
||||
{
|
||||
*info = CLAPACK_NOT_IMPLEMENTED;
|
||||
return 0;
|
||||
}
|
||||
-25
@@ -1,25 +0,0 @@
|
||||
#include "f2c.h"
|
||||
|
||||
static const unsigned char lapack_toupper_tab[] =
|
||||
{
|
||||
0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23,
|
||||
24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45,
|
||||
46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67,
|
||||
68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89,
|
||||
90, 91, 92, 93, 94, 95, 96, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79,
|
||||
80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 123, 124, 125, 126, 127, 128, 129, 130, 131,
|
||||
132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149,
|
||||
150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167,
|
||||
168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185,
|
||||
186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203,
|
||||
204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221,
|
||||
222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239,
|
||||
240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255
|
||||
};
|
||||
|
||||
#define lapack_toupper(c) ((char)lapack_toupper_tab[(unsigned char)(c)])
|
||||
|
||||
int lsame_(char *ca, char *cb)
|
||||
{
|
||||
return lapack_toupper(ca[0]) == lapack_toupper(cb[0]);
|
||||
}
|
||||
Vendored
-27
@@ -1,27 +0,0 @@
|
||||
#include "f2c.h"
|
||||
|
||||
double pow_di(double *ap, int *bp)
|
||||
{
|
||||
double p = 1;
|
||||
double x = *ap;
|
||||
int n = *bp;
|
||||
|
||||
if(n != 0)
|
||||
{
|
||||
if(n < 0)
|
||||
{
|
||||
n = -n;
|
||||
x = 1/x;
|
||||
}
|
||||
unsigned u = (unsigned)n;
|
||||
for(;;)
|
||||
{
|
||||
if((u & 1) != 0)
|
||||
p *= x;
|
||||
if((u >>= 1) == 0)
|
||||
break;
|
||||
x *= x;
|
||||
}
|
||||
}
|
||||
return p;
|
||||
}
|
||||
Vendored
-25
@@ -1,25 +0,0 @@
|
||||
#include "f2c.h"
|
||||
|
||||
int pow_ii(int *ap, int *bp)
|
||||
{
|
||||
int p;
|
||||
int x = *ap;
|
||||
int n = *bp;
|
||||
|
||||
if (n <= 0) {
|
||||
if (n == 0 || x == 1)
|
||||
return 1;
|
||||
return x != -1 ? 0 : (n & 1) ? -1 : 1;
|
||||
}
|
||||
unsigned u = (unsigned)n;
|
||||
for(p = 1; ; )
|
||||
{
|
||||
if(u & 01)
|
||||
p *= x;
|
||||
if(u >>= 1)
|
||||
x *= x;
|
||||
else
|
||||
break;
|
||||
}
|
||||
return p;
|
||||
}
|
||||
Vendored
-22
@@ -1,22 +0,0 @@
|
||||
/* Unless compiled with -DNO_OVERWRITE, this variant of s_cat allows the
|
||||
* target of a concatenation to appear on its right-hand side (contrary
|
||||
* to the Fortran 77 Standard, but in accordance with Fortran 90).
|
||||
*/
|
||||
|
||||
#include "f2c.h"
|
||||
|
||||
int s_cat(char *lp, char **rpp, int* rnp, int *np)
|
||||
{
|
||||
int i, L = 0;
|
||||
int n = *np;
|
||||
|
||||
for(i = 0; i < n; i++) {
|
||||
int ni = rnp[i];
|
||||
if(ni > 0) {
|
||||
memcpy(lp + L, rpp[i], ni);
|
||||
L += ni;
|
||||
}
|
||||
}
|
||||
lp[L] = '\0';
|
||||
return 0;
|
||||
}
|
||||
Vendored
-40
@@ -1,40 +0,0 @@
|
||||
#include "f2c.h"
|
||||
|
||||
/* compare two strings */
|
||||
int s_cmp(char *a0, char *b0)
|
||||
{
|
||||
int la = (int)strlen(a0);
|
||||
int lb = (int)strlen(b0);
|
||||
unsigned char *a, *aend, *b, *bend;
|
||||
a = (unsigned char *)a0;
|
||||
b = (unsigned char *)b0;
|
||||
aend = a + la;
|
||||
bend = b + lb;
|
||||
|
||||
if(la <= lb)
|
||||
{
|
||||
while(a < aend)
|
||||
if(*a != *b)
|
||||
return( *a - *b );
|
||||
else
|
||||
{ ++a; ++b; }
|
||||
|
||||
while(b < bend)
|
||||
if(*b != ' ')
|
||||
return( ' ' - *b );
|
||||
else ++b;
|
||||
}
|
||||
else
|
||||
{
|
||||
while(b < bend)
|
||||
if(*a == *b)
|
||||
{ ++a; ++b; }
|
||||
else
|
||||
return( *a - *b );
|
||||
while(a < aend)
|
||||
if(*a != ' ')
|
||||
return(*a - ' ');
|
||||
else ++a;
|
||||
}
|
||||
return(0);
|
||||
}
|
||||
-71
@@ -1,71 +0,0 @@
|
||||
#include "f2c.h"
|
||||
#include <float.h>
|
||||
#include <stdio.h>
|
||||
|
||||
/* *********************************************************************** */
|
||||
|
||||
double slamc3_(float *a, float *b)
|
||||
{
|
||||
/* -- LAPACK auxiliary routine (version 3.1) -- */
|
||||
/* Univ. of Tennessee, Univ. of California Berkeley and NAG Ltd.. */
|
||||
/* November 2006 */
|
||||
|
||||
/* .. Scalar Arguments .. */
|
||||
/* .. */
|
||||
|
||||
/* Purpose */
|
||||
/* ======= */
|
||||
|
||||
/* SLAMC3 is intended to force A and B to be stored prior to doing */
|
||||
/* the addition of A and B , for use in situations where optimizers */
|
||||
/* might hold one of these in a register. */
|
||||
|
||||
/* Arguments */
|
||||
/* ========= */
|
||||
|
||||
/* A (input) REAL */
|
||||
/* B (input) REAL */
|
||||
/* The values A and B. */
|
||||
|
||||
/* ===================================================================== */
|
||||
|
||||
/* .. Executable Statements .. */
|
||||
|
||||
float ret_val = *a + *b;
|
||||
|
||||
return ret_val;
|
||||
|
||||
/* End of SLAMC3 */
|
||||
|
||||
} /* slamc3_ */
|
||||
|
||||
/* simpler version of slamch for the case of IEEE754-compliant FPU module by Piotr Luszczek S.
|
||||
taken from http://www.mail-archive.com/numpy-discussion@lists.sourceforge.net/msg02448.html */
|
||||
|
||||
#ifndef FLT_DIGITS
|
||||
#define FLT_DIGITS 24
|
||||
#endif
|
||||
|
||||
static const unsigned char lapack_slamch_tab0[] =
|
||||
{
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 2, 0, 0, 0, 0, 0, 0, 3, 4, 5, 6, 7, 0, 8, 9, 0, 10, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 2, 0, 0, 0, 0, 0, 0, 3, 4, 5, 6, 7, 0, 8, 9,
|
||||
0, 10, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
|
||||
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0
|
||||
};
|
||||
|
||||
const double lapack_slamch_tab1[] =
|
||||
{
|
||||
0, FLT_RADIX, FLT_EPSILON, FLT_MAX_EXP, FLT_MIN_EXP, FLT_DIGITS, FLT_MAX,
|
||||
FLT_EPSILON*FLT_RADIX, 1, FLT_MIN*(1 + FLT_EPSILON), FLT_MIN
|
||||
};
|
||||
|
||||
double slamch_(char* cmach)
|
||||
{
|
||||
return lapack_slamch_tab1[lapack_slamch_tab0[(unsigned char)cmach[0]]];
|
||||
}
|
||||
-19
@@ -1,19 +0,0 @@
|
||||
/* xerbla.f -- translated by f2c (version 20061008).
|
||||
You must link the resulting object file with libf2c:
|
||||
on Microsoft Windows system, link with libf2c.lib;
|
||||
on Linux or Unix systems, link with .../path/to/libf2c.a -lm
|
||||
or, if you install libf2c.a in a standard place, with -lf2c -lm
|
||||
-- in that order, at the end of the command line, as in
|
||||
cc *.o -lf2c -lm
|
||||
Source for libf2c is in /netlib/f2c/libf2c.zip, e.g.,
|
||||
|
||||
http://www.netlib.org/f2c/libf2c.zip
|
||||
*/
|
||||
|
||||
#include "f2c.h"
|
||||
|
||||
/* Subroutine */ int xerbla_(char *srname, int *info)
|
||||
{
|
||||
printf("** On entry to %s, parameter number %2i had an illegal value\n", srname, *info);
|
||||
return 0;
|
||||
} /* xerbla_ */
|
||||
Vendored
-752
@@ -1,752 +0,0 @@
|
||||
/* -- translated by f2c (version 20201020 (for_lapack)). -- */
|
||||
|
||||
#include "f2c.h"
|
||||
|
||||
//> \brief \b CGEMM
|
||||
//
|
||||
// =========== DOCUMENTATION ===========
|
||||
//
|
||||
// Online html documentation available at
|
||||
// http://www.netlib.org/lapack/explore-html/
|
||||
//
|
||||
// Definition:
|
||||
// ===========
|
||||
//
|
||||
// SUBROUTINE CGEMM(TRANSA,TRANSB,M,N,K,ALPHA,A,LDA,B,LDB,BETA,C,LDC)
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// COMPLEX ALPHA,BETA
|
||||
// INTEGER K,LDA,LDB,LDC,M,N
|
||||
// CHARACTER TRANSA,TRANSB
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// COMPLEX A(LDA,*),B(LDB,*),C(LDC,*)
|
||||
// ..
|
||||
//
|
||||
//
|
||||
//> \par Purpose:
|
||||
// =============
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> CGEMM performs one of the matrix-matrix operations
|
||||
//>
|
||||
//> C := alpha*op( A )*op( B ) + beta*C,
|
||||
//>
|
||||
//> where op( X ) is one of
|
||||
//>
|
||||
//> op( X ) = X or op( X ) = X**T or op( X ) = X**H,
|
||||
//>
|
||||
//> alpha and beta are scalars, and A, B and C are matrices, with op( A )
|
||||
//> an m by k matrix, op( B ) a k by n matrix and C an m by n matrix.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Arguments:
|
||||
// ==========
|
||||
//
|
||||
//> \param[in] TRANSA
|
||||
//> \verbatim
|
||||
//> TRANSA is CHARACTER*1
|
||||
//> On entry, TRANSA specifies the form of op( A ) to be used in
|
||||
//> the matrix multiplication as follows:
|
||||
//>
|
||||
//> TRANSA = 'N' or 'n', op( A ) = A.
|
||||
//>
|
||||
//> TRANSA = 'T' or 't', op( A ) = A**T.
|
||||
//>
|
||||
//> TRANSA = 'C' or 'c', op( A ) = A**H.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] TRANSB
|
||||
//> \verbatim
|
||||
//> TRANSB is CHARACTER*1
|
||||
//> On entry, TRANSB specifies the form of op( B ) to be used in
|
||||
//> the matrix multiplication as follows:
|
||||
//>
|
||||
//> TRANSB = 'N' or 'n', op( B ) = B.
|
||||
//>
|
||||
//> TRANSB = 'T' or 't', op( B ) = B**T.
|
||||
//>
|
||||
//> TRANSB = 'C' or 'c', op( B ) = B**H.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] M
|
||||
//> \verbatim
|
||||
//> M is INTEGER
|
||||
//> On entry, M specifies the number of rows of the matrix
|
||||
//> op( A ) and of the matrix C. M must be at least zero.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] N
|
||||
//> \verbatim
|
||||
//> N is INTEGER
|
||||
//> On entry, N specifies the number of columns of the matrix
|
||||
//> op( B ) and the number of columns of the matrix C. N must be
|
||||
//> at least zero.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] K
|
||||
//> \verbatim
|
||||
//> K is INTEGER
|
||||
//> On entry, K specifies the number of columns of the matrix
|
||||
//> op( A ) and the number of rows of the matrix op( B ). K must
|
||||
//> be at least zero.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] ALPHA
|
||||
//> \verbatim
|
||||
//> ALPHA is COMPLEX
|
||||
//> On entry, ALPHA specifies the scalar alpha.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] A
|
||||
//> \verbatim
|
||||
//> A is COMPLEX array, dimension ( LDA, ka ), where ka is
|
||||
//> k when TRANSA = 'N' or 'n', and is m otherwise.
|
||||
//> Before entry with TRANSA = 'N' or 'n', the leading m by k
|
||||
//> part of the array A must contain the matrix A, otherwise
|
||||
//> the leading k by m part of the array A must contain the
|
||||
//> matrix A.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] LDA
|
||||
//> \verbatim
|
||||
//> LDA is INTEGER
|
||||
//> On entry, LDA specifies the first dimension of A as declared
|
||||
//> in the calling (sub) program. When TRANSA = 'N' or 'n' then
|
||||
//> LDA must be at least max( 1, m ), otherwise LDA must be at
|
||||
//> least max( 1, k ).
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] B
|
||||
//> \verbatim
|
||||
//> B is COMPLEX array, dimension ( LDB, kb ), where kb is
|
||||
//> n when TRANSB = 'N' or 'n', and is k otherwise.
|
||||
//> Before entry with TRANSB = 'N' or 'n', the leading k by n
|
||||
//> part of the array B must contain the matrix B, otherwise
|
||||
//> the leading n by k part of the array B must contain the
|
||||
//> matrix B.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] LDB
|
||||
//> \verbatim
|
||||
//> LDB is INTEGER
|
||||
//> On entry, LDB specifies the first dimension of B as declared
|
||||
//> in the calling (sub) program. When TRANSB = 'N' or 'n' then
|
||||
//> LDB must be at least max( 1, k ), otherwise LDB must be at
|
||||
//> least max( 1, n ).
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] BETA
|
||||
//> \verbatim
|
||||
//> BETA is COMPLEX
|
||||
//> On entry, BETA specifies the scalar beta. When BETA is
|
||||
//> supplied as zero then C need not be set on input.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in,out] C
|
||||
//> \verbatim
|
||||
//> C is COMPLEX array, dimension ( LDC, N )
|
||||
//> Before entry, the leading m by n part of the array C must
|
||||
//> contain the matrix C, except when beta is zero, in which
|
||||
//> case C need not be set on entry.
|
||||
//> On exit, the array C is overwritten by the m by n matrix
|
||||
//> ( alpha*op( A )*op( B ) + beta*C ).
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] LDC
|
||||
//> \verbatim
|
||||
//> LDC is INTEGER
|
||||
//> On entry, LDC specifies the first dimension of C as declared
|
||||
//> in the calling (sub) program. LDC must be at least
|
||||
//> max( 1, m ).
|
||||
//> \endverbatim
|
||||
//
|
||||
// Authors:
|
||||
// ========
|
||||
//
|
||||
//> \author Univ. of Tennessee
|
||||
//> \author Univ. of California Berkeley
|
||||
//> \author Univ. of Colorado Denver
|
||||
//> \author NAG Ltd.
|
||||
//
|
||||
//> \date December 2016
|
||||
//
|
||||
//> \ingroup complex_blas_level3
|
||||
//
|
||||
//> \par Further Details:
|
||||
// =====================
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> Level 3 Blas routine.
|
||||
//>
|
||||
//> -- Written on 8-February-1989.
|
||||
//> Jack Dongarra, Argonne National Laboratory.
|
||||
//> Iain Duff, AERE Harwell.
|
||||
//> Jeremy Du Croz, Numerical Algorithms Group Ltd.
|
||||
//> Sven Hammarling, Numerical Algorithms Group Ltd.
|
||||
//> \endverbatim
|
||||
//>
|
||||
// =====================================================================
|
||||
/* Subroutine */ int cgemm_(char *transa, char *transb, int *m, int *n, int *
|
||||
k, complex *alpha, complex *a, int *lda, complex *b, int *ldb,
|
||||
complex *beta, complex *c__, int *ldc)
|
||||
{
|
||||
// Table of constant values
|
||||
complex c_b1 = {1.f,0.f};
|
||||
complex c_b2 = {0.f,0.f};
|
||||
|
||||
// System generated locals
|
||||
int a_dim1, a_offset, b_dim1, b_offset, c_dim1, c_offset, i__1, i__2,
|
||||
i__3, i__4, i__5, i__6;
|
||||
complex q__1, q__2, q__3, q__4;
|
||||
|
||||
// Local variables
|
||||
int i__, j, l, info;
|
||||
int nota, notb;
|
||||
complex temp;
|
||||
int conja, conjb;
|
||||
int ncola;
|
||||
extern int lsame_(char *, char *);
|
||||
int nrowa, nrowb;
|
||||
extern /* Subroutine */ int xerbla_(char *, int *);
|
||||
|
||||
//
|
||||
// -- Reference BLAS level3 routine (version 3.7.0) --
|
||||
// -- Reference BLAS is a software package provided by Univ. of Tennessee, --
|
||||
// -- Univ. of California Berkeley, Univ. of Colorado Denver and NAG Ltd..--
|
||||
// December 2016
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// ..
|
||||
//
|
||||
// =====================================================================
|
||||
//
|
||||
// .. External Functions ..
|
||||
// ..
|
||||
// .. External Subroutines ..
|
||||
// ..
|
||||
// .. Intrinsic Functions ..
|
||||
// ..
|
||||
// .. Local Scalars ..
|
||||
// ..
|
||||
// .. Parameters ..
|
||||
// ..
|
||||
//
|
||||
// Set NOTA and NOTB as true if A and B respectively are not
|
||||
// conjugated or transposed, set CONJA and CONJB as true if A and
|
||||
// B respectively are to be transposed but not conjugated and set
|
||||
// NROWA, NCOLA and NROWB as the number of rows and columns of A
|
||||
// and the number of rows of B respectively.
|
||||
//
|
||||
// Parameter adjustments
|
||||
a_dim1 = *lda;
|
||||
a_offset = 1 + a_dim1;
|
||||
a -= a_offset;
|
||||
b_dim1 = *ldb;
|
||||
b_offset = 1 + b_dim1;
|
||||
b -= b_offset;
|
||||
c_dim1 = *ldc;
|
||||
c_offset = 1 + c_dim1;
|
||||
c__ -= c_offset;
|
||||
|
||||
// Function Body
|
||||
nota = lsame_(transa, "N");
|
||||
notb = lsame_(transb, "N");
|
||||
conja = lsame_(transa, "C");
|
||||
conjb = lsame_(transb, "C");
|
||||
if (nota) {
|
||||
nrowa = *m;
|
||||
ncola = *k;
|
||||
} else {
|
||||
nrowa = *k;
|
||||
ncola = *m;
|
||||
}
|
||||
if (notb) {
|
||||
nrowb = *k;
|
||||
} else {
|
||||
nrowb = *n;
|
||||
}
|
||||
//
|
||||
// Test the input parameters.
|
||||
//
|
||||
info = 0;
|
||||
if (! nota && ! conja && ! lsame_(transa, "T")) {
|
||||
info = 1;
|
||||
} else if (! notb && ! conjb && ! lsame_(transb, "T")) {
|
||||
info = 2;
|
||||
} else if (*m < 0) {
|
||||
info = 3;
|
||||
} else if (*n < 0) {
|
||||
info = 4;
|
||||
} else if (*k < 0) {
|
||||
info = 5;
|
||||
} else if (*lda < max(1,nrowa)) {
|
||||
info = 8;
|
||||
} else if (*ldb < max(1,nrowb)) {
|
||||
info = 10;
|
||||
} else if (*ldc < max(1,*m)) {
|
||||
info = 13;
|
||||
}
|
||||
if (info != 0) {
|
||||
xerbla_("CGEMM ", &info);
|
||||
return 0;
|
||||
}
|
||||
//
|
||||
// Quick return if possible.
|
||||
//
|
||||
if (*m == 0 || *n == 0 || (alpha->r == 0.f && alpha->i == 0.f || *k == 0)
|
||||
&& (beta->r == 1.f && beta->i == 0.f)) {
|
||||
return 0;
|
||||
}
|
||||
//
|
||||
// And when alpha.eq.zero.
|
||||
//
|
||||
if (alpha->r == 0.f && alpha->i == 0.f) {
|
||||
if (beta->r == 0.f && beta->i == 0.f) {
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
c__[i__3].r = 0.f, c__[i__3].i = 0.f;
|
||||
// L10:
|
||||
}
|
||||
// L20:
|
||||
}
|
||||
} else {
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
i__4 = i__ + j * c_dim1;
|
||||
q__1.r = beta->r * c__[i__4].r - beta->i * c__[i__4].i,
|
||||
q__1.i = beta->r * c__[i__4].i + beta->i * c__[
|
||||
i__4].r;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
// L30:
|
||||
}
|
||||
// L40:
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
//
|
||||
// Start the operations.
|
||||
//
|
||||
if (notb) {
|
||||
if (nota) {
|
||||
//
|
||||
// Form C := alpha*A*B + beta*C.
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
if (beta->r == 0.f && beta->i == 0.f) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
c__[i__3].r = 0.f, c__[i__3].i = 0.f;
|
||||
// L50:
|
||||
}
|
||||
} else if (beta->r != 1.f || beta->i != 0.f) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
i__4 = i__ + j * c_dim1;
|
||||
q__1.r = beta->r * c__[i__4].r - beta->i * c__[i__4]
|
||||
.i, q__1.i = beta->r * c__[i__4].i + beta->i *
|
||||
c__[i__4].r;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
// L60:
|
||||
}
|
||||
}
|
||||
i__2 = *k;
|
||||
for (l = 1; l <= i__2; ++l) {
|
||||
i__3 = l + j * b_dim1;
|
||||
q__1.r = alpha->r * b[i__3].r - alpha->i * b[i__3].i,
|
||||
q__1.i = alpha->r * b[i__3].i + alpha->i * b[i__3]
|
||||
.r;
|
||||
temp.r = q__1.r, temp.i = q__1.i;
|
||||
i__3 = *m;
|
||||
for (i__ = 1; i__ <= i__3; ++i__) {
|
||||
i__4 = i__ + j * c_dim1;
|
||||
i__5 = i__ + j * c_dim1;
|
||||
i__6 = i__ + l * a_dim1;
|
||||
q__2.r = temp.r * a[i__6].r - temp.i * a[i__6].i,
|
||||
q__2.i = temp.r * a[i__6].i + temp.i * a[i__6]
|
||||
.r;
|
||||
q__1.r = c__[i__5].r + q__2.r, q__1.i = c__[i__5].i +
|
||||
q__2.i;
|
||||
c__[i__4].r = q__1.r, c__[i__4].i = q__1.i;
|
||||
// L70:
|
||||
}
|
||||
// L80:
|
||||
}
|
||||
// L90:
|
||||
}
|
||||
} else if (conja) {
|
||||
//
|
||||
// Form C := alpha*A**H*B + beta*C.
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
temp.r = 0.f, temp.i = 0.f;
|
||||
i__3 = *k;
|
||||
for (l = 1; l <= i__3; ++l) {
|
||||
r_cnjg(&q__3, &a[l + i__ * a_dim1]);
|
||||
i__4 = l + j * b_dim1;
|
||||
q__2.r = q__3.r * b[i__4].r - q__3.i * b[i__4].i,
|
||||
q__2.i = q__3.r * b[i__4].i + q__3.i * b[i__4]
|
||||
.r;
|
||||
q__1.r = temp.r + q__2.r, q__1.i = temp.i + q__2.i;
|
||||
temp.r = q__1.r, temp.i = q__1.i;
|
||||
// L100:
|
||||
}
|
||||
if (beta->r == 0.f && beta->i == 0.f) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__1.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__1.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
} else {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__2.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__2.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
i__4 = i__ + j * c_dim1;
|
||||
q__3.r = beta->r * c__[i__4].r - beta->i * c__[i__4]
|
||||
.i, q__3.i = beta->r * c__[i__4].i + beta->i *
|
||||
c__[i__4].r;
|
||||
q__1.r = q__2.r + q__3.r, q__1.i = q__2.i + q__3.i;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
}
|
||||
// L110:
|
||||
}
|
||||
// L120:
|
||||
}
|
||||
} else {
|
||||
//
|
||||
// Form C := alpha*A**T*B + beta*C
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
temp.r = 0.f, temp.i = 0.f;
|
||||
i__3 = *k;
|
||||
for (l = 1; l <= i__3; ++l) {
|
||||
i__4 = l + i__ * a_dim1;
|
||||
i__5 = l + j * b_dim1;
|
||||
q__2.r = a[i__4].r * b[i__5].r - a[i__4].i * b[i__5]
|
||||
.i, q__2.i = a[i__4].r * b[i__5].i + a[i__4]
|
||||
.i * b[i__5].r;
|
||||
q__1.r = temp.r + q__2.r, q__1.i = temp.i + q__2.i;
|
||||
temp.r = q__1.r, temp.i = q__1.i;
|
||||
// L130:
|
||||
}
|
||||
if (beta->r == 0.f && beta->i == 0.f) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__1.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__1.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
} else {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__2.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__2.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
i__4 = i__ + j * c_dim1;
|
||||
q__3.r = beta->r * c__[i__4].r - beta->i * c__[i__4]
|
||||
.i, q__3.i = beta->r * c__[i__4].i + beta->i *
|
||||
c__[i__4].r;
|
||||
q__1.r = q__2.r + q__3.r, q__1.i = q__2.i + q__3.i;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
}
|
||||
// L140:
|
||||
}
|
||||
// L150:
|
||||
}
|
||||
}
|
||||
} else if (nota) {
|
||||
if (conjb) {
|
||||
//
|
||||
// Form C := alpha*A*B**H + beta*C.
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
if (beta->r == 0.f && beta->i == 0.f) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
c__[i__3].r = 0.f, c__[i__3].i = 0.f;
|
||||
// L160:
|
||||
}
|
||||
} else if (beta->r != 1.f || beta->i != 0.f) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
i__4 = i__ + j * c_dim1;
|
||||
q__1.r = beta->r * c__[i__4].r - beta->i * c__[i__4]
|
||||
.i, q__1.i = beta->r * c__[i__4].i + beta->i *
|
||||
c__[i__4].r;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
// L170:
|
||||
}
|
||||
}
|
||||
i__2 = *k;
|
||||
for (l = 1; l <= i__2; ++l) {
|
||||
r_cnjg(&q__2, &b[j + l * b_dim1]);
|
||||
q__1.r = alpha->r * q__2.r - alpha->i * q__2.i, q__1.i =
|
||||
alpha->r * q__2.i + alpha->i * q__2.r;
|
||||
temp.r = q__1.r, temp.i = q__1.i;
|
||||
i__3 = *m;
|
||||
for (i__ = 1; i__ <= i__3; ++i__) {
|
||||
i__4 = i__ + j * c_dim1;
|
||||
i__5 = i__ + j * c_dim1;
|
||||
i__6 = i__ + l * a_dim1;
|
||||
q__2.r = temp.r * a[i__6].r - temp.i * a[i__6].i,
|
||||
q__2.i = temp.r * a[i__6].i + temp.i * a[i__6]
|
||||
.r;
|
||||
q__1.r = c__[i__5].r + q__2.r, q__1.i = c__[i__5].i +
|
||||
q__2.i;
|
||||
c__[i__4].r = q__1.r, c__[i__4].i = q__1.i;
|
||||
// L180:
|
||||
}
|
||||
// L190:
|
||||
}
|
||||
// L200:
|
||||
}
|
||||
} else {
|
||||
//
|
||||
// Form C := alpha*A*B**T + beta*C
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
if (beta->r == 0.f && beta->i == 0.f) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
c__[i__3].r = 0.f, c__[i__3].i = 0.f;
|
||||
// L210:
|
||||
}
|
||||
} else if (beta->r != 1.f || beta->i != 0.f) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
i__4 = i__ + j * c_dim1;
|
||||
q__1.r = beta->r * c__[i__4].r - beta->i * c__[i__4]
|
||||
.i, q__1.i = beta->r * c__[i__4].i + beta->i *
|
||||
c__[i__4].r;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
// L220:
|
||||
}
|
||||
}
|
||||
i__2 = *k;
|
||||
for (l = 1; l <= i__2; ++l) {
|
||||
i__3 = j + l * b_dim1;
|
||||
q__1.r = alpha->r * b[i__3].r - alpha->i * b[i__3].i,
|
||||
q__1.i = alpha->r * b[i__3].i + alpha->i * b[i__3]
|
||||
.r;
|
||||
temp.r = q__1.r, temp.i = q__1.i;
|
||||
i__3 = *m;
|
||||
for (i__ = 1; i__ <= i__3; ++i__) {
|
||||
i__4 = i__ + j * c_dim1;
|
||||
i__5 = i__ + j * c_dim1;
|
||||
i__6 = i__ + l * a_dim1;
|
||||
q__2.r = temp.r * a[i__6].r - temp.i * a[i__6].i,
|
||||
q__2.i = temp.r * a[i__6].i + temp.i * a[i__6]
|
||||
.r;
|
||||
q__1.r = c__[i__5].r + q__2.r, q__1.i = c__[i__5].i +
|
||||
q__2.i;
|
||||
c__[i__4].r = q__1.r, c__[i__4].i = q__1.i;
|
||||
// L230:
|
||||
}
|
||||
// L240:
|
||||
}
|
||||
// L250:
|
||||
}
|
||||
}
|
||||
} else if (conja) {
|
||||
if (conjb) {
|
||||
//
|
||||
// Form C := alpha*A**H*B**H + beta*C.
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
temp.r = 0.f, temp.i = 0.f;
|
||||
i__3 = *k;
|
||||
for (l = 1; l <= i__3; ++l) {
|
||||
r_cnjg(&q__3, &a[l + i__ * a_dim1]);
|
||||
r_cnjg(&q__4, &b[j + l * b_dim1]);
|
||||
q__2.r = q__3.r * q__4.r - q__3.i * q__4.i, q__2.i =
|
||||
q__3.r * q__4.i + q__3.i * q__4.r;
|
||||
q__1.r = temp.r + q__2.r, q__1.i = temp.i + q__2.i;
|
||||
temp.r = q__1.r, temp.i = q__1.i;
|
||||
// L260:
|
||||
}
|
||||
if (beta->r == 0.f && beta->i == 0.f) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__1.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__1.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
} else {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__2.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__2.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
i__4 = i__ + j * c_dim1;
|
||||
q__3.r = beta->r * c__[i__4].r - beta->i * c__[i__4]
|
||||
.i, q__3.i = beta->r * c__[i__4].i + beta->i *
|
||||
c__[i__4].r;
|
||||
q__1.r = q__2.r + q__3.r, q__1.i = q__2.i + q__3.i;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
}
|
||||
// L270:
|
||||
}
|
||||
// L280:
|
||||
}
|
||||
} else {
|
||||
//
|
||||
// Form C := alpha*A**H*B**T + beta*C
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
temp.r = 0.f, temp.i = 0.f;
|
||||
i__3 = *k;
|
||||
for (l = 1; l <= i__3; ++l) {
|
||||
r_cnjg(&q__3, &a[l + i__ * a_dim1]);
|
||||
i__4 = j + l * b_dim1;
|
||||
q__2.r = q__3.r * b[i__4].r - q__3.i * b[i__4].i,
|
||||
q__2.i = q__3.r * b[i__4].i + q__3.i * b[i__4]
|
||||
.r;
|
||||
q__1.r = temp.r + q__2.r, q__1.i = temp.i + q__2.i;
|
||||
temp.r = q__1.r, temp.i = q__1.i;
|
||||
// L290:
|
||||
}
|
||||
if (beta->r == 0.f && beta->i == 0.f) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__1.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__1.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
} else {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__2.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__2.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
i__4 = i__ + j * c_dim1;
|
||||
q__3.r = beta->r * c__[i__4].r - beta->i * c__[i__4]
|
||||
.i, q__3.i = beta->r * c__[i__4].i + beta->i *
|
||||
c__[i__4].r;
|
||||
q__1.r = q__2.r + q__3.r, q__1.i = q__2.i + q__3.i;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
}
|
||||
// L300:
|
||||
}
|
||||
// L310:
|
||||
}
|
||||
}
|
||||
} else {
|
||||
if (conjb) {
|
||||
//
|
||||
// Form C := alpha*A**T*B**H + beta*C
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
temp.r = 0.f, temp.i = 0.f;
|
||||
i__3 = *k;
|
||||
for (l = 1; l <= i__3; ++l) {
|
||||
i__4 = l + i__ * a_dim1;
|
||||
r_cnjg(&q__3, &b[j + l * b_dim1]);
|
||||
q__2.r = a[i__4].r * q__3.r - a[i__4].i * q__3.i,
|
||||
q__2.i = a[i__4].r * q__3.i + a[i__4].i *
|
||||
q__3.r;
|
||||
q__1.r = temp.r + q__2.r, q__1.i = temp.i + q__2.i;
|
||||
temp.r = q__1.r, temp.i = q__1.i;
|
||||
// L320:
|
||||
}
|
||||
if (beta->r == 0.f && beta->i == 0.f) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__1.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__1.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
} else {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__2.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__2.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
i__4 = i__ + j * c_dim1;
|
||||
q__3.r = beta->r * c__[i__4].r - beta->i * c__[i__4]
|
||||
.i, q__3.i = beta->r * c__[i__4].i + beta->i *
|
||||
c__[i__4].r;
|
||||
q__1.r = q__2.r + q__3.r, q__1.i = q__2.i + q__3.i;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
}
|
||||
// L330:
|
||||
}
|
||||
// L340:
|
||||
}
|
||||
} else {
|
||||
//
|
||||
// Form C := alpha*A**T*B**T + beta*C
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
temp.r = 0.f, temp.i = 0.f;
|
||||
i__3 = *k;
|
||||
for (l = 1; l <= i__3; ++l) {
|
||||
i__4 = l + i__ * a_dim1;
|
||||
i__5 = j + l * b_dim1;
|
||||
q__2.r = a[i__4].r * b[i__5].r - a[i__4].i * b[i__5]
|
||||
.i, q__2.i = a[i__4].r * b[i__5].i + a[i__4]
|
||||
.i * b[i__5].r;
|
||||
q__1.r = temp.r + q__2.r, q__1.i = temp.i + q__2.i;
|
||||
temp.r = q__1.r, temp.i = q__1.i;
|
||||
// L350:
|
||||
}
|
||||
if (beta->r == 0.f && beta->i == 0.f) {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__1.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__1.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
} else {
|
||||
i__3 = i__ + j * c_dim1;
|
||||
q__2.r = alpha->r * temp.r - alpha->i * temp.i,
|
||||
q__2.i = alpha->r * temp.i + alpha->i *
|
||||
temp.r;
|
||||
i__4 = i__ + j * c_dim1;
|
||||
q__3.r = beta->r * c__[i__4].r - beta->i * c__[i__4]
|
||||
.i, q__3.i = beta->r * c__[i__4].i + beta->i *
|
||||
c__[i__4].r;
|
||||
q__1.r = q__2.r + q__3.r, q__1.i = q__2.i + q__3.i;
|
||||
c__[i__3].r = q__1.r, c__[i__3].i = q__1.i;
|
||||
}
|
||||
// L360:
|
||||
}
|
||||
// L370:
|
||||
}
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
//
|
||||
// End of CGEMM .
|
||||
//
|
||||
} // cgemm_
|
||||
|
||||
Vendored
-171
@@ -1,171 +0,0 @@
|
||||
/* -- translated by f2c (version 20201020 (for_lapack)). -- */
|
||||
|
||||
#include "f2c.h"
|
||||
|
||||
//> \brief \b DCOPY
|
||||
//
|
||||
// =========== DOCUMENTATION ===========
|
||||
//
|
||||
// Online html documentation available at
|
||||
// http://www.netlib.org/lapack/explore-html/
|
||||
//
|
||||
// Definition:
|
||||
// ===========
|
||||
//
|
||||
// SUBROUTINE DCOPY(N,DX,INCX,DY,INCY)
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// INTEGER INCX,INCY,N
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// DOUBLE PRECISION DX(*),DY(*)
|
||||
// ..
|
||||
//
|
||||
//
|
||||
//> \par Purpose:
|
||||
// =============
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> DCOPY copies a vector, x, to a vector, y.
|
||||
//> uses unrolled loops for increments equal to 1.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Arguments:
|
||||
// ==========
|
||||
//
|
||||
//> \param[in] N
|
||||
//> \verbatim
|
||||
//> N is INTEGER
|
||||
//> number of elements in input vector(s)
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] DX
|
||||
//> \verbatim
|
||||
//> DX is DOUBLE PRECISION array, dimension ( 1 + ( N - 1 )*abs( INCX ) )
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] INCX
|
||||
//> \verbatim
|
||||
//> INCX is INTEGER
|
||||
//> storage spacing between elements of DX
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[out] DY
|
||||
//> \verbatim
|
||||
//> DY is DOUBLE PRECISION array, dimension ( 1 + ( N - 1 )*abs( INCY ) )
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] INCY
|
||||
//> \verbatim
|
||||
//> INCY is INTEGER
|
||||
//> storage spacing between elements of DY
|
||||
//> \endverbatim
|
||||
//
|
||||
// Authors:
|
||||
// ========
|
||||
//
|
||||
//> \author Univ. of Tennessee
|
||||
//> \author Univ. of California Berkeley
|
||||
//> \author Univ. of Colorado Denver
|
||||
//> \author NAG Ltd.
|
||||
//
|
||||
//> \date November 2017
|
||||
//
|
||||
//> \ingroup double_blas_level1
|
||||
//
|
||||
//> \par Further Details:
|
||||
// =====================
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> jack dongarra, linpack, 3/11/78.
|
||||
//> modified 12/3/93, array(1) declarations changed to array(*)
|
||||
//> \endverbatim
|
||||
//>
|
||||
// =====================================================================
|
||||
/* Subroutine */ int dcopy_(int *n, double *dx, int *incx, double *dy, int *
|
||||
incy)
|
||||
{
|
||||
// System generated locals
|
||||
int i__1;
|
||||
|
||||
// Local variables
|
||||
int i__, m, ix, iy, mp1;
|
||||
|
||||
//
|
||||
// -- Reference BLAS level1 routine (version 3.8.0) --
|
||||
// -- Reference BLAS is a software package provided by Univ. of Tennessee, --
|
||||
// -- Univ. of California Berkeley, Univ. of Colorado Denver and NAG Ltd..--
|
||||
// November 2017
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// ..
|
||||
//
|
||||
// =====================================================================
|
||||
//
|
||||
// .. Local Scalars ..
|
||||
// ..
|
||||
// .. Intrinsic Functions ..
|
||||
// ..
|
||||
// Parameter adjustments
|
||||
--dy;
|
||||
--dx;
|
||||
|
||||
// Function Body
|
||||
if (*n <= 0) {
|
||||
return 0;
|
||||
}
|
||||
if (*incx == 1 && *incy == 1) {
|
||||
//
|
||||
// code for both increments equal to 1
|
||||
//
|
||||
//
|
||||
// clean-up loop
|
||||
//
|
||||
m = *n % 7;
|
||||
if (m != 0) {
|
||||
i__1 = m;
|
||||
for (i__ = 1; i__ <= i__1; ++i__) {
|
||||
dy[i__] = dx[i__];
|
||||
}
|
||||
if (*n < 7) {
|
||||
return 0;
|
||||
}
|
||||
}
|
||||
mp1 = m + 1;
|
||||
i__1 = *n;
|
||||
for (i__ = mp1; i__ <= i__1; i__ += 7) {
|
||||
dy[i__] = dx[i__];
|
||||
dy[i__ + 1] = dx[i__ + 1];
|
||||
dy[i__ + 2] = dx[i__ + 2];
|
||||
dy[i__ + 3] = dx[i__ + 3];
|
||||
dy[i__ + 4] = dx[i__ + 4];
|
||||
dy[i__ + 5] = dx[i__ + 5];
|
||||
dy[i__ + 6] = dx[i__ + 6];
|
||||
}
|
||||
} else {
|
||||
//
|
||||
// code for unequal increments or equal increments
|
||||
// not equal to 1
|
||||
//
|
||||
ix = 1;
|
||||
iy = 1;
|
||||
if (*incx < 0) {
|
||||
ix = (-(*n) + 1) * *incx + 1;
|
||||
}
|
||||
if (*incy < 0) {
|
||||
iy = (-(*n) + 1) * *incy + 1;
|
||||
}
|
||||
i__1 = *n;
|
||||
for (i__ = 1; i__ <= i__1; ++i__) {
|
||||
dy[iy] = dx[ix];
|
||||
ix += *incx;
|
||||
iy += *incy;
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
} // dcopy_
|
||||
|
||||
Vendored
-172
@@ -1,172 +0,0 @@
|
||||
/* -- translated by f2c (version 20201020 (for_lapack)). -- */
|
||||
|
||||
#include "f2c.h"
|
||||
|
||||
//> \brief \b DDOT
|
||||
//
|
||||
// =========== DOCUMENTATION ===========
|
||||
//
|
||||
// Online html documentation available at
|
||||
// http://www.netlib.org/lapack/explore-html/
|
||||
//
|
||||
// Definition:
|
||||
// ===========
|
||||
//
|
||||
// DOUBLE PRECISION FUNCTION DDOT(N,DX,INCX,DY,INCY)
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// INTEGER INCX,INCY,N
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// DOUBLE PRECISION DX(*),DY(*)
|
||||
// ..
|
||||
//
|
||||
//
|
||||
//> \par Purpose:
|
||||
// =============
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> DDOT forms the dot product of two vectors.
|
||||
//> uses unrolled loops for increments equal to one.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Arguments:
|
||||
// ==========
|
||||
//
|
||||
//> \param[in] N
|
||||
//> \verbatim
|
||||
//> N is INTEGER
|
||||
//> number of elements in input vector(s)
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] DX
|
||||
//> \verbatim
|
||||
//> DX is DOUBLE PRECISION array, dimension ( 1 + ( N - 1 )*abs( INCX ) )
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] INCX
|
||||
//> \verbatim
|
||||
//> INCX is INTEGER
|
||||
//> storage spacing between elements of DX
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] DY
|
||||
//> \verbatim
|
||||
//> DY is DOUBLE PRECISION array, dimension ( 1 + ( N - 1 )*abs( INCY ) )
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] INCY
|
||||
//> \verbatim
|
||||
//> INCY is INTEGER
|
||||
//> storage spacing between elements of DY
|
||||
//> \endverbatim
|
||||
//
|
||||
// Authors:
|
||||
// ========
|
||||
//
|
||||
//> \author Univ. of Tennessee
|
||||
//> \author Univ. of California Berkeley
|
||||
//> \author Univ. of Colorado Denver
|
||||
//> \author NAG Ltd.
|
||||
//
|
||||
//> \date November 2017
|
||||
//
|
||||
//> \ingroup double_blas_level1
|
||||
//
|
||||
//> \par Further Details:
|
||||
// =====================
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> jack dongarra, linpack, 3/11/78.
|
||||
//> modified 12/3/93, array(1) declarations changed to array(*)
|
||||
//> \endverbatim
|
||||
//>
|
||||
// =====================================================================
|
||||
double ddot_(int *n, double *dx, int *incx, double *dy, int *incy)
|
||||
{
|
||||
// System generated locals
|
||||
int i__1;
|
||||
double ret_val;
|
||||
|
||||
// Local variables
|
||||
int i__, m, ix, iy, mp1;
|
||||
double dtemp;
|
||||
|
||||
//
|
||||
// -- Reference BLAS level1 routine (version 3.8.0) --
|
||||
// -- Reference BLAS is a software package provided by Univ. of Tennessee, --
|
||||
// -- Univ. of California Berkeley, Univ. of Colorado Denver and NAG Ltd..--
|
||||
// November 2017
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// ..
|
||||
//
|
||||
// =====================================================================
|
||||
//
|
||||
// .. Local Scalars ..
|
||||
// ..
|
||||
// .. Intrinsic Functions ..
|
||||
// ..
|
||||
// Parameter adjustments
|
||||
--dy;
|
||||
--dx;
|
||||
|
||||
// Function Body
|
||||
ret_val = 0.;
|
||||
dtemp = 0.;
|
||||
if (*n <= 0) {
|
||||
return ret_val;
|
||||
}
|
||||
if (*incx == 1 && *incy == 1) {
|
||||
//
|
||||
// code for both increments equal to 1
|
||||
//
|
||||
//
|
||||
// clean-up loop
|
||||
//
|
||||
m = *n % 5;
|
||||
if (m != 0) {
|
||||
i__1 = m;
|
||||
for (i__ = 1; i__ <= i__1; ++i__) {
|
||||
dtemp += dx[i__] * dy[i__];
|
||||
}
|
||||
if (*n < 5) {
|
||||
ret_val = dtemp;
|
||||
return ret_val;
|
||||
}
|
||||
}
|
||||
mp1 = m + 1;
|
||||
i__1 = *n;
|
||||
for (i__ = mp1; i__ <= i__1; i__ += 5) {
|
||||
dtemp = dtemp + dx[i__] * dy[i__] + dx[i__ + 1] * dy[i__ + 1] +
|
||||
dx[i__ + 2] * dy[i__ + 2] + dx[i__ + 3] * dy[i__ + 3] +
|
||||
dx[i__ + 4] * dy[i__ + 4];
|
||||
}
|
||||
} else {
|
||||
//
|
||||
// code for unequal increments or equal increments
|
||||
// not equal to 1
|
||||
//
|
||||
ix = 1;
|
||||
iy = 1;
|
||||
if (*incx < 0) {
|
||||
ix = (-(*n) + 1) * *incx + 1;
|
||||
}
|
||||
if (*incy < 0) {
|
||||
iy = (-(*n) + 1) * *incy + 1;
|
||||
}
|
||||
i__1 = *n;
|
||||
for (i__ = 1; i__ <= i__1; ++i__) {
|
||||
dtemp += dx[ix] * dy[iy];
|
||||
ix += *incx;
|
||||
iy += *incy;
|
||||
}
|
||||
}
|
||||
ret_val = dtemp;
|
||||
return ret_val;
|
||||
} // ddot_
|
||||
|
||||
Vendored
-14369
File diff suppressed because it is too large
Load Diff
Vendored
-444
@@ -1,444 +0,0 @@
|
||||
/* -- translated by f2c (version 20201020 (for_lapack)). -- */
|
||||
|
||||
#include "f2c.h"
|
||||
|
||||
//> \brief \b DGEMM
|
||||
//
|
||||
// =========== DOCUMENTATION ===========
|
||||
//
|
||||
// Online html documentation available at
|
||||
// http://www.netlib.org/lapack/explore-html/
|
||||
//
|
||||
// Definition:
|
||||
// ===========
|
||||
//
|
||||
// SUBROUTINE DGEMM(TRANSA,TRANSB,M,N,K,ALPHA,A,LDA,B,LDB,BETA,C,LDC)
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// DOUBLE PRECISION ALPHA,BETA
|
||||
// INTEGER K,LDA,LDB,LDC,M,N
|
||||
// CHARACTER TRANSA,TRANSB
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// DOUBLE PRECISION A(LDA,*),B(LDB,*),C(LDC,*)
|
||||
// ..
|
||||
//
|
||||
//
|
||||
//> \par Purpose:
|
||||
// =============
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> DGEMM performs one of the matrix-matrix operations
|
||||
//>
|
||||
//> C := alpha*op( A )*op( B ) + beta*C,
|
||||
//>
|
||||
//> where op( X ) is one of
|
||||
//>
|
||||
//> op( X ) = X or op( X ) = X**T,
|
||||
//>
|
||||
//> alpha and beta are scalars, and A, B and C are matrices, with op( A )
|
||||
//> an m by k matrix, op( B ) a k by n matrix and C an m by n matrix.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Arguments:
|
||||
// ==========
|
||||
//
|
||||
//> \param[in] TRANSA
|
||||
//> \verbatim
|
||||
//> TRANSA is CHARACTER*1
|
||||
//> On entry, TRANSA specifies the form of op( A ) to be used in
|
||||
//> the matrix multiplication as follows:
|
||||
//>
|
||||
//> TRANSA = 'N' or 'n', op( A ) = A.
|
||||
//>
|
||||
//> TRANSA = 'T' or 't', op( A ) = A**T.
|
||||
//>
|
||||
//> TRANSA = 'C' or 'c', op( A ) = A**T.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] TRANSB
|
||||
//> \verbatim
|
||||
//> TRANSB is CHARACTER*1
|
||||
//> On entry, TRANSB specifies the form of op( B ) to be used in
|
||||
//> the matrix multiplication as follows:
|
||||
//>
|
||||
//> TRANSB = 'N' or 'n', op( B ) = B.
|
||||
//>
|
||||
//> TRANSB = 'T' or 't', op( B ) = B**T.
|
||||
//>
|
||||
//> TRANSB = 'C' or 'c', op( B ) = B**T.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] M
|
||||
//> \verbatim
|
||||
//> M is INTEGER
|
||||
//> On entry, M specifies the number of rows of the matrix
|
||||
//> op( A ) and of the matrix C. M must be at least zero.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] N
|
||||
//> \verbatim
|
||||
//> N is INTEGER
|
||||
//> On entry, N specifies the number of columns of the matrix
|
||||
//> op( B ) and the number of columns of the matrix C. N must be
|
||||
//> at least zero.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] K
|
||||
//> \verbatim
|
||||
//> K is INTEGER
|
||||
//> On entry, K specifies the number of columns of the matrix
|
||||
//> op( A ) and the number of rows of the matrix op( B ). K must
|
||||
//> be at least zero.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] ALPHA
|
||||
//> \verbatim
|
||||
//> ALPHA is DOUBLE PRECISION.
|
||||
//> On entry, ALPHA specifies the scalar alpha.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] A
|
||||
//> \verbatim
|
||||
//> A is DOUBLE PRECISION array, dimension ( LDA, ka ), where ka is
|
||||
//> k when TRANSA = 'N' or 'n', and is m otherwise.
|
||||
//> Before entry with TRANSA = 'N' or 'n', the leading m by k
|
||||
//> part of the array A must contain the matrix A, otherwise
|
||||
//> the leading k by m part of the array A must contain the
|
||||
//> matrix A.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] LDA
|
||||
//> \verbatim
|
||||
//> LDA is INTEGER
|
||||
//> On entry, LDA specifies the first dimension of A as declared
|
||||
//> in the calling (sub) program. When TRANSA = 'N' or 'n' then
|
||||
//> LDA must be at least max( 1, m ), otherwise LDA must be at
|
||||
//> least max( 1, k ).
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] B
|
||||
//> \verbatim
|
||||
//> B is DOUBLE PRECISION array, dimension ( LDB, kb ), where kb is
|
||||
//> n when TRANSB = 'N' or 'n', and is k otherwise.
|
||||
//> Before entry with TRANSB = 'N' or 'n', the leading k by n
|
||||
//> part of the array B must contain the matrix B, otherwise
|
||||
//> the leading n by k part of the array B must contain the
|
||||
//> matrix B.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] LDB
|
||||
//> \verbatim
|
||||
//> LDB is INTEGER
|
||||
//> On entry, LDB specifies the first dimension of B as declared
|
||||
//> in the calling (sub) program. When TRANSB = 'N' or 'n' then
|
||||
//> LDB must be at least max( 1, k ), otherwise LDB must be at
|
||||
//> least max( 1, n ).
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] BETA
|
||||
//> \verbatim
|
||||
//> BETA is DOUBLE PRECISION.
|
||||
//> On entry, BETA specifies the scalar beta. When BETA is
|
||||
//> supplied as zero then C need not be set on input.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in,out] C
|
||||
//> \verbatim
|
||||
//> C is DOUBLE PRECISION array, dimension ( LDC, N )
|
||||
//> Before entry, the leading m by n part of the array C must
|
||||
//> contain the matrix C, except when beta is zero, in which
|
||||
//> case C need not be set on entry.
|
||||
//> On exit, the array C is overwritten by the m by n matrix
|
||||
//> ( alpha*op( A )*op( B ) + beta*C ).
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] LDC
|
||||
//> \verbatim
|
||||
//> LDC is INTEGER
|
||||
//> On entry, LDC specifies the first dimension of C as declared
|
||||
//> in the calling (sub) program. LDC must be at least
|
||||
//> max( 1, m ).
|
||||
//> \endverbatim
|
||||
//
|
||||
// Authors:
|
||||
// ========
|
||||
//
|
||||
//> \author Univ. of Tennessee
|
||||
//> \author Univ. of California Berkeley
|
||||
//> \author Univ. of Colorado Denver
|
||||
//> \author NAG Ltd.
|
||||
//
|
||||
//> \date December 2016
|
||||
//
|
||||
//> \ingroup double_blas_level3
|
||||
//
|
||||
//> \par Further Details:
|
||||
// =====================
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> Level 3 Blas routine.
|
||||
//>
|
||||
//> -- Written on 8-February-1989.
|
||||
//> Jack Dongarra, Argonne National Laboratory.
|
||||
//> Iain Duff, AERE Harwell.
|
||||
//> Jeremy Du Croz, Numerical Algorithms Group Ltd.
|
||||
//> Sven Hammarling, Numerical Algorithms Group Ltd.
|
||||
//> \endverbatim
|
||||
//>
|
||||
// =====================================================================
|
||||
/* Subroutine */ int dgemm_(char *transa, char *transb, int *m, int *n, int *
|
||||
k, double *alpha, double *a, int *lda, double *b, int *ldb, double *
|
||||
beta, double *c__, int *ldc)
|
||||
{
|
||||
// System generated locals
|
||||
int a_dim1, a_offset, b_dim1, b_offset, c_dim1, c_offset, i__1, i__2,
|
||||
i__3;
|
||||
|
||||
// Local variables
|
||||
int i__, j, l, info;
|
||||
int nota, notb;
|
||||
double temp;
|
||||
int ncola;
|
||||
extern int lsame_(char *, char *);
|
||||
int nrowa, nrowb;
|
||||
extern /* Subroutine */ int xerbla_(char *, int *);
|
||||
|
||||
//
|
||||
// -- Reference BLAS level3 routine (version 3.7.0) --
|
||||
// -- Reference BLAS is a software package provided by Univ. of Tennessee, --
|
||||
// -- Univ. of California Berkeley, Univ. of Colorado Denver and NAG Ltd..--
|
||||
// December 2016
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// ..
|
||||
//
|
||||
// =====================================================================
|
||||
//
|
||||
// .. External Functions ..
|
||||
// ..
|
||||
// .. External Subroutines ..
|
||||
// ..
|
||||
// .. Intrinsic Functions ..
|
||||
// ..
|
||||
// .. Local Scalars ..
|
||||
// ..
|
||||
// .. Parameters ..
|
||||
// ..
|
||||
//
|
||||
// Set NOTA and NOTB as true if A and B respectively are not
|
||||
// transposed and set NROWA, NCOLA and NROWB as the number of rows
|
||||
// and columns of A and the number of rows of B respectively.
|
||||
//
|
||||
// Parameter adjustments
|
||||
a_dim1 = *lda;
|
||||
a_offset = 1 + a_dim1;
|
||||
a -= a_offset;
|
||||
b_dim1 = *ldb;
|
||||
b_offset = 1 + b_dim1;
|
||||
b -= b_offset;
|
||||
c_dim1 = *ldc;
|
||||
c_offset = 1 + c_dim1;
|
||||
c__ -= c_offset;
|
||||
|
||||
// Function Body
|
||||
nota = lsame_(transa, "N");
|
||||
notb = lsame_(transb, "N");
|
||||
if (nota) {
|
||||
nrowa = *m;
|
||||
ncola = *k;
|
||||
} else {
|
||||
nrowa = *k;
|
||||
ncola = *m;
|
||||
}
|
||||
if (notb) {
|
||||
nrowb = *k;
|
||||
} else {
|
||||
nrowb = *n;
|
||||
}
|
||||
//
|
||||
// Test the input parameters.
|
||||
//
|
||||
info = 0;
|
||||
if (! nota && ! lsame_(transa, "C") && ! lsame_(transa, "T")) {
|
||||
info = 1;
|
||||
} else if (! notb && ! lsame_(transb, "C") && ! lsame_(transb, "T")) {
|
||||
info = 2;
|
||||
} else if (*m < 0) {
|
||||
info = 3;
|
||||
} else if (*n < 0) {
|
||||
info = 4;
|
||||
} else if (*k < 0) {
|
||||
info = 5;
|
||||
} else if (*lda < max(1,nrowa)) {
|
||||
info = 8;
|
||||
} else if (*ldb < max(1,nrowb)) {
|
||||
info = 10;
|
||||
} else if (*ldc < max(1,*m)) {
|
||||
info = 13;
|
||||
}
|
||||
if (info != 0) {
|
||||
xerbla_("DGEMM ", &info);
|
||||
return 0;
|
||||
}
|
||||
//
|
||||
// Quick return if possible.
|
||||
//
|
||||
if (*m == 0 || *n == 0 || (*alpha == 0. || *k == 0) && *beta == 1.) {
|
||||
return 0;
|
||||
}
|
||||
//
|
||||
// And if alpha.eq.zero.
|
||||
//
|
||||
if (*alpha == 0.) {
|
||||
if (*beta == 0.) {
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
c__[i__ + j * c_dim1] = 0.;
|
||||
// L10:
|
||||
}
|
||||
// L20:
|
||||
}
|
||||
} else {
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
c__[i__ + j * c_dim1] = *beta * c__[i__ + j * c_dim1];
|
||||
// L30:
|
||||
}
|
||||
// L40:
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
//
|
||||
// Start the operations.
|
||||
//
|
||||
if (notb) {
|
||||
if (nota) {
|
||||
//
|
||||
// Form C := alpha*A*B + beta*C.
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
if (*beta == 0.) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
c__[i__ + j * c_dim1] = 0.;
|
||||
// L50:
|
||||
}
|
||||
} else if (*beta != 1.) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
c__[i__ + j * c_dim1] = *beta * c__[i__ + j * c_dim1];
|
||||
// L60:
|
||||
}
|
||||
}
|
||||
i__2 = *k;
|
||||
for (l = 1; l <= i__2; ++l) {
|
||||
temp = *alpha * b[l + j * b_dim1];
|
||||
i__3 = *m;
|
||||
for (i__ = 1; i__ <= i__3; ++i__) {
|
||||
c__[i__ + j * c_dim1] += temp * a[i__ + l * a_dim1];
|
||||
// L70:
|
||||
}
|
||||
// L80:
|
||||
}
|
||||
// L90:
|
||||
}
|
||||
} else {
|
||||
//
|
||||
// Form C := alpha*A**T*B + beta*C
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
temp = 0.;
|
||||
i__3 = *k;
|
||||
for (l = 1; l <= i__3; ++l) {
|
||||
temp += a[l + i__ * a_dim1] * b[l + j * b_dim1];
|
||||
// L100:
|
||||
}
|
||||
if (*beta == 0.) {
|
||||
c__[i__ + j * c_dim1] = *alpha * temp;
|
||||
} else {
|
||||
c__[i__ + j * c_dim1] = *alpha * temp + *beta * c__[
|
||||
i__ + j * c_dim1];
|
||||
}
|
||||
// L110:
|
||||
}
|
||||
// L120:
|
||||
}
|
||||
}
|
||||
} else {
|
||||
if (nota) {
|
||||
//
|
||||
// Form C := alpha*A*B**T + beta*C
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
if (*beta == 0.) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
c__[i__ + j * c_dim1] = 0.;
|
||||
// L130:
|
||||
}
|
||||
} else if (*beta != 1.) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
c__[i__ + j * c_dim1] = *beta * c__[i__ + j * c_dim1];
|
||||
// L140:
|
||||
}
|
||||
}
|
||||
i__2 = *k;
|
||||
for (l = 1; l <= i__2; ++l) {
|
||||
temp = *alpha * b[j + l * b_dim1];
|
||||
i__3 = *m;
|
||||
for (i__ = 1; i__ <= i__3; ++i__) {
|
||||
c__[i__ + j * c_dim1] += temp * a[i__ + l * a_dim1];
|
||||
// L150:
|
||||
}
|
||||
// L160:
|
||||
}
|
||||
// L170:
|
||||
}
|
||||
} else {
|
||||
//
|
||||
// Form C := alpha*A**T*B**T + beta*C
|
||||
//
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
temp = 0.;
|
||||
i__3 = *k;
|
||||
for (l = 1; l <= i__3; ++l) {
|
||||
temp += a[l + i__ * a_dim1] * b[j + l * b_dim1];
|
||||
// L180:
|
||||
}
|
||||
if (*beta == 0.) {
|
||||
c__[i__ + j * c_dim1] = *alpha * temp;
|
||||
} else {
|
||||
c__[i__ + j * c_dim1] = *alpha * temp + *beta * c__[
|
||||
i__ + j * c_dim1];
|
||||
}
|
||||
// L190:
|
||||
}
|
||||
// L200:
|
||||
}
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
//
|
||||
// End of DGEMM .
|
||||
//
|
||||
} // dgemm_
|
||||
|
||||
Vendored
-370
@@ -1,370 +0,0 @@
|
||||
/* -- translated by f2c (version 20201020 (for_lapack)). -- */
|
||||
|
||||
#include "f2c.h"
|
||||
|
||||
//> \brief \b DGEMV
|
||||
//
|
||||
// =========== DOCUMENTATION ===========
|
||||
//
|
||||
// Online html documentation available at
|
||||
// http://www.netlib.org/lapack/explore-html/
|
||||
//
|
||||
// Definition:
|
||||
// ===========
|
||||
//
|
||||
// SUBROUTINE DGEMV(TRANS,M,N,ALPHA,A,LDA,X,INCX,BETA,Y,INCY)
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// DOUBLE PRECISION ALPHA,BETA
|
||||
// INTEGER INCX,INCY,LDA,M,N
|
||||
// CHARACTER TRANS
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// DOUBLE PRECISION A(LDA,*),X(*),Y(*)
|
||||
// ..
|
||||
//
|
||||
//
|
||||
//> \par Purpose:
|
||||
// =============
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> DGEMV performs one of the matrix-vector operations
|
||||
//>
|
||||
//> y := alpha*A*x + beta*y, or y := alpha*A**T*x + beta*y,
|
||||
//>
|
||||
//> where alpha and beta are scalars, x and y are vectors and A is an
|
||||
//> m by n matrix.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Arguments:
|
||||
// ==========
|
||||
//
|
||||
//> \param[in] TRANS
|
||||
//> \verbatim
|
||||
//> TRANS is CHARACTER*1
|
||||
//> On entry, TRANS specifies the operation to be performed as
|
||||
//> follows:
|
||||
//>
|
||||
//> TRANS = 'N' or 'n' y := alpha*A*x + beta*y.
|
||||
//>
|
||||
//> TRANS = 'T' or 't' y := alpha*A**T*x + beta*y.
|
||||
//>
|
||||
//> TRANS = 'C' or 'c' y := alpha*A**T*x + beta*y.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] M
|
||||
//> \verbatim
|
||||
//> M is INTEGER
|
||||
//> On entry, M specifies the number of rows of the matrix A.
|
||||
//> M must be at least zero.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] N
|
||||
//> \verbatim
|
||||
//> N is INTEGER
|
||||
//> On entry, N specifies the number of columns of the matrix A.
|
||||
//> N must be at least zero.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] ALPHA
|
||||
//> \verbatim
|
||||
//> ALPHA is DOUBLE PRECISION.
|
||||
//> On entry, ALPHA specifies the scalar alpha.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] A
|
||||
//> \verbatim
|
||||
//> A is DOUBLE PRECISION array, dimension ( LDA, N )
|
||||
//> Before entry, the leading m by n part of the array A must
|
||||
//> contain the matrix of coefficients.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] LDA
|
||||
//> \verbatim
|
||||
//> LDA is INTEGER
|
||||
//> On entry, LDA specifies the first dimension of A as declared
|
||||
//> in the calling (sub) program. LDA must be at least
|
||||
//> max( 1, m ).
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] X
|
||||
//> \verbatim
|
||||
//> X is DOUBLE PRECISION array, dimension at least
|
||||
//> ( 1 + ( n - 1 )*abs( INCX ) ) when TRANS = 'N' or 'n'
|
||||
//> and at least
|
||||
//> ( 1 + ( m - 1 )*abs( INCX ) ) otherwise.
|
||||
//> Before entry, the incremented array X must contain the
|
||||
//> vector x.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] INCX
|
||||
//> \verbatim
|
||||
//> INCX is INTEGER
|
||||
//> On entry, INCX specifies the increment for the elements of
|
||||
//> X. INCX must not be zero.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] BETA
|
||||
//> \verbatim
|
||||
//> BETA is DOUBLE PRECISION.
|
||||
//> On entry, BETA specifies the scalar beta. When BETA is
|
||||
//> supplied as zero then Y need not be set on input.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in,out] Y
|
||||
//> \verbatim
|
||||
//> Y is DOUBLE PRECISION array, dimension at least
|
||||
//> ( 1 + ( m - 1 )*abs( INCY ) ) when TRANS = 'N' or 'n'
|
||||
//> and at least
|
||||
//> ( 1 + ( n - 1 )*abs( INCY ) ) otherwise.
|
||||
//> Before entry with BETA non-zero, the incremented array Y
|
||||
//> must contain the vector y. On exit, Y is overwritten by the
|
||||
//> updated vector y.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] INCY
|
||||
//> \verbatim
|
||||
//> INCY is INTEGER
|
||||
//> On entry, INCY specifies the increment for the elements of
|
||||
//> Y. INCY must not be zero.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Authors:
|
||||
// ========
|
||||
//
|
||||
//> \author Univ. of Tennessee
|
||||
//> \author Univ. of California Berkeley
|
||||
//> \author Univ. of Colorado Denver
|
||||
//> \author NAG Ltd.
|
||||
//
|
||||
//> \date December 2016
|
||||
//
|
||||
//> \ingroup double_blas_level2
|
||||
//
|
||||
//> \par Further Details:
|
||||
// =====================
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> Level 2 Blas routine.
|
||||
//> The vector and matrix arguments are not referenced when N = 0, or M = 0
|
||||
//>
|
||||
//> -- Written on 22-October-1986.
|
||||
//> Jack Dongarra, Argonne National Lab.
|
||||
//> Jeremy Du Croz, Nag Central Office.
|
||||
//> Sven Hammarling, Nag Central Office.
|
||||
//> Richard Hanson, Sandia National Labs.
|
||||
//> \endverbatim
|
||||
//>
|
||||
// =====================================================================
|
||||
/* Subroutine */ int dgemv_(char *trans, int *m, int *n, double *alpha,
|
||||
double *a, int *lda, double *x, int *incx, double *beta, double *y,
|
||||
int *incy)
|
||||
{
|
||||
// System generated locals
|
||||
int a_dim1, a_offset, i__1, i__2;
|
||||
|
||||
// Local variables
|
||||
int i__, j, ix, iy, jx, jy, kx, ky, info;
|
||||
double temp;
|
||||
int lenx, leny;
|
||||
extern int lsame_(char *, char *);
|
||||
extern /* Subroutine */ int xerbla_(char *, int *);
|
||||
|
||||
//
|
||||
// -- Reference BLAS level2 routine (version 3.7.0) --
|
||||
// -- Reference BLAS is a software package provided by Univ. of Tennessee, --
|
||||
// -- Univ. of California Berkeley, Univ. of Colorado Denver and NAG Ltd..--
|
||||
// December 2016
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// ..
|
||||
//
|
||||
// =====================================================================
|
||||
//
|
||||
// .. Parameters ..
|
||||
// ..
|
||||
// .. Local Scalars ..
|
||||
// ..
|
||||
// .. External Functions ..
|
||||
// ..
|
||||
// .. External Subroutines ..
|
||||
// ..
|
||||
// .. Intrinsic Functions ..
|
||||
// ..
|
||||
//
|
||||
// Test the input parameters.
|
||||
//
|
||||
// Parameter adjustments
|
||||
a_dim1 = *lda;
|
||||
a_offset = 1 + a_dim1;
|
||||
a -= a_offset;
|
||||
--x;
|
||||
--y;
|
||||
|
||||
// Function Body
|
||||
info = 0;
|
||||
if (! lsame_(trans, "N") && ! lsame_(trans, "T") && ! lsame_(trans, "C"))
|
||||
{
|
||||
info = 1;
|
||||
} else if (*m < 0) {
|
||||
info = 2;
|
||||
} else if (*n < 0) {
|
||||
info = 3;
|
||||
} else if (*lda < max(1,*m)) {
|
||||
info = 6;
|
||||
} else if (*incx == 0) {
|
||||
info = 8;
|
||||
} else if (*incy == 0) {
|
||||
info = 11;
|
||||
}
|
||||
if (info != 0) {
|
||||
xerbla_("DGEMV ", &info);
|
||||
return 0;
|
||||
}
|
||||
//
|
||||
// Quick return if possible.
|
||||
//
|
||||
if (*m == 0 || *n == 0 || *alpha == 0. && *beta == 1.) {
|
||||
return 0;
|
||||
}
|
||||
//
|
||||
// Set LENX and LENY, the lengths of the vectors x and y, and set
|
||||
// up the start points in X and Y.
|
||||
//
|
||||
if (lsame_(trans, "N")) {
|
||||
lenx = *n;
|
||||
leny = *m;
|
||||
} else {
|
||||
lenx = *m;
|
||||
leny = *n;
|
||||
}
|
||||
if (*incx > 0) {
|
||||
kx = 1;
|
||||
} else {
|
||||
kx = 1 - (lenx - 1) * *incx;
|
||||
}
|
||||
if (*incy > 0) {
|
||||
ky = 1;
|
||||
} else {
|
||||
ky = 1 - (leny - 1) * *incy;
|
||||
}
|
||||
//
|
||||
// Start the operations. In this version the elements of A are
|
||||
// accessed sequentially with one pass through A.
|
||||
//
|
||||
// First form y := beta*y.
|
||||
//
|
||||
if (*beta != 1.) {
|
||||
if (*incy == 1) {
|
||||
if (*beta == 0.) {
|
||||
i__1 = leny;
|
||||
for (i__ = 1; i__ <= i__1; ++i__) {
|
||||
y[i__] = 0.;
|
||||
// L10:
|
||||
}
|
||||
} else {
|
||||
i__1 = leny;
|
||||
for (i__ = 1; i__ <= i__1; ++i__) {
|
||||
y[i__] = *beta * y[i__];
|
||||
// L20:
|
||||
}
|
||||
}
|
||||
} else {
|
||||
iy = ky;
|
||||
if (*beta == 0.) {
|
||||
i__1 = leny;
|
||||
for (i__ = 1; i__ <= i__1; ++i__) {
|
||||
y[iy] = 0.;
|
||||
iy += *incy;
|
||||
// L30:
|
||||
}
|
||||
} else {
|
||||
i__1 = leny;
|
||||
for (i__ = 1; i__ <= i__1; ++i__) {
|
||||
y[iy] = *beta * y[iy];
|
||||
iy += *incy;
|
||||
// L40:
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
if (*alpha == 0.) {
|
||||
return 0;
|
||||
}
|
||||
if (lsame_(trans, "N")) {
|
||||
//
|
||||
// Form y := alpha*A*x + y.
|
||||
//
|
||||
jx = kx;
|
||||
if (*incy == 1) {
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
temp = *alpha * x[jx];
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
y[i__] += temp * a[i__ + j * a_dim1];
|
||||
// L50:
|
||||
}
|
||||
jx += *incx;
|
||||
// L60:
|
||||
}
|
||||
} else {
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
temp = *alpha * x[jx];
|
||||
iy = ky;
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
y[iy] += temp * a[i__ + j * a_dim1];
|
||||
iy += *incy;
|
||||
// L70:
|
||||
}
|
||||
jx += *incx;
|
||||
// L80:
|
||||
}
|
||||
}
|
||||
} else {
|
||||
//
|
||||
// Form y := alpha*A**T*x + y.
|
||||
//
|
||||
jy = ky;
|
||||
if (*incx == 1) {
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
temp = 0.;
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
temp += a[i__ + j * a_dim1] * x[i__];
|
||||
// L90:
|
||||
}
|
||||
y[jy] += *alpha * temp;
|
||||
jy += *incy;
|
||||
// L100:
|
||||
}
|
||||
} else {
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
temp = 0.;
|
||||
ix = kx;
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
temp += a[i__ + j * a_dim1] * x[ix];
|
||||
ix += *incx;
|
||||
// L110:
|
||||
}
|
||||
y[jy] += *alpha * temp;
|
||||
jy += *incy;
|
||||
// L120:
|
||||
}
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
//
|
||||
// End of DGEMV .
|
||||
//
|
||||
} // dgemv_
|
||||
|
||||
Vendored
-18599
File diff suppressed because it is too large
Load Diff
Vendored
-186
@@ -1,186 +0,0 @@
|
||||
/* -- translated by f2c (version 20201020 (for_lapack)). -- */
|
||||
|
||||
#include "f2c.h"
|
||||
|
||||
//> \brief \b DISNAN tests input for NaN.
|
||||
//
|
||||
// =========== DOCUMENTATION ===========
|
||||
//
|
||||
// Online html documentation available at
|
||||
// http://www.netlib.org/lapack/explore-html/
|
||||
//
|
||||
//> \htmlonly
|
||||
//> Download DISNAN + dependencies
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.tgz?format=tgz&filename=/lapack/lapack_routine/disnan.f">
|
||||
//> [TGZ]</a>
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.zip?format=zip&filename=/lapack/lapack_routine/disnan.f">
|
||||
//> [ZIP]</a>
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.txt?format=txt&filename=/lapack/lapack_routine/disnan.f">
|
||||
//> [TXT]</a>
|
||||
//> \endhtmlonly
|
||||
//
|
||||
// Definition:
|
||||
// ===========
|
||||
//
|
||||
// LOGICAL FUNCTION DISNAN( DIN )
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// DOUBLE PRECISION, INTENT(IN) :: DIN
|
||||
// ..
|
||||
//
|
||||
//
|
||||
//> \par Purpose:
|
||||
// =============
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> DISNAN returns .TRUE. if its argument is NaN, and .FALSE.
|
||||
//> otherwise. To be replaced by the Fortran 2003 intrinsic in the
|
||||
//> future.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Arguments:
|
||||
// ==========
|
||||
//
|
||||
//> \param[in] DIN
|
||||
//> \verbatim
|
||||
//> DIN is DOUBLE PRECISION
|
||||
//> Input to test for NaN.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Authors:
|
||||
// ========
|
||||
//
|
||||
//> \author Univ. of Tennessee
|
||||
//> \author Univ. of California Berkeley
|
||||
//> \author Univ. of Colorado Denver
|
||||
//> \author NAG Ltd.
|
||||
//
|
||||
//> \date June 2017
|
||||
//
|
||||
//> \ingroup OTHERauxiliary
|
||||
//
|
||||
// =====================================================================
|
||||
int disnan_(double *din)
|
||||
{
|
||||
// System generated locals
|
||||
int ret_val;
|
||||
|
||||
// Local variables
|
||||
extern int dlaisnan_(double *, double *);
|
||||
|
||||
//
|
||||
// -- LAPACK auxiliary routine (version 3.7.1) --
|
||||
// -- LAPACK is a software package provided by Univ. of Tennessee, --
|
||||
// -- Univ. of California Berkeley, Univ. of Colorado Denver and NAG Ltd..--
|
||||
// June 2017
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// ..
|
||||
//
|
||||
// =====================================================================
|
||||
//
|
||||
// .. External Functions ..
|
||||
// ..
|
||||
// .. Executable Statements ..
|
||||
ret_val = dlaisnan_(din, din);
|
||||
return ret_val;
|
||||
} // disnan_
|
||||
|
||||
/* -- translated by f2c (version 20201020 (for_lapack)). -- */
|
||||
|
||||
//> \brief \b DLAISNAN tests input for NaN by comparing two arguments for inequality.
|
||||
//
|
||||
// =========== DOCUMENTATION ===========
|
||||
//
|
||||
// Online html documentation available at
|
||||
// http://www.netlib.org/lapack/explore-html/
|
||||
//
|
||||
//> \htmlonly
|
||||
//> Download DLAISNAN + dependencies
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.tgz?format=tgz&filename=/lapack/lapack_routine/dlaisnan.f">
|
||||
//> [TGZ]</a>
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.zip?format=zip&filename=/lapack/lapack_routine/dlaisnan.f">
|
||||
//> [ZIP]</a>
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.txt?format=txt&filename=/lapack/lapack_routine/dlaisnan.f">
|
||||
//> [TXT]</a>
|
||||
//> \endhtmlonly
|
||||
//
|
||||
// Definition:
|
||||
// ===========
|
||||
//
|
||||
// LOGICAL FUNCTION DLAISNAN( DIN1, DIN2 )
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// DOUBLE PRECISION, INTENT(IN) :: DIN1, DIN2
|
||||
// ..
|
||||
//
|
||||
//
|
||||
//> \par Purpose:
|
||||
// =============
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> This routine is not for general use. It exists solely to avoid
|
||||
//> over-optimization in DISNAN.
|
||||
//>
|
||||
//> DLAISNAN checks for NaNs by comparing its two arguments for
|
||||
//> inequality. NaN is the only floating-point value where NaN != NaN
|
||||
//> returns .TRUE. To check for NaNs, pass the same variable as both
|
||||
//> arguments.
|
||||
//>
|
||||
//> A compiler must assume that the two arguments are
|
||||
//> not the same variable, and the test will not be optimized away.
|
||||
//> Interprocedural or whole-program optimization may delete this
|
||||
//> test. The ISNAN functions will be replaced by the correct
|
||||
//> Fortran 03 intrinsic once the intrinsic is widely available.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Arguments:
|
||||
// ==========
|
||||
//
|
||||
//> \param[in] DIN1
|
||||
//> \verbatim
|
||||
//> DIN1 is DOUBLE PRECISION
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] DIN2
|
||||
//> \verbatim
|
||||
//> DIN2 is DOUBLE PRECISION
|
||||
//> Two numbers to compare for inequality.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Authors:
|
||||
// ========
|
||||
//
|
||||
//> \author Univ. of Tennessee
|
||||
//> \author Univ. of California Berkeley
|
||||
//> \author Univ. of Colorado Denver
|
||||
//> \author NAG Ltd.
|
||||
//
|
||||
//> \date June 2017
|
||||
//
|
||||
//> \ingroup OTHERauxiliary
|
||||
//
|
||||
// =====================================================================
|
||||
int dlaisnan_(double *din1, double *din2)
|
||||
{
|
||||
// System generated locals
|
||||
int ret_val;
|
||||
|
||||
//
|
||||
// -- LAPACK auxiliary routine (version 3.7.1) --
|
||||
// -- LAPACK is a software package provided by Univ. of Tennessee, --
|
||||
// -- Univ. of California Berkeley, Univ. of Colorado Denver and NAG Ltd..--
|
||||
// June 2017
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// ..
|
||||
//
|
||||
// =====================================================================
|
||||
//
|
||||
// .. Executable Statements ..
|
||||
ret_val = *din1 != *din2;
|
||||
return ret_val;
|
||||
} // dlaisnan_
|
||||
|
||||
Vendored
-184
@@ -1,184 +0,0 @@
|
||||
/* -- translated by f2c (version 20201020 (for_lapack)). -- */
|
||||
|
||||
#include "f2c.h"
|
||||
|
||||
//> \brief \b DLACPY copies all or part of one two-dimensional array to another.
|
||||
//
|
||||
// =========== DOCUMENTATION ===========
|
||||
//
|
||||
// Online html documentation available at
|
||||
// http://www.netlib.org/lapack/explore-html/
|
||||
//
|
||||
//> \htmlonly
|
||||
//> Download DLACPY + dependencies
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.tgz?format=tgz&filename=/lapack/lapack_routine/dlacpy.f">
|
||||
//> [TGZ]</a>
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.zip?format=zip&filename=/lapack/lapack_routine/dlacpy.f">
|
||||
//> [ZIP]</a>
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.txt?format=txt&filename=/lapack/lapack_routine/dlacpy.f">
|
||||
//> [TXT]</a>
|
||||
//> \endhtmlonly
|
||||
//
|
||||
// Definition:
|
||||
// ===========
|
||||
//
|
||||
// SUBROUTINE DLACPY( UPLO, M, N, A, LDA, B, LDB )
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// CHARACTER UPLO
|
||||
// INTEGER LDA, LDB, M, N
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// DOUBLE PRECISION A( LDA, * ), B( LDB, * )
|
||||
// ..
|
||||
//
|
||||
//
|
||||
//> \par Purpose:
|
||||
// =============
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> DLACPY copies all or part of a two-dimensional matrix A to another
|
||||
//> matrix B.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Arguments:
|
||||
// ==========
|
||||
//
|
||||
//> \param[in] UPLO
|
||||
//> \verbatim
|
||||
//> UPLO is CHARACTER*1
|
||||
//> Specifies the part of the matrix A to be copied to B.
|
||||
//> = 'U': Upper triangular part
|
||||
//> = 'L': Lower triangular part
|
||||
//> Otherwise: All of the matrix A
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] M
|
||||
//> \verbatim
|
||||
//> M is INTEGER
|
||||
//> The number of rows of the matrix A. M >= 0.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] N
|
||||
//> \verbatim
|
||||
//> N is INTEGER
|
||||
//> The number of columns of the matrix A. N >= 0.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] A
|
||||
//> \verbatim
|
||||
//> A is DOUBLE PRECISION array, dimension (LDA,N)
|
||||
//> The m by n matrix A. If UPLO = 'U', only the upper triangle
|
||||
//> or trapezoid is accessed; if UPLO = 'L', only the lower
|
||||
//> triangle or trapezoid is accessed.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] LDA
|
||||
//> \verbatim
|
||||
//> LDA is INTEGER
|
||||
//> The leading dimension of the array A. LDA >= max(1,M).
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[out] B
|
||||
//> \verbatim
|
||||
//> B is DOUBLE PRECISION array, dimension (LDB,N)
|
||||
//> On exit, B = A in the locations specified by UPLO.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] LDB
|
||||
//> \verbatim
|
||||
//> LDB is INTEGER
|
||||
//> The leading dimension of the array B. LDB >= max(1,M).
|
||||
//> \endverbatim
|
||||
//
|
||||
// Authors:
|
||||
// ========
|
||||
//
|
||||
//> \author Univ. of Tennessee
|
||||
//> \author Univ. of California Berkeley
|
||||
//> \author Univ. of Colorado Denver
|
||||
//> \author NAG Ltd.
|
||||
//
|
||||
//> \date December 2016
|
||||
//
|
||||
//> \ingroup OTHERauxiliary
|
||||
//
|
||||
// =====================================================================
|
||||
/* Subroutine */ int dlacpy_(char *uplo, int *m, int *n, double *a, int *lda,
|
||||
double *b, int *ldb)
|
||||
{
|
||||
// System generated locals
|
||||
int a_dim1, a_offset, b_dim1, b_offset, i__1, i__2;
|
||||
|
||||
// Local variables
|
||||
int i__, j;
|
||||
extern int lsame_(char *, char *);
|
||||
|
||||
//
|
||||
// -- LAPACK auxiliary routine (version 3.7.0) --
|
||||
// -- LAPACK is a software package provided by Univ. of Tennessee, --
|
||||
// -- Univ. of California Berkeley, Univ. of Colorado Denver and NAG Ltd..--
|
||||
// December 2016
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// ..
|
||||
//
|
||||
// =====================================================================
|
||||
//
|
||||
// .. Local Scalars ..
|
||||
// ..
|
||||
// .. External Functions ..
|
||||
// ..
|
||||
// .. Intrinsic Functions ..
|
||||
// ..
|
||||
// .. Executable Statements ..
|
||||
//
|
||||
// Parameter adjustments
|
||||
a_dim1 = *lda;
|
||||
a_offset = 1 + a_dim1;
|
||||
a -= a_offset;
|
||||
b_dim1 = *ldb;
|
||||
b_offset = 1 + b_dim1;
|
||||
b -= b_offset;
|
||||
|
||||
// Function Body
|
||||
if (lsame_(uplo, "U")) {
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = min(j,*m);
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
b[i__ + j * b_dim1] = a[i__ + j * a_dim1];
|
||||
// L10:
|
||||
}
|
||||
// L20:
|
||||
}
|
||||
} else if (lsame_(uplo, "L")) {
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = j; i__ <= i__2; ++i__) {
|
||||
b[i__ + j * b_dim1] = a[i__ + j * a_dim1];
|
||||
// L30:
|
||||
}
|
||||
// L40:
|
||||
}
|
||||
} else {
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
b[i__ + j * b_dim1] = a[i__ + j * a_dim1];
|
||||
// L50:
|
||||
}
|
||||
// L60:
|
||||
}
|
||||
}
|
||||
return 0;
|
||||
//
|
||||
// End of DLACPY
|
||||
//
|
||||
} // dlacpy_
|
||||
|
||||
Vendored
-367
@@ -1,367 +0,0 @@
|
||||
/* -- translated by f2c (version 20201020 (for_lapack)). -- */
|
||||
|
||||
#include "f2c.h"
|
||||
|
||||
//> \brief \b DCOMBSSQ adds two scaled sum of squares quantities.
|
||||
//
|
||||
// =========== DOCUMENTATION ===========
|
||||
//
|
||||
// Online html documentation available at
|
||||
// http://www.netlib.org/lapack/explore-html/
|
||||
//
|
||||
//
|
||||
// Definition:
|
||||
// ===========
|
||||
//
|
||||
// SUBROUTINE DCOMBSSQ( V1, V2 )
|
||||
//
|
||||
// .. Array Arguments ..
|
||||
// DOUBLE PRECISION V1( 2 ), V2( 2 )
|
||||
// ..
|
||||
//
|
||||
//
|
||||
//> \par Purpose:
|
||||
// =============
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> DCOMBSSQ adds two scaled sum of squares quantities, V1 := V1 + V2.
|
||||
//> That is,
|
||||
//>
|
||||
//> V1_scale**2 * V1_sumsq := V1_scale**2 * V1_sumsq
|
||||
//> + V2_scale**2 * V2_sumsq
|
||||
//> \endverbatim
|
||||
//
|
||||
// Arguments:
|
||||
// ==========
|
||||
//
|
||||
//> \param[in,out] V1
|
||||
//> \verbatim
|
||||
//> V1 is DOUBLE PRECISION array, dimension (2).
|
||||
//> The first scaled sum.
|
||||
//> V1(1) = V1_scale, V1(2) = V1_sumsq.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] V2
|
||||
//> \verbatim
|
||||
//> V2 is DOUBLE PRECISION array, dimension (2).
|
||||
//> The second scaled sum.
|
||||
//> V2(1) = V2_scale, V2(2) = V2_sumsq.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Authors:
|
||||
// ========
|
||||
//
|
||||
//> \author Univ. of Tennessee
|
||||
//> \author Univ. of California Berkeley
|
||||
//> \author Univ. of Colorado Denver
|
||||
//> \author NAG Ltd.
|
||||
//
|
||||
//> \date November 2018
|
||||
//
|
||||
//> \ingroup OTHERauxiliary
|
||||
//
|
||||
// =====================================================================
|
||||
/* Subroutine */ int dcombssq_(double *v1, double *v2)
|
||||
{
|
||||
// System generated locals
|
||||
double d__1;
|
||||
|
||||
//
|
||||
// -- LAPACK auxiliary routine (version 3.7.0) --
|
||||
// -- LAPACK is a software package provided by Univ. of Tennessee, --
|
||||
// -- Univ. of California Berkeley, Univ. of Colorado Denver and NAG Ltd..--
|
||||
// November 2018
|
||||
//
|
||||
// .. Array Arguments ..
|
||||
// ..
|
||||
//
|
||||
//=====================================================================
|
||||
//
|
||||
// .. Parameters ..
|
||||
// ..
|
||||
// .. Executable Statements ..
|
||||
//
|
||||
// Parameter adjustments
|
||||
--v2;
|
||||
--v1;
|
||||
|
||||
// Function Body
|
||||
if (v1[1] >= v2[1]) {
|
||||
if (v1[1] != 0.) {
|
||||
// Computing 2nd power
|
||||
d__1 = v2[1] / v1[1];
|
||||
v1[2] += d__1 * d__1 * v2[2];
|
||||
}
|
||||
} else {
|
||||
// Computing 2nd power
|
||||
d__1 = v1[1] / v2[1];
|
||||
v1[2] = v2[2] + d__1 * d__1 * v1[2];
|
||||
v1[1] = v2[1];
|
||||
}
|
||||
return 0;
|
||||
//
|
||||
// End of DCOMBSSQ
|
||||
//
|
||||
} // dcombssq_
|
||||
|
||||
/* -- translated by f2c (version 20201020 (for_lapack)). -- */
|
||||
|
||||
//> \brief \b DLANGE returns the value of the 1-norm, Frobenius norm, infinity-norm, or the largest absolute value of any element of a general rectangular matrix.
|
||||
//
|
||||
// =========== DOCUMENTATION ===========
|
||||
//
|
||||
// Online html documentation available at
|
||||
// http://www.netlib.org/lapack/explore-html/
|
||||
//
|
||||
//> \htmlonly
|
||||
//> Download DLANGE + dependencies
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.tgz?format=tgz&filename=/lapack/lapack_routine/dlange.f">
|
||||
//> [TGZ]</a>
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.zip?format=zip&filename=/lapack/lapack_routine/dlange.f">
|
||||
//> [ZIP]</a>
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.txt?format=txt&filename=/lapack/lapack_routine/dlange.f">
|
||||
//> [TXT]</a>
|
||||
//> \endhtmlonly
|
||||
//
|
||||
// Definition:
|
||||
// ===========
|
||||
//
|
||||
// DOUBLE PRECISION FUNCTION DLANGE( NORM, M, N, A, LDA, WORK )
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// CHARACTER NORM
|
||||
// INTEGER LDA, M, N
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// DOUBLE PRECISION A( LDA, * ), WORK( * )
|
||||
// ..
|
||||
//
|
||||
//
|
||||
//> \par Purpose:
|
||||
// =============
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> DLANGE returns the value of the one norm, or the Frobenius norm, or
|
||||
//> the infinity norm, or the element of largest absolute value of a
|
||||
//> real matrix A.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \return DLANGE
|
||||
//> \verbatim
|
||||
//>
|
||||
//> DLANGE = ( max(abs(A(i,j))), NORM = 'M' or 'm'
|
||||
//> (
|
||||
//> ( norm1(A), NORM = '1', 'O' or 'o'
|
||||
//> (
|
||||
//> ( normI(A), NORM = 'I' or 'i'
|
||||
//> (
|
||||
//> ( normF(A), NORM = 'F', 'f', 'E' or 'e'
|
||||
//>
|
||||
//> where norm1 denotes the one norm of a matrix (maximum column sum),
|
||||
//> normI denotes the infinity norm of a matrix (maximum row sum) and
|
||||
//> normF denotes the Frobenius norm of a matrix (square root of sum of
|
||||
//> squares). Note that max(abs(A(i,j))) is not a consistent matrix norm.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Arguments:
|
||||
// ==========
|
||||
//
|
||||
//> \param[in] NORM
|
||||
//> \verbatim
|
||||
//> NORM is CHARACTER*1
|
||||
//> Specifies the value to be returned in DLANGE as described
|
||||
//> above.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] M
|
||||
//> \verbatim
|
||||
//> M is INTEGER
|
||||
//> The number of rows of the matrix A. M >= 0. When M = 0,
|
||||
//> DLANGE is set to zero.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] N
|
||||
//> \verbatim
|
||||
//> N is INTEGER
|
||||
//> The number of columns of the matrix A. N >= 0. When N = 0,
|
||||
//> DLANGE is set to zero.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] A
|
||||
//> \verbatim
|
||||
//> A is DOUBLE PRECISION array, dimension (LDA,N)
|
||||
//> The m by n matrix A.
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] LDA
|
||||
//> \verbatim
|
||||
//> LDA is INTEGER
|
||||
//> The leading dimension of the array A. LDA >= max(M,1).
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[out] WORK
|
||||
//> \verbatim
|
||||
//> WORK is DOUBLE PRECISION array, dimension (MAX(1,LWORK)),
|
||||
//> where LWORK >= M when NORM = 'I'; otherwise, WORK is not
|
||||
//> referenced.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Authors:
|
||||
// ========
|
||||
//
|
||||
//> \author Univ. of Tennessee
|
||||
//> \author Univ. of California Berkeley
|
||||
//> \author Univ. of Colorado Denver
|
||||
//> \author NAG Ltd.
|
||||
//
|
||||
//> \date December 2016
|
||||
//
|
||||
//> \ingroup doubleGEauxiliary
|
||||
//
|
||||
// =====================================================================
|
||||
double dlange_(char *norm, int *m, int *n, double *a, int *lda, double *work)
|
||||
{
|
||||
// Table of constant values
|
||||
int c__1 = 1;
|
||||
|
||||
// System generated locals
|
||||
int a_dim1, a_offset, i__1, i__2;
|
||||
double ret_val, d__1;
|
||||
|
||||
// Local variables
|
||||
extern /* Subroutine */ int dcombssq_(double *, double *);
|
||||
int i__, j;
|
||||
double sum, ssq[2], temp;
|
||||
extern int lsame_(char *, char *);
|
||||
double value;
|
||||
extern int disnan_(double *);
|
||||
extern /* Subroutine */ int dlassq_(int *, double *, int *, double *,
|
||||
double *);
|
||||
double colssq[2];
|
||||
|
||||
//
|
||||
// -- LAPACK auxiliary routine (version 3.7.0) --
|
||||
// -- LAPACK is a software package provided by Univ. of Tennessee, --
|
||||
// -- Univ. of California Berkeley, Univ. of Colorado Denver and NAG Ltd..--
|
||||
// December 2016
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// ..
|
||||
// .. Array Arguments ..
|
||||
// ..
|
||||
//
|
||||
//=====================================================================
|
||||
//
|
||||
// .. Parameters ..
|
||||
// ..
|
||||
// .. Local Scalars ..
|
||||
// ..
|
||||
// .. Local Arrays ..
|
||||
// ..
|
||||
// .. External Subroutines ..
|
||||
// ..
|
||||
// .. External Functions ..
|
||||
// ..
|
||||
// .. Intrinsic Functions ..
|
||||
// ..
|
||||
// .. Executable Statements ..
|
||||
//
|
||||
// Parameter adjustments
|
||||
a_dim1 = *lda;
|
||||
a_offset = 1 + a_dim1;
|
||||
a -= a_offset;
|
||||
--work;
|
||||
|
||||
// Function Body
|
||||
if (min(*m,*n) == 0) {
|
||||
value = 0.;
|
||||
} else if (lsame_(norm, "M")) {
|
||||
//
|
||||
// Find max(abs(A(i,j))).
|
||||
//
|
||||
value = 0.;
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
temp = (d__1 = a[i__ + j * a_dim1], abs(d__1));
|
||||
if (value < temp || disnan_(&temp)) {
|
||||
value = temp;
|
||||
}
|
||||
// L10:
|
||||
}
|
||||
// L20:
|
||||
}
|
||||
} else if (lsame_(norm, "O") || *(unsigned char *)norm == '1') {
|
||||
//
|
||||
// Find norm1(A).
|
||||
//
|
||||
value = 0.;
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
sum = 0.;
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
sum += (d__1 = a[i__ + j * a_dim1], abs(d__1));
|
||||
// L30:
|
||||
}
|
||||
if (value < sum || disnan_(&sum)) {
|
||||
value = sum;
|
||||
}
|
||||
// L40:
|
||||
}
|
||||
} else if (lsame_(norm, "I")) {
|
||||
//
|
||||
// Find normI(A).
|
||||
//
|
||||
i__1 = *m;
|
||||
for (i__ = 1; i__ <= i__1; ++i__) {
|
||||
work[i__] = 0.;
|
||||
// L50:
|
||||
}
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
i__2 = *m;
|
||||
for (i__ = 1; i__ <= i__2; ++i__) {
|
||||
work[i__] += (d__1 = a[i__ + j * a_dim1], abs(d__1));
|
||||
// L60:
|
||||
}
|
||||
// L70:
|
||||
}
|
||||
value = 0.;
|
||||
i__1 = *m;
|
||||
for (i__ = 1; i__ <= i__1; ++i__) {
|
||||
temp = work[i__];
|
||||
if (value < temp || disnan_(&temp)) {
|
||||
value = temp;
|
||||
}
|
||||
// L80:
|
||||
}
|
||||
} else if (lsame_(norm, "F") || lsame_(norm, "E")) {
|
||||
//
|
||||
// Find normF(A).
|
||||
// SSQ(1) is scale
|
||||
// SSQ(2) is sum-of-squares
|
||||
// For better accuracy, sum each column separately.
|
||||
//
|
||||
ssq[0] = 0.;
|
||||
ssq[1] = 1.;
|
||||
i__1 = *n;
|
||||
for (j = 1; j <= i__1; ++j) {
|
||||
colssq[0] = 0.;
|
||||
colssq[1] = 1.;
|
||||
dlassq_(m, &a[j * a_dim1 + 1], &c__1, colssq, &colssq[1]);
|
||||
dcombssq_(ssq, colssq);
|
||||
// L90:
|
||||
}
|
||||
value = ssq[0] * sqrt(ssq[1]);
|
||||
}
|
||||
ret_val = value;
|
||||
return ret_val;
|
||||
//
|
||||
// End of DLANGE
|
||||
//
|
||||
} // dlange_
|
||||
|
||||
Vendored
-125
@@ -1,125 +0,0 @@
|
||||
/* -- translated by f2c (version 20201020 (for_lapack)). -- */
|
||||
|
||||
#include "f2c.h"
|
||||
|
||||
//> \brief \b DLAPY2 returns sqrt(x2+y2).
|
||||
//
|
||||
// =========== DOCUMENTATION ===========
|
||||
//
|
||||
// Online html documentation available at
|
||||
// http://www.netlib.org/lapack/explore-html/
|
||||
//
|
||||
//> \htmlonly
|
||||
//> Download DLAPY2 + dependencies
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.tgz?format=tgz&filename=/lapack/lapack_routine/dlapy2.f">
|
||||
//> [TGZ]</a>
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.zip?format=zip&filename=/lapack/lapack_routine/dlapy2.f">
|
||||
//> [ZIP]</a>
|
||||
//> <a href="http://www.netlib.org/cgi-bin/netlibfiles.txt?format=txt&filename=/lapack/lapack_routine/dlapy2.f">
|
||||
//> [TXT]</a>
|
||||
//> \endhtmlonly
|
||||
//
|
||||
// Definition:
|
||||
// ===========
|
||||
//
|
||||
// DOUBLE PRECISION FUNCTION DLAPY2( X, Y )
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// DOUBLE PRECISION X, Y
|
||||
// ..
|
||||
//
|
||||
//
|
||||
//> \par Purpose:
|
||||
// =============
|
||||
//>
|
||||
//> \verbatim
|
||||
//>
|
||||
//> DLAPY2 returns sqrt(x**2+y**2), taking care not to cause unnecessary
|
||||
//> overflow.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Arguments:
|
||||
// ==========
|
||||
//
|
||||
//> \param[in] X
|
||||
//> \verbatim
|
||||
//> X is DOUBLE PRECISION
|
||||
//> \endverbatim
|
||||
//>
|
||||
//> \param[in] Y
|
||||
//> \verbatim
|
||||
//> Y is DOUBLE PRECISION
|
||||
//> X and Y specify the values x and y.
|
||||
//> \endverbatim
|
||||
//
|
||||
// Authors:
|
||||
// ========
|
||||
//
|
||||
//> \author Univ. of Tennessee
|
||||
//> \author Univ. of California Berkeley
|
||||
//> \author Univ. of Colorado Denver
|
||||
//> \author NAG Ltd.
|
||||
//
|
||||
//> \date June 2017
|
||||
//
|
||||
//> \ingroup OTHERauxiliary
|
||||
//
|
||||
// =====================================================================
|
||||
double dlapy2_(double *x, double *y)
|
||||
{
|
||||
// System generated locals
|
||||
double ret_val, d__1;
|
||||
|
||||
// Local variables
|
||||
int x_is_nan__, y_is_nan__;
|
||||
double w, z__, xabs, yabs;
|
||||
extern int disnan_(double *);
|
||||
|
||||
//
|
||||
// -- LAPACK auxiliary routine (version 3.7.1) --
|
||||
// -- LAPACK is a software package provided by Univ. of Tennessee, --
|
||||
// -- Univ. of California Berkeley, Univ. of Colorado Denver and NAG Ltd..--
|
||||
// June 2017
|
||||
//
|
||||
// .. Scalar Arguments ..
|
||||
// ..
|
||||
//
|
||||
// =====================================================================
|
||||
//
|
||||
// .. Parameters ..
|
||||
// ..
|
||||
// .. Local Scalars ..
|
||||
// ..
|
||||
// .. External Functions ..
|
||||
// ..
|
||||
// .. Intrinsic Functions ..
|
||||
// ..
|
||||
// .. Executable Statements ..
|
||||
//
|
||||
x_is_nan__ = disnan_(x);
|
||||
y_is_nan__ = disnan_(y);
|
||||
if (x_is_nan__) {
|
||||
ret_val = *x;
|
||||
}
|
||||
if (y_is_nan__) {
|
||||
ret_val = *y;
|
||||
}
|
||||
if (! (x_is_nan__ || y_is_nan__)) {
|
||||
xabs = abs(*x);
|
||||
yabs = abs(*y);
|
||||
w = max(xabs,yabs);
|
||||
z__ = min(xabs,yabs);
|
||||
if (z__ == 0.) {
|
||||
ret_val = w;
|
||||
} else {
|
||||
// Computing 2nd power
|
||||
d__1 = z__ / w;
|
||||
ret_val = w * sqrt(d__1 * d__1 + 1.);
|
||||
}
|
||||
}
|
||||
return ret_val;
|
||||
//
|
||||
// End of DLAPY2
|
||||
//
|
||||
} // dlapy2_
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user