1
0
mirror of https://github.com/nlohmann/json.git synced 2026-07-24 04:33:02 +04:00

Compare commits

..

13 Commits

Author SHA1 Message Date
Angadi56 9a3ebb9456 validate BSON document size against the bytes read (#5287) 2026-07-23 19:51:40 +00:00
Luke Banicevic 3296a3ad8c Docs: clarify there is no official npm package (impersonation/typosquat awareness) (#5299) 2026-07-23 16:02:29 +00:00
Angadi56 3565f40229 reject negative UBJSON/BJData string length (#5284) 2026-07-20 20:32:38 +00:00
Angadi56 1c5a953de5 reject unpaired UTF-16 surrogates in wide-string input (#5276) 2026-07-19 18:33:42 +02:00
dependabot[bot] a03e65420c ⬆️ Bump the codeql-action group with 4 updates (#5275) 2026-07-18 13:59:20 +02:00
dependabot[bot] d6ede37088 ⬆️ Bump actions/stale from 10.3.0 to 10.4.0 (#5277) 2026-07-18 13:58:45 +02:00
Niels Lohmann 722c03495f Extend std::optional null regression coverage (assignment, get_to) (#5269) 2026-07-12 09:16:25 +02:00
Niels Lohmann c197feff81 Extend memcpy fast path to sized sentinels (e.g. std::counted_iterator) (#5268) 2026-07-12 09:16:06 +02:00
Niels Lohmann b2b47c69b1 📝 Document std::pair/std::tuple conversion and C++20 range-view construction (#5271)
Both had zero documentation anywhere in docs/mkdocs/. The tuple/pair
gap was first spotted in the very first git-log audit pass but never
turned into an actionable todo, so it persisted uncaught across four
subsequent passes.

- Document basic positional std::pair/std::tuple <-> json array
  conversion, plus #5016's reference-extraction capability
  (get<std::tuple<T&, T&>>() returning references into the stored
  array elements).
- Document #5205's new json-from-C++20-range-view constructor
  (e.g. nums | std::views::filter(...)).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-07-11 23:58:49 +02:00
Niels Lohmann 6a406ee141 Document that JSON_Diagnostics CMake option doesn't apply to pre-installed packages (#5270)
Closes #3106. set(JSON_Diagnostics ON) before find_package() has no
effect on a package built and installed elsewhere (Homebrew, vcpkg, a
system package, etc.) -- the compile definition is baked into the
exported nlohmann_jsonTargets.cmake at install time and the generated
config script never re-reads that variable. Verified empirically
against the real Homebrew-installed 3.12.0 package: the exported
target carries a fixed $<$<BOOL:OFF>:JSON_DIAGNOSTICS=1>, and the
suggested set(JSON_Diagnostics ON) snippet produces no change in
exception output.

Documents the actual working fix (overriding the imported target's
INTERFACE_COMPILE_DEFINITIONS property after find_package()) and the
multi-target "JSON_DIAGNOSTICS redefined" pitfall reported earlier in
the issue thread.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-11 23:58:18 +02:00
Niels Lohmann ca76c37650 Add iterator+sentinel tests and docs for binary deserializers (#5265)
* Add iterator+sentinel tests and docs for binary deserializers

This commit extends the C++20 ranges support (iterator+sentinel pairs) to the
binary format deserializers from_cbor, from_msgpack, from_ubjson, from_bjdata,
and from_bson, matching what was already done for parse(), accept(), and
sax_parse().

Changes:
- Add istreambuf_sentinel helper to test_utils.hpp for EOF detection in tests
- Add 5 new test cases that read binary files directly via
  std::istreambuf_iterator<char> + sentinel, without pre-buffering
- Update documentation for all 5 from_* functions to document overload (3)
  with SentinelType parameter
- All tests pass; verified against existing test suite data
- Fix potential buffer over-read warning in heterogeneous iterator test

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Merge iterator+sentinel overloads and fix ambiguity/CI issues

Address PR review feedback and CI failures:

- Merge the separate same-type and sentinel-type iterator overloads of
  parse(), accept(), sax_parse(), and the five from_* binary deserializers
  into a single overload with SentinelType defaulted to IteratorType,
  as suggested in review. Applied the same simplification to the
  detail::input_adapter() free functions.
- Fix a latent ambiguity: some compilers (e.g. GCC 4.8) unreliably SFINAE
  the operator!= detection for std::nullptr_t against container/string
  types, making calls like parse(s, nullptr, ...) ambiguous with the
  compatible-input overload. can_compare_ne now explicitly excludes
  std::nullptr_t as a SentinelType.
- Use a named enable_if_t template parameter instead of an unnamed
  function parameter for the SFINAE guard, fixing a clang-tidy
  hicpp-named-parameter/readability-named-parameter failure.
- Update parse.md, accept.md, sax_parse.md, and the five from_*.md pages
  to document the merged overload instead of separate (2)/(3) overloads,
  also fixing an over-160-char line that broke the documentation
  style_check CI job.
- Rework the BSON iterator+sentinel test to parse a BSON file already
  present in the test suite instead of writing/deleting a temp file.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix -Wunneeded-internal-declaration for CustomSentinel in test

CustomSentinel lives in an anonymous namespace (internal linkage), and
the library's parse loop only ever evaluates the iterator-first
direction (it != last), so the reversed-order friend operator!= was
never referenced. Clang's -Weverything flags such unused internal
declarations as an error. Drop the unused overload; the used direction
is enough to satisfy can_compare_ne's either-order detection.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix clang-tidy hicpp-named-parameter and misc-const-correctness

- Drop the unused reversed-order operator!= overload from
  utils::istreambuf_sentinel (only iterator != sentinel is ever
  evaluated) and name the remaining friend's sentinel parameter, fixing
  hicpp-named-parameter/readability-named-parameter.
- Mark the istreambuf_iterator first/last helper variable const in the
  five binary-format sentinel tests, fixing misc-const-correctness.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix clang-tidy misc-const-correctness in heterogeneous sentinel test

json_str is only read via .data()/.size() and never reassigned, so
clang-tidy correctly flags it as const-able. Verified against the exact
CI job (silkeh/clang:dev, ci_clang_tidy target) by running clang-tidy
directly on this file plus the five binary-format sentinel tests
touched by prior commits; all are now clean.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-07-11 15:09:28 +02:00
Niels Lohmann fe2bcc080f Custom BinaryType direct assignment (#5266)
* Fix discussion #4209: custom BinaryType direct assignment and extraction

When a custom BinaryType is configured (other than the default std::vector<uint8_t>),
users can now:
1. Assign values of that type directly to create binary values (not arrays)
2. Extract binary values back to that type with get<>()
3. Extract arrays to that type (for backward compatibility)

Implementation:
- Add is_compatible_binary_type trait to centralize SFINAE condition
- Update to_json to accept custom BinaryType values directly
- Update from_json to handle both binary and array inputs for custom BinaryType
- Add #include <vector> with IWYU comment to from_json.hpp
- Add comprehensive tests for assignment and array extraction
- Update binary_t documentation with example

This is purely additive and invisible to the default nlohmann::json alias, which
continues to treat std::vector<uint8_t> as arrays.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: missing include and const-correctness

- Add #include <vector> to type_traits.hpp for the new
  is_compatible_binary_type trait's std::vector<std::uint8_t> reference
  (caught by cpplint's include-what-you-use check)
- Mark test-local json variables const where never reassigned
  (caught by clang-tidy's misc-const-correctness check)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-07-11 15:09:05 +02:00
Federico Sfriso c60217e801 fix: support constructing json from C++20 range views (#4916) (#5205) 2026-07-11 09:13:04 +02:00
24 changed files with 816 additions and 44 deletions
+8
View File
@@ -15,6 +15,14 @@ guidance.
For vulnerabilities in third-party dependencies or modules, please report them directly to the respective maintainers.
## Unofficial packages
This project does not publish an official npm package. The npm package
[`nlohmann-json`](https://www.npmjs.com/package/nlohmann-json) (or similarly named packages) is not maintained or
endorsed by this project. See the
[package managers documentation](https://json.nlohmann.me/integration/package_managers/#npm) for supported
integration options.
## Additional Resources
- Explore security-related topics and contribute to tools and projects through
+3 -3
View File
@@ -38,14 +38,14 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/init@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with:
languages: c-cpp
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below)
- name: Autobuild
uses: github/codeql-action/autobuild@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/autobuild@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/analyze@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
+1 -1
View File
@@ -43,6 +43,6 @@ jobs:
output: 'flawfinder_results.sarif'
- name: Upload analysis results to GitHub Security tab
uses: github/codeql-action/upload-sarif@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/upload-sarif@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with:
sarif_file: ${{github.workspace}}/flawfinder_results.sarif
+1 -1
View File
@@ -76,6 +76,6 @@ jobs:
# Upload the results to GitHub's code scanning dashboard.
- name: "Upload to code-scanning"
uses: github/codeql-action/upload-sarif@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/upload-sarif@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with:
sarif_file: results.sarif
+1 -1
View File
@@ -61,7 +61,7 @@ jobs:
# Upload SARIF file generated in previous step
- name: Upload SARIF file
uses: github/codeql-action/upload-sarif@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/upload-sarif@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with:
sarif_file: semgrep.sarif
if: always()
+1 -1
View File
@@ -20,7 +20,7 @@ jobs:
with:
egress-policy: audit
- uses: actions/stale@eb5cf3af3ac0a1aa4c9c45633dd1ae542a27a899 # v10.3.0
- uses: actions/stale@1e223db275d687790206a7acac4d1a11bd6fe629 # v10.4.0
with:
stale-issue-label: 'state: stale'
stale-pr-label: 'state: stale'
@@ -51,6 +51,38 @@ represent a byte array in modern C++.
The default values for `BinaryType` is `#!cpp std::vector<std::uint8_t>`.
#### Custom BinaryType behavior
When a custom `BinaryType` is configured (other than the default `#!cpp std::vector<std::uint8_t>`), you can assign
values of that type directly to a `basic_json` instance, and they will automatically be recognized as binary values
rather than arrays:
```cpp
using custom_json = nlohmann::basic_json<
nlohmann::ordered_map, // ObjectType
std::vector, // ArrayType
std::string, // StringType
bool, // BooleanType
std::int64_t, // NumberIntegerType
std::uint64_t, // NumberUnsignedType
double, // NumberFloatType
std::allocator, // AllocatorType
nlohmann::adl_serializer,
std::vector<std::byte> // Custom BinaryType
>;
std::vector<std::byte> data{std::byte{1}, std::byte{2}, std::byte{3}};
custom_json j = data; // Creates a binary value, not an array
assert(j.is_binary());
// Round-tripping works seamlessly
auto extracted = j.get<std::vector<std::byte>>();
assert(extracted == data);
```
This automatic type detection is a convenience feature that only applies to custom (non-default) `BinaryType` configurations.
The default `nlohmann::json` continues to treat `#!cpp std::vector<std::uint8_t>` as arrays for backward compatibility.
#### Storage
Binary Arrays are stored as pointers in a `basic_json` type. That is, for any access to array values, a pointer of the
@@ -38,7 +38,8 @@ When the macro is not defined, the library will define it to its default value.
Diagnostic messages can also be controlled with the CMake option
[`JSON_Diagnostics`](../../integration/cmake.md#json_diagnostics) (`OFF` by default)
which defines `JSON_DIAGNOSTICS` accordingly.
which defines `JSON_DIAGNOSTICS` accordingly. Note this only applies when building the
library from source — see the pre-installed-package caveat on that page.
## Examples
+36
View File
@@ -47,6 +47,28 @@ json j = {{"one", 1}, {"two", 2}};
auto m = j.get<std::map<std::string, int>>(); // {{"one", 1}, {"two", 2}}
```
`#!cpp std::pair` and `#!cpp std::tuple` are also supported, converting positionally to and from a JSON array:
```cpp
json j = {1.0, "hello", 42};
auto t = j.get<std::tuple<double, std::string, int>>(); // {1.0, "hello", 42}
```
!!! info "Extracting references into a tuple"
A tuple type may also hold references (e.g. `#!cpp std::tuple<double&, std::string&>`) to avoid copying: `get`
then returns a tuple of references pointing directly at the elements stored inside the `basic_json` array,
rather than a tuple of copies:
```cpp
json j = {1.0, "hello"};
auto refs = j.get<std::tuple<double&, std::string&>>();
std::get<1>(refs) = "world"; // modifies j[1] in place
```
A referenced type must be one the library actually stores (or an arithmetic type it can convert to/from);
otherwise this is a compile error.
## Implicit conversions
By default, a JSON value implicitly converts to a compatible C++ type, so the explicit `get` call can often be omitted:
@@ -136,6 +158,20 @@ std::vector<int> numbers = {1, 2, 3};
json j = numbers; // [1,2,3]
```
!!! info "Constructing from a C++20 range view"
A `json` array can also be constructed directly from a C++20 range view (`std::ranges::view`), such as the result
of `std::views::filter` or `std::views::transform` -- no intermediate container is needed:
```cpp
std::vector<int> nums{1, 2, 37, 42, 21};
auto filtered = nums | std::views::filter([](int i) { return i > 10; });
json j(filtered); // [37,42,21]
```
This requires [`JSON_HAS_RANGES`](../api/macros/json_has_ranges.md) to be enabled and is unavailable on MinGW due
to incomplete C++20 ranges support there.
## Your own types
The conversions above are built in for standard types. To make the same syntax work for **your own** types, provide
+25
View File
@@ -135,6 +135,31 @@ Enable CI build targets. The exact targets are used during the several CI steps
Enable [extended diagnostic messages](../home/exceptions.md#extended-diagnostic-messages) by defining macro [`JSON_DIAGNOSTICS`](../api/macros/json_diagnostics.md). This option is `OFF` by default.
!!! warning "Does not apply to a pre-installed package"
This option only takes effect when building nlohmann/json from source as part of your own
CMake project (e.g. via [`FetchContent`](#fetchcontent) or [`add_subdirectory`](#external)).
It has **no effect** on a package that was already built and installed elsewhere (Homebrew,
vcpkg, a system package, etc.) — the resulting compile definition is baked into the exported
`nlohmann_jsonTargets.cmake` at install time, and `set(JSON_Diagnostics ON)` before
`find_package()` does not change it (verified against the Homebrew-installed package: the
exported target still carries a fixed `$<$<BOOL:OFF>:JSON_DIAGNOSTICS=1>`, regardless of any
variable set in the consuming project).
To enable extended diagnostics for a pre-installed package, override the imported target's
property directly after `find_package()`:
```cmake
find_package(nlohmann_json REQUIRED)
set_target_properties(nlohmann_json::nlohmann_json PROPERTIES
INTERFACE_COMPILE_DEFINITIONS "JSON_DIAGNOSTICS=1")
```
This only works cleanly when your project is the sole consumer of that imported target. If
nlohmann_json is pulled in from more than one place in your dependency graph with different
`JSON_DIAGNOSTICS` values, you may see a `"JSON_DIAGNOSTICS" redefined` compiler error, since
conflicting `-D` flags can end up on the same compile command line.
### `JSON_Diagnostic_Positions`
Enable position diagnostics by defining macro [`JSON_DIAGNOSTIC_POSITIONS`](../api/macros/json_diagnostic_positions.md). This option is `OFF` by default.
@@ -930,6 +930,12 @@ If you are using [CocoaPods](https://cocoapods.org), you can use the library by
to your podfile (see [an example](https://bitbucket.org/benman/nlohmann_json-cocoapod/src/master/)). Please file issues
[here](https://bitbucket.org/benman/nlohmann_json-cocoapod/issues?status=new&status=open).
## npm
This project does not publish an official [npm](https://www.npmjs.com) package. The npm package
[`nlohmann-json`](https://www.npmjs.com/package/nlohmann-json) (or similarly named packages) is not maintained or
endorsed by this project. Use one of the package managers listed above, or integrate the single header directly.
## ESP-IDF and PlatformIO
There is no official package published to the [ESP-IDF Component Registry](https://components.espressif.com) or the
@@ -19,6 +19,7 @@
#include <unordered_map> // unordered_map
#include <utility> // pair, declval
#include <valarray> // valarray
#include <vector> // vector
#include <nlohmann/detail/exceptions.hpp>
#include <nlohmann/detail/macro_scope.hpp>
@@ -332,6 +333,7 @@ template < typename BasicJsonType, typename ConstructibleArrayType,
!is_constructible_object_type<BasicJsonType, ConstructibleArrayType>::value&&
!is_constructible_string_type<BasicJsonType, ConstructibleArrayType>::value&&
!std::is_same<ConstructibleArrayType, typename BasicJsonType::binary_t>::value&&
!is_compatible_binary_type<BasicJsonType, ConstructibleArrayType>::value&&
!is_basic_json<ConstructibleArrayType>::value,
int > = 0 >
auto from_json(const BasicJsonType& j, ConstructibleArrayType& arr)
@@ -377,6 +379,25 @@ inline void from_json(const BasicJsonType& j, typename BasicJsonType::binary_t&
bin = *j.template get_ptr<const typename BasicJsonType::binary_t*>();
}
template < typename BasicJsonType, typename CompatibleArrayType,
enable_if_t < is_compatible_binary_type<BasicJsonType, CompatibleArrayType>::value,
int > = 0 >
inline void from_json(const BasicJsonType& j, CompatibleArrayType& bin)
{
if (j.is_binary())
{
bin = static_cast<CompatibleArrayType>(*j.template get_ptr<const typename BasicJsonType::binary_t*>());
}
else if (j.is_array())
{
from_json_array_impl(j, bin, priority_tag<3> {});
}
else
{
JSON_THROW(type_error::create(302, concat("type must be binary or array, but is ", j.type_name()), &j));
}
}
template<typename BasicJsonType, typename ConstructibleObjectType,
enable_if_t<is_constructible_object_type<BasicJsonType, ConstructibleObjectType>::value, int> = 0>
inline void from_json(const BasicJsonType& j, ConstructibleObjectType& obj)
@@ -177,8 +177,11 @@ struct external_constructor<value_t::array>
}
template < typename BasicJsonType, typename CompatibleArrayType,
enable_if_t < !std::is_same<CompatibleArrayType, typename BasicJsonType::array_t>::value,
int > = 0 >
enable_if_t < !std::is_same<CompatibleArrayType, typename BasicJsonType::array_t>::value
#if JSON_HAS_RANGES && !defined(__MINGW32__)
&& !is_compatible_range_view<CompatibleArrayType>::value
#endif
, int > = 0 >
static void construct(BasicJsonType& j, const CompatibleArrayType& arr)
{
using std::begin;
@@ -218,6 +221,25 @@ struct external_constructor<value_t::array>
j.set_parents();
j.assert_invariant();
}
// std::ranges does not work properly on MinGW due to incomplete C++20 support
// see https://github.com/nlohmann/json/issues/4916
#if JSON_HAS_RANGES && !defined(__MINGW32__)
template<typename BasicJsonType, typename CompatibleArrayType,
enable_if_t<is_compatible_range_view<std::remove_cvref_t<CompatibleArrayType>>::value, int> = 0>
static void construct(BasicJsonType& j, CompatibleArrayType && arr)
{
j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array;
j.m_data.m_value = value_t::array;
for (auto&& x : std::forward<CompatibleArrayType>(arr))
{
j.m_data.m_value.array->push_back(x);
j.set_parent(j.m_data.m_value.array->back());
}
j.assert_invariant();
}
#endif
};
template<>
@@ -355,19 +377,44 @@ template < typename BasicJsonType, typename CompatibleArrayType,
!is_compatible_object_type<BasicJsonType, CompatibleArrayType>::value&&
!is_compatible_string_type<BasicJsonType, CompatibleArrayType>::value&&
!std::is_same<typename BasicJsonType::binary_t, CompatibleArrayType>::value&&
!is_basic_json<CompatibleArrayType>::value,
!is_compatible_binary_type<BasicJsonType, CompatibleArrayType>::value&&
!is_basic_json<CompatibleArrayType>::value
#if JSON_HAS_RANGES && !defined(__MINGW32__)
&& !is_compatible_range_view<CompatibleArrayType>::value
#endif
,
int > = 0 >
inline void to_json(BasicJsonType& j, const CompatibleArrayType& arr)
{
external_constructor<value_t::array>::construct(j, arr);
}
#if JSON_HAS_RANGES && !defined(__MINGW32__)
template < typename BasicJsonType, typename T,
enable_if_t < is_compatible_range_view<std::remove_cvref_t<T>>::value
&& !is_compatible_string_type<BasicJsonType, std::remove_cvref_t<T>>::value
&& !is_compatible_object_type<BasicJsonType, std::remove_cvref_t<T>>::value
&& !is_basic_json<std::remove_cvref_t<T>>::value, int > = 0 >
inline void to_json(BasicJsonType& j, T && arr)
{
external_constructor<value_t::array>::construct(j, std::forward<T>(arr));
}
#endif
template<typename BasicJsonType>
inline void to_json(BasicJsonType& j, const typename BasicJsonType::binary_t& bin)
{
external_constructor<value_t::binary>::construct(j, bin);
}
template < typename BasicJsonType, typename CompatibleArrayType,
enable_if_t < is_compatible_binary_type<BasicJsonType, CompatibleArrayType>::value,
int > = 0 >
inline void to_json(BasicJsonType& j, const CompatibleArrayType& bin)
{
external_constructor<value_t::binary>::construct(j, typename BasicJsonType::binary_t(bin));
}
template<typename BasicJsonType, typename T,
enable_if_t<std::is_convertible<T, BasicJsonType>::value, int> = 0>
inline void to_json(BasicJsonType& j, const std::valarray<T>& arr)
@@ -163,12 +163,39 @@ class binary_reader
// BSON //
//////////
/*!
@brief Validate a BSON document's declared size against the bytes read.
A BSON document starts with an int32 that counts its own total length in
bytes, including that prefix and the trailing 0x00. The reader is driven
by the terminator rather than the declared length, so without this check a
nested document could declare a length that disagrees with where its
terminator actually falls and quietly hand the bytes in between to the
enclosing document. A well-formed document is at least 5 bytes (the prefix
plus the terminator); the equality also rejects those impossible sizes,
since at least 5 bytes are always consumed.
@param[in] document_start value of chars_read before the size prefix
@param[in] document_size the declared document size
@return whether the declared size matches the number of bytes read
*/
bool check_bson_document_size(const std::size_t document_start, const std::int32_t document_size)
{
if (JSON_HEDLEY_UNLIKELY(document_size < 0 || static_cast<std::size_t>(document_size) != chars_read - document_start))
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(112, chars_read,
exception_message(input_format_t::bson, concat("document size ", std::to_string(document_size), " does not match the number of bytes read (", std::to_string(chars_read - document_start), ")"), "document"), nullptr));
}
return true;
}
/*!
@brief Reads in a BSON-object and passes it to the SAX-parser.
@return whether a valid BSON-value was passed to the SAX parser
*/
bool parse_bson_internal()
{
const std::size_t document_start = chars_read;
std::int32_t document_size{};
get_number<std::int32_t, true>(input_format_t::bson, document_size);
@@ -182,6 +209,11 @@ class binary_reader
return false;
}
if (JSON_HEDLEY_UNLIKELY(!check_bson_document_size(document_start, document_size)))
{
return false;
}
return sax->end_object();
}
@@ -397,6 +429,7 @@ class binary_reader
*/
bool parse_bson_array()
{
const std::size_t document_start = chars_read;
std::int32_t document_size{};
get_number<std::int32_t, true>(input_format_t::bson, document_size);
@@ -410,6 +443,11 @@ class binary_reader
return false;
}
if (JSON_HEDLEY_UNLIKELY(!check_bson_document_size(document_start, document_size)))
{
return false;
}
return sax->end_array();
}
@@ -1846,6 +1884,29 @@ class binary_reader
return get_ubjson_value(get_char ? get_ignore_noop() : current);
}
/*!
@brief reject a negative UBJSON/BJData string length
String and key lengths are written with signed integer markers (i, I, l,
L). A negative value is malformed; without this check get_string() would
silently treat it as an empty string and leave the following bytes to be
misread as the next value. This mirrors the non-negative check the
optimized-container count path already performs in get_ubjson_size_value.
@param[in] len the string length read from the input
@return whether the length is valid (non-negative)
*/
template<typename NumberType>
bool check_ubjson_string_length(const NumberType len)
{
if (JSON_HEDLEY_UNLIKELY(len < 0))
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(113, chars_read,
exception_message(input_format, "string length must not be negative", "string"), nullptr));
}
return true;
}
/*!
@brief reads a UBJSON string
@@ -1883,25 +1944,25 @@ class binary_reader
case 'i':
{
std::int8_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'I':
{
std::int16_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'l':
{
std::int32_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'L':
{
std::int64_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'u':
@@ -395,17 +395,30 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
}
else
{
if (JSON_HEDLEY_UNLIKELY(!input.empty()))
// A supplementary code point is a high surrogate (0xD800..0xDBFF)
// followed by a low surrogate (0xDC00..0xDFFF). A lone low
// surrogate, a high surrogate at the end of the input, or a high
// surrogate followed by any other unit is malformed UTF-16. In
// that case the offending unit is passed through unchanged so the
// UTF-8 decoder rejects it, matching how \uXXXX surrogate escapes
// are handled in the lexer.
bool valid_pair = false;
if (wc <= 0xDBFF && JSON_HEDLEY_UNLIKELY(!input.empty()))
{
const auto wc2 = static_cast<unsigned int>(input.get_character());
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu));
utf8_bytes_filled = 4;
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
{
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu));
utf8_bytes_filled = 4;
valid_pair = true;
}
}
else
if (!valid_pair)
{
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
utf8_bytes_filled = 1;
+83 -1
View File
@@ -13,11 +13,15 @@
#include <tuple> // tuple
#include <type_traits> // false_type, is_constructible, is_integral, is_same, true_type
#include <utility> // declval
#include <vector> // vector
#if defined(__cpp_lib_byte) && __cpp_lib_byte >= 201603L
#include <cstddef> // byte
#endif
#include <nlohmann/detail/iterators/iterator_traits.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#ifdef JSON_HAS_CPP_17
#include <optional> // optional
#endif
#include <nlohmann/detail/meta/call_std/begin.hpp>
#include <nlohmann/detail/meta/call_std/end.hpp>
#include <nlohmann/detail/meta/cpp_future.hpp>
@@ -450,6 +454,51 @@ struct is_constructible_string_type
value_type_t, laundered_type >>::value;
};
// Forward declarations: iteration_proxy.hpp includes this file, so we cannot
// include it here.
template<typename IteratorType> class iteration_proxy;
template<typename IteratorType> class iteration_proxy_value;
// Identifies nlohmann's internal iteration-proxy types. These must be excluded
// before evaluating any std::ranges concept to avoid circular constraints.
template<typename T> struct is_iteration_proxy_type : std::false_type {};
template<typename T> struct is_iteration_proxy_type<iteration_proxy<T>> : std::true_type {};
template<typename T> struct is_iteration_proxy_type<iteration_proxy_value<T>> : std::true_type {};
// In C++26, std::optional satisfies std::ranges::view; exclude it so the
// range-view overload does not hijack the optional serializer.
#ifdef JSON_HAS_CPP_17
template<typename T> struct is_range_view_optional_type : std::false_type {};
template<typename T> struct is_range_view_optional_type<std::optional<T>> : std::true_type {};
#else
template<typename T> struct is_range_view_optional_type : std::false_type {};
#endif
// std::ranges does not work properly on MinGW due to incomplete C++20 support
// see https://github.com/nlohmann/json/issues/4916
#if JSON_HAS_RANGES && !defined(__MINGW32__)
// SafeToCheck guards against types that trigger circular constraints when
// std::ranges::view<T> is evaluated on GCC 12 / libstdc++ 12:
// - iteration_proxy / iteration_proxy_value directly
// - views wrapping the above (e.g. owning_view<iteration_proxy<...>>)
// - views wrapping basic_json (e.g. ref_view<json>) — same circularity
// via json's constructors → is_compatible_array_type → here
// nlohmann's plain range_value_t (iterator_traits-based) is safe to call
// before any std::ranges concept is touched, so we use it for the checks.
template < typename T, bool SafeToCheck =
!is_iteration_proxy_type<T>::value &&
!is_iteration_proxy_type<detected_t<range_value_t, T>>::value &&
!is_basic_json<detected_t<range_value_t, T>>::value &&
!is_range_view_optional_type<T>::value >
struct is_compatible_range_view : std::false_type {};
template<typename T>
struct is_compatible_range_view<T, true>
: std::bool_constant<std::ranges::view<T>> {};
#endif
template<typename BasicJsonType, typename CompatibleArrayType, typename = void>
struct is_compatible_array_type_impl : std::false_type {};
@@ -461,13 +510,38 @@ struct is_compatible_array_type_impl <
is_iterator_traits<iterator_traits<detected_t<iterator_t, CompatibleArrayType>>>::value&&
// special case for types like std::filesystem::path whose iterator's value_type are themselves
// c.f. https://github.com/nlohmann/json/pull/3073
!std::is_same<CompatibleArrayType, detected_t<range_value_t, CompatibleArrayType>>::value >>
!std::is_same<CompatibleArrayType, detected_t<range_value_t, CompatibleArrayType>>::value
// When range-view support is enabled, std::ranges::view types (e.g. std::string_view,
// filter_view) can match BOTH this iterator-based specialization AND the view-based one
// below, causing ambiguity. Exclude views here so the two specializations are mutually
// exclusive: this one handles plain iterable containers, the other handles views.
#if JSON_HAS_RANGES && !defined(__MINGW32__)
&& !is_compatible_range_view<CompatibleArrayType>::value
#endif
>>
{
static constexpr bool value =
is_constructible<BasicJsonType,
range_value_t<CompatibleArrayType>>::value;
};
#if JSON_HAS_RANGES && !defined(__MINGW32__)
template<typename BasicJsonType, typename CompatibleArrayType>
struct is_compatible_array_type_impl <
BasicJsonType, CompatibleArrayType,
enable_if_t < is_compatible_range_view<CompatibleArrayType>::value
&& !std::is_same<detected_t<range_value_t, CompatibleArrayType>, char>::value
&& !std::is_same<detected_t<range_value_t, CompatibleArrayType>, wchar_t>::value >>
{
// CompatibleArrayType is a std::ranges::view here, so std::ranges::range_value_t
// is safe and correctly handles C++20 iterators that may lack classic iterator_traits.
static constexpr bool value =
is_constructible<BasicJsonType,
std::ranges::range_value_t<CompatibleArrayType>>::value;
};
#endif
template<typename BasicJsonType, typename CompatibleArrayType>
struct is_compatible_array_type
: is_compatible_array_type_impl<BasicJsonType, CompatibleArrayType> {};
@@ -559,6 +633,14 @@ template<typename BasicJsonType, typename CompatibleType>
struct is_compatible_type
: is_compatible_type_impl<BasicJsonType, CompatibleType> {};
template<typename BasicJsonType, typename CompatibleArrayType>
struct is_compatible_binary_type
{
static constexpr bool value =
std::is_same<typename BasicJsonType::binary_t::container_type, CompatibleArrayType>::value &&
!std::is_same<typename BasicJsonType::binary_t::container_type, std::vector<std::uint8_t>>::value;
};
template<typename BasicJsonType, typename CompatibleReferenceType>
struct is_compatible_reference_type_impl
{
+240 -16
View File
@@ -189,6 +189,7 @@
#include <unordered_map> // unordered_map
#include <utility> // pair, declval
#include <valarray> // valarray
#include <vector> // vector
// #include <nlohmann/detail/exceptions.hpp>
// __ _____ _____ _____
@@ -3592,6 +3593,7 @@ NLOHMANN_JSON_NAMESPACE_END
#include <tuple> // tuple
#include <type_traits> // false_type, is_constructible, is_integral, is_same, true_type
#include <utility> // declval
#include <vector> // vector
#if defined(__cpp_lib_byte) && __cpp_lib_byte >= 201603L
#include <cstddef> // byte
#endif
@@ -3663,6 +3665,9 @@ NLOHMANN_JSON_NAMESPACE_END
// #include <nlohmann/detail/macro_scope.hpp>
#ifdef JSON_HAS_CPP_17
#include <optional> // optional
#endif
// #include <nlohmann/detail/meta/call_std/begin.hpp>
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
@@ -4212,6 +4217,51 @@ struct is_constructible_string_type
value_type_t, laundered_type >>::value;
};
// Forward declarations: iteration_proxy.hpp includes this file, so we cannot
// include it here.
template<typename IteratorType> class iteration_proxy;
template<typename IteratorType> class iteration_proxy_value;
// Identifies nlohmann's internal iteration-proxy types. These must be excluded
// before evaluating any std::ranges concept to avoid circular constraints.
template<typename T> struct is_iteration_proxy_type : std::false_type {};
template<typename T> struct is_iteration_proxy_type<iteration_proxy<T>> : std::true_type {};
template<typename T> struct is_iteration_proxy_type<iteration_proxy_value<T>> : std::true_type {};
// In C++26, std::optional satisfies std::ranges::view; exclude it so the
// range-view overload does not hijack the optional serializer.
#ifdef JSON_HAS_CPP_17
template<typename T> struct is_range_view_optional_type : std::false_type {};
template<typename T> struct is_range_view_optional_type<std::optional<T>> : std::true_type {};
#else
template<typename T> struct is_range_view_optional_type : std::false_type {};
#endif
// std::ranges does not work properly on MinGW due to incomplete C++20 support
// see https://github.com/nlohmann/json/issues/4916
#if JSON_HAS_RANGES && !defined(__MINGW32__)
// SafeToCheck guards against types that trigger circular constraints when
// std::ranges::view<T> is evaluated on GCC 12 / libstdc++ 12:
// - iteration_proxy / iteration_proxy_value directly
// - views wrapping the above (e.g. owning_view<iteration_proxy<...>>)
// - views wrapping basic_json (e.g. ref_view<json>) — same circularity
// via json's constructors → is_compatible_array_type → here
// nlohmann's plain range_value_t (iterator_traits-based) is safe to call
// before any std::ranges concept is touched, so we use it for the checks.
template < typename T, bool SafeToCheck =
!is_iteration_proxy_type<T>::value &&
!is_iteration_proxy_type<detected_t<range_value_t, T>>::value &&
!is_basic_json<detected_t<range_value_t, T>>::value &&
!is_range_view_optional_type<T>::value >
struct is_compatible_range_view : std::false_type {};
template<typename T>
struct is_compatible_range_view<T, true>
: std::bool_constant<std::ranges::view<T>> {};
#endif
template<typename BasicJsonType, typename CompatibleArrayType, typename = void>
struct is_compatible_array_type_impl : std::false_type {};
@@ -4223,13 +4273,38 @@ struct is_compatible_array_type_impl <
is_iterator_traits<iterator_traits<detected_t<iterator_t, CompatibleArrayType>>>::value&&
// special case for types like std::filesystem::path whose iterator's value_type are themselves
// c.f. https://github.com/nlohmann/json/pull/3073
!std::is_same<CompatibleArrayType, detected_t<range_value_t, CompatibleArrayType>>::value >>
!std::is_same<CompatibleArrayType, detected_t<range_value_t, CompatibleArrayType>>::value
// When range-view support is enabled, std::ranges::view types (e.g. std::string_view,
// filter_view) can match BOTH this iterator-based specialization AND the view-based one
// below, causing ambiguity. Exclude views here so the two specializations are mutually
// exclusive: this one handles plain iterable containers, the other handles views.
#if JSON_HAS_RANGES && !defined(__MINGW32__)
&& !is_compatible_range_view<CompatibleArrayType>::value
#endif
>>
{
static constexpr bool value =
is_constructible<BasicJsonType,
range_value_t<CompatibleArrayType>>::value;
};
#if JSON_HAS_RANGES && !defined(__MINGW32__)
template<typename BasicJsonType, typename CompatibleArrayType>
struct is_compatible_array_type_impl <
BasicJsonType, CompatibleArrayType,
enable_if_t < is_compatible_range_view<CompatibleArrayType>::value
&& !std::is_same<detected_t<range_value_t, CompatibleArrayType>, char>::value
&& !std::is_same<detected_t<range_value_t, CompatibleArrayType>, wchar_t>::value >>
{
// CompatibleArrayType is a std::ranges::view here, so std::ranges::range_value_t
// is safe and correctly handles C++20 iterators that may lack classic iterator_traits.
static constexpr bool value =
is_constructible<BasicJsonType,
std::ranges::range_value_t<CompatibleArrayType>>::value;
};
#endif
template<typename BasicJsonType, typename CompatibleArrayType>
struct is_compatible_array_type
: is_compatible_array_type_impl<BasicJsonType, CompatibleArrayType> {};
@@ -4321,6 +4396,14 @@ template<typename BasicJsonType, typename CompatibleType>
struct is_compatible_type
: is_compatible_type_impl<BasicJsonType, CompatibleType> {};
template<typename BasicJsonType, typename CompatibleArrayType>
struct is_compatible_binary_type
{
static constexpr bool value =
std::is_same<typename BasicJsonType::binary_t::container_type, CompatibleArrayType>::value &&
!std::is_same<typename BasicJsonType::binary_t::container_type, std::vector<std::uint8_t>>::value;
};
template<typename BasicJsonType, typename CompatibleReferenceType>
struct is_compatible_reference_type_impl
{
@@ -5464,6 +5547,7 @@ template < typename BasicJsonType, typename ConstructibleArrayType,
!is_constructible_object_type<BasicJsonType, ConstructibleArrayType>::value&&
!is_constructible_string_type<BasicJsonType, ConstructibleArrayType>::value&&
!std::is_same<ConstructibleArrayType, typename BasicJsonType::binary_t>::value&&
!is_compatible_binary_type<BasicJsonType, ConstructibleArrayType>::value&&
!is_basic_json<ConstructibleArrayType>::value,
int > = 0 >
auto from_json(const BasicJsonType& j, ConstructibleArrayType& arr)
@@ -5509,6 +5593,25 @@ inline void from_json(const BasicJsonType& j, typename BasicJsonType::binary_t&
bin = *j.template get_ptr<const typename BasicJsonType::binary_t*>();
}
template < typename BasicJsonType, typename CompatibleArrayType,
enable_if_t < is_compatible_binary_type<BasicJsonType, CompatibleArrayType>::value,
int > = 0 >
inline void from_json(const BasicJsonType& j, CompatibleArrayType& bin)
{
if (j.is_binary())
{
bin = static_cast<CompatibleArrayType>(*j.template get_ptr<const typename BasicJsonType::binary_t*>());
}
else if (j.is_array())
{
from_json_array_impl(j, bin, priority_tag<3> {});
}
else
{
JSON_THROW(type_error::create(302, concat("type must be binary or array, but is ", j.type_name()), &j));
}
}
template<typename BasicJsonType, typename ConstructibleObjectType,
enable_if_t<is_constructible_object_type<BasicJsonType, ConstructibleObjectType>::value, int> = 0>
inline void from_json(const BasicJsonType& j, ConstructibleObjectType& obj)
@@ -6205,8 +6308,11 @@ struct external_constructor<value_t::array>
}
template < typename BasicJsonType, typename CompatibleArrayType,
enable_if_t < !std::is_same<CompatibleArrayType, typename BasicJsonType::array_t>::value,
int > = 0 >
enable_if_t < !std::is_same<CompatibleArrayType, typename BasicJsonType::array_t>::value
#if JSON_HAS_RANGES && !defined(__MINGW32__)
&& !is_compatible_range_view<CompatibleArrayType>::value
#endif
, int > = 0 >
static void construct(BasicJsonType& j, const CompatibleArrayType& arr)
{
using std::begin;
@@ -6246,6 +6352,25 @@ struct external_constructor<value_t::array>
j.set_parents();
j.assert_invariant();
}
// std::ranges does not work properly on MinGW due to incomplete C++20 support
// see https://github.com/nlohmann/json/issues/4916
#if JSON_HAS_RANGES && !defined(__MINGW32__)
template<typename BasicJsonType, typename CompatibleArrayType,
enable_if_t<is_compatible_range_view<std::remove_cvref_t<CompatibleArrayType>>::value, int> = 0>
static void construct(BasicJsonType& j, CompatibleArrayType && arr)
{
j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array;
j.m_data.m_value = value_t::array;
for (auto&& x : std::forward<CompatibleArrayType>(arr))
{
j.m_data.m_value.array->push_back(x);
j.set_parent(j.m_data.m_value.array->back());
}
j.assert_invariant();
}
#endif
};
template<>
@@ -6383,19 +6508,44 @@ template < typename BasicJsonType, typename CompatibleArrayType,
!is_compatible_object_type<BasicJsonType, CompatibleArrayType>::value&&
!is_compatible_string_type<BasicJsonType, CompatibleArrayType>::value&&
!std::is_same<typename BasicJsonType::binary_t, CompatibleArrayType>::value&&
!is_basic_json<CompatibleArrayType>::value,
!is_compatible_binary_type<BasicJsonType, CompatibleArrayType>::value&&
!is_basic_json<CompatibleArrayType>::value
#if JSON_HAS_RANGES && !defined(__MINGW32__)
&& !is_compatible_range_view<CompatibleArrayType>::value
#endif
,
int > = 0 >
inline void to_json(BasicJsonType& j, const CompatibleArrayType& arr)
{
external_constructor<value_t::array>::construct(j, arr);
}
#if JSON_HAS_RANGES && !defined(__MINGW32__)
template < typename BasicJsonType, typename T,
enable_if_t < is_compatible_range_view<std::remove_cvref_t<T>>::value
&& !is_compatible_string_type<BasicJsonType, std::remove_cvref_t<T>>::value
&& !is_compatible_object_type<BasicJsonType, std::remove_cvref_t<T>>::value
&& !is_basic_json<std::remove_cvref_t<T>>::value, int > = 0 >
inline void to_json(BasicJsonType& j, T && arr)
{
external_constructor<value_t::array>::construct(j, std::forward<T>(arr));
}
#endif
template<typename BasicJsonType>
inline void to_json(BasicJsonType& j, const typename BasicJsonType::binary_t& bin)
{
external_constructor<value_t::binary>::construct(j, bin);
}
template < typename BasicJsonType, typename CompatibleArrayType,
enable_if_t < is_compatible_binary_type<BasicJsonType, CompatibleArrayType>::value,
int > = 0 >
inline void to_json(BasicJsonType& j, const CompatibleArrayType& bin)
{
external_constructor<value_t::binary>::construct(j, typename BasicJsonType::binary_t(bin));
}
template<typename BasicJsonType, typename T,
enable_if_t<std::is_convertible<T, BasicJsonType>::value, int> = 0>
inline void to_json(BasicJsonType& j, const std::valarray<T>& arr)
@@ -7232,17 +7382,30 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
}
else
{
if (JSON_HEDLEY_UNLIKELY(!input.empty()))
// A supplementary code point is a high surrogate (0xD800..0xDBFF)
// followed by a low surrogate (0xDC00..0xDFFF). A lone low
// surrogate, a high surrogate at the end of the input, or a high
// surrogate followed by any other unit is malformed UTF-16. In
// that case the offending unit is passed through unchanged so the
// UTF-8 decoder rejects it, matching how \uXXXX surrogate escapes
// are handled in the lexer.
bool valid_pair = false;
if (wc <= 0xDBFF && JSON_HEDLEY_UNLIKELY(!input.empty()))
{
const auto wc2 = static_cast<unsigned int>(input.get_character());
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu));
utf8_bytes_filled = 4;
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
{
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu));
utf8_bytes_filled = 4;
valid_pair = true;
}
}
else
if (!valid_pair)
{
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
utf8_bytes_filled = 1;
@@ -10549,12 +10712,39 @@ class binary_reader
// BSON //
//////////
/*!
@brief Validate a BSON document's declared size against the bytes read.
A BSON document starts with an int32 that counts its own total length in
bytes, including that prefix and the trailing 0x00. The reader is driven
by the terminator rather than the declared length, so without this check a
nested document could declare a length that disagrees with where its
terminator actually falls and quietly hand the bytes in between to the
enclosing document. A well-formed document is at least 5 bytes (the prefix
plus the terminator); the equality also rejects those impossible sizes,
since at least 5 bytes are always consumed.
@param[in] document_start value of chars_read before the size prefix
@param[in] document_size the declared document size
@return whether the declared size matches the number of bytes read
*/
bool check_bson_document_size(const std::size_t document_start, const std::int32_t document_size)
{
if (JSON_HEDLEY_UNLIKELY(document_size < 0 || static_cast<std::size_t>(document_size) != chars_read - document_start))
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(112, chars_read,
exception_message(input_format_t::bson, concat("document size ", std::to_string(document_size), " does not match the number of bytes read (", std::to_string(chars_read - document_start), ")"), "document"), nullptr));
}
return true;
}
/*!
@brief Reads in a BSON-object and passes it to the SAX-parser.
@return whether a valid BSON-value was passed to the SAX parser
*/
bool parse_bson_internal()
{
const std::size_t document_start = chars_read;
std::int32_t document_size{};
get_number<std::int32_t, true>(input_format_t::bson, document_size);
@@ -10568,6 +10758,11 @@ class binary_reader
return false;
}
if (JSON_HEDLEY_UNLIKELY(!check_bson_document_size(document_start, document_size)))
{
return false;
}
return sax->end_object();
}
@@ -10783,6 +10978,7 @@ class binary_reader
*/
bool parse_bson_array()
{
const std::size_t document_start = chars_read;
std::int32_t document_size{};
get_number<std::int32_t, true>(input_format_t::bson, document_size);
@@ -10796,6 +10992,11 @@ class binary_reader
return false;
}
if (JSON_HEDLEY_UNLIKELY(!check_bson_document_size(document_start, document_size)))
{
return false;
}
return sax->end_array();
}
@@ -12232,6 +12433,29 @@ class binary_reader
return get_ubjson_value(get_char ? get_ignore_noop() : current);
}
/*!
@brief reject a negative UBJSON/BJData string length
String and key lengths are written with signed integer markers (i, I, l,
L). A negative value is malformed; without this check get_string() would
silently treat it as an empty string and leave the following bytes to be
misread as the next value. This mirrors the non-negative check the
optimized-container count path already performs in get_ubjson_size_value.
@param[in] len the string length read from the input
@return whether the length is valid (non-negative)
*/
template<typename NumberType>
bool check_ubjson_string_length(const NumberType len)
{
if (JSON_HEDLEY_UNLIKELY(len < 0))
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(113, chars_read,
exception_message(input_format, "string length must not be negative", "string"), nullptr));
}
return true;
}
/*!
@brief reads a UBJSON string
@@ -12269,25 +12493,25 @@ class binary_reader
case 'i':
{
std::int8_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'I':
{
std::int16_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'l':
{
std::int32_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'L':
{
std::int64_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'u':
+13
View File
@@ -2721,6 +2721,19 @@ TEST_CASE("BJData")
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(v), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing BJData string: expected length type specification (U, i, u, I, m, l, M, L); last byte: 0x31", json::parse_error&);
}
SECTION("negative length")
{
json _;
std::vector<uint8_t> const vi = {'S', 'i', 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vi), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData string: string length must not be negative", json::parse_error&);
CHECK(json::from_bjdata(vi, true, false).is_discarded());
std::vector<uint8_t> const vl = {'S', 'l', 0xFF, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vl), "[json.exception.parse_error.113] parse error at byte 6: syntax error while parsing BJData string: string length must not be negative", json::parse_error&);
CHECK(json::from_bjdata(vl, true, false).is_discarded());
}
SECTION("parse bjdata markers in ubjson")
{
// create a single-character string for all number types
+56
View File
@@ -854,6 +854,62 @@ TEST_CASE("Unsupported BSON input")
CHECK(!json::sax_parse(bson, &scp, json::input_format_t::bson));
}
TEST_CASE("BSON document size mismatch")
{
json _;
SECTION("top-level document declaring more bytes than it contains")
{
// empty object, but the length prefix claims 6 bytes instead of 5
std::vector<std::uint8_t> const input = {0x06, 0x00, 0x00, 0x00, 0x00};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
SECTION("top-level document with a negative size")
{
std::vector<std::uint8_t> const input = {0xFF, 0xFF, 0xFF, 0xFF, 0x00};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size -1 does not match the number of bytes read (5)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
SECTION("embedded document whose size disagrees with its terminator")
{
// the embedded document "d" declares 0x7FFFFFFF bytes but its 0x00
// terminator falls right after {"a":null}; the length prefix would
// otherwise let the following "h" element be read as a member of the
// enclosing document instead of "d"
std::vector<std::uint8_t> const input =
{
0x00, 0x00, 0x00, 0x00, // outer size
0x03, 'd', 0x00, // entry: embedded document "d"
0xFF, 0xFF, 0xFF, 0x7F, // embedded size 0x7FFFFFFF
0x0A, 'a', 0x00, // entry: null "a"
0x00, // embedded end marker
0x08, 'h', 0x00, 0x01, // entry: bool "h" = true
0x00 // outer end marker
};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON document: document size 2147483647 does not match the number of bytes read (8)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
SECTION("embedded array whose size disagrees with its terminator")
{
// array [42] is 12 bytes, but the length prefix claims 13
std::vector<std::uint8_t> const input =
{
0x00, 0x00, 0x00, 0x00, // outer size
0x04, 'a', 0x00, // entry: array "a"
0x0D, 0x00, 0x00, 0x00, // array size 13 (real is 12)
0x10, '0', 0x00, 0x2A, 0x00, 0x00, 0x00, // entry: int32 "0" = 42
0x00, // array end marker
0x00 // outer end marker
};
CHECK_THROWS_WITH_AS(_ = json::from_bson(input), "[json.exception.parse_error.112] parse error at byte 19: syntax error while parsing BSON document: document size 13 does not match the number of bytes read (12)", json::parse_error&);
CHECK(json::from_bson(input, true, false).is_discarded());
}
}
TEST_CASE("BSON numerical data")
{
SECTION("number")
+15
View File
@@ -1782,6 +1782,21 @@ TEST_CASE("std::optional")
"[json.exception.type_error.302] type must be string, but is null", json::type_error&);
CHECK_THROWS_WITH_AS(std::optional<int>(j_null),
"[json.exception.type_error.302] type must be number, but is null", json::type_error&);
// Assignment goes through the same overload resolution as direct
// construction, so it throws for the same reason. This relies on
// basic_json's implicit conversion operator, so it only applies
// when JSON_USE_IMPLICIT_CONVERSIONS is enabled (the default).
#if JSON_USE_IMPLICIT_CONVERSIONS
std::optional<std::string> opt_assign;
CHECK_THROWS_WITH_AS(opt_assign = j_null,
"[json.exception.type_error.302] type must be string, but is null", json::type_error&);
#endif
// get_to() is the correct way to obtain std::nullopt from a JSON null.
std::optional<std::string> opt_get_to = "placeholder";
j_null.get_to(opt_get_to);
CHECK(opt_get_to == std::nullopt);
}
SECTION("string")
+74
View File
@@ -1136,6 +1136,40 @@ TEST_CASE("regression tests 2")
CHECK((decoded == json_4804::array()));
}
SECTION("discussion #4209 - custom BinaryType direct assignment and round-tripping")
{
// Test that assigning a custom BinaryType directly creates a binary value, not an array
const std::vector<std::byte> original{std::byte{1}, std::byte{2}, std::byte{3}};
const json_4804 j = original;
CHECK(j.is_binary());
CHECK(!j.is_array());
// Test round-tripping: extracting the binary value back as the custom container type
const auto extracted = j.get<std::vector<std::byte>>();
CHECK(extracted == original);
// Test that the default json alias behavior is unchanged: std::vector<uint8_t> -> array
const json default_json = std::vector<std::uint8_t> {1, 2, 3};
CHECK(default_json.is_array());
CHECK(!default_json.is_binary());
}
SECTION("discussion #4209 - custom BinaryType extraction from parsed array")
{
// Test that extracting a custom BinaryType from a parsed JSON array still works
// (not just from a binary-typed node)
const auto j = json_4804::parse("[1,2,3]");
CHECK(j.is_array());
CHECK(!j.is_binary());
// Extracting as custom BinaryType should work from arrays
const auto extracted = j.get<std::vector<std::byte>>();
CHECK(extracted.size() == 3);
CHECK(extracted[0] == std::byte{1});
CHECK(extracted[1] == std::byte{2});
CHECK(extracted[2] == std::byte{3});
}
SECTION("issue #5046 - implicit conversion of return json to std::optional no longer implicit")
{
const json jval{};
@@ -1165,6 +1199,46 @@ TEST_CASE("regression tests 2")
}
#endif
#if JSON_HAS_RANGES && !defined(__MINGW32__)
SECTION("issue #4916 - constructing array from C++20 ranges view does not work")
{
std::vector<int> nums{1, 2, 37, 42, 21};
auto filteredNums = nums | std::views::filter([](int i)
{
return i > 10;
});
json const j(filteredNums);
CHECK(j.type() == json::value_t::array);
CHECK(j == json({37, 42, 21}));
}
#endif
// owning_view is not available in libstdc++ < 12
#if JSON_HAS_RANGES && !defined(__MINGW32__) && !(defined(__GLIBCXX__) && _GLIBCXX_RELEASE < 12)
SECTION("issue #4916 - constructing array from prvalue C++20 ranges view (owning_view)")
{
json const j(std::vector<int> {1, 2, 37, 42, 21} | std::views::filter([](int i)
{
return i > 10;
}));
CHECK(j.type() == json::value_t::array);
CHECK(j == json({37, 42, 21}));
}
#endif
#if JSON_HAS_RANGES && !defined(__MINGW32__)
SECTION("issue #4916 - constructing array from C++20 transform view (prvalue elements)")
{
std::vector<int> nums{1, 2, 3};
auto t = nums | std::views::transform([](int i) noexcept
{
return i * 2;
});
json const j(t);
CHECK(j.type() == json::value_t::array);
CHECK(j == json({2, 4, 6}));
}
#endif
}
TEST_CASE_TEMPLATE("issue #4798 - nlohmann::json::to_msgpack() encode float NaN as double", T, double, float) // NOLINT(readability-math-missing-parentheses, bugprone-throwing-static-initialization)
+25
View File
@@ -1862,6 +1862,31 @@ TEST_CASE("UBJSON")
json _;
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(v), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON string: expected length type specification (U, i, I, l, L); last byte: 0x31", json::parse_error&);
}
SECTION("negative length")
{
json _;
std::vector<uint8_t> const vi = {'S', 'i', 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vi), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vi, true, false).is_discarded());
std::vector<uint8_t> const vI = {'S', 'I', 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vI), "[json.exception.parse_error.113] parse error at byte 4: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vI, true, false).is_discarded());
std::vector<uint8_t> const vl = {'S', 'l', 0xFF, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vl), "[json.exception.parse_error.113] parse error at byte 6: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vl, true, false).is_discarded());
std::vector<uint8_t> const vL = {'S', 'L', 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vL), "[json.exception.parse_error.113] parse error at byte 10: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vL, true, false).is_discarded());
// a length of zero remains valid and yields an empty string
std::vector<uint8_t> const v0 = {'S', 'i', 0};
CHECK(json::from_ubjson(v0) == json(""));
}
}
SECTION("array")
+4 -3
View File
@@ -216,15 +216,16 @@ TEST_CASE("Parse with heterogeneous iterator and sentinel types")
// JSON_HAS_CPP_20 (do not remove; see note at top of file)
TEST_CASE("Parse with std::counted_iterator and std::default_sentinel_t")
{
using iterator_type = std::string::const_iterator;
const std::string json_str = R"({"key":"value","array":[1,2,3]})";
const auto len = static_cast<std::iter_difference_t<std::string::const_iterator>>(json_str.size());
const auto len = static_cast<std::iter_difference_t<iterator_type>>(json_str.size());
std::counted_iterator first(json_str.begin(), len);
const std::counted_iterator<iterator_type> first(json_str.begin(), len);
const json j = json::parse(first, std::default_sentinel);
CHECK(j["key"] == "value");
CHECK(j["array"].size() == 3);
std::counted_iterator first2(json_str.begin(), len);
const std::counted_iterator<iterator_type> first2(json_str.begin(), len);
CHECK(json::accept(first2, std::default_sentinel));
}
#endif
+33 -1
View File
@@ -53,6 +53,27 @@ TEST_CASE("wide strings")
std::wstring const w = L"\"\xDBFF";
json _;
CHECK_THROWS_AS(_ = json::parse(w), json::parse_error&);
// the exact message depends on the width of wchar_t: a 16-bit
// wchar_t passes the lone surrogate to the UTF-8 decoder unchanged
// (rejected as a single ill-formed byte at column 2), while a
// 32-bit wchar_t first encodes it as an ill-formed three-byte
// sequence (rejected one byte later, at column 3)
const char* const error_low_surrogate = sizeof(wchar_t) == 2
? "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'"
: "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xED\xB0'";
const char* const error_high_surrogate = sizeof(wchar_t) == 2
? "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'"
: "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xED\xA0'";
// a lone low surrogate cannot start a pair
CHECK_THROWS_WITH_AS(_ = json::parse(std::wstring{L'"', static_cast<wchar_t>(0xDC00), L'"'}), error_low_surrogate, json::parse_error&);
// a high surrogate followed by a non-low-surrogate unit is invalid
CHECK_THROWS_WITH_AS(_ = json::parse(std::wstring{L'"', static_cast<wchar_t>(0xD800), L'a', L'"'}), error_high_surrogate, json::parse_error&);
// a lone low surrogate must not swallow the following unit: pairing
// it with any second unit would produce valid UTF-8, so the error
// has to report an ill-formed byte at the surrogate's own position
CHECK_THROWS_WITH_AS(_ = json::parse(std::wstring{L'"', static_cast<wchar_t>(0xDC00), L'a', L'"'}), error_low_surrogate, json::parse_error&);
}
}
@@ -68,11 +89,22 @@ TEST_CASE("wide strings")
SECTION("invalid std::u16string")
{
if (wstring_is_utf16())
if (u16string_is_utf16())
{
std::u16string const w = u"\"\xDBFF";
json _;
CHECK_THROWS_AS(_ = json::parse(w), json::parse_error&);
// a lone low surrogate cannot start a pair
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xDC00, u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a high surrogate followed by a non-low-surrogate unit is invalid
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xD800, u'a', u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a lone low surrogate must not swallow the following unit: pairing
// it with any second unit would produce valid UTF-8, so the error
// has to report an ill-formed byte at the surrogate's own position
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xDC00, u'a', u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a valid surrogate pair is still decoded (U+1F600)
CHECK(json::parse(std::u16string{u'"', 0xD83D, 0xDE00, u'"'}).get<std::string>() == "\xF0\x9F\x98\x80");
}
}