scann-core
Copyright 2026 Elias Benali (@ebenali) and TheCleaners.

scann-core is a derived work of ScaNN (Scalable Nearest Neighbors). It is
not an official Google product and is not affiliated with, sponsored by or
endorsed by Google. "ScaNN" is used only to identify the software this work
is derived from.

This product includes software from ScaNN, Copyright The Google Research
Authors, licensed under the Apache License, Version 2.0 (see LICENSE).
scann-core's own additions and modifications are Copyright Elias Benali
(@ebenali) and TheCleaners and are licensed under the same Apache License,
Version 2.0.

  Upstream:     https://github.com/google-research/google-research/tree/master/scann
  Extracted at: google-research commit 758b894eb02dc2a7097068031089a2803be147c6
                (2026-09-21), subdirectory scann/

The following files are copies of upstream files at that commit:
  src/**                                  (from scann/scann/..., same paths)
  python/scann/scann_ops/py/*.py          (from scann/scann/scann_ops/py/)
They are unmodified except for these changes. Each modified file says so in
its license header (or, for .inc fragments, which have none, in a first-line
comment), and each change is marked with a "scann-core:" comment where it
isn't a one-line change:
  src/scann/scann_ops/cc/scann.h
      ParseTextProto reports text-format parse errors (upstream ignored
      them, so misspelled or malformed configs were accepted); the
      ReshapeNNResult / ReshapeBatchedNNResult templates are const.
  src/scann/scann_ops/cc/scann.cc
      Initialize() validates dataset shapes (n_points == 0 divided by zero;
      a size that isn't a multiple of n_points silently produced a wrong
      dimensionality); RetrainAndReindex() unlocks the mutex it hands to
      RetrainAndReindexSearcher (upstream destroyed it locked);
      SearchBatchedParallel() rejects batch_size < 1 (division by zero) and
      works without a thread pool (null dereference after SetNumThreads(0)
      or on single-CPU machines); query dimension errors state both
      dimensionalities; a partitioning config with max_spill_centers = 0 is
      rejected (the searcher built, then every search failed with a bare
      RET_CHECK error); Initialize() rejects NaN/infinity in the dataset
      (training a partitioner on it aborted on a QCHECK); batched search
      rejects NaN/infinity queries (upstream checked single queries only);
      leaves_to_search attaches tree parameters only for partitioned
      configs (the int8 brute-force searcher misread them: segfault).
  src/scann/brute_force/scalar_quantized_brute_force.cc
      Searcher-specific parameters are dynamic_cast; parameters of another
      type are ignored (upstream down_cast them unconditionally).
  src/scann/brute_force/bfloat16_brute_force_mutator.cc
      Mutator::Create() gives a searcher without docids an empty docid
      collection, as the int8 mutator does. The bfloat16 leaves of a tree had
      none, so every add, update or delete failed a RET_CHECK after partially
      changing the index.
  src/scann/partitioning/kmeans_tree_partitioner.cc
      ResidualizeToFloat() and OrthogonalityAmplifiedTokenForDatapointBatched()
      bounds-check partition tokens (upstream read out of bounds for a
      NaN/infinity datapoint's token -1).
  src/scann/utils/single_machine_retraining.cc
      RetrainAndReindexSearcher() builds the new searcher from the
      reconstructed dataset and no longer modifies the old one. Upstream
      replaced the live searcher's dataset and docids before building, so
      when retraining failed (e.g. fewer points than leaves) the kept
      searcher's mutator used a freed docid collection (heap-use-after-free).
  src/scann/base/health_stats_collector.h
      Health statistics no longer compare unprojected datapoints with
      projected (PCA/TRUNCATE) centroids (an out-of-bounds read on every
      mutation of such a tree); for projected trees the quantization error
      isn't tracked, as upstream's Initialize() already did.
  src/scann/hashes/internal/lut16_avx2.inc
      The "smart" prefetch of the next partition is skipped when there is
      none (the last partition): upstream did pointer arithmetic on the null
      next_partition pointer (undefined behavior, flagged by UBSan).
  src/scann/projection/projection_factory.cc
      PCA and TRUNCATE projections reject a projected dimensionality outside
      1..input_dim (PCA aborted the process through a CHECK failure,
      TRUNCATE failed with "vector::_M_range_insert").
  src/scann/scann_ops/cc/scann_npy.cc
      Upsert() rejects batch_size < 1 (division by zero) and wrong-sized
      vectors (upstream failed with an uninformative RET_CHECK); Delete()
      re-attaches the mutation thread pool after a retrain, like Upsert();
      Upsert() validates every row (dimensionality, NaN/infinity, update
      index range) before mutating, so a bad row neither crashes a tree
      index nor leaves a batch half-applied.
  src/scann/scann_ops/cc/scann_npy.h, src/scann/scann_ops/cc/scann_npy.cc
      ScannNumpy holds a reader/writer lock: searches and size() shared,
      everything else exclusive, always with the GIL released. Upstream
      relied on the GIL, but searches release it, so a search could run
      concurrently with an upsert or delete; without the GIL, concurrent
      upserts corrupted the index.
  src/scann/scann_ops/cc/python/scann_pybind.cc
      The module declares mod_gil_not_used(), so free-threaded Python keeps
      the GIL disabled when it is imported.
  src/scann/utils/memory_logging.cc
      No longer dereferences empty std::optionals.
  src/scann/data_format/docid_collection.h
      MemoryUsage() counts sizeof(*this), not sizeof(this).
  src/scann/tree_x_hybrid/mutator.h
      UpdateDatapoint() and RemoveDatapoint() skip unused (kInvalidToken)
      assignment slots. With spilling (SOAR) upstream passed them to the
      health-stats collector, which indexed its arrays at -1: heap
      corruption when updating or deleting points in a SOAR index.
      Partition tokens from tokenizing a new vector are bounds-checked (a
      NaN/infinity vector tokenizes to -1; upstream indexed
      leaf_mutators_[-1]). RemoveDatapoint() removes from the base only
      after the leaf removals succeed, and AddDatapoint() undoes its base
      and leaf additions when a leaf add fails (upstream left size() out of
      step with the leaves).
  src/scann/utils/bfloat16_helpers.h
      Bfloat16Decompress() shifts as unsigned; left-shifting a negative
      int16 is undefined behavior in C++17.
  GCC portability (upstream built only with clang, which accepts these as
  extensions or never checks them; GCC >= 13 rejects them). None changes
  clang's results (verified: still bit-identical to upstream):
    src/scann/utils/intrinsics/sse4.h,
    src/scann/distance_measures/many_to_many/many_to_many_impl.inc,
    src/scann/distance_measures/one_to_many/one_to_many_asymmetric_impl.inc,
    src/scann/tree_x_hybrid/internal/utils.cc
      Explicit bit casts between integer and float SIMD vectors (same bits)
      instead of implicit vector conversions or static_cast; in sse4.h the
      vectors were also declared with the wrong type.
    src/scann/distance_measures/one_to_many/one_to_many_symmetric.h,
    src/scann/distance_measures/one_to_many/one_to_many_impl_highway.inc,
    src/scann/base/single_machine_base.h
      Declarations before use, namespace qualification, and `typename` where
      C++17 requires them.
    src/scann/utils/common.h, src/scann/data_format/dataset.h
      SCANN_INLINE_FORWARDING: four virtual methods that forward to the same
      method on another view are not force-inlined under GCC (it rejects
      always_inline on what it sees as recursion).
    src/scann/utils/intrinsics/flags.h,
    src/scann/distance_measures/many_to_many/int8_tile.h,
    src/scann/distance_measures/many_to_many/many_to_many_templates.h
      The AMX kernels (clang's tile builtins) are compiled with clang only;
      upstream assumed any non-clang compiler had them.
  python/scann/scann_ops/py/scann_ops_pybind.py
      upsert() updates the docid bookkeeping only after the index accepted
      the vectors, and delete() validates all docids before changing
      anything (upstream left docids out of sync with the index when either
      raised). The docid bookkeeping is guarded by a reader/writer lock, so
      a search overlapping an upsert or delete maps its results to the
      right docids. delete() on a searcher without docids raises the same
      ValueError as upsert() (upstream: AttributeError).
      search_batched()/search_batched_parallel() map result padding (NaN
      distance) to None instead of docids[0].
docs/algorithms.md is upstream's scann/docs/algorithms.md, modified (added
explanation of asymmetric hashing and anisotropic quantization, and links).
Everything else (build system, the C++ config builder in core/, Rust
bindings, tests, examples, other docs,
python/scann/__init__.py and the empty python/scann/**/__init__.py files)
is new in scann-core. Upstream's own repository has no NOTICE file.

Arm (aarch64) support
---------------------
The Neon/SVE implementations and AArch64 run-time feature detection come
from work by Arm engineers, contributed to ScaNN but not merged upstream as
of 2026-09-26:
  * google-research/google-research PR #3374, "scann: Add AArch64 run-time
    feature detection" (Gerda Zsejke More, gerdazsejke.more@arm.com)
  * lizhang-arm/google-research PR #1, "Arm: Add Neon implementations for
    many to many functions" (Gerda Zsejke More)
  * lizhang-arm/google-research PR #2, "Arm: Add Neon implementation for
    ScaNN indexDatapointNoiseShaped" (Li Zhang, li.zhang2@arm.com)
  * lizhang-arm/google-research PR #3, "scann: Add Arm implementation of
    DenseDotProductInt8Float" (Li Zhang)
Their commits were rebased onto the extraction commit above and kept with
their original authorship. Files they add carry The Google Research
Authors' header, as contributed; files they modify say so in their
headers. On top of them, scann-core:
  - merged the two independently created versions of
    utils/intrinsics/mem_neon.h;
  - moved the <arm_neon.h> include in
    hashes/internal/asymmetric_hashing_impl_neon.cc inside its
    `#if defined(__aarch64__)` guard (x86 builds failed otherwise).
The first commit of PR #3374 also enables x86 CPU feature detection
(PLATFORM_IS_X86), which upstream open-source builds lack; see its commit.

Further aarch64 fixes to upstream files (needed with highway 1.4.0 and a
baseline without the AES extension):
  src/scann/distance_measures/many_to_many/int8_tile.cc
      Includes hwy/foreach_target.h before int8_tile.h, as Highway
      requires; otherwise the N_NEON dispatch target came out empty
      ("use of undeclared identifier 'N_NEON'").
  src/scann/utils/hwy-compact.cc
      Static target HWY_NEON_WITHOUT_AES when AES isn't enabled (the code
      doesn't use AES); HWY_NEON otherwise, as before.

Third-party dependencies are downloaded at configure time (see
cmake/Dependencies.cmake), except two vendored in third_party/ with their
licenses: cnpy (MIT, Copyright (c) Carl Rogers, 2011) and googletest's
gtest_prod.h (BSD 3-Clause, Copyright 2008 Google Inc.). All are distributed
under their own licenses:
  abseil-cpp   Apache License 2.0
  protobuf     BSD 3-Clause
  highway      Apache License 2.0 or BSD 3-Clause (dual-licensed)
  Eigen        MPL 2.0 (some files BSD/Apache/MINPACK; see its COPYING.README)
  cnpy         MIT
  googletest   BSD 3-Clause         (gtest_prod.h only)
  zlib         zlib license
  pybind11     BSD 3-Clause         (Python bindings only)
  cxx          MIT or Apache 2.0    (Rust bindings only, via crates.io)
Binaries that statically link these carry their license obligations.
