scann-core
Copyright 2026 Elias Benali (@ebenali) and TheCleaners.

scann-core is a derived work of ScaNN (Scalable Nearest Neighbors). It is
not an official Google product and is not affiliated with, sponsored by or
endorsed by Google. "ScaNN" is used only to identify the software this work
is derived from.

This product includes software from ScaNN, Copyright The Google Research
Authors, licensed under the Apache License, Version 2.0 (see LICENSE).
scann-core's own additions and modifications are Copyright Elias Benali
(@ebenali) and TheCleaners and are licensed under the same Apache License,
Version 2.0.

  Upstream:     https://github.com/google-research/google-research/tree/master/scann
  Extracted at: google-research commit 758b894eb02dc2a7097068031089a2803be147c6
                (2026-09-21), subdirectory scann/

The following files are copies of upstream files at that commit:
  src/**                                  (from scann/scann/..., same paths)
  python/scann/scann_ops/py/*.py          (from scann/scann/scann_ops/py/)
They are unmodified except for these changes. Each modified file says so in
its license header (or, for .inc fragments, which have none, in a first-line
comment), and each change is marked with a "scann-core:" comment where it
isn't a one-line change:
  src/scann/scann_ops/cc/scann.h
      ParseTextProto reports text-format parse errors (upstream ignored
      them, so misspelled or malformed configs were accepted); the
      ReshapeNNResult / ReshapeBatchedNNResult templates are const;
      default_num_neighbors() and NormalizeDatapoints() accessors (for the
      scann_npy.cc changes below).
      Declares SerializeToDirectory() and LoadArtifactsFromMemory() (see
      scann.cc).
  src/scann/scann_ops/cc/scann.cc
      Initialize() validates dataset shapes (n_points == 0 divided by zero;
      a size that isn't a multiple of n_points silently produced a wrong
      dimensionality); RetrainAndReindex() unlocks the mutex it hands to
      RetrainAndReindexSearcher (upstream destroyed it locked);
      SearchBatchedParallel() rejects batch_size < 1 (division by zero) and
      works without a thread pool (null dereference after SetNumThreads(0)
      or on single-CPU machines); query dimension errors state both
      dimensionalities; a partitioning config with max_spill_centers = 0 is
      rejected (the searcher built, then every search failed with a bare
      RET_CHECK error); Initialize() rejects NaN/infinity in the dataset
      (training a partitioner on it aborted on a QCHECK); batched search
      rejects NaN/infinity queries (upstream checked single queries only);
      leaves_to_search attaches tree parameters only for partitioned
      configs (the int8 brute-force searcher misread them: segfault);
      with spherical partitioning, Initialize() normalizes the dataset's
      rows (upstream only tagged it unit-norm, while later upserts were
      normalized in some configurations: the same vector scored differently
      depending on when it was added), and NormalizeDatapoints() does the
      same for upserts; SearchBatchedParallel() reports a query
      dimensionality mismatch like SearchBatched() (upstream: a bare
      RET_CHECK failure).
      LoadArtifacts() validates an index directory instead of trusting it
      (upstream segfaulted, died of SIGFPE or LOG(FATAL), threw exceptions
      out of Status-returning functions, or loaded mixed files silently):
      it checks ranks and duplicates of assets, assets the config can't
      use, row counts and dimensionalities across all assets, the
      tokenization's length and token range (a bounds-checked loop instead
      of vector::at, which threw), the partitioner's leaf count and center
      dimensionality, the int8 multipliers, hashed codes against the AH
      codebook, and SOAR assets (a null docid collection was passed to
      DenseDataset). SerializeToDirectory() is new: it stages the files,
      then commits them behind an "incomplete" marker manifest that
      LoadArtifacts() rejects, so an interrupted re-serialize can't leave a
      directory that loads a mix of two indexes (upstream's in-place
      Serialize() wrote scann_config.pb first). Without relative_path it
      records absolute asset paths even for a relative directory (upstream
      recorded dir/name, which loaded as dir/dir/name).
      LoadArtifactsFromMemory() is new: LoadArtifacts() from an index
      directory's files held in memory (used by the TensorFlow op in
      tf_op/), with the same validation.
  src/scann/utils/io_npy.h, src/scann/utils/io_npy.cc
      NumpyToVectorAndShape() validates the .npy file (magic, version,
      header length, header dict, dtype kind, size and byte order, C order,
      shape overflow, data size against the file size) with its own parser
      instead of cnpy::parse_npy_header, which read past its buffer for a
      bad header length, threw std::out_of_range for dimensions above
      INT_MAX, and ignored the dtype's kind and byte order. The header
      reader works on any seekable stream, and NumpyBytesToVectorAndShape()
      reads a .npy file's contents from memory.
  src/scann/utils/io_oss_wrapper.cc
      File writes are flushed and checked, so a failed write (e.g. a full
      disk) is reported instead of being lost in the stream's destructor.
  src/scann/tree_x_hybrid/tree_x_hybrid_smmd.cc,
  src/scann/tree_x_hybrid/internal/utils.cc
      A tree with every datapoint deleted reports empty datasets of the
      right dimensionality (float, int8, bfloat16, hashed), so it
      serializes to a directory that loads (upstream wrote none: "dataset,
      hashed_dataset, ... are all null"); the bfloat16 datasets of
      bfloat16 brute-force leaves are merged and serialized (upstream wrote
      no data for such trees).
  src/scann/brute_force/scalar_quantized_brute_force.cc
      Searcher-specific parameters are dynamic_cast; parameters of another
      type are ignored (upstream down_cast them unconditionally).
  src/scann/brute_force/bfloat16_brute_force_mutator.cc
      Mutator::Create() gives a searcher without docids an empty docid
      collection, as the int8 mutator does. The bfloat16 leaves of a tree had
      none, so every add, update or delete failed a RET_CHECK after partially
      changing the index.
  src/scann/partitioning/kmeans_tree_partitioner.cc
      ResidualizeToFloat() and OrthogonalityAmplifiedTokenForDatapointBatched()
      bounds-check partition tokens (upstream read out of bounds for a
      NaN/infinity datapoint's token -1).
  src/scann/utils/single_machine_retraining.cc
      RetrainAndReindexSearcher() builds the new searcher from the
      reconstructed dataset and no longer modifies the old one. Upstream
      replaced the live searcher's dataset and docids before building, so
      when retraining failed (e.g. fewer points than leaves) the kept
      searcher's mutator used a freed docid collection (heap-use-after-free).
      Retraining into spherical partitioning normalizes and tags the
      reconstructed dataset when it isn't tagged unit-norm (upstream failed
      with "Input vectors must be unit L2-norm" for most such indexes); the
      error for a searcher without float data says what it means.
  src/scann/utils/single_machine_retraining.h
      Declares the spherical-partitioning helpers used by the above and by
      scann.cc.
  src/scann/base/health_stats_collector.h
      Health statistics no longer compare unprojected datapoints with
      projected (PCA/TRUNCATE) centroids (an out-of-bounds read on every
      mutation of such a tree); for projected trees the quantization error
      isn't tracked, as upstream's Initialize() already did.
  src/scann/hashes/internal/lut16_avx2.inc
      The "smart" prefetch of the next partition is skipped when there is
      none (the last partition): upstream did pointer arithmetic on the null
      next_partition pointer (undefined behavior, flagged by UBSan).
  src/scann/projection/projection_factory.cc
      PCA and TRUNCATE projections reject a projected dimensionality outside
      1..input_dim (PCA aborted the process through a CHECK failure,
      TRUNCATE failed with "vector::_M_range_insert").
  src/scann/scann_ops/cc/scann_npy.cc
      Upsert() rejects batch_size < 1 (division by zero) and wrong-sized
      vectors (upstream failed with an uninformative RET_CHECK); Delete()
      re-attaches the mutation thread pool after a retrain, like Upsert();
      Upsert() validates every row (dimensionality, NaN/infinity, update
      index range) before mutating, so a bad row neither crashes a tree
      index nor leaves a batch half-applied. The constructor and Upsert()
      reject more than 2^32 - 1 datapoints (upstream truncated the row count
      to 32 bits); Upsert() normalizes vectors for spherical partitioning;
      SearchBatched() with zero queries returns empty (0, k) results
      (upstream failed with a misleading dimensionality error).
  src/scann/scann_ops/cc/scann_npy.h, src/scann/scann_ops/cc/scann_npy.cc
      ScannNumpy holds a reader/writer lock: searches and size() shared,
      everything else exclusive, always with the GIL released. Upstream
      relied on the GIL, but searches release it, so a search could run
      concurrently with an upsert or delete; without the GIL, concurrent
      upserts corrupted the index. Serialize() goes through
      SerializeToDirectory() and takes the pickled docids, which are
      committed with the index.
  src/scann/scann_ops/cc/python/scann_pybind.cc
      The module declares mod_gil_not_used(), so free-threaded Python keeps
      the GIL disabled when it is imported. serialize() has named
      arguments and an optional docids_pkl.
  src/scann/utils/memory_logging.cc
      No longer dereferences empty std::optionals.
  src/scann/data_format/docid_collection.h
      MemoryUsage() counts sizeof(*this), not sizeof(this).
  src/scann/tree_x_hybrid/mutator.h
      UpdateDatapoint() and RemoveDatapoint() skip unused (kInvalidToken)
      assignment slots. With spilling (SOAR) upstream passed them to the
      health-stats collector, which indexed its arrays at -1: heap
      corruption when updating or deleting points in a SOAR index.
      Partition tokens from tokenizing a new vector are bounds-checked (a
      NaN/infinity vector tokenizes to -1; upstream indexed
      leaf_mutators_[-1]). RemoveDatapoint() removes from the base only
      after the leaf removals succeed, and AddDatapoint() undoes its base
      and leaf additions when a leaf add fails (upstream left size() out of
      step with the leaves). Incremental training on a projected tree or
      one with an upper tree is a config error with a clear message
      (upstream: a bare RET_CHECK failure).
      UpdateDatapoint() is all-or-nothing: it first
      appends the new vector to each of its leaves (undone if one fails),
      then updates the base and removes the old leaf entries. Upstream
      updated the base and the assignments step by step and returned at the
      first failing leaf, leaving the datapoint half-updated.
  src/scann/utils/bfloat16_helpers.h
      Bfloat16Decompress() shifts as unsigned; left-shifting a negative
      int16 is undefined behavior in C++17.
  src/scann/projection/chunking_projection.cc,
  src/scann/projection/projection_factory.h
      Chunking projections reject num_dims_per_block, num_blocks or
      input_dim below 1 (upstream divided by zero, SIGFPE, or CHECK-failed,
      aborting the process).
  src/scann/utils/hash_leaf_helpers.cc,
  src/scann/tree_x_hybrid/tree_ah_hybrid_residual.cc
      INT8_LUT16 lookups require 16 clusters per block (with fewer, the
      LUT16 kernels read past the lookup table: heap overflow).
      TrainAsymmetricHashingModel() calls TrainingOptions::Validate(), as
      the tree-AH residual factory does; an invalid AH projection used to
      end in a bare "SCANN_RET_CHECK failure" (null projector). In
      tree_ah_hybrid_residual.cc, a failure to create a query's lookup
      table is returned as a Status (upstream called .value(), which
      throws). The LUT16 searches skip empty leaves: when every searched
      leaf was empty (e.g. all points deleted), num_blocks was 0, and the
      AVX2 kernel divided by it (SIGFPE).
  src/scann/base/internal/single_machine_factory_impl.h,
  src/scann/base/internal/single_machine_factory_impl.cc
      The factory rejects binary distance measures (Hamming, ...) on
      non-binary data, bfloat16 brute force with distances other than dot
      product and squared L2 (both LOG(FATAL) upstream), and a brute-force
      fixed_point_multiplier_quantile outside (0, 1] (undefined behavior
      when quantizing tree leaves). For asymmetric hashing without residual
      quantization, an AH projection without input_dim gets the dataset's
      dimensionality, as the tree-AH residual factory already did with the
      centers' (upstream failed to build a tree with a PCA/TRUNCATE
      projection and non-residual AH, e.g. any squared L2 tree from the
      Python builder's pca()/truncate()).
  src/scann/partitioning/tree_brute_force_second_level_wrapper.cc
      An upper tree with FIXED8 or BFLOAT16 scoring rejects query
      tokenization distances its searchers don't support (LOG(FATAL)).
  src/scann/base/reordering_helper_factory.cc
      Fixed-point reordering rejects a NaN multiplier quantile and a missing
      dataset (upstream: undefined behavior, null reference). So does exact
      (float) reordering (upstream: LOG(FATAL) in ExactReorderingHelper's
      constructor, e.g. for an index directory without dataset.npy).
  src/scann/utils/reduction.h
      The dense and sparse-dense accumulation loops compare remaining
      counts instead of advancing past the end of the input; for an empty
      sparse datapoint that was arithmetic on a null pointer (undefined
      behavior, flagged by UBSan).
  GCC portability (upstream built only with clang, which accepts these as
  extensions or never checks them; GCC >= 13 rejects them). None changes
  clang's results (verified: still bit-identical to upstream):
    src/scann/utils/intrinsics/sse4.h,
    src/scann/distance_measures/many_to_many/many_to_many_impl.inc,
    src/scann/distance_measures/one_to_many/one_to_many_asymmetric_impl.inc,
    src/scann/tree_x_hybrid/internal/utils.cc
      Explicit bit casts between integer and float SIMD vectors (same bits)
      instead of implicit vector conversions or static_cast; in sse4.h the
      vectors were also declared with the wrong type.
    src/scann/distance_measures/one_to_many/one_to_many_symmetric.h,
    src/scann/distance_measures/one_to_many/one_to_many_impl_highway.inc,
    src/scann/base/single_machine_base.h
      Declarations before use, namespace qualification, and `typename` where
      C++17 requires them.
    src/scann/utils/common.h, src/scann/data_format/dataset.h
      SCANN_INLINE_FORWARDING: four virtual methods that forward to the same
      method on another view are not force-inlined under GCC (it rejects
      always_inline on what it sees as recursion).
    src/scann/utils/intrinsics/flags.h,
    src/scann/distance_measures/many_to_many/int8_tile.h,
    src/scann/distance_measures/many_to_many/many_to_many_templates.h
      The AMX kernels (clang's tile builtins) are compiled with clang only;
      upstream assumed any non-clang compiler had them.
  python/scann/scann_ops/py/scann_ops_pybind.py
      upsert() updates the docid bookkeeping only after the index accepted
      the vectors, and delete() validates all docids before changing
      anything (upstream left docids out of sync with the index when either
      raised). The docid bookkeeping is guarded by a reader/writer lock, so
      a search overlapping an upsert or delete maps its results to the
      right docids. delete() on a searcher without docids raises the same
      ValueError as upsert() (upstream: AttributeError).
      search_batched()/search_batched_parallel() map result padding (NaN
      distance) to None instead of docids[0]. upsert() rejects a docid
      listed twice (upstream added a new one twice but mapped it once,
      leaving a duplicate in docids). ScannSearcher keeps a copy of the
      docids list it is given (upstream mutated the caller's list).
      Vectors from other array libraries go through _host_array(): tensors
      that require grad are detached, GPU tensors copied to host memory,
      bfloat16 converted to float32 (upstream: a pybind TypeError).
      serialize() passes the pickled docids to the C++ serialize, which
      commits them with the index (upstream wrote scann_docids.pkl
      afterwards, and left a stale one when the index had no docids);
      load_searcher() rejects a scann_docids.pkl whose length doesn't match
      the index.
  python/scann/scann_ops/py/scann_builder.py
      create_config() raises ValueError for tree(incremental_threshold=...)
      together with pca(), truncate() or upper_tree(), which the C++ side
      can't build (upstream failed only when initializing the searcher).
docs/algorithms.md is upstream's scann/docs/algorithms.md, modified (added
explanation of asymmetric hashing and anisotropic quantization, and links).
python/scann/tf.py is adapted from upstream's
scann/scann/scann_ops/py/scann_ops.py (its builder() docstring, and the
builder / create_searcher / ScannSearcher search API): the searcher is the
pybind one, called through tf.numpy_function instead of a TensorFlow op, and
serialize_to_module() / searcher_from_module() raise NotImplementedError.
tf_op/scann_tf_ops.cc is adapted from upstream's
scann/scann/scann_ops/cc/ops/scann_ops.cc and kernels/scann_ops.cc (the
ScannSearch / ScannSearchBatched ops and their search and result-shaping
semantics): rewritten against TensorFlow's C API, with the index passed as
string tensors of its files instead of a resource, a searcher cache keyed
by index id and a fingerprint of those tensors, and batched results exactly
final_num_neighbors wide when it is given. tf_op/python/scann_tf_ops/
__init__.py is adapted from upstream's scann/scann/scann_ops/py/
scann_ops.py (its API and builder() docstring): the searcher is a tf.Module
holding the index files as variables, and serialize_to_module() /
searcher_from_module() use them.
Everything else (build system, the C++ config builder in core/, Rust
bindings, tests, examples, other docs,
python/scann/__init__.py and the empty python/scann/**/__init__.py files)
is new in scann-core. Upstream's own repository has no NOTICE file.

Arm (aarch64) support
---------------------
The Neon/SVE implementations and AArch64 run-time feature detection come
from work by Arm engineers, contributed to ScaNN but not merged upstream as
of 2026-09-26:
  * google-research/google-research PR #3374, "scann: Add AArch64 run-time
    feature detection" (Gerda Zsejke More, gerdazsejke.more@arm.com)
  * lizhang-arm/google-research PR #1, "Arm: Add Neon implementations for
    many to many functions" (Gerda Zsejke More)
  * lizhang-arm/google-research PR #2, "Arm: Add Neon implementation for
    ScaNN indexDatapointNoiseShaped" (Li Zhang, li.zhang2@arm.com)
  * lizhang-arm/google-research PR #3, "scann: Add Arm implementation of
    DenseDotProductInt8Float" (Li Zhang)
Their commits were rebased onto the extraction commit above and kept with
their original authorship. Files they add carry The Google Research
Authors' header, as contributed; files they modify say so in their
headers. On top of them, scann-core:
  - merged the two independently created versions of
    utils/intrinsics/mem_neon.h;
  - moved the <arm_neon.h> include in
    hashes/internal/asymmetric_hashing_impl_neon.cc inside its
    `#if defined(__aarch64__)` guard (x86 builds failed otherwise).
The first commit of PR #3374 also enables x86 CPU feature detection
(PLATFORM_IS_X86), which upstream open-source builds lack; see its commit.

Further aarch64 fixes to upstream files (needed with highway 1.4.0 and a
baseline without the AES extension):
  src/scann/distance_measures/many_to_many/int8_tile.cc
      Includes hwy/foreach_target.h before int8_tile.h, as Highway
      requires; otherwise the N_NEON dispatch target came out empty
      ("use of undeclared identifier 'N_NEON'").
  src/scann/utils/hwy-compact.cc
      Static target HWY_NEON_WITHOUT_AES when AES isn't enabled (the code
      doesn't use AES); HWY_NEON otherwise, as before.

Third-party dependencies are downloaded at configure time (see
cmake/Dependencies.cmake), except two vendored in third_party/ with their
licenses: cnpy (MIT, Copyright (c) Carl Rogers, 2011) and googletest's
gtest_prod.h (BSD 3-Clause, Copyright 2008 Google Inc.). All are distributed
under their own licenses:
  abseil-cpp   Apache License 2.0
  protobuf     BSD 3-Clause
  highway      Apache License 2.0 or BSD 3-Clause (dual-licensed)
  Eigen        MPL 2.0 (some files BSD/Apache/MINPACK; see its COPYING.README)
  cnpy         MIT
  googletest   BSD 3-Clause         (gtest_prod.h only)
  zlib         zlib license
  pybind11     BSD 3-Clause         (Python bindings only)
  cxx          MIT or Apache 2.0    (Rust bindings only, via crates.io)
Binaries that statically link these carry their license obligations.
