scann-core
Copyright 2026 Elias Benali (@ebenali) and TheCleaners.

scann-core is a derived work of ScaNN (Scalable Nearest Neighbors). It is
not an official Google product and is not affiliated with, sponsored by or
endorsed by Google. "ScaNN" is used only to identify the software this work
is derived from.

This product includes software from ScaNN, Copyright The Google Research
Authors, licensed under the Apache License, Version 2.0 (see LICENSE).
scann-core's own additions and modifications are Copyright Elias Benali
(@ebenali) and TheCleaners and are licensed under the same Apache License,
Version 2.0.

  Upstream:     https://github.com/google-research/google-research/tree/master/scann
  Extracted at: google-research commit 758b894eb02dc2a7097068031089a2803be147c6
                (2026-09-21), subdirectory scann/

The following files are copies of upstream files at that commit:
  src/**                                  (from scann/scann/..., same paths)
  python/scann/scann_ops/py/*.py          (from scann/scann/scann_ops/py/)
They are unmodified except for these changes. Each modified file says so in
its license header (or, for .inc fragments, which have none, in a first-line
comment), and each change is marked with a "scann-core:" comment where it
isn't a one-line change:
  src/scann/scann_ops/cc/scann.h
      ParseTextProto reports text-format parse errors (upstream ignored
      them, so misspelled or malformed configs were accepted); the
      ReshapeNNResult / ReshapeBatchedNNResult templates are const;
      default_num_neighbors() and NormalizeDatapoints() accessors (for the
      scann_npy.cc changes below).
      Declares SerializeToDirectory() and LoadArtifactsFromMemory() (see
      scann.cc).
  src/scann/scann_ops/cc/scann.cc
      Initialize() validates dataset shapes (n_points == 0 divided by zero;
      a size that isn't a multiple of n_points silently produced a wrong
      dimensionality); RetrainAndReindex() unlocks the mutex it hands to
      RetrainAndReindexSearcher (upstream destroyed it locked);
      SearchBatchedParallel() rejects batch_size < 1 (division by zero) and
      works without a thread pool (null dereference after SetNumThreads(0)
      or on single-CPU machines); query dimension errors state both
      dimensionalities; a partitioning config with max_spill_centers = 0 is
      rejected (the searcher built, then every search failed with a bare
      RET_CHECK error); Initialize() rejects NaN/infinity in the dataset
      (training a partitioner on it aborted on a QCHECK); batched search
      rejects NaN/infinity queries (upstream checked single queries only);
      leaves_to_search attaches tree parameters only for partitioned
      configs (the int8 brute-force searcher misread them: segfault);
      with spherical partitioning, Initialize() normalizes the dataset's
      rows (upstream only tagged it unit-norm, while later upserts were
      normalized in some configurations: the same vector scored differently
      depending on when it was added), and NormalizeDatapoints() does the
      same for upserts; SearchBatchedParallel() reports a query
      dimensionality mismatch like SearchBatched() (upstream: a bare
      RET_CHECK failure).
      LoadArtifacts() validates an index directory instead of trusting it
      (upstream segfaulted, died of SIGFPE or LOG(FATAL), threw exceptions
      out of Status-returning functions, or loaded mixed files silently):
      it checks ranks and duplicates of assets, assets the config can't
      use, row counts and dimensionalities across all assets, the
      tokenization's length and token range (a bounds-checked loop instead
      of vector::at, which threw), the partitioner's leaf count and center
      dimensionality, the int8 multipliers, hashed codes against the AH
      codebook, and SOAR assets (a null docid collection was passed to
      DenseDataset). SerializeToDirectory() is new: it stages the files,
      then commits them behind an "incomplete" marker manifest that
      LoadArtifacts() rejects, so an interrupted re-serialize can't leave a
      directory that loads a mix of two indexes (upstream's in-place
      Serialize() wrote scann_config.pb first). Without relative_path it
      records absolute asset paths even for a relative directory (upstream
      recorded dir/name, which loaded as dir/dir/name).
      LoadArtifactsFromMemory() is new: LoadArtifacts() from an index
      directory's files held in memory (used by the TensorFlow op in
      tf_op/), with the same validation.
      RetrainAndReindex() initializes the health statistics before
      releasing the float dataset, as CreateSearcher() does (upstream
      released it first: after a rebalance, trees without float
      reordering reported a quantization error of 0).
  src/scann/utils/io_npy.h, src/scann/utils/io_npy.cc
      NumpyToVectorAndShape() validates the .npy file (magic, version,
      header length, header dict, dtype kind, size and byte order, C order,
      shape overflow, data size against the file size) with its own parser
      instead of cnpy::parse_npy_header, which read past its buffer for a
      bad header length, threw std::out_of_range for dimensions above
      INT_MAX, and ignored the dtype's kind and byte order. The header
      reader works on any seekable stream, and NumpyBytesToVectorAndShape()
      reads a .npy file's contents from memory.
  src/scann/utils/io_oss_wrapper.cc
      File writes are flushed and checked, so a failed write (e.g. a full
      disk) is reported instead of being lost in the stream's destructor.
  src/scann/tree_x_hybrid/tree_x_hybrid_smmd.cc,
  src/scann/tree_x_hybrid/internal/utils.cc
      A tree with every datapoint deleted reports empty datasets of the
      right dimensionality (float, int8, bfloat16, hashed), so it
      serializes to a directory that loads (upstream wrote none: "dataset,
      hashed_dataset, ... are all null"); the bfloat16 datasets of
      bfloat16 brute-force leaves are merged and serialized (upstream wrote
      no data for such trees).
  src/scann/tree_x_hybrid/internal/utils.cc
      DeduplicateDatabaseSpilledResults() merges the duplicates of spilled
      (SOAR) candidates with a stable sort by datapoint index instead of a
      flat_hash_map: the same merged distances, in a deterministic order.
      Upstream emitted them in the hash map's iteration order, which absl
      varies per table, and reordering rounds differently depending on a
      candidate's position, so repeating a search on an unchanged index
      could return distances differing in the last bits.
  src/scann/brute_force/scalar_quantized_brute_force.cc
      Searcher-specific parameters are dynamic_cast; parameters of another
      type are ignored (upstream down_cast them unconditionally).
  src/scann/brute_force/bfloat16_brute_force_mutator.cc
      Mutator::Create() gives a searcher without docids an empty docid
      collection, as the int8 mutator does. The bfloat16 leaves of a tree had
      none, so every add, update or delete failed a RET_CHECK after partially
      changing the index.
  src/scann/partitioning/kmeans_tree_partitioner.cc
      ResidualizeToFloat() and OrthogonalityAmplifiedTokenForDatapointBatched()
      bounds-check partition tokens (upstream read out of bounds for a
      NaN/infinity datapoint's token -1).
  src/scann/utils/single_machine_retraining.cc
      RetrainAndReindexSearcher() builds the new searcher from the
      reconstructed dataset and no longer modifies the old one. Upstream
      replaced the live searcher's dataset and docids before building, so
      when retraining failed (e.g. fewer points than leaves) the kept
      searcher's mutator used a freed docid collection (heap-use-after-free).
      Retraining into spherical partitioning normalizes and tags the
      reconstructed dataset when it isn't tagged unit-norm (upstream failed
      with "Input vectors must be unit L2-norm" for most such indexes); the
      error for a searcher without float data says what it means.
  src/scann/utils/single_machine_retraining.h
      Declares the spherical-partitioning helpers used by the above and by
      scann.cc.
  src/scann/base/health_stats_collector.h
      Health statistics no longer compare unprojected datapoints with
      projected (PCA/TRUNCATE) centroids (an out-of-bounds read on every
      mutation of such a tree). Datapoints are mapped into the centroids'
      space with the partitioner's projection, so projected trees report
      and maintain their quantization error (upstream's Initialize() skipped
      it: 0). When the error can't be computed or maintained because the
      tree has no float dataset (released after the build without float
      reordering), it is reported as NaN; upstream kept the build's sum over
      the changed number of datapoints, or reported 0 after
      InitializeHealthStats().
  src/scann/hashes/internal/lut16_avx2.inc
      The "smart" prefetch of the next partition is skipped when there is
      none (the last partition): upstream did pointer arithmetic on the null
      next_partition pointer (undefined behavior, flagged by UBSan).
  src/scann/projection/projection_factory.cc
      PCA and TRUNCATE projections reject a projected dimensionality outside
      1..input_dim (PCA aborted the process through a CHECK failure,
      TRUNCATE failed with "vector::_M_range_insert").
  src/scann/scann_ops/cc/scann_npy.cc
      Upsert() rejects batch_size < 1 (division by zero) and wrong-sized
      vectors (upstream failed with an uninformative RET_CHECK); Delete()
      re-attaches the mutation thread pool after a retrain, like Upsert();
      Upsert() validates every row (dimensionality, NaN/infinity, update
      index range) before mutating, so a bad row neither crashes a tree
      index nor leaves a batch half-applied. The constructor and Upsert()
      reject more than 2^32 - 1 datapoints (upstream truncated the row count
      to 32 bits); Upsert() normalizes vectors for spherical partitioning;
      SearchBatched() with zero queries returns empty (0, k) results
      (upstream failed with a misleading dimensionality error).
  src/scann/scann_ops/cc/scann_npy.h, src/scann/scann_ops/cc/scann_npy.cc
      ScannNumpy holds a reader/writer lock: searches and size() shared,
      everything else exclusive, always with the GIL released. Upstream
      relied on the GIL, but searches release it, so a search could run
      concurrently with an upsert or delete; without the GIL, concurrent
      upserts corrupted the index. Serialize() goes through
      SerializeToDirectory() and takes the pickled docids, which are
      committed with the index.
  src/scann/scann_ops/cc/python/scann_pybind.cc
      The module declares mod_gil_not_used(), so free-threaded Python keeps
      the GIL disabled when it is imported. serialize() has named
      arguments and an optional docids_pkl.
  src/scann/utils/memory_logging.cc
      No longer dereferences empty std::optionals.
  src/scann/data_format/docid_collection.h
      MemoryUsage() counts sizeof(*this), not sizeof(this).
  src/scann/trees/kmeans_tree/kmeans_tree_node.h
      ApplyAvq() fills the new centres through a mutator of its own
      (upstream cached one in them with GetMutator(), then moved them into
      the node with the mutator still pointing at the moved-from local:
      incremental training later wrote through it, and upserts into an AVQ
      tree with incremental training failed or corrupted memory).
  src/scann/tree_x_hybrid/mutator.h
      UpdateDatapoint() and RemoveDatapoint() skip unused (kInvalidToken)
      assignment slots. With spilling (SOAR) upstream passed them to the
      health-stats collector, which indexed its arrays at -1: heap
      corruption when updating or deleting points in a SOAR index.
      Partition tokens from tokenizing a new vector are bounds-checked (a
      NaN/infinity vector tokenizes to -1; upstream indexed
      leaf_mutators_[-1]). RemoveDatapoint() removes from the base only
      after the leaf removals succeed, and AddDatapoint() undoes its base
      and leaf additions when a leaf add fails (upstream left size() out of
      step with the leaves). Incremental training on a projected tree or
      one with an upper tree is a config error with a clear message
      (upstream: a bare RET_CHECK failure).
      UpdateDatapoint() is all-or-nothing: it first
      appends the new vector to each of its leaves (undone if one fails),
      then updates the base and removes the old leaf entries. Upstream
      updated the base and the assignments step by step and returned at the
      first failing leaf, leaving the datapoint half-updated.
  src/scann/utils/bfloat16_helpers.h
      Bfloat16Decompress() shifts as unsigned; left-shifting a negative
      int16 is undefined behavior in C++17.
  src/scann/projection/chunking_projection.cc,
  src/scann/projection/projection_factory.h
      Chunking projections reject num_dims_per_block, num_blocks or
      input_dim below 1 (upstream divided by zero, SIGFPE, or CHECK-failed,
      aborting the process).
  src/scann/utils/hash_leaf_helpers.cc,
  src/scann/tree_x_hybrid/tree_ah_hybrid_residual.cc
      INT8_LUT16 lookups require 16 clusters per block (with fewer, the
      LUT16 kernels read past the lookup table: heap overflow).
      TrainAsymmetricHashingModel() calls TrainingOptions::Validate(), as
      the tree-AH residual factory does; an invalid AH projection used to
      end in a bare "SCANN_RET_CHECK failure" (null projector). In
      tree_ah_hybrid_residual.cc, a failure to create a query's lookup
      table is returned as a Status (upstream called .value(), which
      throws). The LUT16 searches skip empty leaves: when every searched
      leaf was empty (e.g. all points deleted), num_blocks was 0, and the
      AVX2 kernel divided by it (SIGFPE).
  src/scann/base/internal/single_machine_factory_impl.h,
  src/scann/base/internal/single_machine_factory_impl.cc
      The factory rejects binary distance measures (Hamming, ...) on
      non-binary data, bfloat16 brute force with distances other than dot
      product and squared L2 (both LOG(FATAL) upstream), and a brute-force
      fixed_point_multiplier_quantile outside (0, 1] (undefined behavior
      when quantizing tree leaves). For asymmetric hashing without residual
      quantization, an AH projection without input_dim gets the dataset's
      dimensionality, as the tree-AH residual factory already did with the
      centers' (upstream failed to build a tree with a PCA/TRUNCATE
      projection and non-residual AH, e.g. any squared L2 tree from the
      Python builder's pca()/truncate()).
  src/scann/partitioning/tree_brute_force_second_level_wrapper.cc
      An upper tree with FIXED8 or BFLOAT16 scoring rejects query
      tokenization distances its searchers don't support (LOG(FATAL)).
  src/scann/base/reordering_helper_factory.cc
      Fixed-point reordering rejects a NaN multiplier quantile and a missing
      dataset (upstream: undefined behavior, null reference). So does exact
      (float) reordering (upstream: LOG(FATAL) in ExactReorderingHelper's
      constructor, e.g. for an index directory without dataset.npy).
  src/scann/utils/reduction.h
      The dense and sparse-dense accumulation loops compare remaining
      counts instead of advancing past the end of the input; for an empty
      sparse datapoint that was arithmetic on a null pointer (undefined
      behavior, flagged by UBSan).
  GCC portability (upstream built only with clang, which accepts these as
  extensions or never checks them; GCC >= 13 rejects them). None changes
  clang's results (verified: still bit-identical to upstream):
    src/scann/utils/intrinsics/sse4.h,
    src/scann/distance_measures/many_to_many/many_to_many_impl.inc,
    src/scann/distance_measures/one_to_many/one_to_many_asymmetric_impl.inc,
    src/scann/tree_x_hybrid/internal/utils.cc
      Explicit bit casts between integer and float SIMD vectors (same bits)
      instead of implicit vector conversions or static_cast; in sse4.h the
      vectors were also declared with the wrong type.
    src/scann/distance_measures/one_to_many/one_to_many_symmetric.h,
    src/scann/distance_measures/one_to_many/one_to_many_impl_highway.inc,
    src/scann/base/single_machine_base.h
      Declarations before use, namespace qualification, and `typename` where
      C++17 requires them.
    src/scann/utils/common.h, src/scann/data_format/dataset.h
      SCANN_INLINE_FORWARDING: four virtual methods that forward to the same
      method on another view are not force-inlined under GCC (it rejects
      always_inline on what it sees as recursion).
    src/scann/utils/intrinsics/flags.h,
    src/scann/distance_measures/many_to_many/int8_tile.h,
    src/scann/distance_measures/many_to_many/many_to_many_templates.h
      The AMX kernels (clang's tile builtins) are compiled with clang only;
      upstream assumed any non-clang compiler had them.
  python/scann/scann_ops/py/scann_ops_pybind.py
      upsert() updates the docid bookkeeping only after the index accepted
      the vectors, and delete() validates all docids before changing
      anything (upstream left docids out of sync with the index when either
      raised). The docid bookkeeping is guarded by a reader/writer lock, so
      a search overlapping an upsert or delete maps its results to the
      right docids. delete() on a searcher without docids raises the same
      ValueError as upsert() (upstream: AttributeError).
      search_batched()/search_batched_parallel() map result padding (NaN
      distance) to None instead of docids[0]. upsert() rejects a docid
      listed twice (upstream added a new one twice but mapped it once,
      leaving a duplicate in docids). ScannSearcher keeps a copy of the
      docids list it is given (upstream mutated the caller's list).
      Vectors from other array libraries go through _host_array(): tensors
      that require grad are detached, GPU tensors copied to host memory,
      bfloat16 converted to float32 (upstream: a pybind TypeError).
      serialize() passes the pickled docids to the C++ serialize, which
      commits them with the index (upstream wrote scann_docids.pkl
      afterwards, and left a stale one when the index had no docids);
      load_searcher() rejects a scann_docids.pkl whose length doesn't match
      the index. ScannSearcher counts the calls that may change the index
      (upsert, delete, rebalance), for scann.torch's native backend.
      set_num_threads() documents the thread count's meaning (0.2.1).
  python/scann/scann_ops/py/scann_builder.py
      create_config() raises ValueError for tree(incremental_threshold=...)
      together with pca(), truncate() or upper_tree(), which the C++ side
      can't build (upstream failed only when initializing the searcher).
  The exact L2 -> inner-product reduction (0.2.1; squared L2 search through
  an inner-product index, opt-in):
  src/scann/proto/scann.proto
      ScannConfig.l2_as_dot_product (field 900) and the
      L2AsDotProductConfig message (scale, center). Not in upstream's
      proto: an upstream loader keeps it as an unknown field, and fails on
      the distance_measure such an index is saved with (see scann.cc).
  src/scann/scann_ops/cc/scann.h, src/scann/scann_ops/cc/scann.cc
      With l2_as_dot_product, Initialize() stores the dataset with the
      extra coordinate (center - |x|^2) / (2 scale), filling in unset
      parameters, and runs a DotProductDistance searcher over one more
      dimension; Search(), SearchBatched*() append `scale` to each query and
      return squared L2 distances; ToStoredDatapoints() /
      stored_dimensionality() / TransformsDatapoints() convert upserted
      vectors (and normalize them with spherical partitioning, as
      NormalizeDatapoints() does); RetrainAndReindex() keeps scale and
      center; config() returns the config as given; Serialize() saves the
      parameters and the distance measure "SquaredL2Distance
      [l2_as_dot_product: needs scann-core >= 0.2.1]", and LoadArtifacts()
      accepts it; CreateSearcher() refuses such a config.
  src/scann/scann_ops/cc/scann_npy.cc
      Upsert() converts rows with ToStoredDatapoints() (it normalized them
      for spherical partitioning before).
  python/scann/scann_ops/py/scann_builder.py
      l2_as_dot_product(scale=None, center=None); with it, tree(avq=...,
      soar_lambda=...), residual AH, the AH blocks and pca() are built for
      a dot-product index over one more dimension, and truncate(),
      spherical trees and autopilot() are rejected.
  Autopilot's tuned rules (0.2.1; the builders' default, upstream's rules
  on request):
  src/scann/proto/auto_tuning.proto
      AutopilotTreeAH.rules (UPSTREAM, the default, or TUNED_V1),
      noise_shaping_threshold and allow_l2_as_dot_product (fields 900-902).
      Not in upstream's proto: an upstream loader keeps them as unknown
      fields and applies upstream's rules.
  src/scann/utils/single_machine_autopilot.h,
  src/scann/utils/single_machine_autopilot.cc
      Autopilot() applies the TUNED_V1 rules when the config asks for them
      (AutopilotTreeAhTuned: at most about sqrt(n) leaves, leaves to search
      scaled to them, AH blocks and an anisotropic threshold scaled to the
      dimensionality and to the datapoints' norms, which is recorded in the
      config, tree AVQ, rounded lookup tables); upstream's AutopilotTreeAh
      is unchanged. ComputeAutopilotDataStats(),
      AutopilotChoosesL2AsDotProduct() and ApplyAutopilotL2AsDotProduct()
      are new.
  src/scann/scann_ops/cc/scann.cc
      Initialize() lets the tuned rules build a squared L2 autopilot index
      as l2_as_dot_product (ApplyAutopilotL2AsDotProduct).
  python/scann/scann_ops/py/scann_builder.py
      autopilot(mode, quantize, rules="tuned", allow_l2_as_dot_product=True):
      the tuned rules unless rules="upstream", which emits upstream's
      stanza.
  Performance changes (0.2.1; results unchanged unless noted):
  src/scann/scann_ops/cc/scann.h, src/scann/scann_ops/cc/scann.cc
      Thread pools are sized from scann_core::AvailableCPUs() (the CPU
      affinity mask, capped by the cgroup CPU quota, or SCANN_NUM_THREADS),
      evaluated at each Initialize, instead of the machine's online CPUs
      cached for the whole process; the query pool is started on first use
      (thread-safe) instead of in Initialize; SetNumThreads(n) means n
      workers (n - 1 pool threads plus the caller; upstream: n pool
      threads, while its default was NumCPUs() - 1); RetrainAndReindex()
      trains with the index's training_threads (upstream: the query pool).
      After a build, load or rebalance, the reordering data and the tree
      leaves' AH codes are madvise()d MADV_HUGEPAGE and MADV_COLLAPSE
      (SCANN_HUGEPAGES=0 turns it off).
      SearchBatchedParallel() counts the calling thread as a worker and
      hands out one chunk per fetch (upstream's ParallelForWithStatus<1>
      batched chunks dynamically, leaving workers idle with small pools).
      SearchBatchedRows() (new; SearchBatchedParallel() and ScannNumpy use
      it) searches row-major queries in place, as views, instead of copying
      them per call and again per chunk, checks NaN/infinity per chunk
      inside the parallel region (upstream: serially over all queries
      first), hands each finished chunk to a callback (ScannNumpy writes
      the results into its output arrays there), and searches chunks of
      one query, and on tree indexes chunks of at most 8, with Search()
      (results then match Search(), not a batched search, in the last
      bits).
      GetSearchParameters() reuses one TreeXOptionalParameters per thread and
      leaves_to_search value (upstream: a make_shared and a vector per
      query); result_multiplier() accessor.
  src/scann/scann_ops/cc/scann_npy.cc
      SearchBatched() releases the GIL before touching the queries (no
      copy) and returns arrays allocated before the search and filled by
      the search's chunks (upstream copied the results twice, the second
      time with the GIL held). Search() writes the results into the
      returned arrays directly (upstream: into two vectors, then copied).
      Loading an index (the ScannNumpy(dir, ...) constructor) releases the
      GIL (upstream held it for the whole load). Upsert() has an overload
      for a 2-D float32 array, read in place with the GIL released (the
      list-of-rows one made one Python array object per row), and runs the
      incremental maintenance once per call instead of after every batch.
  src/scann/scann_ops/cc/python/scann_pybind.cc
      upsert() takes a C-contiguous float32 2-D array without conversion,
      and anything else through the list-of-rows overload; named arguments,
      batch_size defaults to 256.
  src/scann/utils/intrinsics/attributes.h
      The AVX2 / AVX-512 / VNNI / AMX target attributes include popcnt.
  src/scann/utils/intrinsics/flags.cc
      The x86 ignore_avx2 / ignore_avx512 / ignore_avx512_vnni / ignore_amx
      flags are honored (upstream read none of them). ignore_amx defaults
      to false, keeping the AMX kernels on as they were in practice
      (upstream's documented default was true, but unread).
  src/scann/tree_x_hybrid/tree_ah_hybrid_residual.h
      ForEachLeafPackedCodes(), used by scann.cc to advise huge pages for
      the leaves' codes.
  src/scann/base/health_stats_collector.h
      Initialize() computes the per-partition quantization error in
      parallel (a pool sized by scann_core::AvailableCPUs(), for 65,536
      datapoints or more) and sums the partitions in order afterwards: the
      same statistics, bit for bit.
  src/scann/utils/reordering_helper_interface.h,
  src/scann/utils/reordering_helper.h, src/scann/base/single_machine_base.cc
      NonEmptyDatasetDimensionality(): the query dimensionality check in
      FindNeighbors() no longer copies the reordering helper's dataset
      shared_ptr on every search (a reference count all searching threads
      shared).
  src/scann/tree_x_hybrid/internal/utils.cc
      DeduplicateDatabaseSpilledResults() sorts with std::sort by (index,
      distance) instead of std::stable_sort by index (a temporary buffer
      per query); same results.
  src/scann/distance_measures/one_to_many/one_to_many_asymmetric_impl.inc
      With a float query on AVX2, OneToManyAsymmetricTemplate computes two
      groups of three rows at once (six independent accumulators; each row
      computed exactly as before), HandleXDims expands each 8-byte half of
      16 int8 values straight from memory, and the paired path prefetches
      each cache line of the next rows once instead of every 16 elements.
  src/scann/utils/fast_top_neighbors.h
      InitLikeNew(): the state a new object gets from Init(), reusing the
      object's arrays when they are large enough (allocated_capacity_);
      ScratchIsRetainable() for scann_core::ScratchLease.
  src/scann/trees/kmeans_tree/kmeans_tree_node.h,
  src/scann/trees/kmeans_tree/kmeans_tree_node.cc,
  src/scann/tree_x_hybrid/tree_ah_hybrid_residual.cc
      FindChildrenWithSpilling(), GetAllDistancesInt8() (dense queries),
      PostprocessDistancesForSpilling(), FindNeighborsImpl() and
      FindNeighborsInternal1() (global top-N) use per-thread scratch objects
      (scann_core::ScratchLease, core/scann_core/scratch.h) for the
      centroid distances, the adjusted query, the top-N buffers, the token
      list and the leaf list, instead of allocating them per query.
  src/scann/hashes/internal/asymmetric_hashing_impl.h,
  src/scann/hashes/internal/asymmetric_hashing_impl.cc
      Lut16DotProductLookupBuilder: the raw float LUT of a dot-product
      model with 16 centers per block and blocks of 1-4 dimensions, from a
      dimension-major copy of the centers, evaluating each entry with the
      same operations as the generic DenseDistanceOneToMany path for the
      build's Highway lane count (and DenseDotProductSse4 for the 16th
      center); Create() verifies that against the generic path on test
      inputs and declines otherwise. ConvertLookupToFixedPoint's truncating
      int8 conversion is branch-free (same values), so it vectorizes.
  src/scann/hashes/asymmetric_hashing2/querying.h
      CreateLookupTable() uses the model's Lut16DotProductLookupBuilder
      (made on first use) for DOT_PRODUCT lookups, reading the blocks
      straight from the query when the projection is plain chunking
      (checked once with a probe query), into a per-thread buffer.
  python/scann/scann_ops/py/scann_ops_pybind.py
      search() on a searcher without docids calls the pybind searcher
      directly (no docid lock, no context manager). upsert() passes the rows
      as one float32 2-D array, and its batch_size defaults to 256, the C++
      default (upstream: 1).
docs/algorithms.md is upstream's scann/docs/algorithms.md, modified (added
explanation of asymmetric hashing and anisotropic quantization, and links).
python/scann/tf.py is adapted from upstream's
scann/scann/scann_ops/py/scann_ops.py (its builder() docstring, and the
builder / create_searcher / ScannSearcher search API): the searcher is the
pybind one, called through tf.numpy_function instead of a TensorFlow op, and
serialize_to_module() / searcher_from_module() raise NotImplementedError.
tf_op/scann_tf_ops.cc is adapted from upstream's
scann/scann/scann_ops/cc/ops/scann_ops.cc and kernels/scann_ops.cc (the
ScannSearch / ScannSearchBatched ops and their search and result-shaping
semantics): rewritten against TensorFlow's C API, with the index passed as
string tensors of its files instead of a resource, a searcher cache keyed
by index id and a fingerprint of those tensors, and batched results exactly
final_num_neighbors wide when it is given. tf_op/python/scann_tf_ops/
__init__.py is adapted from upstream's scann/scann/scann_ops/py/
scann_ops.py (its API and builder() docstring): the searcher is a tf.Module
holding the index files as variables, and serialize_to_module() /
searcher_from_module() use them.
Everything else (build system, the C++ config builder in core/, Rust
bindings, tests, examples, other docs,
python/scann/__init__.py and the empty python/scann/**/__init__.py files)
is new in scann-core. Upstream's own repository has no NOTICE file.

Arm (aarch64) support
---------------------
The Neon/SVE implementations and AArch64 run-time feature detection come
from work by Arm engineers, contributed to ScaNN but not merged upstream as
of 2026-09-26:
  * google-research/google-research PR #3374, "scann: Add AArch64 run-time
    feature detection" (Gerda Zsejke More, gerdazsejke.more@arm.com)
  * lizhang-arm/google-research PR #1, "Arm: Add Neon implementations for
    many to many functions" (Gerda Zsejke More)
  * lizhang-arm/google-research PR #2, "Arm: Add Neon implementation for
    ScaNN indexDatapointNoiseShaped" (Li Zhang, li.zhang2@arm.com)
  * lizhang-arm/google-research PR #3, "scann: Add Arm implementation of
    DenseDotProductInt8Float" (Li Zhang)
Their commits were rebased onto the extraction commit above and kept with
their original authorship. Files they add carry The Google Research
Authors' header, as contributed; files they modify say so in their
headers. On top of them, scann-core:
  - merged the two independently created versions of
    utils/intrinsics/mem_neon.h;
  - moved the <arm_neon.h> include in
    hashes/internal/asymmetric_hashing_impl_neon.cc inside its
    `#if defined(__aarch64__)` guard (x86 builds failed otherwise).
The first commit of PR #3374 also enables x86 CPU feature detection
(PLATFORM_IS_X86), which upstream open-source builds lack; see its commit.

Further aarch64 fixes to upstream files (needed with highway 1.4.0 and a
baseline without the AES extension):
  src/scann/distance_measures/many_to_many/int8_tile.cc
      Includes hwy/foreach_target.h before int8_tile.h, as Highway
      requires; otherwise the N_NEON dispatch target came out empty
      ("use of undeclared identifier 'N_NEON'").
  src/scann/utils/hwy-compact.cc
      Static target HWY_NEON_WITHOUT_AES when AES isn't enabled (the code
      doesn't use AES); HWY_NEON otherwise, as before.

Third-party dependencies are downloaded at configure time (see
cmake/Dependencies.cmake), except two vendored in third_party/ with their
licenses: cnpy (MIT, Copyright (c) Carl Rogers, 2011) and googletest's
gtest_prod.h (BSD 3-Clause, Copyright 2008 Google Inc.). All are distributed
under their own licenses:
  abseil-cpp   Apache License 2.0
  protobuf     BSD 3-Clause
  highway      Apache License 2.0 or BSD 3-Clause (dual-licensed)
  Eigen        MPL 2.0 (some files BSD/Apache/MINPACK; see its COPYING.README)
  cnpy         MIT
  googletest   BSD 3-Clause         (gtest_prod.h only)
  zlib         zlib license
  pybind11     BSD 3-Clause         (Python bindings only)
  cxx          MIT or Apache 2.0    (Rust bindings only, via crates.io)
Binaries that statically link these carry their license obligations.
