# The environment variables this repository reads, in the layout every
# worktree's .env takes.
#
# `just init` provisions .env from this file, and `just init-env` runs that one
# step alone. The layout is always this file's: an existing .env keeps its own
# values and is laid out again to match, and a new one takes the values the
# main worktree's .env holds for the names declared here. A name this file
# does not declare is one nothing reads: it is never copied into a new
# worktree. Uncomment a line to set it.
# .env is gitignored — never commit real tokens.

# ---------------------------------------------------------------------------
# vaultspec-rag configuration
#
# Resolution order, one framework rule every vaultspec package follows:
#   1. Invocation - a CLI flag (noted per variable below).
#   2. Session environment - this process's own environment. Three settings
#      chain to a shared framework name read behind this package's own:
#        VAULTSPEC_RAG_ROOT          -> VAULTSPEC_TARGET_DIR
#        VAULTSPEC_RAG_LOG_LEVEL     -> VAULTSPEC_LOG_LEVEL
#        VAULTSPEC_RAG_STDIO_WATCHDOG -> VAULTSPEC_STDIO_WATCHDOG
#      A session that sets only the shared name configures every vaultspec
#      tool at once; setting this package's own scoped name overrides it for
#      this package alone. No other variable below reads a shared name.
#   3. Workspace .env - credentials only; the eligible name is
#      VAULTSPEC_RAG_TYPESAFE_API_KEY. No other variable in this
#      file is ever read from a .env, and no variant (.env.local and the
#      like) is read at all. The file opens only when both hold: the running
#      interpreter lives inside the workspace, and this package's resolved
#      install mode for that workspace is dependency or dev - never for a
#      globally installed tool. This is a settings file, not a repository
#      surface: repository content may hand you a key, never reconfigure
#      the tool.
#   4. Persisted configuration, where a knob has one (the service status
#      directory's local-only marker, noted per variable below).
#   5. The shipped default named per variable below.
#
# This file and the source are held to each other by guard tests, in both
# directions. Every variable the source reads is declared below exactly once,
# under a description of its own, showing the default it ships with; and
# nothing is declared below, or named anywhere in the source, that the source
# does not read. The EnvVar enum in
# config/_types.py is the authoritative list of the product's variables, and
# the development harness's are in the last section. The only names kept off
# this file are markers one process sets on its own child, which nobody sets
# by hand.
#
# Value parsing, applied when the settings are built:
#   - One boolean vocabulary, everywhere: 1, true, yes or on read TRUE; 0,
#     false, no or off read FALSE (case-insensitive, surrounding whitespace
#     stripped). Anything else is REJECTED with a message naming the
#     variable, the value and the accepted words - a typo such as "treu" is
#     refused rather than quietly read as false.
#   - Integers and floats are parsed with int()/float(); a non-numeric value
#     is rejected the same way.
#   - A blank value means unset for every variable below, whatever its type -
#     a boolean included - and falls through to the next rung, ending at the
#     shipped default. That makes an unexpanded VAR="$UNSET" harmless instead
#     of silently turning a switch off or repointing a managed directory at
#     the working directory.
#   - One exception: VAULTSPEC_RAG_STDIO_WATCHDOG is a protective switch. A
#     blank or unrecognised value there leaves it armed (and warns) rather
#     than falling through, because the safer reading of a typo on a
#     protective switch is to keep the guard up.
#   - One bad value anywhere makes the whole settings object unbuildable, so
#     every rejected setting is reported together when the process starts,
#     not one at a time as each knob happens to be read.
#   - Relative paths resolve against the project root. Use forward slashes on
#     Windows.
# ---------------------------------------------------------------------------

# ---------------------------------------------------------------------------
# Project and data locations
# ---------------------------------------------------------------------------

# Project root directory (default: discovered from the working directory).
# Determines what gets indexed and where .vault/ is resolved from, and names
# the project the stdio MCP server serves. Precedence: --target, then this,
# then the framework-wide VAULTSPEC_TARGET_DIR, then discovery from the
# working directory. A value that is not an enrolled workspace fails the
# command instead of being ignored. The resident service ignores both root
# variables and serves every root at once.
# CLI: --target
# VAULTSPEC_RAG_ROOT=

# RAG data root relative to project root (default: .vault/data/search-data).
# CLI: --data-dir
# VAULTSPEC_RAG_DATA_DIR=.vault/data/search-data

# On-disk store subdirectory relative to the data root (default: qdrant).
# Used by the local backend only; server mode stores under
# VAULTSPEC_RAG_QDRANT_STORAGE_DIR instead.
# CLI: --storage-dir
# VAULTSPEC_RAG_QDRANT_DIR=qdrant

# ---------------------------------------------------------------------------
# Service runtime and logging
# ---------------------------------------------------------------------------

# Service status directory — stores service.json, the local-only marker, the
# provisioned qdrant binary, and logs (default: ~/.vaultspec-rag). Per-host and
# gitignored; never point it at the project tree.
# CLI: --status-dir
# VAULTSPEC_RAG_STATUS_DIR=~/.vaultspec-rag

# Service log filename relative to status dir (default: service.log).
# CLI: --log-file
# VAULTSPEC_RAG_LOG_FILE=service.log

# HTTP service port, also the MCP fast path (default: 8766).
# CLI: --port
# VAULTSPEC_RAG_PORT=8766

# Absolute path to the compiled monitor executable. Unset resolves the
# installed vaultspec-rag-monitor command on PATH; missing binaries fail
# coupled server startup. Build and installation happen before service start.
# VAULTSPEC_RAG_MONITOR_BINARY=

# Absolute path to the vaultspec-rag command the monitor runs its backend
# controls through (default: unset — the executable installed beside the
# monitor binary). Read by the monitor itself, not by the service. Set it when
# the two are installed in different directories; a relative path is refused.
# A monitor the service supervisor launched uses the service's own Python
# environment instead, and ignores this.
# VAULTSPEC_RAG_MONITOR_OWNER=

# Log level: DEBUG, INFO, WARNING, ERROR, CRITICAL (default: WARNING for the
# CLI and the stdio MCP server; the resident daemon's own last rung is INFO,
# since its output is a managed log nobody is watching live). Falls back to
# the framework-wide VAULTSPEC_LOG_LEVEL when unset. A name outside that list
# is refused rather than degraded to the default.
# CLI: --verbose (INFO) / --debug (DEBUG)
# VAULTSPEC_RAG_LOG_LEVEL=WARNING

# Size threshold in bytes at which a managed log file rotates
# (default: 2097152, 2 MiB). The budget applies per source — service.log and
# qdrant.log each get the full allowance, it is not divided between them. One
# generation is sized to the window the log readers scan back over, so raising
# it keeps bytes that no log command or tool will return.
# VAULTSPEC_RAG_MANAGED_LOG_MAX_BYTES=2097152

# Rotated backups retained per managed log source (default: 5). With the two
# defaults above the aggregate on-disk log budget is roughly 24 MiB.
# VAULTSPEC_RAG_MANAGED_LOG_BACKUP_COUNT=5

# ---------------------------------------------------------------------------
# Resident service lifecycle
# ---------------------------------------------------------------------------

# Seconds a project may sit idle before the service evicts its slot
# (default: 1800). Eviction closes that project's store AND stops its file
# watcher, so a project with automatic index updates enabled stops receiving
# them after this long without a search or an index run, until the next
# access re-warms it. The service log records each one as
# "Evicted ProjectSlot <root> (reason=idle)". Raise it to keep watchers alive
# on projects that are used intermittently; lower it to release GPU and store
# resources sooner.
# VAULTSPEC_RAG_SERVICE_IDLE_TTL_SECONDS=1800

# Maximum project slots held open at once (default: 16). Reaching the cap
# evicts the least recently used slot, with the same watcher consequence as
# the idle eviction above.
# VAULTSPEC_RAG_SERVICE_MAX_PROJECTS=16

# Seconds a client waits for a search response (default: 300). Generous
# because a cold query can sit behind model load. A malformed value falls
# back to the default rather than failing the call.
# CLI: --timeout
# VAULTSPEC_RAG_SEARCH_TIMEOUT=300

# Maximum bounded freshness wait a search may request (default: 30 seconds).
# Raise it to permit longer waits for an in-flight publication; lower it to
# cap how long callers can hold service capacity while waiting for a fresh index.
# VAULTSPEC_RAG_SEARCH_FRESHNESS_WAIT_MAX_SECONDS=30

# Seconds a client waits for a lifecycle/admin call (default: 30).
# VAULTSPEC_RAG_ADMIN_TIMEOUT=30

# Seconds a pause waits for in-flight work to drain before it refuses
# (default: 20). Raise it on a service running long indexing jobs, so a pause
# does not refuse work that was about to finish; keep it under
# VAULTSPEC_RAG_ADMIN_TIMEOUT, or the caller's transport deadline fires first
# and the refusal arrives with no envelope to explain what is still holding.
# VAULTSPEC_RAG_PAUSE_DRAIN_TIMEOUT=20

# Seconds a client waits for a reindex call (default: 900). Reindex admits
# every requested domain before it queues, and code admission scans the
# whole tree, so this call is bounded by repository size rather than by a
# lifecycle round trip. Lowering it toward the admin bound makes a large
# tree report a failure for a job the service went on to accept and run.
# VAULTSPEC_RAG_REINDEX_TIMEOUT=900

# Seconds to wait for the managed qdrant server to report ready
# (default: 300). A multi-hundred-GB store with ~170 collections was measured
# at ~131s, so raise this rather than patching the supervisor if a very large
# store times out at startup.
# VAULTSPEC_RAG_QDRANT_READY_TIMEOUT=300

# Collections the managed qdrant server loads at once while it starts
# (default: 2, minimum 1). 1 loads them serially; a higher value shortens the
# startup of a store holding many collections, at the cost of CPU and storage
# pressure while they open. Takes effect on the next managed qdrant start, and
# the applied value is logged. It does not configure a remote server, and it
# leaves shard and segment load concurrency alone.
# VAULTSPEC_RAG_QDRANT_COLLECTION_LOAD_CONCURRENCY=2

# Maximum simultaneously tracked non-terminal (queued or running) index jobs
# (default: 64). Non-terminal records are exact-addressable and so cannot be
# evicted; at the bound the service refuses new admissions rather than growing
# an unbounded registry or silently dropping controllable work.
# VAULTSPEC_RAG_JOB_MAX_NONTERMINAL=64

# Seconds a graceful daemon stop waits for running index jobs to unwind
# (default: 300). Bounds the wait so shutdown can never hang on an
# unacknowledged indexer.
# VAULTSPEC_RAG_JOB_SHUTDOWN_TIMEOUT_SECONDS=300

# ---------------------------------------------------------------------------
# Storage backend selection
# ---------------------------------------------------------------------------

# Supervise the pinned qdrant server binary and route stores at it
# (default: 1). Server mode is the assumed backend — an A/B on a 469k-chunk
# corpus measured a ~54x end-to-end win — so leave it on unless you have a
# reason not to.
# CLI: --qdrant / --no-qdrant
# VAULTSPEC_RAG_QDRANT_SERVER=1

# Use the per-project on-disk store regardless of the server default
# (default: 0). This is the first-class opt-out for CI, offline, and small
# projects, and it always wins: effective server mode is
# qdrant_server AND NOT local_only. `install --local-only` also persists the
# choice to {status_dir}/local-only.json, which outranks the built-in default
# but not this variable.
# CLI: --local-only
# VAULTSPEC_RAG_LOCAL_ONLY=0

# HTTP port of the managed qdrant child (default: 8765). Its gRPC listener
# binds one below that. Keep it clear of VAULTSPEC_RAG_PORT.
# VAULTSPEC_RAG_QDRANT_PORT=8765

# Storage directory for the managed qdrant server
# (default: ~/.vaultspec-rag/qdrant-server/storage). Server storage is shared
# and multi-root — per-root data lives in namespaced collections inside it — so
# it belongs under the managed service directory, never in a project data dir.
# VAULTSPEC_RAG_QDRANT_STORAGE_DIR=~/.vaultspec-rag/qdrant-server/storage

# Absolute path to a qdrant binary you supply yourself (default: unset — use
# the managed install of the pinned release). For a custom build. Must be set
# together with VAULTSPEC_RAG_QDRANT_BINARY_SHA256: the file is hashed and
# compared with that digest before every start, and never run unverified.
# Setting only one of the two is refused for every command.
# VAULTSPEC_RAG_QDRANT_BINARY=

# SHA256 of the binary VAULTSPEC_RAG_QDRANT_BINARY names (default: unset).
# 64 hexadecimal characters, either letter case. Print it with
# `Get-FileHash -Algorithm SHA256 <path>` on Windows, or `sha256sum <path>`
# or `shasum -a 256 <path>` elsewhere. Required with the path and refused
# without it. Read from the process environment only, never from a file in a
# project: a repository must not be able to vouch for a binary.
# VAULTSPEC_RAG_QDRANT_BINARY_SHA256=

# Download the pinned qdrant server binary when a host `server start` finds
# none installed (default: 1). Starting the service is the consent, as it is
# for model weights. Set 0 to fail with the install command instead. A client
# installation never downloads a binary, whatever this says.
# CLI: --qdrant-auto-provision / --no-qdrant-auto-provision
# VAULTSPEC_RAG_QDRANT_AUTO_PROVISION=1

# Release base URL the qdrant server binary is downloaded from
# (default: https://github.com/qdrant/qdrant/releases/download). The archive
# is fetched from {base}/v{version}/{asset}, so a mirror must keep that path
# layout. Must be an https URL with a host; a port and a path prefix are
# allowed, credentials, a query and a fragment are not. The download is
# checked against digests compiled into the tool, which no variable changes.
# Read from the process environment only, never from a file in a project.
# VAULTSPEC_RAG_QDRANT_RELEASE_BASE_URL=https://github.com/qdrant/qdrant/releases/download

# Hosts a qdrant binary download may be redirected to (default: the three
# hosts shown). Comma-separated bare host names, case-insensitive; an entry
# with a scheme, a port or a path is rejected, and so is an empty list. The
# base URL's own host always serves the first request. Setting this replaces
# the default, so a mirror that redirects to its own storage host lists that
# host here. Read from the process environment only.
# VAULTSPEC_RAG_QDRANT_DOWNLOAD_HOSTS=github.com,release-assets.githubusercontent.com,objects.githubusercontent.com

# URL of a remote or externally managed qdrant server (default: unset).
# Setting it selects server mode in the store.
# VAULTSPEC_RAG_QDRANT_URL=

# API key for a remote qdrant server (default: unset). A credential — keep it
# in .env, never in a committed config.
# VAULTSPEC_RAG_QDRANT_API_KEY=

# Vector quantization for newly created collections (default: unset, no
# quantization). Accepts scalar (or int8), turbo, product (or pq). Applied at
# collection-creation time only: changing it leaves existing collections as
# they are, so a rebuild is needed for it to take effect on them.
# VAULTSPEC_RAG_QDRANT_QUANTIZATION=

# ---------------------------------------------------------------------------
# Store write resilience
# ---------------------------------------------------------------------------
# A transient store-write failure (disk pressure, a write-ahead-log stall) is
# retried with bounded exponential backoff before the operation is abandoned.

# Per-operation deadline in seconds before a store write is abandoned
# (default: 120).
# VAULTSPEC_RAG_STORE_OPERATION_TIMEOUT_SECONDS=120

# Retry attempts for a transient store-write failure (default: 5).
# VAULTSPEC_RAG_STORE_WRITE_RETRY_ATTEMPTS=5

# Initial backoff in seconds before the first store-write retry (default: 0.5).
# VAULTSPEC_RAG_STORE_WRITE_RETRY_BASE_SECONDS=0.5

# Maximum backoff in seconds between store-write retries (default: 8).
# VAULTSPEC_RAG_STORE_WRITE_RETRY_MAX_SECONDS=8

# ---------------------------------------------------------------------------
# Scheduled storage maintenance
# ---------------------------------------------------------------------------
# The daemon's maintenance tick reclaims namespaces whose project root has been
# gone for a continuous, persisted grace window. Unknown or unverifiable
# namespaces are never touched, and a point-bearing namespace is archived
# before it is dropped — a failed archive aborts the delete.

# Scheduled auto-prune on/off (default: 1). Server mode only.
# VAULTSPEC_RAG_STORAGE_AUTOPRUNE=1

# Minutes between maintenance cycles (default: 60). Fractional values are
# accepted.
# VAULTSPEC_RAG_STORAGE_AUTOPRUNE_INTERVAL_MINUTES=60

# The three GRACE_HOURS windows below are how long a namespace has to stay
# observably dead before it may be destroyed, and each is rejected below 1
# hour: at 0 the very cycle that first saw a namespace would also be cleared
# to drop it, on a single scan. Do not confuse them with
# ..._EPHEMERAL_IDLE_HOURS at the end of this block, which is a tier's own
# on/off switch and where 0 means the OPPOSITE - that tier reclaims nothing.

# Continuous hours a namespace must be observed orphaned before an EMPTY
# (zero-point) one is reclaimed (default: 24, minimum 1). Any live or
# unverifiable observation resets the clock, so a transiently missing root
# — an unplugged drive, an offline share, a worktree mid-recreate — only ever
# extends protection. Lowering this shortens the window in which a real root
# can come back before its data is dropped.
# VAULTSPEC_RAG_STORAGE_AUTOPRUNE_GRACE_HOURS=24

# Continuous orphan hours before a POINT-BEARING namespace is archived and
# reclaimed (default: 168, one week; minimum 1). The longer of the two tiers
# because this one destroys indexed data.
# VAULTSPEC_RAG_STORAGE_AUTOPRUNE_GRACE_HOURS_DATA=168

# Continuous orphan hours before a TEMP-ROOTED namespace is reclaimed
# (default: 24, minimum 1), replacing both windows above for that class
# whatever it holds. A root created under the OS temp directory and now
# provably gone was a throwaway sandbox that was torn down. The point count
# still selects the tier, so a point-bearing one is still archived before it
# is dropped; only the wait is shorter. Raise this if you index real projects
# out of a temp directory. An absent root on an unreachable volume classifies
# unverifiable rather than orphaned and never draws this window.
# VAULTSPEC_RAG_STORAGE_AUTOPRUNE_GRACE_HOURS_EPHEMERAL=24

# Days a snapshot archive is kept before the retention sweep deletes it
# (default: 30). This is the window in which a wrongly reclaimed namespace can
# still be restored.
# VAULTSPEC_RAG_STORAGE_AUTOPRUNE_ARCHIVE_RETENTION_DAYS=30

# Total size cap in GB on the archive directory (default: 64). Oldest archives
# are evicted first when the cap is exceeded.
# VAULTSPEC_RAG_STORAGE_AUTOPRUNE_ARCHIVE_MAX_GB=64

# Maximum namespaces reclaimed in one cycle (default: 16). Bounds the blast
# radius of a single tick.
# VAULTSPEC_RAG_STORAGE_AUTOPRUNE_MAX_PER_CYCLE=16

# Idle hours after which a temp-rooted namespace whose root still EXISTS is
# treated as dangling (default: 72). Catches harness temp dirs that outlive
# their usefulness but never disappear. Runs through the same empty/data tiers,
# archive gate, and per-cycle cap. 0 disables this tier - the one knob in this
# block where 0 is accepted, and it means reclaim nothing rather than reclaim
# on sight.
# VAULTSPEC_RAG_STORAGE_AUTOPRUNE_EPHEMERAL_IDLE_HOURS=72

# Converge collections created before per-collection preallocation was bounded
# onto the current segment geometry (default: 1). Non-destructive: no points
# move, and it was measured to reclaim 63-84% of disk per collection.
# VAULTSPEC_RAG_STORAGE_RECONCILE=1

# Maximum collections reconciled per cycle (default: 4). Keeps a large drifted
# backend converging over several ticks instead of saturating disk in one.
# VAULTSPEC_RAG_STORAGE_RECONCILE_MAX_PER_CYCLE=4

# Per-collection seconds to wait for the merge to settle (default: 300). A
# collection that has not settled is reported as still converging and retried
# later — never as failed.
# VAULTSPEC_RAG_STORAGE_RECONCILE_BUDGET_SECONDS=300

# ---------------------------------------------------------------------------
# Models
# ---------------------------------------------------------------------------
# Changing a model against an index built with a different one leaves the
# stored vectors in the old model's space. Rebuild the affected index after
# changing any of these. A dense width that disagrees with the dense model is
# rejected by the server on the first upsert, so that mismatch fails loudly
# rather than storing nonsense.

# Dense embedding model (default: Qwen/Qwen3-Embedding-0.6B).
# CLI: --model
# VAULTSPEC_RAG_EMBEDDING_MODEL=Qwen/Qwen3-Embedding-0.6B

# Commit the dense model is fetched and loaded at (default: unset). Unset is
# the pinned state: the default model then uses the commit compiled into the
# tool, and its snapshot is used only when every file matches the SHA256
# compiled in for it. Setting this names another commit, which is honoured
# and reported as unpinned, because no compiled-in digest describes it. A
# model named above that is not the default has no compiled-in commit, so
# unset then means the hub's default branch. Must be a full commit id of 40
# hexadecimal characters; a branch or tag name is rejected.
# VAULTSPEC_RAG_EMBEDDING_MODEL_REVISION=

# Dense vector width; must match the dense model (default: 1024).
# VAULTSPEC_RAG_EMBEDDING_DIMENSION=1024

# Sparse model: Linkup-Platform/linkup-sparseup-embed-v1. Public, pinned to a
# commit and to the SHA256 of every file in it. The adapter supports this
# repository only; another sparse model is refused. It has no revision
# variable: its repository ships the code that builds the model, and a commit
# the environment could change would select code no compiled-in digest covers.
# VAULTSPEC_RAG_SPARSE_MODEL=Linkup-Platform/linkup-sparseup-embed-v1

# Compute SPARSEUP sparse vectors alongside the dense ones (default: 1).
# Turning it off frees the most GPU memory of any single knob here, and drops
# hybrid retrieval back to dense-only — exact-token matches (identifiers, error
# strings) get noticeably weaker.
# VAULTSPEC_RAG_SPARSE_ENABLED=1

# Cross-encoder reranker model (default: BAAI/bge-reranker-v2-m3).
# VAULTSPEC_RAG_RERANKER_MODEL=BAAI/bge-reranker-v2-m3

# Commit the reranker model is fetched and loaded at (default: unset). The
# same rule as the dense model's: unset uses the compiled-in commit and its
# file digests; a value moves the model to a commit nothing vouches for, and
# it is reported as unpinned. A full commit id of 40 hexadecimal characters.
# VAULTSPEC_RAG_RERANKER_MODEL_REVISION=

# Model hub the three models above are downloaded from
# (default: https://huggingface.co). Point it at a mirror that keeps the hub's
# path layout. Must be an https URL with a host; a port and a path prefix are
# allowed, credentials, a query and a fragment are not. It is exported to
# HF_ENDPOINT when each process starts, because the Hub client reads its
# endpoint once, when it is first loaded - so set it before starting the
# service, and restart a running one to change it. It outranks an HF_ENDPOINT
# set directly; left unset, it exports nothing and HF_ENDPOINT is untouched.
# Read from the process environment only, never from a file in a project.
# VAULTSPEC_RAG_HF_ENDPOINT=https://huggingface.co

# Longest one model fetch may take, in seconds, every repository together
# (default: 14400, four hours). It covers `install`, `server warmup` and the
# fetch a host `server start` makes, including time spent waiting for another
# command that is fetching into the same cache. A download that receives too
# little is stopped much sooner; this bounds one that keeps receiving just
# enough. A fetch stopped by it keeps every file that finished, so running the
# command again continues. Raise it for a link that needs longer.
# VAULTSPEC_RAG_MODEL_FETCH_DEADLINE_SECONDS=14400

# Run the cross-encoder rerank pass (default: 1). Off frees GPU memory and
# shortens the query path, at a direct cost in result ordering quality.
# VAULTSPEC_RAG_RERANKER_ENABLED=1

# Reranker forward batch size (default: 32). Lower it if reranking a large
# candidate set approaches the CUDA ceiling.
# VAULTSPEC_RAG_RERANKER_BATCH_SIZE=32

# Token bound for cross-encoder inputs (default: 1024). The reranker scores
# token-bounded full candidate content and its tokenizer truncates each
# (query, content) pair to this length; 1024 covers a 3000-char vault chunk or
# a 1500-char code chunk plus the query. Lowering it truncates real content out
# of the ranking decision.
# VAULTSPEC_RAG_RERANKER_MAX_LENGTH=1024

# Dense-encoder backend (default: torch). "torch" is the only validated path;
# "onnx" is experimental, requires sentence-transformers[onnx-gpu] in an
# onnxruntime-compatible CUDA environment, and degrades back to torch on any
# failure.
# VAULTSPEC_RAG_DENSE_BACKEND=torch

# Cached ONNX model file, relative to the model repo
# (default: onnx/model_O4.onnx). Only consulted when DENSE_BACKEND=onnx.
# VAULTSPEC_RAG_DENSE_ONNX_FILE=onnx/model_O4.onnx

# ---------------------------------------------------------------------------
# Encoding batch sizes and text bounds
# ---------------------------------------------------------------------------
# The encoders halve their batch size and retry on a CUDA out-of-memory error,
# down to a batch of one, so most cards need no tuning. Lower these to reduce
# memory pressure before that automatic backoff has to engage.

# Outer batch size fed to the embedding pipeline (default: 64).
# VAULTSPEC_RAG_EMBEDDING_BATCH_SIZE=64

# Inner encode sub-batch for the VAULT path (default: 32). Inputs are
# heading-aware chunks capped at VAULT_CHUNK_CHARS and length-sorted per slice,
# so padding waste is bounded and a larger sub-batch keeps the tensor cores
# fed.
# VAULTSPEC_RAG_EMBEDDING_ENCODE_BATCH_SIZE=32

# Inner encode sub-batch for the CODEBASE path (default: 32). Code chunks are
# short (<=1500 chars) and length-sorted, so the padding pathology that keeps
# other paths small does not apply here.
# VAULTSPEC_RAG_EMBEDDING_CODE_ENCODE_BATCH_SIZE=32

# Inner encode sub-batch for the DOCUMENT path (default: 12). Deliberately
# smaller: document fragments fill the model's full ~2048-token window, roughly
# 2.7x a vault chunk's token volume, so 32 of them is far more activation
# memory than the vault batch was sized for.
# VAULTSPEC_RAG_EMBEDDING_DOCUMENT_ENCODE_BATCH_SIZE=12

# Estimated token footprint allowed per planned encode bucket (default:
# 24000). The planner cuts a length-sorted slice into buckets whose items x
# padded per-item token estimate stays under this budget, which is what bounds
# activation memory per forward pass; a CUDA out-of-memory error lowers the
# learned ceiling below it for the rest of the run. Lower it to buy headroom on
# a busy or smaller card, at a direct cost in encode throughput, since smaller
# buckets mean more forward passes for the same slice.
# VAULTSPEC_RAG_EMBEDDING_ENCODE_TOKEN_BUDGET=24000

# Independent padded token footprint for sparse encode buckets (default:
# 24000). Dense encoding keeps the budget above. The measured 8192-token
# option avoided sparse OOM retries on long and mixed workloads, but used
# more GPU energy on short inputs; choose it for your measured workload.
# VAULTSPEC_RAG_EMBEDDING_SPARSE_ENCODE_TOKEN_BUDGET=24000

# Chars-per-token ratio the encode bucket planner estimates with (default: 3).
# Deliberately conservative and equal to the document chunking divisor: prose
# tokenises nearer 4 chars per token, but token-dense content such as digit
# tables measures as low as ~1. Raising it plans larger buckets and reads as
# free throughput on prose, while re-admitting out-of-memory errors on exactly
# the dense content the low ratio exists to cover.
# VAULTSPEC_RAG_EMBEDDING_ENCODE_CHARS_PER_TOKEN=3

# Hard cap on the sequence length the dense model may process (default: 2048).
# Prevents the model advertising its full 32k context, which otherwise leaks
# into kernel-selection heuristics and wastes attention memory on padding.
# Raising it inflates memory per forward pass sharply.
# VAULTSPEC_RAG_EMBEDDING_MAX_SEQ_LENGTH=2048

# Character cap applied to each text before encoding (default: 8000, roughly
# 2000 BPE tokens). Text beyond it is truncated, not chunked, so raising this
# without raising EMBEDDING_MAX_SEQ_LENGTH gains nothing.
# VAULTSPEC_RAG_MAX_EMBED_CHARS=8000

# Heading-aware vault chunk budget in characters (default: 3000). One stored
# point per chunk; ~3000 chars is ~750 BPE tokens, well inside the encoder cap
# with the title header prepended. Changing it changes chunk boundaries, so
# rebuild the vault index afterwards.
# VAULTSPEC_RAG_VAULT_CHUNK_CHARS=3000

# Chars-per-token ratio used to turn the model's token window into a document
# chunk budget (default: 3). Lower it if a corpus tokenises denser than plain
# prose and document chunks are being truncated by the encoder.
# VAULTSPEC_RAG_DOCUMENT_CHUNK_CHARS_PER_TOKEN=3

# Characters carried across a document chunk boundary (default: 256), so a
# passage split across two chunks stays retrievable from either side. Changing
# it changes chunk boundaries, so rebuild the document index afterwards.
# VAULTSPEC_RAG_DOCUMENT_CHUNK_OVERLAP_CHARS=256

# ---------------------------------------------------------------------------
# Indexing throughput
# ---------------------------------------------------------------------------

# Worker processes for parallel codebase chunking (default: 0 = auto, resolved
# to the CPU count at run time). tree-sitter parsing is CPU-bound and holds the
# GIL, so this is a process pool. 1 forces the serial in-process path.
# VAULTSPEC_RAG_INDEX_CHUNK_WORKERS=0

# Minimum total source bytes before AUTO worker selection engages the process
# pool (default: 8388608, 8 MiB). Spawn workers cost ~0.3s each, so below this
# the auto path stays serial. An explicit INDEX_CHUNK_WORKERS >= 1 bypasses
# this gate entirely.
# VAULTSPEC_RAG_INDEX_PARALLEL_MIN_BYTES=8388608

# Flush the CUDA caching allocator every N codebase embed slices (default: 8).
# Per-slice flushing forces a device sync each iteration; throttling removes
# most of those syncs while bounding allocator growth to N slices' worth of
# transient activations. 1 restores per-slice flushing.
# VAULTSPEC_RAG_INDEX_CACHE_FLUSH_SLICES=8

# Same cadence for the VAULT embed path (default: 1, flush every slice).
# Per-slice flushing may be load-bearing against allocator fragmentation on a
# CUDA ceiling that counts reserved memory — raise it only after validating
# peak reserved memory on a real full rebuild.
# VAULTSPEC_RAG_VAULT_CACHE_FLUSH_SLICES=1

# Same cadence for the DOCUMENT embed path (default: 1). Same caveat as above.
# VAULTSPEC_RAG_DOCUMENT_CACHE_FLUSH_SLICES=1

# Reuse already-computed dense and sparse vectors from sibling namespaces for
# chunks whose stored content matches byte-for-byte (default: 1). Indexing a
# fork of an already-indexed tree then skips the encode that dominates a
# rebuild. Set to 0 to disable every donor lookup and encode every chunk — the
# A/B lever and the paranoia escape hatch.
# VAULTSPEC_RAG_INDEX_REUSE=1

# ---------------------------------------------------------------------------
# Index resource bounds and memory ceilings
# ---------------------------------------------------------------------------
# These bound one index run's segment and queue geometry, its memory use, and
# its liveness. The defaults suit a managed multi-root service; lower them on a
# smaller host.

# Chunks per index upsert segment (default: 64). A segment is the durable unit,
# so this also sets how much work a failed write can lose.
# VAULTSPEC_RAG_INDEX_SEGMENT_MAX_CHUNKS=64

# Byte cap per index upsert segment (default: 8388608, 8 MiB). Counts text,
# vectors, sparse vectors, payload, and Python container overhead.
# VAULTSPEC_RAG_INDEX_SEGMENT_MAX_BYTES=8388608

# Chunks buffered in the producer-to-consumer index queue (default: 512).
# VAULTSPEC_RAG_INDEX_QUEUE_MAX_CHUNKS=512

# Byte cap on the buffered index queue (default: 134217728, 128 MiB). Reaching
# either queue bound applies backpressure to the producer. The queue must admit
# at least one whole segment, so keep it above the segment bounds.
# VAULTSPEC_RAG_INDEX_QUEUE_MAX_BYTES=134217728

# Seconds without storage-confirmed durable progress before the run is failed
# (default: 900). This is a liveness bound, never a total run deadline: a
# healthy long index continues indefinitely, while a failed write ladder is
# contained.
# VAULTSPEC_RAG_INDEX_NO_PROGRESS_TIMEOUT_SECONDS=900

# Resident-memory ceiling in MiB enforced at index checkpoints
# (default: 16384).
# VAULTSPEC_RAG_INDEX_RSS_CEILING_MIB=16384

# CUDA-memory ceiling in MiB enforced at index checkpoints (default: 0 = auto,
# derived as total device memory minus INDEX_CUDA_HEADROOM_MIB). A positive
# value is an authoritative override that raises OR lowers the effective
# ceiling.
# VAULTSPEC_RAG_INDEX_CUDA_CEILING_MIB=0

# Memory in MiB reserved below the device total when the CUDA ceiling is
# auto-derived (default: 2048). Leaves room for the driver, concurrent search,
# and allocator fragmentation; on a 16 GiB card this yields a ~14 GiB indexing
# ceiling.
# VAULTSPEC_RAG_INDEX_CUDA_HEADROOM_MIB=2048

# Fraction of device memory the process-wide CUDA allocator may reserve
# (default: 0.8). Must be greater than 0 and at most 1; a larger value is
# rejected. Preserves device headroom for concurrent search before model load.
# VAULTSPEC_RAG_INDEX_CUDA_ALLOCATOR_FRACTION=0.8

# Free device memory in MiB required before this process loads model stacks
# onto the GPU (default: 0, meaning derive it). Read once per process before the
# first load; a card with less free than this refuses the load instead of
# starving every consumer on it. Zero derives the floor from the CUDA demand the
# configured support profile declares, so it tracks the workload rather than
# asserting one machine's figures. It does not read your card: model weights
# occupy what they occupy on any device, so what a load needs is a property of
# the models and not of the hardware. Set a positive value to override the
# derivation for your own device.
#
# Whatever you set must cover the resident stack a load creates PLUS the largest
# demand that stack then places on top of its own residency. A floor sized to
# residency alone still leaves enough free memory, on a card already holding one
# tenant, to admit a second onto a device that cannot hold both - which is the
# arrangement this gate exists to refuse. Set it too high and loads the card
# could have served are refused instead; on a small card a floor above total
# memory refuses every load.
# VAULTSPEC_RAG_GPU_ADMISSION_FLOOR_MIB=0

# Named indexing support profile (default: managed-service). Accepts
# managed-service or embedded-local; any other value is rejected on first read.
# VAULTSPEC_RAG_INDEX_SUPPORT_PROFILE=managed-service

# ---------------------------------------------------------------------------
# Concurrency limits
# ---------------------------------------------------------------------------
# Interactive searches and long-running index jobs draw from separate capacity
# limiters, so a reindex can never exhaust the threads that serve searches.
# Saturating a limiter queues callers rather than failing them.

# Search worker limiter (default: 16).
# VAULTSPEC_RAG_SEARCH_CONCURRENCY=16

# Index job limiter (default: 4). Raise it if the host has spare cores.
# VAULTSPEC_RAG_INDEX_JOB_CONCURRENCY=4

# ---------------------------------------------------------------------------
# Automatic index updates (filesystem watcher)
# ---------------------------------------------------------------------------

# Auto-reindex on file change (default: 1). 0 makes the service pull-only, so
# indexes only update when something asks for a reindex.
# CLI: --updates / --no-updates
# VAULTSPEC_RAG_WATCH_ENABLED=1

# Debounce window in milliseconds coalescing change events before a reindex
# (default: 2000). 0 means no delay, not disabled.
# CLI: --update-delay-ms
# VAULTSPEC_RAG_WATCH_DEBOUNCE_MS=2000

# Per-source cooldown in seconds after a completed auto-reindex before another
# may start (default: 30). 0 means no delay, not disabled.
# CLI: --repeat-update-delay-s
# VAULTSPEC_RAG_WATCH_COOLDOWN_S=30

# Lower bound of the adaptive coalescing window in seconds (default: 2).
# Cannot exceed the upper bound below.
# VAULTSPEC_RAG_WATCH_COALESCE_MIN_SECONDS=2

# Upper bound of the adaptive coalescing window in seconds (default: 30).
# Cannot exceed VAULTSPEC_RAG_WATCH_MAXIMUM_FRESHNESS_SECONDS.
# VAULTSPEC_RAG_WATCH_COALESCE_MAX_SECONDS=30

# Maximum adaptive post-success cooling delay in seconds (default: 120).
# VAULTSPEC_RAG_WATCH_COOLING_MAX_SECONDS=120

# Oldest ordinary-pressure event age allowed before admission (default: 300).
# VAULTSPEC_RAG_WATCH_MAXIMUM_FRESHNESS_SECONDS=300

# Maximum interval between service-measurement reevaluations (default: 5).
# VAULTSPEC_RAG_WATCH_MEASUREMENT_REEVALUATION_SECONDS=5

# Pending paths that make a controller immediately admission-ready
# (default: 10000; cannot exceed the durable scope path bound).
# VAULTSPEC_RAG_WATCH_BATCH_PATH_LIMIT=10000

# Most paths one controller's durable exact scope may hold (default: 100000).
# Exceeding it fails closed with a typed refusal.
# VAULTSPEC_RAG_WATCH_SCOPE_MAX_PATHS=100000

# Most bytes one controller's durable exact scope may occupy
# (default: 8388608, 8 MiB). Exceeding it fails closed the same way.
# VAULTSPEC_RAG_WATCH_SCOPE_MAX_BYTES=8388608

# Initial backoff in seconds before retrying a failed auto-reindex
# (default: 30).
# VAULTSPEC_RAG_WATCH_RETRY_BASE_SECONDS=30

# Maximum backoff in seconds between auto-reindex retries (default: 1800).
# VAULTSPEC_RAG_WATCH_RETRY_MAX_SECONDS=1800

# Symmetric random jitter fraction applied to each retry backoff
# (default: 0.1). Keeps several failing sources from retrying in lockstep.
# VAULTSPEC_RAG_WATCH_RETRY_JITTER_FRACTION=0.1

# Consecutive failures before the watch circuit opens and stops retrying a
# persistently failing source (default: 3).
# VAULTSPEC_RAG_WATCH_CIRCUIT_FAILURE_THRESHOLD=3

# ---------------------------------------------------------------------------
# Index integrity
# ---------------------------------------------------------------------------

# Queue a repair when a search finds the served index holding fewer points
# than its published manifest claims (default: 0). On, the service queues one
# deduplicated, non-destructive, failure-safe reindex for that domain. Off, the
# shrink is still detected and reported on the status surface, and nothing
# repairs it until something asks for a reindex.
# VAULTSPEC_RAG_INTEGRITY_AUTO_REPAIR=0

# ---------------------------------------------------------------------------
# Search and ranking
# ---------------------------------------------------------------------------

# Optional hosted query and result classification with TypeSafe Jev.
# A credential, not a setting: read from the process environment first, and a
# workspace-root .env supplies it only under the gate this file's header
# describes (this workspace's own interpreter, dependency or dev mode). The
# resident service is handed the resolved value through its own environment
# rather than reading a file itself. Set only in the environment of the
# process executing search (the server for service-backed searches). This
# sends queries and full candidate content to https://api.typesafe.ai and
# incurs API usage. No separate enable flag.
# A successful classification establishes current credential usability; invalid
# or unfunded credentials retain the existing search pipeline. Authentication
# and payment rejection disable calls until key rotation or process restart;
# transient failures use a cooldown and restore the original search behavior.
# Unset or blank means no TypeSafe calls. Never commit a credential here.
# VAULTSPEC_RAG_TYPESAFE_API_KEY=

# Default vault ranking intent when a search names none
# (default: orientation). "orientation" lifts active decisions and grounding;
# "debugging" inverts toward execution records and audits.
# VAULTSPEC_RAG_VAULT_INTENT_DEFAULT=orientation

# Apply the intent-aware prior to vault results (default: 1). 0 restores the
# bare-reranker ordering.
# VAULTSPEC_RAG_VAULT_INTENT_RANKING_ENABLED=1

# Maximum results of one doc type allowed on a vault page (default: 4).
# 0 disables the cap and lets one document type fill the page.
# VAULTSPEC_RAG_VAULT_INTENT_TYPE_CAP=4

# Code domains dropped from results by default
# (default: worktree,generated). Comma-separated; unrecognised labels are
# ignored so a typo cannot silently widen the policy. Hidden domains stay
# reachable per call via --include-domain.
# VAULTSPEC_RAG_CODE_NOISE_HIDE_DOMAINS=worktree,generated

# Code domains kept visible but ranked below production
# (default: tests,docs,locale,vendored). A domain listed in both sets is
# hidden, not merely demoted.
# VAULTSPEC_RAG_CODE_NOISE_DEMOTE_DOMAINS=tests,docs,locale,vendored

# Score subtracted from a demoted code result (default: 0.3). Calibrated
# rerank scores live in [0, 1], so this is large enough to sink noise below
# production near-ties and small enough to leave a demoted hit recoverable.
# 0 disables demotion.
# VAULTSPEC_RAG_CODE_NOISE_DEMOTE_PENALTY=0.3

# Collapse near-duplicate locale-variant code results to one representative
# (default: 1). Only recognised locale shapes within a score tie window
# collapse.
# CLI: --dedup-locales / --no-dedup-locales
# VAULTSPEC_RAG_DEDUP_LOCALES_DEFAULT=1

# Seconds a parsed vault graph stays cached (default: 300).
# VAULTSPEC_RAG_GRAPH_TTL_SECONDS=300

# ---------------------------------------------------------------------------
# Document preprocessing hooks
# ---------------------------------------------------------------------------

# Kill switch for a root's .vaultragpreprocess.toml rules (default: unset, so
# approved rules run). An ordinary boolean, using the same vocabulary and
# refusal as every other switch above: a false word (0, false, no, off) stops
# every root's rules from running and wins over everything else, so an
# operator can always silence a root's rules; a true word, blank, or unset
# leaves preprocessing on; anything else is refused. Note that a root's
# preprocess config is repo-authored code that runs with your privileges, so
# it runs only after `vaultspec-rag preprocess approve`; no variable approves.
# CLI: --no-preprocess
# VAULTSPEC_RAG_PREPROCESS=

# Cap in bytes on the text a preprocessor may emit for one file
# (default: 10485760, 10 MiB). The source-size cap is relaxed for matched
# files, so this is what keeps a runaway extractor from emitting tens of MB;
# a 12 MB PDF that distils to 40 KB still indexes.
# VAULTSPEC_RAG_PREPROCESS_MAX_EMITTED_BYTES=10485760

# Strip HTML tags to plain text before chunking .html sources (default: 1).
# Raw markup wastes about a third of each chunk's budget and pollutes results
# with navigation boilerplate. Falls back to raw-markup chunking on a parse
# error.
# VAULTSPEC_RAG_HTML_STRIP=1

# ---------------------------------------------------------------------------
# Diagnostics and process lifetime
# ---------------------------------------------------------------------------

# Diagnostic RSS + CUDA memory probe (default: unset, off). Uses the standard
# boolean parsing above in full, rejection included: 1/true/yes/on turn it on,
# 0/false/no/off and an empty value turn it off, and anything else is refused.
# When on, the probe records memory at named index checkpoints and emits a
# structured report.
# VAULTSPEC_RAG_MEMORY_PROBE=

# Ancestor-death backstop for the stdio MCP shim (default: unset, enabled).
# Set to 0, false, no, or off to disable it, after which stdin EOF is the shim's
# only exit path — an orphaned shim then survives its client. Falls back to the
# framework-wide VAULTSPEC_STDIO_WATCHDOG when unset. It accepts the same
# spellings as every other boolean, but resolves the leftovers the other
# way: a blank value and an unrecognised word both leave the backstop ARMED
# and log a warning, rather than falling through like every other switch.
# Losing this by accident is what strands shim processes, so only an explicit
# off-word disarms it.
# VAULTSPEC_RAG_STDIO_WATCHDOG=

# ---------------------------------------------------------------------------
# Third-party Hugging Face variables
# ---------------------------------------------------------------------------
# The variables below are NOT defined by vaultspec-rag. They belong to the
# Hugging Face Hub and Transformers libraries, which vaultspec-rag downloads
# its models through; it reads or forwards them but does not wrap them, and
# their behaviour is documented upstream at
# https://huggingface.co/docs/huggingface_hub/en/package_reference/environment_variables
# They are listed here because the project references them by name and because
# an offline or mirrored host has to set them.

# Hub cache root — where model files are stored
# (Hugging Face default: ~/.cache/huggingface).
# HF_HOME=

# Directory the model snapshots themselves are kept in, when it should not be
# the one under HF_HOME (Hugging Face default: <HF_HOME>/hub).
# HF_HUB_CACHE=

# Hub endpoint, read by the Hub client once, when it is first loaded
# (Hugging Face default: https://huggingface.co). Setting it directly works.
# VAULTSPEC_RAG_HF_ENDPOINT outranks it: when that is set, this is overwritten
# at process start; when it is not, this is left exactly as set.
# HF_ENDPOINT=

# The Hub client's per-read timeout: seconds with no data arriving before it
# abandons one download attempt (Hugging Face default: 10). Not a budget for
# a whole file. vaultspec-rag does not set it. The client retries a file up to
# five more times, counting again whenever data arrives, so a hub that goes
# silent fails a file after about a minute. Raise it for a link that pauses
# for longer than that.
# HF_HUB_DOWNLOAD_TIMEOUT=

# Offline mode: no network access to the Hub. This is the authoritative
# offline switch. Recognised as on for 1, true, yes, or on; install, server
# start and server warmup then fetch no model and report what the cache lacks.
# Loading a model never fetches one, whether or not this is set.
# HF_HUB_OFFLINE=

# Transformers' own offline switch, honoured the same way and with the same
# accepted values. Either variable being set stops every model fetch.
# TRANSFORMERS_OFFLINE=

# Skip on-the-fly conversion of model weights to safetensors. Model loading
# sets this to 1 if it is unset, so conversion never runs implicitly during
# startup.
# DISABLE_SAFETENSORS_CONVERSION=

# Turn off the parallel weight-materialising path. Model loading sets this to
# 1 on Windows if it is unset, where the parallel path was observed corrupting
# weights under concurrent loads; an explicit value here still wins.
# HF_DEACTIVATE_ASYNC_LOAD=

# Suppress the Hub's download progress bars. The search path sets this to 1 if
# it is unset, because its output is a result envelope rather than a console
# somebody is watching.
# HF_HUB_DISABLE_PROGRESS_BARS=

# Suppress Transformers' advisory warnings, set by the search path for the
# same reason.
# TRANSFORMERS_NO_ADVISORY_WARNINGS=

# Transformers log verbosity. The search path sets this to error if it is
# unset, so library chatter cannot reach a result envelope.
# TRANSFORMERS_VERBOSITY=

# ---------------------------------------------------------------------------
# Third-party PyTorch, uv and platform variables
# ---------------------------------------------------------------------------
# Owned by PyTorch, by uv, and by the operating system. Listed because
# vaultspec-rag reads them, never because it defines them: each keeps its
# owner's meaning, including which values that owner recognises.

# PyTorch's documented fallback from an unimplemented Metal operator to the
# processor. Read with PyTorch's own reading of it - the exact value 1 - so a
# word PyTorch ignores is not reported here as a fallback that is in place.
# PYTORCH_ENABLE_MPS_FALLBACK=

# uv's cache directory. Read only to tell an ephemeral cache environment apart
# from an installed tool when explaining a GPU failure, so the remediation
# names the right reinstall.
# UV_CACHE_DIR=

# uv's tool-install directory, read for the same classification.
# UV_TOOL_DIR=

# The Windows temporary directory. Read to decide whether an indexed root was
# somewhere throwaway, which selects how long an orphaned namespace waits
# before storage maintenance may reclaim it.
# TEMP=

# The second Windows temporary-directory name, read alongside TEMP for the
# same decision.
# TMP=

# The POSIX temporary-directory name, read for the same decision.
# TMPDIR=

# ---------------------------------------------------------------------------
# Interpreter environment (read, never set)
# ---------------------------------------------------------------------------
# Not a knob. Your virtualenv activator sets this, and vaultspec-rag only reads
# it to report WHICH interpreter is serving — it appears in the discovery
# pointer, the /health payload, and each job's runtime record, so an operator
# can tell a venv daemon from a system-interpreter one when two disagree.
# Setting it by hand here does not change which interpreter runs; it only makes
# those reports lie.
# VIRTUAL_ENV=

# ---------------------------------------------------------------------------
# Interpreter safe-path mode (set while indexing, honoured when you set it)
# ---------------------------------------------------------------------------
# Python's own switch that keeps the working directory off a process's import
# path. Indexing runs from inside the project being indexed, so vaultspec-rag
# sets this to 1 in its own environment while an indexing worker pool is open
# and puts back what was there when the pool closes; a pool worker then cannot
# import a file from the directory the command was run in. Nothing needs
# setting for that. Export it as 1 in the shell before a command starts to run
# every vaultspec-rag process in safe-path mode; the interpreter reads it once
# at startup, and the tool then leaves it alone.
# PYTHONSAFEPATH=

# ---------------------------------------------------------------------------
# Program search path (read, never set)
# ---------------------------------------------------------------------------
# Not a knob. vaultspec-rag runs a few programs it does not ship - uv, the
# graphics driver's nvidia-smi, the compiled monitor when it is not installed
# beside the commands - and finds them on your search path. Only the ABSOLUTE
# entries are searched. An empty or relative entry means "the directory the
# command was typed in", and commands are typed inside project checkouts, so a
# program is never run from there. Operating-system tools are run from the
# operating system's own directory and do not consult this at all.
# PATH=

# The extensions a program name may carry on Windows, read alongside PATH.
# PATHEXT=

# ---------------------------------------------------------------------------
# Development harness
# ---------------------------------------------------------------------------
# Read by the `just` recipes, the modules under dev/ and tools/, and the test
# harness — never by the installed product. Like everything above except the
# one credential, they are read from the session environment: set them in the
# shell or the CI job that runs the recipe. A line in .env does nothing.

# Stream `just init` progress as NDJSON events on stdout (default: unset,
# human-readable text). Accepts 1, true, yes or on. The recipes take no
# arguments, so this is how a caller asks for machine-readable output.
# VAULTSPEC_INIT_JSON=

# Ignore the initialization stamp and run every phase again
# (default: unset). Accepts 1, true, yes or on.
# VAULTSPEC_INIT_FORCE=

# How much of the local inference stack `just init-python` installs
# (default: full). full installs the gpu dependency group with the CUDA
# runtime torch pulls in; types installs the group without that runtime, so
# type checkers can read the libraries on a host that never runs them; none
# omits the group, which is several gigabytes. Any other value is refused
# rather than read as full.
# VAULTSPEC_INIT_GPU_STACK=full

# Directory the gates write report artifacts into (default: unset, nothing is
# written anywhere). Setting it has the dependency audit write its report
# there and every pytest run write a JUnit file there.
# VAULTSPEC_CI_REPORTS=

# Filename stem of the JUnit report one pytest run writes into the directory
# above (default: unset — a short digest of the invocation, so two lanes in
# one job do not overwrite each other). An explicit --junitxml always wins.
# VAULTSPEC_CI_REPORT_NAME=

# uv's own setting for where a project's environment lives (uv's default:
# .venv in the project). A test session leaves a value set here alone. With
# none set, a session started from any interpreter other than the checkout's
# own .venv names a directory under its temporary tree for as long as it
# runs, so that no uv process a test starts, however indirectly, can create
# or replace the checkout's environment. A session started from the
# checkout's own environment changes nothing.
# UV_PROJECT_ENVIRONMENT=

# Admit the latency and footprint lane into the test recipes (default: unset,
# the lane is reported skipped). Any non-empty value opts in. The lane's
# wall-clock assertions are its system under test, so run it only on a machine
# doing nothing else.
# VAULTSPEC_RAG_PERF_LANE=

# Absolute path to an installed Chromium-family browser for the monitor
# release smoke check (default: unset — google-chrome, chromium,
# chromium-browser and msedge are looked up on PATH). A relative path or a
# missing file is refused; the check never downloads a browser.
# CHROME_BIN=

# The conventional continuous-integration marker (default: unset). Any
# non-empty value makes the monitor dev-server harness run in its CI mode and
# skip registering local proxy names.
# CI=

# Windows per-user application data directory. The monitor dev-server harness
# keeps its service records under it, outside every checkout
# (falls back to ~/AppData/Local).
# LOCALAPPDATA=

# The XDG state directory, which serves the same purpose off Windows
# (falls back to ~/.local/state).
# XDG_STATE_HOME=

# The Windows installation directory. The monitor's offline release probe
# locates the system PowerShell under it and fails without it.
# SYSTEMROOT=

# Encoding of a Python child's standard streams. The harness sets utf-8 when
# this is unset, because a stock Windows console codepage cannot encode the
# characters the toolchain prints; a value set here wins.
# PYTHONIOENCODING=utf-8

# The Windows account name. The citation gate reads it to learn the identity
# of the machine it runs on, so it can refuse that identity appearing in
# tracked files.
# USERNAME=

# The POSIX account name, read by the citation gate for the same purpose.
# USER=

# The POSIX login name, read by the citation gate for the same purpose.
# LOGNAME=
