Metadata-Version: 2.4
Name: schema-sanitizer
Version: 0.2.2
Summary: Data sanitization for CSV, JSON, JSONL, XML, Parquet, and Python objects incremental pipelines.
Keywords: arrow,pyarrow,json,xml,csv,schema,sanitization
Author: bgallan
License-Expression: Apache-2.0
License-File: LICENSE
Classifier: Development Status :: 3 - Alpha
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS :: MacOS X
Classifier: Operating System :: Microsoft :: Windows
Classifier: Programming Language :: C++
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: Implementation :: CPython
Project-URL: Homepage, https://github.com/bgallan/schema-sanitizer
Project-URL: Repository, https://github.com/bgallan/schema-sanitizer
Project-URL: Changelog, https://github.com/bgallan/schema-sanitizer/releases
Project-URL: Issues, https://github.com/bgallan/schema-sanitizer/issues
Requires-Python: >=3.11
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: pre-commit>=3.7; extra == "dev"
Requires-Dist: build>=1.2; extra == "dev"
Requires-Dist: ruff>=0.6.0; extra == "dev"
Requires-Dist: mypy>=1.10; extra == "dev"
Requires-Dist: pyarrow>=14.0.0; extra == "dev"
Requires-Dist: pandas>=2.0; extra == "dev"
Requires-Dist: polars>=0.20; extra == "dev"
Requires-Dist: duckdb>=1.0; extra == "dev"
Provides-Extra: pyarrow
Requires-Dist: pyarrow>=14.0.0; extra == "pyarrow"
Provides-Extra: polars
Requires-Dist: polars>=0.20; extra == "polars"
Provides-Extra: pandas
Requires-Dist: pandas>=2.0; extra == "pandas"
Provides-Extra: duckdb
Requires-Dist: duckdb>=1.0; extra == "duckdb"
Provides-Extra: all
Requires-Dist: pyarrow>=14.0.0; extra == "all"
Requires-Dist: polars>=0.20; extra == "all"
Requires-Dist: pandas>=2.0; extra == "all"
Requires-Dist: duckdb>=1.0; extra == "all"
Description-Content-Type: text/markdown

# schema-sanitizer

**Version 0.2.2:** this project is still being tuned and tested, especially for
generating Parquet files used by BigQuery external tables.

`schema-sanitizer` converts messy CSV, JSON, JSON Lines, NDJSON, XML, and
Parquet data into stable analytical tables or sanitized files. The native C++23
core handles schema inference, scalar/container reconciliation, field
versioning, bounded streaming, and Arrow C Data materialization.

<a id="index"></a>

## Index

- [Install](#install)
- [Public API](#public-api)
- [Input Formats](#input-formats)
- [Input Mode](#input-mode)
- [Shared Parameters](#shared-parameters)
- [Paths And Input Selection](#paths-and-input-selection)
- [Schema And Field Handling](#schema-and-field-handling)
- [String Scalar Parsing](#string-scalar-parsing)
- [Source-Specific Parsing](#source-specific-parsing)
- [Errors And Resources](#errors-and-resources)
- [Configuration Examples](#configuration-examples)
- [Result](#result)
- [ETL Generated Columns](#etl-generated-columns)
- [Schema Reconciliation](#schema-reconciliation)
- [Field Names](#field-names)
- [Timestamp Precision](#timestamp-precision)
- [Depth Limits](#depth-limits)
- [Memory Safety And Tuning](#memory-safety-and-tuning)
- [Filesystems](#filesystems)
- [Example 7](#example-7)
- [Development](#development)
- [License](#license)

## [Install](#index)

```bash
pip install 'schema-sanitizer[pyarrow]'
```

Optional analytical targets:

```bash
pip install 'schema-sanitizer[pandas]'
pip install 'schema-sanitizer[polars]'
pip install 'schema-sanitizer[duckdb]'
pip install 'schema-sanitizer[all]'
```

```python
import schema_sanitizer as ss
```

## [Public API](#index)

All public operations are named `to_*`.

In-memory analytical functions:

| Function | `Result.clean_data` |
|---|---|
| `to_pyarrow(...)` | `pyarrow.Table` |
| `to_pandas(...)` | `pandas.DataFrame` |
| `to_polars(...)` | `polars.DataFrame` |
| `to_duckdb(...)` | DuckDB relation |

File-to-file functions:

| Function | Output |
|---|---|
| `to_csv(input_path, output_path, ...)` | CSV file |
| `to_jsonl(input_path, output_path, ...)` | JSON Lines file |
| `to_parquet(input_path, output_path, ...)` | Parquet file |

```python
events = ss.to_pyarrow(
    "raw/events.jsonl",
    input_format="jsonl",
)

customers = ss.to_pandas(
    "raw/customers.csv",
    input_format="csv",
)

ss.to_parquet(
    "raw/events.jsonl",
    "silver/events.parquet",
    input_format="jsonl",
)
```

All seven functions expose the same input and cleaning options. File-to-file
functions additionally take `output_path`.

## [Input Formats](#index)

`input_format` must always be selected explicitly. The signature default is
`None`, but calling any `to_*` function with `None` raises an error. Neither
`None` nor `"auto"` infers a format from the extension or file contents.

The selected format also validates the source extension. For `.json` files,
choose `"json"` for one document treated as one row or `"json_array"` for a
top-level array of row objects.

| `input_format` | Required extension | Content |
|---|---|---|
| `csv` | `.csv` | Delimited rows |
| `json` | `.json` | One JSON document treated as one source row |
| `json_array` | `.json` | Top-level array containing JSON objects |
| `jsonl` | `.jsonl` | One JSON object per line |
| `ndjson` | `.ndjson` | One JSON object per line |
| `xml` | `.xml` | XML document or streamed `xml_row_tag` elements |
| `parquet` | `.parquet` or `.pq` | Parquet rows |

`jsonl` and `ndjson` use the same newline-delimited JSON parser. Their only
difference is the required extension.

Valid JSONL or NDJSON:

```json
{"a": 1}
{"a": 2}
```

Valid `json_array`:

```json
[
  {"id": 1, "name": "Ana"},
  {"id": 2, "name": "Luis"},
  {"id": 3, "name": "Marta"}
]
```

Every top-level `json_array` element must be an object. The array is split
incrementally into rows instead of being materialized as one nested value.

Passing a mismatched extension fails before ingestion:

```python
# Raises: jsonl requires .jsonl, not .ndjson
ss.to_pyarrow("events.ndjson", input_format="jsonl")
```

## [Input Mode](#index)

`input_mode` accepts:

| Value | Behavior |
|---|---|
| `single_file` | Default. Process exactly one source file. |
| `directory` | Process matching direct child files in deterministic filename order. |

Directory traversal is non-recursive. Files with other extensions and nested
directories are ignored.

```python
table = ss.to_pyarrow(
    "raw/2026-01/",
    input_format="jsonl",
    input_mode="directory",
).clean_data
```

Directory behavior:

- `jsonl` reads only direct `.jsonl` children.
- `ndjson` reads only direct `.ndjson` children.
- `json` reads direct `.json` documents as rows.
- `json_array` flattens each direct `.json` array into rows.
- `csv` removes repeated matching headers and rejects header mismatches.
- `xml` combines direct `.xml` documents and requires a compatible root/row tag.
- `parquet` streams direct `.parquet` and `.pq` children.

Directory mode requires an explicit `input_format`.

## [Shared Parameters](#index)

```python
result = ss.to_pyarrow(
    input_path,
    input_format="jsonl",
    input_mode="single_file",
    schema_mode="additive",
    column_order="alphabetically",
    field_name_policy="lower_alpha",
    timestamp_precision="TIMESTAMP_MICROS",
    parse_integers=False,
    parse_floats=False,
    parse_float_decimal_separator=".",
    parse_float_thousands_separator=",",
    parse_iso_timestamps=False,
    parse_iso_dates=False,
    parse_iso_times=False,
    true_tokens=(),
    false_tokens=(),
    custom_timestamp_patterns=(),
    custom_date_patterns=(),
    custom_time_patterns=(),
    arrow_max_depth=32,
    parquet_max_depth=15,
    scalar_object_key="default_key",
    csv_has_header=True,
    csv_delimiter=",",
    input_text_encoding="utf-8",
    xml_row_tag=None,
    on_error="emit_null_row",
    batch_memory_limit_bytes=None,
    read_chunk_bytes=1024 * 1024,
    schema_registry=None,
)
```

### [Paths And Input Selection](#index)

| Parameter | Default | Accepted values / example | Use |
|---|---|---|---|
| `input_path` | Required | `"events.jsonl"`, `Path("events.csv")`, `"gs://bucket/events.jsonl"` | Source file or directory. Local paths and supported PyArrow filesystem URIs are accepted. |
| `output_path` | Required for file sinks | `"events.parquet"`, `"s3://bucket/events.jsonl"` | Destination used only by `to_csv`, `to_jsonl`, and `to_parquet`. |
| `input_format` | `None` (raises) | `"csv"`, `"json"`, `"json_array"`, `"jsonl"`, `"ndjson"`, `"xml"`, `"parquet"` | Required parser selection. The default `None` and `"auto"` are rejected. The selected format validates the source extension. |
| `input_mode` | `"single_file"` | `"single_file"`, `"directory"` | Process one source file or all matching direct children of one directory. Directory traversal is non-recursive. |

### [Schema And Field Handling](#index)

| Parameter | Default | Accepted values / example | Use |
|---|---|---|---|
| `schema_mode` | `"additive"` | `"additive"`, `"strict"` | `additive` preserves the registry contract and adds compatible fields or versions. `strict` rejects incompatible input and requires a registry-derived schema. |
| `column_order` | `"alphabetically"` | `"alphabetically"`, `"schema_contract_first"` | Order fields recursively. `schema_contract_first` keeps registered fields first and appends new fields deterministically. |
| `field_name_policy` | `"lower_alpha"` | `"lower_alpha"`, `"lower_snake"`, `"preserve"` | Sanitize every field name. `lower_alpha` keeps lowercase `a-z`; `lower_snake` also keeps digits and `_`; `preserve` retains source spelling. |
| `scalar_object_key` | `"default_key"` | `"value"`, `"raw_value"` | Child field used when reconciling a scalar with a struct, for example `5` becomes `{"default_key": 5}`. The name is processed by the selected field-name policy. |
| `arrow_max_depth` | `32` | `8`, `16`, `32` | Maximum expanded Arrow container depth. Structs and lists count; deeper values are flattened to string-compatible output. |
| `parquet_max_depth` | `15` | `8`, `12`, `15` | Maximum Parquet/BigQuery RECORD depth. List wrappers do not add a RECORD level. |
| `schema_registry` | `None` | Python mapping, registry JSON string, or `None` | Previous registry used as the source of truth for incremental conversion and historical reprocessing. `None` starts a new registry. |

### [String Scalar Parsing](#index)

These options apply to string values such as CSV cells, XML text, and quoted
JSON values. Actual JSON numbers and booleans are already typed by JSON syntax
and do not depend on these options.

| Parameter | Default | Accepted values / example | Use |
|---|---|---|---|
| `parse_integers` | `False` | `True`, `False` | Convert integer-looking strings such as `"42"` and `"-7"` to `int64`. |
| `parse_floats` | `False` | `True`, `False` | Convert float-looking strings such as `"12.5"` or `"1,234.56"` to `float64`. |
| `parse_float_decimal_separator` | `"."` | `"."`, `","` | Decimal separator used when `parse_floats=True`. Must be one ASCII punctuation character. |
| `parse_float_thousands_separator` | `","` | `","`, `"."`, `"_"` | Optional grouping separator used when `parse_floats=True`. It must differ from the decimal separator and grouped sections must contain exactly three digits. |
| `true_tokens` | `()` | `("true", "yes", "y")` | Case-insensitive string tokens converted to Boolean `True`. An empty sequence disables custom string-to-Boolean parsing. |
| `false_tokens` | `()` | `("false", "no", "n")` | Case-insensitive string tokens converted to Boolean `False`. True and false token sets must not overlap. |
| `parse_iso_timestamps` | `False` | `True`, `False` | Parse built-in ISO timestamps such as `"2026-01-02T03:04:05Z"` or `"2026-01-02 03:04:05+01:00"`. |
| `parse_iso_dates` | `False` | `True`, `False` | Parse built-in ISO dates in `YYYY-MM-DD` form. |
| `parse_iso_times` | `False` | `True`, `False` | Parse built-in ISO times in `HH:MM:SS` form. |
| `custom_timestamp_patterns` | `()` | `(r"(\d{4})/(\d{2})/(\d{2}) (\d{2}):(\d{2}):(\d{2})",)` | Additional timestamp patterns. Capture groups 1-6 represent year, month, day, hour, minute, and second; optional groups 7 and 8 represent fraction and timezone. |
| `custom_date_patterns` | `()` | `(r"(\d{4})#(\d{2})#(\d{2})",)` | Additional date patterns. Capture groups 1-3 represent year, month, and day. |
| `custom_time_patterns` | `()` | `(r"(\d{2})\|(\d{2})\|(\d{2})",)` | Additional time patterns. Capture groups 1-3 represent hour, minute, and second. |
| `timestamp_precision` | `"TIMESTAMP_MICROS"` | `"TIMESTAMP_MILLIS"`, `"TIMESTAMP_MICROS"`, `"TIMESTAMP_NANOS"` | Arrow and Parquet unit used after timestamp parsing. Microseconds are the BigQuery-compatible default. |

### [Source-Specific Parsing](#index)

| Parameter | Default | Accepted values / example | Use |
|---|---|---|---|
| `csv_has_header` | `True` | `True`, `False` | Treat the first CSV row as field names. In directory mode, repeated matching headers are removed. |
| `csv_delimiter` | `","` | `","`, `";"`, `"\t"`, `"|"` | One-character CSV delimiter. Values containing the delimiter must be quoted according to CSV rules. |
| `input_text_encoding` | `"utf-8"` | `"utf-8"`, `"utf-16"`, `"latin-1"` | Decode text inputs. Python codec names and aliases are accepted and normalized. It does not affect Parquet input. |
| `xml_row_tag` | `None` | `None`, `"row"`, `"item"` | Stream each direct matching XML element as one row. `None` treats the complete XML document as one row. |

### [Errors And Resources](#index)

| Parameter | Default | Accepted values / example | Use |
|---|---|---|---|
| `on_error` | `"emit_null_row"` | `"stop"`, `"skip_row"`, `"emit_null_row"` | Stop immediately, drop an offending row, or retain it while writing null for fields that cannot be materialized. |
| `batch_memory_limit_bytes` | `None` | `64 * 1024 * 1024`, `256 * 1024 * 1024`, `None` | Best-effort native inference/materialization budget per batch or document. Lower values reduce peak memory at a possible throughput cost. |
| `read_chunk_bytes` | `1024 * 1024` | `256 * 1024`, `4 * 1024 * 1024` | Streaming source read-buffer size. Smaller chunks use less transient memory and perform more reads. |

ISO timestamp, date, and time parsing is opt-in. With all three `parse_iso_*`
flags left at `False`, ISO-looking source strings remain strings. The
`custom_*_patterns` options are independent: configured custom patterns are
still applied even when the corresponding built-in ISO parser is disabled.

Float separator options apply only when `parse_floats=True` and only to string
values, including CSV cells and XML text. Real JSON numbers always use JSON's
`.` decimal syntax. Grouping is strict: the default configuration accepts
`"1,234.56"`, while European input can use:

```python
result = ss.to_pyarrow(
    "prices.csv",
    input_format="csv",
    parse_floats=True,
    parse_float_decimal_separator=",",
    parse_float_thousands_separator=".",
)
```

That configuration accepts `"1.234,56"` and `"1234,56"`. Grouped sections
after the first must contain exactly three digits. In comma-delimited CSV,
values containing commas must be quoted.

Enabled string-to-scalar parsers first test the source string unchanged. If
that strict attempt fails, they retry once after removing surrounding ASCII
spaces, tabs, line breaks, form feeds, and vertical tabs. This applies to
integer, float, Boolean-token, ISO temporal, and custom temporal parsing. The
retry does not allocate or modify the source value:

```text
" 123456"       -> 123456       when parse_integers=True
" yes "         -> true         when "yes" is a true token
" 2026-01-02 "  -> 2026-01-02   when parse_iso_dates=True
```

An unmatched string retains its original whitespace, and a whitespace-only
string remains a string. Exact configured values are tested before trimming,
so custom tokens or temporal patterns that intentionally include surrounding
whitespace continue to work.

### [Configuration Examples](#index)

European numeric and semicolon-delimited CSV:

```python
prices = ss.to_pyarrow(
    "prices.csv",
    input_format="csv",
    csv_delimiter=";",
    parse_floats=True,
    parse_float_decimal_separator=",",
    parse_float_thousands_separator=".",
).clean_data
```

Custom Boolean and temporal strings:

```python
events = ss.to_pandas(
    "events.ndjson",
    input_format="ndjson",
    true_tokens=("yes", "active"),
    false_tokens=("no", "inactive"),
    parse_iso_timestamps=True,
    custom_date_patterns=(r"(\d{4})-(\d{2})-(\d{2})",),
).clean_data
```

Strict incremental conversion using an existing registry:

```python
result = ss.to_parquet(
    "raw/events.jsonl",
    "silver/events.parquet",
    input_format="jsonl",
    schema_mode="strict",
    schema_registry=previous_result.schema_registry,
    on_error="stop",
)
```

Memory-first processing of a large directory:

```python
result = ss.to_parquet(
    "raw/2026-01/",
    "silver/2026-01.parquet",
    input_format="jsonl",
    input_mode="directory",
    batch_memory_limit_bytes=64 * 1024 * 1024,
    read_chunk_bytes=256 * 1024,
)
```

`to_csv`, `to_jsonl`, and `to_parquet` return `Result.clean_data is None`.
Analytical functions return their named in-memory object.

## [Result](#index)

Every public function returns `schema_sanitizer.Result`.

| Property | Description |
|---|---|
| `clean_data` | Analytical object, or `None` for file outputs |
| `stats` | Inference, materialization, batching, depth, and error counters |
| `schema_registry` / `schema_registry_json` | Updated registry state |
| `schema_drifts` / `schema_drifts_json` | Drift events generated by this run |

Analytical and file outputs use the same registry-backed native path, so they
produce the same schema and metadata behavior.

## [ETL Generated Columns](#index)

Every analytical and file conversion adds these fixed top-level columns:

| Column | Behavior | Generic first-row value |
|---|---|---|
| `source_file` | Full local/cloud file path, or the input directory path in directory mode | `"gs://example-bucket/raw/2026-06-25/events.jsonl"` |
| `ingestion_timestamp` | Native UTC timestamp captured when the output file materializes | `"2026-06-25T09:05:08.947122Z"` |
| `schema_registry` | Canonical schema and field-version registry serialized as JSON | `{"registry_version":1,"schema_generation":2,...}` |
| `schema_drifts` | Drift events generated for this input, serialized as JSON | `[{"source_path":"amount","output_name":"amount_v2_float",...}]` |

These columns contain values only in the first output row. Remaining rows are
null to avoid repeating large registry payloads.

Generic `schema_registry` value:

```json
{
  "field_name_policy": "lower_snake",
  "registry_version": 1,
  "schema_generation": 2,
  "canonical_schema": {
    "fields": [
      {
        "name": "amount",
        "nullable": true,
        "type": {"kind": "string"}
      }
    ]
  },
  "variants": {
    "amount": {
      "versions": [
        {
          "output_name": "amount",
          "schema": "string",
          "is_most_compatible_current_version": true
        }
      ]
    }
  }
}
```

Generic `schema_drifts` value after `amount` acquires a float alternative:

```json
[
  {
    "detected_at": "2026-06-25T09:05:08.947122Z",
    "source_path": "amount",
    "output_name": "amount_v2_float",
    "drift_type": "new_version_generated",
    "previous_schema": "string",
    "new_schema": "double"
  }
]
```

The names are part of the ETL output contract and cannot be configured. They
are reserved at the top level: conversion fails before writing if the source
schema already contains any of them. A nested source key such as
`payload.source_file` is allowed because it does not conflict with the
generated top-level columns; normal field-name sanitization still applies to
that nested key. Reserved top-level source fields are not renamed or versioned,
since silently doing so would make downstream registry discovery ambiguous.

## [Schema Reconciliation](#index)

The embedded `schema_registry` is the source of truth for incremental
processing. Pass the latest registry to the next conversion:

```python
result = ss.to_parquet(
    "raw/2026-01-09/events.jsonl",
    "silver/2026-01-09/events.parquet",
    input_format="jsonl",
    schema_registry=previous_registry,
)

next_registry = result.schema_registry
```

Before generating a field version, the native merge attempts compatible
reconciliation:

- A singleton can be wrapped into an existing list.
- A scalar can be wrapped into an existing struct under `default_key`.
- Empty objects and lists provide no schema-inference evidence. If no other
  value or registry entry defines the field, the field is omitted. If the
  field is already established, the empty container materializes as null.
- New compatible struct children are added as nullable fields.

This rule applies recursively. Empty nested fields do not create child columns,
affect sibling-name collision handling, trigger strict-schema extra-field
errors, or generate schema drift. Empty elements inside an established list
become null elements so list positions remain stable. Typed Parquet input keeps
its declared columns on the direct Arrow path, but empty container values still
become null and do not create additional type versions.

Irreconcilable drift creates a hybrid
`<original_name>_v<version>_<semantic_type>` field at the lowest incompatible
schema level. The original field remains unsuffixed and is version 1.

```text
sentiment_analysis: struct<...>
sentiment_analysis_v2_struct_array: list<struct<
  magnitude: double,
  magnitude_v2_string: string
>>
```

The numeric component guarantees uniqueness and records discovery order within
the registry. The semantic component describes the new logical type:

| Logical type | Semantic suffix |
|---|---|
| Boolean | `boolean` |
| 64-bit integer | `integer` |
| 64-bit float | `float` |
| String | `string` |
| Timestamp | `timestamp` |
| Date | `date` |
| Time | `time` |
| Struct | `struct` |
| List | `<element_type>_array` |

For example, list types produce `integer_array`, `struct_array`, or
`integer_array_array`. If two incompatible alternatives have the same semantic
type, their numeric versions still keep the columns distinct, such as
`payload_v2_struct_array` and `payload_v3_struct_array`.

Existing exact historical variants are preferred during past-date
reprocessing. Otherwise the newest compatible container is evolved
recursively. Repeating an already known shape does not increment
`schema_generation`.

Materialization routes each non-null source value to exactly one
most-compatible member of its version family. It does not always choose the
latest version:

- Arrays prefer list variants.
- Numeric values prefer numeric scalar variants.
- Ordinary strings prefer string variants.
- Parse-enabled numeric and temporal strings can target typed variants.
- A singleton can target a list variant and be wrapped as one element.
- Exact compatibility wins over fallback string conversion.
- If multiple versions receive the same compatibility score, the highest
  `_vN_...` version wins.
- A null source value leaves every member of the family null.

Given this family:

```text
amount: string
amount_v2_integer: int64
amount_v3_float: double
```

values are routed as follows:

| Source value | Destination |
|---|---|
| `"unknown"` | `amount` |
| `7` | `amount_v2_integer` |
| `2.5` | `amount_v3_float` |
| `"7"` with `parse_integers=True` | `amount_v2_integer` |
| `"7"` with integer parsing disabled | `amount` |
| `null` | All three columns remain null |

For a container family containing `items: struct<...>` and
`items_v2_struct_array: list<struct<...>>`, both an array and a compatible
singleton object go to `items_v2_struct_array`; the singleton is wrapped into a
one-element list. Other family columns are null in that row.

Each drift event receives a native UTC `detected_at` timestamp. The same
timestamp is written to the output file's first-row `ingestion_timestamp`, so
the materialized file and its drift audit events share one conversion time.
Source partition identity remains available through `source_file` and any Hive
partition columns, including during historical reprocessing.

## [Field Names](#index)

`field_name_policy="lower_alpha"` keeps lowercase `a-z` only.
`lower_snake` keeps lowercase letters, digits, and underscores. `preserve`
keeps source names.

Collisions use deterministic suffixes derived from the original dirty key, so
source field order does not change the dirty-key to clean-key mapping.

## [Timestamp Precision](#index)

Accepted values:

- `TIMESTAMP_MILLIS`
- `TIMESTAMP_MICROS` (default)
- `TIMESTAMP_NANOS`

Microseconds are the default because BigQuery external tables support Parquet
timestamp micros. BigQuery does not accept Parquet `TIMESTAMP_NANOS`.

## [Depth Limits](#index)

`arrow_max_depth` counts struct and list containers. `parquet_max_depth` counts
Parquet/BigQuery RECORD levels; list wrappers do not add a RECORD level.

Over-depth nested values are flattened to string-compatible output rather than
allowing unbounded schema expansion.

## [Memory Safety And Tuning](#index)

The pipeline uses replayable streaming sources, bounded inference batches, and
streaming file writers. `batch_memory_limit_bytes` controls the approximate
per-batch budget.

Memory-first settings for large files:

```python
ss.to_parquet(
    "raw/large.jsonl",
    "silver/large.parquet",
    input_format="jsonl",
    batch_memory_limit_bytes=64 * 1024 * 1024,
    read_chunk_bytes=256 * 1024,
)
```

`64 * 1024 * 1024` is 64 MiB.

Trade-offs:

- Lower `batch_memory_limit_bytes` reduces peak memory and may reduce speed.
- Lower `read_chunk_bytes` reduces transient input buffers and increases read calls.
- Parquet decoding enables threads only when the memory budget is large enough.
- Directory mode processes direct child files incrementally rather than loading
  the full directory at once.
- CSV directory normalization holds at most one configured-size source file in
  memory while validating and removing repeated headers.

## [Filesystems](#index)

Input and output paths may be local paths or PyArrow filesystem URIs such as:

```text
file:///data/events.jsonl
s3://bucket/events/2026-01-09/events.jsonl
gs://bucket/events/2026-01-09/events.jsonl
abfs://container/events/2026-01-09/events.jsonl
```

Directory listing uses the same filesystem and is non-recursive.

## [Example 7](#index)

`examples/example_07/07_gcs_jsonl_to_silver_parquet_range_prefix.py` implements
a single-writer GCS-to-Parquet pipeline with:

- CLI-selected `input_format`: `csv`, `json`, `json_array`, `jsonl`, `ndjson`,
  `xml`, or `parquet`
- CLI-selected `input_mode`: `single_file` or non-recursive `directory`
- daily `year=YYYY/month=MM/date=YYYY-MM-DD` partitions
- hourly `year=YYYY/month=MM/date=YYYY-MM-DD/hour=HH` partitions
- source extension validation derived from `input_format`
- integer, float, ISO timestamp, ISO date, and ISO time string parsing enabled
- source discovery and empty/missing partition skipping
- one sanitized Parquet output per logical partition
- embedded registry retrieval through Arrow ADBC
- incremental and random past-date reprocessing
- one final BigQuery external-table create/replace operation

In `directory` mode, all direct files matching `input_format` inside a source
Hive partition are combined into that partition's single output Parquet.
Subdirectories are not scanned.

Daily single-file layout:

```text
source-prefix/year=2026/month=06/date=2026-06-25/events_20260625.json
silver-prefix/year=2026/month=06/date=2026-06-25/events_20260625.parquet
```

Hourly directory layout:

```text
source-prefix/year=2026/month=06/date=2026-06-25/hour=08/*.jsonl
silver-prefix/year=2026/month=06/date=2026-06-25/hour=08/events_20260625_08.parquet
```

Use `--start-hour` and `--end-hour` to restrict the hourly partitions processed
for every selected date. Their defaults are `0` and `23`.

## [Development](#index)

```bash
pip install -e .[dev]
pytest
```

Native build:

```bash
cmake -S . -B build/dev -G Ninja -DCMAKE_BUILD_TYPE=Release
cmake --build build/dev
```

## [License](#index)

Apache License 2.0. See [`LICENSE`](LICENSE).
