Column report · {{ report.source.filename }}{% if report.source.format in ('json', 'jsonl') %} · {{ ds.id }}{% endif %}

{{ column.name or '(unnamed)' }}

  • Column {{ column.position }} of {{ ds.summary.column_count }}
  • {{ column.path }}
  • {{ ds.summary.row_count | count }} {{ 'records' if report.source.format in ('json', 'jsonl') else 'rows' }}
← Full report

01 / Overview

Column overview

Profile, values and formats at a glance.

Analyzable values
{{ column.scan_details.value_count | count if column.scan_details else (ds.summary.row_count - column.missing_count) | count }}in {{ ds.summary.row_count | count }} {{ 'records' if report.source.format in ('json', 'jsonl') else 'rows' }}
Missing
{{ column.missing_percent }}%{{ column.missing_count | count }} values
Distinct
{% if column.distinct_count is none %}limited{% else %}{{ column.distinct_count | count }}{% endif %}{{ column.value_profile.selection | replace('_',' ') }}
Type errors
{{ column.type_error_count | count if column.type_error_count is not none else '–' }}{{ column.type_error_percent ~ '%' if column.type_error_percent is not none else 'Not measured' }}

Interpretation

Inferred type
{{ column.inferred_type }}
Semantic type
{% if column.semantic_type %}{{ column.semantic_type | replace('_',' ') }}{% else %}None{% endif %}
Confidence
{{ ((column.type_confidence * 100) | round(1)) ~ '%' if column.type_confidence is not none else 'Not applicable' }}
Exposure
{{ column.exposure or 'Normal' }}

Value coverage

{{ column.value_profile.sampled_distinct_count | count }} stored distinct values · {{ column.value_profile.selection | replace('_',' ') }}{% if column.distinct_count is none %} · cardinality limited{% endif %}.

{% if column.semantic_type == 'enumeration' %}

{{ 'All visible enumeration values' if column.value_profile.selection == 'complete' else 'Stored enumeration values' }}

{% for item in column.value_profile.values %}
{{ item.value }}{{ item.count | count }}
{% else %}

No visible values.

{% endfor %}{% else %}

The full value list appears below.

{% endif %}

Detected formats

{% for detector in column.detectors %}{% if detector.formats %}

{{ detector.id }} {{ detector.matched_count | count }} matched

{% for fmt in detector.formats %}
{{ fmt.format }}{{ fmt.count | count }} · {{ fmt.percent }}%
{% endfor %}{% endif %}{% else %}

No detector formats.

{% endfor %}{% if column.date_profile and column.date_profile.formats %}

Date formats

{% for fmt in column.date_profile.formats %}
{{ fmt.format }}{{ fmt.count | count }} · {{ fmt.percent }}%
{% endfor %}{% endif %}

Type counts

{% for name, count in column.type_counts.items() %}
{{ name | replace('_',' ') }}{{ count | count }}
{% else %}

No classified values.

{% endfor %}{% if column.scan_details %}

Native types

{% for name, count in column.scan_details.native_types.items() %}
{{ name }}{{ count | count }}
{% endfor %}{% endif %}
{% if issues or diagnostics %}

Observations

{% for issue in issues %}
{{ issue.message }}{{ issue.severity }}{{ issue.count | count }}
{% endfor %}{% for diagnostic in diagnostics %}
{{ diagnostic.message }}{{ diagnostic.code }} · {{ diagnostic.level }}{{ diagnostic.count | count }}
{% endfor %}
{% endif %}

02 / Distribution

Values({{ column.value_profile.values | length }})

{% if column.value_profile.selection == 'complete' %}Every stored distinct value and its cardinality.{% else %}A {{ column.value_profile.selection | replace('_',' ') }} from {{ column.distinct_count | count if column.distinct_count is not none else 'a limited number of' }} distinct values.{% endif %} Counts are occurrences in the analyzed column.

{% if column.value_profile.values %}
{% for item in column.value_profile.values %}
{{ item.value }}{% if item.truncated %}…{% endif %}{{ item.count | count }}
{% endfor %}
{% else %}

No stored values{% if column.exposure == 'hide' %}: values are hidden by the exposure setting{% endif %}.

{% endif %}
{% if column.detectors %}

03 / Detection

Detectors and formats

All detectors that recognized values in this column.

{% for detector in column.detectors %}

{{ detector.id | replace('_',' ') | title }}{% if detector.primary %} primary{% endif %}

Status
{{ detector.status }}
Eligible
{{ detector.eligible_count | count }}
Matched
{{ detector.matched_count | count }} ({{ detector.matched_percent }}%)
Ambiguous
{{ detector.ambiguous_count | count }}
Invalid
{{ detector.invalid_count | count }}
{% if detector.coverage %}
Tested
{{ detector.coverage.tested | count }}
Not tested
{{ detector.coverage.not_tested | count }}
Not matched
{{ detector.coverage.not_matched | count }}
{% endif %}
{% for fmt in detector.formats %}
{{ fmt.format }}{{ fmt.count | count }} · {{ fmt.percent }}%
{% endfor %}{% if detector.evidence %}

Examples by outcome

{% for state in ['matched','ambiguous','invalid','not_matched'] %}{% set values = detector.evidence | attr(state) %}{% if values %}
{{ state | replace('_',' ') }}{{ values | join(' · ') }}
{% endif %}{% endfor %}{% endif %}{% if detector.adaptive %}

Adaptive detection: {{ detector.adaptive.not_tested | count }} values not tested after {{ detector.adaptive.skipped_after | count }} distinct values.

{% endif %}{% if detector.details %}

Detector details

{{ detector.details | tojson(indent=2) }}
{% endif %}
{% endfor %}
{% endif %} {% if column.date_profile %}{% set date = column.date_profile %}

04 / Dates

Date analysis

Ambiguous dates remain unresolved unless an order was configured.

Valid
{{ date.valid_count | count }}{{ date.valid_percent }}%
Ambiguous
{{ date.ambiguous_count | count }}{{ date.ambiguous_percent }}%
Invalid dates
{{ date.invalid_date_count | count }}{{ date.invalid_date_percent }}%
Other values
{{ date.not_date_count | count }}{{ date.not_date_percent }}%

Breakdown · {{ date.status | replace('_',' ') }}

{% for item in date.breakdown %}
{{ item.label }}{{ item.count | count }} · {{ item.percent }}%
{% endfor %}

Formats · {{ date.format_count }}

{% for item in date.formats %}
{{ item.format }}{% if item.order %} · {{ item.order }}{% endif %}{% if item.separator %} · {{ item.separator }}{% endif %}{{ item.count | count }} · {{ item.percent }}%
{% endfor %}

Ambiguity evidence

{% for order, count in date.ambiguity_evidence.items() %}
{{ order }}{{ count | count }}
{% else %}

No unambiguous order evidence.

{% endfor %}

Resolved order: {{ date.resolved_ambiguous_order or 'none' }}{% if date.ambiguous_order_source %} ({{ date.ambiguous_order_source }}){% endif %}

{% endif %} {% if column.numeric %}

05 / Numbers

Numeric analysis

Minimum
{{ column.numeric.minimum | number }}
Maximum
{{ column.numeric.maximum | number }}
Mean
{{ column.numeric.mean | number }}
Median
{{ column.numeric.median | number if column.numeric.median is not none else 'limited' }}Range {{ column.numeric.range | number }}
{% endif %} {% if column.string_profile %}{% set string = column.string_profile %}

06 / Text

String analysis

Lengths and representative values.

Shortest
{{ string.minimum_length }}
Longest
{{ string.maximum_length }}
Mean length
{{ string.mean_length | number }}
Median length
{{ string.median_length | number }}{{ string.status | replace('_',' ') }}{% if string.fixed_length is not none %} · fixed length {{ string.fixed_length }}{% endif %}

Length distribution · {{ string.distinct_length_count }}

{% for item in string.length_distribution %}
{{ item.length }} characters{% if item.examples %}{% for example in item.examples %}{{ example.value }}{% if not loop.last %} / {% endif %}{% endfor %}{% endif %}{{ item.count | count }} · {{ item.percent }}%
{% endfor %}

Representative values

{% for value in string.representative_examples %}
{{ value }}
{% else %}

No visible examples.

{% endfor %}
{% endif %}

07 / Cleaning

Transformations

Normalization preserves the raw values in the data sample.

Stages

{% for stage in column.normalization.stages %}
{{ stage.stage | replace('_',' ') }}{% if not stage.enabled %} · disabled{% endif %}{{ stage.changed_count | count if stage.changed_count is not none else '–' }} changed · {{ stage.changed_percent ~ '%' if stage.changed_percent is not none else '–' }}{{ stage.distinct_count | count if stage.distinct_count is not none else stage.distinct_status }}
{% endfor %}

Variant groups · {{ column.normalization.variant_group_count | count if column.normalization.variant_group_count is not none else column.normalization.variant_group_status }}

{% for group in column.normalization.variant_groups %}
{{ group.key }}{% for variant in group.variants %}{{ variant.value }} ({{ variant.count | count }}){% if not loop.last %} · {% endif %}{% endfor %}{% if group.truncated %} · truncated{% endif %}{{ group.count | count }}
{% else %}

No stored variant groups.

{% endfor %}{% if column.normalization.variant_groups_truncated %}

Additional groups were not stored.

{% endif %}
{% if column.scan_details %}{% set scan = column.scan_details %}

08 / Scan

Scan evidence

Field facts retained from the Scan document.

Presence and missing

First record
{{ scan.first_record if scan.first_record is not none else '–' }}
Occurrences
{{ scan.occurrences | count }}
Parent
{{ scan.presence.parent_type }}
Parents
{{ scan.presence.parent_count | count }}
Present
{{ scan.presence.present | count }}
Absent
{{ scan.presence.absent | count if scan.presence.absent is not none else '–' }}
{% for name, value in scan.missing.components.model_dump().items() %}
{{ name }}{{ value | count if value is not none else '–' }}
{% endfor %}

Missing definition: {{ scan.missing.definition | join(', ') }}

First and last values

{% if scan.first %}
First · record {{ scan.first.record }}{{ scan.first.value }} · {{ scan.first.type }}
{% endif %}{% if scan.last %}
Last · record {{ scan.last.record }}{{ scan.last.value }} · {{ scan.last.type }}
{% endif %}{% if not scan.first and not scan.last %}

No visible values.

{% endif %}

String categories

{% for name, value in scan.strings.model_dump().items() %}
{{ name | replace('_',' ') }}{{ value | count }}
{% endfor %}

String characteristics

{% for name, value in scan.string_characteristics.model_dump().items() %}
{{ name | replace('_',' ') }}{{ value | count }}
{% endfor %}{% if scan.string_lengths %}

Scan length histogram

{% for item in scan.string_lengths.length_histogram %}
{{ item.length }} characters{{ item.count | count }}
{% endfor %}{% endif %}
{% if scan.numeric %}

Scan numeric statistics

Numeric values
{{ scan.numeric.count | count }}
Native / text
{{ scan.numeric.native_count | count }} / {{ scan.numeric.text_count | count }}
Sum
{{ scan.numeric.sum | number }}
Variance
{{ scan.numeric.population_variance | number }}
Standard deviation
{{ scan.numeric.population_std | number }}
Positive / negative / zero
{{ scan.numeric.positive | count }} / {{ scan.numeric.negative | count }} / {{ scan.numeric.zero | count }}
Integral decimals
{{ scan.numeric.integral_decimals | count }}
{% endif %} {% if scan.booleans %}

Boolean values

True{{ scan.booleans.true | count }}
False{{ scan.booleans.false | count }}
{% endif %} {% if scan.temporal %}

Temporal range

{{ scan.temporal.count | count }} recognized; {{ scan.temporal.ambiguous | count }} ambiguous.

{% for kind in scan.temporal.kinds %}
{{ kind.kind | replace('_',' ') }}{{ kind.min }} → {{ kind.max }}{% if kind.years %}{% for year in kind.years %}{{ year.year }}: {{ year.count | count }}{% if not loop.last %} · {% endif %}{% endfor %}{% endif %}{{ kind.count | count }}
{% endfor %}
{% endif %}
{% endif %} {% if ds.preview %}

09 / Raw values

Data sample

First {{ ds.preview | length }} records, with original row numbers.

{% for row in ds.preview %}{% set index = column.position - 1 %}
Row {{ row.row_number }}{% if index in row.absent %}absent{% elif row.values[index] is none %}hidden{% elif row.values[index] == '' %}empty{% else %}{{ row.values[index] }}{% endif %}
{% endfor %}
{% endif %} {% if limits %}

10 / Limits

Limited measures

{% for item in limits %}
{{ item.measure }}{{ item.reason }} · limit {{ item.limit | count }}{% if item.lower_bound is not none %}at least {{ item.lower_bound | count }}{% else %}limited{% endif %}
{% endfor %}
{% endif %}