01 / Overview
Column overview
Profile, values and formats at a glance.
- Analyzable values
- {{ column.scan_details.value_count | count if column.scan_details else (ds.summary.row_count - column.missing_count) | count }}in {{ ds.summary.row_count | count }} {{ 'records' if report.source.format in ('json', 'jsonl') else 'rows' }}
- Missing
- {{ column.missing_percent }}%{{ column.missing_count | count }} values
- Distinct
- {% if column.distinct_count is none %}limited{% else %}{{ column.distinct_count | count }}{% endif %}{{ column.value_profile.selection | replace('_',' ') }}
- Type errors
- {{ column.type_error_count | count if column.type_error_count is not none else '–' }}{{ column.type_error_percent ~ '%' if column.type_error_percent is not none else 'Not measured' }}
Interpretation
- Inferred type
- {{ column.inferred_type }}
- Semantic type
- {% if column.semantic_type %}{{ column.semantic_type | replace('_',' ') }}{% else %}None{% endif %}
- Confidence
- {{ ((column.type_confidence * 100) | round(1)) ~ '%' if column.type_confidence is not none else 'Not applicable' }}
- Exposure
- {{ column.exposure or 'Normal' }}
Value coverage
{{ column.value_profile.sampled_distinct_count | count }} stored distinct values · {{ column.value_profile.selection | replace('_',' ') }}{% if column.distinct_count is none %} · cardinality limited{% endif %}.
{% if column.semantic_type == 'enumeration' %}{{ 'All visible enumeration values' if column.value_profile.selection == 'complete' else 'Stored enumeration values' }}
{% for item in column.value_profile.values %}{{ item.value }}{{ item.count | count }}No visible values.
{% endfor %}{% else %}The full value list appears below.
{% endif %}Detected formats
{% for detector in column.detectors %}{% if detector.formats %}{{ detector.id }} {{ detector.matched_count | count }} matched
{% for fmt in detector.formats %}{{ fmt.format }}{{ fmt.count | count }} · {{ fmt.percent }}%No detector formats.
{% endfor %}{% if column.date_profile and column.date_profile.formats %}Date formats
{% for fmt in column.date_profile.formats %}{{ fmt.format }}{{ fmt.count | count }} · {{ fmt.percent }}%Type counts
{% for name, count in column.type_counts.items() %}No classified values.
{% endfor %}{% if column.scan_details %}Native types
{% for name, count in column.scan_details.native_types.items() %}Observations
{% for issue in issues %}02 / Distribution
Values({{ column.value_profile.values | length }})
{% if column.value_profile.selection == 'complete' %}Every stored distinct value and its cardinality.{% else %}A {{ column.value_profile.selection | replace('_',' ') }} from {{ column.distinct_count | count if column.distinct_count is not none else 'a limited number of' }} distinct values.{% endif %} Counts are occurrences in the analyzed column.
{{ item.value }}{% if item.truncated %}…{% endif %}{{ item.count | count }}No stored values{% if column.exposure == 'hide' %}: values are hidden by the exposure setting{% endif %}.
{% endif %}03 / Detection
Detectors and formats
All detectors that recognized values in this column.
{{ detector.id | replace('_',' ') | title }}{% if detector.primary %} primary{% endif %}
- Status
- {{ detector.status }}
- Eligible
- {{ detector.eligible_count | count }}
- Matched
- {{ detector.matched_count | count }} ({{ detector.matched_percent }}%)
- Ambiguous
- {{ detector.ambiguous_count | count }}
- Invalid
- {{ detector.invalid_count | count }} {% if detector.coverage %}
- Tested
- {{ detector.coverage.tested | count }}
- Not tested
- {{ detector.coverage.not_tested | count }}
- Not matched
- {{ detector.coverage.not_matched | count }} {% endif %}
{{ fmt.format }}{{ fmt.count | count }} · {{ fmt.percent }}%Examples by outcome
{% for state in ['matched','ambiguous','invalid','not_matched'] %}{% set values = detector.evidence | attr(state) %}{% if values %}Adaptive detection: {{ detector.adaptive.not_tested | count }} values not tested after {{ detector.adaptive.skipped_after | count }} distinct values.
{% endif %}{% if detector.details %}Detector details
{{ detector.details | tojson(indent=2) }}{% endif %}04 / Dates
Date analysis
Ambiguous dates remain unresolved unless an order was configured.
- Valid
- {{ date.valid_count | count }}{{ date.valid_percent }}%
- Ambiguous
- {{ date.ambiguous_count | count }}{{ date.ambiguous_percent }}%
- Invalid dates
- {{ date.invalid_date_count | count }}{{ date.invalid_date_percent }}%
- Other values
- {{ date.not_date_count | count }}{{ date.not_date_percent }}%
Breakdown · {{ date.status | replace('_',' ') }}
{% for item in date.breakdown %}Formats · {{ date.format_count }}
{% for item in date.formats %}{{ item.format }}{% if item.order %} · {{ item.order }}{% endif %}{% if item.separator %} · {{ item.separator }}{% endif %}{{ item.count | count }} · {{ item.percent }}%Ambiguity evidence
{% for order, count in date.ambiguity_evidence.items() %}No unambiguous order evidence.
{% endfor %}Resolved order: {{ date.resolved_ambiguous_order or 'none' }}{% if date.ambiguous_order_source %} ({{ date.ambiguous_order_source }}){% endif %}
05 / Numbers
Numeric analysis
- Minimum
- {{ column.numeric.minimum | number }}
- Maximum
- {{ column.numeric.maximum | number }}
- Mean
- {{ column.numeric.mean | number }}
- Median
- {{ column.numeric.median | number if column.numeric.median is not none else 'limited' }}Range {{ column.numeric.range | number }}
06 / Text
String analysis
Lengths and representative values.
- Shortest
- {{ string.minimum_length }}
- Longest
- {{ string.maximum_length }}
- Mean length
- {{ string.mean_length | number }}
- Median length
- {{ string.median_length | number }}{{ string.status | replace('_',' ') }}{% if string.fixed_length is not none %} · fixed length {{ string.fixed_length }}{% endif %}
Length distribution · {{ string.distinct_length_count }}
{% for item in string.length_distribution %}Representative values
{% for value in string.representative_examples %}{{ value }}No visible examples.
{% endfor %}07 / Cleaning
Transformations
Normalization preserves the raw values in the data sample.
Stages
{% for stage in column.normalization.stages %}Variant groups · {{ column.normalization.variant_group_count | count if column.normalization.variant_group_count is not none else column.normalization.variant_group_status }}
{% for group in column.normalization.variant_groups %}{{ group.key }}{% for variant in group.variants %}{{ variant.value }} ({{ variant.count | count }}){% if not loop.last %} · {% endif %}{% endfor %}{% if group.truncated %} · truncated{% endif %}{{ group.count | count }}No stored variant groups.
{% endfor %}{% if column.normalization.variant_groups_truncated %}Additional groups were not stored.
{% endif %}08 / Scan
Scan evidence
Field facts retained from the Scan document.
Presence and missing
- First record
- {{ scan.first_record if scan.first_record is not none else '–' }}
- Occurrences
- {{ scan.occurrences | count }}
- Parent
- {{ scan.presence.parent_type }}
- Parents
- {{ scan.presence.parent_count | count }}
- Present
- {{ scan.presence.present | count }}
- Absent
- {{ scan.presence.absent | count if scan.presence.absent is not none else '–' }}
Missing definition: {{ scan.missing.definition | join(', ') }}
First and last values
{% if scan.first %}{{ scan.first.value }} · {{ scan.first.type }}{{ scan.last.value }} · {{ scan.last.type }}No visible values.
{% endif %}String categories
{% for name, value in scan.strings.model_dump().items() %}String characteristics
{% for name, value in scan.string_characteristics.model_dump().items() %}Scan length histogram
{% for item in scan.string_lengths.length_histogram %}Scan numeric statistics
- Numeric values
- {{ scan.numeric.count | count }}
- Native / text
- {{ scan.numeric.native_count | count }} / {{ scan.numeric.text_count | count }}
- Sum
- {{ scan.numeric.sum | number }}
- Variance
- {{ scan.numeric.population_variance | number }}
- Standard deviation
- {{ scan.numeric.population_std | number }}
- Positive / negative / zero
- {{ scan.numeric.positive | count }} / {{ scan.numeric.negative | count }} / {{ scan.numeric.zero | count }}
- Integral decimals
- {{ scan.numeric.integral_decimals | count }}
Boolean values
Temporal range
{{ scan.temporal.count | count }} recognized; {{ scan.temporal.ambiguous | count }} ambiguous.
{% for kind in scan.temporal.kinds %}09 / Raw values
Data sample
First {{ ds.preview | length }} records, with original row numbers.
{{ row.values[index] }}{% endif %}10 / Limits
Limited measures
{{ item.measure }}{{ item.reason }} · limit {{ item.limit | count }}{% if item.lower_bound is not none %}at least {{ item.lower_bound | count }}{% else %}limited{% endif %}