{% macro logo() -%} {%- endmacro %} {% macro icons() -%} {%- endmacro %} {% macro table_head(label, kind='text', align=false, pill=false, filter=true) -%}
{{ label }}{% if filter %}{% endif %}
{%- endmacro %} {% macro column_name(column) -%}{{ column.position }}{{ column.name or '(unnamed)' }}{%- endmacro %} {% macro percent_cell(value, count, color='var(--attn)', attention=false) -%} {% if value %}{{ value }}%{{ count | count }}{% else %}–{% endif %} {%- endmacro %} {% macro value_examples(column, prefix='col') -%} {% for value in column.examples %}{{ value }}{% endfor %}{% if column.distinct_count > column.examples|length %}{% endif %} {%- endmacro %} {{ icons() }}

Dataset report

{{ source_stem }}{{ source_suffix }}

  • CSV
  • {{ report.summary.row_count | count }} rows
  • {{ report.summary.column_count | count }} columns
  • {{ report.source.size_bytes | size }}
  • {{ report.source.encoding }}
  • Delimiter {{ report.source.delimiter }}
Analysis complete · {{ report.processing_seconds | seconds }} s

01 / Overview

Dataset overview

Size, completeness and the main signal to check first.

Rows
{{ report.summary.row_count | count }}{% if report.summary.empty_row_count %}{{ report.summary.empty_row_count | count }} empty rows{% else %}No empty rows{% endif %}
Columns
{{ report.summary.column_count | count }}{{ report.summary.cell_count | count }} cells analyzed
Missing cells
{{ report.summary.missing_percent }}%{{ report.summary.missing_count | count }} cells
Duplicate rows
{{ report.summary.duplicate_row_count | count }}Beyond the first occurrence
{% if report.date_summary %}Derived · sum of {{ report.summary.date_column_count }} date columns
{{ report.date_summary.ambiguous_count | count }}

ambiguous dates, kept unresolved.

Values that match more than one configured date order. That is {{ report.date_summary.ambiguous_percent }}% of the {{ report.date_summary.present_count | count }} present date values. Tabalyst does not guess the order.

{% if report.date_summary.columns %}{% endif %}See date analysis {% else %}Date analysis
0

date columns detected.

No date-shaped values were identified with the configured rules.

{% endif %}

Quality observations

    {% for issue in report.issues %}
  • {{ issue.severity }}: {% if issue.code in ['missing_values','empty_columns','constant_columns','ambiguous_headers'] %}{{ issue.message }}{% elif issue.code in ['trimmed_values','collapsed_whitespace'] %}{{ issue.message }}{% elif issue.code == 'mixed_types' %}{{ issue.message }}{% else %}{{ issue.message }}{% endif %}{{ issue.count | count }}
  • {% endfor %}

Column types

    {% for name, count in report.summary.inferred_type_counts.items() %}
  • {{ name | replace('_',' ') | title }} {{ count }}
  • {% endfor %}

Semantic types

    {% for name, count in report.summary.semantic_type_counts.items() %}
  • {{ 'No semantic type' if name == 'none' else name | title }} {{ count }}
  • {% endfor %}

Mixed: fewer than {{ (report.config.type_inference.minimum_confidence * 100) | round(1) }}% of present values agree on one type.

02 / Structure

Columns({{ report.summary.column_count }})

Inferred and semantic types per column. With issues: missing values or mixed type.

{{ report.summary.column_count }} of {{ report.summary.column_count }} columns
{{ table_head('Column','list') }}{{ table_head('Missing','num',true) }}{{ table_head('Distinct','num',true) }}{{ table_head('Inferred type','list',false,true) }}{{ table_head('Semantic type','list',false,true) }}{{ table_head('Error','num',true) }}{{ table_head('Examples',filter=false) }}{% for column in report.columns %}{{ column_name(column) }}{{ percent_cell(column.missing_percent,column.missing_count) }}{{ percent_cell(column.type_error_percent,column.type_error_count or 0,'var(--attn)',column.type_error_count != 0) }}{{ value_examples(column) }}{% endfor %}
{{ column.distinct_count | count }}{{ column.inferred_type }}{% if column.semantic_type %}{{ column.semantic_type }}{% else %}–{% endif %}

03 / Cleaning

Transformations({{ report.summary.column_count }})

Normalization changes recorded during analysis. Raw preview values remain unchanged.

{{ report.summary.column_count }} of {{ report.summary.column_count }} columns
{{ table_head('Column','list') }}{{ table_head('Missing','num',true) }}{{ table_head('Trimmed','num',true) }}{{ table_head('Whitespace collapsed','num',true) }}{{ table_head('Distinct','num',true) }}{% for column in report.columns %}{{ column_name(column) }}{{ percent_cell(column.missing_percent,column.missing_count) }}{{ percent_cell(column.normalization.trim_percent,column.normalization.trim_count,'var(--d2)') }}{{ percent_cell(column.normalization.collapse_internal_whitespace_percent,column.normalization.collapse_internal_whitespace_count,'var(--d2)') }}{% endfor %}
{{ column.distinct_count | count }}
{% if numeric_columns %}

04 / Numeric

Numeric analysis({{ report.summary.numeric_column_count }})

Range and distribution statistics for accepted numeric values.

{{ report.summary.numeric_column_count }} of {{ report.summary.numeric_column_count }} columns
{{ table_head('Column','list') }}{{ table_head('Range','num',true) }}{{ table_head('Minimum','num',true) }}{{ table_head('Maximum','num',true) }}{{ table_head('Mean','num',true) }}{{ table_head('Median','num',true) }}{{ table_head('Distinct','num',true) }}{{ table_head('Examples',filter=false) }}{% for column in numeric_columns %}{{ column_name(column) }}{{ value_examples(column,'num') }}{% endfor %}
{{ column.numeric.range | number }}{{ column.numeric.minimum | number }}{{ column.numeric.maximum | number }}{{ column.numeric.mean | number }}{{ column.numeric.median | number }}{{ column.distinct_count | count }}
{% endif %} {% if date_columns %}

05 / Dates

Date analysis({{ report.summary.date_column_count }})

Strict date parsing keeps ambiguous, invalid and non-date values separate.

{{ report.summary.date_column_count }} of {{ report.summary.date_column_count }} date columns
{{ table_head('Column','list') }}{{ table_head('Status','list',false,true) }}{{ table_head('Breakdown',filter=false) }}{{ table_head('Valid','num',true) }}{{ table_head('Ambiguous','num',true) }}{{ table_head('Invalid','num',true) }}{{ table_head('Other','num',true) }}{{ table_head('Distinct','num',true) }}{{ table_head('Formats',filter=false) }}{% for column in date_columns %}{{ column_name(column) }}{{ percent_cell(column.date_profile.valid_percent,column.date_profile.valid_count,'var(--ok)') }}{{ percent_cell(column.date_profile.ambiguous_percent,column.date_profile.ambiguous_count,'var(--attn)',true) }}{{ percent_cell(column.date_profile.invalid_date_percent,column.date_profile.invalid_date_count,'var(--attn)',true) }}{{ percent_cell(column.date_profile.not_date_percent,column.date_profile.not_date_count,'var(--d3)') }}{% endfor %}
{{ column.date_profile.status | replace('_',' ') }}{{ column.distinct_count | count }}{{ column.date_profile.format_count }} variant{{ '' if column.date_profile.format_count == 1 else 's' }}
{% endif %} {% if string_columns %}

06 / Text

String analysis({{ report.summary.string_column_count }})

Length classes, fixed widths and representative values.

{{ report.summary.string_column_count }} of {{ report.summary.string_column_count }} string columns
{{ table_head('Column','list') }}{{ table_head('Class','list',false,true) }}{{ table_head('Fixed','num',true) }}{{ table_head('Min','num',true) }}{{ table_head('Max','num',true) }}{{ table_head('Mean','num',true) }}{{ table_head('Median','num',true) }}{{ table_head('Lengths',filter=false) }}{{ table_head('Distinct','num',true) }}{{ table_head('Examples',filter=false) }}{% for column in string_columns %}{{ column_name(column) }}{% if column.string_profile.fixed_length is not none %}{% else %}{% endif %}{% endfor %}
{{ column.string_profile.status | replace('_',' ') }}{% if column.string_profile.fixed_length is not none %}{{ column.string_profile.fixed_length }}{% else %}–{% endif %}––––{{ column.string_profile.minimum_length }}{{ column.string_profile.maximum_length }}{{ column.string_profile.mean_length }}{{ column.string_profile.median_length }}{% for item in column.string_profile.length_distribution %}{% endfor %}{{ column.distinct_count | count }}{% for value in column.string_profile.representative_examples %}{{ value }}{% endfor %}
{% endif %}

07 / Raw values

Data sample({{ report.preview | length }})

First {{ report.preview | length }} records with original row numbers and raw values.

{% if report.preview %}
{{ report.preview | length }} of {{ report.preview | length }} sample rows
{% for column in report.columns %}{% endfor %}{% for row in report.preview %}{% for value in row.values %}{% if value == '' %}empty{% else %}{{ value }}{% endif %}{% endfor %}{% endfor %}
Row
{{ column.name or '(unnamed)' }}{{ column.inferred_type }}
{{ row.row_number }}
{% else %}

No data records.

{% endif %}

08 / Method

Analysis settings

Rules used for this analysis, so the result can be reproduced.

Scope
All {{ report.summary.row_count | count }} records
Missing-value markers
{% for marker in report.config.missing_values %}{{ marker | tojson }} {% else %}None{% endfor %}
Missing-value comparison
Whitespace trimmed; case-sensitive
Value normalization
Trim: {{ 'enabled' if report.config.normalization.trim else 'disabled' }}; collapse internal horizontal whitespace: {{ 'enabled' if report.config.normalization.collapse_internal_whitespace else 'disabled' }}. Raw preview values are preserved.
Duplicate comparison
Exact raw values across every column
Type inference
At least {{ (report.config.type_inference.minimum_confidence * 100) | round(1) }}% agreement across present values; configured strict date formats, dot-decimal numbers and true/false. Leading-zero identifiers remain text.
Date detection
Orders: {{ report.config.date_detection.orders | join(', ') }}; separators: {% for separator in report.config.date_detection.separators %}{{ separator | tojson }} {% endfor %}; ambiguous order: {{ report.config.date_detection.ambiguous_order or 'inferred only from unambiguous column evidence' }}
Enum candidates
Text columns with at most {{ report.config.enum_detection.maximum_distinct_values }} values when the dataset has at least {{ report.config.enum_detection.minimum_row_count | count }} rows
String lengths
Very short through {{ report.config.string_analysis.very_short_max_length }}; short through {{ report.config.string_analysis.short_max_length }}; medium through {{ report.config.string_analysis.medium_max_length }}; long through {{ report.config.string_analysis.long_max_length }}; otherwise very long. Up to {{ report.config.string_analysis.examples_per_length }} examples per length are retained when the maximum is {{ report.config.string_analysis.length_distribution_max_length }} or less.
Value representation
Complete through {{ report.config.value_examples.full_distribution_max_distinct }} distinct values; otherwise a reproducible sample of {{ report.config.value_examples.candidate_sample_size }}
Row numbering
Data records start at 1, excluding the header. Quoted multiline values count as one record.
Preview selection
First {{ report.config.preview_rows }} records; raw values preserved
JSON format
{{ report.format_version }} - revision {{ report.format_revision }} (experimental)
SOURCE SHA-256{{ report.source.sha256 }}