01 / Overview
Dataset overview
Size, completeness and the main signal to check first.
- Rows
- {{ report.summary.row_count | count }}{% if report.summary.empty_row_count %}{{ report.summary.empty_row_count | count }} empty rows{% else %}No empty rows{% endif %}
- Columns
- {{ report.summary.column_count | count }}{{ report.summary.cell_count | count }} cells analyzed
- Missing cells
- {{ report.summary.missing_percent }}%{{ report.summary.missing_count | count }} cells
- Duplicate rows
- {{ report.summary.duplicate_row_count | count }}Beyond the first occurrence
ambiguous dates, kept unresolved.
Values that match more than one configured date order. That is {{ report.date_summary.ambiguous_percent }}% of the {{ report.date_summary.present_count | count }} present date values. Tabalyst does not guess the order.
{% if report.date_summary.columns %}{% endif %}See date analysis {% else %}Date analysisdate columns detected.
No date-shaped values were identified with the configured rules.
{% endif %}Quality observations
- {% for issue in report.issues %}
- {{ issue.severity }}: {% if issue.code in ['missing_values','empty_columns','constant_columns','ambiguous_headers'] %}{{ issue.message }}{% elif issue.code in ['trimmed_values','collapsed_whitespace'] %}{{ issue.message }}{% elif issue.code == 'mixed_types' %}{{ issue.message }}{% else %}{{ issue.message }}{% endif %}{{ issue.count | count }} {% endfor %}
Column types
- {% for name, count in report.summary.inferred_type_counts.items() %}
- {{ name | replace('_',' ') | title }} {{ count }} {% endfor %}
Semantic types
- {% for name, count in report.summary.semantic_type_counts.items() %}
- {{ 'No semantic type' if name == 'none' else name | title }} {{ count }} {% endfor %}
Mixed: fewer than {{ (report.config.type_inference.minimum_confidence * 100) | round(1) }}% of present values agree on one type.
02 / Structure
Columns({{ report.summary.column_count }})
Inferred and semantic types per column. With issues: missing values or mixed type.
| {{ column.distinct_count | count }} | {{ column.inferred_type }} | {% if column.semantic_type %}{{ column.semantic_type }}{% else %}–{% endif %} | {{ percent_cell(column.type_error_percent,column.type_error_count or 0,'var(--attn)',column.type_error_count != 0) }}{{ value_examples(column) }}
03 / Cleaning
Transformations({{ report.summary.column_count }})
Normalization changes recorded during analysis. Raw preview values remain unchanged.
| {{ column.distinct_count | count }} |
04 / Numeric
Numeric analysis({{ report.summary.numeric_column_count }})
Range and distribution statistics for accepted numeric values.
| {{ column.numeric.range | number }} | {{ column.numeric.minimum | number }} | {{ column.numeric.maximum | number }} | {{ column.numeric.mean | number }} | {{ column.numeric.median | number }} | {{ column.distinct_count | count }} | {{ value_examples(column,'num') }}
05 / Dates
Date analysis({{ report.summary.date_column_count }})
Strict date parsing keeps ambiguous, invalid and non-date values separate.
| {{ column.date_profile.status | replace('_',' ') }} | {{ column.distinct_count | count }} | {{ column.date_profile.format_count }} variant{{ '' if column.date_profile.format_count == 1 else 's' }} {{ column.date_profile.format_count }} variant{{ '' if column.date_profile.format_count == 1 else 's' }} {% for item in column.date_profile.breakdown %} {{ item.label }}{{ item.count | count }} ({{ item.percent }}%) {% endfor %} |
06 / Text
String analysis({{ report.summary.string_column_count }})
Length classes, fixed widths and representative values.
| {{ column.string_profile.status | replace('_',' ') }} | {% if column.string_profile.fixed_length is not none %}{{ column.string_profile.fixed_length }}{% else %}–{% endif %} | {% if column.string_profile.fixed_length is not none %}– | – | – | – | {% else %}{{ column.string_profile.minimum_length }} | {{ column.string_profile.maximum_length }} | {{ column.string_profile.mean_length }} | {{ column.string_profile.median_length }} | {% endif %}{% for item in column.string_profile.length_distribution %}{% endfor %} | {{ column.distinct_count | count }} | {% for value in column.string_profile.representative_examples %}{{ value }}{% endfor %} Lengths {{ column.string_profile.distinct_length_count }} {% for item in column.string_profile.length_distribution | sort(attribute='count', reverse=true) %} {{ item.length }} character{{ '' if item.length == 1 else 's' }}{% for example in item.examples[:3] %}{{ example.value }}{% if not loop.last %} / {% endif %}{% endfor %} {{ item.percent }}%{{ item.count | count }} |
07 / Raw values
Data sample({{ report.preview | length }})
First {{ report.preview | length }} records with original row numbers and raw values.
Row | {% for column in report.columns %}{{ column.name or '(unnamed)' }}{{ column.inferred_type }} | {% endfor %}
|---|---|
| {{ row.row_number }} | {% for value in row.values %}{% if value == '' %}empty{% else %}{{ value }}{% endif %} | {% endfor %}
No data records.
{% endif %}08 / Method
Analysis settings
Rules used for this analysis, so the result can be reproduced.
- Scope
- All {{ report.summary.row_count | count }} records
- Missing-value markers
- {% for marker in report.config.missing_values %}
{{ marker | tojson }}{% else %}None{% endfor %} - Missing-value comparison
- Whitespace trimmed; case-sensitive
- Value normalization
- Trim: {{ 'enabled' if report.config.normalization.trim else 'disabled' }}; collapse internal horizontal whitespace: {{ 'enabled' if report.config.normalization.collapse_internal_whitespace else 'disabled' }}. Raw preview values are preserved.
- Duplicate comparison
- Exact raw values across every column
- Type inference
- At least {{ (report.config.type_inference.minimum_confidence * 100) | round(1) }}% agreement across present values; configured strict date formats, dot-decimal numbers and true/false. Leading-zero identifiers remain text.
- Date detection
- Orders: {{ report.config.date_detection.orders | join(', ') }}; separators: {% for separator in report.config.date_detection.separators %}
{{ separator | tojson }}{% endfor %}; ambiguous order: {{ report.config.date_detection.ambiguous_order or 'inferred only from unambiguous column evidence' }} - Enum candidates
- Text columns with at most {{ report.config.enum_detection.maximum_distinct_values }} values when the dataset has at least {{ report.config.enum_detection.minimum_row_count | count }} rows
- String lengths
- Very short through {{ report.config.string_analysis.very_short_max_length }}; short through {{ report.config.string_analysis.short_max_length }}; medium through {{ report.config.string_analysis.medium_max_length }}; long through {{ report.config.string_analysis.long_max_length }}; otherwise very long. Up to {{ report.config.string_analysis.examples_per_length }} examples per length are retained when the maximum is {{ report.config.string_analysis.length_distribution_max_length }} or less.
- Value representation
- Complete through {{ report.config.value_examples.full_distribution_max_distinct }} distinct values; otherwise a reproducible sample of {{ report.config.value_examples.candidate_sample_size }}
- Row numbering
- Data records start at 1, excluding the header. Quoted multiline values count as one record.
- Preview selection
- First {{ report.config.preview_rows }} records; raw values preserved
- JSON format
- {{ report.format_version }} - revision {{ report.format_revision }} (experimental)
{{ report.source.sha256 }}