Profile format changelog
The generated report.json identifies its contract with two independent fields:
{ "format_version": "0.1.0a", "format_revision": 11}format_version names the experimental compatibility family. It remains
0.1.0a throughout the Tabalyst 0.1.0aX application series.
format_revision is a monotonic integer incremented for meaningful structural or
semantic changes. It does not change for report styling, documentation, performance
improvements, or corrections that preserve the JSON contract.
Beta revisions may be incompatible. No automatic migration is provided yet.
Revision 1 - 2026-09-18
Section titled “Revision 1 - 2026-09-18”- Established the first explicitly tracked alpha profile contract.
- Includes dataset/source metadata, effective configuration, quality summary, issues, raw preview rows, and global column profiles.
- Column profiles include normalization, inferred and semantic types, errors, occurrences, bounded examples, numeric statistics, date analysis, and string length statistics.
- Added
format_versionand monotonicformat_revisionidentifiers.
Revision 2 - 2026-09-23
Section titled “Revision 2 - 2026-09-23”- Materialized per-column
with_issues, defined as missing values or inferred typemixed. - Added dataset-level physical/semantic type distributions and report section counts.
- Added dataset-level date aggregates and per-date-column ambiguity summaries.
- Added present counts and percentages to date profiles, numeric ranges, and relative/representative string-length values used by the report.
Revision 3 - 2026-09-27
Section titled “Revision 3 - 2026-09-27”The report is built on Tabalyst Scan instead of the pandas engine.
configholds the report settings (preview_rows,string_analysis,value_exampleswithoutcandidate_sample_sizeandrandom_seed) andscan, the complete effective scan configuration. The other former settings moved toscan.- Columns of dates with several formats or ambiguous values have
inferred_typedateinstead ofmixed.with_issuesis alsotruefor columns with ambiguous dates, and the newambiguous_datesissue counts them. - Ambiguous dates are no longer resolved from the other values of the column:
ambiguous_order_sourceis onlyconfig, and the newdate_profile.ambiguity_evidencecounts unambiguous values per order. Date formats include ISO date-times, times and month names;orderandseparatorarenullfor formats that are not numeric dates.date_profile.errorswas removed. semantic_typeisdateor the id of the scan’s primary interpretation (enumerationinstead ofenum, and new values such asemail,phoneorpostal_code). Theenumblock was removed. Enumerations count present values, not rows, against their minimum.- New
exposureper column; values of sensitive columns are masked by default inexamples,value_profileandpreview, where a hidden value isnull. distinct_countandnumeric.medianarenullwhen a scan limit stopped them; the newlimited_measuresissue lists such columns.type_confidenceisnullforemptycolumns.- Numeric statistics cover every number of the column, including decimal commas and grouped thousands.
- Normalization counts cover present values only: a whitespace-only cell is missing, not trimmed. Distinct values also apply Unicode composition (NFC).
string_profile.length_distributionitems no longer havedistinct_count; their examples come from the most frequent and the sampled values.- Value samples are drawn by the scan (
scan.limits.max_samples,scan.random_seed), so sampled examples differ from revision 2.
Revision 4 - 2026-09-27
Section titled “Revision 4 - 2026-09-27”The report reads JSON files and lists the detectors of each column.
- The dataset-level fields
summary,date_summary,columns,issuesandpreviewmoved into the newdatasetsarray, one item per dataset with itsidandkind. A CSV file has one dataset,rows. sourcehas a newformatfield;delimiterisnullfor JSON files.- Columns have a new
pathand a newdetectorsarray: the detectors that recognized values, with their counts and formats. - Preview rows have a new
absentarray: the positions of the JSON fields absent from the record. - In JSON datasets, a column is a field holding scalar values, named by its
path; absent fields count as missing, and
cell_countcounts the places a value could be.
Revision 5 - 2026-09-28
Section titled “Revision 5 - 2026-09-28”The report can be built from a scan document, whose records block now holds the preview and the duplicate rows.
- New
summary.duplicate_row_status:complete,limitedwhen more distinct rows thanscan.limits.max_tracked_recordswere read, in which caseduplicate_row_countis a lower bound, ordisabledwhenscan.records.duplicatesisfalse, in which caseduplicate_row_countisnull. - The
duplicate_rowsissue message ends withat leastwhen the count is a lower bound. - The preview size moved from
config.preview_rowstoconfig.scan.records.preview. - For a report built from a scan document,
processing_secondsis the time to read the document and build the profile.
Revision 6 - 2026-09-28
Section titled “Revision 6 - 2026-09-28”The report shows every normalization stage with its variant groups, what the scan could not measure completely, and the structure of JSON datasets.
- Column
normalizationgainsstages(raw, thennfc,trim,collapse_whitespace,casefoldandstrip_accents, each withenabled,changed_count,changed_percent,distinct_countanddistinct_status),variant_group_count,variant_group_status,variant_groupsandvariant_groups_truncated. - New info issue
variant_groups: values written in several ways that normalization compares as equal. - Each dataset has a new
limitsobject:measuresstopped by a scan limit, with their field, reason, limit and proven lower bound;untracked_observations,depth_truncated_observations; and the scandiagnosticsof the dataset and of the whole scan. - JSON datasets have a new
structureobject:record_types,path_count,path_status,max_depth_seenand one item per path, containers included, with presence per parent and array lengths. It isnullfor CSV files.
Revision 7 - 2026-09-28
Section titled “Revision 7 - 2026-09-28”The scan behind the report uses adaptive detection by default: on columns with more than 10,000 distinct values, detectors that recognized none of the first ones stop testing the others.
config.scan.detectiongainswarmup_values(default 10,000;0tests every value) andprobe_interval(default 100).- Semantic types and inferred types are unchanged on columns where a
detector matches at least 95% of the values. Counts of rare matches after
the warm-up may be lower; a
detector_skipped_reactedwarning in the datasetlimits.diagnosticssays when a probe found one. - For a report built from a scan document,
processing_secondsnow adds the scan duration recorded in the document to the time to read it and build the profile, so it is comparable with a report built from the source.
Revision 8 - 2026-09-28
Section titled “Revision 8 - 2026-09-28”The scan behind the report also stops testing rare detectors: those that recognized at most 0.1% of the first 10,000 distinct values of a column, and none of the last 5,000.
config.scan.detectiongainsrare_share(default 0.001;0keeps the behavior of revision 7).- Semantic types and inferred types are unchanged; only counts of detectors
with rare matches may be lower. A
detector_skipped_reactedwarning in the datasetlimits.diagnosticssays when such a detector recognized probed values more often than its warm-up.
Revision 9 - 2026-09-29
Section titled “Revision 9 - 2026-09-29”- Column
scan_detailsretains bounded Scan evidence for standalone column pages: presence, native types, missing components, first/last exposed values, string characteristics and lengths, numeric statistics, boolean counts and temporal ranges. Values obey the Scan exposure setting. - Detector entries gain complete coverage, exposed evidence, exposed details and adaptive detection metadata when available.
- Primary enumerations include every available frequency even when their
cardinality exceeds
value_examples.full_distribution_max_distinct. tabalyst report --detailsgenerates one self-contained HTML page per column under the report stem folder. Details are off by default. This changes report artifacts, not the meaning of other profile fields.
Revision 10 - 2026-09-30
Section titled “Revision 10 - 2026-09-30”config.scanrecords the rules the scan applied, defaults resolved:config.scan.errors.policyisstrictortolerant, nevernull. The setting itself acceptsnull, meaning the default of the source format.config.scan.jsongainsflatten(enabled,separator,max_depth) andarrays(mode). Their defaults keep the previous behavior.- A JSON field holding only objects or arrays kept whole at the
config.scan.json.flatten.max_depthlimit is a column withinferred_typecomplexand theobjectandarraycounts intype_counts. Without a flatten limit, containers stay structure and no column changes. source.formatcan bejsonl, for files ending in.jsonlor.ndjson. Lines excluded under thetolerantpolicy are counted by the existingexcluded_recordsissue, with their line numbers asrow_numbers.config.scan.limitsgainsmax_line_bytes.- The reports built by the commands have one dataset per collection chosen with
Inspect or
--collection, usually one, and no dataset$for the rest of the document. The structure ofdatasetsdoes not change. - Column
pathandnameof JSON fields are joined withconfig.scan.json.flatten.separator.
Revision 11
Section titled “Revision 11”source.formatcan beexcel, for.xlsxand.xlsmworkbooks. The report analyzes one table of the workbook, a datasetrowswhose columns are the cells of its header, as for a CSV file.source.encodingisnullfor a workbook, which has no text encoding. It stays a string for CSV, JSON and JSONL sources.source.delimiterisnull, as for JSON.config.scangainsexcel(dataset_path,header_row), which selects the table of a workbook.