Skip to content
Beta · version 0.6.1 · commands and the JSON format may still change.

JSON profile

Each tabalyst report run writes a JSON profile next to the HTML report, with the same name: orders.html comes with orders.json. For a JSON source, the default names are orders.report.html and orders.report.json, so the profile never replaces the source. The profile holds the complete analysis result; the HTML report is rendered from it. Use it to read Tabalyst results from scripts or other tools.

The format is experimental: it can change incompatibly between releases. Always check format_version and format_revision first. This page describes format 0.1.0a, revision 10.

A shortened profile for a five-row orders.csv:

{
"format_version": "0.1.0a",
"format_revision": 10,
"generated_at": "2026-09-25T22:25:07.873024Z",
"processing_seconds": 0.0098,
"source": {
"filename": "orders.csv",
"format": "csv",
"size_bytes": 274,
"sha256": "7d426449...",
"encoding": "utf-8-sig",
"delimiter": ","
},
"config": { "...": "effective report and scan settings" },
"datasets": [
{
"id": "rows",
"kind": "table",
"summary": {
"row_count": 5,
"column_count": 6,
"missing_count": 2,
"duplicate_row_count": 1,
"with_issues_column_count": 2
},
"date_summary": { "...": "present when a column contains dates" },
"columns": [
{
"id": "column_3",
"name": "amount",
"path": "amount",
"position": 3,
"inferred_type": "mixed",
"type_counts": { "number": 3, "text": 1 },
"type_confidence": 0.75,
"missing_count": 1,
"missing_percent": 20.0,
"with_issues": true,
"distinct_count": 3,
"examples": ["25.00", "12.50", "not available"],
"semantic_type": null,
"detectors": [
{
"id": "number",
"status": "complete",
"primary": false,
"eligible_count": 4,
"matched_count": 3,
"matched_percent": 75.0,
"ambiguous_count": 0,
"invalid_count": 0,
"formats": []
}
]
}
],
"issues": [
{
"code": "duplicate_rows",
"severity": "warning",
"message": "Duplicate rows beyond their first occurrence",
"count": 1,
"column_ids": [],
"row_numbers": [4]
}
],
"preview": [
{
"row_number": 1,
"values": [
"001", "Alice", "12.50", "2026-01-01", "true", "First order"
],
"absent": []
}
],
"limits": {
"measures": [],
"untracked_observations": 0,
"depth_truncated_observations": 0,
"diagnostics": []
},
"structure": null
}
]
}
FieldTypeContent
format_versionstringExperimental format family, "0.1.0a", retained during beta
format_revisionintegerRevision within the family, increased for each structural or semantic change
generated_atstringUTC date and time of the analysis (ISO 8601)
processing_secondsnumberScan and profile time, in seconds; for a report built from a scan document with --scan, the scan duration recorded in the document plus the time to read it and build the profile
sourceobjectThe analyzed file (see below)
configobjectThe effective settings, after merging defaults, the configuration files and command options: string_analysis, value_examples and scan, the complete scan configuration; for a report built from a scan document, scan is the configuration recorded in it. See the configuration
datasetsarrayOne profile per dataset of the file (see below)
FieldContent
filenameFile name, without folder
formatcsv, json, jsonl or excel
size_bytesFile size in bytes
sha256SHA-256 hash of the file, to check that two profiles describe the same file
encodingText encoding used to read the file; null for an Excel workbook
delimiterField delimiter used to read the file; null for JSON, JSONL and Excel files

A CSV file has one dataset, rows, whose records are the rows of the file; so has an Excel workbook, whose rows are those of the table analyzed. A JSON file has one dataset per collection analyzed, such as $.customers[], usually one: the collection selected by Inspect or named with --collection. A JSONL file has one dataset, $[]. Reports built by the commands have no $ dataset for the rest of the document. See the scan format for the datasets of a scan.

FieldTypeContent
idstringDataset identifier: rows, $ or a collection path such as $.customers[]
kindstringtable (CSV), document or collection (JSON)
summaryobjectDataset-level counts (see below)
date_summaryobject or nullDataset-level date counts, when at least one column contains date values
columnsarrayOne profile per column, in file order for CSV and in order of discovery for JSON
issuesarrayDetected problems (see below)
previewarrayThe first records, with raw values; values of sensitive columns are masked or hidden
limitsobjectMeasures stopped by a scan limit, structural truncation and scan diagnostics (see below)
structureobject or nullPaths of a JSON dataset, containers included (see below); null for CSV files

In a JSON dataset, a column is a field that holds strings, numbers, booleans or nulls, named by its path: email, address.city, orders[].total. Objects and arrays themselves are not columns, except the ones kept whole at the scan.json.flatten.max_depth limit, which are columns of type complex. A row is a record of the dataset.

Counts over the whole dataset:

  • size: row_count (records), column_count, cell_count (the places a value could be: one per row and column for CSV; for JSON, the occurrences of each field plus the objects where it is absent);
  • missing values: missing_count, missing_percent;
  • normalization: trim_count, collapse_internal_whitespace_count;
  • structure: duplicate_row_count, empty_row_count, empty_column_count, constant_column_count, with_issues_column_count;
  • duplicate_row_status: complete; limited when more distinct rows than scan.limits.max_tracked_records were read, duplicate_row_count then being a lower bound; or disabled when scan.records.duplicates is false, duplicate_row_count then being null;
  • column families: numeric_column_count, date_column_count, string_column_count;
  • type distributions: inferred_type_counts, inferred_type_percents, semantic_type_counts, semantic_type_percents.

Percentages are numbers from 0 to 100.

Each column profile always contains:

FieldContent
idStable identifier from the position, such as column_3. Use it rather than name, which can be blank or repeated
nameHeader text for CSV; the field path for JSON
pathField path as displayed by the scan: the header text for CSV, a path such as orders[].total for JSON
positionPosition in the file for CSV, in order of discovery for JSON, starting at 1
inferred_typeempty, complex, boolean, integer, number, date, text or mixed. A column of dates is date even with several formats or ambiguous values. complex is a JSON column whose present values are all objects or arrays kept whole at the flatten limit; type_counts then counts object and array
type_countsNumber of present values of each type
type_confidenceShare of present values accepted by the inferred type, from 0 to 1. For mixed columns, the share of the largest type family; null for empty columns
type_error_count, type_error_percentValues outside the inferred type; null for mixed columns
missing_count, missing_percentMissing cells; for JSON, also the objects where the field is absent
with_issuestrue when the column has missing values, its type is mixed or it holds ambiguous dates
normalizationValues changed by each normalization stage and the variant groups (see below); missing cells are not counted
distinct_countNumber of distinct values after normalization, even for a masked column; null when a scan limit stopped the count
examplesA few representative values
value_profileValue occurrences, complete or sampled (see selection)
semantic_typedate for date columns, otherwise the id of the scan’s primary interpretation, such as enumeration, email, phone or postal_code, or null
exposuremask, hide or show for a sensitive column, the way its values appear in examples, value_profile and preview; null otherwise
detectorsThe detectors that recognized values of the column, in scan order (see below); empty when none did
scan_detailsBounded field evidence retained from the exposed Scan result: presence, native types, missing breakdown, first and last values, string characteristics and lengths, numeric statistics, boolean counts and temporal ranges. See below

value_profile.selection is complete when all available frequencies are listed. A primary enumeration keeps this complete list even above value_examples.full_distribution_max_distinct, if the Scan frequency measure is complete. Otherwise the profile uses a bounded diverse or random sample.

Depending on the column, these objects are also present (otherwise null):

FieldPresent forContent
numericinteger and number columns with finite valuesminimum, maximum, range, mean, median (null when a scan limit stopped it), over every number of the column, including decimal commas such as 12,5
date_profileColumns containing date values, whatever their typeValid, ambiguous and invalid counts, detected formats and their breakdown, and ambiguity_evidence: the number of unambiguous values per day-month order (DMY, MDY), shown but never applied. resolved_ambiguous_order is set only by scan.detectors.date.ambiguous_order
string_profiletext columnsLength statistics, length distribution and representative examples

What the scan’s normalization did to the present values of the column. Raw values are never changed.

FieldContent
trim_count, trim_percentValues changed by trimming surrounding whitespace
collapse_internal_whitespace_count, collapse_internal_whitespace_percentValues changed by collapsing repeated internal whitespace
stagesOne item per stage, in order: raw, nfc, trim, collapse_whitespace, casefold, strip_accents (see below)
variant_group_countNormalized values written in at least two raw ways; null unless variant_group_status is complete
variant_group_statuscomplete, limited when a scan limit stopped the count, or not_applicable without text values
variant_groupsThe largest groups (scan.limits.max_variant_groups), each with its comparison key, its occurrence count, distinct_count (raw spellings), variants (value and count, at most scan.limits.max_variants_per_group) and truncated; masked for a sensitive column, empty when hidden
variant_groups_truncatedtrue when more groups exist than are listed

Each item of stages has:

FieldContent
stageStage name; raw describes the values as read
enabledfalse when the stage is turned off in scan.normalization
changed_count, changed_percentValues the stage changed, given the previous enabled stage; null for raw and disabled stages
distinct_countDistinct values after the stage; null unless distinct_status is complete
distinct_statuscomplete, limited, not_applicable (no values) or disabled

A variant group is an analytical equivalence, such as Montréal, montreal and MONTREAL, not proof that the values mean the same thing.

Each item describes what one scan detector recognized in the column. The detectors that matched values or found ambiguous or invalid ones are listed, as well as failed ones.

FieldContent
idDetector id, such as email, date or pattern:order_id
statuscomplete, or failed when the detector stopped with an error
primarytrue for the scan’s primary interpretation, shown as semantic_type
eligible_countValues the detector could examine
matched_count, matched_percentValues recognized, and their share of the eligible values
ambiguous_countValues with more than one reading, such as 01/02/2026
invalid_countValues with the right shape but an impossible content, such as 2026-02-30
formatsFormats found, each with format, count and percent of the eligible values
coverageComplete Scan coverage, including tested, not tested and not matched counts; null for a failed detector
evidenceBounded exposed examples for matched, ambiguous, invalid and not matched values; null for a failed detector
detailsDetector-specific, exposure-controlled Scan metadata
adaptiveAdaptive detection skip and probe metadata, when applicable

The standalone column pages use these Scan facts when present. first and last retain the record number and scalar type; their values, detector evidence and detector details use the same mask, hide or show policy as the rest of the profile. string_lengths, numeric, booleans and temporal are null when the corresponding Scan measure is unavailable. The complete Scan document has other fields and statuses; scan_details is a bounded selection, not a copy of the Scan document.

Each issue has a code, a severity (warning or info), an English message, a count, the affected column_ids and up to 10 row_numbers. Row numbers count the records of the dataset from 1; a CSV header is not counted.

codeSeverityMeaning
duplicate_rowswarningRows repeating an earlier row, beyond its first occurrence
missing_valueswarningMissing cells
empty_columnswarningColumns without any present value
mixed_typeswarningColumns with mixed value types
ambiguous_dateswarningDate values matching more than one day-month order
ambiguous_headerswarningBlank or repeated column names
constant_columnsinfoColumns with a single distinct present value
trimmed_cellsinfoCells changed by trimming surrounding whitespace
collapsed_whitespaceinfoCells changed by collapsing repeated internal whitespace
limited_measuresinfoColumns with measures stopped by a scan limit
excluded_recordswarningRecords excluded by the tolerant error policy, not analyzed
variant_groupsinfoValues written in several ways that normalization compares as equal

An issue is listed only when its count is above zero, except trimmed_cells and collapsed_whitespace, which are always listed.

The first records of the dataset (10 by default, set by scan.records.preview in the configuration), each with its row_number and raw values in column order. The preview size does not affect the analysis, which always reads every record. In sensitive columns, values are masked (mask) or null (hide), as given by the column’s exposure; missing cells stay as read.

For JSON, a field absent from the record is null and its position (from 0) is listed in absent. A field under an array, such as tags[], joins the values of the record with , . A JSON null is the text null.

What the scan could not measure completely in the dataset. A limited measure is never estimated.

FieldContent
measuresOne item per limited measure: column_id and path of its field (both null for a measure of the dataset), measure (its place in the scan result, such as values.cardinality, numeric.median or structure.paths), reason (such as distinct_limit, global_budget, value_too_long, field_limit or record_budget), limit and lower_bound, a proven minimum or null
untracked_observationsValues under paths beyond scan.limits.max_fields, counted but not analyzed
depth_truncated_observationsValues deeper than scan.limits.max_depth, counted but not analyzed
diagnosticsThe scan diagnostics of this dataset and of the whole scan: code, level (error or warning), message, count, dataset (null for the whole scan), path, detector and up to scan.errors.max_locations locations (record, and line for CSV or element for JSON)

The shape of a JSON dataset; null for CSV files.

FieldContent
record_typesNumber of records of each native type, such as {"object": 60}
path_countNumber of paths; null when path_status is limited
path_statuscomplete, or limited when the dataset has more paths than scan.limits.max_fields
max_depth_seenDepth of the deepest analyzed value
fieldsOne item per path, in order of discovery, containers included (see below)

Each item of fields has:

FieldContent
idScan field id
pathPath from the record, such as orders[].total
depthNumber of path segments: orders[].total has depth 3
parentPath of the parent field, null at the record root
native_typesOccurrences of each JSON type: object, array, string, integer, number, boolean, null
occurrencesValues found at the path
parent_typerecord, object, or array for array items
parent_count, present_count, absent_countParents that could hold the field, those that do and those that do not; absent_count is null for array items
present_percentShare of the parents holding the field; null for array items, which count elements
collectionFor an array analyzed as its own dataset, the id of that dataset
arraysFor arrays: count, empty_count, minimum_length, maximum_length, mean_length and item_count; null otherwise
columntrue when the field is also listed in columns
  • format_version names the experimental format family. It stays "0.1.0a" during beta.
  • format_revision increases for each structural or semantic change. It does not change for report styling or performance improvements.
  • Revisions can be incompatible, and no migration is provided. See the profile format changelog.

The profile contains values from the source file, in examples, value_profile and preview; only sensitive columns are masked. Share it as you would share the data.