Skip to content
Beta · version 0.6.1 · commands and the JSON format may still change.

Inspect format

tabalyst inspect data.json writes data.json-inspect.json beside the source: a JSON document that says how Tabalyst understands the source and holds the rules to read it. This page describes format tabalyst.inspect, version 0.1.0a, revision 1, kinds json and excel. The format is experimental: it can change incompatibly between releases. Changes are listed in the Inspect format changelog. For the commands, see Inspect JSON and JSONL files and Inspect Excel workbooks. The kind json is described first; the kind excel follows the same top level with its own detection and config.

{
"format": "tabalyst.inspect",
"format_version": "0.1.0a",
"format_revision": 1,
"inspect": {
"kind": "json",
"tabalyst_version": "0.5.0",
"generated_at": "2026-10-01T05:34:13Z",
"note": "Edit only the \"config\" section. ..."
},
"source": {
"name": "orders.json",
"format": "json",
"size_bytes": 60720,
"sha256": "ece53017..."
},
"detection": {"...": "see below"},
"warnings": [],
"config": {
"structure": {"dataset_path": "$.customers[]"},
"flatten": {"enabled": true, "separator": ".", "max_depth": null},
"arrays": {"mode": "preserve"},
"errors": {"policy": "strict"}
}
}

The sections always come in this order, with config last. The file is UTF-8 with two-space indentation. It holds no absolute path: source.name is a file name, and the other paths point inside the source, such as $.customers[].

SectionWritten byYou edit itEffect on scans
format, format_version, format_revisionTabalystNoA version Tabalyst does not read is refused.
inspectTabalystNoNone.
sourceTabalystNoNone: it describes the source when it was inspected.
detectionTabalystNoNone.
warningsTabalystNoNone.
configTabalyst once, then youYesThe rules tabalyst scan and tabalyst report apply.

For the same source, the same version of Tabalyst and the same settings, two inspections write the same file except inspect.generated_at.

KeyMeaning
inspect.kindThe kind of Inspect: json for JSON, JSONL and NDJSON sources, excel for workbooks.
inspect.tabalyst_versionThe version of Tabalyst that wrote the file.
inspect.generated_atUTC time of the inspection, the only value that varies.
inspect.noteA reminder to edit only config.
source.nameFile name of the source, with its extension.
source.formatjson, jsonl for .jsonl and .ndjson files, or excel for .xlsx and .xlsm files.
source.size_bytesBytes read.
source.sha256SHA-256 of every byte read.

source.sha256 records the source at inspection time. It is never used to decide whether a stored scan can be reused.

KeyMeaning
scope.structureAlways complete: candidates, element counts and element types come from reading the whole source.
scope.detailAlways bounded: fields, depths and nesting come from the first records of each candidate.
scope.limitsThe bounds applied: records (first records observed per candidate, 1,000) and fields (distinct fields tracked per candidate, 1,000).
scope.discovery_max_depthHow many keys deep arrays were searched, from scan.json.discovery_max_depth (default 3).
scope.candidatestruncated when more than 100 arrays were found. Absent otherwise.
root.typeType of the root: object, array, string, number, boolean, null, or lines for JSONL.
candidatesThe candidate collections, in document order.
selectionThe collection proposed, or none.
linesJSONL only. Exact counts of lines: read (non-blank), blank, objects, invalid, not_object. invalid includes lines longer than scan.limits.max_line_bytes.

A candidate collection is the root array, or an array reachable from the root through object keys only, at most scan.json.discovery_max_depth keys deep. Arrays inside the records of a candidate are not candidates: they stay fields of the records. A JSONL file has one candidate, $[].

Key of a candidateMeaning
pathAbsolute path, such as $.customers[].
elementsExact number of elements; for JSONL, the number of object lines.
element_typesExact number of elements per type, among object, array, string, number, boolean, null.
eligibletrue when the candidate has at least one element and every element is an object.
ineligible_reasonOnly when not eligible: empty or non_object_elements.
observationDetail from the first records: records observed, distinct fields, max_depth (keys and []), nested_objects, arrays, and complete, false when a bound cut the observation.

A field that first appears after the observed records is absent from detection; the scan finds it. Observations are never exhaustive when complete is false.

selection.path is the collection proposed, or null. selection.basis says why:

basisMeaning
root_arrayThe root is an array of objects.
jsonl_recordsThe source is JSONL and has at least one object line.
only_eligible_candidateOne candidate is eligible.
dominant_candidateSeveral are eligible and the largest has at least 10 times the elements of the next one, which is given in selection.over.
ambiguousSeveral are eligible and none stands out. Nothing is selected.
no_eligible_candidateNo candidate is eligible, or the root is not an object or an array. Nothing is selected.
candidates_truncatedMore than 100 arrays exist, so the list cannot prove that one stands out. Nothing is selected.

The names of properties play no part. A tie never selects. When nothing is selected, config.structure.dataset_path is null and tabalyst scan and tabalyst report stop until you set it.

A list of entries with code, level (warning or info) and message, plus the keys that apply: path, reason, count, and locations, a list of {"record": n, "line": n} of at most scan.errors.max_locations entries (the count is always complete). Warnings never change how a scan reads the source.

CodeLevelWhenKeys
ambiguous_collectionswarningbasis is ambiguouscount of eligible candidates
no_collectionwarningbasis is no_eligible_candidatereason: scalar_root, no_array or no_eligible_array
candidates_truncatedwarningThe candidate list was cut at 100count kept
candidate_not_eligibleinfoOne per ineligible candidate, for the first 10 onlypath, reason (empty, non_object_elements)
candidate_not_eligible_truncatedinfoMore than 10 candidates are ineligiblecount of ineligible candidates
invalid_lineswarningJSONL lines that are not valid JSON, or too longcount, locations
non_object_lineswarningJSONL lines that are valid JSON but not objectscount, locations
configured_path_not_foundwarningAfter a new inspection, the kept dataset_path is not an array of the sourcepath
source_name_mismatchinfosource.name of the existing file differs from the source inspectedpath: the recorded name

no_collection reasons: scalar_root (the root is not an object or an array), no_array (an object with no array reachable through keys) and no_eligible_array (arrays exist and none is eligible, such as an array of plain values). configured_path_not_found is only raised for a path the candidates can judge: one reachable through keys alone within the discovery depth. A deeper path, or one that crosses an array, is checked by the scan.

The part you edit. Omit a key to keep the default of the layer below (see Which setting wins).

KeyTypeDefault when omittedWritten by Inspect
structure.dataset_pathstring or nullnull: not decidedThe selection, or null
flatten.enabledbooleantrueThe effective value
flatten.separatorstring"."The effective value
flatten.max_depthinteger or nullnull: no limitThe effective value
arrays.modestring"preserve""preserve"
errors.policy"strict", "tolerant" or nullnull: the default of the formatThe resolved value, never null

Rules:

  • structure.dataset_path is an absolute path that starts with $ and ends with [], such as $.customers[] or $["a.b"][]. For a JSONL source it must be $[] or null. Equal paths in different spellings are the same path. Writing null does not select the automatic discovery of tabalyst.scan().
  • flatten.separator is exactly one character that is not a letter, a digit, _, whitespace, a control character or one of [ ] " \ $. It changes the names of fields, never the paths of collections, which always use .. A key that holds the separator is written ["a/b"], so a key a.b and the nested keys a then b stay two distinct fields.
  • flatten.max_depth is null or an integer from 1 to 1,000. Depth counts keys and []: address is 1, address.city is 2, orders[].amount is 3. An object or array at that depth is kept whole as a complex value. enabled: false is a depth of 1.
  • arrays.mode accepts preserve only. ignore and explode are refused.
  • errors.policy is strict for JSON and tolerant for JSONL when it is null.

A config is kept as you wrote it when Inspect runs again: the same keys and values, nothing added. Every setting that influences the dataset, the fields or the records analyzed takes part in the identity of a scan.

For the file beside a source, Tabalyst validates the top-level format, format_version and format_revision, inspect.kind and config; the other sections are informative, so a damaged detection never stops a scan. It refuses, with exit code 2 and the name of the file:

CaseMessage names
Not a JSON object, or another formatThe file
A format_version or format_revision it does not readBoth versions, and the two ways out: edit the file or tabalyst inspect --force
An unknown or missing inspect.kindThe kinds it reads
A missing config, an unknown key at any depth, a value of the wrong type (a number or boolean written as a string included) or a value outside the supported setThe key, such as config.flatten.sepator

There is no migration during the beta, and Tabalyst never replaces an invalid file with an automatic choice.

When a JSON source or a workbook has no Inspect file and nothing on the command line or in a --config file names its collection, tabalyst scan and tabalyst report inspect it and keep a copy of the result in Tabalyst’s local storage, next to the stored scan. The copy is rebuilt whenever the source content, the format version, the version of Tabalyst or the discovery depth (JSON) or header search (Excel) differ. It is disposable and is never read when an Inspect file exists beside the source.

tabalyst inspect sales.xlsx writes sales.xlsx-inspect.json with the same top level, the same zones and the same reading rules as above. Only inspect.kind (excel), source.format (excel), detection and config differ.

{
"format": "tabalyst.inspect",
"format_version": "0.1.0a",
"format_revision": 1,
"inspect": {
"kind": "excel",
"tabalyst_version": "0.5.1",
"generated_at": "2026-10-04T22:36:19Z",
"note": "..."
},
"source": {
"name": "sales.xlsx",
"format": "excel",
"size_bytes": 13344,
"sha256": "35657227..."
},
"detection": {
"scope": {"structure": "complete", "header_scan_rows": 50},
"workbook": {"sheets": 3, "tables": 0},
"candidates": [
{
"path": "$.Orders", "kind": "sheet", "sheet": "Orders",
"visible": true,
"range": "A4:I124", "header_row": 4, "elements": 120, "eligible": true,
"observation": {
"columns": 9, "column_names": ["order_id", "..."],
"blank_rows": 0, "merged_ranges": 0
}
},
{"path": "$.Notes", "kind": "sheet", "sheet": "Notes", "visible": true,
"elements": 0, "eligible": false, "ineligible_reason": "no_header"}
],
"selection": {
"path": "$.Orders", "basis": "dominant_candidate", "over": "$.Regions"
}
},
"warnings": [],
"config": {"structure": {"dataset_path": "$.Orders", "header_row": null}}
}
KeyMeaning
scope.structureAlways complete: every sheet and table is read.
scope.header_scan_rowsRows searched for the header of a sheet, from its first filled row (50).
scope.candidatestruncated when more than 100 candidates were found. Absent otherwise.
workbookThe number of sheets and of named tables.
candidatesThe sheets and named tables, in workbook order; the tables of a sheet follow it.
selectionThe table proposed, or none.
Key of a candidateMeaning
path$.Sheet for a sheet, $.Sheet.Table for a named table; a name that is not a plain identifier is quoted, $["Q1 2026"].
kindsheet or table.
sheet, tableThe names; table only for a named table.
visiblefalse for a hidden sheet, and for the tables on it.
rangeA1 range of the header and the data, such as A4:I124. Absent without a table.
header_rowThe 1-based sheet row of the header. Absent without a table.
elementsFilled data rows; 0 for a candidate with none.
eligibletrue when the candidate has a header and at least one data row.
ineligible_reasonOnly when not eligible: no_cells, no_header, no_data_rows or has_tables.
observationcolumns, column_names (the first 100), blank_rows among the data and merged_ranges in the table. Absent without a table.

A sheet that holds named tables is not eligible (has_tables): its tables are. selection.basis takes the values of the JSON kind that make sense for a table: only_eligible_candidate, dominant_candidate (the largest eligible table has at least 10 times the data rows of the next one, given in selection.over), ambiguous, no_eligible_candidate and candidates_truncated. Hidden sheets compete only when no visible table is eligible. Sheet names play no part, and a tie never selects.

The entries have the keys of the JSON kind. ambiguous_collections, no_collection (reason no_eligible_table), candidates_truncated, candidate_not_eligible, candidate_not_eligible_truncated, configured_path_not_found and source_name_mismatch are raised as for JSON. Five codes are specific to workbooks, all at level warning, with the path of the table and a count:

CodeWhen
blocks_not_splitBlank rows lie among the data of a sheet; count of blank rows.
duplicate_headersThe header repeats names; count of repeated names.
blank_headersThe header leaves cells empty; count of empty cells.
merged_cellsMerged ranges lie in the data; count of ranges.
multi_level_headerMerged cells lie on the header row; count of ranges.
KeyTypeDefault when omittedWritten by Inspect
structure.dataset_pathstring or nullnull: not decidedThe selection, or null
structure.header_rowinteger or nullnull: the detected rownull
  • structure.dataset_path is an absolute path with the workbook as root and one or two keys: $.Sales for a sheet, $.Sales.Orders for a named table of that sheet. Equal paths in different spellings are the same path ($["Sales"]["Orders"] is $.Sales.Orders). A path that ends with [], has three keys or does not start with $ is refused.
  • structure.header_row is a 1-based row number (an integer from 1; a string, a boolean or 0 is refused) that replaces the detected header row of a sheet. It does not apply to a named table. A new file takes it from scan.excel.header_row of a --config file, and writes null when there is none.
  • flatten, arrays and errors do not exist for a workbook, and an unknown key is an error that names it, as in the JSON kind.

tabalyst scan and tabalyst report apply config as the settings excel.dataset_path and excel.header_row of the scan configuration. An Inspect file written for another kind than the source, such as a JSON Inspect file beside a workbook, is not replaced silently: tabalyst inspect --force replaces it.