Beta · version 0.6.1 · commands and the JSON format may still change.
Scan format changelog
A scan document written by tabalyst scan identifies its contract with three
fields:
{ "format": "tabalyst.scan", "format_version": "0.1.0a", "format_revision": 6}format names the kind of document; the JSON profile of tabalyst report has
its own changelog. format_version names the
experimental compatibility family, 0.1.0a retained during beta.
format_revision is a monotonic integer incremented for each meaningful
structural or semantic change. It does not change for performance improvements,
documentation or corrections that keep the contract.
Beta revisions may be incompatible. No automatic migration is provided.
tabalyst report --scan reads the current revision only: scan the source
again to report on an older document.
Revision 6
Section titled “Revision 6”source.formatcan beexcel, for.xlsxand.xlsmworkbooks.source.excelrecords the table read (dataset_path,sheet,table,range,header_row,header) and isnullfor other sources.source.encodingisnullfor a workbook.config.excel(dataset_path,header_row) selects the table of a workbook. Adding it changedconfig_sha256of every configuration, so documents of revision 5 are refused and must be scanned again; a stored scan of an older revision is replaced bytabalyst scanandtabalyst report.- An Excel source has one dataset,
rows, of kindtable, withcolumnfields as for CSV and typed values: integers, numbers, booleans and strings, dates being ISO text.
Revision 5
Section titled “Revision 5”configrecords the rules the scan applied, defaults resolved:errors.policyisstrictortolerant, nevernull; the setting itself acceptsnull, meaning the default of the source format (strictfor CSV and JSON).config_sha256is the hash of that resolved configuration.config.jsongainsflatten(enabled,separator,max_depth) andarrays(mode). Their defaults keep the previous behavior. Whenflatten.enabledisfalse,separatorandmax_depthare recorded as their defaults, since they change nothing.source.formatcan bejsonl: files ending in.jsonlor.ndjsonare read as one dataset$[]whose records are the object lines. Theirtolerantdefault policy excludes a line that is not valid JSON (jsonl_invalid_line), is not an object (jsonl_record_not_object) or exceedslimits.max_line_bytes(jsonl_line_too_long), with locations{record, line}.scope.collectionsisnull.config.limitsgainsmax_line_bytes.- The JSON reader applies
flatten: a container atmax_depthsegments is kept whole (type and array length), its children are not observed and are not counted as truncated. The fielddisplayis joined withflatten.separator. config.json.collectionsis recorded in canonical spelling ($.orders[]for$["orders"][]), so equal paths give equal hashes.- The datasets of a JSON scan made by
tabalyst scanortabalyst reportcome from Inspect:scope.collectionshasmodeexplicitand there is nodocumentdataset$.tabalyst.scan()alone keeps the automatic discovery (modeauto). The document structure does not change. - The identity of a scan is its source SHA-256, its
config_sha256andengine.version. Another engine version is no longer reused as a stored scan, and a source is compared withsource.sha256whatever its modification time;source.modified_atstays informative.
Revision 4
Section titled “Revision 4”- Rare detectors: after the warm-up, a detector that recognized at most
detection.rare_shareof its distinct values (default 0.001, 10 values of a 10,000-value warm-up), and none of those in its second half, is skipped too, like a detector that recognized nothing. A sensitive detector that found only invalid values is never skipped this way. Setdetection.rare_shareto0for the behavior of revision 3. adaptivegainswarmup_reactions: the warm-up values the detector recognized,0for a detector that recognized nothing.- A skipped rare detector reports
detector_skipped_reactedwhen at least 10 probed occurrences were recognized, more often thanrare_shareallows. - New setting in
config:detection.rare_share.
Revision 3
Section titled “Revision 3”- Adaptive detection: after the first
detection.warmup_valuesdistinct values of a field (default 10,000), a detector that recognized none of them skips the other values of the field, except about one indetection.probe_interval(default 100). Skipped values count asnot_tested.numberanddateare never skipped. Setdetection.warmup_valuesto0for the exhaustive detection of revision 2. - Each
completedetector result gainsadaptive:null, or the warm-up size, the values not tested and the index of a newdetector_skipped_reactedwarning when a probe was recognized. - New settings in
config:detection.warmup_valuesanddetection.probe_interval.
Revision 2
Section titled “Revision 2”- Each dataset gains
records: records with missing values, empty records, duplicate records and a preview of the first records, sensitive values exposed as in the rest of the document. - Duplicate records are counted within the budget of the new
limits.max_tracked_recordssetting; beyond it, the count is a lower bound with reasonrecord_budget, and arecord_budgetwarning is reported. - New settings in
config:records.preview,records.duplicates,limits.max_tracked_recordsandlimits.max_listed_records.
Revision 1
Section titled “Revision 1”- First published scan format, written by
tabalyst scan. - Top level: engine and detector versions, scan status, source identity with size, modification time and SHA-256, effective configuration and its SHA-256, scope and diagnostics.
- CSV and JSON datasets, with fields identified by paths, exact presence, native types, string categories and a configurable missing count.
- Value measures: cardinality, frequencies, samples, first and last values, string characteristics and lengths, exact numeric statistics, booleans and dates, wrapped in measure envelopes when they can be limited.
- Normalization version 1 with change counters, distinct values per stage and variant groups.
- Technical type, thirteen built-in detectors, declarative patterns, interpretations and the masking of sensitive values.