Tabalyst Scan
Tabalyst Scan is the analysis engine of Tabalyst. It reads a CSV, JSON, JSONL or Excel file once and describes every field in one JSON document: presence, native types, missing values, frequencies, exact statistics, normalization variants, technical type and the result of every detector.
tabalyst scan customers.csvtabalyst scan orders.json --collection "$.customers[]"tabalyst scan events.jsonltabalyst scan sales.xlsxWithout -o or -d, the scan is stored for reuse in Tabalyst’s local storage.
-o or -d writes a standalone <stem>.scan.json document instead.
Tabalyst Report is built on it, and reuses a stored scan
when the source has not changed.
What it detects
Section titled “What it detects”Numbers with decimal commas, dates, booleans, enumerations, email addresses, URLs, phone numbers, postal codes, currency amounts, percentages, quantities, UUIDs and IP addresses. Values of sensitive fields are masked by default. Memory depends on configurable limits, not on the number of records. Results are written atomically: an interrupted scan never leaves a partial file.
- Scan CSV and JSON files: commands, JSON, JSONL and Excel sources, reuse, settings and the Python API.
- Scan format: the structure of the scan document, and its changelog.
- Manage query caches:
tabalyst cache infoandtabalyst cache clean. - Configuration: the
scansettings. - Tabalyst Inspect: how JSON sources are read before a scan, and for workbooks.