Kolmos replaces dumb byte compression with intelligent mathematical explanations. Discovered formulas and prototype rows store data 1.66×–2.12× smaller than Parquet-zstd.
As datasets accumulate, autonomous background workers elevate chunks up the compression ladder.
Baseline literal compression. Trains custom dictionary models on high-cardinality strings and bitpacks uniform integer columns, achieving fast decompression throughput.
Analyzes column vectors for deterministic mathematical relationships: linear relations (y = mx + c), arithmetic progressions (id = start + i), timestamp offsets, and string patterns.
Identifies representative prototype rows across multi-dimensional table clusters. Records are encoded as compact residual bit-deltas against cluster medoids.
Tested on 100 GB real-world analytical schema (ecommerce events & financial telemetry).