Dataset versioning by blocks

Splits text files into fixed-size blocks, computes a SHA-256 digest per block and per whole file, reports block and unique-block counts, the bytes saved by deduplication and builds a .dvc-like manifest so only the dataset description goes into git. Everything is computed locally and no file is uploaded anywhere.

Fill in the fields, run the tool and review the result. Your input is not added to a public page. Use the learning mode for calculation details.