Subject: scrubbed – a native web-data cleaning pipeline written in D
Shammah Chancellor
shammah.chancellor at gmail.com
Sun Oct 4 03:17:45 UTC 2026
Hey Dheads,
It’s been a while since I’ve had an opportunity to use D and
really show what it can do. But I finally had time for a project
that I think makes D’s strengths shine.
I’ve released scrubbed, an open-source command-line pipeline for
cleaning web data before using it in LLM training or other
text-processing workflows.
It handles encoding repair, HTML main-content extraction, PII
scanning, language identification, and near-duplicate detection.
The goal is to replace a chain of Python tools and environments
with one native binary that is fast and straightforward to deploy.
The project includes benchmarks against comparable Python tools.
Website:
https://schancel.github.io/scrubbed/
Source:
https://github.com/schancel/scrubbed
I’d be interested in feedback on the implementation, performance
work, packaging, or ways the D ecosystem could make the project
better.
More information about the Digitalmars-d-announce
mailing list