Subject: scrubbed – a native web-data cleaning pipeline written in D

Shammah Chancellor shammah.chancellor at gmail.com
Sun Oct 4 03:17:45 UTC 2026


Hey Dheads,

It’s been a while since I’ve had an opportunity to use D and 
really show what it can do. But I finally had time for a project 
that I think makes D’s strengths shine.

I’ve released scrubbed, an open-source command-line pipeline for 
cleaning web data before using it in LLM training or other 
text-processing workflows.

It handles encoding repair, HTML main-content extraction, PII 
scanning, language identification, and near-duplicate detection. 
The goal is to replace a chain of Python tools and environments 
with one native binary that is fast and straightforward to deploy.

The project includes benchmarks against comparable Python tools.

Website:
https://schancel.github.io/scrubbed/

Source:
https://github.com/schancel/scrubbed

I’d be interested in feedback on the implementation, performance 
work, packaging, or ways the D ecosystem could make the project 
better.



More information about the Digitalmars-d-announce mailing list