Subject: scrubbed – a native web-data cleaning pipeline written in D

Shammah Chancellor shammah.chancellor at gmail.com
Mon Oct 5 19:56:41 UTC 2026


On Monday, 5 October 2026 at 10:21:24 UTC, Dejan Lekic wrote:
> On Sunday, 4 October 2026 at 03:17:45 UTC, Shammah Chancellor 
> wrote:
>> Hey Dheads,
>>
>> It’s been a while since I’ve had an opportunity to use D and 
>> really show what it can do. But I finally had time for a 
>> project that I think makes D’s strengths shine.
>>
>> I’ve released scrubbed, an open-source command-line pipeline 
>> for cleaning web data before using it in LLM training or other 
>> text-processing workflows.
>>
>> It handles encoding repair, HTML main-content extraction, PII 
>> scanning, language identification, and near-duplicate 
>> detection. The goal is to replace a chain of Python tools and 
>> environments with one native binary that is fast and 
>> straightforward to deploy.
>>
>> The project includes benchmarks against comparable Python 
>> tools.
>>
>> Website:
>> https://schancel.github.io/scrubbed/
>>
>> Source:
>> https://github.com/schancel/scrubbed
>>
>> I’d be interested in feedback on the implementation, 
>> performance work, packaging, or ways the D ecosystem could 
>> make the project better.
>
> Fantastic! Is it possible, by any chance, to use it as a 
> library? Sure I could just grab code and adapt it to be used as 
> a library but question is whether it is already designed that 
> way or not?

Some parts of it definitely are; I just haven't published them to 
dub yet.

There's a warc-reader, parquet library, working on an s3 client, 
and a few other things. I'm eventually going to publish them once 
they're more well tested.

What part are you interested in using as a library? My thought 
was more that this would integrate as a binary with things like 
DataTrove and spark.


More information about the Digitalmars-d-announce mailing list