Subject: scrubbed – a native web-data cleaning pipeline written in D
Shammah Chancellor
shammah.chancellor at gmail.com
Mon Oct 5 19:56:41 UTC 2026
On Monday, 5 October 2026 at 10:21:24 UTC, Dejan Lekic wrote:
> On Sunday, 4 October 2026 at 03:17:45 UTC, Shammah Chancellor
> wrote:
>> Hey Dheads,
>>
>> It’s been a while since I’ve had an opportunity to use D and
>> really show what it can do. But I finally had time for a
>> project that I think makes D’s strengths shine.
>>
>> I’ve released scrubbed, an open-source command-line pipeline
>> for cleaning web data before using it in LLM training or other
>> text-processing workflows.
>>
>> It handles encoding repair, HTML main-content extraction, PII
>> scanning, language identification, and near-duplicate
>> detection. The goal is to replace a chain of Python tools and
>> environments with one native binary that is fast and
>> straightforward to deploy.
>>
>> The project includes benchmarks against comparable Python
>> tools.
>>
>> Website:
>> https://schancel.github.io/scrubbed/
>>
>> Source:
>> https://github.com/schancel/scrubbed
>>
>> I’d be interested in feedback on the implementation,
>> performance work, packaging, or ways the D ecosystem could
>> make the project better.
>
> Fantastic! Is it possible, by any chance, to use it as a
> library? Sure I could just grab code and adapt it to be used as
> a library but question is whether it is already designed that
> way or not?
Some parts of it definitely are; I just haven't published them to
dub yet.
There's a warc-reader, parquet library, working on an s3 client,
and a few other things. I'm eventually going to publish them once
they're more well tested.
What part are you interested in using as a library? My thought
was more that this would integrate as a binary with things like
DataTrove and spark.
More information about the Digitalmars-d-announce
mailing list