PDF Pipelines: Testing preprocessing methods in our Bulk Extractor workflow

Annie Schweikert, Victor Aguilar | BitCurator Consortium

PDF Pipelines: Testing preprocessing methods in our Bulk Extractor workflow

Stanford University Libraries uses Bulk Extractor to review our digital collections for personally identifiable information. While engaged in a project of reviewing a series of PDF-heavy digital collections, we found that we wanted to test tools that specialized in parsing and extracting text from those PDFs. We are looking at “resetting” our workflow by adding an additional format-specific tool, Kreuzberg, as a preprocessing step to normalize the text extraction step before sending it downstream to Bulk Extractor. We will discuss our goals in testing Kreuzberg, our considerations for adoption, and our comparison benchmarking. In reflecting on and adapting our workflow, our goal is to produce a more transparent, format-aware analysis of our own collections.

SLIDES

This session is part of Lightning Talks Session 1 and begins at 23:55


Cite this resource:
Annie Schweikert, Victor Aguilar. (June 17, 2026). PDF Pipelines: Testing preprocessing methods in our Bulk Extractor workflow. BitCurator Consortium.