JLGet in touch

04 / DOCUMENT AUTOMATION

Bulk work-order PDF extraction

I built a Python desktop application that reads work-order PDFs in batches and consolidates site, technology, equipment and validation data into a structured report. It replaces opening each PDF just to find and copy the same fields.

PythonpdfplumberPandasCustomTkinterCSV / Excel
Application scope61 structured fields per PDFIncluding 21 equipment configuration columns.

My contribution: PDF table extraction, field normalization, batch processing, and a desktop interface with CSV and Excel export. Decision it supports: reviewing many received OTs from one structured report.

Technical approach · extraction and a Python excerpt

The parser uses table positions to preserve site IDs when adjacent PDF cells are blank, with regex fallbacks for inconsistent layouts. Combined equipment configurations are split into separate columns before the batch becomes a Pandas DataFrame. A background thread keeps the desktop interface responsive during extraction.

Excerpt from the extractor’s normalization function:

def strip_all_spaces(val):
    if val is None:
        return ""
    return re.sub(r'\s+', '', str(val))

def split_config(val, expected_len):
    if not val:
        return [""] * expected_len
    parts = [strip_all_spaces(p) for p in str(val).split('/')]
    if len(parts) < expected_len:
        parts.extend([""] * (expected_len - len(parts)))
    return parts[:expected_len]

This excerpt contains no customer data. It shows how a variable-length PDF field becomes a consistent set of columns.

Anonymized desktop PDF extractor showing 47 selected PDFs, 47 extracted rows, 61 fields, a populated results table and CSV and Excel export buttons
Desktop application after a 47-PDF batch. This image is based on a real screen capture; names and operational identifiers have been replaced with fictional examples. View larger
INTERACTIVE DEMO

From PDF batch to usable report

47 synthetic work orders · 12 representative fields
PDF work orders→Read three pages→Split and normalize→CSV or Excel
PDFs in sample batch—
Rows extracted—
Fields in full application61
Fields in this demo12
Consolidated report

Search the extracted rows, then open an OT to see its source sections.

—
PDF fileSiteRegionOTNew radioNew BBUValidation
EXTRACTED RECORD
—
Download the sample reportAll 47 fictional records, with 12 representative fields.

CSV uses UTF-8 with BOM for Excel compatibility.

This browser demo illustrates the output of the desktop application; it does not read uploaded PDFs. The full application extracts 61 fields, previews the results, logs progress and exports CSV or Excel.