SI Back Office

Platform · updated 2026-10-11

PDF documents: getting data, forms and searchable text out of office files

For office managers and administrators with PDF price lists, filled forms and scans: how to tell which kind of PDF you have, how to check what is extracted and where a paid, checked job fits.

Start with what kind of PDF you have

A PDF is a container for pages. Some hold text positioned on the page, some hold fields that people filled in and some hold only pictures of paper. The right method, the right checks and the right promises differ for each, so the first job is to find out which you have. A two-minute test with a text search and a text selection is described in the first guide below.

  • Text-based table or list: the rows can be read from the positions of the text.
  • Fillable form returned with answers: the answers are stored as named fields.
  • Scanned pages: only pictures, so nothing can be searched until text recognition adds a hidden layer.
  • A mix of the above in one file.

The questions to ask in order

The collection linked here puts the questions in the order that saves most effort: does the file have text, is it a table or a form, does it have one layout, what should unusual entries become, how will you know it is complete, and who decides the exceptions. The worked example shows a row count that exposes mistakes a matching total hides.

  • Guide: can this PDF be read as data? Check for a text layer first.
  • Guide: how to check an extraction is right.
  • Guide: where filled-in form answers live and when they are gone.
  • Guide: what text recognition can fix on scans and how to test it.

What documentation says, and what we do with it

Microsoft documents a PDF connector for Excel that returns tables found in a PDF, and notes that rows spanning several lines may need cleaning. The pdfplumber project says its table extraction works best on machine-generated PDFs and offers no text recognition. The pypdf documentation describes reading form values by field name, and notes that flattening removes the field structure. The OCRmyPDF documentation says text recognition output depends on input quality and cannot read handwriting. We cite these only because they explain how PDFs behave. What a buyer receives is a checked spreadsheet or a searchable copy, not a tool.

The paid routes, and what none of them promise

A text-based list of up to 60 pages in one layout becomes a checked spreadsheet from £295. Up to 200 returned fillable forms from one template become one sheet from £245. Up to 500 already-scanned pages can be made searchable from £195. Each month's batch of same-layout documents can be extracted as a standing service from £395 a month. All prices are untested proposals, and payment follows agreed checks and your sign-off.

None of these promise accuracy. Results are reported as counts and sampled comparisons, with every row that could not be read with confidence listed for you to decide. None gives legal, medical, financial or tax advice. The first enquiry never includes confidential or personal documents: send counts, descriptions and an invented or redacted sample, and secure handling is agreed after scoping.

  • This page describes new offers; no delivery history is claimed.
  • Documents with personal, payment or health information are not suitable for these fixed jobs.

Sources and limits

  • Power Query PDF connector (Microsoft Learn) Checked 2026-10-11.
    • Pdf.Tables returns tables found in a PDF; where multi-line rows are not identified properly the data may need cleaning, and similar tables on consecutive pages are combined by default.
  • pdfplumber README Checked 2026-10-11.
    • Table extraction works best on machine-generated PDFs and the project offers no text recognition.
  • pypdf documentation: interactions with PDF forms Checked 2026-10-11.
    • Form field values are read by field name; flattening turns field contents into regular page content; an XFA entry can override the page content.
  • OCRmyPDF introduction Checked 2026-10-11.
    • It adds text layers to scanned image PDFs, depends on input quality and cannot recognise handwriting.