PDF to text

Type conversion
PDF to text
Success
Input(Input stream)
Error
Text
Pages
Page count

Description

Extracts the text layer of a PDF, as one joined string and as an array with one entry per page. Branches on success/error; the error branch is taken for an unreadable or encrypted PDF, and also when the document carries no text at all — a scan or an image-only PDF has nothing to extract.

When to use

Use PDF to text when you need the words out of a PDF and nothing more — an invoice number to look up, a reference to match, a body of text to search. It is exact, repeatable, and costs no AI step coins, which makes it the right first choice whenever the document has a text layer and you know where in it to look.

What it cannot do is understand the page. A PDF stores text as positioned fragments with no reliable reading order, so a multi-column layout, a table, or a form with labels scattered around the page can come out interleaved. If you find yourself writing string surgery to undo that, stop and wire the stream into a Flovello AI step instead: it reads the rendered page, so layout is something it sees rather than something it has to reconstruct.

The two work well together on long documents. Extract here, loop the pages array with For each, and send only the page you care about to the AI step — the AI’s cost scales with pages, so narrowing first is what keeps a 200-page document affordable.

Scans take the error branch. A PDF produced by a scanner or a phone camera holds images, not text, and there is genuinely nothing for this node to return. Rather than succeed with an empty string, it fails with a log line saying so, so a pipeline can branch to the AI step on the error pin and handle both kinds of PDF in one flow. Encrypted and corrupt PDFs take the same branch.

Pins

Input pins

Pin Type Default Notes
Input Input stream — The PDF's bytes. Wire from Open email attachment, Open FTP file, or an HTTP response body. The stream is read fully and closed.

Output pins

Pin Type Notes
Text string The whole document's text, pages joined with a newline. Available on the `success` branch.
Pages string[] One entry per page, in document order. A page with no text layer contributes an empty string rather than being skipped, so an index into this array is always the real page number minus one.
Page count integer How many pages the document has — the length of `pages`.

Execution pins

Pin Direction
In Input
Success Output
Error Output

Example

Read an emailed invoice. After Find emails and For each, Break the attachment and Branch on contentType equals application/pdf. On the true path, wire the attachment into Open email attachment, and its Content stream into PDF to text. On the Success output, feed Text into String contains to check for a purchase order number.

For a fallback that handles scanned invoices too, wire the Error output of PDF to text into a second Open email attachment node fed from the same attachment, and that node’s Content into a Flovello AI step. The second Open is not optional: an input stream can only be read once, so the one PDF to text consumed is spent. Digital PDFs then cost nothing to read, and only the scans spend coins.

See also