PDF to text
Description
Extracts the text layer of a PDF, as one joined string and as an array with one entry per page. Branches on success/error; the error branch is taken for an unreadable or encrypted PDF, and also when the document carries no text at all — a scan or an image-only PDF has nothing to extract.
When to use
Use PDF to text when you need the words out of a PDF and nothing more — an invoice number to look up, a reference to match, a body of text to search. It is exact, repeatable, and costs no AI step coins, which makes it the right first choice whenever the document has a text layer and you know where in it to look.
What it cannot do is understand the page. A PDF stores text as positioned fragments with no reliable reading order, so a multi-column layout, a table, or a form with labels scattered around the page can come out interleaved. If you find yourself writing string surgery to undo that, stop and wire the stream into a Flovello AI step instead: it reads the rendered page, so layout is something it sees rather than something it has to reconstruct.
The two work well together on long documents. Extract here, loop the pages
array with For each, and send only the page you care about to the AI step — the
AI’s cost scales with pages, so narrowing first is what keeps a 200-page document
affordable.
Scans take the error branch. A PDF produced by a scanner or a phone camera holds images, not text, and there is genuinely nothing for this node to return. Rather than succeed with an empty string, it fails with a log line saying so, so a pipeline can branch to the AI step on the error pin and handle both kinds of PDF in one flow. Encrypted and corrupt PDFs take the same branch.
Pins
Input pins
| Pin | Type | Default | Notes |
|---|---|---|---|
| Input | Input stream | — | The PDF's bytes. Wire from Open email attachment, Open FTP file, or an HTTP response body. The stream is read fully and closed. |
Output pins
| Pin | Type | Notes |
|---|---|---|
| Text | string | The whole document's text, pages joined with a newline. Available on the `success` branch. |
| Pages | string[] | One entry per page, in document order. A page with no text layer contributes an empty string rather than being skipped, so an index into this array is always the real page number minus one. |
| Page count | integer | How many pages the document has — the length of `pages`. |
Execution pins
| Pin | Direction |
|---|---|
| In | Input |
| Success | Output |
| Error | Output |
Example
Read an emailed invoice. After Find emails and For each, Break the attachment and
Branch on contentType equals application/pdf. On the true path, wire the
attachment into Open email attachment, and its Content stream into PDF to text.
On the Success output, feed Text into String contains to check for a purchase
order number.
For a fallback that handles scanned invoices too, wire the Error output of PDF to text into a second Open email attachment node fed from the same attachment, and that node’s Content into a Flovello AI step. The second Open is not optional: an input stream can only be read once, so the one PDF to text consumed is spent. Digital PDFs then cost nothing to read, and only the scans spend coins.