The OCR node extracts text from images and document scans using Tesseract (via Tess4J). The recognized text is stored so it can be indexed and searched.
Kind |
|
Applies to |
Image |
Inputs |
None |
Output keys |
|
Requirements |
Tesseract / Tess4J plus its |
Persists to |
|
Configuration
OCRNodeOptions:
| Option | Meaning |
|---|---|
|
Path to the Tesseract |
|
Recognition language(s), e.g. |
Use Cases
-
Searchable scans — index text from screenshots, receipts and scanned documents.
-
Document workflows — surface embedded text for classification or routing.
-
Compliance — find images that contain particular text.