Microsoft Form Recognizer
Tenant-level settings for Azure Form Recognizer, enabling OCR, table extraction, and layout analysis on complex or scanned PDFs.
Last updated
Was this helpful?
Tenant-level settings for Azure Form Recognizer, enabling OCR, table extraction, and layout analysis on complex or scanned PDFs.
The Settings > Configuration > Parser section includes Microsoft Form Recognizer settings that control how the Aisera platform routes PDFs to Azure's Form Recognizer service for advanced parsing. These are tenant-level defaults that can be overridden at the data source level.
Type
Checkbox
Default
Disabled
When enabled, the parser routes qualifying PDFs to Microsoft's Form Recognizer service instead of standard PDF-to-HTML conversion. Form Recognizer produces structured HTML that more accurately reflects complex layouts, including multi-column text, dense tables, and scanned content. Enable this for data sources containing PDFs that standard parsing handles poorly.
This setting requires two additional configurations before it has any effect: the tenant must have a Form Recognizer integration configured, and Pdf Names must specify which PDFs to route. Enabling this checkbox without both in place routes no documents to Form Recognizer.
Additional Azure service charges apply for each document analyzed. Configure Pdf Names carefully to limit scope. By enabling this feature, you agree that Aisera will send documents to Microsoft's Form Recognizer service for analysis.
See also: Pdf Names
Type
Checkbox
Default
Disabled
When enabled, the parser extracts the location and content of images embedded in a PDF alongside the standard Form Recognizer text analysis. For any region where Form Recognizer would otherwise output text overlapping an image boundary, the parser replaces that text with the image itself, embedded inline in the parsed output. This preserves diagrams, charts, and technical illustrations at their original positions rather than losing them to garbled text. Enable this for PDFs where embedded images are integral to the content, such as technical documentation or knowledge base articles with diagrams.
This setting applies only to PDFs with selectable text that also contain embedded images. It does not apply to fully scanned image-only PDFs processed through OCR.
See also: Microsoft Form Recognizer (Additional charges may occur)
Last updated
Was this helpful?
Was this helpful?
