# Highlight Texts as Images in C# with IronOCR
IronOCR's `HighlightTextAndSaveAsImages` method visualizes OCR results by drawing bounding boxes around detected text (characters, words, lines, or paragraphs) and saves them as diagnostic images, enabling developers to validate OCR accuracy and debug recognition issues.
Visualizing OCR results involves rendering bounding boxes around specific text elements that the engine has detected within an image. This process overlays distinct highlights on the exact locations of individual characters, words, lines, or paragraphs, providing a clear map of recognized content.
This visual feedback is crucial for debugging and validating OCR output accuracy, showing what the software has identified and where it has made errors. When working with complex documents or troubleshooting recognition issues, visual highlighting becomes an essential diagnostic tool.
This article demonstrates IronOCR's diagnostic capabilities with its `HighlightTextAndSaveAsImages` method. This function highlights specific sections of text and saves them as images for verification. Whether building a document processing system, implementing quality control measures, or validating your OCR implementation, this feature provides immediate visual feedback on what the OCR engine detects.
*as-heading:2(Quickstart: Highlight Words in Your PDF Instantly)*
This snippet demonstrates IronOCR usage: load a PDF and highlight each word in the document, saving the result as images. Just one line to get visual feedback on your OCR results.
```cs
:title=Highlight PDF Text in One Line
new IronOcr.OcrInput().LoadPdf("document.pdf").HighlightTextAndSaveAsImages(new IronOcr.IronTesseract(), "highlight_page_", IronOcr.ResultHighlightType.Word);
```
<div class="hsg-featured-snippet">
<h3>Minimal Workflow (4 steps)</h3>
<ol>
<li>Download and install the IronOCR C# library</li>
<li>Instantiate OCR engine</li>
<li>Load the PDF document with <code>LoadPdf</code></li>
<li>Using <code>HighlightTextAndSaveAsImages</code> highlight sections of text and save them as images</li>
</ol>
</div>
## How Do I Highlight Text and Save As Images?
Highlighting text and saving it as images is straightforward with IronOCR. Load an existing PDF with `LoadPdf`, then call the `HighlightTextAndSaveAsImages` method to highlight sections of text and save them as images. This technique verifies OCR accuracy and debugs text recognition issues in your documents.
The method takes three parameters: the `IronTesseract` OCR engine, a prefix for the output filename, and an enum from `ResultHighlightType` that dictates the type of text to highlight. This example uses `ResultHighlightType.Paragraph` to highlight text blocks as paragraphs. `HighlightTextAndSaveAsImages`
[[i:(This function uses the output string prefix and appends a page identifier (e.g., "page_0", "page_1") to the output image filename for each page.)]]
This example uses a PDF with three paragraphs.
### What Does the Input PDF Look Like?
<iframe loading="lazy" src="/static-assets/ocr/how-to/highlight-texts-as-images/sample.pdf" width="100%" height="500px">
</iframe>
### How Do I Implement the Highlighting Code?
The example code below demonstrates the basic implementation using the [OcrInput class](https://ironsoftware.com/csharp/ocr/examples/csharp-ocr-input-for-iron-tesseract/).
```csharp
using IronOcr;
IronTesseract ocrTesseract = new IronTesseract();
using var ocrInput = new OcrInput();
ocrInput.LoadPdf("document.pdf");
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_page_", ResultHighlightType.Paragraph);
```
### What Do the Output Images Show?
<div class="content-img-align-center">
<div class="center-image-wrapper" style="width=50%">
<img src="/static-assets/ocr/how-to/highlight-texts-as-images/highlighted-output.png" alt="Webpage with three paragraphs, middle paragraph highlighted with red border showing text selection capability" class="img-responsive add-shadow" />
</div>
</div>
As shown in the output image above, all three paragraphs have been highlighted with a light red box. This visual representation helps developers quickly identify how the OCR engine segments the document into readable blocks.
### What Are the Different ResultHighlightType Options?
The example above used `ResultHighlightType.Paragraph` to highlight text blocks. IronOCR provides additional highlighting options through this enum. Below is a complete list of available types, each serving different diagnostic purposes.
**Character:** Draws a bounding box around every single character detected by the OCR engine. Useful for debugging character recognition or specialized fonts, particularly when working with [custom language files](https://ironsoftware.com/csharp/ocr/examples/ocr-tesseract-custom-languages/).
**Word:** Highlights each complete word identified by the engine. Ideal for validating word boundaries and proper word identification, especially when implementing [barcode and QR reading](https://ironsoftware.com/csharp/ocr/examples/csharp-ocr-barcodes/) alongside text recognition.
**Line:** Highlights every detected text line. Useful for documents with complex layouts requiring line identification verification, such as when processing [scanned documents](https://ironsoftware.com/csharp/ocr/examples/read-scanned-document/).
**Paragraph:** Highlights entire text blocks grouped as paragraphs. Perfect for understanding document layout and verifying text block segmentation, particularly useful when working with [table extraction](https://ironsoftware.com/csharp/ocr/examples/read-table-in-document/).
### How Do I Compare Different Highlight Types?
This comprehensive example demonstrates generating highlights for all different types on the same document, allowing you to compare the results:
```csharp
using IronOcr;
using System;
// Initialize the OCR engine with custom configuration
IronTesseract ocrTesseract = new IronTesseract();
// Configure for better accuracy if needed
ocrTesseract.Configuration.ReadBarCodes = false; // Disable if not needed for performance
ocrTesseract.Configuration.PageSegmentationMode = TesseractPageSegmentationMode.AutoOsd;
// Load the PDF document
using var ocrInput = new OcrInput();
ocrInput.LoadPdf("document.pdf");
// Generate highlights for each type
Console.WriteLine("Generating character-level highlights...");
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_character_", ResultHighlightType.Character);
Console.WriteLine("Generating word-level highlights...");
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_word_", ResultHighlightType.Word);
Console.WriteLine("Generating line-level highlights...");
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_line_", ResultHighlightType.Line);
Console.WriteLine("Generating paragraph-level highlights...");
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_paragraph_", ResultHighlightType.Paragraph);
Console.WriteLine("All highlight images have been generated successfully!");
```
### How Do I Handle Multi-Page Documents?
When processing multi-page PDFs or [multi-frame TIFF files](https://ironsoftware.com/csharp/ocr/examples/csharp-tesseract-multipage-tiff/), the highlighting feature automatically handles each page individually. This is especially useful when implementing [PDF OCR text extraction](https://ironsoftware.com/csharp/ocr/examples/csharp-pdf-ocr/) workflows:
```csharp
using IronOcr;
using System.IO;
IronTesseract ocrTesseract = new IronTesseract();
// Load a multi-page document
using var ocrInput = new OcrInput();
ocrInput.LoadPdf("multi-page-document.pdf");
// Create output directory if it doesn't exist
string outputDir = "highlighted_pages";
Directory.CreateDirectory(outputDir);
// Generate highlights for each page
// Files will be named: highlighted_pages/page_0.png, page_1.png, etc.
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract,
Path.Combine(outputDir, "page_"),
ResultHighlightType.Word);
// Count generated files for verification
int pageCount = Directory.GetFiles(outputDir, "page_*.png").Length;
Console.WriteLine($"Generated {pageCount} highlighted page images");
```
### What Are the Performance Best Practices?
When using the highlighting feature, consider these best practices:
1. **File Size:** Highlighted images can be large, especially for high-resolution documents. Consider the output directory's available space when processing large batches. For optimization tips, see our [fast OCR configuration guide](https://ironsoftware.com/csharp/ocr/examples/tune-tesseract-for-speed-in-dotnet/).
2. **Performance:** Generating highlights adds processing overhead. For production systems where highlights are only needed occasionally, implement them as a separate diagnostic process rather than part of the main workflow. Consider using multithreaded OCR for batch processing.
3. **Error Handling:** Always implement proper error handling when working with file operations:
```csharp
try
{
using var ocrInput = new OcrInput();
ocrInput.LoadPdf("document.pdf");
// Apply image filters if needed for better recognition
ocrInput.Deskew(); // Correct slight rotations
ocrInput.DeNoise(); // Remove background noise
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_", ResultHighlightType.Word);
}
catch (Exception ex)
{
Console.WriteLine($"Error during highlighting: {ex.Message}");
// Log error details for debugging
}
```
### How Does Highlighting Integrate with OCR Results?
The highlighting feature works with IronOCR's [result objects](https://ironsoftware.com/csharp/ocr/examples/results-objects/), allowing you to correlate visual highlights with extracted text data. This is particularly useful when you need to `track OCR progress` or validate specific sections of recognized text. The `OcrResult` class provides detailed information about each detected element, which corresponds directly to the visual highlights generated by this method.
### What If I Encounter Issues?
If experiencing issues with the highlighting feature, consult the [general troubleshooting guide](https://ironsoftware.com/csharp/ocr/troubleshooting/general-troubleshooting-ocr/) for common solutions. For specific highlighting-related problems:
- **Blank output images:** Ensure the input document contains readable text and that the OCR engine is properly [configured for your document type](https://ironsoftware.com/csharp/ocr/examples/csharp-configure-setup-tesseract/). You may need to apply [image optimization filters](https://ironsoftware.com/csharp/ocr/examples/ocr-image-filters-for-net-tesseract/) or `fixing image orientation` to improve recognition.
- **Missing highlights:** Some document types may require specific preprocessing. Try applying [image filters](https://ironsoftware.com/csharp/ocr/examples/ocr-image-filters-for-net-tesseract/) or `fixing image orientation` to improve recognition.
- **Performance issues:** For large documents, consider implementing `multithreading` to improve processing speed. Additionally, review our guide on [fixing low quality scans](https://ironsoftware.com/csharp/ocr/examples/ocr-low-quality-scans-tesseract/) if working with poor quality inputs.
### How Can I Use This for Production Debugging?
The highlighting feature serves as an excellent production debugging tool. When integrated with [abort tokens](https://ironsoftware.com/csharp/ocr/examples/abort-token/) for long-running operations and [timeouts](https://ironsoftware.com/csharp/ocr/examples/timeouts/), you can create a reliable diagnostic system. Consider implementing a debug mode in your application:
```csharp
public class OcrDebugger
{
private readonly IronTesseract _tesseract;
private readonly bool _debugMode;
public OcrDebugger(bool enableDebugMode = false)
{
_tesseract = new IronTesseract();
_debugMode = enableDebugMode;
}
public OcrResult ProcessDocument(string filePath)
{
using var input = new OcrInput();
input.LoadPdf(filePath);
// Apply preprocessing
input.Deskew();
input.DeNoise();
// Generate debug highlights if in debug mode
if (_debugMode)
{
string debugPath = $"debug_{Path.GetFileNameWithoutExtension(filePath)}_";
input.HighlightTextAndSaveAsImages(_tesseract, debugPath, ResultHighlightType.Word);
}
// Perform actual OCR
return _tesseract.Read(input);
}
}
```
### Where Should I Go Next?
Now that you understand how to use the highlighting feature, explore:
- [Creating searchable PDFs](https://ironsoftware.com/csharp/ocr/examples/tesseract-create-searchable-pdf/) from your OCR results
- [Reading specific document types](https://ironsoftware.com/csharp/ocr/tutorials/read-specific-document/) like passports or licenses
- Setting up IronOCR in your development environment with our [getting started guides](https://ironsoftware.com/csharp/ocr/get-started/windows/)
- Implementing 125 international language support for global applications
- Using the Filter Wizard to optimize image processing
For production use, remember to [obtain a license](https://ironsoftware.com/csharp/ocr/licensing/) to remove watermarks and access full functionality.
IronOCR's HighlightTextAndSaveAsImages method visualizes OCR results by drawing bounding boxes around detected text (characters, words, lines, or paragraphs) and saves them as diagnostic images, enabling developers to validate OCR accuracy and debug recognition issues.
Visualizing OCR results involves rendering bounding boxes around specific text elements that the engine has detected within an image. This process overlays distinct highlights on the exact locations of individual characters, words, lines, or paragraphs, providing a clear map of recognized content.
This visual feedback is crucial for debugging and validating OCR output accuracy, showing what the software has identified and where it has made errors. When working with complex documents or troubleshooting recognition issues, visual highlighting becomes an essential diagnostic tool.
This article demonstrates IronOCR's diagnostic capabilities with its HighlightTextAndSaveAsImages method. This function highlights specific sections of text and saves them as images for verification. Whether building a document processing system, implementing quality control measures, or validating your OCR implementation, this feature provides immediate visual feedback on what the OCR engine detects.
Quickstart: Highlight Words in Your PDF Instantly
This snippet demonstrates IronOCR usage: load a PDF and highlight each word in the document, saving the result as images. Just one line to get visual feedback on your OCR results.
new IronOcr.OcrInput().LoadPdf("document.pdf").HighlightTextAndSaveAsImages(new IronOcr.IronTesseract(), "highlight_page_", IronOcr.ResultHighlightType.Word);
C#
3Deploy to test on your live environment
Start using IronOCR in your project today with a free trial
Minimal Workflow (4 steps)
Download and install the IronOCR C# library
Instantiate OCR engine
Load the PDF document with LoadPdf
Using HighlightTextAndSaveAsImages highlight sections of text and save them as images
How Do I Highlight Text and Save As Images?
Highlighting text and saving it as images is straightforward with IronOCR. Load an existing PDF with LoadPdf, then call the HighlightTextAndSaveAsImages method to highlight sections of text and save them as images. This technique verifies OCR accuracy and debugs text recognition issues in your documents.
The method takes three parameters: the IronTesseract OCR engine, a prefix for the output filename, and an enum from ResultHighlightType that dictates the type of text to highlight. This example uses ResultHighlightType.Paragraph to highlight text blocks as paragraphs. HighlightTextAndSaveAsImages
Please note: This function uses the output string prefix and appends a page identifier (e.g., "page_0", "page_1") to the output image filename for each page.
This example uses a PDF with three paragraphs.
What Does the Input PDF Look Like?
How Do I Implement the Highlighting Code?
The example code below demonstrates the basic implementation using the OcrInput class.
using IronOcr;IronTesseract ocrTesseract = new IronTesseract();using var ocrInput = new OcrInput();ocrInput.LoadPdf("document.pdf");ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_page_", ResultHighlightType.Paragraph);
using IronOcr;
IronTesseract ocrTesseract = new IronTesseract();
using var ocrInput = new OcrInput();
ocrInput.LoadPdf("document.pdf");
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_page_", ResultHighlightType.Paragraph);
ImportsIronOcrDim ocrTesseract As New IronTesseract()Using ocrInput As New OcrInput() ocrInput.LoadPdf("document.pdf") ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_page_", ResultHighlightType.Paragraph)EndUsing
Imports IronOcr
Dim ocrTesseract As New IronTesseract()
Using ocrInput As New OcrInput()
ocrInput.LoadPdf("document.pdf")
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_page_", ResultHighlightType.Paragraph)
End Using
What Do the Output Images Show?
As shown in the output image above, all three paragraphs have been highlighted with a light red box. This visual representation helps developers quickly identify how the OCR engine segments the document into readable blocks.
What Are the Different ResultHighlightType Options?
The example above used ResultHighlightType.Paragraph to highlight text blocks. IronOCR provides additional highlighting options through this enum. Below is a complete list of available types, each serving different diagnostic purposes.
Character: Draws a bounding box around every single character detected by the OCR engine. Useful for debugging character recognition or specialized fonts, particularly when working with custom language files.
Word: Highlights each complete word identified by the engine. Ideal for validating word boundaries and proper word identification, especially when implementing barcode and QR reading alongside text recognition.
Line: Highlights every detected text line. Useful for documents with complex layouts requiring line identification verification, such as when processing scanned documents.
Paragraph: Highlights entire text blocks grouped as paragraphs. Perfect for understanding document layout and verifying text block segmentation, particularly useful when working with table extraction.
How Do I Compare Different Highlight Types?
This comprehensive example demonstrates generating highlights for all different types on the same document, allowing you to compare the results:
using IronOcr;using System;// Initialize the OCR engine with custom configurationIronTesseract ocrTesseract = new IronTesseract();// Configure for better accuracy if neededocrTesseract.Configuration.ReadBarCodes = false; // Disable if not needed for performanceocrTesseract.Configuration.PageSegmentationMode = TesseractPageSegmentationMode.AutoOsd;// Load the PDF documentusing var ocrInput = new OcrInput();ocrInput.LoadPdf("document.pdf");// Generate highlights for each typeConsole.WriteLine("Generating character-level highlights...");ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_character_", ResultHighlightType.Character);Console.WriteLine("Generating word-level highlights...");ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_word_", ResultHighlightType.Word);Console.WriteLine("Generating line-level highlights...");ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_line_", ResultHighlightType.Line);Console.WriteLine("Generating paragraph-level highlights...");ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_paragraph_", ResultHighlightType.Paragraph);Console.WriteLine("All highlight images have been generated successfully!");
using IronOcr;
using System;
// Initialize the OCR engine with custom configuration
IronTesseract ocrTesseract = new IronTesseract();
// Configure for better accuracy if needed
ocrTesseract.Configuration.ReadBarCodes = false; // Disable if not needed for performance
ocrTesseract.Configuration.PageSegmentationMode = TesseractPageSegmentationMode.AutoOsd;
// Load the PDF document
using var ocrInput = new OcrInput();
ocrInput.LoadPdf("document.pdf");
// Generate highlights for each type
Console.WriteLine("Generating character-level highlights...");
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_character_", ResultHighlightType.Character);
Console.WriteLine("Generating word-level highlights...");
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_word_", ResultHighlightType.Word);
Console.WriteLine("Generating line-level highlights...");
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_line_", ResultHighlightType.Line);
Console.WriteLine("Generating paragraph-level highlights...");
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_paragraph_", ResultHighlightType.Paragraph);
Console.WriteLine("All highlight images have been generated successfully!");
ImportsIronOcrImportsSystem' Initialize the OCR engine with custom configurationDim ocrTesseract As New IronTesseract()' Configure for better accuracy if neededocrTesseract.Configuration.ReadBarCodes = False ' Disable if not needed for performanceocrTesseract.Configuration.PageSegmentationMode = TesseractPageSegmentationMode.AutoOsd' Load the PDF documentUsing ocrInput As New OcrInput() ocrInput.LoadPdf("document.pdf") ' Generate highlights for each typeConsole.WriteLine("Generating character-level highlights...") ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_character_", ResultHighlightType.Character)Console.WriteLine("Generating word-level highlights...") ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_word_", ResultHighlightType.Word)Console.WriteLine("Generating line-level highlights...") ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_line_", ResultHighlightType.Line)Console.WriteLine("Generating paragraph-level highlights...") ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_paragraph_", ResultHighlightType.Paragraph)EndUsingConsole.WriteLine("All highlight images have been generated successfully!")
Imports IronOcr
Imports System
' Initialize the OCR engine with custom configuration
Dim ocrTesseract As New IronTesseract()
' Configure for better accuracy if needed
ocrTesseract.Configuration.ReadBarCodes = False ' Disable if not needed for performance
ocrTesseract.Configuration.PageSegmentationMode = TesseractPageSegmentationMode.AutoOsd
' Load the PDF document
Using ocrInput As New OcrInput()
ocrInput.LoadPdf("document.pdf")
' Generate highlights for each type
Console.WriteLine("Generating character-level highlights...")
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_character_", ResultHighlightType.Character)
Console.WriteLine("Generating word-level highlights...")
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_word_", ResultHighlightType.Word)
Console.WriteLine("Generating line-level highlights...")
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_line_", ResultHighlightType.Line)
Console.WriteLine("Generating paragraph-level highlights...")
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_paragraph_", ResultHighlightType.Paragraph)
End Using
Console.WriteLine("All highlight images have been generated successfully!")
How Do I Handle Multi-Page Documents?
When processing multi-page PDFs or multi-frame TIFF files, the highlighting feature automatically handles each page individually. This is especially useful when implementing PDF OCR text extraction workflows:
using IronOcr;using System.IO;IronTesseract ocrTesseract = new IronTesseract();// Load a multi-page documentusing var ocrInput = new OcrInput();ocrInput.LoadPdf("multi-page-document.pdf");// Create output directory if it doesn't existstring outputDir = "highlighted_pages";Directory.CreateDirectory(outputDir);// Generate highlights for each page// Files will be named: highlighted_pages/page_0.png, page_1.png, etc.ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, Path.Combine(outputDir, "page_"), ResultHighlightType.Word);// Count generated files for verificationint pageCount = Directory.GetFiles(outputDir, "page_*.png").Length;Console.WriteLine($"Generated {pageCount} highlighted page images");
using IronOcr;
using System.IO;
IronTesseract ocrTesseract = new IronTesseract();
// Load a multi-page document
using var ocrInput = new OcrInput();
ocrInput.LoadPdf("multi-page-document.pdf");
// Create output directory if it doesn't exist
string outputDir = "highlighted_pages";
Directory.CreateDirectory(outputDir);
// Generate highlights for each page
// Files will be named: highlighted_pages/page_0.png, page_1.png, etc.
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract,
Path.Combine(outputDir, "page_"),
ResultHighlightType.Word);
// Count generated files for verification
int pageCount = Directory.GetFiles(outputDir, "page_*.png").Length;
Console.WriteLine($"Generated {pageCount} highlighted page images");
ImportsIronOcrImportsSystem.IODim ocrTesseract As New IronTesseract()' Load a multi-page documentUsing ocrInput As New OcrInput() ocrInput.LoadPdf("multi-page-document.pdf") ' Create output directory if it doesn't exist Dim outputDir AsString = "highlighted_pages"Directory.CreateDirectory(outputDir) ' Generate highlights for each page ' Files will be named: highlighted_pages/page_0.png, page_1.png, etc. ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, Path.Combine(outputDir, "page_"), ResultHighlightType.Word) ' Count generated files for verification Dim pageCount AsInteger = Directory.GetFiles(outputDir, "page_*.png").LengthConsole.WriteLine($"Generated {pageCount} highlighted page images")EndUsing
Imports IronOcr
Imports System.IO
Dim ocrTesseract As New IronTesseract()
' Load a multi-page document
Using ocrInput As New OcrInput()
ocrInput.LoadPdf("multi-page-document.pdf")
' Create output directory if it doesn't exist
Dim outputDir As String = "highlighted_pages"
Directory.CreateDirectory(outputDir)
' Generate highlights for each page
' Files will be named: highlighted_pages/page_0.png, page_1.png, etc.
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract,
Path.Combine(outputDir, "page_"),
ResultHighlightType.Word)
' Count generated files for verification
Dim pageCount As Integer = Directory.GetFiles(outputDir, "page_*.png").Length
Console.WriteLine($"Generated {pageCount} highlighted page images")
End Using
What Are the Performance Best Practices?
When using the highlighting feature, consider these best practices:
File Size: Highlighted images can be large, especially for high-resolution documents. Consider the output directory's available space when processing large batches. For optimization tips, see our fast OCR configuration guide.
Performance: Generating highlights adds processing overhead. For production systems where highlights are only needed occasionally, implement them as a separate diagnostic process rather than part of the main workflow. Consider using multithreaded OCR for batch processing.
Error Handling: Always implement proper error handling when working with file operations:
try{ using var ocrInput = new OcrInput(); ocrInput.LoadPdf("document.pdf"); // Apply image filters if needed for better recognition ocrInput.Deskew(); // Correct slight rotations ocrInput.DeNoise(); // Remove background noise ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_", ResultHighlightType.Word);}catch (Exception ex){Console.WriteLine($"Error during highlighting: {ex.Message}"); // Log error details for debugging}
try
{
using var ocrInput = new OcrInput();
ocrInput.LoadPdf("document.pdf");
// Apply image filters if needed for better recognition
ocrInput.Deskew(); // Correct slight rotations
ocrInput.DeNoise(); // Remove background noise
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_", ResultHighlightType.Word);
}
catch (Exception ex)
{
Console.WriteLine($"Error during highlighting: {ex.Message}");
// Log error details for debugging
}
ImportsSystemTryUsing ocrInput As New OcrInput() ocrInput.LoadPdf("document.pdf") ' Apply image filters if needed for better recognition ocrInput.Deskew() ' Correct slight rotations ocrInput.DeNoise() ' Remove background noise ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_", ResultHighlightType.Word)EndUsingCatch ex AsExceptionConsole.WriteLine($"Error during highlighting: {ex.Message}") ' Log error details for debuggingEndTry
Imports System
Try
Using ocrInput As New OcrInput()
ocrInput.LoadPdf("document.pdf")
' Apply image filters if needed for better recognition
ocrInput.Deskew() ' Correct slight rotations
ocrInput.DeNoise() ' Remove background noise
ocrInput.HighlightTextAndSaveAsImages(ocrTesseract, "highlight_", ResultHighlightType.Word)
End Using
Catch ex As Exception
Console.WriteLine($"Error during highlighting: {ex.Message}")
' Log error details for debugging
End Try
How Does Highlighting Integrate with OCR Results?
The highlighting feature works with IronOCR's result objects, allowing you to correlate visual highlights with extracted text data. This is particularly useful when you need to track OCR progress or validate specific sections of recognized text. The OcrResult class provides detailed information about each detected element, which corresponds directly to the visual highlights generated by this method.
What If I Encounter Issues?
If experiencing issues with the highlighting feature, consult the general troubleshooting guide for common solutions. For specific highlighting-related problems:
Blank output images: Ensure the input document contains readable text and that the OCR engine is properly configured for your document type. You may need to apply image optimization filters or fixing image orientation to improve recognition.
Missing highlights: Some document types may require specific preprocessing. Try applying image filters or fixing image orientation to improve recognition.
Performance issues: For large documents, consider implementing multithreading to improve processing speed. Additionally, review our guide on fixing low quality scans if working with poor quality inputs.
How Can I Use This for Production Debugging?
The highlighting feature serves as an excellent production debugging tool. When integrated with abort tokens for long-running operations and timeouts, you can create a reliable diagnostic system. Consider implementing a debug mode in your application:
public class OcrDebugger{ private readonly IronTesseract _tesseract; private readonly bool _debugMode; publicOcrDebugger(bool enableDebugMode = false) { _tesseract = new IronTesseract(); _debugMode = enableDebugMode; } public OcrResultProcessDocument(string filePath) { using var input = new OcrInput(); input.LoadPdf(filePath); // Apply preprocessing input.Deskew(); input.DeNoise(); // Generate debug highlights if in debug mode if (_debugMode) { string debugPath = $"debug_{Path.GetFileNameWithoutExtension(filePath)}_"; input.HighlightTextAndSaveAsImages(_tesseract, debugPath, ResultHighlightType.Word); } // Perform actual OCR return _tesseract.Read(input); }}
public class OcrDebugger
{
private readonly IronTesseract _tesseract;
private readonly bool _debugMode;
public OcrDebugger(bool enableDebugMode = false)
{
_tesseract = new IronTesseract();
_debugMode = enableDebugMode;
}
public OcrResult ProcessDocument(string filePath)
{
using var input = new OcrInput();
input.LoadPdf(filePath);
// Apply preprocessing
input.Deskew();
input.DeNoise();
// Generate debug highlights if in debug mode
if (_debugMode)
{
string debugPath = $"debug_{Path.GetFileNameWithoutExtension(filePath)}_";
input.HighlightTextAndSaveAsImages(_tesseract, debugPath, ResultHighlightType.Word);
}
// Perform actual OCR
return _tesseract.Read(input);
}
}
Public Class OcrDebugger PrivateReadOnly _tesseract AsIronTesseract PrivateReadOnly _debugMode AsBoolean Public Sub New(Optional enableDebugMode AsBoolean = False) _tesseract = New IronTesseract() _debugMode = enableDebugMode End Sub Public Function ProcessDocument(filePath AsString) AsOcrResultUsing input As New OcrInput() input.LoadPdf(filePath) ' Apply preprocessing input.Deskew() input.DeNoise() ' Generate debug highlights if in debug mode If _debugMode Then Dim debugPath AsString = $"debug_{Path.GetFileNameWithoutExtension(filePath)}_" input.HighlightTextAndSaveAsImages(_tesseract, debugPath, ResultHighlightType.Word) End If ' Perform actual OCR Return _tesseract.Read(input)EndUsing End FunctionEnd Class
Public Class OcrDebugger
Private ReadOnly _tesseract As IronTesseract
Private ReadOnly _debugMode As Boolean
Public Sub New(Optional enableDebugMode As Boolean = False)
_tesseract = New IronTesseract()
_debugMode = enableDebugMode
End Sub
Public Function ProcessDocument(filePath As String) As OcrResult
Using input As New OcrInput()
input.LoadPdf(filePath)
' Apply preprocessing
input.Deskew()
input.DeNoise()
' Generate debug highlights if in debug mode
If _debugMode Then
Dim debugPath As String = $"debug_{Path.GetFileNameWithoutExtension(filePath)}_"
input.HighlightTextAndSaveAsImages(_tesseract, debugPath, ResultHighlightType.Word)
End If
' Perform actual OCR
Return _tesseract.Read(input)
End Using
End Function
End Class
Where Should I Go Next?
Now that you understand how to use the highlighting feature, explore:
Implementing 125 international language support for global applications
Using the Filter Wizard to optimize image processing
For production use, remember to obtain a license to remove watermarks and access full functionality.
Frequently Asked Questions
How does IronOCR highlight text in images?
IronOCR uses the `HighlightTextAndSaveAsImages` method to visualize OCR results. It draws bounding boxes around detected text such as characters, words, lines, or paragraphs, and saves them as images.
Why is text highlighting useful in OCR?
Text highlighting is useful for debugging and validating OCR output accuracy. It provides visual feedback on what the OCR engine has detected, helping to identify recognition errors in complex documents.
How can I implement text highlighting for a PDF document using IronOCR?
You can load a PDF document with `LoadPdf` and use the `HighlightTextAndSaveAsImages` method from IronOCR. This highlights specified text elements and saves the results as images.
What are the different ResultHighlightType options in IronOCR?
IronOCR offers different `ResultHighlightType` options including Character, Word, Line, and Paragraph, each serving various diagnostic purposes for text recognition.
Can IronOCR handle multi-page documents for text highlighting?
Yes, IronOCR can process multi-page documents like PDFs or TIFF files, automatically handling each page individually and generating corresponding highlighted images for each.
What should I consider when using the highlighting feature in production?
When using the highlighting feature, consider file sizes and processing overhead. It's best to implement highlighting as a separate diagnostic process if immediate feedback isn't essential.
How can I address issues with blank output images in IronOCR?
Ensure the input document contains readable text, and that the OCR engine is configured correctly. You may also need image optimization or orientation adjustments to improve results.
What performance best practices should I follow when using IronOCR highlights?
Manage large highlighted images by ensuring enough available space and consider implementing multithreading for improved speed. Proper error handling should also be in place for file operations.
Curtis Chau holds a Bachelor’s degree in Computer Science (Carleton University) and specializes in front-end development with expertise in Node.js, TypeScript, JavaScript, and React. Passionate about crafting intuitive and aesthetically pleasing user interfaces, Curtis enjoys working with modern frameworks and creating well-structured, visually appealing manuals.