# How to Debug OCR in C#
IronOCR enables you to detect OCR failures at the source, assess recognition quality at the word and character level, and monitor long-running jobs in real time. Built-in tools such as diagnostic file logging, a typed exception hierarchy, per-result confidence scoring, and the `OcrProgress` event support these workflows in production pipelines.
This guide walks through working examples for each: enabling diagnostic logging, handling typed exceptions, validating output with confidence scores, monitoring job progress in real time, and isolating errors in batch pipelines.
*as-heading:2(Quickstart: Enable full OCR diagnostic logging)*
Set `LogFilePath` and `LoggingMode` on the `Installation` class before the first `Read` call. Two properties are all it takes to capture Tesseract initialization, language pack loading, and processing details to a log file.
```cs
:title=Enable Full OCR Diagnostics in One Line
IronOcr.Installation.LogFilePath = "ocr.log"; IronOcr.Installation.LoggingMode = IronOcr.Installation.LoggingModes.All;
```
<div class="hsg-featured-snippet">
<h3>Minimal Workflow (5 steps)</h3>
<ol>
<li><a class="js-modal-open" data-modal-id="trial-license-after-download" href="https://nuget.org/packages/IronOcr/">Download a C# library for debugging OCR</a></li>
<li>Set <code>LogFilePath</code> to a writable file path</li>
<li>Set <code>LoggingMode</code> to <code>All</code> for full diagnostic capture</li>
<li>Run your OCR operation and reproduce the issue</li>
<li>Inspect the generated log file for engine warnings and processing details</li>
</ol>
</div>
<br class="clear" />
## How Do I Enable Diagnostic Logging?
The [`Installation`](https://ironsoftware.com/csharp/ocr/object-reference/api/IronOcr.Installation.html) class exposes three logging controls. Set these before calling any `Read` method.
```cs
using IronOcr;
// Write logs to a specific file
Installation.LogFilePath = "logs/ocr_diagnostics.log";
// Enable all logging channels: file + debug output
Installation.LoggingMode = Installation.LoggingModes.All;
// Or pipe logs into your existing ILogger pipeline
Installation.CustomLogger = myLoggerInstance;
```
`LoggingMode` accepts flag values from the [`LoggingModes`](https://ironsoftware.com/csharp/ocr/object-reference/api/IronOcr.Installation.LoggingModes.html) enum:
<div class="content__data-table" data-content-table>
<table>
<caption>Table 1: LoggingModes Options</caption>
<thead>
<tr><th>Mode</th><th>Output Target</th><th>Use Case</th></tr>
</thead>
<tbody>
<tr><td><code>None</code></td><td>Disabled</td><td>Production with external monitoring</td></tr>
<tr><td><code>DebugOutputWindow</code></td><td>IDE debug output window</td><td>Local development</td></tr>
<tr><td><code>File</code></td><td><code>LogFilePath</code></td><td>Server-side log collection</td></tr>
<tr><td><code>All</code></td><td>DebugOutputWindow + File</td><td>Full diagnostic capture</td></tr>
</tbody>
</table>
</div>
The `CustomLogger` property supports any `Microsoft.Extensions.Logging.ILogger` implementation, allowing you to direct OCR diagnostics to Serilog, NLog, or other structured logging sinks in your pipeline. Use [`ClearLogFiles`](https://ironsoftware.com/csharp/ocr/object-reference/api/IronOcr.Installation.html) to remove accumulated log data between runs.
With logging in place, the next step is understanding which exceptions IronOCR can throw and how to handle each one.
## What Exceptions Does IronOCR Throw?
IronOCR defines typed exceptions under the [`IronOcr.Exceptions`](https://ironsoftware.com/csharp/ocr/object-reference/api/IronOcr.Exceptions.html) namespace. Catching these specifically, rather than a blanket catch block, lets you route each failure type to the correct remediation path.
<div class="content__data-table" data-content-table>
<table>
<caption>Table 2: IronOCR Exception Reference</caption>
<thead>
<tr><th>Exception</th><th>Common Cause</th><th>Fix</th></tr>
</thead>
<tbody>
<tr><td><code>IronOcrInputException</code></td><td>Corrupt or unsupported image/PDF</td><td>Validate file before loading into <code>OcrInput</code></td></tr>
<tr><td><code>IronOcrProductException</code></td><td>Internal engine error during OCR execution</td><td>Enable logging, check log output, update to latest NuGet version</td></tr>
<tr><td><code>IronOcrDictionaryException</code></td><td>Missing or corrupt <code>.traineddata</code> language file</td><td>Reinstall the language pack NuGet or set <code>LanguagePackDirectory</code></td></tr>
<tr><td><code>IronOcrNativeException</code></td><td>Native C++ interop failure</td><td>Install <a href="https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist">Visual C++ Redistributable</a>; check AVX support</td></tr>
<tr><td><code>IronOcrLicensingException</code></td><td>Missing or expired license key</td><td>Set <code>LicenseKey</code> before calling <code>Read</code></td></tr>
<tr><td><code>LanguagePackException</code></td><td>Language pack not found at expected path</td><td>Verify <code>LanguagePackDirectory</code> or reinstall the NuGet language package</td></tr>
<tr><td><code>IronOcrAssemblyVersionMismatchException</code></td><td>Mismatched assembly versions after partial update</td><td>Clear NuGet cache, restore packages, ensure all IronOCR packages match</td></tr>
</tbody>
</table>
</div>
Use the following try-catch block to handle each exception type separately, applying exception filters for conditional logging.
### Input
A single-page vendor invoice from IronOCR Solutions to Acme Corporation, loaded via `LoadPdf` into `OcrInput`. It includes four line items, tax, and a grand total - enough text variety to give each exception handler a realistic exercise.
<iframe loading="lazy" src="/static-assets/ocr/how-to/debugging/invoice_scan.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">invoice_scan.pdf: Vendor invoice (#INV-2024-7829) used to demonstrate each typed exception handler in sequence.</p>
```cs
using IronOcr;
using IronOcr.Exceptions;
var ocr = new IronTesseract();
try
{
using var input = new OcrInput();
input.LoadPdf("invoice_scan.pdf");
OcrResult result = ocr.Read(input);
Console.WriteLine($"Text: {result.Text}");
Console.WriteLine($"Confidence: {result.Confidence:P1}");
}
catch (IronOcrInputException ex)
{
// File could not be loaded — corrupt, locked, or unsupported format
Console.Error.WriteLine($"Input error: {ex.Message}");
}
catch (IronOcrDictionaryException ex)
{
// Language pack missing — common in containerized deployments
Console.Error.WriteLine($"Language pack error: {ex.Message}");
}
catch (IronOcrNativeException ex) when (ex.Message.Contains("AVX"))
{
// CPU does not support AVX instructions
Console.Error.WriteLine($"Hardware incompatibility: {ex.Message}");
}
catch (IronOcrLicensingException)
{
Console.Error.WriteLine("License key is missing or invalid.");
}
catch (IronOcrProductException ex)
{
// Catch-all for other IronOCR engine errors
Console.Error.WriteLine($"OCR engine error: {ex.Message}");
Console.Error.WriteLine($"Stack trace: {ex.StackTrace}");
}
```
### Output
#### Success Output
The invoice loads cleanly and the engine returns a character count alongside a confidence score.
<div class="content-img-align-center">
<div class="center-image-wrapper">
<img src="/static-assets/ocr/how-to/debugging/exception-handling-success.png" alt="Terminal output showing successful OCR read of invoice_scan.pdf with character count and confidence score" class="img-responsive add-shadow" />
</div>
</div>
#### Failed Output
<div class="content-img-align-center">
<div class="center-image-wrapper">
<img src="/static-assets/ocr/how-to/debugging/exception-handling-failed.png" alt="Terminal output showing exception thrown when loading a missing PDF file" class="img-responsive add-shadow" />
</div>
</div>
Order catch blocks from most specific to most general. The `when` clause on `IronOcrNativeException` filters for [AVX-related failures](https://ironsoftware.com/csharp/ocr/troubleshooting/sehexception-avx-support/) without catching unrelated native errors. Each handler logs the exception message; the catch-all block also captures the stack trace for post-mortem analysis.
Catching the right exception tells you that something went wrong, but not how well the engine performed when it did succeed. For that, use confidence scores.
## How Do I Validate OCR Output with Confidence Scores?
Every [`OcrResult`](https://ironsoftware.com/csharp/ocr/object-reference/api/IronOcr.OcrResult.html) exposes a `Confidence` property, a value between 0 and 1 representing the engine's statistical certainty averaged across all recognized characters. You can access this at every level of the result hierarchy: [document](https://ironsoftware.com/csharp/ocr/object-reference/api/IronOcr.OcrResult.html), [page](https://ironsoftware.com/csharp/ocr/object-reference/api/IronOcr.OcrResult.Page.html), [paragraph](https://ironsoftware.com/csharp/ocr/object-reference/api/IronOcr.OcrResult.Paragraph.html), [word](https://ironsoftware.com/csharp/ocr/object-reference/api/IronOcr.OcrResult.Word.html), and [character](https://ironsoftware.com/csharp/ocr/object-reference/api/IronOcr.OcrResult.Character.html).
Use a threshold-gated pattern to prevent low-quality results from propagating downstream.
### Input
A thermal receipt with itemized line items, discounts, totals, and a barcode, loaded via `LoadImage`. Its narrow width, monospace font, and faint print make it a practical stress test for per-word confidence thresholds.
<div class="content-img-align-center">
<div class="center-image-wrapper" style="max-width: 320px; margin: 0 auto;">
<img src="/static-assets/ocr/how-to/debugging/receipt.png" alt="Sample thermal receipt from FoodMart showing itemized purchases, totals, and rewards points used as OCR input" class="img-responsive add-shadow" />
</div>
</div>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">receipt.png: Thermal receipt scan used to demonstrate threshold-gated confidence validation and per-word accuracy drill-down for the receipt image</p>
```cs
using IronOcr;
var ocr = new IronTesseract();
using var input = new OcrInput();
input.LoadImage("receipt.png");
OcrResult result = ocr.Read(input);
double confidence = result.Confidence;
Console.WriteLine($"Overall confidence: {confidence:P1}");
// Threshold-gated decision
if (confidence >= 0.90)
{
Console.WriteLine("ACCEPT — high confidence, processing result.");
ProcessResult(result.Text);
}
else if (confidence >= 0.70)
{
Console.WriteLine("FLAG — moderate confidence, queuing for review.");
QueueForReview(result.Text, confidence);
}
else
{
Console.WriteLine("REJECT — low confidence, logging for investigation.");
LogRejection("receipt.png", confidence);
}
// Drill into per-page and per-word confidence for diagnostics
foreach (var page in result.Pages)
{
Console.WriteLine($" Page {page.PageNumber}: {page.Confidence:P1}");
var lowConfidenceWords = page.Words
.Where(w => w.Confidence < 0.70)
.ToList();
foreach (var word in lowConfidenceWords)
{
Console.WriteLine($" Low-confidence word: \"{word.Text}\" ({word.Confidence:P1})");
}
}
```
### Output
<div class="content-img-align-center">
<div class="center-image-wrapper">
<img src="/static-assets/ocr/how-to/debugging/confidence-scoring-output.png" alt="Terminal output showing confidence score, accept/flag/reject decision, and per-word low-confidence drill-down for the receipt image" class="img-responsive add-shadow" />
</div>
</div>
This pattern is essential in pipelines where OCR feeds into data entry, invoice processing, or compliance workflows. The per-word drill-down identifies exactly which regions of the source image caused degradation; you can then apply [image quality filters](https://ironsoftware.com/csharp/ocr/how-to/image-quality-correction/) or [orientation corrections](https://ironsoftware.com/csharp/ocr/how-to/image-orientation-correction/) and re-process. For a deeper look at confidence scoring, see the [confidence levels how-to](https://ironsoftware.com/csharp/ocr/how-to/tesseract-result-confidence/).
For long-running jobs, confidence alone is not enough. You also need to know whether the engine is still making progress, and that is where the `OcrProgress` event comes in.
## How Do I Monitor OCR Progress in Real Time?
For multi-page documents, the `OcrProgress` event on `IronTesseract` fires after each page completes. The `OcrProgressEventsArgs` object exposes progress percent, elapsed duration, total pages, and pages complete. The example uses this three-page quarterly report as input: a structured business document spanning an executive summary, revenue breakdown, and operational metrics.
### Input
A three-page Q1 2024 financial report loaded via `LoadPdf`. Page one covers the executive summary with KPI metrics, page two contains revenue tables by product line and region, and page three covers operational processing volumes - each page type produces distinct per-page timing you can observe in the progress callbacks.
<iframe loading="lazy" src="/static-assets/ocr/how-to/debugging/quarterly_report.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">quarterly_report.pdf: Three-page Q1 2024 financial report (executive summary, revenue breakdown, operational metrics) used to demonstrate real-time `OcrProgress` callbacks per page.</p>
```cs
using IronOcr;
var ocr = new IronTesseract();
ocr.OcrProgress += (sender, e) =>
{
Console.WriteLine(
$"[OCR] {e.ProgressPercent}% complete | " +
$"Page {e.PagesComplete}/{e.TotalPages} | " +
$"Elapsed: {e.Duration.TotalSeconds:F1}s"
);
};
using var input = new OcrInput();
input.LoadPdf("quarterly_report.pdf");
OcrResult result = ocr.Read(input);
Console.WriteLine($"Finished in {result.Pages.Count()} pages, confidence: {result.Confidence:P1}");
```
### Output
<div class="content-img-align-center">
<div class="center-image-wrapper">
<img src="/static-assets/ocr/how-to/debugging/progress-monitoring-output.png" alt="Terminal output showing OcrProgress event callbacks per page with percent complete and elapsed time for a three-page PDF" class="img-responsive add-shadow" />
</div>
</div>
Wire this event into your logging infrastructure to track OCR job duration and detect stalls. If the elapsed duration exceeds a threshold without the progress percent advancing, the pipeline can flag the job for investigation. This is particularly useful for [batch PDF processing](https://ironsoftware.com/csharp/ocr/how-to/input-pdfs/) where a single malformed page can stall the entire job.
Progress monitoring shows execution state, but a file-level failure can still stop the entire batch short if not isolated.
## How Do I Handle Errors in Batch OCR Pipelines?
In production, a single file failure should not halt the entire batch. Isolate errors per file, log failures with context, and produce a summary report at the end. The example processes a folder of scan documents containing an invoice, a purchase order, a service contract, plus one intentionally corrupted file to trigger the error path. A representative sample is shown below:
### Input
A folder of PDFs passed to `Directory.GetFiles` - an invoice, a purchase order, a service contract, and one intentionally corrupted file. The two representative samples below show the document variety the pipeline processes in a single run.
<div class="competitors-section__wrapper-even-1">
<div class="competitors__card" style="width: 48%;">
<iframe loading="lazy" src="/static-assets/ocr/how-to/debugging/batch-scan-01.pdf" width="100%" height="380px"></iframe>
<p class="competitors__download-link" style="color: #181818; font-style: italic;">batch-scan-01.pdf: Invoice for Bright Horizon Ltd. (INV-2024-001) - successful OCR pass.</p>
</div>
<div class="competitors__card" style="width: 48%;">
<iframe loading="lazy" src="/static-assets/ocr/how-to/debugging/batch-scan-02.pdf" width="100%" height="380px"></iframe>
<p class="competitors__download-link" style="color: #181818; font-style: italic;">batch-scan-02.pdf: Purchase order for TechSupply Inc. (PO-2024-042) - second document type in the same run.</p>
</div>
</div>
```cs
using IronOcr;
using IronOcr.Exceptions;
var ocr = new IronTesseract();
Installation.LogFilePath = "batch_debug.log";
Installation.LoggingMode = Installation.LoggingModes.File;
string[] files = Directory.GetFiles("scans/", "*.pdf");
int succeeded = 0, failed = 0;
double totalConfidence = 0;
var failures = new List<(string File, string Error)>();
foreach (string file in files)
{
try
{
using var input = new OcrInput();
input.LoadPdf(file);
OcrResult result = ocr.Read(input);
totalConfidence += result.Confidence;
succeeded++;
Console.WriteLine($"OK: {Path.GetFileName(file)} — {result.Confidence:P1}");
}
catch (IronOcrInputException ex)
{
failed++;
failures.Add((file, $"Input error: {ex.Message}"));
Console.Error.WriteLine($"FAIL: {Path.GetFileName(file)} — {ex.Message}");
}
catch (IronOcrProductException ex)
{
failed++;
failures.Add((file, $"Engine error: {ex.Message}"));
Console.Error.WriteLine($"FAIL: {Path.GetFileName(file)} — {ex.Message}");
}
catch (Exception ex)
{
failed++;
failures.Add((file, $"Unexpected: {ex.Message}"));
Console.Error.WriteLine($"FAIL: {Path.GetFileName(file)} — {ex.GetType().Name}: {ex.Message}");
}
}
// Summary report
Console.WriteLine($"\n--- Batch Summary ---");
Console.WriteLine($"Total: {files.Length} | Passed: {succeeded} | Failed: {failed}");
if (succeeded > 0)
Console.WriteLine($"Average confidence: {totalConfidence / succeeded:P1}");
foreach (var (f, err) in failures)
Console.WriteLine($" {Path.GetFileName(f)}: {err}");
```
### Output
<div class="content-img-align-center">
<div class="center-image-wrapper">
<img src="/static-assets/ocr/how-to/debugging/batch-pipeline-output.png" alt="Terminal output showing batch pipeline results with per-file character counts, confidence scores, one error from a corrupted PDF, and a summary line" class="img-responsive add-shadow" />
</div>
</div>
The outer catch block handles unforeseen errors including network timeouts on shared storage, permission issues, or out-of-memory conditions on large TIFFs. Each failure records the file path and error message for the summary, while the loop continues processing remaining files. The log file at `batch_debug.log` captures engine-level detail for any file that triggers internal diagnostics.
For non-blocking execution in services or web applications, IronOCR supports [`ReadAsync`](https://ironsoftware.com/csharp/ocr/how-to/async/), which uses the same try-catch structure.
If the pipeline runs without errors but the extracted text is still wrong, the root cause is almost always image quality rather than code. Here is how to address that.
## How Do I Debug OCR Accuracy?
If confidence scores are consistently low, the issue is the source image rather than the OCR engine. IronOCR provides preprocessing tools to address this:
- Apply [image quality filters](https://ironsoftware.com/csharp/ocr/how-to/image-quality-correction/) such as sharpen, denoise, dilate, and erode to improve text clarity
- Use [orientation correction](https://ironsoftware.com/csharp/ocr/how-to/image-orientation-correction/) to automatically deskew and rotate scanned documents
- [Adjust the DPI setting](https://ironsoftware.com/csharp/ocr/how-to/dpi-setting/) for low-resolution images before processing
- Use [computer vision](https://ironsoftware.com/csharp/ocr/how-to/computer-vision/) to detect and isolate text regions in complex layouts
- The [IronOCR Utility](https://ironsoftware.com/csharp/ocr/troubleshooting/ironocr-utility/) lets you visually test filter combinations and export the optimal C# configuration
For deployment-specific issues, IronOCR maintains dedicated troubleshooting guides for [Azure Functions](https://ironsoftware.com/csharp/ocr/troubleshooting/azure-functions-deployment/), [Docker and Linux](https://ironsoftware.com/csharp/ocr/troubleshooting/libgdiplus/), and [general environment setup](https://ironsoftware.com/csharp/ocr/troubleshooting/general-troubleshooting-ocr/).
## Where Should I Go Next?
Now that you understand how to debug IronOCR at runtime, explore:
- Navigating [OCR result structure and metadata](https://ironsoftware.com/csharp/ocr/how-to/read-results/) including pages, blocks, paragraphs, words, and coordinates
- Understanding [confidence scoring](https://ironsoftware.com/csharp/ocr/how-to/tesseract-result-confidence/) at every level of the result hierarchy
- Using [async and multithreading](https://ironsoftware.com/csharp/ocr/how-to/async/) with `ReadAsync` for high-throughput pipelines
- Browsing the [full API reference](https://ironsoftware.com/csharp/ocr/object-reference/api/) for the full property list
For production use, remember to [obtain a license](https://ironsoftware.com/csharp/ocr/licensing/) to remove watermarks and access full functionality.
IronOCR enables you to detect OCR failures at the source, assess recognition quality at the word and character level, and monitor long-running jobs in real time. Built-in tools such as diagnostic file logging, a typed exception hierarchy, per-result confidence scoring, and the OcrProgress event support these workflows in production pipelines.
This guide walks through working examples for each: enabling diagnostic logging, handling typed exceptions, validating output with confidence scores, monitoring job progress in real time, and isolating errors in batch pipelines.
Quickstart: Enable full OCR diagnostic logging
Set LogFilePath and LoggingMode on the Installation class before the first Read call. Two properties are all it takes to capture Tesseract initialization, language pack loading, and processing details to a log file.
Set LoggingMode to All for full diagnostic capture
Run your OCR operation and reproduce the issue
Inspect the generated log file for engine warnings and processing details
How Do I Enable Diagnostic Logging?
The Installation class exposes three logging controls. Set these before calling any Read method.
using IronOcr;// Write logs to a specific fileInstallation.LogFilePath = "logs/ocr_diagnostics.log";// Enable all logging channels: file + debug outputInstallation.LoggingMode = Installation.LoggingModes.All;// Or pipe logs into your existing ILogger pipelineInstallation.CustomLogger = myLoggerInstance;
using IronOcr;
// Write logs to a specific file
Installation.LogFilePath = "logs/ocr_diagnostics.log";
// Enable all logging channels: file + debug output
Installation.LoggingMode = Installation.LoggingModes.All;
// Or pipe logs into your existing ILogger pipeline
Installation.CustomLogger = myLoggerInstance;
ImportsIronOcr' Write logs to a specific fileInstallation.LogFilePath = "logs/ocr_diagnostics.log"' Enable all logging channels: file + debug outputInstallation.LoggingMode = Installation.LoggingModes.All' Or pipe logs into your existing ILogger pipelineInstallation.CustomLogger = myLoggerInstance
Imports IronOcr
' Write logs to a specific file
Installation.LogFilePath = "logs/ocr_diagnostics.log"
' Enable all logging channels: file + debug output
Installation.LoggingMode = Installation.LoggingModes.All
' Or pipe logs into your existing ILogger pipeline
Installation.CustomLogger = myLoggerInstance
LoggingMode accepts flag values from the LoggingModes enum:
Table 1: LoggingModes Options
Mode
Output Target
Use Case
None
Disabled
Production with external monitoring
DebugOutputWindow
IDE debug output window
Local development
File
LogFilePath
Server-side log collection
All
DebugOutputWindow + File
Full diagnostic capture
The CustomLogger property supports any Microsoft.Extensions.Logging.ILogger implementation, allowing you to direct OCR diagnostics to Serilog, NLog, or other structured logging sinks in your pipeline. Use ClearLogFiles to remove accumulated log data between runs.
With logging in place, the next step is understanding which exceptions IronOCR can throw and how to handle each one.
What Exceptions Does IronOCR Throw?
IronOCR defines typed exceptions under the IronOcr.Exceptions namespace. Catching these specifically, rather than a blanket catch block, lets you route each failure type to the correct remediation path.
Table 2: IronOCR Exception Reference
Exception
Common Cause
Fix
IronOcrInputException
Corrupt or unsupported image/PDF
Validate file before loading into OcrInput
IronOcrProductException
Internal engine error during OCR execution
Enable logging, check log output, update to latest NuGet version
IronOcrDictionaryException
Missing or corrupt .traineddata language file
Reinstall the language pack NuGet or set LanguagePackDirectory
Verify LanguagePackDirectory or reinstall the NuGet language package
IronOcrAssemblyVersionMismatchException
Mismatched assembly versions after partial update
Clear NuGet cache, restore packages, ensure all IronOCR packages match
Use the following try-catch block to handle each exception type separately, applying exception filters for conditional logging.
Input
A single-page vendor invoice from IronOCR Solutions to Acme Corporation, loaded via LoadPdf into OcrInput. It includes four line items, tax, and a grand total - enough text variety to give each exception handler a realistic exercise.
invoice_scan.pdf: Vendor invoice (#INV-2024-7829) used to demonstrate each typed exception handler in sequence.
using IronOcr;using IronOcr.Exceptions;var ocr = new IronTesseract();try{ using var input = new OcrInput(); input.LoadPdf("invoice_scan.pdf"); OcrResult result = ocr.Read(input);Console.WriteLine($"Text: {result.Text}");Console.WriteLine($"Confidence: {result.Confidence:P1}");}catch (IronOcrInputException ex){ // File could not be loaded — corrupt, locked, or unsupported formatConsole.Error.WriteLine($"Input error: {ex.Message}");}catch (IronOcrDictionaryException ex){ // Language pack missing — common in containerized deploymentsConsole.Error.WriteLine($"Language pack error: {ex.Message}");}catch (IronOcrNativeException ex) when (ex.Message.Contains("AVX")){ // CPU does not support AVX instructionsConsole.Error.WriteLine($"Hardware incompatibility: {ex.Message}");}catch (IronOcrLicensingException){Console.Error.WriteLine("License key is missing or invalid.");}catch (IronOcrProductException ex){ // Catch-all for other IronOCR engine errorsConsole.Error.WriteLine($"OCR engine error: {ex.Message}");Console.Error.WriteLine($"Stack trace: {ex.StackTrace}");}
using IronOcr;
using IronOcr.Exceptions;
var ocr = new IronTesseract();
try
{
using var input = new OcrInput();
input.LoadPdf("invoice_scan.pdf");
OcrResult result = ocr.Read(input);
Console.WriteLine($"Text: {result.Text}");
Console.WriteLine($"Confidence: {result.Confidence:P1}");
}
catch (IronOcrInputException ex)
{
// File could not be loaded — corrupt, locked, or unsupported format
Console.Error.WriteLine($"Input error: {ex.Message}");
}
catch (IronOcrDictionaryException ex)
{
// Language pack missing — common in containerized deployments
Console.Error.WriteLine($"Language pack error: {ex.Message}");
}
catch (IronOcrNativeException ex) when (ex.Message.Contains("AVX"))
{
// CPU does not support AVX instructions
Console.Error.WriteLine($"Hardware incompatibility: {ex.Message}");
}
catch (IronOcrLicensingException)
{
Console.Error.WriteLine("License key is missing or invalid.");
}
catch (IronOcrProductException ex)
{
// Catch-all for other IronOCR engine errors
Console.Error.WriteLine($"OCR engine error: {ex.Message}");
Console.Error.WriteLine($"Stack trace: {ex.StackTrace}");
}
ImportsIronOcrImportsIronOcr.ExceptionsDim ocr = New IronTesseract()TryUsing input = New OcrInput() input.LoadPdf("invoice_scan.pdf") Dim result AsOcrResult = ocr.Read(input)Console.WriteLine($"Text: {result.Text}")Console.WriteLine($"Confidence: {result.Confidence:P1}")EndUsingCatch ex AsIronOcrInputException ' File could not be loaded — corrupt, locked, or unsupported formatConsole.Error.WriteLine($"Input error: {ex.Message}")Catch ex AsIronOcrDictionaryException ' Language pack missing — common in containerized deploymentsConsole.Error.WriteLine($"Language pack error: {ex.Message}")Catch ex AsIronOcrNativeExceptionWhen ex.Message.Contains("AVX") ' CPU does not support AVX instructionsConsole.Error.WriteLine($"Hardware incompatibility: {ex.Message}")Catch ex AsIronOcrLicensingExceptionConsole.Error.WriteLine("License key is missing or invalid.")Catch ex AsIronOcrProductException ' Catch-all for other IronOCR engine errorsConsole.Error.WriteLine($"OCR engine error: {ex.Message}")Console.Error.WriteLine($"Stack trace: {ex.StackTrace}")EndTry
Imports IronOcr
Imports IronOcr.Exceptions
Dim ocr = New IronTesseract()
Try
Using input = New OcrInput()
input.LoadPdf("invoice_scan.pdf")
Dim result As OcrResult = ocr.Read(input)
Console.WriteLine($"Text: {result.Text}")
Console.WriteLine($"Confidence: {result.Confidence:P1}")
End Using
Catch ex As IronOcrInputException
' File could not be loaded — corrupt, locked, or unsupported format
Console.Error.WriteLine($"Input error: {ex.Message}")
Catch ex As IronOcrDictionaryException
' Language pack missing — common in containerized deployments
Console.Error.WriteLine($"Language pack error: {ex.Message}")
Catch ex As IronOcrNativeException When ex.Message.Contains("AVX")
' CPU does not support AVX instructions
Console.Error.WriteLine($"Hardware incompatibility: {ex.Message}")
Catch ex As IronOcrLicensingException
Console.Error.WriteLine("License key is missing or invalid.")
Catch ex As IronOcrProductException
' Catch-all for other IronOCR engine errors
Console.Error.WriteLine($"OCR engine error: {ex.Message}")
Console.Error.WriteLine($"Stack trace: {ex.StackTrace}")
End Try
Output
Success Output
The invoice loads cleanly and the engine returns a character count alongside a confidence score.
Failed Output
Order catch blocks from most specific to most general. The when clause on IronOcrNativeException filters for AVX-related failures without catching unrelated native errors. Each handler logs the exception message; the catch-all block also captures the stack trace for post-mortem analysis.
Catching the right exception tells you that something went wrong, but not how well the engine performed when it did succeed. For that, use confidence scores.
How Do I Validate OCR Output with Confidence Scores?
Every OcrResult exposes a Confidence property, a value between 0 and 1 representing the engine's statistical certainty averaged across all recognized characters. You can access this at every level of the result hierarchy: document, page, paragraph, word, and character.
Use a threshold-gated pattern to prevent low-quality results from propagating downstream.
Input
A thermal receipt with itemized line items, discounts, totals, and a barcode, loaded via LoadImage. Its narrow width, monospace font, and faint print make it a practical stress test for per-word confidence thresholds.
receipt.png: Thermal receipt scan used to demonstrate threshold-gated confidence validation and per-word accuracy drill-down for the receipt image
using IronOcr;var ocr = new IronTesseract();using var input = new OcrInput();input.LoadImage("receipt.png");OcrResult result = ocr.Read(input);double confidence = result.Confidence;Console.WriteLine($"Overall confidence: {confidence:P1}");// Threshold-gated decisionif (confidence >= 0.90){Console.WriteLine("ACCEPT — high confidence, processing result.");ProcessResult(result.Text);}else if (confidence >= 0.70){Console.WriteLine("FLAG — moderate confidence, queuing for review.");QueueForReview(result.Text, confidence);}else{Console.WriteLine("REJECT — low confidence, logging for investigation.");LogRejection("receipt.png", confidence);}// Drill into per-page and per-word confidence for diagnosticsforeach (var page in result.Pages){Console.WriteLine($" Page {page.PageNumber}: {page.Confidence:P1}"); var lowConfidenceWords = page.Words .Where(w => w.Confidence < 0.70) .ToList(); foreach (var word in lowConfidenceWords) {Console.WriteLine($" Low-confidence word: \"{word.Text}\" ({word.Confidence:P1})"); }}
using IronOcr;
var ocr = new IronTesseract();
using var input = new OcrInput();
input.LoadImage("receipt.png");
OcrResult result = ocr.Read(input);
double confidence = result.Confidence;
Console.WriteLine($"Overall confidence: {confidence:P1}");
// Threshold-gated decision
if (confidence >= 0.90)
{
Console.WriteLine("ACCEPT — high confidence, processing result.");
ProcessResult(result.Text);
}
else if (confidence >= 0.70)
{
Console.WriteLine("FLAG — moderate confidence, queuing for review.");
QueueForReview(result.Text, confidence);
}
else
{
Console.WriteLine("REJECT — low confidence, logging for investigation.");
LogRejection("receipt.png", confidence);
}
// Drill into per-page and per-word confidence for diagnostics
foreach (var page in result.Pages)
{
Console.WriteLine($" Page {page.PageNumber}: {page.Confidence:P1}");
var lowConfidenceWords = page.Words
.Where(w => w.Confidence < 0.70)
.ToList();
foreach (var word in lowConfidenceWords)
{
Console.WriteLine($" Low-confidence word: \"{word.Text}\" ({word.Confidence:P1})");
}
}
ImportsIronOcrDim ocr As New IronTesseract()Using input As New OcrInput() input.LoadImage("receipt.png") Dim result AsOcrResult = ocr.Read(input) Dim confidence AsDouble = result.ConfidenceConsole.WriteLine($"Overall confidence: {confidence:P1}") ' Threshold-gated decision If confidence >= 0.9 ThenConsole.WriteLine("ACCEPT — high confidence, processing result.")ProcessResult(result.Text) ElseIf confidence >= 0.7 ThenConsole.WriteLine("FLAG — moderate confidence, queuing for review.")QueueForReview(result.Text, confidence) ElseConsole.WriteLine("REJECT — low confidence, logging for investigation.")LogRejection("receipt.png", confidence) End If ' Drill into per-page and per-word confidence for diagnostics For Each page In result.PagesConsole.WriteLine($" Page {page.PageNumber}: {page.Confidence:P1}") Dim lowConfidenceWords = page.Words _ .Where(Function(w) w.Confidence < 0.7) _ .ToList() For Each word In lowConfidenceWordsConsole.WriteLine($" Low-confidence word: ""{word.Text}"" ({word.Confidence:P1})") Next NextEndUsing
Imports IronOcr
Dim ocr As New IronTesseract()
Using input As New OcrInput()
input.LoadImage("receipt.png")
Dim result As OcrResult = ocr.Read(input)
Dim confidence As Double = result.Confidence
Console.WriteLine($"Overall confidence: {confidence:P1}")
' Threshold-gated decision
If confidence >= 0.9 Then
Console.WriteLine("ACCEPT — high confidence, processing result.")
ProcessResult(result.Text)
ElseIf confidence >= 0.7 Then
Console.WriteLine("FLAG — moderate confidence, queuing for review.")
QueueForReview(result.Text, confidence)
Else
Console.WriteLine("REJECT — low confidence, logging for investigation.")
LogRejection("receipt.png", confidence)
End If
' Drill into per-page and per-word confidence for diagnostics
For Each page In result.Pages
Console.WriteLine($" Page {page.PageNumber}: {page.Confidence:P1}")
Dim lowConfidenceWords = page.Words _
.Where(Function(w) w.Confidence < 0.7) _
.ToList()
For Each word In lowConfidenceWords
Console.WriteLine($" Low-confidence word: ""{word.Text}"" ({word.Confidence:P1})")
Next
Next
End Using
Output
This pattern is essential in pipelines where OCR feeds into data entry, invoice processing, or compliance workflows. The per-word drill-down identifies exactly which regions of the source image caused degradation; you can then apply image quality filters or orientation corrections and re-process. For a deeper look at confidence scoring, see the confidence levels how-to.
For long-running jobs, confidence alone is not enough. You also need to know whether the engine is still making progress, and that is where the OcrProgress event comes in.
How Do I Monitor OCR Progress in Real Time?
For multi-page documents, the OcrProgress event on IronTesseract fires after each page completes. The OcrProgressEventsArgs object exposes progress percent, elapsed duration, total pages, and pages complete. The example uses this three-page quarterly report as input: a structured business document spanning an executive summary, revenue breakdown, and operational metrics.
Input
A three-page Q1 2024 financial report loaded via LoadPdf. Page one covers the executive summary with KPI metrics, page two contains revenue tables by product line and region, and page three covers operational processing volumes - each page type produces distinct per-page timing you can observe in the progress callbacks.
quarterly_report.pdf: Three-page Q1 2024 financial report (executive summary, revenue breakdown, operational metrics) used to demonstrate real-time OcrProgress callbacks per page.
using IronOcr;var ocr = new IronTesseract();ocr.OcrProgress += (sender, e) =>{Console.WriteLine( $"[OCR] {e.ProgressPercent}% complete | " + $"Page {e.PagesComplete}/{e.TotalPages} | " + $"Elapsed: {e.Duration.TotalSeconds:F1}s" );};using var input = new OcrInput();input.LoadPdf("quarterly_report.pdf");OcrResult result = ocr.Read(input);Console.WriteLine($"Finished in {result.Pages.Count()} pages, confidence: {result.Confidence:P1}");
using IronOcr;
var ocr = new IronTesseract();
ocr.OcrProgress += (sender, e) =>
{
Console.WriteLine(
$"[OCR] {e.ProgressPercent}% complete | " +
$"Page {e.PagesComplete}/{e.TotalPages} | " +
$"Elapsed: {e.Duration.TotalSeconds:F1}s"
);
};
using var input = new OcrInput();
input.LoadPdf("quarterly_report.pdf");
OcrResult result = ocr.Read(input);
Console.WriteLine($"Finished in {result.Pages.Count()} pages, confidence: {result.Confidence:P1}");
ImportsIronOcrDim ocr As New IronTesseract()AddHandler ocr.OcrProgress, Sub(sender, e)Console.WriteLine($"[OCR] {e.ProgressPercent}% complete | " & $"Page {e.PagesComplete}/{e.TotalPages} | " & $"Elapsed: {e.Duration.TotalSeconds:F1}s")End SubUsing input As New OcrInput() input.LoadPdf("quarterly_report.pdf") Dim result AsOcrResult = ocr.Read(input)Console.WriteLine($"Finished in {result.Pages.Count()} pages, confidence: {result.Confidence:P1}")EndUsing
Imports IronOcr
Dim ocr As New IronTesseract()
AddHandler ocr.OcrProgress, Sub(sender, e)
Console.WriteLine($"[OCR] {e.ProgressPercent}% complete | " &
$"Page {e.PagesComplete}/{e.TotalPages} | " &
$"Elapsed: {e.Duration.TotalSeconds:F1}s")
End Sub
Using input As New OcrInput()
input.LoadPdf("quarterly_report.pdf")
Dim result As OcrResult = ocr.Read(input)
Console.WriteLine($"Finished in {result.Pages.Count()} pages, confidence: {result.Confidence:P1}")
End Using
Output
Wire this event into your logging infrastructure to track OCR job duration and detect stalls. If the elapsed duration exceeds a threshold without the progress percent advancing, the pipeline can flag the job for investigation. This is particularly useful for batch PDF processing where a single malformed page can stall the entire job.
Progress monitoring shows execution state, but a file-level failure can still stop the entire batch short if not isolated.
How Do I Handle Errors in Batch OCR Pipelines?
In production, a single file failure should not halt the entire batch. Isolate errors per file, log failures with context, and produce a summary report at the end. The example processes a folder of scan documents containing an invoice, a purchase order, a service contract, plus one intentionally corrupted file to trigger the error path. A representative sample is shown below:
Input
A folder of PDFs passed to Directory.GetFiles - an invoice, a purchase order, a service contract, and one intentionally corrupted file. The two representative samples below show the document variety the pipeline processes in a single run.
ImportsIronOcrImportsIronOcr.ExceptionsDim ocr As New IronTesseract()Installation.LogFilePath = "batch_debug.log"Installation.LoggingMode = Installation.LoggingModes.FileDim files AsString() = Directory.GetFiles("scans/", "*.pdf")Dim succeeded AsInteger = 0, failed AsInteger = 0Dim totalConfidence AsDouble = 0Dim failures As New List(Of (FileAsString, ErrorAsString))()For Each file AsStringIn filesTryUsing input As New OcrInput() input.LoadPdf(file) Dim result AsOcrResult = ocr.Read(input) totalConfidence += result.Confidence succeeded += 1Console.WriteLine($"OK: {Path.GetFileName(file)} — {result.Confidence:P1}")EndUsingCatch ex AsIronOcrInputException failed += 1 failures.Add((file, $"Input error: {ex.Message}"))Console.Error.WriteLine($"FAIL: {Path.GetFileName(file)} — {ex.Message}")Catch ex AsIronOcrProductException failed += 1 failures.Add((file, $"Engine error: {ex.Message}"))Console.Error.WriteLine($"FAIL: {Path.GetFileName(file)} — {ex.Message}")Catch ex AsException failed += 1 failures.Add((file, $"Unexpected: {ex.Message}"))Console.Error.WriteLine($"FAIL: {Path.GetFileName(file)} — {ex.GetType().Name}: {ex.Message}")EndTryNext' Summary reportConsole.WriteLine(vbCrLf & "--- Batch Summary ---")Console.WriteLine($"Total: {files.Length} | Passed: {succeeded} | Failed: {failed}")If succeeded > 0 ThenConsole.WriteLine($"Average confidence: {totalConfidence / succeeded:P1}")End IfFor Each failure In failuresConsole.WriteLine($" {Path.GetFileName(failure.File)}: {failure.Error}")Next
Imports IronOcr
Imports IronOcr.Exceptions
Dim ocr As New IronTesseract()
Installation.LogFilePath = "batch_debug.log"
Installation.LoggingMode = Installation.LoggingModes.File
Dim files As String() = Directory.GetFiles("scans/", "*.pdf")
Dim succeeded As Integer = 0, failed As Integer = 0
Dim totalConfidence As Double = 0
Dim failures As New List(Of (File As String, Error As String))()
For Each file As String In files
Try
Using input As New OcrInput()
input.LoadPdf(file)
Dim result As OcrResult = ocr.Read(input)
totalConfidence += result.Confidence
succeeded += 1
Console.WriteLine($"OK: {Path.GetFileName(file)} — {result.Confidence:P1}")
End Using
Catch ex As IronOcrInputException
failed += 1
failures.Add((file, $"Input error: {ex.Message}"))
Console.Error.WriteLine($"FAIL: {Path.GetFileName(file)} — {ex.Message}")
Catch ex As IronOcrProductException
failed += 1
failures.Add((file, $"Engine error: {ex.Message}"))
Console.Error.WriteLine($"FAIL: {Path.GetFileName(file)} — {ex.Message}")
Catch ex As Exception
failed += 1
failures.Add((file, $"Unexpected: {ex.Message}"))
Console.Error.WriteLine($"FAIL: {Path.GetFileName(file)} — {ex.GetType().Name}: {ex.Message}")
End Try
Next
' Summary report
Console.WriteLine(vbCrLf & "--- Batch Summary ---")
Console.WriteLine($"Total: {files.Length} | Passed: {succeeded} | Failed: {failed}")
If succeeded > 0 Then
Console.WriteLine($"Average confidence: {totalConfidence / succeeded:P1}")
End If
For Each failure In failures
Console.WriteLine($" {Path.GetFileName(failure.File)}: {failure.Error}")
Next
Output
The outer catch block handles unforeseen errors including network timeouts on shared storage, permission issues, or out-of-memory conditions on large TIFFs. Each failure records the file path and error message for the summary, while the loop continues processing remaining files. The log file at batch_debug.log captures engine-level detail for any file that triggers internal diagnostics.
For non-blocking execution in services or web applications, IronOCR supports ReadAsync, which uses the same try-catch structure.
If the pipeline runs without errors but the extracted text is still wrong, the root cause is almost always image quality rather than code. Here is how to address that.
How Do I Debug OCR Accuracy?
If confidence scores are consistently low, the issue is the source image rather than the OCR engine. IronOCR provides preprocessing tools to address this:
Apply image quality filters such as sharpen, denoise, dilate, and erode to improve text clarity
For production use, remember to obtain a license to remove watermarks and access full functionality.
Frequently Asked Questions
How can I enable diagnostic logging in IronOCR?
To enable diagnostic logging in IronOCR, set the `LogFilePath` and `LoggingMode` properties on the `Installation` class before the first `Read` call. This setup captures details like Tesseract initialization and language pack loading.
What are some common exceptions thrown by IronOCR?
Common exceptions include `IronOcrInputException` for corrupt images, `IronOcrProductException` for internal engine errors, `IronOcrDictionaryException` for missing language files, and `IronOcrLicensingException` for missing license keys.
How does IronOCR handle confidence scoring?
IronOCR provides a `Confidence` property representing statistical certainty averaged across all recognized characters. This can be used to validate OCR output, allowing you to filter out low-confidence results before further processing.
Can IronOCR monitor OCR progress in real time?
Yes, IronOCR features the `OcrProgress` event, which tracks the progress of processing multi-page documents by providing updates on the number of pages completed and the overall percentage progress.
How can errors be isolated in a batch OCR pipeline using IronOCR?
In batch OCR pipelines, errors are isolated by handling exceptions per file, logging the context of the failures, and producing a summary report. This ensures that a single failure does not halt the entire batch.
What preprocessing tools does IronOCR offer for improving OCR accuracy?
IronOCR offers preprocessing tools like image quality filters (sharpen, denoise), orientation correction to deskew scanned documents, and DPI adjustments for low-resolution images to enhance OCR accuracy.
How can developers handle exceptions using IronOCR?
Developers can handle exceptions using IronOCR by implementing specific catch blocks for different exception types like `IronOcrInputException` and `IronOcrProductException`, directing each to a suitable remediation path.
Is it possible to execute OCR operations asynchronously with IronOCR?
Yes, IronOCR supports asynchronous operations through `ReadAsync`, allowing non-blocking execution in services or web applications, and uses the same try-catch structure as synchronous operations.
What logging modes are available in IronOCR?
IronOCR supports various logging modes such as `None` for disabled logging, `DebugOutputWindow` for debug output in IDEs, `File` for server-side collection, and `All` for full diagnostic capture.
What should be done if OCR confidence scores are low?
If OCR confidence scores are low, the issue might be with the source image. Use IronOCR's preprocessing tools like image quality filters and orientation corrections to improve the source image clarity.
Curtis Chau holds a Bachelor’s degree in Computer Science (Carleton University) and specializes in front-end development with expertise in Node.js, TypeScript, JavaScript, and React. Passionate about crafting intuitive and aesthetically pleasing user interfaces, Curtis enjoys working with modern frameworks and creating well-structured, visually appealing manuals.