# 在C#中使用IronOCR保存可搜尋的PDF
IronOCR使C#開發者能將掃描的檔案和圖像轉換成可搜尋的PDF,使用[OCR技術](https://ironsoftware.com/csharp/ocr/features/),支持以檔案、位元組或流的形式輸出,僅需幾行程式碼。
可搜尋的PDF通常被稱為OCR(光學字元識別)PDF,是一種PDF文件型別,包含掃描圖像和機器可讀文字。 這些PDF是通過對掃描的紙質檔案或圖像進行OCR處理,識別圖像中的文字,並轉換成可選擇和搜尋的文字來建立的。
`ReadDocumentAdvanced`的結果,支持從照片和高級文件OCR工作流建立可搜尋的PDF。 此功能特別有用於將紙質檔案數位化或讓舊版PDF可搜尋,以便更好的文件管理。
*as-heading:2(快速入門:一行程式碼導出可搜尋的PDF)*
設置`SaveAsSearchablePdf(...)`。 這就是IronOCR生成完整可搜尋PDF所需的一切。
```cs
:title=Quickly Make a PDF Searchable with IronOCR
new IronOcr.IronTesseract { Configuration = { RenderSearchablePdf = true } } .Read(new IronOcr.OcrImageInput("file.jpg")).SaveAsSearchablePdf("searchable.pdf");
```
<div class="hsg-featured-snippet">
<h3>最小工作流程(5步)</h3>
<ol>
<li><a class="js-modal-open" data-modal-id="trial-license-after-download" href="https://nuget.org/packages/IronOcr/">下載一個C#程式庫以將結果保存為可搜尋的PDF</a></li>
<li>準備OCR的圖像和PDF文件</li>
<li>將<strong>RenderSearchablePdf</strong>屬性設置為<code>true</code></li>
<li>利用<code>SaveAsSearchablePdf</code>方法輸出一個可搜尋的PDF檔案</li>
<li>將可搜尋的PDF導出為位元組和流</li>
</ol>
</div>
<br class="clear" />
## 如何將OCR結果導出為可搜尋的PDF?
要使用IronOCR將結果導出為可搜尋的PDF,請將`SaveAsSearchablePdf`並提供輸出檔案路徑。
### 輸入
來自哈利波特小說的一頁,掃描為TIFF檔案,通過`OcrImageInput`載入。 該頁面包含密集的印刷文字,是測試可搜尋PDF文字層的真實輸入。
<div class="content-img-align-center">
<div class="center-image-wrapper" style="max-width: 380px; margin: 0 auto;">
<img src="/static-assets/ocr/how-to/searchable-pdf/potter.webp" alt="哈利波特書籍中的一頁顯示第八章"死忌派對",其中包含哈利遇見差點沒頭的奈的文字" class="img-responsive add-shadow" />
</div>
</div>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">potter.tiff:作為OCR輸入使用的掃描小說頁面,生成帶有不可見文字層的可搜尋PDF。</p>
```csharp
using IronOcr;
// Create the OCR engine: defaults to English with balanced speed and accuracy
IronTesseract ocrTesseract = new IronTesseract();
// Required: without this flag the text overlay layer is not built, and SaveAsSearchablePdf produces a plain image PDF
ocrTesseract.Configuration.RenderSearchablePdf = true;
// Wrap the TIFF in OcrImageInput: handles DPI detection and page layout automatically
using var imageInput = new OcrImageInput("Potter.tiff");
// Run OCR; returns a result containing the recognized text and spatial layout data
OcrResult ocrResult = ocrTesseract.Read(imageInput);
// Write the output: the original scanned image is preserved with an invisible text layer on top
ocrResult.SaveAsSearchablePdf("searchablePdf.pdf");
```
### 輸出
<iframe loading="lazy" src="/static-assets/ocr/how-to/searchable-pdf/searchablePdf.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">searchablePdf.pdf:可搜尋的PDF輸出。選擇或搜尋任何單詞以驗證OCR文字層。</p>
生成的PDF嵌入了原始掃描頁面圖像,且每個識別出的字上方都有不可見的文字層。 在檢視器中選擇或搜尋任何單詞,以確認文字層是否存在。
IronOCR對此使用特定字體,這可能會導致渲染文字大小與原始大小略有不同。
在處理[多頁TIFF檔案](https://ironsoftware.com/csharp/ocr/examples/csharp-tesseract-multipage-tiff/)或複雜文件時,IronOCR自動處理所有頁面並將其包含在輸出中。 該程式庫自動處理頁面排序和文字覆蓋位置,確保準確的文字到圖像映射。
### 如何從照片或高級文件掃描建立可搜尋的PDF?
使用`ReadDocumentAdvanced`時,也提供可搜尋的PDF導出功能。 這些方法中的每一種返回一種支持`SaveAsSearchablePdf`的結果型別。
調用這些方法時,您可以選擇性地傳遞`ModelType`。 預設為`Enhanced`在速度的代價下提供更高的準確性。
#### 輸入
一張牆壁壁畫的照片,包含塗鴉文字,通過`LoadImage`載入。 場景包含多個單詞嵌入在現實環境中,這使得它成為用`Enhanced`模型的實用測試。
<div class="content-img-align-center">
<div class="center-image-wrapper" style="max-width: 450px; margin: 0 auto;">
<img src="/static-assets/ocr/how-to/searchable-pdf/photo-input.webp" alt="用作ReadPhoto OCR輸入的包含文字的照片" class="img-responsive add-shadow" />
</div>
</div>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">photo.png:使用增強模型的ReadPhoto所載入的牆壁壁畫照片,用於生成可搜尋的PDF。</p>
```csharp
using IronOcr;
var ocr = new IronTesseract();
using var input = new OcrInput();
input.LoadImage("photo.png");
// ReadPhoto with Enhanced model
OcrPhotoResult photoResult = ocr.ReadPhoto(input, ModelType.Enhanced);
Console.WriteLine(photoResult.Text);
// Save as searchable PDF
byte[] pdfBytes = photoResult.SaveAsSearchablePdfBytes();
File.WriteAllBytes("searchable-photo.pdf", pdfBytes);
```
#### 輸出
<iframe loading="lazy" src="/static-assets/ocr/how-to/searchable-pdf/searchable-photo.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">searchable-photo.pdf:從ReadPhoto中導出的可搜尋PDF。文字層在任何PDF檢視器中支持全文搜尋。</p>
生成的可搜尋PDF包含了識別字上方的不可見文字層。 在PDF檢視器中搜尋"Milk"返回3個匹配項,直接從原始照片中的塗鴉文字提取。
相同的方法適用於`OcrDocAdvancedResult`:
#### 輸入
通過`LoadImage`載入的掃描發票。 它包含結構化字段(供應商名稱、明細項目和總數),`Enhanced`模型識別並作為可搜尋文字層嵌入。
<div class="content-img-align-center">
<div class="center-image-wrapper" style="max-width: 400px; margin: 0 auto;">
<img src="/static-assets/ocr/how-to/searchable-pdf/invoice-input.webp" alt="發票文件作為ReadDocumentAdvanced OCR的輸入" class="img-responsive add-shadow" />
</div>
</div>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">invoice.png:載入到OcrInput並傳入ReadDocumentAdvanced增強模型的掃描發票。</p>
```csharp
using IronOcr;
var ocr = new IronTesseract();
using var input = new OcrInput();
input.LoadImage("invoice.png");
// ReadDocumentAdvanced with Enhanced model
OcrDocAdvancedResult docResult = ocr.ReadDocumentAdvanced(input, ModelType.Enhanced);
byte[] docPdfBytes = docResult.SaveAsSearchablePdfBytes();
File.WriteAllBytes("searchable-doc.pdf", docPdfBytes);
```
#### 輸出
<iframe loading="lazy" src="/static-assets/ocr/how-to/searchable-pdf/searchable-doc.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">searchable-doc.pdf:從ReadDocumentAdvanced中導出的可搜尋PDF。發票欄位是可選擇和可搜尋的。</p>
`ExtensionAdvancedScanException`。
### 處理多頁文件
在進行多頁文件上的PDF OCR操作時,IronOCR按順序處理每一頁並保持原始文件結構。
#### 輸入
一份來自Hartwell Capital Management的11頁年度報告,通過`OcrPdfInput`載入。 使用`Read`調用中處理它們。
<iframe loading="lazy" src="/static-assets/ocr/how-to/searchable-pdf/multi-page-scan.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">multi-page-scan.pdf:用作多頁可搜尋PDF轉換輸入的11頁Hartwell Capital Management年度報告。</p>
```csharp
using IronOcr;
// Create the OCR engine. RenderSearchablePdf is false by default; no need to set it when using OcrPdfInput directly
var ocrTesseract = new IronTesseract();
// Load pages 1–10 (indices 0–9) only; PageIndices avoids loading and OCR-ing the full document unnecessarily
using var pdfInput = new OcrPdfInput("multi-page-scan.pdf", PageIndices: Enumerable.Range(0, 10));
// Run OCR across all selected pages in order
OcrResult result = ocrTesseract.Read(pdfInput);
// Write the searchable PDF; true = apply the input's image filters to the embedded page images in the output
result.SaveAsSearchablePdf("searchable-multi-page.pdf", true);
```
#### 輸出
<iframe loading="lazy" src="/static-assets/ocr/how-to/searchable-pdf/searchable-multi-page.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">searchable-multi-page.pdf:10頁可搜尋PDF輸出。每一頁都有一個不可見的文字層,支持全文搜尋。</p>
生成的PDF包含10頁(原始報告的1-10頁),每頁上方都有文字層,使提取的內容在任何PDF檢視器中可選擇和可搜尋。
### 如何在建立可搜尋的PDF時應用濾鏡?
`SaveAsSearchablePdf`的第二個參數接受布林值來控制是否對嵌入的輸出應用圖像濾鏡。 使用[圖像優化濾鏡](https://ironsoftware.com/csharp/ocr/examples/ocr-image-filters-for-net-tesseract/)可以顯著提高OCR準確性,特別是在處理[低質量掃描](https://ironsoftware.com/csharp/ocr/examples/ocr-low-quality-scans-tesseract/)時。
下面的例子應用了灰度濾鏡,並將`true`作為第二個參數來嵌入經過濾鏡處理的圖像進入可搜尋PDF輸出中。
```cs
using IronOcr;
// Create OCR engine: filters are applied at the OcrInput level, so no configuration changes are needed here
var ocr = new IronTesseract();
var ocrInput = new OcrInput();
// Load the scanned PDF as the OCR source
ocrInput.LoadPdf("invoice.pdf");
// Convert to grayscale: removes color noise that can reduce OCR accuracy on color-printed documents
ocrInput.ToGrayScale();
// Run OCR on the preprocessed input
OcrResult result = ocr.Read(ocrInput);
// Write the searchable PDF; true = embed the grayscale-filtered image rather than the original color scan
result.SaveAsSearchablePdf("outputGrayscale.pdf", true);
```
為獲得最佳效果,考慮使用[濾鏡精靈](https://ironsoftware.com/csharp/ocr/examples/filter-wizard/)自動確定適合您具體文件型別的最佳濾鏡組合。 該工具會分析您的輸入並提出適當的預處理步驟。
### 我如何修正可搜尋PDF中的錯誤字元?
如果文字在PDF中視覺上看起來正確但在搜尋或複製時顯示為損壞字元,問題出在使用於可搜尋文字層中的預設字體。 預設情況下,`SaveAsSearchablePdf`使用Times New Roman,這不完全支持所有Unicode字元。 這影響有重音或非ASCII字元的語言。
為了解決這個問題,請提供一個支持Unicode的字體文件作為第三個參數:
```csharp
result.SaveAsSearchablePdf("output.pdf", false, "Fonts/LiberationSerif-Regular.ttf");
```
您還可以指定自定義字體名稱作為第四個參數:
```csharp
result.SaveAsSearchablePdf("output.pdf", false, "Fonts/LiberationSerif-Regular.ttf", "MyFont");
```
這適用於所有結果型別,包括`OcrDocAdvancedResult`,因此無論哪種讀取方法產生的結果,修正都能起作用。
[[i:(對於最初使用Times New Roman排版的文件,建議使用Liberation Serif,因為它在度量上相容,可以保持原有的間距和佈局。" 出於多語言通用目的,Noto Sans或DejaVu Sans是良好的替代選擇。)]]
在無法寫入文件路徑的情況下,IronOCR還支持將可搜尋的PDF作為位元組陣列或流返回。
<hr />
## 如何將可搜尋的PDF導出為位元組或流?
可搜尋PDF的輸出也可以用位元組或流方式處理,分別使用`SaveAsSearchablePdfStream`方法。 下面的程式碼範例展示了如何使用這些方法。
```csharp
// Return as a byte array: suited for storing in a database or sending in an HTTP response body
byte[] pdfByte = ocrResult.SaveAsSearchablePdfBytes();
// Return as a stream: suited for uploading to cloud storage or piping to another I/O operation without buffering the full file
Stream pdfStream = ocrResult.SaveAsSearchablePdfStream();
```
這些輸出選項在整合雲儲存服務、資料庫或web應用程式時非常有用,特別是在文件系統存取有限的情況下。 以下例子展示了實際應用:
```csharp
using IronOcr;
using System.IO;
public class SearchablePdfExporter
{
public async Task ProcessAndUploadPdf(string inputPath)
{
var ocr = new IronTesseract
{
Configuration = { RenderSearchablePdf = true }
};
// Process the input
using var input = new OcrImageInput(inputPath);
var result = ocr.Read(input);
// Option 1: Save to database as byte array
byte[] pdfBytes = result.SaveAsSearchablePdfBytes();
// Store pdfBytes in database BLOB field
// Option 2: Upload to cloud storage using stream
using (Stream pdfStream = result.SaveAsSearchablePdfStream())
{
// Upload stream to Azure Blob Storage, AWS S3, etc.
await UploadToCloudStorage(pdfStream, "searchable-output.pdf");
}
// Option 3: Return as web response
// return File(pdfBytes, "application/pdf", "searchable.pdf");
}
private async Task UploadToCloudStorage(Stream stream, string fileName)
{
// Cloud upload implementation
}
}
```
### 性能考量
在處理大量文件時,考慮實施[多執行緒OCR操作](https://ironsoftware.com/csharp/ocr/examples/csharp-tesseract-multithreading-for-speed/)來提高吞吐量。 IronOCR支持並行處理,允許您同時處理多個文件:
```csharp
using IronOcr;
using System.Threading.Tasks;
using System.Collections.Concurrent;
public class BatchPdfProcessor
{
private readonly IronTesseract _ocr;
public BatchPdfProcessor()
{
_ocr = new IronTesseract
{
Configuration =
{
RenderSearchablePdf = true,
// Configure for optimal performance
Language = OcrLanguage.English
}
};
}
public async Task ProcessBatchAsync(string[] filePaths)
{
var results = new ConcurrentBag<(string source, string output)>();
await Parallel.ForEachAsync(filePaths, async (filePath, ct) =>
{
using var input = new OcrImageInput(filePath);
var result = _ocr.Read(input);
string outputPath = Path.ChangeExtension(filePath, ".searchable.pdf");
result.SaveAsSearchablePdf(outputPath);
results.Add((filePath, outputPath));
});
Console.WriteLine($"Processed {results.Count} files");
}
}
```
### 進階配置選項
對於更進階的情境,您可以利用[詳細的Tesseract配置](https://ironsoftware.com/csharp/ocr/examples/csharp-configure-setup-tesseract/)來微調OCR引擎針對具體文件型別或語言:
```csharp
var advancedOcr = new IronTesseract
{
Configuration =
{
RenderSearchablePdf = true,
TesseractVariables = new Dictionary<string, object>
{
{ "preserve_interword_spaces", 1 },
{ "tessedit_char_whitelist", "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz" }
},
PageSegmentationMode = TesseractPageSegmentationMode.SingleColumn
},
Language = OcrLanguage.EnglishBest
};
```
這些配置選項同樣適用於所有三種輸出方法:`SaveAsSearchablePdfStream`。 下面的摘要收集了完整的可搜尋PDF方法集及其相應的輸出格式。
## 總結
使用IronOCR建立可搜尋的PDF既簡單又靈活。 無論您需要處理單張照片、由`ReadDocumentAdvanced`進行的高級文件掃描,該程式庫都提供了強大的方法以多種格式生成可搜尋的PDF。 使用`ModelType`參數在標準和增強型ML模型之間選擇以提高準確性。 能夠導出為文件、位元組或流使它可以適應任何應用程式架構,從桌面應用到基於雲的服務。
對於更進階的OCR情境,探索[綜合程式碼範例](https://ironsoftware.com/csharp/ocr/examples/csharp-tesseract-5/)或參考API文件以獲取詳細的方法簽名和選項。
using IronOcr;// Create the OCR engine: defaults to English with balanced speed and accuracyIronTesseract ocrTesseract = new IronTesseract();// Required: without this flag the text overlay layer is not built, and SaveAsSearchablePdf produces a plain image PDFocrTesseract.Configuration.RenderSearchablePdf = true;// Wrap the TIFF in OcrImageInput: handles DPI detection and page layout automaticallyusing var imageInput = new OcrImageInput("Potter.tiff");// Run OCR; returns a result containing the recognized text and spatial layout dataOcrResult ocrResult = ocrTesseract.Read(imageInput);// Write the output: the original scanned image is preserved with an invisible text layer on topocrResult.SaveAsSearchablePdf("searchablePdf.pdf");
using IronOcr;
// Create the OCR engine: defaults to English with balanced speed and accuracy
IronTesseract ocrTesseract = new IronTesseract();
// Required: without this flag the text overlay layer is not built, and SaveAsSearchablePdf produces a plain image PDF
ocrTesseract.Configuration.RenderSearchablePdf = true;
// Wrap the TIFF in OcrImageInput: handles DPI detection and page layout automatically
using var imageInput = new OcrImageInput("Potter.tiff");
// Run OCR; returns a result containing the recognized text and spatial layout data
OcrResult ocrResult = ocrTesseract.Read(imageInput);
// Write the output: the original scanned image is preserved with an invisible text layer on top
ocrResult.SaveAsSearchablePdf("searchablePdf.pdf");
ImportsIronOcr' Create the OCR engine: defaults to English with balanced speed and accuracyDim ocrTesseract As New IronTesseract()' Required: without this flag the text overlay layer is not built, and SaveAsSearchablePdf produces a plain image PDFocrTesseract.Configuration.RenderSearchablePdf = True' Wrap the TIFF in OcrImageInput: handles DPI detection and page layout automaticallyUsing imageInput As New OcrImageInput("Potter.tiff") ' Run OCR; returns a result containing the recognized text and spatial layout data Dim ocrResult AsOcrResult = ocrTesseract.Read(imageInput) ' Write the output: the original scanned image is preserved with an invisible text layer on top ocrResult.SaveAsSearchablePdf("searchablePdf.pdf")EndUsing
Imports IronOcr
' Create the OCR engine: defaults to English with balanced speed and accuracy
Dim ocrTesseract As New IronTesseract()
' Required: without this flag the text overlay layer is not built, and SaveAsSearchablePdf produces a plain image PDF
ocrTesseract.Configuration.RenderSearchablePdf = True
' Wrap the TIFF in OcrImageInput: handles DPI detection and page layout automatically
Using imageInput As New OcrImageInput("Potter.tiff")
' Run OCR; returns a result containing the recognized text and spatial layout data
Dim ocrResult As OcrResult = ocrTesseract.Read(imageInput)
' Write the output: the original scanned image is preserved with an invisible text layer on top
ocrResult.SaveAsSearchablePdf("searchablePdf.pdf")
End Using
using IronOcr;var ocr = new IronTesseract();using var input = new OcrInput();input.LoadImage("photo.png");// ReadPhoto with Enhanced modelOcrPhotoResult photoResult = ocr.ReadPhoto(input, ModelType.Enhanced);Console.WriteLine(photoResult.Text);// Save as searchable PDFbyte[] pdfBytes = photoResult.SaveAsSearchablePdfBytes();File.WriteAllBytes("searchable-photo.pdf", pdfBytes);
using IronOcr;
var ocr = new IronTesseract();
using var input = new OcrInput();
input.LoadImage("photo.png");
// ReadPhoto with Enhanced model
OcrPhotoResult photoResult = ocr.ReadPhoto(input, ModelType.Enhanced);
Console.WriteLine(photoResult.Text);
// Save as searchable PDF
byte[] pdfBytes = photoResult.SaveAsSearchablePdfBytes();
File.WriteAllBytes("searchable-photo.pdf", pdfBytes);
using IronOcr;var ocr = new IronTesseract();using var input = new OcrInput();input.LoadImage("invoice.png");// ReadDocumentAdvanced with Enhanced modelOcrDocAdvancedResult docResult = ocr.ReadDocumentAdvanced(input, ModelType.Enhanced);byte[] docPdfBytes = docResult.SaveAsSearchablePdfBytes();File.WriteAllBytes("searchable-doc.pdf", docPdfBytes);
using IronOcr;
var ocr = new IronTesseract();
using var input = new OcrInput();
input.LoadImage("invoice.png");
// ReadDocumentAdvanced with Enhanced model
OcrDocAdvancedResult docResult = ocr.ReadDocumentAdvanced(input, ModelType.Enhanced);
byte[] docPdfBytes = docResult.SaveAsSearchablePdfBytes();
File.WriteAllBytes("searchable-doc.pdf", docPdfBytes);
一份來自Hartwell Capital Management的11頁年度報告,通過OcrPdfInput載入。 使用Read調用中處理它們。
multi-page-scan.pdf:用作多頁可搜尋PDF轉換輸入的11頁Hartwell Capital Management年度報告。
using IronOcr;// Create the OCR engine. RenderSearchablePdf is false by default; no need to set it when using OcrPdfInput directlyvar ocrTesseract = new IronTesseract();// Load pages 1–10 (indices 0–9) only; PageIndices avoids loading and OCR-ing the full document unnecessarilyusing var pdfInput = new OcrPdfInput("multi-page-scan.pdf", PageIndices: Enumerable.Range(0, 10));// Run OCR across all selected pages in orderOcrResult result = ocrTesseract.Read(pdfInput);// Write the searchable PDF; true = apply the input's image filters to the embedded page images in the outputresult.SaveAsSearchablePdf("searchable-multi-page.pdf", true);
using IronOcr;
// Create the OCR engine. RenderSearchablePdf is false by default; no need to set it when using OcrPdfInput directly
var ocrTesseract = new IronTesseract();
// Load pages 1–10 (indices 0–9) only; PageIndices avoids loading and OCR-ing the full document unnecessarily
using var pdfInput = new OcrPdfInput("multi-page-scan.pdf", PageIndices: Enumerable.Range(0, 10));
// Run OCR across all selected pages in order
OcrResult result = ocrTesseract.Read(pdfInput);
// Write the searchable PDF; true = apply the input's image filters to the embedded page images in the output
result.SaveAsSearchablePdf("searchable-multi-page.pdf", true);
ImportsIronOcr' Create the OCR engine. RenderSearchablePdf is false by default; no need to set it when using OcrPdfInput directlyDim ocrTesseract As New IronTesseract()' Load pages 1–10 (indices 0–9) only; PageIndices avoids loading and OCR-ing the full document unnecessarilyUsing pdfInput As New OcrPdfInput("multi-page-scan.pdf", PageIndices:=Enumerable.Range(0, 10)) ' Run OCR across all selected pages in order Dim result AsOcrResult = ocrTesseract.Read(pdfInput) ' Write the searchable PDF; true = apply the input's image filters to the embedded page images in the output result.SaveAsSearchablePdf("searchable-multi-page.pdf", True)EndUsing
Imports IronOcr
' Create the OCR engine. RenderSearchablePdf is false by default; no need to set it when using OcrPdfInput directly
Dim ocrTesseract As New IronTesseract()
' Load pages 1–10 (indices 0–9) only; PageIndices avoids loading and OCR-ing the full document unnecessarily
Using pdfInput As New OcrPdfInput("multi-page-scan.pdf", PageIndices:=Enumerable.Range(0, 10))
' Run OCR across all selected pages in order
Dim result As OcrResult = ocrTesseract.Read(pdfInput)
' Write the searchable PDF; true = apply the input's image filters to the embedded page images in the output
result.SaveAsSearchablePdf("searchable-multi-page.pdf", True)
End Using
using IronOcr;// Create OCR engine: filters are applied at the OcrInput level, so no configuration changes are needed herevar ocr = new IronTesseract();var ocrInput = new OcrInput();// Load the scanned PDF as the OCR sourceocrInput.LoadPdf("invoice.pdf");// Convert to grayscale: removes color noise that can reduce OCR accuracy on color-printed documentsocrInput.ToGrayScale();// Run OCR on the preprocessed inputOcrResult result = ocr.Read(ocrInput);// Write the searchable PDF; true = embed the grayscale-filtered image rather than the original color scanresult.SaveAsSearchablePdf("outputGrayscale.pdf", true);
using IronOcr;
// Create OCR engine: filters are applied at the OcrInput level, so no configuration changes are needed here
var ocr = new IronTesseract();
var ocrInput = new OcrInput();
// Load the scanned PDF as the OCR source
ocrInput.LoadPdf("invoice.pdf");
// Convert to grayscale: removes color noise that can reduce OCR accuracy on color-printed documents
ocrInput.ToGrayScale();
// Run OCR on the preprocessed input
OcrResult result = ocr.Read(ocrInput);
// Write the searchable PDF; true = embed the grayscale-filtered image rather than the original color scan
result.SaveAsSearchablePdf("outputGrayscale.pdf", true);
ImportsIronOcr' Create OCR engine: filters are applied at the OcrInput level, so no configuration changes are needed hereDim ocr As New IronTesseract()Dim ocrInput As New OcrInput()' Load the scanned PDF as the OCR sourceocrInput.LoadPdf("invoice.pdf")' Convert to grayscale: removes color noise that can reduce OCR accuracy on color-printed documentsocrInput.ToGrayScale()' Run OCR on the preprocessed inputDim result AsOcrResult = ocr.Read(ocrInput)' Write the searchable PDF; True = embed the grayscale-filtered image rather than the original color scanresult.SaveAsSearchablePdf("outputGrayscale.pdf", True)
Imports IronOcr
' Create OCR engine: filters are applied at the OcrInput level, so no configuration changes are needed here
Dim ocr As New IronTesseract()
Dim ocrInput As New OcrInput()
' Load the scanned PDF as the OCR source
ocrInput.LoadPdf("invoice.pdf")
' Convert to grayscale: removes color noise that can reduce OCR accuracy on color-printed documents
ocrInput.ToGrayScale()
' Run OCR on the preprocessed input
Dim result As OcrResult = ocr.Read(ocrInput)
' Write the searchable PDF; True = embed the grayscale-filtered image rather than the original color scan
result.SaveAsSearchablePdf("outputGrayscale.pdf", True)
// Return as a byte array: suited for storing in a database or sending in an HTTP response bodybyte[] pdfByte = ocrResult.SaveAsSearchablePdfBytes();// Return as a stream: suited for uploading to cloud storage or piping to another I/O operation without buffering the full fileStream pdfStream = ocrResult.SaveAsSearchablePdfStream();
// Return as a byte array: suited for storing in a database or sending in an HTTP response body
byte[] pdfByte = ocrResult.SaveAsSearchablePdfBytes();
// Return as a stream: suited for uploading to cloud storage or piping to another I/O operation without buffering the full file
Stream pdfStream = ocrResult.SaveAsSearchablePdfStream();
' Return as a byte array: suited for storing in a database or sending in an HTTP response bodyDim pdfByte AsByte() = ocrResult.SaveAsSearchablePdfBytes()' Return as a stream: suited for uploading to cloud storage or piping to another I/O operation without buffering the full fileDim pdfStream AsStream = ocrResult.SaveAsSearchablePdfStream()
' Return as a byte array: suited for storing in a database or sending in an HTTP response body
Dim pdfByte As Byte() = ocrResult.SaveAsSearchablePdfBytes()
' Return as a stream: suited for uploading to cloud storage or piping to another I/O operation without buffering the full file
Dim pdfStream As Stream = ocrResult.SaveAsSearchablePdfStream()
using IronOcr;using System.IO;public class SearchablePdfExporter{ public async TaskProcessAndUploadPdf(string inputPath) { var ocr = new IronTesseract {Configuration = { RenderSearchablePdf = true } }; // Process the input using var input = new OcrImageInput(inputPath); var result = ocr.Read(input); // Option 1: Save to database as byte array byte[] pdfBytes = result.SaveAsSearchablePdfBytes(); // Store pdfBytes in database BLOB field // Option 2: Upload to cloud storage using stream using (Stream pdfStream = result.SaveAsSearchablePdfStream()) { // Upload stream to Azure Blob Storage, AWS S3, etc. awaitUploadToCloudStorage(pdfStream, "searchable-output.pdf"); } // Option 3: Return as web response // return File(pdfBytes, "application/pdf", "searchable.pdf"); } private async TaskUploadToCloudStorage(Stream stream, string fileName) { // Cloud upload implementation }}
using IronOcr;
using System.IO;
public class SearchablePdfExporter
{
public async Task ProcessAndUploadPdf(string inputPath)
{
var ocr = new IronTesseract
{
Configuration = { RenderSearchablePdf = true }
};
// Process the input
using var input = new OcrImageInput(inputPath);
var result = ocr.Read(input);
// Option 1: Save to database as byte array
byte[] pdfBytes = result.SaveAsSearchablePdfBytes();
// Store pdfBytes in database BLOB field
// Option 2: Upload to cloud storage using stream
using (Stream pdfStream = result.SaveAsSearchablePdfStream())
{
// Upload stream to Azure Blob Storage, AWS S3, etc.
await UploadToCloudStorage(pdfStream, "searchable-output.pdf");
}
// Option 3: Return as web response
// return File(pdfBytes, "application/pdf", "searchable.pdf");
}
private async Task UploadToCloudStorage(Stream stream, string fileName)
{
// Cloud upload implementation
}
}
ImportsIronOcrImportsSystem.IOPublic Class SearchablePdfExporter PublicAsync Function ProcessAndUploadPdf(inputPath AsString) AsTask Dim ocr As New IronTesseractWith { .Configuration = { .RenderSearchablePdf = True } } ' Process the inputUsing input As New OcrImageInput(inputPath) Dim result = ocr.Read(input) ' Option 1: Save to database as byte array Dim pdfBytes AsByte() = result.SaveAsSearchablePdfBytes() ' Store pdfBytes in database BLOB field ' Option 2: Upload to cloud storage using streamUsing pdfStream AsStream = result.SaveAsSearchablePdfStream() ' Upload stream to Azure Blob Storage, AWS S3, etc.AwaitUploadToCloudStorage(pdfStream, "searchable-output.pdf")EndUsing ' Option 3: Return as web response ' Return File(pdfBytes, "application/pdf", "searchable.pdf")EndUsing End Function PrivateAsync Function UploadToCloudStorage(stream AsStream, fileName AsString) AsTask ' Cloud upload implementation End FunctionEnd Class
Imports IronOcr
Imports System.IO
Public Class SearchablePdfExporter
Public Async Function ProcessAndUploadPdf(inputPath As String) As Task
Dim ocr As New IronTesseract With {
.Configuration = { .RenderSearchablePdf = True }
}
' Process the input
Using input As New OcrImageInput(inputPath)
Dim result = ocr.Read(input)
' Option 1: Save to database as byte array
Dim pdfBytes As Byte() = result.SaveAsSearchablePdfBytes()
' Store pdfBytes in database BLOB field
' Option 2: Upload to cloud storage using stream
Using pdfStream As Stream = result.SaveAsSearchablePdfStream()
' Upload stream to Azure Blob Storage, AWS S3, etc.
Await UploadToCloudStorage(pdfStream, "searchable-output.pdf")
End Using
' Option 3: Return as web response
' Return File(pdfBytes, "application/pdf", "searchable.pdf")
End Using
End Function
Private Async Function UploadToCloudStorage(stream As Stream, fileName As String) As Task
' Cloud upload implementation
End Function
End Class
using IronOcr;using System.Threading.Tasks;using System.Collections.Concurrent;public class BatchPdfProcessor{ private readonly IronTesseract _ocr; publicBatchPdfProcessor() { _ocr = new IronTesseract {Configuration = {RenderSearchablePdf = true, // Configure for optimal performanceLanguage = OcrLanguage.English } }; } public async TaskProcessBatchAsync(string[] filePaths) { var results = new ConcurrentBag<(string source, string output)>(); awaitParallel.ForEachAsync(filePaths, async (filePath, ct) => { using var input = new OcrImageInput(filePath); var result = _ocr.Read(input); string outputPath = Path.ChangeExtension(filePath, ".searchable.pdf"); result.SaveAsSearchablePdf(outputPath); results.Add((filePath, outputPath)); });Console.WriteLine($"Processed {results.Count} files"); }}
using IronOcr;
using System.Threading.Tasks;
using System.Collections.Concurrent;
public class BatchPdfProcessor
{
private readonly IronTesseract _ocr;
public BatchPdfProcessor()
{
_ocr = new IronTesseract
{
Configuration =
{
RenderSearchablePdf = true,
// Configure for optimal performance
Language = OcrLanguage.English
}
};
}
public async Task ProcessBatchAsync(string[] filePaths)
{
var results = new ConcurrentBag<(string source, string output)>();
await Parallel.ForEachAsync(filePaths, async (filePath, ct) =>
{
using var input = new OcrImageInput(filePath);
var result = _ocr.Read(input);
string outputPath = Path.ChangeExtension(filePath, ".searchable.pdf");
result.SaveAsSearchablePdf(outputPath);
results.Add((filePath, outputPath));
});
Console.WriteLine($"Processed {results.Count} files");
}
}
ImportsIronOcrImportsSystem.Threading.TasksImportsSystem.Collections.ConcurrentPublic Class BatchPdfProcessor PrivateReadOnly _ocr AsIronTesseract Public Sub New() _ocr = New IronTesseractWith { .Configuration = { .RenderSearchablePdf = True, ' Configure for optimal performance .Language = OcrLanguage.English } } End Sub PublicAsync Function ProcessBatchAsync(filePaths AsString()) AsTask Dim results As New ConcurrentBag(Of (source AsString, output AsString))()AwaitTask.Run(Sub()Parallel.ForEach(filePaths, Sub(filePath)Using input As New OcrImageInput(filePath) Dim result = _ocr.Read(input) Dim outputPath AsString = Path.ChangeExtension(filePath, ".searchable.pdf") result.SaveAsSearchablePdf(outputPath) results.Add((filePath, outputPath))EndUsing End Sub) End Function)Console.WriteLine($"Processed {results.Count} files") End FunctionEnd Class
Imports IronOcr
Imports System.Threading.Tasks
Imports System.Collections.Concurrent
Public Class BatchPdfProcessor
Private ReadOnly _ocr As IronTesseract
Public Sub New()
_ocr = New IronTesseract With {
.Configuration = {
.RenderSearchablePdf = True,
' Configure for optimal performance
.Language = OcrLanguage.English
}
}
End Sub
Public Async Function ProcessBatchAsync(filePaths As String()) As Task
Dim results As New ConcurrentBag(Of (source As String, output As String))()
Await Task.Run(Sub()
Parallel.ForEach(filePaths,
Sub(filePath)
Using input As New OcrImageInput(filePath)
Dim result = _ocr.Read(input)
Dim outputPath As String = Path.ChangeExtension(filePath, ".searchable.pdf")
result.SaveAsSearchablePdf(outputPath)
results.Add((filePath, outputPath))
End Using
End Sub)
End Function)
Console.WriteLine($"Processed {results.Count} files")
End Function
End Class
var advancedOcr = new IronTesseract{Configuration = {RenderSearchablePdf = true,TesseractVariables = new Dictionary<string, object> { { "preserve_interword_spaces", 1 }, { "tessedit_char_whitelist", "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz" } },PageSegmentationMode = TesseractPageSegmentationMode.SingleColumn },Language = OcrLanguage.EnglishBest};
var advancedOcr = new IronTesseract
{
Configuration =
{
RenderSearchablePdf = true,
TesseractVariables = new Dictionary<string, object>
{
{ "preserve_interword_spaces", 1 },
{ "tessedit_char_whitelist", "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz" }
},
PageSegmentationMode = TesseractPageSegmentationMode.SingleColumn
},
Language = OcrLanguage.EnglishBest
};
ImportsIronOcrDim advancedOcr As New IronTesseractWith { .Configuration = New TesseractConfigurationWith { .RenderSearchablePdf = True, .TesseractVariables = New Dictionary(OfString, Object) From { {"preserve_interword_spaces", 1}, {"tessedit_char_whitelist", "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"} }, .PageSegmentationMode = TesseractPageSegmentationMode.SingleColumn }, .Language = OcrLanguage.EnglishBest}
Imports IronOcr
Dim advancedOcr As New IronTesseract With {
.Configuration = New TesseractConfiguration With {
.RenderSearchablePdf = True,
.TesseractVariables = New Dictionary(Of String, Object) From {
{"preserve_interword_spaces", 1},
{"tessedit_char_whitelist", "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"}
},
.PageSegmentationMode = TesseractPageSegmentationMode.SingleColumn
},
.Language = OcrLanguage.EnglishBest
}
這是因為可搜尋文字層中使用的預設字體(Times New Roman)不完全支持所有Unicode字元。要解決此問題,請將Unicode相容的字體文件作為SaveAsSearchablePdf的第三個參數傳入。如果您的文件最初是Times New Roman排版的,並且在使用其他字體時發現間距不一致,請嘗試Liberation Serif,因為它具有相同的字形尺寸,並保持原始佈局。
What advanced configuration options are available in IronOCR for creating searchable PDFs?
IronOCR offers advanced configuration options, including detailed Tesseract configuration, setting Tesseract variables, and choosing different page segmentation modes. These can be tailored to fine-tune OCR for specific document types or languages.