# 使用 IronOCR 在 C# 中保存可搜索 PDF
IronOCR 使 C# 开发人员能够使用 [OCR 技术](https://ironsoftware.com/csharp/ocr/features/)将扫描的文档和图像转换为可搜索的 PDF,只需几行代码即可支持以文件、字节或流的形式输出。
可搜索的 PDF,通常被称为 OCR(光学字符识别)PDF,是一种包含扫描图像和机器可读文本的 PDF 文档。 这些 PDF 文件是通过对扫描的纸质文档或图像执行 OCR 功能创建的,它可以识别图像中的文本,并将其转换为可选择和可搜索的文本。
`ReadDocumentAdvanced`的结果中使用,使得从照片和高级文档OCR工作流创建可搜索的PDF成为可能。 在将纸质档案数字化或使旧版 PDF 文件可搜索以实现更佳文档管理时,此功能尤为实用。
*as-heading:2(快速入门:一行导出可搜索 PDF)*
设置`SaveAsSearchablePdf(...)`。 只需这些步骤,即可使用 IronOCR 生成一份完全可搜索的 PDF 文件。
```cs
:title=Quickly Make a PDF Searchable with IronOCR
new IronOcr.IronTesseract { Configuration = { RenderSearchablePdf = true } } .Read(new IronOcr.OcrImageInput("file.jpg")).SaveAsSearchablePdf("searchable.pdf");
```
<div class="hsg-featured-snippet">
<h3>最小工作流程(5 个步骤)</h3>
<ol>
<li><a class="js-modal-open" data-modal-id="trial-license-after-download" href="https://nuget.org/packages/IronOcr/">下载一个 C# 库,用于将结果保存为可搜索的 PDF 文件。</a></li>
<li>为OCR准备图像和PDF文档</li>
<li>将 <strong>RenderSearchablePdf</strong> 属性设置为 <code>true</code></li>
<li>使用<code>SaveAsSearchablePdf</code>方法输出可搜索的 PDF 文件</li>
<li>将可搜索的 PDF 导出为字节流</li>
</ol>
</div>
<br class="clear" />
## 如何将 OCR 结果导出为可搜索的 PDF?
要使用IronOCR将结果导出为可搜索的PDF,请将`SaveAsSearchablePdf`。
### 输入
从哈利波特小说中的一页,扫描为TIFF文件并通过`OcrImageInput`加载。 该页面包含密集的印刷文本,是用于测试可搜索 PDF 文本层的理想测试素材。
<div class="content-img-align-center">
<div class="center-image-wrapper" style="max-width: 380px; margin: 0 auto;">
<img src="/static-assets/ocr/how-to/searchable-pdf/potter.webp" alt="《哈利·波特》书中展示第八章"死亡纪念日派对"的页面,内容讲述哈利与"几乎无头尼克"的相遇" class="img-responsive add-shadow" />
</div>
</div>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">potter.tiff:作为OCR输入的扫描小说页面,用于生成带有不可见文本图层的可搜索PDF文件。</p>
```csharp
using IronOcr;
// Create the OCR engine: defaults to English with balanced speed and accuracy
IronTesseract ocrTesseract = new IronTesseract();
// Required: without this flag the text overlay layer is not built, and SaveAsSearchablePdf produces a plain image PDF
ocrTesseract.Configuration.RenderSearchablePdf = true;
// Wrap the TIFF in OcrImageInput: handles DPI detection and page layout automatically
using var imageInput = new OcrImageInput("Potter.tiff");
// Run OCR; returns a result containing the recognized text and spatial layout data
OcrResult ocrResult = ocrTesseract.Read(imageInput);
// Write the output: the original scanned image is preserved with an invisible text layer on top
ocrResult.SaveAsSearchablePdf("searchablePdf.pdf");
```
### 输出
<iframe loading="lazy" src="/static-assets/ocr/how-to/searchable-pdf/searchablePdf.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">searchablePdf.pdf:可搜索的 PDF 输出文件。请选中或搜索任意单词,以验证 OCR 文本层。</p>
生成的 PDF 文件将原始扫描的页面图像嵌入其中,并在每个识别出的单词上方叠加了一层不可见的文本层。 在查看器中选择或搜索任意WORD,以确认文本图层是否存在。
IronOCR 在叠加层中使用了特定字体,这可能会导致渲染后的文本大小与原文存在轻微差异。
在处理[多页 TIFF 文件](https://ironsoftware.com/csharp/ocr/examples/csharp-tesseract-multipage-tiff/)或复杂文档时,IronOCR 会自动处理所有页面并将它们包含在输出中。 该库可自动处理页面排序和文本叠加定位,确保文本到图像的映射准确无误。
### 如何从照片或高级文档扫描创建可搜索的 PDF?
使用`ReadDocumentAdvanced`时,也可导出可搜索PDF。 每种方法返回一种支持`SaveAsSearchablePdf`的结果类型。
调用这些方法时,您可以选择传递`ModelType`。 默认是`Enhanced`在精度上更好,但速度较慢。
#### 输入
一幅墙壁壁画内嵌有文字的照片,通过`LoadImage`加载。 场景中包含嵌入在真实环境中的多个单词,使其成为`Enhanced`模型的实际测试。
<div class="content-img-align-center">
<div class="center-image-wrapper" style="max-width: 450px; margin: 0 auto;">
<img src="/static-assets/ocr/how-to/searchable-pdf/photo-input.webp" alt="用作 ReadPhoto OCR 输入的含文字图片" class="img-responsive add-shadow" />
</div>
</div>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">photo.png:通过 ReadPhoto 并使用增强型模型加载的墙面壁画照片,用于生成可搜索的 PDF 文件。</p>
```csharp
using IronOcr;
var ocr = new IronTesseract();
using var input = new OcrInput();
input.LoadImage("photo.png");
// ReadPhoto with Enhanced model
OcrPhotoResult photoResult = ocr.ReadPhoto(input, ModelType.Enhanced);
Console.WriteLine(photoResult.Text);
// Save as searchable PDF
byte[] pdfBytes = photoResult.SaveAsSearchablePdfBytes();
File.WriteAllBytes("searchable-photo.pdf", pdfBytes);
```
#### 输出
<iframe loading="lazy" src="/static-assets/ocr/how-to/searchable-pdf/searchable-photo.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">searchable-photo.pdf:由 ReadPhoto 生成的可搜索 PDF。其文本层支持在任何 PDF 阅读器中进行全文搜索。</p>
生成的可搜索 PDF 文件会在识别出的文字上方叠加一层不可见的文本层。 在 PDF 阅读器中搜索"Milk"会返回 3 个匹配结果,这些结果直接提取自原始照片中的手写文本。
类似方法适用于`OcrDocAdvancedResult`:
#### 输入
通过`LoadImage`加载的扫描发票。 它包含结构化字段(供应商名称、行项目和总计),而`Enhanced`模型识别并将其内嵌为可搜索文本层。
<div class="content-img-align-center">
<div class="center-image-wrapper" style="max-width: 400px; margin: 0 auto;">
<img src="/static-assets/ocr/how-to/searchable-pdf/invoice-input.webp" alt="用作 ReadDocumentAdvanced OCR 输入的发票文档" class="img-responsive add-shadow" />
</div>
</div>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">invoice.png:扫描的发票已加载到 OcrInput 中,并使用增强型模型传递给 ReadDocumentAdvanced。</p>
```csharp
using IronOcr;
var ocr = new IronTesseract();
using var input = new OcrInput();
input.LoadImage("invoice.png");
// ReadDocumentAdvanced with Enhanced model
OcrDocAdvancedResult docResult = ocr.ReadDocumentAdvanced(input, ModelType.Enhanced);
byte[] docPdfBytes = docResult.SaveAsSearchablePdfBytes();
File.WriteAllBytes("searchable-doc.pdf", docPdfBytes);
```
#### 输出
<iframe loading="lazy" src="/static-assets/ocr/how-to/searchable-pdf/searchable-doc.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">searchable-doc.pdf:由 ReadDocumentAdvanced 生成的可搜索 PDF。发票字段可供选择和搜索。</p>
`ExtensionAdvancedScanException`。
### 处理多页文档
在处理多页文档的 PDF OCR 操作时,IronOCR 会按顺序处理每一页,并保持原始文档的结构。
#### 输入
通过`OcrPdfInput`加载的11页Hartwell Capital Management年度报告。 使用`Read`调用中处理这些页面。
<iframe loading="lazy" src="/static-assets/ocr/how-to/searchable-pdf/multi-page-scan.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">multi-page-scan.pdf:一份 11 页的 Hartwell Capital Management 年报,用作多页可搜索 PDF 转换的输入文件。</p>
```csharp
using IronOcr;
// Create the OCR engine. RenderSearchablePdf is false by default; no need to set it when using OcrPdfInput directly
var ocrTesseract = new IronTesseract();
// Load pages 1–10 (indices 0–9) only; PageIndices avoids loading and OCR-ing the full document unnecessarily
using var pdfInput = new OcrPdfInput("multi-page-scan.pdf", PageIndices: Enumerable.Range(0, 10));
// Run OCR across all selected pages in order
OcrResult result = ocrTesseract.Read(pdfInput);
// Write the searchable PDF; true = apply the input's image filters to the embedded page images in the output
result.SaveAsSearchablePdf("searchable-multi-page.pdf", true);
```
#### 输出
<iframe loading="lazy" src="/static-assets/ocr/how-to/searchable-pdf/searchable-multi-page.pdf" width="100%" height="400px"></iframe>
<p style="text-align: center; font-style: italic; color: #555; font-size: 13px; margin-top: 6px;">searchable-multi-page.pdf:10 页的可搜索 PDF 输出文件。每页均包含一个用于全文搜索的不可见文本图层。</p>
生成的 PDF 文件共 10 页(源自原始报告的第 1–10 页),每页均包含一个不可见的文本图层,使提取的内容可在任何 PDF 阅读器中被选中并进行搜索。
### 在创建可搜索 PDF 时如何应用过滤器?
`SaveAsSearchablePdf`第二个参数接受一个布尔值,用于控制是否将图像滤波器应用于嵌入的输出。 使用[图像优化过滤器](https://ironsoftware.com/csharp/ocr/examples/ocr-image-filters-for-net-tesseract/)可以显著提高 OCR 的准确性,尤其是在处理[低质量扫描](https://ironsoftware.com/csharp/ocr/examples/ocr-low-quality-scans-tesseract/)时。
以下示例应用灰度滤波器,并将`true`作为第二个参数传递,以将滤波后的图像嵌入到可搜索PDF输出中。
```cs
using IronOcr;
// Create OCR engine: filters are applied at the OcrInput level, so no configuration changes are needed here
var ocr = new IronTesseract();
var ocrInput = new OcrInput();
// Load the scanned PDF as the OCR source
ocrInput.LoadPdf("invoice.pdf");
// Convert to grayscale: removes color noise that can reduce OCR accuracy on color-printed documents
ocrInput.ToGrayScale();
// Run OCR on the preprocessed input
OcrResult result = ocr.Read(ocrInput);
// Write the searchable PDF; true = embed the grayscale-filtered image rather than the original color scan
result.SaveAsSearchablePdf("outputGrayscale.pdf", true);
```
为获得最佳效果,请考虑使用 [Filter Wizard](https://ironsoftware.com/csharp/ocr/examples/filter-wizard/) 自动确定特定文档类型的最佳过滤器组合。 该工具会分析您的输入,并建议适当的预处理步骤。
### 如何修复可搜索PDF中的错误字符?
如果 PDF 中的文本在视觉上显示正确,但在搜索或复制时显示损坏的字符,则该问题是由可搜索文本层中使用的默认字体引起的。 默认情况下,`SaveAsSearchablePdf`使用Times New Roman,但该字体不完全支持所有Unicode字符。 这会影响包含重音字符或非 ASCII 字符的语言。
要解决这个问题,请提供与 Unicode 兼容的字体文件作为第三个参数:
```csharp
result.SaveAsSearchablePdf("output.pdf", false, "Fonts/LiberationSerif-Regular.ttf");
```
您还可以将自定义字体名称指定为第四个参数:
```csharp
result.SaveAsSearchablePdf("output.pdf", false, "Fonts/LiberationSerif-Regular.ttf", "MyFont");
```
这适用于包括`OcrDocAdvancedResult`在内的所有结果类型,因此这种修复无论是哪种读取方法产生的结果都有效。
对于最初使用 Times New Roman 排版的文档,建议使用 Liberation Serif,因为它在度量上是兼容的,可以保留原始的间距和布局。 对于通用多语言用途,Noto Sans 或 DejaVu Sans 都是不错的选择。
在无法写入文件路径的情况下,IronOCR 还支持将可搜索的 PDF 作为字节数组或流返回。
<hr />
## 如何将可搜索的 PDF 导出为字节或数据流?
可搜索PDF的输出也可以使用`SaveAsSearchablePdfStream`方法分别作为字节或流处理。 下面的代码示例展示了如何使用这些方法。
```csharp
// Return as a byte array: suited for storing in a database or sending in an HTTP response body
byte[] pdfByte = ocrResult.SaveAsSearchablePdfBytes();
// Return as a stream: suited for uploading to cloud storage or piping to another I/O operation without buffering the full file
Stream pdfStream = ocrResult.SaveAsSearchablePdfStream();
```
这些输出选项在与云存储服务、数据库或网络应用程序集成时特别有用,因为文件系统访问可能会受到限制。 以下示例展示了实际应用场景:
```csharp
using IronOcr;
using System.IO;
public class SearchablePdfExporter
{
public async Task ProcessAndUploadPdf(string inputPath)
{
var ocr = new IronTesseract
{
Configuration = { RenderSearchablePdf = true }
};
// Process the input
using var input = new OcrImageInput(inputPath);
var result = ocr.Read(input);
// Option 1: Save to database as byte array
byte[] pdfBytes = result.SaveAsSearchablePdfBytes();
// Store pdfBytes in database BLOB field
// Option 2: Upload to cloud storage using stream
using (Stream pdfStream = result.SaveAsSearchablePdfStream())
{
// Upload stream to Azure Blob Storage, AWS S3, etc.
await UploadToCloudStorage(pdfStream, "searchable-output.pdf");
}
// Option 3: Return as web response
// return File(pdfBytes, "application/pdf", "searchable.pdf");
}
private async Task UploadToCloudStorage(Stream stream, string fileName)
{
// Cloud upload implementation
}
}
```
### 性能考虑
在处理大量文件时,请考虑实施[多线程 OCR 操作](https://ironsoftware.com/csharp/ocr/examples/csharp-tesseract-multithreading-for-speed/)以提高吞吐量。 IronOCR 支持并发处理,允许您同时处理多个文档:
```csharp
using IronOcr;
using System.Threading.Tasks;
using System.Collections.Concurrent;
public class BatchPdfProcessor
{
private readonly IronTesseract _ocr;
public BatchPdfProcessor()
{
_ocr = new IronTesseract
{
Configuration =
{
RenderSearchablePdf = true,
// Configure for optimal performance
Language = OcrLanguage.English
}
};
}
public async Task ProcessBatchAsync(string[] filePaths)
{
var results = new ConcurrentBag<(string source, string output)>();
await Parallel.ForEachAsync(filePaths, async (filePath, ct) =>
{
using var input = new OcrImageInput(filePath);
var result = _ocr.Read(input);
string outputPath = Path.ChangeExtension(filePath, ".searchable.pdf");
result.SaveAsSearchablePdf(outputPath);
results.Add((filePath, outputPath));
});
Console.WriteLine($"Processed {results.Count} files");
}
}
```
### 高级配置选项
对于更高级的应用场景,您可以利用[详细的 Tesseract 配置](https://ironsoftware.com/csharp/ocr/examples/csharp-configure-setup-tesseract/),针对特定文档类型或语言对 OCR 引擎进行微调:
```csharp
var advancedOcr = new IronTesseract
{
Configuration =
{
RenderSearchablePdf = true,
TesseractVariables = new Dictionary<string, object>
{
{ "preserve_interword_spaces", 1 },
{ "tessedit_char_whitelist", "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz" }
},
PageSegmentationMode = TesseractPageSegmentationMode.SingleColumn
},
Language = OcrLanguage.EnglishBest
};
```
这些配置选项同样适用于所有三种输出方法:`SaveAsSearchablePdfStream`。 下方的摘要汇总了所有可搜索的 PDF 处理方法及其对应的输出格式。
## 摘要
using IronOCR 创建可搜索 PDF 既简单又灵活。 无论是需要通过`ReadDocumentAdvanced`进行高级文档扫描,该库均提供强大的方法来生成各种格式的可搜索PDF。 使用`ModelType`参数选择标准和增强ML模型之间的准确性。 以文件、字节或流的形式导出的功能使其能够适应从桌面应用程序到基于云的服务等任何应用程序架构。
对于更复杂的 OCR 场景,请查阅[全面的代码示例](https://ironsoftware.com/csharp/ocr/examples/csharp-tesseract-5/),或参考 API 文档以获取详细的方法签名和选项。
using IronOcr;// Create the OCR engine: defaults to English with balanced speed and accuracyIronTesseract ocrTesseract = new IronTesseract();// Required: without this flag the text overlay layer is not built, and SaveAsSearchablePdf produces a plain image PDFocrTesseract.Configuration.RenderSearchablePdf = true;// Wrap the TIFF in OcrImageInput: handles DPI detection and page layout automaticallyusing var imageInput = new OcrImageInput("Potter.tiff");// Run OCR; returns a result containing the recognized text and spatial layout dataOcrResult ocrResult = ocrTesseract.Read(imageInput);// Write the output: the original scanned image is preserved with an invisible text layer on topocrResult.SaveAsSearchablePdf("searchablePdf.pdf");
using IronOcr;
// Create the OCR engine: defaults to English with balanced speed and accuracy
IronTesseract ocrTesseract = new IronTesseract();
// Required: without this flag the text overlay layer is not built, and SaveAsSearchablePdf produces a plain image PDF
ocrTesseract.Configuration.RenderSearchablePdf = true;
// Wrap the TIFF in OcrImageInput: handles DPI detection and page layout automatically
using var imageInput = new OcrImageInput("Potter.tiff");
// Run OCR; returns a result containing the recognized text and spatial layout data
OcrResult ocrResult = ocrTesseract.Read(imageInput);
// Write the output: the original scanned image is preserved with an invisible text layer on top
ocrResult.SaveAsSearchablePdf("searchablePdf.pdf");
ImportsIronOcr' Create the OCR engine: defaults to English with balanced speed and accuracyDim ocrTesseract As New IronTesseract()' Required: without this flag the text overlay layer is not built, and SaveAsSearchablePdf produces a plain image PDFocrTesseract.Configuration.RenderSearchablePdf = True' Wrap the TIFF in OcrImageInput: handles DPI detection and page layout automaticallyUsing imageInput As New OcrImageInput("Potter.tiff") ' Run OCR; returns a result containing the recognized text and spatial layout data Dim ocrResult AsOcrResult = ocrTesseract.Read(imageInput) ' Write the output: the original scanned image is preserved with an invisible text layer on top ocrResult.SaveAsSearchablePdf("searchablePdf.pdf")EndUsing
Imports IronOcr
' Create the OCR engine: defaults to English with balanced speed and accuracy
Dim ocrTesseract As New IronTesseract()
' Required: without this flag the text overlay layer is not built, and SaveAsSearchablePdf produces a plain image PDF
ocrTesseract.Configuration.RenderSearchablePdf = True
' Wrap the TIFF in OcrImageInput: handles DPI detection and page layout automatically
Using imageInput As New OcrImageInput("Potter.tiff")
' Run OCR; returns a result containing the recognized text and spatial layout data
Dim ocrResult As OcrResult = ocrTesseract.Read(imageInput)
' Write the output: the original scanned image is preserved with an invisible text layer on top
ocrResult.SaveAsSearchablePdf("searchablePdf.pdf")
End Using
输出
searchablePdf.pdf:可搜索的 PDF 输出文件。请选中或搜索任意单词,以验证 OCR 文本层。
生成的 PDF 文件将原始扫描的页面图像嵌入其中,并在每个识别出的单词上方叠加了一层不可见的文本层。 在查看器中选择或搜索任意WORD,以确认文本图层是否存在。
photo.png:通过 ReadPhoto 并使用增强型模型加载的墙面壁画照片,用于生成可搜索的 PDF 文件。
using IronOcr;var ocr = new IronTesseract();using var input = new OcrInput();input.LoadImage("photo.png");// ReadPhoto with Enhanced modelOcrPhotoResult photoResult = ocr.ReadPhoto(input, ModelType.Enhanced);Console.WriteLine(photoResult.Text);// Save as searchable PDFbyte[] pdfBytes = photoResult.SaveAsSearchablePdfBytes();File.WriteAllBytes("searchable-photo.pdf", pdfBytes);
using IronOcr;
var ocr = new IronTesseract();
using var input = new OcrInput();
input.LoadImage("photo.png");
// ReadPhoto with Enhanced model
OcrPhotoResult photoResult = ocr.ReadPhoto(input, ModelType.Enhanced);
Console.WriteLine(photoResult.Text);
// Save as searchable PDF
byte[] pdfBytes = photoResult.SaveAsSearchablePdfBytes();
File.WriteAllBytes("searchable-photo.pdf", pdfBytes);
C#
输出
searchable-photo.pdf:由 ReadPhoto 生成的可搜索 PDF。其文本层支持在任何 PDF 阅读器中进行全文搜索。
生成的可搜索 PDF 文件会在识别出的文字上方叠加一层不可见的文本层。 在 PDF 阅读器中搜索"Milk"会返回 3 个匹配结果,这些结果直接提取自原始照片中的手写文本。
using IronOcr;var ocr = new IronTesseract();using var input = new OcrInput();input.LoadImage("invoice.png");// ReadDocumentAdvanced with Enhanced modelOcrDocAdvancedResult docResult = ocr.ReadDocumentAdvanced(input, ModelType.Enhanced);byte[] docPdfBytes = docResult.SaveAsSearchablePdfBytes();File.WriteAllBytes("searchable-doc.pdf", docPdfBytes);
using IronOcr;
var ocr = new IronTesseract();
using var input = new OcrInput();
input.LoadImage("invoice.png");
// ReadDocumentAdvanced with Enhanced model
OcrDocAdvancedResult docResult = ocr.ReadDocumentAdvanced(input, ModelType.Enhanced);
byte[] docPdfBytes = docResult.SaveAsSearchablePdfBytes();
File.WriteAllBytes("searchable-doc.pdf", docPdfBytes);
在处理多页文档的 PDF OCR 操作时,IronOCR 会按顺序处理每一页,并保持原始文档的结构。
输入
通过OcrPdfInput加载的11页Hartwell Capital Management年度报告。 使用Read调用中处理这些页面。
multi-page-scan.pdf:一份 11 页的 Hartwell Capital Management 年报,用作多页可搜索 PDF 转换的输入文件。
using IronOcr;// Create the OCR engine. RenderSearchablePdf is false by default; no need to set it when using OcrPdfInput directlyvar ocrTesseract = new IronTesseract();// Load pages 1–10 (indices 0–9) only; PageIndices avoids loading and OCR-ing the full document unnecessarilyusing var pdfInput = new OcrPdfInput("multi-page-scan.pdf", PageIndices: Enumerable.Range(0, 10));// Run OCR across all selected pages in orderOcrResult result = ocrTesseract.Read(pdfInput);// Write the searchable PDF; true = apply the input's image filters to the embedded page images in the outputresult.SaveAsSearchablePdf("searchable-multi-page.pdf", true);
using IronOcr;
// Create the OCR engine. RenderSearchablePdf is false by default; no need to set it when using OcrPdfInput directly
var ocrTesseract = new IronTesseract();
// Load pages 1–10 (indices 0–9) only; PageIndices avoids loading and OCR-ing the full document unnecessarily
using var pdfInput = new OcrPdfInput("multi-page-scan.pdf", PageIndices: Enumerable.Range(0, 10));
// Run OCR across all selected pages in order
OcrResult result = ocrTesseract.Read(pdfInput);
// Write the searchable PDF; true = apply the input's image filters to the embedded page images in the output
result.SaveAsSearchablePdf("searchable-multi-page.pdf", true);
ImportsIronOcr' Create the OCR engine. RenderSearchablePdf is false by default; no need to set it when using OcrPdfInput directlyDim ocrTesseract As New IronTesseract()' Load pages 1–10 (indices 0–9) only; PageIndices avoids loading and OCR-ing the full document unnecessarilyUsing pdfInput As New OcrPdfInput("multi-page-scan.pdf", PageIndices:=Enumerable.Range(0, 10)) ' Run OCR across all selected pages in order Dim result AsOcrResult = ocrTesseract.Read(pdfInput) ' Write the searchable PDF; true = apply the input's image filters to the embedded page images in the output result.SaveAsSearchablePdf("searchable-multi-page.pdf", True)EndUsing
Imports IronOcr
' Create the OCR engine. RenderSearchablePdf is false by default; no need to set it when using OcrPdfInput directly
Dim ocrTesseract As New IronTesseract()
' Load pages 1–10 (indices 0–9) only; PageIndices avoids loading and OCR-ing the full document unnecessarily
Using pdfInput As New OcrPdfInput("multi-page-scan.pdf", PageIndices:=Enumerable.Range(0, 10))
' Run OCR across all selected pages in order
Dim result As OcrResult = ocrTesseract.Read(pdfInput)
' Write the searchable PDF; true = apply the input's image filters to the embedded page images in the output
result.SaveAsSearchablePdf("searchable-multi-page.pdf", True)
End Using
输出
searchable-multi-page.pdf:10 页的可搜索 PDF 输出文件。每页均包含一个用于全文搜索的不可见文本图层。
生成的 PDF 文件共 10 页(源自原始报告的第 1–10 页),每页均包含一个不可见的文本图层,使提取的内容可在任何 PDF 阅读器中被选中并进行搜索。
using IronOcr;// Create OCR engine: filters are applied at the OcrInput level, so no configuration changes are needed herevar ocr = new IronTesseract();var ocrInput = new OcrInput();// Load the scanned PDF as the OCR sourceocrInput.LoadPdf("invoice.pdf");// Convert to grayscale: removes color noise that can reduce OCR accuracy on color-printed documentsocrInput.ToGrayScale();// Run OCR on the preprocessed inputOcrResult result = ocr.Read(ocrInput);// Write the searchable PDF; true = embed the grayscale-filtered image rather than the original color scanresult.SaveAsSearchablePdf("outputGrayscale.pdf", true);
using IronOcr;
// Create OCR engine: filters are applied at the OcrInput level, so no configuration changes are needed here
var ocr = new IronTesseract();
var ocrInput = new OcrInput();
// Load the scanned PDF as the OCR source
ocrInput.LoadPdf("invoice.pdf");
// Convert to grayscale: removes color noise that can reduce OCR accuracy on color-printed documents
ocrInput.ToGrayScale();
// Run OCR on the preprocessed input
OcrResult result = ocr.Read(ocrInput);
// Write the searchable PDF; true = embed the grayscale-filtered image rather than the original color scan
result.SaveAsSearchablePdf("outputGrayscale.pdf", true);
ImportsIronOcr' Create OCR engine: filters are applied at the OcrInput level, so no configuration changes are needed hereDim ocr As New IronTesseract()Dim ocrInput As New OcrInput()' Load the scanned PDF as the OCR sourceocrInput.LoadPdf("invoice.pdf")' Convert to grayscale: removes color noise that can reduce OCR accuracy on color-printed documentsocrInput.ToGrayScale()' Run OCR on the preprocessed inputDim result AsOcrResult = ocr.Read(ocrInput)' Write the searchable PDF; True = embed the grayscale-filtered image rather than the original color scanresult.SaveAsSearchablePdf("outputGrayscale.pdf", True)
Imports IronOcr
' Create OCR engine: filters are applied at the OcrInput level, so no configuration changes are needed here
Dim ocr As New IronTesseract()
Dim ocrInput As New OcrInput()
' Load the scanned PDF as the OCR source
ocrInput.LoadPdf("invoice.pdf")
' Convert to grayscale: removes color noise that can reduce OCR accuracy on color-printed documents
ocrInput.ToGrayScale()
' Run OCR on the preprocessed input
Dim result As OcrResult = ocr.Read(ocrInput)
' Write the searchable PDF; True = embed the grayscale-filtered image rather than the original color scan
result.SaveAsSearchablePdf("outputGrayscale.pdf", True)
如果 PDF 中的文本在视觉上显示正确,但在搜索或复制时显示损坏的字符,则该问题是由可搜索文本层中使用的默认字体引起的。 默认情况下,SaveAsSearchablePdf使用Times New Roman,但该字体不完全支持所有Unicode字符。 这会影响包含重音字符或非 ASCII 字符的语言。
// Return as a byte array: suited for storing in a database or sending in an HTTP response bodybyte[] pdfByte = ocrResult.SaveAsSearchablePdfBytes();// Return as a stream: suited for uploading to cloud storage or piping to another I/O operation without buffering the full fileStream pdfStream = ocrResult.SaveAsSearchablePdfStream();
// Return as a byte array: suited for storing in a database or sending in an HTTP response body
byte[] pdfByte = ocrResult.SaveAsSearchablePdfBytes();
// Return as a stream: suited for uploading to cloud storage or piping to another I/O operation without buffering the full file
Stream pdfStream = ocrResult.SaveAsSearchablePdfStream();
' Return as a byte array: suited for storing in a database or sending in an HTTP response bodyDim pdfByte AsByte() = ocrResult.SaveAsSearchablePdfBytes()' Return as a stream: suited for uploading to cloud storage or piping to another I/O operation without buffering the full fileDim pdfStream AsStream = ocrResult.SaveAsSearchablePdfStream()
' Return as a byte array: suited for storing in a database or sending in an HTTP response body
Dim pdfByte As Byte() = ocrResult.SaveAsSearchablePdfBytes()
' Return as a stream: suited for uploading to cloud storage or piping to another I/O operation without buffering the full file
Dim pdfStream As Stream = ocrResult.SaveAsSearchablePdfStream()
using IronOcr;using System.IO;public class SearchablePdfExporter{ public async TaskProcessAndUploadPdf(string inputPath) { var ocr = new IronTesseract {Configuration = { RenderSearchablePdf = true } }; // Process the input using var input = new OcrImageInput(inputPath); var result = ocr.Read(input); // Option 1: Save to database as byte array byte[] pdfBytes = result.SaveAsSearchablePdfBytes(); // Store pdfBytes in database BLOB field // Option 2: Upload to cloud storage using stream using (Stream pdfStream = result.SaveAsSearchablePdfStream()) { // Upload stream to Azure Blob Storage, AWS S3, etc. awaitUploadToCloudStorage(pdfStream, "searchable-output.pdf"); } // Option 3: Return as web response // return File(pdfBytes, "application/pdf", "searchable.pdf"); } private async TaskUploadToCloudStorage(Stream stream, string fileName) { // Cloud upload implementation }}
using IronOcr;
using System.IO;
public class SearchablePdfExporter
{
public async Task ProcessAndUploadPdf(string inputPath)
{
var ocr = new IronTesseract
{
Configuration = { RenderSearchablePdf = true }
};
// Process the input
using var input = new OcrImageInput(inputPath);
var result = ocr.Read(input);
// Option 1: Save to database as byte array
byte[] pdfBytes = result.SaveAsSearchablePdfBytes();
// Store pdfBytes in database BLOB field
// Option 2: Upload to cloud storage using stream
using (Stream pdfStream = result.SaveAsSearchablePdfStream())
{
// Upload stream to Azure Blob Storage, AWS S3, etc.
await UploadToCloudStorage(pdfStream, "searchable-output.pdf");
}
// Option 3: Return as web response
// return File(pdfBytes, "application/pdf", "searchable.pdf");
}
private async Task UploadToCloudStorage(Stream stream, string fileName)
{
// Cloud upload implementation
}
}
ImportsIronOcrImportsSystem.IOPublic Class SearchablePdfExporter PublicAsync Function ProcessAndUploadPdf(inputPath AsString) AsTask Dim ocr As New IronTesseractWith { .Configuration = { .RenderSearchablePdf = True } } ' Process the inputUsing input As New OcrImageInput(inputPath) Dim result = ocr.Read(input) ' Option 1: Save to database as byte array Dim pdfBytes AsByte() = result.SaveAsSearchablePdfBytes() ' Store pdfBytes in database BLOB field ' Option 2: Upload to cloud storage using streamUsing pdfStream AsStream = result.SaveAsSearchablePdfStream() ' Upload stream to Azure Blob Storage, AWS S3, etc.AwaitUploadToCloudStorage(pdfStream, "searchable-output.pdf")EndUsing ' Option 3: Return as web response ' Return File(pdfBytes, "application/pdf", "searchable.pdf")EndUsing End Function PrivateAsync Function UploadToCloudStorage(stream AsStream, fileName AsString) AsTask ' Cloud upload implementation End FunctionEnd Class
Imports IronOcr
Imports System.IO
Public Class SearchablePdfExporter
Public Async Function ProcessAndUploadPdf(inputPath As String) As Task
Dim ocr As New IronTesseract With {
.Configuration = { .RenderSearchablePdf = True }
}
' Process the input
Using input As New OcrImageInput(inputPath)
Dim result = ocr.Read(input)
' Option 1: Save to database as byte array
Dim pdfBytes As Byte() = result.SaveAsSearchablePdfBytes()
' Store pdfBytes in database BLOB field
' Option 2: Upload to cloud storage using stream
Using pdfStream As Stream = result.SaveAsSearchablePdfStream()
' Upload stream to Azure Blob Storage, AWS S3, etc.
Await UploadToCloudStorage(pdfStream, "searchable-output.pdf")
End Using
' Option 3: Return as web response
' Return File(pdfBytes, "application/pdf", "searchable.pdf")
End Using
End Function
Private Async Function UploadToCloudStorage(stream As Stream, fileName As String) As Task
' Cloud upload implementation
End Function
End Class
using IronOcr;using System.Threading.Tasks;using System.Collections.Concurrent;public class BatchPdfProcessor{ private readonly IronTesseract _ocr; publicBatchPdfProcessor() { _ocr = new IronTesseract {Configuration = {RenderSearchablePdf = true, // Configure for optimal performanceLanguage = OcrLanguage.English } }; } public async TaskProcessBatchAsync(string[] filePaths) { var results = new ConcurrentBag<(string source, string output)>(); awaitParallel.ForEachAsync(filePaths, async (filePath, ct) => { using var input = new OcrImageInput(filePath); var result = _ocr.Read(input); string outputPath = Path.ChangeExtension(filePath, ".searchable.pdf"); result.SaveAsSearchablePdf(outputPath); results.Add((filePath, outputPath)); });Console.WriteLine($"Processed {results.Count} files"); }}
using IronOcr;
using System.Threading.Tasks;
using System.Collections.Concurrent;
public class BatchPdfProcessor
{
private readonly IronTesseract _ocr;
public BatchPdfProcessor()
{
_ocr = new IronTesseract
{
Configuration =
{
RenderSearchablePdf = true,
// Configure for optimal performance
Language = OcrLanguage.English
}
};
}
public async Task ProcessBatchAsync(string[] filePaths)
{
var results = new ConcurrentBag<(string source, string output)>();
await Parallel.ForEachAsync(filePaths, async (filePath, ct) =>
{
using var input = new OcrImageInput(filePath);
var result = _ocr.Read(input);
string outputPath = Path.ChangeExtension(filePath, ".searchable.pdf");
result.SaveAsSearchablePdf(outputPath);
results.Add((filePath, outputPath));
});
Console.WriteLine($"Processed {results.Count} files");
}
}
ImportsIronOcrImportsSystem.Threading.TasksImportsSystem.Collections.ConcurrentPublic Class BatchPdfProcessor PrivateReadOnly _ocr AsIronTesseract Public Sub New() _ocr = New IronTesseractWith { .Configuration = { .RenderSearchablePdf = True, ' Configure for optimal performance .Language = OcrLanguage.English } } End Sub PublicAsync Function ProcessBatchAsync(filePaths AsString()) AsTask Dim results As New ConcurrentBag(Of (source AsString, output AsString))()AwaitTask.Run(Sub()Parallel.ForEach(filePaths, Sub(filePath)Using input As New OcrImageInput(filePath) Dim result = _ocr.Read(input) Dim outputPath AsString = Path.ChangeExtension(filePath, ".searchable.pdf") result.SaveAsSearchablePdf(outputPath) results.Add((filePath, outputPath))EndUsing End Sub) End Function)Console.WriteLine($"Processed {results.Count} files") End FunctionEnd Class
Imports IronOcr
Imports System.Threading.Tasks
Imports System.Collections.Concurrent
Public Class BatchPdfProcessor
Private ReadOnly _ocr As IronTesseract
Public Sub New()
_ocr = New IronTesseract With {
.Configuration = {
.RenderSearchablePdf = True,
' Configure for optimal performance
.Language = OcrLanguage.English
}
}
End Sub
Public Async Function ProcessBatchAsync(filePaths As String()) As Task
Dim results As New ConcurrentBag(Of (source As String, output As String))()
Await Task.Run(Sub()
Parallel.ForEach(filePaths,
Sub(filePath)
Using input As New OcrImageInput(filePath)
Dim result = _ocr.Read(input)
Dim outputPath As String = Path.ChangeExtension(filePath, ".searchable.pdf")
result.SaveAsSearchablePdf(outputPath)
results.Add((filePath, outputPath))
End Using
End Sub)
End Function)
Console.WriteLine($"Processed {results.Count} files")
End Function
End Class
var advancedOcr = new IronTesseract{Configuration = {RenderSearchablePdf = true,TesseractVariables = new Dictionary<string, object> { { "preserve_interword_spaces", 1 }, { "tessedit_char_whitelist", "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz" } },PageSegmentationMode = TesseractPageSegmentationMode.SingleColumn },Language = OcrLanguage.EnglishBest};
var advancedOcr = new IronTesseract
{
Configuration =
{
RenderSearchablePdf = true,
TesseractVariables = new Dictionary<string, object>
{
{ "preserve_interword_spaces", 1 },
{ "tessedit_char_whitelist", "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz" }
},
PageSegmentationMode = TesseractPageSegmentationMode.SingleColumn
},
Language = OcrLanguage.EnglishBest
};
ImportsIronOcrDim advancedOcr As New IronTesseractWith { .Configuration = New TesseractConfigurationWith { .RenderSearchablePdf = True, .TesseractVariables = New Dictionary(OfString, Object) From { {"preserve_interword_spaces", 1}, {"tessedit_char_whitelist", "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"} }, .PageSegmentationMode = TesseractPageSegmentationMode.SingleColumn }, .Language = OcrLanguage.EnglishBest}
Imports IronOcr
Dim advancedOcr As New IronTesseract With {
.Configuration = New TesseractConfiguration With {
.RenderSearchablePdf = True,
.TesseractVariables = New Dictionary(Of String, Object) From {
{"preserve_interword_spaces", 1},
{"tessedit_char_whitelist", "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"}
},
.PageSegmentationMode = TesseractPageSegmentationMode.SingleColumn
},
.Language = OcrLanguage.EnglishBest
}
这些配置选项同样适用于所有三种输出方法:SaveAsSearchablePdfStream。 下方的摘要汇总了所有可搜索的 PDF 处理方法及其对应的输出格式。
摘要
using IronOCR 创建可搜索 PDF 既简单又灵活。 无论是需要通过ReadDocumentAdvanced进行高级文档扫描,该库均提供强大的方法来生成各种格式的可搜索PDF。 使用ModelType参数选择标准和增强ML模型之间的准确性。 以文件、字节或流的形式导出的功能使其能够适应从桌面应用程序到基于云的服务等任何应用程序架构。
using IronOCR 创建可搜索 PDF 只需一行代码即可完成: new IronOcr.IronTesseract { Configuration = { RenderSearchablePdf = true }.}.Read(new IronOcr.OcrImageInput("file.jpg")).SaveAsSearchablePdf("searchable.pdf").这展示了 IronOCR 简化的 API 设计。
在可搜索 PDF 中,不可见文本层是如何工作的?
IronOCR 会自动将识别出的文本作为不可见图层定位在 PDF 原始图像之上。这确保了文本与图像的精确映射,使用户能够在保持原始文档视觉外观的同时,对文本进行选择和搜索。IronOCR库通过专用字体和定位算法来实现这一功能。
我能将照片或屏幕截图转换为可搜索的 PDF 文件吗?
是的,ReadPhoto、ReadScreenShot 和 ReadDocumentAdvanced 的处理结果均支持 SaveAsSearchablePdf 功能。每个方法返回的结果类型均支持可搜索 PDF 导出,从而能够轻松地将真实照片、屏幕截图或复杂的文档扫描件转换为可搜索的 PDF 文件。
出现这种情况是因为可搜索文本层使用的默认字体(Times New Roman)无法完全支持所有 Unicode 字符。要解决此问题,请将兼容 Unicode 的字体文件作为 SaveAsSearchablePdf 的第三个参数传入。如果您的文档最初使用 Times New Roman 排版,且发现与其他字体存在间距不一致的情况,请尝试使用 Liberation Serif,因为它具有相同的字形度量,并能保留原始布局。
What advanced configuration options are available in IronOCR for creating searchable PDFs?
IronOCR offers advanced configuration options, including detailed Tesseract configuration, setting Tesseract variables, and choosing different page segmentation modes. These can be tailored to fine-tune OCR for specific document types or languages.