IronOCR 使您只需一行代码就能用 C# 从 PDF 文件中提取文本,支持所有 PDF 版本,并通过其基于 Tesseract 的引擎提供准确的 OCR 结果。
PDF 是 "便携式文档格式 "的缩写。它是由 Adobe 公司开发的一种文件格式,可以保留任何源文件的字体、图像、图形和布局,而与创建这些文件时使用的应用程序和平台无关。 PDF 文件通常用于以一致的格式共享和查看文档,而无需考虑打开它们的软件或硬件。 IronOCR 可处理各种版本的 PDF 文档,从较早的 PDF 1.0 规范到最新的 PDF 2.0 标准。
快速入门:在几秒内OCR一个PDF文件
通过构建一个指向您的 PDF 的 OcrPdfInput 并调用 Read 快速配置 OCR。 本例演示了使用 IronOCR 从 PDF 中提取文本。
1Install IronOCR with NuGet Package Manager
PM > Install-Package IronOcr
Install-Package IronOcr
2复制并运行这段代码。
using var result = new IronOcr.IronTesseract().Read(new IronOcr.OcrPdfInput("document.pdf", PdfContents.TextAndImages));
using var result = new IronOcr.IronTesseract().Read(new IronOcr.OcrPdfInput("document.pdf", PdfContents.TextAndImages));
/* :path=/static-assets/ocr/content-code-examples/how-to/input-pdfs-read-pdf.cs */using IronOcr;// Instantiate IronTesseractIronTesseract ocrTesseract = new IronTesseract();// Add PDFusing var pdfInput = new OcrPdfInput("Potter.pdf");// Perform OCROcrResult ocrResult = ocrTesseract.Read(pdfInput);// Access the extracted textstring extractedText = ocrResult.Text;System.Console.WriteLine(extractedText);
/* :path=/static-assets/ocr/content-code-examples/how-to/input-pdfs-read-pdf.cs */
using IronOcr;
// Instantiate IronTesseract
IronTesseract ocrTesseract = new IronTesseract();
// Add PDF
using var pdfInput = new OcrPdfInput("Potter.pdf");
// Perform OCR
OcrResult ocrResult = ocrTesseract.Read(pdfInput);
// Access the extracted text
string extractedText = ocrResult.Text;
System.Console.WriteLine(extractedText);
ImportsIronOcr' Instantiate IronTesseractDim ocrTesseract As New IronTesseract()' Add PDFUsing pdfInput As New OcrPdfInput("Potter.pdf") ' Perform OCR Dim ocrResult AsOcrResult = ocrTesseract.Read(pdfInput) ' Access the extracted text Dim extractedText AsString = ocrResult.TextSystem.Console.WriteLine(extractedText)EndUsing
Imports IronOcr
' Instantiate IronTesseract
Dim ocrTesseract As New IronTesseract()
' Add PDF
Using pdfInput As New OcrPdfInput("Potter.pdf")
' Perform OCR
Dim ocrResult As OcrResult = ocrTesseract.Read(pdfInput)
' Access the extracted text
Dim extractedText As String = ocrResult.Text
System.Console.WriteLine(extractedText)
End Using
在处理 PDF 时,您可以通过指定内容类型来优化性能。 枚举 PdfContents 允许您针对特定内容:
// For text-only PDFs (faster processing)var textOnlyPdf = new OcrPdfInput("document.pdf", PdfContents.Text);// For image-only PDFs (scanned documents)var imageOnlyPdf = new OcrPdfInput("scanned.pdf", PdfContents.Images);// For mixed content (default)var mixedPdf = new OcrPdfInput("mixed.pdf", PdfContents.TextAndImages);
// For text-only PDFs (faster processing)
var textOnlyPdf = new OcrPdfInput("document.pdf", PdfContents.Text);
// For image-only PDFs (scanned documents)
var imageOnlyPdf = new OcrPdfInput("scanned.pdf", PdfContents.Images);
// For mixed content (default)
var mixedPdf = new OcrPdfInput("mixed.pdf", PdfContents.TextAndImages);
' For text-only PDFs (faster processing)Dim textOnlyPdf = New OcrPdfInput("document.pdf", PdfContents.Text)' For image-only PDFs (scanned documents)Dim imageOnlyPdf = New OcrPdfInput("scanned.pdf", PdfContents.Images)' For mixed content (default)Dim mixedPdf = New OcrPdfInput("mixed.pdf", PdfContents.TextAndImages)
' For text-only PDFs (faster processing)
Dim textOnlyPdf = New OcrPdfInput("document.pdf", PdfContents.Text)
' For image-only PDFs (scanned documents)
Dim imageOnlyPdf = New OcrPdfInput("scanned.pdf", PdfContents.Images)
' For mixed content (default)
Dim mixedPdf = New OcrPdfInput("mixed.pdf", PdfContents.TextAndImages)
如何阅读 PDF 中的特定页面?
从 PDF 文档中读取特定页面时,请指定导入的页面索引号。 为此,在构建 OcrPdfInput 时将页面索引列表传递给 PageIndices 参数。 请注意,页面索引采用从零开始的编号。 在处理只有某些页面包含相关信息的大型文档时,该功能尤其有用。
using IronOcr;using System.Collections.Generic;// Instantiate IronTesseractIronTesseract ocrTesseract = new IronTesseract();// Create page indices listList<int> pageIndices = new List<int>() { 0, 2 };// Add PDFusing var pdfInput = new OcrPdfInput("Potter.pdf", PageIndices: pageIndices);// Perform OCROcrResult ocrResult = ocrTesseract.Read(pdfInput);
using IronOcr;
using System.Collections.Generic;
// Instantiate IronTesseract
IronTesseract ocrTesseract = new IronTesseract();
// Create page indices list
List<int> pageIndices = new List<int>() { 0, 2 };
// Add PDF
using var pdfInput = new OcrPdfInput("Potter.pdf", PageIndices: pageIndices);
// Perform OCR
OcrResult ocrResult = ocrTesseract.Read(pdfInput);
ImportsIronOcrImportsSystem.Collections.Generic' Instantiate IronTesseractPrivate ocrTesseract As New IronTesseract()' Create page indices listPrivate pageIndices As New List(OfInteger)() From {0, 2}' Add PDFPrivate pdfInput = New OcrPdfInput("Potter.pdf", PageIndices:= pageIndices)' Perform OCRPrivate ocrResult AsOcrResult = ocrTesseract.Read(pdfInput)
Imports IronOcr
Imports System.Collections.Generic
' Instantiate IronTesseract
Private ocrTesseract As New IronTesseract()
' Create page indices list
Private pageIndices As New List(Of Integer)() From {0, 2}
' Add PDF
Private pdfInput = New OcrPdfInput("Potter.pdf", PageIndices:= pageIndices)
' Perform OCR
Private ocrResult As OcrResult = ocrTesseract.Read(pdfInput)
using IronOCR 可以直接阅读非连续页面。 只需将所需的页面索引添加到您的列表中,顺序不限。 例如:
// Read pages 1, 3, 5, and 10 (using zero-based indices)List<int> pageIndices = new List<int>() { 0, 2, 4, 9 };// Or use LINQ for range-based selectionvar evenPages = Enumerable.Range(0, 10).Where(x => x % 2 == 0).ToList();
// Read pages 1, 3, 5, and 10 (using zero-based indices)
List<int> pageIndices = new List<int>() { 0, 2, 4, 9 };
// Or use LINQ for range-based selection
var evenPages = Enumerable.Range(0, 10).Where(x => x % 2 == 0).ToList();
ImportsSystem.Collections.GenericImportsSystem.Linq' Read pages 1, 3, 5, and 10 (using zero-based indices)Dim pageIndices As New List(OfInteger)() From {0, 2, 4, 9}' Or use LINQ for range-based selectionDim evenPages = Enumerable.Range(0, 10).Where(Function(x) x Mod2 = 0).ToList()
Imports System.Collections.Generic
Imports System.Linq
' Read pages 1, 3, 5, and 10 (using zero-based indices)
Dim pageIndices As New List(Of Integer)() From {0, 2, 4, 9}
' Or use LINQ for range-based selection
Dim evenPages = Enumerable.Range(0, 10).Where(Function(x) x Mod 2 = 0).ToList()
OCR 引擎将只处理指定的页面,从而显著提高大型文档的性能。
如果指定了无效的页码会怎样?
如果您指定的页面索引超过了文档的页数,IronOCR 将抛出异常。 在处理之前实施错误处理或验证页面计数。 您可以在执行 OCR 之前检查 PDF 的总页数,以确保您的索引有效。
如何 OCR PDF 的特定区域?
通过缩小阅读范围,可以显著提高阅读效率。 为此,请指定导入 PDF 中需要阅读的精确区域。 在下面的代码示例中,IronOCR 只专注于提取章节编号和标题。 这种技术类似于为图像定义 OCR 区域,可以提高速度和准确性。
using IronOcr;using IronSoftware.Drawing;using System;// Instantiate IronTesseractIronTesseract ocrTesseract = new IronTesseract();// Specify crop regionsRectangle[] scanRegions = { new Rectangle(550, 100, 600, 300) };// Add PDFusing (var pdfInput = new OcrPdfInput("Potter.pdf", ContentAreas: scanRegions)){ // Perform OCR OcrResult ocrResult = ocrTesseract.Read(pdfInput); // Output the result to consoleConsole.WriteLine(ocrResult.Text);}
using IronOcr;
using IronSoftware.Drawing;
using System;
// Instantiate IronTesseract
IronTesseract ocrTesseract = new IronTesseract();
// Specify crop regions
Rectangle[] scanRegions = { new Rectangle(550, 100, 600, 300) };
// Add PDF
using (var pdfInput = new OcrPdfInput("Potter.pdf", ContentAreas: scanRegions))
{
// Perform OCR
OcrResult ocrResult = ocrTesseract.Read(pdfInput);
// Output the result to console
Console.WriteLine(ocrResult.Text);
}
ImportsIronOcrImportsIronSoftware.DrawingImportsSystem' Instantiate IronTesseractPrivate ocrTesseract As New IronTesseract()' Specify crop regionsPrivate scanRegions() AsRectangle = { New Rectangle(550, 100, 600, 300) }' Add PDFUsing pdfInput = New OcrPdfInput("Potter.pdf", ContentAreas:= scanRegions) ' Perform OCR Dim ocrResult AsOcrResult = ocrTesseract.Read(pdfInput) ' Output the result to consoleConsole.WriteLine(ocrResult.Text)EndUsing
Imports IronOcr
Imports IronSoftware.Drawing
Imports System
' Instantiate IronTesseract
Private ocrTesseract As New IronTesseract()
' Specify crop regions
Private scanRegions() As Rectangle = { New Rectangle(550, 100, 600, 300) }
' Add PDF
Using pdfInput = New OcrPdfInput("Potter.pdf", ContentAreas:= scanRegions)
' Perform OCR
Dim ocrResult As OcrResult = ocrTesseract.Read(pdfInput)
' Output the result to console
Console.WriteLine(ocrResult.Text)
End Using
如何确定正确的矩形坐标?
找到正确的坐标需要了解 PDF 的坐标系。 Rectangle 构造器接受四个参数:Width 和 Height。 所有测量值均以像素为单位。 带有标尺功能的 PDF 查看器或调试实用程序等工具可以帮助确定准确的坐标。 此外,还可以通过小幅调整反复试验来完善您的选择区域。
Rectangle[] scanRegions = { new Rectangle(50, 50, 200, 100), // Header region new Rectangle(50, 200, 500, 300), // Main content region new Rectangle(50, 550, 200, 50) // Footer region};
Rectangle[] scanRegions = {
new Rectangle(50, 50, 200, 100), // Header region
new Rectangle(50, 200, 500, 300), // Main content region
new Rectangle(50, 550, 200, 50) // Footer region
};
ImportsSystem.DrawingDim scanRegions AsRectangle() = { New Rectangle(50, 50, 200, 100), ' Header region New Rectangle(50, 200, 500, 300), ' Main content region New Rectangle(50, 550, 200, 50) ' Footer region}
Imports System.Drawing
Dim scanRegions As Rectangle() = {
New Rectangle(50, 50, 200, 100), ' Header region
New Rectangle(50, 200, 500, 300), ' Main content region
New Rectangle(50, 550, 200, 50) ' Footer region
}
自定义配置:针对特定文档类型微调OCR设置
4.导出选项:将结果保存为各种格式,包括可搜索的 PDF 和 hOCR HTML
这些功能使 IronOCR 成为满足企业级 PDF 处理要求的全面解决方案。
常见问题解答
如何用 C# 从 PDF 文件中提取文本?
只需一行代码,您就可以使用 IronOCR 从 PDF 文件中提取文本。只需创建一个 IronTesseract 实例,然后使用 OcrPdfInput 的读取方法即可:`using var result = new IronOcr.IronTesseract().Read(new IronOcr.OcrPdfInput("document.pdf", PdfContents.TextAndImages));`.IronOCR 可处理扫描的 PDF(基于图像)和可搜索的 PDF(基于文本)。
哪些 PDF 版本支持文本提取?
IronOCR 支持所有 PDF 版本,从较早的 PDF 1.0 规范到最新的 PDF 2.0 标准。OCR 引擎基于 Tesseract 技术构建,无论您使用的是哪种 PDF 版本,都能确保文本提取的准确性。
我可以只阅读 PDF 中的特定页面而不是整个文档吗?
是的,IronOCR 允许您通过提供页面索引来读取 PDF 中的特定页面。您可以使用 OcrPdfInput 对象指定要从哪些页面提取文本,而不是处理整个文档,从而提高 OCR 处理大型文档的效率。
在 PDF 文件上进行 OCR 的最基本工作流程是什么?
IronOCR 的最小工作流程包括 5 个步骤:1)下载 C# 库;2)准备 PDF 文档;3)使用 PDF 文件路径创建 OcrPdfInput 对象;4)使用读取方法执行 OCR;5)可选择指定页面索引进行选择性读取。
是的,IronOCR 可以有效处理扫描 PDF(基于图像)和可搜索 PDF(基于文本)。基于 Tesseract 的引擎可自动处理不同的 PDF 类型,使其成为从各种 PDF 格式中提取文本的多功能工具,而无需采用不同的方法。
What are the advantages of using region-specific OCR over full-page OCR?
Region-specific OCR improves performance by focusing on relevant areas, enhancing accuracy and reducing noise from irrelevant content. It is especially useful for forms and structured documents.
What advanced OCR features does IronOCR offer for PDFs?
IronOCR provides features such as creating searchable PDFs, multithreading for faster processing, image preprocessing, and support for multiple languages, catering to complex document processing needs.
How do you handle invalid page numbers in IronOCR?
If invalid page numbers are specified in IronOCR, it will throw an exception. It is recommended to implement error handling or verify page counts prior to processing.
Can IronOCR process non-consecutive pages from a PDF?
Yes, IronOCR can handle non-consecutive pages by specifying the desired page indices in a list. This allows for selective processing of pages in any order.