使用LibTiff.NET處理超過2 GB的TIFF文件的OCR

This article was translated from English: Does it need improvement?
Translated
View the article in English

IronOCR將每個圖像載入到一個以32位整數索引的記憶體AnyBitmap緩衝區中,因此無論系統記憶體有多少,最多只能達到大約2 GB。 大於該大小的TIFF無法載入。 使用BitMiracle.LibTiff.NET將過大的文件拆分為小於2 GB的塊,並對每個塊進行OCR。

這是一個完全管理的Magick.NET替代方案:它在原始條帶或磚級別複製頁面,無需解碼或重新編碼步驟。

請注意2 GB的限制是AnyBitmap的架構限制,不是操作系統或記憶體問題。 IronOCR中的原生每頁TIFF流媒體尚不可用。)]

解決方案

在開始之前,確保您已在.NET項目中有IronOCR (IronTesseract)。 以下使用的LibTiff型別位於BitMiracle.LibTiff.Classic命名空間中。

1. 新增LibTiff.NET套件

dotnet add package BitMiracle.LibTiff.NET
dotnet add package BitMiracle.LibTiff.NET
SHELL

2. 新增TiffPageSplitter助手

這個助手在原始條帶或磚級別複製每個頁面,因此像素資料和壓縮會被完全保留。 它一次一個地串流頁面,保持峰值記憶體大約為一個塊,而不是整個文件,並產生每個保持在大小限制內的多頁塊。

/// <summary>
/// Splits a multi-page TIFF into single-page TIFF byte streams without ever
/// holding the whole file in memory. Each page is copied at the raw
/// (still-encoded) strip/tile level, so pixel data and compression are
/// preserved exactly - there is no decode/re-encode step.
///
/// This is the chunking step only. It produces sub-2 GB single-page byte
/// arrays; feeding them to IronOCR (which is where the AnyBitmap 2 GB
/// single-buffer limit lives) is the consumer's job - see TiffOcrExample.
/// </summary>
public static class TiffPageSplitter
{
    // Tags that describe how a page's raw strip/tile data is encoded.
    // With a raw copy nothing is re-encoded, so every one of these must be
    // carried over verbatim or the copied bytes become uninterpretable.
    // Extend this list if your TIFFs carry tags not covered here
    // (e.g. ICC profiles, EXTRASAMPLES for alpha channels).
    private static readonly TiffTag[] ScalarIntTags =
    {
        TiffTag.IMAGEWIDTH,
        TiffTag.IMAGELENGTH,
        TiffTag.BITSPERSAMPLE,
        TiffTag.SAMPLESPERPIXEL,
        TiffTag.COMPRESSION,
        TiffTag.PHOTOMETRIC,
        TiffTag.FILLORDER,
        TiffTag.PLANARCONFIG,
        TiffTag.ORIENTATION,
        TiffTag.RESOLUTIONUNIT,
        TiffTag.PREDICTOR,      // required for LZW / Deflate raw copies
        TiffTag.SAMPLEFORMAT,
        TiffTag.T4OPTIONS,      // CCITT Group 3
        TiffTag.T6OPTIONS,      // CCITT Group 4
        TiffTag.SUBFILETYPE,
    };

    private static readonly TiffTag[] ScalarDoubleTags =
    {
        TiffTag.XRESOLUTION,    // DPI directly affects OCR accuracy
        TiffTag.YRESOLUTION,
    };

    /// <summary>
    /// Lazily yields each page of <paramref name="inputPath"/> as a standalone
    /// single-page TIFF. The source file stays open for the lifetime of the
    /// enumeration and only one page is materialised at a time, so peak memory
    /// is roughly one page rather than the whole file.
    /// </summary>
    public static IEnumerable<byte[]> SplitTiffToPages(string inputPath)
    {
        using (Tiff input = Tiff.Open(inputPath, "r"))
        {
            if (input == null)
                throw new InvalidOperationException($"Could not open TIFF: {inputPath}");

            int pageCount = input.NumberOfDirectories();
            for (int page = 0; page < pageCount; page++)
            {
                input.SetDirectory((short)page);
                yield return ExtractCurrentPage(input);
            }
        }
    }

    private static byte[] ExtractCurrentPage(Tiff input)
    {
        using (var ms = new MemoryStream())
        {
            // Default TiffStream operates on the MemoryStream passed as clientData.
            using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
            {
                if (output == null)
                    throw new InvalidOperationException("Could not create in-memory TIFF.");

                CopyTags(input, output);

                if (input.IsTiled())
                    CopyRawTiles(input, output);
                else
                    CopyRawStrips(input, output);

                output.WriteDirectory();
            }

            return ms.ToArray();
        }
    }

    private static void CopyTags(Tiff input, Tiff output)
    {
        foreach (TiffTag tag in ScalarIntTags)
        {
            FieldValue[] v = input.GetField(tag);
            if (v != null && v.Length > 0)
                output.SetField(tag, v[0].ToInt());
        }
        foreach (TiffTag tag in ScalarDoubleTags)
        {
            FieldValue[] v = input.GetField(tag);
            if (v != null && v.Length > 0)
                output.SetField(tag, v[0].ToDouble());
        }
        // Strip vs tile layout must match the raw data exactly, otherwise the
        // raw bytes won't line up with the declared boundaries.
        if (input.IsTiled())
        {
            output.SetField(TiffTag.TILEWIDTH, input.GetField(TiffTag.TILEWIDTH)[0].ToInt());
            output.SetField(TiffTag.TILELENGTH, input.GetField(TiffTag.TILELENGTH)[0].ToInt());
        }
        else
        {
            FieldValue[] rps = input.GetField(TiffTag.ROWSPERSTRIP);
            if (rps != null && rps.Length > 0)
                output.SetField(TiffTag.ROWSPERSTRIP, rps[0].ToInt());
        }

        // Palette images: the colour map is required to interpret pixel indices.
        FieldValue[] cmap = input.GetField(TiffTag.COLORMAP);
        if (cmap != null && cmap.Length >= 3)
            output.SetField(TiffTag.COLORMAP,
                cmap[0].ToShortArray(), cmap[1].ToShortArray(), cmap[2].ToShortArray());
    }

    private static void CopyRawStrips(Tiff input, Tiff output)
    {
        int stripCount = input.NumberOfStrips();
        int[] byteCounts = input.GetField(TiffTag.STRIPBYTECOUNTS)[0].ToIntArray();
        for (int strip = 0; strip < stripCount; strip++)
        {
            byte[] buffer = new byte[byteCounts[strip]];
            int read = input.ReadRawStrip(strip, buffer, 0, buffer.Length);
            output.WriteRawStrip(strip, buffer, read);
        }
    }

    private static void CopyRawTiles(Tiff input, Tiff output)
    {
        int tileCount = input.NumberOfTiles();
        int[] byteCounts = input.GetField(TiffTag.TILEBYTECOUNTS)[0].ToIntArray();
        for (int tile = 0; tile < tileCount; tile++)
        {
            byte[] buffer = new byte[byteCounts[tile]];
            int read = input.ReadRawTile(tile, buffer, 0, buffer.Length);
            output.WriteRawTile(tile, buffer, read);
        }
    }

    /// <summary>
    /// Lazily yields multi-page TIFF chunks (the equivalent of the old
    /// Magick.NET 100-pages-per-chunk approach). A new chunk is started when
    /// adding the next page would push the chunk past
    /// <paramref name="maxChunkBytes"/>, or when <paramref name="maxPagesPerChunk"/>
    /// is reached - whichever comes first.
    ///
    /// Size is the real guard: page count alone can exceed the 2 GB AnyBitmap
    /// limit on large pages. The byte total here is the encoded (compressed)
    /// size, which is a cheap proxy - validate the cap against your actual
    /// pages, since decoded size can be much larger than encoded.
    /// </summary>
    public static IEnumerable<byte[]> SplitTiffToChunks(
        string inputPath,
        int maxPagesPerChunk = 100,
        long maxChunkBytes = 1_500_000_000L)
    {
        using (Tiff input = Tiff.Open(inputPath, "r"))
        {
            if (input == null)
                throw new InvalidOperationException($"Could not open TIFF: {inputPath}");
            int pageCount = input.NumberOfDirectories();
            int page = 0;
            while (page < pageCount)
            {
                using (var ms = new MemoryStream())
                {
                    using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
                    {
                        if (output == null)
                            throw new InvalidOperationException("Could not create in-memory TIFF.");

                        int pagesInChunk = 0;
                        long chunkBytes = 0;

                        while (page < pageCount && pagesInChunk < maxPagesPerChunk)
                        {
                            input.SetDirectory((short)page);
                            long pageBytes = RawPageByteSize(input);
                            // Stop before exceeding the cap, but always allow at
                            // least one page so a single large page still goes through.
                            if (pagesInChunk > 0 && chunkBytes + pageBytes > maxChunkBytes)
                                break;
                            CopyTags(input, output);
                            if (input.IsTiled())
                                CopyRawTiles(input, output);
                            else
                                CopyRawStrips(input, output);

                            output.WriteDirectory(); // finalise this page as one directory in the chunk
                            chunkBytes += pageBytes;
                            pagesInChunk++;
                            page++;
                        }
                    }

                    yield return ms.ToArray();
                }
            }
        }
    }

    private static long RawPageByteSize(Tiff page)
    {
        TiffTag tag = page.IsTiled() ? TiffTag.TILEBYTECOUNTS : TiffTag.STRIPBYTECOUNTS;
        int[] counts = page.GetField(tag)[0].ToIntArray();
        long total = 0;
        foreach (int c in counts)
            total += c;
        return total;
    }
}
/// <summary>
/// Splits a multi-page TIFF into single-page TIFF byte streams without ever
/// holding the whole file in memory. Each page is copied at the raw
/// (still-encoded) strip/tile level, so pixel data and compression are
/// preserved exactly - there is no decode/re-encode step.
///
/// This is the chunking step only. It produces sub-2 GB single-page byte
/// arrays; feeding them to IronOCR (which is where the AnyBitmap 2 GB
/// single-buffer limit lives) is the consumer's job - see TiffOcrExample.
/// </summary>
public static class TiffPageSplitter
{
    // Tags that describe how a page's raw strip/tile data is encoded.
    // With a raw copy nothing is re-encoded, so every one of these must be
    // carried over verbatim or the copied bytes become uninterpretable.
    // Extend this list if your TIFFs carry tags not covered here
    // (e.g. ICC profiles, EXTRASAMPLES for alpha channels).
    private static readonly TiffTag[] ScalarIntTags =
    {
        TiffTag.IMAGEWIDTH,
        TiffTag.IMAGELENGTH,
        TiffTag.BITSPERSAMPLE,
        TiffTag.SAMPLESPERPIXEL,
        TiffTag.COMPRESSION,
        TiffTag.PHOTOMETRIC,
        TiffTag.FILLORDER,
        TiffTag.PLANARCONFIG,
        TiffTag.ORIENTATION,
        TiffTag.RESOLUTIONUNIT,
        TiffTag.PREDICTOR,      // required for LZW / Deflate raw copies
        TiffTag.SAMPLEFORMAT,
        TiffTag.T4OPTIONS,      // CCITT Group 3
        TiffTag.T6OPTIONS,      // CCITT Group 4
        TiffTag.SUBFILETYPE,
    };

    private static readonly TiffTag[] ScalarDoubleTags =
    {
        TiffTag.XRESOLUTION,    // DPI directly affects OCR accuracy
        TiffTag.YRESOLUTION,
    };

    /// <summary>
    /// Lazily yields each page of <paramref name="inputPath"/> as a standalone
    /// single-page TIFF. The source file stays open for the lifetime of the
    /// enumeration and only one page is materialised at a time, so peak memory
    /// is roughly one page rather than the whole file.
    /// </summary>
    public static IEnumerable<byte[]> SplitTiffToPages(string inputPath)
    {
        using (Tiff input = Tiff.Open(inputPath, "r"))
        {
            if (input == null)
                throw new InvalidOperationException($"Could not open TIFF: {inputPath}");

            int pageCount = input.NumberOfDirectories();
            for (int page = 0; page < pageCount; page++)
            {
                input.SetDirectory((short)page);
                yield return ExtractCurrentPage(input);
            }
        }
    }

    private static byte[] ExtractCurrentPage(Tiff input)
    {
        using (var ms = new MemoryStream())
        {
            // Default TiffStream operates on the MemoryStream passed as clientData.
            using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
            {
                if (output == null)
                    throw new InvalidOperationException("Could not create in-memory TIFF.");

                CopyTags(input, output);

                if (input.IsTiled())
                    CopyRawTiles(input, output);
                else
                    CopyRawStrips(input, output);

                output.WriteDirectory();
            }

            return ms.ToArray();
        }
    }

    private static void CopyTags(Tiff input, Tiff output)
    {
        foreach (TiffTag tag in ScalarIntTags)
        {
            FieldValue[] v = input.GetField(tag);
            if (v != null && v.Length > 0)
                output.SetField(tag, v[0].ToInt());
        }
        foreach (TiffTag tag in ScalarDoubleTags)
        {
            FieldValue[] v = input.GetField(tag);
            if (v != null && v.Length > 0)
                output.SetField(tag, v[0].ToDouble());
        }
        // Strip vs tile layout must match the raw data exactly, otherwise the
        // raw bytes won't line up with the declared boundaries.
        if (input.IsTiled())
        {
            output.SetField(TiffTag.TILEWIDTH, input.GetField(TiffTag.TILEWIDTH)[0].ToInt());
            output.SetField(TiffTag.TILELENGTH, input.GetField(TiffTag.TILELENGTH)[0].ToInt());
        }
        else
        {
            FieldValue[] rps = input.GetField(TiffTag.ROWSPERSTRIP);
            if (rps != null && rps.Length > 0)
                output.SetField(TiffTag.ROWSPERSTRIP, rps[0].ToInt());
        }

        // Palette images: the colour map is required to interpret pixel indices.
        FieldValue[] cmap = input.GetField(TiffTag.COLORMAP);
        if (cmap != null && cmap.Length >= 3)
            output.SetField(TiffTag.COLORMAP,
                cmap[0].ToShortArray(), cmap[1].ToShortArray(), cmap[2].ToShortArray());
    }

    private static void CopyRawStrips(Tiff input, Tiff output)
    {
        int stripCount = input.NumberOfStrips();
        int[] byteCounts = input.GetField(TiffTag.STRIPBYTECOUNTS)[0].ToIntArray();
        for (int strip = 0; strip < stripCount; strip++)
        {
            byte[] buffer = new byte[byteCounts[strip]];
            int read = input.ReadRawStrip(strip, buffer, 0, buffer.Length);
            output.WriteRawStrip(strip, buffer, read);
        }
    }

    private static void CopyRawTiles(Tiff input, Tiff output)
    {
        int tileCount = input.NumberOfTiles();
        int[] byteCounts = input.GetField(TiffTag.TILEBYTECOUNTS)[0].ToIntArray();
        for (int tile = 0; tile < tileCount; tile++)
        {
            byte[] buffer = new byte[byteCounts[tile]];
            int read = input.ReadRawTile(tile, buffer, 0, buffer.Length);
            output.WriteRawTile(tile, buffer, read);
        }
    }

    /// <summary>
    /// Lazily yields multi-page TIFF chunks (the equivalent of the old
    /// Magick.NET 100-pages-per-chunk approach). A new chunk is started when
    /// adding the next page would push the chunk past
    /// <paramref name="maxChunkBytes"/>, or when <paramref name="maxPagesPerChunk"/>
    /// is reached - whichever comes first.
    ///
    /// Size is the real guard: page count alone can exceed the 2 GB AnyBitmap
    /// limit on large pages. The byte total here is the encoded (compressed)
    /// size, which is a cheap proxy - validate the cap against your actual
    /// pages, since decoded size can be much larger than encoded.
    /// </summary>
    public static IEnumerable<byte[]> SplitTiffToChunks(
        string inputPath,
        int maxPagesPerChunk = 100,
        long maxChunkBytes = 1_500_000_000L)
    {
        using (Tiff input = Tiff.Open(inputPath, "r"))
        {
            if (input == null)
                throw new InvalidOperationException($"Could not open TIFF: {inputPath}");
            int pageCount = input.NumberOfDirectories();
            int page = 0;
            while (page < pageCount)
            {
                using (var ms = new MemoryStream())
                {
                    using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
                    {
                        if (output == null)
                            throw new InvalidOperationException("Could not create in-memory TIFF.");

                        int pagesInChunk = 0;
                        long chunkBytes = 0;

                        while (page < pageCount && pagesInChunk < maxPagesPerChunk)
                        {
                            input.SetDirectory((short)page);
                            long pageBytes = RawPageByteSize(input);
                            // Stop before exceeding the cap, but always allow at
                            // least one page so a single large page still goes through.
                            if (pagesInChunk > 0 && chunkBytes + pageBytes > maxChunkBytes)
                                break;
                            CopyTags(input, output);
                            if (input.IsTiled())
                                CopyRawTiles(input, output);
                            else
                                CopyRawStrips(input, output);

                            output.WriteDirectory(); // finalise this page as one directory in the chunk
                            chunkBytes += pageBytes;
                            pagesInChunk++;
                            page++;
                        }
                    }

                    yield return ms.ToArray();
                }
            }
        }
    }

    private static long RawPageByteSize(Tiff page)
    {
        TiffTag tag = page.IsTiled() ? TiffTag.TILEBYTECOUNTS : TiffTag.STRIPBYTECOUNTS;
        int[] counts = page.GetField(tag)[0].ToIntArray();
        long total = 0;
        foreach (int c in counts)
            total += c;
        return total;
    }
}
Imports System
Imports System.Collections.Generic
Imports System.IO

''' <summary>
''' Splits a multi-page TIFF into single-page TIFF byte streams without ever
''' holding the whole file in memory. Each page is copied at the raw
''' (still-encoded) strip/tile level, so pixel data and compression are
''' preserved exactly - there is no decode/re-encode step.
'''
''' This is the chunking step only. It produces sub-2 GB single-page byte
''' arrays; feeding them to IronOCR (which is where the AnyBitmap 2 GB
''' single-buffer limit lives) is the consumer's job - see TiffOcrExample.
''' </summary>
Public NotInheritable Class TiffPageSplitter

    ' Tags that describe how a page's raw strip/tile data is encoded.
    ' With a raw copy nothing is re-encoded, so every one of these must be
    ' carried over verbatim or the copied bytes become uninterpretable.
    ' Extend this list if your TIFFs carry tags not covered here
    ' (e.g. ICC profiles, EXTRASAMPLES for alpha channels).
    Private Shared ReadOnly ScalarIntTags As TiffTag() = {
        TiffTag.IMAGEWIDTH,
        TiffTag.IMAGELENGTH,
        TiffTag.BITSPERSAMPLE,
        TiffTag.SAMPLESPERPIXEL,
        TiffTag.COMPRESSION,
        TiffTag.PHOTOMETRIC,
        TiffTag.FILLORDER,
        TiffTag.PLANARCONFIG,
        TiffTag.ORIENTATION,
        TiffTag.RESOLUTIONUNIT,
        TiffTag.PREDICTOR,      ' required for LZW / Deflate raw copies
        TiffTag.SAMPLEFORMAT,
        TiffTag.T4OPTIONS,      ' CCITT Group 3
        TiffTag.T6OPTIONS,      ' CCITT Group 4
        TiffTag.SUBFILETYPE
    }

    Private Shared ReadOnly ScalarDoubleTags As TiffTag() = {
        TiffTag.XRESOLUTION,    ' DPI directly affects OCR accuracy
        TiffTag.YRESOLUTION
    }

    ''' <summary>
    ''' Lazily yields each page of <paramref name="inputPath"/> as a standalone
    ''' single-page TIFF. The source file stays open for the lifetime of the
    ''' enumeration and only one page is materialised at a time, so peak memory
    ''' is roughly one page rather than the whole file.
    ''' </summary>
    Public Shared Iterator Function SplitTiffToPages(inputPath As String) As IEnumerable(Of Byte())
        Using input As Tiff = Tiff.Open(inputPath, "r")
            If input Is Nothing Then
                Throw New InvalidOperationException($"Could not open TIFF: {inputPath}")
            End If

            Dim pageCount As Integer = input.NumberOfDirectories()
            For page As Integer = 0 To pageCount - 1
                input.SetDirectory(CShort(page))
                Yield ExtractCurrentPage(input)
            Next
        End Using
    End Function

    Private Shared Function ExtractCurrentPage(input As Tiff) As Byte()
        Using ms As New MemoryStream()
            ' Default TiffStream operates on the MemoryStream passed as clientData.
            Using output As Tiff = Tiff.ClientOpen("InMemory", "w", ms, New TiffStream())
                If output Is Nothing Then
                    Throw New InvalidOperationException("Could not create in-memory TIFF.")
                End If

                CopyTags(input, output)

                If input.IsTiled() Then
                    CopyRawTiles(input, output)
                Else
                    CopyRawStrips(input, output)
                End If

                output.WriteDirectory()
            End Using

            Return ms.ToArray()
        End Using
    End Function

    Private Shared Sub CopyTags(input As Tiff, output As Tiff)
        For Each tag As TiffTag In ScalarIntTags
            Dim v As FieldValue() = input.GetField(tag)
            If v IsNot Nothing AndAlso v.Length > 0 Then
                output.SetField(tag, v(0).ToInt())
            End If
        Next
        For Each tag As TiffTag In ScalarDoubleTags
            Dim v As FieldValue() = input.GetField(tag)
            If v IsNot Nothing AndAlso v.Length > 0 Then
                output.SetField(tag, v(0).ToDouble())
            End If
        Next
        ' Strip vs tile layout must match the raw data exactly, otherwise the
        ' raw bytes won't line up with the declared boundaries.
        If input.IsTiled() Then
            output.SetField(TiffTag.TILEWIDTH, input.GetField(TiffTag.TILEWIDTH)(0).ToInt())
            output.SetField(TiffTag.TILELENGTH, input.GetField(TiffTag.TILELENGTH)(0).ToInt())
        Else
            Dim rps As FieldValue() = input.GetField(TiffTag.ROWSPERSTRIP)
            If rps IsNot Nothing AndAlso rps.Length > 0 Then
                output.SetField(TiffTag.ROWSPERSTRIP, rps(0).ToInt())
            End If
        End If

        ' Palette images: the colour map is required to interpret pixel indices.
        Dim cmap As FieldValue() = input.GetField(TiffTag.COLORMAP)
        If cmap IsNot Nothing AndAlso cmap.Length >= 3 Then
            output.SetField(TiffTag.COLORMAP,
                cmap(0).ToShortArray(), cmap(1).ToShortArray(), cmap(2).ToShortArray())
        End If
    End Sub

    Private Shared Sub CopyRawStrips(input As Tiff, output As Tiff)
        Dim stripCount As Integer = input.NumberOfStrips()
        Dim byteCounts As Integer() = input.GetField(TiffTag.STRIPBYTECOUNTS)(0).ToIntArray()
        For strip As Integer = 0 To stripCount - 1
            Dim buffer As Byte() = New Byte(byteCounts(strip) - 1) {}
            Dim read As Integer = input.ReadRawStrip(strip, buffer, 0, buffer.Length)
            output.WriteRawStrip(strip, buffer, read)
        Next
    End Sub

    Private Shared Sub CopyRawTiles(input As Tiff, output As Tiff)
        Dim tileCount As Integer = input.NumberOfTiles()
        Dim byteCounts As Integer() = input.GetField(TiffTag.TILEBYTECOUNTS)(0).ToIntArray()
        For tile As Integer = 0 To tileCount - 1
            Dim buffer As Byte() = New Byte(byteCounts(tile) - 1) {}
            Dim read As Integer = input.ReadRawTile(tile, buffer, 0, buffer.Length)
            output.WriteRawTile(tile, buffer, read)
        Next
    End Sub

    ''' <summary>
    ''' Lazily yields multi-page TIFF chunks (the equivalent of the old
    ''' Magick.NET 100-pages-per-chunk approach). A new chunk is started when
    ''' adding the next page would push the chunk past
    ''' <paramref name="maxChunkBytes"/>, or when <paramref name="maxPagesPerChunk"/>
    ''' is reached - whichever comes first.
    '''
    ''' Size is the real guard: page count alone can exceed the 2 GB AnyBitmap
    ''' limit on large pages. The byte total here is the encoded (compressed)
    ''' size, which is a cheap proxy - validate the cap against your actual
    ''' pages, since decoded size can be much larger than encoded.
    ''' </summary>
    Public Shared Iterator Function SplitTiffToChunks(
        inputPath As String,
        Optional maxPagesPerChunk As Integer = 100,
        Optional maxChunkBytes As Long = 1500000000L) As IEnumerable(Of Byte())
        Using input As Tiff = Tiff.Open(inputPath, "r")
            If input Is Nothing Then
                Throw New InvalidOperationException($"Could not open TIFF: {inputPath}")
            End If
            Dim pageCount As Integer = input.NumberOfDirectories()
            Dim page As Integer = 0
            While page < pageCount
                Using ms As New MemoryStream()
                    Using output As Tiff = Tiff.ClientOpen("InMemory", "w", ms, New TiffStream())
                        If output Is Nothing Then
                            Throw New InvalidOperationException("Could not create in-memory TIFF.")
                        End If

                        Dim pagesInChunk As Integer = 0
                        Dim chunkBytes As Long = 0

                        While page < pageCount AndAlso pagesInChunk < maxPagesPerChunk
                            input.SetDirectory(CShort(page))
                            Dim pageBytes As Long = RawPageByteSize(input)
                            ' Stop before exceeding the cap, but always allow at
                            ' least one page so a single large page still goes through.
                            If pagesInChunk > 0 AndAlso chunkBytes + pageBytes > maxChunkBytes Then
                                Exit While
                            End If
                            CopyTags(input, output)
                            If input.IsTiled() Then
                                CopyRawTiles(input, output)
                            Else
                                CopyRawStrips(input, output)
                            End If

                            output.WriteDirectory() ' finalise this page as one directory in the chunk
                            chunkBytes += pageBytes
                            pagesInChunk += 1
                            page += 1
                        End While
                    End Using

                    Yield ms.ToArray()
                End Using
            End While
        End Using
    End Function

    Private Shared Function RawPageByteSize(page As Tiff) As Long
        Dim tag As TiffTag = If(page.IsTiled(), TiffTag.TILEBYTECOUNTS, TiffTag.STRIPBYTECOUNTS)
        Dim counts As Integer() = page.GetField(tag)(0).ToIntArray()
        Dim total As Long = 0
        For Each c As Integer In counts
            total += c
        Next
        Return total
    End Function
End Class
$vbLabelText   $csharpLabel

maxPagesPerChunk時啟動新塊,以先發生者為準。單一頁面大於限制時,總是允許通過。

3. 將每個塊進行OCR

迭代AnyBitmap

var inputPath = "2gb_benchmark1200.tiff";
var ocr = new IronTesseract();
int chunk = 0;
foreach (byte[] chunkBytes in TiffPageSplitter.SplitTiffToChunks(inputPath, maxPagesPerChunk: 100))
{
    using (var ocrInput = new OcrInput())
    {
        ocrInput.LoadImage(chunkBytes); // loads every page in the chunk
        var result = ocr.Read(ocrInput);
        Console.WriteLine($"Chunk {chunk}: {result.Text?.Length ?? 0} chars");
    }
    chunk++;
}
var inputPath = "2gb_benchmark1200.tiff";
var ocr = new IronTesseract();
int chunk = 0;
foreach (byte[] chunkBytes in TiffPageSplitter.SplitTiffToChunks(inputPath, maxPagesPerChunk: 100))
{
    using (var ocrInput = new OcrInput())
    {
        ocrInput.LoadImage(chunkBytes); // loads every page in the chunk
        var result = ocr.Read(ocrInput);
        Console.WriteLine($"Chunk {chunk}: {result.Text?.Length ?? 0} chars");
    }
    chunk++;
}
Imports IronOcr

Dim inputPath As String = "2gb_benchmark1200.tiff"
Dim ocr As New IronTesseract()
Dim chunk As Integer = 0

For Each chunkBytes As Byte() In TiffPageSplitter.SplitTiffToChunks(inputPath, maxPagesPerChunk:=100)
    Using ocrInput As New OcrInput()
        ocrInput.LoadImage(chunkBytes) ' loads every page in the chunk
        Dim result = ocr.Read(ocrInput)
        Console.WriteLine($"Chunk {chunk}: {If(result.Text?.Length, 0)} chars")
    End Using
    chunk += 1
Next
$vbLabelText   $csharpLabel

傳遞字節陣列保持每個輸入都在限制內。請注意,OcrInput.LoadImage(filePath)目前在文件過大時返回零載入的頁面,而不是拋出明確的錯誤; 這種默默無聞的失敗是一個已知問題,而塊化完全避免了它。

4. 調整塊限制以適應您的資料

調整maxChunkBytes來匹配您的TIFF。 如果在解碼後塊接近2 GB或者記憶體緊張,則降低它們; 增加較小頁面的頁數以降低開銷。

警告maxChunkBytes根據編碼(壓縮)大小衡量,這只是個廉價的代理。 解碼後的大小可能更大,因此解碼後單一非常大的頁面可能仍然超過2 GB。 預設的1.5 GB限制留有餘地。

注意事項和限制

  • 標籤覆蓋範圍:拆分工具僅繼承在ScalarDoubleTags中列出的標籤。 如果您的TIFF使用未涵蓋的標籤,例如ICC輪廓或用於透明通道的EXTRASAMPLES,請擴展這些列表,否則原始複製的字節可能會被誤解。
  • 管理的依賴:LibTiff.NET是完全管理的,不含原生二進制文件,與Magick.NET不同,後者帶有ImageMagick原生庫,增加了包的大小和部署佈局。
  • 精確保留:原始的條帶或磚級別複製避免了Magick.NET方法中的解碼和重新編碼,保持了原始壓縮和像素資料完好無損。

欲進一步了解,請參見NuGet上的BitMiracle.LibTiff.NET

Curtis Chau
技術作家

Curtis Chau擁有Carleton大學的電腦科學學士學位,專精於前端開發,擁有Node.js、TypeScript、JavaScript和React的專業知識。Curtis熱衷於建立直觀且美觀的使用者介面,喜愛使用現代框架並建立結構良好、視覺吸引力的手冊。

除了開發,Curtis對物聯網(IoT)有濃厚的興趣,探索創新的方法來整合硬體和軟體。在空閒時間,他喜歡玩遊戲和建立Discord機器人,結合他對技術的熱愛與創造力。

準備開始了嗎?
Nuget 下載 6,175,195 | 版本: 2026.7 剛剛發布
Still Scrolling Icon

還在滾動?

想要快速證明? PM > Install-Package IronOcr
執行範例 觀看您的圖像轉變為可搜尋文字。