OCR de TIFFs de más de 2 GB con LibTiff.NET
IronOCR carga cada imagen en un búfer en memoria AnyBitmap indexado por un entero de 32 bits, por lo que se limita a aproximadamente 2 GB independientemente de cuánta memoria del sistema esté disponible. Un TIFF más grande que eso no se carga. Divida el archivo sobredimensionado en secciones de menos de 2 GB con BitMiracle.LibTiff.NET y haga OCR para cada sección.
Esta es una alternativa totalmente gestionada a la solución de Magick.NET: copia páginas a nivel de tira o mosaico en bruto, sin paso de decodificación o recodificación.
Solución
Antes de comenzar, asegúrese de tener IronOCR (IronTesseract) en un proyecto .NET. Los tipos de LibTiff usados a continuación pertenecen al espacio de nombres BitMiracle.LibTiff.Classic.
1. Añadir el paquete LibTiff.NET
dotnet add package BitMiracle.LibTiff.NET
dotnet add package BitMiracle.LibTiff.NET
2. Añadir el ayudante TiffPageSplitter
El ayudante copia cada página a nivel de tira o mosaico en bruto, por lo que los datos de píxeles y la compresión se preservan exactamente. Transmite las páginas una a la vez, manteniendo la memoria máxima aproximadamente a un nivel de sección en lugar de todo el archivo, y genera secciones de varias páginas que permanecen bajo el límite de tamaño.
/// <summary>
/// Splits a multi-page TIFF into single-page TIFF byte streams without ever
/// holding the whole file in memory. Each page is copied at the raw
/// (still-encoded) strip/tile level, so pixel data and compression are
/// preserved exactly - there is no decode/re-encode step.
///
/// This is the chunking step only. It produces sub-2 GB single-page byte
/// arrays; feeding them to IronOCR (which is where the AnyBitmap 2 GB
/// single-buffer limit lives) is the consumer's job - see TiffOcrExample.
/// </summary>
public static class TiffPageSplitter
{
// Tags that describe how a page's raw strip/tile data is encoded.
// With a raw copy nothing is re-encoded, so every one of these must be
// carried over verbatim or the copied bytes become uninterpretable.
// Extend this list if your TIFFs carry tags not covered here
// (e.g. ICC profiles, EXTRASAMPLES for alpha channels).
private static readonly TiffTag[] ScalarIntTags =
{
TiffTag.IMAGEWIDTH,
TiffTag.IMAGELENGTH,
TiffTag.BITSPERSAMPLE,
TiffTag.SAMPLESPERPIXEL,
TiffTag.COMPRESSION,
TiffTag.PHOTOMETRIC,
TiffTag.FILLORDER,
TiffTag.PLANARCONFIG,
TiffTag.ORIENTATION,
TiffTag.RESOLUTIONUNIT,
TiffTag.PREDICTOR, // required for LZW / Deflate raw copies
TiffTag.SAMPLEFORMAT,
TiffTag.T4OPTIONS, // CCITT Group 3
TiffTag.T6OPTIONS, // CCITT Group 4
TiffTag.SUBFILETYPE,
};
private static readonly TiffTag[] ScalarDoubleTags =
{
TiffTag.XRESOLUTION, // DPI directly affects OCR accuracy
TiffTag.YRESOLUTION,
};
/// <summary>
/// Lazily yields each page of <paramref name="inputPath"/> as a standalone
/// single-page TIFF. The source file stays open for the lifetime of the
/// enumeration and only one page is materialised at a time, so peak memory
/// is roughly one page rather than the whole file.
/// </summary>
public static IEnumerable<byte[]> SplitTiffToPages(string inputPath)
{
using (Tiff input = Tiff.Open(inputPath, "r"))
{
if (input == null)
throw new InvalidOperationException($"Could not open TIFF: {inputPath}");
int pageCount = input.NumberOfDirectories();
for (int page = 0; page < pageCount; page++)
{
input.SetDirectory((short)page);
yield return ExtractCurrentPage(input);
}
}
}
private static byte[] ExtractCurrentPage(Tiff input)
{
using (var ms = new MemoryStream())
{
// Default TiffStream operates on the MemoryStream passed as clientData.
using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
{
if (output == null)
throw new InvalidOperationException("Could not create in-memory TIFF.");
CopyTags(input, output);
if (input.IsTiled())
CopyRawTiles(input, output);
else
CopyRawStrips(input, output);
output.WriteDirectory();
}
return ms.ToArray();
}
}
private static void CopyTags(Tiff input, Tiff output)
{
foreach (TiffTag tag in ScalarIntTags)
{
FieldValue[] v = input.GetField(tag);
if (v != null && v.Length > 0)
output.SetField(tag, v[0].ToInt());
}
foreach (TiffTag tag in ScalarDoubleTags)
{
FieldValue[] v = input.GetField(tag);
if (v != null && v.Length > 0)
output.SetField(tag, v[0].ToDouble());
}
// Strip vs tile layout must match the raw data exactly, otherwise the
// raw bytes won't line up with the declared boundaries.
if (input.IsTiled())
{
output.SetField(TiffTag.TILEWIDTH, input.GetField(TiffTag.TILEWIDTH)[0].ToInt());
output.SetField(TiffTag.TILELENGTH, input.GetField(TiffTag.TILELENGTH)[0].ToInt());
}
else
{
FieldValue[] rps = input.GetField(TiffTag.ROWSPERSTRIP);
if (rps != null && rps.Length > 0)
output.SetField(TiffTag.ROWSPERSTRIP, rps[0].ToInt());
}
// Palette images: the colour map is required to interpret pixel indices.
FieldValue[] cmap = input.GetField(TiffTag.COLORMAP);
if (cmap != null && cmap.Length >= 3)
output.SetField(TiffTag.COLORMAP,
cmap[0].ToShortArray(), cmap[1].ToShortArray(), cmap[2].ToShortArray());
}
private static void CopyRawStrips(Tiff input, Tiff output)
{
int stripCount = input.NumberOfStrips();
int[] byteCounts = input.GetField(TiffTag.STRIPBYTECOUNTS)[0].ToIntArray();
for (int strip = 0; strip < stripCount; strip++)
{
byte[] buffer = new byte[byteCounts[strip]];
int read = input.ReadRawStrip(strip, buffer, 0, buffer.Length);
output.WriteRawStrip(strip, buffer, read);
}
}
private static void CopyRawTiles(Tiff input, Tiff output)
{
int tileCount = input.NumberOfTiles();
int[] byteCounts = input.GetField(TiffTag.TILEBYTECOUNTS)[0].ToIntArray();
for (int tile = 0; tile < tileCount; tile++)
{
byte[] buffer = new byte[byteCounts[tile]];
int read = input.ReadRawTile(tile, buffer, 0, buffer.Length);
output.WriteRawTile(tile, buffer, read);
}
}
/// <summary>
/// Lazily yields multi-page TIFF chunks (the equivalent of the old
/// Magick.NET 100-pages-per-chunk approach). A new chunk is started when
/// adding the next page would push the chunk past
/// <paramref name="maxChunkBytes"/>, or when <paramref name="maxPagesPerChunk"/>
/// is reached - whichever comes first.
///
/// Size is the real guard: page count alone can exceed the 2 GB AnyBitmap
/// limit on large pages. The byte total here is the encoded (compressed)
/// size, which is a cheap proxy - validate the cap against your actual
/// pages, since decoded size can be much larger than encoded.
/// </summary>
public static IEnumerable<byte[]> SplitTiffToChunks(
string inputPath,
int maxPagesPerChunk = 100,
long maxChunkBytes = 1_500_000_000L)
{
using (Tiff input = Tiff.Open(inputPath, "r"))
{
if (input == null)
throw new InvalidOperationException($"Could not open TIFF: {inputPath}");
int pageCount = input.NumberOfDirectories();
int page = 0;
while (page < pageCount)
{
using (var ms = new MemoryStream())
{
using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
{
if (output == null)
throw new InvalidOperationException("Could not create in-memory TIFF.");
int pagesInChunk = 0;
long chunkBytes = 0;
while (page < pageCount && pagesInChunk < maxPagesPerChunk)
{
input.SetDirectory((short)page);
long pageBytes = RawPageByteSize(input);
// Stop before exceeding the cap, but always allow at
// least one page so a single large page still goes through.
if (pagesInChunk > 0 && chunkBytes + pageBytes > maxChunkBytes)
break;
CopyTags(input, output);
if (input.IsTiled())
CopyRawTiles(input, output);
else
CopyRawStrips(input, output);
output.WriteDirectory(); // finalise this page as one directory in the chunk
chunkBytes += pageBytes;
pagesInChunk++;
page++;
}
}
yield return ms.ToArray();
}
}
}
}
private static long RawPageByteSize(Tiff page)
{
TiffTag tag = page.IsTiled() ? TiffTag.TILEBYTECOUNTS : TiffTag.STRIPBYTECOUNTS;
int[] counts = page.GetField(tag)[0].ToIntArray();
long total = 0;
foreach (int c in counts)
total += c;
return total;
}
}
/// <summary>
/// Splits a multi-page TIFF into single-page TIFF byte streams without ever
/// holding the whole file in memory. Each page is copied at the raw
/// (still-encoded) strip/tile level, so pixel data and compression are
/// preserved exactly - there is no decode/re-encode step.
///
/// This is the chunking step only. It produces sub-2 GB single-page byte
/// arrays; feeding them to IronOCR (which is where the AnyBitmap 2 GB
/// single-buffer limit lives) is the consumer's job - see TiffOcrExample.
/// </summary>
public static class TiffPageSplitter
{
// Tags that describe how a page's raw strip/tile data is encoded.
// With a raw copy nothing is re-encoded, so every one of these must be
// carried over verbatim or the copied bytes become uninterpretable.
// Extend this list if your TIFFs carry tags not covered here
// (e.g. ICC profiles, EXTRASAMPLES for alpha channels).
private static readonly TiffTag[] ScalarIntTags =
{
TiffTag.IMAGEWIDTH,
TiffTag.IMAGELENGTH,
TiffTag.BITSPERSAMPLE,
TiffTag.SAMPLESPERPIXEL,
TiffTag.COMPRESSION,
TiffTag.PHOTOMETRIC,
TiffTag.FILLORDER,
TiffTag.PLANARCONFIG,
TiffTag.ORIENTATION,
TiffTag.RESOLUTIONUNIT,
TiffTag.PREDICTOR, // required for LZW / Deflate raw copies
TiffTag.SAMPLEFORMAT,
TiffTag.T4OPTIONS, // CCITT Group 3
TiffTag.T6OPTIONS, // CCITT Group 4
TiffTag.SUBFILETYPE,
};
private static readonly TiffTag[] ScalarDoubleTags =
{
TiffTag.XRESOLUTION, // DPI directly affects OCR accuracy
TiffTag.YRESOLUTION,
};
/// <summary>
/// Lazily yields each page of <paramref name="inputPath"/> as a standalone
/// single-page TIFF. The source file stays open for the lifetime of the
/// enumeration and only one page is materialised at a time, so peak memory
/// is roughly one page rather than the whole file.
/// </summary>
public static IEnumerable<byte[]> SplitTiffToPages(string inputPath)
{
using (Tiff input = Tiff.Open(inputPath, "r"))
{
if (input == null)
throw new InvalidOperationException($"Could not open TIFF: {inputPath}");
int pageCount = input.NumberOfDirectories();
for (int page = 0; page < pageCount; page++)
{
input.SetDirectory((short)page);
yield return ExtractCurrentPage(input);
}
}
}
private static byte[] ExtractCurrentPage(Tiff input)
{
using (var ms = new MemoryStream())
{
// Default TiffStream operates on the MemoryStream passed as clientData.
using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
{
if (output == null)
throw new InvalidOperationException("Could not create in-memory TIFF.");
CopyTags(input, output);
if (input.IsTiled())
CopyRawTiles(input, output);
else
CopyRawStrips(input, output);
output.WriteDirectory();
}
return ms.ToArray();
}
}
private static void CopyTags(Tiff input, Tiff output)
{
foreach (TiffTag tag in ScalarIntTags)
{
FieldValue[] v = input.GetField(tag);
if (v != null && v.Length > 0)
output.SetField(tag, v[0].ToInt());
}
foreach (TiffTag tag in ScalarDoubleTags)
{
FieldValue[] v = input.GetField(tag);
if (v != null && v.Length > 0)
output.SetField(tag, v[0].ToDouble());
}
// Strip vs tile layout must match the raw data exactly, otherwise the
// raw bytes won't line up with the declared boundaries.
if (input.IsTiled())
{
output.SetField(TiffTag.TILEWIDTH, input.GetField(TiffTag.TILEWIDTH)[0].ToInt());
output.SetField(TiffTag.TILELENGTH, input.GetField(TiffTag.TILELENGTH)[0].ToInt());
}
else
{
FieldValue[] rps = input.GetField(TiffTag.ROWSPERSTRIP);
if (rps != null && rps.Length > 0)
output.SetField(TiffTag.ROWSPERSTRIP, rps[0].ToInt());
}
// Palette images: the colour map is required to interpret pixel indices.
FieldValue[] cmap = input.GetField(TiffTag.COLORMAP);
if (cmap != null && cmap.Length >= 3)
output.SetField(TiffTag.COLORMAP,
cmap[0].ToShortArray(), cmap[1].ToShortArray(), cmap[2].ToShortArray());
}
private static void CopyRawStrips(Tiff input, Tiff output)
{
int stripCount = input.NumberOfStrips();
int[] byteCounts = input.GetField(TiffTag.STRIPBYTECOUNTS)[0].ToIntArray();
for (int strip = 0; strip < stripCount; strip++)
{
byte[] buffer = new byte[byteCounts[strip]];
int read = input.ReadRawStrip(strip, buffer, 0, buffer.Length);
output.WriteRawStrip(strip, buffer, read);
}
}
private static void CopyRawTiles(Tiff input, Tiff output)
{
int tileCount = input.NumberOfTiles();
int[] byteCounts = input.GetField(TiffTag.TILEBYTECOUNTS)[0].ToIntArray();
for (int tile = 0; tile < tileCount; tile++)
{
byte[] buffer = new byte[byteCounts[tile]];
int read = input.ReadRawTile(tile, buffer, 0, buffer.Length);
output.WriteRawTile(tile, buffer, read);
}
}
/// <summary>
/// Lazily yields multi-page TIFF chunks (the equivalent of the old
/// Magick.NET 100-pages-per-chunk approach). A new chunk is started when
/// adding the next page would push the chunk past
/// <paramref name="maxChunkBytes"/>, or when <paramref name="maxPagesPerChunk"/>
/// is reached - whichever comes first.
///
/// Size is the real guard: page count alone can exceed the 2 GB AnyBitmap
/// limit on large pages. The byte total here is the encoded (compressed)
/// size, which is a cheap proxy - validate the cap against your actual
/// pages, since decoded size can be much larger than encoded.
/// </summary>
public static IEnumerable<byte[]> SplitTiffToChunks(
string inputPath,
int maxPagesPerChunk = 100,
long maxChunkBytes = 1_500_000_000L)
{
using (Tiff input = Tiff.Open(inputPath, "r"))
{
if (input == null)
throw new InvalidOperationException($"Could not open TIFF: {inputPath}");
int pageCount = input.NumberOfDirectories();
int page = 0;
while (page < pageCount)
{
using (var ms = new MemoryStream())
{
using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
{
if (output == null)
throw new InvalidOperationException("Could not create in-memory TIFF.");
int pagesInChunk = 0;
long chunkBytes = 0;
while (page < pageCount && pagesInChunk < maxPagesPerChunk)
{
input.SetDirectory((short)page);
long pageBytes = RawPageByteSize(input);
// Stop before exceeding the cap, but always allow at
// least one page so a single large page still goes through.
if (pagesInChunk > 0 && chunkBytes + pageBytes > maxChunkBytes)
break;
CopyTags(input, output);
if (input.IsTiled())
CopyRawTiles(input, output);
else
CopyRawStrips(input, output);
output.WriteDirectory(); // finalise this page as one directory in the chunk
chunkBytes += pageBytes;
pagesInChunk++;
page++;
}
}
yield return ms.ToArray();
}
}
}
}
private static long RawPageByteSize(Tiff page)
{
TiffTag tag = page.IsTiled() ? TiffTag.TILEBYTECOUNTS : TiffTag.STRIPBYTECOUNTS;
int[] counts = page.GetField(tag)[0].ToIntArray();
long total = 0;
foreach (int c in counts)
total += c;
return total;
}
}
Imports System
Imports System.Collections.Generic
Imports System.IO
''' <summary>
''' Splits a multi-page TIFF into single-page TIFF byte streams without ever
''' holding the whole file in memory. Each page is copied at the raw
''' (still-encoded) strip/tile level, so pixel data and compression are
''' preserved exactly - there is no decode/re-encode step.
'''
''' This is the chunking step only. It produces sub-2 GB single-page byte
''' arrays; feeding them to IronOCR (which is where the AnyBitmap 2 GB
''' single-buffer limit lives) is the consumer's job - see TiffOcrExample.
''' </summary>
Public NotInheritable Class TiffPageSplitter
' Tags that describe how a page's raw strip/tile data is encoded.
' With a raw copy nothing is re-encoded, so every one of these must be
' carried over verbatim or the copied bytes become uninterpretable.
' Extend this list if your TIFFs carry tags not covered here
' (e.g. ICC profiles, EXTRASAMPLES for alpha channels).
Private Shared ReadOnly ScalarIntTags As TiffTag() = {
TiffTag.IMAGEWIDTH,
TiffTag.IMAGELENGTH,
TiffTag.BITSPERSAMPLE,
TiffTag.SAMPLESPERPIXEL,
TiffTag.COMPRESSION,
TiffTag.PHOTOMETRIC,
TiffTag.FILLORDER,
TiffTag.PLANARCONFIG,
TiffTag.ORIENTATION,
TiffTag.RESOLUTIONUNIT,
TiffTag.PREDICTOR, ' required for LZW / Deflate raw copies
TiffTag.SAMPLEFORMAT,
TiffTag.T4OPTIONS, ' CCITT Group 3
TiffTag.T6OPTIONS, ' CCITT Group 4
TiffTag.SUBFILETYPE
}
Private Shared ReadOnly ScalarDoubleTags As TiffTag() = {
TiffTag.XRESOLUTION, ' DPI directly affects OCR accuracy
TiffTag.YRESOLUTION
}
''' <summary>
''' Lazily yields each page of <paramref name="inputPath"/> as a standalone
''' single-page TIFF. The source file stays open for the lifetime of the
''' enumeration and only one page is materialised at a time, so peak memory
''' is roughly one page rather than the whole file.
''' </summary>
Public Shared Iterator Function SplitTiffToPages(inputPath As String) As IEnumerable(Of Byte())
Using input As Tiff = Tiff.Open(inputPath, "r")
If input Is Nothing Then
Throw New InvalidOperationException($"Could not open TIFF: {inputPath}")
End If
Dim pageCount As Integer = input.NumberOfDirectories()
For page As Integer = 0 To pageCount - 1
input.SetDirectory(CShort(page))
Yield ExtractCurrentPage(input)
Next
End Using
End Function
Private Shared Function ExtractCurrentPage(input As Tiff) As Byte()
Using ms As New MemoryStream()
' Default TiffStream operates on the MemoryStream passed as clientData.
Using output As Tiff = Tiff.ClientOpen("InMemory", "w", ms, New TiffStream())
If output Is Nothing Then
Throw New InvalidOperationException("Could not create in-memory TIFF.")
End If
CopyTags(input, output)
If input.IsTiled() Then
CopyRawTiles(input, output)
Else
CopyRawStrips(input, output)
End If
output.WriteDirectory()
End Using
Return ms.ToArray()
End Using
End Function
Private Shared Sub CopyTags(input As Tiff, output As Tiff)
For Each tag As TiffTag In ScalarIntTags
Dim v As FieldValue() = input.GetField(tag)
If v IsNot Nothing AndAlso v.Length > 0 Then
output.SetField(tag, v(0).ToInt())
End If
Next
For Each tag As TiffTag In ScalarDoubleTags
Dim v As FieldValue() = input.GetField(tag)
If v IsNot Nothing AndAlso v.Length > 0 Then
output.SetField(tag, v(0).ToDouble())
End If
Next
' Strip vs tile layout must match the raw data exactly, otherwise the
' raw bytes won't line up with the declared boundaries.
If input.IsTiled() Then
output.SetField(TiffTag.TILEWIDTH, input.GetField(TiffTag.TILEWIDTH)(0).ToInt())
output.SetField(TiffTag.TILELENGTH, input.GetField(TiffTag.TILELENGTH)(0).ToInt())
Else
Dim rps As FieldValue() = input.GetField(TiffTag.ROWSPERSTRIP)
If rps IsNot Nothing AndAlso rps.Length > 0 Then
output.SetField(TiffTag.ROWSPERSTRIP, rps(0).ToInt())
End If
End If
' Palette images: the colour map is required to interpret pixel indices.
Dim cmap As FieldValue() = input.GetField(TiffTag.COLORMAP)
If cmap IsNot Nothing AndAlso cmap.Length >= 3 Then
output.SetField(TiffTag.COLORMAP,
cmap(0).ToShortArray(), cmap(1).ToShortArray(), cmap(2).ToShortArray())
End If
End Sub
Private Shared Sub CopyRawStrips(input As Tiff, output As Tiff)
Dim stripCount As Integer = input.NumberOfStrips()
Dim byteCounts As Integer() = input.GetField(TiffTag.STRIPBYTECOUNTS)(0).ToIntArray()
For strip As Integer = 0 To stripCount - 1
Dim buffer As Byte() = New Byte(byteCounts(strip) - 1) {}
Dim read As Integer = input.ReadRawStrip(strip, buffer, 0, buffer.Length)
output.WriteRawStrip(strip, buffer, read)
Next
End Sub
Private Shared Sub CopyRawTiles(input As Tiff, output As Tiff)
Dim tileCount As Integer = input.NumberOfTiles()
Dim byteCounts As Integer() = input.GetField(TiffTag.TILEBYTECOUNTS)(0).ToIntArray()
For tile As Integer = 0 To tileCount - 1
Dim buffer As Byte() = New Byte(byteCounts(tile) - 1) {}
Dim read As Integer = input.ReadRawTile(tile, buffer, 0, buffer.Length)
output.WriteRawTile(tile, buffer, read)
Next
End Sub
''' <summary>
''' Lazily yields multi-page TIFF chunks (the equivalent of the old
''' Magick.NET 100-pages-per-chunk approach). A new chunk is started when
''' adding the next page would push the chunk past
''' <paramref name="maxChunkBytes"/>, or when <paramref name="maxPagesPerChunk"/>
''' is reached - whichever comes first.
'''
''' Size is the real guard: page count alone can exceed the 2 GB AnyBitmap
''' limit on large pages. The byte total here is the encoded (compressed)
''' size, which is a cheap proxy - validate the cap against your actual
''' pages, since decoded size can be much larger than encoded.
''' </summary>
Public Shared Iterator Function SplitTiffToChunks(
inputPath As String,
Optional maxPagesPerChunk As Integer = 100,
Optional maxChunkBytes As Long = 1500000000L) As IEnumerable(Of Byte())
Using input As Tiff = Tiff.Open(inputPath, "r")
If input Is Nothing Then
Throw New InvalidOperationException($"Could not open TIFF: {inputPath}")
End If
Dim pageCount As Integer = input.NumberOfDirectories()
Dim page As Integer = 0
While page < pageCount
Using ms As New MemoryStream()
Using output As Tiff = Tiff.ClientOpen("InMemory", "w", ms, New TiffStream())
If output Is Nothing Then
Throw New InvalidOperationException("Could not create in-memory TIFF.")
End If
Dim pagesInChunk As Integer = 0
Dim chunkBytes As Long = 0
While page < pageCount AndAlso pagesInChunk < maxPagesPerChunk
input.SetDirectory(CShort(page))
Dim pageBytes As Long = RawPageByteSize(input)
' Stop before exceeding the cap, but always allow at
' least one page so a single large page still goes through.
If pagesInChunk > 0 AndAlso chunkBytes + pageBytes > maxChunkBytes Then
Exit While
End If
CopyTags(input, output)
If input.IsTiled() Then
CopyRawTiles(input, output)
Else
CopyRawStrips(input, output)
End If
output.WriteDirectory() ' finalise this page as one directory in the chunk
chunkBytes += pageBytes
pagesInChunk += 1
page += 1
End While
End Using
Yield ms.ToArray()
End Using
End While
End Using
End Function
Private Shared Function RawPageByteSize(page As Tiff) As Long
Dim tag As TiffTag = If(page.IsTiled(), TiffTag.TILEBYTECOUNTS, TiffTag.STRIPBYTECOUNTS)
Dim counts As Integer() = page.GetField(tag)(0).ToIntArray()
Dim total As Long = 0
For Each c As Integer In counts
total += c
Next
Return total
End Function
End Class
SplitTiffToChunks inicia una nueva sección siempre que la siguiente página superaría el total permitido por maxChunkBytes o una vez alcanzado maxPagesPerChunk, lo que ocurra primero. Una sola página más grande que el límite siempre pasa por sí misma.
3. Hacer OCR a cada sección
Itere sobre SplitTiffToChunks y cargue cada matriz de bytes con OcrInput.LoadImage(byte[]) en lugar de la ruta del archivo, para que nada más grande que el límite llegue jamás a AnyBitmap.
var inputPath = "2gb_benchmark1200.tiff";
var ocr = new IronTesseract();
int chunk = 0;
foreach (byte[] chunkBytes in TiffPageSplitter.SplitTiffToChunks(inputPath, maxPagesPerChunk: 100))
{
using (var ocrInput = new OcrInput())
{
ocrInput.LoadImage(chunkBytes); // loads every page in the chunk
var result = ocr.Read(ocrInput);
Console.WriteLine($"Chunk {chunk}: {result.Text?.Length ?? 0} chars");
}
chunk++;
}
var inputPath = "2gb_benchmark1200.tiff";
var ocr = new IronTesseract();
int chunk = 0;
foreach (byte[] chunkBytes in TiffPageSplitter.SplitTiffToChunks(inputPath, maxPagesPerChunk: 100))
{
using (var ocrInput = new OcrInput())
{
ocrInput.LoadImage(chunkBytes); // loads every page in the chunk
var result = ocr.Read(ocrInput);
Console.WriteLine($"Chunk {chunk}: {result.Text?.Length ?? 0} chars");
}
chunk++;
}
Imports IronOcr
Dim inputPath As String = "2gb_benchmark1200.tiff"
Dim ocr As New IronTesseract()
Dim chunk As Integer = 0
For Each chunkBytes As Byte() In TiffPageSplitter.SplitTiffToChunks(inputPath, maxPagesPerChunk:=100)
Using ocrInput As New OcrInput()
ocrInput.LoadImage(chunkBytes) ' loads every page in the chunk
Dim result = ocr.Read(ocrInput)
Console.WriteLine($"Chunk {chunk}: {If(result.Text?.Length, 0)} chars")
End Using
chunk += 1
Next
Pasar la matriz de bytes mantiene cada entrada bajo el límite. Tenga en cuenta que OcrInput.LoadImage(filePath) actualmente devuelve cero páginas cargadas en lugar de lanzar un error claro cuando el archivo es demasiado grande; esta falla silenciosa es un problema conocido, y el fraccionamiento lo evita por completo.
4. Ajustar los límites de sección para sus datos
Ajuste maxPagesPerChunk y maxChunkBytes para que coincidan con sus TIFFs. Redúzcalos si una sección se aproxima a 2 GB una vez decodificada o si la memoria está limitada; aumente el conteo de páginas para páginas más pequeñas para reducir el desperdicio.
Notas y limitaciones
- Cobertura de etiquetas: el divisor solo conserva las etiquetas listadas en
ScalarIntTagsyScalarDoubleTags. Si sus TIFFs usan etiquetas no cubiertas allí, como perfiles ICC oEXTRASAMPLESpara canales alfa, amplíe esas listas o los bytes copiados en bruto pueden ser mal interpretados. - Dependencia gestionada: LibTiff.NET está completamente gestionada sin binarios nativos, a diferencia de Magick.NET, que incluye bibliotecas nativas ImageMagick que añaden tamaño al paquete y huella de implementación.
- Preservación exacta: la copia de tira o mosaico en bruto evita la decodificación y recodificación que realiza la solución Magick.NET, manteniendo la compresión y los datos de píxeles originales intactos.
Para más información, consulte BitMiracle.LibTiff.NET en NuGet.

