OCR of TIFFs Over 2 GB with LibTiff.NET
IronOCR lädt jedes Bild in einen im Speicher befindlichen AnyBitmap Puffer, der durch einen 32-Bit-Integer indexiert wird, sodass er auf etwa 2 GB begrenzt ist, unabhängig davon, wie viel Systemspeicher verfügbar ist. Ein TIFF, das größer ist als das, kann nicht geladen werden. Teilen Sie die zu große Datei in unter 2 GB große Segmente mit BitMiracle.LibTiff.NET auf und führen Sie OCR für jedes Segment aus.
Dies ist eine vollständig verwaltete Alternative zum Magick.NET-Workaround: Es kopiert Seiten auf Rohstreifen- oder Kachel-Ebene, ohne De- oder Rekodierungsschritt.
Lösung
Bevor Sie beginnen, stellen Sie sicher, dass Sie IronOCR (IronTesseract) in einem .NET-Projekt haben. Die unten verwendeten LibTiff-Typen befinden sich im BitMiracle.LibTiff.Classic Namespace.
1. Fügen Sie das LibTiff.NET-Paket hinzu
dotnet add package BitMiracle.LibTiff.NET
2. Fügen Sie den TiffPageSplitter-Helfer hinzu
Der Helfer kopiert jede Seite auf Rohstreifen- oder Kachel-Ebene, sodass die Pixeldaten und die Komprimierung exakt erhalten bleiben. Er streamt Seiten jeweils einzeln, wobei der Spitzenwert des Speichers auf etwa ein Segment statt auf die gesamte Datei beschränkt bleibt, und liefert mehrseitige Segmente, die jeweils unter der Größenobergrenze bleiben.
/// <summary>
/// Splits a multi-page TIFF into single-page TIFF byte streams without ever
/// holding the whole file in memory. Each page is copied at the raw
/// (still-encoded) strip/tile level, so pixel data and compression are
/// preserved exactly - there is no decode/re-encode step.
///
/// This is the chunking step only. It produces sub-2 GB single-page byte
/// arrays; feeding them to IronOCR (which is where the AnyBitmap 2 GB
/// single-buffer limit lives) is the consumer's job - see TiffOcrExample.
/// </summary>
public static class TiffPageSplitter
{
// Tags that describe how a page's raw strip/tile data is encoded.
// With a raw copy nothing is re-encoded, so every one of these must be
// carried over verbatim or the copied bytes become uninterpretable.
// Extend this list if your TIFFs carry tags not covered here
// (e.g. ICC profiles, EXTRASAMPLES for alpha channels).
private static readonly TiffTag[] ScalarIntTags =
{
TiffTag.IMAGEWIDTH,
TiffTag.IMAGELENGTH,
TiffTag.BITSPERSAMPLE,
TiffTag.SAMPLESPERPIXEL,
TiffTag.COMPRESSION,
TiffTag.PHOTOMETRIC,
TiffTag.FILLORDER,
TiffTag.PLANARCONFIG,
TiffTag.ORIENTATION,
TiffTag.RESOLUTIONUNIT,
TiffTag.PREDICTOR, // required for LZW / Deflate raw copies
TiffTag.SAMPLEFORMAT,
TiffTag.T4OPTIONS, // CCITT Group 3
TiffTag.T6OPTIONS, // CCITT Group 4
TiffTag.SUBFILETYPE,
};
private static readonly TiffTag[] ScalarDoubleTags =
{
TiffTag.XRESOLUTION, // DPI directly affects OCR accuracy
TiffTag.YRESOLUTION,
};
/// <summary>
/// Lazily yields each page of <paramref name="inputPath"/> as a standalone
/// single-page TIFF. The source file stays open for the lifetime of the
/// enumeration and only one page is materialised at a time, so peak memory
/// is roughly one page rather than the whole file.
/// </summary>
public static IEnumerable<byte[]> SplitTiffToPages(string inputPath)
{
using (Tiff input = Tiff.Open(inputPath, "r"))
{
if (input == null)
throw new InvalidOperationException($"Could not open TIFF: {inputPath}");
int pageCount = input.NumberOfDirectories();
for (int page = 0; page < pageCount; page++)
{
input.SetDirectory((short)page);
yield return ExtractCurrentPage(input);
}
}
}
private static byte[] ExtractCurrentPage(Tiff input)
{
using (var ms = new MemoryStream())
{
// Default TiffStream operates on the MemoryStream passed as clientData.
using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
{
if (output == null)
throw new InvalidOperationException("Could not create in-memory TIFF.");
CopyTags(input, output);
if (input.IsTiled())
CopyRawTiles(input, output);
else
CopyRawStrips(input, output);
output.WriteDirectory();
}
return ms.ToArray();
}
}
private static void CopyTags(Tiff input, Tiff output)
{
foreach (TiffTag tag in ScalarIntTags)
{
FieldValue[] v = input.GetField(tag);
if (v != null && v.Length > 0)
output.SetField(tag, v[0].ToInt());
}
foreach (TiffTag tag in ScalarDoubleTags)
{
FieldValue[] v = input.GetField(tag);
if (v != null && v.Length > 0)
output.SetField(tag, v[0].ToDouble());
}
// Strip vs tile layout must match the raw data exactly, otherwise the
// raw bytes won't line up with the declared boundaries.
if (input.IsTiled())
{
output.SetField(TiffTag.TILEWIDTH, input.GetField(TiffTag.TILEWIDTH)[0].ToInt());
output.SetField(TiffTag.TILELENGTH, input.GetField(TiffTag.TILELENGTH)[0].ToInt());
}
else
{
FieldValue[] rps = input.GetField(TiffTag.ROWSPERSTRIP);
if (rps != null && rps.Length > 0)
output.SetField(TiffTag.ROWSPERSTRIP, rps[0].ToInt());
}
// Palette images: the colour map is required to interpret pixel indices.
FieldValue[] cmap = input.GetField(TiffTag.COLORMAP);
if (cmap != null && cmap.Length >= 3)
output.SetField(TiffTag.COLORMAP,
cmap[0].ToShortArray(), cmap[1].ToShortArray(), cmap[2].ToShortArray());
}
private static void CopyRawStrips(Tiff input, Tiff output)
{
int stripCount = input.NumberOfStrips();
int[] byteCounts = input.GetField(TiffTag.STRIPBYTECOUNTS)[0].ToIntArray();
for (int strip = 0; strip < stripCount; strip++)
{
byte[] buffer = new byte[byteCounts[strip]];
int read = input.ReadRawStrip(strip, buffer, 0, buffer.Length);
output.WriteRawStrip(strip, buffer, read);
}
}
private static void CopyRawTiles(Tiff input, Tiff output)
{
int tileCount = input.NumberOfTiles();
int[] byteCounts = input.GetField(TiffTag.TILEBYTECOUNTS)[0].ToIntArray();
for (int tile = 0; tile < tileCount; tile++)
{
byte[] buffer = new byte[byteCounts[tile]];
int read = input.ReadRawTile(tile, buffer, 0, buffer.Length);
output.WriteRawTile(tile, buffer, read);
}
}
/// <summary>
/// Lazily yields multi-page TIFF chunks (the equivalent of the old
/// Magick.NET 100-pages-per-chunk approach). A new chunk is started when
/// adding the next page would push the chunk past
/// <paramref name="maxChunkBytes"/>, or when <paramref name="maxPagesPerChunk"/>
/// is reached - whichever comes first.
///
/// Size is the real guard: page count alone can exceed the 2 GB AnyBitmap
/// limit on large pages. The byte total here is the encoded (compressed)
/// size, which is a cheap proxy - validate the cap against your actual
/// pages, since decoded size can be much larger than encoded.
/// </summary>
public static IEnumerable<byte[]> SplitTiffToChunks(
string inputPath,
int maxPagesPerChunk = 100,
long maxChunkBytes = 1_500_000_000L)
{
using (Tiff input = Tiff.Open(inputPath, "r"))
{
if (input == null)
throw new InvalidOperationException($"Could not open TIFF: {inputPath}");
int pageCount = input.NumberOfDirectories();
int page = 0;
while (page < pageCount)
{
using (var ms = new MemoryStream())
{
using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
{
if (output == null)
throw new InvalidOperationException("Could not create in-memory TIFF.");
int pagesInChunk = 0;
long chunkBytes = 0;
while (page < pageCount && pagesInChunk < maxPagesPerChunk)
{
input.SetDirectory((short)page);
long pageBytes = RawPageByteSize(input);
// Stop before exceeding the cap, but always allow at
// least one page so a single large page still goes through.
if (pagesInChunk > 0 && chunkBytes + pageBytes > maxChunkBytes)
break;
CopyTags(input, output);
if (input.IsTiled())
CopyRawTiles(input, output);
else
CopyRawStrips(input, output);
output.WriteDirectory(); // finalise this page as one directory in the chunk
chunkBytes += pageBytes;
pagesInChunk++;
page++;
}
}
yield return ms.ToArray();
}
}
}
}
private static long RawPageByteSize(Tiff page)
{
TiffTag tag = page.IsTiled() ? TiffTag.TILEBYTECOUNTS : TiffTag.STRIPBYTECOUNTS;
int[] counts = page.GetField(tag)[0].ToIntArray();
long total = 0;
foreach (int c in counts)
total += c;
return total;
}
}
SplitTiffToChunks beginnt ein neues Segment, sobald die nächste Seite die Gesamtsumme über maxChunkBytes drücken würde oder sobald maxPagesPerChunk erreicht ist, je nachdem, was zuerst eintritt. Eine einzelne Seite, die größer ist als die Grenze, ist immer alleine zulässig.
3. Führen Sie OCR für jedes Segment aus
Iterieren Sie über SplitTiffToChunks und laden Sie jedes Byte-Array mit OcrInput.LoadImage(byte[]) anstelle des Dateipfads, sodass nichts Größeres als das Limit jemals AnyBitmap erreicht.
var inputPath = "2gb_benchmark1200.tiff";
var ocr = new IronTesseract();
int chunk = 0;
foreach (byte[] chunkBytes in TiffPageSplitter.SplitTiffToChunks(inputPath, maxPagesPerChunk: 100))
{
using (var ocrInput = new OcrInput())
{
ocrInput.LoadImage(chunkBytes); // loads every page in the chunk
var result = ocr.Read(ocrInput);
Console.WriteLine($"Chunk {chunk}: {result.Text?.Length ?? 0} chars");
}
chunk++;
}
Das Übergeben des Byte-Arrays hält jedes Eingabe unter der Grenze. Beachten Sie, dass OcrInput.LoadImage(filePath) derzeit null geladene Seiten zurückgibt, anstatt einen klaren Fehler auszulösen, wenn die Datei zu groß ist; dieser stille Fehler ist ein bekanntes Problem, und das Segmentieren umgeht es vollständig.
4. Passen Sie die Segmentgrenzen für Ihre Daten an
Passen Sie maxPagesPerChunk und maxChunkBytes an Ihre TIFFs an. Senken Sie sie, wenn ein Segment nach dem Dekodieren fast 2 GB erreicht oder wenn der Speicher knapp ist; erhöhen Sie die Seitenanzahl für kleinere Seiten, um den Overhead zu reduzieren.
Hinweise und Einschränkungen
- Tag-Abdeckung: der Splitter überträgt nur die in
ScalarIntTagsundScalarDoubleTagsaufgeführten Tags. Wenn Ihre TIFFs Tags verwenden, die dort nicht abgedeckt sind, wie ICC-Profile oderEXTRASAMPLESfür Alphakanäle, erweitern Sie diese Listen, oder die roh kopierten Bytes könnten falsch interpretiert werden. - Verwaltete Abhängigkeit: LibTiff.NET ist vollständig verwaltet ohne native Binärdateien, im Gegensatz zu Magick.NET, das ImageMagick-native Bibliotheken enthält, die die Paketgröße und den Bereitstellungsaufwand erhöhen.
- Exakte Erhaltung: Die Rohstreifen- oder Kachelkopie vermeidet das De- und Rekodieren, das bei der Magick.NET-Methode auftritt, und erhält die ursprüngliche Komprimierung und Pixeldaten intakt.
For further reading, see BitMiracle.LibTiff.NET on NuGet.

Curtis Chau hat einen Bachelor-Abschluss in Informatik von der Carleton University und ist spezialisiert auf Frontend-Entwicklung mit Expertise in Node.js, TypeScript, JavaScript und React. Leidenschaftlich widmet er sich der Erstellung intuitiver und ästhetisch ansprechender Benutzerschnittstellen und arbeitet gerne mit modernen Frameworks sowie der Erstellung gut strukturierter, optisch ansprechender Handbücher.