LibTiff.NET ile 2 GB'dan Büyük TIFF'lerin OCR İşlemi
IronOCR, her görüntüyü 32-bit bir tamsayı tarafından indekslenen bir bellek içi AnyBitmap arabelleğe yükler, bu nedenle sistem belleği ne kadar çok olursa olsun yaklaşık 2 GB'da sınırlandırılır. Bundan daha büyük bir TIFF yüklenemez. Aşırı büyük dosyayı BitMiracle.LibTiff.NET ile 2 GB altı parçalara ayırın ve her parçanın OCR işlemini gerçekleştirin.
Bu, Magick.NET geçici çözümüne tam yönetilen bir alternatiftir: ham şerit veya döşeme seviyesinde sayfaları kopyalar, kod çözme veya yeniden kodlama adımı yoktur.
Çözüm
Başlamadan önce, .NET projesinde IronOCR'un (IronTesseract) olduğundan emin olun. Aşağıda kullanılan LibTiff türleri BitMiracle.LibTiff.Classic ad alanında bulunur.
1. LibTiff.NET Paketini Ekle
dotnet add package BitMiracle.LibTiff.NET
2. TiffPageSplitter Yardımcısını Ekle
Yardımcı, her sayfayı ham şerit veya döşeme seviyesinde kopyalar, bu nedenle piksel verileri ve sıkıştırma tam olarak korunur. Sayfaları birer birer akışa alır, tepe belleği bütün dosya yerine yaklaşık bir parça düzeyinde tutar ve her biri boyut sınırının altında kalan çok sayfalı parçalar oluşturur.
/// <summary>
/// Splits a multi-page TIFF into single-page TIFF byte streams without ever
/// holding the whole file in memory. Each page is copied at the raw
/// (still-encoded) strip/tile level, so pixel data and compression are
/// preserved exactly - there is no decode/re-encode step.
///
/// This is the chunking step only. It produces sub-2 GB single-page byte
/// arrays; feeding them to IronOCR (which is where the AnyBitmap 2 GB
/// single-buffer limit lives) is the consumer's job - see TiffOcrExample.
/// </summary>
public static class TiffPageSplitter
{
// Tags that describe how a page's raw strip/tile data is encoded.
// With a raw copy nothing is re-encoded, so every one of these must be
// carried over verbatim or the copied bytes become uninterpretable.
// Extend this list if your TIFFs carry tags not covered here
// (e.g. ICC profiles, EXTRASAMPLES for alpha channels).
private static readonly TiffTag[] ScalarIntTags =
{
TiffTag.IMAGEWIDTH,
TiffTag.IMAGELENGTH,
TiffTag.BITSPERSAMPLE,
TiffTag.SAMPLESPERPIXEL,
TiffTag.COMPRESSION,
TiffTag.PHOTOMETRIC,
TiffTag.FILLORDER,
TiffTag.PLANARCONFIG,
TiffTag.ORIENTATION,
TiffTag.RESOLUTIONUNIT,
TiffTag.PREDICTOR, // required for LZW / Deflate raw copies
TiffTag.SAMPLEFORMAT,
TiffTag.T4OPTIONS, // CCITT Group 3
TiffTag.T6OPTIONS, // CCITT Group 4
TiffTag.SUBFILETYPE,
};
private static readonly TiffTag[] ScalarDoubleTags =
{
TiffTag.XRESOLUTION, // DPI directly affects OCR accuracy
TiffTag.YRESOLUTION,
};
/// <summary>
/// Lazily yields each page of <paramref name="inputPath"/> as a standalone
/// single-page TIFF. The source file stays open for the lifetime of the
/// enumeration and only one page is materialised at a time, so peak memory
/// is roughly one page rather than the whole file.
/// </summary>
public static IEnumerable<byte[]> SplitTiffToPages(string inputPath)
{
using (Tiff input = Tiff.Open(inputPath, "r"))
{
if (input == null)
throw new InvalidOperationException($"Could not open TIFF: {inputPath}");
int pageCount = input.NumberOfDirectories();
for (int page = 0; page < pageCount; page++)
{
input.SetDirectory((short)page);
yield return ExtractCurrentPage(input);
}
}
}
private static byte[] ExtractCurrentPage(Tiff input)
{
using (var ms = new MemoryStream())
{
// Default TiffStream operates on the MemoryStream passed as clientData.
using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
{
if (output == null)
throw new InvalidOperationException("Could not create in-memory TIFF.");
CopyTags(input, output);
if (input.IsTiled())
CopyRawTiles(input, output);
else
CopyRawStrips(input, output);
output.WriteDirectory();
}
return ms.ToArray();
}
}
private static void CopyTags(Tiff input, Tiff output)
{
foreach (TiffTag tag in ScalarIntTags)
{
FieldValue[] v = input.GetField(tag);
if (v != null && v.Length > 0)
output.SetField(tag, v[0].ToInt());
}
foreach (TiffTag tag in ScalarDoubleTags)
{
FieldValue[] v = input.GetField(tag);
if (v != null && v.Length > 0)
output.SetField(tag, v[0].ToDouble());
}
// Strip vs tile layout must match the raw data exactly, otherwise the
// raw bytes won't line up with the declared boundaries.
if (input.IsTiled())
{
output.SetField(TiffTag.TILEWIDTH, input.GetField(TiffTag.TILEWIDTH)[0].ToInt());
output.SetField(TiffTag.TILELENGTH, input.GetField(TiffTag.TILELENGTH)[0].ToInt());
}
else
{
FieldValue[] rps = input.GetField(TiffTag.ROWSPERSTRIP);
if (rps != null && rps.Length > 0)
output.SetField(TiffTag.ROWSPERSTRIP, rps[0].ToInt());
}
// Palette images: the colour map is required to interpret pixel indices.
FieldValue[] cmap = input.GetField(TiffTag.COLORMAP);
if (cmap != null && cmap.Length >= 3)
output.SetField(TiffTag.COLORMAP,
cmap[0].ToShortArray(), cmap[1].ToShortArray(), cmap[2].ToShortArray());
}
private static void CopyRawStrips(Tiff input, Tiff output)
{
int stripCount = input.NumberOfStrips();
int[] byteCounts = input.GetField(TiffTag.STRIPBYTECOUNTS)[0].ToIntArray();
for (int strip = 0; strip < stripCount; strip++)
{
byte[] buffer = new byte[byteCounts[strip]];
int read = input.ReadRawStrip(strip, buffer, 0, buffer.Length);
output.WriteRawStrip(strip, buffer, read);
}
}
private static void CopyRawTiles(Tiff input, Tiff output)
{
int tileCount = input.NumberOfTiles();
int[] byteCounts = input.GetField(TiffTag.TILEBYTECOUNTS)[0].ToIntArray();
for (int tile = 0; tile < tileCount; tile++)
{
byte[] buffer = new byte[byteCounts[tile]];
int read = input.ReadRawTile(tile, buffer, 0, buffer.Length);
output.WriteRawTile(tile, buffer, read);
}
}
/// <summary>
/// Lazily yields multi-page TIFF chunks (the equivalent of the old
/// Magick.NET 100-pages-per-chunk approach). A new chunk is started when
/// adding the next page would push the chunk past
/// <paramref name="maxChunkBytes"/>, or when <paramref name="maxPagesPerChunk"/>
/// is reached - whichever comes first.
///
/// Size is the real guard: page count alone can exceed the 2 GB AnyBitmap
/// limit on large pages. The byte total here is the encoded (compressed)
/// size, which is a cheap proxy - validate the cap against your actual
/// pages, since decoded size can be much larger than encoded.
/// </summary>
public static IEnumerable<byte[]> SplitTiffToChunks(
string inputPath,
int maxPagesPerChunk = 100,
long maxChunkBytes = 1_500_000_000L)
{
using (Tiff input = Tiff.Open(inputPath, "r"))
{
if (input == null)
throw new InvalidOperationException($"Could not open TIFF: {inputPath}");
int pageCount = input.NumberOfDirectories();
int page = 0;
while (page < pageCount)
{
using (var ms = new MemoryStream())
{
using (Tiff output = Tiff.ClientOpen("InMemory", "w", ms, new TiffStream()))
{
if (output == null)
throw new InvalidOperationException("Could not create in-memory TIFF.");
int pagesInChunk = 0;
long chunkBytes = 0;
while (page < pageCount && pagesInChunk < maxPagesPerChunk)
{
input.SetDirectory((short)page);
long pageBytes = RawPageByteSize(input);
// Stop before exceeding the cap, but always allow at
// least one page so a single large page still goes through.
if (pagesInChunk > 0 && chunkBytes + pageBytes > maxChunkBytes)
break;
CopyTags(input, output);
if (input.IsTiled())
CopyRawTiles(input, output);
else
CopyRawStrips(input, output);
output.WriteDirectory(); // finalise this page as one directory in the chunk
chunkBytes += pageBytes;
pagesInChunk++;
page++;
}
}
yield return ms.ToArray();
}
}
}
}
private static long RawPageByteSize(Tiff page)
{
TiffTag tag = page.IsTiled() ? TiffTag.TILEBYTECOUNTS : TiffTag.STRIPBYTECOUNTS;
int[] counts = page.GetField(tag)[0].ToIntArray();
long total = 0;
foreach (int c in counts)
total += c;
return total;
}
}
SplitTiffToChunks bir sonraki sayfanın toplamı maxChunkBytes'un üzerine çıkaracağı veya maxPagesPerChunk'a ulaşıldığı zaman yeni bir parça başlatılır, hangisi önce gelirse. Sınırı aşan tek bir büyük sayfa her zaman kendi başına geçmesine izin verilir.
3. Her Parçanın OCR İşlemini Yapın
SplitTiffToChunks üzerinde yineleyin ve her bayt dizisini, dosya yolu yerine OcrInput.LoadImage(byte[]) ile yükleyin; böylece limiti aşan hiçbir şey AnyBitmap'e ulaşmaz.
var inputPath = "2gb_benchmark1200.tiff";
var ocr = new IronTesseract();
int chunk = 0;
foreach (byte[] chunkBytes in TiffPageSplitter.SplitTiffToChunks(inputPath, maxPagesPerChunk: 100))
{
using (var ocrInput = new OcrInput())
{
ocrInput.LoadImage(chunkBytes); // loads every page in the chunk
var result = ocr.Read(ocrInput);
Console.WriteLine($"Chunk {chunk}: {result.Text?.Length ?? 0} chars");
}
chunk++;
}
Bayt dizisini geçmek, her girdiği sınır altında tutar. Şu anda OcrInput.LoadImage(filePath), dosya çok büyük olduğunda açık bir hata atmak yerine sıfır yüklenmiş sayfa döndürmektedir; Bu sessiz hata bilinmektedir ve parçalama bunu tamamen atlatır.
4. Veriniz İçin Parça Limitlerini Ayarlayın
maxPagesPerChunk ve maxChunkBytes değerlerini TIFF'lerinize uygun şekilde ayarlayın. Bir parça kod çözüldüğünde 2 GB'a yaklaşır veya bellek dar olduğunda, bunları azaltın; Aşırı maliyeti azaltmak için daha küçük sayfalar için sayfa sayısını artırın.
Notlar ve Sınırlamalar
- Etiket kapsamı: ayırıcı, yalnızca
ScalarIntTagsveScalarDoubleTags'de listelenen etiketleri taşır. TIFF'leriniz oradaki kapsama dahil edilmeyen etiketler kullanıyorsa, örneğin ICC profilleri veya alfa kanalları içinEXTRASAMPLESgibi, bu listeleri genişletin veya ham kopyalanan baytlar yanlış yorumlanabilir. - Yönetilen bağımlılık: LibTiff.NET tamamen yönetilen, yerel ikili dosyalar içermeyen bir yapıdır, Magick.NET ise paket boyutuna ve dağıtım ayak izine eklenen ImageMagick yerel kütüphaneleri içerir.
- Kesin muhafaza: ham şerit veya döşeme kopyalama, Magick.NET yaklaşımının gerçekleştirdiği kod çözümünü ve yeniden kodlamayı önler, orijinal sıkıştırma ve piksel verilerini bozulmadan tutar.
Daha fazla okuma için BitMiracle.LibTiff.NET on NuGet adresine bakın.

Curtis Chau, Bilgisayar Bilimleri alanında Lisans Derecesine (Carleton Üniversitesi) sahip ve Node.js, TypeScript, JavaScript ve React konularında uzmanlaşmış ön uç geliştirmeyle ilgileniyor. Sezgisel ve estetik açıdan hoş kullanıcı arayüzleri oluşturma tutkunu, Curtis modern çerçevelerle çalışmayı ve iyi yapılandırılmış, görsel olarak çekici kılavuzlar oluşturmayı seviyor.