PyMuPDF (fitz)
A high-performance Python library for PDF manipulation, text extraction, and document processing.
PyMuPDF (fitz)
A high-performance Python library for PDF manipulation, text extraction, and document processing.
Overview
PyMuPDF (imported as fitz) provides comprehensive PDF handling capabilities including reading, creating, and modifying PDF documents. It excels at text extraction, image handling, and annotation management with excellent performance characteristics.
flowchart LR
A[PDF Document] --> B[PyMuPDF]
B --> C[Text Extraction]
B --> D[Image Handling]
B --> E[Page Operations]
B --> F[Annotations]
B --> G[Metadata]
PDF Document Manipulation
Core operations for opening, saving, and managing PDF documents.
Opening Documents
import fitz # PyMuPDF
# Open existing PDF
doc = fitz.open("document.pdf")
# Open password-protected PDF
doc = fitz.open("encrypted.pdf")
doc.authenticate("password")
# Create new empty PDF
doc = fitz.open()
# Open from bytes/stream
pdf_bytes = open("document.pdf", "rb").read()
doc = fitz.open(stream=pdf_bytes, filetype="pdf")
# Open from URL (with requests)
import requests
response = requests.get("https://example.com/doc.pdf")
doc = fitz.open(stream=response.content, filetype="pdf")
Document Properties
# Basic properties
print(f"Pages: {doc.page_count}")
print(f"Metadata: {doc.metadata}")
print(f"Is encrypted: {doc.is_encrypted}")
print(f"Is PDF: {doc.is_pdf}")
print(f"Permissions: {doc.permissions}")
# Table of contents
toc = doc.get_toc() # Returns list of [level, title, page]
# Set metadata
doc.set_metadata({
"title": "My Document",
"author": "Author Name",
"subject": "Subject",
"keywords": "pdf, python",
"creator": "PyMuPDF",
"producer": "PyMuPDF"
})
Saving Documents
# Save to file
doc.save("output.pdf")
# Save with options
doc.save(
"output.pdf",
garbage=4, # Maximum garbage collection
deflate=True, # Compress streams
clean=True, # Clean unused objects
encryption=fitz.PDF_ENCRYPT_AES_256, # Encryption
owner_pw="owner", # Owner password
user_pw="user", # User password
permissions=fitz.PDF_PERM_PRINT # Permissions
)
# Incremental save (faster for small changes)
doc.save("output.pdf", incremental=True, encryption=fitz.PDF_ENCRYPT_KEEP)
# Save to bytes
pdf_bytes = doc.tobytes(garbage=4, deflate=True)
Closing Documents
# Explicit close
doc.close()
# Context manager (recommended)
with fitz.open("document.pdf") as doc:
# Work with document
text = doc[0].get_text()
# Automatically closed
# Check if closed
if doc.is_closed:
print("Document is closed")
Text Extraction Techniques
Various methods for extracting text content from PDFs.
flowchart TD
A[Text Extraction] --> B[Simple Text]
A --> C[Structured Text]
A --> D[Blocks/Lines/Words]
B --> E[get_text]
C --> F[get_text with dict]
D --> G[get_text with blocks]
Basic Text Extraction
page = doc[0] # First page
# Plain text
text = page.get_text()
# Text with layout preservation
text = page.get_text("text", sort=True)
# HTML output
html = page.get_text("html")
# XML output
xml = page.get_text("xml")
# JSON output
json_text = page.get_text("json")
# Extract from all pages
full_text = ""
for page in doc:
full_text += page.get_text()
Structured Text Extraction
# Get text as dictionary with position info
text_dict = page.get_text("dict")
# Structure: blocks -> lines -> spans
for block in text_dict["blocks"]:
if block["type"] == 0: # Text block
for line in block["lines"]:
for span in line["spans"]:
print(f"Text: {span['text']}")
print(f"Font: {span['font']}")
print(f"Size: {span['size']}")
print(f"Colour: {span['color']}")
print(f"Position: {span['bbox']}")
# Get blocks directly
blocks = page.get_text("blocks")
for b in blocks:
x0, y0, x1, y1, text, block_no, block_type = b
if block_type == 0: # Text
print(text)
Word and Line Extraction
# Extract words with positions
words = page.get_text("words")
for w in words:
x0, y0, x1, y1, word, block_no, line_no, word_no = w
print(f"Word: {word} at ({x0}, {y0})")
# Extract with specific area
rect = fitz.Rect(0, 0, 300, 400)
text = page.get_text("text", clip=rect)
# Search for text
text_instances = page.search_for("search term")
for rect in text_instances:
print(f"Found at: {rect}")
Table Extraction
# Find tables on page
tabs = page.find_tables()
for table in tabs:
# Get table data as list of lists
data = table.extract()
# Convert to pandas DataFrame
import pandas as pd
df = table.to_pandas()
# Table properties
print(f"Rows: {table.row_count}")
print(f"Columns: {table.col_count}")
print(f"Bounding box: {table.bbox}")
Page Operations
Adding, deleting, and reordering pages within documents.
Accessing Pages
# By index (0-based)
page = doc[0] # First page
page = doc[-1] # Last page
# Load specific page
page = doc.load_page(0)
# Iterate all pages
for page_num, page in enumerate(doc):
print(f"Page {page_num}: {page.rect}")
# Page properties
print(f"Width: {page.rect.width}")
print(f"Height: {page.rect.height}")
print(f"Rotation: {page.rotation}")
print(f"Media box: {page.mediabox}")
Adding Pages
# Insert blank page
doc.insert_page(
pno=0, # Insert position (0 = first)
width=595, # A4 width in points
height=842, # A4 height in points
text="Page content" # Optional text
)
# Insert page with text
page = doc.new_page(pno=-1, width=595, height=842) # -1 = append
text_point = fitz.Point(50, 50)
page.insert_text(text_point, "Hello World", fontsize=12)
# Copy page from another document
src_doc = fitz.open("source.pdf")
doc.insert_pdf(
src_doc,
from_page=0, # Source start page
to_page=2, # Source end page
start_at=0 # Insert position in target
)
Deleting Pages
# Delete single page
doc.delete_page(0) # Delete first page
# Delete multiple pages
doc.delete_pages([0, 2, 4]) # Delete pages 0, 2, 4
# Delete page range
doc.delete_pages(from_page=5, to_page=10)
# Keep only specific pages (delete others)
pages_to_keep = [0, 1, 5, 6]
all_pages = list(range(doc.page_count))
pages_to_delete = [p for p in all_pages if p not in pages_to_keep]
doc.delete_pages(pages_to_delete)
Reordering and Copying Pages
# Move page (delete + insert)
page = doc[3]
doc.move_page(3, 0) # Move page 3 to position 0
# Copy page within document
doc.copy_page(0, -1) # Copy first page to end
# Full page reorder
new_order = [2, 0, 1, 3] # New page sequence
doc.select(new_order)
# Reverse all pages
doc.select(list(range(doc.page_count - 1, -1, -1)))
# Merge multiple PDFs
doc1 = fitz.open("file1.pdf")
doc2 = fitz.open("file2.pdf")
doc1.insert_pdf(doc2) # Append doc2 to doc1
doc1.save("merged.pdf")
Page Transformation
# Rotate page
page.set_rotation(90) # 0, 90, 180, 270
# Resize page
new_rect = fitz.Rect(0, 0, 612, 792) # Letter size
page.set_mediabox(new_rect)
# Crop page
crop_rect = fitz.Rect(50, 50, 500, 700)
page.set_cropbox(crop_rect)
Annotations and Highlights
Adding and managing annotations, highlights, and comments.
flowchart LR
A[Annotations] --> B[Highlights]
A --> C[Underlines]
A --> D[Strikeouts]
A --> E[Text Notes]
A --> F[Shapes]
A --> G[Free Text]
Text Markup Annotations
# Highlight text
text_instances = page.search_for("important")
for inst in text_instances:
highlight = page.add_highlight_annot(inst)
highlight.set_colors(stroke=(1, 1, 0)) # Yellow
highlight.update()
# Underline text
underline = page.add_underline_annot(rect)
underline.set_colors(stroke=(0, 0, 1)) # Blue
underline.update()
# Strikeout text
strikeout = page.add_strikeout_annot(rect)
strikeout.update()
# Squiggly underline
squiggly = page.add_squiggly_annot(rect)
squiggly.update()
Text and Note Annotations
# Sticky note
note_point = fitz.Point(100, 100)
note = page.add_text_annot(
note_point,
"This is a comment",
icon="Note" # Note, Comment, Help, Insert, Key, etc.
)
note.set_info(title="Author Name")
note.update()
# Free text annotation
rect = fitz.Rect(100, 100, 300, 150)
free_text = page.add_freetext_annot(
rect,
"Free text content",
fontsize=12,
fontname="helv",
text_color=(0, 0, 0),
fill_color=(1, 1, 0.8)
)
free_text.update()
Shape Annotations
# Rectangle
rect = fitz.Rect(100, 100, 200, 200)
annot = page.add_rect_annot(rect)
annot.set_colors(stroke=(1, 0, 0), fill=(1, 0.9, 0.9))
annot.set_border(width=2)
annot.update()
# Circle/Ellipse
annot = page.add_circle_annot(rect)
annot.set_colors(stroke=(0, 0, 1))
annot.update()
# Line
start = fitz.Point(100, 100)
end = fitz.Point(200, 200)
annot = page.add_line_annot(start, end)
annot.set_colors(stroke=(0, 0, 0))
annot.update()
# Polygon
points = [
fitz.Point(100, 100),
fitz.Point(150, 50),
fitz.Point(200, 100)
]
annot = page.add_polygon_annot(points)
annot.set_colors(stroke=(0, 1, 0), fill=(0.9, 1, 0.9))
annot.update()
# Polyline (open polygon)
annot = page.add_polyline_annot(points)
annot.update()
Managing Annotations
# Get all annotations on page
for annot in page.annots():
print(f"Type: {annot.type}")
print(f"Content: {annot.info.get('content', '')}")
print(f"Rect: {annot.rect}")
# Delete annotation
annot = page.annots()[0]
page.delete_annot(annot)
# Update annotation properties
annot.set_info(
content="Updated content",
title="Author",
subject="Subject"
)
annot.set_opacity(0.5)
annot.update()
# Redact (permanently remove content)
rect = fitz.Rect(100, 100, 200, 150)
page.add_redact_annot(rect, fill=(0, 0, 0))
page.apply_redactions() # Apply permanently
Image Extraction and Insertion
Working with images within PDF documents.
Extracting Images
# Get all images on page
image_list = page.get_images()
for img_index, img in enumerate(image_list):
xref = img[0]
# Extract image
base_image = doc.extract_image(xref)
image_bytes = base_image["image"]
image_ext = base_image["ext"]
# Save to file
with open(f"image_{img_index}.{image_ext}", "wb") as f:
f.write(image_bytes)
# Get image with more details
for img in page.get_images(full=True):
xref, smask, width, height, bpc, colorspace, alt_cs, name, filter, invoker = img
print(f"Image: {name}, Size: {width}x{height}")
# Extract all images from document
for page_num in range(doc.page_count):
page = doc[page_num]
for img_index, img in enumerate(page.get_images()):
xref = img[0]
image = doc.extract_image(xref)
# Process image...
Converting Pages to Images
# Page to pixmap (image)
page = doc[0]
pix = page.get_pixmap()
# Save as PNG
pix.save("page.png")
# Higher resolution
pix = page.get_pixmap(matrix=fitz.Matrix(2, 2)) # 2x zoom
# Specific DPI
zoom = 300 / 72 # 300 DPI
mat = fitz.Matrix(zoom, zoom)
pix = page.get_pixmap(matrix=mat)
# With transparency
pix = page.get_pixmap(alpha=True)
# Specific colour space
pix = page.get_pixmap(colorspace=fitz.csGRAY) # Grayscale
# Get as bytes
img_bytes = pix.tobytes("png")
# Convert to PIL Image
from PIL import Image
import io
img = Image.open(io.BytesIO(pix.tobytes("png")))
Inserting Images
# Insert image from file
page = doc[0]
rect = fitz.Rect(100, 100, 300, 300)
page.insert_image(rect, filename="image.png")
# Insert from bytes
with open("image.png", "rb") as f:
img_bytes = f.read()
page.insert_image(rect, stream=img_bytes)
# Insert with options
page.insert_image(
rect,
filename="image.png",
keep_proportion=True, # Maintain aspect ratio
overlay=True, # On top of existing content
rotate=90 # Rotation in degrees
)
# Insert from pixmap
pix = fitz.Pixmap("image.png")
page.insert_image(rect, pixmap=pix)
# Insert SVG
page.insert_image(rect, filename="image.svg")
# Embed image once, use multiple times
xref = page.insert_image(rect, filename="logo.png")
page2 = doc[1]
page2.insert_image(rect, xref=xref) # Reuse embedded image
Image Manipulation
# Create pixmap from scratch
pix = fitz.Pixmap(fitz.csRGB, fitz.IRect(0, 0, 100, 100), 1)
pix.clear_with(255) # White background
# Draw on pixmap
shape = page.new_shape()
shape.draw_rect(fitz.Rect(10, 10, 90, 90))
shape.finish(color=(1, 0, 0), fill=(0.9, 0.9, 1))
shape.commit()
# Resize pixmap
small_pix = fitz.Pixmap(pix, 50, 50) # New width, height
# Convert colour space
gray_pix = fitz.Pixmap(fitz.csGRAY, pix)
Common Use Cases
Practical examples for typical PDF processing tasks.
Report Generation
def create_report(title, sections, output_path):
"""Generate a PDF report with title and sections."""
doc = fitz.open()
# Title page
page = doc.new_page(width=595, height=842)
# Add title
title_point = fitz.Point(50, 100)
page.insert_text(
title_point,
title,
fontsize=24,
fontname="helv"
)
# Add sections
y_pos = 200
for section_title, content in sections.items():
# Section header
page.insert_text(
fitz.Point(50, y_pos),
section_title,
fontsize=14,
fontname="helv"
)
y_pos += 20
# Section content
text_rect = fitz.Rect(50, y_pos, 545, 792)
rc = page.insert_textbox(
text_rect,
content,
fontsize=10,
fontname="helv"
)
y_pos += abs(rc) + 20
# New page if needed
if y_pos > 750:
page = doc.new_page(width=595, height=842)
y_pos = 50
doc.save(output_path)
doc.close()
# Usage
sections = {
"Introduction": "This report covers...",
"Findings": "The analysis revealed...",
"Conclusion": "In summary..."
}
create_report("Annual Report 2024", sections, "report.pdf")
Data Extraction Pipeline
def extract_document_data(pdf_path):
"""Extract structured data from PDF."""
doc = fitz.open(pdf_path)
result = {
"metadata": doc.metadata,
"pages": [],
"tables": [],
"images": []
}
for page_num, page in enumerate(doc):
# Extract text
page_data = {
"number": page_num + 1,
"text": page.get_text(),
"links": []
}
# Extract links
for link in page.get_links():
if link.get("uri"):
page_data["links"].append(link["uri"])
result["pages"].append(page_data)
# Extract tables
for table in page.find_tables():
result["tables"].append({
"page": page_num + 1,
"data": table.extract()
})
# Extract images
for img in page.get_images():
xref = img[0]
image = doc.extract_image(xref)
result["images"].append({
"page": page_num + 1,
"ext": image["ext"],
"size": len(image["image"])
})
doc.close()
return result
PDF Merging and Splitting
def merge_pdfs(pdf_list, output_path):
"""Merge multiple PDFs into one."""
result = fitz.open()
for pdf_path in pdf_list:
doc = fitz.open(pdf_path)
result.insert_pdf(doc)
doc.close()
result.save(output_path)
result.close()
def split_pdf(input_path, output_dir, pages_per_file=10):
"""Split PDF into smaller files."""
import os
doc = fitz.open(input_path)
for start in range(0, doc.page_count, pages_per_file):
end = min(start + pages_per_file, doc.page_count)
new_doc = fitz.open()
new_doc.insert_pdf(doc, from_page=start, to_page=end - 1)
output_path = os.path.join(
output_dir,
f"split_{start + 1}-{end}.pdf"
)
new_doc.save(output_path)
new_doc.close()
doc.close()
Watermarking
def add_watermark(input_path, output_path, watermark_text):
"""Add text watermark to all pages."""
doc = fitz.open(input_path)
for page in doc:
# Calculate centre position
rect = page.rect
text_length = fitz.get_text_length(
watermark_text,
fontname="helv",
fontsize=50
)
x = (rect.width - text_length) / 2
y = rect.height / 2
# Insert diagonal watermark
page.insert_text(
fitz.Point(x, y),
watermark_text,
fontsize=50,
fontname="helv",
color=(0.8, 0.8, 0.8),
rotate=45
)
doc.save(output_path)
doc.close()
Performance Optimisation Strategies
Techniques for improving PyMuPDF processing speed and memory usage.
flowchart TD
A[Performance] --> B[Memory Management]
A --> C[Batch Processing]
A --> D[Selective Loading]
B --> E[Close documents]
B --> F[Use context managers]
C --> G[Process in chunks]
D --> H[Load only needed pages]
Memory Management
# Always close documents
doc = fitz.open("large.pdf")
try:
# Process document
pass
finally:
doc.close()
# Use context managers
with fitz.open("document.pdf") as doc:
# Process document
pass
# Clear pixmaps when done
pix = page.get_pixmap()
# Use pixmap...
pix = None # Allow garbage collection
# Process large documents in chunks
def process_large_pdf(path, chunk_size=50):
doc = fitz.open(path)
for start in range(0, doc.page_count, chunk_size):
end = min(start + chunk_size, doc.page_count)
for page_num in range(start, end):
page = doc[page_num]
# Process page
text = page.get_text()
# Force garbage collection between chunks
import gc
gc.collect()
doc.close()
Selective Processing
# Only load needed pages
doc = fitz.open("document.pdf")
pages_to_process = [0, 5, 10, 15]
for page_num in pages_to_process:
page = doc.load_page(page_num)
# Process page
# Extract specific regions only
rect = fitz.Rect(0, 0, 300, 200) # Top portion only
text = page.get_text("text", clip=rect)
# Skip image extraction if not needed
text = page.get_text() # Faster than get_text("dict")
Optimised Saving
# Compress and clean on save
doc.save(
"output.pdf",
garbage=4, # Maximum garbage collection (0-4)
deflate=True, # Compress streams
clean=True # Clean unused objects
)
# Incremental save for small changes
doc.save("output.pdf", incremental=True)
# Linear save for web viewing
doc.save("output.pdf", linear=True)
# Avoid repeated saves
# Bad: save after each change
for page in doc:
page.insert_text(...)
doc.save("output.pdf") # Slow!
# Good: save once at end
for page in doc:
page.insert_text(...)
doc.save("output.pdf") # Fast!
Batch Processing
import concurrent.futures
from pathlib import Path
def process_pdf(pdf_path):
"""Process single PDF file."""
with fitz.open(pdf_path) as doc:
text = ""
for page in doc:
text += page.get_text()
return {"path": str(pdf_path), "text": text}
def batch_process(pdf_folder, max_workers=4):
"""Process multiple PDFs in parallel."""
pdf_files = list(Path(pdf_folder).glob("*.pdf"))
results = []
with concurrent.futures.ThreadPoolExecutor(max_workers=max_workers) as executor:
futures = {
executor.submit(process_pdf, pdf): pdf
for pdf in pdf_files
}
for future in concurrent.futures.as_completed(futures):
try:
result = future.result()
results.append(result)
except Exception as e:
print(f"Error: {e}")
return results
Caching Strategies
from functools import lru_cache
class PDFProcessor:
def __init__(self, pdf_path):
self.doc = fitz.open(pdf_path)
self._page_cache = {}
def get_page_text(self, page_num):
"""Cache page text for repeated access."""
if page_num not in self._page_cache:
page = self.doc[page_num]
self._page_cache[page_num] = page.get_text()
return self._page_cache[page_num]
def clear_cache(self):
self._page_cache.clear()
def close(self):
self.clear_cache()
self.doc.close()
Troubleshooting Common PDF Issues
Solutions for frequently encountered problems.
Document Opening Issues
# Handle encrypted PDFs
doc = fitz.open("encrypted.pdf")
if doc.needs_pass:
if not doc.authenticate("password"):
raise ValueError("Incorrect password")
# Handle corrupted PDFs
# MuPDF auto-repairs recoverable files on open; FileDataError means unrecoverable.
try:
doc = fitz.open("corrupted.pdf")
if doc.is_repaired: # file had errors that MuPDF fixed on open
print("PDF was auto-repaired on open")
doc.save("repaired.pdf", garbage=4, clean=True) # persist the repair
except fitz.FileDataError as e:
# Open could not parse the file at all — nothing to recover
print(f"PDF is corrupted and could not be opened: {e}")
# Handle non-PDF files
try:
doc = fitz.open("file.xyz")
except Exception as e:
print(f"Cannot open file: {e}")
Text Extraction Issues
# Empty text from scanned PDFs
text = page.get_text()
if not text.strip():
print("PDF may be scanned - use OCR")
# Consider using OCR like pytesseract
pix = page.get_pixmap()
# Pass to OCR engine
# Garbled text (encoding issues)
text = page.get_text("text", flags=fitz.TEXT_PRESERVE_WHITESPACE)
# Missing text due to fonts
text_dict = page.get_text("dict")
for block in text_dict["blocks"]:
if block["type"] == 0:
for line in block["lines"]:
for span in line["spans"]:
if not span["text"] and span.get("font"):
print(f"Font issue: {span['font']}")
# Text in wrong order
text = page.get_text("text", sort=True) # Sort by position
Image Extraction Issues
# No images found
images = page.get_images()
if not images:
# Check for inline images
images = page.get_images(full=True)
# Images may be in form XObjects
xrefs = page.get_xobjects()
# Extracting masked images
for img in page.get_images(full=True):
xref = img[0]
smask = img[1] # Soft mask reference
base = doc.extract_image(xref)
if smask:
# Apply mask
mask = doc.extract_image(smask)
# Combine base and mask...
# Image quality issues
# Use higher resolution
pix = page.get_pixmap(matrix=fitz.Matrix(3, 3)) # 3x zoom
Annotation Issues
# Annotations not visible after save
for annot in page.annots():
annot.update() # Must call update!
doc.save("output.pdf")
# Annotations not appearing
# Check annotation flags
annot.set_flags(fitz.PDF_ANNOT_IS_PRINT) # Make printable
annot.update()
# Redactions not applied
page.add_redact_annot(rect)
page.apply_redactions() # Must apply!
doc.save("output.pdf")
Memory and Performance Issues
# Out of memory with large PDFs
import gc
doc = fitz.open("huge.pdf")
for i in range(doc.page_count):
page = doc[i]
# Process page
del page
if i % 100 == 0:
gc.collect()
doc.close()
# Slow processing
# Avoid: getting full dict for simple text
text_dict = page.get_text("dict") # Slow
# Prefer: simple text extraction
text = page.get_text() # Fast
# Avoid: repeated small saves
# Prefer: batch changes then save once
Quick Reference
Essential Operations
| Operation | Code |
|---|---|
| Open PDF | doc = fitz.open("file.pdf") |
| Get page | page = doc[0] |
| Extract text | text = page.get_text() |
| Save PDF | doc.save("output.pdf") |
| Close PDF | doc.close() |
| Page count | doc.page_count |
| Add page | doc.new_page() |
| Delete page | doc.delete_page(0) |
| Insert image | page.insert_image(rect, filename="img.png") |
| Add highlight | page.add_highlight_annot(rect) |
| Page to image | pix = page.get_pixmap() |
| Search text | rects = page.search_for("text") |
| Get metadata | doc.metadata |
| Merge PDFs | doc1.insert_pdf(doc2) |
Common Rectangle Operations
# Create rectangle
rect = fitz.Rect(x0, y0, x1, y1) # Top-left to bottom-right
# Page rectangle
rect = page.rect
# Normalise rectangle
rect = rect.normalize()
# Check containment
if rect1.contains(rect2):
pass
# Intersection
rect3 = rect1 & rect2
# Union
rect4 = rect1 | rect2
Colour Specifications
# RGB tuples (0-1 range)
red = (1, 0, 0)
green = (0, 1, 0)
blue = (0, 0, 1)
white = (1, 1, 1)
black = (0, 0, 0)
yellow = (1, 1, 0)
# Colour spaces
fitz.csRGB # RGB
fitz.csGRAY # Grayscale
fitz.csCMYK # CMYK
Page Sizes in Points
| Size | Width x Height |
|---|---|
| A4 | 595 x 842 |
| Letter | 612 x 792 |
| Legal | 612 x 1008 |
| A3 | 842 x 1191 |
| A5 | 420 x 595 |
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
| Empty text extraction | Scanned PDF | Use OCR (pytesseract, etc.) |
FileDataError on open |
Corrupted PDF | Recoverable files auto-repair on open (check doc.is_repaired, re-save with clean=True); unrecoverable ones raise |
| Annotations not saved | Missing update() |
Call annot.update() before save |
| Poor image quality | Low resolution | Increase pixmap matrix zoom factor |
| Memory errors | Large documents | Process in chunks, use context managers |
| Password error | Encrypted PDF | Use doc.authenticate(password) |
| Text in wrong order | Complex layout | Use sort=True in get_text() |
| Slow performance | Repeated saves | Batch changes, save once at end |
| Missing images | Inline images | Use get_images(full=True) |
| Garbled characters | Font issues | Check font embedding, try different extraction |