Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

PyMuPDF (fitz)

A high-performance Python library for PDF manipulation, text extraction, and document processing.

PyMuPDF (fitz)

A high-performance Python library for PDF manipulation, text extraction, and document processing.

Overview

PyMuPDF (imported as fitz) provides comprehensive PDF handling capabilities including reading, creating, and modifying PDF documents. It excels at text extraction, image handling, and annotation management with excellent performance characteristics.

PDF DocumentPyMuPDFText ExtractionImage HandlingPage OperationsAnnotationsMetadataPDF DocumentPyMuPDFText ExtractionImage HandlingPage OperationsAnnotationsMetadata

PDF Document Manipulation

Core operations for opening, saving, and managing PDF documents.

Opening Documents

import fitz  # PyMuPDF

# Open existing PDF
doc = fitz.open("document.pdf")

# Open password-protected PDF
doc = fitz.open("encrypted.pdf")
doc.authenticate("password")

# Create new empty PDF
doc = fitz.open()

# Open from bytes/stream
pdf_bytes = open("document.pdf", "rb").read()
doc = fitz.open(stream=pdf_bytes, filetype="pdf")

# Open from URL (with requests)
import requests
response = requests.get("https://example.com/doc.pdf")
doc = fitz.open(stream=response.content, filetype="pdf")

Document Properties

# Basic properties
print(f"Pages: {doc.page_count}")
print(f"Metadata: {doc.metadata}")
print(f"Is encrypted: {doc.is_encrypted}")
print(f"Is PDF: {doc.is_pdf}")
print(f"Permissions: {doc.permissions}")

# Table of contents
toc = doc.get_toc()  # Returns list of [level, title, page]

# Set metadata
doc.set_metadata({
    "title": "My Document",
    "author": "Author Name",
    "subject": "Subject",
    "keywords": "pdf, python",
    "creator": "PyMuPDF",
    "producer": "PyMuPDF"
})

Saving Documents

# Save to file
doc.save("output.pdf")

# Save with options
doc.save(
    "output.pdf",
    garbage=4,           # Maximum garbage collection
    deflate=True,        # Compress streams
    clean=True,          # Clean unused objects
    encryption=fitz.PDF_ENCRYPT_AES_256,  # Encryption
    owner_pw="owner",    # Owner password
    user_pw="user",      # User password
    permissions=fitz.PDF_PERM_PRINT  # Permissions
)

# Incremental save (faster for small changes)
doc.save("output.pdf", incremental=True, encryption=fitz.PDF_ENCRYPT_KEEP)

# Save to bytes
pdf_bytes = doc.tobytes(garbage=4, deflate=True)

Closing Documents

# Explicit close
doc.close()

# Context manager (recommended)
with fitz.open("document.pdf") as doc:
    # Work with document
    text = doc[0].get_text()
# Automatically closed

# Check if closed
if doc.is_closed:
    print("Document is closed")

Text Extraction Techniques

Various methods for extracting text content from PDFs.

Text ExtractionSimple TextStructured TextBlocks/Lines/Wordsget_textget_text with dictget_text with blocksText ExtractionSimple TextStructured TextBlocks/Lines/Wordsget_textget_text with dictget_text with blocks

Basic Text Extraction

page = doc[0]  # First page

# Plain text
text = page.get_text()

# Text with layout preservation
text = page.get_text("text", sort=True)

# HTML output
html = page.get_text("html")

# XML output
xml = page.get_text("xml")

# JSON output
json_text = page.get_text("json")

# Extract from all pages
full_text = ""
for page in doc:
    full_text += page.get_text()

Structured Text Extraction

# Get text as dictionary with position info
text_dict = page.get_text("dict")

# Structure: blocks -> lines -> spans
for block in text_dict["blocks"]:
    if block["type"] == 0:  # Text block
        for line in block["lines"]:
            for span in line["spans"]:
                print(f"Text: {span['text']}")
                print(f"Font: {span['font']}")
                print(f"Size: {span['size']}")
                print(f"Colour: {span['color']}")
                print(f"Position: {span['bbox']}")

# Get blocks directly
blocks = page.get_text("blocks")
for b in blocks:
    x0, y0, x1, y1, text, block_no, block_type = b
    if block_type == 0:  # Text
        print(text)

Word and Line Extraction

# Extract words with positions
words = page.get_text("words")
for w in words:
    x0, y0, x1, y1, word, block_no, line_no, word_no = w
    print(f"Word: {word} at ({x0}, {y0})")

# Extract with specific area
rect = fitz.Rect(0, 0, 300, 400)
text = page.get_text("text", clip=rect)

# Search for text
text_instances = page.search_for("search term")
for rect in text_instances:
    print(f"Found at: {rect}")

Table Extraction

# Find tables on page
tabs = page.find_tables()

for table in tabs:
    # Get table data as list of lists
    data = table.extract()

    # Convert to pandas DataFrame
    import pandas as pd
    df = table.to_pandas()

    # Table properties
    print(f"Rows: {table.row_count}")
    print(f"Columns: {table.col_count}")
    print(f"Bounding box: {table.bbox}")

Page Operations

Adding, deleting, and reordering pages within documents.

Accessing Pages

# By index (0-based)
page = doc[0]          # First page
page = doc[-1]         # Last page

# Load specific page
page = doc.load_page(0)

# Iterate all pages
for page_num, page in enumerate(doc):
    print(f"Page {page_num}: {page.rect}")

# Page properties
print(f"Width: {page.rect.width}")
print(f"Height: {page.rect.height}")
print(f"Rotation: {page.rotation}")
print(f"Media box: {page.mediabox}")

Adding Pages

# Insert blank page
doc.insert_page(
    pno=0,              # Insert position (0 = first)
    width=595,          # A4 width in points
    height=842,         # A4 height in points
    text="Page content" # Optional text
)

# Insert page with text
page = doc.new_page(pno=-1, width=595, height=842)  # -1 = append
text_point = fitz.Point(50, 50)
page.insert_text(text_point, "Hello World", fontsize=12)

# Copy page from another document
src_doc = fitz.open("source.pdf")
doc.insert_pdf(
    src_doc,
    from_page=0,        # Source start page
    to_page=2,          # Source end page
    start_at=0          # Insert position in target
)

Deleting Pages

# Delete single page
doc.delete_page(0)      # Delete first page

# Delete multiple pages
doc.delete_pages([0, 2, 4])  # Delete pages 0, 2, 4

# Delete page range
doc.delete_pages(from_page=5, to_page=10)

# Keep only specific pages (delete others)
pages_to_keep = [0, 1, 5, 6]
all_pages = list(range(doc.page_count))
pages_to_delete = [p for p in all_pages if p not in pages_to_keep]
doc.delete_pages(pages_to_delete)

Reordering and Copying Pages

# Move page (delete + insert)
page = doc[3]
doc.move_page(3, 0)  # Move page 3 to position 0

# Copy page within document
doc.copy_page(0, -1)  # Copy first page to end

# Full page reorder
new_order = [2, 0, 1, 3]  # New page sequence
doc.select(new_order)

# Reverse all pages
doc.select(list(range(doc.page_count - 1, -1, -1)))

# Merge multiple PDFs
doc1 = fitz.open("file1.pdf")
doc2 = fitz.open("file2.pdf")
doc1.insert_pdf(doc2)  # Append doc2 to doc1
doc1.save("merged.pdf")

Page Transformation

# Rotate page
page.set_rotation(90)  # 0, 90, 180, 270

# Resize page
new_rect = fitz.Rect(0, 0, 612, 792)  # Letter size
page.set_mediabox(new_rect)

# Crop page
crop_rect = fitz.Rect(50, 50, 500, 700)
page.set_cropbox(crop_rect)

Annotations and Highlights

Adding and managing annotations, highlights, and comments.

AnnotationsHighlightsUnderlinesStrikeoutsText NotesShapesFree TextAnnotationsHighlightsUnderlinesStrikeoutsText NotesShapesFree Text

Text Markup Annotations

# Highlight text
text_instances = page.search_for("important")
for inst in text_instances:
    highlight = page.add_highlight_annot(inst)
    highlight.set_colors(stroke=(1, 1, 0))  # Yellow
    highlight.update()

# Underline text
underline = page.add_underline_annot(rect)
underline.set_colors(stroke=(0, 0, 1))  # Blue
underline.update()

# Strikeout text
strikeout = page.add_strikeout_annot(rect)
strikeout.update()

# Squiggly underline
squiggly = page.add_squiggly_annot(rect)
squiggly.update()

Text and Note Annotations

# Sticky note
note_point = fitz.Point(100, 100)
note = page.add_text_annot(
    note_point,
    "This is a comment",
    icon="Note"  # Note, Comment, Help, Insert, Key, etc.
)
note.set_info(title="Author Name")
note.update()

# Free text annotation
rect = fitz.Rect(100, 100, 300, 150)
free_text = page.add_freetext_annot(
    rect,
    "Free text content",
    fontsize=12,
    fontname="helv",
    text_color=(0, 0, 0),
    fill_color=(1, 1, 0.8)
)
free_text.update()

Shape Annotations

# Rectangle
rect = fitz.Rect(100, 100, 200, 200)
annot = page.add_rect_annot(rect)
annot.set_colors(stroke=(1, 0, 0), fill=(1, 0.9, 0.9))
annot.set_border(width=2)
annot.update()

# Circle/Ellipse
annot = page.add_circle_annot(rect)
annot.set_colors(stroke=(0, 0, 1))
annot.update()

# Line
start = fitz.Point(100, 100)
end = fitz.Point(200, 200)
annot = page.add_line_annot(start, end)
annot.set_colors(stroke=(0, 0, 0))
annot.update()

# Polygon
points = [
    fitz.Point(100, 100),
    fitz.Point(150, 50),
    fitz.Point(200, 100)
]
annot = page.add_polygon_annot(points)
annot.set_colors(stroke=(0, 1, 0), fill=(0.9, 1, 0.9))
annot.update()

# Polyline (open polygon)
annot = page.add_polyline_annot(points)
annot.update()

Managing Annotations

# Get all annotations on page
for annot in page.annots():
    print(f"Type: {annot.type}")
    print(f"Content: {annot.info.get('content', '')}")
    print(f"Rect: {annot.rect}")

# Delete annotation
annot = page.annots()[0]
page.delete_annot(annot)

# Update annotation properties
annot.set_info(
    content="Updated content",
    title="Author",
    subject="Subject"
)
annot.set_opacity(0.5)
annot.update()

# Redact (permanently remove content)
rect = fitz.Rect(100, 100, 200, 150)
page.add_redact_annot(rect, fill=(0, 0, 0))
page.apply_redactions()  # Apply permanently

Image Extraction and Insertion

Working with images within PDF documents.

Extracting Images

# Get all images on page
image_list = page.get_images()
for img_index, img in enumerate(image_list):
    xref = img[0]

    # Extract image
    base_image = doc.extract_image(xref)
    image_bytes = base_image["image"]
    image_ext = base_image["ext"]

    # Save to file
    with open(f"image_{img_index}.{image_ext}", "wb") as f:
        f.write(image_bytes)

# Get image with more details
for img in page.get_images(full=True):
    xref, smask, width, height, bpc, colorspace, alt_cs, name, filter, invoker = img
    print(f"Image: {name}, Size: {width}x{height}")

# Extract all images from document
for page_num in range(doc.page_count):
    page = doc[page_num]
    for img_index, img in enumerate(page.get_images()):
        xref = img[0]
        image = doc.extract_image(xref)
        # Process image...

Converting Pages to Images

# Page to pixmap (image)
page = doc[0]
pix = page.get_pixmap()

# Save as PNG
pix.save("page.png")

# Higher resolution
pix = page.get_pixmap(matrix=fitz.Matrix(2, 2))  # 2x zoom

# Specific DPI
zoom = 300 / 72  # 300 DPI
mat = fitz.Matrix(zoom, zoom)
pix = page.get_pixmap(matrix=mat)

# With transparency
pix = page.get_pixmap(alpha=True)

# Specific colour space
pix = page.get_pixmap(colorspace=fitz.csGRAY)  # Grayscale

# Get as bytes
img_bytes = pix.tobytes("png")

# Convert to PIL Image
from PIL import Image
import io
img = Image.open(io.BytesIO(pix.tobytes("png")))

Inserting Images

# Insert image from file
page = doc[0]
rect = fitz.Rect(100, 100, 300, 300)
page.insert_image(rect, filename="image.png")

# Insert from bytes
with open("image.png", "rb") as f:
    img_bytes = f.read()
page.insert_image(rect, stream=img_bytes)

# Insert with options
page.insert_image(
    rect,
    filename="image.png",
    keep_proportion=True,  # Maintain aspect ratio
    overlay=True,          # On top of existing content
    rotate=90              # Rotation in degrees
)

# Insert from pixmap
pix = fitz.Pixmap("image.png")
page.insert_image(rect, pixmap=pix)

# Insert SVG
page.insert_image(rect, filename="image.svg")

# Embed image once, use multiple times
xref = page.insert_image(rect, filename="logo.png")
page2 = doc[1]
page2.insert_image(rect, xref=xref)  # Reuse embedded image

Image Manipulation

# Create pixmap from scratch
pix = fitz.Pixmap(fitz.csRGB, fitz.IRect(0, 0, 100, 100), 1)
pix.clear_with(255)  # White background

# Draw on pixmap
shape = page.new_shape()
shape.draw_rect(fitz.Rect(10, 10, 90, 90))
shape.finish(color=(1, 0, 0), fill=(0.9, 0.9, 1))
shape.commit()

# Resize pixmap
small_pix = fitz.Pixmap(pix, 50, 50)  # New width, height

# Convert colour space
gray_pix = fitz.Pixmap(fitz.csGRAY, pix)

Common Use Cases

Practical examples for typical PDF processing tasks.

Report Generation

def create_report(title, sections, output_path):
    """Generate a PDF report with title and sections."""
    doc = fitz.open()

    # Title page
    page = doc.new_page(width=595, height=842)

    # Add title
    title_point = fitz.Point(50, 100)
    page.insert_text(
        title_point,
        title,
        fontsize=24,
        fontname="helv"
    )

    # Add sections
    y_pos = 200
    for section_title, content in sections.items():
        # Section header
        page.insert_text(
            fitz.Point(50, y_pos),
            section_title,
            fontsize=14,
            fontname="helv"
        )
        y_pos += 20

        # Section content
        text_rect = fitz.Rect(50, y_pos, 545, 792)
        rc = page.insert_textbox(
            text_rect,
            content,
            fontsize=10,
            fontname="helv"
        )
        y_pos += abs(rc) + 20

        # New page if needed
        if y_pos > 750:
            page = doc.new_page(width=595, height=842)
            y_pos = 50

    doc.save(output_path)
    doc.close()

# Usage
sections = {
    "Introduction": "This report covers...",
    "Findings": "The analysis revealed...",
    "Conclusion": "In summary..."
}
create_report("Annual Report 2024", sections, "report.pdf")

Data Extraction Pipeline

def extract_document_data(pdf_path):
    """Extract structured data from PDF."""
    doc = fitz.open(pdf_path)

    result = {
        "metadata": doc.metadata,
        "pages": [],
        "tables": [],
        "images": []
    }

    for page_num, page in enumerate(doc):
        # Extract text
        page_data = {
            "number": page_num + 1,
            "text": page.get_text(),
            "links": []
        }

        # Extract links
        for link in page.get_links():
            if link.get("uri"):
                page_data["links"].append(link["uri"])

        result["pages"].append(page_data)

        # Extract tables
        for table in page.find_tables():
            result["tables"].append({
                "page": page_num + 1,
                "data": table.extract()
            })

        # Extract images
        for img in page.get_images():
            xref = img[0]
            image = doc.extract_image(xref)
            result["images"].append({
                "page": page_num + 1,
                "ext": image["ext"],
                "size": len(image["image"])
            })

    doc.close()
    return result

PDF Merging and Splitting

def merge_pdfs(pdf_list, output_path):
    """Merge multiple PDFs into one."""
    result = fitz.open()

    for pdf_path in pdf_list:
        doc = fitz.open(pdf_path)
        result.insert_pdf(doc)
        doc.close()

    result.save(output_path)
    result.close()

def split_pdf(input_path, output_dir, pages_per_file=10):
    """Split PDF into smaller files."""
    import os
    doc = fitz.open(input_path)

    for start in range(0, doc.page_count, pages_per_file):
        end = min(start + pages_per_file, doc.page_count)

        new_doc = fitz.open()
        new_doc.insert_pdf(doc, from_page=start, to_page=end - 1)

        output_path = os.path.join(
            output_dir,
            f"split_{start + 1}-{end}.pdf"
        )
        new_doc.save(output_path)
        new_doc.close()

    doc.close()

Watermarking

def add_watermark(input_path, output_path, watermark_text):
    """Add text watermark to all pages."""
    doc = fitz.open(input_path)

    for page in doc:
        # Calculate centre position
        rect = page.rect
        text_length = fitz.get_text_length(
            watermark_text,
            fontname="helv",
            fontsize=50
        )

        x = (rect.width - text_length) / 2
        y = rect.height / 2

        # Insert diagonal watermark
        page.insert_text(
            fitz.Point(x, y),
            watermark_text,
            fontsize=50,
            fontname="helv",
            color=(0.8, 0.8, 0.8),
            rotate=45
        )

    doc.save(output_path)
    doc.close()

Performance Optimisation Strategies

Techniques for improving PyMuPDF processing speed and memory usage.

PerformanceMemory ManagementBatch ProcessingSelective LoadingClose documentsUse context managersProcess in chunksLoad only neededpagesPerformanceMemory ManagementBatch ProcessingSelective LoadingClose documentsUse context managersProcess in chunksLoad only neededpages

Memory Management

# Always close documents
doc = fitz.open("large.pdf")
try:
    # Process document
    pass
finally:
    doc.close()

# Use context managers
with fitz.open("document.pdf") as doc:
    # Process document
    pass

# Clear pixmaps when done
pix = page.get_pixmap()
# Use pixmap...
pix = None  # Allow garbage collection

# Process large documents in chunks
def process_large_pdf(path, chunk_size=50):
    doc = fitz.open(path)

    for start in range(0, doc.page_count, chunk_size):
        end = min(start + chunk_size, doc.page_count)

        for page_num in range(start, end):
            page = doc[page_num]
            # Process page
            text = page.get_text()

        # Force garbage collection between chunks
        import gc
        gc.collect()

    doc.close()

Selective Processing

# Only load needed pages
doc = fitz.open("document.pdf")
pages_to_process = [0, 5, 10, 15]

for page_num in pages_to_process:
    page = doc.load_page(page_num)
    # Process page

# Extract specific regions only
rect = fitz.Rect(0, 0, 300, 200)  # Top portion only
text = page.get_text("text", clip=rect)

# Skip image extraction if not needed
text = page.get_text()  # Faster than get_text("dict")

Optimised Saving

# Compress and clean on save
doc.save(
    "output.pdf",
    garbage=4,      # Maximum garbage collection (0-4)
    deflate=True,   # Compress streams
    clean=True      # Clean unused objects
)

# Incremental save for small changes
doc.save("output.pdf", incremental=True)

# Linear save for web viewing
doc.save("output.pdf", linear=True)

# Avoid repeated saves
# Bad: save after each change
for page in doc:
    page.insert_text(...)
    doc.save("output.pdf")  # Slow!

# Good: save once at end
for page in doc:
    page.insert_text(...)
doc.save("output.pdf")  # Fast!

Batch Processing

import concurrent.futures
from pathlib import Path

def process_pdf(pdf_path):
    """Process single PDF file."""
    with fitz.open(pdf_path) as doc:
        text = ""
        for page in doc:
            text += page.get_text()
        return {"path": str(pdf_path), "text": text}

def batch_process(pdf_folder, max_workers=4):
    """Process multiple PDFs in parallel."""
    pdf_files = list(Path(pdf_folder).glob("*.pdf"))
    results = []

    with concurrent.futures.ThreadPoolExecutor(max_workers=max_workers) as executor:
        futures = {
            executor.submit(process_pdf, pdf): pdf
            for pdf in pdf_files
        }

        for future in concurrent.futures.as_completed(futures):
            try:
                result = future.result()
                results.append(result)
            except Exception as e:
                print(f"Error: {e}")

    return results

Caching Strategies

from functools import lru_cache

class PDFProcessor:
    def __init__(self, pdf_path):
        self.doc = fitz.open(pdf_path)
        self._page_cache = {}

    def get_page_text(self, page_num):
        """Cache page text for repeated access."""
        if page_num not in self._page_cache:
            page = self.doc[page_num]
            self._page_cache[page_num] = page.get_text()
        return self._page_cache[page_num]

    def clear_cache(self):
        self._page_cache.clear()

    def close(self):
        self.clear_cache()
        self.doc.close()

Troubleshooting Common PDF Issues

Solutions for frequently encountered problems.

Document Opening Issues

# Handle encrypted PDFs
doc = fitz.open("encrypted.pdf")
if doc.needs_pass:
    if not doc.authenticate("password"):
        raise ValueError("Incorrect password")

# Handle corrupted PDFs
# MuPDF auto-repairs recoverable files on open; FileDataError means unrecoverable.
try:
    doc = fitz.open("corrupted.pdf")
    if doc.is_repaired:           # file had errors that MuPDF fixed on open
        print("PDF was auto-repaired on open")
        doc.save("repaired.pdf", garbage=4, clean=True)  # persist the repair
except fitz.FileDataError as e:
    # Open could not parse the file at all — nothing to recover
    print(f"PDF is corrupted and could not be opened: {e}")

# Handle non-PDF files
try:
    doc = fitz.open("file.xyz")
except Exception as e:
    print(f"Cannot open file: {e}")

Text Extraction Issues

# Empty text from scanned PDFs
text = page.get_text()
if not text.strip():
    print("PDF may be scanned - use OCR")
    # Consider using OCR like pytesseract
    pix = page.get_pixmap()
    # Pass to OCR engine

# Garbled text (encoding issues)
text = page.get_text("text", flags=fitz.TEXT_PRESERVE_WHITESPACE)

# Missing text due to fonts
text_dict = page.get_text("dict")
for block in text_dict["blocks"]:
    if block["type"] == 0:
        for line in block["lines"]:
            for span in line["spans"]:
                if not span["text"] and span.get("font"):
                    print(f"Font issue: {span['font']}")

# Text in wrong order
text = page.get_text("text", sort=True)  # Sort by position

Image Extraction Issues

# No images found
images = page.get_images()
if not images:
    # Check for inline images
    images = page.get_images(full=True)

    # Images may be in form XObjects
    xrefs = page.get_xobjects()

# Extracting masked images
for img in page.get_images(full=True):
    xref = img[0]
    smask = img[1]  # Soft mask reference

    base = doc.extract_image(xref)

    if smask:
        # Apply mask
        mask = doc.extract_image(smask)
        # Combine base and mask...

# Image quality issues
# Use higher resolution
pix = page.get_pixmap(matrix=fitz.Matrix(3, 3))  # 3x zoom

Annotation Issues

# Annotations not visible after save
for annot in page.annots():
    annot.update()  # Must call update!
doc.save("output.pdf")

# Annotations not appearing
# Check annotation flags
annot.set_flags(fitz.PDF_ANNOT_IS_PRINT)  # Make printable
annot.update()

# Redactions not applied
page.add_redact_annot(rect)
page.apply_redactions()  # Must apply!
doc.save("output.pdf")

Memory and Performance Issues

# Out of memory with large PDFs
import gc

doc = fitz.open("huge.pdf")
for i in range(doc.page_count):
    page = doc[i]
    # Process page
    del page
    if i % 100 == 0:
        gc.collect()
doc.close()

# Slow processing
# Avoid: getting full dict for simple text
text_dict = page.get_text("dict")  # Slow

# Prefer: simple text extraction
text = page.get_text()  # Fast

# Avoid: repeated small saves
# Prefer: batch changes then save once

Quick Reference

Essential Operations

Operation Code
Open PDF doc = fitz.open("file.pdf")
Get page page = doc[0]
Extract text text = page.get_text()
Save PDF doc.save("output.pdf")
Close PDF doc.close()
Page count doc.page_count
Add page doc.new_page()
Delete page doc.delete_page(0)
Insert image page.insert_image(rect, filename="img.png")
Add highlight page.add_highlight_annot(rect)
Page to image pix = page.get_pixmap()
Search text rects = page.search_for("text")
Get metadata doc.metadata
Merge PDFs doc1.insert_pdf(doc2)

Common Rectangle Operations

# Create rectangle
rect = fitz.Rect(x0, y0, x1, y1)  # Top-left to bottom-right

# Page rectangle
rect = page.rect

# Normalise rectangle
rect = rect.normalize()

# Check containment
if rect1.contains(rect2):
    pass

# Intersection
rect3 = rect1 & rect2

# Union
rect4 = rect1 | rect2

Colour Specifications

# RGB tuples (0-1 range)
red = (1, 0, 0)
green = (0, 1, 0)
blue = (0, 0, 1)
white = (1, 1, 1)
black = (0, 0, 0)
yellow = (1, 1, 0)

# Colour spaces
fitz.csRGB      # RGB
fitz.csGRAY     # Grayscale
fitz.csCMYK     # CMYK

Page Sizes in Points

Size Width x Height
A4 595 x 842
Letter 612 x 792
Legal 612 x 1008
A3 842 x 1191
A5 420 x 595

Common Issues and Solutions

Issue Cause Solution
Empty text extraction Scanned PDF Use OCR (pytesseract, etc.)
FileDataError on open Corrupted PDF Recoverable files auto-repair on open (check doc.is_repaired, re-save with clean=True); unrecoverable ones raise
Annotations not saved Missing update() Call annot.update() before save
Poor image quality Low resolution Increase pixmap matrix zoom factor
Memory errors Large documents Process in chunks, use context managers
Password error Encrypted PDF Use doc.authenticate(password)
Text in wrong order Complex layout Use sort=True in get_text()
Slow performance Repeated saves Batch changes, save once at end
Missing images Inline images Use get_images(full=True)
Garbled characters Font issues Check font embedding, try different extraction