Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Python PyYAML

YAML parsing and serialisation library for Python with full YAML 1.1 support.

Python PyYAML

YAML parsing and serialisation library for Python with full YAML 1.1 support.

Overview

PyYAML is a YAML parser and emitter for Python. It supports YAML 1.1 specification, including advanced features like custom tags, anchors/aliases, and multi-document streams. PyYAML is widely used for configuration files, data serialisation, and inter-process communication due to YAML's human-readable format.

YAML FeaturesAnchors & AliasesMulti-DocumentCustom TagsFlow/Block StyleLoading Methodssafe_loadBasic Types Onlyfull_loadAll YAML Tagsunsafe_loadArbitrary ObjectsPyYAML OperationsdumploadPython ObjectYAML String/FileYAML FeaturesAnchors & AliasesMulti-DocumentCustom TagsFlow/Block StyleLoading Methodssafe_loadBasic Types Onlyfull_loadAll YAML Tagsunsafe_loadArbitrary ObjectsPyYAML OperationsdumploadPython ObjectYAML String/File

Installation

# Install PyYAML
pip install pyyaml

# Official wheels bundle the LibYAML C bindings; building from source
# against libyaml-dev also enables the fast C loader/dumper

Loading YAML Data

Key Concepts

  • safe_load: Recommended for untrusted input; loads only basic Python types
  • full_load: Supports all standard YAML tags but not arbitrary code
  • unsafe_load: Loads arbitrary Python objects; security risk with untrusted data
  • Loader classes: SafeLoader, FullLoader, UnsafeLoader; yaml.load() requires an explicit Loader= argument since PyYAML 6.0

Basic Loading

import yaml

# Load from string
yaml_string = """
name: Application Config
version: 1.0
debug: true
ports:
  - 8080
  - 8443
database:
  host: localhost
  port: 5432
"""

# Safe load (recommended)
data = yaml.safe_load(yaml_string)
print(data['name'])  # Application Config
print(data['ports'])  # [8080, 8443]

# Load from file
with open('config.yaml', 'r') as f:
    config = yaml.safe_load(f)

# Using explicit loader
data = yaml.load(yaml_string, Loader=yaml.SafeLoader)

Loading Multiple Documents

import yaml

# Multi-document YAML (separated by ---)
multi_doc = """
---
name: Document 1
value: 100
---
name: Document 2
value: 200
---
name: Document 3
value: 300
"""

# Load all documents
documents = list(yaml.safe_load_all(multi_doc))
for doc in documents:
    print(f"{doc['name']}: {doc['value']}")

# Load from file with multiple documents
with open('multi.yaml', 'r') as f:
    for doc in yaml.safe_load_all(f):
        process_document(doc)

# Iterate without loading all into memory
with open('large_multi.yaml', 'r') as f:
    for doc in yaml.safe_load_all(f):
        # Process each document individually
        handle(doc)

Safe Loading (Security)

import yaml

# ALWAYS use safe_load for untrusted input
user_input = get_yaml_from_user()
data = yaml.safe_load(user_input)

# SafeLoader only allows:
# - Scalars: str, int, float, bool, None
# - Collections: list, dict
# - Dates and timestamps

# DANGEROUS - never use with untrusted data
# yaml.unsafe_load(user_input)  # Can execute arbitrary code

# Example of malicious YAML (DO NOT USE):
# !!python/object/apply:os.system ['rm -rf /']

# Custom safe loader with additional types
class ExtendedSafeLoader(yaml.SafeLoader):
    pass

# Add constructors for specific safe types only
def construct_decimal(loader, node):
    value = loader.construct_scalar(node)
    return Decimal(value)

ExtendedSafeLoader.add_constructor(
    '!decimal',
    construct_decimal
)

Dumping YAML Data

Key Concepts

  • dump: Serialise Python object to YAML string
  • safe_dump: Only dump basic Python types (safe for sharing)
  • dump_all: Serialise multiple documents
  • Style options: Control output formatting (flow vs block)

Basic Dumping

import yaml

data = {
    'name': 'MyApp',
    'version': '2.0',
    'settings': {
        'debug': False,
        'log_level': 'INFO',
        'max_connections': 100
    },
    'servers': ['web1', 'web2', 'web3']
}

# Basic dump
yaml_string = yaml.dump(data)
print(yaml_string)

# Safe dump (recommended for data exchange)
yaml_string = yaml.safe_dump(data)

# Dump to file
with open('output.yaml', 'w') as f:
    yaml.safe_dump(data, f)

# With formatting options
yaml_string = yaml.dump(
    data,
    default_flow_style=False,  # Block style (default)
    sort_keys=False,           # Preserve insertion order
    indent=2,                  # Indentation
    width=80,                  # Line width
    allow_unicode=True         # Allow Unicode characters
)

Formatting Options

import yaml

data = {
    'users': [
        {'name': 'Alice', 'age': 30},
        {'name': 'Bob', 'age': 25}
    ],
    'config': {'debug': True, 'timeout': 30}
}

# Block style (human-readable); keys are sorted unless sort_keys=False
print(yaml.dump(data, default_flow_style=False))
# config:
#   debug: true
#   timeout: 30
# users:
# - age: 30
#   name: Alice
# - age: 25
#   name: Bob

# Flow style (compact)
print(yaml.dump(data, default_flow_style=True))
# {config: {debug: true, timeout: 30}, users: [{age: 30, name: Alice}, {age: 25, name: Bob}]}

# Explicit start/end markers
print(yaml.dump(data, explicit_start=True, explicit_end=True))
# ---
# ...

# Custom string style
class LiteralStr(str):
    pass

def literal_str_representer(dumper, data):
    return dumper.represent_scalar('tag:yaml.org,2002:str', data, style='|')

yaml.add_representer(LiteralStr, literal_str_representer)

multiline = LiteralStr("""First line
Second line
Third line""")

print(yaml.dump({'text': multiline}))
# text: |-
#   First line
#   Second line
#   Third line

Dumping Multiple Documents

import yaml

documents = [
    {'type': 'config', 'version': 1},
    {'type': 'data', 'items': [1, 2, 3]},
    {'type': 'metadata', 'author': 'Alice'}
]

# Dump all documents (keys sorted by default)
yaml_string = yaml.dump_all(documents)
print(yaml_string)
# type: config
# version: 1
# ---
# items:
# - 1
# - 2
# - 3
# type: data
# ---
# author: Alice
# type: metadata

# Dump to file
with open('multi_output.yaml', 'w') as f:
    yaml.dump_all(documents, f)

# With explicit document markers
yaml_string = yaml.dump_all(
    documents,
    explicit_start=True,
    explicit_end=True
)

Configuration Files

Key Concepts

  • YAML is ideal for configuration due to readability
  • Support for comments (lost on round-trip with PyYAML)
  • Hierarchical structure maps well to application settings
  • Environment-specific configurations

Application Configuration

import yaml
from pathlib import Path

# config.yaml
"""
app:
  name: MyApplication
  version: 1.0.0

server:
  host: 0.0.0.0
  port: 8080
  workers: 4

database:
  driver: postgresql
  host: localhost
  port: 5432
  name: myapp
  pool_size: 10

logging:
  level: INFO
  format: "%(asctime)s - %(name)s - %(levelname)s - %(message)s"
  handlers:
    - console
    - file

features:
  enable_cache: true
  enable_metrics: true
  rate_limit: 100
"""

class Config:
    def __init__(self, config_path: str):
        self.config_path = Path(config_path)
        self._config = self._load_config()

    def _load_config(self) -> dict:
        if not self.config_path.exists():
            raise FileNotFoundError(f"Config file not found: {self.config_path}")

        with open(self.config_path, 'r') as f:
            return yaml.safe_load(f)

    def get(self, key: str, default=None):
        """Get nested config value using dot notation."""
        keys = key.split('.')
        value = self._config

        for k in keys:
            if isinstance(value, dict):
                value = value.get(k)
            else:
                return default

            if value is None:
                return default

        return value

    def reload(self):
        """Reload configuration from file."""
        self._config = self._load_config()

# Usage
config = Config('config.yaml')
print(config.get('server.host'))  # 0.0.0.0
print(config.get('database.pool_size'))  # 10
print(config.get('nonexistent.key', 'default'))  # default

Environment-Specific Configuration

import yaml
import os
from pathlib import Path

def load_config(env: str = None) -> dict:
    """Load configuration with environment overrides."""
    env = env or os.getenv('APP_ENV', 'development')

    # Load base configuration
    base_path = Path('config/base.yaml')
    with open(base_path, 'r') as f:
        config = yaml.safe_load(f)

    # Load environment-specific overrides
    env_path = Path(f'config/{env}.yaml')
    if env_path.exists():
        with open(env_path, 'r') as f:
            env_config = yaml.safe_load(f)
            config = deep_merge(config, env_config)

    # Override with environment variables
    config = apply_env_overrides(config)

    return config

def deep_merge(base: dict, override: dict) -> dict:
    """Deep merge two dictionaries."""
    result = base.copy()

    for key, value in override.items():
        if key in result and isinstance(result[key], dict) and isinstance(value, dict):
            result[key] = deep_merge(result[key], value)
        else:
            result[key] = value

    return result

def apply_env_overrides(config: dict, prefix: str = 'APP') -> dict:
    """Override config values from environment variables."""
    # APP_DATABASE_HOST -> config['database']['host']
    for key, value in os.environ.items():
        if key.startswith(f'{prefix}_'):
            parts = key[len(prefix)+1:].lower().split('_')
            set_nested(config, parts, value)

    return config

def set_nested(d: dict, keys: list, value):
    """Set a nested dictionary value."""
    for key in keys[:-1]:
        d = d.setdefault(key, {})

    # Type conversion
    if value.lower() in ('true', 'false'):
        value = value.lower() == 'true'
    elif value.isdigit():
        value = int(value)

    d[keys[-1]] = value

Anchors and Aliases

Key Concepts

  • Anchors (&): Mark a node for reuse
  • Aliases (*): Reference an anchored node
  • Merge key (<<): Merge mappings into current mapping
  • Reduces repetition and file size
&anchor (Define)*anchor (Reference)&defaults&lt;&lt;: *defaults(Merge)&anchor (Define)*anchor (Reference)&defaults&lt;&lt;: *defaults(Merge)

Using Anchors and Aliases

import yaml

# YAML with anchors and aliases
yaml_content = """
# Define default settings
defaults: &defaults
  adapter: postgres
  host: localhost
  port: 5432

# Reference with alias
development:
  database:
    <<: *defaults
    database: myapp_dev

test:
  database:
    <<: *defaults
    database: myapp_test

production:
  database:
    <<: *defaults
    host: db.example.com
    database: myapp_prod
"""

config = yaml.safe_load(yaml_content)

print(config['development']['database'])
# {'adapter': 'postgres', 'host': 'localhost', 'port': 5432, 'database': 'myapp_dev'}

print(config['production']['database'])
# {'adapter': 'postgres', 'host': 'db.example.com', 'port': 5432, 'database': 'myapp_prod'}

# Anchors for repeated values
yaml_anchors = """
colours:
  primary: &primary "#3498db"
  secondary: &secondary "#2ecc71"

theme:
  header:
    background: *primary
    text: white
  sidebar:
    background: *secondary
    text: *primary
  footer:
    background: *primary
"""

theme = yaml.safe_load(yaml_anchors)
print(theme['theme']['header']['background'])  # #3498db

Creating Anchors Programmatically

import yaml

# Anchors are resolved during loading
# To preserve them, use custom representation

data = {
    'defaults': {
        'timeout': 30,
        'retries': 3
    },
    'service_a': {
        'name': 'Service A',
        'timeout': 30,
        'retries': 3
    },
    'service_b': {
        'name': 'Service B',
        'timeout': 30,
        'retries': 3
    }
}

# Note: PyYAML does not automatically create anchors
# Anchors are a loading/authoring convenience
# Dumped YAML will expand all references

yaml_string = yaml.dump(data)
print(yaml_string)
# Each service has its own copy of timeout/retries

Custom Representers and Constructors

Key Concepts

  • Representer: Converts Python object to YAML node
  • Constructor: Converts YAML node to Python object
  • Tags: Custom YAML tags for type identification
  • Enables serialisation of custom classes
LoadingConstructorYAML StringYAML NodePython ObjectDumpingRepresenterPython ObjectYAML NodeYAML StringLoadingConstructorYAML StringYAML NodePython ObjectDumpingRepresenterPython ObjectYAML NodeYAML String

Custom Representers

import yaml
from datetime import datetime
from decimal import Decimal
from pathlib import Path

# Custom class
class Person:
    def __init__(self, name: str, age: int, email: str):
        self.name = name
        self.age = age
        self.email = email

    def __repr__(self):
        return f"Person({self.name}, {self.age})"

# Representer function
def person_representer(dumper, person):
    return dumper.represent_mapping(
        '!person',
        {
            'name': person.name,
            'age': person.age,
            'email': person.email
        }
    )

# Register representer
yaml.add_representer(Person, person_representer)

# Now we can dump Person objects
person = Person('Alice', 30, 'alice@example.com')
yaml_string = yaml.dump({'user': person})
print(yaml_string)
# user: !person
#   age: 30
#   email: alice@example.com
#   name: Alice

# Representer for built-in types
def decimal_representer(dumper, value):
    return dumper.represent_scalar('!decimal', str(value))

yaml.add_representer(Decimal, decimal_representer)

# Representer for Path objects
def path_representer(dumper, path):
    return dumper.represent_scalar('!path', str(path))

yaml.add_representer(Path, path_representer)

# Dump data with custom types
data = {
    'price': Decimal('19.99'),
    'config_path': Path('/etc/myapp/config.yaml')
}

print(yaml.dump(data))
# config_path: !path '/etc/myapp/config.yaml'
# price: !decimal '19.99'

Custom Constructors

import yaml
from decimal import Decimal
from pathlib import Path
from datetime import datetime

# Constructor for Person class
def person_constructor(loader, node):
    values = loader.construct_mapping(node)
    return Person(
        name=values['name'],
        age=values['age'],
        email=values['email']
    )

# Register constructor
yaml.add_constructor('!person', person_constructor)

# Now we can load Person objects
yaml_string = """
user: !person
  name: Bob
  age: 25
  email: bob@example.com
"""

data = yaml.load(yaml_string, Loader=yaml.FullLoader)
print(data['user'])  # Person(Bob, 25)

# Constructor for Decimal
def decimal_constructor(loader, node):
    value = loader.construct_scalar(node)
    return Decimal(value)

yaml.add_constructor('!decimal', decimal_constructor)

# Constructor for Path
def path_constructor(loader, node):
    value = loader.construct_scalar(node)
    return Path(value)

yaml.add_constructor('!path', path_constructor)

# Safe loader with custom constructors
class CustomSafeLoader(yaml.SafeLoader):
    pass

CustomSafeLoader.add_constructor('!decimal', decimal_constructor)
CustomSafeLoader.add_constructor('!path', path_constructor)

# Use custom safe loader
yaml_string = """
price: !decimal '29.99'
data_dir: !path '/var/data'
"""

data = yaml.load(yaml_string, Loader=CustomSafeLoader)
print(type(data['price']))  # <class 'decimal.Decimal'>

Multi-Constructor Pattern

import yaml

# Constructor for any Python object in a safe way
class SafeObjectLoader(yaml.SafeLoader):
    pass

# Allow specific classes only
ALLOWED_CLASSES = {
    'myapp.models.User': User,
    'myapp.models.Product': Product,
}

def safe_object_constructor(loader, tag_suffix, node):
    class_name = tag_suffix

    if class_name not in ALLOWED_CLASSES:
        raise yaml.YAMLError(f"Class not allowed: {class_name}")

    cls = ALLOWED_CLASSES[class_name]
    values = loader.construct_mapping(node)
    return cls(**values)

# Register multi-constructor for !python/object: prefix
SafeObjectLoader.add_multi_constructor(
    '!python/object:',
    safe_object_constructor
)

Stream Handling

Key Concepts

  • Stream-based parsing for large files
  • Memory-efficient processing
  • Event-based API for fine-grained control
  • Useful for files that don't fit in memory

Processing Large Files

import yaml

# Stream processing for large files
def process_large_yaml(filepath: str):
    """Process YAML documents one at a time."""
    with open(filepath, 'r') as f:
        for doc in yaml.safe_load_all(f):
            yield doc

# Usage
for document in process_large_yaml('large_file.yaml'):
    process_document(document)

# Event-based parsing for very large files
def count_documents(filepath: str) -> int:
    """Count documents without loading into memory."""
    count = 0
    with open(filepath, 'r') as f:
        for event in yaml.parse(f):
            if isinstance(event, yaml.DocumentStartEvent):
                count += 1
    return count

# Stream output for large data
def stream_to_file(data_generator, filepath: str):
    """Stream multiple documents to file."""
    with open(filepath, 'w') as f:
        for item in data_generator:
            yaml.dump(item, f, explicit_start=True)

Event-Based API

import yaml

yaml_content = """
name: Test
items:
  - one
  - two
"""

# Parse to events
events = list(yaml.parse(yaml_content))
for event in events:
    print(type(event).__name__)

# Output:
# StreamStartEvent
# DocumentStartEvent
# MappingStartEvent
# ScalarEvent
# ScalarEvent
# ScalarEvent
# SequenceStartEvent
# ScalarEvent
# ScalarEvent
# SequenceEndEvent
# MappingEndEvent
# DocumentEndEvent
# StreamEndEvent

# Emit from events (round-trip)
yaml_string = yaml.emit(events)

# Compose to nodes (intermediate representation)
with open('config.yaml', 'r') as f:
    for node in yaml.compose_all(f):
        # Access YAML node structure
        if isinstance(node, yaml.MappingNode):
            for key, value in node.value:
                print(f"Key: {key.value}")

Streaming Writer

import yaml

class YAMLStreamWriter:
    """Write YAML documents incrementally."""

    def __init__(self, filepath: str):
        self.filepath = filepath
        self.file = None
        self.first_document = True

    def __enter__(self):
        self.file = open(self.filepath, 'w')
        return self

    def __exit__(self, *args):
        if self.file:
            self.file.close()

    def write_document(self, data):
        """Write a single YAML document."""
        yaml.dump(
            data,
            self.file,
            explicit_start=True,
            default_flow_style=False
        )
        self.first_document = False

# Usage
with YAMLStreamWriter('output.yaml') as writer:
    for i in range(1000):
        writer.write_document({
            'id': i,
            'data': f'Item {i}'
        })

Integration with Dataclasses

Key Concepts

  • Dataclasses provide structured Python objects
  • Combine YAML's readability with type safety
  • Use custom representers/constructors for round-trips
import yaml
from dataclasses import dataclass, field, asdict
from typing import List, Optional
from datetime import datetime

@dataclass
class DatabaseConfig:
    host: str
    port: int = 5432
    name: str = 'default'
    user: str = 'admin'
    password: str = ''

@dataclass
class ServerConfig:
    host: str = '0.0.0.0'
    port: int = 8080
    workers: int = 4
    debug: bool = False

@dataclass
class AppConfig:
    name: str
    version: str
    server: ServerConfig
    database: DatabaseConfig
    tags: List[str] = field(default_factory=list)
    created_at: Optional[datetime] = None

# Custom loader for dataclasses
class DataclassLoader(yaml.SafeLoader):
    pass

def make_dataclass_constructor(cls):
    def constructor(loader, node):
        values = loader.construct_mapping(node, deep=True)
        # Handle nested dataclasses
        hints = getattr(cls, '__annotations__', {})
        for key, hint in hints.items():
            if hasattr(hint, '__dataclass_fields__') and key in values:
                if isinstance(values[key], dict):
                    values[key] = hint(**values[key])
        return cls(**values)
    return constructor

# Register dataclass constructors
for cls in [DatabaseConfig, ServerConfig, AppConfig]:
    tag = f'!{cls.__name__}'
    DataclassLoader.add_constructor(tag, make_dataclass_constructor(cls))

# Custom dumper for dataclasses
class DataclassDumper(yaml.SafeDumper):
    pass

def dataclass_representer(dumper, data):
    tag = f'!{data.__class__.__name__}'
    return dumper.represent_mapping(tag, asdict(data))

for cls in [DatabaseConfig, ServerConfig, AppConfig]:
    DataclassDumper.add_representer(cls, dataclass_representer)

# Usage
config = AppConfig(
    name='MyApp',
    version='1.0.0',
    server=ServerConfig(port=9000, workers=8),
    database=DatabaseConfig(host='db.example.com', name='production'),
    tags=['production', 'critical']
)

# Dump
yaml_string = yaml.dump(config, Dumper=DataclassDumper)
print(yaml_string)

# Load
loaded_config = yaml.load(yaml_string, Loader=DataclassLoader)
print(loaded_config)

Simple Dataclass Integration

import yaml
from dataclasses import dataclass, asdict, fields
from typing import get_type_hints

@dataclass
class Config:
    name: str
    port: int
    debug: bool = False

def dataclass_from_yaml(cls, yaml_string: str):
    """Load YAML into a dataclass."""
    data = yaml.safe_load(yaml_string)
    return cls(**data)

def dataclass_to_yaml(instance) -> str:
    """Dump a dataclass to YAML."""
    return yaml.safe_dump(asdict(instance))

# Usage
yaml_string = """
name: MyService
port: 8080
debug: true
"""

config = dataclass_from_yaml(Config, yaml_string)
print(config.name)  # MyService

# Dump back
output = dataclass_to_yaml(config)
print(output)

Integration with Pydantic

Key Concepts

  • Pydantic provides validation and serialisation
  • YAML as human-readable configuration source
  • Automatic type coercion and validation
  • Excellent for configuration management
import yaml
from pydantic import BaseModel, Field, field_validator
from typing import List, Optional
from datetime import datetime

class DatabaseSettings(BaseModel):
    host: str
    port: int = Field(default=5432, ge=1, le=65535)
    name: str
    user: str
    password: str = Field(repr=False)
    pool_size: int = Field(default=10, ge=1, le=100)

    @field_validator('host')
    @classmethod
    def validate_host(cls, v):
        if not v:
            raise ValueError('Host cannot be empty')
        return v

class ServerSettings(BaseModel):
    host: str = '0.0.0.0'
    port: int = Field(default=8080, ge=1, le=65535)
    workers: int = Field(default=4, ge=1)
    timeout: int = Field(default=30, ge=1)

class AppSettings(BaseModel):
    name: str
    version: str
    environment: str = 'development'
    server: ServerSettings
    database: DatabaseSettings
    features: List[str] = []
    metadata: Optional[dict] = None

def load_yaml_config(path: str, model: type[BaseModel]) -> BaseModel:
    """Load YAML file into Pydantic model with validation."""
    with open(path, 'r') as f:
        data = yaml.safe_load(f)
    return model.model_validate(data)

def save_yaml_config(model: BaseModel, path: str):
    """Save Pydantic model to YAML file."""
    with open(path, 'w') as f:
        yaml.safe_dump(
            model.model_dump(mode='json'),
            f,
            default_flow_style=False,
            sort_keys=False
        )

# Usage
yaml_config = """
name: MyApplication
version: 2.0.0
environment: production
server:
  port: 9000
  workers: 8
  timeout: 60
database:
  host: db.example.com
  port: 5432
  name: myapp_prod
  user: admin
  password: secret123
  pool_size: 20
features:
  - authentication
  - caching
  - metrics
"""

# Load and validate
settings = AppSettings.model_validate(yaml.safe_load(yaml_config))
print(settings.server.port)  # 9000
print(settings.database.pool_size)  # 20

# Validation errors are caught
try:
    invalid_yaml = """
    name: Test
    version: 1.0
    server:
      port: 99999  # Invalid port
    database:
      host: ""  # Invalid empty host
      name: test
      user: admin
      password: pass
    """
    AppSettings.model_validate(yaml.safe_load(invalid_yaml))
except Exception as e:
    print(f"Validation error: {e}")

Pydantic Settings with YAML

from pydantic_settings import BaseSettings
from pydantic import Field
import yaml
from pathlib import Path

class Settings(BaseSettings):
    """Application settings from YAML and environment."""

    app_name: str = 'DefaultApp'
    debug: bool = False
    database_url: str
    secret_key: str = Field(repr=False)

    @classmethod
    def from_yaml(cls, path: str) -> 'Settings':
        """Load settings from YAML file."""
        config_path = Path(path)

        if config_path.exists():
            with open(config_path, 'r') as f:
                yaml_data = yaml.safe_load(f)
            return cls(**yaml_data)

        # Fall back to environment variables
        return cls()

# Usage
settings = Settings.from_yaml('config.yaml')

Performance Considerations

Key Concepts

  • LibYAML C extension significantly faster
  • Choose appropriate loader for use case
  • Stream processing for large files
  • Caching for repeated loads

Using LibYAML (C Extension)

import yaml

# Check if LibYAML is available
print(f"LibYAML available: {yaml.__with_libyaml__}")

# Use C-based loader/dumper for performance
if yaml.__with_libyaml__:
    # 5-10x faster than pure Python
    Loader = yaml.CSafeLoader
    Dumper = yaml.CSafeDumper
else:
    Loader = yaml.SafeLoader
    Dumper = yaml.SafeDumper

# Load with C loader
with open('large_config.yaml', 'r') as f:
    data = yaml.load(f, Loader=Loader)

# Dump with C dumper
yaml_string = yaml.dump(data, Dumper=Dumper)

Performance Comparison

import yaml
import time

def benchmark_loader(yaml_string: str, iterations: int = 1000):
    """Benchmark different loaders."""
    loaders = [
        ('SafeLoader', yaml.SafeLoader),
        ('FullLoader', yaml.FullLoader),
    ]

    if yaml.__with_libyaml__:
        loaders.extend([
            ('CSafeLoader', yaml.CSafeLoader),
            ('CFullLoader', yaml.CFullLoader),
        ])

    for name, loader in loaders:
        start = time.time()
        for _ in range(iterations):
            yaml.load(yaml_string, Loader=loader)
        elapsed = time.time() - start
        print(f"{name}: {elapsed:.3f}s ({iterations/elapsed:.0f} ops/sec)")

# Sample YAML
yaml_string = """
config:
  database:
    host: localhost
    port: 5432
  servers:
    - name: web1
      port: 8080
    - name: web2
      port: 8081
"""

benchmark_loader(yaml_string)
# CSafeLoader is typically 5-10x faster than SafeLoader

Caching Configurations

import yaml
from functools import lru_cache
from pathlib import Path
import hashlib

@lru_cache(maxsize=32)
def load_config_cached(filepath: str) -> dict:
    """Load and cache configuration."""
    with open(filepath, 'r') as f:
        return yaml.safe_load(f)

class ConfigCache:
    """Configuration cache with file change detection."""

    def __init__(self):
        self._cache = {}
        self._hashes = {}

    def load(self, filepath: str) -> dict:
        path = Path(filepath)

        # Calculate file hash
        content = path.read_bytes()
        file_hash = hashlib.md5(content).hexdigest()

        # Return cached if unchanged
        if filepath in self._cache and self._hashes.get(filepath) == file_hash:
            return self._cache[filepath]

        # Load and cache
        config = yaml.safe_load(content)
        self._cache[filepath] = config
        self._hashes[filepath] = file_hash

        return config

    def clear(self):
        self._cache.clear()
        self._hashes.clear()

# Usage
cache = ConfigCache()
config = cache.load('config.yaml')

Memory-Efficient Processing

import yaml

def process_large_yaml_file(filepath: str):
    """Process large YAML file without loading entirely into memory."""

    # For multi-document files
    with open(filepath, 'r') as f:
        for doc in yaml.safe_load_all(f):
            # Process each document individually
            result = process_document(doc)
            yield result
            # Document is garbage collected after processing

def filter_yaml_documents(input_path: str, output_path: str, predicate):
    """Filter documents from a YAML file."""
    with open(input_path, 'r') as infile, open(output_path, 'w') as outfile:
        for doc in yaml.safe_load_all(infile):
            if predicate(doc):
                yaml.dump(doc, outfile, explicit_start=True)

Quick Reference

Operation Code Description
Loading yaml.safe_load(string) Load YAML from string (safe)
yaml.safe_load(file) Load YAML from file object
yaml.safe_load_all(string) Load multiple documents
yaml.load(string, Loader=yaml.FullLoader) Load with specific loader
Dumping yaml.dump(data) Dump to YAML string
yaml.safe_dump(data) Dump basic types only
yaml.dump(data, file) Dump to file object
yaml.dump_all(docs) Dump multiple documents
Formatting yaml.dump(data, default_flow_style=False) Block style output
yaml.dump(data, indent=2) Custom indentation
yaml.dump(data, sort_keys=False) Preserve key order
yaml.dump(data, allow_unicode=True) Allow Unicode chars
Custom Types yaml.add_representer(cls, func) Add custom representer
yaml.add_constructor(tag, func) Add custom constructor
Performance yaml.CSafeLoader C-based loader (faster)
yaml.CSafeDumper C-based dumper (faster)
Security yaml.safe_load() Safe for untrusted input
Never use yaml.unsafe_load() With untrusted data

Common Issues and Solutions

Issue Solution
YAMLError: could not determine constructor Use safe_load() or register custom constructor
Duplicate keys silently overwritten PyYAML keeps the last value without warning; use ruamel.yaml to detect duplicates
TypeError: cannot serialize Add custom representer for the type
UnicodeDecodeError Specify encoding: open(file, encoding='utf-8')
Memory error with large files Use safe_load_all() for streaming or process in chunks
Comments lost after round-trip PyYAML doesn't preserve comments; use ruamel.yaml
Wrong boolean parsing (yes/no) YAML 1.1 treats these as booleans; quote strings
Anchors not created in output Anchors are resolved on load; manually author YAML
None becomes string 'null' This is correct YAML; use None in Python
Numbers with leading zeros Quote them: '007' to prevent octal interpretation
Multiline strings formatting Use | (literal) or > (folded) style
Float precision loss Use Decimal with custom representer/constructor
!!python/object security risk Never use unsafe_load() with untrusted data
Dates parsed as strings safe_load parses ISO dates/timestamps natively; quote values to keep them as strings
C loader not available Use an official wheel, or build PyYAML from source with libyaml-dev installed

Related Topics

The following topics complement PyYAML development:

  1. Pydantic - Data validation and settings management for type-safe YAML configuration
  2. Python Patterns - Design patterns for configuration management and dependency injection
  3. FastAPI - Web framework that benefits from YAML configuration with Pydantic
  4. Docker/Kubernetes - Container orchestration using YAML manifests
  5. Ansible - Infrastructure automation heavily using YAML playbooks
  6. Jinja2 - Template engine for generating dynamic YAML files