Python PyYAML
YAML parsing and serialisation library for Python with full YAML 1.1 support.
Python PyYAML
YAML parsing and serialisation library for Python with full YAML 1.1 support.
Overview
PyYAML is a YAML parser and emitter for Python. It supports YAML 1.1 specification, including advanced features like custom tags, anchors/aliases, and multi-document streams. PyYAML is widely used for configuration files, data serialisation, and inter-process communication due to YAML's human-readable format.
flowchart TB
subgraph "PyYAML Operations"
A[Python Object] -->|dump| B[YAML String/File]
B -->|load| A
end
subgraph "Loading Methods"
C[safe_load] --> D[Basic Types Only]
E[full_load] --> F[All YAML Tags]
G[unsafe_load] --> H[Arbitrary Objects]
end
subgraph "YAML Features"
I[Anchors & Aliases]
J[Multi-Document]
K[Custom Tags]
L[Flow/Block Style]
end
Installation
# Install PyYAML
pip install pyyaml
# Official wheels bundle the LibYAML C bindings; building from source
# against libyaml-dev also enables the fast C loader/dumper
Loading YAML Data
Key Concepts
- safe_load: Recommended for untrusted input; loads only basic Python types
- full_load: Supports all standard YAML tags but not arbitrary code
- unsafe_load: Loads arbitrary Python objects; security risk with untrusted data
- Loader classes: SafeLoader, FullLoader, UnsafeLoader;
yaml.load()requires an explicitLoader=argument since PyYAML 6.0
Basic Loading
import yaml
# Load from string
yaml_string = """
name: Application Config
version: 1.0
debug: true
ports:
- 8080
- 8443
database:
host: localhost
port: 5432
"""
# Safe load (recommended)
data = yaml.safe_load(yaml_string)
print(data['name']) # Application Config
print(data['ports']) # [8080, 8443]
# Load from file
with open('config.yaml', 'r') as f:
config = yaml.safe_load(f)
# Using explicit loader
data = yaml.load(yaml_string, Loader=yaml.SafeLoader)
Loading Multiple Documents
import yaml
# Multi-document YAML (separated by ---)
multi_doc = """
---
name: Document 1
value: 100
---
name: Document 2
value: 200
---
name: Document 3
value: 300
"""
# Load all documents
documents = list(yaml.safe_load_all(multi_doc))
for doc in documents:
print(f"{doc['name']}: {doc['value']}")
# Load from file with multiple documents
with open('multi.yaml', 'r') as f:
for doc in yaml.safe_load_all(f):
process_document(doc)
# Iterate without loading all into memory
with open('large_multi.yaml', 'r') as f:
for doc in yaml.safe_load_all(f):
# Process each document individually
handle(doc)
Safe Loading (Security)
import yaml
# ALWAYS use safe_load for untrusted input
user_input = get_yaml_from_user()
data = yaml.safe_load(user_input)
# SafeLoader only allows:
# - Scalars: str, int, float, bool, None
# - Collections: list, dict
# - Dates and timestamps
# DANGEROUS - never use with untrusted data
# yaml.unsafe_load(user_input) # Can execute arbitrary code
# Example of malicious YAML (DO NOT USE):
# !!python/object/apply:os.system ['rm -rf /']
# Custom safe loader with additional types
class ExtendedSafeLoader(yaml.SafeLoader):
pass
# Add constructors for specific safe types only
def construct_decimal(loader, node):
value = loader.construct_scalar(node)
return Decimal(value)
ExtendedSafeLoader.add_constructor(
'!decimal',
construct_decimal
)
Dumping YAML Data
Key Concepts
- dump: Serialise Python object to YAML string
- safe_dump: Only dump basic Python types (safe for sharing)
- dump_all: Serialise multiple documents
- Style options: Control output formatting (flow vs block)
Basic Dumping
import yaml
data = {
'name': 'MyApp',
'version': '2.0',
'settings': {
'debug': False,
'log_level': 'INFO',
'max_connections': 100
},
'servers': ['web1', 'web2', 'web3']
}
# Basic dump
yaml_string = yaml.dump(data)
print(yaml_string)
# Safe dump (recommended for data exchange)
yaml_string = yaml.safe_dump(data)
# Dump to file
with open('output.yaml', 'w') as f:
yaml.safe_dump(data, f)
# With formatting options
yaml_string = yaml.dump(
data,
default_flow_style=False, # Block style (default)
sort_keys=False, # Preserve insertion order
indent=2, # Indentation
width=80, # Line width
allow_unicode=True # Allow Unicode characters
)
Formatting Options
import yaml
data = {
'users': [
{'name': 'Alice', 'age': 30},
{'name': 'Bob', 'age': 25}
],
'config': {'debug': True, 'timeout': 30}
}
# Block style (human-readable); keys are sorted unless sort_keys=False
print(yaml.dump(data, default_flow_style=False))
# config:
# debug: true
# timeout: 30
# users:
# - age: 30
# name: Alice
# - age: 25
# name: Bob
# Flow style (compact)
print(yaml.dump(data, default_flow_style=True))
# {config: {debug: true, timeout: 30}, users: [{age: 30, name: Alice}, {age: 25, name: Bob}]}
# Explicit start/end markers
print(yaml.dump(data, explicit_start=True, explicit_end=True))
# ---
# ...
# Custom string style
class LiteralStr(str):
pass
def literal_str_representer(dumper, data):
return dumper.represent_scalar('tag:yaml.org,2002:str', data, style='|')
yaml.add_representer(LiteralStr, literal_str_representer)
multiline = LiteralStr("""First line
Second line
Third line""")
print(yaml.dump({'text': multiline}))
# text: |-
# First line
# Second line
# Third line
Dumping Multiple Documents
import yaml
documents = [
{'type': 'config', 'version': 1},
{'type': 'data', 'items': [1, 2, 3]},
{'type': 'metadata', 'author': 'Alice'}
]
# Dump all documents (keys sorted by default)
yaml_string = yaml.dump_all(documents)
print(yaml_string)
# type: config
# version: 1
# ---
# items:
# - 1
# - 2
# - 3
# type: data
# ---
# author: Alice
# type: metadata
# Dump to file
with open('multi_output.yaml', 'w') as f:
yaml.dump_all(documents, f)
# With explicit document markers
yaml_string = yaml.dump_all(
documents,
explicit_start=True,
explicit_end=True
)
Configuration Files
Key Concepts
- YAML is ideal for configuration due to readability
- Support for comments (lost on round-trip with PyYAML)
- Hierarchical structure maps well to application settings
- Environment-specific configurations
Application Configuration
import yaml
from pathlib import Path
# config.yaml
"""
app:
name: MyApplication
version: 1.0.0
server:
host: 0.0.0.0
port: 8080
workers: 4
database:
driver: postgresql
host: localhost
port: 5432
name: myapp
pool_size: 10
logging:
level: INFO
format: "%(asctime)s - %(name)s - %(levelname)s - %(message)s"
handlers:
- console
- file
features:
enable_cache: true
enable_metrics: true
rate_limit: 100
"""
class Config:
def __init__(self, config_path: str):
self.config_path = Path(config_path)
self._config = self._load_config()
def _load_config(self) -> dict:
if not self.config_path.exists():
raise FileNotFoundError(f"Config file not found: {self.config_path}")
with open(self.config_path, 'r') as f:
return yaml.safe_load(f)
def get(self, key: str, default=None):
"""Get nested config value using dot notation."""
keys = key.split('.')
value = self._config
for k in keys:
if isinstance(value, dict):
value = value.get(k)
else:
return default
if value is None:
return default
return value
def reload(self):
"""Reload configuration from file."""
self._config = self._load_config()
# Usage
config = Config('config.yaml')
print(config.get('server.host')) # 0.0.0.0
print(config.get('database.pool_size')) # 10
print(config.get('nonexistent.key', 'default')) # default
Environment-Specific Configuration
import yaml
import os
from pathlib import Path
def load_config(env: str = None) -> dict:
"""Load configuration with environment overrides."""
env = env or os.getenv('APP_ENV', 'development')
# Load base configuration
base_path = Path('config/base.yaml')
with open(base_path, 'r') as f:
config = yaml.safe_load(f)
# Load environment-specific overrides
env_path = Path(f'config/{env}.yaml')
if env_path.exists():
with open(env_path, 'r') as f:
env_config = yaml.safe_load(f)
config = deep_merge(config, env_config)
# Override with environment variables
config = apply_env_overrides(config)
return config
def deep_merge(base: dict, override: dict) -> dict:
"""Deep merge two dictionaries."""
result = base.copy()
for key, value in override.items():
if key in result and isinstance(result[key], dict) and isinstance(value, dict):
result[key] = deep_merge(result[key], value)
else:
result[key] = value
return result
def apply_env_overrides(config: dict, prefix: str = 'APP') -> dict:
"""Override config values from environment variables."""
# APP_DATABASE_HOST -> config['database']['host']
for key, value in os.environ.items():
if key.startswith(f'{prefix}_'):
parts = key[len(prefix)+1:].lower().split('_')
set_nested(config, parts, value)
return config
def set_nested(d: dict, keys: list, value):
"""Set a nested dictionary value."""
for key in keys[:-1]:
d = d.setdefault(key, {})
# Type conversion
if value.lower() in ('true', 'false'):
value = value.lower() == 'true'
elif value.isdigit():
value = int(value)
d[keys[-1]] = value
Anchors and Aliases
Key Concepts
- Anchors (&): Mark a node for reuse
- Aliases (*): Reference an anchored node
- Merge key (<<): Merge mappings into current mapping
- Reduces repetition and file size
flowchart LR
A["&anchor (Define)"] --> B["*anchor (Reference)"]
C["&defaults"] --> D["<<: *defaults (Merge)"]
Using Anchors and Aliases
import yaml
# YAML with anchors and aliases
yaml_content = """
# Define default settings
defaults: &defaults
adapter: postgres
host: localhost
port: 5432
# Reference with alias
development:
database:
<<: *defaults
database: myapp_dev
test:
database:
<<: *defaults
database: myapp_test
production:
database:
<<: *defaults
host: db.example.com
database: myapp_prod
"""
config = yaml.safe_load(yaml_content)
print(config['development']['database'])
# {'adapter': 'postgres', 'host': 'localhost', 'port': 5432, 'database': 'myapp_dev'}
print(config['production']['database'])
# {'adapter': 'postgres', 'host': 'db.example.com', 'port': 5432, 'database': 'myapp_prod'}
# Anchors for repeated values
yaml_anchors = """
colours:
primary: &primary "#3498db"
secondary: &secondary "#2ecc71"
theme:
header:
background: *primary
text: white
sidebar:
background: *secondary
text: *primary
footer:
background: *primary
"""
theme = yaml.safe_load(yaml_anchors)
print(theme['theme']['header']['background']) # #3498db
Creating Anchors Programmatically
import yaml
# Anchors are resolved during loading
# To preserve them, use custom representation
data = {
'defaults': {
'timeout': 30,
'retries': 3
},
'service_a': {
'name': 'Service A',
'timeout': 30,
'retries': 3
},
'service_b': {
'name': 'Service B',
'timeout': 30,
'retries': 3
}
}
# Note: PyYAML does not automatically create anchors
# Anchors are a loading/authoring convenience
# Dumped YAML will expand all references
yaml_string = yaml.dump(data)
print(yaml_string)
# Each service has its own copy of timeout/retries
Custom Representers and Constructors
Key Concepts
- Representer: Converts Python object to YAML node
- Constructor: Converts YAML node to Python object
- Tags: Custom YAML tags for type identification
- Enables serialisation of custom classes
flowchart LR
subgraph "Dumping"
A[Python Object] -->|Representer| B[YAML Node]
B --> C[YAML String]
end
subgraph "Loading"
D[YAML String] --> E[YAML Node]
E -->|Constructor| F[Python Object]
end
Custom Representers
import yaml
from datetime import datetime
from decimal import Decimal
from pathlib import Path
# Custom class
class Person:
def __init__(self, name: str, age: int, email: str):
self.name = name
self.age = age
self.email = email
def __repr__(self):
return f"Person({self.name}, {self.age})"
# Representer function
def person_representer(dumper, person):
return dumper.represent_mapping(
'!person',
{
'name': person.name,
'age': person.age,
'email': person.email
}
)
# Register representer
yaml.add_representer(Person, person_representer)
# Now we can dump Person objects
person = Person('Alice', 30, 'alice@example.com')
yaml_string = yaml.dump({'user': person})
print(yaml_string)
# user: !person
# age: 30
# email: alice@example.com
# name: Alice
# Representer for built-in types
def decimal_representer(dumper, value):
return dumper.represent_scalar('!decimal', str(value))
yaml.add_representer(Decimal, decimal_representer)
# Representer for Path objects
def path_representer(dumper, path):
return dumper.represent_scalar('!path', str(path))
yaml.add_representer(Path, path_representer)
# Dump data with custom types
data = {
'price': Decimal('19.99'),
'config_path': Path('/etc/myapp/config.yaml')
}
print(yaml.dump(data))
# config_path: !path '/etc/myapp/config.yaml'
# price: !decimal '19.99'
Custom Constructors
import yaml
from decimal import Decimal
from pathlib import Path
from datetime import datetime
# Constructor for Person class
def person_constructor(loader, node):
values = loader.construct_mapping(node)
return Person(
name=values['name'],
age=values['age'],
email=values['email']
)
# Register constructor
yaml.add_constructor('!person', person_constructor)
# Now we can load Person objects
yaml_string = """
user: !person
name: Bob
age: 25
email: bob@example.com
"""
data = yaml.load(yaml_string, Loader=yaml.FullLoader)
print(data['user']) # Person(Bob, 25)
# Constructor for Decimal
def decimal_constructor(loader, node):
value = loader.construct_scalar(node)
return Decimal(value)
yaml.add_constructor('!decimal', decimal_constructor)
# Constructor for Path
def path_constructor(loader, node):
value = loader.construct_scalar(node)
return Path(value)
yaml.add_constructor('!path', path_constructor)
# Safe loader with custom constructors
class CustomSafeLoader(yaml.SafeLoader):
pass
CustomSafeLoader.add_constructor('!decimal', decimal_constructor)
CustomSafeLoader.add_constructor('!path', path_constructor)
# Use custom safe loader
yaml_string = """
price: !decimal '29.99'
data_dir: !path '/var/data'
"""
data = yaml.load(yaml_string, Loader=CustomSafeLoader)
print(type(data['price'])) # <class 'decimal.Decimal'>
Multi-Constructor Pattern
import yaml
# Constructor for any Python object in a safe way
class SafeObjectLoader(yaml.SafeLoader):
pass
# Allow specific classes only
ALLOWED_CLASSES = {
'myapp.models.User': User,
'myapp.models.Product': Product,
}
def safe_object_constructor(loader, tag_suffix, node):
class_name = tag_suffix
if class_name not in ALLOWED_CLASSES:
raise yaml.YAMLError(f"Class not allowed: {class_name}")
cls = ALLOWED_CLASSES[class_name]
values = loader.construct_mapping(node)
return cls(**values)
# Register multi-constructor for !python/object: prefix
SafeObjectLoader.add_multi_constructor(
'!python/object:',
safe_object_constructor
)
Stream Handling
Key Concepts
- Stream-based parsing for large files
- Memory-efficient processing
- Event-based API for fine-grained control
- Useful for files that don't fit in memory
Processing Large Files
import yaml
# Stream processing for large files
def process_large_yaml(filepath: str):
"""Process YAML documents one at a time."""
with open(filepath, 'r') as f:
for doc in yaml.safe_load_all(f):
yield doc
# Usage
for document in process_large_yaml('large_file.yaml'):
process_document(document)
# Event-based parsing for very large files
def count_documents(filepath: str) -> int:
"""Count documents without loading into memory."""
count = 0
with open(filepath, 'r') as f:
for event in yaml.parse(f):
if isinstance(event, yaml.DocumentStartEvent):
count += 1
return count
# Stream output for large data
def stream_to_file(data_generator, filepath: str):
"""Stream multiple documents to file."""
with open(filepath, 'w') as f:
for item in data_generator:
yaml.dump(item, f, explicit_start=True)
Event-Based API
import yaml
yaml_content = """
name: Test
items:
- one
- two
"""
# Parse to events
events = list(yaml.parse(yaml_content))
for event in events:
print(type(event).__name__)
# Output:
# StreamStartEvent
# DocumentStartEvent
# MappingStartEvent
# ScalarEvent
# ScalarEvent
# ScalarEvent
# SequenceStartEvent
# ScalarEvent
# ScalarEvent
# SequenceEndEvent
# MappingEndEvent
# DocumentEndEvent
# StreamEndEvent
# Emit from events (round-trip)
yaml_string = yaml.emit(events)
# Compose to nodes (intermediate representation)
with open('config.yaml', 'r') as f:
for node in yaml.compose_all(f):
# Access YAML node structure
if isinstance(node, yaml.MappingNode):
for key, value in node.value:
print(f"Key: {key.value}")
Streaming Writer
import yaml
class YAMLStreamWriter:
"""Write YAML documents incrementally."""
def __init__(self, filepath: str):
self.filepath = filepath
self.file = None
self.first_document = True
def __enter__(self):
self.file = open(self.filepath, 'w')
return self
def __exit__(self, *args):
if self.file:
self.file.close()
def write_document(self, data):
"""Write a single YAML document."""
yaml.dump(
data,
self.file,
explicit_start=True,
default_flow_style=False
)
self.first_document = False
# Usage
with YAMLStreamWriter('output.yaml') as writer:
for i in range(1000):
writer.write_document({
'id': i,
'data': f'Item {i}'
})
Integration with Dataclasses
Key Concepts
- Dataclasses provide structured Python objects
- Combine YAML's readability with type safety
- Use custom representers/constructors for round-trips
import yaml
from dataclasses import dataclass, field, asdict
from typing import List, Optional
from datetime import datetime
@dataclass
class DatabaseConfig:
host: str
port: int = 5432
name: str = 'default'
user: str = 'admin'
password: str = ''
@dataclass
class ServerConfig:
host: str = '0.0.0.0'
port: int = 8080
workers: int = 4
debug: bool = False
@dataclass
class AppConfig:
name: str
version: str
server: ServerConfig
database: DatabaseConfig
tags: List[str] = field(default_factory=list)
created_at: Optional[datetime] = None
# Custom loader for dataclasses
class DataclassLoader(yaml.SafeLoader):
pass
def make_dataclass_constructor(cls):
def constructor(loader, node):
values = loader.construct_mapping(node, deep=True)
# Handle nested dataclasses
hints = getattr(cls, '__annotations__', {})
for key, hint in hints.items():
if hasattr(hint, '__dataclass_fields__') and key in values:
if isinstance(values[key], dict):
values[key] = hint(**values[key])
return cls(**values)
return constructor
# Register dataclass constructors
for cls in [DatabaseConfig, ServerConfig, AppConfig]:
tag = f'!{cls.__name__}'
DataclassLoader.add_constructor(tag, make_dataclass_constructor(cls))
# Custom dumper for dataclasses
class DataclassDumper(yaml.SafeDumper):
pass
def dataclass_representer(dumper, data):
tag = f'!{data.__class__.__name__}'
return dumper.represent_mapping(tag, asdict(data))
for cls in [DatabaseConfig, ServerConfig, AppConfig]:
DataclassDumper.add_representer(cls, dataclass_representer)
# Usage
config = AppConfig(
name='MyApp',
version='1.0.0',
server=ServerConfig(port=9000, workers=8),
database=DatabaseConfig(host='db.example.com', name='production'),
tags=['production', 'critical']
)
# Dump
yaml_string = yaml.dump(config, Dumper=DataclassDumper)
print(yaml_string)
# Load
loaded_config = yaml.load(yaml_string, Loader=DataclassLoader)
print(loaded_config)
Simple Dataclass Integration
import yaml
from dataclasses import dataclass, asdict, fields
from typing import get_type_hints
@dataclass
class Config:
name: str
port: int
debug: bool = False
def dataclass_from_yaml(cls, yaml_string: str):
"""Load YAML into a dataclass."""
data = yaml.safe_load(yaml_string)
return cls(**data)
def dataclass_to_yaml(instance) -> str:
"""Dump a dataclass to YAML."""
return yaml.safe_dump(asdict(instance))
# Usage
yaml_string = """
name: MyService
port: 8080
debug: true
"""
config = dataclass_from_yaml(Config, yaml_string)
print(config.name) # MyService
# Dump back
output = dataclass_to_yaml(config)
print(output)
Integration with Pydantic
Key Concepts
- Pydantic provides validation and serialisation
- YAML as human-readable configuration source
- Automatic type coercion and validation
- Excellent for configuration management
import yaml
from pydantic import BaseModel, Field, field_validator
from typing import List, Optional
from datetime import datetime
class DatabaseSettings(BaseModel):
host: str
port: int = Field(default=5432, ge=1, le=65535)
name: str
user: str
password: str = Field(repr=False)
pool_size: int = Field(default=10, ge=1, le=100)
@field_validator('host')
@classmethod
def validate_host(cls, v):
if not v:
raise ValueError('Host cannot be empty')
return v
class ServerSettings(BaseModel):
host: str = '0.0.0.0'
port: int = Field(default=8080, ge=1, le=65535)
workers: int = Field(default=4, ge=1)
timeout: int = Field(default=30, ge=1)
class AppSettings(BaseModel):
name: str
version: str
environment: str = 'development'
server: ServerSettings
database: DatabaseSettings
features: List[str] = []
metadata: Optional[dict] = None
def load_yaml_config(path: str, model: type[BaseModel]) -> BaseModel:
"""Load YAML file into Pydantic model with validation."""
with open(path, 'r') as f:
data = yaml.safe_load(f)
return model.model_validate(data)
def save_yaml_config(model: BaseModel, path: str):
"""Save Pydantic model to YAML file."""
with open(path, 'w') as f:
yaml.safe_dump(
model.model_dump(mode='json'),
f,
default_flow_style=False,
sort_keys=False
)
# Usage
yaml_config = """
name: MyApplication
version: 2.0.0
environment: production
server:
port: 9000
workers: 8
timeout: 60
database:
host: db.example.com
port: 5432
name: myapp_prod
user: admin
password: secret123
pool_size: 20
features:
- authentication
- caching
- metrics
"""
# Load and validate
settings = AppSettings.model_validate(yaml.safe_load(yaml_config))
print(settings.server.port) # 9000
print(settings.database.pool_size) # 20
# Validation errors are caught
try:
invalid_yaml = """
name: Test
version: 1.0
server:
port: 99999 # Invalid port
database:
host: "" # Invalid empty host
name: test
user: admin
password: pass
"""
AppSettings.model_validate(yaml.safe_load(invalid_yaml))
except Exception as e:
print(f"Validation error: {e}")
Pydantic Settings with YAML
from pydantic_settings import BaseSettings
from pydantic import Field
import yaml
from pathlib import Path
class Settings(BaseSettings):
"""Application settings from YAML and environment."""
app_name: str = 'DefaultApp'
debug: bool = False
database_url: str
secret_key: str = Field(repr=False)
@classmethod
def from_yaml(cls, path: str) -> 'Settings':
"""Load settings from YAML file."""
config_path = Path(path)
if config_path.exists():
with open(config_path, 'r') as f:
yaml_data = yaml.safe_load(f)
return cls(**yaml_data)
# Fall back to environment variables
return cls()
# Usage
settings = Settings.from_yaml('config.yaml')
Performance Considerations
Key Concepts
- LibYAML C extension significantly faster
- Choose appropriate loader for use case
- Stream processing for large files
- Caching for repeated loads
Using LibYAML (C Extension)
import yaml
# Check if LibYAML is available
print(f"LibYAML available: {yaml.__with_libyaml__}")
# Use C-based loader/dumper for performance
if yaml.__with_libyaml__:
# 5-10x faster than pure Python
Loader = yaml.CSafeLoader
Dumper = yaml.CSafeDumper
else:
Loader = yaml.SafeLoader
Dumper = yaml.SafeDumper
# Load with C loader
with open('large_config.yaml', 'r') as f:
data = yaml.load(f, Loader=Loader)
# Dump with C dumper
yaml_string = yaml.dump(data, Dumper=Dumper)
Performance Comparison
import yaml
import time
def benchmark_loader(yaml_string: str, iterations: int = 1000):
"""Benchmark different loaders."""
loaders = [
('SafeLoader', yaml.SafeLoader),
('FullLoader', yaml.FullLoader),
]
if yaml.__with_libyaml__:
loaders.extend([
('CSafeLoader', yaml.CSafeLoader),
('CFullLoader', yaml.CFullLoader),
])
for name, loader in loaders:
start = time.time()
for _ in range(iterations):
yaml.load(yaml_string, Loader=loader)
elapsed = time.time() - start
print(f"{name}: {elapsed:.3f}s ({iterations/elapsed:.0f} ops/sec)")
# Sample YAML
yaml_string = """
config:
database:
host: localhost
port: 5432
servers:
- name: web1
port: 8080
- name: web2
port: 8081
"""
benchmark_loader(yaml_string)
# CSafeLoader is typically 5-10x faster than SafeLoader
Caching Configurations
import yaml
from functools import lru_cache
from pathlib import Path
import hashlib
@lru_cache(maxsize=32)
def load_config_cached(filepath: str) -> dict:
"""Load and cache configuration."""
with open(filepath, 'r') as f:
return yaml.safe_load(f)
class ConfigCache:
"""Configuration cache with file change detection."""
def __init__(self):
self._cache = {}
self._hashes = {}
def load(self, filepath: str) -> dict:
path = Path(filepath)
# Calculate file hash
content = path.read_bytes()
file_hash = hashlib.md5(content).hexdigest()
# Return cached if unchanged
if filepath in self._cache and self._hashes.get(filepath) == file_hash:
return self._cache[filepath]
# Load and cache
config = yaml.safe_load(content)
self._cache[filepath] = config
self._hashes[filepath] = file_hash
return config
def clear(self):
self._cache.clear()
self._hashes.clear()
# Usage
cache = ConfigCache()
config = cache.load('config.yaml')
Memory-Efficient Processing
import yaml
def process_large_yaml_file(filepath: str):
"""Process large YAML file without loading entirely into memory."""
# For multi-document files
with open(filepath, 'r') as f:
for doc in yaml.safe_load_all(f):
# Process each document individually
result = process_document(doc)
yield result
# Document is garbage collected after processing
def filter_yaml_documents(input_path: str, output_path: str, predicate):
"""Filter documents from a YAML file."""
with open(input_path, 'r') as infile, open(output_path, 'w') as outfile:
for doc in yaml.safe_load_all(infile):
if predicate(doc):
yaml.dump(doc, outfile, explicit_start=True)
Quick Reference
| Operation | Code | Description |
|---|---|---|
| Loading | yaml.safe_load(string) |
Load YAML from string (safe) |
yaml.safe_load(file) |
Load YAML from file object | |
yaml.safe_load_all(string) |
Load multiple documents | |
yaml.load(string, Loader=yaml.FullLoader) |
Load with specific loader | |
| Dumping | yaml.dump(data) |
Dump to YAML string |
yaml.safe_dump(data) |
Dump basic types only | |
yaml.dump(data, file) |
Dump to file object | |
yaml.dump_all(docs) |
Dump multiple documents | |
| Formatting | yaml.dump(data, default_flow_style=False) |
Block style output |
yaml.dump(data, indent=2) |
Custom indentation | |
yaml.dump(data, sort_keys=False) |
Preserve key order | |
yaml.dump(data, allow_unicode=True) |
Allow Unicode chars | |
| Custom Types | yaml.add_representer(cls, func) |
Add custom representer |
yaml.add_constructor(tag, func) |
Add custom constructor | |
| Performance | yaml.CSafeLoader |
C-based loader (faster) |
yaml.CSafeDumper |
C-based dumper (faster) | |
| Security | yaml.safe_load() |
Safe for untrusted input |
Never use yaml.unsafe_load() |
With untrusted data |
Common Issues and Solutions
| Issue | Solution |
|---|---|
YAMLError: could not determine constructor |
Use safe_load() or register custom constructor |
| Duplicate keys silently overwritten | PyYAML keeps the last value without warning; use ruamel.yaml to detect duplicates |
TypeError: cannot serialize |
Add custom representer for the type |
UnicodeDecodeError |
Specify encoding: open(file, encoding='utf-8') |
| Memory error with large files | Use safe_load_all() for streaming or process in chunks |
| Comments lost after round-trip | PyYAML doesn't preserve comments; use ruamel.yaml |
Wrong boolean parsing (yes/no) |
YAML 1.1 treats these as booleans; quote strings |
| Anchors not created in output | Anchors are resolved on load; manually author YAML |
None becomes string 'null' |
This is correct YAML; use None in Python |
| Numbers with leading zeros | Quote them: '007' to prevent octal interpretation |
| Multiline strings formatting | Use | (literal) or > (folded) style |
| Float precision loss | Use Decimal with custom representer/constructor |
!!python/object security risk |
Never use unsafe_load() with untrusted data |
| Dates parsed as strings | safe_load parses ISO dates/timestamps natively; quote values to keep them as strings |
| C loader not available | Use an official wheel, or build PyYAML from source with libyaml-dev installed |
Related Topics
The following topics complement PyYAML development:
- Pydantic - Data validation and settings management for type-safe YAML configuration
- Python Patterns - Design patterns for configuration management and dependency injection
- FastAPI - Web framework that benefits from YAML configuration with Pydantic
- Docker/Kubernetes - Container orchestration using YAML manifests
- Ansible - Infrastructure automation heavily using YAML playbooks
- Jinja2 - Template engine for generating dynamic YAML files