Python Markdown
A comprehensive guide to converting Markdown to HTML using the Python-Markdown library.
Python Markdown
A comprehensive guide to converting Markdown to HTML using the Python-Markdown library.
Overview
Python-Markdown is a full-featured Markdown parser that converts Markdown text to HTML. It supports the standard Markdown syntax, numerous extensions for additional features, and allows custom extension development for specialised needs.
flowchart LR
A[Markdown Text] --> B[Python-Markdown]
B --> C{Extensions}
C --> D[Built-in]
C --> E[Third-party]
C --> F[Custom]
D --> G[HTML Output]
E --> G
F --> G
Installation and Basic Usage
Installation
# Install base package
pip install markdown
# Install with common extras
pip install markdown[toc,tables]
# Install popular third-party extensions
pip install pymdown-extensions
pip install markdown-include
Converting Markdown to HTML
import markdown
# Basic conversion
md_text = "# Hello World\n\nThis is **bold** text."
html = markdown.markdown(md_text)
# Output: <h1>Hello World</h1>\n<p>This is <strong>bold</strong> text.</p>
# Using the Markdown class for multiple conversions
md = markdown.Markdown()
html1 = md.convert("First document")
md.reset() # Reset state between documents
html2 = md.convert("Second document")
# Convert file to HTML
with open('input.md', 'r') as f:
html = markdown.markdown(f.read())
# Save to file
with open('output.html', 'w') as f:
f.write(html)
Output Formats
import markdown
# XHTML output (default)
html = markdown.markdown(text, output_format='xhtml')
# HTML5 output
html = markdown.markdown(text, output_format='html')
# Control line breaks
md = markdown.Markdown(output_format='html')
Built-in Extensions
Python-Markdown includes several useful extensions that extend the base Markdown syntax.
graph TB
subgraph "Built-in Extensions"
A[tables] --> B[GFM-style tables]
C[fenced_code] --> D[Code blocks with syntax hints]
E[toc] --> F[Table of contents]
G[footnotes] --> H[Reference footnotes]
I[meta] --> J[Document metadata]
K[attr_list] --> L[HTML attributes]
M[def_list] --> N[Definition lists]
O[abbr] --> P[Abbreviations]
end
Tables Extension
import markdown
text = """
| Header 1 | Header 2 | Header 3 |
|----------|:--------:|---------:|
| Left | Centre | Right |
| Cell | Cell | Cell |
"""
html = markdown.markdown(text, extensions=['tables'])
Fenced Code Extension
import markdown
text = """
```python
def greet(name):
return f"Hello, {name}!"
"""
html = markdown.markdown(text, extensions=['fenced_code'])
With code highlighting (requires Pygments)
html = markdown.markdown( text, extensions=['fenced_code', 'codehilite'], extension_configs={ 'codehilite': { 'css_class': 'highlight', 'linenums': True } } )
### Table of Contents (TOC) Extension
```python
import markdown
text = """
# Main Title
## Section One
Content here.
## Section Two
More content.
### Subsection
Details here.
"""
md = markdown.Markdown(extensions=['toc'])
html = md.convert(text)
# Access generated TOC
toc_html = md.toc
toc_tokens = md.toc_tokens # Structured data
# Configuration options
html = markdown.markdown(
text,
extensions=['toc'],
extension_configs={
'toc': {
'title': 'Contents',
'toc_depth': 3,
'permalink': True,
'permalink_title': 'Link to this section',
'slugify': lambda value, separator: value.lower().replace(' ', separator)
}
}
)
Footnotes Extension
import markdown
text = """
This is a paragraph with a footnote[^1].
[^1]: This is the footnote content.
"""
html = markdown.markdown(text, extensions=['footnotes'])
Metadata Extension
import markdown
text = """
Title: My Document
Author: Jane Smith
Date: 2024-01-15
# Document Content
Body text here.
"""
md = markdown.Markdown(extensions=['meta'])
html = md.convert(text)
# Access metadata
title = md.Meta.get('title', [''])[0]
author = md.Meta.get('author', [''])[0]
Attribute Lists Extension
import markdown
text = """
# Heading {#custom-id .my-class}
A paragraph with custom attributes.
{: .highlight #para1 data-value="test" }
[Link](https://example.com){: target="_blank" rel="noopener" }
"""
html = markdown.markdown(text, extensions=['attr_list'])
Definition Lists Extension
import markdown
text = """
Term 1
: Definition for term 1
Term 2
: Definition for term 2
: Another definition for term 2
"""
html = markdown.markdown(text, extensions=['def_list'])
Multiple Extensions Together
import markdown
html = markdown.markdown(
text,
extensions=[
'tables',
'fenced_code',
'codehilite',
'toc',
'footnotes',
'attr_list',
'meta',
'smarty', # Smart quotes
'nl2br', # Newlines to <br>
],
extension_configs={
'codehilite': {'linenums': False},
'toc': {'permalink': True}
}
)
Third-party Extensions
Popular extensions that add powerful features beyond the built-in set.
PyMdown Extensions
pip install pymdown-extensions
import markdown
text = """
==Highlighted text==
~~Strikethrough~~
- [x] Completed task
- [ ] Pending task
!!! note "Important"
This is an admonition block.
???+ example "Collapsible"
Click to expand this content.
"""
html = markdown.markdown(
text,
extensions=[
'pymdownx.mark', # Highlighting
'pymdownx.tilde', # Strikethrough
'pymdownx.tasklist', # Task lists
'pymdownx.details', # Collapsible blocks
'pymdownx.superfences', # Enhanced fenced code
'pymdownx.tabbed', # Tabbed content
'pymdownx.emoji', # Emoji support
],
extension_configs={
'pymdownx.tasklist': {
'clickable_checkbox': True
}
}
)
Markdown Include
pip install markdown-include
import markdown
# In your markdown file:
# {!path/to/file.md!}
html = markdown.markdown(
text,
extensions=['markdown_include.include'],
extension_configs={
'markdown_include.include': {
'base_path': '/path/to/includes'
}
}
)
Other Useful Extensions
# Install various extensions
# pip install mdx-truly-sane-lists
# pip install markdown-callouts
# pip install markdown-katex
import markdown
html = markdown.markdown(
text,
extensions=[
'mdx_truly_sane_lists', # Better list handling
'markdown_katex', # LaTeX maths
]
)
Custom Extension Development
Create custom extensions to add specialised Markdown syntax.
flowchart TD
A[Extension Class] --> B[extendMarkdown]
B --> C{Processor Type}
C --> D[Preprocessor]
C --> E[BlockProcessor]
C --> F[InlineProcessor]
C --> G[Postprocessor]
D --> H[Modify raw text]
E --> I[Process blocks]
F --> J[Process inline patterns]
G --> K[Modify final output]
Basic Extension Structure
from markdown import Extension
from markdown.preprocessors import Preprocessor
from markdown.inlinepatterns import InlineProcessor
from markdown.postprocessors import Postprocessor
import xml.etree.ElementTree as etree
import re
class MyExtension(Extension):
def __init__(self, **kwargs):
# Define default configuration
self.config = {
'option_name': ['default_value', 'Description of option']
}
super().__init__(**kwargs)
def extendMarkdown(self, md):
# Register processors
md.preprocessors.register(
MyPreprocessor(md), 'my_preprocessor', 25
)
md.inlinePatterns.register(
MyInlineProcessor(r'pattern', md), 'my_inline', 75
)
md.postprocessors.register(
MyPostprocessor(md), 'my_postprocessor', 25
)
# Create extension factory function
def makeExtension(**kwargs):
return MyExtension(**kwargs)
Custom Preprocessor
from markdown.preprocessors import Preprocessor
import re
class VariablePreprocessor(Preprocessor):
"""Replace {{variable}} with values from a dictionary."""
def __init__(self, md, variables):
super().__init__(md)
self.variables = variables
def run(self, lines):
new_lines = []
pattern = re.compile(r'\{\{(\w+)\}\}')
for line in lines:
def replace(match):
var_name = match.group(1)
return self.variables.get(var_name, match.group(0))
new_lines.append(pattern.sub(replace, line))
return new_lines
class VariableExtension(Extension):
def __init__(self, **kwargs):
self.config = {
'variables': [{}, 'Dictionary of variables']
}
super().__init__(**kwargs)
def extendMarkdown(self, md):
variables = self.getConfig('variables')
md.preprocessors.register(
VariablePreprocessor(md, variables),
'variables',
30
)
# Usage
html = markdown.markdown(
"Hello, {{name}}!",
extensions=[VariableExtension(variables={'name': 'World'})]
)
Custom Inline Processor
from markdown.inlinepatterns import InlineProcessor
import xml.etree.ElementTree as etree
class KeyboardInlineProcessor(InlineProcessor):
"""Convert [[key]] to <kbd>key</kbd>."""
def handleMatch(self, m, data):
el = etree.Element('kbd')
el.text = m.group(1)
return el, m.start(0), m.end(0)
class KeyboardExtension(Extension):
def extendMarkdown(self, md):
pattern = r'\[\[([^\]]+)\]\]'
md.inlinePatterns.register(
KeyboardInlineProcessor(pattern, md),
'keyboard',
75
)
# Usage
html = markdown.markdown(
"Press [[Ctrl]]+[[C]] to copy",
extensions=[KeyboardExtension()]
)
# Output: Press <kbd>Ctrl</kbd>+<kbd>C</kbd> to copy
Custom Block Processor
from markdown.blockprocessors import BlockProcessor
import xml.etree.ElementTree as etree
import re
class AlertBlockProcessor(BlockProcessor):
"""Convert ::: alert blocks to styled divs."""
RE_START = re.compile(r'^::: ?(\w+)\s*$')
RE_END = re.compile(r'^:::\s*$')
def test(self, parent, block):
return bool(self.RE_START.match(block.split('\n')[0]))
def run(self, parent, blocks):
block = blocks.pop(0)
lines = block.split('\n')
# Get alert type
match = self.RE_START.match(lines[0])
alert_type = match.group(1)
# Find content
content_lines = []
for i, line in enumerate(lines[1:], 1):
if self.RE_END.match(line):
break
content_lines.append(line)
# Create element
div = etree.SubElement(parent, 'div')
div.set('class', f'alert alert-{alert_type}')
# Parse nested content
self.parser.parseBlocks(div, content_lines)
class AlertExtension(Extension):
def extendMarkdown(self, md):
md.parser.blockprocessors.register(
AlertBlockProcessor(md.parser),
'alert',
75
)
Custom Postprocessor
from markdown.postprocessors import Postprocessor
import re
class LazyLoadPostprocessor(Postprocessor):
"""Add loading='lazy' to all images."""
def run(self, text):
pattern = r'<img ([^>]*)>'
def add_lazy(match):
attrs = match.group(1)
if 'loading=' not in attrs:
attrs += ' loading="lazy"'
return f'<img {attrs}>'
return re.sub(pattern, add_lazy, text)
class LazyLoadExtension(Extension):
def extendMarkdown(self, md):
md.postprocessors.register(
LazyLoadPostprocessor(md),
'lazy_load',
25
)
Common Use Cases
Documentation Generation
import markdown
from pathlib import Path
def generate_docs(source_dir, output_dir):
"""Generate HTML documentation from Markdown files."""
md = markdown.Markdown(
extensions=[
'toc',
'tables',
'fenced_code',
'codehilite',
'meta',
'attr_list',
],
extension_configs={
'toc': {
'permalink': True,
'toc_depth': 3
},
'codehilite': {
'css_class': 'highlight',
'guess_lang': False
}
}
)
source_path = Path(source_dir)
output_path = Path(output_dir)
output_path.mkdir(parents=True, exist_ok=True)
template = """<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>{title}</title>
<link rel="stylesheet" href="styles.css">
</head>
<body>
<nav>{toc}</nav>
<main>{content}</main>
</body>
</html>"""
for md_file in source_path.glob('**/*.md'):
content = md_file.read_text(encoding='utf-8')
html_content = md.convert(content)
# Get metadata
title = md.Meta.get('title', [md_file.stem])[0]
toc = md.toc
# Generate full HTML
full_html = template.format(
title=title,
toc=toc,
content=html_content
)
# Write output
rel_path = md_file.relative_to(source_path)
out_file = output_path / rel_path.with_suffix('.html')
out_file.parent.mkdir(parents=True, exist_ok=True)
out_file.write_text(full_html, encoding='utf-8')
md.reset()
# Usage
generate_docs('docs/source', 'docs/build')
Blog Post Processing
import markdown
from datetime import datetime
import re
def process_blog_post(md_content):
"""Process a blog post with metadata and content."""
md = markdown.Markdown(
extensions=[
'meta',
'toc',
'tables',
'fenced_code',
'codehilite',
'footnotes',
'smarty',
],
extension_configs={
'toc': {'toc_depth': 2}
}
)
html = md.convert(md_content)
# Extract and process metadata
meta = {
'title': md.Meta.get('title', ['Untitled'])[0],
'author': md.Meta.get('author', ['Anonymous'])[0],
'date': md.Meta.get('date', [datetime.now().isoformat()])[0],
'tags': md.Meta.get('tags', []),
'summary': md.Meta.get('summary', [''])[0],
}
# Calculate reading time (average 200 words per minute)
word_count = len(re.findall(r'\w+', md_content))
meta['reading_time'] = max(1, round(word_count / 200))
return {
'meta': meta,
'content': html,
'toc': md.toc
}
# Example blog post
blog_post = """
Title: Getting Started with Python-Markdown
Author: Jane Developer
Date: 2024-01-20
Tags: python, markdown, tutorial
Summary: Learn how to convert Markdown to HTML using Python-Markdown.
# Getting Started with Python-Markdown
In this tutorial, we'll explore the basics of Python-Markdown...
"""
result = process_blog_post(blog_post)
print(f"Title: {result['meta']['title']}")
print(f"Reading time: {result['meta']['reading_time']} min")
API Documentation with Type Hints
import markdown
import inspect
from typing import get_type_hints
def generate_api_docs(module):
"""Generate API documentation from module docstrings."""
md_content = f"# {module.__name__} API Reference\n\n"
for name, obj in inspect.getmembers(module):
if name.startswith('_'):
continue
if inspect.isfunction(obj):
md_content += f"## `{name}`\n\n"
# Get signature
sig = inspect.signature(obj)
md_content += f"```python\n{name}{sig}\n```\n\n"
# Get docstring
if obj.__doc__:
md_content += f"{obj.__doc__}\n\n"
# Get type hints
try:
hints = get_type_hints(obj)
if hints:
md_content += "**Parameters:**\n\n"
for param, type_hint in hints.items():
if param != 'return':
md_content += f"- `{param}`: {type_hint}\n"
md_content += "\n"
except Exception:
pass
return markdown.markdown(
md_content,
extensions=['fenced_code', 'tables', 'toc']
)
Security Considerations
HTML Sanitisation
Python-Markdown does not sanitise HTML by default. Use additional libraries for user-generated content.
import markdown
import bleach
def safe_markdown(text, allowed_tags=None):
"""Convert Markdown to sanitised HTML."""
# Default allowed tags
if allowed_tags is None:
allowed_tags = [
'h1', 'h2', 'h3', 'h4', 'h5', 'h6',
'p', 'br', 'hr',
'strong', 'em', 'code', 'pre',
'ul', 'ol', 'li',
'blockquote',
'a', 'img',
'table', 'thead', 'tbody', 'tr', 'th', 'td',
]
allowed_attrs = {
'a': ['href', 'title', 'rel'],
'img': ['src', 'alt', 'title', 'loading'],
'code': ['class'],
'pre': ['class'],
'*': ['id', 'class'],
}
# Convert to HTML
html = markdown.markdown(
text,
extensions=['tables', 'fenced_code']
)
# Sanitise
clean_html = bleach.clean(
html,
tags=allowed_tags,
attributes=allowed_attrs,
strip=True
)
# Linkify URLs (optional)
clean_html = bleach.linkify(clean_html)
return clean_html
# Usage with user content
user_input = """
# User Post
<script>alert('XSS')</script>
Normal **markdown** content.
<img src="x" onerror="alert('XSS')">
"""
safe_html = safe_markdown(user_input)
# Script tags and event handlers are stripped
Disable Raw HTML
import markdown
import re
# Python-Markdown has no "safe mode" — raw HTML always passes through
# unchanged (the safe_mode/HTML-stripping option was removed in 3.0).
# md_in_html does NOT strip HTML; it parses Markdown *inside* HTML blocks.
# To drop raw HTML, strip it before (or sanitise after — see bleach above).
# Note: bleach is archived/EOL (2023). For new code prefer nh3 (Rust ammonia
# bindings): nh3.clean(html) — actively maintained, faster, safer defaults.
text = "Normal text <script>alert('bad')</script>"
def strip_html(text):
return re.sub(r'<[^>]+>', '', text)
clean_text = strip_html(text)
html = markdown.markdown(clean_text)
Safe Extension Configuration
import markdown
# Avoid unsafe configurations
html = markdown.markdown(
text,
extensions=['codehilite'],
extension_configs={
'codehilite': {
# Don't execute code
'use_pygments': True,
# Limit language guessing
'guess_lang': False,
}
}
)
# Validate external includes
from markdown_include.include import IncludePreprocessor
class SafeIncludePreprocessor(IncludePreprocessor):
def __init__(self, md, config, allowed_paths):
super().__init__(md, config)
self.allowed_paths = allowed_paths
def run(self, lines):
# Validate paths before including
# Implementation would check against allowed_paths
return super().run(lines)
Performance Considerations
Caching Converted Content
import markdown
import hashlib
from functools import lru_cache
# Simple in-memory cache
@lru_cache(maxsize=1000)
def cached_markdown(text_hash, extensions_tuple):
"""Cache converted Markdown by content hash."""
# Note: actual text passed separately for hashing
pass
def convert_with_cache(text, extensions=None):
"""Convert Markdown with caching."""
if extensions is None:
extensions = ['tables', 'fenced_code']
# Create hash of content and extensions
content_hash = hashlib.md5(text.encode()).hexdigest()
cache_key = f"{content_hash}:{','.join(sorted(extensions))}"
# Check cache (use Redis/Memcached in production)
from functools import lru_cache
@lru_cache(maxsize=1000)
def _convert(cache_key, text):
return markdown.markdown(text, extensions=extensions)
return _convert(cache_key, text)
# Redis-based caching for production
import redis
import json
class MarkdownCache:
def __init__(self, redis_url='redis://localhost:6379'):
self.redis = redis.from_url(redis_url)
self.md = markdown.Markdown(
extensions=['tables', 'fenced_code', 'toc']
)
self.ttl = 3600 # 1 hour
def convert(self, text):
cache_key = f"md:{hashlib.sha256(text.encode()).hexdigest()}"
# Check cache
cached = self.redis.get(cache_key)
if cached:
return json.loads(cached)
# Convert and cache
html = self.md.convert(text)
result = {
'html': html,
'toc': self.md.toc
}
self.redis.setex(
cache_key,
self.ttl,
json.dumps(result)
)
self.md.reset()
return result
Batch Processing
import markdown
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
from pathlib import Path
def convert_file(file_path):
"""Convert a single file (for parallel processing)."""
md = markdown.Markdown(
extensions=['tables', 'fenced_code', 'toc']
)
content = Path(file_path).read_text(encoding='utf-8')
html = md.convert(content)
return {
'path': str(file_path),
'html': html,
'toc': md.toc
}
def batch_convert(file_paths, max_workers=4):
"""Convert multiple files in parallel."""
# Use ProcessPoolExecutor for CPU-bound work
with ProcessPoolExecutor(max_workers=max_workers) as executor:
results = list(executor.map(convert_file, file_paths))
return results
# Usage
files = list(Path('docs').glob('**/*.md'))
results = batch_convert(files)
Memory Optimisation
import markdown
def convert_large_file(file_path, chunk_size=1000):
"""Process large files in chunks (where possible)."""
md = markdown.Markdown(
extensions=['tables', 'fenced_code']
)
# For most cases, process entire file
# Markdown needs full context for proper conversion
with open(file_path, 'r', encoding='utf-8') as f:
content = f.read()
html = md.convert(content)
# Clear internal state to free memory
md.reset()
return html
# Reuse Markdown instance for multiple documents
class MarkdownConverter:
def __init__(self, extensions=None):
self.md = markdown.Markdown(
extensions=extensions or ['tables', 'fenced_code']
)
def convert(self, text):
result = self.md.convert(text)
self.md.reset() # Important: reset between documents
return result
def convert_many(self, texts):
results = []
for text in texts:
results.append(self.convert(text))
return results
Profiling Conversion
import markdown
import time
import cProfile
import pstats
def profile_conversion(text, extensions):
"""Profile Markdown conversion performance."""
profiler = cProfile.Profile()
profiler.enable()
start = time.perf_counter()
for _ in range(100):
md = markdown.Markdown(extensions=extensions)
md.convert(text)
elapsed = time.perf_counter() - start
profiler.disable()
stats = pstats.Stats(profiler)
stats.sort_stats('cumulative')
stats.print_stats(10)
print(f"\nTotal time for 100 conversions: {elapsed:.3f}s")
print(f"Average per conversion: {elapsed/100*1000:.2f}ms")
# Test different extension combinations
text = open('large_document.md').read()
print("Minimal extensions:")
profile_conversion(text, ['tables'])
print("\nFull extensions:")
profile_conversion(text, [
'tables', 'fenced_code', 'codehilite',
'toc', 'footnotes', 'attr_list'
])
Quick Reference
Common Operations
| Operation | Code |
|---|---|
| Basic conversion | markdown.markdown(text) |
| With extensions | markdown.markdown(text, extensions=['tables', 'toc']) |
| Configure extension | extension_configs={'toc': {'permalink': True}} |
| Reusable instance | md = markdown.Markdown(); md.convert(text) |
| Reset between docs | md.reset() |
| Access TOC | md.toc |
| Access metadata | md.Meta |
| HTML5 output | output_format='html' |
Built-in Extensions
| Extension | Purpose |
|---|---|
tables |
GFM-style tables |
fenced_code |
Code blocks with language hints |
codehilite |
Syntax highlighting (requires Pygments) |
toc |
Table of contents generation |
footnotes |
Reference-style footnotes |
meta |
YAML-style metadata |
attr_list |
Add HTML attributes |
def_list |
Definition lists |
abbr |
Abbreviations |
smarty |
Smart quotes and dashes |
nl2br |
Convert newlines to <br> |
sane_lists |
Better list handling |
md_in_html |
Markdown inside HTML blocks |
PyMdown Extensions
| Extension | Purpose |
|---|---|
pymdownx.superfences |
Enhanced code blocks |
pymdownx.tabbed |
Tabbed content |
pymdownx.details |
Collapsible blocks |
pymdownx.tasklist |
Task lists with checkboxes |
pymdownx.mark |
Highlighted text |
pymdownx.emoji |
Emoji support |
pymdownx.arithmatex |
LaTeX maths |
pymdownx.critic |
Track changes |
Extension Configuration Examples
extension_configs = {
'toc': {
'permalink': True,
'toc_depth': 3,
'title': 'Contents'
},
'codehilite': {
'css_class': 'highlight',
'linenums': False,
'guess_lang': False
},
'footnotes': {
'BACKLINK_TEXT': '↩'
}
}
Common Issues and Solutions
Issue: Extensions Not Loading
# Problem: Extension not found
# ModuleNotFoundError: No module named 'tables'
# Solution: Use correct extension names (strings)
html = markdown.markdown(text, extensions=['tables']) # Correct
# Not: extensions=[tables] # Wrong - undefined variable
# For third-party extensions, ensure they're installed
# pip install pymdown-extensions
Issue: TOC Not Generating
# Problem: md.toc is empty
# Solution: Use Markdown class instance
md = markdown.Markdown(extensions=['toc'])
html = md.convert(text)
toc = md.toc # Access after conversion
# Not: html = markdown.markdown(text, extensions=['toc'])
# The function doesn't expose toc attribute
Issue: Metadata Not Accessible
# Problem: md.Meta is empty or missing
# Solution 1: Ensure meta extension is enabled
md = markdown.Markdown(extensions=['meta'])
# Solution 2: Ensure metadata is at document start with no blank lines before
text = """Title: My Doc
Author: Me
# Content here
"""
# Solution 3: Check metadata format (no spaces around colon)
# Wrong: Title : My Doc
# Right: Title: My Doc
Issue: Code Highlighting Not Working
# Problem: Code blocks not syntax highlighted
# Solution 1: Install Pygments
# pip install Pygments
# Solution 2: Enable codehilite extension
html = markdown.markdown(
text,
extensions=['fenced_code', 'codehilite']
)
# Solution 3: Specify language in code block
text = """
```python
print("Hello")
"""
Solution 4: Include Pygments CSS in your HTML
pygmentize -S default -f html > pygments.css
### Issue: Reset Not Clearing State
```python
# Problem: Previous document's TOC/meta appearing
# Solution: Always reset between documents
md = markdown.Markdown(extensions=['toc', 'meta'])
html1 = md.convert(doc1)
toc1 = md.toc
md.reset() # Critical!
html2 = md.convert(doc2)
toc2 = md.toc # Now contains only doc2's TOC
Issue: Raw HTML Being Escaped
# Problem: HTML tags appearing as text
# Solution: HTML is allowed by default, check for issues
# 1. Ensure proper blank lines around HTML blocks
text = """
Paragraph before.
<div class="custom">
Content inside
</div>
Paragraph after.
"""
# 2. For inline HTML, no blank lines needed
text = "This is <strong>bold</strong> text."
# 3. For Markdown inside HTML, use md_in_html extension
text = """
<div markdown="1">
# This will be processed as Markdown
</div>
"""
html = markdown.markdown(text, extensions=['md_in_html'])
Issue: Unicode/Encoding Errors
# Problem: UnicodeDecodeError or garbled characters
# Solution: Always specify encoding
with open('document.md', 'r', encoding='utf-8') as f:
text = f.read()
html = markdown.markdown(text)
# When writing output
with open('output.html', 'w', encoding='utf-8') as f:
f.write(html)
Issue: Nested Lists Not Rendering Correctly
# Problem: Nested lists becoming flat
# Solution 1: Use consistent indentation (4 spaces)
text = """
- Item 1
- Nested item
- Another nested
- Item 2
"""
# Solution 2: Use sane_lists extension
html = markdown.markdown(text, extensions=['sane_lists'])
# Solution 3: Use mdx_truly_sane_lists for better handling
# pip install mdx-truly-sane-lists
html = markdown.markdown(text, extensions=['mdx_truly_sane_lists'])
Issue: Tables Not Rendering
# Problem: Table markup appearing as plain text
# Solution 1: Enable tables extension
html = markdown.markdown(text, extensions=['tables'])
# Solution 2: Check table syntax
# - Need header row
# - Need separator row with dashes
# - Pipes on outer edges optional but recommended
text = """
| Header 1 | Header 2 |
|----------|----------|
| Cell 1 | Cell 2 |
"""
# Solution 3: Ensure blank lines before and after table
Issue: Performance with Large Documents
# Problem: Slow conversion for large files
# Solution 1: Minimise extensions
# Only use what you need
html = markdown.markdown(text, extensions=['tables'])
# Solution 2: Disable expensive features
extension_configs = {
'codehilite': {
'guess_lang': False # Disable language guessing
}
}
# Solution 3: Cache results
# See Performance Considerations section
# Solution 4: Process in parallel for multiple files
# See Batch Processing section