Python Bleach
A library for sanitising and cleaning HTML text, commonly used to prevent XSS attacks in user-generated content.
Python Bleach
A library for sanitising and cleaning HTML text, commonly used to prevent XSS attacks in user-generated content.
Deprecated — end of life. Bleach was archived by its maintainers in January 2023 and receives no further security fixes;
6.4.0is effectively the final line. For new work, prefernh3— Python bindings to the Rust ammonia sanitiser, which is faster, actively maintained, and the upstream-recommended successor. This sheet documents bleach for existing codebases; migrate where you can.
Overview
Bleach is a Python library that sanitises HTML by removing or escaping potentially dangerous content while preserving safe markup. It's essential for any application that displays user-generated HTML content, providing protection against cross-site scripting (XSS) attacks.
flowchart LR
A[User Input] --> B[Bleach]
B --> C{Sanitisation}
C --> D[Remove Unsafe Tags]
C --> E[Filter Attributes]
C --> F[Validate URLs]
D --> G[Safe HTML Output]
E --> G
F --> G
G --> H[Display to Users]
Installation and Basic Usage
Installation
pip install bleach
Basic Sanitisation
import bleach
# Default clean() - escapes any tag NOT in bleach.ALLOWED_TAGS;
# tags that ARE in the defaults (e.g. <b>) are kept
dirty = '<script>alert("xss")</script><p>Hello <b>World</b></p>'
clean = bleach.clean(dirty)
# Result: '<script>alert("xss")</script><p>Hello <b>World</b></p>'
# (<p> is escaped - not a default tag; <b> survives - it IS a default tag)
# Allow specific tags
clean = bleach.clean(dirty, tags=['p', 'b'])
# Result: '<script>alert("xss")</script><p>Hello <b>World</b></p>'
HTML Sanitisation
Core Concepts
Bleach's clean() function is the primary method for sanitising HTML. It processes input by:
- Parsing the HTML
- Removing or escaping disallowed tags
- Filtering attributes
- Validating URL protocols in links
flowchart TD
A[Input HTML] --> B[Parse HTML]
B --> C[Walk DOM Tree]
C --> D{Tag Allowed?}
D -->|Yes| E[Check Attributes]
D -->|No| F[Escape/Strip Tag]
E --> G{Attribute Allowed?}
G -->|Yes| H[Validate Value]
G -->|No| I[Remove Attribute]
H --> J[Output Clean HTML]
F --> J
I --> J
Cleaning Options
import bleach
html = '<div class="content"><a href="javascript:alert(1)">Click</a><img src="x" onerror="alert(1)"></div>'
# Default behaviour - escape all tags
result = bleach.clean(html)
# Strip DISALLOWED tags instead of escaping them.
# With default tags, <a> is allowed (its unsafe attrs are dropped);
# <div> and <img> are not in the defaults, so they're removed.
result = bleach.clean(html, strip=True)
# Result: '<a>Click</a>'
# Strip comments
html_with_comments = '<!-- comment --><p>Text</p>'
result = bleach.clean(html_with_comments, strip_comments=True, tags=['p'])
# Result: '<p>Text</p>'
Handling Malformed HTML
import bleach
# Bleach handles malformed HTML gracefully
malformed = '<p>Unclosed paragraph<div>Nested wrong</p></div>'
clean = bleach.clean(malformed, tags=['p', 'div'])
# Bleach uses html5lib parser which fixes structure
Whitelisting and Blacklisting Tags
Default Allowed Tags
import bleach
# Bleach's default allowed tags
print(bleach.ALLOWED_TAGS)
# {'a', 'abbr', 'acronym', 'b', 'blockquote', 'code', 'em', 'i', 'li', 'ol', 'strong', 'ul'}
Custom Tag Whitelist
import bleach
# Basic whitelist
allowed_tags = ['p', 'br', 'strong', 'em', 'a', 'ul', 'ol', 'li']
html = '<p>Hello</p><script>bad()</script><div>World</div>'
clean = bleach.clean(html, tags=allowed_tags)
# Result: '<p>Hello</p><script>bad()</script><div>World</div>'
# Extended whitelist for rich content
rich_tags = [
'p', 'br', 'strong', 'em', 'b', 'i', 'u',
'h1', 'h2', 'h3', 'h4', 'h5', 'h6',
'ul', 'ol', 'li',
'a', 'img',
'blockquote', 'pre', 'code',
'table', 'thead', 'tbody', 'tr', 'th', 'td',
'span', 'div'
]
Blacklisting Approach
Bleach uses a whitelist approach by default. To blacklist specific tags, process after cleaning:
import bleach
import re
def blacklist_tags(html, blacklist):
"""Remove specific tags while keeping others."""
# First, allow most tags
all_tags = ['p', 'div', 'span', 'a', 'img', 'table', 'tr', 'td', 'th',
'ul', 'ol', 'li', 'h1', 'h2', 'h3', 'b', 'i', 'strong', 'em']
# Remove blacklisted tags from allowed list
allowed = [tag for tag in all_tags if tag not in blacklist]
return bleach.clean(html, tags=allowed, strip=True)
# Remove tables but keep other formatting
html = '<p>Text</p><table><tr><td>Data</td></tr></table>'
clean = blacklist_tags(html, ['table', 'tr', 'td', 'th'])
# Result: '<p>Text</p>\nData' (stripped table leaves a stray newline)
Tag-Specific Configurations
import bleach
# Different configurations for different contexts
COMMENT_TAGS = ['p', 'br', 'strong', 'em', 'a']
ARTICLE_TAGS = ['p', 'br', 'strong', 'em', 'a', 'img', 'h2', 'h3', 'ul', 'ol', 'li', 'blockquote', 'pre', 'code']
ADMIN_TAGS = ARTICLE_TAGS + ['table', 'tr', 'td', 'th', 'thead', 'tbody', 'div', 'span']
def sanitise_content(html, content_type='comment'):
tag_sets = {
'comment': COMMENT_TAGS,
'article': ARTICLE_TAGS,
'admin': ADMIN_TAGS
}
return bleach.clean(html, tags=tag_sets.get(content_type, COMMENT_TAGS))
Attribute Filtering
Default Allowed Attributes
import bleach
# Bleach's default allowed attributes
print(bleach.ALLOWED_ATTRIBUTES)
# {'a': ['href', 'title'], 'abbr': ['title'], 'acronym': ['title']}
Custom Attribute Configuration
import bleach
# Attributes as dictionary
attributes = {
'a': ['href', 'title', 'rel'],
'img': ['src', 'alt', 'title', 'width', 'height'],
'p': ['class'],
'*': ['id'] # Allow 'id' on all tags
}
html = '<a href="http://example.com" onclick="bad()">Link</a>'
clean = bleach.clean(html, tags=['a'], attributes=attributes)
# Result: '<a href="http://example.com">Link</a>'
Callable Attribute Filter
import bleach
def filter_attributes(tag, name, value):
"""Custom attribute filter function."""
# Allow href on anchors, but validate protocol
if tag == 'a' and name == 'href':
if value.startswith(('http://', 'https://', 'mailto:')):
return True
return False
# Allow src on images with http(s) only
if tag == 'img' and name == 'src':
return value.startswith(('http://', 'https://'))
# Allow class on any tag, but restrict values
if name == 'class':
# Only allow specific CSS classes
allowed_classes = ['highlight', 'note', 'warning', 'code-block']
return value in allowed_classes
# Allow data attributes
if name.startswith('data-'):
return True
return False
html = '<a href="javascript:alert(1)">Bad</a><a href="https://safe.com">Good</a>'
clean = bleach.clean(html, tags=['a'], attributes=filter_attributes)
# Result: '<a>Bad</a><a href="https://safe.com">Good</a>'
Combined Dictionary and Callable
import bleach
def style_filter(tag, name, value):
"""Filter style attributes to allow only safe properties."""
if name != 'style':
return False
# Allow only specific CSS properties
safe_properties = ['color', 'background-color', 'font-size', 'text-align']
for prop in value.split(';'):
if ':' in prop:
property_name = prop.split(':')[0].strip()
if property_name and property_name not in safe_properties:
return False
return True
# Combine dictionary and callable
attributes = {
'a': ['href', 'title'],
'img': ['src', 'alt'],
'span': style_filter,
'div': style_filter
}
Link Handling and Linkification
URL Protocol Filtering
import bleach
# Default allowed protocols
print(bleach.ALLOWED_PROTOCOLS)
# {'http', 'https', 'mailto'}
# Custom protocols
protocols = ['http', 'https', 'mailto', 'tel', 'ftp']
html = '<a href="javascript:alert(1)">Bad</a><a href="tel:+441234567890">Call</a>'
clean = bleach.clean(
html,
tags=['a'],
attributes={'a': ['href']},
protocols=protocols
)
# Result: '<a>Bad</a><a href="tel:+441234567890">Call</a>'
Linkification with linkify()
import bleach
# Convert URLs to links automatically
text = "Check out https://example.com for more info"
linked = bleach.linkify(text)
# Result: 'Check out <a href="https://example.com" rel="nofollow">https://example.com</a> for more info'
# Linkify email addresses
text = "Contact us at support@example.com"
linked = bleach.linkify(text, parse_email=True)
# Result: 'Contact us at <a href="mailto:support@example.com">support@example.com</a>'
Linkify Callbacks
import bleach
from bleach.linkifier import LinkifyFilter
def set_target(attrs, new=False):
"""Add target="_blank" to external links."""
if new: # This is a new link (from linkification)
attrs[(None, 'target')] = '_blank'
attrs[(None, 'rel')] = 'noopener noreferrer'
return attrs
def shorten_url(attrs, new=False):
"""Shorten displayed URL text."""
if new:
text = attrs.get('_text', '')
if len(text) > 30:
attrs['_text'] = text[:27] + '...'
return attrs
def skip_internal(attrs, new=False):
"""Skip linkification for internal URLs."""
if new:
href = attrs.get((None, 'href'), '')
if 'internal-site.com' in href:
return None # Don't create link
return attrs
# Apply callbacks
text = "Visit https://example.com/very/long/path/to/page for details"
linked = bleach.linkify(
text,
callbacks=[set_target, shorten_url]
)
Combining clean() and linkify()
import bleach
def sanitise_and_linkify(html):
"""Clean HTML and convert URLs to links."""
# First, clean the HTML
allowed_tags = ['p', 'br', 'strong', 'em', 'a', 'ul', 'ol', 'li']
allowed_attrs = {'a': ['href', 'title', 'rel']}
cleaned = bleach.clean(
html,
tags=allowed_tags,
attributes=allowed_attrs,
strip=True
)
# Then linkify any plain URLs
def add_nofollow(attrs, new=False):
attrs[(None, 'rel')] = 'nofollow'
return attrs
return bleach.linkify(
cleaned,
callbacks=[add_nofollow],
skip_tags=['pre', 'code'] # Don't linkify inside code blocks
)
# Usage
user_content = '<p>Check https://example.com</p><script>bad()</script>'
safe_content = sanitise_and_linkify(user_content)
Custom Filters and Callbacks
Using Cleaner Class
import bleach
from bleach import Cleaner
# Create reusable cleaner instance
cleaner = Cleaner(
tags=['p', 'br', 'a', 'strong', 'em'],
attributes={'a': ['href', 'title']},
protocols=['http', 'https'],
strip=True,
strip_comments=True
)
# Use the cleaner
html = '<p>Hello <script>bad()</script><a href="https://example.com">World</a></p>'
clean = cleaner.clean(html)
Custom Filters with html5lib
import bleach
from bleach import Cleaner
from bleach.html5lib_shim import Filter
class AddRelNofollow(Filter):
"""Add rel="nofollow" to all links."""
def __iter__(self):
for token in Filter.__iter__(self):
if token['type'] == 'StartTag' and token['name'] == 'a':
token['data'][(None, 'rel')] = 'nofollow'
yield token
class RemoveEmptyTags(Filter):
"""Remove empty paragraph and div tags."""
def __iter__(self):
buffer = []
for token in Filter.__iter__(self):
if token['type'] == 'StartTag' and token['name'] in ('p', 'div'):
buffer.append(token)
elif token['type'] == 'EndTag' and token['name'] in ('p', 'div'):
if buffer:
start_token = buffer.pop()
# Check if there was content between start and end
# This is a simplified check
continue
else:
if buffer:
for buffered in buffer:
yield buffered
buffer = []
yield token
# Apply custom filters
cleaner = Cleaner(
tags=['p', 'a', 'div', 'strong'],
attributes={'a': ['href']},
filters=[AddRelNofollow]
)
Post-Processing Callbacks
import bleach
import re
def post_process_html(html):
"""Apply additional transformations after cleaning."""
# Convert newlines to <br> tags
html = html.replace('\n', '<br>')
# Auto-link @mentions
html = re.sub(
r'@(\w+)',
r'<a href="/user/\1">@\1</a>',
html
)
# Auto-link #hashtags
html = re.sub(
r'#(\w+)',
r'<a href="/tag/\1">#\1</a>',
html
)
return html
def sanitise_social_content(content):
"""Full pipeline for social media-style content."""
# Clean
clean = bleach.clean(
content,
tags=['p', 'br', 'a', 'strong', 'em'],
attributes={'a': ['href']},
strip=True
)
# Linkify URLs
linked = bleach.linkify(clean)
# Post-process for mentions and hashtags
final = post_process_html(linked)
return final
Common Use Cases
User-Generated Content
import bleach
class ContentSanitiser:
"""Sanitiser for user-generated content."""
# Configuration for different content types
CONFIGS = {
'comment': {
'tags': ['p', 'br', 'strong', 'em', 'a', 'code'],
'attributes': {'a': ['href', 'title']},
'strip': True
},
'post': {
'tags': ['p', 'br', 'strong', 'em', 'a', 'img',
'h2', 'h3', 'ul', 'ol', 'li', 'blockquote',
'pre', 'code'],
'attributes': {
'a': ['href', 'title'],
'img': ['src', 'alt', 'title']
},
'strip': True
},
'bio': {
'tags': ['br', 'strong', 'em', 'a'],
'attributes': {'a': ['href']},
'strip': True
}
}
@classmethod
def sanitise(cls, content, content_type='comment'):
config = cls.CONFIGS.get(content_type, cls.CONFIGS['comment'])
return bleach.clean(content, **config)
@classmethod
def sanitise_and_linkify(cls, content, content_type='comment'):
cleaned = cls.sanitise(content, content_type)
return bleach.linkify(cleaned)
# Usage
user_comment = '<p>Check out <script>alert("xss")</script>https://example.com</p>'
safe_comment = ContentSanitiser.sanitise_and_linkify(user_comment, 'comment')
Web Application Integration
Flask Integration
from flask import Flask, request
from markupsafe import Markup
import bleach
app = Flask(__name__)
def clean_html(html):
"""Template filter for cleaning HTML."""
cleaned = bleach.clean(
html,
tags=['p', 'br', 'strong', 'em', 'a'],
attributes={'a': ['href']},
strip=True
)
return Markup(cleaned)
# Register as template filter
app.jinja_env.filters['clean_html'] = clean_html
@app.route('/comment', methods=['POST'])
def create_comment():
content = request.form.get('content', '')
# Sanitise before storing
safe_content = bleach.clean(
content,
tags=['p', 'br', 'strong', 'em', 'a'],
attributes={'a': ['href']},
strip=True
)
# Store safe_content in database
# ...
return {'status': 'success'}
Django Integration
# forms.py
from django import forms
import bleach
class CommentForm(forms.Form):
content = forms.CharField(widget=forms.Textarea)
def clean_content(self):
content = self.cleaned_data['content']
return bleach.clean(
content,
tags=['p', 'br', 'strong', 'em', 'a'],
attributes={'a': ['href']},
strip=True
)
# templatetags/bleach_tags.py
from django import template
from django.utils.safestring import mark_safe
import bleach
register = template.Library()
@register.filter(is_safe=True)
def bleach_clean(value):
"""Clean HTML in templates."""
cleaned = bleach.clean(
value,
tags=['p', 'br', 'strong', 'em', 'a'],
attributes={'a': ['href']},
strip=True
)
return mark_safe(cleaned)
@register.filter(is_safe=True)
def bleach_linkify(value):
"""Linkify URLs in templates."""
return mark_safe(bleach.linkify(value))
FastAPI Integration
from fastapi import FastAPI, Form
from pydantic import BaseModel, field_validator
import bleach
app = FastAPI()
class Comment(BaseModel):
content: str
@field_validator('content')
@classmethod
def sanitise_content(cls, v):
return bleach.clean(
v,
tags=['p', 'br', 'strong', 'em', 'a'],
attributes={'a': ['href']},
strip=True
)
@app.post('/comments')
async def create_comment(comment: Comment):
# comment.content is already sanitised
return {'content': comment.content}
Markdown with Bleach
import bleach
import markdown
def safe_markdown(text):
"""Convert markdown to HTML and sanitise."""
# Convert markdown to HTML
html = markdown.markdown(
text,
extensions=['fenced_code', 'tables', 'nl2br']
)
# Sanitise the HTML output
allowed_tags = [
'p', 'br', 'strong', 'em', 'a', 'img',
'h1', 'h2', 'h3', 'h4', 'h5', 'h6',
'ul', 'ol', 'li',
'pre', 'code',
'blockquote',
'table', 'thead', 'tbody', 'tr', 'th', 'td',
'hr'
]
allowed_attrs = {
'a': ['href', 'title'],
'img': ['src', 'alt', 'title'],
'code': ['class'], # For syntax highlighting
'*': ['id']
}
return bleach.clean(
html,
tags=allowed_tags,
attributes=allowed_attrs,
strip=True
)
Security Best Practices
Essential Security Rules
flowchart TD
A[User Input] --> B{Sanitise on Input?}
B -->|Recommended| C[Store Raw + Sanitised]
B -->|Alternative| D[Sanitise on Output]
C --> E[Display Sanitised]
D --> E
E --> F{Context-Aware?}
F -->|Yes| G[Secure Display]
F -->|No| H[Potential XSS]
Defence in Depth
import bleach
from html import escape
class SecureContent:
"""Secure content handling with multiple layers."""
# Strict default configuration
STRICT_TAGS = ['p', 'br', 'strong', 'em']
STRICT_ATTRS = {}
@classmethod
def sanitise(cls, content, tags=None, attributes=None):
"""Sanitise with secure defaults."""
# Never allow None - use strict defaults
tags = tags if tags is not None else cls.STRICT_TAGS
attributes = attributes if attributes is not None else cls.STRICT_ATTRS
# Always strip dangerous content
cleaned = bleach.clean(
content,
tags=tags,
attributes=attributes,
protocols=['https', 'mailto'], # No http
strip=True,
strip_comments=True
)
return cleaned
@classmethod
def escape_all(cls, content):
"""Complete escape - no HTML allowed."""
return escape(content)
@classmethod
def sanitise_url(cls, url):
"""Sanitise URL for use in href/src."""
# Only allow safe protocols
safe_protocols = ('https://', 'mailto:', 'tel:')
if not any(url.startswith(p) for p in safe_protocols):
return '#'
# Additional URL validation here
return url
Preventing Common Attacks
import bleach
import re
def prevent_data_exfiltration(attrs, new=False):
"""Prevent data exfiltration via links."""
href = attrs.get((None, 'href'), '')
# Block data: URLs
if href.startswith('data:'):
return None
# Block URLs with credentials
if '@' in href and '://' in href:
return None
return attrs
def sanitise_for_display(content):
"""Full security sanitisation."""
# Step 1: Clean HTML
cleaned = bleach.clean(
content,
tags=['p', 'br', 'strong', 'em', 'a', 'code'],
attributes={'a': ['href', 'title']},
protocols=['https'],
strip=True,
strip_comments=True
)
# Step 2: Linkify with security callbacks
linked = bleach.linkify(
cleaned,
callbacks=[prevent_data_exfiltration],
skip_tags=['code', 'pre']
)
# Step 3: Additional security checks
# Remove any remaining javascript: (belt and braces)
linked = re.sub(r'javascript:', '', linked, flags=re.IGNORECASE)
return linked
Content Security Policy Integration
import bleach
def get_csp_compatible_cleaner():
"""Create cleaner compatible with strict CSP."""
def no_inline_styles(tag, name, value):
"""Remove all inline styles for CSP compliance."""
if name == 'style':
return False
if name == 'class':
# Whitelist specific classes
allowed = ['highlight', 'code', 'note', 'warning']
return value in allowed
return True
return bleach.Cleaner(
tags=['p', 'br', 'strong', 'em', 'a', 'code', 'pre', 'span'],
attributes={
'a': ['href', 'title'],
'span': no_inline_styles,
'code': no_inline_styles
},
protocols=['https', 'mailto'],
strip=True
)
# Usage
csp_cleaner = get_csp_compatible_cleaner()
safe_html = csp_cleaner.clean(user_input)
Performance Considerations
Optimisation Strategies
import bleach
from functools import lru_cache
# Create reusable Cleaner instances (avoid recreating)
COMMENT_CLEANER = bleach.Cleaner(
tags=['p', 'br', 'strong', 'em', 'a'],
attributes={'a': ['href']},
strip=True
)
ARTICLE_CLEANER = bleach.Cleaner(
tags=['p', 'br', 'strong', 'em', 'a', 'img', 'h2', 'h3',
'ul', 'ol', 'li', 'blockquote', 'pre', 'code'],
attributes={'a': ['href'], 'img': ['src', 'alt']},
strip=True
)
def sanitise_comment(content):
"""Use pre-configured cleaner."""
return COMMENT_CLEANER.clean(content)
def sanitise_article(content):
"""Use pre-configured cleaner."""
return ARTICLE_CLEANER.clean(content)
Caching Sanitised Content
import bleach
import hashlib
from functools import lru_cache
@lru_cache(maxsize=1000)
def cached_sanitise(content_hash, content):
"""Cache sanitisation results."""
return bleach.clean(
content,
tags=['p', 'br', 'strong', 'em', 'a'],
attributes={'a': ['href']},
strip=True
)
def sanitise_with_cache(content):
"""Sanitise with caching."""
content_hash = hashlib.md5(content.encode()).hexdigest()
return cached_sanitise(content_hash, content)
Batch Processing
import bleach
from concurrent.futures import ThreadPoolExecutor
# Single cleaner instance for batch processing
BATCH_CLEANER = bleach.Cleaner(
tags=['p', 'br', 'strong', 'em', 'a'],
attributes={'a': ['href']},
strip=True
)
def batch_sanitise(contents, max_workers=4):
"""Sanitise multiple contents in parallel."""
with ThreadPoolExecutor(max_workers=max_workers) as executor:
results = list(executor.map(BATCH_CLEANER.clean, contents))
return results
# Usage
user_posts = ['<p>Post 1</p>', '<p>Post 2</p>', '<p>Post 3</p>']
safe_posts = batch_sanitise(user_posts)
Performance Tips
| Technique | Benefit | When to Use |
|---|---|---|
| Reuse Cleaner instances | Avoid parser reinitialisation | Always |
| Cache results | Avoid re-sanitising same content | Repeated display |
| Batch processing | Parallel sanitisation | Multiple items |
| Minimal tag sets | Faster parsing | When possible |
| Store sanitised + raw | Avoid runtime sanitisation | Database storage |
Quick Reference
Essential Functions
| Function | Purpose | Example |
|---|---|---|
bleach.clean() |
Sanitise HTML | bleach.clean('<script>bad</script>') |
bleach.linkify() |
Convert URLs to links | bleach.linkify('Visit https://example.com') |
bleach.Cleaner() |
Create reusable cleaner | cleaner = Cleaner(tags=['p']) |
Common Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
tags |
list | ALLOWED_TAGS |
Tags to allow |
attributes |
dict/callable | ALLOWED_ATTRIBUTES |
Attributes to allow |
protocols |
list | ALLOWED_PROTOCOLS |
URL protocols to allow |
strip |
bool | False |
Strip disallowed tags |
strip_comments |
bool | False |
Remove HTML comments |
Defaults
# Default allowed tags
ALLOWED_TAGS = {'a', 'abbr', 'acronym', 'b', 'blockquote',
'code', 'em', 'i', 'li', 'ol', 'strong', 'ul'}
# Default allowed attributes
ALLOWED_ATTRIBUTES = {
'a': ['href', 'title'],
'abbr': ['title'],
'acronym': ['title']
}
# Default allowed protocols
ALLOWED_PROTOCOLS = {'http', 'https', 'mailto'}
Common Configurations
# Minimal (comments)
bleach.clean(html, tags=['p', 'br', 'strong', 'em', 'a'],
attributes={'a': ['href']}, strip=True)
# Standard (posts)
bleach.clean(html, tags=['p', 'br', 'strong', 'em', 'a', 'img',
'h2', 'h3', 'ul', 'ol', 'li', 'blockquote'],
attributes={'a': ['href'], 'img': ['src', 'alt']}, strip=True)
# Linkify with security
bleach.linkify(text, callbacks=[lambda a, n: {**a, (None, 'rel'): 'nofollow'}])
Common Issues and Solutions
Issue: Tags Being Escaped Instead of Stripped
Problem: Tags appear as <script> instead of being removed.
Solution: Use strip=True:
# Wrong - escapes tags
result = bleach.clean('<script>bad</script>')
# Result: '<script>bad</script>'
# Correct - strips tags
result = bleach.clean('<script>bad</script>', strip=True)
# Result: 'bad'
Issue: Attributes Being Removed
Problem: Safe attributes like href are being removed.
Solution: Explicitly allow attributes:
# Wrong - uses default attributes
result = bleach.clean('<a href="https://example.com" class="btn">Link</a>',
tags=['a'])
# 'href' kept but 'class' removed (not in defaults)
# Correct - specify needed attributes
result = bleach.clean('<a href="https://example.com" class="btn">Link</a>',
tags=['a'],
attributes={'a': ['href', 'class']})
Issue: Links with javascript: Protocol
Problem: javascript: URLs are being allowed.
Solution: Check protocol configuration:
# Bleach blocks javascript: by default
# If it's passing through, ensure you haven't added it to protocols
# Correct configuration
result = bleach.clean(
'<a href="javascript:alert(1)">Click</a>',
tags=['a'],
attributes={'a': ['href']},
protocols=['https', 'mailto'] # No javascript
)
# Result: '<a>Click</a>'
Issue: Empty Tags in Output
Problem: Empty tags like <p></p> appear after cleaning.
Solution: Post-process or use custom filter:
import re
def remove_empty_tags(html):
"""Remove empty paragraph and div tags."""
pattern = r'<(p|div|span)>\s*</\1>'
while re.search(pattern, html):
html = re.sub(pattern, '', html)
return html
cleaned = bleach.clean(html, tags=['p', 'div'], strip=True)
final = remove_empty_tags(cleaned)
Issue: Linkify Creating Unwanted Links
Problem: URLs in code blocks are being converted to links.
Solution: Use skip_tags:
result = bleach.linkify(
'<code>https://example.com</code>',
skip_tags=['code', 'pre']
)
# URL inside code remains as text
Issue: Performance Degradation
Problem: Sanitisation is slow with many items.
Solution: Reuse Cleaner instances:
# Wrong - creates new cleaner each time
def sanitise(content):
return bleach.clean(content, tags=['p'], strip=True)
# Correct - reuse cleaner
CLEANER = bleach.Cleaner(tags=['p'], strip=True)
def sanitise(content):
return CLEANER.clean(content)
Issue: Unicode/Encoding Errors
Problem: Non-ASCII characters cause errors.
Solution: Ensure proper encoding:
# Ensure string input (not bytes)
if isinstance(content, bytes):
content = content.decode('utf-8')
result = bleach.clean(content)
Issue: HTML Entities Being Double-Escaped
Problem: & becomes &amp;.
Solution: This is expected behaviour for safety. Decode before cleaning if needed:
import html
# If input has HTML entities that should be preserved
content = '& < >'
decoded = html.unescape(content) # '& < >'
result = bleach.clean(decoded, strip=True)
Related Topics
- HTML5lib - Underlying HTML parser used by Bleach
- MarkupSafe - Safe string handling for templates
- Jinja2 - Template engine with autoescaping
- Django Security - Django's built-in XSS protection
- Content Security Policy - Browser-level XSS mitigation
- OWASP XSS Prevention - Comprehensive XSS prevention guidelines