Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Python Bleach

A library for sanitising and cleaning HTML text, commonly used to prevent XSS attacks in user-generated content.

Python Bleach

A library for sanitising and cleaning HTML text, commonly used to prevent XSS attacks in user-generated content.

Deprecated — end of life. Bleach was archived by its maintainers in January 2023 and receives no further security fixes; 6.4.0 is effectively the final line. For new work, prefer nh3 — Python bindings to the Rust ammonia sanitiser, which is faster, actively maintained, and the upstream-recommended successor. This sheet documents bleach for existing codebases; migrate where you can.

Overview

Bleach is a Python library that sanitises HTML by removing or escaping potentially dangerous content while preserving safe markup. It's essential for any application that displays user-generated HTML content, providing protection against cross-site scripting (XSS) attacks.

User InputBleachSanitisationRemove Unsafe TagsFilter AttributesValidate URLsSafe HTML OutputDisplay to UsersUser InputBleachSanitisationRemove Unsafe TagsFilter AttributesValidate URLsSafe HTML OutputDisplay to Users

Installation and Basic Usage

Installation

pip install bleach

Basic Sanitisation

import bleach

# Default clean() - escapes any tag NOT in bleach.ALLOWED_TAGS;
# tags that ARE in the defaults (e.g. <b>) are kept
dirty = '<script>alert("xss")</script><p>Hello <b>World</b></p>'
clean = bleach.clean(dirty)
# Result: '&lt;script&gt;alert("xss")&lt;/script&gt;&lt;p&gt;Hello <b>World</b>&lt;/p&gt;'
# (<p> is escaped - not a default tag; <b> survives - it IS a default tag)

# Allow specific tags
clean = bleach.clean(dirty, tags=['p', 'b'])
# Result: '&lt;script&gt;alert("xss")&lt;/script&gt;<p>Hello <b>World</b></p>'

HTML Sanitisation

Core Concepts

Bleach's clean() function is the primary method for sanitising HTML. It processes input by:

  1. Parsing the HTML
  2. Removing or escaping disallowed tags
  3. Filtering attributes
  4. Validating URL protocols in links
YesNoYesNoInput HTMLParse HTMLWalk DOM TreeTag Allowed?Check AttributesEscape/Strip TagAttribute Allowed?Validate ValueRemove AttributeOutput Clean HTMLYesNoYesNoInput HTMLParse HTMLWalk DOM TreeTag Allowed?Check AttributesEscape/Strip TagAttribute Allowed?Validate ValueRemove AttributeOutput Clean HTML

Cleaning Options

import bleach

html = '<div class="content"><a href="javascript:alert(1)">Click</a><img src="x" onerror="alert(1)"></div>'

# Default behaviour - escape all tags
result = bleach.clean(html)

# Strip DISALLOWED tags instead of escaping them.
# With default tags, <a> is allowed (its unsafe attrs are dropped);
# <div> and <img> are not in the defaults, so they're removed.
result = bleach.clean(html, strip=True)
# Result: '<a>Click</a>'

# Strip comments
html_with_comments = '<!-- comment --><p>Text</p>'
result = bleach.clean(html_with_comments, strip_comments=True, tags=['p'])
# Result: '<p>Text</p>'

Handling Malformed HTML

import bleach

# Bleach handles malformed HTML gracefully
malformed = '<p>Unclosed paragraph<div>Nested wrong</p></div>'
clean = bleach.clean(malformed, tags=['p', 'div'])
# Bleach uses html5lib parser which fixes structure

Whitelisting and Blacklisting Tags

Default Allowed Tags

import bleach

# Bleach's default allowed tags
print(bleach.ALLOWED_TAGS)
# {'a', 'abbr', 'acronym', 'b', 'blockquote', 'code', 'em', 'i', 'li', 'ol', 'strong', 'ul'}

Custom Tag Whitelist

import bleach

# Basic whitelist
allowed_tags = ['p', 'br', 'strong', 'em', 'a', 'ul', 'ol', 'li']
html = '<p>Hello</p><script>bad()</script><div>World</div>'
clean = bleach.clean(html, tags=allowed_tags)
# Result: '<p>Hello</p>&lt;script&gt;bad()&lt;/script&gt;&lt;div&gt;World&lt;/div&gt;'

# Extended whitelist for rich content
rich_tags = [
    'p', 'br', 'strong', 'em', 'b', 'i', 'u',
    'h1', 'h2', 'h3', 'h4', 'h5', 'h6',
    'ul', 'ol', 'li',
    'a', 'img',
    'blockquote', 'pre', 'code',
    'table', 'thead', 'tbody', 'tr', 'th', 'td',
    'span', 'div'
]

Blacklisting Approach

Bleach uses a whitelist approach by default. To blacklist specific tags, process after cleaning:

import bleach
import re

def blacklist_tags(html, blacklist):
    """Remove specific tags while keeping others."""
    # First, allow most tags
    all_tags = ['p', 'div', 'span', 'a', 'img', 'table', 'tr', 'td', 'th',
                'ul', 'ol', 'li', 'h1', 'h2', 'h3', 'b', 'i', 'strong', 'em']

    # Remove blacklisted tags from allowed list
    allowed = [tag for tag in all_tags if tag not in blacklist]

    return bleach.clean(html, tags=allowed, strip=True)

# Remove tables but keep other formatting
html = '<p>Text</p><table><tr><td>Data</td></tr></table>'
clean = blacklist_tags(html, ['table', 'tr', 'td', 'th'])
# Result: '<p>Text</p>\nData'  (stripped table leaves a stray newline)

Tag-Specific Configurations

import bleach

# Different configurations for different contexts
COMMENT_TAGS = ['p', 'br', 'strong', 'em', 'a']
ARTICLE_TAGS = ['p', 'br', 'strong', 'em', 'a', 'img', 'h2', 'h3', 'ul', 'ol', 'li', 'blockquote', 'pre', 'code']
ADMIN_TAGS = ARTICLE_TAGS + ['table', 'tr', 'td', 'th', 'thead', 'tbody', 'div', 'span']

def sanitise_content(html, content_type='comment'):
    tag_sets = {
        'comment': COMMENT_TAGS,
        'article': ARTICLE_TAGS,
        'admin': ADMIN_TAGS
    }
    return bleach.clean(html, tags=tag_sets.get(content_type, COMMENT_TAGS))

Attribute Filtering

Default Allowed Attributes

import bleach

# Bleach's default allowed attributes
print(bleach.ALLOWED_ATTRIBUTES)
# {'a': ['href', 'title'], 'abbr': ['title'], 'acronym': ['title']}

Custom Attribute Configuration

import bleach

# Attributes as dictionary
attributes = {
    'a': ['href', 'title', 'rel'],
    'img': ['src', 'alt', 'title', 'width', 'height'],
    'p': ['class'],
    '*': ['id']  # Allow 'id' on all tags
}

html = '<a href="http://example.com" onclick="bad()">Link</a>'
clean = bleach.clean(html, tags=['a'], attributes=attributes)
# Result: '<a href="http://example.com">Link</a>'

Callable Attribute Filter

import bleach

def filter_attributes(tag, name, value):
    """Custom attribute filter function."""
    # Allow href on anchors, but validate protocol
    if tag == 'a' and name == 'href':
        if value.startswith(('http://', 'https://', 'mailto:')):
            return True
        return False

    # Allow src on images with http(s) only
    if tag == 'img' and name == 'src':
        return value.startswith(('http://', 'https://'))

    # Allow class on any tag, but restrict values
    if name == 'class':
        # Only allow specific CSS classes
        allowed_classes = ['highlight', 'note', 'warning', 'code-block']
        return value in allowed_classes

    # Allow data attributes
    if name.startswith('data-'):
        return True

    return False

html = '<a href="javascript:alert(1)">Bad</a><a href="https://safe.com">Good</a>'
clean = bleach.clean(html, tags=['a'], attributes=filter_attributes)
# Result: '<a>Bad</a><a href="https://safe.com">Good</a>'

Combined Dictionary and Callable

import bleach

def style_filter(tag, name, value):
    """Filter style attributes to allow only safe properties."""
    if name != 'style':
        return False

    # Allow only specific CSS properties
    safe_properties = ['color', 'background-color', 'font-size', 'text-align']

    for prop in value.split(';'):
        if ':' in prop:
            property_name = prop.split(':')[0].strip()
            if property_name and property_name not in safe_properties:
                return False
    return True

# Combine dictionary and callable
attributes = {
    'a': ['href', 'title'],
    'img': ['src', 'alt'],
    'span': style_filter,
    'div': style_filter
}

Link Handling and Linkification

URL Protocol Filtering

import bleach

# Default allowed protocols
print(bleach.ALLOWED_PROTOCOLS)
# {'http', 'https', 'mailto'}

# Custom protocols
protocols = ['http', 'https', 'mailto', 'tel', 'ftp']

html = '<a href="javascript:alert(1)">Bad</a><a href="tel:+441234567890">Call</a>'
clean = bleach.clean(
    html,
    tags=['a'],
    attributes={'a': ['href']},
    protocols=protocols
)
# Result: '<a>Bad</a><a href="tel:+441234567890">Call</a>'

Linkification with linkify()

import bleach

# Convert URLs to links automatically
text = "Check out https://example.com for more info"
linked = bleach.linkify(text)
# Result: 'Check out <a href="https://example.com" rel="nofollow">https://example.com</a> for more info'

# Linkify email addresses
text = "Contact us at support@example.com"
linked = bleach.linkify(text, parse_email=True)
# Result: 'Contact us at <a href="mailto:support@example.com">support@example.com</a>'

Linkify Callbacks

import bleach
from bleach.linkifier import LinkifyFilter

def set_target(attrs, new=False):
    """Add target="_blank" to external links."""
    if new:  # This is a new link (from linkification)
        attrs[(None, 'target')] = '_blank'
        attrs[(None, 'rel')] = 'noopener noreferrer'
    return attrs

def shorten_url(attrs, new=False):
    """Shorten displayed URL text."""
    if new:
        text = attrs.get('_text', '')
        if len(text) > 30:
            attrs['_text'] = text[:27] + '...'
    return attrs

def skip_internal(attrs, new=False):
    """Skip linkification for internal URLs."""
    if new:
        href = attrs.get((None, 'href'), '')
        if 'internal-site.com' in href:
            return None  # Don't create link
    return attrs

# Apply callbacks
text = "Visit https://example.com/very/long/path/to/page for details"
linked = bleach.linkify(
    text,
    callbacks=[set_target, shorten_url]
)

Combining clean() and linkify()

import bleach

def sanitise_and_linkify(html):
    """Clean HTML and convert URLs to links."""
    # First, clean the HTML
    allowed_tags = ['p', 'br', 'strong', 'em', 'a', 'ul', 'ol', 'li']
    allowed_attrs = {'a': ['href', 'title', 'rel']}

    cleaned = bleach.clean(
        html,
        tags=allowed_tags,
        attributes=allowed_attrs,
        strip=True
    )

    # Then linkify any plain URLs
    def add_nofollow(attrs, new=False):
        attrs[(None, 'rel')] = 'nofollow'
        return attrs

    return bleach.linkify(
        cleaned,
        callbacks=[add_nofollow],
        skip_tags=['pre', 'code']  # Don't linkify inside code blocks
    )

# Usage
user_content = '<p>Check https://example.com</p><script>bad()</script>'
safe_content = sanitise_and_linkify(user_content)

Custom Filters and Callbacks

Using Cleaner Class

import bleach
from bleach import Cleaner

# Create reusable cleaner instance
cleaner = Cleaner(
    tags=['p', 'br', 'a', 'strong', 'em'],
    attributes={'a': ['href', 'title']},
    protocols=['http', 'https'],
    strip=True,
    strip_comments=True
)

# Use the cleaner
html = '<p>Hello <script>bad()</script><a href="https://example.com">World</a></p>'
clean = cleaner.clean(html)

Custom Filters with html5lib

import bleach
from bleach import Cleaner
from bleach.html5lib_shim import Filter

class AddRelNofollow(Filter):
    """Add rel="nofollow" to all links."""

    def __iter__(self):
        for token in Filter.__iter__(self):
            if token['type'] == 'StartTag' and token['name'] == 'a':
                token['data'][(None, 'rel')] = 'nofollow'
            yield token

class RemoveEmptyTags(Filter):
    """Remove empty paragraph and div tags."""

    def __iter__(self):
        buffer = []
        for token in Filter.__iter__(self):
            if token['type'] == 'StartTag' and token['name'] in ('p', 'div'):
                buffer.append(token)
            elif token['type'] == 'EndTag' and token['name'] in ('p', 'div'):
                if buffer:
                    start_token = buffer.pop()
                    # Check if there was content between start and end
                    # This is a simplified check
                    continue
            else:
                if buffer:
                    for buffered in buffer:
                        yield buffered
                    buffer = []
                yield token

# Apply custom filters
cleaner = Cleaner(
    tags=['p', 'a', 'div', 'strong'],
    attributes={'a': ['href']},
    filters=[AddRelNofollow]
)

Post-Processing Callbacks

import bleach
import re

def post_process_html(html):
    """Apply additional transformations after cleaning."""
    # Convert newlines to <br> tags
    html = html.replace('\n', '<br>')

    # Auto-link @mentions
    html = re.sub(
        r'@(\w+)',
        r'<a href="/user/\1">@\1</a>',
        html
    )

    # Auto-link #hashtags
    html = re.sub(
        r'#(\w+)',
        r'<a href="/tag/\1">#\1</a>',
        html
    )

    return html

def sanitise_social_content(content):
    """Full pipeline for social media-style content."""
    # Clean
    clean = bleach.clean(
        content,
        tags=['p', 'br', 'a', 'strong', 'em'],
        attributes={'a': ['href']},
        strip=True
    )

    # Linkify URLs
    linked = bleach.linkify(clean)

    # Post-process for mentions and hashtags
    final = post_process_html(linked)

    return final

Common Use Cases

User-Generated Content

import bleach

class ContentSanitiser:
    """Sanitiser for user-generated content."""

    # Configuration for different content types
    CONFIGS = {
        'comment': {
            'tags': ['p', 'br', 'strong', 'em', 'a', 'code'],
            'attributes': {'a': ['href', 'title']},
            'strip': True
        },
        'post': {
            'tags': ['p', 'br', 'strong', 'em', 'a', 'img',
                    'h2', 'h3', 'ul', 'ol', 'li', 'blockquote',
                    'pre', 'code'],
            'attributes': {
                'a': ['href', 'title'],
                'img': ['src', 'alt', 'title']
            },
            'strip': True
        },
        'bio': {
            'tags': ['br', 'strong', 'em', 'a'],
            'attributes': {'a': ['href']},
            'strip': True
        }
    }

    @classmethod
    def sanitise(cls, content, content_type='comment'):
        config = cls.CONFIGS.get(content_type, cls.CONFIGS['comment'])
        return bleach.clean(content, **config)

    @classmethod
    def sanitise_and_linkify(cls, content, content_type='comment'):
        cleaned = cls.sanitise(content, content_type)
        return bleach.linkify(cleaned)

# Usage
user_comment = '<p>Check out <script>alert("xss")</script>https://example.com</p>'
safe_comment = ContentSanitiser.sanitise_and_linkify(user_comment, 'comment')

Web Application Integration

Flask Integration

from flask import Flask, request
from markupsafe import Markup
import bleach

app = Flask(__name__)

def clean_html(html):
    """Template filter for cleaning HTML."""
    cleaned = bleach.clean(
        html,
        tags=['p', 'br', 'strong', 'em', 'a'],
        attributes={'a': ['href']},
        strip=True
    )
    return Markup(cleaned)

# Register as template filter
app.jinja_env.filters['clean_html'] = clean_html

@app.route('/comment', methods=['POST'])
def create_comment():
    content = request.form.get('content', '')

    # Sanitise before storing
    safe_content = bleach.clean(
        content,
        tags=['p', 'br', 'strong', 'em', 'a'],
        attributes={'a': ['href']},
        strip=True
    )

    # Store safe_content in database
    # ...

    return {'status': 'success'}

Django Integration

# forms.py
from django import forms
import bleach

class CommentForm(forms.Form):
    content = forms.CharField(widget=forms.Textarea)

    def clean_content(self):
        content = self.cleaned_data['content']
        return bleach.clean(
            content,
            tags=['p', 'br', 'strong', 'em', 'a'],
            attributes={'a': ['href']},
            strip=True
        )

# templatetags/bleach_tags.py
from django import template
from django.utils.safestring import mark_safe
import bleach

register = template.Library()

@register.filter(is_safe=True)
def bleach_clean(value):
    """Clean HTML in templates."""
    cleaned = bleach.clean(
        value,
        tags=['p', 'br', 'strong', 'em', 'a'],
        attributes={'a': ['href']},
        strip=True
    )
    return mark_safe(cleaned)

@register.filter(is_safe=True)
def bleach_linkify(value):
    """Linkify URLs in templates."""
    return mark_safe(bleach.linkify(value))

FastAPI Integration

from fastapi import FastAPI, Form
from pydantic import BaseModel, field_validator
import bleach

app = FastAPI()

class Comment(BaseModel):
    content: str

    @field_validator('content')
    @classmethod
    def sanitise_content(cls, v):
        return bleach.clean(
            v,
            tags=['p', 'br', 'strong', 'em', 'a'],
            attributes={'a': ['href']},
            strip=True
        )

@app.post('/comments')
async def create_comment(comment: Comment):
    # comment.content is already sanitised
    return {'content': comment.content}

Markdown with Bleach

import bleach
import markdown

def safe_markdown(text):
    """Convert markdown to HTML and sanitise."""
    # Convert markdown to HTML
    html = markdown.markdown(
        text,
        extensions=['fenced_code', 'tables', 'nl2br']
    )

    # Sanitise the HTML output
    allowed_tags = [
        'p', 'br', 'strong', 'em', 'a', 'img',
        'h1', 'h2', 'h3', 'h4', 'h5', 'h6',
        'ul', 'ol', 'li',
        'pre', 'code',
        'blockquote',
        'table', 'thead', 'tbody', 'tr', 'th', 'td',
        'hr'
    ]

    allowed_attrs = {
        'a': ['href', 'title'],
        'img': ['src', 'alt', 'title'],
        'code': ['class'],  # For syntax highlighting
        '*': ['id']
    }

    return bleach.clean(
        html,
        tags=allowed_tags,
        attributes=allowed_attrs,
        strip=True
    )

Security Best Practices

Essential Security Rules

RecommendedAlternativeYesNoUser InputSanitise on Input?Store Raw +SanitisedSanitise on OutputDisplay SanitisedContext-Aware?Secure DisplayPotential XSSRecommendedAlternativeYesNoUser InputSanitise on Input?Store Raw +SanitisedSanitise on OutputDisplay SanitisedContext-Aware?Secure DisplayPotential XSS

Defence in Depth

import bleach
from html import escape

class SecureContent:
    """Secure content handling with multiple layers."""

    # Strict default configuration
    STRICT_TAGS = ['p', 'br', 'strong', 'em']
    STRICT_ATTRS = {}

    @classmethod
    def sanitise(cls, content, tags=None, attributes=None):
        """Sanitise with secure defaults."""
        # Never allow None - use strict defaults
        tags = tags if tags is not None else cls.STRICT_TAGS
        attributes = attributes if attributes is not None else cls.STRICT_ATTRS

        # Always strip dangerous content
        cleaned = bleach.clean(
            content,
            tags=tags,
            attributes=attributes,
            protocols=['https', 'mailto'],  # No http
            strip=True,
            strip_comments=True
        )

        return cleaned

    @classmethod
    def escape_all(cls, content):
        """Complete escape - no HTML allowed."""
        return escape(content)

    @classmethod
    def sanitise_url(cls, url):
        """Sanitise URL for use in href/src."""
        # Only allow safe protocols
        safe_protocols = ('https://', 'mailto:', 'tel:')
        if not any(url.startswith(p) for p in safe_protocols):
            return '#'

        # Additional URL validation here
        return url

Preventing Common Attacks

import bleach
import re

def prevent_data_exfiltration(attrs, new=False):
    """Prevent data exfiltration via links."""
    href = attrs.get((None, 'href'), '')

    # Block data: URLs
    if href.startswith('data:'):
        return None

    # Block URLs with credentials
    if '@' in href and '://' in href:
        return None

    return attrs

def sanitise_for_display(content):
    """Full security sanitisation."""
    # Step 1: Clean HTML
    cleaned = bleach.clean(
        content,
        tags=['p', 'br', 'strong', 'em', 'a', 'code'],
        attributes={'a': ['href', 'title']},
        protocols=['https'],
        strip=True,
        strip_comments=True
    )

    # Step 2: Linkify with security callbacks
    linked = bleach.linkify(
        cleaned,
        callbacks=[prevent_data_exfiltration],
        skip_tags=['code', 'pre']
    )

    # Step 3: Additional security checks
    # Remove any remaining javascript: (belt and braces)
    linked = re.sub(r'javascript:', '', linked, flags=re.IGNORECASE)

    return linked

Content Security Policy Integration

import bleach

def get_csp_compatible_cleaner():
    """Create cleaner compatible with strict CSP."""

    def no_inline_styles(tag, name, value):
        """Remove all inline styles for CSP compliance."""
        if name == 'style':
            return False
        if name == 'class':
            # Whitelist specific classes
            allowed = ['highlight', 'code', 'note', 'warning']
            return value in allowed
        return True

    return bleach.Cleaner(
        tags=['p', 'br', 'strong', 'em', 'a', 'code', 'pre', 'span'],
        attributes={
            'a': ['href', 'title'],
            'span': no_inline_styles,
            'code': no_inline_styles
        },
        protocols=['https', 'mailto'],
        strip=True
    )

# Usage
csp_cleaner = get_csp_compatible_cleaner()
safe_html = csp_cleaner.clean(user_input)

Performance Considerations

Optimisation Strategies

import bleach
from functools import lru_cache

# Create reusable Cleaner instances (avoid recreating)
COMMENT_CLEANER = bleach.Cleaner(
    tags=['p', 'br', 'strong', 'em', 'a'],
    attributes={'a': ['href']},
    strip=True
)

ARTICLE_CLEANER = bleach.Cleaner(
    tags=['p', 'br', 'strong', 'em', 'a', 'img', 'h2', 'h3',
          'ul', 'ol', 'li', 'blockquote', 'pre', 'code'],
    attributes={'a': ['href'], 'img': ['src', 'alt']},
    strip=True
)

def sanitise_comment(content):
    """Use pre-configured cleaner."""
    return COMMENT_CLEANER.clean(content)

def sanitise_article(content):
    """Use pre-configured cleaner."""
    return ARTICLE_CLEANER.clean(content)

Caching Sanitised Content

import bleach
import hashlib
from functools import lru_cache

@lru_cache(maxsize=1000)
def cached_sanitise(content_hash, content):
    """Cache sanitisation results."""
    return bleach.clean(
        content,
        tags=['p', 'br', 'strong', 'em', 'a'],
        attributes={'a': ['href']},
        strip=True
    )

def sanitise_with_cache(content):
    """Sanitise with caching."""
    content_hash = hashlib.md5(content.encode()).hexdigest()
    return cached_sanitise(content_hash, content)

Batch Processing

import bleach
from concurrent.futures import ThreadPoolExecutor

# Single cleaner instance for batch processing
BATCH_CLEANER = bleach.Cleaner(
    tags=['p', 'br', 'strong', 'em', 'a'],
    attributes={'a': ['href']},
    strip=True
)

def batch_sanitise(contents, max_workers=4):
    """Sanitise multiple contents in parallel."""
    with ThreadPoolExecutor(max_workers=max_workers) as executor:
        results = list(executor.map(BATCH_CLEANER.clean, contents))
    return results

# Usage
user_posts = ['<p>Post 1</p>', '<p>Post 2</p>', '<p>Post 3</p>']
safe_posts = batch_sanitise(user_posts)

Performance Tips

Technique Benefit When to Use
Reuse Cleaner instances Avoid parser reinitialisation Always
Cache results Avoid re-sanitising same content Repeated display
Batch processing Parallel sanitisation Multiple items
Minimal tag sets Faster parsing When possible
Store sanitised + raw Avoid runtime sanitisation Database storage

Quick Reference

Essential Functions

Function Purpose Example
bleach.clean() Sanitise HTML bleach.clean('<script>bad</script>')
bleach.linkify() Convert URLs to links bleach.linkify('Visit https://example.com')
bleach.Cleaner() Create reusable cleaner cleaner = Cleaner(tags=['p'])

Common Parameters

Parameter Type Default Description
tags list ALLOWED_TAGS Tags to allow
attributes dict/callable ALLOWED_ATTRIBUTES Attributes to allow
protocols list ALLOWED_PROTOCOLS URL protocols to allow
strip bool False Strip disallowed tags
strip_comments bool False Remove HTML comments

Defaults

# Default allowed tags
ALLOWED_TAGS = {'a', 'abbr', 'acronym', 'b', 'blockquote',
                'code', 'em', 'i', 'li', 'ol', 'strong', 'ul'}

# Default allowed attributes
ALLOWED_ATTRIBUTES = {
    'a': ['href', 'title'],
    'abbr': ['title'],
    'acronym': ['title']
}

# Default allowed protocols
ALLOWED_PROTOCOLS = {'http', 'https', 'mailto'}

Common Configurations

# Minimal (comments)
bleach.clean(html, tags=['p', 'br', 'strong', 'em', 'a'],
             attributes={'a': ['href']}, strip=True)

# Standard (posts)
bleach.clean(html, tags=['p', 'br', 'strong', 'em', 'a', 'img',
             'h2', 'h3', 'ul', 'ol', 'li', 'blockquote'],
             attributes={'a': ['href'], 'img': ['src', 'alt']}, strip=True)

# Linkify with security
bleach.linkify(text, callbacks=[lambda a, n: {**a, (None, 'rel'): 'nofollow'}])

Common Issues and Solutions

Issue: Tags Being Escaped Instead of Stripped

Problem: Tags appear as &lt;script&gt; instead of being removed.

Solution: Use strip=True:

# Wrong - escapes tags
result = bleach.clean('<script>bad</script>')
# Result: '&lt;script&gt;bad&lt;/script&gt;'

# Correct - strips tags
result = bleach.clean('<script>bad</script>', strip=True)
# Result: 'bad'

Issue: Attributes Being Removed

Problem: Safe attributes like href are being removed.

Solution: Explicitly allow attributes:

# Wrong - uses default attributes
result = bleach.clean('<a href="https://example.com" class="btn">Link</a>',
                      tags=['a'])
# 'href' kept but 'class' removed (not in defaults)

# Correct - specify needed attributes
result = bleach.clean('<a href="https://example.com" class="btn">Link</a>',
                      tags=['a'],
                      attributes={'a': ['href', 'class']})

Issue: Links with javascript: Protocol

Problem: javascript: URLs are being allowed.

Solution: Check protocol configuration:

# Bleach blocks javascript: by default
# If it's passing through, ensure you haven't added it to protocols

# Correct configuration
result = bleach.clean(
    '<a href="javascript:alert(1)">Click</a>',
    tags=['a'],
    attributes={'a': ['href']},
    protocols=['https', 'mailto']  # No javascript
)
# Result: '<a>Click</a>'

Issue: Empty Tags in Output

Problem: Empty tags like <p></p> appear after cleaning.

Solution: Post-process or use custom filter:

import re

def remove_empty_tags(html):
    """Remove empty paragraph and div tags."""
    pattern = r'<(p|div|span)>\s*</\1>'
    while re.search(pattern, html):
        html = re.sub(pattern, '', html)
    return html

cleaned = bleach.clean(html, tags=['p', 'div'], strip=True)
final = remove_empty_tags(cleaned)

Issue: Linkify Creating Unwanted Links

Problem: URLs in code blocks are being converted to links.

Solution: Use skip_tags:

result = bleach.linkify(
    '<code>https://example.com</code>',
    skip_tags=['code', 'pre']
)
# URL inside code remains as text

Issue: Performance Degradation

Problem: Sanitisation is slow with many items.

Solution: Reuse Cleaner instances:

# Wrong - creates new cleaner each time
def sanitise(content):
    return bleach.clean(content, tags=['p'], strip=True)

# Correct - reuse cleaner
CLEANER = bleach.Cleaner(tags=['p'], strip=True)

def sanitise(content):
    return CLEANER.clean(content)

Issue: Unicode/Encoding Errors

Problem: Non-ASCII characters cause errors.

Solution: Ensure proper encoding:

# Ensure string input (not bytes)
if isinstance(content, bytes):
    content = content.decode('utf-8')

result = bleach.clean(content)

Issue: HTML Entities Being Double-Escaped

Problem: &amp; becomes &amp;amp;.

Solution: This is expected behaviour for safety. Decode before cleaning if needed:

import html

# If input has HTML entities that should be preserved
content = '&amp; &lt; &gt;'
decoded = html.unescape(content)  # '& < >'
result = bleach.clean(decoded, strip=True)

Related Topics

  • HTML5lib - Underlying HTML parser used by Bleach
  • MarkupSafe - Safe string handling for templates
  • Jinja2 - Template engine with autoescaping
  • Django Security - Django's built-in XSS protection
  • Content Security Policy - Browser-level XSS mitigation
  • OWASP XSS Prevention - Comprehensive XSS prevention guidelines