Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Python Beautiful Soup

A powerful HTML and XML parsing library for extracting data from web pages and structured documents.

Python Beautiful Soup Cheatsheet

A powerful HTML and XML parsing library for extracting data from web pages and structured documents.

Overview

Beautiful Soup provides Pythonic idioms for navigating, searching, and modifying parse trees, making it the go-to library for web scraping and HTML/XML processing in Python.

OperationsParse TreeBeautifulSoup ParserInput SourcesHTML StringXML StringFile ObjectURL Responsehtml.parserlxmlhtml5liblxml-xmlTag ObjectsNavigableStringCommentSearch/FindNavigateExtractModifyOperationsParse TreeBeautifulSoup ParserInput SourcesHTML StringXML StringFile ObjectURL Responsehtml.parserlxmlhtml5liblxml-xmlTag ObjectsNavigableStringCommentSearch/FindNavigateExtractModify

Installation

uv add beautifulsoup4

# Install parsers (optional but recommended)
uv add lxml        # Fast, lenient HTML/XML parser
uv add html5lib    # Extremely lenient, parses like a browser

Parsing HTML/XML

Key Concepts

  • BeautifulSoup object represents the entire parsed document
  • Parsers differ in speed, leniency, and dependencies
  • Document encoding is automatically detected or can be specified
  • Parse tree is built from tags, navigable strings, and comments

Common Patterns

from bs4 import BeautifulSoup

# Parse HTML string
soup = BeautifulSoup(html_string, 'html.parser')

# Parse with lxml (faster)
soup = BeautifulSoup(html_string, 'lxml')

# Parse XML
soup = BeautifulSoup(xml_string, 'lxml-xml')  # or 'xml'

# Parse from file
with open('page.html', 'r', encoding='utf-8') as f:
    soup = BeautifulSoup(f, 'html.parser')

# Parse with specific encoding
soup = BeautifulSoup(html_bytes, 'html.parser', from_encoding='iso-8859-1')

Examples

Choosing the Right Parser

from bs4 import BeautifulSoup

html = """
<html>
    <head><title>Test Page</title></head>
    <body>
        <p>Hello <b>World</p>  <!-- Malformed HTML -->
    </body>
</html>
"""

# html.parser - Built-in, decent speed, lenient
soup = BeautifulSoup(html, 'html.parser')
print(soup.prettify())

# lxml - Fast, lenient (recommended for most cases)
soup = BeautifulSoup(html, 'lxml')
print(soup.prettify())

# html5lib - Slowest but parses exactly like browsers
soup = BeautifulSoup(html, 'html5lib')
print(soup.prettify())

# lxml-xml - For strict XML parsing
xml = '<root><item id="1">Value</item></root>'
soup = BeautifulSoup(xml, 'lxml-xml')
print(soup.prettify())

Parsing from Web Requests

import requests
from bs4 import BeautifulSoup

# Fetch and parse a webpage
response = requests.get('https://example.com')
response.raise_for_status()

# Parse the HTML content
soup = BeautifulSoup(response.content, 'lxml')

# Use response.content (bytes) instead of response.text
# to let BeautifulSoup handle encoding detection

# Access document metadata
print(f"Title: {soup.title.string}")
print(f"Encoding: {soup.original_encoding}")

Handling Encoding Issues

from bs4 import BeautifulSoup

# Let BeautifulSoup detect encoding
soup = BeautifulSoup(html_bytes, 'html.parser')
print(f"Detected encoding: {soup.original_encoding}")

# Force specific encoding
soup = BeautifulSoup(
    html_bytes,
    'html.parser',
    from_encoding='utf-8'
)

# Exclude wrong encodings from detection
soup = BeautifulSoup(
    html_bytes,
    'html.parser',
    exclude_encodings=['iso-8859-7']
)

Finding Elements

Key Concepts

  • find() returns the first matching element or None
  • find_all() returns a list of all matching elements
  • select() uses CSS selectors for powerful, flexible searches
  • Filters can be strings, regex, lists, functions, or True
FiltersStringRegexListFunctionTrueSearch Methodfindfind_allselectselect_oneFirst MatchAll MatchesCSS Selector AllCSS Selector FirstFiltersStringRegexListFunctionTrueSearch Methodfindfind_allselectselect_oneFirst MatchAll MatchesCSS Selector AllCSS Selector First

Common Patterns

from bs4 import BeautifulSoup
import re

# Find by tag name
soup.find('div')
soup.find_all('p')

# Find by attributes
soup.find('div', id='main')
soup.find_all('a', class_='link')
soup.find_all('input', attrs={'type': 'text'})

# Find by CSS selector
soup.select('div.container')
soup.select_one('#header')
soup.select('ul > li')

# Find with regex
soup.find_all('a', href=re.compile(r'^https://'))

# Find with function
soup.find_all(lambda tag: tag.has_attr('data-value'))

Examples

Basic Finding

from bs4 import BeautifulSoup

html = """
<html>
<body>
    <div id="content" class="main">
        <h1>Title</h1>
        <p class="intro">First paragraph</p>
        <p class="body">Second paragraph</p>
        <a href="/page1" class="link">Link 1</a>
        <a href="/page2" class="link external">Link 2</a>
    </div>
</body>
</html>
"""

soup = BeautifulSoup(html, 'lxml')

# Find first element
first_p = soup.find('p')
print(first_p.text)  # "First paragraph"

# Find all elements
all_links = soup.find_all('a')
for link in all_links:
    print(link['href'])

# Find by id
content = soup.find(id='content')
print(content.h1.text)  # "Title"

# Find by class (note: use class_ to avoid Python keyword conflict)
intro = soup.find('p', class_='intro')
print(intro.text)  # "First paragraph"

# Find multiple classes
external_link = soup.find('a', class_='link external')
print(external_link.text)  # "Link 2"

Advanced Filtering

from bs4 import BeautifulSoup
import re

html = """
<div>
    <a href="https://example.com">HTTPS Link</a>
    <a href="http://example.org">HTTP Link</a>
    <a href="/relative/path">Relative Link</a>
    <img src="image1.png" alt="Image 1" />
    <img src="image2.jpg" alt="Image 2" />
    <p data-value="100">Paragraph</p>
</div>
"""

soup = BeautifulSoup(html, 'lxml')

# Find with regex pattern
https_links = soup.find_all('a', href=re.compile(r'^https://'))
print(f"HTTPS links: {len(https_links)}")

# Find with list of values
images = soup.find_all('img', src=re.compile(r'\.(png|jpg)$'))
print(f"Images: {len(images)}")

# Find with custom function
def has_data_attribute(tag):
    return tag.has_attr('data-value')

data_elements = soup.find_all(has_data_attribute)
print(f"Elements with data-value: {len(data_elements)}")

# Find any tag with specific attribute
any_with_alt = soup.find_all(alt=True)
print(f"Elements with alt: {len(any_with_alt)}")

# Find tags by multiple criteria
def complex_filter(tag):
    return (tag.name == 'a' and
            tag.has_attr('href') and
            'example' in tag.get('href', ''))

example_links = soup.find_all(complex_filter)
print(f"Example links: {len(example_links)}")

CSS Selectors

from bs4 import BeautifulSoup

html = """
<div id="wrapper">
    <nav class="menu">
        <ul>
            <li class="active"><a href="/">Home</a></li>
            <li><a href="/about">About</a></li>
            <li><a href="/contact">Contact</a></li>
        </ul>
    </nav>
    <main>
        <article class="post featured">
            <h2>Featured Post</h2>
            <p>Content here</p>
        </article>
        <article class="post">
            <h2>Regular Post</h2>
            <p>More content</p>
        </article>
    </main>
</div>
"""

soup = BeautifulSoup(html, 'lxml')

# Select by ID
wrapper = soup.select_one('#wrapper')

# Select by class
posts = soup.select('.post')
print(f"Posts: {len(posts)}")

# Select by multiple classes
featured = soup.select('.post.featured')
print(f"Featured posts: {len(featured)}")

# Descendant selector
nav_links = soup.select('nav a')
print(f"Nav links: {len(nav_links)}")

# Direct child selector
list_items = soup.select('ul > li')
print(f"List items: {len(list_items)}")

# Attribute selectors
home_link = soup.select_one('a[href="/"]')
print(f"Home link: {home_link.text}")

# Attribute contains
about_links = soup.select('a[href*="about"]')

# Attribute starts with
internal_links = soup.select('a[href^="/"]')

# Nth-child
second_item = soup.select_one('li:nth-child(2)')
print(f"Second item: {second_item.text.strip()}")

# Combining selectors
active_link = soup.select_one('li.active > a')
print(f"Active link: {active_link.text}")

Limiting Results

from bs4 import BeautifulSoup

html = """
<ul>
    <li>Item 1</li>
    <li>Item 2</li>
    <li>Item 3</li>
    <li>Item 4</li>
    <li>Item 5</li>
</ul>
"""

soup = BeautifulSoup(html, 'lxml')

# Limit number of results
first_three = soup.find_all('li', limit=3)
print(f"First three items: {len(first_three)}")

# Find starting from specific element
ul = soup.find('ul')
items = ul.find_all('li', limit=2)

# Recursive search (default is True)
# Set to False to only search direct children
direct_children = soup.find_all('li', recursive=False)  # Returns []
direct_children = ul.find_all('li', recursive=False)     # Returns all li

Navigating the Tree

Key Concepts

  • Parent is the containing element
  • Children are direct descendants only
  • Descendants include all nested elements
  • Siblings are elements at the same level
  • Navigation returns None when element doesn't exist
ParentPrevious SiblingCurrent ElementNext SiblingChild 1Child 2GrandchildParentPrevious SiblingCurrent ElementNext SiblingChild 1Child 2Grandchild

Common Patterns

from bs4 import BeautifulSoup

# Parent navigation
element.parent              # Direct parent
element.parents             # Generator of all parents

# Child navigation
element.children            # Direct children generator
element.descendants         # All descendants generator
element.contents            # Direct children as list

# Sibling navigation
element.next_sibling        # Next sibling (may be whitespace)
element.previous_sibling    # Previous sibling
element.next_siblings       # Generator of following siblings
element.previous_siblings   # Generator of preceding siblings

# Element navigation (skips whitespace)
element.next_element        # Next element in parse order
element.previous_element    # Previous element
element.find_next()         # Next matching element
element.find_previous()     # Previous matching element
element.find_next_sibling() # Next matching sibling

Examples

Parent Navigation

from bs4 import BeautifulSoup

html = """
<html>
<body>
    <div id="outer">
        <div id="inner">
            <p id="target">Target paragraph</p>
        </div>
    </div>
</body>
</html>
"""

soup = BeautifulSoup(html, 'lxml')
target = soup.find(id='target')

# Get direct parent
parent = target.parent
print(f"Direct parent: {parent.name}, id={parent.get('id')}")
# Output: Direct parent: div, id=inner

# Iterate through all parents
for parent in target.parents:
    if parent.name:
        print(f"Parent: {parent.name}")
# Output: Parent: div, Parent: div, Parent: body, Parent: html, Parent: [document]

# Find specific parent
outer_div = target.find_parent('div', id='outer')
print(f"Found outer: {outer_div.get('id')}")

# Find all parent divs
parent_divs = target.find_parents('div')
print(f"Parent divs: {len(parent_divs)}")  # 2

Child Navigation

from bs4 import BeautifulSoup

html = """
<ul id="menu">
    <li>First</li>
    <li>Second
        <ul>
            <li>Nested</li>
        </ul>
    </li>
    <li>Third</li>
</ul>
"""

soup = BeautifulSoup(html, 'lxml')
menu = soup.find(id='menu')

# Get direct children as list
# Note: includes NavigableString (whitespace/text)
contents = menu.contents
print(f"Contents count: {len(contents)}")

# Iterate direct children (excludes NavigableString whitespace)
for child in menu.children:
    if child.name:  # Skip text nodes
        print(f"Child: {child.name}")

# Get all descendants (deep traversal)
for descendant in menu.descendants:
    if hasattr(descendant, 'name') and descendant.name:
        print(f"Descendant: {descendant.name}")

# Find all descendants with specific tag
all_li = menu.find_all('li')  # Includes nested li
print(f"All li elements: {len(all_li)}")  # 4

# Find only direct children
direct_li = menu.find_all('li', recursive=False)
print(f"Direct li children: {len(direct_li)}")  # 3

Sibling Navigation

from bs4 import BeautifulSoup

html = """
<div>
    <h1>Title</h1>
    <p id="first">First paragraph</p>
    <p id="second">Second paragraph</p>
    <p id="third">Third paragraph</p>
    <footer>Footer</footer>
</div>
"""

soup = BeautifulSoup(html, 'lxml')
second = soup.find(id='second')

# Next sibling (may be whitespace NavigableString)
next_sib = second.next_sibling
while next_sib and not hasattr(next_sib, 'name'):
    next_sib = next_sib.next_sibling
print(f"Next sibling: {next_sib.get('id')}")  # third

# Previous sibling
prev_sib = second.previous_sibling
while prev_sib and not hasattr(prev_sib, 'name'):
    prev_sib = prev_sib.previous_sibling
print(f"Previous sibling: {prev_sib.get('id')}")  # first

# Better: use find_next_sibling/find_previous_sibling
next_p = second.find_next_sibling('p')
print(f"Next p: {next_p.get('id')}")  # third

prev_p = second.find_previous_sibling('p')
print(f"Previous p: {prev_p.get('id')}")  # first

# Iterate all following siblings
for sibling in second.next_siblings:
    if hasattr(sibling, 'name') and sibling.name:
        print(f"Following: {sibling.name}")
# Output: Following: p, Following: footer

# Find next/previous element regardless of hierarchy
next_elem = second.find_next('footer')
print(f"Next footer: {next_elem.text}")

Tree Traversal Example

from bs4 import BeautifulSoup

html = """
<table>
    <tr>
        <th>Name</th>
        <th>Age</th>
    </tr>
    <tr>
        <td>Alice</td>
        <td>30</td>
    </tr>
    <tr>
        <td>Bob</td>
        <td>25</td>
    </tr>
</table>
"""

soup = BeautifulSoup(html, 'lxml')
table = soup.find('table')

# Get all rows
rows = table.find_all('tr')

# Process header row
header_row = rows[0]
headers = [th.text for th in header_row.find_all('th')]
print(f"Headers: {headers}")

# Process data rows
data = []
for row in rows[1:]:
    cells = row.find_all('td')
    row_data = {
        headers[i]: cells[i].text
        for i in range(len(cells))
    }
    data.append(row_data)

print(f"Data: {data}")
# Output: [{'Name': 'Alice', 'Age': '30'}, {'Name': 'Bob', 'Age': '25'}]

Extracting Data

Key Concepts

  • text/get_text() extracts visible text content
  • string gets the single child NavigableString
  • strings/stripped_strings generators for text content
  • attrs dictionary of all attributes
  • get() safely retrieves attribute values

Common Patterns

from bs4 import BeautifulSoup

# Get text content
element.text                    # All text, whitespace preserved
element.get_text()              # Same as .text
element.get_text(strip=True)    # Strip whitespace
element.get_text(separator=' ') # Join text with separator
element.string                  # Single child string only

# Get attributes
element['href']                 # Direct access (raises KeyError)
element.get('href')             # Safe access (returns None)
element.get('href', '#')        # With default value
element.attrs                   # All attributes as dict
element.has_attr('class')       # Check attribute exists

# Get name
element.name                    # Tag name as string

Examples

Extracting Text

from bs4 import BeautifulSoup

html = """
<div id="content">
    <h1>Welcome</h1>
    <p>This is a <strong>test</strong> paragraph.</p>
    <ul>
        <li>Item 1</li>
        <li>Item 2</li>
    </ul>
</div>
"""

soup = BeautifulSoup(html, 'lxml')
content = soup.find(id='content')

# Get all text (preserves whitespace)
all_text = content.text
print(repr(all_text))

# Get text with stripped whitespace
clean_text = content.get_text(strip=True)
print(clean_text)

# Get text with custom separator
spaced_text = content.get_text(separator=' ', strip=True)
print(spaced_text)
# Output: "Welcome This is a test paragraph. Item 1 Item 2"

# Iterate over strings
for string in content.strings:
    print(repr(string))

# Get stripped strings (no whitespace-only strings)
text_parts = list(content.stripped_strings)
print(text_parts)
# Output: ['Welcome', 'This is a', 'test', 'paragraph.', 'Item 1', 'Item 2']

# Get string from element with single text child
h1 = soup.find('h1')
print(h1.string)  # "Welcome"

# string returns None if multiple children
p = soup.find('p')
print(p.string)   # None (has multiple children)
print(p.text)     # "This is a test paragraph."

Extracting Attributes

from bs4 import BeautifulSoup

html = """
<div>
    <a href="https://example.com"
       class="link external"
       id="main-link"
       data-value="123"
       title="Example Link">Visit Example</a>
    <img src="image.png" alt="Test Image" />
</div>
"""

soup = BeautifulSoup(html, 'lxml')
link = soup.find('a')

# Direct attribute access
href = link['href']
print(f"URL: {href}")

# Safe access with get()
data_value = link.get('data-value')
print(f"Data value: {data_value}")

# With default
target = link.get('target', '_self')
print(f"Target: {target}")

# Get all attributes
print(f"All attributes: {link.attrs}")
# {'href': 'https://example.com', 'class': ['link', 'external'],
#  'id': 'main-link', 'data-value': '123', 'title': 'Example Link'}

# Note: class is always a list
classes = link.get('class', [])
print(f"Classes: {classes}")  # ['link', 'external']

# Check if attribute exists
if link.has_attr('title'):
    print(f"Title: {link['title']}")

# Get image attributes
img = soup.find('img')
print(f"Image src: {img.get('src')}")
print(f"Image alt: {img.get('alt')}")

Extracting Links and Images

from bs4 import BeautifulSoup
from urllib.parse import urljoin

html = """
<html>
<head>
    <base href="https://example.com/">
</head>
<body>
    <nav>
        <a href="/home">Home</a>
        <a href="/about">About</a>
        <a href="https://external.com">External</a>
    </nav>
    <article>
        <img src="images/photo.jpg" alt="Photo" />
        <p>Read more at <a href="/articles/1">Article 1</a></p>
    </article>
</body>
</html>
"""

soup = BeautifulSoup(html, 'lxml')
base_url = 'https://example.com'

# Extract all links with absolute URLs
links = []
for a in soup.find_all('a', href=True):
    url = a['href']
    absolute_url = urljoin(base_url, url)
    links.append({
        'text': a.get_text(strip=True),
        'url': absolute_url,
        'is_external': not absolute_url.startswith(base_url)
    })

for link in links:
    print(f"{link['text']}: {link['url']} (external: {link['is_external']})")

# Extract all images
images = []
for img in soup.find_all('img'):
    images.append({
        'src': urljoin(base_url, img.get('src', '')),
        'alt': img.get('alt', '')
    })

print(f"\nImages: {images}")

Extracting Structured Data

from bs4 import BeautifulSoup

html = """
<div class="product-list">
    <div class="product" data-id="1">
        <h3 class="name">Laptop</h3>
        <span class="price">$999.99</span>
        <p class="description">High-performance laptop</p>
        <span class="stock in-stock">In Stock</span>
    </div>
    <div class="product" data-id="2">
        <h3 class="name">Mouse</h3>
        <span class="price">$29.99</span>
        <p class="description">Wireless mouse</p>
        <span class="stock out-of-stock">Out of Stock</span>
    </div>
</div>
"""

soup = BeautifulSoup(html, 'lxml')

products = []
for product_div in soup.select('.product'):
    product = {
        'id': product_div.get('data-id'),
        'name': product_div.select_one('.name').get_text(strip=True),
        'price': product_div.select_one('.price').get_text(strip=True),
        'description': product_div.select_one('.description').get_text(strip=True),
        'in_stock': 'in-stock' in product_div.select_one('.stock').get('class', [])
    }
    products.append(product)

for product in products:
    print(f"{product['name']}: {product['price']} - In stock: {product['in_stock']}")

Modifying the Tree

Key Concepts

  • append() adds content to the end of an element
  • insert() adds content at a specific position
  • replace_with() replaces an element with new content
  • decompose() removes element from tree and destroys it
  • extract() removes element but keeps it usable
  • wrap() wraps element in another tag
  • unwrap() replaces tag with its contents

Common Patterns

from bs4 import BeautifulSoup

# Adding content
tag.append(content)             # Add to end
tag.insert(position, content)   # Add at position
tag.insert_before(content)      # Add before element
tag.insert_after(content)       # Add after element

# Removing content
tag.decompose()                 # Remove and destroy
tag.extract()                   # Remove but keep
tag.clear()                     # Remove all children

# Replacing content
tag.replace_with(new_content)   # Replace element
tag.string.replace_with(text)   # Replace text

# Wrapping/unwrapping
tag.wrap(wrapper_tag)           # Wrap in another tag
tag.unwrap()                    # Remove tag, keep contents

# Creating new elements
soup.new_tag('div', id='new')   # Create new tag
NavigableString('text')         # Create text node

Examples

Adding Content

from bs4 import BeautifulSoup
from bs4 import NavigableString

html = """
<div id="container">
    <p>First paragraph</p>
</div>
"""

soup = BeautifulSoup(html, 'lxml')
container = soup.find(id='container')

# Create and append new tag
new_p = soup.new_tag('p')
new_p.string = 'Second paragraph'
container.append(new_p)

# Create tag with attributes
new_link = soup.new_tag('a', href='https://example.com', class_='link')
new_link.string = 'Click here'
container.append(new_link)

# Insert at specific position
header = soup.new_tag('h2')
header.string = 'Section Title'
container.insert(0, header)  # Insert at beginning

# Append string content
container.append(NavigableString(' Some text '))

# Insert before/after
footer = soup.new_tag('footer')
footer.string = 'Footer content'
new_p.insert_after(footer)

print(container.prettify())

Removing Content

from bs4 import BeautifulSoup

html = """
<div id="content">
    <p class="keep">Keep this</p>
    <p class="remove">Remove this</p>
    <script>alert('Remove scripts')</script>
    <style>.remove { color: red; }</style>
    <p class="keep">Keep this too</p>
    <!-- Remove comments -->
</div>
"""

soup = BeautifulSoup(html, 'lxml')

# Remove specific elements
for script in soup.find_all('script'):
    script.decompose()

for style in soup.find_all('style'):
    style.decompose()

# Remove elements by class
for elem in soup.find_all(class_='remove'):
    elem.decompose()

# Remove comments
from bs4 import Comment
for comment in soup.find_all(string=lambda text: isinstance(text, Comment)):
    comment.extract()

# Extract (remove but keep for later use)
content = soup.find(id='content')
kept_paragraphs = []
for p in content.find_all('p', class_='keep'):
    kept_paragraphs.append(p.extract())

print("Extracted paragraphs:")
for p in kept_paragraphs:
    print(f"  - {p.text}")

# Clear all children
# content.clear()  # Would remove all children

print("\nRemaining content:")
print(soup.prettify())

Replacing Content

from bs4 import BeautifulSoup

html = """
<div id="article">
    <h1>Old Title</h1>
    <p>Paragraph with <em>emphasis</em> and <strong>bold</strong> text.</p>
    <a href="http://old-url.com">Old Link</a>
</div>
"""

soup = BeautifulSoup(html, 'lxml')

# Replace entire element
old_h1 = soup.find('h1')
new_h1 = soup.new_tag('h1')
new_h1.string = 'New Title'
old_h1.replace_with(new_h1)

# Replace with multiple elements
em = soup.find('em')
new_content = soup.new_tag('strong')
new_content.string = em.string.upper()
em.replace_with(new_content)

# Replace text content
for strong in soup.find_all('strong'):
    if strong.string:
        strong.string.replace_with(strong.string.upper())

# Replace attribute
link = soup.find('a')
link['href'] = 'https://new-url.com'
link.string = 'New Link'

print(soup.prettify())

Wrapping and Unwrapping

from bs4 import BeautifulSoup

html = """
<div>
    <p>Unwrapped paragraph</p>
    <span><strong>Bold text</strong></span>
</div>
"""

soup = BeautifulSoup(html, 'lxml')

# Wrap element in new tag
p = soup.find('p')
wrapper = soup.new_tag('section', class_='content')
p.wrap(wrapper)

# Wrap multiple elements
# First, create a wrapper
div = soup.find('div')
container = soup.new_tag('article')

# Move children to new container
children = list(div.children)
for child in children:
    if child.name:
        child.wrap(container.new_tag('div'))

# Unwrap (remove tag but keep contents)
span = soup.find('span')
span.unwrap()  # Removes <span>, keeps <strong>Bold text</strong>

print(soup.prettify())

Building Documents from Scratch

from bs4 import BeautifulSoup

# Start with minimal HTML
soup = BeautifulSoup('<html><body></body></html>', 'lxml')

# Add head section
head = soup.new_tag('head')
soup.html.insert(0, head)

# Add title
title = soup.new_tag('title')
title.string = 'Generated Page'
head.append(title)

# Add meta tags
meta = soup.new_tag('meta', charset='utf-8')
head.insert(0, meta)

# Build body content
body = soup.body

# Add header
header = soup.new_tag('header')
h1 = soup.new_tag('h1')
h1.string = 'Welcome'
header.append(h1)
body.append(header)

# Add main content
main = soup.new_tag('main')
for i in range(3):
    p = soup.new_tag('p')
    p.string = f'Paragraph {i + 1}'
    main.append(p)
body.append(main)

# Add footer
footer = soup.new_tag('footer')
footer.string = 'Copyright 2025'
body.append(footer)

# Output formatted HTML
print(soup.prettify())

Common Patterns for Web Scraping

Key Concepts

  • Respect robots.txt and terms of service
  • Rate limiting prevents server overload
  • User-Agent headers identify your scraper
  • Error handling manages network and parsing issues
  • Data validation ensures extracted data quality

Examples

Basic Web Scraper

import requests
from bs4 import BeautifulSoup
import time

class SimpleScraper:
    def __init__(self, base_url, delay=1):
        self.base_url = base_url
        self.delay = delay
        self.session = requests.Session()
        self.session.headers.update({
            'User-Agent': 'Mozilla/5.0 (compatible; MyScraper/1.0)'
        })

    def fetch_page(self, url):
        """Fetch a page with rate limiting and error handling."""
        try:
            response = self.session.get(url, timeout=10)
            response.raise_for_status()
            time.sleep(self.delay)  # Rate limiting
            return BeautifulSoup(response.content, 'lxml')
        except requests.RequestException as e:
            print(f"Error fetching {url}: {e}")
            return None

    def scrape_article(self, url):
        """Extract article data from a page."""
        soup = self.fetch_page(url)
        if not soup:
            return None

        article = soup.find('article') or soup.find(class_='post')
        if not article:
            return None

        return {
            'title': self._get_text(article, 'h1'),
            'content': self._get_text(article, '.content'),
            'author': self._get_text(article, '.author'),
            'date': self._get_text(article, '.date'),
            'url': url
        }

    def _get_text(self, element, selector):
        """Safely extract text from element."""
        found = element.select_one(selector)
        return found.get_text(strip=True) if found else ''

# Usage
scraper = SimpleScraper('https://example.com', delay=2)
article = scraper.scrape_article('https://example.com/article/1')
if article:
    print(f"Title: {article['title']}")

Pagination Handling

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

def scrape_paginated_list(base_url, list_selector, item_selector, max_pages=10):
    """Scrape items from paginated list pages."""
    session = requests.Session()
    session.headers.update({
        'User-Agent': 'Mozilla/5.0 (compatible; MyScraper/1.0)'
    })

    all_items = []
    current_url = base_url
    page = 1

    while current_url and page <= max_pages:
        print(f"Scraping page {page}: {current_url}")

        response = session.get(current_url, timeout=10)
        response.raise_for_status()
        soup = BeautifulSoup(response.content, 'lxml')

        # Extract items from current page
        container = soup.select_one(list_selector)
        if not container:
            break

        items = container.select(item_selector)
        for item in items:
            item_data = {
                'title': item.select_one('h3').get_text(strip=True) if item.select_one('h3') else '',
                'link': urljoin(base_url, item.select_one('a')['href']) if item.select_one('a') else ''
            }
            all_items.append(item_data)

        # Find next page link
        next_link = soup.select_one('a.next, a[rel="next"], .pagination a:last-child')
        if next_link and 'href' in next_link.attrs:
            current_url = urljoin(base_url, next_link['href'])
            page += 1
        else:
            break

    return all_items

# Usage
items = scrape_paginated_list(
    'https://example.com/products',
    '.product-list',
    '.product-item',
    max_pages=5
)
print(f"Scraped {len(items)} items")

Table Extraction

from bs4 import BeautifulSoup
import csv

def extract_table_data(html, table_selector='table'):
    """Extract data from HTML tables."""
    soup = BeautifulSoup(html, 'lxml')
    tables = soup.select(table_selector)
    results = []

    for table in tables:
        # Extract headers
        headers = []
        header_row = table.select_one('thead tr') or table.select_one('tr')
        if header_row:
            headers = [
                th.get_text(strip=True)
                for th in header_row.select('th, td')
            ]

        # Extract rows
        rows = []
        for tr in table.select('tbody tr') or table.select('tr')[1:]:
            cells = [td.get_text(strip=True) for td in tr.select('td')]
            if cells:
                if headers:
                    row = dict(zip(headers, cells))
                else:
                    row = cells
                rows.append(row)

        results.append({
            'headers': headers,
            'rows': rows
        })

    return results

def save_table_to_csv(table_data, filename):
    """Save extracted table data to CSV."""
    if not table_data or not table_data['rows']:
        return

    with open(filename, 'w', newline='', encoding='utf-8') as f:
        writer = csv.DictWriter(f, fieldnames=table_data['headers'])
        writer.writeheader()
        writer.writerows(table_data['rows'])

# Usage
html = """
<table>
    <thead>
        <tr><th>Name</th><th>Email</th><th>Role</th></tr>
    </thead>
    <tbody>
        <tr><td>Alice</td><td>alice@example.com</td><td>Admin</td></tr>
        <tr><td>Bob</td><td>bob@example.com</td><td>User</td></tr>
    </tbody>
</table>
"""

tables = extract_table_data(html)
for i, table in enumerate(tables):
    save_table_to_csv(table, f'table_{i}.csv')
    print(f"Table {i}: {len(table['rows'])} rows")

Handling Dynamic Content Indicators

from bs4 import BeautifulSoup
import re

def extract_json_ld(soup):
    """Extract JSON-LD structured data from page."""
    import json

    scripts = soup.find_all('script', type='application/ld+json')
    data = []

    for script in scripts:
        try:
            json_data = json.loads(script.string)
            data.append(json_data)
        except (json.JSONDecodeError, TypeError):
            continue

    return data

def extract_meta_tags(soup):
    """Extract Open Graph and Twitter meta tags."""
    meta_data = {
        'og': {},
        'twitter': {},
        'standard': {}
    }

    for meta in soup.find_all('meta'):
        # Open Graph
        if meta.get('property', '').startswith('og:'):
            key = meta['property'][3:]
            meta_data['og'][key] = meta.get('content', '')

        # Twitter Cards
        elif meta.get('name', '').startswith('twitter:'):
            key = meta['name'][8:]
            meta_data['twitter'][key] = meta.get('content', '')

        # Standard meta tags
        elif meta.get('name'):
            meta_data['standard'][meta['name']] = meta.get('content', '')

    return meta_data

def extract_data_attributes(soup, selector):
    """Extract all data-* attributes from elements."""
    elements = soup.select(selector)
    results = []

    for elem in elements:
        data_attrs = {
            key[5:]: value  # Remove 'data-' prefix
            for key, value in elem.attrs.items()
            if key.startswith('data-')
        }
        if data_attrs:
            results.append(data_attrs)

    return results

# Usage
html = """
<html>
<head>
    <meta property="og:title" content="Page Title" />
    <meta property="og:description" content="Page description" />
    <meta name="twitter:card" content="summary" />
    <script type="application/ld+json">
    {"@context": "https://schema.org", "@type": "Article", "name": "Test"}
    </script>
</head>
<body>
    <div class="item" data-id="1" data-category="books">Item 1</div>
    <div class="item" data-id="2" data-category="electronics">Item 2</div>
</body>
</html>
"""

soup = BeautifulSoup(html, 'lxml')

# Extract all metadata types
json_ld = extract_json_ld(soup)
meta = extract_meta_tags(soup)
data_attrs = extract_data_attributes(soup, '.item')

print(f"JSON-LD: {json_ld}")
print(f"OG tags: {meta['og']}")
print(f"Data attributes: {data_attrs}")

Robust Scraper with Retry Logic

import requests
from bs4 import BeautifulSoup
import time
import logging
from functools import wraps

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

def retry(max_attempts=3, delay=1, backoff=2):
    """Decorator for retrying failed requests."""
    def decorator(func):
        @wraps(func)
        def wrapper(*args, **kwargs):
            attempts = 0
            current_delay = delay

            while attempts < max_attempts:
                try:
                    return func(*args, **kwargs)
                except Exception as e:
                    attempts += 1
                    if attempts == max_attempts:
                        logger.error(f"Failed after {max_attempts} attempts: {e}")
                        raise
                    logger.warning(f"Attempt {attempts} failed: {e}. Retrying in {current_delay}s")
                    time.sleep(current_delay)
                    current_delay *= backoff

        return wrapper
    return decorator

class RobustScraper:
    def __init__(self, base_url):
        self.base_url = base_url
        self.session = requests.Session()
        self.session.headers.update({
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36'
        })

    @retry(max_attempts=3, delay=1)
    def fetch(self, url):
        """Fetch URL with automatic retry."""
        response = self.session.get(url, timeout=10)
        response.raise_for_status()
        return BeautifulSoup(response.content, 'lxml')

    def safe_select(self, soup, selector, attribute=None):
        """Safely select element and extract data."""
        element = soup.select_one(selector)
        if not element:
            return None

        if attribute:
            return element.get(attribute)
        return element.get_text(strip=True)

    def safe_select_all(self, soup, selector, attribute=None):
        """Safely select all elements and extract data."""
        elements = soup.select(selector)
        if attribute:
            return [elem.get(attribute) for elem in elements if elem.get(attribute)]
        return [elem.get_text(strip=True) for elem in elements]

    def scrape_with_validation(self, url, schema):
        """Scrape page and validate extracted data against schema."""
        soup = self.fetch(url)
        data = {}

        for field, config in schema.items():
            value = self.safe_select(
                soup,
                config['selector'],
                config.get('attribute')
            )

            # Validate required fields
            if config.get('required') and not value:
                logger.warning(f"Missing required field: {field}")

            # Apply transform if specified
            if value and 'transform' in config:
                value = config['transform'](value)

            data[field] = value

        return data

# Usage
scraper = RobustScraper('https://example.com')

schema = {
    'title': {
        'selector': 'h1',
        'required': True
    },
    'price': {
        'selector': '.price',
        'transform': lambda x: float(x.replace('$', '').replace(',', ''))
    },
    'image': {
        'selector': 'img.product-image',
        'attribute': 'src'
    }
}

try:
    data = scraper.scrape_with_validation('https://example.com/product/1', schema)
    print(data)
except Exception as e:
    logger.error(f"Scraping failed: {e}")

Quick Reference

Operation Code Example
Parse HTML BeautifulSoup(html, 'lxml')
Parse XML BeautifulSoup(xml, 'lxml-xml')
Find first soup.find('div', class_='name')
Find all soup.find_all('a', href=True)
CSS selector soup.select('div.class > p')
CSS selector first soup.select_one('#id')
Get text element.get_text(strip=True)
Get attribute element.get('href', '')
Get all attributes element.attrs
Check attribute element.has_attr('class')
Parent element.parent
Children list(element.children)
Descendants element.descendants
Next sibling element.find_next_sibling()
Previous sibling element.find_previous_sibling()
Create tag soup.new_tag('div', id='new')
Append parent.append(new_element)
Insert parent.insert(0, new_element)
Replace element.replace_with(new_element)
Remove element.decompose()
Extract element.extract()
Clear children element.clear()
Wrap element.wrap(wrapper)
Unwrap element.unwrap()
Pretty print soup.prettify()
Find with regex soup.find_all(href=re.compile(r'pattern'))
Find with function soup.find_all(lambda tag: condition)
Limit results soup.find_all('div', limit=5)

Common Issues and Solutions

Issue: Parser Differences Causing Inconsistent Results

Problem: Same HTML produces different parse trees with different parsers

Solution:

from bs4 import BeautifulSoup

html = '<p>Test<p>Another'  # Malformed HTML

# Different parsers handle this differently
# Use lxml for most cases - it's fast and lenient
soup = BeautifulSoup(html, 'lxml')

# For maximum browser compatibility, use html5lib
soup = BeautifulSoup(html, 'html5lib')

# Stick to one parser throughout your project for consistency

Issue: AttributeError When Element Not Found

Problem: NoneType has no attribute when element doesn't exist

Solution:

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, 'lxml')

# Wrong - raises AttributeError if not found
# text = soup.find('div', class_='missing').text

# Right - check for None first
element = soup.find('div', class_='missing')
text = element.text if element else ''

# Or use select_one with conditional
text = (soup.select_one('.missing') or soup.new_tag('div')).get_text()

# Or with walrus operator (Python 3.8+)
if (elem := soup.find('div', class_='target')):
    print(elem.text)

Issue: Whitespace in Sibling Navigation

Problem: next_sibling returns whitespace NavigableString

Solution:

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, 'lxml')
element = soup.find('p')

# Wrong - may return whitespace
next_elem = element.next_sibling

# Right - use find_next_sibling for tags
next_p = element.find_next_sibling('p')
next_any = element.find_next_sibling()  # Any tag, not text

# Or filter manually
next_elem = element.next_sibling
while next_elem and isinstance(next_elem, str):
    next_elem = next_elem.next_sibling

Issue: Class Attribute Returns List

Problem: element['class'] returns list, not string

Solution:

from bs4 import BeautifulSoup

soup = BeautifulSoup('<p class="intro main">Text</p>', 'lxml')
p = soup.find('p')

# class_ is always a list
classes = p.get('class', [])  # ['intro', 'main']

# Join if you need a string
class_string = ' '.join(p.get('class', []))  # 'intro main'

# Check if class exists
if 'intro' in p.get('class', []):
    print('Has intro class')

Issue: Encoding Problems with Special Characters

Problem: Garbled text or encoding errors

Solution:

from bs4 import BeautifulSoup

# Let BeautifulSoup detect encoding from bytes
response = requests.get(url)
soup = BeautifulSoup(response.content, 'lxml')  # Use .content not .text

# Check detected encoding
print(soup.original_encoding)

# Force specific encoding if detection fails
soup = BeautifulSoup(html_bytes, 'lxml', from_encoding='utf-8')

# Output with specific encoding
output = soup.encode('utf-8')

Issue: Memory Issues with Large Documents

Problem: High memory usage when parsing large HTML files

Solution:

from bs4 import BeautifulSoup, SoupStrainer

# Parse only specific parts using SoupStrainer
only_articles = SoupStrainer('article')
soup = BeautifulSoup(html, 'lxml', parse_only=only_articles)

# Or parse only elements with specific class
only_products = SoupStrainer(class_='product')
soup = BeautifulSoup(html, 'lxml', parse_only=only_products)

# Clean up when done
soup.decompose()
del soup

Issue: find_all Not Finding Elements with Multiple Classes

Problem: Can't find elements with exact multiple classes

Solution:

from bs4 import BeautifulSoup

html = '<div class="btn btn-primary large">Button</div>'
soup = BeautifulSoup(html, 'lxml')

# This finds element with 'btn' as one of its classes
soup.find_all(class_='btn')  # Works

# For multiple classes, use CSS selector
soup.select('.btn.btn-primary')  # Works

# Or check all classes
soup.find_all(class_=['btn', 'btn-primary'])  # Matches any

# For exact class match, use custom function
def exact_classes(tag):
    return tag.get('class') == ['btn', 'btn-primary', 'large']

soup.find_all(exact_classes)

Issue: Modifying Tree During Iteration

Problem: Unexpected behaviour when modifying while iterating

Solution:

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, 'lxml')

# Wrong - modifying during iteration
for tag in soup.find_all('script'):
    tag.decompose()  # May skip elements

# Right - convert to list first
for tag in list(soup.find_all('script')):
    tag.decompose()

# Or use copy
from copy import copy
for tag in soup.find_all('script')[:]:  # Slice creates copy
    tag.decompose()

Issue: Relative URLs in Scraped Content

Problem: Extracted URLs are relative, not absolute

Solution:

from bs4 import BeautifulSoup
from urllib.parse import urljoin

base_url = 'https://example.com/page/'
soup = BeautifulSoup(html, 'lxml')

# Convert all relative URLs to absolute
for link in soup.find_all('a', href=True):
    link['href'] = urljoin(base_url, link['href'])

for img in soup.find_all('img', src=True):
    img['src'] = urljoin(base_url, img['src'])

# When extracting, always convert
urls = [
    urljoin(base_url, a['href'])
    for a in soup.find_all('a', href=True)
]

Related Topics

  • Python Requests - HTTP library for fetching web pages
  • Scrapy - Full-featured web scraping framework
  • Selenium - Browser automation for JavaScript-rendered content
  • lxml - Fast XML and HTML processing library
  • CSS Selectors - Understanding selector syntax for web scraping
  • Regular Expressions - Pattern matching for text extraction
  • XPath - Alternative query language for XML/HTML navigation