Python Beautiful Soup
A powerful HTML and XML parsing library for extracting data from web pages and structured documents.
Python Beautiful Soup Cheatsheet
A powerful HTML and XML parsing library for extracting data from web pages and structured documents.
Overview
Beautiful Soup provides Pythonic idioms for navigating, searching, and modifying parse trees, making it the go-to library for web scraping and HTML/XML processing in Python.
flowchart TD
subgraph Input["Input Sources"]
A[HTML String]
B[XML String]
C[File Object]
D[URL Response]
end
subgraph Parser["BeautifulSoup Parser"]
E[html.parser]
F[lxml]
G[html5lib]
H[lxml-xml]
end
subgraph Tree["Parse Tree"]
I[Tag Objects]
J[NavigableString]
K[Comment]
end
subgraph Operations["Operations"]
L[Search/Find]
M[Navigate]
N[Extract]
O[Modify]
end
A & B & C & D --> E & F & G & H
E & F & G & H --> I & J & K
I & J & K --> L & M & N & O
Installation
uv add beautifulsoup4
# Install parsers (optional but recommended)
uv add lxml # Fast, lenient HTML/XML parser
uv add html5lib # Extremely lenient, parses like a browser
Parsing HTML/XML
Key Concepts
- BeautifulSoup object represents the entire parsed document
- Parsers differ in speed, leniency, and dependencies
- Document encoding is automatically detected or can be specified
- Parse tree is built from tags, navigable strings, and comments
Common Patterns
from bs4 import BeautifulSoup
# Parse HTML string
soup = BeautifulSoup(html_string, 'html.parser')
# Parse with lxml (faster)
soup = BeautifulSoup(html_string, 'lxml')
# Parse XML
soup = BeautifulSoup(xml_string, 'lxml-xml') # or 'xml'
# Parse from file
with open('page.html', 'r', encoding='utf-8') as f:
soup = BeautifulSoup(f, 'html.parser')
# Parse with specific encoding
soup = BeautifulSoup(html_bytes, 'html.parser', from_encoding='iso-8859-1')
Examples
Choosing the Right Parser
from bs4 import BeautifulSoup
html = """
<html>
<head><title>Test Page</title></head>
<body>
<p>Hello <b>World</p> <!-- Malformed HTML -->
</body>
</html>
"""
# html.parser - Built-in, decent speed, lenient
soup = BeautifulSoup(html, 'html.parser')
print(soup.prettify())
# lxml - Fast, lenient (recommended for most cases)
soup = BeautifulSoup(html, 'lxml')
print(soup.prettify())
# html5lib - Slowest but parses exactly like browsers
soup = BeautifulSoup(html, 'html5lib')
print(soup.prettify())
# lxml-xml - For strict XML parsing
xml = '<root><item id="1">Value</item></root>'
soup = BeautifulSoup(xml, 'lxml-xml')
print(soup.prettify())
Parsing from Web Requests
import requests
from bs4 import BeautifulSoup
# Fetch and parse a webpage
response = requests.get('https://example.com')
response.raise_for_status()
# Parse the HTML content
soup = BeautifulSoup(response.content, 'lxml')
# Use response.content (bytes) instead of response.text
# to let BeautifulSoup handle encoding detection
# Access document metadata
print(f"Title: {soup.title.string}")
print(f"Encoding: {soup.original_encoding}")
Handling Encoding Issues
from bs4 import BeautifulSoup
# Let BeautifulSoup detect encoding
soup = BeautifulSoup(html_bytes, 'html.parser')
print(f"Detected encoding: {soup.original_encoding}")
# Force specific encoding
soup = BeautifulSoup(
html_bytes,
'html.parser',
from_encoding='utf-8'
)
# Exclude wrong encodings from detection
soup = BeautifulSoup(
html_bytes,
'html.parser',
exclude_encodings=['iso-8859-7']
)
Finding Elements
Key Concepts
- find() returns the first matching element or None
- find_all() returns a list of all matching elements
- select() uses CSS selectors for powerful, flexible searches
- Filters can be strings, regex, lists, functions, or True
flowchart LR
A[Search Method] --> B[find]
A --> C[find_all]
A --> D[select]
A --> E[select_one]
B --> F[First Match]
C --> G[All Matches]
D --> H[CSS Selector All]
E --> I[CSS Selector First]
subgraph Filters
J[String]
K[Regex]
L[List]
M[Function]
N[True]
end
B & C --> Filters
Common Patterns
from bs4 import BeautifulSoup
import re
# Find by tag name
soup.find('div')
soup.find_all('p')
# Find by attributes
soup.find('div', id='main')
soup.find_all('a', class_='link')
soup.find_all('input', attrs={'type': 'text'})
# Find by CSS selector
soup.select('div.container')
soup.select_one('#header')
soup.select('ul > li')
# Find with regex
soup.find_all('a', href=re.compile(r'^https://'))
# Find with function
soup.find_all(lambda tag: tag.has_attr('data-value'))
Examples
Basic Finding
from bs4 import BeautifulSoup
html = """
<html>
<body>
<div id="content" class="main">
<h1>Title</h1>
<p class="intro">First paragraph</p>
<p class="body">Second paragraph</p>
<a href="/page1" class="link">Link 1</a>
<a href="/page2" class="link external">Link 2</a>
</div>
</body>
</html>
"""
soup = BeautifulSoup(html, 'lxml')
# Find first element
first_p = soup.find('p')
print(first_p.text) # "First paragraph"
# Find all elements
all_links = soup.find_all('a')
for link in all_links:
print(link['href'])
# Find by id
content = soup.find(id='content')
print(content.h1.text) # "Title"
# Find by class (note: use class_ to avoid Python keyword conflict)
intro = soup.find('p', class_='intro')
print(intro.text) # "First paragraph"
# Find multiple classes
external_link = soup.find('a', class_='link external')
print(external_link.text) # "Link 2"
Advanced Filtering
from bs4 import BeautifulSoup
import re
html = """
<div>
<a href="https://example.com">HTTPS Link</a>
<a href="http://example.org">HTTP Link</a>
<a href="/relative/path">Relative Link</a>
<img src="image1.png" alt="Image 1" />
<img src="image2.jpg" alt="Image 2" />
<p data-value="100">Paragraph</p>
</div>
"""
soup = BeautifulSoup(html, 'lxml')
# Find with regex pattern
https_links = soup.find_all('a', href=re.compile(r'^https://'))
print(f"HTTPS links: {len(https_links)}")
# Find with list of values
images = soup.find_all('img', src=re.compile(r'\.(png|jpg)$'))
print(f"Images: {len(images)}")
# Find with custom function
def has_data_attribute(tag):
return tag.has_attr('data-value')
data_elements = soup.find_all(has_data_attribute)
print(f"Elements with data-value: {len(data_elements)}")
# Find any tag with specific attribute
any_with_alt = soup.find_all(alt=True)
print(f"Elements with alt: {len(any_with_alt)}")
# Find tags by multiple criteria
def complex_filter(tag):
return (tag.name == 'a' and
tag.has_attr('href') and
'example' in tag.get('href', ''))
example_links = soup.find_all(complex_filter)
print(f"Example links: {len(example_links)}")
CSS Selectors
from bs4 import BeautifulSoup
html = """
<div id="wrapper">
<nav class="menu">
<ul>
<li class="active"><a href="/">Home</a></li>
<li><a href="/about">About</a></li>
<li><a href="/contact">Contact</a></li>
</ul>
</nav>
<main>
<article class="post featured">
<h2>Featured Post</h2>
<p>Content here</p>
</article>
<article class="post">
<h2>Regular Post</h2>
<p>More content</p>
</article>
</main>
</div>
"""
soup = BeautifulSoup(html, 'lxml')
# Select by ID
wrapper = soup.select_one('#wrapper')
# Select by class
posts = soup.select('.post')
print(f"Posts: {len(posts)}")
# Select by multiple classes
featured = soup.select('.post.featured')
print(f"Featured posts: {len(featured)}")
# Descendant selector
nav_links = soup.select('nav a')
print(f"Nav links: {len(nav_links)}")
# Direct child selector
list_items = soup.select('ul > li')
print(f"List items: {len(list_items)}")
# Attribute selectors
home_link = soup.select_one('a[href="/"]')
print(f"Home link: {home_link.text}")
# Attribute contains
about_links = soup.select('a[href*="about"]')
# Attribute starts with
internal_links = soup.select('a[href^="/"]')
# Nth-child
second_item = soup.select_one('li:nth-child(2)')
print(f"Second item: {second_item.text.strip()}")
# Combining selectors
active_link = soup.select_one('li.active > a')
print(f"Active link: {active_link.text}")
Limiting Results
from bs4 import BeautifulSoup
html = """
<ul>
<li>Item 1</li>
<li>Item 2</li>
<li>Item 3</li>
<li>Item 4</li>
<li>Item 5</li>
</ul>
"""
soup = BeautifulSoup(html, 'lxml')
# Limit number of results
first_three = soup.find_all('li', limit=3)
print(f"First three items: {len(first_three)}")
# Find starting from specific element
ul = soup.find('ul')
items = ul.find_all('li', limit=2)
# Recursive search (default is True)
# Set to False to only search direct children
direct_children = soup.find_all('li', recursive=False) # Returns []
direct_children = ul.find_all('li', recursive=False) # Returns all li
Navigating the Tree
Key Concepts
- Parent is the containing element
- Children are direct descendants only
- Descendants include all nested elements
- Siblings are elements at the same level
- Navigation returns None when element doesn't exist
flowchart TD
A[Parent] --> B[Previous Sibling]
A --> C[Current Element]
A --> D[Next Sibling]
C --> E[Child 1]
C --> F[Child 2]
E --> G[Grandchild]
style C fill:#f9f,stroke:#333,stroke-width:2px
Common Patterns
from bs4 import BeautifulSoup
# Parent navigation
element.parent # Direct parent
element.parents # Generator of all parents
# Child navigation
element.children # Direct children generator
element.descendants # All descendants generator
element.contents # Direct children as list
# Sibling navigation
element.next_sibling # Next sibling (may be whitespace)
element.previous_sibling # Previous sibling
element.next_siblings # Generator of following siblings
element.previous_siblings # Generator of preceding siblings
# Element navigation (skips whitespace)
element.next_element # Next element in parse order
element.previous_element # Previous element
element.find_next() # Next matching element
element.find_previous() # Previous matching element
element.find_next_sibling() # Next matching sibling
Examples
Parent Navigation
from bs4 import BeautifulSoup
html = """
<html>
<body>
<div id="outer">
<div id="inner">
<p id="target">Target paragraph</p>
</div>
</div>
</body>
</html>
"""
soup = BeautifulSoup(html, 'lxml')
target = soup.find(id='target')
# Get direct parent
parent = target.parent
print(f"Direct parent: {parent.name}, id={parent.get('id')}")
# Output: Direct parent: div, id=inner
# Iterate through all parents
for parent in target.parents:
if parent.name:
print(f"Parent: {parent.name}")
# Output: Parent: div, Parent: div, Parent: body, Parent: html, Parent: [document]
# Find specific parent
outer_div = target.find_parent('div', id='outer')
print(f"Found outer: {outer_div.get('id')}")
# Find all parent divs
parent_divs = target.find_parents('div')
print(f"Parent divs: {len(parent_divs)}") # 2
Child Navigation
from bs4 import BeautifulSoup
html = """
<ul id="menu">
<li>First</li>
<li>Second
<ul>
<li>Nested</li>
</ul>
</li>
<li>Third</li>
</ul>
"""
soup = BeautifulSoup(html, 'lxml')
menu = soup.find(id='menu')
# Get direct children as list
# Note: includes NavigableString (whitespace/text)
contents = menu.contents
print(f"Contents count: {len(contents)}")
# Iterate direct children (excludes NavigableString whitespace)
for child in menu.children:
if child.name: # Skip text nodes
print(f"Child: {child.name}")
# Get all descendants (deep traversal)
for descendant in menu.descendants:
if hasattr(descendant, 'name') and descendant.name:
print(f"Descendant: {descendant.name}")
# Find all descendants with specific tag
all_li = menu.find_all('li') # Includes nested li
print(f"All li elements: {len(all_li)}") # 4
# Find only direct children
direct_li = menu.find_all('li', recursive=False)
print(f"Direct li children: {len(direct_li)}") # 3
Sibling Navigation
from bs4 import BeautifulSoup
html = """
<div>
<h1>Title</h1>
<p id="first">First paragraph</p>
<p id="second">Second paragraph</p>
<p id="third">Third paragraph</p>
<footer>Footer</footer>
</div>
"""
soup = BeautifulSoup(html, 'lxml')
second = soup.find(id='second')
# Next sibling (may be whitespace NavigableString)
next_sib = second.next_sibling
while next_sib and not hasattr(next_sib, 'name'):
next_sib = next_sib.next_sibling
print(f"Next sibling: {next_sib.get('id')}") # third
# Previous sibling
prev_sib = second.previous_sibling
while prev_sib and not hasattr(prev_sib, 'name'):
prev_sib = prev_sib.previous_sibling
print(f"Previous sibling: {prev_sib.get('id')}") # first
# Better: use find_next_sibling/find_previous_sibling
next_p = second.find_next_sibling('p')
print(f"Next p: {next_p.get('id')}") # third
prev_p = second.find_previous_sibling('p')
print(f"Previous p: {prev_p.get('id')}") # first
# Iterate all following siblings
for sibling in second.next_siblings:
if hasattr(sibling, 'name') and sibling.name:
print(f"Following: {sibling.name}")
# Output: Following: p, Following: footer
# Find next/previous element regardless of hierarchy
next_elem = second.find_next('footer')
print(f"Next footer: {next_elem.text}")
Tree Traversal Example
from bs4 import BeautifulSoup
html = """
<table>
<tr>
<th>Name</th>
<th>Age</th>
</tr>
<tr>
<td>Alice</td>
<td>30</td>
</tr>
<tr>
<td>Bob</td>
<td>25</td>
</tr>
</table>
"""
soup = BeautifulSoup(html, 'lxml')
table = soup.find('table')
# Get all rows
rows = table.find_all('tr')
# Process header row
header_row = rows[0]
headers = [th.text for th in header_row.find_all('th')]
print(f"Headers: {headers}")
# Process data rows
data = []
for row in rows[1:]:
cells = row.find_all('td')
row_data = {
headers[i]: cells[i].text
for i in range(len(cells))
}
data.append(row_data)
print(f"Data: {data}")
# Output: [{'Name': 'Alice', 'Age': '30'}, {'Name': 'Bob', 'Age': '25'}]
Extracting Data
Key Concepts
- text/get_text() extracts visible text content
- string gets the single child NavigableString
- strings/stripped_strings generators for text content
- attrs dictionary of all attributes
- get() safely retrieves attribute values
Common Patterns
from bs4 import BeautifulSoup
# Get text content
element.text # All text, whitespace preserved
element.get_text() # Same as .text
element.get_text(strip=True) # Strip whitespace
element.get_text(separator=' ') # Join text with separator
element.string # Single child string only
# Get attributes
element['href'] # Direct access (raises KeyError)
element.get('href') # Safe access (returns None)
element.get('href', '#') # With default value
element.attrs # All attributes as dict
element.has_attr('class') # Check attribute exists
# Get name
element.name # Tag name as string
Examples
Extracting Text
from bs4 import BeautifulSoup
html = """
<div id="content">
<h1>Welcome</h1>
<p>This is a <strong>test</strong> paragraph.</p>
<ul>
<li>Item 1</li>
<li>Item 2</li>
</ul>
</div>
"""
soup = BeautifulSoup(html, 'lxml')
content = soup.find(id='content')
# Get all text (preserves whitespace)
all_text = content.text
print(repr(all_text))
# Get text with stripped whitespace
clean_text = content.get_text(strip=True)
print(clean_text)
# Get text with custom separator
spaced_text = content.get_text(separator=' ', strip=True)
print(spaced_text)
# Output: "Welcome This is a test paragraph. Item 1 Item 2"
# Iterate over strings
for string in content.strings:
print(repr(string))
# Get stripped strings (no whitespace-only strings)
text_parts = list(content.stripped_strings)
print(text_parts)
# Output: ['Welcome', 'This is a', 'test', 'paragraph.', 'Item 1', 'Item 2']
# Get string from element with single text child
h1 = soup.find('h1')
print(h1.string) # "Welcome"
# string returns None if multiple children
p = soup.find('p')
print(p.string) # None (has multiple children)
print(p.text) # "This is a test paragraph."
Extracting Attributes
from bs4 import BeautifulSoup
html = """
<div>
<a href="https://example.com"
class="link external"
id="main-link"
data-value="123"
title="Example Link">Visit Example</a>
<img src="image.png" alt="Test Image" />
</div>
"""
soup = BeautifulSoup(html, 'lxml')
link = soup.find('a')
# Direct attribute access
href = link['href']
print(f"URL: {href}")
# Safe access with get()
data_value = link.get('data-value')
print(f"Data value: {data_value}")
# With default
target = link.get('target', '_self')
print(f"Target: {target}")
# Get all attributes
print(f"All attributes: {link.attrs}")
# {'href': 'https://example.com', 'class': ['link', 'external'],
# 'id': 'main-link', 'data-value': '123', 'title': 'Example Link'}
# Note: class is always a list
classes = link.get('class', [])
print(f"Classes: {classes}") # ['link', 'external']
# Check if attribute exists
if link.has_attr('title'):
print(f"Title: {link['title']}")
# Get image attributes
img = soup.find('img')
print(f"Image src: {img.get('src')}")
print(f"Image alt: {img.get('alt')}")
Extracting Links and Images
from bs4 import BeautifulSoup
from urllib.parse import urljoin
html = """
<html>
<head>
<base href="https://example.com/">
</head>
<body>
<nav>
<a href="/home">Home</a>
<a href="/about">About</a>
<a href="https://external.com">External</a>
</nav>
<article>
<img src="images/photo.jpg" alt="Photo" />
<p>Read more at <a href="/articles/1">Article 1</a></p>
</article>
</body>
</html>
"""
soup = BeautifulSoup(html, 'lxml')
base_url = 'https://example.com'
# Extract all links with absolute URLs
links = []
for a in soup.find_all('a', href=True):
url = a['href']
absolute_url = urljoin(base_url, url)
links.append({
'text': a.get_text(strip=True),
'url': absolute_url,
'is_external': not absolute_url.startswith(base_url)
})
for link in links:
print(f"{link['text']}: {link['url']} (external: {link['is_external']})")
# Extract all images
images = []
for img in soup.find_all('img'):
images.append({
'src': urljoin(base_url, img.get('src', '')),
'alt': img.get('alt', '')
})
print(f"\nImages: {images}")
Extracting Structured Data
from bs4 import BeautifulSoup
html = """
<div class="product-list">
<div class="product" data-id="1">
<h3 class="name">Laptop</h3>
<span class="price">$999.99</span>
<p class="description">High-performance laptop</p>
<span class="stock in-stock">In Stock</span>
</div>
<div class="product" data-id="2">
<h3 class="name">Mouse</h3>
<span class="price">$29.99</span>
<p class="description">Wireless mouse</p>
<span class="stock out-of-stock">Out of Stock</span>
</div>
</div>
"""
soup = BeautifulSoup(html, 'lxml')
products = []
for product_div in soup.select('.product'):
product = {
'id': product_div.get('data-id'),
'name': product_div.select_one('.name').get_text(strip=True),
'price': product_div.select_one('.price').get_text(strip=True),
'description': product_div.select_one('.description').get_text(strip=True),
'in_stock': 'in-stock' in product_div.select_one('.stock').get('class', [])
}
products.append(product)
for product in products:
print(f"{product['name']}: {product['price']} - In stock: {product['in_stock']}")
Modifying the Tree
Key Concepts
- append() adds content to the end of an element
- insert() adds content at a specific position
- replace_with() replaces an element with new content
- decompose() removes element from tree and destroys it
- extract() removes element but keeps it usable
- wrap() wraps element in another tag
- unwrap() replaces tag with its contents
Common Patterns
from bs4 import BeautifulSoup
# Adding content
tag.append(content) # Add to end
tag.insert(position, content) # Add at position
tag.insert_before(content) # Add before element
tag.insert_after(content) # Add after element
# Removing content
tag.decompose() # Remove and destroy
tag.extract() # Remove but keep
tag.clear() # Remove all children
# Replacing content
tag.replace_with(new_content) # Replace element
tag.string.replace_with(text) # Replace text
# Wrapping/unwrapping
tag.wrap(wrapper_tag) # Wrap in another tag
tag.unwrap() # Remove tag, keep contents
# Creating new elements
soup.new_tag('div', id='new') # Create new tag
NavigableString('text') # Create text node
Examples
Adding Content
from bs4 import BeautifulSoup
from bs4 import NavigableString
html = """
<div id="container">
<p>First paragraph</p>
</div>
"""
soup = BeautifulSoup(html, 'lxml')
container = soup.find(id='container')
# Create and append new tag
new_p = soup.new_tag('p')
new_p.string = 'Second paragraph'
container.append(new_p)
# Create tag with attributes
new_link = soup.new_tag('a', href='https://example.com', class_='link')
new_link.string = 'Click here'
container.append(new_link)
# Insert at specific position
header = soup.new_tag('h2')
header.string = 'Section Title'
container.insert(0, header) # Insert at beginning
# Append string content
container.append(NavigableString(' Some text '))
# Insert before/after
footer = soup.new_tag('footer')
footer.string = 'Footer content'
new_p.insert_after(footer)
print(container.prettify())
Removing Content
from bs4 import BeautifulSoup
html = """
<div id="content">
<p class="keep">Keep this</p>
<p class="remove">Remove this</p>
<script>alert('Remove scripts')</script>
<style>.remove { color: red; }</style>
<p class="keep">Keep this too</p>
<!-- Remove comments -->
</div>
"""
soup = BeautifulSoup(html, 'lxml')
# Remove specific elements
for script in soup.find_all('script'):
script.decompose()
for style in soup.find_all('style'):
style.decompose()
# Remove elements by class
for elem in soup.find_all(class_='remove'):
elem.decompose()
# Remove comments
from bs4 import Comment
for comment in soup.find_all(string=lambda text: isinstance(text, Comment)):
comment.extract()
# Extract (remove but keep for later use)
content = soup.find(id='content')
kept_paragraphs = []
for p in content.find_all('p', class_='keep'):
kept_paragraphs.append(p.extract())
print("Extracted paragraphs:")
for p in kept_paragraphs:
print(f" - {p.text}")
# Clear all children
# content.clear() # Would remove all children
print("\nRemaining content:")
print(soup.prettify())
Replacing Content
from bs4 import BeautifulSoup
html = """
<div id="article">
<h1>Old Title</h1>
<p>Paragraph with <em>emphasis</em> and <strong>bold</strong> text.</p>
<a href="http://old-url.com">Old Link</a>
</div>
"""
soup = BeautifulSoup(html, 'lxml')
# Replace entire element
old_h1 = soup.find('h1')
new_h1 = soup.new_tag('h1')
new_h1.string = 'New Title'
old_h1.replace_with(new_h1)
# Replace with multiple elements
em = soup.find('em')
new_content = soup.new_tag('strong')
new_content.string = em.string.upper()
em.replace_with(new_content)
# Replace text content
for strong in soup.find_all('strong'):
if strong.string:
strong.string.replace_with(strong.string.upper())
# Replace attribute
link = soup.find('a')
link['href'] = 'https://new-url.com'
link.string = 'New Link'
print(soup.prettify())
Wrapping and Unwrapping
from bs4 import BeautifulSoup
html = """
<div>
<p>Unwrapped paragraph</p>
<span><strong>Bold text</strong></span>
</div>
"""
soup = BeautifulSoup(html, 'lxml')
# Wrap element in new tag
p = soup.find('p')
wrapper = soup.new_tag('section', class_='content')
p.wrap(wrapper)
# Wrap multiple elements
# First, create a wrapper
div = soup.find('div')
container = soup.new_tag('article')
# Move children to new container
children = list(div.children)
for child in children:
if child.name:
child.wrap(container.new_tag('div'))
# Unwrap (remove tag but keep contents)
span = soup.find('span')
span.unwrap() # Removes <span>, keeps <strong>Bold text</strong>
print(soup.prettify())
Building Documents from Scratch
from bs4 import BeautifulSoup
# Start with minimal HTML
soup = BeautifulSoup('<html><body></body></html>', 'lxml')
# Add head section
head = soup.new_tag('head')
soup.html.insert(0, head)
# Add title
title = soup.new_tag('title')
title.string = 'Generated Page'
head.append(title)
# Add meta tags
meta = soup.new_tag('meta', charset='utf-8')
head.insert(0, meta)
# Build body content
body = soup.body
# Add header
header = soup.new_tag('header')
h1 = soup.new_tag('h1')
h1.string = 'Welcome'
header.append(h1)
body.append(header)
# Add main content
main = soup.new_tag('main')
for i in range(3):
p = soup.new_tag('p')
p.string = f'Paragraph {i + 1}'
main.append(p)
body.append(main)
# Add footer
footer = soup.new_tag('footer')
footer.string = 'Copyright 2025'
body.append(footer)
# Output formatted HTML
print(soup.prettify())
Common Patterns for Web Scraping
Key Concepts
- Respect robots.txt and terms of service
- Rate limiting prevents server overload
- User-Agent headers identify your scraper
- Error handling manages network and parsing issues
- Data validation ensures extracted data quality
Examples
Basic Web Scraper
import requests
from bs4 import BeautifulSoup
import time
class SimpleScraper:
def __init__(self, base_url, delay=1):
self.base_url = base_url
self.delay = delay
self.session = requests.Session()
self.session.headers.update({
'User-Agent': 'Mozilla/5.0 (compatible; MyScraper/1.0)'
})
def fetch_page(self, url):
"""Fetch a page with rate limiting and error handling."""
try:
response = self.session.get(url, timeout=10)
response.raise_for_status()
time.sleep(self.delay) # Rate limiting
return BeautifulSoup(response.content, 'lxml')
except requests.RequestException as e:
print(f"Error fetching {url}: {e}")
return None
def scrape_article(self, url):
"""Extract article data from a page."""
soup = self.fetch_page(url)
if not soup:
return None
article = soup.find('article') or soup.find(class_='post')
if not article:
return None
return {
'title': self._get_text(article, 'h1'),
'content': self._get_text(article, '.content'),
'author': self._get_text(article, '.author'),
'date': self._get_text(article, '.date'),
'url': url
}
def _get_text(self, element, selector):
"""Safely extract text from element."""
found = element.select_one(selector)
return found.get_text(strip=True) if found else ''
# Usage
scraper = SimpleScraper('https://example.com', delay=2)
article = scraper.scrape_article('https://example.com/article/1')
if article:
print(f"Title: {article['title']}")
Pagination Handling
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
def scrape_paginated_list(base_url, list_selector, item_selector, max_pages=10):
"""Scrape items from paginated list pages."""
session = requests.Session()
session.headers.update({
'User-Agent': 'Mozilla/5.0 (compatible; MyScraper/1.0)'
})
all_items = []
current_url = base_url
page = 1
while current_url and page <= max_pages:
print(f"Scraping page {page}: {current_url}")
response = session.get(current_url, timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.content, 'lxml')
# Extract items from current page
container = soup.select_one(list_selector)
if not container:
break
items = container.select(item_selector)
for item in items:
item_data = {
'title': item.select_one('h3').get_text(strip=True) if item.select_one('h3') else '',
'link': urljoin(base_url, item.select_one('a')['href']) if item.select_one('a') else ''
}
all_items.append(item_data)
# Find next page link
next_link = soup.select_one('a.next, a[rel="next"], .pagination a:last-child')
if next_link and 'href' in next_link.attrs:
current_url = urljoin(base_url, next_link['href'])
page += 1
else:
break
return all_items
# Usage
items = scrape_paginated_list(
'https://example.com/products',
'.product-list',
'.product-item',
max_pages=5
)
print(f"Scraped {len(items)} items")
Table Extraction
from bs4 import BeautifulSoup
import csv
def extract_table_data(html, table_selector='table'):
"""Extract data from HTML tables."""
soup = BeautifulSoup(html, 'lxml')
tables = soup.select(table_selector)
results = []
for table in tables:
# Extract headers
headers = []
header_row = table.select_one('thead tr') or table.select_one('tr')
if header_row:
headers = [
th.get_text(strip=True)
for th in header_row.select('th, td')
]
# Extract rows
rows = []
for tr in table.select('tbody tr') or table.select('tr')[1:]:
cells = [td.get_text(strip=True) for td in tr.select('td')]
if cells:
if headers:
row = dict(zip(headers, cells))
else:
row = cells
rows.append(row)
results.append({
'headers': headers,
'rows': rows
})
return results
def save_table_to_csv(table_data, filename):
"""Save extracted table data to CSV."""
if not table_data or not table_data['rows']:
return
with open(filename, 'w', newline='', encoding='utf-8') as f:
writer = csv.DictWriter(f, fieldnames=table_data['headers'])
writer.writeheader()
writer.writerows(table_data['rows'])
# Usage
html = """
<table>
<thead>
<tr><th>Name</th><th>Email</th><th>Role</th></tr>
</thead>
<tbody>
<tr><td>Alice</td><td>alice@example.com</td><td>Admin</td></tr>
<tr><td>Bob</td><td>bob@example.com</td><td>User</td></tr>
</tbody>
</table>
"""
tables = extract_table_data(html)
for i, table in enumerate(tables):
save_table_to_csv(table, f'table_{i}.csv')
print(f"Table {i}: {len(table['rows'])} rows")
Handling Dynamic Content Indicators
from bs4 import BeautifulSoup
import re
def extract_json_ld(soup):
"""Extract JSON-LD structured data from page."""
import json
scripts = soup.find_all('script', type='application/ld+json')
data = []
for script in scripts:
try:
json_data = json.loads(script.string)
data.append(json_data)
except (json.JSONDecodeError, TypeError):
continue
return data
def extract_meta_tags(soup):
"""Extract Open Graph and Twitter meta tags."""
meta_data = {
'og': {},
'twitter': {},
'standard': {}
}
for meta in soup.find_all('meta'):
# Open Graph
if meta.get('property', '').startswith('og:'):
key = meta['property'][3:]
meta_data['og'][key] = meta.get('content', '')
# Twitter Cards
elif meta.get('name', '').startswith('twitter:'):
key = meta['name'][8:]
meta_data['twitter'][key] = meta.get('content', '')
# Standard meta tags
elif meta.get('name'):
meta_data['standard'][meta['name']] = meta.get('content', '')
return meta_data
def extract_data_attributes(soup, selector):
"""Extract all data-* attributes from elements."""
elements = soup.select(selector)
results = []
for elem in elements:
data_attrs = {
key[5:]: value # Remove 'data-' prefix
for key, value in elem.attrs.items()
if key.startswith('data-')
}
if data_attrs:
results.append(data_attrs)
return results
# Usage
html = """
<html>
<head>
<meta property="og:title" content="Page Title" />
<meta property="og:description" content="Page description" />
<meta name="twitter:card" content="summary" />
<script type="application/ld+json">
{"@context": "https://schema.org", "@type": "Article", "name": "Test"}
</script>
</head>
<body>
<div class="item" data-id="1" data-category="books">Item 1</div>
<div class="item" data-id="2" data-category="electronics">Item 2</div>
</body>
</html>
"""
soup = BeautifulSoup(html, 'lxml')
# Extract all metadata types
json_ld = extract_json_ld(soup)
meta = extract_meta_tags(soup)
data_attrs = extract_data_attributes(soup, '.item')
print(f"JSON-LD: {json_ld}")
print(f"OG tags: {meta['og']}")
print(f"Data attributes: {data_attrs}")
Robust Scraper with Retry Logic
import requests
from bs4 import BeautifulSoup
import time
import logging
from functools import wraps
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)
def retry(max_attempts=3, delay=1, backoff=2):
"""Decorator for retrying failed requests."""
def decorator(func):
@wraps(func)
def wrapper(*args, **kwargs):
attempts = 0
current_delay = delay
while attempts < max_attempts:
try:
return func(*args, **kwargs)
except Exception as e:
attempts += 1
if attempts == max_attempts:
logger.error(f"Failed after {max_attempts} attempts: {e}")
raise
logger.warning(f"Attempt {attempts} failed: {e}. Retrying in {current_delay}s")
time.sleep(current_delay)
current_delay *= backoff
return wrapper
return decorator
class RobustScraper:
def __init__(self, base_url):
self.base_url = base_url
self.session = requests.Session()
self.session.headers.update({
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36'
})
@retry(max_attempts=3, delay=1)
def fetch(self, url):
"""Fetch URL with automatic retry."""
response = self.session.get(url, timeout=10)
response.raise_for_status()
return BeautifulSoup(response.content, 'lxml')
def safe_select(self, soup, selector, attribute=None):
"""Safely select element and extract data."""
element = soup.select_one(selector)
if not element:
return None
if attribute:
return element.get(attribute)
return element.get_text(strip=True)
def safe_select_all(self, soup, selector, attribute=None):
"""Safely select all elements and extract data."""
elements = soup.select(selector)
if attribute:
return [elem.get(attribute) for elem in elements if elem.get(attribute)]
return [elem.get_text(strip=True) for elem in elements]
def scrape_with_validation(self, url, schema):
"""Scrape page and validate extracted data against schema."""
soup = self.fetch(url)
data = {}
for field, config in schema.items():
value = self.safe_select(
soup,
config['selector'],
config.get('attribute')
)
# Validate required fields
if config.get('required') and not value:
logger.warning(f"Missing required field: {field}")
# Apply transform if specified
if value and 'transform' in config:
value = config['transform'](value)
data[field] = value
return data
# Usage
scraper = RobustScraper('https://example.com')
schema = {
'title': {
'selector': 'h1',
'required': True
},
'price': {
'selector': '.price',
'transform': lambda x: float(x.replace('$', '').replace(',', ''))
},
'image': {
'selector': 'img.product-image',
'attribute': 'src'
}
}
try:
data = scraper.scrape_with_validation('https://example.com/product/1', schema)
print(data)
except Exception as e:
logger.error(f"Scraping failed: {e}")
Quick Reference
| Operation | Code Example |
|---|---|
| Parse HTML | BeautifulSoup(html, 'lxml') |
| Parse XML | BeautifulSoup(xml, 'lxml-xml') |
| Find first | soup.find('div', class_='name') |
| Find all | soup.find_all('a', href=True) |
| CSS selector | soup.select('div.class > p') |
| CSS selector first | soup.select_one('#id') |
| Get text | element.get_text(strip=True) |
| Get attribute | element.get('href', '') |
| Get all attributes | element.attrs |
| Check attribute | element.has_attr('class') |
| Parent | element.parent |
| Children | list(element.children) |
| Descendants | element.descendants |
| Next sibling | element.find_next_sibling() |
| Previous sibling | element.find_previous_sibling() |
| Create tag | soup.new_tag('div', id='new') |
| Append | parent.append(new_element) |
| Insert | parent.insert(0, new_element) |
| Replace | element.replace_with(new_element) |
| Remove | element.decompose() |
| Extract | element.extract() |
| Clear children | element.clear() |
| Wrap | element.wrap(wrapper) |
| Unwrap | element.unwrap() |
| Pretty print | soup.prettify() |
| Find with regex | soup.find_all(href=re.compile(r'pattern')) |
| Find with function | soup.find_all(lambda tag: condition) |
| Limit results | soup.find_all('div', limit=5) |
Common Issues and Solutions
Issue: Parser Differences Causing Inconsistent Results
Problem: Same HTML produces different parse trees with different parsers
Solution:
from bs4 import BeautifulSoup
html = '<p>Test<p>Another' # Malformed HTML
# Different parsers handle this differently
# Use lxml for most cases - it's fast and lenient
soup = BeautifulSoup(html, 'lxml')
# For maximum browser compatibility, use html5lib
soup = BeautifulSoup(html, 'html5lib')
# Stick to one parser throughout your project for consistency
Issue: AttributeError When Element Not Found
Problem: NoneType has no attribute when element doesn't exist
Solution:
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, 'lxml')
# Wrong - raises AttributeError if not found
# text = soup.find('div', class_='missing').text
# Right - check for None first
element = soup.find('div', class_='missing')
text = element.text if element else ''
# Or use select_one with conditional
text = (soup.select_one('.missing') or soup.new_tag('div')).get_text()
# Or with walrus operator (Python 3.8+)
if (elem := soup.find('div', class_='target')):
print(elem.text)
Issue: Whitespace in Sibling Navigation
Problem: next_sibling returns whitespace NavigableString
Solution:
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, 'lxml')
element = soup.find('p')
# Wrong - may return whitespace
next_elem = element.next_sibling
# Right - use find_next_sibling for tags
next_p = element.find_next_sibling('p')
next_any = element.find_next_sibling() # Any tag, not text
# Or filter manually
next_elem = element.next_sibling
while next_elem and isinstance(next_elem, str):
next_elem = next_elem.next_sibling
Issue: Class Attribute Returns List
Problem: element['class'] returns list, not string
Solution:
from bs4 import BeautifulSoup
soup = BeautifulSoup('<p class="intro main">Text</p>', 'lxml')
p = soup.find('p')
# class_ is always a list
classes = p.get('class', []) # ['intro', 'main']
# Join if you need a string
class_string = ' '.join(p.get('class', [])) # 'intro main'
# Check if class exists
if 'intro' in p.get('class', []):
print('Has intro class')
Issue: Encoding Problems with Special Characters
Problem: Garbled text or encoding errors
Solution:
from bs4 import BeautifulSoup
# Let BeautifulSoup detect encoding from bytes
response = requests.get(url)
soup = BeautifulSoup(response.content, 'lxml') # Use .content not .text
# Check detected encoding
print(soup.original_encoding)
# Force specific encoding if detection fails
soup = BeautifulSoup(html_bytes, 'lxml', from_encoding='utf-8')
# Output with specific encoding
output = soup.encode('utf-8')
Issue: Memory Issues with Large Documents
Problem: High memory usage when parsing large HTML files
Solution:
from bs4 import BeautifulSoup, SoupStrainer
# Parse only specific parts using SoupStrainer
only_articles = SoupStrainer('article')
soup = BeautifulSoup(html, 'lxml', parse_only=only_articles)
# Or parse only elements with specific class
only_products = SoupStrainer(class_='product')
soup = BeautifulSoup(html, 'lxml', parse_only=only_products)
# Clean up when done
soup.decompose()
del soup
Issue: find_all Not Finding Elements with Multiple Classes
Problem: Can't find elements with exact multiple classes
Solution:
from bs4 import BeautifulSoup
html = '<div class="btn btn-primary large">Button</div>'
soup = BeautifulSoup(html, 'lxml')
# This finds element with 'btn' as one of its classes
soup.find_all(class_='btn') # Works
# For multiple classes, use CSS selector
soup.select('.btn.btn-primary') # Works
# Or check all classes
soup.find_all(class_=['btn', 'btn-primary']) # Matches any
# For exact class match, use custom function
def exact_classes(tag):
return tag.get('class') == ['btn', 'btn-primary', 'large']
soup.find_all(exact_classes)
Issue: Modifying Tree During Iteration
Problem: Unexpected behaviour when modifying while iterating
Solution:
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, 'lxml')
# Wrong - modifying during iteration
for tag in soup.find_all('script'):
tag.decompose() # May skip elements
# Right - convert to list first
for tag in list(soup.find_all('script')):
tag.decompose()
# Or use copy
from copy import copy
for tag in soup.find_all('script')[:]: # Slice creates copy
tag.decompose()
Issue: Relative URLs in Scraped Content
Problem: Extracted URLs are relative, not absolute
Solution:
from bs4 import BeautifulSoup
from urllib.parse import urljoin
base_url = 'https://example.com/page/'
soup = BeautifulSoup(html, 'lxml')
# Convert all relative URLs to absolute
for link in soup.find_all('a', href=True):
link['href'] = urljoin(base_url, link['href'])
for img in soup.find_all('img', src=True):
img['src'] = urljoin(base_url, img['src'])
# When extracting, always convert
urls = [
urljoin(base_url, a['href'])
for a in soup.find_all('a', href=True)
]
Related Topics
- Python Requests - HTTP library for fetching web pages
- Scrapy - Full-featured web scraping framework
- Selenium - Browser automation for JavaScript-rendered content
- lxml - Fast XML and HTML processing library
- CSS Selectors - Understanding selector syntax for web scraping
- Regular Expressions - Pattern matching for text extraction
- XPath - Alternative query language for XML/HTML navigation