Python Debugging Tools
Comprehensive guide to debugging, profiling, and troubleshooting Python applications.
Python Debugging Tools
Comprehensive guide to debugging, profiling, and troubleshooting Python applications.
Overview
Python provides a rich ecosystem of debugging tools, from the built-in pdb debugger to sophisticated profiling utilities. Effective debugging combines interactive debugging for step-through analysis, strategic logging for production monitoring, and profiling for performance optimisation.
flowchart LR
A[Bug Detected] --> B{Type?}
B -->|Logic Error| C[pdb/IDE Debugger]
B -->|Performance| D[cProfile/timeit]
B -->|Production| E[Logging]
C --> F[Fix & Test]
D --> F
E --> F
F --> G[Verify Solution]
Using pdb for Interactive Debugging
The Python Debugger (pdb) allows step-by-step execution and state inspection.
Starting pdb
# Method 1: Insert breakpoint in code
import pdb; pdb.set_trace()
# Method 2: Python 3.7+ built-in breakpoint
breakpoint()
# Method 3: Run script with pdb
# python -m pdb script.py
# Method 4: Post-mortem debugging after exception
import pdb
try:
problematic_function()
except Exception:
pdb.post_mortem()
Essential pdb Commands
# Navigation
n # Next line (step over)
s # Step into function
r # Return from function
c # Continue until next breakpoint
q # Quit debugger
# Inspection
p variable # Print variable value
pp variable # Pretty print
l # List source code around current line
l 1, 20 # List lines 1-20
ll # List entire current function
w # Print stack trace
u # Move up one frame in stack
d # Move down one frame in stack
# Breakpoints
b 42 # Set breakpoint at line 42
b func # Break when entering function
b file.py:42 # Break at line in specific file
b 42, x > 5 # Conditional breakpoint
cl 1 # Clear breakpoint 1
cl # Clear all breakpoints
disable 1 # Disable breakpoint 1
enable 1 # Enable breakpoint 1
# Expression evaluation
!statement # Execute Python statement
interact # Start interactive interpreter
Advanced pdb Usage
# Custom breakpoint behaviour via environment variable
# PYTHONBREAKPOINT=0 python script.py # Disable all breakpoints
# PYTHONBREAKPOINT=ipdb.set_trace python script.py # Use ipdb instead
# Using pdb++ for enhanced debugging (install: pip install pdbpp)
import pdb
pdb.set_trace() # Automatically uses pdb++ if installed
# Remote debugging with remote-pdb
from remote_pdb import RemotePdb
RemotePdb('127.0.0.1', 4444).set_trace()
# Connect: telnet 127.0.0.1 4444
Debugging Workflow Example
def calculate_average(numbers):
breakpoint() # Execution pauses here
total = sum(numbers)
count = len(numbers)
return total / count
# In pdb session:
# (Pdb) p numbers # Check input
# (Pdb) n # Execute sum()
# (Pdb) p total # Verify total
# (Pdb) p count # Check count (catch division by zero)
# (Pdb) c # Continue execution
Logging Module Setup and Usage
Python's logging module provides flexible, production-ready logging.
Basic Configuration
import logging
# Simple configuration
logging.basicConfig(
level=logging.DEBUG,
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s',
datefmt='%Y-%m-%d %H:%M:%S'
)
logger = logging.getLogger(__name__)
# Log levels (in order of severity)
logger.debug('Detailed information for diagnosing')
logger.info('Confirmation that things work')
logger.warning('Something unexpected happened')
logger.error('Serious problem occurred')
logger.critical('Program may not continue')
Advanced Configuration with Handlers
import logging
from logging.handlers import RotatingFileHandler, TimedRotatingFileHandler
# Create logger
logger = logging.getLogger('myapp')
logger.setLevel(logging.DEBUG)
# Console handler with INFO level
console_handler = logging.StreamHandler()
console_handler.setLevel(logging.INFO)
console_format = logging.Formatter('%(levelname)s - %(message)s')
console_handler.setFormatter(console_format)
# File handler with DEBUG level and rotation
file_handler = RotatingFileHandler(
'app.log',
maxBytes=10*1024*1024, # 10MB
backupCount=5
)
file_handler.setLevel(logging.DEBUG)
file_format = logging.Formatter(
'%(asctime)s - %(name)s - %(levelname)s - %(funcName)s:%(lineno)d - %(message)s'
)
file_handler.setFormatter(file_format)
# Add handlers to logger
logger.addHandler(console_handler)
logger.addHandler(file_handler)
Structured Logging with Extra Fields
import logging
import json
class JSONFormatter(logging.Formatter):
def format(self, record):
log_data = {
'timestamp': self.formatTime(record),
'level': record.levelname,
'message': record.getMessage(),
'module': record.module,
'function': record.funcName,
}
# Include extra fields
if hasattr(record, 'user_id'):
log_data['user_id'] = record.user_id
if hasattr(record, 'request_id'):
log_data['request_id'] = record.request_id
return json.dumps(log_data)
# Usage
logger.info('User action', extra={'user_id': 123, 'request_id': 'abc-456'})
Logging Configuration via Dictionary
import logging.config
LOGGING_CONFIG = {
'version': 1,
'disable_existing_loggers': False,
'formatters': {
'standard': {
'format': '%(asctime)s [%(levelname)s] %(name)s: %(message)s'
},
},
'handlers': {
'default': {
'level': 'INFO',
'formatter': 'standard',
'class': 'logging.StreamHandler',
'stream': 'ext://sys.stdout',
},
'file': {
'level': 'DEBUG',
'formatter': 'standard',
'class': 'logging.FileHandler',
'filename': 'debug.log',
'mode': 'a',
},
},
'loggers': {
'': { # Root logger
'handlers': ['default', 'file'],
'level': 'DEBUG',
'propagate': True
},
'myapp.module': { # Specific module logger
'handlers': ['file'],
'level': 'DEBUG',
'propagate': False
},
}
}
logging.config.dictConfig(LOGGING_CONFIG)
Context Managers for Temporary Log Levels
import logging
from contextlib import contextmanager
@contextmanager
def temporary_log_level(logger, level):
"""Temporarily change log level."""
old_level = logger.level
logger.setLevel(level)
try:
yield
finally:
logger.setLevel(old_level)
# Usage
with temporary_log_level(logger, logging.DEBUG):
logger.debug('This will be logged')
Common Debugging Patterns
Print Debugging (Enhanced)
# Basic debug print with context
def debug_print(var_name, var_value):
import inspect
frame = inspect.currentframe().f_back
print(f"[{frame.f_code.co_filename}:{frame.f_lineno}] {var_name} = {var_value!r}")
# Using icecream for better debug output (pip install icecream)
from icecream import ic
ic.configureOutput(prefix='DEBUG| ')
x = 42
ic(x) # Output: DEBUG| x: 42
# f-string debugging (Python 3.8+)
name = "debug"
value = 42
print(f"{name=}, {value=}") # Output: name='debug', value=42
Decorator for Function Tracing
import functools
import logging
logger = logging.getLogger(__name__)
def trace(func):
"""Decorator to trace function calls."""
@functools.wraps(func)
def wrapper(*args, **kwargs):
args_repr = [repr(a) for a in args]
kwargs_repr = [f"{k}={v!r}" for k, v in kwargs.items()]
signature = ", ".join(args_repr + kwargs_repr)
logger.debug(f"Calling {func.__name__}({signature})")
try:
result = func(*args, **kwargs)
logger.debug(f"{func.__name__} returned {result!r}")
return result
except Exception as e:
logger.exception(f"{func.__name__} raised {e.__class__.__name__}")
raise
return wrapper
@trace
def calculate(x, y, operation='add'):
if operation == 'add':
return x + y
return x - y
Assert for Development Debugging
def process_data(items):
# Validate preconditions
assert items is not None, "Items cannot be None"
assert len(items) > 0, f"Expected non-empty list, got {len(items)} items"
result = []
for item in items:
processed = transform(item)
# Validate invariants during development
assert processed is not None, f"transform() returned None for {item}"
result.append(processed)
# Validate postconditions
assert len(result) == len(items), "Output length mismatch"
return result
# Disable assertions in production: python -O script.py
Debug Context Manager
import time
import traceback
from contextlib import contextmanager
@contextmanager
def debug_context(name, log_memory=False):
"""Context manager for debugging code blocks."""
start_time = time.perf_counter()
if log_memory:
import tracemalloc
tracemalloc.start()
print(f"[DEBUG] Entering: {name}")
try:
yield
except Exception as e:
print(f"[DEBUG] Exception in {name}: {e}")
traceback.print_exc()
raise
finally:
elapsed = time.perf_counter() - start_time
print(f"[DEBUG] Exiting: {name} (took {elapsed:.4f}s)")
if log_memory:
current, peak = tracemalloc.get_traced_memory()
tracemalloc.stop()
print(f"[DEBUG] Memory: current={current/1024:.1f}KB, peak={peak/1024:.1f}KB")
# Usage
with debug_context("data_processing", log_memory=True):
result = heavy_computation()
Profiling Code
Using cProfile
import cProfile
import pstats
from pstats import SortKey
# Profile a function
def profile_function():
cProfile.run('my_function()', 'output.prof')
# Analyse results
stats = pstats.Stats('output.prof')
stats.strip_dirs()
stats.sort_stats(SortKey.CUMULATIVE)
stats.print_stats(20) # Top 20 functions
# Profile from command line
# python -m cProfile -o output.prof script.py
# python -m cProfile -s cumulative script.py
# Context manager for profiling
from contextlib import contextmanager
@contextmanager
def profile_block(output_file=None):
"""Profile a code block."""
profiler = cProfile.Profile()
profiler.enable()
try:
yield profiler
finally:
profiler.disable()
if output_file:
profiler.dump_stats(output_file)
else:
stats = pstats.Stats(profiler)
stats.strip_dirs()
stats.sort_stats(SortKey.CUMULATIVE)
stats.print_stats(10)
# Usage
with profile_block():
expensive_operation()
Using timeit for Micro-benchmarks
import timeit
# Time a single statement
time_taken = timeit.timeit(
'sum(range(1000))',
number=10000
)
print(f"Average: {time_taken/10000*1000:.3f}ms")
# Time with setup code
time_taken = timeit.timeit(
stmt='sorted(data)',
setup='import random; data = [random.randint(0, 1000) for _ in range(100)]',
number=1000
)
# Compare implementations
def method1():
return [x**2 for x in range(1000)]
def method2():
return list(map(lambda x: x**2, range(1000)))
t1 = timeit.timeit(method1, number=10000)
t2 = timeit.timeit(method2, number=10000)
print(f"List comprehension: {t1:.4f}s")
print(f"Map: {t2:.4f}s")
# Using repeat for more reliable results
results = timeit.repeat(
'sorted(data)',
setup='import random; data = [random.randint(0, 1000) for _ in range(100)]',
number=1000,
repeat=5
)
print(f"Min: {min(results):.4f}s, Mean: {sum(results)/len(results):.4f}s")
Memory Profiling
# Using tracemalloc (built-in)
import tracemalloc
tracemalloc.start()
# Your code here
data = [x**2 for x in range(100000)]
snapshot = tracemalloc.take_snapshot()
top_stats = snapshot.statistics('lineno')
print("Top 10 memory allocations:")
for stat in top_stats[:10]:
print(stat)
# Using memory_profiler (pip install memory-profiler)
from memory_profiler import profile
@profile
def memory_intensive_function():
a = [1] * (10 ** 6)
b = [2] * (2 * 10 ** 7)
del b
return a
# Run with: python -m memory_profiler script.py
Line Profiler for Detailed Analysis
# Install: pip install line_profiler
# Add decorator to functions to profile
@profile
def slow_function(n):
total = 0
for i in range(n):
total += sum(range(i))
return total
# Run with: kernprof -l -v script.py
# Or programmatically
from line_profiler import LineProfiler
def profile_lines(func, *args, **kwargs):
profiler = LineProfiler()
profiler.add_function(func)
profiler.enable_by_count()
result = func(*args, **kwargs)
profiler.disable_by_count()
profiler.print_stats()
return result
Exception Handling Best Practices
Proper Exception Hierarchy
flowchart TD
A[BaseException] --> B[Exception]
A --> C[SystemExit]
A --> D[KeyboardInterrupt]
B --> E[ValueError]
B --> F[TypeError]
B --> G[RuntimeError]
B --> H[Custom Exceptions]
Creating Custom Exceptions
class ApplicationError(Exception):
"""Base exception for application."""
pass
class ValidationError(ApplicationError):
"""Raised when validation fails."""
def __init__(self, field, message):
self.field = field
self.message = message
super().__init__(f"{field}: {message}")
class DatabaseError(ApplicationError):
"""Raised for database operations."""
def __init__(self, operation, original_error):
self.operation = operation
self.original_error = original_error
super().__init__(f"Database {operation} failed: {original_error}")
# Usage
raise ValidationError("email", "Invalid format")
Exception Handling Patterns
import logging
logger = logging.getLogger(__name__)
# Pattern 1: Specific exception handling
def read_config(path):
try:
with open(path) as f:
return json.load(f)
except FileNotFoundError:
logger.warning(f"Config not found: {path}, using defaults")
return {}
except json.JSONDecodeError as e:
logger.error(f"Invalid JSON in {path}: {e}")
raise ValidationError("config", f"Invalid JSON: {e}")
# Pattern 2: Exception chaining
def fetch_user(user_id):
try:
response = api.get(f'/users/{user_id}')
return response.json()
except requests.RequestException as e:
raise DatabaseError("fetch_user", e) from e
# Pattern 3: Context manager for cleanup
class DatabaseConnection:
def __enter__(self):
self.conn = connect()
return self.conn
def __exit__(self, exc_type, exc_val, exc_tb):
self.conn.close()
if exc_type is not None:
logger.error(f"Error during database operation: {exc_val}")
return False # Don't suppress exceptions
# Pattern 4: Retry with exponential backoff
import time
def retry_operation(func, max_retries=3, base_delay=1):
for attempt in range(max_retries):
try:
return func()
except (ConnectionError, TimeoutError) as e:
if attempt == max_retries - 1:
raise
delay = base_delay * (2 ** attempt)
logger.warning(f"Attempt {attempt + 1} failed, retrying in {delay}s: {e}")
time.sleep(delay)
Exception Logging Best Practices
import logging
import sys
logger = logging.getLogger(__name__)
# Log with full traceback
try:
risky_operation()
except Exception as e:
logger.exception("Operation failed") # Includes traceback
# or
logger.error("Operation failed", exc_info=True)
# Log and re-raise
try:
process_data()
except ValueError as e:
logger.error(f"Invalid data: {e}")
raise
# Global exception handler
def global_exception_handler(exc_type, exc_value, exc_traceback):
if issubclass(exc_type, KeyboardInterrupt):
sys.__excepthook__(exc_type, exc_value, exc_traceback)
return
logger.critical("Unhandled exception", exc_info=(exc_type, exc_value, exc_traceback))
sys.excepthook = global_exception_handler
Using warnings Module
import warnings
# Issue warnings for deprecated code
def old_function():
warnings.warn(
"old_function is deprecated, use new_function instead",
DeprecationWarning,
stacklevel=2
)
return new_function()
# Control warning behaviour
warnings.filterwarnings('error', category=DeprecationWarning) # Treat as error
warnings.filterwarnings('ignore', message='.*experimental.*') # Suppress specific
# Temporarily catch warnings
with warnings.catch_warnings(record=True) as w:
warnings.simplefilter("always")
result = legacy_function()
if w:
print(f"Warnings caught: {[str(warning.message) for warning in w]}")
Using IDE Debugging Features
VS Code Configuration
// .vscode/launch.json
{
"version": "0.2.0",
"configurations": [
{
"name": "Python: Current File",
"type": "debugpy",
"request": "launch",
"program": "${file}",
"console": "integratedTerminal",
"justMyCode": false,
"env": {
"PYTHONPATH": "${workspaceFolder}"
}
},
{
"name": "Python: FastAPI",
"type": "debugpy",
"request": "launch",
"module": "uvicorn",
"args": ["main:app", "--reload", "--port", "8000"],
"jinja": true,
"env": {
"DEBUG": "true"
}
},
{
"name": "Python: Pytest",
"type": "debugpy",
"request": "launch",
"module": "pytest",
"args": ["-v", "-s", "${file}"]
},
{
"name": "Python: Attach",
"type": "debugpy",
"request": "attach",
"connect": {
"host": "localhost",
"port": 5678
}
}
]
}
PyCharm Run Configurations
# Enable debugpy for remote debugging
import debugpy
# Allow other computers to attach
debugpy.listen(("0.0.0.0", 5678))
print("Waiting for debugger attach...")
debugpy.wait_for_client() # Optional: pause until debugger connects
# Continue with your code
main()
Breakpoint Types and Strategies
flowchart TD
A[Breakpoint Types] --> B[Line Breakpoint]
A --> C[Conditional]
A --> D[Logpoint]
A --> E[Exception]
B --> F["Pause at specific line"]
C --> G["Pause when condition true<br/>e.g., i == 100"]
D --> H["Log without pausing<br/>e.g., 'x = {x}'"]
E --> I["Pause on exception<br/>caught or uncaught"]
Common IDE Debug Features
# Conditional breakpoint expression examples
user.id == 42
len(items) > 100
'error' in response.text
request.method == 'POST'
# Watch expressions
len(data)
user.__dict__
[x for x in items if x.active]
sum(order.total for order in orders)
# Evaluate expressions in debug console
>>> locals()
>>> dir(object)
>>> type(variable).__mro__
>>> import sys; sys.getsizeof(data)
Debug Configuration for Tests
# pytest with debugging
# pytest --pdb # Drop into pdb on failure
# pytest -x --pdb # Stop on first failure, enter pdb
# pytest --trace # Drop into pdb at the start of each test
# pytest.ini configuration
"""
[pytest]
addopts = --tb=short
filterwarnings =
error
ignore::DeprecationWarning
"""
# Debug specific test
# pytest tests/test_module.py::test_function -v -s
Quick Reference
| Tool | Purpose | Command/Usage |
|---|---|---|
breakpoint() |
Insert debugger pause point | breakpoint() in code |
pdb |
Interactive debugging | python -m pdb script.py |
logging |
Production-ready logging | logging.getLogger(__name__) |
cProfile |
Function-level profiling | python -m cProfile -s cumulative script.py |
timeit |
Micro-benchmarking | timeit.timeit('code', number=1000) |
tracemalloc |
Memory profiling | tracemalloc.start() |
line_profiler |
Line-by-line profiling | kernprof -l -v script.py |
memory_profiler |
Memory per line | python -m memory_profiler script.py |
icecream |
Enhanced print debugging | from icecream import ic; ic(var) |
debugpy |
VS Code remote debugging | debugpy.listen(5678) |
pdb Command Quick Reference
| Command | Action |
|---|---|
n |
Next line (step over) |
s |
Step into function |
r |
Return from function |
c |
Continue to next breakpoint |
p var |
Print variable |
pp var |
Pretty print |
l |
List source code |
b 42 |
Breakpoint at line 42 |
b func |
Breakpoint at function |
w |
Show stack trace |
q |
Quit debugger |
Logging Levels
| Level | Value | Usage |
|---|---|---|
DEBUG |
10 | Detailed diagnostic information |
INFO |
20 | Confirmation of expected behaviour |
WARNING |
30 | Unexpected but handled situation |
ERROR |
40 | Serious problem, function failed |
CRITICAL |
50 | Program may not continue |
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
| Breakpoint not hit | Code path not executed | Verify control flow reaches breakpoint |
| pdb hangs on import | Circular import | Use python -m pdb instead of inline breakpoint |
| Logs not appearing | Logger level too high | Set logger and handler to DEBUG level |
| Profile shows wrong function | C extensions not profiled | Use line_profiler for Python-only code |
| Memory leak not found | Objects held by globals | Use gc.get_referrers() to find references |
| Exception swallowed | Bare except: clause |
Always specify exception type or log in except |
| Debug too slow | justMyCode enabled |
Set "justMyCode": false in launch.json |
| Remote debug won't attach | Firewall blocking port | Check port 5678 is open, verify host binding |
| Conditional breakpoint fails | Expression syntax error | Test expression in Python REPL first |
| Logging duplicate messages | Multiple handlers added | Check handler list with logger.handlers |
Debugging Anti-patterns to Avoid
# BAD: Bare except clause
try:
something()
except:
pass # Swallows all exceptions including KeyboardInterrupt
# GOOD: Specific exception handling
try:
something()
except ValueError as e:
logger.error(f"Value error: {e}")
raise
# BAD: Print debugging in production
print(f"DEBUG: {variable}") # Gets lost, no context
# GOOD: Structured logging
logger.debug("Processing item", extra={'item_id': item.id})
# BAD: Catching and re-raising without context
except Exception as e:
raise Exception(str(e)) # Loses original traceback
# GOOD: Chain exceptions
except Exception as e:
raise ApplicationError("Processing failed") from e
Performance Debugging Workflow
flowchart TD
A[Identify Slow Code] --> B[Profile with cProfile]
B --> C{Hotspot Found?}
C -->|Yes| D[Analyse with line_profiler]
C -->|No| E[Check I/O and Database]
D --> F[Optimise Algorithm]
E --> G[Add Caching/Async]
F --> H[Benchmark with timeit]
G --> H
H --> I{Fast Enough?}
I -->|No| B
I -->|Yes| J[Deploy]