Python Requests
A powerful and elegant HTTP library for Python that simplifies making web requests and handling responses.
Python Requests Cheatsheet
A powerful and elegant HTTP library for Python that simplifies making web requests and handling responses.
Overview
The requests library is the de facto standard for making HTTP requests in Python, providing a simple and intuitive API for interacting with web services, APIs, and websites.
flowchart LR
subgraph Client["Python Application"]
A[requests.get/post/etc]
end
subgraph Request["HTTP Request"]
B[Headers]
C[Body/Data]
D[Parameters]
E[Authentication]
end
subgraph Server["Web Server"]
F[API Endpoint]
end
subgraph Response["HTTP Response"]
G[Status Code]
H[Headers]
I[Content/JSON]
end
A --> B & C & D & E
B & C & D & E --> F
F --> G & H & I
G & H & I --> A
Installation
pip install requests
Making GET and POST Requests
Key Concepts
- GET requests retrieve data from a server without modifying it
- POST requests send data to a server to create or update resources
- Response object contains status code, headers, and content
- HTTP methods include GET, POST, PUT, DELETE, PATCH, HEAD, OPTIONS
Common Patterns
import requests
# Basic GET request
response = requests.get('https://api.example.com/data')
# Basic POST request
response = requests.post('https://api.example.com/data', data={'key': 'value'})
# Other HTTP methods
response = requests.put('https://api.example.com/data/1', data={'key': 'updated'})
response = requests.delete('https://api.example.com/data/1')
response = requests.patch('https://api.example.com/data/1', data={'field': 'value'})
response = requests.head('https://api.example.com/data')
response = requests.options('https://api.example.com/data')
Examples
Simple GET Request
import requests
# Fetch data from an API
response = requests.get('https://api.github.com/users/octocat')
# Check if request was successful
if response.status_code == 200:
print(f"Status: {response.status_code}")
print(f"Content-Type: {response.headers['Content-Type']}")
print(f"Data: {response.json()}")
else:
print(f"Request failed with status: {response.status_code}")
POST Request with Form Data
import requests
# Send form data
payload = {
'username': 'testuser',
'email': 'test@example.com'
}
response = requests.post(
'https://api.example.com/users',
data=payload
)
print(f"Status: {response.status_code}")
print(f"Response: {response.text}")
Custom Headers
import requests
headers = {
'User-Agent': 'MyApp/1.0',
'Accept': 'application/json',
'X-Custom-Header': 'custom-value'
}
response = requests.get(
'https://api.example.com/data',
headers=headers
)
Handling Query Parameters and JSON Data
Key Concepts
- Query parameters are key-value pairs appended to URLs after
? - JSON data is sent in the request body with proper Content-Type header
- Response JSON can be parsed directly using
.json()method - URL encoding is handled automatically by requests
Common Patterns
import requests
# Query parameters
params = {'search': 'python', 'page': 1, 'limit': 10}
response = requests.get('https://api.example.com/search', params=params)
# Send JSON data
json_data = {'name': 'test', 'value': 123}
response = requests.post('https://api.example.com/data', json=json_data)
# Parse JSON response
data = response.json()
Examples
Query Parameters
import requests
# Multiple query parameters
params = {
'q': 'requests library',
'sort': 'stars',
'order': 'desc',
'per_page': 5
}
response = requests.get(
'https://api.github.com/search/repositories',
params=params
)
# The actual URL becomes:
# https://api.github.com/search/repositories?q=requests+library&sort=stars&order=desc&per_page=5
data = response.json()
for repo in data.get('items', []):
print(f"{repo['name']}: {repo['stargazers_count']} stars")
Sending and Receiving JSON
import requests
# POST JSON data
payload = {
'title': 'New Post',
'body': 'This is the content',
'userId': 1
}
response = requests.post(
'https://jsonplaceholder.typicode.com/posts',
json=payload # Automatically sets Content-Type: application/json
)
# Parse JSON response
created_post = response.json()
print(f"Created post ID: {created_post['id']}")
print(f"Title: {created_post['title']}")
Handling Different Response Types
import requests
response = requests.get('https://api.example.com/data')
# Text content
text_content = response.text
# Binary content (images, files)
binary_content = response.content
# JSON content
json_content = response.json()
# Check encoding
print(f"Encoding: {response.encoding}")
Authentication (Basic, OAuth)
Key Concepts
- Basic Authentication sends username and password encoded in Base64
- Bearer Token authentication uses tokens in the Authorisation header
- OAuth 2.0 is a standard protocol for authorisation
- API Keys can be sent as headers, query parameters, or in the body
flowchart TD
A[Authentication Methods] --> B[Basic Auth]
A --> C[Bearer Token]
A --> D[OAuth 2.0]
A --> E[API Key]
B --> B1[Username + Password]
C --> C1[JWT/Access Token]
D --> D1[Authorisation Code]
D --> D2[Client Credentials]
E --> E1[Header/Query Param]
Common Patterns
import requests
from requests.auth import HTTPBasicAuth, HTTPDigestAuth
# Basic authentication
response = requests.get(url, auth=('username', 'password'))
# or
response = requests.get(url, auth=HTTPBasicAuth('username', 'password'))
# Bearer token
headers = {'Authorization': 'Bearer <token>'}
response = requests.get(url, headers=headers)
# Digest authentication
response = requests.get(url, auth=HTTPDigestAuth('username', 'password'))
Examples
Basic Authentication
import requests
from requests.auth import HTTPBasicAuth
# Method 1: Tuple shorthand
response = requests.get(
'https://api.example.com/user',
auth=('username', 'password')
)
# Method 2: HTTPBasicAuth class
response = requests.get(
'https://api.example.com/user',
auth=HTTPBasicAuth('username', 'password')
)
if response.status_code == 200:
print("Authentication successful")
print(response.json())
elif response.status_code == 401:
print("Authentication failed")
Bearer Token Authentication
import requests
# Using an access token
access_token = 'your_access_token_here'
headers = {
'Authorization': f'Bearer {access_token}',
'Content-Type': 'application/json'
}
response = requests.get(
'https://api.example.com/protected-resource',
headers=headers
)
print(response.json())
OAuth 2.0 Client Credentials Flow
import requests
# Step 1: Get access token
token_url = 'https://auth.example.com/oauth/token'
client_id = 'your_client_id'
client_secret = 'your_client_secret'
token_response = requests.post(
token_url,
data={
'grant_type': 'client_credentials',
'client_id': client_id,
'client_secret': client_secret,
'scope': 'read write'
}
)
token_data = token_response.json()
access_token = token_data['access_token']
# Step 2: Use access token for API requests
headers = {'Authorization': f'Bearer {access_token}'}
api_response = requests.get(
'https://api.example.com/data',
headers=headers
)
print(api_response.json())
API Key Authentication
import requests
api_key = 'your_api_key_here'
# Method 1: As a header
response = requests.get(
'https://api.example.com/data',
headers={'X-API-Key': api_key}
)
# Method 2: As a query parameter
response = requests.get(
'https://api.example.com/data',
params={'api_key': api_key}
)
Session Management
Key Concepts
- Sessions persist parameters across requests (cookies, headers, auth)
- Connection pooling reuses TCP connections for better performance
- Cookies are automatically handled and persisted within a session
- Session context manager ensures proper resource cleanup
flowchart LR
subgraph Session["requests.Session()"]
A[Persistent Cookies]
B[Shared Headers]
C[Connection Pool]
D[Authentication]
end
Session --> E[Request 1]
Session --> F[Request 2]
Session --> G[Request 3]
E & F & G --> H[Same Server]
Common Patterns
import requests
# Create a session
session = requests.Session()
# Set default headers for all requests
session.headers.update({'User-Agent': 'MyApp/1.0'})
# Set default authentication
session.auth = ('username', 'password')
# Make requests using the session
response = session.get('https://api.example.com/data')
# Close the session when done
session.close()
Examples
Basic Session Usage
import requests
# Using session as context manager (recommended)
with requests.Session() as session:
# Set session-wide headers
session.headers.update({
'User-Agent': 'MyApp/1.0',
'Accept': 'application/json'
})
# First request - login
login_response = session.post(
'https://example.com/login',
data={'username': 'user', 'password': 'pass'}
)
# Subsequent requests automatically include session cookies
profile_response = session.get('https://example.com/profile')
settings_response = session.get('https://example.com/settings')
print(f"Profile: {profile_response.json()}")
print(f"Settings: {settings_response.json()}")
Session with Persistent Authentication
import requests
session = requests.Session()
# Set up authentication for all requests
session.auth = ('api_user', 'api_password')
# Set base URL pattern
base_url = 'https://api.example.com'
try:
# All requests use the same authentication
users = session.get(f'{base_url}/users').json()
posts = session.get(f'{base_url}/posts').json()
comments = session.get(f'{base_url}/comments').json()
print(f"Users: {len(users)}")
print(f"Posts: {len(posts)}")
print(f"Comments: {len(comments)}")
finally:
session.close()
Managing Cookies
import requests
with requests.Session() as session:
# Make a request that sets cookies
session.get('https://example.com')
# View all cookies
print("Cookies:")
for cookie in session.cookies:
print(f" {cookie.name}: {cookie.value}")
# Manually set a cookie
session.cookies.set('custom_cookie', 'custom_value', domain='example.com')
# Clear specific cookie
session.cookies.clear(domain='example.com', path='/', name='custom_cookie')
# Clear all cookies
session.cookies.clear()
Error Handling and Retries
Key Concepts
- Status codes indicate request success or failure (2xx, 4xx, 5xx)
- Exceptions are raised for network errors, timeouts, etc.
- Retries can be implemented using adapters with backoff strategies
- Timeouts prevent requests from hanging indefinitely
flowchart TD
A[Make Request] --> B{Network Error?}
B -->|Yes| C[ConnectionError]
B -->|No| D{Timeout?}
D -->|Yes| E[Timeout Exception]
D -->|No| F{Status Code}
F -->|2xx| G[Success]
F -->|4xx| H[Client Error]
F -->|5xx| I[Server Error]
C --> J[Retry Logic]
E --> J
I --> J
J --> A
Common Patterns
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
# Basic error handling
try:
response = requests.get(url, timeout=10)
response.raise_for_status() # Raises HTTPError for 4xx/5xx
except requests.exceptions.RequestException as e:
print(f"Error: {e}")
# Configure retries
retry_strategy = Retry(
total=3,
backoff_factor=1,
status_forcelist=[429, 500, 502, 503, 504]
)
adapter = HTTPAdapter(max_retries=retry_strategy)
session = requests.Session()
session.mount("https://", adapter)
Examples
Comprehensive Error Handling
import requests
from requests.exceptions import (
RequestException,
ConnectionError,
HTTPError,
Timeout,
TooManyRedirects
)
def make_safe_request(url):
try:
response = requests.get(url, timeout=(5, 30)) # (connect, read) timeout
response.raise_for_status()
return response.json()
except ConnectionError:
print("Failed to connect to the server")
except Timeout:
print("Request timed out")
except TooManyRedirects:
print("Too many redirects")
except HTTPError as e:
print(f"HTTP error occurred: {e.response.status_code}")
if e.response.status_code == 404:
print("Resource not found")
elif e.response.status_code == 401:
print("Authentication required")
elif e.response.status_code == 403:
print("Access forbidden")
elif e.response.status_code >= 500:
print("Server error - try again later")
except RequestException as e:
print(f"An error occurred: {e}")
return None
# Usage
data = make_safe_request('https://api.example.com/data')
if data:
print(data)
Implementing Retries with Backoff
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
def create_session_with_retries(
retries=3,
backoff_factor=0.5,
status_forcelist=(500, 502, 503, 504)
):
"""Create a session with automatic retry logic."""
session = requests.Session()
retry_strategy = Retry(
total=retries,
read=retries,
connect=retries,
backoff_factor=backoff_factor,
status_forcelist=status_forcelist,
allowed_methods=["HEAD", "GET", "OPTIONS", "POST"]
)
adapter = HTTPAdapter(max_retries=retry_strategy)
session.mount("http://", adapter)
session.mount("https://", adapter)
return session
# Usage
session = create_session_with_retries()
try:
response = session.get('https://api.example.com/data', timeout=10)
response.raise_for_status()
print(response.json())
except requests.exceptions.RequestException as e:
print(f"Request failed after retries: {e}")
finally:
session.close()
Custom Retry Decorator
import requests
import time
from functools import wraps
def retry_request(max_retries=3, delay=1, backoff=2):
"""Decorator for retrying requests with exponential backoff."""
def decorator(func):
@wraps(func)
def wrapper(*args, **kwargs):
retries = 0
current_delay = delay
while retries < max_retries:
try:
return func(*args, **kwargs)
except requests.exceptions.RequestException as e:
retries += 1
if retries == max_retries:
raise
print(f"Attempt {retries} failed: {e}")
print(f"Retrying in {current_delay} seconds...")
time.sleep(current_delay)
current_delay *= backoff
return None
return wrapper
return decorator
@retry_request(max_retries=3, delay=1, backoff=2)
def fetch_data(url):
response = requests.get(url, timeout=10)
response.raise_for_status()
return response.json()
# Usage
try:
data = fetch_data('https://api.example.com/data')
print(data)
except requests.exceptions.RequestException as e:
print(f"All retries failed: {e}")
File Uploads and Downloads
Key Concepts
- Multipart form data is used for file uploads
- Streaming allows handling large files without loading into memory
- Progress tracking can be implemented for large transfers
- Content-Disposition header provides filename information
Common Patterns
import requests
# Upload a file
with open('file.txt', 'rb') as f:
files = {'file': f}
response = requests.post(url, files=files)
# Download a file (with closes the connection when done)
with requests.get(url, stream=True) as response:
with open('downloaded_file', 'wb') as f:
for chunk in response.iter_content(chunk_size=8192):
f.write(chunk)
Examples
File Upload
import requests
# Simple file upload
def upload_file(url, file_path):
with open(file_path, 'rb') as f:
files = {'file': f}
response = requests.post(url, files=files)
return response
# Upload with additional data
def upload_file_with_data(url, file_path, metadata):
with open(file_path, 'rb') as f:
files = {'file': (file_path, f, 'application/octet-stream')}
data = {'metadata': metadata}
response = requests.post(url, files=files, data=data)
return response
# Multiple file upload
def upload_multiple_files(url, file_paths):
files = []
file_handles = []
try:
for path in file_paths:
f = open(path, 'rb')
file_handles.append(f)
files.append(('files', (path, f, 'application/octet-stream')))
response = requests.post(url, files=files)
return response
finally:
for f in file_handles:
f.close()
# Usage
response = upload_file('https://api.example.com/upload', 'document.pdf')
print(f"Upload status: {response.status_code}")
File Download with Progress
import requests
from tqdm import tqdm # pip install tqdm
def download_file(url, filename):
"""Download a file with progress bar."""
with requests.get(url, stream=True) as response:
response.raise_for_status()
# Get total file size from headers
total_size = int(response.headers.get('content-length', 0))
with open(filename, 'wb') as f:
with tqdm(total=total_size, unit='B', unit_scale=True) as pbar:
for chunk in response.iter_content(chunk_size=8192):
if chunk:
f.write(chunk)
pbar.update(len(chunk))
print(f"Downloaded: {filename}")
# Simple download without progress
def download_file_simple(url, filename):
"""Download a file without progress tracking."""
with requests.get(url, stream=True) as response:
response.raise_for_status()
with open(filename, 'wb') as f:
for chunk in response.iter_content(chunk_size=8192):
if chunk:
f.write(chunk)
return filename
# Usage
download_file('https://example.com/largefile.zip', 'largefile.zip')
Streaming Large Files
import requests
def stream_large_file(url, output_path, chunk_size=1024*1024):
"""Stream download large files in chunks (1MB default)."""
with requests.get(url, stream=True) as response:
response.raise_for_status()
with open(output_path, 'wb') as f:
downloaded = 0
for chunk in response.iter_content(chunk_size=chunk_size):
if chunk:
f.write(chunk)
downloaded += len(chunk)
print(f"Downloaded: {downloaded / (1024*1024):.2f} MB", end='\r')
print(f"\nComplete: {output_path}")
# Upload large file in chunks
def upload_large_file(url, file_path):
"""Upload a large file using streaming."""
def file_generator(file_path, chunk_size=8192):
with open(file_path, 'rb') as f:
while True:
chunk = f.read(chunk_size)
if not chunk:
break
yield chunk
response = requests.post(url, data=file_generator(file_path))
return response
Quick Reference
| Operation | Code Example |
|---|---|
| GET request | requests.get(url) |
| POST request | requests.post(url, data={'key': 'value'}) |
| POST JSON | requests.post(url, json={'key': 'value'}) |
| Query parameters | requests.get(url, params={'q': 'search'}) |
| Custom headers | requests.get(url, headers={'User-Agent': 'App'}) |
| Basic auth | requests.get(url, auth=('user', 'pass')) |
| Bearer token | requests.get(url, headers={'Authorization': 'Bearer token'}) |
| Timeout | requests.get(url, timeout=10) |
| Timeout (connect, read) | requests.get(url, timeout=(5, 30)) |
| Session | session = requests.Session() |
| Raise for status | response.raise_for_status() |
| JSON response | response.json() |
| Text response | response.text |
| Binary response | response.content |
| Status code | response.status_code |
| Response headers | response.headers |
| Cookies | response.cookies |
| Stream response | requests.get(url, stream=True) |
| File upload | requests.post(url, files={'file': open('f.txt', 'rb')}) |
| Disable SSL verify | requests.get(url, verify=False) |
| Custom CA bundle | requests.get(url, verify='/path/to/cert.pem') |
| Proxy | requests.get(url, proxies={'https': 'http://proxy:8080'}) |
| Disable redirects | requests.get(url, allow_redirects=False) (default is True for GET; HEAD defaults to False) |
Common Issues and Solutions
Issue: SSL Certificate Verification Errors
Problem: requests.exceptions.SSLError: certificate verify failed
Solution:
import requests
# Option 1: Specify CA bundle
response = requests.get(url, verify='/path/to/ca-bundle.crt')
# Option 2: Disable verification (NOT recommended for production)
import urllib3
urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning)
response = requests.get(url, verify=False)
# Option 3: Update certifi package
# pip install --upgrade certifi
Issue: Connection Timeout vs Read Timeout
Problem: Need different timeouts for connection and reading
Solution:
import requests
# Tuple format: (connect_timeout, read_timeout)
response = requests.get(url, timeout=(3.05, 27))
# 3.05 seconds to establish connection
# 27 seconds to wait for response
Issue: Handling Rate Limiting (429 Status)
Problem: API returns 429 Too Many Requests
Solution:
import requests
import time
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
def parse_retry_after(value, default=60):
"""Retry-After may be delay-seconds or an HTTP-date (RFC 9110)."""
if value is None:
return default
try:
return max(0, int(value))
except ValueError:
try:
dt = parsedate_to_datetime(value)
return max(0, (dt - datetime.now(timezone.utc)).total_seconds())
except (TypeError, ValueError):
return default
def request_with_rate_limit(url, max_retries=5):
for attempt in range(max_retries):
response = requests.get(url)
if response.status_code == 429:
retry_after = parse_retry_after(response.headers.get('Retry-After'))
print(f"Rate limited. Waiting {retry_after} seconds...")
time.sleep(retry_after)
continue
return response
raise Exception("Max retries exceeded due to rate limiting")
Issue: Memory Issues with Large Responses
Problem: Out of memory when downloading large files
Solution:
import requests
# Use streaming to avoid loading entire response into memory
# (with ensures the connection is released when done)
with requests.get(url, stream=True) as response:
with open('large_file.zip', 'wb') as f:
for chunk in response.iter_content(chunk_size=8192):
f.write(chunk)
Issue: JSON Decode Errors
Problem: json.decoder.JSONDecodeError when parsing response
Solution:
import requests
response = requests.get(url)
# Check content type before parsing
if 'application/json' in response.headers.get('Content-Type', ''):
try:
data = response.json()
except ValueError as e:
print(f"Invalid JSON: {e}")
print(f"Response text: {response.text[:200]}")
else:
print(f"Unexpected content type: {response.headers.get('Content-Type')}")
Issue: Character Encoding Problems
Problem: Garbled text in response
Solution:
import requests
response = requests.get(url)
# Check detected encoding
print(f"Detected encoding: {response.encoding}")
# Override encoding if needed
response.encoding = 'utf-8'
text = response.text
# Or use apparent_encoding
response.encoding = response.apparent_encoding
text = response.text
Issue: Proxy Configuration
Problem: Need to route requests through a proxy
Solution:
import requests
proxies = {
'http': 'http://proxy.example.com:8080',
'https': 'http://proxy.example.com:8080'
}
# With authentication
proxies = {
'http': 'http://user:password@proxy.example.com:8080',
'https': 'http://user:password@proxy.example.com:8080'
}
response = requests.get(url, proxies=proxies)
# Or set environment variables
# export HTTP_PROXY="http://proxy.example.com:8080"
# export HTTPS_PROXY="http://proxy.example.com:8080"
Issue: Handling Redirects
Problem: Need to track or disable redirects
Solution:
import requests
# Redirects are followed by default (allow_redirects=True) for all
# methods except HEAD, which defaults to allow_redirects=False
# Disable redirects
response = requests.get(url, allow_redirects=False)
# Access redirect history
response = requests.get(url)
for redirect in response.history:
print(f"Redirected: {redirect.status_code} -> {redirect.url}")
print(f"Final URL: {response.url}")
Related Topics
- aiohttp - Asynchronous HTTP client/server for Python using asyncio
- httpx - Modern HTTP client with async support and HTTP/2 capabilities
- urllib3 - Powerful HTTP client library underlying requests
- REST API Design - Best practices for designing and consuming RESTful APIs
- OAuth 2.0 and JWT - Deep dive into modern authentication and authorisation
- Web Scraping with BeautifulSoup - Parsing HTML responses and extracting data