Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Python Requests

A powerful and elegant HTTP library for Python that simplifies making web requests and handling responses.

Python Requests Cheatsheet

A powerful and elegant HTTP library for Python that simplifies making web requests and handling responses.

Overview

The requests library is the de facto standard for making HTTP requests in Python, providing a simple and intuitive API for interacting with web services, APIs, and websites.

HTTP ResponseWeb ServerHTTP RequestPython Applicationrequests.get/post/etcHeadersBody/DataParametersAuthenticationAPI EndpointStatus CodeHeadersContent/JSONHTTP ResponseWeb ServerHTTP RequestPython Applicationrequests.get/post/etcHeadersBody/DataParametersAuthenticationAPI EndpointStatus CodeHeadersContent/JSON

Installation

pip install requests

Making GET and POST Requests

Key Concepts

  • GET requests retrieve data from a server without modifying it
  • POST requests send data to a server to create or update resources
  • Response object contains status code, headers, and content
  • HTTP methods include GET, POST, PUT, DELETE, PATCH, HEAD, OPTIONS

Common Patterns

import requests

# Basic GET request
response = requests.get('https://api.example.com/data')

# Basic POST request
response = requests.post('https://api.example.com/data', data={'key': 'value'})

# Other HTTP methods
response = requests.put('https://api.example.com/data/1', data={'key': 'updated'})
response = requests.delete('https://api.example.com/data/1')
response = requests.patch('https://api.example.com/data/1', data={'field': 'value'})
response = requests.head('https://api.example.com/data')
response = requests.options('https://api.example.com/data')

Examples

Simple GET Request

import requests

# Fetch data from an API
response = requests.get('https://api.github.com/users/octocat')

# Check if request was successful
if response.status_code == 200:
    print(f"Status: {response.status_code}")
    print(f"Content-Type: {response.headers['Content-Type']}")
    print(f"Data: {response.json()}")
else:
    print(f"Request failed with status: {response.status_code}")

POST Request with Form Data

import requests

# Send form data
payload = {
    'username': 'testuser',
    'email': 'test@example.com'
}

response = requests.post(
    'https://api.example.com/users',
    data=payload
)

print(f"Status: {response.status_code}")
print(f"Response: {response.text}")

Custom Headers

import requests

headers = {
    'User-Agent': 'MyApp/1.0',
    'Accept': 'application/json',
    'X-Custom-Header': 'custom-value'
}

response = requests.get(
    'https://api.example.com/data',
    headers=headers
)

Handling Query Parameters and JSON Data

Key Concepts

  • Query parameters are key-value pairs appended to URLs after ?
  • JSON data is sent in the request body with proper Content-Type header
  • Response JSON can be parsed directly using .json() method
  • URL encoding is handled automatically by requests

Common Patterns

import requests

# Query parameters
params = {'search': 'python', 'page': 1, 'limit': 10}
response = requests.get('https://api.example.com/search', params=params)

# Send JSON data
json_data = {'name': 'test', 'value': 123}
response = requests.post('https://api.example.com/data', json=json_data)

# Parse JSON response
data = response.json()

Examples

Query Parameters

import requests

# Multiple query parameters
params = {
    'q': 'requests library',
    'sort': 'stars',
    'order': 'desc',
    'per_page': 5
}

response = requests.get(
    'https://api.github.com/search/repositories',
    params=params
)

# The actual URL becomes:
# https://api.github.com/search/repositories?q=requests+library&sort=stars&order=desc&per_page=5

data = response.json()
for repo in data.get('items', []):
    print(f"{repo['name']}: {repo['stargazers_count']} stars")

Sending and Receiving JSON

import requests

# POST JSON data
payload = {
    'title': 'New Post',
    'body': 'This is the content',
    'userId': 1
}

response = requests.post(
    'https://jsonplaceholder.typicode.com/posts',
    json=payload  # Automatically sets Content-Type: application/json
)

# Parse JSON response
created_post = response.json()
print(f"Created post ID: {created_post['id']}")
print(f"Title: {created_post['title']}")

Handling Different Response Types

import requests

response = requests.get('https://api.example.com/data')

# Text content
text_content = response.text

# Binary content (images, files)
binary_content = response.content

# JSON content
json_content = response.json()

# Check encoding
print(f"Encoding: {response.encoding}")

Authentication (Basic, OAuth)

Key Concepts

  • Basic Authentication sends username and password encoded in Base64
  • Bearer Token authentication uses tokens in the Authorisation header
  • OAuth 2.0 is a standard protocol for authorisation
  • API Keys can be sent as headers, query parameters, or in the body
AuthenticationMethodsBasic AuthBearer TokenOAuth 2.0API KeyUsername + PasswordJWT/Access TokenAuthorisation CodeClient CredentialsHeader/Query ParamAuthenticationMethodsBasic AuthBearer TokenOAuth 2.0API KeyUsername + PasswordJWT/Access TokenAuthorisation CodeClient CredentialsHeader/Query Param

Common Patterns

import requests
from requests.auth import HTTPBasicAuth, HTTPDigestAuth

# Basic authentication
response = requests.get(url, auth=('username', 'password'))
# or
response = requests.get(url, auth=HTTPBasicAuth('username', 'password'))

# Bearer token
headers = {'Authorization': 'Bearer <token>'}
response = requests.get(url, headers=headers)

# Digest authentication
response = requests.get(url, auth=HTTPDigestAuth('username', 'password'))

Examples

Basic Authentication

import requests
from requests.auth import HTTPBasicAuth

# Method 1: Tuple shorthand
response = requests.get(
    'https://api.example.com/user',
    auth=('username', 'password')
)

# Method 2: HTTPBasicAuth class
response = requests.get(
    'https://api.example.com/user',
    auth=HTTPBasicAuth('username', 'password')
)

if response.status_code == 200:
    print("Authentication successful")
    print(response.json())
elif response.status_code == 401:
    print("Authentication failed")

Bearer Token Authentication

import requests

# Using an access token
access_token = 'your_access_token_here'

headers = {
    'Authorization': f'Bearer {access_token}',
    'Content-Type': 'application/json'
}

response = requests.get(
    'https://api.example.com/protected-resource',
    headers=headers
)

print(response.json())

OAuth 2.0 Client Credentials Flow

import requests

# Step 1: Get access token
token_url = 'https://auth.example.com/oauth/token'
client_id = 'your_client_id'
client_secret = 'your_client_secret'

token_response = requests.post(
    token_url,
    data={
        'grant_type': 'client_credentials',
        'client_id': client_id,
        'client_secret': client_secret,
        'scope': 'read write'
    }
)

token_data = token_response.json()
access_token = token_data['access_token']

# Step 2: Use access token for API requests
headers = {'Authorization': f'Bearer {access_token}'}
api_response = requests.get(
    'https://api.example.com/data',
    headers=headers
)

print(api_response.json())

API Key Authentication

import requests

api_key = 'your_api_key_here'

# Method 1: As a header
response = requests.get(
    'https://api.example.com/data',
    headers={'X-API-Key': api_key}
)

# Method 2: As a query parameter
response = requests.get(
    'https://api.example.com/data',
    params={'api_key': api_key}
)

Session Management

Key Concepts

  • Sessions persist parameters across requests (cookies, headers, auth)
  • Connection pooling reuses TCP connections for better performance
  • Cookies are automatically handled and persisted within a session
  • Session context manager ensures proper resource cleanup
requests.Session()Persistent CookiesShared HeadersConnection PoolAuthenticationRequest 1Request 2Request 3Same Serverrequests.Session()Persistent CookiesShared HeadersConnection PoolAuthenticationRequest 1Request 2Request 3Same Server

Common Patterns

import requests

# Create a session
session = requests.Session()

# Set default headers for all requests
session.headers.update({'User-Agent': 'MyApp/1.0'})

# Set default authentication
session.auth = ('username', 'password')

# Make requests using the session
response = session.get('https://api.example.com/data')

# Close the session when done
session.close()

Examples

Basic Session Usage

import requests

# Using session as context manager (recommended)
with requests.Session() as session:
    # Set session-wide headers
    session.headers.update({
        'User-Agent': 'MyApp/1.0',
        'Accept': 'application/json'
    })

    # First request - login
    login_response = session.post(
        'https://example.com/login',
        data={'username': 'user', 'password': 'pass'}
    )

    # Subsequent requests automatically include session cookies
    profile_response = session.get('https://example.com/profile')
    settings_response = session.get('https://example.com/settings')

    print(f"Profile: {profile_response.json()}")
    print(f"Settings: {settings_response.json()}")

Session with Persistent Authentication

import requests

session = requests.Session()

# Set up authentication for all requests
session.auth = ('api_user', 'api_password')

# Set base URL pattern
base_url = 'https://api.example.com'

try:
    # All requests use the same authentication
    users = session.get(f'{base_url}/users').json()
    posts = session.get(f'{base_url}/posts').json()
    comments = session.get(f'{base_url}/comments').json()

    print(f"Users: {len(users)}")
    print(f"Posts: {len(posts)}")
    print(f"Comments: {len(comments)}")
finally:
    session.close()

Managing Cookies

import requests

with requests.Session() as session:
    # Make a request that sets cookies
    session.get('https://example.com')

    # View all cookies
    print("Cookies:")
    for cookie in session.cookies:
        print(f"  {cookie.name}: {cookie.value}")

    # Manually set a cookie
    session.cookies.set('custom_cookie', 'custom_value', domain='example.com')

    # Clear specific cookie
    session.cookies.clear(domain='example.com', path='/', name='custom_cookie')

    # Clear all cookies
    session.cookies.clear()

Error Handling and Retries

Key Concepts

  • Status codes indicate request success or failure (2xx, 4xx, 5xx)
  • Exceptions are raised for network errors, timeouts, etc.
  • Retries can be implemented using adapters with backoff strategies
  • Timeouts prevent requests from hanging indefinitely
YesNoYesNo2xx4xx5xxMake RequestNetwork Error?ConnectionErrorTimeout?Timeout ExceptionStatus CodeSuccessClient ErrorServer ErrorRetry LogicYesNoYesNo2xx4xx5xxMake RequestNetwork Error?ConnectionErrorTimeout?Timeout ExceptionStatus CodeSuccessClient ErrorServer ErrorRetry Logic

Common Patterns

import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

# Basic error handling
try:
    response = requests.get(url, timeout=10)
    response.raise_for_status()  # Raises HTTPError for 4xx/5xx
except requests.exceptions.RequestException as e:
    print(f"Error: {e}")

# Configure retries
retry_strategy = Retry(
    total=3,
    backoff_factor=1,
    status_forcelist=[429, 500, 502, 503, 504]
)
adapter = HTTPAdapter(max_retries=retry_strategy)
session = requests.Session()
session.mount("https://", adapter)

Examples

Comprehensive Error Handling

import requests
from requests.exceptions import (
    RequestException,
    ConnectionError,
    HTTPError,
    Timeout,
    TooManyRedirects
)

def make_safe_request(url):
    try:
        response = requests.get(url, timeout=(5, 30))  # (connect, read) timeout
        response.raise_for_status()
        return response.json()

    except ConnectionError:
        print("Failed to connect to the server")
    except Timeout:
        print("Request timed out")
    except TooManyRedirects:
        print("Too many redirects")
    except HTTPError as e:
        print(f"HTTP error occurred: {e.response.status_code}")
        if e.response.status_code == 404:
            print("Resource not found")
        elif e.response.status_code == 401:
            print("Authentication required")
        elif e.response.status_code == 403:
            print("Access forbidden")
        elif e.response.status_code >= 500:
            print("Server error - try again later")
    except RequestException as e:
        print(f"An error occurred: {e}")

    return None

# Usage
data = make_safe_request('https://api.example.com/data')
if data:
    print(data)

Implementing Retries with Backoff

import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

def create_session_with_retries(
    retries=3,
    backoff_factor=0.5,
    status_forcelist=(500, 502, 503, 504)
):
    """Create a session with automatic retry logic."""
    session = requests.Session()

    retry_strategy = Retry(
        total=retries,
        read=retries,
        connect=retries,
        backoff_factor=backoff_factor,
        status_forcelist=status_forcelist,
        allowed_methods=["HEAD", "GET", "OPTIONS", "POST"]
    )

    adapter = HTTPAdapter(max_retries=retry_strategy)
    session.mount("http://", adapter)
    session.mount("https://", adapter)

    return session

# Usage
session = create_session_with_retries()

try:
    response = session.get('https://api.example.com/data', timeout=10)
    response.raise_for_status()
    print(response.json())
except requests.exceptions.RequestException as e:
    print(f"Request failed after retries: {e}")
finally:
    session.close()

Custom Retry Decorator

import requests
import time
from functools import wraps

def retry_request(max_retries=3, delay=1, backoff=2):
    """Decorator for retrying requests with exponential backoff."""
    def decorator(func):
        @wraps(func)
        def wrapper(*args, **kwargs):
            retries = 0
            current_delay = delay

            while retries < max_retries:
                try:
                    return func(*args, **kwargs)
                except requests.exceptions.RequestException as e:
                    retries += 1
                    if retries == max_retries:
                        raise
                    print(f"Attempt {retries} failed: {e}")
                    print(f"Retrying in {current_delay} seconds...")
                    time.sleep(current_delay)
                    current_delay *= backoff

            return None
        return wrapper
    return decorator

@retry_request(max_retries=3, delay=1, backoff=2)
def fetch_data(url):
    response = requests.get(url, timeout=10)
    response.raise_for_status()
    return response.json()

# Usage
try:
    data = fetch_data('https://api.example.com/data')
    print(data)
except requests.exceptions.RequestException as e:
    print(f"All retries failed: {e}")

File Uploads and Downloads

Key Concepts

  • Multipart form data is used for file uploads
  • Streaming allows handling large files without loading into memory
  • Progress tracking can be implemented for large transfers
  • Content-Disposition header provides filename information

Common Patterns

import requests

# Upload a file
with open('file.txt', 'rb') as f:
    files = {'file': f}
    response = requests.post(url, files=files)

# Download a file (with closes the connection when done)
with requests.get(url, stream=True) as response:
    with open('downloaded_file', 'wb') as f:
        for chunk in response.iter_content(chunk_size=8192):
            f.write(chunk)

Examples

File Upload

import requests

# Simple file upload
def upload_file(url, file_path):
    with open(file_path, 'rb') as f:
        files = {'file': f}
        response = requests.post(url, files=files)
    return response

# Upload with additional data
def upload_file_with_data(url, file_path, metadata):
    with open(file_path, 'rb') as f:
        files = {'file': (file_path, f, 'application/octet-stream')}
        data = {'metadata': metadata}
        response = requests.post(url, files=files, data=data)
    return response

# Multiple file upload
def upload_multiple_files(url, file_paths):
    files = []
    file_handles = []

    try:
        for path in file_paths:
            f = open(path, 'rb')
            file_handles.append(f)
            files.append(('files', (path, f, 'application/octet-stream')))

        response = requests.post(url, files=files)
        return response
    finally:
        for f in file_handles:
            f.close()

# Usage
response = upload_file('https://api.example.com/upload', 'document.pdf')
print(f"Upload status: {response.status_code}")

File Download with Progress

import requests
from tqdm import tqdm  # pip install tqdm

def download_file(url, filename):
    """Download a file with progress bar."""
    with requests.get(url, stream=True) as response:
        response.raise_for_status()

        # Get total file size from headers
        total_size = int(response.headers.get('content-length', 0))

        with open(filename, 'wb') as f:
            with tqdm(total=total_size, unit='B', unit_scale=True) as pbar:
                for chunk in response.iter_content(chunk_size=8192):
                    if chunk:
                        f.write(chunk)
                        pbar.update(len(chunk))

    print(f"Downloaded: {filename}")

# Simple download without progress
def download_file_simple(url, filename):
    """Download a file without progress tracking."""
    with requests.get(url, stream=True) as response:
        response.raise_for_status()

        with open(filename, 'wb') as f:
            for chunk in response.iter_content(chunk_size=8192):
                if chunk:
                    f.write(chunk)

    return filename

# Usage
download_file('https://example.com/largefile.zip', 'largefile.zip')

Streaming Large Files

import requests

def stream_large_file(url, output_path, chunk_size=1024*1024):
    """Stream download large files in chunks (1MB default)."""
    with requests.get(url, stream=True) as response:
        response.raise_for_status()

        with open(output_path, 'wb') as f:
            downloaded = 0
            for chunk in response.iter_content(chunk_size=chunk_size):
                if chunk:
                    f.write(chunk)
                    downloaded += len(chunk)
                    print(f"Downloaded: {downloaded / (1024*1024):.2f} MB", end='\r')

    print(f"\nComplete: {output_path}")

# Upload large file in chunks
def upload_large_file(url, file_path):
    """Upload a large file using streaming."""
    def file_generator(file_path, chunk_size=8192):
        with open(file_path, 'rb') as f:
            while True:
                chunk = f.read(chunk_size)
                if not chunk:
                    break
                yield chunk

    response = requests.post(url, data=file_generator(file_path))
    return response

Quick Reference

Operation Code Example
GET request requests.get(url)
POST request requests.post(url, data={'key': 'value'})
POST JSON requests.post(url, json={'key': 'value'})
Query parameters requests.get(url, params={'q': 'search'})
Custom headers requests.get(url, headers={'User-Agent': 'App'})
Basic auth requests.get(url, auth=('user', 'pass'))
Bearer token requests.get(url, headers={'Authorization': 'Bearer token'})
Timeout requests.get(url, timeout=10)
Timeout (connect, read) requests.get(url, timeout=(5, 30))
Session session = requests.Session()
Raise for status response.raise_for_status()
JSON response response.json()
Text response response.text
Binary response response.content
Status code response.status_code
Response headers response.headers
Cookies response.cookies
Stream response requests.get(url, stream=True)
File upload requests.post(url, files={'file': open('f.txt', 'rb')})
Disable SSL verify requests.get(url, verify=False)
Custom CA bundle requests.get(url, verify='/path/to/cert.pem')
Proxy requests.get(url, proxies={'https': 'http://proxy:8080'})
Disable redirects requests.get(url, allow_redirects=False) (default is True for GET; HEAD defaults to False)

Common Issues and Solutions

Issue: SSL Certificate Verification Errors

Problem: requests.exceptions.SSLError: certificate verify failed

Solution:

import requests

# Option 1: Specify CA bundle
response = requests.get(url, verify='/path/to/ca-bundle.crt')

# Option 2: Disable verification (NOT recommended for production)
import urllib3
urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning)
response = requests.get(url, verify=False)

# Option 3: Update certifi package
# pip install --upgrade certifi

Issue: Connection Timeout vs Read Timeout

Problem: Need different timeouts for connection and reading

Solution:

import requests

# Tuple format: (connect_timeout, read_timeout)
response = requests.get(url, timeout=(3.05, 27))
# 3.05 seconds to establish connection
# 27 seconds to wait for response

Issue: Handling Rate Limiting (429 Status)

Problem: API returns 429 Too Many Requests

Solution:

import requests
import time
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime

def parse_retry_after(value, default=60):
    """Retry-After may be delay-seconds or an HTTP-date (RFC 9110)."""
    if value is None:
        return default
    try:
        return max(0, int(value))
    except ValueError:
        try:
            dt = parsedate_to_datetime(value)
            return max(0, (dt - datetime.now(timezone.utc)).total_seconds())
        except (TypeError, ValueError):
            return default

def request_with_rate_limit(url, max_retries=5):
    for attempt in range(max_retries):
        response = requests.get(url)

        if response.status_code == 429:
            retry_after = parse_retry_after(response.headers.get('Retry-After'))
            print(f"Rate limited. Waiting {retry_after} seconds...")
            time.sleep(retry_after)
            continue

        return response

    raise Exception("Max retries exceeded due to rate limiting")

Issue: Memory Issues with Large Responses

Problem: Out of memory when downloading large files

Solution:

import requests

# Use streaming to avoid loading entire response into memory
# (with ensures the connection is released when done)
with requests.get(url, stream=True) as response:
    with open('large_file.zip', 'wb') as f:
        for chunk in response.iter_content(chunk_size=8192):
            f.write(chunk)

Issue: JSON Decode Errors

Problem: json.decoder.JSONDecodeError when parsing response

Solution:

import requests

response = requests.get(url)

# Check content type before parsing
if 'application/json' in response.headers.get('Content-Type', ''):
    try:
        data = response.json()
    except ValueError as e:
        print(f"Invalid JSON: {e}")
        print(f"Response text: {response.text[:200]}")
else:
    print(f"Unexpected content type: {response.headers.get('Content-Type')}")

Issue: Character Encoding Problems

Problem: Garbled text in response

Solution:

import requests

response = requests.get(url)

# Check detected encoding
print(f"Detected encoding: {response.encoding}")

# Override encoding if needed
response.encoding = 'utf-8'
text = response.text

# Or use apparent_encoding
response.encoding = response.apparent_encoding
text = response.text

Issue: Proxy Configuration

Problem: Need to route requests through a proxy

Solution:

import requests

proxies = {
    'http': 'http://proxy.example.com:8080',
    'https': 'http://proxy.example.com:8080'
}

# With authentication
proxies = {
    'http': 'http://user:password@proxy.example.com:8080',
    'https': 'http://user:password@proxy.example.com:8080'
}

response = requests.get(url, proxies=proxies)

# Or set environment variables
# export HTTP_PROXY="http://proxy.example.com:8080"
# export HTTPS_PROXY="http://proxy.example.com:8080"

Issue: Handling Redirects

Problem: Need to track or disable redirects

Solution:

import requests

# Redirects are followed by default (allow_redirects=True) for all
# methods except HEAD, which defaults to allow_redirects=False

# Disable redirects
response = requests.get(url, allow_redirects=False)

# Access redirect history
response = requests.get(url)
for redirect in response.history:
    print(f"Redirected: {redirect.status_code} -> {redirect.url}")

print(f"Final URL: {response.url}")

Related Topics

  • aiohttp - Asynchronous HTTP client/server for Python using asyncio
  • httpx - Modern HTTP client with async support and HTTP/2 capabilities
  • urllib3 - Powerful HTTP client library underlying requests
  • REST API Design - Best practices for designing and consuming RESTful APIs
  • OAuth 2.0 and JWT - Deep dive into modern authentication and authorisation
  • Web Scraping with BeautifulSoup - Parsing HTML responses and extracting data