The threading module in Python facilitates the execution of concurrent IO-bound tasks. In Python's execution model, a Global Interpreter Lock (GIL) ensures only one thread executes Python byetcode at a time. Consequently, multi-threading is most effective for operations involving network calls, file I/O, or other blocking tasks, where threads can yield control during wait periods, simulating parallelism.
Thread Creation and Execution
A simple demonstration uses a Flask server endpoint that introduces a delay.
# Flask Server
from flask import Flask
import time
app = Flask(__name__)
@app.route('/delay/<int:seconds>')
def delayed_response(seconds):
time.sleep(seconds)
return f'Response after {seconds} seconds'
if __name__ == '__main__':
app.run()
The client code compares sequential and concurrent execution.
import threading
import time
import requests
BASE_URL = 'http://localhost:5000/delay/'
def fetch_url(url_path):
resp = requests.get(BASE_URL + url_path)
print(resp.text)
def concurrent_requests():
thread_list = []
for i in range(10):
t = threading.Thread(target=fetch_url, args=('2',))
thread_list.append(t)
t.start()
for th in thread_list:
th.join()
def sequential_requests():
for i in range(5):
fetch_url('2')
if __name__ == '__main__':
start = time.time()
concurrent_requests()
print(f'Concurrent duration: {time.time() - start:.2f}s')
start = time.time()
sequential_requests()
print(f'Sequential duration: {time.time() - start:.2f}s')
The Thread constructor accepts a target function and its arguments. The start() method initiates the thread's activity, while join() blocks the calling thread until the target thread completes.
Daemon Threads
Daemon threads terminate when the main program exits. They are suitable for background tasks like monitoring.
import threading
import time
def background_task():
while True:
print(f'Background tick: {time.time()}')
time.sleep(5)
def main_operation():
daemon_thread = threading.Thread(target=background_task, daemon=True)
daemon_thread.start()
# Main work here
time.sleep(15)
print('Main operation finished')
if __name__ == '__main__':
main_operation()
Thread Attributes and Control
Thread objects provide metaadta and control methods.
worker = threading.Thread(target=fetch_url, args=('1',), name='FetchWorker')
print(worker.name) # FetchWorker
print(worker.is_alive()) # False before start()
worker.start()
print(worker.is_alive()) # True
print(worker.isDaemon()) # False
worker.setDaemon(True) # Convert to daemon
Threading Module Utilities
Useful functions for introspection.
print(f'Active threads: {threading.active_count()}')
print(f'Current thread: {threading.current_thread().name}')
for t in threading.enumerate():
print(f'Thread: {t.name}')
Capturing Exceptions and Results
Standard Thread objects do not propagate exceptions or return values to the main thread. Subclassing Thread allows custom handling.
import threading
import traceback
class ResultThread(threading.Thread):
def __init__(self, func, *args, **kwargs):
super().__init__()
self.function = func
self.arguments = args
self.keyword_args = kwargs
self.output = None
self.error = None
self.stack_trace = None
def run(self):
try:
self.output = self.function(*self.arguments, **self.keyword_args)
except Exception as e:
self.error = e
self.stack_trace = traceback.format_exc()
def get_result(self):
if self.error:
raise self.error
return self.output
# Usage
def task_that_may_fail(x):
if x == 0:
raise ValueError('Invalid input')
return 10 / x
th1 = ResultThread(task_that_may_fail, 5)
th2 = ResultThread(task_that_may_fail, 0)
th1.start()
th2.start()
th1.join()
th2.join()
print(th1.get_result()) # 2.0
print(th2.error) # ValueError instance
print(th2.stack_trace) # Traceback string
Synchronization with Locks
Shared resource access requires synchronization to prevent race conditions.
shared_counter = 0
counter_lock = threading.Lock()
def unsafe_increment():
global shared_counter
for _ in range(100000):
shared_counter += 1
def safe_increment():
global shared_counter
for _ in range(100000):
counter_lock.acquire()
shared_counter += 1
counter_lock.release()
def test_lock():
global shared_counter
shared_counter = 0
t1 = threading.Thread(target=safe_increment)
t2 = threading.Thread(target=safe_increment)
t1.start()
t2.start()
t1.join()
t2.join()
print(f'Final counter (with lock): {shared_counter}')
The acquire(blocking=True, timeout=None) method attempts to obtain the lock. release() frees it. Use locked() to check status.
Thread Pools with concurrent.futures
The ThreadPoolExecutor manages a pool of worker threads.
from concurrent.futures import ThreadPoolExecutor, as_completed
import requests
def query_api(endpoint):
resp = requests.get(f'https://api.example.com/{endpoint}')
return resp.json()
with ThreadPoolExecutor(max_workers=4) as pool:
futures = [pool.submit(query_api, f'item/{i}') for i in range(10)]
for future in as_completed(futures):
try:
data = future.result(timeout=5)
print(f'Received: {data}')
except Exception as exc:
print(f'Task generated an exception: {exc}')
The Future object provides state and result inspection.
result(timeout=None): Returns the callable's return value or re-raises any exception.done(): ReturnsTrueif the task finished (successfully or with an exception).exception(timeout=None): Returns the expection raised, orNone.cancel(): Attempts to cancel the task if not running.cancelled(): Checks if the task was cancelled.running(): Checks if the task is currently executing.
Determining Optimal Thread Pool Size
Identifying the ideal number of threads for a pool involves benchmarking. For IO-bound workloads, start with a small pool and incrementally increase size while monitoring system metrics (CPU, memory, I/O wait) and total execution time. The goal is to find the point where adding more threads yields diminishing returns or degrades performance due to increased context-switching overhead.