AsyncIO & High-Concurrency Python for AI Services
Build high-throughput, non-blocking AI pipelines using Python AsyncIO, TaskGroups, Semaphores for LLM rate-limit management, and async streaming.
Key Takeaways
- LLM API calls are I/O bound (waiting 1–5 seconds for token generation); synchronous Python wastes CPU cycles blocking on single requests
- AsyncIO allows a single Python process to handle thousands of concurrent LLM API requests simultaneously on an event loop
- `asyncio.Semaphore` strictly throttles concurrent outgoing requests to prevent HTTP 429 Rate Limit errors from API providers
- Python 3.11+ `asyncio.TaskGroup` guarantees structured concurrency, ensuring no leaked background coroutines if an exception occurs
The Diagnostic Context
When processing 1,000 customer documents through an LLM, making synchronous calls that take 2 seconds each takes over 33 minutes. With AsyncIO and rate-limited concurrency, the exact same workload completes in under 30 seconds.
The Core Technique
Synchronous vs Asynchronous LLM Processing
gantt
title Synchronous vs AsyncIO Concurrency (10 LLM Requests)
dateFormat X
axisFormat %s sec
section Synchronous (Blocking)
Request 1 (2s) :0, 2
Request 2 (2s) :2, 4
Request 3 (2s) :4, 6
Request 4 (2s) :6, 8
Total 20s :crit, 8, 20
section AsyncIO (Concurrent with Concurrency = 5)
Req 1-5 Parallel :active, 0, 2
Req 6-10 Parallel:active, 2, 4
Finished in 4s :done, 4, 4
Structured Concurrency with CODE / PROMPTasyncio.TaskGroup
& CODE / PROMPTSemaphore
asyncio.TaskGroup
Semaphore
Below is a production-grade async batch processor with concurrency throttling:
import asyncio
from httpx import AsyncClient
async def process_document_with_ai(client: AsyncClient, sem: asyncio.Semaphore, doc_id: str, text: str) -> dict:
async with sem: # Throttles max concurrent requests to 10
# Simulating non-blocking async HTTP call to LLM API:
response = await client.post(
"https://api.openai.com/v1/chat/completions",
json={"model": "gpt-4o-mini", "messages": [{"role": "user", "content": f"Summarize: {text}"}]},
headers={"Authorization": "Bearer $OPENAI_API_KEY"},
timeout=30.0
)
data = response.json()
return {"doc_id": doc_id, "summary": data["choices"][0]["message"]["content"]}
async def batch_process_all_documents(documents: list[tuple[str, str]]) -> list[dict]:
semaphore = asyncio.Semaphore(10) # Max 10 concurrent calls
results = []
async with AsyncClient() as client:
async with asyncio.TaskGroup() as tg:
tasks = [
tg.create_task(process_document_with_ai(client, semaphore, doc_id, text))
for doc_id, text in documents
]
# All tasks completed cleanly:
results = [t.result() for t in tasks]
return results
Try This Right Now
Run an AsyncIO experiment: Write a script that uses `asyncio.gather()` to fetch 5 mock endpoints concurrently with `asyncio.sleep(1)`. Compare total elapsed time against synchronous `time.sleep(1)` executed in a standard `for` loop.
Tip: Knowledge only becomes capability once you run the prompt yourself.