How to Use the OpenAI API: Prompt Design, Streaming, and Cost Optimization
Prerequisites
- OpenAI API key
- Python 3.11+ with
openaipackage installed - Basic understanding of HTTP APIs
Step 1: Install and Authenticate
pip install openai
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
# Or use a .env file with python-dotenv
Never hardcode API keys. Use environment variables, a secrets manager, or your framework’s credentials system.
Step 2: Prompt Design Principles
Be Specific and Give Context
# ❌ Vague
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Write a function to sort data."}]
)
# ✅ Specific
response = client.chat.completions.create(
model="gpt-4o",
messages=[{
"role": "user",
"content": """Write a Python function that sorts a list of dictionaries
by a given key. Requirements:
- Use type hints
- Handle missing keys gracefully
- Sort in descending order by default
- Include a docstring with examples"""
}]
)
Use System Messages for Behaviour
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "system",
"content": "You are an expert PostgreSQL DBA. Answer questions about indexing, query optimization, and performance tuning. When providing SQL, include EXPLAIN output expectations."
},
{
"role": "user",
"content": "Why is my query doing a sequential scan on a table with 10 million rows?"
}
]
)
Chain of Thought Prompting
For complex reasoning, ask the model to think step by step:
response = client.chat.completions.create(
model="gpt-4o",
messages=[{
"role": "user",
"content": """Debug this error and explain your reasoning step by step:
NoMethodError: undefined method name' for nil:NilClass app/controllers/users_controller.rb:15:in show'
Step 1: Identify the error type
Step 2: Locate the problematic line
Step 3: Determine why the object is nil
Step 4: Propose a fix with code"""
}]
)
Step 3: Streaming Responses
For long generations, stream tokens to show progress:
stream = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Explain the CAP theorem in detail."}],
stream=True
)
for chunk in stream:
content = chunk.choices[0].delta.content
if content:
print(content, end="", flush=True)
With FastAPI:
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
import json
app = FastAPI()
@app.post("/chat")
async def chat(prompt: str):
async def generate():
stream = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
stream=True
)
for chunk in stream:
content = chunk.choices[0].delta.content
if content:
yield f"data: {json.dumps({'content': content})}\n\n"
yield "data: [DONE]\n\n"
return StreamingResponse(generate(), media_type="text/event-stream")
Step 4: Error Handling and Retries
The OpenAI API can return rate limits, timeouts, and server errors. Handle them gracefully:
import time
from openai import RateLimitError, APIError, APITimeoutError
MAX_RETRIES = 3
BASE_DELAY = 1 # seconds
def chat_with_retry(messages, retries=MAX_RETRIES):
for attempt in range(retries):
try:
response = client.chat.completions.create(
model="gpt-4o",
messages=messages,
timeout=30
)
return response.choices[0].message.content
except RateLimitError:
if attempt == retries - 1:
raise
delay = BASE_DELAY * (2 ** attempt) # Exponential backoff
time.sleep(delay)
except APITimeoutError:
if attempt == retries - 1:
raise
time.sleep(BASE_DELAY)
except APIError as e:
if e.status_code >= 500: # Server error, retry
if attempt == retries - 1:
raise
time.sleep(BASE_DELAY * (2 ** attempt))
else: # Client error (400, 401), don't retry
raise
Step 5: Cost Optimization
Use Smaller Models When Possible
| Model | Cost (per 1M tokens) | Best for |
|---|---|---|
| GPT-4o | $2.50 input / $10 output | Complex reasoning |
| GPT-4o-mini | $0.15 input / $0.60 output | Most tasks |
| GPT-3.5 Turbo | $0.50 input / $1.50 output | Simple tasks |
Default to GPT-4o-mini unless you need GPT-4o’s reasoning.
Cache Responses
import hashlib
import json
cache = {}
def cached_chat(messages):
cache_key = hashlib.md5(json.dumps(messages).encode()).hexdigest()
if cache_key in cache:
return cache[cache_key]
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=messages
)
content = response.choices[0].message.content
cache[cache_key] = content
return content
In production, use Redis or a database for caching.
Limit Token Usage
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
max_tokens=500, # Limit response length
temperature=0.3 # Lower = more deterministic, often shorter
)
print(f"Tokens used: {response.usage.total_tokens}")
Step 6: Rate Limit Handling
OpenAI has tiered rate limits. Check your current usage:
# Track rate limits from response headers
response = client.chat.completions.create(model="gpt-4o-mini", messages=messages)
# For sync API, use the openai library's built-in retry:
from openai import OpenAI
client = OpenAI(
api_key=os.environ["OPENAI_API_KEY"],
max_retries=3 # Auto-retry on rate limit and server errors
)
Verification
# Quick test
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Say 'Hello, BackendMedic!' and nothing else."}]
)
print(response.choices[0].message.content)
# Hello, BackendMedic!
print(f"Cost: {response.usage.total_tokens} tokens")
# Cost: ~10 tokens (fraction of a cent)
Summary
- Be specific in prompts — context and constraints improve output quality
- Use system messages to set behaviour consistently
- Stream responses for better UX on long generations
- Implement exponential backoff for rate limits and server errors
- Default to GPT-4o-mini for cost efficiency; reserve GPT-4o for complex tasks
- Cache responses for repeated prompts
- Never hardcode API keys
References
Advertisement