How to Use the OpenAI API: Prompt Design, Streaming, and Cost Optimization

Prerequisites

  • OpenAI API key
  • Python 3.11+ with openai package installed
  • Basic understanding of HTTP APIs

Step 1: Install and Authenticate

pip install openai
import os
from openai import OpenAI

client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
# Or use a .env file with python-dotenv

Never hardcode API keys. Use environment variables, a secrets manager, or your framework’s credentials system.

Step 2: Prompt Design Principles

Be Specific and Give Context

# ❌ Vague
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Write a function to sort data."}]
)

# ✅ Specific
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{
        "role": "user",
        "content": """Write a Python function that sorts a list of dictionaries
by a given key. Requirements:
- Use type hints
- Handle missing keys gracefully
- Sort in descending order by default
- Include a docstring with examples"""
    }]
)

Use System Messages for Behaviour

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {
            "role": "system",
            "content": "You are an expert PostgreSQL DBA. Answer questions about indexing, query optimization, and performance tuning. When providing SQL, include EXPLAIN output expectations."
        },
        {
            "role": "user",
            "content": "Why is my query doing a sequential scan on a table with 10 million rows?"
        }
    ]
)

Chain of Thought Prompting

For complex reasoning, ask the model to think step by step:

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{
        "role": "user",
        "content": """Debug this error and explain your reasoning step by step:

NoMethodError: undefined method name' for nil:NilClass app/controllers/users_controller.rb:15:in show'


Step 1: Identify the error type
Step 2: Locate the problematic line
Step 3: Determine why the object is nil
Step 4: Propose a fix with code"""
    }]
)

Step 3: Streaming Responses

For long generations, stream tokens to show progress:

stream = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Explain the CAP theorem in detail."}],
    stream=True
)

for chunk in stream:
    content = chunk.choices[0].delta.content
    if content:
        print(content, end="", flush=True)

With FastAPI:

from fastapi import FastAPI
from fastapi.responses import StreamingResponse
import json

app = FastAPI()

@app.post("/chat")
async def chat(prompt: str):
    async def generate():
        stream = client.chat.completions.create(
            model="gpt-4o",
            messages=[{"role": "user", "content": prompt}],
            stream=True
        )
        for chunk in stream:
            content = chunk.choices[0].delta.content
            if content:
                yield f"data: {json.dumps({'content': content})}\n\n"
        yield "data: [DONE]\n\n"

    return StreamingResponse(generate(), media_type="text/event-stream")

Step 4: Error Handling and Retries

The OpenAI API can return rate limits, timeouts, and server errors. Handle them gracefully:

import time
from openai import RateLimitError, APIError, APITimeoutError

MAX_RETRIES = 3
BASE_DELAY = 1  # seconds

def chat_with_retry(messages, retries=MAX_RETRIES):
    for attempt in range(retries):
        try:
            response = client.chat.completions.create(
                model="gpt-4o",
                messages=messages,
                timeout=30
            )
            return response.choices[0].message.content
        except RateLimitError:
            if attempt == retries - 1:
                raise
            delay = BASE_DELAY * (2 ** attempt)  # Exponential backoff
            time.sleep(delay)
        except APITimeoutError:
            if attempt == retries - 1:
                raise
            time.sleep(BASE_DELAY)
        except APIError as e:
            if e.status_code >= 500:  # Server error, retry
                if attempt == retries - 1:
                    raise
                time.sleep(BASE_DELAY * (2 ** attempt))
            else:  # Client error (400, 401), don't retry
                raise

Step 5: Cost Optimization

Use Smaller Models When Possible

ModelCost (per 1M tokens)Best for
GPT-4o$2.50 input / $10 outputComplex reasoning
GPT-4o-mini$0.15 input / $0.60 outputMost tasks
GPT-3.5 Turbo$0.50 input / $1.50 outputSimple tasks

Default to GPT-4o-mini unless you need GPT-4o’s reasoning.

Cache Responses

import hashlib
import json

cache = {}

def cached_chat(messages):
    cache_key = hashlib.md5(json.dumps(messages).encode()).hexdigest()
    if cache_key in cache:
        return cache[cache_key]

    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=messages
    )
    content = response.choices[0].message.content
    cache[cache_key] = content
    return content

In production, use Redis or a database for caching.

Limit Token Usage

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=messages,
    max_tokens=500,           # Limit response length
    temperature=0.3            # Lower = more deterministic, often shorter
)

print(f"Tokens used: {response.usage.total_tokens}")

Step 6: Rate Limit Handling

OpenAI has tiered rate limits. Check your current usage:

# Track rate limits from response headers
response = client.chat.completions.create(model="gpt-4o-mini", messages=messages)

# For sync API, use the openai library's built-in retry:
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["OPENAI_API_KEY"],
    max_retries=3  # Auto-retry on rate limit and server errors
)

Verification

# Quick test
response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Say 'Hello, BackendMedic!' and nothing else."}]
)
print(response.choices[0].message.content)
# Hello, BackendMedic!

print(f"Cost: {response.usage.total_tokens} tokens")
# Cost: ~10 tokens (fraction of a cent)

Summary

  • Be specific in prompts — context and constraints improve output quality
  • Use system messages to set behaviour consistently
  • Stream responses for better UX on long generations
  • Implement exponential backoff for rate limits and server errors
  • Default to GPT-4o-mini for cost efficiency; reserve GPT-4o for complex tasks
  • Cache responses for repeated prompts
  • Never hardcode API keys

References


Advertisement