Learn how to use the OpenAI Batch API to cut costs by 50% for bulk processing tasks. Covers setup, request formatting, job creation, monitoring, error handling, and troubleshooting.
This guide covers how to use the OpenAI Batch API to reduce costs for large-scale, asynchronous processing tasks. It is intended for developers who need to send many requests (e.g., classification, summarization, data extraction) without the latency or expense of real-time API calls. You will learn the concepts, setup, request formatting, cost savings, and error handling needed to run batch jobs reliably.
Before you start, make sure you have the following:
The OpenAI Batch API allows you to submit a large number of requests as a single file. The API processes these requests asynchronously, typically within 24 hours. The primary benefit is a 50% cost reduction compared to real-time API calls. This makes it ideal for workloads where immediate responses are not required, such as:
custom_id, the API endpoint (e.g., /v1/chat/completions), the HTTP method (POST), and the request body.batch.According to the official documentation, the Batch API offers a 50% discount on token usage compared to the standard API. This discount applies to both input and output tokens. For example, if a real-time request costs $0.01, the same request via the batch API costs $0.005. This can lead to significant savings for high-volume workloads.
If you haven't already, install the official OpenAI Python library:
pip install openai
Set your API key as an environment variable. This is more secure than hardcoding it in your scripts.
export OPENAI_API_KEY="your-api-key-here"
Alternatively, you can pass the key directly when initializing the client, but this is not recommended for production.
The batch request file must be in JSONL format. Each line is a JSON object representing a single request. The structure is as follows:
{"custom_id": "request-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-4o", "messages": [{"role": "user", "content": "Hello, world!"}]}}
{"custom_id": "request-2", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-4o", "messages": [{"role": "user", "content": "What is the capital of France?"}]}}
custom_id: A string that you define to identify each request. This ID is returned in the output file so you can match results to requests. It must be unique within the file.method: The HTTP method. For the Chat Completions API, this is always POST.url: The API endpoint. For chat completions, use /v1/chat/completions. For the Responses API, use /v1/responses.body: The request body as a JSON object. This is the same payload you would send in a real-time request, including model, messages, max_tokens, temperature, etc.You can include different models or parameters for each request. For example:
{"custom_id": "request-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-4o", "messages": [{"role": "user", "content": "Summarize this article."}], "max_tokens": 100}}
{"custom_id": "request-2", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Translate to French: Hello."}], "temperature": 0.3}}
.jsonl extension.custom_id must be unique across all requests in the file. Duplicate IDs will cause the batch to fail.Once your JSONL file is ready, upload it using the Files API. The purpose must be set to batch.
from openai import OpenAI
client = OpenAI()
batch_file = client.files.create(
file=open("batch_requests.jsonl", "rb"),
purpose="batch"
)
print(batch_file.id) # e.g., "file-abc123"
curl https://api.openai.com/v1/files \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F purpose="batch" \
-F file="@batch_requests.jsonl"
This returns a file object with an id field. Save this ID; you will need it to create the batch job.
With the file ID, you can now create a batch job.
batch_job = client.batches.create(
input_file_id=batch_file.id,
endpoint="/v1/chat/completions",
completion_window="24h"
)
print(batch_job.id) # e.g., "batch-xyz789"
input_file_id: The ID of the uploaded file.endpoint: The API endpoint to use. Must match the url field in your JSONL file. Options are /v1/chat/completions and /v1/responses.completion_window: The maximum time to wait for completion. Currently, only "24h" is supported.curl https://api.openai.com/v1/batches \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input_file_id": "file-abc123",
"endpoint": "/v1/chat/completions",
"completion_window": "24h"
}'
You can poll the batch status to see when it completes.
batch_status = client.batches.retrieve(batch_job.id)
print(batch_status.status)
validating: The input file is being validated.in_progress: The batch is being processed.completed: The batch has finished successfully.failed: The batch failed. Check the errors field for details.expired: The batch took longer than the completion_window.cancelling: A cancellation request is in progress.cancelled: The batch was cancelled.You can set up a webhook to be notified when the batch completes. This is more efficient than polling. To do this, you need to configure a webhook endpoint in your OpenAI dashboard and specify it when creating the batch. The official documentation does not provide a direct parameter for this in the batch creation call, but you can use the Webhooks API to listen for batch completion events.
Once the batch status is completed, you can download the output file.
# Get the output file ID
batch = client.batches.retrieve(batch_job.id)
output_file_id = batch.output_file_id
# Download the file content
file_content = client.files.content(output_file_id).read()
# Save to a file
with open("batch_output.jsonl", "wb") as f:
f.write(file_content)
The output file is also in JSONL format. Each line corresponds to a request and includes the custom_id, the HTTP status code, and the response body.
{"id": "batch_req_abc123", "custom_id": "request-1", "response": {"status_code": 200, "request_id": "req_xyz789", "body": {"id": "chatcmpl-123", "object": "chat.completion", "created": 1712345678, "model": "gpt-4o-2024-05-13", "choices": [{"index": 0, "message": {"role": "assistant", "content": "Hello! How can I help you?"}, "finish_reason": "stop"}], "usage": {"prompt_tokens": 10, "completion_tokens": 7, "total_tokens": 17}}}, "error": null}
{"id": "batch_req_def456", "custom_id": "request-2", "response": {"status_code": 200, "request_id": "req_abc123", "body": {"id": "chatcmpl-456", "object": "chat.completion", "created": 1712345679, "model": "gpt-4o-2024-05-13", "choices": [{"index": 0, "message": {"role": "assistant", "content": "The capital of France is Paris."}, "finish_reason": "stop"}], "usage": {"prompt_tokens": 12, "completion_tokens": 6, "total_tokens": 18}}}, "error": null}
If a request failed, the error field will contain details instead of being null.

Errors can occur at multiple stages: file upload, batch creation, processing, or result download. The official error codes guide provides a comprehensive list of possible errors. Here are the most relevant ones for batch processing:
401 - Invalid Authentication: Your API key is invalid or revoked. Check your key and generate a new one if necessary.401 - Incorrect API key provided: There is a typo or extra space in your API key. Clear your browser cache or generate a new key.401 - You must be a member of an organization to use the API: Your account is not part of an organization. Contact support or ask your organization manager to invite you.401 - IP not authorized: Your request IP does not match the configured IP allowlist. Update your allowlist or send from the correct IP.403 - Country, region, or territory not supported: You are accessing the API from an unsupported location. See the supported countries page for more information.429 - Rate limit reached for requests: You are sending requests too quickly. Pace your requests and implement exponential backoff.429 - You exceeded your current quota, please check your plan and billing details: You have run out of credits or hit your maximum monthly spend. Buy more credits or increase your limits.500 - The server had an error while processing your request: Issue on OpenAI's servers. Retry after a brief wait and check the status page.503 - The engine is currently overloaded, please try again later: Servers are experiencing high traffic. Retry after a brief wait.503 - Slow Down: A sudden increase in your request rate is impacting service reliability. Reduce your request rate to its original level, maintain it for at least 15 minutes, then gradually increase.When using the Python library, you may encounter these exceptions:
APIConnectionError: Issue connecting to OpenAI services. Check your network settings, proxy configuration, SSL certificates, or firewall rules.APITimeoutError: Request timed out. Retry after a brief wait.AuthenticationError: Your API key or token was invalid, expired, or revoked. Check your key and generate a new one if needed.BadRequestError: Your request was malformed or missing required parameters. Check the error message for specifics and review the API documentation.ConflictError: The resource was updated by another request. Retry the update.InternalServerError: Issue on OpenAI's side. Retry after a brief wait and contact support if it persists.NotFoundError: Requested resource does not exist. Ensure you are using the correct resource identifier.PermissionDeniedError: You don't have access to the requested resource. Check your API key, organization ID, and resource ID.RateLimitError: You have hit your assigned rate limit. Pace your requests and implement exponential backoff.UnprocessableEntityError: Unable to process the request despite the format being correct. Try the request again.The official documentation recommends handling errors programmatically. Here is a Python example:
import openai
from openai import OpenAI
client = OpenAI()
try:
# Make your OpenAI API request here
response = client.responses.create(
model="gpt-4o",
input="Hello world"
)
except openai.APIError as e:
# Handle API error here, e.g. retry or log
print(f"OpenAI API returned an API Error: {e}")
pass
except openai.APIConnectionError as e:
# Handle connection error here
print(f"Failed to connect to OpenAI API: {e}")
pass
except openai.RateLimitError as e:
# Handle rate limit error (we recommend using exponential backoff)
print(f"OpenAI API request exceeded rate limit: {e}")
pass
If a batch job fails, you can check the errors field on the batch object. This field contains an array of error objects, each with a code and message. Common batch errors include:
invalid_json: The input file contains malformed JSON.invalid_custom_id: A custom_id is missing or duplicated.invalid_endpoint: The endpoint in the batch creation does not match the url field in the JSONL file.rate_limit_exceeded: The batch exceeded the rate limit for the endpoint.timeout: The batch took longer than the completion_window.Batch API costs are 50% less than real-time API costs. For example, if the real-time price for gpt-4o is $5.00 per 1M input tokens, the batch price is $2.50 per 1M input tokens. Output tokens are similarly discounted. You can find the exact pricing for each model on the pricing page.
For failed requests within a batch, you cannot retry individual requests. Instead, you must create a new batch with only the failed requests. To do this, parse the output file, identify requests where the error field is not null, extract their custom_id and original body, and create a new JSONL file with those requests. Then upload and create a new batch.
You can also use the Batch API with the Responses API endpoint (/v1/responses). The process is identical, except the url field in your JSONL file should be /v1/responses, and the body should follow the Responses API format.

If your batch job stays in the validating status for more than a few minutes, it may indicate an issue with the input file. Check the file for malformed JSON or duplicate custom_id values. You can also try uploading a smaller file to test.
This means one or more lines in your JSONL file are not valid JSON. Use a JSON validator to check each line. Common issues include trailing commas, unescaped quotes, or missing brackets.
Your batch job exceeded the rate limit for the endpoint. This is rare because batch jobs have their own rate limits, but it can happen if you have many concurrent batches. Wait for some batches to complete before submitting new ones.
If only some requests failed, check the error field in the output file. Common errors include:
invalid_model: The model specified in the request body is not available or does not exist.invalid_messages: The messages array is malformed or missing required fields.context_length_exceeded: The request exceeded the model's context window.To handle these, you can modify the request body and resubmit the failed requests in a new batch.
If you encounter 401 or AuthenticationError, follow these steps:
If you encounter APIConnectionError or APITimeoutError:
If an error persists, contact OpenAI support via chat with the following information:
You can also post in the OpenAI Community Forum, but be sure to omit any sensitive information.
Now that you understand the basics of the Batch API, you can explore more advanced topics:
gpt-4o-mini vs. gpt-4o) to balance cost and quality for your specific use case./v1/responses endpoint.Learn how to write prompts for ChatGPT that reliably generate working code. Covers core concepts, setup, and specific patterns for different coding tasks.
Learn how to use the OpenAI Vision API to analyze images with GPT models. Covers setup, sending images via URL, Base64, or file ID, controlling detail levels, cost calculation, and troubleshooting.
Learn how to stream ChatGPT responses in web applications using the OpenAI API. This guide covers setup, implementation in JavaScript and Python, event handling, and troubleshooting for real-time text generation.
Learn how to use OpenAI's text embedding models for semantic search, clustering, recommendations, and classification. This guide covers concepts, setup, API usage, dimension reduction, and practical tips from official documentation and community experience.
Learn how to handle OpenAI API rate limits (429) and quota errors. Covers error types, exponential backoff, Python library exceptions, and troubleshooting steps for production applications.
Learn how to use OpenAI Structured Outputs to guarantee JSON schema conformance from GPT models. Covers function tool definitions, strict mode, tool call handling, and best practices for both Chat Completions and Responses APIs.
Workflows from the Neura Market marketplace related to this ChatGPT resource