
摘要本文記錄 GPT-6 Luna 模型通過 Batch API 進行批量異步推理的完整接入過程涵蓋模型能力與規(guī)格、開發(fā)環(huán)境搭建、批量任務提交與輪詢、推理強度配置、多模態(tài)請求、結構化輸出、計費規(guī)則與成本估算、任務管理與錯誤處理并附完整批量文本分類示例幫助讀者快速掌握 Batch API 的高效用法。目錄① 模型能力與規(guī)格能力一覽模型規(guī)格② 開發(fā)環(huán)境搭建前置要求安裝 SDK初始化客戶端Batch API 限流③ Batch API 基本流程準備批量請求文件上傳文件并創(chuàng)建批量任務④ 輪詢任務狀態(tài)與獲取結果輪詢狀態(tài)下載并解析結果⑤ 推理強度配置推理強度選擇參考⑥ 多模態(tài)批量請求⑦ 結構化輸出⑧ 計費規(guī)則與成本估算Batch 定價短上下文 ≤ 272K tokens按 Standard 50% 計長上下文階梯單請求輸入 272K tokens成本估算示例提示詞緩存⑨ 任務管理與錯誤處理取消未完成的任務列出歷史任務常見錯誤排查結果校驗⑩ 完整示例批量文本分類參考文檔記錄 GPT-6 Luna 模型通過 Batch API 進行批量異步推理的接入過程涵蓋環(huán)境準備、批量任務提交、結果輪詢、參數(shù)調優(yōu)等場景的代碼示例。① 模型能力與規(guī)格GPT-6 Luna 是 GPT-6 系列中面向高吞吐量任務的模型支持文本和圖像輸入、文本輸出具備可配置的推理強度和工具調用能力。能力一覽能力說明多模態(tài)輸入文本、圖像、文件PDF推理控制reasoning effort: none / low / medium / high / xhigh / max工具調用Function Calling、Web Search、File Search、Computer Use結構化輸出JSON Schema 約束提示詞緩存自動緩存最低 1024 token 前綴Batch API異步批量處理50% 折扣模型規(guī)格項目參數(shù)模型 IDgpt-6-luna上下文窗口1,050,000 tokens最大輸出128,000 tokens輸入模態(tài)文本、圖像、文件輸出模態(tài)文本知識截止2026 年 5 月 18 日默認推理強度medium② 開發(fā)環(huán)境搭建前置要求Python 3.9OpenAI API Keypip安裝 SDKpip install openai初始化客戶端import os from openai import OpenAI client OpenAI( api_keyos.environ.get(OPENAI_API_KEY) )Batch API 限流TierRPMTPMBatch 隊列上限Free不支持——Tier 1500500,0005,000,000Tier 25,0002,000,00020,000,000Tier 35,0004,000,00040,000,000Tier 410,00010,000,0001,000,000,000Tier 530,000180,000,00015,000,000,000③ Batch API 基本流程Batch API 用于異步處理大量請求按 Standard 費率的50%計費24 小時內完成。流程分三步提交 → 輪詢 → 取結果。準備批量請求文件Batch API 要求先上傳一個 JSONL 文件每行一個請求import json 構造批量請求數(shù)據(jù) requests [ { custom_id: task-001, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: user, content: 請用一句話總結人工智能在醫(yī)療領域的應用} ], max_completion_tokens: 200 } }, { custom_id: task-002, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: user, content: 請用一句話總結區(qū)塊鏈技術的核心思想} ], max_completion_tokens: 200 } }, { custom_id: task-003, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: user, content: 請用一句話總結邊緣計算與云計算的區(qū)別} ], max_completion_tokens: 200 } } ] 寫入 JSONL 文件 with open(batch_requests.jsonl, w, encodingutf-8) as f: for req in requests: f.write(json.dumps(req, ensure_asciiFalse) \n) print(f已生成 {len(requests)} 條請求)上傳文件并創(chuàng)建批量任務# 1. 上傳 JSONL 文件 batch_file client.files.create( fileopen(batch_requests.jsonl, rb), purposebatch ) 2. 創(chuàng)建批量任務 batch_job client.batches.create( input_file_idbatch_file.id, endpoint/v1/chat/completions, completion_window24h ) print(f批量任務 ID: {batch_job.id}) print(f狀態(tài): {batch_job.status})④ 輪詢任務狀態(tài)與獲取結果輪詢狀態(tài)import time batch_id batch_job.id while True: batch client.batches.retrieve(batch_id) print(f狀態(tài): {batch.status} | f已完成: {batch.request_counts.completed} | f失敗: {batch.request_counts.failed} | f總計: {batch.request_counts.total}) if batch.status in (completed, failed, cancelled, expired): break time.sleep(30) # 每 30 秒檢查一次下載并解析結果if batch.status completed: # 下載結果文件 result_content client.files.content(batch.output_file_id) result_text result_content.text # 逐行解析 JSONL for line in result_text.strip().split(\n): result json.loads(line) custom_id result[custom_id] response result[response] if response[status_code] 200: content response[body][choices][0][message][content] print(f[{custom_id}] {content}) else: print(f[{custom_id}] 錯誤: {response}) # 如果有失敗記錄下載錯誤文件 if batch.error_file_id: error_content client.files.content(batch.error_file_id) print(f錯誤詳情:\n{error_content.text})Batch 輸入和結果文件保留30 天過期自動刪除。⑤ 推理強度配置GPT-6 Luna 支持 6 檔推理強度Batch 請求中通過reasoning_effort參數(shù)控制# 構造不同推理強度的請求 requests [ { custom_id: ftask-{i:03d}, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: user, content: f分析以下代碼的時間復雜度并說明理由{code_snippet}} ], reasoning_effort: effort, # none / low / medium / high / xhigh / max max_completion_tokens: 1024 } } for i, (effort, code_snippet) in enumerate([ (none, def f(n): return n * 2), (low, def f(n): return sum(range(n))), (medium, def f(n): return [i*j for i in range(n) for j in range(n)]), (high, def f(n): return sorted([randint(0,n) for _ in range(n*n)])), ]) ] with open(batch_reasoning.jsonl, w, encodingutf-8) as f: for req in requests: f.write(json.dumps(req, ensure_asciiFalse) \n)推理強度選擇參考effort適用場景token 消耗none簡單分類、格式轉換、提取最低low基礎問答、短文本摘要低medium默認值通用任務中high代碼分析、多步推理高xhigh復雜邏輯、數(shù)學證明較高max最深推理耗時最長最高注意Chat Completions API 中使用 Function Calling 時reasoning_effort只能設為none。需要同時使用工具和推理時改用 Responses API。⑥ 多模態(tài)批量請求Batch 請求同樣支持圖像輸入將圖片轉為 base64 后放入消息體import base64 def image_to_data_url(image_path): with open(image_path, rb) as f: b64 base64.b64encode(f.read()).decode(utf-8) return fdata:image/jpeg;base64,{b64} 批量圖片分類請求 image_files [img1.jpg, img2.jpg, img3.jpg] requests [] for i, img_path in enumerate(image_files): data_url image_to_data_url(img_path) requests.append({ custom_id: fimg-{i:03d}, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [{ role: user, content: [ {type: text, text: 請用 JSON 格式輸出這張圖片的類別和置信度格式: {category: ..., confidence: 0.0}}, {type: image_url, image_url: {url: data_url}} ] }], max_completion_tokens: 200, response_format: {type: json_object} } }) with open(batch_vision.jsonl, w, encodingutf-8) as f: for req in requests: f.write(json.dumps(req, ensure_asciiFalse) \n)⑦ 結構化輸出通過response_format約束輸出為 JSON Schemarequest { custom_id: extract-001, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: user, content: 從以下文本提取人名、日期和金額2026年9月22日張三向李四轉賬5000元。} ], max_completion_tokens: 256, response_format: { type: json_schema, json_schema: { name: extraction, strict: True, schema: { type: object, properties: { people: {type: array, items: {type: string}}, date: {type: string}, amount: {type: number} }, required: [people, date, amount] } } } } }⑧ 計費規(guī)則與成本估算Batch 定價短上下文 ≤ 272K tokens按 Standard 50% 計計費項Batch 單價/百萬 tokens輸入$0.05緩存命中輸入$0.005緩存寫入$0.0625輸出$0.25長上下文階梯單請求輸入 272K tokens計費項Batch 長上下文單價輸入$0.10緩存命中輸入$0.01緩存寫入$0.125輸出$0.375階梯按單次請求的輸入 token 數(shù)判定不是 batch 內所有請求的合計。成本估算示例假設 1000 條請求每條輸入 2000 tokens、輸出 300 tokens無緩存num_requests 1000 input_tokens_per_req 2000 output_tokens_per_req 300 total_input num_requests * input_tokens_per_req # 2,000,000 total_output num_requests * output_tokens_per_req # 300,000 Batch 短上下文 input_cost (total_input / 1_000_000) * 0.05 # $0.10 output_cost (total_output / 1_000_000) * 0.25 # $0.075 total_cost input_cost output_cost print(f輸入費用: ${input_cost:.4f}) print(f輸出費用: ${output_cost:.4f}) print(f總費用: ${total_cost:.4f}) 輸入費用: $0.1000 輸出費用: $0.0750 總費用: $0.1750提示詞緩存緩存自動生效無需改代碼。命中條件前綴 ≥ 1024 tokensTTL 5–10 分鐘最長 1 小時緩存命中按 $0.005/M 計費Batch 短上下文適合重復使用相同系統(tǒng)提示詞的場景# 所有請求共享相同的系統(tǒng) prompt長前綴只有 user 內容不同 system_prompt 你是一個專業(yè)的文本分類助手。請將輸入文本分類到以下類別之一科技、財經(jīng)、體育、娛樂、教育、健康。只輸出類別名稱不要解釋。 * 20 # 確保超過 1024 tokens requests [ { custom_id: fcls-{i:04d}, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: system, content: system_prompt}, {role: user, content: text} ], max_completion_tokens: 10, reasoning_effort: none } } for i, text in enumerate(texts_to_classify) ] 首次請求寫入緩存后續(xù)命中緩存輸入費用降至 1/10⑨ 任務管理與錯誤處理取消未完成的任務# 取消批量任務 cancelled client.batches.cancel(batch_id) print(f取消后狀態(tài): {cancelled.status})列出歷史任務# 列出最近的批量任務 batches client.batches.list(limit10) for b in batches.data: print(fID: {b.id} | 狀態(tài): {b.status} | f完成: {b.request_counts.completed}/{b.request_counts.total} | f創(chuàng)建時間: {b.created_at})常見錯誤排查錯誤原因解決方案invalid_api_keyAPI Key 錯誤檢查環(huán)境變量file_too_largeJSONL 文件超限拆分為多個 batchrate_limit_exceeded超出 Batch 隊列上限升級 tier 或分批提交batch_expired超過 24h 窗口未完成檢查隊列負載減少單次請求量invalid_request請求體格式錯誤校驗 JSONL 每行的body結構結果校驗# 下載結果后逐條校驗 results [] for line in result_text.strip().split(\n): entry json.loads(line) if entry[response][status_code] 200: results.append(entry) else: # 記錄失敗請求后續(xù)重試 print(f失敗: {entry[custom_id]} - {entry[response]}) 校驗 JSON 結構化輸出 for r in results: try: data json.loads(r[response][body][choices][0][message][content]) except json.JSONDecodeError: print(fJSON 解析失敗: {r[custom_id]})⑩ 完整示例批量文本分類將以上步驟整合為一個完整的批量分類流程import os import json import time from openai import OpenAI client OpenAI(api_keyos.environ.get(OPENAI_API_KEY)) 1. 準備數(shù)據(jù) texts [ 蘋果發(fā)布新款 MacBook Pro搭載 M5 芯片, 美聯(lián)儲宣布降息 50 個基點, 中國隊獲得乒乓球世錦賽團體冠軍, 某明星宣布退出娛樂圈, 教育部發(fā)布新一輪課程改革方案, 研究發(fā)現(xiàn)每天步行 8000 步可顯著降低心血管風險, ] system_prompt 你是一個新聞分類助手。請將輸入文本分類到以下類別之一科技、財經(jīng)、體育、娛樂、教育、健康。只輸出類別名稱。 2. 構造 JSONL requests [ { custom_id: fnews-{i:03d}, method: POST, url: /v1/chat/completions, body: { model: gpt-6-luna, messages: [ {role: system, content: system_prompt}, {role: user, content: text} ], max_completion_tokens: 10, reasoning_effort: none } } for i, text in enumerate(texts) ] with open(classify_batch.jsonl, w, encodingutf-8) as f: for req in requests: f.write(json.dumps(req, ensure_asciiFalse) \n) 3. 上傳 創(chuàng)建任務 batch_file client.files.create( fileopen(classify_batch.jsonl, rb), purposebatch ) batch_job client.batches.create( input_file_idbatch_file.id, endpoint/v1/chat/completions, completion_window24h ) print(f任務已提交: {batch_job.id}) 4. 輪詢 while True: batch client.batches.retrieve(batch_job.id) print(f[{batch.status}] {batch.request_counts.completed}/{batch.request_counts.total}) if batch.status in (completed, failed, cancelled, expired): break time.sleep(15) 5. 輸出結果 if batch.status completed: result_text client.files.content(batch.output_file_id).text for line in result_text.strip().split(\n): entry json.loads(line) content entry[response][body][choices][0][message][content] print(f{entry[custom_id]}: {content})參考文檔模型文檔developers.openai.com/api/docs/models/gpt-6-lunaBatch APIdevelopers.openai.com/api/docs/batch定價developers.openai.com/api/docs/pricingGPT-6 使用指南developers.openai.com/api/docs/guides/latest-model以上內容基于 2026 年 9 月的 API 版本整理具體參數(shù)以官方文檔為準。