如何处理LLM API调用中的各种错误(超时、限流、格式错误等)?给出健壮的重试策略。
📋 面试问题
如何处理LLM API调用中的各种错误(超时、限流、格式错误等)?给出健壮的重试策略。
✅ 期望回答
import asyncio, time
from openai import AsyncOpenAI, APIError, RateLimitError, APITimeoutError
client = AsyncOpenAI(max_retries=0) # 关闭SDK内置重试,用自定义重试
async def robust_call(prompt: str, max_retries: int = 3):
for attempt in range(max_retries):
try:
resp = await asyncio.wait_for(
client.chat.completions.create(
model='gpt-4',
messages=[{'role':'user','content':prompt}],
temperature=0.3,
max_tokens=1000
),
timeout=30.0 # 请求超时
)
return resp.choices[0].message.content
except RateLimitError:
# 429: 令牌/请求数超限 -> 指数退避+等待
wait = 2 ** attempt + random.uniform(0, 1)
await asyncio.sleep(wait)
except APITimeoutError:
# 服务端超时 -> 重试(最多2次)
if attempt < max_retries - 1: continue
raise
except APIError as e:
if e.status_code >= 500: # 5xx服务端错误 -> 重试
await asyncio.sleep(2 ** attempt)
else: # 4xx客户端错误 -> 不重试,直接报错
raise
raise Exception('Max retries exceeded')
输出解析错误处理:
def safe_json_parse(response: str):
try:
return json.loads(response)
except json.JSONDecodeError:
# 尝试提取markdown代码块中的JSON
match = re.search(r'
json\n(.*?)\n``', response, re.DOTALL)
if match:
return json.loads(match.group(1))
raise ValueError(f'无法解析JSON: {response}')
``
Key Takeaways:
- 429/5xx -> 指数退避重试;4xx -> 不重试
- 输出解析包裹try-catch+后备逻辑
- 整个函数加超时保护