TS TypeSafe 文档中文版 原文 ↗

置信度门控路由

把置信度当作第二个维度。答案告诉你是什么;置信度告诉你是否该行动。

TypeSafe 最强大的特性之一是置信度。通过有意识地用置信度来门控决策,你可以构建既可靠又安全的系统。

示例:语音银行指令

让我们设想你正在构建一个语音银行界面,让用户可以用语音与自己的账户交互。虽然你总是希望对用户意图的解读保持合理的置信度,但有些操作风险更高,因此要求更高的置信度阈值。

%%{init: {"fontFamily": "Inter, sans-serif", "flowchart": {"rankSpacing": 35, "wrappingWidth": 300, "subGraphTitleMargin": {"top": 12, "bottom": 36}}}}%% flowchart LR command["语音银行指令"] subgraph req["TypeSafe 评估<br/>该问题"] intent["<b>Choice:</b> 意图"] end command -- "一次请求<br/>指令 + 意图<br/>问题" --> req req -- "一次响应<br/>意图答案 +<br/>置信度" --> gate{"<b>置信度足够高吗?</b><br/>你的代码"} gate -- "低于 0.6<br/>或其他意图" --> human["转给支持人员"] gate -- "check_balance<br/>至少 0.6" --> balance["显示余额"] gate -- "approve_transfer<br/>0.6 到 0.85" --> confirm["请用户确认"] gate -- "approve_transfer<br/>高于 0.85" --> approve["批准转账"]

第 1 步:确定用户的意图

questions
{
  "questions": {
    "intent": {
      "type": "choice",
      "instructions": "What action is the user requesting?",
      "criteria": {
        "check_balance": "Check the balance of an account",
        "approve_transfer": "Approve the pending transfer request",
        "other": "Something else"
      }
    }
  }
}

这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

第 2 步:置信度门控路由

python
action = response.answers["intent"]
 
# Below 0.6 confidence on any action, route to a human
if action.confidence < 0.6:
    route_to_support_agent(account_id)
 
elif action.choice == "check_balance":
    # Low stakes. 0.6 confidence is sufficient.
    show_balance(account_id)
 
elif action.choice == "approve_transfer":
    if action.confidence > 0.85:
        # High stakes, but high confidence. Safe to act automatically.
        approve_transfer(account_id)
    else:
        # High stakes, moderate confidence. Verify intent first.
        ask_user_to_confirm("Just to confirm: you would like to approve this transfer, is that correct?")
 
else:
    route_to_support_agent(account_id)
 

0.6 这个下限能捕获模型真正不确定的任何情况。在这个下限之上,每种操作类型都有自己的阈值,依据是对错误分类采取行动的后果。以 0.6 的置信度查询余额是可以的,因为最坏的情况不过是用户得听一遍余额播报。但批准转账需要非常高的置信度(>0.85),否则系统应当请用户确认。

关于如何在你的系统中思考置信度,详见置信度。