TS TypeSafe 文档中文版 原文 ↗

Choice

Choice 是一种 System One 问题类型,用于从一个已定义的集合中选出唯一一个选项。答案包含被选中的选项、每个选项的概率,以及置信度。

当答案是一组固定选项中的一个时,就使用 Choice。例如,某个工单由哪个团队处理、某个商品属于哪个类别,或者一段代码是用哪种语言写的。如果答案是某个谱系上的一个位置,请使用 Score。如果是「是」或「否」,请使用 Noul。选择问题类型对这三者做了比较。

Choice 的答案是被选中的选项,位于 choice 中。模型还会在 probabilities 中返回每个选项的概率,并为被选中的选项返回一个 confidence 值。

示例问题:

text
"What programming language is this code written in"
  → options: python, javascript, typescript, go, rust, other
 
"What type of meeting is this based on the title and description"
  → options: standup, planning, retrospective, one on one, brainstorm, none of the above
 
"Which product category does this item belong to"
  → options: electronics, clothing, home garden, food and beverage
 

请求结构

发往 TypeSafe API 的 POST 请求体有特定的结构。顶层有三个字段:state,即要评估的内容;model;以及 questions,一个从你选择的问题 id 到问题对象的映射。每个 Choice 问题有以下字段:

  • type:始终为 "choice"。
  • instructions:模型要回答的问题。
  • criteria:答案选项,以映射表示。每个键是一个选项名称,每个值是该选项的描述。

下面是一个请求,其状态是一家在线鞋店的一条支持工单,问题是应该由哪个团队处理它:

request
{
  "state": "My running shoes arrived in the wrong size. Can I swap them for a size 10?",
  "selectedModels": [
    "jev-latest"
  ],
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this?",
      "criteria": {
        "returns": "Exchanges, wrong or damaged items",
        "shipping": "Delivery status, delays, lost packages",
        "billing": "Charges, invoices, payment problems"
      }
    }
  }
}

这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

问题 id 由你选择,这里是 department。答案会以相同的 id 返回。模型永远看不到问题 id。选项名称和它们的描述都会发送给模型,所以要写出能把各个选项彼此区分开的描述。

我们的客户端 SDK提供带类型的问题。在 Python 中,同一个问题就是一个 Choice:

python
from typesafe_sdk import Choice, TypeSafeClient
 
with TypeSafeClient() as client:
    response = client.system_one(
        state="My running shoes arrived in the wrong size. Can I swap them for a size 10?",
        questions={
            "department": Choice(
                instructions="Which team should handle this?",
                criteria={
                    "returns": "Exchanges, wrong or damaged items",
                    "shipping": "Delivery status, delays, lost packages",
                    "billing": "Charges, invoices, payment problems",
                },
            ),
        },
    )
 
    print(response.answers["department"].choice)
 

使用 system_one 方法或 https://api.typesafe.ai/v1/systemone 端点来调用 System One 模型。model 字段选择由哪个模型处理该请求。如何用 TypeSafe 构建介绍了应该在代码中的什么地方调用它。

使用我们的某个客户端 SDK,或者直接调用 HTTP API。如果由编码智能体为你编写集成,请先安装 TypeSafe agent skill,这样它就知道请求和响应的结构。

注意

instructions 以及 criteria 中的每一项都可以是字符串、对象或数组。先从字符串开始。当一段描述需要好几种指引时,就使用对象,比如某个选项涵盖什么、不涵盖什么,以及一些示例。参见下文的结构化的 instructions 和 criteria以及 API 参考。

响应结构

响应的 answers 中为每个问题各有一项,位于请求所用的 id 之下。下面是上面那个示例请求的响应:

json
{
  "model": "jev-1.13.0",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "returns",
      "confidence": 1.0,
      "probabilities": {
        "shipping": 0.0,
        "returns": 1.0,
        "billing": 0.0
      }
    }
  },
  "usage": {
    "input_tokens": 328,
    "output_tokens": 34
  }
}
 

除了 type,每个 Choice 答案还有三个值:

  • choice:概率最高的选项。
  • probabilities:跨越所有选项的完整概率分布。所有值之和为 1。
  • confidence:一个 0 到 1 的数字,由 probabilities 的分布情况计算得出。形状平坦、概率分散在多个选项上意味着低置信度。在单个选项上出现单一峰值意味着高置信度。

这条工单很容易,所以全部概率都落在 returns 上,置信度为 1.0。一条同时提到尺码不对和退款未到的工单会把概率分摊在 returns 和 billing 之间,置信度也会下降。

最佳实践:一次调用中问多个问题

把代码可能需要的每个 Choice 问题都放在一个请求里,而不是每个问题发一次请求。所有问题都是并行评估的。增加问题几乎不会改变响应时间,而代码可以忽略它不需要的答案。多余的问题仍然会消耗 token。一起问多个问题对此做了完整说明;下一节展示了在一次调用中问五个 Choice 问题。

同样的逻辑也适用于单个 Choice 问题内部的选项。一个 Choice 问题最多接受 255 个选项,而每增加一个选项只会多消耗几个 token,所以要把完整的团队、类别或商品列表交给模型,而不是只给一个候选清单。当这个列表可能无法覆盖所有输入时,加上一个 other 或 none of the above 选项,这样模型就能表示其他选项都不合适。

要通过很深的层级体系或很大的分类体系来对文档进行分类时,就把 Choice 问题逐层串联起来。层级分类实践手册展示了如何对 Choice 概率做束搜索,在每一层保留最好的 K 条候选路径,而不是认定单条贪婪路径。

一个更复杂的示例

上面的基础示例把一条工单路由到一个团队。一个更大的支持系统可能还需要退货原因、配送问题、客户想要什么,以及客户的语气。

下面的请求针对一条比第一条更含糊的工单问了五个 Choice 问题:它牵涉三个团队,而且没有说明客户想要什么。

request
{
  "state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges of $120 on my card. What are you going to do about this?",
  "selectedModels": [
    "jev-latest"
  ],
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this?",
      "criteria": {
        "returns": "Exchanges, wrong or damaged items",
        "shipping": "Delivery status, delays, lost packages",
        "billing": "Charges, invoices, payment problems"
      }
    },
    "return_reason": {
      "type": "choice",
      "instructions": "If the customer wants to return something, why?",
      "criteria": {
        "wrong_size": "The item doesn't fit",
        "wrong_item": "A different product was delivered",
        "damaged": "The item arrived broken or faulty",
        "changed_mind": "The item is fine, the customer no longer wants it",
        "other": "A return reason that fits none of the above"
      }
    },
    "shipping_issue": {
      "type": "choice",
      "instructions": "If this is a shipping problem, which kind is it?",
      "criteria": {
        "not_delivered": "The package never arrived",
        "delayed": "The package is late but still on its way",
        "wrong_address": "The package went to the wrong place",
        "damaged_in_transit": "The package arrived damaged",
        "other": "A shipping problem that fits none of the above"
      }
    },
    "requested_resolution": {
      "type": "choice",
      "instructions": "What does the customer want to happen?",
      "criteria": {
        "exchange": "Swap the item for a different one",
        "refund": "Money back",
        "replacement": "The same item sent again",
        "information": "Just an answer, no action needed"
      }
    },
    "tone": {
      "type": "choice",
      "instructions": "What is the customer's tone?",
      "criteria": {
        "calm": null,
        "frustrated": null,
        "angry": null
      }
    }
  }
}

这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

这些 Choice 问题中有两个是推测性的:return_reason 只在 department 是 returns 时才有意义,shipping_issue 只在它是 shipping 时才有意义。tone 问题使用了 null 描述,因为这些选项名称本身就足够清楚。

TypeSafe 的响应:

json
{
  "model": "jev-1.13.0",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "returns",
      "confidence": 0.42,
      "probabilities": {
        "shipping": 0.04,
        "billing": 0.35,
        "returns": 0.61
      }
    },
    "return_reason": {
      "type": "choice",
      "choice": "wrong_size",
      "confidence": 1.0,
      "probabilities": {
        "other": 0.0,
        "wrong_size": 1.0,
        "changed_mind": 0.0,
        "damaged": 0.0,
        "wrong_item": 0.0
      }
    },
    "shipping_issue": {
      "type": "choice",
      "choice": "delayed",
      "confidence": 0.67,
      "probabilities": {
        "wrong_address": 0.0,
        "other": 0.26,
        "not_delivered": 0.0,
        "damaged_in_transit": 0.0,
        "delayed": 0.74
      }
    },
    "requested_resolution": {
      "type": "choice",
      "choice": "refund",
      "confidence": 0.2,
      "probabilities": {
        "replacement": 0.34,
        "refund": 0.4,
        "information": 0.02,
        "exchange": 0.24
      }
    },
    "tone": {
      "type": "choice",
      "choice": "frustrated",
      "confidence": 0.76,
      "probabilities": {
        "frustrated": 0.84,
        "angry": 0.16,
        "calm": 0.0
      }
    }
  },
  "usage": {
    "input_tokens": 589,
    "output_tokens": 212
  }
}
 

每个问题都独立地针对这条工单作答:

  • department 的答案是 returns,概率为 0.61,但由于重复扣费,billing 也占 0.35。这条工单属于两个团队,被拆分的 0.42 置信度正反映了这一点。
  • return_reason 是 wrong_size,置信度为 1.0,这是意料之中的,因为工单里明确说了这一点。
  • shipping_issue 的答案在 delayed 和 other 之间分摊。这是一个推测性问题,而 department 的结果并不是 shipping,所以代码可以忽略它,如下面的示例代码片段所示。
  • requested_resolution 的答案偏向 refund,为 0.40,replacement 和 exchange 分摊了剩下的大部分,置信度为 0.20。重复扣费暗示要退款,尺码不对暗示要换货,而客户从未说明自己想要哪一种。
  • tone 的答案是 frustrated,概率为 0.84,置信度为 0.76。

下面的示例代码读取它需要的答案,忽略其余的,并把低置信度的答案当作询问而不是行动的理由:

python
from typesafe_sdk import Choice, TypeSafeClient
 
TRIAGE_QUESTIONS = {
    "department": Choice(
        instructions="Which team should handle this?",
        criteria={
            "returns": "Exchanges, wrong or damaged items",
            "shipping": "Delivery status, delays, lost packages",
            "billing": "Charges, invoices, payment problems",
        },
    ),
    "return_reason": Choice(
        instructions="If the customer wants to return something, why?",
        criteria={
            "wrong_size": "The item doesn't fit",
            "wrong_item": "A different product was delivered",
            "damaged": "The item arrived broken or faulty",
            "changed_mind": "The item is fine, the customer no longer wants it",
            "other": "A return reason that fits none of the above",
        },
    ),
    "shipping_issue": Choice(
        instructions="If this is a shipping problem, which kind is it?",
        criteria={
            "not_delivered": "The package never arrived",
            "delayed": "The package is late but still on its way",
            "wrong_address": "The package went to the wrong place",
            "damaged_in_transit": "The package arrived damaged",
            "other": "A shipping problem that fits none of the above",
        },
    ),
    "requested_resolution": Choice(
        instructions="What does the customer want to happen?",
        criteria={
            "exchange": "Swap the item for a different one",
            "refund": "Money back",
            "replacement": "The same item sent again",
            "information": "Just an answer, no action needed",
        },
    ),
    "tone": Choice(
        instructions="What is the customer's tone?",
        criteria={"calm": None, "frustrated": None, "angry": None},
    ),
}
 
 
def triage(ticket: str) -> None:
    with TypeSafeClient() as client:
        response = client.system_one(
            state=ticket,
            questions=TRIAGE_QUESTIONS,
        )
    answers = response.answers
 
    department = answers["department"]
    if department.confidence < 0.3:
        # Not clear which team to send to. Let a person decide.
        send_to_manual_triage(ticket)
        return
 
    if department.choice == "returns":
        # return_reason answer is only used here
        assign(ticket, team="returns", issue=answers["return_reason"].choice)
    elif department.choice == "shipping":
        # shipping_issue answer is only used here
        assign(ticket, team="shipping", issue=answers["shipping_issue"].choice)
    else:
        assign(ticket, team="billing")
 
    # A second team with a real share of the probability gets a copy
    for team, probability in department.probabilities.items():
        if team != department.choice and probability > 0.25:
            notify(ticket, team=team)
 
    resolution = answers["requested_resolution"]
    if resolution.confidence < 0.5:
        # The customer hasn't said what they want. Ask, don't guess.
        ask_customer_what_they_want(ticket)
    elif resolution.choice == "refund":
        flag_for_refund_approval(ticket)
 
    if answers["tone"].choice == "angry":
        flag_for_senior_agent(ticket)
 

对于上面的工单,这段代码把工单分配给 returns 团队,问题为 wrong_size;因为 billing 的 0.35 份额超过了 0.25 的阈值,所以抄送一份给 billing 团队;又因为解决方式的置信度 0.20 低于 0.5,所以询问客户想要什么。这段代码没有使用 shipping_issue 的答案。

一次请求,五个答案,而路由逻辑就是普通的 if 语句。如果你之后需要知道客户的语言,或者这条工单涉及哪个商品,就往 TRIAGE_QUESTIONS 里再加一个 Choice 问题;请求数量仍然是一次。

智能家居助手演示在一次调用中用一长串 Choice 问题来评估每个用户请求:请求类别、房间、设备以及动作。这些问题大多与任何一个具体请求都无关,代码会把它们忽略掉。

结构化的 instructions 和 criteria

先从每个选项一行描述开始。当两个选项很相似、模型总是把它们弄混时,就用对象而不是字符串来描述每一个。给它一些字段,说明这个选项涵盖什么、什么反而属于邻近的另一个选项,以及几个示例输入。

下面这两个答案选项 return_policy 和 return_status 很容易混淆。关于其中任何一个的工单都可能提到退货和退款,所以每个选项都说明了自己不适用于什么。

request
{
  "state": "I sent the shoes back a week ago. When do I get my money?",
  "selectedModels": [
    "jev-latest"
  ],
  "questions": {
    "return_topic": {
      "type": "choice",
      "instructions": {
        "question": "Which returns topic is the customer asking about?",
        "focus": "Classify the information the customer wants."
      },
      "criteria": {
        "return_policy": {
          "what": "Whether and how an item can be returned",
          "not_for": "Progress of a return already sent",
          "examples": [
            "Can I return shoes I've worn once?",
            "How long do I have to return an order?"
          ]
        },
        "return_status": {
          "what": "Progress of a return already sent",
          "not_for": "Whether and how an item can be returned",
          "examples": [
            "Has my return arrived yet?",
            "When will my refund be paid?"
          ]
        }
      }
    }
  }
}

这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

响应是 return_status,置信度为 1.0:

json
{
  "model": "jev-1.13.0",
  "answers": {
    "return_topic": {
      "type": "choice",
      "choice": "return_status",
      "confidence": 1.0,
      "probabilities": {
        "return_policy": 0.0,
        "return_status": 1.0
      }
    }
  },
  "usage": {
    "input_tokens": 407,
    "output_tokens": 32
  }
}
 

字段名 question、focus、what、not_for 和 examples 不属于 API 的一部分,也都不是保留字。它们由你选择,就像你选择选项名称一样。模型会连同取值一起看到这些名称,所以要使用能标示其内容的简短名称。