Score
Score 是一种 System One 问题类型,用于按照有序的、带描述的等级来给内容评分。答案包含一个 score、每个等级的概率以及 confidence。
当答案是你能够分步描述的某个谱系上的一个位置时,就使用 Score。例如,一个 bug 有多严重、一位客户有多满意,或者一位候选人的 Python 经验有多少。如果答案是固定的一组选项之一,且它们之间没有顺序,请使用 Choice。如果是「是」或「否」,请使用 Noul。选择问题类型对这三者做了比较。
Score 的答案在 score 中是你各个等级上的一个位置,它可以落在两个等级之间。模型还会在 probabilities 中返回每个等级的概率,并为该答案返回一个 confidence 值。
这个控件可以交互查看不同等级分布下的评分与概率。本地静态版无法运行它,请到官网原页面体验。 打开官网原页面
每个等级前面的数字是位置,在等级一节中有说明。
请求结构
发送到 TypeSafe API 的 POST 请求体与其他任何问题类型一样有三个顶层字段:state,即要评估的内容;model;以及 questions。每个 Score 问题有以下字段:
type:始终为"score"。instructions:模型要回答的问题。即它在给什么评分。criteria:一个有序的等级描述数组,从量表的低端到高端。至少应有两个等级;API 最多接受 10 个。
下面是一个请求,其状态是一份 bug 报告,问题是这个 bug 有多严重:
{
"state": "The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
"selectedModels": [
"jev-latest"
],
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
}
}
}
这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。
问题 id 由你选择,这里是 bug_severity。这个 id 不会发送给模型。答案会以相同的 id 返回。
等级
criteria 中的每一项都是一个等级:即可能答案谱系上的一个点,用文字描述。等级的编号是它在 criteria 数组中的位置,从 0 开始,因此上面三项分别是等级 0、1 和 2。数组的顺序就是编号。
模型只会拿到这些描述,别的什么都没有,并且每个等级都会针对该状态单独评判。
响应中的 score 是等级谱系上的一个位置。对于三级量表,它的取值范围是 0 到 2,并且可以落在两个等级之间。
我们的客户端 SDK提供带类型的问题。在 Python 中,同一个问题就是一个 Score:
from typesafe_sdk import Score, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
state="The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
questions={
"bug_severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
),
},
)
print(response.answers["bug_severity"].score)
使用 system_one 方法或 https://api.typesafe.ai/v1/systemone 端点来调用 System One 模型。model 字段选择由哪个模型处理该请求。如何用 TypeSafe 构建介绍了应该在代码中的什么地方调用它。
使用我们的某个客户端 SDK,或者直接调用 TypeSafe API。如果由编码智能体为你编写集成,请先安装 TypeSafe agent skill,这样它就知道请求和响应的结构。
响应结构
响应的 answers 中为每个问题各有一项,位于请求所用的 id 之下。下面是上面那个示例请求的响应:
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.43,
"confidence": 0.35,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": {
"0": 0.0,
"1": 0.57,
"2": 0.43
}
}
},
"usage": {
"input_tokens": 332,
"output_tokens": 18
}
}
每个 Score 答案有五个值:
type:TypeSafe 问题的类型。probabilities:每个等级的概率,以等级编号的字符串为键。所有值之和为 1。score:等级编号数轴上的位置,从 0 到最高等级编号(这里是 2)。它等于每个等级编号乘以其概率后相加:0 x 0.0 + 1 x 0.57 + 2 x 0.43 = 1.43。legend:每个等级编号映射回它的描述。confidence:一个 0 到 1 的数字,由probabilities的分布情况计算得出。概率集中在单个等级上意味着高置信度。概率分散在多个等级上意味着低置信度。
1.43 的 score 意味着模型在等级 1 和等级 2 之间摇摆,略偏向等级 1。这与报告相符:导出功能坏了,切换到 Chrome 对大多数客户来说是一种变通办法,但对只用 Safari 的那些客户不是。模型把 0.57 放在「存在变通办法」上,把 0.43 放在「没有变通办法」上,而 confidence 是 0.35,因为它摇摆不定。
使用 Python SDK 时,ScoreAnswer 有 score、confidence、probabilities 和 legend 这些带类型的字段。SDK 中 probabilities 和 legend 以整数等级为键,而不是字符串。
解读 Score
我们来看看 score 如何随不同输入而变化。例如,使用上面请求中的问题及其等级:
"How severe is the reported issue?"
→ 0: Cosmetic; no impact to functionality
→ 1: Broken or degraded feature, but workaround exists
→ 2: Blocking issue; no workaround exists
我们可以看到不同的 bug 报告如何改变 score:
| 设置页面上导出按钮错位了几个像素。 | 0.0 | 1.0 | 1.0 | 0.0 | 0.0 |
在这些例子中,confidence 为 1.0 意味着返回的分布把全部概率都放在一个等级上。这描述的是模型的答案,而不是对答案正确性的保证。
score 是各等级编号按概率加权的平均值。在第三个和第四个例子中,概率被分摊在等级 1 和等级 2 上。等级 2 上的权重越大,score 就越高。它并不衡量没有变通办法的客户所占的比例。
不同的分布可以产生相同的 score。score 为 1.0 可能意味着全部概率都在等级 1 上,也可能意味着等级 0 和等级 2 上各占一半。要区分这些情况,请把 probabilities 和 confidence 与 score 一起解读。
带小数的 score 是一个位置。你可以用它按严重程度对报告排序,或者在代码只需要一个结果时把它四舍五入到最近的等级。我们的实体对齐实践手册给出了一个四舍五入到最近等级以做出决策的例子。
Score 上的低置信度通常意味着三种情况之一。这些等级对该状态而言相互重叠,该问题在衡量不止一件事,或者状态提供的信息不足以定位它。我们的 Confidence 文档介绍了如何在代码中使用它。
编写好的等级
描述情境,而不是程度。「功能损坏或降级,但存在变通办法」给了模型可以拿状态去对照的东西。「中等严重」则没有。具体的描述能帮助模型区分各个等级。要用已知的示例来检验答案;单凭更高的置信度并不能说明某个描述更好。
每个等级都是单独评估的。模型看不到某个等级的编号,也看不到它相邻的等级,所以「比上一个等级更糟」对模型毫无意义,描述或 instructions 里的数字也帮不上忙。下面是当等级只有数字时,对上面表格中那条按钮错位的报告会发生什么:
instructions: "Rate severity from 0 to 2, where 2 is worst"
criteria: ["0", "1", "2"]
→ score 0.55, confidence 0.33, probabilities 0: 0.45, 1: 0.55, 2: 0.0
同一条报告在采用三个带描述的等级时,得分是 0.0,置信度为 1.0。而只有数字时,模型没有任何可对照的东西,于是把概率分摊在 0 和 1 之间。
在你能够清楚区分描述的范围内,尽可能多地使用等级,最多 10 个。三个就很好。不要添加你无法清楚区分描述的等级。
让每个 Score 问题只涉及一个维度。如果某个描述写的是「守时、聪明且有经验」,那么这个问题就在衡量三件事,而一个在某方面高、在另一方面低的输入就无法被定位。置信度会下降,score 的意义也会变小。把它拆成每件事一个 Score 问题,然后在代码中把它们组合起来,如下一节所示。
如果你量表的顶端有一个罕见的极端情况需要你采取不同的处理方式,就给它一个单独的等级。一个以「非常愤怒」结尾的情感量表可以加上「辱骂或威胁」。如果没有这个等级,两类消息可能都会得到接近顶端的 score。单靠 score 可能无法区分它们。
如果完全没有中间状态,而答案是少数几个离散类别之一,那就改用 Choice,或者把问题拆成几个 Noul 问题。用自己的数据来检验你的等级很重要。同一个量表的两种措辞在你的数据上可能表现不同。
把复杂的判断拆成几个 Score 问题
一个复杂的判断,即依赖好几件事的判断,最好拆成每件事一个 Score 问题。然后你可以在代码中把 TypeSafe 返回的这些 Score 组合起来,做出该判断。有些 Score 问题可能比其他问题更重要,所以给每个 Score 问题一个表示其相对重要性的权重。权重由你决定。当组合结果与你团队会做出的决定不符时,在代码中修改权重并重新运行。把这些 Score 问题放在一个请求里发送。它们会被并行评估。增加问题几乎不会改变响应时间,只会多花几个问题 token;参见把多个问题一起提问。
下面的请求就是上面表格中那张转圈(spinner)工单,并补充了一些上下文。它提出三个 Score 问题:这个 bug 有多严重、这位客户有多沮丧,以及这份报告能给工程师提供多少可用的信息。
{
"state": "Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.",
"selectedModels": [
"jev-latest"
],
"questions": {
"severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language or threatening to leave"
]
},
"report_quality": {
"type": "score",
"instructions": "How much does the report give an engineer to work with?",
"criteria": [
"No detail; just says something is broken",
"Names the feature but no steps or environment",
"Steps to reproduce or environment, but not both",
"Steps to reproduce and environment"
]
}
}
}
这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。
TypeSafe 的响应:
{
"model": "jev-1.13.0",
"answers": {
"severity": {
"type": "score",
"score": 1.24,
"confidence": 0.64,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": {
"0": 0.0,
"1": 0.76,
"2": 0.24
}
},
"frustration": {
"type": "score",
"score": 1.28,
"confidence": 0.58,
"legend": {
"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language or threatening to leave"
},
"probabilities": {
"0": 0.0,
"1": 0.72,
"2": 0.28
}
},
"report_quality": {
"type": "score",
"score": 3.0,
"confidence": 1.0,
"legend": {
"0": "No detail; just says something is broken",
"1": "Names the feature but no steps or environment",
"2": "Steps to reproduce or environment, but not both",
"3": "Steps to reproduce and environment"
},
"probabilities": {
"0": 0.0,
"1": 0.0,
"2": 0.0,
"3": 1.0
}
}
},
"usage": {
"input_tokens": 468,
"output_tokens": 43
}
}
每个问题都会针对该工单单独作答,并给出一个 score:
severity是 1.24,置信度为 0.64。与开头的例子解读相同:导出功能坏了,有些人有变通办法。frustration是 1.28,置信度为 0.58。措辞是礼貌的,但「第三次」和「我受够了」把一部分 score 推向了最高等级,于是模型在「沮丧但礼貌」和「非常愤怒」之间分出了 0.72 和 0.28。对这张工单来说,这两个等级相互重叠,这就是置信度处于中等水平的原因。report_quality是 3.0,置信度为 1.0。复现步骤和浏览器版本都写明了。
这三个量表的长度不同,所以在组合之前,要先对每个 score 做归一化。四级量表返回 0 到 3,三级量表返回 0 到 2,因此一个量表上的最高分比另一个量表上的最高分更大。把每个 score 除以它的最高等级编号 len(criteria) - 1,让每个 score 都落在 0 到 1 上。这样权重才名副其实:严重程度 0.6、沮丧程度 0.3,意味着严重程度的分量是它的两倍。
下面的 TypeSafe Python SDK 代码提出这三个问题,对每个 score 做归一化,并用一个示例优先级计算把它们组合起来:
from typesafe_sdk import Score, TypeSafeClient
TRIAGE_QUESTIONS = {
"severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
),
"frustration": Score(
instructions="How frustrated is the customer?",
criteria=[
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language or threatening to leave",
],
),
"report_quality": Score(
instructions="How much does the report give an engineer to work with?",
criteria=[
"No detail; just says something is broken",
"Names the feature but no steps or environment",
"Steps to reproduce or environment, but not both",
"Steps to reproduce and environment",
],
),
}
def normalized(answers, question_id: str) -> float:
"""Put a score on 0 to 1 by dividing by its top level number."""
top_level = len(TRIAGE_QUESTIONS[question_id].criteria) - 1
return answers[question_id].score / top_level
def priority(ticket: str) -> float:
with TypeSafeClient() as client:
response = client.system_one(
state=ticket,
questions=TRIAGE_QUESTIONS,
)
answers = response.answers
severity = normalized(answers, "severity")
frustration = normalized(answers, "frustration")
report_quality = normalized(answers, "report_quality")
# A detailed report helps an engineer investigate, so it raises priority a little.
return 0.6 * severity + 0.3 * frustration + 0.1 * report_quality
对于上面那个示例响应,归一化后的分数是:严重程度 0.62、沮丧程度 0.64、报告质量 1.0。优先级为 0.6 × 0.62 + 0.3 × 0.64 + 0.1 × 1.0 = 0.664,四舍五入为 0.66。
权重就放在你的代码里,所以你能确切看到这个数字是如何得出的,并在排序与你团队会做的决定不符时修改它。如果之后你需要更多 Score 问题,把它们加进 TRIAGE_QUESTIONS 即可。请求数量仍然是一个。这种把一个复杂判断拆成若干个独立 Score、再在代码中用权重把它们组合起来的技术,称为复合评分模式。
结构化的等级描述
先从每个等级的基本文本描述开始。当模型在你认为很明确的输入上总是给出介于相邻两个等级之间的分数时,就给每个等级一个对象而不是字符串,其中一个字段说明该等级涵盖什么,另一个字段给出几个示例情境。在每个等级上使用相同的字段名,这样模型才能做同类比较。
下面的请求就是我们之前用过的那张转圈工单,但每个等级上都带了示例:
{
"state": "Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too.",
"selectedModels": [
"jev-latest"
],
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
{
"what": "Cosmetic; no impact to functionality",
"examples": [
"typo in a label",
"misaligned icon"
]
},
{
"what": "Broken or degraded feature, but workaround exists",
"examples": [
"export fails in one browser but works in another"
]
},
{
"what": "Blocking issue; no workaround exists",
"examples": [
"cannot log in",
"data loss"
]
}
]
}
}
}
这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。
响应:
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.09,
"confidence": 0.87,
"legend": {
"0": {
"what": "Cosmetic; no impact to functionality",
"examples": [
"typo in a label",
"misaligned icon"
]
},
"1": {
"what": "Broken or degraded feature, but workaround exists",
"examples": [
"export fails in one browser but works in another"
]
},
"2": {
"what": "Blocking issue; no workaround exists",
"examples": [
"cannot log in",
"data loss"
]
}
},
"probabilities": {
"0": 0.0,
"1": 0.91,
"2": 0.09
}
}
},
"usage": {
"input_tokens": 379,
"output_tokens": 18
}
}
使用纯字符串时,这张工单得分为 1.11,置信度为 0.84。加上示例后,它的得分是 1.09,置信度为 0.87,变化很小,因为纯字符串本来就已经把它放得很准。当纯字符串让模型左右摇摆时,效果会更大,如下表所示。
示例会引导模型,只有当它们看起来像你的真实输入时才有帮助。下表是最开头那条 Safari 报告配三组不同的等级对象:
| 等级描述 | score |
confidence |
|---|---|---|
| 纯字符串:没有带示例的对象 | 1.43 | 0.35 |
| 添加了 examples 数组,其中是有用的示例:「在一个浏览器中导出失败,但在另一个中正常」 | 1.03 | 0.96 |
| 添加了 examples 数组,其中是与浏览器无关的示例:「搜索失败,但浏览分类仍然正常」 | 1.43 | 0.35 |
在这个对比中,匹配的示例把几乎全部概率都集中到了一个等级上。不相关的示例则返回与纯字符串相同的结果。更高的置信度并不能确定哪个答案是正确的。选择那些预期等级已知的示例,然后在单独的输入上测试修改后的描述,确认没问题后再保留。