TS TypeSafe 文档中文版 原文 ↗

Noul

Noul 问题要求 TypeSafe 模型评估一个是/否问题,并返回答案为「是」的概率。

当答案只有「是」或「否」时,就使用 Noul。例如:这条消息是否要求退款,这份简历是否提到分布式系统,这条评论是否包含个人数据。如果答案是若干选项之一,请使用 Choice。如果它是某个谱系上的一个位置,请使用 Score。选择问题类型对这三者做了比较。

Noul 的答案是一个数字,表示答案为「是」的概率,其中 0 表示否,1 表示是。

请求结构

发往 TypeSafe API 的 POST 请求体与其他任何问题类型一样,有三个顶层字段:state,即要评估的内容;model;以及 questions。每个 Noul 问题有以下字段:

  • type:始终为 "noul"。
  • instructions:模型要回答的是/否问题,或者供它判断的一句陈述。
  • criteria:可选。一个对象,包含 true 和 false 两个描述,说明「是」和「否」分别意味着什么。

下面是一个请求,其状态是一条支持消息,两个问题分别是客户是否想要人工服务,以及他们之前是否联系过支持团队:

request
{
  "state": "I have asked three times now. Can I please just talk to a real person?",
  "selectedModels": [
    "jev-latest"
  ],
  "questions": {
    "is_human_escalation": {
      "type": "noul",
      "instructions": "Is the customer asking for a human agent?"
    },
    "is_repeat_contact": {
      "type": "noul",
      "instructions": "Has the customer contacted support about this before?",
      "criteria": {
        "true": "Mentions a prior attempt, ticket, or that they have asked before",
        "false": "No sign of any previous contact"
      }
    }
  }
}

这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

问题 id 由你选择,这里是 is_human_escalation 和 is_repeat_contact。这些 id 不会发送给模型。每个答案都以相同的 id 返回。第一个问题只依赖 instructions。第二个问题额外使用 criteria,说明什么算「是」、什么算「否」。

使用 Python SDK 时,同样的问题写成 Noul 对象:

python
from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient
 
with TypeSafeClient() as client:
    response = client.system_one(
        model="jev-latest",
        state="I have asked three times now. Can I please just talk to a real person?",
        questions={
            "is_human_escalation": Noul(
                instructions="Is the customer asking for a human agent?",
            ),
            "is_repeat_contact": Noul(
                instructions="Has the customer contacted support about this before?",
                criteria=NoulCriteria(
                    true="Mentions a prior attempt, ticket, or that they have asked before",
                    false="No sign of any previous contact",
                ),
            ),
        },
    )
 
    print(response.answers["is_human_escalation"].noul)
    print(response.answers["is_repeat_contact"].noul)
 

system_one 方法和 https://api.typesafe.ai/v1/systemone 端点都以 TypeSafe 的 AI 模型 System One 命名。如何用 TypeSafe 构建介绍了在代码中的何处使用它。

如果你使用编码 agent,请先安装 TypeSafe agent skill,这样它就能了解请求与响应的结构。

注意

instructions 可以是字符串、对象或数组。先从字符串开始。当问题需要附带数据时——例如用一条记录来与状态比对——或者当问题的一部分由你的代码生成时,就使用对象。在问题中使用结构说明了结构在什么时候有帮助,下面的示例则展示了用代码构建问题的方式。

响应结构

响应中的 answers 为每个问题提供一个条目,键就是请求中的 id:

json
{
  "model": "jev-1.13.0",
  "answers": {
    "is_human_escalation": {
      "type": "noul",
      "noul": 0.99
    },
    "is_repeat_contact": {
      "type": "noul",
      "noul": 0.93
    }
  },
  "usage": {
    "input_tokens": 360,
    "output_tokens": 39
  }
}
 

这里的两个答案都接近 1。客户说「talk to a real person」(想和真人说话),所以 is_human_escalation 为 0.99。「I have asked three times now」符合 is_repeat_contact 的 true 描述,所以为 0.93。

解读 Noul

这个数字同时是答案和确定性。接近 1 的值是强烈的「是」。接近 0 的值是强烈的「否」。接近 0.5 的值意味着模型给「是」和「否」的概率相近。

下表展示了 jev-1.13.0 对不同客户消息在 is_human_escalation 问题上的记录答案:

状态 noul
谢谢,修好了! 0.02
我怎么重置密码? 0.07
我今天就要解决这件事,不惜一切代价。 0.26
你是机器人吗? 0.40
有没有办法就我的账单找个人谈谈? 0.84
我已经问过三次了。能不能让我和一个真人说话? 0.99

前两条和后两条都很明确。「我今天就要解决这件事」很紧急,但从未要求找人工,得分为 0.26。「你是机器人吗?」暗示想要人工服务却没有直接提出要求,模型几乎均分概率,为 0.40。这两类消息都需要在你的代码中基于阈值来做出决定。

与 Choice 或 Score 不同,Noul 没有单独的 confidence 值。Noul 的概率分布只有两个结果——是和不——所以单个 noul 值就能完整描述它。Choice 或 Score 把概率分散到多个选项或等级上,confidence 则概括这种分散程度。

最常见的情况是,你的代码把 noul 按阈值转换成布尔值:

python
wants_human = response.answers["is_human_escalation"].noul > 0.9
 
if wants_human:
    route_to_agent(ticket)
else:
    route_to_bot(ticket)
 

阈值设在哪里取决于判断错误的代价。当「是」和「否」同样容易处理时,用 0.5。当对错误的「是」采取行动代价高昂时(例如呼叫值班人员或发放退款),就调高阈值。当漏掉一个真实的「是」代价高昂时(例如未能标记出安全问题),就调低阈值。处于中间的值可以交给人工,而不是走任何一条代码路径。这与 置信度页面为 Choice 和 Score 答案描述的三种分流方式相同。

Noul 的值在 0 到 1 之间,但它并不是你所问之事的一个刻度。它是答案为「是」的概率。如果问题实际上关乎程度,这个值并不衡量程度。下面针对四位候选人问了「Is the candidate strong in Python?」,旁边是一个含四个等级的 Score:没有经验、有些熟悉、在工作中经常使用、深厚专长。

候选人 Noul:「候选人在 Python 方面是否很强?」 Score:「候选人的 Python 经验有多少?」
我的经验在 Java 和 Go。我没有用过 Python。 0.03 0.0(没有经验)
除了主要的 Java 工作之外,我偶尔用 Python 写些小脚本。 0.14 1.0(有些熟悉)
上一份工作中我每天使用 Python,持续了两年,主要是数据管道。 0.81 2.05(在工作中经常使用)
我八年来每天都写 Python,包括维护一个大型 Django 代码库。 0.92 2.89(深厚专长)

Noul 判断的是单一命题——「强」——这些值就是它为真的可能性。你可以在代码中自行在 0 到 1 的范围内划分等级,比如把 0.3 到 0.7 视为「有一些经验」,但模型看不到这些等级,因此答案中没有任何内容是针对它们做出的判断。中间值可能意味着中等经验,也可能意味着情况不明,而候选人之间的间隔也不是你选择的。Score 则逐个判断每个等级描述,因此每位候选人都落在你写下的某个等级上或其附近,返回的概率显示了模型如何在不同等级之间分配它的判断。如果你不认同,改写某个等级的描述再跑一次。选择问题类型解释了这一区别。

编写 Noul 问题

每个 Noul 只问一个是/否问题。如果一个问题包含两个条件,比如「客户是否既生气又要求退款?」,模型就必须同时判断两者,这个值的意义就会减弱。请改为问两个 Noul,并在代码中组合它们。

措辞要让高值意味着「是」。「这条消息是否包含个人数据?」很清晰。「这条消息是否不含个人数据?」则把含义反转了,之后读取它的代码会得到相反的结果。

陈述句和疑问句同样有效。对于「客户正在要求退款」,接近 1 的值表示该陈述为真。用你自己的数据把两种措辞都试一遍,看看哪种效果更好。

让「是」与「否」的边界没有歧义。「这位候选人有任何 Python 经验吗?」效果很好,因为「任何」不留中间地带。当边界比较微妙时,就添加带 true 和 false 描述的 criteria,正如上面的 is_repeat_contact 问题那样。对大多数 Noul 来说指令已经足够,所以把带与不带 criteria 的问题都试一下,保留在你的文档上给出更好答案的那一种。

良好实践:每次调用问多个问题

对于一组条件清单,在一次请求中问多个 Noul 问题:每个条件一个问题,由代码决定组合的含义。问题会并行评估,所以增加 Noul 几乎不会改变响应时间。同时问多个问题对此有更详细的说明。

在代码中处理多个 Noul 答案

上面的两问题请求已经足以让代码对消息进行路由。下面的例子在客户要求人工时升级到人工,在他们此前联系过时提高优先级。任一问题出现中间值时,就交给审核人员,而不是走某条代码路径:

python
from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient
 
SUPPORT_QUESTIONS = {
    "is_human_escalation": Noul(
        instructions="Is the customer asking for a human agent?",
    ),
    "is_repeat_contact": Noul(
        instructions="Has the customer contacted support about this before?",
        criteria=NoulCriteria(
            true="Mentions a prior attempt, ticket, or that they have asked before",
            false="No sign of any previous contact",
        ),
    ),
}
 
YES = 0.8
NO = 0.2
 
 
def route(message: str) -> None:
    with TypeSafeClient() as client:
        response = client.system_one(
            model="jev-latest",
            state=message,
            questions=SUPPORT_QUESTIONS,
        )
    answers = response.answers
 
    wants_human = answers["is_human_escalation"].noul
    repeat = answers["is_repeat_contact"].noul
 
    if NO < wants_human < YES or NO < repeat < YES:
        # The model isn't sure either way. Let a person decide.
        send_to_review(message)
        return
 
    priority = "high" if repeat > YES else "normal"
    if wants_human > YES:
        route_to_agent(message, priority=priority)
    else:
        route_to_bot(message, priority=priority)
 

对于上面的消息,is_human_escalation 的 noul 答案值为 0.99,is_repeat_contact 为 0.93,所以代码会以高优先级把它路由给人工。而消息「我怎么重置密码?」在两个问题上都是 0.07,会被路由给机器人。

阈值存在于你的代码中。如果审核人员看到的待审消息太多,就缩小 NO 和 YES 之间的间隔。如果错误路由通过得太多,就扩大它。如果你之后需要知道消息是否提到付款,或者是否包含个人数据,就再往 SUPPORT_QUESTIONS 里加一个 Noul。请求数量仍然只有一次。

结构化 instructions

Instructions 可以是对象而不是字符串,把问题放在一个字段中,把补充数据放在其他字段中。在问题中使用结构介绍了它在什么时候有帮助。这里它用于一个由代码构建的问题:把一份刚收到的简历与候选人数据库中可能是同一个人的记录进行比对。每条记录原样放入 potential_duplicate 字段,question 对每条记录都相同,所有记录在一次请求中检查。由代码生成的问题键包含每条记录的数据库 ID:

request
{
  "state": {
    "resume": {
      "name": "John Smith",
      "location": "Oakland, CA",
      "summary": "Backend engineer with eight years of Python and Go experience.",
      "experience": [
        {
          "employer": "Google",
          "title": "Senior Backend Engineer",
          "years": "2021-2025"
        },
        {
          "employer": "Microsoft",
          "title": "Software Engineer",
          "years": "2017-2021"
        }
      ]
    }
  },
  "selectedModels": [
    "jev-latest"
  ],
  "questions": {
    "same_as_record_18": {
      "type": "noul",
      "instructions": {
        "potential_duplicate": {
          "name": "Jon Smith",
          "location": "Oakland, CA",
          "last_employer": "Google"
        },
        "question": "Is the resume for the same person as `potential_duplicate`?"
      }
    },
    "same_as_record_42": {
      "type": "noul",
      "instructions": {
        "potential_duplicate": {
          "name": "John Smith",
          "location": "Austin, TX",
          "last_employer": "Lone Star Freight"
        },
        "question": "Is the resume for the same person as `potential_duplicate`?"
      }
    },
    "same_as_record_77": {
      "type": "noul",
      "instructions": {
        "potential_duplicate": {
          "name": "John Smithers",
          "location": "Oakland, CA",
          "last_employer": "Bay Health Clinic"
        },
        "question": "Is the resume for the same person as `potential_duplicate`?"
      }
    }
  }
}

这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

响应:

json
{
  "model": "jev-1.13.0",
  "answers": {
    "same_as_record_18": {
      "type": "noul",
      "noul": 0.74
    },
    "same_as_record_42": {
      "type": "noul",
      "noul": 0.09
    },
    "same_as_record_77": {
      "type": "noul",
      "noul": 0.08
    }
  },
  "usage": {
    "input_tokens": 535,
    "output_tokens": 58
  }
}
 

每个答案都是这份简历属于该条记录所对应之人的概率。记录 18 的姓名拼写不同,但地点和雇主一致,得分为 0.74。记录 42 姓名相同,但在不同城市、雇主也不同,得分为 0.09。记录 77 姓名相似、地点相同,但雇主不同,得分为 0.08。像在代码中处理多个 Noul 答案那样,在你的代码中为每个值设定阈值,并把中间值交给人工。

使用 Python SDK 时,问题由候选人记录构建。问题文本固定,记录变化:

python
from typesafe_sdk import Noul, TypeSafeClient
 
SAME_PERSON = "Is the resume for the same person as `potential_duplicate`?"
 
 
def duplicate_questions(candidates: list[dict]) -> dict[str, Noul]:
    """One Noul per candidate record, all asking the same question."""
    return {
        f"same_as_record_{candidate['id']}": Noul(
            instructions={
                "potential_duplicate": {
                    "name": candidate["name"],
                    "location": candidate["location"],
                    "last_employer": candidate["last_employer"],
                },
                "question": SAME_PERSON,
            },
        )
        for candidate in candidates
    }
 
 
def find_duplicates(resume: dict, candidates: list[dict]) -> list[str]:
    with TypeSafeClient() as client:
        response = client.system_one(
            model="jev-latest",
            state={"resume": resume},
            questions=duplicate_questions(candidates),
        )
    return [
        question_id
        for question_id, answer in response.answers.items()
        if answer.noul > 0.7
    ]
 

结构化数据抽取级联实践手册使用结构化 instructions 来验证抽取出的记录。每个字段都得到同一组问题。每个问题的 instructions 对象把问题文本放在 main_question 属性中。此外还有 field_spec 和 extracted_field 属性,它们随每个字段而变化。

实践手册中的 Noul

看看我们的实践手册,了解使用 Noul 问题的应用:

  • 并行问题在一次请求中对一篇文章运行一个包含 13 个问题的合规检查清单。
  • 自洽性:nouls依据一个 15 个问题的评分量表为一份保险理赔评分,并衡量这些值在多次运行之间的稳定程度。
  • 重排序直接使用概率本身,而不是阈值:每个「查询—候选」对一个 Noul,然后按该值对候选排序。
  • 逐行搜索把一个用于找出匹配行的 Choice 与一个检查文档中是否存在答案的 Noul 配对使用。
  • 结构还原对每对行问一个 Noul——换行是否把一个句子拆开了——从而从纯文本重建段落。