TS TypeSafe 文档中文版 原文 ↗

逐行搜索

为 GitHub 的服务条款构建语义搜索。在一次请求中,用一个 Choice 问题针对一段自然语言查询给 218 个行 id 打分,并用一个 Noul 问题检查文档中是否包含答案。

你手上有 GitHub 的服务条款,以及一个关于它的自然语言问题。你需要找到能回答这个问题的 那些行,还需要一种方法来检测文档在什么时候没有答案。随附的查询会把带直接答案的行 排在最前面。exists 阈值则把其余情况归类为缺失或部分。最终你得到 find(),它返回 exists 概率以及每一行的一个相关性分数。

一个查询扫描文档,并揭示附着在匹配行上的答案

搜索后端由三部分组成:

  1. 给每一行打上一个 ID 标签,以便 TypeSafe 能指向它。
  2. 用一个 Choice 问题,按这些行 ID 回答该查询的好坏程度对它们排序。Choice 问题的概率总和始终为 1,所以即使没有任何一行回答了查询,也总有一行排在 第一位。
  3. 在同一次请求中,用一个 Noul 问题检查文档中到底有没有答案。

环境准备

获取 TypeSafe API 密钥

在 TypeSafe 控制台创建一个密钥并导出它:

bash
export TYPESAFE_API_KEY="your-key-here"
 

安装依赖

bash
pip install "typesafe-sdk>=0.5.7" cooksafe \
  --extra-index-url https://pypi.typesafe.ai/
 

JsonCache 会重放随附的 API 响应,因此下面的步骤无需 API 密钥、也不产生任何花费即可 运行。若要让请求真正发出,请设置 TYPESAFE_API_KEY 并删除 json_cache.json。

创建脚本

从导入和客户端开始编写 semantic_search.py:

python
import os
import urllib.request
from pathlib import Path
 
from cooksafe import JsonCache
from typesafe_sdk import Choice, Noul, NoulCriteria, TypeSafeClient
 
TYPESAFE_MODEL = "jev-1.12"
 
client = TypeSafeClient(
    api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"), timeout=120.0
)
json_cache = JsonCache(Path("json_cache.json"))
 

第 1 步:给每一行打上 ID

测试文档是 GitHub 的服务条款,被切分成 218 个条款,因此每个搜索结果都指向一行 可引用的文本。

向 semantic_search.py 中添加:

python
GIST = (
    "https://gist.githubusercontent.com/eugene-shvarts/900632789a24983d5678ffd508dd01f6"
    "/raw/cf9c2ab422d568deade949ef0a06bed6896964b9/github-tos.txt"
)
 
 
@json_cache
def fetch_document(url: str) -> str:
    request = urllib.request.Request(
        url, headers={"User-Agent": "typesafe-cookbook/1.0"}
    )
    with urllib.request.urlopen(request) as response:
        return response.read().decode()
 
 
LINES = fetch_document(GIST).splitlines()
 

缓存避免了重复下载,splitlines() 留下一个包含 218 个字符串的列表。

现在给每一行加上一个短 ID 前缀,再把这些行拼回一个文档。模型用这些 ID 来指向 它的答案。

python
def line_id(i: int) -> str:
    return f"L{i:03d}"
 
 
DOCUMENT = "\n".join(f"{line_id(i)}| {line}" for i, line in enumerate(LINES))
 

DOCUMENT 现在看起来是这样:

text
L052| You own Your Content. If you post Content you did not create, you are responsible for...
L053| You grant us and other Users the licenses in Sections D.4–D.8. These licenses apply...
L054| 4. License Grant to Us
 

第 2 步:询问答案在哪里

一个 Choice 问题会为每个选项返回一个概率。把行 ID 用作 选项, 「挑一个选项」就变成了「指向一行」。

python
def where_question(query: str) -> Choice:
    return Choice(
        instructions=f'Which line of the document contains the answer to: "{query}"?',
        criteria={line_id(i): None for i in range(len(LINES))},
    )
 

选项描述是 None,因为文档本身已经包含了每个 ID 对应的文本。查询放在 instructions 里;状态在各次搜索之间 保持不变。

说明

行窗口,第二个问题在该窗口内对这些行排序。

第 3 步:检查是否存在答案

Choice 概率的总和始终为 1,所以即使文档没有回答该问题, 也总有一行排在第一位。仅凭排序无法区分真正的答案和最接近的 不相关行。

所以在同一次请求中再问一个问题:

python
def exists_question(query: str) -> Noul:
    return Noul(
        instructions=f'Does any line of the document address or answer: "{query}"?',
        criteria=NoulCriteria(
            true="At least one line of the document states or directly implies the answer",
            false="No line of the document addresses this",
        ),
    )
 

与 Choice 概率不同,Noul 概率不依赖其他选项, 所以当文档没有答案时,它可以降到接近零。

第 4 步:在一次请求中发送两个问题

system_one 方法在一次传递中回答两个问题。状态只发送一次,所以 加上存在性检查只需很少的额外输出。

一个带标签的文档和用户问题进入一次 TypeSafe 请求。一个 Choice 问题给
每一行打分,同时一个 Noul 问题检查是否存在答案。随后本地代码对
这些行排序并应用文档的判定结果。

python
@json_cache
def _find(
    model: str,
    state: str,
    where: Choice,
    exists: Noul,
) -> dict:
    response = client.system_one(
        state=state,
        questions={"where": where, "exists": exists},
        model=model,
    )
    probabilities = response.answers["where"].probabilities
    return {
        "exists": response.answers["exists"].noul,
        "relevance": [probabilities.get(line_id(i), 0.0) for i in range(len(LINES))],
    }
 
 
def find(query: str) -> dict:
    return _find(
        TYPESAFE_MODEL,
        DOCUMENT,
        where_question(query),
        exists_question(query),
    )
 

relevance 列表按文档顺序为每一行保留一个分数。

第 5 步:读取结果

两段本地代码完成收尾:verdict() 把原始的 exists 概率 转换为三种状态,中间一种对应部分回答;show() 把 relevance 渲染成条形图,让排序在终端里可读。

python
FOUND, ABSENT = 0.7, 0.35  # present answers typically read >=0.9, absent <=0.05
 
 
def verdict(exists: float) -> str:
    if exists >= FOUND:
        return "answered in this document"
    return "not in this document" if exists < ABSENT else "partially addressed"
 
 
def show(query: str, top: int = 4) -> dict:
    result = find(query)
    print(f'"{query}"')
    print(f"  exists {result['exists']:.2f} -> {verdict(result['exists'])}")
    ranked = sorted(
        range(len(LINES)), key=lambda i: result["relevance"][i], reverse=True
    )
    for i in ranked[:top]:
        bar = "#" * max(1, round(result["relevance"][i] * 12))
        preview = LINES[i][:58].rstrip()
        print(f"  {line_id(i)}  {result['relevance'][i]:.2f}  {bar:<12}  {preview}")
    return result
 

这些阈值能区分下面的示例,但在生产环境中使用之前, 请针对你自己的文档调整它们。

问两个有直接答案的问题、一个没有答案的问题,以及一个有 部分答案的问题,总共四个。

python
print(f"{len(LINES)} lines, {len(DOCUMENT):,} characters\n")
show("who owns the code I upload?")
print()
show("can GitHub kick me off the platform without warning?")
print()
show("do I have to take disputes to arbitration?", top=2)
print()
show("can minors use GitHub with parental permission?", top=2)
 
text
218 lines, 43,980 characters
 
"who owns the code I upload?"
  exists 0.98 -> answered in this document
  L052  0.95  ###########   You own Your Content. If you post Content you did not crea
  L046  0.02  #             Short version: You own content you create, but you allow u
  L051  0.02  #             3. Ownership and License Grants
  L217  0.01  #             Questions about the Terms of Service? Contact us through t
 
"can GitHub kick me off the platform without warning?"
  exists 0.97 -> answered in this document
  L168  0.97  ############  GitHub has the right to suspend or terminate your access t
  L167  0.03  #             3. GitHub May Terminate
  L000  0.00  #             Effective date: April 27, 2026 · A. Definitions
  L001  0.00  #             Short version: We use these basic terms throughout the agr
 
"do I have to take disputes to arbitration?"
  exists 0.14 -> not in this document
  L205  0.86  ##########    Except to the extent applicable law provides otherwise, th
  L168  0.02  #             GitHub has the right to suspend or terminate your access t
 
"can minors use GitHub with parental permission?"
  exists 0.46 -> partially addressed
  L029  0.90  ###########   You must be age 13 or older. While we are thrilled to see
  L012  0.07  #             “User,” “You,” and “Your” refer to the individual person,
 

这些分数意味着什么

前两个查询返回直接答案,以及验证这些答案所需的源行。

另外两个则说明为什么存在性检查很重要:

  • 仲裁: 排序给最接近的那一行打了 0.86 分,但 exists 只有 0.14。答案并不在文档中。
  • 父母许可: 年龄规则排在第一,但它并没有回答父母 许可是否会改变这条规则。结果是部分涉及。

排序告诉你该去哪里找;exists 分数告诉你这个结果 是否回答了该问题。

在你自己的文档上试试

在 TypeSafe Playground 中打开这份带标签的合同,针对同一段文本编辑 问题。若要搜索你自己的文档,替换 fetch_document() 中的 URL;脚本的其余每一 行都基于 LINES 工作。