TS TypeSafe 文档中文版 原文 ↗

如何使用 TypeSafe 构建

设计 AI 驱动的软件时,让代码保持掌控,并交给 System One 做狭窄而结构化的决策。

System One 是 TypeSafe 用于构建 AI 驱动软件(而非智能体)的模型。它不生成代码,也不自行选择下一步动作。它提供可嵌入软件中的 AI 原语,因此代码始终保持掌控,而模型负责对非结构化数据做出常识性判断。

说明

概要: 构建一个普通的软件工作流,只在需要 AI 的地方插入 System One。

  • 把控制流、确定性规则和副作用保留在代码中。
  • 把宽泛的判断拆解为狭窄、带类型的问题,并配有明确的 instructions 和 criteria。
  • 只给每个问题提供它所需的上下文。
  • 用概率和置信度来决定是直接行动、请求复核,还是上报升级。
  • 把相互独立的问题放在一起提问,然后在代码中组合它们的答案。

三种软件架构

TypeSafe 是为构建 AI 驱动的软件 而设计的:在这类软件中,代码拥有工作流,AI 负责狭窄而结构化的决策。

传统代码是由简单软件原语构成的复杂决策树。由于每个原语都可靠,开发者可以把它们组合成更高层的抽象。

智能体会处理指令并选择自己的下一步。当有人监控整个过程时,这种方式运转良好,但每一次循环都多了一次偏离正轨的机会。

代码负责确定性工作并拥有控制流。模型只出现在系统需要可编程常识或需要解读非结构化数据的地方。每个 AI 任务都保持原子化并受到约束。

传统软件、智能体和 AI 驱动的软件,展示为三种不同的系统架构。 传统软件、智能体和 AI 驱动的软件,展示为三种不同的系统架构。

是什么让 System One 可组合

结构化

System One 在构造上就是类型安全的。决策和概率符合你的代码所期望的结构化软件类型与 JSON schema,因此它从不需要从生成的散文中还原取值。

并行

各个问题会被独立且并行地评估。一个原语的结果不会变成改变另一个原语结果的隐藏上下文。

可比较

输出可以排序,并能驱动智能的 if 语句、阈值和比较。

快速

大多数查询在约 100 毫秒内完成。System One 足够快,可用于实时请求路径和用户界面。

校准过的置信度

RLCD 通过校准过的概率来表达不确定性,而不是倾向于过度自信。

自洽

System One 的设计目标是在反复评估中返回稳定的答案。参见自洽性实践手册。

由于每个输出都被约束在所提供的选项之内,模型会返回这些选项上的完整概率分布,而不是凭空造出 schema 之外的值。TypeSafe 的目标是实现超过 100 倍的智能与速度及成本之比;其底层押注是:更便宜的智能会创造出大得多的需求。

设计一个 System One 工作流

  1. 能用代码就用代码

    把确定性工作保留在代码中。它可靠且廉价。当软件工作流能表达同样的行为时,避免使用智能体的 while 循环。

    示例:把确定性规则保留在代码中
    python
    days_overdue = (today - invoice.due_date).days
     
    if days_overdue > 30:
        route_to_collections(invoice)
     

    浏览 System One 模式,了解把模型决策与代码组合起来的有界做法。

  2. 拆解输入状态

    只包含与当前问题相关的上下文。这有助于模型避免干扰和上下文腐化。当最新信息可以来自你自己的知识库时,不要依赖存储在模型权重中的知识。

    示例:只发送相关的上下文
    request
    {
      "state": {
        "ticket_message": "My flight was cancelled. Can I get a refund?",
        "refund_policy": "Cancelled flights are eligible for a full refund."
      },
      "selectedModels": [
        "jev-latest"
      ],
      "questions": {
        "policy_supports_refund": {
          "type": "noul",
          "instructions": "Does the refund policy support the refund requested in the ticket?"
        }
      }
    }

    这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

  3. 在输入状态中使用结构

    为 state 和 questions 字段使用嵌套 JSON。当这样能消除歧义时,让问题指向具体的值,并在问题内部为每个路径加上反引号。

    示例:引用嵌套值

    使用带反引号的点号加索引路径,让问题指向某个具体的嵌套值,例如 support.tickets[0].message。

    request
    {
      "state": {
        "support": {
          "tickets": [
            {
              "message": "I was charged twice for order A-104."
            },
            {
              "message": "How do I reset my password?"
            }
          ]
        },
        "commerce": {
          "orders": [
            {
              "id": "A-104",
              "charges": [
                {
                  "amount_usd": 49,
                  "status": "captured"
                },
                {
                  "amount_usd": 49,
                  "status": "captured"
                }
              ]
            }
          ]
        },
        "account": {
          "security": {
            "password_reset": "Email a reset link to the address on file."
          }
        }
      },
      "selectedModels": [
        "jev-latest"
      ],
      "questions": {
        "duplicate_charge": {
          "type": "noul",
          "instructions": "Do `support.tickets[0].message` and `commerce.orders[0].charges` indicate a duplicate charge?"
        },
        "password_reset_supported": {
          "type": "noul",
          "instructions": "Can `account.security.password_reset` resolve the request in `support.tickets[1].message`?"
        }
      }
    }

    这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

  4. 拆解问题

    尽可能提出最明确、最狭窄、最具体、最原子化的问题。把复杂或定义不清的问题拆成多个独立问题,每个问题只评估一个属性。

    说明

    这大概是本指南中最重要的概念。宽泛的问题会把多个判断藏在同一个答案背后。原子化的问题会把这些判断暴露出来,让你能在代码中检查、调整和组合它们。

    示例:拆解垃圾信息检测
    One broad question (bad)
    {
      "state": {
        "message": {
          "sender": {
            "display_name": "Acme Payroll",
            "email": "rewards@claim-bonus.example"
          },
          "subject": "Urgent: claim your employee bonus",
          "body": "You have been selected for a $1,000 bonus. Confirm your payroll password today to receive it.",
          "links": [
            {
              "text": "Claim bonus",
              "url": "http://claim-bonus.example/acme"
            }
          ]
        }
      },
      "selectedModels": [
        "jev-latest"
      ],
      "questions": {
        "is_spam": {
          "type": "noul",
          "instructions": "Is `message` spam?"
        }
      }
    }

    这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

    Decomposed questions (good)
    {
      "state": {
        "message": {
          "sender": {
            "display_name": "Acme Payroll",
            "email": "rewards@claim-bonus.example"
          },
          "subject": "Urgent: claim your employee bonus",
          "body": "You have been selected for a $1,000 bonus. Confirm your payroll password today to receive it.",
          "links": [
            {
              "text": "Claim bonus",
              "url": "http://claim-bonus.example/acme"
            }
          ]
        }
      },
      "selectedModels": [
        "jev-latest"
      ],
      "questions": {
        "requests_credentials": {
          "type": "noul",
          "instructions": "Does `message.body` ask the recipient to provide a password or other login credential?"
        },
        "offers_unexpected_reward": {
          "type": "noul",
          "instructions": "Does `message.body` claim the recipient received an unexpected prize, payment, or reward?"
        },
        "creates_time_pressure": {
          "type": "noul",
          "instructions": "Does `message.subject` or `message.body` pressure the recipient to act quickly?"
        },
        "sender_identity_mismatch": {
          "type": "noul",
          "instructions": "Does the organization named in `message.sender.display_name` conflict with the domain in `message.sender.email`?"
        },
        "link_domain_mismatch": {
          "type": "noul",
          "instructions": "Does the domain in `message.links[0].url` conflict with the organization named in `message.sender.display_name`?"
        },
        "disguises_link_destination": {
          "type": "noul",
          "instructions": "Does `message.links[0].text` conceal or misrepresent the destination in `message.links[0].url`?"
        }
      }
    }

    这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

    示例:核查工具调用轨迹
    One broad question (bad)
    {
      "state": {
        "request": {
          "text": "What's the weather in Seattle tomorrow in Fahrenheit?",
          "location": "Seattle, WA",
          "date": "2026-09-03",
          "unit": "fahrenheit"
        },
        "available_tools": {
          "geocode_city": {
            "description": "Resolve a city to latitude and longitude.",
            "parameters": {
              "city": "string"
            }
          },
          "get_weather": {
            "description": "Get the forecast for coordinates and a date.",
            "parameters": {
              "latitude": "number",
              "longitude": "number",
              "date": "YYYY-MM-DD",
              "unit": [
                "fahrenheit",
                "celsius"
              ]
            }
          }
        },
        "trace": {
          "tool_calls": [
            {
              "id": "call_1",
              "name": "geocode_city",
              "arguments": {
                "city": "Seattle, WA"
              }
            },
            {
              "id": "call_2",
              "name": "get_weather",
              "arguments": {
                "latitude": 47.6062,
                "longitude": -122.3321,
                "date": "2026-09-03",
                "unit": "celsius"
              }
            }
          ],
          "tool_results": [
            {
              "tool_call_id": "call_1",
              "output": {
                "latitude": 47.6062,
                "longitude": -122.3321
              }
            }
          ]
        }
      },
      "selectedModels": [
        "jev-latest"
      ],
      "questions": {
        "tool_calls_are_correct": {
          "type": "noul",
          "instructions": "Is `trace.tool_calls` correct for `request` and `available_tools`?"
        }
      }
    }

    这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

    Decomposed questions (good)
    {
      "state": {
        "request": {
          "text": "What's the weather in Seattle tomorrow in Fahrenheit?",
          "location": "Seattle, WA",
          "date": "2026-09-03",
          "unit": "fahrenheit"
        },
        "available_tools": {
          "geocode_city": {
            "description": "Resolve a city to latitude and longitude.",
            "parameters": {
              "city": "string"
            }
          },
          "get_weather": {
            "description": "Get the forecast for coordinates and a date.",
            "parameters": {
              "latitude": "number",
              "longitude": "number",
              "date": "YYYY-MM-DD",
              "unit": [
                "fahrenheit",
                "celsius"
              ]
            }
          }
        },
        "trace": {
          "tool_calls": [
            {
              "id": "call_1",
              "name": "geocode_city",
              "arguments": {
                "city": "Seattle, WA"
              }
            },
            {
              "id": "call_2",
              "name": "get_weather",
              "arguments": {
                "latitude": 47.6062,
                "longitude": -122.3321,
                "date": "2026-09-03",
                "unit": "celsius"
              }
            }
          ],
          "tool_results": [
            {
              "tool_call_id": "call_1",
              "output": {
                "latitude": 47.6062,
                "longitude": -122.3321
              }
            }
          ]
        }
      },
      "selectedModels": [
        "jev-latest"
      ],
      "questions": {
        "geocode_tool_is_relevant": {
          "type": "noul",
          "instructions": "Is `trace.tool_calls[0].name` an appropriate tool for resolving `request.location`?"
        },
        "geocode_location_matches": {
          "type": "noul",
          "instructions": "Does `trace.tool_calls[0].arguments.city` match `request.location`?"
        },
        "geocode_arguments_match_schema": {
          "type": "noul",
          "instructions": "Does `trace.tool_calls[0].arguments` conform to `available_tools.geocode_city.parameters`?"
        },
        "geocode_result_matches_call": {
          "type": "noul",
          "instructions": "Does `trace.tool_results[0].tool_call_id` match `trace.tool_calls[0].id`?"
        },
        "weather_tool_is_relevant": {
          "type": "noul",
          "instructions": "Is `trace.tool_calls[1].name` an appropriate tool for answering `request.text`?"
        },
        "weather_arguments_match_schema": {
          "type": "noul",
          "instructions": "Does `trace.tool_calls[1].arguments` conform to `available_tools.get_weather.parameters`?"
        },
        "weather_uses_geocoded_coordinates": {
          "type": "noul",
          "instructions": "Do the coordinates in `trace.tool_calls[1].arguments` match those in `trace.tool_results[0].output`?"
        },
        "weather_date_matches": {
          "type": "noul",
          "instructions": "Does `trace.tool_calls[1].arguments.date` match `request.date`?"
        },
        "weather_unit_matches": {
          "type": "noul",
          "instructions": "Does `trace.tool_calls[1].arguments.unit` match `request.unit`?"
        }
      }
    }

    这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

  5. 在问题中使用结构

    让问题保持简短。instructions 和 criteria 通常是字符串,对于简短、无歧义的问题,一个字符串就够了。它们也可以是对象或数组。把问题放在一个字段里,把引导该问题的数据放在其他字段里。

    结构在以下情形中会有帮助:

    • 问题需要上下文或示例。一长句背景信息或一组示例输入应当放在问题旁边的具名字段中,这样你的代码就能增补或替换它们,而无需重写问题。
    • 问题的一部分来自你的代码。当某个值来自数据库时,把它放进自己的字段,而不是拼接进字符串模板。
    • 多个问题有相似的 instructions。一个请求接收一个状态,并且可以包含多个问题。添加补充数据有助于让问题彼此区分。
    示例:引用来自你的代码的记录

    这个 Noul 会把状态中的一份简历与候选人数据库中的一条记录进行比较。该记录按原样放进 potential_duplicate,问题则按名称引用它。

    questions
    {
      "state": {
        "resume": {
          "name": "John Smith",
          "location": "Oakland, CA",
          "summary": "Backend engineer with eight years of Python and Go experience.",
          "experience": [
            {
              "employer": "Google",
              "title": "Senior Backend Engineer",
              "years": "2021-2025"
            },
            {
              "employer": "Microsoft",
              "title": "Software Engineer",
              "years": "2017-2021"
            }
          ]
        }
      },
      "selectedModels": [
        "jev-latest"
      ],
      "questions": {
        "same_as_record_18": {
          "type": "noul",
          "instructions": {
            "potential_duplicate": {
              "name": "John Smith",
              "location": "Oakland, California",
              "last_employer": "Google"
            },
            "question": "Is the resume for the same person as `potential_duplicate`?"
          }
        }
      }
    }

    这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

    来源为代码的 "potential_duplicate" 数据可能随时间变化。"question" 使用反引号引用它。

    criteria 内部的描述也可以是对象。对于 Choice,每个选项的描述都可以是一个对象,说明该选项涵盖什么、什么属于另一个选项,以及几个示例。在各选项之间使用相同的字段名,这样模型就能直接比较它们。

    示例:定义对比式的 Choice criteria
    questions
    {
      "state": "How many disposable virtual cards can I make per day?",
      "selectedModels": [
        "jev-latest"
      ],
      "questions": {
        "card_help_topic": {
          "type": "choice",
          "instructions": {
            "question": "Which disposable virtual card topic is the user asking about?",
            "focus": "Classify the information the user wants."
          },
          "criteria": {
            "get_disposable_virtual_card": {
              "what": "Purpose, eligibility, or setup",
              "not_for": "Quantity, transaction, or merchant restrictions",
              "examples": [
                "How can I get a disposable virtual card?",
                "What are disposable cards for?"
              ]
            },
            "disposable_card_limits": {
              "what": "Quantity, transaction, or merchant restrictions",
              "not_for": "Purpose, eligibility, or setup",
              "examples": [
                "How many disposable cards can I make per day?",
                "Where can I use a disposable card?"
              ]
            }
          }
        }
      }
    }

    这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。

    每种问题类型的页面都有一个完整示例:

    • Noul 会把一份简历与多条候选人记录进行比较,每条记录一个问题,问题在代码中构建。
    • Choice 用「每个选项涵盖什么、不适用于什么、以及示例」来描述两个容易混淆的选项。
    • Score 为每个等级给出描述和示例情形。

    结构化数据抽取级联实践手册 展示了共享措辞的情形:对一条抽取记录的每个字段提出同一组问题。

    简短、无歧义的问题或 criteria 可以继续用字符串。当结构能把本会糊在一起的指引分开时,就加上结构。关于接受结构的全部位置,参见进阶:结构。

  6. 提出大量问题

    在同一个请求中针对同一个状态提出许多狭窄、独立的问题。这就是用该 API 最大化每美元效能与智能的方式:问题会并行运行,代码可以组合它们的信号,而无需增加串行的模型往返。

    参见推测式扇出模式和并行问题实践手册。

  7. 在代码中组合问题输出(或喂给经典 ML 模型)

    用确定性规则或加权和来组合相互独立的答案。若要做学习式组合,可以把这些概率作为特征,输入下游的经典机器学习模型。

    示例:用加权分数组合信号
    python
    answers = response.answers
     
    # Combine independent signals into one application-specific score.
    quality = (
        0.4 * answers["answers_request"].noul
        + 0.4 * answers["citations_are_supported"].noul
        + 0.2 * (1 - answers["contradicts_context"].noul)
    )
     

    复合评分 展示了如何在组合各个判断的同时保留它们。如果你没有为下游模型准备标签,可以用一组昂贵的推理模型来生成标签;AutoResearch 实践手册 展示了如何在 System One 的输出上训练一个经典模型。

  8. 按不确定性路由

    让代码对置信度高和置信度低的答案采取不同动作。把不确定的案例上报给人工或更昂贵的推理模型。通过在你的数据上绘制置信度与准确率的关系来测试阈值。

    示例:按置信度路由
    python
    answer = response.answers["card_help_topic"]
     
    if answer.confidence < 0.8:
        route_to_human_review(ticket)
    else:
        route_to_handler(answer.choice, ticket)
     

    关于如何选择阈值并使其与每个动作的风险相匹配,参见置信度和置信度门控路由。

提示

拆解并不需要更多的往返。针对同一个状态的问题会并行运行。

把它们组合起来

这个支持工单工作流把确定性工作保留在代码中,只发送相关的结构化上下文,在同一个请求中评估许多原子化问题,并用明确的置信度门控来组合答案。

triage_ticket.py
from typesafe_sdk import Choice, Noul, NoulCriteria, Score, TypeSafeClient
 
 
def triage_ticket(ticket, customer):
    # Handle deterministic states without calling a model.
    if ticket["status"] == "closed":
        return "no_action"
 
    open_orders = [
        order for order in customer["orders"] if order["status"] != "delivered"
    ]
 
    # Include only the structured context needed by the questions below.
    state = {
        "ticket": {
            "message": ticket["message"],
            "sender": ticket["sender"],
            "links": ticket["links"],
        },
        "customer": {
            "plan": customer["plan"],
            "open_orders": open_orders,
        },
        "policy": {
            "sensitive_credentials": ["password", "security code", "API key"],
        },
    }
 
    # Ask structured, atomic questions together so they run in parallel.
    questions = {
        "topic": Choice(
            instructions={
                "question": "Which team should handle `ticket.message`?",
                "focus": "Classify the customer's primary request.",
            },
            criteria={
                "billing": {
                    "what": "Charges, invoices, refunds, or subscriptions",
                    "not_for": "Order tracking or account access",
                    "examples": ["I was charged twice", "Where is my refund?"],
                },
                "orders": {
                    "what": "Order status, delivery, cancellation, or returns",
                    "not_for": "Charges or account access",
                    "examples": ["Where is my order?", "Cancel my shipment"],
                },
                "account": {
                    "what": "Login, profile, permissions, or security",
                    "not_for": "Charges or order tracking",
                    "examples": ["Reset my password", "I cannot sign in"],
                },
            },
        ),
        "requests_credentials": Noul(
            instructions={
                "question": "Does the message request a sensitive credential?",
                "compare": [
                    "`ticket.message`",
                    "`policy.sensitive_credentials`",
                ],
                "focus": "Look for a request to disclose the credential itself.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Asks the recipient to disclose a listed credential",
                    "examples": [
                        "Reply with your password",
                        "Send us your API key",
                    ],
                },
                false={
                    "what": "Does not ask the recipient to disclose a credential",
                    "not_for": "A legitimate instruction to reset a credential",
                    "examples": ["Use this link to reset your password"],
                },
            ),
        ),
        "sender_identity_mismatch": Noul(
            instructions={
                "question": "Does the claimed sender identity conflict with its domain?",
                "compare": [
                    "`ticket.sender.display_name`",
                    "`ticket.sender.email`",
                ],
                "focus": "Compare the named organization with the email domain.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Claims an organization unrelated to the email domain",
                    "examples": ["Acme Payroll sent from claim-bonus.example"],
                },
                false={
                    "what": "The identity and domain agree or make no conflicting claim",
                    "examples": ["Acme Payroll sent from acme.example"],
                },
            ),
        ),
        "unexpected_reward": Noul(
            instructions={
                "question": "Does the message announce an unexpected reward?",
                "inspect": "`ticket.message`",
                "focus": "Look for an unsolicited prize, payment, or reward claim.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Announces an unrequested prize, payment, or reward",
                    "examples": ["You were selected for a $1,000 bonus"],
                },
                false={
                    "what": "Contains no reward claim or discusses an expected payment",
                    "not_for": "A customer asking about a known refund or payroll deposit",
                    "examples": ["When will my approved refund arrive?"],
                },
            ),
        ),
        "refund_requested": Noul(
            instructions={
                "question": "Does the customer explicitly request a refund or credit?",
                "inspect": "`ticket.message`",
                "focus": "Require a requested remedy, not a billing complaint alone.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Directly asks for money back or an account credit",
                    "examples": ["Please refund the duplicate charge"],
                },
                false={
                    "what": "Does not ask for a refund or credit",
                    "not_for": "A complaint or billing question without a requested remedy",
                    "examples": ["Why was I charged twice?"],
                },
            ),
        ),
        "mentions_open_order": Noul(
            instructions={
                "question": "Does the message refer to a supplied open order?",
                "compare": [
                    "`ticket.message`",
                    "`customer.open_orders`",
                ],
                "focus": "Match an order id or other identifying details.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Refers to an open order by id or identifying details",
                    "examples": ["Where is order A-104?"],
                },
                false={
                    "what": "Does not identify any supplied open order",
                    "not_for": "A generic order question with no matching details",
                    "examples": ["How long does shipping usually take?"],
                },
            ),
        ),
        "frustration": Score(
            instructions={
                "question": "How frustrated does the customer appear?",
                "inspect": "`ticket.message`",
                "focus": "Judge expressed frustration, not issue severity.",
            },
            criteria=[
                {
                    "what": "Calm and matter-of-fact",
                    "signals": ["Neutral wording", "No complaint about the experience"],
                },
                {
                    "what": "Frustrated but civil",
                    "signals": ["Expresses annoyance", "Remains constructive"],
                },
                {
                    "what": "Very angry or threatening to leave",
                    "signals": ["Hostile language", "Threatens cancellation or churn"],
                },
            ],
        ),
    }
 
    with TypeSafeClient() as client:
        response = client.system_one(
            state=state,
            questions=questions,
        )
 
    # Compose independent spam signals with weights controlled by code.
    answers = response.answers
    spam_risk = (
        0.45 * answers["requests_credentials"].noul
        + 0.30 * answers["sender_identity_mismatch"].noul
        + 0.25 * answers["unexpected_reward"].noul
    )
 
    # Escalate uncertain judgments instead of guessing.
    spam_is_uncertain = 0.4 < spam_risk < 0.6
    if spam_is_uncertain or answers["topic"].confidence < 0.75:
        return route_to_human_review(ticket)
    if spam_risk >= 0.6:
        return quarantine_as_spam(ticket)
 
    # Let code decide which speculative answers matter on this path.
    if answers["topic"].choice == "billing":
        return route_to_billing(
            ticket,
            refund_requested=answers["refund_requested"].noul >= 0.7,
        )
    if answers["topic"].choice == "orders":
        return route_to_orders(
            ticket,
            mentions_open_order=answers["mentions_open_order"].noul >= 0.7,
        )
 
    priority = (
        "high"
        if answers["frustration"].confidence >= 0.7
        and answers["frustration"].score >= 1.5
        else "normal"
    )
    return route_to_account_support(ticket, priority=priority)