如何使用 TypeSafe 构建
设计 AI 驱动的软件时,让代码保持掌控,并交给 System One 做狭窄而结构化的决策。
System One 是 TypeSafe 用于构建 AI 驱动软件(而非智能体)的模型。它不生成代码,也不自行选择下一步动作。它提供可嵌入软件中的 AI 原语,因此代码始终保持掌控,而模型负责对非结构化数据做出常识性判断。
概要: 构建一个普通的软件工作流,只在需要 AI 的地方插入 System One。
- 把控制流、确定性规则和副作用保留在代码中。
- 把宽泛的判断拆解为狭窄、带类型的问题,并配有明确的 instructions 和 criteria。
- 只给每个问题提供它所需的上下文。
- 用概率和置信度来决定是直接行动、请求复核,还是上报升级。
- 把相互独立的问题放在一起提问,然后在代码中组合它们的答案。
三种软件架构
TypeSafe 是为构建 AI 驱动的软件 而设计的:在这类软件中,代码拥有工作流,AI 负责狭窄而结构化的决策。
传统代码是由简单软件原语构成的复杂决策树。由于每个原语都可靠,开发者可以把它们组合成更高层的抽象。
智能体会处理指令并选择自己的下一步。当有人监控整个过程时,这种方式运转良好,但每一次循环都多了一次偏离正轨的机会。
代码负责确定性工作并拥有控制流。模型只出现在系统需要可编程常识或需要解读非结构化数据的地方。每个 AI 任务都保持原子化并受到约束。
是什么让 System One 可组合
System One 在构造上就是类型安全的。决策和概率符合你的代码所期望的结构化软件类型与 JSON schema,因此它从不需要从生成的散文中还原取值。
各个问题会被独立且并行地评估。一个原语的结果不会变成改变另一个原语结果的隐藏上下文。
输出可以排序,并能驱动智能的 if 语句、阈值和比较。
大多数查询在约 100 毫秒内完成。System One 足够快,可用于实时请求路径和用户界面。
RLCD 通过校准过的概率来表达不确定性,而不是倾向于过度自信。
System One 的设计目标是在反复评估中返回稳定的答案。参见自洽性实践手册。
由于每个输出都被约束在所提供的选项之内,模型会返回这些选项上的完整概率分布,而不是凭空造出 schema 之外的值。TypeSafe 的目标是实现超过 100 倍的智能与速度及成本之比;其底层押注是:更便宜的智能会创造出大得多的需求。
设计一个 System One 工作流
- 能用代码就用代码
把确定性工作保留在代码中。它可靠且廉价。当软件工作流能表达同样的行为时,避免使用智能体的
while循环。示例:把确定性规则保留在代码中
python days_overdue = (today - invoice.due_date).days if days_overdue > 30: route_to_collections(invoice)浏览 System One 模式,了解把模型决策与代码组合起来的有界做法。
- 拆解输入状态
只包含与当前问题相关的上下文。这有助于模型避免干扰和上下文腐化。当最新信息可以来自你自己的知识库时,不要依赖存储在模型权重中的知识。
示例:只发送相关的上下文
request { "state": { "ticket_message": "My flight was cancelled. Can I get a refund?", "refund_policy": "Cancelled flights are eligible for a full refund." }, "selectedModels": [ "jev-latest" ], "questions": { "policy_supports_refund": { "type": "noul", "instructions": "Does the refund policy support the refund requested in the ticket?" } } }这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。
- 在输入状态中使用结构
为
state和questions字段使用嵌套 JSON。当这样能消除歧义时,让问题指向具体的值,并在问题内部为每个路径加上反引号。示例:引用嵌套值
使用带反引号的点号加索引路径,让问题指向某个具体的嵌套值,例如
support.tickets[0].message。request { "state": { "support": { "tickets": [ { "message": "I was charged twice for order A-104." }, { "message": "How do I reset my password?" } ] }, "commerce": { "orders": [ { "id": "A-104", "charges": [ { "amount_usd": 49, "status": "captured" }, { "amount_usd": 49, "status": "captured" } ] } ] }, "account": { "security": { "password_reset": "Email a reset link to the address on file." } } }, "selectedModels": [ "jev-latest" ], "questions": { "duplicate_charge": { "type": "noul", "instructions": "Do `support.tickets[0].message` and `commerce.orders[0].charges` indicate a duplicate charge?" }, "password_reset_supported": { "type": "noul", "instructions": "Can `account.security.password_reset` resolve the request in `support.tickets[1].message`?" } } }这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。
- 拆解问题
尽可能提出最明确、最狭窄、最具体、最原子化的问题。把复杂或定义不清的问题拆成多个独立问题,每个问题只评估一个属性。
说明这大概是本指南中最重要的概念。宽泛的问题会把多个判断藏在同一个答案背后。原子化的问题会把这些判断暴露出来,让你能在代码中检查、调整和组合它们。
示例:拆解垃圾信息检测
One broad question (bad) { "state": { "message": { "sender": { "display_name": "Acme Payroll", "email": "rewards@claim-bonus.example" }, "subject": "Urgent: claim your employee bonus", "body": "You have been selected for a $1,000 bonus. Confirm your payroll password today to receive it.", "links": [ { "text": "Claim bonus", "url": "http://claim-bonus.example/acme" } ] } }, "selectedModels": [ "jev-latest" ], "questions": { "is_spam": { "type": "noul", "instructions": "Is `message` spam?" } } }这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。
Decomposed questions (good) { "state": { "message": { "sender": { "display_name": "Acme Payroll", "email": "rewards@claim-bonus.example" }, "subject": "Urgent: claim your employee bonus", "body": "You have been selected for a $1,000 bonus. Confirm your payroll password today to receive it.", "links": [ { "text": "Claim bonus", "url": "http://claim-bonus.example/acme" } ] } }, "selectedModels": [ "jev-latest" ], "questions": { "requests_credentials": { "type": "noul", "instructions": "Does `message.body` ask the recipient to provide a password or other login credential?" }, "offers_unexpected_reward": { "type": "noul", "instructions": "Does `message.body` claim the recipient received an unexpected prize, payment, or reward?" }, "creates_time_pressure": { "type": "noul", "instructions": "Does `message.subject` or `message.body` pressure the recipient to act quickly?" }, "sender_identity_mismatch": { "type": "noul", "instructions": "Does the organization named in `message.sender.display_name` conflict with the domain in `message.sender.email`?" }, "link_domain_mismatch": { "type": "noul", "instructions": "Does the domain in `message.links[0].url` conflict with the organization named in `message.sender.display_name`?" }, "disguises_link_destination": { "type": "noul", "instructions": "Does `message.links[0].text` conceal or misrepresent the destination in `message.links[0].url`?" } } }这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。
示例:核查工具调用轨迹
One broad question (bad) { "state": { "request": { "text": "What's the weather in Seattle tomorrow in Fahrenheit?", "location": "Seattle, WA", "date": "2026-09-03", "unit": "fahrenheit" }, "available_tools": { "geocode_city": { "description": "Resolve a city to latitude and longitude.", "parameters": { "city": "string" } }, "get_weather": { "description": "Get the forecast for coordinates and a date.", "parameters": { "latitude": "number", "longitude": "number", "date": "YYYY-MM-DD", "unit": [ "fahrenheit", "celsius" ] } } }, "trace": { "tool_calls": [ { "id": "call_1", "name": "geocode_city", "arguments": { "city": "Seattle, WA" } }, { "id": "call_2", "name": "get_weather", "arguments": { "latitude": 47.6062, "longitude": -122.3321, "date": "2026-09-03", "unit": "celsius" } } ], "tool_results": [ { "tool_call_id": "call_1", "output": { "latitude": 47.6062, "longitude": -122.3321 } } ] } }, "selectedModels": [ "jev-latest" ], "questions": { "tool_calls_are_correct": { "type": "noul", "instructions": "Is `trace.tool_calls` correct for `request` and `available_tools`?" } } }这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。
Decomposed questions (good) { "state": { "request": { "text": "What's the weather in Seattle tomorrow in Fahrenheit?", "location": "Seattle, WA", "date": "2026-09-03", "unit": "fahrenheit" }, "available_tools": { "geocode_city": { "description": "Resolve a city to latitude and longitude.", "parameters": { "city": "string" } }, "get_weather": { "description": "Get the forecast for coordinates and a date.", "parameters": { "latitude": "number", "longitude": "number", "date": "YYYY-MM-DD", "unit": [ "fahrenheit", "celsius" ] } } }, "trace": { "tool_calls": [ { "id": "call_1", "name": "geocode_city", "arguments": { "city": "Seattle, WA" } }, { "id": "call_2", "name": "get_weather", "arguments": { "latitude": 47.6062, "longitude": -122.3321, "date": "2026-09-03", "unit": "celsius" } } ], "tool_results": [ { "tool_call_id": "call_1", "output": { "latitude": 47.6062, "longitude": -122.3321 } } ] } }, "selectedModels": [ "jev-latest" ], "questions": { "geocode_tool_is_relevant": { "type": "noul", "instructions": "Is `trace.tool_calls[0].name` an appropriate tool for resolving `request.location`?" }, "geocode_location_matches": { "type": "noul", "instructions": "Does `trace.tool_calls[0].arguments.city` match `request.location`?" }, "geocode_arguments_match_schema": { "type": "noul", "instructions": "Does `trace.tool_calls[0].arguments` conform to `available_tools.geocode_city.parameters`?" }, "geocode_result_matches_call": { "type": "noul", "instructions": "Does `trace.tool_results[0].tool_call_id` match `trace.tool_calls[0].id`?" }, "weather_tool_is_relevant": { "type": "noul", "instructions": "Is `trace.tool_calls[1].name` an appropriate tool for answering `request.text`?" }, "weather_arguments_match_schema": { "type": "noul", "instructions": "Does `trace.tool_calls[1].arguments` conform to `available_tools.get_weather.parameters`?" }, "weather_uses_geocoded_coordinates": { "type": "noul", "instructions": "Do the coordinates in `trace.tool_calls[1].arguments` match those in `trace.tool_results[0].output`?" }, "weather_date_matches": { "type": "noul", "instructions": "Does `trace.tool_calls[1].arguments.date` match `request.date`?" }, "weather_unit_matches": { "type": "noul", "instructions": "Does `trace.tool_calls[1].arguments.unit` match `request.unit`?" } } }这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。
- 在问题中使用结构
让问题保持简短。
instructions和criteria通常是字符串,对于简短、无歧义的问题,一个字符串就够了。它们也可以是对象或数组。把问题放在一个字段里,把引导该问题的数据放在其他字段里。结构在以下情形中会有帮助:
- 问题需要上下文或示例。一长句背景信息或一组示例输入应当放在问题旁边的具名字段中,这样你的代码就能增补或替换它们,而无需重写问题。
- 问题的一部分来自你的代码。当某个值来自数据库时,把它放进自己的字段,而不是拼接进字符串模板。
- 多个问题有相似的 instructions。一个请求接收一个状态,并且可以包含多个问题。添加补充数据有助于让问题彼此区分。
示例:引用来自你的代码的记录
这个 Noul 会把状态中的一份简历与候选人数据库中的一条记录进行比较。该记录按原样放进
potential_duplicate,问题则按名称引用它。questions { "state": { "resume": { "name": "John Smith", "location": "Oakland, CA", "summary": "Backend engineer with eight years of Python and Go experience.", "experience": [ { "employer": "Google", "title": "Senior Backend Engineer", "years": "2021-2025" }, { "employer": "Microsoft", "title": "Software Engineer", "years": "2017-2021" } ] } }, "selectedModels": [ "jev-latest" ], "questions": { "same_as_record_18": { "type": "noul", "instructions": { "potential_duplicate": { "name": "John Smith", "location": "Oakland, California", "last_employer": "Google" }, "question": "Is the resume for the same person as `potential_duplicate`?" } } } }这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。
来源为代码的 "potential_duplicate" 数据可能随时间变化。"question" 使用反引号引用它。
criteria内部的描述也可以是对象。对于 Choice,每个选项的描述都可以是一个对象,说明该选项涵盖什么、什么属于另一个选项,以及几个示例。在各选项之间使用相同的字段名,这样模型就能直接比较它们。示例:定义对比式的 Choice criteria
questions { "state": "How many disposable virtual cards can I make per day?", "selectedModels": [ "jev-latest" ], "questions": { "card_help_topic": { "type": "choice", "instructions": { "question": "Which disposable virtual card topic is the user asking about?", "focus": "Classify the information the user wants." }, "criteria": { "get_disposable_virtual_card": { "what": "Purpose, eligibility, or setup", "not_for": "Quantity, transaction, or merchant restrictions", "examples": [ "How can I get a disposable virtual card?", "What are disposable cards for?" ] }, "disposable_card_limits": { "what": "Quantity, transaction, or merchant restrictions", "not_for": "Purpose, eligibility, or setup", "examples": [ "How many disposable cards can I make per day?", "Where can I use a disposable card?" ] } } } } }这个示例是可交互的;到官网原页面可以直接在 Playground 里运行。
每种问题类型的页面都有一个完整示例:
- Noul 会把一份简历与多条候选人记录进行比较,每条记录一个问题,问题在代码中构建。
- Choice 用「每个选项涵盖什么、不适用于什么、以及示例」来描述两个容易混淆的选项。
- Score 为每个等级给出描述和示例情形。
结构化数据抽取级联实践手册 展示了共享措辞的情形:对一条抽取记录的每个字段提出同一组问题。
简短、无歧义的问题或 criteria 可以继续用字符串。当结构能把本会糊在一起的指引分开时,就加上结构。关于接受结构的全部位置,参见进阶:结构。
- 提出大量问题
- 在代码中组合问题输出(或喂给经典 ML 模型)
用确定性规则或加权和来组合相互独立的答案。若要做学习式组合,可以把这些概率作为特征,输入下游的经典机器学习模型。
示例:用加权分数组合信号
python answers = response.answers # Combine independent signals into one application-specific score. quality = ( 0.4 * answers["answers_request"].noul + 0.4 * answers["citations_are_supported"].noul + 0.2 * (1 - answers["contradicts_context"].noul) )复合评分 展示了如何在组合各个判断的同时保留它们。如果你没有为下游模型准备标签,可以用一组昂贵的推理模型来生成标签;AutoResearch 实践手册 展示了如何在 System One 的输出上训练一个经典模型。
- 按不确定性路由
拆解并不需要更多的往返。针对同一个状态的问题会并行运行。
把它们组合起来
这个支持工单工作流把确定性工作保留在代码中,只发送相关的结构化上下文,在同一个请求中评估许多原子化问题,并用明确的置信度门控来组合答案。
from typesafe_sdk import Choice, Noul, NoulCriteria, Score, TypeSafeClient
def triage_ticket(ticket, customer):
# Handle deterministic states without calling a model.
if ticket["status"] == "closed":
return "no_action"
open_orders = [
order for order in customer["orders"] if order["status"] != "delivered"
]
# Include only the structured context needed by the questions below.
state = {
"ticket": {
"message": ticket["message"],
"sender": ticket["sender"],
"links": ticket["links"],
},
"customer": {
"plan": customer["plan"],
"open_orders": open_orders,
},
"policy": {
"sensitive_credentials": ["password", "security code", "API key"],
},
}
# Ask structured, atomic questions together so they run in parallel.
questions = {
"topic": Choice(
instructions={
"question": "Which team should handle `ticket.message`?",
"focus": "Classify the customer's primary request.",
},
criteria={
"billing": {
"what": "Charges, invoices, refunds, or subscriptions",
"not_for": "Order tracking or account access",
"examples": ["I was charged twice", "Where is my refund?"],
},
"orders": {
"what": "Order status, delivery, cancellation, or returns",
"not_for": "Charges or account access",
"examples": ["Where is my order?", "Cancel my shipment"],
},
"account": {
"what": "Login, profile, permissions, or security",
"not_for": "Charges or order tracking",
"examples": ["Reset my password", "I cannot sign in"],
},
},
),
"requests_credentials": Noul(
instructions={
"question": "Does the message request a sensitive credential?",
"compare": [
"`ticket.message`",
"`policy.sensitive_credentials`",
],
"focus": "Look for a request to disclose the credential itself.",
},
criteria=NoulCriteria(
true={
"what": "Asks the recipient to disclose a listed credential",
"examples": [
"Reply with your password",
"Send us your API key",
],
},
false={
"what": "Does not ask the recipient to disclose a credential",
"not_for": "A legitimate instruction to reset a credential",
"examples": ["Use this link to reset your password"],
},
),
),
"sender_identity_mismatch": Noul(
instructions={
"question": "Does the claimed sender identity conflict with its domain?",
"compare": [
"`ticket.sender.display_name`",
"`ticket.sender.email`",
],
"focus": "Compare the named organization with the email domain.",
},
criteria=NoulCriteria(
true={
"what": "Claims an organization unrelated to the email domain",
"examples": ["Acme Payroll sent from claim-bonus.example"],
},
false={
"what": "The identity and domain agree or make no conflicting claim",
"examples": ["Acme Payroll sent from acme.example"],
},
),
),
"unexpected_reward": Noul(
instructions={
"question": "Does the message announce an unexpected reward?",
"inspect": "`ticket.message`",
"focus": "Look for an unsolicited prize, payment, or reward claim.",
},
criteria=NoulCriteria(
true={
"what": "Announces an unrequested prize, payment, or reward",
"examples": ["You were selected for a $1,000 bonus"],
},
false={
"what": "Contains no reward claim or discusses an expected payment",
"not_for": "A customer asking about a known refund or payroll deposit",
"examples": ["When will my approved refund arrive?"],
},
),
),
"refund_requested": Noul(
instructions={
"question": "Does the customer explicitly request a refund or credit?",
"inspect": "`ticket.message`",
"focus": "Require a requested remedy, not a billing complaint alone.",
},
criteria=NoulCriteria(
true={
"what": "Directly asks for money back or an account credit",
"examples": ["Please refund the duplicate charge"],
},
false={
"what": "Does not ask for a refund or credit",
"not_for": "A complaint or billing question without a requested remedy",
"examples": ["Why was I charged twice?"],
},
),
),
"mentions_open_order": Noul(
instructions={
"question": "Does the message refer to a supplied open order?",
"compare": [
"`ticket.message`",
"`customer.open_orders`",
],
"focus": "Match an order id or other identifying details.",
},
criteria=NoulCriteria(
true={
"what": "Refers to an open order by id or identifying details",
"examples": ["Where is order A-104?"],
},
false={
"what": "Does not identify any supplied open order",
"not_for": "A generic order question with no matching details",
"examples": ["How long does shipping usually take?"],
},
),
),
"frustration": Score(
instructions={
"question": "How frustrated does the customer appear?",
"inspect": "`ticket.message`",
"focus": "Judge expressed frustration, not issue severity.",
},
criteria=[
{
"what": "Calm and matter-of-fact",
"signals": ["Neutral wording", "No complaint about the experience"],
},
{
"what": "Frustrated but civil",
"signals": ["Expresses annoyance", "Remains constructive"],
},
{
"what": "Very angry or threatening to leave",
"signals": ["Hostile language", "Threatens cancellation or churn"],
},
],
),
}
with TypeSafeClient() as client:
response = client.system_one(
state=state,
questions=questions,
)
# Compose independent spam signals with weights controlled by code.
answers = response.answers
spam_risk = (
0.45 * answers["requests_credentials"].noul
+ 0.30 * answers["sender_identity_mismatch"].noul
+ 0.25 * answers["unexpected_reward"].noul
)
# Escalate uncertain judgments instead of guessing.
spam_is_uncertain = 0.4 < spam_risk < 0.6
if spam_is_uncertain or answers["topic"].confidence < 0.75:
return route_to_human_review(ticket)
if spam_risk >= 0.6:
return quarantine_as_spam(ticket)
# Let code decide which speculative answers matter on this path.
if answers["topic"].choice == "billing":
return route_to_billing(
ticket,
refund_requested=answers["refund_requested"].noul >= 0.7,
)
if answers["topic"].choice == "orders":
return route_to_orders(
ticket,
mentions_open_order=answers["mentions_open_order"].noul >= 0.7,
)
priority = (
"high"
if answers["frustration"].confidence >= 0.7
and answers["frustration"].score >= 1.5
else "normal"
)
return route_to_account_support(ticket, priority=priority)