复核引用
通过与源文档核对,捕捉错误或凭空捏造的引用。一个 TypeSafe Choice 问题判断引文的上下文是否支持该主张,它的置信度可以把这条引用标记出来供人工复核。
大语言模型(LLM)在回答问题时附上引用:对每个主张,给出一段源文档的小节以及它所依据的引文。其中一些引用是错误的或凭空捏造的:引文可能根本不在文档中,也可能逐字出现在文档里,但其上下文说的却与主张相反。
人工核对一条引用很慢:找到文档,在文档里找到引文,然后读足够多的上下文来判断它是否支持该主张。
为了自动完成这项核对,我们首先用普通的字符串匹配找出缺失的引文,然后用一个 Choice 问题读取每一条幸存引文的上下文,判断它是否支持该主张。
下面,来自一个 LLM 关于 RFC 7519(JSON Web Token)的回答的八条引用经过这项检查。其中四条准确的引用以 0.93 或更高的置信度返回 verified。四个人为植入的失败全部被捕捉到:一条捏造的引文、一个被反驳的主张,以及两条被送去人工处理的不受支持引用。
check_citation() 是你在这里构建的函数,它接收一份源文档和一条引用,返回四种判定之一:verified、unsupported、contradicted 或 fabricated。它还会返回一个置信度,用来标记出需要人工查看的那些。
环境准备
pip install ipython "typesafe-sdk>=0.5.7" cooksafe --extra-index-url https://pypi.typesafe.ai/
然后设置 TYPESAFE_API_KEY。每次 API 调用都会缓存到 json_cache.json,该文件随实践手册一起提供,因此重新运行时会重放已发布的数字,而不是调用 API。删除该文件即可全部实时运行。
下面的数字来自 2026-08-16 的 jev-1.12。
import json
import os
import re
from pathlib import Path
from time import perf_counter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
AUTO_ACCEPT = 0.8 # start high for more human review as you build trust in the model
client = TypeSafeClient(
api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"),
base_url=os.environ.get("TYPESAFE_ENDPOINT"),
timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))
加载源文档与引用
源文档是 RFC 7519(JSON Web Token),从 rfc-editor.org 获取,并作为 rfc7519.txt 提交在本实践手册旁边。下面的代码去掉页眉和页脚,然后把文本切分为带编号的小节。
citations.json 中的八条引用由一个 LLM 针对该 RFC 撰写。四条是准确的;另外四条被我们改成了无法通过检查。
def load_source() -> str:
"""RFC 7519 verbatim, minus the page headers and footers that interrupt its paragraphs."""
lines = []
for line in Path("rfc7519.txt").read_text().splitlines():
bare = line.lstrip("\f")
if re.match(r"Jones, et al\.\s.*\[Page \d+\]$", bare):
continue
if re.match(r"RFC 7519\s+JSON Web Token \(JWT\)\s+May 2015$", bare):
continue
lines.append(bare)
return re.sub(r"\n{3,}", "\n\n", "\n".join(lines))
def split_sections(source: str) -> dict[str, str]:
"""Map each numbered section ("4.1.3") to its text, split on the RFC's header lines."""
boundary = re.compile(r"(?m)^(?:(\d+(?:\.\d+)*)\. .+|Appendix [A-Z]\..*)$")
marks = list(boundary.finditer(source))
sections = {}
for mark, nxt in zip(marks, marks[1:] + [None]):
if mark.group(1) is None: # an appendix header only terminates the section before it
continue
sections[mark.group(1)] = source[mark.start() : nxt.start() if nxt else len(source)].strip()
return sections
SOURCE = load_source()
SECTIONS = split_sections(SOURCE)
CITATIONS = json.loads(Path("citations.json").read_text())
print(f"{len(SOURCE):,} characters, {len(SECTIONS)} numbered sections, {len(CITATIONS)} citations")
print("\nA citation with a quote:")
print(json.dumps(CITATIONS[1], indent=2))
print("\nA claim-only citation:")
print(json.dumps(next(c for c in CITATIONS if c["quote"] is None), indent=2))
58,365 characters, 45 numbered sections, 8 citations
A citation with a quote:
{
"id": "aud_reject",
"claim": "If a validator does not find itself in a token's audience list, it has to reject the token.",
"quote": "If the principal processing the claim does not identify itself with a value in the \"aud\" claim when this claim is present, then the JWT MUST be rejected.",
"section": "4.1.3"
}
A claim-only citation:
{
"id": "iat_future",
"claim": "The \"iat\" claim requires validators to reject tokens whose issue time is in the future.",
"quote": null,
"section": "4.1.6"
}
在源文档中查找每条引文
不在源文档中的引文就是捏造的,发现这一点不需要任何模型。对空白和弯引号做归一化,使引文在 RFC 的换行之间仍能匹配,然后把它作为子串查找。匹配结果还能说明引文来自哪一小节,而那一小节就是下一步模型要读的文本。
一条引用可以只指明某个小节而不从其中引用任何内容。这种情况下没有什么可匹配的,因此直接取该引用所指明的小节,交给模型。
def normalize(text: str) -> str:
"""Collapse whitespace and fold curly quotes, so a quote matches across line wraps."""
table = str.maketrans({"“": '"', "”": '"', "‘": "'", "’": "'"})
return re.sub(r"\s+", " ", text.translate(table)).strip()
def find_quote(sections: dict[str, str], quote: str) -> str | None:
"""The number of the section that contains the quote verbatim, or None."""
needle = normalize(quote)
for number in sorted(sections, key=lambda n: [int(p) for p in n.split(".")]):
if needle in normalize(sections[number]):
return number
return None
def locate(sections: dict[str, str], citation: dict) -> tuple[str, str | None]:
"""Step 1 for one citation: a status, plus the section step 2 will read."""
if citation["quote"] is None:
return "section-only", sections[citation["section"]]
number = find_quote(sections, citation["quote"])
if number is None:
return "missing", None
return "found", sections[number]
for citation in CITATIONS:
status, section = locate(SECTIONS, citation)
where = f"section of {len(section):,} chars" if section else "not in the source"
print(f"{citation['id']:<18}{status:<14}{where}")
epoch_seconds found section of 3,122 chars
aud_reject found section of 761 chars
sig_reporting missing not in the source
clock_skew found section of 529 chars
exp_required found section of 529 chars
pii_encryption found section of 1,653 chars
iat_future section-only section of 270 chars
duplicate_names found section of 918 chars
验证源文档是否支持该主张
到这里仍然带有引文的引用,都与源文档逐字匹配。但这还不够:引文可能准确,而建立在它之上的主张仍然错误。要判断这一点,需要引文的上下文,也就是第 1 步找到的那一小节。
对每条幸存的引用问一个 Choice 问题,涵盖小节与主张之间的三种关系。
概率最高的选项就是判定结果,而 AUTO_ACCEPT(上面代码中的 0.8)决定如何处理它:
- 置信度达到或高于 0.8:判定结果自行生效;
- 低于 0.8:在有任何东西依据该判定行动之前,由人工确认。
先把阈值设高,随着你观察到模型在自己的文档上的表现,再逐步调低。
QUESTIONS = {
"relation": Choice(
instructions="How does the section relate to the claim?",
criteria={
"supports": "The section states the claim or directly implies that it is true",
"contradicts": "The section states the opposite of the claim or implies it is false",
"says_nothing": "The section does not address what the claim asserts, either way",
},
),
}
RELATION_TO_VERDICT = {
"supports": "verified",
"contradicts": "contradicted",
"says_nothing": "unsupported",
}
@json_cache
def ask(claim: str, section: str) -> dict:
started = perf_counter()
response = client.system_one(
state={"claim": claim, "section": section},
questions=QUESTIONS,
model=TYPESAFE_MODEL,
)
answer = response.answers["relation"]
return {
"choice": answer.choice,
"probabilities": answer.probabilities,
"confidence": answer.confidence,
"seconds": round(perf_counter() - started, 2),
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
def verdict(status: str, answer: dict | None) -> dict:
"""Fold step 1 and step 2 into one of the four labels, plus an auto-or-review flag."""
if status == "missing":
# confidence None: no model was called, so there is no model confidence to report
return {"verdict": "fabricated", "confidence": None, "auto": True}
return {
"verdict": RELATION_TO_VERDICT[answer["choice"]],
"confidence": answer["confidence"],
"auto": answer["confidence"] >= AUTO_ACCEPT,
}
def check_citation(sections: dict[str, str], citation: dict) -> dict:
status, section = locate(sections, citation)
answer = ask(citation["claim"], section) if section is not None else None
return {"id": citation["id"], "status": status, "answer": answer, **verdict(status, answer)}
检查每一条引用
八条引用全部经过同一项检查:
print(f"{'citation':<18}{'quote':<14}{'relation':<14}{'conf':>6} {'verdict':<13}{'action':>7}")
for citation in CITATIONS:
result = check_citation(SECTIONS, citation)
answer = result["answer"]
relation = answer["choice"] if answer else "-"
conf = f"{answer['confidence']:.2f}" if answer else "-"
action = "auto" if result["auto"] else "review"
print(
f"{result['id']:<18}{result['status']:<14}{relation:<14}{conf:>6}"
f" {result['verdict']:<13}{action:>7}"
)
citation quote relation conf verdict action
epoch_seconds found supports 0.93 verified auto
aud_reject found supports 0.95 verified auto
sig_reporting missing - - fabricated auto
clock_skew found supports 0.99 verified auto
exp_required found contradicts 0.99 contradicted auto
pii_encryption found says_nothing 0.27 unsupported review
iat_future section-only says_nothing 0.56 unsupported review
duplicate_names found supports 0.99 verified auto
四条引用返回 verified,一条返回 fabricated,一条返回 contradicted,两条返回 unsupported。
epoch_seconds、aud_reject、clock_skew和duplicate_names就是那四条准确的引用。 它们全部以 0.93 或更高的置信度返回verified,远高于AUTO_ACCEPT。sig_reporting从未到达模型。它的引文不在 RFC 中,所以仅凭字符串匹配就把它标记为fabricated。exp_required逐字引用了第 4.1.4 节,而同一节写着「Use of this claim is OPTIONAL」,所以它是contradicted,置信度为 0.99。pii_encryption和iat_future以 0.27 和 0.56 返回unsupported,两者都低于 阈值,因此都交给了人工。pii_encryption说明了为什么仅靠字符串匹配还不够: 它的引文逐字出现在源文档中,而它来自的那一小节对主张什么也没说。
要把它用于你自己的数据,请替换 rfc7519.txt 和 citations.json。
load_source() 和 split_sections() 是针对 RFC 的版式编写的,因此其他形态的文档需要自己的解析逻辑。
归一化之后字符串匹配是精确的:被截断或略微改写的引文都会返回 fabricated。若要容忍不严谨的引用方式,生产系统需要改用模糊匹配。
在 playground 中打开它
该链接包含一条引用的主张和小节,以及那个问题。打开它即可在浏览器中实时运行同样的调用。
example = next(c for c in CITATIONS if c["id"] == "exp_required")
_, example_section = locate(SECTIONS, example)
playground_link = make_playground_link(
{"claim": example["claim"], "section": example_section}, QUESTIONS, models=[TYPESAFE_MODEL]
)
display(Markdown(f"🔗 [Open one citation's claim + section in the TypeSafe playground]({playground_link})"))