知识图谱实体对齐
判断来自两份啤酒目录的 450 个候选配对中,哪些描述的是同一款产品。整项决策由一个 TypeSafe Score 问题承担,因为它的三个等级恰好就是对一对实体可以做的三件事:合并它、让它保持未关联,或把它交给策展人。这里没有任何需要拟合的阈值,同时还有三个 Noul 问题搭同一趟请求的便车,用来告诉策展人两个数据源在哪个字段上存在分歧。
知识图谱中的一个关键问题是判断一个新进入的实体是否与已有实体重复,尤其是当手头只有来自不同来源的自然语言时。给定潜在的重复配对,单个 TypeSafe Score 就能判断每一对是否重复,或者是否值得策展人进一步细看。
假设两个数据源描述的是同一批事物中相互重叠的部分,而你需要知道一侧的哪一条记录与另一侧的哪一条是同一个事物。知识图谱把这些记录称为实体,并保存关于每个实体所记录的事实。某个廉价但粗略的初筛已经比较过这两个数据源,并挑出了 450 对值得细看的配对。剩下的工作就是对每一对做出判断。
不恰当地合并两个实体的代价更高,因为关于其中任一实体的每一条事实现在都用来描述合并后的实体,而且与任一方相关联的东西也都会被带上。事后要撤销,就得弄清哪条事实来自哪里。而漏掉一个匹配只会留下一条重复记录,因此这项判断需要第三种选择:那些既不适合合并、也不适合丢弃的配对。
这项判断是一个 Score 问题,三种结果各占一个等级:
- 不同产品 — 让两个实体保持未关联
- 相关,但可能不是同一个 — 交给策展人决定
- 同一产品 — 合并它们
我们使用 Score 问题,是因为想把语义标签(即 score criteria)直接附着到每一种结果上,包括中间那种结果。Noul 问题只能通过对它的输出做阈值判断来间接实现这一点,而 Choice 问题会丢失这三种结果之间的有序关系。
接下来,对于我们想考察的实体的每个字段,关于这些字段是否匹配的 Noul 问题可以搭同一趟请求的便车。如果 score 既没有落在「同一产品」等级、也没有落在「不同产品」等级,这些 noul 就为策展人提供更详细的信息。
最终你得到一个 route(),它接收一对候选实体,返回三种结果之一,而且没有任何需要你针对自己的数据去拟合的阈值。
环境准备
pip install matplotlib ipython "typesafe-sdk>=0.5.7" cooksafe --extra-index-url https://pypi.typesafe.ai/
然后设置 TYPESAFE_API_KEY。每次调用都会缓存到 json_cache.json,该文件随本实践手册一起提供,因此重新渲染会重放已发布的数字,而不会调用 API。删除该文件即可全部实时重跑。
下面的数字来自 2026-08-11 的 jev-1.12。
import json
import os
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
import matplotlib
import matplotlib.pyplot as plt
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Noul, Score, TypeSafeClient
matplotlib.use("Agg") # headless render
TYPESAFE_MODEL = "jev-1.12"
MAX_WORKERS = 6 # small pool; the public endpoint rate-limits above roughly eight
client = TypeSafeClient(
api_key=os.environ.get(
"TYPESAFE_API_KEY", "cache-only"
), # keyless kernels replay the cache
base_url=os.environ.get("TYPESAFE_ENDPOINT"),
timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))
加载候选配对
这些配对来自一个已发布的基准数据集,即 Magellan 集合中的 Beer 数据:两份从不同网站抓取的啤酒目录,已经由那次粗略初筛缩减到 450 对。每个实体带有四个字段:name、brewery、style 和酒精含量。每对还带有 known_same_as,即该基准自己的答案。
文本完全按发布时的样子保留,未做预处理:从未被还原成字符的 HTML 实体、被拆成独立单词的撇号、少数解码错误的字符。
每一对都会发出一次请求,因此你的花费取决于交给你的配对数量,而不是任一数据源的规模。
PAIRS = json.loads(Path("candidate_pairs.json").read_text(encoding="utf-8"))
BY_ID = {pair["id"]: pair for pair in PAIRS}
print(f"{len(PAIRS)} candidate pairs. The first one, as the model will see it:")
print(json.dumps({k: PAIRS[0][k] for k in ("entity_a", "entity_b")}, indent=2)[:420])
450 candidate pairs. The first one, as the model will see it:
{
"entity_a": {
"name": "C N Red Imperial Red Ale",
"brewery": "Redwood Lodge",
"style": "American Amber / Red Ale",
"abv": "8.10 %"
},
"entity_b": {
"name": "Kinetic Infrared Imperial Red Ale",
"brewery": "Kinetic Brewing Company",
"style": "American Strong Ale",
"abv": "9.30 %"
}
}
为每对候选实体提一个 Score 问题和三个 Noul 问题
两个实体都放进同一个状态,作为 entity_a 和 entity_b,因此这些问题问的是这一对,而不是单独的任一方。四个问题都在同一次请求中发出。
下面的三段等级描述就是整项决策:每个等级就是一种结果。这个文件里任何地方都没有阈值常量。你还可以在看到任何一个 score 之前就写好这些描述,而这一点对于你必须拟合的数字来说并不成立。
中间那个等级值得仔细撰写。在这里它覆盖了变体、特别版,以及可能指代任一产品的名称,这样这些情况会交到策展人手里,而不会被合并或丢弃。
OUTCOME 为三种结果命名。合并这一结果叫作 assert sameAs,因为 sameAs 是记录两个实体为同一事物的标准方式,而写入这样一条记录就是合并实际发生的方式。
四个字段中有三个各配一个 Noul 问题:name、brewery 和 style。酒精含量没有,因为比较两个数字属于算术;如果你需要它,用代码算就行。要在另一类数据上使用它,你只需重写 QUESTIONS 和 LEVELS。除此之外,只有那两个打印结果的函数知道啤酒这回事,它们会给出字段名。
LEVELS = [
"They describe two different products.",
"They describe closely related products that may or may not be the same one: "
"a variant, a special edition, or a name that could plausibly refer to either.",
"They describe one and the same product.",
]
OUTCOME = {0: "leave unlinked", 1: "curator queue", 2: "assert sameAs"}
QUESTIONS = {
"link_state": Score(
instructions="How do the two entity descriptions relate as products?",
criteria=LEVELS,
),
"same_name": Noul(
instructions="Do the two entities state the same beer name?",
),
"same_brewery": Noul(
instructions="Are the two entities from the same brewery?",
),
"same_style": Noul(
instructions="Do the two entities describe the same beer style?",
),
}
@json_cache
def score(pair_id: str) -> dict:
"""One request about one candidate pair -> the score plus the three noul answers."""
pair = BY_ID[pair_id]
response = client.system_one(
state={"entity_a": pair["entity_a"], "entity_b": pair["entity_b"]},
questions=QUESTIONS,
model=TYPESAFE_MODEL,
)
link = response.answers["link_state"]
return {
"score": link.score,
"probabilities": link.probabilities,
"confidence": link.confidence,
"properties": {
k: response.answers[k].noul for k in QUESTIONS if k != "link_state"
},
# tokens and requests are the durable units; don't cache a derived cost
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
def route(score_value: float) -> str:
"""The whole decision rule: the nearest level names the outcome."""
return OUTCOME[min(int(score_value + 0.5), len(LEVELS) - 1)]
def show(pair_id: str) -> None:
pair, result = BY_ID[pair_id], score(pair_id)
print(
f"{pair_id} score {result['score']:.2f} confidence {result['confidence']:.2f}"
f" -> {route(result['score'])}"
)
for side in ("entity_a", "entity_b"):
e = pair[side]
print(f" {e['name'][:44]:<46}{e['brewery'][:30]:<32}{e['style'][:22]}")
nouls = result["properties"]
print(
f" name {nouls['same_name']:.2f} brewery {nouls['same_brewery']:.2f} "
f"style {nouls['same_style']:.2f}"
)
四对。c446 是同一款产品,c427 是两款。另外两对因不同原因落在中间等级:c100 的名称和酒厂都相同,但两个数据源对它的风格的措辞不同;而 c428 则把一款啤酒与它的水果加啤酒花变体配在了一起。
for pair_id in ("c446", "c427", "c100", "c428"):
show(pair_id)
print()
c446 score 1.94 confidence 0.92 -> assert sameAs
Thomas Hooker Old Marley Barleywine Thomas Hooker Brewing Company American Barleywine
Thomas Hooker Old Marley Barleywine Thomas Hooker Brewing Company Barley Wine
name 0.97 brewery 0.99 style 0.81
c427 score 0.03 confidence 0.95 -> leave unlinked
Frost Quake Bourbon Barrel Aged Barley Wine Wellington County Brewery American Barleywine
Lompoc Bourbon Barrel Aged Proletariat Red A Lompoc Brewing Amber Ale
name 0.02 brewery 0.09 style 0.08
c100 score 1.30 confidence 0.27 -> curator queue
Belle Gueule Rousse Brasseurs R.J. American Amber / Red A
Belle Gueule Rousse Brasseurs RJ Amber Lager/Vienna
name 0.95 brewery 0.94 style 0.35
c428 score 1.10 confidence 0.77 -> curator queue
Ambleside Amber Ale Bridge Brewing Company American Amber / Red A
Bridge Ambleside Amber Ale - Pomegranate & G Bridge Brewing Company Amber Ale
name 0.63 brewery 0.98 style 0.74
对每一对候选实体做路由
# 450 candidate pairs, one request each; a small pool keeps a live run to a few minutes.
with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
scored = list(pool.map(lambda pair: score(pair["id"]), PAIRS))
scores = [result["score"] for result in scored]
by_outcome: dict[str, list[str]] = {name: [] for name in OUTCOME.values()}
for pair, s in zip(PAIRS, scores):
by_outcome[route(s)].append(pair["id"])
SURFACE, INK, INK2, MUTED = "#fcfcfb", "#0b0b0b", "#52514e", "#898781"
GRID, AXIS, BLUE, ORANGE = "#e1e0d9", "#c3c2b7", "#2a78d6", "#eb6834"
BINS, TOP = 20, len(LEVELS) - 1
counts = [0] * BINS
for s in scores:
counts[min(int(s / TOP * BINS), BINS - 1)] += 1
centers = [(i + 0.5) / BINS * TOP for i in range(BINS)]
queued = [c if route(x) == "curator queue" else 0 for c, x in zip(counts, centers)]
settled = [c if route(x) != "curator queue" else 0 for c, x in zip(counts, centers)]
fig, ax = plt.subplots(figsize=(7.2, 3.6), facecolor=SURFACE)
ax.set_facecolor(SURFACE)
for side in ("top", "right"):
ax.spines[side].set_visible(False)
for side in ("left", "bottom"):
ax.spines[side].set_color(AXIS)
ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
ax.set_axisbelow(True)
ax.grid(axis="y", color=GRID, linewidth=0.8)
ax.bar(
centers, settled, width=TOP / BINS * 0.9, color=BLUE, label="settled automatically"
)
ax.bar(
centers, queued, width=TOP / BINS * 0.9, color=ORANGE, label="sent to the curator"
)
for edge in (0.5, 1.5):
ax.axvline(edge, color=INK2, linewidth=1, linestyle="--")
ax.set_xticks([0, 0.5, 1, 1.5, 2])
ax.set_xticklabels(["0\ndifferent", "0.5", "1\nrelated", "1.5", "2\nsame"])
ax.set_xlabel("score for the pair", color=INK2, fontsize=9)
ax.set_ylabel("candidate pairs", color=INK2, fontsize=9)
ax.set_title(
f"{len(PAIRS)} candidate pairs, scored once each",
loc="left",
color=INK,
fontsize=11,
)
ax.legend(frameon=False, labelcolor=INK2, fontsize=9)
display(fig)
plt.close(fig)
for name in ("assert sameAs", "curator queue", "leave unlinked"):
n = len(by_outcome[name])
print(f"{name:<16}{n:>5} ({n / len(PAIRS):>5.1%})")
assert sameAs 40 ( 8.9%)
curator queue 50 (11.1%)
leave unlinked 360 (80.0%)
route() 改变答案的那两个 score 值就是分界点。大多数配对都能自动落定:360 对低于下分界点,40 对高于上分界点,剩下 50 对留给策展人。
在这个数据集上,score 并没有整齐地落在整数上。大多数落在 0.25 附近。两款毫无共同点的啤酒仍可能共用同一个风格名称,它们的酒厂名也可能看起来相似,因此模型会把一部分概率分给中间等级,而不是一点都不给。决定一对配对结果的是它落在分界点的哪一侧。它离某个等级有多近并不参与其中。
两个分界点附近的拥挤程度并不相同。有九对落在上分界点 1.5 的 0.1 范围内,而这个分界点决定的是哪些配对会被合并进图谱。有四十七对落在下分界点 0.5 的同样距离内,而这个分界点只决定策展人是否会看到这一对。这两个数字都不是你需要调的东西。它们都取决于你如何措辞各个等级,而正是中间等级的措辞在「交给策展人」和「保持未关联」之间移动着配对。
在演练场中打开它
下面的演练场链接会打开 c428,它的得分是 1.10,被送到了策展人那里。它把 Ambleside Amber Ale 与 Bridge Ambleside Amber Ale - Pomegranate & Galena Hops 配在一起:酒厂相同,酒精含量相同。四个问题都会随链接一起带上。
playground_link = make_playground_link(
{"entity_a": BY_ID["c428"]["entity_a"], "entity_b": BY_ID["c428"]["entity_b"]},
QUESTIONS,
models=[TYPESAFE_MODEL],
)
display(
Markdown(
f"🔗 [Open this pair + questions in the TypeSafe playground]({playground_link})"
)
)