Release Candidate — TerminusDB 12.1
This page documents functionality in the upcoming TerminusDB 12.1 release. Details may change before the final release.
Entity resolution is the process of determining whether records from one or more datasets refer to the same real-world entity. Versioned Search provides two endpoints for this: /duplicates for near-duplicate detection within a population, and /candidates for raw k-nearest-neighbour gathering across two record sets.
The mental model: gather wide, then filter tight
Entity resolution runs in two stages:
The golden rule: you can only filter what you gathered. Every tau must be ≤ threshold. The engine rejects tau > threshold — you cannot filter looser than you gathered. So gather generously, then tighten.
Distance scale
All values are on the reference cosine scale [0, 1]: 0 = identical, 0.5 = unrelated (orthogonal), 1 = opposite. Smaller = closer = more similar.
A threshold of 0.3 means "only consider neighbours within 0.3 distance". A tau of 0.15 means "only accept as a match if within 0.15".
Duplicate detection
Surface near-duplicate groups within a single population:
curl -u admin:root 'http://localhost:7372/duplicates?domain=admin/star_wars&commit=c1&threshold=0.05'Each result is a { "group": [ {"id"[, "snippet"]}, … ], "distance": <0..1> }, sorted nearest-first. /duplicates is always bounded — it never runs an unbounded all-pairs scan.
Cross-set entity resolution
Match records across two sets — for example, matching Abt.com products against Buy.com products:
curl -u admin:root 'http://localhost:7372/duplicates?domain=admin/er&commit=c1&threshold=0.1&doc_type=Abt&target_doc_type=Buy&snippet=true'doc_type/doc_iddefine the set side.target_doc_type/target_doc_iddefine the target side.- The engine gathers cross-set neighbours and filters by cardinality-aware tau gates.
The tuning knobs
k — candidate breadth
How many nearest neighbours to gather per record. This is your ceiling on how many matches a single record can have.
- Set
kto the maximum plausible number of pairs per record in your domain. If a set record realistically matches at most 2–3 target records,k = 3–5is plenty. - Keep it low.
kdoes not need to grow with how many chunks a document has — the engine filters to cross-document neighbours, so a smallkis safe. - Higher
kmeans more candidates considered (more recall headroom) but more noise to filter and slightly more compute. Start low, raise only if you find true matches being cut off at the candidate stage.
Rule of thumb: start at k = 5. Drop to 2–3 for near-1:1 data. Raise toward 10 only if a record legitimately matches many.
threshold — gather ceiling
The maximum distance at which a neighbour is even considered. Anything beyond this is never gathered.
- Set
thresholdwide enough to catch all plausible matches, but not so wide that you drown in noise. - A good starting point is
0.3— generous enough for fuzzy matches, tight enough to exclude clearly unrelated records. - If you see true matches being missed (false negatives), widen
threshold. If you see too many false positives at the gather stage, tighten it.
tau gates — filter tightness
After gathering, the engine applies cardinality-aware filters:
tau_one_to_one— the reciprocal core. Both sides rank each other as nearest. This is the tightest gate and catches high-confidence 1:1 matches.tau_one_to_many— set-side extras. A set record legitimately matches multiple target records.tau_many_to_one— target-side extras. A target record legitimately matches multiple set records.
Each tau must be ≤ threshold. Start with all tau values equal to threshold (effectively no extra filtering), then tighten the tau gates to improve precision while maintaining recall.
Candidates: raw KNN gather
For cases where you want the raw k-nearest-neighbour pairs without the tau filtering stage, use /candidates:
curl -u admin:root -X POST 'http://localhost:7372/candidates' \
-H 'Content-Type: application/json' \
-d '{
"domain": "admin/er",
"commit": "c1",
"set_doc_types": ["Abt"],
"target_doc_types": ["Buy"],
"k": 5,
"threshold_set": 0.3,
"threshold_target": 0.3,
"include": "embeddings,content"
}'This returns the raw gathered pairs with optional embeddings and content, so you can apply your own filtering logic downstream.
A worked example
Say you have two product catalogues — Abt and Buy — and want to find which Abt products match which Buy products.
- Index both catalogues in the same data product, with
doc_typedistinguishing them (e.g.AbtandBuy). - Start with
threshold=0.3,k=5and call/duplicateswithdoc_type=Abt&target_doc_type=Buy. - Inspect the results. If you see obvious false positives, tighten
thresholdto0.2or lower thetaugates. - If you see false negatives (known matches missing), widen
thresholdto0.35or raisekto7. - Iterate until the F1 score on a known-good sample is acceptable.
The key insight: tuning is a two-stage process. First, set k and threshold to gather generously. Then, tighten the tau gates to filter precisely. You can only filter what you gathered — so get the gather stage right first.
Next: API Reference.