Release Candidate — TerminusDB 12.1
This page documents functionality in the upcoming TerminusDB 12.1 release. Details may change before the final release.
Every chunk gets one embedding by default. That is enough for search and similarity — the two things the engine does out of the box. But some workloads need a second, separate embedding vector per chunk, computed with a different role, to support projection visualizations, near-duplicate grouping, and entity resolution with clustering distances. This page covers how to turn that on and what happens when you do.
The option
Clustering is a schema-level flag. You set it in the @context document — the first object in your schema — under @metadata.terminusdb.options. The value is an array of strings, not a boolean or a single string:
[
{
"@type": "@context",
"@base": "terminusdb:///data/",
"@schema": "terminusdb:///schema#",
"@metadata": {
"terminusdb": {
"options": ["store_clustering"]
}
}
},
{
"@id": "Article",
"@type": "Class",
"@key": { "@type": "Lexical", "@fields": ["title"] },
"title": "xsd:string",
"body": "xsd:string",
"@metadata": {
"embedding": {
"query": "query($id: ID){ Article(id: $id) { title body } }",
"template": "{{title}}. {{body}}"
}
}
}
]A few things to keep in mind:
- The array can hold other option strings. Currently,
"store_clustering"and"store_indices"are supported."store_indices"enables automatic indexing on every commit, while"store_clustering"enables the secondary clustering embeddings described on this page. - If
@metadata,terminusdb, oroptionsis missing — or the string just is not there — clustering is off. The default is alwaysfalse. - TerminusDB reads this at commit time. On schema-change commits, it checks the flag directly from the in-memory schema objects. On document-only commits, it skips the check entirely and relies on the Rust-side cache to decide whether indexing is needed. No restart, no environment variable.
What changes at index time
With clustering on, each chunk is embedded twice:
- Once with
EmbeddingRole::Document— the primary vector, ANN-indexed, used by/searchand/similar. - Once with
EmbeddingRole::Clustering— stored in a separateclustering_embeddingcolumn, never ANN-indexed, retrieved on demand by/embeddingsand/candidates.
That is double the embedding calls. For a 2,000-document corpus with one chunk each, you go from 2,000 model invocations to 4,000. The trade-off is bandwidth and latency at index time for richer downstream queries.
With clustering off, the second column is filled with zero vectors of the same dimensionality. No second call is made. The column always exists — it is just empty.
What changes at query time
The /embeddings endpoint returns both sets when clustering is enabled:
curl -u admin:root 'http://localhost:7373/api/plugin/search-embeddings/admin/my_db/local/branch/main'{
"doc_embeddings": {
"terminusdb:///data/Article/1": [0.12, -0.04, 0.33, ...]
},
"clustering_embeddings": {
"terminusdb:///data/Article/1": [0.08, 0.21, -0.15, ...]
},
"store_clustering": true,
"served_commit": "abc123"
}The store_clustering field in the response tells you whether clustering data is present. When false, clustering_embeddings is an empty object.
The /candidates endpoint uses the clustering column for a dual KNN gather. This produces a clustering_distance alongside the standard document distance — two signals that, together, separate true duplicates from semantic neighbors that merely happen to be close.
Checking the current state
The index status endpoint reports whether clustering is active for a domain:
curl -u admin:root 'http://localhost:7373/api/index/admin/my_db/local/branch/main'The engine.store_clustering field is true, false, or null. The null case means the engine has no indexed data for that domain yet — it has never heard of it.
Turning it on after the first index
Clustering is set on first push. If you indexed without it and later want to enable it, a full reindex is required. There is no in-place upgrade path — the clustering column would be all zeros for every existing chunk, and the only way to populate it is to re-embed everything.
The steps are:
- Update the schema's
@metadata.terminusdb.optionsto include"store_clustering". - Delete the domain's search footprint from the engine.
- Re-push the full index.
Going the other direction — turning clustering off after it was on — also requires a reindex. The zero-vector fallback only applies to data indexed from scratch with clustering disabled.
Next: History & Branching.