Versioned Search Quickstart

Open inAnthropic

Release Candidate — TerminusDB 12.1

This page documents functionality in the upcoming TerminusDB 12.1 release. Details may change before the final release.

Prerequisites

  • Docker and Docker Compose installed
  • No external API keys required — the default embedding model runs locally on CPU

What you'll achieve By the end of this guide, you will have the search engine running, one commit indexed, and a semantic search returning results — all over HTTP with curl.

Bring up the stack, index one commit, and run a search. The whole thing takes about five minutes once the embedding model is pulled.

1. Start the stack

One command — no clone, no build. All three services use pre-built images:

Example: Bash
curl -fsSL https://raw.githubusercontent.com/terminusdb-org/vectorlink/main/docker-compose.quickstart.yml \
  | docker compose -p vectorlink-quickstart -f - up -d

This starts three services on CPU, with no external network after the one-time model pull:

  • terminusdb — a TerminusDB server on :6365 (with the indexer plugin that auto-pushes to VectorLink)
  • vectorlink — the search engine (HTTP API on :7372)
  • embeddings — a local embedding model server (Ollama serving nomic-embed-text-v2-moe)

Stop everything with docker compose -p vectorlink-quickstart down.

Wait until the engine is ready:

Example: Bash
curl -fsS http://localhost:7372/health/ready | jq
# { "ready": true, "index": true, "search": true }

index: true means it can accept pushes. search: true means the embedding backend is warm. Poll until both are true — don't sleep a fixed amount.

Example: Bash
until curl -fsS http://localhost:7372/health/ready | jq -e '.index and .search' >/dev/null; do
  echo "waiting for engine…"; sleep 2
done
echo "ready"

2. Index a commit

Find where the engine is up to (empty on first use):

curl -u admin:root 'http://localhost:7372/last-indexed?domain=admin/star_wars&branch=main'
# { "branch": "main", "commit": null, "version": 0 }

Push a small NDJSON delta as commit c1 — one operation per line. Each Inserted carries the rendered text the engine will chunk and embed:

curl -u admin:root -X POST \
  'http://localhost:7372/push?domain=admin/star_wars&branch=main&target_commit=c1' \
  -H 'Content-Type: application/x-ndjson' \
  -d '{"op":"Inserted","id":"terminusdb:///star-wars/People/20","string":"This person is named Yoda. A wise old Jedi master, small and green."}
{"op":"Inserted","id":"terminusdb:///star-wars/Species/8","string":"The Mon Calamari are an amphibious species resembling squid, known as skilled starship engineers."}'
# task-7f3a9c

The string field above is hand-written for this tutorial. In production, TerminusDB generates it automatically from your schema's embedding template — a GraphQL query selects fields, a Handlebars template renders them to text. For example, a Person class with name and bio fields might use "template": "{{name}}. {{bio}}", producing the same kind of plain-text string you see above. See TerminusDB Push Indexing for the full schema setup.

Poll the task until complete:

Example: Bash
curl -u admin:root 'http://localhost:7372/check?task_id=task-7f3a9c'
# { "status": "Complete", "indexed_documents": 2, "skipped": [] }
Run the example above first to populate this value.

GET — simple and cacheable:

curl -u admin:root 'http://localhost:7372/search?domain=admin/star_wars&commit=c1&q=wise+old+man&snippet=true'
# [
#   {
#     "id": "terminusdb:///star-wars/People/20",
#     "distance": 0.0823,
#     "chunk": { "index": 0, "count": 1, "token_start": 0, "doc_token_len": 14, "location": 0.0,
#       "snippet": "This person is named Yoda. A wise old Jedi master, small and green." }
#   }
# ]

POST — structured JSON:

curl -u admin:root -X POST 'http://localhost:7372/search' \
  -H 'Content-Type: application/json' \
  -d '{"domain":"admin/star_wars","commit":"c1","q":"who are the squid people","snippet":true}'
# [
#   {
#     "id": "terminusdb:///star-wars/Species/8",
#     "distance": 0.0941,
#     "chunk": { "index": 0, "count": 1, "token_start": 0, "doc_token_len": 16, "location": 0.0,
#       "snippet": "The Mon Calamari are an amphibious species resembling squid, known as skilled starship engineers." }
#   }
# ]

Expected (hybrid, the default): People/20 (Yoda) tops "wise old man". Species/8 (Mon Calamari) tops "squid people" — even though neither rendered text contains those exact words. The snippet field in each hit shows the chunk text that was actually matched, making it easy to see why the engine returned that document. That is the semantic payoff.

What just happened

  1. You pushed two documents as an NDJSON stream to the engine.
  2. The engine chunked each document's text, embedded the chunks with the local model, and stored them as a versioned snapshot tagged commit:c1.
  3. You searched that snapshot. The engine embedded your query, ran hybrid (vector + full-text) search, deduplicated chunk hits back to documents, and returned the nearest match first.

In production, TerminusDB performs steps 1–2 automatically after each commit. The curl calls above are how you drive the engine standalone for testing or debugging.

Asymmetric model prefixes (automatic)

The default embedding model — nomic-embed-text-v2-moe — is an asymmetric model. It expects different text prefixes depending on whether you are indexing a document or running a query:

  • Documents being indexed get search_document: prepended automatically.
  • Search queries get search_query: prepended automatically.
  • Duplicate detection and entity resolution use clustering: automatically.

You never write search_query: or search_document: yourself — the engine applies the correct prefix based on which endpoint you call. Getting this wrong is the most common silent quality killer with asymmetric models (e5, bge, and nomic all share this failure mode), so the engine handles it for you.

Cleanup

Delete the indexed domain so you can rerun the quickstart from scratch:

curl -u admin:root -X DELETE 'http://localhost:7372/domain?domain=admin/star_wars'

Next: TerminusDB Push Indexing.

Was this helpful?