Getting Started

Tutorial: Building with CleaveDB

Performance and Edge Cases

Graph traversal and vector search do more than scan structured fields, so they have different performance characteristics. A graph query spends work resolving connected paths; a semantic query spends compute comparing vectors. Knowing where that work happens helps you write queries that stay focused and set sensible expectations for newly written documents.

Keep graph patterns selective

A pattern may hold intermediate paths in memory as it follows bonds. In a highly connected graph, each additional hop can multiply the number of possible paths. Anchor the pattern with a selective condition so the engine can start from a smaller set of nodes and prune irrelevant paths early.

CleaveQLExample · Anchor a path at a specific brand
FIND PATTERN users AS u
  LINKED VIA "purchased" TO products AS p
  LINKED VIA "manufactured_by" TO brands AS b
  WHERE b.name = "Acme Corp"

The brand condition narrows the destination before CleaveDB resolves full purchase paths. Apply similarly selective conditions to whichever node has the strongest useful constraint for your question.

Plan for semantic search to arrive after writes

Semantic search depends on document vectors generated by a background worker. A SCOOP read can see a newly written document immediately, while the semantic index may take time to catch up. The search phrase is embedded and compared against the vectors, so semantic queries can also be more compute-intensive than ordinary field filtering, especially as a collection grows.

Be aware of graph changes during traversal

Bonds can expire while background cleanup is running. The guide describes this as a possible time-of-check/time-of-use race: a bond may be observed just before its expiry cleanup runs. CleaveDB re-verifies graph paths so returned results remain structurally sound, but applications should still treat relationship results as a view of graph state at query time.