Tutorial: Building with CleaveDB
Semantic Search
Keyword search expects the user’s words to resemble the words stored in a document. Semantic search starts with an idea instead. A person looking for “a rainy day” may be interested in documents about storms, wet weather, or a quiet afternoon indoors—even if those exact words never appear together. FIND’s quoted-phrase IN form asks CleaveDB to retrieve documents by similarity of meaning.
FIND "a rainy day" IN articlesHow the similarity search works
CleaveDB creates dense vector embeddings from document text fields using its bundled ONNX model. A background worker processes those embeddings after a document is written. For a semantic query, the model turns the search phrase into a vector, CleaveDB compares it with the stored vectors using cosine similarity, and the closest results are ranked by semantic distance. The result can surface related content even when it does not share the query’s exact keywords.
Account for indexing time and thresholds
Embedding generation happens in the background, so semantic search is eventually consistent: a newly written document may be available to a structured read before its vector is ready. The short FIND form ranks available matches but does not accept an inline similarity threshold. When the request needs hard field conditions or an explicit threshold alongside meaning, use the guided semantic form documented under SCOOP meaning search.
