Core Concepts

Vector Embeddings

Modern AI applications require semantic similarity search, but managing a separate vector database introduces network latency and synchronization nightmares. CleaveDB solves this by embedding a native ONNX Transformer directly into the database engine.

Zero-Dependency ONNX

You do not need to call external APIs (like OpenAI) to vectorize strings. The database natively embeds text on POUR and natively embeds prompts on FIND MEANING using a bundled Transformer model.

Hardware SIMD Vectorization

Cosine similarity calculations run at the bare-metal hardware level. By leveraging C++ AVX-512 intrinsic instructions (_mm512_dp_ps), CleaveDB compares thousands of high-dimensional vectors in a single clock cycle.

Vector Embeddings Architecture

Semantic Querying (MEANING)

Retrieving semantically similar documents is as simple as using the MEANING keyword. CleaveQL allows you to specify a floating-point THRESHOLD to filter out low-confidence matches. Because the ONNX model is deeply integrated with the query parser, there is no need to join against an external vector index.

When a query containing the MEANING keyword reaches the execution engine, the natural language prompt is first tokenized using a bundled BPE (Byte-Pair Encoding) tokenizer. The tokens are then fed through the native ONNX Transformer model in real-time to generate a dense embedding vector. The Rust storage layer then performs a highly optimized nearest-neighbor scan across the bucket, utilizing advanced indexing structures (such as HNSW - Hierarchical Navigable Small World graphs) to bypass exhaustive linear scanning. This allows the database to instantly return documents that conceptually match the user's intent, even if they share zero exact keywords with the prompt.

Auto-Embedding on POUR

Because the ONNX Transformer is embedded directly in the storage engine, you don't need to manually calculate embeddings in your application layer. When you insert or upsert a document using the POUR command, CleaveDB automatically vectorizes the text fields in the background during the transaction lifecycle, ensuring your data and embeddings are always perfectly synchronized.

To prevent heavy machine-learning workloads from blocking the primary Write-Ahead Log (WAL), the embedding process is decoupled into a dedicated asynchronous worker pool. When a document is written to disk, its raw text fields are queued in an in-memory lock-free ring buffer. The background workers consume these text chunks, pass them through the Transformer model, and silently update the document's hidden vector metadata in the HNSW index. This architecture guarantees that high-throughput ingestion pipelines remain unaffected by the computational overhead of deep learning inference.

Hybrid Search (Exact + Semantic)

Vector search is rarely used in isolation. CleaveDB allows you to seamlessly combine exact scalar filtering (like matching specific IDs, dates, or boolean categories) with fuzzy semantic matching in a single execution phase. The engine handles the complex logic of intersecting exact B-Tree lookups with high-dimensional vector similarities automatically.

Under the hood, the query optimizer employs a sophisticated Cost-Based Optimizer (CBO) to determine the most efficient execution path. If the exact scalar filters are highly selective (e.g., matching a single user ID), the engine will execute the B-Tree lookup first, retrieving a tiny subset of candidate documents. It then performs a brute-force SIMD cosine similarity scan only on those candidates, bypassing the vector index entirely. Conversely, if the scalar filters are broad, the engine will traverse the HNSW vector index first, applying the scalar constraints as post-filtering conditions during the graph walk. This dynamic execution strategy guarantees sub-millisecond latencies regardless of data distribution.