Tutorial: Building with CleaveDB
PEER INTO ATTENTION
CleaveDB integrates native Transformer attention layers directly into its hybrid document-graph engine. The PEER INTO ATTENTION diagnostic command inspects the operational health, hardware acceleration, execution providers, and tensor dimensions of the integrated Semantic Relevance Attention (SRA) subsystem.
Syntax and command
Run the standalone diagnostic statement in any CleaveQL session:
PEER INTO ATTENTIONThis operation returns instantaneous metadata from the ONNX Runtime session without executing any document scans or disk I/O.
Structured telemetry response
When the AI acceleration subsystem is online, PEER INTO ATTENTION returns a comprehensive telemetry object:
{
"status": "ok",
"attention_stats": {
"model": "all-MiniLM-L6-v2",
"onnx_version": "1.16.0",
"providers": [
"CPUExecutionProvider"
],
"input_names": [
"input_ids",
"attention_mask",
"token_type_ids"
],
"output_names": [
"last_hidden_state"
],
"dim": 384,
"hw_acceleration": "CPUExecutionProvider",
"custom_metadata": {}
}
}The response confirms that the neural transformer is loaded in memory and ready to calculate dense semantic embeddings.
Diagnostic fields breakdown
| Property | Value / Format | Description |
|---|---|---|
| model | all-MiniLM-L6-v2 | Quantized 8-bit Transformer model bundled directly in CleaveDB |
| dim | 384 | Output vector embedding dimensionality used for cosine similarity ranking |
| providers | ["CPUExecutionProvider"] | Active runtime backends; leverages AVX and AVX-512 SIMD vector extensions |
| input_names | input_ids, attention_mask, token_type_ids | Expected input tensor signatures fed by the embedded tokenizer |
| output_names | last_hidden_state | Transformer hidden states pooled to produce normalized 1D sentence vectors |
| hw_acceleration | CPUExecutionProvider / CUDA | Primary hardware execution provider currently handling tensor operations |
Bundled offline architecture
Unlike databases that rely on external third-party embedding APIs (introducing network latency and API token costs), CleaveDB bundles its quantized 8-bit model (~22 MB) directly inside the attention/model/ directory.
- 100% Offline: No internet connectivity or remote API calls required during indexing or search.
- Zero API Latency: Embeddings are computed in microseconds via local CPU SIMD vector units.
- Mean Pooling & Normalization: Raw output tensors undergo automated mean pooling and L2 normalization to compute dot-product cosine similarity.
Graceful fallback mode
If optional machine learning dependencies (such as ONNX Runtime or tokenizers) are absent from the host environment, CleaveDB degrades gracefully to exact keyword matching without failing database queries:
{
"status": "ok",
"attention_stats": {
"status": "offline",
"reason": "Missing ML dependencies"
}
}In fallback mode, standard document storage, ACID transactions, and graph traversals continue uninterrupted.
