Getting Started

Tutorial: Building with CleaveDB

PEER INTO ATTENTION

CleaveDB integrates native Transformer attention layers directly into its hybrid document-graph engine. The PEER INTO ATTENTION diagnostic command inspects the operational health, hardware acceleration, execution providers, and tensor dimensions of the integrated Semantic Relevance Attention (SRA) subsystem.

Syntax and command

Run the standalone diagnostic statement in any CleaveQL session:

CleaveQLExample · Inspect attention engine status
PEER INTO ATTENTION

This operation returns instantaneous metadata from the ONNX Runtime session without executing any document scans or disk I/O.

Structured telemetry response

When the AI acceleration subsystem is online, PEER INTO ATTENTION returns a comprehensive telemetry object:

CleaveQLExample · Attention diagnostics report
{
  "status": "ok",
  "attention_stats": {
    "model": "all-MiniLM-L6-v2",
    "onnx_version": "1.16.0",
    "providers": [
      "CPUExecutionProvider"
    ],
    "input_names": [
      "input_ids",
      "attention_mask",
      "token_type_ids"
    ],
    "output_names": [
      "last_hidden_state"
    ],
    "dim": 384,
    "hw_acceleration": "CPUExecutionProvider",
    "custom_metadata": {}
  }
}

The response confirms that the neural transformer is loaded in memory and ready to calculate dense semantic embeddings.

Diagnostic fields breakdown

PropertyValue / FormatDescription
modelall-MiniLM-L6-v2Quantized 8-bit Transformer model bundled directly in CleaveDB
dim384Output vector embedding dimensionality used for cosine similarity ranking
providers["CPUExecutionProvider"]Active runtime backends; leverages AVX and AVX-512 SIMD vector extensions
input_namesinput_ids, attention_mask, token_type_idsExpected input tensor signatures fed by the embedded tokenizer
output_nameslast_hidden_stateTransformer hidden states pooled to produce normalized 1D sentence vectors
hw_accelerationCPUExecutionProvider / CUDAPrimary hardware execution provider currently handling tensor operations

Bundled offline architecture

Unlike databases that rely on external third-party embedding APIs (introducing network latency and API token costs), CleaveDB bundles its quantized 8-bit model (~22 MB) directly inside the attention/model/ directory.

  • 100% Offline: No internet connectivity or remote API calls required during indexing or search.
  • Zero API Latency: Embeddings are computed in microseconds via local CPU SIMD vector units.
  • Mean Pooling & Normalization: Raw output tensors undergo automated mean pooling and L2 normalization to compute dot-product cosine similarity.

Graceful fallback mode

If optional machine learning dependencies (such as ONNX Runtime or tokenizers) are absent from the host environment, CleaveDB degrades gracefully to exact keyword matching without failing database queries:

CleaveQLExample · Offline fallback state
{
  "status": "ok",
  "attention_stats": {
    "status": "offline",
    "reason": "Missing ML dependencies"
  }
}

In fallback mode, standard document storage, ACID transactions, and graph traversals continue uninterrupted.