CleaveDB 3.9.0 Architecture
Hybrid Relational - Document - Graph with Transformer Attention
CleaveDB is engineered from the ground up for extreme performance. It completely bypasses traditional embedded storage engines in favor of a bespoke, high-performance Rust core, seamlessly bridged with a distributed Go coordinator, Python/PyO3 interpreter, and native ONNX transformers.

Clients & Connection Handling
CleaveDB supports a wide variety of clients communicating through different protocols. The primary entry points are the TCP Client (port 8300), WebSocket Client (port 8301), and the HTTP REST API (port 8302). Official SDKs are provided for Python and Node.js (via npm).
All connections flow through a unified authentication and connection handling layer before reaching the core server.
TCP Server (cleavedb_server.py)
The frontend TCP server manages the connection lifecycle and security context. Key responsibilities include:
- Auth & Session Mgmt: Validating credentials and establishing user sessions with full context.
- Multi-Tenant: Ensuring all requests operate within the correct, isolated tenant namespace.
- Cron Worker: Handling scheduled asynchronous background tasks.
- DLS (Policy Engine): Intercepting and enforcing Document-Level Security rules dynamically.
Query Interpreter & CleaveQL Engine
Queries are passed from the TCP server to the Query Interpreter (powered by Python + PyO3), which bridges the gap to the core CleaveQL Engine. The query execution flows through five distinct stages:
- Parse: The Parser builds an Abstract Syntax Tree (AST) from the CleaveQL string.
- Plan: The Optimizer generates an efficient execution plan based on cost.
- Execute: Executor workers run the plan against the storage engine.
- Filter (Policies): The Policy Evaluator applies DLS filtering (e.g. dynamic masking, field rules).
- Return Results: The final, properly secured payload is returned to the client.
Backend Services
Go Coordinator
Distributed multi-shard management.
- Cluster management
- Fault tolerance
Rust Storage Engine
The custom high-performance core.
- AVX-512 / AVX2 (C++)
- B+Tree / LSM hybrid
- WAL + Compression
- Page cache + Buffer pool
Vector & AI
Native semantic capabilities.
- Transformer embeddings
- ONNX Runtime (~22MB RAM model)
- Semantic search
Storage Layer
The foundational storage layer handles all physical persistence and indexing:
- Documents: Namespaced strictly by tenant (e.g.,
products:david.laptop) to ensure isolated, controlled access. - Graph Bonds: Native functional relationships connecting documents (e.g., owner, friend, bond) to enable JOIN-free traversal.
- Indexes & Vectors: Fast retrieval structures including B+Tree and HNSW-like vector indexes for high-speed similarity search.
- Metadata: Internal schema and system data.
Security (DLS) in Action
The Document-Level Security engine operates synchronously with the Query Interpreter to guarantee that sensitive data is protected at the engine level. It enforces:
- Tenant Isolation
- RBAC (Role Based) & GBAC (Graph Based)
- Field Rules & Dynamic Masking
- Strict Policy filtering
- Session Context evaluation
- Complete removal of sensitive data (not just hiding it)
