Getting Started

Tutorial: Building with CleaveDB

Aggregation (DISTILL) — Overview

When querying collections at scale, returning individual documents over the network just to compute totals or statistics is slow and inefficient. CleaveQL introduces the DISTILL command to execute high-performance aggregation pipelines directly on the storage engine without materializing or transferring individual records.

Key capabilities of DISTILL

  • Mathematical Reductions: Compute TOTAL, SUM, AVERAGE (or AVG), MIN, MAX, and SPREAD (range between max and min) on numeric attributes.
  • Document Counting: Use COUNT or TALLY to count all documents in a bucket, or count only documents containing a specific field.
  • Simultaneous Multi-Metrics: Compute multiple comma-separated aggregations in a single query pass.
  • Selective Filtering: Filter input documents with WHERE clauses before reduction.
  • Categorical Grouping: Group records by a categorical field with GROUP BY and compute metrics per group.
  • Custom Aliases: Re-label result keys cleanly using AS alias for direct API consumption.

Explore the Aggregation topics

Step through the lessons below to learn all canonical DISTILL aggregation forms: