Deep Dive: Read Path¶
This document traces the HyperbyteDB query path from HTTP request to InfluxDB-compatible JSON response. It covers TimeseriesQL parsing, ClickHouse SQL translation against native MergeTree tables, chDB execution, tombstone filtering, and result formatting.
Table of Contents¶
- Overview
- HTTP Query Handler
- TimeseriesQL Parser
- Query Dispatch
- SELECT Execution
- Tombstone Injection
- chDB Execution
- Result Formatting
- Cluster Query Behavior
- Metrics
1. Overview¶
Client GET/POST /query
|
v
+------------------------------------+
| HTTP Handler (query.rs) |
| - Extract q, db, epoch, params |
| - Substitute bind parameters |
| - Delegate to QueryService |
+------------------------------------+
|
v
+------------------------------------+
| QueryServiceImpl::execute_query() |
| - Parse TimeseriesQL |
| - Timeout wrapper |
| - Dispatch each statement |
+------------------------------------+
|
v (SELECT statements)
+------------------------------------+
| execute_measurement_query() |
| - Resolve native table name |
| - Translate AST → ClickHouse SQL |
| - Inject tombstone predicates |
| - Execute via chDB |
| - Parse JSONEachRow → Series |
+------------------------------------+
|
v
+------------------------------------+
| chDB (embedded ClickHouse) |
| SELECT FROM MergeTree tables |
| Returns JSONEachRow |
+------------------------------------+
Key source files: adapters/http/query.rs, application/query_service.rs, timeseriesql/, adapters/chdb/query_adapter.rs.
2. HTTP Query Handler¶
File: src/adapters/http/query.rs
Entry points¶
GET /query?q=...&db=...POST /querywith form-encoded or query-string parameters
POST merges form body parameters into query string parameters (body takes precedence).
Parameters¶
| Parameter | Purpose |
|---|---|
q | Query string (required) |
db | Default database |
epoch | Timestamp format: ns, us, ms, s, or RFC3339 strings |
chunked | Enable chunked response for large result sets |
pretty | Pretty-print JSON |
$param | Bind parameter substitution in query text |
Timeout¶
The handler wraps execution in tokio::time::timeout using server.query_timeout_secs.
3. TimeseriesQL Parser¶
Module: src/timeseriesql/ (Influx-compatible query language)
The parser is a hand-rolled recursive descent parser in parser.rs. It splits multi-statement input on ; and dispatches on the first keyword:
| First token | Statement type |
|---|---|
SELECT | Measurement query |
SHOW | Metadata introspection |
CREATE / DROP / ALTER | DDL |
DELETE | Tombstone mutation |
SET / GRANT / REVOKE | Auth admin |
SELECT parsing builds a SelectStatement AST with fields, FROM clause, WHERE, GROUP BY (including time(5m) buckets), ORDER BY, LIMIT/OFFSET, SLIMIT/SOFFSET, fill options, and subqueries.
Expression parsing uses precedence climbing for OR, AND, comparisons, arithmetic, and function calls.
4. Query Dispatch¶
File: src/application/query_service.rs
execute_query() parses the input into Vec<Statement> and dispatches each:
| Statement | Handler |
|---|---|
SHOW DATABASES | metadata.list_databases() |
SHOW MEASUREMENTS | metadata.list_measurements() |
SHOW TAG KEYS/VALUES | Metadata indexes |
SHOW FIELD KEYS | MeasurementMeta |
CREATE DATABASE / DROP DATABASE | Metadata + cluster Raft mutation |
DELETE | Store tombstone + replicate |
SELECT | handle_select() → execute_measurement_query() |
SHOW and DDL statements never touch chDB tables directly — they read or mutate the RocksDB metadata store (and Raft log in cluster mode).
5. SELECT Execution¶
Function: execute_measurement_query()
Steps¶
- Resolve measurement — Extract name from FROM clause; handle regex measurements by querying metadata for matches and UNION ALL across tables.
- Retention policy — Default RP from metadata when omitted.
- Native table name — Build the physical table identifier via
domain/chdb_naming(same naming as the write path). - Translate —
to_clickhouse::translate_native_table(stmt, &table, column_mapping)converts the AST to ClickHouse SQL targeting the MergeTree table directly (notfile()over external files). - Tombstones — Append
AND NOT (predicate)for each stored tombstone (see §6). - Execute —
QueryPort::execute_sql()runs the SQL through chDB. - Parse results — Transform JSONEachRow into InfluxDB v1 series format.
Translation highlights¶
| TimeseriesQL | ClickHouse |
|---|---|
GROUP BY time(5m) | toStartOfInterval(time, INTERVAL 5 MINUTE) AS __time |
mean("field") | avg("field") |
now() - 1h | now64(9) - INTERVAL 1 HOUR |
fill(null) | ORDER BY __time WITH FILL … |
| Subqueries | Inline SELECT |
The internal alias __time avoids collision with the raw time column and is renamed back to time in the result parser.
6. Tombstone Injection¶
DELETE statements do not immediately remove rows from MergeTree tables. Instead:
- The WHERE predicate is converted to a ClickHouse SQL fragment.
- A tombstone record is stored in metadata:
tombstone:{db}:{measurement}:{uuid}. - In cluster mode, the DELETE mutation is replicated via Raft.
On SELECT, inject_tombstone_predicates() loads all tombstones for the measurement and appends exclusion predicates to the WHERE clause. Deleted data is hidden at query time without rewriting stored rows.
7. chDB Execution¶
Adapter: ChdbQueryAdapter (adapters/chdb/query_adapter.rs)
Port: QueryPort
Session model¶
chDB runs inside spawn_blocking. HyperbyteDB opens separate query and write connection pools to the same session_data_path (chdb.query_pool_size and chdb.write_pool_size, default 4 each). The read path uses the query pool only; each connection has its own ChdbClient mutex, so concurrent query tasks overlap when query_pool_size > 1. Ingest/flush uses the write pool, so heavy queries do not block inserts. Tune server.max_concurrent_queries (≥ query_pool_size) to cap in-flight query tasks.
Output format¶
Queries use FORMAT JSONEachRow. Each line is a JSON object:
8. Result Formatting¶
The query service transforms JSONEachRow into InfluxDB v1 series format:
- Parse each line as a JSON object.
- Rename
__timeback totime. - Convert ClickHouse datetime strings to nanosecond timestamps.
- Apply the
epochparameter (ns/us/ms/sintegers or RFC3339 strings). - Group rows by tag combination into separate
SeriesResultobjects. - Apply SLIMIT/SOFFSET for series-level pagination.
SELECT INTO¶
When the query includes an INTO clause, results are written back through the ingestion path as new points in the target measurement.
9. Cluster Query Behavior¶
Full-copy cluster (default)¶
When [sharding] enabled = false, each node executes queries against its local chDB tables. There is no distributed scatter-gather for data queries — clients should route to a healthy peer or use a load balancer.
Schema mutations (CREATE DATABASE, DELETE tombstones, CQ definitions) go through Raft for ordering and are replicated to all peers.
Sharded cluster (experimental)¶
When [sharding] enabled = true, SELECT and SHOW statements on sharded measurements are scatter-gather:
- Resolve the measurement's shard map and select overlapping regions.
- Translate TimeseriesQL to ClickHouse SQL and inject per-region
series_idpredicates. - For each region, execute locally (if this node is an Active peer) or POST to
/internal/shard/query//internal/shard/metadataon Active region peers. - Merge partial results on the coordinator into InfluxDB-compatible JSON.
The scatter loop tries Active peers in order (self → primary → replicas), respects scatter_max_peer_attempts and scatter_peer_timeout_ms, and fails over to replicas when the primary is unreachable. Stale shard epochs abort without retry.
For replication, sync, and sharding control-plane details, see Deep Dive: Clustering.
10. Metrics¶
| Metric | Type | Description |
|---|---|---|
hyperbytedb_query_requests_total | counter | Query requests received |
hyperbytedb_query_errors_total | counter | Failed queries |
hyperbytedb_query_duration_seconds | histogram | End-to-end query latency |
When sharding is enabled, additional metrics track region selection and scatter behavior — see Deep Dive: Clustering.
When enabled, StatementSummary records normalized query text and timing for GET /api/v1/statements.
Related documents¶
- Architecture — hexagonal overview
- Write path — how data reaches MergeTree tables
- API & TimeseriesQL Reference — HTTP endpoints and syntax