cqlite-flight — Round 12 Field Validation

Observability round · flight round12@sha256:2433dde8 · connector 0.14.3 · Trino 481 · Cassandra 5.0 RF=3 · 3× i4i.xlarge · ~1.93M partitions/node, 2 SSTable gens · 2026-07-15

Verdict: ALL GREEN — every R11b baseline held, and the round's headline (the #2419 in-process saturation gauges) is visible under overload where R11b showed only zeros. The #2436 large-cell fix produced a legitimate row-count increase. No restarts, no OOM, no regressions.

1. Regression parity vs R11b — held or better

MetricR11bRound 12Bar
warm LIMIT 52.3–3.0s2.3–2.5s≈ JDBC floorPASS
warm LIMIT 1002.30s2.43s≈ floorPASS
point-read (pk=)2–3s2.3–3.5s≈ floorPASS
warm throughput (8-thr)~34 qps~33 qps (679 rows/s)≥ parityPASS
warm p50 / p99227 / 366ms242 / 363ms (steady)≤300 / ≤500msPASS
count(*) wall66.2s61.1s≥ parityPASS
count(*) rows1,939,2861,947,775 (+8,489)#2436: may increaseFIX
index parses (#2412)absent (0)absent (0)zero coldPASS

#2436 large-cell fix confirmed: identical data returned +8,489 rows vs R11b — rows with a ≥~1MB single cell that were previously silently dropped now read correctly. A fix, not a discrepancy.

RPC requests/sec by method — do_get ok ramps through the test phases, errors near zero
RPC requests/sec by method — do_get ok ramps through the test phases, errors near zero
RPC duration p50 by method
RPC duration p50 by method
RPC duration p99 by method
RPC duration p99 by method
Rows/sec streamed by method
Rows/sec streamed by method

2. #2419 saturation gauges — the headline VISIBLE

During the 80-thread scan overload, the new in-process gauges produced a legible saturation trace — the exact gap flagged in R11b (where admission_in_use read a flat 0). All rose under load and returned to baseline after drain (drift-free).

GaugeR11bR12 peak under 80-thrPost-drain
cqlite_flight_blocking_tasks_in_useflat 0 (invisible)80
cqlite_merge_egress_channel_depth35050
cqlite_flight_admission_in_use_ratio (limit 64)flat 0 (invisible)120
cqlite_proc_threads93baseline
cqlite_proc_fds135baseline
cqlite_proc_rss_bytes~603 MBbaseline

Note: OTLP export cadence is ~15s, so the gauges sample coarsely in VictoriaMetrics — peaks are captured but the trace is sparse (points, not lines). The in-process gauge is ~2s; a shorter export interval would render a smoother curve. Success criterion (overload visible, not flatlined) is met.

merge_egress_channel_depth — backpressure spikes under load (peaks ~1850/1150/270) then returns to 0 (drift-free)
merge_egress_channel_depth — backpressure spikes under load (peaks ~1850/1150/270) then returns to 0 (drift-free)
blocking_tasks_in_use + admission_in_use_ratio — rise under overload, settle to baseline
blocking_tasks_in_use + admission_in_use_ratio — rise under overload, settle to baseline
process RSS + thread count during load
process RSS + thread count during load

3. Fan-out & memory — bounded, no leak

8-thread load fans across all 3 flight pods (94–377m CPU). Memory bounded and matches R11b's summary-only profile.

PointMemory (per flight pod)
idle (pre-query)3 Mi (all pods)
warm steady-state (8-thr)~24–344 Mi (busy pod 341Mi)
peak (80-thr overload)~603 MB (heaviest); others far lower
OOMKills / restarts0 / 0
RPC ok vs error — client-visible errors flat at zero across all phases
RPC ok vs error — client-visible errors flat at zero across all phases
Flight error rate (% of requests)
Flight error rate (% of requests)

4. #2452 snapshot grace-sweep probe

=== #2452 snapshot grace-sweep probe ===
post-burst count was: 738 (at 22:50:33)
count now (>10min after burst, before probe query): 738
--- running ONE query against keyvalue (burst table) to trigger lazy sweep ---
query ok
count after probe query + 20s: 18
delta: -720 (negative = grace-sweep pruned; #2452 mechanism confirmed)

5. Resilience

CheckResult
Flight pod restarts (whole round, incl. 80-thr overload)0
OOMKills0
do_get status (whole round)8413 ok / 102 error (1.2%) — aborted/superseded splits, 0 client-visible
Digest pin (all 3 pods)round12 @ 2433dde8