← Skarve product page

Saved benchmark evidence for d15b0bc. Release publication is pending. This local report includes public-safe tables and no source pixels. Links to unpublished release documents are shown as text.

Final coalescing sentinel: b7 versus integrated d15

All 30 fresh process lifecycles, 60 complete queries and 30 matched exact output pairs pass. Seven of ten latency medians improve; three regress. All 12 individual latency losses remain below. This is a whole-build comparison with b7, not an ablation that isolates coalescing from the other integrated changes. No new engine operation, build, test or source-pixel read was performed for this report.

Control is b7bccd79e62c3d8e87cf5a2a84319207d98f4e19, installed library SHA-256 1d1e4d3b0366d3fa769c009c04b936745c67276f16c29ec287bad91fcf4efba6. Candidate is d15b0bccdf87654e6e6bd221ce114fb4f129e42d, library 45cea7221247d815f1cf15ad0fe279ed2dd1b9253b675c033a3954a75eb7b67e. These are the frozen qualified installed artifacts. The saved records bind them; this analysis does not rerun artifact qualification.

The generated40 raster is 1025 × 1031 with 40 distinct bands, in two existing physical representations: band_group=1, independent-band band payloads (24,731,327 B), and band_group=40, row_group_v1 payloads (43,567,822 B). The real36 raster is the authorized 512 × 512 native crop with 36 distinct bands, independent-band band_group=1 payloads (28,933,939 B). All have 128-pixel chunks, byte_delta_v1, DEFLATE level 3 and stored summaries. This is not genuine real40 evidence. Exact paths, source hashes, logical digests and writer provenance are in the diagnostic JSON.

Five cells each have three matched fresh rounds. Singlewide uses one polygon and ordinary carve; mixed uses one ordinary eight-polygon cleave per phase. The cold geometry-generation frame is the full source dimension stated above; the distinct retained Q2 frame shifts by (+3.25, −1.5) pixels on the same registered source and engine. These are frames for the frozen polygons, not claims that every query covers the full raster. HTTP5 is 5 ms configured request delay and 64 MiB/s body service; HTTP0 has zero configured delay and the same body rate. The fixed native policy is native_grid_planar_fractional, with sum/support/mean/min/max; no mask expression or median/quantile output is requested.

All latency results. Milliseconds are source registration/open through fully consumed Q, or distinct retained Q2 through consumption; n=3 per runtime and cell/state. Positive change is a regression. The JSON retains unrounded observations and median/min/max for every reported counter; three observations are not a p95 or a significance test.

CellStateb7 median msd15 median msChangeLosing pairs
Generated40 band singlewide, HTTP5Cold7,378.3025,621.634−23.81%0/3
Generated40 band singlewide, HTTP5Retained7,283.4325,483.826−24.71%0/3
Generated40 row singlewide, HTTP5Cold900.368887.297−1.45%1/3
Generated40 row singlewide, HTTP5Retained807.397755.156−6.47%1/3
Real36 band singlewide, HTTP5Cold3,235.0362,381.422−26.39%0/3
Real36 band singlewide, HTTP5Retained55.62858.717+5.55%2/3
Generated40 band mixed batch, HTTP0Cold2,425.4422,497.802+2.98%3/3
Generated40 band mixed batch, HTTP0Retained2,385.8282,433.524+2.00%3/3
Generated40 row mixed batch, HTTP0Cold1,354.3901,327.469−1.99%1/3
Generated40 row mixed batch, HTTP0Retained1,287.4581,280.563−0.54%1/3

Primary CPU and physical transfer. CPU is measured process CPU during the primary operation, not wall time minus I/O. Body bytes are identical between runtimes in each of the 30 matched pairs. All primary accepted byte intervals are nonrepeating within that operation; retained operations may legitimately reread bytes from Q1.

Cell/stateb7 → d15 phase CPU msb7 → d15 physical requestsBody bytes per runtime
G40 band single, cold1,021.806 → 788.4041,131 → 85110,946,276
G40 band single, retained967.766 → 710.9981,122 → 84210,402,532
G40 row single, cold398.741 → 424.82439 → 3218,960,789
G40 row single, retained358.707 → 343.69230 → 2318,417,045
Real36 single, cold484.276 → 431.013437 → 29321,941,738
Real36 single, retained44.768 → 47.7642 → 20
G40 band batch, cold1,473.329 → 1,518.6022,050 → 2,05018,400,167
G40 band batch, retained1,447.280 → 1,468.7642,122 → 2,12218,666,597
G40 row batch, cold830.759 → 806.40161 → 6132,456,581
G40 row batch, retained747.238 → 738.83255 → 5533,292,354

The band-batch regression is observed with unchanged work volumes. All 12 batch pairs have identical ordered physical request method/status/range/body sequences. For every pair, raw encoded bytes, decoded bytes, decoder calls, materialized cells, normalized output writes and native output allocation counters match. Cold/retained band batches decode 2,040/2,120 chunks and 157,824,000/164,377,600 raw bytes, writing 284,083,200/295,879,680 normalized bytes. Each uses 51/53 windows. Candidate prefetch calls, admissions, ranges, planning and fetch time are measured zero in all batch phases. Shared-mask passes and scratch are also zero; the requested statistics do not activate median/quantile reuse. These facts exclude extra fetches, extra decoded cells or active prefetch work as an explanation for these particular losses. They do not prove that inactive branches or other whole-build effects are costless.

Band-batch counter, median msCold b7 → d15Retained b7 → d15
Consumer read_decode_ms1,968.844 → 2,033.3951,951.650 → 1,992.426
Consumer read_window_validation_ms19.955 → 24.25921.430 → 23.089
Source read_ms1,753.162 → 1,777.9361,756.035 → 1,785.455
Source decode_ms85.265 → 86.91788.725 → 89.321
Source predictor_decode_ms73.980 → 79.15971.253 → 71.367
Source normalization_ms59.453 → 65.63058.883 → 59.572
Consumer reduction_ms346.232 → 346.027330.263 → 333.593

The larger measured movement is in the aggregate read/decode side, alongside higher phase CPU and somewhat longer physical server service time. These timers overlap and their medians can come from different rounds; they cannot be summed into a causal decomposition. The saved counters do not identify an individual instruction, allocation, scheduler effect or compiler/code-layout effect. Those remain hypotheses, not diagnoses or optimization recommendations. Row-batch medians improve here, with two individual losses retained; that does not erase earlier batch regressions in other checkpoints.

Prefetch evidence and the retained real36 loss. Candidate generated40 band singlewide saves 280 physical requests per phase; row singlewide saves seven per phase; real36 cold saves 144. Those counts equal the candidate's planned request savings. Encoded/decoded work and transfer bytes are unchanged in every matched pair. Candidate source prefetch planning medians are 15.213/0.352 ms for generated40 band cold/retained, 14.831/0.281 ms for row, and 7.682/0 ms for real36. Cold planning is higher here; the saved aggregate does not subdivide that cost. Prefetch fetch time overlaps the reader/consumer timers and must not be added to them as extra latency.

Retained real36 incurs zero GETs, raw decodes, materialization or new output allocation, with two conditional guard HEADs in both runtimes. Both retain 63,740,916 decoded-cache bytes and perform the same 12 boundary-tile / 166,788 positive-cell / four summary-tile work. Candidate prefetch calls and queue occupancy are zero; its consumer boundary-planning timer is 0.064 ms. The median loss is +3.089 ms, with primary CPU +2.996 ms. The counters therefore do not support blaming network transfer, decoding or that 0.064 ms planning timer for the full loss. No further causal isolation was performed.

Lifecycle costs and memory. These medians include both operations in one process. Wall and process CPU include imports, engine creation, post-query source inspection, close and receipt work. Peak RSS is one process high-water mark, not a per-query allocation count. External HTTP server CPU is not separately measured.

Cellb7 → d15 process wall msb7 → d15 process CPU msb7 → d15 peak RSS MiB
G40 band singlewide14,978.160 → 11,421.7982,298.148 → 1,805.872133.211 → 134.320
G40 row singlewide1,865.861 → 1,775.801906.259 → 890.576130.184 → 129.641
Real36 singlewide3,475.157 → 2,645.468702.229 → 662.476120.246 → 120.184
G40 band mixed5,147.179 → 5,266.4553,246.916 → 3,316.063141.688 → 141.938
G40 row mixed2,786.022 → 2,767.8471,727.757 → 1,704.875132.500 → 132.145

Both runtimes retain zero native raw intermediate allocation. Copy counters are unchanged per pair: independent-band paths report zero copied_bytes; grouped-row paths report 91,750,400 B per singlewide phase, and 157,824,000/164,377,600 B for cold/retained batch restoration, with a 2,304 B native row-scratch peak. This is repeated bounded row-scratch copying, not a retained intermediate of that total size. Native normalized outputs remain allocated and charged, as shown above. Maximum observed RSS is 148,865,024 B; encoded-cache occupancy is at most 4 MiB. Candidate prefetch scratch is a reported 8 MiB bound, not measured RSS; peak demanded encoded-cache charge is 4,003,430 B. No cache capacity was enlarged for this comparison.

Physical accounting and limits. Recomputed from all completed server timelines: b7 has 21,177 requests (21,087 GET + 90 HEAD), d15 19,023 (18,933 GET + 90 HEAD). Each transfers 550,452,237 B. The 2,154-request reduction is entirely GETs. Each runtime includes 30 post-query inspection HEADs outside primary clocks, zero bytes; those remain in the lifecycle totals. All 30 HTTP traces reconcile with record totals, CSV totals and the final source remote counters. There are zero HTTP errors, incomplete bodies, handler failures or source invalidations. All observed requests are conditional. Maximum lifecycle usage is 4,174 requests and 65,748,935 body bytes; maximum single response is 1,375,742 B. These remain below the frozen 8,192-request, 1 GiB job, 768 MiB source and 4 MiB range limits. Child address space remains 2 GiB, one worker, timeout 120 s.

There are 54 complete logical phase traces and six truncated retained band-batch traces. Those six retain 4,096 cumulative entries while aggregate logical reads report 4,168. Their logical trace bytes/counts are lower bounds; no complete logical range claim is made from them. Full physical traces are complete in all cases. For the 54 complete logical traces, raw record counts and lengths also match decoder-call and raw-encoded-byte counters. raw_prefetch logical bytes overlap later raw reads served from cache; adding both would overstate physical traffic.

Every individual latency loss, zero-based frozen round, milliseconds:

Cell/stateRoundb7d15Increase
G40 row single/cold0900.368924.99524.626
G40 row single/retained0852.631877.67125.041
Real36 single/retained172.89577.5964.701
Real36 single/retained253.94158.7174.777
G40 band batch/cold02,446.3592,497.80251.443
G40 band batch/cold12,425.4422,522.01096.568
G40 band batch/cold22,372.3022,406.76834.466
G40 band batch/retained02,410.5312,429.67219.141
G40 band batch/retained12,384.7212,433.52448.803
G40 band batch/retained22,385.8282,438.77252.944
G40 row batch/cold11,285.3261,361.16075.834
G40 row batch/retained01,287.4581,296.1378.678

The inputs are the completed sentinel freeze, original audit summary and its five hash-bound result files. Freeze SHA-256 is 5b54757c1e822f8216a8bcc3ff34624cfddc2c60bce4a5b9d0a799c0e6bf463f; complete receipt is 73f89629eaef4c63409013010da47957bfb7d1d1e3115af6b746918665045b36. Diagnostic JSON has SHA-256 8c0702bb90779ee58a667e5bd975cda9fd12141fda1098ce2aee54832c4e9cbb and records all 30 matched deltas, 60 compact operation costs, 30 lifecycle costs, source/runtime pins and saved-input hashes without duplicating full answers. The stdlib-only derivation helper checks raw record/task/runtime identities, recomputes transport partitions, rechecks exact serialized answer bits and reconciles saved medians. It refuses to overwrite its JSON output. Two initial helper-only guard stops (CSV field-size default; treating truncated logical traces as complete) produced no diagnostic output; the correction retains truncation explicitly. Original data, audits and frozen helpers are unchanged.

Fresh process/source lifecycle does not mean OS-cache cold: no page-cache eviction is claimed. This finite sentinel does not establish universal performance, WAN/provider behavior, genuine real40 performance or the isolated cost of each integrated change.