Evaluation groups

A

Semantic compilation

Schema, mapping, migration, plan compilation, and cached reuse

B

Execution performance

Native SQL, compiled SQL, federation, materialization, refresh, and provenance

C

Integration complexity

Mapping count, transformation code, repair surface, and propagation cost

D

Migration correctness and human study

Seeded semantic faults, recovery, plan correctness, and blinded scoring

E

Incremental materialization

Update latency, recomputation, freshness, recovery, and equivalence

F

Distributed behavior

Replay, convergence, conflicts, recovery, staleness, and causal ordering

G

Versioned workspaces

Storage, checkout, diff, merge, historical queries, and isolation

H

Federation and communication-aware routing

Planning, latency, transfer, calls, ablation, and prediction error

Results require traceable source data

Every future chart must link to its harness revision, dataset, environment, raw runs, statistical analysis, and failed trials. The protocol makes that evidence contract explicit before measurement.

Read the benchmark planSee the related limitations