A result is only useful
with its context.
This is the proposed publication standard for evaluation records, not a claim that every product has already met it.
Define the question
Specify the capability, operating conditions, metric, baseline and acceptance threshold. Separate intended behaviour from directly measured behaviour.
Freeze the comparison
Version the dataset, test harness, product revision and configuration before the result is interpreted. Retain held-out cases and document any exclusions.
Make the result reproducible
Store raw outputs and the exact commands needed to recreate the run. Report the population measured, relevant uncertainty and the difference from the baseline.
Publish the boundaries
Record failed cases, untested conditions and known limitations. A result for one workload, provider or repository size does not establish general performance.