# Comparison protocol Frozen before measuring the Go candidate. The reference is original application logic at `b74e8ed`, launched through the isolated `adapters.settings` configuration. Python runtime: CPython 3.14.6; database: PostgreSQL 17.10. ## Contract The external Go harness restores a deterministic fixture before each case and compares HTTP status, relevant headers, JSON/text output, and all five blog tables. The checked-in captures are observations, reviewed against source and replayed against Django before Go implementation. Capturing again is an explicit command; verification never rewrites expected results. Normalization is limited to repository path prefixes in tracebacks, UTC-equivalent timestamps, validated generated request-time clocks, unordered tags, and equal-date groups. Contents and multiplicity remain exact. Generated IDs remain exact after sequence reset. Date order outside ties remains exact. The reference renderer truncates timestamp microseconds to milliseconds; database clocks retain microseconds. Transport `Date`, `Server`, content length, connection framing, and JSON whitespace are not contract fields. Framework error bodies remain in scope. ## Workloads and resources - Primary server: Gunicorn 23, two workers with four threads each, persistent DB connections, eight maximum active handler/database slots. Candidate: Go with `GOMAXPROCS=2` and eight database connections. No coverage or SQL tracing in timings. - Small fixture: 3 users, 3 tags, 4 posts, 3 comments. Full fixture: deterministic synthetic data with 1,000 users, 50 tags, 100,000 posts, and 500,000 comments. It reproduces scale and skew, not Faker's literal text or exact original distribution. - All eight endpoints; concurrency 1, 8, 32 for the small workload. Full-data broad reads may use fewer requests and lower concurrency to bound baseline cost and memory use; label those results separately, never substitute them for a measured high-concurrency result. - Five repetitions per declared workload, three warm-up requests per repetition, fixed requests and dataset, full response-body consumption. Errors/timeouts count. - Latency covers HTTP exchange and full body consumption. Report p50, p95, completed requests/second, bytes, failures, CPU seconds per completed request, peak aggregate PSS and RSS across the master and every worker. Linux clock ticks are 100/second. Memory is sampled every 25ms; peaks between samples can be missed. Database server resource use is shared and is not included in application-process CPU/PSS. - Stop verified workloads in `review_analysis` before timing. Run applications separately and record residual load. No machine-wide cache flush or claims of a cold OS/database cache. Warm-up does not prove every PostgreSQL page is resident. ## No-regression decision Evaluate each workload/concurrency independently, not a favorable aggregate. Use repetition-level results, report medians, and a deterministic 10,000-resample bootstrap of the difference of repetition means with a 95% interval. For latency, CPU/request, and peak PSS/RSS, the interval's upper bound must be at or below zero. For throughput, its lower bound must be at or above zero. No additional failures are permitted. Zero-resolution CPU samples and overlapping intervals are explicitly inconclusive, not a proved pass. Increase sample duration for unresolved cases. The goal is no regressions under declared conditions, not proof over every input or machine. Report any unresolved gate. Do not loosen thresholds after observing candidate results. Full-data samples with too few observations for meaningful tail estimation are descriptive and must be labeled that way. ## Preparation correction, protocol v2 The first measurements exposed unequal physical database state. The original seed ran ANALYZE but not VACUUM; small fixture resets did neither. During the Go full-user count timings, comments still required heap visibility checks. Automatic vacuum completed afterward. The longer reference run had already benefited from that maintenance. Initial failed measurements remain checked in. Protocol `vacuum-analyze-v2` explicitly runs VACUUM ANALYZE on the five task tables outside timing, before warm-up and measurement. It does not change shared server configuration, flush caches, or disable durability. Paired reruns use the same preparation and request counts for both targets; the comparator rejects different preparation versions. No acceptance threshold changed. The original broad-read comparisons remain separately labeled as protocol v1. An untimed wait sampler observed row-lock and WAL waits on the mutating detail endpoint. Later unchanged-server measurements did not reproduce the initial 250ms stalls. Those stalls cannot be attributed conclusively to a Go code defect or to table maintenance from this evidence alone. ## Attribution limits Initial results measure the complete rewrite. The final optional Django query optimization control helps separate query-count improvements from remaining runtime/driver/serialization costs; it still does not isolate language alone. ## Known measurement limits The CPU and memory sampler itself consumes resources. Both targets use the same collection procedure, but more worker processes require more `/proc` reads. No absolute latency SLO or production capacity claim follows from this laptop test. The fixture is synthetic. A strict compatibility suite is finite and cannot prove equivalence for every possible malformed request or concurrency interleaving.