# A compatible Go implementation, with measured improvements The eight endpoints now have a parallel Go implementation and an independent HTTP, SQL and database test harness. All 66 characterization cases pass both implementations. All six full-data read responses match after the agreed narrow normalization. Original Django application source remains unchanged. The declared compatibility-profile comparison passes 32 endpoint/concurrency groups, including latency, throughput, errors, application CPU and memory. This includes corrected paired reruns, not a claim that every initial experiment passed. Full listing dropped from about 93 seconds to 263 ms on the synthetic dataset. Most of that gain comes from eliminating repeated queries, not evidence that Python is incapable of serving this application efficiently. ## Endpoint measurements Values below are means of five repetition-level medians, in milliseconds. Small data uses 1,000 requests/repetition at concurrency one under preparation protocol v2. Broad full reads use one measured request/repetition under v1, so those three values describe full-response duration and cannot establish production tail latency. | Workload | Django small | Go small | Django full | Go full | | --- | ---: | ---: | ---: | ---: | | Posts | 3.774 | 0.185 | 93,148 | 262.773 | | Search, python | 3.128 | 0.209 | 18,717 | 136.728 | | Tag, python | 3.570 | 0.227 | 17,001 | 70.774 | | Post detail | 5.119 | 1.389 | 117.218 | 2.186 | | User | 1.914 | 0.138 | 2.806 | 0.947 | | Email lookup | 1.884 | 0.141 | 2.872 | 1.009 | | Create post | 4.866 | 2.642 | 5.391 | 2.226 | | Create comment | 3.136 | 1.405 | 3.561 | 1.480 | The full detail/create/comment pairs use 200 requests/repetition under v1; the full user/email pairs use 1,000 under v2. No cross-protocol comparison is used. Results at concurrency 8 and 32, p95, throughput, individual latencies, bytes, CPU ticks, memory samples and bootstrap intervals are in [raw](raw/) and [comparisons](comparisons/). `bin/compare` reproduces the declared gate decisions. On the hot-post small workload at concurrency 32, mean repetition p95 fell from 127.5 ms to 38.7 ms in the corrected pair. The existing lost-update race remains; this was not achieved by silently making the counter atomic. ## SQL work explains the largest gains Untimed full-fixture SQL adapters recorded these statement counts per request: | Workload | Django | Go | | --- | ---: | ---: | | Posts | 180,205 | 2 | | Search | 36,007 | 2 | | Tag | 33,558 | 3 | | Detail | 282 | 4 | | User / email | 3 | 1 | | Create with one tag | 4 | 4 | | Comment | 3 | 3 | The Go implementation joins authors and batch-loads tags, joins comment authors, and obtains user counts in one query. It retains the existing schema and indexes. It preserves the original partial-write sequence and read/modify/full-row-save counter behavior. Unspecified tag and equal-date ordering may change. Evidence: [Django SQL](django-sql.jsonl), [Go SQL](go-sql.jsonl), and [SELECT plans](diagnostics/go-full-v2-plans.json). For example, the prepared full user query took 1.215 ms inside EXPLAIN ANALYZE, with no heap fetches for the count scans. Instrumented query durations are diagnostics, not HTTP benchmark timings. Recorded counts include connection-initialization SQL where present. The original listing capture includes two driver type lookups in addition to application queries. ## Application resources For small concurrency-one workloads, average per-repetition peak PSS was about 107 to 109 MiB for Django and 11.6 to 12.4 MiB for Go. Application CPU/request fell from 1.506 to 3.316 ms to 0.076 to 0.254 ms, depending on endpoint. These are process-tree measurements, including the Gunicorn master and both workers. For full listing, average sampled peak PSS fell from 1,241 MiB to 165 MiB, about 86.7%. Full responses remain large; the rewrite does not make unbounded response sizes safe at arbitrary concurrency. Shared PostgreSQL CPU and memory are excluded, and sampled peaks can miss short-lived allocation peaks. ## What the first failed runs taught us Initial comparisons failed for some writes and full user/email lookups. The files without `v2` remain checked in, including their failing comparisons. The initial seed only ran ANALYZE, and small resets did not refresh statistics. Database visibility state therefore depended on when autovacuum happened. Go's full-user measurements occurred before comment-table vacuum, while the much longer reference workload had already benefited from it. V2 explicitly vacuums/analyzes the five task tables outside timing for both implementations and rejects mismatched preparation versions. Its paired count timings pass the unchanged thresholds. The initial write stalls reached roughly 250 ms. Untimed sampling found row-lock and WAL waits. An unchanged-server repeat did not reproduce those stalls, and the controlled paired reruns pass. This evidence does not conclusively identify the initial stalls' cause. It would be wrong to describe the rerun as a Go code fix or to discard the failed results. See [diagnostics](diagnostics/) and the full [measurement protocol](methodology.md). ## Developer workflow and operations The native path builds both Go modules with `bin/build`. `harness/harness` initializes an owned test database and loads deterministic data; `go-service/contentd migrate` applies the schema and `go-service/contentd` serves the API. Python/uv are needed only for the reference workflow. There is no container, VM, separate database installation, hardcoded password, or manual schema SQL step in the native path. See [development instructions](../DEVELOPMENT.md). Production defaults to validated configuration, explicit schema version/checksum, generic server errors, exact host checking, bounded connections and admission, body/header limits, deadlines, health/readiness and JSON logs. Migrations are an explicit transactional command, never a side effect of startup. A systemd example and rollout/rollback instructions are in [deploy](../deploy/README.md). [Black-box operational checks](operations.json) verify health/readiness, host/proxy policy, error suppression, both known-length and chunked body limits, preserved partial writes, and successful SIGTERM draining of an in-flight blocked request. Unit tests also cover admission rejection, failed readiness and panic suppression. The native check command passed locally. CI configuration is included but has not been run on a hosted runner. No systemd service was installed; local unit validation reported the expected absent `/opt/contentd/current/contentd` executable. These are demonstrated workflow properties, not invented setup-time percentages. No clean-laptop timing comparison or production deployment/capacity test was done. The final production profile also passes all 24 small-data groups against the same reference pairs, with JSON access logging and safeguards enabled. These additional artifacts are named `go-production-v2-small-*` and `production-v2-small-*`. They do not replace the compatibility-profile full-data evidence. The final [verification log](final-check.log) records the locally completed checks, including all three 66-case suites after adding the bonus variant. ## Behavior, coverage and exclusions - Shared contract: [66-case matrix](contract-matrix.md), including database and identity-sequence effects. [Full equivalence](full-equivalence.json) includes fingerprints for 100,000 posts, 500,000 comments and 243,116 tag links. - Django coverage: 98/101 application statements and both measured branches; API/schema code is fully covered. [Machine-readable report](django-coverage.json). - Go compatibility revision: 86.5% statements. Final expanded application: 82.9% across external compatibility, production and migration runs from one instrumented build. [Final report](go-final-coverage.out). Config failures, uncommon storage failures and forced-shutdown paths remain incompletely covered. Python line/branch and Go statement percentages are not interchangeable. - Original smoke tests: three pass. Both Go modules pass race tests and vet. - Admin, authentication/authorization, functional bug fixes, pagination, live-data cutover and non-Linux support are deliberately excluded. Known behavior problems are [documented](known-bugs.md), not silently repaired. ## Reproducibility and limits Reference application: `b74e8ed`. Compatibility implementation and initial full-data evidence: `25fb1b7`. Corrected paired evidence: `e472e95`. The operational profile is in `918e02a`; compatibility SQL and business behavior remain the same. Benchmark applications use CPython 3.14.6/Gunicorn 23 or Go 1.26.4, PostgreSQL 17.10, Linux 6.18.44 on an Intel Core Ultra 5 135H with 18 logical CPUs and about 30 GiB RAM. Gunicorn uses two workers/four threads each; Go uses GOMAXPROCS=2; both have eight database slots. This is a laptop microbenchmark, not an isolated production host. The full fixture is deterministic synthetic data with fixed PRNG/time inputs, not literal Faker output. The API remains unpaginated. Broad full-data concurrency was not swept because the original workload is expensive. Short fast-endpoint runs do not establish sustainable capacity. Caches were not globally flushed. Functional tests cannot prove equivalence for all inputs or concurrent interleavings. ## Final bonus: query-optimized Django After the core deliverables were complete, a separate [Django control](../experiments/optimized/README.md) added eager authors, tag prefetching and field-limited projections to the three collection endpoints. The other five handlers remain the original functions. All 66 contract cases pass, and [all six full responses match Go](optimized-equivalence.json). This final matched pair uses v2 preparation, three measured requests per repetition, three warm-ups, five repetitions, concurrency one, and the same CPU/connection budgets. Values are means of repetition medians, not reliable tail estimates. | Full workload | Optimized Django, ms | Go, ms | Duration ratio, Django/Go | | --- | ---: | ---: | ---: | | Posts | 8,732.025 | 252.461 | 34.6× | | Search | 1,549.918 | 131.187 | 11.8× | | Tag | 1,425.492 | 62.768 | 22.7× | All three matched comparisons pass latency, throughput, error, CPU and memory gates. Mean per-repetition peak PSS for listing was 781 MiB for optimized Django and 167 MiB for Go. Listing application CPU/request was 8,186 ms versus 259 ms. Evidence: `raw/optimized-v2-full-*`, `raw/go-bonus-v2-full-*`, and `comparisons/bonus-v2-full-*`. SQL diagnostics are in `optimized-sql.jsonl`. The optimized control issues two application queries for listing/search and three for tag filtering. Its first listing/search diagnostic captures each total four, including two driver type lookups on opening the worker's connection. The original roughly 93-second listing is therefore not an inherent Django limit: ordinary query optimization reduced it to about 8.7 seconds here. That before/after observation spans different preparation/request protocols and is descriptive, not another strict paired gate. The optimized-Django/Go comparison above is matched. Equal query counts still do not isolate the programming language. Django constructs ORM instances and uses its own validation/serialization pipeline; SQL shape, drivers, allocation patterns and JSON encoding also differ. The measured result is that this Go implementation remains faster and smaller than this optimized Django variant for these workloads. It is not proof that Python cannot be optimized further.