fintual-backend-devops-go / plans / 2026-09-18-behavior-harness-and-go-rewrite.md
2026-09-18-behavior-harness-and-go-rewrite.md
Raw

Behavior harness and Go rewrite

Date: 2026-09-18 Status: implementation and bonus complete; final verification and merge pending Baseline revision: b74e8ed (interview)

Why?

The assignment asks for improvements to developer experience, performance, and production readiness. The user wants a Go implementation and an independent test and benchmark harness that captures the existing Django behavior before the rewrite. The harness should make compatibility, database effects, SQL work, and performance improvements inspectable against the same workloads.

The user prefers using their existing PostgreSQL installation. Requiring another server installation, Docker, or a VM would work against that preference.

This file is the living plan and progress record. Update it as decisions settle, work completes, or evidence changes the plan. Implementation is underway.

Settled direction

  • Preserve the Django implementation as the reference for a parallel Go implementation.
  • Build the harness before the replacement implementation.
  • Exercise applications over HTTP and inspect database state independently of their implementation language. Do not import application code into contract tests.
  • Use a standalone Go harness with ordinary tests, an HTTP client, and PostgreSQL access.
  • Test all eight documented application endpoints.
  • Exclude Django admin and its authentication UI from the replacement. The user accepts treating it as unused default scaffolding. Retain useful OpenAPI docs.
  • Do not test, benchmark, set coverage targets for, or reimplement admin. Its absence is an explicit scope exclusion, not a parity failure or missing feature.
  • Compare controlled fixtures and equivalent workloads before claiming improvements.
  • Document and report existing functional bugs; fixing them is out of scope.
  • Preserve observable behavior while improving implementation performance. Do not silently change default pagination, write atomicity, or other behavior.
  • Use the existing local PostgreSQL server. No Docker or VM requirement.
  • Keep shared PostgreSQL configuration unchanged. Use diagnostic adapters with a common SQL report format, direct database inspection, and query plans.
  • The user tracks effort. Do not impose a time budget or maintain an agent time ledger.
  • Track planning decisions, implementation progress, and verification in this file.
  • Quantify and document improvements in a reproducible comparison report. Include regressions, unchanged results, limitations, and raw evidence.
  • Deliver a short NOTES.md explaining what was done and why, deliberate exclusions, and the most valuable next step if another day were available. Keep claims aligned with completed work and actual measurements. Include AI disclosure and transcripts.
  • Treat current responses as the specification, including framework-specific error responses. Do not silently exclude framework debug output, whether plaintext or HTML, from compatibility testing; capture the actual response format first.
  • Require no performance regressions on any in-scope endpoint. Statistical noise needs a defined measurement policy; unexplained or inconclusive results do not pass.
  • CPU and memory are also strict no-regression requirements, not merely report fields.
  • Compare unchanged Django under a production WSGI server for primary performance results; separately record the documented development-server experience.
  • Normalize only specifically identified, justified variable values. Preserve stable error content and verify the meaning of normalized values.
  • Enforce only specified ordering. Preserve chronological order for posts/comments, compare equal-date groups without imposing a tie-breaker, and compare tag contents including multiplicity without requiring incidental order. Query/index optimization may change unspecified order without violating compatibility.
  • Complete behavioral compatibility before production hardening. After compatibility passes, add and verify an explicit production profile while retaining a compatibility profile; document every configuration-dependent behavioral difference.
  • Verify Linux only. Native Go binary and systemd deployment artifacts are the target.
  • Run an optimized-Django comparison only as the final bonus step after the core harness, replacement, verification, operational deliverables, and report are done.
  • Before benchmarks, stop the resource-heavy workloads in the exact tmux session review_analysis and verify their exit. The user explicitly authorized this preparation. Do not stop them during planning or terminate unrelated processes.

Verified starting facts

  • The original repository has three pytest smoke tests and no detailed behavioral spec.
  • The README, Django Ninja schemas, and implementation describe existing behavior.
  • The assignment does not explicitly require Python, but asks for focused improvements to the existing prototype. A rewrite adds scope and needs evidence and justification.
  • The assignment requests NOTES.md, AI-use disclosure, and chat transcripts, and asks candidates not to open a PR against the original repository.
  • Go 1.26.4, uv 0.11.29, mise 2026.8.14, and psql 17.10 are available.
  • PostgreSQL server 17.10 accepts local socket connections. No new server is needed.
  • Default Python is 3.13.14; the application requires Python >=3.14. Whether a suitable alternate interpreter is already installed remains to be checked.
  • pg_stat_statements is available on disk, but shared_preload_libraries is empty. Using it would require a shared-server configuration change and restart.
  • No dependencies have been installed, applications started, or databases changed during this planning session.
  • The non-shallow local Git history contains one original root commit, b74e8ed, which added the complete application. There is no separate admin/auth addition or explanation in the commit message. Only local planning commits follow it.
  • Django admin uses standard configuration and has no blog models registered. Django's built-in authentication users are separate from blog.User. The eight blog endpoints have no authentication configuration or request-user checks.
  • The README does not document administration or superuser setup. Leftover Django scaffolding is the likely explanation for admin, but author intent is unproven.

Design tree

  1. Compatibility policy, settled Q1
    • Report bugs instead of fixing them; preserve observable behavior.
    • Error fidelity, timestamp handling, ordering, and nondeterministic outcomes still need explicit comparison rules.
  2. Replacement scope, settled Q2
    • Replace the eight blog endpoints and retain useful OpenAPI documentation.
    • Do not rebuild Django admin or its authentication UI.
  3. Effort limit, settled Q3
    • The user tracks effort; do not impose a time cap or track hours.
  4. SQL observation strategy, settled Q4
    • No shared-server configuration changes or restart.
    • Diagnostic adapters collect SQL behind a common report format.
    • Per-request attribution, query budgets, and instrumentation overhead still need acceptance criteria.
  5. Comparison policy and delivery target, settled round 2
    • Q5: current behavior is the specification; include error response fidelity.
    • Q6: optimized Django is a final bonus experiment, not a prerequisite.
    • Q7: no endpoint performance regressions.
    • Q8: Linux.
  6. Comparison mechanics, settled round 3
    • Q9: narrow normalization accepted; concrete project inventory requested.
    • Q10: production WSGI server for Django's primary performance reference.
    • Q11: production hardening only after complete compatibility verification.
    • Q12: CPU and memory also must not regress.
  7. Execution design, consolidated for plan confirmation
    • Q13 settled: unspecified ordering is not part of the compatibility contract.
    • Isolated database ownership, fixture lifecycle, and reset safeguards.
    • Application startup and baseline preservation.
    • Case format, deterministic assertions, and database state comparisons.
    • Benchmark datasets, workload sizes, concurrency, resource measurements, thresholds, and reproducibility.
    • Deployment target details and production acceptance checks.
    • Evidence artifacts, CI, and final submission contents.

Coverage proposal

Track implementation coverage and contract coverage as complementary evidence. Coverage collection is diagnostic and must run separately from timed benchmarks.

  • Python: launch Django under coverage.py, disable the development autoreloader, exercise it through the HTTP harness, and stop it in a way that saves coverage. Collect line and branch coverage with HTML and machine-readable reports. Account for subprocess collection if the selected baseline server uses multiple workers.
  • Go: build the application with coverage instrumentation, run it with a task-owned GOCOVERDIR, exercise it through the same harness, and ensure counters flush on clean shutdown or through an explicit diagnostic flush. Convert coverage data with go tool covdata and render it with go tool cover. Go's standard coverage measures statements, not Python-style branch coverage.
  • Keep Django and Go percentages separate. A cross-language percentage comparison does not establish behavioral equivalence or test quality.
  • Maintain a contract matrix of endpoints and behavior cases, including validation, errors, ordering, and database effects. This is the shared coverage measure.
  • Distinguish application coverage from coverage of the harness itself. Merely running the harness with go test -cover does not instrument external applications.
  • Review uncovered relevant paths for missing cases. Do not invent an arbitrary percentage target or require framework/generated-code coverage.

Included in the execution plan. Cover the in-scope application behavior, exclude admin, and keep diagnostic coverage collection separate from benchmark timing.

Round 1: decisions

Q1: What does compatibility require?

Decision: document and report existing bugs instead of fixing them. The user considers functional bug fixes out of scope. Characterization tests should record observable behavior, including deterministic failure side effects, without silently changing the expected behavior for Go.

Implication: default pagination and atomic-write fixes are not authorized by the performance goal. Concurrency races and other nondeterministic behavior need a reporting policy; tests must not require an inherently unreliable outcome.

Q2: What are we replacing?

Recommendation: the eight application endpoints and useful OpenAPI documentation. Exclude rebuilding Django's administration UI, which currently has no blog models registered. This keeps the replacement focused while retaining API discoverability.

Alternative: include administration behavior too. This broadens parity but adds authentication, sessions, forms, and UI work beyond the assignment's stated goals.

Decision: exclude admin. The user accepts assuming it was included by default and is not needed. History supports that interpretation but does not prove author intent. The API's lack of authentication is distinct from admin's built-in login.

Q3: Is there a hard effort limit?

Decision: the user will track effort. Ignore the soft time budget for planning; do not maintain an agent time ledger or ask further time-budget questions.

Q4: How should we observe SQL?

Recommendation: keep the existing PostgreSQL instance configuration unchanged and use minimal implementation-specific SQL diagnostic adapters behind a common report format, plus direct database inspection and query plans. Contract tests remain external. Adapters add maintenance and their instrumentation must be disabled or accounted for during performance measurements.

Alternative: enable pg_stat_statements on the existing server. This offers native aggregate statistics but requires a configuration change and restart affecting the shared instance. It does not, by itself, attribute SQL to individual HTTP requests.

Decision: accepted. Keep shared PostgreSQL configuration unchanged and use the diagnostic-adapter approach. No shared-server restart or configuration change.

Round 2: decisions

Q5: What counts as the same response?

Decision: the current implementation is the specification. The user answered yes to including framework-specific error output. The earlier proposal to exclude Python-specific debug pages is not accepted. Capture these responses and keep them in scope; any relaxation requires an explicit decision.

Q9 resolves the policy for dynamic values: identify variation between reference runs and permit only narrow, documented handling. The concrete inventory below distinguishes known source behavior from runtime facts still to be measured.

Q6: Do we need an optimized Django control?

Decision: do the optimized-Django comparison as the last bonus step, after everything else is complete. The user is interested in measuring how much of the gain comes from Go. Keep the reference untouched and the experimental variant separate. Update the report with the bonus results afterward.

Until that control exists, claim end-to-end rewrite improvements only, not a measured language-only speedup. Even with the control, document remaining server, driver, serialization, and configuration differences.

Q7: What makes performance acceptable?

Decision: strict, no performance regressions on any endpoint. The goal is to improve all endpoints, not offset a slower endpoint with aggregate wins elsewhere. Do not declare success if a regression remains. Define repeatability and confidence rules before evaluating Go so measurement noise is not mistaken for evidence.

Q8: Which native platforms must the workflow support?

Decision: Linux. Verify native development and deployment with existing PostgreSQL, a production binary, and systemd deployment artifacts. Do not claim tested macOS or Windows support. Creating artifacts does not authorize installing a system service on the user's workstation.

NOTES.md editorial decisions

  • The rewrite is an exercise-specific choice: a small application, freedom to choose the language, interest in Go, and a chance to demonstrate a reusable characterization and measurement harness. The user also values making this fun.
  • Say explicitly that a rewrite would not be the automatic choice for an inherited production system. Explain the exercise's bounded scope and use evidence.
  • Do not claim Python developer experience or Django performance is inherently unfixable. The observed N+1 queries have concrete Django fixes; the optional final control experiment can quantify those. Preserve the user's preference without presenting an opinion as a measured technical fact.
  • Deliberate exclusions include admin, authentication/authorization, functional bug fixes, contract changes such as default pagination, unnecessary domain changes, another PostgreSQL installation, containers/VMs, and non-Linux verification.
  • The primary work does not modernize the Python development workflow or optimize the Django application. Minimal baseline launch/diagnostic support is allowed; Django query optimization is reserved for the final bonus experiment.
  • Answer the "another day" prompt using the strongest remaining limitation revealed by the work. It is a prioritization question, not a claim about time required. Do not invent missing work or say no possible future improvement exists before measuring the system. Known-bug remediation is a possible follow-up outside scope.
  • Keep NOTES.md short. Put measurements, full bug descriptions, and reproduction details in the report and link them rather than duplicating them.

Round 3: decisions

Q9: May comparisons normalize proven nondeterminism?

Recommendation: repeat identical sequences against freshly restored Django fixtures to identify values that change between Django runs. Permit only named, justified normalizations and assertions for those fields, such as clock times and server ports. Keep stable error output in scope. Do not strip all HTML, error messages, or timestamps.

Benefit: avoid declaring Django incompatible with itself. Cost: each exception must be reviewed and recorded; normalization cannot establish byte-for-byte equality.

Decision: accepted. The user requests a concrete inventory of variable values in this project. Record source-confirmed facts separately from framework behavior that must be established by repeated baseline execution.

Q10: Which Django server is the performance reference?

Recommendation: record the README's runserver experience, but use a documented production WSGI server for the primary no-regression comparison, with unchanged application logic and explicit resource/connection budgets. This is a server-launch configuration, not the query-optimized Django bonus variant.

Benefit: avoid attributing development-server overhead to Go. Cost: an additional baseline launch dependency and configuration, both disclosed in the report.

Decision: accepted. Do not reopen the server-choice question.

Q11: May production configuration differ from the compatibility profile?

Recommendation: characterize current debug behavior in a compatibility profile; verify production separately with debug output disabled, explicit hosts, validated configuration, and bounded server/connection behavior. List every intentional configuration-dependent difference. Do not call differing production error output identical to the current DEBUG=True reference.

Benefit: address production readiness without quietly erasing baseline behavior. Cost: two explicit configuration profiles and verification of both. Operational hardening must not become a pretext for fixing excluded application bugs.

Decision: accepted with strict ordering: finish matching the Python implementation and verifying compatibility first, then perform production hardening. Preserve the compatibility profile and report production-profile differences explicitly.

Q12: Which metrics must not regress?

Recommendation: gate each declared endpoint/workload on p50 and p95 latency, throughput at specified concurrency, and errors/timeouts. Measure CPU and memory and expose any regressions in the report. Use repeated runs and a predeclared statistical rule; inconclusive comparisons are not proof of non-regression.

Benefit: enforce the user's requirement against specific measurements. Cost: requiring CPU and memory also to never increase would further constrain choices, such as trading memory for faster responses; that policy needs an explicit decision.

Decision: CPU and memory also must not regress. This overrides the recommendation to treat them as report-only metrics. Report and investigate any regression; do not waive it merely because latency improves.

Define CPU consumption per completed equivalent workload and memory measurement scope before running the comparison. Include all Django worker processes rather than comparing one Python worker with the entire Go process. Count errors and timeouts, and do not interpret less completed work as a resource improvement.

Concrete variability inventory

Read-only source findings, not yet observed HTTP captures:

Value Source Comparison policy
Post detail view_count blog/api.py:66-68 Reset fixtures and assert the exact increment; never normalize the counter away
Post detail updated_at blog/models.py:32, saved by GET Assert request-time generation and equality with the stored value
Newly created row timestamps blog/models.py:10,19,31,43 Fixed fixture dates; validate new timestamps and database consistency
Generated post/comment IDs blog/migrations/0001_initial.py, create endpoints Reset sequences and compare exact IDs for sequential cases; follow returned IDs under concurrency
Full-seed dates seed.py:34-38,86-95 Fixed PRNG seed still uses current time; seed once and restore the same snapshot for both targets
Tags and tied-date ordering blog/api.py:36,44,53,60,77,84 Q13 permits unspecified order; preserve explicit date ordering and exact contents/multiplicity

The POST responses expose IDs/title rather than generated timestamps, but subsequent GETs and direct database assertions expose those timestamps. Seeded post updated_at also uses automatic time generation, so separate seed runs are not identical.

Framework/server candidates requiring capture:

  • The HTTP server may generate time-dependent Date and server-dependent Server headers. These are not set in application source.
  • Error diagnostics may contain paths, locations, request URLs/ports, or other environment-dependent content. Do not assume all errors are Django HTML pages: Django Ninja may produce plaintext tracebacks in debug mode.
  • Missing tags and duplicate-email lookups provide concrete unhandled-error cases to characterize. Keep their status/body and partial-write behavior in scope.
  • There are no application-generated request IDs or random JSON tokens in the inspected source. Do not invent normalization rules for nonexistent fields.

Ordering caveat: posts sort only by creation date, comments only by creation date, and tags have no explicit order. Source does not guarantee tie order. Runtime variability has not yet been measured; repeated observations cannot establish a database ordering guarantee that the query does not express.

Round 4: decision

Q13: Must incidental ordering match?

Recommendation: preserve explicit date ordering, compare tied-date groups without assuming a tie-breaker, and compare tags by membership. Include explicit tie cases and verify their contents. This follows the ordering expressed by the reference queries rather than sorting every response indiscriminately.

Benefit: compatible query/index changes are not rejected for unspecified ordering. Cost: this is an explicit exception to matching each observed array byte-for-byte. If incidental order must match too, keep strict array comparison and record the resulting restrictions and fragility instead of claiming the order is guaranteed.

Decision: accepted. Unspecified ordering is unspecified, and optimizations may change it. The user notes that the same freedom would be necessary when optimizing the original Django implementation. Do not introduce a new public ordering promise.

Tests must still catch missing or duplicated items, incorrect date ordering, and changed item values. Compare only equal-date groups without order; do not sort an entire date-ordered response in a way that hides a broken sort. For tags, compare the complete collection with multiplicity rather than a set that hides duplicates.

Comparison report

Required deliverable: a report quantifying and documenting achieved improvements, linked to reproducible commands and raw evidence. A claimed improvement must have a defined measurement or a verifiable before/after property. Do not assign invented percentages to qualitative changes or promise that every metric will improve.

Proposed evidence matrix:

Area Evidence
Behavior Shared contract cases, database effects, parity results, known-bug register
HTTP performance Per-workload latency distribution, throughput, errors/timeouts, response bytes
Database work SQL calls per request, normalized statements, adapter timing, selected execution plans
Resources Process CPU, memory, connections, process count, artifact size under defined conditions
Developer setup Required tools, documented commands/manual steps, measured setup/start/seed times
Operations Demonstrated config validation, startup, readiness, shutdown, logs, and deployment procedure
Coverage Separate Python line/branch and Go statement reports plus shared behavior matrix

Measurement controls:

  • Record application revisions, dependency/runtime versions, hardware, server configuration, fixture identity, request mix, concurrency, repetitions, and caches.
  • Use equal response semantics and equal starting data for comparative workloads.
  • Separate coverage and SQL diagnostic runs from uninstrumented timing runs.
  • Report adapter-measured SQL time as such; it is not pure PostgreSQL execution time.
  • Run implementations separately; reset mutating workloads to equivalent starting states. Do not claim a cache is cold unless the procedure establishes that.
  • Report absolute values and relative changes with run-to-run variation. Include failures and timeouts; a failed or truncated response is not a speed improvement.
  • Distinguish observed facts from predictions and document remaining bottlenecks.
  • Measure setup time separately from engineering effort, which the user tracks.

Use a versioned Markdown report with machine-readable result files and generated figures where they clarify a comparison. No publication or external hosting is implied. The no-regression requirement is settled; predeclare the statistical measurement procedure after establishing baseline repeatability, before evaluating Go. Do not introduce a permitted regression percentage during implementation.

Execution design

This is the proposed concrete execution sequence for final review. Routine filenames and tool flags can change during implementation; contract and acceptance policies above cannot be silently relaxed.

Repository and isolation

  • Keep the original Django source available at b74e8ed; use a worktree and commit implementation milestones there, then merge completed work back per host rules.
  • Put the independent Go harness in harness/, the Go application in go-service/, native operating artifacts in deploy/, and durable reports/evidence in reports/. Neither Go module may import the other. Reuse fixtures and external protocol definitions rather than importing implementation validation/serialization logic.
  • Keep baseline launch/configuration overrides and SQL diagnostics in a clearly separate adapter. Do not optimize or correct reference application logic.
  • Use dedicated, explicitly named task databases on the existing PostgreSQL server. Validate ownership and a task marker before any fixture reset or database removal. Never reset unrelated data, roles, server configuration, or global statistics.
  • Build one fixture snapshot and restore equivalent independent target databases. Reset sequences consistently. Writes and view-count mutations cannot leak from one case or benchmark repetition into another.
  • Store transient captures under ~/codex-tmp when permitted and clean up task-owned files after exporting evidence. Keep database credentials and raw environment dumps out of commits and reports.

Phase 1: runnable reference and fixtures

  1. Install/use a Python version satisfying the lockfile and project requirements.
  2. Establish the original application with isolated database settings and apply its migrations. Run the three provided smoke tests and record any baseline failures.
  3. Create a small deterministic contract fixture covering authors, publication status, tags, comments, duplicate emails, and edge-value inputs. Use fixed times and explicit identities where possible.
  4. Establish a reusable full dataset at the original seed scale. Record schema and data fingerprints and table counts. If the seed is too slow, measure/report that behavior, and generate equivalent data in harness tooling without calling a faster generator the unchanged original seed. Preserve documented distributions.
  5. Capture repeated reference responses and database changes before deciding the exact normalization manifest. Save source-confirmed and observed findings separately.

Exit evidence: reference startup command, smoke-test results, fixture identities, reproducible restoration, and reviewed variable-field inventory.

Phase 2: independent behavior and SQL harness

  • Keep the caller interface small: target base URL, database connection, fixture identity, run mode, and artifact destination. Ordinary Go test cases define scenarios; avoid inventing a general-purpose test language.
  • Cover all eight endpoints: successful reads/writes, empty results, missing IDs, filtering/search semantics, draft behavior, ordering, duplicates, malformed input, validation/coercion, and failure side effects. Exclude admin entirely.
  • Compare statuses, relevant headers, response bodies, and expected database rows after request sequences. Check clock values and generated IDs with specific assertions; do not remove entire classes of data from comparison.
  • Capture HTML error responses too. Reference instability must be measured first. Stable response differences remain failures under the agreed Q5 policy.
  • Record reproducible existing bugs in a separate register with triggering requests, responses, and database effects. Concurrency races get diagnostic evidence; do not demand a particular lost-update count from an inherently variable run.
  • SQL adapters emit request-associated statement fingerprints, counts, durations, errors, and optional sanitized parameter samples. Disable them in timed runs. Obtain selected plans on isolated data; EXPLAIN ANALYZE executes a statement and must not mutate benchmark starting state or unrelated data.
  • Run Python application coverage from external scenarios. Once Go exists, use the same cases to collect Go application coverage, separately from timing workloads.

Exit evidence: baseline characterization suite passes, known-bug register exists, SQL reports identify the slow patterns, and normalization rules are inspectable.

Phase 3: performance baseline

  • Immediately before timed benchmark work, resolve the exact review_analysis tmux session and identify its workload process trees. Stop those workloads and verify their descendants have exited before proceeding. Prefer graceful termination first; any forced termination must remain limited to the verified target processes. Preserve unrelated sessions and workloads. Record the cleanup and residual system load. Apply this preflight to Django, Go, and the final bonus benchmark runs.
  • Measure all eight endpoints on representative successful workloads; include empty, selective, broad, and heavily related-object reads plus write workloads.
  • Use small and full datasets. Bound baseline experiments with explicit timeouts; retain timeout/error observations rather than deleting slow samples.
  • Use documented warm-up, repeated measurements, and a concurrency sweep. Start at 1, 8, and 32 in-flight requests, adjusting only for demonstrated machine limits and documenting any adjustment before candidate comparison.
  • Record p50/p95 latency, completed throughput, errors, timeouts, response sizes, process-tree CPU per completed equivalent workload, and aggregate process memory. Report peak memory and distinguish RSS from proportional/shared-memory accounting.
  • Give both implementations explicit CPU, worker, connection, and request budgets. Record the production WSGI launch configuration. Include all Python workers.
  • Define a reproducible statistical comparison rule from baseline variability before looking at candidate results. No post-hoc endpoint exclusions or tolerance changes. Inconclusive measurements remain unresolved rather than a claimed pass.
  • Measure database work separately from application process resources, on the same PostgreSQL installation with no unrelated server configuration changes.

Exit evidence: repeatable baseline artifacts, workload definitions, resource budgets, and a predeclared no-regression evaluation rule.

Phase 4: compatible Go implementation

  • Implement the eight endpoints with explicit SQL using pgx and preserve the existing schema and contract. Add indexes only with query-plan evidence and verify their write overhead. Do not use silent response truncation as a speedup.
  • Match query/validation/error semantics before optimizing. Keep uncovered cases visible; do not edit expected outputs to make Go pass.
  • Preserve deterministic documented bugs and failure side effects. Do not introduce atomic increments or transactions that silently fix behavior excluded by Q1.
  • Improve database access through equivalent joins/batched loading and compatible indexes, then measure remaining allocation, serialization, and concurrency costs.
  • Run the entire external suite against both implementations. Compare SQL work, collect separate coverage, and run the predeclared performance workloads.
  • Require no endpoint latency/throughput/error regression and no CPU/memory regression for equivalent completed work under the declared conditions. Investigate failures; do not claim universal speed guarantees beyond tested workloads.

Exit evidence: compatibility results, known-bug parity, Go coverage, and a passing comparative performance/resource report. Unresolved mismatches prevent completion.

Phase 5: production profile and developer workflow

Start only after Phase 4's compatibility gate passes.

  • Add explicit environment configuration, startup validation, connection limits, timeouts, graceful shutdown, health/readiness checks, and useful structured logs.
  • Disable debug disclosure in the production profile and validate host/proxy handling where applicable. Document configuration-dependent differences from compatibility.
  • Provide a native Linux build/run/test workflow using existing PostgreSQL, with explicit migration commands and small/full fixture commands. A Go-only developer path must not require installing Django just to initialize a fresh Go database; derive and verify application-schema migrations from the reference schema.
  • Provide a systemd unit and deployment/rollback instructions. Validate artifacts and run local smoke/failure checks without installing a permanent host service.
  • Verify missing configuration, unavailable database, readiness, shutdown under in-flight work, and restart behavior in the isolated task environment.
  • Preserve the compatibility profile and rerun parity to catch accidental changes.

Exit evidence: documented native workflow, operational checks, production artifact, and explicit profile-difference inventory.

Phase 6: report and submission

  • Produce a concise comparison report with all endpoint results, resource results, SQL evidence, known bugs, coverage, methodology, and limitations. Link raw machine-readable artifacts and exact reproduction commands.
  • Finalize short NOTES.md against completed work, including the exercise-specific rewrite rationale, exclusions, and an evidence-backed next priority.
  • Include AI disclosure and actual chat transcripts. Do not fabricate transcripts or substitute an undisclosed summary; establish export availability before delivery.
  • Include CI definitions for the external contract/smoke workflow on PostgreSQL. Leave machine-dependent performance gates in the controlled benchmark workflow; do not make noisy hosted-runner timings a false regression authority.
  • Package a reproducible Git deliverable. Do not open a PR against the source repo.
  • Update this progress tracker with commands, results, commits, and artifact paths.

Exit evidence: reviewable submission with no unsubstantiated completion or speed claims.

Phase 7: final bonus comparison

Only after the preceding phases are complete, create a small separate Django variant with equivalent query-loading improvements for the principal read workloads. Reuse the same harness and budgets. Append results explaining database-query wins versus remaining runtime/driver/serialization differences, and update NOTES.md accordingly.

Progress checklist

The design questions are resolved. The execution stages and acceptance evidence above await the user's shared-understanding confirmation before implementation.

  • Resolve grilling rounds Q1-Q13 and document the execution plan.
  • Confirm shared understanding with the user before implementation.
  • Establish the isolated Django reference environment and fixture lifecycle.
  • Build the external behavior harness and review the characterization cases.
  • Add separate Django and Go coverage collection and a shared behavior matrix.
  • Capture a reproducible Django performance and SQL baseline.
  • Implement Go against the agreed contract and report existing bugs without fixing them.
  • Compare behavior, database effects, and equivalent performance workloads.
  • After compatibility passes, complete the native Linux production workflow and verify its explicitly different configuration profile.
  • Produce the quantified comparison report and link its raw evidence.
  • Finish NOTES.md, evidence, AI disclosure and visible transcript extract.
  • Final verification and submission packaging; merge the implementation worktree.
  • Last bonus step: compare a query-optimized Django variant, then append its measurements and attribution limits to the report and update NOTES.md.

Progress log

2026-09-18

  • Created this dated plan and progress tracker at the user's request.
  • Read the grill-me and grilling skills. The grilling session resolves independent decisions in rounds and requires shared-understanding confirmation before execution.
  • Completed read-only environment discovery. Verified the existing PostgreSQL 17.10 server and identified the pg_stat_statements restart dependency.
  • Prepared round 1. No application implementation or benchmark results yet.
  • Added a proposal for external-test-driven application coverage after the user asked whether coverage can and should be tracked.
  • Recorded round 1 answers: report bugs without fixes, user-owned effort tracking, and SQL diagnostics without shared PostgreSQL configuration changes.
  • Completed the requested read-only Git-history investigation of Django admin. Recorded evidence and the limits of conclusions about author intent.
  • Added the user's requirement for a quantified improvement report and proposed an evidence matrix covering performance, behavior, resources, setup, and operations.
  • Settled Q2 after the user accepted excluding admin as unused default scaffolding.
  • Recorded the user's explicit clarification: no admin tests or reimplementation; admin is outside the harness, benchmark, coverage, and replacement scope.
  • Prepared round 2 for response parity, performance attribution and acceptance, and native platform support. Application implementation remains unstarted.
  • Recorded Q5-Q8: include existing error output, defer optimized Django until the final bonus step, require no endpoint regressions, and target Linux.
  • Expanded the NOTES.md deliverable and created an explicitly marked planning draft. Recorded the rewrite rationale and the need to base future-work suggestions on evidence.
  • Prepared round 3 for nondeterminism, fair server configuration, production-profile differences, and the exact dimensions of the no-regression requirement.
  • Recorded Q9-Q12: narrow normalization accepted, production WSGI reference confirmed, production hardening deferred until parity, and CPU/memory included in the strict no-regression requirement.
  • Started read-only source inspection for the concrete output-variability inventory requested by the user. No baseline requests have run yet.
  • Completed the source inventory. Recorded specific timestamp/ID/state handling, corrected the assumption that all framework errors are HTML, and identified the remaining decision about unspecified ordering.
  • Expanded the execution plan with phase-specific acceptance evidence, fixture isolation, workload measurements, production sequencing, and submission artifacts.
  • Recorded the user's instruction to stop workloads in review_analysis before benchmarks. No processes have been stopped during planning.
  • Settled Q13: unspecified ordering may change under optimization. Added precise collection-comparison rules that preserve explicit ordering, contents, and duplicates.
  • Consolidated the plan for review. All design questions asked so far are resolved; baseline-dependent runtime facts remain execution tasks, not assumed results.

Verification record

Implementation started

  • User approved execution after Q13. Created branch implement/go-harness in the sibling backend-devops-go worktree. Original Django application files remain intact.

  • Found installed CPython 3.14.6 at /usr/bin/python3.14; no Python installation needed.

  • Confirmed local PostgreSQL access. All task data will use dedicated databases.

  • Original smoke tests: 3 passed. External characterization: 53 cases captured and replayed successfully against both runserver and production Gunicorn.

  • Django application coverage: API 100%, schemas 100%, models 91%; combined 97%. The uncovered model paths are string-display methods, outside the HTTP interface.

  • Narrow clock handling now reflects measured Django millisecond rendering versus database microsecond storage. Framework 500 responses are plaintext tracebacks.

  • Stopped the three verified review_analysis workload processes with SIGTERM and verified exit. Both tmux shells remain available; unrelated workloads were untouched.

  • Small-fixture baseline completed at concurrency 1, 8, and 32, five repetitions of 200 requests per endpoint, with zero errors. Raw data is in reports/raw/.

  • Loaded the deterministic full fixture: 100,000 posts, 500,000 comments, 243,116 links. Its text/distribution differs from Faker and is explicitly labeled.

  • Began the bounded full-scale baseline. Go compatibility source is being written; no candidate benchmark or production hardening has occurred yet.

  • Revision 25fb1b7: all 66 characterization cases pass both implementations, including sequence state and partial failures. Full-table fingerprints and all six full-data read responses match. Evidence: reports/full-equivalence.json.

  • Go statement coverage is 86.5% for the compatibility revision. Python and Go coverage reports and request-associated SQL evidence are checked in separately.

  • Initial uninstrumented Go measurements show large read improvements, but the strict gate fails on some writes and full-data user counts. Failed candidate runs and comparisons remain in reports/raw/ and reports/comparisons/.

  • Untimed PostgreSQL sampling found transaction/tuple lock waits and WAL sync during concurrent hot-row reads. A subsequent unchanged-server run did not reproduce the earlier 250ms stalls. Root cause of those initial stalls remains unproven; do not describe a measured rerun as an application fix.

  • Found a preparation defect: only the full seed analyzed tables, and neither fixture controlled visibility-map state. Background vacuum occurred after the fast Go full-user timings, unlike the much longer Django run. Added explicit task-table VACUUM ANALYZE outside timings for both targets, versioned the preparation protocol, and began paired reruns with 1,000 requests/repetition. Cross-protocol comparisons are rejected. Thresholds and exclusions are unchanged.

  • Corrected paired measurements pass all 24 small endpoint/concurrency groups and the two full-data user-count groups, including CPU and memory. The six other full-data groups pass their original paired comparison. bin/compare selects these declared artifacts and exits successfully. Broad full reads have only five single-request observations, so their tail estimates remain descriptive.

  • Phase 4 gate passed. Proceeding to the explicitly different production profile, native migrations, systemd example, operational tests, and documentation.

  • Native setup and idempotent migrations verified. Existing untracked schemas are refused. Added a default production profile, bounded requests/connections, host checking, debug suppression, health/readiness, JSON logs and graceful drain. Black-box SIGTERM test completed a blocked request and exited zero.

  • bin/check passed original smoke tests, unit/race/vet checks, migration and operational tests, and all 66 cases against both implementations. CI uses this native command but has not run on a hosted runner. No systemd service installed.

  • Final Go external coverage is 82.9%, including new operational and migration code; retain the earlier 86.5% compatibility-only report separately.

  • Post-hardening production profile passes all 24 small-data comparison groups with normal access logging enabled. Core report, NOTES, deployment/development instructions and visible transcript extract are written. Proceeding to final bonus query-optimized Django control; original Django source stays untouched.

  • Bonus complete: query-optimized Django passes all 66 cases and all six full-data responses match Go. Paired v2 full listing/search/tag measurements pass all Go gates. Mean repetition medians are 8,732/1,550/1,425 ms for optimized Django and 252/131/63 ms for Go. Report includes attribution limits and raw comparisons. Added the optional variant to the native check workflow for future maintenance.