# Evidence-backed software factory pilot

- **Date:** 2026-09-02
- **Status:** Continuation active; trial 1/5 completed as PR #333, which remains open for a separate human merge decision
- **Branch:** `codex/software-factory-pilot-20260902` from `origin/dev` at `181445d948610ebd7b665d7333a4f6d0cbc7e7a3`

## Hypothesis

> StratoFusion can make agent-assisted changes easier and safer to review by combining structured intake, task isolation, deterministic checks, and private runtime evidence, while retaining its existing architecture, CI, CASA gates, and human merge authority.

This is a falsifiable pilot. A `do not adopt` result is valid.

## Inputs and interpretation

The external workflow is useful as a menu of ideas, not a repository operating contract. The creator's [software factory video](https://youtu.be/blI10_91xgA) and [skills repository](https://github.com/michaelshimeles/skills) propose isolated task branches, service boundaries, runtime evidence, before/after handoff, and iterative AI review. The video's description identifies Cursor as its sponsor. Its speed, throughput, and rebuild claims are therefore treated as creator/vendor self-report, not independent evidence.

Primary product documentation confirms that Cursor Cloud Agents use isolated VMs and separate branches and can produce screenshots, videos, and logs. Those capabilities are not being installed or evaluated in this pilot. Greptile markets an iterative `5/5` review loop and currently lists a free individual tier plus a paid Pro tier. Neither the score nor the service is part of this pilot. See [Cursor Cloud Agents](https://cursor.com/docs/cloud-agent), [Greptile's workflow description](https://www.greptile.com/independence), and [Greptile pricing](https://www.greptile.com/pricing).

GitHub supports structured issue forms and repository pull-request templates. It also supports workflow artifacts and private-repository attachments, but neither mechanism makes unsafe evidence safe: evidence still needs deliberate redaction and access control. This pilot keeps generated evidence local and performs no uploads. See [issue and pull-request templates](https://docs.github.com/en/communities/using-templates-to-encourage-useful-issues-and-pull-requests/about-issue-and-pull-request-templates), [private attachments](https://docs.github.com/en/get-started/writing-on-github/working-with-advanced-formatting/attaching-files), and [workflow artifacts](https://docs.github.com/en/actions/concepts/workflows-and-actions/workflow-artifacts).

## Existing StratoFusion controls

- `AGENTS.md`, the public AI operating protocol, and focused repo-local skills already define architecture, safety, testing, and documentation boundaries.
- Feature work normally targets `dev`; `main` is the production/release branch.
- CI already runs lint/typecheck, two deterministic Vitest shards, worker checks, a production build, the CASA security gate, and anonymous loopback Playwright security journeys. The aggregate `build-and-test` check remains authoritative.
- Playwright defaults to a fresh loopback production server and Chromium context. Remote origins and credential-backed suites are capability-gated.
- `.auth/`, `test-results/`, `playwright-report/`, `.tmp/`, and local environment files are already ignored.
- Existing screenshot/video utilities are purpose-built for Google OAuth verification and use credentials or mutation-capable flows. Reusing them for routine PR evidence would widen scope and risk.
- The repository had no issue forms or pull-request template at the pilot baseline.

## Recent pull-request baseline

Sample: the eight most recently merged, non-Dependabot pull requests available on 2026-09-02. Data comes from GitHub PR bodies, commits, comments, changed-file metadata, and check runs. `Explicit acceptance criteria` means a separately stated, testable acceptance-criteria section, not an inferred summary. `Review/revision signal` reports observable commits and reviews; it is not elapsed review effort.

| PR                                                                                          | Base <- head                                                | Files | Explicit acceptance criteria | Reported verification                                                 | Runtime or visual evidence                                                             | Privacy/deployment boundary                                                                | CI result                  | Review/revision signal              | Automated reviewer                                        |
| ------------------------------------------------------------------------------------------- | ----------------------------------------------------------- | ----: | ---------------------------- | --------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ | -------------------------- | ----------------------------------- | --------------------------------------------------------- |
| [#332](https://github.com/rikster/stratofusion/pull/332) Google OAuth verification response | `dev` <- `cdx-gqc007v/google-oauth-verification-response`   |    28 | No                           | Exact commands/results, suite counts, build, Playwright, CASA         | Evidence package described; no PR-linked before/after artifact                         | Explicit: no dashboard mutation, deployment, credential use, video upload, or Google reply | All reported checks passed | 1 commit; no GitHub review objects  | None observable                                           |
| [#330](https://github.com/rikster/stratofusion/pull/330) reusable local-VM images           | `dev` <- `cdx-c644i0q/local-vm-multi-machine-release`       |     6 | No                           | Commands plus a GitHub Actions run and OCI revision labels            | Runtime/image identity values reported; no visual artifact                             | Test-key and immutable-release boundary stated                                             | All reported checks passed | 1 commit; no GitHub review objects  | None observable                                           |
| [#329](https://github.com/rikster/stratofusion/pull/329) infrastructure management links    | `dev` <- `cdx-c644i0q/infrastructure-management-tools`      |    25 | No                           | Focused tests, lint, typecheck, Compose validation, live local E2E    | Live local observations described; no linked screenshot/video                          | Read-only UI, private URLs, no host controls                                               | All reported checks passed | 2 commits; no GitHub review objects | None observable                                           |
| [#328](https://github.com/rikster/stratofusion/pull/328) admin MFA freshness                | `dev` <- `cdx-gq007v/casa-admin-mfa-freshness`              |    18 | No                           | Exact pass counts across focused/full suites, build, Playwright, CASA | Redacted runtime proof explicitly left outstanding                                     | Explicitly no certification/readiness claim or production proof                            | All reported checks passed | 1 commit; no GitHub review objects  | None observable                                           |
| [#327](https://github.com/rikster/stratofusion/pull/327) two-admin MFA evidence             | `dev` <- `cdx-gq007v/casa-mfa-two-admin-evidence`           |    10 | No                           | Exact pass/skip/environment-limited results and reruns                | Redacted evidence recorded in repository files; negative-boundary evidence outstanding | Detailed exclusions for identities, tokens, claims, IP/device data                         | All reported checks passed | 2 commits; no GitHub review objects | CodeRabbit was requested; completed review not observable |
| [#326](https://github.com/rikster/stratofusion/pull/326) runner-platform framing            | `main` <- `cdx-c644i0q/generalize-runner-platform-framing`  |     3 | No                           | Social validation, focused content tests, check, diff check           | Not applicable and not stated                                                          | Non-endorsement boundary stated                                                            | All reported checks passed | 2 commits; no human review objects  | CodeRabbit summary; a later run was rate-limited          |
| [#325](https://github.com/rikster/stratofusion/pull/325) CI article punctuation             | `main` <- `cdx-c644i0q/remove-em-dashes-and-publish-social` |     5 | No                           | Content tests, social validation, check                               | Not applicable and not stated                                                          | Not stated                                                                                 | All reported checks passed | 1 commit; no human review objects   | CodeRabbit summary                                        |
| [#324](https://github.com/rikster/stratofusion/pull/324) tighten CI article                 | `main` <- `cdx-c644i0q/tighten-ci-article`                  |     6 | No                           | Content tests, social validation, check                               | Not applicable and not stated                                                          | Runner-time caveat stated                                                                  | All reported checks passed | 1 commit; no human review objects   | CodeRabbit comment was rate-limited                       |

### Baseline observations

- 0/8 PRs used an explicit acceptance-criteria section.
- 8/8 reported verification and 8/8 finished with all sampled CI checks passing.
- 4/8 described runtime or repository evidence; none linked a consistent before/after evidence package in the PR body.
- 6/8 stated at least one privacy, deployment, or claim boundary; two small editorial PRs did not state one.
- Observable commits ranged from one to two. GitHub exposed no human review objects in the sample, so review effort and true review-cycle counts cannot be determined.
- CodeRabbit participation was observable on four PRs, including rate-limit-only responses. An AI reviewer status or score is not evidence of correctness.
- The sample is small, single-author, and skewed toward security, deployment, and editorial work. It cannot establish causal time savings.

## Proposed pilot slice

1. Add two structured issue forms that collect implementation-ready facts while warning against secrets, customer data, private identifiers, filenames, storage, and production logs.
2. Add one concise PR template that connects purpose, acceptance criteria, architecture/safety boundaries, verification, private evidence, privacy review, recovery, and human authority.
3. Add one experimental, opt-in repo skill that layers worktree isolation and evidence handoff onto the existing StratoFusion feature workflow.
4. Add a TypeScript/Playwright helper that captures two loopback pages or normalizes two existing PNG files into a gitignored task directory, writes sanitized metadata, and uploads nothing.
5. Prove the helper only against inert loopback fixture pages and deterministic tests.
6. Record observed pilot results, then prepare unpublished article and social drafts only if the evidence supports them.

## Pre-implementation go/no-go checklist

| Gate                                                                                         | Result               | Evidence                                                                                                                                  |
| -------------------------------------------------------------------------------------------- | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- |
| Isolated worktree and unique branch from latest `origin/dev`                                 | Go                   | Worktree and branch listed above; shared checkout remains untouched                                                                       |
| No overlap with relevant open PRs or another worktree                                        | Go                   | Open PR #331 is a `main` release candidate touching unrelated release/CASA/product files; no open pilot branch exists                     |
| Existing stack can implement the helper without a dependency                                 | Go                   | Node, TypeScript, PNPM, Playwright, and Chromium capture APIs already exist                                                               |
| Product/runtime behavior remains unchanged                                                   | Go                   | Scope is repository workflow files, local scripts/tests, docs, and drafts only                                                            |
| Secrets, provider state, auth, billing, rclone, database, and production remain out of scope | Go                   | Helper uses a fresh browser context or existing local PNGs, limits browser requests to read-only loopback traffic, and performs no upload |
| CI and human approval remain authoritative                                                   | Go                   | No workflow, branch protection, reviewer, merge, or deployment setting changes are proposed                                               |
| Pilot can be removed cleanly                                                                 | Go                   | New templates/skill/helper/docs/drafts are separable; package and index entries can be reverted                                           |
| Baseline can prove time or review-effort improvement                                         | No-go for that claim | No baseline review timings exist; the pilot will report completeness and evidence output, not time saved                                  |

**Decision before workflow edits:** proceed with the narrow pilot. Do not claim adoption success or time savings.

## Explicit exclusions

This pilot does not install or configure Greptile, Cursor Cloud Agents, Bezalel, Eve, or another agent platform; replace CodeRabbit or CI; add an automatic reviewer loop; approve or merge changes; change branch protection or repository settings; copy the external `code-structure` skill; use a public evidence host; copy environment files, auth state, provider credentials, secrets, or browser sessions; run rclone, provider, billing, OAuth, migration, deployment, production, or credential-backed E2E actions; change customer-facing behavior; or publish/schedule external content.

## Pilot results

The pilot produced the proposed repository slice without changing product behaviour, CI workflow definitions, provider integrations, auth, billing, rclone, databases, deployment, production, branch protection, or merge authority. After the ready-for-review handoff, two narrow transitive dependency overrides were added to remediate newly reported production advisories rather than weakening CASA.

- Two issue forms and one pull-request template now collect the goal, testable acceptance criteria, safety implications, verification, evidence, privacy review, rollback, and unresolved decisions.
- The skill is explicitly experimental and absent from the default skill-loading strategy.
- The helper rejects remote target URLs, embedded credentials, sensitive query-key names, remote browser subrequests, and non-read-only browser requests. It uses a fresh Chromium context, blocks service workers and downloads, and uploads nothing.
- The deterministic smoke run produced `before.png` (23,152 bytes), `after.png` (23,584 bytes), `manifest.json`, and `report.md` beneath the ignored `.tmp/pr-evidence/software-factory-pilot-smoke/` directory.
- The final manifest recorded only `/before` and `/after`, safe labels, a 1440x900 viewport, timestamp, commit `c3b856e37761e68bd404f48eeb35773a16c61b08`, and the smoke command. Manual image and metadata review found no credentials, user paths, account IDs, provider data, customer filenames, or customer data.
- The first focused test run found a real defect: a generic token-shaped CLI value was not redacted from command metadata. The sanitiser and regression coverage were updated. The final focused helper suite passed 9/9 tests.
- The article and social records remain drafts. Neither is loaded as a published article, approved for social publishing, scheduled, or posted.
- Follow-up remediation pins `browserslist` 4.28.8 and `postcss-selector-parser` 6.1.4 through the existing PNPM override mechanism. It adds no direct dependency and removes the two high and one low production advisories reported by CASA.

### Verification results

| Check                                                                     | Result                        | Notes                                                                                                                                                                                                                                                                                                                                                   |
| ------------------------------------------------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `pnpm env:guard`                                                          | Pass                          | Windows-native PowerShell environment; no WSL                                                                                                                                                                                                                                                                                                           |
| `pnpm evidence:pr -- --help`                                              | Pass                          | Usage and local-only boundary rendered                                                                                                                                                                                                                                                                                                                  |
| `pnpm test scripts/pr-evidence/core.test.ts`                              | Pass after one useful failure | Final: 9 tests passed; initial generic token-redaction miss fixed and covered                                                                                                                                                                                                                                                                           |
| `pnpm evidence:smoke`                                                     | Pass                          | Four-file local evidence package complete; images visually inspected                                                                                                                                                                                                                                                                                    |
| Focused ESLint and Prettier checks                                        | Pass                          | `scripts/pr-evidence/*.ts` and touched article test formatted; no lint warnings                                                                                                                                                                                                                                                                         |
| `pnpm social:validate`                                                    | Pass                          | 6 repository social records validated, including the new draft                                                                                                                                                                                                                                                                                          |
| `pnpm test scripts/social/content-store.test.ts src/lib/articles.test.ts` | Pass                          | 9 tests passed; new article draft explicitly remains outside published results                                                                                                                                                                                                                                                                          |
| `pnpm check`                                                              | Pass                          | Repository lint and TypeScript checks passed                                                                                                                                                                                                                                                                                                            |
| `pnpm audit --prod --json`                                                | Pass after remediation        | Zero production advisories reported after applying the two transitive overrides.                                                                                                                                                                                                                                                                        |
| `pnpm casa:security`                                                      | Pass after remediation        | Web and worker production audits, static scans, tracked-change secret scan, file-size check, deployment/supply-chain policy, and assessor handoff policy passed. Live database checks were skipped because `DATABASE_URL` was intentionally absent. Generated evidence changes were discarded.                                                          |
| GitGuardian Security Checks                                               | Pass after owner disposition  | GitGuardian classified the deliberately synthetic CLI-redaction fixture in commit `c3b856e3` as a generic CLI secret. The value was not a credential, was replaced in the current tree with a dynamically assembled fixture, and the sole maintainer dismissed incident `36837173` as a false positive. A fresh scan passed on closing head `0acd16c3`. |
| `pnpm build`                                                              | Pass after remediation        | The production compile and generate phases completed successfully with the patched dependency graph.                                                                                                                                                                                                                                                    |

No credential-backed, provider, OAuth, billing, rclone, database, deployment, or production check was run because those surfaces are outside the pilot and the necessary credentials were intentionally unavailable.

## Evaluation

| Measure                     | Baseline                                                | Pilot result                                                                                                                                                                                                                                                                          | Confidence                                       | Adopt?                              |
| --------------------------- | ------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | ----------------------------------- |
| Intake completeness         | 0/8 sampled PRs had explicit acceptance criteria        | Issue and PR templates now request testable acceptance criteria and exclusions; actual author completion is unmeasured                                                                                                                                                                | Medium for structure, none for future compliance | Continue                            |
| Reviewer comprehension      | Verification was strong, but handoff structure varied   | The sole maintainer rated the evidence `partly` helpful: structured acceptance, safety boundaries, and exact checks helped verification, while fixture screenshots and duplicated result maintenance added little value. This is self-review, not independent evidence.               | Low                                              | Continue with revisions             |
| Runtime evidence usefulness | No consistent linked evidence package in 8 PR bodies    | Smoke package is deterministic, visually legible, named consistently, and honest about missing evidence                                                                                                                                                                               | Medium for mechanics, low for real PR usefulness | Continue                            |
| First-pass verification     | 8/8 sampled PRs had passing CI                          | One initial helper-test failure exposed a redaction defect. A later CASA advisory failure was remediated with patched transitive versions. Closing CI run `33628767467` passed every repository job, including CASA, GitGuardian, and aggregate `build-and-test`, on head `0acd16c3`. | High for observed commands                       | Continue with the failures recorded |
| Added maintenance burden    | No template/helper maintenance                          | 889 lines across the helper, test and skill, plus a 43-line canonical testing section, concise templates and drafts; repeated-use burden is unknown                                                                                                                                   | High for line count, low for ongoing effort      | Continue narrowly                   |
| Privacy/security risk       | Evidence practices were task-specific                   | Loopback/read-only request policy, fresh context, metadata redaction, ignored output and manual pixel review are explicit; deny-list sanitisation and screenshot pixels remain residual risks                                                                                         | Medium                                           | Continue only with human inspection |
| Vendor/tooling cost         | Existing GitHub, CodeRabbit, PNPM, and Playwright stack | No direct dependency or paid service added; two transitive packages were pinned to patched versions through existing PNPM overrides                                                                                                                                                   | High                                             | Continue                            |

## Recommendation

**Continue the pilot** for three to five suitable, opt-in pull requests. Do not adopt it unchanged yet. Measure template completion, clarification comments, whether reviewers use the evidence, sensitive-data incidents, maintenance effort, and first-pass check outcomes. A timed comparison is required before making a time-savings claim.

Continuation measurements are recorded in `SOFTWARE_FACTORY_PILOT_SCORECARD.md`. PR #333 completed trial 1/5 with a `partly` helpful sole-maintainer self-review, zero confirmed privacy disclosures, unmeasured maintenance time, and all closing checks passing. That result cannot establish independent reviewer comprehension. Trials 2 and 3 should time pilot-specific maintenance, use deterministic output for nonvisual changes, and reserve screenshots for visual changes.

The decision is intentionally reversible:

| Outcome              | Exact repository action                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    |
| -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Adopt unchanged      | Retain `.github/ISSUE_TEMPLATE/bug_report.yml`, `.github/ISSUE_TEMPLATE/feature_or_workflow.yml`, `.github/pull_request_template.md`, `.gitignore`, `package.json`, `docs/ai-skills/INDEX.md`, `docs/ai-skills/16-evidence-backed-delivery-pilot.md`, `scripts/pr-evidence/capture.ts`, `scripts/pr-evidence/cli.ts`, `scripts/pr-evidence/core.ts`, `scripts/pr-evidence/git.ts`, `scripts/pr-evidence/smoke.ts`, `scripts/pr-evidence/core.test.ts`, `public/docs/TESTING.md`, `src/lib/articles.test.ts`, and this report. Retain the article/social files as drafts until separate editorial approval. |
| Adopt with revisions | Retain the same files while changes are reviewed. Update the templates, skill, helper, test and `public/docs/TESTING.md` together; remove only a superseded file after its replacement and migration notes are accepted. Keep this report as the original measurement record.                                                                                                                                                                                                                                                                                                                              |
| Continue the pilot   | Retain the same repository files and keep `content/articles/drafts/testing-a-small-ai-software-factory.md` and `content/social/drafts/testing-a-small-ai-software-factory.md` unpublished. Retain local `.tmp/pr-evidence/<task>/` packages only as long as their private review is active, then delete them locally.                                                                                                                                                                                                                                                                                      |
| Reject and remove    | Retain only this report as the decision record. Remove the three `.github` template files, `docs/ai-skills/16-evidence-backed-delivery-pilot.md`, all six implementation/test files under `scripts/pr-evidence/`, both draft content files, and local `.tmp/pr-evidence/`. Revert only the pilot additions in `.gitignore`, `package.json`, `docs/ai-skills/INDEX.md`, `public/docs/TESTING.md`, and `src/lib/articles.test.ts`.                                                                                                                                                                           |
