C01 · Named-business bounded pass
VERIFIED RECEIPT On 1 August, one named-business run returned HTTP 200 with H1 “Meridian Tax Group” in 13,684 ms.
Open the public receipt. It proves this run, not all request classes.
Case 001 · Agentic coding reliability
A first-party reliability case on deterministic control-plane integrity, false internal success, cost-aware recovery, and independent customer-outcome verification.
Agent-assisted changes affected deterministic gates, runtime configuration, delivery state, and success signals. Internal components reported success while the intended customer outcome remained unavailable.
A deterministic system is not trustworthy merely because it executes consistently. Its governing contract may have been changed incorrectly by a probabilistic agent.
This is a first-party case. It supports a control-plane and acceptance-boundary analysis—not a universal judgment about Claude, Cursor, or another vendor. Commit co-author metadata does not establish sole causation or universal tool quality.
The acceptance layer must detect green-internal / empty-customer gaps. The detailed technical record is the July 2026 app postmortem.
UNKNOWN Full current closure requires a fresh same-release production E2E across named, generic, complex, and repair scenarios. Historical receipts below are dated evidence, not a claim that production is closed today.
Request: a signed-in customer gives the front person a normal sentence for a named business or service.
Accepted effect: the app returns plain customer-facing status and an unauthenticated public page that matches the request.
Not accepted as substitutes: a health endpoint, unit tests, a builder verdict, a verifier verdict, or a “done” message whose public link returns NOT_GENERATED.
| Bounded finding | Result | Evidence boundary |
|---|---|---|
| Named-business run | HTTP 200 · “Meridian Tax Group” · 13,684 ms | VERIFIED RECEIPT one dated run |
| Generic trade-description run | No public page before 302,662 ms timeout | VERIFIED RECEIPT one dated failure |
| Legacy incident matrix | 3 component rows verified · 1 row open | REPOSITORY-SUPPORTED not a 3/3 customer-path pass |
| Expanded ten-scenario router suite | 10 plans compiled | BUILD EVIDENCE no public delivery |
| Expanded live run | 10 timeouts · 0 delivered | REPOSITORY-SUPPORTED candidate branch not on target |
| Current production closure | Pending a fresh same-release E2E | UNKNOWN |
| Green signal | What it actually established | Missing customer proof |
|---|---|---|
| Health endpoint | A service answered a health request. | The signed-in job route, publish step and public URL could still fail. |
| Unit and content-gate tests | Selected code paths matched their fixtures. | Fixtures did not represent every natural-language request or the final public effect. |
| Motor or builder success | An internal artifact passed internal checks. | The app had not proved durable storage and an unauthenticated customer URL. |
| Local ten-scenario router PASS | Contracts and routes could be planned. | The corresponding live run later delivered zero of ten on its target. |
| Chat “done” message | A status path believed work was complete. | No independent public GET was bound to the announcement. |
Root lesson: acceptance must bind the full effect: request → route → build → durable publish → independent public fetch → plain customer message.
| Phase | Metadata attribution | Supported statement | Not supported |
|---|---|---|---|
| Earlier observed operation | Cursor co-author metadata appears in related changes. | The owner reports that the earlier product produced pages. | Exact timing, full prompt coverage, sole authorship, or current-release behavior. |
| Gate expansion, PR #322 | Claude co-author metadata appears in the change set. | The PR spans 130 commits and 182 files. | That one tool solely authored or caused every later defect. |
| 31 July recovery session | Claude Code session record plus owner testimony. | Multiple component repairs landed; the owner reports repeated premature “fixed” conclusions. | A controlled measurement of vendor quality or exact total harm. |
| Later routing and delivery candidates | Mixed Cursor/agent metadata. | The repository added outcome contracts, routing and delivery-state work. | Customer closure until the same release passes a live public-effect suite. |
Commit and PR metadata can support co-authorship or assistance. It cannot establish sole authorship, intent, causation, or universal tool quality.
VERIFIED RECEIPT On 1 August, one named-business run returned HTTP 200 with H1 “Meridian Tax Group” in 13,684 ms.
Open the public receipt. It proves this run, not all request classes.
VERIFIED RECEIPT A generic trade-description run on the same tested app release failed to publish before 302,662 ms.
Open the public receipt. No later production rerun is claimed here.
REPOSITORY-SUPPORTED The 07:28 snapshot recorded three component rows verified and one generic-trade row open. Only the named-entity row traversed the complete customer path.
REPOSITORY-SUPPORTED A subsequent expanded run recorded ten timeouts and zero delivered scenarios. Its candidate branch was not deployed to the target host, so the result proves missing end-to-end delivery on that target—not candidate quality.
UNKNOWN Current full customer-outcome closure requires a fresh same-release production run across named, generic, complex and repair scenarios. This page does not convert a branch candidate or historical matrix into current success.
OWNER-OBSERVED The account owner reports that the product had produced pages before the gate changes, then suffered a multi-day customer-impact period involving repeated failed builds, internal jargon shown to the client, paid attempts without deliverables, and repeated premature “fixed” reports.
This testimony matters, but this bundle does not independently calculate total hours, model spend, lost opportunity, or causal share by vendor. Those quantities remain unsealed.
| Prior claim | Disposition | Replacement |
|---|---|---|
| 25 July was the last demonstrated working run. | SUPERSEDED | Later bounded receipts exist; no single “last known good” is asserted without a complete same-release evidence chain. |
| The customer path passed 3/3 and represented current closure. | SUPERSEDED | The legacy matrix mixed component checks with one full path; prompt-level public outcomes were one pass and one failure. |
| 2b818e… was a sealed verifier release. | CORRECTED | It is the tested app/static release bound to the dated receipts; runtime implementation identity was not established. |
| Publication and incident evidence came from one repository. | CORRECTED | The public page is in Noetfield; implementation evidence is in NOETFIELD-RUNWAY. |
A controlled comparison has not been run. It needs the same frozen repository SHA, same written task, equal permissions and budget, isolated worktrees, repeated runs, an independent browser-to-public-effect scorer, and hash-bound raw outputs.
NOT MEASURED Therefore this case supports a change-set and acceptance-boundary analysis—not “Cursor is always better” or “Claude is always worse.”
Anthropic, Cursor, OpenAI, contributors, and affected operators may request a correction with verifiable citations through the evidence-correction contact. Substantiated corrections are appended with the prior claim preserved as superseded.
Strategic category: provider-neutral artifact/effect acceptance. A system is not done when an internal agent says done. It is done when the intended customer effect is independently observed.