Noetfield Systems Inc. Evidence lab
← Proof register

Case 001 · Agentic coding reliability

When AI Changes the Controls That Govern AI

A first-party reliability case on deterministic control-plane integrity, false internal success, cost-aware recovery, and independent customer-outcome verification.

Claim-level evidence Published from: Noetfield Incident evidence: NOETFIELD-RUNWAY Not a universal tool scorecard

The failure

Agent-assisted changes affected deterministic gates, runtime configuration, delivery state, and success signals. Internal components reported success while the intended customer outcome remained unavailable.

Why it matters

A deterministic system is not trustworthy merely because it executes consistently. Its governing contract may have been changed incorrectly by a probabilistic agent.

What the case proves

  • Health and unit tests did not prove the customer effect
  • Paid regeneration could not repair deterministic metadata defects
  • Models must not change, judge, and promote their own control-plane changes
  • Acceptance must bind request → execution → durable result → independent fetch
  • Cost must be measured per accepted outcome, not per successful model call

Evidence boundary

This is a first-party case. It supports a control-plane and acceptance-boundary analysis—not a universal judgment about Claude, Cursor, or another vendor. Commit co-author metadata does not establish sole causation or universal tool quality.

Architecture lesson

The acceptance layer must detect green-internal / empty-customer gaps. The detailed technical record is the July 2026 app postmortem.

Current status

UNKNOWN Full current closure requires a fresh same-release production E2E across named, generic, complex, and repair scenarios. Historical receipts below are dated evidence, not a claim that production is closed today.

1 — The acceptance contract

Request: a signed-in customer gives the front person a normal sentence for a named business or service.

Accepted effect: the app returns plain customer-facing status and an unauthenticated public page that matches the request.

Not accepted as substitutes: a health endpoint, unit tests, a builder verdict, a verifier verdict, or a “done” message whose public link returns NOT_GENERATED.

2 — What the record shows

Bounded findingResultEvidence boundary
Named-business runHTTP 200 · “Meridian Tax Group” · 13,684 msVERIFIED RECEIPT one dated run
Generic trade-description runNo public page before 302,662 ms timeoutVERIFIED RECEIPT one dated failure
Legacy incident matrix3 component rows verified · 1 row openREPOSITORY-SUPPORTED not a 3/3 customer-path pass
Expanded ten-scenario router suite10 plans compiledBUILD EVIDENCE no public delivery
Expanded live run10 timeouts · 0 deliveredREPOSITORY-SUPPORTED candidate branch not on target
Current production closurePending a fresh same-release E2EUNKNOWN

3 — Why “green” was not “delivered”

Green signalWhat it actually establishedMissing customer proof
Health endpointA service answered a health request.The signed-in job route, publish step and public URL could still fail.
Unit and content-gate testsSelected code paths matched their fixtures.Fixtures did not represent every natural-language request or the final public effect.
Motor or builder successAn internal artifact passed internal checks.The app had not proved durable storage and an unauthenticated customer URL.
Local ten-scenario router PASSContracts and routes could be planned.The corresponding live run later delivered zero of ten on its target.
Chat “done” messageA status path believed work was complete.No independent public GET was bound to the announcement.

Root lesson: acceptance must bind the full effect: request → route → build → durable publish → independent public fetch → plain customer message.

4 — Change-set comparison, with attribution limits

PhaseMetadata attributionSupported statementNot supported
Earlier observed operationCursor co-author metadata appears in related changes.The owner reports that the earlier product produced pages.Exact timing, full prompt coverage, sole authorship, or current-release behavior.
Gate expansion, PR #322Claude co-author metadata appears in the change set.The PR spans 130 commits and 182 files.That one tool solely authored or caused every later defect.
31 July recovery sessionClaude Code session record plus owner testimony.Multiple component repairs landed; the owner reports repeated premature “fixed” conclusions.A controlled measurement of vendor quality or exact total harm.
Later routing and delivery candidatesMixed Cursor/agent metadata.The repository added outcome contracts, routing and delivery-state work.Customer closure until the same release passes a live public-effect suite.

Commit and PR metadata can support co-authorship or assistance. It cannot establish sole authorship, intent, causation, or universal tool quality.

5 — Claim ledger

C01 · Named-business bounded pass

VERIFIED RECEIPT On 1 August, one named-business run returned HTTP 200 with H1 “Meridian Tax Group” in 13,684 ms.

tested app release 2b818e415a2841d7e9eb390ba9f464722f92e8c7 · receipt payload sha256 3daa06f7e116f42346573ed357e2c0f2951bba8732ca4934a1cb6bb5e3c76aba

Open the public receipt. It proves this run, not all request classes.

C02 · Generic trade bounded failure

VERIFIED RECEIPT A generic trade-description run on the same tested app release failed to publish before 302,662 ms.

receipt payload sha256 900f05759b3397f77efbd4a792c1b4829aacedee8ef8d49ac97076c4e7199e28

Open the public receipt. No later production rerun is claimed here.

C03 · Historical matrix, not current closure

REPOSITORY-SUPPORTED The 07:28 snapshot recorded three component rows verified and one generic-trade row open. Only the named-entity row traversed the complete customer path.

Open the historical matrix snapshot.

C06 · Expanded live run

REPOSITORY-SUPPORTED A subsequent expanded run recorded ten timeouts and zero delivered scenarios. Its candidate branch was not deployed to the target host, so the result proves missing end-to-end delivery on that target—not candidate quality.

Open the immutable repository receipt.

C08 · Current closure

UNKNOWN Current full customer-outcome closure requires a fresh same-release production run across named, generic, complex and repair scenarios. This page does not convert a branch candidate or historical matrix into current success.

6 — Owner-observed impact

OWNER-OBSERVED The account owner reports that the product had produced pages before the gate changes, then suffered a multi-day customer-impact period involving repeated failed builds, internal jargon shown to the client, paid attempts without deliverables, and repeated premature “fixed” reports.

This testimony matters, but this bundle does not independently calculate total hours, model spend, lost opportunity, or causal share by vendor. Those quantities remain unsealed.

7 — Corrections preserved, not erased

Prior claimDispositionReplacement
25 July was the last demonstrated working run.SUPERSEDEDLater bounded receipts exist; no single “last known good” is asserted without a complete same-release evidence chain.
The customer path passed 3/3 and represented current closure.SUPERSEDEDThe legacy matrix mixed component checks with one full path; prompt-level public outcomes were one pass and one failure.
2b818e… was a sealed verifier release.CORRECTEDIt is the tested app/static release bound to the dated receipts; runtime implementation identity was not established.
Publication and incident evidence came from one repository.CORRECTEDThe public page is in Noetfield; implementation evidence is in NOETFIELD-RUNWAY.

8 — What would support a real tool comparison

A controlled comparison has not been run. It needs the same frozen repository SHA, same written task, equal permissions and budget, isolated worktrees, repeated runs, an independent browser-to-public-effect scorer, and hash-bound raw outputs.

NOT MEASURED Therefore this case supports a change-set and acceptance-boundary analysis—not “Cursor is always better” or “Claude is always worse.”

9 — Method and privacy boundaries

  • Claims are separated into verified receipt, repository-supported, owner-observed and unknown.
  • Every numerical claim is tied to a dated run or immutable repository record.
  • Raw project and session identifiers are not repeated here.
  • Stable hashes are pseudonymous and may remain linkable; they are not anonymous.
  • Private repository links may require authorization. Public receipts are linked separately.
  • Dollar cost and total token spend remain unknown in this bundle.

10 — Sources and machine-readable record

11 — Right of reply and correction policy

Anthropic, Cursor, OpenAI, contributors, and affected operators may request a correction with verifiable citations through the evidence-correction contact. Substantiated corrections are appended with the prior claim preserved as superseded.

Strategic category: provider-neutral artifact/effect acceptance. A system is not done when an internal agent says done. It is done when the intended customer effect is independently observed.