Noetfield Systems Inc. Evidence lab
← Evidence register

Evidence Lab · Single source of truth

Agent Execution Assurance

Autonomous AI execution, bounded by explicit authority and judged by a verifier that is separate from the worker that produced the result. This page is the one place where the demonstrations, the receipts, the measured cost, and the honest gaps are kept together — so that what you are told matches what you can open.

LIVE ALPHA · FOUNDER-OPERATED First-party evidence No external customers or revenue yet Not a certification

How to read this page

Noetfield sells acceptance, so this page holds itself to the standard it sells. Every factual line below carries one of five grades, and nothing is stated without one.

GradeWhat it means
VERIFIED RECEIPTBound to a dated run with a published, hash-identified artifact you can fetch yourself.
REPOSITORY-SUPPORTEDRecorded in an immutable repository or CI record, but not a full customer-path pass.
OWNER-OBSERVEDReported by the operator. Real, but not independently measured here.
NOT MEASUREDNot yet instrumented. Stated as unknown rather than estimated.
CONCEPTA depiction of the intended product. Not a recording, not evidence, and never counted as a demonstration.

A claim with no grade is not a claim. If you find one on this page, it is a defect — report it and it will be corrected with the prior version preserved.

Recorded demonstrations

Each case is one bounded run, recorded end to end. Cases are added here as they are filmed; a case appears only once its receipt exists.

Proof Case 001 — The Refusal Cycle TAKE v2

The system builds a website and an automation from one plain request, and refuses to invent a fact it was never given.

VERIFIED RECEIPT One bounded run, recorded end to end by the machine on production, no cuts: a single plain request produced a published website in eight seconds and a delivered automation beside it. The refusal beat — a stop with a stated reason instead of a plausible invention — is evidenced separately on the receipt wall below. The recording ships with its run receipt.

Proof Case 002 — The Front Door TAKE v1

One sentence asks for a website and the automation behind it. The plan answers both halves — before an account exists.

VERIFIED RECEIPT Seventeen seconds, one continuous take on production, no cuts: the request typed character by character, then the plan naming the page sections and the automation it would add. The recording is hash-identified (sha256 0895eb00…39cb08c) and published with its poster and captions. This one needs no receipt to be believed — you can reproduce it yourself in under a minute at app.noetfield.com, with no account. What it does not show: the build, the published page, or the automation running — those are Case 001's ground, and they begin after sign-up.

Proof Case 004 — A Delivered Page TAKE v1

The finished result: a page built from one plain request and published to a public address, read top to bottom.

VERIFIED RECEIPT One plain request in, a published public address out, in 17 seconds — delivery confirmed by fetching the live page twice and comparing body hashes, which matched. Filmed at that public address, unedited. The recording is hash-identified (sha256 ad47c2cc…c8773767). Boundary, stated plainly: this films the delivered result, not the build, and it is a different run from Case 002 — it does not continue it. What it does not show: the work between the request and the page.

Proof Case 005 — Determinism & Binding TAKE v1

A hundred runs collapse to one verdict hash, six controls move it, and a still-valid deploy token refuses an artifact that changed by one byte.

VERIFIED RECEIPT Thirty-one seconds, one continuous take on production, no cuts: two deployments blocked at EXIT 20 and EXIT 40, a clean artifact clearing all eight gates, then one hundred executions under perturbed paths, timestamps, file ordering, locales and timezones — returning a single verdict hash. Six negative controls each move that hash, which is what separates a deterministic function from one that returns a constant. The recording is hash-identified (sha256 daaaf41f…e06115b7). This one needs no receipt to be believed — run it yourself at proof.noetfield.com, with no account, and compare your verdict hash to the one on screen. What it does not show: a customer outcome. Nothing is built for anyone here; the artifacts under test are fixtures carried by the lab. Cases 001, 002 and 004 are that ground.

Proof Case 006 — One Sentence, Two Deliverables TAKE v1

A client asks for a landing page and the automation behind it in one sentence; the front person hires a team, publishes the page, provisions the workflow, and hands over both.

VERIFIED RECEIPT One request, typed as a client writes it, asking for two different products at once: a landing page and an automation for its booking form. The front person hires its team in the open — front person, builder, delivery orchestrator, each shown as it completes. The page is live and serving at a public address in five seconds; the automation half then runs in view (integrator, workflow provisioner, quality check) and hands over a downloadable workflow file at two minutes. Both delivered from one sentence, inside a hard three-cent ceiling per job that refuses a call before it can exceed it. Filmed as a production-faithful cockpit animatic replay. The recording is hash-identified (sha256 9e70ae78…06e102d). The recording ships with its run receipt. Boundary, stated plainly: this is one customer-outcome run, not a success rate; it does not show the model’s reasoning, and it does not show the workflow executing inside the client’s own n8n — the file is delivered for them to import.

Proof Case 007 — Summit Ridge Combined Live TAKE v2

Summit Ridge Physical Therapy — full customer journey from front door through signup, workspace chat, and cockpit to delivered page plus automation.

VERIFIED RECEIPT The full customer path on production: front-door wizard (Business ownerMore customers), account creation, workspace at /app/new/ with the combined Summit Ridge prompt typed and sent to Front Person, cockpit with team roster and live ticker while work runs in flight, then edited cuts to delivery. Public page: live site (HTTP 200, 22,030 bytes). Cockpit: open project. Verified workflow in Files. ElevenLabs narration (~82s film). Hash: sha256 ee38c698…99853. Supersedes v1 output-only recap. Receipt: /proof/lab/case-007-run-receipt.json.

Proof Case 008 — Summit Ridge Full Customer Path TAKE v1

Summit Ridge Physical Therapy — the complete customer journey from front door through signup, workspace chat, and cockpit to delivered page plus automation.

VERIFIED RECEIPT The full customer path on production: front-door wizard, account creation, workspace at /app/new/ with the combined Summit Ridge prompt typed and sent, then inside the cockpit — Front Person chat rail, team roster, live ticker, tasks, documents, and website panels. Edited time cuts to delivery: public page at app.noetfield.com/v1/site/i-need-a-professional-landing-page-for-summit-ridge-physical-3 (HTTP 200, 22,030 bytes) and verified workflow file in Files. Cockpit: open project. ElevenLabs business narration (~83s film). Hash: sha256 c63d05c6…29f9d9. Receipt: /proof/lab/case-008-run-receipt.json.

Proof Case 009 — Northwind Dental Founder Path TAKE v1

Northwind Dental Studio — startup founder path from front door through workspace to delivered page plus appointment automation.

VERIFIED RECEIPT A new case, new vertical, new wizard UX on production: Startup founderA page to test demand → combined ask for Northwind Dental Studio (Calgary) — landing page with services, team, and appointment request form, plus automation (email + Google Sheet on each booking). Full customer path filmed: front door, signup, workspace chat, cockpit with team in flight, then delivered page and verified workflow in Files. Page sealed at 4m 31s; workflow at 5m 46s. Public page: live site (HTTP 200, 16,109 bytes, H1 “Northwind Dental Studio”). Cockpit: open project. Pro-grade edited session with ElevenLabs narration (~84s film). Hash: sha256 1ddd8b8a…53b71b. Receipt: /proof/lab/case-009-run-receipt.json. Boundary: one live harness run; workflow delivered for import, not shown in client n8n.

Product vision

These films show the product as it is intended to work at full scale. They are made, not recorded — nothing in them is a capture of a live run, none of it is evidence, and none of it is claimed as available today. What exists now is above and on the receipt wall; what is intended is here. The roadmap states which is which.

Noetfield in full production

The intended product: work arriving planned, executed, verified and evidenced at scale, with the operator supervising rather than assembling.

CONCEPT A depiction of the intended product, not a recording and not evidence. Read it as direction, and hold the company to the graded claims elsewhere on this page.

The receipt wall

These are published, dated runs. Each links to an artifact hosted outside this page so the claim can be checked without trusting the page.

RunRecorded resultGrade and boundary
Northwind Dental founder path, 7 August HTTP 200 · “Northwind Dental Studio” · 16,109 bytes · verified workflow · founder wizard · 271 s page · 346 s total VERIFIED RECEIPT full customer-path film — new vertical, new UX
Summit Ridge full customer path, 7 August HTTP 200 · Summit Ridge page · 22,030 bytes · verified workflow · front door → signup → workspace → cockpit VERIFIED RECEIPT one live production full-path film run
Summit Ridge combined live, 7 August HTTP 200 · “Summit Ridge Physical Therapy” · 21,961 bytes · verified workflow · 174 s page · 204 s total VERIFIED RECEIPT one live production combined-ask run; output recap (see Case 008 for full path)
Named-business request, 1 August HTTP 200 · public page headed “Meridian Tax Group” · 13,684 ms VERIFIED RECEIPT one dated run, not all request classes
Generic trade-description request No public page before a 302,662 ms timeout VERIFIED RECEIPT one dated failure, preserved deliberately
Expanded ten-scenario live run 10 timeouts · 0 delivered on that target REPOSITORY-SUPPORTED candidate branch was not deployed to the target host
Governed replacement, run CS2-GISF-2026-07-14T005554Z Verification failed · bounded repair applied · independent re-check passed · human promotion VERIFIED RECEIPT first-party run, not external deployment
Current production closure Fresh same-release runs exist and passed; a success rate does not NOT MEASURED single runs, counted — never divided

NOT MEASURED The last row used to read “pending a fresh same-release run”, which stopped being true once the 7 August rows above it landed. It is corrected rather than removed, because the honest gap moved rather than closing: fresh runs on the current release exist and passed, and that is still a count of single runs, not a rate. Every row here is one run. None of them divides.

A rate needs a fixed prompt set, a stated number of attempts, every outcome counted including the failures, and the raw results published. That has not been done, so no percentage appears anywhere on this page. When it is, it will arrive as its own row with its method stated — not as an average quietly computed from the rows above.

Named-business receipt payload sha256 3daa06f7e116f42346573ed357e2c0f2951bba8732ca4934a1cb6bb5e3c76aba · generic-trade receipt payload sha256 900f05759b3397f77efbd4a792c1b4829aacedee8ef8d49ac97076c4e7199e28 · tested app release 2b818e415a2841d7e9eb390ba9f464722f92e8c7

Cost economics

Cost is measured per accepted outcome, not per successful model call. A run that produced no customer effect still consumed spend, and counting only the successful calls is the most common way agent economics are misreported.

Per-job spend ceiling $0.03 Enforced in the customer model policy and asserted by a test, so it cannot drift quietly.
Per-project ceiling $0.25 The bound a whole project is held under, including its repair rounds.
Cost per accepted outcome Measured, not published Recorded per terminal job since 6 August. No public receipt carries it yet, so no figure is claimed here.
Total token spend Not measured Still unknown in the published evidence bundles.

REPOSITORY-SUPPORTED The ceilings are constants in the customer model policy with a test pinning them, so they are checkable without trusting this page: the policy and the test that asserts them.

NOT PUBLISHED Spend per job is now written down at the end of every terminal job. It is deliberately not quoted here: until a public receipt carries the figure, a number on this page would rest on nothing a reader could open — which is the failure this page exists to avoid. What changed on 6 August is worth stating plainly, because it was worse than a gap: the ledger had always had a column for cost and nothing filled it, so 45 model-backed jobs read exactly $0.000000. That is not a low price. It is an absent measurement, and it is the kind of zero that flatters a company until someone asks.

What actually stands between a model and a customer

The controls below are the ones the acceptance argument rests on. They are described as design commitments; the receipts above are what show whether they held on a given run.

ControlWhat it enforces
Explicit authority before actionA run executes inside a stated authority envelope with recorded alternatives and a decision, not on an implied instruction.
Exact-candidate bindingThe thing being judged is pinned to exact commit digests, so a later tree cannot be substituted for the one that passed.
Separate verificationThe verifier is distinct from the worker that produced the change. A model does not judge and promote its own work.
Bounded repair with safe stopsRepair is scoped and terminates. A failure that cannot be repaired inside the bound stops and is recorded as a failure.
Human promotionA passing verification produces a candidate, not a release. Promotion is a human act.
Receipt for every accepted outcomeAcceptance binds request → execution → durable result → independent fetch, and preserves the record.

REPOSITORY-SUPPORTED These controls are implemented and exercised first-party. Case 001 in the register records a period in which internal signals reported success while the customer effect was absent — that case is kept public precisely because it is the failure mode these controls exist to catch.

Roadmap, with status

Stated so that a reader can tell today's capability from the intended one without asking.

Phase 1 — Prompt to production

LIVE ALPHA

One plain request produces a planned, executed, verified and deployable result: a published site and an accompanying automation definition, with the evidence preserved. Founder-operated, in the open alpha at app.noetfield.com. Delivery is demonstrated on named requests and has recorded failures on others; both are on the receipt wall above.

Phase 2 — Hosted execution

PLANNED

The automation runs where it is provisioned, rather than being handed over as a file to import. Provisioned endpoints, managed credentials, and the same acceptance and receipt layer applied to a running workflow instead of a generated one.

Phase 3 — TrustField

PRODUCT VERTICAL · SYNTHETIC DEMONSTRATIONS

Governed execution applied to regulated case operations: FINTRAC workflows for Canadian MSBs, fintechs, PSPs and digital-asset operators, with the human FILE / NO-FILE / ESCALATE decision preserved rather than automated away. TrustField is a product vertical of Noetfield Systems Inc., not a separate venture. Public synthetic demonstrations at trustfield.ca and the product page.

What this page does not claim

  • No external customers, revenue, or third-party production deployments.
  • No certification. A receipt is a record, not an accreditation, and this page is not SOC 2 or any equivalent.
  • No claim that current end-to-end delivery is closed. That requires a fresh same-release run and is marked NOT MEASURED until it exists.
  • No universal comparison of any AI coding tool or model vendor. No controlled comparison has been run.
  • No dollar cost or token spend figures, because none are yet instrumented.

Machine-readable record and sources

Correction policy

Any reader — investor, customer, vendor or contributor — may request a correction with verifiable citations through the evidence-correction contact. Substantiated corrections are appended, and the prior claim is preserved as superseded rather than quietly removed.

A system is not done when an agent says done. It is done when the intended effect is independently observed, and the record of both the pass and the failure survives.