Azdaja

A live product-path receipt

Three exact first-party harness assertions

Claude Fable synthesized the programs. Azdaja executed them locally against complete deterministic inputs. One synthetic live run per task returned 3/3 exact. This is not benchmark accuracy or independent replication.

3/3exact results
3provider calls
3local Monty executions
0recursive or semantic subcalls

This run tests a concrete Azdaja contract. The model receives a bounded root prompt and writes a generic inspection program rather than the answer. The complete source remains in the local evaluator. Azdaja runs the program, checks the runtime path, and returns one exact result or fails.

The three tasks

TaskExact resultRoot promptProvider time
Duplicate-sensitive build-log aggregationAnswer: 1312,739 B17.570 s
Repository blocker and source pathsrc/module_07777.rs|AZD-777712,460 B22.054 s
Color on the final matching catalog recordColor: cerulean12,770 B23.872 s

Provenance map

EvidenceMeasurement sourceReceipt or artifact SHA-256Scope
Live Fable suite2514d26966b92bd3153c519bd3e4e9d62152a590ab848b77a2c84729d0c38cbaa86b508One provider-generated observation per synthetic task
Provider-free acceptancefe91f98ebcc0632ef458a90b90eb0b95ed3343df85e66d9ed22633173e2ce3fdb1ec08bDeterministic scripted transport and retained raw log
Shared release binaryRecorded by both receipts3111c880bcf2cbf738282bf826bd5649e175fa5a0efc0b3ace5b077425f0921aBinary identity only, not a security attestation
Offline verifier bundlemanifest.json → source_commitEvery listed artifact is byte-counted and SHA-256-boundNarrow first-party reproducible evidence

What the receipt checks

  • Each scenario recorded exactly one configured Claude Fable CLI call.
  • Each model response is published and contains neither its expected answer nor its answer-specific constant.
  • Each program ran once in the local Monty evaluator with zero recursive or semantic subcalls.
  • Exact scanners found no 100-byte source span in any provider prompt.
  • No provider prompt contained a repository, input, or scratch host path.
  • The verifier binds the receipt to the source manifest and release binary hashes and rejects altered totals.

Verify the recorded evidence offline

The provider-generated live observation is already recorded. This command does not reproduce that model call. It verifies both retained receipts, the provider-free raw log, fixture specifications, deterministic invariants, verifier sources, and this public page against one SHA-256 manifest. It performs no provider call, downloads nothing, and fails closed on changed files, claims, answers, totals, paths, or source bindings.

./proof/reproduction/run.sh

Inspect the manifestRead the verifierReproduction guide

This is narrow first-party reproducible evidence, not an independent replication or complete third-party verification capsule. Exact deterministic checks are separated from timestamped live observations, and the provider-free v2 input-digest limitation is explicit.

Inspect or replay it

The machine-readable receipt was generated on 2 September 2026 and binds source commit 2514d26 plus the release binary hash. It includes all three model-authored programs, command templates, prompt and response hashes, runtime counters, exact outputs, environment versions, source hashes, and explicit claim limits.

Read the resultInspect the receiptRun the harness

python3 bench/live_fable_suite/verify.py \
  bench/results/live-fable-suite.json \
  --binary target/release/azdaja

Claim boundary

This is one synthetic live run per task on one model and one provider route. There is no baseline arm, no repeated trial, and no subscription token-usage record in text mode. It is product-path evidence, not a benchmark, leaderboard result, or superiority claim.