System 04 Creative evidence control plane
Know what worked.
Know why it might have.
A governed workspace for creative teams that separates observed performance from inference, catches invalid comparisons, detects fatigue, and drafts the next controlled test.
Live replay
Creative decision room
Awaiting run
Turn campaign results into bounded learning.
Choose an evidence state, optionally change the challenger, and inspect the decision receipt.
Graph running
Verifying source receipts
The review did not complete.
Check the request and try again.
Synthetic creative portfolio
Creative comparison
The challenger leads the registered KPI in a valid comparison.
Decision risks
0 findingsClaim contract
Observed · Inferred · RecommendationHuman checkpoint
Approve the draft experiment?
The decision is stored as a receipt. No external action runs.
Expected risk found across 13 golden scenarios.
Expected invalid, hold, iterate, or scale result.
Evidence binding, policy gates, and zero mutations.
Deterministic public analysis with no paid inference.
Historical baseline
A forecast that admits what it does not know.
The included ridge model uses 400 training rows and 100 holdout rows from the licensed UCI Facebook Metrics dataset. It is a transparent historical baseline, not a campaign promise, causal judge, or Florida-specific model.
- Holdout interaction MAE101.44
- Interval coverage0.80
- Dataset period2014
System architecture
Deterministic gates around agent collaboration.
Agents organize the investigation. Typed rules own measurement validity, evidence hashes, claim boundaries, permissions, and the final action gate.
Typed contracts first
Pydantic rejects impossible metric relationships before they enter the graph. OpenAPI keeps browser, Postman, and integration contracts aligned.
State is inspectable
Parallel analysis, checkpoints, branching, and the approval interrupt are explicit. A free-form group chat would be harder to test and govern.
Learning needs lineage
Tenant runs, evidence, model artifacts, outcomes, decisions, usage, and audit records remain queryable and deployable beyond one process.
Implementation
A deployable product, not a scripted screen.
Run it with one Docker command, replace replay fixtures with approved aggregate connectors, and move checkpoints to PostgreSQL without changing the graph contract.
Open the API contract ↗Evaluation and observability
Every decision leaves evidence.
Evaluation measures
Detection, decision accuracy, schemas, evidence binding, injection resistance, measurement gates, exposure bounds, model disclosure, and approval safety.
Metric families
Latency, graph duration, active runs, agent nodes, policy outcomes, findings, fatigue, attribution defects, model forecasts, errors, quotas, and audit events.
Claim classes
Observed, inferred, and recommendation. Every material statement keeps its evidence references and limitations.