Evidence Scout
Separates supported facts, inference, and open questions.
INPUT · source packetIndependent release evidence
Run reproducible black-box checks for contracts, evidence, approval gates, live boundaries, idempotency, and latency.
01 / LIVE EVALUATION
Ready. The lab calls only the three allowlisted public deployments.
02 / RESULT
03 / CREWAI TEACHING TRIAL
This teaching view mirrors the four-agent crew deployed in CrewAI AMP. A deterministic no-key model proves the live handoffs and named human-review boundary; it is not a model-quality benchmark.
Ready to show the handoff.
Separates supported facts, inference, and open questions.
INPUT · source packetTurns evidence into roles, handoffs, gates, and receipts.
CONTEXT · evidence ledgerDrafts the personal post and a distinct company adaptation.
CONTEXT · ledger + mapChecks every claim and stops at human review.
CONTEXT · complete draftExecution timeline
Sequential processTruth boundary: this page replays the architecture in the browser and never claims that animation is an AI run. The deployed AMP automation is the source for live execution status and traces. Its deterministic teaching model proves orchestration without an external model key; it does not prove generative quality. The hosted run ends with a review packet; the locked gate is enforced by having no publishing tool or connected social account.
04 / SCENARIO BRIEF
Agent demos often validate only the happy path. This lab calls three allowlisted public systems, verifies seven contract and governance controls, repeats requests for idempotency, probes forbidden live mode, measures latency, and exposes evidence for every score.
The original interface only showed a passing baseline. Missing evidence, approval bypass, and latency breach can now be injected explicitly and each lowers the score.
Target names resolve through a code allowlist. This prevents the public service from becoming a server-side request-forgery tool.
Privacy, provider quality, semantic task success, and real-data mutations require authenticated project-specific evaluation.
05 / ARCHITECTURE
06 / DECISIONS
The public URL, headers, latency, and boundary behavior must work outside the repository.
An approval bypass or missing evidence has one repeatable result. A model judge is reserved for calibrated qualitative output.
The public evaluator has no reason to fetch arbitrary hosts.
Idempotency is observable only when the same key is used more than once.
The bounded request surface is small enough for the platform runtime.
Current checks finish synchronously. Scheduled suites and long model graders are the trigger for a queue.
07 / EVALUATION
08 / OBSERVABILITY
09 / NEXT GATE