{
  "schemaVersion": "1.0",
  "title": "Live product end-to-end qualification",
  "description": "Use this procedure to prove that one identified AIP build can interact safely with one identified deployment of Cal.diy, Hermes Agent, Chatwoot, Dify, CrewAI, or Twenty. A valid campaign exercises a real product boundary in an isolated envi",
  "canonical": "https://getaip.org/docs/testing/live-product-e2e",
  "route": "/docs/testing/live-product-e2e",
  "source": "docs/testing/live-product-e2e.md",
  "protocol": "Agent Interoperability Protocol",
  "protocolVersion": "1.0",
  "section": "Qualification and Evidence",
  "documentType": "Qualification evidence",
  "language": "en",
  "revision": {
    "lastReviewedRevision": "d7cce13d1d555644d04a4d73c66c95b113737635",
    "documentationSourceRevision": "9192fef3695ad294994f2712f6d156241e5e92fb",
    "basis": "frontmatter"
  },
  "downloads": {
    "md": "/docs/download/testing/live-product-e2e.md",
    "txt": "/docs/download/testing/live-product-e2e.txt",
    "json": "/docs/download/testing/live-product-e2e.json",
    "pdf": "/docs/download/testing/live-product-e2e.pdf"
  },
  "content": {
    "format": "text/markdown",
    "markdown": "---\ntitle: Live product end-to-end qualification\ndescription: Design and evaluate an isolated live campaign for any maintained connector without overstating the result\nkind: procedure\naudience: operator\nappliesTo: \"1.x\"\nwritingStandard: \"aip-docs/1.0\"\nlastReviewedRevision: \"d7cce13d1d555644d04a4d73c66c95b113737635\"\n---\n\n# Live product end-to-end qualification\n\nUse this procedure to prove that one identified AIP build can interact safely\nwith one identified deployment of Cal.diy, Hermes Agent, Chatwoot, Dify,\nCrewAI, or Twenty. A valid campaign exercises a real product boundary in an\nisolated environment, retains redacted evidence, and removes every disposable\neffect before it reports `PASS`.\n\nThis page describes how to design and review a campaign for GetAIP `2.0.0`\nsource revision `d7cce13d1d555644d04a4d73c66c95b113737635`. The protected\ncontrolled product-fleet gate is not an isolated-live external-provider\ncampaign.\n\n## Know what the source can run\n\nThe pinned source does not contain a general live suite for all six products.\nUse only the entry points in this table, or record a separately reviewed\ncampaign driver as an external artifact.\n\n| Product | Source-owned product-facing entry point | Evidence boundary |\n| --- | --- | --- |\n| Cal.diy | Complete isolated Docker qualification and a narrower ignored reservation test | The complete campaign covers governed booking, replay, webhook, restart, and persistence; the ignored test covers reservation, replay, and release only |\n| Hermes Agent | Deployed-stack operator and delegation binary | Exercises MCP and native approval paths against identified Hermes endpoints and a controlled downstream AIP capability |\n| Chatwoot | No dedicated live-product driver at the pinned revision | Deterministic connector tests do not prove a live Chatwoot deployment |\n| Dify | No dedicated live-product driver at the pinned revision | Deterministic connector tests do not prove a live Dify deployment |\n| CrewAI | Ignored exact-upstream sidecar test | Proves the Rust connector and real CrewAI library boundary; the supplied qualification registry uses a deterministic network-free language model |\n| Twenty | No dedicated live-product driver at the pinned revision | Deterministic connector tests do not prove a live Twenty deployment |\n\nThe optional five-product fleet verifier uses controlled upstream behavior and\ndoes not include Twenty. It is useful AIP integration evidence, but it is not\nlive-provider evidence for any product.\n\n## Define the claim before the run\n\nWrite one sentence that names the product, provider deployment, connector\nartifact, AIP artifact, campaign revision, and operations in scope. Do not use\nthe broader phrase “the connector is qualified” when the run covers only a\nsubset of operations.\n\nFor example:\n\n> The identified AIP and connector artifacts completed the documented Cal.diy\n> booking campaign against the identified isolated provider deployment, with\n> cleanup confirmed, at the recorded time.\n\nThe campaign must not expand its claim to another provider version, another\nconnector artifact, untested operations, production safety, or portable\nperformance.\n\n## Complete the safety and identity preflight\n\nDo not begin a mutating step until every item below has an owner and a recorded\nvalue.\n\n| Check | Required record |\n| --- | --- |\n| AIP identity | Full source revision, dirty-tree state, build or image digest, configuration digest, and manifest digest |\n| Connector identity | Connector artifact digest, exact upstream pin, capability-set digest, and declared operation surface |\n| Provider identity | Product version or source revision, deployment or image digest, enabled feature flags, and migration state |\n| Isolation | Dedicated non-production tenant, account, workspace, application, crew, or calendar with no production data |\n| Mutation budget | Exact objects the run may create, update, cancel, or delete, plus a cleanup owner and deadline |\n| Principals | Separate requester and approver identities where approval is tested; least-privilege product identity |\n| Credentials | Protected file or secret-store references; record only names, scopes, expiry, and fingerprints |\n| Network | Endpoint names, trust roots, proxy path, redirect policy, and relevant allowlists without secret values |\n| Time | UTC clock source and measured skew for signatures, deadlines, leases, and replay windows |\n| Evidence | New owner-only directory, redaction rules, retention period, and reviewer |\n\nConfirm that every planned mutation is reversible. If a product operation has\nno reliable cleanup path, use a disposable deployment that can be destroyed as\na unit. Never discover this constraint after writing shared customer data.\n\n## Run one complete campaign\n\nUse the same sequence for every product, adapting only the product operations.\n\n1. Record all identities and prove that the isolated target is ready.\n2. Run the connector's deterministic tests and the applicable conformance\n   checks. Retain their results separately from live evidence.\n3. Perform authenticated, non-mutating discovery and baseline reads.\n4. Submit one governed mutation through AIP, including approval when required.\n5. Repeat the same action, idempotency key, or transaction identity and verify\n   that it does not create an additional external effect.\n6. Exercise the supported cancellation, timeout, or recovery boundary without\n   inventing behavior the connector does not declare.\n7. When webhooks are supported, verify the raw-body signature, freshness,\n   duplicate suppression, and rejection of invalid authentication.\n8. Restart the relevant AIP or connector process and verify durable status,\n   provider references, and recovery behavior.\n9. Remove or revert every created provider object and verify its absence or\n   terminal cleanup state through an independent read.\n10. Redact, hash, inventory, and review the retained evidence before assigning\n    a result.\n\nRecord unsupported operations as `NOT_RUN`; do not replace them with similar\noperations or silently omit them.\n\n## Apply the product-specific minimum\n\nA product campaign is complete only when it covers the applicable minimum in\nthis table. The connector documentation defines the exact inputs, exclusions,\nand product-specific cleanup steps.\n\n| Product | Minimum live assertions |\n| --- | --- |\n| Cal.diy | Ready health; availability read; mutation-free booking plan; independent approval; one booking; identical idempotent replay; signed webhook acceptance and replay rejection; restart recovery; booking or reservation cleanup |\n| Hermes Agent | Endpoint health and discovery; streaming result; independent approval; governed AIP tool lifecycle; scoped delegation; out-of-scope rejection; cancellation or terminal status; reconnect or restart observation |\n| Chatwoot | Account-scoped read; reversible conversation, private-note, status, or handoff mutation; idempotent replay; authenticated webhook; event-loop prevention; object or state cleanup |\n| Dify | Application parameter discovery; applicable synchronous and streaming execution; cancellation; human-input or pause state when supported; task recovery; applicable knowledge mode; cleanup of disposable inputs and tasks |\n| CrewAI | Sidecar authentication; exact registered crew discovery; real CrewAI crew execution; stream and terminal status where enabled; cancellation; replay behavior; restart fencing; one authoritative writer; separate identification of any external model provider |\n| Twenty | OpenAPI-derived core-record read; reversible record lifecycle; metadata read; webhook HMAC and nonce replay rejection; credential redaction; proof that the connector did not use the excluded GraphQL or legacy API-key paths; record cleanup |\n\nIf the product cannot supply one applicable assertion, report the gap and stop\nat `BLOCKED` or `FAIL`. A deterministic substitute cannot fill a live-product\ngap.\n\n## Use only existing source-owned entry points\n\nRun commands from the root of the pinned AIP checkout. The examples below show\nthe available interfaces; operators must supply their own protected values and\nmust satisfy the preflight before execution.\n\n### Complete Cal.diy campaign\n\n```sh\nexamples/cal-diy-qualification/qualify.sh\n```\n\nThe script owns an isolated Compose project and evidence directory. Review its\ndestructive scope and configured paths before running it.\n\n### Narrow Cal.diy reservation probe\n\n```sh\nAIP_E2E_CAL_DIY_URL=\"https://isolated-cal.example/api\" \\\nAIP_E2E_CAL_DIY_API_KEY=\"$CAL_DIY_TEST_API_KEY\" \\\nAIP_E2E_CAL_DIY_EVENT_TYPE_ID=\"42\" \\\nAIP_E2E_CAL_DIY_SLOT_START=\"2050-01-01T10:00:00Z\" \\\ncargo test -p aip-connector-cal-diy \\\n  live_cal_diy_reservation_round_trip_uses_the_aip_connector \\\n  -- --ignored --nocapture\n```\n\nThis probe checks ready health, one reservation, identical replay output, and\nrelease. It does not replace the complete Cal.diy campaign.\n\n### Exact-upstream CrewAI sidecar probe\n\n```sh\nAIP_E2E_CREWAI_URL=\"http://127.0.0.1:8090\" \\\nAIP_E2E_CREWAI_CREW_ID=\"support-qualification\" \\\nAIP_E2E_CREWAI_BEARER_TOKEN=\"$CREWAI_TEST_TOKEN\" \\\ncargo test -p aip-connector-crewai \\\n  exact_upstream_sidecar_executes_a_real_crewai_crew \\\n  -- --ignored --nocapture\n```\n\nThe supplied upstream image installs CrewAI from the exact recorded source\ncheckout. Its qualification registry uses a deterministic, network-free\nlanguage model. The result therefore proves a real CrewAI library mapping, not\na live external model-provider integration.\n\n### Hermes deployed-stack probe\n\n```sh\nGETAIP_SERVER_MCP_ACCESS_TOKEN=\"$MCP_TEST_TOKEN\" \\\nGETAIP_SERVER_NATIVE_BEARER_TOKEN=\"$APPROVER_TEST_TOKEN\" \\\ncargo run -p aip-connector-hermes-agent \\\n  --bin getaip-server-hermes-operator-smoke -- \\\n  --mcp-url \"https://isolated-aip.example/mcp\" \\\n  --server-url \"https://isolated-aip.example\" \\\n  --endpoint-id \"hermes-qualification\"\n```\n\nRepeat `--endpoint-id` to check more than one deployed Hermes endpoint. This\ndriver requires different MCP and native approval credentials and emits a JSON\nreport. Redact the report before retention.\n\nThere is no equivalent source-owned live command for Chatwoot, Dify, or Twenty\nat the pinned revision. A campaign for those products needs a separately\nreviewed driver and must retain its source revision with the evidence.\n\n## Retain a reviewable evidence package\n\nStore raw evidence in an access-controlled location and publish only a\nredacted report. The retained package must include:\n\n- a manifest of AIP, connector, upstream, provider, configuration, and\n  campaign identities;\n- the written claim, scope, operator, reviewer, UTC start, and UTC end;\n- a step ledger with command or request identity, expected invariant, result,\n  and evidence path;\n- redacted requests, responses, stream records, webhooks, and provider reads;\n- approval, action, transaction, idempotency, provider-reference, and audit\n  correlation identifiers;\n- before-and-after provider state plus cleanup verification;\n- process restart, cancellation, timeout, and recovery observations;\n- stdout and stderr, exit status, environment-name inventory, and redaction\n  log;\n- a checksum manifest covering every retained artifact;\n- the final result and every unresolved limitation.\n\nDo not retain bearer tokens, API keys, cookie values, signing secrets, raw\ncustomer data, or complete secret-bearing environment dumps.\n\n## Assign the result\n\nUse exactly one result state.\n\n| Result | Meaning |\n| --- | --- |\n| `PASS` | Every applicable planned assertion passed, evidence is complete, and cleanup was independently confirmed |\n| `FAIL` | An assertion was exercised and produced a contradictory or unsafe result |\n| `BLOCKED` | A prerequisite, safe mutation, cleanup path, or required observation was unavailable |\n| `NOT_RUN` | No campaign was attempted for the identified artifact set |\n\nA skipped ignored test is `NOT_RUN`, not `PASS`. A successful process exit is\ninsufficient when evidence is incomplete. Cleanup failure makes the campaign a\n`FAIL`, even when all product operations previously succeeded.\n\n## Interpret existing reports narrowly\n\nThe retained [Cal.diy report](cal-diy-isolated-live.md) and\n[Hermes Agent report](hermes-agent-isolated-live.md) describe historical\ncampaigns and identify the artifacts they evaluated. They do not qualify the\ncurrent pinned AIP revision by inheritance. Use them as report-shape examples\nand as evidence only for their recorded identities and scope.\n\n## Related documentation\n\n- [Testing and qualification](README.md)\n- [Connector upstream baselines](../connectors/upstream-baselines.md)\n- [Conformance reference](../reference/conformance.md)\n- [Implementation status](../reference/implementation-status.md)\n- [Connector documentation](../connectors/README.md)\n",
    "text": "Live product end-to-end qualification\n\nUse this procedure to prove that one identified AIP build can interact safely\nwith one identified deployment of Cal.diy, Hermes Agent, Chatwoot, Dify,\nCrewAI, or Twenty. A valid campaign exercises a real product boundary in an\nisolated environment, retains redacted evidence, and removes every disposable\neffect before it reports PASS.\n\nThis page describes how to design and review a campaign for GetAIP 2.0.0\nsource revision d7cce13d1d555644d04a4d73c66c95b113737635. The protected\ncontrolled product-fleet gate is not an isolated-live external-provider\ncampaign.\n\nKnow what the source can run\n\nThe pinned source does not contain a general live suite for all six products.\nUse only the entry points in this table, or record a separately reviewed\ncampaign driver as an external artifact.\n\n| Product | Source-owned product-facing entry point | Evidence boundary |\n\n| Cal.diy | Complete isolated Docker qualification and a narrower ignored reservation test | The complete campaign covers governed booking, replay, webhook, restart, and persistence; the ignored test covers reservation, replay, and release only |\n| Hermes Agent | Deployed-stack operator and delegation binary | Exercises MCP and native approval paths against identified Hermes endpoints and a controlled downstream AIP capability |\n| Chatwoot | No dedicated live-product driver at the pinned revision | Deterministic connector tests do not prove a live Chatwoot deployment |\n| Dify | No dedicated live-product driver at the pinned revision | Deterministic connector tests do not prove a live Dify deployment |\n| CrewAI | Ignored exact-upstream sidecar test | Proves the Rust connector and real CrewAI library boundary; the supplied qualification registry uses a deterministic network-free language model |\n| Twenty | No dedicated live-product driver at the pinned revision | Deterministic connector tests do not prove a live Twenty deployment |\n\nThe optional five-product fleet verifier uses controlled upstream behavior and\ndoes not include Twenty. It is useful AIP integration evidence, but it is not\nlive-provider evidence for any product.\n\nDefine the claim before the run\n\nWrite one sentence that names the product, provider deployment, connector\nartifact, AIP artifact, campaign revision, and operations in scope. Do not use\nthe broader phrase “the connector is qualified” when the run covers only a\nsubset of operations.\n\nFor example:\n\nThe identified AIP and connector artifacts completed the documented Cal.diy\nbooking campaign against the identified isolated provider deployment, with\ncleanup confirmed, at the recorded time.\n\nThe campaign must not expand its claim to another provider version, another\nconnector artifact, untested operations, production safety, or portable\nperformance.\n\nComplete the safety and identity preflight\n\nDo not begin a mutating step until every item below has an owner and a recorded\nvalue.\n\n| Check | Required record |\n\n| AIP identity | Full source revision, dirty-tree state, build or image digest, configuration digest, and manifest digest |\n| Connector identity | Connector artifact digest, exact upstream pin, capability-set digest, and declared operation surface |\n| Provider identity | Product version or source revision, deployment or image digest, enabled feature flags, and migration state |\n| Isolation | Dedicated non-production tenant, account, workspace, application, crew, or calendar with no production data |\n| Mutation budget | Exact objects the run may create, update, cancel, or delete, plus a cleanup owner and deadline |\n| Principals | Separate requester and approver identities where approval is tested; least-privilege product identity |\n| Credentials | Protected file or secret-store references; record only names, scopes, expiry, and fingerprints |\n| Network | Endpoint names, trust roots, proxy path, redirect policy, and relevant allowlists without secret values |\n| Time | UTC clock source and measured skew for signatures, deadlines, leases, and replay windows |\n| Evidence | New owner-only directory, redaction rules, retention period, and reviewer |\n\nConfirm that every planned mutation is reversible. If a product operation has\nno reliable cleanup path, use a disposable deployment that can be destroyed as\na unit. Never discover this constraint after writing shared customer data.\n\nRun one complete campaign\n\nUse the same sequence for every product, adapting only the product operations.\n1. Record all identities and prove that the isolated target is ready.\n2. Run the connector's deterministic tests and the applicable conformance\n   checks. Retain their results separately from live evidence.\n3. Perform authenticated, non-mutating discovery and baseline reads.\n4. Submit one governed mutation through AIP, including approval when required.\n5. Repeat the same action, idempotency key, or transaction identity and verify\n   that it does not create an additional external effect.\n6. Exercise the supported cancellation, timeout, or recovery boundary without\n   inventing behavior the connector does not declare.\n7. When webhooks are supported, verify the raw-body signature, freshness,\n   duplicate suppression, and rejection of invalid authentication.\n8. Restart the relevant AIP or connector process and verify durable status,\n   provider references, and recovery behavior.\n9. Remove or revert every created provider object and verify its absence or\n   terminal cleanup state through an independent read.\n10. Redact, hash, inventory, and review the retained evidence before assigning\n    a result.\n\nRecord unsupported operations as NOTRUN; do not replace them with similar\noperations or silently omit them.\n\nApply the product-specific minimum\n\nA product campaign is complete only when it covers the applicable minimum in\nthis table. The connector documentation defines the exact inputs, exclusions,\nand product-specific cleanup steps.\n\n| Product | Minimum live assertions |\n\n| Cal.diy | Ready health; availability read; mutation-free booking plan; independent approval; one booking; identical idempotent replay; signed webhook acceptance and replay rejection; restart recovery; booking or reservation cleanup |\n| Hermes Agent | Endpoint health and discovery; streaming result; independent approval; governed AIP tool lifecycle; scoped delegation; out-of-scope rejection; cancellation or terminal status; reconnect or restart observation |\n| Chatwoot | Account-scoped read; reversible conversation, private-note, status, or handoff mutation; idempotent replay; authenticated webhook; event-loop prevention; object or state cleanup |\n| Dify | Application parameter discovery; applicable synchronous and streaming execution; cancellation; human-input or pause state when supported; task recovery; applicable knowledge mode; cleanup of disposable inputs and tasks |\n| CrewAI | Sidecar authentication; exact registered crew discovery; real CrewAI crew execution; stream and terminal status where enabled; cancellation; replay behavior; restart fencing; one authoritative writer; separate identification of any external model provider |\n| Twenty | OpenAPI-derived core-record read; reversible record lifecycle; metadata read; webhook HMAC and nonce replay rejection; credential redaction; proof that the connector did not use the excluded GraphQL or legacy API-key paths; record cleanup |\n\nIf the product cannot supply one applicable assertion, report the gap and stop\nat BLOCKED or FAIL. A deterministic substitute cannot fill a live-product\ngap.\n\nUse only existing source-owned entry points\n\nRun commands from the root of the pinned AIP checkout. The examples below show\nthe available interfaces; operators must supply their own protected values and\nmust satisfy the preflight before execution.\n\nComplete Cal.diy campaign\n\nexamples/cal-diy-qualification/qualify.sh\n\nThe script owns an isolated Compose project and evidence directory. Review its\ndestructive scope and configured paths before running it.\n\nNarrow Cal.diy reservation probe\n\nAIPE2ECALDIYURL=\"https://isolated-cal.example/api\" \\\nAIPE2ECALDIYAPIKEY=\"$CALDIYTESTAPIKEY\" \\\nAIPE2ECALDIYEVENTTYPEID=\"42\" \\\nAIPE2ECALDIYSLOTSTART=\"2050-01-01T10:00:00Z\" \\\ncargo test -p aip-connector-cal-diy \\\n  livecaldiyreservationroundtripusestheaipconnector \\\n  -- --ignored --nocapture\n\nThis probe checks ready health, one reservation, identical replay output, and\nrelease. It does not replace the complete Cal.diy campaign.\n\nExact-upstream CrewAI sidecar probe\n\nAIPE2ECREWAIURL=\"http://127.0.0.1:8090\" \\\nAIPE2ECREWAICREWID=\"support-qualification\" \\\nAIPE2ECREWAIBEARERTOKEN=\"$CREWAITESTTOKEN\" \\\ncargo test -p aip-connector-crewai \\\n  exactupstreamsidecarexecutesarealcrewaicrew \\\n  -- --ignored --nocapture\n\nThe supplied upstream image installs CrewAI from the exact recorded source\ncheckout. Its qualification registry uses a deterministic, network-free\nlanguage model. The result therefore proves a real CrewAI library mapping, not\na live external model-provider integration.\n\nHermes deployed-stack probe\n\nGETAIPSERVERMCPACCESSTOKEN=\"$MCPTESTTOKEN\" \\\nGETAIPSERVERNATIVEBEARERTOKEN=\"$APPROVERTESTTOKEN\" \\\ncargo run -p aip-connector-hermes-agent \\\n  --bin getaip-server-hermes-operator-smoke -- \\\n  --mcp-url \"https://isolated-aip.example/mcp\" \\\n  --server-url \"https://isolated-aip.example\" \\\n  --endpoint-id \"hermes-qualification\"\n\nRepeat --endpoint-id to check more than one deployed Hermes endpoint. This\ndriver requires different MCP and native approval credentials and emits a JSON\nreport. Redact the report before retention.\n\nThere is no equivalent source-owned live command for Chatwoot, Dify, or Twenty\nat the pinned revision. A campaign for those products needs a separately\nreviewed driver and must retain its source revision with the evidence.\n\nRetain a reviewable evidence package\n\nStore raw evidence in an access-controlled location and publish only a\nredacted report. The retained package must include:\n• a manifest of AIP, connector, upstream, provider, configuration, and\n  campaign identities;\n• the written claim, scope, operator, reviewer, UTC start, and UTC end;\n• a step ledger with command or request identity, expected invariant, result,\n  and evidence path;\n• redacted requests, responses, stream records, webhooks, and provider reads;\n• approval, action, transaction, idempotency, provider-reference, and audit\n  correlation identifiers;\n• before-and-after provider state plus cleanup verification;\n• process restart, cancellation, timeout, and recovery observations;\n• stdout and stderr, exit status, environment-name inventory, and redaction\n  log;\n• a checksum manifest covering every retained artifact;\n• the final result and every unresolved limitation.\n\nDo not retain bearer tokens, API keys, cookie values, signing secrets, raw\ncustomer data, or complete secret-bearing environment dumps.\n\nAssign the result\n\nUse exactly one result state.\n\n| Result | Meaning |\n\n| PASS | Every applicable planned assertion passed, evidence is complete, and cleanup was independently confirmed |\n| FAIL | An assertion was exercised and produced a contradictory or unsafe result |\n| BLOCKED | A prerequisite, safe mutation, cleanup path, or required observation was unavailable |\n| NOTRUN | No campaign was attempted for the identified artifact set |\n\nA skipped ignored test is NOTRUN, not PASS. A successful process exit is\ninsufficient when evidence is incomplete. Cleanup failure makes the campaign a\nFAIL, even when all product operations previously succeeded.\n\nInterpret existing reports narrowly\n\nThe retained Cal.diy report (cal-diy-isolated-live.md) and\nHermes Agent report (hermes-agent-isolated-live.md) describe historical\ncampaigns and identify the artifacts they evaluated. They do not qualify the\ncurrent pinned AIP revision by inheritance. Use them as report-shape examples\nand as evidence only for their recorded identities and scope.\n\nRelated documentation\n• Testing and qualification (README.md)\n• Connector upstream baselines (../connectors/upstream-baselines.md)\n• Conformance reference (../reference/conformance.md)\n• Implementation status (../reference/implementation-status.md)\n• Connector documentation (../connectors/README.md)\n"
  },
  "integrity": {
    "algorithm": "sha256",
    "sourceDigest": "0aa0abd085f83ff5a2679e31aa869c6de6ac9134fb41811ca876f871604575b4"
  }
}
