ADACTION
DOCK

ACTIONDOCK FIELD NOTES

How to Retry Failed AI Agent Webhooks Without Repeating API Writes

Recover failed AI-agent webhook deliveries: repair the receiver, retry only the notification, deduplicate stable IDs, and verify job state before resuming work.

An amber notification envelope travels toward a receiving tray beside an untouched mechanism enclosed in teal glass.

To recover a failed AI-agent webhook, fix the receiving endpoint and redeliver the notification—not the underlying API write. Verify the redelivered message, deduplicate its stable delivery ID, and read authenticated job state before resuming the workflow. A failed callback does not mean the write failed, and it does not prove the receiver did nothing.

In ActionDock, the signed-in workspace owner can retry a terminal failed callback from Callback → Recent deliveries. That queues the original event again without rerunning its job, repeating approval, or sending another provider request. The distinction is essential when the agent is waiting for news about an operation that may already be complete.

Separate the three states before touching Retry

A webhook connects systems with different records of progress. Diagnose the record that actually failed.

Layer What its state tells you Recovery belongs here
ActionDock job What the gateway recorded about the approved operation Read the job; reconcile an unknown provider outcome separately
Callback delivery Whether the receiver acknowledged the notification over HTTP Repair the endpoint and retry the notification if eligible
Receiver processing Whether your application stored and processed the event Inspect your receipt store and work queue; recover local processing

For example, an approved PATCH /inventory/items/42 might update a reorder threshold successfully while the callback endpoint returns 503. Submitting the inventory write again would address the wrong failure. Redelivering the stored success event lets the receiver catch up with the existing job.

The opposite case matters too: the receiver might persist the event, then lose the connection before its 2xx reaches ActionDock. The delivery can look failed even though local processing is underway. A durable duplicate check keeps recovery from starting that work twice.

This separation is common webhook practice, not a special provider integration. GitHub’s failed-delivery guide tells operators to inspect the delivery failure and redeliver the notification. Retry policies differ: GitHub does not automatically redeliver failed webhooks, while ActionDock has bounded automatic retries.

1. Read the job and the delivery diagnostics

Start with the known job ID. Use authenticated GET /v1/jobs/{id} or MCP get_job; do not infer its state from a missing callback. Keep the original job ID and submission idempotency key while investigating.

Then open Callback → Recent deliveries and record the event ID, delivery ID, last HTTP response or error, Attempts this cycle, and any next scheduled attempt. In ActionDock, non-2xx responses and transport failures trigger automatic retries, which stop after eight attempts per cycle. Startup recovery may repeat an interrupted attempt, and an owner can start another cycle, so eight is not a lifetime delivery limit.

If the delivery is still pending or in flight, it is not eligible for manual retry. Refresh rather than creating a competing notification. If it is already delivered, inspect the receiver’s queue: HTTP acknowledgment does not establish that downstream agent processing finished.

The dashboard shows the latest outcome for each delivery, not an immutable history of every attempt. Preserve relevant receiver logs for your investigation, without copying signing secrets or sensitive payloads into a support ticket.

2. Repair and test the receiver first

Use the failure to choose a concrete repair. Check a wrong URL or route for 404, access restrictions or signature configuration for 401 and 403, and application or storage failures for 5xx. For a connection error, investigate public DNS, HTTPS reachability, and the certificate. These are diagnostic starting points, not proof of a particular cause.

If a handler waits for an agent to reason before responding, move that work behind durable acceptance. Verify the request, atomically store its receipt and queue entry, then return 2xx promptly. GitHub’s webhook best practices recommend asynchronous processing so long-running work does not hold the delivery connection open. ActionDock’s callback sender uses a ten-second socket inactivity timeout, not a total delivery deadline.

For signature failures, preserve the raw request body and check the configured signing secret. Do not bypass verification to clear the backlog. The callback verification guide covers the exact HMAC input, timestamp checks, and atomic receipt storage; this recovery procedure assumes those checks remain enabled.

Confirm that the configured callback is enabled and points to the repaired receiver. Coordinate secret changes with that receiver before retrying old deliveries. Changing or removing a callback does not cancel a notification already being sent.

3. Queue only the failed notification

For a terminal failed delivery, choose Retry notification. ActionDock requires an enabled, configured callback and the signed-in owner’s dashboard session. An agent API key cannot perform this operation, and there is no MCP notification-retry tool.

The owner-session HTTP endpoint is POST /v1/dashboard/callback-deliveries/{deliveryId}/retry, with a same-origin Origin header. Its successful response is:

{"retried":true}

The status is 202: the retry is durably queued, not confirmed delivered. A missing or foreign-workspace delivery returns 404. A pending, in-flight, or delivered notification—or a missing or disabled callback—returns 409 callback_retry_unavailable. Refresh the state and configuration after a conflict; do not treat it as permission to rerun the job.

Queueing resets the new cycle’s attempt count to zero. The previous HTTP response and error remain visible until another delivery result arrives, so an old diagnostic immediately after queueing is not necessarily a new failure.

ActionDock retries the original event envelope with the same event ID and delivery ID, but sends it using the current callback URL and signing secret. Check that the current destination is authorized to receive that historical payload, especially if the receiver changed during the outage.

4. Accept a duplicate without duplicating work

Each send calculates a timestamp from the current server time and recomputes its signature. The event’s creation time still describes the original event; it is not the timestamp used to authenticate the new HTTP request. Check freshness against X-ActionDock-Timestamp, and deduplicate the stable X-ActionDock-Delivery, not the signature.

A verified delivery already stored by your receiver should receive a successful acknowledgment without another work enqueue. If the earlier receipt exists but its worker failed, repair that local queue or processing record instead. Notification redelivery is not a substitute for a receiver’s internal recovery mechanism.

Stable identifiers on redelivery are also documented by GitHub. Stripe’s duplicate-event guidance likewise recommends remembering processed event IDs. Those are examples of receiver design, not claims that their headers or retry schedules are ActionDock’s contract.

Keep the receipt claim and initial queue insertion atomic. A process-local set disappears on restart; a separate “check, then insert” can race. Your receiver must remain safe under at-least-once delivery, including a lost acknowledgment and two concurrent copies of the same notification.

5. Verify delivery and resume from current state

Refresh the delivery record and confirm the receiver acknowledged it. Then find the same delivery ID in your durable receipt store and check that its processing completed. These are separate acceptance gates.

Before a consequential next step, read the authenticated job again. The callback wakes the agent; it is not authorization for an arbitrary follow-up write. A new integration.execute proposal still needs its own owner approval. The durable actions workflow explains how callbacks and polling fit around that boundary.

If the job is execution_unknown, redelivering its notification leaves it unknown. Check the provider’s authoritative state and follow the unknown-outcome reconciliation procedure. Neither a delivered callback nor a gateway execution record independently proves the provider’s current external state.

ActionDock controls only actions routed through its MCP or HTTP API. This notification recovery does not control direct provider access, orchestrate the agent’s entire workflow, or guarantee exactly-once execution in another system.

A short outage recovery checklist

Use the ActionDock callback contract for the message format and polling fallback. The recovery target is an acknowledged notification and one durable receiver work item—not a second attempt at an already approved API operation.