Use polling when you need a simple client that asks for progress. Use webhooks when your server should react to completed work without repeatedly checking every job. For a workflow that must recover from missed notifications, combine webhooks with occasional status reconciliation of unresolved jobs.
The important decision is who remembers unfinished work. Neither an HTTP callback nor a polling loop replaces a durable record of the job ID and your application's next action.
Choose by the way your application runs
| Application | Useful starting point | Recovery responsibility |
|---|---|---|
| Short script or local tool | Poll an existing job with a deadline | Save the job ID when the script stops waiting |
| Backend with an HTTPS receiver | Webhook-driven processing | Persist accepted notifications and handle duplicates |
| Background workflow with recovery requirements | Webhooks plus scheduled status checks | Reconcile unresolved jobs and track downstream completion |
Polling is often easier to inspect during initial integration: each response belongs to a request your application just made. The Python media tutorial demonstrates that flow. As the number of simultaneously active jobs grows, account for request volume and your account's applicable limits rather than copying a fixed polling interval from another application.
Webhooks move notification delivery into a separate HTTP interaction. They are useful when the user has closed the browser or the original API request has long since returned. They also require a receiver whose deployment, authentication, persistence and monitoring you own.
Use the correct webhook contract
Tornado's webhook reference describes distinct integration families. Dashboard organization endpoints use events such as job.completed and have a delivery_id. The explicit per-job webhook surface uses a different event contract. Do not reuse a parser simply because both payloads describe a completed job.
Select verification appropriate to the endpoint's configured mode. A simple shared key and an HMAC signature are different checks. Do not accept a missing verification value, or reconstruct a signed body from parsed JSON when the signature contract requires the original bytes.
Keep this boundary separate from business processing. Authenticate and validate the request, durably accept the work, then acknowledge it promptly. Run lengthy media analysis in a worker. Returning success before anything durable exists can lose work if the receiver stops immediately afterward.
Recover when no callback arrives
A missing notification is an observation about your receiver, not proof of media failure. Keep a record of submitted jobs that still need an outcome. For records whose last successful observation is old enough for your policy, look up the same job through the status endpoint.
Tornado exposes GET /jobs/{id} with API-key authentication. Use the returned status to decide what happens next. A completed result leads to delivery inspection; an unsuccessful final outcome needs its own handling. A failed status request leaves the result unresolved. It does not justify creating a replacement job.
The following example selects jobs for a reconciliation pass. Its input is an application-owned record, not a webhook payload. Times are synthetic seconds on one common timeline; the 60-second threshold is an example policy, not a recommended API rate or delivery guarantee.
ACTIVE = {"Pending", "Processing"}
FINAL = {"Completed", "Failed", "Warning", "Skipped",
"Cancelled", "CancelledByAdmin"}
def recovery_action(record, now, stale_after):
if stale_after <= 0:
raise ValueError("stale_after must be positive")
if not record.get("job_id"):
return "reconcile_submission"
status = record.get("status")
if status in FINAL:
return "handle_recorded_outcome"
if status not in ACTIVE and status is not None:
return "review_unknown_status"
last = record.get("last_observed_at")
if last is None or now - last >= stale_after:
return "lookup_existing_job"
return "wait"
job = {"job_id": "example-job", "status": "Processing",
"last_observed_at": 100}
assert recovery_action(job, 159, 60) == "wait"
assert recovery_action(job, 160, 60) == "lookup_existing_job"
assert recovery_action({"job_id": "example-job"}, 160, 60) == "lookup_existing_job"
assert recovery_action({}, 160, 60) == "reconcile_submission"
assert recovery_action(dict(job, status="Completed"), 160, 60) == "handle_recorded_outcome"
assert recovery_action(dict(job, status="Unexpected"), 160, 60) == "review_unknown_status"
All six assertions were executed locally. The function selects an action; it sends no requests and creates no jobs. In production, persist timestamps consistently, handle clock anomalies, limit reconciliation concurrency and schedule retries separately from the last successful observation. A failing lookup must not produce an immediate tight retry loop.
Let both paths converge on one outcome handler
A webhook and a status lookup can discover completion at nearly the same time. Route both through a guarded application transition so they cannot independently launch the same downstream action.
Notification deduplication and business-action deduplication solve different problems. A delivery identifier can recognize the same notification again. Your business record determines whether a particular processing step for that job has already been accepted or finished.
Use a transaction or another appropriate concurrency mechanism for that transition. If saving local state and sending a downstream request are separate operations, document the failure window between them. An outbox or a consumer's documented idempotency mechanism may help; neither should be claimed to provide exactly-once behavior without checking its actual contract.
Do not let an older progress observation overwrite a recorded final outcome. If observations conflict, preserve enough evidence to investigate instead of blindly applying whichever arrived last. Legitimate retries also need explicit attempt identity so a later attempt is not suppressed by an earlier result.
Measure gaps, not just successful callbacks
Monitor how many accepted jobs remain unresolved, how old they are, whether your receiver rejects verification, and whether accepted downstream tasks finish. A healthy callback count can coexist with a small set of jobs that never reach your consumer.
Exercise receiver downtime, duplicate notifications, a completion discovered by polling, and interruption after durable acceptance. Use fixtures first; a successful local test does not establish live delivery latency or provider availability.
Start with the first media workflow, then add the notification path your application needs. Keep the worker restart recovery pattern underneath both approaches, and validate the delivered file before treating media completion as business completion.