Blog/Recover a media batch with partial failures
OverviewAll posts
Tutorial

Recover a media batch with partial failures

Track each media item, distinguish timeouts from failures, and resume the right stage without resubmitting successful jobs.

VTornado API team
Covered in this article
Persist an item manifest
Classify partial outcomes
Resume only the failing stage
6 min reading time
Published September 28, 2026
VProduct guides by Velys Software

Recover a media batch by tracking each input independently and resuming from its last known stage. Do not resubmit every source because one item failed or because your application stopped waiting. A mixed result needs an item-level recovery plan.

Here, “batch” means a group of media tasks owned by your application. The design works even when you submit sources as separate jobs; it does not assume that a provider's bulk endpoint makes the entire group atomic.

Create a manifest before submitting work

Give the group an internal run ID and each input an item ID. Store the intended source and output settings, submission state, remote job ID when known, and downstream outcome. Persist this record before distributing work to background workers.

Keep the item identity separate from the source URL. The same source can legitimately require two different outputs. Conversely, repeating a failed network call should not automatically create a new business item.

A useful manifest distinguishes “not submitted” from “submission outcome unknown.” If the connection drops after a remote service accepts work, an empty job-ID field alone does not prove that no job exists.

Separate the stages of success

For each item, distinguish acceptance, media completion, usable delivery, and downstream processing. An API request can succeed while a job later fails. A completed media job can also produce a file that the next service cannot access or accept.

Item observationRecovery starting point
Job ID known; still activeContinue checking that existing job
Completed; delivery not yet consumedInspect the delivered file and access path
Media succeeded; analysis failedResume analysis from retained media
Terminal unsuccessful resultInspect its reason before choosing a retry
Submission response was lostReconcile the uncertain submission

Tornado's job status reference distinguishes active, completed, and unsuccessful outcomes. Preserve those distinctions in your manifest rather than reducing everything to a single failed-batch flag.

Build a recovery plan without creating jobs

The following Python example classifies already collected observations. It does not call Tornado API, enqueue work, or decide whether a failure is retryable. Its records are application data, not a batch-response schema.

ACTIVE = {"Pending", "Processing"}
STOPPED = {"Failed", "Warning", "Skipped", "Cancelled", "CancelledByAdmin"}

def next_action(item):
    status = item.get("status")
    if status == "Completed":
        return "check_delivery"
    if status in ACTIVE:
        return "resume_status_lookup"
    if status in STOPPED:
        return "review_reason"
    return "reconcile_unknown"

items = [
    {"item_id": "a", "status": "Completed"},
    {"item_id": "b", "status": "Processing"},
    {"item_id": "c", "status": "Failed"},
    {"item_id": "d"},
]
plan = {item["item_id"]: next_action(item) for item in items}
assert plan == {
    "a": "check_delivery",
    "b": "resume_status_lookup",
    "c": "review_reason",
    "d": "reconcile_unknown",
}
print(plan)

The example and assertion were executed locally on these synthetic records. That checks the classification shown, not a live batch integration. The conservative result for Failed is review, not automatic resubmission; unknown status also stays visible.

In production, use your durable item ID and validated status observations. An in-memory dictionary is only a demonstration. If multiple workers can update an item, protect transitions with a transaction or compare-and-set operation appropriate to your database.

Treat timeouts as uncertainty

A client deadline means your application stopped waiting. It does not establish that the remote operation stopped. When you already have a job ID, resume status lookup instead of sending the source again.

For an ambiguous creation attempt, use the provider's documented reconciliation or idempotency contract where applicable. Do not invent a new request key for every retry of the same logical operation. AWS's discussion of safe retries explains why request identity matters when an earlier response is missing; its API guarantees should not be assumed to apply unchanged to another service.

The Tornado Python tutorial demonstrates continuing with an existing job ID. Store that ID durably rather than relying on console output.

Retry the failing stage with a bounded policy

A private or unavailable source, an expired access link, and a transient network error require different responses. First classify the cause using the information available. Correct configuration problems before repeating work, and respect cancellation as an intentional outcome.

For operations you have determined are safe to retry, record an attempt count, next eligible attempt time, and stopping condition. Keep retry traffic within your account and application limits. Avoid a loop that immediately resubmits every failed item whenever the batch summary is refreshed.

If delivery already succeeded, preserve its object reference. An analysis timeout should ordinarily lead to investigation of that analysis task, not automatic re-ingestion. Where the downstream request itself has an uncertain outcome, reconcile it before starting another copy.

Report progress without hiding failures

Show separate counts for active items, usable results, unsuccessful outcomes, and unresolved submissions. Define your batch completion policy explicitly: all items terminal, all required outputs accepted, or another rule your product needs.

A batch can be finished while still containing failures. Let users retrieve successful results without making them wait for every problematic source, where your product permits partial delivery. Keep the unsuccessful items visible with enough context for a deliberate recovery action.

Do not calculate a success rate by silently excluding unknown outcomes. Record the observation window and denominator so a report can be interpreted later.

Verify interruption and partial recovery

Before scaling, exercise a mixed set of synthetic outcomes, concurrent updates, a worker restart after submission, and a lost downstream response. Check that completed items are not automatically recreated and that uncertainty survives a restart.

Retain old and new attempt identifiers when a deliberate retry creates new remote work. This makes costs and results traceable without overwriting the original evidence.

Start with a small authorized workload and a durable manifest. Use the failure and retry guide to classify unsuccessful items, then resume only the stage that needs work. A recoverable batch is one whose individual outcomes remain understandable after interruption.