Blog/Resume media jobs after a worker restart
OverviewAll posts
Tutorial

Resume media jobs after a worker restart

Persist remote job IDs with Python and SQLite, recover interrupted media workers, and handle uncertain submissions without blindly creating duplicate jobs.

VTornado API team
Covered in this article
Persist submission attempts and job IDs
Recover state with SQLite
Reconcile uncertain submissions
7 min reading time
Published October 2, 2026
VProduct guides by Velys Software

Resume an interrupted media worker from a durably stored job ID. A process restart does not tell you whether remote media processing stopped, finished, or never began. First recover the record of your submission, then check the existing job before deciding whether new work is needed.

This guide implements the local persistence step missing from a simple polling script. It uses Python and SQLite to retain a submission attempt across connection closure and reopening. The same boundary matters when your application uses another database.

Record the attempt before contacting the API

Give each intended submission an application-owned attempt ID. Associate it with the business item, source and requested output settings in your application. Persist the attempt before sending the creation request, then persist the returned remote job ID as soon as you receive it.

These are two separate writes around a network operation. They cannot become a single atomic transaction just because the local database supports transactions. A worker can stop after the API accepts a request but before the job ID reaches your database.

That leaves two recovery cases:

Durable recordNext step after restart
Remote job ID savedLook up that existing job
Attempt exists, remote ID unknownReconcile the uncertain submission

An unknown remote ID is not permission to resubmit. The worker might have stopped before sending anything, or after acceptance. Preserve that uncertainty instead of treating both cases as a failed job.

Save and recover an acknowledged job ID

The example below requires Python 3.12 or later. It creates a temporary, file-backed database, records two synthetic attempts, attaches a synthetic job ID to one, closes the connection, and recovers the records through a new connection. It makes no API requests.

import sqlite3
from contextlib import closing
from pathlib import Path
from tempfile import TemporaryDirectory


def connect(path):
    return sqlite3.connect(path, autocommit=False)


def begin_attempt(db, attempt_id):
    with db:
        db.execute(
            "INSERT INTO attempts (attempt_id) VALUES (?)",
            (attempt_id,),
        )


def remember_job(db, attempt_id, job_id):
    if not job_id:
        raise ValueError("A nonempty remote job ID is required")
    with db:
        changed = db.execute(
            "UPDATE attempts SET job_id = ? "
            "WHERE attempt_id = ? AND (job_id IS NULL OR job_id = ?)",
            (job_id, attempt_id, job_id),
        ).rowcount
        if changed != 1:
            raise ValueError("Missing attempt or conflicting remote job ID")


def recovery_plan(db):
    return [
        (attempt_id, "lookup" if job_id else "reconcile", job_id)
        for attempt_id, job_id in db.execute(
            "SELECT attempt_id, job_id FROM attempts ORDER BY attempt_id"
        )
    ]


with TemporaryDirectory() as directory:
    path = Path(directory) / "jobs.sqlite3"
    with closing(connect(path)) as db:
        with db:
            db.execute(
                "CREATE TABLE attempts ("
                "attempt_id TEXT PRIMARY KEY NOT NULL, job_id TEXT)"
            )
        begin_attempt(db, "item-a-attempt-1")
        begin_attempt(db, "item-b-attempt-1")
        # In your application, call this after receiving the real job ID.
        remember_job(db, "item-a-attempt-1", "example-job-a")

    # A fresh connection represents the worker recovering its local state.
    with closing(connect(path)) as db:
        plan = recovery_plan(db)
        assert plan == [
            ("item-a-attempt-1", "lookup", "example-job-a"),
            ("item-b-attempt-1", "reconcile", None),
        ]
        print(plan)

This exact example was executed locally. Its assertion checks that committed records survive connection closure and reopening, and that an uncertain attempt remains uncertain. It does not simulate a machine crash, test storage durability under power loss, or establish any remote API guarantee.

The Python SQLite reference documents transaction handling and parameter binding. Here, with db commits successful transactions or rolls them back on an exception; closing separately closes the connection. SQL values use placeholders rather than string interpolation.

The temporary directory makes the demonstration disposable. A real worker must store its database on persistent storage that survives container replacement. Treat a failed database write as a recovery problem; do not proceed as though the attempt or job ID was saved.

Continue with the remote job you already created

For each lookup record, call the documented job status endpoint, GET /jobs/{id}, with your API key in the required header. The Python media tutorial shows the request and polling flow using an existing job ID.

Inspect the returned status rather than assuming that every recovered job needs polling forever. Active work needs later status checks. Completed work needs delivery inspection. An unsuccessful terminal outcome needs review of its reason. A failed status request by itself does not prove that the media job failed.

Keep the API key out of job records and diagnostic output. Store only the identifiers and operational information needed for recovery, with access controls appropriate to your application.

Keep uncertain submissions out of automatic retry loops

A reconcile record should enter an investigation path. Use available request records and documented provider support or reconciliation mechanisms to establish what happened. This example does not assume a Tornado creation-request idempotency key or a lookup-by-client-attempt-ID feature.

If you deliberately create replacement work after investigating, assign it a new attempt record linked to the same business item. Retain the original uncertain attempt. Overwriting it would hide the possibility that two remote jobs exist.

The local primary key prevents insertion of the same attempt twice; it does not guarantee exactly-once remote execution. Likewise, remember_job rejects replacement of a known job ID with a different one, but cannot remove the network-to-database uncertainty window.

Persist the handoff after media processing

A remote job ID is the starting point, not your entire workflow state. Record when your application accepts the delivered result and when the downstream consumer finishes. Otherwise, a restarted worker may send a completed file to the consumer twice.

Persist a stable object reference where available instead of relying only on a temporary access URL. If a recovered link no longer works, follow the expired URL versus missing file guide before deciding to regenerate media.

Multiple workers also need a claim or lease mechanism and guarded state transitions. The demonstration has one worker and intentionally supplies neither a queue nor a distributed locking design. Test concurrent claims, database failures and lost downstream responses before deploying your own recovery loop.

Start by adding durable attempt records to one workflow, then verify restart recovery using controlled interruptions. Use the partial batch recovery guide when that workflow expands to multiple items, or build your first Tornado media workflow to connect submission, status and delivery from the beginning.