Delivering media into Amazon S3 is only part of a downstream processing workflow. Before starting transcription, indexing or editing, your application should identify the delivered object, check it through its own storage access, and record which object the consumer will read.
This guide focuses on that handoff. It complements the Tornado S3 configuration reference, which covers bucket setup and credentials, and the storage and retention guide. It does not assume that successful job completion proves your downstream application can read or process the file.
Separate delivery configuration from consumer access
Configure the destination used by your actual submission path. Tornado's dashboard storage configuration and its older per-API-key S3 route are different surfaces; do not assume updating one changes every integration. Follow the current S3 reference for the required fields and permissions, then validate the path you intend to use.
Keep the bucket private. Give the downstream application its own appropriate read access rather than copying Tornado's upload credentials into every service. Uploading, reading and managing retention are different responsibilities.
For your first end-to-end check, use one authorized source and a recognizable business item ID. Record the remote job ID as well. This gives you a way to trace the input, delivered object and consumer result without exposing credentials or temporary signed URLs in logs.
Identify the exact object before reading it
Use the completed job's delivery information and your configured destination to establish the bucket and full object key. Do not rebuild the key from a video title, append a folder prefix twice, or assume that a filename is globally unique.
A useful application record contains the business item ID, job ID, destination, complete object key, observation time and consumer state. These are fields in your own record, not a proposed Tornado API response schema.
For versioned buckets, retain a specific version identifier when your integration supplies or retrieves one. Otherwise, decide how your application prevents another writer from replacing the key between inspection and consumption. A saved path is not automatically an immutable file identity.
Inspect metadata with your consumer's credentials
Amazon S3's HeadObject operation reads object metadata without downloading the body. With an already configured AWS CLI profile, the following template requests metadata for an existing object in a general purpose bucket:
aws s3api head-object \
--bucket YOUR_BUCKET \
--key 'YOUR_EXACT_OBJECT_KEY' \
--output json \
--no-cli-pager
Replace both placeholders. Use the same authorized identity your consumer will use, or an equivalent test identity. The command is a documentation-checked template; it was not run against a live customer bucket for this article.
The AWS CLI reference describes the response fields and permission requirements. Inspect ContentLength, ContentType, ETag and, when available, VersionId. A successful metadata request does not validate the media's codec, duration or decodability.
Without bucket-list permission, a missing object may produce a 403 response. Do not conclude from that response alone that delivery failed or that adding public access is the fix. Check identity, destination and exact key first.
Make the next action explicit
The example below applies an application policy to metadata already collected from a successful HEAD response. It does not connect to AWS, validate credentials or decode media. The byte count and MIME type are synthetic fixtures.
def inspect_metadata(metadata, expected_bytes=None):
size = metadata.get("ContentLength")
if type(size) is not int or size <= 0:
return "review_size"
if expected_bytes is not None and size != expected_bytes:
return "review_size_mismatch"
content_type = metadata.get("ContentType")
if content_type not in {"video/mp4", "audio/mpeg", "audio/mp4"}:
return "inspect_media_type"
return "validate_media_with_consumer"
sample = {"ContentLength": 4096, "ContentType": "video/mp4"}
assert inspect_metadata(sample, 4096) == "validate_media_with_consumer"
assert inspect_metadata(sample, 8192) == "review_size_mismatch"
assert inspect_metadata({"ContentLength": 0}) == "review_size"
assert inspect_metadata({"ContentLength": "4096"}) == "review_size"
assert inspect_metadata({"ContentLength": 4096}) == "inspect_media_type"
These five assertions were executed locally. The accepted MIME types are an illustrative application policy, not Tornado's complete format list. Adapt the policy to your consumer. Missing or unfamiliar metadata triggers inspection rather than automatic deletion or media resubmission.
Only set expected_bytes from a trustworthy measurement of the same output. A source page's reported size, a predicted output size and a billing quantity are not interchangeable with the delivered object's byte count. Even an exact size match does not prove that the contents are correct.
Do not treat every ETag as a file hash
An ETag is not a universal MD5 checksum. Multipart uploads and encryption choices affect what it represents. AWS's object integrity documentation distinguishes full-object and composite checksums.
If integrity verification is required, choose a documented checksum algorithm and scope, and compare values that describe the same bytes. Do not compare a composite checksum with a locally calculated whole-file digest and interpret the mismatch as corruption. This article does not claim that Tornado exposes an end-to-end checksum contract for every delivery.
Metadata inspection is a gate before processing. Your consumer still needs to open the file and validate its own requirements. For video dimensions, use the delivered resolution guide; for transcription inputs, check audio output compatibility.
Record consumer completion separately
After the consumer accepts the object, record its task identifier and eventual result alongside the media job ID. If the consumer times out, investigate that operation before fetching the source again. The existing S3 object may remain perfectly usable.
Choose retention around your recovery needs. Deleting media immediately after enqueueing downstream work can make a later retry impossible. Conversely, keeping every file forever without a policy obscures ownership and costs. Define which stage authorizes cleanup and which records remain afterward.
Start with one complete path: submit, observe delivery, inspect the exact object, process it, and record the result. Then exercise denied reads, unexpected metadata and downstream failure using controlled fixtures. The first media workflow guide connects the API steps; the worker restart guide explains how to retain job identity when the process stops.