A PBS member station recently lost access to roughly seventy years of archived footage when the cloud storage vendor it depended on shuttered without adequate warning. The content — news segments, documentary footage, local history recordings spanning back decades — was not recoverable. The vendor was gone. The data was gone with it.
This is not primarily a story about cloud contracts or vendor due diligence, though both matter. At the file-format level, it is a precise and brutal illustration of what “stored” actually means, and why the container a file lives in — the codec it was encoded with, the wrapper format it was packaged into, the platform-specific metadata that accompanied it — determines whether that content survives its original infrastructure or dies with it.
What “Archived” Actually Means Technically
The word archive implies permanence. In practice, it describes a relationship between a file and the system that can read it.
Broadcast media archives are especially vulnerable here because video formats are not self-contained the way a plain-text document is. A .mov file wrapping ProRes 422 footage is only as readable as the decoder available to the system attempting to open it. The same is true for MXF containers — the Material Exchange Format used heavily in broadcast production — which can wrap dozens of different codec streams. An MXF file containing IMX-encoded video from twenty years ago may open cleanly in one NLE today and fail silently in another. The container is readable; the codec stream inside it may not decode without legacy libraries.
Older broadcast formats compound this further. Betacam SP, for example, had to be digitized into something — and whatever choices were made during that initial digitization (codec, bit depth, color space tagging, whether the container embedded timecode correctly) are baked into the archive copy. If the digitization was done into a proprietary format tied to a specific hardware encoder or software platform that has since vanished, the resulting files may be nominally present while being functionally inaccessible.
The Specific Failure Mode Cloud Storage Introduces
Local storage fails predictably: drives degrade, RAID arrays drop members, tape decays. These failures are gradual enough that monitoring catches them, and the response is restoration from another copy. Cloud vendor failure is categorically different — it can be instantaneous and total.
When a vendor folds, several things can happen simultaneously:
- Access revoked before data transfer is possible. Providers shutting down under financial stress do not always provide the sixty- or ninety-day migration windows their contracts promised. The practical window is whatever the billing cycle allows.
- Proprietary transcoding pipelines disappear with the vendor. If the service had ingested files and re-encoded them into its own internal format for storage or delivery, the original may never have left the vendor’s infrastructure. What looked like a backup was actually a transform.
- Metadata schemas become unresolvable. Broadcast archives carry substantial embedded and sidecar metadata — rights information, timecode, scene descriptions, technical parameters. If a vendor’s platform stored metadata in a proprietary database layer rather than embedding it into the files themselves, that information is lost even if the raw media somehow survives.
The last point is underappreciated. Embedded metadata — XMP sidecar files, ID3-style fields baked into container headers, SMPTE-standard timecode tracks — survives a vendor change because it travels with the file. Metadata in a vendor’s proprietary database does not. It exists only as long as the vendor does.
Format Choices That Survive Infrastructure Changes
This distinction between portable and platform-dependent storage is where format decisions made years earlier either protect an archive or hollow it out.
Open, well-documented container formats have a structural advantage here. An MXF file following SMPTE published specifications can, in principle, be decoded by any conforming implementation, not just the one that created it. An MP4 container wrapping H.264 is readable by essentially any modern media software. The risk is the codec stream inside, not the container — proprietary codec variants narrow the field of compatible decoders significantly.
For photographic archives specifically — and many broadcast archives include enormous libraries of still images alongside video — the equivalent question is whether still images were stored as open formats or as proprietary RAW variants. A CR2 file from a Canon camera body that was discontinued a decade ago depends on decoder libraries that someone must continue to maintain. DNG, by contrast, was designed around this exact survivability problem: it embeds the raw sensor data alongside a fully developed JPEG reference, a copy of the original proprietary RAW if one existed, and detailed metadata describing the sensor characteristics needed to render the file correctly. The ongoing formalization of DNG as a standard — covered in more detail in our article on After Over 20 Years, DNG Becomes the Official RAW Standard — is relevant precisely because standardization means the specification is maintained independent of any single vendor.
The Archive Is Not the Backup
One pattern that likely contributed to the PBS station’s loss: treating the cloud archive as the authoritative copy rather than a redundant one. This is an easy trap. Cloud storage is reliable enough day-to-day that it feels like permanence. It is not.
A genuine archival strategy has at least two legs that are independent of each other — meaning they cannot both fail for the same reason at the same time. A cloud storage account and its geo-redundant replica within the same vendor’s infrastructure are not two independent legs. They are one leg with two feet. The vendor’s insolvency takes both.
The 3-2-1 rule is old enough to feel obvious and ignored often enough to keep being relevant: three copies, on two different media types, with one copy offsite. The “offsite” requirement exists to guard against a single physical disaster. The “different media types” requirement exists to guard against a single failure mode — including a vendor’s business failure.
Formats matter here too. An offline backup written to LTO tape in a well-documented codec buys time. A backup consisting of files in a proprietary format tied to playback software that the organization no longer licenses does not, even if the tape survives.
What an Organization Can Actually Do Now
If you manage a media archive — broadcast or photographic — the actionable audit is not complicated, but it requires checking specifics rather than assuming general health:
- Identify every proprietary format in the archive. Proprietary codecs, platform-specific containers, and software-locked metadata stores are the risk surface. Open, documented formats are not risk-free, but the risk is meaningfully different.
- Verify that metadata is embedded in files, not only in a database. Export a sample file and open it in a different application. If the metadata travels with the file, it will appear. If it requires the original platform to display, you have a dependency problem.
- Confirm that your cloud storage and your local or tape backup cannot both be wiped by the same event. Vendor insolvency, a security breach, or an account compromise should be on that list alongside fire and flood.
- Check codec decodability now, not when you need the file. Open a selection of your oldest files in software other than whatever created them. Silent decode failures — files that open but display corrupted frames or dropped audio — are more common than outright unopenable files.
- Document what you have. An inventory of formats, codecs, container types, and storage locations is a prerequisite for any migration or recovery plan. Discovering that the archive includes formats no one recognized is a discovery best made on a Tuesday afternoon, not during a recovery.
The PBS station’s loss is a specific, documented instance of something the file-format and archival communities have been warning about for years. The mechanism is not mysterious. Formats and platforms become inaccessible when their supporting infrastructure disappears — and the only reliable hedge is making sure the files themselves, and the metadata inside them, are not dependent on that infrastructure surviving. Our broader File Formats coverage returns to this dependency problem repeatedly, because it runs through nearly every format decision a working organization makes.
The footage is gone. The lesson, at minimum, should not be.