NVIDIA announced a major expansion to NVIDIA AI for Media on 9 September 2026, ahead of the IBC conference running 11 to 14 September in Amsterdam. The headline number is that the Synthetic Video Detector (SVD) NIM microservice now reaches 99.3% accuracy on text-to-video content and 97.7% on image-to-video content, and that Wowza will push it into a streaming platform running more than 35,000 video deployments across over 170 countries.
Those accuracy figures are real and they come from NVIDIA. They are also measured along a different axis than the accuracy figures NVIDIA published for the same model seven weeks earlier, and the axis that disappeared is the one that decides whether the detector works on video you actually encounter. That is the story worth reading carefully, and it sits inside a bundle that also ships four separate ways to put synthetic content into video.
What NVIDIA Actually Shipped
AI for Media is NVIDIA's collection of GPU-accelerated SDKs, NIM microservices, playbooks and blueprints for media and entertainment workflows. The IBC expansion touches eight named capabilities plus two frameworks. Several are callable today from the AI for Media model catalogue on build.nvidia.com, which is the practical entry point if you want to test any of this rather than read about it.
| Capability | What it does | Adds or detects synthesis | Named partner |
|---|---|---|---|
| Synthetic Video Detector (SVD) | Frame-by-frame classifier score for whether a clip contains synthetic content | Detects | Wowza, Dalet, TwelveLabs |
| Video Frame Generation (VFG) | Generates new frames between original frames, 2x or 4x frame rate | Adds | Ross Video (Rio Replay) |
| LipSync NIM | Rewrites mouth movement to match a different audio track | Adds | NDI |
| Video Super Resolution (VSR) | Upscales video, reduces noise, blur and compression artifacts, adds 10-bit support | Adds | Available via Video Effects SDK and NIM |
| TrueHDR | Converts SDR to HDR in real time, up to approximately 2,000 nits | Adds | Combinable with VSR and VFG |
| 3D Body Pose | 2D and 3D joint locations from a single camera, no marker rig | Neither | Vizrt (virtual studios) |
| Active Speaker Detection NIM | Identifies who is speaking, no diarization required, adds voice activity detection | Neither | Content Localization workflow |
| Studio Voice Microphone Profiles | Suppresses noise and reverb, then shapes speech into a chosen mic character | Adds | Built on Studio Voice NIM |
| Holoscan for Media plus MXL | Open reference architecture for software-defined live production, now with Media Exchange Layer | Framework | Demo at EBU Stand 10.D21 |
| Sports Intelligence Playbooks | Frameworks for fine-tuning NVIDIA open models on proprietary sports footage | Framework | Machina Sports, Wowza |
Read the middle column again. In a single announcement NVIDIA shipped one tool that flags video as synthetic and four that synthesize parts of video: VFG invents frames that were never captured, LipSync rewrites the mouth, TrueHDR invents dynamic range the camera never recorded, and VSR invents pixels. NVIDIA even notes that VSR, VFG and TrueHDR "can be combined within a single video-effects pipeline." That is not a contradiction on NVIDIA's part, but it does raise a question the announcement never answers, and I come back to it below.

The Accuracy Number Changed Axis
SVD was not new at IBC. NVIDIA first announced the Synthetic Video Detector NIM microservice at SIGGRAPH in July 2026. That post reported accuracy this way: "In NVIDIA testing, the model's accuracy reached up to 92% on uncompressed video, 87% at 15% compression and 82% at 50% compression." It also gave latency: 1080p video in as little as 22 milliseconds on NVIDIA RTX systems and approximately 30 milliseconds on L40 GPUs.
The IBC post reports accuracy this way: "Since its initial release, SVD's accuracy has reached 99.3% for text-to-video content and 97.7% for image-to-video content, with especially large gains on difficult image-to-video cases."
Both statements can be true at once, because they are not measuring the same thing. July broke the number down by how degraded the input video was. September breaks it down by how the fake was generated. September says nothing about compression at all.
That matters more than it might sound. Compression is not an edge case in video, it is the default state of video. A clip that reaches a newsroom, a livestream ingest, a social repost or a compliance queue has been compressed, usually several times, usually badly. In July NVIDIA disclosed that its own detector fell from 92% to 82% as compression went from none to 50%. If that sensitivity still exists, a 99.3% figure quoted without a compression level does not tell an editorial team what the model will do on the footage they receive. If it no longer exists, that would be the more impressive claim of the two, and it is not made.
None of this makes the improvement fake. Text-to-video and image-to-video are meaningful categories, image-to-video genuinely is the harder detection problem, and NVIDIA calling out "especially large gains on difficult image-to-video cases" is a specific, falsifiable claim rather than marketing filler. The point is narrower: 99.3% and 92% are not two points on one line, and anyone treating the jump as a straight 7-point improvement is reading a chart that does not exist.
NVIDIA itself is careful here in a way most coverage will not be. The company positions SVD as helping teams "assess the probability of whether footage is authentic or AI-generated," giving them "another point of analysis in their review process." That is a hedge, and it is the correct one. A classifier score is evidence, not a verdict.

The Detector Lands in Distribution, Not in the Edit
The most consequential part of this announcement is not the accuracy number, it is where SVD is being installed. All three named integrations put the detector upstream or downstream of the creator, never in the creator's hands.
Dalet is building SVD into a cloud-hosted verification workflow for news organizations, so editorial teams submit footage and review scores and metadata inside a Dalet interface. TwelveLabs announced general availability of Compliance by TwelveLabs, which screens content against regional and custom compliance standards and uses SVD to add frame-level authenticity signals to that same screening pass. And Wowza will distribute SVD through its Video Intelligence Framework, analyzing live feeds in real time for detected objects, scenes and signs of AI generation.
The Wowza number is the one to sit with. Wowza Streaming Engine powers more than 35,000 video deployments in over 170 countries, and the framework can run on premises, at the edge, in the cloud, in hybrid setups or fully air-gapped. Synthetic-content scoring is moving from a forensic tool that a specialist invokes deliberately into a property of the streaming infrastructure itself, applied by default, at ingest, to everything.
If you generate or heavily process video, this changes your position without changing your workflow. You do not opt in and you do not see the score. Someone in a compliance queue does.

The Question NVIDIA Does Not Answer
Here is the practical gap. NVIDIA now sells, in one product family, tools that add generated frames to real footage and a tool that scores footage for generated content. It does not say whether the first set trips the second.
Take the flagship example from the announcement. Ross Video is integrating VFG into its Rio Replay platform for AI-assisted slow motion in sports production, currently supporting 6x slow-motion generation with development underway toward 8x interpolation. At 6x, the large majority of frames in that replay were generated by a model rather than captured by a camera. The footage is authentic in every sense a viewer cares about: a real match, real players, a real moment. Structurally, though, most of its frames are synthetic.
Does SVD flag it? NVIDIA does not say. The detector is described as producing a classifier score for whether a clip "contains synthetic content," and a 6x interpolated replay unambiguously contains synthetic content while being an honest record of a real event. The same question applies to a LipSync-dubbed interview, a VSR-upscaled archive clip and a TrueHDR conversion of an SDR master. All four are legitimate broadcast operations. All four insert model output into the pixel stream.
This is not a gotcha, it is the central unsolved problem in synthetic-media detection, and it is why the provenance camp exists at all. But it is a live operational question for anyone whose output passes through a Wowza ingest, and the announcement that ships both halves is the natural place to have addressed it.
Detection Versus Provenance, Same Day
The timing here is worth noting. NVIDIA's detection expansion landed the same day Apple presented its own answer to the same question from the opposite direction. Apple's approach, covered in our piece on Reference Image and photo provenance, signs authenticity at the moment of capture. NVIDIA's approach infers it at the moment of inspection.
The two are not competitors so much as complements with different failure modes. Provenance is cryptographic and near-certain when present, but it only covers content captured by participating hardware, and metadata is trivially stripped by any re-encode or repost. Detection works on anything, including a screenshot of a screenshot, but it is probabilistic, degrades with compression, and cannot distinguish "generated" from "legitimately processed." Anthropic's text watermarking approach sits in a third position again, embedding the signal in the output rather than the container.
A newsroom that wants a real answer in 2027 will run all three, and will still occasionally be wrong.
What the Sports Numbers Actually Claim
NVIDIA Sports Intelligence Playbooks are the other substantial release here, and they carry the announcement's most eye-catching statistic. They give leagues and media companies structured frameworks for fine-tuning NVIDIA open models on their own footage and annotations, spanning data preparation, fine-tuning, inference, evaluation, optimization and deployment, and pulling in Nemotron, NeMo AutoModel, Megatron Bridge and NIM microservices. The documentation is public.
The claimed result: multiple-choice accuracy rising from approximately 53% to 94%, and open-ended evaluation from approximately 5.7% to 66%. Those are large jumps. They also come with a qualifier NVIDIA states plainly and most summaries will drop, which is that the evaluation used "previously unseen footage using question formats similar to those used in training."
Unseen footage is a genuine test. Question formats similar to training is a much softer condition, and it is the difference between a model that understands a sport and a model that has learned the shape of the questions asked about it. A 53% to 94% jump on matched question formats is a real and useful engineering result for anyone building a sports analytics product on their own library. It is not evidence of general sports understanding, and NVIDIA does not claim it is.
Wowza is already applying this, fine-tuning vision language models including Cosmos 3 and Nemotron through the playbooks to detect sports-specific moments in live streams. That is the intended shape of the product: not a general sports model, but a pipeline for turning a rights holder's private archive into a narrow model competitors cannot replicate.

The Rest of the Stack
Two infrastructure pieces round this out. Holoscan for Media, NVIDIA's open reference architecture and developer toolkit for software-defined live production, now integrates Media Exchange Layer, an open way for software-based media functions to exchange live video, audio and data across a distributed environment. NVIDIA has also brought its Content Localization technologies into that toolkit as a reference workflow covering captions, translated audio, dubbing, synchronized video and localized graphics, with AI-Media, CAMB.AI, Chyron and Panjaya each covering a piece.
On the localization side, NDI is using LipSync for real-time translation and lip-synced dubbing inside existing broadcast workflows, generating multiple language versions from one media stream. The LipSync release improves facial occlusion handling and better preserves teeth, lip and facial textures. Active Speaker Detection no longer requires speaker diarization for multiple audio tracks, adds voice activity detection, and expands deployment support through a gRPC interface and broader GPU compatibility.
For a small creator these are less immediately usable than the detector, but the Studio Voice Microphone Profiles are the exception. Suppressing room reverb and then shaping the result into a chosen microphone character is squarely a podcasting and streaming feature, and it is the one item in this bundle aimed at a desk rather than a broadcast facility.
What to Do With This
If you produce AI-generated or heavily processed video, assume it is being scored. Not by a specialist deciding to check, but automatically, at ingest, by infrastructure your distributor already runs. Label generated content yourself before a classifier does it for you with less nuance and no context.
If you build media tooling, the detector, LipSync and Active Speaker Detection NIMs are the pieces you can wire in now rather than evaluate later. If you are testing SVD for a real workflow, test it on compressed footage at the compression levels your pipeline actually produces, and treat the 99.3% figure as an upper bound until you have your own numbers on your own material.
Frequently asked questions
Is the NVIDIA Synthetic Video Detector available to use today?
Yes. SVD is a NIM microservice listed in NVIDIA's model catalogue at build.nvidia.com, alongside the LipSync and Active Speaker Detection NIMs. It was first announced at SIGGRAPH in July 2026, and the IBC announcement covers accuracy improvements and new partner distribution rather than a first release.
How accurate is the detector really?
NVIDIA reports 99.3% on text-to-video content and 97.7% on image-to-video content as of the IBC announcement. In July the company reported up to 92% on uncompressed video, 87% at 15% compression and 82% at 50% compression. The September figures do not state a compression level, so the two sets are not directly comparable and the compression sensitivity disclosed in July is not addressed either way.
Will AI-upscaled or frame-interpolated video get flagged as synthetic?
NVIDIA does not say. VFG, VSR, TrueHDR and LipSync all insert model-generated content into the pixel stream, and SVD scores clips for whether they contain synthetic content. A 6x slow-motion replay built with VFG is an honest record of a real event in which most frames were generated. Whether the detector distinguishes that from a fabricated clip is not addressed in the announcement.
How does this compare with C2PA content credentials?
They solve the same problem from opposite ends. C2PA and similar provenance schemes sign content at capture, which is near-certain when the signature survives but useless once metadata is stripped or the content was never signed. Detection works on any file but returns a probability rather than a proof, and is sensitive to compression and to legitimate processing. Serious verification workflows will use both.
What is the practical difference between the AI for Media SDKs and the NIM microservices?
The SDKs, such as the Video Effects SDK that carries VSR, are libraries you integrate into an application. NIM microservices are containerised inference endpoints you call over the network, deployable on premises, at the edge, in the cloud, in hybrid setups or air-gapped. Several capabilities including VSR ship both ways, so the choice is about where you want the inference to run.
Do the Sports Intelligence Playbooks require NVIDIA hardware?
Effectively yes. The playbooks are built around NVIDIA open models and the NVIDIA fine-tuning and inference stack, including Nemotron, NeMo AutoModel, Megatron Bridge and NIM microservices, and the documentation and repository are published by NVIDIA. The frameworks are open and readable, but the intended path runs on NVIDIA accelerated computing.