NVIDIA announced a major expansion to NVIDIA AI for Media on 9 September 2026, ahead of the IBC conference running 11 to 14 September in Amsterdam. The headline number is that the Synthetic Video Detector (SVD) NIM microservice now reaches 99.3% accuracy on text-to-video content and 97.7% on image-to-video content, and that Wowza will push it into a streaming platform running more than 35,000 video deployments across over 170 countries.

Those accuracy figures are real and they come from NVIDIA. They are also measured along a different axis than the accuracy figures NVIDIA published for the same model seven weeks earlier, and the axis that disappeared is the one that decides whether the detector works on video you actually encounter. That is the story worth reading carefully, and it sits inside a bundle that also ships four separate ways to put synthetic content into video.

What NVIDIA Actually Shipped

AI for Media is NVIDIA's collection of GPU-accelerated SDKs, NIM microservices, playbooks and blueprints for media and entertainment workflows. The IBC expansion touches eight named capabilities plus two frameworks. Several are callable today from the AI for Media model catalogue on build.nvidia.com, which is the practical entry point if you want to test any of this rather than read about it.

CapabilityWhat it doesAdds or detects synthesisNamed partner
Synthetic Video Detector (SVD)Frame-by-frame classifier score for whether a clip contains synthetic contentDetectsWowza, Dalet, TwelveLabs
Video Frame Generation (VFG)Generates new frames between original frames, 2x or 4x frame rateAddsRoss Video (Rio Replay)
LipSync NIMRewrites mouth movement to match a different audio trackAddsNDI
Video Super Resolution (VSR)Upscales video, reduces noise, blur and compression artifacts, adds 10-bit supportAddsAvailable via Video Effects SDK and NIM
TrueHDRConverts SDR to HDR in real time, up to approximately 2,000 nitsAddsCombinable with VSR and VFG
3D Body Pose2D and 3D joint locations from a single camera, no marker rigNeitherVizrt (virtual studios)
Active Speaker Detection NIMIdentifies who is speaking, no diarization required, adds voice activity detectionNeitherContent Localization workflow
Studio Voice Microphone ProfilesSuppresses noise and reverb, then shapes speech into a chosen mic characterAddsBuilt on Studio Voice NIM
Holoscan for Media plus MXLOpen reference architecture for software-defined live production, now with Media Exchange LayerFrameworkDemo at EBU Stand 10.D21
Sports Intelligence PlaybooksFrameworks for fine-tuning NVIDIA open models on proprietary sports footageFrameworkMachina Sports, Wowza

Read the middle column again. In a single announcement NVIDIA shipped one tool that flags video as synthetic and four that synthesize parts of video: VFG invents frames that were never captured, LipSync rewrites the mouth, TrueHDR invents dynamic range the camera never recorded, and VSR invents pixels. NVIDIA even notes that VSR, VFG and TrueHDR "can be combined within a single video-effects pipeline." That is not a contradiction on NVIDIA's part, but it does raise a question the announcement never answers, and I come back to it below.

Grid of matte 3D cubes with one highlighted, representing the NVIDIA AI for Media capability bundle
The IBC expansion touches eight named capabilities plus two frameworks, only one of which detects synthetic content.

The Accuracy Number Changed Axis

SVD was not new at IBC. NVIDIA first announced the Synthetic Video Detector NIM microservice at SIGGRAPH in July 2026. That post reported accuracy this way: "In NVIDIA testing, the model's accuracy reached up to 92% on uncompressed video, 87% at 15% compression and 82% at 50% compression." It also gave latency: 1080p video in as little as 22 milliseconds on NVIDIA RTX systems and approximately 30 milliseconds on L40 GPUs.

The IBC post reports accuracy this way: "Since its initial release, SVD's accuracy has reached 99.3% for text-to-video content and 97.7% for image-to-video content, with especially large gains on difficult image-to-video cases."

Both statements can be true at once, because they are not measuring the same thing. July broke the number down by how degraded the input video was. September breaks it down by how the fake was generated. September says nothing about compression at all.

That matters more than it might sound. Compression is not an edge case in video, it is the default state of video. A clip that reaches a newsroom, a livestream ingest, a social repost or a compliance queue has been compressed, usually several times, usually badly. In July NVIDIA disclosed that its own detector fell from 92% to 82% as compression went from none to 50%. If that sensitivity still exists, a 99.3% figure quoted without a compression level does not tell an editorial team what the model will do on the footage they receive. If it no longer exists, that would be the more impressive claim of the two, and it is not made.

None of this makes the improvement fake. Text-to-video and image-to-video are meaningful categories, image-to-video genuinely is the harder detection problem, and NVIDIA calling out "especially large gains on difficult image-to-video cases" is a specific, falsifiable claim rather than marketing filler. The point is narrower: 99.3% and 92% are not two points on one line, and anyone treating the jump as a straight 7-point improvement is reading a chart that does not exist.

NVIDIA itself is careful here in a way most coverage will not be. The company positions SVD as helping teams "assess the probability of whether footage is authentic or AI-generated," giving them "another point of analysis in their review process." That is a hedge, and it is the correct one. A classifier score is evidence, not a verdict.

Two separate groups of matte 3D bars, showing that the July and September accuracy figures are different measurements
July measured accuracy by compression level. September measures it by generation method. They are not two points on one line.

The Detector Lands in Distribution, Not in the Edit

The most consequential part of this announcement is not the accuracy number, it is where SVD is being installed. All three named integrations put the detector upstream or downstream of the creator, never in the creator's hands.

Dalet is building SVD into a cloud-hosted verification workflow for news organizations, so editorial teams submit footage and review scores and metadata inside a Dalet interface. TwelveLabs announced general availability of Compliance by TwelveLabs, which screens content against regional and custom compliance standards and uses SVD to add frame-level authenticity signals to that same screening pass. And Wowza will distribute SVD through its Video Intelligence Framework, analyzing live feeds in real time for detected objects, scenes and signs of AI generation.

The Wowza number is the one to sit with. Wowza Streaming Engine powers more than 35,000 video deployments in over 170 countries, and the framework can run on premises, at the edge, in the cloud, in hybrid setups or fully air-gapped. Synthetic-content scoring is moving from a forensic tool that a specialist invokes deliberately into a property of the streaming infrastructure itself, applied by default, at ingest, to everything.

If you generate or heavily process video, this changes your position without changing your workflow. You do not opt in and you do not see the score. Someone in a compliance queue does.

Matte 3D nodes linked in sequence with one ringed in orange, representing detection placed inside the distribution pipeline
Synthetic-content scoring is moving into the streaming pipeline itself, applied at ingest rather than invoked by a specialist.

The Question NVIDIA Does Not Answer

Here is the practical gap. NVIDIA now sells, in one product family, tools that add generated frames to real footage and a tool that scores footage for generated content. It does not say whether the first set trips the second.

Take the flagship example from the announcement. Ross Video is integrating VFG into its Rio Replay platform for AI-assisted slow motion in sports production, currently supporting 6x slow-motion generation with development underway toward 8x interpolation. At 6x, the large majority of frames in that replay were generated by a model rather than captured by a camera. The footage is authentic in every sense a viewer cares about: a real match, real players, a real moment. Structurally, though, most of its frames are synthetic.

Does SVD flag it? NVIDIA does not say. The detector is described as producing a classifier score for whether a clip "contains synthetic content," and a 6x interpolated replay unambiguously contains synthetic content while being an honest record of a real event. The same question applies to a LipSync-dubbed interview, a VSR-upscaled archive clip and a TrueHDR conversion of an SDR master. All four are legitimate broadcast operations. All four insert model output into the pixel stream.

This is not a gotcha, it is the central unsolved problem in synthetic-media detection, and it is why the provenance camp exists at all. But it is a live operational question for anyone whose output passes through a Wowza ingest, and the announcement that ships both halves is the natural place to have addressed it.

Detection Versus Provenance, Same Day

The timing here is worth noting. NVIDIA's detection expansion landed the same day Apple presented its own answer to the same question from the opposite direction. Apple's approach, covered in our piece on Reference Image and photo provenance, signs authenticity at the moment of capture. NVIDIA's approach infers it at the moment of inspection.

The two are not competitors so much as complements with different failure modes. Provenance is cryptographic and near-certain when present, but it only covers content captured by participating hardware, and metadata is trivially stripped by any re-encode or repost. Detection works on anything, including a screenshot of a screenshot, but it is probabilistic, degrades with compression, and cannot distinguish "generated" from "legitimately processed." Anthropic's text watermarking approach sits in a third position again, embedding the signal in the output rather than the container.

A newsroom that wants a real answer in 2027 will run all three, and will still occasionally be wrong.

What the Sports Numbers Actually Claim

NVIDIA Sports Intelligence Playbooks are the other substantial release here, and they carry the announcement's most eye-catching statistic. They give leagues and media companies structured frameworks for fine-tuning NVIDIA open models on their own footage and annotations, spanning data preparation, fine-tuning, inference, evaluation, optimization and deployment, and pulling in Nemotron, NeMo AutoModel, Megatron Bridge and NIM microservices. The documentation is public.

The claimed result: multiple-choice accuracy rising from approximately 53% to 94%, and open-ended evaluation from approximately 5.7% to 66%. Those are large jumps. They also come with a qualifier NVIDIA states plainly and most summaries will drop, which is that the evaluation used "previously unseen footage using question formats similar to those used in training."

Unseen footage is a genuine test. Question formats similar to training is a much softer condition, and it is the difference between a model that understands a sport and a model that has learned the shape of the questions asked about it. A 53% to 94% jump on matched question formats is a real and useful engineering result for anyone building a sports analytics product on their own library. It is not evidence of general sports understanding, and NVIDIA does not claim it is.

Wowza is already applying this, fine-tuning vision language models including Cosmos 3 and Nemotron through the playbooks to detect sports-specific moments in live streams. That is the intended shape of the product: not a general sports model, but a pipeline for turning a rights holder's private archive into a narrow model competitors cannot replicate.

Stepped matte 3D platforms representing fine-tuning a general model into a narrow domain model
The playbooks turn a rights holder's private archive into a narrow model competitors cannot replicate.

The Rest of the Stack

Two infrastructure pieces round this out. Holoscan for Media, NVIDIA's open reference architecture and developer toolkit for software-defined live production, now integrates Media Exchange Layer, an open way for software-based media functions to exchange live video, audio and data across a distributed environment. NVIDIA has also brought its Content Localization technologies into that toolkit as a reference workflow covering captions, translated audio, dubbing, synchronized video and localized graphics, with AI-Media, CAMB.AI, Chyron and Panjaya each covering a piece.

On the localization side, NDI is using LipSync for real-time translation and lip-synced dubbing inside existing broadcast workflows, generating multiple language versions from one media stream. The LipSync release improves facial occlusion handling and better preserves teeth, lip and facial textures. Active Speaker Detection no longer requires speaker diarization for multiple audio tracks, adds voice activity detection, and expands deployment support through a gRPC interface and broader GPU compatibility.

For a small creator these are less immediately usable than the detector, but the Studio Voice Microphone Profiles are the exception. Suppressing room reverb and then shaping the result into a chosen microphone character is squarely a podcasting and streaming feature, and it is the one item in this bundle aimed at a desk rather than a broadcast facility.

What to Do With This

If you produce AI-generated or heavily processed video, assume it is being scored. Not by a specialist deciding to check, but automatically, at ingest, by infrastructure your distributor already runs. Label generated content yourself before a classifier does it for you with less nuance and no context.

If you build media tooling, the detector, LipSync and Active Speaker Detection NIMs are the pieces you can wire in now rather than evaluate later. If you are testing SVD for a real workflow, test it on compressed footage at the compression levels your pipeline actually produces, and treat the 99.3% figure as an upper bound until you have your own numbers on your own material.

Frequently asked questions

Is the NVIDIA Synthetic Video Detector available to use today?

Yes. SVD is a NIM microservice listed in NVIDIA's model catalogue at build.nvidia.com, alongside the LipSync and Active Speaker Detection NIMs. It was first announced at SIGGRAPH in July 2026, and the IBC announcement covers accuracy improvements and new partner distribution rather than a first release.

How accurate is the detector really?

NVIDIA reports 99.3% on text-to-video content and 97.7% on image-to-video content as of the IBC announcement. In July the company reported up to 92% on uncompressed video, 87% at 15% compression and 82% at 50% compression. The September figures do not state a compression level, so the two sets are not directly comparable and the compression sensitivity disclosed in July is not addressed either way.

Will AI-upscaled or frame-interpolated video get flagged as synthetic?

NVIDIA does not say. VFG, VSR, TrueHDR and LipSync all insert model-generated content into the pixel stream, and SVD scores clips for whether they contain synthetic content. A 6x slow-motion replay built with VFG is an honest record of a real event in which most frames were generated. Whether the detector distinguishes that from a fabricated clip is not addressed in the announcement.

How does this compare with C2PA content credentials?

They solve the same problem from opposite ends. C2PA and similar provenance schemes sign content at capture, which is near-certain when the signature survives but useless once metadata is stripped or the content was never signed. Detection works on any file but returns a probability rather than a proof, and is sensitive to compression and to legitimate processing. Serious verification workflows will use both.

What is the practical difference between the AI for Media SDKs and the NIM microservices?

The SDKs, such as the Video Effects SDK that carries VSR, are libraries you integrate into an application. NIM microservices are containerised inference endpoints you call over the network, deployable on premises, at the edge, in the cloud, in hybrid setups or air-gapped. Several capabilities including VSR ship both ways, so the choice is about where you want the inference to run.

Do the Sports Intelligence Playbooks require NVIDIA hardware?

Effectively yes. The playbooks are built around NVIDIA open models and the NVIDIA fine-tuning and inference stack, including Nemotron, NeMo AutoModel, Megatron Bridge and NIM microservices, and the documentation and repository are published by NVIDIA. The frameworks are open and readable, but the intended path runs on NVIDIA accelerated computing.