
Guides
Synthetic media verification: a working guide for reporters
A working guide to synthetic media verification for reporters, from preserving the original file to responsibly bounding what you can honestly publish.
What to take away
- Preserve the best available file, the post URL and the upload time before you test anything.
- Write the claim as testable fieldswho, what, where, when, and which event it is tied to.
- Keep provenance records, forensic observations, detector scores and source testimony in separate lanes.
- Do not call a file synthetic because it lacks metadata or Content Credentials.
- Publish what the evidence establishes, what it suggests, and what stays unresolved.
Synthetic media is a family of things, not one thing. It covers several kinds of synthetic media:
- generated images
- cloned voices
- face swaps
- lip-synced video
- altered backgrounds
- composites stitched from real and generated parts
That range is why "is this AI?" is the wrong first question. A newsroom has to establish what file it holds, where the file came from, what it claims to show, and whether the evidence supports that claim.
Define the claim before inspecting pixels
Write the claim as testable fields:
Claim fields to define
- Speaker or subjectwho appears?
- Actionwhat was said or done?
- Placewhere did it happen?
- Timerecorded and posted when?
- Eventwhat real-world event surrounds it?
- File historyoriginal, export, screen recording, repost?
| Field | Question |
|---|---|
| Speaker or subject | Who is said to appear or speak? |
| Action | What are they said to have done or said? |
| Place | Where did it happen? |
| Time | When was it recorded, and when was it posted? |
| Event | What real-world event supposedly surrounds it? |
| File history | Is this an original, an export, a screen recording, or a repost? |
Separate the findings these fields produce. A clip can be authentic and falsely dated. A real voice recording can be laid over unrelated images. A generated picture can illustrate a true event while being passed off as documentary evidence.
Each of those is a different verdict, and each needs different wording in print.
Preserve the evidence package
Save the highest-quality file you can get without re-encoding it. Record the URL, the account name and identifier, the post text, and visible engagement. Also record upload time with its time zone, and your download time.
Evidence package to preserve
- Save highest-quality file without re-encoding
- Record URL, account name and identifier
- Record post text and visible engagement
- Record upload time with time zone
- Capture whole page, not just media
- Hash every file for later comparison
- Read detector retention and privacy terms
Capture the whole page, not just the media. Note any edits or platform labels applied since publication. If the publisher hands over an original through a secure channel, keep that copy separate and write down how it reached you.
Hash every file so later copies can be compared. Keep working copies apart from preserved originals. Before uploading sensitive unpublished material to a public detector, read that tool's retention and privacy terms.
Build four evidence lanes
1. Source and distribution
Find the earliest known appearance. Trace reposts backward, compare captions, and ask whether the first visible account had access to the claimed scene or person.
Four evidence lanes
- Source and distribution: earliest appearance, reposts, captions
- Event and context: schedules, weather, daylight, architecture
- Provenance and file history: metadata, signed records
- Media forensics: continuity, lighting, reflections, alignment
Contact that account through an independently confirmed channel. Ask for the original file, the capture device, an approximate time, and the circumstances. Do not reveal every detail you expect to hear back.
2. Event and context
Check schedules, weather, daylight, and architecture. Check language, clothing, official statements, and live streams. Check other independent recordings. Context often settles a claim without deciding how the pixels were made.
A person shown somewhere else at the claimed time disproves the caption even if the production method stays unknown.
3. Provenance and file history
Inspect embedded metadata and any signed provenance record. Ask whether the credential validates, who signed it, what actions it describes, and whether it belongs to this exact file.
Missing metadata is ordinary after platform processing. It is not evidence of fabrication.
4. Media forensics
Examine these for continuity:
- frame continuity
- lighting
- reflections
- perspective
- lip and sound alignment
- background noise
- compression
- abrupt edits
- repeated regions Compare suspicious details across adjacent frames and against known authentic material.
A single odd hand, shadow, or waveform is a lead, not a verdict. When several such leads point the same way, you have something worth reporting. When they conflict, say so.
The National Institute of Standards and Technology's report on technical approaches to synthetic content transparency separates provenance, labeling and watermarking, detection, testing, and auditing. No single method answers every verification question.
Use detectors as measured instruments
Before you rely on a detector, record five things. Record its version, the media types it supports, and its stated training scope. Also record what its output actually means and its known error rates. Run the best available file, not a screenshot, when the original exists.
Record before trusting a detector
- Detector version
- Supported media types
- Stated training scope
- What the output actually means
- Known error rates
Compare tools that use different methods. Do not treat several opaque scores as independent confirmation of each other.
A score reading "82 percent AI" may describe a model output, not an 82 percent probability that the claim is false. Cropping, compression, subtitles, noise removal, and platform transcoding all move the number. Test an authentic comparison file from the same device or distribution path where you can.
The field measures itself this way. The NIST Open Media Forensics Challenge scores whether a system can decide an image or video was manipulated, then whether it locates the altered region and names the manipulation type, using benchmark datasets on shared evaluation infrastructure.
A tool worth quoting has been measured on something, and the measurement has conditions attached.
Reach a bounded finding
Useful conclusions include:
Choosing a bounded finding
Is the original available?
verify against the stated event
unverified because source unavailable
- verified authentic for the stated event;
- authentic media used with false context;
- edited, with the consequential alteration identified;
- synthetic or partly synthetic, supported by several independent signals;
- unverified because the source or original is unavailable; or
- inconclusive because credible signals conflict.
Name the decisive evidence. Avoid "looks fake," "an AI detector confirmed it," and "no metadata proves it was generated." Where a finding could harm an identifiable person, seek a second review and give that person a fair chance to respond.
Publish a verification note
Set out the exact claim. Set out the material examined. State whether you had the original or a repost. List the tools and versions. Give independent context. Give source contact. Give limitations. Give your evidence cutoff.
State your revision policy. Keep detector output in the reporting file, but give readers the reasoning rather than a dashboard of unexplained numbers. When a claim survives to this stage, the remaining disputes usually follow a handful of familiar patterns, and it helps to know common visual verification problems before you publish.
Common questions
Is every deepfake fully generated?
No. A deepfake can alter part of authentic media, such as a face or a voice, while leaving the rest untouched. That distinction changes what you can safely write.
Does missing metadata prove a file is fake?
No. Social platforms, editing software, messaging services, and screenshots routinely strip or rewrite metadata. Treat its absence as a gap in the record, not as a finding.
Can one detector settle the question?
No. Treat a detector score as one reproducible signal and test it against provenance, context, source access, and forensic observations. Where those disagree, report the disagreement.
What if the source will not provide the original?
Report that limitation plainly. You may still verify context, but your conclusion should not imply access to evidence you never examined.







