Documentary & AI
The AI Vérité Workflow
Most conversations about generative video start from the wrong end: the prompt. You type a description, you get a plausible moving image, and the only question left is whether it looks convincing. An AI vérité workflow inverts that. It starts from something that actually happened (handheld footage, unscripted conversation, real room noise) and uses generative tools only to show what the camera could not: a memory, an interior state, an event that predates the recording. The real material is the anchor. Everything synthetic hangs off it.
The short version: AI vérité treats unscripted audio and video as the source of truth, mines the transcript for what the subject is actually saying beneath the words, and generates imagery for the parts of their experience no camera was present for. The technique is genuinely useful for the archival gap and for interior states. Its whole burden is keeping the seam between recorded and generated legible to the viewer, and that is a disclosure problem, not a technical one.
Where the name comes from, and why it matters
Cinéma vérité emerged when equipment got light enough to follow people around: portable cameras, sync sound, small crews, no lighting setup, no script. The aesthetic (handheld, available light, long takes, the crew occasionally audible) was a by-product of the method rather than a style someone designed. The claim underneath it was epistemic. This happened, we were there, here is the tape.
That claim is the thing generative video does not have and cannot manufacture. A model produces an image that resembles the world without having been pointed at any part of it. The word for what documentary footage has and synthetic footage lacks is indexicality. The physical causal link between the thing and its recording. Light bounced off a real face and hit a sensor.
So pairing the two words is deliberate and slightly provocative. An AI vérité workflow does not pretend the generated material carries the same authority as the recorded material. It uses the recorded material to earn authority for the piece as a whole, and then spends some of that authority on sequences that are openly interpretive. Whether that trade is honest depends entirely on how clearly the viewer can tell which is which. A point worth holding onto through everything below.
Stage one: the groundwork layer
Everything downstream depends on capture, and the capture is deliberately ordinary documentary practice.
Run-and-gun shooting. Handheld camera, available light, unscripted interaction. Small enough that the subject stops performing. The aesthetic matters here for a second reason beyond authenticity: a handheld frame with real motion, real grain, and imperfect exposure gives you a rich set of characteristics to match against later. Immaculate tripod-locked footage is harder to blend with generated material, because generated material rarely looks that clean and any mismatch reads immediately.
High-fidelity audio as the source of truth. This is the load-bearing decision. Natural dialogue, ambient soundscape, spontaneous interview (captured properly, close, monitored on headphones) becomes the spine that drives the entire edit. Everything else is derived from it or answerable to it.
Audio earns that role for a practical reason. It carries the words, which are what a language model can analyse. It carries prosody, which is where emotion actually lives. And it is far cheaper to capture well than picture. A workflow that treats audio as the primary asset also inherits the ordinary discipline of good field recording (get the microphone close, control the room, record ambience separately) and those habits pay twice here, because the ambient bed later becomes the reference for any synthetic sound.
Stage two: semantic analysis and story extraction
Between capture and generation sits the stage that distinguishes this workflow from prompt-driven video, and it is mostly reading.
Transcription. Field audio goes through a speech-to-text model. The output is a working draft, not a record: proper nouns, technical vocabulary, and sentence boundaries are exactly where recognition fails, and those are the high-information words. Correcting the transcript against the audio is not optional if anything downstream will depend on it, and the corrected transcript is worth publishing alongside the finished piece for the same reasons it is worth making.
Subtext mining. A language model reads the corrected transcript looking for emotional beats, recurring metaphors, implicit themes, and, most usefully, the concrete images buried in how people actually talk. Subjects describe their inner lives in pictures constantly and without noticing. Someone says the house felt like it was holding its breath. Someone describes a job as being underwater. Those are visual prompts the subject wrote themselves, and they are far better source material than anything a director invents, because they came from the person the film is about.
This is the step where the workflow is doing something other than automation. The model is not deciding what the film means. It is surfacing candidates from a transcript longer than anyone wants to read four times. The selection stays human, and it should. An LLM will confidently identify a theme in material that has none, and a director who accepts that reading is now making a film about a pattern nobody said.
Stage three: generative synthesis
Only now does anything get generated, and it is generated against extracted material rather than invented from a blank prompt.
Prompt-driven world-building. Video and image models produce sequences from the themes and images mined in stage two. The memory being described, the metaphor being used, the place that no longer exists. The prompts are downstream of the subject's own words, which is what keeps the imagery tethered to the film rather than merely decorative. Current tooling in this space is moving quickly enough that naming products dates fast. The trajectory of generative video matters more than this month's leader.
Sound augmentation. Generative audio reconstructs environments that were never recorded. A factory floor that closed in 1988, a room as someone remembers it rather than as it sounds now. The constraint that keeps this from feeling pasted-on is cadence: synthetic sound should sit inside the acoustic character of the real recording, matching its noise floor, its reverberation, its general roughness. A pristine synthetic ambience under a recording made in a live room announces itself instantly.
Synthetic voice deserves separate treatment and more caution than the rest of this stage. Reconstructing a voice raises questions of consent that are independent of any platform policy, and doing it for someone who has not agreed, including someone who has died, is a decision to make deliberately and disclose plainly. Our guide on consent and disclosure for AI voice covers the ground. The summary is that technical possibility settles nothing here.
Stage four: integration and style transfer
Generated sequences that look like generated sequences sitting next to documentary footage produce a film with two textures and no relationship between them. This stage is about the join.
Matching the camera. Control networks and style-transfer models carry characteristics from the real footage into the synthetic material: camera motion, grain structure, depth cues, colour and lighting behaviour. The aim is not to make generated shots indistinguishable from recorded ones. It is to make them feel as though they belong to the same film, which is a lower and more achievable bar, and an ethically safer one.
The craft problem here is identical to a much older one: making two cameras cut together. The same discipline applies, match the black point, then the white point, then midtones, then colour balance, and judge it on scopes rather than by eye after an hour in a dark room. Our guide to matching shots from different cameras is directly applicable, because a generative model is, for grading purposes, just another source with its own colour behaviour.
Assembly. The finished timeline moves between observed reality and interpretive sequence. How hard that transition should land is the central creative decision of the form. A hard cut asserts the difference. A slow dissolve blurs it. Both are legitimate. Only one of them is honest by default, and the other requires the film to make its method clear somewhere else.
Why filmmakers are reaching for it
The archival gap. Documentary has always struggled with events that happened before anyone was filming, or where filming was impossible. The traditional answers are dramatised reenactment, slow pans across photographs, or an empty landscape with narration over it. All three are conventions audiences have long since learned to read as "we do not have footage of this." Generated imagery is a fourth option, and for some subjects a better one.
Emotional subjectivity. Observational documentary is very good at what a person does and comparatively poor at what they are experiencing. It can show a face and let you infer. AI vérité lets the film put the interior state on screen next to the exterior one, not as an assertion of fact, but as an illustration of what the subject just described.
Reach for small teams. The world-building that used to require a visual effects budget is now available to one person with a laptop. That is a genuine democratisation, and it is also why the next few years will produce a great deal of work in this form, much of it bad. Access to a technique arrives long before fluency in it.
The problems worth naming
An honest account of this workflow has to sit with what it costs, because the failure modes are not hypothetical.
The seam is the whole ethical question. Documentary's authority is borrowed from the fact that a camera was present. Generated sequences borrow that authority without earning it. If a viewer cannot tell which is which, the film is trading on a claim it is not entitled to make, regardless of intent. This is why disclosure is not a compliance checkbox bolted on at delivery but a design constraint that shapes the edit. Our guides on when to label content as AI-made and on content credentials and provenance cover the practical mechanics.
The pull toward generating rather than reporting. Generation is fast, cheap, and always available. Getting the interview, returning to the location, finding the archive, and earning a subject's trust are none of those things. There is a real gravitational pull toward filling gaps with synthesis instead of doing the work, and the resulting film will look finished while being about less than it claims.
Consent extends past the recording. A subject who agreed to be filmed has not thereby agreed to have their childhood home, their memories, or their inner life visualised by a model and shown to an audience. That is a separate conversation, held before generation, in terms they can actually evaluate. The same care that applies to interviewing someone after a traumatic event applies here, and for the same reason: consent given without understanding what will be made is not meaningful consent.
A house style is already forming. The drifting, painterly, slightly liquid dreamscape is becoming this form's equivalent of the slow pan across a photograph. A convention that reads less as a creative choice and more as a signal that the budget ran out. Anything that becomes instantly recognisable as "the AI bit" has stopped doing the job it was introduced to do.
It does not survive contact with journalism. Everything above assumes an interpretive documentary. In reporting, where the audience is being told what is the case, the calculus is different and much stricter. The standards in our guides on verifying found video and interviewing for journalism do not bend to accommodate generated imagery.
If you want to try it
Four things separate work in this form that holds up from work that does not.
Capture more than you think you need, and capture audio properly. The whole method is downstream of the recording. A thin, poorly recorded interview yields a thin transcript, which yields generic extracted themes, which yields generic imagery.
Let the subject's language drive the prompts. The best generated sequences in this form are illustrations of something a person actually said. The worst are a director's idea of what the film should look like, laundered through the subject.
Decide your disclosure posture before you edit, not after. Whether the seam is hard or soft, whether there is an on-screen indication, whether the method is explained in the film or only in the description. These change the edit. Deciding at delivery means retrofitting honesty onto a structure built without it.
Keep a human decision layer, and document it. Note what was generated, from which prompt, derived from which passage of transcript. Six months later, when someone asks whether a sequence represents something the subject said or something the film inferred, that record is the only thing that can answer. The same discipline described in our AI-assisted editing workflow applies with more weight here, because the stakes are representational rather than merely technical.
The interesting thing about AI vérité is not that it makes generated video respectable by attaching it to real footage. It is that it inverts the default relationship between the two. The recording leads, the model serves, and the film is still accountable to something that happened. That is a meaningfully different proposition from prompt-first video, and it is worth defending on those terms rather than on the novelty of the images.
Read more
This article develops an original analysis of an emerging hybrid documentary practice. The techniques described are in active development across the field. Tool names date quickly, so the guides below focus on the underlying craft and the disclosure questions rather than on specific products.
- A Practical AI-Assisted Edit: From Raw Footage to Rough Cut
- When to Label Content as AI-Made
- Content Credentials (C2PA): Proving Your Footage Is Real
- AI Voice Cloning: Consent, Contracts, and Disclosure
- Documentary Storytelling Structure
- Field Recording Basics: Gear, Wind, and Technique
- Matching Shots From Different Cameras
- The Future of AI Video Generation