Tutorials Video

Using Pro Tools and Media Composer Speech-to-Text Together

Advanced · ~25 min

Overview

Avid put speech-to-text into both halves of its post pipeline: Media Composer gained PhraseFind AI and ScriptSync AI, and Pro Tools gained a Speech-to-Text engine of its own. Both run entirely on your machine with no cloud involvement. The thing nobody tells you up front is that they do not talk to each other, neither can read the other's transcript, in either direction. Once you know that, the combined workflow becomes obvious and genuinely fast. This guide covers what each feature actually does, how they fit either side of the AAF handoff, and how to avoid planning around an integration that does not exist.

What You Need

  • Media Composer with the PhraseFind AI and/or ScriptSync AI options licensed. They are separate paid add-ons, not bundled
  • Pro Tools Studio or Ultimate at 2025.6 or later. The Speech-to-Text engine is a separate installer download
  • Enough local CPU headroom, since all transcription runs on your own machine
  • Dialogue recorded well enough to transcribe. This is the real prerequisite
  • A settled AAF handoff routine between picture and sound
  • Realistic expectations about accuracy on overlapping speech and heavy accents

Steps

1

Transcribe in Media Composer as an ingest step, not an editorial one

Run PhraseFind AI indexing across the rushes when the media lands, before anyone starts cutting. It is a background cost that pays back the moment an editor needs to find a line, and doing it at ingest means the assistant absorbs the processing time rather than the editor waiting mid-session.

2

Use PhraseFind AI to find dialogue, ScriptSync AI to structure it

They solve different problems. PhraseFind AI is search: type a phrase, get every clip where it was spoken, and start editing directly from the result. ScriptSync AI is alignment: it matches takes against a script in the Script window so you can work through coverage line by line. Documentary and interview work leans on the first. Scripted work leans on the second.

3

Cut picture as normal. The transcript is a finding aid, not the edit

Both features accelerate locating material. Neither changes how the timeline is assembled. Treat the transcript as a way to stop scrubbing, then edit on performance and rhythm as you always would. Editors who try to cut from the text alone produce sequences that read correctly and play flat.

4

Hand off to Pro Tools by AAF, expecting no transcript to travel with it

Export the AAF as usual. It carries media references, edits, levels, and pan information. It does not carry Media Composer's transcript database, and Pro Tools cannot import transcription data from another application. Assume the sound side starts with no text and plan the schedule accordingly.

5

Re-transcribe on the Pro Tools side, and let it run while you set up

Analyse the dialogue tracks in Pro Tools to generate its own transcript. This is the step people resent until they see what it buys: text sitting in a track view alongside the waveform, staying attached to clips as they are trimmed and moved, plus a session- wide Transcript window you can search. Start it going while you build routing and templates.

6

Use the Pro Tools transcript for the jobs that are actually hard in audio post

Finding one line across forty tracks of production sound. Locating every instance of a word that needs replacing. Checking dialogue against a script for ADR spotting. Navigating a long interview without scrubbing. These are the wins, not re-editing the story, which is settled by the time the AAF arrives.

7

Keep dialogue in WAV or AIFF if the session will move between machines

Pro Tools stores transcription data inside WAV and AIFF files themselves, so the text travels with the audio. Other formats such as MXF rely on a session-side database instead, which does not follow the media if someone hands off only the files. On a multi-editor or multi-facility job that difference decides whether the transcript survives the move.

Pro Tips

  • Both engines run locally with no internet requirement, which is why they are usable on confidential material where a cloud transcription service would not be permitted.
  • Transcription quality is governed by the recording. Close, clean, one-speaker-at-a-time dialogue transcribes well. A crowded room with overlapping speech defeats both engines equally.
  • PhraseFind AI now works from real transcription rather than the older phonetic matching, so results are text you can read rather than fuzzy sound-alike hits.
  • Budget the transcription pass into the schedule as an ingest task on the picture side and a session-prep task on the sound side. It is not instant on long-form material.
  • You cannot currently adjust word timing in the Pro Tools transcript, so treat word boundaries as approximate when using them to place edits.

What You'll Learn

Below: what each of the three features actually does, what the AAF handoff does and does not carry, where the transcript data physically lives on each side, and the limitations worth knowing before you build a schedule around any of this.

Three features, three different jobs

FeatureApplicationWhat it doesBest for
PhraseFind AIMedia ComposerIndexes dialogue across media and finds every clip where a phrase was spoken, returning readable text you can edit straight fromDocumentary, interview, reality — anywhere the material is unscripted and large
ScriptSync AIMedia ComposerAligns takes against a script in the Script window so coverage can be worked line by lineScripted drama and comedy, and anything with a shooting script to work against
Speech-to-TextPro ToolsTranscribes clips into a track view beside the waveform, keeps text attached as clips move, and adds a session-wide searchable Transcript windowDialogue editing, ADR spotting, and navigating long-form sessions

PhraseFind AI and ScriptSync AI are separate paid options for Media Composer rather than included features. Avid kept the same licensing and activation model the pre-AI versions used. Pro Tools Speech-to-Text is available to Studio and Ultimate customers as a separate installer, from version 2025.6 onward. Earlier versions cannot display transcriptions at all.

The handoff: what AAF carries, and what it does not

This is the fact that changes how you plan, so it is worth stating plainly.

An AAF from Media Composer to Pro Tools carries the things it has always carried: media references, the edit itself, clip boundaries, levels, pans, and enough structure to rebuild the sequence on the sound side. It does not carry the Media Composer transcript database.

Nor is there a back door. Avid documents that Media Composer's auto-created transcripts cannot currently be exported, and separately that Pro Tools cannot import transcription data from other applications. Both halves of the round trip are closed. Whether that changes in a future release is Avid's call. Today it does not work.

So "using them together" does not mean sharing a transcript. It means running two independent local transcription passes over the same dialogue, on either side of a handoff that was never designed to carry text, and accepting the second pass as a cost of the workflow rather than a mistake you are making.

The practical consequence is scheduling. If a sound editor expects text to arrive with the AAF, they will discover otherwise at exactly the wrong moment. Build the Pro Tools transcription pass into session prep the same way you build in conform checking.

Where the transcript actually lives, and why it matters on a shared job

The two applications store their text differently, and the Pro Tools behaviour has a real operational consequence.

Pro Tools embeds transcription data inside WAV and AIFF files themselves. The text is part of the audio file, so it travels wherever the file goes, another machine, another editor, another facility. For other formats, notably MXF, the transcript lives in a session-side database instead.

That distinction decides whether your transcript survives a handoff. Send someone a folder of WAVs and the text goes with them. Send MXF media without the session and the text stays behind. On a job where dialogue is passed between editors, that is a reason to prefer WAV or AIFF for anything you expect to be transcribed.

Multi-channel and field recorder files are supported with channel-specific options, which matters for production sound where each character is on their own iso track and you want the transcript per channel rather than mashed together.

On the Media Composer side, transcripts are held in a management tool within the application. More recent releases added editable transcripts with word-level timing retained, shared transcript databases across a network so a cutting room can index once and share, and support for proxy workflows.

Both run locally, which is the underrated part

Avid states explicitly that both engines are local pre-trained models: Pro Tools Speech-to-Text "does not require internet access" and "any files you analyse are kept entirely on your local machine," and the Media Composer AI features are "not cloud-based solutions. All processing is local to your system."

This is more consequential than it sounds. A large amount of professional dialogue material cannot lawfully or contractually be uploaded to a third-party transcription service, unaired drama, embargoed documentary interviews, anything under an NDA, anything involving vulnerable contributors. Local processing removes that objection entirely, which is frequently the difference between being able to use transcription on a job and not.

The trade is that transcription competes for the same CPU as everything else you are doing. On long-form material it is not instantaneous, which is the argument for running it as an ingest or prep task rather than on demand mid-session.

Both sides support automatic detection across roughly twenty-one languages, so mixed-language material does not need to be separated by hand before indexing.

Limitations worth knowing before you plan around it

No transcript interchange, in either direction. Covered above, and the most important one. Do not design a pipeline that assumes text moves with the AAF.

No third-party transcript import into Pro Tools. If you already pay for a transcription service and have accurate, human-corrected text, you cannot currently bring it in. The Pro Tools transcript has to be generated by Pro Tools.

Word timing is not editable in Pro Tools. You can correct what a word says, but not where the engine decided it started and ended. Treat word boundaries as approximate when using the transcript to navigate to an edit point rather than as frame-accurate markers.

Accuracy tracks recording quality, not licence tier. Overlapping dialogue, heavy background, distant microphones, and strong accents degrade both engines. No amount of processing recovers speech that was not clearly captured, which is the same conclusion reached in our guide to how speech recognition actually works.

Licensing is not uniform. The Media Composer features are paid add-ons. The Pro Tools feature is limited to Studio and Ultimate. A facility can easily end up with the capability on one side of the pipeline and not the other, which is worth checking before promising a workflow to a client.

Where this fits against the wider text-based editing picture

Avid is not alone here, transcript-driven editing has become a standard feature across the major applications, and the general technique is covered in our guide to text-based video editing.

What distinguishes the Avid implementation is the combination of local processing and presence on both sides of a traditional broadcast pipeline. Most transcript-editing tools are cloud services attached to a single NLE. Having the same capability in the picture-cutting application and the audio post application, both running offline, suits exactly the kind of long-form, confidential, deadline-bound work Avid pipelines already dominate.

The honest limitation is that it is two capabilities rather than one workflow. Until transcript interchange exists, the integration is conceptual. The same job done twice with two tools, connected by a handoff that does not know the text exists. That is still a large net gain over scrubbing, and it is worth understanding as it actually is rather than as the marketing implies.

For the disciplined side of working with AI-assisted post more generally (what to verify, what to document, and where human judgement has to stay) see our AI-assisted editing workflow.

Where This Fits

This guide covers one specific part of video production. The wider picture. The full arc from planning through camera, light, sound, edit, and delivery, and how each stage constrains the next, is in Video Production Fundamentals: The Complete Guide, which frames the discipline as a whole and links out to the detailed guides underneath it, including this one. If you are starting from scratch rather than solving a specific problem, read that first and come back here.

FAQ

Q: Can Media Composer send its transcript to Pro Tools?
A: No. Avid documents that Media Composer's auto-created transcripts cannot currently be exported, and separately that Pro Tools cannot import transcription data from other applications. The AAF handoff carries media, edits, levels, and pans, not text. You run a second transcription pass in Pro Tools.

Q: Which versions and tiers do I need?
A: Pro Tools Speech-to-Text requires Pro Tools Studio or Ultimate at version 2025.6 or later, installed as a separate download. Earlier versions cannot display transcriptions. On the Media Composer side, PhraseFind AI and ScriptSync AI are separate paid options rather than bundled features, using the same licensing model as the pre-AI versions.

Q: Does any of this upload my audio to the cloud?
A: No. Both are local pre-trained models. Avid states that the Pro Tools engine does not require internet access and that analysed files stay entirely on your machine, and that the Media Composer AI features are not cloud-based. That is what makes them usable on material under NDA or embargo, where a cloud service would not be permitted.

Q: What is the difference between PhraseFind AI and ScriptSync AI?
A: PhraseFind AI is search, find every clip where a phrase was spoken and edit from the results, which suits unscripted material. ScriptSync AI is alignment, match takes against a script in the Script window so you can work through coverage line by line, which suits scripted production. They are separate options solving different problems.

Q: Why does file format matter for the Pro Tools transcript?
A: Because WAV and AIFF store the transcription data inside the audio file itself, so the text travels with the media to any machine. Other formats such as MXF keep it in a session-side database instead, which does not follow the files on their own. If dialogue will be handed between editors or facilities, prefer WAV or AIFF.

Q: Can I import a transcript I already paid a service to produce?
A: Not into Pro Tools currently. Import of transcription data from other applications is not supported. If you have accurate human-corrected text from a transcription vendor, it can still inform the edit as a reference document, but it cannot populate the in-application transcript.

Q: Is it accurate enough to rely on?
A: For finding material, yes, and that is what it is for. Accuracy tracks recording quality rather than licence tier: clean, close, one-speaker-at-a-time dialogue transcribes well while overlapping speech and distant microphones degrade it. Treat it as a fast way to stop scrubbing, not as a verbatim record.

Translate this page

Machine translation provided by Google Translate, on Google’s servers. We do not check these translations and they will get technical terms wrong. The English page is the authoritative one. Following a link sends this page’s address to Google. Your browser may also offer to translate this page itself, which keeps the request on your device.