Video Editing
How to Edit Video With OpenAI Codex: Setup, Word Timing and MCP
What is specific to Codex: its install and plans, timing edits to words, checking frames with image input, scripted runs, and MCP settings. Every ffmpeg command below is one we ran.
Key takeaways
- Codex edits video by running tools on your computer: ffmpeg, Whisper for word timings, or an editor's MCP server.
- The Codex CLI is listed from the ChatGPT Plus plan up, or with an API key at API rates. Free and Go list only the desktop app, subject to rollout.
- Word-timed edits work from a transcript JSON. In our test a title landed on its word to the frame.
- Put house rules in AGENTS.md and show Codex the result with -i frames.
- Codex stops waiting for an MCP tool after 60 seconds by default; raise tool_timeout_sec for editors.
1Can OpenAI Codex edit videos?
Yes. Codex is a coding agent that runs commands on your computer, so it edits video by running ffmpeg, writing code for a renderer, or calling a video editor's tools over MCP. It does not watch video. It reads transcripts, command output and still frames, so every reliable Codex workflow turns sound into timed words and pictures into images it can be shown.
Our overview of AI agents editing video compares the routes, and our Claude Code guide walks through a full ffmpeg pipeline whose commands apply unchanged. This guide covers what differs with Codex, plus one technique Codex creators share on r/codex: timing edits to individual words.
2What do you need to edit video with Codex?
The Codex CLI, a ChatGPT plan that includes it (or an API key), and ffmpeg. OpenAI's Codex CLI page gives four install methods (checked October 2026):
# macOS, Linux curl -fsSL https://chatgpt.com/codex/install.sh | sh # Windows, in a new PowerShell window powershell -ExecutionPolicy ByPass -c "irm https://chatgpt.com/codex/install.ps1 | iex" # or with a package manager npm install -g @openai/codex brew install --cask codex codex login # browser sign-in with your ChatGPT plan codex login status # shows which sign-in method is active
OpenAI's authentication docs say you sign in with ChatGPT for subscription access or with an API key, in which case Codex uses standard API pricing instead of plan credits.
Which plan? The Codex pricing page lists “Codex on the web, in the CLI, in the IDE extension, and on iOS” starting with Plus. Pro includes everything in Plus, and the Business and Enterprise cards include Codex too. Free and Go list Codex in the desktop app, “subject to rollout” (checked October 2026). Check that page for current prices.
Then install ffmpeg: brew install ffmpeg on a Mac, and the Windows package is in our Claude Code guide. We ran the commands below on ffmpeg 9.0.2 from Homebrew and checked the Codex flags against Codex CLI 0.153.3. That ffmpeg build lacks the drawtext and subtitles filters, so our title below is an image overlay.
3How do you sync video edits to spoken words with Codex?
Make a JSON transcript with a start and end time for every word, then tell Codex what to do at which word.
The creator of a two-part “I Edited This Video 100% With Codex” series on r/codex describes this: a transcript JSON with word-level timestamps, then instructions tied to words, rendered with Remotion. We tested the idea with plain ffmpeg.
Our test clip: we recorded one sentence with macOS's say command (“Most videos never get finished. Today we launch a faster way to cut them.”) and put it under a 1920x1080, 30 fps test pattern: 4.49 s in total. The goal: show a “LAUNCH DAY” title from the word “launch” onward.
Step 1: word timestamps
# 1. Audio for Whisper: mono, 16 kHz ffmpeg -i talk.mp4 -vn -ac 1 -ar 16000 talk.wav
For the transcript we used faster-whisper, an MIT-licensed reimplementation of OpenAI's Whisper that runs on your computer, with the base model and its word_timestamps option. (We read the WAV with numpy because our Python audio decoder failed; the README passes a file path.)
# words.py (pip install faster-whisper)
import json, wave
import numpy as np
from faster_whisper import WhisperModel
with wave.open("talk.wav") as f:
audio = np.frombuffer(f.readframes(f.getnframes()), np.int16) / 32768
model = WhisperModel("base", device="cpu", compute_type="int8")
segments, _ = model.transcribe(audio.astype(np.float32), language="en", word_timestamps=True)
words = [{"word": w.word.strip(), "start": round(w.start, 2), "end": round(w.end, 2)}
for s in segments for w in s.words]
json.dump({"words": words}, open("transcript.json", "w"), indent=1)Part of what it wrote, 14 words in all:
{"word": "finished,", "start": 1.18, "end": 1.74}
{"word": "today", "start": 2.02, "end": 2.24}
{"word": "launch", "start": 2.48, "end": 2.72}
...
{"word": "them.", "start": 3.88, "end": 4.12}The cloud alternative: OpenAI's speech-to-text guide says to use whisper-1 with timestamp_granularities[] when you need word timestamps “for captioning and video editing” (checked October 2026). That uploads the audio to OpenAI and bills your API key.
Step 2: check the timings against the audio
# 2. Check the word edges against the audio itself ffmpeg -hide_banner -nostats -i talk.wav -af silencedetect=noise=-35dB:d=0.1 -f null - # silence_start: 1.722438 # silence_end: 2.061063 # silence_start: 4.275437
The pause between the sentences measured 1.72 to 2.06 s. Whisper put the end of “finished” at 1.74 and the start of “today” at 2.02: both within about 0.04 s of the measured edges. The last word was looser. Whisper ended “them” at 4.12, but the sound ran to 4.28, so a cut placed on that word's end would have clipped 0.16 s of speech (our arithmetic). Whisper also heard the full stop after “finished” as a comma, which does not matter for timing but would for captions.
Step 3: the prompt and the edit
The instruction we would give Codex:
transcript.json has word timings (word, start, end in seconds). Put title.png centered, 120 px from the bottom of talk.mp4, from the start of the word "launch" to the end. One ffmpeg command, output synced.mp4. Then save one frame just before and one just after that word starts to ./frames, at 640 px wide.
And the commands it should produce, which we ran ourselves rather than through Codex, so the output here is ffmpeg's, not a model's:
# 3. What that prompt should produce ffmpeg -i talk.mp4 -i title.png -filter_complex "[0:v][1:v]overlay=x=(W-w)/2:y=H-h-120:enable='gte(t,2.48)'[v]" -map "[v]" -map 0:a -c:v libx264 -crf 18 -c:a copy synced.mp4 ffmpeg -ss 2.4 -i synced.mp4 -frames:v 1 -vf scale=640:-2 frames/at_2.4.png ffmpeg -ss 2.5 -i synced.mp4 -frames:v 1 -vf scale=640:-2 frames/at_2.5.png
The result: the frame at 2.4 s shows the plain test pattern, and the frame at 2.5 s shows the title. At 30 fps, 2.48 s falls between frames (2.48 x 30 = 74.4), so the title first appears on frame 75, at 2.50 s: 0.02 s after the word starts.
4How do you make Codex check its own video edits?
Give it house rules in AGENTS.md and pictures of the result with image input. Codex reads AGENTS.md before doing any work, and the CLI attaches frames with -i.
The r/codex creator above passes on a tip they credit to an OpenAI engineer:
“close the loop with the agent. Have it review its own output, looking at the images and iterating on itself.”
They saved time with a script that renders only certain frames for Codex to review. OpenAI's image inputs docs show the CLI taking one or more PNG or JPEG files with -i, comma-separated (checked October 2026). For our overlay:
codex -i frames/at_2.4.png,frames/at_2.5.png "The first frame is 0.08 s before the word 'launch', the second 0.02 s after. Is the title absent in the first and fully visible in the second?"
Name what each frame should show, as that page advises; a yes or no question beats “does this look right?”.
House rules in AGENTS.md
OpenAI's AGENTS.md docs say Codex reads a global file in ~/.codex, then files from the project root down to the current folder, with closer files taking precedence, up to 32 KiB combined by default (checked October 2026). Our starter for a video folder:
# AGENTS.md (in the video project folder) - Originals live in ./footage. Never write there; outputs go to ./out. - Transcripts live in ./work/transcript.json: word, start, end in seconds. - Time every overlay, cut and caption from transcript.json, never by ear. - One ffmpeg operation per command. Re-encode cuts. - After any visual change, save frames just before and after it.
Grow it every time you reject an edit. In an r/codex thread about a first Codex video edit, one commenter made the same point about iterating:
“You cant one shot perfection, but you can guide multiple 60-80% shots into a 95+% "good enough" range.”
5Can you run Codex video edits from a script?
Yes, with codex exec, which runs one task without the interactive interface. Give it write access explicitly, because it starts read-only. OpenAI's non-interactive mode docs say it runs in a read-only sandbox until you pass --sandbox workspace-write, that --json streams events (command runs, MCP tool calls) as JSON Lines, and that it expects a Git repository unless you pass --skip-git-repo-check (checked October 2026).
cd ~/videos/launch-teaser # a git repo: codex exec expects one codex exec --sandbox workspace-write --json \ "Follow AGENTS.md. Make a 1080x1920 copy of footage/talk.mp4 in ./out and report its duration with ffprobe." \ > run.jsonl
The log records every command Codex ran, so a bad cut is easy to trace. Put media folders in .gitignore so Git versions the rules and transcripts, not the footage. Our advice: run one clip by hand until AGENTS.md is stable, then batch.
6How do you connect Codex to a video editor over MCP?
Add the editor's MCP server with codex mcp add or a [mcp_servers] table in ~/.codex/config.toml, and raise its tool timeout. OpenAI's MCP docs say the ChatGPT desktop app, Codex CLI and IDE extension share this configuration, and give the command form:
codex mcp add <server-name> --env VAR1=VALUE1 -- <stdio server-command> codex mcp list
The same page gives tool_timeout_sec a default of 60 seconds (checked October 2026). Transcribing an interview or rendering a preview takes longer, and Codex stops waiting. Raise it per server:
# ~/.codex/config.toml [mcp_servers.my-editor] command = "/path/to/editor-mcp" args = [] tool_timeout_sec = 3600 # the default is 60 seconds
The pricing page adds that every MCP server uses more of your limit, so disable idle ones. Which editors offer a server that works with Codex is covered in our comparison of video editing MCP servers.
Example: Linkeddit Studio
Disclosure: we make Linkeddit Studio, a desktop video editor for Mac and Windows. With the Codex CLI installed, pick Codex in Studio's Agent panel and it runs there with Studio's tools attached. For your own terminal, Studio's Connect agent screen shows a config.toml block with your install's paths and tool_timeout_sec set to 3600, because transcribing a long clip takes minutes.
The word-timing idea from section 3 is built in: one Studio tool returns every spoken word with its start and end, and another cuts runs of words as one undo step. For checking, Codex can render up to 8 frames exactly as the export will show them and measure a clip's loudness in LUFS, its peak, and where people speak. Keep the app open while the agent works.
7How is CapCut's Codex plugin different?
CapCut's plugin connects Codex to CapCut's online service rather than to tools on your computer, and it is not yet available in the US. CapCut's codex ffmpeg video editing page links to its CapCut x Codex page, which says you install the plugin in Codex, generate images and video or upload a draft, refine the timeline in chat or the CapCut editor, and search templates. The same page says it is “Now available outside the U.S.; U.S. launch coming soon.” (both checked October 2026).
Its setup prompt, pasted into a ChatGPT desktop task, fetches an install runbook (plugin version 0.2.0). The runbook connects Codex to an MCP server at a capcut.com address, describes the editing tool as turning “uploaded footage into a finished video”, and checks your network location before sign-in: if the network appears to be in the United States, the plugin stays installed but “authorization is unavailable” (checked October 2026).
Our summary of the three ways to edit with Codex covered here (we make Linkeddit Studio):
| ffmpeg run by Codex | CapCut plugin | Editor over MCP (Studio) | |
|---|---|---|---|
| Where the footage goes | Stays on your computer | Uploaded, per CapCut's runbook | Stays on your computer |
| What you end up with | New video files | An editable CapCut draft | A project you can open and edit |
| What Codex can check | Frames and numbers you have it extract | Not stated by CapCut | Up to 8 rendered frames; LUFS, peak, speech |
| Software cost | Free | Not stated on the pages above | $99 one-time, no credits |
| US sign-in | Not applicable | Not yet, per CapCut | Not applicable |
8What goes wrong when you edit video with Codex?
Beyond the timeout, permissions and loose word ends covered above: shell quoting, invented flags, and the limits of the model.
- Windows quoting. OpenAI's Windows app docs say the desktop app's agent runs commands in PowerShell by default. Filter strings full of quotes, commas and colons break there first; say which shell you use in AGENTS.md, or switch the agent to WSL2.
- Invented ffmpeg flags. Models struggle most with long filter graphs, as our overview covers. Ask for one operation per command, or build the command with our ffmpeg command generator and have Codex run it.
- Content edits, not just properties. In the same first-edit thread, one reply argued that adding audio or subtitles and changing speed or color is relatively straightforward, and that “the real limitation is editing the actual video content itself”.
Our test before trusting any Codex setup: one short clip, one word-timed edit, two frames that prove it landed.
Let your own Codex edit a real timeline
Frequently asked questions
Can OpenAI Codex edit videos?+
Yes, through tools on your computer: ffmpeg, code for renderers such as Remotion, or a video editor's MCP tools. It does not watch video; it works from transcripts, command output and still frames.
Which ChatGPT plan do I need to use the Codex CLI?+
OpenAI's Codex pricing page lists Codex in the CLI starting with Plus; Pro includes everything in Plus, and Business and Enterprise include Codex too. Free and Go list Codex in the desktop app, subject to rollout. An API key works too, at standard API rates.
Does Codex video editing work on Windows?+
Yes. OpenAI publishes a PowerShell installer for the CLI, and the Windows desktop app runs the agent in PowerShell by default, with WSL2 as an option. Tell Codex which shell it uses, since quoting of ffmpeg filters differs.
How does Codex know when a word is spoken?+
From a transcript with word-level timestamps, made on your computer with Whisper (we used faster-whisper) or with OpenAI's whisper-1 API. Codex reads the JSON and puts the times into ffmpeg or Remotion.
Is CapCut's Codex plugin available in the United States?+
Not yet. CapCut's page says it is now available outside the U.S., with a U.S. launch coming soon, and its install runbook (plugin version 0.2.0) skips sign-in when your network appears to be in the US (both checked October 2026).
Why does my Codex MCP tool call stop after a minute?+
OpenAI's MCP docs give tool_timeout_sec a default of 60 seconds. Raise it for that server in config.toml when tools transcribe or render.