Video Editing

How to Edit a Podcast Video: Sync, Switch, Level, Export

Most podcast editing guides are written for audio. Video adds cameras that start at different times, deciding who is on screen, a loudness target, and vertical clips that cut people out. Here is the order to work in, with every command run and measured.

By Linkeddit·October 5, 2026·12 min read

Key takeaways

  • Order: sync, rough cut, angles, silence and filler cleanup, loudness, captions, export, clips. Sync before the first cut.
  • Sync by audio: line each camera's scratch sound up with the mic. In our test cross-correlation recovered a 1.300 s offset, re-checked at 0.000 s.
  • Apple Podcasts documents -16 LKFS plus or minus 1 dB, true peak no higher than -1 dBFS. Spotify's -14 LUFS page is about music, and YouTube publishes no loudness number we could find.
  • Single-pass ffmpeg loudnorm missed: asked for -16, it measured -14.8 LUFS on our file. Two-pass linear mode hit -16.0.
  • A center crop to 1080x1920 keeps under a third of a 16:9 frame. With two people on screen, each vertical clip needs its own framing decision.

1What is the order for editing a video podcast?

Sync the cameras to the mic audio, make the rough cut, choose angles, clean up silence and fillers, set loudness, add captions, export the episode, then cut clips from the finished edit. Sync is easy on raw, full-length files and painful once anything has been cut.

  1. Sync. Line up every camera and mic track on one timeline.
  2. Rough cut. Drop the pre-show, breaks and known tangents.
  3. Angles. Decide who is on screen, cut by cut.
  4. Cleanup. Long pauses and fillers, at a level the picture can hide.
  5. Loudness. Measure the mix and bring it to one target.
  6. Captions, export, clips. Last, because each depends on the final timing.

On time, The Podcast Host says a 30 minute show that only gets a top and tail “it'll take no longer than 15 minutes to edit” (checked October 2026), but that page is about audio. For video, the only numbers we found are self-reports. One host of a two-person show that records with a hardware switcher wrote:

“I'd say the total edit time is now about 3:1 to get the audio AND video output.”
r/podcasting

That is one show, not a rule. If you are new, a reply in an r/podcasting advice thread is sensible: “Start slowly. Upload the first few episodes with minimal or no editing”.

2How do you sync multicam podcast footage by audio?

Record scratch audio on every camera, then line each camera's sound up with the mic recording. The shift that makes the waveforms match is the offset; apply it and picture and mic are in sync. A clap helps you check by eye, but speech lines up on its own. Editors' sync by audio features rest on the same idea.

We tested it. We generated a 20 second “mic” track at 48 kHz, then a “camera” copy that starts 1.3 seconds late at 30 percent of the level, muxed onto a 1920x1080 test picture. Cross-correlation tests every shift and picks the best match. This script does it with numpy and ffmpeg:

import subprocess, sys
import numpy as np

def load(path, sr=8000):
    raw = subprocess.run(["ffmpeg", "-v", "error", "-i", path, "-ac", "1",
                          "-ar", str(sr), "-f", "f32le", "-"],
                         capture_output=True, check=True).stdout
    return np.frombuffer(raw, dtype=np.float32)

sr = 8000
a, b = load(sys.argv[1], sr), load(sys.argv[2], sr)
n = len(a) + len(b)
corr = np.fft.irfft(np.fft.rfft(b, n) * np.conj(np.fft.rfft(a, n)), n)
lag = int(np.argmax(corr))
if lag > n // 2:
    lag -= n
print(f"{sys.argv[2]} lags {sys.argv[1]} by {lag / sr:.3f} s")
python3 offset.py mic.wav cam.mp4
# cam.mp4 lags mic.wav by 1.300 s

It found the offset exactly, despite the quieter camera copy. We applied it two ways and re-ran the check, with ffmpeg 9.0.2:

# A: delay the mic file (our pick)
ffmpeg -i mic.wav -af adelay=1300 mic_delayed.wav
# re-check: mic_delayed.wav lags cam.wav by 0.000 s

# B: offset the mic while muxing
ffmpeg -i cam.mp4 -itsoffset 1.3 -i mic.wav -map 0:v -map 1:a \
  -c:v copy -c:a aac -shortest synced.mp4
# ffprobe: audio start_time 1.278667, not 1.300

Both lined up. One catch with B: the AAC audio stream starts at 1.278667 seconds, about 21 ms earlier than asked, because AAC encoding adds a short priming delay. Our re-check only read 0.000 s once we resampled from the true start of the stream; a plain extract reported -1.279 s, which is a measurement artifact, not a sync error. Delaying the mic file with adelay avoids that question entirely, which is why we prefer it.

3How often should you cut between speakers?

Cut to the person who is talking, go wide when people overlap or laugh together, and hold a reaction shot when the listener's face adds something. Those are our editorial rules, not data. For scale, Descript's Automatic Multicam help page (checked October 2026) offers a cutaway every 30 seconds (“Occasional”) or every 10 seconds (“Frequent”) during monologues, or a “Show only active speaker” style with no cutaways.

  • Cut on the sentence, not the syllable. Switch as a new speaker starts a thought, not on every short interjection.
  • Avoid very short shots. In our experience a shot under about a second reads as a glitch.
  • Keep a wide. B&H's video podcast editing guide (checked October 2026) suggests at least one medium or close-up shot of each speaker, plus a wide to establish the space. The wide is also where to go when nobody owns the moment.
  • Use the switch to hide audio edits. A cut in the sound looks natural if the picture changes angle at the same frame.

That is why more cameras make editing easier, not just prettier. From the r/podcasting advice thread:

“For solo episodes I would definitely try to get at least two camera to record. You can hide a lot of edits this way.”
r/podcasting

With only one camera, B&H recommends shooting in 4K or higher so you can crop in. Cut between the full frame and a tighter crop and you have a second angle.

4How do you remove silence and filler words from a video podcast?

Shorten long pauses and remove the fillers that break a sentence, but cut less than you would in audio, because every cut is visible. The Podcast Host makes the same point: “you really can't go as granular as audio, as it'll look jumpy and jarring”. The r/podcasting comment quoted above adds: “I hate jump cuts. But making the audio sound perfect does mean the video takes a lot more time to match.”

We measured both elsewhere. Our guide to removing silence from video tests what each threshold and minimum length actually removes, and the silence detector finds the pauses in your browser. The guide to removing filler words shows why cutting at transcript word times can leave half an um behind. Do both on the synced multicam timeline, so each cut takes every camera and mic with it.

5What loudness should a video podcast be?

Master to about -16 LUFS integrated with true peak at or below -1 dBFS. That is Apple's documented target, and it is the only podcast loudness figure we found on a platform's own pages. LKFS and LUFS are the same scale under two names.

PlatformWhat its own page says (checked October 2026)
Apple PodcastsAround -16 dB LKFS, +/- 1 dB tolerance, true peak not above -1 dB FS, measured per ITU-R BS.1770-5
SpotifyNormalizes tracks to -14 dB LUFS; this page is in Spotify for Artists and is about music
YouTubeNo loudness number on the YouTube Help pages we checked

Sources: Apple Podcasts audio requirements and Spotify loudness normalization. Apple says to do this before encoding, because compression can clip a signal whose true peak is too high. Check your own file with the loudness checker, which measures in the browser.

In ffmpeg, how you run loudnorm matters. We measured our 20 second test file first:

ffmpeg -hide_banner -nostats -i mic.wav -af ebur128=peak=true -f null -
#   I: -6.6 LUFS   True peak: 0.9 dBFS

# pass 1: measure only
ffmpeg -i mic.wav -af loudnorm=I=-16:TP=-1.5:LRA=11:print_format=json -f null -
#   "input_i": "-6.61", "input_tp": "0.88", "input_lra": "3.20",
#   "input_thresh": "-16.87", "target_offset": "-1.16"

Then we normalized it to -16 two ways and measured each output with the same ebur128 command:

# single pass (the trap)
ffmpeg -i mic.wav -af "loudnorm=I=-16:TP=-1.5:LRA=11" -ar 48000 out.wav
#   measured: I: -14.8 LUFS, peak -5.2 dBFS  (normalization_type: dynamic)

# pass 2: feed pass 1 back in, linear mode
ffmpeg -i mic.wav -af "loudnorm=I=-16:TP=-1.5:LRA=11:measured_I=-6.61:\
measured_TP=0.88:measured_LRA=3.20:measured_thresh=-16.87:offset=-1.16:\
linear=true" -ar 48000 n-16.wav
#   measured: I: -16.0 LUFS, peak -8.5 dBFS  (normalization_type: linear)

The single pass landed at -14.8 LUFS, 1.2 LU louder than asked, because without measured values loudnorm works in dynamic mode and adjusts as it goes. The two-pass run reported linear mode and hit -16.0 LUFS; the same two-pass command with I=-14 measured -14.0. That is one synthetic file; real speech differs, so measure your output.

6What export settings should a podcast video use?

For YouTube, export at the frame rate you recorded, with 48 kHz stereo audio; YouTube recommends 384 kbps for stereo and 8 Mbps video for 1080p SDR at 24 to 30 fps. Source: YouTube's recommended upload encoding settings (checked October 2026). For an audio-only RSS feed, Apple “strongly recommend[s] using AAC instead of MP3”. Codec, container and HDR choices are covered in best export settings for YouTube, and the YouTube export settings tool gives the numbers for your resolution and frame rate.

7How do you add captions to a video podcast?

Transcribe the final edit, and if you have separate mic tracks, transcribe each one on its own: where two people talk over each other, a mixed track loses words. We tested it with openai-whisper (base.en, word timestamps) on a 9.6 second clip of two macOS say voices that overlap for about 1.6 seconds. On the mixed track, Whisper wrote one transcript with no speaker labels, turned the host's “editing video podcasts” into “editing a video” and dropped the guest's opening “Right,”. Each voice transcribed alone kept every word, though “syncing” came back as “sinking”, so proofread either way.

Burn captions into clips; give the full episode a separate file viewers can turn off. Our Whisper subtitles guide covers the setup, and the caption reading speed checker flags lines that go by too fast.

8How do you turn a podcast video into Shorts and clips?

Cut each clip from the finished episode, crop it to 1080x1920, and decide its framing by hand whenever two people share the shot. YouTube's aspect ratio help (checked October 2026) asks you not to add padding or black bars, because the player adapts to your video's shape. This is the center crop we ran on our synced test file:

ffmpeg -ss 5 -t 8 -i synced.mp4 \
  -vf "crop=ih*9/16:ih:(iw-ih*9/16)/2:0,scale=1080:1920,setsar=1" \
  -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a aac -b:a 192k clip_9x16.mp4

ffprobe -v error -select_streams v:0 \
  -show_entries stream=width,height,sample_aspect_ratio,duration \
  -of csv=p=0 clip_9x16.mp4
# 1080,1920,1:1,8.000000

The output was exactly 1080x1920, square pixels, 8.000 seconds with matching audio. From a 1920 pixel wide frame, ffmpeg kept a strip 608 pixels wide, under a third of the picture: room for one person in the middle. Two people side by side lose at least one. Automated tools struggle here too. A producer who tested a run of AI clipping tools wrote:

“The moment there are two speakers and a camera, the wheels come off: vertical reframing that actually follows whoever's talking, captions that don't need a cleanup pass, cuts that don't butcher the timing.”
r/podcasting

Treat the vertical clip as its own edit: crop to the speaker and switch crops when the speaker changes. A YouTube commenter under a video podcast tutorial agreed: “I believe the short & tall videos benefit from their own multicam edits as the composition changes from landscape to vertical.” The aspect ratio calculator gives crop sizes for other source resolutions.

9Which editor should you use for a video podcast?

Pick an editor with multicam if you have more than one camera. Only what each vendor documents is listed, checked October 2026. We make Linkeddit Studio, so its row is the vendor talking.

EditorWhat the vendor documents for podcasts
DaVinci ResolveResolve 21 feature list: creating multicam clips on desktop; Multicam SmartSwitch marked Studio
DescriptAutomatic Multicam: you assign each audio track to a camera, it switches angles and adds cutaways; Center active speaker (Beta) reframes one clip with combined video and audio
RiversidePricing page lists in-person multi-cam recording and unlimited text-based editing
Adobe PremiereNot verified: we could not confirm Adobe's documentation, so we list nothing
CapCutPodcast maker page covers voice recording, volume, pitch and speed; no multicam listed there
Linkeddit StudioNo multicam, no sync by audio, no speaker detection; transcript cleanup, LUFS normalize, burned-in captions, 1080x1920

Sources: the DaVinci Resolve Studio 21 feature list (September 2026), Descript's Automatic Multicam and Center active speaker pages, which say the reframe does not work on multi-track sequences, Riverside pricing and CapCut's podcast maker. The Resolve list does not settle what the free edition includes; our DaVinci Resolve vs CapCut comparison covers what each free tier holds back.

10Where does Linkeddit Studio fit in a podcast workflow?

After the camera decisions are made. Linkeddit Studio has no multicam, does not sync tracks by audio and does not detect speakers, so it cannot switch angles for you. Do sync and angle switching in a multicam editor first, then bring the edited episode in. What Studio does with it:

  • Pauses and fillers from the transcript: Tighten pauses and Remove filler words buttons show what will go, and cut only when you press Remove.
  • Loudness. The agent measures LUFS, peak and speech regions, then normalizes a clip or track to a target; -16 is the podcast preset, and the export limiter keeps peaks under -1 dBFS. It can also duck music under speech.
  • Captions from on-device Whisper, burned in, in plain, highlight, pop, karaoke or box styles.
  • Clips. 1080x1920 projects for Shorts, exported as MP4 (H.264 and AAC), or the episode sound alone as .m4a.

Your own Claude Code or Codex does the editing inside Studio, and every edit is undoable. It costs $99 once, no subscription. Every command above runs free with ffmpeg.

Finish the episode after the multicam edit

Linkeddit Studio is a desktop editor for Mac and Windows where Claude Code or Codex edits your footage. Tighten pauses and fillers from the transcript, normalize to -16 LUFS, burn in captions and cut 1080x1920 clips. Files stay on your disk. $99 once, no subscription.
See Linkeddit Studio

Frequently asked questions

How long does it take to edit a video podcast?+

There is no benchmark, only self-reports. One two-person show on r/podcasting reports about 3:1 for audio and video with a hardware switcher. The Podcast Host says a 30 minute audio show that only gets a top and tail takes up to 15 minutes.

How do I sync separate audio and video for a podcast?+

Line up each camera's scratch audio with the mic recording before you make any cuts. In our test, cross-correlation found a 1.300 second offset and delaying the mic by it left a 0.000 second residual.

What loudness should a podcast be for YouTube and Spotify?+

Apple Podcasts documents about -16 LKFS, plus or minus 1 dB, with true peak no higher than -1 dBFS. Spotify's -14 LUFS page is written for music, and we found no YouTube loudness number.

Should I cut to the person talking or keep a wide shot?+

Mostly cut to the speaker, go wide when people talk over each other, and hold a reaction shot when it adds something. Descript's automatic multicam presets cut away every 10 or 30 seconds.

Can I edit a video podcast with one camera?+

Yes. Shoot in 4K or higher, as B&H suggests, and cut between the full frame and a tighter crop.

How do I make vertical clips without cutting a speaker out?+

A center crop of a 16:9 frame keeps under a third of the width, enough for one person. For two, crop to whoever is talking and recut when the speaker changes.