Field SOP
Field SOP

AI Podcast Production SOP: From Topic to Distribution

A one-person, zero-recording podcast pipeline across topic, script, TTS synthesis, polish, editing, and distribution, with 3 copy-paste prompts, pitfalls, and FAQ.

Published July 31, 20267 min read
<!-- ai-podcast-production-sop | sop | AI Podcast Production SOP: From Topic to Distribution -->

The thing that stops most people from starting a podcast isn't a lack of ideas-it's that recording is painful. No gear, no quiet room, tongue-tied reading a script, and on playback it's 40% "um" and "ah." AI flattens that barrier: a model writes the script, TTS synthesizes the voice, you edit by deleting words, and distribution is one RSS link. One person, no face, no mic, no gear, can ship a weekly show.

This is the full pipeline from topic to distribution: topic, script, TTS synthesis, polish, editing, distribution-six steps, each with real tools and copy-paste prompts, plus a pitfalls section and FAQ at the end. Pricing and capabilities are as of 2026-07-31 and subject to the official sites-TTS and editing tools change prices often, so verify before you commit.


1. Approach: The AI Podcast Pipeline at a Glance

The core shift is replacing "recording" with "generation + editing." Six steps, each with its tools:

StageWhat you doMainstream tools
TopicAI finds ideas, validates audienceLLM + Xiaoyuzhou trends
ScriptWrite a speakable, oral scriptLLM (with a prompt)
SynthesisTTS voice / voice cloningElevenLabs, Doubao TTS, Jimeng
PolishMake it spoken, prevent misreadsLLM + manual listen-back
EditingRemove filler, caption, write shownotesDescript, Jianying, Tongyi Tingwu
DistributionPublish to podcast directoriesXiaoyuzhou + Apple Podcasts + Spotify

Selection in one line: for English/multilingual or top-tier voice cloning, go ElevenLabs; for Chinese naturalness and API pricing, Doubao TTS (Volcengine) is more cost-effective; for a free all-in-one, Jianying's text-to-speech + smart spoken-word cut gets you through; to edit audio by editing text like a doc, use Descript.


2. Topic: Lock the Topic with AI

Don't pick topics on a hunch. Have the model batch-generate ideas based on your show's positioning, then pick the one with the sharpest pain point that you can fully unpack in one episode. The key is feeding the model your positioning and audience-otherwise it returns generic titles.

text
You are a podcast topic strategist. My show is about "[positioning, e.g. AI tools in practice]" for an audience of "[e.g. professionals who want to save time with AI]".
Based on the keyword "[keyword]", give me 5 topics, each with:
1. Title (<=20 chars, conversational, with a hook)
2. One-line pain point (why listeners must hear this)
3. Outline (3-5 points)
4. Suggested length (minutes)
5. One sample opening line
End by naming which is most likely to break out and why.

Pick criteria: a concrete pain point (not "AI matters" filler), something you can fully unpack in one episode, and a actionable takeaway. Two of three-check, and it's decided.


3. Script: Write Something You Can Speak

The biggest script trap is writing an "article"-sentences too long, awkward when TTS reads them, weird phrasing. A podcast script needs short sentences, spoken register, and room for tone.

text
You are a podcast script editor. Based on the topic "[topic]", write a single-host spoken-word script, about [8] minutes. Requirements:
- Pure spoken register, short sentences, max 25 words each, easy to read aloud
- Open with a 30-second hook (question / contrast / story)
- Middle in 3 segments, linked by transition lines
- End with one actionable step + a teaser for next episode
- Mark [pause][emphasis][light laugh] tone cues for TTS tuning
- No written-style phrasing, no "firstly secondly lastly", no subheadings
Output format: [Open]...[Seg 1]...[Seg 2]...[Seg 3]...[Close]

Read it silently once after drafting-fix wherever your tongue trips. Skip this and TTS will double-expose every awkward spot.


4. Synthesis: TTS Voice and Voice Cloning

With the script locked, into TTS. Three tools, different positions, pick by need:

  • ElevenLabs (international top pick): strongest voice cloning and emotional expression. Free tier ~10,000 characters/month (non-commercial, attribution required); Starter $5/mo (30k chars, instant cloning + commercial license); Creator $22/mo (100k chars, professional cloning). Credit-based, ~1 credit ≈ 2 characters.
  • Doubao TTS (Volcengine): Chinese TTS + high-fidelity voice cloning that intelligently predicts emotion and intonation, billed per API call with a free allowance for new users-good for Chinese podcast volume. Requires applying for appid/token/secret_key.
  • Jimeng dubbing (ByteDance Jimeng AI): leans toward video dubbing-Video 3.5 Pro supports simultaneous video+audio generation with lip-sync, ideal for captioned podcast clips; not the first pick for pure audio.

Key tip: don't dump the whole script at once. Generate in segments (<=30s each) and stitch-long text tends to drift on pauses, tail tones, and emotion. Brand names, acronyms, and numbers misread often; fix them with the tool's pronunciation dictionary or SSML, or mark [read as xxx].

To clone your own voice: both ElevenLabs Professional cloning and Descript Overdub need about 10 minutes of clean samples (no background noise, single speaker, natural pace)-dirty samples yield muddy clones.


5. Polish and Listen-Back: Make AI Sound Human

A first TTS pass is ~90% "robotic"-weird pause positions, stiff "therefore/moreover" from written phrasing, mispronounced brand names. This step fixes exactly that.

text
You are a podcast spoken-polish editor. Rewrite the passage below so it "reads aloud like a real person":
[paste passage]
Rules:
- Swap written words for spoken ones ("therefore"->"so", "moreover"->"also", "however"->"but")
- Break long sentences so each is one breath
- Add [read as xxx] after numbers / brand names / acronyms to prevent misreads
- Add 2-3 [pause] markers and filler words ("right?", "think about it") for natural rhythm
- Keep the original meaning, add no new points
Output: the polished text + one line on what changed.

After polishing you must put on headphones and listen segment by segment-don't just read the text. Check: number readings, brand names, proper nouns, pause naturalness. Fix problems back in the script, then re-synthesize just that segment-don't redo the whole thing.


6. Editing: Descript / Jianying + Shownotes

Editing isn't about effects-it's removing filler, captioning, and writing show notes.

  • Descript (text-based editing): transcribes audio to text; deleting text deletes audio, so you cut like editing a doc. Creator ~$24/mo includes Overdub (clone your own voice to fix flubs), filler removal (one-click "um/ah" delete), and studio sound (denoise + brighten). Top pick for English podcasts.
  • Jianying / CapCut (free all-in-one): text-to-speech (pick a voice or clone) + smart spoken-word cut (AI flags filler, edit by text) + vocal separation + audio denoise + loudness normalization. Zero-cost full chain for Chinese podcasts.
  • Tongyi Tingwu (Alibaba): speech-to-text at 0.6 RMB/hour; LLM capabilities (summary/chapters/mind map) from 0.22 RMB/hour, stacking per capability. Use it to transcribe the final cut into shownotes and chapter timestamps for distribution.

Edit order: denoise + normalize loudness first -> cut filler and flubs -> add intro/outro and transitions -> export MP3 128kbps (~30MB for 30 min, accepted everywhere) -> transcribe shownotes with Tongyi Tingwu.


7. Distribution: Xiaoyuzhou + Apple Podcasts + Spotify

Podcast distribution runs on one mechanism: a hosting platform generates an RSS feed, which you submit to directories.

  1. Host on Xiaoyuzhou: in the creator dashboard, create a show, upload audio + cover (3000x3000 square PNG) + show info; Xiaoyuzhou generates an RSS link (Show settings -> Multi-platform sync). The default pick in China, free.
  2. Apple Podcasts: sign in to podcasters.apple.com (Apple Podcasts Connect) with an Apple ID, "Add a new show," paste the RSS, submit for review-1-5 business days. The directory with the widest reach.
  3. Spotify: go to Spotify for Creators (podcasters.spotify.com), paste RSS to claim the show, or use the host's direct-submit button.

Key traps: the cover must be square and high-res (3000x3000 passes best); the RSS show description, category, and author fields can't be empty or Apple rejects; after each episode, update once on the host and all three platforms sync automatically-no re-uploading.


8. Pitfalls from the Trenches

Pitfall 1: Dumping the whole script into TTS at once. Long text drifts on pauses, emotion, and tail tones. Fix: cut by segment, <=30s each, generate and stitch, and you can redo single segments.

Pitfall 2: Brand names / numbers / acronyms misread. ElevenLabs often gets brand-name stress and number readings wrong. Use the pronunciation dictionary or SSML, or mark [read as xxx] so you know what to fix later.

Pitfall 3: Voice clone sample too short or noisy. The clone comes out muddy and off. Samples need 10+ minutes, clean, single-speaker, natural pace. No pro gear? Record in a closet-clothes dampen echo.

Pitfall 4: Using Jimeng dubbing as pure TTS. Jimeng's strength is simultaneous video+audio (lip-sync); exporting audio alone is roundabout. Use ElevenLabs / Doubao for pure audio; save Jimeng for captioned clips.

Pitfall 5: Apple Podcasts cover rejected. Nine times out of ten the image isn't square or the resolution is too low. Use a 3000x3000 square PNG, fill in the RSS description/category/author, and resubmit.

Pitfall 6: Using the free TTS tier commercially. ElevenLabs' free tier is non-commercial (attribution required); commercial use starts at Starter at minimum. Check each tool's license for commercial use-don't assume.

Pitfall 7: Publishing audio with no shownotes. Distributed to Apple with no show notes, listeners scroll past. Transcribe shownotes + chapter timestamps with Tongyi Tingwu and paste them in.


9. FAQ

Q1: Can one person with no gear or budget make a podcast? Yes. Script with an LLM, voice with Jianying's text-to-speech or ElevenLabs' free allowance, edit in Jianying, distribute on Xiaoyuzhou-zero recording, zero gear. But voice cloning and commercial use cost money.

Q2: ElevenLabs or Doubao TTS-how to choose? For English/multilingual or cloning your own voice at high quality, ElevenLabs; for Chinese naturalness and high-volume API calls, Doubao (Volcengine) pay-as-you-go is more cost-effective. For a Chinese podcast, try Doubao first.

Q3: Will listeners skip an AI-voiced podcast? High-fidelity TTS (ElevenLabs / Doubao high-fidelity cloning) is mostly imperceptible; pure robotic voices get skipped instantly. Label the show "AI-assisted voice"-if naturalness is high, retention is fine.

Q4: What if Apple Podcasts rejects my submission? Usually the cover (not square / low-res) or RSS missing fields (description/category/author). Go back to the Xiaoyuzhou dashboard, fill them in, swap to a 3000x3000 square image, and resubmit.

Q5: How big should each episode's audio file be, and what format? Export MP3, 128kbps, ~30MB for 30 minutes. Xiaoyuzhou / Apple / Spotify all accept MP3-don't upload WAV (too big, slow upload, and the platform re-compresses anyway).


References

This article is AI-assisted and human-edited. Last updated: 2026-07-31

FAQ

Can one person with no gear or budget make a podcast?
Yes. Script with an LLM, voice with Jianying's text-to-speech or ElevenLabs' free allowance, edit in Jianying, distribute on Xiaoyuzhou-zero recording, zero gear. But voice cloning and commercial use cost money.
ElevenLabs or Doubao TTS-how to choose?
For English/multilingual or cloning your own voice at high quality, ElevenLabs; for Chinese naturalness and high-volume API calls, Doubao (Volcengine) pay-as-you-go is more cost-effective. For a Chinese podcast, try Doubao first.
Will listeners skip an AI-voiced podcast?
High-fidelity TTS (ElevenLabs / Doubao high-fidelity cloning) is mostly imperceptible; pure robotic voices get skipped instantly. Label the show "AI-assisted voice"-if naturalness is high, retention is fine.
What if Apple Podcasts rejects my submission?
Usually the cover (not square / low-res) or RSS missing fields (description/category/author). Go back to the Xiaoyuzhou dashboard, fill them in, swap to a 3000x3000 square image, and resubmit.
How big should each episode's audio file be, and what format?
Export MP3, 128kbps, ~30MB for 30 minutes. Xiaoyuzhou / Apple / Spotify all accept MP3-don't upload WAV (too big, slow upload, and the platform re-compresses anyway).

Related

Field SOP

AI Digital Human Creation SOP: A Repeatable Workflow from Script to Final Cut

Breaks AI digital human creation into a six-step repeatable workflow: pick the tool by use case (HeyGen/D-ID/Synthesia/Colossyan/DeepBrain plus China's Tencent Zhiying/Guiji Intelligent), write the talking-head script (with prompt template), pick or customize the avatar, lock the voice before driving lip-sync, post-process subtitles/editing/compliance, and publish with platform adaptation. Includes 5 pitfalls (avatar licensing/lip-sync drift/multilingual voice/long-video cost/compliance labels) and 5 FAQs. Representative workflow, not a single-tool hands-on test; features subject to official sites.

Aug 7, 20268 min read