Field SOP
Field SOP

AI Podcast Production SOP: From Topic to Distribution

A one-person, zero-recording podcast pipeline across topic, script, TTS synthesis, polish, editing, and distribution, with 3 copy-paste prompts, pitfalls, and FAQ.

Published July 31, 20267 min read
<!-- ai-podcast-production-sop | sop | AI Podcast Production SOP: From Topic to Distribution -->

The thing that stops most people from starting a podcast isn't a lack of ideas-it's that recording is painful. No gear, no quiet room, tongue-tied reading a script, and on playback it's 40% "um" and "ah." AI flattens that barrier: a model writes the script, TTS synthesizes the voice, you edit by deleting words, and distribution is one RSS link. One person, no face, no mic, no gear, can ship a weekly show.

This is the full pipeline from topic to distribution: topic, script, TTS synthesis, polish, editing, distribution-six steps, each with real tools and copy-paste prompts, plus a pitfalls section and FAQ at the end. Pricing and capabilities are as of 2026-07-31 and subject to the official sites-TTS and editing tools change prices often, so verify before you commit.


1. Approach: The AI Podcast Pipeline at a Glance

The core shift is replacing "recording" with "generation + editing." Six steps, each with its tools:

StageWhat you doMainstream tools
TopicAI finds ideas, validates audienceLLM + Xiaoyuzhou trends
ScriptWrite a speakable, oral scriptLLM (with a prompt)
SynthesisTTS voice / voice cloningElevenLabs, Doubao TTS, Jimeng
PolishMake it spoken, prevent misreadsLLM + manual listen-back
EditingRemove filler, caption, write shownotesDescript, Jianying, Tongyi Tingwu
DistributionPublish to podcast directoriesXiaoyuzhou + Apple Podcasts + Spotify

Selection in one line: for English/multilingual or top-tier voice cloning, go ElevenLabs; for Chinese naturalness and API pricing, Doubao TTS (Volcengine) is more cost-effective; for a free all-in-one, Jianying's text-to-speech + smart spoken-word cut gets you through; to edit audio by editing text like a doc, use Descript.


2. Topic: Lock the Topic with AI

Don't pick topics on a hunch. Have the model batch-generate ideas based on your show's positioning, then pick the one with the sharpest pain point that you can fully unpack in one episode. The key is feeding the model your positioning and audience-otherwise it returns generic titles.

text
You are a podcast topic strategist. My show is about "[positioning, e.g. AI tools in practice]" for an audience of "[e.g. professionals who want to save time with AI]".
Based on the keyword "[keyword]", give me 5 topics, each with:
1. Title (<=20 chars, conversational, with a hook)
2. One-line pain point (why listeners must hear this)
3. Outline (3-5 points)
4. Suggested length (minutes)
5. One sample opening line
End by naming which is most likely to break out and why.

Pick criteria: a concrete pain point (not "AI matters" filler), something you can fully unpack in one episode, and a actionable takeaway. Two of three-check, and it's decided.


3. Script: Write Something You Can Speak

The biggest script trap is writing an "article"-sentences too long, awkward when TTS reads them, weird phrasing. A podcast script needs short sentences, spoken register, and room for tone.

text
You are a podcast script editor. Based on the topic "[topic]", write a single-host spoken-word script, about [8] minutes. Requirements:
- Pure spoken register, short sentences, max 25 words each, easy to read aloud
- Open with a 30-second hook (question / contrast / story)
- Middle in 3 segments, linked by transition lines
- End with one actionable step + a teaser for next episode
- Mark [pause][emphasis][light laugh] tone cues for TTS tuning
- No written-style phrasing, no "firstly secondly lastly", no subheadings
Output format: [Open]...[Seg 1]...[Seg 2]...[Seg 3]...[Close]

Read it silently once after drafting-fix wherever your tongue trips. Skip this and TTS will double-expose every awkward spot.


4. Synthesis: TTS Voice and Voice Cloning

With the script locked, into TTS. Three tools, different positions, pick by need:

  • ElevenLabs (international top pick): strongest voice cloning and emotional expression. Free tier ~10,000 characters/month (non-commercial, attribution required); Starter $5/mo (30k chars, instant cloning + commercial license); Creator $22/mo (100k chars, professional cloning). Credit-based, ~1 credit ≈ 2 characters.
  • Doubao TTS (Volcengine): Chinese TTS + high-fidelity voice cloning that intelligently predicts emotion and intonation, billed per API call with a free allowance for new users-good for Chinese podcast volume. Requires applying for appid/token/secret_key.
  • Jimeng dubbing (ByteDance Jimeng AI): leans toward video dubbing-Video 3.5 Pro supports simultaneous video+audio generation with lip-sync, ideal for captioned podcast clips; not the first pick for pure audio.

Key tip: don't dump the whole script at once. Generate in segments (<=30s each) and stitch-long text tends to drift on pauses, tail tones, and emotion. Brand names, acronyms, and numbers misread often; fix them with the tool's pronunciation dictionary or SSML, or mark [read as xxx].

To clone your own voice: both ElevenLabs Professional cloning and Descript Overdub need about 10 minutes of clean samples (no background noise, single speaker, natural pace)-dirty samples yield muddy clones.


5. Polish and Listen-Back: Make AI Sound Human

A first TTS pass is ~90% "robotic"-weird pause positions, stiff "therefore/moreover" from written phrasing, mispronounced brand names. This step fixes exactly that.

text
You are a podcast spoken-polish editor. Rewrite the passage below so it "reads aloud like a real person":
[paste passage]
Rules:
- Swap written words for spoken ones ("therefore"->"so", "moreover"->"also", "however"->"but")
- Break long sentences so each is one breath
- Add [read as xxx] after numbers / brand names / acronyms to prevent misreads
- Add 2-3 [pause] markers and filler words ("right?", "think about it") for natural rhythm
- Keep the original meaning, add no new points
Output: the polished text + one line on what changed.

After polishing you must put on headphones and listen segment by segment-don't just read the text. Check: number readings, brand names, proper nouns, pause naturalness. Fix problems back in the script, then re-synthesize just that segment-don't redo the whole thing.


6. Editing: Descript / Jianying + Shownotes

Editing isn't about effects-it's removing filler, captioning, and writing show notes.

  • Descript (text-based editing): transcribes audio to text; deleting text deletes audio, so you cut like editing a doc. Creator ~$24/mo includes Overdub (clone your own voice to fix flubs), filler removal (one-click "um/ah" delete), and studio sound (denoise + brighten). Top pick for English podcasts.
  • Jianying / CapCut (free all-in-one): text-to-speech (pick a voice or clone) + smart spoken-word cut (AI flags filler, edit by text) + vocal separation + audio denoise + loudness normalization. Zero-cost full chain for Chinese podcasts.
  • Tongyi Tingwu (Alibaba): speech-to-text at 0.6 RMB/hour; LLM capabilities (summary/chapters/mind map) from 0.22 RMB/hour, stacking per capability. Use it to transcribe the final cut into shownotes and chapter timestamps for distribution.

Edit order: denoise + normalize loudness first -> cut filler and flubs -> add intro/outro and transitions -> export MP3 128kbps (~30MB for 30 min, accepted everywhere) -> transcribe shownotes with Tongyi Tingwu.


7. Distribution: Xiaoyuzhou + Apple Podcasts + Spotify

Podcast distribution runs on one mechanism: a hosting platform generates an RSS feed, which you submit to directories.

  1. Host on Xiaoyuzhou: in the creator dashboard, create a show, upload audio + cover (3000x3000 square PNG) + show info; Xiaoyuzhou generates an RSS link (Show settings -> Multi-platform sync). The default pick in China, free.
  2. Apple Podcasts: sign in to podcasters.apple.com (Apple Podcasts Connect) with an Apple ID, "Add a new show," paste the RSS, submit for review-1-5 business days. The directory with the widest reach.
  3. Spotify: go to Spotify for Creators (podcasters.spotify.com), paste RSS to claim the show, or use the host's direct-submit button.

Key traps: the cover must be square and high-res (3000x3000 passes best); the RSS show description, category, and author fields can't be empty or Apple rejects; after each episode, update once on the host and all three platforms sync automatically-no re-uploading.


8. Pitfalls from the Trenches

Pitfall 1: Dumping the whole script into TTS at once. Long text drifts on pauses, emotion, and tail tones. Fix: cut by segment, <=30s each, generate and stitch, and you can redo single segments.

Pitfall 2: Brand names / numbers / acronyms misread. ElevenLabs often gets brand-name stress and number readings wrong. Use the pronunciation dictionary or SSML, or mark [read as xxx] so you know what to fix later.

Pitfall 3: Voice clone sample too short or noisy. The clone comes out muddy and off. Samples need 10+ minutes, clean, single-speaker, natural pace. No pro gear? Record in a closet-clothes dampen echo.

Pitfall 4: Using Jimeng dubbing as pure TTS. Jimeng's strength is simultaneous video+audio (lip-sync); exporting audio alone is roundabout. Use ElevenLabs / Doubao for pure audio; save Jimeng for captioned clips.

Pitfall 5: Apple Podcasts cover rejected. Nine times out of ten the image isn't square or the resolution is too low. Use a 3000x3000 square PNG, fill in the RSS description/category/author, and resubmit.

Pitfall 6: Using the free TTS tier commercially. ElevenLabs' free tier is non-commercial (attribution required); commercial use starts at Starter at minimum. Check each tool's license for commercial use-don't assume.

Pitfall 7: Publishing audio with no shownotes. Distributed to Apple with no show notes, listeners scroll past. Transcribe shownotes + chapter timestamps with Tongyi Tingwu and paste them in.


9. FAQ

Q1: Can one person with no gear or budget make a podcast? Yes. Script with an LLM, voice with Jianying's text-to-speech or ElevenLabs' free allowance, edit in Jianying, distribute on Xiaoyuzhou-zero recording, zero gear. But voice cloning and commercial use cost money.

Q2: ElevenLabs or Doubao TTS-how to choose? For English/multilingual or cloning your own voice at high quality, ElevenLabs; for Chinese naturalness and high-volume API calls, Doubao (Volcengine) pay-as-you-go is more cost-effective. For a Chinese podcast, try Doubao first.

Q3: Will listeners skip an AI-voiced podcast? High-fidelity TTS (ElevenLabs / Doubao high-fidelity cloning) is mostly imperceptible; pure robotic voices get skipped instantly. Label the show "AI-assisted voice"-if naturalness is high, retention is fine.

Q4: What if Apple Podcasts rejects my submission? Usually the cover (not square / low-res) or RSS missing fields (description/category/author). Go back to the Xiaoyuzhou dashboard, fill them in, swap to a 3000x3000 square image, and resubmit.

Q5: How big should each episode's audio file be, and what format? Export MP3, 128kbps, ~30MB for 30 minutes. Xiaoyuzhou / Apple / Spotify all accept MP3-don't upload WAV (too big, slow upload, and the platform re-compresses anyway).


References

This article is AI-assisted and human-edited. Last updated: 2026-07-31

FAQ

Can one person with no gear or budget make a podcast?
Yes. Script with an LLM, voice with Jianying's text-to-speech or ElevenLabs' free allowance, edit in Jianying, distribute on Xiaoyuzhou-zero recording, zero gear. But voice cloning and commercial use cost money.
ElevenLabs or Doubao TTS-how to choose?
For English/multilingual or cloning your own voice at high quality, ElevenLabs; for Chinese naturalness and high-volume API calls, Doubao (Volcengine) pay-as-you-go is more cost-effective. For a Chinese podcast, try Doubao first.
Will listeners skip an AI-voiced podcast?
High-fidelity TTS (ElevenLabs / Doubao high-fidelity cloning) is mostly imperceptible; pure robotic voices get skipped instantly. Label the show "AI-assisted voice"-if naturalness is high, retention is fine.
What if Apple Podcasts rejects my submission?
Usually the cover (not square / low-res) or RSS missing fields (description/category/author). Go back to the Xiaoyuzhou dashboard, fill them in, swap to a 3000x3000 square image, and resubmit.
How big should each episode's audio file be, and what format?
Export MP3, 128kbps, ~30MB for 30 minutes. Xiaoyuzhou / Apple / Spotify all accept MP3-don't upload WAV (too big, slow upload, and the platform re-compresses anyway).

Related

Field SOP

No app switching: JianYing Hub from storyboard to final cut

A hands-on SOP for JianYing Hub that starts with a fitness check, giving three cases where it fits (you need a full asset-to-final-cut chain, you want batches of the same storyboard for A/B tests, or you dislike repeated export and import across apps) and four where it does not (you only lack one generated shot, you need fine multi-cam editing, your assets and data must stay local, or your budget must be precisely predictable), and noting that if data-on-premises or avoiding ecosystem lock-in is the priority you should look at an open-source route instead. It then walks the official 14-step onboarding path all the way to multi-track refinement: confirm the entry and asset source (keeping the same JianYing account so Jimeng and Xiaoyunque assets import in one click), generate key assets and a storyboard script, wire assets to script, let asset management auto-complete characters, props and scenes, batch-generate storyboard prompts (with two copy-ready prompt templates, one for product ads and one for talking-head video), pick models such as Seedance 2.5 to generate each storyboard clip, preview the assembled cut, use "More Editing" to jump straight into JianYing's multi-track timeline, apply local edits, AI extend, and AI post-editing (upscaling, frame interpolation, color grading, removal, vocal separation), call Skills through the JianYing Assistant in text, and import ByteDance assets back in. It adds three batch-production methods (deriving multiple versions from one product for A/B, a talking-head batch pipeline, and size and platform adaptation), a line-by-line pre-launch checklist and a six-row pitfall table covering script drift, local edits breaking consistency, audio-video desync after AI extend, human review of proper nouns in subtitle fixes, version and font mismatches after import, and rigid pacing from over-reliance on auto-assembly. Pricing, free quota, concurrency, credits, regional availability and duration or resolution limits are all unpublished and stated as following the official app, and JianYing Hub is noted as a closed-source commercial product that is neither open source nor free.

Sep 22, 20268 min read
Field SOP

Same voice, 2.3s lag: ship Qwen3.8 live-translate in your app

A deployment SOP for Qwen3.8-LiveTranslate that starts with a fitness check: if you need a conversation to be understood as it happens, use simultaneous interpretation, and if you can translate slowly afterwards, start with offline. It then pins down the two most commonly misused figures, that 2.3 seconds is average lag (LAAL) rather than end-to-end first-packet latency, and that 60 is recognition input while 29 is speech output, two different units with the remaining 31 text-only. Credentials and environment come next: keys belong only in environment variables, never in source, repositories or front-end bundles, since committing one puts it into version history and the shipped bundle, and the only correct response to a leak is to revoke the old key immediately and rebuild and rotate it. It also warns up front about the trap that a browser cannot set an Authorization header during a WebSocket handshake, so the front end must never connect directly and the correct shape is a server holding the secret. The article then walks through a minimal streaming loop and an event-driven WebSocket skeleton, opening a long connection, receiving session.created, pushing audio and draining events, covering speaker diarization with voice reproduction, same-frame bilingual output and long-context disambiguation, and closes with concurrency limits, cost accounting, observability, failure fallbacks and a pre-launch checklist. Pricing, rate limits, concurrency and regional availability that are not public are all marked as following the Qwen AI platform and Alibaba Cloud Bailian documentation rather than invented.

Sep 21, 20267 min read
Field SOP

One Sentence to a Live App With a Database: Qoder Sites SOP

A hands-on SOP for building and shipping with Qoder Sites from a single sentence: version checks (desktop v0.3.3 or newer, CLI v1.1.54 or newer) and the /sites entry, a five-part prompt template (goal, data model, interaction, style, deploy now) with two copy-ready build prompts, first publish covering preview, sharing and permissions, three post-publish permission checks, three persistence checks for the auto-provisioned database plus a checklist form, then binding Cloud Agents in one sentence to give the page agent ability (including when to bind and when not to). It closes with three landing scenarios (prototypes, dashboards, campaign pages), four boundaries to avoid, and a four-pitfall quick check. Database type, capacity, billing and regional rules that are not public are all marked as following the client interface and official announcements rather than invented.

Sep 20, 20267 min read