Podcast Distribution
Turn articles or written text into audio episodes and distribute them to Spotify and Apple Podcasts
Your articles become episodes
From an article you've written, Sorabun can generate the script, the audio, and the RSS feed for distribution. No recording or editing required. You can also create an episode without an article. Just write the text on the spot.
Steps
- Set your text-to-speech API key under "Settings > Podcast Generation (TTS)" (Fish Audio or OpenAI)
- In "Sorabun > Podcast Distribution", use "Create episode" to pick an article or write text
- On the same screen, enter the show information (title, description, cover image, category)
- Register the generated RSS feed URL with Spotify, Apple Podcasts, and other platforms
Creating from text
Sometimes you want to put out an episode that isn't quite worth turning into an article (an announcement, an answer to a question you received, a seasonal greeting). For episodes like these, choose "From text" in "Sorabun > Podcast Distribution > Create episode" and write a title and the text.
- Just write down what you want to say. Bullet-point notes work fine too, since the AI restructures it into a spoken-language script
- 100 to 10,000 characters. The episode length is set to match how much you write
- You can give this one episode its own cover art (if not set, the show's cover is used)
- You can also switch the speaking style (solo narration / dialogue) for just this episode
The finished episode is listed alongside episodes made from articles and distributed through the same RSS feed. Layering BGM works the same way too.
When creating from an article, the length is fixed at 8-12 minutes, because articles tend to run a similar length.
When creating from text, the text might be only 300 characters. If you told the AI "make it 8-12 minutes" anyway, it would start saying things that were never written just to fill the time, and listeners have no way to tell that from the real thing. So instead we match the length to the text, and also tell the script writer: "don't add anything that isn't in the original text; if there isn't enough material, let the episode end short."
They're stored in a dedicated container that doesn't show up on the site or in the admin post list. Since they aren't articles, showing them in a list or site search would lead readers to an "article" they can't actually open.
Because of this, you also delete them from "Sorabun > Podcast Distribution." Use "Delete" on each row to remove the episode along with its audio. You can check the original text at any time from "View original text" on the same row.
Creating from an existing article
Choose "From article" in "Create episode" to pick from the site's posts and pages and turn one into audio (the 200 most recent, newest first). This covers not only articles Sorabun created but also articles you published before. If the article has a featured image, it's used as that episode's cover art.
Detailed settings
Under "Settings > Podcast Generation (TTS)" you can fine-tune how the show is produced.
| Item | Setting key |
|---|---|
| Text-to-speech service | podcast_provider |
| Speaking style (solo narration / dialogue) | podcast_mode |
| Show name | podcast_show_name |
| Opening/closing boilerplate | podcast_opening / podcast_ending |
| Speaker names | podcast_speaker_a / podcast_speaker_b |
| Fish Audio voices | fish_voice_a / fish_voice_b |
| OpenAI TTS voices | openai_voice_a / openai_voice_b |
| Extra instructions for the script | podcast_custom_prompt |
| Add BGM (jingle) | podcast_bgm |
| How BGM is layered (lead-in, tail, overlap, ducking amount) | podcast_mix_lead / podcast_mix_tail / podcast_mix_overlap / podcast_mix_duck |
marin and cedar sound the most natural. alloy and echo are older-generation voices and sound flat.
What you write in "Extra instructions for the script" is passed not only to the script writer but also directly to the narration AI's delivery. Instructions like "speak in a calm, lower voice" or "don't talk too fast" take effect as written, so if a reading sounds flat, add a note here.
Choosing a speaking style
- Solo narration: a narration format, suited to shows that convey information plainly
- Dialogue: a back-and-forth between two speakers, easy to listen to and suited to longer content
Showing a player on the article page
Enable "Add player to article" in the show settings and an audio player appears below the article. Readers can then choose to read or listen.
Once you register the RSS feed URL with a platform, future episodes are picked up automatically.
BGM (jingle)
"Settings > Podcast Generation (TTS) > Add BGM (jingle)" attaches BGM before and after the narration. It's on by default.
Just playing the same jingle at the start and end of the show makes it feel a lot more like a real program. No extra tools are needed, so it works on shared hosting too, and it's applied automatically to scheduled posts as well.
Choose the BGM from the media library under "Settings > Podcast Generation (TTS)." You can set separate opening and closing tracks, and you'll get a warning on the spot if the length or format doesn't fit.
This is as far as automatic processing goes. Since it just attaches BGM before and after, it does not lay BGM under the voice (ducking).
Layering BGM under the voice (manual)
If just attaching BGM isn't enough, press "Layer BGM" for that episode and it re-mixes the BGM under the voice. The button sits next to "Play" on each row of the episode list in "Sorabun > Podcast Distribution" (it only appears once BGM is set).
Here's what happens when you press it.
- The opening BGM plays alone, then automatically dips when the voice comes in (ducking)
- It returns to its original volume once the voice ends, then fades out
- The closing BGM overlaps briefly with the end of the voice before it plays
- Because the BGM is converted to match the narration's format, you don't get the brief break at the seam that can happen with the simple attach-only version
You control how it's layered under "Settings > Podcast Generation (TTS) > How BGM is layered."
| Item | Meaning | Default |
|---|---|---|
| Lead-in | Time from when the BGM starts to when the voice comes in | 2 sec |
| Tail | How long the BGM lingers after the voice ends | 10 sec |
| Overlap | How long the closing BGM overlaps the end of the voice | 2 sec |
| BGM under the voice | How many dB to lower the BGM while the voice is playing | 20 dB |
Since we can't assume the server has audio-processing tools (ffmpeg) installed, the audio is assembled in the browser you have the admin screen open in. For a 20-minute episode this takes about 1-2 minutes, and you need to keep this screen open the whole time.
If you close it partway through, the audio currently live is untouched. The replacement only happens once the export finishes completely and the server has received it, so you can always try again later.
This button only repaints an episode that's already finished. Nothing happens unless you press it. Automatic article generation and scheduled posting still complete and publish, with BGM attached before and after, exactly as before.
Before v2.20.0, "audio finishing" worked the other way around: finishing interrupted automatic generation partway through, and episodes got stuck "waiting to finish" without BGM, so scheduled posts would go live without BGM. Because of that, we removed it once already.
The narration audio from before layering is kept, so you can change the layering settings and redo the finish as many times as you like (no need to redo text-to-speech, and it costs nothing). Since the audio URL doesn't change when it's replaced, your published RSS feed keeps working as-is.
Layering requires converting the MP3 back into a waveform and re-encoding it to MP3. Simple attaching involves no conversion, so this is the one place where you lose a generation of quality. Listen to both and decide whether it's worth the trade for ducking.
For episodes made before v2.28.0, it can be unclear where the BGM starts and ends (for example, if the BGM was reconfigured afterward). Pressing the button on those episodes shows "Cannot layer," so re-generate the audio first and then try again.
About volume
Sorabun doesn't measure and normalize loudness (it lowers the level only enough to avoid clipping when layering). That's because Spotify and Apple Podcasts each re-normalize loudness on playback, to roughly -14 LUFS and -16 LUFS respectively. Since the platform does the adjusting, matching it precisely on our end wouldn't show up in the result. It would only add processing weight and more settings.
If you want to dial in the volume precisely, run the exported MP3 through an audio-finishing service like Auphonic or your own editing software. Replace the file in the media library and your published RSS feed keeps working as-is.
We haven't brought back the server-side ffmpeg finishing path. Shared hosting typically doesn't have it installed, which would break the main path, and maintaining it alongside the browser-based version would leave no way to tell which pipeline actually produced a given file.