Hey Product Hunt I m Parsons, founder of PodcastorAI.
Making a video podcast is still painfully slow. Even when your content is ready, you may need a camera, lighting, a cohost s schedule, captions, visuals, and hours of editing. One hour of finished content can easily take around five hours to produce.
That s why we built PodcastorAI.
Osaurus
I can't believe (or maybe I can?) how much work it takes to produce a video podcast... (trust me, I'm trying!).
Honestly, I think @jackbogdan and I are ready to outsource our visages to AI and just focus on the banter and chitchat.
I'm ready to give @PodcastorAI a try for our next episode!
PodcastorAI
@jackbogdan @chrismessina This means a lot coming from you, Chris 🙏 And yes — 'keep the banter, outsource the visages' is exactly the point. Whenever you and Jack are ready, send a photo each and we'll build your twins so your next episode is basically just the two of you talking, minus the camera setup. Would love to see how it turns out.
PodcastorAI
@jackbogdan @chrismessina Thank you so much for hunting us, Chris! 🙌 Since you’re experiencing the video podcast workflow firsthand, your feedback will be especially valuable to us. We can’t wait to see what you and Jack create with PodcastorAI—please send any questions or honest feedback our way as you try it!
PodcastorAI
@tehreem_fatima5 Thank you so much, Tehreem! 🙏 This means a lot.
Damn, nice one. Congrats on the launch. “NotebookLM can write a podcast; Podcastor puts you on screen” is a very clear way to explain it. The digital twin side is interesting, but I’d want to know how much control you have over the final performance. Can you adjust pacing, tone, pauses and host reactions after generation, or do you mostly accept what comes back? That feels important for educators and experts in particular. A technically polished avatar is useful, but only if it still sounds like the person behind the content.
PodcastorAI
@os_ishmael Thanks, and that’s a great point. If you use Podcastor’s TTS, you can precisely adjust pacing and add pauses to shape the delivery. For uploaded NotebookLM audio, we can extract the script so you can revise the spoken content rather than simply accepting the original output.
Fine-grained control over host reactions isn’t available yet, although the current lip-sync is already very strong. We’re now developing digital twins based on real recorded video, along with a new model that drives emotions and body language from the tone and context of the content. This is especially important for educators and experts, where preserving the creator’s authentic delivery matters as much as visual polish.
The digital twin angle solves the camera problem, but the workflow after the episode seems just as interesting for creators. When someone turns one podcast into YouTube, clips, and social posts, how do you preserve the speaker's intent while adapting the edit and framing for each format? That is where a fast production shortcut either creates leverage or just creates more review work.
PodcastorAI
@wesc That’s exactly the balance we’re thinking about. Repurposing should create leverage, not add another layer of review work. We’re currently developing AI clipping so creators can turn a full episode into platform-ready short-form content while preserving the speaker’s original context and intent. It’s not available yet, but we’re working to bring it to Podcastor soon.
@parsons_wu_real That sounds like the right product boundary. Preserving the source context while adapting the cut is more valuable than simply producing more clips, especially when the creator still needs to stand behind the final message. I will be interested to see how you expose that context in the editing workflow when the clipping feature ships.
the two-host episode format is the part I'm curious about. with one AI host it's just pacing/lip-sync against your own audio, but with two, is there actual back-and-forth timing between them, interruptions, reacting to what the other one just said, or is each host's track generated separately and then interleaved so it reads more like alternating monologues than a real conversation?
PodcastorAI
@galdayan Right now, the two hosts speak in turns, so there aren’t overlapping lines or natural interruptions yet. The conversation is structured as a real back-and-forth across formats like Deep Dive, Debate, and Storytelling, rather than simply stitching together two unrelated monologues. We’re already developing overlapping speech, which should make reactions and conversational timing feel much more natural, and we expect to support it soon.
@parsons_wu_real makes sense as the sane starting point. once overlapping speech ships, how are you deciding who "wins" an overlap - is that baked into the script generation ahead of time (the model writes in an interruption on purpose), or does it need actual real-time turn-taking logic that reacts to timing as the audio is generated? feels like the second one is a much harder problem than the first.
30 minutes for a single video is genuinely surprising — that alone covers most real episodes. One thing I'm curious about, though: most video-generation models I've used are pretty careful about how they handle a real person's likeness as reference. Where does Podcastor draw that line? Concretely — could I upload a photo of Elon or Tim Cook and bring them on as an "AI guest" on my show, or is likeness locked to the account owner / consented identities only? For a product built on digital twins, that consent boundary feels like it matters as much as the tech itself.
PodcastorAI
@kryptonite_wei You’re absolutely right that the consent boundary matters as much as the technology. Podcastor requires the person’s explicit consent before their likeness can be used to create a digital twin. So no, you couldn’t simply upload a photo of Elon Musk or Tim Cook and add them as an “AI guest.” We also use celebrity detection to help prevent unauthorized use of public figures and protect everyone’s likeness rights.
Congrats! Biggest feedback: get away from TTS voices that sound like they live in the San Francisco bubble trying to sound “nice” and “safe.” It's tiring to hear ai voices doing with constant uptalk, smile voice, forced reassurance, over-enunciation, dramatic pauses, or sometimes exaggerated “authority.” Add voices from other eras, cultures, cities, classes, and subcultures. ADD PLAIN SOUNDING VOICES that are earthy and grounded without sounding like they are "smiling"! Until then the human psyche gets really fatigued fast by listening to this stuff
PodcastorAI
@archvalmiki This is thoughtful feedback, and we agree that “polished” AI voices can quickly become tiring when they all share the same cadence and personality. Authenticity requires much more variety in accent, rhythm, age, culture, and social context, including voices that simply sound plain, grounded, and human. Today, creators can also upload their own recorded audio to preserve their natural delivery completely. We’re taking this feedback seriously as we expand our TTS options.
JoggAI
An MCP integration would be really useful for sending a transcript from another tool directly into PodcastorAI.
PodcastorAI
@rhinogo You're reading our minds 👀 We've actually been building an MCP server for exactly this — pipe a transcript straight from your tool of choice into Podcastor, no copy-paste tango. Consider your idea officially on the roadmap (and mostly already there 😏)
PodcastorAI
@rhinogo Great suggestion! This is exactly the kind of seamless workflow we want to build—moving content from the tools you already use directly into PodcastorAI. Which tools would you most like us to support first?