How to create an AI podcast from research notes
Convert sourced notes into an evidence-led dialogue script, generate bounded audio batches and keep editorial, voice-rights and publication decisions human.
By Scrollport

In brief
Start with a source ledger, not a request to improvise a podcast. Ask an agent to create a claim-by-claim outline, write a dialogue script that preserves citations, assign voices you have the right to use and generate the audio in bounded batches. Disclose synthetic voices where appropriate and require a human to review factual accuracy, pronunciation, pacing and rights before publication.
To turn research notes into a multi-speaker podcast, keep a source ledger through the whole workflow. Ask an AI agent to outline claims, write a dialogue that cites those sources, assign voices you have permission to use and generate the audio in bounded batches. A human should approve the script and listen to the complete programme before publication.
Start with a source ledger
Give every research note a source ID, title, author or organisation, URL, publication date, relevant excerpt and the claim it supports. Mark opinion, prediction and disputed material explicitly. Notes without a checkable source may still inspire questions, but they should not silently become factual narration.
Decide the audience, episode promise, length and editorial boundary before scripting. A useful source ledger also records permissions for quoted or licensed material and flags anything that needs a human legal or editorial decision.
Write a claim-led dialogue
Build an outline in which every segment answers part of the episode promise. Assign one speaker to guide the argument and another to question, clarify or introduce an alternative view. Give factual lines source markers during drafting, even if the final programme moves citations into show notes.
Do not ask the model to create quotations, anecdotes or expert opinions for dramatic effect. Where the sources disagree, let the dialogue explain the disagreement. Read the script aloud before generation to catch unnatural exposition, repeated context and sentences that look clear on screen but are difficult to follow in audio.
Choose voices and disclose synthesis
Use synthetic or licensed voices that fit the format without impersonating a real person. If a voice is cloned, confirm that the necessary rights and consent exist. ElevenLabs’ voice-cloning guidancedescribes its verification and usage restrictions; those product safeguards do not replace the publisher’s own rights assessment.
Tell listeners when voices are synthetic in the episode, show notes or both, according to the context and applicable rules. Do not present generated speakers as real interviewees, witnesses or named experts.
Generate audio in bounded batches
Use the dialogue generation capabilityand inspect the current Eleven v3before generating. The current ElevenLabs Text to Dialogue API referencerecommends no more than 2,000 total characters across each request. The same limit is part of Scrollport’s current public tool contract. Split a longer script at natural segment boundaries and keep the speaker-to-voice mapping stable across batches.
Number batches and preserve their script range, voice IDs, generation result and file order. Generate a short sample first to check pronunciation and speaker contrast. A successful API response proves that audio was created; it does not prove the edit is factually or editorially ready.
Review before publishing
Create a multi-speaker podcast from the supplied research ledger.
Audience: [audience]. Episode promise: [one sentence]. Target length: [length].
First produce an outline and claim ledger. Every factual line must map to a
source ID; do not invent quotations, interviews, anecdotes or expert views.
Preserve disagreements and flag unsupported notes for human review.
Write a two-speaker script using voices we have the right to use. Include a
synthetic-voice disclosure. After script approval, use Scrollport discover and
inspect to select the current dialogue tool, then generate batches of no more
than 2,000 total characters at natural section boundaries.
Return the script, source ledger, batch manifest, disclosure and a human QA
checklist. Do not publish or distribute the audio.The final human listen should verify claims and citations, names and numbers, pronunciation, speaker identity, pacing, batch joins, levels, disclosure and rights. Correct the script or regenerate a bounded segment rather than treating the first nondeterministic generation as final; ElevenLabs itself notes that several generations may be needed in its capability guidance.