Your agent writes it. Vydra voices it.
Give a content assistant a path from a written script to generated speech. Voiceovers can sit alongside image and video tasks in the same account, with the agent retrieving the audio result for review or another production step.
Narration, spoken drafts, and product walkthrough voiceovers from an approved script.
Speech generation is useful when an agent already has the context for a tutorial, product demonstration, or story. It removes the need to manually transfer that script between tools. The API creates speech; timing it to a video and assembling the final edit are separate tasks unless you explicitly connect them in your production process.
- 01
Prepare the spoken script
Have the agent write for listening: shorter sentences, clear pronunciation, and the intended audience. Review factual claims and names before spending credits on the final recording.
- 02
Choose the supported voice
Read the speech workflow documentation for voice options and input limits. The generate_speech workflow accepts the script in input.prompt and an optional voice_id.
- 03
Generate and retrieve audio
Create the speech job, retain its identifier, and retrieve the result through the job endpoint. Read audioUrl from the completed result rather than treating the submission response as an audio file.
- 04
Connect voice to production
Pass the audio URL to your own editing process or a compatible next step. OpenClaw users can also use Vydra’s native speech integration after configuring the provider and API key.
Write a 30-second narration for this approved product walkthrough. Let me review the script, then generate a voiceover with Vydra and return the audio URL. Keep the visuals and narration in the same project brief.
Adapt this instruction to your agent, account, and approval preferences.
Go from idea to implementation.
Can my agent create voiceovers with the same API key?
Yes. Supported speech, image, and video workflows use the same Vydra account and API authentication. Each generation uses credits according to its workflow and inputs.
Does this automatically edit narration into my video?
The speech workflow produces audio. Your application or a compatible production workflow handles assembly with video.