Browse tutorials

Speech-to-Text & Text-to-Speech

Speech to Text and Text to Speech use the same create, result, and history pattern, but they solve opposite tasks:

  • Speech to Text starts with a recording and produces a transcript.
  • Text to Speech starts with a script and produces spoken audio.

Choose the first when you need searchable, reusable words from audio. Choose the second when written material needs a voiceover, narration draft, character line, lesson, product demo, or spoken prototype.

Choose the correct direction

Starting materialGoalToolResult
Meeting, interview, podcast, voice note, or songRead, search, quote, or caption the speechSpeech to TextFull transcript, timestamped segments, and TXT download
Script, narration, dialogue, announcement, or lessonHear and export a spoken performanceText to SpeechPlayable and downloadable audio

The transcript or generated speech is saved as a separate history item. The source song or recording is not replaced.

Speech to Text converts a song or upload into a transcript, while Text to Speech converts an edited script into spoken audio.
Choose the direction from the material you already have.
The live Speech to Text workspace with source audio and transcript settings on the left and a timestamped speaker transcript on the right.
Select or upload the source on the left, then review the generated transcript and download control on the right.

Speech to Text

Create a transcript

  1. Open Speech to Text.
  2. Choose My songs or Upload audio.
  3. Select a completed playable song, or upload a supported audio file under 50 MB.
  4. Preview the source and confirm that it is the intended recording.
  5. Open More options and choose the language, speaker, sound-note, channel, and timestamp settings.
  6. Review the duration-based credit estimate.
  7. Select Create transcript.
  8. When processing finishes, review both Full transcript and Timestamped segments.
  9. Select Copy or Download TXT after correcting the text in your working document.

The upload field accepts common audio formats shown in the interface, including MP3, WAV, M4A, FLAC, OGG, and WEBM. A file with a supported extension can still fail if it is empty, corrupted, larger than 50 MB, or encoded as non-audio content.

Prepare audio for a cleaner transcript

Transcription is most reliable when the words are easier to hear than everything around them.

Better sourceHarder source
Speakers close to a microphoneDistant room recording
Stable volume without clippingVery quiet speech or distorted peaks
Limited overlap between speakersSeveral people speaking at once
Low music and background noiseLoud music under dialogue
Complete recording with no missing sectionsAbrupt cuts, dropouts, or long corrupted areas
Separate channels for isolated speakersA multichannel file with mixed speakers on every channel

Listen to the beginning, a middle section, and the ending before submitting. If speech is only on one side of a stereo recording or the wrong file was selected, fix the source first rather than relying on transcript settings.

Understand every transcript option

ControlWhat it changesWhen to use it
My songsUses a completed playable song from your libraryLyrics, spoken-word songs, podcasts, or audio already saved in the product
Upload audioUses a local audio fileMeetings, interviews, voice notes, lectures, and external recordings
LanguageAuto-detects or specifies the spoken languageSet it manually when the recording is short, accented, noisy, or repeatedly detected incorrectly
Identify speakersAdds speaker identities to conversational speechOne mixed recording containing two or more speakers
Number of speakersUses automatic count or a selected count from 2–6Set a known count when automatic speaker detection splits one person or merges different people
Include sound notesIncludes non-speech events such as audible reactions or musicCaptions, interviews, accessibility review, and scene logging
Separate audio channelsTreats channels as separate sourcesCalls or studio recordings where each speaker is isolated on a different channel
TimestampsChooses none, word-level, or character-level timingSelect only the timing detail required by the next workflow

Identify speakers or separate channels

These two settings cannot be active together.

  • Use Identify speakers when several people share one mixed microphone or mono/stereo mix.
  • Use Separate audio channels when the recording already places each participant on a different channel.
  • Use neither for a clean single-speaker recording when speaker labels are unnecessary.

If you turn on one option, the interface turns the other off. Speaker count is only available while Identify speakers is enabled.

Choose the timestamp detail

SettingResultGood choice for
NonePlain transcript without detailed timingNotes, summaries, search, and article drafts
WordEach recognized word carries timing dataCaption preparation, quote location, and audio review
CharacterTiming for individual charactersSpecialized alignment workflows that truly need it

Word timestamps are the practical default for most timed work. More detail does not make the words more accurate; it only makes each time marker more precise.

Use sensible option combinations

One-person voice note

  • Language: auto-detect or choose the known language
  • Identify speakers: off
  • Include sound notes: off unless reactions matter
  • Separate audio channels: off
  • Timestamps: None for notes, Word for review

Two-person interview on one recorder

  • Identify speakers: on
  • Number of speakers: 2
  • Include sound notes: on when laughter, applause, or music matters
  • Separate audio channels: off
  • Timestamps: Word

Remote call with one person per channel

  • Identify speakers: off
  • Separate audio channels: on
  • Timestamps: Word

Song lyrics or spoken-word audio

  • Choose the known language when possible
  • Turn speaker identification off unless distinct performers must be separated
  • Keep sound notes on if instrumental or non-speech cues are useful
  • Use Word timestamps when locating lyrics in the audio

Understand credits and processing

Credits are estimated from the source duration. The current rate is 4 credits per hour, calculated from the recording length with a 1-credit minimum. The number beside the button is the final estimate for the selected source.

After submission, the result area shows a processing state while the recording is analyzed. Do not submit the same file repeatedly because the first view has not finished. If the request fails, the task follows the normal refund path.

Review and export the transcript

The completed result shows the source name, detected language, word count, speaker count, full transcript, and available timestamped segments. Use the audio and timestamps together when checking uncertain passages.

Before publishing, verify:

  • names, product terms, acronyms, dates, numbers, and URLs;
  • punctuation and paragraph boundaries;
  • which speaker said each sentence;
  • overlapping or interrupted speech;
  • sound notes that are irrelevant, missing, or placed incorrectly;
  • words near music, applause, laughter, or background noise;
  • the beginning and ending, where clipped syllables are easy to miss.

Copy places the plain transcript on the clipboard. Download TXT creates a text file that includes available timestamped segments. Editing a downloaded copy does not update the saved history result.

Fix common transcription problems

ProblemWhat to do
The wrong language was detectedChoose the spoken language manually and transcribe again
Two speakers were mergedEnable Identify speakers and set the known speaker count
One speaker was split into several labelsSet the correct speaker count or turn speaker identification off
Channel separation produces confusing outputUse it only when people are truly isolated by channel; otherwise use Identify speakers
Names or specialist terms are wrongCorrect them during review; use cleaner context around the term in the source when possible
Words are missing during overlapReduce overlap in the source edit or treat the result as a draft for manual correction
The transcript contains too many bracketed soundsTurn off Include sound notes when non-speech events are not needed
The upload is rejectedConfirm it is audio, not empty, in a supported format, and under 50 MB
The estimate cannot be calculated immediatelyWait for the browser to read the upload duration before submitting

Text to Speech

The live Text to Speech workspace with script, voice, voice settings, credits, and Create speech controls beside the preview and voice cards.
Edit the script and choose the voice on the left; use the right-side preview to compare voices before generation.

Create spoken audio

  1. Open Text to Speech.
  2. Enter or paste the script in Enter Text.
  3. Keep the text at or below 5,000 characters.
  4. Open Select voice, preview suitable voices, and choose one that fits the intended language and delivery.
  5. Choose a Download format.
  6. Open More options to adjust Speaking speed, Voice stability, and Enhance speaker clarity.
  7. Check the character-based credit estimate.
  8. Select Create speech.
  9. Play the complete result, then download it or revise the script and generate a new version.

The trash button clears only the current text field. It does not remove completed speech from History.

Choose the voice before fine-tuning settings

Voice choice has a larger effect than small slider changes. Preview several voices with your intended audience in mind:

  • accent and language fit;
  • perceived age and vocal character;
  • calm, energetic, intimate, authoritative, or conversational tone;
  • clarity on names and specialist vocabulary;
  • suitability for narration, dialogue, product demos, or learning content.

A good voice for a calm documentary may not suit a fast interface demo. Preview samples help narrow the list, but generate a short line from your actual script before committing to a long section.

Format the script for speech

The engine reads the text you provide. Spelling, punctuation, capitalization, paragraph breaks, and context all influence the result.

Write for the ear

Written version:

Our platform supports rapid iteration, collaborative review, asset management, and export, enabling teams to move efficiently from an initial concept to a finished result.

Spoken version:

Start with an idea. Create a first version, review it with your team, and keep every asset in one place. When it is ready, export the final result.

Shorter sentences give the voice clearer breathing points and make mistakes easier to isolate.

Use punctuation to guide pacing

  • Commas create short separations inside a sentence.
  • Periods and paragraph breaks create stronger boundaries.
  • Question marks help mark a question's intonation.
  • An em dash can introduce a deliberate turn or interruption.
  • Ellipses can suggest hesitation, so use them only when that delivery fits.

Do not fill a script with repeated punctuation. If a pause must be predictable, split the material into shorter generations and place the clips on a timeline.

Prepare names, numbers, and symbols

Unusual names, acronyms, URLs, emoji, and dense symbols can be read unpredictably. Write the spoken form when accuracy matters:

Written for displaySafer spoken version
2048two thousand forty-eight
8:30 PMeight thirty P M
5 kmfive kilometers
A/B testA B test
a difficult proper namea clear phonetic respelling that sounds correct when read aloud

Generate a short pronunciation test before using the term throughout a long script.

Add performance direction carefully

Supported bracketed cues can influence delivery, for example:

[whispering] We should keep this part quiet.
[excited] The final mix is ready!
Well... [sigh] let's try that one more time.

The selected voice still needs to suit the requested performance. Use a small number of cues, keep them close to the relevant line, and test them in a short passage. Do not put production notes in the text unless you want them interpreted as performance direction or possibly spoken.

Understand every speech control

ControlAvailable settingEffect
VoiceAvailable previewable voicesChanges the core speaker identity, accent, and character
Download formatMP3 standard, MP3 high quality, MP3 small file, or Opus high qualitySets the audio encoding for the new result
Speaking speed0.70×–1.20×Values below 1.00 slow delivery; values above 1.00 speed it up
Voice stability0–1Lower values allow more expressive variation; higher values favor consistency
Enhance speaker clarityOn or offAsks for stronger speaker similarity and clarity; the audible change can be subtle

Tune speed

Start at 1.00×. Lower the value for dense instructions or a deliberate narration. Raise it slightly for short announcements or energetic content. Extreme values can reduce naturalness, so edit the sentence length before pushing the slider to an endpoint.

Tune stability

Start from the default 0.50.

  • Lower stability when the read is too flat and the script needs more emotional variation.
  • Raise stability when pacing, tone, or pronunciation changes too much between sentences.
  • If the voice becomes monotone, reduce stability slightly or choose a voice with a delivery closer to your goal.
  • If the voice becomes chaotic, noisy, or inconsistent, raise stability and simplify performance cues.

The same text and settings can still produce small differences between generations. Save any take you want to keep before experimenting further.

Use speaker clarity

Keep Enhance speaker clarity on when preserving the selected voice identity is more important than the smallest possible processing time. Turn it off and compare when the result sounds strained or when you are testing whether the setting is responsible for an artifact. The difference may be subtle, so compare the same short sentence.

Choose the download format

FormatBest forTradeoff
MP3 standardGeneral voiceovers and previewsBalanced size and quality
MP3 high qualityFinal compressed deliveryLarger file
MP3 small fileFast sharing and prototypesMore audible compression
Opus high qualityModern apps and playback systems that support OpusConfirm editor and platform compatibility

Choose the format before generation. An existing history item keeps the format that was created at the time.

Understand text-to-speech credits

The current estimate uses the trimmed character count: 10 credits per 1,000 characters, calculated proportionally and rounded up, with a 1-credit minimum. The 5,000-character maximum would therefore display up to 50 credits. Check the live estimate after your final script edit.

For long material, generate a short test first. Once the voice, pronunciation, and settings are correct, split the script at natural paragraph or scene boundaries. This makes revisions smaller and avoids paying to recreate an entire long passage because of one line.

Review and refine the result

Listen with the script visible and check:

  • every name, number, acronym, and technical term;
  • pauses at commas, sentences, and paragraph breaks;
  • changes in accent, volume, tone, or speed;
  • unexpected breaths, clicks, noise, or extra speech;
  • emotional cues and whether they fit the surrounding lines;
  • silence at the beginning and ending;
  • whether the voice still sounds clear under music or sound effects.

When one sentence fails, rewrite or generate that section separately. Keep the good take and assemble approved sections in AudioMass – Audio Editor or another timeline instead of repeatedly regenerating everything.

Fix common speech problems

ProblemWhat to change
The voice is too flatChoose a more suitable voice, lower stability slightly, and improve punctuation or performance cues
The delivery is too chaoticRaise stability, remove excess cues, and shorten the passage
A name is mispronouncedWrite a phonetic alternative and test it in one short sentence
Numbers or symbols sound wrongSpell out the spoken form instead of using digits or symbols
The voice becomes too fast or slow mid-passageSplit long text, simplify punctuation, and move speed closer to 1.00×
The accent changesChoose a voice suited to the script language and generate shorter sections
Extra noise or speech appearsRaise stability, remove conflicting cues, and retry a shorter passage
The Create button is disabledAdd non-empty text and keep it within 5,000 characters
A voice preview does not playTry another preview, check browser audio permission, then select and run a short generation

Use History without losing useful work

Each tool has its own History. Search by title or content, switch between newest and oldest order, open earlier results, copy or download them, load more items, and delete unwanted entries. Deleting a history item is different from clearing the current input field.

For sound design around a generated voice, continue with Generate Sound Effects. To edit, layer, and export completed audio, see Audio Editing, MIDI & Exports.