How to Prepare Audio So Transcription Is Worth Reading
Mic distance, noise, sample rate, and language settings that make speech-to-text useful — and the edits to do locally before you spend a cloud credit.
Updated 2026-10-05.
The transcript is only as clean as the take
Speech-to-text does not invent a studio recording. It guesses words from the sound you give it. A phone on a table in a cafeteria, two people talking over each other, and a Bluetooth headset with dropouts will produce a transcript you have to rewrite. Ten minutes of setup — closer mic, one speaker, a quieter room — saves more editing time than any prompt tweak.
On Vocix, transcription is a cloud feature with a daily free quota. Trim and denoise, when you use them as ordinary audio edits, can run locally first. Spend the cloud credit on the passage you actually need, not on twenty minutes of room tone and “is this recording?”.
What to fix before you upload
Cut the file down to the speech. A local trim from the first real sentence to the last one removes the long silent head and tail that add cost and errors. If the track has a loud hum, a denoise pass can help mild hiss. It will not repair clipping, where the waveform is already flattened because the mic was too hot. Clipped audio stays clipped.
One speaker at a time is easier than a debate. If you recorded a conversation, say the names out loud at the start or plan to label speakers yourself afterward. Most general transcription does not reliably write a screenplay with character names unless the product explicitly diarizes speakers.
Pick the language you actually spoke. Asking for English on a Hindi interview, or the reverse, produces confident nonsense. If the recording switches languages, transcribe the sections separately when you can.
- Listen once on headphones and note the useful range.
- Trim locally to that range and export or keep the shortened audio.
- Denoise only if the noise is steady and the voice is still intelligible.
- Set the transcription language to match the speaker.
- Proofread names, numbers, and anything you would be embarrassed to publish wrong.
Sample rate, mono, and file size
Speech does not need music-master settings. A 16-bit mono file around 16 kHz or 44.1 kHz is plenty for words. Stereo interview files are fine too; you do not have to downmix unless the size is the problem. What hurts is a tiny bitrate from a voice-note app that already sounds like a phone in a pocket, or a file so long that you included the whole meeting when you needed one answer.
MP3, M4A, and WAV are all reasonable inputs. WAV is larger and closer to the original. MP3 is smaller and usually good enough for speech if it was encoded at a sane bitrate. If the voice already sounds watery, converting to WAV will not put the missing detail back.
Keep confidential recordings off cloud transcription unless you are allowed to upload them. Local trim and export do not require that upload. Medical, legal, and client audio often belong in offline software under a written policy, not in a free web quota.
How to read the result
Treat the transcript as a draft. Models mis-hear names, medicine, course codes, and street addresses. Read it against the audio for any sentence you will quote. Punctuation is a guess. Paragraph breaks are a convenience, not a legal record of who paused.
If the text is wildly wrong, do not immediately spend another credit on the same file. Check language, trim out the worst noise, and confirm you uploaded the audio track rather than a music bed. A second pass helps when the first file was the wrong clip. It rarely helps when the recording itself is unintelligible.
For a voiceover you will generate rather than record, write the script first and use text-to-speech, then trim the result. That path is the opposite of transcription: text in, audio out, still quota-limited, still something you should listen to before you publish.