
Interview transcription turns recorded speech into text you can quote, code and search. The decision that shapes everything else is the style: full verbatim, clean verbatim or edited. Pick that first, format the transcript consistently, then check the AI draft against the audio in two passes.
This guide covers each step, plus the parts most guides skip: privacy and anonymization, and interviews where people switch languages mid-sentence.
Choose your interview transcription style first
There are three common styles, and the right one depends on what you will do with the text. Choosing after you have transcribed means redoing work.
Otter's guide to interview transcripts describes the three styles this way:
| Style | What it keeps | Typical use |
|---|---|---|
| Full verbatim | Filler words ("um", "uh"), false starts, overlapping speech, pauses, non-verbal cues | Conversation analysis, discourse research, legal depositions |
| Clean (intelligent) verbatim | All meaningful content and non-verbal cues; removes fillers, repetitions and false starts | Journalism and general qualitative research |
| Edited | Rewrites phrasing, polishes grammar, reshapes sentences | Internal synthesis, UX reports, public-facing articles |
If you are unsure, start with clean verbatim. It stays faithful to what the person said and is readable enough to code or quote. You can always move toward edited later. You cannot recover pauses and false starts from a transcript that never recorded them, unless you go back to the audio.
One rule across all three: write down which style you used in the transcript header. A reader who sees a quote with no "ums" should know whether that is the speaker's style or your edit.
Match the style to your research method
Your methodology, not your preference, decides how much detail the transcript needs. ParrotNotes' qualitative guide maps it out:
- Content and thematic analysis need the participants' actual words with light notation.
- IPA and narrative interviews need hesitations and emotional cues captured, because those carry meaning for the analysis.
- Conversation analysis demands full Jefferson notation: timed pauses, overlaps, latching and pitch.
For journalists and UX researchers the question is simpler. Will anyone read the transcript as evidence, or only as a source for quotes and findings? If it is evidence, stay closer to verbatim. If it feeds a report, clean verbatim plus timestamps back to the audio is usually enough.
Be careful with a mixed approach. If most of your interviews are clean verbatim and a few are full verbatim, your coding will treat hesitation inconsistently. Pick one standard per project.
A workflow from recording to first draft
Good transcripts start with good audio. An AI draft from a clean recording needs far less repair than one from a noisy room.
Atter AI's interview guide lays out the AI workflow: record clear audio in a quiet space, upload the file, apply automated speaker diarization, then verify against the audio. Here is that sequence with the extra steps we would add:
- Set up the recording. Use a quiet room, put the microphone close to each speaker and test it for a minute before you begin. For remote interviews, ask each person to use headphones so the other voice does not leak into their microphone.
- State the basics on tape. Say the date, the interview code and that the participant consents to recording. This doubles as a record and as a check that the file is the right one.
- Name the file consistently. For example,
2026-10-08_P07_interview.m4a. Use the participant code, not the person's name. - Pick the transcription language. If the interview is not in English, set the right language, or use a tool that detects it. Wrong language settings are a common cause of unusable drafts.
- Upload and generate the draft. Keep the untouched machine output as its own file.
- Check speaker labels. Diarization often confuses speakers when people interrupt or sound alike. Fix labels before you fix words, since a wrong label changes who said what.
- Run the two-pass check described below.
- Save the final version separately from the raw draft.
If you are comparing ways to get from audio to text at all, our post on how to transcribe audio to text compares five methods.
How to format an interview transcript
A proper interview transcript has header metadata, consistent speaker labels, timestamps and standard tags for unclear speech. Otter's guide lists these as the standard elements: speaker labels (initials such as SJ:, or role tags such as I: and P:), timestamps at set intervals of 5 to 10 minutes or at speaker turns, header metadata, and standardized tags for inaudibles and overlapping speech.
Header
Put the basics at the top:
- Interview code (not the participant's name)
- Date and approximate length
- Interviewer
- Language(s) spoken
- Transcription style (for example, "clean verbatim")
- Who produced and checked the transcript
Speaker labels
Pick one scheme and keep it. I: and P: works for one interviewer and one participant. With several participants, use P1:, P2: or role tags. Avoid real names if you plan to anonymize.
Timestamps
Interval timestamps are lighter to read. Speaker-turn timestamps make it easier to jump back to the audio. If you expect to quote people or check disputed passages, use speaker-turn timestamps.
Tags for unclear speech
Choose a small set and use it everywhere. For example:
[inaudible 00:14:32]for speech you cannot make out[unclear: "budget"?]for a best guess[overlap]where two people talk at once[laughs]or[long pause]where it affects meaning
Do not guess silently. A confident but wrong word is worse than a visible gap.
The two-pass quality check
Run one pass against the audio and one pass without it. Convert Audio to Text's guide describes the method: in pass one, read along while listening at 1.0x to 1.25x speed and correct names, negations and technical terms. In pass two, read the text on its own to catch contextual blind spots and formatting errors.
The two passes catch different problems. Listening catches what the machine misheard. Reading alone catches sentences that do not make sense, because you are no longer being carried along by the audio.
Where AI drafts fail
Fluent text can hide errors. These are the places to slow down in pass one:
- Negations. "Can" versus "cannot", "did" versus "didn't". A dropped negation reverses a finding. Listen again every time a sentence carries a decision, a complaint or a refusal.
- Names and organizations. Check spelling against your participant list.
- Numbers, dates and money. Verify each against the audio.
- Technical terms and acronyms. Build a short glossary and search the transcript for variants.
- Text that appears where the audio is unclear. If the transcript reads smoothly over a stretch where you hear noise, treat it with suspicion. Replace it with an
[inaudible]tag unless you can confirm the words.
Keep an audit trail
Keep the raw AI draft, the corrected version and a short change log: what you changed, where and why. For research, this lets you show that a quote reflects the recording and not the machine's guess. For journalism, it protects you if someone disputes a quotation. Even a spreadsheet with timestamp, original text and correction is enough.
Consent, anonymization and participant privacy
Treat transcripts as personal data from the moment you record. Get consent that covers recording, storage and who will see the transcript, and then reduce what the transcript reveals.
The rules that apply to you come from your ethics board (IRB), your institution, your employer or data protection law such as GDPR. Ask them before you start, not after the interviews are done. What follows is a practical baseline, not legal advice.
- Remove direct identifiers. Names, phone numbers, email addresses, exact job titles and employers can all point to one person.
- Check indirect identifiers. A small town, a rare role and a specific event together can identify someone even without a name.
- Use pseudonyms or codes. Replace names with
P07or an invented name, consistently, across all transcripts. - Store the key separately. The file linking codes to real names should live apart from the transcripts, with tighter access.
- Remember that audio is identifying. A voice can identify a person. Decide how long you keep recordings and delete them when your protocol says so.
- Check your tool's data handling. Before uploading, find out where files are stored, whether they are used to train models, and whether you can delete them.
Anonymize in a copy. Keep the original transcript under restricted access until you no longer need it.
Accents, code-switching and cross-border interviews
Most transcription guides assume one language and a standard accent. Real interviews in Southeast Asia and Nigeria often do not fit that. People switch between Burmese and English, or between English and Yoruba or Hausa, in the middle of a sentence.
Practical steps:
- Choose the right language setting. If you pick English for an interview that is half Burmese, the Burmese portions come out as gibberish or get skipped.
- Decide how to render switches. Keep each language in its original script and mark the switch if it matters to your analysis. If you add translations, put them in brackets or a separate column so the original words stay visible.
- Use a glossary of terms. Local place names, organizations and borrowed words need consistent spelling.
- Give accented speech extra time in pass one. Accent is a common cause of misheard names and numbers.
- Consider a bilingual reviewer. For anything quoted or analyzed closely, a second person who speaks both languages should check key passages.
We go deeper on this in why Burmese-English code-switching breaks transcription tools and in our guide to multilingual meeting transcription.
Where Loka Note fits
Loka Note is an AI meeting-notes tool that also accepts uploaded audio and video files, so you can use it to produce a first-draft transcript of a recorded interview. It is built Burmese-first and handles mixed Burmese-English speech, and it also supports English, Thai, Vietnamese, Chinese (Simplified), Yoruba and Hausa. The language you pick is checked against the audio, so an interview recorded under the wrong setting is transcribed in the language actually spoken.
On privacy, customer audio and transcripts are used only for transcription and summarization and never to train AI models. You can delete any recording or your whole account at any time, and recordings are retained for six months. Details are on our security page.
Pricing is pay as you go, starting at $1.99 for 60 minutes, with top-up minutes valid for six months. An optional Unlimited plan costs $19.99 a month. Whatever tool produced the draft, the checks above still apply: verify names, negations and numbers against the audio before you quote anything.
Try it on one recorded interview: create a Loka Note account
Frequently asked questions
What is the difference between verbatim and clean verbatim transcription?
Full verbatim keeps everything: filler words like 'um' and 'uh', false starts, overlapping speech, pauses and non-verbal cues. Clean verbatim, also called intelligent verbatim, removes fillers, repetitions and false starts while keeping all meaningful content. Clean verbatim is the usual standard for journalism and general qualitative research.
How do you format an interview transcript properly?
Start with header metadata such as date, interviewer and participant code. Label each speaker consistently, using initials or role tags like I: and P:. Add timestamps at set intervals or at speaker turns, and use the same tags every time for inaudible passages and overlapping speech.
How do qualitative researchers transcribe interviews for thematic analysis?
Thematic and content analysis generally need the participant's actual words with light notation, so clean or lightly annotated verbatim is usually enough. Methods like IPA and narrative interviews also need hesitations and emotional cues. Conversation analysis is the exception and requires full Jefferson notation.
Can AI transcribe interviews accurately with multiple speakers?
AI tools can apply automatic speaker labeling and produce a usable first draft, but the labels and the words both need checking against the audio. Pay special attention to overlapping speech, dropped negations, names and numbers. Treat the AI output as a draft, never as the final record.
How should you handle consent, anonymization, and participant privacy in interview transcripts?
Get consent for recording and for how the transcript will be stored and shared. Replace direct identifiers with pseudonyms or codes, and keep the key that links codes to real names in a separate, restricted file. Follow the rules set by your ethics board or data protection officer, and delete audio when your protocol says to.
Sources
- 1How to Transcribe Interviews with AI (2026)transcription.atter-ai.com
- 2Interview Transcription: The Complete 2026 Guideconvertaudiototext.com
- 3A Guide to Writing Interview Transcripts: Styles, Workflow, and Template | Otter.aiotter.ai
- 4Qualitative Interview Transcription Guide | ParrotNotesparrotnotes.app


