Skip to content

transcription

Interview Transcription: Styles, Formatting and a Two-Pass Check

How to choose a transcription style, format speaker labels and timestamps, verify an AI draft in two passes, and handle anonymization and mixed-language audio.

Loka Team9 min read
Simple editorial illustration of two speech bubbles flowing into a typed page with a magnifying glass over one line of text

Interview transcription turns recorded speech into text you can quote, code and search. The decision that shapes everything else is the style: full verbatim, clean verbatim or edited. Pick that first, format the transcript consistently, then check the AI draft against the audio in two passes.

This guide covers each step, plus the parts most guides skip: privacy and anonymization, and interviews where people switch languages mid-sentence.

Choose your interview transcription style first

There are three common styles, and the right one depends on what you will do with the text. Choosing after you have transcribed means redoing work.

Otter's guide to interview transcripts describes the three styles this way:

StyleWhat it keepsTypical use
Full verbatimFiller words ("um", "uh"), false starts, overlapping speech, pauses, non-verbal cuesConversation analysis, discourse research, legal depositions
Clean (intelligent) verbatimAll meaningful content and non-verbal cues; removes fillers, repetitions and false startsJournalism and general qualitative research
EditedRewrites phrasing, polishes grammar, reshapes sentencesInternal synthesis, UX reports, public-facing articles

If you are unsure, start with clean verbatim. It stays faithful to what the person said and is readable enough to code or quote. You can always move toward edited later. You cannot recover pauses and false starts from a transcript that never recorded them, unless you go back to the audio.

One rule across all three: write down which style you used in the transcript header. A reader who sees a quote with no "ums" should know whether that is the speaker's style or your edit.

Match the style to your research method

Your methodology, not your preference, decides how much detail the transcript needs. ParrotNotes' qualitative guide maps it out:

  • Content and thematic analysis need the participants' actual words with light notation.
  • IPA and narrative interviews need hesitations and emotional cues captured, because those carry meaning for the analysis.
  • Conversation analysis demands full Jefferson notation: timed pauses, overlaps, latching and pitch.

For journalists and UX researchers the question is simpler. Will anyone read the transcript as evidence, or only as a source for quotes and findings? If it is evidence, stay closer to verbatim. If it feeds a report, clean verbatim plus timestamps back to the audio is usually enough.

Be careful with a mixed approach. If most of your interviews are clean verbatim and a few are full verbatim, your coding will treat hesitation inconsistently. Pick one standard per project.

A workflow from recording to first draft

Good transcripts start with good audio. An AI draft from a clean recording needs far less repair than one from a noisy room.

Atter AI's interview guide lays out the AI workflow: record clear audio in a quiet space, upload the file, apply automated speaker diarization, then verify against the audio. Here is that sequence with the extra steps we would add:

  1. Set up the recording. Use a quiet room, put the microphone close to each speaker and test it for a minute before you begin. For remote interviews, ask each person to use headphones so the other voice does not leak into their microphone.
  2. State the basics on tape. Say the date, the interview code and that the participant consents to recording. This doubles as a record and as a check that the file is the right one.
  3. Name the file consistently. For example, 2026-10-08_P07_interview.m4a. Use the participant code, not the person's name.
  4. Pick the transcription language. If the interview is not in English, set the right language, or use a tool that detects it. Wrong language settings are a common cause of unusable drafts.
  5. Upload and generate the draft. Keep the untouched machine output as its own file.
  6. Check speaker labels. Diarization often confuses speakers when people interrupt or sound alike. Fix labels before you fix words, since a wrong label changes who said what.
  7. Run the two-pass check described below.
  8. Save the final version separately from the raw draft.

If you are comparing ways to get from audio to text at all, our post on how to transcribe audio to text compares five methods.

How to format an interview transcript

A proper interview transcript has header metadata, consistent speaker labels, timestamps and standard tags for unclear speech. Otter's guide lists these as the standard elements: speaker labels (initials such as SJ:, or role tags such as I: and P:), timestamps at set intervals of 5 to 10 minutes or at speaker turns, header metadata, and standardized tags for inaudibles and overlapping speech.

Put the basics at the top:

  • Interview code (not the participant's name)
  • Date and approximate length
  • Interviewer
  • Language(s) spoken
  • Transcription style (for example, "clean verbatim")
  • Who produced and checked the transcript

Speaker labels

Pick one scheme and keep it. I: and P: works for one interviewer and one participant. With several participants, use P1:, P2: or role tags. Avoid real names if you plan to anonymize.

Timestamps

Interval timestamps are lighter to read. Speaker-turn timestamps make it easier to jump back to the audio. If you expect to quote people or check disputed passages, use speaker-turn timestamps.

Tags for unclear speech

Choose a small set and use it everywhere. For example:

  • [inaudible 00:14:32] for speech you cannot make out
  • [unclear: "budget"?] for a best guess
  • [overlap] where two people talk at once
  • [laughs] or [long pause] where it affects meaning

Do not guess silently. A confident but wrong word is worse than a visible gap.

The two-pass quality check

Run one pass against the audio and one pass without it. Convert Audio to Text's guide describes the method: in pass one, read along while listening at 1.0x to 1.25x speed and correct names, negations and technical terms. In pass two, read the text on its own to catch contextual blind spots and formatting errors.

The two passes catch different problems. Listening catches what the machine misheard. Reading alone catches sentences that do not make sense, because you are no longer being carried along by the audio.

Where AI drafts fail

Fluent text can hide errors. These are the places to slow down in pass one:

  • Negations. "Can" versus "cannot", "did" versus "didn't". A dropped negation reverses a finding. Listen again every time a sentence carries a decision, a complaint or a refusal.
  • Names and organizations. Check spelling against your participant list.
  • Numbers, dates and money. Verify each against the audio.
  • Technical terms and acronyms. Build a short glossary and search the transcript for variants.
  • Text that appears where the audio is unclear. If the transcript reads smoothly over a stretch where you hear noise, treat it with suspicion. Replace it with an [inaudible] tag unless you can confirm the words.

Keep an audit trail

Keep the raw AI draft, the corrected version and a short change log: what you changed, where and why. For research, this lets you show that a quote reflects the recording and not the machine's guess. For journalism, it protects you if someone disputes a quotation. Even a spreadsheet with timestamp, original text and correction is enough.

Treat transcripts as personal data from the moment you record. Get consent that covers recording, storage and who will see the transcript, and then reduce what the transcript reveals.

The rules that apply to you come from your ethics board (IRB), your institution, your employer or data protection law such as GDPR. Ask them before you start, not after the interviews are done. What follows is a practical baseline, not legal advice.

  1. Remove direct identifiers. Names, phone numbers, email addresses, exact job titles and employers can all point to one person.
  2. Check indirect identifiers. A small town, a rare role and a specific event together can identify someone even without a name.
  3. Use pseudonyms or codes. Replace names with P07 or an invented name, consistently, across all transcripts.
  4. Store the key separately. The file linking codes to real names should live apart from the transcripts, with tighter access.
  5. Remember that audio is identifying. A voice can identify a person. Decide how long you keep recordings and delete them when your protocol says so.
  6. Check your tool's data handling. Before uploading, find out where files are stored, whether they are used to train models, and whether you can delete them.

Anonymize in a copy. Keep the original transcript under restricted access until you no longer need it.

Accents, code-switching and cross-border interviews

Most transcription guides assume one language and a standard accent. Real interviews in Southeast Asia and Nigeria often do not fit that. People switch between Burmese and English, or between English and Yoruba or Hausa, in the middle of a sentence.

Practical steps:

  • Choose the right language setting. If you pick English for an interview that is half Burmese, the Burmese portions come out as gibberish or get skipped.
  • Decide how to render switches. Keep each language in its original script and mark the switch if it matters to your analysis. If you add translations, put them in brackets or a separate column so the original words stay visible.
  • Use a glossary of terms. Local place names, organizations and borrowed words need consistent spelling.
  • Give accented speech extra time in pass one. Accent is a common cause of misheard names and numbers.
  • Consider a bilingual reviewer. For anything quoted or analyzed closely, a second person who speaks both languages should check key passages.

We go deeper on this in why Burmese-English code-switching breaks transcription tools and in our guide to multilingual meeting transcription.

Where Loka Note fits

Loka Note is an AI meeting-notes tool that also accepts uploaded audio and video files, so you can use it to produce a first-draft transcript of a recorded interview. It is built Burmese-first and handles mixed Burmese-English speech, and it also supports English, Thai, Vietnamese, Chinese (Simplified), Yoruba and Hausa. The language you pick is checked against the audio, so an interview recorded under the wrong setting is transcribed in the language actually spoken.

On privacy, customer audio and transcripts are used only for transcription and summarization and never to train AI models. You can delete any recording or your whole account at any time, and recordings are retained for six months. Details are on our security page.

Pricing is pay as you go, starting at $1.99 for 60 minutes, with top-up minutes valid for six months. An optional Unlimited plan costs $19.99 a month. Whatever tool produced the draft, the checks above still apply: verify names, negations and numbers against the audio before you quote anything.

Try it on one recorded interview: create a Loka Note account

Frequently asked questions

What is the difference between verbatim and clean verbatim transcription?

Full verbatim keeps everything: filler words like 'um' and 'uh', false starts, overlapping speech, pauses and non-verbal cues. Clean verbatim, also called intelligent verbatim, removes fillers, repetitions and false starts while keeping all meaningful content. Clean verbatim is the usual standard for journalism and general qualitative research.

How do you format an interview transcript properly?

Start with header metadata such as date, interviewer and participant code. Label each speaker consistently, using initials or role tags like I: and P:. Add timestamps at set intervals or at speaker turns, and use the same tags every time for inaudible passages and overlapping speech.

How do qualitative researchers transcribe interviews for thematic analysis?

Thematic and content analysis generally need the participant's actual words with light notation, so clean or lightly annotated verbatim is usually enough. Methods like IPA and narrative interviews also need hesitations and emotional cues. Conversation analysis is the exception and requires full Jefferson notation.

Can AI transcribe interviews accurately with multiple speakers?

AI tools can apply automatic speaker labeling and produce a usable first draft, but the labels and the words both need checking against the audio. Pay special attention to overlapping speech, dropped negations, names and numbers. Treat the AI output as a draft, never as the final record.

How should you handle consent, anonymization, and participant privacy in interview transcripts?

Get consent for recording and for how the transcript will be stored and shared. Replace direct identifiers with pseudonyms or codes, and keep the key that links codes to real names in a separate, restricted file. Follow the rules set by your ethics board or data protection officer, and delete audio when your protocol says to.

Sources

  1. 1How to Transcribe Interviews with AI (2026)transcription.atter-ai.com
  2. 2Interview Transcription: The Complete 2026 Guideconvertaudiototext.com
  3. 3A Guide to Writing Interview Transcripts: Styles, Workflow, and Template | Otter.aiotter.ai
  4. 4Qualitative Interview Transcription Guide | ParrotNotesparrotnotes.app

Keep reading