Skip to content
All posts
Guides

Transcribing research interviews for a thesis: verbatim vs clean verbatim, timing, and getting the file into NVivo or ATLAS.ti

Every methods textbook tells you to transcribe your interviews and pick a convention. None of them tell you which parts of that convention an automatic transcript already gets right, or how many hours the review pass actually takes. Here is both, plus the timestamp trick that gets a transcript into NVivo or ATLAS.ti without retyping it.

Kalima Team10 min read

For most qualitative theses, a straightforward verbatim transcript is enough: every word the speaker used, cleaned of false starts if your convention calls for it, with speaker labels and enough timing information to get back to the audio. An automatic transcript gets you a meaningful part of that for free, gets a second part approximately right, and does not attempt a third part at all. This guide goes through which is which, gives timing figures from the literature instead of a vendor promise, and walks the file from upload to a citation you can put in your findings chapter.

A note before any of it: your department's own style guide or your supervisor's expectation always wins over a textbook default. The conventions below are the common ground most guidance agrees on, not a substitute for asking.

Verbatim, clean verbatim, or something stricter

Qualitative methods literature generally distinguishes two workable levels for a thesis, plus a third that almost nobody needs.

Verbatim transcription captures everything: false starts, filler words, repeated words, incomplete sentences. It preserves how something was said, not only what was said, which matters when your analysis is about language use itself, hesitation, or discourse.

Clean verbatim (also called "intelligent verbatim" in some transcription-industry usage) removes filler words such as "um" and "like" where they carry no analytic weight, smooths obvious false starts, and keeps direct speech otherwise intact. This is the default most content-analysis and thematic-analysis guidance assumes, because the fillers add transcription time without adding codeable content.

Conversation-analytic transcription, in the tradition associated with Jeffersonian notation, records pitch, overlap, timed pauses to the tenth of a second, and prosody. No automatic tool produces this, at any price point, because it requires a trained human ear making judgment calls a speech recognition model was never built to make. If your methodology chapter cites conversation analysis or discourse analysis as your analytic frame, plan on transcribing that material by hand.

Decide which one your project needs before you transcribe a single interview, write the decision into your methods chapter, and apply it consistently. Switching convention halfway through a data set is the most common and most avoidable transcription mistake in a thesis.

What an automatic transcript already gets right, and what it does not

Feature of a usable transcriptDoes automatic transcription deliver it?What you still do
Verbatim wording, not a summaryYesNothing. This is the core job a speech recognition engine does.
Speaker separationYes, with generic labelsRename "Speaker 1" and "Speaker 2" to your participant codes once; it applies everywhere.
Paragraph breaks per turnMostlyShort backchannels ("mhm", "right") sometimes land inside the other speaker's turn. Split them.
TimestampsYesTurn on or off at export, matched to your citation style.
Filler words and false startsIncluded by defaultRemove them if your convention is clean verbatim; keep them if it is strict verbatim. Either way, this is a judgment call the software cannot make for you.
Pauses, timed to the secondNoAdd by hand where your convention requires it.
Emphasis, volume, laughterNoAdd by hand on the listen-back pass.
Overlapping speech marked as suchNo, and speaker separation itself can misfire hereMark manually; expect to re-listen to any moment where two people spoke at once.
Uncertain or inaudible wordsPartiallyThe model tends to guess a plausible word rather than leave a gap. Flag anything you are not sure of while you listen back.

The pattern is consistent: anything that is about what was said comes out largely correct. Anything that is about how it was said, timing, emphasis, overlap, non-verbal sound, is not attempted and has to be added by a human who was either in the room or is listening carefully afterwards.

How long transcription actually takes, and why the numbers disagree

Several figures circulate, and they disagree because they are measuring different things.

SourceTime per hour of audioWhat it covers
"(Re)thinking transcription strategies", Journal of Business Research, 20236 to 10 hoursmanual verbatim transcription by a trained researcher
Commonly cited range in qualitative methods guidance4 to 6 hoursmanual clean verbatim, experienced transcriber
Conversation-analytic (Jeffersonian) transcriptionCommonly reported at 20 hours or morefull prosodic and timing detail

(Journal of Business Research figure from ScienceDirect, accessed September 2026.)

The spread depends on five things: audio quality, number of speakers, how technical the vocabulary is, which convention you picked, and your own typing speed. None of the published ranges is wrong; they are measuring different starting points.

Automatic transcription changes where the time goes rather than eliminating it. Processing itself takes minutes. The review pass, listening back to check names, technical terms, and any passage the model was unsure about, plus adding whatever your convention requires beyond plain wording, commonly runs a few hours per hour of interview audio rather than the 4 to 10 hours of transcribing from scratch. Budget accordingly rather than assuming the software finishes the job. For eight hour-long interviews, that is the difference between roughly 32 to 80 hours done entirely by hand and a review pass measured in single-digit hours per interview plus processing time.

The review time drops further if you treat it as your first analytic pass rather than a separate chore. You are listening to the material again either way; correcting the transcript and starting to notice patterns can happen in the same sitting.

Three ways to get from recording to transcript, compared honestly

RouteCostTime to a transcriptBest for
Transcribe it yourselfYour own time only4 to 10 hours per audio hour, by the figures aboveA handful of interviews, unusual audio quality, a methodology that needs full manual control
A transcription serviceCommonly priced per audio minute, figures vary by turnaround and accuracy tier; get a current quote before budgetingThe service's stated turnaround, typically one to a few daysLarger studies with funding, audio a model handles poorly (heavy accents, multiple overlapping speakers, poor recording quality)
Automatic transcription (local or cloud)Free to a modest monthly planMinutes of processing, then your review passMost single-researcher qualitative projects with clear audio and a small number of speakers

If your ethics approval or your institution's data policy requires that audio never leave your own machine, a fully local, open-source option is the correct answer regardless of what an online service costs, because "processed locally" and "processed in the cloud" are not interchangeable under most research ethics frameworks. Check that requirement before you pick a tool, not after you have uploaded a file.

From upload to a transcript you can work with

If the recording already exists, on a phone, a dedicated recorder, or pulled from a video call, file upload is the starting point. Kalima accepts common audio formats (MP3, WAV, M4A, AAC, FLAC, OGG, Opus, including WhatsApp voice notes, plus WebM and AIFF) and video containers (MP4, MOV, WebM), extracting the audio in the browser so only sound is uploaded from a video file.

Limits scale with plan, and they are the one place a free tier gets genuinely tight for interview-length recordings: 50 MB and 30 minutes per file on Free, up to 500 MB and 180 minutes on Plus, and 1 GB and 300 minutes on Pro, with a hard cap of five hours per file on every plan and a weekly upload count that resets Monday at 00:00 UTC. A single 60-minute interview has to be split into two files on the free plan, and each half counts as a separate upload. One or two interviews fit comfortably; a full study is easier on a paid plan. Full detail is in supported formats and limits.

Choose Post-Process+ rather than Realtime when you record the interview directly in Kalima. It shows no text while you talk, processes the whole recording afterwards, and delivers the highest accuracy and the best speaker separation of the two modes, at the same transcription cost, because the engine has the full recording to work with rather than a live stream. See recording modes. For a semi-structured interview you are conducting, not reading along with, live text adds nothing.

Speaker separation is automatic and starts with generic labels. Click the speaker field on any segment and type your participant code, "P1" or "R" for researcher; the change applies to every segment from that voice and to every export. If separation gets a passage wrong, Check voices re-listens on your own device, downloading a 38 MB model once and taking about 40 seconds per hour of audio, and lets you correct misattributed passages. See speakers and diarization. One limit worth knowing: identity does not carry across separate recordings, so the same participant can get a different label in a second session.

From transcript to appendix

1
Export the text. Open the export dialog and choose Copy for AI for plain Markdown you can paste into Word, turning off the AI instruction line and keeping speaker names on; decide on timestamps according to your department's citation convention. For analysis software, export SRT or VTT instead.
2
De-identify. Replace names, places, employers and anything else identifying with a bracketed placeholder: [P1], [City], [Employer]. Keep the key linking placeholders to real identities in a separate file, not in the appendix, per your ethics approval.
3
Format for submission. 1.5 line spacing, a blank line at each speaker change, page numbers, and continuous line numbering through Word's line-number feature if your department cites by line rather than by page.
4
Cite consistently. A common convention is participant code plus line or timestamp, for example "P1, lines 45 to 78". Check whether your department wants the interview date included in the citation as well.
5
Check your department's own template last. Institutional transcription guidance often adds requirements a general methods textbook does not: a header with date and location, a specific notation for overlapping speech, or a required disclosure statement about how the transcript was produced.

Getting the transcript into NVivo, ATLAS.ti or MAXQDA

All three major qualitative analysis packages import timestamped subtitle formats. Export SRT or VTT rather than plain text, import it through the general document or transcript import in your software rather than a vendor-specific list (Kalima will not appear by name in any of these tools' vendor picklists, and that is fine, since the timestamp format is what the software actually reads), and link the original media file afterwards so a coded segment can jump back to the audio. That link is worth the extra step: it is the fastest way to catch a transcription error during coding rather than after your findings are written.

Say before you record that you are recording, why, how long the file will be kept, and how it will be de-identified, and get agreement, ideally in writing or recorded as part of the interview itself if your ethics protocol allows that. This is not a formality: recording someone's spoken word without their knowledge is a criminal offence in a number of jurisdictions, and it is very likely to fall outside your institutional ethics approval regardless of what the law technically permits. Ready-made wording for the spoken announcement and for a written consent line is in meeting recording consent wording.

For the methods section, state plainly what happens to the recording: Kalima encrypts connections, stores data in the EU, never uses audio or transcripts to train models, keeps sharing off by default, and lets you export and delete your own data on request. What this guide cannot claim, and does not, is that using any particular tool automatically satisfies your institution's ethics or data-protection requirements. Confirm the specifics, storage location, data processing terms, retention, with your supervisor and your institution's data protection or research ethics office before your first interview, not after. If your ethics approval requires that audio never leave your institution's network, a local tool is the right answer, not an online one.

Frequently asked questions

Do I need to transcribe every "um" and false start?

Only if your convention is strict verbatim. Clean verbatim, the more common choice for thematic and content analysis, removes fillers that carry no analytic content and smooths obvious false starts. Decide once, apply it to every interview, and say which you chose in your methods chapter.

Is an automatic transcript good enough for a thesis?

As a starting point, yes; as a final product, not on its own. Wording, speaker separation and timestamps come through largely correct. Pauses, emphasis, overlapping speech and non-verbal sound do not, and reviewers generally expect a listen-back pass on any thesis transcript regardless of how it was produced.

How long will eight one-hour interviews actually take?

By hand, using the commonly cited figures, somewhere between roughly 32 and 80 hours. With automatic transcription and a careful review pass, expect a few hours of review per interview hour plus processing time, considerably less than transcribing from scratch, but not zero. Audio quality and how technical the vocabulary is will move the estimate more than any tool choice does.

Can I import a Kalima transcript into NVivo or ATLAS.ti?

Yes, by exporting SRT or VTT and using your software's general transcript or subtitle import rather than a named vendor list. Both major formats carry timestamps that the analysis software reads automatically, and linking the source audio afterwards lets you jump from a coded segment back to the recording.

Can I record an interview without telling the participant?

Do not. Beyond the ethical problem, recording someone's spoken word without their knowledge is a criminal offence in several jurisdictions and will very likely put you outside your institution's ethics approval. Announce it before you start, and get the participant's agreement on record.

Try Kalima on your next call.

Live transcription and translation, no bot in the meeting. Free to start.

Start for free

Keep reading.