Transcribing research interviews for a thesis: verbatim vs clean verbatim, timing, and getting the file into NVivo or ATLAS.ti
Every methods textbook tells you to transcribe your interviews and pick a convention. None of them tell you which parts of that convention an automatic transcript already gets right, or how many hours the review pass actually takes. Here is both, plus the timestamp trick that gets a transcript into NVivo or ATLAS.ti without retyping it.
For most qualitative theses, a straightforward verbatim transcript is enough: every word the speaker used, cleaned of false starts if your convention calls for it, with speaker labels and enough timing information to get back to the audio. An automatic transcript gets you a meaningful part of that for free, gets a second part approximately right, and does not attempt a third part at all. This guide goes through which is which, gives timing figures from the literature instead of a vendor promise, and walks the file from upload to a citation you can put in your findings chapter.
A note before any of it: your department's own style guide or your supervisor's expectation always wins over a textbook default. The conventions below are the common ground most guidance agrees on, not a substitute for asking.
Verbatim, clean verbatim, or something stricter
Qualitative methods literature generally distinguishes two workable levels for a thesis, plus a third that almost nobody needs.
Verbatim transcription captures everything: false starts, filler words, repeated words, incomplete sentences. It preserves how something was said, not only what was said, which matters when your analysis is about language use itself, hesitation, or discourse.
Clean verbatim (also called "intelligent verbatim" in some transcription-industry usage) removes filler words such as "um" and "like" where they carry no analytic weight, smooths obvious false starts, and keeps direct speech otherwise intact. This is the default most content-analysis and thematic-analysis guidance assumes, because the fillers add transcription time without adding codeable content.
Conversation-analytic transcription, in the tradition associated with Jeffersonian notation, records pitch, overlap, timed pauses to the tenth of a second, and prosody. No automatic tool produces this, at any price point, because it requires a trained human ear making judgment calls a speech recognition model was never built to make. If your methodology chapter cites conversation analysis or discourse analysis as your analytic frame, plan on transcribing that material by hand.
Decide which one your project needs before you transcribe a single interview, write the decision into your methods chapter, and apply it consistently. Switching convention halfway through a data set is the most common and most avoidable transcription mistake in a thesis.
What an automatic transcript already gets right, and what it does not
| Feature of a usable transcript | Does automatic transcription deliver it? | What you still do |
|---|---|---|
| Verbatim wording, not a summary | Yes | Nothing. This is the core job a speech recognition engine does. |
| Speaker separation | Yes, with generic labels | Rename "Speaker 1" and "Speaker 2" to your participant codes once; it applies everywhere. |
| Paragraph breaks per turn | Mostly | Short backchannels ("mhm", "right") sometimes land inside the other speaker's turn. Split them. |
| Timestamps | Yes | Turn on or off at export, matched to your citation style. |
| Filler words and false starts | Included by default | Remove them if your convention is clean verbatim; keep them if it is strict verbatim. Either way, this is a judgment call the software cannot make for you. |
| Pauses, timed to the second | No | Add by hand where your convention requires it. |
| Emphasis, volume, laughter | No | Add by hand on the listen-back pass. |
| Overlapping speech marked as such | No, and speaker separation itself can misfire here | Mark manually; expect to re-listen to any moment where two people spoke at once. |
| Uncertain or inaudible words | Partially | The model tends to guess a plausible word rather than leave a gap. Flag anything you are not sure of while you listen back. |
The pattern is consistent: anything that is about what was said comes out largely correct. Anything that is about how it was said, timing, emphasis, overlap, non-verbal sound, is not attempted and has to be added by a human who was either in the room or is listening carefully afterwards.
How long transcription actually takes, and why the numbers disagree
Several figures circulate, and they disagree because they are measuring different things.
| Source | Time per hour of audio | What it covers |
|---|---|---|
| "(Re)thinking transcription strategies", Journal of Business Research, 2023 | 6 to 10 hours | manual verbatim transcription by a trained researcher |
| Commonly cited range in qualitative methods guidance | 4 to 6 hours | manual clean verbatim, experienced transcriber |
| Conversation-analytic (Jeffersonian) transcription | Commonly reported at 20 hours or more | full prosodic and timing detail |
(Journal of Business Research figure from ScienceDirect, accessed September 2026.)
The spread depends on five things: audio quality, number of speakers, how technical the vocabulary is, which convention you picked, and your own typing speed. None of the published ranges is wrong; they are measuring different starting points.
Automatic transcription changes where the time goes rather than eliminating it. Processing itself takes minutes. The review pass, listening back to check names, technical terms, and any passage the model was unsure about, plus adding whatever your convention requires beyond plain wording, commonly runs a few hours per hour of interview audio rather than the 4 to 10 hours of transcribing from scratch. Budget accordingly rather than assuming the software finishes the job. For eight hour-long interviews, that is the difference between roughly 32 to 80 hours done entirely by hand and a review pass measured in single-digit hours per interview plus processing time.
The review time drops further if you treat it as your first analytic pass rather than a separate chore. You are listening to the material again either way; correcting the transcript and starting to notice patterns can happen in the same sitting.
Three ways to get from recording to transcript, compared honestly
| Route | Cost | Time to a transcript | Best for |
|---|---|---|---|
| Transcribe it yourself | Your own time only | 4 to 10 hours per audio hour, by the figures above | A handful of interviews, unusual audio quality, a methodology that needs full manual control |
| A transcription service | Commonly priced per audio minute, figures vary by turnaround and accuracy tier; get a current quote before budgeting | The service's stated turnaround, typically one to a few days | Larger studies with funding, audio a model handles poorly (heavy accents, multiple overlapping speakers, poor recording quality) |
| Automatic transcription (local or cloud) | Free to a modest monthly plan | Minutes of processing, then your review pass | Most single-researcher qualitative projects with clear audio and a small number of speakers |
If your ethics approval or your institution's data policy requires that audio never leave your own machine, a fully local, open-source option is the correct answer regardless of what an online service costs, because "processed locally" and "processed in the cloud" are not interchangeable under most research ethics frameworks. Check that requirement before you pick a tool, not after you have uploaded a file.
From upload to a transcript you can work with
If the recording already exists, on a phone, a dedicated recorder, or pulled from a video call, file upload is the starting point. Kalima accepts common audio formats (MP3, WAV, M4A, AAC, FLAC, OGG, Opus, including WhatsApp voice notes, plus WebM and AIFF) and video containers (MP4, MOV, WebM), extracting the audio in the browser so only sound is uploaded from a video file.
Limits scale with plan, and they are the one place a free tier gets genuinely tight for interview-length recordings: 50 MB and 30 minutes per file on Free, up to 500 MB and 180 minutes on Plus, and 1 GB and 300 minutes on Pro, with a hard cap of five hours per file on every plan and a weekly upload count that resets Monday at 00:00 UTC. A single 60-minute interview has to be split into two files on the free plan, and each half counts as a separate upload. One or two interviews fit comfortably; a full study is easier on a paid plan. Full detail is in supported formats and limits.
Choose Post-Process+ rather than Realtime when you record the interview directly in Kalima. It shows no text while you talk, processes the whole recording afterwards, and delivers the highest accuracy and the best speaker separation of the two modes, at the same transcription cost, because the engine has the full recording to work with rather than a live stream. See recording modes. For a semi-structured interview you are conducting, not reading along with, live text adds nothing.
Speaker separation is automatic and starts with generic labels. Click the speaker field on any segment and type your participant code, "P1" or "R" for researcher; the change applies to every segment from that voice and to every export. If separation gets a passage wrong, Check voices re-listens on your own device, downloading a 38 MB model once and taking about 40 seconds per hour of audio, and lets you correct misattributed passages. See speakers and diarization. One limit worth knowing: identity does not carry across separate recordings, so the same participant can get a different label in a second session.
From transcript to appendix
Getting the transcript into NVivo, ATLAS.ti or MAXQDA
All three major qualitative analysis packages import timestamped subtitle formats. Export SRT or VTT rather than plain text, import it through the general document or transcript import in your software rather than a vendor-specific list (Kalima will not appear by name in any of these tools' vendor picklists, and that is fine, since the timestamp format is what the software actually reads), and link the original media file afterwards so a coded segment can jump back to the audio. That link is worth the extra step: it is the fastest way to catch a transcription error during coding rather than after your findings are written.
The consent and data-handling paragraph your ethics board will want
Say before you record that you are recording, why, how long the file will be kept, and how it will be de-identified, and get agreement, ideally in writing or recorded as part of the interview itself if your ethics protocol allows that. This is not a formality: recording someone's spoken word without their knowledge is a criminal offence in a number of jurisdictions, and it is very likely to fall outside your institutional ethics approval regardless of what the law technically permits. Ready-made wording for the spoken announcement and for a written consent line is in meeting recording consent wording.
For the methods section, state plainly what happens to the recording: Kalima encrypts connections, stores data in the EU, never uses audio or transcripts to train models, keeps sharing off by default, and lets you export and delete your own data on request. What this guide cannot claim, and does not, is that using any particular tool automatically satisfies your institution's ethics or data-protection requirements. Confirm the specifics, storage location, data processing terms, retention, with your supervisor and your institution's data protection or research ethics office before your first interview, not after. If your ethics approval requires that audio never leave your institution's network, a local tool is the right answer, not an online one.
Frequently asked questions
Do I need to transcribe every "um" and false start?
Only if your convention is strict verbatim. Clean verbatim, the more common choice for thematic and content analysis, removes fillers that carry no analytic content and smooths obvious false starts. Decide once, apply it to every interview, and say which you chose in your methods chapter.
Is an automatic transcript good enough for a thesis?
As a starting point, yes; as a final product, not on its own. Wording, speaker separation and timestamps come through largely correct. Pauses, emphasis, overlapping speech and non-verbal sound do not, and reviewers generally expect a listen-back pass on any thesis transcript regardless of how it was produced.
How long will eight one-hour interviews actually take?
By hand, using the commonly cited figures, somewhere between roughly 32 and 80 hours. With automatic transcription and a careful review pass, expect a few hours of review per interview hour plus processing time, considerably less than transcribing from scratch, but not zero. Audio quality and how technical the vocabulary is will move the estimate more than any tool choice does.
Can I import a Kalima transcript into NVivo or ATLAS.ti?
Yes, by exporting SRT or VTT and using your software's general transcript or subtitle import rather than a named vendor list. Both major formats carry timestamps that the analysis software reads automatically, and linking the source audio afterwards lets you jump from a coded segment back to the recording.
Can I record an interview without telling the participant?
Do not. Beyond the ethical problem, recording someone's spoken word without their knowledge is a criminal offence in several jurisdictions and will very likely put you outside your institution's ethics approval. Announce it before you start, and get the participant's agreement on record.
Related
- Recording without transcript, transcript without a recording: getting from an existing file to text.
- Meeting recording consent wording you can copy
- File upload and supported formats and limits
- Speakers and diarization and recording modes
- Kalima for education and pricing
Try Kalima on your next call.
Live transcription and translation, no bot in the meeting. Free to start.
Keep reading.

Your Company Blocked AI Notetakers. What You Can and Cannot Do Next
A blocked notetaker is either a switched-off transcription policy, a lobby that refused a bot, or a network that refuses the service itself. Only one of those is about something joining your call, and only one is a question you can answer without asking somebody first.

Can you record a job interview? Consent, state law and the AI Act question nobody answers
Five vendor blogs answer this with a state list and a consent form. None of them mention that a voice recording can trigger a biometric statute, none cover an EU employer interviewing a US candidate or the reverse, and none say what happens to the recording after the decision is made.

The Google Meet Transcript Exists and You Still Cannot Open It
You were in the meeting, the transcript was made, and Drive still shows you a Request access screen. Meet ties the transcript to the Calendar invite, cuts off meetings with more than 200 invitees, and leaves external guests out entirely. Here is each rule, and why the file came out in the wrong language.