The recording exists. That was the hard part, and you did it.
Ninety-four minutes with your grandmother, on a Sunday afternoon, about the year the family moved and what the house was like before it was sold. She talked about people you have never heard of. She switched languages three times in one sentence when she got to the part about the wedding. At one point she stopped for eleven seconds and then said something quietly that you did not catch at the time and still have not caught.
The file has been on your phone for four months.
This is where most oral history projects die, and the reason is always the same. Somebody worked out that transcribing ninety minutes by hand takes between six and nine hours, tried it for twenty minutes, and stopped.
The recording is the document, not the transcript
Before anything practical, one distinction that changes every decision that follows.
In most transcription work the text replaces the audio. Nobody keeps the recording of a project status meeting once the minutes are written. The audio was a means.
Oral history does not work that way. The recording is the primary source. It carries the pause before a difficult answer, the change in register when the speaker moves from public account to private one, the moment somebody starts crying and keeps talking. None of that is in the text, and no transcription convention has ever successfully captured it.
What the transcript does is make the recording usable. It lets you find the passage about the move without listening to all ninety-four minutes. It lets you quote accurately. It lets someone who cannot spend an afternoon listening still learn what is in the collection.
The transcript is a door into the recording. It is not a replacement for the room.
Once you accept that, the question changes from “how do I produce a perfect transcript” to “how do I produce a good enough door quickly enough that the project finishes.” Those have very different answers, and only the second one gets projects done.
Four things that make these recordings genuinely hard
Automatic transcription is now good enough for clean audio of standard speech. Oral history interviews are frequently none of those things, and it is worth knowing where the trouble sits before you are disappointed by it.
Language mixing inside a single sentence. A speaker moves between Hindi and English, or Bengali and English, or three languages in a household that has lived in four places, and does it mid clause without pausing. Speech models are built to identify a language and transcribe it. A sentence that is genuinely two languages at once sits outside that assumption, and what comes back is usually the dominant language rendered correctly and the embedded phrases rendered as whatever the model thought was closest.
Names. Every name in the interview is a proper noun the system has probably never encountered: villages, communities, festivals, relatives, the name of a shop that closed in 1974. These are also the highest value content in the entire recording, because they are what makes it a specific family’s history rather than a general account of a period. The error rate is highest precisely where the value is highest.
Older voices. Lower volume, more breath, longer pauses, dentures, tremor, and a tendency to trail off at the end of a clause. All of it reduces accuracy. Sitting closer with a better microphone helps more than any software choice you will make.
Regional varieties of English. Indian English is not a deviation from some standard, it is a variety with its own established patterns, and models trained mostly on other varieties handle it unevenly. The same is true of the dozens of regional languages that do not have dedicated model support. A tool that lists fifty languages is listing major languages. Your grandmother’s Awadhi or Bhojpuri will get transcribed as the nearest thing the model knows, which is not the same as being transcribed.
None of this makes the machine useless. It makes it a first pass rather than a result.
What the machine is genuinely good at
The bulk. The dull ninety percent that is somebody speaking clearly about something ordinary, which is most of any interview and all of the part you were never going to hand transcribe.
Timestamps, which turn the text into an index into the audio. This is the single most useful output and the one people overlook.
Speaker separation, which matters more than expected once a second family member wanders into the room and joins the conversation halfway through.
And searchability, which is the whole point. You do not need a flawless transcript to be able to search for the name of a village and find the four places it comes up.
The two pass method
The workflow that actually finishes looks like this.
First pass by machine, to get every word of the recording into text with timestamps. Upload the file to a tool that will transcribe audio recordings to text and let it run. The steps take about as long as making tea.
Upload the audio. M4A from a phone recorder, MP3, WAV or FLAC all work, and so do video files if you filmed the interview rather than only recording it. There is no cap on the length of a single file, which matters here because oral history interviews run long and splitting a ninety minute recording into three parts is exactly the kind of friction that stalls a project.
Set the language or let it detect. Around 50 languages are supported. For a mixed language interview, set it to the dominant language rather than relying on detection, and expect to fix the embedded phrases yourself.
Let it process. You get punctuated text, speakers labelled automatically, and a summary. Accuracy runs around 95% on clean audio and lower on the material described above, which is honest rather than discouraging.
Then export as TXT, DOCX, PDF or Markdown. Take DOCX if a human correction pass is coming, because that is where you will do it.
Second pass by a human, with the audio playing. This is not optional and it is not proofreading. You are listening to the recording with the transcript open and fixing four specific categories.
Every name, of every kind. Ask the narrator to spell them if they are still available to ask. This is the single highest value hour you will spend on the project.
Every passage where the language switched. Restore what was actually said in the language it was said in.
Every place the machine produced fluent text that is wrong. This is the dangerous error, more than gaps. A gap announces itself. A plausible wrong sentence does not, and it will be quoted by somebody in fifteen years.
Every marked inaudible section, checked once with headphones and better volume before you accept it.
Budget roughly one to one and a half times the recording length for this pass. For ninety minutes that is a long evening, not a lost week, and it is a tenth of what full manual transcription would have cost.
Decide your conventions before you start, not after
Every oral history project makes the same set of decisions, and making them halfway through means going back to the beginning.
Verbatim or clean read. Verbatim keeps the false starts, the repetitions and the fillers. Clean read removes them for readability. Both are defensible. Verbatim is standard for anything intended for research use, because the hesitation before an answer is data. Clean read is kinder if the transcript is going to be read by the family. If you are unsure, do verbatim and produce a cleaned version later, because you can always remove and never restore.
How to mark uncertainty. Adopt one convention and use it without exception. Square brackets with a question mark for a guessed word, and a bracketed note with a timestamp for anything genuinely unclear, is enough.
How to handle language switches. At minimum, note where they happen. Better, transcribe in the language spoken and add a translation in brackets. Never silently translate everything into English, which erases the fact that the switch occurred, and the switch is often the most interesting thing in the passage.
What to do about things said off the record. Some of the most important material in family interviews arrives after the formal part ends. Decide in advance whether that is in scope and tell the narrator what you decided.
The ethics part, which is not optional
Oral history has a longer tradition of getting this right than most fields, and the practices are worth borrowing whole.
Informed consent, gathered before recording and recorded on the audio itself. Explain what the recording is for, where it will live, and who will be able to hear it.
The narrator’s right to review. Standard practice, and reflected in the Oral History Association’s published principles, is that the person interviewed gets to see the transcript and can restrict or withdraw parts of it. This is not a formality. People say things about living relatives that they would not say if they had thought about who might read it.
Clarity about what happens to the file afterward. If it goes into an archive, say which one. If it stays in the family, say who holds it.
And basic care with the material while you are working on it. Recordings of living people talking about their families are sensitive by any reasonable definition. It is worth using tools that encrypt files in transit and at rest, do not share data with third parties, and let you delete a recording outright when you are done rather than leaving it in somebody’s cloud indefinitely. Vomo does those things and deletion is per recording, which is what you want when a narrator asks you to remove one interview and not the rest.
What the transcript makes possible
Once the text exists and has been corrected, the work you actually wanted to do becomes possible.
You can build an index of names, places and dates, which is what turns one interview into something that can be cross referenced against another.
You can find the passages worth quoting without relistening. Three good quotations from ninety-four minutes is a normal yield, and finding them by ear takes an afternoon each time.
You can ask questions of the text directly. Vomo lets you query a transcript, so asking what was said about the move, or which years come up, returns the relevant passages rather than making you scroll. For one interview that saves an hour. For a family collection of eight, it is the difference between a collection and a pile.
And you can deposit it somewhere. Archives and community history projects generally want the audio and a transcript together, and a rough but honest transcript with marked uncertainties is far more useful to them than no transcript at all.
Where this will still let you down
The machine cannot tell you who somebody meant when they said “he” for four minutes without naming anyone. Only the family can.
It cannot restore what was mumbled. If it was not audible in the room, it is not recoverable from the file, and the honest move is to mark it and move on rather than guess.
It will not catch the moment somebody decided not to answer. That lives in the pause, and the pause lives in the audio.
Which is the argument for keeping the audio, always, alongside whatever text you produce. Storage is cheap. The recording is not repeatable, and in this particular kind of work that is not a figure of speech.
Start with the file you already have
The ninety-four minutes on your phone is worth more than the four interviews you are planning and have not scheduled.
Convert it this week. On cost, so it does not become the excuse: the free tier covers 30 minutes of transcription a week, enough to test the method on a section and see what the error pattern looks like for your particular speaker, and unlimited is $1.92 a week if you want to process the whole thing in one sitting.
Then do the correction pass while your grandmother is still available to ask how the village name is spelled.
That is the part with a deadline, and it is not the transcription.
Write and Win: Participate in Creative writing Contest & International Essay Contest and win fabulous prizes.