Transcription for journalists
Journalists choosing transcription should weigh four things in order: source confidentiality (where the audio is processed and stored), accuracy on real-world interview audio, whether the tool keeps timestamps for quote verification, and how it fits an interview-to-story workflow. For confidential or off-the-record recordings, on-device or on-prem transcription that never uploads the audio is the safest choice; polished cloud services are fine only when the material is not sensitive.
Confidentiality comes first
Protecting a source can be a legal and ethical obligation, and where the recording is processed decides how much control you keep. The Reporters Committee for Freedom of the Press notes that reporter’s-privilege protections vary widely by jurisdiction — so minimising who holds the raw audio reduces your exposure.
- Local / on-device — the recording is transcribed on your own laptop or server and never leaves it. Best for confidential or unpublished interviews.
- On-prem — a self-hosted pipeline on hardware you control; same benefit, scaled to a team or a back-catalogue.
- Cloud — the audio is uploaded to a vendor. Convenient and often accurate, but the file now lives on someone else’s infrastructure under their terms and retention policy.
Cloud tools such as Otter.ai and Rev publish SOC2 and, in Rev’s case, HIPAA commitments, which help with vendor diligence — but they do not change the fact that the audio leaves your environment. For a deeper comparison of the two models, see on-prem vs cloud transcription.
Accuracy, timestamps and speaker labels
- Accuracy — leading engines report word error rates around 3-5% on clean speech, but street noise, accents and crosstalk push that higher. Learn how WER is measured so vendor numbers don’t mislead you.
- Word-level timestamps — essential for journalism. They let you jump back to the exact second and verify a quote against the recording before it runs.
- Diarization — speaker labels matter for multi-source interviews and panels so you attribute quotes correctly. Always spot-check them on overlapping speech.
- Language coverage — a foreign-language fixer interview needs an engine that handles the language well; coverage ranges widely, from a handful of languages to several dozen. Check the catalog for a tool’s actual list rather than assuming.
- Verification — treat every automatic transcript as a draft; re-listen to anything you intend to quote.
Build, buy, or run it locally
For an occasional interview, a local desktop app (built on open Whisper models) transcribes on your own machine with no upload and no per-minute bill — see the best free and open-source transcription ranking. A newsroom with a growing archive of recordings has a different problem: keeping hundreds of interviews searchable and private at once.
That is where an on-prem pipeline fits. Self-hosting Whisper with diarization keeps audio in-house and can turn the archive into a searchable knowledge base. NoParrot is one on-prem option that adds diarization and a local search layer; the trade-off, as with any self-hosted tool, is setup effort versus the confidentiality you gain. Weigh it with the build-vs-buy guide and the best on-prem transcription ranking before committing.
Frequently asked questions
What should journalists look for in transcription software?
Prioritise source confidentiality (where audio is processed and stored), accuracy on accents and crosstalk, and a workflow that keeps timestamps so quotes can be verified against the recording. On-device or on-prem processing matters most for confidential or off-the-record interviews.
Is cloud transcription safe for confidential sources?
Cloud transcription uploads the recording to a third party, which can conflict with a promise of confidentiality and with legal shield protections. For sensitive interviews, local or on-prem transcription keeps the audio on your own machine, so nothing leaves your control.
How accurate is automatic transcription for interviews?
Leading engines report roughly 3-5% word error rate on clean speech, but accuracy drops with accents, background noise and crosstalk. Always treat the transcript as a draft: re-listen to any passage you plan to quote before publishing.
Do journalists need speaker labels?
Yes for multi-person interviews and panels. Speaker diarization tags who said what, so you can attribute quotes correctly and find a source's remarks quickly. Verify the labels — no diarization is perfect on overlapping speech.