Simply Transcribe

Engineering ·

How we made a 64-minute recording transcribe in 24 seconds on iPhone

Benchmarks measured on a real iPhone 14 Pro — no simulator, no debug build. Here's what we measured, what changed, and what we still don't claim.

The setup

Our benchmark file is 3,838 seconds — about 64 minutes — of real speech, transcribed through the production pipeline: voice activity detection, window planning, the on-device Parakeet v3 model, and the transcript merger. The engine splits it into 306 windows and 14 speech regions, and produces 855 sentence segments.

Every number below is from a physical iPhone 14 Pro running a Release-optimized build. Simulator and debug numbers are fine for validating the harness, but they can't establish iPhone performance — so we never use them here.

The problem: converting audio twice

The app originally recorded at 48 kHz and converted audio to the model's 16 kHz format during transcription — in parts of the pipeline, twice. For the 64-minute file, that conversion alone cost 25–29 seconds of a roughly 71-second action.

The change: store audio in the format the model eats

New recordings are now captured directly as 16 kHz mono audio — the exact format the model consumes — and older recordings are updated to that format when you retranscribe them. Preparation happens during recording instead of at transcription time.

Measured on the same iPhone 14 Pro, same file:

That's 29–34% faster end to end, with the transcript bit-identical in every run: same 855 segments, same timestamps, same speech regions. As a bonus, the stored audio is one-third the size — about 3.8 MB per minute instead of 11.5.

What we don't claim

A competitor reports transcribing a 35-minute file in about 18 seconds — on a Mac with an M4 Pro. That's a different device class with a much larger thermal envelope, and our 24 seconds covers a file nearly twice as long on a phone. We don't claim parity; we publish our own device, our own file, and our own method, so the numbers are checkable.

Two runs also aren't a stable bound — thermal state moves mobile benchmarks, and we say so in our test notes. The next milestones are sustained-capture power measurements and more device tiers.

Why it matters

On-device transcription means your recording never leaves the phone — but it only feels magical if it's fast. Cutting a third of the wait brings an hour-long lecture down to a coffee-sip, with nothing uploaded and no subscription meter running.

← Blog