mp3→midi runs in your browser · nothing is uploaded

Video to MIDI: What Happens to the Audio Track in an MP4

Last updated 2 October 2026

The video is thrown away. The page opens the MP4 as a container, lifts the audio track out of it, and transcribes that — the picture is never decoded, so its size and resolution do not matter. We tested this: one 20-second line, packaged fourteen different ways. Eleven of the packages converted, and all eleven returned the same 40 notes over the same range, D4 to A4, in a 432-byte MIDI file. The MP4, MOV and M4V routes placed every note 20 to 30 milliseconds later than a plain WAV did; MP3, M4A, WebM and MKV did not. Three packages failed, each in under half a second, each with an error message that says exactly what was wrong.

This page is not a description of how container formats work in general. We built the files, ran them through this site's converter one at a time, and compared the note tables it produced.

What we put in

One 20-second monophonic line at 120 BPM, 44.1 kHz, generated by us so the ground truth is known by construction rather than estimated: 40 quarter notes, 5 distinct pitches, MIDI 62 to 69 — D4 to A4 — with a sung shape, meaning sustained notes, a gentle 5.2 Hz vibrato and formant-weighted harmonics. Then we encoded that one master fourteen ways and converted each encode separately.

What came out

Eleven of the fourteen converted. Every one of the eleven returned 40 notes, spanning D4 to A4, written into a 432-byte file — the same note count, the same range, the same file size as the WAV.

What went inFile sizeNotesPitch range.mid sizeTime
WAV 44.1 kHz mono (the master)1,764,044 B40D4 – A4432 B34.0 s
MP3 192 kbps481,532 B40D4 – A4432 B28.5 s
M4A, AAC-LC 128 kbps324,389 B40D4 – A4432 B28.1 s
MP4, AAC-LC 128 kbps @ 44.1 kHz, 640×3601,037,093 B40D4 – A4432 B28.5 s
MP4, AAC-LC 128 kbps @ 48 kHz, 640×3601,041,182 B40D4 – A4432 B28.7 s
MP4, AAC-LC 128 kbps @ 44.1 kHz, 1920×10805,098,367 B40D4 – A4432 B27.0 s
MOV, AAC-LC 128 kbps1,037,144 B40D4 – A4432 B26.9 s
M4V, AAC-LC 128 kbps1,040,967 B40D4 – A4432 B26.8 s
WebM, Opus 96 kbps796,917 B40D4 – A4432 B26.7 s
WebM, Vorbis 128 kbps579,749 B40D4 – A4432 B26.6 s
MKV, H.264 + AAC 128 kbps1,290,230 B40D4 – A4432 B27.5 s

One file per row, one conversion per row, desktop Chrome, headless, software renderer, one file at a time and nothing else running in the tab. Wall-clock time is decode plus analysis plus MIDI writing. The test signal is synthetic and deterministic — generated by us, so the ground truth is known by construction rather than estimated. These are measurements from one machine and one signal, not a general accuracy claim.

Read the file sizes down the left and the note counts down the middle together. The input spans 324,389 bytes to 5,098,367 bytes — a factor of 15.7 — and the output does not move at all: 40 notes, D4 to A4, 432 bytes, eleven times in a row. Not one encoding added a note, lost a note or moved a pitch by more than 2 cents.

Two of the rows produced byte-for-byte identical MIDI files: the MP4 at 44.1 kHz and the M4V. Same audio, same container family, same file — the MIDI writer produced the same bytes, which is the strongest form of "the container did not matter" available.

What actually happens to the audio track

This is the part that no other converter page describes, because it is a property of how this page is built rather than of audio in general. When you drop an MP4, M4V or MOV, here is the sequence:

  1. The file name is checked first. Only .mp4, .m4v and .mov are routed to the container reader. Everything else goes to your browser's own audio decoder instead — which is why WebM and MKV work here even though the container reader never sees them.
  2. The reader is fetched on demand. The MP4 reader is a 185,285-byte script that is not loaded until you actually hand it one of those three extensions. If you only ever convert MP3s, you never download it.
  3. The file is parsed and its audio tracks are listed. If the list is empty, it stops right there and tells you the video has no audio track to convert. This is the failure that takes under a second.
  4. The first audio track is selected and its frames are extracted, in batches of up to 100,000 samples at a time. The video track is left alone; it is never decoded, never scaled, never touched.
  5. The codec configuration is located inside the file, in an esds box — that is where an MP4 keeps the sample rate, channel count and profile of its audio. If the box cannot be found, the reader falls back to the track metadata instead.
  6. The raw frames are re-wrapped into a fresh stream. Each frame gets a new 7-byte ADTS header carrying the codec parameters found in step 5, and the reassembled stream is handed to your browser's audio decoder. This is the step that explains the timing shift below: the rebuilt stream carries the samples, but not the container's edit list.
  7. Extraction is polled for up to 15 seconds, because the reader library has no "finished" event to wait on. It simply checks whether the frames that have arrived match the number the file says it has.
  8. Only then does the audio reach the transcriber, down-mixed to mono and resampled to 44,100 Hz, which is the rate the model expects. From this point on, a video file and an MP3 file are indistinguishable to the rest of the page.

Two consequences follow directly. First, the video resolution cannot affect your MIDI — by the time anything musical happens, the picture is already gone. Second, the audio codec is the thing that decides whether it works at all, not the container: the reader will happily extract frames from any MP4, but the browser still has to be able to decode what those frames contain.

The one real difference: a uniform 22 to 30 ms shift

The notes were identical, but their positions were not. Comparing every note against the WAV master:

RouteMean onset shiftSpread across the 40 notesMax pitch differenceNotes matching the WAV exactly
MP3 192 kbps0 ms20 ms1 cent35 / 40
M4A, AAC-LC0 ms20 ms1 cent35 / 40
WebM, Opus0 ms10 ms1 cent22 / 40
WebM, Vorbis0 ms10 ms1 cent16 / 40
MP4, 44.1 kHz+30 ms10 ms2 cents0 / 40
MP4, 48 kHz+22 ms10 ms2 cents0 / 40
MP4, 1920×1080+29 ms10 ms2 cents0 / 40
MOV+29 ms10 ms2 cents0 / 40
M4V+30 ms10 ms2 cents0 / 40

The MP4, MOV and M4V routes are uniformly late by roughly a quarter of a tick grid — every note, not just the first one, and with a spread of 10 milliseconds or less across all 40. It does not accumulate and it does not grow over the length of the file. The range is 22 to 30 milliseconds depending on the sample rate of the audio track: 21.75 ms for the 48 kHz AAC track and 29.50 ms for the 44.1 kHz one. At the 480 ticks per quarter note this page writes, that is 21 to 28 ticks.

The cause is step 6 above, and it is worth stating plainly because it is the only thing on this page you might need to correct for. The MP4 route rebuilds the audio stream from the raw codec frames and hands that to the browser. A rebuilt stream has the samples but not the container's edit list, and the edit list is where a file records that a decoder should skip the encoder's priming samples at the start. Without it, the audio starts a little late — and the measured 22 to 30 milliseconds is the same order as one AAC frame of priming, 1,024 samples, which is 23.2 milliseconds at 44.1 kHz.

M4A is also AAC in an MP4-style container, and it shows zero shift, because it is handed to the browser whole rather than rebuilt. That contrast is the cleanest evidence that the shift belongs to the rebuilding step and not to AAC itself.

What to do about it. If the MIDI only needs to be musically right, do nothing — 30 milliseconds is still under a 64th note at 120 BPM and no listener will hear it. If it needs to line up with the video frame-accurately, you have two options: select all notes in your DAW and nudge the whole track earlier by the amount measured here, or extract the audio track to WAV or MP3 first and convert that instead, which removes the shift entirely. The second option costs you one export step.

What does not convert, and what it tells you

Three of the fourteen packages failed. All three stopped in a fraction of a second — compare that with the roughly 28 seconds a real conversion takes, and you can see that the failure happens before any transcription work begins. We tested four more containers afterwards, on the same audio. Failure times vary by a few hundred milliseconds from run to run; the values below are from a single run on one machine.

FileResultMessageTime
MP4 with an MP3 audio trackFailedThe audio track in that video could not be decoded (unsupported codec).628 ms
MP4 with an AC-3 audio trackFailedThe audio track in that video could not be decoded (unsupported codec).228 ms
MP4 with no audio trackFailedThat video file has no audio track to convert.196 ms
AVI, MPEG-4 video + MP3 audioFailedThis browser could not decode that file. Try converting it to WAV or MP3 first.335 ms
FLV, H.264 + AACFailedThis browser could not decode that file. Try converting it to WAV or MP3 first.230 ms
WebM with no audio trackFailedThis browser could not decode that file. Try converting it to WAV or MP3 first.274 ms
MKV, H.264 + AACConverted40 notes, D4 – A4, 432 B file27.5 s

Notice the two different messages, because they mean different things. "The audio track in that video could not be decoded" comes from the container reader: it found the audio track, extracted the frames, and the browser then refused the codec. "This browser could not decode that file" comes from the other path entirely: the file was never recognised as a container this page reads, so it went straight to the browser's audio decoder and stopped there. AVI and FLV are in the second group; an AC-3 track inside a perfectly ordinary MP4 is in the first.

MKV is the surprise of the set. It is not one of the three extensions the container reader handles, yet it converted — because the browser itself can decode Matroska audio, and so the file went through the same path a WebM does. If you are choosing a container to keep a recording in and the audio inside it is AAC, Opus or Vorbis, MKV is a safe choice here even though nothing on the page promises it.

A video with two audio tracks: you get the first one

Plenty of real files carry more than one audio track — a film with an original-language track and a dub, a concert recording with stereo and surround, a screen capture with system audio and a microphone. There is no track picker on this page, so we built a file to find out which one it reads. Audio track 1 held the D4–A4 line; audio track 2 held the same line an octave up, D5 to A5. They cannot be confused.

The output was the 40 notes of track 1, D4 to A4, in 33.0 seconds. Track 2 contributed nothing — not a note, not a fragment. The rule is simply the first audio track in the file, which is what step 4 above says in code and what the measurement confirms in practice. If the track you want is not first, extract it before converting; there is no setting to change here.

Should you convert the video, or extract the audio first?

The honest answer is that it depends on one thing: whether you need the timing to be exact.

What you do not need to do is convert the video to MP3 first because you assume a video is harder. It is not. It is a container with an audio track inside it, and that audio track is the only part of the file that reaches the transcriber.

What the file it writes actually is

Unchanged by any of the above, because the container stops mattering the moment decoding finishes:

FormatStandard MIDI File, format 0 — one track, every note in it
Resolution480 ticks per quarter note
TempoMeasured from the note onsets and written into the file; 500,000 µs per quarter note (120 BPM) only when no steady pulse can be found
Time signature4/4, a constant
Track namemp3 to midi
ChannelsOne — a video's second audio track does not become a second MIDI track
Pitch bend range±2 semitones, set explicitly with an RPN 0 message, centre 8192
Velocityclamp(amplitude × 127, 1, 127), carried over from the model's own confidence value rather than measured from the audio
Note-off velocityFixed at 0x40 (64) on every note
Shortest note0.02 s — anything shorter is clamped up to it
Timingticks = seconds × 8 × BPM, so positions are absolute times rather than bar numbers — 960 ticks per second at the 120 BPM default

That timing row is worth one more sentence, because it is what makes the timing shift survivable. Notes are written as absolute times converted to ticks, not snapped to a grid, so a uniform offset stays uniform — every note is late by the same amount, and shifting the whole track in your DAW fixes all 40 of them at once. If the positions had been quantised, the same shift would have moved some notes onto the wrong side of a grid line and mangled the rhythm instead.

Everything above happens inside your browser tab. The container reader is fetched only when you use it, the transcriber is a small ONNX model that is downloaded after the first paint and cached, the audio is down-mixed to mono and resampled to 44,100 Hz before the model sees it, and nothing is uploaded — which is why a 5 GB video with a 20-second line in it costs you nothing but the wait.

Frequently asked questions

Does converting a video to MIDI work the same as converting the audio?

Yes, with one measurable difference. In our test, the same 20-second line packaged as an MP4, an MOV, an M4V, a WebM and an MKV all returned exactly the same 40 notes over the same range, D4 to A4, and all wrote a 432-byte MIDI file. The only difference was timing: the MP4, MOV and M4V versions placed every note 20 to 30 milliseconds later than the plain WAV did. MP3, M4A and WebM had no such shift. So the container does not change which notes you get, but the MP4 route is a few tens of milliseconds late.

Which video formats can be converted to MIDI?

MP4, M4V and MOV are read by this page's own container reader, which pulls the audio track out of the file. WebM and MKV happen to work too, but through your browser's own decoder rather than that reader. AVI and FLV do not work: in our test both failed in under 400 milliseconds with the message "This browser could not decode that file." The audio codec matters more than the container. AAC-LC, Opus, Vorbis and MP3 audio all converted here; an MP4 carrying an AC-3 track failed with "The audio track in that video could not be decoded (unsupported codec)."

Why is my converted MIDI shifted slightly compared to the video?

Because the MP4 reader rebuilds the audio stream before decoding it, and in doing so it does not apply the container's edit list — the part that tells a decoder to skip the encoder's priming samples. We measured the result: every note came out 20 to 30 milliseconds late compared with the same audio converted as a WAV, uniformly, with a spread of 10 milliseconds or less across all 40 notes. At the 480 ticks per quarter note this page writes, that is 21 to 28 ticks. If you need the MIDI to line up with the video exactly, shift the whole track earlier in your DAW — about 30 milliseconds for a 44.1 kHz track and about 22 for a 48 kHz one — or convert an extracted audio file instead of the video.

Do I need to extract the audio from the video first?

Not for the notes — the result is the same 40 notes either way. There are two reasons to extract anyway. The first is timing: extracting to WAV or MP3 avoids the 20 to 30 millisecond shift the MP4 route introduces. The second is control: a two-hour video will be read from beginning to end, and extracting a section first means you convert only the part you care about. There is one reason not to bother, and it is the interesting one: the audio is decoded exactly once here, whereas exporting it to another file first means decoding it, re-encoding it and decoding it again.

My video file is enormous. Does the video size matter?

The video stream is never decoded, so its size costs you almost nothing. We converted the identical audio inside a 640x360 file (1,037,093 bytes) and inside a 1920x1080 file (5,098,367 bytes) — the second file was 4.92 times larger and converted in 27.0 seconds against the first file's 28.5 seconds. The difference is noise. What does cost time is the length of the audio, because the model works through the audio in fixed windows; a two-hour video is a two-hour audio job.

My video has more than one audio track. Which one gets used?

The first one in the file, and the others are ignored entirely. We built an MP4 with two audio tracks — track 1 holding a line from D4 to A4, track 2 holding the same line an octave higher, from D5 to A5 — and converted it. The output was the 40 notes of track 1, D4 to A4, with nothing at all from track 2. There is no track picker on the page, so if the track you want is not first, extract it before converting.

What happens if the video has no audio track at all?

You get a clear message rather than a broken file. A silent MP4, a screen recording with the microphone muted, a video whose audio was stripped: in our test it failed in under 200 milliseconds with "That video file has no audio track to convert." Nothing was downloaded, nothing was half-written, and no notes appeared. The same is true of the other failure modes — they all stop early and say what went wrong, which is why the failures above took a fraction of a second while a real conversion took around 28 seconds.

Related: which audio format converts best, why the tempo is wrong when you open the file, and what happens when the audio contains drums.