mp3→midi runs in your browser · nothing is uploaded

How Long Does Audio-to-MIDI Conversion Take?

Last updated 8 October 2026

Budget about 1.4 seconds of conversion for every second of audio, plus a one-time start-up of 12 to 14 seconds. Measured on this machine, with the engine already loaded: a 30-second clip took 44 seconds, one minute took 84 seconds, two minutes took 3 minutes 1 second, and three minutes took 4 minutes 41 seconds. The price per second of audio is not constant, and the reason is specific: the engine works in 8-second windows and feeds each one a second of context on either side, so it processes 1.2 to 1.24 times the audio you gave it.

Those figures come from seven conversions run back to back in the same browser tab, on audio built to be exactly 5, 10, 20, 30, 60, 120 and 180 seconds long and to contain the same music at the same density throughout — two notes per second in every file. Timing starts the moment the file is handed to the page and stops when the download link appears.

The short version Time is almost entirely inference, and inference is almost entirely proportional to audio length. What breaks the proportionality is the windowing: 8 seconds of output per window, 10 seconds of audio fed, and the length rounded up to whole windows. A 9-second clip costs more model work than the arithmetic suggests, a 5-second clip costs less, and a new tab costs about 12 seconds before any of this begins.

The measured table

Desktop Chrome, no other work on the machine, engine loaded and warm. "Audio fed" is what the engine's window planner actually sends to the model — it is the length of the file plus the context around each window.

Audio lengthWindowsAudio fedAnalysisTotal waitms per second of audio
5 s15.0 s5.6 s6.4 s1,126
10 s212.0 s13.6 s13.9 s1,358
20 s324.0 s28.4 s29.1 s1,421
30 s436.0 s43.4 s44.1 s1,446
60 s874.0 s83.7 s84.4 s1,395
120 s15148.0 s180.5 s181.4 s1,504
180 s23224.0 s280.1 s281.1 s1,556

Read the last two columns together, because they are the whole answer. Total wait rises faster than audio length — from 1,126 ms per second of audio on a 5-second clip to 1,556 ms on a three-minute one. But the number of seconds the model is actually asked to analyse rose from 5.0 to 224.0 over the same range, and the cost per second of that stayed between 1,126 and 1,250 ms the whole way. Nothing is slowing down as files get longer. The engine is simply being handed more audio than the file contains.

One repeat, to show how much of this is machine noise rather than structure: a second, independent three-minute conversion in a fresh tab took 4 minutes 20 seconds instead of 4 minutes 41. That is an 8% spread on identical input, so treat every number here as a bracket, not a stopwatch reading.

Where the time actually goes

The page reports four phases, and only one of them matters. These are the same seven conversions, broken down.

Decoding the file0.1 to 0.5 s. Turning your MP3, WAV or M4A into raw samples. Grows with file length, stays trivial.
Preparing the audio0.02 to 0.22 s. Downmixing to mono and resampling to the 44,100 Hz the model expects. Effectively free.
Analysis99.6% of the time on the three-minute file — 280.1 of 281.1 seconds. This is the neural network running on the CPU.
Writing the MIDI0 ms that the clock can resolve. A 359-note file is a few kilobytes; encoding it is instant.

Everything outside the model — reading the file, resampling it, writing the .mid — came to between 0.3 and 1.0 seconds per conversion, whatever the length. If you were hoping to speed things up by feeding the tool a smaller file format, that is where the hope dies: the decode step you would be optimising is 0.2% of the wait.

Reason one: start-up is a fixed cost you pay per page, not per file

Before the tool can convert anything it has to fetch and initialise its model. That cost does not scale with your audio at all, which is why short clips feel disproportionately slow.

SituationWait before the tool is ready
Browser that has never loaded the model13.9 s
New tab, same browser, model already cached11.6 – 11.9 s
Second file in a tab that has already converted something0 s

The gap between the first two rows is the download: about 2 seconds on this connection. Almost all of the start-up is not fetching but initialising — five ONNX models are compiled and five inference sessions are created, sequentially, before the tool will accept a file. The weights themselves are 51,054,395 bytes, about 48.7 MiB, and they are held in the browser's cache, which is why the second row is only slightly faster than the first and why a repeat visit a week later skips the download entirely.

The practical consequence is that the arithmetic people do in their heads is wrong at the short end. A 5-second clip does not take a twelfth of the time of a one-minute file. It takes 6.4 seconds against 84.4 — because the one-minute file amortises the same start-up over twelve times as much audio.

Reason two: the engine works in 8-second windows with context

This is the part that shows up in the numbers and nowhere else. The engine does not analyse your file as one signal. It cuts it into windows of 8 seconds, and for each window it sends 1 second of audio from before and after as context, so the model can hear what is coming. Only the middle 8 seconds of each window are kept as output.

Two things follow. First, the audio fed to the model is longer than your file. Second, the file length is rounded up to a whole number of windows, so crossing an 8-second boundary adds a full extra window's worth of work at once. Here is what the planner does to a range of lengths:

Audio lengthWindowsAudio fed to the modelFed ÷ your audio
8 s18.0 s1.000
9 s211.0 s1.222
16 s218.0 s1.125
24 s328.0 s1.167
32 s438.0 s1.188
64 s878.0 s1.219

Look at the first two rows. Going from an 8-second clip to a 9-second clip adds 12.5% more audio and 37.5% more model work, because the ninth second forces a second window and that second window arrives with its own context. A 9-second clip is relatively more expensive than a 16-second one. This is the only place in the whole pipeline where the cost is genuinely discontinuous, and it is a boundary you can see: at 120 BPM in 4/4, one window is exactly four bars, and the context is half a bar on each side.

The same effect explains the slow drift in the per-second price. A file that fits in one window pays no context at all — 1.000×. As files get longer, the proportion of audio that is context settles toward the 25% ceiling, so the ratio climbs from 1.000 through 1.200, 1.233 and 1.244 and stays there. Measured, on a clean set of single-file runs, the cost per second of audio fed was 1.09 s at 10 seconds, 1.16 s at 60 seconds and 1.15 s at 180 seconds — flat, while the cost per second of your audio rose from 1.31 s to 1.44 s over the same span. The drift is entirely accounting.

Reason three: your machine, not the file

Everything above was measured on one desktop. To see how much of the result is hardware, we ran two of the same files again with the CPU throttled to a quarter of its speed — an emulated slow device, not a physical phone, and it should be read as a direction rather than a specification.

Audio lengthFull speedCPU throttled to 25%Ratio
10 s13.9 s25.7 s1.85×
30 s44.1 s87.2 s1.98×

The throttled run came out about twice as slow, not four times — 1.85× on the 10-second file and 1.98× on the 30-second one. The shape of the curve was unchanged: the 30-second file cost 3.2 times the 10-second file at full speed and 3.4 times when throttled. So if you are on a slower machine, scale the whole table and keep the relationships.

What moved least was start-up. For the same cached-model case, the engine took 14.9 seconds to become ready with the CPU throttled against 11.6 to 11.9 seconds at full speed — about 25% slower rather than twice as slow. Most of that wait is fetching and creating five inference sessions rather than computing, so it is the part of the pipeline a slow CPU hurts least.

What the progress bar is really telling you

The bar moves once per window — 23 steps for a three-minute file, 4 for a 30-second one — and the time estimate next to it is derived from how long the work has taken against how far the bar has got. That method has one structural flaw: progress is reported just before a window runs, not after it finishes, so the estimate is always computed from one window less work than the bar claims.

Here is the three-minute run, panel against reality:

ProgressPanel saidActually remainingError
31%~2:313:5335% low
46%~2:333:0417% low
68%~1:411:519% low
87%~47s0:448% high
98%~12s0:05164% high

The estimate converges as the run goes on and then overshoots at the end, for the same reason in both directions: it assumes every window costs the same, and the first window costs more than the rest while the last one is usually short. On the 10-second file the flaw is at its worst — the panel said ~1s left at 55% while 13.6 seconds of work remained, because the estimate was computed from the six milliseconds of elapsed time between the analysis starting and the first window being handed off.

Two details of the display are worth knowing before you read anything into it. The bar never shows 0% or 100% during analysis; it is mapped across 12% to 98%. And any estimate under one second is printed as "~1s left" — a floor in the formatting, not a measurement. If you see that, you know the estimate has collapsed, not that the job is nearly done.

Practical answers

Frequently asked questions

How long does a three-minute song take to convert?

Two runs of the same three-minute file took 4 minutes 41 seconds and 4 minutes 20 seconds on the machine we tested, with the engine already loaded. That is about 1.5 seconds of work for every second of audio. If the tab is new, add the one-time start-up: 11.6 to 11.9 seconds when the model files are already in the browser's cache, 13.9 seconds on a browser that has never loaded them. Run-to-run variation on the same file was about 8%, so treat these as a bracket rather than a promise.

Why does a 10-second clip take 14 seconds?

Because the engine never processes your file in one piece. It splits the audio into 8-second windows and feeds each window with 1 second of context on either side, so a 10-second file becomes two windows carrying 12 seconds of audio. Measured: 13.6 seconds of analysis for a 10-second file, against 5.6 seconds for a 5-second file that fits in a single window. The doubling is not in your file, it is in the windowing.

Does conversion time grow in proportion to the length of the audio?

Only above about a minute, and even then not exactly. The cost per second of audio was 1.13 s at 5 seconds of audio, 1.36 s at 10 seconds, 1.40 s at 60 seconds and 1.56 s at 180 seconds. The reason the price per second rises is that the model is fed 1.2 to 1.24 times the audio you supplied, because every 8-second window carries a second of context on each side. The cost per second of audio the model actually sees stayed between 1.09 and 1.25 seconds across every file we ran.

Why does the progress bar move in steps instead of smoothly?

Because there is nothing to report between windows. The engine hands off one 8-second window at a time and reports progress once per window, so a three-minute file moves the bar in 23 steps and a 30-second file in 4. There is no finer-grained signal available — the inference inside a window is a single opaque call.

Why is the time estimate so wrong at the beginning?

Because progress is reported before the window it refers to has run, and the estimate divides elapsed time by progress. On a three-minute file the panel said ~2:31 left at 31% while 3 minutes 53 seconds of work actually remained — 35% low. On a 10-second file it said ~1s left while 13.6 seconds of work remained. The estimate becomes accurate at about 83% of the way through, then overshoots, because the final window is usually a short one: the last window of the three-minute file carried 4.5 seconds of audio against roughly 12 seconds for a full one. Any estimate below one second is displayed as ~1s left.

How long does it take on a phone?

We did not test a phone, and we are not going to invent a number for one. What we can offer is a directional measurement: with the CPU throttled to a quarter of its speed, the same 10-second file went from 13.9 to 25.7 seconds and the 30-second file from 44.1 to 87.2 seconds — roughly twice the wall clock, not four times. The shape of the curve was unchanged, so on slower hardware expect every number on this page to scale, with the one-time start-up staying closest to constant because most of it is downloading and initialising rather than computing.

Can I make a conversion finish faster?

Four things help, in order of how much they save. Convert only the section you need: time tracks audio length almost linearly, so a 30-second excerpt is a fifth of the wait of a three-minute file. Keep the tab open and convert the next file in it, because the start-up is paid once per page load, not once per file. Come back to the same browser rather than a fresh one, because the model weights stay in the cache and the second visit skips the download. And trim to just under a multiple of 8 seconds rather than just over: a file of 9 seconds needs two windows and 11 seconds of model work, while an 8-second file needs one window and 8 seconds. No setting in the panel changes the speed — the clean-up controls re-run on notes the model has already produced, so they never touch the engine.

Related: how long a file you can convert before the browser gives up, what local conversion costs you in speed and saves you in privacy, and whether a denser file takes longer.