Batch Audio to MIDI: Why One at a Time, and How to Go Faster
Last updated 9 October 2026
There is no batch mode here, and it is not a gap in the interface — the tool is built as one file in, one file out. Handed three files at once, it converts one. The lever you do have is the tab: measured on this machine, five 10-second files converted one after another in a single tab took 66.0 seconds, while the same files converted with the page reloaded between each one cost 24.7 to 25.8 seconds each — because every fresh page spends 12.1 seconds building the transcription engine before it can touch any audio.
Everything below was measured by driving the real page in a real browser. The test file is a 10-second WAV (44,100 Hz mono, two notes per second, 19 notes detected), converted repeatedly. Conversion time is timed from the moment the file is handed to the page to the moment the download link appears; start-up is timed separately, from page load to the moment the panel reports the engine ready.
One file in, one file out — checked three ways
The page gives you three ways to hand it a file, and all three take exactly one.
| The file picker | An <input type="file"> that lists 17 accepted types and carries no multiple attribute. That attribute is what tells the browser's picker to allow a multi-selection, so the picker itself will not let you choose a second file. |
| The drop target | The page-wide drop handler reads e.dataTransfer.files[0]. Anything else in the dropped list is ignored. |
| Paste | The paste handler reads e.clipboardData.files[0] — again, one file. |
The selection is then read the same way. The shipped page script contains five references to files[0] and zero references to files[1] or anything above it. For anyone who wants to check: the whole selection path is if (fileInput.files && fileInput.files[0]) handleFile(fileInput.files[0]);, in 40,927 bytes of page script.
To separate the interface from the code we ran the decisive version of the test. We gave the input a multiple attribute at runtime — something the page never sets — and handed it three files at once. The input then held all three files at the moment its change event fired, and the page still converted one: one analysis, one result, 19 notes. So it is not only the picker refusing; the code below it is written for a single file too.
That also means there is nothing to queue into. There is one engine per page, one result panel and one download button, and a second conversion reuses those slots rather than adding to a list.
The measured batch: five files, one tab
Engine reported ready, then five conversions of the same 10-second file, back to back, in that tab.
| File | Conversion time | Notes written | .mid size |
|---|---|---|---|
| 1 | 16.2 s | 19 | 244 B |
| 2 | 13.2 s | 19 | 244 B |
| 3 | 12.1 s | 19 | 244 B |
| 4 | 12.4 s | 19 | 244 B |
| 5 | 12.1 s | 19 | 244 B |
| Total | 66.0 s | — | — |
The first file is the slowest, and it is not the start-up: the engine had already reported itself ready 12.5 seconds earlier. It cost 16.2 seconds against an average of 12.5 for the four that followed. After that the cost is flat — 12.1, 12.4, 12.1 — so nothing accumulates. The fifth file in a batch costs the same as the second, and each conversion writes its own 244-byte file.
What "one at a time" really costs: the start-up is paid per page
Now the same file, but with the page loaded fresh each time. Start-up is measured on its own, before any file is handed over.
| File | Start-up first | Conversion | Total |
|---|---|---|---|
| 1 | 12.8 s | 13.0 s | 25.8 s |
| 2 | 11.8 s | 13.0 s | 24.8 s |
| 3 | 11.7 s | 13.0 s | 24.7 s |
The conversion itself barely moves — 13.0 seconds every time. What changes is the 11.7 to 12.8 seconds in front of it, which is 48% of the 25.1-second average. In a tab that has already converted something, the same file costs 12.5 seconds. Keeping the page open is therefore worth about 12.7 seconds per file — half the wall clock, on every file except the first.
That start-up is not mainly a download. The model weights are 51,054,395 bytes — about 48.7 MiB — and the browser caches them, so a return visit skips the transfer almost entirely: a repeat visit moves 118,352 bytes against 56,228,831 bytes on the first visit, roughly 1/475 of it. The wait is dominated by setting up the engine, which is why it barely improves on the second visit even though the download has already been paid for.
The trap: starting the next file deletes the previous result
A batch is only worth anything if you keep the output, and this is where the single result slot bites. The download link is a temporary blob URL, and the page revokes it the moment the next file arrives — before it even starts converting.
Measured: after a conversion the download link resolved to the 244-byte .mid. We started the next conversion without saving. 1.2 seconds later, fetching the first file's link failed outright — Failed to fetch — and the result panel was hidden again while the new conversion ran.
The rule that follows is short: click Download before you hand over the next file. The panel is going to be cleared either way; the file you already saved is not.
Handing over a second file mid-run does not run them in parallel
While a conversion runs, the dashed drop box is made unclickable, so you cannot open the file picker from it. But the page-wide drop and paste handlers carry no such guard, and a script can set the input directly — so "what if a second file arrives mid-run?" is a real question, and we tested it: a 30-second file, then a 10-second file handed over 1.5 seconds later.
The second conversion did not begin until the first had finished. The second file's first line of progress is timestamped 53.5 seconds into the page's life, and at that moment the first conversion was still reporting "Writing MIDI…". Total wall clock for the pair was 57.8 seconds — the sum of the two, not the longer of the two. Two results appeared, in order, and the panel ended up showing the second one; the first was already unreachable by the rule above. No errors were logged, and both files converted normally on their own terms.
So there is no way to overlap work inside one tab. Feeding the next file early saves you nothing; it only costs you the result you had.
Two tabs at once? Not a way to double throughput
Since one tab converts one file at a time, the obvious next idea is two tabs. Two things happen, and only the first is good news. The second tab is ready much faster — 5.3 seconds against 13.0 — because the model weights are already in the browser's cache by then. But both tabs build a full engine: the same 48.7 MiB of weights plus their own set of inference sessions. On one machine they compete for the same CPU.
| Situation | File | Time |
|---|---|---|
| One tab open, nothing else running | 10 s | 12.1 – 13.2 s |
| Two tabs open, the second idle, converting in the first | 10 s | 25.1 s and 29.2 s |
| Two tabs, one file each, started at the same moment | 10 s | 28.1 s / 17.0 s |
Two files at the same time took 28.1 seconds of wall clock, against about 24.5 seconds for the same two files one after another in a single tab (12.1 and 12.4 seconds). Two tabs was not faster — it was slightly slower, and the two tabs did not even split the machine evenly: 28.1 seconds against 17.0 for the same work.
The control run matters here. With a single tab open and nothing else running, the same file converted in 13.2, 12.8 and 12.5 seconds. So the 25-to-29-second figures are the cost of the second engine, not a slow machine.
Do the files affect each other? No — and the output is nearly, but not exactly, repeatable
We converted the same 10-second file five times in one tab and compared the .mid files event by event.
| Run | Time | Bytes | Events | Notes | Compared with run 1 |
|---|---|---|---|---|---|
| 1 | 12.5 s | 244 | 46 | 19 | baseline |
| 2 | 12.4 s | 244 | 46 | 19 | 1 event differs |
| 3 | 12.4 s | 244 | 46 | 19 | identical |
| 4 | 12.2 s | 244 | 46 | 19 | identical |
| 5 | 12.1 s | 244 | 46 | 19 | identical |
Four of the five files were byte-identical. The fifth differed in exactly one event: a note-on that landed 10 ticks later than in the other four — 10.4 ms at the 120 BPM the file declares. Every run reported the same analysis: 19 notes, mono mode, tempo 120 BPM, confidence 0.55. The 46 events are the 19 notes and their note-offs, four controller messages that declare the pitch-bend range, and the meta events that carry the tempo, time signature and track name.
Two conclusions, and they point in opposite directions. Nothing carries over between conversions — nothing from file 1 leaks into file 5, which is what you want from a batch. But the model is not bit-exact either, so expect a note to land a few milliseconds differently once in a while. Do not read a one-off difference as batch corruption; check it against a single conversion of the same file before you blame the run order.
The manual batch recipe
- Open the converter once and leave that tab open. The engine start-up is per page load — about 12 seconds. Ten files in one tab pay it once; ten reloads pay it ten times.
- Convert one file, then download it before you touch anything else. The previous download link dies the moment the next file arrives.
- Feed the next file by clicking the drop box. It becomes clickable again as soon as the previous conversion ends, and dropping a file anywhere on the page or pasting one works too.
- Do not reload between files, and do not open a second tab to run the next file in parallel. Everything you gain by batching is the start-up you avoid, and two engines on one machine fight over the same CPU — measured, that made both tabs slower than doing the two files in sequence.
- Trim each file to the section you need. Conversion time tracks audio length, so a 30-second excerpt is a fifth of the wait of a three-minute file.
- Save as you go, not at the end. Only one result is held at a time, so the only file you can recover after a long batch is one you already downloaded.
- If you are automating it: set the file input directly and wait for the download link to change, then read the blob. That is how every timing on this page was taken.
Nothing here is metered. There is no upload, no account and no daily limit, so a 20-file batch has no quota to run into — the only thing it costs is wall-clock time on your own machine, at roughly 12.5 seconds per 10 seconds of audio.
Frequently asked questions
Can I convert several MP3 files to MIDI at once?
No. The converter takes one file per conversion. The file input declares 17 accepted types but no multiple attribute, so the system file picker will not let you select a second file, and the selection logic reads files[0] — the shipped script contains five references to files[0] and none at all to files[1] or beyond. To separate the interface from the code we gave the input a multiple attribute at runtime and handed it three files: the input held all three at the moment its change event fired, and the page still converted one. There is no queue to join, either — one engine per page, one result panel, one download button.
Why can't I select more than one file in the picker?
Because the input element does not carry the multiple attribute, and that attribute is what tells the browser's file picker to allow a multi-selection. It is not a browser limitation or a permission you can grant — a file input without multiple is defined to accept exactly one file. The same one-file rule runs through the other two ways of feeding the tool: the page-wide drop handler reads e.dataTransfer.files[0], and the paste handler reads e.clipboardData.files[0].
How much slower is it to convert files one at a time?
For a 10-second file, about 12.5 seconds per file if you keep one tab open and feed files into it, against 24.7 to 25.8 seconds per file if the page is reloaded between files. The difference is the engine start-up, which a fresh page pays before it can touch any audio: 11.7 to 12.8 seconds in three runs, or 48% of the average 25.1-second total. Measured: five 10-second files back to back in one tab took 66.0 seconds in total; the same file converted in a fresh tab each time cost 25.8, 24.8 and 24.7 seconds.
Why did my previous MIDI file disappear?
Because the download link is a temporary blob URL that the page revokes the moment the next file arrives, and the result panel is hidden again while the new conversion runs. Measured: after a conversion the download link resolved to a 244-byte .mid file; we started the next conversion without saving, and 1.2 seconds later fetching the first file's link failed outright with Failed to fetch. The rule that follows is simple — click Download before you hand over the next file.
What happens if I start a second file while the first is still converting?
It does not run in parallel. We converted a 30-second file and then handed the tool a 10-second file 1.5 seconds later. The second conversion did not begin until the first had finished: the second file's first line of progress is timestamped 53.5 seconds into the page's life, at which point the first conversion was still reporting Writing MIDI. Total wall clock for both was 57.8 seconds, which is the sum of the two rather than the longer of the two. Two results appeared in order and the panel ended up showing the second one — the first was already unreachable by then.
Do the conversions affect each other?
No. We converted the same 10-second file five times in one tab and compared the .mid files event by event. All five were 244 bytes and contained 46 events — 19 notes, their note-offs, four controller messages and the meta events. Four of the five were byte-identical, and the fifth differed in exactly one event: a note-on that landed 10 ticks later, which is 10.4 ms at the 120 BPM the file declares. Nothing carries over from one file to the next. The model is simply not bit-exact, so do not read a one-off difference as batch corruption.
What is the fastest way to convert 20 files?
Keep one tab open and feed files into it one after another, saving each .mid before you start the next conversion. That pays the roughly 12-second engine start-up once instead of 20 times, which is worth about 12.7 seconds per file, and it avoids the trap where starting the next file revokes the previous download link. Trim each file to the section you need, because conversion time tracks audio length. And do not open a second tab to try to double the throughput: two engines on one machine share the same CPU, and in our test two files converted at the same time took 28.1 seconds against about 24.5 seconds for the same two files run one after another in a single tab. There is no upload, no account and no daily limit, so the only cost of a long batch is wall-clock time on your own machine.
Related: how long a single conversion takes and why, how long a file you can convert before the browser gives up, and what a cloud converter would cost you instead.