Local vs Cloud Audio-to-MIDI: Privacy, Speed, File Limits and Cost
Last updated 3 October 2026
We put a network recorder on this page and converted files through the real file input while logging every request. Across the whole session — a WAV, an MP3, an M4A and a 5.0 MB MP4 — the total uploaded request body was 0 bytes. Then we switched the network off at the browser level and converted two more files anyway: a 324,389-byte M4A in 26,284 ms and a 5,098,367-byte MP4 in 26,679 ms, both returning the same 40 notes over D4–A4 in a 432-byte MIDI file, both reporting 0 bytes down and 0 bytes up. The price of that is paid once, up front: the first visit downloads 56.2 MB of engine, and after that a return visit costs 118,352 bytes.
This page is a measurement, not a comparison of marketing pages. We did not sign up for any cloud converter and we are not quoting one — the right-hand column below describes the general shape of an upload-and-wait service, and the left-hand column is what we recorded on this one.
The comparison, with the local side measured
| Local — this page, measured | A cloud converter — general architecture | |
|---|---|---|
| Where your file goes | Nowhere. 0 bytes of request body across a full session of conversions | It is the request payload. The file leaves your device and lands on someone else's disk |
| Where the work happens | Your CPU, inside the tab, one file at a time | The vendor's machines, usually in a queue shared with other users |
| What it costs to start | 56.2 MB downloaded once, then cached | Nothing to download — but every job waits its turn |
| Recurring cost | Nothing is counted. There is no server-side job to meter | Almost always an allowance — per day or per month — with the free tier the tightest |
| Account | None. Nothing to sign up for and nothing to close | Commonly required, and often required just to get the file back |
| Size limit | None imposed by the page. Bounded by your device's memory | Bounded by the vendor's upload limit, which is usually stated in megabytes |
| Speed | 26.3–28.0 s per 20 s of audio on a 2-core, 4 GB machine. No queue | Depends on queue depth and their hardware. Nothing you can inspect or change |
| Offline | Works, once the engine has been fetched (measured, see below) | Impossible by construction — the work is not on your machine |
| What the operator sees | Nothing about your audio. No copy exists to be kept | The file, its name, its size, and whatever your account identifies |
The right-hand column is architecture, not a benchmark. No number in it is a measurement of any particular service — we did not test one. The left-hand column is what the recorder in this tab logged.
How the measurement was taken
The page was loaded from a local copy, and a network recorder was attached to the tab for the entire session. It logged every request with its method, its URL, the bytes actually transferred, and — the column that matters here — the size of any request body. An upload is a request with a body. Across the whole session the sum of those bodies was 0 bytes.
Files were handed to the page the same way a person hands it a file: through the file input, one at a time, waiting for the result or the error to appear. All conversions shared one page instance, so the 51 MB of model weights were loaded once rather than once per file. The machine reports 2 logical cores and 4 GB of device memory.
What the first visit downloads, and what the second one costs
The first visit made 16 requests and downloaded 56,228,831 bytes — 56.2 MB. A second, independent run measured 56,229,173 bytes across the same 16 requests. Here is where it goes:
| What | Requests | Bytes | Source |
|---|---|---|---|
| Model weights — five ONNX files | 5 | 51,085,616 | Hugging Face CDN |
| ONNX Runtime Web — one script, one module, one 4.7 MB wasm binary | 3 | 4,846,282 | jsDelivr |
| Google's analytics tag script | 1 | 178,581 | googletagmanager.com |
| This page's own HTML, CSS and four scripts | 6 | 118,352 | this site |
| Analytics beacon | 1 | 0 | google-analytics.com |
The largest single file is one of the model weights at 20,918,738 bytes. The whole engine — weights plus runtime — is 55,931,898 bytes, or 99.5% of what a first visit downloads. The page's own files, the part you actually see, are 118,352 bytes.
Now the second visit. Reloading the page produced 10 requests and 118,352 bytes — and those bytes are entirely this page's own files. The model weights: 0 bytes. The 4.7 MB runtime wasm binary: 0 bytes, served from the browser's disk cache. So a return visit costs 475 times less than the first one, and the number that matters for a person converting a file every day is the 118,352.
That 51 MB does have to live somewhere on your machine. After the engine loaded, the browser reported 51,055,714 bytes of storage used against a quota of 10,788,473,954 bytes — 0.47% of what the browser is willing to give this site.
The decisive test: convert with the network switched off
Privacy claims are cheap, so here is the version that cannot be faked. Once the engine was warm, the network was disabled at the browser level and conversions were run anyway:
| File | Size | Network | Result | Data in / out | Time |
|---|---|---|---|---|---|
| WAV, 20 s | 1,764,044 B | online | 40 notes, 432 B .mid | 0 B / 0 B | 28,037 ms |
| MP3 192 kbps | 481,532 B | online | 40 notes, 432 B .mid | 0 B / 0 B | 26,497 ms |
| M4A, AAC-LC | 324,389 B | disabled | 40 notes, 432 B .mid | 0 B / 0 B | 26,284 ms |
| MP4, 1920×1080 | 5,098,367 B | online (first video) | 40 notes, 432 B .mid | 185,478 B / 0 B | 27,699 ms |
| MP4, 1920×1080 | 5,098,367 B | disabled | 40 notes, 432 B .mid | 0 B / 0 B | 26,679 ms |
| WAV, 20 s | 1,764,044 B | disabled | 40 notes, 432 B .mid | 0 B / 0 B | 26,654 ms |
Every row produced the same 40 notes over the same range, D4 to A4, written into the same 432-byte file. The input sizes span 324,389 to 5,098,367 bytes — a factor of 15.7 — and the output does not move.
Three things follow from that table, and they are the whole answer to the local-versus-cloud question:
- The transcription does not need a server. If any part of it were remote, a 5.0 MB video could not have converted with the network down.
- File size barely affects the time. The 5.0 MB video took 26,679 ms; the 1.76 MB WAV took 28,037 ms. What costs time is the length of the audio — 20 seconds in every row — because the model works through it in fixed windows, not because of how many bytes the container happens to hold.
- There is no per-file cost to ration. Nothing was counted, because there is no server-side job to count. A limit only makes sense when somebody else's machine is doing the work.
What leaves your device, precisely
Being exact matters here, so: the page is not silent on the network. Inside the conversion windows themselves the log is nearly empty — a single request per file, which is the browser reading back the output blob it just created, at 0 bytes. Over the whole session, the only outbound request that was not a GET was this site's own analytics beacon: a POST with an empty body, whose parameters ride in the URL. It appears on every page of this site, it fired twice in the log, and in neither case did it carry anything derived from the file being converted.
Two consequences worth stating plainly. First, your audio contributes 0 bytes to that, whether it is a 324 KB M4A or a 5.0 MB video. Second, blocking it changes nothing about the conversion — the offline rows above were converted with the network switched off entirely.
What the video path costs that the audio path does not
One honest wrinkle came out of the offline test, and it is the kind of thing a comparison table hides. The MP4 reader — a 185,285-byte script — is not part of the engine and is not downloaded with it. It is fetched only when you first hand the page a video file.
In our first session, no video was converted until the network had already been cut. The result: the MP4 failed in 386 ms with "MP4 reader failed to load." — a failure of the download, not of the transcription. A second session converted one video while online first (185,478 bytes over the wire for that reader), and then converted the same 5.0 MB file with the network disabled: 26,679 ms, 40 notes, 432 bytes, 0 bytes in and out. So if you convert video on a plane, do it once while you still have a connection.
This is also the answer to a reasonable question: why not cache everything up front? Because the reader is dead weight for the majority of visits that only ever convert audio, and the page fetches it at the moment it becomes necessary rather than making every visitor pay for it.
What local conversion genuinely costs you
The local side of the table is not free, and the price is worth naming rather than burying:
- The first visit needs a connection. With the network off, loading the page fails outright. 56.2 MB has to arrive before the first conversion can happen.
- Your CPU is the ceiling. 26.3–28.0 seconds per 20 seconds of audio, on a machine reporting 2 cores and 4 GB. A faster machine is a faster conversion; a slow one is a slow conversion, and there is no server to fall back on.
- One file at a time. There is no queue and no parallelism, because there is no pool of machines to spread work across.
- Memory, not a quota, is the real limit. The browser's JavaScript heap sat between 8.7 and 15.3 MB after a conversion, with the model's own tensors living outside it in the WebAssembly heap. Long files are where this bites.
- 51 MB of your browser storage. 0.47% of the quota here, but it is storage you are lending, not the vendor's disk.
Set against that: nothing is uploaded, nothing is metered, there is no account, and there is no vendor whose pricing page can change your workflow next month.
Which one should you actually use
- Use local when the audio is something you would rather not hand to a stranger — an unreleased demo, a client's voice memo, a rough take. The offline test is the guarantee: it ran with the network down.
- Use local when you convert a handful of files rather than hundreds. 56.2 MB once, then 118,352 bytes per return visit, is cheaper than an account you will use twice.
- Use local when you want to convert now and not after a queue. The 20-second file above finished in 26 seconds; nobody else's job was ahead of it.
- Use a cloud service when you need something this page does not do — stem separation into several instruments, a batch of a hundred files, a rendered score rather than a .mid. Those are real reasons, and they are not about privacy.
- Use a cloud service if your machine is genuinely the bottleneck. A phone converting a five-minute song will be slower than a rented server doing the same thing; that is the trade you are making.
What the file it writes actually is
Identical in both modes, because none of the above touches the writer:
| Format | Standard MIDI File, format 0 — one track, every note in it |
| Resolution | 480 ticks per quarter note |
| Tempo | Measured from the note onsets and written into the file; 500,000 µs per quarter note (120 BPM) only when no steady pulse can be found |
| Time signature | 4/4, a constant |
| Track name | mp3 to midi |
| Pitch bend range | ±2 semitones, set explicitly with an RPN 0 message, centre 8192 |
| Velocity | clamp(amplitude × 127, 1, 127), carried over from the model's own confidence value rather than measured from the audio |
| Note-off velocity | Fixed at 0x40 (64) on every note |
| Shortest note | 0.02 s — anything shorter is clamped up to it |
| Timing | ticks = seconds × 8 × BPM, so positions are absolute times rather than bar numbers — 960 ticks per second at the 120 BPM default |
Everything on this page happens inside your browser tab. The engine is a set of ONNX models executed with ONNX Runtime Web, the audio is down-mixed to mono and resampled to 44,100 Hz before the model sees it, and the whole chain — decode, analyse, write — runs on your machine, which is exactly why it keeps working when the network does not.
Frequently asked questions
Is my audio file uploaded anywhere when I convert it?
No. We attached a network recorder to the page and converted four files through the real file input — a WAV, an MP3, an M4A and a 5.0 MB MP4 — while logging every request, its method and the size of its body. The total uploaded request body across the entire session was 0 bytes. The only non-GET request in the whole run was this site's own analytics beacon, which carries an empty body and is not related to your audio. Your file never becomes a request payload.
Does this converter still work with no internet connection?
Yes, once the engine has been downloaded once. We disabled the network at the browser level and converted anyway: a 324,389-byte M4A finished in 26,284 milliseconds, and a 5,098,367-byte MP4 finished in 26,679 milliseconds, both returning the same 40 notes over D4 to A4 in a 432-byte MIDI file. Both of those conversions reported 0 bytes downloaded and 0 bytes uploaded. A first visit is the exception: with the network off, the page itself cannot load, because the 51 MB engine has to arrive before anything can be transcribed.
How much does the first visit download?
56.2 MB, in 16 requests, measured in two independent runs. Of that, 51,085,616 bytes are the five model weights pulled from a Hugging Face CDN, 4,846,282 bytes are ONNX Runtime Web from jsDelivr, 178,581 bytes are Google's analytics tag script, and 118,352 bytes are this page's own HTML, CSS and JavaScript. A return visit downloads only the 118,352 bytes of page files — the model weights and the runtime come from the browser cache, so the second visit costs about 475 times less than the first.
Is there a file size limit or a daily conversion limit?
Nothing on the page counts your conversions or caps your file size, because there is no server doing the work to meter. The limit is your own machine. On a 2-core, 4 GB machine, a 20-second file took 26.3 to 28.0 seconds regardless of whether the input was a 324 KB M4A or a 5.0 MB video — what drives the time is the length of the audio, not the size of the file. The engine occupies 51,055,714 bytes of browser storage, which is 0.47% of the 10.0 GiB quota the browser reports.
What is the downside of converting locally instead of in the cloud?
Your own machine does the work, and that has three visible costs. The first visit has to download the engine, 56.2 MB, before the first conversion can happen. There is no queue and no parallelism — one file at a time, on your CPU, and a slow laptop is a slow conversion. And long files are bounded by your device's memory rather than by a server's capacity. What you get in exchange is that nothing is metered, nothing is uploaded, and no account exists to be closed.
Why does the engine need 51 MB when the page itself is tiny?
Because the page's own files total 118,352 bytes and the model is a separate download. The transcriber is five ONNX files totalling 51,085,616 bytes, fetched after the first paint and then cached by the browser, which is why the interface appears immediately and the engine note says it is still loading. Inlining 51 MB into the HTML would delay the first paint by the whole download, so the engine is fetched in the background instead.
Do I need an account or a subscription to convert a file?
No account, no email, no key, and no per-conversion counter, because there is no server-side job to attach any of those to. The measured evidence is the offline test: a 5.0 MB video converted with the network disabled and 0 bytes uploaded, which is not possible if a remote service is doing the transcription or gating the download.
Related: how long a file you can convert in a browser, what happens to the audio track in an MP4, and why the tempo is wrong when you open the file.