Record to MIDI: What Works and What Does Not
Last updated 6 October 2026
Short answer: this tool cannot record. There is no microphone button, and no code behind one — we searched all 984,704 bytes of JavaScript the page loads and found zero references to any capture API, and the file input carries no capture attribute. What does work is the two-step route: record with something else, save the file, then drop the file in. We tested that route on the two file types real recorders actually produce — an AAC voice memo and a browser Opus recording — and both came back as 40 notes and a 432-byte .mid, byte-identical to each other and identical to the studio WAV of the same performance except for a single event at the very end.
Everything below is measured on this tool, in a browser, on the engine this site uses for vocals, humming and single melodies. The numbers are reproducible and the method is described so you can check the shape of it yourself.
First, the part that does not exist
Plenty of pages that convert audio let you sing into them. This one does not, and the useful thing to do is establish that as a fact rather than an impression, because a missing feature is otherwise easy to argue about.
We loaded the page in a headless browser, waited for the engine to finish loading, and then read two things: the input element as it exists in the live DOM, and the source of every script the page pulled in.
| JavaScript loaded by the page | 984,704 bytes total — 91,867 bytes of the site's own code (including 4,663 bytes of inline script) and 892,837 bytes of third-party code (the analytics tag and the ONNX runtime) |
| References to a capture API in that JavaScript | 0 — searched for getUserMedia, MediaRecorder, navigator.mediaDevices, enumerateDevices, createMediaStreamSource and AudioWorkletNode |
| Same check on the live build | Its own app.js is 74,028 bytes; capture-API matches: 0 |
| The input element | <input type="file">, accept list of 17 extensions, capture attribute present? No |
| Buttons on the page | Download .mid · Play preview · Convert another · Re-run clean-up · Reset to defaults |
| Does the visible text mention recording? | No — the words "record" and "microphone" do not appear in the page's text at all |
| How audio actually gets in | Four ways, all of them files: the file picker, drag-and-drop anywhere on the page, pasting a file from the clipboard, and the "Convert another" button which reopens the picker |
That capture attribute is the small detail worth understanding. A file input with capture set tells a phone to open the camera or the microphone directly. Without it — which is the case here — the phone opens its file browser instead. So the absence of that one attribute is the mechanical reason this page will never ask for microphone permission: it has nothing to ask with.
The route that does work
Recording elsewhere and converting the file is not a workaround with a catch. The accept list on the input already covers what recorders write, so there is no conversion step and no re-export:
- Record. Any recorder will do — a phone voice-memo app, a browser recorder, or a desktop capture tool. Get as close to the source as you can and keep it to one line at a time.
- Save it as a file you can find. On a phone, that usually means sharing or exporting the memo out of the recorder app and onto the same device you are converting on.
- Drop the file into the converter. Pick it, drag it onto the page, or paste it. It is read locally, so nothing is uploaded while you wait.
- Check the piano-roll preview before downloading. The note count and the pitch range are printed on screen. If the range is not the range you sang, the recording is the problem, not the converter.
We ran that route end to end
The honest question is whether a recording made this way survives contact with the converter, or whether the compressed audio a recorder writes costs you something. So we took one performance — a 20-second hummed line, 40 quarter notes at 120 BPM spanning D4 to A4 — and saved it three ways: as the original uncompressed WAV, as an AAC recording at 128 kbps in an M4A container, and as an Opus recording at 96 kbps in a WebM container. The last two are what a phone voice memo and a browser recorder hand you. Each was converted once, in the same browser session, through the real file input.
| File | Size | Notes returned | Pitch range | Output .mid | Time |
|---|---|---|---|---|---|
| WAV, 44.1 kHz mono (baseline) | 1,764,044 B | 40 | D4–A4 | 432 B | 27,361 ms |
| AAC 128 kbps in M4A, 48 kHz mono (phone voice memo) | 324,604 B | 40 | D4–A4 | 432 B | 26,712 ms |
| Opus 96 kbps in WebM, 48 kHz mono (browser recorder) | 335,144 B | 40 | D4–A4 | 432 B | 26,539 ms |
Three files, one performance, 40 notes out of every one of them, the same pitch range, and the same 432-byte output. The recording is 5.4 times smaller than the WAV and the note content is unchanged: the voice memo is 18.4% of the WAV's bytes and the browser recording is 19.0%.
The timing is worth a sentence too. Twenty seconds of audio took between 26.5 and 27.4 seconds to convert on this machine, so the work runs at roughly 1.3 times the length of the recording. The compressed inputs were not measurably faster or slower than the uncompressed one — the cost tracks how long the audio is, not how big the file is.
The byte-level comparison, because "the same" is a claim
"Both produced 40 notes" is a weak statement. The stronger one is what the files actually contain, so we read all three back and compared them event by event. A .mid file here holds 88 events: 40 note-on messages, 40 note-off messages, 4 controller messages that declare the pitch-bend range, and 4 meta events — the tempo, the time signature, the track name and the end-of-track marker.
- The two recordings produced byte-identical files. The AAC output and the Opus output have the same 432 bytes and the same SHA-256 hash. Not similar — identical.
- Against the WAV, the difference is confined to the last two events in the file. The final note-off sits at tick 19,190 in the recorded versions and tick 19,200 in the WAV version — 10 ticks. At the 120 BPM written into the file, one tick is about 1.04 ms, so the gap is about 10.4 milliseconds.
- The other 86 events are identical — every note-on, every note-off, every controller message, the tempo, the time signature and the track name. The second of the two events that move is the end-of-track marker, which always sits immediately after the last note-off and so travels with it.
The cause is not mysterious: an AAC encoder adds a little padding at the start of the stream, so the decoded audio ends a fraction of a millisecond to a few milliseconds earlier than the original. The converter saw a marginally shorter buffer and wrote the last note-off marginally earlier. A ten-millisecond difference in the release of the final note is not audible and is not worth chasing.
There is a useful side effect of running this comparison: the WAV output in this run has the same hash as an earlier, independent run of the same WAV in a different test — which is how we know the pipeline is deterministic. Same input, same output, every time.
What each recorder hands you
The format you end up with depends on what you record with, and all of the common ones are already in the accept list, so none of them needs converting first.
| Recorder | What it writes | Does it go straight in? |
|---|---|---|
| iPhone or iPad voice memo | M4A holding AAC audio | Yes — tested, 40 notes out |
| Browser recorder (MediaRecorder) | WebM holding Opus audio | Yes — tested, byte-identical to the AAC result |
| Android recorder | Usually M4A/AAC; sometimes OGG holding Opus | Yes — both are in the accept list |
| Desktop capture tool, e.g. Audacity | WAV by default | Yes — uncompressed, the shortest path to analysis |
| Screen or system-audio capture | MP4, MOV or WebM | Yes — the audio track is read and the video ignored |
If you have a choice, WAV is the safest pick, but the reason is narrower than people expect. It is not that compressed audio damages the transcription — in our test the compressed recordings produced the same notes as the uncompressed one. It is that WAV removes one thing that could go wrong, and it is the format every recorder can produce.
Where recording, rather than converting, is what fails
None of the failures below are the converter's fault, and all of them are visible in the piano-roll preview before you download anything.
- A room, not a source. A microphone in a room records the room as well as the performance. Reflections and background noise are pitched information, and a model looking for a melody has to find it in that. Record close to the instrument or the mouth, in the quietest space you have, and record one line at a time.
- Too quiet, or clipping. Aim for a healthy level that does not touch the top of the meter. A recording so quiet that the melody is near the noise floor, or so loud that it distorts, is damaged before the converter ever sees it.
- Several things at once. The engine on this page follows a single line. A phone held up in front of a band records a mix, and a mix is the hard case for a melody engine. This site offers separate engines for piano, guitar, bass and other instruments, and a separate route for a full mix — but for a recording, the cleanest fix is to record one part at a time.
- Waiting on a long take. Conversion runs at roughly 1.3 times the length of the audio and the tab has to stay open throughout. Record a short excerpt first, check the roll, then commit to the whole thing.
What the file it writes actually is
Whatever route the audio took to get in, the writer behaves the same way, and we confirmed this by reading the three output files back byte by byte.
| Format | Standard MIDI File, format 0 — one track, every note in it |
| Resolution | 480 ticks per quarter note |
| Analysis sample rate | 44,100 Hz, mono, after decoding and down-mixing — a 48 kHz stereo recording is resampled and mixed down before the model sees it |
| Tempo | Measured from the note onsets and written into the file; 500,000 µs per quarter note (120 BPM) only when no steady pulse can be found |
| Time signature | 4/4, a constant |
| Track name | mp3 to midi |
| Pitch-bend range | ±2 semitones, declared with an RPN 0 message, centre 8192 |
| Velocity | clamp(amplitude × 127, 1, 127) — on all 120 notes we measured here it came out as 90, because the engine reported the same confidence value for every one of them |
| Note-off velocity | Fixed at 0x40 (64) on every note |
| Shortest note | 0.02 s — anything shorter is clamped up to it |
| Pitch bends written | None, on every file we measured here — the range is declared but no bend data is emitted |
| Audio content | None. The file carries note events only — 0 bytes of samples |
One number in that table is worth pulling out, because it is the one people are surprised by when they record themselves singing. Every note carried the same velocity — 90 — across all three files, because the engine reports the same confidence value (0.71) for each note it emits. There is no dynamic information in the output. If you recorded a phrase with a crescendo in it, the .mid will not contain the crescendo; you will draw that in your DAW afterwards.
Frequently asked questions
Can I record straight into this converter with my microphone?
No. There is no microphone path anywhere in the tool. We searched every byte of JavaScript the page loads — 984,704 bytes, of which 91,867 bytes are the site's own code — for getUserMedia, MediaRecorder, navigator.mediaDevices, enumerateDevices and createMediaStreamSource, and the count is zero. The same check on the live build's own 74,028-byte app.js is also zero. The only way audio gets in is a file: a picker, a drag-and-drop, or a paste from the clipboard.
Why is there no record button?
Because the engine takes a decoded audio buffer, not a live stream. A microphone gives you a stream that has to be recorded, encoded into a container, saved, and then decoded before analysis can start — a separate pipeline with its own permission prompt, its own length limits and its own failure modes. This tool is file in, file out, and it says so. The input element carries an accept list of 17 extensions and no capture attribute, which is exactly why tapping the drop zone opens a file browser instead of the microphone. The page has five buttons and none of them records: Download .mid, Play preview, Convert another, Re-run clean-up and Reset to defaults.
I recorded a voice memo on my phone. Will it convert?
Yes. We made a mono 48 kHz AAC recording at 128 kbps in an M4A container — the shape a phone voice-memo app writes — from a 20-second hummed line of 40 quarter notes, and converted it. It came back as 40 notes, a pitch range of D4 to A4, and a 432-byte .mid file. A browser recorder's output in Opus at 96 kbps in a WebM container produced the identical result, byte for byte. Both are in the accept list of the file input, so both go straight in with no re-encoding step.
Which file does each recorder give me, and does it convert?
An iPhone voice memo writes an M4A holding AAC audio; a browser recorder using MediaRecorder writes WebM holding Opus; Android's recorder usually writes M4A, sometimes OGG holding Opus; a desktop capture tool such as Audacity writes WAV by default. All of those are in the accept list, and the ones we tested — AAC in M4A and Opus in WebM — both converted on the first try, with no error and no re-encoding step. WAV is the safest choice if you have the option, only because it is the shortest path to the analysis: it is already uncompressed.
The two recordings converted to identical bytes. Does recording quality not matter?
It matters, but not through the file format. The AAC voice memo and the Opus recorder file produced the same 432 bytes because the signal underneath them was the same clean line — the codecs had nothing to damage. What damages a recording is what a microphone picks up: room reflections, background noise, a source that is too quiet, clipping, or several instruments at once. Those change the audio the model sees, and no container choice compensates for them. Record close, record clean, record one line, and check the piano-roll preview before you import.
Why did the last note come out slightly short?
We compared the .mid files event by event. The output from the AAC recording and the output from the WAV of the same performance differ in exactly two of their 88 events, and both are the last two in the file: the final note-off sits at tick 19,190 instead of tick 19,200, and the end-of-track marker follows it. That is 10 ticks, which at the 120 BPM written into the file is about 10.4 milliseconds. Everything before those two events — all 86 of them — is identical. The cause is the small amount of padding an AAC encoder adds, which shifts the end of the decoded audio slightly. A ten-millisecond difference at the very end of a take is not something you will hear.
How long can a recording be?
There is no length limit imposed by the page, but there is a practical one, and it is set by your device rather than by the recorder. In our tests a 20-second file took between 26.5 and 27.4 seconds to convert, so the work runs at roughly 1.3 times the length of the audio — a ten-minute recording is a thirteen-minute wait, and the tab has to stay open for all of it. For a quick idea, record 20 or 30 seconds, convert that, and only then commit to the full take.
Related: what leaves your device when you convert, how to read the piano-roll preview and spot a bad conversion, and why a recording with drums in it comes out wrong.