mp3→midi runs in your browser · nothing is uploaded

Record to MIDI: What Works and What Does Not

Last updated 6 October 2026

Short answer: this tool cannot record. There is no microphone button, and no code behind one — we searched all 984,704 bytes of JavaScript the page loads and found zero references to any capture API, and the file input carries no capture attribute. What does work is the two-step route: record with something else, save the file, then drop the file in. We tested that route on the two file types real recorders actually produce — an AAC voice memo and a browser Opus recording — and both came back as 40 notes and a 432-byte .mid, byte-identical to each other and identical to the studio WAV of the same performance except for a single event at the very end.

Everything below is measured on this tool, in a browser, on the engine this site uses for vocals, humming and single melodies. The numbers are reproducible and the method is described so you can check the shape of it yourself.

First, the part that does not exist

Plenty of pages that convert audio let you sing into them. This one does not, and the useful thing to do is establish that as a fact rather than an impression, because a missing feature is otherwise easy to argue about.

We loaded the page in a headless browser, waited for the engine to finish loading, and then read two things: the input element as it exists in the live DOM, and the source of every script the page pulled in.

JavaScript loaded by the page984,704 bytes total — 91,867 bytes of the site's own code (including 4,663 bytes of inline script) and 892,837 bytes of third-party code (the analytics tag and the ONNX runtime)
References to a capture API in that JavaScript0 — searched for getUserMedia, MediaRecorder, navigator.mediaDevices, enumerateDevices, createMediaStreamSource and AudioWorkletNode
Same check on the live buildIts own app.js is 74,028 bytes; capture-API matches: 0
The input element<input type="file">, accept list of 17 extensions, capture attribute present? No
Buttons on the pageDownload .mid · Play preview · Convert another · Re-run clean-up · Reset to defaults
Does the visible text mention recording?No — the words "record" and "microphone" do not appear in the page's text at all
How audio actually gets inFour ways, all of them files: the file picker, drag-and-drop anywhere on the page, pasting a file from the clipboard, and the "Convert another" button which reopens the picker

That capture attribute is the small detail worth understanding. A file input with capture set tells a phone to open the camera or the microphone directly. Without it — which is the case here — the phone opens its file browser instead. So the absence of that one attribute is the mechanical reason this page will never ask for microphone permission: it has nothing to ask with.

Why it is built that way The engine's input is a decoded audio buffer. A microphone hands you a live stream, which has to be recorded to a blob, encoded into a container, and decoded again before analysis can begin — a separate pipeline with its own permission prompt, its own length limits and its own ways to fail. This tool is file in, file out. Adding a recorder would mean adding that whole second pipeline, and a page that claims nothing leaves your device is a bad place to bolt one on casually.

The route that does work

Recording elsewhere and converting the file is not a workaround with a catch. The accept list on the input already covers what recorders write, so there is no conversion step and no re-export:

  1. Record. Any recorder will do — a phone voice-memo app, a browser recorder, or a desktop capture tool. Get as close to the source as you can and keep it to one line at a time.
  2. Save it as a file you can find. On a phone, that usually means sharing or exporting the memo out of the recorder app and onto the same device you are converting on.
  3. Drop the file into the converter. Pick it, drag it onto the page, or paste it. It is read locally, so nothing is uploaded while you wait.
  4. Check the piano-roll preview before downloading. The note count and the pitch range are printed on screen. If the range is not the range you sang, the recording is the problem, not the converter.

We ran that route end to end

The honest question is whether a recording made this way survives contact with the converter, or whether the compressed audio a recorder writes costs you something. So we took one performance — a 20-second hummed line, 40 quarter notes at 120 BPM spanning D4 to A4 — and saved it three ways: as the original uncompressed WAV, as an AAC recording at 128 kbps in an M4A container, and as an Opus recording at 96 kbps in a WebM container. The last two are what a phone voice memo and a browser recorder hand you. Each was converted once, in the same browser session, through the real file input.

FileSizeNotes returnedPitch rangeOutput .midTime
WAV, 44.1 kHz mono (baseline)1,764,044 B40D4–A4432 B27,361 ms
AAC 128 kbps in M4A, 48 kHz mono (phone voice memo)324,604 B40D4–A4432 B26,712 ms
Opus 96 kbps in WebM, 48 kHz mono (browser recorder)335,144 B40D4–A4432 B26,539 ms

Three files, one performance, 40 notes out of every one of them, the same pitch range, and the same 432-byte output. The recording is 5.4 times smaller than the WAV and the note content is unchanged: the voice memo is 18.4% of the WAV's bytes and the browser recording is 19.0%.

The timing is worth a sentence too. Twenty seconds of audio took between 26.5 and 27.4 seconds to convert on this machine, so the work runs at roughly 1.3 times the length of the recording. The compressed inputs were not measurably faster or slower than the uncompressed one — the cost tracks how long the audio is, not how big the file is.

The byte-level comparison, because "the same" is a claim

"Both produced 40 notes" is a weak statement. The stronger one is what the files actually contain, so we read all three back and compared them event by event. A .mid file here holds 88 events: 40 note-on messages, 40 note-off messages, 4 controller messages that declare the pitch-bend range, and 4 meta events — the tempo, the time signature, the track name and the end-of-track marker.

The cause is not mysterious: an AAC encoder adds a little padding at the start of the stream, so the decoded audio ends a fraction of a millisecond to a few milliseconds earlier than the original. The converter saw a marginally shorter buffer and wrote the last note-off marginally earlier. A ten-millisecond difference in the release of the final note is not audible and is not worth chasing.

There is a useful side effect of running this comparison: the WAV output in this run has the same hash as an earlier, independent run of the same WAV in a different test — which is how we know the pipeline is deterministic. Same input, same output, every time.

What each recorder hands you

The format you end up with depends on what you record with, and all of the common ones are already in the accept list, so none of them needs converting first.

RecorderWhat it writesDoes it go straight in?
iPhone or iPad voice memoM4A holding AAC audioYes — tested, 40 notes out
Browser recorder (MediaRecorder)WebM holding Opus audioYes — tested, byte-identical to the AAC result
Android recorderUsually M4A/AAC; sometimes OGG holding OpusYes — both are in the accept list
Desktop capture tool, e.g. AudacityWAV by defaultYes — uncompressed, the shortest path to analysis
Screen or system-audio captureMP4, MOV or WebMYes — the audio track is read and the video ignored

If you have a choice, WAV is the safest pick, but the reason is narrower than people expect. It is not that compressed audio damages the transcription — in our test the compressed recordings produced the same notes as the uncompressed one. It is that WAV removes one thing that could go wrong, and it is the format every recorder can produce.

Where recording, rather than converting, is what fails

None of the failures below are the converter's fault, and all of them are visible in the piano-roll preview before you download anything.

What the file it writes actually is

Whatever route the audio took to get in, the writer behaves the same way, and we confirmed this by reading the three output files back byte by byte.

FormatStandard MIDI File, format 0 — one track, every note in it
Resolution480 ticks per quarter note
Analysis sample rate44,100 Hz, mono, after decoding and down-mixing — a 48 kHz stereo recording is resampled and mixed down before the model sees it
TempoMeasured from the note onsets and written into the file; 500,000 µs per quarter note (120 BPM) only when no steady pulse can be found
Time signature4/4, a constant
Track namemp3 to midi
Pitch-bend range±2 semitones, declared with an RPN 0 message, centre 8192
Velocityclamp(amplitude × 127, 1, 127) — on all 120 notes we measured here it came out as 90, because the engine reported the same confidence value for every one of them
Note-off velocityFixed at 0x40 (64) on every note
Shortest note0.02 s — anything shorter is clamped up to it
Pitch bends writtenNone, on every file we measured here — the range is declared but no bend data is emitted
Audio contentNone. The file carries note events only — 0 bytes of samples

One number in that table is worth pulling out, because it is the one people are surprised by when they record themselves singing. Every note carried the same velocity — 90 — across all three files, because the engine reports the same confidence value (0.71) for each note it emits. There is no dynamic information in the output. If you recorded a phrase with a crescendo in it, the .mid will not contain the crescendo; you will draw that in your DAW afterwards.

Frequently asked questions

Can I record straight into this converter with my microphone?

No. There is no microphone path anywhere in the tool. We searched every byte of JavaScript the page loads — 984,704 bytes, of which 91,867 bytes are the site's own code — for getUserMedia, MediaRecorder, navigator.mediaDevices, enumerateDevices and createMediaStreamSource, and the count is zero. The same check on the live build's own 74,028-byte app.js is also zero. The only way audio gets in is a file: a picker, a drag-and-drop, or a paste from the clipboard.

Why is there no record button?

Because the engine takes a decoded audio buffer, not a live stream. A microphone gives you a stream that has to be recorded, encoded into a container, saved, and then decoded before analysis can start — a separate pipeline with its own permission prompt, its own length limits and its own failure modes. This tool is file in, file out, and it says so. The input element carries an accept list of 17 extensions and no capture attribute, which is exactly why tapping the drop zone opens a file browser instead of the microphone. The page has five buttons and none of them records: Download .mid, Play preview, Convert another, Re-run clean-up and Reset to defaults.

I recorded a voice memo on my phone. Will it convert?

Yes. We made a mono 48 kHz AAC recording at 128 kbps in an M4A container — the shape a phone voice-memo app writes — from a 20-second hummed line of 40 quarter notes, and converted it. It came back as 40 notes, a pitch range of D4 to A4, and a 432-byte .mid file. A browser recorder's output in Opus at 96 kbps in a WebM container produced the identical result, byte for byte. Both are in the accept list of the file input, so both go straight in with no re-encoding step.

Which file does each recorder give me, and does it convert?

An iPhone voice memo writes an M4A holding AAC audio; a browser recorder using MediaRecorder writes WebM holding Opus; Android's recorder usually writes M4A, sometimes OGG holding Opus; a desktop capture tool such as Audacity writes WAV by default. All of those are in the accept list, and the ones we tested — AAC in M4A and Opus in WebM — both converted on the first try, with no error and no re-encoding step. WAV is the safest choice if you have the option, only because it is the shortest path to the analysis: it is already uncompressed.

The two recordings converted to identical bytes. Does recording quality not matter?

It matters, but not through the file format. The AAC voice memo and the Opus recorder file produced the same 432 bytes because the signal underneath them was the same clean line — the codecs had nothing to damage. What damages a recording is what a microphone picks up: room reflections, background noise, a source that is too quiet, clipping, or several instruments at once. Those change the audio the model sees, and no container choice compensates for them. Record close, record clean, record one line, and check the piano-roll preview before you import.

Why did the last note come out slightly short?

We compared the .mid files event by event. The output from the AAC recording and the output from the WAV of the same performance differ in exactly two of their 88 events, and both are the last two in the file: the final note-off sits at tick 19,190 instead of tick 19,200, and the end-of-track marker follows it. That is 10 ticks, which at the 120 BPM written into the file is about 10.4 milliseconds. Everything before those two events — all 86 of them — is identical. The cause is the small amount of padding an AAC encoder adds, which shifts the end of the decoded audio slightly. A ten-millisecond difference at the very end of a take is not something you will hear.

How long can a recording be?

There is no length limit imposed by the page, but there is a practical one, and it is set by your device rather than by the recorder. In our tests a 20-second file took between 26.5 and 27.4 seconds to convert, so the work runs at roughly 1.3 times the length of the audio — a ten-minute recording is a thirteen-minute wait, and the tab has to stay open for all of it. For a quick idea, record 20 or 30 seconds, convert that, and only then commit to the full take.

Related: what leaves your device when you convert, how to read the piano-roll preview and spot a bad conversion, and why a recording with drums in it comes out wrong.