How Many Notes at Once Before Transcription Breaks Down
Last updated 5 October 2026
Short answer: one. The engine measured here returns a single note per onset, so the second simultaneous note is already gone — and the third never appeared at all. We ran 28 real conversions covering 588 chord onsets: 3,348 notes went in and 532 came back, two notes sounded at the same instant at exactly 3 of those onsets, and three or more sounded together at none of them. Detection measured against the chord therefore falls as roughly 1 divided by the number of notes: 100.0% for a single line, 45.8% for two notes, 2.8% for a three-note triad. This is not a bug you can tune around with settings — it is what a melody-following model does with a chord.
Everything below is measured on this tool, in a browser, on the engine this site uses for vocals, humming and single melodies. The numbers are reproducible and the method is described so you can check the shape of it yourself.
The measurement, in one table
We built ten test signals. Each one plays twelve chords, two seconds apart, in a four-chord progression. The chords are stacked in close position — diatonic thirds climbing from C4, the way a keyboard player voices a chord — and only one thing changes between signals: how many notes are in the chord.
| Notes per chord | Chord | Notes returned | Recall against the chord | Of the notes returned, how many were real chord tones | Onsets it responded to |
|---|---|---|---|---|---|
| 1 | C4 | 12 | 100.0% | 12 of 12 | 12 of 12 |
| 2 | C4 E4 | 11 | 45.8% | 11 of 11 | 11 of 12 |
| 3 | C4 E4 G4 | 11 | 2.8% | 1 of 11 | 11 of 12 |
| 4 | C4 E4 G4 B4 | 10 | 4.2% | 2 of 10 | 10 of 12 |
| 5 | + D5 | 11 | 15.0% | 9 of 11 | 11 of 12 |
| 6 | + F5 | 11 | 4.2% | 3 of 11 | 11 of 12 |
| 7 | + A5 | 11 | 3.6% | 3 of 11 | 11 of 12 |
| 8 | + C6 | 8 | 1.0% | 1 of 8 | 8 of 12 |
| 10 | + D6 F6 | 10 | 2.5% | 3 of 10 | 10 of 12 |
| 12 | + G6 B6 | 7 | 1.4% | 2 of 7 | 7 of 12 |
"Recall against the chord" counts every note of the chord that the tool reproduced, divided by the number of notes played. "Onsets it responded to" counts the chord changes that produced any note at all — twelve is the maximum.
Read the "notes returned" column. Twelve chords were played at every level, and the tool returned between 7 and 12 notes in total — roughly one per chord, no matter whether the chord had one note in it or twelve. That is the whole finding in one column: the number of notes coming out tracks the number of onsets, not the number of notes going in.
Why the ceiling is one note and not a number
This is not a limit that a bigger file, a better bitrate or a different setting would move, because it is not a limit of the analysis — it is the design of the model.
The engine is GAME Small v1.0.3 (OpenVPI, models licensed CC BY-NC-SA 4.0), a singing-voice transcription model. It is the model this site offers for vocals, humming and single melodies, and its job is to follow one voice through a recording and report the notes that voice sings. A model built to track one line returns one line. Hand it a chord and it will pick a line out of the chord and follow that.
You can see the same thing in the numbers it emits. Every note the model returns carries an amplitude of 0.71 — we checked all 532 notes across the three test sets and the value was 0.71 every time, with no variation at all. The writer turns that into a velocity of 90 through clamp(amplitude × 127, 1, 127). There is no amplitude information in the output that could distinguish the loud note of a chord from the quiet ones, because the model is not separating the notes in the first place.
Some detail on how the engine runs, since it explains the shape of the result. Before analysis the audio is decoded, down-mixed to mono and resampled to 44,100 Hz. The engine loads five files totalling 51,054,395 bytes — three model files of 10,805,372, 20,906,659 and 19,335,163 bytes, plus two small configuration files — from a pinned source with a SHA-256 check on each, and runs them locally through ONNX Runtime Web on the CPU. Nothing about the audio is sent anywhere.
Two things we got wrong first, and what they tell you
Our first test used chords half a second long, and the engine returned exactly one note per chord at every level. That looked like it might be a speed limit, so we re-ran the entire experiment with two-second chords — the same chord duration used in an earlier test in which the tool did transcribe chords successfully. The result did not change.
Our second guess was voicing. The first two sets spread the notes of each chord evenly across a two-octave band, which produces wide, octave-doubled voicings; the third set used close-position thirds, the way a chord is actually played. The result did not change either.
All three sets are in the totals at the top of this page. Twenty-eight conversions, 588 chord onsets, 3,348 notes played, 532 returned, two notes at once at 3 onsets, three notes at once never. The shape held across half-second chords and two-second chords, across open voicings and close voicings, across every level from one note to twelve.
What happens above eight notes: the line stops tracking the audio
The half-second set is useful here because it ran forty chord onsets per level instead of twelve, so the failures are easier to see. Below six notes it returned exactly one note per onset, every time. Above that, the single line it follows loses the audio in two specific ways.
| Notes per chord | Notes returned per onset | Recall against the chord | Range it reported |
|---|---|---|---|
| 1 | 1.000 | 100.0% | MIDI 69 to 74 |
| 2 | 1.000 | 50.0% | MIDI 57 to 62 |
| 3 | 1.000 | 33.3% | MIDI 60 to 69 |
| 4 | 1.000 | 6.3% | MIDI 57 to 59 |
| 5 | 1.000 | 20.0% | MIDI 57 to 62 |
| 6 | 0.950 | 10.8% | MIDI 57 to 74 |
| 8 | 0.675 | 2.8% | MIDI 32 to 89 |
| 10 | 0.825 | 1.3% | MIDI 37 to 84 |
| 12 | 0.875 | 1.9% | MIDI 33 to 84 |
The first failure is dropped onsets. From six notes upward the returned count drops below one per onset — 0.950, then 0.675 — meaning some chord changes produced nothing at all. At eight notes it answered 27 of 40 onsets; at twelve, 35 of 40.
The second failure is worse. At eight simultaneous notes the engine returned MIDI 32 (G#1) when the lowest note in the audio was MIDI 57 (A3). That is 25 semitones below anything in the recording. The notes it hands you at that density are not wrong versions of what you played; they are pitches that were never there. Treat the output as unusable from about eight simultaneous notes upward.
One thing that did not change with density: the time it takes. On the 24-second files the whole range from one note to twelve notes took 31.3 to 32.5 seconds. How much is happening in the audio does not affect how long the conversion runs — the length of the audio does.
The same result on real recordings
Synthetic chords are a fair test but they are still synthetic, so it is worth asking whether the same ceiling shows up in real music. We already had the data: an earlier test put six public-domain piano recordings through this same engine and recorded how many notes were sounding at the same instant at every point in each transcription.
| Recording | Notes returned | Most notes sounding at once | Share of the piece with two notes at once | Share with three or more |
|---|---|---|---|---|
| Canon in F minor | 43 | 2 | 0.6% | 0% |
| Fugue in A minor, B. 144 | 29 | 2 | 0.4% | 0% |
| Mazurka Op. 7 no. 3 | 21 | 2 | 0.2% | 0% |
| Nocturne Op. 55 no. 1 | 15 | 2 | 0.17% | 0% |
| Waltz Op. 69 no. 2 | 24 | 2 | 0.32% | 0% |
| Cello Sonata Op. 65, III. Largo | 13 | 2 | 0.15% | 0% |
Six piano pieces, 145 notes returned in total, and three or more notes sounding at once in none of them. Two notes overlapped for between 0.15% and 0.6% of each piece. This is dense solo piano music — Chopin, where chords are the norm — and the transcription came out as a single line every time. That is the same ceiling, measured on real recordings rather than synthesised ones, and it is the strongest confirmation on this page.
As with any measurement of ours, this is what our tool did on our test material. It is not a general accuracy claim about audio-to-MIDI transcription, and it is not a promise about your file.
What to do with this
- Give a melody engine one melody. Vocals, a hummed line, a single lead instrument, a bass line — anything monophonic comes back complete, and that is the case the engine was built for. The single-note test returned 12 of 12 notes at 12 of 12 onsets.
- Check the returned note count against the number of onsets. This is the fastest diagnosis available to you. If a piece has roughly one note per chord change, you are looking at a melody line rather than a chord, and no amount of cleanup in the DAW will recover notes that were never written to the file.
- Match the engine to the source. This site offers separate engines for piano, guitar, bass and other instruments, and a separate route for a full mix. Those are different models, and nothing on this page describes them — but if what you have is a chordal instrument, a melody engine is not the right choice for it.
- Do not chase this with file settings. Bitrate, container and sample rate change how cleanly the audio is presented to the model; none of them change how many notes the model reports. We tested half-second chords and two-second chords, open voicings and close voicings, and the ceiling was identical.
- Distrust the output above eight simultaneous notes. Once the reported pitches leave the range of the audio, you are editing fiction. Cut the section, or transcribe it by hand.
What the file it writes actually is
None of the above changes the writer. Whatever the engine returns, the file is built the same way every time:
| Format | Standard MIDI File, format 0 — one track, every note in it |
| Resolution | 480 ticks per quarter note |
| Analysis sample rate | 44,100 Hz, mono, after decoding and down-mixing |
| Tempo | Measured from the note onsets and written into the file; 500,000 µs per quarter note (120 BPM) only when no steady pulse can be found |
| Time signature | 4/4, a constant |
| Track name | mp3 to midi |
| Pitch-bend range | ±2 semitones, declared with an RPN 0 message, centre 8192 |
| Velocity | clamp(amplitude × 127, 1, 127), from the model's own confidence value — which was 0.71, and so 90, on every note of all 532 we measured |
| Note-off velocity | Fixed at 0x40 (64) on every note |
| Shortest note | 0.02 s — anything shorter is clamped up to it |
| Audio content | None. The file carries note events only — 0 bytes of samples |
There is a small detail worth noticing in the file sizes. The one-note chords produced a 180-byte file; the twelve-note chords produced a 138-byte file. Adding notes to the input made the output smaller, because fewer notes came back. When the output shrinks as the input gets busier, you are watching this page's finding happen.
Frequently asked questions
How many notes can sound at once before the transcription misses them?
One. The engine measured here returns a single note per onset, so the second simultaneous note is already missing. Over 28 real conversions and 588 chord onsets we played 3,348 notes and got 532 back. Two notes sounded at the same instant at exactly 3 of those 588 onsets, and three or more never happened once. Detection measured against the chord therefore falls as roughly 1 divided by the number of notes: 100.0% for a single line, 45.8% for two notes, 2.8% for a three-note triad.
Why does a chord come back as a single note?
Because the engine is a singing-voice transcription model, not a chord model. It is the model this site offers for vocals, humming and single melodies, and its job is to follow one voice through a recording. A model built to track one line returns one line. This is not a file-size or file-format problem and it does not change with the length of the notes: we tested half-second chords and two-second chords, notes spread across two octaves and notes stacked in close position, and the result was the same at every one of the 28 levels we measured.
Which note does it keep when the input is a chord?
Usually the lowest one, but not reliably. With two notes sounding, 11 of the 11 notes it returned were real members of the chord and 8 of those 11 were the bottom note. With a three-note triad it went wrong in a different way: only 1 of the 11 notes it returned was a member of the chord at all. Once three or more notes sound together, which one survives is not something you can predict from the input.
Do longer notes or slower chords help it find more notes?
No, and we tested that directly because it was our first guess. The first run used half-second chords and the engine returned exactly one note per chord at every level. We then re-ran the whole experiment with two-second chords — the same duration used in an earlier test where the tool transcribed chords successfully — and the answer was unchanged: one note per onset, at every level. Chord speed was not the limiting factor.
Why does it start returning notes I never played?
Above about six simultaneous notes the single line it follows stops tracking the audio. Two things happen. First, chord onsets get dropped: at eight notes it responded to only 8 of 12 onsets, and at twelve notes only 7 of 12. Second, the pitches it reports leave the range you played. On the half-second test, where the lowest note in the audio was MIDI 57 (A3), the engine returned MIDI 32 (G#1) — 25 semitones below anything in the recording. At that density the pitch information is not usable.
Why do all the notes in the file have the same velocity?
Because the engine reports the same confidence value for every note it emits — an amplitude of 0.71, which the writer turns into a velocity of 90 through clamp(amplitude times 127, 1, 127). We checked this across all 532 notes the three test sets returned and the value was 0.71 every single time. There is no amplitude information in the output that could tell you which note of a chord was the loud one, which is another way of seeing that the model is not separating voices.
Can I still get the chords out of a song?
Yes, by not asking a melody engine to do it. Give the melody engine audio that contains one line at a time, or use the engine this site offers for the source type you actually have — piano, guitar, bass and other instruments are handled by separate polyphonic models, and a full mix has its own route. The measurements on this page describe the singing-voice engine only; nothing here says anything about those other models. Whichever route you take, check the note count and the piano-roll preview before you import, because a returned note count near the number of onsets is the signature of a melody line rather than a chord.
Related: how to read the piano-roll preview and spot a bad conversion, what happened when we converted public-domain recordings, and why drum tracks come out wrong.