toneglass
A closer listen

FROM SOUND TO A MUSICAL NOTE

Hear the family.

A bright overtone can win the spectrum.
A family of overtones can reveal the pitch.

Toneglass looks for clear, related, persistent frequencies. It estimates a common fundamental, then names the nearest note. This is a transparent harmonic heuristic—not a complete simulation of hearing.

  1. 01Measureshort FFT windows
  2. 02Find peaksabove nearby noise
  3. 03Groupinteger multiples
  4. 04Followstable over time
  5. 05NameHz → note + cents
01

One sound, many bright lines.

Try making the eighth harmonic louder.

SYNTHETIC EXAMPLE · NO AUDIO522.5 Hz × 1, 2, 3, 4, 8
Frequency ↑ · logarithmicIllustrative time →
STRONGEST SINGLE PEAKC84,180 Hz
HARMONIC FAMILYC5522.5 Hz

These synthetic peaks run through the app’s actual harmonic-grouping function. The picture illustrates changing levels; it is not a recording. A strong 4,180 Hz peak contributes evidence for 522.5 Hz because it is 8 × that frequency.

02

Measure contrast before trusting a peak.

A narrow line matters more than a broad hill of noise.

A narrow peak rises 24 decibels above a broad local noise floor. Peaks need at least 12 decibels of prominence.24 dB abovelocal floorbroad washmedian of nearby binsfrequency →
Illustrative spectrum · digital level, not calibrated sound pressure.

A 16,384-sample FFT separates the sound into frequencies. At 48 kHz, each window spans about 341 ms, with bins about 2.93 Hz apart.

For each local maximum, the detector estimates the nearby noise floor using a median, leaving out the peak’s centre. A candidate must rise at least 12 dB above that floor.

It refines the frequency between bins by fitting the peak and its two neighbours. This gives a smoother estimate; it does not magically add frequency resolution.

03

Let related peaks vote together.

Normalize their levels, then give lower harmonics more weight.

Every admitted peak proposes possible fundamentals by division. Other peaks join a proposal when they lie near integer multiples of it. Unrelated peaks compete as separate families.

Levels become amplitudes relative to the strongest admitted peak. Each contribution is adjusted for local contrast, then divided by √harmonic order. A bright eighth harmonic helps, but gets less weight than an equally strong fundamental.

CONTRIBUTIONS FROM THE EXAMPLE ABOVErelative amplitude ÷ √order

The bars show weighted contributions, not probabilities. This example gives every peak the same strong local contrast. The final score also rewards prominence and multiple supporting harmonics.

The exact normalization and scoring rules +

Keep up to 48 prominent peaks; ignore those more than 35 dB below the strongest of them. Let M be that strongest level, L the peak level, P its prominence and h its harmonic order.

w = 10^((L − M) / 20) × clamp(P / 18, 0.4, 1) / √h

Sum w across the family. Score = M + 20 log₁₀(sum w) + 0.7 × weighted mean prominence + a gentle A-weighting term + a support bonus (1.5 per additional harmonic, capped at 6). Prominence is capped at 36 dB in the mean.

Seeds are divided by 1…10; matches may use orders 1…12. Match tolerance is the larger of 1.5 FFT bins or 1.2% of the expected harmonic frequency. These are engineering choices, not established perceptual thresholds.

If the fundamental is present, use its measured frequency. Otherwise, average each member’s frequency divided by its order, using the weights above. Normalizing helps compare relative structure as level changes, but microphone response, masking and absolute cutoffs still matter.

04

Don’t invent an octave below.

Harmonic matching needs guardrails.

MISSING FUNDAMENTAL
1f2f3f4f

Enough support to infer f.

At least three different orders must agree, with the lowest at order 3 or below. An absent fundamental is labelled Inferred.

FALSE SUBHARMONIC
2f4f6f

These also fit 1, 2, 3 × 2f.

If all orders share a divisor, reject the lower proposal. There is no independent evidence here for f instead of 2f.

A more complete parent family can also outrank its octave subset: the subset gets a 6-point penalty when a sufficiently strong parent has at least three members, including two the subset cannot explain.

05

Wait for persistence. Then name the note.

Frequency comes first; the musical label comes last.

0 msobserve70observe140observe210eligible

A candidate needs four consecutive observations and at least 180 ms of persistence. Nearby estimates are smoothed. The current leader stays selected if it is within 3 score points of the challenger, reducing flicker. The timeline shows nominal updates; FFT windows overlap.

Finally, the estimate maps to equal-tempered notes at A4 = 440 Hz. 522.5 Hz → C5, about 2.5 cents flat. The raw frequency is never snapped to C. Middle C is C4 = 261.63 Hz; C5 = 523.25 Hz.

Frequency, note number and cents +
note number = 69 + 12 × log₂(frequency / 440)

Round the note number to select the nearest semitone. Cents = 100 × (unrounded note number − rounded note number). A semitone is 100 cents.

The spectrum uses browser Blackman windowing and 0.55 magnitude smoothing. The tracker blends 45% of the previous frequency with 55% of the new estimate. Stability adds delay; it is not evidence that a brief or changing sound has no pitch.

06

What this does—and doesn’t—model.

Useful evidence, with the assumptions visible.

Mel processing?
None. Mel filterbanks pool energy into perceptually spaced bands; they do not by themselves identify a shared fundamental. Our frequency axis is logarithmic, not Mel.
Human sensitivity?
A gentle A-weighting adjustment contributes to ranking: one quarter of the A-weighting value, bounded from −35 to +2 dB before scaling. This is a rough bias, not calibrated loudness or a full model of masking and pitch salience.
Noise removal?
We reject insufficiently prominent peaks. No source separation or audio denoising is performed. Colour, contrast and fade controls only change the picture; frequency range and selection mode change detection.
What can fool it?
Unrelated sources with coincidentally aligned peaks, inharmonic machines, weak fundamentals and microphone processing. The strongest harmonic family need not be the tone a particular listener attends to. Other candidates remain available.
What was checked?
The pump recording used during development produced a steady estimate around 522.5 Hz, matching your listening report. Synthetic tests cover other pitches, missing fundamentals and noise. The fundamental was present in this clip; it did not need to be inferred. This is a development example, not independent perceptual validation.
The level meter?
It estimates RMS power in the fundamental’s narrow frequency band, in dBFS (relative to digital full scale). It excludes distant overtones, but nearby noise can contribute. Microphone gain affects the reading; it is not calibrated dB SPL. Inferred, absent fundamentals have no measured level. The thin marker holds the recent high for 1.2 seconds before falling.
Where does audio go?
Live audio stays in device memory. There is no recording or upload. The app requests microphone processing off, but the browser may not honour every request. The original pump recording is not hosted here.
How the pump recording was checked +

Offline spectral analysis followed the low line over the recording: its median was about 522.51 Hz, with a central 90% range of 520.46–523.27 Hz. Peaks near 1,045, 1,568, 2,091 and 4,180 Hz supported the same family as their relative levels changed.

Replaying 466 spectrum frames through the grouped detector gave C5 on all 463 frames after the three-frame startup. An offline autocorrelation cross-check also usually found this pitch, but sometimes chose the octave below. Autocorrelation is not part of the live detector. This recording helped develop the heuristic, so the replay is a regression check, not a held-out test.

How the fundamental level is measured +

Sum the linear spectral powers of the nearest FFT bin and three bins on either side. Double this positive-frequency power and divide by the Blackman window’s mean-square gain (0.3046), then take 10 log₁₀. This estimates band-limited RMS relative to full scale: a steady sine of peak amplitude 1 reads approximately −3.01 dBFS.

level ≈ 10 log₁₀(2 × Σ 10^(bin dB / 10) / 0.3046)

The seven-bin band is about 20.5 Hz wide at 48 kHz sampling. Existing spectral smoothing makes the meter settle gradually. It is not a true-peak or whole-input meter; its held marker tracks the band RMS estimate. Colour thresholds are display conventions, not acoustic safety limits. See the Web Audio FFT definition.

The piano highlights the nearest semitone from C3 through B5. Notes outside this range are labelled above or below the keyboard, never folded into another octave. Blue C4 and the A4 violin reference stay fixed.

Back to listening