Open methodology

What we measure—and what we don’t

This page describes the current browser-based method, its quality checks, and the limits you should keep in mind when reading a result.

Short version

Voice Gender Test reports acoustic characteristics of one recording. It does not identify a speaker’s gender, diagnose a condition, or predict how every listener will perceive a voice.

Processing happens locally

After you actively grant microphone permission, the page samples five seconds of sound through the Web Audio API. Short analysis frames are processed in memory. The application does not use an upload endpoint, database, or browser storage for audio or results.

The four reported dimensions

Median pitchThe median estimated fundamental frequency (F0) across usable voiced frames, measured in hertz.
Pitch movementThe interquartile range of usable pitch estimates after converting them to semitones.
Spectral brightness proxyThe median spectral centroid of usable frames. It describes energy balance, not vocal-tract anatomy.
Recording qualityAn engineering check based on level, clipping, valid voiced-frame coverage, and periodicity.

Pitch estimation

The site uses a YIN-style difference function and cumulative mean normalized difference to estimate the repeating period of voiced sound. Frames that are too quiet, outside the expected speaking range, or insufficiently periodic are discarded. A median is used so a small number of incorrect frames have less influence than they would on a simple average.

The current search range is an engineering boundary of 65–500 Hz. It is not a male/female classification range. Voices and speaking styles can fall outside it.

Why there is no gender percentage

Vocal gender perception is influenced by more than fundamental frequency. Resonance, intonation, voice quality, articulation, language, nonverbal communication, context, and listener expectations can all matter. The American Speech-Language-Hearing Association recommends a multidimensional, person-centered approach and cautions against overemphasizing speaking pitch.

A browser recording cannot represent all of those factors, and this version has not been validated on a population dataset. A “78% feminine” or “92% masculine” result would therefore imply a level of calibration the product does not have.

Quality gates

The analyzer declines to show metrics when the sample is too quiet, visibly clipped, or lacks enough reliable voiced frames. These thresholds were chosen to prevent obviously poor signals from becoming precise-looking results. They are tested against generated signals, but they are not clinical cutoffs.

Sources of variation

For repeat comparisons, use the same sentence, microphone, room, distance, and comfortable speaking volume.

References and further reading

Ready for a careful snapshot?

Record the same phrase in a quiet room, then read each dimension separately.

Open the voice test