GestureSynth.me

Inside the browser instrument

How Hand Tracking Becomes Music

Follow one camera frame through local hand-landmark inference, gesture mapping, chord construction and Web Audio synthesis. This technical guide explains the system that powers the playable instrument above.

hand tracking synthesizer

Inside the Hand Tracking Synthesizer

GestureSynth.me uses four local stages: a video-only camera stream, MediaPipe hand landmarks, deterministic musical mapping and Web Audio sound. Each finger count has a defined meaning, so an output note can be traced to a key, scale degree and chord rule.

The page above this guide is the working instrument, so the explanation describes production behavior rather than a concept demo. Gesture mode uses stable poses for chords. Theremin mode uses continuous hand height for pitch and volume. Both modes share the same camera and tracker, while their mapping and synthesis paths differ after landmarks are available.

Play the instrument ↑
Landmarks
Up to two hands, in browser
Inference rate
24 FPS or 15 FPS low-power
Sound engine
Web Audio oscillators and filters
Camera frames
Not uploaded by the app

1. Camera Access Starts with a Deliberate Action

The application does not ask for a camera while you read the page. Pressing Start calls the browser media API with video enabled and audio explicitly disabled. The preferred capture size is 640 by 480 pixels with the user-facing camera. The browser remains responsible for displaying its permission prompt, remembering your choice and showing its camera-use indicator.

Once permission is granted, the stream is attached to a muted inline video element. Stop session cancels the frame loop, stops every media track, closes the detector and removes the stream. Closing the page triggers the same cleanup.

2. MediaPipe Estimates Hand Landmarks Locally

After the camera is ready, the app loads a MediaPipe Hand Landmarker model and its WebAssembly runtime from this website. It first asks the browser for GPU inference and falls back to CPU if that setup fails. The detector runs in video mode, looks for at most two hands and returns landmark coordinates plus a left-or-right handedness label for each result.

Inference uses requestAnimationFrame but is capped at about 24 detections per second, or 15 in low-performance mode. It pauses while the page is hidden. The interface records detection time locally, and each new result replaces the previous in-memory result.

3. Landmark Geometry Becomes a Small Set of Controls

The musical layer compares landmark positions with ordinary geometry. A non-thumb finger counts as raised when its tip sits above its middle joint. Thumb direction follows handedness. Wrist position produces normalized tilt, and its vertical coordinate provides hand height.

A short stabilizer prevents a single noisy frame from changing a chord. A new discrete value must remain consistent for roughly 100 milliseconds before it is committed, while a brief tracking dropout receives about 50 milliseconds of grace. Continuous controls use audio smoothing as well. These choices trade a small amount of latency for fewer accidental changes.

4. Two Hands Divide Harmony and Expression

In Gesture mode, the left hand chooses a scale degree from one through seven. One to five raised digits map directly to degrees one to five; two special index-and-pinky shapes select six or seven depending on the left thumb. Left-wrist tilt switches between major and minor chord worlds when automatic mode selection is active.

The right hand supplies the voicing and expression. One raised finger selects a triad, two select first inversion, three select a major or minor seventh, and four select a dominant or diminished seventh. Right-hand height controls volume, wrist tilt opens or closes the low-pass filter, and an extended right thumb lowers the voicing by one octave. The split lets one hand answer which harmony while the other answers how it should sound.

5. Chords Are Calculated from Music Rules

The chord builder begins with the selected key and converts the scale degree to a MIDI root. Major and minor modes choose a four- or three-semitone third, while the fifth sits seven semitones above the root. The selected voicing then determines whether the output contains a doubled triad, first inversion, seventh or diminished shape. Octave-down subtracts twelve semitones from every note.

Each MIDI number becomes a frequency using equal temperament around A4 at 440 hertz. The same pose changes after the key changes because it describes a musical role, not fixed frequencies. Labels, note names and the sound engine share that result.

6. Web Audio Produces and Shapes the Sound

The audio engine creates an AudioContext after a user action. Gesture chords use four sawtooth oscillators whose gains and frequencies glide toward new targets. They pass through a resonant low-pass filter, master gain and dynamics limiter to control peaks.

Theremin mode follows a separate path. Right-hand height is mapped across roughly 65 to 1,200 hertz and drives a sine oscillator; left-hand height controls amplitude. When the user changes modes, the inactive path fades to zero. Muting, master volume and the panic action also update gain nodes directly, so silence does not depend on stopping the camera tracker.

7. The Sequencer Stores Musical Events, Then Renders Locally

The loop sequencer captures musical frames such as frequencies, volume, filter position, mode and theremin pitch. It does not store webcam pixels. Four tracks can hold those events at tempo-based steps, with mute, solo and track-level volume. Playback schedules temporary oscillator envelopes so a saved step sounds consistent even when the hands have moved elsewhere.

Export uses an OfflineAudioContext for a stereo 44.1 kHz WAV. If unavailable, MediaRecorder can create WebM audio. Both paths use synthesized Web Audio rather than the microphone. A temporary local URL downloads the Blob and is then released.

8. Lighting, Occlusion and Audio Hardware Set Practical Limits

Reliable tracking needs the hands to be visually distinct. Soft front lighting, visible wrists and space between the hands help the model maintain landmarks. Backlighting, motion blur, patterned backgrounds and one hand covering the other can cause temporary loss or an incorrect handedness label. Moving deliberately gives the stabilizer time to commit the intended pose.

Perceived latency combines camera exposure, inference, browser scheduling and audio output. Closing graphics-heavy tabs and choosing low-performance mode can help an overloaded device. Bluetooth headphones may add delay even when detection is fast, so wired output is a better timing reference. The instrument is designed for playful browser performance; it is not calibrated stage hardware, a biometric system or a clinical motion tool.

9. The Privacy Boundary Follows the Architecture

The application does not intentionally upload camera frames, hand landmarks, loops or exported audio. It asks for no microphone and exposes no raw-landmark API. Hosting still receives ordinary requests for pages, scripts, the model and other assets.

To inspect the boundary, open the Network panel, press Start and move both hands. Model assets may load, but video frames and landmarks should not appear as outbound requests. Press Stop and confirm the camera indicator turns off.

Answers before you play

How the Hand Tracking Synthesizer Works

Does hand tracking run on a server?

The application loads its MediaPipe model and WebAssembly runtime from GestureSynth.me, then performs video inference in the browser. The application does not intentionally send camera frames or landmark results to a tracking server.

Why does a pose need to be held briefly?

A discrete value must remain stable for about 100 milliseconds before it is committed. That debounce rejects one-frame mistakes and gives a short dropout grace period when tracking flickers.

What causes delay between movement and sound?

Camera frame rate, model inference, the 100-millisecond pose stabilizer, browser load and the selected audio output all contribute. Bluetooth output can add delay after the synthesizer has already generated the sound.

Does the tracker identify faces or people?

No identity feature is part of the musical pipeline. The interface uses hand landmark geometry to derive musical controls and does not need a face profile or biometric identity.

Can I export the raw tracking data?

The public instrument exports rendered loop audio, not camera frames or raw landmark arrays. It does not expose a general-purpose tracking API.