← All work

Selective Hearing / AI & apps

Audio separation study · playable example

Choose what stays in the mix.

A model trained from scratch estimates sirens, dog barks, or door knocks within a recording. Hear one prepared mixture, then isolate, prioritize, or attenuate the sound you choose.

My role
Study design / model training / audio demo
When
October 2026
Built with
Python · PyTorch · Conditional U-Net · FSD50K · React

Hear the difference

Try it yourself ↓

Recorded example · model output

A synthetic alarm, a dog bark, a door knock, and running water.

Choose a sound
Compare versions
Original mix

The full recording, with all four sounds.

Switch versions at the same point in the recording.

Precomputed model outputs, not live inference. All versions share one playback gain.

Sound credits

Outcome

Across 300 held-out mixtures and 900 class queries, the U-Net improved gain-sensitive SDR by 7.06 dB on the 525 target-present queries, averaged across three seeds, versus 4.95 dB for a spectral-template baseline. SI-SDR improved by 3.42 dB. Seed SDR improvements ranged from 6.83 to 7.43 dB; absolute target SDR was 3.62 dB, so errors remain. The other 375 queries measure leakage when the requested source is absent.

Developed with AI assistance. The experiment recovers labeled recordings in synthetic mixtures; FSD50K sources can already contain other sounds. It does not establish clean separation in everyday recordings or real-time microphone performance. The demo is one two-second validation example with precomputed outputs from seed 0. Its siren-class source is a synthetic aircraft warning. The study-wide scores describe the held-out comparison, not this preset.

Held-out target SDR improvement by sound class: the three-seed U-Net average exceeds the spectral-template baseline, with a separate reference that uses known source tracks
Original final-test figure (Spanish) · gain-sensitive SDR improvement on target-present queries. Dots show three trained seeds; the known-source reference is unavailable in deployment. These scores do not describe the preset above. · Open full size ↗

The challenge

Turning down a recording makes everything quieter. I wanted to estimate one sound class separately so its level could change while retaining the rest of the scene, and to measure how much a small network actually recovers.

Approach

Keep recordings and authors apart

I selected 1,412 FSD50K recordings: 868 for training, 184 for validation, and 360 reserved for test. No uploader, source ID, or exact original/PCM duplicate crosses those partitions. Two-second scenes mix background audio with zero to three target classes, including queries for sounds that are absent.

Train three seeds before testing

A 318,361-parameter U-Net uses a class embedding to estimate a spectrogram mask. Two learning-rate pilots precede three 4,000-step training runs. Validation selects checkpoints; the demo uses the run with median validation loss, chosen before final test scoring. The three main training and validation loops took 26.5 minutes locally.

Separate once, then change the balance

The remainder is the input minus the estimated target. Isolate plays only the estimate; prioritize keeps it while lowering the remainder to 25%; attenuate plays only the remainder. The preset uses saved model outputs with a common playback scale, making it easy to compare modes without running a model in the browser.

Sources & project context
Next projectLoRA intent adaptation ↗