Selective Hearing / AI & apps
Audio separation study · playable example
Choose what stays in the mix.
A model trained from scratch estimates sirens, dog barks, or door knocks within a recording. Hear one prepared mixture, then isolate, prioritize, or attenuate the sound you choose.
- My role
- Study design / model training / audio demo
- When
- October 2026
- Built with
- Python · PyTorch · Conditional U-Net · FSD50K · React
Hear the difference
Try it yourself ↓Recorded example · model output
A synthetic alarm, a dog bark, a door knock, and running water.
The full recording, with all four sounds.
Switch versions at the same point in the recording.
Outcome
Across 300 held-out mixtures and 900 class queries, the U-Net improved gain-sensitive SDR by 7.06 dB on the 525 target-present queries, averaged across three seeds, versus 4.95 dB for a spectral-template baseline. SI-SDR improved by 3.42 dB. Seed SDR improvements ranged from 6.83 to 7.43 dB; absolute target SDR was 3.62 dB, so errors remain. The other 375 queries measure leakage when the requested source is absent.
Developed with AI assistance. The experiment recovers labeled recordings in synthetic mixtures; FSD50K sources can already contain other sounds. It does not establish clean separation in everyday recordings or real-time microphone performance. The demo is one two-second validation example with precomputed outputs from seed 0. Its siren-class source is a synthetic aircraft warning. The study-wide scores describe the held-out comparison, not this preset.

The challenge
Turning down a recording makes everything quieter. I wanted to estimate one sound class separately so its level could change while retaining the rest of the scene, and to measure how much a small network actually recovers.
Approach
Keep recordings and authors apart
I selected 1,412 FSD50K recordings: 868 for training, 184 for validation, and 360 reserved for test. No uploader, source ID, or exact original/PCM duplicate crosses those partitions. Two-second scenes mix background audio with zero to three target classes, including queries for sounds that are absent.
Train three seeds before testing
A 318,361-parameter U-Net uses a class embedding to estimate a spectrogram mask. Two learning-rate pilots precede three 4,000-step training runs. Validation selects checkpoints; the demo uses the run with median validation loss, chosen before final test scoring. The three main training and validation loops took 26.5 minutes locally.
Separate once, then change the balance
The remainder is the input minus the estimated target. Isolate plays only the estimate; prioritize keeps it while lowering the remainder to 25%; attenuate plays only the remainder. The preset uses saved model outputs with a common playback scale, making it easy to compare modes without running a model in the browser.