
Junwon Lee 이준원
I am a Ph.D. student at KAIST Music and Audio Computing Lab advised by Prof. Juhan Nam.
My research focuses on audio-centric multimodal understanding and generative AI technologies for content creators.
Recently, I have been particularly interested in controllable and interactive audio generation that leverages the reasoning, generative, and agentic capabilities of multimodal LLMs.
Previously, I was a research intern at NAVER AI Lab and Gaudio Lab.
Selected Publications
- All
- -
- Audio
- -
- Music
- -
- Language
- -
- Vision
- -
- Generation
- -
- Evaluation
Hear What Matters! Text-conditioned Selective Video-to-Audio Generation
Junwon Lee, Juhan Nam, Jiyoung Lee
CVPR 2026
- Generates only the sound of selected source specified via text prompt, for compositional video-to-audio generation.
- Introduces a self-supervised distillation training scheme for a text-adaptive video encoder, incorporating video mixing and cross-attention from video features to supplementary token embeddings and text features.
- Achieves SoTA performance on the proposed evaluation benchmark, outperforming baselines including one that leverages object segmentation.
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
Yoonjin Chung*, Pilsun Eu*, Junwon Lee, Keunwoo Choi, Juhan Nam, Ben Sangbae Chon (* equal contribution)
ICML 2025 Workshop on Machine Learning for Audio
- Proposes Kernel Audio Distance (KAD), a distribution-free, unbiased, and computationally cheaper alternative to FAD.
- Aligns more strongly with human perceptions of audio quality than FAD on general audio datasets.
- Role: Project lead, experiment design, evaluation, analysis, toolkit implementation.
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
Junwon Lee, Jaekwon Im, Dabin Kim, Juhan Nam
IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), 2025
- Proposes a two-stage framework, i.e., video-to-loudness and loudness-and-text-to-audio stages, for highly synchronized Foley sound generation.
- Video2RMS, a module that predicts sound loudness, is trained using discretized loudness bins with label smoothing.
- RMS2Sound, implemented as a ControlNet-based text-to-audio diffusion model, generates sound conditioned on the predicted loudness and text description.
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
Junwon Lee*, Modan Tailleur*, Laurie M. Heller*, Keunwoo Choi*, Mathieu Lagrange*, Brian McFee, Keisuke Imoto, Yuki Okamoto
Audio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation
- Benchmarking the compositional generation capabilities of text-to-audio models, including DCASE challenge participants and other state-of-the-art models.
- Role: Task design, baseline system and evaluation implementation, result analysis.
Correlation of Fr ́echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
Modan Tailleur*, Junwon Lee*, Mathieu Lagrange, Keunwoo Choi, Laurie M. Heller, Keisuke Imoto, and Yuki Okamoto (* equal contribution)
32nd European Signal Processing Conference (EUSIPCO), 2024
- Demonstrates that the choice of domain-specific audio embeddings for FAD calculation affects the correlation between FAD scores and perceptual ratings.
- For environmental sounds, PANNs-Wavegram-LogMel achieves the strongest correlation with perceptual ratings for both audio quality and category fit.
- Role: Experiment design, toolkit implementation, analysis.
T-FOLEY: A Controllable Waveform-Domain Diffusion Model For Temporal-Event-Guided Foley Sound Synthesis
Yoonjin Chung*, Junwon Lee*, and Juhan Nam (* equal contribution)
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024
- Introduces a diffusion model that incorporates temporal cues alongside text as conditioning for Sketch2Sound Foley generation.
- Proposes audio RMS as a temporal event feature and Block-FiLM as an effective and efficient conditioning mechanism.
- Role: Designing and implementing the conditioning method, experiment design, evaluation, analysis.