Junwon Lee 이준원
drawing by Joonhyung.

Junwon Lee 이준원

I am a Ph.D. student at KAIST Music and Audio Computing Lab  advised by Prof. Juhan Nam.
My research focuses on audio-centric multimodal understanding and generative AI technologies for content creators.
Recently, I have been particularly interested in controllable and interactive audio generation that leverages the reasoning, generative, and agentic capabilities of multimodal LLMs.
Previously, I was a research intern at NAVER AI Lab and Gaudio Lab.

News

  • Feb, 2026 | SelVA accepted to CVPR 2026!Link
  • Mar, 2025 | KAD Toolkit (kadtk) released!Link
  • Mar, 2025 | Starting my PhD @ KAIST GSAI, MAC Lab!Link
  • | If you have any questions about joining our lab, please contact me via email :)

Selected Publications

  • All
  • -
  • Audio
  • -
  • Music
  • -
  • Language
  • -
  • Vision
  • -
  • Generation
  • -
  • Evaluation
Hear What Matters! Text-conditioned Selective Video-to-Audio Generation

Junwon Lee, Juhan Nam, Jiyoung Lee

CVPR 2026

#Generation #Audio #Vision 

  • Generates only the sound of selected source specified via text prompt, for compositional video-to-audio generation.
  • Introduces a self-supervised distillation training scheme for a text-adaptive video encoder, incorporating video mixing and cross-attention from video features to supplementary token embeddings and text features.
  • Achieves SoTA performance on the proposed evaluation benchmark, outperforming baselines including one that leverages object segmentation.
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation

Yoonjin Chung*, Pilsun Eu*, Junwon Lee, Keunwoo Choi, Juhan Nam, Ben Sangbae Chon (* equal contribution)

ICML 2025 Workshop on Machine Learning for Audio

#Generation #Audio #Evaluation 

  • Proposes Kernel Audio Distance (KAD), a distribution-free, unbiased, and computationally cheaper alternative to FAD.
  • Aligns more strongly with human perceptions of audio quality than FAD on general audio datasets.
  • Role: Project lead, experiment design, evaluation, analysis, toolkit implementation.
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound

Junwon Lee, Jaekwon Im, Dabin Kim, Juhan Nam

IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), 2025

#Generation #Audio #Vision 

  • Proposes a two-stage framework, i.e., video-to-loudness and loudness-and-text-to-audio stages, for highly synchronized Foley sound generation.
  • Video2RMS, a module that predicts sound loudness, is trained using discretized loudness bins with label smoothing.
  • RMS2Sound, implemented as a ControlNet-based text-to-audio diffusion model, generates sound conditioned on the predicted loudness and text description.
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation

Junwon Lee*, Modan Tailleur*, Laurie M. Heller*, Keunwoo Choi*, Mathieu Lagrange*, Brian McFee, Keisuke Imoto, Yuki Okamoto

Audio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation

#Generation #Audio #Evaluation 

  • Benchmarking the compositional generation capabilities of text-to-audio models, including DCASE challenge participants and other state-of-the-art models.
  • Role: Task design, baseline system and evaluation implementation, result analysis.
Correlation of Fr ́echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant

Modan Tailleur*, Junwon Lee*, Mathieu Lagrange, Keunwoo Choi, Laurie M. Heller, Keisuke Imoto, and Yuki Okamoto (* equal contribution)

32nd European Signal Processing Conference (EUSIPCO), 2024

#Generation #Audio #Evaluation 

  • Demonstrates that the choice of domain-specific audio embeddings for FAD calculation affects the correlation between FAD scores and perceptual ratings.
  • For environmental sounds, PANNs-Wavegram-LogMel achieves the strongest correlation with perceptual ratings for both audio quality and category fit.
  • Role: Experiment design, toolkit implementation, analysis.
T-FOLEY: A Controllable Waveform-Domain Diffusion Model For Temporal-Event-Guided Foley Sound Synthesis

Yoonjin Chung*, Junwon Lee*, and Juhan Nam (* equal contribution)

Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024

#Generation #Audio 

  • Introduces a diffusion model that incorporates temporal cues alongside text as conditioning for Sketch2Sound Foley generation.
  • Proposes audio RMS as a temporal event feature and Block-FiLM as an effective and efficient conditioning mechanism.
  • Role: Designing and implementing the conditioning method, experiment design, evaluation, analysis.