Publications

Selected publications can be found on the home page.

  • All
  • -
  • Audio
  • -
  • Music
  • -
  • Language
  • -
  • Vision
  • -
  • Generation
  • -
  • Evaluation
TuneJury: An Open Metric for Improving Music Generation Preference Alignment

Yonghyun Kim, Junwon Lee, Haiwen Xia, Yinghao Ma, Junghyun Koo, Koichi Saito, Yuki Mitsufuji, Chris Donahue

arxiv preprint

#Generation #Music #Evaluation 

  • An instance-level music reward model (2.8M), trained on 17.5K human preference pairs in Ranknet-style.
  • Role: Leveraging reward signals during model training and inference stage.
AdaTT: Text-Guided Instrument Timbre Transfer with Target-Adaptive Structural Control

Dabin Kim, Junwon Lee, Juhan Nam

Interspeech 2026

#Generation #Music 

  • Adaptively determine the control strength of pitch and loudness conditions for each target instrument, enabling text-guided generative timbre transfer.
  • Role: Advising Dabin Kim, model architecture design, evaluation, analysis.
CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation

Gyubin Lee, Junwon Lee, Juhan Nam

CVPR 2026 Sight and Sound Workshop

#Generation #Audio #Vision 

  • A two-phase inference-time guidance method for counterfactual video and text conditions, leveraging semantic cues from text and temporal cues from video for video-text-to-audio generation.
  • Role: Advising Gyubin Lee, experiment design, evaluation, analysis.
Hear What Matters! Text-conditioned Selective Video-to-Audio Generation

Junwon Lee, Juhan Nam, Jiyoung Lee

CVPR 2026

#Generation #Audio #Vision 

  • Generates only the sound of selected source specified via text prompt, for compositional video-to-audio generation.
  • Introduces a self-supervised distillation training scheme for a text-adaptive video encoder, incorporating video mixing and cross-attention from video features to supplementary token embeddings and text features.
  • Achieves SoTA performance on the proposed evaluation benchmark, outperforming baselines including one that leverages object segmentation.
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation

Yoonjin Chung*, Pilsun Eu*, Junwon Lee, Keunwoo Choi, Juhan Nam, Ben Sangbae Chon (* equal contribution)

ICML 2025 Workshop on Machine Learning for Audio

#Generation #Audio #Evaluation 

  • Proposes Kernel Audio Distance (KAD), a distribution-free, unbiased, and computationally cheaper alternative to FAD.
  • Aligns more strongly with human perceptions of audio quality than FAD on general audio datasets.
  • Role: Project lead, experiment design, evaluation, analysis, toolkit implementation.
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound

Junwon Lee, Jaekwon Im, Dabin Kim, Juhan Nam

IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), 2025

#Generation #Audio #Vision 

  • Proposes a two-stage framework, i.e., video-to-loudness and loudness-and-text-to-audio stages, for highly synchronized Foley sound generation.
  • Video2RMS, a module that predicts sound loudness, is trained using discretized loudness bins with label smoothing.
  • RMS2Sound, implemented as a ControlNet-based text-to-audio diffusion model, generates sound conditioned on the predicted loudness and text description.
Sound Scene Synthesis at the DCASE 2024 Challenge

Mathieu Lagrange, Junwon Lee, Modan Tailleur, Laurie M. Heller, Keunwoo Choi, Brian McFee, Keisuke Imoto, Yuki Okamoto

arxiv preprint

#Generation #Audio #Evaluation 

  • Report on text-to-audio generation systems and evaluation results from the DCASE 2024 Sound Scene Synthesis Challenge.
  • Role: Organizing committee member.
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation

Junwon Lee*, Modan Tailleur*, Laurie M. Heller*, Keunwoo Choi*, Mathieu Lagrange*, Brian McFee, Keisuke Imoto, Yuki Okamoto

Audio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation

#Generation #Audio #Evaluation 

  • Benchmarking the compositional generation capabilities of text-to-audio models, including DCASE challenge participants and other state-of-the-art models.
  • Role: Task design, baseline system and evaluation implementation, result analysis.
CONMOD: Controllable Neural Frame-based Modulation Effects

Gyubin Lee, Hounsu Kim, Junwon Lee, Juhan Nam

Proceedings of the 27th International Conference on Digital Audio Effects (DAFx24)

#Audio 

Correlation of Fr ́echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant

Modan Tailleur*, Junwon Lee*, Mathieu Lagrange, Keunwoo Choi, Laurie M. Heller, Keisuke Imoto, and Yuki Okamoto (* equal contribution)

32nd European Signal Processing Conference (EUSIPCO), 2024

#Generation #Audio #Evaluation 

  • Demonstrates that the choice of domain-specific audio embeddings for FAD calculation affects the correlation between FAD scores and perceptual ratings.
  • For environmental sounds, PANNs-Wavegram-LogMel achieves the strongest correlation with perceptual ratings for both audio quality and category fit.
  • Role: Experiment design, toolkit implementation, analysis.
T-FOLEY: A Controllable Waveform-Domain Diffusion Model For Temporal-Event-Guided Foley Sound Synthesis

Yoonjin Chung*, Junwon Lee*, and Juhan Nam (* equal contribution)

Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024

#Generation #Audio 

  • Introduces a diffusion model that incorporates temporal cues alongside text as conditioning for Sketch2Sound Foley generation.
  • Proposes audio RMS as a temporal event feature and Block-FiLM as an effective and efficient conditioning mechanism.
  • Role: Designing and implementing the conditioning method, experiment design, evaluation, analysis.
Foley Sound Synthesis In Waveform Domain With Diffusion Model

Yoonjin Chung, Junwon Lee, and Juhan Nam

DCASE 2023 Challenge Task 7 Foley Sound Synthesis Technical Report (15th, 1st model w/o phase reconstruction model)

#Generation #Audio 

Music Playlist Title Generation Using Artist Information

Haven Kim, Seungheon Doh, Junwon Lee, and Juhan Nam

AAAI-23 Workshop on Creative AI Across Modalities

#Generation #Language #Music 

Music Playlist Title Generation: A Machine-Translation Approach

Seungheon Doh, Junwon Lee, and Juhan Nam

2nd Workshop on Natural Language Processing for Music and Spoken Audio (NLP4MusA), 2021

#Generation #Language #Music 

  • Frames music playlist title generation as a sequence-to-sequence machine translation task.
  • Trains an encoder-decoder model to generate phrase-level titles from sequences of music track IDs, incorporating data augmentation techniques.
  • Role: Dataset preprocessing and filtering, evaluation, analysis.