Publications
Selected publications can be found on the home page.
- All
- -
- Audio
- -
- Music
- -
- Language
- -
- Vision
- -
- Generation
- -
- Evaluation
TuneJury: An Open Metric for Improving Music Generation Preference Alignment
Yonghyun Kim, Junwon Lee, Haiwen Xia, Yinghao Ma, Junghyun Koo, Koichi Saito, Yuki Mitsufuji, Chris Donahue
arxiv preprint
- An instance-level music reward model (2.8M), trained on 17.5K human preference pairs in Ranknet-style.
- Role: Leveraging reward signals during model training and inference stage.
AdaTT: Text-Guided Instrument Timbre Transfer with Target-Adaptive Structural Control
Dabin Kim, Junwon Lee, Juhan Nam
Interspeech 2026
- Adaptively determine the control strength of pitch and loudness conditions for each target instrument, enabling text-guided generative timbre transfer.
- Role: Advising Dabin Kim, model architecture design, evaluation, analysis.
CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation
Gyubin Lee, Junwon Lee, Juhan Nam
CVPR 2026 Sight and Sound Workshop
- A two-phase inference-time guidance method for counterfactual video and text conditions, leveraging semantic cues from text and temporal cues from video for video-text-to-audio generation.
- Role: Advising Gyubin Lee, experiment design, evaluation, analysis.
Hear What Matters! Text-conditioned Selective Video-to-Audio Generation
Junwon Lee, Juhan Nam, Jiyoung Lee
CVPR 2026
- Generates only the sound of selected source specified via text prompt, for compositional video-to-audio generation.
- Introduces a self-supervised distillation training scheme for a text-adaptive video encoder, incorporating video mixing and cross-attention from video features to supplementary token embeddings and text features.
- Achieves SoTA performance on the proposed evaluation benchmark, outperforming baselines including one that leverages object segmentation.
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
Yoonjin Chung*, Pilsun Eu*, Junwon Lee, Keunwoo Choi, Juhan Nam, Ben Sangbae Chon (* equal contribution)
ICML 2025 Workshop on Machine Learning for Audio
- Proposes Kernel Audio Distance (KAD), a distribution-free, unbiased, and computationally cheaper alternative to FAD.
- Aligns more strongly with human perceptions of audio quality than FAD on general audio datasets.
- Role: Project lead, experiment design, evaluation, analysis, toolkit implementation.
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
Junwon Lee, Jaekwon Im, Dabin Kim, Juhan Nam
IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), 2025
- Proposes a two-stage framework, i.e., video-to-loudness and loudness-and-text-to-audio stages, for highly synchronized Foley sound generation.
- Video2RMS, a module that predicts sound loudness, is trained using discretized loudness bins with label smoothing.
- RMS2Sound, implemented as a ControlNet-based text-to-audio diffusion model, generates sound conditioned on the predicted loudness and text description.
Sound Scene Synthesis at the DCASE 2024 Challenge
Mathieu Lagrange, Junwon Lee, Modan Tailleur, Laurie M. Heller, Keunwoo Choi, Brian McFee, Keisuke Imoto, Yuki Okamoto
arxiv preprint
- Report on text-to-audio generation systems and evaluation results from the DCASE 2024 Sound Scene Synthesis Challenge.
- Role: Organizing committee member.
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
Junwon Lee*, Modan Tailleur*, Laurie M. Heller*, Keunwoo Choi*, Mathieu Lagrange*, Brian McFee, Keisuke Imoto, Yuki Okamoto
Audio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation
- Benchmarking the compositional generation capabilities of text-to-audio models, including DCASE challenge participants and other state-of-the-art models.
- Role: Task design, baseline system and evaluation implementation, result analysis.
CONMOD: Controllable Neural Frame-based Modulation Effects
Gyubin Lee, Hounsu Kim, Junwon Lee, Juhan Nam
Proceedings of the 27th International Conference on Digital Audio Effects (DAFx24)
Correlation of Fr ́echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
Modan Tailleur*, Junwon Lee*, Mathieu Lagrange, Keunwoo Choi, Laurie M. Heller, Keisuke Imoto, and Yuki Okamoto (* equal contribution)
32nd European Signal Processing Conference (EUSIPCO), 2024
- Demonstrates that the choice of domain-specific audio embeddings for FAD calculation affects the correlation between FAD scores and perceptual ratings.
- For environmental sounds, PANNs-Wavegram-LogMel achieves the strongest correlation with perceptual ratings for both audio quality and category fit.
- Role: Experiment design, toolkit implementation, analysis.
T-FOLEY: A Controllable Waveform-Domain Diffusion Model For Temporal-Event-Guided Foley Sound Synthesis
Yoonjin Chung*, Junwon Lee*, and Juhan Nam (* equal contribution)
Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024
- Introduces a diffusion model that incorporates temporal cues alongside text as conditioning for Sketch2Sound Foley generation.
- Proposes audio RMS as a temporal event feature and Block-FiLM as an effective and efficient conditioning mechanism.
- Role: Designing and implementing the conditioning method, experiment design, evaluation, analysis.
Foley Sound Synthesis In Waveform Domain With Diffusion Model
Yoonjin Chung, Junwon Lee, and Juhan Nam
DCASE 2023 Challenge Task 7 Foley Sound Synthesis Technical Report (15th, 1st model w/o phase reconstruction model)
Music Playlist Title Generation Using Artist Information
Haven Kim, Seungheon Doh, Junwon Lee, and Juhan Nam
AAAI-23 Workshop on Creative AI Across Modalities
Music Playlist Title Generation: A Machine-Translation Approach
Seungheon Doh, Junwon Lee, and Juhan Nam
2nd Workshop on Natural Language Processing for Music and Spoken Audio (NLP4MusA), 2021
- Frames music playlist title generation as a sequence-to-sequence machine translation task.
- Trains an encoder-decoder model to generate phrase-level titles from sequences of music track IDs, incorporating data augmentation techniques.
- Role: Dataset preprocessing and filtering, evaluation, analysis.