Haopeng Chen
PhD Student in Computer Science · University of Mississippi
Publications
Leveraging Diverse Semantic-Based Audio Pretrained Models for Singing Voice Conversion
Xueyao Zhang, Zihao Fang, Yicheng Gu, Haopeng Chen, Lexiao Zou, Junan Zhang, Liumeng Xue, Zhizheng Wu
2024 IEEE Spoken Language Technology Workshop (SLT), 2024
Abstract
Singing Voice Conversion (SVC) is a technique that enables any singer to perform any song. To achieve this, it is essential to obtain speaker-agnostic representations from the source audio, which poses a significant challenge. A common solution involves utilizing a semantic-based audio pretrained model as a feature extractor However, the degree to which the extracted features can meet the SVC requirements remains an open question. This includes their capability to accurately model melody and lyrics, the speaker-independency of their underlying acoustic information, and their robustness for in-the-wild acoustic environments. In this study, we investigate the knowledge within classical semantic-based pretrained models in much detail. We discover that the knowledge of different models is diverse and can be complementary for SVC. Based on the above, we design a Singing Voice Conversion framework based on Diverse Semantic-based Feature Fusion (DSFF-SVC). Experimental results demonstrate that DSFF-SVC can be generalized and improve various existing SVC models, particularly in challenging real-world conversion tasks. Our demo website is available at https://diversesemanticsvc.github.io/.
Amphion: an Open-Source Audio, Music, and Speech Generation Toolkit
Xueyao Zhang, Liumeng Xue, Yicheng Gu, …, Haopeng Chen, …, Zhizheng Wu
2024 IEEE Spoken Language Technology Workshop (SLT), 2024
Abstract
Amphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields. It presents a unified framework that includes diverse generation tasks and models, with the added bonus of being easily extendable for new incorporation. The toolkit is designed with beginner-friendly workflows and pre-trained models, allowing both beginners and seasoned researchers to kick-start their projects with relative ease. The initial release of Amphion v0.1 supports a range of tasks including Text to Speech (TTS), Text to Audio (TTA), and Singing Voice Conversion (SVC), supplemented by essential components like data preprocessing, state-of-the-art vocoders, and evaluation metrics. This paper presents a high-level overview of Amphion. Amphion is open-sourced at https://github.com/open-mmlab/Amphion.
Experience
- Zhipu AI (aka. Z.ai)May 2026 – Aug 2026LLM Research InternHangzhou, China
Education
- The University of MississippiAug 2024 – Present
PhD in Computer Science
Advisor: Bo Wang
- The Chinese University of Hong Kong, ShenzhenSep 2020 – Jul 2024
B.Eng. in Computer Science and Engineering
- Dean's List, Academic Year 2022–2023
- Bowen Scholarship (Merit-based), 2020–2024
- Undergraduate Research Award, 2023
- University of California, BerkeleyAug 2022 – Dec 2022
Visiting Student in Computer Science
Teaching
- Engr 691: Deep Learning — Foundations and FrontiersSpring 2025
Teaching Assistant · University of Mississippi
- CSci 543: Data MiningFall 2024
Teaching Assistant · University of Mississippi
Talks
- Can AI See in the Dark? — UDAPoseApr 2026
Research and Scholarly Activity Showcase, University of Mississippi
- Advances and Challenges in Human Pose Estimation in the Era of Foundation ModelsOct 2025
Guest Lecture, Computer Vision, University of Mississippi
- Human Pose and Motion Analysis under Low VisibilityNov 2024
AI Task Force Meeting, University of Mississippi
Service
- Reviewer — NeurIPS (The Fortieth Annual Conference on Neural Information Processing Systems)2026
- Reviewer — WACV (IEEE/CVF Winter Conference on Applications of Computer Vision)2026