About Me个人简介
I am a doctoral student at Nakadai Lab, Institute of Science Tokyo. My research interests include speech/audio generative models, audio source separation, and spatial audio. 🤖
我是东京科学大学 Nakadai Lab 的博士生。我的研究方向包括语音/音频生成模型、音源分离与空间音频等,欢迎合作~ 🤖
News动态
- Our new preprint UNITE-AUDIO is now available. 我们的新论文 UNITE-AUDIO 现已公开。 [Demo Page]
- Two papers have been accepted by APSIPA ASC 2026. 两篇论文已被 APSIPA ASC 2026 录用。
Project项目
-
UNITE-AUDIO: Joint Learning of Continuous Tokenization and Latent Flow Matching for Text-to-Audio Generation
-
Unsupervised Single-Channel Audio Separation with Diffusion Source Priors
-
Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance
-
Real Grid RIR Dataset
Publication论文
In Submission 投稿中
- R. Shi, K. Li, Y. Wang, et al. “UNITE-AUDIO: Joint Learning of Continuous Tokenization and Latent Flow Matching for Text-to-Audio Generation”. In submission, 2026. [arXiv]
- R. Shi, Y. Wang, H. Song, et al. “Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS”. In submission, 2026. [arXiv]
First Author 第一作者
- R. Shi, C. Li, J. Wang, et al. “Unsupervised Single-Channel Audio Separation with Diffusion Source Priors”. Proceedings of the AAAI Conference on Artificial Intelligence, 2026. [arXiv]
- R. Shi, B. Yen, and K. Nakadai. “Distance Based Single-Channel Target Speech Extraction”. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025. [arXiv]
- R. Shi, Z. Lin, B. Yen, et al. “Single-Channel Target Speech Extraction Utilizing Distance and Room Clues”. European Signal Processing Conference (EUSIPCO), 2025. [arXiv]
- J. Wang*, R. Shi*, B. Yen, et al. “Single-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments”. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025. [arXiv] * equal contribution
- R. Shi, K. Li, Y. Wang, et al. “Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance”. Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2026. [arXiv]
- R. Shi, C. Li, J. Li, et al. “Exploring Efficient Waveform Diffusion Models for Foley Sound Generation”. Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2026. [arXiv]
- R. Shi, K. Itoyama, and K. Nakadai. “Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning”. Vocal Interactivity in-and-between Humans, Animals and Robots (VIHAR), 2024. [arXiv]
- R. Shi, S. Yang, Y. Chen, et al. “CNN-Transformer for Visual Tactile Fusion Applied in Road Recognition of Autonomous Vehicles”. Pattern Recognition Letters, 2023. [DOI]
- R. Shi, S. Yang, J. Lu, et al. “Road Profile Reconstruction Based on Recurrent Neural Network Embedded with Attention Mechanism”. SAE Technical Paper, 2024. [DOI]
- R. Shi, S. Yang, Y. Chen, et al. “Road Recognition for Autonomous Vehicles Based on Intelligent Tire and SE-CNN”. Intelligent Systems and Pattern Recognition, 2022. [DOI]
Contribution 合作论文
- J. Wang, R. Shi, J. Li, et al. “Manifold-Optimization-Based 3D Sound Source Mapping with Unknown Camera-Microphone Array Relative Pose”. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026. [DOI]
- J. Wang, R. Shi, Y. Kang, et al. “Fast-SDE: Efficient Single-Microphone Sound Source Distance Estimation in Reverberant Environments”. arXiv preprint, 2026. [arXiv]
Education教育背景
Institute of Science Tokyo, 东京科学大学(东京工业大学), Ph.D. in Systems and Control Engineering 系统与控制工程,博士
Beihang University, 北京航空航天大学, M.S. in Vehicle Engineering 车辆工程,硕士
Jilin University, 吉林大学, B.Eng. in Vehicle Engineering 车辆工程,本科