About Me个人简介

I am a doctoral student at Nakadai Lab, Institute of Science Tokyo. My research interests include speech/audio generative models, audio source separation, and spatial audio. 🤖

我是东京科学大学 Nakadai Lab 的博士生。我的研究方向包括语音/音频生成模型、音源分离与空间音频等,欢迎合作~ 🤖

News动态

  • Our new preprint UNITE-AUDIO is now available. 我们的新论文 UNITE-AUDIO 现已公开。 [Demo Page]
  • Two papers have been accepted by APSIPA ASC 2026. 两篇论文已被 APSIPA ASC 2026 录用。

Project项目

  • UNITE-AUDIO: Joint Learning of Continuous Tokenization and Latent Flow Matching for Text-to-Audio Generation

    Overview of the UNITE-AUDIO training framework

    Joint learning of continuous audio representations and latent flow matching for text-to-audio generation, with Flow-GRPO post-training.

    联合学习连续音频表征与潜在 Flow Matching 的文本到音频生成方法,并采用 Flow-GRPO 进行后训练。

  • Unsupervised Single-Channel Audio Separation with Diffusion Source Priors

    Diffusion source prior audio separation framework

    Open-source implementation for the AAAI 2026 paper on unsupervised single-channel audio separation across speech-sound, sound-sound, and speech-speech mixtures.

    AAAI 2026 论文的开源实现,面向 speech-sound、sound-sound 与 speech-speech 等单通道音频分离任务。

  • Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance

    Speaker-embedding guided speech separation framework

    Code for unsupervised single-channel speech separation with diffusion source models and speaker-embedding guidance.

    基于扩散源模型和说话人嵌入引导的无监督单通道语音分离代码。

  • Real Grid RIR Dataset

    Microphone and loudspeaker setup for the Grid RIR dataset

    A real room impulse response dataset collected at grid source locations in a meeting room, useful for distance estimation, spatial audio separation, and target extraction.

    会议室网格位置采集的真实房间脉冲响应数据集,适用于距离估计、空间音频分离和目标提取等研究。

Publication论文

In Submission 投稿中

  1. R. Shi, K. Li, Y. Wang, et al. “UNITE-AUDIO: Joint Learning of Continuous Tokenization and Latent Flow Matching for Text-to-Audio Generation”. In submission, 2026. [arXiv]
  2. R. Shi, Y. Wang, H. Song, et al. “Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS”. In submission, 2026. [arXiv]

First Author 第一作者

  1. R. Shi, C. Li, J. Wang, et al. “Unsupervised Single-Channel Audio Separation with Diffusion Source Priors”. Proceedings of the AAAI Conference on Artificial Intelligence, 2026. [arXiv]
  2. R. Shi, B. Yen, and K. Nakadai. “Distance Based Single-Channel Target Speech Extraction”. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025. [arXiv]
  3. R. Shi, Z. Lin, B. Yen, et al. “Single-Channel Target Speech Extraction Utilizing Distance and Room Clues”. European Signal Processing Conference (EUSIPCO), 2025. [arXiv]
  4. J. Wang*, R. Shi*, B. Yen, et al. “Single-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments”. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025. [arXiv] * equal contribution
  5. R. Shi, K. Li, Y. Wang, et al. “Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance”. Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2026. [arXiv]
  6. R. Shi, C. Li, J. Li, et al. “Exploring Efficient Waveform Diffusion Models for Foley Sound Generation”. Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2026. [arXiv]
  7. R. Shi, K. Itoyama, and K. Nakadai. “Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning”. Vocal Interactivity in-and-between Humans, Animals and Robots (VIHAR), 2024. [arXiv]
  8. R. Shi, S. Yang, Y. Chen, et al. “CNN-Transformer for Visual Tactile Fusion Applied in Road Recognition of Autonomous Vehicles”. Pattern Recognition Letters, 2023. [DOI]
  9. R. Shi, S. Yang, J. Lu, et al. “Road Profile Reconstruction Based on Recurrent Neural Network Embedded with Attention Mechanism”. SAE Technical Paper, 2024. [DOI]
  10. R. Shi, S. Yang, Y. Chen, et al. “Road Recognition for Autonomous Vehicles Based on Intelligent Tire and SE-CNN”. Intelligent Systems and Pattern Recognition, 2022. [DOI]

Contribution 合作论文

  1. J. Wang, R. Shi, J. Li, et al. “Manifold-Optimization-Based 3D Sound Source Mapping with Unknown Camera-Microphone Array Relative Pose”. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026. [DOI]
  2. J. Wang, R. Shi, Y. Kang, et al. “Fast-SDE: Efficient Single-Microphone Sound Source Distance Estimation in Reverberant Environments”. arXiv preprint, 2026. [arXiv]

Education教育背景

Institute of Science Tokyo, 东京科学大学(东京工业大学), Ph.D. in Systems and Control Engineering 系统与控制工程,博士

Beihang University, 北京航空航天大学, M.S. in Vehicle Engineering 车辆工程,硕士

Jilin University, 吉林大学, B.Eng. in Vehicle Engineering 车辆工程,本科