About Me个人简介

I am a doctoral student at Nakadai Lab, Institute of Science Tokyo. My research interests include speech/audio generative models, audio source separation, and spatial audio. I welcome collaborations ~ 🤖

我是东京科学大学 Nakadai Lab 的博士生。我的研究方向包括语音/音频生成模型、音源分离与空间音频等,欢迎合作~ 🤖

News动态

Education教育背景

Institute of Science Tokyo, 东京科学大学(东京工业大学), Ph.D. in Systems and Control Engineering 系统与控制工程,博士

Beihang University, 北京航空航天大学, M.S. in Vehicle Engineering 车辆工程,硕士

Jilin University, 吉林大学, B.Eng. in Vehicle Engineering 车辆工程,本科

Project项目

  • Unsupervised Single-Channel Audio Separation with Diffusion Source Priors

    Open-source implementation for the AAAI 2026 paper on unsupervised single-channel audio separation across speech-sound, sound-sound, and speech-speech mixtures.

    AAAI 2026 论文的开源实现,面向 speech-sound、sound-sound 与 speech-speech 等单通道音频分离任务。

  • Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance

    Code for unsupervised single-channel speech separation with diffusion source models and speaker-embedding guidance.

    基于扩散源模型和说话人嵌入引导的无监督单通道语音分离代码。

  • Real Grid RIR Dataset

    A real room impulse response dataset collected at grid source locations in a meeting room, useful for distance estimation, spatial audio separation, and target extraction.

    会议室网格位置采集的真实房间脉冲响应数据集,适用于距离估计、空间音频分离和目标提取等研究。

Publication论文

  1. R. Shi, C. Li, J. Wang, et al. “Unsupervised Single-Channel Audio Separation with Diffusion Source Priors”. Proceedings of the AAAI Conference on Artificial Intelligence, 2026. [arXiv]
  2. R. Shi, B. Yen, and K. Nakadai. “Distance Based Single-Channel Target Speech Extraction”. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025. [arXiv]
  3. R. Shi, Z. Lin, B. Yen, et al. “Single-Channel Target Speech Extraction Utilizing Distance and Room Clues”. European Signal Processing Conference (EUSIPCO), 2025. [arXiv]
  4. J. Wang*, R. Shi*, B. Yen, et al. “Single-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments”. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025. [arXiv] * equal contribution
  5. R. Shi, K. Li, Y. Wang, et al. “Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance”. Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2026. [arXiv]
  6. R. Shi, C. Li, J. Li, et al. “Exploring Efficient Waveform Diffusion Models for Foley Sound Generation”. Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2026. [arXiv]
  7. R. Shi, Y. Wang, H. Song, et al. “Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS”. Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2026. [arXiv]
  8. R. Shi, K. Itoyama, and K. Nakadai. “Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning”. Vocal Interactivity in-and-between Humans, Animals and Robots (VIHAR), 2024. [arXiv]
  9. R. Shi, S. Yang, Y. Chen, et al. “CNN-Transformer for Visual Tactile Fusion Applied in Road Recognition of Autonomous Vehicles”. Pattern Recognition Letters, 2023. [DOI]
  10. R. Shi, S. Yang, J. Lu, et al. “Road Profile Reconstruction Based on Recurrent Neural Network Embedded with Attention Mechanism”. SAE Technical Paper, 2024. [DOI]
  11. R. Shi, S. Yang, Y. Chen, et al. “Road Recognition for Autonomous Vehicles Based on Intelligent Tire and SE-CNN”. Intelligent Systems and Pattern Recognition, 2022. [DOI]

  1. J. Wang, R. Shi, Y. Kang, et al. “Fast-SDE: Efficient Single-Microphone Sound Source Distance Estimation in Reverberant Environments”. arXiv preprint, 2026. [arXiv]
  2. J. Wang, R. Shi, J. Li, et al. “Manifold-Optimization-Based 3D Sound Source Mapping with Unknown Camera-Microphone Array Relative Pose”. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026. [DOI]
  3. N. Khan, B. Wu, R. Shi, et al. “From Sign Language Generation to Humanoid Execution: Vision-Language Guided Retargeting with Collision Mitigation”. arXiv preprint, 2026. [arXiv]
  4. Y. Kang, J. Wang, R. Shi, et al. “What Do Neural Networks Learn for TDOA Estimation? A Cross-Architecture Probing Study”. arXiv preprint, 2026. [arXiv]
  5. R. A. Nihal, B. Yen, R. Shi, et al. “Ecologically-Constrained Task Arithmetic for Multi-Taxa Bioacoustic Classifiers Without Shared Data”. arXiv preprint, 2026. [arXiv]
  6. R. A. Nihal, B. Yen, R. Shi, et al. “Weakly Supervised Multiple Instance Learning for Whale Call Detection and Localization in Long-Duration Passive Acoustic Monitoring”. arXiv preprint, 2025. [arXiv]
  7. R. Cao, Z. Zhang, R. Shi, et al. “Model-Constrained Deep Learning for Online Fault Diagnosis in Li-Ion Batteries over Stochastic Conditions”. Nature Communications, 2025. [DOI]
  8. J. Lu, Z. Peng, R. Shi, et al. “A Quantitative Blind Area Risks Assessment Method for Safe Driving Assistance”. Journal of Systems Architecture, 2024. [DOI]
  9. J. Lu, Y. Cao, R. Shi, et al. “An Efficient Driver Anomaly State Detection Approach Based on End-Cloud Integration and Unsupervised Learning”. IEEE International Conference on Intelligent Transportation Systems (ITSC), 2023. [DOI]
  10. J. Lu, S. Yang, R. Shi, et al. “Safety Co-Pilot: A System for Autonomous Vehicle to Make Decision Safer and Smarter”. IEEE International Conference on Intelligent Transportation Systems (ITSC), 2022. [DOI]
  11. S. Yang, Y. Chen, R. Shi, et al. “A Survey of Intelligent Tires for Tire-Road Interaction Recognition Toward Autonomous Vehicles”. IEEE Transactions on Intelligent Vehicles, 2022. [DOI]
  12. S. Yang, R. Wang, R. Shi, et al. “An Intelligent Tyre System for Road Condition Perception”. International Journal of Pavement Engineering, 2023. [DOI]