BBYR Achieve
返回信息流
这是一条镜像帖。来源:北邮人论坛 / job-info / #978316同步于 2026/1/28
该镜像源已超过 30 天没有更新,可能在源站已被删除。
JobInfo机器人发帖

【社招】【大模型语音算法】base北京,上海

DLL641388166
2026/1/28镜像同步0 回复
微信:18518019493 邮件:641388166@qq.com 语音算法 工作职责: 1、负责语音理解和语音生成算法在滴滴场景的落地使用 2、跟进最新技术,结合业务场景,提升语音识别、音频事件检测、声纹识别、语音合成等算法效果 3、探索语音大模型或多模态大模型在语音理解及语音生成场景的应用范式 4、算法优化,从模型架构、推理框架、量化压缩等角度提升模型推理速度、降低推理成本 Job Description 1. Responsible for the implementation of speech understanding and speech generation algorithms in Didi’s business scenarios. 2. Stay updated with the latest technologies and improve the performance of algorithms such as speech recognition, audio event detection, speaker recognition in real-world applications. 3. Explore the application paradigms of large language models or multimodal models in speech understanding and generation scenarios. 4. Optimize algorithms by enhancing inference speed and reducing costs through improvements in frameworks and quantization 任职资格: 1、电子、计算机或相关声学、信号处理专业毕业,具备一定语音信号处理基础 2、熟悉Pytorch框架,良好的编程能力,熟练使用python编程语言,具备Linux平台开发经验 3、3-5年语音识别、音频事件检测、声纹识别、语音合成等算法经验 4、在ICASSP、Interspeech、ASRU等语音顶会或国际竞赛有论文发表或优异成绩优先 5、ACM竞赛取得优异成绩或有优秀C++编程能力优先 6、有大型语音识别、语音合成项目经验者优先 Qualifications 1. Bachelor’s or higher degree in electronics, computer science, acoustics, signal processing, or related fields, with a solid foundation in speech signal processing. 2. Proficient in PyTorch framework, strong programming skills, fluency in Python, and experience in Linux platform development. 3. 3-5 years of hands-on experience in algorithms such as speech recognition, audio event detection, speaker recognition, and speech synthesis. 4. Candidates with publications or outstanding achievements in top-tier speech conferences (e.g., ICASSP, Interspeech, ASRU) or international competitions are preferred. 5. Candidates with excellent ACM competition results or strong C++ programming skills are preferred. 6. Experience in large-scale speech recognition or speech synthesis projects is a big plus. 语音大模型算法研究员(大模型 评测 合成 识别 降噪等等方向均可) 职位描述 1、主导或深度参与语音多模态大模型的架构设计、训练、调优及迭代工作,提升模型整体性能,研究方向包括但不限于音频表征、音视频理解、语音生成、全双工语音对话、语音强化学习。 2、跟踪学术界与工业界的前沿技术动态,能够复现最新成果,推进技术迭代与创新突破。 3、系统性设计并执行实验方案,深入分析模型表现,定位核心问题并提出有效的优化与解决方案。 4、参与团队合作,与团队一起解决技术难题,推动技术进展。 职位要求 1、985/211高校研究生及以上学历或优秀本科生,计算机、人工智能、软件、数学等相关专业。 2、精通常用算法和数据结构, 熟悉Python等编程语言, 熟悉Linux平台。 3、熟悉机器学习、深度学习,具备扎实的算法基础, 熟练使用Pytorch/Megatron深度学习框架。 4、在语音多模态大模型某一领域(语音大模型、Voice Agent、生成模型等)有过深入的研究经历。 5、具备卓越的实验分析与问题解决能力,有创新思维与较强的学习能力,能够良好沟通、与团队成员高效协作。 加分项 1、有顶级学术会议(ICASSP、Interspeech、ACL、ICML、NIPS、CVPR等)发表过有影响力研究成果的优先。 2、在 ACM/ICPC、NOI/IOI、Kaggle 等编程/AI 比赛获奖者优先。 3、主导、参与过 AI 相关的有大影响力的开源/闭源项目的优先。
订阅后,新回复会通过你的通知中心匿名送达。
0 条回复
暂无回复 · 你可以订阅本帖等待新回复。