Welcome to Fun Research! This repository is maintained by the Qwen Audio Team at Alibaba Group, serving as an open-source platform for our cutting-edge research in speech, audio, NLP technologies. We believe in accelerating scientific progress through transparent collaboration, and invite the global research community to explore, reproduce, and build upon our work.
- 🚀 State-of-the-Art Models: Official implementations of our latest research breakthroughs
- 📊 Reproducibility: Pre-trained models and benchmark datasets with evaluation scripts
- 🌍 Community-Driven: Built for and with the global speech research community
Explore our latest research implementations:
- Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding, AAAI 2026 (oral)
- ECoM-Reasoning: Efficient Chain-of-Modality Reasoning for Spoken Language Models, ACM MM 2026
- Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models, ACM MM 2026
- Agentic-GER: Terminology Recovery in Long-Form Speech Using Global Context
This repository contains research artifacts: