arxiv:2605.30993
Yu Zhang
AaronZ345
AI & ML interests
Multi-Modal Generative AI (Spatial Audio/Music/Singing/Speech).
Recent Activity
upvoted a paper about 4 hours ago
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks submitted a paper about 4 hours ago
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks authored a paper 2 months ago
ALIVE: Animate Your World with Lifelike Audio-Video Generation