🎉 One paper accepted to COLM 2026! Agent Q-Mix - Selecting the Right Action for LLM Multi-Agent Systems through Reinforcement Learning