AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
An Luo, Jin Du, Xun Xian, et al.
AgentDS is a benchmark that tests how well AI agents and humans working together can solve real-world data science problems across industries like healthcare, retail, and manufacturing. The research found that while AI agents alone perform poorly on these domain-specific tasks, human-AI collaboration produces the best results, showing that human expertise remains crucial even as AI becomes more advanced.
AI agentsbenchmarkinghuman-AI collaborationlarge language models