NeurIPS Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Poster
in
Workshop: Statistical Frontiers in LLMs and Foundation Models

Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Siyuan Wang · Zhuohan Long · Zhihao Fan · Xuanjing Huang · zhongyu wei

Keywords: [ dynamic LLM evaluation ] [ benchmark self-evolving ] [ evaluation bias ]

[ Abstract ] [ Project Page ]

[ OpenReview]

Sat 14 Dec 3:45 p.m. PST — 4:30 p.m. PST

Abstract:

This paper presents a benchmark self-evolving framework to dynamically evaluate rapidly advancing Large Language Models (LLMs). We utilize a multi-agent system to reframe new evolving instances with high confidence that extend existing benchmarks. Towards a more scalable, robust and fine-grained evaluation, we implement six reframing operations to construct evolving instances testing LLMs against diverse queries, shortcut biases and probing their problem-solving sub-abilities. With this framework, we extend datasets across general and specific tasks, through various iterations. Experimental results show a performance decline in most LLMs against their original results under scalable and robust evaluations, offering a more accurate reflection of model capabilities alongside our fine-grained evaluation. Besides, our framework widens performance discrepancies both between different models and within the same model across various tasks, facilitating more informed model selection for specific tasks. Overall, this sustainable framework contributes the research community for continuously evolving benchmarks alongside LLM development. \footnote{Code and data are available at \url{https://anonymous.4open.science/r/Self-Evolving-Benchmark }.}

Chat is not available.

Poster in Workshop: Statistical Frontiers in LLMs and Foundation Models

Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Siyuan Wang · Zhuohan Long · Zhihao Fan · Xuanjing Huang · zhongyu wei

Poster
in
Workshop: Statistical Frontiers in LLMs and Foundation Models