Model Arena: Ranking Language Models by Head-to-Head Battles

A single score on a single benchmark tells you little about how a model behaves across work. MYND Model Arena is built around a different frame. Models compete head to head, each battle updates a rating, and the leaderboard is the result of many battles. The README describes it as an evolutionary benchmarking platform inspired by biology, with tournament-style ELO ranking.

How it is put together

It is a Turborepo monorepo. The dashboard is SvelteKit 2 with Svelte 5 and Chart.js. The API is Fastify 4 with TypeScript, JWT auth, rate limiting and auto-generated Swagger docs. Data lives in PostgreSQL 16 through Drizzle ORM, and Redis holds benchmark job queues and cached results. WebSockets stream benchmark progress live to the dashboard.

The README lists benchmark dimensions for reasoning, coding with HumanEval and MBPP, creativity, factual accuracy and agent tasks, and says you can define your own suites. Model calls go through the OpenAI API, so GPT-4o, GPT-4 and o1 are available out of the box.

Where head-to-head ranking comes from

Pairwise battles are not my invention. The Chatbot Arena paper, submitted in March 2024, describes an open platform that evaluates LLMs by human preference. A later paper on vote rigging describes the same setup as pairwise battles where users vote for the preferred response from two randomly sampled anonymous models. That second paper also shows why such rankings need care: crowdsourced votes can be manipulated.

A difference worth stating

In Chatbot Arena the judges are people voting. In Model Arena as the README describes it, the battles run across standardized benchmark suites and are scored by the engine. So the ratings answer a different question: how models rank on the suites you chose, not which answers a crowd preferred. Neither is a full picture, and I would use Model Arena to compare models on tasks that match my own work.

What I have not published

I have not released a leaderboard from this system, so there are no rankings to quote. The repository is the platform. Any numbers come from whoever runs it, with their own suites and keys.

Sources

MYND Model Arena README: github.com/yethikrishna/mynd-model-arena. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, arxiv.org/abs/2403.04132. Improving Your Model Ranking on Chatbot Arena by Vote Rigging, arxiv.org/html/2501.17858v2.

← back to the journal