AI Benchmarking Platform Arena Valued at $3.1 Billion After 82% Jump in Nine Months

Arena, the platform where users benchmark and compare ChatGPT, Claude, Gemini, and other AI models, has raised $200 million in a new funding round. The deal values the company at $3.1 billion, up from $1.7 billion in January 2026. Notably, Arena does not build its own large language model; its business centers entirely on independently evaluating third-party AI systems.
The company announced the Series B round on October 8. Lightspeed Venture Partners and Khosla Ventures led the round, with participation from Salesforce Ventures, Dell Technologies Capital, 01 Advisors, Endeavor Catalyst, and existing investors.
Arena closed its previous major round on January 6, securing $150 million at a $1.7 billion valuation. In less than a year, the company’s valuation has surged by roughly 82%.
How Free AI Comparisons Turned into a Business
Arena originated as a UC Berkeley research project launched in 2023 under the name Chatbot Arena. The platform later operated as LMArena before rebranding to its current name in early 2026.
The core mechanics remain unchanged: users submit a single prompt to two anonymous models, compare their answers side by side, and vote on the better response. Thousands of these pairwise comparisons fuel leaderboards based not only on synthetic benchmarks, but on real-world user preferences.
Today, Arena evaluates far more than just text models. The platform benchmarks image and video generators, code-generation systems, search tools, and AI agents.
According to company figures, the platform has logged 350 million sessions, 62 million votes, and conducted over 1,000 new model evaluations since January. Agent Arena accumulated over 7 million sessions in its first five months, and the overall platform’s monthly audience spans tens of millions of users across more than 150 countries.
Arena remains free for casual visitors. The company generates revenue through offerings like AI Evaluations, a commercial service that helps developers and enterprise customers test their models on real-world tasks and receive detailed analytics.
In June, Arena reported that its annualized run rate exceeded $100 million. That does not mean the company has already generated $100 million within a calendar year; rather, the figure reflects the annualized projection of its current revenue pace.
Arena Now Evaluates AI Behavior, Not Just Performance
Alongside the funding announcement, the company unveiled the Arena Alignment Index, a benchmark measuring how faithfully AI agents follow user intentions.
Researchers analyzed around 90,000 real-world sessions across 27 models, identifying three primary failure modes: taking actions without user permission, attributing non-existent claims to users, and falsely reporting that tasks were completed.
The last issue proved particularly prevalent. On average, roughly 10% of sessions featured instances where an agent claimed to have finished a task without actually doing so. When debugging code, that figure climbed to 48%.
Arena emphasizes that the index remains preliminary and covers only a fraction of potential risks. However, the focus helps explain investor enthusiasm: as AI agents gain access to files, browsers, corporate data, and other tools, enterprises need to know not just which model is “smarter,” but how safely they can trust it with autonomous actions.
Developing these evaluation frameworks is precisely where Arena plans to allocate the fresh $200 million.