The creators of Design Arena have raised $7.9 million in seed funding to expand a platform that uses human preferences to evaluate the design abilities of artificial intelligence models.
The financing was secured by Intelligence, the company behind Design Arena. Index Ventures led the round, with participation from Conviction, A*, Valkyrie and other investors.
The startup says Design Arena has attracted 5.3 million users worldwide and is generating $60 million in annual recurring revenue. That figure would equal approximately $5 million in recurring revenue per month when annualized.
The reported revenue is about 7.6 times larger than the company’s new funding round, an unusual ratio for a young startup that began as a student project in 2025.
Design Arena originated a few weeks before co-founder Grace Li and several college friends graduated in 2025.
The group was initially attempting to build an AI-powered game engine. Although available models could produce games that functioned correctly, the resulting experiences were often not enjoyable.
That problem exposed a weakness in conventional AI evaluation. A system can determine whether software runs, but measuring whether a website is attractive, a game is entertaining or an image has the right visual style requires human judgment.
The founders began developing a system that could collect subjective feedback from large numbers of people.
About one week after identifying the commercial opportunity, the team reportedly signed its first major agreement with a frontier AI laboratory.
Design Arena allows users to enter a prompt and compare outputs produced by different AI models.
The platform supports websites, images, games, mobile applications, presentations, data visualizations, 3D designs, video, audio and several other creative formats.
Instead of showing users which company produced each result, Design Arena presents the outputs anonymously. Users select the option they prefer through a series of head-to-head comparisons.
A typical tournament samples 4 models and conducts 5 individual battles. The process produces a complete ranking of the 4 outputs while generating multiple preference signals for the platform’s leaderboards.
Removing model names is intended to reduce brand bias. Users can judge the quality of an output without knowing whether it came from OpenAI, Google, Anthropic, xAI or another developer.
The consumer interface helps users locate the model that produces the design they prefer. However, Intelligence’s main commercial opportunity comes from selling evaluation data and feedback to AI companies.
Model developers can use the rankings to understand which outputs people find attractive, useful or enjoyable.
This information can complement automated benchmarks, which typically test whether a model completes a technical task correctly. Automated tests may determine whether a generated website loads, but they may not reliably determine whether its layout looks professional or whether its design matches user expectations.
Design Arena’s voting system is intended to measure those subjective qualities at scale.
Each user interaction produces information about the relative strengths of multiple models. Millions of interactions can reveal patterns that may not appear in smaller laboratory evaluations.
Users must log in before receiving their final output, giving Intelligence the ability to examine how design preferences vary across locations and change over time.
The company has observed that preferred design styles are not necessarily universal. For example, users in some Asian markets may favor more visually dense or maximalist dashboard designs than users in other regions.
Such differences could help AI laboratories customize models for particular markets rather than relying on a single global definition of good design.
The data could also show how preferences change as new visual trends emerge. A model that performs well on evaluations created one year earlier may not satisfy users whose expectations have shifted.
Design Arena’s reported scale is significant for a company launched roughly one year earlier.
Its 5.3 million users represent more than 4 times the 1.3 million users previously reported by Yupp, another human-feedback platform that shut down after raising $33 million.
Intelligence’s $60 million in claimed annual recurring revenue works out to approximately $11.32 for every reported Design Arena user. The actual business model is more concentrated because enterprise AI laboratories, rather than ordinary users, appear to generate most of the revenue.
The company has raised approximately 13 cents in seed financing for every $1 of reported annual recurring revenue.
These calculations do not reveal profitability, customer concentration or operating costs. Intelligence has also not disclosed its valuation or identified all of its paying enterprise customers.
Design Arena is entering a growing market for independent AI evaluation.
LM Arena, which applies a comparable voting approach to text-based AI responses, raised $150 million in Series A funding in January 2026. That round was approximately 19 times larger than Intelligence’s $7.9 million seed financing.
However, the failure of Yupp demonstrates that attracting users and AI companies does not guarantee a sustainable business.
Yupp raised $33 million and reached more than 1.3 million users before closing less than one year after launch. Its shutdown showed that evaluation platforms must convert user activity into dependable enterprise contracts while controlling the cost of providing access to multiple AI models.
Design Arena may have an advantage if its reported $60 million revenue run rate is maintained. The company’s focus on visual quality also gives it a more specialized position than platforms primarily comparing text responses.
Traditional benchmarks have played a central role in measuring the progress of AI models. However, developers may optimize systems specifically for popular tests, reducing the value of benchmark scores as indicators of real-world performance.
Crowdsourced platforms attempt to address this weakness by continuously collecting new prompts and preference data from real users.
Design Arena’s methodology requires models to accumulate enough comparisons before their scores are treated as reliable. Models with limited data are marked as preliminary, while some rankings require approximately 200 pairwise comparisons before reaching greater statistical reliability.
Human evaluation also presents challenges. User preferences can be inconsistent, coordinated voting can influence results and popular design styles are not always the most functional or accessible.
Intelligence will need to demonstrate that its data remains representative as the platform grows.
The startup’s broader argument is that AI progress should not be measured only through coding accuracy, reasoning tests or technical completion rates.
As models become capable of producing websites, applications, images, games and presentations, qualities such as visual balance, originality and enjoyment will become more commercially important.
Those qualities are difficult to express through a simple automated score.
Design Arena is attempting to convert millions of subjective judgments into structured evaluation data that AI laboratories can use to improve their models.
With $7.9 million in fresh funding, 5.3 million reported users and a claimed $60 million annual revenue run rate, Intelligence is betting that human taste will become one of the technology industry’s most valuable AI benchmarks.
Share your thoughts about this article.
Be the first to post a comment!