The company behind DesignArena, an AI tool that ranks AI-generated outputs through A/B comparisons, has raised $7.9 million in a seed round led by Index Ventures. The startup, called Intelligence, plans to use the funding to expand a service already used by 5.3 million people worldwide. The round was announced Monday, and it marks a bet that human judgment, not just automated scoring, will decide which AI models win.
From a college project to a frontier lab deal
Intelligence started a few weeks before graduation in 2025. The founders were college friends working on an AI game engine. They hit a wall quickly. The models could build functional games, but the games were not fun. That gap led them to a simple conclusion: human judgment was necessary for evaluating fun.
Grace Li, co-founder of Intelligence and DesignArena, described the moment of realization. "It was the missing bottleneck for a lot of these models to make improvements in the design space," she said. The team pivoted from building games to building a tool that lets people rank AI outputs. The shift paid off fast. "About a week later, we closed our first major deal with a frontier lab, and the rest is kind of history," Li added.
DesignArena works like a model router with a ChatGPT-style window. Users pick a format, from websites and images to a dozen other visual styles. They then rank outputs through A/B choices until a best-to-worst order is established. That ranking data becomes training signal for media-generating models. The enterprise side of the business provides endless instant feedback for those models, which is where the real value sits.
A business growing on taste
The consumer side of DesignArena is free, but it feeds the commercial engine. Users must log in to get output, which lets the company track taste changes over time and across continents. Li noted that web dashboards in Asia tend to have a more maximalist design style. That kind of granular preference data is gold for AI labs trying to match local aesthetics.
The site is currently generating $60 million in annual recurring revenue, according to Li. That figure is striking for a seed-stage company. It suggests strong demand for human-led evaluation, especially as automated benchmarks face growing scrutiny. Automated benchmarks can be gamed or manipulated, and last week a Hugging Face breach demonstrated exactly how that manipulation can happen. The incident underscored why labs are turning to crowdsourced judgment as a complement to machine scoring.
The seed round also included participation from Conviction, represented by Sarah Guo and Mike Vernal, as well as A* and Valkyrie. The investor lineup signals confidence in the human feedback market, even as that market has shown mixed results for other players.
A market with winners and losers
Not every startup in this space has fared well. Yupp, a competitor in the human feedback space, shut down earlier this year after raising $33 million from a16z crypto's Chris Dixon. Yupp had over 1.3 million users at the time of its collapse, less than a year after launching. The failure shows that crowdsourced human feedback is not an automatic winning market. Execution, timing, and product fit matter as much as the underlying idea.
Yet other startups seem to be thriving. LM Arena, which focuses on text-based human evaluation, raised $150 million in a Series A round in January. That raise came just four months after LM Arena launched its paid product. The contrast between Yupp's shutdown and LM Arena's growth suggests the market is real but selective. Investors are willing to back human evaluation startups, but they are picky about which ones.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
DesignArena's trajectory sits somewhere between those two extremes. It has scale, revenue, and now venture backing. But it also faces the same question that killed Yupp: can a consumer-facing ranking tool sustain long-term engagement? Li argues the answer is yes, because users are indifferent to which models they rank. They want the best output, not a specific brand. That indifference makes the data more honest and the feedback loop more valuable.
Why human feedback matters now
The broader context is a shift in how AI models are evaluated. Automated benchmarks have long been the standard, but they have a known weakness: they can be gamed. The Hugging Face breach last week was a stark reminder of that vulnerability. If a benchmark can be manipulated, the scores it produces are worthless. Human feedback offers a different kind of signal, one based on actual preference rather than pattern matching.
For media-generating models, that signal is critical. A model can ace a technical benchmark and still produce ugly websites or awkward images. Only human eyes can judge aesthetics, fun, and style. DesignArena taps into that need by turning millions of casual users into a distributed evaluation team. The company tracks taste changes over time and across continents, building a picture of what people actually want from AI output.
Li framed the opportunity in terms of what the models are missing. The design space, she said, was the bottleneck. That framing explains why a seed-stage company can already pull in $60 million in annual recurring revenue. The demand is not speculative. It is coming from frontier labs that need better feedback loops now.
What the round means for the sector
The $7.9 million seed round validates the market for human feedback in AI, at least as far as investors are concerned. Index Ventures leading the round gives it weight. Conviction, A*, and Valkyrie joining adds further credibility. The money will likely go toward scaling the platform, improving the ranking experience, and expanding the enterprise side of the business.
The timing is notable. Yupp's failure earlier this year raised questions about the viability of the human feedback model. LM Arena's $150 million Series A in January answered some of those questions. DesignArena's seed round, announced Monday, reinforces the answer. The market is not dead. It is consolidating around companies that can execute.
DesignArena's user base of 5.3 million people gives it a moat that Yupp never had. Its $60 million in annual recurring revenue gives it a runway that most seed-stage startups can only dream of. And its early deal with a frontier lab, closed about a week after the company started, shows that demand was there from day one.
The road ahead is not without risk. Crowdsourced feedback can suffer from quality issues, and taste is a moving target. But the company's ability to track preference shifts across continents and over time gives it a unique vantage point. If human judgment is indeed the missing bottleneck for AI improvement, DesignArena is positioned to be the pipeline that delivers it.

