Intelligence has raised $7.9 million to scale Design Arena, a platform that turns users’ choices between competing AI outputs into preference data. The bet is commercially significant: as model builders approach diminishing returns from automated benchmarks and conventional training data, the companies that can measure what people actually like may become strategic suppliers to the frontier AI market.
The most valuable data in artificial intelligence may no longer be the data that helps a model produce a technically correct answer. It may be the data that explains why one answer, image, interface, or game feels better than another.
That is the business opportunity behind Design Arena, a product from the company Intelligence. The startup has raised a $7.9 million seed round led by Index Ventures to build a commercial platform around a problem that has become increasingly important as generative AI systems improve: measuring taste.
Design Arena allows users to enter prompts for visual work, including websites and images, and then compare competing outputs in an A-versus-B format. Users repeatedly select the result they prefer, allowing the platform to rank outputs based on collective feedback. The experience resembles an AI model router, where several systems compete to answer the same request. But the strategic value for Intelligence is not primarily the comparison interface. It is the stream of preference data generated by those decisions.
The company says Design Arena is used by 5.3 million people worldwide and generates $60 million in annual recurring revenue. It also says frontier AI labs pay for access to scalable feedback from users. Those figures are company claims, and the business is still early enough that its long-term durability remains unproven. But the underlying market signal is clear: AI developers are beginning to treat human judgment as a product in its own right.
That creates a new competitive question for the industry. If the next generation of models depends on understanding not only what is correct but what users prefer, will frontier labs build preference measurement internally, or will they rely on specialist platforms that aggregate feedback across millions of people?
The answer will shape where value accumulates in the AI stack.
From model comparison to data infrastructure
At the surface level, Design Arena offers a familiar consumer experience. A user submits a prompt and sees two outputs. The user chooses the stronger result, and the process continues. This format is simple enough to encourage participation and structured enough to generate data that can be analyzed.
The distinction between the two sides of the business is important. One side is the visible product: an interactive comparison tool for people who want to test or use different AI systems. The other is the less visible commercial layer: a dataset recording how users evaluate those systems.
That dataset can be useful in several ways. It can help model developers identify which outputs win in specific categories. It can reveal whether a model performs better on website design than image generation, or whether a particular system produces outputs that appeal to one market but not another. It can also track changes over time, showing whether a model update improves user preference or merely raises a score on a conventional benchmark.
This is a more complicated problem than measuring factual accuracy. A mathematical answer is often right or wrong. A generated interface can be functional while still feeling clumsy. An image can satisfy the prompt while appearing generic. A game can operate correctly but fail to hold a player’s attention. These judgments are contextual, subjective, and often difficult to express in a written instruction.
A-versus-B testing provides a practical way to capture them. Instead of asking users to assign an abstract score, it asks them to make a relative decision. That reduces the burden on participants and creates a consistent format for aggregating opinions.
For model developers, the output is potentially more actionable than a broad survey. The data is tied to actual model outputs, prompts, and choices. It can be segmented by task, geography, user cohort, or other attributes that a platform is able to collect and analyze responsibly. It may show not simply that users dislike an output, but which alternative they selected instead.
The commercial attraction is therefore straightforward. An AI lab can spend heavily on training compute and still struggle to determine whether a new model is better in the ways that matter to customers. A preference platform offers an external measurement layer. It can provide continuous feedback after a model is released, when users are interacting with products in real-world conditions rather than within a fixed internal test set.
This is why the business should be understood as more than an AI comparison website. Its potential role is closer to data infrastructure for model evaluation and product development.
Why taste is becoming a technical problem
The AI industry has traditionally relied on benchmarks to compare models. Benchmarks are useful because they are standardized, repeatable, and easy to communicate. A company can announce that its model scored a certain percentage on a coding, reasoning, or language test. Investors, developers, and customers can use those figures to form a rough view of progress.
But benchmarks have limitations, particularly as models become capable of performing well across a wide range of standard tests. They can become targets for optimization, lose their ability to distinguish leading systems, or measure narrow abilities that do not translate into product adoption.
They also struggle with subjective output. There is no universal answer to which of two landing pages has better visual hierarchy, which generated portrait feels more natural, or which creative concept is more compelling. A benchmark can check whether a page includes required elements. It cannot easily determine whether the page communicates trust or looks like something a customer would actually use.
That gap matters because generative AI is moving into markets where preference drives purchasing decisions. Companies are using AI to create marketing material, product designs, customer-facing interfaces, video, entertainment, and software. In these categories, technical compliance is only the starting point. The product must also produce an outcome that people value.
The challenge is especially acute for companies selling AI tools to businesses. A model can be cheaper, faster, or more accurate in a narrow technical sense and still lose customers if its outputs require too much editing or fail to match a brand’s aesthetic. Human preferences become a proxy for commercial usefulness.
Preference data is already part of the model development process through methods such as reinforcement learning from human feedback. But traditional feedback collection can be expensive and operationally heavy. Human evaluators may be hired to assess responses against detailed criteria, often in controlled environments. That approach can produce high-quality labels, but it may not capture the behavior of a broad population using a product for its intended purpose.
A platform such as Design Arena approaches the problem from the opposite direction. It seeks scale and repeated judgments from a large user base. The benefit is breadth: more choices, more categories, and potentially more natural reactions. The cost is that the data may be noisier, less representative, and harder to interpret.
The market will determine whether the scale advantage outweighs those limitations.
A new supplier class for frontier labs
If Design Arena’s model works, it could occupy a position between consumer software and enterprise AI infrastructure. The platform attracts users with an engaging comparison experience, while selling the resulting intelligence to model developers.
That structure has an important strategic advantage. The platform can generate feedback continuously rather than through occasional commissioned studies. Every new model, prompt type, and product category can create another opportunity for comparison. If enough users participate, the platform may build a large historical record of what people prefer and how those preferences change.
The data could be valuable to several types of customers.
Frontier labs could use it to evaluate model releases, identify weaknesses, and prioritize training. Application companies could compare models for specific workflows rather than relying on general-purpose rankings. Enterprise buyers could use preference results to select systems for design, content generation, or interface development. Advertising and creative firms could potentially use similar signals to evaluate concepts before investing in production.
The strongest commercial position would come from becoming a standard testing layer across multiple AI providers. If developers consistently send their outputs through the same independent platform, Design Arena could create a common reference point for quality. Its data would become more valuable as more models participate because users would be comparing a wider range of alternatives.
That is a classic network effect, although it is not guaranteed. More users generate more feedback. More feedback attracts more model developers. More participating models create more useful comparisons for users. If the cycle works, the platform can compound its advantage.
However, network effects in data businesses are often weaker than they appear. A rival can create a similar interface. A large AI lab can recruit users directly. And customers may not want to share sensitive prompts or outputs with an independent intermediary. The defensibility depends on the quality, scale, and uniqueness of the data, not simply on having a popular website.
Intelligence’s reported revenue is an important signal because it suggests the company has already found willingness to pay. A claimed $60 million in annual recurring revenue would be substantial for a seed-stage company. But the source and composition of that revenue matter. The durability of the business will depend on whether customers are signing long-term contracts, purchasing recurring access, or making more experimental payments while the market is still developing.
A preference-data company is most valuable when its data changes a customer’s decision. If labs use the platform only to generate interesting rankings, the business may remain a media product. If the feedback affects model training, product launches, and resource allocation, it becomes part of the AI development workflow.
The comparison with LM Arena
Design Arena is not alone in turning model comparisons into a product. LM Arena, which uses a similar concept for text-model comparisons, raised $150 million in a Series A in January. Its financing demonstrates that investors see substantial value in platforms that aggregate human judgments about AI systems.
The two businesses operate in adjacent but distinct areas. Text-model comparisons often focus on writing quality, reasoning, coding, instruction following, and general usefulness. Design Arena is oriented toward visual outputs and the broader category of aesthetic judgment.
That difference could matter. Text-based systems can often be evaluated with a combination of human review and automated tests. Visual and creative outputs are more difficult to reduce to objective metrics. A platform with expertise in visual preference may therefore become particularly valuable as image, website, video, and game-generation systems compete for users.
At the same time, the comparison raises a strategic issue: whether preference platforms will consolidate or specialize. One possibility is that a few large companies become general-purpose evaluation networks across text, images, video, audio, and software. Another is that specialized platforms develop deeper expertise in individual categories.
Specialization can improve the quality of feedback. A user evaluating a website may need different criteria from someone comparing two chat responses. The prompts, interface, and user population all influence the meaning of a preference. Design Arena’s focus on visual work could help it build a better dataset for those applications than a general platform.
Consolidation, however, may offer stronger economics. AI labs may prefer a single supplier that can deliver feedback across modalities and provide unified reporting. Companies that control broad user communities could also spread the cost of data collection across multiple products.
The competitive advantage will depend on how much context each platform can capture without making the experience too cumbersome. A binary vote is easy to collect, but a vote without information about the user’s goal can be difficult to interpret.
The warning from Yupp
The market also has an important counterexample. TechCrunch reported that Yupp, another human-evaluation startup, shut down earlier in 2026 after raising $33 million and attracting more than 1.3 million users.
Yupp’s failure does not disprove the preference-data thesis. It does show that user scale alone is not enough. A platform can attract participants and still fail to create a sustainable business if the data is not differentiated, customer demand is inconsistent, or the economics of collecting and processing feedback are unfavorable.
The example should make investors and customers more cautious about headline user numbers. A large audience is valuable only if it produces reliable, relevant, and repeatable judgments. Users may participate for entertainment without providing feedback that a model developer can use. They may also behave differently on a public platform than they would in a commercial workflow.
Retention is another issue. Comparison products can be engaging at first, particularly when users are curious about which model wins. But novelty may fade. The platform must give users a reason to return, whether through better creative tools, more useful model access, social features, rewards, or a continuing role in shaping products they use.
The business also needs to balance participation incentives with data quality. If users are paid, they may complete more evaluations but rush through them. If they are not paid, the most active participants may be hobbyists whose preferences do not represent the customers that AI companies are targeting. If the platform optimizes for volume, it may produce a large dataset with limited decision-making value.
Design Arena’s reported revenue suggests it may have addressed some of these challenges, at least commercially, but the Yupp shutdown illustrates how quickly the model can break down.
The representativeness problem
Human preference data is not automatically objective simply because it comes from people. It reflects who participates, what they understand, and what incentives they have.
A platform used by millions can still overrepresent particular countries, age groups, technical communities, or online cultures. A design that wins among frequent AI users may not win among the broader audience for a retail website. An image preferred by one market may perform poorly in another. Preferences can also vary by device, language, profession, and familiarity with generative tools.
That creates both a challenge and a potential advantage. If Design Arena can identify differences among user segments, it may provide more useful data than a single global score. A model developer could learn that one system produces outputs preferred by professional designers while another is more attractive to casual consumers. It could test whether an interface works better in one region or whether visual styles are becoming less popular over time.
But segmentation requires careful methodology. The platform must know which attributes are relevant, collect them responsibly, and avoid turning uncertain behavioral patterns into definitive conclusions. It must also protect user privacy and manage the rights associated with prompts, uploaded material, and generated outputs.
Trust will be a decisive factor in enterprise sales. AI labs may want preference data, but they will also worry about data provenance, consent, manipulation, and confidentiality. A customer comparing unreleased models or proprietary workflows may not want those interactions exposed to a public audience. The platform must demonstrate that its data is not only plentiful but controlled.
There is also the possibility of coordinated voting. If users, developers, or online communities can influence rankings, public leaderboards may become targets for gaming. A preference platform would need safeguards against duplicate accounts, automated participation, strategic voting, and changes in user behavior caused by the visibility of previous results.
Automated benchmarks are vulnerable to gaming, but human systems are vulnerable in different ways. The relevant question is not whether one method is perfectly reliable. It is whether combining human preference with technical tests produces a better picture of product quality.
Will labs buy the capability or build it?
The largest strategic threat to Design Arena comes from its customers. Frontier labs have enormous incentives to own the most valuable parts of the development pipeline. If preference data becomes critical, they may decide to build the capability themselves.
Large labs already have access to millions of users through chatbots, productivity software, developer tools, and creative applications. They can collect feedback directly from those products. They can run A-versus-B experiments, measure user engagement, and connect preferences to downstream behavior such as retention, editing, sharing, or purchase.
Internal systems offer advantages that an independent platform cannot easily match. A lab can observe the full user journey, not just a single vote. It can test outputs in the context where they will be used. It can link feedback to product metrics and make rapid changes to models and interfaces.
An independent platform still has a reason to exist if it can offer cross-model neutrality. A company that develops one model cannot easily provide a trusted comparison across its competitors. Design Arena could also aggregate feedback from users who interact with multiple systems, creating a common evaluation environment that no individual lab can reproduce.
The likely outcome may be a hybrid market. Labs will collect large volumes of product feedback internally while purchasing external data for cross-model comparisons, independent validation, specialized categories, and geographic or demographic segments they cannot reach efficiently.
Synthetic evaluation is another competitive threat. AI systems can already be used to judge other AI outputs. Synthetic judges are cheap, fast, and easy to scale. They can evaluate millions of examples at a fraction of the cost of human participation.
But synthetic evaluation inherits the biases and limitations of the systems generating it. A model may be able to explain why an image appears polished without knowing whether real users would choose it. Automated judges can also converge on stylistic preferences that are easy to describe rather than qualities that produce genuine engagement.
Human feedback will remain most valuable where the target is difficult to formalize. The question is whether its incremental value justifies the cost. If synthetic systems become sufficiently predictive of human behavior, preference platforms will face pressure on pricing. If they do not, human judgment could become a scarce input for high-stakes model development.
The economics of higher-level labeling
Traditional data-labeling businesses often rely on large workforces performing repetitive tasks. Their value comes from supplying labor at scale and meeting quality requirements across massive datasets. Preference platforms operate at a higher level of abstraction. Instead of asking a worker to identify an object or transcribe a sentence, they ask a person to make a judgment about quality.
That can support higher pricing, but it also creates more complex operational demands. The company must understand what each choice means, design comparisons that reduce ambiguity, and determine how many votes are needed before a result is reliable. It must account for disagreement rather than treating every preference as an unquestionable label.
The output may be more valuable than a conventional annotation because it can influence a model’s style, usability, and market positioning. Yet it may also be harder to standardize. A preference dataset is not simply a collection of facts. It is a record of decisions made under particular conditions.
For Intelligence, the key execution challenge will be converting those decisions into products that AI developers can use. Customers may want raw rankings, but they may also need diagnostic reports, evaluation APIs, training datasets, and tools for testing new models privately. The company will need to move from an engaging front end to repeatable enterprise infrastructure.
Its seed financing from Index Ventures gives it capital to pursue that transition. The funding can support product development, data-quality systems, customer acquisition, and expansion into additional creative categories. But it also raises expectations. A company with meaningful reported revenue must demonstrate that growth is not dependent on novelty or a small number of customers.
What success would look like
Design Arena does not need to replace automated benchmarks to become valuable. It needs to establish that human preference adds information that labs cannot obtain as efficiently elsewhere.
The strongest proof would come from customers using its data in measurable ways: selecting a model for a production workflow, improving a model release, reducing editing time, increasing user retention, or identifying regional differences that internal testing missed. Recurring contracts would matter more than one-off evaluations. High retention among both users and enterprise buyers would show that the platform is part of an ongoing process rather than a temporary experiment.
The company will also need to prove that its rankings are stable and interpretable. If the same outputs produce radically different results depending on minor changes to the interface or participant mix, customers may struggle to rely on the data. If the results remain consistent across appropriately designed tests, the platform’s credibility will rise.
The broader market is moving in its favor. AI systems are expanding into domains where taste, usefulness, and emotional response matter. Model quality is becoming less about whether a system can generate something and more about whether it generates the right thing for a particular person and situation.
That shift creates demand for better measurement. It also creates a contest over who owns the feedback loop.
Design Arena’s wager is that the best feedback will come from a broad community making repeated choices across competing systems. Frontier labs may dispute that assumption, preferring internal telemetry or synthetic judges. Investors may question whether the data remains differentiated as more platforms enter the market. The failure of Yupp shows that a large audience and significant funding do not guarantee a viable business, while LM Arena’s financing shows that the category can attract substantial capital when investors believe the evaluation layer is becoming strategically important.
The outcome will depend on execution rather than the appeal of the concept alone. Intelligence must keep users engaged, maintain trustworthy data practices, demonstrate value to paying labs, and defend its position against companies with larger distribution or deeper technical resources.
If it succeeds, the company could help establish a new class of AI supplier: a marketplace and measurement network for human judgment. Its core product would not be a model, a chip, or a dataset scraped from the internet. It would be evidence about what people choose when confronted with competing machine-generated results.
That evidence may become increasingly important as the industry moves beyond raw capability. The next competitive advantage in AI will not be determined solely by which model can generate the most sophisticated output. It may belong to the companies that can learn fastest which outputs people actually want.