MatrAIx is a new evaluation infrastructure designed to simulate the world with 8.3 billion persona agents, aiming to address the high cost, slow speed, and limited scalability of human evaluation for AI systems and digital products. The project, led by over 200 scientists from Harvard, MIT, and more than 40 researchers from OpenAI, Anthropic, Google DeepMind, and xAI, offers a scalable alternative to traditional user studies.
The infrastructure comprises three core components: the Persona-8B database with 8.3 billion persona records across 1,290 categorical dimensions; the MatrAIx Playground with four environments (Survey, AI Chatbot, Web, and App); and 1,010 application tasks spanning over 25 domains including Commerce, Software, Finance, and Healthcare.
In a series of 18,189 evaluation trials across eight representative tasks, personas powered by Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5 demonstrated varied feedback on product decisions and preferences, such as hesitation after price increases, willingness to continue after AI assistant failures, and latency tolerance.
Validation studies showed high persona adherence: in a 400-trial controlled study, declared behavior was expressed or correctly suppressed in 91.5% of trials. Additionally, human and LLM judges evaluated the extraction quality of human-grounded personas, confirming the reliability of the approach.
MatrAIx provides an end-to-end solution for evaluating AI systems with diverse simulated users, potentially transforming how digital products are tested at scale.