
OpenAI is fundamentally shifting the landscape of AI evaluation with its new GDPval suite, designed to measure model performance on real-world, economically valuable tasks [1]. Moving beyond abstract academic benchmarks, this framework assesses AI capabilities across 44 occupations within nine major U.S. economic sectors. At the heart of GDPval is a methodology grounded in practical utility: blinded pairwise comparisons. In this evaluation method, a human expert reviews two outputs side-by-side without knowing their source – for instance, which was created by an AI – and simply chooses the better one, providing a direct and unbiased judgment of quality. This approach replaces abstract scores with direct, qualitative judgments on authentic deliverables. To facilitate broader research, OpenAI has also released a 220-task...








