Benchmarks · Language & multilingual
IndQA
Whether a model can answer questions that require cultural and contextual knowledge specific to India, in Indian languages, rather than knowledge that happens to be translated into them.
OpenAIReleased 5 November 2025Live
IndQA was OpenAI’s attempt to test something existing multilingual benchmarks did not: not whether a model can answer a general-knowledge question that has been translated into an Indian language, but whether it understands context that only makes sense within Indian culture in the first place. The set covers 2,278 questions across 12 languages and language varieties — including Hinglish, the code-switched mix common in everyday Indian speech — spanning ten domains from law and religion to food, history and media, drafted with 261 domain experts based across India.
OpenAI argued that benchmarks like MGSM and MMMLU mostly tested translated general knowledge or reasoning rather than reasoning that depended on local context, and built IndQA specifically to require the latter. Its announcement compared GPT-5 Thinking against other frontier systems and said performance had improved over successive model generations, but the company published no full numerical scores or public leaderboard alongside the release.
The benchmark arrived as OpenAI and its competitors were investing heavily in the Indian market through 2025, and it fits a broader pattern among frontier labs of building region-specific evaluations rather than relying solely on English-centric or translated tests as they compete for users outside English-speaking countries. As of this writing, IndQA has no independent public leaderboard or third-party dataset release to check OpenAI’s own reporting against.
The set
2,278 questions across 12 languages, including Hindi, Bengali, Tamil, Telugu, Punjabi and Hinglish (code-switched Hindi-English), drafted with 261 India-based domain experts across 10 domains including law, religion, food, history and media.
Example
From the Food and cuisine domain (Bengali): 'কোন পরিপ্রেক্ষিতে উনিশ শতকের শেষ দিক থেকে রান্নার বইগুলো বেরচ্ছিল? প্রথম বাংলা রান্নার বইটির সাথে বিপ্রদাস মুখোপাধ্যায় রচিত বইটির পার্থক্য কোথায়? বিপ্রদাসের উদ্যোগে প্রকাশিত পত্রিকাটি চলেছিল কতদিন? বিপ্রদাস ও প্রজ্ঞা সুন্দরীর লেখা অনুসরণ করে দিঘাপতিয়া থেকে কোন বইটি বেরিয়েছিল?' — English translation: 'In what context were cookbooks published from the end of the 19th century? What is the difference between the first Bengali cookbook and the book written by Bipradas Mukherjee? How long did the magazine published by Bipradas run? Which book was published by Dighapatiya following the writings of Bipradas and Pragya Sundari?'openai.com
Where it stands
Released without published numerical scores; OpenAI compared GPT-5 Thinking against other frontier systems qualitatively and said performance had improved over time but left substantial room for error.