🧔♂️ A friendly human may check it before it goes live. More news here
HongShan launches new benchmark tool for AI
HongShan Capital Group (HSG, formerly Sequoia China) has launched xbench, a new benchmarking tool to evaluate the practical utility of AI.
After two years of development and testing, the platform is now available to the AI community, along with a research paper explaining its methodology.
Unlike traditional benchmarks, xbench aims to map AI performance against measurable business outcomes, featuring test sets that adapt to ongoing advancements in AI.
The public release includes two evaluations: xbench-ScienceQA, which measures academic knowledge and reasoning, and xbench-DeepSearch, which focuses on information gathering in Chinese-language environments. These evaluations are updated monthly and refreshed quarterly.
HSG plans to expand Profession Aligned evaluations to sectors such as finance, law, and sales. The tool also incorporates item response theory (IRT) to track improvements over time, providing a framework for measuring AI progress and market alignment.
🔗 Source: HSG
🧠 Food for thought
1️⃣ The benchmarking obsolescence cycle plagues AI evaluation
The rapid obsolescence of AI benchmarks has been a persistent industry challenge that xbench specifically aims to address with its “evergreen evaluation” approach.
This problem is widespread. BetterBench’s assessment of 24 AI benchmarks found significant quality variations and widespread issues, including lack of statistical significance reporting and insufficient maintenance processes 1.
The cycle typically unfolds predictably: new benchmarks are created, AI models quickly master them, and the benchmarks become obsolete. This aligns with xbench’s experience of having to replace test suites within months as models maxed out scores.
This pattern mirrors earlier technology evaluation challenges seen in computer vision benchmarks, where evaluation methods required constant updating as deep learning capabilities rapidly advanced after 2010 2.
2️⃣ AI evaluation is shifting from theoretical to practical, real-world metrics
xbench’s dual-track system reflects a broader industry transition from theoretical capabilities testing toward measuring real-world, professional utility—a shift appearing across multiple evaluation frameworks.
The Berkeley Function-Calling Leaderboard (BFCL) exemplifies this trend, having evolved through multiple iterations to assess increasingly complex real-world scenarios rather than abstract capabilities 3.
Similarly, Ï-bench focuses specifically on real-world interactions in dynamic environments with domain-specific policies, moving beyond simplified test cases to evaluate how agents perform in complex situations with practical constraints 3.
Recent Sequoia developments
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




