How Braintrust CEO uses AI evals to kill the standard PRD
This article summarizes an episode of Aakash Gupta’s video series featuring Ankur Goyal, CEO of Braintrust.

Ankur Goyal, CEO of Braintrust / Photo credit: wearefoundermode
Building AI requires a completely different approach to product management. Braintrust, a platform designed to help companies test and evaluate AI models, is leading a shift away from traditional planning methods. Founder and CEO Ankur Goyal warns that using vague product requirements documents (PRDs) to guide AI development is an operational risk.
Because AI behaves unpredictably, Goyal believes that standard written descriptions must be replaced with measurable tests. By turning abstract ideas into executable code, product teams can mathematically prove whether their software actually does what users need it to do, completely replacing the traditional PRD.
Turning product plans into tests
The rise of AI has exposed a flaw in how digital products are traditionally managed. For years, the product requirement document was the standard for building software, acting as a written guide for what a product should do. However, because it relies on human language, it inevitably causes costly confusion.
Goyal explains that the standard product plan has evolved from a subjective guide into a mathematical proof:
- The past: An unstructured written document explaining how to build a feature, leaving too much room for interpretation.
- The present: A measurable test (an eval) that proves if the software actually works, allowing engineers to build effective tools for industries they don’t personally work in.
Creating a lasting competitive advantage
While establishing clear tests solves the immediate communication problem, building an AI product carries an entirely new long-term risk. Because the underlying technology advances so rapidly, the specific AI model a company relies on today will likely be outdated in just a few months.
Many companies mistakenly believe that their unique combination of specific models and clever prompts gives them a permanent competitive edge. Focusing exclusively on this method creates a fragile business. The true lasting value lies in the testing infrastructure itself.
Goyal states, “If you believe that the way that you’ve wired together your agent today is your differentiator, you’re actually highly likely to fail because that’s probably going to change in a couple months. On the other hand, if you build really good evals, then you’ve built something that has a little bit more durability to it.”
Using simple checks for initial feedback
To actually build that durable testing infrastructure, engineering teams need to understand exactly where the testing process begins. The evaluation pipeline almost always starts with simple, gut-feeling checks that many technical leaders mistakenly ignore because they do not seem highly scientific.
While this initial gut check is often dismissed as an unimportant preliminary step, Goyal argues it is the foundation of the entire testing process.
Goyal notes, “I actually think vibe checks are a form of eval. When you do a vibe check, you are using your AI product and then using a scoring function which is your brain to try to intuit whether the result is good or bad. And if it’s not very good, then you might tweak the prompt.”
Adapting tests based on the end user
This early testing works exceptionally well when engineering teams are building tools designed for themselves, as the feedback loop is instantaneous and highly accurate. However, this simple method fails when developers attempt to build software for industries they do not personally work in.
Setting up a strict testing baseline
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.







