With the arrival of generative AI, evaluations (“evals”) have become necessary to protect businesses from potentially negative outcomes stemming from the probabilistic nature of large language models (LLMs). Without proper AI evals, companies risk customer churn, legal liability, and failed launches. AI evaluations measure performance of an AI model and its output. It aids in alerting companies when that performance degrades. In this era where AI-powered products and features ship frequently, poor quality AI output in your product can mean the difference between being the market leader in your category for an AI feature or product, or falling behind.

Claude Code's Cautionary AI Evals Tale

Between August and September 2025, several people complained on X about the quality of Claude Code, an AI coding assistant from Anthropic that could have been mitigated much earlier, reducing damage, if more thorough evals were in place. The complaints weren't just from power users. According to Anthropic in a September 2025 public blog post, they had three separate bugs that reduced Claude Code's performance. They had trouble identifying these issues as three distinct problems. The trouble in identifying these three bugs delayed the fixes, which caused multiple users to switch to other products. Some people switched to OpenAI's Codex, Cursor, or other AI coding assistants.

This week, I switched from Claude Code to OpenAI Codex. I have no idea what happened to Claude Code over the last two weeks, but as of now, Codex is producing better quality code more regularly than it.
Mike Endale, co-founder and vice president of BLEN, on X, September 19, 2025

Customer churn isn't good, and it's especially felt in the world of AI products and features where one company can feel like a clear winner today, and then seem behind months later.

Anthropic cited that more thorough evals would have helped them find, understand, and resolve the issues faster. As a part of their response, they stated in a blog article that they would be improving their evals. The goal is to make it easier for them to identify AI output issues that fall outside of the expected quality and guardrails much faster in the future, hopefully preventing users from losing trust and leaving the product.

AI Evals Impact Business and Product Success

While evals are a standard for many companies, including Anthropic, the amount of effort spent in evals is uneven across companies, and some companies even go without them. Insufficient or incomplete evals pose several risks to the business. When companies break customer trust, they risk:

  • Slower Acquisition: If enough word of mouth or bad press around a product's reliability and quality get around, prospective customers may be reluctant to try yet another AI-powered product when the market is flooded with them.
  • Adoption Doesn't Stick: Customers may still sign up for a product, especially if it's from a trusted brand like Anthropic. But if the product doesn't work as promised, customers never get the full value. New customers can be quick to try other tools that align with their expectations.
  • Harming Conversion: Companies that have a freemium model may kill their opportunity to convert customers who think the free version is fine, but not “good enough” to pay for. This is especially true for premium priced products. Customers will expect consistent outcomes and delightful experiences for a premium price tag.
  • Creating Churn: When a product or feature doesn't fulfill its promise, it creates churn. There are a lot of AI alternatives for nearly every AI-powered product or feature for customers to choose from. Once customers lose faith in your product, it is hard to bring them back if they find something that solves their problem better. In the case of SaaS products, in a year, there likely will be even more alternatives for them to switch to when their contract is up.
  • Legal or Financial Impact: In 2022, an Air Canada chatbot hallucinated an incorrect answer to a customer's question about applying a bereavement discount to an airline ticket. Air Canada argued it was not responsible for what its chatbot said, however, the courts in Canada ruled in favor of the customer. Air Canada had to pay the customer damages to cover the extra cost incurred of being unable to apply the discount, despite what their chatbot said. Evals can reduce the possibility of outputs that result in legal or financial impact to the business.

Anthropic is currently third in the AI coding assistant market, represents 17.4% of the market, according to a September 2025 CB Insights report. Having a couple months where users publicly complain about the product on platforms like X and Reddit, and even switch from Claude Code to other AI coding assistants, can create setbacks in increasing marketshare. This is why evals are critical to supporting the business success of AI products and features.

Implement AI Evals Early

It's important to catch potential issues early to prevent issues impacting users, or slowdowns in the development process.

When we built an agent for our own platform, the golden dataset plus internal dogfooding surfaced issues long before rollout. These evals and datapoints gave us evidence to fix logic checks and tone guidance early, preventing thousands of bad customer interactions.
Aman Khan, Head of Product at Arize

Leverage Multiple Types of AI Evals

Make sure the evals are robust enough to make it clear what has gone wrong, and catch it as quickly as possible. This means leveraging multiple methods of evals. This makes it easier to tell when certain dimensions of AI output have shifted outside the expected level of quality. This can range from evaluating:

  • Accuracy: Assessing the output for correctness and relevance.
  • Task Completion and Usefulness: Determining if the user's expectations for the completion of the task were met.
  • Style and Tone: Assessing that the style and tone of the output is consistent.
  • Safety, Compliance and Bias: Validating that the output isn't creating unsafe output, violating specific compliance guidelines, especially for highly regulated industries, and sometimes bias.

Strong AI evals are not a luxury for “more mature” companies and products. They are a critical part of protecting your business and your customer experience. Evals are how some of the best AI products spot hidden issues before they become public failures. As AI becomes more embedded in business and work, discipline around creating good evals will become a silent champion of AI products that consistently work well and drive value.

Related

Working on something this touches? Let's talk about it.

Get in Touch