AI Agent Evaluations: A QA Blueprint
Table of Contents
1. Introduction & Core Philosophy
Evaluations (“evals”) might sound like another complex paradigm you need to master, but if you have Quality Assurance (QA) experience, you already know 80% of the discipline.
Simply put, evals are validating the output of an agent. An agent at the end of the day is a computer program that runs and generates outputs. An eval is deciding what the quality of that output is. The non-deterministic nature of LLMs/AI agents is not a barrier that is new. Non-determinism has always been a factor that QA has had to deal with, such as UI, UX, or performance of something. You can get hard metrics to measure the load time of a page, but if it doesn’t “feel” right then there is a problem, and that problem can be logged.
And since evals are so close to traditional testing, it’s best to still try to conform to the industry jargon of today, rather than trying to make these new concepts sound exotic. The most egregious examples of this are the terms offline and online evals.
Put simply, online evals are when you test on user data in production. Offline evals are when you use a static dataset to run your tests on. We already have terminology for this. Online evals are ad hoc testing, and offline evals are regression suites. By inserting the online/offline dynamic, it only confuses people, especially your business stakeholders. We’ve spent thousands of years refining language to convey ideas simply and quickly, and there’s no need to buck that trend.
Admittedly, “online” evals are tricky to define since it’s a blend of production testing and shadow testing. However, if you use the term “online”, that already has connotations with being connected to a physical network. But since agents are already connected to a network, it only causes confusion.
Since AI agents are so new, the testing procedures will be advancing rapidly. You can already see this with the discourse about AI evaluations. Most content out there will focus on how to evaluate user prompts, which is what most chatbots are doing. And this is ok, but as agents are given more responsibility and deal with actual systems, using user prompts as an eval guide starts to become weak. Instead, the focus needs to be on the actual agent output itself.
2. The Strategic Approach: Planning & Independence
Before you can even start to measure the agent quality, you have to decide what the quality should look like. Start with some basic metrics that can be easily explained or conceptualised.
As an example, the number of turns taken to get to a given outcome. Or, when given a set of instructions, the agent correctly returns 75% of the items.
This will require an understanding of what data you actually intend to generate, and this can change over the course of a project. This means that the evals will need to change alongside the requirements. Rigidly sticking to any success metrics you define will not get you the quality outcomes you are looking for.
{
"testcase_name": "sample_test",
"tags": ["ci","deterministic"],
"description": "useful description here",
"turns": [
{"user_message": "process my file"},
{"file": "report.pdf"},
{"user_message": "add some annotations"}
],
"expected": {
"tags": ["pdf","has_title", "has_annotations"],
"status": "completed"
}
}
To help adapt to that change, the test/eval team needs to be semi-independent, just like a traditional testing team from the development team. Crossover and communication are encouraged, but the test team needs to have the capability to define what success or quality is and stick by it. The quality conversation may involve the product team since they will be the end users and most likely signing off.
3. Client Dynamics & Human Feedback
For agents that are being built to solve a business process problem, understanding the needs of the business becomes even more important. This has always been a factor, most notably during integration testing. But with evals, it raises the stakes even higher, since the eval team needs to be able to understand the agent output and its quality. It’s no longer possible to simply rely on business requirements handed down and a simple box-checking exercise.
The eval team needs to be part of the discovery process and understand what the business is looking for. Then the eval team can write the quality expectations ahead of time, similar to Test Driven Development.
When designing the feedback mechanism, it’s best to avoid too many options for the users to select (paralysis by analysis). Use a familiar 1 to 5 star scale where a 5 means “no issues” rather than “incredible experience”. Setting this standard early manages stakeholder expectations around non-deterministic AI.
1 = 0-20% of expected data is pulled across
2 = 40% of expected data is pulled across
3 = 60% of expected data is pulled across
4 = 80% of expected data is pulled across
5 = 100% of expected data is pulled across
4. Scope & Test Design Philosophy
Prioritise code-based asserts such as regex checks, substring matches, etc. Using LLM evaluators will add cost and time which need to be managed carefully as the test suite grows. There are also throughput considerations, as some providers such as Gemini/Google have a limit of 5 LLM requests per project (this can be raised by a direct request).
Example of LLM as a judge
Only respond with these enums: accepted, cannot-process, not-accepted.
If the conversation only contains INIT then that is accepted.
If the user indicates that he/she wants to make major changes that would be a not-accepted. The language can be like: completely change this, reword everything, update from scratch, delete all of it, etc.
If the user wants to update/reword a small or moderate amount, then that is considered accepted. The language can be like: reword this, make the language more X, tweak this, swap these sections, update X section to be Y etc.
If the data being presented cannot be parsed, or there is no data then that is cannot-process.
Finally if there doesn't appear to be any edits made, that is accepted.
As it is the simplest to set up task advancement should be the first priority, where you can set up a short series of turns 3-4 long, then ensure that the agent arrives at the expected outcome. Having this in place early will help with regression down the line as more system instructions get added in. It requires only a handful of test cases to catch annoying regressions, but it does require monitoring. If there are any process changes such as adding in a mandatory step, then the evals will need to be updated.
The most important part of the evals is the initial test data that gets set up. There are some key ideas to help achieve this:
- Setup ideal input data from real-world applications.
- Talk with the users to understand what they look for in the output.
- Keep turns short to minimise noise from agents.
With these core principles in place, the eval writing is straightforward and any frontier model will easily be able to fill in the gaps. When the LLMs have clearly defined data, it is very easy for them to generate reasonable tests.
5. Technical Execution & Automation
Handling the lifecycle of evals is no different to a test suite. As the system evolves, maintenance will be required to keep accurate results coming in. If the test data is clear, then this maintenance can be easily handled by LLMs.
When using LLMs, it is necessary to keep a close eye on them, since they tend to overcomplicate both the tests and the eval framework. “Read the code” is starting to become a more common saying, and it doubly applies here. The LLMs tend to want to add extra plugins or add in complexity, and it’s up to the eval engineer to ensure that the evals are usable both for machines and for people who need to maintain them.