Building a Rovo Agent is the easy part. Knowing whether it actually works the way you intended, consistently, across hundreds of different user queries, is a harder problem. An agent that performs well in a handful of manual tests can still fail in ways that are difficult to predict when it goes live across an entire organization. Incomplete answers, misinterpreted questions, responses that technically address the query but miss the point entirely. These failures erode user trust fast and slow down AI adoption at exactly the moment organizations are trying to accelerate it.
Atlassian has released Rovo Agent Evaluations, a dedicated evaluation workspace inside Rovo Studio that gives agent builders a systematic way to test, measure, and improve agent quality before deployment. For IT admins, Atlassian administrators, and anyone responsible for the quality of AI in their organization's Atlassian environment, this is one of the most practically useful governance tools Atlassian has released for Rovo to date.
Rovo Agent Evaluations sits inside Rovo Studio as a dedicated tab on any published agent. It gives builders three distinct ways to validate how their agent behaves, each suited to a different testing scenario and a different level of preparation required.
Response Accuracy is the most rigorous evaluation method and the right starting point for agents handling critical, high-stakes workflows. It works by comparing your agent's actual responses against a set of reference responses you define in advance.
To use it, you upload a CSV containing questions and their ideal responses. The evaluation runs your agent against every question in that set and uses an LLM to compare the agent's responses to your reference answers. Each response receives a pass or fail judgment along with qualitative feedback explaining where and why the response diverged from the expected answer.
This method is particularly valuable for agents handling HR policies, IT support FAQs, product knowledge bases, onboarding guidance, and internal process documentation where accuracy is not a matter of opinion. It moves organizations from a position of believing their agent is working to having documented evidence that it passes a defined percentage of reference tests. For organizations with governance or compliance requirements around AI deployment, that documentation is meaningful.
Not every agent has a defined set of correct answers. For many service and question-and-answer agents, the more important question is simply whether the user's request was resolved, regardless of the exact wording of the response.
Resolution Rate addresses this by running your agent against a set of questions without requiring you to provide expected answers. You upload a CSV of questions only, the agent responds as usual, and an LLM scores each interaction as Resolved or Unresolved based on how well the response addressed the question.
This method is ideal for agents where resolution quality matters more than response precision. It gives organizations fast, actionable signal about where their agent is succeeding and where it is falling short, without the overhead of building and maintaining a full reference answer library.
Manual Testing is the most flexible evaluation method and the most useful for exploratory testing and sanity checking after instruction changes. You upload a CSV of questions with no expected answers required, run them all at once, and review the generated responses in a single consolidated view.
This method is particularly well suited to catching unexpected behavior after updating agent instructions, exploring how the agent handles edge cases, and spotting patterns of strength or weakness across a broad range of query types. It is not designed to produce a definitive quality score but rather to give builders a fast, comprehensive look at how the agent behaves across a diverse set of scenarios before making changes permanent.
Before Rovo Agent Evaluations, testing an agent before deployment meant manually chatting with it and hoping you covered enough ground to catch the obvious failures. For agents handling a narrow, well-defined set of use cases, that approach was manageable. For agents deployed at scale across large organizations handling hundreds of different query types, it was not sufficient.
The evaluation workspace changes this in three meaningful ways:
Rather than an afterthought, evaluation is now a standard part of the agent development workflow. Organizations can build test sets that reflect the actual queries their agents will receive in production and run those same tests after every significant change to agent instructions or knowledge sources.
Instead of a one-off check before launch, teams can monitor how agent quality trends as the environment evolves. Instruction changes, new knowledge sources, and updated scenarios can all be validated against the same test set to confirm improvements and catch regressions.
When a user reports that an agent gave a wrong or incomplete answer, that query can be added to the test set along with the correct expected response. The test set grows more comprehensive over time and becomes a living record of every known failure mode the organization has encountered and addressed.
Getting started requires no additional setup beyond having a published agent in Rovo Studio.
Go to your agent in Rovo Studio and look for the Evaluations tab on the published agent view.
For Response Accuracy testing your CSV needs two columns: the question and the ideal response. For Resolution Rate and Manual Testing a single column of questions is all you need.
Upload your CSV, select the evaluation type, and kick off the run. Results appear in a consolidated view showing pass and fail judgments for Response Accuracy, resolved and unresolved scores for Resolution Rate, and generated responses for Manual Testing.
Atlassian has indicated that dataset versioning and live conversation review are both in active development. Dataset versioning will allow teams to track how test sets evolve over time and compare evaluation results across different versions of both the test set and the agent. Live conversation review will extend the evaluation capability to real user interactions happening through the JSM portal and other surfaces, giving organizations visibility into how their deployed agents are actually performing with real users rather than synthetic test queries.
Both capabilities will make the evaluation workspace significantly more powerful for organizations managing multiple agents across complex environments.
Building effective test sets, interpreting evaluation results, and using those results to systematically improve agent quality requires both Atlassian platform expertise and a clear understanding of what good agent behavior looks like in your specific organizational context. Getting this right from the start means faster time to value and fewer quality issues reaching users in production.
As an Atlassian Platinum Solution Partner and World Class Software Development Finalist, 2024-2025, Atlas Bench helps organizations implement, govern, and continuously improve Rovo Agents across the full Atlassian platform. We help agent builders design effective evaluation frameworks, configure Rovo Studio governance processes, and ensure that the agents your organization deploys are tested, trusted, and ready to scale.
Interested in learning more? Contact Atlas Bench to book a complimentary consultation with our Certified Atlassian Experts.
Try Rovo Today and Transform the Way You Work
Atlas Bench is an Atlassian Platinum Solution Partner with deep Cloud Specialization, helping Data Center customers move beyond migration and fully transform on the Atlassian Cloud Platform. We deliver Atlassian’s System of Work end to end through four collections: Strategy, Teamwork, Service, and Software, connecting planning to execution, aligning teams around outcomes, modernizing service delivery, and accelerating software development.
We are also recognized as a World Class Software Development Finalist, 2024-2025, and we bring AI leadership to every engagement. By leveraging Rovo for automation and purpose built agents, we reduce manual work, unlock knowledge across tools, and speed up decision making so organizations realize measurable value faster in cloud.