Skip to content
The blog

Rovo Agent Evaluations Are Now Available in Rovo Studio

See how Atlassian's new Rovo Agent Evals let you test agent accuracy, and validate behavior at scale before deploying to your organization.

contents
  1. What Rovo Agent Evaluations Does
  2. Why This Matters for Enterprise Atlassian Environments
  3. How to Get Started With Rovo Agent Evaluations
  4. Tips for building an effective test set
  5. What Is Coming Next
  6. The Role of Atlas Bench

Building a Rovo Agent is the easy part. Knowing whether it actually works the way you intended, consistently, across hundreds of different user queries, is a harder problem. An agent that performs well in a handful of manual tests can still fail in ways that are difficult to predict when it goes live across an entire organization. Incomplete answers, misinterpreted questions, responses that technically address the query but miss the point entirely. These failures erode user trust fast and slow down AI adoption at exactly the moment organizations are trying to accelerate it.

Atlassian has released Rovo Agent Evaluations, a dedicated evaluation workspace inside Rovo Studio that gives agent builders a systematic way to test, measure, and improve agent quality before deployment. For IT admins, Atlassian administrators, and anyone responsible for the quality of AI in their organization's Atlassian environment, this is one of the most practically useful governance tools Atlassian has released for Rovo to date.

Rovo Evaluation interface showing two completed evaluations, with scores, simulated conversations, and run times displayed in a results table

What Rovo Agent Evaluations Does

Rovo Agent Evaluations sits inside Rovo Studio as a dedicated tab on any published agent. It gives builders three distinct ways to validate how their agent behaves, each suited to a different testing scenario and a different level of preparation required.

1. Response Accuracy: Test Against the Answers You Expect

Response Accuracy is the most rigorous evaluation method and the right starting point for agents handling critical, high-stakes workflows. It works by comparing your agent's actual responses against a set of reference responses you define in advance.

To use it, you upload a CSV containing questions and their ideal responses. The evaluation runs your agent against every question in that set and uses an LLM to compare the agent's responses to your reference answers. Each response receives a pass or fail judgment along with qualitative feedback explaining where and why the response diverged from the expected answer.

This method is particularly valuable for agents handling HR policies, IT support FAQs, product knowledge bases, onboarding guidance, and internal process documentation where accuracy is not a matter of opinion. It moves organizations from a position of believing their agent is working to having documented evidence that it passes a defined percentage of reference tests. For organizations with governance or compliance requirements around AI deployment, that documentation is meaningful.

2. Resolution Rate: Score Whether a Request Was Actually Resolved

Not every agent has a defined set of correct answers. For many service and question-and-answer agents, the more important question is simply whether the user's request was resolved, regardless of the exact wording of the response.

Resolution Rate addresses this by running your agent against a set of questions without requiring you to provide expected answers. You upload a CSV of questions only, the agent responds as usual, and an LLM scores each interaction as Resolved or Unresolved based on how well the response addressed the question.

This method is ideal for agents where resolution quality matters more than response precision. It gives organizations fast, actionable signal about where their agent is succeeding and where it is falling short, without the overhead of building and maintaining a full reference answer library.

3. Manual Testing: Review Behavior Across a Wide Range of Scenarios

Manual Testing is the most flexible evaluation method and the most useful for exploratory testing and sanity checking after instruction changes. You upload a CSV of questions with no expected answers required, run them all at once, and review the generated responses in a single consolidated view.

This method is particularly well suited to catching unexpected behavior after updating agent instructions, exploring how the agent handles edge cases, and spotting patterns of strength or weakness across a broad range of query types. It is not designed to produce a definitive quality score but rather to give builders a fast, comprehensive look at how the agent behaves across a diverse set of scenarios before making changes permanent.

Why This Matters for Enterprise Atlassian Environments

Before Rovo Agent Evaluations, testing an agent before deployment meant manually chatting with it and hoping you covered enough ground to catch the obvious failures. For agents handling a narrow, well-defined set of use cases, that approach was manageable. For agents deployed at scale across large organizations handling hundreds of different query types, it was not sufficient.

The evaluation workspace changes this in three meaningful ways:

1. Testing becomes systematic and repeatable.

Rather than an afterthought, evaluation is now a standard part of the agent development workflow. Organizations can build test sets that reflect the actual queries their agents will receive in production and run those same tests after every significant change to agent instructions or knowledge sources.

2. Quality can be tracked over time.

Instead of a one-off check before launch, teams can monitor how agent quality trends as the environment evolves. Instruction changes, new knowledge sources, and updated scenarios can all be validated against the same test set to confirm improvements and catch regressions.

3. Real-world failures become future test cases.

When a user reports that an agent gave a wrong or incomplete answer, that query can be added to the test set along with the correct expected response. The test set grows more comprehensive over time and becomes a living record of every known failure mode the organization has encountered and addressed.

How to Get Started With Rovo Agent Evaluations

Getting started requires no additional setup beyond having a published agent in Rovo Studio.

Step 1: Open the Evaluations tab.

Go to your agent in Rovo Studio and look for the Evaluations tab on the published agent view.

Step 2 : Prepare your test set as a CSV file.

For Response Accuracy testing your CSV needs two columns: the question and the ideal response. For Resolution Rate and Manual Testing a single column of questions is all you need.

Step 3 : Upload, select, and run.

Upload your CSV, select the evaluation type, and kick off the run. Results appear in a consolidated view showing pass and fail judgments for Response Accuracy, resolved and unresolved scores for Resolution Rate, and generated responses for Manual Testing.

Tips for building an effective test set

  • Mix critical queries, common queries, and edge cases so scores reflect the full range of real usage
  • Keep prompts short and unambiguous with any needed context included in the prompt itself
  • Treat the test set as a living document and fold real-world failures back in as new test cases to prevent regressions

What Is Coming Next

Atlassian has indicated that dataset versioning and live conversation review are both in active development. Dataset versioning will allow teams to track how test sets evolve over time and compare evaluation results across different versions of both the test set and the agent. Live conversation review will extend the evaluation capability to real user interactions happening through the JSM portal and other surfaces, giving organizations visibility into how their deployed agents are actually performing with real users rather than synthetic test queries.

Both capabilities will make the evaluation workspace significantly more powerful for organizations managing multiple agents across complex environments.

The Role of Atlas Bench

Building effective test sets, interpreting evaluation results, and using those results to systematically improve agent quality requires both Atlassian platform expertise and a clear understanding of what good agent behavior looks like in your specific organizational context. Getting this right from the start means faster time to value and fewer quality issues reaching users in production.

As an Atlassian Platinum Solution Partner and World Class Software Development Finalist, 2024-2025, Atlas Bench helps organizations implement, govern, and continuously improve Rovo Agents across the full Atlassian platform. We help agent builders design effective evaluation frameworks, configure Rovo Studio governance processes, and ensure that the agents your organization deploys are tested, trusted, and ready to scale.

Interested in learning more? Contact Atlas Bench to book a complimentary consultation with our Certified Atlassian Experts.

Try Rovo Today and Transform the Way You Work

rovo logo

Atlas Bench is an Atlassian Platinum Solution Partner with deep Cloud Specialization, helping Data Center customers move beyond migration and fully transform on the Atlassian Cloud Platform. We deliver Atlassian’s System of Work end to end through four collections: Strategy, Teamwork, Service, and Software, connecting planning to execution, aligning teams around outcomes, modernizing service delivery, and accelerating software development.

We are also recognized as a World Class Software Development Finalist, 2024-2025, and we bring AI leadership to every engagement. By leveraging Rovo for automation and purpose built agents, we reduce manual work, unlock knowledge across tools, and speed up decision making so organizations realize measurable value faster in cloud.

The blog, weekly.

One email a week with what we published. No drip sequence, and you can leave in a click.

Get an agent readiness assessment

Fixed scope. You get a findings report across identity, platform, and governance, an ownership gap analysis, and a sequenced plan for closing it.