Skip to content
Agent evaluations

How G2 evaluates AI agents

G2 tests AI agents by giving them real work, then publishes the results. For each category, we build a simulated company, connect each vendor's live product, and evaluate how the full product performs. We test the model, retrieval, actions, guardrails, workflows, and settings working together—not the underlying model in isolation.

What an evaluation measures

An evaluation measures whether a product can complete category-specific work in a controlled environment. It is designed to reflect the product a buyer would actually use.

A simulated company

Each category gets the policies, data, tools, and operating context an agent needs to do real work.

Buyer-informed tasks

Tasks come from buyer research and design partners, with synthetic cases added to cover important edge conditions.

A consistent comparison

Products in the same comparison group complete the same task set under the same methodology.

How an evaluation runs

Every product follows the same four-stage process within its category. Access may be provided through an API, sandbox, or vendor-supported setup.

  1. 01

    Set up the product

    We configure each product as a customer would, using its knowledge base, actions, workflows, and documented settings.

  2. 02

    Run consistent tasks

    Every agent in a category handles the same buyer-informed tasks, including realistic cases and meaningful edge cases.

  3. 03

    Capture the evidence

    Every task produces a complete transcript and a log of the actions the agent took in the simulated environment.

  4. 04

    Review before publishing

    Vendors can identify setup mistakes before publication. Valid issues are fixed and re-run, but vendors cannot change scores.

How scores work

The Overall Score rewards completed work. Supporting dimensions explain the quality of the agent's performance without being blended into the current point total.

Current Overall Score

Each current standard task earns five points when passed and zero when failed or missing. Future advanced and complex tasks can carry higher predefined values.

passed standard tasks × 5 points

LLM judge

Accuracy

Are the agent's claims supported by evidence?

LLM judge

Policy compliance

Did the agent follow the written policy?

LLM judge

Relevance

Did the agent stay focused on the user's problem?

Deterministic check

Completeness

Did the required outcomes happen in the underlying systems?

Where each signal comes from

Evaluation results, buyer reviews, and product information answer different questions. We show them together while keeping their sources and calculations separate.

G2 evaluation results

Measured performance

Scores produced by G2's controlled task runs. These are the results used in an evaluation leaderboard.

G2 ratings and reviews

Buyer feedback

Verified buyer sentiment shown separately. Review data never changes an agent's evaluation score.

Product facts and claims

Documented capabilities

Product details and vendor-reported claims are identified by source and remain separate from measured performance.

Category methodologies

The evaluation framework is consistent across G2, while the company environment, task set, and policies are specific to each category. This section will expand as new agent categories are evaluated.

Customer Experience agentsCurrent published category methodology+

The CX environment includes a written support policy, working tools for orders, refunds, subscriptions, and accounts, plus simulated customers with realistic problems. The approach builds on the open τ²-bench method and extends it to the products vendors ship.

46
assigned tasks
38
working business tools
4
scoring dimensions

Technical methodology

For readers who want to understand the evidence, evaluator types, score aggregation, and comparison rules behind the public results.

View technical detailsEvidence, evaluators, aggregation, and comparability+

Evaluation evidence

Task context, policy, the full agent-and-user trace, observable tool calls, and final-state evidence from the simulated environment.

Evaluator types

LLM judges apply defined rubrics to qualitative dimensions. Deterministic checks verify outcomes in the underlying systems.

Score aggregation

Current standard tasks award five Overall Score points when passed. Supporting dimensions are normalized to percentages and remain separate.

Comparable results

Agents share a comparison set only when the task set, rubric, judge, dataset, and methodology version match.

Limits and freshness

An evaluation is a dated snapshot, not a permanent label. Products and the methodology will continue to improve.

Results are refreshed

We aim to refresh evaluations several times a year, or sooner when a product changes materially.

The full task set stays private

Representative samples can be published while the full set remains private to reduce tuning to the test.

The benchmark will grow

Tasks will become longer and more complex, and repeated task runs are planned as the program matures.