How G2 evaluates AI agents
G2 tests AI agents by giving them real work, then publishes the results. For each category, we build a simulated company, connect each vendor's live product, and evaluate how the full product performs. We test the model, retrieval, actions, guardrails, workflows, and settings working together—not the underlying model in isolation.
What an evaluation measures
An evaluation measures whether a product can complete category-specific work in a controlled environment. It is designed to reflect the product a buyer would actually use.
A simulated company
Each category gets the policies, data, tools, and operating context an agent needs to do real work.
Buyer-informed tasks
Tasks come from buyer research and design partners, with synthetic cases added to cover important edge conditions.
A consistent comparison
Products in the same comparison group complete the same task set under the same methodology.
How an evaluation runs
Every product follows the same four-stage process within its category. Access may be provided through an API, sandbox, or vendor-supported setup.
- 01
Set up the product
We configure each product as a customer would, using its knowledge base, actions, workflows, and documented settings.
- 02
Run consistent tasks
Every agent in a category handles the same buyer-informed tasks, including realistic cases and meaningful edge cases.
- 03
Capture the evidence
Every task produces a complete transcript and a log of the actions the agent took in the simulated environment.
- 04
Review before publishing
Vendors can identify setup mistakes before publication. Valid issues are fixed and re-run, but vendors cannot change scores.
How scores work
The Overall Score rewards completed work. Supporting dimensions explain the quality of the agent's performance without being blended into the current point total.
Current Overall Score
Each current standard task earns five points when passed and zero when failed or missing. Future advanced and complex tasks can carry higher predefined values.
passed standard tasks × 5 points
Accuracy
Are the agent's claims supported by evidence?
Policy compliance
Did the agent follow the written policy?
Relevance
Did the agent stay focused on the user's problem?
Completeness
Did the required outcomes happen in the underlying systems?
Where each signal comes from
Evaluation results, buyer reviews, and product information answer different questions. We show them together while keeping their sources and calculations separate.
G2 evaluation results
Measured performance
Scores produced by G2's controlled task runs. These are the results used in an evaluation leaderboard.
G2 ratings and reviews
Buyer feedback
Verified buyer sentiment shown separately. Review data never changes an agent's evaluation score.
Product facts and claims
Documented capabilities
Product details and vendor-reported claims are identified by source and remain separate from measured performance.
Category methodologies
The evaluation framework is consistent across G2, while the company environment, task set, and policies are specific to each category. This section will expand as new agent categories are evaluated.
Customer Experience agentsCurrent published category methodology+
The CX environment includes a written support policy, working tools for orders, refunds, subscriptions, and accounts, plus simulated customers with realistic problems. The approach builds on the open τ²-bench method and extends it to the products vendors ship.
- 46
- assigned tasks
- 38
- working business tools
- 4
- scoring dimensions
Technical methodology
For readers who want to understand the evidence, evaluator types, score aggregation, and comparison rules behind the public results.
View technical detailsEvidence, evaluators, aggregation, and comparability+
Evaluation evidence
Task context, policy, the full agent-and-user trace, observable tool calls, and final-state evidence from the simulated environment.
Evaluator types
LLM judges apply defined rubrics to qualitative dimensions. Deterministic checks verify outcomes in the underlying systems.
Score aggregation
Current standard tasks award five Overall Score points when passed. Supporting dimensions are normalized to percentages and remain separate.
Comparable results
Agents share a comparison set only when the task set, rubric, judge, dataset, and methodology version match.
Limits and freshness
An evaluation is a dated snapshot, not a permanent label. Products and the methodology will continue to improve.
Results are refreshed
We aim to refresh evaluations several times a year, or sooner when a product changes materially.
The full task set stays private
Representative samples can be published while the full set remains private to reduce tuning to the test.
The benchmark will grow
Tasks will become longer and more complex, and repeated task runs are planned as the program matures.