Loading
Prepare for Email to Become the Default Login ExperienceRead More
Agentforce and Einstein Generative AI
Create Custom Scorers

Create Custom Scorers

Custom scorers help you test your AI against specific criteria like accuracy, tone, and brand voice. By using an LLM-as-judge, you can automatically score and review outputs to make sure that your agents' responses consistently meet your specific goals and quality standards.

Required Editions

green checkmark

This article applies to:

New Testing Center in Agentforce Studio (Beta)
red crossmark

This article doesn’t apply to:

Legacy Agentforce Testing Center in Setup
Available in: Lightning Experience
Available in: Enterprise, Performance, Unlimited, and Developer Editions. Required add-on licenses vary by agent type.
User Permissions Needed
To create tests in the Testing Center:

Manage Agentforce Grids

AND

Manage Agentforce Testing

Every agent has a purpose or goal, so evaluating an agent's performance requires identifying metrics tailored to that purpose. Along with default scorers, custom scorers let you define specific criteria to assess your agent's effectiveness beyond simple pass or fail checks. With custom scorers you can check that your AI agents consistently reflect your brand voice, meet quality expectations, and deliver the intended sentiment. Or, use custom scorers when you expect a structured output like a JSON, and a text example in the Expected Response field isn't sufficient to capture your validation criteria.

  1. To add or review your scorers from the test suite, click Select Scorers.
  2. Click Add and select LLM Judge.

    What is an LLM judge?

    An LLM judge (or LLM-as-judge) is when one large language model (LLM) evaluates and scores the outputs of another. A judge LLM receives a prompt that defines the task and outlines the scoring criteria like factual accuracy, relevance, coherence, and faithfulness to source. With these resources and guidelines, the LLM judge determines the expected response and compares it to the agent response. Based on this comparison, the judge generates scores, rankings, or written feedback. The LLM as judge model serves as a scalable, automated, and objective evaluation tool for tasks like scoring summaries or ranking responses. We’ve carefully designed our LLM-as-judge prompts to give you the most accurate and useful test results.
  3. When you create a prompt for a custom scorer, you can tailor several key elements.
    • Determine which AI model serves as the judge
    • Add in Salesforce resources
    • Save multiple template versions
    • Set the threshold score for your passing criteria
  4. Save your scorer.

After you save your custom scorer, it's automatically selected for your test suite. Once Agentforce finishes generating tests, they appear under Test Cases. Depending on the size and complexity of the request, this process can take anywhere from a few minutes to several hours. You may need to refresh the testing suite to see your test cases.

When the test cases appear, click Run Test. Agentforce begins running the chosen scorers on the test cases.

 
Loading
Salesforce Help | Article