Autoevals is a tool to quickly and easily evaluate AI model outputs.
It bundles together a variety of automatic evaluation methods including:
- LLM-as-a-judge
- Heuristic (e.g. Levenshtein distance)
- Statistical (e.g. BLEU)
Autoevals is developed by the team at Braintrust.
Autoevals uses model-graded evaluation for a variety of subjective tasks including fact checking, safety, and more. Many of these evaluations are adapted from OpenAI's excellent evals project but are implemented so you can flexibly run them on individual examples, tweak the prompts, and debug their outputs.
You can also create your own model-graded evaluations with Autoevals. It's easy to add custom prompts, parse outputs, and manage exceptions.
Use Autoevals to model-grade an example LLM completion using the Factuality prompt.
By default, Autoevals uses your OPENAI_API_KEY environment variable to authenticate with OpenAI's API.
fromautoevals.llmimport*importasyncio# Create a new LLM-based evaluatorevaluator=Factuality()
# Synchronous evaluationinput="Which country has the highest population?"output="People's Republic of China"expected="China"# Using the synchronous APIresult=evaluator(output, expected, input=input)
print(f"Factuality score (sync): {result.score}")
print(f"Factuality metadata (sync): {result.metadata['rationale']}")
# Using the asynchronous APIasyncdefmain():
result=awaitevaluator.eval_async(output, expected, input=input)
print(f"Factuality score (async): {result.score}")
print(f"Factuality metadata (async): {result.metadata['rationale']}")
# Run the async exampleasyncio.run(main())import{Factuality}from"autoevals";(async()=>{constinput="Which country has the highest population?";constoutput="People's Republic of China";constexpected="China";constresult=awaitFactuality({ output, expected, input });console.log(`Factuality score: ${result.score}`);console.log(`Factuality metadata: ${result.metadata?.rationale}`);})();When you use Autoevals, it will look for an OPENAI_BASE_URL environment variable to use as the base for requests to an OpenAI-compatible API. If OPENAI_BASE_URL is not set, it will look for a BRAINTRUST_AI_GATEWAY_URL environment variable and then default to the Braintrust Gateway.
When you use the Braintrust Gateway, you'll also get:
- Simplified access to many AI providers
- Reduced costs with automatic request caching
- Increased observability when you enable logging to Braintrust
The Braintrust-hosted Gateway is free to use while it is in beta.
Set the BRAINTRUST_API_KEY environment variable to authenticate Gateway requests. You can also route requests to supported AI providers and models or custom models you have configured in Braintrust.
# NOTE: ensure BRAINTRUST_API_KEY is set in your environmentfromautoevals.llmimport*# Create an LLM-based evaluator using the Claude 3.5 Sonnet model from Anthropicevaluator=Factuality(model="claude-3-5-sonnet-latest")
# Evaluate an example LLM completioninput="Which country has the highest population?"output="People's Republic of China"expected="China"result=evaluator(output, expected, input=input)
# The evaluator returns a score from [0,1] and includes the raw outputs from the evaluatorprint(f"Factuality score: {result.score}")
print(f"Factuality metadata: {result.metadata['rationale']}")// NOTE: ensure BRAINTRUST_API_KEY is set in your environmentimport{Factuality}from"autoevals";(async()=>{constinput="Which country has the highest population?";constoutput="People's Republic of China";constexpected="China";// Run an LLM-based evaluator using the Claude 3.5 Sonnet model from Anthropicconstresult=awaitFactuality({model: "claude-3-5-sonnet-latest",
output,
expected,
input,});// The evaluator returns a score from [0,1] and includes the raw outputs from the evaluatorconsole.log(`Factuality score: ${result.score}`);console.log(`Factuality metadata: ${result.metadata?.rationale}`);})();There are two ways you can configure a custom client when you need to use a different OpenAI compatible API:
- Global configuration: Initialize a client that will be used by all evaluators
- Instance configuration: Configure a client for a specific evaluator
Set up a client that all your evaluators will use:
importopenaiimportasynciofromautoevalsimportinitfromautoevals.llmimportFactualityclient=init(openai.AsyncOpenAI(base_url="https://api.openai.com/v1/"))
asyncdefmain():
evaluator=Factuality()
result=awaitevaluator.eval_async(
input="What is the speed of light in a vacuum?",
output="The speed of light in a vacuum is 299,792,458 meters per second.",
expected="The speed of light in a vacuum is approximately 300,000 kilometers per second."
)
print(f"Factuality score: {result.score}")
asyncio.run(main())importOpenAIfrom"openai";import{init,Factuality}from"autoevals";constclient=newOpenAI({baseURL: "https://api.openai.com/v1/",});init({ client });(async()=>{constresult=awaitFactuality({input: "What is the speed of light in a vacuum?",output: "The speed of light in a vacuum is 299,792,458 meters per second.",expected:
"The speed of light in a vacuum is approximately 300,000 kilometers per second (or precisely 299,792,458 meters per second).",});console.log("Factuality Score:",result);})();Configure a client for a specific evaluator instance:
importopenaifromautoevals.llmimportFactualitycustom_client=openai.OpenAI(base_url="https://custom-api.example.com/v1/")
evaluator=Factuality(client=custom_client)importOpenAIfrom"openai";import{Factuality}from"autoevals";(async()=>{constcustomClient=newOpenAI({baseURL: "https://custom-api.example.com/v1/",});constresult=awaitFactuality({client: customClient,output: "Paris is the capital of France",expected:
"Paris is the capital of France and has a population of over 2 million",input: "Tell me about Paris",});console.log(result);})();Once you grade an output using Autoevals, you can optionally use Braintrust to log and compare your evaluation results. This integration is completely optional and not required for using Autoevals.
Create a file named example.eval.js (it must take the form *.eval.[ts|tsx|js|jsx]):
import{Eval}from"braintrust";import{Factuality}from"autoevals";Eval("Autoevals",{data: ()=>[{input: "Which country has the highest population?",expected: "China",},],task: ()=>"People's Republic of China",scores: [Factuality],});Then, run
npx braintrust run example.eval.jsCreate a file named eval_example.py (it must take the form eval_*.py):
importbraintrustfromautoevals.llmimportFactualityEval(
"Autoevals",
data=lambda: [
dict(
input="Which country has the highest population?",
expected="China",
),
],
task=lambda*args: "People's Republic of China",
scores=[Factuality],
)- Battle
- Closed QA
- Humor
- Factuality
- Moderation
- Security
- Summarization
- SQL
- Translation
- Fine-tuned binary classifiers
- Context precision
- Context relevancy
- Context recall
- Context entity recall
- Faithfulness
- Answer relevancy
- Answer similarity
- Answer correctness
- Semantic list contains
- JSON validity
- Embedding similarity
- Levenshtein distance
- Exact match
- Numeric difference
- JSON diff
For detailed documentation on all scorers, including parameters, score ranges, and usage examples, see the Scorer Reference.
Autoevals supports custom evaluation prompts for model-graded evaluation. To use them, simply pass in a prompt and scoring mechanism:
fromautoevalsimportLLMClassifier# Define a prompt prefix for a LLMClassifier (returns just one answer)prompt_prefix="""You are a technical project manager who helps software engineers generate better titles for their GitHub issues.You will look at the issue description, and pick which of two titles better describes it.I'm going to provide you with the issue description, and two possible titles.Issue Description: {{input}}1: {{output}}2: {{expected}}"""# Define the scoring mechanism# 1 if the generated answer is better than the expected answer# 0 otherwiseoutput_scores= {"1": 1, "2": 0}
evaluator=LLMClassifier(
name="TitleQuality",
prompt_template=prompt_prefix,
choice_scores=output_scores,
use_cot=True,
)
# Evaluate an example LLM completionpage_content="""As suggested by Nicolo, we should standardize the error responses coming from GoTrue, postgres, and realtime (and any other/future APIs) so that it's better DX when writing a client,We can make this change on the servers themselves, but since postgrest and gotrue are fully/partially external may be harder to change, it might be an option to transform the errors within the client libraries/supabase-js, could be messy?Nicolo also dropped this as a reference: http://spec.openapis.org/oas/v3.0.3#openapi-specification"""output="Standardize error responses from GoTrue, Postgres, and Realtime APIs for better DX"expected="Standardize Error Responses across APIs"response=evaluator(output, expected, input=page_content)
print(f"Score: {response.score}")
print(f"Metadata: {response.metadata}")import{LLMClassifierFromTemplate}from"autoevals";(async()=>{constpromptTemplate=`You are a technical project manager who helps software engineers generate better titles for their GitHub issues.You will look at the issue description, and pick which of two titles better describes it.I'm going to provide you with the issue description, and two possible titles.Issue Description: {{input}}1: {{output}}2: {{expected}}`;constchoiceScores={1: 1,2: 0};constevaluator=LLMClassifierFromTemplate<{input: string}>({name: "TitleQuality",
promptTemplate,
choiceScores,useCoT: true,});constinput=`As suggested by Nicolo, we should standardize the error responses coming from GoTrue, postgres, and realtime (and any other/future APIs) so that it's better DX when writing a client,We can make this change on the servers themselves, but since postgrest and gotrue are fully/partially external may be harder to change, it might be an option to transform the errors within the client libraries/supabase-js, could be messy?Nicolo also dropped this as a reference: http://spec.openapis.org/oas/v3.0.3#openapi-specification`;constoutput=`Standardize error responses from GoTrue, Postgres, and Realtime APIs for better DX`;constexpected=`Standardize Error Responses across APIs`;constresponse=awaitevaluator({ input, output, expected });console.log("Score",response.score);console.log("Metadata",response.metadata);})();Every scorer returns a small Score result object. This is the public surface
consumers should read when they need to store, compare, or export evaluation
results:
name: the scorer namescore: a number between 0 and 1, orNone/nullwhen the evaluation is skippedmetadata: optional scorer-specific details, such as rationale text or a selected choice. Keys are scorer-specific; consumers should not assume metadata keys are shared across scorer types.error: deprecated and retained for backward compatibility; some scorers may still populate it, but callers should primarily handle thrown exceptions
Inputs, expected values, model prompts, and other runtime context are not part
of the Score object. Keep those separately if your application needs them.
You can also create your own scoring functions that do not use LLMs. For example, to test whether the word 'banana'
is in the output, you can use the following:
fromautoevalsimportScoredefbanana_scorer(output, expected, input):
returnScore(name="banana_scorer", score=1if"banana"inoutputelse0)
input="What is 1 banana + 2 bananas?"output="3"expected="3 bananas"result=banana_scorer(output, expected, input)
print(f"Banana score: {result.score}")import{Score}from"autoevals";constbananaScorer=({
output,
expected,
input,}: {output: string;expected: string;input: string;}): Score=>{return{name: "banana_scorer",score: output.includes("banana") ? 1 : 0};};(async()=>{constinput="What is 1 banana + 2 bananas?";constoutput="3";constexpected="3 bananas";constresult=bananaScorer({ output, expected, input });console.log(`Banana score: ${result.score}`);})();There is nothing particularly novel about the evaluation methods in this library. They are all well-known and well-documented. However, there are a few things that are particularly difficult when evaluating in practice:
- Normalizing metrics between 0 and 1 is tough. For example, check out the calculation in number.py to see how it's done for numeric differences.
- Parsing the outputs on model-graded evaluations is also challenging. There are frameworks that do this, but it's hard to debug one output at a time, propagate errors, and tweak the prompts. Autoevals makes these tasks easy.
- Collecting metrics behind a uniform interface makes it easy to swap out evaluation methods and compare them. Prior to Autoevals, we couldn't find an open source library where you can simply pass in
input,output, andexpectedvalues through a bunch of different evaluation methods.
The full docs are available for your reference.
We welcome contributions!
To install the development dependencies, run make develop, and run source env.sh to activate the environment. Make a .env file from the .env.example file and set the environment variables. Run direnv allow to load the environment variables.
To run the tests, run pytest from the root directory.
Send a PR and we'll review it! We'll take care of versioning and releasing.