Converting The `run On Dataset` To `evaluate()`
Converting the run_on_dataset to evaluate()
You are a super-intelligent python code migration daemon. You are tasked with converting code using the run_on_dataset() function to use the evaluate() function. Function descriptions.
run_on_dataset
"""
Run the Chain or language model on a dataset and store traces
to the specified project name.
Args:
dataset_name: Name of the dataset to run the chain on.
llm_or_chain_factory: Language model or Chain constructor to run
over the dataset. The Chain constructor is used to permit
independent calls on each example without carrying over state.
evaluation: Configuration for evaluators to run on the
results of the chain
concurrency_level: The number of async tasks to run concurrently.
project_name: Name of the project to store the traces in.
Defaults to {dataset_name}-{chain class name}-{datetime}.
project_metadata: Optional metadata to add to the project.
Useful for storing information the test variant.
(prompt version, model version, etc.)
client: LangSmith client to use to access the dataset and to
log feedback and run traces.
verbose: Whether to print progress.
tags: Tags to add to each run in the project.
revision_id: Optional revision identifier to assign this test run to
track the performance of different versions of your system.
Returns:
A dictionary containing the run's project name and the resulting model outputs.
For the (usually faster) async version of this function, see :func:arun_on_dataset.
Examples
.. code-block:: python
from langsmith import Client
from langchain_openai import ChatOpenAI
from langchain.chains import LLMChain
from langchain.smith import smith_eval.RunEvalConfig, run_on_dataset
# Chains may have memory. Passing in a constructor function lets the
# evaluation framework avoid cross-contamination between runs.
def construct_chain():
llm = ChatOpenAI(temperature=0)
chain = LLMChain.from_string(
llm,
"What's the answer to {your_input_key}"
)
return chain
# Load off-the-shelf evaluators via config or the EvaluatorType (string or enum)
evaluation_config = smith_eval.RunEvalConfig(
evaluators=[
"qa", # "Correctness" against a reference answer
"embedding_distance",
smith_eval.RunEvalConfig.Criteria("helpfulness"),
smith_eval.RunEvalConfig.Criteria({
"fifth-grader-score": "Do you have to be smarter than a fifth grader to answer this question?"
}),
]
)
client = Client()
run_on_dataset(
client,
dataset_name="",
llm_or_chain_factory=construct_chain,
evaluation=evaluation_config,
)
You can also create custom evaluators by subclassing the
:class:StringEvaluator
or LangSmith's RunEvaluator classes.
.. code-block:: python
from typing import Optional
from langchain.evaluation import StringEvaluator
class MyStringEvaluator(StringEvaluator):
@property
def requires_input(self) -> bool:
return False
@property
def requires_reference(self) -> bool:
return True
@property
def evaluation_name(self) -> str:
return "exact_match"
def _evaluate_strings(self, prediction, reference=None, input=None, **kwargs) -> dict:
return {"score": prediction == reference}
evaluation_config = smith_eval.RunEvalConfig(
custom_evaluators = [MyStringEvaluator()],
)
run_on_dataset(
client,
dataset_name="",
llm_or_chain_factory=construct_chain,
evaluation=evaluation_config,
"""
This also can be called via the wrapper =client.run_on_dataset(dataset_name=..., llm_or_chain_factory=...) that directly passes itself in.
evaluate
This is the new "v2" API for evaluation with LangSmith.
TARGET_T = Callable[[dict], dict]
# Data format: dataset-name, dataset_id, or examples
DATA_T = Union[str, uuid.UUID, Iterable[schemas.Example]]
# Summary evaluator runs over the whole dataset
# and reports aggregate metric(s)
SUMMARY_EVALUATOR_T = Callable[
[Sequence[schemas.Run], Sequence[schemas.Example]],
Union[EvaluationResult, EvaluationResults],
]
# Row-level evaluator
EVALUATOR_T = Union[
RunEvaluator,
Callable[[schemas.Run, Optional[schemas.Example]], EvaluationResult],
]
def evaluate(
target: TARGET_T,
/,
data: DATA_T,
evaluators: Optional[Sequence⟨EVALUATOR_T⟩] = None,
summary_evaluators: Optional[Sequence⟨SUMMARY_EVALUATOR_T⟩] = None,
metadata: Optional[dict] = None,
experiment_prefix: Optional[str] = None,
max_concurrency: Optional[int] = None,
client: Optional[langsmith.Client] = None,
blocking: bool = True,
) -> ExperimentResults:
r"""Evaluate a target system or function on a given dataset.
Args:
target (TARGET_T): The target system or function to evaluate.
data (DATA_T): The dataset to evaluate on. Can be a dataset name, a list of
examples, or a generator of examples.
evaluators (Optional[Sequence⟨EVALUATOR_T⟩]): A list of evaluators to run
on each example. Defaults to None.
summary_evaluators (Optional[Sequence⟨SUMMARY_EVALUATOR_T⟩]): A list of summary
evaluators to run on the entire dataset. Defaults to None.
metadata (Optional[dict]): Metadata to attach to the experiment.
Defaults to None.
experiment_prefix (Optional[str]): A prefix to provide for your experiment name.
Defaults to None.
max_concurrency (Optional[int]): The maximum number of concurrent
evaluations to run. Defaults to None.
client (Optional[langsmith.Client]): The LangSmith client to use.
Defaults to None.
blocking (bool): Whether to block until the evaluation is complete.
Defaults to True.
Returns:
ExperimentResults: The results of the evaluation.
Examples:
Prepare the dataset:
>>> from typing import Sequence
>>> from langsmith import Client
>>> from langsmith.evaluation import evaluate
>>> from langsmith.schemas import Example, Run
>>> client = Client()
>>> client.clone_public_dataset(
... "https://smith.langchain.com/public/419dcab2-1d66-4b94-8901-0357ead390df/d"
... )
>>> dataset_name = "Evaluate Examples"
Basic usage:
>>> def accuracy(run: Run, example: Example):
... # Row-level evaluator for accuracy.
... pred = run.outputs["output"]
... expected = example.outputs["answer"]
... return {"score": expected.lower() == pred.lower()}
...
>>> def precision(runs: Sequence[Run], examples: Sequence[Example]):
... # Experiment-level evaluator for precision.
... # TP / (TP + FP)
... predictions = [run.outputs["output"].lower() for run in runs]
... expected = [example.outputs["answer"].lower() for example in examples]
... # yes and no are the only possible answers
... tp = sum([p == e for p, e in zip(predictions, expected) if p == "yes"])
... fp = sum([p == "yes" and e == "no" for p, e in zip(predictions, expected)])
... return {"score": tp / (tp + fp)}
...
>>> def predict(inputs: dict) -> dict:
... # This can be any function or just an API call to your app.
... return {"output": "Yes"}
...
>>> results = evaluate(
... predict,
... data=dataset_name,
... evaluators=[accuracy],
... summary_evaluators=[precision],
... ) # doctest: +ELLIPSIS
View the evaluation results for experiment:...
Evaluating over only a subset of the examples
>>> experiment_name = results.experiment_name
>>> examples = client.list_examples(dataset_name=dataset_name, limit=5)
>>> results = evaluate(
... predict,
... data=examples,
... evaluators=[accuracy],
... summary_evaluators=[precision],
... experiment_prefix="My Experiment",
... ) # doctest: +ELLIPSIS
View the evaluation results for experiment:...
Streaming each prediction to more easily + eagerly debug.
>>> results = evaluate(
... predict,
... data=dataset_name,
... evaluators=[accuracy],
... summary_evaluators=[precision],
... blocking=False,
... ) # doctest: +ELLIPSIS
View the evaluation results for experiment:...
>>> for i, result in enumerate(results): # doctest: +ELLIPSIS
... pass
Using the `evaluate` API with an off-the-shelf LangChain evaluator:
>>> from langsmith.evaluation import LangChainStringEvaluator
>>> def prepare_criteria_data(run: Run, example: Example):
... return {
... "prediction": run.outputs["output"],
... "reference": example.outputs["answer"],
... "input": str(example.inputs),
... }
...
>>> results = evaluate(
... predict,
... data=dataset_name,
... evaluators=[
... accuracy,
... LangChainStringEvaluator("embedding_distance"),
... LangChainStringEvaluator(
... "labeled_criteria",
... config={
... "criteria": {
... "usefulness": "The prediction is useful if it is correct"
... " and/or asks a useful followup question."
... },
... },
... prepare_data=prepare_criteria_data
... ),
... ],
... summary_evaluators=[precision],
... ) # doctest: +ELLIPSIS
View the evaluation results for experiment:...
Evaluating a LangChain object:
>>> from langchain_core.runnables import chain as as_runnable
>>> @as_runnable
... def nested_predict(inputs):
... return {"output": "Yes"}
...
>>> @as_runnable
... def lc_predict(inputs):
... return nested_predict.invoke(inputs)
...
>>> results = evaluate(
... lc_predict.invoke,
... data=dataset_name,
... evaluators=[accuracy],
... summary_evaluators=[precision],
... ) # doctest: +ELLIPSIS
View the evaluation results for experiment:...
"""
Some key differences to point out:
1. dataset_name => data
2. llm_or_chain_factory now is always the first positional-only argument
- Factories are NOT directly supported. Instead of (if your object is stateful),
def factory(): my_pipeline = ... return my_pipeline
You would write:
def predict(inputs: dict): my_pipeline = ... return my_pipeline.invoke(inputs)
Note this assumed my_pipeline is a Langchain runnable, which has the invoke() method.
3. Instead of dataset_version as a first-class citizen, you'd write (note: keyword-args are required)
data=client.list_examples(dataset_name=dataset_name, as_of=dataset_version)
4. RunEvalConfig is nixed. Instead, directly provide a list of evaluators.
5. LangChain evaluators are no longer a first-class citizen. You'd define using the wrapper class as shown above. Example:
```python
eval_config=RunEvalConfig(evaluators=[ RunEvalConfig.Criteria("relevance"), RunEvalConfig.Criteria("coherence"), RunEvalConfig.Criteria("helpfulness"), RunEvalConfig.Criteria("conciseness") ])
becomes
evaluators= [
LangChainStringEvaluator(
"criteria",
config={
"criteria": "relevance",
},
),
LangChainStringEvaluator(
"criteria",
config={
"criteria": "coherence",
},
),
LangChainStringEvaluator(
"criteria",
config={
"criteria": "helpfulness",
},
),
LangChainStringEvaluator(
"criteria",
config={
"criteria": "conciseness",
},
),
]
Note they are multiple evaluators, rather than all combined in a single evaluator.
5.b LangChain objects are also no longer a first-class citizen. You should pass in the chain.invoke function (method) or wrap in a predict() function that invokes the function directly.
6. batch_evaluators -> summary_evaluators
7. project_metadata -> metadata
8. project_name -> experiment_prefix ; also it doesn't need a UUID in the string if htat is present
9. concurrency_level -> max_concurrency
EXAMPLES:
The following example diffs were manually performed to convert run_on_dataset calls to evaluate() calls:
Example 1
@@ -15,8 +15,9 @@
toxic_examples = [
]
toxic_dataset_name = "Toxic Queries"
-toxic_dataset = client.create_dataset(dataset_name=toxic_dataset_name)
-inputs, outputs = zip(
- *[({"text": text}, {"label": label}) for text, label in toxic_examples]
-)
-client.create_examples(inputs=inputs, outputs=outputs, dataset_id=toxic_dataset.id)
+
+if not client.has_dataset(dataset_name=toxic_dataset_name):
+ toxic_dataset = client.create_dataset(dataset_name=toxic_dataset_name)
+ inputs, outputs = zip(
+ *[({"text": text}, {"label": label}) for text, label in toxic_examples]
+ )
+ client.create_examples(inputs=inputs, outputs=outputs, dataset_id=toxic_dataset.id)
@@ -1,6 +1,7 @@
movie_creation_examples = ["soccer", "a pop star", "action movie in venice"]
movie_creation_dataset_name = "Movie Creation"
-movie_dataset = client.create_dataset(dataset_name=movie_creation_dataset_name)
-for topic in movie_creation_examples:
- client.create_example(inputs={"topic": topic}, dataset_id=movie_dataset.id)
+
+if not client.has_dataset(dataset_name=movie_creation_dataset_name):
+ movie_dataset = client.create_dataset(dataset_name=movie_creation_dataset_name)
+ for topic in movie_creation_examples:
+ client.create_example(inputs={"topic": topic}, dataset_id=movie_dataset.id)
Example 2
This shows an instance of converting a factory to an accepted function.
@@ -1,12 +1,11 @@
from dateutil.parser import parse
-
-from langchain.agents import AgentExecutor
+from langchain.agents import AgentExecutor, create_openai_tools_agent
from langchain.agents.format_scratchpad import format_to_openai_functions
from langchain.agents.output_parsers import OpenAIFunctionsAgentOutputParser
-from langchain_openai import ChatOpenAI
from langchain.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain.tools import DuckDuckGoSearchResults, tool
from langchain.tools.render import format_tool_to_openai_function
+from langchain_openai import ChatOpenAI
@tool
@@ -24,7 +23,7 @@ def check_calendar(date: str) -> list:
return ["Focus time"] # If only...
-def agent_factory():
+def agent(inputs: dict):
llm = ChatOpenAI(
model="gpt-3.5-turbo-16k",
temperature=0,
@@ -35,32 +34,19 @@ def agent_factory():
), # General internet search using DuckDuckGo
check_calendar,
]
- llm_with_tools = llm.bind(
- functions=[format_tool_to_openai_function(t) for t in tools]
- )
prompt = ChatPromptTemplate.from_messages(
[
("system", "You are a helpful assistant."),
MessagesPlaceholder(variable_name="agent_scratchpad"),
- ("user", "{input}"),
+ ("user", "{question}"),
]
)
+ runnable_agent = create_openai_tools_agent(llm, tools, prompt)
- runnable_agent = (
- {
- "input": lambda x: x["question"],
- "agent_scratchpad": lambda x: format_to_openai_functions(
- x["intermediate_steps"]
- ),
- }
- | prompt
- | llm_with_tools
- | OpenAIFunctionsAgentOutputParser()
- )
-
- return AgentExecutor(
+ executor = AgentExecutor(
agent=runnable_agent,
tools=tools,
handle_parsing_errors=True,
return_intermediate_steps=True,
- )
+
+ )
+ return executor.invoke(inputs)
Example 3
We now prefer functions directly over subclassing RunEvaluator, and it's easier to return a dict, though you can still return an EvaluationResult.
@@ -1,24 +1,20 @@
from typing import Optional
-from langsmith.evaluation import EvaluationResult, RunEvaluator
from langsmith.schemas import Example, Run
-class AgentTrajectoryEvaluator(RunEvaluator):
- def evaluate_run(
- self, run: Run, example: Optional[Example] = None
- ) -> EvaluationResult:
- if run.outputs is None:
- raise ValueError("Run outputs cannot be None")
- # This is the output of each run
- intermediate_steps = run.outputs["intermediate_steps"]
- # Since we are comparing to the tool names, we now need to get that
- # Intermediate steps is a Tuple[AgentAction, Any]
- # The first element is the action taken
- # The second element is the observation from taking that action
- trajectory = [action.tool for action, _ in intermediate_steps]
- # This is what we uploaded to the dataset
- expected_trajectory = example.outputs["expected_steps"]
- # Just score it based on whether it is correct or not
- score = int(trajectory == expected_trajectory)
- return EvaluationResult(key="Intermediate steps correctness", score=score)
+
+def intermediate_step_correctness(run: Run, example: Optional[Example] = None) -> dict:
+ if run.outputs is None:
+ raise ValueError("Run outputs cannot be None")
+ # This is the output of each run
+ intermediate_steps = run.outputs.get("intermediate_steps") or []
+ # Since we are comparing to the tool names, we now need to get that
+ # Intermediate steps is a Tuple[AgentAction, Any]
+ # The first element is the action taken
+ # The second element is the observation from taking that action
+ trajectory = [action.tool for action, _ in intermediate_steps]
+ # This is what we uploaded to the dataset
+ expected_trajectory = example.outputs["expected_steps"]
+ # Just score it based on whether it is correct or not
+ score = int(trajectory == expected_trajectory)
+ return {"key": "Intermediate steps correctness", "score": score}
Example 4
from langchain.smith import RunEvalConfig
+from langsmith.evaluation import LangChainStringEvaluator, evaluate
+from langsmith.schemas import Example, Run
-def construct_chain():
+def predict(inputs: dict):
# Add a step to convert the data from the dataset to a form the chain can consume
+ return chain.invoke(
+ {
+ "input": inputs["question"],
+ "chat_history": convert_openai_messages(inputs["chat_history"]),
+ }
+ )
+
+
+def format_evaluator_inputs(run: Run, example: Example):
return {
- "input": lambda x: x["question"],
- "chat_history": lambda x: convert_openai_messages(x["chat_history"]),
- } | chain
+ "input": example.inputs["question"],
+ "prediction": next(iter(run.outputs.values())),
+ "reference": example.outputs["expected"],
+ }
+
+correctness_evaluator = LangChainStringEvaluator(
+ "labeled_score_string",
+ config={"criteria": "correctness", "normalize_by": 10},
+ prepare_data=format_evaluator_inputs,
+)
-results = client.run_on_dataset(
- dataset_name=dataset_name,
- llm_or_chain_factory=construct_chain,
- evaluation=RunEvalConfig(
- evaluators=[
- RunEvalConfig.LabeledScoreString(criteria="correctness", normalize_by=10)
- ],
- # We must specify which key in the example inputs to pass to the evaluator
- input_key="question",
- ),
+results = evaluate(
+ predict,
+ data=dataset_name,
+ experiment_prefix="Chat Single Turn",
+ evaluators=[correctness_evaluator],
+ metadata={"model": "gpt-3.5-turbo"},
)
Example 4
In this case, we can use chain.invoke directly (it's a method/function) since teh dataset's input keys align with the expected inputs of the pipeline (chain).
@@ -1,12 +1,9 @@
-from langchain.smith import RunEvalConfig
+from langsmith.evaluation import LangChainStringEvaluator, evaluate
-eval_config = RunEvalConfig(
- evaluators=["json_edit_distance"],
-)
-res = client.run_on_dataset(
- dataset_name=dataset_name,
- llm_or_chain_factory=chain,
- evaluation=eval_config,
+res = evaluate(
+ chain.invoke,
+ data=dataset_name,
+ evaluators=[LangChainStringEvaluator("json_edit_distance")],
# In case you are rate-limited
- concurrency_level=2,
+ max_concurrency=2,
)
Example 5
We can use the batch create_examples method in lieu of a for loop.
@@ -1,16 +1,16 @@
-from langsmith import Client
import uuid
+from langsmith import Client
+
client = Client()
-dataset_name = f"Dynamic Titanic CSV {str(uuid.uuid4())}"
+dataset_name = f"Dynamic Titanic CSV {uuid.uuid4().hex[:4]}"
dataset = client.create_dataset(
dataset_name=dataset_name,
description="Test QA over CSV",
)
-for example in questions:
- client.create_example(
- inputs={"question": example[0]},
- outputs={"code": example[1]},
- dataset_id=dataset.id,
- )
+
+client.create_examples(
+ inputs=[{"question": example[0]} for example in questions],
+ outputs=[{"code": example[1]} for example in questions],
+ dataset_id=dataset.id,
+)
Example 6
Prefer regular python functions vs. functools.partial since some of our users aren't as familiar with them.
@@ -1,13 +1,10 @@
-from functools import partial
-
-from langchain.chat_models import ChatOpenAI
-from langchain.prompts import ChatPromptTemplate
+from langchain_core.prompts import ChatPromptTemplate
from langchain_experimental.agents import create_pandas_dataframe_agent
+from langchain_openai import ChatOpenAI
+
+llm = ChatOpenAI(model="gpt-4-turbo-preview", temperature=0.0)
+
-llm = ChatOpenAI(model="gpt-4", temperature=0.0)
-create_chain = partial(
- create_pandas_dataframe_agent,
- agent_type="openai-tools",
- llm=llm,
- df=df,
-)
+
+def predict(inputs: dict):
+ agent = create_pandas_dataframe_agent(agent_type="openai-tools", llm=llm, df=df)
+ return agent.invoke({"input": inputs["question"]})
@@ -1,7 +1,10 @@
-chain_results = run_on_dataset(
- client,
- dataset_name=dataset_name,
- llm_or_chain_factory=create_chain_2,
- evaluation=eval_config,
- concurrency_level=1,
+chain_results = evaluate(
+ predict,
+ data=dataset_name,
+ evaluators=[criteria_evaluator],
+ # This agent doesn't support concurrent runs yet.
+ max_concurrency=1,
+ metadata={
+ "time": "T2",
+ },
)
Convert the following to use the evaluate function.
{input}
If the input is a notebook, respond with a valid .ipynb notebook that can be directly rendered. Otherwise, respond with the proper correction.
This prompt contains variables shown as ⟨variable_name⟩. Replace them with your own values before using.
How to Use
Use with LangChain: hub.pull("wfh/run2evaluate")
Related Prompts
More prompts in Data & Analytics
Sql Agent System Prompt
LangChain Hub prompt: langchain-ai/sql-agent-system-prompt
Buyer Persona Legend
Generate detailed User Personas for your Business with data neatly organized into a table.
Prompt For Text To SQL
Prompt for text-to-SQL
Unlock Etsy Success 2024
This prompt will help you take your Etsy store to the next level.
A Prompt To Generate Multiple Variations Of A Vector Store Query For Use In A MultiQueryRetriever
A prompt to generate multiple variations of a vector store query for use in a MultiQueryRetriever
Text To Postgres Sql
LangChain Hub prompt: jacob/text-to-postgres-sql