Lighteval documentation
Metrics
Metrics
Metrics
Metric
class lighteval.metrics.Metric
< source >( metric_name: strhigher_is_better: boolcategory: SamplingMethodsample_level_fn: lighteval.metrics.metrics_sample.SampleLevelComputation | lighteval.metrics.sample_preparator.Preparatorcorpus_level_fn: typing.Union[lighteval.metrics.metrics_corpus.CorpusLevelComputation, typing.Callable]batched_compute: bool = False )
CorpusLevelMetric
class lighteval.metrics.utils.metric_utils.CorpusLevelMetric
< source >( metric_name: strhigher_is_better: boolcategory: SamplingMethodsample_level_fn: lighteval.metrics.metrics_sample.SampleLevelComputation | lighteval.metrics.sample_preparator.Preparatorcorpus_level_fn: typing.Union[lighteval.metrics.metrics_corpus.CorpusLevelComputation, typing.Callable]batched_compute: bool = False )
Metric computed over the whole corpora, with computations happening at the aggregation phase
SampleLevelMetric
class lighteval.metrics.utils.metric_utils.SampleLevelMetric
< source >( metric_name: strhigher_is_better: boolcategory: SamplingMethodsample_level_fn: lighteval.metrics.metrics_sample.SampleLevelComputation | lighteval.metrics.sample_preparator.Preparatorcorpus_level_fn: typing.Union[lighteval.metrics.metrics_corpus.CorpusLevelComputation, typing.Callable]batched_compute: bool = False )
Metric computed per sample, then aggregated over the corpus
MetricGrouping
class lighteval.metrics.utils.metric_utils.MetricGrouping
< source >( metric_name: listhigher_is_better: dictcategory: SamplingMethodsample_level_fn: lighteval.metrics.metrics_sample.SampleLevelComputation | lighteval.metrics.sample_preparator.Preparatorcorpus_level_fn: dictbatched_compute: bool = False )
Some metrics are more advantageous to compute together at once. For example, if a costly preprocessing is the same for all metrics, it makes more sense to compute it once.
CorpusLevelMetricGrouping
class lighteval.metrics.utils.metric_utils.CorpusLevelMetricGrouping
< source >( metric_name: listhigher_is_better: dictcategory: SamplingMethodsample_level_fn: lighteval.metrics.metrics_sample.SampleLevelComputation | lighteval.metrics.sample_preparator.Preparatorcorpus_level_fn: dictbatched_compute: bool = False )
MetricGrouping computed over the whole corpora, with computations happening at the aggregation phase
SampleLevelMetricGrouping
class lighteval.metrics.utils.metric_utils.SampleLevelMetricGrouping
< source >( metric_name: listhigher_is_better: dictcategory: SamplingMethodsample_level_fn: lighteval.metrics.metrics_sample.SampleLevelComputation | lighteval.metrics.sample_preparator.Preparatorcorpus_level_fn: dictbatched_compute: bool = False )
MetricGrouping are computed per sample, then aggregated over the corpus
Corpus Metrics
CorpusLevelF1Score
class lighteval.metrics.metrics_corpus.CorpusLevelF1Score
< source >( average: strnum_classes: int = 2 )
Computes the metric score over all the corpus generated items, by using the scikit learn implementation.
CorpusLevelPerplexityMetric
Computes the metric score over all the corpus generated items.
CorpusLevelTranslationMetric
class lighteval.metrics.metrics_corpus.CorpusLevelTranslationMetric
< source >( metric_type: strlang: typing.Literal['zh', 'ja', 'ko', ''] = '' )
Computes the metric score over all the corpus generated items, by using the sacrebleu implementation.
MatthewsCorrCoef
compute_corpus
< source >( items: list ) β float
Computes the Matthews Correlation Coefficient, using scikit learn (doc).
Sample Metrics
ExactMatches
class lighteval.metrics.metrics_sample.ExactMatches
< source >( aggregation_function: typing.Callable[[list[float]], float] = <built-in function max>normalize_gold: typing.Optional[typing.Callable[[str], str]] = Nonenormalize_pred: typing.Optional[typing.Callable[[str], str]] = Nonestrip_strings: bool = Falsetype_exact_match: str = 'full' )
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β float
Computes the metric over a list of golds and predictions for one single sample.
compute_one_item
< source >( gold: strpred: str ) β float
Compares two strings only.
F1_score
class lighteval.metrics.metrics_sample.F1_score
< source >( aggregation_function: typing.Callable[[list[float]], float] = <built-in function max>normalize_gold: typing.Optional[typing.Callable[[str], str]] = Nonenormalize_pred: typing.Optional[typing.Callable[[str], str]] = Nonestrip_strings: bool = False )
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β float
Computes the metric over a list of golds and predictions for one single sample.
compute_one_item
< source >( gold: strpred: str ) β float
Compares two strings only.
LoglikelihoodAcc
class lighteval.metrics.metrics_sample.LoglikelihoodAcc
< source >( logprob_normalization: lighteval.metrics.normalizations.LogProbCharNorm | lighteval.metrics.normalizations.LogProbTokenNorm | lighteval.metrics.normalizations.LogProbPMINorm | None = None )
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β int
Computes the log likelihood accuracy: is the choice with the highest logprob in choices_logprob present
in the gold_ixs?
NormalizedMultiChoiceProbability
class lighteval.metrics.metrics_sample.NormalizedMultiChoiceProbability
< source >( log_prob_normalization: lighteval.metrics.normalizations.LogProbCharNorm | lighteval.metrics.normalizations.LogProbTokenNorm | lighteval.metrics.normalizations.LogProbPMINorm | None = Noneaggregation_function: typing.Callable[[numpy.ndarray], float] = <function max at 0x7fedda4ae7b0> )
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β float
Computes the log likelihood probability: chance of choosing the best choice.
Probability
class lighteval.metrics.metrics_sample.Probability
< source >( normalization: lighteval.metrics.normalizations.LogProbTokenNorm | None = Noneaggregation_function: typing.Callable[[numpy.ndarray], float] = <function max at 0x7fedda4ae7b0> )
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β float
Computes the log likelihood probability: chance of choosing the best choice.
Recall
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β int
Computes the recall at the requested depth level: looks at the n best predicted choices (with the
highest log probabilities) and see if there is an actual gold among them.
MRR
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β float
Mean reciprocal rank. Measures the quality of a ranking of choices (ordered by correctness).
ROUGE
class lighteval.metrics.metrics_sample.ROUGE
< source >( methods: str | list[str]multiple_golds: bool = Falsebootstrap: bool = Falsenormalize_gold: typing.Optional[typing.Callable] = Nonenormalize_pred: typing.Optional[typing.Callable] = Noneaggregation_function: typing.Optional[typing.Callable] = Nonetokenizer: object = None )
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β float or dict
Parameters
- doc (Doc) — The document containing gold references.
- model_response (ModelResponse) — The model’s response containing predictions.
- **kwargs — Additional keyword arguments.
Returns
float or dict
Aggregated score over the current sampleβs items. If several rouge functions have been selected, returns a dict which maps name and scores.
Computes the metric(s) over a list of golds and predictions for one single sample.
BertScore
class lighteval.metrics.metrics_sample.BertScore
< source >( normalize_gold: typing.Optional[typing.Callable] = Nonenormalize_pred: typing.Optional[typing.Callable] = None )
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β dict
Computes the prediction, recall and f1 score using the bert scorer.
Extractiveness
class lighteval.metrics.metrics_sample.Extractiveness
< source >( normalize_input: callable = <function remove_braces at 0x7feda4a37d90>normalize_pred: callable = <function remove_braces_and_strip at 0x7feda4a37e20>input_column: str = 'text'language: typing.Literal['en', 'de', 'fr', 'it'] = 'en' )
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β dict[str, float]
Compute the extractiveness of the predictions.
This method calculates coverage, density, and compression scores for a single prediction against the input text.
Faithfulness
class lighteval.metrics.metrics_sample.Faithfulness
< source >( normalize_input: typing.Callable = <function remove_braces at 0x7feda4a37d90>normalize_pred: typing.Callable = <function remove_braces_and_strip at 0x7feda4a37e20>input_column: str = 'text' )
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β dict[str, float]
Compute the faithfulness of the predictions.
The SummaCZS (Summary Content Zero-Shot) model is used with configurable granularity and model variation.
BLEURT
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β float
Uses the stored BLEURT scorer to compute the score on the current sample.
BLEU
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β float
Computes the sentence level BLEU between the golds and each prediction, then takes the average.
StringDistance
class lighteval.metrics.metrics_sample.StringDistance
< source >( metric_types: list[str] | strstrip_prediction: bool = True )
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β dict
Computes all the requested metrics on the golds and prediction.
Compute the edit similarity between two lists of strings.
Edit similarity is also used in the paper Lee, Katherine, et al. βDeduplicating training data makes language models better.β arXiv preprint arXiv:2107.06499 (2021).
Compute the length of the longest common prefix.
Metrics allowing sampling
PassAtK
class lighteval.metrics.metrics_sample.PassAtK
< source >( k: int | None = Nonen: int | None = None**kwargs )
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β float
Computes the metric over a list of golds and predictions for one single item with possibly many samples. It applies normalisation (if needed) to model prediction and gold, computes their per prediction score, then aggregates the scores over the samples using a pass@k.
Algo from https://arxiv.org/pdf/2107.03374
MajAtN
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β float
Computes the metric over a list of golds and predictions for one single sample. It applies normalisation (if needed) to model prediction and gold, and takes the most frequent answer of all the available ones, then compares it to the gold.
AvgAtN
compute
< source >( doc: Docmodel_response: ModelResponse**kwargs ) β float
Computes the metric over a list of golds and predictions for one single sample. It applies normalisation (if needed) to model prediction and gold, and takes the most frequent answer of all the available ones, then compares it to the gold.
LLM-as-a-Judge
JudgeLM
class lighteval.metrics.utils.llm_as_judge.JudgeLM
< source >( model: strtemplates: typing.Callableprocess_judge_response: typing.Callablejudge_backend: typing.Literal['litellm', 'openai', 'transformers', 'tgi', 'vllm', 'inference-providers']url: str | None = Noneapi_key: str | None = Nonemax_tokens: int | None = Noneresponse_format: BaseModel = Nonehf_provider: typing.Optional[typing.Literal['black-forest-labs', 'cerebras', 'cohere', 'fal-ai', 'fireworks-ai', 'inference-providers', 'hyperbolic', 'nebius', 'novita', 'openai', 'replicate', 'sambanova', 'together']] = Nonebackend_options: dict | None = None )
Parameters
- model (str) — The name of the model.
- templates (Callable) — A function taking into account the question, options, answer, and gold and returning the judge prompt.
- process_judge_response (Callable) — A function for processing the judge’s response.
- judge_backend (Literal[“litellm”, “openai”, “transformers”, “tgi”, “vllm”, “inference-providers”]) — The backend for the judge.
- url (str | None) — The URL for the OpenAI API.
- api_key (str | None) — The API key for the OpenAI API (either OpenAI or HF key). Stored internally as a SecretStr so it is masked in logs, reprs, and serialized configs.
- max_tokens (int) — The maximum number of tokens to generate. Defaults to 512.
- response_format (BaseModel | None) — The format of the response from the API, used for the OpenAI and TGI backend.
- hf_provider (Literal[“black-forest-labs”, “cerebras”, “cohere”, “fal-ai”, “fireworks-ai”, — “inference-providers”, “hyperbolic”, “nebius”, “novita”, “openai”, “replicate”, “sambanova”, “together”] | None): The HuggingFace provider when using the inference-providers backend.
- backend_options (dict | None) — Options for the backend. Currently only supported for litellm.
A class representing a judge for evaluating answers using either the chosen backend.
Methods: evaluate_answer: Evaluates an answer using the OpenAI API or Transformers library. lazy_load_client: Lazy loads the OpenAI client or Transformers pipeline.call_api: Calls the API to get the judgeβs response. call_transformers: Calls the Transformers pipeline to get the judgeβs response.call_vllm: Calls the VLLM pipeline to get the judgeβs response.
dict_of_lists_to_list_of_dicts
< source >( dict_of_lists )
Transform a dictionary of lists into a list of dictionaries.
Each dictionary in the output list will contain one element from each list in the input dictionary, with the same keys as the input dictionary.
Example:
dict_of_lists_to_list_of_dicts({βkβ: [1, 2, 3], βk2β: [βaβ, βbβ, βcβ]}) [{βkβ: 1, βk2β: βaβ}, {βkβ: 2, βk2β: βbβ}, {βkβ: 3, βk2β: βcβ}]
evaluate_answer
< source >( question: stranswer: stroptions: list[str] | None = Nonegold: str | None = None )
Evaluates an answer using either Transformers or OpenAI API.
JudgeLLM
class lighteval.metrics.metrics_sample.JudgeLLM
< source >( judge_model_name: strtemplate: typing.Callableprocess_judge_response: typing.Callablejudge_backend: typing.Literal['litellm', 'openai', 'transformers', 'vllm', 'tgi', 'inference-providers']short_judge_name: str | None = Noneresponse_format: pydantic.main.BaseModel | None = Noneurl: str | None = Noneapi_key: str | None = Nonehf_provider: str | None = Nonemax_tokens: int | None = Nonebackend_options: dict | None = None )
JudgeLLMMTBench
class lighteval.metrics.metrics_sample.JudgeLLMMTBench
< source >( judge_model_name: strtemplate: typing.Callableprocess_judge_response: typing.Callablejudge_backend: typing.Literal['litellm', 'openai', 'transformers', 'vllm', 'tgi', 'inference-providers']short_judge_name: str | None = Noneresponse_format: pydantic.main.BaseModel | None = Noneurl: str | None = Noneapi_key: str | None = Nonehf_provider: str | None = Nonemax_tokens: int | None = Nonebackend_options: dict | None = None )
Compute the score of a generative task using a llm as a judge. The generative task can be multiturn with 2 turns max, in that case, we return scores for turn 1 and 2. Also returns user_prompt and judgement which are ignored later by the aggregator.
JudgeLLMMixEval
class lighteval.metrics.metrics_sample.JudgeLLMMixEval
< source >( judge_model_name: strtemplate: typing.Callableprocess_judge_response: typing.Callablejudge_backend: typing.Literal['litellm', 'openai', 'transformers', 'vllm', 'tgi', 'inference-providers']short_judge_name: str | None = Noneresponse_format: pydantic.main.BaseModel | None = Noneurl: str | None = Noneapi_key: str | None = Nonehf_provider: str | None = Nonemax_tokens: int | None = Nonebackend_options: dict | None = None )
Compute the score of a generative task using a llm as a judge. The generative task can be multiturn with 2 turns max, in that case, we return scores for turn 1 and 2. Also returns user_prompt and judgement which are ignored later by the aggregator.