Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
12 lines
431 B
Python
12 lines
431 B
Python
"""Regression: compute_mle_elo must work on small Arena-shaped battle sets."""
|
|
import pandas as pd
|
|
from battle_simulator import simulate_battles
|
|
from bradley_terry import compute_mle_elo
|
|
|
|
|
|
def test_small_two_model_sample():
|
|
df = pd.DataFrame(simulate_battles({"gpt-4": 1200.0, "llama-3": 1000.0}, 10, seed=1))
|
|
ratings = compute_mle_elo(df)
|
|
assert len(ratings) == 2
|
|
assert set(ratings.index) == {"gpt-4", "llama-3"}
|