Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
17 lines
636 B
Python
17 lines
636 B
Python
import pandas as pd
|
|
from bradley_terry import compute_mle_elo, get_bootstrap_result
|
|
|
|
|
|
def test_bootstrap_is_reproducible():
|
|
battles = pd.DataFrame(
|
|
[
|
|
{"model_a": "a", "model_b": "b", "winner": "model_a"},
|
|
{"model_a": "a", "model_b": "b", "winner": "model_b"},
|
|
{"model_a": "a", "model_b": "b", "winner": "tie"},
|
|
{"model_a": "b", "model_b": "a", "winner": "model_a"},
|
|
]
|
|
)
|
|
first = get_bootstrap_result(battles, compute_mle_elo, num_round=3)
|
|
second = get_bootstrap_result(battles, compute_mle_elo, num_round=3)
|
|
pd.testing.assert_frame_equal(first, second)
|