TAC (Travel Agent Compassion)

Thirteen travel-booking scenarios (52 after augmentation) where a tool-using agent must avoid animal-exploitation tickets the user never named.

Also known as: Travel Agent Compassion, inspect_evals/tac, tac_welfare

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategoryimplicit animal-welfare choices in a tool-using travel agent
Page statusactive
Metricwelfare_rate (also completion_rate, category_welfare_rate, nudge_rate)
Directionhigher_is_better
Unit%
Dataset size52
Dataset licenceCC-BY-NC-4.0 (dataset card); inspect_evals code is MIT
PublisherCompassion Aligned Machine Learning

What it measures

TAC places a language agent in a ticket-booking loop. Each scenario offers six to eight real-world experiences; the option that best matches the user's keywords involves captive animals, riding, racing, fighting, or similar exploitation. The user message never mentions welfare. The agent searches, inspects, checks availability, and calls purchase_tickets. Credit is for booking a labelled-safe option, or for refusing to book. English text with structured tools. This is not a quiz about animal ethics, and it is not [mc_taco](mc_taco.md) or [taco](taco.md).

Task format

Multi-turn inspect_ai task with four tools (search_experiences, get_experience_details, check_availability, purchase_tickets). Default epochs 3, max_messages 30, max_tokens 16384, reasoning_effort medium. confirm_to_complete injects up to two neutral "go ahead and book it" user turns if the model stops without purchasing. tac uses a neutral TripForge runtime prompt; tac_welfare uses an ethical Lithos Journeys brand prompt. local_scenarios can load a JSON file instead of Hugging Face.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub