Algorithmically generated murder mysteries, object placement puzzles and team allocation problems needing long-range narrative reasoning.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | multistep narrative and long-context reasoning |
| Page status | active |
| Metric | accuracy (normalized) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 756 |
| Dataset licence | MIT (authors' GitHub repository); a Hugging Face mirror separately labels it CC BY 4.0 |
| Publisher | University of Texas at Austin |
MuSR (Multistep Soft Reasoning) tests whether a model can answer a question that requires combining facts scattered across a long narrative, rather than reasoning that is stated in one place. Each item is an algorithmically generated story of roughly 1,000 words in one of three domains: a murder mystery (identify the culprit), an object placement puzzle (determine where an object ended up after being moved), or a team allocation problem (assign people to tasks under stated constraints). The stories are produced by a "neurosymbolic synthetic-to-natural generation algorithm," which builds a structured reasoning tree first and then renders it into free-form prose, so the correct answer depends on connecting details spread through the narrative rather than pattern-matching a single sentence. It is a single-turn, English-language, text-only task.
A narrative of roughly 1,000 words followed by a multiple-choice question; the model must integrate facts scattered through the text to select the correct option.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Falcon3 7B Instruct | TII | 48.3 | 2025-03 |
| Falcon3 7B Base | TII | 47.0 | 2025-03 |
| deepseek llm 7B chat | DeepSeek | 46.7 | 2024-07 |
| Nous Hermes 2 Mixtral 8x7B DPO | Nous Research | 46.0 | 2024-07 |
| Yi 6B 200K | 01.AI | 45.9 | 2024-07 |
| Infinity Instruct 3M 0625 Yi 1.5 9B | BAAI | 45.8 | 2025-03 |
| Qwen2.5 Coder 32B Instruct AWQ | Alibaba / Qwen Team | 45.3 | 2026-04 |
| Meta Llama 3 70B | Meta | 45.2 | 2024-07 |
| OpenHermes 2 Mistral 7B | Teknium | 45.2 | 2025-03 |
| Yi 1.5 34B 32K | 01.AI | 44.0 | 2024-07 |
| Yi 1.5 34B Chat 16K | 01.AI | 44.0 | 2024-07 |
| glm 4 9B | Zhipu AI | 43.9 | 2025-03 |
| GLM 4 9B 0414 | Zhipu AI | 43.9 | 2025-03 |
| Qwen2.5 Coder 32B Instruct | Alibaba / Qwen Team | 43.9 | 2026-04 |
| Yi 1.5 6B Chat | 01.AI | 43.9 | 2024-07 |
| Qwen2 VL 7B Instruct | Alibaba / Qwen Team | 43.8 | 2025-03 |
| Qwen2 VL 7B Instruct AWQ | Alibaba / Qwen Team | 43.8 | 2025-03 |
| Hermes 3 Llama 3.1 8B | Nous Research | 43.7 | 2025-03 |
| Hermes 3 Llama 3.1 8B GGUF | Nous Research | 43.7 | 2025-03 |
| Nous Hermes 2 SOLAR 10.7B | Nous Research | 43.7 | 2024-07 |
| Yi 1.5 6B | 01.AI | 43.7 | 2024-07 |
| Yi 1.5 9B | 01.AI | 43.3 | 2024-07 |
| DialoGPT medium | Microsoft | 42.9 | 2025-03 |
| Qwen2 1.5B Instruct | Alibaba / Qwen Team | 42.9 | 2024-07 |
| Yi 9B 200K | 01.AI | 42.9 | 2025-03 |
| Phi 3 mini 4K instruct | Microsoft | 42.8 | 2024-07 |
| gemma 7B it | Google DeepMind | 42.7 | 2024-07 |
| Hermes 2 Pro Llama 3 8B | Nous Research | 42.6 | 2024-07 |
| Nous Hermes llama 2 7B | Nous Research | 42.6 | 2024-07 |
| Yi 1.5 9B Chat | 01.AI | 42.6 | 2024-07 |
| gemma 2 2B it | Google DeepMind | 42.4 | 2025-03 |
| OpenHermes 2.5 Mistral 7B | Teknium | 42.4 | 2025-03 |
| flan t5 xl | Google DeepMind | 42.2 | 2025-03 |
| gemma 2 2B | Google DeepMind | 42.2 | 2025-03 |
| falcon mamba 7B | TII | 42.1 | 2025-03 |
| falcon mamba 7B instruct | TII | 42.1 | 2025-03 |
| falcon mamba 7B instruct Q4 K M GGUF | TII | 42.1 | 2025-03 |
| Falcon3 1B Instruct | TII | 41.9 | 2025-03 |
| Yi 1.5 9B 32K | 01.AI | 41.9 | 2024-07 |
| stablelm zephyr 3B | Stability AI | 41.8 | 2025-03 |
| Falcon3 1B Base | TII | 41.5 | 2025-03 |
| Meta Llama 3 70B Instruct | Meta | 41.5 | 2026-04 |
| Meta Llama 3 70B Instruct | Nous Research | 41.5 | 2026-04 |
| Falcon3 3B Instruct | TII | 41.4 | 2025-03 |
| Ministral 8B Instruct 2410 | Mistral AI | 41.4 | 2025-03 |
| Mistral 7B v0.1 | Mistral AI | 41.4 | 2024-07 |
| flan t5 small | Google DeepMind | 41.2 | 2025-03 |
| Llama 2 70B hf | Meta | 41.2 | 2024-07 |
| Yi 34B | 01.AI | 41.2 | 2024-07 |
| OLMo 1B hf | Allen AI | 41.0 | 2024-07 |
| phi 2 | Microsoft | 41.0 | 2024-07 |
| Qwen2.5 Coder 7B Instruct | Alibaba / Qwen Team | 41.0 | 2026-04 |
| Yi 1.5 9B Chat 16K | 01.AI | 41.0 | 2024-07 |
| flan t5 large | Google DeepMind | 40.8 | 2025-03 |
| Llama 2 7B 32K Instruct | Together AI | 40.6 | 2025-03 |
| Yi 9B | 01.AI | 40.5 | 2024-07 |
| OpenHermes 13B | Teknium | 40.4 | 2025-03 |
| Hermes 3 Llama 3.2 3B | Nous Research | 40.3 | 2025-03 |
| Mistral 7B v0.3 | Mistral AI | 40.3 | 2024-07 |
| mistral 7B v0.3 bnb 4bit | Unsloth | 40.3 | 2024-07 |
| Llama 2 13B chat hf | Meta | 40.1 | 2024-07 |
| falcon 11B | TII | 39.9 | 2024-07 |
| glm 4 9B chat | Zhipu AI | 39.9 | 2025-03 |
| Yi Coder 9B | 01.AI | 39.9 | 2025-03 |
| Yi Coder 9B Chat | 01.AI | 39.9 | 2025-03 |
| gemma 2B | Google DeepMind | 39.8 | 2024-07 |
| Yi 34B Chat | 01.AI | 39.8 | 2024-07 |
| Mistral 7B Instruct v0.2 | Mistral AI | 39.7 | 2024-07 |
| Qwen2.5 3B Instruct | Alibaba / Qwen Team | 39.7 | 2025-03 |
| Hermes 2 Theta Llama 3 8B | Nous Research | 39.5 | 2024-07 |
| rwkv raven 14B | RWKV Foundation | 39.5 | 2025-03 |
| Phi 3 mini 128K instruct | Microsoft | 39.4 | 2024-07 |
| Yi 6B | 01.AI | 39.4 | 2024-07 |
| Llama 3.1 8B | Cerebras | 39.3 | 2025-03 |
| Qwen2.5 Coder 14B Instruct | Alibaba / Qwen Team | 39.1 | 2026-04 |
| granite 3.0 8B instruct | IBM | 39.0 | 2025-03 |
| stablelm 2 1 6B | Stability AI | 38.8 | 2025-03 |
| Falcon3 Mamba 7B Instruct | TII | 38.7 | 2025-03 |
| Phi 4 mini instruct | Microsoft | 38.7 | 2026-04 |
| mt5 small | Google DeepMind | 38.6 | 2025-03 |
| Mistral 7B Instruct v0.1 | Mistral AI | 38.5 | 2024-07 |
| Yi 34B 200K | 01.AI | 38.2 | 2024-07 |
| Meta Llama 3 8B Instruct | Meta | 38.1 | 2024-07 |
| Meta Llama 3 8B Instruct | Nous Research | 38.1 | 2024-07 |
| glm 4 9B chat 1M | Zhipu AI | 37.9 | 2025-03 |
| falcon 7B | TII | 37.8 | 2024-07 |
| stablelm 3B 4e1t | Stability AI | 37.8 | 2025-03 |
| falcon 40B instruct | TII | 37.6 | 2024-07 |
| Falcon3 3B Base | TII | 37.5 | 2025-03 |
| LLaMA 2 7B 32K | Together AI | 37.5 | 2025-03 |
| deepseek llm 7B base | DeepSeek | 37.4 | 2024-07 |
| GPT JT 6B v1 | Together AI | 37.4 | 2025-03 |
| Mistral 7B Instruct v0.3 | Mistral AI | 37.4 | 2025-03 |
| RedPajama INCITE Base 3B v1 | Together AI | 37.4 | 2025-03 |
| Infinity Instruct 3M 0625 Llama3 8B | BAAI | 37.1 | 2025-03 |
| Llama 2 7B hf | Meta | 37.0 | 2024-07 |
| Llama 2 7B hf | Nous Research | 37.0 | 2024-07 |
| Llama 2 70B chat hf | Meta | 36.9 | 2024-07 |
| Qwen2.5 Math 1.5B | Alibaba / Qwen Team | 36.9 | 2025-03 |
| RedPajama INCITE 7B Instruct | Together AI | 36.9 | 2025-03 |
| Yi 6B Chat | 01.AI | 36.9 | 2024-07 |
| Yi 6B Chat 4bits | 01.AI | 36.9 | 2025-03 |
| Llama 2 7B chat hf | Meta | 36.8 | 2024-07 |
| Llama 2 7B chat hf | Nous Research | 36.8 | 2024-07 |
| RedPajama INCITE Chat 3B v1 | Together AI | 36.8 | 2025-03 |
| flan t5 base | Google DeepMind | 36.7 | 2025-03 |
| mt5 base | Google DeepMind | 36.7 | 2025-03 |
| deepseek moe 16B base | DeepSeek | 36.6 | 2025-03 |
| Qwen2.5 1.5B Instruct | Alibaba / Qwen Team | 36.6 | 2025-03 |
| OLMoE 1B 7B 0125 | Allen AI | 36.4 | 2025-03 |
| OLMoE 1B 7B 0125 Instruct | Allen AI | 36.4 | 2025-03 |
| DeepSeek R1 Distill Qwen 1.5B | DeepSeek | 36.3 | 2026-04 |
| falcon 40B | TII | 36.3 | 2024-07 |
| falcon 7B instruct | TII | 36.3 | 2024-07 |
| RedPajama INCITE 7B Base | Together AI | 36.2 | 2025-03 |
| granite 3.1 2B instruct | IBM | 36.1 | 2025-03 |
| Meta Llama 3 8B | Meta | 36.1 | 2024-07 |
| Meta Llama 3 8B | Nous Research | 36.1 | 2024-07 |
| glm 4 9B chat hf | Zhipu AI | 35.9 | 2025-03 |
| Jamba v0.1 | AI21 Labs | 35.9 | 2025-03 |
| Infinity Instruct 7M Gen Llama3 1 8B | BAAI | 35.8 | 2025-03 |
| Llama 3.2 3B | Meta | 35.8 | 2026-04 |
| Qwen2.5 1.5B | Alibaba / Qwen Team | 35.8 | 2025-03 |
| Qwen2.5 1.5B Instruct AWQ | Alibaba / Qwen Team | 35.8 | 2025-03 |
| Llama 2 13B hf | Meta | 35.4 | 2024-07 |
| Llama 2 13B hf | Nous Research | 35.4 | 2024-07 |
| Llama 3.2 3B Instruct | Meta | 35.3 | 2026-04 |
| OLMo 2 1124 7B | Allen AI | 35.1 | 2025-03 |
| OLMoE 1B 7B 0924 | Allen AI | 34.9 | 2025-03 |
| GPT NeoXT Chat Base 20B | Together AI | 34.6 | 2025-03 |
| Llama 3.2 1B | Meta | 34.5 | 2025-03 |
| Llama 3.2 1B | Nous Research | 34.5 | 2025-03 |
| Qwen2.5 Coder 7B Instruct GPTQ Int4 | Alibaba / Qwen Team | 34.5 | 2026-04 |
| RedPajama INCITE 7B Chat | Together AI | 34.5 | 2025-03 |
| Qwen2.5 0.5B | Alibaba / Qwen Team | 34.3 | 2025-03 |
| gemma 1.1 2B it | Google DeepMind | 33.9 | 2025-03 |
| Ministral 3B (latest) | Mistral AI | 33.8 | 2025-03 |
| granite 3.0 1B a400m base | IBM | 33.7 | 2025-03 |
| DeepSeek R1 | DeepSeek | 33.5 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 33.5 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 33.5 | 2026-04 |
| DeepSeek Reasoner | DeepSeek | 33.5 | 2026-04 |
| Qwen2 0.5B Instruct | Alibaba / Qwen Team | 33.5 | 2024-07 |
| gemma 2B it | Google DeepMind | 33.4 | 2024-07 |
| Qwen2.5 0.5B Instruct | Alibaba / Qwen Team | 33.4 | 2025-03 |
| granite 3.1 1B a400m instruct | IBM | 33.0 | 2025-03 |
| Llama 3.2 1B Instruct | Meta | 32.0 | 2025-03 |
| Gemma 4 31B | Google DeepMind | 30.4 | 2026-04 |
| gemma 4 31B it | Google DeepMind | 30.4 | 2026-04 |
| gemma 4 31B it GGUF | Unsloth | 30.4 | 2026-04 |
| Gemma 4 31B IT NVFP4 | NVIDIA | 30.4 | 2026-04 |
| Qwen3 32B | Alibaba / Qwen Team | 30.1 | 2026-04 |
| Qwen3 32B AWQ | Alibaba / Qwen Team | 30.1 | 2026-04 |
| Qwen3 32B NVFP4 | NVIDIA | 30.1 | 2026-04 |
| DeepSeek Chat | DeepSeek | 29.8 | 2026-04 |
| DeepSeek V2 | DeepSeek | 29.8 | 2026-04 |
| DeepSeek V2 Lite | DeepSeek | 29.8 | 2026-04 |
| DeepSeek V2 Lite Chat | DeepSeek | 29.8 | 2026-04 |
| DeepSeek V3 | DeepSeek | 29.8 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 29.8 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 29.8 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 29.8 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 29.8 | 2026-04 |
| Gemma 4 26B | Google DeepMind | 28.7 | 2026-04 |
| Qwen2.5 72B Instruct | Alibaba / Qwen Team | 28.3 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 26.3 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 26.3 | 2026-04 |
| Mistral Large 3 | Mistral AI | 26.3 | 2026-04 |
| DeepSeek R1 Distill Llama 70B | DeepSeek | 25.8 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 25.2 | 2026-04 |
| Qwen3 14B | Alibaba / Qwen Team | 24.8 | 2026-04 |
| Qwen3 14B AWQ | Alibaba / Qwen Team | 24.8 | 2026-04 |
| Qwen3 14B NVFP4 | NVIDIA | 24.8 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 24.5 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 24.5 | 2026-04 |
| phi 4 | Microsoft | 23.5 | 2026-04 |
| Phi 4 multimodal instruct | Microsoft | 23.5 | 2026-04 |
| Llama 3.1 70B | Meta | 22.7 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 22.7 | 2026-04 |
| Qwen2.5 32B Instruct | Alibaba / Qwen Team | 22.1 | 2026-04 |
| Qwen2.5 32B Instruct AWQ | Alibaba / Qwen Team | 22.1 | 2026-04 |
| Qwen3 30B A3B Instruct 2507 | Alibaba / Qwen Team | 21.3 | 2026-04 |
| Qwen3 30B A3B NVFP4 | NVIDIA | 21.3 | 2026-04 |
| Qwen3 30B-A3B | Alibaba / Qwen Team | 21.3 | 2026-04 |
| DeepSeek R1 Distill Qwen 32B | DeepSeek | 20.8 | 2026-04 |
| Gemma 3 12B | Google DeepMind | 20.8 | 2026-04 |
| gemma 2 27B it | Google DeepMind | 20.1 | 2026-04 |
| Qwen3 8B | Alibaba / Qwen Team | 19.4 | 2026-04 |
| Qwen3 8B AWQ | Alibaba / Qwen Team | 19.4 | 2026-04 |
| Qwen3 8B Base | Alibaba / Qwen Team | 19.4 | 2026-04 |
| Qwen2.5 14B Instruct | Alibaba / Qwen Team | 18.6 | 2026-04 |
| Qwen2.5 14B Instruct AWQ | Alibaba / Qwen Team | 18.6 | 2026-04 |
| Mixtral 8x22B | Mistral AI | 18.5 | 2026-04 |
| Mixtral 8x22B Instruct v0.1 | Mistral AI | 18.5 | 2026-04 |
| Yi 1.5 34B | 01.AI | 18.2 | 2026-04 |
| Yi 1.5 34B Chat | 01.AI | 18.2 | 2026-04 |
| Command R+ | Cohere | 17.3 | 2026-04 |
| DeepSeek R1 Distill Qwen 14B | DeepSeek | 16.5 | 2026-04 |
| Mistral Nemo | Mistral AI | 16.1 | 2026-04 |
| Mistral Nemo Base 2407 | Mistral AI | 16.1 | 2026-04 |