Relay-Bench Reveals Frontier LLM Blind Spot: Multi-Domain Reasoning Collapses to 43%
Key Takeaways
- ▸Frontier LLMs including GPT-5.5 show dramatic performance collapse (40-point drop) when reasoning must chain across multiple domains rather than staying within a single domain
- ▸Relay-Bench exposes a critical blind spot in LLM evaluation: strong single-domain benchmark scores may not reflect fragility when problems require cross-domain reasoning
- ▸The benchmark remains unsaturated with GPT-5.5 at 43.3%, suggesting substantial room for improvement but also indicating fundamental challenges in how current LLMs approach complex multi-step reasoning
Summary
Researchers have introduced Relay-Bench, a new unsaturated benchmark that measures frontier language models' ability to reason across multiple domains in a single prompt. The results expose a dramatic capability gap: while frontier LLMs achieve strong performance on single-domain tasks, their accuracy collapses when reasoning must chain across different domains. OpenAI's GPT-5.5 (xHigh), the leading model tested, scores just 43.3% on Relay-Bench—a stark contrast to its performance on domain-specific tasks.
The benchmark consists of 2-13 chained subproblems spanning visual reasoning, coding, mathematics, information extraction, problem-solving, general knowledge, and data analysis. All models are given full access to tools including code execution and web search. The test set introduces complexity through prompt encoding and context bloat to simulate real-world reasoning demands. The 40-point performance drop (from typical single-domain scores around 83% to 43% multi-domain) suggests that frontier LLMs may be overfitting to narrow task distributions rather than developing robust cross-domain reasoning.
This finding carries significant implications for AI deployment and safety. If leading models struggle severely when reasoning spans multiple domains—a common requirement in real-world applications—current benchmarks may be masking fundamental limitations that only emerge in complex scenarios.
- Real-world deployment of frontier LLMs in mission-critical applications may be riskier than current benchmarks suggest if cross-domain reasoning is required
Editorial Opinion
Relay-Bench provides a crucial reality check for the AI industry. The dramatic performance collapse when reasoning spans domains challenges the narrative that frontier models have solved core reasoning problems. If GPT-5.5 drops 40 points moving from single-domain to multi-domain tasks, our models may be far narrower than marketing suggests—a sobering reminder that published benchmarks don't capture the full picture of model capabilities or limitations.


