BotBeat
...
← Back

> ▌

Google / AlphabetGoogle / Alphabet
RESEARCHGoogle / Alphabet2026-08-05

Half of Small Language Model Benchmark Failures Actually Contained the Right Answer, Study Finds

Key Takeaways

  • ▸Over half of Gemma's benchmark failures on SQL tasks contained correct answers written in prose, not wrong reasoning
  • ▸Gemma's true CRM reasoning performance is approximately 50% rather than the measured 39.7%, with the gap attributable to tool-use formatting, not model capability
  • ▸When restricted to submitted answers, Gemma's performance is within 2 points of Gemini-3.1-flash-lite, contradicting the headline 28-point performance gap
Source:
Hacker Newshttps://neurometric.substack.com/p/the-glass-is-half-correct-half-our↗

Summary

A detailed technical analysis of Google's small language models on CRM benchmark tasks reveals that poor performance scores mask stronger underlying reasoning capabilities. The study tested three Google models—Gemini-3.6-flash, Gemini-3.1-flash-lite, and Gemma-4-E4B-it—on 340 Salesforce CRM tasks using two different tool interfaces, producing 2,040 agent rollouts. While raw benchmark scores suggested a stark 28-point gap between the small model Gemma and Gemini-3.1-flash-lite (39.7% vs 67.6%), deeper analysis uncovered a critical finding: over 50% of Gemma's measured "failures" actually contained correct answers—the model simply never invoked the required submit_answer tool.

In 171 Gemma trials on SQL tasks, the model completed its analysis and articulated the correct answer in natural language prose but failed to trigger the submission tool. When researchers examined these 250 prose non-submissions across the dataset, 44-52% contained the factually correct answer depending on how strictly the researchers evaluated the responses. This formatting gap, rather than reasoning quality, accounts for roughly half the observed performance differential. When the analysis restricted comparison to trials that actually submitted answers through the proper interface, Gemma performed within 2 percentage points of Gemini-3.1-flash-lite—a margin well within statistical noise. The hosted Gemini models showed near-perfect tool compliance, making them appear far more capable when benchmarks conflate instruction-following with reasoning ability.

  • Benchmark frameworks that depend on specific tool-calling conventions may systematically underestimate small model reasoning while overvaluing instruction-following discipline

Editorial Opinion

This finding exposes a methodological risk in LLM benchmarking: confusing 'wrong answer' with 'unconventional submission format.' The tight coupling between reasoning evaluation and harness compliance creates a misleading picture of model competence, particularly for small models that may prioritize cost and latency over rigid interface adherence. As small models proliferate in edge and embedded applications, the industry needs benchmarks that decouple reasoning quality from harness discipline—otherwise we systematically discount models that work correctly but differently.

Large Language Models (LLMs)AI AgentsMachine LearningData Science & Analytics

More from Google / Alphabet

Google / AlphabetGoogle / Alphabet
UPDATE

Google Cloud Enhances Filestore with Colossus Backend for AI Agent Workloads

2026-08-05
Google / AlphabetGoogle / Alphabet
FUNDING & BUSINESS

Google Elevates Demis Hassabis to Chief Scientist as DeepMind Reshapes Leadership

2026-08-05
Google / AlphabetGoogle / Alphabet
INDUSTRY REPORT

Can Reddit fend off a new wave of AI SEO spam?

2026-08-04

Comments

Suggested

Google / AlphabetGoogle / Alphabet
UPDATE

Google Cloud Enhances Filestore with Colossus Backend for AI Agent Workloads

2026-08-05
OpenAIOpenAI
UPDATE

Microsoft Defaults GitHub Copilot to OpenAI's GPT-5.6 Sol for Staff

2026-08-05
OpenAIOpenAI
PRODUCT LAUNCH

OpenAI Launches ChatGPT Work: An AI Agent for Knowledge Workers Across Enterprise Tools

2026-08-05
← Back to news
© 2026 BotBeat
AboutPrivacy PolicyTerms of ServiceContact Us