The floor benchmark
Best Buy Bench
Which model actually works inside Imagine App — measured on real sales-floor questions with objectively checkable answers, through the app’s full tool stack. No judge model, no vibes.
The trade
Score vs. what it burns
Up and left wins. Cost axis is logarithmic — the gaps are bigger than they look.
measured $ per question →
Ranking
Every model, every cell
- 98%
Claude Sonnet 5Best answers
$0.03/question · 11.6s median
52 of 53 questions passedTop score. The one model that never flubbed a comparison or a product fact — worth the burn for the hard calls.
Easy18/18Medium17/17Hard17/18Search17/18Compare17/17Product Q&A18/18 - 96%
Gemini 3 Flash PreviewThe step-up
$0.01/question · 9.1s median
51 of 53 questions passedSonnet-class score at less than half Sonnet’s burn. Perfect on search. The value pick when the default fumbles.
Easy18/18Medium16/17Hard17/18Search18/18Compare17/17Product Q&A16/18 - 96%
Gemini 3.5 Flash
$0.04/question · 13.2s median
51 of 53 questions passedTies Flash Preview on every cell — at 3.6x the cost and slower. Still selectable, no longer recommended.
Easy18/18Medium16/17Hard17/18Search18/18Compare17/17Product Q&A16/18 - 96%
Claude Opus 4.8not offered in app
$0.11/question · 8s median
51 of 53 questions passedScores below Sonnet at 4x Sonnet’s price. Frontier weight buys nothing on floor work — cut from the app.
Easy18/18Medium16/17Hard17/18Search16/18Compare17/17Product Q&A18/18 - 87%
Gemini 3.1 Flash LiteThe default
$0.003/question · 3.3s median
46 of 53 questions passedThe economics champion: 9x cheaper and 3x faster than anything that beats it. Hard questions are its ceiling.
Easy16/18Medium16/17Hard14/18Search16/18Compare15/17Product Q&A15/18 - 83%
Mercury 2not offered in app
$0.003/question · 6.7s median
44 of 53 questions passedThe diffusion challenger: perfect on easy, strong on product facts — then craters on medium and comparisons. No vision.
Easy18/18Medium11/17Hard15/18Search15/18Compare12/17Product Q&A17/18
Disqualified
0% — and not because they’re dumb
Every request this app makes requires providers that neither store nor train on the conversation (Best Buy’s data terms demand it). These models have no compliant endpoint, so OpenRouter refuses every call — they can never answer here, at any price.
Nemotron 3 Ultra (free)
Every request refused: no provider offers it under zero-data-retention.
Gemma 4 31B (free)
Same wall: zero ZDR-compliant endpoints, so it can never answer in this app.
Methodology
How the sausage is measured
What it asks
53 questions a real floor employee would ask, across a 3×3 grid: three kinds of work (catalog search, product comparison, product Q&A) at three difficulties. Every question was grounded against the live Best Buy catalog the day it was written — “What’s the most popular 65-inch TV?” has one verifiable answer set, and the checks carry the evidence.
How it runs
Every model gets the identical harness: the app’s real system prompt, the full tool registry (catalog search, product analysis, comparison, web search, store stock, cart), the live Best Buy catalog, and the same zero-data-retention provider routing production uses. If it fails here, it fails in the app — same code path.
How it scores
Pass/fail per question, checked mechanically: the answer must reference the right SKU or contain the grounded fact (typography-normalized so “240 Hz” ≡ “240Hz”). No LLM judge anywhere. Cost is the USD actually billed per question; speed is the median end-to-end answer time including tool calls.
How to read it
53 questions means one question ≈ 2 points — treat gaps of a question or two as noise, not ranking. This is a dated snapshot (2026-07-08): the catalog drifts, models update, and results are harness-dependent by design — this page answers “how does it do in THIS app,” nothing broader.
Run 2026-07-08 · 53 questions · ~$11.02 total inference spend · cells: easy 18 / medium 17 / hard 18 · results never edited in place — re-runs get a new date.