The floor benchmark

Best Buy Bench

Which model actually works inside Imagine App — measured on real sales-floor questions with objectively checkable answers, through the app’s full tool stack. No judge model, no vibes.

2026-07-0853 questionslive Best Buy catalogproduction agent stack
Top scoreClaude Sonnet 598%$0.03/question
Best valueGemini 3.1 Flash Lite87%$0.003/question
Sweet spotGemini 3 Flash Preview96%$0.01/question

The trade

Score vs. what it burns

Up and left wins. Cost axis is logarithmic — the gaps are bigger than they look.

$0.003$0.01$0.03$0.10

measured $ per question →

Claude Sonnet 5Gemini 3 Flash PreviewGemini 3.5 FlashClaude Opus 4.8Gemini 3.1 Flash LiteMercury 2

Ranking

Every model, every cell

  1. Claude Sonnet 5Best answers

    $0.03/question · 11.6s median

    98%
    52 of 53 questions passed

    Top score. The one model that never flubbed a comparison or a product fact — worth the burn for the hard calls.

    Easy18/18
    Medium17/17
    Hard17/18
    Search17/18
    Compare17/17
    Product Q&A18/18
  2. Gemini 3 Flash PreviewThe step-up

    $0.01/question · 9.1s median

    96%
    51 of 53 questions passed

    Sonnet-class score at less than half Sonnet’s burn. Perfect on search. The value pick when the default fumbles.

    Easy18/18
    Medium16/17
    Hard17/18
    Search18/18
    Compare17/17
    Product Q&A16/18
  3. Gemini 3.5 Flash

    $0.04/question · 13.2s median

    96%
    51 of 53 questions passed

    Ties Flash Preview on every cell — at 3.6x the cost and slower. Still selectable, no longer recommended.

    Easy18/18
    Medium16/17
    Hard17/18
    Search18/18
    Compare17/17
    Product Q&A16/18
  4. Claude Opus 4.8not offered in app

    $0.11/question · 8s median

    96%
    51 of 53 questions passed

    Scores below Sonnet at 4x Sonnet’s price. Frontier weight buys nothing on floor work — cut from the app.

    Easy18/18
    Medium16/17
    Hard17/18
    Search16/18
    Compare17/17
    Product Q&A18/18
  5. Gemini 3.1 Flash LiteThe default

    $0.003/question · 3.3s median

    87%
    46 of 53 questions passed

    The economics champion: 9x cheaper and 3x faster than anything that beats it. Hard questions are its ceiling.

    Easy16/18
    Medium16/17
    Hard14/18
    Search16/18
    Compare15/17
    Product Q&A15/18
  6. Mercury 2not offered in app

    $0.003/question · 6.7s median

    83%
    44 of 53 questions passed

    The diffusion challenger: perfect on easy, strong on product facts — then craters on medium and comparisons. No vision.

    Easy18/18
    Medium11/17
    Hard15/18
    Search15/18
    Compare12/17
    Product Q&A17/18

Disqualified

0% — and not because they’re dumb

Every request this app makes requires providers that neither store nor train on the conversation (Best Buy’s data terms demand it). These models have no compliant endpoint, so OpenRouter refuses every call — they can never answer here, at any price.

Nemotron 3 Ultra (free)

Every request refused: no provider offers it under zero-data-retention.

0%

Gemma 4 31B (free)

Same wall: zero ZDR-compliant endpoints, so it can never answer in this app.

0%

Methodology

How the sausage is measured

What it asks

53 questions a real floor employee would ask, across a 3×3 grid: three kinds of work (catalog search, product comparison, product Q&A) at three difficulties. Every question was grounded against the live Best Buy catalog the day it was written — “What’s the most popular 65-inch TV?” has one verifiable answer set, and the checks carry the evidence.

How it runs

Every model gets the identical harness: the app’s real system prompt, the full tool registry (catalog search, product analysis, comparison, web search, store stock, cart), the live Best Buy catalog, and the same zero-data-retention provider routing production uses. If it fails here, it fails in the app — same code path.

How it scores

Pass/fail per question, checked mechanically: the answer must reference the right SKU or contain the grounded fact (typography-normalized so “240 Hz” ≡ “240Hz”). No LLM judge anywhere. Cost is the USD actually billed per question; speed is the median end-to-end answer time including tool calls.

How to read it

53 questions means one question ≈ 2 points — treat gaps of a question or two as noise, not ranking. This is a dated snapshot (2026-07-08): the catalog drifts, models update, and results are harness-dependent by design — this page answers “how does it do in THIS app,” nothing broader.

Run 2026-07-08 · 53 questions · ~$11.02 total inference spend · cells: easy 18 / medium 17 / hard 18 · results never edited in place — re-runs get a new date.

See the lineup these scores pickedHelp me choose a model