Last reviewed
Correct answer: D. A faster, lighter model
Explanation
Match the model to the job. OpenAI's own latency guidance is blunt about the main lever: "smaller models usually run faster (and cheaper), and when used correctly can even outperform larger models". For a short, well-defined request — classify this, pull that field out, tidy this sentence — a lighter model answers sooner, costs less, and produces output you will struggle to tell apart.
The cost gap is not marginal. Inside a single generation the cost-optimised variant can run at a fraction of the flagship's price per token, so a task firing thousands of times a day is exactly where this decision is really made.
The trap is treating the most capable model as the safe default. It is only safe if latency and spend do not matter, and in anything people actually use, they do. The better habit is to start light, check that quality holds on your own inputs, and step up only where it visibly fails. That inverts most people's instinct, and it is the reason the lighter tiers exist at all.
Sources
“smaller models usually run faster (and cheaper), and when used correctly can even outperform larger models”
“GPT-5.6 Luna is designed for cost-sensitive, high-volume workloads.”
Practise 4 questions on this topic
Take GPT Models Overview — Timed Test (4 questions) — scored instantly, explanation for every question, no login.