• Newsletter
  • Posts
  • 🔵Four flagship models in eight weeks. Here's how to actually choose.

🔵Four flagship models in eight weeks. Here's how to actually choose.

Fable 5, GPT-5.6, Opus 5, Gemini 3.6 — what changed, and a routing framework that survives the next release

👋 Welcome

Between June 9 and July 24, Anthropic shipped three models and OpenAI shipped three. Google shipped two more. If you set your AI stack in May, it is now out of date — and probably overpriced.

The latest frontier models are remarkably close on headline benchmarks—and surprisingly different in real work.

Claude Fable 5 can sustain difficult analysis for hours. GPT‑5.6 Sol is particularly strong at turning an assignment into a working product or finished deliverable. The newer Claude Opus 5 offers much of Fable’s capability at half the price. Gemini remains compelling for Google-connected and multimodal work, while Kimi K3 is pushing open models closer to the frontier.

The practical conclusion: there is no single best model. There is a best model for each type of work.

Where each model genuinely wins

Claude Fable 5 — the hardest, longest jobs.

Use it for: multi-day migrations, due diligence across hundreds of pages of filings, the analysis behind a board decision.

Fable 5 is designed for long, ambiguous assignments: the type of work where the model must absorb many documents, decide what matters, maintain a line of reasoning and revise its own conclusions.

In business, that could mean:

  • analysing dozens of customer interviews and identifying the real causes of churn;

  • reviewing financial reports, market data and management assumptions before an investment decision;

  • comparing contracts and policies across jurisdictions;

  • investigating an unfamiliar codebase or completing a large migration;

  • reconstructing a product or interface from screenshots and visual material.

Independent testing placed Fable ahead of GPT‑5.6 Sol in overall analytical quality and complex knowledge work, although the gap was small. Anthropic also reports strong performance in finance, document reasoning, charts and long-running software work.

But Fable is expensive: $10 per million input tokens and $50 per million output tokens. It also requires 30-day data retention, and some sensitive cybersecurity, biology and chemistry requests may be routed to another Claude model.

Business verdict: use Fable when a better answer could materially change an important decision—not for routine work.

Claude Opus 5 — the practical business default.

Released July 24 at half Fable's price, it posts 79.2% on SWE-bench Pro versus Sol's 64.6%, and 30.2% on ARC-AGI-3 versus Sol's 7.8% — a near-4x gap on problems that cannot be pattern-matched from training data. Use it for: day-to-day engineering, novel analytical problems, anything where requirements are underspecified.

Opus 5 arrived shortly after Fable and may be the most important model for ordinary professional work.

Anthropic describes it as approaching Fable-level intelligence at half the price: $5 per million input tokens and $25 per million output tokens. Early evaluations show particular strength in:

  • root-cause analysis and difficult debugging;

  • financial modelling and numerical reasoning;

  • legal review and contract redlining;

  • complex research;

  • code review;

  • checking and correcting its own work;

  • questioning flawed requirements rather than following them blindly.

Unlike Fable, normal Opus 5 access does not carry the special 30-day retention requirement. Its safety filters are also expected to intervene far less frequently.

Business verdict: before paying for Fable, try Opus 5. For strategy, finance, legal review and careful decision support, it may provide the best balance of judgment, cost and reliability.

GPT-5.6 Sol — the execution specialist.

Independent evaluation put Sol one point below Fable 5 on overall intelligence at roughly one-third the cost per task. Use it for: CI/CD and DevOps agents, competitive research sweeps, client-facing decks.

Sol’s strength is not simply answering questions. It works particularly well inside an agent environment where it can inspect files, use tools, browse, write code, run tests and create finished artifacts.

Good business applications include:

  • building an internal dashboard from a written requirement;

  • automating a manual reporting process;

  • investigating and fixing a production bug;

  • creating and testing a functional prototype;

  • producing a board presentation or analytical spreadsheet;

  • researching a market and converting the findings into an action plan;

  • testing a website across different screen sizes;

  • carrying a multi-step assignment from brief to deliverable.

At launch, GPT‑5.6 Sol led Artificial Analysis’s coding-agent evaluation and came within one point of Fable on its broader intelligence index—at roughly one-third of the estimated cost per task. Its API price is $5 per million input tokens and $30 per million output tokens.

Hands-on reports also point to an important difference: Fable may reason more deeply, but Sol can be easier to collaborate with and more willing to adjust when a plan is not working.

Business verdict: Sol is the strongest general choice when AI must not only analyse the work but actively complete it.

Where the other models fit

Gemini 3.6 Flash — the volume workhorse. Same input price as its predecessor, 17% lower output price, and roughly 17% fewer output tokens to do the same work. It beats Google's own flagship on most coding and agentic tests. Use it for: anything Tier 1 or Tier 2 at scale.

Gemini 3.1 Pro remains attractive for large multimodal inputs and workflows connected to Google Search, Maps or Workspace. It is useful when an assignment combines documents, images, video and audio. However, the Pro model remains a preview product, so businesses should be cautious about making it the sole foundation of a critical production process.

Kimi K3 is strategically interesting because it combines frontier-level ambitions, a one-million-token context window and open weights. It may matter for organisations that want greater control over deployment or long-term model independence. But it is still new, and its governance, infrastructure and production maturity require additional due diligence.

For classification, extraction, translation, summarisation and routing at scale, frontier models are usually unnecessary. GPT‑5.6 Luna now costs only $0.20 per million input tokens and $1.20 per million output tokens. Gemini Flash-Lite offers a similar high-volume lane.

What's on the board

All carry roughly a one-million-token context window. OpenAI cut Terra and Luna prices on July 30 — Luna by 80%. If you budgeted in early July, re-run your numbers.

The four-tier routing framework

Route on task shape, not on leaderboards. Three properties decide the tier: how long the task runs, how expensive a wrong answer is, and how much of the bill is output tokens.

Tier 1 — High volume, low stakes. Classification, extraction, tagging, routing, translation, first-pass summaries. Models: Luna, Gemini 3.6 Flash, Haiku. Examples: triaging a support inbox, normalising free-text survey responses, tagging 4,000 product records.

Tier 2 — Everyday professional work. Meeting notes into decisions and owners, stakeholder updates, SQL queries, code review, slide outlines, first drafts. Models: Terra, Sonnet 5, Gemini 3.6 Flash. This is where 70–80% of real business volume lives — and where most teams accidentally overpay by defaulting to a flagship.

Tier 3 — Hard, consequential, one-shot. The forecast a VP will present. The architecture decision. The analysis where being wrong costs a quarter. Models: Opus 5, GPT-5.6 Sol. At this tier, per-token price is irrelevant next to the value of getting it right.

Tier 4 — Long-running autonomy. Agents that plan, execute, verify and self-correct over hours or days. Models: Fable 5, Sol Ultra. Anthropic's own guidance is to use this tier selectively and judge it by completed tasks, not launch-day scores.

AI Automations: 🌐[https://cmasterai.com]

Contact us at [[email protected]]

Reply

or to participate.