Skip to content

Published issue · September 7, 2026

Don't default everything to Astra

Three launches in four days. The trap is the $10/$50 stampede.

Indexed — Monday, September 7, 2026

Three model launches landed in four days. Two of them sit at $10/$50. One kept the cheap Flash sticker. Here’s how to route without lighting your agent bill on fire.

The Signal

Anthropic shipped Claude Fable 5.1 on September 1. Same list price as Fable 5: $10 per million input tokens and $50 per million output. The real change is cache reads at $0.25 per million — a 75% cut — which Anthropic says lands about 25% cheaper on typical workloads and up to about 45% on cache-heavy agent loops (Anthropic, Claude docs). Anthropic's own routing note: start with Opus 5 for most work; escalate to Fable when high-effort Opus still fails your evals.

Google shipped Gemini 3.8 Flash on September 2 at the same introductory sticker as 3.7 Flash: $0.75 / $3.75 through December 31, 2026, then $1.50 / $7.50 on January 1, 2027 (Google blog). That is a quality bump at the workhorse tier, not a new price cut. Google also says 3.8 “works harder” and can burn more tokens at higher effort — so sticker parity is not a free lunch on cost-per-task. (We already covered the 3.7 Flash intro rate on August 25. Don't re-migrate for the sticker alone.)

OpenAI released GPT-6 Astra on September 3–4. API id gpt-6-astra, standard short-context pricing $10 / $50, cached input $1, and any request over 272K input tokens reprices the whole job to $20 / $75 (OpenAI model card / pricing). It is rolling into ChatGPT paid plans, Azure, Bedrock, and GitHub Copilot (Pro+, Max, Business, Enterprise). Enterprise workspaces ship with Astra off by default.

What it means for you: the week looks like a frontier stampede. Your bill only moves if you change the default. Flash stays the volume lane. Opus stays the Claude default. Fable and Astra are escalate-on-failure models sitting on the same $10/$50 sticker — and Fable's cache math is friendlier for long agent loops that actually reuse context.

One catch from the social lane: HN operators are already arguing cost-per-task, not SOTA tables. Greg Brockman's line that got quoted hard this week — price per task is what matters — is the right scoreboard. If Astra or Fable finishes the job in fewer tokens, the higher sticker can still win. If it doesn't, you just bought a more expensive failure.

The Stack

A three-lane route you can set this week without ripping the stack.

Keep volume agents on Gemini Flash. Use Gemini 3.8 Flash (or stay on 3.7 if your evals are flat) for high-volume, well-scoped agent steps. Log tokens per successful task for one week before you call it a win — 3.8 can spend more thinking tokens even when the rate card matches.

Default Claude to Opus 5; escalate to Fable 5.1 on failed evals. Opus is half the list price of Fable ($5/$25 vs $10/$50). Turn on Fable where your agent actually reuses a big repo or brief prefix and the $0.25 cache reads pay for themselves.

Compare frontier runs in Cursor, not in vibes. Cursor (featured in the AI Tools Index) lets you pick the model per request. Run the same failing job through Sol or Opus, then Astra or Fable, and keep the cheaper model that still passes your check.

One decision today: name the single workflow that already fails quality on your cheap lane, and put a 30-minute block on the calendar to A/B it. Don't retune the workflows that already clear.

Prompt of the Day

Routing pages bury the decision. This prompt forces a lane call.

You are my model router. I will give you: (1) one concrete workflow, (2) my current model and its in/out $/MTok, (3) average monthly input and output tokens for that workflow, (4) candidate models with pricing, including cache-read rates if I have them, (5) my pass/fail criteria. Estimate monthly cost for each candidate. Flag any candidate where list price is higher but cost-per-successful-task might still be lower if token use drops. Tell me which lane to use as default, which to keep as escalate-on-failure, and the smallest test I should run before changing production. Ask for missing numbers before you calculate. End with one line: stay / test / switch.

Paste real numbers from the billing dashboard. The last line is what you send your finance lead.

Operators inside AI Freedom Lab are already arguing which jobs deserve the $10/$50 tier this week — faster than waiting for another launch post.

Tool of the Day

Cursor — the AI-first code editor built for pair programming with AI. Featured in the AI Tools Index, freemium, starts around $20.

Why it earns the slot on a routing week: model picker per request. You can push the same multi-file job through Gemini Flash, Claude Opus/Fable, and GPT-6 Astra without changing editors, then keep the cheaper pass.

Beyond that: codebase chat, Tab autocomplete, and inline edits with Cmd+K. The practical move for operators today is a side-by-side on one failing workflow, not a fleet-wide model swap.

Do this one thing: pick your worst agent failure from last week, run it twice in Cursor on two models, and keep the one that passes at the lower cost-per-task.

Quick Hits

Fable 5.1 and Mythos 5.1 share weights; Mythos stays on trusted-access programs for cyber and life sciences. Don't plan production on Mythos availability.

Gemini 3.8 Flash Cyber is Fairwind-only for trusted defenders. Same intelligence family, different mitigations — not your default Flash SKU.

Astra's cybersecurity Preparedness rating is Critical. Advanced defensive cyber workflows sit behind Daybreak-style trusted access. Treat that as access friction, not a free upgrade.

Requests over 272K input tokens on Astra reprice the entire request, not just the overflow. Cap context or budget the long-context lane on purpose.

Independent benches and vendor benches still disagree on who “won” the week. Artificial Analysis commentary in circulation this week has Fable ahead on some indexes while Astra's pitch leans cost-per-task. Run your own eval.

HN operators keep reporting overnight frontier runs that balloon into bloated monorepos. A model upgrade does not replace anti-bloat prompts and sub-agent commit gates.

Closing the loop

This week did not cut your Flash bill again. It raised the cost of a lazy default. Keep the cheap lane for volume, escalate to Fable or Astra only when a named job fails, and score the swap on cost per successful task.

Start with the workflow that already fails — not the one that already works.

Join AI Freedom Lab — builders shipping with AI: https://www.skool.com/aifreedomlab

Today's Tool of the Day: Cursor on the AI Tools Index: https://aitoolind.io/tools/cursor

Missed an issue? Browse the Indexed archive: https://indexednewsletter.com/archive

Prompt resources

Build from the prompt library.

Use these Indexed guides when an issue or tool mention needs a practical next step.

Text prompt pillar

Best ChatGPT prompts for work

Reusable prompts for briefs, research, decisions, emails, SOPs, and weekly operator workflows.

Read guide

Image prompt pillar

Best ChatGPT image prompts

Visual prompts for product shots, ads, thumbnails, social posts, and AI workflow content.

Read guide

Examples pillar

AI prompt examples

Copyable AI prompt examples for workflow documentation, operations, research, marketing, sales, reporting, and automation.

Read guide

Marketing workflows

ChatGPT prompts for marketing

Campaign, positioning, landing-page, email, and content prompts for practical marketing work.

Read guide

Small business workflows

ChatGPT prompts for small business

Prompts for offers, customer replies, SOPs, local marketing, hiring, and weekly business reviews.

Read guide

Research workflows

ChatGPT prompts for research

Prompts for source synthesis, competitor reviews, market scans, interview notes, and evidence checks.

Read guide

More from Indexed

Published issue

Who owns your agent loop?

OpenAI Agents API and Cursor Projects landed the same week. The question is not another model. It is who runs the standing work.

Read issue

Published issue

Your agent costs just got cut in half

Gemini 3.7 Flash runs at $0.75/$3.75 per million tokens through Dec. 31, 2026. Plus DeepSeek's open-source agent harness.

Read issue