Field Notes · AI Tutorials

GPT-6 Astra Effort Levels: Which One Should You Use?

GPT-6 Astra effort levels tested from low to Ultra to see which setting delivers the best results for the time and tokens.

Mark Kashef6 min readWatch the video
GPT-6 Astra effort levels thumbnail comparing low, medium, high, and Ultra reasoning settings

GPT-6 Astra effort levels run from low up through Ultra, and after testing all seven on the same research and build task, Mark Kashef found medium is the best default for day-to-day work. Low was not faster or cheaper than the higher tiers in his test. Ultra spent 22 million tokens without producing a meaningfully better outcome. High and Extra High land close to medium in quality, so they only earn the extra time on tasks that actually need the depth.

TL;DR

  • Astra Medium finished the full research-to-prototype task in 30 minutes on about 10 million tokens, the fastest and cheapest run of the seven tested.
  • Astra Low took 37 minutes and roughly 15 million tokens, slower and pricier than Medium, which undercuts the idea that low effort is automatically cheap.
  • Astra Ultra burned 22 million tokens across 42 minutes and still landed on a business idea no more developed than Medium's.
  • GPT-5.6 Sol on high needed only 32 minutes and 6 million tokens, testing OpenAI product lead Tibo's claim that Astra Low matches Sol High.

How GPT-6 Astra Effort Levels Actually Compare

GPT-6 Astra effort levels change how long the model reasons before it acts, not just how smart the output looks. Mark ran the exact same prompt across low, medium, high, extra high, max, and Ultra, plus GPT-5.6 Sol on high as a control, using Codex to spin up seven parallel threads from one prompt.

Every thread got the same job: research a SaaS opportunity for small service business owners on X and Reddit, pick one idea, sketch the plan in an Excalidraw diagram, then build a working prototype and website. Same context, same tools, same scope, so any gap in the results comes from the effort setting alone.

Join Early AI-dopters, 1,300 people mastering Claude Code and Codex. Mark breaks down the full test, plus more effort-level findings, inside the community.

Does Low Effort Really Match Sol on High?

No, not in this test. The claim came from Tibo, OpenAI's core products lead, who posted that Astra on low performs better than Sol on high for the same work.

Astra Low took 37 minutes and close to 15 million tokens, then came back with a cleaning-business idea that had no evidence pulled from X and nothing beyond what was asked. Sol High, running the older model, finished in 32 minutes on just 6 million tokens with a usable business plan and a website Mark called "unbelievably ugly" but functional. If low effort was meant to be the cheap option, it did not act like one here.

What Do You Get From Medium vs High Effort?

Medium and high produce nearly identical business plans and websites, so the extra time on high mostly buys a longer research pass. Here is how all seven runs compared on the same task.

Effort LevelTimeTokens SpentResult
Low37 min~15MThinnest plan, no X sources
Medium30 min~10MClean breakdown, fastest and cheapest
High~40 min~14MMost detailed research, similar plan to medium
Max46 min~10MMost polished website, referral-based moat
Ultra42 min~22MMost expensive, no clearer business idea
Sol High (control)32 min~6MCheapest overall, ugly but functional

Token spend does not climb in a straight line with effort either. High used less tokens than Low, and Max used less than half of Ultra while taking longer to get there. Mark's read: pick medium first, and only step up when a specific task shows it needs the extra pass.

Is Astra Ultra or Max Worth the Extra Tokens?

Rarely, based on this test. Ultra spent 22 million tokens, by far the most of any run, and produced a business idea (automating remodeling change orders) that was no more actionable than what Medium found in a third of the tokens.

Max came closer to earning its cost. It shipped the most polished-looking website of the group and a referral-based business moat that felt more thought through. But Mark still could not point to a result that justified 46 minutes and the token bill over just re-prompting Medium or High with more specific direction.

How Do You Test Multiple Effort Levels Without Manual Setup?

You ask Codex to do the setup for you instead of opening threads by hand one at a time. Mark's prompt told Codex to take one task, spin up several threads, rename each for its effort level, and run them all at the same time.

Codex created and labeled the threads, kicked off every run in parallel, and let Mark track tokens and time per thread from one screen. The same pattern works for comparing prompts, models, or reasoning effort settings on any task where you want a side-by-side instead of a guess, and it applies past Astra to the broader question of AI reasoning effort levels on any model that offers the setting.

Mark's Final Thoughts

Astra is not lazy at low effort, it is just inconsistent, and more tokens do not reliably buy a better answer. Medium on GPT-6 Astra is Mark's daily driver now, fast enough to not think about and thorough enough for real work. Save high, extra high, max, and Ultra for the tasks that actually demonstrate they need it, the same way GPT-6 Astra computer use is worth reaching for only once a task genuinely needs a hand on the screen.

GPT-6 Astra Effort Levels FAQs

What are the GPT-6 Astra effort levels?

Astra ships with low, medium, high, extra high, max, and Ultra. Each level controls how much the model reasons and iterates before answering, and higher levels burn more tokens and take longer, though not in a straight line.

Is GPT-6 Astra on low as good as GPT-5.6 Sol on high?

Not in Mark's test. Astra Low took 37 minutes and about 15 million tokens, more than Sol High's 32 minutes and 6 million tokens, and Sol still produced a usable, if uglier, result.

Does higher effort mean better results on Astra?

Not reliably. Extra High, Max, and Ultra all cost more time and tokens than Medium without a clear jump in output quality on the same research and build task.

Which Astra effort level should I use for daily work?

Medium. It finished in 30 minutes on about 10 million tokens, the fastest and cheapest run in the test, and produced a business plan and website as complete as High's.

How do you test multiple Astra effort levels at once?

Ask Codex to spin up several threads from the same prompt, rename each one for its effort level, and run them in parallel. Codex handles the setup and you compare the results side by side.

By Mark Kashef

All field notes