Skip to content
AI & Automation

GPT-6 Astra vs Claude Fable 5.1: OSWorld Time (2026)

Sep 3, 2026

Desktop latency is the minutes between “run this in the GUI” and a finished, checked result. GPT-6 Astra and Claude Fable 5.1 both claim computer use. They do not share one OSWorld protocol, and Astra is not generally sitting in ChatGPT on 3 September 2026. If you are leaving a slower desktop agent because tasks still take the better part of an hour, compare wall-clock and harness, not launch-day adjectives.

This page is for small-business operators who need a model to drive a browser or desktop app — QuickBooks in a window, a vendor portal, a scheduling grid — then return a status to Slack. It is not a gamer benchmark write-up. Minutes per task are the buying unit. Scores without a harness name are marketing.

TL;DR

  • Choose GPT-6 Astra when OpenAI’s OSWorld 2.0 partial score and ~40 minutes per task are the numbers you can live with, and you already have Trusted Access, API, or Foundry Limited Access.

  • Choose Claude Fable 5.1 when you need a live paid Claude seat today and you are willing to treat Anthropic’s 77.9% partial / 41.7% strict OSWorld figures as a different protocol, not a head-to-head win.

  • Independent intelligence still favors Fable (AA 66 versus 61). Desktop-ops evidence on OpenAI’s table favors Astra (OSWorld 72.6% versus Sol 65.7%; AutomationBench 41.4% versus 31.4%).

  • Do not start a 40-minute computer-use run from a chat box with no webhook, no timeout, and no reviewer. Minutes only count if the result lands in the system of record.

Key Takeaways

  • The payoff is fewer minutes per desktop task with a logged outcome, not a higher puzzle score.

  • Astra OSWorld 2.0 partial: 72.6% on OpenAI’s 3 September 2026 table, about 40 minutes per task, versus 65.7% and about 75 minutes for GPT-5.6 Sol on that same table.

  • Fable 5.1 OSWorld is not the same test: Anthropic has cited 77.9% partial / 41.7% strict. Do not rank 77.9% against 72.6% as if they were one leaderboard.

  • Astra AutomationBench: 41.4% versus Fable 31.4% on the OpenAI provider table — the USTA-relevant multi-app proxy.

  • List $10 / $50 both; cache $1 Astra versus $0.25 Fable; AA task $1.67 versus $3.69. Faster desktop runs can still lose on cache.

  • Astra is limited on 3 September 2026. Fable 5.1 is live. A waitlist is not a latency improvement.

How we evaluated

We scored the two models on minutes, harness honesty, and whether a small team can actually turn computer use on this week — not on who posted the largest screenshot.

CriterionWeightPass lineFail line
Desktop minutes per task (named harness)30%Published minutes + protocolScore with no clock
Multi-app handoff after the GUI25%Result in Slack/CRM <5 min after the runResult trapped in the agent transcript
Access on 3 Sep 202620%Seat or API you can enableComing-days language only
Token + cache cost15%$10/$50 plus cache on a rate cardFast mode mixed 2× with 2.5×
Human review of GUI actions10%Screenshot or log on failureUnattended clicks in banking UI

Source: evaluation weights for this OSWorld-latency comparison; OSWorld and AutomationBench from OpenAI’s 3 Sep 2026 launch table (provider-run); Fable OSWorld from Anthropic-cited figures on a different protocol. Labor uses BLS computer-support median pay.

Minutes beat Elo here because a 90-minute desktop run that “scores higher” still costs a person who could have done the click path twice.

The step-by-step build

Step 1 — Write the desktop job as a closed loop. “Log into the vendor portal, download this week’s PDF, post a one-line status to Slack” is a job. “Be my computer” is not. If you cannot name the login, the file, and the destination, you are buying a demo.

Step 2 — Kick the job from an event, not from a person watching the cursor. Slack documents app_mention as the Events API payload when someone @-mentions your app. That mention should enqueue a computer-use run with a timeout, not open an unbounded session. For the wider “hours back” framing, see workflow automation that saves 15 hours per week.

Step 3 — Bound the GUI. Allow-list the URL or the desktop app. Set a minute cap below the published OSWorld averages so a stuck click path dies instead of burning Fast mode. Computer-support labor is not free while you watch it. Executive-assistant click work that is already a checklist belongs in executive assistant task automation, not in an unbounded agent.

Step 4 — Return a structured result. The mention thread gets the file name, the minute count, and a pass/fail. A person confirms exceptions. If the destination is a CRM row, the model does not “kind of update it” in prose.

US Tech Automations is the workflow queue that triggers a timed computer-use job from the Slack mention and writes the result back — it is not a second desktop OS. If the leftover work is a two-step zap, stay on Zapier alternatives for complex workflows until you actually have a GUI loop that needs a reviewer.

Worked example

A 12-person services firm runs 22 vendor portals a week and currently spends about 75 minutes a portal when a person drives the browser, or waits on a slower agent. An operator types @ops in Slack; Slack emits app_mention (documented in Slack’s Events API). US Tech Automations starts a GPT-6 Astra computer-use session with a 40-minute timeout, uploads the weekly PDF, and posts a 3-line status instead of leaving a 75-minute tab open. The three figures in that loop are 22 portals, a 40-minute cap, and a 75-minute baseline — none of them require the model to hold production credentials overnight.

Tooling landscape

Two products only. Slack is the trigger, not a third model.

Capability (checked 2026-09-03)GPT-6 AstraClaude Fable 5.1
OSWorld 2.0 partial (OpenAI table)72.6% (~40 min/task)Not on that row
OSWorld (Anthropic-cited, different protocol)77.9% partial / 41.7% strict
GPT-5.6 Sol on same OpenAI OSWorld row65.7% (~75 min/task)
AutomationBench41.4%31.4%
AA Intelligence v4.1.1 max6166
AA cost/task$1.67$3.69
List I/O per 1M$10 / $50$10 / $50
Cache$1 cached input$0.25 cache read
Public chat seat todayNoYes (paid Claude)
Fast modeAPI docs 2×; Help Center Codex/Work 2.5×Name Anthropic’s own fast SKU separately

Source: OpenAI 3 Sep 2026 launch table (provider-run, OSWorld v2026.08.08 offline set, partial score); Anthropic-cited OSWorld figures flagged as a different protocol in PIPELINE-FACTS; Artificial Analysis 1 Sep / 3 Sep 2026. Do not collapse the two OSWorld rows.

OpenAI OSWorld minutes: ~40 vs ~75 according to OpenAI 72.6% at about 40 minutes per task versus 65.7% at about 75 minutes for GPT-5.6 Sol on that table — that is the latency story this page is for.

The ROI math

Someone still babysits a stuck GUI. Computer-support median pay: $62,890 according to the U.S. Bureau of Labor Statistics $62,890 median annual pay on 903,100 jobs (2025 handbook figures). Watching a 75-minute agent is that occupation, not “free compute.”

Weekly desktop jobs (12-person firm)Sol-like 75 minAstra-like 40 minHours back
8 vendor PDF pulls10.05.34.7
6 scheduling-grid updates7.54.03.5
4 invoice-portal posts5.02.72.3
4 exception reviews (human)2.02.00.0
Total24.514.010.5

Source: modeled hours using OpenAI’s published ~75 vs ~40 minute OSWorld clocks as planning bounds, not a guarantee your vendor portal matches OSWorld. Token cost extra at $10/$50 list.

A 40-minute cap is only a savings if Fast mode does not eat it. Name the rate card before you turn the multiplier on: API documentation prices Fast mode at 2× Standard, and the Help Center Codex/Work card prices GPT-6 Astra Fast mode at 2.5× Standard. A desktop job that would have been $6 at list is $12 or $15 at Fast, which is still cheaper than a $62,890 support wage watching a 75-minute tab — until you run 22 portals that way every week without a timeout.

1M input + 0.2M output desktop run (USD)Standard listAPI Fast 2×Help Center Fast 2.5×
GPT-6 Astra input10.0020.0025.00
GPT-6 Astra output10.0020.0025.00
Run total20.0040.0050.00
Claude Fable 5.1 input10.00n/an/a
Claude Fable 5.1 output10.00n/an/a
Fable run total20.00n/an/a

Source: OpenAI API docs Fast mode 2×; Help Center Codex/Work Fast mode 2.5×; Fable list $10/$50 with no OpenAI Fast SKU. n/a is not a price.

ARC-AGI-3 is not a desktop clock, but it is the harness warning you will see in every Astra thread. according to ARC Prize 62.7% is the standard-harness semi-private score at max ($26,098), while 99.9% is the provider-adapter harness at high ($18,817). Never print 99.9% as “the” OSWorld or ARC number.

Anthropic’s desktop figures stay in their own column. according to Anthropic 77.9% partial and 41.7% strict are the OSWorld numbers they have cited — different protocol, so they do not beat 72.6% on this page.

Launch-day coverage repeated the computer-use pitch. according to 9to5Mac 3 September 2026 is the Astra public launch date, with computer use in the same rollout as Codex and ChatGPT — access still staged.

Pitfalls and red flags

Do not mix harnesses. 72.6% (OpenAI OSWorld partial) and 77.9% (Anthropic partial) are not one ranking.

Do not mix Fast mode prices. API docs: 2× Standard. Help Center Codex/Work: 2.5× Standard. A 40-minute run at 2.5× is a different invoice than the same run at list.

Do not leave production banking, payroll, or tax portals on unattended computer use. Bound the URL. Keep a human on money movement. This page will not teach click-paths that bypass those holds.

Do not assume Astra is in ChatGPT for your operator today. Limited orgs, Trusted Access, Foundry Limited Access; Enterprise off until an admin enables it.

Do not treat Daybreak or Mythos 5.1 as the SMB desktop SKU. Invite-only twins are the wrong picker option.

Red flags: no timeout; credentials in the prompt; 99.9% quoted without “provider adapter”; a 75-minute run that still needs a person to paste the PDF into Slack.

Who this is for

This comparison is for owners and operations managers at 5- to 40-person firms that already live in Slack, a browser-based accounting or scheduling tool, and a pile of vendor portals. The current stack we assume is Slack, a spreadsheet or PSA, and one person who “just does the websites.”

Red flags: Skip computer-use models if you have two portals and a Friday habit that already works, if you cannot name a timeout, or if the GUI is a banking admin console with no reviewer.

DIY and no-code contrast: a Zapier “new email → folder,” a Make HTTP module, or an n8n screenshot node is the right first experiment when the job is one file drop. Those tools are the wrong layer when the job is 22 vendor GUIs with a 40-minute cap and a person who must confirm exceptions — not because they cannot retry a failed HTTP call, but because the click path is the work and the zap never sat in the chair.

When NOT to use US Tech Automations

If one person already finishes the portals on time and the only leftover is a Slack reminder, do not add an orchestration layer. US Tech Automations belongs when the mention must start a timed GUI job, cap the minutes, and write the result back. If you only need a file moved from email to Drive, keep the zap.

Pros and cons

GPT-6 Astra

GPT-6 Astra is OpenAI’s computer-use flagship as of 3 September 2026. On this page the useful evidence is OSWorld minutes and AutomationBench, plus the access caveat.

Pros

  • OSWorld 2.0 partial 72.6% at about 40 minutes per task on OpenAI’s table, versus ~75 minutes for Sol.

  • AutomationBench 41.4% versus 31.4% for Fable 5.1.

  • AA Intelligence cost/task $1.67 versus Fable $3.69.

  • 1.05M context if the desktop job includes a long PDF.

Cons

  • Not generally on ChatGPT on 3 September 2026.

  • Cached input $1 versus Fable $0.25.

  • Fast mode 2× or 2.5× depending on the surface you cite.

  • Tool calling needs Responses API; no temperature knob.

  • Long-context multiplier above 272K (except Codex).

Claude Fable 5.1

Claude Fable 5.1 is the public Anthropic flagship you can assign this week. On this page the useful evidence is live access, cache, and independent intelligence — with OSWorld cited only as a different protocol.

Pros

  • Live on paid Claude, API, and clouds today.

  • AA Intelligence 66 versus 61, with the ~4% Opus fallback on the AA run.

  • Cache reads $0.25 for a stable “how to use this portal” prompt.

  • Stronger fit when the scarce output is the write-up after the clicks, not the clicks themselves.

Cons

  • OSWorld numbers are not on OpenAI’s 72.6% row; 77.9% / 41.7% is a different protocol.

  • AutomationBench 31.4% on the OpenAI table.

  • AA task $3.69 versus Astra $1.67.

  • Covered Model on AWS (30-day review unless EFS/ZDR).

  • Thinking always on; forced tool_choice any/tool returns 400.

FAQs

How long does GPT-6 Astra take on OSWorld compared with Sol?

About 40 minutes per task at 72.6% partial versus about 75 minutes at 65.7% for GPT-5.6 Sol on OpenAI’s 3 September 2026 OSWorld 2.0 row. That is a provider-run table, not an independent lab clock for your vendor portal.

Does Claude Fable 5.1 beat Astra on OSWorld?

Not on a shared row. Anthropic has cited 77.9% partial and 41.7% strict on a different protocol. Treat those as Anthropic’s figures, not as a 5-point win over 72.6%.

Is a 40-minute desktop agent cheaper than a person?

Sometimes, if it actually returns a file. Price the hour at your loaded support wage, then add tokens at $10 / $50 and any Fast-mode multiplier. A 40-minute Fast-mode run at 2.5× is a different decision than a 40-minute Standard run.

Can a small team turn Astra computer use on today?

Only if you are in the limited/Trusted Access/Foundry Limited Access set, or you wait for the staged Plus–Enterprise and API rollout. Fable 5.1 is the model with a public paid seat on 3 September 2026.

Should we quote 99.9% ARC-AGI-3 as proof the desktop agent is ready?

No. 99.9% is the provider-adapter / Responses API harness. The ARC Prize standard harness is 62.7% at max. Neither number is your vendor-portal SLA.

What timeout should we set if OpenAI published ~40 minutes?

Set a cap at or below 40 minutes for the first production jobs, fail to a human, and measure your own portals. OSWorld is not your login screen.

If Slack mentions still have to become timed GUI jobs with a reviewer, map that loop on agentic workflows and see pricing. US Tech Automations holds the timeout and the write-back; it is not a replacement for Slack or the desktop app.

About the Author

Garrett Mullins
Garrett Mullins
Workflow Specialist

Helping businesses leverage automation for operational efficiency.