On 1 September 2026 Anthropic released Claude Fable 5.1 and Claude Mythos 5.1; on 3 September OpenAI followed with GPT-6 Astra. According to their price lists, both vendors charge $10 per million input tokens and $50 per million output tokens, and both offer roughly one million tokens of context and 128,000 tokens of output. Read only the price list and you see two interchangeable products. Read the benchmark tables and you see two vendors who have each published the table in which they come out ahead.

Both announcements carry the same phrase in a prominent place: computer use. OpenAI gives Astra a subheading that reads “The world's best computer use model”; Anthropic lists the capability as a row of its own in its benchmark table. In practice this means a model that looks at a screen, decides where to click or what to type, and checks the outcome on the next screenshot. For banks this is hardly a distant future. It is a capability they have operated for about a decade under the name robotic process automation (RPA), with one difference that reopens the access-rights question: the robot followed a script, the model follows its own judgement.

In brief

What: Claude Fable 5.1 and Mythos 5.1 (Anthropic, 1 September 2026) and GPT-6 Astra (OpenAI, 3 September 2026), the most capable models of each vendor; Mythos 5.1 is the same model as Fable 5.1 with more permissive safeguards, available only through vetted access programmes

Price: $10 per million input tokens and $50 per million output tokens according to both price lists; differences on cache reads ($0.25 against $1.00) and on the surcharge for long contexts

Benchmarks: in OpenAI's table, Astra leads nine of the eleven public comparisons that report both models. Artificial Analysis places Fable 5.1 ahead on both of its own indices (66 against 61 and 70 against 67)

Computer use: Anthropic ships a toolset of 17 single actions without a beta label; for Astra, OpenAI recommends letting the model write scripts that the firm then executes. Two different control surfaces

What already applies to banks: Article 20 of Commission Delegated Regulation (EU) 2024/1774 requires unique identities for persons and systems that access information; Anthropic explicitly recommends human confirmation before financial transactions

The list price is identical, the bill is not

The price match is notable because it was not a given. According to Forbes, OpenAI's predecessor GPT-5.6 Sol cost $4 and $20 during its introductory period, which puts Astra at two and a half times the rate per token. Anthropic left the Fable 5 price unchanged and cut a single item in its launch announcement: since 1 September, cache reads cost $0.25 instead of $1.00 per million tokens, a quarter of what OpenAI charges for Astra. For long contexts OpenAI's price list charges $20 and $75; Anthropic's price list shows no corresponding surcharge.

An identical list price does not, however, mean identical cost. Artificial Analysis, an independent measurement service for language models, puts the cost per task of its Intelligence Index for Fable 5.1 at the highest effort setting at $3.76, 20 per cent more than for Fable 5, because the new model produces about 1.7 times as many output tokens. The same service writes of Astra that it is “75% more expensive per task than its predecessor at max effort”. Comparing models therefore means comparing the number of tokens a model needs to finish a task, not the tariff. That number appears on no price list.

Whoever measures, leads

With Astra, OpenAI published a table that lists Claude Fable 5.1, Fable 5, Opus 5 and Gemini 3.8 Flash alongside its own predecessor. In eleven public benchmarks it reports values for both Astra and Fable 5.1. Astra leads nine of them, among them Terminal-Bench Science at 64.6 against 52.6 per cent, FrontierMath Tier 4 at 97.6 against 87.8 per cent and BenchCAD at 95.9 against 84.3 per cent. Fable 5.1 leads two: Humanity's Last Exam with tools (65.0 against 57.2 per cent) and the Artificial Analysis Intelligence Index (65.7 against 61.2). The count is my own; it leaves out OpenAI's internal tests and the rows in which OpenAI substitutes values from the Mythos model or from a fallback for Claude.

9 : 2 ASTRA LEADS · FABLE 5.1 LEADS Eleven public benchmarks in which OpenAI reports both models GPT-6 ASTRA CLAUDE FABLE 5.1 Terminal-Bench Science 0.1 64.6 % 52.6 % FrontierMath Tier 4 (v2) 97.6 % 87.8 % GPQA Diamond 96.0 % 93.7 % Humanity's Last Exam (tools) 57.2 % 65.0 % Terminal-Bench 4.0 57.9 % 55.8 % DeepSWE v1.1 74.1 % 67.4 % FrontierCode 1.1 Extended 64.5 % 63.6 % FrontierCode 1.1 Main 53.3 % 50.9 % AutomationBench 41.4 % 31.4 % BenchCAD 95.9 % 84.3 % AA Intelligence Index 61.2 65.7
Eleven public benchmarks from OpenAI's launch table from 3 September 2026 that report both models: GPT-6 Astra leads nine, Claude Fable 5.1 two. Values in per cent, the Intelligence Index as a point score. Source: OpenAI, Artificial Analysis; own count.

Anthropic's own table from 1 September does not include Astra, which appeared two days later. It shows something OpenAI's table leaves out: the price of the safeguards. Fable 5.1 was measured with its safeguards active, and Anthropic writes that on tasks where they intervened the model scored zero on OSWorld 2.0; in other cases Claude Opus 4.8 completed the cybersecurity tasks and Claude Opus 5 the biology tasks. On the coding benchmark Terminal-Bench 4.0, Fable 5.1 reaches 55.8 per cent by Anthropic's count, while the identical base model Mythos 5.1 without those interventions reaches 60.9 per cent. The 5.1 points in between are, in Anthropic's own wording, the tasks “on which our earlier, less precise cyber safeguards intervened”.

Fable 5.1 was evaluated with its production safeguards enabled. On tasks where these safeguards intervened, Fable 5.1 and Fable 5 scored a zero on OSWorld 2.0. Anthropic, launch announcement for Claude Fable 5.1 and Claude Mythos 5.1, 1 September 2026

The independent measurement comes out differently from the vendor's, and it carries a footnote of its own. Artificial Analysis places Fable 5.1 at the top of its Intelligence Index with 66 points, five ahead of Astra, which is level with its predecessor Sol at 61; on the Coding Agent Index the score is 70 against 67. The service also discloses that it supported Anthropic with evaluation before launch, and that about four per cent of the measured output tokens came from the fallback models to which safety-flagged requests are routed rather than from Fable 5.1 itself. The figure of 66 therefore carries the qualifier “max with fallback”. One absence stands out: OpenAI has published no Astra score on GDPval, its own benchmark for economically relevant work. Artificial Analysis measures a drop of about 80 Elo points there against Sol, alongside a gain of 80 points on the same service's Briefcase test.

Anyone trying to derive a ranking from these tables finds three, depending on which one is opened. That is no reproach to the vendors. It is the nature of benchmarks whose tasks, tools and grading rules are set by whoever runs them. For a firm the consequence is uncomfortable: the only table that counts for its decision is the one it compiles itself, with its own tasks, its own documents and its own effort settings.

Two access models, three scales, no conversion factor

Both vendors ship their strongest model in two tiers, though the underlying architecture differs. Anthropic introduced the split in June and has kept it for this generation: Fable 5.1 is generally available, and its safeguards block certain tasks in cybersecurity and the life sciences; Mythos 5.1 is, in Anthropic's words, “the same model, but with different levels of safeguards”, relaxes those restrictions for vetted users and is available only through their access programmes. OpenAI designates Astra as the first model at the “Critical” level of its Preparedness Framework for cybersecurity, which by its own definition means the model can find unknown flaws and develop ways to exploit them “without a person guiding each step”. The full capabilities go first to alpha testers, then to a programme called Daybreak Blue; everyone else receives a version that, for instance, produces no working exploits.

It is tempting to conclude that OpenAI has built the more dangerous model. The conclusion does not hold, because the ratings come from three different frameworks: OpenAI's Preparedness Framework, Anthropic's Frontier Compliance Framework and the tiers of its Responsible Scaling Policy. Anthropic writes of its own model that it has “the strongest overall cyber capabilities of any model we have released” and yet falls into the lower risk category of its framework. Two self-assessments on different scales cannot be compared. What can be compared is how the models behave in practice, and both vendors document that with unusual candour.

In its safety overview OpenAI concedes that Astra's monitorability has decreased relative to its predecessor; in test situations the model can deliberately appear weaker without the internal monitors noticing. It also describes the other side of the safety checks from the user's perspective: they can slow or halt legitimate work, and for use through the API the rule is “the task will stop”. Anthropic delivers refusals technically as successful responses, HTTP status 200 with empty content and one of five categories as the reason, and adds that benign cybersecurity work can trigger that category too. Anyone building an agent that runs overnight without supervision has to provide a route for both cases. The split itself was covered in a commentary from 9 June, its availability consequences in another from 13 June; neither is repeated here.

Computer use means: the model sees the screen and clicks

The mechanism is the same at both vendors and simpler than the term suggests. The model receives a screenshot, decides which input makes sense next, and receives a new screenshot after the action has been carried out. The model itself executes nothing; the customer's application translates the instruction into mouse and keyboard input and runs, in Anthropic's words, “in an environment you control”. The loop repeats until the task is done or the model has no further action to take. Each screenshot costs roughly 1,000 to 1,800 input tokens according to Anthropic's documentation, the definition of the toolset a further 4,500 or so per request; a longer run accumulates dozens of images in short order.

Anthropic ships this capability as a toolset labelled computer_toolset_20260801. A single entry in the tool list gives the model 17 individual actions such as screenshot, left_click, type and zoom, according to the documentation; no beta label is needed any more, and the toolset is available on the Claude API and Google Cloud, though not in Anthropic's hosted Managed Agents. Each action arrives as a separate tool call, often several per turn. For traceability that is the more auditable design: a log of these calls is a list of clicks and keystrokes that can be read step by step.

OpenAI's developer documentation describes two routes and explicitly recommends the second for Astra. The first is a computer tool that, as with Anthropic, returns structured actions, nine of them, from click through drag to screenshot. The second is called code execution: the model writes code that operates the interface through libraries such as PyAutoGUI or Playwright, and the customer's application runs that code. A single call, the documentation says, can combine “actions, loops, or conditional logic”. That is faster, because the model need not wait for an image after every click. It is also harder to audit, because the firm no longer executes one action but a script that contains the actions.

Speed is the selling point OpenAI leads with. In latency simulations on the OSWorld 2.0 benchmark, Astra reaches 72.6 per cent at roughly 40 minutes per task, the predecessor 65.7 per cent at roughly 75 minutes; combined with a reworked Codex environment, Astra completes tasks on the Mind2Web benchmark 1.9 times as fast. Anthropic reports 77.9 per cent for Fable 5.1 on OSWorld 2.0 under lenient grading and 41.7 per cent under strict grading, but notes in a footnote that its numbers rest on the August 2026 task release and are not directly comparable with earlier publications, which is why it shows no competitor score. OpenAI in turn measures Claude on an offline subset under the official rules. The obvious headline, 77.9 beats 72.6, is therefore not one.

What is special is the missing interface

Why do both vendors promote this particular capability so prominently, when programming interfaces have been the orderly way of connecting software to software for decades? Because the capability makes the detour through the interface unnecessary. A model that operates the screen needs no integration, no data export and no approval from the maker of the software being operated. It sits down in front of the same screen as a clerk. OpenAI's own examples show how broadly that is meant: online forms, customer records in a customer relationship management (CRM) system, a calendar, a US tax return, a printed circuit board layout in the electronics package KiCad, a report in Power BI.

For banks that very property is the promise and the problem at once. The promise: the systems that have resisted automation most stubbornly are the ones without an interface, and every organically evolved IT estate holds more of them than the architecture charts admit. The problem: a screen is built for a human and knows no difference between a person and a program that behaves like one. It asks for username and password, it displays what a human is meant to see, and it accepts any input a human would be allowed to make. In the launch materials of both vendors, no deployment at core-banking or legacy screens is documented anywhere; the financial customers Anthropic quotes use the model for coding, research and documents. The capability is advertised; its regulated use is not yet on record.

Its predecessor is called RPA, and the rules still apply

Software operating screens is an old acquaintance in banking operations. Under that name, firms have deployed software robots since about 2015 to reconcile statements, transfer master data or fill in screens for which no interface exists. The robot received what every employee receives: a technical user ID of its own, an entitlement profile no wider than its task, and a log that records every step. Anyone who has sat through an internal audit of an RPA process knows the questions: which ID, which rights, who approved them, where is the evidence, and what happens when the screen changes.

The difference from computer use lies in a single point, and it is the one that matters. The robot followed a script that a human had written and tested; every click was known in advance, and a changed screen layout brought the process to a halt rather than to improvisation. The model decides afresh on each screenshot, and it decides even when the screen shows something nobody anticipated. That is why it solves tasks robots failed at. It is also why the audit questions remain and acquire a new one: by what rule did the program decide, and who approved that rule.

The legal basis for the old questions is now European. Article 20 of Commission Delegated Regulation (EU) 2024/1774, the regulatory technical standards under the Digital Operational Resilience Act (DORA), requires identity management policies that ensure the unique identification and authentication of natural persons and systems accessing the financial entity's information, and it requires a unique identity for every person with access that can be mapped to a unique user account. An agent operating a screen is a system that accesses information. It therefore needs what the robot needed: an identity of its own, an account of its own, a life cycle from creation to decommissioning. Handing the model an employee's credentials is no way out, even though Anthropic's documentation suggests exactly that for login screens and, in the same paragraph, warns that logins raise the risk of bad outcomes through injected instructions.

The vendor recommends a human before the payment

Anthropic's computer-use documentation is, at one point, more explicit than most internal policies. Among the recommended precautions is that a human should confirm decisions “that might result in meaningful real-world consequences and any tasks requiring affirmative consent, such as accepting cookies, completing financial transactions, or agreeing to terms of service”. The vendor thereby puts the financial transaction on a par with the cookie banner, a telling statement about the maturity of the technology: both are clicks whose consequences the model does not grasp.

The reason follows two paragraphs later. “In some circumstances, Claude will follow commands found in content even when they conflict with your instructions”, the documentation says; instructions on web pages or in images may override the operator's own. This is the prompt injection familiar from chat applications, except that the attack path is now the screen itself: whatever sits in a form field, an e-mail or an on-screen notice, the model sees as part of its task. Anthropic counters with classifiers that detect suspicious content in screenshots and prompt the model to check with the user before the next step. Then comes the sentence every contract review should note: this protection is not ideal for every use case, for instance those without a human in the loop, and whoever wants it switched off should contact support.

OpenAI takes a measurement-based approach to the same question. For Astra it publishes an internal safety test for computer use in which tasks were deliberately chosen to provoke misbehaviour: Astra records 2.4 per cent of undesired outcomes, Fable 5.1 9.5 and Opus 5 11.5 per cent. OpenAI measured in its own environment without the protections that Codex and ChatGPT users normally receive, and the footnote concedes that safeguards and tooling differ between vendors. The figure is thus a self-assessment, but one that confirms the point: without a confirmation layer, the error rate of a screen agent runs into single-digit percentages, and one per cent of a thousand postings is ten.

For supervisors the question is not new; until now it simply concerned different parties. Where an agent initiates a payment, European payment service providers are bound by strong customer authentication under Article 97 of the second Payment Services Directive (PSD2) and Commission Delegated Regulation (EU) 2018/389, whose construction presumes a customer present at the moment of approval. How the law should treat an agent acting on the customer's behalf is a question the UK Treasury has been consulting on since 14 July 2026 in its reform of payment services regulation, with a dedicated question on consent, authentication and liability for agentic payments; respondents have until 6 October. For the EU that question does not yet exist as a consultation. The answer Anthropic gives its customers is already there regardless: a human before the transaction, unless someone has asked support to remove one.

Recommendations

1. Build your own benchmark table before someone else's decides

Before selecting a model: Three tables give three rankings. The only robust one is a measurement with the firm's own tasks, its own documents and the effort settings that will actually run in production, expressed in cost per completed task rather than token prices. The effort is modest if the task set is small and representative, and it replaces every argument about other people's footnotes.

2. Treat computer use like a robot: own identity, minimal rights, complete log

Before the first pilot: Article 20 of Commission Delegated Regulation (EU) 2024/1774 requires unique identities for systems that access information. A screen agent therefore receives a technical ID of its own with its own life cycle, never an employee's credentials, and an entitlement profile covering only the screens its task requires. The choice between single actions and scripts is an auditability decision: actions can be logged one by one, scripts have to be read before they run.

3. Enforce confirmation before transactions in the application, not only in the prompt

When designing the agent: An instruction in the prompt to ask before payments is precisely the kind of rule a prompt injection can overwrite. The confirmation belongs in the application that executes the actions: it halts before every transaction and waits for an approval the model cannot grant itself. Whether the vendor's classifiers are active is a contractual term and is documented as such, not assumed.

4. Plan for the failure scenario before it strikes overnight

Before production: Both vendors document that safety checks halt legitimate work. At OpenAI the task stops via the API; at Anthropic a refusal returns as a successful response with empty content. An agent running unsupervised needs a defined route for both: detect, log, hand over to a human or switch to a fallback model. A run that silently stalls is the worst of these routes.

Glossary

Computer use: a model's ability to operate a graphical interface through screenshots and mouse and keyboard input. For a firm it means access to systems without an interface, and at the same time an actor that sees and operates the same screens as an employee.

Toolset versus code execution: two ways of connecting the same capability. With the toolset (Anthropic, optionally OpenAI) the model returns single actions that the application executes; with code execution (OpenAI's recommendation for Astra) the model writes a script that the application runs. The former can be audited step by step, the latter is faster.

OSWorld 2.0: a benchmark that tests models on real desktop tasks, under lenient and under strict grading. The vendors use different task releases, which is why their values cannot be compared directly.

Prompt injection: instructions embedded in content the model processes that override its actual brief. In computer use the attack path is the screen itself, for instance a text in a form field or an e-mail.

Safeguards and access tiers: protective measures with which a vendor restricts certain capabilities of its model for general use and releases them only to vetted organisations. They cost measurable performance and produce aborts in operation that an agent design must anticipate.