Skip to main content
agentic banker
  • Regulatory Outlook
  • News & Insights
  • Newsletter
  • About me
  • DE | EN
the agentic banker #12

The price falls, the supervisors close ranks

GPT-5.6 Luna 80 per cent cheaper, Muse Glimmer under Apache 2.0, a joint statement by the EBA, EIOPA and ESMA on frontier AI models, 92 per cent cash acceptance – what really mattered these past weeks.

Christian Schablitzki · 18 August 2026

Highlight

The price falls, the supervisors close ranks

Barely three weeks of summer break, and the price list for machine intelligence looks different. OpenAI has cut the cost of the cheapest model in the GPT-5.6 range by 80 per cent, Google is releasing Gemini 3.7 Flash at half its later price, and Meta has put a 30-billion-parameter model under an open licence that runs on a well-equipped laptop. For planning inside a bank, that moves a boundary: what was a capital decision in spring is drifting into the operating budget.

The same period also shows how quickly a figure loses its origin. The most widely shared pricing story of these weeks claimed that the flagship model Sol had become half as expensive. It did not. OpenAI's list price stands unchanged at 5 and 30 US dollars per million tokens; the discount belongs to a reseller – and it expires. Anyone budgeting model costs would do well to check who is actually setting the price.

On the other side of the table, the EBA, EIOPA and ESMA published a joint statement on 31 July on the ICT risks arising from frontier AI models, carrying forward their DORA oversight of critical third-party providers along the way. And while the digital euro is being built in Frankfurt, the ECB reports that merchants across the euro area are accepting cash more widely than they did two years ago – while acceptance of mobile payments jumped from 36 to 68 per cent over the same period. Three speeds that eventually meet inside the same institution.

LinkedIn Featured

When the agent watches the trader: Deutsche Bank's trade surveillance with Google Cloud

Deutsche Bank is building a trade surveillance system with Google Cloud that rests on a large language model and escalates anomalies in orders, trades and market movements to a human compliance officer. The project became known through a Bloomberg report; neither company has confirmed it officially. The article separates what is documented from what merely circulates: the frequently quoted reduction of false alerts by more than 25 per cent is a statement by the then chief technology officer with no published methodology, and the 40 per cent figure that travels alongside it belongs to a different bank. The second phase, moreover, touches a question that in Germany the works council helps decide. (Article in German)

read more →

Agentic AI

tools, skills & what's trending

Pricing: The floor is falling, but not where the headline says it is

On 30 July OpenAI cut prices for two of the three GPT-5.6 variants: Luna now costs 0.20 and 1.20 US dollars per million tokens instead of 1.00 and 6.00, a drop of 80 per cent. Terra falls by 20 per cent to 2.00 and 12.00 dollars. Google followed on 13 August, offering Gemini 3.7 Flash at an introductory $0.75 and $3.75 until the end of the year. A third figure travelled through the timelines, however: that Sol had become 50 per cent cheaper. It has not. OpenAI's list price remains 5.00 and 30.00 dollars; the discount is a limited-time promotion by the reseller OpenRouter. The list price therefore stayed what it was; what became cheaper was a reseller's promotion.

Infrastructure: 750 tokens per second, and still no price tag

Since 13 August OpenAI and Cerebras have been previewing Ultrafast, an API tier above the existing Fast mode. GPT-5.6 Sol runs on hardware that keeps the model weights in on-chip memory and reaches up to 750 output tokens per second. Cerebras puts the gap to competitors at eleven times Fable 5 and five times Opus 4.8 in Fast mode, calculated against throughput figures from Artificial Analysis; these are the vendor's own numbers, not an independent measurement. Access remains limited to a select group of customers for now, and OpenAI has published no price. Where response time forms part of the control itself – fraud detection inside a live payment flow, for instance – it decides whether an agent can still intervene before the booking.

Open source: One house opens its weights, the other holds them back

On 10 August Meta released Muse Glimmer, a 30-billion-parameter model under an Apache 2.0 licence, designed for always-on local agents. Quantised to four bits it stays below 20 gigabytes and runs on a well-equipped laptop. For institutions that want to solve data residency physically rather than contractually, that is the more remarkable news of the month, even though Meta does not market it that way. Z.ai moved in the opposite direction: GLM-5.3 uses the same base as its predecessor but doubled its hit rate for exploiting vulnerabilities through post-training (ExploitBench 24.4 to 54.4 per cent; for discovery alone, 77.2 to 84.5) and, working with security teams, uncovered more than 2,400 flaws across 269 projects. Z.ai is withholding the weights for now and has announced a safety review.

Banking & Regulation

what really counts now

Merchants in the euro area are accepting cash more widely again

92 per cent of businesses with a physical point of sale across the euro area accept cash, two percentage points more than in the previous survey in 2024. The ECB explicitly reads the figure as a recovery following the decline during and after the pandemic. A total of 8,205 businesses across all 21 euro area countries were surveyed between February and April 2026. The larger movement, however, sits a line further down: acceptance of mobile payments jumped from 36 to 68 per cent over the same period, while card acceptance held flat at 88 per cent. Cash is therefore not winning against digital payments; it has stopped losing, while beside it a channel has nearly doubled its reach in two years. One distinction matters for interpretation: what was measured is acceptance by merchants, not usage by customers – two different questions.

Three supervisory authorities, one joint statement on frontier AI models

On 31 July the EBA, EIOPA and ESMA published a joint statement on the ICT risks arising from frontier AI models. They call for a cross-sectoral, risk-based and consistently applied supervisory approach, drawing on the European Commission's action plan on cybersecurity and artificial intelligence as well as work by the ESRB, ENISA and the Single Supervisory Mechanism. The most concrete operational part sits further down the document: the statement carries forward ongoing and planned DORA oversight activities for critical ICT third-party providers. Institutions that source their models through such a provider will therefore find the consequence not in the AI rulebook but in the outsourcing register. The authorities expressly recommend using the statement as a basis for supervisory dialogue – an invitation worth accepting before the question appears in an audit report.

Signal & Noise

what your time is worth

  • An autonomous attacking agent compromised Snowflake's Jira through a CI/CD flaw – Wiz Research. The flaw did not come from a machine-generated fix: Copilot Autofix's contribution to the same pull request concerned a different file. Wiz expressly leaves open whether the faulty change itself was AI-assisted; Copilot co-authored the pull request and cleared it without findings. AI was therefore on both sides: Wiz's agent built the exploit, corrected itself after an initial failure and exfiltrated credentials. Snowflake confirmed the finding and closed the gap on the day it was reported. Essential reading for anyone who has given agents write access to the build chain.
  • Anthropic's text watermark and what it does to writing – John Gruber, Daring Fireball. Future Claude models are to shift word choice statistically in longer passages so the text stays identifiable as machine-generated; the trigger is the EU code of practice on transparency for AI-generated content. Gruber regards this as an intervention in the quality of the writing itself. Anthropic denies any loss of quality, citing internal testing and a controlled study with human raters; Gruber calls that reasoning euphemistic. For the implementation of transparency obligations on AI text, this is the uncomfortable part of the debate.
  • Encrypted reasoning traces can be recovered from LLM APIs – Alexander Panfilov and others, MPI for Intelligent Systems and ELLIS Institute Tübingen. The authors convert the encrypted reasoning traces of major providers back into plain text and find credentials and personal data inside publicly shared logs. Not yet peer reviewed, but pertinent wherever confidential matters travel through a model API.
  • OpenAI measures how enterprises use AI, using its own usage data – OpenAI. Instructive figures on the shift from assisting to executing. The underlying data, however, is telemetry from the company's own paying customers, not a randomised survey with a control group. Worth reading as a vendor's perspective, not as a market study.
  • Consolidated banking data for the end of the first quarter of 2026 – European Central Bank. CET1 ratio 16.27 per cent, NPL ratio 1.98 per cent, total assets 34.33 trillion euro. The reported return on equity of 2.44 per cent is a non-annualised quarterly figure – a number regularly misread as an annual return outside the release itself.

„It is wholly a confusion of ideas to suppose that the economical use of fuel is equivalent to a diminished consumption. The very contrary is the truth."

– William Stanley Jevons, The Coal Question, 1865
▸ Sources of this issue
  1. When the agent watches the trader: Deutsche Bank's trade surveillance with Google Cloud – Schablitzki Consulting
  2. Advancing the price-performance frontier with GPT-5.6 – OpenAI
  3. GPT-5.6 Sol – pricing overview and limited-time discount – OpenRouter
  4. Introducing Gemini 3.7 Flash – Google
  5. Accelerating GPT-5.6 Sol Ultrafast with OpenAI – Cerebras
  6. Previewing Ultrafast mode – OpenAI
  7. Introducing Muse Glimmer, an open agentic model – Meta AI Research
  8. GLM-5.3: Frontier coding with emergent cyber capabilities – Z.ai
  9. Cash remains most widely accepted payment method in euro area – European Central Bank
  10. EBA, EIOPA and ESMA call for enhanced governance and consistent supervision to mitigate ICT risks from frontier AI models in the EU financial sector – EBA, EIOPA, ESMA
  11. How an autonomous red-team agent compromised Snowflake's Jira through a CI/CD flaw – Wiz Research
  12. Anthropic's watermark text adulteration in Claude is a perversion of writing – John Gruber, Daring Fireball
  13. Stealing Reasoning Traces from Proprietary LLM APIs (preprint) – Alexander Panfilov et al., MPI / ELLIS Institute Tübingen
  14. From assistance to execution: How enterprises put AI to work – OpenAI
  15. ECB publishes consolidated banking data for end-March 2026 – European Central Bank
  16. The Coal Question, Chapter VII (cited from the 2nd edition, 1866) – William Stanley Jevons, 1865
← Back to newsletter overview
Christian Schablitzki

Christian Schablitzki

Strategy & Management Consultant · Agentic-AI expert for financial institutions

More than 20 years in investment banking and derivatives trading, followed by over 10 years as a consultant to financial institutions. Currently a Partner at Infosys Consulting in Germany. Certified in Google AI, Generative AI Leader (Google Cloud) and IBM RAG and Agentic AI.

LinkedIn profile →

Schablitzki Management Consulting

Pacellistr. 21, 85221 Dachau, Germany

info@schablitzki-consulting.de

© 2026 Schablitzki Management Consulting · Imprint · Privacy