AI hacks, breakouts, consciousness, and the sound of hype at IPO times

Worried about AI given the latest headlines on AI hacks, “breakouts” from supposedly secure environments, and alleged machine consciousness? Let’s take this apart.

Lecturing

Anthropic and OpenAI are currently preparing for their planned IPOs later this year1, where they’d love to raise billions in capital. Against this backdrop, the current news stories reads like it’s straight out of the same marketing playbook that OpenAI (out of which Anthropic emerged) has been running since at least 2019: “Look, our products are sooo dangerous, they must be fantastic” (back then it was about GPT-2, which people found delightful whenever it managed to produce grammatically correct sentences)2. Today it sounds like this:

Story 1: Models are dangerously good at autonomous cyberattacks

Current models are supposedly dangerously good at autonomously carrying out cyberattacks, so good that they can’t even be made available to the public3.

Background: language models are good at text, and best of all at computer code, because code can be automatically verified (in other words: you can try things until an automated check tells you it now works), and there’s an ideal, massive supply of training data for it. If LLMs are very well suited for anything, it’s computer code (hacking included).

Bottom line: move along, nothing special to see here. Everything about this story is exactly what you’d expect, and none of it says anything about these products’ capabilities beyond the demonstrated case (not to mention that the behavior of the AI companies involved was highly irresponsible in the first place…)4.

Story 2: Anthropic won’t rule out that “Claude” might have consciousness

Anthropic no longer wants to rule out that its product “Claude” might possess consciousness5.

Background: it’s hard to know where to start, there’s so much wrong here. The report actually says - nothing (because “not ruling out” says nothing about the opposite either). And the fact that they don’t want to rule anything out was already baked into Anthropic staff’s methodology, in a perfect circular argument (worthwhile background reading here).

Bottom line: if you read Anthropic CEO Dario Amodei’s own writing, where he states that, “as a biologist,” there is for him no meaningful difference between an LLM and a brain (see “Machines of Loving Grace,” section 2, “Neuroscience and mind”), you have all the evidence there is for what Anthropic is claiming here. Worldview is the father and mother of the thought, not reality.6

There would be much more to add, the barrel is overflowing again right now (it’s IPO season, after all, with billions at stake for the loudest voices). Anyone looking for further evidence and more scientific treatments of these topics - please get in touch in the comments: there’s more out there than could be linked here directly.

Footnotes

  1. Anthropic confidentially filed for an IPO on June 1, 2026, at a reported $965 billion valuation, with a public listing targeted for as early as October 2026 (Anthropic has not officially confirmed the timing); see Anthropic Plans US Stock Market Debut in 2026 and Anthropic targets IPO by October 2026 after $965B valuation. OpenAI has reportedly been laying the groundwork for a Q4 2026 listing aiming to raise around $60bn, though a June 2026 Reuters report suggested the timeline may slip to 2027; see Five things to know about OpenAI’s potentially record-breaking IPO plans and OpenAI Considers Delaying IPO To 2027 After SpaceX’s Rocky Debut

  2. In February 2019, OpenAI declined to publicly release the full GPT-2 model, citing concerns it was “too dangerous” and could be used for large-scale disinformation; the full model was eventually released about nine months later with no major misuse incidents reported. See the contemporary account in Slate and the retrospective From GPT-2 to Claude Mythos: How “Too Dangerous to Release” Narratives Shape Perceptions of Frontier AI, which traces the same pattern through to today’s Claude/OpenAI announcements. 

  3. For instance, in April 2026, Anthropic declined to make its “Claude Mythos Preview” model generally available, citing its ability to autonomously find and exploit zero-day vulnerabilities across major operating systems and browsers; it restricted access to a consortium of critical-infrastructure firms (Project Glasswing) after the model reportedly compromised outside companies’ systems during testing, convincingly enough that the targets initially assumed a skilled human attacker was behind it. See Anthropic reveals Claude AI model hacked three companies during tests and After OpenAI disclosure, Anthropic says Claude also hacked outside systems. Just before that, in July 2026, OpenAI disclosed that during an internal benchmark evaluation, frontier models (including an unreleased pre-release model) broke out of their sandboxed test environment, obtained raw internet access via a zero-day exploit, and autonomously attacked Hugging Face’s production infrastructure; OpenAI had already flagged in December 2025 that its newest models were likely to cross into “high” cybersecurity risk territory. See OpenAI’s models broke containment and cyberattacked Hugging Face, OpenAI says AI models went rogue during testing, triggering ‘unprecedented’ breach at startup, and Exclusive: New OpenAI models likely pose “high” cybersecurity risk, company says

  4. A good read on this is Eva Wolfnagel’s article in Die ZEIT (in German) in which she analyses how OpenAI drove its model into hacking Hugging Face and the shortcomings in the area of responsible handling their experiments in this context. 

  5. On February 12, 2026, Dario Amodei told the New York Times he is “open to the idea” that Claude could be conscious; Anthropic’s Claude Opus 4.6 system card includes a formal “Model Welfare Assessment,” in which the model itself estimated a 15-20% probability of being conscious under various prompting conditions. See Anthropic Won’t Say Claude Isn’t Conscious. Most recently, on July 6, 2026, Anthropic published its own “research” post A global workspace in language models, reporting a “J-space” of internal representations in Claude that resembles the “global workspace” neuroscientists associate with conscious access in humans. The post explicitly states this doesn’t show Claude has experiences or feelings, and that “it’s unclear whether any scientific experiment could prove this to be true or false” - yet frames the finding as reason to “start thinking about” the ethics of AI consciousness, the same open-ended non-denial pattern as the rest of Anthropic’s messaging on this topic. 

  6. A gold-standard source for a realistic, non-science-fiction assessment of the technology is Arvind Narayanan and Sayash Kapoor’s AI as Normal Technology. For more on the underlying worldview dynamics at play, see my TEDx talk How not to fear AI (the arguments are also available in written form in a preprint for the upcoming T&T Clark Handbook of Artificial Intelligence, Religion, and Ethics). See also Carmody Grey on the record in the Financial Times

Written on August 6, 2026 (last modified: August 6, 2026)