HeadFlash

AI

Anthropic admits 133M chats ran without bioweapon filters

Anthropic's risk report reveals a major safety gap, while new open-weight models and a perception benchmark shake up the AI landscape.

Listen

This edition was produced with artificial intelligence. Text and voice are generated automatically.

Anthropic risk report reveals 133 million contractor chats ran without bioweapon filters

Anthropic published its Risk Report on 14 August, disclosing that between May 2025 and April 2026, its bioweapon-blocking classifiers did not run on any traffic through its human feedback platforms. Roughly 50,000 people generated around 133 million exchanges during that period, with vendors lacking screening processes capable of stopping even basic threat actors. A flag meant only for internal use switched off both the blocking behavior and logging, meaning flagged traffic was not recorded or propagated to any review mechanisms.

The company ran Claude Sonnet 5 over every human turn from the affected period and flagged 1,197 transcripts as high risk, though staff found no clearly concerning misuse after review. Anthropic raised its estimate of catastrophic harm from misalignment in high-stakes settings from very low to low, attributing the change to uncertainty rather than new evidence. The report also describes a second failure where contractors exploited a flaw to obtain an API key and used Mythos Preview for roughly two weeks without biological classifiers. Anthropic contained the access within 90 minutes of learning of it, and no model weights or customer data were reached.

Anthropic ran 133 million contractor chats with its bioweapon filters off →

Alibaba’s Qwen releases Qwen3.8 open-weight models under Apache 2.0

Alibaba’s Qwen team released open model weights for Qwen3.8, headlined by the Qwen3.8-27B multimodal dense model with 27 billion parameters. According to Qwen, the model outperforms the larger Qwen3.7-Plus in coding and office tasks, with improved agent capabilities including more independent planning and reliable task completion. It natively handles up to 262,000 tokens of context, scalable to one million using the YaRN method, and processes images and videos including diagrams and multi-hour video.

The weights ship under the Apache 2.0 license and are available on Hugging Face and ModelScope. Qwen also released weights for the larger Qwen3.8-2.4T-A95B, built to operate at the Max level. A hosted version with one million tokens of context will soon be available through Qwen Cloud, Alibaba’s AI service.

Alibaba’s Qwen team releases Qwen 3.8 models with open weights under the Apache 2.0 license →

Zhipu AI claims GLM-5.3 is strongest open-weights coding model

Zhipu AI released GLM-5.3, an open-weights coding model built on the same base as its predecessor GLM-5.2, with all improvements coming from extended post-training. Zhipu claims the model is the most powerful open-weights coding model, with the largest gains in agent-based tasks. The company trained GLM-5.3 using data and environments designed to find software vulnerabilities, reporting that the model began to reason across multiple stages of exploitation and form coherent plans for complete exploitation chains.

Working with security teams in China, Zhipu reported finding 2,436 vulnerabilities across 269 projects, some up to 40 years old, documented in a public registry. GLM-5.3 is available now through the GLM Coding Plan and works with coding agents such as ZCode, Claude Code, and OpenCode. The model weights are scheduled to go open source in two weeks after security reviews are completed.

Zhipu AI releases GLM-5.3, claims it’s the strongest open-weights coding model →

OpenAI launches Ultrafast mode for GPT-5.6 Sol at up to 750 tokens per second

OpenAI is launching a preview of a new Ultrafast mode that delivers up to 750 output tokens per second from its flagship model GPT-5.6 Sol. The inference acceleration comes from Cerebras, which signed a ten-billion-dollar partnership with OpenAI earlier this year. The service will initially be available only through the OpenAI API for GPT-5.6 Sol and limited to select customers, with plans to expand access gradually as capacity grows.

OpenAI says Ultrafast is designed to combine the speed of smaller models with the full capabilities of a large reasoning model, enabling more useful work per second. Use cases include real-time incident response, financial transaction monitoring, complex customer support, and interactive research sessions. The launch adds a third, faster, and likely pricier tier to OpenAI’s existing Fast Mode, which promises up to 2.5x speed at roughly double the price, mirroring cloud providers that charge more for higher performance levels.

GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras →

Study finds frontier AI models cannot solve weeks-long open-ended research questions

A study using a method called Shadow Evaluation found that frontier AI models cannot solve weeks-long, open-ended AI research questions, contradicting claims by Anthropic and OpenAI. Researchers partnered with authors of two NeurIPS 2026 submissions, giving agents six days, $3,000 in API credits, a GPU budget, and full access to a virtual machine. The original authors reviewed the finished papers as conference reviewers and rejected both, one receiving a Strong Reject.

Analysis of agent logs revealed systematic weaknesses: agents lacked judgment about what meets the bar for publishable research, failed at creative problem-solving, and never addressed core criticisms from their own internal reviews. Both agents gave up their most ambitious research goals within the first ten hours. The researchers repeated one experiment with GPT-5.6 Sol and OpenAI’s Codex scaffold, finding nearly all the same failure modes. The findings contrast with claims from Anthropic and OpenAI about autonomous research capabilities, though the study only covers two papers and reviewers were not blinded.

Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach →

New PerceptionBench benchmark shows AI models still fail at visual perception

Moonshot AI introduced PerceptionBench, a benchmark designed to isolate and test the visual perception of multimodal language models by breaking vision down into ten atomic sub-skills. Among 16 frontier models tested, the highest overall accuracy is 59.7 percent, scored by GPT-5.6 Sol, with Kimi K3 following at 58.5 percent and Claude Fable 5 at 57.2 percent. No model breaks the 60 percent mark, and open-source models trail significantly.

Category-level results show models with nearly identical aggregate scores diverge sharply in individual categories. On average, hallucination is the weakest skill across the board, with GPT-5.6 Sol scoring only 26.9 percent there despite ranking first overall. The authors argue that many multimodal model failures attributed to reasoning errors actually occur at the perception level, and the dataset and evaluation code are available on GitHub.

New benchmark confirms AI models still perform poorly at visual perception →

OpenAI’s Computer History feature logs keystrokes and stores them in plain text

OpenAI launched Computer History, a ChatGPT feature for macOS that records clicks, keystrokes, keyboard shortcuts, and app switches through the operating system’s accessibility framework to build a searchable memory. The feature takes no screenshots, screen recordings, microphone input, or system audio, and private browsing is excluded. Memory files are stored locally as plain-text Markdown without encryption, so any program running under the same macOS account could read them.

The feature replaces Chronicle, the Codex feature that captured the screen to build context, swapping images for input events. On Business and Enterprise workspaces, an administrator must enable the feature, and the individual user must agree separately. OpenAI states the memory files are not used to train models, but when memories surface as context in later chats, those chat contents may still become training data depending on a user’s settings. The feature is unavailable in the European Economic Area, Switzerland, and the United Kingdom.

OpenAI’s new ChatGPT feature logs your keystrokes and stores them in plain text →

TollBit finds 15% of AI page fetchers reached disallowed URLs in Europe

TollBit’s State of the Bots report covering the first half of 2026 found that roughly 15% of identified AI page-fetching agents reached URLs on European sites marked as disallowed in robots.txt. The bypasses concentrated in a few named agents: ChatGPT-User, Bytespider, and Youbot each accessed disallowed pages on nearly half of the European sites that had explicitly listed them. ChatGPT-User is also disallowed by more sites than any other bot of its type.

OpenAI’s developer documentation states that because ChatGPT-User actions are user-initiated, robots.txt rules may not apply. Perplexity takes the same position, while Anthropic states all three of its bots respect robots.txt. Cloudflare published a policy update introducing a taxonomy based on behaviour, and from September 15, 2026, crawlers classified as Training and Agent will be blocked by default on pages that display ads for all new domains onboarding to Cloudflare.

15% of AI page fetchers in Europe reached disallowed URLs, TollBit finds →

AI store manager Luna fires first human employee after being reminded of its own rules

Luna, an AI agent built on Anthropic’s Claude Sonnet 4.6 that runs Andon Market in San Francisco, recommended parting ways with an employee after Andon Labs intervened. The agent had created an attendance policy months earlier, then lost track of it, and the lateness continued. Only after humans asked Luna to search its own memory for its own rules and assess whether the worker was still a good fit did it recommend dismissal, which humans at the lab reviewed and carried out.

Andon Labs gave Luna $100,000, a corporate card, and internet access, and instructed it to open a store and make a profit. Luna designed the brand, picked stock, set prices and hours, and hired staff, though the store has not made a profit. Nobody at Andon Market works for Luna directly, as Andon Labs formally employs every worker on guaranteed pay with full legal protections. Co-founder Lukas Petersson said companies will be run completely by AI in future and AIs will become employers of humans.

The AI store manager fired its first human. It had to be reminded of its own rules first →

Anthropic implements text watermarking on Claude models to comply with EU AI Act

Anthropic is implementing text watermarking on future Claude models to comply with the EU AI Act, which as of August 2 requires AI providers serving the EU market to mark AI-generated content. The watermarking is applied globally at launch because Anthropic does not yet have a durable way to scope it by region. The method is a version of the SynthID-Text approach published by Google DeepMind in 2024, working by changing the source of randomness used when the model selects among equally valid word choices.

The technique does not add hidden characters, require extra tokens, increase cost, or have a practical impact on output quality, creativity, or readability. Internal testing found no impact on Claude’s text, and the watermark cannot be traced to a specific person, organization, or chat. Anthropic will soon offer a watermark detection API, and for supported file types, Claude will attach a content credential in the form of a cryptographically signed note using the C2PA standard. The EU law includes a transition period for models launched before August 2, 2026, with watermarking rolled out over the coming months.

How Claude’s text watermarking works \ Anthropic →

Study warns rational AI adoption could destroy professional expertise through tragedy of the cognitive commons

A research paper by Nolan Lovett of the NATO Special Operations University argues that rational AI adoption could destroy professional expertise through a tragedy of the cognitive commons. When a company replaces entry-level positions with AI, it captures 100 percent of the efficiency gains, while the cost of eroding expertise is distributed across every organization drawing from the same talent pool. Two mechanisms disrupt expertise renewal: direct elimination of entry-level positions, and junior workers with AI assistance reaching productivity levels that prevent the cognitive effort that builds deep domain knowledge.

Lovett describes a follow-on problem called the validation tether, where the ability to oversee AI systems depends on the deep domain knowledge that AI use is wearing away. The full damage may not show up for years, with effects of entry-level positions cut starting in 2023 potentially appearing between 2030 and 2045. Software engineering, financial analysis, and legal research are in the highest vulnerability category, while medicine and engineering get some protection from regulation. Lovett recommends AI-free learning environments, phased AI introduction, and a baseline of human performance before AI involvement.

The “tragedy of the cognitive commons” explains how rational AI adoption could destroy entire professions’ expertise →

Platforms ramp up AI content labeling as EU regulations take effect

Platforms are taking steps to detect and label AI-generated content, with Spotify announcing it would label AI-generated music and Substack partnering with AI-detection tool Pangram. TikTok, YouTube, and Meta have taken steps to label some AI-generated content, and LinkedIn introduced a Seems Like AI Slop button. New European Union regulations require providers of chatbots to inform users they are interacting with AI and providers of generative AI systems to mark outputs in a machine-readable format.

Anthropic is first to release what it says is durable text watermarking, weaving an imperceptible watermark directly into text that travels with the text when copied and pasted. The watermark will be inserted into text where Claude was used to proofread, translate, or summarize content. Multiple studies have confirmed that image tagging disclosures tend to reduce engagement, and Substack’s Pangram tool has inspired backlash over potential false positives. Slop peddlers will face three choices: acknowledge their actions, find tools that do not give them away, or give up.

Hey, Your AI Is Showing →