HeadFlash

AI

Apple Rushes Patches, Anthropic Expands, Senate Drafts AI Agent Bill

Apple rushes security fixes; California inks Anthropic deal; Senate bill creates trusted AI agent list; Anthropic poaches Google talent; CEO-Bench tests models.

Listen

Apple Rushes Security Patches Amid AI-Powered Hacking Risks

Apple released iOS 26.5.2, iPadOS 26.5.2, and macOS 26.5.2 with security fixes addressing vulnerabilities in the kernel, WebKit, and WebRTC. The updates had first been made available through the iOS 26.6, iPadOS 26.6, and macOS Tahoe 26.6 betas, meaning Apple decided to release them to the public earlier than originally planned. Apple told Reuters the move is a direct response to the ability of artificial intelligence to speed the development of malicious hacking tools, requiring a reduction in the time between when updates are first made public and when they are put into customers’ hands. Apple added that while there was no evidence that any of the newly patched vulnerabilities had been exploited, it still decided to release the fixes early to reduce the time attackers would have to exploit them. This accelerated patch cycle highlights a broader industry shift as AI lowers the barrier for creating exploit code, forcing faster reactions from even the most security-conscious vendors.

Apple accelerates security updates in response to AI-powered hacking risks →

California Signs Deal with Anthropic to Expand Government AI Use

California Governor Gavin Newsom signed a deal with Anthropic to expand government use of the company’s AI models. The agreement makes Anthropic’s Claude model available to state employees through a discounted, self-service platform with pay-per-use pricing rather than a flat enterprise license. Anthropic products already support several state services: the digital assistant Poppy, a public engagement platform for AI impact on jobs, California DMV customer service, Medicaid workflows at the Department of Health Care Services, and a cybersecurity partnership with the California Department of Technology that uses Claude to spot vulnerabilities in state code. Newsom stated that „AI should not replace the human work of government. It should help our workers move faster, solve problems more effectively, and deliver better results for Californians.” A designated official, Liana Given, said the state will seek similar discount deals with other AI companies and tech providers. Notably, the federal supply chain risk designation that the Pentagon placed on Anthropic in March 2025 did not come up during negotiations; Given said „it just didn’t come up.” Her department is formalizing recommendations for Newsom on that part of his executive order, expected by next month.

Newsom, Anthropic ink deal to expand government use →

Senate Draft Bill Proposes Federally Vetted AI Agent List

A new Senate draft bill, the Artificial Intelligence Access, Gatekeeper Exchange, and Nondiscriminatory Transfer (AI AGENT) Act, led by Sen. Mark Warner, D-Va., would establish a federally vetted list of AI agent software providers. End users of large online platforms with more than 50 million customers or subscribers per month would have the right to choose at least one AI agent provider that complies with security and identity standards developed by the Federal Trade Commission. The FTC would certify independent bodies to vet AI agent vendors for baseline protections in privacy, data security, and acting in the user’s interest. Providers must link each AI agent to its human operator’s identity and include built-in controls that let users grant or revoke permission for the agent to act on their behalf. The FTC cannot bar platforms from using non-compliant providers but can deregister violators from the list. Warner stated that „consumers deserve a real choice” and that the draft is a step toward a federal framework promoting innovation, protecting consumers, and maintaining U.S. leadership. Morgan Stanley estimated that nearly one-in-four (23%) Americans made purchases using AI over a 30-day period and that agentic shoppers could account for potentially hundreds of billions of dollars in online commerce by 2030. The bill acknowledges that AI agents can still be unreliable, make absurd purchases, leak sensitive data, or act contrary to a user’s interest, underscoring the need for safe solutions that verify accountable human identities behind AI activity.

Warner bill would create federally vetted list for secure, trustworthy AI agents →

Anthropic Poaches Google Gemini Stars Adler and Pritzel

Jonas Adler and Alexander Pritzel, both key contributors to Google’s Gemini AI model, are leaving Alphabet Inc. for rival Anthropic PBC. Adler worked on Google’s AI coding effort; Pritzel was involved in training AI systems. Their departures follow the loss of Nobel laureate John Jumper, who is also joining Anthropic, and star researcher Noam Shazeer, who is going to OpenAI. Alphabet shares closed down slightly after falling as much as 1.2% during trading Wednesday. The exits highlight pressure from two startups nearing IPOs, offering employees a payday by signing before an IPO. DeepMind engineers are nearly 11 times more likely to leave for Anthropic than the reverse, according to a 2025 industry analysis by venture capital firm SignalFire. Anthropic, which both competes and partners with Google, is exploring life sciences and healthcare applications. It recently raised a funding round at a $965 billion valuation and is considering an IPO as soon as fall 2025. AI researchers in the UK are often subject to lengthy non-compete agreements enforceable under British law; Jumper would likely not begin work at Anthropic until next year. A Google spokesperson pointed to DeepMind CEO Demis Hassabis’s remarks: „There’s a lot of talent movement between all the leading labs and we win our fair share of the top talent.”

Google hit by new AI brain drain as Anthropic poaches top Gemini talent →

Princeton CEO-Bench: Only 3 of 14 AI Models Turn Profit Running Simulated Company

Princeton University’s Z-Lab developed CEO-Bench, a benchmark that placed 14 AI systems in full operational control of a simulated SaaS company for 500 simulated days. Each agent received $1 million in seed capital and accessed a Python interface with 34 tools and 19 database tables, allowing them to set pricing, allocate R&D budgets, choose ad channels, scale infrastructure, and staff support, with a simulated social network for customer feedback. The environment introduced realistic delays and hidden signals. Of the 14 systems, only three finished above the $1 million starting balance: Claude Fable 5 posted about $47.15 million (a 47x return), Claude Opus 4.8 about $27.8 million, and GPT-5.5 about $21.3 million. Five models went bankrupt before day 500. A purely rule-based heuristic script with no language-model calls finished fourth with about $15.76 million, outperforming every other LLM. The researchers identified four capabilities that distinguished high performers: uncovering hidden signals, forecasting cash flow accurately, adapting quickly when competitors moved, and planning ahead with scenario-based thinking. Running models inside coding-agent frameworks degraded performance, suggesting domain-specific scaffolding is needed. Despite its dominant result, Claude Fable 5 declined to respond multiple times during its runs, citing safety constraints — the same model the US government ordered pulled offline in mid-June under export controls. The authors described the task as a „hell-difficulty” long-horizon exercise requiring „steering intelligence.”

Most AI Models Would Run Your Company Into the Ground, Princeton’s CEO-Bench Finds →