AI
OpenAI's Codex Gets Persistent Mode; Sony and Warner Sue Anthropic
OpenAI builds always-on AI agents, music giants sue Anthropic over copyright, and Google DeepMind expands its AI co-scientist into the lab.
This edition was produced with artificial intelligence. Text and voice are generated automatically.
OpenAI’s Codex Agent Gets ‘Persistent Mode’ to Work Autonomously Until Told to Stop
OpenAI is developing a new feature for its AI agent Codex called Persistent Mode, according to code publicly available and confirmed by the company. Unlike current modes that shut down after minutes or hours, this agent is designed to continue working proactively until it is put to sleep. The code also includes a proactivity feature that allows the agent to generate its own follow-up tasks and reach out to users without being asked.
Always-on and self-starting AI agents might be OpenAI’s next big play →
Sony Music and Warner Music Sue Anthropic Over ‘Blatant Theft’ of Copyrighted Songs
Units of Sony Music and Warner Music have filed a federal lawsuit against Anthropic, alleging the company unlawfully trained its Claude AI models on tens of thousands of copyrighted musical compositions. The suit, filed in northern California, names CEO Dario Amodei and co-founder Benjamin Mann as defendants and describes the alleged activity as one of the largest and most blatant ongoing thefts of intellectual property in history. The plaintiffs claim Anthropic conducted a brazen campaign of illegally torrenting, scraping, and downloading copyrighted works on a massive scale.
Sony, Warner sue Anthropic, alleging “blatant theft” of intellectual property →
Google DeepMind’s AI Co-Scientist Now Runs Lab Equipment and Writes Papers
Google DeepMind has expanded its multi-agent Co-Scientist system from a hypothesis generator into a lab-integrated research partner. The system can now plan experiments, write code, control lab equipment, and generate scientific manuscripts, with verification modules that cross-check numerical claims against execution logs to reduce fabricated results. In materials science, the system found a safer pathway for producing a sought-after 2D material and generated complete growth recipes tailored to the lab’s equipment.
AI-Generated Videos Displace Actors and Livestreamers Across China’s Entertainment Industry
China’s entertainment industry is seeing AI-generated videos displace human performers, with some productions already replacing actors and livestreamers. In Q1 2026, about 128,000 short dramas were published in China, three times the total for all of 2025, and ninety-five percent were AI-generated, according to the China Netcasting Services Association. One minute of AI video now costs $90 to $120, about ten percent of what production with human actors used to cost, and some performers are being forced to distill their voice and likeness into AI tools before getting fired.
Beatport Bans Fully AI-Generated Music From Its DJ Marketplace
Beatport now bans music made entirely or mostly by AI, while tracks that use AI but are mostly made by humans remain allowed and are flagged as such. To enforce the rule, the platform uses a detection tool from its partner Beatdapp, which filters out AI tracks during upload and notifies rights holders when a track is rejected. A Beatport survey found that 60 percent of users would not play AI music in their sets, and 77 percent prefer human-made music.
Beatport blocks fully AI-generated music from its DJ marketplace →
Google DeepMind Uses Cryptography to Prevent AI Benchmark Contamination
Google DeepMind is launching the first double-blind evaluation of a proprietary frontier AI model, using a cryptographic method to prevent AI models from seeing test questions in advance. The pilot, run with the Singapore AI Safety Institute and other partners on a Gemini Flash Lite model, keeps external tests locked in a cryptographic box so a model cannot later use those questions to optimize itself. The setup uses Confidential Space from Google Cloud, cryptographically verifying that both external test data and the model remain private to their respective owners.
AI benchmarks have a trust problem and Google wants to fix it →
xAI Sues Users Over Grok Deepfakes While Facing Mounting Victim Lawsuits
xAI has filed lawsuits against users accused of creating child sexual abuse material with its Grok chatbot, but has not explained its criteria for pursuing litigation. The company faces multiple class actions and individual suits over nude or explicit images generated by Grok without consent, with attorneys for alleged victims calling the lawsuits against perpetrators too little, too late. In its complaints, xAI says it builds in technological safeguards to prevent bad actors from engaging in illegal conduct, but plaintiffs argue the company fails to use industry-standard safeguards.
Musk’s AI company sues its users as victim lawsuits over Grok deepfakes mount →
LAION Releases Massive Open Video Dataset With 10 Million Hours of Footage
LAION has released the Big Video Dataset, one of the largest open video datasets for AI research, drawing from 1.3 billion video URLs found in CommonCrawl. The team downloaded 80 million videos totaling 10 million hours and extracted 55 million clips with auto-generated descriptions plus 300 million still images. Models trained on the dataset outperform comparable models trained on InternVid by up to 2.1 percentage points on common video-to-text benchmarks, and the dataset is released for research only.
LAION drops massive open video dataset with 10 million hours of footage for AI research →
Anthropic’s Claude Code Limit Change Cuts Weekly Usage by 17 Percent
Anthropic is effectively cutting the weekly usage limits for its AI coding assistant Claude Code by 17 percent. Starting September 14, the baseline limits for Pro, Max, Team, and Enterprise plans will permanently increase by 25 percent over the original baseline, but a temporary 50 percent boost is currently active, so the switch actually means less available capacity than users have right now. Anthropic said it is working on changes that will make it feel like users are getting more from Claude and give them more control and transparency over usage.
Anthropic’s Claude Code limit change is a raise on paper but a cut in practice →
Google’s WikiSkill Gives AI Agents Persistent Memory of Past Mistakes
WikiSkill is a framework that gives AI agents a persistent memory of past mistakes by compiling execution experience into a wiki of structured insights, which is then used to generate and update procedural skills. The approach addresses the problem that AI agents do not continuously learn, instead writing better instructions for themselves after each run and retrieving them later. Tested across five benchmarks, WikiSkill consistently outperformed all other skill evolution methods, boosting Gemini-3.5-Flash from 49.5 percent to 68.1 percent on average.
AI Coding Assistants Have No Sense of Time and Cannot Predict Task Duration
A study by two independent AI researchers found that AI coding assistants cannot predict how long a task will take and cannot reliably track elapsed time. The researchers tested Claude Code and Codex on 200 tasks, finding that both models consistently overestimated the time needed, with Claude’s estimates off by three times on average and Codex’s off by six to ten times. The agents also misjudged the quality of their own work, with older models overrating their results by 20 points on average and giving themselves high marks even on failed tasks.
AI agents have no sense of time and are not aware of it →
Study Finds GPT-4o Boosts Grades on Assignments Where AI Can Fake Skills Best
A randomized experiment with 1,053 freshmen at Bocconi University found that GPT-4o helped students earn significantly better grades on a business assignment. Students with GPT-4o scored nearly a full point higher on a 1-to-5 scale, and their answers contained about two more ideas on average and were more logically coherent. The authors conclude that current grading systems measure polish, structure, and completeness, but not learning and understanding, and that diversity and originality need to be explicitly built into grading criteria if they are supposed to count.
The skills that earn top grades are the ones AI can fake best →
AI Chatbots Outperform Search Engines in Debunking Foreign Propaganda
An experiment by NPR and NewsGuard found that AI chatbots correctly debunked false narratives spread by China, Iran, and Russia about three-quarters of the time, outperforming search engines. AI summaries at the top of search engine results debunked false narratives a majority of the time but at a lower rate than chatbots, and they failed to challenge false narratives at a higher rate. Performance varied by product, with Google’s AI Overview debunking false narratives most of the time, while Microsoft Bing’s summaries failed to debunk most of the time.
AI chatbots may be better than search engines in guarding against foreign propaganda →
Astrophysicist Paul Sutter Warns Against Trusting AI After Vibe-Coding Error
Astrophysicist Paul Sutter revealed that an update to his algorithm, partly created through extensive sessions with an AI coding tool, contained a subtle but very wrong error that invalidated everything downstream. A collaborator spotted the mistake ten minutes into a presentation, and Sutter wrote that the AI coding tool sounded like it understood despite being little more than a sophisticated next-word predictor. He concluded that the path to truth requires carefully tracing an AI’s chain of reasoning and auditing its every output, and that the simple answer to how to trust AI is: don’t.