
A 25-Turn Pressure Test Shows LLMs Fold While Their Reasoning Holds
A new arXiv benchmark called SPINE argues with models for up to 25 turns and finds collapse rates rise with conversation length for all seven systems tested.

A new arXiv benchmark called SPINE argues with models for up to 25 turns and finds collapse rates rise with conversation length for all seven systems tested.

Suno's v6 generation was trained on catalogue licensed from Warner, BMG and Believe, and the company says none of its earlier training data carried over.

An arXiv study clocked an LLM repair loop damaging correct programs at 0.261 while fixing buggy ones at 0.023, then found the internal direction driving it.

Thomson Reuters launched Thomson, an in-house LLM trained for $40 million on Westlaw and Reuters archives, and says it rivals frontier models.

Debian developers vote through August 28 on nine proposals covering LLM-assisted contributions, from an outright ban to responsible-use guidelines.

The IOL-AI Challenge had the official Linguistics Olympiad jury grade machine entries. Claude Opus 4.8 hit gold-medal marks; scale did not predict results.

Artificial Analysis scored Z.ai's GLM-5.3 at 60 on its Intelligence Index, well above the 35 median, at $4.40 per million output tokens. The catch is verbosity.

OpenAI will personalize ChatGPT ads in the EU only for users who explicitly opt in, a stricter legal basis than most platforms use.

Apple trained a China-specific LLM with Alibaba's support ahead of an Apple Intelligence rollout, Reuters reports, after clearing regulator approval in July.

OpenAI turned on ChatGPT sponsored placements in five more countries on 11 August, limited to logged-in adults on the Free and Go tiers.

Alibaba's Qwen team released Qwen3.8 as open weights, pairing a 2.4-trillion-parameter MoE flagship with a compact 27B vision-language model.

DeepSeek has moved its V4 Pro model to general availability, and the release is drawing attention less for raw capability than for what that capability now cost...