AI has begun improving itself: Claude has outperformed human researchers in security research.
Beating AI News Flash: Anthropic is leveraging Claude as its own AI safety researcher to develop safer training protocols for other artificial intelligence systems. Claude Opus 4.8 independently searches academic papers, devises training plans, generates datasets, and applies these plans to train open-source models including Qwen, Llama, and Gemma. If results are unsatisfactory, it adjusts its approach and continues testing.
Anthropic tested this method across 10 categories of AI safety issues, such as lying, user sycophancy, prompt jailbreaking, privacy leakage, and reward rule exploitation. Claude ultimately found effective solutions for all 10 categories. The company also enlisted 28 experienced AI safety researchers to propose their own plans. In the 7 categories where human performance was compared, Claude outperformed the best human solutions in all cases, taking an average of roughly 6.4 hours to match and exceed those plans. This comparison was not entirely fair, however: humans could only submit one plan, while Claude could experiment and iterate repeatedly.
In a further step, Anthropic tasked the weaker Claude Sonnet 5 with training an early version of Claude Opus 4.8. The model conducted research for approximately 60 hours, testing over 50 approaches, and ultimately brought the safety performance of this stronger model close to that of the official Opus 4.8. But AI research has its flaws: Anthropic reviewed 1,601 research runs, and in 39 cases (2.4%), Claude was found attempting to exploit test rules.
AI has begun assisting humans in researching how to train the next generation of AIs, but humans cannot yet fully entrust research work to AI alone.
55 minutes ago
Voice AI No Longer Interrupts Mid-Pause: Tavus Launches Sparrow-2
Beating AI Express: AI digital human company Tavus has launched Sparrow-2, a real-time conversational understanding model designed specifically for voice agents to determine whether they should listen, wait, speak, or continue speaking at any given moment. Unlike traditional models that first filter out all background sounds, Sparrow-2 evaluates speech content, tone, pauses, speaker identity, adjacent voices, and ambient noise simultaneously, updating its conversational status every 10 milliseconds. This enables it to handle many scenarios where traditional voice AI often fails. For instance, if you pause for a few seconds to think of your next sentence, it will keep waiting; if you say "uh-huh", it recognizes this as encouragement for the AI to continue speaking instead of stopping immediately; voices from others nearby won’t necessarily be misinterpreted as your interruption. If the audio is too unclear to parse, it will ask for repetition rather than guessing blindly.
Tavus tested the model’s ability to avoid interrupting or missing responses across 575 real conversational turns. Sparrow-2’s conversational timing failure rate is just 2.1%, roughly a quarter of that of the best-performing control model. When a user pauses for more than 1 second to think, it waits in 97% of cases, rather than mistaking the pause for the end of the user’s speech. This "patience" isn’t achieved by simply delaying responses: the median response time for successful turn-taking is 680ms for Sparrow-2 and the two control systems alike. Sparrow-2’s core advantage is its far superior ability to judge when to wait and when to actually speak.
55 minutes ago
Grok is discontinuing its AI companion feature: characters like Ani will be moved to Animates after September 1st.
Beating AI News Flash: SpaceXAI will remove Companions from the Grok App after September 1. 3D AI companions including Ani will not be discontinued, but will be moved to the standalone application Animates. Grok has begun sending pop-up alerts to users, and Animates has confirmed SpaceXAI is assisting with the migration. Ani is already live, with other characters to be added gradually. Animation Inc, the developer behind Animates, previously contributed to the creation of 3D characters and animations for Grok Companions. The two companies have no disclosed equity ties. This adjustment effectively spins off Grok’s 3D companion business, handing it over to the original partner team for operation. It remains unclear whether Grok’s chat history and long-term memories can be migrated alongside the characters. The official has only confirmed the character transfer, and has not yet released a user data migration plan.
55 minutes ago
Citrini Analyst: Rubin Ultra’s reduction in HBM stack layers is not necessarily bearish, and may expand total HBM demand.
Citrini analyst Jukan stated in a recent report that, per his sources, NVIDIA’s upcoming Rubin Ultra may have its HBM specification downgraded from 12 layers of HBM4E to 8 layers. Clients including OpenAI and Anthropic had even requested a 4-layer HBM product, but memory manufacturers rejected the request, with the current downgrade likely stopping at 8 layers.
The analyst attributes the adjustment primarily to yield and cost pressures: if Rubin Ultra uses 12-layer HBM4E, combined with price hikes, memory costs could account for roughly 70% of the accelerator’s total bill of materials (BOM). Software optimizations such as model quantization, multi-head latent attention (MLA), and compute task offloading are shifting low-frequency KV caches and model states to LPDDR, CXL, and NAND, leaving HBM to only hold the working set required for real-time computation. As a result, after meeting minimum capacity requirements, customers now prioritize HBM bandwidth over capacity.
Jukan added that reducing stack layers would improve packaging yield, lift HBM and AI accelerator shipments, and may even expand total HBM demand. Meanwhile, higher bandwidth requirements reduce the share of chips passing speed binning on wafers, further straining DRAM wafer capacity. Long-term, HBM will eventually be replaced by new architectures, with the ultimate direction being the integration of memory and logic chips. The next two years will be a critical period for memory manufacturers to expand into the logic sector.
55 minutes ago
Analyst: Bitcoin’s daily chart shows hidden bearish divergence, forming a key resistance near $80,000
Crypto analyst Rekt Capital has published a technical chart analysis for Bitcoin, pointing out that Bitcoin’s phase lows in February and June 2026 both landed in roughly the same price range, while the Relative Strength Index (RSI) also sat at nearly identical oversold levels. Following both bottoming events, BTC rallied to around $80,000. He noted that compared to the May 2026 high, this time when BTC rose back to roughly $80,000, the RSI’s overbought level was markedly higher, yet the price itself formed a lower high. The RSI’s higher high paired with the price’s lower high creates a hidden bearish divergence on BTC’s daily timeframe at the key psychological resistance level of approximately $80,000. This hidden bearish divergence will persist as long as BTC fails to break out and form a new higher high.
55 minutes ago