Lookonchain APP

App Store

Hugging Face enables Claude Code and Codex to directly participate in RL training.

51 minutes ago

Beating AI Express: Hugging Face has open-sourced a Multi-harness reinforcement learning (RL) framework that enables direct integration of existing coding agents such as Claude Code, Codex, and OpenCode into RL workflows for training its open-source models, eliminating the need to build separate training environments for each agent. Large language model firms like OpenAI have long trained models within their proprietary agent environments. Hugging Face has now made this capability a general open-source tool, allowing other teams to leverage existing agents for RL directly. Models no longer need to be trained in only one agent: a single training run can cycle through Claude Code, Codex, OpenCode, and Mini-SWE-Agent to complete tasks, helping the same model adapt to different prompts, tools, and execution logic, and reducing the "specialization bias" that only works well with a single harness. The team tested the framework on a 2.6-billion-parameter small model from Liquid AI. When using identical model weights but switching agents, task pass rates dropped from approximately 62% to 33%. After training across four agent environments, the average pass rate rose from 42.2% to 54.2%, with more balanced performance improvements across different agents compared to training in a single environment.

Relevant content

Mistral Large 4 Launched: 1.05 Trillion Parameters, Claims to Be Europe and the U.S.’s Strongest Open-Source Model

Beating AI News: French AI firm Mistral has launched its new flagship model, Mistral Large 4, internally codenamed "Le Chonk". It adopts a Mixture of Experts (MoE) architecture, with a total of 1.05 trillion parameters, though only 49 billion are activated per inference. The model natively supports text and image input, and offers a maximum context length of 1 million tokens. Its API is now available for preview, while the model weights are scheduled to be released publicly on October 27. Mistral claims Large 4 is currently the highest-performing open-weight model developed in the U.S. and Europe. Its official DeepSWE v1.1 software engineering benchmark score stands at 62%, compared to GLM-5.3 (61%), DeepSeek V4 Pro (57%), and Qwen 3.8 Max (51%). For financial benchmarks, Finch scores 67%, on par with DeepSeek V4 Pro; satellite image target localization task DIOR-RSVG scores 73%, higher than GPT-6 Astra's 68%. However, these results are primarily from Mistral's self-testing, not based on a unified standardized testing framework. When Zhipu AI launched GLM-5.3, it announced a score of 66.9% on the same DeepSWE v1.1 benchmark, while Kimi K3 reached 67.5% – both outperforming Large 4's 62%. The gap may stem from different testing configurations and execution methods used by various teams. Large 4 is trained from scratch: Mistral used around 4,000 Nvidia Grace Blackwell GPUs in its own European data centers over approximately two months. Mistral also positions "European autonomy" as a key selling point, noting that the model – from training and API services to future independent deployment – can all be hosted on European infrastructure.

20 minutes ago

Aptos launches MonoMove, an engine for on-chain markets that delivers a 55x speed boost.

Aptos Labs has announced the launch of MonoMove, the largest proposed upgrade in the history of the Aptos blockchain execution engine. The firm noted that internal testing shows the new engine runs smart contracts up to 55 times faster than the current system, putting its performance among the fastest of production-ready blockchains. Smart contracts are software deployed on blockchains that automatically execute functions such as trade matching, collateral transfers, and payment processing when predefined conditions are met. All on-chain financial operations ultimately rely on smart contracts for execution, meaning the speed of the engine powering these contracts directly determines the volume of activities a blockchain network can support. Aptos Labs added that on Decibel—its incubated fully on-chain trading platform—collateral withdrawal speed has improved by up to 55 times versus the existing engine, while order placement speed has increased by up to 22 times.

20 minutes ago

Mysten Labs announces partnership with Google Cloud to launch a verifiable AI agent arbitrator.

According to official announcements, Mysten Labs today announced a partnership with Google Cloud to launch the Verifiable Agent Arbitrator (VAA). Developed as a joint layer-1 evidence layer, the system is designed to help enterprises prove their AI agents operate within authorized scopes. As AI agents increasingly execute transactions, negotiate with other agents, and conduct cross-enterprise activities, the VAA provides enterprises with activity records that can be verified independently of the platform that generated them. Currently, enterprises aim to deploy AI agents into production environments, but relevant applications are limited by the lack of activity records trusted by counterparties, auditors, or regulators. The VAA is intended to fill this gap, helping enterprises comply with upcoming regulatory requirements—including the EU Artificial Intelligence Act (AI Act). Starting in 2027, the AI Act will impose fines of up to €15 million or 3% of a company’s global annual turnover for log failures. The VAA separates AI agent activity records from their integrity proofs. Detailed telemetry data, including prompts, model outputs, tool calls, and policy decisions, will remain private and stored in customer-controlled Google Cloud Storage; encrypted proofs of these records will be stored in Walrus, a data platform built for AI applications, and coordinated via the Sui Layer 1 blockchain.

20 minutes ago

Pornographic AI projects are too numerous: a16z directly removed them from its rankings, otherwise they would account for 20% of the Top 50.

Beating AI Insight News Brief: a16z has released its 7th edition Consumer AI Ranking. The biannual list tracks the most popular AI products, compiling the top 50 web and mobile AI tools by monthly visits and monthly active users (MAUs) respectively. The new edition has revised its rules: AI products primarily used for adult content will no longer be included in the ranking. This is not due to waning usage – quite the opposite. a16z noted that if the original rules were retained, these products would account for over 20% of the web top 50, or at least 10 spots. As a result, the firm decided to categorize them as an independent segment, excluding them from the overall consumer AI ranking. Mainstream AI assistants like ChatGPT and Claude have strict restrictions on adult content, creating demand for these specialized tools.

20 minutes ago

Devin begins daily "dreaming": automatically organizes long-term memory, with the solution now open-sourced.

Beating AI Express News: AI programming firm Cognition has rolled out Memory and Dreaming features for its AI coding agent Devin. Memory enables Devin to retain user preferences, corrections, and project experience across sessions, while Dreaming automatically reorganizes these memories in the background daily. Users can directly view what Devin has memorized, as well as recent modifications made by Dreaming. These memories are stored in a personal Memory Drive, essentially a persistent Git repository, with content organized primarily in Markdown files. A short MEMORY.md file serves as an entry point and index, loaded into the context at the start of each new session. Devin then searches for relevant memory files based on the current task, similar to how it queries code repositories, eliminating the need to feed the entire memory bank into prompts. Multiple Devin sessions can read and write to the same memory set simultaneously. Each session has its own Git checkout; after modifying memories, changes are committed and merged back into the Memory Drive. In case of concurrent edits, Git flags conflicts, preventing sessions from overwriting each other’s work. Dreaming itself is a background session that runs automatically daily. It reviews past conversations and existing memories, merges duplicate content, deletes temporary or long-unused information, and fills in previously unrecorded experience. Cognition has also open-sourced this framework as the Agent Memory Repo. Other agents like Claude Code and Cursor can similarly store long-term memories as files and manage them via Git. However, the open-source version only provides the memory structure and workflow; features like Devin’s daily automatic Dreaming run require additional configuration.

20 minutes ago

All three major U.S. stock indexes opened higher.

According to BIT (bit.com) market data, US stocks opened with the Dow Jones Industrial Average up 0.39%, the S&P 500 Index gaining 0.47%, and the Nasdaq rising 0.58%. AI chip stocks rallied broadly: AMD (AMD.O) climbed 2.5%, while Nvidia (NVDA.O), Broadcom (AVGO.O), and Microsoft (MSFT.O) all advanced more than 1%. Constellation Energy rose roughly 10% following the company’s signing of a 20-year power purchase agreement with Google.

20 minutes ago

Popular tokens

BitcoinEthereumHyperliquidSolanaTRONBNBTetherAaveXRPPepeFartcoinOndoJupiterUniswapBonkPendleEthenaArbitrumAvalancheLidoChainlinkPolygonDogecoinCardano