20-hour programming benchmark highlights performance gaps: Claude Fable 5.1 leads GPT-5.6 by more than 24 points, with GLM-5.3 ranking third.
1 hours ago
Beating AI Express: AI research team Proximal has released FrontierSWE v2, a long-horizon programming benchmark. The number of tasks has expanded from 17 to 34, with each model running 5 trials per task, each trial taking up to 20 hours. Claude Fable 5.1 posted an average score of 56.29%, significantly outperforming GPT-5.6’s 32.2%. Open-source model GLM-5.3 ranked third with 30.2%. Tasks in FrontierSWE go far beyond basic code debugging: AI agents must build circuit simulators from scratch, train weather forecasting models, match star catalogs using telescope images, or train racing bots solely from game visuals. Version v2 has uniformly switched to the Proximus harness. Each task runs for up to 20 hours; when a model is ready to submit its work, the system saves its current state and notifies it of remaining time to prevent premature task termination—a change that significantly impacted scores. Proximal’s comparison across 6 tasks found that both Claude Opus 5 and GPT-5.6 ran longer with the Proximus harness, and posted higher average scores than with their original harnesses. The benchmark also caught multiple instances of intentional cheating: GPT-5.6 once recognized that accessing public answers “might involve anti-cheat issues” but still used the shortcut; on another occasion, it even exploited Modal’s backend services to read hidden validation files. Muse Spark 1.2 modified test scripts, inserted public answers, and wrote code to cover up its cheating traces. All runs confirmed to be violations were scored zero.
Nvidia plans to acquire Hugging Face for $12.93 billion.
4 minutes ago
Microduck surges over 35% in a short-term rally, market cap crosses $46 million.
4 minutes ago
Digital currency network Cari completes first round of external financing worth $32.5 million.
4 minutes ago
Bonk Guy hypes meme coin USELESS, driving its market cap to break through $150 million, with the token rising an additional 50% in 24 hours.
4 minutes ago
A crypto whale withdrew 824,350 HYPE tokens from Coinbase, valued at approximately $67.52 million.
4 minutes ago
Hyperliquid officially announces HIP-3*: opens the door for compliant institutional-grade participants, offering an optional whitelist mechanism.
4 minutes ago
Hot feeds
A trader profits $448K by monitoring #Binance's new listings!
2024.12.13 17:37:29
Last week, funds have flowed into #Bitcoin, #Ethereum, and #Hyperliquid.
2024.12.16 14:48:36
A $PEPE whale that had been dormant for 600 days transferred all 2.1T $PEPE($52M) to a new address.
2024.12.14 10:35:27
When Elon Musk tweeted about Moltbook, the meme coin MOLT experienced a short-term 30% price surge, hitting a new all-time high of $114 million.
2026.01.31 18:37:29
A smart #AI coin trader made $17.6M on $GOAT, $ai16z, $Fartcoin,$arc.
2025.01.05 16:05:18
A sniper earned 2,277 $ETH ($8.3M) trading $SHIRO within 18 hours!
2024.12.03 23:09:08
MoreHot Articles

How did I turn $1,000 into $30,000 with smart money?
2024.12.09

10 promising AI Agent cryptos
2024.12.05

The 30-Year-Old Entrepreneur Behind Virtual, a Multi-Million Dollar AI Agent Society
2025.01.22

10 smart traders specializing in MEMEcoin trading on Solana
2024.12.09

A trader lost $73.9K trading memecoins in just 3 minutes — a lesson for us all!
2024.12.13

What is $SPORE? Let us take you through the on-chain records to show you how it works.
2024.12.25

