The DFlash team expands to Mac, with its 27B model running at 144 tokens per second.
56 minutes ago
Beating AI News (from Dongcha): Inco AI has open-sourced its local inference engine Splash, which runs Qwen3.8-27B on an M5 Max MacBook Pro at up to 144 tokens per second. LM Studio has since integrated Splash, and version 0.4.25 is available for direct use. Inco’s earlier DFlash technology has been adopted by SGLang, vLLM, TensorRT-LLM, and llama.cpp. Companies including Meta, NVIDIA, Xiaomi, and Poolside have also integrated DFlash to accelerate their small models. DFlash works by first having a small model predict multiple tokens in parallel, then submitting them to a large model for batch validation, reducing the computation required for token-by-token generation. Splash extends this optimization to its entire local inference engine. It currently only supports Qwen3.8-27B and Qwen3.6-35B-A3B, with each model having dedicated GPU kernels, memory configurations, and DFlash 2 small models. While its model coverage is limited, it delivers aggressive performance optimizations. The 144 tokens per second is the peak demo performance on the M5 Max. On the 48GB M5 Pro used by Inco for cross-device testing, Qwen3.8-27B hits 74 tokens per second for a single short-context run; with 4 concurrent runs, total throughput reaches 170 tokens per second, 3.9 times that of the second-place alternative. This is why Splash is particularly well-suited for agent use cases. When an agent launches multiple subtasks simultaneously, total concurrent throughput is often more critical than the speed of a single conversation. Splash is open-source and not tied exclusively to LM Studio. It currently requires an M3 or newer Mac, macOS 26.4 or later, and at least 36GB of unified memory.
Jiang Zhuoer: Short selling is employed to reduce drawdowns following spot price increases, generating a 34% coin-denominated profit after several months of trading.
4 minutes ago
Hyperliquid’s Open Interest Hits New One-Year High, Drawing Liquidity From Traditional Assets and Prediction Markets
4 minutes ago
Bitget has launched USDT-margined STONK perpetual contracts.
4 minutes ago
Bitget has launched six stock perpetual contracts, including SDGR, PATH and others.
4 minutes ago
Former Hong Kong bank manager sentenced to four years for US$1.6 billion in fake letters of credit and crypto-related bribes.
4 minutes ago
Vietnam and Austria are collaborating to strengthen cryptocurrency regulation, with plans to issue the first batch of licenses for crypto service providers in 2026.
4 minutes ago
Hot feeds
A trader profits $448K by monitoring #Binance's new listings!
2024.12.13 17:37:29
Last week, funds have flowed into #Bitcoin, #Ethereum, and #Hyperliquid.
2024.12.16 14:48:36
A $PEPE whale that had been dormant for 600 days transferred all 2.1T $PEPE($52M) to a new address.
2024.12.14 10:35:27
When Elon Musk tweeted about Moltbook, the meme coin MOLT experienced a short-term 30% price surge, hitting a new all-time high of $114 million.
2026.01.31 18:37:29
A smart #AI coin trader made $17.6M on $GOAT, $ai16z, $Fartcoin,$arc.
2025.01.05 16:05:18
A sniper earned 2,277 $ETH ($8.3M) trading $SHIRO within 18 hours!
2024.12.03 23:09:08
MoreHot Articles

How did I turn $1,000 into $30,000 with smart money?
2024.12.09

10 promising AI Agent cryptos
2024.12.05

The 30-Year-Old Entrepreneur Behind Virtual, a Multi-Million Dollar AI Agent Society
2025.01.22

10 smart traders specializing in MEMEcoin trading on Solana
2024.12.09

A trader lost $73.9K trading memecoins in just 3 minutes — a lesson for us all!
2024.12.13

What is $SPORE? Let us take you through the on-chain records to show you how it works.
2024.12.25

