Lookonchain APP

App Store

25 AI models collectively lose to humans: AI still struggles to understand much everyday common sense

57 minutes ago

Beating AI News: Scale Labs, the research arm of Scale AI, has partnered with Elorian to launch a new benchmark called Humanity’s Sixth Sense (HSS), designed to test whether AI can infer implicit information from images and videos. The benchmark includes 522 open-ended questions covering 288 images and 234 videos, such as whether two vehicles can pass each other, why a woman suddenly slows down while chasing a bus, and who holds more sway in a given scenario. The research team evaluated 25 multimodal models and had 20 human participants complete the tasks. Humans achieved an accuracy rate of 93.1%, while the top-performing model, GPT-6 Astra, scored only 53.6% even when using maximum inference effort. GPT-6.1 Sol and Claude Opus 5.5 followed with scores of 46.6% and 44.6% respectively, while the median score across all models was just 30.9%. Researchers analyzed 8,573 model failures and found that 94% were linked to missing key clues, misidentifying objects, or failing to infer implicit relationships in visuals. Only around 5% of errors were categorized as logical reasoning mistakes. Social comprehension proved particularly challenging: 21 out of the 25 models performed worst on these types of questions. Video-based tasks were also generally more difficult than image-based ones. Increasing inference effort does not always help. Models consumed an average of around 4,000 inference tokens per question, and for some tasks, longer reasoning times correlated with lower performance. The research team also tested letting agents zoom in, crop, and re-examine visuals, which led to improved results, though they still lagged far behind human performance.

Relevant content

GlobalFoundries and TSMC Sign $2 Billion Silicon Interposer Agreement

GlobalFoundries (GFS) and TSMC have signed a $2 billion silicon interposer agreement with an initial 5-year term. According to market data from BIT (bit.com), GlobalFoundries (GFS)’s pre-market gain on US stocks has widened to over 5%.

3 minutes ago

Google's US shares reversed pre-market losses, now up 1%.

According to market data from BIT (bit.com), Google's US stock flipped from negative to positive in pre-market trading, now rising 1%. At the Gemini at Work 2026 conference, Google Cloud CEO Thomas Kurian unveiled Gemini Agent, positioned as a single universal work agent. Users only need to state their goal rather than specific instructions to accomplish knowledge work, content creation, image and media generation, as well as code writing and execution.

3 minutes ago

Jev tops Ramp Software’s Dark Horse Chart, with its enterprise adoption rate rising by 1 percentage point in one month.

Beating AI Express: US enterprise payment platform Ramp has released its latest software procurement ranking, based on real transactions from over 70,000 businesses on its platform. TypeSafe AI topped the "Explosive Growth" list with its Jev model. In the market share growth ranking, it also placed second, trailing only Anthropic, while OpenAI ranked fourth. Jev, launched less than a month ago, has seen an approximate 1 percentage point rise in enterprise adoption. It is an AI model specifically designed to help software make judgments, not generating chat responses, but handling classification, selection, and scoring tasks. It charges $0.042 per million input tokens, with outputs provided free of charge. Low-cost models or model access platforms including DeepSeek, Featherless AI, and OpenRouter also feature on the list. However, demand for high-priced AI tools among enterprises is also growing. Video generation platforms Runway and Higgsfield have both entered the top ten in market share growth, and are currently primarily used in advertising and social media marketing. Ara Kharazian, chief economist at Ramp, stated that low-cost models may lower enterprises' AI call costs, but rising demand for image and video generation could push up overall AI spending.

3 minutes ago

STRK breaks above $0.06, rallying 18.1% in 24 hours.

According to HTX market data, STRK (Starknet) rallied 18.1% in 24 hours, decisively breaking above the $0.06 mark. On the news front, StarkWare CEO Eli Ben-Sasson proposed at Token2049 that blockchains need post-quantum readiness and cryptographic agility to counter threats from quantum computing and AI. Ben-Sasson noted Starknet is considering multiple options, including converting to a Layer 1 (L1) chain to independently advance security migration efforts. Earlier developments include: On October 2, the strkBTC incentive program covered bridging fees for the first 100 bridged BTC and distributed weekly faucet rewards; On October 5, the v0.14.4 mainnet launched, expanding SNIP-36 single-proof capacity to approximately 1.1 billion L2 gas.

3 minutes ago

Perplexity open-sources a new retrieval model: its 0.6B small model can directly query indexes built by the 9B model.

Beating AI News: Perplexity has open-sourced two multimodal embedding models, pplx-embed-v2-late, with sizes of 0.6B and 9B parameters respectively. The two models are vector-compatible, allowing developers to use the larger 9B model for building knowledge bases while leveraging the smaller 0.6B model for daily searches—eliminating the need to run the large model for every query. The new models support text, image, and PDF page retrieval. Traditional embedding models typically compress content into a single vector, which often leads to detail loss. In contrast, the new models retain a 128-dimensional vector for each token, enabling search queries to match relevant segments within documents directly. When processing PDFs, PPTs, and scanned materials, the models can convert page images straight to vectors, bypassing OCR text extraction while preserving charts, tables, and layout information. Both models are trained from the same 18B teacher model and share a common vector space. In official ViDoRe v3 image retrieval tests, the 0.6B model scored 62.3% when used for both database building and queries; switching to the 9B model for database building and the 0.6B model for queries boosted the score to 63.5%; the full 9B model achieved 65.2%. The model weights are now available on Hugging Face under the MIT license.

3 minutes ago

Hong Kong will take appropriate law enforcement actions as needed to crack down on unlicensed payment platforms.

Regarding some payment platforms that are registered solely as fintech companies without holding any financial licenses, Hong Kong’s Secretary for Financial Services and the Treasury Chan Ho-lim said that appropriate enforcement actions will be taken as needed under the relevant statutory regulatory framework to maintain financial stability and safeguard the rights and interests of customers and the public. Under Hong Kong’s Payment Systems and Stored Value Facilities Ordinance, it is illegal for any individual or entity to issue or operate stored value facilities in Hong Kong without a license, unless a statutory exemption applies. If cases of suspected unlicensed operation or violations of the Payment Ordinance are identified, the Hong Kong Monetary Authority (HKMA) will directly intervene to handle them, collaborate with other relevant regulators based on the nature of the case, and refer cases to law enforcement departments when necessary.

3 minutes ago

Popular tokens

BitcoinEthereumHyperliquidSolanaTRONBNBTetherAaveXRPPepeFartcoinOndoJupiterUniswapBonkPendleEthenaArbitrumAvalancheLidoChainlinkPolygonDogecoinCardano