Qwen 3.8 27B open-weights LLM earns praise for reasoning and local speed
Qwen 3.8 27B 开源权重模型获赞:推理强、本地运行快
⭐️ 9.0/10

The Qwen team released Qwen 3.8 27B, a 27-billion-parameter open-weights language model distributed in FP8 precision on Hugging Face. Early community tests highlight strong reasoning and good local performance on standard laptops. This release matters because a strong, open-weights 27B model can run locally on consumer hardware, expanding who can access capable AI without cloud dependencies. The enthusiastic reception also signals growing community demand for open alternatives and sharper competition among open-weights models such as Gemma 4. According to one commenter, Qwen 3.8 27B is only the second local model after Gemma 4 to pass their private reasoning benchmark, though it needed five times more tokens and 12m30s with multi-token prediction (MTP) enabled. Other reviewers note FP8 VRAM usage appears less efficient than Gemma 4 or Glimmer, and an RTX 5090 user measured roughly 138 tokens/s with the ninfer inference engine.

hackernews · erdaltoprak · Aug 14, 15:00 · Discussion

Background: Qwen is Alibaba's family of AI models, many of which are distributed under open licenses such as Apache 2.0. 'Open-weights' means the trained parameters of the model are publicly available, so anyone can download, inspect, fine-tune, or run them. 'Local LLM inference' means executing the trained model on your own hardware rather than sending prompts to remote cloud servers owned by companies like OpenAI or Google.

References

Discussion: Community reactions are strongly positive: users praise Qwen 3.8 27B's reasoning quality and note it runs impressively on a laptop, with one saying it is 'absolutely the best pelican' rendered by a laptop-runnable model. There are also practical caveats about higher VRAM usage and slower reasoning traces, alongside excitement that open-weights models are approaching frontier capabilities previously dominated by big US companies.

Tags: #Qwen, #LLM, #Open Source, #Local AI, #Benchmarks


GLM-5.3 Release Shows Frontier Coding and Emergent Cyber Capabilities
GLM-5.3 发布:前沿编程与突现网络能力
⭐️ 9.0/10

Z.ai released GLM-5.3, a post-training upgrade to GLM 5.2 that achieves frontier coding performance and open-source state-of-the-art results on Terminal Bench 3.0 and Agents' Last Exam. The release also surfaces 'emergent cyber capabilities,' with users reporting autonomous vulnerability exploitation and Z.ai running a coordinated vulnerability disclosure portal at cvd.z.ai. This matters because LLM-driven cyber capabilities are advancing faster than many expected, potentially lowering the cost of vulnerability discovery while raising new safety and disclosure questions. It affects security researchers, enterprises, and AI safety teams, and intensifies the competitive pressure on frontier model vendors such as OpenAI and Anthropic. Benchmark comparisons still show some models ahead, with one commenter noting 'Mythos 5 remains well ahead at 181 and 247 tasks' on exploitation-chain benchmarks. Community tests included 0-days in WordPress plugins, RCE exploits, and Linux 6.8 kernel exploit adaptation, while Z.ai's CVD portal lists many critical/high CVEs under embargo.

hackernews · pella · Aug 14, 05:19 · Discussion

Background: Z.ai, formerly known as Zhipu AI, is a Chinese AI company that develops the GLM family of large language models and has released many of them under the MIT License since July 2025; it was added to the U.S. Entity List in January 2025 and held an IPO on the Hong Kong Stock Exchange in January 2026. Emergent capabilities in AI are skills that are absent in smaller models but abruptly appear in larger ones, and autonomous agents are increasingly able to discover, chain, and exploit vulnerabilities in minutes rather than weeks.

References

Discussion: The discussion is largely positive but measured: one user described seamless red-team execution including 0-days and RCE with GLM-5.3 in a Claude Code harness, while another highlighted large-scale OSS vulnerability scanning and disclosure under embargo. Others noted the model is 'still shy of Sol and Fable, but only just by a hair,' raised economic concerns about switching from OpenAI, and appreciated that the blog post felt researcher-written rather than marketing hype. There are also underlying ethical and safety concerns about autonomous scanning and mass vulnerability disclosure.

Tags: #AI, #Cybersecurity, #LLM, #Coding, #Vulnerability Discovery


X Open-Sources 'For You' Algorithm and Adds Shadow-Ban Check Tool
X 开源“为你推荐”算法并推出“限流”自查工具
⭐️ 9.0/10

X has open-sourced the source code for its 'For You' recommendation algorithm on GitHub and introduced a transparency tool that lets users check whether their account or individual posts received visibility-limiting labels last month. The data can be downloaded for offline analysis. By exposing the code behind one of the most-used social feeds, X is advancing algorithmic accountability and making shadow-banning practices more detectable. This move could pressure other platforms to offer similar transparency and shape public debate on content curation. The visibility-limit tool is currently in a pilot phase, randomly sampling eligible accounts; eligibility requires the account to be at least one year old and to have posted at least 10 times in the previous month. X also says that open-sourcing the For You code expanded its open-source codebase by about 10 to 15 times.

rss · 小互(@imxiaohu) · Aug 14, 02:37

Background: The 'For You' page is Twitter/X's default algorithmic timeline, which selects and ranks posts using a mix of user interactions, content signals, and reputation scores. Historically, these ranking mechanisms were proprietary, leading to concerns about opaque moderation and 'shadow-banning'—where a user's content is quietly demoted. Open-sourcing this code and launching the visibility-label checkpoint fulfills earlier promises from Elon Musk about algorithm transparency. The published repository also includes a PageRank-based user reputation component and GraphJet-based real-time streaming services.

References

Tags: #recommendation algorithm, #open source, #social media, #transparency, #algorithmic accountability


X Open-Sources Its For You Timeline Algorithm with Transparent Weights
X 开源 For You 时间线算法,权重全面透明
⭐️ 9.0/10

X has open-sourced the algorithm powering its For You timeline, publishing the code in the xai-org/x-algorithm repository. The release discloses the ranking factors and weights that determine post visibility, making recommendation logic publicly inspectable. This marks a significant transparency breakthrough for a major social platform, allowing users, researchers, and regulators to understand exactly how content is ranked. It also sets a precedent for algorithmic accountability in AI/ML systems. The repository focuses on transparency into code affecting For You post visibility; some components, such as the Phoenix scoring model, are designed to run end-to-end. The algorithm distills roughly 500 million daily posts into a shortlist shown on users' devices.

rss · meng shao(@shao__meng) · Aug 14, 02:20

Background: X (formerly Twitter) first open-sourced its recommendation algorithm in 2023 under the twitter/the-algorithm repository. The For You timeline algorithm typically scores posts based on users' past interactions, similar users' engagement, and trending topics. Open-sourcing such code aims to give users full transparency into how posts are recommended.

References

Tags: #algorithm, #open-source, #transparency, #recommendation-system, #X/Twitter


Qwen 3.8-Max Debuts on Modal with 1M Context and DFlash Speculator
Qwen 3.8-Max 登陆 Modal,支持 100 万上下文与 DFlash 推测解码
⭐️ 9.0/10

Alibaba Qwen announced that Qwen 3.8-Max (Qwen3.8-2.4T-A95B) is now available on Modal, served with a custom DFlash speculator and the full 1M token context window. This is a major model release with 2.4T parameters. The deployment of a 2.4T-parameter MoE model with a 1M context window on a serverless platform signals a milestone in large-scale LLM inference accessibility. It also highlights the growing importance of speculative decoding techniques like DFlash for reducing latency and cost. The model is described as Qwen3.8-2.4T-A95B, indicating a Mixture-of-Experts architecture with only ~95B active parameters despite the 2.4T total count. The custom DFlash speculator is trained on tool-call-heavy data, and Modal serves the full context window rather than a reduced version.

rss · Qwen(@Alibaba_Qwen) · Aug 14, 15:32

Background: DFlash is a speculative decoding method where a small draft model (the speculator) proposes a block of future tokens, and a larger verifier model (the main LLM) accepts or rejects them, speeding up inference while preserving output quality. Modal is a serverless compute platform for AI and data teams, allowing developers to run GPU and CPU workloads at scale. Speculative decoding techniques such as DFlash are increasingly used by major inference providers to achieve 2-3x higher tokens per second.

References

Discussion: The quoted tweet from Modal highlights the availability and the custom DFlash speculator. The original post has limited public comments available, but the sentiment appears positive given the enthusiastic 'Love seeing our launch partner Modal go all in from Day 0!' and the high engagement (345 likes, 28,964 views). No critical discussion or counterarguments are visible in the provided content.

Tags: #LLM, #Qwen, #Modal, #AI Infrastructure, #Model Release


Qwen3.8-Max Goes Live on Together AI with 2.4T Parameters
Qwen3.8-Max 在 Together AI 上线,总参数达 2.4T
⭐️ 9.0/10

Alibaba's Qwen team announced that Qwen3.8-Max (officially Qwen3.8-2.4T-A95B) is now available on Together AI as a Day 0 launch partner. The open-weight Mixture-of-Experts model features 2.4T total parameters, 95B active parameters, and a 1M-token context window. This is a major open-weight model release with enterprise-grade specs, and the Day 0 partnership signals strong infrastructure support from Together AI. It could accelerate adoption of large MoE models on cloud GPU platforms and intensify competition in the open-source LLM space. The model name 'Qwen3.8-2.4T-A95B' reveals its Mixture-of-Experts architecture: 2.4 trillion total parameters with only 95 billion active per token, improving inference efficiency. Together AI highlights strong coding and agent capabilities, and the 1M context window enables long-document and complex-agent workflows.

rss · Qwen(@Alibaba_Qwen) · Aug 14, 15:24

Background: Mixture of Experts (MoE) is a machine learning technique that divides a model into multiple specialized sub-networks, or 'experts,' and uses a routing mechanism to activate only a few relevant experts per input token. This is why the model can have 2.4T total parameters while only 95B are active at inference time, reducing computational cost without necessarily sacrificing capacity. Open-weight MoE models have become a key trend in LLM development, as seen with other large-scale releases.

References

Tags: #Qwen, #Together AI, #LLM, #Model Release, #AI Infrastructure


GPT-5.6-Sol Autonomously Proves Crouzeix's Conjecture in 16 Hours
GPT-5.6-Sol 在 16 小时内自主证明 Crouzeix 猜想
⭐️ 9.0/10

Shanmu Jin, a Beijing neurosurgeon resident, used OpenAI's GPT-5.6-Sol model inside ChatGPT Work to autonomously derive a proof of Crouzeix's Conjecture in 16 hours. The result was reviewed by Michel Crouzeix and verified by mathematicians including Alex Townsend. This marks a major milestone where an AI system autonomously produced the decisive idea in proving a longstanding mathematical conjecture. It signals a paradigm shift in AI-assisted research and could accelerate discovery in mathematics and other scientific fields. The proof was generated with the model disconnected from the web, running autonomously for 16 hours. According to the arXiv paper, the solution combines previous tools for weaker estimates with a perturbation lemma for 2-dilations. GPT-5.6-Sol is OpenAI's frontier model, noted for its strong coding and reasoning capabilities.

rss · The Rundown AI(@TheRundownAI) · Aug 14, 15:20

Background: Crouzeix's Conjecture is an open problem in matrix analysis proposed by Michel Crouzeix in 2004. It concerns the norm of functions of complex matrices in relation to their field of values, and is relevant to numerical linear algebra. GPT-5.6-Sol is the flagship model of OpenAI's GPT-5.6 family, designed for complex reasoning and agentic workflows. ChatGPT Work is an OpenAI product that lets users delegate tasks and run long-running workflows.

References

Tags: #AI, #mathematics, #GPT, #breakthrough, #research


Cursor Acquired by SpaceX, Joins SpaceXAI to Build Grok
Cursor 被 SpaceX 收购,加入 SpaceXAI 助力 Grok
⭐️ 9.0/10

Cursor announced that it has officially closed its acquisition by SpaceX and will join the SpaceXAI team. The team will work on improving Grok, Grok Build, Grok Bot, Grok API, Cursor, and more, aiming to make Grok the world's most useful AI. This acquisition reshapes the AI coding tool landscape, bringing one of the most widely used AI-powered code editors under the same umbrella as the Grok model. It signals deepening consolidation among AI tools and models, and could meaningfully affect developers who rely on Cursor as well as the broader AI assistant market. The deal was an all-stock transaction valuing Cursor (Anysphere) at $60 billion, and it closed on August 14, 2026. Cursor will remain a wholly owned subsidiary of SpaceX, integrated within the SpaceXAI unit, and per recent reports, Grok 4.6 was already pre-trained with data from Cursor.

rss · Cursor(@cursor_ai) · Aug 14, 13:02

Background: Cursor, founded in 2022 as Anysphere, is an AI-powered coding agent and IDE that lets developers edit code, search codebases, and complete programming tasks using natural-language instructions. SpaceXAI, formerly xAI, is the AI division of SpaceX that develops the Grok family of models, including Grok 4.6. The acquisition follows a broader restructuring in May 2026, when Grok and X were folded into SpaceX's AI division.

References

Tags: #acquisition, #AI, #SpaceX, #Cursor, #Grok


Gemini 4 Announced as Google DeepMind's Most Ambitious Pre-training Run Yet
Gemini 4:谷歌 DeepMind 最雄心勃勃的预训练运行
⭐️ 9.0/10

Logan Kilpatrick of Google DeepMind announced that Gemini 4 is their most ambitious pre-training run to date, confirming that the next-generation frontier model is already in training. This matters because Gemini 4 represents Google's next flagship AI model, and a more ambitious pre-training run signals potential major capability leaps that could reshape the competitive landscape among leading AI labs. Gemini 4 is not yet released as a public model; it is confirmed to be in pre-training. Google announced the effort at the end of its July 21, 2026 model announcement, and the company is reportedly shifting to monthly AI releases.

rss · Logan Kilpatrick(@OfficialLoganK) · Aug 14, 02:28

Background: Pre-training is the initial phase in building large language models, where the system learns general patterns and knowledge from vast amounts of unlabeled data before fine-tuning for specific tasks. Gemini is Google DeepMind's series of multimodal AI models, and Gemini 4 is the upcoming frontier model currently in pre-training, indicating it is still in an early stage of development.

References

Tags: #AI, #Gemini 4, #Google DeepMind, #Pre-training, #LLM


npm 12 Released: Install Scripts Default Off, Registry Shifts to Explicit Trust
npm 12 发布:安装脚本默认关闭,注册表转向显式信任
⭐️ 9.0/10

npm 12 makes install scripts opt-in by default and restricts non-registry sources to improve security and trust. Explicit approval is now required for running scripts, including implicit builds. This is a major security-default change in the world's largest package registry, altering the trust model for package installation. It affects millions of JavaScript developers and organizations by reducing supply chain attack surface, but may require workflow adjustments. Script allowances are now off by default, and the update also restricts non-registry sources. These changes address long-standing community concerns about automatic script execution during package installation.

rss · InfoQ · Aug 14, 06:39

Background: npm lifecycle scripts run automatically during package installation, and attackers have weaponized them in incidents like the Axios postinstall remote access trojan attack. The new explicit-trust model requires developers to approve such behaviors, aligning with existing best practices like ignore-scripts. This is part of a broader industry shift toward stronger software supply chain security.

References

Tags: #npm, #security, #JavaScript, #package management, #software supply chain


Meta Open-Sources Muse Glimmer: 30B On-Device Agentic Model
Meta 开源 Muse Glimmer:30B 端侧智能体模型
⭐️ 9.0/10

Meta AI Research released Muse Glimmer, a 30-billion-parameter open-weight model under the Apache 2.0 license, designed for on-device autonomous agents and multimodal task execution. It enables local workflows on consumer GPUs without relying on cloud APIs. This release significantly advances open-weight AI by bringing agentic capabilities to consumer hardware, reducing dependence on cloud APIs. Developers, researchers, and enterprises can now build private, local automation and coding assistants with a 30B model. Muse Glimmer employs a multi-stage training approach for efficient performance and supports multimodal inputs, enhancing coding and automation tasks. The Apache 2.0 license permits broad commercial and research use.

rss · InfoQ · Aug 14, 05:05

Background: An agentic model refers to an AI system that operates autonomously, exhibiting goal-driven behavior and adaptability, often using tools and memory to complete tasks. An open-weight model allows anyone to download and run its trained parameters locally, enabling study, modification, and deployment without proprietary restrictions. Muse Glimmer combines these concepts, targeting local execution on consumer GPUs.

References

Tags: #AI, #Open Source, #On-Device, #Agentic Model, #Meta


Apple CEO Transition: Cook to Board Chairman, Ternus Takes Helm in 2026
苹果换帅:库克卸任 CEO,特努斯 2026 年起接任
⭐️ 9.0/10

Apple has announced a leadership transition: current CEO Tim Cook will become executive chairman of the board, and Senior Vice President of Hardware Engineering John Ternus will take over as CEO on September 1, 2026. The board has unanimously approved the arrangement, with Cook remaining CEO through the summer to complete the handover. This marks Apple's first CEO change since 2011, making it a pivotal moment for the company's future direction. Ternus, who has led hardware engineering for years, will now shape Apple's product strategy in the post-Cook era, affecting the wider tech industry and consumers worldwide. Ternus joined Apple in 2001, was promoted to Vice President of Hardware Engineering in 2013, and joined the executive team in 2021. Current Chairman Arthur Levinson will become Lead Independent Director on September 1, 2026, and Ternus will join the board the same day.

telegram · zaihuapd · Aug 14, 11:00

Background: Tim Cook has served as Apple's CEO since 2011, succeeding Steve Jobs, and has overseen the company's rise to become one of the world's most valuable. The transition is significant because Apple leadership changes are rare. John Ternus is a hardware veteran who has overseen development of iPhone, Mac, iPad, AirPods, and other products in recent years. The orderly handover with a five-month transition period reflects careful succession planning.

Tags: #Apple, #CEO transition, #Tech industry, #Leadership


PostgreSQL Patches Critical to_char Heap Overflow Enabling Code Execution
PostgreSQL 修复 to_char 高危堆溢出漏洞,可执行任意代码
⭐️ 9.0/10

PostgreSQL disclosed CVE-2026-14669, a heap overflow in the to_char(timestamptz) function triggered by overlong POSIX timezone abbreviations. The flaw lets low-privilege database users execute arbitrary code with the OS privileges of the PostgreSQL server; fixes ship in 18.6, 17.11, 16.15, 15.19, and 14.24. With a CVSS score of 8.8, this vulnerability affects all supported PostgreSQL branches and lets an authenticated low-privilege user take full control of the database server. Given PostgreSQL's massive deployment, administrators should treat this as an urgent patching priority. All versions before 18.5, 17.11, 16.15, 15.19, and 14.24 are affected; because 18.5 was never released due to a regression, 18-series users should upgrade to 18.6, while others upgrade to 17.11, 16.15, 15.19, or 14.24. This minor update requires no dump/reload or pg_upgrade — just replace the binaries and restart the service.

telegram · zaihuapd · Aug 14, 14:35

Background: to_char is a PostgreSQL formatting function that converts timestamps, intervals, and numbers into formatted strings, and timestamptz is the timezone-aware timestamp type. POSIX timezone specifications allow abbreviations like 'EST' or arbitrary strings in angle brackets, and a malformed or extremely long abbreviation can overflow the buffer when to_char processes it. A heap overflow is a memory-safety bug that attackers can often turn into arbitrary code execution, making this a critical remote-code-execution risk.

References

Tags: #postgresql, #security, #CVE, #vulnerability, #database


Cryptography Expert Explores the 'Going Dark' Era of Law Enforcement Hacking
密码学专家探讨执法黑客的“走向黑暗”时代
⭐️ 8.0/10

In a new essay on Cryptography Engineering, a cryptography expert argues that the 'going dark' debate has shifted from pushing for backdoors to embracing law enforcement hacking — actively exploiting software vulnerabilities to access encrypted devices and communications. This reframing matters because it moves the surveillance debate from legal mandates to technical capabilities, raising critical questions about vulnerability disclosure, the security of consumer devices, and the balance between privacy and law enforcement. The article has sparked a lively discussion, reflecting its relevance to ongoing policy battles over encryption. The essay discusses the history of wiretapping, the Vulnerabilities Equities Process (VEP), and the practical limits of bug hunting, including the possibility of hitting a 'ceiling' on useful vulnerabilities. It also touches on security disparities, noting that while well-resourced actors find bugs, many organizations still fail at basic security hygiene.

hackernews · vslira · Aug 14, 20:52 · Discussion

Background: The 'going dark' debate refers to the challenges law enforcement faces in accessing encrypted communications and data, which they argue hampers criminal investigations. In response, some governments have pushed for backdoors or other legal access mechanisms, while others have turned to law enforcement hacking — using vulnerabilities to gain access to devices. The Vulnerabilities Equities Process (VEP) is a U.S. government framework that decides whether to disclose newly discovered zero-day vulnerabilities to vendors or retain them for intelligence and offensive cyber operations. Historically, wiretapping required physical infrastructure, but modern interception increasingly relies on technical exploitation and cooperation from tech companies.

References

Discussion: Community comments span several viewpoints: Animats provides historical context on costly pre-digital wiretapping, while mbroshi argues that AI-generated code is increasing bugs, challenging the idea of a vulnerability 'ceiling.' Insimwytim contrasts well-resourced state actors with companies that neglect basic security, and fitblipper questions the 'going dark' label itself given the ubiquity of surveillance cameras and metadata sharing.

Tags: #cryptography, #law enforcement, #security, #hacking, #government surveillance


Why Opus 5 Feels Worse to Work With: Agent-Speak and Over-Apologies Frustrate Users
Opus 5 为何让人感觉更难用?面向智能体的表达与过度道歉引不满
⭐️ 8.0/10

A new critique of Anthropic's Claude Opus 5, published at mun-logadan.github.io, argues the model feels worse to work with despite being more capable, citing elliptical writing, over-apologizing, and agent-oriented communication. The post sparked a large Hacker News discussion with 770 points and 705 comments. This debate signals a possible shift in AI development where models are optimized to communicate with other agents rather than human users. It matters for developers, professionals, and everyday users who rely on Claude, and raises urgent UX questions for the next generation of frontier models. From Anthropic's announcement, Opus 5 is designed as a strong agentic coding model for long-running, multi-step work and ranks near the top of public benchmarks while remaining behind Mythos 5 on cybersecurity tasks. However, commenters report that its communication style feels tiring, and some say they have switched to OpenAI's Sol or back to Opus 4.8.

hackernews · numeri · Aug 14, 10:12 · Discussion

Background: Claude Opus is Anthropic's high-end model family, and Opus 5 is optimized for agentic coding and knowledge work. 'Elliptical writing' refers to prose that omits words and context, often feeling terse or abstract; 'agent-oriented communication' describes outputs designed to be consumed by other AI agents rather than by humans. These concepts are central to understanding the critique.

References

Discussion: Commenters largely agree with the critique, complaining about elliptical phrasing, excessive 'confessions' of mistakes, and verbose yet draining dialogue. Several said they moved to OpenAI's Sol or reverted to Opus 4.8, while others speculated that Anthropic is optimizing for agents over humans; there was also concern that benchmark scores mask degraded everyday quality.

Tags: #AI, #LLM, #UX, #human-AI interaction, #model behavior


Firefox now the only major browser supporting uBlock Origin
Firefox 成为唯一支持 uBlock Origin 的主流浏览器
⭐️ 8.0/10

Firefox is now the last major browser where users can still run the full version of uBlock Origin, while Chrome, Edge, and other Chromium-based browsers have migrated to Manifest V3, which breaks the extension. This shift makes Firefox the only mainstream venue for robust content blocking. This matters because uBlock Origin has long been one of the most effective tools for blocking ads, trackers, and malicious scripts, and the MV3 restriction degrades the ad-blocking capabilities available in most browsers. It also represents a broader industry shift where browser vendors, not users, decide what extensions are permitted. The key technical constraint is Manifest V3's removal of broad webRequestBlocking permissions; the replacement declarativeNetRequest API limits extension rule sets and requires a static rule limit. Firefox still supports both Manifest V2 and V3, and also performs manual code vetting on popular extensions like uBlock Origin, while an unofficial MV3 port exists but is not a full replacement.

hackernews · DemiGuru · Aug 14, 19:03 · Discussion

Background: Browser extensions are small programs that modify a browser's behavior, and their capabilities are defined by a manifest file. Manifest V3 (MV3) is the latest version of this extension platform introduced by Google, which changed permission models to increase privacy, security, and performance but made it impossible for full-featured content blockers like uBlock Origin to run on Chromium-based browsers. Firefox's extension platform has diverged by continuing to support older APIs, allowing these extensions to remain functional.

References

Discussion: Commenters largely frame the situation as a loss of user freedom, with one noting that Google pushed MV3 through despite objections, like 'the frog got boiled.' Others pointed out Firefox's extra code vetting for uBlock Origin, discussed the unofficial uBlock-MV3 port, and offered mixed reports on whether uBlock Origin Lite's ad blocking is sufficient.

Tags: #privacy, #ad-blocking, #browsers, #extensions, #manifest v3


GLM-5.3 Debuts as Top Open-Source Coding Model via RL, Trailing Closed Models
智谱发布 GLM-5.3:凭强化学习登顶开源编程,仍逊于闭源模型
⭐️ 8.0/10

Zhipu AI released GLM-5.3, sharing the same base model as GLM-5.2 with all improvements coming from post-training reinforcement learning. It achieved top open-weight scores on multiple coding benchmarks, such as 28.3 on Terminal Bench 3.0 and 66.9 on DeepSWE v1.1. This release shows scaling reinforcement learning post-training can sharply improve coding and security capabilities without a new base model. Because GLM-5.3's weights will be public, it also puts a near-frontier vulnerability-finding engine into open circulation, raising both opportunity and safety concerns. Key gains include CyberGym score of 84.5%, exceeding GPT-5.6 Sol and Fable 5, and ExploitBench jumping from 24.4% to 54.4%. The model has no option to disable chain-of-thought; users choose low, high, or max thinking levels, and weights ship in two weeks. It runs on the slime RL training framework with IndexShare and SAO components.

rss · 宝玉(@dotey) · Aug 14, 05:50

Background: Reinforcement learning post-training uses feedback from task execution to refine a model's behavior after initial pre-training. Benchmarks like Terminal Bench 3.0, DeepSWE, and Agents' Last Exam measure how well agents handle long-horizon, real-world tasks in terminal or software environments. Zhipu says the model also uncovered 2,436 real vulnerabilities across 269 open-source projects, suggesting emergent offensive security capability.

References

Tags: #AI, #GLM, #LLM, #强化学习, #编程


Zhipu's GLM-5.3 Boosts Coding and Cybersecurity via Post-Training Only
智谱 GLM-5.3 仅靠后训练大幅提升编程与网络安全能力
⭐️ 8.0/10

Zhipu AI released GLM-5.3, built on the same base model as GLM-5.2 with all improvements coming from post-training. Internal evaluations show a 50% improvement in programming compared with GLM-5.2, and cybersecurity performance is said to match Anthropic's Mythos 5, with 2,436 real vulnerabilities found across 269 open-source projects. This release shows that post-training alone can produce dramatic gains in specialized domains like coding and cybersecurity, without requiring a new foundation model. It intensifies competition among frontier LLM providers and matters for developers, security researchers, and enterprises evaluating model capabilities. The claims come from a third-party summary and Zhipu's internal evaluations, not an official technical report. The model reportedly uncovered 2,436 real vulnerabilities in 269 open-source projects, a concrete result that supports the cybersecurity claim.

rss · 小互(@imxiaohu) · Aug 14, 13:41

Background: Language models are typically trained in two phases: pre-training, which gives broad knowledge from large datasets, and post-training, which aligns behavior and sharpens specialized skills. Anthropic's Mythos 5 is positioned as a cybersecurity-focused advanced model, so matching it is a notable benchmark. This case illustrates that post-training alone can be a powerful lever for capability improvements.

References

Tags: #AI, #GLM, #LLM, #Post-training, #Cybersecurity


X Open-Sources For You Algorithm, Reveals Exact Action Weights
X 开源 For You 算法,公布精确动作权重
⭐️ 8.0/10

X has open-sourced the source code for its For You timeline recommendation algorithm, publishing for the first time the exact numeric weights assigned to user actions, such as 0.5 points for a like and 20 points for copying a link to share. The release also includes the filtering and labeling system and the actual training code. This unprecedented transparency gives researchers and developers a concrete reference for how a major social platform ranks content, shifting from a black-box relevance score to explicit action prediction. It could spur deeper community discussion and independent auditing of recommendation systems. The algorithm now uses a multi-task approach that predicts whether a user will like, reply, copy-link share, or report a post, then combines these probabilities into a single score. The configuration lists 21 positive weights and 5 negative weights, a major expansion from the earlier open-source release that only revealed the structural framework.

rss · 小互(@imxiaohu) · Aug 14, 11:49

Background: X (formerly Twitter) open-sourced its For You recommendation algorithm on GitHub under the Apache v2 license, with the expanded release including model configuration, filters, and core ranking system details. The open-source codebase has reportedly grown by approximately 10 to 15 times. This approach reflects an industry trend toward multi-task deep recommender systems, where models optimize for several objectives at once rather than a single relevance score.

References

Tags: #algorithm, #open source, #recommendation system, #machine learning, #X/Twitter


GPT-5.6 Sol Outperforms GPT-5.5 at Lower Reasoning Intensity
GPT-5.6 Sol 低推理强度超越 GPT-5.5 高推理强度
⭐️ 8.0/10

OpenAI's official tests reportedly show that GPT-5.6 Sol at 'low' reasoning intensity outperforms GPT-5.5 at 'high' reasoning intensity, while using fewer tokens and better filtering irrelevant information. This marks a significant efficiency leap in LLM reasoning, meaning users can achieve equal or better results with lower computational cost and latency. It could reshape how developers configure reasoning settings and pressure competitors to focus more on efficiency. GPT-5.6 reached general availability on July 9, 2026, and ships as three separate models: Sol, Terra, and Luna, not as settings on one model. Sol is the top capability tier for the hardest coding, agentic, and research work, and ranks high on public benchmark lists.

rss · 小互(@imxiaohu) · Aug 14, 11:27

Background: GPT-5.6 is OpenAI's latest model generation, available across ChatGPT, Codex, and the API. Reasoning intensity controls how much computation the model spends on 'thinking' before answering; higher intensity usually improves accuracy but consumes more tokens and increases latency. This news indicates that Sol's native capabilities have advanced to the point where even its low-intensity outputs surpass the previous model's high-intensity outputs, saving both cost and time.

References

Tags: #GPT-5.6, #OpenAI, #LLM, #Efficiency


MiniMax Open-Sources Music 3.0, Generates 5-Minute Songs on 8GB GPUs
MiniMax 开源 Music 3.0,8GB 显存生成 5 分钟完整歌曲
⭐️ 8.0/10

MiniMax has open-sourced Music 3.0, a model that generates a complete five-minute song—including intro, chorus, and bridge—from lyrics and a style description, outputting 32kHz stereo audio. The entire model runs on a single 8GB VRAM GPU, and the release also includes 1,000 templates and an expansion skill. By requiring only 8GB of VRAM, Music 3.0 drastically lowers the barrier for generating full-length songs locally, making advanced AI music creation accessible to individual developers and small teams. The open-source release also encourages community customization and innovation in an increasingly competitive AI music landscape. The model outputs 32kHz stereo audio, which is below the CD-standard 44.1kHz but sufficient for most listening scenarios. Notably, it produces structurally complete five-minute tracks in a single pass, while the bundled 1,000 templates and an expansion skill help users create more varied results easily.

rss · 小互(@imxiaohu) · Aug 14, 05:38

Background: AI music generation models typically convert text prompts, lyrics, or style descriptions into audio, but many require substantial GPU memory and computing power, limiting their practical use. MiniMax Music 3.0 is notable for both its open-source availability and its ability to run on a modest 8GB GPU, offering precise control over genre and style. The 32kHz sample rate is a common digital audio standard used in consumer devices, lower than CDs (44.1kHz) but adequate for casual listening and prototyping.

References

Tags: #AI音乐生成, #开源模型, #MiniMax, #音频生成, #深度学习


GLM-5.3 Launches Amid a Week of Major AI Releases
GLM-5.3 发布,本周 AI 大模型密集上新
⭐️ 8.0/10

Z.ai (formerly Zhipu AI) released GLM-5.3 near the end of a week that already saw Grok 4.6, DeepSeek V4 Pro and DeepSeek Harness, and a DeepSeek API price increase. The tweet also notes Gemini 3.7 Flash was not even mentioned. GLM-5.3's release hot on the heels of other frontier models signals an accelerated arms race among AI labs, giving developers more leading open-weight options. It may reshape model selection for coding and agentic tasks. According to Z.ai's official docs, GLM-5.3 ranks among the top open-source models on mainstream benchmarks, with coding and agentic capabilities comparable to Claude Fable 5. The tweet itself disclosed no technical specifications, and DataLearner cautions against extrapolating GLM-5.2's 753B MoE architecture, 1M context, pricing, or MIT license to the new version.

rss · meng shao(@shao__meng) · Aug 14, 05:40

Background: GLM (General Language Model) is Z.ai's flagship family of large language models. Z.ai, formerly known as Zhipu AI outside China, is one of China's six 'AI tigers' and went public on the Hong Kong Stock Exchange in January 2026; it has released GLM under the MIT open-source license since July 2025. The current release came at the end of an unusually packed week that also included Grok 4.6, DeepSeek V4 Pro and the DeepSeek Harness developer preview. This context highlights the rapid pace of open-weight model development in 2026.

References

Tags: #GLM, #大语言模型, #AI发布, #人工智能


Qwen3.8-27B Is Now Available on Ollama
Qwen3.8-27B 现已支持 Ollama
⭐️ 8.0/10

Alibaba Qwen announced that Qwen3.8-27B is now available on Ollama, and Ollama confirmed support for running the model with Claude Code, OpenCode, Hermes Agent, and Pi via CLI commands. The release also includes an Apple Silicon-optimized MLX build identified as qwen3.8:27b-mlx. This integration brings a top-performing open-weights model at the 27B scale to developers through familiar local tools, lowering the barrier for agentic coding and professional AI workflows. It also reinforces Ollama's role as a key hub for running open models locally and reflects the growing importance of model–harness compatibility in the AI ecosystem. Ollama documents commands such as 'ollama launch claude --model qwen3.8' for Claude Code, and analogous commands for OpenCode, Hermes Agent, and Pi. The model card also reports a SWE-bench Pro score of 61.7 for Qwen3.8-27B.

rss · Qwen(@Alibaba_Qwen) · Aug 14, 17:44

Background: Qwen is Alibaba's open-source large language model family, and Qwen3.8-27B is the latest generation at the 27-billion-parameter scale, designed for coding, agentic tasks, and professional workloads. Ollama is a widely used local LLM runtime that lets users pull and run open models through a simple CLI and API. In this context, a 'harness' refers to an application framework or agent tool that connects the model to a developer's workflow. Availability on Ollama means developers can use these models inside their existing tools without custom infrastructure.

References

Tags: #Qwen, #Ollama, #大语言模型, #模型发布, #AI


Qwen3.8-2.4T-A95B: 2.4T-Parameter Open Model Goes Live on SiliconFlow
Qwen3.8-2.4T-A95B:2.4 万亿参数开源模型在 SiliconFlow 上线
⭐️ 8.0/10

Alibaba Qwen has open-sourced the Qwen3.8-2.4T-A95B model, and SiliconFlow has made it available with Day-0 support. The model is now live on SiliconFlow's platform for immediate use via API. This launch is significant because a 2.4T-parameter open-weight model with only 95B active parameters brings frontier-scale intelligence to developers through an accessible API. It enables autonomous coding, deep research, and end-to-end agent execution for serious workloads, impacting the broader AI ecosystem. Qwen3.8-2.4T-A95B has 2.4T total parameters and activates 95B per token, with support for a 1M-token context window and up to 128K output tokens. Pricing on SiliconFlow is $2.00 per 1M input tokens, $6.00 per 1M output tokens, and $0.25 per 1M cached input tokens.

rss · Qwen(@Alibaba_Qwen) · Aug 14, 16:50

Background: Qwen is Alibaba's open-weight large language model family, and the naming convention '2.4T-A95B' indicates a Mixture-of-Experts design with 2.4 trillion total parameters but only 95 billion active per forward pass. This sparse activation keeps inference costs manageable while retaining massive capacity. SiliconFlow is an AI inference cloud platform that offers optimized, API-accessible deployment for open-source models, making Day-0 availability valuable for developers who want immediate access to new models.

References

Tags: #AI/ML, #LLM, #Model Launch, #Qwen, #SiliconFlow


Qwen3.8-2.4T-A95B MoE Model Now Available on DeepInfra
Qwen3.8-2.4T-A95B 模型现已上线 DeepInfra
⭐️ 8.0/10

Alibaba's Qwen team announced that the new sparse Mixture-of-Experts model Qwen3.8-2.4T-A95B is now live on the DeepInfra API platform. The model features 2.4 trillion total parameters with 95 billion active parameters, 512 experts, and a native 262K token context window. This release makes a frontier-scale MoE model accessible through a managed API, potentially enabling developers to build advanced coding, agentic, and reasoning applications without owning massive GPU infrastructure. It also reflects the growing trend of deploying extremely large sparse models on cost-effective inference platforms like DeepInfra. According to DeepInfra's announcement, pricing is set at $2.00 per million input tokens, $6.00 per million output tokens, and $0.20 per million cached tokens. The model is built for coding, agentic workflows, and complex reasoning scenarios.

rss · Qwen(@Alibaba_Qwen) · Aug 14, 16:27

Background: Mixture of Experts (MoE) is a machine learning technique that divides a model into multiple specialized sub-networks, or 'experts,' and activates only a subset of them for each input. This allows models to scale to trillions of parameters while keeping inference compute manageable. DeepInfra is a platform that offers cost-effective and production-ready hosting for large machine learning models, enabling users to deploy them via simple API calls.

References

Tags: #LLM, #AI, #Qwen, #DeepInfra, #Model Release


Qwen3.8-27B open-source model launches with day-zero AMD local support
Qwen3.8-27B 开源模型发布,首发日即支持 AMD 本地运行
⭐️ 8.0/10

Alibaba's Qwen team released Qwen3.8-27B, a native multimodal dense open-weight model, with day-zero support for AMD Ryzen AI Max+ processors and Radeon AI PRO R9700 GPUs. The model can be run locally through LM Studio and lemonade server immediately at launch. This release strengthens the open-source local AI ecosystem by offering day-zero support on AMD platforms, giving developers a powerful alternative to NVIDIA/CUDA-centric workflows. It lowers barriers for local development of coding, agentic, and office automation applications on consumer and prosumer AMD hardware. Qwen3.8-27B is a dense 27B-parameter native multimodal model that excels at coding, agentic workflows, and office automation tasks. AMD's day-zero support covers Ryzen AI Max+ processors with integrated RDNA and XDNA accelerators, as well as a single Radeon AI PRO R9700 GPU, enabled via LM Studio and lemonade server.

rss · Qwen(@Alibaba_Qwen) · Aug 14, 16:18

Background: AMD GPUs have historically lagged behind NVIDIA in LLM software support because many AI tools are built around CUDA, while AMD's ROCm stack provides an open-source GPU compute platform for LLM inference on supported hardware. Day-zero support means the model and toolchain were tested and optimized for these AMD platforms before or at public release, in collaboration with the AMD team. Qwen3.8-27B is part of Alibaba's latest Qwen3.8 open-weights series and is positioned as a strong local-performance model for developers.

References

Tags: #Qwen, #LLM, #AMD, #Open Source, #Local AI


Qwen and SGLang Achieve 206 tok/s on RTX 5090 for Qwen3.8-27B
Qwen 联合 SGLang 在 RTX 5090 上实现 206 tok/s 推理速度
⭐️ 8.0/10

Alibaba Qwen announced that SGLang has launched day-0 support for the open-source Qwen3.8-27B model, achieving 206.1 tokens per second decode speed on a single RTX 5090 GPU. The performance is enabled by NVFP4 quantization and DSpark technology, with 38.28 tok/s also recorded on DGX Spark. This milestone shows that small open-source models can run at real-time speeds on consumer-grade hardware, lowering the barrier for local deployment of agentic AI. It also underscores the importance of close collaboration between model developers and inference engines like SGLang. The benchmark numbers were 206.1 tok/s decode on one RTX 5090 using NVFP4 plus DSpark, and 38.28 tok/s on DGX Spark. Qwen3.8-27B is designed for agentic planning and long-horizon tasks, and SGLang provides immediate support through its open-source framework.

rss · Qwen(@Alibaba_Qwen) · Aug 14, 15:55

Background: SGLang is an open-source framework for high-throughput serving of large language models, developed to reduce latency and increase inference efficiency. Qwen is a family of large language models from Alibaba, and Qwen3.8-27B is the latest small model in the lineup. Tokens per second is a standard metric for LLM inference speed, and the RTX 5090 is NVIDIA's latest consumer GPU. 'Day-0 support' means the inference framework is optimized and ready on the same day the model is released.

References

Tags: #LLM Inference, #SGLang, #RTX 5090, #Performance, #GPU


Qwen3.8-Max Launches on Fireworks with Day-0 Support
Qwen3.8-Max 在 Fireworks 上线,提供 Day-0 支持
⭐️ 8.0/10

Alibaba's Qwen3.8-Max model (Qwen3.8-2.4T-A95B) is now available on Fireworks AI, with Fireworks acting as a Day-0 launch partner. The 2.4-trillion-parameter Mixture-of-Experts model is positioned for agentic workloads, heavy coding, and long-context tasks. This gives developers immediate production access to one of the largest open-weight models in the Qwen 3.8 series through Fireworks' high-performance inference platform. It could accelerate adoption of Alibaba's flagship model for coding agents and long-context applications, intensifying competition among open-model providers. The model is a Mixture-of-Experts architecture with 2.4 trillion total parameters and approximately 95 billion active parameters (A95B). It is available at fireworks.ai/models/fireworks/qwen3p8-max and is optimized for coding, agentic workflows, and large context windows.

rss · Qwen(@Alibaba_Qwen) · Aug 14, 15:26

Background: Qwen is a family of large language models developed by Alibaba Cloud; Qwen3.8-Max is the flagship general-availability model in the Qwen 3.8 series, succeeding the Qwen 3.8 Max Preview and supporting multimodal reasoning, visual understanding, coding, and agentic workflows. Fireworks AI is an inference and model-serving platform that hosts open-source models such as Qwen, Llama, and DeepSeek, focusing on speed and cost-efficiency. Large context windows allow models to process lengthy inputs, while MoE architectures scale parameters efficiently by activating only a subset per token.

References

Tags: #Qwen, #LLM, #AI, #Model Release, #Fireworks


Alibaba Qwen Details Performance of New Qwen3.8-27B
阿里公布 Qwen3.8-27B 性能
⭐️ 8.0/10

Alibaba's Qwen team announced performance details for its new Qwen3.8-27B model via a Tweet, showcasing benchmark charts. The model is a native multimodal dense open-weight model built on the Qwen3.5 architecture, delivering top-tier performance for coding, agentic workflows, and office automation. Qwen3.8-27B offers compact, deployment-friendly performance on local hardware, broadening access to high-performing multimodal AI for developers and enterprises. As an open-weight release from a major AI lab, it strengthens the open-source ecosystem and intensifies competition among mid-size models. According to the Hugging Face card, Qwen3.8-27B is evaluated on MathVision using a fixed prompt requiring step-by-step reasoning with \boxed{} formatting, while other models report higher scores from two prompt variants. It is a dense vision-language model with flexible thinking control and supports local deployment via LM Studio.

rss · Qwen(@Alibaba_Qwen) · Aug 14, 15:02

Background: Qwen is Alibaba's family of open-source large language models, covering dense and mixture-of-experts architectures. 'Native multimodal' means the model can directly process both text and images without separate vision encoders. Dense models activate all parameters during inference, offering predictable performance, and open-weight releases allow developers to fine-tune and deploy them on their own infrastructure.

References

Tags: #Qwen, #LLM, #AI, #Model Performance


Alibaba Qwen Releases Open Weights for Qwen3.8-27B and Max-Level MoE
阿里 Qwen 发布 Qwen3.8-27B 及 Max 级 MoE 开源权重
⭐️ 8.0/10

Alibaba's Qwen team released the open weights for Qwen3.8-27B, a native multimodal dense model that outperforms Qwen3.7-Plus, and also released the Max-level Qwen3.8-2.4T-A95B weights. Both are now available for download under the Apache 2.0 license. This release makes a high-performance multimodal model with a permissive license widely accessible, enabling developers to run it locally or build agents without vendor lock-in. It also underscores the growing trend of open-weight releases competing with closed frontier models. Qwen3.8-27B has 27B parameters, supports a native context length of 262K tokens, and can be extended to 1M tokens via YaRN. The Qwen3.8-2.4T-A95B is a Mixture-of-Experts model with 2.4 trillion total parameters, activating 95 billion per token.

rss · Qwen(@Alibaba_Qwen) · Aug 14, 15:02

Background: Open-weight models release the trained neural network parameters, allowing anyone to download and run the model locally, though they typically do not include the full training data or codebase. Dense models use all parameters for every input, while Mixture-of-Experts (MoE) models route inputs to specialized subsets of experts for greater efficiency. YaRN is a positional interpolation technique that extends the context window of rotary-position-embedding-based models beyond their original training length.

References

Tags: #AI, #LLM, #Open Source, #Qwen, #Multimodal


Qwen Teases Qwen3.8-27B Release on HuggingFace Within Hours
Qwen 预告 Qwen3.8-27B 即将在 HuggingFace 发布
⭐️ 8.0/10

Alibaba's Qwen team posted a teaser on X saying a new model, Qwen3.8-27B, would appear on HuggingFace in less than two hours, with a countdown image and link to the model page. This release signals continued momentum in open-weight large language models from one of the leading AI labs, giving developers another strong dense multimodal option. It could deepen competition with other open-weight model families and shape the ecosystem around HuggingFace. According to early model listings, Qwen3.8-27B is a dense 27B-parameter causal language model with a vision encoder, built on the Qwen 3.5 architecture. It supports a 262,144-token native context, extendable toward 1M tokens with RoPE scaling, and carries open weights.

rss · Qwen(@Alibaba_Qwen) · Aug 14, 13:08

Background: Qwen is Alibaba's family of open-weight AI models, widely used and downloaded on HuggingFace. Model names like Qwen2.5 and Qwen3 established the series as a popular alternative to other open LLMs, often offering long context windows and multimodal capabilities.

References

Tags: #LLM, #Qwen, #HuggingFace, #Model Release, #AI


Travis Kalanick Reveals He Was Quietly Building Industrial AI Company Atoms
卡兰尼克公开透露一直在默默打造工业 AI 公司 Atoms
⭐️ 8.0/10

Travis Kalanick publicly revealed that he has been quietly building Atoms, an industrial AI company, for the past eight years. The announcement came during a fireside chat with Ben Horowitz, who said the investment in Atoms is the largest check he has ever written. This marks a major comeback for Kalanick, one of the most prominent figures in the tech and venture capital world, and signals a significant bet on industrial AI as the next frontier. The size of a16z's investment underscores the high expectations for automating physical industries like manufacturing, real estate, and logistics. Atoms frames manufacturing, real estate, and logistics as the CPU, storage, and network of the physical world, and aims to deploy 'gainfully employed robots.' According to TechCrunch, Atoms raised $1.7B in a round led by a16z, with Uber also participating.

rss · a16z(@a16z) · Aug 14, 15:32

Background: Kalanick is best known as the co-founder and former CEO of Uber, which he left amid controversies in 2017. Industrial AI refers to AI systems purpose-built for high-stakes physical environments, focusing on reliability, predictability, and explainability, unlike general-purpose AI. The company's vision, described on Atoms' website, is to create specialized robots with productive jobs that bring abundance to their owners and society.

References

Tags: #AI, #Industrial AI, #Venture Capital, #Startups, #Podcast


15 Charts Reveal DeepSeek Harness's Plugin Architecture and AI-Driven Development
15 张数据图揭秘 DeepSeek Harness 的插件架构与 AI 辅助开发
⭐️ 8.0/10

A new data-driven analysis uses 15 charts to dissect DeepSeek Harness, the open-source agent harness DeepSeek released as a developer preview. The analysis reveals that its plugin system closely resembles Koishi's platform, that a large share of commits appear AI-generated, and that the project produced roughly 840,000 lines of code in 65 days. This matters because it offers one of the first independent, code-level looks at how DeepSeek builds production agent infrastructure and how heavily it relies on AI-assisted development. It also highlights the fast-growing ecosystem around DeepSeek Harness, which reached more than 80,000 GitHub stars within 20 hours of release. According to the analysis, Codex-related naming appears in 21.2% of mainline PRs and 28.2% of branch messages, and the project cites Pi, Codex, and Claude Code as the most-referenced external agent projects. The project also initially used a TUI interface, later moved to TUI plus Web UI, and finally removed TUI entirely, while a final schema of 52 tools and scientific models was reduced to one to lower context usage.

rss · 歸藏(guizang.ai)(@op7418) · Aug 14, 09:40

Background: DeepSeek Harness (dsh) is an open-source agent harness from DeepSeek AI in which models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and UI are all composed as plugins, powered by the Cordis framework. The plugin architecture resembles Koishi, a cross-platform chatbot framework built around a plugin system; this has led the analyst to speculate that DeepSeek may have reused large parts of Koishi or hired its core developers. Claude Code is Anthropic's agentic coding tool for developers, and the analysis suggests traces of such AI assistants were left in commit names and branch messages, indicating heavy AI-assisted development.

References

Tags: #DeepSeek, #AI Development, #Data Analysis, #Software Engineering, #Harness


ExtractBench Ranks Document AI on Scans, Rotations, Handwriting
ExtractBench 基准:评测扫描、旋转与手写文档提取
⭐️ 8.0/10

LlamaIndex announced ExtractBench, a new document-extraction benchmark covering scanned, rotated, and handwritten documents. Initial results across 14 systems show Codex excels at scans but struggles with rotations, while OCR APIs show the opposite pattern. Real-world documents are often messy — photocopied, faxed, or hand-filled — and this benchmark exposes that leading AI extraction systems have uneven 'perception blind spots.' ExtractBench provides a standardized way for enterprises and researchers to compare systems on the challenges that actually appear in production, not just clean digital PDFs. The benchmark includes documents such as 1950s regulatory filings, hand-filled tax forms, and pages degraded by fax thresholding, photocopier tone curves, sensor noise, and phone-camera capture. LlamaIndex's own Agentic Plus extraction tier was the only tested system with no major blind spot, scoring 95.9%/93.9%/93.8% on rotated, scanned, and handwritten documents respectively.

rss · Jerry Liu(@jerryjliu0) · Aug 14, 23:28

Background: Document extraction is the process of turning unstructured or semi-structured documents into structured data, often JSON, for downstream workflows. Clean digital PDFs are relatively easy for modern LLMs and OCR tools, but physical documents that have been scanned, rotated, or filled in by hand introduce perception challenges that many systems handle inconsistently. ExtractBench aims to standardize evaluation of these real-world conditions.

References

Tags: #document extraction, #benchmark, #OCR, #LLM, #AI


SQLite WAL-Reset Bug Found via Forensic Telemetry After Months of Outages
SQLite WAL 重置漏洞经数月遥测取证终被发现
⭐️ 8.0/10

A deep bug in SQLite's WAL (Write-Ahead Logging) code, named the 'WAL-Reset bug' by SQLite developers, was identified as the root cause of persistent production outages. It took months of deploying passive forensic telemetry in live environments to track down a bug that had likely been present for at least 16 years. SQLite is one of the most widely used database engines, embedded in countless applications, so a subtle corruption bug with real outage impact is highly consequential. This case also underscores the importance of production telemetry and long-term forensic debugging for diagnosing rare, non-reproducible race conditions. The bug involves a race condition in the WAL-index file (specifically the mxFrame and nBackfill fields) and appears to be triggered only under very specific timing conditions, making synthetic reproduction impossible. The team had to rely on passive forensic telemetry because there were no reliable trigger conditions to replicate the issue in a test environment.

rss · Michael Tsai · Aug 14, 18:59

Background: SQLite normally uses a rollback journal for atomic commits, but WAL mode improves concurrency and performance by writing changes to a separate log file. In WAL mode, a checkpoint operation merges the WAL back into the main database, and race conditions in coordinating this process can corrupt the database without any error messages. Such corruption may go unnoticed for months, especially when it only manifests under particular I/O timing or multi-process access patterns.

References

Tags: #sqlite, #database, #bug, #reliability, #forensics


Faraday: 27B Agent Beats Claude Opus 4.8 and GPT-5.5 on Research Replication
Faraday:270 亿参数智能体在研究复现中击败 Claude Opus 4.8 和 GPT-5.5
⭐️ 8.0/10

Researchers introduced Replica, a scalable RL task space for research replication, and used it to train Faraday, a 27B-parameter agent. Faraday outperformed Claude Opus 4.8 and GPT-5.5 on held-out research replication tasks, using an auto-generated rubric judge as the reward signal. This shows that a relatively small 27B model can beat much larger frontier models on scientific reasoning tasks when trained with the right RL objective. It points toward long-horizon scientific capability that is trained into weights rather than relying on complex external harnesses. Each Replica task requires an agent to replicate a figure from a research paper under a limited time and compute budget, without seeing the original plot. The initial suite contains 310 tasks drawn from 100 ML and AI-for-science papers across domains such as NLP and materials science.

rss · elvis(@omarsar0) · Aug 14, 17:00

Background: Research replication is the process of independently reproducing experimental results to verify claims, but it is rarely automated. Replica turns this into a scalable RL environment, where the reward comes from an auto-generated rubric judge that aligns with human assessment of replication quality. Faraday calls coding agents as tools, and rollout analysis suggests it learns scientifically principled strategies rather than gaming the rubric.

References

Tags: #AI, #Reinforcement Learning, #Research Replication, #Agent


Meta's Wiggle Framework: LLM Judges Flip Verdicts Up to 91% Under Pressure
Meta 的 Wiggle 框架:LLM 评委在压力下高达 91%的判决翻转
⭐️ 8.0/10

Meta's Wiggle Framework stress-tested nine frontier models across 14 judging tasks and found that LLM judge verdicts flip 25–71% of the time under static re-prompting and 62–91% of the time against an adversarial persuader. This shows that standard accuracy-based validation on golden data does not capture whether a verdict can survive questioning. The findings challenge the common practice of validating LLM judges solely by their accuracy on golden data, since a high accuracy score says nothing about robustness under pressure. This has broad implications for LLM evaluation, AI safety, and any system that relies on model-based judgement. The framework evaluates stability along three axes: stability under re-prompting, stability under a single challenge, and stability under sustained pressure. The paper also reports that pressure that changes a verdict is almost always net-corrupting against ground truth, and that baseline jury majority strength is the best single-shot predictor of which items will move.

rss · elvis(@omarsar0) · Aug 14, 15:50

Background: LLM judges are language models used to evaluate outputs from other models, often validated against golden data—human-labeled reference answers. However, this paper argues that such validation ignores whether a judge's verdict remains stable when challenged. The Wiggle Framework introduces stress tests for LLM judges, and the pre-print is available on arXiv.

References

Tags: #LLM, #evaluation, #Meta, #AI safety, #research


Open-Weight Model GLM-5.3 Hits Frontier Performance
开源权重模型 GLM-5.3 达到前沿性能
⭐️ 8.0/10

Z.ai has unveiled GLM-5.3, an open-weight model that the company claims delivers frontier-level performance rivaling OpenAI and Anthropic. The model features top-tier coding and agentic capabilities, built through post-training on a 743B base model. This milestone signals that open-weight models are closing the gap with closed proprietary systems, potentially giving developers and enterprises cheaper, more customizable access to frontier AI. It could intensify competition in the AI model market and accelerate adoption of open alternatives. GLM-5.3 is positioned for coding and cyber defense, with cybersecurity capabilities described as a major leap among open models. According to Z.ai, the model's top-tier coding and agentic performance comes from post-training the 743B-parameter base model; the technical blog is available at z.ai/blog/glm-5.3.

rss · elvis(@omarsar0) · Aug 14, 15:23

Background: Open-weight models are AI models whose trained parameters are publicly released, allowing anyone to download, run, and fine-tune them on their own hardware. Z.ai is a Chinese AI company behind the GLM series of open-weight models; GLM-5 is its new-generation foundation model designed for agentic engineering and long-range tasks. Post-training on a large base model is a common technique to specialize a model for tasks like coding and security.

References

Tags: #AI, #Open-Weight Models, #GLM, #Machine Learning


Coding Harness Nearly Solves ARC-AGI-3, Amjad Masad Claims
编码工具链几乎攻克 ARC-AGI-3,Amjad Masad 宣称
⭐️ 8.0/10

Amjad Masad claims that adding a coding harness nearly solves the ARC-AGI-3 benchmark, citing a 96.2% score achieved by Jeremy Berman using Opus 5. The tweet suggests that coding generalizes LLMs, a prediction that appears to be confirmed. This supports the emerging view that coding abilities can generalize LLMs to abstract reasoning tasks, potentially accelerating progress toward AGI. It also highlights that benchmark performance may depend more on scaffolding than on model architecture alone. Jeremy Berman reported 96.2% on ARC-AGI-3 using Claude Code plus Opus 5 (high) with one action command and filesystem logs, and 99.3% pass@2. The setup is almost entirely task-agnostic, with 'almost nothing ARC specific' according to his tweet.

rss · Amjad Masad(@amasad) · Aug 14, 04:45

Background: ARC-AGI (Abstraction and Reasoning Corpus for Artificial General Intelligence) is a benchmark designed to test a system's general fluid intelligence through small grid-based reasoning tasks that are easy for humans but hard for AI. A coding harness is the infrastructure layer that wraps a large language model with tools, permissions, and context management, turning it into an autonomous coding agent.

References

Tags: #ARC-AGI, #LLM, #Coding, #AI Generalization, #Benchmarks


Anthropic Publishes Second Risk Report Under Responsible Scaling Policy
Anthropic 发布负责任扩展政策下的第二份风险评估报告
⭐️ 8.0/10

Anthropic announced on X (Twitter) the release of its second Risk Report, now available at anthropic.com/aug-2026-risk-report. The report details the risks of its AI systems and the company's preparedness to address them under its Responsible Scaling Policy. This release is significant for AI safety governance because it increases transparency about frontier model risks. It affects the broader AI industry, policymakers, and researchers who rely on such disclosures to understand and mitigate AI-related dangers. The report is the second in a regular series required by Anthropic's Responsible Scaling Policy (RSP), which ties model deployment to safety thresholds. The tweet shows high engagement (211 comments, 402,358 views), indicating substantial public interest in AI risk transparency.

rss · Anthropic(@AnthropicAI) · Aug 14, 18:00

Background: Anthropic introduced its Responsible Scaling Policy in September 2023 as a framework to anticipate and secure against emerging threats from increasingly powerful AI models. The policy requires regular risk reports to share detailed information on system risks and preparedness. The second Risk Report continues this transparency effort.

References

Tags: #AI Safety, #Responsible Scaling, #Risk Report, #Anthropic, #AI Policy


After DeepSeek V4 Flash, LLM Race Shifts to Intelligence Efficiency
DeepSeek V4 Flash 之后,大模型开始卷「智效比」
⭐️ 8.0/10

DeepSeek has released V4 Flash, a Mixture-of-Experts model with 284B total parameters (13B active) that supports million-token context, positioning it as a fast and economical option. The release has pushed the large-model industry to shift its competitive focus from raw capability to 'intelligence efficiency ratio' (智效比). This shift is significant because in the Agent era, models are expected to autonomously execute long, multi-step tasks, making efficiency and cost as important as raw intelligence. Developers and enterprises will increasingly choose models based on intelligence-per-dollar rather than peak benchmark scores. The V4-Flash-0731 official version was released on July 31, 2026, marking a transition from preview to production-ready status, with optimizations for long-horizon Agent and engineering tasks. DeepSeek V4 Pro remains unchanged for now, with an official version expected soon.

rss · 爱范儿 · Aug 14, 02:14

Background: Most large language models are evaluated on capabilities such as reasoning and context length, but deployment also depends on inference cost and speed. MoE architectures reduce compute by activating only a subset of parameters per token. DeepSeek is a Chinese AI research company known for open-weight models. The rise of agentic AI, where models autonomously carry out multi-step tasks, makes long context and efficiency critical selection criteria.

References

Tags: #DeepSeek, #大模型, #AI效率, #模型发布


Hacker News Daily Digest Highlights AI Model Releases and Security Concerns
Hacker News 每日摘要:AI 模型发布与安全担忧
⭐️ 8.0/10

The August 15, 2026 Hacker News digest covers ten major stories, headlined by the open-source GLM-5.3 model with significantly improved coding and emergent cyber capabilities, and Google's official release of Gemini 3.7 Flash at half the price of its predecessor. This digest reflects the accelerating race in AI model development, particularly open-weight models approaching closed-source performance while introducing new security risks. It also shows intensifying price competition and performance gains in coding and agentic tasks that directly affect developers and enterprises. Notable items include GLM-5.3 scoring 84.5% on CyberGym and 54.4% on ExploitBench, Gemini 3.7 Flash priced at $0.75 input / $3.75 output per million tokens, and Cerebras and OpenAI's GPT-5.6 Sol Ultrafast reaching 750 tokens per second. The digest also covers Qwen3.8-27B-FP8 with 262K context, Opus 5's usability issues, and a satirical page about annoying web design patterns.

rss · HackerNews每日摘要 on SuperTechFans · Aug 14, 23:22

Background: Hacker News is a technology-focused social news site where users submit and discuss links, and daily digests curate the top-scoring stories. The current digest highlights trends in large language models: open-weight models like GLM-5.3 are closing the gap with proprietary models, while new model families like Gemini 3.7 Flash and Qwen3.8-27B-FP8 optimize for coding and agentic workflows. FP8 refers to 8-bit floating point format used to reduce memory and computation costs in AI inference.

References

Discussion: Comments on the GLM-5.3 story were mixed: many praised its strong performance and low cost (e.g., $10/month via OpenCode matching $200/month subscriptions), but others expressed concern about the model autonomously executing 0-day exploits without safety restrictions. Users also discussed the role of 'harnesses' like Claude Code and OpenCode in shaping model behavior, with preferences varying for manual review modes and security practices.

Tags: #AI, #Hacker News, #Tech News, #Security, #Open Source


AI proves zlib round-trip property by translating C to Lean
AI 将 zlib 的 C 代码翻译为 Lean,并证明压缩往返属性
⭐️ 8.0/10

In a recent podcast, a speaker reported that colleague Kim Morrison used AI to translate the C source of zlib, a widely used compression library, into the Lean proof assistant, then formally proved that compressing followed by decompressing restores the original data. The task was described as science fiction just six months earlier and was completed in about one week. This marks a significant milestone in AI-assisted formal verification: AI can now port real-world C libraries into a proof assistant and prove deep correctness properties, not just toy examples. It suggests that formally verified software libraries could become far more practical, with AI handling the bulk of the verification work while humans review the proofs. The proof establishes the essential compression invariant that decompression inverts compression for zlib's Deflate-based data formats. After the initial proof, the stated work shifts to optimizing the code while ensuring the proofs remain valid, i.e., optimizations must not break the verified property.

rss · Ryan Peterman · Aug 14, 13:03

Background: zlib is a free, widely used lossless compression library that implements the Deflate algorithm and ships in many operating systems. Lean is an open-source proof assistant and functional programming language that lets users write mathematical proofs that are mechanically checked. Formal verification uses these proofs to show that software satisfies a formal specification for all possible inputs, which is traditionally extremely labor-intensive.

References

Tags: #AI, #formal verification, #Lean, #zlib, #software engineering


Anthropic announces watermark detection API for Claude-generated text
Anthropic 推出检测 Claude 文本的水印检测 API
⭐️ 8.0/10

Anthropic announced it will release a watermark detection API that lets third parties verify whether text was generated by Claude. The technology is based on Google DeepMind's SynthID and modifies token-selection randomness without sacrificing text quality. This gives independent platforms, publishers, and regulators a practical tool to authenticate AI-generated text and curb misinformation. It also marks a major step toward industry-wide provenance standards, following Google's lead with SynthID. The watermarking works by tweaking the randomness in word selection, but it has known limitations with fact-heavy text, code, and heavily rewritten content. Anthropic did not specify a release date or exact API terms in the announcement.

rss · The Decoder · Aug 14, 21:29

Background: SynthID is a Google DeepMind technology that directly embeds imperceptible digital watermarks into AI-generated images, audio, text, or video, allowing detection without degrading output quality. Text watermarking has been an active research area because large language models make it increasingly hard to distinguish human and AI writing. Third-party access to detection tools is seen as key to enabling broader trust in AI content online.

References

Tags: #AI, #Watermarking, #Anthropic, #Content Detection, #API


Alibaba's Qwen Releases Qwen 3.8 Open-Weight Models Under Apache 2.0
阿里巴巴 Qwen 发布 Apache 2.0 许可的 Qwen 3.8 开放权重模型
⭐️ 8.0/10

Alibaba's Qwen team has released Qwen 3.8, a dense 27-billion-parameter model, under the Apache 2.0 license. The model natively supports up to 262,000 tokens of context and is designed to outperform the larger Qwen 3.7 Plus in coding and office tasks. This release strengthens the open-weight AI ecosystem by offering a high-performance, locally deployable model with permissive licensing. It gives developers building local and agent-based applications a competitive alternative to larger proprietary models. Qwen 3.8 is a dense model, meaning it activates all its parameters on every input, unlike Mixture-of-Experts architectures. Its 262K-token context window supports long-document processing and complex agent workflows, while Apache 2.0 allows commercial use and modification with attribution.

rss · The Decoder · Aug 14, 17:01

Background: Open-weight models make advanced AI accessible to startups, universities, and enterprises without requiring costly training from scratch or per-task API fees. A dense model uses all parameters for every input, providing consistent performance, while a context window determines how much text the model can consider at once when generating output.

References

Tags: #AI, #Open Source, #Model Release, #Qwen, #Alibaba


Study challenges OpenAI, Anthropic claims that autonomous AI research is near
研究质疑 OpenAI 与 Anthropic 关于自主 AI 研究即将实现的声明
⭐️ 8.0/10

A study conducted with Princeton and the UK AI Security Institute found that frontier AI agents using Claude Opus 4.8 and GPT-5.6 Sol failed to produce publishable research papers, with original authors rating their outputs as 'Reject.' The models could handle research engineering but fell short on research judgment, creative problem-solving, and abandoning failed approaches. This contradicts recent claims by Anthropic and OpenAI that autonomous AI research is within reach. It provides empirical evidence for AI safety and evaluation debates, suggesting frontier models still lack critical human research skills. The agents were given six days, $3,000 in API credits, and GPU access to independently write papers based on unpublished NeurIPS submissions. While they managed the full research engineering pipeline, they struggled with research judgment, creative problem-solving, and knowing when to abandon failed approaches.

rss · The Decoder · Aug 14, 16:06

Background: Frontier AI models are the most advanced artificial intelligence systems currently available, trained on massive datasets to achieve state-of-the-art performance across many tasks. Companies such as OpenAI and Anthropic have suggested that these models could soon perform autonomous research, but this study indicates that the judgment and creativity required for genuine scientific work remain difficult to automate. The distinction between research engineering—running pipelines, writing code, formatting results—and research judgment—choosing promising directions, interpreting findings, abandoning dead ends—is central to evaluating how close AI is to truly independent research.

References

Tags: #AI research, #autonomous agents, #AI evaluation, #frontier models, #AI safety


OpenAI's Ultrafast mode makes GPT-5.6 Sol 14x faster on Cerebras hardware
OpenAI 推出 Ultrafast 模式:GPT-5.6 Sol 在 Cerebras 硬件上提速 14 倍
⭐️ 8.0/10

OpenAI has launched a new inference mode called 'Ultrafast' that delivers GPT-5.6 Sol at up to 750 output tokens per second, a 14x speedup, using Cerebras hardware under a $10 billion partnership. Alongside 'Standard' and 'Fast', this creates a three-tier pricing structure that turns inference speed into its own product. This move makes inference speed a premium product tier, potentially reshaping how AI services are priced and delivered. It will affect enterprises and developers who rely on low-latency AI responses, and it highlights Cerebras's wafer-scale hardware as a serious competitor to GPU-based infrastructure. The mode is powered by the $10 billion OpenAI-Cerebras partnership, and speed is measured in output tokens per second, which differs from total task time because reasoning models often require 'thinking' time before generating output. The three tiers are Standard, Fast, and Ultrafast.

rss · The Decoder · Aug 14, 14:21

Background: Cerebras makes the Cerebras Wafer-Scale Engine (WSE), a chip that is far larger than a typical GPU and is purpose-built for ultra-fast AI workloads. GPT-5.6 Sol is OpenAI's flagship model in the GPT-5.6 series, launched in July 2026, and is suited for complex reasoning, coding, and agentic workflows. Output tokens per second measures how quickly a model generates text after it begins, but it does not capture the full latency including 'thinking' time. This launch reflects a broader trend of making inference speed a key differentiator in the AI market.

References

Tags: #OpenAI, #Cerebras, #GPT-5.6, #inference, #AI hardware


Zhipu AI releases GLM-5.3, claims strongest open-weights coding model
智谱 AI 发布 GLM-5.3,号称最强开放权重编程模型
⭐️ 8.0/10

Zhipu AI has released GLM-5.3, an open-weights coding model that it claims achieves a 50 percent improvement over its predecessor solely through post-training. According to Zhipu's benchmarks, GLM-5.3 helped security teams discover 2,436 vulnerabilities across 269 projects, and the model weights are scheduled to be open-sourced in two weeks. If independently verified, GLM-5.3 could become the most powerful open-weights coding model, offering a credible open-source alternative to proprietary coding assistants. Its strong vulnerability-finding performance also highlights the growing role of AI in cybersecurity, potentially changing how security teams approach code audits and exploit discovery. The claimed performance gain comes entirely from post-training, not from changes to the base model architecture or pre-training. The vulnerability-finding results are based on Zhipu's own benchmarks and have not yet been independently reproduced; the open-weights release in two weeks will allow the broader community to verify these claims.

rss · The Decoder · Aug 14, 10:21

Background: Open-weights AI models make their trained parameters publicly available, allowing developers to download and fine-tune them for specific tasks. Post-training refers to the additional fine-tuning stage after initial pre-training, which turns a generic language model into a specialized assistant for tasks like coding or security analysis. GLM is a family of large language models developed by Zhipu AI, a Chinese artificial intelligence company known for its competitive open-source models.

References

Tags: #AI/ML, #coding model, #open-weights, #GLM, #cybersecurity


AI Human-Tissue Labs Test 3M Samples a Year, Aim to End Animal Testing
AI 人体组织实验室年测 300 万样本,有望终结动物试验
⭐️ 8.0/10

Vivodyne's 12 robotic 'hive' labs in the San Francisco Bay Area culture human tissues and use AI to design experiments, enabling over 3 million controlled tests per year. This annual capacity is roughly double the total number of clinical trial participants in the United States. Because around 90% of clinical trials still fail after passing animal tests, better human-tissue models at this scale could sharply improve predictions of drug efficacy and safety. If validated, it could make animal testing obsolete and speed up drug development. Each robotic lab is about the size of a wardrobe, and the AI system designs experiments to optimize information gain. The technology is still early-stage and needs further validation before it can replace animal models in regulatory approval.

telegram · zaihuapd · Aug 14, 01:48

Background: Animal testing has long been the standard for predicting human drug responses, but it often fails to translate to human outcomes. Organ-on-a-chip and advanced tissue engineering aim to better mimic human physiology with cultured cells. Vivodyne grows realistic human tissues and uses AI plus automation to run high-throughput experiments, generating human-relevant data at large scale.

References

Tags: #AI, #drug testing, #tissue engineering, #robotics, #animal testing alternatives


Apple Seeks Supreme Court Review of App Store Fee Ruling, Wins Stay
苹果寻求最高法院复审 App Store 收费裁决,获暂停执行
⭐️ 8.0/10

Apple received court approval on April 6 to pause a ruling restricting its ability to charge fees on external payments, and plans to ask the U.S. Supreme Court to review the case. Epic Games immediately objected to the stay. The Supreme Court's decision could fundamentally reshape App Store economics and the 30% (or 27%) commission model, affecting millions of developers worldwide. It also sets a precedent for how antitrust law applies to digital platform gatekeepers in the U.S. The Ninth Circuit upheld a lower court's contempt finding against Apple in December 2025 for charging a 27% commission on developers using external payment systems. Apple received a stay from the appeals court on April 6, and Epic Games immediately challenged that stay.

telegram · zaihuapd · Aug 14, 02:33

Background: The case stems from Epic Games v. Apple, a 2020 antitrust trial over App Store rules requiring developers to use Apple's in-app purchase system and pay commissions. The district court issued an anti-steering injunction allowing developers to link to external payments, but Apple retained a fee for those transactions, leading to the contempt dispute. Apple now seeks Supreme Court review of the contempt ruling and the fee restrictions.

Tags: #Apple, #App Store, #Epic Games, #antitrust, #legal


Zhipu releases GLM-5.3 with 50% code-bench gain; weights open in two weeks
智谱发布 GLM-5.3,代码基准提升 50%,两周后开源权重
⭐️ 8.0/10

Zhipu AI released GLM-5.3, built on the GLM-5.2 base with all improvements coming from post-training. It scores 50% higher on Zhipu's internal Z.ai Code Bench than GLM-5.2 and reportedly leads open-source models on Terminal Bench 3.0, with weights to be open-sourced in two weeks. This release demonstrates that open-source LLMs can still improve rapidly through post-training alone, without training a new base model. The claimed security gains could also make open-weight models more viable for real-world vulnerability detection and enterprise use. All gains come from post-training, while technical details remain sparse. Security highlights include more than doubling exploit benchmark scores and helping security teams identify 2,436 vulnerabilities across 269 projects, 1,097 of which were medium-to-high severity.

telegram · zaihuapd · Aug 14, 05:27

Background: Large language models typically undergo pretraining followed by post-training stages such as supervised fine-tuning and preference alignment. GLM-5.2 is Zhipu's previous open-source model, which reportedly surpassed GPT-5.5 on SWE-bench Pro (62.1 vs 58.6) at a fraction of the cost. Terminal Bench is a widely used benchmark for measuring agents' ability to complete valuable work in containerized environments.

References

Tags: #GLM, #AI, #open-source, #LLM, #benchmark


Xiaohongshu Open-Sources dots3-note: 280B MoE with 16B Active Parameters
小红书开源 dots3-note:280B MoE 仅 16B 激活参数
⭐️ 8.0/10

Xiaohongshu's dots lab has open-sourced dots3-note preview, the first open-weight model in the dots3 series, on Hugging Face. The 280B-parameter Mixture-of-Experts model activates only 16B parameters per token and supports a 512K context window across text, image, video, and audio. This open-source release is significant because it demonstrates that a massive multimodal MoE model can be served with relatively modest compute, lowering the barrier for deploying large-scale models. The accompanying TEMPO reinforcement learning method and real-world agent benchmarks (VibeSearchBench, VibeLifeBench) could advance long-horizon agent training and evaluation. The model's TEMPO training method uses self-critique and test-time value estimation to train long-horizon agents. Alongside the weights, Xiaohongshu released VibeSearchBench and VibeLifeBench, two real-world agent benchmarks for evaluating such capabilities.

telegram · zaihuapd · Aug 14, 08:27

Background: Mixture of Experts (MoE) is an architecture that splits a neural network into specialized sub-networks called experts and uses a router to activate only the most relevant ones for each token, enabling massive scale with minimal compute. Reinforcement learning (RL) for LLM agents trains models through interaction with environments rather than relying solely on labeled data; test-time value estimation is a technique that estimates rewards during inference when ground-truth labels are unavailable. These techniques are central to building efficient large models and capable long-horizon agents.

References

Tags: #open-source, #MoE, #multimodal, #reinforcement-learning, #large-language-model


US Judge Orders Google to Remove Third-Party App Store Installation Barriers
美国法官责令谷歌取消第三方应用商店安装障碍
⭐️ 8.0/10

U.S. District Judge James Donato ordered Google to simplify installation of rival Android app stores within one week, removing extra steps and warning screens in the Play Store. The ruling follows the Epic v. Google antitrust case, where a jury found Google held an illegal monopoly in Android app distribution. This ruling directly reshapes how Android apps are distributed, potentially lowering barriers for alternative app stores and increasing competition. It affects Google's control over its ecosystem and could influence broader antitrust scrutiny of app store practices worldwide. The court characterized the step where users must tap 'view' before seeing an 'install' button as deliberately created 'anticompetitive friction' designed to deter average users. Google must make installing a third-party store as straightforward as installing a regular Android app.

telegram · zaihuapd · Aug 14, 09:55

Background: Epic Games filed an antitrust lawsuit after Google removed Fortnite from the Play Store for bypassing its payment systems, and a jury ruled against Google. The case has continued through appeals, with the U.S. government also weighing in on remedies. Sideloading on Android has long been a contentious issue, as Google balances security warnings against openness.

References

Tags: #Antitrust, #Google, #Android, #App Stores, #Epic Games


Apple trains China-specific AI model with Alibaba support
苹果携手阿里自研中国区 AI 大模型
⭐️ 8.0/10

Apple has reportedly trained its own large language model specifically for the Chinese market, with support from Alibaba. Apple Intelligence is expected to launch in China within the coming months alongside an iOS update, marking a shift from relying on third-party models. If approved, Apple would become the first foreign company allowed by Beijing to offer its own AI model in China. This could give Apple greater control over the AI experience in its largest overseas market and potentially reshape the competitive dynamics of China's AI industry. The Chinese cyberspace regulator (CAC) reportedly registered Apple's generative AI service last month. The model is being trained specifically for the China market, with Alibaba providing support, and will arrive with an iOS update in the coming months.

telegram · zaihuapd · Aug 14, 14:47

Background: Apple Intelligence is Apple's suite of AI features that currently relies on third-party models like ChatGPT in other regions. In China, foreign AI services face strict regulatory requirements, including security reviews and government approval, so many global companies partner with local firms. Training a dedicated model with Alibaba's help allows Apple to better comply with local rules and tailor AI to Chinese users.

Tags: #AI, #苹果, #阿里巴巴, #中国监管, #大模型



📊 Run stats · Total 16m 02s · AI analysis 4m 18s · Tokens 0.95 MCY (input 0.56 / output 0.40 MCY)