OpenAI introduces mathematician using GPT-5.6 to solve unsolvable problems
OpenAI 介绍数学家使用 GPT-5.6 解决无解数学问题 ⭐️ 10.0/10
OpenAI announced that mathematician Bartosz (@nasqret) is using GPT-5.6 to solve previously unsolvable mathematics problems, as demonstrated in a video posted by the company. This marks a major leap in AI reasoning capabilities, suggesting that large language models can now assist in groundbreaking mathematical research beyond human ability. It could accelerate discoveries in fields reliant on complex proofs. GPT-5.6 comes in three versions: Luna, Terra, and Sol, with Sol being the most capable. The video shows Bartosz collaborating with the model, though the specific version used is not disclosed.
rss · OpenAI(@OpenAI) · Jul 9, 20:04
Background: GPT-5.6 is a large language model released by OpenAI in July 2026, following a preview in June. AI theorem proving has been an active research area, with systems like Meta's neural theorem prover solving IMO problems, but GPT-5.6 claims to tackle unsolvable problems.
References
Tags: #AI, #Mathematics, #GPT-5.6, #Breakthrough, #OpenAI
OpenAI Announces ChatGPT Work Agent with Codex and GPT-5.6
OpenAI 发布 ChatGPT Work 智能体,结合 Codex 与 GPT-5.6 ⭐️ 10.0/10
OpenAI introduced ChatGPT Work, a new AI agent in ChatGPT powered by Codex and GPT-5.6, capable of autonomously taking actions across apps and files and completing complex projects over extended periods. This represents a major leap in AI productivity tools, enabling autonomous task execution and end-to-end project completion without constant human supervision, potentially transforming workflows across industries. ChatGPT Work is available on mobile, web, and desktop, allowing users to start tasks on one device and continue on another. It leverages Codex, OpenAI's software engineering agent, and GPT-5.6, which comes in three variants: Luna, Terra, and Sol.
rss · OpenAI(@OpenAI) · Jul 9, 17:41
Background: ChatGPT Work combines ChatGPT, a conversational AI assistant, with Codex, an agent designed for software engineering tasks. GPT-5.6 is OpenAI's latest large language model family, with Sol being the most capable variant. This integration allows the agent to reason, plan, and execute multi-step operations across applications, similar to an autonomous assistant.
Tags: #OpenAI, #ChatGPT, #Codex, #AI agent, #GPT-5.6
OpenAI launches GPT-5.6 family: Sol, Terra, Luna
OpenAI 推出 GPT-5.6 系列模型:Sol、Terra、Luna ⭐️ 10.0/10
OpenAI has begun rolling out the GPT-5.6 family of models, including Sol (flagship), Terra (balanced), and Luna (fastest/cheapest), across ChatGPT, Codex, and the API. The models promise state-of-the-art performance in coding, knowledge work, cybersecurity, and science with improved efficiency and lower cost. This release marks a significant advancement in AI capabilities, offering tiered options that balance intelligence and cost. It sets a new benchmark for language models and could accelerate adoption across industries requiring high-performance AI for coding, research, and security tasks. Sol is the most capable model, achieving state-of-the-art results on benchmarks like ARC-AGI-3. Terra offers strong performance at lower cost, while Luna is optimized for speed and efficiency. The rollout began as a limited preview on June 26, 2026, with general availability on July 9, 2026.
rss · OpenAI(@OpenAI) · Jul 9, 17:30
Background: GPT-5.6 is OpenAI's latest large language model family, succeeding GPT-5.5. It comes in three tiers to serve different use cases: Sol for complex reasoning and agentic tasks, Terra for everyday work, and Luna for cost-sensitive applications. The models incorporate improved intent understanding and safety measures. This release continues OpenAI's tradition of iterative model updates, with each version pushing the frontier of AI capabilities.
References
Discussion: Community reactions are mixed: some developers praise Sol's performance on benchmarks like ARC-AGI-3, while others note that in practical coding tests, Terra performs similarly to GPT-5.5 and slightly behind Sonnet 5. There is also discussion about the exclusion of Fable 5 from certain evaluations, with one commenter calling it a 'winner by default.' Comparisons between Codex and Claude Code are common, with users seeking consensus on which tool is better.
Tags: #OpenAI, #GPT-5.6, #AI models, #ChatGPT, #Codex
EU Parliament Passes Chat Control 1.0 via Procedural Trick
欧盟议会通过程序性手段批准聊天控制 1.0 ⭐️ 9.0/10
On July 9, 2026, the EU Parliament approved Chat Control 1.0, allowing US tech companies to scan private messages without a warrant or suspicion, after a motion to reject the law failed to achieve the required absolute majority of 361 votes, despite a majority of voting MEPs opposing it (314 against, 276 in favor). This decision enables mass surveillance of private communications on platforms like Instagram, Discord, and Gmail, undermining encryption and privacy for millions of users. It sets a precedent for EU surveillance laws and pressures messaging services to implement client-side scanning, which could be expanded to other content. The law was passed using a procedural maneuver: the vote was held on the last day before the summer break, and under urgency procedure, an absolute majority of all MEPs (rather than a simple majority of those present) was needed to block it. The scanning authorization lasts until 2028 and applies to direct messages and emails, while public posts and cloud storage remain unaffected.
hackernews · rapnie · Jul 9, 11:03 · Discussion
Background: Chat Control is an EU regulation aimed at detecting child sexual abuse material (CSAM) in private communications. The first version (1.0) relies on client-side scanning, where messages are scanned on the user's device before encryption or after decryption, raising serious privacy and security concerns. Critics argue it introduces a backdoor to encryption and violates fundamental rights, potentially leading to broader surveillance.
References
Discussion: Comments express outrage at the parliamentary tactics, noting that the vote was scheduled before the summer break to ensure low attendance, and that the law passed despite a majority voting against it. Commenters view this as a step toward totalitarianism and criticize the EU's legitimacy, with some pointing to Roberta Metsola's role in forcing the vote under urgency procedure.
Tags: #privacy, #surveillance, #EU, #encryption, #policy
OpenAI Unveils GPT-Live: Full-Duplex Voice with Real-Time UI
OpenAI 发布 GPT-Live:全双工语音加实时 UI ⭐️ 9.0/10
OpenAI released GPT-Live, a full-duplex voice model that supports simultaneous speaking and listening, real-time UI generation, and can delegate complex reasoning to GPT-5.5 in the background. GPT-Live could transform AI assistants by enabling natural conversational interaction with dynamic visual feedback, potentially disrupting traditional voice assistants like Siri and signaling a new paradigm for AI interfaces. The model outperforms Advanced Voice Mode in human evaluations and benchmarks like GPQA and τ³-Voice Telecom; it also features adjustable reasoning intensity (Instant, Medium, High) and enhanced safety guardrails for voice scenarios.
rss · 小互(@imxiaohu) · Jul 9, 02:29
Background: Full-duplex voice architecture allows AI to process input and generate output simultaneously, enabling natural interruptions and overlapping speech. GPT-5.5 is OpenAI's latest reasoning model, designed for deeper cognitive tasks. Real-time UI generation means the AI can create interactive visual elements on the fly, effectively acting as a personal assistant.
References
Tags: #语音模型, #OpenAI, #全双工, #AI助理, #GPT-5
OpenAI unifies Codex and ChatGPT into one desktop app with GPT-5.6
OpenAI 将 Codex 与 ChatGPT 整合为统一的桌面应用,并引入 GPT-5.6 ⭐️ 9.0/10
OpenAI announced a new desktop app that combines Codex and ChatGPT, introducing ChatGPT Work — a new agent powered by GPT-5.6 that can take actions across apps and files, along with a Chrome extension and revamped in-app browser. This integration marks a significant step in unifying AI-assisted coding and general-purpose AI assistance into a single platform, potentially streamlining workflows for developers and power users. The introduction of GPT-5.6 with advanced capabilities in coding, science, and cybersecurity could redefine productivity tools. The new ChatGPT Work agent can stay with a project for hours, turning a goal into finished work. The desktop app also features faster Computer Use powered by GPT-5.6, which is available as three model variants: Sol (flagship), Terra (lower-cost), and Luna (fastest and most cost-efficient).
rss · OpenAI Developers(@OpenAIDevs) · Jul 9, 17:48
Background: Codex is an AI coding agent from OpenAI that translates natural language into code and can handle tasks like pull requests, refactors, and migrations. ChatGPT is a general-purpose conversational AI. Combining them into one app aims to provide a seamless experience for both coding and general productivity tasks. GPT-5.6 is OpenAI's latest model family, previewed on June 26, 2026, with enhanced capabilities in coding, science, and cybersecurity, along with a robust safety stack.
References
Tags: #OpenAI, #Codex, #ChatGPT, #GPT-5.6, #Desktop App
Meta Releases Muse Spark 1.1 with Strong Agentic Abilities
Meta 发布 Muse Spark 1.1,具备强大智能体能力 ⭐️ 9.0/10
Meta CEO Mark Zuckerberg announced the release of Muse Spark 1.1, a new model with strong agentic and coding capabilities, available via the new Meta Model API at a very low price. This model's strong agentic performance at a low price could disrupt the AI market, offering a cost-effective alternative for developers and enterprises, especially in tool-use and long-horizon tasks. Muse Spark 1.1 leads on 4 out of 6 agent benchmarks, including MCP Atlas (88.1) and JobBench (54.7, up from 17.0), but lags in coding and multimodal tasks compared to GPT-5.5 and Opus 4.8.
rss · meng shao(@shao__meng) · Jul 10, 00:45
Background: Large language models (LLMs) are increasingly used for agentic tasks, such as tool calling and multi-step reasoning. Benchmarks like MCP Atlas and JobBench evaluate models' ability to use real-world tools and complete professional tasks. Sub-agents allow a model to delegate subtasks to parallel processes, improving efficiency.
References
Discussion: Community comments highlight the model's extremely low pricing ($1.25/$4.5 per million tokens) and its potential as a spoiler in the AI race. However, one commenter noted that the Terminal-Bench evaluation might be disqualified due to resource limits. Others shared hands-on experiences and discussed the competitive landscape.
Tags: #AI, #Meta, #LLM, #Agentic, #Coding
OpenAI CEO announces new best model and blog post
OpenAI CEO 宣布最新最强模型及博文 ⭐️ 9.0/10
Sam Altman, CEO of OpenAI, announced on X that the company has released its best model yet, along with a blog post describing it, with a link to openai.com/index/gpt-5-6/. This announcement signals a potential major breakthrough in AI capabilities, likely surpassing GPT-4, and could set new benchmarks in language model performance affecting numerous industries. The tweet explicitly calls it 'the best model we have ever produced' and the blog post 'one of the best we have ever produced'; the link suggests the model may be named GPT-5 or GPT-6, but official naming is unconfirmed.
rss · Sam Altman(@sama) · Jul 9, 17:10
Background: OpenAI is a leading AI research organization known for its GPT series of large language models. GPT-4, released in March 2023, was a significant milestone. A new model would represent the next generation in AI language understanding and generation.
Tags: #OpenAI, #GPT-5, #AI, #Language Model, #Announcement
Open-source model turns images into interactive 3D worlds
开源模型将图像转为可交互 3D 世界 ⭐️ 9.0/10
An open-source model now allows users to transform any static image into an infinite, fully interactive 3D world running at 720p and 60fps, with multiplayer support and dynamic actions like shooting and casting spells. A 1.3 billion parameter variant can run on consumer-grade GPUs and is available on Hugging Face. This represents a major leap in AI-generated interactive environments, democratizing the creation of immersive 3D worlds from single images without specialized hardware or game engines. It could revolutionize game prototyping, virtual tourism, and creative storytelling by enabling real-time, multiplayer exploration with minimal computational cost. The model supports infinite generation of 720p 60fps scenes and includes interactive capabilities such as shooting, casting spells, and summoning new elements. The 1.3B parameter variant is optimized for deployment on consumer GPUs, and both model weights and the harness (agent framework) are fully open-source.
rss · Paul Couvert(@itsPaulAi) · Jul 9, 13:42
Background: AI world models generate dynamic environments from static inputs, but earlier models (e.g., Oasis, Genie) were limited to basic navigation without gameplay interactions. The term 'harness' refers to the software framework that connects the AI model to user input and tools, enabling actions beyond simple movement. This new model integrates such a harness to support complex interactions like shooting and spellcasting, marking a shift from passive viewing to active gameplay.
References
Tags: #open source, #AI, #interactive world, #image to 3D, #Hugging Face
GPT-5.6 Now Available in Microsoft 365 Copilot
GPT-5.6 现已加入 Microsoft 365 Copilot ⭐️ 9.0/10
Microsoft CEO Satya Nadella announced that OpenAI's GPT-5.6 model is now integrated into Microsoft 365 Copilot, bringing stronger reasoning, multi-step agentic capabilities, and higher-quality outputs across Copilot Chat, Cowork, M365 apps, GitHub, and Foundry. This update significantly enhances productivity AI for millions of Microsoft 365 users, enabling more complex tasks like multi-step analysis and content creation. It also reinforces the deep partnership between Microsoft and OpenAI, setting a new benchmark for AI-assisted work. GPT-5.6 introduces a 'Work IQ' feature and is available immediately in Copilot Chat, Cowork, Microsoft 365 apps, GitHub, and Foundry. The model improves reasoning and output quality without sacrificing efficiency.
rss · Satya Nadella(@satyanadella) · Jul 9, 20:35
Background: Microsoft 365 Copilot is a generative AI assistant integrated into Microsoft's productivity suite, built on OpenAI's GPT models. It was launched in 2023 as a replacement for Cortana and has since become a central AI tool for work, offering features like chat, content creation, and collaboration. GPT-5.6 is the latest version of OpenAI's language model, featuring advanced reasoning and multi-step task handling.
References
Tags: #AI, #Microsoft, #GPT-5.6, #Copilot, #Productivity
AlloyDB Proxy Models Replace LLM Calls with Local Inference
AlloyDB 代理模型用本地推理替代 LLM 调用 ⭐️ 9.0/10
Google has shipped AlloyDB AI functions GA with a proxy model architecture that trains a lightweight local model from LLM outputs, enabling queries to run at database speed without external calls. Smart batching delivers up to 2,400x throughput improvement compared to standard LLM calls. This breakthrough significantly reduces latency and cost for AI-powered database queries by eliminating reliance on external LLM services. It sets a new standard for integrating AI inference directly within databases, potentially accelerating adoption of AI features in data-intensive applications. The proxy model reaches 100,000 rows per second in preview, though benchmark numbers apply only to the AI.IF function in internal testing. The training process uses LLM outputs to create a lightweight model that runs entirely within AlloyDB.
rss · InfoQ · Jul 9, 08:00
Background: Traditionally, AI-powered database queries require sending data to external LLM APIs, which introduces network latency and cost. Proxy models address this by training a small, efficient model locally on the database to mimic the LLM's behavior for specific tasks. This approach, inspired by proxy model architectures in LLM serving, allows AlloyDB to perform AI inference without external calls. The AlloyDB Auth Proxy is a separate utility for secure connections, not to be confused with the AI proxy model.
References
Tags: #database, #AI, #machine learning, #proxy models, #AlloyDB
Verification scaling emerges as new axis for LLM advances
验证扩展成为大模型进步的新维度 ⭐️ 9.0/10
Stanford AI Lab introduced LLM-as-a-Verifier, a framework that achieves state-of-the-art results on Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench by scaling verification. The method also improves reinforcement learning sample efficiency by using fine-grained feedback as a proxy for task progress. This work demonstrates that scaling verification—rather than just model size or data—can significantly boost performance, opening a new scaling axis for AI research. The fine-grained feedback mechanism enables more efficient reinforcement learning and test-time scaling, which could reduce computational costs and improve agent reliability. The framework uses finer scoring granularity (e.g., 1-20 instead of 1-5), takes the expectation over the full log-probability distribution of score tokens, and scales repeated evaluation and criteria decomposition. These techniques enable more nuanced and effective verification signals.
rss · Stanford AI Lab(@StanfordAILab) · Jul 9, 23:45
Background: Large language models (LLMs) are often used as judges or verifiers to evaluate outputs from other models. Traditionally, verification uses coarse scores (e.g., 1-5) and limited criteria. LLM-as-a-Verifier scales verification by increasing granularity and repeating evaluations, drawing an analogy to how inference-time compute (e.g., 'chain-of-thought') scales model reasoning. This work shows that verification itself can be scaled to improve performance independently of model size.
References
- GitHub - llm-as-a-verifier/llm-as-a-verifier · GitHub
- Cameron R. Wolfe, Ph.D. в X: «Strongly recommend the LLM-as-a-Verifier writeup. Biggest takeaway for me is that increasing scoring granularity makes the verifier more effective. This indicates that LLM judges / verifiers are developing new (and better) capabilities. This did not work well 1-2 years ago. https://t.co/4W7h3CREvn» / X
Tags: #LLM, #Verification, #Scaling, #AI Research, #Reinforcement Learning
TypeScript 7.0 Released with Go Rewrite, Up to 12x Faster Builds
TypeScript 7.0 发布:Go 重写带来最高 12 倍速度提升 ⭐️ 9.0/10
Microsoft has officially released TypeScript 7.0, a native version rewritten in Go that delivers 8–12x faster full builds and supports shared-memory multithreading. Users can install it via npm, and editors can use LSP support for the new language server. This major performance improvement dramatically reduces compile times for large TypeScript codebases, benefiting millions of developers and impacting the entire JavaScript ecosystem. The Go rewrite also demonstrates a new direction for compiler infrastructure in web development. The new version introduces the --checkers and --builders flags to customize parallelism. However, embedded language toolchains like Vue and Svelte still require the older version due to unfinished APIs.
telegram · zaihuapd · Jul 9, 04:01
Background: TypeScript is a typed superset of JavaScript that compiles to plain JavaScript. Previously, the TypeScript compiler was written in TypeScript itself, leading to performance bottlenecks. Rewriting it in Go, a systems language known for speed and concurrency, allows much faster compilation through native code and multithreading.
References
Tags: #TypeScript, #Microsoft, #performance, #Go rewrite, #release
Ant Group Open-Sources LingBot-Video, First MoE Embodied Video Model
蚂蚁集团开源 LingBot-Video,首个 MoE 具身视频基模 ⭐️ 9.0/10
Ant Group's LingBot has open-sourced LingBot-Video, the world's first Mixture-of-Experts (MoE) embodied video generation foundation model, with 30 billion total parameters and only 3 billion activated during inference, achieving 3x efficiency over dense models. This is a major breakthrough in embodied AI, combining video generation with robotics: the model can be used for robot action prediction, simulation data generation, and world model research, and its open-source release under Apache 2.0 will accelerate research and development in the field. LingBot-Video uses a DiT+MoE architecture and was trained on a 70,000-hour embodied data engine covering dexterous manipulation, robot locomotion, and first-person interaction. It also incorporates a multi-dimensional reinforcement learning reward system that emphasizes physical plausibility and task completion.
telegram · zaihuapd · Jul 9, 04:30
Background: MoE (Mixture of Experts) is an AI model architecture that uses multiple specialized submodels (experts) and a gating mechanism to activate only a subset of parameters per input, enabling larger model capacity with lower computational cost. Embodied AI integrates AI into physical systems that interact with the real world, such as robots. DiT (Diffusion Transformer) replaces the traditional U-Net backbone in diffusion models with a transformer, improving scalability and generation quality.
References
Tags: #embodied AI, #MoE, #video generation, #open source, #robotics
DJI EV50 Sets UAV Altitude Record at 8,861m on Everest
大疆 EV50 飞越珠峰 8861 米创纪录 ⭐️ 9.0/10
DJI's unreleased EV50 vertical takeoff and landing (VTOL) drone flew to 8,861 meters on Mount Everest during the 'Peak Mission' scientific expedition, setting a world record for the highest flight altitude in its class. It also collected real atmospheric profile data above 8,000 meters. This achievement demonstrates the extreme altitude capability of VTOL drones, potentially accelerating applications in high-altitude logistics, scientific research, and low-altitude cargo delivery. It also showcases DJI's technological leadership in the drone industry. The EV50 is a composite-wing VTOL drone that can take off vertically and transition to fixed-wing cruise. During the 12-day mission, it completed 32 sorties, climbing 3,730 meters continuously, and still had 30% battery remaining on return. DJI targets the EV50 for cargo transport over distances of 100 km or more.
telegram · zaihuapd · Jul 9, 06:00
Background: A composite-wing drone combines the advantages of fixed-wing aircraft (long endurance, high speed) and multirotors (vertical takeoff/landing, hover). Traditional multirotors have limited range and altitude due to battery constraints, while fixed-wing drones require runways. VTOL drones like the EV50 offer a flexible solution, making them ideal for missions in challenging terrain such as the Himalayas. DJI has a long history of testing drones on Everest, dating back to 2009.
References
Tags: #大疆, #无人机, #高海拔飞行, #物流, #珠峰科考
OpenAI Begins GPT-5.6 Livestream Launch Event
OpenAI 开始 GPT-5.6 发布会直播 ⭐️ 9.0/10
OpenAI has started a livestream for the launch of GPT-5.6, a next-generation large language model family that includes variants named Sol, Terra, and Luna. This event marks a major milestone in AI development, as GPT-5.6 promises stronger capabilities in coding, science, and cybersecurity, potentially advancing automation and research across industries. The livestream is hosted at openai.com/live, and the event began with a 10-minute countdown. GPT-5.6 is previewed as a family of models with an advanced safety stack, building upon the previous GPT-5.5 release from April 2026.
telegram · zaihuapd · Jul 9, 17:02
Background: OpenAI's GPT series are large language models (LLMs) that generate human-like text and perform complex tasks. GPT-5.5, codenamed 'Spud,' introduced notable benchmark scores and was used in cybersecurity initiatives. GPT-5.6 aims to further push the boundaries of AI capability and safety.
References
Tags: #OpenAI, #GPT-5.6, #AI, #LLM, #live
Tencent's Hy3 Small LLM Rivals DeepSeek
腾讯 Hy3 小型 LLM 可与 DeepSeek 媲美 ⭐️ 8.0/10
Tencent has released Hy3, a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters, which has quickly risen to the top of OpenRouter's rankings, sparking community discussion and comparisons with DeepSeek models. Hy3 demonstrates that a relatively small model can achieve capabilities comparable to much larger models like DeepSeek V4 Pro, potentially making high-performance LLMs more accessible for local deployment and cost-effective inference. Hy3 has 295B total parameters but only 21B active per token thanks to its MoE architecture, and a preview version was released in April 2026 with enhanced context learning and instruction-following. A free tier is available on OpenRouter until July 21, 2026.
hackernews · andai · Jul 9, 15:27 · Discussion
Background: Mixture-of-Experts (MoE) models activate only a subset of parameters per token, enabling efficiency gains while retaining large model capacity. OpenRouter is a unified API marketplace that aggregates LLMs from various providers. DeepSeek's V4 Flash and V4 Pro are popular competing models with similar parameter scales.
References
Discussion: Community members noted Hy3 topped OpenRouter rankings initially but has since fallen, with some questioning its advantage over competitors. Others praised its surprising capability for its small size, suggesting it could become a popular local model, while comparisons with DeepSeek Flash V4 drew interest.
Tags: #AI/ML, #Large Language Model, #Tencent, #OpenRouter, #LLM comparison
PostgreSQL Rewritten in Rust Passes All Regression Tests
用 Rust 重写 PostgreSQL 并通过全部回归测试 ⭐️ 8.0/10
A Rust rewrite of PostgreSQL, called pgrust and generated largely by LLMs, now passes 100% of the PostgreSQL regression tests, as announced by the author on GitHub. This project demonstrates the potential of using LLMs to rewrite complex, decades-old database systems, which could lead to more modern, performant, and safe implementations while preserving compatibility. The rewrite is largely generated by LLMs, resulting in over 7,000 commits in less than a month, but remains experimental and is licensed under AGPL instead of the original PostgreSQL license.
hackernews · SweetSoftPillow · Jul 9, 06:18 · Discussion
Background: PostgreSQL is a mature, open-source relational database known for its reliability and extensibility. Rewriting it in Rust, a systems language emphasizing safety and performance, could address memory safety issues and improve concurrency. The use of LLMs to generate the bulk of the code is a novel approach that raises questions about code review, licensing, and maintainability.
Discussion: The community is impressed but cautious. The author explains the project is experimental and a new version is being developed. Users suggest mirroring queries to test correctness and express concerns about the AGPL license change and the difficulty of reviewing LLM-generated code. Some question the value of rewrites and note potential compatibility issues.
Tags: #Rust, #PostgreSQL, #LLM, #Database, #Open Source
US Army logistics fragility in future war
美军后勤在下次战争中的脆弱性 ⭐️ 8.0/10
An article argues that the US Army's logistics system is too fragile and will fail in a major conflict due to over-reliance on air and sea lines and complex supply chains. This analysis challenges the Army's modernization priorities and highlights a critical vulnerability that could determine the outcome of future wars, affecting strategic planning and budgeting. The article references the tooth-to-tail ratio concept, historical examples like Fabian strategy, and current production rates compared to WWII to illustrate dwindling industrial capacity.
hackernews · baud147258 · Jul 9, 13:24 · Discussion
Background: Military logistics involves moving, supplying, and maintaining forces. The tooth-to-tail ratio compares combat units (tooth) to support units (tail). Modern logistics rely heavily on secure transport lines and just-in-time delivery, which can be disrupted by adversaries using drones or precision strikes.
Discussion: Comments largely agree with the article's thesis, drawing parallels to the Ukraine war and historical Fabian strategies. Some discuss alternative delivery methods like orbital drops, while others lament the decline in mass production capability compared to WWII.
Tags: #military logistics, #strategy, #defense, #supply chain
GLM 5.2 Matches Human Accuracy in Bookkeeping Benchmark
GLM 5.2 在记账测试中精度接近人类 ⭐️ 8.0/10
GLM 5.2, a new language model from Z.AI, achieved near-human accuracy in a bookkeeping benchmark, correctly coding transactions from bank feeds and invoices. The benchmark, however, excluded tasks such as finding invoices and handling ambiguous circumstances. This demonstrates that LLMs can automate core bookkeeping tasks with high accuracy, potentially reducing costs for businesses. However, liability and scope limitations remain significant barriers to adoption. The model has a 1M-token context window and is designed for long-horizon tasks. The benchmark focused only on transaction coding, not the full range of human bookkeeper responsibilities.
hackernews · adamkurkiewicz · Jul 9, 18:29 · Discussion
Background: GLM 5.2 is a flagship model from Z.AI for long-horizon tasks, succeeding GLM 5.1. The authors of the benchmark (Toot Books) aim to automate bookkeeping but note that the model cannot handle unannotated invoices or situations requiring judgment.
Discussion: Commenters highlighted that the benchmark excluded critical human tasks like finding invoices and handling edge cases, and raised concerns about liability if an LLM makes errors. Some expressed distrust due to lack of transparency about the startup behind the model.
Tags: #AI, #LLM, #benchmarking, #bookkeeping, #automation
OpenAI launches GPT-5.6 family, merges ChatGPT and Codex, introduces ChatGPT Work agent
OpenAI 发布 GPT-5.6 系列,合并 ChatGPT 与 Codex,推出 ChatGPT Work 智能体 ⭐️ 8.0/10
OpenAI has publicly released the GPT-5.6 family of models (Sol, Terra, Luna), merged ChatGPT and Codex into a single desktop application, and introduced a new agent feature called ChatGPT Work that can execute complex tasks across multiple applications. The updates also include a 'Sites' feature for sharing visualizations as web pages, enhanced browser capabilities, and new developer features for Codex. This unified platform strategy positions OpenAI as a direct competitor to established productivity suites like Google Workspace and Microsoft 365, aiming to lock in enterprise customers ahead of its IPO. The GPT-5.6 models with tiered pricing and the autonomous agent capabilities signal a shift from simple chatbots to comprehensive AI-powered work assistants. GPT-5.6 Sol introduces 'Ultra Mode' that internally spawns parallel subagents for complex tasks, available only to Pro and Enterprise users. The ChatGPT and Codex merger allows users to switch between 'Chat', 'Work', and 'Codex' modes while sharing context and history, with Codex now supporting direct file editing, PR reviews, and an Ultra mode for advanced coding. The ChatGPT Work agent can connect to Google Drive, Slack, email, and other tools, perform scheduled tasks, and operate on web and mobile.
rss · 宝玉(@dotey) · Jul 9, 18:39
Background: GPT-5.6 is the latest generation of OpenAI's large language models, with three tiers: Sol (flagship, high cost), Terra (mid-range, half the price of previous top model), and Luna (lightweight, low cost). The 'Ultra Mode' in Sol uses multiple sub-agents to work in parallel on complex tasks, a step towards more autonomous AI systems. OpenAI's application strategy consolidates its tools into a single platform to better compete with Google and Microsoft in the productivity software market. The company has filed for an IPO, with enterprise customers contributing 40% of revenue.
References
Tags: #OpenAI, #GPT-5.6, #AI, #Product Update, #LLM
Native SDK vs Flutter vs RN: Deep Framework Comparison
Native SDK 与 Flutter、RN 的深度框架对比 ⭐️ 8.0/10
A detailed analysis reveals that Native SDK uses a custom rendering engine like Flutter rather than native widgets like React Native, but employs a hybrid strategy by delegating scroll physics, context menus, and IME to the OS. It also compares Native SDK (Zig) to GPUI (Rust) across language, UI description, state management, and positioning. This comparison clarifies the trade-offs between major cross-platform UI approaches, helping developers choose the right framework for their needs. It also highlights emerging patterns like hybrid rendering and deterministic state management, which could influence future framework design. Native SDK's rendering is purely custom (no native UI components), but it delegates scroll physics, context menus, system dialogs, tray, and text input (IME) to the OS. Compared to GPUI: Native SDK uses Zig and separate .native template files, while GPUI uses Rust with UI code in Render trait implementations; Native SDK employs strict Elm architecture (Model/Msg/Update) for state management, whereas GPUI uses Entity system with reactive context.
rss · 宝玉(@dotey) · Jul 9, 04:34
Background: Cross-platform UI frameworks allow developers to build apps that run on multiple operating systems with a single codebase. React Native (RN) bridges JavaScript components to native platform widgets, while Flutter uses its own Skia-based rendering engine. GPUI is a Rust-based UI framework used in the Zed code editor, designed for high-performance text rendering and collaboration. Input Method Editors (IME) are system components that enable text input for complex scripts like Chinese.
References
Tags: #cross-platform, #UI framework, #Flutter, #React Native, #native SDK
Grok-4.5 reaches #3 in Code Arena: Frontend
Grok-4.5 在 Code Arena 前端基准测试中升至第三名 ⭐️ 8.0/10
SpaceXAI released Grok-4.5, its first model trained specifically for coding and agents, which jumped from #62 to #3 in the Code Arena: Frontend benchmark, tying with GLM-5.2 (Max) and Claude Opus 4.8 (Thinking). This significant improvement demonstrates Grok-4.5's enhanced frontend coding capabilities, making it a competitive tool for developers and narrowing the gap with top models like Claude and GLM. It also underscores the trend toward specialized coding models optimized for real-world tasks. The Code Arena benchmark measures human preference on open-ended coding and design tasks, not just static tests. Grok-4.5 also achieved #2 in content creation tools, simulations, gaming, and reference-based design.
rss · Arena.ai(@lmarena_ai) · Jul 9, 19:29
Background: Code Arena is a human-curated benchmark with 397 high-quality samples across 40 categories, designed to emulate real-world coding complexity. It evaluates how well models generate shippable code from user queries, bridging the gap between automated metrics and human judgment. Grok-4.5 is SpaceXAI's latest model, trained using Cursor and optimized for coding efficiency and cost.
References
Tags: #AI, #Grok, #Benchmark, #Code Arena, #Frontend
Vercel Labs Open-Sources Native SDK for Desktop Apps
Vercel Labs 开源桌面应用开发工具包 Native SDK ⭐️ 8.0/10
Vercel Labs has open-sourced Native SDK, a toolkit that allows developers to build native desktop applications using a declarative UI paradigm similar to the web, but without a browser or WebView. The SDK is built on a custom rendering engine written in Zig and supports macOS, Windows, and Linux. This addresses a long-standing trade-off in desktop development by combining the developer experience of web frameworks (declarative UI, hot reload) with native performance and a tiny binary size. It could challenge established frameworks like Electron and Tauri by offering a lighter, faster alternative. The SDK uses the Elm Architecture pattern with Model, Msg, update, and View components, implemented in Zig. In development mode, .native markup files are parsed at runtime with ~2s hot reload; in release mode, markup is compiled at comptime into the binary, eliminating the parser.
rss · meng shao(@shao__meng) · Jul 9, 01:06
Background: Desktop app development typically involves a trade-off between native solutions (Cocoa/Win32/GTK/Qt) that offer performance but poor cross-platform and web technologies (Electron/Tauri) that offer ease of development but bundle a heavy browser engine. Zig is a systems programming language focused on simplicity and performance, often used as a C replacement. The Elm Architecture is a pattern for structuring GUI applications that emphasizes a unidirectional data flow.
References
Tags: #Vercel, #Desktop Development, #Native SDK, #Declarative UI, #Zig
Local AI Harness Could Become the OS for Frontier Models
本地 AI 框架可能成为前沿模型的操作系统 ⭐️ 8.0/10
Aravind Srinivas, CEO of Perplexity AI, tweeted that a sufficiently advanced local AI harness, model, and runtime could become the primary gateway for using frontier AI models, with those frontier models acting as advisors within the harness, effectively turning the harness into an operating system for consuming frontier tokens. This vision challenges the current paradigm where frontier models are accessed primarily via the cloud, suggesting a future where local AI serves as the user's primary interface and orchestrator of powerful remote models. It could democratize access to frontier AI and shift control from centralized providers to users. The tweet specifically mentions that the harness would become an 'OS for using frontier tokens,' implying that local and remote models would coexist, with the local harness managing context, tools, and interactions. The comment cites 18 replies and 13 retweets, suggesting a focused discussion among AI practitioners.
rss · Aravind Srinivas(@AravSrinivas) · Jul 9, 22:42
Background: An AI harness, also known as an agent harness, connects an AI model to external tools, memory, and environments, enabling autonomous agent behaviors. Frontier models are the most capable general-purpose AI systems, typically accessed via API due to their massive scale. The idea of a local harness that can orchestrate both local and remote models is an emerging concept in AI infrastructure.
References
Tags: #AI, #Local AI, #Frontier Models, #AI Infrastructure, #Perplexity
1X's NEO Hand with 25 Tendons Achieves Human-Like Dexterity
1X 的 NEO 手部拥有 25 个肌腱,实现类人灵活性 ⭐️ 8.0/10
1X has unveiled its NEO humanoid robot hand with 25 tendon-driven degrees of freedom, demonstrating fine manipulation tasks such as screwing light bulbs, holding screws, and zipping zippers. This breakthrough removes a key hardware bottleneck in humanoid robotics, enabling robots to perform dexterous tasks previously limited to humans, and moves closer to general-purpose household robots. The hand features 25 actuated degrees of freedom with tendon-driven transmission, achieving near human-level speed and precision. It is designed specifically for the NEO humanoid platform.
rss · 歸藏(guizang.ai)(@op7418) · Jul 10, 00:26
Background: Traditional robot hands lack the dexterity for fine manipulation. Tendon-driven hands mimic human tendons by using cables to pull joints, allowing flexible and powerful motion. The NEO hand's 25-doF design approaches the human hand's capability, which is a significant engineering achievement.
References
Tags: #robotics, #dexterous manipulation, #robotic hand, #hardware engineering
Ollama Raises $65M Series B, Used by 8.9M Developers
Ollama 完成 6500 万美元 B 轮融资,拥有 890 万开发者 ⭐️ 8.0/10
Ollama announced a $65 million Series B funding round, with 8.9 million developers and 85% of Fortune 500 companies using its platform, all with just 14 employees. This significant investment validates the growing demand for open-source AI platforms, especially for local model deployment, and signals that Ollama is becoming the standard for developers to run open models easily. The Series B round was led by unnamed investors, and the funding will be used to scale open models and expand the platform. Ollama supports a wide range of open-weight models like DeepSeek, Qwen, and Gemma through a command-line interface and REST API.
rss · Y Combinator(@ycombinator) · Jul 9, 17:29
Background: Ollama is an open-source platform that simplifies running large language models on local computers. It provides tools for model management, a local REST API, and integrations with coding assistants. This funding follows the broader trend of enterprises adopting open-source AI for data privacy and customization.
Tags: #funding, #open-source, #AI, #developer tools
Tencent Open-Sources BrowserSkill for AI Agent-Browser Bridging
腾讯开源 BrowserSkill,连接 AI 智能体与浏览器 ⭐️ 8.0/10
Tencent has open-sourced BrowserSkill, a local bridge tool that enables AI agents like Cursor, Claude Code, and Codex to control a user's already-logged-in browser via shell commands without interrupting normal work. This tool addresses a critical gap for AI coding assistants needing browser interaction—such as filling forms or scraping data—while reusing existing sessions and adding a human-in-the-loop for security, making automation safer and more practical for developers. BrowserSkill uses a local daemon that communicates via WebSocket to a browser extension, executing tasks in a separate 'Agent Window' without affecting other tabs. It supports Chrome and Edge on macOS, Linux, and Windows, and triggers human takeover for CAPTCHAs or confirmation dialogs.
rss · Geek(@geekbb) · Jul 9, 06:25
Background: AI coding assistants can invoke shell commands to automate tasks, but they lack direct access to browsers that require authentication or complex interactions. BrowserSkill bridges this by allowing agents to use the user's existing browser session securely, with human verification for sensitive operations, rather than relying on traditional automation tools like Selenium or Puppeteer that cannot reuse sessions.
Tags: #AI agents, #browser automation, #open-source, #Tencent, #developer tools
DeepMind paper outlines four paths from AGI to ASI
DeepMind 论文提出 AGI 到 ASI 的四条路径 ⭐️ 8.0/10
Google DeepMind published a paper discussing four potential pathways from Artificial General Intelligence (AGI) to Artificial Superintelligence (ASI): scaling compute and data, replacing the Transformer architecture, recursive self-improvement, and collective intelligence via specialized agents. This paper provides a structured roadmap for the AI field, framing the transition to superintelligence as a series of iterative upgrades rather than a single breakthrough. It helps researchers and policymakers understand possible directions and risks on the path to ASI. The paper argues that ASI will likely emerge not from a single sudden event but from accelerating iterative improvements. Four distinct paths are identified: continued scaling, architectural innovation beyond Transformers, recursive self-improvement where AI accelerates its own development, and swarms of specialized agents forming collective superintelligence.
rss · AI Will(@FinanceYF5) · Jul 9, 02:29
Background: Artificial General Intelligence (AGI) refers to AI that can perform any intellectual task a human can, while Artificial Superintelligence (ASI) surpasses human capability across all domains. The Transformer architecture, introduced in 2017, is the foundation of most modern large language models. Recursive self-improvement (RSI) describes a scenario where an AI system redesigns itself, potentially leading to an intelligence explosion. The paper provides a systematic analysis of how the AI field might navigate the transition from AGI to ASI.
Tags: #AGI, #ASI, #DeepMind, #AI scaling, #recursive self-improvement
Yann LeCun: Open source AI key to sovereignty
Yann LeCun:开源 AI 是主权的关键 ⭐️ 8.0/10
Yann LeCun tweeted that the biggest risk of AI is the concentration of power in a few dominant providers of proprietary AI assistants, and that the only solution to AI sovereignty is open source foundation models. This statement from a leading AI researcher underscores the growing debate over AI governance, emphasizing the need for decentralized control to prevent power imbalance. It could influence policy discussions on open source AI and national AI strategies. The tweet received 72 comments, 239 retweets, and 1559 likes, indicating substantial engagement. LeCun specifically contrasts proprietary AI assistants with open source foundation models, aligning with his longstanding advocacy for open AI research.
rss · Yann LeCun(@ylecun) · Jul 9, 06:20
Background: Foundation models are large AI models trained on vast amounts of unlabeled data that can be adapted to a wide range of downstream tasks. AI sovereignty refers to the ability of a nation or organization to control its own AI infrastructure, data, and decision-making, rather than relying on external providers. Open source foundation models allow for transparency, customization, and independence from proprietary vendors.
References
Tags: #AI, #open source, #AI governance, #Yann LeCun, #proprietary AI
Super App Era Has Arrived
超级应用时代已到来 ⭐️ 8.0/10
The author declares that the super app era has begun, signaling a paradigm shift in how apps are designed and used. This matters because super apps integrate multiple services (messaging, payments, e-commerce) into one platform, reshaping user expectations and business models across industries. While no specific app is mentioned, the tweet's high engagement (1773 likes, 196 comments) reflects strong community interest; examples like WeChat, Gojek, and Grab illustrate the super app model.
rss · Logan Kilpatrick(@OfficialLoganK) · Jul 9, 19:43
Background: A super app is a mobile application that offers a wide range of services, such as messaging, payments, and e-commerce, within a single platform. Originating in Asia with apps like WeChat, the concept has gained global traction as companies seek to create integrated ecosystems.
References
Tags: #super app, #technology trends, #software engineering, #platform evolution
NVIDIA Flex-Forcing: Unified Bidirectional and Autoregressive Video Generation
NVIDIA Flex-Forcing:统一双向与自回归视频生成 ⭐️ 8.0/10
NVIDIA AI research team released Flex-Forcing, a video generation method that trains a single model to perform both bidirectional and autoregressive generation, allowing users to switch between methods at inference time based on their compute budget. This unification addresses a key trade-off between structure preservation and generation speed, potentially enabling more flexible and efficient video generation for applications like content creation and simulation. Flex-Forcing trains a single model on both generation paradigms, letting users choose any point along the spectrum at inference. This contrasts with prior work that requires separate models for each approach.
rss · NVIDIA AI(@NVIDIAAI) · Jul 9, 21:24
Background: Current video generation methods fall into two main categories: bidirectional diffusion models that attend to all frames simultaneously (good structure, slow) and autoregressive models that generate frame by frame (fast, but accumulate error). Flex-Forcing trains one model to support both, offering a flexible trade-off.
References
Tags: #video generation, #deep learning, #diffusion models, #autoregressive models, #NVIDIA AI
Llama.cpp seamlessly integrated into Zed editor v1.10
Llama.cpp 无缝集成到 Zed 编辑器 1.10 版本 ⭐️ 8.0/10
Llama.cpp, an open-source library for local LLM inference, has been seamlessly integrated into Zed editor version 1.10 with automatic model discovery, allowing developers to run AI assistance locally without relying on remote APIs. This integration makes local AI inference accessible directly within a popular code editor, reducing latency, enhancing privacy, and lowering costs for developers who want AI assistance without sending code to external servers. The integration is "completely seamless" with automatic model discovery, meaning Zed can locate and load compatible models without manual configuration. It leverages Llama.cpp's efficient C/C++ implementation for high-performance inference on consumer hardware.
rss · Julien Chaumond(@julien_c) · Jul 9, 12:27
Background: Llama.cpp is an open-source C/C++ library that enables running large language models (like Llama) locally on CPUs and GPUs with minimal setup, serving as the core engine for tools like Ollama and LM Studio. Zed is a high-performance, multiplayer code editor written in Rust, created by former Atom developers. This integration allows developers to use AI features (e.g., code completion, chat) directly within Zed using locally running models.
References
Tags: #Llama.cpp, #Zed editor, #AI integration, #local inference, #developer tools
Modal CTO: Kubernetes Not Fit for AI, Proposes Agent-Era Cloud
Modal CTO:K8s 不适合 AI,提出 Agent 时代云 ⭐️ 8.0/10
Modal CTO Akshat Bubna argues that Kubernetes is fundamentally unsuitable for bursty, compute-intensive AI workloads and proposes a new cloud paradigm built for agents, using self-provisioning runtimes and code-as-configuration instead of YAML. This challenges the dominant use of Kubernetes for AI infrastructure, suggesting that a purpose-built platform can significantly improve performance and developer experience for AI and agent workloads. It may influence how cloud providers design AI-specific infrastructure in the Agent era. Modal uses GPU snapshots, speculative decoding (DFlash), and sandbox environments to support elastic inference and agent sandboxing. The platform routes workloads across 17 cloud providers for optimal capacity and cost.
rss · 跨国串门儿计划 · Jul 9, 01:46
Background: Kubernetes is a container orchestration system originally designed for web services with gradual scaling, not for the sudden GPU bursts required by AI workloads. Modal is a serverless platform built from the ground up for AI, offering fast autoscaling and instant container boot times. Speculative decoding is an inference optimization where a smaller model proposes candidate tokens that a larger model verifies in parallel, reducing latency by 2-3x without sacrificing output quality.
References
Tags: #AI Infrastructure, #Kubernetes, #Cloud Computing, #AI Agents, #Podcast
AlphaEvolve now widely available on Google Cloud
AlphaEvolve 现已广泛部署于 Google Cloud ⭐️ 8.0/10
Google Cloud has made AlphaEvolve, its Gemini-powered AI agent for designing advanced algorithms, widely available after a private preview period. This rollout allows any Google Cloud customer to use AlphaEvolve to solve complex optimization problems. AlphaEvolve automates the discovery of optimal algorithms for challenging tasks like microchip design, logistics routing, and drug discovery, potentially saving significant time and improving efficiency. This could transform how organizations approach optimization in various industries. AlphaEvolve is a Gemini-powered coding agent that explores vast algorithm search spaces beyond traditional coding methods. It has already been used internally at Google to improve infrastructure efficiency and is now available to all Google Cloud customers.
rss · The Keyword · Jul 9, 16:00
Background: Designing efficient algorithms for complex problems is difficult because the search space of possible implementations is enormous. AlphaEvolve addresses this by using AI to automatically generate and test algorithms based on natural language descriptions of the problem. Google Cloud customers can now leverage this capability to optimize their own operations.
Tags: #Google Cloud, #AI, #AlphaEvolve, #optimization, #algorithms
Google Backs Open Health Stack Foundation for Global Health
谷歌支持成立开放健康栈基金会,推动全球健康 ⭐️ 8.0/10
Google announced the Open Health Stack Software Foundation (OHS-SF), established under the Linux Foundation with WHO and other partners, to advance open-source digital health. This foundation will provide community-governed digital public goods, lowering barriers for countries to build interoperable health systems, especially for marginalized populations. OHS-SF is supported by Google, WHO, UNICEF, and other global health organizations, focusing on open standards and interoperability for digital health.
rss · The Keyword · Jul 9, 08:00
Background: Digital health systems often face fragmentation and lack of interoperability, hindering effective care. The Open Health Stack is Google's set of open-source tools for building health applications. The new foundation puts these tools under community governance to ensure they are widely accessible and sustainable.
References
Tags: #health, #open source, #software foundation, #global health, #Google
OpenAI Fixes 18-Year-Old GNU libunwind Bug Using Epidemiology
OpenAI 用流行病学方法修复 18 年历史的 GNU libunwind 漏洞 ⭐️ 8.0/10
OpenAI discovered two unrelated bugs in ChatGPT's data infrastructure: silent hardware corruption on an Azure host and an 18-year-old race condition in GNU libunwind's setcontext function. They fixed both by applying population-level crash analysis instead of examining individual core dumps. This fix resolves a longstanding bug in a widely-used library, benefiting many applications that rely on libunwind. The novel debugging approach—treating crash data like epidemiological data—could inspire new techniques for diagnosing elusive bugs. The race condition had a one-instruction vulnerability window, making it extremely hard to detect. The two bugs masqueraded as one, complicating diagnosis. OpenAI collaborated with libunwind maintainers to implement the fix.
rss · InfoQ · Jul 9, 10:15
Background: GNU libunwind is a portable C API for determining the call chain of program threads and resuming execution, used in debugging, exception handling, and profiling. A race condition is a timing-dependent software bug that can cause intermittent failures. Population-level crash analysis aggregates crash data from many instances to identify patterns, similar to how epidemiologists study disease outbreaks.
References
Tags: #OpenAI, #GNU libunwind, #debugging, #bug fix, #Azure
Google Cloud Run sandboxes public preview for secure AI code execution
Google Cloud Run 沙箱公开预览,安全运行 AI 代码 ⭐️ 8.0/10
Google Cloud announced the public preview of Cloud Run sandboxes, a native, secure, and ultra-fast runtime environment for executing untrusted code and agent workloads, launching in milliseconds. This addresses a major pain point for developers who need to safely run AI-generated code without risking host applications or cloud credentials. It enables use cases like LLM code interpreters and headless browsers without complex sandboxing infrastructure. Sandboxes can be spawned near-instantly within existing Cloud Run service instances, as demonstrated by an example that started, executed, and stopped 1,000 sandboxes with an average latency of 500ms. The feature was announced at the WeAreDevelopers World Congress.
rss · Cloud Blog · Jul 9, 16:30
Background: Cloud Run is Google Cloud's serverless container platform. Sandboxing is a technique to isolate untrusted code from the host system. Previously, developers had to build complex infrastructure using container clusters or pay for specialized third-party microVM runtimes like Firecracker. Cloud Run sandboxes provide a simpler, integrated solution.
Tags: #Google Cloud, #Cloud Run, #sandbox, #AI-generated code, #security
TRACE: Self-Improving AI Agent Achieves 73.2% on SWE-bench
TRACE:自我改进 AI 代理在 SWE-bench 上达到 73.2% ⭐️ 8.0/10
Stanford AI Lab introduced TRACE, a self-improvement approach where an agent identifies missing capabilities behind its failures and trains itself to address them. The TRACE-trained Qwen3.6-27B model achieved 73.2% on SWE-bench Verified, outperforming much larger models like Codex 5.2 and GLM 5 while using less than a quarter of the training rollouts required by GRPO and GEPA. This breakthrough demonstrates that targeted self-improvement can achieve state-of-the-art results with significantly less computational cost, challenging the trend of scaling up model size. It points toward a more efficient paradigm for advancing AI capabilities in code generation and complex task solving. The TRACE approach identified specific missing capabilities by analyzing failure patterns and then trained on targeted data to address those gaps. The model used is Qwen3.6-27B, a 27-billion parameter model, which outperformed larger models including GPT-5.2-Codex and Claude 4.5 Sonnet on SWE-bench Verified.
rss · Stanford AI Lab(@StanfordAILab) · Jul 9, 23:45
Background: SWE-bench Verified is a human-filtered subset of 500 instances from SWE-bench, created in collaboration with OpenAI, to evaluate AI models on resolving real GitHub issues from popular Python repositories. GRPO (Group Relative Policy Optimization) and GEPA (Genetic-Pareto) are alternative training algorithms that rely on reinforcement learning with verifiable rewards. The concept of recursive self-improvement, where an AI system improves its own abilities, has been explored in AI safety and capability research.
References
Tags: #self-improvement, #AI, #SWE-bench, #machine learning, #code generation
Stanford D2D Exposes Hidden LLM Biases by Amplifying Them
斯坦福 D2D 通过放大隐性偏见来检测 LLM 偏见 ⭐️ 8.0/10
Stanford AI Lab researchers introduced Distill to Detect (D2D), a method that surfaces hidden biases in fine-tuned LLMs by distilling the shift between the suspect model and its base into a KV-cache prefix adapter, thereby amplifying subtle biases into detectable text. D2D addresses a critical blind spot in AI safety: hidden biases that only surface on unknown topics. This method enables auditors to detect biases they don't know to look for, improving trustworthiness of fine-tuned LLMs. D2D is a gray-box auditing method that trains a low-capacity adapter on benign prompts to concentrate Fisher-weighted bias signals while mitigating noise. In experiments, it achieved up to 100% detection rate for certain stealth biases.
rss · Stanford AI Lab(@StanfordAILab) · Jul 9, 23:30
Background: Fine-tuning large language models (LLMs) can inadvertently introduce subtle biases that evade standard auditing because they may only activate on specific topics. Traditional methods require prior knowledge of the bias to search for it. D2D amplifies such biases by distilling the model shift into a small cartridge, making them visible to existing tools.
References
Tags: #AI safety, #bias detection, #LLM interpretability, #machine learning, #fine-tuning
Cloudflare: Use ML-DSA Now, Don't Wait for New Algorithms
Cloudflare:立即使用 ML-DSA,无需等待新算法 ⭐️ 8.0/10
Cloudflare published a blog post arguing that while NIST evaluates nine new post-quantum signature algorithms, organizations should adopt ML-DSA now as the best currently available standard. This matters because post-quantum cryptography is urgently needed to protect against future quantum threats, and delaying adoption could leave systems vulnerable. ML-DSA is a NIST-approved general-purpose digital signature algorithm based on lattice cryptography, specified in FIPS 204, and is designed to replace RSA and ECC signatures.
rss · The Cloudflare Blog · Jul 9, 14:00
Background: Quantum computers may eventually break widely used cryptographic algorithms like RSA. NIST has been leading a standardization process to select quantum-resistant algorithms, releasing ML-DSA (CRYSTALS-Dilithium) as one of the first standards in August 2024. However, nine additional candidate algorithms are being evaluated for future standardization.
References
Tags: #post-quantum cryptography, #signature algorithms, #ML-DSA, #NIST, #cryptography
The Human Role in AI Outer Loop Engineering
人工干预在 AI 外循环中的关键作用 ⭐️ 8.0/10
A recent article explains why human intervention is crucial in the outer loop of AI and machine learning systems, emphasizing that loop engineering requires a human at the boundary. This insight is significant because as AI systems become more autonomous, maintaining human oversight in the outer loop ensures safety, alignment, and effective engineering, impacting AI engineers and developers. The article uses the concept of 'loop engineering' and distinguishes between inner loops (e.g., model training automation) and outer loops (deployment, monitoring, feedback), where human judgment is essential.
rss · Elevate · Jul 9, 14:31
Background: In machine learning, an inner loop refers to the iterative process of training and hyperparameter tuning, while the outer loop involves the broader system context including deployment, monitoring, and human feedback. Loop engineering is the practice of designing these loops to incorporate AI agents effectively.
References
Tags: #AI Engineering, #Human-in-the-loop, #Machine Learning, #Software Engineering, #Loop Engineering
National Supercomputing Internet Core Node Launches in Zhengzhou with 100K+ Domestic AI GPUs
国家超算互联网核心节点郑州上线,提供超 10 万卡国产算力 ⭐️ 8.0/10
On July 9, 2026, the core node of the National Supercomputing Internet officially went online in Zhengzhou, providing over 100,000 domestic AI computing cards (GPUs). This represents the largest single domestic AI computing resource pool ever connected to the National Supercomputing Internet platform. This milestone significantly boosts China's domestic AI computing infrastructure, reducing reliance on foreign GPUs and enabling large-scale AI model training and scientific computing. It strengthens the national computing power grid, which is critical for digital economy and national security. The core node in Zhengzhou serves as the hub for operations, management, and resource scheduling, integrating supply-demand matching and industry incubation services. It is part of the National Supercomputing Internet, which as of early 2026 had connected over 30 national supercomputing and AI computing centers across 14 provinces.
telegram · zaihuapd · Jul 9, 07:00
Background: The National Supercomputing Internet is a national initiative launched by the Ministry of Science and Technology to connect dispersed supercomputing centers into a unified computing resource network. It aims to provide seamless access to advanced computing resources (including AI training) for research and industry. An "AI computing card" (e.g., domestic GPUs) is a specialized accelerator optimized for AI workloads like matrix multiplications, distinct from general-purpose CPUs or gaming GPUs.
References
Tags: #超算, #AI算力, #国产算力, #基础设施, #HPC
OpenAI and US War Department Amend Contract to Ban Citizen Surveillance
OpenAI 与美国战争部修订合同,禁止监控本国公民 ⭐️ 8.0/10
OpenAI and the U.S. Department of War (the secondary title for the Department of Defense) have agreed to amend their AI collaboration contract to explicitly prohibit the use of AI systems for surveillance of American citizens. The amendment was proposed by OpenAI CEO Sam Altman to address concerns over AI-enabled mass surveillance. This move sets a precedent for AI companies restricting military use of their technology, potentially influencing industry norms and rebuilding public trust in AI ethics. It also highlights the increasing tension between AI innovation and government surveillance capabilities. The amended clauses explicitly prohibit deliberate surveillance of US citizens using AI systems and forbid tracking individuals through commercially acquired personally identifiable information. The contract has not been formally signed yet, and Anthropic previously ended a similar agreement with the War Department over comparable concerns.
telegram · zaihuapd · Jul 9, 13:22
Background: The U.S. Department of War was originally a cabinet department (1789-1947) overseeing the Army and early naval affairs, later replaced by the Department of Defense. In September 2025, President Trump signed an executive order authorizing 'Department of War' as a secondary title for the DoD. Anthropic, an AI safety-focused company founded by former OpenAI staff, previously ended its contract with the War Department due to surveillance concerns, highlighting ongoing tensions in military AI collaborations.
Tags: #OpenAI, #AI ethics, #surveillance, #US military, #AI policy
📊 Run stats · Total
17m 13s· AI analysis3m 42s· Tokens0.86 MCY(input0.59/ output0.27MCY)