AI Researcher Tian Yuandong Launches Recursive, a $4.65B Superintelligence Startup
AI 研究员田渊栋创立 Recursive,估值 46.5 亿美元的超智能公司
⭐️ 9.0/10

Tian Yuandong, former Meta FAIR director, officially announced his new company Recursive (Recursive Superintelligence) as co-founder, with over $650 million in funding and a $4.65 billion valuation, aiming to build recursive self-improving superintelligence. This represents a paradigm shift towards autonomous AI that can improve itself, potentially accelerating scientific discovery and technological progress. The high-profile team and significant funding indicate strong industry confidence in this approach. The company is named Recursive Superintelligence, with Richard Socher as CEO, and co-founders include Tim Rocktäschel, Jeff Clune, Alexey Dosovitskiy, Caiming Xiong, Josh Tobin, and Tim Shi. The core idea is that since AI is code and AI can now write code, the self-improvement loop can be closed.

rss · meng shao(@shao__meng) · May 14, 00:56

Background: Recursive self-improvement (RSI) is a concept where an AI system can rewrite its own code to enhance its capabilities, potentially leading to an intelligence explosion. This idea has been discussed in theoretical AI safety and AGI contexts for decades. The Recursive team aims to turn this theoretical concept into practical systems, building on recent advances in coding agents and automated research loops.

References

Tags: #AI, #AGI, #Startup, #Deep Learning, #Funding


Google reports first known real-world AI-crafted zero-day exploit
谷歌报告首个已知 AI 制作零日漏洞实例
⭐️ 9.0/10

Google reports the first known instance of an AI being used to craft a zero-day exploit in the wild.

rss · Hacker News: Newest · May 14, 02:05

Tags: #AI, #zero-day exploit, #cybersecurity, #threat intelligence, #Google


Android integrates on-device MCP for Gemini agentic actions
Android 集成设备端 MCP,让 Gemini 实现代理操作
⭐️ 9.0/10

Google announced at the Android Show that Android will have native Model Context Protocol (MCP) built into the OS, enabling apps to expose functions via the @AppFunction annotation for Gemini and other agents to perform cross-app actions locally on-device. This eliminates the need for cloud dependency in agentic AI actions on mobile, enabling faster, more private, and offline-capable AI experiences, potentially revolutionizing how users interact with apps through AI agents. The system works like MCP but runs entirely on-device without server or network round-trips, and is available on Android 16+ with an Early Access Program for testing.

rss · Philipp Schmid(@_philschmid) · May 13, 13:00

Background: The Model Context Protocol (MCP) is an open standard introduced by Anthropic to standardize how AI systems integrate with external tools and data. By embedding MCP natively into Android, Google allows any app to declare its functions as tools for AI agents, enabling seamless cross-app automation.

References

Tags: #Android, #Gemini, #MCP, #on-device AI, #agents


Xiaomi Unveils OneVL: One-Step Latent Reasoning for Autonomous Driving
小米发布 OneVL:一步式潜空间推理自动驾驶框架
⭐️ 9.0/10

Xiaomi released OneVL, a one-step latent space reasoning framework that unifies Vision-Language-Action (VLA) models and world models for autonomous driving, achieving state-of-the-art results on multiple benchmarks and being fully open-sourced. OneVL represents a breakthrough in autonomous driving by combining VLA and world model reasoning into a single latent space, significantly reducing inference latency to 0.24s (only 5.4% of autoregressive VLA) while surpassing explicit Chain-of-Thought methods. This could accelerate real-time deployment of intelligent driving systems. OneVL uses latent tokens to encode physical causal structures and driving intentions, with dual auxiliary decoders predicting future frames and readable CoT during training, which are removed at inference for one-step parallel generation. It achieved a PDM-score of 88.84 on NAVSIM, the first latent CoT method to surpass explicit CoT (88.29).

telegram · zaihuapd · May 13, 10:33

Background: Vision-Language-Action (VLA) models combine visual understanding, language reasoning, and action prediction for autonomous driving, while world models simulate future states. Traditional Chain-of-Thought (CoT) reasoning generates intermediate text, which is slow and loses spatial-temporal structure. Latent space reasoning compresses reasoning into compact vector representations, enabling faster inference while preserving rich structure.

References

Tags: #autonomous driving, #VLA, #world model, #latent space reasoning, #open source


AI Makes Personal Software as Easy as Customizing Emacs
AI 让个人软件制作像定制 Emacs 一样简单
⭐️ 8.0/10

A new article argues that AI, particularly large language models, has made building personal software as straightforward as customizing Emacs, ushering in an era of highly personalized tools. This shift empowers individuals to create bespoke software for their own needs, reducing reliance on prepackaged professional applications and democratizing software development. The article lists several application categories—such as podcast apps, feed readers, and note-taking tools—where AI can produce better-than-replacement-grade results, though not necessarily the globally best.

hackernews · rdslw · May 13, 07:06 · Discussion

Background: Emacs is a highly extensible text editor with over 10,000 built-in commands and a Lisp extension language, allowing users to customize nearly every aspect. The term 'Emacsification' metaphorically describes making software as customizable as Emacs. The article uses this metaphor to illustrate how AI enables anyone to craft personal software solutions.

References

Discussion: Hacker News users tptacek and dang strongly agree, with tptacek listing categories where AI already excels and dang coining the 'dot emacs' phrase. shaokind shares a mixed experience, noting that Emacs setup was brittle across platforms, while SoftTalker links the concept to the original vision of home computing.

Tags: #AI, #software development, #personalization, #LLMs, #Emacs


Princeton Ends 133-Year Honor Code, Mandates Proctored Exams
普林斯顿结束 133 年荣誉制度,改为监考考试
⭐️ 8.0/10

Princeton University has mandated proctored in-person exams for the first time in 133 years, ending its long-standing honor code that allowed unproctored exams. The decision, driven by concerns over AI-enabled cheating, was passed by the faculty in May 2026. This shift reflects a broader societal transition from high-trust to low-trust institutions, especially in academia, as AI tools like ChatGPT make cheating easier and harder to detect. It may prompt other universities to reconsider their honor codes and proctoring policies. Under the old honor code, students took exams without proctors and were required to report violations themselves. The new policy introduces active proctoring and likely device confiscation during exams, addressing concerns that nearly 30% of students admitted to cheating.

hackernews · bookofjoe · May 13, 20:12 · Discussion

Background: An honor code is a set of rules that relies on students to uphold academic integrity without external monitoring. Princeton's honor code, established in 1893, was one of the oldest in the U.S. and allowed unproctored exams, with violations adjudicated by a student-run body. The rise of generative AI has made it easier to cheat on exams and assignments, challenging the effectiveness of trust-based systems.

Discussion: Comments show mixed reactions: some alumni recall the honor code fondly, while others argue that proctoring is necessary in the age of AI. Critics lament the shift from a high-trust society to a low-trust one, citing that 30% of students admitted cheating under the old system. There is also discussion about the impracticality of relying on students to report peers.

Tags: #education, #AI cheating, #academic integrity, #proctoring, #honor code


Developer Migrates from GitHub to Forgejo
开发者从 GitHub 迁移到 Forgejo
⭐️ 8.0/10

A developer documented their migration from GitHub to Forgejo, citing desires for decentralization and greater control over their repositories and data. This highlights the growing trend of developers moving away from centralized platforms like GitHub towards self-hosted or federated alternatives, driven by concerns about AI scraping, corporate control, and the spirit of Git's decentralized design. Forgejo is an open-source self-hosted Git forge, a fork of Gitea, supporting features like bug tracking, code review, and continuous integration. The author's move may involve losing social graph and collaboration history, though tools like GitSocial can help.

hackernews · jorijn · May 13, 12:54 · Discussion

Background: Git was designed as a decentralized version control system, but GitHub became a dominant centralized platform offering convenience and features. Forgejo is an open-source forge that enables self-hosting and federation, aligning with Git's original decentralized spirit. It is a community-driven fork of Gitea, hosted on Codeberg.

References

Discussion: Comments discuss the importance of federation for decentralized git hosting, concerns about AI scraping, and tools like GitSocial for cross-forge collaboration. Some users note that mirrors on GitHub may die if not maintained. Overall sentiment is supportive of decentralization.

Tags: #Git, #Forgejo, #GitHub, #Self-hosting, #Open Source


Anthropic caps programmatic usage in Claude plans, adds separate credits
Anthropic 限制 Claude 订阅的程序化调用,新增独立额度
⭐️ 8.0/10

Starting June 15, Anthropic will replace unlimited programmatic usage in Claude Pro and Max subscriptions with a small, fixed monthly credit for SDK calls, including Claude Agent SDK, claude -p, and third-party tools like OpenClaw. Pro subscribers get only $20 in credits, while Max 20x users get $200, which matches their subscription cost. This change eliminates the cost advantage that heavy Claude Code users enjoyed by running high-volume agent workflows under a flat subscription fee, and effectively forces them toward pay-as-you-go API pricing. It also disrupts third-party tools like OpenClaw and Conductor that relied on shared subscription quotas. The credits cover Claude Agent SDK, claude -p, Claude Code GitHub Actions, and third-party apps built on the Agent SDK, but not interactive Claude Code in terminal/IDE or web/mobile chat. Pro: $20, Max 5x: $100, Max 20x: $200, Team standard: $20/seat, Team advanced: $100/seat. Excess usage will be billed at API rates.

rss · 宝玉(@dotey) · May 13, 20:47

Background: Claude subscriptions (Pro, Max, Team) offer users a monthly quota of interactive chat and Claude Code usage within rate limits. Programmatic usage refers to automated calls via APIs, SDKs, or command-line tools. Many developers used shared subscription credentials to run heavy automated workflows with third-party tools like OpenClaw and Conductor, effectively getting more value than the subscription cost. Anthropic's new policy separates programmatic and interactive usage, capping the former.

References

Tags: #Claude, #Anthropic, #Subscription, #API, #Pricing


LandingAI's Pre-Parse Page Classification API
LandingAI 推出解析前页面分类 API
⭐️ 8.0/10

LandingAI has launched ADE Classify, an API that performs page-level classification on PDFs before expensive document parsing, assigning labels and reasoning to each page. This approach reduces computational waste and extraction hallucinations by routing only relevant pages to downstream pipelines, solving a core inefficiency in enterprise document processing. The API accepts a custom JSON class list with semantic descriptions, evaluates pages concurrently, returns labels with reasoning strings, and marks unknown pages with a suggested class.

rss · meng shao(@shao__meng) · May 13, 14:31

Background: Enterprise document pipelines often receive mixed PDFs containing various document types, such as invoices, bank statements, and IDs. Traditional processing sends the entire document through OCR, parsing, and LLM extraction, wasting resources on irrelevant pages and causing errors when extraction agents encounter unexpected content.

References

Tags: #AI, #文档处理, #PDF, #API, #企业应用


Warning: Phishing Email Posing as X Security Alert
警告:冒充 X 官方安全提醒的钓鱼邮件
⭐️ 8.0/10

A user has reported receiving a phishing email that impersonates an official security alert from X (formerly Twitter), urging recipients not to click any links. This phishing attack could compromise many X accounts by tricking users into entering their credentials on a fake login page, leading to account takeover and potential data breaches. The email claims to be from X official security alert but is itself a phishing attempt. The warning emphasizes not to click any links in the email.

rss · meng shao(@shao__meng) · May 13, 10:30

Background: Phishing is a form of social engineering where attackers deceive individuals into revealing sensitive information, such as passwords, by impersonating a trusted entity. These attacks have become increasingly sophisticated, often mimicking official communications. This specific case targets X users, leveraging the platform's security alert format to appear legitimate.

References

Tags: #phishing, #security, #X, #Twitter, #social engineering


a16z: Incumbents Bet on Data Layer with Headless Agents
a16z:现有企业押注数据层,转向无头代理
⭐️ 8.0/10

a16z partner Seema Amble published an analysis arguing that as traditional system-of-record incumbents adopt headless agents, they are betting the data layer remains the primary source of value, while startups can compete on proprietary data, action layer ownership, and real-world execution. This insight highlights a fundamental strategic divide in AI agent architecture: incumbents protect their data moats, while startups must innovate on action and execution to differentiate. It shapes how the next generation of enterprise software will be built and who captures value. The analysis defines 'headless agent' as an AI agent without a user interface that operates on predefined business signals. It also notes that the next generation of systems of record will be agentic, capturing context, initiating work, and recording data exhaust.

rss · a16z(@a16z) · May 13, 22:09

Background: A system of record (SOR) is an authoritative data source for critical business information, like a CRM for customer data. A headless agent is an AI agent with no graphical interface; it receives inputs and triggers actions via APIs. Incumbents like Salesforce or SAP have massive data stores, while startups can build agents that execute real-world tasks.

References

Tags: #headless agents, #system of record, #data layer, #startups, #AI agents


Salesforce Headless Bet: Data Layer Defensibility
Salesforce 无头化:数据层成为防御优势
⭐️ 8.0/10

Salesforce announced it would open its APIs and launch a headless product, betting that its value lies in the data layer rather than the UI. a16z's Seema Amble published an analysis on where defensibility moves in the agentic era. This signals a major shift in software architecture: as AI agents interact via APIs, the UI becomes less critical, and data ownership becomes the primary moat. Businesses may need to rethink their product strategy to focus on data defensibility. Headless architecture decouples the frontend from the backend, exposing core services via APIs. Salesforce's move is part of a broader industry trend where companies like Shopify and Kontent.ai also offer headless solutions.

rss · a16z(@a16z) · May 13, 15:45

Background: Headless architecture separates the presentation layer from the backend, allowing content and data to be delivered across multiple channels via APIs. The agentic era refers to the rise of AI agents that autonomously perform tasks, often interacting with software through APIs rather than graphical interfaces. In this context, the value of traditional UI-heavy software diminishes, while the underlying data becomes the key competitive advantage.

References

Tags: #API, #Headless, #Agentic Era, #Salesforce, #Data Defensibility


Notion Launches Developer Platform with CLI, Workers, Agents SDK
Notion 发布开发者平台:CLI、Workers、Agent SDK
⭐️ 8.0/10

Notion introduced its Developer Platform, featuring a CLI tool called ntn, serverless Workers, database sync, webhooks, agent tools, an External Agents API, and the Notion Agents SDK. This marks a major expansion of Notion from a productivity tool into a full-fledged platform for building custom workflows and integrations, enabling developers and AI agents to extend Notion's functionality deeply. The platform includes the Notion CLI (ntn) for terminal-based interactions, Workers that run code on Notion's infrastructure, and the Notion Agents SDK for integrating Notion's agents into external applications. Database sync allows two-way data synchronization with any data source, while webhooks and the External Agents API enable event-driven automation.

rss · Notion(@NotionHQ) · May 13, 18:09

Background: Notion is a popular all-in-one workspace for notes, databases, and collaboration. Previously, its API was limited to basic CRUD operations. The new Developer Platform introduces serverless compute (Workers), a CLI, and advanced agent capabilities, positioning Notion as a competitive platform for workflow automation and AI integration.

References

Tags: #Notion, #Developer Platform, #API, #Automation, #Agent Tools


Sam Altman on AI Model Tradeoffs: Price/Speed vs Price/Intelligence
Sam Altman 谈 AI 模型权衡:价格/速度 与 价格/智能
⭐️ 8.0/10

Sam Altman tweeted about the anxiety of not using the smartest AI model, noting that sometimes he is okay with slow responses and suggested focusing on a price/speed tradeoff rather than a price/intelligence tradeoff. This insight from OpenAI's CEO highlights a growing industry debate about balancing cost, speed, and intelligence in LLM deployment, which could influence how developers choose models for real-world applications. The tweet received 1,614 replies and 4,186 likes, indicating strong community engagement. Altman specifically contrasts 'price/speed tradeoff' with 'price/intelligence tradeoff,' suggesting that for some use cases, speed may be more valuable than raw intelligence.

rss · Sam Altman(@sama) · May 13, 18:17

Background: Large language models (LLMs) like OpenAI's GPT series vary in cost, speed, and intelligence. Faster models are often cheaper but less capable, while smarter models are more expensive and slower. Developers must choose between these tradeoffs based on application needs.

References

Tags: #AI, #LLM, #cost efficiency, #model selection, #Sam Altman


LangChain Unveils SmithDB: Distributed DB for Agent Observability
LangChain 发布 SmithDB:专为智能体可观测性设计的分布式数据库
⭐️ 8.0/10

At the Interrupt! conference, LangChain announced SmithDB, a purpose-built distributed database for agent observability, claiming up to 12x faster performance with full portability. As AI agent traces grow exponentially, traditional databases struggle with scalability and query performance; SmithDB addresses this critical bottleneck, enabling faster debugging and monitoring of complex agent systems. SmithDB is part of the LangSmith platform and delivers up to 12x faster performance compared to existing solutions, with full data portability across environments.

rss · LangChain(@LangChainAI) · May 13, 20:22

Background: LangChain is a popular framework for building applications with large language models (LLMs). Agent observability refers to the ability to monitor and trace the inputs, outputs, and internal steps of AI agents, which is crucial for debugging and optimization. Traditional databases were not designed for the high-frequency, multi-step trace data generated by modern agents, creating a need for specialized storage.

References

Tags: #LangChain, #SmithDB, #Agent Observability, #Distributed Database


Inth Launches c15t: Cookie Banners That Don't Hurt Core Web Vitals
Inth 推出 c15t:不损害 Core Web Vitals 的 Cookie 横幅
⭐️ 8.0/10

Inth has launched c15t, an open-source, headless consent standard for cookie banners that is designed to not degrade Core Web Vitals. It offers native SDKs, internationalization (i18n), and full customization to match any website's UI. Cookie banners often harm web performance by causing layout shifts and delaying page load, negatively impacting Core Web Vitals scores. c15t provides a developer-friendly solution that maintains compliance without compromising user experience, making it valuable for performance-conscious websites. c15t is a headless consent engine that can be quickly integrated via npx @c15t/cli and is fully customizable to match a website's brand. It is open-source under the MIT license and available on GitHub for community contributions.

rss · Y Combinator(@ycombinator) · May 13, 18:00

Background: Core Web Vitals are Google's set of metrics that measure user experience, including Largest Contentful Paint (LCP), First Input Delay (FID), and Cumulative Layout Shift (CLS). Many third-party cookie banners hurt CLS by introducing asynchronous scripts and layout shifts. c15t addresses this by being lightweight and bundled directly into the website's code, minimizing performance impact.

References

Discussion: The community has responded positively, with developers on Vercel's forum calling c15t 'our cookie banner prayers answered.' The tweet from Y Combinator received significant engagement, indicating strong interest among web developers.

Tags: #web performance, #cookie banners, #open source, #Core Web Vitals, #launch


Paul Graham on Silicon Valley and Building Startup Hubs
Paul Graham 谈硅谷与创业中心建设
⭐️ 8.0/10

Paul Graham spoke at Y Combinator's Stockholm event on April 29, 2026, discussing why founders should move to Silicon Valley and how to build successful startup hubs elsewhere. These insights guide founders on the importance of geographic concentration for startup success and provide actionable advice for building vibrant startup ecosystems globally. The talk includes timestamps covering topics like serendipitous meetings, investor speed in the Valley, the Dropbox story, and Silicon Valley's pay-it-forward culture. Graham also explores whether Stockholm could become the Silicon Valley of Europe.

rss · Y Combinator(@ycombinator) · May 13, 14:32

Background: Silicon Valley is historically the premier startup hub due to its dense network of investors, talent, and a culture of paying it forward. Paul Graham is a co-founder of Y Combinator, a top startup accelerator, making his views influential in the startup world. The speech was part of YC's global outreach to foster startup ecosystems beyond the Valley.

Tags: #startups, #silicon valley, #y combinator, #startup hubs, #paul graham


Figure Live Streams 8-Hour Shift of F.03 Robots Sorting Packages
Figure 直播 F.03 人形机器人 8 小时分拣包裹
⭐️ 8.0/10

Figure AI is live streaming an 8-hour shift of its F.03 humanoid robots sorting packages, demonstrating practical, continuous automation in a logistics setting. This real-world demonstration highlights the progress of humanoid robots in performing repetitive tasks for extended periods, potentially transforming warehouse and logistics operations with reliable automation. The live stream shows multiple F.03 robots sorting packages over a full work shift, with no indication of failures or interventions. The F.03 is the third-generation humanoid robot from Figure, capable of general-purpose tasks controlled by the Helix AI model.

rss · The Rundown AI(@TheRundownAI) · May 13, 17:57

Background: Figure AI, founded in 2022, develops humanoid robots powered by AI. The F.03 is its latest generation, following the F.01 and F.02. The company also created the Helix vision-language-action model, which can command up to two robots simultaneously. This demo showcases progress toward commercial viability in logistics.

References

Tags: #Robotics, #Humanoid Robots, #Automation, #Logistics, #AI


Anthropic surpasses OpenAI in enterprise AI spending
Anthropic 在企业 AI 支出上超越 OpenAI
⭐️ 8.0/10

In April 2024, Anthropic overtook OpenAI in enterprise AI spending among U.S. businesses, capturing 34.4% market share compared to OpenAI's 32.3%, according to Ramp's AI Index. This marks a significant shift in the enterprise AI landscape, indicating that businesses are increasingly adopting Anthropic's Claude models over OpenAI's GPT models, and suggests that safety-focused AI offerings may be gaining traction. Over the past year, Anthropic's business adoption quadrupled, while OpenAI's grew by only 0.3%. The data is from Ramp, a corporate spending management platform, tracking paid AI subscriptions among U.S. businesses.

rss · The Rundown AI(@TheRundownAI) · May 13, 15:45

Background: Anthropic is an AI safety company founded by former OpenAI employees, known for its Claude series of large language models. The Ramp AI Index measures AI adoption among U.S. businesses by analyzing spending data from its platform, providing insights into which AI tools enterprises are paying for.

References

Tags: #AI, #Enterprise, #Market Share, #Anthropic, #OpenAI


Vercel AI Gateway Reveals Production AI Usage Trends
Vercel AI Gateway 揭示生产级 AI 使用趋势
⭐️ 8.0/10

Vercel's AI Gateway data shows Google leading in production AI scale, Anthropic dominating in coding and spending, OpenAI growing rapidly since GPT-4o release, and open-source models gaining ground. This provides rare visibility into real-world AI adoption across major providers, helping developers and businesses make strategic decisions about which models to use for production workloads. The AI Gateway, now in beta, offers a single endpoint to access models from multiple providers with improved uptime and no lock-in, but Vercel notes it is not yet ready for production use or migrating existing projects.

rss · Guillermo Rauch(@rauchg) · May 13, 21:15

Background: Vercel is a cloud platform for frontend developers, and its AI Gateway is a proxy layer that routes requests to various AI providers, enabling observability and analytics. The data shared by CEO Guillermo Rauch reflects aggregate usage patterns from developers using the gateway, offering a snapshot of the competitive landscape in AI model deployment.

References

Tags: #AI, #Vercel, #Production AI, #AI Agents, #Industry Trends


RubricEM: Meta-RL with Rubric-Guided Policy Decomposition
RubricEM:基于评分引导的元强化学习策略分解
⭐️ 8.0/10

RubricEM introduces a rubric-guided meta-reinforcement learning framework that decomposes agent policy into planning, research, review, and answer stages, enabling learning beyond verifiable rewards through stagewise policy decomposition and reflection-based meta-policy evolution. This approach addresses the limitation of verifiable rewards in long-horizon tasks like research report generation, where final answers cannot be easily graded, by providing dense credit assignment and structured guidance through rubrics. RubricEM decomposes the agent's policy into four stages: planning, research, review, and answer, each guided by a rubric. It also incorporates reflection-based meta-policy evolution to iteratively improve the policy across episodes.

rss · AK(@_akhaliq) · May 13, 12:53

Background: Meta-reinforcement learning aims to enable agents to learn to learn, adapting quickly to new tasks. Verifiable rewards work well for tasks with clear-cut answers (e.g., math problems) but fail for open-ended tasks like writing reports. Rubrics provide a structured set of criteria to evaluate partial progress and guide learning, making them useful for complex, long-horizon tasks.

References

Tags: #meta-RL, #reinforcement learning, #AI research, #policy decomposition, #machine learning


LangChain Launches LangSmith Engine for Automated Trace Analysis
LangChain 推出 LangSmith Engine 自动分析追踪
⭐️ 8.0/10

LangChain CEO Harrison Chase announced the launch of LangSmith Engine, an AI agent that sits on top of LangSmith traces, automatically identifies issues, and proactively suggests actionable code changes or evaluators to add. This enhances developer productivity by automating the debugging and evaluation of LLM applications, reducing manual effort in identifying root causes and improving reliability. It positions LangSmith as a more proactive observability tool in the competitive LLM monitoring space. LangSmith Engine runs continuously in the background and can propose specific code changes and suggest new evaluators based on trace analysis. It is accessible via smith.langchain.com and likely integrates with existing LangSmith workflows.

rss · Harrison Chase(@hwchase17) · May 13, 20:17

Background: LangSmith is a platform for building, testing, and monitoring LLM applications, offering features like tracing, evaluation, and debugging. Traces are end-to-end records of requests flowing through an AI system, composed of spans representing individual steps. Previously, developers had to manually review traces to find issues; LangSmith Engine automates this analysis.

References

Tags: #LangSmith, #AI, #LLM, #tooling, #monitoring


Enterprise excitement surges for OpenAI Codex adoption
企业对采用 OpenAI Codex 热情高涨
⭐️ 8.0/10

OpenAI president Greg Brockman reported great excitement from enterprises wanting to adopt Codex, citing that 2,000 developers reached out in just three hours after an OpenAI Developers post. This signals strong enterprise demand for AI-powered coding tools, validating Codex as a key platform for automating software development tasks and potentially accelerating enterprise AI adoption. OpenAI Developers noted that 2,000 developers reached out in 3 hours, and Brockman's tweet highlights enterprise interest specifically. Codex has grown to over 2 million weekly active users by March 2026.

rss · Greg Brockman(@gdb) · May 13, 23:47

Background: OpenAI Codex is an AI-powered coding agent that automates software engineering tasks such as feature building, refactors, and migrations. As of March 2026, it had over 2 million weekly active users and is positioned as an enterprise agent platform beyond coding.

References

Tags: #OpenAI, #Codex, #enterprise AI, #AI coding


Perplexity AI Announces Secure Computer with Hardware-Isolated Sandboxes
Perplexity AI 推出具有硬件隔离沙箱的安全计算机
⭐️ 8.0/10

Perplexity AI announced a new computer architecture where every task runs in a hardware-isolated sandbox with VPC-level storage and compute separation, and agents use short-lived proxy tokens instead of raw API keys for authentication. This approach significantly enhances security by preventing cross-task contamination and reducing the risk of API key leaks, which could set a new standard for secure execution environments in AI agent platforms. The sandboxing uses microVM technology (e.g., Firecracker) to provide hardware-level isolation, and VPC-level separation means each task's storage and compute are isolated at the virtual network level, similar to cloud VPCs.

rss · Perplexity(@perplexity_ai) · May 13, 17:05

Background: Traditional sandboxing relies on containers or processes, which share the host kernel and may have weaker isolation. Hardware-isolated sandboxes, like microVMs, boot a separate kernel for each task, providing stronger security. VPC-level separation is commonly used in cloud computing to create isolated network environments for different customers.

References

Tags: #security, #sandboxing, #VPC, #authentication, #Perplexity AI


Grafana Pyroscope 2.0: Continuous Profiling at Scale
Grafana Pyroscope 2.0:大规模持续性能分析
⭐️ 8.0/10

Grafana Labs launched Pyroscope 2.0, a rearchitected open-source continuous profiling database with single write paths, stateless query processing, and native OpenTelemetry Protocol (OTLP) support. This release makes continuous profiling practical at scale by reducing storage costs and query latency, aligning with the industry trend toward OpenTelemetry-native observability. Key architectural improvements include a unified write path for all profile types, stateless query processing that reduces operational complexity, and full OTLP ingestion, enabling seamless integration with existing OpenTelemetry pipelines.

rss · InfoQ · May 13, 08:00

Background: Continuous profiling is a method that captures performance data (e.g., CPU usage, memory allocations) from applications in real-time, unlike traditional ad-hoc profiling. OpenTelemetry Protocol (OTLP) is a vendor-agnostic standard for transmitting telemetry data (traces, metrics, logs) between components. Pyroscope 2.0 builds on these concepts to provide a scalable, cost-effective profiling solution.

References

Tags: #continuous profiling, #observability, #open source, #Grafana, #Pyroscope


Reflecting on the End of Finetuning in AI
反思 AI 中微调时代的终结
⭐️ 8.0/10

A reflective article from Latent Space considers whether finetuning is becoming obsolete as a dominant paradigm in AI. This discussion challenges the conventional reliance on finetuning and may influence how researchers and practitioners approach model adaptation in the future. The article uses a quiet day to reflect on the trajectory of finetuning, though specific technical details are not provided in the summary.

rss · Latent.Space · May 13, 02:47

Background: Finetuning is a common technique in deep learning where a pre-trained model is further trained on a specific task to improve performance. Recent trends like prompt engineering and retrieval-augmented generation have offered alternatives, leading some to question finetuning's long-term dominance.

Tags: #AI, #finetuning, #machine learning, #deep learning


Anthropic and SpaceX Partner to Boost Claude AI Compute Capacity
Anthropic 与 SpaceX 合作提升 Claude AI 算力
⭐️ 8.0/10

Anthropic has signed an agreement to rent the entire computing capacity of SpaceX's Colossus 1 data center, gaining access to over 220,000 Nvidia GPUs and more than 300 megawatts of power. In response, Anthropic has doubled the 5-hour rate limits for Claude Code across paid plans and removed peak-hour restrictions for Pro and Max users, while also significantly increasing API rate limits for Claude Opus. This partnership dramatically expands Anthropic's compute capacity, enabling faster scaling of Claude AI services and potentially accelerating development of advanced AI models. It also highlights the growing strategic importance of massive GPU clusters controlled by entities like SpaceX. Colossus 1 is an AI data center in Memphis, Tennessee, built by SpaceX to support Elon Musk's ventures including X and SpaceX itself. The deal gives Anthropic exclusive access to the entire facility, which includes over 220,000 Nvidia GPUs and over 300 MW of power capacity.

telegram · zaihuapd · May 14, 00:57

Background: Anthropic is the developer of the Claude family of AI models, competing with OpenAI and Google. Claude Code is a developer tool that acts as an AI programmer in the terminal, understanding entire codebases and assisting with coding, modification, and refactoring. The partnership with SpaceX provides Anthropic with a massive increase in compute resources needed to train and run large language models.

References

Tags: #Anthropic, #Claude, #算力合作, #AI, #SpaceX