トップの記事
4 engineering patterns behind the strongest AI Agents Challenge submissions
The recent Google for Startups AI Agents Challenge revealed that the most successful multi-agent systems rely on foundational software engineering patterns rather than just raw model power. Winning architectures consist…
PineMail
Catch dev emails. Let your AI agent read them. / Pine Mail is a tiny, self-hosted SMTP mail catcher for development environments — like Mailtrap or Mailpit, but written in Rust and built with AI-agent-driven testing in…
Learning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA
arXiv:2609.01687v1 Announce Type: new Abstract: Grounded question answering systems should answer only when the supplied evidence supports the answer. In multi-hop QA, this requirement is difficult because partial evide…
ChatGPT Work Mode High Error Rates
Status: Resolved All impacted services have now fully recovered. Affected components Image Generation (Operational) GPTs (Operational) ChatGPT Work (Operational) Voice mode (Operational) Agent (Operational) Codex in Cha…
Significant-Gravitas/AutoGPT
AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
OpenMusicPrompt.com
All-in-one AI Studio: Analyze, Generate & Perfect AI Music / Stop guessing prompts. OpenMusicPrompt is your all-in-one AI music workspace. Unlock the full potential of AI music with: 🎵 Analyze: Extract the "soul" & sty…
VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages
arXiv:2609.01788v1 Announce Type: new Abstract: Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing p…
METR Report on OpenAI / Hugging Face Hacking Incident
score=77, comments=64
huggingface/transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Watermarks Remover
Clean unwanted marks from images in your browser / Watermarks Remover is a simple browser tool for cleaning unwanted marks from images and preparing visuals for reuse. Try quick edits at
Induction and Inquiry via Probabilistic Reasoning over Language and Code
arXiv:2609.01815v1 Announce Type: new Abstract: How humans grow and maintain abstract knowledge from the sparse, streaming noisy data of experience is a longstanding challenge in cognitive science. Any computational acc…
Impersonating IT support: how threat actors turn a remote session into enterprise-wide access
Microsoft Threat Intelligence observed a human-operated intrusion campaign that abuses Microsoft Teams external collaboration to impersonate IT support, gain remote access, and deploy a Node.js-based implant. Learn how…
HKUDS/nanobot
Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
Originkit
Animated component library for building modern websites / OriginKit is a growing library of production-ready animated components, sections, and interactions for building modern websites faster. Customize everything visu…
When Does Information Sharing Improve Decentralized Discovery? Aggregation, Independent Rescue, and Equilibrium Selection
arXiv:2609.01814v1 Announce Type: new Abstract: Information sharing can improve a pooled estimate while eliminating independent rescue actions. This paper separates those effects in exact finite discovery models. A cent…
Elevated errors for Claude Sonnet 5
Sep 2 , 21:44 UTC Resolved - The issue affecting Claude Sonnet 5 has been resolved. Impact occurred from 2:05pm PT / 21:05 UTC to 2:19pm PT / 21:19 UTC. Sep 2 , 21:17 UTC Investigating - We are investigating elevated er…
affaan-m/ECC
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
January AI
Stripe for healthcare / Clinical-grade infrastructure designed for consumer-grade experiences. January AI puts metabolic & nutrition intelligence into your health platform. Available via APIs or SDKs. Sign up, create an…
AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking
arXiv:2609.01828v1 Announce Type: new Abstract: Spoken dialogue state tracking recovers slot-value pairs from speech, where ASR errors concentrate in entity values and persist across turns, making it both a generation a…
OpenLeash Adds a Human Check to Risky AI Agent Actions
The security tool intercepts potentially dangerous agent actions, blocking clear threats and requesting human approval when intent is uncertain. The post OpenLeash Adds a Human Check to Risky AI Agent Actions appeared f…
langgenius/dify
Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
Buoylog
AI-drafted changelogs from your GitHub PRs / Buoylog turns the GitHub PRs you merge into AI-categorized release notes, publishing them to a hosted changelog page with an in-app widget and email digests.
PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation
arXiv:2609.01658v1 Announce Type: new Abstract: Retrieval-Augmented Generation enhances Large Language Models by grounding responses in external knowledge, but multi-hop reasoning remains vulnerable to error propagation…
Elevated errors creating new accounts
Status: Resolved The issue affecting new account creation has been resolved.