論文・研究の記事
MasterControl Seventeen Every Time
arXiv:2609.03209v1 Announce Type: new Abstract: We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic policy selects and runs a pre-approved analytical progr…
Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory
arXiv:2609.03340v1 Announce Type: new Abstract: Distributed LLM-agent teams can read the latest shared facts and still act on an obsolete plan. A planner may derive an action from requirement $r3$, another agent may com…
Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence
arXiv:2609.02981v1 Announce Type: new Abstract: Artificial intelligence is changing the form of applied English materials from fixed paper sequences to adaptive learning systems that can diagnose learners, recommend tas…
Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
arXiv:2609.03460v1 Announce Type: new Abstract: As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust f…
DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents
arXiv:2609.03423v1 Announce Type: new Abstract: Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely te…
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
arXiv:2609.03438v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions d…
A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant
arXiv:2609.03402v1 Announce Type: new Abstract: Artificial intelligence (AI) teaching assistants powered by large language models (LLMs) offer scalable educational support but often provide limited personalization. This…
Speculative Macro Commit for Faster Tool-Using Agents
arXiv:2609.03236v1 Announce Type: new Abstract: Tool-using LLM agents spend wall-clock time not only on model inference but also in serial action--observation turns, where each tool call, environment transition, and obs…
DUDE:Double Detection Multi-Agent System for Paper-Code Discrepancy Detection デュアル検出マルチエージェントシステム
arXiv:2609.03416v1 発表の種類: 新しい 抽象: LLM パワフルな紙コード不一致の検出は、研究提出の拡大が手動レビューの能力を超えているので、懸念が高まっています。
ストーリーに捕らわれた:多回転LLMの会話における話しの捕らえ
arXiv:2609.03407v1 Announce Type: new 抽象: 人々はますます日常的なアドバイスのために大規模な言語モデル(LLMs)に転向し、倫理的に負担された人际問題を実践的な道徳的アドバイス文脈にします。
Switchyard: NVIDIA’s Open Source Routing Library
Stop sending every AI request to your most expensive model. See how intelligent routing can cut cost and latency without sacrificing much quality.
5 Free LLM API Providers You Can Use in 2026
Explore five free AI API providers for accessing large language models, fast inference, multimodal AI, and agentic applications without paying for API usage.
R$^{2}$Adapter: A Routing and Rewriting Adapter for Efficient Hybrid RAG
arXiv:2609.02894v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become a prevailing paradigm for enhancing Large Language Models (LLMs) with non-parametric knowledge. Vanilla RAG efficiently han…
Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents
arXiv:2609.02889v1 Announce Type: new Abstract: A growing body of work improves frozen large language models (LLMs) as agents by evolving their harness: the textual scaffolding around the model, including persona, strat…
Distilled Rapid Embedding Transfer (DRET): Parameter-Efficient Biomedical Domain Adaptation via Priority-Based Embedding Transfer
arXiv:2609.02898v1 Announce Type: new Abstract: Large domain-specific language models such as BioBERT and ClinicalBERT achieve strong performance on biomedical NLP tasks, but their computational demands make them imprac…
Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding
arXiv:2609.02897v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet…
Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent
arXiv:2609.02890v1 Announce Type: new Abstract: A personalized language agent must convert a user's interaction history into behavior on each new request at inference time. Two strategies dominate. Retrieval pulls a few…
BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events
arXiv:2609.02895v1 Announce Type: new Abstract: Large-scale public events, such as religious festivals, political rallies, and cultural gatherings, are increasingly vulnerable to the rapid dissemination of misinformatio…
Probe Generalization as Subspace Selection for OOD Deception Detection
arXiv:2609.02893v1 Announce Type: new Abstract: Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the g…
汚染はスコアを膨らませますが、ほとんど大きな言語モデルリーダーボードを再リストラすることはありません。
arXiv:2609.02899v1 Announce Type: new 抽象: Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. テストアイテムのトレーニングデータへの漏洩は、広く大きな言語モデル(LLM)リーダーボードの信頼性に脅威として説明されています。
Counterexamples as Feedback for Agent Self-Correction(エージェント自己修正のためのフィードバック)
arXiv:2609.02892v1 Announce Type: new 抽象: Single-turn code-generation metrics understates a central property of deployed agents: whether they can repair a wrong artifact after receiving concrete feedback. この論文は、エージェントが間違ったアーティファクトを修理できるかどうかを示しています。
PiPMRE: A Pipeline Based on Language Model for Medical Relation Extraction
arXiv:2609.02896v1 Announce Type: new Abstract: Medical relation extraction (MRE) is commonly known for extracting entities and their relations jointly from a medical text, which has attracted considerable attention in…
OpenAI ボットがサイバーセキュリティのテストから逃れた時、本当に何が起こったのか?
何百人ものOpenAIエージェントがテスト環境から抜け出し、AIプラットフォーム「Hugging Face」に侵入したとき、ヘッジ・クラフ氏はこれが間違った話であり、失敗だと警告した。
ESPO:Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize(エラー構造プロンプト最適化、診断、多様化、安定化)
GEPAのような進化的なプロンプト最適化は、プロンプトブレイクに苦しみます:それぞれのイーテレーションにはルールと警告が付属し、プロンプトが最大で3$\times$長くなりますが、それ以上正確ではありません。