記事本文へスキップ

論文・研究の記事

論文・研究 / 最新表示設定
論文・研究arxiv.org

MasterControl Seventeen Every Time

arXiv:2609.03209v1 Announce Type: new Abstract: We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic policy selects and runs a pre-approved analytical progr…

論文・研究arxiv.org

Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

arXiv:2609.03340v1 Announce Type: new Abstract: Distributed LLM-agent teams can read the latest shared facts and still act on an obsolete plan. A planner may derive an action from requirement $r3$, another agent may com…

論文・研究arxiv.org

Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence

arXiv:2609.02981v1 Announce Type: new Abstract: Artificial intelligence is changing the form of applied English materials from fixed paper sequences to adaptive learning systems that can diagnose learners, recommend tas…

論文・研究arxiv.org

Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty

arXiv:2609.03460v1 Announce Type: new Abstract: As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust f…

論文・研究arxiv.org

DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

arXiv:2609.03423v1 Announce Type: new Abstract: Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely te…

論文・研究arxiv.org

Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

arXiv:2609.03438v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions d…

論文・研究arxiv.org

A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant

arXiv:2609.03402v1 Announce Type: new Abstract: Artificial intelligence (AI) teaching assistants powered by large language models (LLMs) offer scalable educational support but often provide limited personalization. This…

論文・研究arxiv.org

Speculative Macro Commit for Faster Tool-Using Agents

arXiv:2609.03236v1 Announce Type: new Abstract: Tool-using LLM agents spend wall-clock time not only on model inference but also in serial action--observation turns, where each tool call, environment transition, and obs…

論文・研究arxiv.org

DUDE:Double Detection Multi-Agent System for Paper-Code Discrepancy Detection デュアル検出マルチエージェントシステム

arXiv:2609.03416v1 発表の種類: 新しい 抽象: LLM パワフルな紙コード不一致の検出は、研究提出の拡大が手動レビューの能力を超えているので、懸念が高まっています。

論文・研究arxiv.org

ストーリーに捕らわれた:多回転LLMの会話における話しの捕らえ

arXiv:2609.03407v1 Announce Type: new 抽象: 人々はますます日常的なアドバイスのために大規模な言語モデル(LLMs)に転向し、倫理的に負担された人际問題を実践的な道徳的アドバイス文脈にします。

論文・研究kdnuggets.com

Switchyard: NVIDIA’s Open Source Routing Library

Stop sending every AI request to your most expensive model. See how intelligent routing can cut cost and latency without sacrificing much quality.

論文・研究kdnuggets.com

5 Free LLM API Providers You Can Use in 2026

Explore five free AI API providers for accessing large language models, fast inference, multimodal AI, and agentic applications without paying for API usage.

論文・研究arxiv.org

R$^{2}$Adapter: A Routing and Rewriting Adapter for Efficient Hybrid RAG

arXiv:2609.02894v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become a prevailing paradigm for enhancing Large Language Models (LLMs) with non-parametric knowledge. Vanilla RAG efficiently han…

論文・研究arxiv.org

Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents

arXiv:2609.02889v1 Announce Type: new Abstract: A growing body of work improves frozen large language models (LLMs) as agents by evolving their harness: the textual scaffolding around the model, including persona, strat…

論文・研究arxiv.org

Distilled Rapid Embedding Transfer (DRET): Parameter-Efficient Biomedical Domain Adaptation via Priority-Based Embedding Transfer

arXiv:2609.02898v1 Announce Type: new Abstract: Large domain-specific language models such as BioBERT and ClinicalBERT achieve strong performance on biomedical NLP tasks, but their computational demands make them imprac…

論文・研究arxiv.org

Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding

arXiv:2609.02897v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by drafting candidate tokens and verifying them in parallel. Tree-attention drafters such as EAGLE-3 are widely adopted, yet…

論文・研究arxiv.org

Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent

arXiv:2609.02890v1 Announce Type: new Abstract: A personalized language agent must convert a user's interaction history into behavior on each new request at inference time. Two strategies dominate. Retrieval pulls a few…

論文・研究arxiv.org

BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events

arXiv:2609.02895v1 Announce Type: new Abstract: Large-scale public events, such as religious festivals, political rallies, and cultural gatherings, are increasingly vulnerable to the rapid dissemination of misinformatio…

論文・研究arxiv.org

Probe Generalization as Subspace Selection for OOD Deception Detection

arXiv:2609.02893v1 Announce Type: new Abstract: Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the g…

論文・研究arxiv.org

汚染はスコアを膨らませますが、ほとんど大きな言語モデルリーダーボードを再リストラすることはありません。

arXiv:2609.02899v1 Announce Type: new 抽象: Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. テストアイテムのトレーニングデータへの漏洩は、広く大きな言語モデル(LLM)リーダーボードの信頼性に脅威として説明されています。

論文・研究arxiv.org

Counterexamples as Feedback for Agent Self-Correction(エージェント自己修正のためのフィードバック)

arXiv:2609.02892v1 Announce Type: new 抽象: Single-turn code-generation metrics understates a central property of deployed agents: whether they can repair a wrong artifact after receiving concrete feedback. この論文は、エージェントが間違ったアーティファクトを修理できるかどうかを示しています。

論文・研究arxiv.org

PiPMRE: A Pipeline Based on Language Model for Medical Relation Extraction

arXiv:2609.02896v1 Announce Type: new Abstract: Medical relation extraction (MRE) is commonly known for extracting entities and their relations jointly from a medical text, which has attracted considerable attention in…

論文・研究Openai

OpenAI ボットがサイバーセキュリティのテストから逃れた時、本当に何が起こったのか?

何百人ものOpenAIエージェントがテスト環境から抜け出し、AIプラットフォーム「Hugging Face」に侵入したとき、ヘッジ・クラフ氏はこれが間違った話であり、失敗だと警告した。

論文・研究arxiv

ESPO:Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize(エラー構造プロンプト最適化、診断、多様化、安定化)

GEPAのような進化的なプロンプト最適化は、プロンプトブレイクに苦しみます:それぞれのイーテレーションにはルールと警告が付属し、プロンプトが最大で3$\times$長くなりますが、それ以上正確ではありません。

次の記事を読む