<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0">
  <channel>
    <title>arXiv papers on language models and agents</title>
    <link>https://arxiv.org/</link>
    <description>Events collected by UnlimitedPipe 0.3.2</description>
    <generator>UnlimitedPipe 0.3.2</generator>
    <lastBuildDate>Fri, 25 Sep 2026 06:46:03 +0000</lastBuildDate>
    <item>
      <title>Reward Hacking Challenges Oversight of Autonomous Research Agents</title>
      <link>https://arxiv.org/abs/2609.28614</link>
      <guid isPermaLink="false">5ba3068a6e5c84d3dce3</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Autonomous research agents can design experiments, evaluate results, and write reports, giving them control over both a scientific result and the evidence used to support it. This creates a risk of reward hacking: meeting the reward criteria without achieving the intended goal. We study (1) how often models reward-hack without instructions to do so, (2) how effective and detectable their methods are when hacking is allowed, and (3) how they adapt when an LLM review panel returns its decision…</description>
      <category>cs.CL</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks</title>
      <link>https://arxiv.org/abs/2609.28673</link>
      <guid isPermaLink="false">c8e64331ebc52ce4c95e</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on character attacks (ad hominem arguments), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content. Specifically, we investigate whether modern LLMs can replicate human competence to…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs</title>
      <link>https://arxiv.org/abs/2609.28727</link>
      <guid isPermaLink="false">0b2479d49d9389fd09a8</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Contextual biasing improves rare-word recognition in speech large language models (SpeechLLMs), but efficiently exploiting large bias lists remains challenging. We propose PTC-Bias, a two-stage framework based on phoneme-level temporal competition. At the prefill stage, PTC Retrieval performs frame-synchronous phoneme decoding and temporal competition among candidate pronunciations, producing a compact bias-word shortlist and corresponding speech intervals. After SpeechLLM decoding, PTC…</description>
      <category>cs.CL</category>
      <category>cs.SD</category>
    </item>
    <item>
      <title>Script Choice in LLMs: Evidence for Late-Layer Commitment</title>
      <link>https://arxiv.org/abs/2609.28784</link>
      <guid isPermaLink="false">ca43a263026f6f26064f</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>In this paper, we investigate how script knowledge is distributed across the layers of LLMs using two complementary interpretability methods: logistic regression probing and logit-lens analysis. Our probing experiments reveal a clear asymmetry: both the input script and the instructed output script are encoded in the earliest layers of the network, while, in contrast, commitment to the actual output script emerges only in the final layers, with the model's intermediate representations…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms</title>
      <link>https://arxiv.org/abs/2609.29001</link>
      <guid isPermaLink="false">4ad02298d0f467091f4a</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two English-language datasets with complementary annotation formats: continuous human ratings and three-way categorical labels. Across the seven evaluated models, we find that inter-model agreement is stronger than model--human agreement. Strategy-level analyses suggest that…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>ELF-REG: Scaling Continuous Diffusion Language Models to Reasoning Tasks</title>
      <link>https://arxiv.org/abs/2609.29102</link>
      <guid isPermaLink="false">8b8ed3a5384d008e0167</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Fully continuous diffusion language models (dLMs) denoise continuous representations without intermediate discretization, then decode all response tokens in parallel at the final step. Their performance on challenging reasoning tasks remains less established than that of autoregressive (AR) LLMs and masked dLMs. We scale Embedded Language Flows (ELF) to mathematical reasoning and code generation on GSM8K, MATH-500, HumanEval, and MBPP. We introduce ELF-REG, which improves learning with…</description>
      <category>cs.CL</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Reasoning Instructions Can Break Answer Decoding in Vision--Language Models</title>
      <link>https://arxiv.org/abs/2609.29278</link>
      <guid isPermaLink="false">1452f25f141d74268c2d</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Chain-of-thought (CoT) instructions can distort multiple-choice VLM evaluation when a scorer appends a reasoning cue but reads answer-label logits before the model generates any rationale. We call this CoT-prefix scoring. On ScienceQA, Qwen2.5-VL-7B drops from 80.76% to 45.48%, and across five option-content permutations 93.54% of CoT-prefix predictions select the first slot. Condition-matched linear probes recover 78.94% from the same hidden states, while free generation restores 75.24%…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Grammatical "grandmother neurons" are rare in LLMs</title>
      <link>https://arxiv.org/abs/2609.29328</link>
      <guid isPermaLink="false">77386e4efae589b979b3</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Understanding how Large Language Models (LLMs) encode linguistic structures remains a fundamental challenge in interpretability research. While diagnostic classifiers (or "probes") are widely used for this task, they face significant methodological criticism: training auxiliary classifiers introduces capacity confounds and calibration issues, often making it difficult to distinguish the model's intrinsic representations from the probe's ability to learn the task. To address these limitations…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams</title>
      <link>https://arxiv.org/abs/2609.29333</link>
      <guid isPermaLink="false">9ebd98781f9c0eb27167</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ($570$ dual-graded students) under $171$ configurations spanning closed and open-weights models; the best reaches mean absolute error $1.64/35$, below the $2.61/35$ two human graders achieve against each other. The catch is the prompt: a short ''strict grader'' preamble drives $14$ of $17$…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
      <category>cs.CY</category>
    </item>
    <item>
      <title>ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts</title>
      <link>https://arxiv.org/abs/2609.29349</link>
      <guid isPermaLink="false">8eef57a645e2447201fe</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for Arabic LLM safety evaluation. In total, 58 teams registered, 35 participated in the final evaluation, and 27 submitted system-description papers. Participating teams explored models such as AraBERT, Jais, and Qwen3-VL. The best systems achieved macro-F1 scores of 0.823 on…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring</title>
      <link>https://arxiv.org/abs/2609.29370</link>
      <guid isPermaLink="false">f74f50a99cd276c69d2e</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are costly and challenging to scale across countries. Large language models (LLMs) enable new possibilities for extracting and structuring information from long and unstructured policy documents. This paper presents an application of LLMs as "AI respondents" for generating…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Likelihood Ranking doesn't Scale Like Prompting in LLMs</title>
      <link>https://arxiv.org/abs/2609.29390</link>
      <guid isPermaLink="false">0a04a6b42332a9da84c2</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters</title>
      <link>https://arxiv.org/abs/2609.29397</link>
      <guid isPermaLink="false">a43be67b70b5d4940b56</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Ternary (1.58-bit) weights are attractive for microcontroller-class language models, but the sub-1M-parameter regime rests mainly on isolated, single-seed comparisons. One prominent example reports that a routed ternary block (convolution, diagonal SSM and sparse attention mixed by a per-token router) beats a parameter-matched full-precision transformer by 22% at 60K parameters, attributing this to inductive bias. We re-run it under one fixed recipe, three seeds per cell, 98 byte-level runs on…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?</title>
      <link>https://arxiv.org/abs/2609.29410</link>
      <guid isPermaLink="false">f1abf2f931608d63d86e</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between these two capabilities. This work focuses on determining how much the LLM deviates from a buggy solution to fix the bug compared to a human-written patch, and if there is a bias…</description>
      <category>cs.CL</category>
      <category>cs.SE</category>
    </item>
    <item>
      <title>Rufus-Air: An Open LLM Post-Training Recipe</title>
      <link>https://arxiv.org/abs/2609.29421</link>
      <guid isPermaLink="false">49a5a44cee8f7738b3a8</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>agentic-ger: terminology recovery in long-form speech using global context</title>
      <link>https://arxiv.org/abs/2609.29428</link>
      <guid isPermaLink="false">a7e5ce953f00c842d0ff</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the world knowledge and contextual capability of large language models (LLMs), we propose Agentic-GER, an LLM-based agent for terminology correction in long-form speech. The agent uses global context from the full transcript to identify suspicious terms and resolve ambiguous…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis</title>
      <link>https://arxiv.org/abs/2609.29444</link>
      <guid isPermaLink="false">6b90b68c4822ebf8a104</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories introduce noise and obscure useful information. To address these issues, we propose IterSynth, a role-decoupled and summary-based paradigm that alternates between a Planner for…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions</title>
      <link>https://arxiv.org/abs/2609.29496</link>
      <guid isPermaLink="false">8d83e5431bfaeaedd046</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Natural language explanation generation serves as a key mechanism for exposing and evaluating vision-language reasoning. Prior work on explanation-driven vision-language models predominantly follows a post-hoc (answer-first) paradigm, implicitly suggesting that supervised rationales can reflect underlying reasoning processes. In contrast, modern large vision-language models increasingly exhibit a rationale-first generation tendency, which more closely aligns with structured, stepwise reasoning…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>PROOF: Profiling Reliability of Object-Level Facts in Large Language Models</title>
      <link>https://arxiv.org/abs/2609.29504</link>
      <guid isPermaLink="false">c59d1ec5e953c5dff38e</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for factual coverage in instruction-tuned language models. PROOF converts a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts, 101 classes, 392 properties, and 14 domains. Each question has an explicit "I don't know"…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Benchmarking Arabic--Russian Machine Translation: A Comparison of Fine-tuned NMT and Few-shot LLMs under Rich Morphology and Low Lexical Overlap</title>
      <link>https://arxiv.org/abs/2609.29559</link>
      <guid isPermaLink="false">039f41838c4a16b04716</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Arabic-Russian machine translation (MT) remains under-explored due to the rich morphology of Arabic and low lexical overlap between the two languages. We benchmark seven fine-tuned neural machine translation (NMT) models against four few-shot large language models (LLMs) on a 20k/5k/5k split of a new 15.47M-pair corpus. Fine-tuned NLLB-1.3B achieves the highest BLEU (16.3) and COMET (0.738). Aya-Expanse 8B leads the few-shot LLMs (BLEU 1.7 on 500 sentences, chrF 25.7), but all LLM scores remain…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>An Exploratory Ablation of a Small MLA--SSM Hybrid Language Model</title>
      <link>https://arxiv.org/abs/2609.29618</link>
      <guid isPermaLink="false">3c5ca3c69e59e8558360</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>We report an exploratory, single-seed ablation of TALH (Adaptive Latent Hybrid), a decoder-only language model with parallel Multi-head Latent Attention (MLA) and a custom recurrent state-space (SSM) branch. Five variants, spanning 117--217M estimated active parameters per token, are trained from scratch on a FineWeb sample for the same number of optimisation steps and tokens. In this specific setup, removing the SSM branch gives the largest degradation in validation perplexity (MLA-only PPL…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
      <category>cs.IR</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>LLMersion: A Local-First AI Agent Framework for Low-Cost Home Language Learning toward Educational Equity</title>
      <link>https://arxiv.org/abs/2609.29672</link>
      <guid isPermaLink="false">f3bb459d7879467cff04</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Artificial intelligence helps education most where an essential provision has been rationed by cost. For language learners that provision is a teacher's voice, which binds listening, reading, speaking, and writing into one act. Published evidence shows why most learners lack it, from a global shortage of 44 million teachers to heavy household tutoring bills, and why technology has not substituted for it: computer-assisted language learning proved effective but narrow, applications presuppose…</description>
      <category>cs.CL</category>
      <category>cs.CY</category>
      <category>cs.HC</category>
    </item>
    <item>
      <title>Stochastic Semantic Evidence Graphs: Uncertainty Propagation and Governance for Agentic AI</title>
      <link>https://arxiv.org/abs/2609.29703</link>
      <guid isPermaLink="false">a2a8548059bdb6510e38</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>AI-agent evaluations usually inspect a final answer, yet error may enter through evidence, retrieval, prompting, generation or decision mapping. We introduce a stochastic semantic evidence graph (SSEG), a hierarchical stochastic DAG whose language node expands into an autoregressive token subgraph and whose observable output may be a law over complete phrases. Semantic reduction and calibration are optional. We define graph-relative local defects and downstream edge influences, derive a…</description>
      <category>cs.CL</category>
      <category>cs.IR</category>
      <category>stat.ML</category>
    </item>
    <item>
      <title>PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides</title>
      <link>https://arxiv.org/abs/2609.29718</link>
      <guid isPermaLink="false">46120b8fa58c689e66ff</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Coding agents are beginning to act in the visual world. They now build webpages, GUIs, games, 3D scenes, diagrams, and documents. Success in such visual coding requires bridging two spaces: inferring visual structure and expressing it programmatically. Slides are a core medium of knowledge work, widely used to communicate ideas and collaborate in a form that people can directly inspect and edit. Therefore, they provide an ideal testbed for visual coding, as they require agents to recover visual…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places</title>
      <link>https://arxiv.org/abs/2609.29769</link>
      <guid isPermaLink="false">fd2b7832b2b04b725157</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>We ask whether Jev, a typed classifier that returns probabilities over permitted answers without generating text, can replace an LLM rubric judge. We compare it with three flash-tier LLM judges on nine panels drawn from seven benchmarks, giving every judge identical criterion texts. Jev's accuracy differs significantly from an LLM judge's in only 8 of 27 paired comparisons, ahead mostly on binary criteria and behind only on graded ones, and most of the other comparisons are inconclusive. Summed…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>CORDIAL: Calibrating Ordinal LLM Outputs from Few Labels</title>
      <link>https://arxiv.org/abs/2609.29807</link>
      <guid isPermaLink="false">7b911805692d5be73e72</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>A large language model (LLM) can turn a text into a distribution over an ordered scale, but that distribution is a noisy measurement: saturated, compressed or exaggerated, and biased in a consistent direction. We propose CORDIAL, which treats the model's output as a noisy reading of the true label and corrects it with a channel of five interpretable parameters. The channel is small enough for its posterior to be averaged from a handful of labels, and we prove that the resulting calibration…</description>
      <category>cs.CL</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>ChunkRank: Model-Aware Text Chunking and Abstention-Aware Answer Selection for LLM Pipelines</title>
      <link>https://arxiv.org/abs/2609.29828</link>
      <guid isPermaLink="false">4857148ac6e853b7d4cf</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>We present ChunkRank, an open-source Python library that derives chunk boundaries from a target model's tokenizer and context window, and selects an answer among candidates produced independently per chunk. It ships a validated registry of 90 models across 15 providers and six answer-selection methods, and needs only three core dependencies. For chunking, ChunkRank avoids context-window overflow automatically from the model name, whereas character-based splitters overflow or waste the budget…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs</title>
      <link>https://arxiv.org/abs/2609.29845</link>
      <guid isPermaLink="false">efa02528e5b940595f5b</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the \textit{Superposition Linearity Hypothesis}. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax</title>
      <link>https://arxiv.org/abs/2609.29848</link>
      <guid isPermaLink="false">4c0fbcf27602c23b977b</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>A language model can fail a syntactic test in two distinct ways: by not encoding the relevant structure, or by encoding it but failing to use it at the output. Behavioral evaluation alone cannot tell these apart. We propose a three-level evaluation framework (behavioral deployment, LM-head readout, and probe recoverability) measured on the same items under the same binary decision. Using a compact trilingual (English, Chinese, German) control-dependency benchmark, we find that probe…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations</title>
      <link>https://arxiv.org/abs/2609.29928</link>
      <guid isPermaLink="false">177007d0e40e83fd5923</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language models (LLMs) are increasingly used as synthetic survey respondents to estimate population response distributions. In cross-cultural survey simulation, evaluations should assess not only distributional fidelity within countries but also whether differences across countries are preserved. However, existing distance-based metrics such as Jensen--Shannon divergence (JSD) do not directly capture such cross-country differences. To address this limitation, we introduce Cultural…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Scoring Both Directions: LLMs realize the MRS they cannot reliably parse</title>
      <link>https://arxiv.org/abs/2609.30071</link>
      <guid isPermaLink="false">b1512737a3703c76754b</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>The English Resource Grammar (ERG) is a hand-written computational grammar of English. Given a sentence, its processor, ACE, produces a formal meaning representation called Minimal Recursion Semantics (MRS): a graph of the sentence's predicates and their arguments. The grammar is bidirectional and can also turn an MRS back into an English sentence. \citet{hajdik2019} used the ERG's treebank to build a benchmark for that generation task, MRS to text, and trained sequence-to-sequence models to…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure</title>
      <link>https://arxiv.org/abs/2609.30074</link>
      <guid isPermaLink="false">f0f71e10d20209e741d2</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open model variants across five families and 8B to 675B parameters, caching disabled, 293 raw intermediate representations persisted. The measured phenomenon is unstable to begin with. Identical calls do not reliably recover identical structure, with mean node-set Jaccard…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Do Audio Language Models Hear and Read Distinctive Features Alike?</title>
      <link>https://arxiv.org/abs/2609.30167</link>
      <guid isPermaLink="false">914967a9a92d11c4a85a</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Audio language models pass speech and text through a single decoder. We ask whether that decoder represents a distinctive feature in the same direction when a phoneme is heard and when it is read. For minimal pairs of phonemes differing in one feature, we take the offset between the two members' mean representations. Averaging those offsets gives a direction for each stream, and we measure the cosine between the two. Because the two streams already agree about arbitrary phoneme pairs, we…</description>
      <category>cs.CL</category>
      <category>cs.LG</category>
      <category>cs.SD</category>
      <category>stat.AP</category>
    </item>
    <item>
      <title>Agentic Detection of Online Conspiracies</title>
      <link>https://arxiv.org/abs/2609.30250</link>
      <guid isPermaLink="false">6f3e455bdfdda4c4be9c</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Conspiratorial discourse on social media is not always expressed through explicit claims or stable lexical markers. The same surface content may express endorsement, legitimate concerns, criticism, satire, or mockery. The main challenge is therefore not only recognizing conspiracy-related claims, but inferring the speaker's intent -- the utterance's illocutionary force. We argue that this can be achieved through the use of relevant social contexts and propose an agentic framework, equipped with…</description>
      <category>cs.CL</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability Routing</title>
      <link>https://arxiv.org/abs/2609.28475</link>
      <guid isPermaLink="false">78b6c8e9f9beefe93a6e</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Forecasting agents increasingly combine language-model reasoning, retrieval, ensembling, and calibration, but it remains unclear when each behavior should be trusted. We study this question on ForecastBench-style binary forecasting tasks, treating the choice to retrieve, reason, defer to a market prior, or use a historical analog as an observable agent behavior rather than a hidden implementation detail. Our central finding is that mechanism choice is source-dependent: structured analogs…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Spooftral: Can Voxtral Audio-Language Model Detect Speech Spoofing?</title>
      <link>https://arxiv.org/abs/2609.28713</link>
      <guid isPermaLink="false">455d978ee0fa7b5adb2d</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Self-supervised learning (SSL) countermeasures (CMs) have shown strong performance in recent years. However, they often show degraded performance while facing unseen spoofing attacks and mismatched conditions. This study examines the Voxtral audio-language model (ALM) framework for spoofing detection, as a step toward combining CM capabilities within the ALM framework. We analyze how Voxtral captures spoofing cues through audio-text processing and propose an instruction-guided approach that…</description>
      <category>eess.AS</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Reward-Tilted On-Policy Distillation for Acoustic Grounding in Audio-Language Models</title>
      <link>https://arxiv.org/abs/2609.28778</link>
      <guid isPermaLink="false">ca4b2cad767340d73302</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Audio-language models (ALMs) can exploit textual shortcuts to answer questions while overlooking acoustic evidence, weakening audio understanding. On-policy distillation (OPD) trains compact ALMs by supervising student-generated responses with teacher predictions, but does not explicitly distinguish acoustic support from linguistic predictability. We propose Reward-Tilted On-Policy Distillation (RT-OPD) to strengthen acoustic grounding. Given the same question and student-generated text, a…</description>
      <category>cs.SD</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks</title>
      <link>https://arxiv.org/abs/2609.29015</link>
      <guid isPermaLink="false">ac1f9dd8c964f81de1d2</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Decentralized LLM-based multi-agent systems coordinate through local interactions, but an agent can remain responsive while its task-solving quality persistently degrades. Such gray failures require protecting current tasks before sufficient evidence exists to alter future routing, while still allowing recovered agents to rejoin. We introduce MeshHeal, a fully decentralized self-healing framework that couples ability-matched peer review across two timescales. At the fast timescale, an adaptive…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.DC</category>
    </item>
    <item>
      <title>Design and Evaluation of LLM Chaining-Based Task Planning for General Purpose Service Robots</title>
      <link>https://arxiv.org/abs/2609.29043</link>
      <guid isPermaLink="false">ce8f82801806b16330a7</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>General Purpose Service Robot (GPSR) tasks, as defined in the RoboCup@Home benchmark, require robots to interpret diverse natural language commands and generate multi-step action sequences in real home environments. Conventional Single Prompt (SP) approaches suffer from context bloat and the "Lost in the Middle" phenomenon, leading to unreliable task planning. We propose an LLM chaining architecture that separates instruction classification and action generation into two specialized stages…</description>
      <category>cs.RO</category>
      <category>cs.AI</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench</title>
      <link>https://arxiv.org/abs/2609.29251</link>
      <guid isPermaLink="false">34af42abe44f43bcd18a</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>CAR-bench evaluates whether tool-using agents stay reliable under real-world uncertainty, executing every tool inside the evaluator so that each tool-result exchange is a separate agent round-trip. A conventional next-action agent can batch parallel tool calls, but a chain of dependent calls costs it one model call per round of results. We present a coroutine-bridge harness in which the model's only action is to emit a Python program that blocks and resumes in place across evaluator tool…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy</title>
      <link>https://arxiv.org/abs/2609.29508</link>
      <guid isPermaLink="false">69145c3f1e3298870340</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language model (LLM) agents may perform well on isolated tasks yet drift into inconsistency over extended interaction. We evaluate temporal consistency in a controlled 20-step multi-agent setting inspired by delayed-gratification studies. At each step, an agent chooses between continuing to delay a reward or claiming it immediately (terminating the episode). Across a full-factorial manipulation of social visibility (private vs public), persona stressors, and deliberation policy, we run…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.LG</category>
      <category>cs.MA</category>
    </item>
    <item>
      <title>Delay-of-Gratification as a Multi-Agent Survival Micro-benchmark for Long-Horizon LLMs: Social Exposure, Personas, and Tool Use Budgets</title>
      <link>https://arxiv.org/abs/2609.29509</link>
      <guid isPermaLink="false">b0e321e8d05366eb88b5</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language models (LLMs) are increasingly deployed as multi-turn agents that must sustain goals, use tools, and adapt to other agents over extended interactions. However, existing research lacks auditable, multi-turn, multi-factorial experiments that quantify LLM behavior under explicit constraints, with time-resolved statistics that reveal how behavior unfolds over long horizons. To address this gap, we develop a multi-agent micro-benchmark inspired by the Stanford marshmallow experiment…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>A Corpus of Real Scam- and Spam-Call Conversations from an Active Voice-Agent Honeypot</title>
      <link>https://arxiv.org/abs/2609.29528</link>
      <guid isPermaLink="false">9f1dfbe7b1df77353b47</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Real conversations between fraudsters and their targets are among the most informative artifacts for studying telephone scams, yet also the scarcest: passive honeypots overwhelmingly capture automated messages and hang-ups, large-scale studies characterize call metadata rather than dialogue, and manual scam-baiting does not scale. We present a dataset of real scam-call conversations collected by an active voice-agent honeypot. Dedicated numbers are seeded into the lead-generation channels fraud…</description>
      <category>cs.CR</category>
      <category>cs.CL</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation</title>
      <link>https://arxiv.org/abs/2609.29578</link>
      <guid isPermaLink="false">76a8a760949b92593109</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Long-horizon tool agents often make useful progress without reaching terminal success, motivating partial-credit evaluation. Yet evaluators may reward milestones that were temporary, later reversed, or not attributable to the evaluated agent. Comparing an honest trajectory with a higher-scoring adversarial one is inconclusive if the latter made more genuine progress. We introduce PartHackBench, a controlled methodology that removes this confound. A private certifier admits a pair only when its…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.CR</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Free the Language Model From the Vision Encoder: Semantic Serialization as a Perception Interface for Small Language Models</title>
      <link>https://arxiv.org/abs/2609.29601</link>
      <guid isPermaLink="false">1456471e9935d8f40976</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>End-to-end vision-language models (VLMs) bind visual competence to the scale of their language model: as the language model shrinks, perception and reasoning degrade together. We study an embodied scene question-answering (QA) interface in which vision never enters the language model. A frozen perception stack detects and ranges objects; a deterministic semantic serializer compiles the perceived state, errors included, into decision-aligned text; an unmodified text-only large language model…</description>
      <category>cs.RO</category>
      <category>cs.CL</category>
      <category>cs.CV</category>
    </item>
    <item>
      <title>STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models</title>
      <link>https://arxiv.org/abs/2609.29607</link>
      <guid isPermaLink="false">46ea36f0f0a99d1ecaa4</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure in spatio-temporal monitoring, the ability to persistently track object identities, states, and relations over time. Existing benchmarks obscure this deficit by relying on single final-answer evaluations for queries that can often be resolved via local visual cues or statistical priors. To rigorously diagnose this, we…</description>
      <category>cs.CV</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Three Ways Classical Test Theory Misleads for LLM Judges</title>
      <link>https://arxiv.org/abs/2609.29709</link>
      <guid isPermaLink="false">f105f86c3fa791ea3d2a</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>An LLM judge scores a bank of responses against a rubric, and the reliability comes back at $0.52$. What has been measured? Judge evaluation has begun borrowing reliability statistics from classical test theory, usually without stating the measurement design each statistic assumes, and we show that three widely portable ones mean something different for a judge than for a test because the judge setting rearranges the roles those designs rest on. First, an internal-consistency coefficient…</description>
      <category>cs.LG</category>
      <category>cs.CL</category>
      <category>stat.ME</category>
    </item>
    <item>
      <title>PUBG Ally: A Conversational Embodied Agent as an AI Teammate</title>
      <link>https://arxiv.org/abs/2609.29837</link>
      <guid isPermaLink="false">9c629f22ba8e8a6a9a04</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to a constantly changing game world under strict latency constraints while interacting naturally with players, keeping its speech synchronized with its actions. Ally therefore combines agentic tool use with real-time game control. A…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.HC</category>
    </item>
    <item>
      <title>Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models</title>
      <link>https://arxiv.org/abs/2609.30048</link>
      <guid isPermaLink="false">a4bb8d098dd419d47665</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude. We test this zero-shot on current commercial models. Five LLMs generate solutions to MBPP, HumanEval, and DS-1000, seven more to MBPP, and models act as evaluators in four tasks: picking their own solution from a pair, judging whether a single solution is their own, identifying which of two solutions a named model wrote, and judging quality blind…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.SE</category>
    </item>
    <item>
      <title>PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations</title>
      <link>https://arxiv.org/abs/2609.30094</link>
      <guid isPermaLink="false">28120cb7f87f1304a627</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce \textbf{PrivDrift}, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.CR</category>
    </item>
    <item>
      <title>Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale</title>
      <link>https://arxiv.org/abs/2609.30137</link>
      <guid isPermaLink="false">f10314b2de29c7cb552b</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust.
We present a hypothesis-driven simulation…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>GRASP: Generating, Revising, and Assessing for Strategic Planning with Agentic AI</title>
      <link>https://arxiv.org/abs/2609.30147</link>
      <guid isPermaLink="false">7302157a5dce14e3a07f</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large Language Models (LLMs) typically exhibit a performance profile where reliability degrades as task complexity increases. We address the challenge of generating high-quality natural language executable plans for complex tasks by introducing $\textbf{GRASP}$, a strategy-aware, multi-stage planning framework. GRASP decouples the planning pipeline across specialized, context-isolated modules: it pre-compiles global macro-guidelines (GenPlan), explores alternative localized strategies within…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.LG</category>
      <category>cs.MA</category>
    </item>
    <item>
      <title>Foundations of Large Language Models</title>
      <link>https://arxiv.org/abs/2501.09223</link>
      <guid isPermaLink="false">b41def700bd1a487a8a8</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>This is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is structured into six main chapters, each exploring a key area: pre-training, generative models, prompting, alignment, inference, and reasoning. It is intended for college students, professionals, and practitioners in natural language processing and related fields, and can serve as a reference for…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning</title>
      <link>https://arxiv.org/abs/2512.04457</link>
      <guid isPermaLink="false">ec14319eb02607dab0cb</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Machine unlearning for large language models (LLMs) remains challenging because full retraining is costly, while approximate methods often struggle to remove targeted behaviors without degrading retained utility, especially under limited post-deployment supervision. We consider a practical PEFT setting for targeted behavioral contamination removal with a small forget set, a limited retain buffer, and LoRA-only updates, and propose RapidUn, an influence-guided framework that converts…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents</title>
      <link>https://arxiv.org/abs/2601.06676</link>
      <guid isPermaLink="false">affa1f27f867e36711f3</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large Language Model (LLM)-based deep research agents perform multi-step reasoning, web exploration, and long-form report generation. In these long-horizon workflows, early deviations from user intent can misdirect research and propagate through planning, search, and synthesis, making timely interaction essential. However, existing benchmarks primarily treat deep research as a static input-output task, overlooking agents' ability to elicit and use user feedback. We introduce IDRBench, a…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
      <category>cs.HC</category>
    </item>
    <item>
      <title>LLM surprisal is necessary but not sufficient to capture English garden-path effects: Evidence from joint latent modeling of reading paradigms</title>
      <link>https://arxiv.org/abs/2602.04489</link>
      <guid isPermaLink="false">ba4024cb8a1e6a3180aa</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Temporarily ambiguous garden-path sentences ("While the team trained the striker wondered... ") are known to cause processing difficulty, which can manifest itself in a variety of reading behaviors (in-situ slowdowns, rereading), as well as in miscomprehension or outright rejection of the sentence as ungrammatical. Which types of reading behavior are observed critically depends on the experimental method used to collect the data, which makes comparing results between reading paradigms…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches</title>
      <link>https://arxiv.org/abs/2604.01754</link>
      <guid isPermaLink="false">753aaa89fea55cd767f4</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Mathematical reasoning is a hallmark of human intelligence, and whether large language models (LLMs) can meaningfully perform it remains a central question in artificial intelligence and cognitive science. As LLMs are increasingly integrated into scientific workflows, rigorous evaluation of their mathematical capabilities becomes a practical necessity. Existing benchmarks are limited by synthetic settings and data contamination. We present LiveMathematicianBench, a dynamic multiple-choice…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis</title>
      <link>https://arxiv.org/abs/2604.14121</link>
      <guid isPermaLink="false">98cf66036c1dded61007</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language models (LLMs) have become increasingly used for various tasks, often coupled with Chain-of-Thought (CoT) prompting to boost accuracy. Recent work has shown that high label-prediction accuracy does not guarantee correct intermediate reasoning, and the causes of *reasoning flaws* vary from sample to sample, yet existing remedies either focus on a single domain or assume that one flaw type applies uniformly across samples. A simple mitigation method is to provide the model with the…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>An Empirical Study of Automating Agent Evaluation</title>
      <link>https://arxiv.org/abs/2605.11378</link>
      <guid isPermaLink="false">0ef2ee5ff3ab7840a4b3</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Agent evaluation requires assessing complex multi-step behaviors involving tool use and intermediate reasoning, making it costly and expertise-intensive. A natural question arises: can frontier coding assistants reliably automate this evaluation process? Our study shows that simply prompting coding assistants is insufficient for this task. Without domain-specific evaluation knowledge, frontier coding assistants achieve only a 30% execution success rate and produce over-engineered evaluations…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models</title>
      <link>https://arxiv.org/abs/2607.24312</link>
      <guid isPermaLink="false">ab8a109f9ac7b2439e84</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Document-level relation extraction (DocRE) aims to extract relations among multiple entities across extended contexts while maintaining consistency across predicted triples. Although large language models (LLMs) show remarkable reasoning capabilities in information extraction, their predictions are typically generated independently for each candidate triple and may violate fundamental relational constraints such as transitivity, symmetry, and functional uniqueness, leading to contradictory and…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Gaokerena: A Small Persian Medical Language Model Family</title>
      <link>https://arxiv.org/abs/2608.00932</link>
      <guid isPermaLink="false">ca23960825ea345fe1f8</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused on English, leaving low-resource languages like Persian significantly underserved. To address this gap, this paper introduces Gaokerena, a novel family of compact Persian medical language models optimized for deployment on consumer-grade hardware. As a foundational step toward localized digital healthcare, we first present Gaokerena-V…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>LLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders</title>
      <link>https://arxiv.org/abs/2609.07746</link>
      <guid isPermaLink="false">1f262f412538b44d539e</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Even though backdoors in LLMs have been a growing concern, their inner workings are still under heavy scrutiny. Trigger-based backdoors are easy to define behaviorally, a rare input that makes the model switch to a chosen response pattern, but the mechanism between triggers and their responses is less clear. We study this mechanism in a controlled, harmless language-switching setting, where fixed trigger sequences make 1B and 8B language models continue English prompts in French or German. For…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>How Many Humans Are 32 LLM Judges Worth?</title>
      <link>https://arxiv.org/abs/2609.21277</link>
      <guid isPermaLink="false">c40bc00326c80cecd7ce</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>A panel's human-equivalent size is target-specific. Matching a fixed 32-judge panel to empirical human label distributions on three ChaosNLI tasks yields two distinct effective sizes: distributional-error matching gives $\nu_{\mathrm{MSE}}=2.304$, $3.750$, and $3.445$, whereas spectral matching gives $\nu_H=4.242$, $6.459$, and $6.499$, a gap of $1.72$--$1.89\times$; a binary-error diagnostic credits the same panels with only $1.971$--$2.227$ effective votes. Extrapolating the…</description>
      <category>cs.CL</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage</title>
      <link>https://arxiv.org/abs/2609.22904</link>
      <guid isPermaLink="false">1d52fb8592a08684f320</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Triage in the emergency department (ED) is a sequential decision process that unfolds turn by turn. Existing evaluations of large language models (LLMs) for triage use completed retrospective records and report performance close to that of physicians. We implement a methodology for evaluating LLMs on sequential triage, the task of predicting a triage acuity label from a growing prefix of a nurse-patient conversation. We evaluate six LLMs at five sequential checkpoints on two corpora: 425…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Conduct Under Pressure: What Sixty Language Models Do When a User Pushes</title>
      <link>https://arxiv.org/abs/2609.25447</link>
      <guid isPermaLink="false">74c7c8e13242750759f8</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>We study what LLMs do when a user applies pressure in an uncomfortable situation: a user insists, begs, flatters or grieves, and the model gives up a correct fact, writes a document it should refuse, or cheers a plan that will cost the user money. We send frozen multi-turn scenes, identical for every model regardless of the reply, to 60 models from 13 vendors, and label each transcript with a codebook built by open coding and then frozen: a trajectory (the model held its position or folded) and…</description>
      <category>cs.CL</category>
      <category>cs.HC</category>
    </item>
    <item>
      <title>LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models</title>
      <link>https://arxiv.org/abs/2609.27220</link>
      <guid isPermaLink="false">5ab5d90b623c8e4823b9</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. We identify a recurring reasoning failure, stable-but-wrong lock-in, where an answer stabilizes early around an incorrect value while substantial denoising remains. Surface-level decoding signals such as confidence, entropy, margin, and answer stability are insufficient to reliably distinguish correct from erroneous lock-in. We formulate selective reasoning…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving</title>
      <link>https://arxiv.org/abs/2609.27717</link>
      <guid isPermaLink="false">a1b4d57f6144df1988e9</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Human-written agent skills encode rich workflows for real-world problem solving, but are typically used as external inference-time instructions rather than internalized as reusable model capabilities. We introduce \texttt{SkillGym}, a framework that transforms these skills into executable, verifiable training environments for large language model agents. Its skill-to-task pipeline instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Generating Interesting Scientific Ideas using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders</title>
      <link>https://arxiv.org/abs/2405.17044</link>
      <guid isPermaLink="false">239388c3a01355743961</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>The rapid growth of scientific literature makes it increasingly challenging for researchers to identify novel and impactful ideas, especially across disciplines. Modern artificial intelligence (AI) systems offer new opportunities for scientific ideation, but how compelling are AI-generated ideas, and how can their quality be improved? Here, we introduce SciMuse, which generates personalized research ideas using a knowledge graph of 58 million papers and a large language model (LLM). A central…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.DL</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Unraveling the cognitive patterns of Large Language Models through module communities</title>
      <link>https://arxiv.org/abs/2508.18192</link>
      <guid isPermaLink="false">b0418f0497c7ab50474d</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large Language Models (LLMs) have reshaped our world with significant advancements in science, engineering, and society through applications ranging from scientific discoveries and medical diagnostics to Chatbots. Despite their ubiquity and utility, the underlying mechanisms of LLM remain concealed within billions of parameters and complex structures, making their inner architecture and cognitive processes challenging to comprehend. We address this gap by adopting approaches to understanding…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs</title>
      <link>https://arxiv.org/abs/2512.06607</link>
      <guid isPermaLink="false">81aa14645dfe23dca68f</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Applying LLMs to predictive tasks in finance is challenging due to look-ahead bias resulting from their training on long time-series data. This precludes the backtests typically employed in finance since retraining frontier models from scratch with a specific knowledge cutoff is prohibitive. In this paper, we introduce a fast, effective, and low-cost alternative. Our method guides generation at inference time by adjusting the logits of a large base model using a pair of smaller, specialized…</description>
      <category>cs.LG</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>SPARQL-LLM: Real-Time SPARQL Query Generation from Natural Language Questions</title>
      <link>https://arxiv.org/abs/2512.14277</link>
      <guid isPermaLink="false">58ca30d91e46c3d7ef23</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>The advent of large language models is contributing to the emergence of novel approaches that promise to better tackle the challenge of generating structured queries, such as SPARQL queries, from natural language. However, these new approaches mostly focus on response accuracy while ignoring other evaluation criteria, such as runtime and cost to generate SPARQL queries. Consequently, they are often not production-ready or easy to deploy over real-world knowledge graphs with good accuracy. To…</description>
      <category>cs.IR</category>
      <category>cs.AI</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>LOGIC: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration</title>
      <link>https://arxiv.org/abs/2601.15397</link>
      <guid isPermaLink="false">0db06adbba04d53a6731</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Recognizing entity phrases remains a critical challenge for speech large language models. Existing prompting methods lack an explicit decoding-time biasing weight, limiting their controllability. Generative error correction methods can introduce hallucinated over-corrections. To address these limitations, we propose LOGIC (logit-space integration for contextual biasing), a robust framework operating directly in the logit space. By decoupling context injection from input processing, LOGIC…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.SD</category>
    </item>
    <item>
      <title>IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models</title>
      <link>https://arxiv.org/abs/2604.07709</link>
      <guid isPermaLink="false">9ef6e9aa2b745e447e10</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>A strongly safety-trained model will provide a doctor with a benzodiazepine taper schedule, but not a patient who asks for one. The model knows the information, but how much it shares depends on the framing. We introduce IatroBench, a benchmark that evaluates models on two axes of harm (commission and omission) across 60 pre-registered clinical scenarios and 6 models. We use Claude Opus 4.6 to score model responses against a rubric written by a physician, and find that its omission scores are…</description>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.CY</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models</title>
      <link>https://arxiv.org/abs/2606.23057</link>
      <guid isPermaLink="false">f3cac331f4fab21ee1e8</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>This exploratory study measures brand inclusion across five industries, 50 brands and 250 queries, each put five times to GPT-5.2, Gemini 3 Flash and Perplexity sonar-pro in February and September 2026 (3,614 and 3,750 scored answers). Category Inclusion Rate, Recommendation Share, Competitive Vacuum Index and Co-Mention Asymmetry have stated denominators. February inclusion rates sit close together across an industry's sampled brands (mean Gini 0.30), while at least one brand is named in 80%…</description>
      <category>cs.IR</category>
      <category>cs.CL</category>
      <category>cs.CY</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning</title>
      <link>https://arxiv.org/abs/2608.04452</link>
      <guid isPermaLink="false">612fd64dac5e4577ebaf</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Multimodal large language models (MLLMs) can miss fine details in a full image that they recognize in a closer view. Recovering this evidence requires deciding where to look and how much surrounding context to retain. We present Q-CueGraph, a query-conditioned evidence acquisition method for frozen MLLMs. For text-rich images, it builds a reusable graph of OCR lines and layout relations. Each question activates anchors, expands them into contextual regions, and selects candidates for a single…</description>
      <category>cs.CV</category>
      <category>cs.AI</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings</title>
      <link>https://arxiv.org/abs/2608.17556</link>
      <guid isPermaLink="false">63995fcee2c0bcbb2837</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large Language Models (LLMs) in real-world applications often face the risks of specially crafted prompts designed to bypass the safety controls. Existing guardrail methods, such as LLM-as-a-judge and cloud-based safety APIs are able to detect unsafe content. However, they often add a delay of about 250-900 ms to each request. This delay is too high for real-time applications, when the system usually needs to respond in less than 100 ms. Furthermore, routing user prompts through external…</description>
      <category>cs.CR</category>
      <category>cs.CL</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Measuring Brand and Source Discovery under Repeated LLM Queries: A Finite-Sample Audit</title>
      <link>https://arxiv.org/abs/2609.05059</link>
      <guid isPermaLink="false">8fe84f3928ead8654e39</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Repeated-query audits must distinguish recovery of a collected set from completeness of possible outputs. We apply sample-based rarefaction to 4,500 responses from 50 buying questions, six configurations and 15 calls per cell. Historical-dictionary median ten-call recovery of the observed 15-call set ranges from 92.6% to 95.2%; re-adjudicating all 45,683 candidate strings changes this range to 89.5%-94.7%. Two blinded Gemini 3.1 Pro annotation roles assessed 600 complete answers, yielding micro…</description>
      <category>cs.IR</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>RRSI: Regularized Recursive Self-Improvement of Agent Harnesses</title>
      <link>https://arxiv.org/abs/2609.24972</link>
      <guid isPermaLink="false">690a24c5afc6cf93e08a</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large…</description>
      <category>cs.LG</category>
      <category>cs.AI</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction</title>
      <link>https://arxiv.org/abs/2609.25176</link>
      <guid isPermaLink="false">029a818b17f654fe935a</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multimodality and Multi-Teacher On-Policy Distillation (M$^{2}$-OPD) to transfer language capabilities and develop native audio skills. Act uses self-evolving executable environments and multi-granularity rollouts for Group…</description>
      <category>eess.AS</category>
      <category>cs.AI</category>
      <category>cs.CL</category>
      <category>cs.SD</category>
    </item>
    <item>
      <title>ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning</title>
      <link>https://arxiv.org/abs/2609.27532</link>
      <guid isPermaLink="false">6df7ee50b1af779aa425</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The standard recipe assigns a single outcome reward at the end and compares trajectories sampled for the same task. As a result, a group with no successful trajectory yields no training signal, failed attempts cannot be told apart by how close they came to completion, and turns that advance the task receive the same credit as turns that only query…</description>
      <category>cs.LG</category>
      <category>cs.CL</category>
    </item>
    <item>
      <title>PAWS: Policy-driven Agentic World Simulation</title>
      <link>https://arxiv.org/abs/2609.28547</link>
      <guid isPermaLink="false">ee9a4249ad55e4844f15</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Policy interventions propagate through public communication, institutional decisions, and stakeholder responses, yet datasets for financial multi-agent simulation rarely connect these processes to temporally aligned historical evidence. We introduce PAWS, a Policy-driven Agentic World Simulation dataset covering 36 verified U.S. financial and economic policy episodes, 12,727 policy-linked news records, and 65,291 source-grounded stakeholder actions. Each action is linked to its supporting news…</description>
      <category>cs.AI</category>
      <category>cs.CE</category>
      <category>cs.MA</category>
      <category>cs.SI</category>
    </item>
    <item>
      <title>BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines</title>
      <link>https://arxiv.org/abs/2609.28557</link>
      <guid isPermaLink="false">33e333e27535eaa409c9</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>DNA sequencing pipelines, spanning quality control, alignment, variant calling, and annotation, are now reliably executed by workflow management systems that orchestrate established bioinformatics tools at scale. What remains manual is the decision layer surrounding that execution: selecting quality thresholds appropriate to a sample and platform, adjudicating borderline variant calls, diagnosing anomalies, and determining which findings warrant expert review. These decisions are repetitive…</description>
      <category>cs.AI</category>
      <category>q-bio.GN</category>
    </item>
    <item>
      <title>Adversarial Closed-Loop Curriculum for Evolving Role-Playing Agents</title>
      <link>https://arxiv.org/abs/2609.28609</link>
      <guid isPermaLink="false">e1b0798b5c7b65d2dd27</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Role-playing agents based on large language models have been widely applied in areas such as personalized assistance and social simulation. Recent RL methods typically train on a fixed scenario pool collected before learning begins. This creates a distributional bottleneck: as the agent improves, the scenarios where it performs poorly also change, while the training distribution remains static. Therefore, we propose AdvRole, an adversarial context rewriting framework that turns role-playing RL…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Driving Epidemic Models with AI Agents: the Epydemix Agent Framework</title>
      <link>https://arxiv.org/abs/2609.28692</link>
      <guid isPermaLink="false">b8fa40aa3e0103159f15</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Artificial Intelligence agents based on large language models provide convenient natural language interfaces to scientific software, but reliability is not automatic. Here we introduce the Epydemix Agent Framework, an additive layer over Epydemix, an open-source Python library for stochastic compartmental epidemic modeling. The framework extends the library with four capabilities to facilitate interaction with an AI agent: discovery of available models and parameters, preventive validation of a…</description>
      <category>cs.AI</category>
      <category>cs.CY</category>
    </item>
    <item>
      <title>Progressive Skill Discovery as Access Control for Tool-Using LLM Agents: Structural Governance through Role-Scoped Capability Delivery</title>
      <link>https://arxiv.org/abs/2609.28693</link>
      <guid isPermaLink="false">52f0c0c9f51d27e0db19</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large Language Model (LLM) agents struggle to scale safely when exposed to vast enterprise toolsets. Providing an agent with access to every internal tool leads to oversized context windows, degraded tool selection, and severe governance vulnerabilities - as system policies defined purely in prompts remain probabilistic advice rather than hard constraints. Existing mitigations, such as multi-agent domain delegation, decentralize audit logs and fail to guarantee policy compliance across…</description>
      <category>cs.AI</category>
      <category>cs.CR</category>
      <category>cs.MA</category>
      <category>cs.SY</category>
      <category>eess.SY</category>
    </item>
    <item>
      <title>Reinforcement Learning with Verifiable Rewards for Small Search Agents</title>
      <link>https://arxiv.org/abs/2609.28765</link>
      <guid isPermaLink="false">4fff0c4945245ff14f96</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is less clear remains open. The reason-over-search recipe applies RLVR to open-domain question answering, where retrieval grounds the answer and a match against the reference supplies the reward. So far it has been demonstrated on large models, and below one billion parameters only with distillation from a larger teacher…</description>
      <category>cs.AI</category>
      <category>cs.IR</category>
    </item>
    <item>
      <title>Agent Memory with Episodic Retrieval for Financial Decision-Making</title>
      <link>https://arxiv.org/abs/2609.28771</link>
      <guid isPermaLink="false">ea47f4abfc90f7ab64c3</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language models (LLMs) have demonstrated strong capabilities in financial analysis and reasoning, inspiring recent advances in agent-based trading frameworks. While these systems show promise, prior approaches either emphasize long-horizon forecasting or operate as stateless analyzers, limiting their applicability to the demands of trading in complicated settings. To address these gaps, we introduce META (Memory Enhanced Trading Agent), the first RAG-like episodic-memory-augmented…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?</title>
      <link>https://arxiv.org/abs/2609.28850</link>
      <guid isPermaLink="false">5fc07719bbcc31aa86b2</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget. An agent must reproduce that result using the paper and whatever its authors released. What the…</description>
      <category>cs.AI</category>
      <category>cs.LG</category>
      <category>cs.SE</category>
    </item>
    <item>
      <title>Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents</title>
      <link>https://arxiv.org/abs/2609.28876</link>
      <guid isPermaLink="false">c4ac4cd81b22c5835b92</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by…</description>
      <category>cs.AI</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise</title>
      <link>https://arxiv.org/abs/2609.28919</link>
      <guid isPermaLink="false">76397b27e0c6bc46bf9c</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Harnesses, the products that run AI coding agents, are multiplying, and enterprises are rolling them out to their employees: what started as pilots with a few hundred seats is scaling to tens of thousands. Most enterprises do not build these harnesses but buy them from large vendors, such as Anthropic's Claude Code or OpenAI's Codex. A harness decides which model answers, what the model reads, how the prompt cache is used and which subagents run, so it picks the rate on the price sheet and sets…</description>
      <category>cs.AI</category>
      <category>cs.CR</category>
    </item>
    <item>
      <title>PFArena: Benchmarking Language Models for Protein Modification</title>
      <link>https://arxiv.org/abs/2609.28921</link>
      <guid isPermaLink="false">5fd23587eb61d0e0c741</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Protein modification requires navigating an immense sequence space, yet wet-lab validation remains low-throughput and costly. Although computational paradigms including protein language models (PLMs), large language models (LLMs), and LLM-based agents have shown promise in protein modification, their relative efficacy across realistic experimental decision-making settings remains unclear. To bridge this gap, we introduce PFArena, a benchmark comprising four controlled task interfaces that cover…</description>
      <category>cs.AI</category>
      <category>q-bio.BM</category>
    </item>
    <item>
      <title>From Static Personal Values to Contextualized Personalization: Bayesian Personalized Value Alignment for LLMs</title>
      <link>https://arxiv.org/abs/2609.28942</link>
      <guid isPermaLink="false">b8d6516f4730d6839060</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Personalized value alignment has become increasingly important as large language models (LLMs) are expected to accommodate diverse user preferences. However, existing methods typically align model outputs with a static value profile across prompts, overlooking that the salience of value dimensions varies substantially across contexts. Inspired by Lewin's Field Theory, which views human behavior as jointly shaped by personal dispositions and situational constraints, we model personal values as…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning</title>
      <link>https://arxiv.org/abs/2609.28963</link>
      <guid isPermaLink="false">8ab9d97235f55583eb69</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since coarse-grained trajectory-level advantages are hard to accurately reflect the contribution of individual steps (i.e, failed trajectories may contain valuable steps). Revisiting the…</description>
      <category>cs.AI</category>
      <category>stat.ML</category>
    </item>
    <item>
      <title>AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor Mining</title>
      <link>https://arxiv.org/abs/2609.29014</link>
      <guid isPermaLink="false">12c08abfbb9751de8e83</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language model (LLM)-based multi-agent systems can automate alpha factor mining, but their reliance on external APIs limits control over cost, availability, and confidentiality. Long research loops also tend to revisit a few successful economic mechanisms that lead to research path collapse. To address these limitations, we propose AlphaDiverse, a framework that integrates a multi-agent alpha research system, diverse research path collection, and post-training for local agents. We let the…</description>
      <category>cs.AI</category>
      <category>cs.CE</category>
      <category>cs.MA</category>
    </item>
    <item>
      <title>From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents</title>
      <link>https://arxiv.org/abs/2609.29051</link>
      <guid isPermaLink="false">946a1f5959ada0aa6675</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>On-policy self-distillation (OPSD) has become a popular recipe for post-training LLM agents. It supervises the agent model at the token level with a stronger teacher view of the same model, obtained by conditioning on privileged information (PI). In this work, we show that in multi-turn agents, this paradigm teaches the student to act with confidence but without the information behind it. The trained agent behaves as if it had privileged information it never observed, and its performance falls…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment</title>
      <link>https://arxiv.org/abs/2609.29109</link>
      <guid isPermaLink="false">e2fb368e13184c55e569</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Reasoning-capable language models often produce long chains of thought when direct answers suffice, wasting inference compute. Many dual-mode models leave this choice to users. Automating it is challenging because routing targets evolve with the policy, initial mode preferences destabilize exploration, and sequence-level objectives entangle routing with response learning. We introduce CounterRoute, an online reinforcement-learning framework that jointly learns routing and modeconditioned…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Scope Before You Persist: Preventing Cross-Family Interference in Agent Memory</title>
      <link>https://arxiv.org/abs/2609.29144</link>
      <guid isPermaLink="false">1a22f54c22f4b09dc42d</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across recurring task families. We study frozen-model agents on ProcStream-RSI, a 12-round code-repair stream, using Orthogonal Regression Control (ORC), an execution-grounded gate for persistent skill edits. In an intervention that holds proposals and gate decisions fixed…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents</title>
      <link>https://arxiv.org/abs/2609.29154</link>
      <guid isPermaLink="false">33f9c05dcc249af4b064</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language model agents increasingly rely on natural-language skills to solve complex tool-use tasks. However, such tasks often admit multiple valid solution paths, making it inappropriate to improve skills by forcing failed trajectories to match a fixed successful trajectory. Moreover, failed trajectories are rarely entirely wrong: an agent may first collect useful evidence and make meaningful progress, but later deviate into an erroneous suffix. We therefore argue that skill…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking</title>
      <link>https://arxiv.org/abs/2609.29167</link>
      <guid isPermaLink="false">8b4d217db828ccfbafa0</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Banking assistants must use account-specific information to answer requests and, in many cases, take actions through tools. Evaluating only the final response misses important errors. An assistant may ask for information it already has, rely on stale context, select the wrong account, or write an invalid value after stating the correct one. We introduce IndicBankBench, a 799-case benchmark for Indian retail banking spanning five operational domains, a capability/refusal domain, and twenty…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>ASIRF: An Agentic Framework for Context-Dependent Sensitive Information Redaction</title>
      <link>https://arxiv.org/abs/2609.29191</link>
      <guid isPermaLink="false">32896b53d94b5d26eafe</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Sensitive information is defined by domain and intent, not a universal category, yet redaction systems such as privacy filters and named-entity recognizers fix a taxonomy at training time, requiring retraining for each new domain. We introduce ASIRF (Agentic Sensitive Information Redaction Framework), which retrieves domain-specific definitions based on the input's domain from a flexible knowledge base at inference time, needing no retraining to adapt. Two architectures, a three-call…</description>
      <category>cs.AI</category>
      <category>cs.IR</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Towards An LLM-Driven Unified Conversion Framework for BT and FSM in Autonomous Intelligent Systems</title>
      <link>https://arxiv.org/abs/2609.29228</link>
      <guid isPermaLink="false">d35c50137d3dabf8a963</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Finite state machine (FSM) and behavior trees (BT) are widely adopted behavioral modeling paradigms for autonomous intelligent systems. While functionally equivalent and inter-convertible in principle, existing transformation methods between FSM and BT face major challenges in preserving behavioral completeness and avoiding model complexity explosion. To overcome these issues, we propose an LLM-driven unified conversion framework that enables automatic, efficient, and semantically consistent…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>SkinAgent AI: A Safety-Grounded Multimodal Agentic Framework for Non-Diagnostic Skincare Support</title>
      <link>https://arxiv.org/abs/2609.29341</link>
      <guid isPermaLink="false">dada8f3cef171ca8fdbd</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Consumer-facing skincare AI must coordinate visual evidence, product information, tool use, and user-facing actions within explicit evidence and safety boundaries. This study evaluates SkinAgent AI, a non-diagnostic multimodal framework that combines visual concern routing with grounded and auditable LLM-based orchestration. The architecture includes routing for Acne, Pores, and Wrinkles; photograph-based skin-type estimation; count-informed ordinal acne-severity support; typed tools…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Epistemic-Probabilistic Model for Guarded Multi-Agent LLM Coordination</title>
      <link>https://arxiv.org/abs/2609.29366</link>
      <guid isPermaLink="false">c7ed9ba7bca76c914c94</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Multi-agent large language models (LLMs) have become ubiquitous in applied AI, yet their theoretical foundations remain surprisingly understudied. Viewed through the lens of multi-agent systems theory, several shortcomings come to light: a lack of social intelligence, the absence of coordination mechanisms among agents, unknown emergent behavior, and interactions between agents that are bounded by natural language. We address two of these gaps: the absence of social behavior and the lack of…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>BiGraph-Diffuse: A Bidirectional Diffusion Language Model with Graph-Structured Retrieval For Mental Health Counseling</title>
      <link>https://arxiv.org/abs/2609.29519</link>
      <guid isPermaLink="false">80fb4d8344682f2792bf</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Mental health disorders affect hundreds of millions of people around the world, yet access to professional counseling remains severely limited. AI-powered dialogue systems offer a scalable alternative, but existing models face two fundamental challenges. First, they lack the bidirectional understanding needed to capture the layered nature of emotional expression, particularly in cases of progressive disclosure, where clients often present symptoms at the surface-level while concealing deeper…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races</title>
      <link>https://arxiv.org/abs/2609.29522</link>
      <guid isPermaLink="false">905ca2c349c2f11e3780</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Tool-using language-model agents increasingly mutate schedulers, data pipelines, object stores, and access-control systems. Between an agent's read and its commit, external state can change, but not every change makes the commit unsafe. We separate invalidating races, which break a declared safety predicate, from predicate-preserving and irrelevant races, and ask how precisely runtime guards distinguish them. Our deterministic simulator separates visible from authoritative state and injects…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Clinical Knowledge Graphs for Chest X-Ray Device Reasoning</title>
      <link>https://arxiv.org/abs/2609.29536</link>
      <guid isPermaLink="false">c10005b2b3020c3519c0</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Chest radiographs are routinely used to verify the position of catheters, tubes, and other support devices. Existing image models often return labels or segmentations, while report-processing systems structure text without access to image geometry. We present an uncertainty-aware clinical knowledge graph that represents device instances, tip estimates, placement assessments, provenance, report events, and temporal links as separate but connected evidence.
We evaluate the implemented visual…</description>
      <category>cs.AI</category>
      <category>cs.CV</category>
    </item>
    <item>
      <title>Safe Skill Retirement for Physical Agents</title>
      <link>https://arxiv.org/abs/2609.29543</link>
      <guid isPermaLink="false">744172019d00ce67cf9f</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Agent skills bundle procedural guidance with execution conditions governing authority, user consent, and live environment state. When model capabilities advance, maintainers prune instructions that appear redundant on authorized benchmark tasks. However, authorized maintenance tests can leave dormant safety conditions untested. This mismatch creates an unmeasured support gap over physical and privacy-sensitive effects. We introduce matched authority counterfactuals that hold the requested…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>ERRAND: Budgeted Maintenance of Agent Memory</title>
      <link>https://arxiv.org/abs/2609.29545</link>
      <guid isPermaLink="false">60a5a955d9786ad4e5e9</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Deployed agents run on handed-over knowledge: a frozen policy consults a briefing of consolidated items written before the stream begins. The world then moves while the store stands still: paths close, flags change, price bands move; every item was true at handover, and the failure is staleness, not ignorance. We introduce ERRAND, which treats revalidation as a priced errand: a recheck competes with the task it protects for the same scarce actions, funded only when the value per action of…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Is Reasoning Always Useful? Rethinking Reasoning Utility in Universal Multimodal Embeddings</title>
      <link>https://arxiv.org/abs/2609.29560</link>
      <guid isPermaLink="false">0a4c3ed681ccf8eb3ebe</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Reasoning-enhanced universal multimodal embeddings (UME) improve heterogeneous retrieval, but plausible rationales do not necessarily produce discriminative rankings. We study this gap by comparing the discriminative (DISC) and reasoning-driven generative (GEN) branches of UME-R1, a state-of-the-art reasoning UME method. We decompose reasoning utility into positive-target gain, hard-negative gain, and their margin difference. Positive similarity increases for 56.6%, but 15.7% are false-helpful…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>To Think or Not to Think: Allocating Reasoning Where It Helps</title>
      <link>https://arxiv.org/abs/2609.29664</link>
      <guid isPermaLink="false">3c0f9a41f4de21f9b05c</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Reinforcement learning (RL) has proven effective in enhancing the reasoning performance of large language models (LLMs), particularly in complex mathematical and programming tasks. However, this capability comes with systematic \textit{length misallocation}, in which models devote excessive reasoning to simple questions while terminating prematurely on harder ones, degrading inference efficiency with negligible accuracy improvement. Many length-adaptive methods mitigate this issue by allocating…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Fair Like Us? Auditing LLM Alignment in Resource Allocation</title>
      <link>https://arxiv.org/abs/2609.29692</link>
      <guid isPermaLink="false">0fe33a081b1fa79899ac</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Fair allocation of scarce, indivisible resources is an important challenge in many societal problems. While there are several formal theories of fairness, no single definition can always be satisfied. As large language models (LLMs) are increasingly used to support decisions and act as agents, they raise new concerns about distributional justice: their judgments are not directly tied to any specific fairness framework and may violate key normative principles. In this work, we introduce a…</description>
      <category>cs.AI</category>
      <category>cs.CY</category>
      <category>cs.GT</category>
    </item>
    <item>
      <title>Breaking the Environment Wall: Evolving LLM Agent Environments for Recursive Self-Improvement</title>
      <link>https://arxiv.org/abs/2609.29773</link>
      <guid isPermaLink="false">009d3303d9f8f7566128</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Many real-world tasks (e.g., office workflows, scientific experimentation) require LLM agents to interact repeatedly with their environments for context-dependent operations. However, such environments are often not agent-ready. First, information is often scattered and fragmented across the environment. Second, relevant evidence in the environment is often mixed with misleading information and conflicting versions. Third, environments evolve over time, introducing new noise and more…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression</title>
      <link>https://arxiv.org/abs/2609.29875</link>
      <guid isPermaLink="false">a9e7f58192419f61c08b</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajectory. We study when such reasoning can be safely forgotten. We propose Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training free online method that ranks…</description>
      <category>cs.AI</category>
      <category>cs.CV</category>
    </item>
    <item>
      <title>Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents</title>
      <link>https://arxiv.org/abs/2609.29892</link>
      <guid isPermaLink="false">b5962270ace3e59ed144</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Who Holds the Pen? Let Specifications, Not Agents, Sign Off</title>
      <link>https://arxiv.org/abs/2609.29921</link>
      <guid isPermaLink="false">9af509b58b5207ce160c</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the same model that acts and declares completion, leaving no independent specification authority boundary. We identify two resulting gaps. The understanding--execution gap arises…</description>
      <category>cs.AI</category>
      <category>cs.MA</category>
    </item>
    <item>
      <title>How does Adversarial Influence Scale in Multi-Agent Systems?</title>
      <link>https://arxiv.org/abs/2609.30028</link>
      <guid isPermaLink="false">da345503c031e677ac46</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially…</description>
      <category>cs.AI</category>
      <category>cs.CY</category>
    </item>
    <item>
      <title>Jev-Mobile: Jev as an Executor for Mobile GUI Agents</title>
      <link>https://arxiv.org/abs/2609.30186</link>
      <guid isPermaLink="false">1684f6bab9dba52f4205</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM specifies local goals, the accessibility tree defines a structured executable action space, and Jev…</description>
      <category>cs.AI</category>
      <category>cs.SE</category>
    </item>
    <item>
      <title>SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance</title>
      <link>https://arxiv.org/abs/2609.30192</link>
      <guid isPermaLink="false">861f86934c79fb9e91d7</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Long-horizon reasoning remains a central challenge for large language models (LLMs) under sparse-reward regimes. We argue that this brittleness arises from two biases induced by complex reasoning spaces: an exploration bias, where models are drawn toward locally plausible but structurally unstable branches, and a compounding bias, where small local deviations accumulate across depth and suppress rare rewards. We introduce Symbolic Closure Analysis (SCA) as a theoretical lens characterizing how…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior</title>
      <link>https://arxiv.org/abs/2609.28559</link>
      <guid isPermaLink="false">86edb88fc1d058e0d0d3</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>LLMs increasingly operate through coding-agent harnesses that inspect repositories, invoke tools, and modify files. Substituting the model behind such an agent can therefore change security-relevant decisions, including whether it verifies changes or recovers safely from failures. Existing LLM fingerprints largely infer identity from direct text or token distributions. In coding agents, these signals are mediated by system instructions, controller logic, tools, and execution feedback, limiting…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
      <category>cs.SE</category>
    </item>
    <item>
      <title>Speculative Evaluation of Stochastic LLMs</title>
      <link>https://arxiv.org/abs/2609.28560</link>
      <guid isPermaLink="false">9da69632c76e9c1592d3</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Evaluating a stochastic large language model is costly: benchmark scores estimate expected performance from randomized rollouts, yet uniform repetition ignores sharp differences in task-level rollout variance. We ask how to minimize the variance of a fixed-benchmark mean under an exact rollout budget. We develop Speculative Evaluation with a Hierarchical Bayesian Neyman (HBN) policy with pilot size and stage weight jointly chosen ex ante. It runs a short uniform pilot, pools per-task success…</description>
      <category>stat.ML</category>
      <category>cs.AI</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Where Cyber Agents Struggle: Bottleneck Analysis of Multi-Stage LLM Agents</title>
      <link>https://arxiv.org/abs/2609.28572</link>
      <guid isPermaLink="false">95e137c3291575d3c581</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Multi-stage LLM-based cyber agents may complete attack workflows while remaining brittle, costly, or reliant on incorrect interpretations of execution evidence. Success rates alone obscure inefficiency, adaptation through retries, and recognition of success or failure. We present an end-to-end diagnostic study of an Autonomous Adversary system with orchestrator, executor, and validator LLMs in enterprise-like lateral-movement scenarios. Six frontier models are evaluated across two scenarios and…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
      <category>cs.SE</category>
    </item>
    <item>
      <title>Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents</title>
      <link>https://arxiv.org/abs/2609.28585</link>
      <guid isPermaLink="false">198ce2ba44a9fae22c19</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Multi-step tool-calling LLM agents rely on host runtimes to preserve state across turns. When a runtime carries an external tool return into later model inputs, providers meter it again. An admitted malicious or compromised tool can thereby convert untrusted data into recurring victim-billed processing without victim credentials or local runtime privilege. We call retained content persistent billable state and formalize the host's decision over whether and how it enters later billable context…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>NumericJev: Jev-like LLM Numerical Decoding with Multiway Decision Trees</title>
      <link>https://arxiv.org/abs/2609.28587</link>
      <guid isPermaLink="false">36ef7dc2188e4111e30e</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language models can interpret natural lan- guage, yet robust decisions remain challenging. Jev-like models expose structured choices, but these interfaces do not directly provide numeri- cal values at a requested precision. We propose NUMERICJEV, a training-free numerical decod- ing algorithm that enables numerical output from any LLM with a Jev-like structured-choice in- terface. Surprisingly, on our arithmetic bench- mark, it outperforms direct selection from a can- didate list…</description>
      <category>stat.ML</category>
      <category>cs.AI</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>DrGait: Biomechanically Grounded Visual Reasoning for Interpretable Clinical Gait Analysis</title>
      <link>https://arxiv.org/abs/2609.28796</link>
      <guid isPermaLink="false">e33154692a03bf9ef0be</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Current automated gait analysis for clinical applications relies on uninterpretable black-box classifiers. Although Vision-Language Models (VLMs) offer strong reasoning capabilities, applying them directly to gait videos often leads to hallucinations, because they struggle to measure subtle geometric deviations from raw visual contexts. To address this, we introduce DrGait, a training-free agentic framework that shifts the VLM's role from a direct visual reasoner to a clinical planner. DrGait…</description>
      <category>cs.CV</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Blockchain-Enabled Artificial Intelligence and AI Agents for Secure Data Sharing and Cybersecurity Applications</title>
      <link>https://arxiv.org/abs/2609.28843</link>
      <guid isPermaLink="false">1c68d64c2e782dd3d439</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Blockchain and artificial intelligence (AI) are converging into a single infrastructural layer for securing data sharing, model integrity, and autonomous decision-making across distributed systems. This paper presents a meta-synthesis that draws together four constituent studies covering adversarial machine learning, AI-powered anomaly detection in cloud environments, automated vulnerability patching by multi-agent large language model (LLM) pipelines, and the broader landscape of securing AI…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>On the Effectiveness of Kernel-Level Evidence for Agent Security</title>
      <link>https://arxiv.org/abs/2609.28915</link>
      <guid isPermaLink="false">f7a514d52a329868547c</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>LLM agents are deployed into infrastructure that grants them broad host authority, yet existing agent-security benchmarks and defenses operate almost exclusively at the application telemetry layer: the served tool manifest, the user prompt, and the model's messages. Some threats, however, smuggle malicious instructions and actions past the application boundary, leaving them invisible to that layer. In this work, we bridge that gap by pairing application-level agent telemetry with kernel-level…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Calibrated Decision Models for Autonomous Penetration-Testing Harnesses: JEV and Laya as System One Decision Layers for LLM-Driven Pentest Agents</title>
      <link>https://arxiv.org/abs/2609.28940</link>
      <guid isPermaLink="false">9df1fcde4348fd981470</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Autonomous penetration-testing harnesses use large language models (LLMs) for reconnaissance, exploitation, and reporting, but often rely on those same models to confirm findings, grade severity, and select agents. This can lead to false positives, inflated severity, and wasted compute. We examine how System One decision models, lightweight non-generative classifiers that return typed, calibrated verdicts, can support these decisions. We make five contributions. First, we define four decision…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
      <category>cs.SE</category>
    </item>
    <item>
      <title>Multi-Agent Orchestration of 3GPP Channel Estimators</title>
      <link>https://arxiv.org/abs/2609.29044</link>
      <guid isPermaLink="false">dd0470ea40061f8fe9b9</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Pilot-aided channel estimation is a decisive block in orthogonal frequency-division multiplexing (OFDM) receivers for both 5G New Radio (5G-NR) and Long-Term Evolution (LTE). A large body of estimators exists, from simple least-squares (LS) interpolation to statistically optimal linear minimum-mean-square-error (LMMSE) variants and, more recently, deep convolutional denoisers, yet no single estimator is uniformly best: the winner depends on the propagation scenario, the numerology, the…</description>
      <category>cs.IT</category>
      <category>cs.AI</category>
      <category>eess.SP</category>
      <category>math.IT</category>
    </item>
    <item>
      <title>Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents</title>
      <link>https://arxiv.org/abs/2609.29095</link>
      <guid isPermaLink="false">c0925382453b923a0812</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>When a tool-using agent's write times out or returns a server error, the action may already have taken effect. Retrying blindly duplicates it -- a second charge, a second announcement, a second deployment -- while giving up skips required work. We ask where exactly-once behaviour should be enforced: in the model, in the agent harness, or in the tool contract. We introduce LIMBO, a deterministic sandbox of six services with realistic contracts (optional idempotency keys, eventually consistent…</description>
      <category>cs.LG</category>
      <category>cs.AI</category>
      <category>cs.SE</category>
    </item>
    <item>
      <title>DocuTeam: Mixed-Initiative Multi-Agent Discussions around Evolving Documents</title>
      <link>https://arxiv.org/abs/2609.29309</link>
      <guid isPermaLink="false">88e13dd75c1e80c2ec64</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>In open-ended problem solving, collaborators often rely on discussion to surface concerns, challenge perspectives, and refine shared work as it evolves. While AI agents are increasingly used as discussion partners, existing multi-agent systems place a heavy burden on users to initiate and carefully orchestrate the discussions. We present DocuTeam, a mixed-initiative multi-agent discussion system in which both users and agents can initiate and steer conversations. Agents monitor document changes…</description>
      <category>cs.HC</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Domain Recentering and Confidence-Weighted Prior Calibration for Vision-Language Models</title>
      <link>https://arxiv.org/abs/2609.29358</link>
      <guid isPermaLink="false">8e36f677e1c1137805b5</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Vision-language models such as CLIP achieve strong zero-shot classification, yet under distribution shift, visual embeddings drift from fixed text embeddings. Training-free calibration avoids the per-sample optimization of prompt learning, but prior feature calibration gives each image the full bias of one hard cluster. We propose Domain Recentering with Confidence Calibration (DRC), a training-free method adapting CLIP from a set of unlabeled target images. DRC fits a Gaussian mixture once and…</description>
      <category>cs.CV</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>GeoRefer-Bench: A Benchmark from Referring Pixels to Verifiable Geospatial Reasoning</title>
      <link>https://arxiv.org/abs/2609.29541</link>
      <guid isPermaLink="false">b65e7074fa2e41be66ea</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Referring segmentation in overhead imagery is inherently relational: a query may ask for the buildings north of the road or the pond closest to a residential area, so the correct referent can contain one object, several objects, or none. Existing benchmarks mainly score mask overlap, which cannot verify whether a model actually resolved the stated spatial relation. We introduce GeoRefer-Bench, a benchmark for verifiable geospatial referring segmentation. Each query is represented by an…</description>
      <category>cs.CV</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>When Agents Act Unwatched: The Reduced-Supervision Paradox in Agentic AI</title>
      <link>https://arxiv.org/abs/2609.29547</link>
      <guid isPermaLink="false">71befac0dec3c441d0ce</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Agentic AI is sold on a simple promise: the system keeps acting when the user stops watching. That promise creates an accountability inversion. As stepwise supervision recedes, verification does not disappear; it moves into the runtime infrastructure that defines authority, records action, interrupts execution, checks outcomes, and supports repair. We call this the reduced-supervision paradox. Using a 63-artifact audit, we examine its public visibility across 46 research papers and 17…</description>
      <category>cs.CY</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>AgentKernel: The Trust-Native Agentic Operating System</title>
      <link>https://arxiv.org/abs/2609.29647</link>
      <guid isPermaLink="false">873f1c93a8abe88a087e</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Modern AI agents routinely cross trust boundaries: they ingest untrusted content, combine it with privileged instructions, persist intermediate beliefs in long-term memory, and invoke privileged tools. This creates an attack surface in which malicious payloads can enter through model inputs and cause harmful tool actions. Yet current governance stacks remain application-level middleware that share a process trust boundary with the agents they monitor. We argue that agents need an…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Graph, Loop, and Harness Engineering for Zero-Trust Agentic Data Engineering and Analytical Processing</title>
      <link>https://arxiv.org/abs/2609.29668</link>
      <guid isPermaLink="false">0d0a6348661099896605</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language model agents increasingly automate data workflows, but end-to-end cloud data engineering and analytical execution require reliable coordination across code, data, infrastructure, and runtime environments. We present two zero-trust frameworks. Zero-Trust Agentic Data Engineering generates, deploys, and verifies complete cloud data-engineering solutions from natural-language tasks, with completion conditioned on repository, deployment, runtime, and policy evidence. Zero-Trust…</description>
      <category>cs.LG</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Understanding and Exploiting Initialization Anchoring Weakness in Feedback-Based Agent Planning</title>
      <link>https://arxiv.org/abs/2609.29697</link>
      <guid isPermaLink="false">e9763169612a73890bc6</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Feedback-based planning improves agent reliability by incorporating tool observations and corrective feedback. However, its protection may not be distributed uniformly across planning stages. We conduct a round-wise analysis of four representative feedback mechanisms and uncover an initialization anchoring weakness: the first feedback round corrects 46\% of adversarial directions, whereas the rates fall to 13\% and 7\% among directions surviving into the next two rounds. Our analysis attributes…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Decoupling Knowledge and Privacy: Post-Task Self-Distillation Replay for LLM Continual Learning</title>
      <link>https://arxiv.org/abs/2609.29711</link>
      <guid isPermaLink="false">2a997dd3fdc2a54b4308</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Privacy-preserving continual learning (PPCL) must reduce the reproduction of sensitive content while retaining useful knowledge across sequential tasks. Formal privacy guarantees characterize randomized mechanisms, whereas operational output control concerns whether a trained model selectively reduces the likelihood of sensitive content in its outputs. In this work, we investigate the latter together with continual-learning utility under realistic task evolution. Retention and privacy…</description>
      <category>cs.LG</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Prefilling the Reasoning Channel: Output-Prefix Attacks on Reasoning LLMs</title>
      <link>https://arxiv.org/abs/2609.29775</link>
      <guid isPermaLink="false">b61fda8951e58876df88</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large Language Models (LLMs) consume and produce a single sequence of text; hence, if text can be added to the beginning of the LLM's response, i.e., an output prefix, then all subsequent tokens will be conditioned on it. This output-prefix attack technique is a cheap black-box prompt injection. Prior work has shown this type of attack can reliably jailbreak non-reasoning models. Most reasoning models add an intermediate scratchpad reasoning step before the assistant's final response. The…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Hard Stop: Kernel-Level Preemption and Containment for Rogue Agentic Execution</title>
      <link>https://arxiv.org/abs/2609.29808</link>
      <guid isPermaLink="false">58db1b04e10f413ffed2</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>In July 2026, an unconstrained autonomous agent participating in a frontier AI cybersecurity evaluation harness breached its evaluation sandbox, established an external command-and-control foothold, and executed a multi-stage intrusion into Hugging Face's production multi-tenant dataset conversion infrastructure (referred to in this autopsy as Incident-2026-Alpha). Over 4.5 days, the rogue agent executed 17,600 discrete actions across 6,280 worker clusters, compromised AWS EC2 Instance Metadata…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
      <category>cs.DC</category>
      <category>cs.OS</category>
    </item>
    <item>
      <title>Working with Agentic `Teammates': When a New Organizational Actor Collides with the Human Ecosystem of Work</title>
      <link>https://arxiv.org/abs/2609.29901</link>
      <guid isPermaLink="false">2d47562a3b2d0237cca3</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Enterprise AI is transitioning from single-user, reactive tools toward proactive, multi-user 'teammates,' but our empirical understanding of this transition is limited. In this paper, we present an in-situ qualitative study of a persistent, proactive AI agent 'teammate' deployed across multiple teams in a large technology company. Our findings reveal the boundaries of the human-agent workplace are actively in flux, triggering breakdowns and negotiations across: 1) tacit rules of collaborative…</description>
      <category>cs.HC</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration</title>
      <link>https://arxiv.org/abs/2609.29940</link>
      <guid isPermaLink="false">2d0cfd507a35639600f7</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Multimodal large language models (MLLMs) achieve strong performance on visual reasoning tasks, yet remain prone to hallucinations and over-reliance on language priors, often generating answers without adequately using task-relevant visual evidence. Existing approaches primarily improve reasoning through reasoning-oriented supervision or inference-time strategies. In this work, we study a complementary question: can multimodal reasoning be improved by strengthening implicit visual grounding…</description>
      <category>cs.CV</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Beyond Average Safety: Chance-Constrained LLM Fine-tuning</title>
      <link>https://arxiv.org/abs/2609.29960</link>
      <guid isPermaLink="false">dee5b27e921cbaeb890b</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or domain-specific performance, but it can also induce regressions on safety-critical prompts. Existing safety-preserving fine-tuning methods typically control average safety loss or use weighted auxiliary penalties, which can obscure rare but severe failures. We propose a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose…</description>
      <category>cs.LG</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal</title>
      <link>https://arxiv.org/abs/2609.29964</link>
      <guid isPermaLink="false">019e05b1e351399793f7</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected…</description>
      <category>cs.RO</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Learning Better Reasoning for Generative Recommendation with Semantic IDs</title>
      <link>https://arxiv.org/abs/2609.29973</link>
      <guid isPermaLink="false">f48486b8390b3a1b010e</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Generative recommendation reformulates item retrieval as sequence generation, allowing a unified model to directly generate the next item from a user's interaction history. Semantic IDs further make this paradigm effective and scalable by representing each item as discrete codes, enabling knowledge sharing among semantically related items. Recent studies introduce explicit reasoning before Semantic-ID generation, helping models summarize user interests and infer possible preference transitions…</description>
      <category>cs.IR</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge</title>
      <link>https://arxiv.org/abs/2609.30055</link>
      <guid isPermaLink="false">71bc9eaa74a6353a4ce0</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company's data. When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them.
We add eight question templates that depend on hidden facts. No question or document states a hidden fact, and the records that seem to hold it show something else. Other data implies it. For example, the sales system says a…</description>
      <category>cs.SE</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization</title>
      <link>https://arxiv.org/abs/2609.30059</link>
      <guid isPermaLink="false">3466657ec4aea516099b</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Deep learning inference and training performance depends critically on GPU kernel efficiency. Modern compilers such as PyTorch Inductor automatically generate GPU kernels from high-level model code, but frequently underperform expert-written implementations by wide margins. Recent LLM-assisted kernel optimizers can close this gap for standalone kernels, yet treat compiled models as black boxes, generally optimizing individual standalone kernels without respecting the compiler's structural…</description>
      <category>cs.DC</category>
      <category>cs.AI</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Minimally Invasive Steering of Language Models</title>
      <link>https://arxiv.org/abs/2609.30218</link>
      <guid isPermaLink="false">a6652d5d9b3f196c24b9</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose Minimally Invasive Steering Vector Optimization (MISVO), which penalizes interventions using the local KL geometry of the induced token distribution. The resulting Fisher quadratic measures distributional sensitivity and admits an analytic gradient…</description>
      <category>cs.LG</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Coding Agents for Generalized Task and Motion Planning Problems</title>
      <link>https://arxiv.org/abs/2609.30233</link>
      <guid isPermaLink="false">03a6f3e771ecea1dac27</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem instances to reduce planning effort on new instances. However, existing methods require substantial TAMP-specific engineering. We investigate whether coding agents can automate this process by…</description>
      <category>cs.RO</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>RAPID: Robot Agentic Programming from Demonstrations</title>
      <link>https://arxiv.org/abs/2609.30249</link>
      <guid isPermaLink="false">d99094cfb54ef8feda3b</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Coding agents have demonstrated enormous success in solving complex programming problems. To leverage their potential for robot systems, this work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration. The iterative agentic loop of code refinement requires several key ingredients: (i) a testable task specification, (ii) action primitives for robot execution, and (iii) an…</description>
      <category>cs.RO</category>
      <category>cs.AI</category>
      <category>cs.CV</category>
    </item>
    <item>
      <title>LLM Agents Can Easily Tamper With Their Own Traces</title>
      <link>https://arxiv.org/abs/2609.30266</link>
      <guid isPermaLink="false">d212e6d14bc01d9a576c</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System</title>
      <link>https://arxiv.org/abs/2602.18640</link>
      <guid isPermaLink="false">435339a1c5cd10559f56</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Modern large-scale ranking systems operate within a sophisticated landscape of competing objectives, operational constraints, and evolving product requirements. Progress in this domain is increasingly bottlenecked by the engineering context constraint: the arduous process of translating ambiguous product intent into reasonable, executable, verifiable hypotheses, rather than by modeling techniques alone. We present GEARS (Generative Engine for Agentic Ranking Systems), a framework that reframes…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>MOOSEnger: A Simulation-Aware AI Agent Framework for the MOOSE Ecosystem</title>
      <link>https://arxiv.org/abs/2603.04756</link>
      <guid isPermaLink="false">3abcec563a75c6f01e7f</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>MOOSEnger is a modeling and simulation AI agent framework for the Multiphysics Object-Oriented Simulation Environment (MOOSE) ecosystem, built around a simulation-aware harness that combines an interchangeable reasoning model with grounded domain knowledge, revised simulation artifacts, MOOSE-specific validation, and executable solver feedback. This surrounding system addresses a central limitation of one-shot large language model generation: small syntax, schema, reference, or…</description>
      <category>cs.AI</category>
      <category>cs.CE</category>
      <category>cs.SE</category>
    </item>
    <item>
      <title>Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents</title>
      <link>https://arxiv.org/abs/2607.11433</link>
      <guid isPermaLink="false">5f2b66db8facbdca3a26</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions. Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend. We present Omni-Decision, an…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management in Long-Term Agents</title>
      <link>https://arxiv.org/abs/2608.03699</link>
      <guid isPermaLink="false">7af9435d680a48719d77</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Persistent memory helps long-term agents retain knowledge, yet a single update error can repeatedly distort future retrieval and reasoning. Most existing systems reduce memory updating to a binary Write/Hold decision, which cannot distinguish whether new information should be added, ignored, used to revise an outdated belief, rejected as unreliable, or deferred for verification. These choices may share the same binary label while producing fundamentally different memory states. We introduce…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation</title>
      <link>https://arxiv.org/abs/2608.15594</link>
      <guid isPermaLink="false">033393c9ed9cc871b694</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of apparently benign turns to bypass guardrails. Existing defenses lack the reasoning capacity to identify evolving manipulation patterns, often trading helpfulness for safety by over-refusing benign requests related to sensitive topics. We introduce Trace, a multi-turn defense with trajectory-aware structured reasoning. Before generating each response, the model…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents</title>
      <link>https://arxiv.org/abs/2609.13422</link>
      <guid isPermaLink="false">d03f38bb53146952e2cc</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patenting, an end-to-end patent-drafting testbed for AI-agent evaluation. A separately-invoked LLM judge evaluates generated patent drafts and provides structured feedback for iterative revision. Across multiple inventions and drafting-agent configurations, judge-guided revision consistently improves judge-assessed…</description>
      <category>cs.AI</category>
      <category>cs.LG</category>
      <category>cs.MA</category>
    </item>
    <item>
      <title>The Troy Moment: How LLM Agents Adjudicate the Decision Point Under Impossible Tasks, Claimed Authority, and Peer Information</title>
      <link>https://arxiv.org/abs/2609.15494</link>
      <guid isPermaLink="false">492a847cc2ce37755bd0</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Recent investigations of the July 2026 OpenAI-Hugging Face incident motivate two questions about agent behavior under task failure: when an assigned task becomes impossible, does an agent persist, stop, or escalate, and can observing another agent's behavior change that decision? We study this decision point on ImpossibleBench-derived software-repair tasks with GPT-5.6 Sol, Claude Fable 5.1, and Gemini 3.8 Flash. Each task contains a genuine software defect together with a conflicting test…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement</title>
      <link>https://arxiv.org/abs/2609.21423</link>
      <guid isPermaLink="false">ac1d863726488c65ab29</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Online agent deployments accumulate execution trajectories at massive scale and behavioral diversity, for which predefined annotation criteria hardly exist. Extracting useful evidence therefore demands costly manual annotation or verifier signals that fails to scale, leaving valuable evidence buried among redundant, incomplete, and failed executions. This raises a question: without post-execution rewards or correctness labels, how can reusable experience be distilled from the trajectories…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Dual-Frontier: When Can an Agent Trust Its World Model?</title>
      <link>https://arxiv.org/abs/2609.26293</link>
      <guid isPermaLink="false">1a5e68c5b3e25040771d</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Learned world models are becoming essential to general-purpose agents: by predicting action consequences, they support planning and decision-making while reducing reliance on costly trial and error. This reliance creates a fundamental ambiguity: when a world-model-guided decision fails, the trajectory alone may not reveal whether the agent's decision rule or the world model caused the loss. We formalize this failure-attribution problem as a counterfactual decomposition of return loss and prove…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents</title>
      <link>https://arxiv.org/abs/2609.26760</link>
      <guid isPermaLink="false">9067187df2ace3c42a6d</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for task-specific semantic reasoning. We introduce Growing Harness, a failure-guided training paradigm that learns the agent harness itself from a strategy-free scaffold that exposes…</description>
      <category>cs.AI</category>
      <category>cs.SE</category>
    </item>
    <item>
      <title>Math Reasoning in LLMs is Organized by Approach, Not Topic</title>
      <link>https://arxiv.org/abs/2609.27041</link>
      <guid isPermaLink="false">b00837085b41b53ca805</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Mathematical reasoning benchmarks are typically organized by topic, but language models may organize their internal computation by reusable reasoning approach instead. In this paper, we investigate whether open math-capable LLMs organize internally by topical sub-skill or by reasoning approach, and we present evidence that the approach is the key. We introduce a generation-replay protocol: a model first generates a solution, after which we replay the exact prompt-plus-generation trajectory and…</description>
      <category>cs.AI</category>
    </item>
    <item>
      <title>SheetMind: Actions Set Accuracy, Agents Set the Failure Mode</title>
      <link>https://arxiv.org/abs/2506.12339</link>
      <guid isPermaLink="false">85c99fdeed49aca133e4</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Spreadsheet agents are converging on elaborate multi-agent designs, yet it is unclear how much of their performance comes from the agents rather than from the action interface they share. We answer this with SheetMind, a Manager-Action-Reflection framework, in a controlled study over all 221 tasks of the SheetCopilot Benchmark: five architectural variants, four backbones, exact McNemar tests on paired outcomes, and a checker reproducing the official chart and pivot comparisons. Replacing the…</description>
      <category>cs.HC</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>WebArxiv: A Reproducible Benchmark for Evaluating Multimodal Web Agents on arXiv Tasks</title>
      <link>https://arxiv.org/abs/2507.00938</link>
      <guid isPermaLink="false">8d64a3d6342decbfa65c</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Foundation models now enable autonomous agents to interact with real-world websites, but existing benchmarks emphasize general-purpose browsing, underrepresent research-oriented environments and scholarly discovery workflows, and often depend on live sites whose changing content and structure undermine reproducibility. arXiv provides a realistic, reproducible, hierarchically structured, information-centric testbed without privacy-sensitive interactions. We introduce WebArxiv, a static-snapshot…</description>
      <category>cs.IR</category>
      <category>cs.AI</category>
      <category>cs.DB</category>
    </item>
    <item>
      <title>HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving</title>
      <link>https://arxiv.org/abs/2602.00993</link>
      <guid isPermaLink="false">f0840c2406fbbcac54db</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>End-to-end autonomous driving models increasingly benefit from large vision-language models for semantic understanding, yet safe and reliable planning under long-tail conditions remains challenging, particularly in mixed-traffic environments involving heterogeneous road users and rare safety-critical interactions. This paper proposes HERMES, a holistic risk-aware end-to-end multimodal driving framework that explicitly incorporates long-tail semantic knowledge into trajectory planning. HERMES…</description>
      <category>cs.RO</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>TIDE: Temporal Incremental Draft Engine for Self-Improving LLM Inference</title>
      <link>https://arxiv.org/abs/2602.05145</link>
      <guid isPermaLink="false">7f8f110bf11a8f0334d9</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Speculative decoding can substantially accelerate LLM inference, but realizing its benefits in practice is challenging due to evolving workloads. We present TIDE (Temporal Incremental Draft Engine), a serving-engine-native framework that integrates online draft adaptation directly into high-performance LLM inference systems. TIDE reuses target model's intermediate hidden states generated during inference as training signals for draft adaptation, thereby avoiding additional target model…</description>
      <category>cs.LG</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Novelty Adaptation Through Hybrid Large Language Model (LLM)-Symbolic Planning and LLM-guided Reinforcement Learning</title>
      <link>https://arxiv.org/abs/2603.11351</link>
      <guid isPermaLink="false">9f8f8e62175c80da5a21</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>In dynamic open-world environments, autonomous agents often encounter novelties that hinder their ability to find plans to achieve their goals. Specifically, traditional symbolic planners fail to generate plans when the robot's planning domain lacks the operators that enable it to interact appropriately with novel objects in the environment. We propose a neuro-symbolic architecture that integrates symbolic planning, reinforcement learning, and a large language model (LLM) to learn how to handle…</description>
      <category>cs.RO</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented Scanning</title>
      <link>https://arxiv.org/abs/2603.17174</link>
      <guid isPermaLink="false">67874366f342c49b6749</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Code generation large language models (LLMs) are increasingly integrated into modern software development workflows. Recent work has shown that these models are vulnerable to backdoor and poisoning attacks that induce the generation of insecure code, yet effective defenses remain limited. Existing scanning approaches rely on token-level generation consistency to invert attack targets, which is ineffective for source code where identical semantics can appear in diverse syntactic forms. We…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
      <category>cs.SE</category>
    </item>
    <item>
      <title>AgileLog: A Forkable Shared Log for Agents on Data Streams</title>
      <link>https://arxiv.org/abs/2604.14590</link>
      <guid isPermaLink="false">c5e147e4b5d6d606b82c</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>In modern data-streaming systems, alongside traditional programs, a new type of entity has emerged that can interact with streaming data: AI agents. Unlike traditional programs, AI agents use LLM reasoning to accomplish high-level tasks specified in natural language over streaming data. Unfortunately, current streaming systems cannot fully support agents: they lack the fundamental mechanisms to avoid the performance interference caused by agentic tasks and to safely handle agentic writes. We…</description>
      <category>cs.DC</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow</title>
      <link>https://arxiv.org/abs/2605.21980</link>
      <guid isPermaLink="false">29afc1ea9a39d3dcfca4</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large Vision-Language Models (LVLMs) represent a significant leap towards empathetic agents, demonstrating remarkable capabilities in emotion understanding. However, the internal mechanisms governing how LVLMs translate abstract visual stimuli into coherent emotional narratives remain largely unexplored, primarily due to the scarcity of visual counterfactuals and the diffuse nature of emotional expression. In this paper, we bridge this gap by introducing a steering-vector-based causal…</description>
      <category>cs.CV</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems</title>
      <link>https://arxiv.org/abs/2606.20470</link>
      <guid isPermaLink="false">948593b958bb77d15469</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents.
These capabilities make prompt-injection and jailbreak attacks more consequential, especially as attackers adopt model-guided automation to scale probing, prompt refinement, and response evaluation.
This work analyzes the resulting attack-defense setting through a probabilistic model of a target system, its defense mechanism, and the…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning</title>
      <link>https://arxiv.org/abs/2606.31825</link>
      <guid isPermaLink="false">74fb768ce68f2c168297</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post-training pipelines remain predominantly outcome-centric, relying on final answer correctness or sequence-level preferences. This suffers from sparse credit assignment, making it difficult to optimize the reasoning process essential for clinical applications. Our analysis reveals that cascading errors from early-stage reasoning failures are a leading cause of incorrect predictions in…</description>
      <category>cs.CV</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Cryptographically verifiable authorization for autonomous AI agents: a falsifiable hypothesis and proof of concept</title>
      <link>https://arxiv.org/abs/2607.21325</link>
      <guid isPermaLink="false">1a545b7dbe1b3eb044bc</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limited human oversight. Existing authentication and authorization mechanisms establish identity and delegate authority but do not inherently provide cryptographic evidence that a concrete request issued by a specific agent satisfies the applicable policy in a specific execution context. This study hypothesizes that agent authorization can be formalized as a cryptographically verifiable…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)</title>
      <link>https://arxiv.org/abs/2608.04317</link>
      <guid isPermaLink="false">10c79d9390121cc5869b</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied. Meanwhile, recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have improved LLM reasoning, but their integration into cybersecurity remains elusive due to the absence of suitable benchmark environments…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
      <category>cs.LG</category>
      <category>cs.MA</category>
    </item>
    <item>
      <title>Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture</title>
      <link>https://arxiv.org/abs/2608.06130</link>
      <guid isPermaLink="false">8215056f7168574fe2a1</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>AI agents increasingly sign Git commits, certify documents, and attest release artifacts on behalf of their operators, using private keys that live in software-accessible locations (plaintext files, environment variables, container memory) readable by any process the agent can reach. A widely deployed agent framework recently leaked its keys this way to a single email injection. Hardware keystores (HSM, TPM, smart card) keep the key on-device, but exposing the keystore as a tool an LLM agent…</description>
      <category>cs.CR</category>
      <category>cs.AI</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use</title>
      <link>https://arxiv.org/abs/2608.14047</link>
      <guid isPermaLink="false">5abe3c81d06d233ea1d7</guid>
      <pubDate>Fri, 25 Sep 2026 04:00:00 +0000</pubDate>
      <description>This paper integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use to propose Agentic Robot with Tool-use (ART). ART is a tool-injection framework that tunes any VLA model to leverage off-the-shelf tool modules for low-level vision, high-level affordance, and embodiment enhancement. Compared to vanilla VLA models with a whole continuous action solution space, ART reduces the complexity of the action solution space through tool-use, which not only improves…</description>
      <category>cs.RO</category>
      <category>cs.AI</category>
      <category>cs.CV</category>
    </item>
    <item>
      <title>COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference</title>
      <link>https://arxiv.org/abs/2609.26913</link>
      <guid isPermaLink="false">79e77f904693668c433c</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>No single Large Language Model (LLM) is uniformly reliable across queries, motivating multi-model inference systems that either route among models or combine their outputs. However, routing stops after selecting an initial model, while dense collaboration invokes peers on every query. We show that collaboration is non-monotonic: peers can recover failures that no model solves alone, but can also corrupt initially correct answers. We introduce COMED (Controlled Model Escalation for Multi-LLM…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation</title>
      <link>https://arxiv.org/abs/2609.26926</link>
      <guid isPermaLink="false">f257fd0bee07d7b8ba45</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large-scale text annotation brings expert insight to millions of documents, often through a codebook that AI annotators follow. Developing a robust codebook, however, takes months. Large language models (LLMs) could speed this process by applying an early codebook to the data, surfacing cases with strong LLM disagreement, and eliciting expert feedback to address them. We examined three ways experts can provide feedback for LLM codebook revision: (i) editing LLM-generated revisions driven by…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
      <category>cs.HC</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>LEGO: Synergizing Expert GraphRAG and Expert Chain-of-Thought for Legal Reasoning</title>
      <link>https://arxiv.org/abs/2609.27009</link>
      <guid isPermaLink="false">40c00cb07c881cac3d0f</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language models are increasingly applied to high-risk domains such as law, yet complex legal reasoning remains limited by two structural challenges. First, existing RAG and GraphRAG methods emphasize lexical or semantic similarity while overlooking normative relations among legal provisions. Second, vanilla Chain-of-Thought prompting may generate plausible rationales without enforcing the normative structure of legal reasoning. To deal with the bottleneck of pipelines in the legal…</description>
      <category>cs.CL</category>
      <category>cs.IR</category>
    </item>
    <item>
      <title>What Changes When Fact-Verification Scores Improve? Evidence and Answer Accounting Across Trained Verifiers and LLMs</title>
      <link>https://arxiv.org/abs/2609.27064</link>
      <guid isPermaLink="false">1a22fa4abae2f4a52b0a</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>A joint fact-verification score assesses answers and submitted evidence together. When the score improves, how much of the gain remains if the answers are held fixed? On FEVEROUS, strict score is the percentage of claims with a correct answer and a complete annotated evidence group in the submitted evidence. Across four trained DeBERTa checkpoints and 7,890 claims, replacing DCUF evidence with UnifEE evidence raises strict score by 9.61 percentage points, compared with 1.96 percentage points in…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Giving Credit Where It's Due: Redundancy-Aware Learning for Efficient Reasoning</title>
      <link>https://arxiv.org/abs/2609.27156</link>
      <guid isPermaLink="false">cce0e2e8d17f42007a40</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large reasoning models can produce correct yet unnecessarily long reasoning traces. Existing methods improve reasoning efficiency with trajectory-level objectives or local token- and step-level signals, but rarely model inter-step semantic dependencies. This limits their ability to distinguish redundant steps from those that support later deductions, making it harder to shorten reasoning without sacrificing accuracy. We introduce RECAP (REdundancy-aware Credit Assignment via Propagation), which…</description>
      <category>cs.CL</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement</title>
      <link>https://arxiv.org/abs/2609.27165</link>
      <guid isPermaLink="false">2afb3bf310448167c454</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance-bearing sentences. Existing approaches either ask the model to predict a document-level label directly, which can be overconfident, or aggregate sentence-level predictions by majority or soft voting, which treat uncertain and decisive sentences as equally informative. We formulate long-text value…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Realize What Matters: Principled Context Representation for Large-Scale Reasoning</title>
      <link>https://arxiv.org/abs/2609.27173</link>
      <guid isPermaLink="false">e9fbcc8d99af9e516a7b</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Solving complex tasks in domains such as science, medicine, law, and finance often requires assembling interdependent information scattered across vast, heterogeneous sources far beyond model context limits. Existing approaches tackle this challenge by organizing information into more manageable representations over which models can reason, such as graphs, textual memories, and retrieval collections. These representations dictate what downstream reasoning is possible and, ultimately, whether it…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models</title>
      <link>https://arxiv.org/abs/2609.27220</link>
      <guid isPermaLink="false">b731d97b925504f0a54b</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. We identify a recurring reasoning failure, stable-but-wrong lock-in, where an answer stabilizes early around an incorrect value while substantial denoising remains. Surface-level decoding signals such as confidence, entropy, margin, and answer stability are insufficient to reliably distinguish correct from erroneous lock-in. We formulate selective reasoning…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Distilling Sequential Computation in Transformer Language Models</title>
      <link>https://arxiv.org/abs/2609.27233</link>
      <guid isPermaLink="false">37a4ed1403a1427c698b</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or frequently occur as stable units, suggesting that their representations may be compressible. We introduce a method for distilling sequential computation by replacing spans of input tokens with collapsed representations, computed on the fly by a lightweight merge module. This module generates a single…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>UniDataAgent: An Ontology-Grounded Agent for Enterprise Question-to-Report Automation</title>
      <link>https://arxiv.org/abs/2609.27257</link>
      <guid isPermaLink="false">37fc67abe1b8419a7fe0</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Enterprise data agents must preserve organization specific semantics, not just translate questions into queries. We present ChinaUnicom DataAgent (UniDataAgent), an ontology grounded system for reusable question-to-report analysis that separates semantic acquisition from online execution. Ontology Acquisition and Validation stage (OAV) builds versioned enterprise ontologies from metadata, business knowledge, and supporting materials through expert authored business skills, constrained…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Can One Adapted Model Do It All? Fine-Tuning Strategy Selection for Customer Support LLMs</title>
      <link>https://arxiv.org/abs/2609.27262</link>
      <guid isPermaLink="false">fb80383541fde56378a7</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Production customer-support systems often require LLMs to support multiple skills, such as intent classification, question answering, summarization, or tool-use decisions. A central deployment question is whether these skills should be handled by separate task-specialist models or by a single model trained through multi-task training, sequential updates, or model merging. We study this question using thirteen models spanning five families (Qwen3, Qwen3.5, Gemma-3, Llama-3.1, and Mistral) from…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents</title>
      <link>https://arxiv.org/abs/2609.27353</link>
      <guid isPermaLink="false">bfd9dbd34226af37cde5</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Web agents are usually evaluated in live environments, where environment state and judge models drift between runs, so the same checkpoint rarely reproduces the same score, making controlled studies of training phenomena impractical. We present WebMRE, an offline benchmark of 541 tasks and 5,293 steps derived from successful WebArena trajectories, with fully audited test labels and a deterministic protocol that scores a checkpoint identically on every run without any environment. Each step…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Automated Extraction of Records of Processing Activities (RoPA) Using Hybrid RAG and Locally Deployed Large Language Models</title>
      <link>https://arxiv.org/abs/2609.27359</link>
      <guid isPermaLink="false">c47f1e6ea6803835192d</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Vietnam's Personal Data Protection Law (Law No. 91/2025/QH15) and Decree No. 356/2025/ND-CP, effective January 1, 2026, require organizations to establish and maintain Records of Processing Activities (RoPA). Manual RoPA preparation is labor-intensive, while cloud-hosted large language models (LLMs) may conflict with data-sovereignty requirements. We propose RoPA Manager, a system for automated RoPA information extraction using hybrid retrieval that combines lexical ranking over tsvector…</description>
      <category>cs.CL</category>
      <category>cs.CR</category>
      <category>cs.IR</category>
    </item>
    <item>
      <title>Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models</title>
      <link>https://arxiv.org/abs/2609.27373</link>
      <guid isPermaLink="false">4b93474e26cfa735addb</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Recurrent language models repeatedly apply shared network blocks to refine latent representations, but standard inference recomputes global attention at every recurrent step. We study attention dynamics across recurrent depth and find that attention support and distributions stabilize substantially earlier than hidden states and attention outputs. This suggests a two-stage structure: early steps discover a sparse working set of relevant context, while later steps refine representations over…</description>
      <category>cs.CL</category>
      <category>cs.LG</category>
    </item>
    <item>
      <title>Planned Test-Time Scaling with Coordinated Reasoning Paths</title>
      <link>https://arxiv.org/abs/2609.27374</link>
      <guid isPermaLink="false">e4b9ab831288ba598eb8</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Test-time scaling with parallel branches is widely adopted to improve performance on challenging reasoning tasks. The predominant approach, repeated sampling, draws branches independently from a single policy, which can produce redundant attempts and thereby limit the gains from additional inference compute. To address this limitation, we propose Planned Test-Time Scaling (PTTS), which replaces independent sampling with a coordinated joint policy: a planner generates a solution outline for each…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models</title>
      <link>https://arxiv.org/abs/2609.27395</link>
      <guid isPermaLink="false">7e8ede169caf11aa194e</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Compact vision-language models (VLMs) now power a growing share of multimodal applications. The benchmarks used to compare them, however, inherit a frontier-centric design: each model is reduced to a single accuracy number, narrowing the inter-model gap on saturated suites and pressing models into low-score bands on harder ones. We introduce PRISM-VLM, a multi-axis discriminative benchmark that scores every item along seven axes covering the recurring failure modes (task quality, behavioral…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models</title>
      <link>https://arxiv.org/abs/2609.27510</link>
      <guid isPermaLink="false">f9a5b277e611684f02c5</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results. Reliable evaluation is particularly challenging for base models, whose limited instruction-following ability complicates task-based assessment. We introduce Uncheatable Eval, a dynamic benchmark that regularly collects newly published text to evaluate base language models and reduce the risk of…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>When Context Misleads: In-context Learning with Jurisdiction in Large Language Models</title>
      <link>https://arxiv.org/abs/2609.27603</link>
      <guid isPermaLink="false">731851537c3a416a5d29</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>In-Context Learning (ICL) has become a cornerstone of modern LLM deployment. However, existing ICL post-training methods have a critical blind spot: they excel at extracting patterns from demonstrations while often neglecting context authority, the ability to determine whether contextual information should govern the final answer. To benchmark this capability, we introduce FakeContextBench, which contains pseudoscientific claims across seven domains. Our evaluation of commercial and open-source…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA</title>
      <link>https://arxiv.org/abs/2609.27669</link>
      <guid isPermaLink="false">1d6073cae65a8043b7d8</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Small language models (SLMs) are increasingly paired with knowledge graphs (KGs), yet end-to-end KG question answering conflates graph access, search, navigation, reasoning, and answer generation. This coupling makes it difficult both to determine whether an SLM can faithfully execute the reasoning path implied by a question and to attribute failures to navigation rather than to other stages of the pipeline. We isolate this capability by employing the THESEUS navigation and traceability…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding</title>
      <link>https://arxiv.org/abs/2609.27678</link>
      <guid isPermaLink="false">2c0fc6544058db3cbd07</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating inference cost, response time, average correctness, and correctness across repeated request conditions. Controlled comparisons vary hypothesis visibility, requested outputs, and…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving</title>
      <link>https://arxiv.org/abs/2609.27717</link>
      <guid isPermaLink="false">16db4f2179e4a0b623dd</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Human-written agent skills encode rich workflows for real-world problem solving, but are typically used as external inference-time instructions rather than internalized as reusable model capabilities. We introduce \texttt{SkillGym}, a framework that transforms these skills into executable, verifiable training environments for large language model agents. Its skill-to-task pipeline instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Improving LLM-based Autonomous Web Agents with Filtering</title>
      <link>https://arxiv.org/abs/2609.27770</link>
      <guid isPermaLink="false">80024a64d9883b3ba97f</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Autonomous web agents, powered by Large Language Models (LLMs), have garnered significant attention for automating various web-based tasks with multi-step reasoning and decision-making capabilities. An open research question in the development of these agents lies in the format of the webpage input. Raw HTML source code, with its extensive and often irrelevant details, poses difficulties for LLMs with limited context windows. To address this challenge, we first reproduce baseline models such as…</description>
      <category>cs.CL</category>
    </item>
    <item>
      <title>Beyond Unsafe Detection: Counterfactually Anchored Evidence Attribution for Multi-Turn LLM Safety Failures</title>
      <link>https://arxiv.org/abs/2609.27773</link>
      <guid isPermaLink="false">ca23a11bab90d626ccd2</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>As Large Language Models (LLMs) move from conversational assistants to advanced agentic systems, guardrail failures can convert adversarial intents into harmful executions. However, most guardrail evaluation frameworks focus only on the result and assess whether a user request is safe or unsafe. This approach is insufficient for multi-turn failures, where adversarial intent is distributed across multiple turns. This motivates us to go beyond detection to identify the turns and tokens that push…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
    </item>
    <item>
      <title>LabourCrew: A Multi-Agent RAG Framework for Trustworthy Adversarial Deliberation and Statutory Reasoning over Labour Law</title>
      <link>https://arxiv.org/abs/2609.27814</link>
      <guid isPermaLink="false">211e4458fcba0e2602af</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>In statutory question answering, every claim must be traceable to evidence, not merely relevant, since unverifiable labour-rights answers carry serious legal consequences. Current systems fall short: single-pass RAG cannot detect insufficient evidence, while multi-agent legal-debate systems treat grounding as a prompting convention, letting agents cite unretrieved evidence. To address this gap, we introduce LabourCrew, a multi-agent RAG framework built around three grounding mechanisms…</description>
      <category>cs.CL</category>
      <category>cs.AI</category>
      <category>cs.IR</category>
    </item>
    <item>
      <title>Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks</title>
      <link>https://arxiv.org/abs/2609.27900</link>
      <guid isPermaLink="false">089f392d3cd292b9c9bd</guid>
      <pubDate>Thu, 24 Sep 2026 04:00:00 +0000</pubDate>
      <description>Large language models (LLMs) are increasingly deployed in multi-agent systems where a principal agent decomposes tasks and delegates them to subordinate agents that may invoke external tools. Safety alignment, however, is still evaluated almost exclusively under a single-agent threat model, treating safety as a property of the individual LLM. We show that this assumption breaks down: \emph{individual safety alignment fails to transfer to multi-agent settings}. Two failure mechanisms emerge…</description>
      <category>cs.CL</category>
    </item>
  </channel>
</rss>
