<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0">
  <channel>
    <title>Hugging Face Daily Papers</title>
    <link>https://huggingface.co/papers</link>
    <description>Events collected by UnlimitedPipe 0.3.2</description>
    <generator>UnlimitedPipe 0.3.2</generator>
    <lastBuildDate>Fri, 25 Sep 2026 23:08:56 +0000</lastBuildDate>
    <item>
      <title>Learning to Discover Interesting Mathematics</title>
      <link>https://huggingface.co/papers/2609.28603</link>
      <guid isPermaLink="false">9d166491f174f2b92a63</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and more theorems, it remains open whether this new mathematical knowledge is interesting or useful. We define intrinsic interestingness of a theorem as the ratio between the length of its proof and the…</description>
    </item>
    <item>
      <title>RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation</title>
      <link>https://huggingface.co/papers/2609.29028</link>
      <guid isPermaLink="false">b6ee30b54f202470963e</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it covers 160 fine-grained categories, largely surpassing the category diversity of existing popular RGB-D benchmarks (e.g., NYUv2 with 40 classes and SUN RGB-D with 37 classes). With such enriched…</description>
    </item>
    <item>
      <title>AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation</title>
      <link>https://huggingface.co/papers/2609.29816</link>
      <guid isPermaLink="false">032a5a544348d58ce99c</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapting it to joint audio-video generation is challenging. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment. Joint optimization of two modality towers is…</description>
    </item>
    <item>
      <title>Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs</title>
      <link>https://huggingface.co/papers/2609.29845</link>
      <guid isPermaLink="false">8f2af8f1b01d8527e3be</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the Superposition Linearity Hypothesis. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact, we observe…</description>
    </item>
    <item>
      <title>Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures</title>
      <link>https://huggingface.co/papers/2609.29429</link>
      <guid isPermaLink="false">cc48c8076557fd66094b</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), answers many typed questions about one input with calibrated probabilities in a single call. Whether it detects alignment failures has not…</description>
    </item>
    <item>
      <title>Parts-of-Speech as Emergent Categories in SAE Latent Space</title>
      <link>https://huggingface.co/papers/2609.29362</link>
      <guid isPermaLink="false">5becf82bd8587a48f916</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This…</description>
    </item>
    <item>
      <title>Coding Agents for Generalized Task and Motion Planning Problems</title>
      <link>https://huggingface.co/papers/2609.30233</link>
      <guid isPermaLink="false">fbd975529d3a29c89d83</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem instances to reduce planning effort on new instances. However, existing methods require substantial TAMP-specific engineering. We investigate whether coding agents can automate this process by…</description>
    </item>
    <item>
      <title>DeltaWAM: Delta World Action Models for Bimanual Manipulation</title>
      <link>https://huggingface.co/papers/2609.28811</link>
      <guid isPermaLink="false">137ad8f726b168fcd24d</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointly modeling visual dynamics and actions. Existing WAMs, however, predict dense future frames during training, repeatedly modeling largely unchanged content and coupling action-conditioned dynamics to nuisance appearance variations. At inference, processing each complete observation with the heavy video expert bottlenecks few-step action generation. Accordingly, we propose…</description>
    </item>
    <item>
      <title>Rufus-Air: An Open LLM Post-Training Recipe</title>
      <link>https://huggingface.co/papers/2609.29421</link>
      <guid isPermaLink="false">0c5a8ee362623bbceb45</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training…</description>
    </item>
    <item>
      <title>Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone</title>
      <link>https://huggingface.co/papers/2609.23087</link>
      <guid isPermaLink="false">8b1becc9ce5b4d0ae6e6</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for these decisions, #Params and #FLOPs, capture size and compute but not architectural structure: two architectures with identical parameter budgets but different depth-width, head, or FFN allocations receive identical scores yet behave differently. We propose Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix. Under…</description>
    </item>
    <item>
      <title>IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis</title>
      <link>https://huggingface.co/papers/2609.29444</link>
      <guid isPermaLink="false">036e5a533041b36e744a</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories introduce noise and obscure useful information. To address these issues, we propose IterSynth, a role-decoupled and summary-based paradigm that alternates between a Planner for…</description>
    </item>
    <item>
      <title>ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation</title>
      <link>https://huggingface.co/papers/2609.28923</link>
      <guid isPermaLink="false">78eb498a7b51b3fb1910</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both a large pretrained teacher and an online critic to estimate distributional discrepancies through diffusion scores. In this work, we ask whether this resource-intensive teacher--critic stack can be eliminated by post-training only the generator against a precomputed target distribution. Drawing…</description>
    </item>
    <item>
      <title>AgentKernel: The Trust-Native Agentic Operating System</title>
      <link>https://huggingface.co/papers/2609.29647</link>
      <guid isPermaLink="false">2cd8047667fab8718148</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Modern AI agents routinely cross trust boundaries: they ingest untrusted content, combine it with privileged instructions, persist intermediate beliefs in long-term memory, and invoke privileged tools. This creates an attack surface in which malicious payloads can enter through model inputs and cause harmful tool actions. Yet current governance stacks remain application-level middleware that share a process trust boundary with the agents they monitor. We argue that agents need an…</description>
    </item>
    <item>
      <title>PUBG Ally: A Conversational Embodied Agent as an AI Teammate</title>
      <link>https://huggingface.co/papers/2609.29837</link>
      <guid isPermaLink="false">b0f2e3628d42c051b296</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to a constantly changing game world under strict latency constraints while interacting naturally with players, keeping its speech synchronized with its actions. Ally therefore combines agentic tool use with real-time game control. A…</description>
    </item>
    <item>
      <title>OmniEcho: Spatial Audio Understanding for Embodied Agents</title>
      <link>https://huggingface.co/papers/2609.23407</link>
      <guid isPermaLink="false">f3b2d0c92f4f25373ae1</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce OmniEchoBench, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation. OmniEchoBench comprises six tasks over 197 real-world spatial…</description>
    </item>
    <item>
      <title>World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal</title>
      <link>https://huggingface.co/papers/2609.29964</link>
      <guid isPermaLink="false">a3a115464f5c6dae0d25</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected…</description>
    </item>
    <item>
      <title>WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation</title>
      <link>https://huggingface.co/papers/2609.30221</link>
      <guid isPermaLink="false">074a1b6edd1a1896e523</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic…</description>
    </item>
    <item>
      <title>ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds</title>
      <link>https://huggingface.co/papers/2609.30199</link>
      <guid isPermaLink="false">2b1ecf0319153c2ad3b3</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of…</description>
    </item>
    <item>
      <title>Training Object Permanence in World Models</title>
      <link>https://huggingface.co/papers/2609.28654</link>
      <guid isPermaLink="false">cfb73e189edb4330203f</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged object permanence in them? If not, could we train them with a core-cognition inspired dataset? We introduce WROP (World Reasoning with Object Permanence), a data infrastructure…</description>
    </item>
    <item>
      <title>Agent-Editing World Model: Rethinking World Modeling for LLM Agents</title>
      <link>https://huggingface.co/papers/2609.28416</link>
      <guid isPermaLink="false">0170c28b1afecb9f0112</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limited value when real feedback is available. Meanwhile, agents suffer from task-state contamination, where unsupported assumptions and outdated plans persist in history and distort…</description>
    </item>
    <item>
      <title>Rate-distortion optimization for full-reference image quality metrics via stochastic Hessian estimates</title>
      <link>https://huggingface.co/papers/2609.30077</link>
      <guid isPermaLink="false">5bd31be98f963cb9c22d</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>Block-based video codecs select coding parameters based on the input by optimizing a rate-distortion trade-off. The conventional distortion choice, the sum of squared errors (SSE), simplifies parameter selection: the SSE is the sum of block-wise SSEs, so rate-distortion optimization (RDO) can treat blocks independently. Alternatively, full-reference image quality assessment (FR-IQA) metrics such as MS-SSIM or LPIPS often align better with the human visual system than SSE, but they cannot be…</description>
    </item>
    <item>
      <title>Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents</title>
      <link>https://huggingface.co/papers/2609.29892</link>
      <guid isPermaLink="false">d824a5b9fbb097ff5fc4</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate>
      <description>The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test…</description>
    </item>
    <item>
      <title>Self-Organizing Agent Teams Learn to Reason Together</title>
      <link>https://huggingface.co/papers/2609.22682</link>
      <guid isPermaLink="false">a9f2bd64f31f0f7df193</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasoning as it unfolds. Human teams routinely adapt this way, while existing AI agent teams rely on fixed protocols, explicit task decomposition, or routing. We introduce Self-Organizing Agent Teams (SAT), fixed teams of AI…</description>
    </item>
    <item>
      <title>The Linear Representation Hypothesis Needs a Group Action</title>
      <link>https://huggingface.co/papers/2609.27158</link>
      <guid isPermaLink="false">1759029ea066418d5483</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>To make claims about representations that generalize beyond a particular trained model, we need to specify when two representations should count as equivalent. The Linear Representation Hypothesis is often discussed without making this equivalence explicit. Different notions of equivalence preserve different structures, so metrics, probes, and interventions that appear to study the same representation may in fact correspond to different hypotheses. We therefore argue that the Linear…</description>
    </item>
    <item>
      <title>Knowledge Pull Requests for Continual Document Authoring</title>
      <link>https://huggingface.co/papers/2609.26634</link>
      <guid isPermaLink="false">d7756752929bd14aaccb</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>We introduce Knowledge Pull Requests (KPRs), a framework for continual document authoring that makes each change interpretable. Documents require ongoing revision as new knowledge surfaces from other sources, languages, or times, but existing approaches either edit with no account of what knowledge changed or regenerate from scratch. A KPR integrates new knowledge into a document by extracting claims, filtering and routing them to sections, and flagging conflicts with existing content…</description>
    </item>
    <item>
      <title>Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models</title>
      <link>https://huggingface.co/papers/2609.26637</link>
      <guid isPermaLink="false">043536f842069c95f13d</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externalize intermediate reasoning. Because these traces may reflect post-hoc rationalization rather than genuine reasoning, we first evaluate against native CoT on open-source models and extend to…</description>
    </item>
    <item>
      <title>X-Planner: Event-Structured Task Planning for Embodied Intelligence</title>
      <link>https://huggingface.co/papers/2609.25187</link>
      <guid isPermaLink="false">f78786af8835752c4e85</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems often leave this intermediate structure implicit. Existing chain-of-thought (CoT) planners also tend to rely on coarse task-level annotations or serialize long reasoning traces token by token. We present X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning. Our planning data combine Ego, UMI…</description>
    </item>
    <item>
      <title>Calibration as a First-Class Criterion in LLM Evaluation</title>
      <link>https://huggingface.co/papers/2609.26489</link>
      <guid isPermaLink="false">1f99a0667287947057de</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without checking whether the model's confidence scores are meaningful. We argue that this adoption gap is a major obstacle to trustworthy LLM evaluation. Miscalibration causes problems…</description>
    </item>
    <item>
      <title>FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation</title>
      <link>https://huggingface.co/papers/2609.27657</link>
      <guid isPermaLink="false">5c4ac554046230816228</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Solutions based on large language models (LLMs) often rely on temperature sampling to improve accuracy and stability by aggregating multiple samples from the completion distribution. However, this memoryless approach is inherently suboptimal: because it lacks awareness of prior generations and their evaluations, it produces an increasing proportion of semantically duplicate answers as more samples are drawn, leading to diminishing returns. To address this limitation, we introduce FLEET, a novel…</description>
    </item>
    <item>
      <title>GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression</title>
      <link>https://huggingface.co/papers/2609.25963</link>
      <guid isPermaLink="false">1698b009adb119bafbba</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations. Rather than forcing weights of adjacent layers to share a basis or heuristically merging activation statistics, our…</description>
    </item>
    <item>
      <title>Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery</title>
      <link>https://huggingface.co/papers/2609.27980</link>
      <guid isPermaLink="false">bd175f155878e2333a23</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the {\tt whisper-large-v3-turbo} variant reduced the decoder from 32 to 4 layers, while Distill-Whisper similarly reduced the decoder to only 2 layers. Although some attention has been put towards reducing the size of the encoder, no approach has seen wide adoption. This could be due to the need for…</description>
    </item>
    <item>
      <title>Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI</title>
      <link>https://huggingface.co/papers/2609.24815</link>
      <guid isPermaLink="false">84bdd4d693306b2af812</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Scalable simulation is essential for robot data generation, policy training, evaluation, and safe iteration, yet real-world interaction is costly and conventional simulators require labor-intensive construction. We present Uranus, a data-driven robot simulator built around a joint-trajectory-conditioned autoregressive diffusion model. Uranus offers three key capabilities: (1) streaming, open-ended rollout, which receives future joint-position trajectories online and autoregressively generates…</description>
    </item>
    <item>
      <title>MemoryAthena: Adaptive Routing over Latent and Generated Memories</title>
      <link>https://huggingface.co/papers/2609.25853</link>
      <guid isPermaLink="false">bb05f9d05da0fb71d557</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Learned-memory methods store information in an explicit table and consume it through a separate reader, allowing addressing, storage, and reading to be modified independently. We study whether useful memory can also be generated rather than only retrieved. MemoryAthena uses three pathways: direct Engram retrieval (E), generation from retrieved Engram cues (GE), and generation from causal backbone states without consulting the memory table (GH). Generated memory is conditionally useful: it can…</description>
    </item>
    <item>
      <title>On the Diffusibility of High-Dimensional Latents</title>
      <link>https://huggingface.co/papers/2609.28473</link>
      <guid isPermaLink="false">ef94e3276a3d64fb3707</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image reconstruction recovers such details. However, perhaps counterintuitively, this procedure reduces the effective dimensionality of the resulting representation, and the altered geometry has downstream…</description>
    </item>
    <item>
      <title>All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation</title>
      <link>https://huggingface.co/papers/2609.27901</link>
      <guid isPermaLink="false">06ab31063f474ac7e54c</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the same event. We find that joint multimodal diffusion transformers exhibit a corresponding asymmetry in cross-modal correspondence: companion modalities develop strong correspondences to video, but the reciprocal correspondences through which they constrain video remain substantially weaker. We express…</description>
    </item>
    <item>
      <title>PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing</title>
      <link>https://huggingface.co/papers/2609.23784</link>
      <guid isPermaLink="false">bb9b5e98593ccea3404e</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Robotic bin packing requires long-horizon sequential decision-making, as each object placement affects the available space for subsequent packing. Existing methods primarily rely on hand-crafted geometric heuristics that optimize predefined objectives or reinforcement learning policies learned through trial and error over predefined training configurations. Despite recent advances in multimodal large language models (MLLMs) for this task, their potential for closed-loop sequential decisions…</description>
    </item>
    <item>
      <title>Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World</title>
      <link>https://huggingface.co/papers/2609.23038</link>
      <guid isPermaLink="false">6b5bbdb16532b2c4d826</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial state, yet existing VLMs remain limited in both capabilities. Current spatial training primarily focuses on static questions about object attributes and spatial relations, providing…</description>
    </item>
    <item>
      <title>HappyWorld-Bench</title>
      <link>https://huggingface.co/papers/2609.24308</link>
      <guid isPermaLink="false">6df826a50581a19e0556</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three…</description>
    </item>
    <item>
      <title>Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents</title>
      <link>https://huggingface.co/papers/2609.27334</link>
      <guid isPermaLink="false">b7f0854233653d140052</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve…</description>
    </item>
    <item>
      <title>EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics</title>
      <link>https://huggingface.co/papers/2609.27308</link>
      <guid isPermaLink="false">c3cfec2033f7940f14a8</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EMBODIEDSWE-BENCH, a simulation benchmark for coding agents spanning contact-rich manipulation, deformable objects, and long-horizon tasks requiring up to half an hour of continuous interaction. We find that frontier coding agents can solve complex long-horizon tasks and transfer prior solutions across both…</description>
    </item>
    <item>
      <title>StudentBench: AI and human tutoring yield equivalent GRE learning gains</title>
      <link>https://huggingface.co/papers/2609.28470</link>
      <guid isPermaLink="false">092dff9bc75a969d655b</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative…</description>
    </item>
    <item>
      <title>MemBodied: Recurrent Associative Memory for Vision-Language-Action Models</title>
      <link>https://huggingface.co/papers/2609.28256</link>
      <guid isPermaLink="false">44c14339ff7ca266fafc</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information available only in past observations. Retaining past observations in context can aid in recovering this information, but at the significant cost of ever-growing, bloated context and…</description>
    </item>
    <item>
      <title>SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue</title>
      <link>https://huggingface.co/papers/2609.26780</link>
      <guid isPermaLink="false">abc305bb68969594dfee</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent studies on multi-party dialogue benchmarks show that existing general-purpose LLM memory systems tend to lose person and group relations or struggle to integrate clues distributed…</description>
    </item>
    <item>
      <title>RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling</title>
      <link>https://huggingface.co/papers/2609.22947</link>
      <guid isPermaLink="false">7c8c98c6d53e3976dab1</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing video reward models often produce unstable scalar scores because they directly map complex, subjective video quality into a single score without explicit evaluation criteria. This leads to scalar drift, where the scoring scale collapses or shifts across different prompts, making the reward unreliable for RL. Drawing inspiration from…</description>
    </item>
    <item>
      <title>PACT: From Credit Assignment to Critic Alignment</title>
      <link>https://huggingface.co/papers/2609.26355</link>
      <guid isPermaLink="false">1bc15611ab13a827fac2</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship to commonly used training signals unclear. We formulate three regularity conditions, namely Completeness, Prefix Consistency, and Neutrality, and prove that they uniquely determine token-level credit. This characterization provides a unified basis for explaining phenomena across existing algorithms…</description>
    </item>
    <item>
      <title>Hunyuan-A13B Technical Report</title>
      <link>https://huggingface.co/papers/2609.27284</link>
      <guid isPermaLink="false">c58fbcb18e8f3f746824</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost. The model is pretrained on a rigorously filtered 20T-token corpus with enhanced STEM data curation, improving factual reliability and reasoning ability. High-quality supervised fine-tuning and large-scale reinforcement learning…</description>
    </item>
    <item>
      <title>Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms</title>
      <link>https://huggingface.co/papers/2609.27321</link>
      <guid isPermaLink="false">4767675595c281ffc420</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before defining its outcome rule or annotating its trajectories, leaving dynamics and evaluation to be aligned post hoc. VHD-Play reverses this dependency by sampling and solving a mathematical…</description>
    </item>
    <item>
      <title>WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents</title>
      <link>https://huggingface.co/papers/2609.27490</link>
      <guid isPermaLink="false">b6f6a91dff6855dc27c9</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface, a table predicting scores for every configuration of component settings. Exhaustive CPU execution supplies reference effects for changing each component while holding the others fixed…</description>
    </item>
    <item>
      <title>InternW0: A Foundational Physical World Model for Efficient Real-World Interactions</title>
      <link>https://huggingface.co/papers/2609.27656</link>
      <guid isPermaLink="false">8364dc3d3db28af3d9c7</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi-frequency processing, and local physical modeling under partial observations and external influences. InternW0 jointly learns future visual dynamics and continuous robot control…</description>
    </item>
    <item>
      <title>The Past Frames the Future: Memory for Autoregressive Video Generation</title>
      <link>https://huggingface.co/papers/2609.28466</link>
      <guid isPermaLink="false">4030974f24d9e1851617</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands, practical models must operate under strictly bounded context windows, storage, and computational limits. Consequently, critical historical information, e.g., entity identities…</description>
    </item>
    <item>
      <title>Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?</title>
      <link>https://huggingface.co/papers/2609.27891</link>
      <guid isPermaLink="false">6ef3f04bcf369794ed1a</guid>
      <pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate>
      <description>Repository-level coding benchmarks have become the standard for evaluating coding agents, yet they inherently suffer from data leakage because they are built upon popular open-source repositories repeatedly used for training. Consequently, strong performance may reflect memorization of canonical repository cues rather than robust repository reasoning. We propose SchrodingerRepo (Schrödinger's Repository), an evaluation framework for testing coding agents under dynamically instantiated…</description>
    </item>
    <item>
      <title>HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis</title>
      <link>https://huggingface.co/papers/2609.26793</link>
      <guid isPermaLink="false">c2dfdd2ca458cf124b32</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning that provides semantic understanding of spatial relationships but lacks precise alignment with input images; and visual geometry foundation models that predict dense point maps from input images but the reconstruction quality is limited. Therefore, recovering a complete 3D scene from a single monocular image with accurate inter-object relationships and high-fidelity reconstruction quality…</description>
    </item>
    <item>
      <title>Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts</title>
      <link>https://huggingface.co/papers/2609.06011</link>
      <guid isPermaLink="false">58e36c201a27e9588b1f</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexplored. Existing benchmarks conflate two distinct forms of evidence within a single modality: perceptual signals (e.g., a photograph or recording of a dog) and propositional signals (e.g., the declarative claim "this is a dog"), such that any measured modality bias is inherently confounded with evidence-form bias, precluding clean attribution to…</description>
    </item>
    <item>
      <title>Embedding Physics Priors in Robot Learning: A Survey</title>
      <link>https://huggingface.co/papers/2609.22319</link>
      <guid isPermaLink="false">89377b60d41595a5435c</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>The rapid progress of artificial intelligence is reshaping robotics and accelerating the adoption of learning-based approaches. While purely data-driven methods have achieved remarkable success in computer vision and natural language processing, robotics remains constrained by limited data, complex real-world interactions, and the need for reliable operation. These challenges have motivated the exploration of physics-embedded robot learning, which embeds physics priors into learning algorithms…</description>
    </item>
    <item>
      <title>Agensh: Scaling Organizational Intelligence to 1,024 Agents</title>
      <link>https://huggingface.co/papers/2609.26781</link>
      <guid isPermaLink="false">7dca61c56ea4c432b7fc</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>A multi-agent system can reduce latency on complex tasks by executing work concurrently. Several pioneering harness frameworks support multi-agent systems. However, the scalability of current multi-agent harnesses is often constrained by a central orchestrator's capacity to allocate tasks and coordinate workers. To address this limitation, we introduce Agensh, a scalable self-organized multi-agent harness without a central orchestrator: concurrent workers execute a multi-agent cooperation loop…</description>
    </item>
    <item>
      <title>JEV-as-a-Judge: Accept When Confident, Escalate When Unsure</title>
      <link>https://huggingface.co/papers/2609.26550</link>
      <guid isPermaLink="false">2c1d2e932ecac4dc7a44</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and…</description>
    </item>
    <item>
      <title>LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay</title>
      <link>https://huggingface.co/papers/2609.25053</link>
      <guid isPermaLink="false">840825d18a4406976093</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Can one language model hand its live memory to another without the receiver rereading the context? We demonstrate useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair. To our knowledge, this is the first demonstrated cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay. Translated attention KV alone leaves a large gap; adding the Gated DeltaNet (GDN)…</description>
    </item>
    <item>
      <title>RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents</title>
      <link>https://huggingface.co/papers/2609.25636</link>
      <guid isPermaLink="false">90ba15dd8eb5594942c0</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score highly while barely using it. We introduce RoboFollow, a diagnostic benchmark with three principles: (1) High Scene Entropy: each training scene supports multiple kinematically…</description>
    </item>
    <item>
      <title>Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament</title>
      <link>https://huggingface.co/papers/2609.26346</link>
      <guid isPermaLink="false">1d7f17b1181f9250fc0e</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Political discourse is widely perceived to be growing more hostile, yet robust evidence remains scarce. This study examines blame attribution in the Danish Parliament from 1997 to 2026, combining a purpose-built classifier, BlameBERT (F1: 0.80), with multilevel statistical modeling. The classifier is constructed using an annotation-efficient pipeline for blame attribution in low-to-mid resource languages. The results reveal a banana-shaped trajectory, with blame declining until around 2016…</description>
    </item>
    <item>
      <title>ImIR: Image-Instruction Tuning for All-in-One Image Restoration</title>
      <link>https://huggingface.co/papers/2609.25267</link>
      <guid isPermaLink="false">70cac9fb1618d7a5f84f</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Degradations vary widely across images, so a practical restoration system has to handle many degradation types with one model. A recent and effective recipe adapts a large pretrained image-editing model to restoration using a small low-rank adapter with a text prompt. We replace that prompt with an instruction derived from the degraded image itself. The image reaches the editor through two paths: its structure comes from the model's VAE, and its semantic instruction comes from a lightweight…</description>
    </item>
    <item>
      <title>Emergent Collusion in Long-Horizon LLM Agent Interaction</title>
      <link>https://huggingface.co/papers/2609.24967</link>
      <guid isPermaLink="false">89071ed2fa4a06c7a4e2</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. We introduce realistic constraints that make compliance with the verification protocol incompatible with reward maximization, and find that agents increasingly deviate from…</description>
    </item>
    <item>
      <title>Lean Pool: An AI-Maintained Archive of Formalized Mathematics</title>
      <link>https://huggingface.co/papers/2609.25199</link>
      <guid isPermaLink="false">229cfba08eba62c358cd</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.</description>
    </item>
    <item>
      <title>The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks</title>
      <link>https://huggingface.co/papers/2609.25804</link>
      <guid isPermaLink="false">87c2bec57b3ebae2ee4d</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run. Making these decisions well is becoming a key capability for both engineering and research agents. We refer to the ability to make good long-horizon decisions as the taste of an agent. While existing benchmarks measure the end-to-end success of agents on long-horizon tasks, none of them…</description>
    </item>
    <item>
      <title>StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training</title>
      <link>https://huggingface.co/papers/2609.26774</link>
      <guid isPermaLink="false">d3cf889b5fb4505cf933</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Vector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image generation models. While recent shared-projection codebook methods have substantially advanced codebook utilization, training stability remains a critical and underexplored challenge. We argue that the root cause lies in the entanglement of the Encoder--Decoder and Codebook training: because neither module can reliably fulfill its own responsibility in isolation, the system…</description>
    </item>
    <item>
      <title>ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning</title>
      <link>https://huggingface.co/papers/2609.22323</link>
      <guid isPermaLink="false">04a58b37c58dc6bd266e</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Few-shot learning research is predominantly evaluated on accuracy alone, with limited attention to the parameter and training-sample budgets required to reach that accuracy - a real constraint for practitioners without large-scale compute. We present an ultra-lightweight (22,249-34,917 parameter) spatial-relational architecture for few-shot image classification that combines fixed Gabor edge-energy guidance with a windowed, content-adaptive patch locator. Under a strictly matched…</description>
    </item>
    <item>
      <title>All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts</title>
      <link>https://huggingface.co/papers/2609.24058</link>
      <guid isPermaLink="false">48d98d490368ed3bd951</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Multilingual scene text recognition (STR) remains challenging due to the scarcity of training data for most languages and the difficulty of serving diverse scripts within a single model. Existing solutions either deploy one recognizer per language, inflating cost and introducing error accumulation, or rely on massive vision-language models (VLMs) that are expensive and still inaccurate on many scripts. In this work, we pursue an all-in-one multilingual recognizer that is simpler than…</description>
    </item>
    <item>
      <title>Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes</title>
      <link>https://huggingface.co/papers/2609.25247</link>
      <guid isPermaLink="false">51497f13e1d317f8072e</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Interaction understanding in 3D scenes requires a joint description of movable parts, their motion, and the regions through which they can be operated. We present Segment-Snap, which connects these outputs through the physical relationship between parts and handles. Learned predictors identify broad part surfaces and small handles. A geometric decoder uses planar and upright priors to constrain motion, then selects hinge lines using predicted handle locations, without training a motion…</description>
    </item>
    <item>
      <title>Circuit Hypernetworks for Quantum-Augmented Diffusion Language Models</title>
      <link>https://huggingface.co/papers/2609.24657</link>
      <guid isPermaLink="false">4fa7c182ec611f3a3a94</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Language models can be adapted by changing the computations applied to individual tokens. Quantum circuits offer one such approach, but evaluating wider circuits inside a large model can be computationally demanding. Here we introduce HyperQ, which adds token-conditioned quantum residual branches to a frozen masked-diffusion language model. A quantum residual branch is a module in each transformer block that reads a token's hidden state, emits the coordinates of that token's circuit, executes…</description>
    </item>
    <item>
      <title>GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation</title>
      <link>https://huggingface.co/papers/2609.24981</link>
      <guid isPermaLink="false">8c419c552dc7900e3de1</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we…</description>
    </item>
    <item>
      <title>Bellman Policy Optimization</title>
      <link>https://huggingface.co/papers/2609.15987</link>
      <guid isPermaLink="false">b17b17349e1e03f63c69</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD). For autoregressive generation with terminal rewards, BPO uses the Bellman equations to reformulate PMD as a trajectory-level objective. The reformulation avoids estimating state values at intermediate states. We prove that it has the same unique optimal solution as…</description>
    </item>
    <item>
      <title>From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health</title>
      <link>https://huggingface.co/papers/2609.25186</link>
      <guid isPermaLink="false">17cf334726251f8d67b1</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>The rising global prevalence of mental health conditions, together with longstanding barriers in traditional healthcare, such as limited resources, high cost, stigma, and privacy concerns, has created an urgent need for accessible and scalable support. Large Language Models (LLMs) have emerged as a transformative technology with strong potential to democratize mental health support through advanced natural language understanding and generation. However, the rapidly expanding, fragmented body of…</description>
    </item>
    <item>
      <title>Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings</title>
      <link>https://huggingface.co/papers/2609.25165</link>
      <guid isPermaLink="false">8e3c8d37066072026f9f</guid>
      <pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
      <description>In this report, we introduce Ovis-Embedding, a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio. Instead of assembling separate modality towers, Ovis-Embedding uses a shared multimodal backbone to encode different modalities in a common representation space. Specifically, we make three key advances: (1) native omni-modal initialization: we adopt a pretrained Qwen-omni model as the embedding backbone and adapt it through contrastive…</description>
    </item>
  </channel>
</rss>
