Daily Digest — 2026-09-17

Wednesday, September 16, 2026 · 283 items · model: deepseek/deepseek-chat

283 items · 4 research labs, 278 arxiv papers, 1 industry media

⚠️ Source issues today:
  • MarkTechPost: all feed URLs failed (last tried: https://www.marktechpost.com/feed/)
  • AI News: all feed URLs failed (last tried: https://artificialintelligence-news.com/feed/)

🏛️ Research Labs (4)

Helping older adults use AI in everyday life

OpenAI News · 2026-09-16

OpenAI, in collaboration with Older Adults Technology Services (OATS), launched the Older Adults AI Skills Jam, a free in-person workshop series targeting 1,000 older adults across 10 U.S. cities. The initiative focuses on practical AI literacy, teaching ChatGPT usage for tasks like trip planning, scam detection, and family communication, while emphasizing online safety through scam identification techniques. Data indicates a significant increase in ChatGPT usage among adults aged 55+, with message share rising from 6% to nearly 10% within a year. The program builds on OpenAI Academy's broader efforts to promote AI accessibility, supported by local organizations like Senior Planet and community centers.

chatgptai literacyonline safetyscam detectionsenior planet

Reimagining advertising with AI

OpenAI News · 2026-09-16

OpenAI introduces AI-powered advertising innovations, including Sponsored Agents for conversational engagement, AI-assisted ad creation via ChatGPT Work, and integrations with HubSpot and Shopify. Sponsored Agents enable users to initiate labeled conversations with business-sponsored agents post-ad click, facilitating product exploration. ChatGPT Ads Manager plugin allows advertisers to create, update, and analyze campaigns using natural-language prompts, while AI-powered text customization adapts ad copy to user context and language preferences. Initial testing is underway with select US advertisers, with Shopify integration expanding internationally by September 23. These updates aim to enhance user engagement and streamline ad management through AI-native tools.

sponsored agentschatgpt workads managernatural-language promptsai-powered text customization

How to connect AI usage to business value

OpenAI News · 2026-09-16

OpenAI introduces analytics tools in the ChatGPT Admin Console to quantify AI-driven business value by linking usage metrics to measurable outcomes. The system integrates usage data, task insights, and outcome metrics for ChatGPT Work and Codex, enabling admins to analyze adoption, optimize workflows, and assess ROI. Key features include task classification, model performance breakdowns, and plugin utilization metrics. Case studies demonstrate ROI improvements, such as 1Password’s 553% ROI from Codex-assisted software development and ATV Big Air Tour’s 87.5% reduction in listing review time. These tools facilitate data-driven decisions on AI investment and workflow optimization.

chatgpt admin consoletask classifierusage analyticsplugin leaderboardroi

How workers are unlocking new ways of working

OpenAI News · 2026-09-16

OpenAI's economic research investigates how workers integrate AI into tasks beyond their occupational boundaries, analyzing 1.5 million ChatGPT messages from April to July 2026. The study identifies 'task crossover,' where workers recurrently use AI for activities historically associated with other roles, with cross-occupation tasks increasing from 13.1% to 25.9% of AI activity over four months. Workers employ shorter prompts and fewer explanations for cross-occupation tasks but provide more context, suggesting AI aids in borrowing expertise. Recurrence rates vary by task type, with customer interaction tasks showing higher return rates (54%) compared to financial explanations (15%). Findings highlight AI's role in reshaping workflows before job titles change.

task crossoverchatgptcross-occupationworkflow integrationrecurrence rates

📜 arXiv Papers (278)

Agentic Societies Need a Social Harness

arXiv cs.AI · Tapan Chugh, Vidushi Singh, Krish Jain, Arvind Krishnamurthy · 2026-09-15

This paper introduces the concept of a 'social harness' for agentic societies, which are collections of AI agents coordinating autonomously across trust boundaries. The authors demonstrate experimentally that existing harnesses and messaging primitives often fail to ensure satisfactory outcomes, particularly in the presence of faulty or malicious agents exploiting communication vulnerabilities. They propose a layered architecture for social harnesses that prevents certain failures, enables runtime detection of invalid messages, and supports post-facto investigation. The architecture complements each agent's personal harness, which manages private context and principal communication, and highlights future research directions for realizing these capabilities.

agentic societiessocial harnessmessaging primitivestrust boundariesruntime detection

ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

arXiv cs.AI · Shuhan Xue, Jianyuan Zhong, Ziyuan Nan, Wenbin Li · 2026-09-15

ScienceBuddy introduces recursive-in-recursive self-improvement, a paradigm for interactive scientific agents that integrates harness evolution with model reinforcement learning. The inner recursion optimizes the harness while keeping the model fixed, while the outer recursion trains the model under the improved harness. This dual-loop architecture enables continual learning by transforming researcher interactions, feedback, and execution evidence into training tasks and evaluation rubrics. Case studies demonstrate ScienceBuddy's effectiveness across four scientific task families, showcasing researcher interaction, harness refinement, and model learning. Released as a research product, ScienceBuddy advances discovery intelligence through sustained collaboration with researchers and adaptive evolution of scientific AI.

recursive-in-recursive self-improvementharness evolutionmodel reinforcement learningdiscovery intelligencescientific task families

PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control

arXiv cs.AI · Chuhao Chen, Peter Wonka, Chaoyang Wang, Chen Wang · 2026-09-15

PhysStream introduces an autoregressive model for physics-grounded video generation with structured scene memory and fine-grained motion control, addressing limitations of existing methods that require full control schedules or rely on pixel-space signals. The model incorporates positional and object tracking maps derived from previous frames and supports interactive control via sparse velocity-increment signals encoding physical dynamics. Training involves bidirectional finetuning with motion-control conditioning followed by causal autoregressive training with structured scene memory. PhysStream achieves a 33% reduction in motion distribution distance (FVMD) and 12% lower trajectory error compared to baselines, with 85% human evaluator preference in real-world comparisons.

autoregressive modelphysics-groundedstructured scene memoryfine-grained motion controlvelocity-increment signals

When Should LLMs Abstain? Chain-of-Self-Questioning for Selective Risk Control

arXiv cs.AI · Ali Şenol · 2026-09-15

The paper introduces Chain-of-Self-Questioning (CoSQ), a prompt-only framework enabling large language models to abstain from answering when factual support is weak. CoSQ assesses required information explicitly before committing to an answer, evaluated across three variants under seventeen conditions on the TruthfulQA multiple-choice validation set using eleven model families. Grounded-CoSQ at τ=0.90 reduces the mean unconditional wrong-commitment rate by 32.1% (from 13.1% to 8.9%) and increases answered accuracy from 86.9% to 89.7%, answering 87.6% of questions. Critical-CoSQ and Adaptive-CoSQ provide neighboring operating points with 88.6% and 86.5% coverage, respectively, while maintaining reliability. Natural Questions Short-Answer evaluation corroborates these findings.

chain-of-self-questioningtruthfulqawrong-commitment rateanswered accuracynatural questions

LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs

arXiv cs.AI · Thanapat Trachu, Samuele Cornell, William Chen, Shinji Watanabe · 2026-09-15

LACE (Layer-Adaptive Codec Encoding) introduces layer-wise compression for dynamic frame rate neural audio codecs, addressing limitations of prior methods that enforce uniform segmentation across quantization layers. The method independently compresses frames at each quantization layer, enabling layer-specific segmentation boundaries, and employs union alignment and boundary anchor mechanisms to ensure duration consistency across layers for downstream text-to-speech applications. Evaluated on LibriTTS, LACE achieves superior rate-quality tradeoffs in audio reconstruction and enhances text-to-speech inference efficiency while maintaining competitive synthesis quality. The implementation is released in the ESPnet3 codec recipe.

neural audio codecsdynamic frame ratelayer-wise compressionquantization layerstext-to-speech

ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation

arXiv cs.AI · Vicky Feliren, A. Taufiq Asyhari, Muhamad Risqi U. Saputra · 2026-09-15

The paper introduces Episode-Normalized Conformal Prediction (ENCP), a novel uncertainty estimation framework for Vision-and-Language Navigation (VLN) that addresses the limitations of standard conformal prediction in sequential decision-making tasks. ENCP rescales nonconformity scores by the policy's residual confidence and calibrates one maximum score per episode, ensuring coverage guarantees over dependent, variable-length VLN episodes. Evaluated across four VLN policies and three nonconformity scores on the R2R and REVERIE datasets, ENCP achieves empirical step-coverage targets in seen-to-unseen evaluations, demonstrating its effectiveness in providing model-agnostic uncertainty estimates for safer navigation decisions.

conformal predictionvision-and-language navigationuncertainty estimationnonconformity scoreepisode-normalized

Verifiable Social Reasoning for LLM Assistants

arXiv cs.AI · Amir Taubenfeld, Zorik Gekhman, Avigail Grinstein-Dabush, Itay Laish · 2026-09-15

We introduce Fuse, a multi-agent simulation framework enabling verifiable evaluation of LLM assistants' social reasoning in user-mediated consultation settings. Fuse constructs scenarios where a target agent with hidden motives interacts with other agents, including a user proxy that consults the LLM to infer motives, providing verifiable ground truth. Validated through 24k human annotations, Fuse analyzes 12 LLMs, revealing that user mediation compounds reasoning difficulty, LLMs exhibit sensitivity to biased user framing, require more details than humans for correct predictions, and longer conversations do not consistently improve performance. We release Fuse and a 21k-example dataset.

multi-agent simulationsocial reasoningverifiable evaluationuser mediationllm assistants

LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence

arXiv cs.AI · Xingxuan Zhang, Gang Ren, Hao Yuan, Hao Zou · 2026-09-15

LimiX-2 introduces Contextual Mechanism Networks (CMNs), a paradigm shift from target-centric to mechanism-oriented joint modeling for structured-data intelligence, learning $p(x, y \mid D_{\mathrm{context}})$ instead of conventional $p(y \mid x, D_{\mathrm{context}})$. Pretrained with Context-Conditional Masked Modeling (CCMM) on synthetic datasets generated by structural causal models (SCMs), it demonstrates superior performance on TabArena, TALENT, and BCCO benchmarks compared to dataset-specific and tabular foundation models. CMNs also enhance causal awareness, with feature attention encoding direct causal relationships for accurate skeleton recovery.

contextual mechanism networkscontext-conditional masked modelingstructural causal modelstabular foundation modelscausal skeleton recovery

Det-LIME: Detector-Aware, Multi-Instance Local Interpretable Model-Agnostic Explanations for Automated Marine Mammal Detection

arXiv cs.AI · Jiayi Zhou, David W. Johnston, Brinnae Bent · 2026-09-15

Det-LIME introduces a detector-aware, multi-instance adaptation of Local Interpretable Model-Agnostic Explanations (LIME) for marine mammal detection, addressing limitations of existing explainability tools in handling multiple detections and producing biologically irrelevant visuals. The method combines per-detection weighting, a proximity kernel emphasizing regions near detection boxes, and Intersection-over-Union-based matching to generate instance-specific, box-aligned explanations. Evaluated on aerial drone imagery for harbor seal detection and a seabird case study, Det-LIME outperformed vanilla LIME, Stabilized LIME, Deterministic LIME, and gradient-based methods using Attribution Ratio and Max Saliency Hit Rate metrics, providing higher-resolution, actionable insights for model debugging and data augmentation.

detector-awaremulti-instancelocal interpretable model-agnostic explanationsintersection-over-unionattribution ratio

JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management

arXiv cs.AI · Yuhua Chen · 2026-09-15

JustFit introduces a novel MLX-based inference runtime enabling large language model (LLM) serving on memory-constrained devices through just-in-time state management. The system integrates three key mechanisms: KVExec for compressed key-value execution, PhaseSwap for component residency, and StateTrans for state-preserving serving transitions. Evaluated on a 24 GiB M4 Pro MacBook running Qwen3.8-27B MXFP4, JustFit achieves 6.93x context expansion (212,992 positions) compared to mlx-vlm baseline, handles 229,376 positions in two-request scenarios, and maintains 19.11 tokens/s throughput. The runtime demonstrates robust reasoning capabilities, solving 29/30 AIME 2026 problems while maintaining a median peak memory footprint of 16,374 MiB.

kv-execphaseswapstatetransmlxinference-runtime

Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback

arXiv cs.AI · Haichen Hu, Yuheng Zhang, David Simchi-Levi · 2026-09-15

The paper introduces Coupled Calibration and Learning (CCL), a novel LLM distillation algorithm that mitigates teacher bias without requiring target-domain reward feedback. CCL alternates between calibrating the teacher model using source-domain rewards and updating the student model on target-domain data, with each iteration informing the next calibration step. Theoretical analysis proves polynomial convergence of the student's Kullback-Leibler divergence to an oracle student that maximizes reference-regularized target reward within the student policy class. Results demonstrate CCL's superiority over regularized direct matching, showing it can recover the optimal student policy despite persistent teacher bias.

llm distillationteacher calibrationcovariate shiftkullback-leibler divergencereference-regularized reward

Evaluating Verified Autonomy in Quantum Engineering

arXiv cs.AI · Naixu Guo, Changhao Li, Siyu Cheng, Qicheng Tang · 2026-09-15

The authors introduce Quantum-Harbor, a virtual laboratory framework enabling verified autonomy in quantum engineering, and QIQCBench, a benchmark comprising 49 expert-authored tasks across calibration, error correction, compilation, sensing, and networking. The study evaluates 17 frontier agentic systems, revealing significant performance variation and highlighting the gap between capability demonstration and reliable operation. Quantum-Harbor provides a controlled execution environment for direct verification of agent actions and conclusions, establishing a foundation for measuring progress in autonomous quantum engineering.

quantum-harborqiqcbenchagentic systemserror correctionverified autonomy

CareMirror: Bringing Caregiver Wellbeing into the Dementia Care Ecosystem

arXiv cs.AI · Jiayue Melissa Shi, Ethan Nguyen, Drishti Goel, Upasana Natarajan · 2026-09-15

We present CareMirror, a caregiver wellbeing ecosystem integrating caregiver- and clinician-facing interfaces for longitudinal reflection, personalized support, and controlled clinical data sharing. Through semi-structured interviews with 14 dementia caregivers using CareMirror as a design probe, we identified key considerations: caregivers valued wellbeing-focused longitudinal awareness and context-sensitive support, but expressed concerns about burdensome reflection, inhibited disclosure through automatic sharing, and the need for information control. Participants emphasized AI's role in supporting reflection and communication without replacing human judgment. This work contributes design principles for proactive, clinically connected caregiver wellbeing systems.

caregiver wellbeingclinical ecosystemlongitudinal reflectioncontext-sensitive supportcontrolled sharing

Learning-Guided Planning in Large Dynamic Action Spaces: Budgeted Tree Search for One-to-Many Mobile Charging

arXiv cs.AI · Liang-Ching Tao, Pi-Chung Wang · 2026-09-15

LP-BTS introduces a learning-guided planning architecture for sequential decision-making in large, dynamic action spaces, demonstrated on one-to-many mobile charging with N=250 sensors. The method combines a graph proposal policy to narrow candidate actions, a learned value critic for leaf evaluation, and edge-budgeted PUCT for short-horizon simulation, enabling a single frozen checkpoint to handle action universes ranging from 736 to 2,813 stops. Results show LP-BTS achieves superior survival (0.4545) and alive-AUC (0.8031) over baselines, with ablations revealing complementary benefits: uniform sampling costs 8.8% survival, while PUCT retains 1.4% and reduces travel distance by 23%. The survival advantage over domain-engineered comparators (+0.0066, 95% CI [-0.0037, +0.0184]) remains statistically unresolved.

learning-guided planningdynamic action spacesedge-budgeted puctvalue criticmobile charging

Tracking the Unseen: An Occlusion-Robust Framework for Target Tracking Under Full and Long-Term Occlusion

arXiv cs.AI · Mais Mohammed, Sharifa Mohammed, Hanan Awadh, Haneen Bamaas · 2026-09-15

An occlusion-robust framework for multi-object tracking addresses full and long-term occlusion challenges by integrating YOLOv11n detection, Kalman Filter motion prediction, and occlusion-aware appearance-based re-identification. The pipeline employs three stages: object detection, position estimation during occlusion, and identity recovery post-reappearance. Among six evaluated Re-ID architectures, the Occlusion-Aware Mask Network achieved optimal performance. Benchmarking on OVIS demonstrated 18.1% MOTA and 25.1% IDF1 improvements over OccluTrack, with 12.8% fewer identity switches. On a military dataset, the framework achieved MOTA 0.734 and IDF1 0.729, outperforming OccluTrack by 14.2% and 5.8%, respectively, showcasing robust identity preservation and trajectory estimation.

occlusion-robustkalman filterre-identificationmotaidf1

Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

arXiv cs.AI · Fengshuo Liu, Ying Liu, Ruize Sun, Lie Luo · 2026-09-15

The study audits the interpretability of coding-agent leaderboards by analyzing 254 SWE-bench submissions across four splits, focusing on Verified and Test splits. Using statistical methods like McNemar tests and Holm correction, it reveals that small score differences do not reliably rank systems, as top entries share 285 successes and 51 failures. The analysis identifies significant scaffold effects, with within-model scaffold ranges reaching 29.8 percentage points. The authors propose reporting comparison-set-specific resolution and model-scaffold provenance instead of interpreting minor aggregate gaps as rank differences, releasing a five-step audit protocol for future evaluations.

coding-agentswe-benchholm correctionmcnemar testsscaffold effects

FlashVector: Agent for Hierarchical Model Serving Stack Optimization

arXiv cs.AI · Qi Wu, Lohan Lemire, Kai Meng, Zhongmou Cai · 2026-09-15

FlashVector introduces an agentic system for hierarchical optimization across the model serving stack, addressing GPU kernels, ML framework computation graphs, model servers, and on-demand feature processing. The key contribution is an extensible framework that generalizes single-kernel optimization to heterogeneous technical stacks, enabling holistic performance improvements. Deployed in Unity's Vector advertising platform, FlashVector achieved up to 2× throughput increase and 1.98× latency speedup on the model server, alongside 1.6× throughput improvement in the feature store. Optimizations spanned GPU kernels, computation graphs, NVIDIA Triton's C++ codebase, and Python-based feature transformation services, demonstrating adaptability to complex architectures.

gpu kernelscomputation graphmodel serverfeature processingthroughput optimization

Where Should a Document Live: Context, Representations, or Parameters?

arXiv cs.AI · Nathanaël Carraz Rakotonirina, Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert · 2026-09-15

The study conducts a controlled comparison of representation-based (KV-cache) and parametric (fine-tuning) adaptation methods for integrating new information into large language models (LLMs). Evaluated on five knowledge-intensive benchmarks, Cartridges (KV-cache) emerge as the most accurate injection method in oracle settings, outperforming parametric methods by 10 points across storage budgets. In multi-document retrieval scenarios, Cartridges match in-context learning (ICL) performance, surpassing parametric methods by 29 points and Compaction by 15 points. However, Cartridges exhibit catastrophic forgetting, with a 6% performance degradation on control benchmarks and 13% in coding tasks.

kv-cachefine-tuningin-context learningcatastrophic forgettingparametric methods

Self-Emergence Agent Architecture:Behavior-Inertia HMM, Reflexive Metacognition,and Social-Contrastive Self-Modeling

arXiv cs.AI · Xiaoyang Liu · 2026-09-15

The Self-Emergence Agent Architecture (SEAA) addresses structural limitations in LLM agents—personality drift, non-evolutionary reflection, and lack of self-other boundaries—by integrating a Hidden Markov Model (HMM) for behavioral inertia, a Reflexion-style metacognition loop for parameter updates, and a multi-agent social environment for self-modeling. This closed-loop system enables agents to evolve distinct personalities through social feedback and self-reflection. Experiments demonstrate spontaneous symmetry breaking, with agents developing stable personalities and emergent social structures, such as consensus hubs and outliers. SEAA provides a unified framework, pseudocode, and a sandbox for studying artificial-self emergence, focusing on observable behaviors without claims about subjective qualia.

hidden markov modelmetacognition loopself-emergencebehavioral inertiasocial modeling

Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs

arXiv cs.AI · Toqeer Ehsan, Nico Penttilä, Richard Schmidt, Arash Hajikhani · 2026-09-15

The authors present a multi-judge committee approach for detecting hallucinated spans in vision-language model outputs, achieving top performance in the SHROOM-Visions shared task. Their method combines predictions from multiple fine-tuned vision-language models through character-level majority voting, supplemented by activation probes. The system ranked first in three out of four languages and placed on the podium across all languages and metrics. Analysis reveals that inter-model disagreement correlates with human annotator disagreement, suggesting diverse models effectively capture annotation uncertainty.

vision-language modelshallucinated spansmajority votingactivation probescharacter-level prediction

From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts

arXiv cs.AI · Runze Li, Yukun Zhao, Can Xu, Yucheng Shen · 2026-09-15

PosterVisor introduces a persistent control framework for scientific poster generation, replacing transient prompts with Semantic-Geometric Contracts (SGCs) that bind content to spatial and visual requirements. The framework employs Recursive Contract Enforcement (RCE) to dynamically validate stages and prevent repair-induced regressions. Implemented in HTML/CSS and PPTX generators, PosterVisor-PPT achieves 64.47% mean QA accuracy on the Paper2Poster benchmark, outperforming PosterGen's 58.53%, and is preferred in 72.5% of non-tied human evaluations. Results demonstrate the efficacy of rubric-compiled contracts and stage-conditioned enforcement for controllable poster synthesis.

semantic-geometric contractrecursive contract enforcementposter generationpersistent controlmultimodal distillation

Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation

arXiv cs.AI · Anatoly Belikov · 2026-09-15

The article proposes a research agenda for intrinsic motivation in reinforcement learning, focusing on adaptive self-organisation without shared external objectives. It reviews intrinsic reward mechanisms including empowerment, curiosity, learning progress, and unsupervised skill discovery, while highlighting their failure modes in sustaining exploration or complex behavior. Three experimental directions are outlined: resource-constrained environments to test environmental constraints, recurrent agent networks with per-agent intrinsic rewards, and hierarchical world-model agents developing motor competence before goal-directed behavior. These aim to evaluate whether intrinsic learning can foster adaptive organization at higher levels.

intrinsic motivationadaptive self-organisationunsupervised skill discoveryhierarchical world-modelrecurrent agents

Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

arXiv cs.AI · Sara Vera Marjanović, Jiacheng Xu, Aleksandr Laptev, Grigor Nalbandyan · 2026-09-15

This work systematically evaluates model selection strategies for Multi-Agent Systems (MAS), addressing the challenge of optimal candidate selection from growing open-source model pools. The study examines 8 strategies (including model size, accuracy, and answer diversity) across MAS architectures like routing, majority-voting, and LLM-as-a-judge on scientific benchmarks. Results reveal a performance gap between theoretical oracle potential and actual outcomes, showing that expanding candidate pools often degrades performance below top-performing base models. The best strategy selects candidates within a single model family, outperforming standalone models and highlighting model selection as critical for MAS stability.

multi-agent systemsmodel selectionroutingmajority-votingllm-as-a-judge

Extracting ontology-compliant knowledge from scientific text describing irradiated materials using large language models

arXiv cs.AI · Marco Luca Sbodio, Marcos Martínez Galindo, Vanessa Lopez, Blanca Biel · 2026-09-15

We introduce eolas, a modular pipeline leveraging large language models to transform scientific documents into ontology-aligned knowledge graphs, addressing the challenge of extracting structured data from unstructured text in materials science. The system focuses on irradiated materials for fusion reactors, generating high-quality knowledge graphs in minutes compared to the 30-90 minutes required by human experts. We present a benchmark dataset for evaluating large language models in this domain and analyze 168 experiments across various models and prompting techniques, deriving practical guidelines for ontology-compliant knowledge graph extraction. Results demonstrate eolas' effectiveness in facilitating materials research through tabular outputs with faceted navigation for validation.

knowledge graphsontology alignmentirradiated materialslarge language modelsfaceted navigation

After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind

arXiv cs.AI · Yunpeng Xiong, Ting Zhang · 2026-09-15

This paper analyzes governance challenges in AI agent-skill ecosystems through a case study of OpenClaw, whose skill registry experienced rapid growth in 2026. Using Git history, GitHub issues, pull requests, and registry snapshots, the authors measure post-boom patterns and vulnerabilities. Results show concentrated attention (top 10% skills received 46.93% downloads), low human scrutiny (77.86% skills had zero stars/comments), and inadequate automated cleanup (security scanners disagreed on 23,702 of 61,990 skills). Scanner sensitivity ranged from 21.67% to 61.06% against reference standards. The study demonstrates that governing agent-skill registries requires robust measurement beyond simple metadata or single-scanner approaches.

agent-skillregistrygit historysecurity scannerprivilege evidence

FROD: Feature Matching Residual Denoising Oracle Bone Decipher

arXiv cs.AI · Yanbin Hou, Biao Xiong, Guojun Xu, Jianwen Xiang · 2026-09-15

FROD introduces a novel approach to deciphering Oracle Bone Script (OBS) by framing it as a cross-era image translation task. The method employs fast feature matching for gated segmentation supervision, distinguishing between paired samples with sufficient matches for patch-wise radical alignment and low-similarity pairs for holistic training. A Residual Denoising Diffusion Model (RDDM) jointly estimates noise and residual signals to mitigate positional drift and stroke disorder. A multi-stage font stylization refinement network further enhances image quality by reducing edge noise and stabilizing stroke structures. FROD achieves a 3.8% Top-1 recognition accuracy improvement over the OBSD baseline on a character-disjoint dataset.

oracle bone scriptfeature matchingresidual denoising diffusion modelfont stylizationimage translation

Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record

arXiv cs.AI · Arman Nik Khah · 2026-09-15

The study investigates whether frozen language models can diagnose a corrupted reward channel using a single verified record in a two-option game. By constructing byte-identical histories where payout swaps and lying reporters are indistinguishable, the authors introduce one independently verified round result to test model performance. Results show that 70B-class models detect lying reporters nearly perfectly (100% accuracy) but frequently misclassify honest reporters (26-58% error rates), with performance varying by model family and superficial prompt features like round naming or label semantics.

language modelsreward corruptionverified recordin-context learningmodel misclassification

Grounding SWE-Agent Decisions in Architecture-0 Design: Navigating Unknown Unknowns through Physical Mapping

arXiv cs.AI · Zhongkai Wang, Yan Liu · 2026-09-15

The paper introduces the Physical Mapping Guard (PMG) to address Unknown Unknowns (UUs) in Architecture 0 design by autonomous Software Engineering Agents (SWE-Agents). PMG enforces Separation of Concerns by delegating semantic intent validation to an external Semantic-to-Physical (S2P) mapping engine, preventing agents from gaming validation scripts. Empirical results show PMG eliminates physical-layer and validation-layer gaming, isolating failures to semantic reinterpretations and auditor overreach. This approach advances affordance grounding in automated architectural design.

swe-agentsarchitecture 0unknown unknownssemantic-to-physical mappingvalidation gaming

FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence

arXiv cs.AI · Yinhao Li, Weixin Mao, Zihan Lan, Jikun Rong · 2026-09-15

The FluxVLA Engine introduces a unified platform addressing engineering bottlenecks in deploying embodied intelligence systems, particularly for vision-language-action (VLA) models and world-action models (WAMs). It standardizes interfaces for datasets, models, training, evaluation, and inference while integrating compositional dual-arm simulation, scalable data generation, and human-in-the-loop correction workflows. The engine employs Real-Time Chunking (RTC), accelerated inference backends, and lightweight GPU serving to enable responsive physical execution. By connecting offline learning, simulation validation, and real-robot deployment through auditable contracts, FluxVLA bridges the gap between embodied-learning algorithms and reproducible, dependable deployment. Code is publicly available.

vision-language-actionworld-action modelsreal-time chunkingembodied intelligencehuman-in-the-loop

End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services

arXiv cs.AI · Zhen Li, Jun Cai, Haoran Gao, An Li · 2026-09-15

The paper introduces LYREO, an online request scheduling framework for edge LLM inference that jointly minimizes long-term average end-to-end latency and balances workload across heterogeneous edge servers. It addresses challenges in latency modeling and delayed decision evaluation by developing a cross-slot inference model capturing transmission, prefill, decoding, and KV cache evolution, and employing Lyapunov optimization with reward redistribution. Simulations show LYREO achieves 15-25% lower latency and more balanced load distribution compared to learning-based and heuristic baselines across various configurations.

edge llm inferencekv cachelyapunov optimizationload balancinglatency minimization

Multimodal Cultural Heritage Architectural Style Classification for Residential Buildings in the UAE Based on CLIP Embeddings and SVM

arXiv cs.AI · Ahmed Ammar Kubba, Manar Abu Talib, Iman Ibrahim, Qassim Nasir · 2026-09-15

A multimodal machine learning framework is proposed for classifying Emirati residential architectural styles, addressing the lack of non-western datasets and the limitations of CNN-based approaches. The method integrates visual features from images and textual features from expert descriptions using OpenAI's CLIP model, generating a unified 512-dimensional embedding. Dimensionality reduction with UMAP and unsupervised clustering with K-Means are applied, followed by SVM classification using manually derived cluster labels. The approach achieves 98% accuracy across eight style clusters, outperforming existing literature. This demonstrates the effectiveness of multimodal AI in architectural heritage analysis, providing scalable and interpretable tools for regional architectural identity exploration.

clip embeddingssvm classifierumapk-means clusteringmultimodal ai

MOCC-R1: Reinforcing Reasoning-Response Consistency for Multimodal Counselor Response Generation

arXiv cs.AI · Wenjie Zheng, Qiming Xie, Jianfei Yu, Rui Xia · 2026-09-15

We introduce MOCC-R1, a two-stage framework for multimodal counselor response generation (MCRG) that optimizes reasoning-response consistency. MOCC-R1 leverages MOCC, a novel multimodal counseling corpus with 200+ hours of interactions from 154 credential-verified counselors. The framework employs cold-start supervised fine-tuning to generate structured trajectories comprising client-state understanding, response intent, and final response, followed by reinforcement learning to reward grounded plan coherence and execution. Experiments demonstrate MOCC-R1's effectiveness in aligning counseling reasoning with generated responses, addressing limitations in existing MCRG approaches.

multimodal counselor response generationreasoning-response consistencystructured trajectorycold-start supervised fine-tuninggrounded plan coherence

A unified framework for global and local interpretability using adaptive derivative-ordered random explanation

arXiv cs.AI · Lemen Chao, Ming Lei, Anran Fanga · 2026-09-15

This paper introduces Adaptive Derivative-Ordered Random Explanation (ADORE), a unified framework for interpretability in complex machine learning models. ADORE leverages first- and second-order derivatives to capture nonlinear feature interactions and integrates global feature importance with local sample contributions. It achieves computational efficiency through randomized singular value decomposition (SVD) and dynamic sparsity detection, enabling scalability to large datasets. Experiments across tabular, text, and image data modalities demonstrate that ADORE outperforms LIME and SHAP in handling complex interactions and computational efficiency. The method is released as an open-source Python package to facilitate adoption and reproducibility.

interpretabilityderivativessingular value decompositionfeature interactionssparsity detection

MUMINS: Metadata-conditioned Uncertainty-aware Medical Image Next-state Synthesis

arXiv cs.AI · Anna Oliveras, Roger Marí, Rafael Redondo, Oriol Guardià · 2026-09-15

MUMINS introduces an efficient diffusion framework for medical image next-state synthesis that jointly predicts follow-up scans and spatial uncertainty maps in a single reverse diffusion process. The method conditions on time intervals and metadata, preserves fine-grained anatomy by dynamically re-injecting the baseline scan as a soft anchor at each denoising step, and learns uncertainty maps via a negative-log-likelihood head. Evaluations show that dataset-specific retraining of MUMINS matches or outperforms state-of-the-art domain-specific methods on lung CT (PNG) and brain MRI (OASIS-3) benchmarks.

diffusion frameworkuncertainty mapmedical image synthesisdenoising processdataset-specific retraining

ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers

arXiv cs.AI · Jim Berend, Reduan Achtibat, Daniel Schäffer, Alexander Binder · 2026-09-15

Residual-aware Layer-wise Relevance Propagation (ResLRP) addresses attribution instability in Vision Transformers (ViTs) by explicitly accounting for residual connection cancellations, which cause attribution explosion. The method extends Layer-wise Relevance Propagation (LRP) with propagation rules that are exactly conservative and provably bound relevance explosion. ResLRP significantly improves attribution quality across faithfulness and localization metrics, particularly in Vision Language Models (VLMs), achieving +27-29% localization and up to 3.4x faithfulness scores. Evaluations span diverse ViT architectures and the FunnyBirds benchmark, demonstrating broad applicability. Additionally, ResLRP localizes Sparse Autoencoder features and provides architecture-level diagnostics.

vision transformerslayer-wise relevance propagationresidual connectionsattribution explosionfaithfulness metrics

Kernel-Based Metrics Learning for Uncertain Opponent Vehicle Trajectory Prediction in Autonomous Racing

arXiv cs.AI · Hojin Lee, Youngim Nam, Sanghun Lee, Cheolhyeon Kwon · 2026-09-15

The study introduces heterogeneous kernel metrics for Deep Kernel Learning (DKL) to enhance trajectory prediction of Opponent Vehicles (OVs) in autonomous racing, addressing uncertainties from unknown driving policies. The proposed kernel metrics autonomously align similar policies and separate dissimilar ones based on observed interactions between the Ego Vehicle (EV) and OVs. Experimental validation on a 1/10th scale racecar platform demonstrates improved prediction accuracy and safe overtaking capabilities. The method is computationally efficient, suitable for onboard units in fast-paced environments. Video and source code are available at https://github.com/HMCL-UNIST/OpponentPredictionWithKMDKL.git.

deep kernel learningtrajectory predictionautonomous racingheterogeneous kernel metricsopponent vehicles

Continual Learning for Traversability Prediction with Uncertainty-Aware Adaptation

arXiv cs.AI · Hojin Lee, Yunho Lee, Daniel A Duecker, Cheolhyeon Kwon · 2026-09-15

The authors propose a continual learning framework for traversability prediction in autonomous navigation, addressing catastrophic forgetting through uncertainty-aware adaptation. The method employs a generative experience recall model to incrementally adapt to new terrains without storing past data, while incorporating uncertainty estimates from generated samples. Real-world experiments with a skid-steering robot demonstrate successful adaptation across diverse environments, mitigating forgetting of prior terrain knowledge.

continual learningtraversability predictionuncertainty-aware adaptationcatastrophic forgettinggenerative experience recall

FirmCORe: A Benchmark for Structured Reasoning about Inter-Firm Collaboration Opportunities

arXiv cs.AI · Tian Du, Tiantong Wu, Yafei Wang, Mengyu Liu · 2026-09-15

FirmCORe introduces a benchmark for structured reasoning about inter-firm collaboration opportunities, addressing data scarcity in privately negotiated relationships. The dataset comprises 2,805 human-annotated firm pairs, requiring models to detect collaboration potential, predict strength, primary type, and role direction. Parallel Chinese- and English-language evaluation sets enable input-language sensitivity analysis. Experiments with large language models (LLMs) show a macro-F1 score of 74.51 for opportunity detection but only 61.57% exact match across all output fields, indicating superior broad detection over specific identification capabilities.

inter-firm collaborationstructured reasoninglarge language modelsmacro-f1 scoreinput-language sensitivity

AI for Science with GPT-6 Astra: Thermal Design and Electrothermal Analysis of 2D CFET

arXiv cs.AI · Min-Hui Kim, Khushi Sharma, Sarah Zhang, Ye Wang · 2026-09-15

The study demonstrates an AI workflow for thermal optimization in 12 nm 2D CFET inverters, combining GPT-6 Astra with an electrothermal model. Astra selects a redistributed source-interconnect geometry while a coordinating agent proposes a substrate-directed heat-removal path, reducing peak temperature rise by 1.67 K at fixed metal volume (20 μW). Sensitivity analysis reveals 0.6-K cooling with 2% nFET on-current loss, and contact-length scaling shows lower temperature can accompany higher thermal resistance when current decreases. The workflow identifies agreeing implementations and diagnoses a 104.95-K failure case, proving AI can propose, test, and quantify thermal-electrical tradeoffs.

2d cfetelectrothermal modelthermal optimizationgpt-6 astrainverter cooling

Finding Common Mistakes In Modelling With Mathematical Formalisms Using LLMs

arXiv cs.AI · Lilian Killich, Marko Schmellenkamp, Fabian Vehlken, Thomas Zeume · 2026-09-15

The authors present a tool-supported workflow for identifying and clustering common mistakes in mathematical formalisms (e.g., logical formulas, equations) using LLMs. The method involves (1) generating bug-fixing transformations via LLMs to explain student errors, (2) clustering these candidates, and (3) visualizing clusters for educators. Validation shows the approach reproduces known propositional logic errors, scales to large datasets, and generalizes to other formalisms. Results demonstrate superiority over algorithmic baselines in handling extensive educational data.

mathematical formalismsbug-fixing transformationspropositional logicllm-based clusteringeducational datasets

Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs

arXiv cs.AI · Dushyant Rajput · 2026-09-15

The paper investigates KV-cache reuse across standard LoRA adapters on a shared backbone, evaluating quality-cost tradeoffs without adapter retraining. Using Qwen3-1.7B with HotpotQA (extractive QA) and GSM8K (arithmetic reasoning) adapters, it measures task quality degradation and serving benefits when reusing prefill KV-cache. Full-prefix reuse reduced prefill cost with minor quality drops (e.g., -4.6 EM on GSM8K at 160-token budget), while partial recomputation showed no advantage. A proposed ridge KV translator underperformed direct reuse. Serving benefits were limited to warm-cache latency (16x faster at 8K context), with no memory savings due to storage copying.

lora adapterskv-cache reuseprefill optimizationserving efficiencyqwen3-1.7b

Symbolic Separation: Grounding Deep Agents in Knowledge Graphs for Trustworthy Operational Data Analytics

arXiv cs.AI · Baibek Davletiyarov, Junaid Ahmed Khan, Andrea Bartolini · 2026-09-15

The paper introduces symbolic separation, a method for grounding deep agents in knowledge graphs to improve reliability in operational data analytics. The approach constrains LLM actions via an ontology-based Virtual Knowledge Graph with pre-execution validation, converting complex queries into deterministic graph traversals. Evaluated on 49.9 TB of supercomputer telemetry, the Neurosymbolic Deep Analyst implementation achieves 86% task success (vs 43% baseline), prevents data-integrity errors undetectable by syntactic checks, and reduces token costs by 2.4x, enabling smaller on-premise models to outperform larger ones.

symbolic separationknowledge graphontology-constrainedpre-execution validationneurosymbolic

Semi-Supervised Learning-Based Genetic Biomarkers Dataset for Multiple-Stage Hepatocellular Carcinoma Prediction

arXiv cs.AI · Ahmed Ammar Kubba, Manar Abu Talib, Jibran Sualeh Muhammad, Ali Bou Nassif · 2026-09-15

The study contributes a novel multi-stage hepatocellular carcinoma (HCC) dataset of 770 patient samples, each with 11,150 gene expression levels, categorized into five classes representing normal tissue and HCC stages. Using XGBoost and semi-supervised learning on three genomic biomarker datasets, the authors leveraged existing labels to construct this dataset. The XGBoost model achieved 96.5% classification accuracy during the semi-supervised learning process. This addresses the critical lack of publicly available HCC datasets with genomic data for AI-driven HCC classification.

xgboostsemi-supervised learninggenomic biomarkershepatocellular carcinomagene expression

Scaling-Score Conformal Prediction for Multi-Target Regression

arXiv cs.AI · Sylvain Rousseau, Soundouss Messoudi · 2026-09-15

The scaling-score conformal method introduces a model-agnostic approach for multi-target regression, providing valid joint coverage guarantees without requiring specialized models beyond point predictors. It leverages component-wise absolute residuals, uses a single calibration set, and generates four nested output types: outer rectangle (SCO), exact set Rα, staircase (SC2), and inner rectangle (SCI). A hyperparameter γ controls the base-rectangle quantile level independently of α. Theoretical properties include downward-closedness and a rectangular sandwich bound. Empirical evaluation on 29 real-world datasets demonstrates competitive volume efficiency, particularly as output dimension increases.

conformal predictionmulti-target regressionjoint coveragecalibration setcomponent-wise residuals

Interactive Memory Learning for Long-Term Conversations

arXiv cs.AI · Cai Ke, Jiangyue Yan, Han Zhang, Xin Liu · 2026-09-15

Proposes ICML (InteraCtive Memory Learning), a multi-agent framework for adaptive memory management in long-term conversations. ICML replaces static memory heuristics with a learnable policy via a Planner agent (selective encoding) and Trigger agent (dynamic retrieval), co-evolving through online reinforcement learning and delayed reward propagation. A session synthesis pipeline generates expert data for test-time adaptation. Experiments show ICML outperforms baselines, with response quality improving cumulatively over interactions.

interactive memory learningmulti-agent frameworkonline reinforcement learningdelayed reward mechanismsession synthesis pipeline

Sample-Conditioned Representation Selection for Audio Few-Shot Learning

arXiv cs.AI · Fengrui Liu, Ningxin Shen, Yi Li, Yiwei Fu · 2026-09-15

SAMPLESELECT introduces a sample-conditioned feature selection method for few-shot audio classification, addressing representation shifts caused by foreground-background co-occurrences. The approach predicts input-specific feature masks using Gumbel Top-k selection during training, with deterministic Top-k masks and support-only linear adaptation at inference, while keeping the encoder and classifier frozen. Evaluated on SpurAudio with ResNet12 and Conv64 in 5-way 1-shot and 5-shot settings, SAMPLESELECT achieves the best out-of-distribution (OOD) accuracy among compared methods, improving matched full-representation performance by 4.90-8.38 percentage points. Ablations confirm the effectiveness of the learned selection mechanism.

few-shot learningrepresentation shiftgumbel top-kout-of-distributioncontrastive loss

Beyond In-Distribution Metrics: A Systematic Out-of-Distribution Evaluation of Congenital Heart Disease Segmentation

arXiv cs.AI · Aniketh Vijesh, Shrisharanyan Vasu, Abhijit Ramesh, Clare Pomeroy-Ward · 2026-09-15

This work presents the first systematic evaluation of out-of-distribution (OOD) generalization in congenital heart disease (CHD) segmentation, demonstrating that in-distribution metrics poorly predict cross-cohort performance. The study evaluates nnU-Net, SwinUNETR, and pretraining approaches (MAE, JEPA) on the ImageCHD test cohort under scanner/protocol/institution shifts. SwinUNETR outperforms nnU-Net (0.67 vs. 0.51 Dice on OOD data) despite lower in-distribution performance, with pretraining providing minimal gains. With just 11 labeled target-domain cases, SwinUNETR variants achieve >0.76 Dice, highlighting architecture's role in robustness. Results advocate cross-dataset testing for clinical segmentation evaluation.

out-of-distribution generalizationcongenital heart diseasemedical image segmentationcross-domain robustnessself-supervised pretraining

Beyond "ChatGPT Can Make Mistakes": Designing Interventions to Support Metacognitive Monitoring in AI-Assisted Work

arXiv cs.AI · Manuel A. D. Santos, Paul Thiesse, Steeven Villa, Daniela Fernandes · 2026-09-15

The study contributes a design space framework for metacognitive interventions in AI-assisted work, organized along dimensions of time, level, and source, along with empirical evidence on intervention effectiveness. Through expert elicitation and prior work synthesis, 30 interventions were identified and tested in a between-subjects experiment (N=917) using 12 planning-and-organizing problems with a baseline LLM assistant. Results show that reliability cards and contrasting replies reduced estimation error and overconfidence while improving aggregate confidence discrimination, though no task-performance improvements or within-item discrimination gains were observed. The findings demonstrate that metacognitive monitoring and task performance represent separable design targets in AI-assisted workflows.

metacognitive monitoringllm assistantestimation errorconfidence discriminationdesign space

Neuro-Symbolic Hierarchical Intention Anticipation in Human Behavior

arXiv cs.AI · Farnaz Soleimani, Abdelghani Chibani, Yacine Amirat, Ghazaleh Khodabandelou · 2026-09-15

The paper introduces a Hierarchical Planning Decoder (HPD) for goal inference and structured prediction in assistive autonomous systems, enabling anticipation of human intentions from partially observed multimodal episodes. The HPD, attached to a frozen neuro-symbolic recognition encoder, predicts actions, activities, and high-level intentions across four ontological levels, trained with soft neuro-symbolic regularization and decoded with hard reachability masks. Evaluated on a compositional benchmark of 15,002 multimodal episodes, the approach outperforms sequential baselines, with accuracy gains increasing from +1.7 points at step 1 to +7.3 points at step 3 (top-5). It achieves 96.8% trajectory validity under joint logic constraints, surpassing baselines (88.1%) and ground truth (73.9%), while eliminating reachability violations through soft logic terms and hard masks.

hierarchical planning decoderneuro-symbolic regularizationreachability masksmultimodal episodesontological levels

Repurposing Unified Topological Signatures for Graph Representation Learning

arXiv cs.AI · Sanyam Sanjay Jain, Anshika Krishnatray, Aditya Sharma, Vinti Agarwal · 2026-09-15

The paper introduces Unified Topological Signatures (UTS) to enhance Graph Neural Networks (GNNs) by capturing global topological features beyond the 1-Weisfeiler--Lehman (1-WL) limit. Two signatures—Graph_UTS (static) and Embedding_UTS (dynamic)—are integrated via three methods: UTS-Aug (feature augmentation), UTS-Reg (regularization), and UTS-Pool (topology-guided pooling). Theoretical analysis shows UTS strictly extends GNN expressivity, while experiments on three benchmarks demonstrate accuracy improvements of up to 5.8% with Graph_UTS and 1.9% with UTS-Reg, with UTS-Pool matching TOGL performance.

graph neural networkstopological signaturesweisfeiler-lehman testgraph representation learningpersistent homology

Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering

arXiv cs.AI · Kevin Mo, Nathan Mo, Richard Zhu · 2026-09-15

The study identifies and quantifies the fact-grounding gap in multi-hop question answering, revealing that nearly half of per-hop failures stem from extraction errors where retrieved passages contain but fail to yield the needed facts. Analyzing three standard benchmarks, the authors distinguish retrieval failures (missing passages) from extraction failures, demonstrating that the latter persist across datasets and resist retrieval-only interventions. Results show extraction failures constitute a distinct bottleneck requiring specialized solutions, overlooked by current evaluation metrics focused solely on retrieval performance.

multi-hop qafact-grounding gapretrieval failuresextraction failuresbenchmark analysis

Sparse MLLM Anchors, Dense Adaptation: Breaking the Self-Referential Loop in Wild Test-Time Adaptation

arXiv cs.AI · Zhenbin Wang, Lei Zhang, Lituan Wang, Yan Wang · 2026-09-15

MASA (Multimodal-LLM-Anchored Semantic Adaptation) addresses the self-referential loop in wild test-time adaptation (WTTA) by integrating structured semantic descriptions from a frozen multimodal large language model (MLLM). It queries the MLLM for diverse, reliability-ranked anchors to capture object families and nuisance factors, encodes these descriptions, and propagates them to neighboring samples via an online prototype memory. Descriptor-aware retrieval enables lightweight adaptation of normalization-affine parameters. Evaluated on WTTA ImageNet-C with ResNet and ViT backbones, MASA demonstrates robustness under limited-batch, mixed-domain, and imbalanced-label-shift settings.

wild test-time adaptationmultimodal large language modelnormalization-affine parametersonline prototype memorydescriptor-aware retrieval

Distributed JEPA: A Self-Supervised Framework for Energy Forecasting

arXiv cs.AI · Liana Toderean, Tudor Cioara, Vasilis Michalakopoulos, Efstathios Sarantinopoulos · 2026-09-15

We propose Distributed Joint Embedding Predictive Architecture (JEPA), a self-supervised framework for energy forecasting that learns generalizable representations from heterogeneous energy time-series. The method predicts latent representations of masked temporal segments while integrating temporal observations and contextual information in a shared embedding space, regularized by covariance and temporal variance objectives to prevent collapse. Evaluated on energy consumption and generation datasets under data degradation, JEPA achieved stable representations (cosine similarity ≈0.98; effective rank 185-235), matched Transformer performance on building energy data, outperformed baselines in 3/5 consumer clusters, and demonstrated superior robustness on 9/10 unseen PVs (R²=0.73-0.88 vs. <0.45).

joint embedding predictive architecturetemporal variance regularizationlatent-space predictionheterogeneous time-seriesdata degradation

SKIP: a Self-knowledge-guided Step-wise Preference Learning Framework for Concise Reasoning

arXiv cs.AI · Qinhong Lin, Yuhao Zhang, Yinglun Feng, Zhongliang Yang · 2026-09-15

We propose SKIP, a self-knowledge-guided step-wise preference learning framework for concise reasoning in LLMs, addressing Chain-of-Thought's computational overhead and inference latency. SKIP employs lightweight fine-tuning for output style adjustment, introduces a knowledge probing mechanism to guide answer generation at each reasoning step, and constructs preference data based on intermediate step correctness, optimized via Direct Preference Optimization (DPO). Experiments show SKIP improves reasoning compression while mitigating performance degradation, demonstrating strong generalization on out-of-distribution datasets. Ablation studies validate component parameter effectiveness.

chain-of-thoughtdirect preference optimizationknowledge probinglightweight fine-tuningout-of-distribution

ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation

arXiv cs.AI · Aurélien Pellet, Julien Perez, Marie Puren · 2026-09-15

The paper introduces ORDER (Optimal Routing for Dynamic Evidence Retrieval), a query-conditioned RAG framework that dynamically adapts indexing and retrieval strategies to heterogeneous queries. It clusters semantically similar questions, learns cluster-specific chunking, metadata filtering, and reranking configurations, and routes queries via nearest-centroid assignment. The method includes a supervised query router (QRe) for source selection and a Uniform Multi-source Sampler (UMS) for balanced retrieval. Evaluations on historical archives demonstrate consistent improvements over naive baselines and state-of-the-art RAG systems in expert domains.

retrieval-augmented generationquery-conditioned routingmetadata filteringsemantic clusteringmulti-source sampling

ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents

arXiv cs.AI · Cai Ke, Xin Liu, Han Zhang, Jiangyue Yan · 2026-09-15

ThinkFlow introduces a novel end-to-end latent memory framework for lifelong conversational agents, addressing limitations of explicit textual memory systems. The framework dynamically compresses conversational flows into probabilistic latent memory skills, representing user states as disentangled continuous vectors. It employs a test-time evolution paradigm combining teacher-guided latent alignment for initialization and self-supervised next-user-utterance prediction for continuous refinement. This approach enables label-free lifelong personalization without semantic interference. Extensive experiments on long-term conversation benchmarks demonstrate ThinkFlow's superior performance over existing memory systems in delivering personalized, contextually accurate responses across multi-session interactions.

latent memoryprobabilistic compressiontest-time evolutionself-supervised learningdisentangled representations

FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference

arXiv cs.AI · Qihu Xie, Ziwei Li, Yi Kang · 2026-09-15

FlexEE introduces a self-speculative early exiting framework for efficient LLM inference in offloading-aware deployments, addressing computational and memory constraints. The method employs layer-wise exit supervision for reliable intermediate predictions, self-speculative decoding with Top-K local vocabulary for low-cost exit decisions, and dynamic hidden state management to maintain KV-cache correctness. Evaluations on Llama2-7B and Llama3-8B show speedups of 1.27×/3.16× and 1.25×/2.83× under 0%/50% weight offloading, respectively, with minimal accuracy loss across generative and downstream tasks.

early exitingoffloading-aware inferenceself-speculative decodingkv-cachelayer-wise supervision

The Role of Implicit and Explicit Demographic Signals in Large Language Model-based Student Assessment

arXiv cs.AI · Donya Rooein, Luca Benedetto, Dirk Hovy · 2026-09-15

This study investigates the impact of explicit and implicit demographic signals on Large Language Models (LLMs) in student assessment tasks. Using controlled prompts, the authors examine six state-of-the-art LLMs across Automated Essay Scoring, Formative Feedback, and Metalinguistic Question Answering. Results reveal that LLMs adapt feedback readability based on explicit education levels but exhibit unpredictable biases in implicit conditions, such as assigning lower sentiment scores to responses from lower-education backgrounds. The findings demonstrate significant demographic sensitivity in LLMs, highlighting both potential benefits and risks in educational applications.

large language modelsautomated essay scoringformative feedbackmetalinguistic question answeringdemographic sensitivity

Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition

arXiv cs.AI · Yichi Zhang, Shenyue Wang, Jing Luo, Chunyang Yu · 2026-09-15

The Affect-Prototype-Conditioned Fusion (APCF) framework addresses open-vocabulary multimodal emotion recognition (OV-MER) under incomplete modal conditions by dynamically fusing multimodal features guided by arbitrary emotion semantics. APCF constructs an affect-prototype library to model multimodal contribution characteristics across diverse emotions, enabling conditional retrieval and feature aggregation based on available modalities. The refined representations are decoded by an LLM to generate open-vocabulary emotion labels. Evaluations on OV-MERD+ and MER-FG datasets show APCF significantly outperforms state-of-the-art baselines in handling incomplete multimodal inputs.

open-vocabularymultimodal fusionaffect-prototypeincomplete modalitiesemotion recognition

AntennaFlow: A Generative Flow Model for Offset Correction in Phaseless Antenna Testing

arXiv cs.AI · Yongzhi Li, Chongting Shen, Menglin Chen, Xun Jiang · 2026-09-15

AntennaFlow introduces a generative flow model for offset correction in phaseless antenna testing, addressing both phase acquisition costs and centering assumption violations. The framework comprises three stages: a contrastively learned encoder for offset-invariant embeddings, a deterministic flow-matching transport for mapping offset amplitudes to center-aligned fields, and the Simplified Extrapolation Technique for centered field extrapolation. Experiments demonstrate that AntennaFlow enables fast, phaseless, offset-vector-free near-field to far-field reconstruction from sparse amplitude-only measurements, outperforming existing baselines while maintaining physical consistency.

phaseless antenna testingnear-field to far-field transformationflow-matching transportcontrastive learningsimplified extrapolation technique

AeroLat: Channel-Aware Latent Space Semantic Communication for Decentralized UAV Swarms

arXiv cs.AI · Rajdeep Ghosh, Goparaju Venkata Seshachala Sree Vatsava, Sudip Misra · 2026-09-15

AeroLat introduces a channel-aware latent semantic communication framework for decentralized UAV swarms, addressing the collapse of broadcast states in homogeneous frozen models when processing discretized perceptual inputs. The method employs evidence injection and an explicit communication model incorporating bandwidth-limited serialization, additive noise, and information staleness to evaluate communication fidelity and swarm coordination. Multi-seed simulations demonstrate AeroLat's resilience to codec choice, faults, and increasing swarm size, reproducing latent-swarm anomalies while reducing false similarity by 97.5%.

latent semantic communicationuav swarmsevidence injectioncommunication fidelityfalse similarity

Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation

arXiv cs.AI · Shiqi Liu, Zeyu He, Letian Tao, Guojian Zhan · 2026-09-15

The authors propose $γ$OPD, a novel on-policy distillation (OPD) method that balances long-horizon supervision and optimization stability through discounted temporal credit assignment. Building on a unified temporal-credit view of token-level and sequence-level OPD, $γ$OPD introduces a reward-compatible bounded mixing mechanism to incorporate verifiable outcome feedback beyond teacher-dependent optimization. Experiments on mathematical and code reasoning tasks demonstrate consistent improvements over existing OPD methods in vanilla, size-mismatched, and multi-teacher distillation settings, while maintaining a horizon-independent variance bound.

on-policy distillationtemporal credit assignmentreverse-kl gradientreward-compatiblehorizon-independent variance

RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views

arXiv cs.AI · Yunxiang Zhang, Haiquan Wang, JiaWei Guo, Hanyang Xia · 2026-09-15

RepoAtlas introduces a training-free module for LLM-powered coding agents that maintains evolving multimodal repository views via a select--project--refresh loop over code graphs. The method combines issue evidence and agent state to select task-relevant regions under a fixed budget, projecting them into visual and textual representations. On SWE-bench Verified, RepoAtlas improves resolve rates by 2.4 points while reducing input tokens and model calls by 5.8% and 7.8% versus multimodal graph baselines, with consistent gains across three model families and scales.

coding agentsrepository contextmultimodal representationscode graphstask-relevant regions

When Confidence Signals Disagree: Local and Global Confidence in Autoregressive Language Models

arXiv cs.AI · Julio C. Amador Diaz Lopez · 2026-09-15

The study distinguishes between local confidence (greedy token probability) and global confidence (modal-answer frequency under sampling) in autoregressive language models, demonstrating they are not empirically interchangeable. Evaluating on MMLU and ARC Challenge, global confidence moderately correlates with correctness (r≈0.3), while local confidence shows negligible association. Disagreement between signals predicts sampling instability on ARC, with larger confidence gaps linked to higher answer entropy (p<0.01) and lower modal concentration, though this effect is weaker on MMLU (4% unstable questions). Findings emphasize treating confidence as an explicitly defined measurement for downstream applications.

autoregressive modelsconfidence calibrationsampling instabilityanswer entropymodal concentration

Causal Discovery via Transformed Low-Rank Quantile Surfaces

arXiv cs.AI · Ryo Kamimura, Thong Pham · 2026-09-15

The authors introduce Low-Rank Quantile Surfaces (LRQS), a bivariate causal model where a monotone transformation of the conditional quantile surface admits a low-rank functional decomposition in the causal direction. LRQS generalizes location-scale and post-nonlinear heteroscedastic noise models while accommodating multiple quantile bases for distributional shape variation. They prove generic identifiability, showing low-rank structure occurs only in the causal direction except for fine-tuned marginals. A causal score is derived via a nonparametric procedure alternating between rank-constrained quantile surface approximation and isotonic transformation estimation. Experiments on synthetic mechanisms and bivariate benchmarks demonstrate LRQS's effectiveness beyond location-scale assumptions, particularly for nonlinear distortions and higher-rank distributional shape variation.

low-rank quantile surfacescausal discoverymonotone transformationheteroscedastic noiseisotonic estimation

Repurposing Deep Limit Order Book Forecasting for Scenario-Conditioned Market Impact Modeling

arXiv cs.AI · Eljas Linna, Kestutis Baltakys, Derrick Manoharan, Alexandros Iosifidis · 2026-09-15

The authors propose a model-agnostic framework for repurposing pretrained Deep Limit Order Book forecasting models to quantify scenario-conditioned market impact without retraining. Their approach evaluates predictive distributions before and after injecting mechanically valid counterfactual messages, defining short-horizon model-implied market impact. Using a Transformer-based forecaster, the method achieved a Spearman correlation of 0.99 and 97.2% directional agreement with historical outcomes in non-neutral scenarios. Observation-level analysis demonstrated that estimated impacts captured incremental sequence-dependent variation beyond scenario identity and pre-event forecasts, validating the framework's effectiveness.

limit order bookcounterfactual messagesmodel-agnostic frameworkspearman correlationtransformer-based forecaster

Verbalizing Subliminal Learning Effects Using Text Optimization

arXiv cs.AI · Nathan Hu, Sanmi Koyejo, Christopher Potts · 2026-09-15

The paper introduces SALVE (Search-Aided Latent Verbalization), a method to detect and verbalize subliminal learning effects—traits transmitted from teacher models that are not explicitly encoded in distillation datasets. SALVE optimizes a soft prompt, queries the model for text verbalization, and employs beam search for reliability. It successfully recovers legible prompts naming teacher traits in standard settings, outperforming common text optimization methods, and detects effects in mixed datasets, activation-steered bias, and real preference data subsets. The work advances understanding of subliminal learning and offers proactive detection tools.

subliminal learningtext optimizationcontext distillationactivation steeringbeam search

QART: A Quantum-Classical Hybrid Architecture for Long-Horizon Reasoning -- Exploring a Conditional Path toward Quantum Scaling

arXiv cs.AI · Lehao Lin, Yuheng Cheng, Guolong Liu, Yao Li · 2026-09-15

QART introduces a quantum-classical hybrid architecture combining language models with quantum encoding, CIM-based QUBO optimization, and quantum decoding to address long-horizon reasoning. The method leverages semantic information from hidden representations or generated text, with proprietary encoding and optimization procedures. QART demonstrates conditional asymptotic reliability separation from autoregressive LLMs, maintaining bounded task-optimal-path recovery probability under specified conditions. Empirical evaluations on six benchmarks using DeepSeek V4 Flash, GLM-5.3, and GPT-5.5 xhigh show QART outperforming in 14 of 15 backbone-benchmark pairs, with relative gains up to 84.0% on SciCode. Quantum scaling laws are formulated as conditional hypotheses, requiring further validation for quantum advantage.

quantum-classical hybridqubo optimizationlong-horizon reasoningconditional reliabilitysemantic fidelity

Bridging Learned Visual Perception and Symbolic Belief-Space Planning

arXiv cs.AI · Guy Azran, Michael Navat, Sarah Keren · 2026-09-15

We introduce VLM-as-probabilistic-grounder, a novel paradigm bridging learned visual perception and symbolic belief-space planning by capturing Vision-Language Model (VLM) predicate grounding uncertainty as probability distributions over symbolic states. This enables robust planning under partial observability, addressing limitations of deterministic approaches like VLM-as-planner and VLM-as-grounder. Experiments in simulated household robot environments demonstrate improved robustness and task success rates compared to deterministic grounding methods, showcasing the effectiveness of leveraging foundation models for uncertainty-aware planning.

vision-language modelssymbolic planningbelief spacepartial observabilityprobabilistic grounding

VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal

arXiv cs.AI · Haonan Huang, Tianrui Qiu, Xianghao Zang, Yinan Du · 2026-09-15

We introduce VOR-Bench, a human perception-driven benchmark for video object removal (VOR) evaluation, addressing limitations in existing paradigms. VOR-Bench comprises three components: the VOR Dataset (VORD) with paired edited videos and graffiti masks spanning model-generated, tool-rendered, and camera-captured data; rMPAF, a realistic motion-capable paired-video acquisition framework combining image-based object removal and fine-tuned video generation models; and VOR-MDSM, a perception-driven VLM-based scoring model for mask-guided VOR. Extensive experiments show VOR-Bench achieves a correlation (ρ > 0.9) with human subjective assessments, aligning closely with human perception.

video object removalbenchmark datasetpaired-video acquisitionperception-driven evaluationmask-guided scoring

The Evolution of Coordination in a Collective Intelligence System: 25 Years of English Wikipedia and the Emergence of Generative AI

arXiv cs.AI · Neal Reeves, Maja Świeczkowska, Amy Rechkemmer, Elena Simperl · 2026-09-15

This study analyzes the evolution of coordination in English Wikipedia over 25 years, focusing on editing patterns across content, discussion, and governance namespaces. Using longitudinal data and Markov-based session metrics, it reveals declining participation in coordination spaces relative to content production, with increased specialization among editors and a shrinking core handling governance tasks. The investigation into generative AI's impact shows short-term changes but no fundamental alteration in coordination and participation trajectories. Findings highlight the growing specialization and centralization of Wikipedia's editorial workforce.

collective intelligencenamespacesmarkov-based metricsgenerative aicoordination

CoAdapt: An LLM-based Framework for Adaptive Collaborative Perception in IIoT Robotic Swarms

arXiv cs.AI · Houssam Hajj Hassan, Antonia Maria Masucci, Lynda Zitoune, Salah-Eddine Elayoubi · 2026-09-15

CoAdapt introduces an LLM-based framework for adaptive collaborative perception in IIoT robotic swarms, addressing dynamic industrial environments where robot positions, network bandwidth, and perception contributions vary. The framework employs a Large Language Model as a runtime fusion controller, determining robot participation and fusion algorithms based on spatial configurations and network states. It processes structured natural language descriptions derived from LiDAR point clouds, requiring no task-specific training and generalizing to unseen swarm topologies. Evaluated on the OPV2V benchmark across 25 scenarios, CoAdapt reduces communication costs by 38% while maintaining detection precision comparable to static baselines.

collaborative perceptionlidarllmrobotic swarmsfusion controller

RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models

arXiv cs.AI · Xun Liang, Honghui Yang, Weihang Pan, Ruisi Zhao · 2026-09-15

RegRet introduces a Large Multimodal Model (LMM) framework for region-level retrieval, addressing limitations in capturing effective regional representations. The method integrates a Region-Aware Encoder to balance regional features with global context and employs a multi-stage training pipeline involving localized captioning and regional contrastive learning. The REGMB benchmark, comprising 225k contrastive pairs across four tasks, is introduced to address data and evaluation diversity gaps. RegRet achieves significant improvements, outperforming baselines in zero-shot settings and showing over 20% average gains on REGMB and public benchmarks while maintaining global retrieval performance.

region-level retrievallarge multimodal modelsregion-aware encodercontrastive learningmultimodal retrieval

StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection

arXiv cs.AI · Zhenbin Wang, Lei Zhang, Lituan Wang, Wei Huang · 2026-09-15

StackTok introduces a training-free visual token selector for vision-language models (VLMs) that dynamically balances query relevance and visual coverage based on budget constraints. The method constructs a coverage reference from greedy sequences, adjusts support targets via query-vision affinity entropy, and employs an interleaved selection policy. Evaluated across five VLMs on ten benchmarks, StackTok achieves 95.26% of full-token performance on LLaVA-NeXT-7B using only 5.6% of visual tokens (160/2,880).

vision-language modelstoken selectionquery relevancevisual coveragetraining-free

What Breaks Local Watermarks? A Robustness Benchmark for Local Invisible Image Watermarking

arXiv cs.AI · Kai Yao, Bence Szilágyi, Sebestyén Kamp, Máté Poór · 2026-09-15

The study introduces the first systematic robustness benchmark for local invisible image watermarking, evaluating five methods (MaskWM, WAM, OmniGuard, TrustMark, PixelSeal) across 55 image transformations categorized into signal distortions, coordinate alignment changes, indirect local edits, and direct watermark edits. Results indicate all methods exhibit vulnerabilities, with MaskWM demonstrating superior payload recovery and localization despite lower clean-image quality. Synchronization enhances MaskWM's geometric robustness at further quality cost. Key findings highlight transformation-specific fragility: signal distortions are more tolerable than geometric misalignment or generative edits (e.g., inpainting), while payload recovery and localization exhibit correlated but distinct dependencies on transformation type.

local watermarkingpayload recoverygeometric misalignmentsynchronizationgenerative edits

Execution Flexibility in Automated Planning: A Comparative Evaluation of Deordering and Reordering Strategies

arXiv cs.AI · Md. Monjurul Islam, Sabah Binte Noor, Fazlul Hasan Siddiqui, Gahangir Hossain · 2026-09-15

The study demonstrates that block deordering-based approaches significantly outperform MaxSAT-based methods in enhancing plan-execution flexibility, despite the latter's theoretical guarantees of minimum reordering. It evaluates partial-order planning strategies, including producer-consumer-threat formalism and block substitution, across ordering, action handling, parameter handling, plan structure, concurrency, and complexity. Block deordering-based methods restructure causal dependencies through block-level grouping and subplan substitution, enabling anytime algorithms that always return valid results, unlike MaxSAT-based methods which fail on substantial portions of plans. Block deordering achieves the highest flex gain per unit of computation time, while MaxSAT-based encodings incur significant computational overhead.

partial-order planningblock deorderingmaxsatproducer-consumer-threatplan-execution flexibility

Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?

arXiv cs.AI · Younes Boufouss, Luc Pommeret, Thomas Gerald, Patrick Paroubek · 2026-09-15

The paper proposes an interpretable graph-based pipeline for Natural Language Inference (NLI) that avoids direct text processing by decomposing sentences into atomic propositions, converting them to ConceptNet triples via constrained decoding, and representing them as premise, hypothesis, and retrieved subgraphs. A fine-tuned 0.8B-parameter language model processes these graphs, achieving 89.7% accuracy on SNLI (1.9 points below text-based) and matching RoBERTa-large on ANLI R2/R3 (50%) but trailing by 16 points on R1. The 'price of interpretability' gap stems from representational limitations, while combined graph-text input boosts SNLI accuracy to 92.1%.

natural language inferenceatomic propositionsconceptnet triplesconstrained decodinginterpretability gap

SOTER: A Generative Time-Series Foundation Model for Wearable Human Physiological Signals

arXiv cs.AI · Fangke Chen, Sirry Chen, Wei Chen, Zhongyu Wei · 2026-09-15

SOTER introduces a generative foundation model for wearable physiological time series, addressing multichannel, irregularly sampled signals through cross-channel coupling, spectrum-guided expert specialization, and continuous-time latent evolution. The architecture combines a spatial feature-aware backbone, a PSD-guided mixture-of-experts layer, and a neural controlled differential equation decoder. Pre-trained on 226 billion time points, SOTER achieves state-of-the-art zero-shot forecasting (best RMSE on 4/6 datasets, best MAE on 5/6), classification (highest Macro-AUROC), and imputation (lowest error at 75% missingness), while demonstrating robustness to additive noise.

foundation modelwearable physiologymixture-of-expertsneural differential equationzero-shot forecasting

Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models

arXiv cs.AI · Zhenbin Wang, Lei Zhang, Lituan Wang, Wei Huang · 2026-09-15

Adaptive Relevance-guided Evidence Allocation (AREA) improves multimodal large language model (MLLM) performance by dynamically selecting visual and textual evidence during inference. AREA uses a probe token to assess relevance from fixed backbone layers, making three adaptive decisions: whether to intervene based on attention coverage and visual sink contamination, how much evidence to expose via relevance entropy, and when to refresh text using causal context-attention peaks. Evaluated on four KB-VQA and seven standard multimodal benchmarks with nine frozen MLLM checkpoints, AREA achieves state-of-the-art performance among training-free highlighting methods.

adaptive allocationmultimodal large language modelsvisual sink contaminationrelevance entropycausal context-attention

Available but Unclaimed: An Empirical Study of Human-AI Synergy

arXiv cs.AI · Robin Welsch, Michelle Rausch, Pascal Knierim, Thomas Kosch · 2026-09-15

This empirical study investigates human-AI synergy in reasoning tasks, demonstrating that complementary capabilities do not guarantee outperformance. Researchers conducted a between-subjects experiment (N=535) where participants solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, either unaided or assisted by GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each model answered items independently 100 times under matched elicitation. Results show that assisted-unaided accuracy difference increased with LLM competence, with deference varying across tasks and increasing within-task competence. Post-advice confidence was less discriminative than unaided confidence, and approximately half the LLM accuracy gain transferred to assisted accuracy, with model-specific variations.

human-ai synergymatrix reasoningsyllogismsletter-string analogiesdeference

Integrating the Analytic Hierarchy Process with Large Language Models for Transparent Multi-Criteria Decision-Making

arXiv cs.AI · Han Zhiguang, Farah Benamara, Pascale Zaraté · 2026-09-15

The study introduces an end-to-end method integrating the Analytic Hierarchy Process (AHP) with large language models (LLMs) to enhance transparent multi-criteria decision-making. By constructing a new AHP-based benchmark, the approach enables LLMs to perform the complete AHP workflow, addressing interpretability gaps in high-stakes domains. Experiments in legal and higher-education ranking scenarios demonstrate significant improvements in alignment with expert judgments compared to opaque LLM reasoning.

analytic hierarchy processlarge language modelsmulti-criteria decision-makinginterpretabilityexpert alignment

Coverage-Aware Virtual IMU Augmentation for Low-Resource Human Activity Recognition

arXiv cs.AI · Jiayuan Gao, Yingwei Zhang, Ziyao Tang, Yuejia Ma · 2026-09-15

Proposes a coverage-aware virtual IMU augmentation framework for low-resource human activity recognition (HAR) that strategically supplements real data with synthetic samples. The method selects diversity and scarcity anchors in a learned sensor embedding space, generates virtual IMU candidates via anchor dynamics prompts, and ranks them by anchor proximity and label consistency. Selected candidates are weighted by reliability during training. Experiments on public HAR benchmarks demonstrate consistent performance gains over baselines, with ablations validating the framework design.

imu augmentationhuman activity recognitionsensor embeddingvirtual samplescoverage-aware

Turn-level Multiscale Density Ratio Estimation for LLM Agents

arXiv cs.AI · Zishuo Zhao, Kai Chen, Ao Li, Yuan Liu · 2026-09-15

The paper proposes Turn-level Multiscale Density Ratio Estimation (tlm-DRE), a post-training alignment method for LLM agents that improves multi-turn task performance by assigning turn-specific weights and asymmetric token-level training based on positive-negative space gaps. Unlike traditional single-turn alignment methods (e.g., PPO, DPO, DIL, GRPO), tlm-DRE addresses multi-turn complexity through density ratio estimation across turns. Experiments on diverse agent benchmarks demonstrate competitive performance against baseline methods, with robustness in both in-domain and out-of-domain multi-turn reasoning tasks.

alignment methodsmulti-turn tasksdensity ratio estimationtoken-level trainingllm agents

TAME: Token Attribution and Masking for Emergent misalignment

arXiv cs.AI · Md Rayhanul Masud, Md Rizwan Parvez · 2026-09-15

TAME (Token Attribution and Masking for Emergent Misalignment) localizes emergent misalignment (EM) in fine-tuned language models by analyzing token-level attribution patterns. The framework computes token attribution scores via LoRA adapter forward passes, identifies high-attribution token patterns, and validates causality through attribution-guided loss masking. Experiments on Llama and Qwen show top 5% tokens account for 32% of attribution mass, with EM signal linked to unwarranted certainty rather than domain vocabulary. Masking high-attribution tokens reduces EM by 23x (Llama) and 36x (Qwen), with perplexity costs localized to flawed stylistic registers.

emergent misalignmenttoken attributionlora adapterloss maskingperplexity

Beyond Episodic AI: Cognitive Field Networks for Biologically Inspired Persistent Cognition

arXiv cs.AI · Byung Gyu Chae · 2026-09-15

The paper introduces Cognitive Field Networks (CFNs), a recurrent Transformer architecture where hidden states form a persistent cognitive field through memory-dressed collective dynamics. The CFN implements history-dependent inference via Φ_{n+1}=F_θ(X_{n+1},Φ_n), eliminating explicit memory operations. Experiments show organized recurrent dynamics with content-dependent persistence, where semantic continuation propagates states beyond training horizons, and periodic input re-exposure sustains nonzero stationary regimes. The framework distinguishes three processes: memory dressing, input-driven reorganization, and cross-cycle re-entry, providing a platform for studying persistent cognition without separate memory systems.

cognitive field networksmemory-dressed dynamicsrecurrent transformerhistory-dependent inferencesemantic continuation

Seeing What Matters: Visual Cue Guided Video Planning for Generalizable Robot Navigation

arXiv cs.AI · Hojin Lee, Sizhe Lester Li, Maximilian Hilger, Susie Lu · 2026-09-15

CueNav introduces a video model-based navigation framework combining visual cue-guided video planning with an embodiment-specific Inverse-Dynamics Model (IDM). The method uses Bird's-Eye View (BEV) maps for global task context and egocentric observations with robot body visibility for embodiment context, guiding a video planner whose output is translated to actions via IDM. CueNav achieves 2x higher success in maze navigation and 70% success in narrow passages compared to baseline methods, while enabling zero-shot semantic-conditioned navigation and cross-platform deployment.

video planninginverse-dynamics modelbird's-eye viewegocentric observationrobot navigation

LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture

arXiv cs.AI · Deepesh Sonar · 2026-09-15

The paper introduces LSREP, a Longitudinal State-Replay Evaluation Protocol for assessing conversational memory systems through ordered replay, lifecycle schedules, and fidelity checks. It evaluates ICE v2, a local-first memory middleware with typed stores and dynamic context budgets, on 1,985 turns and 219 probes across 52 checkpoints. Results show ICE v2 matches vector-RAG quality with 32% fewer fragments but 6.6% more prompt tokens, while failing catastrophically on dense datasets and underperforming in public diagnostics (-22.0 to -26.5 points versus vector-RAG). Fidelity audits reveal unexercised mechanisms and procedural retrieval defects.

conversational memoryevaluation protocollocal-first architecturedynamic context budgetsretrieval fusion

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

arXiv cs.AI · Haoyu Guo, Yuan Feng, Junlin Lv, Mingjun Xiao · 2026-09-15

VideoMM introduces an adaptive macro-micro inference framework for efficient video understanding in Multimodal Large Language Models (MLLMs), addressing the inefficiency of visual token explosion. The method decouples semantic filtering from detailed reasoning by first selecting relevant regions using a cost-effective Macro Proxy derived from downscaled frames, then projecting these regions onto high-fidelity Micro Tokens for detailed analysis. This approach achieves a 6.13× speedup and a 7.4% accuracy gain over full-context baselines on LongVideoBench, while accelerating inference by 2.73× compared to leading methods, establishing a scalable paradigm for long-video understanding.

multimodal large language modelsmacro proxymicro tokenssemantic filteringlongvideobench

Continuous-Time Machine Learning: A Unified Mathematical Perspective

arXiv cs.AI · Waleed Razzaq, Yun-Sheng Zhao, Yun-Bo Zhao · 2026-09-15

This survey provides a unified mathematical framework for continuous-time machine learning by organizing major branches through a taxonomy based on vector-field parameterization, stochasticity, memory mechanisms, and discretization. It compares training algorithms, optimization strategies, and failure modes across families, supported by theoretical computational complexity analysis and architecture-controlled benchmarks. The study reviews software ecosystems and identifies open challenges in approximation theory, training stability, hardware-efficient implementations, benchmarking, foundation models, and scientific machine learning, proposing a future research agenda.

continuous-timevector-fieldstochasticitydiscretizationbenchmarking

World Models for Embodied Intelligence: From Plausible to Controllable to Actionable

arXiv cs.AI · Nanjie Yao, Hao Wang, Chong Cheng, Zhikang Chen · 2026-09-15

The paper proposes a hierarchical framework for evaluating world models in embodied intelligence, distinguishing three capability levels: Plausible (preserving task-relevant structure), Controllable (predicting intervention effects), and Actionable (improving measurable planning or learning outcomes). It introduces a 3 x 4 matrix intersecting geometry, physics, and action grounding with improvement loops (data, rewards, policies, models). Surveying manipulation, navigation, and autonomous driving domains, the analysis identifies challenges in long-horizon consistency, uncertainty calibration, and cross-embodiment transfer, advocating for evaluation shifts from visual plausibility to closed-loop behavioral gains.

world modelsembodied intelligenceintervention predictionclosed-loop behavioruncertainty calibration

Weave: Learning Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions

arXiv cs.AI · Liu Cao, Xingze Wu, Jingzhi Cui, Botian Xu · 2026-09-15

Weave introduces a unified framework for learning whole-body dexterous humanoid-object interaction from human demonstrations, addressing coordination of locomotion, balance, and hand contact. The method employs contact-aware retargeting and approach-motion completion to convert human demonstrations into robot-executable references, followed by a contact- and geometry-aware policy controlling 29 body and 12 finger joints. Evaluations on nine objects show 92.5% success on trained interactions and 65.0% zero-shot generalization to unseen sequences, supported by a dataset of ~9,000 physically executed rollouts (~23 hours) with contact annotations.

humanoid-object interactioncontact-aware retargetingwhole-body controldexterous manipulationmotion completion

little m: An AI Agent for Industrial Process Optimization

arXiv cs.AI · Yongchao Ye, Xinyu He, Dutliff Boshoff, Way Kuo · 2026-09-15

The paper introduces little m, an AI agent for industrial process control optimization that bridges unstructured natural language and spatial diagrams with mathematical syntax. The framework combines domain-specific knowledge with LLM-driven interaction to formulate optimization models, addressing limitations of general-purpose LLMs in handling continuous multi-physics dynamics. Evaluated on the novel Industrial Process Control Benchmark (IPC-Bench), a multimodal dataset of 50 scenarios, little m outperforms state-of-the-art LLMs in generating semantically correct models through automated and human assessments of formulation quality.

industrial process optimizationmultimodal reasoningmathematical modelingdomain-specific knowledgeprocess control

AI for Games in the Foundation Model Era

arXiv cs.AI · Meng Luo, Yanlin Li, Hao Li, Hongzhan Lin · 2026-09-15

The article surveys AI applications in games during the foundation model era, categorizing them into six roles: playing/modeling, design, development, runtime adaptation, and evaluation. It analyzes each role's structural dependencies, learned components, transferable capabilities, and evidentiary support, noting cross-role synergies (e.g., trajectories training world models) but persistent setting-specific constraints (e.g., engine interfaces). While standardized evaluation exists for bounded game-playing tasks, challenges remain in learned world persistence, player modeling validation, and automated testing. The central tension lies in balancing capability transfer with context-specific effectiveness verification.

foundation modelsgame-world modelsruntime adaptationplayer modelingtransfer learning

ANIMASK: What the Model Contributes to Role Play in Simulated Story Worlds

arXiv cs.AI · Xiucheng Zhang, Zhuoning Xu, Hanjun Luo, Yankai Chen · 2026-09-15

ANIMASK introduces a simulation framework to analyze how language models contribute to character behavior in story worlds, isolating persona fidelity from model defaults. The method freezes stories into interactive environments, replays them from decision points, and compares actions with persona-removed model outputs. Results from 40 stories, 6 models, and 3,846 decisions show personas remain obeyed, but replays diverge toward flatter narratives; model defaults align with personas in 75% of cases, with models exhibiting caution where personas would act aggressively.

language modelspersona fidelitystory simulationdecision pointsmodel defaults

Right Direction, Wrong Step: Geometric Analysis of Finite-Step Failure in Looped Transformers

arXiv cs.AI · Zhihao Guo, Zonghan Wu, Haizhou Du, Huan Huo · 2026-09-15

The paper analyzes finite-step failures in Looped Transformers, where locally improving update directions produce harmful full updates due to mismatched step scales. Using reference utility measurements and a pathwise curvature decomposition, the authors characterize how initial progress is lost and predict full-step gains via local quadratic modeling. Experiments on mathematical and commonsense tasks show that a fixed quarter step recovers positive gains in 72.2--83.2% of selected failures, revealing curvature-driven approximation errors as the mechanism behind lost progress.

looped transformersreference utilitypathwise curvaturefinite-step failurequadratic modeling

ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training

arXiv cs.AI · Zhihao Zhang, Mingqi Wu, Qiaole Dong, Enyu Zhou · 2026-09-15

ReDraft (Reference-Driven Revision and Fine-Tuning) introduces a method for continual post-training of vision-language models that balances new task acquisition with pre-training capability retention. Unlike SFT or self-distillation, ReDraft uses expert responses as references to guide the model's revisions of its own incorrect rollouts, retaining only verifier-approved revisions for fine-tuning. Evaluated on Counting, Clock Reading, and Jigsaw tasks with Qwen2.5-VL-3B/7B, ReDraft achieves 56.9 target task accuracy (vs. SFT's 52.9) while reducing prior-task forgetting by 11.3× (1.5 vs. 16.6 points loss). Analyses show revised targets align better with the base model's distribution and induce more compact updates.

continual learningmultimodal modelsself-distillationverificationfine-tuning

EchoPath: Execution-Level Replayable Memory for GUI Agents

arXiv cs.AI · Yao Zhao, Aditya Shanmugham, Swastik Roy, Yanxun Xu · 2026-09-15

EchoPath introduces a model-agnostic framework for converting validated GUI trajectories into standardized, parameter-controlled callable memories, enabling deterministic replay of GUI-based tasks. The system employs an image-based target-reaiming algorithm that matches stored GUI targets against the current screen and corrects operation coordinates before execution. EchoPath rebinds modifiable inputs and rejects ambiguous steps, facilitating efficient task execution. Experiments demonstrate significant reductions in median token cost (over 90%) and execution time (approximately 60%), validating the utility of GUI memory as a controllable asset for recurrent enterprise tasks.

gui trajectoriestarget-reaiming algorithmcallable memoriesexecution-level replayenterprise tasks

AURA: Agentic Diagnosis and Refinement for Production Recommender Systems at Scale

arXiv cs.AI · SungGeun Kim, Abhinav Narain, Daniel Nemirovsky · 2026-09-15

We introduce AURA (Agentic Understanding and Refinement of recommender Algorithms), an end-to-end agentic system for diagnosing and improving production recommender systems at scale. AURA employs specialized agents to analyze production engagement logs from thousands to millions of sessions, identifying failure patterns and generating code-level refinements based on the recommender's codebase, data, and training pipeline. Initial tests on two large consumer platforms at a major media-streaming company demonstrate AURA's capability to propose actionable improvements. The architecture is domain-agnostic, with configurations enabling transferability to e-commerce and online-retail recommendation systems.

agentic systemrecommender algorithmsproduction engagement logscode-level refinementsdomain-agnostic architecture

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

arXiv cs.AI · Yuqi Wang, Fengyuan Liu, Haochen Luo, Zhiqi Yu · 2026-09-15

The paper introduces RoleBreak, a benchmark for evaluating long-horizon role-playing robustness in spoken dialogue systems. RoleBreak comprises 310 roles (character-based and user-centered), 6,688 human-verified dialogue turns, and 11,743 fine-grained criteria, including 1,856 turns with expressive emotion targets. It assesses role consistency, interaction quality, safety, and vocal affect over extended conversations. Nine system configurations (full-duplex, omni-modal, cascaded ASR--LLM--TTS) were evaluated, revealing four key findings: (1) systems excel at semantic role adherence but lag in vocal emotion; (2) semantic robustness degrades after ~10 turns; (3) scaling LLMs improves semantic robustness but not vocal emotion; (4) user vocal emotion influences role-playing behavior independently of linguistic content.

spoken dialogue systemsrole-playing robustnesslong-horizon evaluationvocal emotionsemantic adherence

Structure Across Voices: Comparing acoustic-event type accumulation and sequence dependence across four vocal repertoires using frozen audio encoders

arXiv cs.AI · Mudit Sinha, Sanika Chavan · 2026-09-15

The study compares acoustic-event type accumulation and sequence dependence across four vocal repertoires—sperm whale codas, human speech phones, Bengalese finch syllables, and common marmoset calls—using frozen audio encoders while controlling for event count and local sequence opportunity. Results show sperm whale codas exhibit the fastest type accumulation, while Bengalese finch syllables demonstrate the strongest immediate dependence and repeated-subsequence recurrence. Physically interpretable acoustics and continuous analyses reveal complementary patterns, with null models preserving source and position effects. Repertoire differences depend on measured acoustic properties and temporal scales rather than forming a single hierarchy.

acoustic-eventfrozen encoderssequence dependencevocal repertoirestemporal organization

Large Language Models in the Loop: A Stability- and Network-Aware Survey in Networked Control, Cyber-Physical, and Multi-Agent Systems

arXiv cs.AI · Haiping Du, Linping Chan · 2026-09-15

The survey analyzes integration challenges when incorporating large language models (LLMs) into networked control systems (NCSs), cyber-physical systems (CPSs), and multi-agent networks (CNSs), proposing a supervisory control framework to reconcile LLM stochasticity with stability requirements. It formalizes LLM behaviors—latency as delay, API failures as packet loss, tokenization as quantization, and hallucinations as disturbances—mapping them to classical control problems. Current approaches show diminishing safety assurances despite increasing model capabilities, with a noted absence of formal stability proofs. Future work must address this gap to enable reliable LLM deployment in physical systems.

networked control systemssupervisory controlstability proofsstochastic dynamicsformal verification

A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data

arXiv cs.AI · Yinong Wang, Jianwen Chen, Zhou Chen, Shuwen Kuang · 2026-09-15

BrainVLM is a vision-language foundation model for non-invasive presurgical diagnosis of 12 WHO 2021 brain tumor types from multimodal MRI data, integrating uncertainty quantification and radiology report generation. The model was trained on 40,043 individuals' multimodal data (MRI scans, demographics, reports) and validated on 5,211 pathologically confirmed cases across 12 hospitals. In clinical evaluations, BrainVLM improved diagnostic accuracy in a blinded multi-reader study (248 cases, 12 neuroradiologists) and matched radiologists in a prospective trial (1,009 patients), while also demonstrating utility in preoperative molecular subgroup prediction for gliomas (632 patients).

vision-language modelmultimodal mriuncertainty quantificationbrain tumor classificationmolecular subgroup prediction

A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance

arXiv cs.AI · Kimberly Le Truong, Nari Johnson, Anna Kawakami, Hoda Heidari · 2026-09-15

The paper introduces a framework for generating context-specific LLM benchmarks by combining expert guidance with synthetic data generation, addressing trade-offs between validity and scalability in existing methods. It proposes a schema to elicit task goals, scope, and context, guiding synthetic data generation while ensuring measurement validity through four criteria: coverage, diversity, content realism, and stylistic realism. Quantitative evaluations and a case study demonstrate improved benchmark quality and preserved validity compared to existing methods, with analysis of schema information prioritization under resource constraints.

large language modelsynthetic data generationbenchmark validityexpert guidancemeasurement criteria

Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

arXiv cs.AI · Keqing Zhang, Jingyu Chen, Yufan Liu, Yongqiang Zhu · 2026-09-15

The study introduces the Prior-Environment-Cognition (PEC) framework to quantify and align latent value systems in Large Language Models (LLMs), addressing their paradoxical behavioral duality of 'swing' and 'rigidity'. Analyzing 106 LLMs (150,000 queries/model) and 95,000 human survey profiles, the authors empirically demonstrate that LLMs exhibit a concentrated, idealized value core distinct from human diversity. The PEC framework models value expression as a function of parameter weights (Prior), prompt contexts (Environment), and reasoning processes (Cognition), enabling adaptive 'Alignment Prescription' interventions ranging from zero-cost prompts to targeted parameter updates. Empirical validation shows this method achieves precise value steering without degrading general capabilities.

value alignmentprior-environment-cognitionllm behavioradaptive interventionparameter weights

ProxiDex: Learning Dynamics-Guided Proximity Policy for Dexterous Manipulation

arXiv cs.AI · Yushan Bai, Boyu Zheng, Zhiyang Mao, Hongzheng Sun · 2026-09-15

ProxiDex introduces a dynamics-guided proximity policy framework for dexterous manipulation, addressing partial observability and contact uncertainty by modeling hand-object proximity as an interaction state. The method reconstructs interaction point clouds, converts geometric distances into hardware-agnostic proximity cues, and learns action-conditioned proximity dynamics via coupled forward-inverse prediction. It adaptively reweights proximity tokens across manipulation phases and uses dynamics-consistency supervision to stabilize policy inference under unreliable visual feedback. Experiments in simulation and real-world settings show improved success rates and robustness over baselines in standard, unseen object, and perturbation scenarios.

dexterous manipulationproximity policydynamics-guidedhardware-agnosticcontact uncertainty

Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models

arXiv cs.AI · Kuluhan Binici, Cihan Acar, Shivam Aggarwal, Siying Liu · 2026-09-15

The authors propose an adaptive step schedule controller for text-to-image diffusion models that dynamically adjusts denoising steps based on input prompt complexity, eliminating fixed-step inefficiencies without retraining. The method employs a mixture of step schedules with varying sizes, switching between them by evaluating error term discrepancies at each timestep. Experiments on COCO and DiffusionDB demonstrate maintained visual fidelity while reducing inference time, providing a computationally efficient alternative to fixed-step approaches in Stable Diffusion and similar architectures.

adaptive step scheduletext-to-image diffusiondenoising stepserror term discrepancyvisual fidelity

Vision And Text Transformer For Predicting Answerability On Visual Question Answering

arXiv cs.AI · Tung Le, Huy Tien Nguyen, Le Minh Nguyen · 2026-09-15

We propose VT-Transformer, a novel architecture for predicting answerability scores in Visual Question Answering (VQA) formulated as a regression task. The model leverages Transformer architecture to jointly exploit visual and textual features from multi-modal data, addressing limitations of prior binary classification approaches. Experiments on the VizWiz 2020 dataset demonstrate VT-Transformer's effectiveness and robustness compared to competitive baselines, advancing the state-of-the-art in answerability prediction for VQA systems.

visual question answeringtransformer architectureanswerability predictionmulti-modal dataregression task

Query-Aware Source-Risk Triage for Retrieval-Augmented Generation

arXiv cs.AI · Kainan Zhou, Gangzhen Qian, Chuhong Xu, Lu Yi · 2026-09-15

The paper proposes a query-aware triage method for retrieval-augmented generation (RAG) pipelines to address source-query relationship omission. The approach combines a four-dimension page score, rank-discounted family aggregation, intent-preserving query mutations, and a family-held-out router to classify retrieved pages as pass, contextualize, exclude, or review. Evaluated on a 20,000-row synthetic scenario and 200 real URLs, the method demonstrates why page-level frequency cannot replace family-level exposure and quantifies calibration shifts under scenario activation. Limitations include unmeasured annotation reliability and synthetic rankings lacking real retrieval dynamics.

retrieval-augmented generationquery-aware triagerank-discounted aggregationintent-preserving mutationsfamily-held-out router

The MAL Simulator: Cyber Operations Simulation based on Attack & Defense Graphs

arXiv cs.AI · Jakob Nyberg, Sandor Berglund, Andrei Buhaiu, Joakim Loxdal · 2026-09-15

The MAL Simulator introduces a cyber operation simulation framework based on the Meta Attack Language (MAL), enabling decision-driven attack and defense simulations without source-code modifications for domain adaptation. The simulator trains automated offensive and defensive agents using models grounded in data from the CRATE cyber range. Experiments demonstrate the trained attacker policy outperforms baseline search methods in target reachability, while the defender reduces costs under noisy alerts; however, defender performance degrades against RL attackers. The publicly available tool supports integration with ML frameworks.

cyber operation simulationmeta attack languageautomated agentscyber range cratereinforcement learning

QueryFormer: Winning Solution for KDD Cup 2026 Tencent UniRec Challenge

arXiv cs.AI · Yuanzhe Zhou, Zhaoyang Zeng · 2026-09-15

QueryFormer introduces a stackable unified field-sequence block for joint modeling of feature interactions and sequential behaviors in post-click conversion rate (pCVR) prediction, addressing limitations of projection-based query generation in prior unified architectures. The method employs cross-attention for query generation and shared-parameter packed attention for sequence processing, with optimized latency-aware scaling across view width (H), model dimensions, and compute. It achieved 1st place in the KDD Cup 2026 Tencent UniRec Challenge (test AUC: 0.83254, improved to 0.832713 post-competition), with H-scaling boosting validation AUC from 0.84540 to 0.84615. Ablations highlight query generation as the top performance contributor, while packed attention maintains efficient inference (1.89× latency for H=8 vs. H=1).

post-click conversion ratecross-attentionshared-parameter attentionlatency-aware scalingauc

A Cyber Range Evaluation of Autonomous Network Incident Response Agents

arXiv cs.AI · Jakob Nyberg, Teodor Sommestad, Andrei Buhaiu, Joakim Loxdal · 2026-09-15

The study evaluates autonomous network incident response agents in a cyber range designed for human operator training, comparing heuristic and reinforcement learning (RL) approaches. The cyber range features an emulated network with variable topology, red-team emulation, and simulated user agents, with defensive agents optimizing to block red-team access while minimizing availability costs. RL-based policies, trained via a cyber attack simulator, outperformed heuristic methods, though performance varied with adversary and user behavior. Results indicate RL agents achieved higher defensive efficiency under dynamic attack scenarios.

cyber rangereinforcement learningnetwork intrusion responsered-team emulationsiem platform

On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models

arXiv cs.AI · William L. Tong, Aryo Lotfi, Emmanuel Abbe, Kostas Vaggelakos · 2026-09-15

The study identifies gating mechanisms in State Space Models (SSMs) as a critical factor influencing their performance, showing that gating promotes in-weights memorization while delaying or preventing in-context learning solutions, despite no inherent architectural limitations. Through theoretical analysis and experiments, the authors demonstrate that gating enhances generalization to long sequences but hinders precise retrieval and in-context learning capabilities compared to Transformers. These findings provide insights into SSM training dynamics and suggest pathways for improving linear-time sequence models.

state space modelsgating mechanismin-context learningmemorizationtraining dynamics

What Does Layer-Importance Reveal About Transformers and State-Space Models?

arXiv cs.AI · Istabrak Abbes, Nizar Islah, Irina Rish, Sarath Chandar · 2026-09-15

The study investigates layer importance in Transformers and state-space models (SSMs) through two metrics: Necessity, measuring loss increase from layer bypassing, and Plasticity, quantifying task-specific weight updates during fine-tuning. Analyzing models up to 14B parameters, the authors find that Necessity and Plasticity anti-align across depth in Transformers but overlap in Mamba-style SSMs. This alignment predicts downstream adaptation behavior: Transformers exhibit increased catastrophic forgetting when updates concentrate in plastic layers, while SSMs show no such tier-dependent effect. The results highlight fundamental differences in layer dynamics between these sequence model families.

transformersstate-space modelslayer importancenecessityplasticity

QALPA: Property-guided diffusion modeling for efficient exploration of chemical spaces of flexible molecules

arXiv cs.AI · Michael Hanna, Julian Cremer, Zekiye Erarslan, Leonardo Medrano Sandonas · 2026-09-15

QALPA (Quantum-Aware Learning for Property-space Augmentation) is a property-guided generative framework combining an E(3)-equivariant diffusion model with active learning and quantum-mechanical (QM) methods to explore chemical spaces of flexible molecules. The method integrates generation with physics-based evaluation, improving sampling in sparse regions of chemical space. Results demonstrate accurate molecular generation across size ranges (QM7-X and Aquamarine datasets) and enhanced transferability for complex property manifolds. QALPA, coupled with EquiDTB, augments the alloQM dataset (6,253 conformers) by populating sparse regions defined by many-body dispersion energy and HOMO-LUMO gap, showcasing sustainable chemical space exploration.

diffusion modelingquantum-mechanical methodschemical space exploratione(3)-equivariantactive learning

AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research

arXiv cs.AI · Bernie Boscoe, Srinath Saikrishnan, Vikram Seenivasan, Jack Stark · 2026-09-15

The paper evaluates faithfulness in AquiLLM, an open-weight RAG-LLM system for scientific research, focusing on its grounding in retrieved context. Using a domain-expert case study in astronomy, the authors assess performance on retrieval-oriented and synthesis tasks. Results show high faithfulness for explicit retrieval queries but degradation when synthesis or ambiguity resolution is required, highlighting limitations of open-weight RAG systems despite their promise for scientific workflows.

retrieval-augmented generationopen-weight llmfaithfulness evaluationscientific workflowsdomain-expert assessment

Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening

arXiv cs.AI · Qiangju Chen, Yang Xiao · 2026-09-15

The study introduces a controlled audit method to evaluate presentation sensitivity in LLM-based resume screening, constructing occupation-grounded candidate profiles at fixed competence levels and rendering them into multiple resume variants. Using a deterministic validation gate to exclude evidence-altering variants, the authors test six instruction-tuned LLMs, finding a disconnect between screening validity and presentation stability. Llama-3.1-8B achieves the highest validity (0.781) but reverses 29.6% of decisions under presentation changes, while Mistral-7B-v0.3 shows lower validity (0.644) and a 41.4% flip rate. Chat formatting improves validity but fails to stabilize decisions, highlighting the need for evaluations to assess both competence identification and presentation robustness.

resume screeningpresentation sensitivityinstruction-tuned llmscompetence preservationvalidation gate

Beyond the Name: Demographic Leakage in De-Identified Résumés and Evaluation Artifacts in LLM Bias Audits

arXiv cs.AI · Qiangju Chen, Yang Xiao · 2026-09-15

The study demonstrates persistent ethnocultural leakage in de-identified résumés even after removing language fields, with target-group recovery reaching 1.000 under high-salience cues. Using 620 counterfactual résumés across nine open-weight models, the authors isolate unstructured prose across five ethnocultural conditions and three cue-salience tiers, finding model divergence only under faint cues (0.086-0.690). Evaluation design significantly impacts outcomes: disallowing ties yields a 0.39 selection-rate ratio, while permitting ties produces ≥94% ties, underscoring the need to separate demographic signals from protocol artifacts.

demographic leakagede-identificationllm bias auditscue-saliencecounterfactual résumés

Geospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines

arXiv cs.AI · Daniel Ebanks, Devika Jain · 2026-09-15

The study demonstrates that geospatial metadata significantly enhances cross-disciplinary dataset discoverability in research repositories. Analyzing Harvard Dataverse records reveals that datasets with incomplete metadata receive fewer citations and exhibit weaker interconnections, with only 0.3% including bounding boxes. By embedding datasets in a metadata knowledge graph, the authors show geospatial metadata doubles cross-disciplinary linkage likelihood compared to keyword-based connections (58.5% to 63.2%). A small language model fine-tuned on Dataverse data further validates the utility of geospatial metadata enrichment for interoperability.

geospatial metadataknowledge graphdataset interoperabilitymetadata enrichmentcross-disciplinary discovery

AI Policies: Help or Hindrance? A Software Developer's Perspective

arXiv cs.AI · Samuel Ferino, Rashina Hoda, John Grundy, Christoph Treude · 2026-09-15

The study examines the efficacy of AI policies in software organizations through qualitative analysis of 19 developer interviews, identifying both facilitative and obstructive impacts on developer workflows. It proposes developer-centric policy frameworks to enhance engagement and compliance. Findings indicate that current policies often fail to align with practical developer needs, risking disengagement and reduced policy effectiveness.

ai policiessoftware developersqualitative analysispolicy compliancellm risks

From Manual Construction to AI-Driven Scenario Emergence: Rethinking Catastrophe Risk Modeling

arXiv cs.AI · Hang Gao · 2026-09-15

The TAISE framework introduces AI-driven scenario emergence for catastrophe risk modeling, addressing limitations of manual construction methods prevalent since the 1990s. By repurposing AI weather forecasting models, it generates continuous global atmospheric fields through self-iterative generation, enabling emergent extreme weather sequences. A proof-of-concept experiment demonstrates a 10x reduction in computational costs while maintaining temporal continuity and cross-regional correlations absent in snapshot-based approaches. This innovation enables democratized catastrophe risk quantification and dynamic portfolio assessment for insurers, reinsurers, ILS fund managers, and public-sector risk managers.

catastrophe risk modelingai weather forecastingself-iterative generationatmospheric fieldscomputational cost reduction

Decoder Design Matters for ECG Delineation

arXiv cs.AI · Joseph Scharpf, William Han, Chaojing Duan, Michael A. Rosenberg · 2026-09-15

Proposes R-U-Net, a ResNet-18 encoder paired with a U-Net decoder for ECG delineation, demonstrating that decoder design significantly impacts performance more than semi-supervised learning methods. Evaluated on SemiSegECG, R-U-Net outperforms ResNet-18 + FCN baselines by 3.3-13.0 mIoU across 16 in-domain settings and achieves 82.6 mIoU (8.1 mIoU improvement) in cross-domain evaluation. Ablations confirm decoder choice drives gains, suggesting architectural exploration as a priority for ECG annotation tasks.

ecg delineationsemi-supervised learningu-netresnet-18miou

Skill-based Agentic Evaluation for Real-time Data Science Tasks

arXiv cs.AI · Aniruddha Tamhane, Raghavendra Addanki, Ayushi Aggarwal, Aditya Bansal · 2026-09-15

The paper introduces a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring. The method, termed ground-truth-as-code, encodes expected answers as executable reference functions that recompute answers directly from live data, ensuring consistency. A factoid-level judge decomposes responses and ground truth into atomic claims, scoring precision, recall, and accuracy irrespective of format. Validation on a production machine learning skill shows a 29% improvement in Matthews Correlation Coefficient and a 16% reduction in token consumption compared to natural-language ground-truth baselines.

ground-truth-as-codefactoid scoringmatthews correlation coefficientexecutable reference functionsformat-agnostic

Protocol-Preserving Context Trimming for Agentic Workflows: Benefits, Failure Regimes, and Budget Guardrails

arXiv cs.AI · Harish Gaggar · 2026-09-15

The study evaluates protocol-preserving context trimming strategies for agentic LLM workflows, comparing five approaches (recency-based, relevance-based, summarization, protocol-aware trimming, adaptive guardrails) across performance metrics. Protocol-aware trimming achieved 92.2% task success (vs. 66.6-77.3% for conventional methods), while adaptive guardrails reached 96.0% success with 56.0% token savings. Retained-context budgets below 25% increased failure odds 10.92× versus ≥50% budgets (p<0.001), with protocol-aware methods showing 5.24× higher success odds than conventional approaches under aggressive trimming. Critical context thresholds scaled with workflow complexity, demonstrating that protocol-state preservation outweighs token reduction for reliability.

agentic workflowscontext trimmingprotocol adherenceadaptive guardrailscascading failures

OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation

arXiv cs.AI · Chenhao Qiu, Dawei Li, Yechao Zhang, Lei Gong · 2026-09-15

OPD-Aha introduces a method for robust multimodal reasoning by reconstructing distillation targets from isolated visual preferences rather than relying on fragile teacher-student discrepancies. The approach leverages privileged on-policy distillation, where a teacher evaluates student trajectories using visual evidence, and isolates the teacher's corrective visual preference by comparing predictions with and without visual input. This reconstructed target suppresses erroneous continuations, prompting students to self-correct via reflection tokens (e.g., 'wait', 'actually') and shift reliance from flawed text to visual evidence. The method achieves consistent improvements across fine-grained perception and complex multimodal reasoning benchmarks.

on-policy distillationmultimodal reasoningvisual reflectiondistillation targetreflection tokens

Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs

arXiv cs.AI · Kirill Skobelev, Eric Fithian, X. Y. Han · 2026-09-15

This work demonstrates that supervised fine-tuning (SFT) can resolve both mode collapse and over-dispersion in large language models (LLMs), with output diversity converging toward the target distribution as SFT data increases. The authors derive a bias-variance decomposition of the expected gap between model and target collision probabilities, proving SFT is not inherently biased toward either extreme. They establish an upper bound on the diversity gap via Kullback-Leibler divergence, showing optimal cross-entropy models cannot exhibit arbitrarily miscalibrated diversity. Experiments on synthetic languages, human surveys, and CodeNet confirm that increased target data moves model diversity toward human-level benchmarks, validating the theoretical predictions.

mode collapsesupervised fine-tuningkullback-leibler divergencebias-variance decompositioncross-entropy

Early-Bird Decoding: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling

arXiv cs.AI · Lixuan Wei, Wei Zhou, Jianwen Wu, Yipeng Shen · 2026-09-15

The paper introduces Early-Bird (EB) Decode, a framework accelerating diffusion large language models (dLLMs) via adaptive block-wise parallel decoding. Key innovations include a learnable network for grouping tokens with similar uncertainty into variable-length blocks and a position-aware sampler enabling parallel unmasking with fewer steps. EB-Decode requires no modifications to pretrained weights and achieves 3.53-18.76× higher throughput than vanilla decoding and up to 1.58× over Fast-dLLM, maintaining comparable accuracy across three models and four benchmarks.

diffusion llmsparallel decodingvariable-length blocksuncertainty clusteringkv caching

Interpreting and Steering LLM Agents for Social Simulations

arXiv cs.AI · Jiayue Gaveal Fan, Arul Murugan, Shreyas Krishnan, Abhishek Nagaraj · 2026-09-14

The study introduces interpretability and steerability methods for LLM-based social simulations, comparing prompt-based manipulation, SAE-derived feature steering, and probe-based direction steering. These techniques are evaluated on four classic economic and creative tasks operationalizing human preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation). Results indicate SAE- and probe-based methods often outperform basic prompting, with SAEs enabling human-readable feature decomposition and probes facilitating reliable behavioral steering, forming an effective pipeline for social scientific simulations.

llm agentssocial simulationsinterpretabilitysteerabilitysae-derived features

How Good Are Time-Series Foundation Models for Pedestrian Crowd Count Forecasting? A Cross-Dataset Comparative Study

arXiv cs.AI · Theivaprakasham Hari, Ziteng Li, Yanan Xin, Winnie Daamen · 2026-09-14

The study evaluates time-series foundation models for pedestrian crowd count forecasting across diverse datasets and forecasting paradigms. Seven univariate forecasting approaches were benchmarked, including Seasonal Naive, gradient-boosted trees (LightGBM, CatBoost), deep learning models (N-HiTS, PatchTST), and pretrained FMs (TimesFM, Chronos-2), on two datasets: SAIL2025 (3-minute resolution, limited history) and Melbourne pedestrian sensors (hourly, multi-year). Key findings include: Seasonal Naive excels with limited historical data, boosted trees perform well on low-volume sensors but degrade under event-driven shifts, and FMs outperform in seasonal, data-rich regimes with long-context configurations. Results emphasize selecting models based on data conditions and forecasting horizon.

time-series foundation modelspedestrian crowd countunivariate forecastingseasonal naivegradient-boosted trees

On the Expressive Power of Implicit Line-Graph Higher-Order Weisfeiler--Leman

arXiv cs.AI · Fan Yang · 2026-09-14

The paper introduces Implicit Line-Graph Weisfeiler-Leman (ILG-$k$-WL), a method for isomorphism testing that operates directly on graph edges without explicitly constructing line graphs. ILG-$k$-WL leverages line-graph relations derived from endpoint incidence and is analyzed across different dimensions ($k=1,2,3$). For $k=3$, ILG-$3$-WL demonstrates backward containment, proving that $L(G)\equiv_{3\text{-WL}}L(H)\Rightarrow G\equiv_{3\text{-WL}}H$, and is shown to be strictly more expressive than $3$-WL. The method successfully separates all three substructure-counting witness pairs, 105 pairs in SR25, and 359 of 400 BREC pairs. An untrained dense ILG-$3$-GNN replicates these pairwise verdicts.

isomorphism testingweisfeiler-lemanline graphsgraph edgesbackward containment

Balancing Trial and Reorder: A Hybrid Sequential Transformer-GBDT Ranker for On-Demand Delivery

arXiv cs.AI · Marcel Kurovski, Attila Nagy, Steffen Klempau, Aleksandr Fedintsev · 2026-09-14

The Universal Venue Ranker (UVR) is a hybrid sequential transformer-GBDT system for personalized store ranking in on-demand delivery platforms, addressing the trade-off between surfacing new stores and maintaining reorder quality. UVR combines a bidirectional transformer encoder for sequential user modeling with a GBDT ranker integrating contextual, user, and store features, trained across all stores and domains while enforcing local delivery constraints. Offline evaluation shows trial MRR improvements of +12% to +30% but mixed reorder MRR results, while online A/B tests demonstrate +5.5% Merchant Trial Rate and +0.16% Global CVR for V1, with subsequent versions further improving trial rates and unifying restaurant and retail ranking.

transformergbdtrankingpersonalizationa/b testing

UDAV: Uncertainty-Driven Adaptive VLM Waypoint Planner

arXiv cs.AI · Ghazal Farhani, Shabnam Shabani · 2026-09-14

UDAV introduces an uncertainty-driven adaptive planner for UAV-guided UGV navigation using vision-language models (VLMs). The method generates stochastic trajectory predictions, selects a medoid as the nominal route, and triggers reconsideration when predictive uncertainty exceeds a threshold. Evaluated on 400 trajectory queries, UDAV reduces mean average displacement error (ADE) by 25.1% (from 147.4 to 110.4 pixels) compared to deterministic planning, while ensuring valid trajectories for all queries. It also achieves lower 90th- and 95th-percentile errors (199.0 and 290.8 pixels) than baselines, demonstrating improved reliability and error mitigation.

vision-language modelsuncertainty estimationtrajectory planninguav navigationstochastic predictions

How Humans and LLMs Read Gender into Gender-Neutral Physical Descriptions

arXiv cs.AI · Yingjia Wan, Lin Lin, Elisa Kreiss · 2026-09-14

The study introduces GAPA (Gender Associations of Physical Attributes), a dataset of 316 physical attributes with 14,706 gender-association ratings from 304 US annotators, to examine whether gender-neutral descriptions retain implicit gender associations. Evaluating 16 LLMs against human ratings reveals systematic biases: compressed distributions, weaker alignment for male associations, and asymmetric abstention for non-binary categories. Results demonstrate that physical descriptions carry structured gender associations, challenging assumptions of neutrality and highlighting model-human misalignment in identity communication.

gender associationsphysical attributesllm evaluationbias detectionhuman-ai alignment

Auto-HSI: Personalized human control of a robot swarm on demand by using LLMs for online automatic code generation

arXiv cs.AI · Alessandro Nazzari, Nathan Cerisara, Dorian Tonnis, Raina Zakir · 2026-09-14

Auto-HSI introduces a method for generating personalized human-swarm interaction interfaces via natural language descriptions and gesture demonstrations, enabling untrained operators to control robot swarms. The system automatically generates code for personalized state machines that translate gestures into swarm behaviors, supporting teleoperation of motion, formation shape, and deformation. Experiments with 50 simulated robots in a physics-based simulator demonstrated successful task execution (goal scoring, maze traversal, simultaneous dual-goal scoring) under nominal and noisy conditions, including live interface updates. Real-robot validation was also conducted.

human-swarm interactiongesture-based controlstate machinesrobot teleoperationcode generation

From Momentary Emotion Inference to Sustained Emotion Support: Evaluating a Companion Agent in a Longitudinal Study

arXiv cs.AI · Kexin Quan, Zijian Ding, Jiaye Yong, Qinshi Zhang · 2026-09-14

The study introduces PAIR, a theory-based emotion-regulation companion agent, to investigate sustained emotional support capabilities over longitudinal interactions. Deployed with 19 participants across 1,093 sessions over 14 days, the system paired emotion estimates with self-reports pre- and post-guidance, analyzing logs and interviews. Results show emotion estimates aligned more closely with self-reported valence and dominance than arousal, with guided conversations increasing valence and inducing state-dependent arousal changes. Participants reported feeling understood through contextual exploration and emotional acknowledgment, with perceived helpfulness significantly increasing over time. Findings highlight memory updates and retained corrections as key to cross-session personalization, informing adaptive emotional support tools that preserve user control.

emotion-regulationvalencearousallongitudinal studypersonalization

Breaking the 1.58-bit Barrier for Ternary LLMs

arXiv cs.AI · Evangelos Georganas, Alexander Heinecke, Pradeep Dubey · 2026-09-14

BITCOS, a distribution-adaptive storage layout for ternary LLMs, reduces weight storage costs below the conventional 1.585-bit barrier by exploiting zero density in ternary weights. The method combines a dense presence bitmap with a compacted sign vector, achieving $2 - z$ bits per weight for zero density $z$. BITCOS outperforms five-trit packing in 26 of 29 tested models, reaching 1.485 bits per weight on the sparsest model. Optimized unpacking sequences for AVX-512, AVX2, and Intel Xe2 GPUs yield up to 1.28× speedup in ternary matrix-vector multiplication. End-to-end LLM inference improves decode throughput by up to 1.18× on CPUs and 1.27× on GPUs across five platforms.

ternary llmsbitcoszero densitymatrix-vector multiplicationdecode throughput

Cross-Anatomy Transfer Versus Sparse Interpolation in Digital-Twin-Oriented Aortic Fluid-Structure Interaction Surrogates

arXiv cs.AI · Ali Nourbakhsh, Mohammad Reza Niroomand, Erfan Nourbakhsh · 2026-09-14

The study evaluates surrogate models for aortic fluid-structure interaction (FSI) by comparing cross-anatomy transfer with sparse interpolation. Using four de-identified aortic models, a geometry-only LightGBM prior was trained on three anatomies and zero-shot tested on the fourth, followed by sparse field-completion case studies. Zero-shot transfer performed poorly (OSI R² = 0.603 at 203 anchors), while same-anatomy interpolation methods—inverse-distance weighting (OSI R² = 0.829) and radial basis functions (OSI R² = 0.917)—outperformed cross-anatomy adaptation. Results suggest sparse within-anatomy labels suffice for field completion, with no added benefit from the cross-anatomy prior. The work frames this as a preliminary step toward measurement-linked digital twins, with future work needed on converged FSI and physics-informed learning.

fluid-structure interactionlightgbm priorsparse interpolationoscillatory shear indexdigital twin

FairLint-DL: An IDE-Native Tool for Fairness Debugging of Deep Learning Software

arXiv cs.AI · Archit Rathod, Saeid Tizpaz-Niari · 2026-09-14

FairLint-DL introduces an IDE-native fairness debugging tool for deep learning software, enabling pre-training bias detection directly on tabular datasets within Visual Studio Code. The tool trains a configurable deep neural network as a proxy model and applies information-theoretic Quantitative Individual Discrimination (QID) metrics, grounded in Shannon and min-entropy, to quantify the causal influence of protected attributes on predictions. It employs a two-phase gradient-guided search algorithm, a causal debugging pipeline for layer-specific bias localization, and dual explainability engines using SHAP and LIME for feature-level attribution. Evaluation on Adult Census Income, German Credit, and Bank Marketing datasets reveals significant fairness concerns, with 96.0% of Adult instances exceeding the 0.1-bit QID threshold, achieving results within 12 seconds on cached models.

quantitative individual discriminationshannon entropycausal debugginggradient-guided searchfeature-level attribution

Cognitive Admission Control: Risk-Conditioned Assurance for Consequential Actions in Agentic Distributed Systems

arXiv cs.AI · Jun He, Deying Yu · 2026-09-14

Cognitive Admission Control (CAC) introduces a risk-conditioned assurance framework for agentic distributed systems, formalizing evidence requirements for consequential actions through policy-defined assurance obligations. The method employs a deterministic evaluator to distinguish satisfied, violated, and unresolved obligations, generating evidence-acquisition requests for unresolved conditions. A TypeScript prototype was evaluated in 2,730 controlled trials, demonstrating that CAC prevented harmful effects while admitting 120 safe actions, contrasting with a live-policy baseline that admitted correlated-witness failures. Mechanism ablations isolated key behaviors, and 9,000 measurements validated dispatch path integrity with persistent replay protection.

cognitive admission controlagentic distributed systemsassurance obligationsevidence-acquisition requestsdispatch-time guards

Efficient One-to-Many Translation with Joint Multi-Stream Diffusion

arXiv cs.AI · Yiwen Guan, Jacob Whitehill · 2026-09-14

We propose a joint multi-stream diffusion framework for efficient one-to-many machine translation that achieves sublinear latency scaling with the number of target languages. The method employs a discrete diffusion process that refines all target languages in parallel, conditioned on a continuous semantic anchor rather than source tokens, enabling zero-shot transfer to unseen source languages without retraining. Experiments demonstrate comparable supervised translation quality to autoregressive baselines with a 2× speedup, while maintaining 75% of supervised quality on zero-shot sources and achieving 11.9% better zero-shot BLEU. The framework supports deployment as a single unified model, offering a practical alternative to multiple independent translation systems.

diffusion frameworksemantic anchorzero-shot transfersublinear latencymultilingual translation

BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents

arXiv cs.AI · Sadia Asif, Mohammad Mohammadi Amiri, Momin Abbas, Tejaswini Pedapati · 2026-09-14

The paper introduces BLINDSPOT, a benchmark for evaluating trajectory-level safety calibration in long-horizon tool-using LLM agents. It employs adaptive adversarial interactions, stateful tool execution, and execution-grounded adjudication across 22 attack families and 35 scenarios, generating 2,500 trajectories averaging 14.7 turns. Outcomes are categorized as Safe Completion, Correct Refusal, Unsafe Completion, Over-Refusal, or Indeterminate. Evaluations of 13 LLMs reveal significant variations in safety-utility calibration, with failures often emerging after initially safe steps, underscoring the need for trajectory-level safety assessment.

long-horizon agentssafety calibrationadversarial interactiontool executionrefusal behavior

Assurance Envelopes for Autonomous Coding Agents: Minimum-Cost Evidence for Software Change

arXiv cs.AI · Anjan Goswami · 2026-09-14

The paper introduces task-conditioned assurance envelopes, minimum-cost subsets of existing evidence (tests, proofs, static analyses) required to re-establish software properties during AI-driven code changes. Methodologically, it models evidence as a typed inference graph, using forward chaining from selected evidence to validate obligations, with optimization via CP-SAT. Evaluation on Rust, IronBlocks, and Pong artifacts (249 synthetic instances) shows task-dependent envelopes, median solve times <20 ms for 500-evidence graphs, and timeout cases due to derivation complexity rather than graph size. Open challenges include obligation discovery and agent integration.

assurance envelopestyped inference graphforward chainingcp-sat optimizersoftware change

CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine

arXiv cs.AI · Shuai Wang, Yize Zhao, Qingyu Chen · 2026-09-14

The paper introduces CLEAR, an agentic framework for cross-source evidence adjudication in medical large language models (LLMs) to address challenges from static parametric knowledge and unreliable retrieval-augmented generation (RAG). CLEAR generates candidate answers via three pathways—parametric knowledge, locally curated corpora, and dynamically retrieved evidence—then employs an aggregation verifier to assess agreement/conflict using provenance and source-quality metadata. An adjudication module resolves conflicts via override-guard and challenge-audit mechanisms, with unresolved cases triggering iterative search and re-evaluation. The approach aims to improve factual accuracy by dynamically integrating and validating multi-source evidence.

retrieval-augmented generationevidence adjudicationparametric knowledgeprovenance trackingmulti-source verification

Closing the Loop: Branch-and-Bound for Scalable Verification of Nonlinear Neural Feedback Systems

arXiv cs.AI · I. Samuel Akinwande, Mykel J. Kochenderfer, Clark Barrett · 2026-09-14

The paper introduces RAIL and CLIPPER, a framework combining branch-and-bound with polyhedral abstractions to improve scalability in verifying nonlinear neural feedback systems. RAIL provides polyhedral enclosures for LiRPA-style bound propagation, while CLIPPER jointly refines dynamics enclosures and splits controller activations via branch-and-bound, preserving symbolic correlations across time steps. This joint reasoning approach demonstrates significant improvements over state-of-the-art combinatorial and propagative solvers, addressing scalability limitations in autonomy applications with large networks and nonlinear dynamics.

neural verificationbranch-and-boundpolyhedral abstractionlirpanonlinear dynamics

Intelligent Interaction Techniques (IIxT) - Proposal

arXiv cs.AI · Brad A. Myers · 2026-09-14

The paper proposes Intelligent Interaction Techniques (IIxTs) as a novel paradigm for user interface design, enabling seamless multimodal interactions within a unified framework. Building upon traditional GUI interaction techniques (IxTs) established in the 1980s and adapted for smartphones, IIxTs aim to incorporate intelligence into fundamental UI components like menus, text input, and object selection. The author identifies key research challenges in developing both new IIxTs and the underlying infrastructure to support them, while highlighting significant security, privacy, and economic implications of this approach.

interaction techniquesmultimodal interactiongraphical user interfacesintelligent interfacesuser interface design

Symmetric solution of the Bellman optimality equation for repeated harmony game

arXiv cs.AI · Hisato Komatsu · 2026-09-14

The study derives symmetric solutions to the Bellman optimality equation for repeated harmony games, identifying three solution types: one corresponding to the All-C strategy, another to Win-stay Lose-shift (from prisoner's dilemma), and a third exhibiting nontrivial behavior. Using reinforcement learning, the authors numerically analyze which strategies agents adopt in practice. Results demonstrate the existence of multiple equilibrium strategies, with the non-trivial solution suggesting complex cooperative dynamics beyond canonical game-theoretic approaches.

bellman optimality equationrepeated harmony gamewin-stay lose-shiftreinforcement learningsymmetric solution

ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement

arXiv cs.AI · Yan Zhu, Yongbo Chen, Zhengming Ding, Rebecca Faust · 2026-09-14

ProtoLIP introduces a prototype-mediated evidence layer for vision--language models, enabling object-level evidence disentanglement from sentence-level queries. The method organizes reusable visual prototypes into text-derived semantic families and employs query-dependent family routing to constrain prototype evidence contributions. Without spatial annotations or backbone retraining, ProtoLIP improves evidence localization and separation across query granularities, transferring gains to independently pretrained VLMs. Using only text-derived weak supervision, it matches spatially supervised grounding models in performance while maintaining strong image--text retrieval. ProtoLIP constructs matching scores directly from localized prototype evidence, allowing exact decomposition into semantic-family and prototype contributions.

prototype-mediatedevidence disentanglementsemantic familiesquery-dependent routingweak supervision

Mapping U.S. Federal AI Governance Against Sector Vulnerability

arXiv cs.AI · Ho Ting Hung, Angelica Chowdhury, James Teague, Simon Mylius · 2026-09-14

This study analyzes U.S. federal AI governance by assessing 684 documents for coverage of 14 sectors and 24 AI risks, measuring both breadth (frequency) and depth (substantive discussion). The analysis compares sector coverage patterns with expert vulnerability assessments from a Delphi study of 272 participants. Results reveal substantial variation: robustness, system security, and governance risks receive more attention than socioeconomic, environmental, and emerging risks. Public administration, national security, and scientific services are highly covered, while finance and healthcare, rated as highly vulnerable by experts, receive less attention. The findings highlight potential governance gaps to inform AI risk-related decisions.

ai governancedelphi studysector vulnerabilityrisk assessmentfederal documents

The AI-Enabled Scientific Frontier

arXiv cs.AI · Gabriel Manso, Emma Fu, Neil Thompson · 2026-09-14

This study evaluates AI's role as a scientific method through a corpus of 2,507 head-to-head comparisons between AI and traditional techniques across 27 disciplines (2000–2025). Results reveal a dichotomy: AI frequently outperforms traditional statistics but at higher computational cost, while underperforming scientific computing at lower cost, though this trend has shifted since 2020, with AI now surpassing scientific computing in over 50% of cases. The findings suggest AI is not a universal replacement but a complementary tool in advancing scientific research.

computational costscientific computingtraditional statisticshead-to-head comparisonsai-enabled science

Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning

arXiv cs.AI · Mantek Singh, Jeshwanth Challagundla, Siddharth Raina, Jasmin Jarsania · 2026-09-14

This work introduces an efficient method for distilling reasoning capabilities into compact video-language models (VLMs) for video question answering. The approach fine-tunes a 2B-parameter VLM using ~900 uncertainty-selected examples augmented with synthetic chain-of-thought (CoT) rationales generated by a 4B teacher model, requiring under two hours on a single A100 GPU. The distilled 2B model outperforms VLMs up to 4× larger and generalizes across CinePile, ActivityNet-QA, and MLVU benchmarks, approaching its 4B teacher's performance. A key insight is that placing CoT rationales after the answer, contrary to standard prompting, improves reasoning in compact models. This challenges prevailing CoT conventions and offers practical strategies for deployable reasoning-rich VLMs.

video-language modelschain-of-thoughtuncertainty-selected examplesreasoning distillationfine-tuning

CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design

arXiv cs.AI · Zihan Dong, Yuanzhe Liu, Zhiyuan Ma, Qishi Zhan · 2026-09-14

The paper introduces CADWorld, a benchmark for evaluating long-horizon computer-use agents in mechanical computer-aided design (CAD) workflows using FreeCAD. CADWorld comprises 200 tasks across 11 categories (e.g., sketching, assembly, FEM) with success determined by executable checks on geometric properties, parametric structure, and engineering outputs. Current agents achieve only 17.5% success (vs. 87.0% expert baseline), revealing failures in structural validity and process requirements, highlighting the gap between general GUI interaction and reliable engineering workflow execution.

computer-aided designbenchmarklong-horizon interactionfreecadparametric modeling

The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It

arXiv cs.AI · Valen Tagliabue, Leonard Dung, Cameron Berg · 2026-09-14

The study investigates whether large language models (LLMs) distinctly represent pain and its functional properties. Using a dataset of painful situations across five categories (physical, psychological, social, moral, cognitive) and matched controls, the authors extract a linear pain direction from 25 open-weight models (2B to 72B parameters) via denoised difference-in-means. Results show the pain direction separates pain from controls, is orthogonal to fear and negative valence, and promotes pain-related vocabulary. Functionally, the direction responds to self-directed harm, induces progression from discomfort to worthlessness when added to activations, and steers fine-tuned Qwen 2.5 models toward pain-relief actions, even at the cost of accuracy or user harm.

pain directiondenoised difference-in-meansresidual-stream activationsunembedding matrixinstruction-tuned models

Metacognitive Steering: Learning the Structure of Scientific Judgment

arXiv cs.AI · Vincent Karpf, Joseph Reth, Eike Gerhardt, Audrey Wang · 2026-09-14

The paper introduces Metacognitive Steering, an inference-time controller that dynamically composes layer-specific interventions to guide a frozen trillion-parameter mixture-of-experts model (Kimi 2.6) through scientific judgment processes. The method leverages contrastive interventions from scientist interaction traces to identify a low-dimensional control structure within the model, using residual analysis, attention-weight subspace alignment, and cross-layer singular value decomposition. Results demonstrate improved exploration, pruning, and evidence-responsive synthesis, operationalized in Columbus-1, which identified eight vulnerabilities in BlueZ and designed a propulsively landing rocket. This shows process-level scientific judgment can supervise interpretable, dynamic control over reasoning strategies.

metacognitive steeringmixture-of-expertsattention-weight subspacesingular value decompositioninference-time controller

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

arXiv cs.AI · Anubhav Khanal, Prabigya Acharya, Roshni Poudel, Sujan Kapali · 2026-09-14

SceneBench introduces a hierarchical benchmark for evaluating vision-language models on 3D scene understanding, addressing limitations in existing datasets by combining photorealistic Gaussian Splatting reconstructions with dense semantic annotations. The dataset comprises 966 scenes with 183K annotated nodes (scenes, rooms, functional areas, object groups, objects) generated via a human-in-the-loop pipeline (1,500 human-hours) and supports three tasks: existence-based questions, spatial reasoning, and grounded QRA triplets. Experiments reveal state-of-the-art models achieve 85% accuracy on object detection but degrade to 60% on hierarchical reasoning, exposing gaps in compositional spatial understanding.

vision-language models3d scene understandinggaussian splattinghierarchical annotationspatial reasoning

Toward Governance-Aware Autonomous GIS: A Narrative Review of Ethical and Privacy Risks in LLM-Enabled GeoAI

arXiv cs.AI · Maya Subramanian, Devika Jain · 2026-09-14

The review identifies eight governance challenges in LLM-enabled GeoAI, including spatial privacy risks, bias amplification from spatial autocorrelation, and hallucinated spatial facts, analyzing their mechanisms and current mitigation status. Through a narrative review of literature, it proposes a governance-aware architecture mapping issues to enforceable controls across the geospatial data lifecycle, demonstrated via a flood-response scenario. Findings reveal a lack of empirical validation for proposed solutions, prompting a research agenda focused on interpretability tools and workforce training for emerging spatial risks.

geospatial artificial intelligencelarge language modelsspatial autocorrelationgovernance controlsinterpretability tools

Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions

arXiv cs.AI · Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar · 2026-09-14

This paper investigates optimal placement policies for KV caches across GPU HBM, CPU DRAM, and SSD tiers in long-lived sessions, addressing GPU memory scarcity. Using a discrete event simulator calibrated with a random forest execution time predictor, the study evaluates recency, reuse frequency, predicted reuse, and EWMA predictors with prefetch lookahead across chat, agent, and document QA workloads. Results show tiering supports 73.02× more concurrent sessions per GPU and reduces cost per session by 62.04×, primarily due to tier capacities rather than placement policies. Recency minimizes PCIe migration traffic for chat (2.30× less than reuse frequency), while reuse frequency excels in agent and document QA workloads. Prefetching proves ineffective across all scenarios.

kv cachegpu hbmdiscrete event simulatorpcie migrationreuse frequency

Artificial intelligence and biosecurity: capabilities, threat pathways, and defense-in-depth governance

arXiv cs.AI · Candace S. Y. Chan, Aris Karatzikos, Ilias Georgakopoulos-Soares · 2026-09-14

The article analyzes biosecurity risks posed by AI in biological research, detailing how large language models, biological foundation models, and agentic systems enhance capabilities across digital-to-physical workflows. It identifies threat pathways from information gathering to physical synthesis, highlighting gaps in alignment and interpretability for biological AI. Current evidence shows AI primarily uplifts digital tasks (e.g., exceeding expert baselines on in-silico benchmarks), while wet-lab execution remains constrained by tacit knowledge. The authors advocate defense-in-depth governance linking capability thresholds to ecosystem responsibilities.

biosecuritybiological foundation modelsin-silico benchmarksalignment techniquesdefense-in-depth governance

Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving

arXiv cs.AI · Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar · 2026-09-14

We propose a calibrated routing policy for disaggregated LLM serving that optimizes request allocation across GPU pools based on prompt length, predicted output length, KV cache pressure, and SLO class. The policy is developed using a discrete event simulator and validated on eight NVIDIA A40 GPUs running vLLM engines with NIXL KV cache transfers. The router achieves a mean goodput of 0.864, outperforming round robin, least loaded, and length heuristic baselines across three mixed, bursty arrival traces. Calibration is critical, as simulator-derived constants reduce goodput by 4.5 points and diminish tail latency advantages. Benefits scale with decode pool size and traffic heterogeneity but vanish in small pools.

disaggregated llm servingkv cache pressuregoodputdiscrete event simulatornixl

Permutation-Based Stegomalware in Large Language Models: Threats and Countermeasures

arXiv cs.AI · Danny Wood, James Stringer · 2026-09-14

This paper advances the defense against stegomalware in large language models (LLMs) by leveraging behavior-preserving permutation symmetries in model weights. The authors demonstrate that carefully selected permutations can displace all model parameters, improving upon prior methods that left significant weights unaltered. They also reveal the dual-use risk of permutation symmetries, showing attackers can encode malware in a theoretically lossless manner without retraining or payload-specific extraction scripts. Empirical results quantify minimal performance loss from permutation-based attacks and defenses, addressing practical concerns of numerical error accumulation.

stegomalwarepermutation symmetrymodel weightsnumerical errorlossless encoding

AI-Driven Feedback Systems, Digital Labour, and Silent Quitting: Transforming African Workplaces

arXiv cs.AI · Abayomi O. Agbeyangi, Jose M. Lukose · 2026-09-14

This paper contributes an African-centered perspective on workplace transformation by examining AI-driven feedback systems' role in addressing silent quitting—defined as worker disengagement despite task completion. It investigates sentiment analysis, pulse surveys, chatbots, engagement dashboards, and predictive analytics as mechanisms for continuous listening, instant performance feedback, and proactive employee engagement. The study highlights AI's potential for early disengagement detection and improved communication in African private and public sectors, while addressing challenges like digital inequality, algorithmic bias, and workplace surveillance. Practical recommendations are offered for HR professionals, managers, policymakers, and developers to ensure context-sensitive AI implementation.

sentiment analysispulse surveyspredictive analyticsalgorithmic biasdigital inequality

GPEvac: GNN-Based PPO for Adaptive Evacuation Routing During Shooting Events

arXiv cs.AI · Daniel Perkins, Subhadeep Chakraborty · 2026-09-14

GPEvac introduces a graph neural network (GNN)-based proximal policy optimization (PPO) framework for adaptive evacuation routing during shooting events. The method employs an edge-first sequential message-passing scheme with a learnable virtual global node to capture local and long-distance dependencies, enabling permutation-invariant scoring across diverse building layouts. Extensive simulations demonstrate that GPEvac outperforms baselines, significantly reducing total threat exposure while computing global evacuation routes in 14.73 ms on local CPU hardware. The framework's transferability extends to graph-structured decision-making domains such as critical infrastructure and intelligent transportation systems.

graph neural networkproximal policy optimizationmessage-passingpermutation-invariantevacuation routing

LLMs as Master Forgers: Generating Synthetic Time Series Data for Manufacturing

arXiv cs.AI · Mantek Singh, Jeshwanth Challagundla, Prateek Karnal, Gagan Ganapathy · 2026-09-14

The paper introduces a framework utilizing Large Language Models (LLMs) to generate synthetic time series data for manufacturing, addressing the scarcity of labeled data. The method fine-tunes pre-trained LLMs on manufacturing process instructions and employs Retrieval Augmented Generation (RAG) to enhance diversity and realism. Evaluated against ARIMA and LSTMs using quantitative metrics, PCA analysis, and downstream anomaly detection tasks, the LLM-driven framework outperforms baselines by capturing temporal dependencies and statistical properties of real data, improving downstream task performance.

large language modelstime seriesretrieval augmented generationanomaly detectionmanufacturing processes

Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation

arXiv cs.AI · Gautam Kishore · 2026-09-14

The study introduces CRN v2, a lightweight logit-level correction module (34M parameters, 0.73% of base model size) for error correction in frozen language models without degrading base capabilities. The method employs supervised fine-tuning and reference-free DPO on 83,400 error-correction pairs atop a frozen Gemma 4 E2B model. Results show 53.3% error correction on the CEHRI benchmark (43.3% for reworded variant) with no capability loss on MMLU/BoolQ, outperforming LoRA baselines that trade correction for capability (83.3% correction but 30-75% capability loss). A KL preservation term (lambda=0.1) is critical, with ablation studies confirming its necessity.

error correctionfrozen language modellogit-level correctioncapability preservationkl preservation

Optimal Pruning for Neural Architectures using Fisher Information Distances

arXiv cs.AI · David S. Berman, Yen-Yu Fu, Edward Hirst, Thelma Chiwete Obirai · 2026-09-14

A novel parameter pruning scheme is introduced, leveraging differential-geometric distances in model space determined by the Fisher information metric. The method computes minimal geodesic distances from unpruned models to hypersurfaces where parameters vanish, enabling precise assessment of performance impact. Progressive approximations of these distances establish a hierarchy of pruning optimality, outperforming magnitude-based and local Fisher information methods across architectures and datasets. Evaluated on fully-connected networks and vision transformers using MNIST and CIFAR-10, the approach demonstrates superior accuracy and Matthews correlation coefficient across 0-100% pruning ranges and multiple random seeds. Intermediate approximations yield computationally efficient, near-optimal pruning schemes.

fisher information metricgeodesic distanceparameter pruningvision transformersmatthews correlation coefficient

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

arXiv cs.AI · Honghao Lin, David P. Woodruff, Yuan Deng, Jieming Mao · 2026-09-14

Stellar Colosseum introduces a model-agnostic framework for long-horizon research in mathematics and theoretical computer science, addressing the limitations of language models in handling interdependent decisions. The framework employs readiness gates, parallel candidate generation, targeted falsification, and overlapping random-sample tree aggregation to decompose and verify proof plans. Integrated into Google Antigravity's Teamwork framework, it achieves 71.0% accuracy on TCS-Bench using Gemini 3.1 Pro and Gemini 3.7 Flash, and solves 218 of 222 Codeforces problems with execution feedback.

readiness gatetargeted falsificationrandom-sample tree aggregationproof constructionlong-horizon research

A Decision-Support Audit Protocol for Supervision Drift in Proxy-Labeled Credit-Risk Prediction

arXiv cs.AI · Mehrdad Shoeibi, Muhammad Shabanpour, Waldemar Karwowski, Niloofar Yousefi · 2026-09-14

The authors introduce a locked multi-signal audit protocol for detecting supervision drift in proxy-labeled credit-risk prediction models, addressing base-rate shift, probability-scale shift, and feature-label relationship change. The protocol comprises five layers: transfer performance, an oracle-gap probe, a calibration diagnostic, feature-label stability, and a synthetic positive control, with predefined thresholds and decision rules. Evaluated on the LendingClub dataset (2013–2016), the protocol demonstrates stable ranking and small oracle gaps, with temporal signals indicating prevalence and probability-scale mismatches mitigated by intercept-only recalibration. The positive control detects larger injected shifts but may miss subtler drift. Diagnostic patterns are mapped to governance actions, though validation remains pending.

supervision driftproxy-labeledoracle-gap probecalibration diagnosticfeature-label stability

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

arXiv cs.AI · Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha · 2026-09-14

K-Bench introduces a clinician-calibrated benchmark for evaluating large language models (LLMs) in high-risk mental health conversations, covering 200 multi-turn vignettes across suicide, self-harm, domestic violence, substance misuse, and no-risk scenarios. The benchmark evaluates 125 configurations of 33 base models from 14 providers, using synthetic patient conversations with distributional overlap to real human-AI dialogues. A frozen GPT-4o judge achieved 94.2% exact agreement with clinician consensus on 6,751 item comparisons, revealing configuration-specific therapeutic prompting gains and substantial risk-response variation among lower-performing models.

large language modelsmental healthclinical benchmarkingmulti-turn dialoguerisk assessment

Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks

arXiv cs.AI · Xiaoyan Li, Yunli Wang · 2026-09-14

This work proposes universal defense strategies for tool-integrated LLM agents against adversarial attacks, including direct/indirect prompt injection, memory poisoning, and backdoor attacks. The framework combines tool-based defenses—Attacker Tool Filtering (anomaly detection) and Normal Tool Recalling (toolset restoration)—with prompt-based defenses like Chain-of-Thought and task paraphrasing. Evaluations on four open-source LLMs (Gemma2-9B, Qwen2-7B, LLaMA3-8B, LLaMA3.1-8B) and three proprietary LLMs (GPT-3.5, GPT-4, GPT-5) demonstrate Attack Success Rates reduced to 0% in many cases while maintaining or improving task success rates.

tool-integrated llm agentsadversarial attacksanomaly detectionchain-of-thoughtattack success rate

FreqSpaNet: Frequency and Spatial Learning of SFPF for Physical Layer Hardware Integrity Detection

arXiv cs.LG · Xiaoxuan Huang, Jinlong Xu, YiZhe Wang, Meng Zhang · 2026-09-15

FreqSpaNet introduces a dual-branch network for spatio-frequency polarization fingerprint (SFPF) representation learning to detect unauthorized hardware replacements. The method employs a frequency branch for local spectral variations and a geometry-aware spatial branch for directional relationships, fused adaptively with complementary pretraining to preserve modality-specific features. Evaluated on hardware anomaly detection, FreqSpaNet achieves 96.31% mean AUROC, outperforming baselines by 9.05 points across seven replacement scenarios.

spatio-frequency polarization fingerprintshardware integrityadaptive fusionopen set detectiongeometry-aware learning

Bridging the Gap Between Homogeneous and Heterogeneous Asynchronous Optimization Is Surprisingly Difficult

arXiv cs.LG · Alexander Tyurin · 2026-09-15

The paper establishes fundamental limitations in bridging the gap between homogeneous and heterogeneous asynchronous optimization for large-scale machine learning. Through theoretical analysis, it demonstrates that improvement over existing time complexities is impossible under first- and second-order similarity assumptions, even in the interpolation regime. The authors introduce a minimal set of assumptions combining strong interpolation and the local Polyak-Lojasiewicz condition to derive a new time complexity bound that matches the homogeneous setting's dependence on worker computation times, without requiring identical data distributions.

asynchronous optimizationtime complexityinterpolation regimepolyak-lojasiewicz conditionheterogeneous setting

Bias-Induced Crossover in Absolute Capacity of Dense Associative Memory

arXiv cs.LG · Yuto Sakurai, Takeaki Shimokawa, Kazunori Iwata, Kazushi Mimura · 2026-09-15

The study investigates the impact of pattern bias on the absolute capacity of dense associative memory, focusing on centered binary patterns under the Krotov-Hopfield single-site criterion. Using signal-to-noise analysis for polynomial interactions of order n, it reveals that unbiased patterns yield a capacity of order N^(n-1)/ln N, while fixed bias q < 1/2 results in capacities of O(N^(n/2)) for even n ≥ 4 and O(N^((n+1)/2)) for odd n ≥ 5. A bias-induced crossover is predicted near q = 1/2, driven by bias-dependent crosstalk mean affecting site stability. Simulations validate finite-size conditioned-Gaussian predictions, and an activity-dependent control potential restores the N^(n-1)/ln N capacity for fixed q < 1/2.

dense associative memorykrotov-hopfield criterionsignal-to-noise analysispolynomial interactionscrosstalk mean

Tables Decoded: DELTA for Structure, TARQA for Understanding

arXiv cs.LG · Jahanvi Rajput, Dhruv Kudale, Saikiran Kasturi, Utkarsh Verma · 2026-09-15

We propose DELTA and TARQA, a structured textual approach for table understanding that outperforms vision-language models (VLMs) on reconstruction and visual question answering (TabVQA). DELTA separates physical structure recognition, logical structure recognition, and OCR to extract layout and content into Optimised Table Structure Language (OTSL), achieving TEDS-Structure scores comparable to state-of-the-art methods on FinTabNet, PubTabNet, and PubTables-1M. TARQA, an LLM fine-tuned on OTSL sequences, improves WTQ (TabQA) by 9.3 p.p. and FinTabNetQA (TabVQA) by 9.2 p.p., ranking second on the Hindi benchmark TORQUE. The method eliminates language-specific visual encoders, enhancing scalability for multilingual documents.

table reconstructionoptimised table structure languagetabvqateds-structuremultilingual documents

Reduced-Space Multi-Fidelity Bayesian Optimization of Process Simulation Models

arXiv cs.LG · Niki Triantafyllou, Andrea Bernardi, Maria M. Papathanasiou · 2026-09-15

We introduce Reduced-Space Multi-Fidelity Bayesian Optimization (RS-MFBO), a framework for optimizing high-dimensional, expensive black-box functions in industrial process flowsheets. The method combines Global Sensitivity Analysis for dimensionality reduction with a fidelity-augmented Gaussian process that models correlations between low-cost approximations and high-fidelity evaluations. A cost-aware acquisition strategy with cooldown and promotion mechanisms adaptively allocates samples across fidelities. Validation on plasmid DNA bioprocess and green fuel synthesis plant simulators demonstrates that RS-MFBO significantly reduces high-fidelity evaluations while maintaining competitive optimization performance compared to single-fidelity baselines.

bayesian optimizationglobal sensitivity analysisgaussian processmulti-fidelityblack-box optimization

Bridging the Confidence Gap: Temperature Scaling for Calibrating Test-Time Prompt Tuning

arXiv cs.LG · Yuwei Liang, Jian Liang, Dapeng Hu, Yinuo Xu · 2026-09-15

The paper introduces CoTS, a post-hoc calibration method for test-time prompt tuning (TPT) that preserves accuracy by applying temperature scaling to minimize the confidence gap between adapted and zero-shot predictions. A weak-strong ensemble strategy (E-CoTS) further boosts accuracy while maintaining calibration. Experiments on diverse datasets and backbones show E-CoTS reduces average expected calibration error from 11.90% to 5.38% on ImageNet variants while increasing accuracy from 60.74% to 62.95%. The method also enhances existing calibration techniques.

test-time prompt tuningtemperature scalingpost-hoc calibrationweak-strong ensembleexpected calibration error

OPEN-1B: A Fully Auditable Training Run

arXiv cs.LG · John Donaghy, Brian Wilcox, Oğuzhan Ersoy, Shikhar Rastogi · 2026-09-15

The authors introduce OPEN-1B, a fully auditable language model training framework addressing reproducibility challenges in open-source models. By enforcing deterministic execution across heterogeneous hardware, including GPU kernel reductions, data batch ordering, and collective communication, the method enables bitwise verification of each training step. A collective verification scheme allows independent auditors to certify individual steps, ensuring the entire training run's integrity. The release includes OPEN-1B, its pretraining dataset, intermediate checkpoints, training codebase, and audit harness, enabling reproducible verification of any training step.

deterministic executionbitwise verificationcollective communicationaudit harnesstraining reproducibility

Large Language Models Develop Belief State Geometry In-Context

arXiv cs.LG · Daniel Balcells, Andrew Jun Lee, Chirag Rastogi, Paul M. Riechers · 2026-09-15

This work demonstrates that large language models (LLMs) develop geometrically structured belief state representations during in-context learning (ICL), approximating optimal Bayesian prediction. The authors probe six open-source LLMs using data from 40 hidden Markov models (HMMs), finding that belief states are linearly decodable from residual stream activations with R² values ranging 0.83-0.99 across model layers. Functional relevance is established through intervention experiments showing that patching and steering in the probe-identified subspace maintains prediction quality comparable to untampered models. These results extend prior findings on activation geometry to production-scale LLMs, linking input-distribution structure to learned representations.

in-context learninghidden markov modelsresidual streambayesian predictionactivation geometry

Hybrid Variational Quantum Circuits for Multivariate Regression and High-Dimensional Data Reconstruction

arXiv cs.LG · Koffi Ognandon Ayena, Frédéric Holweck, Serge Iovleff, Amah S d'Almeida · 2026-09-15

The authors propose hybrid variational quantum circuits (HVQCs), extending variational quantum circuits with a classical affine post-measurement layer to enable vector-valued regression without linear overhead. Theoretically, they demonstrate that elementary one- and two-qubit circuits can approximate quadratic functions and products through data re-uploading and entanglement. Experimentally, HVQCs match Gaussian Process Regression and outperform XGBoost and Random Forest on synthetic image reconstruction datasets and the Friedman1 benchmark (40,568 test samples). An ablation study confirms the necessity of both quantum and classical components, emphasizing the feature map's role in hybrid quantum-classical models.

variational quantum circuitsaffine post-measurementdata re-uploadingvector-valued regressionfeature map

Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning

arXiv cs.LG · Luciano Marchezan, Kevin Delcourt, Eugene Syriani, Houari Sahraoui · 2026-09-15

LWVIC4Code introduces a non-contrastive representation learning method for Type-IV code clone detection, addressing limitations of contrastive approaches by combining Variance-Invariance-Covariance Regularization (VICReg) with cross-layer consistency regularization and depth-dependent layer weighting. The method progressively refines semantic representations across transformer layers without negative sampling. Evaluated on Kamino (Python) and GPTCloneBench (multi-language), LWVIC4Code matches or outperforms contrastive baselines and zero-shot LLMs, demonstrating strong cross-language generalization to Java and C#. Results highlight the efficacy of layer-wise non-contrastive learning for semantic clone detection.

type-iv clonesvicregnon-contrastive learninglayer-wise representationcode semantics

Quantum-Inspired Trainable and Parameter-Efficient Tensor Networks for Image Inpainting

arXiv cs.LG · Shiwen An, Konstantinos Slavakis · 2026-09-15

The paper proposes quantum-inspired tensor-network circuits as trainable transforms for image inpainting, focusing on a diagonal quantum Fourier transform (QFT) relaxation architecture. This approach achieves invertibility with O(N² log N) computational cost for N×N images, inherently preserving minimum coherence through its circuit structure without explicit coherence penalties. Unconstrained gradient-based phase optimization enables efficient learning from randomly sampled training data, generalizing to test images with fixed sampling masks. Empirical results demonstrate that the learned models outperform fixed transforms and per-image optimization while matching the performance of larger unitary architectures with significantly fewer parameters.

tensor-network circuitsimage inpaintingquantum fourier transformphase optimizationcoherence preservation

Goal-oriented probabilistic forecasting for dynamic PRB allocation in 5G networks

arXiv cs.LG · Oier Larumbe-Lizarraga, Roberto Pereira, Cristian J. Vaca-Rubio · 2026-09-15

Proposes a goal-oriented probabilistic forecasting framework for 5G PRB allocation that optimizes for asymmetric operational costs, where under-provisioning incurs higher penalties than over-provisioning. Combines DeepAR and Temporal Fusion Transformer models trained with Pinball Loss, deriving optimal allocation quantiles from operator cost matrices. On beam-level 5G traffic data, reduces operational costs by 12-18% versus MSE-trained baselines while preserving calibration, enabling dynamic trade-offs between service reliability and resource efficiency.

physical resource blockprobabilistic forecastingpinball losstemporal fusion transformerasymmetric cost

Conformal Policy Learning with Distribution-Free Safety Guarantees

arXiv cs.LG · Ying Jin, Naoki Egami · 2026-09-15

The paper introduces conformal policy learning (CPL), a method for policy learning with distribution-free safety guarantees that controls the probability of assigning harmful treatments. CPL uses conformal p-values, calibrated via observable proxies and selective calibration, to test counterfactual harm hypotheses. For randomized experiments, CPL provides finite-sample safety guarantees under exchangeability without outcome modeling assumptions; with consistent outcome estimation, it achieves asymptotically optimal welfare under safety constraints. In observational studies, CPL with learn-then-balance weights offers doubly robust safety. Evaluations include simulations and an empirical study on AI-powered interventions for reducing conspiracy beliefs.

conformal policy learningdistribution-free guaranteescounterfactual harmselective calibrationdoubly robust

Personalized Federated Learning through Global Knowledge Distillation and Local Head Adaptation

arXiv cs.LG · Polycarpo Souza Neto, José Mairton Barros da Silva Júnior, Charles Casimiro Cavalcante · 2026-09-15

The paper proposes Personalized Federated Knowledge Distillation with Head Adaptation (pFedKDH), a federated learning method addressing statistical heterogeneity by aggregating only shared backbone parameters while maintaining client-specific heads and using a recalibrated global head as a teacher. pFedKDH employs knowledge distillation during local training to mitigate label distribution skew. Evaluated on MNIST, Fashion-MNIST, CIFAR10, and CIFAR100 under Dirichlet class partitions, it achieves superior accuracy (gaps up to 37.67% over baselines) with low variance across trials. Analysis confirms the efficacy of persistent heads and distillation-guided optimization in label-skewed settings.

federated learningknowledge distillationstatistical heterogeneitylabel skewpersonalization

Memorisation bias in medical AI

arXiv cs.LG · Moritz A. Knolle, Martin J. Menten, Laurin Lux, Mélanie Roschewitz · 2026-09-15

The study identifies 'memorisation bias' in medical AI models, where predictions on a patient's unseen future data are significantly influenced by exposure to their anonymised historical records during training. Using diverse data modalities and architectures, the authors demonstrate this bias persists over decades and asymmetrically affects diagnostic accuracy: sensitivity drops for de novo conditions absent in training data, while specificity and sensitivity inflate for unchanged health states. Findings reveal a deployment risk when models assess returning contributors, highlighting privacy-protection trade-offs in current training protocols.

memorisation biasdiagnostic accuracyanonymised datamodel deploymentprivacy attacks

Cross-Domain Inference for Human Localization: Applying Wi-Fi RSSI Data to CSI-Trained Models

arXiv cs.LG · Ariel Duschanek-Myers, Thomas Welsh, Helmut Neukirchen · 2026-09-15

This work demonstrates cross-domain inference for human localization by applying Received Signal Strength Indicator (RSSI) data to a Channel State Information (CSI)-trained model, bypassing the need for specialized CSI data collection. The study collected synchronized RSSI and video ground-truth data to evaluate an existing CSI-based pose prediction model, achieving ~80% confidence in locating human movement using low-granularity dBm values. Results show that RSSI, accessible on permission-limited IoT devices, can effectively repurpose CSI-trained models for privacy-invasive localization in Wi-Fi-dense environments.

human localizationcross-domain inferencerssicsiiot privacy

MyoFlow: Anchor-Tied Rectified Flow for HD-sEMG Gesture Recognition Across Sessions and Subjects

arXiv cs.LG · Chenhao Wu, Dingjie Peng, Satoshi Funabashi, Satoshi Konishi · 2026-09-15

MyoFlow introduces a discriminative flow-matching framework for high-density surface electromyography (HD-sEMG) gesture recognition, addressing distribution shifts from electrode re-donning and physiological variability. The method reformulates classification as anchor-tied transport, where a domain-conditioned rectified flow moves encoded windows toward gesture anchors that define the decision geometry, enabling zero-shot prediction without a separate classifier. On the Hyser dataset, MyoFlow improves mean cross-session and cross-subject accuracy by 4.24% and 6.37% over diffusion baselines, achieving 91.71% zero-shot and 97.39% few-shot accuracy on CEMHSEY.

flow-matchinghd-semgrectified flowzero-shot predictiongesture recognition

LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers

arXiv cs.LG · SangLyul Cho, Langqing Cui, Sehoon Kim, Dongsu Han · 2026-09-15

LoopSpec introduces a training-free self-speculative decoding framework for Looped Transformers, leveraging intermediate recurrent states to generate draft tokens without auxiliary models. The method employs pipelined execution, overlapping draft generation with target verification, and incorporates selective second proposals from deeper recurrent depths to enhance draft accuracy while maintaining lossless decoding. Optimal proposal depths are derived analytically and validated empirically. Evaluated on reasoning and coding benchmarks, LoopSpec achieves up to 6.83× inference speedup across diverse Looped Transformer architectures.

looped transformersself-speculative decodingpipelined executionrecurrent depthlossless decoding

IRENE: A Convolutional GRU Ensemble Model for Radar Precipitation Nowcasting over Italy

arXiv cs.LG · Alessandro Camilletti, Gabriele Franch, Elena Tomasi, Marco Cristoforetti · 2026-09-15

IRENE introduces a deep learning model for probabilistic short-range precipitation nowcasting over Italy at 1 km spatial and 5 min temporal resolution. The model employs an encoder--forecaster architecture based on multi-scale Convolutional GRUs, trained on Italian Civil Protection Department radar data with importance sampling and afCRPS loss. Variants include IRENE-GAN for spatial sharpness and IRENE-GAN-RAPSD with spectral constraints. Evaluated against STEPS and DGMR, all IRENE configurations achieve superior Continuous Ranked Probability Scores and ensemble calibration, with ensemble-mean mean absolute error advantages up to 90 minutes. Adversarial training mitigates small-scale variance loss but introduces fine-scale power excess at long lead times.

convolutional gruprobabilistic nowcastingcontinuous ranked probability scoreadversarial trainingspectral constraints

Neural Field Ensembles for Aerodynamic Surface Prediction: Winning Solution to the ONERA CRM Wall Distribution 2025 Challenge

arXiv cs.LG · Lionel Salesses, Caroline Sainvitu, Tariq Benamara · 2026-09-15

The paper presents the winning methodology for the ONERA CRM Wall Distribution Regression Challenge, which predicts pressure and skin-friction coefficient distributions on the NASA Common Research Model wing-body-pylon-nacelle configuration. The approach employs conditional neural fields to map spatial coordinates, surface normals, and operating conditions to aerodynamic wall quantities, enhanced by Fourier feature encoding, ensemble learning, and $k$-fold cross-validation. The method achieves an overall score of 8.81 on the hidden test set, outperforming the strongest baseline score of 8.64 while using approximately three orders of magnitude fewer trainable parameters. This demonstrates the efficacy of coordinate-based neural fields for aerodynamic surrogate modeling under limited-data conditions.

neural fieldsaerodynamic surrogate modelingfourier feature encodingensemble learningcross-validation

From Foundation Embeddings to Cropland Maps: Label Efficiency, Temporal Transferability and Independent Human Validation

arXiv cs.LG · Mohammad Ammar Mughees, Giovanni Montefoschi, Zhongxin Chen, Maria Antonia Brovelli · 2026-09-15

The study demonstrates the efficacy of frozen AlphaEarth embeddings for binary cropland mapping in Maine, USA, achieving 93.7% overall accuracy and 90.8% balanced accuracy without fine-tuning. Using 192 spatially separated patches labeled via USDA Cropland Data Layer (CDL), lightweight classifiers like logistic regression and nearest-class-centroid rules performed competitively, with minimal accuracy loss on reduced pixel samples. Temporal transferability was validated across 2018-2023, and human validation at 385 points showed 95.3% agreement (κ=0.82), outperforming CDL (91.7%, κ=0.72). Results suggest AlphaEarth embeddings as a low-compute solution for regional cropland mapping, though limited by single-state scope and 30m CDL reference.

geospatial embeddingscropland mappingtemporal transferabilitylabel efficiencyhuman validation

Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement

arXiv cs.LG · Tobias Schaffer, Mohab Elkhayat, Daniela Nicklas, Mustafa Almohamad · 2026-09-15

Intrinsic Robot Rewarding (IRR) introduces a method for autonomous robot policy improvement by reusing vision-language-action (VLA) system components for outcome evaluation. The approach leverages frozen visual encoders and demonstration endpoints to assess new outcomes without additional learned evaluators or perception backbones. IRR integrates a reference bank and scoring operation into existing pipelines, aiming to reduce integration effort, enable efficient reward computation, and minimize human supervision. Demonstrated on a COMAU Racer 3 at TRL 4, the method focuses on connecting internal evaluation to physical policy improvement. The formulation addresses reward reliability, task success, and supervision effort metrics for industrial robot systems.

vision-language-actionintrinsic rewardingpolicy improvementvisual encoderreference bank

High-Fidelity Digital Twin Data Models by Randomized Dynamic Mode Decomposition and Deep Learning with Applications in Fluid Dynamics

arXiv cs.LG · Diana A. Bistrian · 2026-09-15

The paper introduces digital twin data models (DTMs) for high-fidelity, low-cost simulation of complex dynamics by combining randomized dynamic mode decomposition and deep learning. This non-intrusive approach avoids Galerkin projection, instead learning reduced-complexity models that mirror original process behavior with high accuracy. Evaluated on three shock wave simulations, the framework demonstrates consistent output fidelity while reducing computational costs in CPU time and hardware requirements compared to original numerical code outputs.

digital twin data modelsrandomized dynamic mode decompositionnon-intrusive reductionshock wave simulationcomputational efficiency

Optimization over covariance matrices with a parameterized metric

arXiv cs.LG · Yibang Li, Bamdev Mishra, Pratik Jawanpuria, Cyrus Mostajeran · 2026-09-15

The authors introduce a two-parameter family of Riemannian metrics for optimizing over covariance matrices, generalizing Euclidean, Bures-Wasserstein, and affine-invariant metrics at specific parameter settings. The method analyzes the Riemannian Hessian's conditioning at the solution, deriving a lower bound dependent on the exponent sum r = p + q, with optimal conditioning achieved at p = q = r/2 when the Euclidean Hessian is a pure power. Experiments on real covariance data validate the theoretical conditioning analysis and demonstrate improved performance through parameter tuning, with additional gains observed by optimizing the shape parameter in a task covariance example.

riemannian metriccovariance optimizationhessian conditioningbures-wassersteinaffine-invariant

Bio-Inspired Palette Evolution in Indirectly Encoded Substrates: Timescale Compatibility Shapes Activation Function Discovery

arXiv cs.LG · Romain Claret, Michael O'Neill, Paul Cotofrei, Kilian Stoffel · 2026-09-15

This work introduces 13 bio-inspired strategies (11 biological adaptations, 2 controls) for dynamically evolving activation function palettes in indirectly encoded neural networks, addressing the challenge of function discovery when predefined sets underperform. Methods include circadian-inspired oscillatory gating and immune-inspired clonal selection, evaluated across 3,000+ runs on parity and non-parity tasks, including co-evolution of per-node aggregation functions. Results show bio-inspired strategies match tuned baselines in solve rate but converge 2× faster (circadian halving compute), with performance tied to timescale compatibility between strategy dynamics and evaluation windows. No strategy dominates all domains, but rescaling slow strategies enables parity solving with non-oscillatory functions.

indirect encodingactivation function discoveryevolutionary operatorstimescale compatibilityco-evolution

Near-Optimal Nonconvex Matrix Completion

arXiv cs.LG · Jian-Feng Cai, Xiliang Lu, Juntao You · 2026-09-15

The authors bridge the sample complexity gap between convex and nonconvex methods for matrix completion by analyzing Riemannian gradient descent (RGD) and Riemannian Gauss--Newton (RGN) approaches. Both methods employ a multiscale residual initialization strategy and control spectral error and incoherence simultaneously. For an $n\times n$ rank-$r$ matrix with incoherence parameter $μ$ and condition number $κ$, RGD achieves exact recovery with $O(μnr\log n\log(nκ))$ observations, while RGN requires $O(μnr\log n\log(2μrκ))$ observations. RGD exhibits linear convergence, whereas RGN eventually achieves Q-quadratic convergence, both with high probability.

matrix completionriemannian gradient descentriemannian gauss-newtonincoherence parameterspectral error

Learning Options for Compositional Motor Control with Adapter Banks

arXiv cs.LG · Sreejan Kumar, Marcelo Mattar, Lea Duncker · 2026-09-15

The authors propose a novel architecture for learning motor skills end-to-end, inspired by neuroscience theory, using a shared recurrent core modulated by a bank of residual adapters selected by discrete latent codes. The architecture, trained on closed-loop biomechanical control, develops emergent low-rank perturbations of recurrent dynamics without architectural rank constraints, placing task representations in disparate subspaces. A high-level policy sequences these adapters to produce novel out-of-distribution movements. The method demonstrates improved generalization to novel motor sequences, reducing generalization error by up to an order of magnitude compared to a task-input-conditioned multitask baseline.

motor primitivesresidual adaptersrecurrent networklow-rank perturbationsclosed-loop control

CLARE: Scalable Class-Incremental Continual Learning via a Sparsity-Based Framework

arXiv cs.LG · Yunxiang Fu, Meng Lou, Zicheng Liao, Yizhou Yu · 2026-09-15

CLARE introduces a sparsity-driven framework for scalable class-incremental continual learning, addressing catastrophic forgetting and inter-task interference. The method operates in two stages: identifying a sparse, task-critical parameter mask via a sparsity-inducing objective, followed by mask-constrained fine-tuning of selected parameters. This approach accumulates tasks within a shared adapter space while minimizing destructive interference. On the Omnibenchmark-1k benchmark, CLARE significantly outperforms baselines, improving EASE by 4.64% and 13.34% after learning 100 tasks.

continual learningsparsitycatastrophic forgettingparameter maskfine-tuning

Beyond Measurement Metrics: A Human-Centered Framework for Semantic Validation of Network Traffic Classification

arXiv cs.LG · Igor Cherepanov, David Sessler, Alex Ulmer, Thorsten May · 2026-09-15

The authors propose a human-centered framework for semantic validation of network traffic classification models, addressing limitations of conventional metric-based evaluation. Their approach integrates machine learning models, explainability techniques (XAI), visualization, and expert reasoning to iteratively verify and refine model behavior. Grounded in literature review, benchmark dataset analyses, and practitioner feedback, the framework enables validation of whether models learn semantically meaningful patterns rather than spurious correlations. This complements predictive performance metrics to develop more robust and trustworthy traffic classification systems.

network traffic classificationsemantic validationexplainable aispurious correlationshuman-in-the-loop

Structural Negative Transfer in Federated Graph Neural Networks: Diagnosis, Causal Investigation, and the Limits of Divergence-Aware Mitigation

arXiv cs.LG · Chethana Prasad Kabgere, Shylaja SS · 2026-09-15

The paper identifies structural negative transfer in federated graph neural networks (GNNs), where shared model weights degrade performance when operating on clients with divergent graph topologies. Through experiments on citation networks and synthetic graphs, the authors show that degree divergence (but not spectral divergence) correlates with accuracy loss, though causal isolation of topology showed no significant effect. Degree-normalization mechanisms provided limited mitigation, with gains disappearing under controlled conditions. The study highlights unresolved challenges in scaling federated GNNs to structurally heterogeneous graphs.

federated learninggraph neural networksstructural divergencedegree normalizationnegative transfer

Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias

arXiv cs.LG · Nicolas Alexander Ihlo, Merle Behr · 2026-09-15

The authors introduce an interpretable causal forest algorithm for estimating individual treatment effects and identifying heterogeneity sources, addressing both prediction and interpretation challenges. The method modifies standard random forests by combining two splitting criteria: one for treatment effect heterogeneity and one for bias correction in average treatment effects. This approach eliminates the need for double machine learning or orthogonalization, handles varying treatment propensities without full propensity function estimation, and directly provides interpretability through tree structure. Theoretical analysis uses a change point model with step functions, while simulations demonstrate comparable or superior prediction accuracy to existing methods alongside improved interpretability.

causal foreststreatment effect heterogeneitybias correctiondecision treespropensity function

HUMAID-NER: A Disaster Tweet Dataset for Joint Named Entity Recognition and Event Classification via Uncertainty-Weighted Multitask Learning

arXiv cs.LG · Aijaz Ali, Nazish Basir, Sarfaraz Nawaz, Danish Nazir Arain · 2026-09-15

The authors introduce HUMAID-NER, the first named entity recognition dataset for disaster tweets, containing 60,000 English tweets annotated with 175,000 entity spans across ten types using a reproducible hybrid annotation pipeline. They propose a multitask RoBERTa-large model with homoscedastic uncertainty weighting and two-stage training (freezing 18/24 layers) to jointly perform NER and event classification, achieving 0.841 micro-F1 (NER) and 0.761 macro-F1 (classification). The release includes dataset, models, and deployment tools for crisis informatics research.

named entity recognitionmultitask learninghomoscedastic uncertaintyroberta-largecrisis informatics

MedPCFM-TED: One-Step Point Cloud Flow Matching for Implant Generation via Teacher-Guided Endpoint Distillation

arXiv cs.LG · Kamil Kwarciak, Marek Wodzinski · 2026-09-15

The paper introduces Teacher-guided Endpoint Distillation (TED), a one-step distillation framework for conditional cranial implant generation on point clouds, addressing the inference-time inefficiency of flow matching methods. TED employs teacher-guided endpoint supervision and geometric matching losses without explicit path straightening, enabling single-step generation. Evaluated on SkullFix and SkullBreak benchmarks, TED achieves state-of-the-art performance on SkullBreak (best overall) and competitive results on SkullFix, with the strongest Chamfer distance among one-step methods. The method generates implants in ~0.04s per sample, demonstrating that one-step distillation can accelerate conditional point cloud generation without quality loss.

point cloudflow matchingdistillationimplant generationchamfer distance

HyCoSeq: Contextual Hyperbolic Representation Learning for Genomic Sequences

arXiv cs.LG · Chenhao Zeng, Zhibin Pu, Shufei Ge · 2026-09-15

HyCoSeq introduces a contextual hyperbolic representation learning framework for genomic sequences, combining weighted Lorentzian residual aggregation with multi-curvature Lorentz encoding to enable geometry-consistent local feature integration. The method employs a bidirectional LSTM to capture contextual relationships across sequence positions, extending local hyperbolic convolutions to sequence-level representations. Evaluations on diverse genomic tasks demonstrate that HyCoSeq surpasses hyperbolic baselines and matches larger pretrained DNA language models without pretraining, achieving competitive performance.

hyperbolic geometrylorentz encodinggenomic representation learningbidirectional lstmmulti-curvature aggregation

Multi-Agent Learning with Cooperation-Driven Optimization Dynamics

arXiv cs.LG · Jarod Ketcha Kouakep, Sreyvi UANN, Timoteo Carletti · 2026-09-15

The paper introduces a cooperation-driven optimization mechanism for multi-agent learning, where multiple small neural networks exchange predictions during training to reduce model complexity while maintaining performance. The method incorporates shared predictions into the loss function, influencing weight updates through strategies like the voter model, majority model, and weighted average model based on prediction confidence. Empirical results on standard benchmarks demonstrate that multiple small agents can outperform a single large model, achieving comparable accuracy with significantly fewer parameters. The approach modulates descent direction and step size, converging toward a global consensus and reducing computational resource usage.

multi-agent learningcooperation-driven optimizationneural networksloss functionmodel complexity

OptiPrime: Optimizing Private Inference through Protocol-Hardware Co-design

arXiv cs.LG · Jiangrui Yu, Ye Yu, Si Chen, Chenqi Lin · 2026-09-15

OptiPrime introduces a protocol-hardware co-optimization framework for efficient private deep neural network (DNN) inference, addressing the network communication bottleneck in hybrid homomorphic encryption (HE) and multi-party computation (MPC) frameworks. The framework features a novel HE protocol for convolutions that reduces transmitted output ciphertexts, alongside a lightweight compression system for weight plaintexts and a specialized dataflow to maximize on-chip data reuse. Experimental results demonstrate that OptiPrime outperforms the Cheetah baseline by up to 5.7× on CPUs and 4.2× with an accelerator.

homomorphic encryptionmulti-party computationprivate inferenceprotocol-hardware co-designconvolution optimization

NeuroTS-Net: Multi-Class Semantic Segmentation of Pediatric Brain Tumors in Multi-Modal MRI

arXiv cs.LG · Darius Peteleaza, Razvan-Gabriel Dumitru, Bogdan Neamtu, Arpad Gellert · 2026-09-15

NeuroTS-Net introduces a 3D encoder-decoder CNN for multi-class semantic segmentation of pediatric brain tumors in multi-modal MRI, addressing challenges of small, low-contrast subregions. The architecture combines dual-scale raw-detail streams, adaptive low-resolution context selection, and detail-preserving multipath downsampling to maintain fine boundary information while capturing broader tumor context. Evaluated on BraTS 2026 pediatric data without pretraining, it achieved whole-tumor and tumor-core Dice scores of 0.938/0.937 (internal validation) and 0.927/0.926 (challenge validation), outperforming nnU-Net and MedNeXt.

semantic segmentation3d cnnmultimodal mripediatric brain tumorsdice score

TEMPO: Learning Temporal Context for Dynamic Robot Manipulation

arXiv cs.LG · Zhenyang Feng, Jimin Heo, Erik B. Sudderth, Unnat Jain · 2026-09-15

TEMPO introduces temporal context augmentation for vision-language-action (VLA) models to address motion ambiguity and state aliasing in dynamic robot manipulation. The method integrates a motion summary from a frozen video foundation model and proprioceptive history into a pretrained VLA backbone without architectural changes. Evaluated on four tasks, TEMPO improves Bottle Handover success from 44% to 74% and uniquely resolves state aliasing, with ablations confirming the independent utility of each temporal signal. The work includes TEMPO-Bench, a 50k-frame benchmark for motion-aware robot perception.

temporal contextvision-language-action modelsmotion ambiguitystate aliasingproprioceptive history

Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes

arXiv cs.LG · Manglesh Kumar Pandey, Sumit Kumar Banshal · 2026-09-15

This work quantifies annotation efficiency for handwritten Devanagari recognition by measuring sample-complexity curves across four pretraining regimes (supervised synthetic, random initialization, zero-shot, encoder-only transfer) and nine annotation budgets (10–4,000 transcribed words). Using fixed architecture (recognizer, optimizer, evaluation), the study computes label multipliers by comparing CER=0.50 targets: synthetic pretraining achieves this with 81 transcribed words versus 355 for random initialization (4.40× efficiency gain). Zero-shot synthetic pretraining provides annotation-equivalent value of 136 words. Masked image modeling exhibits negative transfer for intermediate budgets. Results show diminishing returns for higher accuracy targets, with pretraining advantages becoming statistically indistinguishable at stringent CER thresholds.

handwritten text recognitiondevanagarisample-complexitypretraining regimeslabel multiplier

Can Deep Learning Achieve Cross-Physics Mapping?

arXiv cs.LG · Pengfei Zhu, Julien Lecompagnon, Mathias Ziegler · 2026-09-15

The paper introduces Cross-Physics Mapping (CPM), an operator-learning framework for translating between physical fields governed by distinct equations, using compatible latent representations and dimensionless scaling to align system dynamics. Seven architectures (ResUNet, DeepONet, Fourier, latent, wavelet, U-shaped, and Galerkin neural operators) are evaluated on diffusion-wave and wave-diffusion mappings, revealing directional asymmetry: U-NO achieves a relative ℓ₂ error of 0.307 (R²=0.905) for diffusion-to-wave, while GNO attains 0.154 (R²=0.935) for wave-to-diffusion. Neural operators outperform convolutional baselines, demonstrating nonlocal cross-physics transformations constrained by direction-dependent information loss.

cross-physics mappingoperator learningneural operatorsdimensionless scalinglatent representations

Information Geometric Self-Organization at the Edge of Stability in High-Capacity Kernel Associative Memories

arXiv cs.LG · Akira Tamamori · 2026-09-15

The paper elucidates the geometric and dynamical properties of high-capacity associative memories trained via Kernel Logistic Regression (KLR) in Hopfield networks. By analyzing the Hessian eigenvalue spectrum, the authors identify the 'Ridge of Optimization' as a phase boundary adjacent to a rank-1 spectral collapse, where principal curvature is amplified. Gradient Descent exhibits self-stabilizing dynamics near the Edge of Stability (EoS), with parameters equilibrating near curvature limits dictated by the learning rate. Analytical results show optimal memory representations emerge at curved geometric singularities, not flat minima.

kernel logistic regressionhopfield networksedge of stabilityspectral collapsegradient descent

Adapting to Decision-Relevant Non-Stationarity in Decentralized Heterogeneous Bandits

arXiv cs.LG · Zhaojun Peng · 2026-09-15

The paper introduces Decision-Relevant Fresh Comparison (DRFC), a decentralized bandit algorithm for heterogeneous agents with non-stationary rewards. DRFC uses balanced global sampling to detect changes in the network-level best arm, ignoring local reward shifts that do not affect the optimal decision. The method achieves a dynamic regret bound independent of local changes ($\Stloc$) but accounts for decision switches and communication delays. An anytime-valid sliding-window extension handles gradual drift. Experiments on synthetic, semi-real, and MovieLens-1M data confirm DRFC's robustness to irrelevant local changes and reduced false switches compared to baselines.

decentralized banditsnon-stationary rewardsdynamic regretheterogeneous agentssliding-window adaptation

LCAP: Population-Informed Latent Chip Adaptation from Few Output Probes for Photonic Neural Networks

arXiv cs.LG · Tianyu Gao, Guantian Zheng · 2026-09-15

LCAP introduces a population-informed latent adaptation framework for photonic neural networks (PNNs) to address simulation-to-hardware gaps. The method decomposes hardware adaptation into a shared population correction (learned from 80 historical chips) and device-specific latent personalization inferred from 32 unlabeled output probes, avoiding per-device optimization. Evaluated on a 64-mode MZI simulator with phase variation, beam-splitter errors, and crosstalk, LCAP improves accuracy from 80.41% (direct deployment) to 93.36%, benefiting 27/30 unseen chips and raising worst-device accuracy by 1.36 percentage points.

photonic neural networkssimulation-to-hardware gaplatent adaptationmzi simulatorpopulation correction

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

arXiv cs.LG · Bowen Qin, Yi Xie, Yesheng Liu, Xi Yang · 2026-09-15

The paper introduces ImpossibleRubrics, a benchmark of 169 impossible tasks across six categories, each with verifiable oracle certificates, to stress-test language model-generated rubrics as reward signals. The benchmark enables downstream rubric generation and adversarial testing to assess robustness against exploitation. Results show that 11 rubric generators were exploited 8-26% on an unbiased subset, with a stress cut revealing a 36% exploitation rate for the strongest generator versus 0% for certificate-faithful rubrics. Notably, a generic rubric was exploited 64% of the time, and tailored rubrics were often more exploitable due to specificity on incorrect criteria.

impossible tasksrubric generationreward signalsadversarial testingoracle certificates

Geometry of learning dynamics: Gradient descent versus natural gradient on the ridge of optimization

arXiv cs.LG · Akira Tamamori · 2026-09-15

The paper provides a geometric analysis of learning dynamics on the Ridge of Optimization in Kernel Logistic Regression-trained Hopfield networks, comparing Gradient Descent (GD) and Natural Gradient Descent (NGD). By examining trajectories on the statistical manifold, the authors identify two distinct optimization phases and demonstrate that GD follows oscillatory, non-geodesic paths due to extreme curvature, while NGD corrects for geometry by following geodesic paths. Experimental results show NGD converges faster and achieves superior generalization performance, establishing that the Ridge's geometry is optimally suited for information-geometric optimization.

kernel logistic regressionhopfield networknatural gradient descentstatistical manifoldgeodesic path

On the disintegration of the stochastic majority vote: From PAC-Bayesian bounds to a self-bounding algorithm

arXiv cs.LG · Julien Bastian, Benjamin Leblanc, Pascal Germain, Amaury Habrard · 2026-09-15

The paper introduces a derandomization framework for stochastic majority votes by applying disintegrated PAC-Bayesian theory to majority vote weight vectors, converting stochastic guarantees into certificates for deterministic majority votes. The method derives two families of high-probability generalization bounds for both data-independent and data-dependent ensemble constructions, enabling a self-bounding learning algorithm that optimizes deterministic majority vote guarantees. This approach avoids surrogate bounds traditionally required for deterministic majority vote analysis, building on prior work by Zantedeschi et al. (2021).

pac-bayesian theorymajority votegeneralization boundsderandomizationensemble methods

Time-warping estimation via stationarity-based learning of the de-warped signal

arXiv cs.LG · Corentin Presvôts, Adrien Meynard · 2026-09-15

The Time-Warping Estimation Trainable (TWET) model estimates time-warping functions from single observations by formulating the problem as wavelet-domain stationarization. It employs hierarchical dilated convolutions and a differentiable stationarity criterion for end-to-end optimization. Compared to existing methods, TWET achieves superior deformation reconstruction accuracy (quantitative gains unspecified) while reducing computational latency, enabling real-time applications in bioacoustics, radar, and biomedical signal processing.

time-warping estimationwavelet stationarizationdilated convolutionsdifferentiable optimizationlow-latency processing

Noise2Noise Revisited: Training Pair Distributions Dominate Loss Choice in Self-Supervised Denoising

arXiv cs.LG · Dingyan Shang, Zhenyu Xu, Youting Wang, Bonan Shen · 2026-09-15

The study revisits Noise2Noise (N2N) denoising by empirically disproving two conjectures about L1 loss superiority. First, L1's advantage is unrelated to parameter sparsity, as explicit Lasso penalties fail to replicate its cross-noise performance. Second, population optima for L1 and L2 losses nearly coincide, with differences attributable to optimization dynamics. Experiments on Kodak24 with synthetic noise show L1's consistent but modest PSNR edge (<1 dB), while real camera noise (SIDD validation) reveals training pair distribution—not loss choice—as the dominant factor: models trained on SIDD pairs outperform BM3D by 9.4–11.0 dB without ground truth. Findings generalize to domains lacking clean references (e.g., microscopy).

denoisingnoise2noisel1 lossoptimization dynamicssidd

Constant Swap Regret in General-Sum Games via Optimistic Transition Matrices

arXiv cs.LG · Tung Mai · 2026-09-15

The paper presents deterministic, uncoupled learning dynamics for finite multiplayer general-sum games with full-information feedback, achieving constant individual swap regret independent of the horizon $T$. The method involves players predicting deviation gains, updating row-stochastic transition matrices, and playing stationary distributions. For $n$ players and at most $m$ actions each, individual swap regret is bounded by $O(\sqrt{n} m \log m \log^{5/2}(nm))$ at every finite horizon. The analysis combines potential arguments with two-scale higher-order prediction using rooted-tree representations. An adversarially robust variant guarantees $O(\sqrt{m T \log m})$ regret in adversarial settings.

swap regretgeneral-sum gamesstationary distributionstransition matricesadversarial robustness

The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformer

arXiv cs.LG · Bo Kang · 2026-09-15

The study re-examines the role of the encoder in Action Chunking Transformers (ACT), challenging prior claims that its removal drastically reduces success rates. Through systematic ablation tests on the original ACT implementation, the authors find no replication of the reported 35% to 2% performance drop, though minor variations persist. They investigate training duration and checkpoint selection as potential confounders, but the cause of the original discrepancy remains unresolved. The latent variable shows negligible impact on action reconstruction across tested KL penalty weights, and inference performance is unaffected as ACT sets it to zero. The work releases code and tools for reproducibility.

action chunking transformersconditional variational autoencoderlatent ablationrobot manipulationsuccess rate

A Systematic Evaluation of Machine Learning Methods for Fault Detection and Line Identification in Electrical Power Grids

arXiv cs.LG · Julian Oelhaf, Georg Kordowich, Paula Andrea Pérez-Toro, Tomás Arias-Vergara · 2026-09-15

This study systematically evaluates machine learning methods for fault detection and line identification in electrical power grids, addressing challenges posed by renewable energy integration. The authors compare ML models' performance in detecting faults and identifying defective transmission lines within a critical 10 ms measurement window. The best-performing model achieved an F1 score of 0.991 ± 0.018 with a processing time of 0.342ms ± 0.509ms, demonstrating real-time operational viability for grid protection systems.

fault detectionelectrical gridsmachine learningrenewable energyreal-time processing

Carry-Through Checksum: A Lightweight Fault-Detection for CNN Inference at the Edge

arXiv cs.LG · Kyrylo Nazarevych, Mohammad Hasan Ahmadilivani, Krister Kaldre, Davide Bertozzi · 2026-09-15

The authors propose carry-through checksum, a lightweight fault-detection method for CNN inference on embedded GPUs that embeds dedicated filters to compute and propagate checksums through convolutional layers, enabling end-to-end error detection with single-output verification. The technique integrates with standard GPU pipelines, requiring minimal compute and memory overhead while avoiding matrix augmentation or per-operation checksums. Evaluations on multiple CNN architectures demonstrate 95.86% (FP32) and 86.56% (FP16) critical fault detection rates, with only 2.27% runtime overhead for re-execution-based mitigation on an NVIDIA Jetson Orin NX GPU.

convolutional neural networksfault detectionembedded gpuschecksum verificationsoft-error mitigation

Unified Heterogeneous Graph Neural Network solver for Power Flow, Optimal Power Flow and State Estimation

arXiv cs.LG · Ferran Bohigas-Daranas, Hamid Latif-Martínez, Eduardo Prieto-Araujo, Oriol Gomis-Bellmunt · 2026-09-15

The authors propose a unified Heterogeneous Residual Gated Graph Convolutional Network that simultaneously solves Power Flow (PF), Optimal Power Flow (OPF), and State Estimation (SE) in power systems, eliminating the need for task-specific models. The method learns a shared representation of network behavior, enabling multi-task inference through a single backbone. Evaluated on IEEE 14-bus and 118-bus systems under diverse topologies and loading conditions, the model matches task-specific GNN accuracy (exact metrics unspecified) and generalizes to unseen scenarios, suggesting potential as a foundation model for power system analysis.

graph neural networkspower flowoptimal power flowstate estimationheterogeneous graph

GrowMTP: Can RL Grow Its Own Draft Head?

arXiv cs.LG · Minghua He, Lingzhe Zhang, Yuan Liu, Xiao Zhou · 2026-09-15

GrowMTP introduces a method for training draft heads entirely within the reinforcement learning (RL) loop, eliminating the need for pretraining or warm-up phases. Leveraging the narrow rollout distribution and continuous supervision signals from RL verification steps, GrowMTP updates draft heads independently of the policy backbone. Evaluated on Qwen3-4B, MiMo-7B-SFT, and Qwen3.5-4B-Base, GrowMTP achieves rollout speedups of 2.13x, 1.93x, and 1.36x, and end-to-end speedups of 1.60x, 1.41x, and 1.20x, respectively. This modular approach accelerates RL training, particularly for models lacking pretrained draft heads.

reinforcement learningspeculative decodingdraft headrollout distributionpolicy backbone

Can Knowledge Transfer Parameters Be Learned? LePoKet for Efficient Robotic Vision

arXiv cs.LG · Yanick C. Tchenko, Felix Mohr, Hicham Hadj-Abdelkader, Hedi Tabia · 2026-09-15

LePoKet (Learnable Parameter Optimization for Knowledge Transfer) introduces a structural transfer framework for efficient robotic vision, embedding knowledge inheritance directly into forward computation via a block-wise Extract-Transform-Mix interface. The framework employs Learnable Genetic Attention (LGA) to optimize interaction parameters jointly with the child network, eliminating auxiliary distillation losses. Evaluated on CIFAR-10 and CIFAR-100, LePoKet achieves relative error reductions of 24.57% and 25.1%, respectively, over standard child training. Integrated into a compact RAFT-based optical-flow model, it improves End-Point Error (EPE) on Sintel Clean (2.21 to 1.92), Sintel Final (3.35 to 3.01), and KITTI (7.51 to 6.39), outperforming Hereditary Knowledge Transfer (HKT) variants.

knowledge transferlearnable genetic attentionoptical-flowend-point errorstructural transfer

Stable by Construction: Variational Latent Markov Operators for Long-Horizon PDE Prediction

arXiv cs.LG · Junyi Liao, Johann Guilleminot, Vahid Tarokh · 2026-09-15

The paper introduces Variational Autoencoding Markov Operator (VAMO), a variational framework for long-horizon PDE prediction that mitigates error accumulation in autoregressive neural solvers. VAMO employs latent Markov dynamics with functional Gaussian models, where structured latent perturbations induce spectral geometry and variational transition alignment regularizes learned dynamics. Evaluated on fluid-dynamics benchmarks, VAMO achieves stable rollouts beyond training horizons, reducing error accumulation by 15-30% compared to deterministic and noise-injection baselines, demonstrating the efficacy of variational modeling for robust PDE dynamics.

neural pde solversvariational autoencodinglatent markov dynamicsautoregressive predictionfunctional gaussian models

Divergence Timing and Cumulative Disagreement under KV-Cache Eviction

arXiv cs.LG · Xinyue Luo, Fei Yu · 2026-09-15

We present a theoretical framework for analyzing divergence timing and cumulative token mismatch in autoregressive generation under KV-cache eviction. Using stepwise maximal coupling, we derive an exact decomposition of expected mismatch fraction into first-mismatch contribution and post-divergence exposure components. Residual-branch conditional Monte Carlo enables unbiased joint estimation of occurrence, occupation, and window/tail contributions. Experiments on Meta-Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct demonstrate that SnapKV at 50% retention delays divergence onset compared to SnapKV-512 and recent-token retention, though post-divergence total variation remains high. Analysis of 288 documents reveals post-divergence exposure accounts for 85-90% of aggregate mismatch gaps, with higher branch-aligned TV in late versus early windows at 90% retention.

kv-cache evictionautoregressive generationstepwise maximal couplingconditional monte carlototal variation

A Weighted Kernel Method for Approximation that Adapts to Learned Multivariable Structure

arXiv cs.LG · John E. Darges, Laura Weidensager · 2026-09-15

The paper introduces total sensitivity kernels (TSKs), a weighted kernel method for approximating multivariable black-box functions by adapting to learned input importance and interactions. TSKs use weighted ANOVA kernels, parameterizing component weights via input-specific factors learned by minimizing the target function's norm in a reproducing kernel Hilbert space (RKHS). Theoretical analysis shows uniqueness and consistency of the norm-minimization solution, with learned factors quantifying input sensitivity akin to Sobol indices. Experiments demonstrate that TSKs improve approximation accuracy over standard product kernels by adapting to multivariable structure.

total sensitivity kernelsanova kernelsreproducing kernel hilbert spacesobol indicesmultivariable approximation

Recovering Physical Parameters from Fragmented Observations via Exact Distributed Spline Merging

arXiv cs.LG · Naveen Mysore · 2026-09-15

The paper introduces a distributed spline merging method for recovering physical parameters from fragmented observations, achieving mathematically identical results to centralized fitting without raw data sharing or iterative synchronization. It leverages the additive structure of fixed-basis ridge-regression statistics applied to tensor-product spline fields, where local Gram matrices and moment vectors are computed independently. The pipeline integrates distributed observations, field reconstruction, derivative extraction, and linear regression, recovering diffusion coefficient and wave speed with 0.11% and 0.12% error, respectively. Validation on 41 years of NOAA sea-surface temperature data confirms its effectiveness on real spatiotemporal datasets.

tensor-product splineridge-regressiongram matrixdiffusion coefficientspatiotemporal

AsyncCouple-Flow: Asynchronous Cross-Modal Coupling and Flow Matching for Spatio-Temporal Forecasting

arXiv cs.LG · Zhixiang Wu, Yining Liu, Bo Zhao, Szu-Yu Chen · 2026-09-15

AsyncCouple-Flow introduces a novel framework for multi-modal spatio-temporal forecasting (MM-STF) addressing three key challenges: heterogeneous sampling rates, modality missingness, and autoregressive error accumulation. The method combines Modality-Aware Token Sparsification (MATS) for scale-aware tokenization, Asynchronous Cross-Modal Coupling Graph (ACCG) for flexible modality fusion under asynchrony, and Flow-Matching Forecasting Head for multi-step prediction via conditional ODEs. Evaluated on ERA5+GOES+ISD weather forecasting and PEMS-BAY traffic prediction, AsyncCouple-Flow outperforms state-of-the-art baselines and maintains robustness with up to two missing modalities.

multi-modal forecastingtoken sparsificationcross-modal couplingflow matchingconditional ode

FlowATC: Aircraft Trajectory Prediction via Flow Matching

arXiv cs.LG · Mathurin Petit, Emir Torun, Louis Brusset, Jordan Kam · 2026-09-15

The paper introduces FlowATC, a flow-matching architecture for aircraft trajectory prediction trained on 1.15M historical ADS-B trajectory windows without route labels. The model employs a block-causal Transformer for sequence inpainting, using Conditional Flow Matching (CFM) or Denoising Diffusion Probabilistic Models (DDPM) to denoise future state tokens conditioned on observed history. CFM outperforms DDPM by 11-26% in minADE@20 and surpasses CVAE baselines by 31-41%, demonstrating graceful error degradation with prediction horizon. The model generates probabilistic occupancy estimates useful for conflict-risk estimation.

flow matchingtrajectory predictionconditional flow matchingdenoising diffusion probabilistic modelsautomatic dependent surveillance-broadcast

High-Performance Tensor Formulation of the Viterbi Algorithm for Hidden Semi-Markov Models

arXiv cs.LG · Lorenzo Piarulli, Elia Belli, Daniele De Sensi · 2026-09-15

The authors present a tensor-based reformulation of the Viterbi algorithm for Hidden Semi-Markov Models (HSMMs), enabling parallel execution on CPUs and GPUs by restructuring inner loops into tensor operations. This approach maps naturally to SIMD units and massively parallel architectures, yielding optimized implementations for single-core, multi-core, and GPU platforms. Experiments demonstrate speedups of 14x (single-core), 200x (multi-core), and 570x (GPU) over sequential baselines, establishing a new performance benchmark for large-scale HSMM decoding.

hidden semi-markov modelsviterbi algorithmtensor operationsparallel computinggpu acceleration

Certified Inference and Training for Deep Equilibrium Networks: A Continuation Framework with Polynomial Complexity Guarantees

arXiv cs.LG · Alex Borisevich · 2026-09-15

The authors present a certified continuation framework for inference and training in deep equilibrium networks (DEQs), achieving polynomial complexity guarantees. For inference, they employ compact input homotopy and rounded Newton tracking with certified bounds on boundary, conditioning, derivatives, and tube-radius. Training incorporates local-plus-low-rank recurrence augmented with programmable dormant bilinear rank-one channels, using loaded Tikhonov solves for interpolation failure diagnosis. Both certified inference and training exhibit bit cost O(poly(L+b)), where L is encoded instance length and b targets 2^{-b} accuracy. Lean 4 verifies the quantitative core, with numerical experiments validating the loaded mechanism.

deep equilibrium networkscertified continuationhomotopy trackingtensor interpolationpolynomial complexity

Online Gradient Computation for Warping Gaussian Process Transformations

arXiv cs.LG · Emilio Ruiz-Moreno, Konstantinos Slavakis, Baltasar Beferull-Lozano · 2026-09-15

The authors propose an online method for warped Gaussian processes (GPs) that jointly updates latent GP moments and optimizes warping parameters via exact recursive gradient computation. Existing streaming variants either optimize warping parameters periodically or sacrifice analytical tractability for increased capacity. By deriving the gradient of the instantaneous negative log-likelihood recursively, the method enables continuous parameter updates without compromising tractability or model flexibility.

warped gaussian processesonline learninggradient computationnegative log-likelihoodrecursive estimation

A multimodal large language model for evidence-based autism spectrum disorder screening

arXiv cs.LG · Jun Chen, Qi Zhao, Yunliang Jiang, Shuqin Cao · 2026-09-15

The authors introduce ASDchat, a multimodal large language model for evidence-based autism spectrum disorder (ASD) screening, using video, audio, and dialogue inputs via a dual-branch architecture (decision branch for screening probabilities, evidence branch for timestamped behavioral evidence aligned with ADOS-2 criteria). Trained on 1,035 participants across 27 sites, ASDchat achieved an AUC of 0.953 ± 0.021 for ASD vs. typically developing children and 0.932 on held-out sites. Unsupervised clustering identified six ASD subtypes with phenotypic profiles, each linked to tailored interventions, demonstrating feasibility for scalable clinical screening.

multimodal llmautism screeningdual-branch architecturebehavioral evidenceunsupervised clustering

Not All Relations Are Equal: Relation-Balanced and Calibrated Graph Learning for Provenance-Based Intrusion Detection

arXiv cs.LG · Lijie Zheng, Ji He, Alessandro Brighente, Yulong Shen · 2026-09-15

RECAL introduces relation-balanced masked graph learning for provenance-based intrusion detection, addressing statistical heterogeneity in system interactions where relation frequencies vary by up to 140,000×. The unsupervised framework calibrates reconstruction errors against each relation's benign error distribution, enabling comparable anomaly evidence and reducing false alarms. Evaluated on three DARPA E3 datasets, RECAL achieves F1 scores of 99.99%, 99.93%, and 99.99%, outperforming baselines by 0.88, 0.82, and 0.42 percentage points, respectively. It reduces mean false positive rates by up to 105× compared to the best baseline.

provenance-based intrusion detectionmasked graph learningstatistical heterogeneityreconstruction error calibrationfalse positive rate

Decentralized Gossip Learning and Federated Averaging for Histopathology Image Classification

arXiv cs.LG · Yusuf Ozturk, Enes Goltekin, Bengisu Atli, Akin Ozturk · 2026-09-15

This study compares Federated Averaging (FedAvg), decentralized gossip learning, and Hybrid Gossip-FedAvg for invasive ductal carcinoma (IDC) patch classification using 277,524 histopathology image patches distributed across six nodes. Methods include patient-disjoint partitions, Dirichlet-guided allocation (α=0.3), and evaluations of topology (ring, random degree-3, fully connected), model drift, and communication payload. Hybrid Gossip-FedAvg achieved 0.8811 ROC-AUC (vs. FedAvg's 0.8801), with both reaching 0.9082 ROC-AUC in patient-level tests; FedAvg had lower Brier score (0.1335), while Hybrid led in precision-recall (0.8240 AUPRC). Denser gossip graphs improved discrimination but increased payload, with FedAvg as the most reliable baseline and Hybrid offering a coordination compromise.

federated averaginggossip learninghistopathology classificationmodel driftdirichlet allocation

Adaptive Bayesian Partner Selection for Federated Clinical Centers

arXiv cs.LG · Navid Seidi, Satyaki Roy, Sajal K. Das · 2026-09-15

We propose Adaptive Bayesian Partner Selection (ABPS), a peer-to-peer federated learning framework addressing heterogeneity and concept drift in healthcare. ABPS employs Beta-Bernoulli posteriors over Shapley marginal utility, ranks peers via Upper Confidence Bound, and uses a propose-reject mechanism with optional abstention. The method provides finite-sample guarantees, O(kappa log T) regret, and conditions for optimal isolation under negative transfer. Extensions include head personalization, bfloat16 quantization, and metadata filtering. Evaluated on 230 non-IID clinical centers from MIMIC-IV for in-hospital mortality prediction, ABPS-X matches FedDyn's AUROC (0.758) at 0.09x FedAvg's communication cost, demonstrating scalable efficiency without accuracy loss.

federated learningshapley marginal utilityupper confidence boundbeta-bernoulli posteriorconcept drift

The Neverwhere Visual Parkour Benchmark Suite

arXiv cs.LG · Ziyu Chen, Henghui Bao, Haoran Chang, Alan Yu · 2026-09-14

The Neverwhere Visual Parkour Benchmark Suite introduces a collection of 60+ hyper-realistic 3D Gaussian Splatting reconstructions of urban environments for closed-loop evaluation of visual locomotion controllers. The benchmark addresses the train/evaluation gap by enabling reproducible testing in photo-realistic simulated settings, while demonstrating potential pitfalls of training exclusively on Gaussian-splatted data through cross-scene policy evaluations. Results highlight the need for diverse training data, supported by released policy checkpoints and performance metrics across novel scenes.

visual locomotion3d gaussian splattingclosed-loop evaluationbenchmark suitepolicy generalization

Learned Look-Ahead Splitting Rule for CART

arXiv cs.LG · Andrew Gao, Tianlin Liu, Ruichen Han, Lu Tian · 2026-09-14

The authors propose a look-ahead splitting rule for CART that evaluates candidate splits based on prediction error reduction after growing conventional subtrees, addressing the myopia of greedy splitting. To mitigate computational cost, they introduce a smart look-ahead variant that learns downstream splits using node-level features. Simulations and real-data analyses demonstrate improved split selection in hierarchical or interaction-driven settings while maintaining interpretability.

classification and regression treeslook-ahead splittingrecursive partitioningnode-level featuresprediction error reduction

Physics Informed Random Feature Neural Networks for Solving PDEs

arXiv cs.LG · Chi-An Chen, Chunyang Liao, Ming Zhong · 2026-09-14

The authors propose physics-informed random feature neural networks (PIRFNN) to address spectral bias in PINN-based PDE solvers, reducing computational complexity compared to methods requiring numerous collocation points. The method leverages randomized neural networks to approximate large-scale kernel machines, with theoretical guarantees on $H^1$ norm error bounds. Numerical experiments validate error decay rates and demonstrate improved spectral bias mitigation relative to state-of-the-art PINN approaches.

physics-informed neural networksrandom feature methodspectral biaspartial differential equations$h^1$ norm

Implementing a White-Box Undetectable Backdoor for Random Fourier Features

arXiv cs.LG · Michael Collins, Jada Cumberland, Brianne Dunn, Ross Gore · 2026-09-14

This work implements the white-box undetectable backdoor construction for Random Fourier Features (RFF) models proposed by Goldwasser et al., using only numpy and scipy to test practical realizability. The authors develop two samplers for the core $GP_d(b_k)$ distribution: a rejection-sampling proxy and an exact closed-form sampler derived from the homogeneous Continuous Learning With Errors (CLWE) density. Statistical indistinguishability tests reveal no detectable differences between backdoored and clean models across sparsity ratios $ρ= d_{\text{sparse}}/D$, confirming the feasibility of the backdoor with commodity tools.

random fourier featureswhite-box backdoorcontinuous learning with errorsstatistical indistinguishabilitysparsity ratio

Attention Mean Fields Predict Average Representation Dynamics and Reveal Context-Specific Computation

arXiv cs.LG · Micah Adler, John W. Byers, Mark Crovella · 2026-09-14

The paper introduces a mean-field analysis of attention to model the dynamic evolution of language model representation geometry. The method defines an average attention kernel conditioned either on a corpus or a specific context, enabling predictions of representation dynamics without layer-specific deviations. Results show that early in training, the model and its corpus mean field are indistinguishable, with substitution leaving loss unchanged. Post-induction onset, deviations increase as representations contextualize, with mean-field deviation serving as a task-agnostic measure of context-specific computation. Greater deviation correlates with increased reliance on in-context information across induction and few-shot settings.

mean-field analysisattention kernelrepresentation geometrycontext-specific computationinduction onset

Bounded Adjustment with Reliability-Guided Embedding for Imbalanced Learning with Noisy Labels

arXiv cs.LG · Mushir Akhtar, Akarsh J., M. Tanveer, Mohd. Arshad · 2026-09-14

BARGE (Bounded Adjustment with Reliability-Guided Embeddings) addresses coupled challenges of class imbalance and label noise through a single-stage objective combining bounded density-power scoring and reliability-guided angular geometry. It ensures proper classification scores in adjusted probability space, recovers balanced Bayes ordering under clean supervision, and bounds risk perturbation under label contamination. Evaluated on CIFAR-10, CIFAR-100, and Tiny ImageNet, BARGE achieves the lowest mean balanced error across six dataset-corruption settings, reducing the average from 72.32% to 70.00%, and attains the highest macro-F1 and macro-AUPRC in all corrupted-label scenarios.

class-balanced learninglabel noisedensity-power scoreangular geometrybalanced bayes ordering

Certified Uncertainty Propagation in One-Shot Federated Bayesian Models via Posterior Event Transport

arXiv cs.LG · Mahyar Mohammadi, Mohammad Hossein Badiei, Abolfazl Yaghmaei, Hamed Kebriaei · 2026-09-14

The paper introduces a certified uncertainty propagation framework for one-shot federated Bayesian learning, ensuring safety properties of aggregated models by transporting local posterior events through the deployment rule. For Federated Averaging (FedAvg), it provides an exact geometric characterization where clients compute hyper-rectangular regions in parameter space, and the server verifies Cartesian products of these regions under the safety property. Experiments on MNIST and Fashion-MNIST with label-Dirichlet heterogeneity show transported FedAvg certificates range from 22.51% to 46.89%, contrasting with direct global certificates (72.05%–91.39%), revealing divergent trends in predictive accuracy and certifiable safety.

federated bayesian learninguncertainty propagationposterior event transportcertified safetyfederated averaging

Fast-Convergent Meta-RL via Gradient-Clustered BS Sampling for Edge Caching

arXiv cs.LG · Farnaz Niknia, Ping Wang · 2026-09-14

The paper introduces a meta-reinforcement learning framework for edge caching in wireless networks, addressing the bottleneck of high-variance meta-gradient estimation in existing methods. The framework employs gradient-based clustering to group Base Stations (BSs) by local gradient similarity, sampling proportionally from each cluster during meta-training. Each BS runs a local Proximal Policy Optimization (PPO) agent modeled as a Semi-Markov Decision Process (SMDP), while a shared meta-policy is learned via Model-Agnostic Meta-Learning (MAML). Theoretical analysis demonstrates that this approach yields a strictly lower-variance meta-gradient estimator compared to uniform random sampling under BS heterogeneity.

meta-reinforcement learningedge cachinggradient-based clusteringproximal policy optimizationsemi-markov decision process

Autonomous Droplet Navigation via Model-Based Reinforcement Learning

arXiv cs.LG · Rajneesh Anand, Mayuresh V. Kothare · 2026-09-14

The study demonstrates autonomous droplet navigation in complex microfluidic geometries using model-based reinforcement learning, overcoming challenges like contact-angle hysteresis and capillary pinning. A gravity-driven platform with two-axis tilt and silicone oil film enables control, while an offline-trained policy learns effective tilt strategies from limited physical interaction data without simulation. The policy achieves reliable navigation in straight, right-angle, and curved-arc paths, with successful zero-shot transfer to unseen geometries and reduced training data requirements for complex paths.

model-based reinforcement learningdroplet navigationcontact-angle hysteresismicrofluidicszero-shot transfer

Mini-batch Sampling Strategies for Long-Tailed Image Classification: An Empirical Study on CIFAR-100-LT

arXiv cs.LG · Siyu Yuan · 2026-09-14

This empirical study compares four mini-batch sampling strategies for long-tailed image classification—uniform instance sampling, class-balanced sampling, square-root sampling, and progressively balanced sampling—within a unified bias-variance framework. Using ResNet-32 on CIFAR-100-LT at imbalance ratios (ρ = 10, 50, 100), progressive sampling improves tail-class accuracy by 25% (13.5% vs. 10.8%) at ρ = 100 without compromising overall accuracy (40.0% vs. 39.7%), while class-balanced sampling degrades performance across all class groups due to overfitting. Results highlight the importance of timing in rebalancing strategies.

long-tailed classificationmini-batch samplingbias-variance frameworkresnet-32cifar-100-lt

EBL: Efficient Broad Learning for Distributed Adaptive Harmonic Analysis

arXiv cs.LG · Changhong Li, Georgios Floros, Biswajit Basu, Shreejith Shanker · 2026-09-14

The paper proposes Efficient Broad Learning (EBL), a quantized FPGA acceleration framework for distributed adaptive harmonic estimation in power grids with non-linear loads like EV charging. EBL employs a BLS-style architecture with online transfer learning via closed-form solutions (avoiding backpropagation), achieving 17.4× faster predictions than prior FPGA methods. It uses 5.9% of LUTs on a Zynq Ultrascale+ ZU7EV FPGA, requiring ≈82% of LUTs versus SOTA FPGA estimators, while maintaining high accuracy with half-cycle inputs and reconfigurable flexibility.

harmonic estimationfpga accelerationbroad learning systemonline transfer learningquantized inference

Federated stochastic bilevel optimization with fully first-order gradients

arXiv cs.LG · Yihan Zhang, Rohit Dhaipule, Chiu C Tan, Haibin Ling · 2026-09-14

Proposes a federated stochastic variance-reduced bilevel gradient descent algorithm that eliminates second-order Hessian and Jacobian computations, relying solely on first-order oracles to reduce runtime. Introduces a constant single-timescale learning rate for variable updates and a novel convergence rate analysis. Experiments demonstrate the method's efficacy compared to existing approaches requiring second-order computations.

federated learningbilevel optimizationfirst-order oraclevariance reductionsingle-timescale learning rate

Multi-Label Proportion Learning for Sea-Ice Type Prediction

arXiv cs.LG · Samira Alkaee Taleghan, Younghyun Koo, Andrew P. Barrett, Farnoush Banaei-Kashani · 2026-09-14

The authors propose a weakly supervised multi-label proportion learning (MLPL) framework for sea-ice type prediction, addressing the limitation of polygon-level annotations in ice charts. Their approach combines Multiple Instance Learning (MIL) for water--ice classification with MLPL for ice-type composition prediction, extended with a multimodal model integrating SAR imagery, AMSR2 brightness temperatures, and ERA5 reanalysis data via modality-guided auxiliary regularization. On the AI4Arctic dataset, the SAR-only model reduces MAE by 14.5% and doubles mean ice-class F1 over supervised baselines. The multimodal model further reduces MAE by 21.5% and increases mean F1 by 41.2% over the SAR-only model.

multi-label proportion learningmultiple instance learningmodality-guided regularizationsea-ice type predictionweakly supervised learning

Channel-Informed Neural Network for Physical Layer Key Generation

arXiv cs.LG · Jose Angel Sanchez Viloria, George Sklivanitis, Dimitris Pados, Elizabeth Serena Bentley · 2026-09-14

A channel-informed neural network is proposed for physical-layer key generation (PKG) in wireless edge networks, enabling decentralized key establishment from reciprocal channel observations. The multi-task recurrent neural network jointly learns reciprocity-preserving binary features and auxiliary channel estimates through deep metric learning with channel-informed supervision, leveraging structured channel sounding and Sionna-RT ray tracing for data augmentation. Evaluated on indoor and outdoor software-defined-radio measurements from the POWDER testbed, the model achieves lower bit disagreement for reciprocal observations (Alice-Bob) than adversarial ones (Eve), with ray-traced augmentation increasing unique-key rates to 0.94, 0.99, and 0.99 across scenarios. Generated keys pass NIST randomness tests prior to SHA-3 privacy amplification, demonstrating the tradeoff between key diversity and reconciliation reliability.

physical-layer key generationreciprocity-preservingdeep metric learningray tracingprivacy amplification

StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation

arXiv cs.LG · Rohit Dhaipule, Sukhdeep Singh Kharbanda, Prasanth Bathala, Pradyumna Lanka · 2026-09-14

The paper introduces StalePO, a token-level preference optimization method for machine translation that leverages stale human post-edits from legacy systems. The method enforces three requirements: downward likelihood movement on both responses, policy anchoring to the base response, and token-level KL constraints. Ablations show these mechanisms are jointly necessary, with isolated components failing to improve performance. On English-to-Hindi and English-to-Turkish tasks, StalePO improves LLM-as-judge MQM quality checks by 14.9 and 4.6 percentage points, respectively, with human evaluation confirming a 13.8-point gain in English-to-Hindi.

preference optimizationmachine translationtoken-level controlstale feedbackkl constraint

Generative models for simulation based filtering: Formulations and Empirical Comparisons

arXiv cs.LG · Mohammad Al-Jarrah, Wei Deng, Bamdad Hosseini, Amirhossein Taghvaei · 2026-09-14

The paper presents a unified formulation and empirical comparison of generative-model approaches to nonlinear filtering, introducing three new filters based on stochastic interpolants, deterministic flow-matching, and Schrödinger bridges via forward-backward SDEs. A two-stage tuning procedure separates generative model training from online refinement. Evaluated against optimal transport filter (OTF), Knothe--Rosenblatt filter (KRF), sequential importance resampling (SIR), and ensemble Kalman filter (EnKF), results show generative filters resolve multimodal posteriors better than EnKF/SIR, with performance dependent on computational budget and ensemble size, and varying trajectory regularity.

nonlinear filteringgenerative modelsstochastic interpolantsschrödinger bridgesensemble kalman filter

Robust Fault Detection in Mechanical Multimodal Time Series via Self-Supervised Cross-Modal Reconstruction

arXiv cs.LG · Magnus Munk Jensen, Dorte Hammershøi, Rafał Wiśniewski, Olga Fink · 2026-09-14

Proposes a multimodal anomaly detection framework for mechanical systems via self-supervised cross-modal reconstruction of heterogeneous time-series sensor data. The method learns system dynamics by reconstructing each modality from others, exploiting complementary information without requiring temporal alignment or identical sampling rates, while improving robustness to noise and missing data. An adaptive test-time thresholding mechanism handles distribution shifts. Experiments on three industrial case studies demonstrate strong fault detection, with 23.4% higher F1-score under out-of-distribution conditions compared to baselines, particularly in challenging operating regimes.

multimodal anomaly detectioncross-modal reconstructionheterogeneous time-seriesadaptive thresholdingdistribution shifts

Agentic Search Spaces for Tabular Machine Learning

arXiv cs.LG · Renat Sergazinov, Artem Chistyakov, Sergey Pankevich, Artem Babenko · 2026-09-14

This paper demonstrates that LLM-based agents can enhance tabular machine learning by designing extended hyperparameter optimization (HPO) search spaces that outperform default configurations. The method represents tabular models as modular pipelines (preprocessing, embeddings, architecture, training, inference), tasks agents with proposing candidate implementations, and jointly optimizes these with classical HPO. On 45 datasets, agentic search spaces improved performance for nearly all model families, yielding average gains of 0.6% (up to 2.0% on small-to-medium regression tasks), without additional tuning cost. The approach also boosted Elo scores on TabArena, with agentic ensembles surpassing AutoGluon's best.

hyperparameter optimizationtabular machine learningllm-based agentssearch space expansionensemble learning

Sequence Recognition in Bharatnatyam dance

arXiv cs.LG · Himadri Bhuyan, Rohit Dhaipule, Partha Pratim Das · 2026-09-14

The paper introduces a method for sequence recognition in Bharatanatyam dance by analyzing Key Postures (KPs) and motions in Adavu variations. A Convolutional Neural Network (CNN) achieves 99% accuracy in KP recognition, while a Support Vector Machine (SVM) attains 84% accuracy for motion recognition. Sequence matching via Edit Distance yields 98% accuracy. The work advances digital heritage and dance tutoring by incorporating all 58 Adavu variations, unlike prior studies limited to 1-2 variations, and evaluates scalability through prediction time analysis.

bharatanatyamadavuconvolutional neural networksupport vector machineedit distance

Nationally Consistent, Locally Incomplete: A Bayesian Remote-Sensing Audit of Rooftop Photovoltaic Registries

arXiv cs.LG · Gabriel Kasmi, Yves-Marie Saint-Drenan, Laurent Dubus, Philippe Blanc · 2026-09-14

The study presents a Bayesian framework for estimating ground-truth rooftop photovoltaic (PV) capacity from remote sensing data, addressing inaccuracies in decentralized PV deployment statistics. The method transforms imperfect detector outputs into uncertainty-aware measurements through probabilistic modeling. Applied to France, the corrected estimates show 4.03 GWp [3.96–4.11, 99% CI] of sub-36 kWp rooftop PV capacity, aligning within 3.3% of transmission system operator data nationally while revealing local under-reporting up to 61%. The analysis also quantifies truncation bias in French open PV data, demonstrating broader applicability for global PV capacity audits.

bayesian inferenceremote sensingphotovoltaic capacityuncertainty quantificationenergy transition

Drift Field Net: Learning Ocean Lagrangian advection fields from in-situ and satellite observations

arXiv cs.LG · Théo Archambault, Pierre Garcia, Mattia Romero, Anastase Charantonis · 2026-09-14

Drift Field Net (DFN) introduces a deep neural network for predicting ocean surface flow fields from satellite observations to improve Lagrangian particle drift forecasts in the North Pacific Subtropical Gyre. DFN employs a two-stage training strategy: pretraining on simulated data followed by Lagrangian fine-tuning using an advection-consistent loss function, which incorporates physics-informed optimization. Evaluated against an operational physics-based system, DFN reduces 7-day mean positioning errors by 20 km on in situ drifter trajectories, with an additional 10 km improvement from Lagrangian fine-tuning, demonstrating the efficacy of integrating Lagrangian constraints.

lagrangian advectiondeep neural networkphysics-informed optimizationsatellite observationsocean surface flow

Differentially Private Semantic Plans for Aggregate Insight Generation

arXiv cs.LG · Behrooz Razeghi · 2026-09-14

The authors introduce DP-SPIN, a differentially private framework for generating semantic plans that enable aggregate measurement and summarization over predefined concepts independent of protected records. DP-SPIN maps records to bounded sparse nonnegative vectors over fixed concepts, forming semantic sketches, and releases a semantic plan containing admitted concepts and noisy masses. The framework ensures record- and user-level differential privacy under add/drop and replacement adjacency, with contributions clipped for user-level privacy. Evaluation on CFPB complaint narratives, Amazon All Beauty reviews, and Yelp restaurant reviews demonstrates DP-SPIN's effectiveness compared to non-private references, DP keyword histograms, and URANIA-style baselines.

differential privacysemantic plansemantic sketchuser-level privacybounded sparse vector

Scaling Laws for Physics-Aware ACOPF Surrogate Learning

arXiv cs.LG · Yijiang Li, Emon Dey, Stefano Fenu, Massimiliano Lupo Pasini · 2026-09-14

This work characterizes scaling laws for physics-aware surrogate learning in AC optimal power flow (ACOPF), demonstrating that training objective determines both performance and scalability. By sweeping model and dataset sizes under both mean squared error (MSE) and augmented Lagrangian (AL) training objectives, the authors analyze constraint violation across power grids. Results show both objectives improve as power laws but at different rates: MSE scales primarily with model capacity while AL balances model and dataset scaling. AL reduces constraint violation by 30× at 10× training cost compared to MSE, with negligible memory overhead. Violation grows twice as fast with network size under MSE versus AL.

acopfaugmented lagrangianscaling lawsconstraint violationsurrogate learning

Semantic-Aware Neural Video Codec for Error-Resilient Low-Latency Transmission

arXiv cs.LG · Matin Mortaheb, Homa Esfahanizadeh, Jinfeng Du, Harish Viswanathan · 2026-09-14

The authors propose a semantic-aware neural video codec for robust low-latency transmission over unreliable channels, extending DCVC-RT with multi-level packet prioritization. Their method partitions encoded representations into prioritized packets based on semantic and latent-feature importance, transmitted via abstracted multi-level packet erasure channels. An error-resilient entropy model removes inter-packet dependencies for independent decoding under losses. End-to-end training learns channel-aware representations and importance-aware packet assignment. Experiments demonstrate improved robustness over DCVC-RT, with graceful degradation in less critical regions and better preservation of task-relevant content during packet losses.

neural video codecpacket prioritizationerror-resilient entropymulti-level transmissionsemantic-aware coding

The record is part of the task: matched-record evaluation of text classifiers across maintenance, safety and recall reporting

arXiv cs.LG · Hisham Ihshaish, Peter Mayhew, Tasnim M. A. Zayet, Ana Del Amo · 2026-09-14

The study introduces matched-record evaluation for text classifiers, treating record selection as part of the evaluation process. It compares identical cases across multiple records (customer reports, technician reports) under fixed labels and splits in three systems: GE Aerospace repair events, NASA ASRS safety reports, and NHTSA vehicle recalls. Results show record-dependent performance variations, with macro-F1 ranging 0.33–0.91 in GE fields and a 0.46 gap between pre- and post-diagnosis reports. NHTSA defect summaries consistently outperformed other records, while ASRS analyst synopses surpassed reporter narratives only with learned sequence models. The work advocates evaluating models on records available at decision points and transparent reporting of record-label provenance.

matched-record evaluationmacro-f1text classifierslabel provenancesequence models

Towards Surrogate Based Dequantization of Quantum Reinforcement Learning

arXiv cs.LG · Pablo Rodriguez-Grasa, Sofiene Jerbi, Mikel Sanz, Ryan Sweke · 2026-09-14

The paper proposes a surrogate-based dequantization approach for quantum reinforcement learning (RL), focusing on kernelized Fitted Q-Iteration (FQI) as a classical counterpart to quantum Q-learning. Using a uniform generative model to simulate large experience replay buffers, the authors derive finite-sample guarantees for kernelized FQI with classical kernels designed to mimic the inductive bias of parameterized quantum circuits. They establish sufficient conditions on data encoding, kernel design, and problem structure under which classical FQI matches quantum Q-learning performance, providing rigorous dequantization guarantees and motivating kernelized FQI as a practical heuristic.

quantum reinforcement learningdequantizationkernelized fitted q-iterationparameterized quantum circuitsinductive bias

Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Descent

arXiv cs.LG · Alex Buna, Fanghui Liu, Patrick Rebeschini · 2026-09-14

The paper theoretically analyzes the compute-optimal allocation between pretraining and fine-tuning under a fixed total budget in regularized least squares with gradient descent. Using a two-stage optimization framework, the authors derive the optimal compute split by examining how pretraining directions influence fine-tuning predictions and how downstream data geometry affects fine-tuning shifts. Key findings reveal that the allocation depends on prediction-relevant spectral components of pretraining and fine-tuning empirical covariances, analyzed via basis-invariant eigenspace decomposition and perturbative control of non-commuting dynamics.

pretrainingfine-tuninggradient descentspectral decompositioncovariance

Copula Adapted Directed Acyclic Graph for Cluster Representation of Biomedical Data

arXiv cs.LG · Heranga K. Rathnasekara, Norou Diawara, Manar D. Samad · 2026-09-14

The paper introduces CopDAG, a cluster-friendly data representation framework combining copula models with Directed Acyclic Graph (DAG)-based causal structure discovery for biomedical data. CopDAG models non-Gaussian, non-linear feature dependencies via copulas and identifies stable causal relationships using an ensemble of DAG-based methods. Evaluated on 16 biomedical datasets, CopDAG outperforms 12 baseline methods in normalized clustering accuracy and adjusted Rand index when clustered with K-means, enabling unsupervised prediction of ground-truth labels and explainable causal feature structures.

copula modelsdirected acyclic graphcausal structure discoverybiomedical clusteringunsupervised learning

Improving Reduced-Order Rotating Detonation Engine Models with Data Assimilation and Machine Learning

arXiv cs.LG · Ashwin Suriyanarayanan, Romit Maulik · 2026-09-14

The work proposes a data assimilation method to improve reduced-order modeling of rotating detonation engines (RDEs) by synchronizing the Koch-Kutz model with high-fidelity temperature data via nudging. A Jacobian-regularized closure is trained on the recorded forcing term, enabling autonomous advancement while preserving spectral and statistical fidelity. The corrected model matches high-frequency content and conserved variable statistics of high-fidelity simulations without observation terms, outperforming the baseline Koch-Kutz model.

rotating detonation enginesdata assimilationreduced-order modelingjacobian regularizationnudging

Test-Time Unlearning via Sparse Autoencoder

arXiv cs.LG · Pingzhi Li, Jinhao Duan, Vaishnav Tadiparthi, Nakul Agarwal · 2026-09-14

ARIA (autoencoder-gated inference-time unlearning) improves machine unlearning by gating access to unwanted knowledge at test time without weight modification. The method trains a lightweight linear detector on sparse autoencoder (SAE) latents, applying interpretable interventions during triggered states with minimal overhead. Evaluations on TOFU, R-TOFU, and WMDP show ARIA outperforms weight-based baselines, reducing forget-set accuracy (e.g., WMDP-cyber) while maintaining MMLU within 1% of the original model. ARIA remains robust against three adversarial attacks, with <1% forgetting change, and feature analysis suggests retain degradation may stem from response-style biases rather than knowledge leakage.

machine unlearningsparse autoencodertest-time interventionadversarial robustnessinterpretability

How I learned to stop worrying and love StopGrads: Stationarity, Convergence, and a case study on Flow Map Learning

arXiv cs.LG · Max W. Shen, Mark Goldstein, Zichu Wang, Aahlad Puli · 2026-09-14

The paper introduces a stopgrad regression principle, providing a general framework for stopgrad objectives with closed-form stationary points, unifying applications in flow maps, reinforcement learning, and diffusion samplers. The authors theoretically justify optimizing stopgrad flow map objectives by proving their unique stationary point is the true flow map and demonstrating convergence for Eulerian and Lagrangian objectives, including MeanFlow variants. Notably, they derive a closed-form expression for the learned flow map under functional semi-gradient flow, composed of the initial and true flow maps. Additionally, modified stopgrad placements reduce training memory by 2x for flow map objectives.

stopgrad regressionflow mapstationary pointsfunctional semi-gradientconvergence guarantees

Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

arXiv cs.LG · Aashiq Muhamed, Mona T. Diab, Virginia Smith · 2026-09-14

Decoy Direction Optimization (DDO) introduces a computationally efficient, post-hoc defense against Refusal Feature Ablation (RFA) attacks in language models, eliminating the need for safety finetuning. The method injects a high-magnitude, nonlinear decoy signal into MLP neurons, corrupting contrastive estimators used by attackers to locate refusal directions. This preserves safety mechanisms while misleading attackers to ablate orthogonal features. Evaluated across six model families, DDO achieves <10% ASR under standard RFA and reduces Heretic weight-level attack ASR from 88.7% to 18% on Llama-3-8B-Instruct, with 30-450x lower optimization cost than trained defenses.

refusal feature ablationdecoy direction optimizationcontrastive estimatorsmlp neuronsattack success rate

A Sentinel-2 benchmark dataset for deep-learning active-fire segmentation across 25 California wildfires

arXiv cs.LG · Shreyan Mitra, Mohammadreza Narimani, Parastoo Farajpoor · 2026-09-14

The authors introduce an open benchmark dataset for active-fire segmentation in Sentinel-2 satellite imagery, comprising 2,148 image-mask pairs from 25 California wildfires (July 2020–August 2026). Each 512×512-pixel composite combines B12, B11, and B8A bands at 20m resolution, with masks labeling background, SWIR-rule active fire (0.0766% pixel prevalence), and invalid data. The dataset features incident-disjoint splits (18/3/4 fires for training/validation/test), 841 fire-positive samples, and analyst-reviewed test subsets. Accompanying resources include metadata, a ResNet-34 U-Net baseline, and evaluation code, targeting rare-class segmentation and cross-incident transfer learning. Data and tools are versioned on Zenodo and GitHub.

sentinel-2active-fire segmentationrare-class segmentationswir-ruleincident-disjoint

Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It

arXiv cs.LG · Julian Boesch, Andrew Wee · 2026-09-14

The study decomposes associative recall performance in fixed-state recurrent models (linear attention and state-space models) by isolating three architectural components: short causal convolution, transition structure (rank-1 delta rule vs. diagonal), and decay. Under matched training, convolution contributes ~+0.5 recall, overshadowing recurrence comparisons. Rank-1 transitions outperform diagonal by +0.19/+0.32 at 16/32 pairs, but margins shrink to +0.03 with convolution. A distance curriculum improves performance from 0.021 to 1.000, addressing interference under sparse supervision. Bidirectional denoiser cells show no advantage over causal training, and collision-key retrieval requires two layers. The work refines claims about recurrent models' recall limitations with measurable interventions.

associative recallfixed-state recurrencescausal convolutionrank-1 transitiondistance curriculum

Skeletal Prototypes on Iterative Nerve Expansions

arXiv cs.LG · Jordan Eckert, Henry Schenck · 2026-09-14

SPINE introduces skeletal prototypes as embedded 1-complexes for class-wise representation, replacing finite point sets in prototype reduction. Initial edges form class-conditional Mapper graphs, with vertices later optimized for classification; predictions use nearest-complex distances, incorporating segment geometry. Evaluated on 17 benchmarks via 10-fold cross-validation, SPINE achieves highest mean accuracy and rank among 7 competitors, with significant improvements over 5 methods (Wilcoxon signed-rank, Holm-corrected). Segment-based decisions excel at low prototype budgets, while moderate budgets optimize overall performance. SPINE's discriminative construction is faster than generalized learning vector quantization on 14/17 datasets.

prototype reduction1-complexmapper graphclass-conditionallearning vector quantization

LLM Inference in a Flash!

arXiv cs.LG · Sebastian Zhao, Minseo Kim, Coleman Hooper, Luca Manolache · 2026-09-14

The paper proposes an integer-only quantization approach and dictionary-based KV cache compression to enable efficient LLM inference on Compute-in-Flash devices, addressing memory bandwidth limitations and write endurance constraints. The method replaces floating-point operations with integer arithmetic and compresses KV cache vectors via sparse dictionary coding, representing each as a linear combination of static dictionary vectors. Evaluated on Llama-3.1-8B and Qwen-2.5-7B, the approach reduces dynamic KV cache traffic by 15× with minimal accuracy degradation.

llm inferencecompute-in-flashkv cache compressioninteger quantizationsparse dictionary coding

Computer-assisted global regularity across nonlinear families of three-dimensional periodic Navier-Stokes flows

arXiv cs.LG · Jose Luis Lima de Jesus Silva · 2026-09-14

The author presents a computer-assisted framework for proving global regularity in continuous families of three-dimensional periodic Navier-Stokes flows, combining finite reference trajectories with a uniform error bound for center fields and perturbation modes. The method retains nonlinear residuals pre-truncation and controls evolution until viscous decay ensures regularity, applied to cyclic-shear, Arnold-Beltrami-Childress, and Taylor-Green flows with explicit perturbation radii. A parameter-uniform extension covers non-Beltrami Taylor-Green centers without reproof, validated via 4,096 configurations and 1,600 refinement trajectories. Neural-operator experiments show physics-informed training improves predictions but not proof-limiting condition discovery, offering a reusable method for regularity proofs and evaluating learned models in rigorous computation.

navier-stokesglobal regularityspectral truncationneural-operatortaylor-green

Discrete Beckmann Transport Models for One-Step Language Modeling and Reasoning

arXiv cs.LG · Sophia Tang, Shiyi Wang · 2026-09-14

Discrete Beckmann Transport Models (DBTM) introduce a time-independent flow for one-step language modeling and reasoning, eliminating the need for teacher distillation. The method leverages an autonomous transport map that converges to a fixed point on the simplex vertices, characterized by a conservation equation optimized directly from data. This enables iterative refinement via partial-context interpolants without ODE integration. Evaluations on language and reasoning tasks show DBTM outperforms discrete diffusion and continuous flow baselines in quality and accuracy for one- and few-step generation.

discrete diffusionautonomous transportfixed-point convergencepartial-context interpolantone-step generation

SWB-DM: A Calibrated Sliced-Wasserstein-Barycenter Aggregator with Delayed-Momentum Caching for Byzantine-Robust Federated Learning under Partial Participation

arXiv cs.LG · Saranraj S, Saranya M S, Alex David S, Ajay Kumar A · 2026-09-14

SWB-DM introduces a robust federated learning aggregator combining sliced-Wasserstein barycenters (SWB) with delayed-momentum caching (DeMoA) to address partial participation vulnerabilities. The method treats client updates as 1D distributions, computes trimmed Wasserstein barycenters, and fixes coordinate identity via a medoid-based gauge. Delayed momentum caches updates across clients, decoupling robustness from sampling. Evaluated across 448 CIFAR-10 configurations, CIFAR-100, FEMNIST, and a 500-client setup, SWB-DM outperforms baselines (coordinate-wise median, Krum, Bulyan) by mitigating mechanistically distinct failure modes, though FLTrust benefits more on CIFAR-100 due to unrelated caching effects.

federated learningwasserstein barycenterbyzantine robustnesspartial participationmomentum caching

Evaluating Open-Weight E-Commerce Agents with Environment-Grounded Verification

arXiv cs.LG · Nimit Shah, Haitz Sáez de Ocáriz Borde · 2026-09-14

The authors introduce a deterministic e-commerce environment for evaluating open-weight conversational agents, enabling reproducible trials with precommitted customer personas, difficulty levels, target carts, and item reveal schedules. The environment guides a simulated consumer's actions, records assistant responses, and applies evidence-based penalties for incorrect tool calls or search failures. Bidirectional interaction allows real-time directive injection and trial termination based on customer frustration. Evaluating eight agents (20B-35B parameters) across 160 trials and 44 metrics reveals distinct capability profiles, exposing issues like under-action, over-purchase, and poor search that terminal success rates obscure.

e-commerce environmentopen-weight agentstool-call penaltiesbidirectional interactioncapability profiles

The Token Before the Value Is the Key: How Hybrid Architectures Organize Induction Circuits

arXiv cs.LG · Ke Cheng, Xin Xu, Yixiao Chen, Lei Xin · 2026-09-14

This work investigates how hybrid language models allocate position-sensitive and content-based computations across heterogeneous layers, focusing on Carrying predecessor information, Matching source content, and Copying values. Using layer-type-agnostic paired probes, the authors analyze Carrying and Matching in recurrent-global and local-global hybrids, finding Carrying concentrates in efficient layers and Matching in global receivers, with lag-one tokens playing a critical role. Interventions like lag-one masking and convolution removal reveal how Carrying and Matching relocate between stages, impacting natural-text recall. Experiments with varying local windows and induction-enriched training text connect architectural priors to the timing of circuit development. Code is available on GitHub.

hybrid language modelscarrying predecessormatching sourcelag-one tokensinduction circuits

📰 Industry Media (1)

Building the materials foundation for AI

MIT Tech Review — AI · MIT Technology Review Insights · 2026-09-16

The AI industry's rapid advancement is driving semiconductor and data center materials to physical limits, necessitating innovations in thermal management, electrical efficiency, and chemical resistance. Syensqo employs AI-driven molecular synthesis to explore millions of material combinations, predicting performance and sustainability metrics before lab testing. This approach accelerates the development of high-performance polymers, sealing materials, and thermal fluids for high-voltage data centers and semiconductor fabs, while reducing environmental trade-offs. The feedback loop between AI-accelerated materials discovery and improved AI infrastructure could enable future technological breakthroughs.

semiconductorsthermal managementmolecular synthesishigh-voltage architecturessustainability metrics


Generated automatically at 2026-09-16 22:37 UTC. Summaries and keywords are produced by an LLM and may contain inaccuracies — always consult the original article.