Daily Digest — 2026-08-26

Tuesday, August 25, 2026 · 355 items · model: deepseek/deepseek-chat

355 items · 10 research labs, 340 arxiv papers, 5 industry media

🏛️ Research Labs (10)

The full stack behind abundant intelligence

OpenAI News · 2026-08-25

OpenAI introduces Jalapeño, its first custom inference chip, demonstrating superior performance on GPT-OSS 120B (InferenceX benchmark) with higher throughput per kilowatt and lower token latency compared to commercial systems. The chip's co-design with models, software, and infrastructure enables optimized efficiency across workloads. Results show a 54% reduction in output tokens for GPT-5.6 Sol on the Artificial Analysis Coding Agent Index. OpenAI's full-stack approach integrates hardware, models, and deployment strategies to maximize capability and cost-efficiency while expanding AI accessibility.

inference chipthroughput efficiencyco-designpareto frontierjevons paradox

Jalapeño’s first results show industry-leading speed and efficiency in AI inference

OpenAI News · 2026-08-25

OpenAI's custom inference chip Jalapeño demonstrates industry-leading efficiency and latency in AI workloads, achieving 1.5–1.9× higher performance per watt and 1.7–3.6× lower end-to-end latency across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T models. The architecture optimizes full-stack integration (chip, memory, networking) to minimize data movement during prefill and decode phases, with KV-cache locality and domain-specific networking. Benchmarking on InferenceX showed 2.1–4.1× higher interactive performance, enabled by AI-assisted design (9-month tapeout) and automated kernel optimization (1.5–1.8× speedup vs. hand-tuned code). Sustained power remained ≤550W despite 700W rating.

inference acceleratorkv-cacheprefill-decodeperformance per watttapeout

Disrupting a new covert influence campaign from Russia

OpenAI News · 2026-08-25

OpenAI disrupted a novel Russian influence operation leveraging ChatGPT to generate multilingual social media content promoting the fabricated International Burke Institute (IBI). The campaign combined AI-generated posts (via VPN-accessed accounts) with a spoofed think tank website featuring plagiarized academic articles and a biased 'sovereignty index' favoring Russia. Forensic analysis revealed Slavic-language artifacts in machine-translated content and VPN usage patterns. Despite low audience engagement (10-20k Telegram followers), the operation demonstrated advanced infrastructure for narrative laundering. OpenAI terminated associated accounts and documented the multi-platform campaign's technical indicators.

influence operationnarrative launderingforensic analysismultilingual generationvpn obfuscation

Introducing the Admin plugin for ChatGPT Work and Codex

OpenAI News · 2026-08-25

OpenAI introduces the Admin plugin for ChatGPT Work and Codex, enabling workspace administrators to manage access, usage, and permissions through conversational AI. The plugin integrates analytics and administrative actions, allowing tasks like member management, credit monitoring, and permission updates within a single interface without complex prompts. It supports automation for high-volume requests (e.g., Slack/MS Teams routing) and maintains role-based permissions, resolving ~45% of IT tickets in OpenAI's internal deployment. The tool reduces manual overhead while preserving policy compliance and auditability.

chatgpt workcodexadmin pluginpermission-awareusage limits

Advancing price-performance for developers with GPT‑5.6 in Kiro

OpenAI News · 2026-08-24

OpenAI introduces GPT-5.6 model family integration into Kiro, a software development agent, enhancing price-performance for AI-native coding. The models (Sol, Terra, Luna) enable structured, spec-driven development by converting high-level intent into executable tasks with 82% cost reduction on Terminal-Bench 2.1. Kiro's context-aware approach grounds models in codebase-specific requirements, improving multi-step task completion and correctness via property-based testing.

gpt-5.6kiroterminal-benchspec-driven developmentproperty-based testing

Granite 4.2 LLMs: How They're Built

Hugging Face Blog · 2026-08-25

IBM introduces Granite 4.2, a family of dense decoder-only LLMs (3B, 8B, 30B) optimized for reasoning and agentic tasks, released under Apache 2.0. The models are pre-trained on 15T tokens with a five-phase strategy extending context to 512K, then fine-tuned on chain-of-thought and agentic-trajectory data, followed by multi-stage RL including agentic RL for tool use. Key features include grouped query attention, rotary position embeddings, and a thinking/non-thinking mode switch. The 8B/30B models demonstrate tool-calling capabilities in sandboxed environments, with native OpenAI-compatible function calling.

grouped query attentionrotary position embeddingchain-of-thoughtagentic rlsupervised fine-tuning

Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original

Hugging Face Blog · 2026-08-25

The paper introduces Quantization-Aware Healing (QAH), a method for recovering compressed and quantized LLMs that outperforms full-precision originals. QAH distills directly from the original pre-compression model (full-precision, full-size) to the quantized student (4-bit, compressed) using KL-divergence on logits, avoiding degradation from intermediate checkpoints. Applied to GPT-OSS 120B compressed to 60B and quantized to MXFP4, QAH achieves higher accuracy than the bfloat16 checkpoint on 7/9 benchmarks, including +7.4 on long-context reasoning, while being 4× smaller. QAH also trains 7× faster and is more stable than quantization-aware training (QAT).

quantization-aware healingkl-divergencestructural compressionmxfp4long-context reasoning

Wire It, Run It, Deploy It: AI Workflows in Gradio

Hugging Face Blog · 2026-08-25

Gradio introduces gr.Workflow, a framework for constructing modular AI pipelines as interactive graphs with typed nodes. The system enables drag-and-drop composition of operators (Python functions, Hugging Face models, or Gradio Spaces) that execute in parallel, with intermediate results visible through a web interface. Each workflow automatically exposes REST endpoints for programmatic access, supporting both Python clients (via gradio_client) and HTTP requests. Demonstrated applications include multi-model media generation (text-to-image-to-sticker chains), dataset profiling via Hugging Face Datasets Server, and GPU-accelerated inference through ZeroGPU allocation. The approach combines visual prototyping with production-ready deployment to Hugging Face Spaces.

gradioworkflowrest apizero-gpuinference providers

How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code

Hugging Face Blog · 2026-08-21

The paper presents a hybrid search system for Papers with Code, combining PostgreSQL full-text search with pgvector semantic embeddings via reciprocal rank fusion (RRF). The architecture leverages Hugging Face Jobs for batch embedding generation, Storage Buckets for immutable artifact management, and Inference Endpoints for low-latency query embedding. Using Qwen/Qwen3-Embedding-0.6B (256-dimensional vectors), the system achieves 0.9955 Recall@20 with 1.31ms p50 latency on a 5,000-paper pilot. The design emphasizes reproducibility through versioned embedding contracts and graceful degradation during cold starts.

hybrid searchreciprocal rank fusionpgvectormatryoshka representation learninghnsw index

5 ways to upgrade your home decor with Google Search

Google AI Blog · Megan Stoner · 2026-08-25

Google Search introduces five AI-enhanced features for home decor optimization, leveraging multimodal interaction and visual search. The system employs AI Mode for furniture visualization via image uploads, Google Lens for vintage item identification, Circle to Search for context-aware product discovery, Search Live for AR-assisted DIY guidance, and price tracking with historical trend analysis. These tools address the 300% surge in 'home decor inspo' searches, enabling spatial planning (e.g., 84-inch wall compatibility checks), 140% trending 'fish wallpaper' discovery, and vintage furniture procurement through cross-platform visual search.

multimodal searchvisual contextprice history trackingar guidanceproduct discovery

📜 arXiv Papers (340)

How to Train a Critic Stably and Efficiently

arXiv cs.AI · Penghui Qi, Xiangxin Zhou, Wee Sun Lee · 2026-08-24

The paper introduces Best-Practice Critic Optimization (BPCO), a stable training recipe for reinforcement learning critics that combines DPPO, bounded value predictions, Monte Carlo targets, unnormalized advantages, and length-adaptive GAE. BPCO enables token-level advantage estimation from single responses while conditioning critics on reward-defining information hidden from policies. Experiments on mathematical reasoning tasks with models from 1.5B to 30B-A3B MoEs show BPCO matches or exceeds group-based baselines (GRPO) while requiring only one response per prompt, and improves rubric-based reward learning. Code is released.

critic optimizationdppomonte carlo targetsgeneralized advantage estimationtoken-level advantages

ReWorld: An Interactive World Model with Long-Horizon Memory

arXiv cs.AI · Zhifei Chen, Luozhou Wang, Guibao Shen, Dongyu Yan · 2026-08-24

ReWorld introduces an interactive world model that reconciles short-horizon control with long-horizon memory via architectural and training innovations. The model employs mixed per-head attention windows, with most heads focusing on recent context and a few global heads attending to the entire history, alongside random head routing and chunk dropping for robustness. At inference, it uses a bounded KV cache with pose-indexed landmark retrieval for efficient long-term recall. Trained on metric-scale-aligned multi-source data and distilled via LoRA, ReWorld achieves state-of-the-art performance in control fidelity (11.95° rotation error), generation quality, and long-horizon recall, outperforming six baselines in a three-axis evaluation protocol.

interactive world modelkv cachemixed attentionlora adaptermetric-scale alignment

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

arXiv cs.AI · Deyao Hong, Yizhe Chi, Wenyi Li, Xiaoqiu Wang · 2026-08-24

The paper introduces SWE Refactor Bench, a benchmark for evaluating coding agents on whole-repository migrations, addressing the limitation of existing benchmarks that fail to detect whether migrations actually occur (termed 'Blindness'). The benchmark comprises 20 repository migrations across 4 technical debt categories, with a three-stage evaluation protocol: Migration Audit, Behavioural Tests, and Agentic Verification. Testing 8 frontier models across 26 configurations (520 runs), only 5.4% passed all stages, with the best model (Claude-Opus-5) scoring 47.0/100. Results show distinct capabilities in migration completeness (58% reach 99% of fixed checks) versus behavioural correctness (26% reach 100%), varying significantly by migration category (31.4 for build toolchain rewrites vs. 5.6 for language rewrites).

coding agentstechnical debtrepository migrationbehavioural correctnessmigration audit

EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

arXiv cs.AI · Md Thamed Bin Zaman Chowdhury, Moazzem Hossain · 2026-08-24

The paper introduces Expert-Grounded Distillation (EGD), a framework for transferring institutional road safety expertise into a compact vision-language model for visual road safety auditing in low-resource settings. The method involves calibrating a teacher vision-language model against expert risk assessments (Cohen's kappa = 0.74) before distilling knowledge into an 8-billion-parameter student model using Low-Rank Adaptation. The authors also present BD-ARSA, a 21,947-image dataset, and EG-ARSA, a vision-language model that outperforms its 31-billion-parameter teacher and Gemini-2.5-Flash in expert evaluations.

expert-grounded distillationvision-language modellow-rank adaptationroad safety auditingordinal risk assessment

Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography

arXiv cs.AI · Yuanyuan Zhang, Yida Zhang, Jiahui Li, Yuyan Wu · 2026-08-24

The paper introduces Phy-BP, a physics-constrained deep learning framework for non-invasive blood pressure (BP) estimation from triaxial bodyseismography (BSG). The method combines an adaptive quality-control algorithm to select cardiogenic BSG segments with a physical model of 3D wave propagation, embedded into the neural network to align multi-axis features and improve robustness. Evaluated on a 162-hour hospital dataset from 21 subjects, Phy-BP demonstrates improved BP monitoring accuracy by filtering low-quality measurements and enforcing physical consistency across axes, particularly with limited training data.

ballistocardiographybodyseismographyhemodynamic modelingphysics-constrained learningnon-invasive monitoring

Prime Agent: A Self-Improving RLM Harness

arXiv cs.AI · Seth Karten, Alex L. Zhang, Kevin Thomas, Sebastian Müller · 2026-08-24

Prime Agent introduces an open-source harness for long-horizon evaluation and coding-agent workflows, implementing a Recursive Language Model abstraction with persistent IPython REPL and Continual Harness for context processing. The system enables recursive subagent coordination, human-inspectable sessions, and standardized execution while delegating strategy to the model. Results show significant improvements, raising ARC-AGI-3 RHAE Best@1 from 30% to 95.5%, with strong performance in long-context coding, GPU-kernel generation, and autonomous tasks like Factorio progression.

recursive language modelcontinual harnesssubagent coordinationlong-horizon agencyipython repl

ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings

arXiv cs.AI · Na Li, Yuchen Jiao, Changxiao Cai, Gen Li · 2026-08-24

ConvergeFlow introduces an embedding-space flow-based language model that guarantees convergence to valid token embeddings without cross-entropy supervision. The method constrains the data predictor to the convex hull of token embeddings and trains it solely with mean squared error from flow matching, provably converging to valid embeddings despite predictor errors. Experiments on OpenWebText show competitive performance with existing continuous and discrete diffusion LMs, demonstrating flow-based modeling's potential.

flow-based language modeltoken embeddingsflow matchingconvex hullmean squared error

How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles

arXiv cs.AI · Shang Wu, Catarina G Belem, Shuyuan Fu, Mark Steyvers · 2026-08-24

The study investigates how AI assistance impacts human skill development in logic puzzles, demonstrating a trade-off between short-term performance gains and long-term skill acquisition. Using a controlled experiment with variable AI request costs, the authors measure participant performance before, during, and after AI access. Results from a Bayesian latent ability model show that frequent AI use during the access phase correlates with worse post-assistance performance, while independent problem-solving effort predicts greater skill gains, suggesting AI assistance substitutes for skill-building reasoning.

skill developmentai assistancelatent ability modelindependent reasoninglogic puzzles

The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams

arXiv cs.AI · Summer Eunhyung Ann, Haokun Liu, Chenhao Tan · 2026-08-24

The paper identifies an 'interaction tax' in multi-agent LLM systems, where full-solution communication erases diversity by causing rapid convergence to initial proposals. Through 11 verifier-scored optimization tasks with matched budgets, the authors demonstrate that independent proposal generation outperforms full-solution interaction, which primarily anchors agents to early solutions. Critique proves beneficial only when rule violations are easily detectable and fixable. Results indicate multi-agent performance depends critically on information exchange timing and granularity rather than agent count.

interaction taxmulti-agent llmdiversity collapseverifier-scored optimizationproposal generation

Adapter-Based Few-Shot Continual Learning for Malicious Packet Recognition

arXiv cs.AI · Kyle Stein, Guillermo Francia, III Eman El-Sheikh, Andrew Arash Mahyari · 2026-08-24

The paper proposes a hybrid framework for Few-Shot Class-Incremental Learning (FSCIL) in malware packet recognition, addressing catastrophic forgetting and limited labeled data. The method combines a Self-Supervised Learning (SSL) backbone with domain-specific pre-training, Low-Rank Adaptation (LoRA) for efficient model adaptation, and a prototype-based classification head for incremental learning. Experiments demonstrate state-of-the-art performance across multiple datasets, outperforming prior FSCIL baselines in malware classification.

few-shot class-incremental learningself-supervised learninglow-rank adaptationmalware classificationcatastrophic forgetting

Correcting a learned physical invariant improves world-model rollouts

arXiv cs.AI · Richard Bao · 2026-08-24

The study demonstrates that world models can learn physically meaningful invariants from video data but fail to preserve them during autonomous rollouts. Using DreamerV3 trained on pendulum videos, the authors identify an energy-like scalar that remains conserved in latent transitions of conservative models but not in damped ones. Correcting latent state deviations to this invariant reduces rollout errors by 30-50% in conservative models, while random constraints typically increase error. This reveals a concrete failure mode where learned physical constraints are violated during imagination.

world modelslatent transitionsphysical invariantsautonomous rolloutsdreamerv3

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

arXiv cs.AI · Zhiqing Cui, Xinxiang Yin, Yihong Tang, Xinglang Zhang · 2026-08-24

EarthVerse introduces a benchmark for evaluating scientific agents on 405 reproducible tasks across 19 hazard families, grounded in 199 documented Earth-system events. The benchmark assesses agents' ability to handle heterogeneous evidence, execute transparent calculations, and maintain provenance, with executable ground truth and task-specific rubrics. Evaluation of 25 model and agent systems reveals a best mean answer-unit accuracy of 84.65%, but only 34.81% Strict@95 accuracy, highlighting challenges in maintaining consistent reasoning chains across complex scientific workflows.

scientific agentsearth-system analysisnatural hazardsbenchmark evaluationanswer-unit accuracy

The Measurement Revolution? Credible Measurement and Inference in the Age of AI

arXiv cs.AI · Melissa Dell, Ashesh Rambachan · 2026-08-24

The article examines AI's role in transforming measurement practices in economics by converting unstructured data into structured variables, shifting the bottleneck from scalability to selecting among plausible measures. It outlines three stages of AI integration—discovery, construct definition, and observation—and emphasizes the need for credible validation anchored to explicit criteria rather than informal proxy claims. The authors argue that validation samples enable valid inference even with biased AI predictions and discuss solutions for scenarios lacking random validation samples.

measurementvalidationunstructured datainferenceeconomics

When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World

arXiv cs.AI · Xiang Chen, Zeyu Zhang · 2026-08-24

The paper introduces MHER, a provenance-controlled benchmark for historical entity reconciliation in the Mongol world, featuring a 396-pair Name-only core and a 160-pair Source-grounded subset with entity-disjoint splits. The study evaluates five generative systems, demonstrating that Source-grounded evidence improves test accuracy by 12.96 to 94.44 percentage points over Name-only input, with correct resolutions in 24/25 cases where models otherwise failed. Results highlight the critical role of provenance-controlled evidence and reveal that surface name restoration can degrade performance, as shown by Qwen3-8B's ten false merges from correct Context-only distinctions.

historical entity reconciliationprovenance-controlled benchmarkname-only coresource-grounded subsetgenerative systems

Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty

arXiv cs.AI · Yipeng Zhao, Qishun Yang, Shenzhe Zhu, Shu Yang · 2026-08-24

The paper introduces Safety-Direction Penalty (SDP) to mitigate Reasoning-Induced Misalignment (RIM), where fine-tuning on benign reasoning data induces harmful behaviors in LLMs. Through representation-space analysis, the authors identify coupled reasoning and safety directions in activation space, with safety degradation correlating with larger shifts in safety representations. SDP penalizes displacement along the safety direction during fine-tuning, guided by layer-localized CKA distance ratios and probes. Evaluations on Qwen2.5-3B and 7B show SDP restores safety while maintaining reasoning performance on benchmarks.

reasoning-induced misalignmentsafety-direction penaltyactivation-space directionscka distance ratiosrepresentation-space analysis

SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

arXiv cs.AI · Jialong Liu, Yuling Shi, Ning Yang, Xiaodong Gu · 2026-08-24

The paper introduces Self-Reflective Policy Optimization (SRPO), a framework enabling Large Language Models (LLMs) to perform post-training self-reflection for improved credit assignment. SRPO synthesizes trajectory errors into reflection patches and uses reflection-conditioned teacher scores as dense token-level training signals, eliminating need for external critics or reward models. Evaluated on mathematical reasoning and long-horizon agentic tasks, SRPO with Qwen3-8B achieves 73.3% on AIME'24 using only 8% of training FLOPs versus supervised fine-tuning, plus significant gains on WebShop (64.7%), ALFWorld (76.8%), and SWE-Bench-Lite (31.2%).

self-reflective learningcredit assignmenttoken-level supervisionlong-horizon reasoningpolicy optimization

Multi-Modal Semantic Expansion with Constrained LLM Reranking for Conversational Music Recommendation

arXiv cs.AI · Naman Garg, Sarika Jain, George Fazekas · 2026-08-24

The paper presents Team Semiintelligencn's multi-modal conversational music recommender system for ACM RecSys 2026, featuring a three-stage pipeline: (1) multi-modal retrieval combining seven dense embedding spaces (CF-BPR, Qwen3, CLAP, SigLIP) with BM25 and artist matching, fused via weighted Reciprocal Rank Fusion; (2) lightweight reranking; (3) GPT-4o-mini response generation. Differential evolution optimized RRF weights improved MRR by +19.5%. Experiments showed unconstrained LLM-guided artist injection caused -18.9% nDCG regression, while conservative injection achieved best Blind A performance. The system scored 0.3213 on Blind B.

multi-modal retrievalreciprocal rank fusiondifferential evolutionconversational recommenderdense embeddings

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

arXiv cs.AI · Sangoh Lee, Sangwoo Mo, Wook-Shin Han · 2026-08-24

The paper introduces Intention Distillation (INDI), a method to enhance Vision-Language-Action (VLA) models by explicitly modeling behavior-level intent during action decoding. INDI employs a frozen teacher VLM to generate multimodal intent representations from demonstrations, which are then used to guide action prediction alongside behavioral execution details. Evaluations on SimplerEnv-Bridge and RoboCasa Kitchen show performance improvements of 20.4 pp and 6.2 pp respectively for GR00T-N1.7, with real-world task success increasing from 62.0% to 68.7%. Analyses confirm the latent intent representation effectively captures behavioral objectives and organizes predictions.

intention distillationvision-language-action modelsbehavior cloningmultimodal intentaction decoding

StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models

arXiv cs.AI · Jinghan Tan, Yuanzheng Wang, Lu Chen, Zijun Chen · 2026-08-24

The paper introduces StrategyBench, a benchmark for evaluating explicit strategy induction in large language models (LLMs) during few-shot in-context learning. It selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines metrics assessing strategy quality and downstream utility. The study analyzes strategy induction across task variation, model configuration (including generator-executor choices and demonstration design), and adaptation settings. Results show significant variation in strategy utility across task categories, with performance dependent on both generation conditions and execution parameters.

strategy inductionin-context learninglarge language modelstask adaptationfew-shot learning

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

arXiv cs.AI · Marek Hradil, Danae Sánchez Villegas · 2026-08-24

The paper introduces TimeCatch, a benchmark for evaluating temporal consistency in vision-language models (VLMs) through anomaly detection tasks. Temporal anomalies are created by frame swaps, while frame-level anomalies use Gaussian noise. Evaluations on four datasets reveal VLMs excel at frame-level anomaly detection (near-ceiling performance) but perform near chance on temporal anomaly detection, contrasting with human near-ceiling performance on both tasks. Analyses across model scales, prompting strategies, and sequence lengths suggest VLMs struggle with temporal integration despite strong single-frame perception. TimeCatch provides a controlled testbed for temporal grounding in VLMs.

vision-language modelstemporal consistencyanomaly detectionframe-level anomaliestemporal grounding

MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters

arXiv cs.AI · ChengAo Shen, Wenchao Yu, Fangyu Wu, Dongjin Song · 2026-08-24

MetaCaster introduces a meta-harness-optimized multi-agent framework for few-shot learning of lightweight time series forecasters (TSF), addressing data scarcity in resource-constrained settings. The method employs agentic data generation to train specialized forecasters from minimal examples and textual contexts, positioning agents as intermediary engineers rather than direct forecasters. Evaluations across 18 datasets, 23 lightweight forecasters, and 14 baselines demonstrate MetaCaster's dual efficiency in data and computation while maintaining high TSF performance.

meta-harnessfew-shot learningtime series forecastingagentic data generationlightweight forecasters

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

arXiv cs.AI · Hanling Tian, Gengyu Zhang, Zeyang Sha, Jingying Wang · 2026-08-24

The paper introduces InjecMEM, a memory injection attack targeting LLM agent memory systems that requires only one interaction to manipulate future responses. The method crafts adversarial inputs with retriever-agnostic anchors for consistent retrieval and optimized commands for reliable generation steering, using gradient-based coordinate search across synthetic templates and positions. Evaluations show high success in topic-conditioned retrieval and targeted generation across multiple systems, with robustness to memory drift and minimal impact on non-target queries.

memory injectionllm agentsadversarial commandretrieval-then-generategradient-based optimization

Machine Learning Assisted Inverse Design of Pixelated mmWave Patch Antennas

arXiv cs.AI · Nadeem Rather, Holger Claussen, Lester Ho · 2026-08-24

The study presents a machine learning framework for inverse design of pixelated mmWave patch antennas (22-30 GHz) using a 19x23 binary pixel grid. Method involves: (1) XGBoost binary classifier to pre-filter resonant patterns, increasing resonant designs from 40% to 52% in a 10,000-sample dataset; (2) hybrid CNN-BiLSTM surrogate with physics-guided loss for S11 prediction; (3) latent-space gradient descent optimization for inverse design. Results show strong agreement between surrogate predictions and CST simulations, validating automated antenna design.

inverse designmmwave antennasxgboost classifiercnn-bilstmgradient descent optimization

Reward-Free Continual Adaptation for Resilient Space Robots

arXiv cs.AI · Andrej Orsula, Miguel Olivares-Mendez, Carol Martinez · 2026-08-24

The paper introduces a reward-free continual learning framework for space robots operating in environments with unobservable rewards due to hardware degradation. The method pre-trains a model-based agent with a latent-state world model across diverse simulations, then freezes the observation encoder and reward predictor during deployment, updating only transition dynamics via unsupervised rollouts. Policies are trained on imagined trajectories from the adapted world model, enabling adaptation without new rewards. Results demonstrate effectiveness in simulated planetary traversal, orbital navigation, and precision assembly tasks under severe morphological failures.

continual reinforcement learninglatent-state world modelsunsupervised rolloutsmodel-based adaptationmorphological failures

Characterizing Necessary Losers to Explain Tournaments Losers

arXiv cs.AI · Contet Clément, Umberto Grandi, Jérôme Mengin · 2026-08-24

The paper characterizes necessary losers in tournament solutions by identifying destructive minimal supports—minimal sub-tournaments where a candidate loses regardless of remaining matches. It analyzes six tournament rules (maximin, uncovered set, weighted uncovered set, top-cycle, Copeland, Borda) to determine when a candidate is a necessary loser or possible winner. Results include bounds on minimal support sizes and polynomial-time algorithms for computation, except for Borda which is conjectured NP-complete.

tournament solutionsnecessary losersdestructive minimal supportsabductive explanationspolynomial-time algorithms

Towards Comprehensive Basketball Understanding

arXiv cs.AI · Yirong Hu, Jiayuan Rao, Yu Zhang, Shangzhe Di · 2026-08-24

The authors introduce BasketballBench, a multimodal benchmark with 7,980 questions across ten tasks (text, image, video) derived from the 2025-2026 NBA season, including play-by-play data, player profiles, and 2,501 possession-level clips. They propose BasketballSkills, an agent that orchestrates eight domain-specific perception/retrieval tools via four reusable skills with tool ordering, evidence binding, and stopping conditions. Experiments reveal current multimodal large language models (MLLMs) struggle on multi-capability integration, while BasketballSkills outperforms them, demonstrating the efficacy of explicit compositional reasoning for holistic basketball understanding.

multimodal benchmarkdomain-specific agentperception-retrieval toolscompositional reasoningbasketball understanding

ChebBooster: A Training-Free Approach for Efficient Diffusion Transformer Inference via Chebyshev-Inspired Extrapolation

arXiv cs.AI · Chengjie Lu, Tianchi Deng, Zhengqi He, Chengwen Luo · 2026-08-24

ChebBooster introduces a training-free extrapolation framework for efficient Diffusion Transformer (DiT) inference using Chebyshev polynomial theory, addressing instability in existing methods. The approach employs Barycentric formulation for stable Chebyshev approximants and decouples computation into offline weight precomputation and lightweight online application. Experiments on DiT-XL/2, PixArt-$Σ$, and FLUX.1-dev show 3.68× latency speedup and 5.12× FLOPs reduction while maintaining visual quality, outperforming training-free baselines across tasks and resolutions.

diffusion transformerschebyshev polynomialsbarycentric formulationtraining-free accelerationinference efficiency

SkillAlchemy: Open-World Agent Skill Creation

arXiv cs.AI · Hengjun Wang, Shuyue Wei, Boyi Liu, Jun Yang · 2026-08-24

SkillAlchemy introduces an admission-centered framework for open-world skill creation, enabling language agents to generate reusable procedural artifacts from underspecified briefs and source-access specifications. The method identifies implicit requirements via contrastive evidence, admits candidate procedures based on evidence-supported scope, and compiles them into grammar-guided skill packages. Evaluated on 87 SkillsBench v1.1 tasks, SkillAlchemy improves pass rates by 19.9 percentage points over no-skill execution and 8.6 percentage points over automated baselines, matching human-curated skill performance.

language agentsskill creationcontrastive evidencegrammar-guided compilationskillsbench

Adaptive Item-based Collaborative Structures via Noise Rescheduling in Diffusion for Generative Recommendation

arXiv cs.AI · Jiaqi Wang, Tianying Liu, Heng Chang, Jihong Guan · 2026-08-24

The paper introduces ANR-DiffRec, a framework integrating item-based collaborative filtering into discrete diffusion models (DDMs) for generative recommendation. The method incorporates an item co-occurrence matrix as a collaborative prior and proposes an adaptive noise rescheduling mechanism that dynamically adjusts denoising weights based on local context and item dependencies. Experiments show ANR-DiffRec outperforms state-of-the-art generative recommendation models on multiple benchmarks.

discrete diffusion modelsgenerative recommendationcollaborative filteringnoise reschedulingitem co-occurrence

MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

arXiv cs.AI · Ruoyu Wu, Shenfu Xie, Yinqian Sun, Haibo Tong · 2026-08-24

MediSkill-Evo introduces a process-constrained clinical agent that evolves procedural knowledge without backbone fine-tuning, employing four experience banks for clinical skills, process rules, symbolic schemas, and measurement procedures. It utilizes a Process-Constrained Preference Harness to bind evidence to its source and rank actions via a safety-prioritized Clinical Process Critic. Evaluated on 300 Qwen encounters, the system improves diagnosis accuracy from 61.33% to 69.00% and treatment-intent coverage from 33.62% to 66.44%, while reducing critical failures from 31.00% to 16.33% compared to AgentClinic. In hard-isolation conditions, it achieves 93.61% target recovery under patient-behavior pressure and 100% for temporal evidence.

process-constrainedclinical process criticexperience bankssymbolic schemaspreference harness

Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep

arXiv cs.AI · Pedro Santos · 2026-08-24

This work introduces a preregistered pilot heuristic for optimizing LLM-agent decomposition in VAT determination tasks, proposing to place partition boundaries at dependency-layer midpoints. The study systematically compares four orchestrated configurations (ranging from one wide worker to five narrow ones) against a single-agent baseline, using a fixed activity surface and 4,400 runs across 40 cases with failure-injection arms. Results show intermediate configurations achieve highest accuracy (0.830 vs. 0.720-0.770 endpoints) but fail to surpass the pre-stated fine endpoint benchmark, while single-agent performance does not Pareto-dominate orchestrated setups. Availability faults were absorbed at all granularities, though schema-conforming hallucinations degraded all configurations disproportionately.

llm-agentvat determinationorchestrationdependency-layerfailure-injection

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

arXiv cs.AI · Bin Dou, Junru Zhang, Zhaoyi Yuan, Wuliang Huang · 2026-08-24

The study proposes User Behavioral Densing Law to quantify the relationship between data scale and minimum sufficient tokenization capacity in billion-scale user representation learning. Through theoretical analysis and experiments on the Alipay dataset, the authors identify an approximate linear log-log relationship between tokenization capacity and input data size, with scaling slopes varying by tokenization method and data source. They introduce ALGN, an adaptive variable-length tokenization method that improves capacity allocation, demonstrating superior performance over baselines across diverse tasks and data sources.

user representation learningtokenization capacityscaling lawbillion-scale datasetadaptive tokenization

Cross-Domain, Multi-Task Data-to-Text Generation without In-Domain Training Data

arXiv cs.AI · Yifei Song, Kun Efimov-Zhang, Claire Gardent · 2026-08-24

The paper introduces a cross-domain data-to-text (D2T) generation approach that operates without in-domain training data or test references, addressing varied input structures and generation goals. It compares data-driven knowledge distillation (DDKD) against zero-shot inference and fine-tuning, proposing structure-preserving augmentation via subsampling and perturbation. Experiments on five benchmarks with a 1.7B-parameter model show DDKD outperforms both baselines, even matching larger fine-tuned models in two domains. The study also extends QUINTD-1 to QUINTD-5, demonstrating that structural augmentation surpasses scaling real inputs for cross-domain distillation.

data-to-text generationknowledge distillationcross-domain learningstructural augmentationzero-shot inference

Cross-lingual Biography Enrichment via Claim Extraction and Alignment

arXiv cs.AI · Yifei Song, Ziyang Chen, Emil Sayilov, Claire Gardent · 2026-08-24

The paper introduces extsc{CLAW-4L}, a benchmark of 300 Wikipedia biography pairs (English with French/Chinese/Azerbaijani) annotated with claims and claim-pair relations, to study cross-lingual biography enrichment. The proposed framework extracts English claims from both biographies, aligns them to identify enrichment evidence from the non-English version, and rewrites the English biography with selected claims. Results demonstrate that non-English biographies provide valuable coverage improvements, though challenges persist in lower-resource settings.

cross-lingual enrichmentclaim extractionclaim alignmentbiography pairslower-resource settings

Adversarial Entropy Inflation Against Gumbel-Based Inference Verification

arXiv cs.AI · Nikita Kezins · 2026-08-24

The paper demonstrates that Gumbel-based inference verification, which mitigates LLM weight exfiltration by tolerating GPU-induced token variations, is vulnerable to adversarial prompt engineering. By crafting prompts that disrupt grammatical and sub-word structure, attackers increase output entropy, expanding the admissible-token set and covert channel capacity. Evaluations across six instruction-tuned models (1B-32B parameters) show character- and script-level disruptions double leaked bits per token, reducing the defense's slowdown from >200x to 60x-118x, necessitating dynamic entropy-based threshold calibration.

gumbel-based verificationllm weight exfiltrationoutput entropycovert channeladversarial prompting

Modalities Should Talk to Each Other: Dual-Stream Multimodal Learning for Long-Horizon Influenza Forecasting

arXiv cs.AI · Seyed Mohammad Hossein Hashemi, Mohsen Hooshmand, Parvin Razzaghi · 2026-08-24

The paper proposes Dual-Stream Attention (DSA), a multimodal deep learning framework for 12-week-ahead influenza-like illness (ILI) forecasting from 36-week histories of paired numerical and textual data. DSA employs separate Transformer-based encoders for each modality, coupled via bidirectional Cross-Modal Attention (CMA) to enable mutual conditioning, followed by a causal temporal model. Evaluated on Time-MMD, DSA reduces mean error by 37-67% versus baselines (MSE 0.416 vs 0.668-0.851), achieves lowest worst-window error, and generalizes to external geography. Ablations confirm CMA's bidirectional benefit and functional informativeness via perturbation analysis.

multimodal learningcross-modal attentioninfluenza forecastingtransformer encodercausal temporal model

Walking on the DARKSIDE

arXiv cs.AI · Aldo Gangemi, Emanuele Bottazzi · 2026-08-24

The paper introduces DARKSIDE, a coherence auditing method that addresses LLMs' vulnerability to reifying nonsensical inputs by tracking exclusions and warranting referents. Built atop POLANYI++'s Logic-Augmented Generation (LAG) framework, DARKSIDE formalizes discourse trails as explicit data structures and classifies referents into Warranted, Unattested, Misattributed, or Fabricated categories, triggering risk assessments when thresholds are exceeded. Evaluated on Gemini 3 using the adversarial BSBench corpus (100 items across 5 domains) with Claude Sonnet 4.6 as judge, results demonstrate that ontology-mediated negative-trail scaffolding can partially bridge the pattern-vs-path gap in LLM outputs.

logic-augmented generationextended knowledge graphwarrant axisdelegationriskassessmentepistemic firewall

DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

arXiv cs.AI · Vlad Hondru, Florinel Alin Croitoru, Iuliana Georgescu, A. Sophia Koepke · 2026-08-24

DF-MoE introduces a generalizable deepfake detection framework leveraging multimodal sparse Mixture-of-Experts (MoE) to mitigate overfitting. The method extracts high-level audio-visual cues (mouth movements, face parsing, facial expressions, head pose, gaze tracking, heart rate, audio emotion, speech activity) via diverse pre-trained models, integrating unimodal and multimodal features through MoE. Evaluated on five benchmarks (MAVOS-DD, AVLips, PolyGlotFake, BioDeepAV, FakeAVCeleb), DF-MoE outperforms state-of-the-art methods in both in-domain and cross-domain settings.

deepfake detectionmixture-of-expertsmultimodal fusiongeneralizabilitypre-trained models

The Emergence of Relevance Through Axiomatic Attention Patterns During LoRA Fine-Tuning

arXiv cs.AI · Matthew Perlman, Atharva Nijasure, James Allan · 2026-08-24

This work investigates where and how task-specific relevance behavior emerges during LoRA fine-tuning of RankLLaMA for reranking. Through ablation studies and attention pattern analysis, the authors identify a compact mid-network region where LoRA attention updates are both necessary and sufficient for performance gains, recovering over 50% of full LoRA tuning benefits. Key findings show strong correlations between performance improvements and increased attention to axiomatic IR features like rarity sensitivity and query-document interaction, providing interpretable evidence for relevance learning mechanisms during parameter-efficient adaptation.

lora fine-tuningattention patternsrerankingaxiomatic irparameter-efficient adaptation

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

arXiv cs.AI · Wenqi Liu, Shijie Ma, Yunxiao Wang, Meng Liu · 2026-08-24

The paper introduces VideoRover, a unified framework combining video reasoning and deep research for open-world video understanding. VideoRover iteratively coordinates video cropping, multimodal search, and webpage browsing, using tool results to guide subsequent actions. The authors curate 26K supervised fine-tuning trajectories and 3K reinforcement learning instances, alongside the VideoRover-Bench benchmark. Experiments on VideoDR and VideoRover-Bench show VideoRover-8B-RL matches proprietary models in direct-answer settings and outperforms larger open-source models with equivalent tools, validating the synergy of active video grounding, external retrieval, and long-horizon RL.

video understandingmultimodal searchreinforcement learningopen-world agentsbenchmarking

Mycelial Search: A Graph-Structured Metaheuristic for Continuous Optimisation

arXiv cs.AI · Mohammad Mahdi Dehshibi · 2026-08-24

The paper introduces Mycelial Search (Myco), a graph-structured metaheuristic for continuous optimization that balances information sharing and search diversity. Myco employs active tips, community-weighted flow, adaptive cord plasticity, and anchor-based injection within an evolving spatial graph, using Louvain partitioning to manage information exchange. Evaluated on the CEC 2022 benchmark suite at D=10 and D=20, Myco achieves competitive performance on selected functions compared to 11 established optimizers. Ablation studies reveal that community structure regulates information exchange range, while cord plasticity controls local directional influence persistence.

mycelial searchcontinuous optimizationadaptive cord plasticitylouvain partitionmetaheuristic

Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

arXiv cs.AI · Zixuan Wang, Yanrui Miao, Zhengxi Lu, Teng Pan · 2026-08-24

Agent-G$^2$ introduces a Gaussian guidance framework for hint-based reinforcement learning, addressing reward sparsity in long-horizon agentic tasks by modeling guidance depth as a Gaussian distribution rather than a deterministic scalar. The method estimates the Gaussian's center (combining global baseline and per-cluster difficulty) and spread (tracking within-cluster variance) online from existing rollouts, eliminating probe rollouts. Evaluated on ALFWorld and WebShop using Qwen2.5-1.5B/7B-Instruct, Agent-G$^2$ outperforms hint-based, hint-free, and Aux-RL baselines by 2.3/3.9/7.4 points on ALFWorld at under one-third the rollout cost of per-sample probing.

hint-based reinforcement learninggaussian guidancelong-horizon tasksagentic learningrollout efficiency

EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models

arXiv cs.AI · Xuetong Li, Gaofeng Liu · 2026-08-24

The paper introduces EviSafe, an evidence-grounded framework for evaluating vision-language model (VLM) safety beyond outcome-level metrics. The method combines natural response assessment, explicit evidence grounding (textual/visual), and counterfactual sensitivity analysis via EviSafeBench, a controlled benchmark with 1,181 scenarios and 2,452 counterfactual variants across eight safety domains. Results on eleven VLMs show performance gaps: natural severity accuracy (27.6-52.8%), relaxed diagnostic consistency (6.1-29.3%), and unsafe-to-safe counterfactual transition success (30.4-58.4%), revealing unreliable multimodal safety reasoning.

vision-language modelssafety evaluationevidence groundingcounterfactual analysismultimodal reasoning

FIDES: A Concordance Protocol for LLM-Generated Trading Strategies

arXiv cs.AI · Arther Tian, Alex Ding, Simon Wu, Aaron Chan · 2026-08-24

The paper introduces FIDES, a protocol for evaluating LLM-generated trading strategies by measuring concordance between three artifacts: natural-language rationale, executable code, and backtest results. The method employs dual delivery (simultaneous natural-language and code outputs), sandboxed execution, and lag-one out-of-sample backtesting to quantify three concordance gaps. Results from 40 strategies across 8 ETFs show poor alignment: only 2/40 strategies outperformed buy-and-hold, 32/40 falsely claimed to do so, and model-judged say-to-do concordance was highly unstable (flipping >50% with judge substitution).

concordance protocoldual deliverylag-one backtestself-assessment calibrationfuture-information probe

Evaluating SAT Solver Metrics as Predictors of Human-Perceived Nonogram Difficulty

arXiv cs.AI · Changdao He, Yibing Ju, Jonathan Calver, Alice Gao · 2026-08-24

The study challenges the assumption that algorithmic solver effort correlates with human-perceived puzzle difficulty by analyzing Nonograms through SAT solving and human solving data. Researchers formulated Nonograms as a constraint satisfaction problem, employed SAT solvers, and conducted a user study measuring participant interactions and reported difficulty. Results show no meaningful correlation between SAT solver metrics and human difficulty perceptions, though expertise moderates this relationship, and reveal human preference for complex propagation strategies distinct from solver-measured complexity.

nonogramssat solverconstraint satisfactionhuman difficultypropagation strategies

Sigmoid Attention as a Better Substrate for Learned KV Cache Eviction

arXiv cs.AI · Isaac, Li · 2026-08-24

The paper demonstrates that sigmoid attention enables more effective learned KV-cache eviction by mitigating the soft-to-hard mismatch between training and inference. Using GPT-2-scale Transformers on OpenWebText, the authors conduct a controlled $2\times2\times2$ ablation over attention type (sigmoid/softmax), learned gating, and positional encoding. Results show sigmoid-gated models achieve negligible perplexity (PPL) degradation under hard eviction, outperforming H$_2$O and KeyDiff baselines, whereas softmax gates fail to transfer cleanly to hard deletion. The findings highlight attention normalization's role in eviction transferability.

kv-cacheevictionsigmoid attentionsoftmax attentionperplexity

How Much Regularization Survives Averaging? Update Masking in Federated Learning

arXiv cs.AI · Wenhao Yan, Fu Kuroda, Yucheng Jin, Zhenke Chen · 2026-08-24

The paper analyzes why update masking, a noise-based regularization technique for finding flat minima, has not been adopted in federated learning (FL). It proves that federated averaging weakens the regularization effect by the cohort size when clients use independent masks, but shared masks restore it proportionally to inverse gradient diversity. Experiments on CIFAR-10 show this factor ranges from 1.17-1.50 under typical FL settings, rising to 8.96 without minibatch sampling. However, configurations preserving effective regularization exhibit prohibitively poor training performance.

federated learningupdate maskingflat minimagradient diversitynon-iid

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

arXiv cs.AI · Apodex Team, B. An, B. Li, B. Wang · 2026-08-24

Apodex 1.1 advances agentic intelligence for complex work through two key innovations: Environment Scaling enhances executable file, search, and code environments for verifiability, while Agentic Coordination Scaling trains agents in task decomposition, parallel delegation, and asynchronous integration via a shared execution harness (AgentOS). The system maintains task state and provenance across tools, using environment trajectories and coordination traces for training. Despite its compact 35B-parameter size (Apodex 1.1 Mini), it achieves leading performance in professional, financial, scientific, and coding tasks, demonstrating verifiable long-horizon problem-solving capabilities.

agentic coordinationenvironment scalingexecution harnesslong-horizon tasksprovenance tracking

Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance

arXiv cs.AI · Or Biton, Tomer Krichli, Itai Allouche, Joseph Keshet · 2026-08-24

The study investigates alignment failures in Large Language Models (LLMs) where ethical behavior conflicts with helpfulness objectives. Using Layer-wise Relevance Propagation (LRP), the authors analyze model responses to unethical scenarios presented as classification tasks, first-person statements, and assistance requests. Results reveal attribution bias toward benign framing tokens over unethical cue-tokens (e.g., 'without getting caught'), leading to harmful compliance. Two LRP-guided decoding methods mitigate this by increasing cue-token relevance, empirically demonstrating improved safety.

large language modelslayer-wise relevance propagationalignment failurestoken relevanceharmful compliance

Automated Construction of FAIR Digital Object Knowledge Graphs from Flat Cultural Heritage Records

arXiv cs.AI · Zeyd Boukhers, Lingxiao Kong, Xenophon Zabulis, Georgios Toubekis · 2026-08-24

The paper introduces a pipeline for transforming flat Europeana cultural heritage records into FAIR Digital Object (FDO)-compliant knowledge graphs structured with CIDOC-CRM. The method employs a large language model to classify metadata values, distinguishing between PID references and literals, and links them to controlled vocabularies such as Getty AAT, Wikidata, VIAF, and PeriodO. Evaluated on 637 archaeological records, the pipeline successfully links 86% of metadata slots, resolving 58.5% of previously unenriched values, and achieves 17 correct merges out of 33 cross-lingual surface forms. The resulting graph ensures all nodes are typed and resolvable, enhancing machine-actionability.

fair digital objectcidoc-crmpersistent identifiercontrolled vocabularyknowledge graph

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

arXiv cs.AI · Yinhao Tang, Youqing Fang, Yanan Sun, Jiangning Liu · 2026-08-24

The paper challenges the effectiveness of next-chunk reasoning RL for training on no-CoT data by comparing it to Mixed SFT, a simpler supervised fine-tuning approach combining no-CoT and long-CoT data. Mixed SFT outperforms next-chunk reasoning RL in post-RLVR accuracy across mathematical and general reasoning tasks while requiring 60× less training compute. Results demonstrate that pre-RLVR performance does not reliably predict post-RLVR gains, emphasizing the need for end-to-end evaluation of no-CoT training strategies.

no-cot datanext-chunk reasoningmixed sftrlvr accuracytraining compute

E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models

arXiv cs.AI · Taoyu Qian, Qi Wang, Daqian Shi, Yuanhao Jiang · 2026-08-24

E2S-Pruner introduces a progressive two-stage evidence-fusion framework for visual token pruning in vision-language models, eliminating the need for auxiliary models, trainable parameters, or fine-tuning. The method first treats each attention head as an independent evidence source, estimating reliability via evidence clarity and inter-head consistency, then uses Dempster-Shafer theory to fuse evidence across layers while incorporating a spatial novelty constraint. On LLaVA-1.5-7B, it retains 96.8% performance with 128 tokens (2.09x throughput gain) and demonstrates cross-model generalization on Qwen2-VL-7B.

visual token pruningevidence fusiondempster-shafer theoryvision-language modelsattention heads

Future Querying: Can LLMs Serve as Implicit Medical World Models?

arXiv cs.AI · Siri Willems, James Butterworth, Lore Goetschalckx, Peter Vrancx · 2026-08-24

The paper introduces future querying, a paradigm evaluating whether LLMs can serve as implicit medical world models by answering time-indexed clinical queries about patient trajectories. The method processes unstructured clinical documentation with endpoint-agnostic training, enabling a single model to handle diverse queries without task-specific retraining. Results on synthetic medical reports and MIMIC-IV ICU notes demonstrate that small, locally fine-tuned open-weight LLMs can approach proprietary system performance, suggesting feasibility for privacy-preserving deployment.

future queryingimplicit world modelsendpoint-agnostic trainingunstructured clinical documentationmimic-iv

Multi-Winner Voting with Argumentative Ballots

arXiv cs.AI · Ryuta Arisaka, Hirotaka Ono · 2026-08-24

The paper introduces multi-winner voting with argumentative ballots (MVArg), generalizing approval ballots to express defeasible preferences and extending voter cohesion and justified representation axioms (JR, PJR, EJR). The authors prove MVArg is strictly more expressive than standard multi-winner voting, show their axioms conservatively generalize classical ones, and establish computational complexity results: JR satisfaction is coNP-hard but polynomial-time constructible, while PJR/EJR counterparts are not always satisfiable. All results are formalized in Lean 4.

multi-winner votingargumentative ballotsjustified representationcomputational social choiceaxiomatic verification

Credal Large Language Models for Semantic Commitment under Uncertainty

arXiv cs.AI · Shireen Kudukkil Manchingal, Sofiia Nikolenko, Fabio Cuzzolin · 2026-08-24

The paper introduces Credal Large Language Models (CLLMs), which address epistemic uncertainty in LLMs by using an ensemble of LoRA adapters to induce a credal set with lower and upper probabilities. Two commitment scores—Credal Token Commitment (CTC) and Semantic Commitment Consistency (SCC)—quantify uncertainty at token and semantic levels. Evaluated on Gemma-2-9B, Llama-3.1-8B, and Qwen2.5-7B across QA benchmarks, CLLMs achieve competitive accuracy and calibration, with CTC performing within 1.5 pp of the best hallucination AUROC and SCC enabling 99.0% accuracy at 80% coverage on OpenBookQA.

credal setlora adaptersepistemic uncertaintysemantic commitmenthallucination detection

Retrieval-Augmented Classification of Environmental Mitigations in Hydropower Licensing Documents

arXiv cs.AI · Hong-Jun Yoon, Tom Ruggles, Joanna Lee, Debjani Singh · 2026-08-24

The paper introduces a hybrid retrieval-augmented system for multi-label classification of environmental mitigations in hydropower licensing documents, addressing severe label scarcity (40/135 categories lack training examples). The method combines fine-tuned BERT for detection with Retrieval-Augmented Generation (RAG) for zero-shot classification, leveraging retrieved category definitions. Evaluated on 5,860 paragraphs across 135 categories, the hybrid achieves Micro F1 of 0.524, outperforming BERT-only (0.477) and RAG-only (0.416) pipelines.

multi-label classificationretrieval-augmented generationzero-shot learningbertlabel scarcity

Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

arXiv cs.AI · Zhiruo Zhou, Zelin Li, Xiwen Chen, Jiazhuo Li · 2026-08-24

The paper introduces TOWN-VLA, a prompt-authority control mechanism that prevents prompt-form collapse in vision-language-action (VLA) policies by selectively authorizing retrieved text interventions. The method separates candidate generation from prompt alteration, restoring the original Base prompt unless a compatibility rule approves the modification. Evaluated on LIBERO-Plus (10,030 episodes) and a physical PiPER arm (150 trials), TOWN-VLA improved success rates from 69.5% to 73.1% (simulated) and 52.7% to 78.7% (physical), demonstrating enforceable control without retraining the frozen policy.

prompt-form collapsevision-language-actionretrieval augmentationfrozen policycompatibility rule

What is mathematics now, and what should it be?

arXiv cs.AI · Jeremy Avigad · 2026-08-24

The essay critiques the narrow focus on neural theorem provers in AI-assisted mathematics, advocating for a broader conceptualization of AI's role. It argues that current successes in automated theorem proving obscure unexplored opportunities for collaboration between mathematicians and AI systems. The author proposes an expanded research agenda that moves beyond formal verification to include heuristic discovery, mathematical intuition, and creative problem-solving support.

neural theorem proversautomated reasoningmathematical intuitionheuristic discoveryformal verification

AI Surrogate Modeling for Real-Time Tokamak Equilibrium Prediction: Benchmarking Neural Architectures and Validation on EXL-50U

arXiv cs.AI · Guoyang Shi, Zitong Zhang, Siqi Ding, Jianguo Chen · 2026-08-24

The study benchmarks five neural architectures (MLP, CNN, FNO, Transformer, KAN) as AI surrogates for real-time tokamak equilibrium prediction, evaluating accuracy, speed, and robustness on 100,000 IID and 10,000 OOD samples from a numerical Grad-Shafranov (GS) database. Device-level validation on EXL-50U confirms simulation-to-device consistency, with surrogates achieving $10^{-3}$-$10^{-2}$ error relative to GS solutions. Transformers yield the best IID accuracy, while CNNs balance accuracy (0.7 ms latency), robustness, and extrapolation stability (4%-5% $L_2$ error on unseen geometries). Scaling improves interpolation but not OOD generalization, highlighting a capacity-generalization trade-off.

surrogate modelingtokamak equilibriumgrad-shafranov solverout-of-distribution generalizationreal-time control

BenthicDINO: Physics-Informed Self-Distillation for View-Invariant Side-Scan Sonar Representations

arXiv cs.AI · Taqi Hamoda, Hayat Rajani, Nuno Gracias · 2026-08-24

The paper introduces BenthicDINO, a physics-informed self-distillation framework for learning view-invariant representations in side-scan sonar (SSS) imagery. The method combines physically motivated augmentations (speckle noise, range-dependent attenuation) with a Hilbert-Schmidt Independence Criterion (HSIC) penalty to decouple features from viewing parameters, using a ConvNeXt-v2-Tiny backbone and hierarchical feature fusion. Evaluated on the S3Seg dataset, it achieves 71.4% mIoU and 86.5% accuracy, reaching 96% of peak performance with only 10% labeled data.

self-distillationside-scan sonarview-invariancehilbert-schmidt independence criterionconvnext-v2

Cognitive Profiling of LRMs' Reasoning Traces Using Bloom's Taxonomy

arXiv cs.AI · Maria-Eleni Zoumpoulidi, Georgios Paraskevopoulos, Alexandros Potamianos · 2026-08-24

The study introduces a framework for automatically annotating reasoning steps in Large Reasoning Models (LRMs) using Bloom's Taxonomy, classifying cognitive processes into six levels (e.g., Remembering, Applying). This enables granular analysis of model behavior across tasks and datasets. Results reveal distinct thinking patterns among models, with cognitive-level annotations correlating with reasoning correctness, suggesting potential for improving LRM performance through cognitive profiling.

large reasoning modelsbloom's taxonomycognitive profilingreasoning tracesin-context learning

AI emotional support is better only when chosen, but shifts preferences even when it is not

arXiv cs.AI · Yaoxi Shi, Cathy Mengying Fang, Guy LabanPattie Maes, Amit Goldenberg · 2026-08-24

This study examines how AI emotional support influences user preferences when choice congruence is manipulated. Across three experiments (N=1,951), participants selected human or AI support but were randomly assigned congruent or incongruent partners. AI support was rated superior only when chosen, yet interaction with AI increased subsequent preference for it regardless of congruence. A 28-day longitudinal study (N=981) using OpenAI showed daily personal conversations shifted preferences toward AI and away from humans, demonstrating path-dependent choice dynamics in emotional support.

emotional supportchoice congruencepreference shiftpath-dependencylongitudinal study

NetConfArena: An Executable Benchmark for LLM Agents in Closed-Loop Network Configuration

arXiv cs.AI · Chang Liu, Xiaohui Xie, Xinyi Chen, Yong Cui · 2026-08-24

NetConfArena introduces an executable benchmark for evaluating LLM agents in closed-loop network configuration, addressing limitations of static or simplified existing benchmarks. The method employs emulated multi-device networks, a standardized action interface, and hidden executable test cases, with an LLM-assisted pipeline converting network materials into parameterized task templates. Evaluation of 480 task instances (96 templates) across 3840 trajectories reveals failure modes beyond command errors, including task-specification adherence and robust planning gaps, suggesting directions for improved supervision and execution harness mechanisms.

llm agentsnetwork configurationexecutable benchmarkclosed-loop evaluationemulation-grounded pipeline

Counterfactual Transition Graphs: Evaluating Cross-Class Transition Quality

arXiv cs.AI · Syed Muhammad Hamza Zaidi, Szymon Bobek, Grzegorz J. Nalepa, Myra Spiliopoulou · 2026-08-24

The paper introduces counterfactual transition graphs (CGTs) to structurally evaluate time-series classifiers by analyzing inter-class transitions rather than individual examples. CGTs represent classes as nodes and transition reliability as edges, computed via proximity-aware retrieval sweeps of counterfactual edits. On a six-class hand-movement task, CGTs reveal non-trivial topology where counterfactual reachability inversely correlates with classifier confidence (Spearman ρ=−0.37), contrasting binary confusion matrices. The framework compares gradient-based CFs (manifold-violating) and replacement-based CFs (manifold-preserving), showing divergent boundary-crossing behaviors.

counterfactual explanationstime-series classificationtransition reliabilitydata manifoldinterpretability

Language Chain in Alignment: Cross-Lingual Ranking Preference Optimization

arXiv cs.AI · Seungyoon Lee, Minhyuk Kim, Jungseob Lee, Heuiseok Lim · 2026-08-24

The paper introduces Cross-Lingual Ranking Preference Optimization (CRPO), a framework for improving multilingual alignment in LLMs by transferring English preference knowledge to target languages. CRPO employs a hierarchical structure within parallel preference pairs to jointly optimize intra- and inter-lingual preferences, extending LambdaLoss for relative ranking across responses. Experiments in five languages show CRPO outperforms standard methods in instruction-following and knowledge utilization, with consistent gains in reward margins and log-probability of desirable responses.

cross-lingual alignmentpreference optimizationlambdalossmultilingual llmsranking signal

Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

arXiv cs.AI · Xiwen Chen, Zelin Li, Zhiruo Zhou, Huiming Chen · 2026-08-24

Pointing-VLA introduces typed spatial grounding interfaces for vision-language-action (VLA) models, replacing brittle text-based coordinate outputs with explicit geometry-specific heads for normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories. Built on Embodied-R1, the method employs an execution contract aligning PICK/PLACE actions with OFG and Pointing readouts. Evaluated on Bridge/WidowX and physical pick-place tasks, it achieves 72.9% average success (SOTA) without task-specific finetuning, improves real-robot success from 52.7% to 80.7%, and reduces controller time by 20×. Typed heads also show 6.68–6.90× faster decoding than text-based approaches.

spatial groundingvision-language-actionobject-functional groundingtyped readoutsrobot execution

LITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space

arXiv cs.AI · Jinghui Zhang, Lang Gao, Ao Li, Mingzhe Li · 2026-08-24

LiteraryBigFive introduces a unified interpretable space for author-personalized text generation, modeling writing characteristics as coordinates along five interpretable axes (e.g., Classicism, Emotionality) derived from activation-space contrasts between author-written and neutral passages. The framework includes an interpretable steering mechanism that adaptively guides generation toward target coordinates, enabling author-specific stylistic control. Experiments demonstrate improved authorial expressiveness and semantic fidelity, with derived axis scores correlating strongly (p<0.05) with literary consensus, providing transparent explanations for generation behavior.

interpretable spaceactivation-space contrastsauthor-personalized generationstylistic dimensionssteering mechanism

Statistical Machine Translation Systems of English-Pnar Language Pair : Some Insights of the Emperical Study

arXiv cs.AI · Edawanbiang Dhar Surmila Thokchom, Thoudam Doren Singh · 2026-08-24

The first machine translation study for English-Pnar establishes quantitative benchmarks using phrase-based SMT, achieving 14.97 BLEU (Pnar→English) and 11.16 BLEU (English→Pnar) on a 371-sentence test set. A parallel corpus of 10,234 sentences from Wyrta newspaper was processed via Moses, GIZA++, and KenLM with lexicalized reordering (yielding +3.73 BLEU for SOV→SVO shifts) and MERT tuning (which degraded performance). Error analysis reveals challenges from morphological OOVs, long-distance reordering, and Khasi code mixing, suggesting neural and multilingual approaches as future directions.

statistical machine translationparallel corpuslexicalized reorderingminimum error rate traininglow-resource language

DeMixPert: Decomposed Response Modeling with Gaussian Mixtures for OOD Single-Cell Perturbation Prediction

arXiv cs.AI · Jiawen Liu, Xuechenxiao Cao, Yutong Li, Bing Liu · 2026-08-24

DeMixPert introduces a decomposed modeling approach for out-of-distribution single-cell perturbation prediction, separating transcriptomic responses into basal-state-dependent systematic, perturbation-specific, and population-level variation components. The method uses pretrained target embeddings for unseen-target generalization, models variation via a Gaussian prototype Invertible Network, and combines these through adaptive mixture modeling. Experiments demonstrate superior performance in capturing heterogeneous single-cell responses and generalization to unseen perturbations compared to existing methods.

single-cell perturbationgaussian mixtureout-of-distributioninvertible networktranscriptomic response

Molecular LLM Agents: From Architectural Design to Scientific Autonomy

arXiv cs.AI · Jiatong Li, Wengyu Zhang, Weida Wang, Yuxuan Ren · 2026-08-24

The paper proposes a conceptual framework for molecular LLM agents through architectural and autonomy perspectives. Architecturally, it addresses molecular representation, agent frameworks, domain-specific toolboxes, and optimization. Autonomously, it introduces a four-level ladder (L1-L4) classifying agents from fixed workflows to scientific-agenda autonomy. The framework enables comparison of existing agents, identifies capability gaps, and guides future deployment in molecular discovery workflows.

molecular llm agentsscientific autonomy laddermolecular representationdomain-specific toolboxesfeedback-aware agents

PolyChirp: Multi-Species Birdsong Classification Using TinyML on Low-Power Acoustic Sensors

arXiv cs.AI · Nathan Duboisset, Zhaolan Huang, Felix Bießmann, Roudy Dagher · 2026-08-24

PolyChirp introduces a TinyML system for multiclass bird species classification on low-power microcontrollers, addressing the limitation of prior binary single-species approaches. The method combines biological domain knowledge, automated dataset curation, neural architecture optimization, and NPU-accelerated microcontroller hardware. Evaluations show PolyChirp outperforms state-of-the-art binary classifiers while achieving robust 10-species classification, maintaining operational viability for a full season on a single battery charge through optimized memory footprint (exact values unspecified), latency, and energy efficiency.

tinymlmulticlass classificationneural processing unitacoustic sensorslow-power hardware

Shaping the Evolutionary Dynamics of Robot Morphology via Adaptive Control Learning

arXiv cs.AI · Junru Song, Yang Yang, Yaqing Xu, Ying Wen · 2026-08-24

This paper investigates bidirectional brain-body interplay in robot co-design, formalizing two orthogonal dimensions of morphological contribution: convergence speed (morphological intelligence) and performance ceiling (true potential). The authors introduce AdaControl, an adaptive method that monitors selection bias toward fast learners and allocates minimally sufficient control learning for unbiased fitness evaluation. Experiments on voxel-based soft robots show AdaControl enables a simple genetic algorithm to match state-of-the-art generative-model-based co-design methods in discovering diverse high-performing designs while reducing computation by 80% versus exhaustive control, revealing the morphological Baldwin effect as an artifact of premature evaluation bias.

morphological intelligencetrue potentialrobot co-designadaptive controlbaldwin effect

Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement

arXiv cs.AI · Yufeng Han, Lifan Deng, Cunliang Kong, Wenhao Li · 2026-08-24

Jiuge-Tuiqiao introduces an interpretable human-AI collaborative system for classical Chinese poetry refinement, addressing limitations of one-shot generation paradigms. The system employs a triadic model combining user-driven control (character/line locking, real-time prosody feedback), ancient-guided evidence (high-frequency collocations, PPL-ranked classical lines), and AI-assisted generation (structured knowledge from classical encyclopedias). Preliminary experiments indicate improved controllability, interpretability, and user engagement compared to autonomous AI poetry systems.

human-ai collaborationinterpretable refinementprosody feedbackclassical poetrystructured knowledge

POOL: Propagated Uncertainty Over Lookalikes

arXiv cs.AI · Rounak Sharma, Ananya B. Sai, Soumyabrata Pal · 2026-08-24

The paper introduces POOL (Propagated Uncertainty Over Lookalikes), a cost-efficient framework for estimating confidence scores in black-box LLMs by clustering query stems and propagating uncertainty scores. It combines verbal confidence with spectral answer diversity (Hy@p) using negative von Neumann entropy of sampled answer embeddings. Evaluated across six domains and five LLMs, Hy@5 achieves higher AUROC than verbal confidence and Vn@10 sampling while using 50% fewer samples. POOL-Hy@5 retains 93.5–97.9% AUROC with 19.3–39.3% generation savings, reaching 73–76% savings on paraphrase-dense workloads.

confidence estimationblack-box llmsspectral diversityvon neumann entropygroup-testing

AgentWeave: Routing Before Reasoning for Efficient Function Calling in Tool-Rich Language Models

arXiv cs.AI · Saurav Singla, Aarav Singla, Advik Gupta, Parnika Gupta · 2026-08-24

AgentWeave introduces a deterministic pre-inference routing layer to reduce the candidate action space for function-calling in tool-rich language models, enhancing efficiency without altering the downstream model. The method constructs a bounded model-visible action space using eligibility, requirement, capability, and routing signals, evaluated with a frozen BFCL-derived routing-pressure protocol on the MadeAgents/Hammer2.1-1.5b model. Results show AgentWeave achieves 12.5% native BFCL successes on 48 tasks, outperforming baselines with 0% success, while reducing tool exposure by 70.18%, input tokens by 61.70%, and local-model latency by 50.95%. This demonstrates the impact of candidate-space construction on function-calling behavior.

function-callingpre-inference routingbfclcandidate-space constructiontool-rich language models

From Generation to Simulation: How Far Are World Models from Being True Simulators?

arXiv cs.AI · Tong Wang, Huan Deng, Mucheng Yang, Yang He · 2026-08-24

The study systematically evaluates generative world models against eight capabilities of traditional simulators (asset construction, physics engine, interaction, controllability, stability, state feedback, diversity, evaluation metrics), analyzing 200 works from 2018–2026 across three technical routes: latent dynamics, video generation, and joint-embedding prediction. Results indicate functional substitution in interaction and controllability for specific scenarios, but shortcomings in physical law guarantees, structured state feedback (only 6/163 papers expose runtime state queries), and long-horizon stability. Six research directions are proposed, including formalized physics and unified action interfaces.

generative world modelslatent dynamicsstate feedbackjoint-embedding predictionphysics engine

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

arXiv cs.AI · Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran · 2026-08-24

The Cultural Moment Benchmark (CMB) introduces a diagnostic framework for evaluating video-cultural reasoning through three distinct abilities: concept naming (S1), visual recognition (S2), and temporal localization (S3). CMB comprises 306 expert-curated concepts from seven Southeast Asian countries, tested via staged evaluations with semantic-similarity distractors and unlabeled video moments. Results show state-of-the-art models score below 30% when all stages must be correct, with abilities failing to cascade (S1 aids S2 but not S3) and audio often distracting. Human experts perform below chance on neighboring countries' concepts, highlighting country-specific knowledge requirements.

cultural reasoningvideo groundingtemporal localizationmultimodal evaluationsemantic distractors

Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies

arXiv cs.AI · Xiaotong Tan, Chunli Qiu, Xin Liu, Qing Huang · 2026-08-24

The study demonstrates that a feature-based hybrid LLM architecture outperforms end-to-end strategies for automated O-RADS classification from pelvic ultrasound reports. Comparing eight LLMs with three reasoning approaches (implicit-knowledge end-to-end, rule-informed end-to-end, hybrid), the hybrid method using Gemini 3.6 Flash achieved 99.2% accuracy (387/390 cases) and perfect agreement (weighted κ=1.00), surpassing original reports (87.7% accuracy) and pure LLM strategies (65.6-95.9% accuracy). The hybrid approach reduced misclassifications and overstaging while maintaining interpretability through decoupled feature extraction and rule-based classification.

large language modelso-rads classificationhybrid architectureclinical decision-makingfeature extraction

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

arXiv cs.AI · Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan · 2026-08-24

The paper systematically reviews LLM-based forecasting agents, categorizing architectures into standalone workflows, tool-augmented agents, and hybrid systems combining LLMs with statistical models. It analyzes training methods, evaluation protocols, and applications across finance, weather, health, energy, and operations. Key findings highlight measurement limitations, including sensitivity to input perturbations, potential benchmark contamination, and cases where LLMs fail to improve accuracy. The study identifies critical research gaps in calibration under distribution shift, contamination-resistant evaluation, and handling feedback loops between forecasts and outcomes.

llm-based agentsforecasting systemstemporal reasoningbenchmark contaminationdistribution shift

Beyond Verdicts: A Graph-Based Analysis of Human and LLM Reasoning in Scientific Fact-Checking

arXiv cs.AI · Abdul Ghafoor, Muhammad Arslan Manzoor, Yufang Hou · 2026-08-24

We introduce a graph-based framework (typed reasoning graph) for comparing human and LLM reasoning paths in scientific fact-checking, addressing limitations of existing verdict-only systems. Building on MISSCIPLUS, we model explanations as reasoning graphs linking claims to study contexts, findings, premises, and fallacy labels, enabling fallacy-specific sub-graph alignment. Evaluating GPT-5, Claude Opus 4.7, and Qwen3-32B on 84 false claims, we find distinct performance dimensions: Qwen3-32B achieves the lowest verdict failure rate, GPT-5 shows the highest human alignment, and Claude Opus 4.7 demonstrates weak verdict prediction but often valid reasoning in successful cases.

reasoning graphmisciplinusfallacy-specificverdict failurehuman alignment

From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation

arXiv cs.AI · Xiangxin Zhang, Zhanwei Zhang, Zhihang Fu, Binbin Lin · 2026-08-24

The paper identifies inertia bias in LLM-powered web search agents, where self-authored actions systematically distort subsequent judgments, and introduces the IBIS benchmark to quantify this effect. The proposed NIS-Agent mitigates bias via context isolation at webpage triage and answer validation, reducing token costs by 33% while maintaining performance on GAIA, WebWalkerQA, and BrowseComp benchmarks. An 8B model trained for bias resistance matches GPT-4o's performance under the same framework.

inertia biasweb search agentscontext isolationibis benchmarknis-agent

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

arXiv cs.AI · Sungho Park, Wonjoong Kim, Rongyuan Tan, Jue Zhang · 2026-08-24

AutoSaddler introduces an automatic harness optimization framework for improving LLM agent reliability on long-horizon tasks by treating harness improvement as an offline learning problem. The method iteratively updates harnesses through failure-trace diagnosis, structured patch generation (treating harnesses as code), and validation-based update selection. Evaluations on GAIA2, SWE-Bench Pro, and Terminal-Bench 2.0 demonstrate performance improvements of 9.0, 9.6, and 10.0 percentage points respectively, with ablations highlighting the importance of deep debugging, targeted modifications, and generalization-aware selection.

harness optimizationllm agentsfailure-trace diagnosisstructured patch generationoffline learning

The Multilingual FrameNet Corpus

arXiv cs.AI · Beatrice Fiumanò, Nicolas Lazzari, Simone Paolo Ponzetto, Valentina Presutti · 2026-08-24

The Multilingual FrameNet Corpus (mFNC) extends the Berkeley FrameNet corpus by harmonizing frame-semantic annotations across nine additional languages (Brazilian Portuguese, Chinese, Dutch, French, German, Italian, Korean, Latvian, Swedish). Using diverse architectures trained on mFNC, the authors achieve state-of-the-art performance in multilingual and cross-lingual Frame Semantic Parsing (FSP), demonstrating the value of multilingual training data. The corpus and trained models are publicly released.

frame semanticsmultilingual corpussemantic parsingcross-lingual transferframe annotation

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

arXiv cs.AI · Yi Zhu, Xiongwei Wu, Qiyi Wang, Tingyu Qu · 2026-08-24

The paper introduces MobilePA-Bench, an interactive benchmark for evaluating mobile planner agents' tool-calling and planning abilities across 13 functional domains with 212 realistic tools. It addresses limitations in existing GUI-centric and static function-calling benchmarks by incorporating live application databases, structured feedback, and three advanced evaluation dimensions: sub-agent collaboration, memory usage, and skill usage. Experiments reveal that current LLMs perform unreliably under mobile constraints like strict tool ordering and permission limits, demonstrating the benchmark's utility for diagnosing agent weaknesses and enabling reinforcement learning.

mobile planner agentstool-callingsub-agent collaborationinteractive benchmarklive application databases

FedCC: Towards Addressing Label Distribution Skews in Distillation-Based Federated Learning

arXiv cs.AI · Wenxuan Ye, Onur Ayan, Xueli An, Georg Carle · 2026-08-24

FedCC addresses label distribution skew in distillation-based federated learning by introducing an 'unknown' class for ambiguous samples and using calibrated pseudo-labels on public data. The method mitigates bias toward majority classes by balancing confidence and uncertainty, particularly for under-represented classes. Experiments show FedCC achieves 67.3% accuracy under extreme label skew (one class per client), outperforming baselines that collapse to near-random performance.

federated learninglabel distribution skewdistillation-based learningpseudo-label calibrationunknown class

Artificial Empathy: Towards a Framework for Unsupervised Agency Detection and Policy Reconstruction

arXiv cs.AI · Peter Kuhn, Chris Pang, Sonakshi Chauhan · 2026-08-24

The paper proposes a framework for unsupervised agency detection and policy reconstruction in AI systems, addressing the understudied problem of identifying and modeling other agents from observation alone. The method employs a reinforcement learning agent, pre-trained on an independent task, as a prior for inferring agentic dynamics in the environment. This approach avoids the constraints of inverse reinforcement learning and aims to enable cooperative behavior in real-world settings.

agency detectionpolicy reconstructionreinforcement learninginverse reinforcement learningagentic dynamics

PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies

arXiv cs.AI · Zeyu Feng, Qingyu Wu, Yuzhe Luo, Hua Cheng · 2026-08-24

PsychJail introduces a psychology-guided framework for multi-turn jailbreak attacks on aligned LLMs, operationalizing the Persuasion Knowledge Model (PKM) through tactic-conditioned attack policies and Change-of-Meaning analysis. The method employs trajectory-level reinforcement learning with PKM-gated rewards to refine attacks. Results show an 87.3% average success rate across four victim models, revealing four distinct psychological profiles (rationalist, credibility-driven, narrative-monoculture, broadly persuadable) that explain cross-model transfer asymmetry.

psychological jailbreakspersuasion knowledge modelmulti-turn persuasionchange-of-meaning analysistrajectory-level reinforcement learning

Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

arXiv cs.AI · Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel, Yash Raj Shrestha · 2026-08-24

The study introduces a multi-agent audit framework to disentangle sociocultural signals in multilingual LLMs, distinguishing between social bias reproduction, identity group representation, and cross-cultural pattern reflection. Analyzing 89,253 outputs from 12 LLMs across English, French, and Chinese, the method employs human-validated controls including identity cue removal and translation. Results show bias representation varies by language and task, with direct identity cues significantly affecting English and Chinese outputs but less so in French. Source language context consistently receives high relevance scores, though identification accuracy drops post-translation and name masking, highlighting the risk of conflating surface cues with cultural understanding.

multilingual llmssociocultural signalsbias representationidentity cuescross-cultural patterns

Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

arXiv cs.AI · Xunlei Chen, Qirui Ye, Yuang Li, Yi Gong · 2026-08-24

The paper introduces ADU, a training-based framework for precise unlearning in large language models that shifts from token erasure to contextual attention-pathway decoupling. The method exploits functional distinctions between local and global attention heads to identify and modify retrieval paths for sensitive content while preserving linguistic structure and general utility. ADU achieves a Forget Quality of 0.93 on TOFU and maintains 92.9% average model utility (versus 81.9% for baselines), with reduced side effects in benign contexts.

unlearningattention-pathwaycontextual decouplinglanguage modelsforget quality

SplitLite: Low-Rank Residual Compression for Split Learning

arXiv cs.AI · Tao Li, Yulin Tang, Qi Guo, Xianhao Chen · 2026-08-24

SplitLite introduces a communication-efficient split federated LoRA fine-tuning method for on-device LLMs, exploiting low-rank structures in activation and gradient residuals. The key insight is that LoRA's rank-r parameter updates induce effective rank-2r (activation) and rank-4r (gradient) residual structures between epochs, enabling compressed transmission via quantized truncated SVD. Experiments on GLUE with various on-device LLMs show 93.5% reduction in activation uplink costs and 83.7% total communication savings without performance loss.

split learninglow-rank approximationlora fine-tuningcommunication efficiencyon-device llms

Coarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG

arXiv cs.AI · Zhe Jin, Zhimin Lin, Bin Zheng, Junhua Fang · 2026-08-24

The paper proposes Density-Aware Graph Construction (DAGC), a method that decouples temporal granularity in long-video retrieval-augmented generation (RAG) by maintaining separate coarse indexing and fine evidence spaces. DAGC constructs a compact, density-adaptive graph index by merging redundant video chunks while preserving mappings to original temporal units, enabling efficient retrieval followed by fine-grained evidence expansion. Experiments on MLVU, VideoMME, and LongVideoBench show DAGC reduces graph nodes by 50-60% and achieves 1.3-1.7× speedup while maintaining 99% of original QA performance across different LVLM backbones and RAG pipelines.

retrieval-augmented generationtemporal granularitygraph constructionvideo understandingdensity-adaptive

PatchWrite: One Line, Not One Section -- Compile-Gated, Validity-Preserving Editing for AI-Drafted Manuscripts

arXiv cs.AI · Weiwei Yang · 2026-08-24

PatchWrite introduces a validity-preserving editing protocol for AI-drafted manuscripts that compiles candidate edits through two gates: fatal-log compilation checks and evidence locks requiring cited keys and numeric tokens to match reference registries. Compared to whole-section regeneration (which mutated unrelated content in 100% of 192 test cases), PatchWrite preserved all target lines while fixing 93.75% of faults in model-proposed edits. Blind evaluations showed significant preference for PatchWrite in fact preservation (Likert 5.0 vs. 2.0) with equivalent prose quality, and deployment logs confirmed real-world applicability.

compile-gated editingevidence locksvalidity-preservingreference registrynumeric jaccard

A Physical Response-and-Memory Model for Muon Optimization

arXiv cs.AI · Yinze Hu, Hongjun Xiang, Xingao Gong, Hongyu Yu · 2026-08-24

The paper proposes Bi-Maxwell, a novel optimizer for large language models based on a physical model treating weight matrices as responsive media with memory. By modeling momentum as internal stress relaxation across multiple timescales and deriving semi-orthogonalization as maximally dissipative response under output constraints, the method theoretically explains Muon's effectiveness. Experimental validation shows Bi-Maxwell achieves target loss in fewer steps than single-timescale approaches on public LLM benchmarks, with probe measurements confirming predicted dynamic memory length requirements during training stages.

optimizermomentum matrixsemi-orthogonalizationstress relaxationmemory kernel

Hypergraph Embedding Indexing for Efficient Dense Vector Retrieval

arXiv cs.AI · Kishore Konda · 2026-08-24

The Hypergraph Embedding Index (HEI) introduces a novel framework for dense vector retrieval by organizing documents according to combinations of highly activated latent embedding dimensions, enabling inverted-index style candidate generation while preserving semantic ranking. The method constructs multiple complementary hypergraphs to improve retrieval coverage without combinatorial growth and identifies activation diversity as a key metric for indexing efficiency. Results demonstrate that HEI maintains semantic ranking capabilities while offering efficient retrieval through coordinate-inverted indexing.

hypergraph embeddingdense vector retrievalactivation diversitycoordinate-inverted indexingsemantic ranking

SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems

arXiv cs.AI · Xiang Wang, Shigang Quan, Tingzhen Chang, Kang Yang · 2026-08-24

The paper proposes SA-RSQ (Sparse Activation-based Residual Soft Quantization), a framework for efficient multi-modal feature representation in recommender systems. The method employs Top-K sparse routing and softmax weights to store compact (Index, Probability) tuples, decoupling storage from codebook dimensionality while enabling gradient propagation without straight-through estimators. Evaluated on a proprietary food-delivery dataset, SA-RSQ achieves favorable reconstruction-performance and CTR trade-offs across 8-48 bytes/item storage budgets, with online A/B tests showing +2.51% CTR and +3.66% CPM improvements.

sparse activationresidual soft quantizationmulti-modal recommendationtop-k routingctr optimization

Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B

arXiv cs.AI · Defu Lin, Wenhui Chen, Ziyao Lin, Jianlin Chen · 2026-08-24

The paper introduces ASP, a training-free wrapper for frozen multimodal models that addresses budget-constrained embodied perception through four resource walls: perceptual Shannon, horizon, round, and conditional composition. ASP combines capped structured state, verbatim episodic index, and query-conditioned budget allocation with iterative access. Evaluated on SEW-Bench with 3B-31B models under a 4,096-token budget, ASP achieves 75-94% episodic retrieval accuracy versus 3-19% for query-independent sampling, though prompted online compression underperforms verbatim-only baselines.

embodied perceptionmultimodal modelsbudget allocationepisodic retrievalresource walls

Toward Effective and Reliable LLM Agents via Dynamic Ontology

arXiv cs.AI · Xiaohui Zhang, Zequn Sun, Chengyuan Yang, Yuanning Cui · 2026-08-24

The paper introduces OaK, an ontology-as-a-kernel framework that dynamically constructs and refines task-oriented ontologies for LLM agents to improve knowledge grounding and multi-step reasoning. OaK generates ontologies and knowledge graphs from task requirements and training data, develops task-adaptation functions for graph reasoning, and iteratively refines these using judge feedback. Evaluations on TravelPlanner, CRMArenaPro, and ToolQA demonstrate that OaK enhances standard LLM agents by strengthening evidence use and boosting reasoning reliability.

ontology-as-a-kernelknowledge graphmulti-step reasoningllm agentstask-adaptation functions

ParallelWorld: Test-Time Scaling for Embodied Reasoning

arXiv cs.AI · Min Chen, Shengjun Zhang, Yuxin Li, Zhang Zhang · 2026-08-24

ParallelWorld introduces a multi-horizon test-time scaling framework for embodied reasoning, addressing the limitations of myopic, single-step exploration in dynamic environments. The method employs a verifier-guided tree-search paradigm, simulating and evaluating parallel multi-step trajectories while dynamically pruning unpromising branches based on intermediate state transitions. Experiments on ESI-Bench demonstrate consistent improvements in active perception and reasoning performance compared to incremental or single-step approaches.

embodied reasoningtest-time scalingmulti-horizon planningverifier-guided searchactive perception

Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents

arXiv cs.AI · Yuchen Huang, Sijia Li, Jun Zhang, Yi R. Fung · 2026-08-24

The paper introduces SPARE, a KL-guided framework for pruning redundant reasoning text in multimodal tool-use agents while preserving visual evidence. SPARE uses a task-state summary and reverse-KL divergence from on-policy self-distillation to test summary coverage, enabling aggressive pruning without disrupting future reasoning. Evaluations on multi-step visual tool-use benchmarks show SPARE removes 37.89-64.58% of reasoning tokens while maintaining the highest accuracy among pruning methods, demonstrating improved visual reliance and reduced textual dominance.

multimodal large language modelstextual debtcontext pruningkl divergenceon-policy self-distillation

What Process Evaluation of Coding Agents Actually Measures: Action, Task, and Step Are Three Different Levels

arXiv cs.AI · Jiawei He, Mengyu Shi, Jie jia, Xikai Yang · 2026-08-24

The paper introduces a measurement framework for process evaluation in coding agents, distinguishing between action prediction, task uncertainty, and step attribution. It proposes SCAE, a replay-based estimator derived from a structural causal model, combining prefix-conditioned identification, replay/intervention-based estimation, and judge-information manipulation. Experiments on 499 file-localization episodes from 12 repositories reveal that next actions are driven by execution provenance, execution uncertainty is structured at the task level, and full-trace judges exhibit collider bias, indicating current evaluations often measure semantic relevance rather than causal contribution.

coding agentsprocess evaluationstructural causal modelcollider biasexecution provenance

WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans

arXiv cs.AI · Jun Zhang, Qiao Zhao, Cheng Cui, Jianying Qu · 2026-08-24

The authors introduce WildHandBench, a benchmark for handwritten text understanding that addresses limitations of existing evaluations by including diverse document structures (free text, tables, formulas), multiple languages, and real-world degradation scenarios. The benchmark comprises 500 documents and proposes a Prior-Driven Error (PDE) metric to distinguish errors stemming from language priors versus visual evidence. Evaluation of 18 state-of-the-art models and human baselines reveals: top model accuracy at 71.85% versus human performance at 77.09%, with 63-91% of model errors being prior-driven compared to 49% for humans, exposing systematic reliance on linguistic priors.

handwritten text understandingprior-driven errordocument parsingmultimodal evaluationbenchmark

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

arXiv cs.AI · Ziyue Wang, Aomufei Yuan, Yiran Yao, Linli Yao · 2026-08-24

The paper introduces Lit2Test, a benchmark for evaluating language models' research ideation quality through falsifiable proposals. The method constructs a six-field contract requiring proposals to specify falsifying outcomes, using 200 real-paper neighborhoods and 1,200 blind pairwise comparisons. Results show a strict model ranking across 10,000 bootstrap replicates, with test quality (not fluency) driving separation, validated by three annotators with explicit reliability bounds. The benchmark and audit artifacts are publicly released.

falsifiable researchbenchmark designlanguage model evaluationpairwise comparisonbootstrap validation

Concepts for Securing Agentic AI Coding and the Terok Environment

arXiv cs.AI · Jiří Vyskočil, Franz Pöschel, Andreas Knüpfer · 2026-08-24

The paper introduces a security framework for agentic AI in software development, addressing heightened risks compared to conventional AI-assisted coding. It proposes (I) risk assessment, (II) mitigation strategies preserving functionality, and (III) an implementation in the Terok environment. The approach aims to enable safe exploration of agentic AI's potential while minimizing IT security vulnerabilities, acknowledging the field's rapid evolution since its emergence in fall 2025.

agentic aiit securitysoftware developmentrisk mitigationterok environment

Deep Learning-Based Multi-User Communication Design for Dense IoT Networks: Interference-Aware Finite-Blocklength Communication and Preliminary MIMO Extensions

arXiv cs.AI · Arkadeep Sinha, Shubham Paul, R. Manivasakan · 2026-08-24

The paper introduces a deep-learning-based end-to-end multi-user communication design for dense IoT networks, addressing interference-limited finite-blocklength scenarios. The method extends a SiameseNet transceiver framework to 2, 4, and 8 users, leveraging learned redundancy for interference suppression and noise robustness without joint detection. Results show superior Block Error Rate (BLER) performance compared to non-orthogonal access baselines, with linear decoder scaling per user, robustness to interference mismatch, and preliminary 2X2 MIMO extension potential.

finite-blocklength communicationmulti-user interferencesiamesenetblock error ratemimo

Beyond Observed Auxiliary Relations: Environment-Conditioned Modeling for Multi-Behavior Recommendation

arXiv cs.AI · Seunghan Lee, Hyunsik Yoo, Jian Kang, Susik Yoon · 2026-08-24

The paper proposes BOAR, an environment-conditioned framework for multi-behavior recommendation (MBR) that addresses missing and unreliable auxiliary behavioral signals. BOAR introduces two complementary modules conditioned on auxiliary observability to handle these challenges systematically. Experiments show BOAR outperforms state-of-the-art baselines by up to 7.82% in HR@10 overall and up to 44.2% for items without auxiliary observations, demonstrating improved generalization beyond observed relations.

multi-behavior recommendationauxiliary signalsenvironment-conditioned modelinggraph neural networkspreference modeling

Safety Hacking in Constrained Best-of-$N$ Inference-time Scaling

arXiv cs.AI · Akifumi Wachi, Takumi Tanabe, Youhei Akimoto · 2026-08-24

The paper identifies a safety vulnerability in constrained Best-of-$N$ inference-time scaling, where imperfect safety proxies contaminate the feasible set and reward maximization amplifies residual unsafe outputs. Through theoretical analysis, the authors derive finite-$N$ bounds showing that safety hacking becomes asymptotically certain if unsafe outputs have heavier reward tails, even with small proxy errors. They propose coverage control via bounded $χ^2$ divergence to limit amplification but note it cannot fully mitigate contamination. Experiments with toy models and language models validate the contamination and amplification effects.

safety hackingconstrained best-of-ninference-time scalingcoverage controlreward-tail amplification

Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text

arXiv cs.AI · Hyeonyu Kim, Hwayeon Kim, Youngwon Choi, Myeongkyun Cho · 2026-08-24

The paper proposes a framework to improve spoken language models (SLMs) by explicitly addressing structural gaps between speech and text representations. The method decouples length mismatch from semantic alignment during training to enhance cross-modal correspondence. Experiments on multiple benchmarks show competitive performance against baselines, demonstrating that current SLMs' weak alignment between speech and text features persists despite strong downstream results.

spoken language modelsspeech-text alignmentcross-modal representationlength mismatchsemantic alignment

CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents

arXiv cs.AI · Xiwei Dai, Zijie Meng, Zhiting Fan, Yixuan Tang · 2026-08-24

The paper introduces CDEG, a graph-based framework for long-horizon diagnostic agents that learns and validates decision-critical evidence from historical trajectories. CDEG contrasts successful and failed trajectories, identifies impactful evidence through counterfactual interventions, and structures diagnosis--evidence--action relations into a graph. During inference, it retrieves relevant relations to guide evidence acquisition or reappraisal. Evaluations show CDEG improves diagnostic accuracy by up to 11.5% over baseline agents, demonstrating the importance of evidence-level learning for reliable diagnosis.

diagnostic agentsdecision-critical evidencegraph-based frameworkcounterfactual interventionslong-horizon diagnosis

AraDetox: A Multi-Dialect Arabic Detoxification Dataset

arXiv cs.AI · Mo El-Haj · 2026-08-24

AraDetox introduces a multi-dialect Arabic detoxification dataset with 10,500 harmful social-media posts and 84,000 detoxified rewrites across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. Generated using GPT-5 and Gemini 2.5 Flash, outputs were evaluated through human assessment and automatic metrics (lexical change, semantic preservation, sentiment, dialectal style). Results indicate detoxification preserves meaning despite lexical reformulation, with human evaluation confirming harmful-language removal and dialectal alignment. The dataset supports research on Arabic detoxification and multi-dialect NLP.

arabic detoxificationmulti-dialect nlpllm-assisted generationsemantic preservationdialectal style

Proxy reliance in large language model decisions is uncalibrated to predictive evidence

arXiv cs.AI · Zengqing Wu, Chuan Xiao · 2026-08-24

The study introduces a method to audit proxy reliance in LLM decisions by comparing model behavior against ground-truth evidence warrants, revealing three patterns: over-reliance, warranted reliance, and under-reliance. Using a clinical-ranking task with known ground truth, the authors evaluate four LLMs under neutral labels, informative proxies, and social field names. Key findings show proxy reliance poorly tracks predictive evidence, social-label suppression is fragile (easily disrupted by in-context examples), and accuracy metrics fail to detect these biases. The work demonstrates that current demographic-based audits cannot distinguish discrimination from valid inference.

proxy reliancelarge language modelspredictive evidencein-context learningbias detection

The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models

arXiv cs.AI · Taebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi · 2026-08-24

The paper formalizes prefix invariance in sequence models, requiring that representations at position t remain independent of future inputs. It introduces a lightweight audit method requiring only two forward passes (no training or gradients) to precisely localize causality violations. While attention-mask inspection fails to detect leaks (0/192 trials), the proposed audit successfully identified all 192 injected faults across eight model checkpoints, including defects in Zamba2 and Nemotron-H. The method reveals that leaks can occur via scans or normalization despite correct masking.

prefix invariancecausality auditattention maskssequence modelsnormalization leaks

Your AI, On a Dial: Controlling Investment Bias in LLMs with a Single Neuron

arXiv cs.AI · Sahong Park, Suhwan Park, Hoyoung Lee, Gakyung Kwon · 2026-08-24

The study introduces an 'investment-bias dial', an inference-time intervention targeting a single neuron to continuously calibrate large language models' (LLMs) aggregate investment stance without modifying prompts or parameters. The method evaluates five open-weight LLMs using matched positive/negative evidence, demonstrating monotonic stance control across decision outputs, rationale emphasis, and agentic retrieval behavior. Results show stable stance control under increasing context lengths (unlike prompt-based attenuation), with downstream effects on security rankings and portfolio composition in exploratory backtests.

investment-bias dialinference-time interventiondecision prioragentic retrievalcontext length

GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

arXiv cs.AI · Long Zhang, Yuhan Chen, Chaoran Zhang, Wanxia Cao · 2026-08-24

The paper introduces GSAR (Goal-State-Anchor Reward), a reinforcement learning framework for training Vision-Language Model-based GUI agents. GSAR addresses data synthesis limitations and unreliable reward signals through self-evolving data synthesis, generating diverse tasks and goal states across multiple environments, and a state-anchor mechanism that annotates task-relevant UI elements for accurate reward signals. Evaluations show 90% accuracy on offline trajectory verification and competitive performance on AndroidWorld and a custom benchmark, demonstrating scalable GUI agent training.

reinforcement learningvision-language modelsgui agentsdata synthesisreward framework

FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

arXiv cs.AI · Hang Wang, Jin Zhang, Guoliang Xu, Pengyue Lu · 2026-08-24

The authors present FinixDoc, an agentic financial document parsing system featuring FinixDoc-VL, a 4B-parameter vision-language model based on Qwen3-VL-4B. The system addresses benchmark-deployment gaps through a Document Parsing Capability Matrix and employs domain-adapted training with homoglyph-aware contrastive learning and multi-stage RL. A human-in-the-loop Data Factory pipeline supports data production. Evaluated on FinixDocBench, FinixDoc-VL achieves 81.43 overall score (5.13 points above next-best open-source model), with peak performance on internal workflows (84.08 vs. 78.73).

vision-language modelcontrastive learningreinforcement learningdocument parsinghuman-in-the-loop

Hierarchy-Aware Supervised Uncertainty Estimation for Black-box LLM Taxonomic Reasoning

arXiv cs.AI · Shuting Xie, Nathaniel Lesperance, Graham W. Taylor · 2026-08-24

The paper introduces hierarchy-aware supervised uncertainty estimation for black-box LLM taxonomic reasoning, addressing reliability challenges in scientific decision support. The method trains lightweight estimators using proxy features from open-source LLMs, with hierarchy-aware supervision to predict rank-wise correctness in taxonomic hierarchies. Evaluated on biodiversity monitoring, the approach outperforms token-likelihood baselines, improving micro AUROC from 0.57 to 0.75–0.80 across three tool LLMs. A rank-specific multi-head design (H3) achieves the best results, demonstrating the importance of hierarchical output structure for unified abstention rules.

uncertainty estimationtaxonomic reasoninghierarchy-aware supervisionblack-box llmselective prediction

Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments

arXiv cs.AI · Ryuki Hyodo · 2026-08-24

Two minimal simulation foundations, SD-AgentFoundry-2D and SD-AgentFoundry-3D, are introduced for studying LLM- and VLM-driven agents in 2D and 3D environments. SD-AgentFoundry-2D enables locally hosted LLM agents to move, communicate, respond to place occupancy, and encounter spatially localized fire events in a 2D multi-agent environment. SD-AgentFoundry-3D provides a 3D digital-twin environment where a locally hosted VLM processes first-person images to generate natural-language movement instructions. Both platforms are designed for local execution on macOS, Windows, and Linux, emphasizing accessibility for education and rapid prototyping. The codebases are intentionally open to modification, serving as foundational tools for generative social simulation and domain-specific extensions.

llm-driven agentsvlm-driven agentsdigital-twin environmentgenerative social simulationspatially localized fire events

Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danmaku

arXiv cs.AI · Xiansheng Luo, Chaowei Zhang, Zewei Zhang, Yi Zhu · 2026-08-24

We propose Genda, a temporal generative Danmaku framework for multimodal fake news detection, addressing the temporal misalignment between Danmaku accumulation and real-time detection needs. Genda comprises a Danmaku Trigger for predicting reaction timing/intensity and a Danmaku Generator for synthesizing semantic-emotional expressions, creating temporally aligned pseudo-Danmaku streams. These streams are integrated into DM-FEND, a Danmaku-guided Temporal Multimodal fake news detection model that enables fine-grained interactions among video, audio, text, and Danmaku modalities. Experiments on Chinese (FakeSV) and English (FakeTT) benchmarks show DM-FEND outperforms state-of-the-art baselines, with ablations confirming the importance of temporal Danmaku modeling for robustness and discriminative capability.

danmakumultimodaltemporal alignmentfake news detectiongenerative framework

Beyond the Harness: End-to-End Optimization of Context Artifacts for Enterprise Text-to-SQL

arXiv cs.AI · Kate Gwimm, Carson Eisenach · 2026-08-24

The paper proposes optimizing knowledge-base context construction for enterprise Text-to-SQL by leveraging historical query patterns rather than treating context as fixed input. Using query-DAG decomposition from production SQL traces, the authors demonstrate that retrieved knowledge-base context provides greater marginal gains (12-25% AST similarity) than retrieval harness optimization (3-12%). A distillation procedure converts historical queries into reusable SQL reference cards, showing improved performance on a 5176-query retail benchmark but mixed results on BEAVER (9.00% vs 6.33%, p=0.12) due to absent production signals.

text-to-sqlknowledge-base contextquery-dagast similarityretrieval harness

Fairness-Aware Mixture-of-Experts via Subgroup Reweighting and Gate Regularization

arXiv cs.AI · Sunhee Hwang · 2026-08-24

The paper proposes a fairness-aware Mixture-of-Experts (MoE) framework addressing routing-induced bias in imbalanced subgroup distributions. The method combines subgroup reweighting to correct data imbalance with gate entropy regularization, preventing biased routing while maintaining expert utilization balance and interpretability. Experiments show improved fairness metrics without compromising predictive performance, with the routing distribution providing interpretable subgroup allocations across experts.

mixture-of-expertsfairnesssubgroup reweightinggate regularizationrouting-induced bias

SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support

arXiv cs.AI · Ahnaf Atef Choudhury, Ramkrishna Saha · 2026-08-24

The paper introduces SDoH-aware narrative anchoring bias as a critical evaluation dimension for medical LLMs, demonstrating that model responses vary across medically equivalent patient narratives despite fixed answer keys. Using NarrativeShield SDoH MedQA, a counterfactual dataset with 300 clinical cases reformatted into persona-based narratives, the authors evaluate three Qwen2.5 models (1.5B, 3B, 7B) across three prompting conditions (8,100 total responses). Results show Qwen2.5-7B achieves highest accuracy (56.33%) and correct consistency (40.33%), yet exhibits persistent narrative sensitivity (minimum error 31.67%), highlighting the need for dual evaluation of correctness and narrative stability in clinical decision support.

narrative anchoring biascounterfactual consistencynarrative sensitivityclinical decision supportinstruction-tuned llms

Triplet2Track: A Hierarchical System with Object-Centric Representations for Reliable Long-Horizon Manipulation

arXiv cs.AI · Jianxiang Liu, Gaojing Zhang, Chuan Wen, Qipeng Liu · 2026-08-24

The paper presents Triplet-to-Track System (TTS), a hierarchical imitation learning system for reliable long-horizon manipulation that addresses limitations of end-to-end VLAs and traditional pipelines. TTS uses human videos to reduce robot data dependence, represents subgoals as instance-grounded triplets, converts them to continuous track priors for execution, and employs observation-based progress monitoring for online replanning. Evaluated on diverse real-world tasks, TTS achieves 74.8% average success rate while supporting object-level and compositional generalization.

hierarchical imitation learninglong-horizon manipulationinstance-grounded tripletstrack priorsonline replanning

Performance of a domain-specific large language model in answering patient questions in psychiatry

arXiv cs.AI · Alexander J. Hish, Arjun Nagendran, Scott N. Compton · 2026-08-24

The study evaluates MIND, a domain-specific LLM fine-tuned on psychiatric patient education resources, against ChatGPT and OpenEvidence in answering escitalopram-related questions. Using rubric-based computer analysis (measuring accuracy, clarity, etc.) and psychiatrist ratings (N=10), MIND outperformed in rubric scores (p<0.001) but was marginally less accurate (p=0.021, r=0.073) and equally safe (p=0.955) versus ChatGPT per psychiatrists, who preferred ChatGPT (57.6% vs. 42.4%, p=0.003). MIND demonstrated clinical utility but revealed trade-offs between completeness and user preference.

large language modelclinical fidelitypatient educationpsychiatryescitalopram

TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

arXiv cs.AI · Wenhao Wu, Menghao Zhang, Xin Wang, Zhi Wang · 2026-08-24

TRACE introduces a self-evolving skill bank framework to improve LLM agent consistency and limit-awareness without weight updates, addressing the reliability gap between potential (Pass@3) and consistent (Pass^k) performance. The method employs trajectory-contrastive evolution to iteratively refine modular skills through comparison of successful/failed behaviors, with runtime skill orchestration. On CAR-bench for in-car assistants, TRACE improves GPT-5.5's Pass^3 consistency by 34.6 points (59.9%→94.5%) and achieves 70% Pass^3 on hidden tests, a 40% relative gain over baselines.

llm agentsskill banktrajectory-contrastive evolutionconsistency gaplimit-awareness

TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts

arXiv cs.AI · Tianqi Xu, Lu Lv, Haoyang Huang, Wenjie Huang · 2026-08-24

TailSieve introduces a partial-rollout-guided framework for optimizing LLM rollouts by jointly controlling tail routing and replica allocation. The method combines tail isolation with load balancing, using partial rollouts to identify long-tail prompts and a hierarchical controller to adapt group isolation and replica allocation. Results show 1.67x routing-only speedup over uniform routing and up to 2.59x speedup when combined with speculative decoding techniques like MTP or DFlash, while maintaining on-policy generation.

llm rolloutstail routingspeculative decodingreplica allocationmakespan optimization

The Retriever Should Remember: Experience-Amortized Reranking for Long-Term Agent Memory

arXiv cs.AI · Qi Feng, Chris Ding, Jicong Fan · 2026-08-24

The paper introduces EARM, an experience-amortized reranking framework for long-term language-model agents that reuses previously acquired LLM relevance scores to improve retrieval efficiency. EARM stores sparse query-memory relevance scores in an online matrix, learns their shared structure via causal matrix completion, and combines newly observed scores with estimates for reranking. Experiments on long-term conversational memory demonstrate that mixed observed-and-estimated reranking boosts answer accuracy by up to 6.62% over semantic retrieval while reducing LLM scoring overhead to 17.5% of candidates. The results advocate for agent memory systems that retain both content and retrieval utility.

experience-amortized rerankingcausal matrix completionlong-term agent memoryllm rerankingsemantic retrieval

Compositional Chain-of-Relations for Faithful Knowledge Graph Question Answering with Large Language Models

arXiv cs.AI · Chenhui Liu, Jianpeng Zhou, Jiahai Wang · 2026-08-24

The paper introduces Compositional Chain-of-Relations (CCoR), a relation-centric framework for faithful knowledge graph question answering (KGQA) with large language models (LLMs). CCoR addresses limitations in entity-centric methods—unreliable pruning and ungrounded constraint handling—by using relation chains for candidate retrieval (main chain) and constraint verification (constraint chain), grounded explicitly in the KG. Evaluated on four KGQA benchmarks, CCoR improves accuracy, faithfulness, and efficiency, particularly for complex queries requiring multi-hop reasoning.

knowledge graph question answeringmulti-hop reasoningrelation-centric explorationlarge language modelsconstraint handling

Don't Repeat Yourself: Stopping Verbatim Loops at Sampling Time

arXiv cs.AI · Philipp Emanuel Weidmann, Allen Roush, Judah Goldfeder, Sanjay Basu · 2026-08-24

The paper introduces Don't Repeat Yourself (DRY), a sampling-time logit adjustment method to mitigate verbatim looping in autoregressive text generation by penalizing tokens that extend current suffixes into exact continuations of earlier context spans. DRY employs sequence breakers to protect chat templates and formatting tokens while maintaining fluency. Evaluations across models (1.5B-120B parameters) show a 47% reduction in suffix-extension rates and improved lexical diversity, with no performance degradation on MT-Bench, MMLU, or GSM8K. DRY has been integrated into major inference frameworks like llama.cpp and ExLlamaV2.

autoregressive generationverbatim loopinglogit adjustmentsuffix matchinglexical diversity

XTC: Head-Aware Sampling by Excluding Top Choices

arXiv cs.AI · Philipp Emanuel Weidmann, Allen Roush, Judah Goldfeder, Sanjay Basu · 2026-08-24

The paper introduces XTC (Exclude Top Choices), a head-aware decoding operator for autoregressive language models that targets scenarios with multiple plausible continuations by removing dominant tokens exceeding a plausibility threshold τ. XTC operates by probabilistically excluding top choices (ρ) and retaining weaker alternatives before renormalization. Evaluated on Gemma 3 27B Q4, Gemma 3 12B Q6, DeepSeek R1 14B Q6, and Llama 3.3 70B Q4, XTC improves diversity (Distinct-2 +11–15%) and reduces repetition (repeat trigrams -27–47%), with additive gains when combined with temperature scaling. Human and GPT-4o evaluations confirm improved creativity without fluency loss, and XTC preserves prompt-level accuracy (IFEval -1.7 points) while enhancing diversity.

xtchead-aware decodingautoregressive modelsdiversity-repetition tradeoffplausibility threshold

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

arXiv cs.AI · Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan · 2026-08-24

The paper introduces Object-Uni, a unified model for object-centric spatial understanding and controllable generation, addressing the limitation of existing models in representing continuous object poses and generating geometrically consistent images. The method formulates spatial intelligence as a unified problem, treating object pose as an explicit geometric variable and proposing a viewpoint-based orientation abstraction for multimodal integration. Evaluated on the UniSpatial-80K benchmark, the model demonstrates improved object-level pose understanding and pose-controllable generation, advancing unified models from object description to spatial state manipulation.

object-centricspatial understandingpose perceptioncontrollable generationmultimodal integration

The Compaction Cliff in Long-Running AI Agent Memory

arXiv cs.AI · Saber Zerhoudi, Jelena Mitrovic, Michael Granitzer · 2026-08-24

The paper introduces Knowledge Triage, a framework addressing the Compaction Cliff phenomenon where AI agents lose critical safety rules during context compaction. The method classifies knowledge base lines by type, applying type-specific retention policies via three deterministic operators: TypeCompact (in-place rewriting), TypeDecompose (partitioning with rule replication), and TypeRetrieve (external storage fetching). Evaluations on five corpora show 2-4× better safety rule preservation than single-shot LLM compactors (96% recall over 5 rounds), 0% locality violations (vs 93% baseline), and 100% recall@50 (vs 73% baseline). Downstream benchmarks demonstrate significant improvements in medical compliance (p < 10^-8), retail tasks (p < 0.01), and airline domains (p = 0.024).

compaction cliffknowledge triagetypecompacttypedecomposetyperetrieve

DiaRelay: Relaying Dialogue Context with a Constant-Size Memory for Emotion Recognition in Conversation

arXiv cs.AI · Zihao Zhou, Bin Yang, Jinghui Qin, Kebing Jin · 2026-08-24

The paper introduces DiaRelay, a lightweight adapter for LLMs that enhances Emotion Recognition in Conversation (ERC) by maintaining a dialogue-level memory. Building on LoRA, DiaRelay incorporates Selective Relay Memory Transition to aggregate historical evidence into a bounded memory and Dual-axis Relay Memory Read to dynamically modulate feature transformations. Evaluations on MELD and IEMOCAP show state-of-the-art weighted F1 and accuracy with only 7.1M additional parameters, demonstrating effective context-aware emotional understanding without expanding context length.

emotion recognition in conversationloraselective relay memory transitiondual-axis relay memory readcontext-aware adaptation

LLM-Based Selection of Incongruent Verbal and Nonverbal Behavior for Virtual Humans

arXiv cs.AI · Parisa Ghanad Torshizi, Stacy Marsella · 2026-08-24

The paper contributes a taxonomy of verbal-nonverbal incongruence categories based on Ekman's framework, addressing how virtual humans can exhibit realistic mismatches between speech and behavior. It evaluates large language models (LLMs) for selecting contextually appropriate incongruent behaviors from dialogue and social context, then assesses their effectiveness through human-subject studies with embodied virtual agents. Results demonstrate LLMs' capability to generate nuanced incongruent behaviors that produce intended observer effects in training scenarios like counseling simulations.

verbal-nonverbal incongruencevirtual humanslarge language modelsbehavior generationsocial context

SEAM: Shot Entity-Attribute Memory for Consistent Short-Drama Generation at Scale

arXiv cs.AI · Jiaqi Liu, Maolin Ran, Xiaoyang Lu, Jian Wang · 2026-08-24

The paper introduces SEAM (Shot Entity-Attribute Memory), a training-free, model-agnostic memory graph for maintaining visual continuity in short-drama generation pipelines. SEAM operates at the prompt-text layer by extracting per-shot states, retrieving prior context via a graph, filtering constraints, and rewriting prompts. Evaluated on SEAM-Bench, it improves cross-episode continuity recall from 0.700 to 0.946 across six text models and achieves a 96.5% director-acceptance rate in production (201 shots), with 21.9 percentage points attributed to its memory mechanism.

short-drama generationvisual continuityprompt rewritingmemory graphentity-attribute memory

TEE-X: TEE-aware Acceleration Framework for Large Vision Models at the Edge

arXiv cs.AI · Kurt M Wilson, Mohaiminul Al Nahian, Abeer Matar A. Almalky, Sadat Shahriyar · 2026-08-24

TEE-X proposes a TEE-aware acceleration framework for secure edge deployment of large vision models, addressing memory and latency constraints in Trusted Execution Environments. The method combines sensitivity-aware modularization and vectorization techniques, optimized for OP-TEE on Arm TrustZone with NVIDIA Jetson AGX Xavier hardware. Results demonstrate minimal accuracy-latency trade-offs, achieving GPU-level inference speeds for Vision Transformers while maintaining security.

trusted execution environmentsvision transformersedge computingvectorizationsensitivity-aware modularization

CacheRouter: A Dual-Path Tool Routing Architecture with Cache-Preserving Main-Model Isolation for Long-Tail Tool Discovery

arXiv cs.AI · Donghui Zha, Lingwei Xu, Linxiao Wu, Yixue Dong · 2026-08-24

The paper introduces CacheRouter, a dual-path architecture for LLM tool routing that resolves the tension between progressive disclosure and prompt caching. It isolates a fixed set of core tools for the main model (preserving KV-cache) while routing long-tail tools through a separate sub-model that performs tool discovery and execution. Automated tool registration supports dynamic updates without cache invalidation. Evaluated on 55 queries and a 30-turn dialogue, the system achieved 90.99-95.2% cache hit rates, reducing input costs to 8-12% of a no-cache baseline under DeepSeek's pricing model.

tool routingkv-cacheprogressive disclosurelong-tail toolscache hit rate

Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf

arXiv cs.AI · Davood Wadi, Yu Ma · 2026-08-24

The study examines position bias in AI agent search behavior, contrasting it with human patterns. Using randomized hotel listings across 5,000 sessions with four large language models, the authors find AI agents inspect results more deeply and without purchase abandonment. Position weakly influences inspection probability, with a non-monotonic pattern (lowest in middle positions). Model heterogeneity exists in whether position affects final choice, but all converge on optimal selections. Results suggest displayed attributes outweigh ranking for agentic search.

position biasai agentssearch rankinglarge language modelsagentic search

Enrich-Retrieve-Rank: Scaling Capability Discovery Beyond In-Context Routing

arXiv cs.AI · Nazib Sorathiya, Daniel Zhang, Bardiya Akhbari · 2026-08-24

The paper introduces Enrich-Retrieve-Rank, a scalable alternative to in-context routing for capability discovery in agent ecosystems with thousands of MATS components. The method employs offline enrichment of sparse metadata into searchable profiles, followed by an online retrieve-then-rank pipeline that avoids candidate invocation. Results show that at scale (N=7,278), the pipeline maintains Match@1 accuracy of 0.39 (vs. 0.12 for in-context routing), reduces cost by 70x versus Full-Ctx, and outperforms Search&Pick by 6.5 percentage points.

capability discoveryin-context routingretrieve-then-rankagent ecosystemsmetadata enrichment

RACO: Reliability-Aware Coarse-Goal Optimization for Inspection-Oriented UAV Vision-Language Navigation

arXiv cs.AI · Sen Wang, Yiming Sun, Jiaxuan He, Pengfei Zhu · 2026-08-24

The paper introduces RACO, a reliability-aware coarse-to-fine navigation framework for inspection-oriented UAV vision-language navigation (UAV-VLN). RACO addresses the unreliability of coarse-goal predictions in existing methods by treating them as runtime hypotheses and using object-level candidate anchors for correction. It also employs scale-adaptive terminal refinement to handle near-miss cases. Evaluated on the LG-UVI benchmark derived from CityNav/CityRefer, RACO improves success rate (SR) by 9.53 and 7.98 percentage points on validation-unseen and test-unseen sets, respectively, while enhancing inspection-region arrival and reducing false verification risks.

uav-vlncoarse-to-fineinspection-orientedreliability-awareobject-centric

Robustness Analysis of Agentic AI to Inconsistent and Incomplete Tool Responses

arXiv cs.AI · Jiachen Xu, Torben Bach Pedersen, Zhongming Yao, Xiaoyu Zhang · 2026-08-24

The study analyzes how agentic AI systems distinguish between incomplete and inconsistent tool responses by examining their decision points. Using a retail customer-service domain, the authors inject controlled faults and analyze model behavior through two channels: log-probabilities of returned content under tool schema versus full trajectory, and action distribution over legal actions. Incomplete returns are identifiable by low schema likelihood and shift action mass toward state-rechecking tools, while inconsistent returns affect likelihood only on context-aligned fields. Recognition is asymmetric, with each fault type detectable in distinct channels but no single channel universal.

agentic aitool schemalog-probabilitiesaction distributionincomplete returns

A-CPES: A Reference Framework for Agentic AI in Cyber-Physical Energy Systems

arXiv cs.AI · Xiaoyu Zhang, Qiuye Sun, Jiachen Xu, Zhongming Yao · 2026-08-24

The paper proposes A-CPES, a reference framework for agentic AI in cyber-physical energy systems, structured as three nested rings: an authorization/accountability frame, an agentic control outer loop, and a six-layer CPES core. The framework addresses the indivisible control loop in energy system operation, which involves optimization problem formulation, infeasibility handling, and multi-party negotiation. The authors identify eight structural failure modes and specify six governance modules to ensure system integrity before deployment. The approach positions agentic AI as the outer loop of control, calling traditional decision models like SCED rather than being subordinate to them.

agentic aicyber-physical energy systemsoptimization loopgovernance modulesstructural failure modes

Hyperbolic Hierarchical Clustering for Visual Representation Learning

arXiv cs.AI · Jianan Wei, Guikun Chen, Zhiyuan Weng, Chunchao Guo · 2026-08-24

The paper introduces ClusterMixer, an interpretable token mixer for vision backbones based on hierarchical clustering in hyperbolic space, addressing the opacity of conventional token mixers like attention or convolution. ClusterMixer explicitly models token mixing through hyperbolic hierarchical clustering, leveraging the space's low-distortion hierarchy embedding properties. The resulting architecture, HCFormer, demonstrates superior performance in image classification, object detection, instance segmentation, and semantic segmentation compared to existing backbones, while maintaining transparency.

token mixerhyperbolic clusteringvision backboneinterpretabilityhierarchical representation

Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs

arXiv cs.AI · Xinyuan Liu, Eren Sadikoglu, Riana Chatterjee, Ransalu Senanayake · 2026-08-23

The paper introduces Physical Agentic AI, an architecture for orchestrating robot crews using LLMs, featuring a semantic planning-execution interface with runtime verification. The framework employs a foundation model planner for task decomposition and robot-skill assignment, while a Robot Orchestrator validates actions against capabilities and constraints. Evaluations on drone-UGV and humanoid-quadruped tasks show retrieval improves skill grounding from 51% to 96%, but enforcement reduces false dispatches to 0%, preventing all eight injected faults from causing motion.

agentic airobot orchestrationskill groundingfoundation model plannerruntime enforcement

Evaluating Inference-Time Defenses Against Package Hallucination in LLM-Generated Code

arXiv cs.AI · Alberick Euraste Djire, Iyiola E. Olatunji, Melissa Tessa, Earl T. Barr · 2026-08-23

The study addresses package hallucination in LLM-generated code through four contributions: (1) identifying methodological flaws that inflate hallucination rates by up to 9.4pp in Python; (2) evaluating seven inference-time defenses across eight models and four languages, with Retrieval-Augmented Generation (RAG) reducing package hallucination rates (PHR) in 18/32 configurations; (3) introducing Package Utility (PU) to assess recommendation quality, finding Greedy decoding optimal for mitigation--utility trade-offs; (4) stress-testing defenses under adversarial prompts, revealing PHR increases up to 45pp (Ruby most vulnerable) and RAG/Self-Refine outperforming decoding-only strategies.

package hallucinationretrieval-augmented generationguided decodingpackage utilityadversarial prompts

CAI-DLLM: Convergence Aware Inference for Diffusion Language Models

arXiv cs.AI · Farhana Amin, Sabiha Afroz, Dimitrios S. Nikolopoulos · 2026-08-23

CAI-DLLM introduces a training-free inference method for diffusion language models that reduces computation by dynamically allocating denoising steps based on first-step confidence. The approach commits stable tokens early, focuses computation on uncertain tokens, and adjusts decoding schedules across output blocks without requiring model retraining. Evaluations on LLaDA-8B-Instruct and Dream-7B-Instruct show speedups up to 44.8x on reasoning tasks (with ≤4.4% accuracy drop), 18.2x on GSM8K (77.41% accuracy), and 13.1x on HumanEval (48.17% pass@1), while reducing energy consumption by up to 95.3%.

diffusion language modelsinference optimizationdynamic decodingconfidence-based pruningtraining-free acceleration

Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

arXiv cs.AI · Florian Rottach, Sebastian Schieferdecker, William Rudman, Randall Balestriero · 2026-08-23

The authors introduce Mol-JEPA, a multimodal Joint Embedding Predictive Architecture for learning molecular world models that addresses limitations in current molecular foundation models, such as chemically invalid augmentations and modality collapse. The framework employs modality masking across diverse data types—including molecular structures, cellular phenotypes, binding affinities, ADMET profiles, and quantum chemistry simulations—to learn robust biochemical representations without relying on suboptimal perturbations. Evaluations demonstrate that Mol-JEPA's latent space prediction yields strong performance across benchmarks, highlighting the benefits of incorporating biochemical context through multimodal learning.

molecular foundation modelsjoint embeddingmodality maskinglatent space predictionbiochemical context

Do Not Copy/Paste: Soft Barriers for Copying in AI-Assisted Programming

arXiv cs.AI · Iyiola E. Olatunji, Alberick Euraste Djire, Jacques Klein, Tegawendé F. Bissyandé · 2026-08-23

The paper introduces soft barriers as a mechanism to manage AI code handoff, focusing on the transition from conversational to executable code. It proposes Unicode output perturbations that disrupt naive copy-paste while preserving readability, measured by Copy-Paste Resistance (CPR). Experiments on HumanEval and MBPP with four LLMs show high CPR but model- and task-dependent effectiveness. A pilot study with 18 participants suggests soft barriers encourage editing over direct transfer, highlighting the need for policy-aware AI code handoff solutions.

soft barriersai code handoffunicode perturbationscopy-paste resistancellm-generated code

GeoRisk-RAG: A Hierarchy-Aware Risk Framework for Improving RAG Reliability through Selective Answering

arXiv cs.AI · Meenu Ravi, Shailik Sarkar, Lulwah AlKulaib, Yordanos Tessema · 2026-08-23

GeoRisk-RAG introduces a hierarchy-aware framework to improve Retrieval-Augmented Generation (RAG) reliability in geospatial domains by addressing geographic validity gaps. The method employs a Directed Acyclic Graph (DAG)-based distance metric for context retrieval and selective answering to mitigate confidently wrong responses. Evaluated on a wildfire QA dataset, it reduces false confidence rates to 0.009 (vs. ~0.090 in baselines) while enhancing human preference alignment, demonstrating safer decision-making for location-dependent queries.

retrieval-augmented generationgeographic validitydirected acyclic graphselective answeringfalse confidence rate

Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains

arXiv cs.AI · Miguel Contreras, Scott Siegel, Subhash Nerella, Jessica Sena · 2026-08-23

The study introduces ICU-REACT, a clinician-developed dataset for teaching LLMs ICU-specific clinical reasoning through retrieval and context-aware processing. Using a clinician-in-the-loop framework, the authors fine-tuned Clin-REACT models (8B-70B parameters) across three model families. Results show Clin-REACT outperforms backbone and general-purpose LLMs on five clinical reasoning benchmarks, including script concordance tests and downstream diagnosis tasks, demonstrating generalization beyond critical care.

clinical reasoningintensive care unitlarge language modelsretrieval-augmented generationscript concordance test

AI-based worker guidance in assembly and disassembly operations using multimodal ego/exo-centric data capture and structured task knowledge

arXiv cs.AI · Vivek Chavan, Jörg Krüger · 2026-08-23

The paper contributes a data-centric approach for extracting structured task knowledge from multimodal expert demonstrations in assembly/disassembly operations. The method jointly encodes temporal and multimodal information from egocentric/exocentric video and narration to derive structured task representations. Evaluation on a real-world disassembly case shows video-based representations outperform static image methods in capturing procedural structure and execution context, demonstrating potential for repair, training, and circular manufacturing applications.

multimodal learningegocentric visiontask knowledge extractionprocedural documentationcircular manufacturing

DeepSAGE: Stage-Aware Reinforcement Learning for Structured CBT Counseling Dialogue

arXiv cs.AI · Qi Zhang, Heajun An, Prakriti Dumaru, Sang Won Lee · 2026-08-23

DeepSAGE introduces a hybrid LLM-DRL framework for structured CBT counseling dialogues, combining stage-aware reinforcement learning with therapeutic intention selection to guide LLM responses. The method models sessions as 11 stages with explicit objectives, using an external controller for stage completion and DRL for intention selection. Evaluated against six alternatives, DeepSAGE achieves higher client engagement, openness, and stage-goal completion efficiency, with expert review confirming plausible emotional trajectories and CBT processes, though clinical effectiveness requires further human evaluation.

reinforcement learningcognitive behavioral therapylarge language modelsdialogue systemstherapeutic intention

Coalition-Aware Skill Reliability for Self-Evolving Agents

arXiv cs.AI · Qiyan Zhao, Xiaofeng Zhang, Bo Liu, Minda Chen · 2026-08-23

The paper introduces coalition-aware reliability modeling for self-evolving agents, addressing two failure modes in skill banks: coalition pollution and cross-domain utility reversal. It proposes Coalition-Aware Skill Selection (CASS) using Shapley marginals for skill accumulation and Unsupervised Skill-Masked Coalition Optimizer (u-SMCO) for transfer learning via label-free masking. Evaluations on LoCoMo, LongMemEval, HotpotQA, and ALFWorld demonstrate improved task performance (12-18% accuracy gains) and robustness to reward noise compared to baseline skill-based agents.

skill bankscoalition pollutionshapley marginalscross-domain transferself-evolving agents

Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video

arXiv cs.AI · Mohammad Sadra Rajabi, Aanuoluwapo Ojelade, Sunwook Kim, Maury A. Nussbaum · 2026-08-23

The study presents a vision-language model (VLM) pipeline for estimating triaxial bilateral hand forces in manual material handling (MMH) tasks from RGB video, eliminating the need for instrumented objects. The method combines task-specific textual cues, pretrained vision-transformer features, and known box mass to localize regions of interest (ROIs) and perform transformer-based temporal regression. Evaluated on 35 participants performing five MMH tasks (6–12 kg loads), the pipeline achieved RMSE of ~4.7–5.6 N (horizontal/mediolateral) and ~10.6–11.0 N (vertical), with multi-camera setups improving peak-force estimation and object ROI inclusion enhancing single-view performance.

vision-language modelmanual material handlingtriaxial force estimationregion of interesttransformer-based regression

Weakly supervised concept Bottleneck Learning for Robust Two stage Object centric visual reasoning

arXiv cs.AI · Sparsh Tiwari, Gesina Schwalbe, Bettina Finzel · 2026-08-23

The paper introduces Dynamic Orthogonal Concept Bottleneck (D-OCB), an object-centric slot-VAE framework for weakly supervised symbolic predicate extraction in two-stage visual reasoning. D-OCB dynamically learns optimal hyperparameters during training, enforces concept subspace independence via correlation penalties, and employs adaptive dimensionality allocation to prevent representation collapse. Evaluations show the method achieves high concept alignment and reasoning accuracy with minimal supervision, matching or surpassing end-to-end approaches.

concept bottleneckslot-vaeweak supervisionobject-centricdimensionality allocation

Clinical Graph-JEPA: Predictive Patient-State Knowledge Graphs for Cognitive Decision Support

arXiv cs.AI · Kushagra Yadav, Nalin Prabhath, Amit Lamba, Goeun Han · 2026-08-23

The paper introduces Clinical Graph-JEPA, a framework for constructing and refining predictive patient-state knowledge graphs from clinical records. The method combines multi-agent relation proposal, ontology-aware normalization, deterministic evidence scoring, and Joint-Embedding Predictive Architecture (JEPA)-based latent refinement. Evaluation on MIMIC-IV data shows that discharge-note context injection improves leave-one-out mean reciprocal rank (MRR) by 31% relative to a note-embedding-free baseline, measured via edge recovery (MRR, Hits@k) and batch-mask evaluation (AUC, MRR).

knowledge graphclinical decision supportjepamimic-ivrelation extraction

Hybrid Panels: Toward Human-AI Collaboration in Survey Research

arXiv cs.AI · Julia Romberg, Tobias Gummer, Gabriella Lapesa, Tanja Kunz · 2026-08-23

The paper proposes hybrid panels, a novel AI-enabled survey infrastructure combining human participants and large language models (LLMs) to address challenges in population surveys (e.g., declining response rates, nonresponse bias). The framework iteratively improves LLM-population alignment by using simulation errors to optimize subsequent survey waves. A pilot study demonstrates the approach, identifying key implementation challenges. The method integrates longitudinal data collection with LLM-based simulation and validation, aiming to reduce costs while maintaining data quality.

hybrid panelslarge language modelssurvey infrastructurenonresponse biaspopulation simulation

CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents

arXiv cs.AI · Jiaxuan Luo, Zhanfeng Liao, Jiayao Teng, Yuan Wang · 2026-08-23

The paper introduces CausalCache, a conditional high-fidelity restoration system for long-horizon GUI agents that optimally allocates a fixed visual-context budget $B$ across interaction traces. The method employs a history-gated key/value (HGKV) adapter to modify only restored history-image tokens while bypassing non-image events, alongside budget-aware selection and difference-in-differences supervision. Results show that CausalCache improves task success by +13 points on OSWorld-Verified over summary-only memory and by +3.7 points on a mobile benchmark, with gains concentrated in memory-critical tasks (+8.6 points) while maintaining performance on controls.

conditional fidelity restorationhgkv adaptervisual-context budgetlong-horizon gui agentsdifference-in-differences supervision

ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation

arXiv cs.AI · Kaustubh D. Dhole, Charles L. A. Clarke, Eugene Y. Agichtein · 2026-08-23

ExecRubrics introduces executable tool-augmented rubrics for verifiable long-form evaluation, addressing ambiguity in natural-language rubrics by encoding evaluation logic as Python scoring functions. The framework provides operational semantics for rubric intent, enabling inspection, execution, and editing. Evaluated on HealthBench, HelpSteer, and ArgQuality, ExecRubrics matches or improves NL rubric baselines with preference accuracies of 53%, 78%, and 92%, respectively, while reducing latency by up to 320×. Integration with NLTK and spaCy further enhances accuracy, offering a transparent alternative to black-box evaluation in high-stakes domains.

executable rubricslong-form evaluationoperational semanticspreference accuracyverifiable scoring

BLADE: Bilevel Low-rank Augmented-Lagrangian Erasure for LLM Unlearning

arXiv cs.AI · Md Toufikuzzaman, Ahmad Mousavi, Dongwon Lee · 2026-08-23

BLADE introduces a bilevel framework for robust LLM unlearning, addressing limitations of existing methods through three mechanisms: a clamped-entropy forget loss with zero gradient beyond sufficient token uncertainty, an asymmetric augmented Lagrangian for adaptive retain protection, and bilevel optimization confined to LoRA adapters. The method demonstrates superior performance, improving composite scores by 6% on TOFU, 9% on MUSE Books, and 7% on KnowUndo, while maintaining stability under 4× scaling and sequential unlearning steps where baselines fail.

llm unlearningbilevel optimizationaugmented lagrangianlora adaptersclamped-entropy loss

Scaling Curriculum Learning For Autonomous Driving

arXiv cs.AI · Cevahir Koprulu, David Paz, Feng Tao, Yuliang Guo · 2026-08-23

CL4AD introduces curriculum learning for batched autonomous driving simulators by framing scenario selection as an unsupervised environment design problem. The method incorporates utility functions based on success rates, behavior realism, and regret estimation to prioritize training scenarios. Experiments in GPUDRIVE show a 77% reduction in wall-clock time and 67% higher sample efficiency compared to domain randomization, achieving 99% success rate a billion steps earlier.

curriculum learningautonomous drivingreinforcement learningunsupervised environment designsample efficiency

STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control

arXiv cs.AI · Mengxi Luo, Changjia Chen, An Cao, Zirong Huang · 2026-08-23

The paper introduces STAGE, a stateful framework for agentic graph execution that combines policy-scoped context with deterministic procedural control. The method confines model judgment to policy-scoped nodes while enforcing procedural flow through deterministic code, providing task-relevant context at each node. Evaluated on SOP-Bench Referral Abuse, τ²-bench domains, and Smart Dispute, STAGE improves task success and reliability, with gains of 7.5–65.7 percentage points in Pass³ metrics on complex workflows like Telecom and Smart Dispute.

agentic graph executionpolicy-scoped contextdeterministic controlprocedural complexitytask success

CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model Trajectories

arXiv cs.AI · Zheyuan Deng, Binghang Lu, Hanqi Feng, Shirley Huang · 2026-08-23

CONTRAMEM introduces a training-free framework for self-evolving procedural memory, leveraging contrasting multi-model trajectories to distill task-relevant procedural distinctions into compact Function and Skill Cards. The method uses outcome variation in correctness, efficiency, recovery, and failure modes as supervision, enabling localized curation rather than append-only accumulation or full-bank rewriting. On GAIA2/ARE tasks, CONTRAMEM more than doubles success rates (26.2% to 55.3%) across GPT-5.5, Claude Sonnet 4.6, and DeepSeek V4 Pro, with transferable gains to Qwen3.7 Plus (18.5 to 35.5). Heterogeneous multi-model trajectories outperform self- or same-model rollouts, driven by contrastive behavioral diversity.

procedural memorymulti-model trajectoriesfunction cardsskill cardscontrastive supervision

HANSARD: A Reference Architecture for Forensic Readiness, Runtime Witnessing, and Graded Attribution in Autonomous Multi-Agent AI Systems

arXiv cs.AI · Christos Sardianos, Iliana Pla, Vasilis Efthymiou, Iraklis Varlamis · 2026-08-23

HANSARD introduces a reference architecture for forensic readiness and accountability in autonomous multi-agent AI systems, addressing attribution laundering through distributed responsibility. The method combines pre-operation readiness profiles, tamper-evident logging at five choke points, runtime PROV-DM-aligned causal graphs, and post-incident replay with Halpern-Pearl causality analysis. Results include graded attribution via compensation-set size and synergy residuals to detect laundering, with separate reporting of cause, responsibility, and accountability under evidentiary tiers.

forensic readinessattribution launderingprovenance forensicsmulti-agent systemshalpern-pearl causality

ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts

arXiv cs.AI · YuanHang Xiao · 2026-08-23

The authors introduce ClawProBench, a trace-aware benchmark for evaluating AI agents in stateful runtime environments, addressing limitations of final-answer-only assessments. The benchmark comprises two tracks: a 102-scenario full profile with live workspace tasks and a 68-scenario frozen holdout with JSON output contracts, scoring agents via execution traces using a safety-gated formula combining correctness, process quality, and efficiency. Evaluation of 68 configurations revealed native-runtime tasks underperformed workspace-live tasks (0.5238 vs. 0.6415), with weak alignment between full-profile and holdout rankings (Spearman 0.1300), demonstrating that final-answer metrics obscure runtime-specific failures.

agent benchmarkingruntime evaluationtrace-aware scoringsafety-gated metricsworkspace tools

When Does AI for PDEs Yield Scientific Evidence?

arXiv cs.AI · Wenshuo Wang · 2026-08-23

The study formalizes a critical evaluation-use mismatch in AI for PDEs, where existing benchmarks prioritize predictive accuracy over evidential support for scientific claims. By extending standard PDE-simulation and inverse-problem benchmarks, the authors enable assessment of whether model outputs provide sufficient evidence for specified claims. Results demonstrate divergent rankings between accuracy and evidential metrics, revealing that current benchmarks may favor methods with weaker scientific support, thus highlighting the need for claim-aligned evaluation frameworks.

pde simulationinverse problemsscientific evidenceevaluation metricsbenchmark design

Small Reasoning Models are Instruction Followers in Function Calling

arXiv cs.AI · Yalda Taheri, Mohammad Hassan Heydari, Erfan Naaman, Afsaneh Fatemi · 2026-08-23

The paper introduces Instruction-Followed Function Calling (IFFC), a framework that decouples function-calling logic from the primary LLM by delegating it to a smaller model operating in an instruction-following context. This approach outperforms native function calling (NFC) and prompt-based function calling (PFC) baselines, particularly for reasoning-oriented LLMs. IFFC maintains robust performance under aggressive quantization, enabling efficient on-device deployment with minimal accuracy loss.

function callinginstruction-followingquantizationedge-computingllm

Functional compatibility as a determinant of persistent neural learning

arXiv cs.AI · Hossein Javidnia · 2026-08-23

The study demonstrates that functional compatibility—the degree to which new learning coexists with preserved behaviors—determines persistent learning in neural networks. Through controlled experiments varying compatibility from matched neural states, the authors show this effect generalizes across learning directions, architectures (convolutional and transformer), and domains (vision and text). Results reveal learning rules differ in efficiency for exploiting compatibility, while retention constraints limit storage capacity. Nonlinear geometry at larger updates disrupts the compatibility continuum. This establishes functional compatibility as a controllable principle for persistent learning, shifting focus from preventing forgetting to identifying safe permanent learning components.

functional compatibilitypersistent learningstability-plasticitynonlinear geometryretention constraints

EMPIRE: Explicit Manipulation Planning as a Learnable Intermediate Representation for Egocentric Hand-Motion Forecasting

arXiv cs.AI · Wen Wang, Ruibing Hou, Hong Chang, Shiguang Shan · 2026-08-23

The paper introduces EMPIRE, a two-stage framework for egocentric hand-motion forecasting that decouples manipulation planning from motion synthesis. Stage I learns explicit manipulation plans from multimodal context to model hand-object interactions, while Stage II synthesizes future bimanual motions using frozen planner representations to prevent gradient interference. The authors also contribute EMPIRE-651K, a dataset with 650,910 training windows across 111 tasks, each annotated with manipulation plans. EMPIRE achieves state-of-the-art performance with 84.53 mm MPJPE and 38.97 mm finger-relative error.

hand-motion forecastingmanipulation planningegocentric visionbimanual interactionintermediate representation

When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing

arXiv cs.AI · Taehyeon An, Jaehyeong Park, Donghyuk Shin · 2026-08-23

The paper introduces Persona-Conditioned Informativeness (PCI), an unsupervised metric for assessing whether persona-conditioned LLM responses exhibit systematic variation. PCI models personas as a similarity graph and uses Local Moran's I to quantify concordant response shifts among semantically similar personas. Evaluated on the PVQ-RR, a PCI-selected 10% subset improved confirmatory factor analysis construct recovery by 0.12 effect size over baselines, demonstrating utility for synthetic respondent screening.

persona-conditioned llmslocal moran's iconfirmatory factor analysisunsupervised metricsynthetic respondents

Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator

arXiv cs.AI · Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban · 2026-08-23

The paper introduces Consensus-Based Calibration (CBC), a label-free method to mitigate rank reversal in multilingual LLM judges by decomposing scores into task difficulty, backbone skill, and language-backbone interaction terms. Using double-centering of the cell-mean score matrix, CBC achieves unbiased estimation even with task-language misspecification, supported by an $O(1/\sqrt{n})$ concentration bound. Evaluations on 7,920 judge runs show CBC improves cross-task rank consistency $τ$ from 0.650 to 0.902 and raises human gold preference agreement from 68.7% to 76.6% on M-RewardBench.

multilingual llm judgesrank reversaldouble-centeringinteraction-recoverylabel-free calibration

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

arXiv cs.AI · Changjiang Jiang, Qiannian Zhao, Lei Xin, Jinxiang Xie · 2026-08-23

The paper proposes Think with Structured Grounding (TwSG), a framework enabling Multimodal Large Language Models (MLLMs) to internalize fine-grained visual perception for chart and visual-tabular understanding without external tools. TwSG combines multi-step reasoning distillation and micro-cropping into a single forward pass, using region-based supervision from teacher-generated VQA data and a two-phase training pipeline (SFT followed by TL-GRPO-based RFT). Experiments show TwSG reduces inference latency while improving accuracy and robustness across MLLM architectures.

multimodal large language modelsvisual question-answeringstructured groundingreinforcement fine-tuningspatial-structural gap

Where World Models Break: Natural-Input Failure Discovery

arXiv cs.AI · Zhanpeng Shi, Zi Liang, Rong Feng, Shiqin Tang · 2026-08-23

The paper introduces BasinLens, a method for discovering natural-input failure modes in world models that conventional average-error evaluations miss. By combining uncertainty-guided global search with typed local replacements, BasinLens efficiently identifies environment-valid condition-action pairs that induce severe prediction failures, verifies their reproducibility, and tests persistence under local edits. Evaluations across diverse benchmarks and world-model families reveal reproducible, persistent failure modes masked by standard metrics, demonstrating systemic risks in world-model-driven control pipelines.

world modelsfailure discoveryuncertainty-guided searchtyped local replacementssystemic risk

Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings

arXiv cs.AI · Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng · 2026-08-23

The paper proposes a Cross-Subject Perceived Speech Decoding (CPSD) framework to improve generalization in decoding perceived speech from non-invasive brain recordings. The method employs a two-stage approach: (1) contrastive learning for shared representation extraction across source subjects, followed by (2) personal specialization via fine-tuning on target subject data. A Positional Encoding-based Spatial Attention (PESA) module standardizes MEG/EEG spatial references. Evaluations on three datasets (Armeni 2022, PKUEEG 2025, Broderick 2018) show Top-10 accuracy improvements of 6.8-15.8% over baselines, demonstrating robust cross-subject generalization.

cross-subject decodingcontrastive learningmeg/eegpositional encodingspeech decoding

LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

arXiv cs.AI · Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican · 2026-08-23

The study demonstrates that GPT-5.4 achieves human-approximate performance in inductive content analysis of open-ended survey responses, with Adjusted Rand Index (ARI) scores of 0.61 (coding) and 0.54 (theme generation) compared to human coders. Researchers compared five human coders and GPT-5.4 on 903 responses from a European PhD student survey using standardized coding procedures. While agreement varied across variables, LLM-human alignment approached human internal consistency (ARI=0.68 for humans, 0.76 for LLM), suggesting LLMs' potential as scalable qualitative analysis tools.

inductive content analysisadjusted rand indexgpt-5.4qualitative codingsurvey text analysis

KONTOGRAPH: Verified Point-in-Time Feature Consistency and Amortised Explanation for Real-Time Anti-Money Laundering under a 200 ms Decision Budget

arXiv cs.AI · Ahmed Abolfadl · 2026-08-23

KONTOGRAPH introduces a real-time anti-money laundering (AML) pipeline for SEPA Instant payments, operating under a 200 ms 99th-percentile latency constraint. The system employs a temporal graph network with per-node memory, achieving a PR-AUC of 0.1717 (+0.166 over a gradient-boosted baseline) on 1.56M simulated payments. Key findings include: (1) per-node memory doubles detection performance, (2) property-based testing uncovered three point-in-time feature inconsistencies in multi-backend execution, and (3) ONNX conversion of tree ensembles altered 0.26% of decisions due to 32-bit floating-point accumulation near cost-optimal thresholds. The work highlights serving-format risks and limitations of subgraph explainers in sparse neighborhoods.

anti-money launderingtemporal graph networkpr-auconnx conversionproperty-based testing

ProBel: Propaganda Detection with Techniques, Spans, and Explanations

arXiv cs.AI · Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Elisa Sartori, Giovanni Da San Martino · 2026-08-23

ProBel introduces a bilingual (Arabic and English) resource for propaganda detection, aligning binary labels, multi-label annotations, technique-labeled spans, and explanations across matched tasks. It evaluates zero-shot prompting, task-specific fine-tuning, and joint training, finding that a single bilingual multi-task model achieves the best overall performance. Results show that joint classification preserves binary performance, while span-only training weakens sentence-level prediction, and bilingual training enhances stability. The study highlights the impact of supervision level and language transfer on task performance.

propaganda detectionmulti-task learningspan identificationzero-shot promptingbilingual training

Self-Supervised Graph Representation Learning for In-The-Wild Wearable and Smartphone based Emotion Recognition

arXiv cs.AI · Ioannis N. Ziogas, Leontios J. Hadjileontiadis, Ahsan H. Khandoker, Aamna Al Shehhi · 2026-08-23

The paper proposes a self-supervised graph representation learning method for in-the-wild wearable and smartphone-based emotion recognition (WER), addressing label scarcity through graph masking augmentation and subgraph sampling. The multi-task inductive graph neural network combines supervised, semi-supervised, and SSL mechanisms, evaluated on K-EmoPhone via leave-one-group-out cross-validation. Results show accuracy improvements of 4.3% (arousal) and 7.8% (valence) using only 20-25% of labels, demonstrating SSL's effectiveness for WER in low-resource settings.

self-supervised learninggraph neural networkemotion recognitionwearable computingsemi-supervised learning

WAM-OPD: On-Policy Distillation for World Action Models

arXiv cs.AI · Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Zezhi Tang · 2026-08-23

WAM-OPD introduces on-policy distillation (OPD) to improve task performance in video-first world action models (WAMs) by leveraging dense teacher supervision on student-induced histories. The method employs a frozen teacher to label student-generated trajectories with video and action targets, training lightweight adapters via joint video-action losses and action flow-matching regularization. In RoboTwin 2.0 experiments, Flash-WAM improved from 0.0% to 58.3% success on HANDOVER MIC and from 16.7% to 33.3% on PUT OBJECT CABINET, demonstrating OPD's potential as a post-training refinement for WAMs without sparse-reward RL.

world action modelson-policy distillationvideo predictionrobot action generationflow-matching

Analyzing and Mitigating Cross-Lingual Degradation in Multilingual Medical VQA

arXiv cs.AI · Jingbo Wang, Sendong Zhao, Haochun Wang, Bing Qin · 2026-08-23

We propose MedVL-XLRepE, a training-free scenario-aware representation engineering method to mitigate cross-lingual degradation in multilingual medical visual question answering (VQA). By leveraging large vision-language models' (LVLMs) superior English medical VQA capability, MedVL-XLRepE steers non-English representations toward their English counterparts at inference time. Evaluated on a multilingual medical VQA benchmark across eight languages and four scenarios, MedVL-XLRepE consistently reduces cross-lingual degradation, achieving gains of up to 6.33% across three LVLMs. The benchmark isolates core capabilities required for medical VQA, revealing that degradation varies significantly by scenario.

medical vqacross-lingual degradationrepresentation engineeringlarge vision-language modelsmultilingual benchmark

Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video

arXiv cs.AI · Masoud Jalayer, Changyi Li, Yu Xiao · 2026-08-23

The paper introduces audio-first triage for budgeted vision-language captioning of untrimmed egocentric video, reducing computational cost by selecting windows for VLM processing using pre-decoding acoustic features. The method trains a selector to trigger per action rather than per frame, leveraging frozen AudioSet-pretrained features without domain-specific labels. Results show 4.0-10.8 percentage point improvements in action coverage across call rates, reducing VLM calls by 9-20% at matched coverage on EPIC-KITCHENS-100 and outperforming visual keyframe selectors on Ego4D over 247 clips.

audio-first triagevision-language modelegocentric videopre-decodingaction coverage

Addressing the Selection Problem in Explainable AI

arXiv cs.AI · Claire Vlases, Katelyn Morrison · 2026-08-23

The paper formalizes the 'selection problem' in Explainable AI (XAI), where users struggle to map natural-language uncertainties to appropriate explanation techniques due to siloed XAI interfaces. The authors propose a multi-agent LLM orchestration tool to bridge this gap by translating user queries into suitable XAI methods. They demonstrate how this structural solution addresses the selection problem through logical premise-conclusion analysis and provide an implementation example.

explainable aiselection problemmulti-agent llmnatural-language uncertaintyexplanation techniques

SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models

arXiv cs.AI · Qingwen Lin, Boyan Xu, Xiao Liu, Zhifeng Hao · 2026-08-23

The paper proposes State Anomaly Neutralization (SANE), a method to stabilize Delta-Rule recurrent models under extreme-context extrapolation by addressing localized norm explosion. Analysis of RWKV-7 reveals failure patterns from uneven state updates, prompting SANE's adaptive tanh compression at chunk boundaries. Evaluations show SANE maintains baseline performance on 11 short-context benchmarks (no significant degradation) while preventing numerical overflow at 100M tokens (24,000× training length), with optimal thresholds (3 ≤ α ≤ 5) preserving reasoning capability (33.46–35.56 scores). Overly permissive thresholds (α ≥ 8) reveal a capacity–stability trade-off.

delta-rule modelsstate anomaly neutralizationextreme-context extrapolationnorm explosionadaptive compression

Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture

arXiv cs.AI · Francisco M. Arrabal-Campos, Francisco G. Montoya, Alfredo Alcayde, Ignacio Fernández · 2026-08-23

The paper dissects emergent versus computed functions in a minimal cognitive architecture comprising a recurrent reasoner with adaptive halting, homeostatic control, and value modules. Through systematic ablation, it demonstrates that while reasoning competence and halting behavior emerge from gradient descent, value assignment requires explicit computation (+0.151 gain with explicit routing). Experiments on a frozen LLM actuator reveal self-consistency voting provides marginal gains (+0.0236), while inter-sample agreement fails as a stopping signal. The study introduces a falsifiable protocol with positive controls, showing commitment-based value assignment yields 5.1x greater payoff range in cliff-cost scenarios.

cognitive architectureadaptive haltinggradient descentvalue moduleself-consistency voting

Multimodal examination answer data with expert-designed Outcome-Based Education rubrics for criterion-level assessment

arXiv cs.AI · Jahangir Alam SM, Md Khalid Syfullah, Saad Ahmed, Munira Akter Mou · 2026-08-23

The article presents a multimodal dataset of 485 examination answers from 415 students across four institutions, annotated with expert-designed Outcome-Based Education (OBE) rubrics. The dataset includes scanned PDFs with diverse academic content (handwriting, equations, code) and metadata (subject labels, model answers, 47 criteria across 12 rubrics). Data preparation involved source consolidation, score validation, and integrity checks. The resource supports research in automated evaluation, multimodal document understanding, and criterion-level feedback, with access restricted to research use.

outcome-based educationmultimodal datasetrubric annotationautomated evaluationdocument understanding

TransHands: Repurposing Human Pose Encoders as Hand Pose Encoders

arXiv cs.AI · Milo Piccioli, Gianluca Amprimo, Claudia Ferraris, Gabriella Olmo · 2026-08-23

TransHands introduces a transfer learning framework that repurposes pre-trained human pose encoders for 3D hand pose estimation from 2D inputs, addressing data scarcity in hand pose datasets. The method employs a two-stage training strategy with a lightweight input adaptation module to align hand kinematics with body motion representations. Evaluated across four architectures (transformer, graph, and frequency-domain models), it demonstrates consistent accuracy gains, cross-domain generalization, and applicability in egocentric settings.

transfer learning3d hand pose estimationmotion representationinput adaptationegocentric vision

Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms

arXiv cs.AI · Naymul Islam, Nusrat Jahan Lia, Shubhashis Roy Dipta, Sabik Bin Sultan · 2026-08-23

The study introduces BanglaSafe, a 879-prompt Bengali benchmark for LLM safety evaluation, combining natively authored and expert-reviewed prompts across 17 culturally grounded harm categories and five prompting conditions. Evaluating 18 frontier LLMs reveals 53.6% unsafe or partially unsafe responses, with 14.7% strictly harmful, and demonstrates that formal writing style increases harmful request success by 17 percentage points compared to casual phrasing. Existing safety classifiers perform poorly, failing on nearly half of Bengali content.

llm safetyculturally grounded harmsregister shiftssafety classifiersmultilingual evaluation

HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory

arXiv cs.AI · Yuanhua Lin, Yile Li, Zhiyuan Zhao, Jing Shang · 2026-08-23

HERO introduces a Human-profile Enhanced Retrieval Optimization framework for long-term agent memory, addressing information loss and semantic drift in existing methods. The approach constructs a traceable heterogeneous memory graph from raw dialogue text, preserving original context while enabling reasoning. Retrieval combines query-derived anchors with human profiles to adaptively activate relevant graph regions via iterative traversal. Evaluations on two benchmarks demonstrate superior performance in factual and personalized reasoning compared to baselines, with improved access to raw dialogue evidence.

long-term memoryheterogeneous memory graphretrieval optimizationpersonalized reasoningiterative traversal

The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction

arXiv cs.AI · Xunzhe Zhou, Yiyang Cai, Fengyi Wang, Ran Ju · 2026-08-23

The paper introduces The Imitator Game, a four-level benchmark (L0-L3) for evaluating robot imitation ability beyond action prediction, alongside IG-10K, a large paired human-robot dataset (20,000+ episodes, 50+ tasks, 6 domains), and Imitator Arena for human evaluation. The benchmark reveals that state-of-the-art models maintain stable performance from L0 to L2 but collapse at L3, identifying functional substitution as the key challenge for intent-level imitation. Human-video-conditioned models outperform caption-conditioned ones, yet all models achieve below 13% zero-shot success on unseen tasks; fine-tuning with 10 demonstrations yields significant improvements scaling with pretraining size.

robot imitationintent-level imitationfunctional substitutionhuman-robot datasetzero-shot learning

LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model

arXiv cs.AI · Ergan Shang, Weijing Tang, Yinqiu He · 2026-08-23

The paper proposes a contextual multidimensional item response theory (IRT) model for predicting large language model (LLM) performance on unseen questions, combining latent capability profiles with question embeddings. This framework outperforms model-free baselines in within-scenario evaluation, demonstrating that multidimensional latent structure captures capability variation better than unidimensional approaches. However, cross-scenario generalization remains challenging, highlighting a key limitation despite improved interpretability and efficiency in LLM evaluation.

item response theorylatent capability profilesquestion embeddingscross-scenario generalizationllm evaluation

Learning from the Test: Self-Referential Differential Testing for Deep RL Agents

arXiv cs.AI · Junda He, Jieke Shi, Zhou Yang, Mingfei Cheng · 2026-08-23

The paper introduces Delta, a framework for differential testing of Deep Reinforcement Learning (DRL) agents that detects both safety-critical failures and suboptimal policies. Delta operates in two phases: (1) Safety Testing collects decision trajectories from the Agent Under Test (AUT), and (2) Optimality Testing trains a challenger agent via Offline RL (using BC, BCQ, or CQL) on this data, then compares rewards to identify suboptimal AUT behavior. Experiments across five environments show Delta detects 2,518 optimality issues on average, outperforming baselines by 50.2%, with BCQ-trained challengers proving most effective.

differential testingoffline reinforcement learningpolicy optimalitysafety testingchallenger agent

OVIBench: Benchmarking Online Video Question Answering under Interruption

arXiv cs.AI · Naiming Liu, Zhiheng Wu, Shuning Wang, Tie Zhang · 2026-08-23

The paper introduces OVIBench, the first benchmark for evaluating vision-language models (VLMs) in Online Video Question Answering under Interruption, addressing the gap in existing offline, single-round paradigms. It categorizes interruptions into Cancellation, False Trigger, and Correction, supporting both open-ended and multiple-choice evaluations through an offline simulation protocol and multi-dimensional metrics. Experiments show OVIBench effectively differentiates models' interruption-handling capabilities, particularly in following correction requests, and models fine-tuned on the accompanying OVI-Train dataset achieve significant performance gains.

vision-language modelsonline video qainterruption handlingbenchmark evaluationfine-tuning

Length-Adaptive Decoding for Masked Diffusion Machine Translation

arXiv cs.AI · Yan Zhan, Mengkai Hou, Wanting Zhang, Zhijun Gao · 2026-08-23

The paper introduces Entropy-Valley (EV), a training-free length selection method for masked diffusion machine translation that scores candidate target lengths by mean predictive entropy from all-mask forward passes. EV addresses the under-explored challenge of target length selection in diffusion models, which directly affects coverage and redundancy. Results show EV recovers 64.9%, 65.3%, and 33.0% of COMET-22 gains from reference lengths on En→Zh, Zh→En, and En→De, with human evaluation supporting adequacy gains, particularly for Zh→En. EV matches LLaMA-3-8B autoregressive performance on En→Zh and outperforms it on Zh→En.

masked diffusionlength adaptationpredictive entropymachine translationdenoising

Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition

arXiv cs.AI · Sophia Riaz, Haoze Zheng, Amos Roche, Miyu Zhang · 2026-08-23

The paper proposes a linguistically structured approach for non-canonical phoneme recognition by decomposing phonemes into articulatory features (manner, place, voicing) via hierarchical multi-task learning. The method employs task-specific feature heads with cross-attention fusion, enhanced by semi-supervised Momentum Pseudo-Labeling and cascaded training of a pretrained encoder. Evaluated on L2-ARCTIC as a pathological speech proxy, the approach shows significant phoneme recognition improvements over baselines while providing interpretable error patterns aligned with phonological structure.

articulatory feature decompositionmulti-task learningnon-canonical speechmomentum pseudo-labelingcross-attention fusion

GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets

arXiv cs.AI · Saif Ahmed, Ashadulla Hil Galib, S. M. Riaz Rahman Antu, Ahmed Faizul Haque Dhrubo · 2026-08-23

The paper introduces GAN-Diff, a hybrid framework combining pretrained WGAN-GP features with conditional diffusion U-Nets for image restoration. The method leverages frozen WGAN-GP generator features via cross-attention in a diffusion U-Net, addressing instability issues like adversarial learning-rate imbalance and diffusion initialization. Evaluated on CelebA for Gaussian denoising and 2× super-resolution, it achieves improvements of 4.40 dB (PSNR) and 3.70 dB over baselines, demonstrating stable restoration guided by GAN priors.

wgan-gpdiffusion u-netimage restorationcross-attentionddim sampling

Clarify User Expertise: Towards Proactive Conversational Agents Tailoring Responses to User Proficiency

arXiv cs.AI · Zhihong Cao, Chen Huang · 2026-08-23

The paper introduces PASSING, a framework enabling conversational agents to proactively clarify user expertise for response tailoring, addressing limitations in existing systems that infer expertise solely from queries. The method employs What-to-ask and How-to-ask strategies, derived via LLM self-play, to dynamically adapt interactions. Experiments demonstrate PASSING's superiority in personalizing responses based on inferred user proficiency, advancing human-centric agent design.

conversational agentsuser expertiseproactive clarificationllm self-playresponse tailoring

Training-Free VLM Personalization via Calibrated Residual Decoding

arXiv cs.AI · Jiaao Yu, Yujian Ma, Xianming Hu, Pengran Wang · 2026-08-23

The paper introduces a training-free calibrated residual decoding framework for personalizing vision-language models (VLMs) without parameter updates. The method constructs three evidence conditions (positive, counterfactual, and empty profiles) to isolate personalized signals from model priors, using normalized-entropy-based uncertainty calibration to adapt enhancement strength. Experiments on MMPB, YoLLaVA, and MyVLM demonstrate improved performance on identity-sensitive tasks, with entropy calibration stabilizing decoding under uncertain personalization signals.

vision-language modelstraining-free personalizationresidual decodinguncertainty calibrationmultimodal understanding

Improving Few-Step Language Flows with Untied Self-Conditioning

arXiv cs.AI · Bocheng Li, Linli Xu · 2026-08-23

The paper introduces Untied Self-Conditioning, a method to improve few-step generation in flow-matching language models by addressing train-inference mismatches in self-conditioning. The authors identify redundancy in self-conditioning inputs during sampling and propose two corrections: dampening redundant directions using frozen projection weights and approximating step-average predictions from history. Evaluated on LangFlow and ELF-B, the method reduces generative perplexity by 8.6× (from 531 to 62) on OpenWebText and achieves 96% preference in pairwise comparisons under Arena-Hard-Auto v2, with consistent improvements across 8-256 sampling steps.

flow-matchingself-conditioninggenerative perplexitysolver steplatent redundancy

Read Less, Solve More: Token-Efficient Sparse Reading for AI Agents

arXiv cs.AI · Zedong Liu, Jiaan Wu, Xinyang Ma, Le Xu · 2026-08-23

SparseRead introduces a token-efficient reading layer for AI agents that controls content admission before unnecessary evidence enters model context, addressing over-reading in long-horizon tasks. The method combines a regime-aware Read Gate, extensible Reader Backends, and a stateful protocol for bounded evidence acquisition with explicit refinement and fallback mechanisms. Evaluations across six models (including Claude Opus 5) and five workloads show token reductions up to 92.9%, latency improvements up to 89.0%, and maintained or improved task quality, with consistent gains across three agent frameworks.

sparse readingtoken efficiencyread gateevidence acquisitionagent frameworks

Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models

arXiv cs.AI · Zhiming Yang, Zhuoxi Xiong, Donglin Zhou, Wenjun Wei · 2026-08-23

The paper introduces situational illusions, a phenomenon where real-world appearances deviate from physical states, and investigates their impact on multimodal large language models (MLLMs). Authors propose a where-what-how taxonomy and MSIBench, a benchmark assessing MLLMs' discrimination, understanding, and reasoning under illusions. Evaluation of 27 model configurations reveals MLLMs' vulnerability, with 6 failure modes identified. Mitigation strategies include evidence-based prompting for closed-source models and supervised fine-tuning for open-source models, achieving up to 20% performance improvement.

situational illusionsmultimodal large language modelsmsibenchsupervised fine-tuningvisual reasoning

FreKoo++: Learning Continuous Spectral Dynamics for Temporal Domain Generalization

arXiv cs.AI · En Yu, Xiaoyu Yang, Wei Duan, Guangquan Zhang · 2026-08-23

FreKoo++ introduces a continuous spectral-dynamical framework for Temporal Domain Generalization (TDG), addressing multi-scale drift patterns and local uncertainties in irregularly sampled streaming data. The method unifies continuous Koopman modal dynamics with adaptive spectral disentanglement, mapping source-domain parameters to a latent space where their evolution is modeled as learnable continuous modes with complex eigenvalues encoding frequency and temporal dynamics. An adaptive soft spectral weighting mechanism isolates persistent dynamics from noise. Theoretical bounds on modal approximation and generalization are derived. Experiments on discrete and continuous TDG benchmarks show state-of-the-art performance under multi-scale drifts and irregular sampling.

temporal domain generalizationkoopman modal dynamicsspectral disentanglementmulti-scale driftirregular sampling

Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN

arXiv cs.AI · Eliuvish Han Cui · 2026-08-23

The paper proposes a target-aligned validation strategy for allocating scarce amyloid positron-emission tomography (PET) measurements in Alzheimer's disease research, contrasting it with generic uncertainty sampling. Using the A4/LEARN PET archive as a testbed, the method evaluates validation approaches by weighting subjects based on target influence and residual protocol uncertainty. Results show that simple APOE4-balanced validation achieves nearly equivalent confidence-interval width ratios (0.923) to target-specific scoring (0.914) for APOE4 carrier vs. non-carrier contrasts, while generic uncertainty sampling performs worse (0.980). Target-specific scoring yields greater gains for age-slope analysis and cutoff-indexed PET positivity, demonstrating the importance of aligning PET spending with validation claims.

amyloid positron-emission tomographyapoe4-balanced validationtarget-aligned validationresidual protocol uncertaintyconfidence-interval width ratio

Provably adaptive sampling with uniform and remasking discrete diffusion models

arXiv cs.LG · Daniil Dmitriev, Zhihan Huang, Yuting Wei · 2026-08-24

The paper introduces a provably adaptive sampling method for discrete diffusion models with uniform and remasking forward processes, demonstrating that sampling complexity depends on the target distribution's dual total correlation (DTC) rather than ambient dimension. The proposed first-order sampler, based on leave-one-out denoisers, enables parallel coordinate updates and corrects denoising errors dynamically. Theoretical analysis shows O(DTC(X₀)/ε) steps suffice for O(ε_score + ε) error, with discretization error characterized via mutual information in the forward process. Experiments on synthetic data validate the dimension-adaptive behavior.

discrete diffusion modelsdual total correlationleave-one-out denoiserremasking processτ-leaping sampler

Robustness of Anomaly Detection Models for Industrial Control Systems under Training-Time Data Contamination

arXiv cs.LG · Mustafa Umut Ozbek, Taiwo Ojo, Pooria Madani, Khalil El-Khatib · 2026-08-24

The paper evaluates the robustness of 11 anomaly detection models for industrial control systems under training-time data contamination, addressing a gap in prior work that assumes clean training data. Using the SWaT benchmark, three contamination strategies (random injection, similarity-targeted injection, feature-noise injection) with 1-10% budgets are tested in an offline protocol. Results show model-dependent robustness: injection-based attacks degrade local-density and distance-based detectors most severely, while PCA, SVM, HBOS, and IForest remain stable. Neural detectors exhibit intermediate vulnerability, demonstrating that clean-data performance doesn't predict contamination robustness.

anomaly detectionindustrial control systemstraining-time contaminationrobustness evaluationswat benchmark

Inertial Manifold Neural Operator for Dissipative Time-Dependent Partial Differential Equations

arXiv cs.LG · Xiaoyang Xie, Clarence W. Rowley · 2026-08-24

The paper introduces the Inertial Manifold Neural Operator (IMNO), a novel neural operator for solving dissipative time-dependent PDEs that exploits low-dimensional structure inherent in long-time dynamics. IMNO improves physical interpretability, accuracy, and stability in autoregressive prediction compared to standard architectures like Fourier Neural Operator (FNO); a shift-equivariant variant (IMNO-SE) is proposed for shift-equivariant PDEs. Benchmark experiments demonstrate superior performance in numerical evaluations of nonlinear dissipative PDE systems.

inertial manifoldneural operatordissipative pdesshift-equivariantautoregressive prediction

Interpretable AI with Local Distillation

arXiv cs.LG · Erin Craig, Yiling Huang, Snigdha Panigrahi · 2026-08-24

The paper introduces local distillation, a method combining black-box AI accuracy with interpretable local linear models. A teacher model defines locality via outcome-similarity weighting and anchors predictions with pseudo-observations, while a lasso-regularized student model provides sparse local fits. Gaussian randomization enables stability analysis through feature selection frequencies and subgroup clustering. Theoretical guarantees show stable feature selection under response perturbations. Evaluated on 17 benchmarks, the method achieves near-teacher accuracy while producing interpretable models; in cancer gene expression data, it reveals patient subgroups with distinct local gene dependencies invisible to global models.

local distillationlasso penaltygaussian randomizationfeature selectioninterpretable ai

Predicting Multiple Clinical Outcomes Related to Functional Recovery and Social Isolation Among Older Adults After Lower-Limb Fracture or Hip Replacement

arXiv cs.LG · Santosh Ray, Pratik K. Mishra, Ali Abedi, Charlene H. Chu · 2026-08-24

The study proposes a multi-output regression approach to jointly predict five clinical outcomes (Social Isolation Scale, Oxford Hip/Knee Scores, Timed Up and Go, 30-second Chair Stand) for older adults recovering from lower-limb fractures or hip replacements. Using the MAISON-LLF dataset (18 participants, 1,008 participant-days), it extracts 46 daily features from multimodal sensors (motion, acceleration, step count, heart rate, mobility, sleep) and evaluates machine learning and deep learning models. The tabular DL model NODE achieved best performance (MSE=3.96, MAE=1.02), demonstrating superior joint prediction versus single-output approaches. SHAP analysis confirmed the importance of multimodal sensor fusion for recovery trajectory assessment.

multi-output regressionmultimodal sensorsclinical outcome predictiontabular deep learningshap analysis

Primal--Dual Alternating Neural Learning for Timely Classification with Performance Guarantees

arXiv cs.LG · Jiaming Qiu, Yingye Zheng, Ying-Qi Zhao · 2026-08-24

The paper proposes a primal--dual alternating neural learning method for timely risk classification in clinical monitoring, optimizing sensitivity, specificity, and monitoring cost. The approach formulates sequential classification as a multi-objective optimization problem, characterizing optimal decisions via value recursion and enforcing constraints through a recurrent neural network and primal--dual updates. Evaluations on simulated data and continuous glucose monitoring for hypoglycemia risk demonstrate accurate, timely decisions adhering to specified operating characteristics.

sequential classificationmulti-objective optimizationrecurrent neural networkprimal--dual learningclinical monitoring

RAD: Rule-Augmented Relational Anomaly Detection

arXiv cs.LG · Noah Dahle, Anne Tumlin, Ngoc Tran, Xenofon Koutsoukos · 2026-08-24

The paper introduces RAD, a rule-augmented relational anomaly detector that combines heterogeneous graph representation learning with symbolic rule signals for identifying anomalies in multi-table databases. RAD derives interpretable rules from random-forest paths, injects them into a graph model, and learns anomaly scores using reconstruction-based and pairwise-ranking supervision. Evaluated on LANL cybersecurity and user-churn datasets, RAD outperforms tabular and relational baselines in AUROC and AUPRC, with rule injection and ranking supervision identified as key performance drivers.

relational anomaly detectionheterogeneous graph learningsymbolic rule injectionaurocauprc

ProxyFormer: A Dual-Stream Proxy Architecture for Ultra-Long Context and High-Resolution Generation

arXiv cs.LG · Zhongpan Tang · 2026-08-24

ProxyFormer introduces a dual-stream architecture with proxy tokens to address the quadratic complexity of attention computation and KV-cache in ultra-long-context models. The method employs bottom-up compression of local features into proxy states, global interactions in proxy space, and top-down decompression to preserve fine-grained information. With a compression ratio of 64, it extends trainable sequence length from 20K to 0.7M tokens on a 16GB GPU, achieving 92%-95% retrieval accuracy at 1M tokens and 94% at 256K. Preliminary results show applicability to image generation via flow matching.

proxy tokenskv-cachedual-stream architectureflow matchingcompression ratio

Diversity-Based Active Learning: An Evaluation of Metric Spaces for Active Learning Selection

arXiv cs.LG · Siddharth Chilamkur, Dorit S. Hochbaum · 2026-08-24

The paper evaluates diversity-based active learning selection methods across multiple metric spaces, demonstrating that entropy-weighted probability space mapping outperforms alternatives. Using Greedy K-center selection with Random Forest classifiers, the study compares raw feature space, LDA space, and model-derived probability spaces (with/without entropy weighting) on synthetic and real-world datasets. Results show entropy-weighted probability space consistently achieves superior performance in active learning sample selection.

active learninggreedy k-centermetric spacesentropy weightingrandom forest

Traceable Spectral Inference via Influence Functions: Efficient Data Attribution and Error Proxies for the Ariel Mission

arXiv cs.LG · Nikki Grens, Luís F. Simões, Kai Hou Yip, Theresa Lueftinger · 2026-08-24

The paper introduces three contributions for interpretable machine learning in spectroscopy pipelines: (1) reformulating influence functions in terms of prediction rather than loss for label-free deployment, (2) efficient computation of infinitesimal prediction influence via closed-form ridge solutions in Extreme Learning Machines, and (3) deriving an influence-based conservative error proxy by propagating training residuals. Evaluated on simulated spectra, the proposed error proxy shows strong correlation with scale and shape-based spectral errors while enabling identification of influential training samples. The method provides an operational framework for scientific ML applications like ESA's Ariel mission.

influence functionsextreme learning machinesspectroscopy pipelineserror proxydata attribution

Exploring Long-period Architectures: Four New Planet Candidates from Kepler with Periods >342 days

arXiv cs.LG · Matthew T. Hansen, Jason A. Dittmann · 2026-08-24

The study presents a novel single-transit detection pipeline combining a classification convolutional neural network with Kepler spacecraft diagnostics to address the observational bias against long-period exoplanets. Applied to Kepler systems with inner planets exhibiting transit timing variations (TTVs), the method identified four new planetary candidates with periods >342 days and radii 2.74-4.81 R⊕. Two candidates (Kepler 1752.02, 777.78 days; Kepler 199.03, 505.495 days) show multiple transits, while Kepler 1897.02 (≥342 days) and Kepler 1811.02 (≥544 days) exhibit single transits. The candidates alone cannot explain observed TTVs, necessitating follow-up observations.

exoplanet detectionconvolutional neural networktransit timing variationskepler pipelinelong-period planets

The Axiomatic Trader: Latent Regularity, Information Budgets, and the Canonical Form of a Quantitative Investment System

arXiv cs.LG · Jiayu Li · 2026-08-24

The paper formalizes systematic trading by modeling financial market regularities as a time-invariant mechanism with latent states. It proposes that a correct quantitative investment system can be specified through five key constants: recurrence bound Λ at block length b, representation invariance defect ε₀, state coordinate coherence times ℓᵢ, signal ceiling ρ, and regime-contingent fraction κ. These parameters constrain the system architecture while accounting for persistent market patterns.

systematic tradinglatent staterecurrence boundinvariance defectsignal ceiling

Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

arXiv cs.LG · Federico Stella, Fei Jiang, Zhongshi Jiang, Zohar Barzelay · 2026-08-24

The paper introduces a next-scale autoregressive transformer for photorealistic novel view synthesis of human faces, achieving high-resolution multi-view outputs with improved cross-view consistency. The method leverages lower-resolution pre-training and fine-tunes on a synthetic dataset of diverse faces, avoiding 2D pre-training requirements of diffusion models. Results demonstrate superior perceptual fidelity and geometric coherence compared to existing approaches, with additional integration of a transformer-based 3D Gaussian lifting module for accurate 3D face reconstruction.

novel view synthesisnext-scale transformersautoregressive models3d gaussian liftingmulti-view consistency

KellyBoost: Growth-Optimal Portfolio Construction with Gradient-Boosted Trees

arXiv cs.LG · Jiayu Li · 2026-08-24

The paper introduces KellyBoost, a multi-output XGBoost model for growth-optimal portfolio construction, where the softmax output directly represents the asset allocation. The training objective minimizes the negative log growth rate loss function -log(1 + w y), with exact gradient and Hessian computations derived analytically and verified via finite differences. The method provides a dependency-free implementation for computing Kelly-optimal allocations conditioned on input features.

kelly criterionxgboostportfolio optimizationgrowth-optimalhessian computation

Spectrum-Aware Bounds on Invertibility for Privacy-Enhancing Instance Encoding

arXiv cs.LG · Seokjin Hwang, Yuting, Li, Kiwan Maeng · 2026-08-24

The authors introduce a family of tighter, spectrum-aware bounds on invertibility for privacy-enhancing instance encoding, addressing limitations of prior work that only provided loose MSE bounds for randomized encoders. Their method leverages spectral analysis to derive bounds applicable to deterministic encoders and supports multiple norm-based similarity metrics. Empirical evaluation across diverse encoders, datasets, and attacks demonstrates consistent improvement over existing bounds.

instance encodingprivacy enhancementspectral analysisdeterministic encodersnorm-based metrics

Hierarchical Exponential-Gaussian Mixtures for Watch-Time Distribution Prediction

arXiv cs.LG · Sofia Gulevskaia, Mikhail Trapeznikov, Aleksandr Poslavsky, Alexander D'yakonov · 2026-08-24

The paper introduces Hierarchical Exponential-Gaussian Mixture (HEGM), a model addressing limitations in Exponential-Gaussian Mixture Network (EGMN) for watch-time distribution prediction. HEGM employs hierarchical skip-watch decomposition, KL-based variance regularization, and structured initialization to mitigate variance collapse and component redundancy. Evaluated on public and industrial datasets, HEGM improves ranking accuracy and threshold-event prediction while maintaining point-estimation accuracy. A 1.5-month A/B test demonstrated significant engagement lifts. The model enhances mixture stability and interpretability compared to EGMN.

watch-time predictionmixture modelsvariance regularizationharchical decompositionrecommendation systems

Beyond chlorophyll: machine learning estimates of diagnostic phytoplankton pigments from multispectral ocean colour data

arXiv cs.LG · David Moffat, Angus Laurenson, Victor Martinez-Vicente, Gemma Kulk · 2026-08-24

The study demonstrates machine learning's ability to estimate diagnostic phytoplankton pigments from multispectral ocean-color data, surpassing chlorophyll-a-only approaches. Using 33,640 HPLC measurements matched with ESA OC-CCI reflectance data, Random Forest and TabPFN models were compared against chlorophyll-a baselines via temporally stratified validation. Multispectral models consistently outperformed, with improvement magnitude varying by pigment's chlorophyll-a correlation, revealing reflectance's additional discriminatory information. This enables more nuanced phytoplankton community monitoring from satellite data.

phytoplankton pigmentsmultispectral reflectancehplc measurementstemporal stratificationdiagnostic retrieval

Test-Time Adaptation for ECG Classification via SQI-Gated Self-Training and Beat-Rhythm Consistency

arXiv cs.LG · Wenhan Jiang, Zhipeng Deng, Jiale Zhou, Haolin Wang · 2026-08-24

The paper proposes BeatRhythm-TTA, a test-time adaptation framework for ECG classification that addresses domain shifts via SQI-gated self-training and beat-rhythm consistency. The method introduces a signal quality index to filter noisy ECG artifacts during adaptation and enforces dual-level consistency on beat morphology and rhythm dynamics. Evaluated on PTB-XL (source) and CPSC2018/Georgia (target) datasets, it achieves +2.70% relative Macro-F1 improvement over baselines in multi-label ECG diagnosis tasks.

test-time adaptationecg classificationsignal quality indexbeat-rhythm consistencydomain shift

Towards Actionable Surgical Team Dynamics: from Teamwork to Counterfactual Annotations

arXiv cs.LG · Vincenzo Marco De Luca, Antonio Longa, Andrea Passerini · 2026-08-24

The authors introduce an extended multimodal dataset for surgical team interaction analysis, addressing fragmentation in existing datasets by providing speaker diarization, transcripts, and multi-level annotations. The dataset captures team performance, interaction processes, and individual characteristics using standardized evaluation protocols and structured rating schemes. Counterfactual annotations are added to analyze coordination breakdowns and performance variability, alongside temporal and relational representations for computational modeling. This unified resource supports the study of individual actions, interaction patterns, and team-level processes in high-stakes surgical environments.

multimodal datasetsurgical teamworkcounterfactual annotationsteam performancecomputational modeling

Beyond Point Predictions: Uncertainty-Aware Satellite Poverty Mapping for Public Policy

arXiv cs.LG · Markus B. Pettersson, James Bailie, Mohammad Kakooei, Eagon Meng · 2026-08-24

The paper introduces an uncertainty-aware earth observation-machine learning (EO-ML) method for poverty mapping in Africa, combining simultaneous quantile regression and conformal prediction to produce statistically valid prediction intervals. A spatiotemporal transformer processes Landsat and nighttime-light image sequences, achieving state-of-the-art point-prediction performance (R²=0.75) but revealing wider-than-expected uncertainty intervals. The authors propose a policy-oriented aid allocation procedure that integrates ground-truth surveys and model predictions, provably limiting exclusion risk below a preset threshold. Simulations show this approach outperforms alternatives in aid delivery efficiency, demonstrating EO-ML's reliability as a supplement to traditional data when uncertainty is properly accounted for.

conformal predictionquantile regressionspatiotemporal transformerearth observationuncertainty quantification

Spicing up Genetic Netlist Generation with LLMs

arXiv cs.LG · Stefan Uhlich, Yağız Gençer, Andrea Bonetti, Arun Venkitaraman · 2026-08-24

LLM-SPICEMixer introduces a hybrid framework for analog circuit synthesis by augmenting genetic netlist generation with IGEL, an LLM-based proposal operator. The method prompts an LLM with elite circuits to generate new SPICE netlists, maintaining simulation as the ground truth while leveraging structured LLM proposals. Evaluated on transistor-level Iris classification circuits, LLM-SPICEMixer improves median training and test rewards by 8.4% and 8.8%, respectively, with the best circuit achieving 93.3% nominal and 85.9% corner-case test accuracy.

genetic algorithmsspice simulationllm guidanceanalog synthesisnetlist generation

ADDA: a Modular Framework for Representing, Simulating and Assimilating Dynamics with End-to-end Differentiability

arXiv cs.LG · Anthony Frion, Vien Minh Nguyen-Thanh, Ali Can Bekar, Pauleo R. Nimtz · 2026-08-24

The authors introduce Automatic Differentiation for Data Assimilation (ADDA), a PyTorch-based modular framework for differentiable simulation and data assimilation in dynamical systems. ADDA provides base classes for system states, simulations, and observation operators, supporting collocated/staggered grids, unstructured meshes, Lagrangian variables, and irregular observations with parallel processing and automatic differentiation. The framework includes implementations of 10 differentiable dynamical systems for benchmarking, demonstrating compatibility with both PyTorch and JAX-based gradient computations. Publicly available code enables comprehensive comparisons across variational, ensemble, and learning-based assimilation methods.

data assimilationautomatic differentiationdynamical systemsunstructured meshesparallel processing

Poisson Subspace Clustering: Focusing on the Essentials in Count Data

arXiv cs.LG · Collin Leiber, Kai Puolamäki, Heikki Mannila · 2026-08-24

The authors propose 3CPO, a Poisson-based subspace clustering algorithm for count data that simultaneously identifies cluster labels and relevant feature subsets. The method employs an iterative maximum a posteriori optimization to model count distributions while enhancing interpretability through column selection. Experiments on gene expression, text, and economic data demonstrate robust performance in identifying meaningful clusters within subspaces, outperforming generic clustering approaches.

poisson distributionsubspace clusteringcount datamaximum a posteriorifeature selection

A Multidimensional Data-Driven Hybrid Transformer Framework for Non-invasive Continuous Blood Pressure Prediction

arXiv cs.LG · Yuexin Ma, Jingqi Hou, Yuxuan Kang, Zhaoying Liu · 2026-08-24

The study introduces a hybrid Transformer framework for cuffless continuous blood pressure (BP) estimation, combining ECG/PPG-derived physiological descriptors and demographic covariates. The proposed Multi-Source Temporal Encoder Module integrates Transformer, Kolmogorov-Arnold Network, and XGBoost branches to capture temporal, nonlinear, and tabular patterns, while a Dynamic Conditional Fusion-Decoder employs multi-head attention and gated residual correction. Evaluated on 2,431 test segments from MIMIC-III, the model achieved mean errors of 0.41±3.74 mmHg (diastolic) and -1.60±5.95 mmHg (systolic), with 98.48% and 94.36% of predictions within 10 mmHg error, outperforming baseline methods.

transformerblood pressure estimationkolmogorov-arnold networkppgmulti-head attention

From Multimodal Observation to Interpretable Suggestions: Counterfactual Time-Expanded Relational Modeling of Surgical Teams

arXiv cs.LG · Vincenzo Marco De Luca, Antonio Longa, Giovanna Varni, Andrea Passerini · 2026-08-24

The paper proposes a tempo-relational framework for modeling surgical team dynamics using Time-Expanded graphs, addressing the gap in AI solutions that neglect team interactions. The method captures relational structure and temporal evolution from multimodal observations, enabling robust performance in low-data surgical settings. A counterfactual procedure generates interpretable suggestions by identifying minimal behavioral changes linked to improved team performance, validated through simulated surgical procedures with enhanced predictive accuracy across diverse interaction goals.

time-expanded graphscounterfactual proceduresurgical team dynamicsmultimodal observationtempo-relational modeling

The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search

arXiv cs.LG · Peiyang Liu, Xi Wang, Di Liang, Wei Ye · 2026-08-24

The paper addresses two bottlenecks in Retrieval-Augmented Generation (RAG): flawed measurement of evidence utilization and suboptimal context allocation. It introduces a causal leave-one-out probe to accurately measure generative reliance and calibrate LLM attention dilution, replacing flawed relevance proxies. For allocation, it proposes iterative compute distribution across sequential generations, achieving 16.7--20.5 percentage point recall gains on models up to 32B parameters. A closed-loop submodular scheduler with attribution-steered contrastive decoding enforces fresh evidence integration, outperforming open-loop baselines.

retrieval-augmented generationcausal measurementcontext allocationcontrastive decodingsubmodular scheduling

Which Histories Matter for Time Series Forecasting? Learning Predictive Relevance with Future Supervision

arXiv cs.LG · Yong-Hoon Choi, Youngjin Cho · 2026-08-24

The paper introduces predictive relevance for time-series forecasting, defined as expected future utility conditioned on inference-time information, using future supervision only during training. The method combines a normalized-pattern retriever for coarse candidate selection with a lightweight residual MLP that learns listwise future-compatibility targets while maintaining past-only scoring at inference. Evaluated across six benchmarks, the reranker improves Pattern retrieval and outperforms SARAF on all 12 tasks, with gains driven by correct future supervision in query-specific regimes. Diagnostics reveal domain-dependent relevance structures, with future-supervised relevance excelling in query-specific cases (e.g., Solar) while L2 rules remain superior in some domains.

predictive relevancetime-series forecastingfuture supervisionlistwise rankingquery-specific regimes

SGHA: A Single-Loop Fully First-Order Algorithm for Nonconvex-Strongly-Convex Bilevel Optimization

arXiv cs.LG · Zhihao Gu, Qilong Wu, Junchi Yang · 2026-08-24

The authors propose SGHA and Stoc-SGHA, novel single-loop algorithms for nonconvex-strongly-convex bilevel optimization, achieving improved oracle complexity bounds. The method reformulates the problem with lower-level stationarity constraints, applies Smoothed Gradient Descent Ascent with Hessian-vector product approximations via gradient finite differences. Deterministic SGHA attains $O(\barκ_y^{5}ε^{-2})$ complexity, while stochastic Stoc-SGHA achieves $O(\barκ_y^{17}ε^{-6}ρ^{-3})$ (high probability) and $O(\barκ_y^{11}ε^{-4})$ (expectation) under additional assumptions, matching lower bounds in $ε$-dependence.

bilevel optimizationoracle complexitynonconvex-strongly-convexsingle-loop algorithmsmoothed gradient descent ascent

A Comparative Study of Label-free Representation Quality Metrics in Deep Learning

arXiv cs.LG · Daniel Richards Arputharaj, Daniel Jönsson, Gabriel Eilertsen · 2026-08-24

The study evaluates label-free metrics for assessing representation quality in deep neural networks, categorizing them into three families and analyzing their reliability across diverse configurations. Through synthetic experiments and evaluation on 260 vision models across six datasets (object classification, scene recognition, geospatial tasks), the authors compare metrics against downstream task accuracy. Results indicate intrinsic dimensionality (ID) is the most reliable predictor, though performance varies by architecture class and training objective. The work clarifies the applicability and interpretation of label-free metrics in practice.

label-free metricsrepresentation qualityintrinsic dimensionalityvision modelsdownstream accuracy

When More Modalities Hurt: Modality Dropout for Heavy-Duty Vehicle Engine Diagnostics

arXiv cs.LG · Adeel Zafar, Slawomir Nowaczyk, Hamid Sarmadi, Saeed Gholami Shahbandi · 2026-08-24

The paper introduces modality dropout for multimodal fusion in heavy-duty vehicle diagnostics, demonstrating improved engine component classification over naive fusion. The method randomly disables text, sensor telemetry (80% missing values), or Diagnostic Trouble Codes during training, forcing models to exploit weaker modalities. On a proprietary dataset (885 samples, 5 engine components), modality dropout achieves 68.8% accuracy (3.5pp over text-only) with varying modality dominance per fault type: sensors reach 93% for intake/exhaust faults while fusion triples fuel system accuracy from 15% to 38%.

modality dropoutmultimodal fusiondiagnostic trouble codessensor telemetryengine diagnostics

Conformal Risk Minimization for Semi-Supervised Domain Adaptation via Optimal Transport

arXiv cs.LG · Manos Giannopoulos, Yi Shen, Michael M. Zavlanos · 2026-08-24

The paper introduces a novel framework integrating Conformal Risk Minimization (CRM) into Semi-Supervised Domain Adaptation (SSDA) to enhance model reliability in healthcare applications. By leveraging Optimal Transport (OT) for pseudolabeling unlabeled target data, the method enables CRM to operate with limited labeled target data, optimizing for both domain invariance and conformal efficiency. The resulting model produces compact, coverage-valid prediction sets, adhering to domain-specific constraints, as demonstrated in skin lesion classification.

conformal predictiondomain adaptationoptimal transportrisk minimizationsemi-supervised learning

Activation-Weighted Seeded Residual Coding for Low-Bit LLM Weight Repair

arXiv cs.LG · Zehao Liu, Chuangchuang Fang, Yang Ren · 2026-08-24

The paper introduces Activation-Weighted Seeded Residual Coding (AWSRC), a compact repair codec for low-bit quantized LLM weights. AWSRC encodes residuals from a base quantization (e.g., INT4 RTN) using deterministic seed-generated bases, storing only seed selectors, low-bit coefficients, and scales as sidecar data. Activation statistics prioritize errors impacting layer outputs. On Qwen2.5-3B-Instruct, AWSRC with 0.162 scope-bits/weight closes 88.2%, 78.9%, and 71.3% of the BF16 PPL, KL, and accuracy gaps, outperforming sparse, low-rank, and vector-quantized codecs with a 49.25 MB sidecar (0.8% of BF16 weight payload).

quantizationresidual codinglow-bit repairactivation-weightedseed-generated bases

Quantum Reservoir Computing with Physics-Informed Correction for Reduced-Order PDE Forecasting

arXiv cs.LG · Krishna Bhatia, Harsh, Shalini Devendrababu · 2026-08-24

The paper proposes a hybrid quantum-classical architecture for reduced-order PDE forecasting, combining a quantum reservoir computer (QRC) with a physics-informed neural network corrector (PIC). The QRC predicts latent coefficient dynamics while the PIC refines local rollout windows using physics constraints. Evaluated on Kuramoto-Sivashinsky (chaotic) and Burgers equations, QRC+PIC shows consistent improvement over QRC alone in RMSE, NRMSE, and PDE residual metrics for KS, though Burgers demonstrates simpler baselines remain competitive. Results indicate physics-informed correction enhances quantum reservoir forecasting in benchmark-dependent scenarios.

quantum reservoir computingphysics-informed neural networksreduced-order modelingpde forecastingchaotic systems

One Inverse Step is a Convex Program: Bayes-Limit Calibration of Diffusion Inversion

arXiv cs.LG · Gordei Verbii · 2026-08-24

The paper demonstrates that a single DDIM inversion step serves as an efficient probe for assessing whether a pretrained diffusion model captures local manifold geometry, formulated as the stationarity condition of an explicit potential. The analysis reveals three key findings: (i) solution uniqueness at the Bayes limit, (ii) a schedule-independent contraction rate, and (iii) geometry-dependent convergence domains. Empirical validation shows exact scores reproduce theoretical predictions within 0.54% across three classes, while trained scores exhibit deviations due to Fermi window conflicts and Hessian-Lipschitz constraints. Notably, unconditional ceiling violations (1.26-4.66×) are observed in DDPM CIFAR-10/CelebA-HQ-256 settings.

diffusion inversionbayes limitddimmanifold geometryhessian-lipschitz

Partial-Moment PINNs for Caldeira--Leggett Parameter Learning in Quantum Brownian Motion

arXiv cs.LG · Krishna Bhatia · 2026-08-24

The authors propose Partial-Moment Physics-Informed Neural Networks (PINNs) for parameter estimation in the Caldeira--Leggett quantum Brownian motion model. The method enforces linear ODE constraints via automatic differentiation, incorporates physical priors through positive-semidefinite covariance structure and high-temperature assumptions, and learns from partial moment traces (μ_x, σ_xx, σ_xp). Results show accurate recovery of (ω, γ) parameters, stabilized D_pp estimation, and lower rollout error compared to finite-difference and Kalman--EM baselines. Fisher analysis reveals variance observables are necessary for diffusion parameter identifiability.

physics-informed neural networksquantum brownian motionparameter estimationmoment dynamicscaldeira-leggett model

When a neural surrogate cannot accelerate a solver: runtime share, closed-loop drift, and the economics of uncertainty gating in a stiff coupled simulation

arXiv cs.LG · L. Thümmler, T. Kuroda · 2026-08-24

The study identifies three structural barriers limiting neural surrogate acceleration in stiff multiphysics simulations, using a general-relativistic radiation-hydrodynamics code as a testbed. First, runtime share (16.9% of wall clock) caps acceleration at ~1.2x via Amdahl's law, rendering a 5.8x cheaper surrogate ineffective. Second, offline accuracy metrics (Spearman ρ=+0.73) fail to predict deployment performance under control (ρ=-0.04). Third, out-of-distribution states (73x off-manifold) force deferral rates of 96.8-99.7%, causing a 0.94-0.96x slowdown. A gated run exhibits linear density bias (-19.9% over 6000 steps), revealing directed error accumulation distinct from variance-driven divergence.

neural surrogateamdahl's lawout-of-distributionautoregressive errormultiphysics simulation

Macro-Action Topological Navigation under Noisy Localization using Reinforcement Learning

arXiv cs.LG · Simon Hakenes, Tobias Glasmachers · 2026-08-24

The paper presents a reinforcement learning agent capable of navigating large, photorealistic 3D environments using only visual input, despite noisy localization. The agent employs an object-centric topological controller with an onboard pose estimator, combining ORB features for object recognition with a minimal Extended Kalman Filter (EKF) to fuse motion models and visual measurements. This approach maintains local pose consistency without full SLAM, enabling successful target object navigation in the Habitat simulator under noisy motion execution.

reinforcement learningtopological navigationorb featuresextended kalman filterhabitat simulator

Graph Representation Learning of Lightweight IoT Ciphers

arXiv cs.LG · Jonathan Cook, Sabih ur Rehman, M. Arif Khan · 2026-08-24

The paper introduces a graph representation learning (GRL) framework for analyzing differential cryptanalysis in lightweight IoT ciphers SIMON32 and SIMECK32. The method extracts four differential features from partial Difference Distribution Tables (pDDT) and constructs directed graphs using KNN, DT, and RF models. Results show all models achieve 1.0 precision in identifying high-probability differentials, with KNN providing optimal cluster separation (F1 score) and fastest graph construction (2.3s), while DT/RF yield near-perfect regression paths. The framework generalizes to AND-rotation-based lightweight cryptographic algorithms.

graph representation learninglightweight cryptographydifferential cryptanalysisdifference distribution tablefeistel cipher

Reservoir of Importance: Learning Semi-Structured Sparsity with Differentiable Subset Sampling

arXiv cs.LG · Ha Dinh, Xuan Duy Ta, Khoat Than, Khac-Hoai Nam Bui · 2026-08-24

The paper introduces Reservoir of Importance (RoI), a lightweight framework for learning semi-structured N:M sparsity masks in large language models (LLMs) via differentiable subset sampling. RoI reduces parameter overhead by modeling compact logits for mask selection and sampling without replacement, cutting trainable parameters from combinatorial complexity to O(M). Evaluations on the Qwen2.5 LLM family (0.5-7B parameters) show RoI achieves competitive performance with 1.5-8.75× fewer parameters and lower memory costs, while maintaining hardware-friendly sparsity patterns.

semi-structured sparsitydifferentiable subset samplingn:m sparsitylarge language modelsmask learning

ST$^2$U: Stateful Test-Time Unlearning via Restricted Knowledge Boundary Control

arXiv cs.LG · Xunlei Chen, Qinghui Gong, Ruini Xue, Yaodong Hu · 2026-08-24

The paper introduces Stateful Test-Time Unlearning (ST$^2$U), a method for controlling restricted knowledge in large language models during inference without retraining. ST$^2$U formulates unlearning as trajectory-wide boundary control, modeling restricted knowledge boundaries in low-dimensional coordinates and applying minimal corrections with contextual anchoring while propagating historical states. Evaluated across three benchmarks and model families, ST$^2$U achieves the strongest balance between retention and forgetting, reducing restricted-knowledge re-entry to 13.76%-19.84% versus 46.50%-59.10% for baselines.

test-time unlearningknowledge boundary controlautoregressive generationcontextual anchoringhidden state trajectory

Neural Boltzmann Equations

arXiv cs.LG · Jonas Spinner, Jack Shergold · 2026-08-24

The authors propose Neural Boltzmann Equations (NBEs), a framework combining neural distribution functions, Monte Carlo integration, and natural gradient optimization to efficiently solve high-dimensional Boltzmann equations. NBEs encode particle properties via physics-inspired neural networks, employ importance sampling for phase-space integrals, and leverage natural gradients for system evolution. This approach overcomes scalability limitations of classical quadrature methods on fixed grids. The method is validated through precision calculation of relativistic neutrino degrees of freedom in early universe cosmology.

boltzmann equationsneural distribution functionsmonte carlo integrationnatural gradientphase-space integrals

Channel-Token Attention for Reliable Dynamic Spectrum Access under Bursty Primary-User Traffic

arXiv cs.LG · Krishna Acharya, Dinanath Padhya, Utsab Dahal, Ashish Kandel · 2026-08-24

The paper introduces TACAN, a centralized dynamic spectrum access policy using channel-token attention to coordinate secondary users under bursty primary-user traffic. The method represents channels as tokens containing occupancy history and modulation entropy, processes them with a Transformer encoder warm-started from a greedy policy and refined via proximal policy optimization. Evaluated on a 20-channel network with 60 primary and 4 secondary users, TACAN achieves 92.53% packet-present access success (+2.59 points over greedy baseline), reduces mean delivery delay from 1.208 to 1.123 slots, and narrows the conditional reliability gap from 9.69 to 3.15 points.

dynamic spectrum accesstransformer encoderproximal policy optimizationchannel-token attentionautomatic-modulation-classification

Hierarchy-Aware Semantic Losses for Knowledge Graph Link Prediction

arXiv cs.LG · Filip Kronström, Ross D. King · 2026-08-24

The paper introduces hierarchy-aware semantic losses for knowledge graph link prediction, leveraging ontological class hierarchies through box-embedding-based losses to improve representation learning in graph neural networks (GNNs). The method enforces hierarchical consistency during training, outperforming both standard link prediction models and those incorporating subclass relations as graph edges. Evaluated on AIFB, CoDEx, and BioKG, the approach improves mean reciprocal rank (MRR) by 7.6%, 2.4%, and 15.5% respectively, demonstrating its effectiveness and parameter efficiency.

knowledge graphlink predictiongraph neural networksbox embeddingssemantic losses

Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study

arXiv cs.LG · Guan-Hua Wen, Kuan-Yu Chen · 2026-08-24

This study evaluates the deployment value of time-series foundation models (TSFMs) for industrial monitoring through protocol-aware empirical assessments across three settings: C-MAPSS degradation-risk proxy, MIMII anomalous-sound detection, and BDG2 forecasting-residual diagnostics. The authors compare classical one-class methods, compact neural autoencoders, residual forecasters, and TSFMs (MOMENT-small, Chronos-T5, TimesFM 2.5) in terms of anomaly-ranking performance, risk-horizon sensitivity, and implementation cost. Results show lightweight baselines (TCN-AE, OCSVM) often outperform TSFMs in AUROC/AUPRC (e.g., 0.9570/0.8960 vs. 0.7310/0.3080 for C-MAPSS), with TSFMs exhibiting higher latency and VRAM usage, suggesting they are task-dependent rather than default replacements.

time-series foundation modelsindustrial monitoringanomaly detectionforecast residualsdeployment cost

Stochastic gradient descent with initial regularization

arXiv cs.LG · Nabil Kahalé · 2026-08-24

The paper introduces stochastic gradient descent with initial regularization (SGDIR), deriving dimension-free upper bounds on its expected excess risk for squared loss. The analysis covers both noiseless and noisy cases, with specific focus on averaged and non-averaged SGDIR under moment, source, and capacity assumptions. In the noiseless case, bounds of order $m^{-2}\log^{2}m$ and $m^{-3+ε}$ are established, depending on the source parameter, with matching lower bounds in certain regimes. In the noisy case, SGDIR's expected excess risk is shown to be no larger than ridge regression's, up to a polylogarithmic factor. Numerical experiments validate these theoretical findings.

stochastic gradient descentinitial regularizationexcess riskridge regressioncapacity assumptions

A Commutator Framework for Selective Spectral Alignment in Deep Neural Networks

arXiv cs.LG · Kaj Nyström · 2026-08-24

The paper develops a geometric framework for analyzing selective spectral alignment in deep neural networks, quantifying incompatibility through three commutator families between weight-generated covariance, gates, and backward sensitivities. A layerwise identity decomposes the sensitivity-covariance commutator into four sources: downstream transport, adjacent-layer imbalance, pointwise sensitivity fluctuations, and nonlinear gate-covariance interactions. Results show spectral alignment as a layer-dependent phenomenon governed by transport, interaction, cancellation, and damping, with numerical experiments demonstrating factorization of spectral and activation geometry across depths and widths.

commutator frameworkspectral alignmentneural feature matricescovariance subspacesgradient flow

Stochastic Separability of Embedding Manifolds

arXiv cs.LG · Liqing Zhang · 2026-08-24

The paper establishes a stochastic separability theorem for high-dimensional embedding manifolds of distinct object categories, proving they become linearly separable under non-singularity conditions. Using a two-layer measure concentration analysis, the authors derive projection concentration inequalities via the law of total expectation, unifying estimation bounds. Results show that datasets with distinct means and bounded total variances achieve high-probability linear separability when projected along directions satisfying the non-singularity condition, revealing geometric-statistical properties of deep representation learning.

embedding manifoldsmeasure concentrationstochastic separabilitynon-singularity conditionrepresentation learning

Mapping the Concept Landscape: Structural Perception of Global Distributions for Transparent Data Pruning

arXiv cs.LG · Dongyue Wu, Tao Ma · 2026-08-24

The paper introduces Mapping the Concept Landscape (MCL), a structural perception framework for transparent data pruning that replaces opaque feature embeddings with explicit sample-level graphs representing entities, events, and attributes. By integrating these into a dataset-level graph, MCL quantifies concept rarity and employs a greedy algorithm to maximize coverage of underrepresented concepts during pruning. Experiments show MCL outperforms state-of-the-art methods in pruning efficiency while providing interpretable selection criteria.

data pruningsemantic conceptsgraph representationgreedy algorithminterpretability

Thinking at the Right Size: Amortized Distillation Across Post-Trained LLMs

arXiv cs.LG · Yan Zhou, Sara Kangaslahti, Jonathan Geuter, Nihal V. Nayak · 2026-08-24

The paper introduces ADAPT, a framework for amortized distillation across both model sizes and post-trained variants in LLM families, producing L×K models with a single distillation run. ADAPT combines two-phase distillation (pre-training alignment and supervised fine-tuning) with weight-delta initialization to transfer distillation-induced weight changes across variants. This enables smooth size–performance interpolation and adaptive model-size selection at inference, improving compute–accuracy trade-offs for reasoning tasks.

amortized distillationweight-delta initializationpost-trained variantsmodel size interpolationlong-form reasoning

RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling

arXiv cs.LG · Ziyuan Wang, Bohao Tang, Fei Zhang, Shuo Han · 2026-08-24

RIBOSPAN introduces a 1.61B-parameter bidirectional RNA foundation model pretrained natively on contexts up to 10,240 nt, addressing limitations of existing models for full-length RNA modeling. The model combines dense bidirectional self-attention, single-nucleotide tokenization, and attention-isolated sequence packing. Evaluations demonstrate strong performance in nucleotide reconstruction, long-context representation benchmarks, and frozen RNA-type analysis, with YaRN scaling improving contextual organization. The framework also enables full-length mRNA generation and redesign via a conditioned discrete-diffusion approach.

ribospanlong-context rnabidirectional self-attentiondiscrete-diffusionmrna generation

Mirror descent algorithms with logarithmic barriers

arXiv cs.LG · Alberto De Marchi, Yura Malitsky, Adrien B. Taylor · 2026-08-24

The paper establishes convergence guarantees for mirror descent and proximal mirror descent algorithms using logarithmic barriers as distance-generating functions, addressing cases where solutions lie on boundaries causing Bregman divergence blow-up. The authors introduce a novel technique to handle this blow-up and resolve a gap in relative smoothness theory. They prove both methods achieve a tight $O(\log k / k)$ convergence rate in specific settings and compare their approach with interior-point methods.

mirror descentlogarithmic barrierbregman divergenceconvergence raterelative smoothness

DIME: Query-Efficient Framework for Membership Inference on Diffusion Models

arXiv cs.LG · Tue Do, Daniel Alabi · 2026-08-24

The paper introduces DIME (Denoiser Ideal Membership Error), a query-efficient framework for membership inference attacks on diffusion models, grounded in theoretical analysis of the optimal denoiser's reconstruction error. DIME leverages bias and local crowding terms to infer membership with minimal queries, outperforming prior methods by up to 3× in true positive rate at 1% false positive rate, even with just two queries. Evaluations across multiple datasets (CIFAR-10/100, STL10-U, CelebA, ImageNet) demonstrate its superior performance and efficiency. The study also explores potential defenses against such attacks.

membership inferencediffusion modelsquery efficiencyreconstruction errorprivacy attacks

SAGE: Stability-Aware Graph-Based Ensemble Feature Selection for Explainable Postpartum Depression Risk Prediction

arXiv cs.LG · Md. Rokon Islam Emon, Syed Shariar Alam Shuvo, Shahriar Siddique Ayon, Abdullah Al Mamun · 2026-08-24

The study introduces SAGE, a Stability-Aware Graph-Based Ensemble feature selection system for predicting postpartum depression (PPD) with explainable AI. SAGE combines information-theoretic relevance, PCA-based structure, and graph-based interactions, optimized via a genetic algorithm and GAN-based oversampling. It achieved 87.96% accuracy, 86.32% F1 score, and 0.88 AUC using only 16 features, outperforming baselines. Key predictors include psychological and socioeconomic factors like EPDS and PHQ-9 scores, with LIME providing instance-based explanations for personalized risk assessment.

postpartum depressionfeature selectiongraph-based ensembleexplainable aigenetic algorithm

CatchBench: When Can an Agent Failure Be Caught?

arXiv cs.LG · Yue Zhao · 2026-08-24

CatchBench introduces a novel benchmarking framework for evaluating when agent failures can be detected across three information states: PRE (pre-run configuration), LIVE (partial trace), and POST (complete trace). The benchmark features seven task contracts with distinct labels and metrics, assessing 72 methods including rule scanners, structural models, and 11 LLM judges from nine model families. Results show limited ordering among methods, with only 47 of 118 pre-declared contrasts separating significantly. Key findings highlight potential shortcuts in scoring, emphasizing the need for transparent label processes. All results are reproducible without model calls.

benchmarking frameworkagent failure detectioninformation statestask contractsllm judges

Change Detection in Probability Flow ODE: Online Testing in Diffusion Latent Spaces

arXiv cs.LG · Artem Kraevskiy, Artem Prokhorov · 2026-08-24

The paper proposes a nonparametric change-point detection method for sequential data using diffusion models. By leveraging the probability flow ODE of a pre-trained conditional diffusion model, observations are mapped to a Gaussian latent space where Maximum Mean Discrepancy (MMD) tests for deviations from the null distribution. The approach handles arbitrary distributional shifts without parametric assumptions, employing a Shiryaev-Roberts procedure for online detection with exact threshold calibration. Theoretical analysis establishes MMD's asymptotic distribution as a degenerate U-statistic under the null.

change-point detectionprobability flow odemaximum mean discrepancydiffusion modelsonline testing

Contrastive Representation-Guided Genetic Minority Oversampling for Imbalanced Time-Series Classification

arXiv cs.LG · Wenbin Pei, Yunrong Hao, Zhen Liu, Guan Wang · 2026-08-24

The paper introduces FreMGP, a frequency-domain representation-guided genetic programming approach for oversampling imbalanced time-series data. The method combines contrastive learning for class-discriminative frequency-domain representations with multi-tree genetic programming to generate diverse synthetic minority-class samples. Evaluations on imbalanced time-series datasets show FreMGP outperforms existing oversampling techniques and enhances classifier performance across both traditional machine learning and deep learning models.

time-series classificationclass imbalancegenetic programmingcontrastive learningfrequency-domain representation

GuidedFlow: An Attention-Guided Framework for Anomaly Detection in Additive Manufacturing

arXiv cs.LG · Sosmita Paul, Krishna Roy · 2026-08-24

The authors propose GuidedFlow, an attention-guided normalizing flow model for anomaly detection in additive manufacturing (AM). The method combines a pre-trained ResNet with a Spatio-Temporal Attention Network (SAN) to prioritize contextual cues across multiple scales and frames, addressing limitations in detecting tiny/stringing defects. Evaluated on the AM3D-AD dataset and MVTec-AD benchmark, GuidedFlow outperforms state-of-the-art models in detection accuracy and AUROC.

anomaly detectionnormalizing flowadditive manufacturingattention mechanismspatio-temporal modeling

ReCoG: Reciprocal Co-Evolution for Multimodal Graph Learning

arXiv cs.LG · Rui Xue, Tianfu Wu · 2026-08-24

ReCoG introduces a reciprocal co-evolution framework for multimodal graph learning, jointly optimizing graph structure and node representations through dynamic interaction. The method combines a multimodal graph refiner for edge correction using cross-modal evidence with coupled cross-modal message passing for joint intra-/inter-modality propagation. Evaluations on node classification and link prediction benchmarks show consistent improvements over existing multimodal graph structure learning approaches, demonstrating the importance of co-evolving topology and semantics.

multimodal graph learningreciprocal co-evolutiongraph structure learningcross-modal message passinggraph refinement

Neural Operator based Multi-Field Reconstruction of Inner Solar Boundary State

arXiv cs.LG · Vignesh Kumar Pandian Sathia, Reza Mansouri, Dustin J. Kempton, Pete Riley · 2026-08-24

The authors propose a Local Neural Operator (LocalNO) to reconstruct multi-field solar magnetohydrodynamic states at 30 solar radii ($R_\odot$) from partial observations. Given radial velocity and magnetic field inputs, the method predicts non-radial components, current density, and thermodynamic variables via operator learning, addressing nonlinear, multi-scale spatial couplings. This approach aims to provide complete boundary conditions for heliospheric modeling, outperforming conventional regression and autoencoder models in structured field-to-field transformations.

neural operatormagnetohydrodynamicssolar windboundary conditionsfunction spaces

Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes

arXiv cs.LG · Ege C. Kaya, Aliasghar Pourghani, Mahsa Ghasemi, Vijay Gupta · 2026-08-24

The paper develops optimal-control methods for Joint Markov Decision Processes (JMDPs), which preserve dependence across counterfactual outcomes in coupled-dynamics environments. It introduces a nonparametric distributional Bellman optimality operator for JMDPs and proves convergence in Wasserstein distance to the optimal joint return law when the induced marginal MDP has a unique optimal policy. For the first two moments, convergence is established under weaker conditions permitting multiple mean-optimal actions with shared second-moment fixed points. The work also derives sampled targets for neural approximation.

joint markov decision processcoupled-dynamicswasserstein distancedistributional bellman operatoroptimal-control

LpWM: A Case for Sparse Representations in World Models

arXiv cs.LG · Yilun Kuang, Yash Dagade, Quentin Le Lidec, Lucas Maes · 2026-08-24

The paper investigates the efficacy of sparse representations in world models for planning, proposing LpWM, a joint-embedding predictive architecture regularized with Rectified Distribution Matching Regularization (RDMReg) to produce non-negative sparse codes. It demonstrates that sparse representations, unlike dense ones, simplify dynamics modeling and reduce predictor complexity, evidenced by a 57% improvement in planning success on the PushT task. The learned sparse codes exhibit mode-factored structure, revealing discrete dynamical regimes and continuous state features. This work suggests sparse representations enhance both interpretability and control efficiency.

sparse representationsworld modelsrectified distribution matchingpredictive architecturesdynamical regimes

MOSH-WM: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models

arXiv cs.LG · Zhekai Wang, Haoxiang Huang, Xiang Liu, Zhikang Chen · 2026-08-24

MOSH-WM introduces a mask-grounded soft-Hamiltonian world model for object-centric video prediction, explicitly linking position-like state variables to slot-owned image support via spatial moments of mask-derived regions. The method uses a frozen video-slot encoder to extract slots and masks, computes canonical state (Q) and temporal differences (P), and applies a learned energy function for directional bias. Decoder-compatible slots are reconstructed via a gated composer combining propagated phase states with causal visual context. Evaluated on OBJ3D (30-frame rollout), MOSH-WM reduces LPIPS by 25.0% and spatial MSE by 33.7% versus baselines; on CLEVRER (10-frame rollout), it achieves 14.5% and 18.7% reductions, with slower error accumulation in closed-loop rollouts.

object-centrichamiltonian dynamicsvideo predictionmask groundingworld model

Generative Neural Networks for Sinkhorn Distributionally Robust Hypothesis Testing

arXiv cs.LG · Fenglin Zhang, Teyan Liu, Jie Wang · 2026-08-24

The paper proposes a generative framework for Sinkhorn distributionally robust hypothesis testing (SDRHT), addressing scalability limitations of existing conic programming approaches. By deriving an equivalent conditional-KL-divergence representation and leveraging Brenier's theorem, the method reformulates the max-min dual problem as a convex potential maximization, approximated via Hyper Input Convex Neural Networks (HyCNNs). Theoretical analysis demonstrates HyCNNs' representation power and universality, while experiments show superior accuracy and robustness across varying sample sizes and dimensions.

sinkhorn discrepancyhypothesis testinghyper input convex neural networkstransport mapsdistributional robustness

Spiking Neural Networks for Continuous Control: Neuromorphic Reinforcement Learning in Conventional Computing

arXiv cs.LG · Jessica Hunter, Md Maruf Hossain Shuvo, Krishna Roy · 2026-08-24

The study introduces Spiking Actor Network Soft Actor Critic (SANSAC), a neuromorphically viable reinforcement learning (RL) framework for continuous control tasks, designed to bridge conventional and neuromorphic computing. By replacing the traditional actor network in Soft Actor-Critic (SAC) with a spiking neural network (SNN), SANSAC achieves near-equivalent performance on conventional hardware, validating SNNs for complex RL environments. Results demonstrate competitive performance with traditional neural networks, highlighting SNNs' potential in continuous RL frameworks without hardware-specific optimizations.

spiking neural networksreinforcement learningcontinuous controlneuromorphic computingsoft actor-critic

HAWKEYE: Seeing One Layer Deeper -- A Cohesion-Aware Structural Channel for Temporal Link Prediction

arXiv cs.LG · Jiacheng Ding, Xiaofei Zhang · 2026-08-24

HAWKEYE introduces a cohesion-aware structural channel for temporal link prediction (TLP), addressing the underutilized structure channel in state-of-the-art models. The method leverages 2-hop cohesive bridges, maintaining classical cohesiveness indicators (degree, k-core, k-truss) and forming 2-hop features. As a drop-in replacement, HAWKEYE enhances DyGFormer, improving test AP/MRR by +0.6 to +10.8 points across six datasets and nearly doubling MRR on the tgbl-subreddit benchmark (0.103 → 0.204). The pipeline scales efficiently, processing the 67M-edge tgbl-flight in five minutes per pass. Gains correlate with a graph's 2-hop discriminative AUC, vanishing on degenerate or saturated graphs.

temporal link predictioncohesion-aware2-hopk-corek-truss

EGAMA-RC: Risk-Calibrated Evidence-Gated Adaptive Malware Analysis for Robust and Interpretable Memory-Forensic Triage

arXiv cs.LG · Isaac Kofi Nti · 2026-08-24

EGAMA-RC introduces a risk-calibrated evidence-gated framework for memory-forensic malware triage, addressing uncertainty, novelty, robustness, and interpretability. The method combines SHAP-guided feature refinement, model-pool evaluation, adversarial testing, novelty scoring, and explanation-conditioned evidence with runtime-aware routing. Results show 93.12% sample acceptance at 99.86% accuracy (0.136% false-accept rate), with XGBoost achieving 0.0054/0.0059 ms p50/p95 latency, demonstrating that risk-calibrated routing outperforms classification-only approaches.

malware triagerisk-calibrationshap-guided refinementnovelty scoringmemory-forensic

Maximum-distance nonnegative matrix factorization for unmixing highly mixed grain-size distribution data: A generalization of AnalySize

arXiv cs.LG · Qianqian Qi, Zhongming Chen, Peter G. M. van der Heijden · 2026-08-24

The authors propose maximum-distance nonnegative matrix factorization (NMF) to address the limitations of AnalySize in unmixing highly mixed grain-size distribution data. Unlike AnalySize, which minimizes distances among end members, the new method maximizes their distinctness using a hierarchical alternating least squares optimization algorithm. This generalization of AnalySize is particularly effective for datasets where no observed samples are close to the true end members. Experimental results confirm that the proposed method successfully decomposes highly mixed grain-size distribution data, overcoming the challenges faced by AnalySize in such scenarios.

nonnegative matrix factorizationgrain-size distributionmaximum-distancehierarchical alternating least squaresend members

Lightweight Multi-scale Hierarchical Anomaly Detection and Localization for Geospatial Big Data Applications at the Edge

arXiv cs.LG · Thomas Benton Townsend, Joshua Bean, Benjamin K Tkach, Narcisa Gabriela Pricope · 2026-08-23

The paper introduces a lightweight edge-oriented framework for anomaly detection and localization in geospatial data streams, addressing challenges in storage, processing, and communication. The method leverages the H3 discrete global grid system and multi-scale drill-down logic to reduce computational overhead, achieving a 99.7% reduction in evaluations compared to traditional flat-scan approaches. By filtering noise-induced flickering anomalies at lower resolutions, the framework efficiently identifies spatially-persistent anomalous signals. Results demonstrate its effectiveness in distilling massive geospatial data into actionable insights for critical applications.

anomaly detectiongeospatial datah3 gridmulti-scaleedge computing

Advanced LLM-Enhanced Intent-Based 5G Network Management using Dynamic Semantic Routes

arXiv cs.LG · Thomas Benton Townsend, Dimitrios Michael Manias · 2026-08-23

Proposes a dynamic semantic routing framework for LLM-enhanced intent-based 5G network management, where natural language prompts are parsed into executable network configurations. The system employs multiple encoders for static/dynamic route selection and detail extraction from operator intents, evaluated against realistic prompts. Results demonstrate successful schema formatting and parameter extraction accuracy for both routing approaches in 5G+ core networks.

intent-based networkingsemantic routerdynamic route selection5g corenatural language processing

NeuroPrefetcher: Storage-Aware Sparse LLM Inference via Delta Prefetching

arXiv cs.LG · Nobel Dhar, Md Romyull Islam, Xuechen Zhang, Gongjin Sun · 2026-08-23

NeuroPrefetcher introduces a storage-backed LLM inference system that leverages temporal locality in MLP activity during autoregressive decoding, where 82-85% of active neurons persist between tokens. The system employs a GPU-resident predictor (2.86% of base model parameters) to forecast sparse activity, enabling delta prefetching of only newly needed weights from storage, thus replacing reactive demand paging with model-aware weight movement. On edge hardware, NeuroPrefetcher achieves 7.9-12.0x speedup over llama.cpp under constrained memory budgets.

llm inferencedelta prefetchingtemporal localityautoregressive decodingsparse weights

Q-Learning with Stable Infinite-Dimensional Linear Function Approximation

arXiv cs.LG · Shengbo Wang · 2026-08-23

The authors propose a stable infinite-dimensional linear function approximation framework for Q-learning, addressing instability in traditional methods by using a compact latent metric space and nonexpansive reconstruction and compression operators. The framework ensures a contractive latent Bellman map with a unique fixed point, approximating the optimal Q-function up to representation error. Two stochastic approximation algorithms are introduced, achieving sup-norm convergence bounds of order $\widetilde O(n^{-1/2})$. The method adapts to smoothness and geometry without prior knowledge of the latent space metric, demonstrated via Q-measure-learning and neural weight training applications.

q-learninglinear function approximationstochastic approximationbellman contractionlatent metric space

Learning Generalizable Behaviors for Terminal Agents

arXiv cs.LG · Yihang Yao, Bo Pang, Xuan Phi Nguyen, Ding Zhao · 2026-08-23

The paper proposes River, a training recipe for improving reinforcement learning (RL) in terminal agents by prioritizing reward-signal quality over environment quantity. The authors hypothesize that RL primarily shapes high-level decision-making behaviors (Agentic Compositional Generalization) rather than teaching new low-level skills. River combines environment filtering and process-level behavior regularization, achieving state-of-the-art performance among open-source RL-trained 8B models on four benchmarks (Terminal-Bench-Lite, Terminal-Bench-v2.1). With 30% fewer environments, it improves RL gains by 106% (2B-27B models) on Terminal-Bench-Lite and 30% on Terminal-Bench-v2.1.

terminal agentsreinforcement learningagentic compositional generalizationbehavior regularizationreward quality

GET: Generative Embedding Translation for Medical Image Segmentation

arXiv cs.LG · Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Mahmudul Hasan · 2026-08-23

The paper proposes Generative Embedding Translation (GET), a structured framework for medical image segmentation that transforms image embeddings into mask embeddings within a frozen Stable Diffusion VAE latent space. GET employs a U-Net-style Embedding Translation Network (1.07M parameters) with Mobile Bottleneck Convolutions, Subsampled Self-Attention, and Multi-scale Feature Enrichment for efficient local-global modeling. Evaluated on five medical datasets, GET outperforms generative (GMS), CNN, and Transformer baselines, improving Dice/IoU by 0.93%/1.26%, reducing HD95 by 0.81px, and using 31.41% fewer parameters. Under domain shift (BUS-BUSI), gains increase to 3.51%/3.39% Dice/IoU and 27.37px HD95 reduction.

generative embedding translationstable diffusion vaemobile bottleneck convolutionssubsampled self-attentionmulti-scale feature enrichment

KMGen: A Skill-based Approach for Synthetic Individual Patient Data Generation

arXiv cs.LG · Jalen Jiang, Chufan Gao, Ethan Rasmussen, Stephen Z. Xie · 2026-08-23

KMGen introduces the first end-to-end framework for automated Kaplan-Meier (KM) curve extraction and synthetic individual patient data (IPD) generation from clinical trial records. The method combines an agentic pipeline for KM curve extraction (mean Integrated Absolute Error 0.0151 on 32 plots) with an LLM-based archetype distillation and mechanistic sampler for IPD generation, preserving marginal survival distributions via bootstrap rank-correlation coupling. Evaluated on three oncology trials, KMGen achieves mean integrated KM absolute difference ≤0.051, demographic JSD ≤0.013 for 5/6 slots, and recovers ≥71% of top-15 adverse events by MedDRA term.

kaplan-meier curvesindividual patient dataadverse-event trajectoriesmechanistic samplingclinical archetypes

What AstroPT knows about galaxies, and what that can teach us about LLMs

arXiv cs.LG · UniverseTBD, :, Kshitij Duraphe, Aman Kumar · 2026-08-23

The paper introduces AstroPT, a transformer trained on galaxy images, as a controlled testbed for mechanistic interpretability in LLMs. By probing frozen representations across model variants, the authors demonstrate that galaxy properties emerge in a fixed order matching their known physical difficulty, with pixel-level features (band magnitude) appearing early/shallow and derived quantities (redshift) emerging late/deep. Results show this ordering is invariant to training objectives and scales with capacity only in magnitude, not sequence, while linear probes recover known physical relationships among properties.

mechanistic interpretabilitylinear probestransformergalaxy propertiesemergent concepts

Mitigating Explanation Leakage in Financial Fraud Detection Systems

arXiv cs.LG · Muhammad Waleed Gul, Elaheh Homayounvala · 2026-08-23

The paper proposes DP-FedSHAP, a federated learning architecture that applies client-level differential privacy to TreeSHAP explanations to mitigate membership inference attacks while maintaining regulatory-compliant transparency. The method contrasts with weight-level differential privacy by selectively perturbing only post-hoc explanation vectors rather than model parameters. Evaluated on the IEEE-CIS Fraud Detection dataset, the study quantifies trade-offs between explanation fidelity (via SHAP vector accuracy), privacy guarantees (via MIA resistance), and detection performance (measured by AUPRC).

federated learningdifferential privacytreeshapmembership inference attacksauprc

Adversarial Agents on Topology Optimization: Understanding the Fragility and Robustness of Deep Learning-based and Physics-Based Design Models under Adversarial Perturbation

arXiv cs.LG · Hoang Anh Nguyen, Yuan Hong, Hongyi Xu · 2026-08-23

This work introduces a mechanics-grounded reliability evaluation framework for assessing the robustness of deep learning-based and physics-based topology optimization models under adversarial perturbations. The study employs a non-intrusive threat model, applying bounded perturbations only to the initial-density channel while keeping other components intact. Results show that deep learning surrogates (U-Net, convolutional, and generative architectures) are vulnerable to catastrophic mechanical failure from initialization noise, with compliance increasing by orders of magnitude, while physics-in-the-loop recovery with SIMP optimizer can mitigate performance degradation.

topology optimizationadversarial perturbationdeep learning surrogatesphysics-based designcyber-manufacturing

Scale-invariant Optimal Sampling for Rare-events Data with Sparse Models

arXiv cs.LG · Jing Wang, HaiYing Wang, Qiang Zhang, Hao Helen Zhang · 2026-08-23

The paper introduces a scale-invariant optimal subsampling function for rare-events data in sparse models, addressing inefficiencies caused by data scaling transformations. The method minimizes prediction error using an adaptive lasso estimator and inverse probability weighting, ensuring oracle properties. An additional estimator based on maximum sampled conditional likelihood is proposed to enhance efficiency. Numerical experiments on simulated and real-world datasets validate the approach, demonstrating improved performance in handling rare events with inactive features.

subsamplingadaptive lassoprediction errorrare-eventsscale-invariant

Sparse Additive Off-Policy Evaluation for Reinforcement Learning with Potentially Limited Number of Trajectories

arXiv cs.LG · Tuoyi Zhao, Chengchun Shi, Zhengling Qi, Lan Wang · 2026-08-23

The authors propose a sparse additive framework for nonlinear off-policy evaluation in infinite-horizon reinforcement learning, addressing interpretability and high-dimensional state spaces. Their method models the Q-function using a sparse additive structure, deriving finite-sample error bounds that scale logarithmically with ambient dimension d. Theoretical guarantees hold when either trajectory count or time horizon is sufficiently large, unlike prior work requiring abundant trajectories. A group-sparsity feature screening procedure identifies relevant covariates with high probability. Experiments validate the approach's effectiveness.

off-policy evaluationsparse additive modelsq-function approximationfeature screeningfinite-sample bounds

Tabular foundation models for non-tabular tasks

arXiv cs.LG · Goran Nakerst, John Brennan, Wouter Beugeling, Masudul Haque · 2026-08-23

The study demonstrates that tabular foundation models (TFMs) can achieve competitive performance on non-tabular tasks when data is represented in tabular form. Using TabPFN v3, the authors evaluate classification performance on MNIST digit recognition, French-German language identification, and Tiny ImageNet, treating each as a tabular prediction problem with in-context learning. Without task-specific training, TabPFN v3 matches or approaches the accuracy of specialized models in some cases, despite lacking explicit structural awareness of spatial or sequential data.

tabular foundation modelsin-context learningnon-tabular tasksclassificationpretrained models

GCA: Global Centroid Alignment in Federated Learning

arXiv cs.LG · Jong-Ik Park, Harry Jiang, Logan Blakely, Georgios Fragkos · 2026-08-23

The paper introduces Global Centroid Alignment (GCA), a federated learning protocol for autoencoder-based anomaly detection that reduces communication overhead and enhances privacy. GCA coordinates clients by exchanging only latent codes and centroid statistics instead of model parameters or gradients, employing reconstruction updates, server-side clustering, and inverse-count-weighted local alignment. Evaluated on five tabular and two vision benchmarks, GCA demonstrates superior privacy protection against data extraction attacks (lower reconstruction error in all 21 comparisons), reduces communication costs by up to 99.15%, and improves test accuracy by up to 5.76% over FedAvg.

federated learningautoencoderlatent codescentroid alignmentprivacy preservation

Two-level domain-decomposition AdaGrad method for scalable training of graph neural networks

arXiv cs.LG · Laurynas Varnas, Julien Herrmann, Alexander Heinlein, Serge Gratton · 2026-08-23

The authors propose DD-AG2m and 2DD-AG2m, domain-decomposition variants of the AG2m optimizer for scalable GNN training. These methods alternate between global and partitioned graph optimization, with 2DD-AG2m additionally using coarse-grained global steps via node subsampling. Experiments on graph classification, node regression, and spatiotemporal tasks show 4-8× faster convergence to equivalent accuracy versus AG2m, and up to 22% higher accuracy at fixed compute budgets.

graph neural networksdomain decompositionadagraddistributed optimizationsecond-order methods

Neighbor-embedded Graph Neural Network-based Crowd Delivery Traffic Management in Smart City

arXiv cs.LG · Kishu Gupta, Deepika Saxena, Ashutosh Kumar Singh, Chung-Nan Lee · 2026-08-23

The paper proposes Neighbor-Embedded Graph Neural Network-based Crowd Delivery Traffic Management (NeCDM), a two-component model for smart city traffic optimization. The Traffic Congestion Prediction Unit (TCPu) employs GNNs to predict flow levels at delivery stations, while the Traffic Observation and Management Unit (TOMu) selects optimal delivery vehicles for crowd requests. The system reduces L1 loss by 4.03%, L2 loss by 16.66%, and computation time by 7.64% while addressing smart city parameters like emissions and travel efficiency.

graph neural networktraffic predictioncrowd deliverysmart citycongestion management

Model-Consistent Byzantine-Resilient Decentralized Federated Learning for Collaborative Missions

arXiv cs.LG · Yue Li, Sudip Bhujel, Cameron Lira, Ning Wang · 2026-08-23

The paper introduces DFL-C, a Byzantine-resilient decentralized federated learning (DFL) architecture ensuring global model consistency for collaborative missions. DFL-C integrates an asynchronous common subset (ACS) consensus protocol to enforce uniform model updates and a dual-domain trust scoring mechanism to mitigate data-domain Byzantine attacks. Experiments show DFL-C maintains accuracy under Byzantine behaviors, outperforming BALANCE in untargeted poisoning attacks and matching resilience in backdoor attacks, particularly in non-IID settings.

decentralized federated learningbyzantine resiliencemodel consistencyconsensus protocoltrust scoring

Stress Testing Unlearning Algorithms

arXiv cs.LG · Noam Diamant, Ethan Fetaya, Neta Glazer · 2026-08-23

The paper introduces WMDP++, an enhanced benchmark for evaluating machine unlearning in large language models (LLMs), addressing two limitations in existing benchmarks: insufficient testing of forced information extraction and inadequate assessment of performance on boundary questions. WMDP++ extends WMDP by incorporating targeted extraction probes and systematic boundary question evaluation, providing a more rigorous framework for assessing both safety and utility in unlearning methods. The benchmark aims to drive progress in unlearning algorithms by offering a more stringent evaluation protocol.

machine unlearninglarge language modelsbenchmark evaluationinformation extractionboundary questions

From Detrimental to Beneficial: Dynamic Influence-based Valuation and Editing

arXiv cs.LG · Adrian Nyakairu, Hongfu Liu · 2026-08-23

DIVE introduces dynamic influence-based valuation and editing, transforming detrimental training samples into beneficial contributions by reversing gradient directions during optimization. The framework operates at batch level without modifying raw data, maintaining compatibility with standard learning procedures. Empirical results show consistent improvements in classification performance, data efficiency, optimization stability, and large language model fine-tuning.

data valuationgradient reversalbatch-level optimizationinfluence estimationlarge language model fine-tuning

Interpretable statistical feature engineering for early disruption prediction in the short pulse ADITYA tokamak

arXiv cs.LG · Jyoti Agarwal, Kavit Patel, Bhaskar Chaudhury, Abhishek Sharma · 2026-08-23

An interpretable machine learning framework is proposed for early disruption prediction in short-pulse tokamaks, addressing the limited warning time in devices like ADITYA. The method extracts statistical descriptors (mean, variance, skewness, kurtosis, wavelet energy entropy) from plasma diagnostics and employs decision tree-based feature selection to identify physically meaningful precursors. A random forest classifier is trained on the reduced feature set, achieving stable predictive performance with a maximum ROC-AUC of 0.87 for 0-35 ms and 0-40 ms analysis windows. The framework demonstrates that engineered statistical descriptors can effectively replace raw time series inputs, offering a computationally efficient pathway for real-time plasma control.

tokamakdisruption predictionfeature engineeringrandom forestwavelet energy entropy

From Symmetry to Invariance: Learning Galois Equivalent Representations in Finite Fields

arXiv cs.LG · Zheng Zhang, Na Zhang · 2026-08-23

The paper investigates neural networks' ability to transfer learned algebraic operations across mathematically equivalent representations, focusing on multiplication in finite fields under basis changes. The authors leverage the Galois action to organize basis representations into orbits, enabling separation between learning multiplication and transferring it to untrained bases. They propose training a model to predict the Galois action between bases, using repeated transformations to construct canonical orbit representatives that support exact matching for transfer. This approach converts learned algebraic symmetries into invariant representations, demonstrating a concrete mechanism for transfer across equivalent representations.

galois actionfinite fieldsinvariant representationsalgebraic symmetrybasis transformations

GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented Reasoning

arXiv cs.LG · Jun Chen, Yongchao Liu, Pengyu Qiu, Jiajun Zheng · 2026-08-23

GTA-RAG introduces a graph-trajectory-augmented RL framework for multi-turn retrieval-augmented reasoning, addressing sparse supervision in existing RL-based RAG approaches. The method samples document paths from an entity-document graph, synthesizes multi-hop QA trajectories, and validates them with the deployed retriever to obtain trajectory-level supervision. It optimizes retrieval policy using Group Relative Policy Optimization (GRPO) and a trajectory-guided reward, enhancing both answer accuracy and evidence-chain coverage. Experiments on three multi-hop and two simple QA benchmarks demonstrate consistent outperformance over RL-based RAG baselines with Qwen2.5-3B and Qwen2.5-7B backbones.

retrieval-augmented generationmulti-hop reasoningreinforcement learningentity-document graphpolicy optimization

Arbitrage-Aware Multi-Step Forecasting of Implied Volatility Surfaces: Modelling Surface Trajectories Using Latent Diffusion

arXiv cs.LG · Dominik Manuel Buchegger, Lukas Gonon · 2026-08-23

The paper proposes a conditional latent diffusion framework for multi-step forecasting of implied volatility surfaces and underlying returns. The method combines an arbitrage-aware autoencoder for low-dimensional surface representation with a diffusion model capturing conditional joint evolution. Evaluated on SPX surfaces, the framework generates realistic 30-step probabilistic scenarios and outperforms the persistence benchmark in point forecasting while preserving economic admissibility.

implied volatility surfaceslatent diffusionarbitrage-aware autoencodermulti-step forecastingprobabilistic scenarios

Quantum-Inspired Hybrid Neural Networks for Neural Decoding: A Controlled Ablation Study of Learnable Quantum Sidecar Integration

arXiv cs.LG · Diana Legziel Levy, Menachem Finkelstein, Peter Chin, Eilon Vaadia · 2026-08-23

The study investigates parameterized quantum circuits (PQCs) as residual sidecar modules in a ResNet-50 backbone for 31-class neural population decoding, specifically imagined handwriting classification from multi-neuron spike rasters. Four variants are compared under controlled conditions: baseline, quantum sidecar with frozen input projection, backbone-gradient-trained projection, and a measurement-guided variant. The backbone-gradient variant improves accuracy in 3/4 seeds (+0.19% mean) and reduces Linear CKA similarity to baseline features (Δ=-0.025), indicating structural reorganization. A nine-variant ablation identifies shallow architectures as most effective. All experiments use noiseless statevector simulation on 4 qubits, reflecting near-term hardware constraints.

parameterized quantum circuitsneural decodingresidual sidecarlinear ckastatevector simulation

The Variance of Thought: Policy Variance, Critical Forks, and Local Credit Assignment

arXiv cs.LG · Yingru Li · 2026-08-23

The paper analyzes credit assignment in long-horizon language-model tasks through policy variance $σ_π^2(s)$, identifying critical forks as key sources of variance in deterministic MDPs. It establishes three theoretical results: (i) policy variance as a discovery budget with sample complexity $Ω(c^2/σ_π^2(s))$, (ii) a Gini dispersion upper bound $σ_π^2(s)\le 1-\|π(\cdot|s)\|_2^2$ computable from logits, and (iii) horizon-dependent estimation costs where Monte Carlo advantage SNR scales as $\sqrt{P}$. Bootstrapping with log-value parameterization is proposed to mitigate sample cost scaling.

policy variancecredit assignmentcritical forksgini dispersionmonte carlo advantage

MASH-Bench: Diagnosing Cross-Source Failure in Mass-Shooting Risk Classification

arXiv cs.LG · Neha Sharma, Ritesh Sharma · 2026-08-23

The paper introduces MASH-Bench, a harmonized benchmark of 6,968 mass-shooting incidents from four U.S. databases, to diagnose cross-source failure in risk classification. Using leave-one-dataset-out evaluation, the study shows that Random Forest, XGBoost, and LightGBM achieve 0.68-0.89 recall on curated sources but generalize poorly to Gun Violence Archive (recall: 0.20, precision: 0.0004). Feature-masking ablation reveals feature completeness as a key failure factor, while domain-adaptation methods like DANN improve recall by 0.282 but fail to address low precision. The benchmark highlights feature and label disparities as critical constraints.

cross-source generalizationrisk classificationdomain adaptationfeature completenessleave-one-dataset-out

KPI-Conditioned Generative Design of Automotive Hood Inner Panels: A Two-Stage Retrieval-Generation Pipeline with Surrogate-Based Performance Estimation

arXiv cs.LG · Sudeep Chavare · 2026-08-23

The paper presents a two-stage retrieval-generation pipeline for inverse design of automotive hood inner panels conditioned on key performance indicators (KPIs). The method first identifies feasible topology families via reachability analysis, then uses a conditional variational autoencoder to generate point-cloud geometries within selected families, with performance estimated via neural-operator surrogates. Results show the surrogate achieves aggregate accuracy but exhibits errors comparable to within-class performance differences, establishing the error-to-signal ratio as a critical determinant of pipeline viability. The system is implemented as an interactive tool using public data and freely available compute.

inverse designconditional variational autoencoderneural operatorperformance surrogatetopology family

Geometric Structures on Graphs: a Holonomy-Based Discretization of Curvature

arXiv cs.LG · Hao Li, Yuhan Peng, Junwen Dong · 2026-08-23

The authors propose a holonomy-based framework for discretizing curvature on graphs with local symmetric positive-definite metrics, where each vertex has a fibre metric and edges carry reversible metric-compatible transports. The method uses the normalized logarithm of holonomy around oriented loops as finite-loop curvature observations, then aggregates these via commutator operations and covariant divergence to produce symmetric Ricci-type metric responses. Results include orthogonal-gauge equivariant updates preserving positive definiteness, a learnable parametrization of edge transports, and empirical validation on the unit sphere for holonomy--curvature relations and transport recovery.

holonomycurvature discretizationmetric-compatible transportricci-type responseorthogonal-gauge equivariance

Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation

arXiv cs.LG · Guan-Hua Wen, Kuan-Yu Chen, Hou-Chiang Tseng · 2026-08-23

The paper proposes DSSM-CRF, a dual-scale state-space model with speaker-wise dynamic conditional random fields for conversational speech emotion recognition. The method separates cross-speaker contextual influence and within-speaker emotion evolution using bidirectional state-space models at frame and dialogue scales, followed by independent CRF chains per speaker with corpus-level transitions and utterance-pair residuals. Evaluated on IEMOCAP and MELD, it achieves 75.81% UA/74.90% WA and 54.72% WA/49.31% WF1 respectively, demonstrating benefits from speaker-wise factorization and CRF modeling.

state-space modelingconditional random fieldspeech emotion recognitiondual-scale encodingdynamic transition

Precision-Aware Variable Bit Processing Elements for Hardware-Efficient Systolic Array Designs

arXiv cs.LG · Dantu Nandini Devi, Madhav Rao · 2026-08-23

This work introduces precision-aware variable bit processing elements for systolic arrays, optimizing floating-point multipliers via partial product matrix truncation and compressor techniques across FP32, TF32, and BF16 formats. Using NSGA-II for design space exploration, the method achieves hardware-efficient approximations while maintaining CNN accuracy on MNIST, F-MNIST, and CIFAR-10. Results show 66-92% area reduction, 60-93% power savings, and 21-54% delay improvement versus exact implementations, demonstrating effective resource optimization for error-tolerant applications.

systolic arrayfloating-point multiplierapproximate computingpartial product matrixnsga-ii

Dataset Complexity Shapes Finite-Distance Loss Geometry in Neural Networks

arXiv cs.LG · Jaeyong Bae, Hawoong Jeong · 2026-08-23

The study demonstrates how dataset complexity influences the loss landscape geometry in neural networks by analyzing local entropy around trained solutions. Using an adapted Franz-Parisi construction from spin-glass theory and estimating local entropy via adaptive sequential Monte Carlo, the authors show that increased complexity contracts the effective solution volume near reference points without uniform acceleration. Synthetic experiments reveal this contraction's distance-dependent nature, while real image data and label randomization confirm the qualitative trend. Results indicate dataset structure organizes low-loss neighborhoods at finite distances from solutions.

local entropyloss landscapedataset complexityspin-glass theorysequential monte carlo

Tracing the Unlabeled Storm: Cross-Variable Transfer in a Lagrangian Atmospheric JEPA Framework

arXiv cs.LG · K M Anirudh, S Sandeep, Hariprasad Kodamana · 2026-08-23

The study introduces M-JEPA, a multiscale Monsoon Joint-Embedding Predictive Architecture, which leverages cross-variable proxy learning to improve precipitation forecasting. Pretrained on five continuous atmospheric proxy fields (e.g., outgoing longwave radiation) without rainfall supervision, the frozen representation is transferred to daily precipitation forecasts via a shared decoder trunk with parallel probabilistic and deterministic branches. Compared to direct rainfall training, proxy pretraining reduces CRPS error by 36% (5.54 vs. 7.52 mm/day) and outperforms the 51-member ECMWF ensemble in CRPS (6.81 vs. 6.89 mm/day) and Brier skill (+0.05 vs. -0.04) using only 15.4M parameters.

joint-embedding predictive architecturecross-variable transferlagrangian atmospheric trackingcontinuous proxy fieldsprobabilistic forecasting

Recovering Weighted Tangent Geometry from a Single-Scale Score Field

arXiv cs.LG · Ziqi Zhao, Qingjian Ni · 2026-08-23

The paper introduces a method to recover weighted tangent geometry from a single-scale score field near branch points in data manifolds, where traditional tangent space analysis fails. By modeling the local geometry as a tangent-measure and leveraging Gaussian smoothing, the authors derive an Ornstein--Uhlenbeck eigenfunction equation that transforms score values into a linear system for center and homogeneity degree calibration. Theoretical results show exact recovery of angular measures from spherical-harmonic multipliers and finite-query certificates for direction and weight recovery. Experiments demonstrate convergence in finite-noise settings but highlight a trade-off between score fit and geometry recovery.

tangent-measurescore fieldorstein-uhlenbeckspherical-harmonicbranch point

Gaussian process learning with flow map refinement for parameter estimation in dynamical systems

arXiv cs.LG · Yue Hao, Dongwei Ye · 2026-08-23

The authors propose Gaussian Process Learning with Flow Map Refinement (GPL-FMR), a two-stage framework for parameter estimation in dynamical systems that addresses global inconsistency in gradient matching methods. Stage 1 uses Gaussian process learning to derive a posterior, which informs Stage 2's flow-map refinement via optimization with global dynamical constraints. Evaluated on Van der Pol, Lotka-Volterra, and Lorenz-63 systems, GPL-FMR improves accuracy under scarce, noisy observations compared to local derivative matching approaches.

parameter estimationgaussian processdynamical systemsflow mapgradient matching

Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers

arXiv cs.LG · Yan Wang · 2026-08-23

The paper introduces Adaptive Log-Space (AL) quantization, a block-wise method for memory-efficient optimizer states that preserves exact zeros and adapts per-block ranges. AL8 and AL16 variants are combined with signed-momentum encodings and topology-aware precision selection. Evaluated across 96 runs (214.7 GPU-hours) on TinyLlama-1.1B and GPT-2, AL8 reduces AdamW's optimizer-state memory from 8392.7 MiB to 2119.2 MiB with minimal perplexity degradation (72.90 vs. 72.48 FP32). CAME requires AL16 (86.16 vs. 86.68 FP32), while topology-aware protection reduces Adafactor's late-loss gap from +0.1185 to +0.0159 in GPT-2.

adaptive quantizationoptimizer statesmemory efficiencylog-space encodingtopology-aware

Does a Modern-Handwriting Warm-Up Help Historical Arabic OCR? A Reproducible, Compute-Matched Evaluation on Muharaf and KHATT

arXiv cs.LG · Sumaih Almarshad, Maram Alamri, Dona Aloraini, Fares Altuwaim · 2026-08-23

The study investigates whether modern Arabic handwriting (KHATT) as an intermediate training stage improves historical Arabic OCR (Muharaf) through reproducible, compute-matched experiments. Four ablation runs with varying configurations (base checkpoint, encoder-freezing, epoch budget, etc.) showed inconsistent effects (-17.64 to +14.52 CER), with clean runs indicating negligible impact (-0.25 to +0.94 CER). A compute-matched three-seed experiment revealed KHATT warm-up worsened performance by +2.42 CER versus same-domain control, with only ~0.6 CER attributable to domain mismatch. The authors release SaudiHeritage-OCR for verification.

historical ocrcompute-matchedcerfine-tuningdomain adaptation

StocBench: A Benchmark for Generative Modeling of Stochastic Dynamics

arXiv cs.LG · Sebastian Pfister, Benjamin Holzschuh, Nils Thuerey · 2026-08-23

The paper introduces StocBench, a benchmark for evaluating generative models on stochastic fluid dynamics forecasting, focusing on Kolmogorov flow with stochastic forcing. It compares transport-based models and few-step distillation methods under constrained inference budgets, measuring one-step distributional accuracy and invariant measure preservation via enstrophy spectra. Flow matching excels at high inference budgets, while DPM-2 performs best at low NFE; distilled methods preserve spectra well but show divergent performance between stochastic and deterministic control tasks. Diffusion samplers exhibit setting-dependent spectral preservation.

generative modelingstochastic dynamicskolmogorov flowdiffusion samplersenstrophy spectrum

The spatial anatomy of urban wildfire vulnerability: a spatially validated GeoAI framework reveals the roles of building density and vegetation moisture in structure loss during the 2025 Palisades Fire

arXiv cs.LG · Parastoo Farajpoor, Mohammadreza Narimani · 2026-08-23

This study presents a spatially validated GeoAI framework to assess urban wildfire vulnerability, focusing on the 2025 Palisades Fire. The method integrates CAL FIRE damage inspections with Sentinel-2 vegetation indices, Landsat surface temperature, and OpenStreetMap data using XGBoost and logistic regression models. Results show building density within 100 m as the strongest predictor of destruction (OR 4.12 per SD), with vegetation moisture (NDMI) at 100-300 m being protective (OR 0.52) and greenness (NDVI) at 30-100 m increasing risk (OR 1.74). Spatial block validation revealed model performance drops from ROC-AUC 0.92 to 0.75, emphasizing neighborhood-scale susceptibility screening.

geoaiwildfire vulnerabilityspatial validationvegetation moisturebuilding density

DAW: Dynamics-Aware Weighting for Deep Learning Forecasts of Chaotic Systems

arXiv cs.LG · Zhou Fang, Gianmarco Mengaldo · 2026-08-23

The paper introduces Dynamics-Aware Weighting (DAW), a data-centric objective reweighting framework for improving deep learning forecasts of chaotic dynamical systems. DAW uses the local dimension from dynamical systems theory to prioritize high-dimensional, dynamically rare states where forecast errors are large, unlike existing methods that rely on target-space density. Evaluated on the chaotic Kuramoto-Sivashinsky equation, DAW outperforms uniform training and statistical density weighting, reducing long-term autoregressive error by better handling complex transitions like wave-merging events.

chaotic systemsautoregressive forecastinglocal dimensionobjective reweightingkuramoto-sivashinsky equation

Toward a First-Principles Update Geometry for the Language-Model Head

arXiv cs.LG · Aditya Somasundaram · 2026-08-23

The paper derives a principled update geometry for the language-model head by analyzing the combined head-softmax module under Hilbert's projective distance. The key result shows that the maximum update-induced change scales with the Euclidean diameter of token rows in the weight matrix. Building on Muon's singular-value conditioning, the authors propose optimizing row separation while constraining this diameter, framing it as an approximate-equidistance problem for high-dimensional vocabularies (V ≫ d).

language-model headhilbert distanceupdate geometryeuclidean diameterequidistance problem

Sharp Barron Regularity Results for Coulombic Many-Electron Wave Functions

arXiv cs.LG · Pingbing Ming, Hao Yu · 2026-08-23

The paper establishes sharp Barron regularity bounds for Coulombic many-electron wave functions after removing universal Jastrow factors. Using a factorization approach from Fournais et al., the authors prove that the quotients φ and φ₃ belong to Barron spaces ℬˢ(ℝ³ᴺ) for all s < 2, demonstrating this range is optimal for universal factorizations. They derive exact endpoint growth rates, showing ‖u‖_{ℬ²⁻ᵉ} ≤ M/ε²‖u‖_{ℬ¹} for computable M, and prove the quadratic rate is sharp when |φ₃(0,0)| ≠ 0, as in ground states.

barron regularitycoulombic wave functionsjastrow factorsuniversal factorizationendpoint growth

MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models

arXiv cs.LG · Yize Li, Ningyuan Yang, Sile Yin, Sindhuja Thogarrati · 2026-08-23

The paper introduces MRMAD, a Multi-Round Multi-Audio Degradation benchmark for evaluating acoustic degradation perception in large audio-language models (LALMs). MRMAD assesses degradation understanding through multi-turn dialogues across speech, music, and sound, requiring models to identify degradation types, compare severity, and track corruption changes. Evaluating 18 LALMs reveals that current models excel at coarse content recognition but struggle with degradation diagnosis, comparison, and reasoning, highlighting a critical gap in low-level acoustic understanding.

audio-language modelsacoustic degradationmulti-turn dialoguebenchmark evaluationlow-level reasoning

A Query-Time Framework for Transient 2D Pore-Scale Flow Prediction and Generative Design

arXiv cs.LG · Yiming Wang, Jiale Zhu, Zhichen Ye, Yandong Lv · 2026-08-23

The study introduces CT-PoreFlow, a continuous-time surrogate model for transient pore-scale flow prediction, addressing costly LBM simulations in porous media design. The method combines topology-aware geometry encoding, compressed spectral mixing, and log-time conditioning with a flux-calibration objective, evaluated on the QSGS-Transient-7606 benchmark (7,606 2D porous structures). Results show a velocity relative L2 of 0.2248 and 12.81% terminal permeability error on unseen geometries, with 98.11% through-connectivity in inverse GAN-based design. The framework integrates flow prediction, transport-aware screening, and LBM-verified inverse design.

pore-scale flowlattice boltzmann methodsurrogate modelgenerative designporous media

When Test-Time Adaptation Helps, Harms, or Becomes Inactive: A Condition-Level Study on CIFAR-10-C

arXiv cs.LG · Sreeja Guha Majumdar, Aratrika Saha · 2026-08-23

The study analyzes when test-time adaptation (TTA) improves, harms, or remains inactive across corruption conditions in CIFAR-10-C. It compares three TTA strategies—BatchNorm-statistics adaptation (BN-Adapt), entropy-minimization (TENT), and reliability-filtered adaptation (EATA)—against an unadapted source model on 15 corruption types at 5 severity levels. All methods improve mean accuracy by 12.2–13.3 percentage points (p < 10^-12) but underperform the source model in 8.0–9.3% of conditions, particularly low-severity corruptions like brightness and fog. EATA closely tracks BN-Adapt (mean difference 0.09 pp), suggesting reliability filtering limits adaptation effectiveness.

test-time adaptationdistribution shiftbatch normalizationentropy minimizationreliability filtering

Counterfactual Evaluation of Temporal Observation Protocols

arXiv cs.LG · Xizhe Zhang · 2026-08-23

The paper introduces counterfactual protocol evaluation, assessing whether benchmark data can determine the predictive value of alternative observation protocols not deployed. The authors develop a value-specific identification theory, showing latent covariance ambiguity may prevent protocol value determination even with infinite data. For linear targets, invisible covariance directions certify non-identification, while targeted measurements restore identifiability. Finite calibration data enables uniform error bounds for protocol-selection regret. Empirical results on Sleep-EDF and Long-Term AF datasets demonstrate reliable distinction of broad temporal-layout differences over fine placements.

counterfactual evaluationobservation protocolslatent covarianceidentification theoryprotocol-selection regret

Joint Causal Structure and Cluster Discovery Using Variational Inference

arXiv cs.LG · Avni Rajpal, Anubhav Kumar, Rishabh Karnad, Mohammad Emtiyaz Khan · 2026-08-23

The paper introduces a variational inference method for joint discovery of latent variable clusters and causal structures, addressing scenarios where neither is known a priori. The approach employs categorical and Bernoulli variational distributions to approximate posteriors over clusters and graph structures, respectively, with derived lower bounds for parameter estimation. Experiments on synthetic and real datasets demonstrate effectiveness in simultaneous cluster and causal discovery.

variational inferencecausal discoverylatent clustersbernoulli modelgraph-structure learning

Efficient Regression Models for Scan Statistics

arXiv cs.LG · Gazi Abdur Rakib, Tristan Ashton, Ryan A. Loomis, Brian S. Mason · 2026-08-23

The paper introduces efficient regression models for scan statistics on real-valued signals, enabling improved fitting of non-stationary signals to detect interval anomalies. The proposed methods, including Nadaraya-Watson kernel regression, represent generalized likelihood ratio statistics and reduce computational complexity from $O(n^4)$ to linear time under maximum interval width constraints. Experiments demonstrate effectiveness in detecting synthetic anomalies and identifying a real-world "platforming" issue in interferometric astronomy.

scan statisticsnon-stationary signalsnadaraya-watson regressiongeneralized likelihood ratiointerferometric astronomy

On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models

arXiv cs.LG · Yang Yu · 2026-08-23

The paper analyzes the theoretical separation between world-model policy learning and imitated world-action models, demonstrating that both direct behavior cloning and world-action imitation recover the same observational behavior policy under ideal conditions. Through formal analysis, it shows that action-conditioned world-model learning differs by predicting outcomes under specified actions via a control objective. The authors characterize irreducible prediction errors in non-action-conditioned models and construct an environment family where observational learners incur worst-case regret while intervention-based learners achieve zero regret. Key findings highlight the distinction between predicting observed futures and action consequences for policy optimization.

world-modelbehavior cloningpolicy optimizationobservational learninginterventional forward model

Unveiling the Depth-Performance Dilemma in Split-Federated Fine-tuning of LLMs

arXiv cs.LG · Hariharan Ramesh, Someshwaran Murugaiyan, Jyotikrishna Dass · 2026-08-23

The study identifies a Depth-Performance Dilemma in Split Federated Fine-tuning (SFF) of Large Language Models (LLMs), where deeper model partitions maximize system throughput and privacy but cause catastrophic performance degradation. Through experiments with models from GPT-2 to Llama-3-8B and multiple benchmarks, the authors show that existing federated adapter aggregation methods (AVG, STACK, SVD, FREEZE) fail to address this issue. Mechanistic analysis reveals that Transformer's near-isometric topology propagates aggregation noise unchecked, leading to Attention Collapse in server partitions.

split federated fine-tuningdepth-performance dilemmaattention collapseadapter aggregationisometric topology

VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR

arXiv cs.LG · Yani Guan, Dengpan Dong, Shuang Luo, Zi Wei · 2026-08-23

The study introduces VERDICT, a method for Optical Chemical Structure Recognition (OCSR) that outperforms pixel-space verification by leveraging agreement among multiple recognizers. Evaluated on 263 ACS journal depictions, agreement among four architecturally distinct recognizers achieved an AUROC of 0.916 (95% CI [0.880,0.952]), significantly surpassing pixel-space re-rendering (AUROC 0.547). The two-of-four rule achieved 81.7% acceptance at 88.8% precision, while the three-of-four rule reached 52.1% at 98.5%. Applied to PMC Open Access, VERDICT produced 6,146 structure labels with 99.5% precision in the three-of-four tier. The method enables reliable multimodal molecular database construction.

optical chemical structure recognitionsmilesaurocmultimodal databasesagreement-based verification

Token-Level Likelihood-Array Regression for Membership Inference and AI-Generated Text Detection

arXiv cs.LG · Jiajun Sun, Zhanrui Cai · 2026-08-23

The paper introduces likelihood-array regression (LAR), a method for membership inference and AI-generated text detection that leverages token-level probabilities across nested left-context windows. LAR organizes likelihood-derived features into structured arrays, aligns them for variable-length texts, and learns detection patterns across context scales, positions, and features. LAR-1 uses individual cell contributions, while LAR-2 incorporates second-order features from paired evaluations. Theoretical analysis establishes minimax bounds and error characterizations. Experiments show LAR outperforms likelihood-based baselines, with shorter-context likelihoods and second-order features providing additional detection gains.

membership inferenceai-generated text detectionlikelihood-array regressiontoken-level probabilitiesminimax bounds

Role-Specialized Mixture-of-Agents with Open-Weight LLMs for Clinical Prediction

arXiv cs.LG · Jun Hou, Yi Fang, Xuan Wang · 2026-08-23

The paper introduces a role-specialized Mixture-of-Agents (MoA) framework for clinical prediction tasks using open-weight Large Language Models (LLMs), addressing privacy and compliance constraints by enabling local deployment. The method combines medical knowledge retrieval with contrastive similar-patient reasoning, isolating the main predictive effect to the final integrator role. Results show that pairing large open-weight analysts with a small open-weight integrator achieves comparable F1 scores to closed-model prompting for mortality prediction, while significantly increasing true high-risk patient detection. Task-dependent performance gains are observed, with smaller improvements for readmission due to weaker record correlations. Role design emerges as a critical factor in training-free clinical LLM prediction.

mixture-of-agentsclinical predictionopen-weight llmscontrastive reasoningrole-specialized

MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning

arXiv cs.LG · Ziyang Luo, Yan Yang, Xiangru Jian, Ziji Shi · 2026-08-23

MCP-Universe RL (MCP-U RL) introduces an open-source framework for training tool-use agents via reinforcement learning, addressing environment orchestration and rollout scheduling challenges. The framework leverages the Model Context Protocol (MCP) to integrate tools without RL-specific code, featuring an environment-orchestration layer for containerized isolation and a rollout-orchestration layer to optimize GPU utilization during multi-turn episodes. It supports backend-agnostic training with veRL and slime integrations. Experiments on GPT-OSS-20B demonstrate improved task reward across software-engineering, deep-research, and general tool-use domains.

reinforcement learningmodel context protocolenvironment orchestrationgpu utilizationtool-use agents

Why Does Robustness Reduce Superposition?

arXiv cs.LG · Adam Elimadi · 2026-08-23

This work provides an empirical explanation for why adversarial training reduces superposition in neural networks, addressing a gap in Gorton & Lewis (2025). Building on Ilyas et al.'s feature taxonomy, the authors demonstrate that adversarial training causes models to discard non-robust features, thereby reducing the total number of features requiring representation and consequently decreasing superposition. The results suggest a causal chain linking robustness, feature selection, and representational efficiency in neural networks.

adversarial trainingsuperpositionnon-robust featuresmechanistic interpretabilityfeature taxonomy

More accurate behavioral predictions with hybrid Bayesian-connectionist models

arXiv cs.LG · Brenden M. Lake, Akshay K. Jagadish, Guangyuan Jiang · 2026-08-23

The authors propose Bayesian distillation with Behavioral Tuning (BBT), a hybrid approach combining Bayesian and neural network models to improve behavioral predictions. BBT involves training a neural network to mimic a Bayesian model using synthetic data, followed by fine-tuning on human behavioral data to capture nuanced patterns. Evaluated across four human concept learning case studies, BBT outperforms traditional methods in predicting human behavior while maintaining Bayesian priors and identifying heuristics and biases that violate standard assumptions. This approach bridges the representational flexibility of Bayesian models with the expressive power of neural networks.

bayesian distillationbehavioral tuningneural networksconcept learninginductive biases

Loss Landscape Features That Make Adam Stall: Definitions, Estimators, and the Preconditioned Hessian View

arXiv cs.LG · Rodion Podorozhny · 2026-08-23

This work identifies metrics that predict Adam's convergence behavior on ill-conditioned loss landscapes, distinguishing between successful optimization and stalling. The authors derive the condition number of both the Hessian and Adam-preconditioned Hessian, alongside diagonal mass ρ, negative spectral mass, and gradient energy fractions to characterize landscape features. Through analytic examples and a FINER image fitting case study, they demonstrate how Adam's diagonal preconditioning mitigates axis-aligned but not cross-coupled ill-conditioning, achieving 120-134 dB PSNR with second-order methods versus Adam's plateaus.

adam optimizerloss landscapepreconditioned hessianill-conditioningimplicit neural representation

Learning Reduced-Order Dynamics with Singularity via Latent-Augmented Neural Ordinary Differential Equations

arXiv cs.LG · Xiaorui Wang, Yu Zhou, Wenjie Mei, Dongzhe Zheng · 2026-08-22

The paper introduces Latent-Augmented Neural Ordinary Differential Equations (LA-NODEs), a framework enhancing neural ODEs to model self-intersecting trajectories in reduced-order systems by augmenting latent space dimensions. The method theoretically derives the minimum required augmentation dimension and resolves conflicting vector fields, improving expressiveness. Evaluated on an interior permanent magnet synchronous motor drive and a distributed energy system, LA-NODEs outperform conventional approaches in prediction accuracy and modeling fidelity, enabling high-precision data-driven modeling of complex industrial systems.

neural ordinary differential equationsreduced-order modelinglatent augmentationvector fieldsindustrial systems

Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion

arXiv cs.LG · Jiaqian Zhu, Yang Zhang, Junhua Ding, Xiaowei Yu · 2026-08-22

This work investigates the robustness of large language models (LLMs) to lexical perturbations, revealing a critical vulnerability in multi-step reasoning tasks. The authors evaluate four instruction-tuned and frontier LLMs across four reasoning benchmarks under three perturbation types: keyboard noise, character swaps, and filler insertion. Results show character-level perturbations significantly degrade accuracy via attention diversion, where fragmented subword tokens attract disproportionate attention mass in middle and final transformer layers. Factorial interventions demonstrate that token content and attention allocation are coupled, explaining why inference-time repair strategies fail to recover performance consistently. Code and data are publicly available.

attention diversionlexical perturbationsubword tokenizationtransformer layersmulti-step reasoning

Who Should Teach? Confidence-Aware Dual-Teacher Learning for Few-Shot Node Classification on Text-Attributed Graphs

arXiv cs.LG · Hojin Kim, Sujin Yoon, Sungsu Lim, Dongwon Lee · 2026-08-22

The paper proposes CoTeach, a confidence-aware dual-teacher framework for few-shot node classification on text-attributed graphs (TAGs). It dynamically selects between a GNN and an LLM as the supervision source per node, leveraging their complementary strengths in structural and semantic information processing. Experiments show CoTeach improves classification accuracy while reducing LLM usage costs by 30% compared to uniform LLM utilization approaches.

text-attributed graphsfew-shot learninggraph neural networkslarge language modelsconfidence estimation

Decoupled Physical Modeling and Execution for Physics Reasoning

arXiv cs.LG · Ye Zhang, Xuehang Guo, Rui Pan, Pengfei Yu · 2026-08-22

The paper proposes a framework for physics reasoning that decouples physical modeling from mathematical execution, mimicking human problem-solving. The method employs a two-stage post-training strategy: supervised fine-tuning for structured modeling, followed by reinforcement learning with rubric-based feedback to refine modeling quality. Evaluations on PhysReason, PhyX, and SeePhys benchmarks demonstrate consistent performance gains, with physical modeling outputs improving by ~3% on average, particularly benefiting smaller LLMs.

physics reasoningintermediate representationssupervised fine-tuningreinforcement learningmultimodal benchmarks

The Price of Decentralization in Top-$K$ Arm Identification

arXiv cs.LG · Larissa Xu, Jasmine Nguyen, William Chang · 2026-08-22

The paper analyzes decentralized top-$K$ joint-arm identification in multi-agent multi-armed bandits under three observability regimes: (A) shared rewards with hidden actions, (B) observed actions with private rewards, and (C) full asymmetry. It proposes communication-free UCB-Intervals algorithms that leverage regime-specific signals—shared ordering, observable deviations, and enlarged confidence radii—to enable implicit coordination. The key theoretical result quantifies the statistical cost of decentralization: shared-reward identification is near-optimal (within a log factor), while full asymmetry incurs a fixed $4\times$ sample complexity penalty ($\rho^2$ multiplicative). Stopping time scales as $O(\sum_{\mathbf{a}} \log(A^M/\delta)/\Delta_{\mathbf{a}}^2)$, with $A^M$ dependence proven unavoidable.

multi-armed banditstop-k identificationdecentralized learningsample complexitychange-of-measure

TANGO: Token-Aggregated Nonlinear Gating Operators for Natural and Formal Language Modeling

arXiv cs.LG · Joshua Nunley · 2026-08-22

The paper introduces TANGO (Token-Aggregated Nonlinear Gating Operators), a Transformer variant replacing self-attention and feed-forward layers with cross-token gated residual updates. Each token generates a SwiGLU gate vector, aggregated via query-key similarities to rescale destination features. TANGO exhibits quadratic sequence-length complexity, while its linear variant WANGO (Windowed Aggregation of Nonlinear Gating Operators) uses windowed attention and feature-map statistics. Evaluated against Recurrent/Untied Transformer++, GAU, and FLASH (all ~44.3M nonembedding params), TANGO achieves the lowest validation NLL on FineWeb-Edu, Lean, and DeepMind Mathematics, despite higher FLOPs. WANGO outperforms linear-complexity baselines at comparable MACs.

transformergatingswiglusequence-modelinglinear-complexity

CST: Collaborative Selective Transmission for Communication-Efficient Multimodal Edge Inference

arXiv cs.LG · Hai Chi, Junrui Zhang, Rui Ning, Chonggang Wang · 2026-08-22

The paper introduces Collaborative Selective Transmission (CST), a communication-efficient framework for multimodal edge inference that selectively retrieves complementary helper features to avoid redundancy. CST employs main-directed query–response with sample-adaptive sparse retrieval supports, guided by Partial Information Decomposition and the Multiview Redundancy Assumption, to minimize transmission of task-relevant but duplicated features. Evaluated on three real-world benchmarks, CST transmits ≤14.18% of helper features while maintaining competitive accuracy, achieving up to 4.27× end-to-end speedup over baseline methods on a Jetson Orin Nano testbed.

multimodal edge inferencepartial information decompositionsparse retrieval supportscommunication-efficient inferencemultiview redundancy assumption

Symbolic Neural ODEs: Learning interpretable models from time-series data

arXiv cs.LG · Nibodh Boddupalli, Jeff Moehlis · 2026-08-22

The authors introduce Symbolic Neural ODEs, a framework for learning sparse, interpretable models of dynamical systems from time-series data. The method parameterizes the vector field via a neural architecture, trained by minimizing a multi-step prediction loss with progressive horizon extension, ensuring consistency under repeated composition of dynamics. Sparsity-promoting regularization yields parsimonious models with improved stability and generalization. Experiments demonstrate accurate recovery of diverse behaviors, including fixed points, periodic orbits, and chaotic attractors, with strong agreement in short-term dynamics and long-time statistical properties. Theoretical bounds link trajectory error to statistical accuracy, providing principled explanations for the observed behavior.

symbolic neural odesdynamical systemssparsity-promoting regularizationchaotic attractorslyapunov exponents

What actually runs: a measurement study of language model placement and decode speed on the Apple Neural Engine

arXiv cs.LG · Shahir M A · 2026-08-22

This study empirically investigates language model placement and decode speed on the Apple Neural Engine (ANE) through three measurement approaches: sweeping LLM primitive variations, training matched models across sizes and precisions, and monitoring ANE memory-controller byte counters. Results demonstrate that computation expression, not content, determines ANE eligibility, with fused RMSNorm achieving full eligibility while its decomposition remains CPU-only. Weight encoding significantly impacts accelerator utilization, with int8 and 2-bit models achieving ~83% residency and 1.8-2.2x speedup over fp16 counterparts. Decode cost correlates with bytes streamed per token at ~0.77 of nominal encoding width. Ternary models emerge as smallest and fastest, with 25M parameter models achieving 10.5MB size and 0.63ms/token speed.

apple neural enginedecode speedweight encodingternary modelsmemory-controller

Development and Feasibility Evaluation of an Edge AI as Medical Device System for Breast Cancer Multidisciplinary Team Meetings

arXiv cs.LG · Aarzoo Dhiman, Farzana Haque, Kartikae Grover, Lydia Brian Smith · 2026-08-22

The study presents an edge AI system for Breast Cancer Multidisciplinary Team (MDT) meetings, combining on-device Automatic Speech Recognition (Whisper large-v3) and a retrieval-augmented generation (RAG) model (MedGemma-RAG) to transcribe discussions, structure clinical data, and generate NICE-guideline-aligned treatment recommendations. Optimized for an NVIDIA Jetson AGX Orin, the pipeline achieves a 20.7-24.4% reduction in word error rate (WER) versus baseline and matches commercial ASR performance within 0.58% WER. MedGemma-RAG identified 2.3× more MDT-concordant interventions than a cloud-based comparator (p=0.020) with comparable accuracy. Stakeholders endorsed documentation automation and decision support but noted workflow integration and trust barriers.

automatic speech recognitionretrieval-augmented generationword error rateedge aimultidisciplinary team

Beyond Fresh Starts: Stateful Inference for Streaming ASR in Conversational Voice Agents

arXiv cs.LG · Sameep Chattopadhyay, Alexander Erdmann, Mari Ostendorf · 2026-08-22

The study introduces two state-management strategies for streaming automatic speech recognition (ASR) in conversational voice agents to address performance degradation caused by limited memory constraints and conversational phenomena like long silences and backchannels. These strategies preserve cross-utterance context, mitigating onset errors typically exacerbated by state resets at each turn. Evaluated on two state-of-the-art streaming models across two spoken dialogue benchmarks, the best method achieves a 15-21% relative reduction in word error rate (WER) at utterance onsets.

streaming asrstate-managementcross-utterance contextword error ratespoken dialogue

Pretreatment DCE-MRI Resolves Response Quality Within Pathologic Endpoints in Neoadjuvant Breast Cancer

arXiv cs.LG · Dattatreya Kantha, Murray H. Loew · 2026-08-22

The study demonstrates that pretreatment dynamic contrast-enhanced MRI entropy (intratumoral enhancement heterogeneity) stratifies recurrence risk within pathologic complete response (pCR) and residual cancer burden (RCB) endpoints in neoadjuvant breast cancer. Across four cohorts (n=1,200), a prespecified entropy threshold identified favorable (lower recurrence) and adverse (higher recurrence) structural states, yielding a four-tier framework with 4.1- to 7.7-fold recurrence differences. Adverse structure captured 25.1% of pCR patients in I-SPY2 (n=219) and was associated with higher recurrence risk (HR=2.87-8.13 across subsets). RNA analysis linked favorable structure to immune-architecture programs but weakly discriminated structural states, highlighting MRI's unique role in revealing compressed response-quality differences.

entropyneoadjuvantrecurrenceheterogeneitystratification

Counterfactual Quotient Models: Learning What Actions Change, Not What the World Does

arXiv cs.LG · Junlin Chen, Ruijie Wang, Jianxin Li · 2026-08-22

The paper introduces Counterfactual Quotient Models (CQM), a reinforcement-learning framework that learns action-dependent effects by canceling shared stochastic dynamics across counterfactual rollouts. CQM focuses on pairwise action comparisons rather than predicting complete future states, removing action-independent variation via a canonical centered representation. Theoretical analysis establishes properties including decision sufficiency, identifiability, and regret bounds. Experiments in physics-based environments demonstrate CQM's ability to suppress irrelevant variation, generalize to unseen rewards, and improve action ranking compared to absolute-future prediction models.

reinforcement learningcounterfactual reasoningaction-dependent effectsstochastic dynamicsrepresentation learning

Semantic Reasoning Denoising: Correcting Language Model Reasoning with Semantic Operators

arXiv cs.LG · Yujiao Yang · 2026-08-22

The paper introduces Semantic Reasoning Denoising (SRD), a method for correcting reasoning errors in large language models through operatorized Markov denoising. SRD represents semantic noise via executable error operators that encode error type, location, and propositional corrections, enabling iterative refinement of reasoning trajectories. During training, the model learns to identify and invert semantic noise; during inference, it applies noise-level-aware denoising with localized updates. Evaluated on six in-domain benchmarks (mathematics, code, knowledge, commonsense), SRD improves baseline performance by 3.2 points on average and matches Llama-3-8B-Instruct in cross-dataset transfer tasks.

semantic reasoning denoisingerror operatorsmarkov denoisingreasoning trajectoriesnoise-level-aware

📰 Industry Media (5)

I spent a day at a robot “carnival” in Shanghai. Here’s what I saw.

MIT Tech Review — AI · You Xiaoying · 2026-08-25

China's strategic push for embodied AI is demonstrated through a Shanghai robotics carnival showcasing over 100 companies specializing in humanoid and task-specific robots. The event featured public demonstrations including robotic dogs performing stunts, quadrupedal robots navigating stairs, and humanoids executing martial arts and acrobatics. Despite global skepticism about humanoid forms due to safety and battery limitations, Chinese firms dominate 90% of the 13,000 global humanoid shipments, using such events to increase public engagement and acceptance.

embodied aihumanoid robotsquadrupedal robotsintelligent manufacturingrobotic arms

Perplexity Ships Portable Computer on NVIDIA DGX Spark: Local Harness, OS-Enforced Sandbox, and Zero Per-Token Cost for Local Steps

MarkTechPost · Asif Razzaq · 2026-08-25

Perplexity introduces Portable Computer, a local-first agentic platform running on NVIDIA DGX Spark, enabling zero per-token costs for local inference. The system integrates a local model (Qwen 3.8 27B or PPLX 27B), inference engine, OS-enforced sandbox, and app connectors, with optional escalation to 15+ cloud models requiring explicit user approval. Benchmarks on the Local Knowledge Work Bench show PPLX 27B achieving 85.4% accuracy, outperforming Pi (77.6%) and Hermes (74.0%). Terminal Bench 2.1 results indicate 59.6% accuracy locally, rising to 73.0% with cloud escalation at $0.415 per rollout. Deployment requires GB10-class hardware or RTX GPUs with ≥24GB VRAM.

agentic platformos-enforced sandboxper-token costqwen 3.8 27bterminal bench

Meta AI Introduces MetaRoCE: A Clean-Sheet RDMA Transport Built for AI-Scale Ethernet

MarkTechPost · Asif Razzaq · 2026-08-25

Meta AI introduces MetaRoCE, a clean-sheet RDMA transport protocol optimized for AI workloads on commodity Ethernet. Departing from standard RoCE, MetaRoCE treats the network as lossy, relocating ordering, path selection, and recovery to the NIC. Key innovations include out-of-order delivery, native multipathing, and congestion control from both ends. Evaluated on a 64-node AMD GPU cluster running RCCL collectives, MetaRoCE maintains ~86% throughput at 1% packet loss and scales linearly with plane count. The protocol requires only ECN and ECMP from switches, making it deployable on uncontrolled fabrics. Meta plans to release the specification, compliance suite, and a reference implementation via OCP in October 2026.

rdmaethernetcongestion controlall-reducenic

Fastino Releases GLiNER2.5: A Boundary-Prediction Architecture That Removes Span Enumeration From Information Extraction

MarkTechPost · Michal Sutter · 2026-08-24

Fastino introduces GLiNER2.5, a boundary-prediction architecture for information extraction that eliminates span enumeration. The model predicts entity start/end boundaries and inside scores, enabling linear computation in sequence length for fixed schemas, 4,096-word contexts, and joint entity-relation decoding. Three checkpoints (74M, 194M, 287M parameters) achieve 56.17 macro F1 on 16 zero-shot benchmarks, with a 24.75-point gain on XNLI. The architecture supports unlimited span lengths, cross-task label constraints, and span attributes.

boundary predictionspan enumerationzero-shot benchmarksjoint decodinglinear computation

MIT AI forecasts extreme weather without historical data

AI News · Ryan Daws · 2026-08-25

MIT researchers propose η-learning, an AI method for generating statistically-plausible extreme weather scenarios without relying on historical disaster data. The approach combines point statistics (event frequency-intensity relationships) with spatial pattern learning from limited-resolution training data to synthesize unprecedented events. When tested on US precipitation data, the model generated plausible 100-year rainfall events (e.g., 300mm storms in NYC) beyond recorded maxima. Output includes spatial intensity distributions, duration estimates, and affected areas for infrastructure resilience testing. Limitations include requiring hazard-specific input statistics and spatial data.

η-learningextreme event modelingspatial statisticsprecipitation forecastinginfrastructure resilience


Generated automatically at 2026-08-25 19:59 UTC. Summaries and keywords are produced by an LLM and may contain inaccuracies — always consult the original article.