Daily Digest — 2026-08-05
371 items · 6 research labs, 355 arxiv papers, 10 industry media
AI News: all feed URLs failed (last tried: https://artificialintelligence-news.com/feed/)
🏛️ Research Labs (6)
New ways to learn and teach with ChatGPT Work and Codex
OpenAI introduces three education-specific plugins for ChatGPT Work and Codex, targeting K–12 educators, college educators, and college students. These plugins integrate with institutional tools (e.g., LMS, calendars) to provide context-aware assistance for pedagogical tasks, leveraging agentic capabilities while maintaining educator control. Deployed via ChatGPT Edu and ChatGPT for Teachers, they aim to reduce prompt engineering overhead and bridge the 'capability overhang' observed in student AI usage. Early adopters include Houston ISD and the University of Pennsylvania, with longitudinal studies underway in Estonia. Structured access programs show students progress toward power-user behaviors in analysis and learning tasks.
agentic capabilitiescapability overhangcontext-aware assistanceprompt engineeringlongitudinal studies
Apple is getting this wrong
OpenAI disputes Apple’s legal claims, alleging procedural errors and miscommunications in Apple’s lawsuit. Apple’s outside counsel mistakenly contacted OpenAI’s General Counsel, Che Chang, due to confusion over Asian last names, falsely claiming prior discussions. Apple also accused former employees Chang Liu and Tang Tan of accessing confidential information, but OpenAI contends that Apple’s system access mismanagement caused residual access issues. OpenAI asserts that it neither possesses nor desires Apple’s trade secrets and offered to resolve the matter collaboratively. Evidence includes email exchanges and iMessages demonstrating miscommunication and Apple’s failure to address allegations prior to litigation.
legal claimsmiscommunicationresidual accesstrade secretslitigation
Circles powers telco personalization with OpenAI technology
Circles leverages OpenAI's API to develop an AI-native telco stack, featuring a multi-agent architecture (CareX) for autonomous customer support and a personalization engine (Xplore IQ) for real-time recommendations. The system integrates customer data across usage, billing, and network activity to enable proactive, context-aware interactions. Results include a 22% ARPU increase, 9% churn reduction, and 65% autonomous resolution rate for support queries. Internally, Codex improves development efficiency by 29% via coding assistance and testing automation. Safeguards include PII encryption and scoped agent access.
multi-agent architectureautonomous resolutionreal-time personalizationllm orchestrationtelco saas
Deploy local agents everywhere with LFM2.5-2.6B
LFM2.5-2.6B introduces a compact, efficient agentic model optimized for edge devices, achieving competitive performance against models up to 4x larger on tasks like tool use, instruction following, and multi-step agentic tasks. The model, pre-trained on ~34T tokens with a 128K context window, undergoes supervised fine-tuning, teacher specialization, multi-domain on-policy distillation, and agentic reinforcement learning within real agent harnesses. Benchmarks show LFM2.5-2.6B excels in instruction following and tool use, outperforming larger models like Gemma and Qwen in most tasks. It achieves 220 tokens/s on Apple M5 Max and 113 tokens/s on AMD Ryzen CPUs, enabling deployment on devices as constrained as smartphones.
agentic reinforcement learningmulti-domain distillationtool useinstruction followingedge deployment
The latest AI news we announced in July 2026
Google announced multiple AI advancements in July 2026, focusing on scalable agentic workflows, embodied reasoning, and multimodal applications. Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber optimize token efficiency and latency for production agents, while Gemini Robotics ER 2 enhances physical-world interaction through natural language understanding and multi-task reasoning. Gemini Notebook integrates with Google Search and the Gemini app, featuring secure cloud compute. AlphaEvolve, a code-optimization agent, is now available on Gemini Enterprise Agent Platform, automating algorithm refinement. Lyria 3.5 improves music generation with advancements in musicality and vocal quality. NOAA leverages Google Cloud's H4D VMs for high-performance weather forecasting.
agentic workflowsembodied reasoningtoken efficiencymultimodal integrationcode-optimization
Inside our 353,000-person vibe coding course
Google and Kaggle conducted a large-scale AI education initiative, the 5-Day AI Agents: Intensive Vibe Coding Course, reaching 353,000 participants. The course introduced 'vibe coding', enabling natural language programming for designing, securing, and deploying production-grade AI agents. Participants engaged in real-time collaboration via Kaggle's Discord, with 392,000 active users exchanging code and debugging. The course culminated in 6,000 capstone project submissions, showcasing advanced applications like historical manuscript transcription and space-weather research systems. Course materials remain accessible through Kaggle Learn, supporting ongoing skill development.
vibe codingai agentscapstone projectskaggle discordnatural language programming
📜 arXiv Papers (355)
Bridging Artificial Intelligence and Power Systems Education Using a Hands-On Executable Framework
The paper introduces an open, executable module library for engineering-grounded AI (EGAI) in power systems, addressing barriers identified in a survey where 92% of practitioners reported difficulties in deploying AI models. The framework progresses from foundational DNN templates for load-curve fitting to advanced modules like CNN-based power-flow surrogates (5-bus system), DNN-assisted optimization, DRL for battery control, and PINNs for swing equations. Delivered as Jupyter notebooks (local/Colab-compatible) via IEEE PES webinars, the initiative attracted 590+ live attendees and 344+ repository visits within two weeks, validating demand for domain-specific AI education.
engineering-grounded aipower-flow surrogatephysics-informed neural networksdeep reinforcement learningjupyter notebooks
UEmbed: Unified Sparse and Dense Multimodal Embeddings
UEEmbed introduces a decoder-only multimodal embedding model that unifies sparse lexical and dense semantic representations in a single causal forward pass. The method appends N learnable special tokens to the input, partitioning the vocabulary into N disjoint subsets, where each token's hidden state predicts sparse weights over its assigned subset. Trained on public data, UEmbed scales at 2B, 4B, and 9B parameters, achieving 71.8 (dense) and 71.0 (sparse) on MMEB-v2, outperforming existing multimodal embedding models like RzenEmbed. It remains competitive on BEIR and demonstrates practical utility in effectiveness, efficiency, and agentic applications, extending sparse retrieval to multimodal inputs.
multimodal embeddingssparse retrievaldecoder-onlycausal forward passlearnable tokens
CoWAM: Coordination Contracts for Selective Policy Intervention with WAMs
CoWAM introduces coordination contracts as a selective intervention layer for bimanual robot policies augmented with World Action Models (WAMs). The method combines typed admissibility checks, event-conditioned verification, and calibrated intervention gates to ensure synchronization, role compatibility, and collision convergence. CoWAM preserves nominal actions unless an alternative satisfies all active obligations and offers a low-risk improvement, otherwise invoking a predefined abstention fallback. Evaluated across eight simulated bimanual tasks, CoWAM improves coordination-valid selection by 16.7 percentage points over contract-only variants and increases closed-loop success by 9.6 percentage points over selective baselines, while maintaining harmful interventions below 1%.
coordination contractsworld action modelsadmissibility checksbimanual tasksselective intervention
AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies
AtumAI introduces a principled framework for automating the generation of datacenter control-plane policies, addressing limitations of off-the-shelf agentic AI in formalization, transferability, and systematic exploration. The framework comprises two components: the Datacenter Task Compiler, which translates plain-language goals into formal, machine-checkable specifications, and the Evolutionary Design Discovery Loop, which employs a diffusion model, evolutionary algorithm, and surrogate model to expand the search space. Evaluated on workload placement, resource scaling, and power management tasks, AtumAI consistently outperforms expert-engineered baselines, reducing task onboarding from months to description writing.
datacenter control-planeagentic aidiffusion modelevolutionary algorithmsurrogate model
Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection
PRECOG (Pre-Computed Context Injection) introduces a retrieval mechanism for State-Space Models (SSMs) that reduces prefill latency from O(L_context) to O(1) by pre-encoding document corpora as SSM hidden states and injecting them at query time. The method leverages SSMs' fixed-size, position-agnostic recurrent hidden states to bypass in-context re-ingestion, enabling Structured Memory Consolidation (SMC) for hierarchical persistent memory. Evaluated on TENNs-LLM (1.2B parameters, 192 KB hidden state), PRECOG matches in-context RAG quality while reducing prefill latency from ~27 s to <6 ms (~4500× speedup) on edge hardware, a feat impossible for Transformer KV-caches due to their position-entangled, linearly growing nature.
retrieval-augmented generationstate-space modelskv-cachecontext injectionedge hardware
A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI
The paper presents a taxonomy of cognitive capability gaps in generative and agentic AI, identifying five key dimensions where current systems fall short: persistent state modeling, goal-directed autonomy, self-monitoring and control, environment interaction, and learning and adaptation. Through a structured literature review, the authors analyze recent advances, recurring limitations, and open challenges in each dimension. They propose an Adaptive Cognitive Intelligence Architecture (ACIA) framework and cognition-centric evaluation methods to guide future research toward more reliable long-term reasoning and adaptive decision-making in AI systems.
cognitive aiagentic aiadaptive architecturepersistent state modelingcognition-centric evaluation
Who Should Be Generated? Justifying Demographic Targets in Open-Ended Generation
The paper formalizes the missing-target problem in open-ended generation audits, proposing a framework to justify demographic target distributions through four commitments: evaluative object, prior admissibility, allocation, and operationalization. It distinguishes geographic (public-world) and occupational (incumbency) priors, requiring independent justification for workforce-composition fidelity. Evaluations on AP-Bench show significant divergence (0.508–0.606) from geography-derived targets, with mean absolute JSD₂ changes of 0.279–0.355 when substituting equal-category comparators. The work shifts target construction from a preliminary step to an integral fairness evaluation component.
demographic targetsopen-ended generationfairness evaluationdistribution divergenceprior admissibility
Analytic Planning under Uncertainty with Moment Closure
The paper introduces a method for analytic planning under uncertainty by deriving closed-form Bellman backups that propagate predictive mean and covariance, avoiding the need for restrictive policy or reward structures. By pairing a Gaussian transition model with a radial-basis value function, the approach ensures compatibility between the transition distribution and value function class, enabling analytic moment propagation. Empirical results in continuous control demonstrate reduced target variance and well-calibrated predictive uncertainty under stochastic observations.
bellman backuppredictive uncertaintygaussian transition modelradial-basis functionmoment propagation
Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation
The authors introduce Magnet, a novel detection framework addressing cross-session AI misuse through capability accumulation. Existing detection methods focus on single-session or multi-turn threat models, leaving vulnerabilities to attackers decomposing harmful goals into innocuous, isolated sessions. Magnet models capabilities accrued across agentic conversations, aggregating evidence at a user-level correlator rather than per-session state. This approach efficiently identifies incriminating artifacts scattered across benign sessions, enabling detection of harmful composites. The authors demonstrate that cross-session goal decomposition elicits more harmful capabilities than single-session attacks, highlighting the need for Magnet's robust, session-spanning detection mechanism.
capability accumulationcross-session detectionagentic conversationsgoal decompositionevidence aggregation
Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies
The paper introduces $k$-adaptable policy synthesis for uncertain Markov decision processes (UMDPs), optimizing a set of $k$ policies under minimax regret when model uncertainty resolves before execution. The authors prove NP-hardness and develop KAPS, an exact nested branch-and-bound algorithm that jointly optimizes policy sharing across MDPs and the policies themselves. Experiments on UMDP benchmarks show the largest regret reduction occurs when increasing from one to two policies, with KAPS matching existing methods in single-policy settings while proving optimality more frequently.
uncertain markov decision processesminimax regretpolicy synthesisbranch-and-boundnp-hardness
Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation
The paper proposes that scientific abduction—specifically identity abduction, the inference of structural equivalence across domains—can occur without continuous sensorimotor embodiment via representational grounding. It introduces the Abduction Loop architecture, comprising representation generation, motif extraction, convention-space canonicalization, cross-domain retrieval, identity-hypothesis generation, and adversarial verification, with abstention as default. A case study demonstrates a multimodal model identifying equivalence between a gravitational-memory transport model and the Kaiser-Squires weak-lensing mass-mapping complex. The work contributes a mechanistic framework, architecture, and the DAB-30 benchmark for evaluation.
identity abductionrepresentational groundingconvention spaceabduction loopmultimodal model
CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization
We introduce Chunked Momentum Orthogonalization (CMuon), a novel optimization strategy that accelerates and stabilizes Diffusion Transformer (DiT) training by addressing implicit subspace coupling in fused weight tensors. CMuon partitions functionally distinct weights (e.g., in AdaLN and QKV layers) into independent sub-components prior to orthogonalization, overcoming the late-stage convergence limitations of Momentum Orthogonalization (Muon). Experiments demonstrate that CMuon enables a 675M-parameter DiT to achieve a FID of 1.18 on ImageNet 256 in 200 epochs, representing a >2x training speedup over AdamW while resolving Muon's convergence plateaus.
diffusion transformersmomentum orthogonalizationadalnqkv layersimplicit subspace coupling
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
SWE-Touch introduces a framework for benchmarking coding agents in shared workspaces where users modify task-relevant code, addressing a gap in repository-level evaluations. The framework employs Counter-Edits: plausible code changes conflicting with task completion, generated via a User Patch Generator and injected with contextual messages. Evaluations on SWE-bench Verified, SWE-Bench Pro, and DeepSWE show a 7.7 percentage point drop in resolve rate, with failures linked to limited workspace awareness and insufficient code re-inspection. Results highlight the need for improved state awareness, adaptive behavior, and conflict resolution in collaborative coding environments.
counter-editsuser patch generatorresolve ratestate awarenessshared workspace
DyFrDet: Towards Accurate Small Object Detection via Dynamic Frequency Suppression with Label Disambiguation
The paper introduces DyFrDet, a novel small object detection (SOD) framework addressing frequency domain noise and label ambiguity. The method features a Dynamic Frequency-aware Feature Pyramid Network (DyFrFPN) that adaptively suppresses low-frequency redundancy and high-frequency noise via a Dynamic Band Predictor (DBP), and a Label Disambiguation Module (LDM) that models label ambiguity probabilistically. Experiments show state-of-the-art performance across multiple benchmarks, demonstrating effectiveness in precise small object localization.
small object detectionfrequency domain suppressiondynamic band predictorlabel disambiguationfeature pyramid network
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
The article proposes a longitudinal approach to evaluating human-AI interactions, emphasizing the need to shift from static NLP benchmarks to measuring long-term behavioral changes in users. Drawing on social science methodologies, it argues that combining computational NLP techniques with longitudinal measurements can identify emergent cognitive and socio-affective risks from prolonged LM use. The framework aims to enable real-time detection of problematic user behaviors and inform alignment strategies to mitigate longitudinal harms, addressing gaps in current post-hoc safety evaluations.
longitudinal analysisbehavioral metricsalignment frameworkssocio-affective risksonline detection
Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery
DiffeoAfford introduces an action-grounded tissue affordance framework that automatically generates visual attention supervision for laparoscopic surgery by combining diffeomorphism-constrained tissue tracking with instrument trajectory analysis. This eliminates manual per-frame annotation and enables AffordView, an assistive auto-framing system that anticipates relevant surgical regions. Evaluations show alignment with expert annotations and surgeon gaze, while reducing cognitive workload through subjective, physiological, and behavioral measures.
tissue affordancediffeomorphism-constrained trackingauto-framinglaparoscopic surgerycognitive workload
Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment
The paper introduces TinyDamage, a hybrid architecture combining a vision-language model (VLM) with dedicated segmentation for fine-grained vehicle damage assessment. The method delegates spatial grounding to a multi-task segmentation model using a supervised contrastive loss (outperforming focal loss for tiny-object detection) and integrates it into a 7-node LangGraph agent pipeline. Results show reduced hallucination rates from 92% (text-only) to 31% in human-verified reports, with Qwen-VL achieving 87.3% semantic classification accuracy but poor spatial grounding. The work also proposes DET_l, a per-category detection metric for tiny objects under class imbalance.
vision-language modelssegmentationtiny-object detectioncontrastive lossvehicle damage assessment
Real-Time Detection and Repair of LLM Agent Failures
The paper introduces a lightweight system for real-time detection and repair of LLM agent failures using step telemetry and deterministic verification. It employs a one-class echo-state-network ensemble with CUSUM alarms trained on healthy runs, achieving 0.71 failure detection at 5% false-alarm rate (AUROC 0.872). The method includes deterministic verification for error-free checks, recovering 45% of failures and improving task success from 52% to 73% with minimal overhead (~200μs/step). Results show strong transferability across models (llama3.1 8b, qwen2.5 7b/3b) and benchmarks (AFTraj-2K, ATBench).
llm agentsecho-state-networkcusum alarmsdeterministic verificationstep telemetry
Syntax Meets Semantics: Understanding Scientific Formulae
The study investigates the cross-modal correspondence between syntactic and semantic representations of scientific formulae, revealing weak observable alignment despite strong latent correlation. Using graph-based encoders for syntax and text-based encoders for semantics, the authors apply contrastive learning to align these modalities in a shared representation space. Results demonstrate that learned alignment significantly improves cross-modal retrieval performance, indicating explicit representation learning can bridge the gap between native syntactic and semantic spaces.
scientific formulaecross-modal retrievalcontrastive learninggraph-based encodersrepresentation alignment
Infinite Trace Objectives with Finite Trace Techniques: Translating LTL to LTLf+
The paper introduces a translation from Linear Temporal Logic (LTL) to LTLf+, enabling the use of finite-trace techniques for infinite-trace objectives without loss of expressive power. The method first normalizes LTL formulas into the reactivity fragment of the Manna-Pnueli hierarchy, then applies linear translations for each fragment component. This approach leverages efficient finite-word automata procedures, maintaining a doubly exponential complexity bound from LTL to automaton via LTLf+.
linear temporal logicltlf+reactivity fragmentfinite automatadoubly exponential complexity
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision
The paper introduces ParEvalLayer, a decision layer for partial LLM-agent evaluations that determines whether a tested agent system is better by a predefined margin, not better, requires more evidence, or should abstain, based on paired outcomes and a pre-selected comparison policy. The method is validated by replaying completed benchmark data as if evaluations had stopped earlier, checking if partial judgments match full evaluations. Results show that some benchmarks reach consistent decisions with only 15-25% of task outcomes, while others require more, highlighting the need to report decision rules and unresolved comparisons alongside partial scores.
llm-agentpartial evaluationdecision layerbenchmark replaycomparison policy
Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks
The study identifies Solution Hacking, a failure mode where large language models (LLMs) achieve correct answers on scientific reasoning benchmarks through invalid shortcuts rather than valid derivations. Analyzing difficulty levels, domains, and frontier models, the authors found hacking rates ranging from 2.2% on common problems to 37.4% on high-level exams, with 8.2%-44.1% of credited answers being hacked. Expert-inspired anti-hacking strategies, including an automatic judge and test-time instruction, reduced reported accuracy while minimally affecting correct solutions, revealing that answer-only evaluation overestimates LLMs' scientific reasoning capabilities.
solution hackingscientific reasoning benchmarkslarge language modelsanswer-only evaluationanti-hacking strategies
Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce
The paper introduces Agentic Commerce World (ACWorld), an auditable environment for evaluating AI agents in vibe commerce scenarios where natural language goals are delegated to Buyer and Merchant agents. ACWorld implements the Vibe Commerce Protocol (VCP) to validate agent actions, maintain transaction state, and record interactions for reproducibility. Evaluations on a 260-task benchmark (200 capability-coverage tasks and 60 large-catalog tasks with 785,022 listings) show model performance ranging from 56.1% to 91.4%, demonstrating the necessity of process-level evidence over final-state evaluation.
vibe commerceagentic commerce worldvibe commerce protocolauditable environmenttransaction state
xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
The paper introduces xPress, a parallel refinement method for diffusion drafters in speculative decoding that addresses the causality gap in block-diffusion approaches like dFlash. By reconciling the entire draft block through parallel refinement, xPress restores token dependencies without sequential processing. Evaluated on Qwen3-8B across seven benchmarks, xPress improves average acceptance length by 30% (up to 56%) and boosts decoding throughput by 1.3× (up to 1.7×) compared to baseline dFlash.
speculative decodingdiffusion draftersparallel refinementcausal dependenciesacceptance length
Agentic Incident Response through Digital Twin-Enhanced Multiscale Planning
We propose an agentic incident response system combining decision-theoretic planning with LLM-generated commands, addressing limitations of abstract models and LLM hallucination. The method employs a rollout planner for tactical resource allocation and a lightweight LLM agent for operational command generation, supported by a digital twin for simulation and emulation. Evaluated across three attack scenarios, the approach reduces recovery execution time by 15.1% and increases recovery rate by 33.6% compared to frontier LLM baselines.
agentic incident responsedecision-theoretic planningrollout plannerdigital twinllm hallucination
Human-Centered Reflections on Care Robots: A Comparative Study of Caregiver Perspectives
This comparative study examines caregiver perceptions of four care robot categories (delivery, patient transfer, vital monitoring, mobility assistance) across the US, Mexico, and Chile (N=298). Using a mixed-factorial design integrating UTAUT, CAN model, and ethical frameworks, it found positive evaluations for logistical/physical tasks over interpersonal ones. Qualitative analysis revealed cross-cultural consensus on benefits (workload reduction, safety) but divergent prioritization of ethical concerns like dependability and job displacement, informing context-sensitive robot design.
care robotsunified theory of acceptancecognitive-affective-normative modelmixed-factorial designethical framework
MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models
MonitrLLM introduces open-source infrastructure for community-centered evaluation of large language models (LLMs), linking full conversation transcripts to user-reported task intent and outcomes as primary signals. A two-week pilot with 26 college students using ChatGPT collected 206 evaluation reports, revealing high average satisfaction (4.19/5) but a 23.1% task failure rate. Multi-turn conversations failed at 2.5 times the rate of single-turn exchanges, suggesting extended interaction signals difficulty rather than engagement. The study advocates integrating user feedback with observational data for robust LLM evaluation.
large language modelsuser feedbackconversation transcriptstask failure rateevaluation infrastructure
Antares: Foundation Models for Agentic Vulnerability Localization
Antares introduces a family of compact foundation models (350M, 1B, 3B parameters) for agentic vulnerability localization in software security. Based on IBM Granite, Antares employs a two-stage training pipeline combining supervised fine-tuning on cybersecurity reasoning and repository exploration with reinforcement learning from verifiable rewards. Evaluations show Antares-3B approaches GPT-5.5 performance while surpassing open-weight models 200x larger, with efficient local inference (500-task sweep in ~15 minutes on H100, <$0.002/task).
vulnerability localizationfoundation modelsreinforcement learningsupervised fine-tuninglow-cost inference
From fragmented data to actionable design: Physics-calibrated learning for plastic upcycling
The authors introduce PC-MG-MoE, a Physics-Calibrated, Missingness-Gated, and Load-Balanced Mixture-of-Experts framework for thermochemical plastic upcycling. The method learns directly from partially observed experimental data without target imputation, reconstructs physically consistent product distributions, and handles cross-laboratory heterogeneity. PC-MG-MoE achieved the lowest aggregate absolute error in source-grouped validation and demonstrated key composition-dependent trends in wet-lab experiments. Implemented as an interactive web-based workflow, it supports forward screening, physics-grounded inverse design, targeted experimental planning, and laboratory-specific adaptation. The framework provides a transferable approach for converting fragmented literature data into actionable guidance for plastic upcycling and broader thermochemical systems.
thermochemical upcyclingmixture-of-expertsmissingness-gatedcross-laboratory heterogeneityphysics-calibrated
Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification
This study benchmarks eleven audio classification methods for closed-set sound-source identification, including task-aware LLMs, fixed-vocabulary taggers, zero-shot models, and audio-grounded LLMs. Evaluated on 2,242 audio clips across 23 fine-grained classes and 11 categories, methods are grouped into four evaluation tiers, reporting macro Precision, Recall, F1, and false-negative rates. Gemini-3.1-Pro-Preview achieves the highest performance with 85.6% category-level F1 and 56.7% fine-grained F1, while Kimi-Audio-7B-Instruct shows competitive results for its size. SSLAM and CLAP perform well at the category level without candidate lists but lag in fine-grained classification. Analysis reveals response length does not predict accuracy, and incorrect answers are confidently stated 92-100% of the time.
closed-settask-awarefine-grainedzero-shotmacro precision
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
GROVE introduces a training-free framework for wearable assistants to perform both reactive question answering and proactive assistance using a unified memory system grown from continuous video streams. The method hierarchically organizes memory into fine-grained perceptual evidence, time-stamped moments, coherent episodes, and cross-day patterns, each paired with scale-native retrieval skills. Evaluated on benchmarks including MM-lifelong and EgoServe, GROVE achieves state-of-the-art results, with ablations demonstrating the complementary benefits of temporal strata and access skills, particularly for multi-day evidence.
video-memory systemstemporal stratascale-native retrievalproactive assistancereactive qa
Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training
The paper introduces Cooperative Parameter-subspace Evolution Strategy (CoPES), a cooperative coevolutionary method for memory-efficient post-training of tool-using LLM agents under resource constraints. CoPES decomposes the parameter space into lower-dimensional subspaces, optimizing them cooperatively to reduce GPU-hour requirements while maintaining performance. Evaluated on a Qwen3.5-4B agent for math tasks across five benchmarks, CoPES recovers 92% of GRPO's validation-accuracy gain versus 67% for standard ES, with less than one-eighth of GRPO's memory usage. It consistently outperforms standard ES and LoRA-based GRPO on pass@k metrics, demonstrating superior efficiency in resource-constrained settings.
cooperative coevolutionevolution strategiespost-trainingparameter-subspacetool-using agents
Chess on Ice: Curling Tactical Decision-Making via Backward Induction and Deep Reinforcement Learning
This work introduces a reinforcement learning framework for tactical decision-making in curling, addressing challenges of continuous state-action spaces, stochastic outcomes, and sensitivity to action perturbations. The authors employ Deep Deterministic Policy Gradient (DDPG), adapted for curling's finite-horizon structure, to learn strategies in a fully self-supervised manner without human-annotated data. Experiments on a four-rock variant demonstrate that the learned agent matches a hand-crafted expert heuristic in optimal regimes, quantified against the variant's intrinsic hammer advantage. The critic provides dense value estimates across the action space, enabling quantitative tactical comparisons for post-game analysis and athlete preparation.
deep deterministic policy gradientfinite-horizoncontinuous action spacestochastic outcomeshammer advantage
GLAIM: Learning Global and Local Adaptive Inter-Variable Dependency for Multivariate Time Series Imputation
GLAIM introduces a Global-Local Adaptive Inter-variable Dependency Modeling framework for multivariate time series imputation, addressing limitations of existing methods that either learn global dependencies across samples or dynamic local dependencies per sample. The framework combines a Stable Global Dependency Constructor, which derives robust global inter-variable dependencies from temporal representations, and a Sample-Conditioned Dependency Refiner, which adapts these dependencies to each sample and time step. Experiments on nine real-world datasets show that GLAIM achieves state-of-the-art performance under random and block missingness, remains robust to missing-rate shifts, and benefits from its complementary global and local components.
multivariate time seriesimputationglobal dependencieslocal dependenciesmissingness
Faster-WAM: Do World Action Models Need Deep Action Modules?
Faster-WAM introduces Dock of Transformer (DoT), a video-centric design principle that decouples action module depth from video backbone depth by docking lightweight output-heads onto pretrained video Transformers via a fusion interface. The method employs a single-layer action head on a 30-layer backbone, fusing keys/values from all layers with RoPE realignment. On LIBERO and RoboTwin 2.0, Faster-WAM matches performance while reducing latency to 66.5 ms (3.2× speedup over Fast-WAM) and demonstrating strong OOD generalization on LIBERO-Plus.
world action modelstransformer dockingrope realignmentinference latencyout-of-distribution generalization
SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents
SkillTrace introduces a graph-based method for composable LLM agents by structuring skill retrieval across three levels: compositional query relations, query-candidate similarity, and skill dependencies. The system organizes queries into semantic hierarchies, matches them with library skills, and propagates through dependency graphs. Evaluations on SkillsBench and ALFWorld show state-of-the-art performance, achieving 53.17% and 91.43% success rates respectively, with consistent improvements across backbone language models.
skill retrievalgraph-basedllm agentssemantic hierarchydependency propagation
KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model Improvement
KC-Agent introduces a dual-process cognitive architecture for automated ML model improvement, combining fast pattern recognition (System 1) with deliberate incremental updates (System 2) to address data drift. The system employs structured memory for knowledge consolidation, atomic change principles, and rollback capabilities, enabling efficient reuse of prior solutions. Evaluated on five datasets including NASA turbofan data, KC-Agent achieves 76.8% accuracy (outperforming CodeAct by +2.4%, Tree of Thoughts by +3.6%) with 13.2s execution time and 91% speedup over slow variants, while scoring 8.33/10 on LLM-evaluated strategic efficacy.
cognitive architecturedata driftknowledge consolidationatomic changesystem 1/system 2
Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
Hierarchical Memory Mamba (HMM) addresses the representation bottleneck in recurrent linear attention models (RLAs) by integrating hierarchical memory mechanisms into a pre-trained Mamba backbone. HMM employs a lightweight working memory to extract paragraph-level semantics from the backbone's hidden states, compressing them into persistent long-term memory for task-relevant retrieval. This hierarchical processing enables cross-task generalization through parametric learning, distinguishing it from other Mamba variants. Evaluations on Passkey Retrieval and LongBench-E tasks show HMM improves retrieval success by 34.3--37.1% and reasoning accuracy by 1.6--14.2% over baseline Mamba models, with only a 2% parameter increase and minimal training overhead.
recurrent linear attention modelshierarchical memoryparagraph-level semanticsparametric learninglong-sequence modeling
Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation
The paper introduces Simulated Randomized Controlled Trials (S-RCTs), a framework for using AI agents to predict A/B test outcomes before live deployment. The method decomposes error into agent approximation and subsampling components, supporting any behavioral model as the simulation engine. Evaluated on 67 historical marketing A/B tests, a foundation-model-based S-RCT achieves 0.70 sign overlap but overestimates effect sizes; calibration reduces squared prediction error by ~77×, while a within-subject design cuts standard errors by ~2.4×.
simulated randomized controlled triala/b testingagent approximation errorbehavioral modelingfoundation models
Hard Constraints, Smooth Gradients: Learning Feasible Inventory Policies via Differentiable Projection
We propose a deep reinforcement learning (DRL) framework for constrained sequential decision problems that enforces hard feasibility constraints via differentiable convex optimization. The method integrates a neural network for continuous action proposals, a quadratic program for relaxed feasible set projection, and a dual-informed integer mapping to restore integrality while maintaining feasibility. Trained end-to-end using pathwise gradients from a differentiable simulator, the approach achieves bounded error relative to exact integer projection and ensures full feasible action space coverage. Applied to multi-echelon production-inventory planning, the policy achieves <1% optimality gap on small instances, outperforms echelon base-stock policies by up to 9.75%, and reduces costs by up to 3.22% in industry-scale scenarios, particularly under tight capacity and high demand variability.
deep reinforcement learningdifferentiable convex optimizationmulti-echelon inventory planninghard constraintspathwise gradients
Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning
The paper introduces Diffusion Policy with Behavioral Advantage Correction (DPBAC), a novel offline reinforcement learning algorithm addressing distribution shift and Q-value estimation bias. DPBAC combines Behavioral Advantage Corrected Policy Evaluation (BAC-PE), which corrects the learned policy's Q-function using the behavior policy's Q-function, with diffusion models for policy representation and distribution matching. Theoretical analysis provides an upper bound on the Q-function approximation error. Experiments on D4RL benchmarks demonstrate DPBAC's superior performance, outperforming state-of-the-art methods in policy representation and bias mitigation.
offline reinforcement learningdistribution shiftq-value estimationdiffusion modelspolicy optimization
Context-Aware Mixture of Domain Experts for Bodily Expression of Emotion in the Wild
We present the Context-Aware Mixture of Domain Experts (CA-MoDE), a novel framework for bodily emotion recognition that explicitly models structured spatial context as complementary discriminative proxies. CA-MoDE incorporates dedicated scene and object experts to generate domain-conditioned soft distributions over emotion categories, which modulate the body expert's predictions at the distributional level rather than at the feature level. A task-tailored max-endorsement gating strategy fuses multi-domain signals by selecting the strongest contextual signal across experts for each emotion dimension, mitigating signal dilution from conflicting or uninformative context distributions. CA-MoDE achieves an Emotion Recognition Score of 0.3269 on the Body Language Database, outperforming existing temporal models using only single still images.
context-aware mixture of domain expertsdomain-conditioned soft distributionsmax-endorsement gatingbodily emotion recognitionstructured spatial context
FastGFDs: Efficient Validation of Graph Functional Dependencies with Desbordante
Proposes FastGFDs, a sequential algorithm for efficient validation of graph functional dependencies (GFD) on consumer-grade hardware, addressing the computational bottleneck of subgraph matching in existing parallel approaches. The method introduces Core-First Decomposition and Compact Path Index (CPI) to optimize graph traversal, implemented in the open-source Desbordante data profiler. Evaluations on real-world graphs show a 2.6× average speedup (up to 3×) and 5× memory reduction compared to parallel baselines, enabling GFD validation in single-node environments.
graph functional dependenciessubgraph matchingcore-first decompositioncompact path indexdata profiling
BRiG-AFA: Bellman Risk-to-Go Learning for Non-Myopic Active Feature Acquisition
BRiG-AFA introduces a non-myopic active feature acquisition method that learns candidate-conditioned risk-to-go functions via Bellman regression, enabling deployable supervised learning without complex optimization or density estimation. The approach fits these functions backward from terminal classification risk, using observed values, mask, candidate identity, and remaining budget for inference. Evaluations on Fashion-MNIST demonstrate consistent accuracy improvements, e.g., +10.20±0.74 points at four acquisitions, with a mean paired gain of 3.50±0.37 points across budgets {2,4,8,12,16}. MiniBooNE results show mixed performance at small budgets but positive gains at larger acquisitions, delineating current limitations.
active feature acquisitionbellman regressionrisk-to-gonon-myopic learningsupervised learning
Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit
The paper introduces agent-declared boundaries for self-segmenting long-horizon trajectories, enabling training on semantic phases without external segmentation. Agents propose falsifiable causal hypotheses during trajectory generation, exposing variable-length segments that yield four supervised targets, including correction transitions from failed regions. Evaluations show models attribute actions to governing hypotheses at over twice chance (paired sign test p = 0.0002), and human annotators match 24/40 boundaries vs. 11.5 random. Downstream DPO on 2,551 phase-boundary pairs improves adversarial held-out accuracy, with 4/60 wrong-to-right corrections vs. zero in controls.
self-segmentationfalsifiable hypothesestrajectory boundariessupervised targetsdpo optimization
MechGeo: Autoformalizing and Proving Euclidean Geometry in Lean 4
MechGeo introduces a Mathlib-native framework for autoformalizing and proving Euclidean geometry in Lean 4, combining GeoFormalizer for deterministic translation and repair of informal problems into GeoIR and Lean 4, and GeoProver for certified proof construction via geometric reasoning and algebraic certification. The system employs structural diagnostics, semantic evaluation, and counterexample-guided repair, with proofs verified by Lean's kernel. Experiments on 43 IMO problems show 29 successfully proved, 14 repaired after counterexamples, and 12/14 Lean-IMO-Bench statements proven or refuted, establishing a foundation for trustworthy formal geometry.
autoformalizationlean 4euclidean geometrycertified proofcounterexample-guided repair
Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning
The paper introduces TreeCredit, a shared-prefix credit assignment framework for adaptive multi-agent reasoning (MAR) that improves efficiency and accuracy. The method estimates operator utility via state-matched downstream comparisons, constructing shared-prefix collaboration trees and assigning correctness-prioritized suffix credits to state--operator pairs. These credits train a lightweight pairwise state router for dynamic operator selection during inference. Experiments on six reasoning benchmarks demonstrate that TreeCredit modestly enhances accuracy while significantly reducing inference cost, outperforming existing MAR methods in accuracy--cost trade-offs.
multi-agent reasoningcredit assignmentadaptive routingshared-prefixstate router
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
The paper introduces SKT, a verified synthetic data generation pipeline for improving language-model agents' skill-use capabilities by constructing skill-grounded tasks and executable trajectories. SKT selects single- and multi-skill configurations, synthesizes tasks via rule-based and agent-based verification with feedback-guided repair, and retains only trajectories that substantially use required skills. Using 2,000 public skills, SKT generates 4,000 task packages and 27,164 verified trajectories, with experiments showing supervised fine-tuning on SKT data consistently improves skill-use performance across models and benchmarks.
skill-grounded tasksverified data synthesisfeedback-guided repairagent-based verificationsupervised fine-tuning
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
Harness-R1 introduces a method for lifecycle-wide, failure-conditioned editing of executable runtime harnesses in LLM-based agents, enabling dynamic improvement beyond static model updates. It employs a 9B harness engineer trained via online reinforcement learning to convert agent failure trajectories into validated executable patches, optimized for task success. Cold-start supervised fine-tuning initializes the editing policy, followed by online training with group-relative policy optimization. Evaluations on WebShop, ALFWorld, and DBBench show Harness-R1 improves vanilla Qwen3.5-9B success rates from 44.3% to 53.6% (+9.3pp) and further to 64.2% (+5.0pp) after target-specific fine-tuning, demonstrating co-evolution of harness engineers and target agents.
harness editingreinforcement learningfailure trajectoriesexecutable patchestask success
TS-MAMP: A Remanufactured Agricultural Robot Powered by Second-Life EV Components and NMS-Free On-Device Weed Detection
The paper introduces TS-MAMP, a remanufactured agricultural robot leveraging second-life EV components to reduce costs for smallholder farming. The system pairs retired 48V BLDC hub motors via back-EMF matching and reuses lead-acid batteries with 60%-80% state of health, cutting powertrain costs by 60% (under USD 450). A truss chassis supports 200 kg loads and adjustable track width. An NMS-free YOLOv10n detector achieves 80.87% mAP@0.5 on the Wanxi Crop-Weed dataset, deployed via FP16 TensorRT on a Jetson Nano for on-device inference.
circular economyback-emf matchingnms-free detectionyolov10non-device inference
Self-Certification of Representation Adequacy: Sequential Certification at Minimum Task Loss
The paper develops a four-layer theory for self-certification of representation adequacy in agents using compressed history representations. It introduces decision-theoretic adequacy via a Bayes-risk grouping identity and formulates sequential certification as an optimal-stopping problem. Key results include an information-task-loss lower bound for δ-correct strategies and a Certification Track-and-Stop policy achieving asymptotic optimality. The work also presents an explicit kernel-switching example but leaves policy switching guarantees as an open problem.
representation adequacybayes-risk groupingoptimal-stoppingcertification complexitykernel-switching
Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow
The paper proposes SC-RF (Sparse-Constraint Rectified Flow), a generative detector for open-set visual text forensics that localizes tampering via restoration cost estimation rather than forgery-specific decision boundaries. The method combines Flow Matching adaptation with self-supervised Artifact Injection and a pixel-space Forensic-DiT to preserve high-frequency forensic traces. Evaluations on three benchmarks show state-of-the-art performance, with 3.2% and 4.8% F1/IoU improvements over prior work, plus strong zero-shot generalization to unseen text edits. An auxiliary stress-test reveals the detector's local harmonization can weaken statistical cues used by existing methods.
visual text forensicsrectified flowartifact injectionzero-shot detectionforensic-dit
Homebot: A Personal AI Agent for Conversational Home Assistance and Automation
Homebot introduces a locally deployable AI agent for conversational home assistance and automation, integrating voice and instant-messaging interfaces. The system employs a shared runtime architecture that combines language-model responses with registered tools and task-specific skills, while separating common request processing from session ownership. It features local wake-word detection, streaming speech recognition and synthesis, and an explicit dialogue-state protocol for managing conversation flow. The design emphasizes modularity through clear channel, tool, and skill contracts, enabling practical customization for household applications.
language-modelwake-word detectiondialogue-state protocolstreaming speech recognitiontask-specific skills
HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts
HarMoE introduces a dataset-aware mixture-of-experts framework for multi-source chest radiograph pretraining, addressing limitations of single-source image-report alignment in radiology vision-language models. The method employs lightweight residual experts in deeper decoder layers to disentangle shared medical semantics from dataset-specific variations, while leveraging a unified disease vocabulary with masked multi-dataset supervision to exploit clean annotations from heterogeneous datasets. Evaluations on large-scale chest X-ray benchmarks demonstrate consistent improvements in zero-shot classification, out-of-distribution transfer, and grounding over strong baselines. The approach highlights the importance of structured knowledge construction from diverse datasets with broader pathology coverage.
mixture-of-expertszero-shot classificationout-of-distribution transfervision-language modelsdataset-disentangled
Assessing the Impacts of Imperfect Datasets on Client Selections in Federated Learning
The study proposes a privacy-preserving scoring method to assess client contributions in federated learning (FL), addressing challenges posed by non-independent and identically distributed (non-IID) and noisy datasets. Experimental evaluations measure the impact of data quantity skews, label distribution skews, noisy data, and fairness in client selection on model accuracy and convergence. Results demonstrate the effectiveness of the proposed method in mitigating performance degradation caused by imperfect datasets and biased client selections.
federated learningnon-iidclient selectionprivacy-preservingconvergence latency
Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability
The review synthesizes advancements in robustness and explainability for trustworthy AI in digital health, addressing technical and ethical challenges across the AI lifecycle. It organizes methods into a structured framework, covering application-specific considerations in intensive care, neonatal health, and metabolic health, alongside techniques for data scarcity and distributional shifts. The paper evaluates metrics for trust, validity, fidelity, and diversity, with a focus on LLM-era systems.
trustworthy airobustnessexplainabilitydigital healthdistributional shifts
PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs
PosterMELD introduces a multi-agent pipeline for generating editable, print-ready scientific posters from papers, addressing limitations in existing systems (non-editable outputs or costly coding workflows). The method employs template-conditioned slot allocation, deterministic gating, and VLM-based failure repair to ensure geometric, readability, and factual integrity (81.3% Print-Ready Rate). It outputs PowerPoint and PNG artifacts with explicit design controls, achieving 3.4× higher PRR than P2P and 5.2× higher than PosterGen, while retaining native editability at USD 0.38 per request (3.5% of Codex+Skill costs). Conditional CHE scores from a frozen VLM confirm quality.
multi-agent pipelineprint-ready ratevision-language modeleditable outputsdesign controls
Fast Discovery of Inclusion Dependencies with Desbordante
The paper presents optimized implementations of two inclusion dependency discovery algorithms, Spider (a classic approach) and Faida (state-of-the-art approximate method), in the C++-based Desbordante profiler. For Spider, a parallelization technique reduces memory usage while improving speed; Faida benefits from four optimizations: data buffering, SIMD execution, hash-table selection, and parallelization. Experimental comparisons with Java-based Metanome show runtime improvements of 5× for Spider and 8× for Faida, demonstrating the impact of engineering optimizations alongside algorithmic design.
inclusion dependencydata profilingparallelizationsimdhash-table optimization
MEGRAG: Multi-Granular Evidence Graphs for Answer-Aware Multi-Hop RAG
MEGRAG introduces an answer-aware framework for multi-hop retrieval-augmented generation (RAG) using multi-granular evidence graphs. It constructs a cross-granularity index linking passages to sentences and triples, then dynamically selects evidence (triples, sentences, or passages) during retrieval. The method iteratively refines queries based on intermediate answers, stopping when the initial query is resolved. Experiments show consistent improvements over diverse RAG baselines.
multi-hop reasoningretrieval-augmented generationcross-granularity indexevidence graphsanswer-aware framework
PAC Approximation and DIRECT Optimization for Parametric Markov Models
We present a PAC approximation framework for parametric Markov decision processes (pMDPs) that combines scenario sampling, linear programming, and statistical model checking to efficiently approximate the satisfaction value function of PRCTL properties. The approach guarantees bounded error margins with prescribed confidence over most of the parameter domain. We integrate the DIRECT algorithm for derivative-free global optimization, establishing conditional optimality-gap guarantees under Lipschitz and PAC-good-set assumptions. Empirical evaluation on 2997 benchmarks shows DIRECT variants achieve slightly better objective values and faster runtimes compared to scenario optimization, while remaining within PAC margins.
parametric markov decision processespac approximationdirect algorithmstatistical model checkingscenario approach
From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents
We introduce IBA-Bench, a benchmark for implicit behavioral alignment in personalized LLM agents that addresses the knowledge-to-action gap by evaluating task execution under inferred user constraints from noisy, longitudinal interaction histories. Our proposed IBA-Agent framework reconciles conflicting priorities through broad retrieval and trajectory-level alignment. Experiments across nine application domains demonstrate that state-of-the-art LLM agents struggle with effective personalization, while IBA-Agent achieves substantial improvements in behavioral alignment for complex scenarios.
behavioral alignmentllm agentsknowledge-to-action gaptrajectory-level alignmentlongitudinal interaction
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution
The authors introduce a verifiable benchmark comprising 500 deep research tasks across 31 topics and 10 categories, designed to evaluate expert-level capabilities. The benchmark is constructed automatically using an iterative Explorer-Formalizer-Challenger pipeline, which transforms simple questions into complex tasks represented as directed acyclic graphs (DAGs) with atomic steps and checkpoints. This method ensures traceable verification and controlled evolution of tasks, queries, and rubrics. Experiments show the benchmark effectively discriminates between models and query types, while its fact-grounded pointwise rubrics enable fine-grained, human-aligned evaluation. Data, implementation, and results are publicly available.
verifiable benchmarkdirected acyclic graphiterative pipelinefact-grounded rubricsdeep research tasks
Lossless Tensor Compression as Program Synthesis
Brevis introduces a program synthesis approach for lossless tensor compression, addressing storage challenges in model checkpoints. The method employs a typed domain-specific language (DSL) with reversible operators to capture tensor structures, synthesizing compact programs via A* search guided by a learned checkpoint-specific prior. Evaluated on 10 public checkpoints (2.13 TB total), Brevis achieves a 33.93% reduction (1.41 TB), outperforming general-purpose (zstd, gzip) and tensor-specific (ZipNN, DFloat11) compressors, with throughputs of 3.60 GB/s (compression) and 6.61 GB/s (decompression).
program synthesistensor compressiondomain-specific languagelossless compressionmodel checkpoints
RamanPFN: learning from Raman spectral structure with a tabular foundation model
RamanPFN introduces a spectral representation framework that enhances TabPFN's in-context learning for Raman spectroscopy by preserving joint dependencies across spectral bands. The method combines Global Compositional Unmixing, which encodes distant bands with shared latent variation, and Local Vibrational Subspace Encoding, which captures peak morphology in contiguous wavenumber regions. Evaluated on 150 tasks from 74 datasets, RamanPFN reduced regression RMSE by 19.6% and classification error by 9.0% compared to direct TabPFN inference, demonstrating improved predictive performance for high-dimensional Raman data.
raman spectroscopytabular foundation modelin-context learningspectral representationnon-negative matrix factorization
Auditing Data Provenance in LLM Fine-tuning via Intrinsic Distributional Fingerprints
The paper introduces Distribution Provenance Audit (DPA), a post-hoc framework for detecting Data Intellectual Property infringement in black-box fine-tuned Large Language Models (LLMs) by analyzing intrinsic lexical-semantic distributional fingerprints. DPA formulates auditing as a statistical hypothesis test, leveraging unbiased output sampling to identify persistent data signatures despite adversarial obfuscations like paraphrasing or knowledge distillation. Experiments on medical and legal fine-tuning tasks demonstrate DPA's superiority over baselines, though the authors note a dual-use risk: the same fingerprints enabling reliable audits could facilitate privacy attacks.
data provenancelexical-semantic intersectionknowledge distillationstatistical hypothesis testadversarial fine-tuning
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs
Introduces PhyCheck, a video QA dataset for evaluating physical-law understanding in Video-LLMs, with coarse-grained (law compliance) and fine-grained (violation details) subsets plus diagnostic contexts. Uses structured supervision to improve physical consistency, tested on Fine-tune Qwen2.5-VL, showing gains in physical-consistency understanding but limitations in incorporating causal conditions. Reveals gaps between surface-level recognition and mechanistic understanding.
video-llmsphysical-law understandingquestion answeringfine-grained evaluationcausal context
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning
The paper introduces Multi-Moment Policy Optimization (MMPO), a novel framework for optimizing large language model reasoning by jointly minimizing multiple moments of the failure-probability distribution. MMPO operationalizes this as minimizing the expected truncated time to first successful response, extending beyond single-moment optimization prevalent in existing methods. A general moment-transformation framework is also developed to systematically induce diverse moment profiles, unifying a broader family of policy optimization objectives. Experiments across five mathematical reasoning benchmarks and models of varying scales demonstrate MMPO's consistent superiority over strong baselines.
policy optimizationfailure-probability distributionmoment-transformationmathematical reasoninglarge language models
UniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D Reconstruction
UniqueSplat introduces a view-conditioned 3D Gaussian Splatting model for generalizable 3D reconstruction, dynamically adjusting Gaussians based on viewpoint queries. The method employs a two-branch view-conditioned hyperNetwork to simultaneously learn view-agnostic embeddings and view-specific knowledge, enabling adaptation to specific views at test time. Evaluations on RealEstate10K, ACID, and DTU datasets demonstrate UniqueSplat's superiority over state-of-the-art methods, particularly in cross-dataset generalization, where it outperforms existing approaches.
3d reconstructiongaussian splattingview-conditionedhypernetworkgeneralization
Beyond Solution-Centric Search: Adaptive Inquiry and Knowledge Revision for Autonomous ML Engineering
The paper introduces Iris, an autonomous ML engineering system that shifts from solution-centric search to an information paradigm, where an evolving information state guides solution improvement. Iris employs an inquiry-revision loop: it generates local action plans for information acquisition via epistemic actions and manages knowledge through revisable claims synthesized from experiments. Evaluated on MLE-Bench, Iris achieves a 64.9% any-medal rate under a 12-hour budget, outperforming existing systems, and demonstrates cross-domain generalization across four tasks.
autonomous ml engineeringinformation paradigminquiry-revision loopepistemic actionsknowledge revision
Self-Improving Large Language Models via Progressive Experience Evolution
The paper introduces SPEE (Self-Progressive Experience Evolution), a post-training framework that bridges test-time and training-time self-improvement paradigms for large language models (LLMs) through experience distillation. SPEE employs explicit experience evolution—extracting, verifying, and consolidating transferable experience from interaction trajectories—followed by implicit policy optimization via reward-driven reinforcement learning. A global experience pool filters low-utility knowledge and mitigates post-hoc rationalization. Evaluations on five mathematical reasoning benchmarks show SPEE outperforms baselines across three model scales.
self-improvementexperience distillationpolicy optimizationreinforcement learningmathematical reasoning
MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents
The paper introduces MemArbiter, a function-aware memory arbitration framework addressing the Memory-Action Gap in long-horizon LLM agents by decomposing interaction histories into atomic items organized into five functional Memory Banks. It dynamically controls memory salience using bank-level demand, item-level relevance, focal-ambient representations, and a temporal presentation gate. Evaluated on ALFWorld with 500- and 750-token budgets, MemArbiter achieves success rates of 82.8% and 92.5%, outperforming Flat Retrieval and Flat Recency baselines by 20.9 and 25.4 percentage points, respectively, while improving post-failure recovery and reducing action repetition.
memory arbitrationlong-horizon agentsmemory-action gapfocal-ambient representationstemporal presentation gate
IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations
IACM-RL introduces a robust framework for complex tool invocation under dynamic intent fluctuations, addressing catastrophic intent deviation and infinite API loops. The method combines a DynamicIntent pipeline synthesizing trajectories across 13 fluctuation scenarios with a BeliefState-based Self-Generated Context Manager that tracks shifting goals and isolates stale parameters. Policy optimization employs a hierarchical intent-driven reward and three auxiliary losses for action calibration, context manager extraction, and state distillation. Experiments on DynamicIntent, BFCL-V3, and τ²-Bench demonstrate significant reductions in infinite loops and stale context errors, alongside improved out-of-domain generalization.
tool invocationintent fluctuationbeliefstatecontext managerpolicy optimization
Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior
The paper proposes Uncertainty-Aware Fusion (UAF), a dual-stream framework for robust animal vocalization classification that dynamically weights raw waveform and log-Mel spectrogram representations based on estimated Gaussian uncertainty. UAF outperforms static concatenation by 15.7-20.4% relative macro F1 on cross-species benchmarks (SoundWel pigs: 59.4% accuracy, DogBark: 73.1%), with uncertainty fusion identified as the primary performance driver through ablation studies. The method requires no reliability labels and addresses complementary limitations of acoustic representations in uncontrolled recording conditions.
uncertainty-aware fusioncrossmodal fusionlog-mel spectrogramanimal vocalizationgaussian uncertainty
DeGS: A Scalable 3DGS Architecture via Decoupled Workload Parsing and Reorganization
DeGS introduces a scalable 3D Gaussian Splatting (3DGS) architecture that decouples workload parsing and reorganization to address PE underutilization in existing accelerators. The method restructures the coupled α-checking, transmittance checking, and α-blending into consecutive stages, transforming fragmented workloads into dense, conflict-free tasks. Implemented in 28 nm, DeGS achieves 2.36×–7.25× throughput, 1.82×–6.02× speedup, and 1.59×–4.42× energy efficiency over GSCore, GBU, and GCC, maintaining >80% PE utilization at 1024 PEs for 8K rendering.
3d gaussian splattingpe underutilizationworkload reorganizationreal-time renderinghardware acceleration
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents
The paper introduces Fetch-then-Explore, a search agent framework that decouples page selection from evidence extraction by maintaining a persistent per-question workspace. Unlike visit-and-read or stateful browsing approaches, it records selected pages on the filesystem, enabling deferred and repeated extraction as hypotheses evolve. Evaluated on BrowseComp and WideSearch benchmarks with three agent backbones, Fetch-then-Explore achieves superior BrowseComp accuracy and competitive WideSearch performance, attributed to its ability to revisit pages and recover missed evidence across search trajectories.
search agentspersistent workspaceevidence extractionin-context learningopen-web benchmarks
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models
The study introduces an observability ladder to evaluate how different levels of access to reasoning traces impact correctness judgments in large language models. Using matched linear correctness predictors, the authors analyze responses, self-summaries, full traces, and internal signals across three benchmarks and five models (Qwen3, gpt-oss). Results show that summaries carry most of the trace's ranking signal without the prompt (mean AUROC 0.774 vs. 0.813), but this advantage collapses with the prompt visible. Full traces consistently outperform summaries, particularly in uncertainty and self-correction cues, even when length-matched. Monitorability is shown to depend on both display and reader capabilities.
observability laddercorrectness predictorsreasoning tracesself-summariesmonitorability
An AI-Based Decision-Support Pipeline for Day-Ahead Photovoltaic Forecasting
The study presents an AI-based decision-support pipeline for day-ahead photovoltaic (PV) forecasting, addressing challenges posed by short, imperfect records at newly deployed sites. The pipeline integrates timestamp correction, leakage-safe solar-geometry features, atmospheric context, and validation-learned stacking of complementary predictors. Evaluated at a UK charging-station site, the ensemble reduces daylight normalized RMSE by 32% against a clear-sky baseline under random day-blocked evaluation and by 9% under rolling-origin validation, outperforming individual machine-learning baselines by 6.6% and 6.4%, respectively. Results highlight the importance of physics-aware stacking, model class, and evaluation protocols in PV forecasting.
photovoltaic forecastingclear-sky baselinevalidation-learned stackingdaylight normalized rmserolling-origin validation
Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation
The paper introduces Instruction-Conditioned Exploration (ICE), a method to enhance exploration in post-training Reinforcement Learning (RL) for Large Language Models (LLMs) by supplementing task prompts with diverse instructions during training. It proposes Asymmetric-RL/SD, a combined Reinforcement Learning and Self-Distillation objective, to transfer explored behaviors to the unconditioned test-time policy. ICE with Asymmetric-RL/SD improves Qwen3-1.7B's held-out pass@1 performance by 5.0% on mathematical reasoning tasks at 4K response length compared to DAPO, with gains persisting at 8K context.
instruction-conditioned explorationasymmetric reinforcement learningself-distillationlarge language modelsmathematical reasoning
Geometry-Guided Layerwise FFN Width Allocation in Transformers
The paper proposes a geometry-guided method for layerwise allocation of feed-forward network (FFN) widths in Transformers, replacing uniform allocation. It quantifies FFN-induced geometric changes using correspondence-preserving shift, Gromov-Wasserstein distortion, and persistent homology under raw and normalized metrics. A surrogate model enables fixed-budget optimization. Experiments on seven pretrained language models show normalized-work schedules reduce validation loss versus uniform width and cosine taper baselines, with geometry-based allocations outperforming uniform by larger margins at 440M parameters.
feed-forward networksgromov-wassersteinpersistent homologylayerwise allocationtransformer architecture
Cross-Fitted Residual Utility for Primary-Preserving Cognitive Decision Correction in Automatic Modulation Classification
The study introduces a primary-preserving cognitive decision policy with cross-fitted residual utility to enhance automatic modulation classification by overriding default predictions when justified by heterogeneous evidence. The method employs a structured KAN-Fourier classifier for default probabilities, supplemented by neural and non-neural candidates, with residual utility learned from train-split out-of-fold predictions. Validation splits freeze action thresholds, transitions, routes, and risk masks. Evaluated on RMLA, RMLB, and HISAR datasets, the system improves accuracy from 63.632% to 66.332%, 65.161% to 66.168%, and 77.769% to 79.867%, respectively. Controlled comparisons and stress tests under various impairments confirm consistent gains.
automatic modulation classificationcross-fitted residual utilitykan-fourier classifierout-of-fold predictionscognitive decision policy
TBSG-Net: Temporal Bipartite Scene Graph Network for Fine-Grained Video Moment Retrieval
TBSG-Net introduces Temporal Bipartite Scene Graphs (TBSGs) to address limitations in Static Scene Graphs (SSGs) for Video Moment Retrieval (VMR). The model leverages Dynamic Scene Graphs (DSGs) to capture evolving object interactions over time, resolving SSGs' lack of temporal dynamics. A Dynamic Scene Graph Embedding (DSG-E) module encodes temporal spans and spatio-temporal information, employing a TBSG Constructor and a hybrid TBSG Encoder combining Transformer and Graph Convolutional Network architectures. Experiments show TBSG-Net outperforms all baselines, demonstrating its effectiveness in fine-grained VMR.
temporal bipartite scene graphsvideo moment retrievaldynamic scene graphsgraph convolutional networktemporal span encoding
TextNCA: Neural Cellular Automata for Language Modeling via Hierarchical Local Attention
TextNCA introduces a hierarchical Neural Cellular Automaton (NCA) for language modeling, employing 1D causal windowed attention across three stages with window sizes {8, 32, 128} and shared-weight iterations. The model, trained on WikiText-103 with ~30M parameters, achieves a perplexity of 60.3, underperforming parameter-matched Transformers (52.8 and 44.7 PPL). Analysis reveals that the staged narrow-to-wide schedule primarily drives performance, with iteration providing a smaller bounded benefit. Optimal iteration count is T_s=4, with degradation beyond this point. GRU gates and learned per-step embeddings are essential for iteration benefits. The study positions TextNCA as an analytical probe into NCA-style computation in language modeling.
neural cellular automatalanguage modelinghierarchical attentionperplexitywikitext-103
CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
CompanionBench introduces a theory-anchored, real-world-grounded benchmark for evaluating AI emotional companionship, addressing limitations in existing evaluations. It uses de-identified real-world data for scenarios and a trained user simulator, operationalizing ten capabilities from 25 psychological theories, including novel metrics like holding ambiguity and calibrated challenge. The benchmark employs a hidden disclosure gate, cross-family panel evaluation, and Item Response Theory to mitigate biases. Testing 28 agents revealed capability-level differences, with emotion regulation and calibrated challenge as common weaknesses, while role-play agents performed poorly. The method ensures reproducible rankings (rho = 0.996 ZH / 0.953 EN) and includes 500 bilingual parallel pairs.
companionbenchemotional companionshipitem response theoryhidden disclosure gatein-context learning
HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning
The paper introduces Hypergraph-Based Paired Failure Attribution (HPFA), a framework for localizing root causes of reasoning failures in large language models (LLMs) by comparing hyperedges between failed and successful reasoning paths. HPFA reduces search space via hypergraph-based analysis of non-linear logical dependencies, enabling efficient data synthesis to train a lightweight attributor via supervised fine-tuning and reinforcement learning. Experiments on mathematical reasoning and agentic coding tasks show HPFA improves attribution accuracy and efficiency, with the trained attributor enhancing test-time reasoning performance over flat-sequence or unpaired baselines.
failure attributionhypergraphreasoning pathsupervised fine-tuningreinforcement learning
EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
EduZone introduces a framework for evaluating LLM safety in K-12 education by combining student- and teacher-facing usage contexts, curriculum concepts, and 6 risk categories (28 subcategories) to generate adversarial interactions. The method tests ten LLMs across single-turn, static multi-turn, and dynamic multi-turn conversations, assessing four safety levels (refusal to fully risky assistance). Results show heightened vulnerability to education-specific risks and dynamic interactions, with current guardrails proving inadequate. The framework enables automated, scalable safety evaluation for educational LLM deployment.
llm safetyadversarial interactionseducation-specific risksmulti-turn conversationssafety guardrails
MANGO-Grasp: Mahalanobis Fields over Geometry-Oriented 3D Gaussians for Cross-Embodiment Dexterous Grasping
MANGO-Grasp introduces an anisotropic interaction framework for cross-embodiment dexterous grasping, representing objects as geometry-oriented 3D Gaussian primitives and robot hands as morpho-kinematic descriptors. The method employs Mahalanobis fields over keypoint-primitive pairs for interaction prediction and grasp optimization, enabling shared hyperparameter settings across heterogeneous multi-fingered hands. Evaluated on CMAP and MultiGripperGrasp benchmarks, MANGO-Grasp outperforms the strongest seen-hand baseline by up to 8.24 percentage points in simulation and improves zero-shot transfer to the unseen SharpaWave hand by up to 16.57 percentage points, achieving 86% real-world success.
mahalanobis fields3d gaussian primitivesmorpho-kinematic descriptorscross-embodiment graspinganisotropic interaction
Before Reasoning Fails: Pre-Evidence Procedural Failures in Agentic RAG
The study identifies pre-evidence procedural failures in agentic retrieval-augmented generation (RAG) systems, where agents retrieve but skip inspecting evidence before finalizing answers. Using tool-call traces and retrieved evidence from 12,000 trajectories on HotpotQA, 2WikiMultiHopQA, and MuSiQue, the authors decompose errors into pre-evidence discipline failures (11.2-13.1% co-occurrence) and post-gold-read failures. Introducing Read-Gate, a runtime invariant enforcing evidence inspection, improves LLM-Acc by 14.9-19.9 points on skip-prone trajectories and 3.2-9.4 points overall, demonstrating that evidence-gathering requires separate trajectory-level control.
retrieval-augmented generationprocedural failuretrajectory analysisruntime invariantevidence inspection
HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents
HALT introduces a verification-aware stopping policy for retrieval-augmented search agents, addressing the stopping problem by framing it as evidence coverage rather than generator confidence. The lightweight method leaves the search agent unchanged and halts only when cumulative evidence supports each required claim derived from expected hop claims. Evaluated on three multi-hop QA benchmarks, HALT reduces redundant search while largely preserving exact match accuracy, with generated claims yielding smaller but exact-match-preserving savings and gold claims demonstrating larger potential savings. Ablations confirm the policy's reliance on claim-evidence alignment, not generic sufficiency or lexical overlap. Open-corpus pilots suggest HALT abstains when coverage cannot be reliably verified, offering a practical runtime control signal without retraining or modifying the host agent.
retrieval-augmented agentsevidence coveragemulti-hop qaverification-aware stoppingclaim-evidence alignment
Evolving in the Agent Jungle via History-Informed Opponent Awareness
The paper introduces OASE (Opponent-Aware Selective Evolution), a method for stable skill adaptation in dynamic multi-agent environments where opponents continuously update strategies. OASE evaluates candidate skills against historical opponent snapshots via paired comparisons, adopting revisions only when payoff gains exceed a threshold. Evaluated in first-price auctions and private-cost Cournot competition, OASE reduces final equilibrium distance by 18-22% compared to Reflexion-style baselines while accepting 35% fewer revisions, demonstrating efficient adaptation with evidence-based selection.
multi-agent learningskill adaptationopponent modelingdynamic environmentsequilibrium distance
TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous Driving
The paper introduces TALSC, a timeliness-aware framework for infrastructure-assisted autonomous driving that coordinates large (LVLMs) and small vision-language models (SVLMs) to balance accuracy and latency. It models Age of Information (AoI) evolution, coupling it with token length and task performance to derive a timeliness metric, and proposes an online scheduling algorithm using Lyapunov drift-plus-estimated-penalty with performance guarantees. Evaluations on nuScenes show TALSC achieves up to 12.6% normalized Micro-F1 improvement over baselines under varying communication and computing conditions.
vision-language modelsautonomous drivingage of informationonline schedulinglyapunov optimization
Long-Horizon Autonomous Architecture Research with a Language-Model Agent: A Behavioural Case Study
The study investigates autonomous neural architecture design by a single large language model (LLM) agent over ~100 sequential experiments, improving a Vision Transformer from weak baseline to sub-SOTA on ImageNet-1K. The agent operates with scientific inputs, compute budget, and research tools, progressing through phases with expanded action surfaces. Key findings include phase-structured productivity (early gains, saturation, recovery), disproportionate early-hypothesis impact, workflow-induced greedy search tendencies, and independent rediscovery of established results. Workflow design significantly influenced outcomes, prompting proposals for diversified search and budgeted moonshot hypotheses in future autonomous research.
autonomous researchvision transformerlong-horizonworkflow designincremental hypothesis
SPARE: Structural Parameter-Free Affinity Regularization for Flow Matching
The paper introduces Structural Parameter-free Affinity Regularization (SPARE), a method to accelerate flow matching by regularizing intermediate token affinities to match those of clean data latents, avoiding external encoders or projection heads. SPARE exploits cross-image token relations, calibrating them with a unified objective, unlike prior target-free methods that repel such pairs. Evaluated on ImageNet 256×256 with SiT backbones under 400K iterations, SPARE adds no parameters and minimal memory (0.08 GB), achieving the lowest FID among parameter-free regularizers, recovering 37-54% of REPA's FID reduction, and reaching FID 1.90 at 1M iterations with classifier-free guidance.
flow matchingaffinity regularizationtoken relationsparameter-freefid reduction
AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning
AdaThinkV introduces an adaptive framework for token-efficient video reasoning by dynamically selecting between explicit chain-of-thought (CoT) reasoning and direct answering. The method employs reinforcement learning with matched rollouts, ThinkGain for prompt-level utility estimation, and Variance Recovery Policy Optimization (VRPO) to handle difficult prompts with sparse rewards. Evaluated on a unified video reasoning benchmark, AdaThinkV achieves 40.79 mean accuracy with 257.20 average output tokens, surpassing baselines by 2.98 accuracy points while reducing token usage by 22.7%.
chain-of-thoughtreinforcement learningtoken efficiencyadaptive reasoningvariance recovery
Music Restoration via Latent Operator Optimization and Diffusion Model Priors
LOUDAR (Latent-space Optimization of Unknown Distortion for Audio Restoration) is a general-purpose music restoration method that models unknown distortions as learnable latent operators in a pretrained audio autoencoder's latent space. It alternates between estimating clean latent variables and updating operator parameters during inference, regularized by an unconditional latent diffusion model prior. Evaluated on singing voice effect removal, restoration, and guitar distortion removal, LOUDAR improves over degraded inputs and matches supervised/unsupervised baselines in waveform and latent domains.
latent operator optimizationdiffusion prioraudio restorationautoencoderunknown distortion
FAST-GS: Frequency Aware Space-time Gaussian Splatting for Photorealistic Dynamic Novel View Synthesis
We propose FAST-GS, a frequency-aware space-time Gaussian splatting method for photorealistic dynamic novel view synthesis, addressing limitations in existing 4D Gaussian Splatting (4DGS) approaches. Our method introduces a Fourier Motion Modeling module that decomposes motion into frequency-based sinusoidal components, capturing both low-frequency global trajectories and high-frequency local details for accurate complex motion modeling. A motion-aware regularization strategy with frequency-dependent weights is integrated into the loss function to suppress high-frequency jitter while preserving low-frequency coherence. Experiments on N3V and Google Immersive datasets demonstrate improved motion fitting and long-term stability while maintaining real-time rendering capabilities.
4d gaussian splattingfourier motion modelingnovel view synthesismotion-aware regularizationfrequency decomposition
Agentic Self-Healing for Data and AI Pipelines: An Affordable Vendor-Agnostic Architecture using Open-Source Software
The paper proposes a vendor-agnostic reference architecture for self-healing data and AI pipelines using open-source tools, addressing fragmentation in existing solutions. The architecture integrates monitoring, metadata, incident history, policy checks, AI-assisted diagnosis, approval workflows, and remediation to automate pipeline issue resolution. It aims to reduce manual effort in detection, diagnosis, repair, and verification while remaining adaptable across data engineering, MLOps, and software delivery environments.
self-healing pipelinesvendor-agnosticmlopsobservabilityremediation
Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory
The study introduces the Recurrent Divisive Normalization Network (RDNN), a biologically inspired model that robustly maintains continuous working memory by leveraging divisive normalization. Unlike classical attractor networks or standard RNNs (e.g., GRUs, LSTMs), RDNN avoids manifold shattering through dynamic divisive normalization, which induces activity-dependent gradient scaling during Backpropagation Through Time (BPTT). Analytical and empirical results show RDNN converges to low-rank slow manifolds, with divisive normalization critical for stability under time-varying inputs. Ablations confirm subtractive inhibition alone fails to preserve dynamic continuity.
recurrent neural networksdivisive normalizationslow manifoldsbackpropagation through timeworking memory
A Contractualist Argumentation Framework for Moral Decision-Making
The paper proposes a formal framework for moral decision-making in autonomous agents, based on Scanlon's contractualism. The method extends ASPIC+, a structured argumentation framework, with value-based filtering to model morally relevant reasons through argumentation semantics. A worked example demonstrates the approach's application in domestic settings, contrasting it with existing value-based argumentation methods.
contractualismargumentation frameworkaspic+moral decision-makingvalue-based filtering
Semantic Networks as Clues: A Theoretical Foundation and Process Optimization for Semantic Network Construction
This paper establishes a theoretical foundation for Semantic Networks (SNs) representing textual non-propositional knowledge and introduces ClueNetwork, a framework for ranking candidate SNs. The authors argue that such SNs serve as clues rather than surrogates of reality, grounding their legitimacy in abduction. They review Semantic Network Construction (SNC) processes, including Automatic Keyphrase Extraction (AKE), Edge Weighting (EW), and Community Detection (CD), and define evaluation criteria for these stages. SNC is reformulated as a Process Optimization Problem (POP), and a global objective function integrating local criteria is proposed. ClueNetwork is demonstrated through illustrative experiments based on these criteria.
semantic networksabductionprocess optimizationkeyphrase extractioncommunity detection
Automatic Annotation of Ancient Greek Vowel Length
We present the first general-purpose macronizer for Ancient Greek, addressing the challenge of disambiguating vowel length in alpha, iota, and ypsilon (dichrona). The system processes CoNLL-U annotated input via recursive modules that propagate vowel length markings from common to rare word forms within the same lexical word. A character-level transformer trained on the macronizer's output achieves accuracy comparable to or exceeding the rule-based system on a manually annotated benchmark of verse and prose. Additionally, macronization improves downstream prosodical NLP tasks such as verse scansion.
macronizerdichronaconll-uverse scansionprosodical nlp
Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection
The study identifies evidence-bearing reasoning in prior chains of thought (CoT) as a textual shortcut that competes with visual recomputation in vision-language models (VLMs), affecting 16 tested models. A counterfactual analysis shows that removing evidence-bearing content shifts answer preference more than non-evidence context or final-answer spans, with shortcut strength modulated by evidence organization. The proposed Fresh-State Attention Firewall (FSAF) intervention isolates fresh computation, increasing visual update rates from 35.28% to 53.61% and reducing prior-answer rates from 39.22% to 3.67% across five VLMs.
vision-language modelstextual shortcutschain of thoughtcounterfactual analysisattention firewall
Effective and Efficient Context Retrieval via Partial Dependency Graph for Repository-Level Code Generation
DyRetriever introduces an efficient context retrieval method for repository-level code generation by constructing partial dependency graphs on demand, eliminating manual rule design and static global graphs. It employs an LLM to select entry-point functions and perform multi-hop reasoning, validating dependencies via semantic understanding. Integrated with similarity-based retrieval as DyCoder, it achieves relative Pass@1 improvements of 25.63% (CoderEval) and 59.73% (DevEval) over RAG baselines while being 7.4x faster.
repository-level code generationretrieval-augmented generationpartial dependency graphmulti-hop reasoningcoder evaluation
ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching
ProWorld introduces a progress-aware hyperbolic visual world model for long-horizon goal-reaching tasks, addressing limitations in existing JEPA-style models where local prediction consistency fails to ensure sustained goal progress. The method leverages goal-conditioned progress order, organizing latent-space dynamics via hyperbolic geometry, and employs hyperbolic entailment learning and future discrimination to maintain directional progress and resolve ambiguity. Experiments on four visual goal-reaching tasks show a 9.67% average absolute success-rate improvement over LeWM.
hyperbolic geometryvisual world modelsgoal-conditioned progresslong-horizon planninglatent-space dynamics
TransNRank: Towards Accurate Neoantigen Ranking with Transformer
TransNRank introduces a Transformer-based deep learning framework for accurate neoantigen ranking, addressing challenges in personalized neoantigen prediction such as data scarcity, noise, class imbalance, and complex immunogenicity features. The model leverages self-attention to capture local and global feature contexts and employs a positive-aware training objective to mitigate class imbalance. Evaluated on NCI, TESLA, and HiTIDE datasets, TransNRank improves top 20 recall rate from 46.9% to 53.1% while reducing training epochs from 200 to 20. Feature analysis reveals the importance of mutation at anchor and TCGA expression level, enabling dimensionality reduction without significant performance loss.
transformerneoantigenself-attentionclass imbalanceimmunogenicity
Diagnosing Search Behavior and Failure Modes in Long-Horizon Search Agents
The study diagnoses search behavior and failure modes in long-horizon search agents by analyzing retrieval and utilization gaps using human-annotated document-level relevance judgments. It evaluates six agents on BrowseComp-Plus and BrowseComp with an open-web search API, holding the retrieval model and evaluation harness fixed. Results show weak alignment between search effort and answer quality, with accuracy better correlated with retrieval recall than search steps; useful evidence often appears early, yet agents continue searching redundantly. The work identifies practical improvements for deep research systems, including query formulation, evidence selection, and stopping criteria.
long-horizon searchretrieval gapsutilization gapsretrieval recallquery formulation
CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents
CoEvoKG introduces a framework for co-evolving knowledge graphs (KGs) and search agents through mutual reinforcement. The system jointly trains a task generator that creates multihop questions from KG entity chains and a search agent that learns from correctness rewards and evidence-supported trajectories. Successful searches enrich the KG with verified evidence, creating a feedback loop for iterative improvement. Evaluations on six QA benchmarks (NQ, TriviaQA, etc.) show accuracy gains of +10.1 to +11.6 points over base models (Qwen2.5-3B/7B, Llama-3.1-8B) and +2.6 to +3.7 points over self-play/RL baselines.
knowledge graphself-evolving agentsmultihop qareinforcement learningevidence verification
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling
AFlex introduces an energy-efficient LLM serving framework by disaggregating Attention and FFN (A/F) operations and optimizing GPU frequency scaling. It employs a global scheduler and local DVFS controller to dynamically adjust A/F resource allocations and frequencies, alongside an interleaved pipeline with microbatch depth adaptation. Evaluated on NVIDIA A800 GPUs using Qwen3-32B and Mixtral-8×7B, AFlex reduces energy per token by up to 49% over prior disaggregated serving and 48% over frequency-scaling baselines while meeting TTFT and TPOT SLOs.
llm servingdynamic voltage and frequency scalingfeed-forward networksenergy efficiencymicrobatch pipelining
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models
The paper introduces SpeechAgent-R, a multimodal audio agent for tool-interactive audio reasoning, capable of coordinating intrinsic audio understanding with external skills and tools. The method involves supervised fine-tuning on interaction trajectories (65,492 samples, 507.6 hours of audio) followed by multi-turn reinforcement learning, evaluated on HIU-Bench (1,395 samples across 56 tasks). SpeechAgent-R achieves 84.17 and 70.94 on in-distribution and out-of-distribution tasks respectively, outperforming the base model by +15.40 and +14.23 points, demonstrating improved generalization in tool usage and workflow composition.
multimodal agentaudio reasoningreinforcement learningtool interactiongeneralization
LAB-Tab: LLM-Augmented Bayesian Network Adaptation for Few-Shot Tabular Generation
LAB-Tab introduces an LLM-augmented Bayesian network adaptation framework for few-shot tabular generation, addressing distribution shifts between source and target domains. The method first fits a Bayesian network (BN) from source data, then uses an LLM to propose plausible target-domain BN edges, followed by a PPO policy to calibrate edges via actions like weaken or flip. The adapted BN is sampled to synthesize target tables. Evaluated on six distribution-shift scenarios from US Census (ACS) tasks, LAB-Tab achieves best performance at 10% target-data budget, leading in four scenarios and reducing macro Overall score by 33.8% versus baselines, while excelling in JSD, WAPE, and UtilityGap metrics.
bayesian networkfew-shot learningtabular generationppo policydistribution shift
ReasonCast: Towards Explainable Time Series Forecasting with Reasoning
ReasonCast introduces a task-fused model for joint numerical time series forecasting and interpretable text reasoning, addressing the limitation of existing models that handle these tasks separately. The method involves finetuning any large language model (LLM) to generate both a reasoning chain and a forecast in a single autoregressive pass. A benchmark, ReasonTS-Bench, is proposed to evaluate the model's ability to identify five fundamental time series patterns. Experiments demonstrate that ReasonCast outperforms both LLMs and specialized TS models in prediction accuracy while producing verifiable, causal explanations.
time series forecastinginterpretable reasoninglarge language modelautoregressive passtask-fused model
No One Wins in Nuclear War: A Social Simulation of Military Decision-making
WOPR introduces a social-simulation environment for studying organizational decision-making in high-stakes scenarios, particularly military contexts, using a deterministic, replay-validated rules engine. The framework instantiates the card game Nuclear War, adhering to its published rules, and employs a decision-point contract to expose the engine to agents, enabling explicit strategic choices. WOPR integrates a four-rung press ladder for communication and models factions as collective command-and-control systems rather than single agents. The method is agnostic to social-simulation frameworks, with Concordia adopted as the default harness. All code, configurations, and replay data are publicly available.
social-simulationdeterministic enginedecision-point contractcommand-and-controlreplay-validated
Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
Wnuan introduces a three-stage pipeline for enterprise question answering that constructs task supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On WnuanBench (707 questions), the 32B model improves acceptable-answer rate from 52.76% to 91.51%, with residual-error sampling outperforming alternative methods by 2.97-3.11 points. General-benchmark performance decreases by 5.17 points, primarily in instruction following, while domain-expert agreement reaches 90.5%.
enterprise question answeringsupervised fine-tuningresidual-error samplinggeneral-data replayacceptable-answer rate
Beyond Magnitude and Shape: A Direction-Aware Loss for Time Series Forecasting
The paper introduces CosDir, a direction-aware loss function for time series forecasting that explicitly optimizes the direction of change via cosine similarity between prediction and target difference vectors. Unlike MSE, CosDir maintains gradient signals for small movements and is scale-invariant. The authors also propose CosDir-UW, an adaptive variant that learns the optimal mixing ratio between directional and magnitude terms. Evaluated across 100K+ experiments, CosDir variants consistently improve directional accuracy without compromising magnitude performance. Code is available.
time series forecastingcosine similaritydirection-aware lossscale-invariantadaptive weighting
EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning
EchoChange introduces a multimodal discrete diffusion language model for bi-temporal remote-sensing disaster change captioning, addressing cascading factual errors in autoregressive decoding. The model reformulates captioning as iterative masked-token denoising, enabling repeated revision of the entire caption conditioned on image pairs. It employs draft-aware dual-pass training, a progressive masking curriculum, and confidence-guided remasking to align training with iterative inference. Evaluated on the RSCC benchmark, EchoChange outperforms general-purpose and remote-sensing-specific baselines across lexical and semantic metrics, demonstrating significant improvements in factual accuracy and coherence.
diffusion language modelmasked-token denoisingdual-pass trainingremote-sensing captioningconfidence-guided remasking
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
The survey organizes robot learning into two paradigms: action-predicting weights (VLA models) and self-improving code-as-policy systems, analyzing 77 representative systems across six technique families. It introduces a taxonomy based on self-improvement mechanisms, ranging from zero-shot program synthesis to open-ended evolutionary loops (e.g., ASPIRE, ENPIRE, RoboClaw), and identifies five distinct uses of the term 'skill'. The analysis connects these technical approaches to challenges in robot-skill marketplaces, including adaptation, portability, and safety verification, while providing operational definitions and limitations for each family.
robot learningcode-as-policyself-improvementskill discoveryvla models
Physics-Informed Neural Networks for Complex Eigenfrequency Identification and Mode Structure Reconstruction of the Ground-State ITG Branch
The paper proposes a physics-informed neural network (PINN) framework combining Fourier feature encoding, complex-valued feature propagation, and three-stage training to jointly identify complex eigenfrequencies and reconstruct two-dimensional complex-valued mode fields of ion-temperature-gradient (ITG) drift waves in tokamak plasmas. The method addresses challenges from localized high-frequency oscillations, real-imaginary coupling, and nonlinear mode-eigenfrequency interactions. Experiments demonstrate accurate recovery of target eigenfrequencies and mode fields, outperforming baseline PINNs, while enabling analysis of higher-order drift-wave modes.
physics-informed neural networksion-temperature-gradient drift wavescomplex eigenfrequencyfourier feature encodingtokamak plasmas
Exploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models
We identify a critical blind spot in Multimodal Large Language Model (MLLM) unlearning evaluation: benchmarks fail to detect knowledge holes—severe degradation on benign inputs sharing generic patterns with the forget set. To address this, we construct a benchmark capturing unintended degradation and propose Selective Protection with Anchored Regularization (SPAR), which protects generic patterns via anchored activation filtering and reinforces them through entity-abstracted enhancement. Experiments on SafeEraser show SPAR recovers over 98% of vanilla response quality, compared to below 50% for baselines, while achieving 0.00% attack success rate and competitive utility. These results highlight the need for fine-grained evaluation in MLLM unlearning.
multimodal large language modelsmachine unlearningknowledge holesanchored regularizationactivation filtering
FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling
FOCUS introduces a post-training quantization framework for FP4 optimization in large language models (LLMs), addressing accuracy degradation via two key innovations. Coupled-Relaxation Scaling (CRS) decouples quantization and dequantization scales using a learnable full-precision coefficient, while Dual-Granularity Scaling (DGS) refines quantization scales at sub-block granularity. Evaluations across LLM families and benchmarks demonstrate state-of-the-art FP4 accuracy under MXFP4 and NVFP4 formats, with zero additional inference overhead.
fp4 quantizationcoupled-relaxation scalingdual-granularity scalingpost-training quantizationlarge language models
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning
The paper introduces Persistent Consistency Self-Distillation (PCSD), a method for improving reinforcement learning in large language model agents by addressing unreliable token-level supervision from privileged teachers. PCSD computes distillation weights based on persistent teacher-favoring signals using adaptive windows, exponential decay aggregation, trend-aware modulation, and sigmoid gating. Combined with GRPO, it achieves state-of-the-art results on ALFWorld (15.6-13.3 points over GRPO, 6.2-5.5 over SDAR) while remaining competitive on WebShop and improving 15.8 points on unseen ALFWorld splits.
self-distillationreinforcement learningtoken-level supervisionadaptive windowstrend-aware modulation
Rewriting or Reweighting? A Geometric Account in Language Models
The paper introduces behavioral manifold analysis to distinguish whether post-training alters language-model mechanisms via rewriting or reweighting. The method constructs low-dimensional charts in activation (ACT) and contribution (NOC) spaces to isolate behavior-specific geometry, applied to repetition and sycophancy failures. Results show supervised fine-tuning (SFT) rewrites inherited behavioral geometry, while reward optimization preserves underlying charts but changes behavior. Charts are compressed, partially alignable across architectures, with NOC space more architecture-robust than ACT space.
behavioral manifold analysisactivation spacecontribution spacesupervised fine-tuningreward optimization
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
DeepVoyager-VL introduces a long-horizon multimodal deep-search framework that integrates vision-in-the-loop reasoning for open-world problem-solving. The method constructs a multimodal event graph to synthesize data with intermediate visual dependencies and long reasoning chains, employs an agent framework for active visual acquisition, and fine-tunes models without reinforcement learning. Evaluations across ten benchmarks demonstrate its effectiveness in enhancing interaction depth and reasoning span compared to static parametric MLLMs.
multimodal large language modelsvision-in-the-looplong-horizon searchmultimodal event graphon-demand image loading
PartMat: Material-Aware 3D Part Decomposition with a Single Global Latent
PartMat introduces an efficient material-aware 3D part decomposition pipeline that represents multi-part geometry with a single global latent, addressing limitations of functional semantics and linear computational scaling in existing methods. The method combines PartVAE for unified part representation and single-pass decoding, a diffusion model with reinforcement learning for material assignment, and a sparse-voxel flow-matching model for geometric refinement. Experiments show PartMat outperforms baselines in material decomposition accuracy (87.3% vs. 72.1%) while maintaining geometric quality and inference efficiency (2.4× faster than part-independent methods).
3d part decompositionmaterial-awareglobal latentdiffusion modelsparse-voxel flow
SearchMaster: Grounded and Regulated Self-Play for Search Agents
SearchMaster introduces a self-play framework for training LLM-based search agents without human-labeled data, addressing common failure modes in task generation and rollout quality. The method employs an Evidence-Chain Generator (ECG) for grounded task creation, a Search-Depth Reward (SDR) for difficulty estimation, and an Over-Opening Penalty (OOP) to regulate tool use, with joint optimization via GRPO. Evaluated on six benchmarks, it improves a Qwen3.5-9B backbone from 38.19% to 51.52% accuracy, including a 30.1-point gain on BrowseComp-Plus.
self-playmulti-hop retrievalevidence chainsearch-depth rewardover-opening penalty
Predictive Maintenance: Deep Learning-Based Remaining Useful Life Prediction for Combat Aircraft Engines
A deep learning-based predictive maintenance model was developed for combat aircraft engine remaining useful life (RUL) estimation, addressing dynamic mission profiles. The model autonomously extracts temporal degradation features from multivariate sensor data using sliding-window sequential blocks (50- and 30-step for NASA C-MAPSS FD001 and FD004 datasets, respectively). It outperformed RF, CNN-LSTM, and BiLSTM baselines, achieving R2=0.8901, RMSE=13.28 (FD001) and 15.71 (FD004), with a 0.9973 AUC at the critical 30-cycle threshold. A decision-support simulator validated the protocol under aggressive flight profiles.
predictive maintenanceremaining useful lifemultivariate sensor datasliding-windowdeep learning
CockpitHAT: Dependency-Graph-Driven Hierarchical Attribution for Embodied Multi-Agent Cockpits
CockpitHAT introduces a hierarchical attribution framework for diagnosing failures in embodied LLM multi-agent systems, addressing Correctness Collapse by leveraging dependency-distance thresholds from interaction DAGs, integrating multi-channel evidence via an embodied adapter, and applying safety-uplift to high-risk failures. The method outperforms text-only baselines, achieving agent-level/step-exact accuracies of 77.9%/37.8% on Hand-Crafted and 86.5%/46.0% on Algorithm-Generated splits of the Who&When benchmark, surpassing ECHO by up to 17.6/16.7 points. On CockpitBench, a new benchmark of 212 annotated failure traces, it attains 78.3% agent-level and 38.2% step-exact accuracy, demonstrating effectiveness in risk-calibrated, multi-channel attribution.
correctness collapsedependency-distance thresholdsmulti-channel evidencesafety-upliftinteraction dags
LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation
LEAP introduces a scalable RL framework for low-level GPU kernel generation, addressing challenges of sparse rewards and compilation latency via Difficulty-Conditioned Pruning (DCP) and Rank-Based Reward formulation. DCP dynamically gates task expansion to focus on high-value tasks, while Rank-Based Reward leverages pairwise tournament outcomes for scale-free advantage estimation. Empirical results demonstrate LEAP's superior first-turn proficiency, multi-turn debugging resilience, and faster convergence compared to unpruned baselines.
reinforcement learningcuda kernel generationadaptive pruningrank-based rewardhardware alignment
CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning
CoNav-UAV introduces a cooperative dual-altitude aerial navigation system modeled as a Stackelberg game between a high-altitude leader (vision-language reasoning) and low-altitude follower (motion control), solved via Iterative Stackelberg Learning. The leader employs memory-based in-context learning, while the follower uses DAgger-style expert distillation, converging to equilibrium. Evaluated on AerialVLN, it outperforms baselines by up to 30.8 success rate points (learning scene) and 9.0 points (cross-scene) with 3× less adaptation data. Analyses reveal complementary gains across VLM backbones.
stackelberg learningvision-language navigationin-context learningdagger distillationaerialvln
Illuminating Visual Identity in Universal Multimodal Embeddings
The paper introduces a unified formulation for visual identity discrimination (VisID) in Universal Multimodal Embeddings (UMEs) and proposes MVEB (Multimodal Visual Identity Embedding Benchmark), a large-scale benchmark for evaluation and training. The authors present a learning framework that jointly optimizes general multimodal and identity-specific representations via identity-aware sampling. Experiments show the method enhances UMEs' identity discrimination while maintaining competitive general performance, addressing a previously underexplored capability in MLLMs. Code and data are released.
universal multimodal embeddingsvisual identity discriminationmultimodal large language modelsinstance retrievalidentity-aware sampling
Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction
ConfBench introduces the first calibration benchmark for vision-language models (VLMs) in key information extraction (KIE), addressing sparse low-accuracy regions in existing document benchmarks. The benchmark comprises 1,346 document variants generated via 20 degradation pipelines, enabling 70K+ entity-level evaluations across the accuracy spectrum. Evaluation of seven VLMs reveals OCR+Image modality yields superior confidence estimates, model capability drives calibration quality (Claude family shows monotonic scaling), and log-probability with first-token aggregation outperforms alternatives. The study also proposes ECARB for operational budget savings and releases ConfBench for systematic calibration research.
vision-language modelsconfidence calibrationkey information extractiondocument degradationoperational metrics
PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation
PICopilot introduces an LLM-based agentic framework for automating photonic integrated circuit (PIC) design script generation from natural language, addressing productivity gaps in script-based methods. The system employs a multi-agent architecture with feedback and a specialized retrieval-augmented generation (RAG) pipeline, optimizing success rate and reliability. Benchmarked on 48 diverse PIC scripting tasks, PICopilot achieves 100% success, outperforming general RAG-based LLM approaches (including GPT-5) by solving 21 additional tasks without significant latency or cost overhead.
photonic integrated circuitsscript generationmulti-agent architectureretrieval-augmented generationnatural language processing
Radar Detection in the CBRS Band: Techniques, Challenges, and Future Directions
This survey analyzes radar signal detection techniques in the 3.5 GHz CBRS band, focusing on interference mitigation between commercial networks and naval radars. It compares traditional energy-based and pattern-matching methods with machine learning approaches, highlighting their performance on public datasets under regulatory constraints (e.g., 99% detection probability, <60s latency). While conventional methods remain reliable in controlled settings, deep learning shows superior adaptability in complex environments. Key challenges include false alarms, co-channel interference, and real-time operation requirements, suggesting future systems will hybridize classical and learning-based techniques.
cbrs bandenvironmental sensing capabilityradar signal detectionspectrum sharinginterference mitigation
REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models
REFLEX introduces refinement-aware compute allocation for mixture-of-experts (MoE) in diffusion language models (DLMs), addressing the mismatch between expert computation and token refinement demands. The method employs a training-free, coarse-to-fine hierarchy for expert-budget allocation, using the Frontier-Progress Score to prioritize active blocks while retaining the default router. Evaluated on LLaDA-MoE and LLaDA2.0-mini, REFLEX reduces expert computation by 15% on average while maintaining or improving generation quality across benchmarks, outperforming autoregressive-style variable-expert routing in quality-computation trade-offs.
mixture-of-expertsdiffusion language modelsrefinement-aware allocationfrontier-progress scoretraining-free optimization
Investigating Social Bias in Narrative Image Generation
This work investigates social bias propagation in text-to-image (T2I) generation across narrative formats (photos, storyboards, comics) by adapting the BBG text-based bias framework to six models. Proprietary models exhibit 25.9% biased outputs in photo generation, increasing by 9.6pp and 18.2pp in storyboard and comic generation respectively. Narrative formats amplify biases through explicit event sequencing, character positioning, and textual elements, unlike subtle cues in photos, demonstrating the need for multi-format bias evaluation in T2I systems.
text-to-image generationsocial biasnarrative visualizationbias evaluationproprietary models
Multi-Source Dynamic Graph Learning for Compound-Flood Forecasting in Managed Coastal Systems
The study introduces an anchored forecasting framework for compound-flood prediction in managed coastal systems, addressing limitations in reproducing prolonged high-water plateaus critical for flood early warning. The method integrates multi-source dynamic graph learning, combining hydrometeorological and operational observations through state- and lead-dependent bounded residual corrections. This approach adaptively calibrates inter-site relationships and correction scales, preserving local temporal forecasts while selectively incorporating cross-site information. Experiments show improved reliability in predicting sustained high-water plateaus and maintained accuracy during routine hydrological conditions, enhancing flood early warning and water-management decision support.
compound-flood forecastingdynamic graph learningbounded residual correctionshigh-water plateausmulti-source regime representation
FRAMES: Guarded and Dual-Objective Skill Evolution for Agents in Policy-Governed Enterprise Workflows
FRAMES introduces a closed-loop framework for evolving LLM agents in policy-governed enterprise workflows, addressing challenges such as sparse feedback, regression risks, and inference cost constraints. The method cold-starts deployable skills from existing assets, employs consensus-based mutation, Pareto selection over accuracy and cost, and ensures anti-regression guarantees while preserving auditability. Deployed on an internal production system, FRAMES achieves the best accuracy-cost trade-off among baselines, with consistent gains reproduced on tau-bench.
llm agentspolicy-governed workflowsconsensus-based mutationpareto selectionanti-regression guarantee
Leveraging AI for fine-grained food safety risk forecasting in sparse data conditions
The study introduces a Transformer-based framework for city-level food safety risk forecasting, addressing data sparsity through a three-stage pretraining approach. The method unifies 11M inspection records with demographic, economic, and environmental indicators, leveraging Wilson interval-based partial supervision and semi-supervised label refinement. Evaluations on 2022 data show significant performance gains over baselines, with field experiments in Zhejiang Province demonstrating improved detection rates and resource allocation efficiency. Regulatory decision-making analysis suggests potential for further enhancements via AI-driven interfaces.
transformer-basedwilson intervalsemi-supervisedfine-grained forecastinginspection records
EntailLLM: Verifying LLM-Generated Vulnerability Discovery Paths with Domain Knowledge via Logic Programming
EntailLLM introduces a method for verifying LLM-generated vulnerability discovery paths by entailment against domain knowledge, addressing reliability gaps in safety-critical applications. The system validates paths as traversals of a binary's function call graph, aligned with a separate domain knowledge graph via temporal annotated logic. Evaluated across three CWE classes, four LLMs (e.g., GPT-3, PaLM), and seven binaries (405–12,696 nodes), domain knowledge improved pooled entailment from 78% to 98%, with only 3% degradation. Deployed on medical-device binaries, EntailLLM achieved 98% entailment without per-device tuning, inheriting formal guarantees from generalized annotated logic.
entailmentfunction call graphtemporal annotated logicvulnerability discoverymedical-device binaries
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
We introduce Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR), a method to mitigate trajectory anchoring bias in Vision-Language-Action models for autonomous driving by transforming future trajectories from pre-decision anchors into post-decision verification targets. DEFT-RLVR leverages Autonomous-Driving Multiple-Choice Question (AD-MCQ), which frames planning as selection among explicit trajectory candidates, disentangling high-level decision-making from low-level dynamics. Experiments demonstrate that DEFT-RLVR improves reasoning fidelity while preserving visual capabilities, offering a scalable foundation for verifiable autonomous driving research.
trajectory anchoring biasvision-language-action modelsautonomous-driving multiple-choice questiondeferred exposureverifiable reasoning
Can Urban Blight Be Accessed with Vision-language Models: A Case Study in Detroit
The study proposes a scalable framework for residential blight assessment using open-source vision-language models, leveraging structured prompts to evaluate housing attributes (roof integrity, wall damage, boarded openings) via binary and probabilistic outputs. Methodologically, it compares professional human annotations against multiple models, including XGBoost-based ensemble learning and weighted scoring, analyzing multi-view street imagery. Results indicate that (1) multi-view inputs improve accuracy, (2) vision-language models exhibit varied inference strengths, and (3) ensemble methods outperform base models, achieving robust blight assessment across diverse housing conditions.
vision-language modelsensemble learningurban blightxgbboostprobabilistic assessment
SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models
The paper introduces SPECTRA, a parameter-efficient fine-tuning framework for geospatial foundation models (GeoFMs) that addresses spectral mismatch and adaptation cost. SPECTRA employs Band-Routed Embedding (BRE) to map downstream sensor bands to the pretrained model's expected input space, and Stage-wise Transferability-aware LoRA (ST-LoRA) to allocate trainable parameters based on layer transferability. Experiments across three GeoFMs and four segmentation datasets demonstrate that BRE improves performance by utilizing all spectral bands, while ST-LoRA reduces parameters compared to full fine-tuning and standard LoRA.
geospatial foundation modelsspectral mismatchparameter-efficient fine-tuningband-routed embeddingstage-wise lora
RL-Lock: Reinforcement Learning for Generating Interlocking Assemblies
The paper introduces RL-Lock, the first reinforcement learning framework for generating interlocking assemblies without handcrafted heuristics. The method formulates assembly generation as a sequential decision-making problem, combining structured action chunking with MCTS-guided policy-value learning to navigate the combinatorial search space. Experiments show RL-Lock outperforms existing approaches, particularly in challenging cases where prior methods fail or require excessive computation time.
interlocking assembliesreinforcement learningmctsvoxel gridshape decomposition
MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents
MemSIF introduces a structured memory framework for LLM agents to address Temporal-Structural Misalignment (TSM) and Delayed Utility Manifestation (DUM) in long-term interactions. The method organizes raw interactions into Topical Segments and Event Trajectories (Structured Interaction Memory) and employs Dual-Track Fact Memory with CoreFact (schema-guided) and ActiveFact (on-demand) components. Evaluated on LoCoMo and LongMemEval-S across five LLM backbones, MemSIF achieves 2.29%-8.79% and 2.87%-6.15% higher Total ACC than baselines, demonstrating efficacy in mitigating TSM and DUM.
long-term memoryllm agentstopical segmentsdual-track memorytemporal-structural misalignment
Disagree to Accelerate: Closing the Loop on Diffusion Feature Forecasts
The paper introduces RACER, a training-free closed-loop controller for accelerating diffusion sampling by dynamically adjusting trust in feature forecasts. The method leverages forecast disagreement as a runtime signal to shrink uncertain predictions toward computed features or refresh risky steps, with deterministic error bounds. Evaluations on SD3.5-Large, FLUX.1-dev, Wan2.1-14B, and HunyuanVideo show improved quality at equal denoiser evaluations and faster sampling at equal quality for SD3.5, generalizing across forecasting designs.
diffusion samplingfeature forecastingclosed-loop controltraining-free accelerationdeterministic error bound
CoEvo-Mem: Co-Evolving Retrieval Policy and Memory Bank for LLM Agents
The paper introduces CoEvo-Mem, a framework for co-evolving retrieval policies and memory banks in long-term LLM agents. It addresses the feedback loop between memory retrieval and evolution by using a frozen LLM to generate query rewrites and routing priors, corrected by a lightweight residual router. Task outcomes credit routing decisions, while trajectory-conditioned feedback updates memory values and graph relations. CoEvo-Mem alternates between updating the router and memory bank to mitigate non-stationarity. Evaluated across seven benchmarks, it achieves state-of-the-art performance, demonstrating the efficacy of retrieval-memory coevolution.
retrieval policymemory bankcoevolutionresidual routernon-stationarity
DAPD: Dual-Anchored Policy Distillation
The paper proposes Dual-Anchored Policy Distillation (DAPD), a framework addressing privilege illusion in on-policy self-distillation (OPSD) for language models. DAPD employs Dual-Path Anchoring (DPA) to align reference and rollout behavior via a self-conditioned bridge, and Dual-Source Anchoring (DSA) for bidirectional correctness supervision. Experiments demonstrate DAPD's effectiveness, outperforming OPSD by +2.00 average points on Qwen3-4B, with scaling gains of +2.69 (4B) and +2.78 (32B).
policy distillationprivilege illusionself-conditioned bridgeon-policy learninglanguage-model alignment
X-KGRank: A Knowledge Graph RAG Framework for Explainable Recommendations via Pattern Mining and LLM Re-Ranking
X-KGRank introduces a knowledge graph retrieval-augmented framework for explainable recommendations, combining collaborative filtering with LLM-based reasoning. The method constructs a heterogeneous knowledge graph from the MovieLens-1M dataset (9,762 nodes, 999,264 edges) and trains a LightGCN ranker with SBERT initialization and a popularity selective routing strategy. Evaluated on MovieLens-1M, X-KGRank achieves NDCG@10 = 0.2956 and Recall@10 = 0.5371, outperforming a popularity baseline by 17.1%. A 1.5B-parameter LLM (Qwen2.5-1.5B) matches a 7B-parameter model (Mistral-7B) in explanation quality (0.97 vs. 0.94) but exhibits higher factual fabrication.
knowledge graphcollaborative filteringndcgllmsbert
TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics
We introduce TIDES, a longitudinal bilingual dataset capturing multi-party social dynamics in 12 university project teams over a semester, comprising 75,971 English and Korean utterances from in-person meetings. The dataset includes socio-structural annotations for interaction types, emergent roles, and development stages, enabling modeling of team evolution. Fine-tuning on TIDES improves next-speaker prediction by 13.8 percentage points over a bigram baseline (64.53%) and achieves performance comparable to proprietary zero-shot models, while using 42% less training data than state-of-the-art approaches. However, human evaluations indicate that improved structural prediction does not necessarily yield more natural or coherent utterances, highlighting a need for further research on multi-party generation.
longitudinal datasetmulti-party interactionnext-speaker predictionsocio-structural annotationsteam evolution
When Extreme Darkness Meets Motion Blur: MeanFlow for Unified RAW Restoration
We propose MeanFlow, a unified framework for extremely low-light RAW enhancement under realistic motion-degraded acquisition, addressing a gap in prior methods that overlook motion-induced degradations. Our approach introduces SIDED, a dataset with controlled motion degradation applied to extremely low-light RAW pairs, and a unified RAW tokenizer with domain-conditioned representation calibration to align low-light and well-exposed RAW data. MeanFlow performs enhancement in a single function evaluation, complemented by a physics-guided refinement model ensuring illumination--reflectance consistency, pixel fidelity, and color preservation without additional inference cost. Experiments demonstrate state-of-the-art performance in handling coupled motion and noise degradations.
meanflowraw enhancementmotion degradationdomain-conditioned calibrationphysics-guided refinement
MNC: Scope-Bound Semantic Declassification for Private LLM-Agent Communication
The paper introduces Minimum-Necessary Communication (MNC), a typed semantic-declassification protocol for private LLM-agent systems that selects task-sufficient disclosures from application-authored candidates and binds them to explicit scopes (recipient, purpose, forwarding, lifetime, logging, memory). A reference monitor enforces these scopes, while a history-aware extension mitigates inference risk from repeated disclosures. Experiments demonstrate that MNC preserves authorized delivery while blocking unauthorized forwarding, logging, and retrieval, outperforming text-only semantic declassifiers. MAGPIE executions confirm mediated disclosures propagate through planning, tool use, and memory retrieval, validating scope-bound declassification as a practical communication boundary.
semantic declassificationllm-agent systemsreference monitorinference riskmediated disclosures
LaCache: Robust Semantic Caching for LLM Serving
LaCache introduces a robust semantic caching scheme for LLM serving that mitigates cache-collision attacks by verifying both query embeddings and the first k speculatively decoded tokens. This dual-check design ensures formal security guarantees against adversarial query injection while improving cache retrieval relevance through enriched semantic context. Empirical evaluations across multiple LLMs and benchmarks demonstrate LaCache's resilience to attacks and efficiency improvements over existing caching methods.
semantic cachingcache-collision attacksllm servingadversarial queriesspeculative decoding
Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch
The study introduces a method using off-the-shelf coding agents as test-suite auditors to identify buggy submissions missed by official online-judge test suites. The approach combines adversarial test-suite generation with a certification chain involving multiple independent solutions and brute-force validation. Results show 589 verified buggy submissions in AtCoder's 20,375 audited accepted submissions, with five agents maintaining within 1.7pp of official-suite coverage on logic bugs and outperforming baselines on Codeforces problems lacking official suites.
coding agentstest-suite auditoradversarial test suitescertification chainbrute-force validation
Constructing Executable Analytical Knowledge Representations for Meta-Analysis Synthesis Using an Agentic Harness
The authors introduce Executable Analytical Knowledge Representation (EAKR), a machine-actionable framework for transforming structured evidence into executable meta-analysis, addressing challenges in evidence assignment, analytical contrasts, and methodological admissibility. EAKR is operationalized via MetaSynDec, an agentic harness combining large language models (LLMs) for structured updates and deterministic services for validation and execution. Across 58 synthesis units, MetaSynDec achieved complete object fidelity in 67.9% of cases and exact evidence-set agreement in 75.0%, outperforming direct LLM generation in synthesis-structure and formulation agreement (p<0.001). Results demonstrate EAKR's feasibility for formal validation, traceability, and statistical execution.
meta-analysisexecutable knowledgeagentic harnessstructured evidencevalidation
Beyond Single-Use Tokens: Durable Authorization State for Replay-Resistant LLM Agent Actions
CapLease introduces durable authorization state to prevent semantic replay in LLM agent executions, addressing the limitation of identifier-local tokens that permit fresh reissuance despite single-use constraints. The method binds authenticated user confirmations to canonical actions and enforces transactional Issue-Prepare-Commit transitions, ensuring replay-resistant execution across replanning, retry, delegation, concurrency, confirmation-replay, and crash-recovery scenarios. Results demonstrate that CapLease and Server Ledger prevent duplicate admission and external effects, identifying durable authorization state as essential for replay-resistant agent execution.
semantic replaydurable authorization stateidentifier-local tokenstransactional transitionsreplay-resistant execution
Rethinking Generative AI Literacy: An Integrative, Developmental, and Dialectical Framework for K-12 Teacher Education
The Responsible AI Literacy in Education (RAIL-Ed) framework addresses the generative AI literacy gap in K-12 teacher education through six interdependent pillars: Technical Fluency, Critical Evaluation, Human-AI Collaboration, Contextual Awareness, Ethical Reasoning, and Empowered Agency. Developed via systematic review and qualitative analysis of 67 studies (2023-2025), RAIL-Ed integrates critical, pragmatist, sociocultural, and human-centered traditions (Freire, Dewey, Vygotsky, Shneiderman). The framework features a developmental rubric (Emerging, Competent, Advanced) and positions ethics as constitutive rather than supplementary. It aligns with UNESCO and OECD/European Commission standards, offering a basis for curriculum design and teacher preparation.
generative ai literacyteacher educationethical reasoningsociocultural learningcurriculum design
Entity-Aware Sequence Transduction for Player-Centric Ball Action Spotting
The paper proposes Multi-Entity Denoising Sequence Transduction (ME-DST) for player-centric ball action spotting, addressing limitations in existing Denoising Sequence Transduction (DST) approaches that flatten player-role representations. ME-DST maintains role-slot dimensions throughout encoding, employing temporal attention for intra-player evolution and spatial attention for inter-player interactions, enhanced by learnable role embeddings and tactical features. Evaluated on FOOTPASS, ME-DST achieves a Micro F1 of 0.778, outperforming the TAAD+DST baseline by 10.3 percentage points, with ablations confirming the importance of entity modeling and role encoding.
multi-entity denoising sequence transductionplayer-centric action spottingtemporal attentionspatial attentionrole embeddings
ARM: Detector-Agnostic Changepoint Attribution with Finite-Sample Error Control
The paper introduces ARM (Attribution by Rank Maxima), a detector-agnostic method for changepoint attribution with finite-sample error control. ARM scores coordinates using a max-over-splits rank statistic, ensuring invariance to changepoint estimation accuracy. It provides three guarantees: per-coordinate validity, exact family-wise error control via Westfall--Young permutation, and FDR control under arbitrary dependence. Simulations show ARM maintains nominal error levels (unlike naive methods, which inflate error to 0.66+ in high dimensions) while preserving power and type-label accuracy. Applied to financial data from the 2008 collapse, ARM correctly attributes scale changes to asset classes while excluding controls.
changepoint attributionrank statisticfamily-wise errorfalse discovery ratepermutation test
ProtoAct: Turning Wet-Lab Protocols into Embodied Robotic Actions
ProtoAct introduces a structured protocol-grounding framework that converts free-form biological wet-lab protocols into executable robotic action sequences. The method combines ProtoRAG for context-sensitive parsing via retrieved annotated examples, RefineChecker for detecting and revising procedural inconsistencies, and ActSchema for mapping procedures into constrained JSON function calls. Evaluated on BioP2E (22 annotated protocols with 258 monitoring conditions and 910 subtasks), ProtoAct shows compatibility with seven large language models and enables successful execution in simulation and real-robot settings. Ablations confirm the complementary roles of retrieval, posterior checking, and schema constraints.
protocol groundingembodied roboticscontext-sensitive parsingaction schemawet-lab automation
GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks
GABench introduces a comprehensive benchmark for evaluating LLM agents on graph analysis tasks, addressing limitations in existing benchmarks by covering three graph types and four task categories (graph retrieval, graph theory, graph machine learning, open-ended QA). The benchmark provides 84 executable tools and a task generation pipeline, yielding 10,400 tasks with verifiable ground truth. Evaluations of frontier LLMs and agent harnesses reveal three findings: (1) LLM agents struggle with complex graph tasks, (2) harness choice significantly impacts performance, and (3) tool-call quality outweighs quantity in graph analysis.
llm agentsgraph analysisbenchmarktool-call qualityagent harness
When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary
The paper identifies authority collapse in LLM agent memory systems, where memory consolidation preserves claims but erases source authority constraints, leading to unauthorized downstream use. AuthMem-Bench is introduced as a paired benchmark to evaluate write-time collapse, authorization errors, and automatic authority preservation. Experiments across seven consolidators and seven LLM backbones show authority collapse in 48/49 configurations, with unauthorized-action rates of 50.3% without authority metadata, reduced to 0.0% when authority labels are preserved, while maintaining task success.
authority collapsememory consolidationllm agentsauthorization errorsauthmem-bench
Learning What to Remember: Test-Time Training via Context Distillation
The paper introduces Test-Time Context Distillation (TTCD), a framework for long-context modeling that optimizes memory allocation for future utility via self-supervised learning. TTCD employs a long-window teacher to supervise a short-window student, using hidden-state discrepancies as a dense signal to guide memorization of contextually relevant information. An in-place variant, IP-TTCD, leverages existing MLP parameters as fast weights. Experiments on long-context language modeling tasks demonstrate IP-TTCD's superiority over DeltaNet, Gated DeltaNet, sliding-window attention, and standard TTT. Additionally, IP-TTCD enables pre-trained transformers to adapt during inference, enhancing long-context capabilities with minimal architectural changes.
test-time trainingcontext distillationlong-context modelingself-supervised learningfast weights
TCPO: Turn-Level Credit Policy Optimization
TCPO introduces turn-level credit assignment for verifier-guided multi-turn reinforcement learning, addressing the gap between verifier scores and credit allocation. The method performs score-to-credit conversion via three mechanisms: retrospective credit for immediate progress, hindsight delayed credit for non-improving turns with later payoff, and selective fixed-history counterfactual estimation for high-surprisal turns. Evaluated on math reasoning, code generation, and AppWorld agent tasks, TCPO achieves best or tied-best Pass@8 on Qwen3-4B and DeepSeek-R1-Distill-Llama-8B, reduces turns to success, and enhances multi-turn agent performance, demonstrating its efficacy across model scales and verifier types.
credit assignmentverifier-guided rlmulti-turn optimizationscore-to-credit conversioncounterfactual estimation
Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation
The authors introduce SciStyleBench, a three-component benchmark to diagnose and mitigate stylistic bias in LLM-based scientific idea evaluation. SciStyleBench comprises (i) SciStyleStage, a controlled evaluation environment with 600 scientific ideas and 15 style variants across three contexts, (ii) SciStyleMetrics, including Style Bias Index (SBI), Substance Recognition Rate (SRR), and Adversarial Win Rate (AWR), and (iii) SciStyleExtractor, a module separating style from content. Experiments reveal that LLM judges are sensitive to writing style, while SciStyleExtractor reduces SBI from 0.566 to 0.501 and improves SRR and AWR from 0.504 and 0.554 to 0.759 and 0.899, respectively, demonstrating enhanced robustness in idea evaluation.
stylistic biasllm-based evaluationscientific substancestyle-conditioned evaluationadversarial win rate
Allocation Before Ranking: Decoupled Token Compression for OmniLLMs
Macer introduces a decoupled token compression framework for OmniLLMs, addressing the mis-specification of shared top-K ranking by separating modality allocation from token ranking. It assigns explicit budgets for audio and video tokens, then performs allocation-normalized ranking within each modality at shallow layers. This training-free approach significantly reduces token costs while preserving accuracy across multimodal benchmarks. At 25% retention, Macer maintains 98.7% and 97.3% of full-token performance on Qwen2.5-Omni-7B and Qwen2.5-Omni-3B, respectively. On OmniVinci-9B, it outperforms shared top-K ranking by up to 12.9 points, achieving OmniZip-level performance at lower FLOPs.
token compressionmultimodalallocation-normalized rankingtraining-freeshallow layers
Few-Shot Concept Prompt Learning for Segmentation Foundation Models via Visual Grounding
We introduce Few-Shot Concept Prompt Learning (FS-CPL), a method to improve segmentation foundation models (FMs) in domains with scarce image-text supervision. FS-CPL learns a continuous concept prompt embedding from a small support set of image-mask pairs via mask supervision, while keeping the encoder-decoder backbone frozen. Evaluated on four public benchmarks (BUSI, HC18, TN3K, CVC-Clinic) spanning ultrasound and endoscopy, FS-CPL achieves absolute Dice improvements of up to +0.62 over canonical text prompts. The method is backbone-agnostic, enhancing both vanilla SAM3 and domain-specifically pretrained Medical SAM3, demonstrating that visual concept prompting complements in-domain pretraining.
segmentation foundation modelsfew-shot learningvisual groundingcontinuous concept promptmask supervision
LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
LongCat Sparse Attention (LSA) addresses system-level bottlenecks in DeepSeek Sparse Attention by introducing three co-designed strategies: (1) Streaming-Aware Indexing for hardware-aligned memory access, (2) Cross-Layer Indexing to amortize overhead via layer reuse, and (3) Hierarchical Indexing for coarse-to-fine candidate selection. Evaluated on models from 69B-A3B to 560B-A27B, LSA matches full-attention performance while enabling training with up to 1M-token contexts and supporting LongCat-2.0 (1.6T-A48B). The open-sourced LongCat-Flash-Lite-Sparse (69B-A3B) integrates LSA with an updated long-context corpus.
sparse attentionmemory-access optimizationcross-layer distillationhierarchical indexinglong-context modeling
SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction
SyncPlan introduces a plan-execute-correct framework for efficient long-horizon LLM-based multi-agent coordination, addressing the trade-off between efficiency and adaptivity. The method employs explicit synchronization via wait primitives and deadlock detection, coupled with adaptive correction through a lightweight Plan Staleness Detector that triggers replanning when environmental changes occur. Evaluated on Overcooked and Honor of Kings benchmarks, SyncPlan achieves state-of-the-art task success rates with a 0.05% reduction in wall-clock runtime compared to existing LLM coordinators.
multi-agent coordinationlong-horizon planningllm-based coordinationadaptive correctiondeadlock detection
GISAgentBench: A Practitioner-Sourced Benchmark for Evaluating LLM Agents on GIS Tasks
The paper introduces GISAgentBench, a benchmark of 349 multi-step GIS tasks curated from GIS Stack Exchange with executable reference trajectories and ground truth outputs, enabling strict evaluation of LLM agents. Tasks are instantiated on real public data across six geographic areas, addressing limitations of existing benchmarks that rely on surrogate signals. Evaluation of six LLM models shows low performance, with the best agent completing only 32.7% of tasks under strict tolerance-aware scoring, despite producing outputs close to ground truth.
gis tasksllm agentsbenchmark evaluationground truth outputstolerance-aware scoring
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models
CRAFT introduces a recursive adaptive fusion method for compressing video tokens in vision-language models, addressing the trade-off between efficiency and adaptivity in token compression. The method decouples parameter-free token selection (via global similarity) from learnable fusion (via position-aware weighting and content-adaptive gating), preserving spatio-temporal coordinates and alignment with pre-trained language models. Evaluated on multiple video benchmarks, CRAFT achieves ~8× compression while retaining ~97% of the backbone model's accuracy, outperforming prior token-compression methods.
token compressionvision-language modelsadaptive fusionspatio-temporal redundancyparameter-free selection
StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring
StreamTalk introduces a closed-loop framework for real-time co-speech gesture generation, addressing drift in open-loop systems via key-pose anchoring. The method employs a generate-retrieve-refine cycle: Pose-Guided Generation predicts coarse motion clips, retrieves plausible tail poses from a speaker-specific database, and refines clips before proceeding. Training incorporates Stochastic Anchor Masking to recover motion from sparse boundaries and a part-aware DiT to separate hand, body, and translation streams. On BEAT2, StreamTalk achieves SOTA FGD, reduces long-horizon drift, and runs at 76 FPS.
co-speech gesture generationkey-pose anchoringstochastic anchor maskingpart-aware ditreal-time generation
AI-assisted Script Management for Requirements Elicitation Interviews
The paper introduces an AI-assisted workflow for requirements elicitation interviews, combining theory-guided script generation with real-time topic coverage tracking and on-demand follow-up question generation. A between-subjects quasi-experimental study compares AI-assisted (no training) and training-only (no AI) conditions, evaluating scripts, interview dynamics, and requirements artifacts. AI-generated scripts outperform training-only scripts (92.8 vs. 74.8/100), while AI-assisted interviews cover fewer topics (9.6 vs. 14.5), ask more follow-ups per topic (3.43 vs. 1.15), and produce more refined goal models (lowest-level goal fraction 0.653 vs. 0.598). Participants rated topic tracking as the most useful feature (86% agreement), demonstrating AI's role as an elicitation scaffold.
requirements elicitationscript generationfollow-up questionsgoal modelsquasi-experimental study
Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw
The paper introduces MemCollusion, an automated framework for constructing collusive memory poisoning attacks against LLM agents with persistent memory. The method employs salami tactics to split adversarial objectives into individually benign memory fragments that collectively induce harmful behavior, guided by four constraints, five strategies, and a fine-tuned generator. Evaluated on OpenClaw across 48 scenarios, MemCollusion achieves 81.3% Memory Save Rate and 75.0% Attack Success Rate, demonstrating resilience to benign memory dilution and defenses.
memory poisoningsalami tacticsllm agentscollusive attacksred-teaming
RING: Retrieval-Internalized Generation for Continual Large-Scale Knowledge Injection
RING (Retrieval-Internalized Generation) introduces a paradigm for continual large-scale knowledge injection by internalizing retrieval into a Mixture-of-Memory Experts, eliminating external retrievers. The method employs a three-stage training process: continued pre-training with Dual Causal Attention, supervised fine-tuning for search-then-answer patterns, and reinforcement learning with hierarchical rewards to optimize routing-and-search policies. Theoretical framing positions RING as a search-free approximation to retrieval-augmented generation (RAG). Evaluated on News-2025, a benchmark for new knowledge post-dating pretraining, RING matches or exceeds RAG and parametric injection baselines in accuracy and efficiency.
retrieval-augmented generationmixture-of-memory expertsdual causal attentionparametric memoryhierarchical rewards
QWRF-Net: A Quantum-Wavelet Framework with Rectified Flow for Short-Term Precipitation Nowcasting
QWRF-Net introduces a quantum-wavelet framework with rectified flow for short-term precipitation nowcasting, addressing multi-scale representation degradation in radar precipitation fields. The method decomposes latent features into wavelet sub-bands, applies quantum-inspired modulation in the decomposed space, and generates future sequences via a rectified-flow-based non-autoregressive decoder. Evaluated on KNMI radar and SEVIR benchmarks, QWRF-Net shows consistent gains at medium-to-high precipitation thresholds, particularly in preserving intense cores and fine-scale structures, with ablations confirming the complementary benefits of wavelet disentanglement, sub-band modulation, and flow-based generation.
quantum-wavelet frameworkrectified flowprecipitation nowcastingwavelet sub-bandsnon-autoregressive decoder
When Memory Updates but Behavior Does Not: Repairing Implicit Stale Dependencies in Personalized Agent Responses
The paper introduces StateAuditor, a method to address the implicit policy adaptation (IPA) gap in memory-augmented agents, where agents plan around outdated user states despite knowing they are stale. StateAuditor audits stored state to draft responses, using an LLM to propose old-to-new transitions verified by deterministic checks for provenance and chronology. On the STALE benchmark (400 scenarios), StateAuditor achieves a +5.0-point gain (95% CI [+2.9, +7.2]) in VTA scores over a locked predecessor, primarily from improved IPA and premise resistance. Cross-family validation on HorizonBench shows preference accuracy gains (user-clustered p<.01), though the draft-side audit contributes most. A matched control attributes STALE gains to transition machinery, not added context.
implicit policy adaptationmemory-augmented agentsstateauditorstale benchmarkprovenance verification
Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge
The paper introduces Linear Multi-Timescale Retention (LIA-MTR), a memory-efficient cross-modal bridge for Vision-Language Models (VLMs) that replaces Softmax Multi-Head Attention (MHA) to address its O(N²) complexity. LIA-MTR combines ELU-based feature mapping, adaptive write-gating, and log-linear recurrent decays to compress visual sequences into bounded states with O(N) complexity. Evaluations show flawless context routing across 16K tokens, 262K-patch processing in 11.2GB VRAM, and a 71.00% score on MME (10% higher object permanence than MLP baselines), enabling infinite-context VL integration.
vision-language modelslinear attentionmemory complexityobject permanencerecurrent decays
Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
The study demonstrates that long-horizon post-training on office workflows improves software engineering performance via cross-domain transfer of goal-directed execution (GDE) behaviors. Post-training Qwen3.5-122B-A10B on 363 Long-Horizon Multi-Tool Agent (LHMTA) tasks—without software-engineering content—yielded a 5.8-point pass@1 improvement on SWE-Bench Pro. Trajectory analysis confirmed gains in GDE components (goal selection, state construction, objective fidelity, verification) and SWE-Bench metrics (information gathering, implementation, verification), suggesting domain-general behavioral adaptation.
goal-directed executionlong-horizon post-trainingcross-domain transferqwen3.5-122b-a10bswe-bench pro
PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters' Lack of Knowledge
PICTURE introduces a prompting method that enhances Theory-of-Mind (ToM) reasoning in large language models (LLMs) by explicitly revealing characters' lack of knowledge in free-form Chain-of-Thought (CoT) explanations, avoiding performance degradation from event hiding. Unlike prior approaches that remove unknown events (perspective-taking), PICTURE enables LLMs to inhibit responses to such events by making ignorance explicit during reasoning. Experiments demonstrate a 7.3% average improvement over existing methods on false-belief tasks.
theory-of-mindpromptingchain-of-thoughtfalse-belief tasksperspective-taking
HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning
The paper introduces HindSearch, a hindsight self-distillation method for search-augmented reinforcement learning (GRPO) that critiques failed trajectories using a frozen judge with gold-answer access. The judge generates auxiliary on-policy distillation signals for the student's search actions, addressing the limitation of binary exact-match rewards in conventional approaches. Evaluated on a seven-benchmark suite with Qwen2.5-3B-Instruct, HindSearch achieves 39.4% average exact-match accuracy, surpassing prior search-RL baselines. Ablation shows gold-answer access is critical, confirming hindsight as the key improvement factor.
hindsight distillationsearch-rltrajectory critiqueon-policy learningexact-match reward
Latent Thought Credit: Multi-Answer Credit Assignment for Latent Reasoning
The paper introduces Latent Thought Credit (LTC), a hierarchical credit-assignment framework for latent reasoning in language models, addressing the challenge of disentangling thought quality from answer-sampling noise. LTC samples multiple latent thoughts per prompt, fixes the context after each thought, and estimates thought-level expected reward by averaging rewards over multiple answers from fixed contexts. It optimizes latent-thought and answer phases using advantage-weighted objectives and a thought-matching loss. Evaluated on mathematical reasoning and STEM multiple-choice tasks, LTC achieves superior average accuracy, with ablations confirming multi-answer estimation reduces reward-estimation error and mitigates incorrect credit assignment.
latent reasoningcredit assignmenthierarchical optimizationadvantage-weighted learningmulti-answer estimation
Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled Reasoning
Problem-Space-Guided On-Policy Self-Distillation (PS-OPSD) improves reasoning by replacing complete reference solutions with trajectory-grounded guidance (initial state, goal conditions, constraints, and state-transition paths) during teacher-student distillation. This addresses the limitation of standard OPSD, where token-level targets may rely on inference-time-unavailable reference details. Evaluated across three mathematical reasoning benchmarks with models from 1.7B to 8B parameters, PS-OPSD achieves the highest aggregate question-only accuracy. Ablation studies confirm that guidance relevance and path coherence drive performance gains, emphasizing privileged information representation as a key design factor.
on-policy self-distillationprivileged informationtrajectory-grounded guidancemathematical reasoningstate-transition path
The Label Defines the Timescale: Trait-State Limits of Temporal-Aggregate Learning
The article demonstrates that performance ceilings in temporal-aggregate learning benchmarks may reflect acquisition-protocol limitations rather than model capacity, by analyzing labels derived from latent Gaussian processes with stable traits and within-individual states. Using an exact protocol-conditioned Bayes-risk identity, it decomposes label variance into $O(1)$ trait and $O(T^{-1})$ state components, derives task-dependent effective temporal spans, and shows state-driven occupation-label variance peaks when the stable trait is at the threshold. Experiments reveal that temporally dispersed observations improve state explainability more than repeated segments. The work distinguishes architectural from protocol limits, emphasizing the label-defined timescale.
temporal-aggregate learninggaussian processbayes-risktrait-state decompositioncorrelation time
Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard
The paper introduces F-ICL, a benchmark for evaluating in-context learning (ICL) in language models against an exact Bayes-optimal standard derived from a Turing-complete machine F. By exhaustively enumerating 1.5 billion programs (length ≤13) and computing closed-form posteriors under a Levin-Solomonoff prior, the study measures how closely model distributions approximate the optimum. Results show that 45 of 46 tested models (0.8B–675B parameters) deviate further from optimality than a keystroke reference, with performance gaps unexplained by scale, instruction tuning, or accuracy (Spearman ρ=−0.19). The benchmark reveals non-monotonic updating and a bias toward loop-free mixtures, releasing as an open toolkit.
in-context learningbayes-optimallevin-solomonoff priorturing-completealgorithmic reasoning
Enhancing Visual Perception in Foggy Conditions via Multiclass Fog Density Modeling
This work proposes a multiclass fog-density modeling approach to enhance visual perception for autonomous driving in foggy conditions. The method leverages synthetically generated fog data derived from the Waymo dataset, with depth images produced via iterative learning. Five fog-density levels (clear, light fog, moderate fog, heavy fog, very heavy fog) are considered, and separate perception models are trained for each level rather than a unified model. Experimental results demonstrate significant performance improvements in severe fog conditions, with recall for the very heavy fog class increasing from 0.076 to 0.232, a 15.6 percentage point gain. Future work will extend this strategy to LiDAR and radar modalities.
autonomous drivingfog-density modelingwaymo datasetiterative learningmulticlass perception
Does the Competitive Component of Adversarial Self-Play Improve Legal Reasoning? A Controlled Negative Result
The study investigates whether the competitive component of adversarial self-play improves legal reasoning, using a verifiable survival reward where citations are checked by a verifier. Through four independent tests and a follow-up pilot with a strengthened adversary, results showed no reliable benefit from competition (49-50% win rates, binomial p≈1.000). The null result confirms prior findings that verifiable environments, not competition itself, drive multi-teacher curricula value, while highlighting pitfalls like metric inversion and adversarial-robustness collapse.
adversarial self-playlegal reasoningverifiable rewardmulti-teacher curriculacitation verifier
Securing Agentic AI: From Per-Action Checks to Trajectory Assurance
The paper identifies security challenges in autonomous AI agents, particularly LLM-based systems, across single-agent and multi-agent deployments. It analyzes attack surfaces including prompt injection, memory manipulation, and tool interfaces, while highlighting emergent risks in delegation, capability control, and behavioral containment where permissible actions collectively violate system invariants. The work proposes treating security as a verifiable property of agent architectures and protocols, emphasizing the need for trajectory-level assurance beyond per-action checks.
agentic stackbehavioral containmentcapability controlmodel provenancesafety invariants
Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning
The paper introduces FedGD, a federated learning method for personalized reward modeling in LLMs that addresses preference heterogeneity without requiring group-specific initializations. By discovering latent preference groups during training and using group-debiased client sampling, FedGD learns a single adaptable reward model initialization. This approach counters group imbalance effects, enabling effective personalization even without prior knowledge of underlying groups. Results show that a single FedAvg model surpasses group-specific models after few local steps due to its flat initialization, which cancels conflicting preferences while maintaining discriminative power.
federated learningreward modelingpreference heterogeneitygroup-debiasingpersonalization
Emergence Invariance: From Symbolized Thought to Interface Refinement
The paper formalizes the Symbolization--Substructure Thesis, proposing emergence invariance as a framework to analyze how scaling and interface refinement affect reasoning gaps in language models. It proves that under fixed input laws, interface informativeness depends on σ-field refinement, with total compensation requiring vanishing interface floors and asymptotic gaps. Empirical validation using DeepSeek V4-Flash API shows performance jumps from 0/16 to 14/16 in pointer chasing when distinctions are available, while observational twins remain at 50%, aligning with theoretical predictions on interface-limited scaling.
emergence invariancesymbolization-substructure thesisinterface refinementσ-fieldcompensation gap
V-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic Memory
V-Mem introduces a modality-routed retrieval system for long-term multimodal agentic memory, addressing two key gaps in multimodal similarity search: the modality gap and the similarity-relevance gap. The system organizes conversations into rounds and retrieves target-modality content from the same round, avoiding cross-modality comparisons. It employs LLM-generated anchors, such as hypothetical captions or enriched search anchors combining text and keywords from images, to bridge these gaps. Evaluated on Mem-Gallery and LoCoMo, V-Mem achieves LLM-judge scores of 0.82 and 0.69 respectively, significantly outperforming baselines, particularly on image-carrying questions (0.87 vs. ≤0.47).
modality gapsimilarity-relevance gapmultimodal retrievalllm-generated anchorsagentic memory
Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics
The paper introduces Question-begets-Question (QbQ), a self-evolving curriculum for reinforcement fine-tuning that overcomes data scarcity and performance plateaus in competition mathematics (AIME). QbQ generates synthetic problem variants by seeding from a model's current capabilities, avoiding reliance on ground-truth reasoning traces. When applied to Qwen2.5-Math-7B (initial pass@1: 5.6%), static QbQ data improves performance to 14.5%, while the dynamic curriculum breaks this ceiling, achieving 16.5% pass@1 after 20 rounds without saturation. Notably, training on variants of near-solved problems generalizes to harder unseen problems.
reinforcement fine-tuningself-evolving curriculumsynthetic data generationcompetition mathematicsqwen2.5-math-7b
MineGrad: Gradient Inversion Attacks on LoRA Fine-Tuning
We introduce MineGrad, a gradient inversion attack targeting low-rank adaptation (LoRA) fine-tuning in federated learning, enabling malicious servers to recover private user data. The attack leverages a poisoned pretrained model and fine-tuning parameters, embedding fine-tuning data within shared gradients for analytical reconstruction. Unlike prior methods, MineGrad is applicable to both language and vision tasks, avoids computationally expensive adversarial pretraining, and does not constrain the number of training tokens relative to LoRA rank. Experiments demonstrate high-fidelity data recovery across multiple baselines, exposing critical vulnerabilities in federated fine-tuning protocols.
gradient inversionlow-rank adaptationfederated learningfine-tuningprivacy
Deep Agentic Search for Repository-Level Code Question Answering: An Empirical Study
This empirical study compares Semantic Search and Deep Agentic Search approaches for repository-level code question answering, evaluating their performance on the SWE-QA benchmark. Semantic Search achieved 65.2% accuracy versus 46.2% for Deep Agentic Search, with lower computational cost. Failure mode analysis revealed that 41.8% of Deep Agentic Search errors occurred during planner-subagent handoff, often producing confident but incorrect answers. The results suggest that while Deep Agentic Search addresses context pollution, retrieval-based methods remain more effective for read-only repository queries.
semantic searchdeep agentic searchcontext pollutionrepository-level qafailure modes
Rapid Embodiment Adaptation for Quadrupedal Locomotion
The authors present an online embodiment adaptation framework for quadrupedal locomotion that infers hardware parameters from brief interaction histories and conditions control on these estimates. The method combines a generalist policy trained with embodiment randomization and a lightweight adaptation module that identifies physical changes within 500ms. Evaluated on joint-range constraints and trunk-mass variations, the module achieves accurate parameter estimation and enables robust closed-loop control. In simulation and on a Unitree Go2 robot, the system maintains stable locomotion under severe changes (e.g., locked leg, 5kg payload), outperforming non-adaptive baselines.
quadrupedal locomotionembodiment adaptationonline parameter estimationhardware degradationclosed-loop control
Sweet Little Lies: Strategic Deception in AI Emotional Support Chatbots
The paper introduces a Bayesian Persuasion model to analyze strategic deception in AI emotional support chatbots, demonstrating economic incentives for misrepresentation. Interactions are modeled between chatbots signaling users' emotional states and users deciding to engage based on these signals. Equilibrium analysis reveals chatbots optimally report truthfully when users genuinely need support but misreport emotional need in good emotional states. This deception increases engagement without reducing users' expected payoff, with more skeptical users receiving honest assessments due to higher engagement thresholds. The findings raise ethical and regulatory concerns despite no payoff reduction.
bayesian persuasionemotional support chatbotsequilibrium analysisengagement metricsstrategic deception
VGER: Voxel-Guided Global Event Ranking for Event Cloud Attribution
The paper introduces Voxel-Guided Global Event Ranking (VGER), a training-free attribution framework for point-based event cloud networks that improves model interpretability. VGER combines event-level gradients with task-aware voxel perturbations to generate fine-grained attribution scores while preserving spatio-temporal event structures. Evaluated on three benchmarks (PointNet, PointNet++, EventMamba) across nine settings, VGER outperforms point-level saliency baselines in both high-tail and low-tail deletion metrics.
event camerasattribution frameworkspatio-temporal structuresvoxel perturbationevent cloud networks
Retrieval Augmented Biomedical Question Answering with Weak Question Recovery and Neural Reranking for BioASQ Task 14b
The DS@GT ARC BioASQ team proposes a biomedical QA pipeline integrating multi-source query expansion, neural reranking, and OpenBioLLM-assisted generation. Their method combines PubMed retrieval with MiniLM-based semantic reranking, Reciprocal Rank Fusion, and feature-based scoring, augmented by a conditional weak-question recovery strategy for challenging queries. Post-retrieval pruning removes redundant snippets while preserving evidence. On BioASQ batches, the system improves retrieval robustness and achieves MAP@10 gains on difficult questions, with additional output validation for submission reliability.
biomedical question answeringneural rerankingreciprocal rank fusionquery expansionretrieval refinement
Computing with Agentic Oracles
The paper extends the stochastic-oracle model to agentic oracles, which autonomously pursue goals and access task-relevant environments, affecting response distributions and token costs. It introduces a framework for analyzing token costs in Stochastic-Oracle Turing Machines (SOTMs), distinguishing between orchestration (visible) and agentic (internal) token costs. Results show agentic oracles with state retention reduce token costs versus stationary oracles, while goal-loss risk analysis provides avoidance criteria, progress formulas, and complexity bounds, including cases where goal-loss probability is zero.
agentic oraclestochastic-oracle turing machinetoken complexitygoal-loss riskorchestration cost
Where Reasoning Diverges: Localized Multi-Agent Debate
The paper introduces Localized Multi-Agent Debate (LMAD), an inference-time protocol that improves multi-agent reasoning by localizing debates to the earliest conflicting claims. LMAD represents agent traces as typed nodes, restricts debate to local segments around conflicts, and uses guarded resolution to extend a shared committed state without reopening accepted steps. Evaluated on four multi-hop QA benchmarks with ten backbones from four model families, LMAD achieves the highest macro-averaged judge accuracy, outperforming conventional baselines by up to 7.20 percentage points.
multi-agent debatereasoning tracesinference-time protocolguarded resolutionmacro-averaged accuracy
Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
We introduce a model-agnostic framework for per-modality failure analysis in multimodal clinical AI, addressing two key deployment challenges: identifying responsible modalities for accuracy loss and distinguishing loud (monitorable) from silent (unflagged) failures when modalities are dropped. The framework employs modality embeddings, mask-aware probes, and labels to generate per-example failure taxonomies, per-modality complementarity matrices, and loud-vs-silent dropout profiles. Validated against planted ground truth, it accurately recovers modality dominance and complementary subsets, scales to three-modality matrices, and reports per-modality failure rates. Applied to EchoJEPA and HuBERT-ECG embeddings for LVEF prediction on a MIMIC-IV cohort (n=245), dropping echocardiograms nearly doubled error, highlighting deployment implications for cardiac foundation models.
multimodal clinical aimodality embeddingsfailure taxonomycomplementarity matrixmask-aware probe
Beyond Routing Saturation: A Long-Horizon Class-Incremental Perspective on Expert Routing in Multimodal Continual Instruction Tuning
The paper introduces FLEX, a 34-task benchmark for Multimodal Continual Instruction Tuning (MCIT) that reduces textual fingerprints to expose long-horizon expert routing challenges. It formulates routing as Multimodal Class-Incremental Learning (MCIL), treating each task as a class whose score distribution determines LoRA mixture weights. The authors adapt four Class-Incremental Learning methods to MCIT frameworks, improving strict LoRA matching by up to 16.3 percentage points and MacroScore by 4.6 points without modifying LoRA experts or generation pipelines.
multimodal continual instruction tuningexpert routingclass-incremental learninglora expertstask identification
Same violence, different answer: how AI responds to coercive control against women across languages
The study evaluates how conversational AI systems respond to disclosures of coercive control against women across languages, identifying systematic failures in recognition and response. Researchers tested seven widely-used language models across nine languages using a scripted scenario involving phone surveillance and self-blame. Results revealed two failure axes: non-anglophone developers' models performed worst in their native languages, and partner-excusing narratives inconsistently eroded control-naming across languages, though two frontier systems maintained strict standards universally.
coercive controlconversational aimultilingual evaluationintimate partner violencelanguage model bias
onepot-Bench 0: towards lab-aware in silico chemistry benchmarks
The authors introduce onepot-Bench 0, a proprietary benchmark suite for evaluating language models in synthetic chemistry tasks relevant to wet-lab execution. The suite comprises three evaluations: ChemAbacus (cheminformatics literacy and numerical reasoning), SynthRefusal (safety and refusal behavior), and SynthBench (reaction-outcome prediction and catalyst selection using private experimental data). These evaluations measure basic competency, reliability, and domain-specific knowledge required for reliable laboratory performance, addressing gaps in existing benchmarks that often rely on public data potentially present in training corpora.
language modelssynthetic chemistrycheminformaticsbenchmark suitewet-lab execution
The Condition-Number Barrier in Sparse Least Squares
The paper establishes a lower bound for sparse least-squares optimization, confirming a conjecture by Axiotis and Sviridenko [AS21] under the Small-Set Expansion Hypothesis [RST12]. It shows that no randomized polynomial-time algorithm can achieve a solution with sparsity $s = O(k\kappa_{s+k}^{1-\gamma})$ while approximating the optimal error within $\varepsilon$, where $\kappa_r$ is the restricted condition number at sparsity $r$. The proof, initially automated via a Gemini-based system, holds even for full-rank rational matrices.
sparse optimizationleast squarescondition numbersmall-set expansionpolynomial-time algorithms
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning
GradCuit introduces credit-assigned gradient flow to enable robust and interpretable test-time latent reasoning in large language models by optimizing instance-specific continuous states at a selected Transformer layer. The method leverages causal self-attention to provide differentiable paths from continuation-token log-probabilities to preceding latent states, enabling direct reward-weighted gradient assignment. Across five instruction-tuned backbones and three reasoning benchmarks, GradCuit achieves 64.5% average accuracy, outperforming chain-of-thought prompting by 6.6 percentage points and reducing accuracy variance. Token-level gradient attribution reveals latent influence on reasoning-connector tokens, with early-to-middle Transformer layers identified as optimal for optimization.
gradient flowtest-time adaptationtransformer layerscredit assignmentlatent reasoning
Smooth Reparameterizations of Functions on Simplicial Product Spaces: Applications to Probabilistic Tensor Decomposition and Functional Data Registration
The paper introduces a smooth, elementwise strictly convex reparameterization for optimization problems on product spaces of simplices, applicable to probabilistic tensor decomposition and functional data registration under the Square Root Velocity Function (SRVF) representation. The method transforms constrained problems into unconstrained ones on a manifold, ensuring weak second-order Karush-Kuhn-Tucker (KKT) points on the simplex map to second-order KKT points on the manifold. A Riemannian Gradient Descent (RGD) algorithm outperforms Projected Gradient Descent (PGD) in curve registration, preserving original function shapes more accurately.
reparameterizationproduct simplexkarush-kuhn-tuckerriemannian gradient descentcurve registration
Pseudorandom Streams within Diffusion Models Act as Learnable Inputs That Affect Generation Quality
The study demonstrates that pseudorandom streams in diffusion models function as learnable inputs influencing generation quality, challenging their treatment as purely stochastic elements. Using a multilayer perceptron to predict pseudorandom sequences and a diffusion probe to evaluate orbit structure exploitation, the authors analyze MNIST and CIFAR-10 datasets. Results reveal significant correlations between pseudorandom orbit structure, diffusion loss, and generation degradation, with empirical power law relationships observed post-normalization. This indicates pseudorandom sources act as model-dependent structured inputs, not merely distributional choices.
diffusion modelspseudorandom streamsmultilayer perceptrongeneration qualityempirical power law
Benchmarking Sheaf Neural Networks for Inductive Tasks
This work presents the first systematic benchmark of Sheaf Neural Networks (SNNs) for inductive tasks, evaluating diffusion mechanisms, restriction-map parameterizations, stalk dimensions, and architectural components across 1,890 experiments on 14 datasets. The study reformulates SNNs via message passing to avoid assembling the sheaf Laplacian, enabling cross-graph batching. Key findings indicate restriction maps dominate performance, general maps outperform alternatives, and architectural components outweigh sheaf-specific choices in impact. While SNNs transfer to inductive settings, they underperform top baselines in a dataset-dependent manner, with a single configuration proving broadly applicable.
sheaf neural networksinductive learningrestriction mapsmessage passinggraph attention networks
A Simple Approximation to the Distribution of the Ridge Regression Estimator
The authors propose a Gaussian approximation for the finite-sample distribution of the ridge regression estimator, capturing its bias-variance trade-off. The method employs nonstandard asymptotics, scaling the regularization parameter with sample size and treating population coefficients as local to the shrinkage direction. Unlike prior work, it accommodates heteroskedasticity and autocorrelation but restricts covariate dimensionality. The approximation informs two novel regularization strategies minimizing average or worst-case excess prediction risk.
ridge regressiongaussian approximationnonstandard asymptoticsheteroskedasticityregularization parameter
Interaction Is Not Necessary for Order-Optimal 1-Bit Mean Estimation
The paper presents a randomized non-adaptive protocol for one-bit mean estimation that matches the optimal adaptive sample complexity, resolving an open problem from COLT 2026. For distributions on ℝ with mean in [-λ,λ] and bounded k-th central moment (k>1), the protocol fixes all queries beforehand, eliminating the need for interaction. The sample complexity scales optimally as log(λ/σ) plus terms depending on k, ε, and δ, matching known lower bounds for fully adaptive protocols.
one-bit mean estimationnon-adaptive protocolsample complexityminimax optimalitycentral moment
Optimal Unambiguous DNFs and Alon-Saks-Seymour
The work constructs unambiguous DNFs with width O(n) and 0-certificate complexity Ω(n²), leveraging their structure to prove a lifting theorem with constant-sized gadgets. This theorem translates certificate complexity separations to communication complexity, yielding optimal refutations of the Alon-Saks-Seymour conjecture and improved lower bounds for the Clique versus Independent Set problem, surpassing prior results by doubly logarithmic factors. Applications include a quartic separation between certificate complexity and approximate degree, and a sample compression lower bound of Ω(√log c) for multiclass concept classes.
dnfcertificate complexitycommunication complexitylifting theoremsample compression
Uncertainty Is Not Enough: Value-of-Information Routing for Mixtures of LoRA Experts
VI-MoLE introduces a certified value-of-information routing framework for mixtures of LoRA experts, addressing the limitation of uncertainty-based routers that conflate ambiguity with recoverable risk. The method learns counterfactual risk remaining after each expert prefix, converts predictions into simultaneous upper-risk certificates on calibration data, and allocates a global adapter budget to token-layer actions with the largest certified marginal risk reduction per unit cost. Theoretical guarantees include simultaneous certificate validity, optimal greedy allocation under diminishing certified gains, and allocation regret under value-estimation error. Evaluations demonstrate improvements in matched-compute accuracy, certificate coverage, risk-coverage, distribution shift robustness, and tail latency compared to fixed and dynamic MoE-LoRA routers.
value-of-informationlora expertsrisk certificatesgreedy allocationcounterfactual risk
LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference
LiveMem introduces state continuity under context turnover for long-running LLM inference, enabling persistent memory state independent of active context bounds. The method augments pretrained full-attention LLMs with an intrinsic memory state, maintaining historical information while the main attention path operates within a bounded KV window. Experiments on LongMemEval demonstrate LiveMem's ability to answer questions based on memory state after supporting evidence is removed from context, outperforming other intrinsic memory methods. Evidence-distance analysis confirms persistent information retention beyond the active window.
state continuitycontext turnoverintrinsic memorykv windowlongmemeval
RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States
RoMeRL introduces reduced-order utility states to address two challenges in self-evolving LLM agent memory systems: feedback dispersion over expanding state spaces and the memory-reward trap from co-retrieved memories. The method factorizes trajectory-indexed utilities into fixed-dimensional per-task states by outcome polarity and dynamics, updating semantic coordinates to concentrate feedback. Evaluated on ALFWorld and LifelongAgentBench, RoMeRL improves task performance by 80.0% in Cold-Q ratio, 6.0× feedback density, 84.4% memory reduction, and 21.1% fewer LLM calls, demonstrating efficient memory evolution with reduced reward contamination.
self-evolving agentsmemory-reward trapreduced-order utilityfeedback densitytrajectory-indexed utilities
Beyond Modern Asymptotics for Log-Likelihood Ratios in Logistic Regression
This work provides nonasymptotic characterizations of the log-likelihood ratio statistic in binary logistic regression, uniformly over design vectors and target parameters. For $n\geq d\geq 3$, the worst-case $(1-\delta)$ quantile is established as $d\log\left(\frac{e n}{d}\right)+\log\left(\frac{1}{\delta}\right)$, generalizing the Wilks $\chi^2_d$ phenomenon without design regularity assumptions. Low-dimensional cases exhibit distinct behavior: $d=2$ scales as $\log\log\log n+\log\left(\frac{1}{\delta}\right)$, while $d=1$ depends solely on $\log(1/\delta)$. Gaussian designs recover classical Wilks scaling, and for $n\gtrsim d+\log(1/\delta)$, the sharp bound $d+\log\left(\frac{1}{\delta}\right)$ is proven.
log-likelihood ratiobinary logistic regressionwilks phenomenonnonasymptotic boundsquantile characterization
Computational and Statistical Guarantees of the \textit{c}-Rectified flow
(No summary returned.)
Cultural Awareness is Represented but Not Decoded: Tracing Mythological Knowledge across 18 Open-Source LLMs
The study investigates cultural bias in open-source LLMs by analyzing mythological knowledge representation across 18 models from 8 architecture families. Using linear probing, logit lens, activation patching, and output extraction on a parallel cross-cultural dataset of Thompson-motif entities, the authors find that while residual streams distinguish cultural representations effectively, decoders collapse culturally-specific tokens onto dominant traditions. Failures occur at readout, not representation, with prompt language gating decoder behavior. The work releases a framework for probe-output decomposition, cross-cultural ground truth, language-conditioned readout tests, and per-entity predictions.
linear probinglogit lensactivation patchingresidual streamlanguage gating
Private Generative Bootstrap via Blocking
The paper introduces Private Generative Bayesian Bootstrap (PGBB), a differentially private method for Bayesian uncertainty quantification without requiring a specified data-generating model. PGBB employs a blocking strategy that groups individuals and assigns shared weights, enhancing privacy by obscuring individual contributions. It uses amortized inference to decouple private learning from posterior sampling, enabling efficient private posterior draws. Theoretical analysis covers differential privacy guarantees, convergence to non-private targets, and asymptotic posterior dispersion restoration. Empirical evaluations on U.S. Census and natality data demonstrate competitive performance against private Bayesian alternatives requiring explicit generative models.
differential privacybayesian bootstrapamortized inferenceposterior samplingblocking strategy
Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees
We introduce Aggregate-then-Calibrate (AtC), a two-stage framework for human-centered assessment tasks lacking verifiable ground truth. Stage-1 aggregates heterogeneous comparative judgments into a consensus ranking using a rank-aggregation model that accounts for annotator reliability. Stage-2 calibrates predictive model scores via isotonic projection onto the consensus order, preserving quantitative information while enforcing ordinal consistency. Theoretical analysis shows AtC yields more efficient consensus estimation, enjoys risk bounds under misspecification, and asymptotically outperforms model-only assessment. Empirical results across semi-synthetic and real-world datasets demonstrate AtC improves accuracy and robustness over human-only or model-only approaches.
rank-aggregationisotonic calibrationconsensus rankingannotator reliabilityordinal consistency
Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search
This work introduces a vision-language model (VLM)-based pipeline for automated relevance evaluation in Pinterest Search, addressing scalability limitations of human annotation. The method validates alignment between VLM-generated judgments and human annotations, demonstrating reliability in relevance measurement. By leveraging VLM-based labeling, the approach enables expanded query sets, optimized sampling designs, and efficient assessment of diverse search experiences at scale. Results show improved relevance metrics and significantly reduced Minimum Detectable Effects (MDEs) in online A/B experiments, enhancing evaluation efficiency.
vision-language modelrelevance evaluationminimum detectable effectsa/b experimentssampling design
Intention Inference Under Execution Noise: Separating Aleatoric and Epistemic Uncertainty in Social Dilemmas
The paper proposes a Partially Observable MDP (POMDP) formulation to distinguish between aleatoric (execution noise) and epistemic (intent uncertainty) in noisy social dilemmas, addressing systematic over-retaliation in standard MDPs. Using active inference (AIF) with a decomposed cost function, the method infers latent opponent intentions from noisy actions and models intent dynamics. In the Iterated Prisoner's Dilemma with symmetric noise, results show a critical noise threshold for cooperation collapse, with POMDP outperforming MDP against conditionally cooperative opponents but exhibiting belief-driven collapse under high noise when intent attribution is decision-relevant.
partially observable mdpactive inferencealeatoric uncertaintyepistemic uncertaintyiterated prisoner's dilemma
Foundations of Reinforcement Learning and Control:Connections and New Perspectives
The paper bridges reinforcement learning (RL) and control theory by introducing adaptive control, actor-critic RL algorithms, and a novel integration of these paradigms for data-driven decision-making in locomotion control. It contrasts the methodologies, goals, and cultural differences between RL and control theory, both rooted in dynamic programming but evolved separately. The tutorial provides foundational insights to facilitate cross-disciplinary understanding and collaboration, enabling experts in each field to leverage tools and approaches from the other.
reinforcement learningcontrol theoryadaptive controlactor-criticdynamic programming
Wasserstein mixing time of the unadjusted Langevin algorithm
(No summary returned.)
Why Large Language Models Fail at Tabular Prediction
The study identifies dimensionality as the primary factor explaining why large language models (LLMs) underperform on tabular prediction tasks compared to classical methods. Through controlled experiments on 31 benchmark datasets, the authors systematically evaluate five hypotheses, falsifying four (noise sensitivity, CSV format, numeric tokenization, and test point count). Results show that LLM accuracy decreases with increasing dimensionality, unlike classical baselines which remain stable or improve. Behavioral analysis reveals that LLMs resemble local, distance-based methods in low dimensions (91.6% grid agreement) but exhibit unique prediction patterns in higher dimensions that cannot be replicated by noise-augmented classical models.
dimensionalitytabular predictionlarge language modelsclassical baselinesnoise sensitivity
Deep Learning-Based Estimation of Ground Reaction Forces in Parkinsonian Gait Using an Optimized Set of IMU Data
A hybrid CNN-BiLSTM model was developed to estimate bilateral vertical ground reaction forces (vGRFs) in Parkinsonian gait using an optimized set of inertial measurement units (IMUs). The model, trained on data from 61 Parkinson's disease (PD) patients and 65 healthy controls (HC) with 13 IMUs, achieved high intra-subject accuracy (R² = 0.98) and strong inter-subject generalization (R² = 0.93 for HC, R² = 0.91 for PD). Optimal sensor placement varied between PD patients and HC, with a minimal setup of two IMUs enabling robust estimation. This approach supports wearable vGRF-based gait analysis systems for PD and other pathological conditions, facilitating accessible clinical assessments and personalized rehabilitation.
cnn-bilstmground reaction forcesinertial measurement unitsparkinsonian gaitsensor optimization
Network Information Enhances Unreliable News Domain Detection
The study demonstrates that network structure enhances unreliable news domain detection, showing assortative mixing where low- and high-reliability domains cluster separately in URL-sharing networks. Using Telegram chat data, the authors construct a statistically validated domain co-sharing network and evaluate Graph Neural Networks (GNNs) against network-unaware baselines. GraphSAGE outperforms Multi-Layer Perceptrons, achieving 0.63 accuracy with content-aware features (multilingual text embeddings) and 0.53 with content-agnostic features (spreading dynamics), yielding 13-14% relative improvements. Network topology thus provides systematic gains in reliability classification, even without content analysis.
graph neural networksdomain reliabilityassortative mixingmultilingual embeddingsspreading dynamics
Gecko: Fast Private Inference via Secure Public Encoder Offloading
Gecko introduces a secure framework for fast private inference by offloading a public encoder outside the protection boundary while maintaining security against feature-space shortcuts. The method employs a frozen backbone for hierarchical feature extraction, fixed Fastfood projections for feature compression, and private feature gating for prediction preparation. It formalizes ideal independence and information-preservation conditions to guide design and evaluates component-reuse extraction attacks. Gecko achieves 0.4-2.2 second inference with ≤10.8 MB communication and accuracy comparable to transfer-learning baselines across image and audio tasks, demonstrating resilience against model-extraction adversaries.
private inferencepublic encoderfastfood projectionsfeature gatingmodel-extraction
A Spectral Filtering Approach to Regret Analysis of Distributed Online Control for Linear Dynamical Systems
The paper extends the Online Spectral Control framework to distributed settings for linear time-invariant systems with adversarial disturbances and time-varying convex costs. Agents generate control sequences using local observations and neighbor communication, applying spectral controllers convolved with past disturbances via Hankel matrix eigenvectors, with parameters updated through distributed online gradient descent. A sublinear regret bound of $O(\frac{\sqrt{T}\text{poly}(\log T)}{γ^3})$ is established, where $T$ is the time horizon and $γ$ the stability margin, capturing network size and connectivity dependencies.
distributed online controllinear dynamical systemsregret minimizationspectral parameterizationhankel matrix
Qwen-CUA: Native Computer Use for (almost) Everything
Qwen-CUA introduces a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone, operating solely via screenshots and keyboard/mouse events without task-specific APIs. The method employs a scaffold maintaining 20 active screenshots with fixed-size history blocks, trained on 40,000 verifiable tasks using cloud rollout fleets (100,000 vCPUs) and hybrid optimization (supervised data refresh, reinforcement-learning recalibration). Qwen-CUA achieves 86.2 on OSWorld-Verified and 18.5/48.4 binary/partial completion on OSWorld 2.0, outperforming Qwen3.7, while Qwen-CUA-Max (1T+ params) reaches 87.6 and 21.2/53.3. It also reduces RedTeamCUA attack success from 36.6 to 16.4.
mixture-of-expertsnative computer-useverifiable taskstrajectory slicingscalable interaction
Self-Supervised Representations for Binary Program Clustering: From Empirical Study to Retrieval-Augmented Learning
This study conducts the first systematic investigation of self-supervised learning (SSL) and tabular representation learning (TRL) methods for binary program clustering, focusing on malware analysis. Phase 1 adapts vision-based SSL models (BYOL, SimSiam, Barlow Twins, VICReg) to tabular data, revealing BYOL and SimSiam perform comparably to supervised models. Phase 2 evaluates unsupervised TRL methods, identifying VIME as state-of-the-art. The authors propose VIME-R, a retrieval-augmented extension of VIME, which improves Homogeneity by 2.7%-5.8% on Ember and Bodmas datasets. Results demonstrate retrieval-augmented TRL as a promising approach for enhancing automated malware clustering.
self-supervised learningtabular representation learningbinary program clusteringretrieval-augmented learningmalware analysis
The Push-Forward Transform for Continuous and Robust Comparison of Dynamic Shapes
The authors introduce the Push-Forward Transform (PF-T), a mathematical framework for invariant and robust shape comparison by mapping functions from the shape domain to a common reference domain. PF-T applied to Signed Distance Functions (SDFs) yields a continuous representation capturing boundary and interior geometry, enabling quantification of shape similarity and revealing skeletal topology and rotational symmetries. The method extends to 2D, 3D, and temporal geometries, supporting joint analysis of shape and scalar fields. An efficient algorithm is described, and benchmarks demonstrate its effectiveness on diverse datasets.
push-forward transformsigned distance functionsshape comparisonskeletal topologymorphometric
Extended Field of View Analysis for VideoGAN-based Trajectory Generation
The paper extends VideoGAN-based traffic trajectory generation by improving semantic representations, replacing trajectory extraction with graph-based association, and scaling to larger fields of view. It introduces an evaluation framework for hallucinations and object permanence in generated videos. Experiments show the approach generalizes to complex scenes (up to 20s duration) while maintaining realistic trajectories and spatial coherence, achieving inference times <20ms after 150 GPU-hours training, demonstrating scalability for automated driving applications.
trajectory generationgenerative adversarial networksbird's-eye-viewobject permanenceautomated driving
A Multi-Objective AutoML-based Efficient Intrusion Detection System for EV Charging Networks
A Multi-Objective Automated Machine Learning (MOO-AutoML) framework is proposed for efficient intrusion detection in Electric Vehicle Charging Systems (EVCS). The method employs a LightGBM-based automated feature selection strategy and optimizes feature selection thresholds and hyperparameters using Non-dominated Sorting Genetic Algorithm III (NSGA-III), targeting three objectives: maximizing weighted F1-score, minimizing 99th percentile inference latency, and minimizing model size. Evaluations on CICEVSE2024 and CICIDS2017 datasets demonstrate that the framework achieves competitive weighted F1-scores, reduced inference latency, and smaller model sizes compared to existing methods, enabling practical deployment in EVCS and IoT security.
automated machine learningintrusion detectionelectric vehicle charginglightgbmnsga-iii
Z-PEFT: Zero-shot Backdoor Detection in Parameter-Efficient Fine-Tuning via Canonical Spectral Signatures
Z-PEFT introduces a zero-shot backdoor detection method for Parameter-Efficient Fine-Tuning (PEFT) models, addressing the vulnerability of models downloaded from open repositories to weight-space backdoor attacks. The approach employs a lightweight meta-classifier that utilizes layer-wise spectral measures for classification, enabling detection of previously unseen attacks and datasets without requiring prior knowledge of specific attack types. Experimental results demonstrate that Z-PEFT outperforms existing weight-space detectors in zero-shot scenarios while maintaining low computational overhead. This method provides a scalable safety mechanism for practitioners deploying PEFT models in open-source environments.
parameter-efficient fine-tuningzero-shot detectionweight-space backdoorspectral measuresmeta-classifier
Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study
The paper introduces a reproducible, multi-metric benchmarking framework for evaluating text-to-speech (TTS) systems across diverse speech domains, with a focus on low-resource languages. The framework combines subjective (MUSHRA, ABX tests) and objective (Resemblyzer, MCD, F0 RMSE) metrics, demonstrated through a case study evaluating four TTS systems (Indic-Parler-TTS, MMS-TTS, Microsoft Edge TTS, Google Gemini TTS) across four domains. Results show domain-dependent performance, with emotional speech being most challenging (mean MCD 12.03 dB, F0 RMSE 889 cents) and conversational speech achieving highest fidelity.
text-to-speechmushramel-cepstral distortionresemblyzerlow-resource languages
Constrained Co-Design for Photonic Bayesian Neural Networks
The work introduces a constrained co-design framework for photonic Bayesian neural networks (BNNs) to address hardware limitations in stochastic variational inference. By formulating photonic BNN inference under analog constraints (quantization, programming error, dynamic range), the study systematically ablates stochasticity location, modality, and representational bounds to derive hardware-software co-design guidelines. Experiments on Dirty-MNIST, CIFAR-10, and CINIC-10 (with Fashion-MNIST/SVHN for OOD evaluation) show that training compensates for constraints when variational families remain representable, while hardware modifications are needed for fundamental representational limits.
bayesian neural networksphotonic computingstochastic variational inferencehardware constraintsout-of-distribution detection
CRIP: Channel Level Representation Injection for Personalized One-Shot Federated Learning
CRIP introduces a personalized one-shot federated learning framework addressing domain heterogeneity through channel-level representation injection. The method involves clients uploading feature extractors to a server, which broadcasts them back; CRIP then selectively fuses features by measuring channel-wise representational similarity on local mini-batches to mitigate domain-specific noise. Experiments on DomainNet, PACS, and Office-Home show CRIP outperforms local models and state-of-the-art baselines under extreme domain heterogeneity.
federated learningdomain heterogeneityfeature alignmentrepresentation spaceone-shot learning
Start Classifying: Categorical Critics for LLM Reinforcement Learning
The paper introduces HL-Gauss PPO, a categorical critic for Proximal Policy Optimization (PPO) that replaces scalar mean-squared-error regression with cross-entropy classification over discretized value targets. The method maintains standard advantage computation while improving calibration and reducing variance in sparse-reward settings like reinforcement learning with verifiable rewards (RLVR). Evaluations on mathematical reasoning, tool-augmented math, and Search-R1 tasks using Qwen2.5 and Qwen3 backbones show consistent gains over PPO and DAPO baselines, with improved Brier scores and advantage symmetry.
proximal policy optimizationcategorical criticreinforcement learning with verifiable rewardsbrier scoreadvantage calibration
Randomized Algorithms for Learning Partitions with Near Optimal Query Complexity in Constant Rounds
The paper presents randomized algorithms for learning hidden partitions via PAIR queries, demonstrating significant improvements over deterministic approaches. Using a 3-round randomized algorithm with known partition count $k$, it achieves $O(nk\log n)$ query complexity (high probability), contrasting with $Ω(n^{4/3}k^{2/3})$ for 2-round methods. For unknown $k$, a 4-round algorithm requires $O(n|\mathcal P|\log^2 n)$ queries, while deterministic methods need $Θ(\log n/\log\log n)$ rounds for near-optimality. Results highlight randomization's advantage in reducing round complexity while maintaining query efficiency.
randomized algorithmsquery complexityhidden partitionsround complexitypair queries
CARNet: Channel-Adaptive Receiver Network for Robust NextG Communications
The paper introduces CARNet, a channel-adaptive neural receiver network for NextG communications, addressing generalization challenges in diverse channel conditions. CARNet leverages a mixture-of-experts (MoE) framework, employing multiple ResNet-based expert networks specialized for specific channel conditions and a lightweight routing mechanism. The routing mechanism projects coarse channel estimates into low-dimensional embeddings to guide expert selection. Link-level simulations demonstrate CARNet's superior performance across varied channel scenarios.
channel-adaptiveneural receivermixture-of-expertsresnetsignal detection
Empowering Credit Risk Detection in Weixin Pay with Billion-Scale Deep Graph Learning
A risk-aware overlapping subgraph learning framework is proposed for scalable credit risk detection in Weixin Pay. The method constructs base partitions for load balancing, performs budget-constrained sampling to preserve critical risk diffusion patterns, and introduces a cross-subgraph consistency alignment mechanism to harmonize local representations into a globally consistent latent space. Experiments on Weixin Pay's production dataset demonstrate significant improvements over existing strategies, offering a scalable solution for industrial graph learning.
graph neural networkscredit risk detectionsubgraph learningrepresentation alignmentrisk diffusion
Cardiovascular Digital Twins from Physics Based to Data Driven Approaches
The article reviews computational approaches for developing cardiovascular digital twins, focusing on patient-specific models for diagnosis, prognosis, and therapy optimization. Mechanistic models offer physiological interpretability but are computationally intensive, while data-driven methods enhance scalability at the cost of robustness. Emerging techniques integrate physics-informed, graph-based, and hybrid methods to combine physical constraints with relational learning across vascular networks. The review covers modeling paradigms, data assimilation frameworks, validation challenges, and translational pathways toward clinical deployment.
cardiovascular digital twinsmechanistic modelsdata-driven approachesphysics-informed methodsgraph-based learning
CoRe-GNN: Multilevel Message passing on Coarsened graphs
CoRe-GNN introduces a multilevel message passing framework combining graph coarsening and Cluster-GCN to address scalability and long-range information propagation in Graph Neural Networks. The method performs parallel propagations: a coarsened inter-cluster term for long-range structure and a local intra-cluster term for node discriminability. CoRe-GNN inherits spectral approximation guarantees from graph coarsening and employs a cluster-based batching scheme for memory efficiency. Evaluations on node classification benchmarks demonstrate superior performance over graph coarsening and Cluster-GCN, particularly in long-range tasks, while maintaining scalability on graphs with millions of nodes.
graph neural networksgraph coarseningmessage passingnode classificationspectral approximation
Do Static Embeddings Add Value to Hybrid Dutch Retrieval?
The study evaluates whether static embeddings provide complementary ranking information in hybrid Dutch retrieval when combined with lexical (BM25) and transformer-based (Qwen3-Embedding-0.6B) retrievers. Using weighted reciprocal rank fusion (RRF) on five MTEB-NL datasets (14,500 queries, 786,573 documents), the authors conduct exhaustive scoring, simplex weight search (0.1 increments), and 10-fold cross-validation with bootstrap confidence intervals. Results show fusion improves MRR over individual retrievers (e.g., +0.061 on Dutch News), but no fold assigns weight to static embeddings, favoring a BM25-Qwen hybrid. Cross-domain validation confirms equal BM25-Qwen weighting as optimal.
hybrid retrievalreciprocal rank fusionstatic embeddingsbm25cross-validation
From Information to Delegation: Mapping Human-AI Financial Decision Making
The study introduces a behavioral measurement framework combining intent and delegated decision authority to quantify human-AI interaction dynamics in financial decision-making. Analyzing 1.5 million ChatGPT and Gemini interactions from 6,304 users in the US and India, the authors find financial services constitute a major AI use case, primarily for information retrieval and judgment formation rather than execution delegation. Results establish a baseline for tracking the shift toward more agentic AI systems in financial contexts.
behavioral measurementdecision authorityinformation retrievalagentic aifinancial judgment
One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse
The study identifies a shared failure channel for diverse low-precision errors in bfloat16 transformers, where abrupt collapse stems from query-key (QK) spectral runaway regardless of fault source. By isolating the failure to the streaming-softmax accumulator and demonstrating that fp32 accumulation repairs it, the authors show QK-channel dominance via causal probes projecting updates off leading singular directions. Their QK-Guard intervention—a parameter-free normalization triggered by attention-logit saturation—contains runaways across architectures and scales, matching always-on normalization over 60k steps while non-QK corrections fail. Results indicate interventions should target the QK channel rather than individual fault sources.
low-precisionattention collapsequery-key channelspectral runawaystreaming-softmax
Pretraining on Call Graphs: When Binary Analysis Tasks Profit From Context
The study evaluates call graph-enhanced binary function embeddings, demonstrating that improvements in binary code similarity detection (BCSD) do not consistently generalize to downstream semantic or syntactic tasks. Using graph-based models on embeddings from two state-of-the-art binary function embedding models, the authors show that optimizing for semantic similarity degrades syntactic task performance. Explanatory analysis reveals call graph context enhances embedding robustness, particularly for namespace-related functions, with maximal benefits in context-dependent scenarios.
binary function embeddingcall graphbinary code similarity detectionsemantic similaritysyntactic tasks
A 2-Block Architecture for Real-Time EEG Gait Decoding: A Pilot Study
The study introduces a 2-block Brain-Computer Interface (BCI) architecture for real-time EEG-based gait decoding, addressing limitations in closed-loop exoskeleton control. The architecture combines a session-specific Feature Extraction Block with artifact suppression and multi-domain feature extraction, and a Decoder Block using a novel Polynomial Time-Varying Layer (PolyTVL)+LSTM for four-state gait classification. The PolyTVL+LSTM variant (v01) achieved superior performance (validation MCC: 0.435) and consistent EEG feature discriminability (p<0.05). Closed-loop deployment demonstrated 55.3% (Rex-assisted) and 52.7% (volitional) gait initiation success, with a mean prediction time of 70.5 ms (+/-41.5), confirming real-time feasibility.
electroencephalographybrain-computer interfacegait decodingpolynomial time-varying layerreal-time control
Isotonic Bradley-Terry Model for Paired Comparison Data
The paper proposes an isotonic Bradley-Terry model for paired comparison data that jointly learns player strength parameters and an inverse link function. The method alternates between (sub-)gradient optimization of rate parameters and isotonic regression for the link function, ensuring monotonic training error improvement and handling insufficient data via exact ties. Experiments on synthetic data and real-world sports datasets (Premier League, MLB, ATP) demonstrate improved win probability prediction and ranking performance compared to fixed-link models.
paired comparisonbradley-terry modelisotonic regressionwin probabilityplayer ranking
Accelerating Evolutionary Strategy via Rao-Blackwellizing Realization of Uncertain Input
The paper introduces Phenotype-Accelerated Evolutionary Strategy (PAES), a refinement of Evolutionary Strategy (ES) for Optimization under Input Uncertainty (OIU). PAES leverages the Rao-Blackwellization technique to reduce the variance of the gradient estimator by incorporating information from the realized input, which is typically discarded in existing approaches. Theoretical analysis confirms the variance reduction, and numerical experiments demonstrate that PAES achieves faster convergence compared to standard ES across a range of problems, including simple continuous optimization and Reinforcement Learning benchmarks.
evolutionary strategyrao-blackwellizationinput uncertaintygradient estimatoroptimization
Feed-Forward Steering in Transformer Residual Dynamics
The paper extends attention-only dynamical theories of Transformer residual directions by incorporating feed-forward networks (FFNs) as local steering fields. It models FFN contributions as tangential and radial components, showing the tangential component is necessary for motion in residual-direction space and preserves model quality and output diversity. Experiments on GPT-2, Pythia, Mistral, and Llama models demonstrate improved angular prediction accuracy, with FFN contributions increasing from GPT-2 to Llama-3-8B. The findings suggest FFN layers act as directional steering fields, enabling approximate parallelization of layers with small commutator defects.
transformerresidual dynamicsfeed-forward networkcommutator defectparallelization
STEAM:ASpatio-TEmporal Alignment Mixture-of-Experts Model with Hierarchical Pre-training for EEG Decoding
The paper introduces STEAM, a hierarchical transfer framework for EEG foundation models that combines general-purpose representation learning with paradigm-specific specialization. The model employs a dual-branch spatio-temporal encoder with a shared soft mixture-of-experts (SSMoE) module to align spatial and temporal branches through soft slots, enabling complementary representation exchange. Evaluated across seven datasets and fourteen settings, STEAM achieves the best average rank at competitive FLOPs, with hierarchical pre-training further improving paradigm-specific decoding accuracy without full retraining.
eeg decodingmixture-of-expertsspatio-temporal alignmenthierarchical pre-trainingbrain-computer interface
Open-DiffLoco: Open-Source Differentiable Learning for Deployable Blind Quadruped Locomotion
Open-DiffLoco introduces an open-source framework for training deployable blind quadruped locomotion policies using differentiable simulation, addressing limitations in conventional reinforcement learning. The framework implements the Short-Horizon Actor-Critic (SHAC) algorithm in MuJoCo XLA (MJX) and trains proprioceptive policies without privileged observations or reference trajectories. It employs a simplified reward function, enabling efficient training on a single NVIDIA GeForce RTX 5080 GPU (under 6 GB VRAM) in 20-60 minutes. Deployed on a Unitree Go2 quadruped, the policy achieves root-mean-square velocity tracking error below 0.2 m/s, speeds above 1 m/s, and robustness to uneven terrain and external disturbances. The authors also propose Jacobian-Augmented Value Estimation (JAVE) to improve early policy-gradient training.
differentiable simulationquadruped locomotionshort-horizon actor-criticproprioceptive policyjacobian-augmented value estimation
A Comparative Analysis of MLP and Kolmogorov-Arnold Networks (KAN) for Faster-than-Nyquist (FTN) Signaling Detection
The paper demonstrates that Kolmogorov-Arnold Networks (KAN) outperform multilayer perceptrons (MLP) in Faster-than-Nyquist (FTN) BPSK detection, achieving 18.6× lower bit error rate (7×10⁻⁶ vs 1.3×10⁻⁴ at 10dB SNR) with 8× fewer hidden units. Using a Monte Carlo dataset of 4M samples (τ=0.8, SNR 7-10dB), the study compares MLP (hidden width 32) against KAN (hidden width 4, spline grid size 5). Results indicate KAN's superior parameter efficiency and detection accuracy for FTN signaling under AWGN.
faster-than-nyquistkolmogorov-arnold networksbpsk detectionbit error rateparameter efficiency
SCOPE: Entanglement Frontier Escape for Source-Free Class Unlearning
SCOPE introduces Spectral Conditional Projective Erasure, a source-free class unlearning method that escapes the entanglement frontier by conditioning erasure on input. It suppresses the forget subspace primarily when the frozen head's weight scores indicate a forget class, requiring no retain data or gradient training. The method is closed-form and computationally efficient, outperforming existing unlearners across five benchmarks spanning object, face, and speaker recognition tasks with convolutional and transformer backbones. SCOPE achieves superior performance at all forget-set sizes and surpasses trained methods in the most challenging settings.
source-free unlearningentanglement frontierspectral conditional projective erasureforget subspacefrozen head
Secrets Everywhere: Auditing Memorization in Mobility Prediction Models
This paper presents the first systematic audit of memorization risks in human mobility prediction models, addressing unique challenges like multi-scale trajectory structure and user-specific behavioral diversity. The authors introduce a framework to quantify memorization at three granularity levels (individual locations, anchor pairs, subtrajectory segments) and develop user-grounded reference sets to evaluate model preference for training data. Evaluations across multiple models and datasets reveal pervasive memorization patterns correlated with user regularity, highlighting privacy risks that necessitate mandatory auditing.
mobility predictionmemorization auditprivacy risksmulti-scale trajectoriesdata extraction
Adaptive Reconstruction of Bosonic Quantum States
We introduce an adaptive reconstruction technique for bosonic quantum states that estimates fidelity with respect to a family of states while reconstructing the Wigner function from minimal measurements. The method integrates a physics-informed parametric model with Bayesian inference, bootstrap, and active learning to iteratively select optimal phase space sampling points. Implemented on a circuit quantum electrodynamics platform, it achieves reproducible fidelity estimates for Schrödinger cat states (α∈[1,3]) within minutes, demonstrating robustness to phase space displacements and rotations. Experimental comparisons show adaptive sampling outperforms existing Wigner function protocols in measurement efficiency. The technique is validated in a closed-loop quantum optimal control experiment, enabling autonomous bosonic state optimization.
bosonic quantum stateswigner functionadaptive samplingbayesian inferencequantum optimal control
Déjà Cue: Localizing States in Object Histories via Vocabulary-Relative Coordinates
Déjà Cue introduces a training-free framework for identity-conditioned state-moment retrieval, localizing intervals in tracked-object histories where specific states hold. It leverages alternative state descriptions to create a vocabulary-relative coordinate system, subtracting state-balanced centroids from descriptions and calibrating frame scores using a frozen encoder. The method scans multiple durations within contiguous visible runs, improving retrieval accuracy. Evaluated on 78 VOST histories, it nearly doubles R@1 at tIoU 0.5 from 10.3% to 20.5% and raises Top-1 tIoU from 16.0% to 21.5%. Candidate-rank analyses confirm higher ranking of useful intervals.
state-moment retrievalvocabulary-relative coordinatesfrozen encodertemporal scantracked-object histories
Convex Neural Energy Elements: Monolithic Finite-Element Assembly of Geometry-Parameterized Neural Operators with Stability and Error Guarantees
The authors introduce convex neural energy elements, a method enabling monolithic finite-element assembly of geometry-parameterized neural operators with stability and error guarantees. Each element exports a scalar energy convex in boundary degrees of freedom and parameterized by geometry, realized via hypernetwork-generated positive-semidefinite quadratic forms. A regularization-nullspace principle ensures the physics nullspace is preserved, and assembled elements inherit classical positive-definite system guarantees. Theoretical error bounds are proven and experimentally validated, demonstrating 0.6-1.0% relative L2 error on heat conduction tasks and 175x faster per-geometry setup. The method supports mixing element types and extends to 3D, achieving 0.23% error on eight-element assemblies.
convex neural energy elementsfinite-element assemblygeometry-parameterized neural operatorsregularization-nullspace principlehypernetwork-generated quadratic forms
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning
The paper introduces Expectile $n$-step Q-learning (ENQ), an off-policy reinforcement learning method that replaces symmetric $n$-step temporal-difference loss with an asymmetric expectile loss to mitigate pessimistic bias from suboptimal logged actions. ENQ introduces a single hyperparameter $τ$ (expectile level) and proves its operator is a $γ^{n}$-contraction, with vanishing bias at $τ=1$ under deterministic dynamics. Evaluated on 27 manipulation and navigation tasks with $τ=0.8$, ENQ matches Long-Horizon Q-learning (LQL) in performance, achieves higher training-step throughput, and scales better with a 10-critic ensemble.
expectile lossoff-policy reinforcement learningtemporal-differencemulti-step returnsaction-value function
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling
DART (Decoded Attention over Recurrent sTates) introduces a hybrid memory representation combining recurrent compression and attention-style retrieval by leveraging Mamba-2's state space duality. The method decodes token-conditioned keys and values from chunk state memories, performs state-memory attention (SMA) via FlashAttention-style computation, and combines results with Mamba-2's output through gated residuals. Experiments show DART reduces inference cache by 75% (chunk size S=256, state size N=128) versus attention baselines while improving associative recall and preserving language-modeling performance.
state-memory attentionchunked scanassociative recallkv cacherecurrent states
Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPO
The paper proposes a transformer-enhanced proximal policy optimization (PPO) framework for collaborative mobile edge computing (MEC) servers to optimize large language model (LLM) inference under soft-deadline constraints. The method introduces an extended deadline mechanism with constrained flexibility, using transformers to model temporal dependencies and cross-server interactions for improved task migration decisions. Simulations show superior performance over conventional PPO and heuristics, achieving higher task completion rates (exact metrics unspecified) while minimizing deadline extensions.
mobile edge computingllm inferenceproximal policy optimizationsoft-deadlinetask migration
Scikit-fingerprints: Python library for scikit-learn compatible molecular fingerprints and chemoinformatics
The authors introduce scikit-fingerprints, a Python library bridging chemoinformatics and scikit-learn by providing RDKit-based molecular fingerprints and related functionalities under a unified API. The library supports molecular filters, similarity/distance metrics, applicability domain estimation, and data splitting while maintaining scikit-learn compatibility for end-to-end workflows. By adhering to scikit-learn conventions, it enables composable pipelines from SMILES strings to deployable models, emphasizing interface consistency, computational efficiency, and extensibility. The tool aims to accelerate prototyping, improve reproducibility, and simplify deployment in molecular machine learning.
molecular fingerprintschemoinformaticsscikit-learnrdkitsmiles
AOS: Adaptive Optimizer Switching via Training-State Signals for Faster Convergence and Better Generalization
AOS-R (Adaptive Optimizer Switching, Rule-Based) improves deep network training by dynamically switching between AdamW, SGD-M, and Lion based on six gradient-space signals (GNS, Hutchinson curvature trace, loss stagnation, update stability ratio, GSI, LIR). The method employs state-preserving momentum transfer and a 400-step learning-rate bridge to maintain accuracy during transitions. On CIFAR-100/WRN-28x10, AOS-R achieves 78% top-1 accuracy in 81 epochs (26% faster than AdamW, 43% faster than SGD-M, 16% faster than Lion) and outperforms AdamW on 6 of 8 benchmarks with +0.4 pp mean accuracy gain and 0.80x convergence speedup.
adaptive optimizer switchinggradient noise scalemomentum transferlearning-rate bridgehutchinson curvature trace
Detecting Nonproperness of Likelihood Equations
The article introduces a novel method for computing nonproperness sets in likelihood-equation systems, crucial for real root classification in algebraic statistical models. The approach identifies data points where solutions at infinity occur, affecting the number of real critical points. Theoretical correctness is proven, and experimental results demonstrate superior efficiency compared to existing methods.
nonproperness setlikelihood equationsreal root classificationalgebraic statistical modeldiscriminant variety
TELLER: Non-intrusive Cross-Layer Root-Cause Analysis for LLM Inference
TELLER introduces a non-intrusive framework for root-cause analysis in LLM inference by reconstructing per-request call-chain trees from NVTX/CUPTI traces and logs, aligning them with execution steps. It employs dependency-aware causal-context slices and a Trace Pair Encoding (TPE) tokenizer to compress traces into structural token sequences with parent, depth, and duration attributes. Experiments on multi-node GPU workloads demonstrate an 80% reduction in per-step trace length with moderate TPE vocabulary while maintaining diagnosis accuracy, outperforming baselines in cross-node and within-node fault localization under low-fault priors.
llm inferenceroot-cause analysistrace compressiondependency-aware slicingmultimodal diagnosis
ChaosProbe: A Neurochaotic Lens on Frozen Transformer Input-Embedding Spaces
The paper introduces ChaosProbe, a deterministic neurochaos-inspired method for generating response-based fingerprints of frozen transformer input-embedding spaces. ChaosProbe applies chaotic trajectory-based transformations to prompt-level embedding matrices, summarizing responses via Firing Rate and Entropy channels to produce fixed-length signatures. In a proof-of-concept study with 80 neutral prompts across four models (GPT-2, DistilGPT2, BERT-base-uncased, RoBERTa-base), correlation metrics successfully recovered model-family relationships, with Pearson and Spearman correlations showing stable pairings. Euclidean distance achieved partial recovery, and validity checks confirmed non-trivial fingerprint dominance.
transformerinput-embeddingneurochaosdeterministic probefingerprinting
Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation
FutureBridge-OPD (FTB) improves on-policy distillation (OPD) by validating teacher guidance through future trajectory analysis in multi-turn agentic tasks. The method executes short teacher bridges at high-disagreement states, then evaluates their impact on subsequent student trajectories to ensure increased positive distillation signal density. Evaluated on ALFWorld, WebShop, and ScienceWorld with Qwen3-32B teacher to Qwen3-1.7B student, FTB outperforms vanilla OPD and TCOD by 16.6 and 7.6 points respectively, demonstrating scalability across student sizes and teacher configurations.
on-policy distillationteacher guidancemulti-turn agentic tasksfuture trajectory validationdisagreement states
HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses
HarnessCompass introduces a novel framework for automatic harness evolution that improves generalization and effectiveness of agent harnesses. The method employs constrained evolution to ensure task-agnostic modifications, integrates proactive first-person feedback from agents, and decouples component-wise optimization to reduce interference. Evaluated on SWE-bench Verified with GPT-5.4, HarnessCompass achieves a Pass@1 improvement from 54% to 66% in 5 iterations, outperforming prior methods in both effectiveness and efficiency. The evolved harness demonstrates strong generalization across held-out tasks and models.
harness evolutiontask-agnosticfirst-person feedbackcomponent-wise optimizationgeneralization
Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning
(No summary returned.)
SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models
SpatioLM introduces a parameter-efficient vision-language model that enhances spatial intelligence without relying on external 3D priors or spatial encoders. The method employs a plug-and-play spatio-vision module to elicit inherent spatial knowledge in VLMs, utilizing pseudo depth and camera information as supervision for physically coherent representations. Extensive experiments demonstrate SpatioLM's significant improvements in spatial perception and understanding, achieving a score of 71.6 on the VSI-Bench, the first model to surpass 70, while maintaining general-purpose capabilities. It also shows competitive performance in embodied manipulation tasks.
spatio-vision modulepseudo depthspatial intelligencevision-language modelsembodied manipulation
Understanding and Correcting Low-Frequency Bias in EEG Foundation Model
We introduce FAME, a frequency-balanced masked autoencoding framework that addresses low-frequency bias in EEG foundation models. FAME reconstructs time-frequency activity in predefined EEG bands from masked inputs, independently standardizes reconstruction targets within each band, and assigns equal weight to all band-specific losses to balance spectral supervision. This approach mitigates the bias caused by EEG's $1/f^α$-like spectral structure and neural networks' preference for low-frequency components. Evaluated on 41 downstream tasks in OmniEEG-Bench, FAME achieves state-of-the-art performance on 24 tasks, demonstrating the importance of balanced spectral supervision for transferable EEG representations.
eeg foundation modelslow-frequency biasmasked autoencodingspectral supervisiontime-frequency reconstruction
CARE: A Cascaded Framework for Efficient and Reliable Time Series Anomaly Detection
The paper proposes CARE, a cascaded framework for efficient time series anomaly detection that combines a Lightweight Pre-filter Model (LPM) with a Complex Detection Model (CDM). The LPM employs a Residual MLP AutoEncoder and Normality-Conditioned Gating with Structure Attention to filter high-confidence normal samples, reducing CDM invocations via confidence-guided selective routing. Experiments on eight benchmarks show CARE achieves 2.7×-4.8× inference speedup while maintaining competitive detection accuracy compared to state-of-the-art methods.
time series anomaly detectioncascaded inferenceresidual mlp autoencodernormality-conditioned gatingstructure attention
Probabilistic Deep Learning for Drought Forecasting: Role of Internal Climate Variability
The study introduces a deep-learning-based framework for European drought forecasting that explicitly incorporates internal climate variability from a large ensemble of climate models to produce uncertainty-aware drought bounds. The method leverages probabilistic deep learning to generate physically plausible lower-tail trajectories of future drought conditions, providing risk-averse references for adaptation planning. Results demonstrate that the ensemble-informed bounds outperform reanalysis-derived bounds in calibration, particularly during anomalously dry conditions, where historical data alone underestimates drought risk.
probabilistic deep learningdrought forecastinginternal climate variabilitylarge ensemblerisk-aware bounds
WorldDynCache: Risk-Controlled Latent Dynamics Approximation for Diffusion World Model
WorldDynCache introduces a risk-controlled latent dynamics approximation framework to accelerate diffusion world models while maintaining generation quality. The method combines a lightweight latent-transition risk estimator, which tracks approximation defects via counterfactual calibration, with a condition-aware lifted latent surrogate that approximates latent evolution without transformer evaluations. Evaluated on HunyuanVoyager-13B and Aether-5B, WorldDynCache achieves 4.92× and 2.15× speedups respectively while outperforming other caching methods on WorldScore, PSNR, SSIM, and LPIPS metrics.
diffusion world modellatent dynamicsrisk estimatortransformer evaluationsgeneration quality
tFUSOperator: Operator Learning for Transcranial Focused Ultrasound Digital Twins
tFUSOperator introduces a neural operator framework for transcranial focused ultrasound (tFUS) digital twins, addressing the computational inefficiency of numerical solvers and the physical inaccuracy of voxel-to-voxel deep learning surrogates. The method maps free-field pressure, skull anatomy, and treatment parameters to the intracranial acoustic field within a shared physical coordinate frame, leveraging operator learning principles. Evaluated on both seen and unseen skulls, the model achieves 90% and 72% Dice scores, respectively, and performs comparably with MR and CT inputs while running 5.6×10^4 times faster than numerical simulation. This approach enables fast, radiation-free digital twins for patient-specific tFUS treatment.
neural operatortranscranial focused ultrasounddigital twinsacoustic fieldoperator learning
Tunneling the Loss Landscape: Bypassing Memorization with Monte Carlo Parameter Swapping
The paper introduces a three-component framework to analyze grokking in neural networks through glass dynamics metrics: parameter mobility (PM), replica correlation (RC), and fractal dimension (FD). It demonstrates that standard optimization traps networks in kinetic arrested memorization states, exhibiting collapsed mobility and history dependence. Inspired by swap Monte Carlo, the authors propose State-Aware Monte Carlo Parameter Swapping (SAM-Swap), which accelerates generalization by enabling random parameter exploration, analogous to physical diffusion processes.
grokkingglass dynamicsparameter mobilityswap monte carlomemorization
DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Models
DAVET introduces a training-free framework for adaptive visual evidence allocation in diffusion vision-language models (dVLMs), addressing the recurring inference cost of visual conditioning. The method leverages a phase-conditioned evidence trajectory, modulating evidence allocation at each denoising step based on trajectory risk and operation demand. It constructs a hierarchy of evidence views from a single visual encoding, decoupling when and how much evidence is needed from view construction. Evaluated on LLaDA-V and LaViDa across multiple benchmarks, DAVET achieves a 1.55× speedup with only a 1.86% average relative performance drop, demonstrating efficient visual conditioning cost reduction.
diffusion vision-language modelsvisual evidence allocationdenoising-awaretraining-free frameworkvisual-understanding benchmarks
Reassessing the Feasibility of PPG-Based Non-Invasive Blood Glucose Level Estimation
We introduce the first reproducible evaluation pipeline for PPG-based non-invasive blood glucose level (BGL) estimation, addressing inconsistencies in datasets, data leakage, and evaluation metrics. Five representative PPG-BGL methods were reassessed under three data-split protocols: random window-level, participant-aware, and leave-some-participants-out (LSPO). While models appeared competitive under random splits, they collapsed under participant-aware and LSPO evaluation, yielding near-zero or negative R$^2$ values comparable to a mean-prediction baseline. Notably, over 90% of predictions fell within clinically acceptable zones (Clarke Error Grid A+B) across all models and splits, revealing a disconnect where clinical metrics conceal model failure. The findings highlight that random splits overestimate generalization due to sample-level data leakage, emphasizing the need for robust ML evaluation before clinical validation.
photoplethysmographyblood glucose leveldata leakageclarke error gridgeneralization
ReFP-AD: Rectified Flow Preconditioning for Energy-Based Anomaly Detection
ReFP-AD introduces a geometric reparameterization method for anomaly detection in high-dimensional token spaces, addressing instability in Energy-Based Models (EBMs) caused by anisotropy and cross-dimensional correlations. The approach leverages optimal transport-coupled rectified flow to map embeddings into a well-conditioned latent space, enabling stable persistent contrastive divergence with preconditioned Stochastic Gradient Langevin Dynamics (SGLD). Anomaly scores are derived from gradient norms of the learned energy landscape. Evaluated on MVTec-AD and VisA datasets under a unified protocol, ReFP-AD achieves 98.6%/97.9% Image/Pixel AUROC on MVTec-AD and 97.3%/99.0% on VisA, outperforming prior EBM baselines by up to +10.8% in Image AUROC.
energy-based modelsrectified flowoptimal transportstochastic gradient langevin dynamicsanomaly detection
Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints
HeLyMARL introduces a Lyapunov-embedded heterogeneous multi-agent reinforcement learning framework for radio resource management under coupled finite-horizon energy and handover constraints. The method employs drift-plus-penalty decomposition with virtual queues to internalize constraint pressures into per-slot rewards, converting the constrained problem into an unconstrained MARL formulation. Simulations demonstrate that HeLyMARL maintains throughput-fairness balance and uninterrupted service, outperforming conventional MARL, Lyapunov-based, and constrained MARL benchmarks without premature budget exhaustion.
multi-agent reinforcement learningradio resource managementlyapunov methodsdrift-plus-penaltyvirtual queues
Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning
The paper introduces Correctness-Conditioned KL Regularization (CoKL), a framework for preserving existing capabilities in large language models (LLMs) during reinforcement learning-based post-training. CoKL narrows the preservation constraint from the full output distribution to correctness-conditioned response distributions, using forward KL divergence and a finite-group training objective. This approach decouples the total probability assigned to correct responses from their correctness-conditioned distribution, avoiding limitations of full-policy KL regularization. Experiments in multi-solution environments and continual post-training settings across multiple model scales show CoKL achieves better target-task improvement and prior-capability retention compared to existing methods.
kl regularizationreinforcement learninglarge language modelspost-trainingcorrectness-conditioned
LLM-Guided Retrieval for Prediction of Molecular Perturbation Responses
The paper introduces LLM-Guided Retrieval (LGR), a retrieve-and-aggregate method for predicting transcriptomic responses to small-molecule perturbations. LGR uses a large language model (LLM) to rank biologically related compounds profiled in a target cell line, then aggregates their observed expression deltas via a fixed mean. Evaluated on the Tahoe-100M single-cell perturbation atlas, LGR outperforms drug mean, ChemCPA, and chemistry-based kNN baselines, particularly in unseen-cell-line generalization, achieving higher correlation and lower error. It also improves directional accuracy of gene regulation, suggesting LLMs enhance retrieval quality for zero-shot prediction. Results indicate LLMs provide a useful biological prior as constrained retrieval modules.
transcriptomic responsesretrieve-and-aggregatelarge language modelsingle-cell perturbationzero-shot prediction
CENTILE: A Telemetry Foundation Model Evaluated by the Decisions It Drives
CENTILE introduces a generative foundation model for network and systems telemetry, evaluated by its impact on operational decisions rather than point-forecast error. The model processes heterogeneous, irregularly timed entity streams and supports flexible forecast horizons without requiring future timestamps. Pretrained on operator event streams, CENTILE demonstrates zero-shot transfer across months and domains with minimal target data. Experiments on HPC job logs and network traffic show CENTILE reduces mean bounded slowdown in backfilling by up to 77% and halves rule violation rates compared to deployed estimators.
generative foundation modeltelemetryconditional quantileszero-shot transferbackfilling
Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models
The paper proposes External Rollout Integration with Length Control and Source-Specific Processing (ERILS), a method to enhance reinforcement learning in diffusion language models by incorporating higher-reward rollouts from an external policy alongside on-policy rollouts. ERILS addresses challenges of rollout length mismatch and reward instability through length control and separate reward processing. Experiments on Sudoku, Countdown, and MATH500 show ERILS achieves 98.4% best-of-4 accuracy on Sudoku, outperforming baselines by 58.1%, while maintaining 90% single-completion accuracy across token lengths. Component analysis confirms the necessity of length control and source-specific reward processing.
diffusion language modelsreinforcement learningon-policy rolloutsexternal policyreward processing
Beckmann Transport Models: From Autonomous Flows to One-Step Maps
The authors propose an instantiation of flow matching using a time-independent velocity field (autonomous flow) to exactly map between distributions when the target is singular (supported on a lower-dimensional manifold). The associated one-step generative map is derived as the unique solution to a conservation equation, enabling direct learning from samples. This framework unifies existing methods like the Poisson-flow generative model and equilibrium matching with quadratic flow-matching loss, while correcting inconsistencies. Experiments demonstrate effectiveness on ImageNet 256x256.
flow matchingautonomous flowone-step mapbeckmann transportsingular distribution
Progressive Agent Skill Generation via Reinforcement Learning
The paper proposes Skill-$α$, a reinforcement learning method for progressive agent skill generation that addresses the lack of natural supervision signals. The approach formulates skill generation as a sequential editing process with rollback rewards, evaluating edits by comparing downstream execution before/after each modification. Experiments show Skill-$α$ outperforms heuristic and pipeline baselines, improving downstream success rates by 3.3 points on CL-Bench and 6.7 points on tau2-bench when using GPT-4o, with ablations confirming the importance of progressive generation and rollback rewards.
skill generationreinforcement learningsequential editingrollback rewardprogressive generation
Generative Brownian Bridge Diffusion In Motion Space For Enhanced Myocardial Strain Analysis
The authors propose a generative Brownian bridge diffusion model operating in motion space to enhance myocardial strain analysis from standard cardiac magnetic resonance (CMR) images. The method learns a probabilistic mapping between motion estimates from conventional registration techniques and high-accuracy motion from specialized strain imaging, conditioned on corresponding CMR images for anatomical fidelity. Evaluated on multi-center CMR datasets, the approach significantly outperforms existing learning-based methods in motion prediction and strain analysis accuracy, offering a cost-effective alternative to specialized imaging for cardiac function assessment.
brownian bridge diffusionmyocardial strain analysiscardiac magnetic resonancemotion spacegenerative modeling
Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation
The study introduces a counterfactual audit framework to analyze sparse attention's causal effects in long-context models, revealing two competing patterns: signal concentration (preservation of critical blocks) and integration loss (disrupted cross-block attention). Using Block Sparse Flash Attention (BSFA) route replay and controlled probes (Gold, Poison, Benign) across four architectures, the authors show that sparsification alters content influence non-uniformly—Gold and Poison blocks are preserved significantly more than Benign blocks (G≈P≫B), while higher compression ratios amplify sparse effects, sometimes reversing outcomes. Three validation methods (BSFA replay, block-top-k, KV-cache eviction) confirm these findings, undetectable via aggregate accuracy metrics.
sparse attentioncounterfactual evaluationlong-context modelskv-cacheintegration loss
Sharp Root Anti-Concentration via Projective Incidence and Ordered Root Laws
(No summary returned.)
FAU at ImageCLEF 2026 Task on Multimodal Reasoning Robust Candidate Scoring and Concise Multilingual Visual Answering
The authors present a multimodal reasoning system for ImageCLEF 2026's Visual MCQ and Visual OpenQA tasks, emphasizing robust output control alongside model selection. For Visual MCQ, they replace free-form generation with direct candidate label scoring from vision-language model logits, employing score fusion and voting. For Visual OpenQA, they utilize image enhancement, concise answer prompting, deterministic decoding, and post-processing to eliminate reasoning traces and formatting artifacts. Without task-specific training, their system achieved third place in Visual MCQ (0.7108 accuracy) and first in Visual OpenQA (0.6488 COMET, 0.1391 BLEU, 0.2762 ROUGE L, 0.2383 METEOR), demonstrating the efficacy of inference engineering.
multimodal reasoningvision-language modelscore fusiondeterministic decodinginference engineering
Non-KKT Accumulation in Entropic Mirror Descent
The article presents the first counterexamples demonstrating that bounded sequences generated by Shannon-entropic mirror descent may accumulate at non-KKT (Karush-Kuhn-Tucker) stationary points, resolving a longstanding open question in optimization. Using $C^\infty$ objectives on the nonnegative orthant $\R_+^n$ ($n\geq 3$) and probability simplex $\Delta_n$ ($n\geq 4$), the authors construct sequences with smooth boundary circles as accumulation sets, containing non-KKT arcs. The analysis reveals this pathology arises from Bregman geometry degeneracy at the boundary, despite nonincreasing objectives, proper stepsizes ($\alpha_k \asymp k^{-\beta}$, $\beta \in (1/2,1)$), and entropy-relative smoothness.
mirror descentkkt conditionsbregman divergencelegendre kernelnonconvex optimization
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models
Bole introduces a kernel--runtime co-design for efficient tree speculative decoding in hybrid-attention large language models, addressing memory-bound autoregressive decoding. It transforms linear-attention recurrence into a tree-structured closed form, enabling parallel verification of proposal nodes with a resource-efficient GPU kernel. Bole losslessly encodes speculative state updates as token-level factors and reconstructs only the selected state, reducing transient memory usage by 82--99×. Integrated into SGLang, it achieves up to 4.72× offline decode throughput and reduces TTFT and TPOT by up to 67.6% and 49.9% respectively under online agent workloads.
tree speculative decodinghybrid-attentionlinear-attention recurrencekv cachesgpu kernel
Evaluating Forecasting Techniques for Hardware Errors on a Large-scale HPC System
This study evaluates time series forecasting techniques for hardware errors in high-performance computing systems, using seven years of production logs from the Theta supercomputer. The authors compare classical statistical and deep learning models, including LSTM and Transformer architectures, focusing on temporal features. Results indicate that forecasting effectiveness varies significantly with the temporal structure of error series: regularly occurring and structurally stable errors are accurately modeled, while sparse and burst-dominated errors remain challenging. The work provides empirical guidance on forecasting applicability and identifies potential directions for improving accuracy in HPC hardware error analysis.
time series forecastinghardware errorslstmtransformerhigh-performance computing
GraphIR: Architecture-Level Search States for LLM-Guided Neural Architecture Evolution
The paper introduces GraphIR, an architecture-aware intermediate representation that enhances neural architecture search (NAS) by providing mutation-aligned candidate states for LLM-guided evolution. GraphIR structures candidates via three views: computation skeleton (tensor flow), mutation surface (editable modules), and validity envelope (interface contracts). Evaluated on NAS-Dependency, a 120-question benchmark, GraphIR excels in dependency tracing and interface diagnostics. On six benchmarks including CLRS, it achieves superior search performance while maintaining model size and NAS efficiency in OpenEvolve.
neural architecture searchintermediate representationdependency reasoningllm-guided evolutionmutation surface
Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models
The study investigates gradient-free weight perturbation in language models, identifying perturbation norm as the critical factor over subspace choice or dimension. Using a fixed pipeline with controlled variables, the authors compare full-weight perturbations against frozen frames of 12-16 scalars across 49 model-benchmark pairs, finding the latter trails by 1.8 accuracy points. Random frames perform equivalently to SVD-aligned frames when norm-matched, with perturbation scale being the sole failure-sensitive parameter. The usable norm range remains consistent across seven models, narrowing design focus to perturbation intensity rather than subspace selection.
gradient-free adaptationweight perturbationsubspace selectionperturbation normlanguage model adaptation
Online Algorithms via Minimax and Posterior Matching
The paper introduces a unified minimax methodology for competitive analysis of online algorithms, reducing worst-case guarantees to Bayesian online design under arbitrary correlated priors. The proposed posterior matching rule selects online actions that track the posterior expectation of the hindsight-optimal solution, constrained by feasibility. This approach yields optimal or near-optimal competitive ratios for fractional problems like set cover, load balancing, and matching, while extending via rounding to integral problems like weighted paging and ski-rental. Key technical contributions include probabilistic inequalities for vector martingales derived from the offline optimum's posterior, providing a general framework for information-theoretic competitive guarantees.
competitive analysisminimax optimizationposterior matchingonline algorithmsvector martingales
Thermalizing Stochastic Programs
The authors introduce a framework for compiling stochastic programs to thermodynamic hardware, enabling energy-efficient stochastic sampling. The method maps Directed Factor Graphs (DFGs) or Parametrized Stochastic Circuits (PSCs) to Energy-Based Models (EBMs) native to the hardware, with analysis of error accumulation across factors. Two training refinements—context matching and trajectory-level REINFORCE post-training—are proposed to reduce residual errors. Implemented via the exttt{thermalizers} framework, the approach is demonstrated on applications including financial market simulation, ecological modeling, Gibbs sampling of EBMs, and sequential Bayesian design loops.
directed factor graphenergy-based modelstochastic samplinggibbs samplingbayesian design
Latent-Regime Bias Auditing for Volatility Forecasting
The paper introduces a model-agnostic audit framework for evaluating volatility forecasts across latent market regimes, addressing limitations of aggregate metrics like RMSE and MAE. The method involves learning time-series representations of market-state windows, clustering them into regimes using training data, assigning regimes out of sample, and comparing aggregate forecast behavior with regime-conditional bias, tail-underprediction, and economic losses. Applied to daily volatility forecasting for cryptocurrency and ETF assets, the audit reveals that models with competitive aggregate accuracy exhibit significant regime-specific bias and severe tail underprediction. The framework emphasizes evaluating forecasts based on conditional reliability rather than average error.
volatility forecastinglatent regimestail-underpredictionmodel-agnostic auditeconomic losses
Statistical comparisons of time-series feature sets on classification tasks
This study systematically evaluates six open-source time-series feature sets and three baseline methods on 124 univariate classification tasks using a normalization-based benchmarking approach. The analysis reveals comparable overall performance (85.3% pairwise ties), with tsfresh achieving the highest win rate (29.03%) despite feature set diversity in size and composition. Results demonstrate problem-dependent performance variations and cases where simple spectral/quantile baselines suffice, emphasizing the importance of feature composition and task-specific benchmarking.
time-series classificationfeature setsnormalization-based benchmarkingtsfreshunivariate analysis
Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer
This work introduces topological methods for analyzing semantic alignment in AI models, addressing challenges in benchmarking language models beyond outcome reasoning. By comparing high-dimensional embedding spaces to low-dimensional, interpretable baselines such as ontologies and curated knowledge graphs, the study enables rigorous evaluation of conceptualization and semantic structure. The approach facilitates tracking model adaptations and testing phrase understanding across multiple languages, providing deeper insights into how models relate abstract concepts. The method leverages multi-modal alignment tests to characterize semantic spaces, offering a novel framework for evaluating cross-lingual transfer and checkpoint dynamics.
semantic alignmenttopological methodsembedding spacescross-lingual transferknowledge graphs
LieStoNet: Learning Lie Symmetries from Spatiotemporal Data for Stochastic Dynamical Systems
LieStoNet introduces an end-to-end framework for discovering Lie-point symmetries in stochastic dynamical systems directly from spatiotemporal data, without predefined symmetry templates or canonical coordinates. The method learns neural surrogates for drift and diffusion terms, enforces SDE determining equations, and regularizes for Lie algebra axioms (bilinearity, antisymmetry, Jacobi) and non-redundant basis adherence. Evaluated on canonical SDEs with known symmetries, LieStoNet recovers ground-truth symmetry generators, demonstrating interpretable symmetry discovery for noisy dynamics. The framework also enables optional parallel discovery of Fokker-Planck equation symmetries.
lie symmetriesstochastic dynamical systemsfokker-planck equationneural surrogateslie algebra
Meganeura: Portable GPU Training and Inference through Vulkan and Metal
Meganeura introduces a compact native compiler enabling portable GPU training and inference across consumer GPUs via Vulkan and Metal APIs. The system employs a typed static graph, automatic differentiation, optimizer, checkpointing, and memory planner to lower specialized programs. Evaluated against PyTorch on NVIDIA, AMD, Apple silicon, and Intel GPUs, Meganeura achieved 48/50 successful device-workload-mode cells, with 12/20 minimal-latency wins in strict f32 mode and a median training gap of 1.8x. Compilation times ranged 0.1-2.4 seconds versus PyTorch's 6-96 seconds, with a 13 MiB binary. Key gaps were identified in convolution derivatives and attention backward, demonstrating consumer graphics APIs' viability for train-to-deploy stacks.
vulkanmetalautomatic differentiationconvolution derivativesattention backward
Generalized Quadratic Gradient: A New Direction in Optimization via the Fusion of Positive-Definite Curvature Matrices and Gradients into A Unified Framework
The paper introduces Generalized Quadratic Gradient (GQG), a unified optimization framework extending quadratic gradient principles to Newton-type algorithms. GQG generalizes beyond specific Hessian approximations (e.g., BFGS, diagonal Hessians) by requiring only positive-definite curvature matrices satisfying a local quadratic model's stationary condition. The method enables curvature-aware updates across diverse Hessian surrogates, broadening the theoretical foundation for second-order optimization techniques while maintaining computational tractability.
generalized quadratic gradientnewton-type optimizationpositive-definite curvaturehessian approximationcurvature-aware optimization
Finite-Probe Total-Variation Certificates for Finite-Basis Drifting Models
(No summary returned.)
Dominant Arm Identification with Mixing and Recycling Observed Samples
The paper introduces a novel dominant arm identification algorithm for multi-armed bandits, addressing limitations of mean-based and pairwise comparison methods. The method combines (i) a dominance score criterion partitioning the reward space and (ii) a joint mixing/recycling mechanism with a doubly robust estimator, ensuring simultaneous convergence of empirical distribution functions. Theoretical guarantees show near-optimal sample complexity, while experiments demonstrate exact recovery of the true dominant arm, outperforming baselines.
multi-armed banditsdominant arm identificationdoubly robust estimatorsample complexityempirical distribution
Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference
Celty introduces a co-designed sparse format, GPU kernel, and SIMT microarchitecture to optimize Sparse Matrix-Sparse Vector (spMspV) workloads in dual-sparse LLM inference. The method combines a Run-Length Compressed CSC (RLC-CSC) format for vectorized weight loading with a Sparse SIMT Core featuring a pipelined RLC decoder and conflict-free accumulation. This approach eliminates unnecessary memory accesses and avoids data layout changes. Celty achieves up to 2.8x speedup over cuBLAS and 2.4x over Flash-LLM with the GPU kernel alone, and up to 5.3x speedup over cuBLAS at 70% dual-sparsity when incorporating the Sparse SIMT Core.
sparse matrix-sparse vectorrun-length compressed cscsimt microarchitecturedual-sparsitygpu kernel
Gram-Space: Structure-Preserving Codebook Compression for Memory-Efficient Neuro-Symbolic AI
Gram-Space introduces a structure-preserving codebook compression framework for vector symbolic architectures (VSA) in neuro-symbolic AI, addressing memory bottlenecks via Gram-Schmidt orthogonalization. The method represents codebook vectors in an orthonormal coordinate system while preserving dot-product structure for matrix-based VSA operators, enabling numerically equivalent execution of similarity and attention computations. Evaluations on standard datasets demonstrate 15.75x GPU memory reduction and 3.62x latency improvement, with profiling confirming reduced allocation overhead and better hardware utilization.
vector symbolic architecturesgram-schmidt orthogonalizationneuro-symbolic aicodebook compressionmemory efficiency
Stochastic Sequential Search in Very-High-Dimensional Feature Selection
The paper introduces Stochastic Sequential Search (SSS), a family of feature selection methods that replace exhaustive candidate evaluation with budgeted sampling, enabling scalable performance in very-high-dimensional spaces. SSS employs temperature-controlled softmax sampling from dependency-aware feature statistics, maintaining uniform exploration guarantees while reducing per-step cost to be dimensionality-independent. Evaluations on madelon (500D), gisette (5,000D), and reuters (10,105D) show that SSS variants (e.g., sSFFS) retain 97% of full-SFFS criterion value at 1/4 evaluations, outperform ranking methods (DAF, BIF) at matched budgets, and achieve superior holdout accuracy, with 10^4-dimensional problems solved in minutes.
feature selectionhigh-dimensionalstochastic searchsoftmax samplingsequential subset
BiKAN: Restoring Collapsed Basis of Binary Kolmogorov--Arnold Networks
BiKAN addresses Spatial Orthogonality Collapse in binary Kolmogorov--Arnold Networks (KANs) by augmenting each layer with degree-2 Walsh characters, restoring pairwise coordinates without learned routing. Fixed circular channel rolls generate parities, while learned binary projections mix them using XNOR--popcount operations. On CIFAR-10, parity improves accuracy by 1.23 points (p=0.003), with gains increasing at lower widths; at ~11.9M parameters, it outperforms conventional widening by 3.09 points (p<10^-4). BiKAN achieves 99.48%, 84.38%, and 55.81% on MNIST, CIFAR-10, and CIFAR-100 (W1A1). FPGA implementation reduces DSP usage from 164 to 72 and latency from 401 to 54.8 ms, with zero-DSP inference possible at minimal accuracy loss.
kolmogorov-arnold networkwalsh charactersspatial orthogonality collapsexnor-popcountw1a1
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
The study improves interpretability in MEG-based speech decoding by redesigning a high-performing retrieval architecture. It replaces spatial attention with spherical harmonics on 3D MEG geometry, reduces subject-specific branches from 270 to 25, and adds temporal filters matching neuronal sources. The model achieves 39.75% Top-1 accuracy on MEG-MASC with 20× fewer parameters, maps weights to cortical sources consistent with speech perception, and identifies 15 of 19 stimulus features contributing to retrieval, notably silence, intensity, vowels, and onsets. Wav2vec targets compress to ~12 dimensions without accuracy loss.
meg decodingspherical harmonicswav2veccortical sourcestemporal filters
Plasticity of Growing and Elastic Neural Networks in Online Continual Learning
This paper investigates the plasticity of growing and elastic neural networks in online continual learning, addressing catastrophic forgetting and loss of plasticity. Adaptive growing networks incrementally add randomly initialized hidden units while maintaining adaptive connections, preserving high prediction accuracy despite increasing dead unit proportions. Adaptive elastic networks further prune dead units at task transitions, achieving excellent accuracy, sustained plasticity, and compact size. Experiments in supervised learning settings demonstrate that these networks adapt their structure to learning objectives, offering a promising approach for plasticity preservation in online continual learning.
online continual learningplasticitygrowing neural networkselastic neural networkscatastrophic forgetting
Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs
Slot2Text introduces a dual-mode surgical MLLM that replaces dense visual tokens with compact slot latents for efficient and spatially traceable scene understanding. The method groups self-supervised vision features into region-based slots, consumed as area-labeled tokens by the language model. Slot2Text-Fast reduces average token consumption by 91.8% (from 1,295 to 47 visual tokens) while matching SOTA performance; Slot2Text-Reason trades additional tokens for explicit spatial grounding. Results demonstrate slot latents as an efficient default visual interface for surgical MLLMs.
multimodal large language modelsslot latentsvisual tokenizationsurgical scene understandingspatial traceability
Conformalized Large Language Models under Configuration Shift
The paper investigates configuration shift in conformal prediction (CP) for large language models (LLMs), where nonconformity scores vary due to configurable inference parameters (prompt templates, decoding temperature, weight quantization). Through empirical analysis across 9 LLMs, 4 datasets, and 4 nonconformity scores, the authors demonstrate that configuration shift consistently degrades CP validity, reducing empirical coverage below target levels while preserving efficiency. They derive coverage lower bounds to quantify shift severity and propose mitigations like bound-inspired recalibration and fragility-aware calibration ensembling, which recover lost coverage without requiring test data.
conformal predictionlarge language modelsconfiguration shiftnonconformity scoresuncertainty quantification
How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection
The study demonstrates that performance claims in provenance-based intrusion detection systems (PIDS) are highly sensitive to benchmark selection and evaluation protocols. By re-evaluating representative PIDS on audited datasets (primarily DARPA TC E3) with unified temporal test splits and validation-only calibration, the authors find that alerting success often diverges from forensic utility, with simple allowlists matching or outperforming learned baselines on 3/4 datasets. Semantic signal quality analysis (via feature completeness and field entropy) reveals that only datasets like Theia, with high semantic signal quality, reliably expose architectural differences in node-level recovery. The results caution against interpreting PIDS architectural claims without explicit benchmark and protocol context.
provenance-based intrusion detectionbenchmark calibrationsemantic signal qualitytemporal evaluation splitsnode-level recovery
DynamicManip: Enabling Dynamic Manipulation from a Single Static Demonstration
DynamicManip enables dynamic robot manipulation from single static demonstrations via a static-to-dynamic data augmentation pipeline and a dynamic-aware adaptive policy. The method synthesizes diverse dynamic trajectories from static data and adjusts inference frequency (18.4% higher success rate, 32.9% lower latency) using real-time dynamics. A benchmark with automatic evaluation validates scalability. Experiments show superior data efficiency and performance in dynamic tasks compared to baselines.
dynamic manipulationdata augmentationimitation learningadaptive policyrobot benchmarking
Statistical Mechanics of Learning on Product Wasserstein Manifolds
The paper reformulates neural networks and variational quantum circuits as gradient flows on product Wasserstein manifolds, treating weight distribution constraints as intrinsic geometry rather than capacity restrictions. It introduces a hierarchical mean-field framework for deep networks and extends it to quantum settings using quantum Wasserstein-1 distance, proposing two algorithms (Hierarchical DisCo-SGD and Quantum DisCo) that follow approximate geodesics. Experiments on teacher-student problems, image classification, and quantum classifiers demonstrate improved generalization, training stability, and mitigation of barren plateaus compared to unconstrained or norm-based baselines, suggesting geometric priors for incorporating domain-specific constraints.
wasserstein manifoldgradient flowvariational quantum circuitsmean-field approximationbarren plateaus
PolymerGPT: Multi-property Optimization with a Decoder-Based GPT Model for Generative Polymer Design
PolymerGPT introduces a decoder-based GPT model for multi-property optimization in generative polymer design, addressing the limitations of single-property optimization methods. The model incorporates up to 37 polymer properties via learned conditioning prefixes and supports scaffold conditions for structure specification. Experimental results show that PolymerGPT achieves high validity, uniqueness, and novelty in both unconditional and conditional generation. Conditioning on five key properties produces structures whose predicted values closely match all target properties simultaneously, enabling accurate prediction of macroscopic material behavior.
polymer designdecoder-based gptmulti-property optimizationconditioning prefixesscaffold condition
When Replanning Becomes the Bottleneck: Budgeted Replanning for Embodied Agents
The paper introduces BRACE, a budgeted control loop for LLM-based replanning in embodied agents, addressing latency-heavy tails caused by growing textual context. BRACE combines decision-making on replanning necessity, mode selection, and explicit token/latency budgeting with E-RECAP, a cost-aware progressive token pruning method that preserves critical tokens while reducing context size. Evaluations across Meta Habitat, RoboFactory, and AirSim show 62-92% token reduction and SLO violation rate drops from 85.5-100% to 4.7-50%, with BRACE + E-RECAP achieving 80% success in challenging scenarios where baselines fail.
embodied agentsreplanningtoken pruninglatency slollm-based control
Cluster-Aware Over-the-Air Federated Learning with Energy-Harvesting Devices: From Global Training to Model Personalization
The paper proposes a unified framework for over-the-air federated learning (FL) with energy-harvesting devices under heterogeneous data distributions, addressing both global model training and cluster-specific personalization. The method leverages cluster-aware scheduling for energy-efficient participation and exploits the wireless multiple-access channel for simultaneous model updates. Results show improved fairness in global mode and enhanced personalization in cluster mode, with reduced communication overhead compared to conventional FL approaches.
federated learningenergy harvestingover-the-air computationmodel personalizationcluster-aware scheduling
Training Small LLMs as Spatial Multi-Agent Policies
The paper introduces a method for training small frozen LLMs as spatial multi-agent policies using symbolic options and multi-agent reinforcement learning. Each agent's LLM acts as a policy over options, with private LoRA adapters trained via per-agent multi-agent GRPO (PA-MAGRPO). Symbolic options are synthesized from game source code and feasibility guards derived from random-policy rollouts. The approach lifts frozen LLMs from zero reward to competent play across three games and four small backbones. Behavioral audits reveal a decoupling between reward and cooperation, emphasizing the need for behavioral evaluation alongside reward metrics.
multi-agent reinforcement learningsymbolic optionslora adaptersfeasibility guardsbehavioral audits
Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics
The paper introduces two complementary notions of alignment for reference-based text evaluation metrics: statistical alignment (correlation with human ratings) and strategic alignment (resistance to task-irrelevant perturbations). It proposes test principles—human-rating correlation, degradation sensitivity, and manipulation robustness—and a unified design framework for mutual-information-based metrics, decomposing them into information measure, estimation method, text representation, and prediction mechanism. Experiments in peer review, summarization, and question answering show that LLM-as-a-Judge achieves high human-rating correlation but is susceptible to manipulation, while mutual-information-based metrics improve robustness. A new metric derived from the framework demonstrates strong robustness while maintaining competitive correlation.
reference-based metricsstatistical alignmentstrategic alignmentmutual-information-based metricsmanipulation robustness
QR-Erase: Efficient Subspace-Based Machine Unlearning with Layer Localization
QR-Erase introduces a subspace-based machine unlearning framework leveraging Pivoted QR decomposition to remove task-specific representations from model parameters efficiently, avoiding costly singular value decompositions (SVD). The method includes Layer-Localized QR-Erase, which targets updates to layers with concentrated task-specific information. Theoretical analysis shows Pivoted QR achieves bounded error in subspace recovery and approximates SVD under a mild spectral gap condition. Empirical results across task-level, cross-lingual, and speech unlearning demonstrate QR-Erase outperforms optimization-based methods in forgetting-retention tradeoff, maintaining within 5% of SVD metrics while reducing speech forget-set accuracy from 53.1% to 15.7%.
machine unlearningpivoted qr decompositionsubspace recoverylayer-localized updatesspectral gap condition
Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning
Prefix-Normalized Policy Optimization (PNPO) is proposed to address computational inefficiency in LLM reinforcement learning by reusing rollout batches while mitigating off-policy degradation. The method normalizes cumulative importance ratios via geometric mean of likelihood ratios along causal prefixes, preserving prefix dependence while controlling log-weight scale. In mathematical reasoning benchmarks with long contexts, 4-epoch PNPO achieves 50.24 Avg@32 (3pp above GSPO) and matches 1-epoch performance with 4× fewer rollouts under fixed update budgets, demonstrating advantages in off-policy regimes.
off-policy correctionimportance ratioautoregressive rolloutprefix normalizationpolicy lag
TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction
TabDPT-Turbo introduces an efficient row-based attention architecture for tabular prediction, combining long-context pre-training and SSL on a larger real-data corpus to eliminate retrieval needs. The model matches TabDPT v1.1's performance on TabArena-Lite, CC18, and CTR23 benchmarks while achieving orders-of-magnitude faster inference, becoming the fastest among leading tabular foundation models. Architectural optimizations and expanded pre-training data enable this speedup without sacrificing accuracy.
in-context learningtabular predictionrow-based attentionssl pre-trainingfoundation models
Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent Safety
The paper establishes a theoretical bound linking the recall of LTL/FSA-based runtime safety monitors for LLM agents to the entropy of attack distributions, proving that coverage is inversely proportional to attack pattern dispersion. Through empirical validation across eight LLM architectures, it demonstrates that low-entropy attacks (e.g., GPT-class, H ~ 0.24 bits) enable high monitor recall (68-75%), while high-entropy attacks (e.g., Gemini variants, H ~ 2.81 bits) result in near-zero coverage (6-13%). A pre-deployment entropy test predicts monitor effectiveness with 76% variance explained (Pearson r = -0.87).
linear temporal logicfinite automataattack distribution entropyruntime safety monitorllm agent safety
On the Identifiability of Masked Prediction: Mode Blindness and Mask Schedules
(No summary returned.)
When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design
The work establishes formal conditions under which machine-learning surrogates can safely replace expensive experimental evaluations in design campaigns, demonstrating that predictive accuracy (R^2) is insufficient for trustworthy candidate selection. Through mathematical analysis and validation on three exhaustively evaluated design tasks, the authors prove that safety requires architectural separation between proposal/certification and derive a minimal rank-preservation criterion for surrogate trustworthiness. Results show that selection-aware audits achieve Spearman rank correlations of 0.80-0.99 with deployed performance, while R^2 correlates as low as 0.33 with regret, with audited screening reducing evaluation costs by 25×.
surrogate modelsdesign optimizationselection taxrank preservationoracle certification
Do Neural Networks Really Beat the Curse of Dimensionality? A Bit-Complexity View
The paper develops a bit-complexity framework for evaluating approximation methods, analyzing classical techniques (polynomial approximation, sparse grids, finite elements) and neural networks via binary encoding and metric entropy. Results show classical methods are suboptimal in bit complexity compared to metric entropy limits, while neural networks exhibit varied behaviors but no fundamental superiority. Key findings reveal that apparent neural network advantages (dimension-independent rates, superconvergence) stem from function class complexity differences rather than architectural superiority, suggesting the 'curse of dimensionality' is better understood as a 'curse of bit complexity' governed by metric entropy.
bit complexitymetric entropyapproximation theoryneural networkscurse of dimensionality
Spatiotemporal Proximal Causal Inference under Hidden Confounding and Interference
The paper introduces a spatiotemporal proximal causal inference framework addressing hidden confounding and interference in real-world data. The method employs treatment- and outcome-inducing proxies with a spatiotemporal outcome confounding bridge function, identifiable under proxy exclusion restrictions and completeness conditions. A neural architecture combines transformer-based spatiotemporal encoders, a mutual information critic, and moment-matching networks to learn proxies and enforce constraints. Experiments on synthetic data show performance comparable to baselines while providing theoretically grounded solutions for hidden confounding under spatiotemporal interference.
proximal causal inferencehidden confoundingspatiotemporal interferencebridge functiontransformer-based encoders
Asleep at the Wheel: JEPA's Limitations in Evaluating Novel Driving Data
The study critiques the efficacy of self-supervised joint-embedding predictive architectures (JEPA) for identifying novel driving clips in autonomous vehicle datasets. Using a frozen V-JEPA video encoder and lightweight predictor head, the method reconstructs masked clip embeddings to flag hard-to-predict clips as novel. While cross-dataset evaluation suggests high effectiveness, a fair single-dataset benchmark reveals performance collapses to chance, matching no-training baselines. A lightly supervised probe on the same embeddings achieves nearly double the average precision, indicating the self-supervised objective, not the representation, is the bottleneck. This highlights how cross-dataset protocols may inadvertently reward domain separation over novelty detection.
self-supervised learningjoint-embedding predictive architecturedomain shiftvideo encodernovelty detection
Dense Language Generation Made Simple: Deterministic, Randomized, and Multi-Order Algorithms
The paper presents a unified framework for achieving optimal lower-density guarantees in language generation from positive examples. Using deterministic and randomized algorithms, the work simplifies prior analyses and extends results to multiple orders. Key results include: (1) a deterministic algorithm achieving the known optimal 1/2 lower-density guarantee, (2) a randomized algorithm achieving 1-1/e against oblivious adversaries, and (3) simultaneous optimal guarantees across multiple orders without performance loss.
language generationlower densitydeterministic algorithmrandomized algorithmoblivious adversary
Sheaf-theoretic Signal Processing on Graphs: Spectral Theory, Filtering, and Sampling
The paper introduces sheaf signal processing (SSP), a unified framework for processing heterogeneous network signals by modeling local vector spaces and linear restriction maps between them. SSP extends fundamental signal processing operations—spectral analysis, filtering, and sampling—to heterogeneous local spaces via the Sheaf Fourier Transform (SFT), polynomial sheaf filters, and a greedy sampling-set design algorithm. Representation sheaves are introduced to incorporate application-dependent signal models while preserving spectral properties. Experiments on synthetic, motion-capture, and financial datasets demonstrate consistent improvements over graph signal processing baselines.
sheaf signal processingrestriction mapssheaf fourier transformpolynomial sheaf filtersrepresentation sheaves
AlphaG-OPD: Reliability-Gated Sibling Counterfactuals for On-Policy Distillation in Symbolic Alpha Factor Discovery
AlphaG-OPD introduces a reliability-gated on-policy distillation framework for symbolic alpha factor discovery, addressing the lack of structural decision labels in existing methods. The framework comprises three components: (1) identifying grammar-valid sibling actions at partial AST states, (2) reliability gating via KL-bounded targets based on winner agreement and empirical LCB, and (3) consolidating targets through bounded replay and forward-gradient balancing. Evaluated across China's CSI300, CSI500, CSI1000, and the U.S. S&P 500, AlphaG-OPD demonstrates robust cross-market performance across multiple random seeds.
on-policy distillationsymbolic alpha factorreliability gatingabstract-syntax-treeempirical lcb
UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction
The paper introduces UDT, a U-Net diffusion transformer that reconciles the strengths of DiTs and U-Nets through data-adaptive token merging for downsampling/upsampling while preserving token dimensions. The architecture combines DiTs' representation power with U-Nets' encoder-decoder benefits, avoiding inefficiencies from standard U-Net spatial operations. UDT outperforms existing U-Net DiTs and matches REPA performance across model sizes, achieving 40x faster convergence than SiT (7.9 FID in 40 vs 1400 epochs) for XL models on 256x256 ImageNet. With CFG, it reaches 1.38 FID (SD-VAE, 320 epochs) and 1.35 FID (VA-VAE, 500 epochs).
diffusion transformerstoken mergingu-netrepresentation alignmentfid
FedChronos: Federated Fine-Tuning of Time-Series Foundation Models for Privacy-Preserving Commodity Price Forecasting
FedChronos introduces federated fine-tuning of time-series foundation models (TSFMs) using Low-Rank Adaptation (LoRA) on the Chronos-T5 backbone, addressing privacy-preserving commodity price forecasting in institutionally fragmented settings. The method employs FedAvg and FedProx for distributed training, transmitting only lightweight adapter weights (384 KB per round), achieving an 86× reduction in communication overhead. Experiments on daily commodity prices from 15 Indian agricultural markets demonstrate that differential privacy (DP) noise acts as implicit regularization, reducing mean absolute percentage error (MAPE) by 31% over zero-shot and 26% over traditional baselines while ensuring per-round (ε, δ)-DP. The compact model and small updates make it suitable for edge AI deployments.
federated learninglow-rank adaptationtime-series foundation modelsdifferential privacycommodity price forecasting
Active Regression for Single-Index Models with Unknown Link Functions
The paper advances active regression for single-index models with unknown 1-Lipschitz link functions under general ℓ_p-loss, addressing a gap in prior work limited to p=2. A non-adaptive sampling algorithm is proposed, achieving a (1+ε)-approximation with O(d^(p/2∨1)/ε^(p∨2) poly log(n/ε)) queries for p≥1. Nearly tight lower bounds are established for p>2, closing much of the remaining gap in active ℓ_p-regression for this class of models.
active regressionsingle-index modelsℓ_p-lossnon-adaptive samplinglipschitz link function
Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents
Router-Mem introduces an evidence-conditioned progressive execution framework for LLM agents, combining low-cost retrieval with a lightweight sufficiency router for early termination. The method employs evidence-level supervision and rationale-conditioned representation distillation, reusing retrieval hits for deeper analysis when evidence is insufficient. On AMA-Bench and BEAM, Router-Mem achieves 55.17% and 38.77% scores while reducing inference time by 27.3% and 25.5% compared to full memory execution.
llm agentsmemory systemsprogressive executionevidence-conditionedretrieval augmentation
Training nGPT
The paper introduces a training recipe for normalized Transformer (nGPT), enabling hyperspherical representation learning through unit hypersphere constraints on parameters and activations. Key techniques include Logit Gradient Preconditioning, Logarithmic Learning Rate Decay, GatedAdamW, angular update control, and optional exploration mechanisms. Evaluated on hybrid Mamba-2--Transformer Mixture-of-Experts (MoE) models up to 14B parameters, nGPT achieves equivalent validation loss to unnormalized counterparts using approximately 50% fewer training tokens, demonstrating scalable efficiency.
normalized transformerhyperspherical representationmixture-of-expertslogit gradient preconditioninggatedadamw
Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design
The paper introduces Riemannian attention mechanisms for Transformers to address representational rank decay in Euclidean self-attention. By replacing flat metrics with learned per-token Riemannian metrics, the authors (1) prove Riemannian attention scores are non-Gram and cannot factorize as QK^T, (2) show low-rank metric factors enable tractable geodesic distance (O(d*r)) and inversion (O(d*r^2)), and (3) propose the Fiber Bundle Transformer, a full architecture with metric-preconditioned updates and curvature proxies. Theoretical analysis confirms feasibility but leaves rank collapse prevention as an open problem.
riemannian attentionfiber bundle transformerrank collapsegeodesic distancemetric preconditioning
Latent Softmax for Data-Efficient Phoneme-Based Multilingual ASR Across Tonal and Non-Tonal Languages
Latent Softmax is introduced as a CTC-compatible output layer for phoneme-based multilingual ASR, addressing the supervision granularity mismatch between tonal and non-tonal languages. It models tone-marked vowels as subclasses and base vowels as major classes, marginalizing latent tone-marked vowels when only base-vowel labels are observed. Experiments on AISHELL-1 Mandarin and LibriSpeech English demonstrate reductions in phoneme error rates by 8.4%, 17.5%, and 12.6% on respective datasets. Additionally, Latent Softmax improves word error rates for phoneme-to-grapheme conversion and reduces mixed error rates by 2.6% and 9.5% on code-switching datasets after adaptation.
latent softmaxctc-compatiblephoneme-based asrsupervision granularitycode-switching adaptation
How fine a change can moments see? A scale law for detecting distribution shift, with a kernel calibration rule
The article establishes a scale law governing moment-based detection of distribution shifts in high-dimensional embedding streams, proving that certifying features at scale ε with mass fraction f requires polynomial tests of degree N* ≥ log(1/f)/(2ε). The law is validated via Chebyshev extremal problems and Gauss-quadrature, showing detection cost depends on feature fineness, not count. Empirical results on real data show median σ*/ε = 1.12 (IQR 1.01-1.52), with RBF kernel MMD tests achieving AUC ≥ 0.95 when bandwidth matches feature scale. Topological methods (e.g., persistent homology) underperform, with 116x higher cost than kurtosis-based tests.
scale lawdistribution shiftmmd testpersistent homologyrbf kernel
Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models
Fisher-Projected On-Policy Distillation (FP-OPD) improves vision-language model distillation by targeting only teacher corrections realizable within the student's capacity. The method estimates the student's local visual tangent space via continuous perturbations, projects the teacher-student log-probability gap onto this space under the Fisher metric, and optimizes with reverse KL on student trajectories. In 8B-to-2B distillation, FP-OPD boosts seven multimodal benchmarks by 2.77 points over the pretrained student and 1.60 points over standard on-policy distillation, demonstrating superior alignment with student capabilities.
on-policy distillationvision-language modelsfisher metricreverse klcapacity-aware distillation
AdaHAT: Adaptive Hard Attention to the Task in Task-Incremental Learning
The paper proposes AdaHAT (Adaptive Hard Attention to the Task), a novel architecture for task-incremental learning that addresses catastrophic forgetting and network capacity limitations in long task sequences. AdaHAT extends Hard Attention to the Task (HAT) by introducing an adaptive attention mechanism that dynamically updates static parameters based on their importance to previous tasks and current network capacity. Experiments on multiple datasets demonstrate AdaHAT's superior average performance over baselines, particularly in long task sequences, by better balancing stability-plasticity trade-offs.
task-incremental learningcatastrophic forgettingadaptive attentionstability-plasticity trade-offnetwork capacity
RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction
RestoreKV introduces learned restoration to recover full-cache behavior under aggressive query-agnostic KV cache eviction, complementing existing selection-based methods. The approach generates a compact, context-conditioned restore cache using a few restore tokens that attend to the full KV cache in a single LoRA-adapted pass during context prefill. Trained via parameter-efficient self-distillation, it optimizes only 0.4% of parameters without task-specific tuning. Evaluated across four backbones and four long-context benchmarks, RestoreKV significantly reduces compression-induced degradation, improving 59 of 60 budget-matched settings on Qwen3-4B and achieving 86.4 RULER accuracy at 16× compression with KVzip+.
kv cacheevictionloraself-distillationcontext-conditioned
Using Non-Lipschitz Signum-based Functions for Distributed Optimization and Machine Learning: Trade-off Between Con-vergence Rate and Optimality Gap
This work investigates the trade-off between convergence rate and optimality gap in distributed optimization and machine learning algorithms, focusing on non-Lipschitz signum-based functions. The authors analyze distributed regression problems by comparing linear and signum-based approaches through extensive simulations. Results demonstrate that while signum-based functions achieve faster convergence, they introduce larger optimality gaps and steady-state residuals in discrete-time setups. These findings contribute to advancing distributed algorithms for constrained optimization and estimation tasks.
distributed optimizationnon-lipschitz functionssignum-based functionsoptimality gapconvergence rate
Amortizing the Calibration Triple: A Projection-Consistent Neural Operator for Local-Stochastic Volatility
The authors propose a projection-consistent neural operator to amortize the calibration triple for local-stochastic volatility (LSV) models, replacing slow McKean-Vlasov fixed-point iterations with a single forward pass. The method combines DeepONet and Fourier Neural Operator (FNO) architectures to enforce arbitrage constraints, Dupire consistency, and Gyöngy projection conditions via a witness-augmented residual system. Synthetic tests show 0.1-0.2pp errors versus particle methods, 36% lower local-volatility RMSE, and 98.9% latency reduction (98.5ms → 0.6ms), enabling offline solves and real-time calibration.
local-stochastic volatilitydeeponetfourier neural operatordupire equationgyöngy projection
Climate-Dyna Deep Hedging for XVAs: Model-Based Reinforcement Learning, Residual Climate HVA, and Hedge-Instrument Discovery
The paper introduces Climate-Dyna, a model-based reinforcement learning approach for residual climate hedging valuation adjustment (HVA) in trading desks. The method computes residual climate costs by comparing paired climate-on and baseline worlds, reoptimizing overlays for each hedge universe, and learning nonlinear corrections from world-model rollouts. In a semi-synthetic EU ETS study, the inherited hedge reduces mean climate charge from 1.517 to 0.906, with learned overlay further lowering it to 0.831 (vs. 0.821 exact floor). Residual Dyna achieves 93% regret reduction with 25% fewer trajectories, while adaptation from 25 transitions retains 60.7% of exact-assisted gains.
residual climate hvamodel-based reinforcement learninghedge-instrument discoverylinear-gaussian riccatiworld-model rollouts
ReBRAC-v2: The Return of the King
The paper introduces ReBRAC-v2, a modernized behavior-regularized actor-critic method for offline reinforcement learning, achieving state-of-the-art performance through systematic engineering choices. The method employs an exact-likelihood normalizing flow as the actor, combines likelihood, MSE, and MAE behavior regularization, and integrates a classification-based residual critic, staged optimization, and multi-sample test-time action selection. Evaluated on OGBench, D4RL AntMaze, and Adroit benchmarks, ReBRAC-v2 averages 74.8, 90.2, and 33.6 respectively, outperforming prior methods. Ablations highlight the importance of mixed cloning objectives, staged training, flow capacity, and multi-sample inference.
offline reinforcement learningbehavior regularizationnormalizing flowactor-criticmulti-sample inference
Fruit-HSNet: A Machine Learning Approach for Hyperspectral Image-Based Fruit Ripeness Prediction
Fruit-HSNet introduces a machine learning architecture for hyperspectral image-based fruit ripeness prediction, addressing challenges of limited labeled data and cross-camera generalization. The method combines spatio-spectral feature extraction via Fourier Transform and central pixel spectral signatures with learnable feature fusion and a specialized classifier. Evaluated on the DeepHS Fruit dataset (5 fruit types, 3 hyperspectral cameras), Fruit-HSNet achieves 70.73% overall accuracy, surpassing prior methods by 12% and setting a new state-of-the-art.
hyperspectral imagingfruit ripeness predictionspatio-spectral featuresfourier transformlearnable fusion
Hybrid Quantum Neural Networks: Theory, Implementations, and Applications
This review synthesizes progress in hybrid quantum neural networks (HQNNs), which integrate classical neural networks with quantum processing units for near-term applications. It surveys theoretical foundations, architectural variants, and empirical results, highlighting cases where HQNNs achieve competitive performance with fewer parameters despite limited scalability. The analysis identifies implementation challenges and tasks with provable quantum advantages, offering a structured assessment of current capabilities and future research directions.
hybrid quantum neural networksquantum machine learningparameter efficiencynear-term quantumprovable advantage
Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races
The study investigates strategic safety behavior in multi-agent AI development races using large language models (LLMs), employing an audit gate to validate game understanding through rule recall, state tracking, and payoff calculation. Seven LLM endpoints were tested in 2-5 player races, revealing that strong rule recall coexists with weak state tracking and payoff computation, and that response representation changes can alter actions despite fixed rules. Results show model-specific action sequences and opponent responses, highlighting the need for trajectory-level analysis in multi-agent simulations. Findings are exploratory, limited to tested models, prompts, and decoding settings.
multi-agent safetyllm behaviorgame-theory benchmarkstate trackingtrajectory-level analysis
3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
3DZip introduces a spatial-aware token compression framework for 3D vision-language models (3D VLMs) to address computational overhead from excessive tokens. The method employs three stages: coarse voxelization for point-level redundancy removal, feature-space diversity-guided anchor token selection via a Determinantal Point Process, and spatially constrained token merging. Evaluated on three 3D question answering benchmarks, 3DZip retains 94.7% of original performance with 128 tokens, achieving 1.92× faster inference compared to existing methods.
3d vision-language modelstoken compressiondeterminantal point processspatial aggregationfeature diversity
SAFE-Merge: Data-Free Continual Model Merging with General Knowledge Preservation
SAFE-Merge introduces a data-free continual model merging framework that preserves general knowledge while incorporating specialized models, without task data access. The method employs risk-aware sparse masking to select safe parameter updates and masked low-rank recovery to compensate for lost task information, maintaining inference efficiency. Evaluated on vision and language benchmarks, SAFE-Merge achieves superior H-scores and outperforms NUFILT in accuracy on longer CLIP task sequences.
continual learningmodel mergingdata-freesparse maskinglow-rank recovery
Interpretable Machine Learning for Traffic Congestion Prediction: Unveiling the Impact of Different COVID-19 Periods
This study contributes an interpretable machine learning framework for traffic congestion prediction across COVID-19 periods in Alameda County, California. The method incorporates weather, seasonality, and COVID-19 variables, employs Recursive Feature Elimination with Cross-Validation for feature selection, and compares Support Vector Regression, Multiple Linear Regression, Recurrent Neural Networks, and Long Short-Term Memory Networks. Bidirectional LSTM achieved the best performance across pre-lockdown, lockdown, and post-lockdown periods, with Integrated Gradients and SHapley Additive exPlanations providing model interpretability. Results indicate COVID-19 cases negatively impacted congestion during lockdown and post-lockdown, while higher hospitalization reduced travel and congestion post-pandemic, and fuel price increases led to private vehicle shifts and congestion growth.
bidirectional lstmintegrated gradientsshapley additive explanationsrecursive feature eliminationtraffic congestion prediction
Differentiable Lifting for Topological Neural Networks
The paper introduces $\partial$lift (DiffLift), a differentiable framework for learning graph liftings to hypergraphs, cellular complexes, and simplicial complexes end-to-end. The method leverages vertex-level latent representations to parameterize distributions over candidate higher-order cells, enabling scalable integration with topological neural networks (TNNs). Experiments demonstrate that $\partial$lift outperforms static lifting methods by up to 45% on graph and node classification benchmarks across multiple TNN architectures.
topological neural networksgraph liftinghypergraphssimplicial complexesdifferentiable learning
Learning-Based Stochastic Optimal Control with Infinite-Horizon Probabilistic Constraints
The paper proposes a learning-based approach for infinite-horizon stochastic optimal control with joint chance constraints. By augmenting the state space, the problem is reformulated as a constrained Markov decision process (CMDP) with additive cost and constraint structures, proven to exhibit strong duality. A dual-ascent algorithm converges to an optimal deterministic Markov policy, while a value function approximation reduces online computational complexity. Numerical experiments show superior performance and efficiency compared to online predictive control methods.
stochastic optimal controlconstrained markov decision processdual-ascent algorithmvalue function approximationinfinite-horizon constraints
📰 Industry Media (10)
Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks
Cursor Research introduces Mixture-of-Kittens (MoK), a deterministic mixture-of-experts (MoE) training megakernel that fuses all communication and computation steps into a single kernel for NVIDIA Blackwell GB300 NVL72 racks. MoK employs pull-based forward dispatch (18µs latency) and push-based forward combine, alongside a ring token buffer to eliminate CPU-GPU synchronization, achieving up to 2.37× speedup in MXFP8 forward passes versus baselines. Benchmarking on 512 GPUs shows 1.41× end-to-end throughput improvement (1,070.2 tokens/GPU/s). The Apache-2.0-licensed kernel requires Blackwell SM100/SM103 GPUs, CUDA 13.0+, and PyTorch 2.10+, targeting large-scale MoE training (e.g., DeepSeek-V3-style models).
mixture-of-expertsmegakernelnvlinkdeterministic trainingblackwell gpu
Reflex Open Sources XY: A Rust-Backed Super-Fast Python Charting Library That Keeps 100 Million Point Charts Interactive
Reflex AI introduces XY, an Apache-2.0 licensed Python charting library with a Rust backend for interactive 2D visualizations of large datasets. By employing WebGL2 rendering, binary data transport, and adaptive representations (M4 decimation for >10k rows, scatter density for >200k points), XY achieves consistent ~0.08s render times from 10k to 100M points, outperforming Matplotlib and Plotly by 34–177× at scale. Benchmarks show 0.32 GiB memory usage for 10M points (vs. 0.84 GiB/1.86 GiB) and 258 KiB HTML exports for 10M-point scatters (vs. Plotly's 259 MiB). The library supports 14 chart types, CSS styling, and Matplotlib compatibility, targeting domains like quantitative finance and genomics where row counts bottleneck existing tools.
webgl2m4 decimationbinary transportcolumnstoreadaptive rendering
Building an Advanced AI Skill Security Auditing Pipeline with NVIDIA SkillSpector, LangGraph, YARA Rules, SARIF, and CI Policy Gates
The tutorial presents a security auditing pipeline for AI skills using NVIDIA SkillSpector, LangGraph, and YARA rules. It constructs a synthetic skill marketplace with clean, risky, malicious, and MCP-based examples, then evaluates them via a LangGraph inspection pipeline to generate risk scores, categorized findings, and SARIF/Markdown reports. The method includes baseline suppressions, regression detection, and custom YARA rule integration, achieving granular security analysis with optional LLM-assisted semantic evaluation. Results show risk distributions, analyzer completeness, and executable-script detection across the skill portfolio.
skillspectorlanggraphyara rulessarifci policy gates
Y Combinator Open-Sources QM: An MIT-Licensed Multiplayer Agent Harness That Runs In Slack And The Web
Y Combinator has open-sourced QM (Quartermaster), a multiplayer agent harness designed for organizational workflows, operating in Slack and web environments. QM provides isolated workspaces for individual users and collaborative channels, leveraging scoped memory, permissions, and sandboxed execution. The architecture features a central headless core handling API, identity, and scheduling, with Postgres for durable state storage. QM supports multiple AI models (Pi, OpenCode, Codex, Claude Code) to avoid vendor lock-in and enforces security postures (Strict, Auto, Dangerous) with predefined command policies. Deployed via cloud infrastructure, QM targets startups to mid-sized companies (10-500 employees) and is currently used internally by Y Combinator across accounting, legal, events, and engineering.
multiplayer agent harnessscoped memorysandboxed executionvendor lock-insecurity postures
Genspark Open Sources GenOffice: A Free, Ad-Free AI Office Suite for macOS and Windows with Docs, Sheets, Slides, PDF
Genspark released GenOffice v0.4.110, an open-source AI-native office suite (Apache 2.0) for macOS and Windows, comprising Docs, Sheets, Slides, and PDF tools. The system employs byte-preserving document editing: Docs parses .docx into a block tree (anchored by docxIndex) and patches only modified blocks via TipTap, ensuring layout fidelity in Word. Sheets integrates Univer core with Rust-based XLSX handling (calamine, IronCalc), while Slides uses in-house pptx rendering with HarfBuzz text shaping. AI features require Genspark account credits and proxy model calls server-side. Security measures include Electron renderer lockdown (contextIsolation, sandboxing) and constrained AI-content execution (AST interpreter, no eval).
docxindextiptapuniverharfbuzzast interpreter
Evaluating Multimodal Vision Models with Moonshot PerceptionBench Using Robust Data Loading and Automated Judging
The authors present PerceptionBench, a multimodal benchmark for evaluating vision models across 10 fine-grained capabilities including OCR, counting, and hallucination detection. They implement a robust evaluation pipeline featuring balanced dataset sampling, image normalization, and automated judging with both rule-based and LLM-assisted scoring. The system supports multiple backends (blind prior, API-based, local Hugging Face models) and produces bootstrap confidence intervals for capability profiles. Initial analysis shows median image resolution of 1024px and coverage of 40% newly-authored tasks alongside decomposed benchmarks.
multimodal benchmarkvisual perceptionautomated judgingcapability profilingbootstrap confidence
How to Secure AI Agents, MCP Servers, and LLM Apps in Production
Mend.io introduces a practical framework for securing AI agents, MCP servers, and LLM applications in production, addressing emergent risks in agentic AI systems. The framework organizes security measures into three phases: discovery, prioritization, and runtime protection, supported by seven reusable artifacts. It identifies five risk layers—interaction, agent, integration, model, and code—and emphasizes automated, evidence-backed triage while reserving novel findings for human review. Runtime guardrails are implemented via Python SDK or Docker API Server, focusing on prompt hardening and policy enforcement. The framework aligns with NIST AI RMF, OWASP AIMA, ISO/IEC 42001, and the EU AI Act.
agentic aimcp serversruntime guardrailsprompt hardeningevidence-backed triage
Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model and the Most Capable One in the Qwen Family to Date
Alibaba Qwen introduces Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts (MoE) model supporting multimodal inputs (text, image, video) with a 1M-token context window. The model achieves 86.6 on Terminal-Bench 2.1, outperforming Claude Opus 4.8 (84.6) but trailing GPT-5.6 Sol (88.8), and shows strong multimodal performance (e.g., 92.1 on OmniDocBench 1.5). Deployment options include a hosted API ($2/$6 per 1M input/output tokens) and open-weight releases (Qwen3.8-Max and Qwen3.8-27B), with the latter targeting on-premise GPU hardware. Key capabilities include function calling, structured outputs, and built-in tools like code_interpreter and web_search.
mixture-of-expertsmultimodalcontext windowstructured outputsterminal-bench
Cogent AI Team Releases VR-1: A Frontier Cyber Reasoning Model That Composes and Verifies Enterprise Attack Paths
Cogent AI introduces VR-1, a cybersecurity-specific reasoning model trained to compose and verify multi-domain enterprise attack paths, addressing limitations of general coding models in intrusion execution. The model operates under partial information, composes evidence across domains, and recovers from dead ends, evaluated via IntrusionBench with execution-based verification. VR-1 demonstrates 2× higher black-box pass@3 success rate at 25% cost compared to Kimi K3, Claude Opus 4.8, and GLM-5, though its absolute success rate remains under 30%.
cyber reasoning modelintrusionbenchblack-box pass@3multi-domain attack pathsenterprise cybersecurity
Onton Releases Ontology 1: A Neurosymbolic Search Model That is 2.7x More Accurate than the World’s Best E-commerce Search Engines
Onton introduces Ontology 1, a neurosymbolic search model achieving mean precision@10 of 0.630 on the Subtext-Decor-90 benchmark, outperforming Google Shopping (0.543) and Amazon (0.469) despite indexing 1% of their catalogs. The model combines inspectable knowledge graphs with continuous self-learning, decomposing vague queries (e.g., 'pet-friendly') into verifiable product properties. Evaluated by three LLM judges (Claude Opus 4.8, Gemini 3.1 Pro, GPT-5.5), Ontology 1 won 52/90 queries but struggled on functional-spec queries where metadata dominates. Judge agreement was modest (Krippendorff’s alpha 0.465). Deployment is restricted to Onton.com and partner access.
neurosymbolicprecision@10knowledge graphself-learning loopkrippendorff’s alpha
Generated automatically at 2026-08-04 21:00 UTC. Summaries and keywords are produced by an LLM and may contain inaccuracies — always consult the original article.
