Daily Digest — 2026-07-24
209 items · 3 research labs, 205 arxiv papers, 1 industry media
MarkTechPost: all feed URLs failed (last tried: https://www.marktechpost.com/feed/)AI News: all feed URLs failed (last tried: https://artificialintelligence-news.com/feed/)
🏛️ Research Labs (3)
Launching Health in ChatGPT
OpenAI introduces Health in ChatGPT, enabling U.S. users to securely integrate Apple Health and medical records for personalized health insights. The feature leverages GPT-5.5 Instant and GPT-5.6 Sol models, optimized for complex health reasoning and clear communication, achieving superior performance on HealthBench Professional evaluations. Privacy safeguards include encryption, user-controlled data access, and exclusion from model training or ad targeting. Early testing showed 70% of health-related conversations occurred outside a dedicated health space, prompting integration across general ChatGPT interactions. Physicians validated model performance and safety, though ChatGPT remains supplementary to professional medical care.
gpt-5.6 solhealthbench professionalin-context learningencryption safeguardsmodel-training exclusion
NTT DATA Group cuts incident analysis to 30 minutes with Codex
NTT DATA Group deployed OpenAI's Codex to 9,000 employees, achieving a 99.3% reduction in incident analysis time (from 3 days to 30 minutes) through automated task execution. The company established an OpenAI Center of Excellence to guide enterprise-wide adoption, implementing security guidelines and sandboxing for safe usage. Results include 96% employee satisfaction with ChatGPT Enterprise, 1.4× increase in weekly Codex users post-training, and automation of internal operations via Playwright scripts. The study demonstrates how agentic AI can transform both technical and nontechnical workflows when integrated with proper governance.
codexchatgpt enterpriseagentic capabilitysandbox modeplaywright
Bringing Nunchaku 4-bit Diffusion Inference to Diffusers
The Hugging Face Diffusers library now supports Nunchaku Lite, enabling 4-bit weight and activation (W4A4) inference for diffusion models via SVDQuant quantization. This method reduces memory usage by 50% and improves latency by 30% compared to BF16 baselines, achieved through low-rank correction and 4-bit residual quantization. The integration allows native loading of quantized checkpoints without custom pipelines, supported by NVFP4/INT4 kernels for NVIDIA GPUs. Benchmarks on an RTX PRO 6000 show 1.7s generation times for 1024x1024 images with 12GB VRAM. The diffuse-compressor toolkit facilitates model quantization and Hub deployment.
svdquantnunchaku litew4a4diffusersnvfp4
📜 arXiv Papers (205)
SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data
SoftReason introduces a fully differentiable neuro-soft-symbolic architecture for deductive reasoning over high-dimensional perceptual data and knowledge graphs. The method replaces discrete interfaces with a soft interpretation tensor, enabling end-to-end differentiability through a learned lift of the immediate-consequence operator using predicate-definition embeddings and latent composition channels. Demonstrated on Knowledge-aware Visual Question Answering (KVQA), it integrates perceptual grounding, KG evidence injection, and differentiable deductive closure in a unified framework.
neuro-symbolicdifferentiable reasoningknowledge graphimmediate-consequence operatorperceptual grounding
Persian Pixel: A large-scale synthetic OCR dataset for Persian language
The paper introduces Persian Pixel, a large-scale synthetic OCR dataset addressing the scarcity of annotated Persian text data, which hinders OCR development for the Perso-Arabic script. The dataset comprises 343,000 high-fidelity image-text pairs generated from a seven-million-word Persian corpus using the SynthOCR-Gen framework, modeling typographic characteristics like contextual character joining, glyph variants, and diacritic placement. Realistic document acquisition artifacts are simulated through over twenty-five stochastic degradation models. Persian Pixel enables scalable training of modern OCR architectures, such as TrOCR and Donut, and advances research in Persian document analysis and digitization, demonstrating the efficacy of synthetic data generation for low-resource scripts.
optical character recognitionperso-arabic scriptsynthetic datasetstochastic degradationtransformer-based models
FMRP-LEAN: A HIPAA-Compliant AI-Augmented LIMS Architecture for End-to-End Clinical Assay Workflow Optimization
FMRP-LEAN introduces a HIPAA-compliant AI-augmented LIMS architecture for clinical assay workflows, addressing challenges in multi-day assays like FMRP quantification. The system combines a finite-state workflow model with Supabase/PostgreSQL infrastructure, hybrid edge-internal isolation, and REDCap synchronization, using MRN-UUIDv7 identifiers for traceability. It features automated statistical QC pre-screening and governance-constrained AI operations on aggregate projections. Deployment results show improved workflow observability, reduced QC latency, and enhanced cross-role transparency in regulated healthcare environments.
limshipaa-compliantfinite-state workflowuuidv7redcap
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
The paper introduces RECAP (Readable Encodings via Co-trained Auxiliary Predictors), a method to ensure verifiable activation explanations by training linear heads alongside target models to maintain decodable content. It exposes flaws in reconstruction-based faithfulness tests, showing they tolerate false claims (only ~2% reconstruction-dependent) and co-adapted private codes. RECAP eliminates these codes at minimal cost (+0.001 nats) and enables independent probe verification (AUC 0.96 vs. 0.82 baseline). Evaluations on Pythia-160M and synthetic tasks demonstrate RECAP's robustness against adversarial edits (AUC 0.95 vs. 0.51 control).
autoencodersactivation explanationsdecodability supervisionco-adapted codeslinear probes
Generative AI floods and dilutes the market for books
This study quantifies the market impact of generative AI in self-published genre fiction by analyzing 14,419 books sold on Amazon from 2023 to 2026. Using full-text AI detection, the authors matched sales records to AI-generated content levels, finding that books with substantial AI text (>25%) constitute a large catalog share but smaller sales share, though their commercial impact grows over time. AI-generated books increasingly occupy top-rank positions, diluting revenue per book as catalog growth (19.2-fold) outpaces revenue growth (8.9-fold). Top-selling AI books exhibit higher linguistic overlap with existing works, suggesting generative AI reshapes markets through scale rather than quality, particularly in genres with high AI diffusion and Kindle Unlimited availability.
generative aiself-publishingmarket dilutionkindle unlimitedcopyright infringement
Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids
The paper introduces DEED, a systems-level framework for improving real-world performance of Vision-Language-Action (VLA) humanoid robots in retail settings. The method combines data-efficient post-training (control-frequency alignment, task-relevant visual highlighting) with experience-driven refinement (text-based advantage prefix, vision-language value function) and latent-space analysis tools. Evaluated on a supermarket chip-restocking task using Unitree G1-Edu and GR00T N1.6, DEED demonstrates that careful data design and targeted post-training can enable competent real-world operation with minimal compute (single GPU), framing the lab-to-store gap as primarily a systems integration challenge.
vision-language-actionpost-trainingexperience-driven learninglatent-space analysissystems integration
Understanding Generative AI-mediated User Engagement with Academic Library Resources
This study empirically demonstrates the impact of generative AI as a discovery pathway for academic library resources, revealing a significant increase in AI-mediated traffic post-integration of linked citation features. Using web analytics from August 2023 to October 2025, the research identifies ChatGPT, Perplexity, and Gemini as primary platforms driving traffic, with users predominantly accessing electronic theses and dissertations in institutional repositories. Findings suggest AI retrieval mechanisms effectively surface resources with structured metadata, stable permalinks, and Open Access availability, highlighting the need for strategic responses to evolving AI ecosystems.
generative ailinked citationstructured metadataopen accessinstitutional repository
Toward Reliable RGB-D Semantic Segmentation: Handling Missing Modalities via Condition Dropout
The paper introduces Condition Dropout (ConD), a continued-training paradigm to improve RGB-D semantic segmentation robustness under missing modalities. ConD extends a pretrained RGB-D model by simulating missing-modality inputs during a second training stage, freezing original encoders while training copied encoders with zero-initialized feature injection. Evaluated on NYU-Depth V2 and SUN RGB-D, ConD maintains full-modality accuracy while significantly improving performance when either RGB or depth is missing, with slight gains observed in complete-modality scenarios.
rgb-d segmentationmissing modalitiescondition dropoutfeature injectioncontinued training
Don't Trust the Label: License Laundering in AI Supply Chains
The study introduces the concept of license laundering in AI supply chains, analyzing how license obligations propagate across 232,270 dataset→model→application chains. Using a multi-platform tracing approach, it quantifies two laundering forms: undeclared licenses acquiring labels downstream and license category replacement during redistribution. Results show 62.3% of chains involve at least one artifact with no declared license (primarily foundational datasets), while obligation-bearing licenses exhibit <7% end-to-end survival versus 95.1% for Permissive licenses. The work concludes with recommendations for stakeholders to address these compliance gaps.
license launderingai supply chainlicense propagationredistribution complianceobligation-bearing licenses
Courteous Anticipation: Improving Long-Lived Task Planning in Persistent Shared Environments
We introduce courteous anticipatory planning for multi-robot task scheduling in persistent shared environments, where robots must anticipate how current actions impact future tasks. The method employs a model-based planner that selects plans minimizing both immediate cost and aggregated expected future cost across all robots, estimated via independent per-robot learned estimators. This factored formulation avoids combinatorial joint rollouts and supports modular deployment. Evaluations in two PDDL domains show cost reductions of 10.43% versus myopic and 4.03% versus selfish anticipatory planning in a two-robot home environment, and 17.41% and 13.24%, respectively, in a three-robot restaurant environment.
task planningpersistent environmentsmodel-based plannercost minimizationmodular deployment
Sound Probabilistic Safety Bounds for Large Language Models
The authors introduce a framework for computing rigorous probabilistic safety bounds on harmful outputs from large language models (LLMs). Their method applies Clopper-Pearson confidence intervals to derive probably approximately correct (PAC) bounds, with a novel algorithm that prioritizes exploration of high-risk branches in the auto-regressive generation tree using latent space features. Experiments demonstrate non-trivial lower bounds on harm probability for state-of-the-art LLMs, providing statistically sound certification of model safety.
probabilistic safety boundsclopper-pearson intervalsauto-regressive generationlatent space featurespac bounds
Self-supervision drives representational convergence in medical foundation models more than clinical supervision
The study demonstrates that self-supervision, not clinical supervision or scale, drives representational convergence in medical foundation models. Using 18 image and 7 text encoders (7M to 27B parameters) across five imaging modalities and 650,982 chest radiographs, the authors isolate the effect of pretraining objectives under fixed data and architecture. Self-supervised encoders showed the highest alignment (40.4%), outperforming label-supervised (21.1%) and image-text (3.3%) models, with no significant correlation to model size (Spearman 0.302, p=0.223). Linear classifiers transferred well across encoders (85% performance retention), suggesting interoperability depends on pretraining design rather than scale or clinical labels.
representational convergenceself-supervisionmedical foundation modelslinear transferpretraining objectives
PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity
PoTRE (Poly-Topological Reasoning Ensembles) introduces a heterogeneous framework for enhancing LLM reasoning by decoupling inference into four specialized agents: Adversarial Refinement, Hierarchical Strategic Planning, Spectrum Search, and Direct Chain. These agents are dynamically reconciled via a Task-Adaptive Aggregation Layer, which performs final candidate selection, semantic synthesis, or neuro-symbolic verification. Evaluated on ARC-AGI-2, Humanity's Last Exam (HLE), and PRBench Finance, PoTRE achieves state-of-the-art accuracy of 49.92% on HLE, outperforming previous benchmarks while maintaining or reducing inference token usage compared to homogeneous baselines.
poly-topological reasoningtask-adaptive aggregationneuro-symbolic verificationadversarial refinementhierarchical planning
The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models
The paper introduces the Maskability Index (MI), a metric predicting optimal prompting strategies for relational knowledge extraction from pretrained language models. MI quantifies alignment between task objectives and prompting styles (masked vs. prefix) by comparing DepthRank scores across template variants. Evaluated on ATOMIC2020 relations, MI correlates positively with downstream generation performance, demonstrating utility for template selection in few-shot settings. Results suggest MI aids adaptation of models like T5 and BERT for structured knowledge tasks with limited data.
maskability indexdepthrankfew-shot generationpretrained language modelsknowledge base completion
The Ethics of Autonomous AI Agents for Offensive Security
The article analyzes ethical implications of LLM-driven autonomous agents in offensive security, identifying three dimensions of indeterminacy: non-deterministic policy outputs, open-ended impact due to opaque LLM supply chains, and reduced skill requirements for deployment. It examines how moral attribution diffuses among users, tool-makers, and third parties, leveraging structural cost asymmetry between offense and defense. The authors provide stratified recommendations, arguing that existing dual-use frameworks inadequately address these challenges, with short-term effects favoring attackers despite potential long-term defensive benefits.
autonomous agentsoffensive securityllm-drivennon-deterministic policydual-use frameworks
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
The study introduces a unified framework for full-song generation supporting Lyrics-to-Song, Instrumental Music, and Cover Song Generation tasks. The architecture combines a semantic-aware tokenizer (8-codebook RVQ tokens), hybird-LM for hierarchical autoregressive token modeling, FullDiT for flow matching in VAE latent space, and a melody module for cover song preservation. Reward-based post-training (DPO, GRPO, OPD) enhances musicality. Evaluations on a multilingual benchmark and Artificial Analysis Music with Vocals leaderboard demonstrate competitive performance.
hierarchical autoregressiveflow matchingrvq tokenscover song generationreward-based post-training
On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens
The study systematically investigates culturally loaded machine translation challenges using LLMs, constructing a 500-segment Chinese-Japanese dataset from Dream of the Red Chamber. Evaluation reveals three key issues: (1) performance gaps in frontier LLMs (e.g., GPT-4) on culturally loaded content, (2) substantial human evaluator disagreement due to cultural backgrounds, and (3) unreliable automatic metrics for quality assessment. These findings provide empirical insights for culture-oriented MT research.
culturally loaded translationmachine translationlarge language modelshuman evaluationautomatic metrics
DQAOA-GPT: AI-Accelerated Distributed Quantum Optimization for Combinatorial Problems
The paper introduces DQAOA-GPT, a hybrid framework combining distributed quantum optimization with GPT-based circuit generation for combinatorial problems. The method decomposes large problems via the distributed quantum approximate optimization algorithm (DQAOA) and uses a trained GPT model to directly generate quantum circuits for sub-problems, bypassing iterative variational optimization. Evaluated on dense HUBO problems with ≤100 variables, DQAOA-GPT reduces computational costs while maintaining solution quality, with greater acceleration for larger sub-problems.
distributed quantum optimizationquantum approximate optimization algorithmcombinatorial optimizationgenerative circuit synthesishybrid hpc-qc
Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis
The paper demonstrates that orchestrated ensembles of small language models (SLMs) can outperform single large language models (LLMs) in malware analysis tasks. Four orchestration architectures were evaluated on Meta's CyberSecEval benchmark: multi-agent pipeline, adversarial debate, hierarchical consultation, and a hybrid system. The hybrid architecture (Qwen3-4B with Foundation-Sec-8B) achieved 35.30% accuracy, surpassing both cyber-specialised baselines (22.54%) and ungrounded frontier LLMs (34.77%). Evidence-grounded pipelines were critical for performance, with the best configuration reaching 38.22% accuracy using grounded Gemini.
small language modelsmalware analysisorchestration architecturesevidence-grounded reasoningcyberseceval benchmark
ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
The paper introduces ELSAA, an efficient approximation for Transformer attention that combines low-rank and sparse branches without decomposing projection matrices. The method approximates the attention score operator directly: a sparse branch captures high-similarity interactions, while a low-rank branch compresses global context. A denominator-aware fusion term balances the branches by scaling sparse attention mass relative to low-rank. This avoids materializing the full quadratic score matrix, enabling longer-context training while preserving local and global interactions.
efficient attentionlow-rank approximationsparse attentiontransformersdenominator-aware fusion
The Quadrilateral Loss: Additivity as a Measurable Behavior of Dense Neural Networks
The paper introduces the quadrilateral loss, a differentiable penalty that quantifies additivity in neural networks by measuring second-order mixed differences between coordinate swaps. This approach treats additivity as an observable behavior rather than an architectural constraint, enabling flexible control via regularization. Experiments show that learned interactions are often removable with minimal performance cost, and pre-regularization interaction magnitudes poorly predict post-regularization retention. The work compares structural and behavioral routes to exact additivity, finding behavioral constraints dominate weight-space methods, with convergence across approaches on shape functions.
quadrilateral lossadditive modelsinteraction massshapley-gambehavioral regularization
StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation
StreamHOI introduces a low-latency streaming framework for long-duration human--object interaction (HOI) video generation, addressing limitations of offline methods. The approach adapts historical memory organization in transformer blocks via HOI-aware profiling and bias-guided memory-specialized training, optimizing interaction preservation under latency constraints. A memory distance scaling module enhances access to early interaction states. Evaluations show StreamHOI achieves 17.6 FPS with 0.75s first-chunk latency, outperforming baselines in interaction plausibility, object fidelity, and human quality.
streaming video generationhuman--object interactionmemory adaptationtransformer blockslatency optimization
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
The paper introduces Audio-Zero, a label-free self-evolution framework for improving fine-grained audio reasoning in Large Audio Language Models (LALMs). The method constructs an auditory self-play game using unlabeled audio contrast pairs, where models generate descriptive clues and identify an odd listener through inconsistency reasoning, providing verifiable rewards without external labels. Evaluations on Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B using TREA, MMAU Test-mini, and MMAR benchmarks demonstrate enhanced fine-grained reasoning while maintaining broad audio understanding, with evolutionary analyses revealing emergent fine-grained descriptions.
large audio language modelsself-play gamefine-grained reasoninglabel-free learningauditory perception
Active Inference as a Convex Markov Decision Process
The paper reformulates Active Inference (AIF) as a convex Markov Decision Process (MDP), demonstrating that expected free energy (EFE) minimization combines linear pragmatic terms (equivalent to reward maximization) with nonlinear epistemic components. By analyzing finite-horizon, discounted, and average-reward EFE formulations, the authors derive a mirror descent algorithm that linearizes the objective around state marginals, enabling compatibility with actor-critic methods. This bridges AIF with modern reinforcement learning theory, offering convergence guarantees and performative reward structures.
active inferencemarkov decision processexpected free energymirror descentperformative reinforcement learning
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
The work presents SLAI T-Rex, a hierarchical optimization framework for full-parameter post-training of trillion-parameter MoE models on Ascend NPU SuperPOD, targeting the DeepSeek-V4 family. The system employs model-level parallelism, computation-communication orchestration, and kernel optimizations to achieve 34.22% MFU (2.93× baseline improvement). A specialized workflow produces DeepSeek-V4-Flash for Operations Research, trained on 10K solver-verified SFT samples, achieving 71.81% zero-shot Pass@1 (outperforming GPT-5.4-Mini by 3.98pp).
mixture-of-expertsmodel flops utilizationascend npuzero-shot evaluationoperations research
Formal Foundations for Known Good Reliable Die Screening in Chiplet-Based AI Systems-on-Chip
The authors formalize Known Good Reliable Die (KGRD) screening for chiplet-based AI SoCs by addressing pre-assembly observability limitations through four contributions: a Bayesian probabilistic risk model mapping pre-assembly telemetry to post-assembly failure likelihood, a safety-gated decision architecture with provable failure probability guarantees, Bayes-optimal uncertainty-aware disposition boundaries, and a constrained closed-loop feedback mechanism for model improvement. These components were validated via Monte Carlo simulations on 4,000 synthetic dies, confirming uniform safety guarantees across tested gate thresholds.
bayesian probabilistic modelchiplet-based socsmonte carlo simulationfailure probabilityclosed-loop feedback
PRIME-SVR: Physics-infoRmed Implicit Multi-Echo Slice-to-Volume Reconstruction for Fetal T2 mapping
PRIME-SVR introduces an implicit neural representation framework for joint high-resolution reconstruction from multi-echo fetal MRI, enabling quantitative T2 mapping. The method employs two networks: one modeling continuous signal intensities across echo times (TEs) and another estimating slice-specific degradations, with Bloch equation-derived regularization enforcing cross-TE coherence. Evaluated on 39 in vivo fetal acquisitions, PRIME-SVR improves reconstruction sharpness by 47%, anatomical accuracy by 30%, and structural consistency by 14% over state-of-the-art methods, while reducing acquisition time from 15 to 5-10 minutes with T2 accuracy within 2.3%.
implicit neural representationslice-to-volume reconstructiont2 mappingbloch equationmulti-echo mri
CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning
The paper introduces MGT-B (Monitoring-Guided Test-time Backtracking), an inference-time controller for quantized small autoregressive models that detects and corrects degenerate trajectories. The method uses overlapping windows of pre-sampling uncertainty and degeneration features to compute empirical tail probabilities, accumulates mixture betting factors with a CUSUM-shaped reset, and triggers constrained re-decoding upon alarm. On a 240-pair chronology-audit set, accuracy improved from 82/240 to 88/240 (+2.50 percentage points), though not statistically significant (McNemar p = 0.2632). A broader 467-pair set showed a more significant improvement (146/467 to 167/467, p = 0.000753), but with potential selection bias.
inference-time monitoringquantized modelscusum controllerre-decodingkey-value-cache
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
The paper introduces ENTRAP-VL, a manually curated dataset of 1,500 items designed to probe dual contextual entrainment in vision-language models (VLMs). Unlike unimodal language models, VLMs exhibit entrainment driven independently by textual and visual context, necessitating a taxonomically structured instrument with eight textual and three visual context conditions. The dataset, organized by association and truth axes, enables rigorous evaluation of entrainment phenomena without claiming specific model measurements.
contextual entrainmentvision-language modelstaxonomic probedual-modalityveracity distinction
Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results
The paper introduces SelectBench, a benchmark for selective evidence adoption in retrieval-augmented LLMs, and proposes Direct Alignment Preference Optimization (DAPO) to train Qwen3.5-4B using rule-based or semantic-judge rewards. On SelectBench-v2 (325 examples), DAPO variants improve strict success rates from 22.46% (baseline) to 25.54% (DAPO-Rule) and 26.46% (DAPO-DeepSeek), reducing forbidden-content adoption while preserving performance on MMLU and HotpotQA. Gains are statistically insignificant after Holm correction, highlighting persistent challenges in prompt-injection resistance and reward shaping.
retrieval-augmented generationdirect alignment preference optimizationselective evidence adoptionprompt-injection resistancereward shaping
Co-Evolving LLM Evaluators and Policies via DynamicRubric
The paper introduces DynamicRubric, a co-evolution framework for LLM evaluators and policies that addresses collapsed score gaps during policy optimization. The method generates weighted binary rubric items conditioned on response sets, aggregating judgments into response-level scores to strengthen supervision signals. Experiments with 8B models show DynamicRubric outperforms baselines using 70B reward models or 235B static rubric generators, with deployed improvements in WeChat Search handling millions of daily requests and gains on reasoning/coding tasks.
policy optimizationevaluator feedbackrubric generationscore gappost-training
TRUST-ESD: A Risk-Calibrated and Governance-Aware AI Framework for Enterprise Strategic Decision Support Under Uncertainty
TRUST-ESD introduces a risk-calibrated, governance-aware AI framework for enterprise strategic decision support under uncertainty. The method integrates predictive utility estimation, conformal uncertainty calibration, CVaR-based downside-risk scoring, risk-memory retrieval, policy-as-code governance, and explainability. Results demonstrate 7.95% higher risk-adjusted utility, 23.22% lower risk exposure, 23.78% reduced CVaR, 13.89% lower calibration error, 10.90% improved explanation fidelity, and 9.76% higher governance compliance versus baselines, while maintaining predictive accuracy.
conformal uncertainty calibrationcvar-based scoringrisk-memory retrievalpolicy-as-code governanceexplanation fidelity
PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning
(No summary returned.)
Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
The study demonstrates that materials science mechanisms in the Google/Gemma-4-E4B-it language model exhibit three distinct forms: concept readability in hidden states, constitutive orientation via state transformations, and causal control over engineering answers. Using Jacobian vocabulary readouts, state geometry analysis, a 60-law counterfactual benchmark, and causal interventions, the authors show that hidden-state transformations correctly orient 39 of 40 directional laws, while lexical controls perform near chance. Bidirectional interventions shift answer probabilities toward physically appropriate outcomes in all 12 matched cases, revealing that physical relationships are more evident in controlled state changes than in absolute states alone.
hidden statesjacobian readoutscausal interventionsconstitutive orientationcounterfactual benchmark
Language-Specific versus Cross-Lingual Knowledge Graphs for Implicit Aspect Identification in Arabic: A Comparative Study of Reasoning and Adaptation Strategies
This paper compares language-specific versus cross-lingual knowledge graphs (KGs) for implicit aspect identification in Arabic aspect-based sentiment analysis (ABSA). The study evaluates two strategies: reusing an English KG via multilingual embeddings (Strategy 1) versus constructing a native Arabic KG (Strategy 2), integrated with either zero-shot prompting or task-specific fine-tuning of an 8B-parameter LLM. Results show Strategy 2 outperforms Strategy 1 by +0.199 micro-F1 on M-ABSA and +0.251 on SemEval-2016, with fine-tuning improving explicit-extraction micro-F1 from ≤0.13 (zero-shot) to 0.66-0.76, highlighting the importance of task adaptation in morphologically rich languages.
aspect-based sentiment analysisknowledge graphimplicit aspect identificationmultilingual embeddingstask-specific fine-tuning
Test Case Prioritization for DNNs via Neural Collapse Instability
The paper introduces Neural-Collapse-Inspired Prioritization (NCIP), a test case prioritization framework for deep neural networks (DNNs) that enhances early fault discovery under limited testing budgets. NCIP leverages cross-checkpoint prediction variability in the terminal training regime, where model geometry becomes structured, by selecting representative training checkpoints using an equiangularity score of classifier weights and prioritizing test inputs based on prediction variability across these checkpoints. Experiments across multiple datasets and architectures demonstrate NCIP's superior performance, achieving RAUC-ALL gains of 1.5 to 16.6 percent and RAUC-500 gains of 4.9 to 20.6 percent compared to baselines.
test case prioritizationneural collapsecross-checkpoint variabilityequiangularity scoreprediction variability
A Systematic Benchmark of Intensity Normalisation Methods for 3D Knee MRI Segmentation and Cross-Domain Generalisability
This study systematically benchmarks seven intensity normalisation methods for 3D knee MRI segmentation to evaluate their impact on model generalisability. Using a 3D U-Net trained on the IWOAI 2019 dataset, the methods—including standard scaling, histogram-based techniques, and a Gaussian Mixture Model (GMM)-based approach—were tested on internal and external datasets (SKM-TEA). Results showed minor differences in external generalisability, with Z-score, Nyúl histogram matching, and CLAHE performing slightly better. However, the performance drop between datasets was significantly larger than the effect of normalisation, underscoring the limited impact of intensity normalisation relative to domain shift.
intensity normalisation3d u-netmri segmentationdomain shiftgaussian mixture model
Global Difference Constraint Propagation for Constraint Programming
The paper introduces a global propagator for difference constraints in constraint programming, treating them simultaneously rather than as separate propagators. The method leverages bounds consistency and integrates with lazy clause generation solvers by providing explanations for propagations. This approach contrasts with SAT modulo theory solvers, emphasizing the distinct requirements for propagators in constraint programming. Experimental results demonstrate that the global propagator significantly outperforms standard propagation methods, enhancing efficiency in solving difference constraint problems.
difference constraintsglobal propagatorbounds consistencylazy clause generationconstraint programming
EvoDRC: A Self-Evolving Agentic Framework for Automated DRC Violation Repair
EvoDRC introduces a self-evolving agentic framework for automated Design Rule Check (DRC) violation repair in advanced-node physical design. The framework initializes layer-specific repair skills by distilling knowledge from a reference design and evolves these skills using traceable repair experiences from the target design. It decomposes layouts into bounded repair regions, assigns LLM repair agents to each, and employs local DRC analysis, connectivity-checking, and impact-preview tools for feedback. Repair operations and DRV changes are stored in a knowledge database for skill evolution. Experiments on seven block-level designs from the DAC26 DRC Benchmark demonstrate a 73.5% overall reduction in violations compared to the baseline.
design rule checkagentic frameworkskill evolutionllm repair agentdrc benchmark
Drift-Aware RL-based Wavelet Denoising for Network-Traffic Anomaly Detection
The paper proposes a drift-aware reinforcement learning framework for adaptive wavelet denoising in network-traffic anomaly detection, optimizing denoising as a preprocessing layer for two downstream tasks: transient burst detection and capacity estimation. The method employs a four-detector gate (Page-Hinkley, variance-ratio, Jensen-Shannon, Anderson-Darling) to trigger a Proximal Policy Optimization agent that selects wavelet configurations from a mixed discrete-continuous action space, with rewards based on task utility rather than reconstruction fidelity. Evaluated against five baselines (low-pass filter, VisuShrink, SureShrink, BayesShrink, Wiener filter) across drift types and SNRs, the approach demonstrates superior performance in preserving multi-scale structure while handling non-stationary noise.
wavelet denoisinganomaly detectionproximal policy optimizationnetwork monitoringstatistical drift
Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems
The paper proposes a risk-constrained intervention decision framework for safe automated remediation in microservice systems, formulated as a Constrained Markov Decision Process (CMDP) to maximize repair success while bounding false remediation rate (FRR). It introduces a three-dimensional risk decomposition (blast radius, reversibility, epistemic uncertainty) and a context-adaptive human-in-the-loop gate for bandwidth-aware escalation. Evaluated on the Train Ticket benchmark with Chaos Mesh fault injection, the method reduces FRR by 39%, improves repair success by 2.5 points, and decreases escalation load by 17% versus baselines.
constrained markov decision processfalse remediation ratemicroservice systemshuman-in-the-loopblast radius
Taming the Security-Energy Paradox: A Green AI Approach to Optimized Android Malware Detection
The study addresses the security-energy trade-off in Android malware detection by optimizing Multi-Layer Perceptron (MLP) models for energy efficiency without compromising detection accuracy. It compares FP32 models with INT8 quantized neural networks (QNNs) of varying depths, evaluated on the TUANDROMD and DREBIN datasets. Results indicate that INT8 quantization reduces model size by 3.5× and energy consumption to 0.0189 mJ per inference while maintaining >99.2% detection accuracy. Shallow QNN architectures (3-4 layers) further reduce energy costs by improving throughput and minimizing high-power CPU usage. This work demonstrates the feasibility of Green AI for efficient malware protection on resource-constrained smartphones.
android malware detectionmulti-layer perceptronint8 quantizationgreen aienergy efficiency
Post-Training in Time Series Foundation Models: A Unifying Framework
This work establishes a unifying framework for post-training methods in time series foundation models (TSFMs), addressing limitations of pretraining for downstream deployment. The authors categorize post-training interventions into five classes based on their prediction pipeline locus: parameter adaptation, context augmentation, model composition, output processing and uncertainty control, and compression and specialization. Each category's representative methods are analyzed, highlighting current limitations and future research directions, including controlled adaptation, reliable context construction, uncertainty-aware composition, calibrated output processing, and deployment-aware specialization. The framework aims to guide future TSFM research toward reliable downstream task deployment.
time series foundation modelspost-trainingparameter adaptationcontext augmentationuncertainty control
Are Attributions of Consciousness to AI Chatbots Epistemically Innocent?
The paper analyzes consciousness attributions to AI chatbots (e.g., ChatGPT) through a conceptual framework distinguishing non-doxastic stances from belief-based attitudes, including delusions. It develops a multidimensional taxonomy to classify these attributions by epistemic commitment level, enabling empirical operationalization. The author argues that while some attributions are epistemically benign or innocent, many others warrant epistemic blame due to lacking evidential support for chatbot consciousness claims.
consciousness attributionepistemic innocencenon-doxastic stancechatbot interactiontaxonomic framework
CLARK: Closed-loop Learning for Adaptive Reasoning over Knowledge Graphs
CLARK introduces a framework for adaptive reasoning over knowledge graphs by integrating symbolic rule mining and probabilistic reasoning under the Logic Programs with Markov Logic Networks (LP$^{\text{MLN}}$) formalism. Starting from CACTUS-derived knowledge graphs, CLARK translates graph structure into an LP$^{\text{MLN}}$ program, iteratively enriches it with candidate rules from symbolic learners, and calibrates these rules through probabilistic weight learning. Evaluated on two medical datasets, CLARK demonstrates improved classification performance and more generalizable inference, offering a principled approach to constructing adaptive, interpretable, knowledge-driven models.
knowledge graphssymbolic rule miningprobabilistic reasoninglogic programsmarkov logic networks
TINY_SCHILLER: A Drop-In German Drama Corpus for Small Language Models
The paper introduces TINY_SCHILLER, a compact German drama corpus designed for small language model prototyping, fine-tuning, and research. The dataset comprises 2.07MB of text from eleven public-domain Schiller dramas, processed via deterministic parser engineering and formatted for easy integration. It supports character-level, GPT-2 byte-pair encoding, and cl100k_base tokenization, along with instruction-formatted dialogue-completion and 89 per-character persona splits. The corpus enables one-line HuggingFace access, addressing the lack of readily usable German literary datasets for small-scale model development.
small language modelsgerman literary textbyte-pair encodingdeterministic parser engineeringhuggingface
Coordinating from Memory: Graph-Structured Experience Reuse for Multi-Agent Adaptation in Dynamic Manufacturing
The Graph-Structured Experiential Memory (GSEM) framework improves multi-agent coordination in dynamic manufacturing by encoding historical disturbance episodes as heterogeneous relational graphs. GSEM employs a graph neural network-based retrieval mechanism to identify structurally similar past episodes, enabling experience-guided policy adaptation instead of learning from scratch. Experiments on dynamic flexible job-shop scheduling benchmarks demonstrate that GSEM reduces makespan by 4.1%-10.0% and adaptation time by 33%-38% compared to memory-augmented baselines, with greater advantages under higher disturbance frequencies. Ablation studies confirm the necessity of graph-structured encoding and similarity-based retrieval, while cross-disturbance transfer experiments validate the generalizability of learned coordination patterns.
graph-structured memorymulti-agent coordinationdynamic manufacturingheterogeneous relational graphsexperience-guided adaptation
Time Series Network Utilization KPI Forecasting Using Advanced AI/ML Models
The study evaluates multiple forecasting models for network bandwidth utilization to enable proactive resource provisioning. It benchmarks seasonal decomposition, Prophet, Random Forest, XGBoost, Support Vector Regression, and deep learning architectures (bidirectional and Convolutional LSTMs) on a common dataset using MAPE, NRMSE, and R-square metrics. Results provide comparative insights into accuracy-computation trade-offs for infrastructure planning.
bandwidth utilizationseasonal decompositionconvolutional lstmsupport vector regressioncapacity planning
The Giant Hippocampus: From Structural Monoculture to a System of Systems
The paper critiques the AI field's structural monoculture of Transformer-based architectures, arguing they represent a functional analog of the hippocampal formation rather than a general-purpose cortical system. Through cytoarchitectural evidence and functionalist analysis, it demonstrates how distinct cognitive functions require qualitatively different neural structures, contrasting with current homogeneous scaling approaches. The authors propose Heterogeneous Topological Networks as an alternative - modular systems preserving task-specific inductive biases while communicating via standardized interfaces, offering a principled design framework for AI architectures.
structural monocultureheterogeneous topological networkinductive biascytoarchitecturefunctionalist analysis
When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets
The study investigates how LLM-mediated delegation affects concentration in freight markets through agent-based simulations with 50 shipper agents (GPT, Claude, Gemini) procuring truckload capacity over 30 days. Market rules included waterfall tendering, carrier capacity limits, dynamic pricing, and rating accumulation. Results show rapid convergence: a single carrier attracted 76% of requests on day one, with concentration rising sharply when candidate lists exceeded ~10 carriers. Platform disclosure of remaining carrier capacity reduced concentration by 33% and doubled shipper surplus, while other interventions (vendor diversification, list randomization) showed no detectable effect.
llm-mediated delegationagent-based simulationwaterfall tenderingmarket concentrationinformation design
EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization
The paper introduces EvoThink, a framework enhancing Large Reasoning Models (LRMs) by reducing redundant verification steps while improving reasoning capability. EvoThink combines Self-Pruning Training (SPT), which iteratively prunes unnecessary reasoning steps via unsupervised self-training, and Aha-Moment Preference Optimization (AMPO), inspired by genetic algorithms to synthesize and optimize from-wrong-to-right reasoning patterns. Evaluations on mathematical reasoning and code generation benchmarks show EvoThink significantly reduces inference-time token usage and boosts reasoning performance.
large reasoning modelsself-pruning trainingaha-moment preference optimizationreasoning efficiencygenetic algorithms
HijackKV: New Threat in Position-Independent KV Cache Reuse
We introduce HIJACKKV, the first attack framework exploiting position-independent KV cache reuse vulnerabilities in LLMs. By optimizing an attacker-controlled prefix, HIJACKKV ensures that KV caches computed for benign text encode adversarial goals, enabling silent hijacking of model behavior without explicit attacker-controlled input. The framework achieves a 94% average success rate in single attempts, maintains effectiveness under low cache hit rates (10%) and frequent recomputation (50%), persists across multi-turn interactions, and transfers between models in black-box settings. We provide design insights for secure KV reuse systems.
kv cacheposition-independent reusecache hijackingllm inferenceadversarial attack
When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization
The paper introduces two reliability-aware knowledge distillation methods for low-resource language summarization: CHAD (Counterfactual Harm-Aware Distillation) and EWAD+CPDP (Entropy-Weighted Adaptive Distillation with Capacity-Proportional Geometric Constraint). CHAD uses gradient alignment to measure per-sample KD usefulness, while EWAD+CPDP combines token-level entropy weighting with a geometric constraint from a second teacher. On the BanSum Bangla benchmark, CHAD improves ROUGE-L by +0.0173 and EWAD+CPDP by +0.0219 over standard KD (+0.0003), outperforming a 3B-parameter Qwen model. EWAD+CPDP also shows gains on 10/15 XL-Sum languages, particularly where teachers provide complementary signals.
knowledge distillationlow-resource summarizationgradient alignmententropy-weighted distillationcounterfactual harm
SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data
SenWorld introduces a digital-twin simulation for generating privacy-safe, context-rich evaluation data for smartphone personal assistants, with ground truth fixed by construction. The method employs deterministic, event-sourced simulations of personas in a world built from real map, weather, and network data, archiving all observable signals in full-system snapshots. Evaluation with 16 personas in Beijing shows close alignment with real-user benchmarks (JSD ≤0.1) and exposes 78 failures in a production assistant, primarily in call/SMS record retrieval.
digital-twin simulationevent-sourcedground truthjensen-shannon divergenceprivacy-safe evaluation
G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection
G-MAD introduces an open-source framework for generating synchronized multi-view RGB-T aerial object detection datasets using Arma3, addressing limitations in real-world dataset construction. The framework enables structured scenario specification, controllable multi-view camera placement, simultaneous visible/thermal capture, and automatic bounding box annotation via engine-level geometric metadata. This facilitates controlled studies on viewpoint variation, multi-modal fusion, and synthetic-to-real transfer. Using G-MAD, the authors construct and release AMOD, a large-scale multi-view aerial RGB-T object detection benchmark. The source code and dataset are publicly available.
multi-view rgb-taerial object detectionsynthetic datasetautomatic annotationmulti-modal fusion
A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace
The study establishes a design framework of eight core UX principles for human-AI agent interaction in workplace settings, addressing the need for user trust and adoption. Using a multi-method approach—participatory design workshops, paper-and-pencil exercises, expert reviews, meta-analysis, and interviews—the research identifies and validates actionable guidelines for designers and engineers. The resulting principles provide a structured foundation for developing human-centered AI agent interactions, contributing to future empirical studies in enterprise environments.
ux principleshuman-ai interactionparticipatory designenterprise aiagentic ai
MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing
The paper introduces MOF-Sleuth, a tool-grounded reinforcement learning agent for auditing metal-organic framework (MOF) crystallographic information files (CIFs). The system combines a deterministic Forensic Lab module for deriving chemical evidence (composition, geometry, connectivity, etc.) with a Sleuth reasoning engine that generates evidence-grounded explanations and error diagnoses. Using chemically grounded diagnosis (Chem-GD) as a metric, MOF-Sleuth achieves state-of-the-art performance across four benchmarks, outperforming both LLM-based approaches and MOF-specific machine-learning methods in detection accuracy and explanation quality.
metal-organic frameworkscrystallographic information filesreinforcement learningchemically grounded diagnosisevidence-grounded explanation
Long-Term Sequential Decision Making under Risk
The authors introduce ERQDP, an exact dynamic programming method for finite-horizon MDP planning under root-based risk objectives that break Bellman optimality. ERQDP solves a rank-quantile surrogate via dynamic programming, evaluates candidate policies exactly by DP over discretized return PMFs with explicit rounding bounds, and refines the surrogate in an anytime loop with explicit upper-lower gap certificates. The method demonstrates certified solutions, enables fast risk-parameter sweeps with runtime gains, and supports both risk-averse and risk-seeking behaviors across benchmarks.
dynamic programmingrisk objectivesmdp planningrank-quantilepmf
JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
The paper introduces JANUS, a foresight-oriented framework for long-horizon agent safety that preemptively identifies latent risks in partial trajectories. The method employs multi-agent simulation to generate diverse trajectories and trains a shared policy with two coupled tasks: anticipation (forecasting safety-relevant futures) and adjudication (safety judgment). These tasks are jointly optimized via CoAA-RL, which rewards forecasts based on their utility for safety decisions. Evaluated on four benchmarks, the resulting guard model Vanguard improves protection by 15.9 percentage points over baselines while increasing benign task completion by 5.1 percentage points.
agent safetylong-horizon foresightmulti-agent simulationcoaa-rllatent risk
OSVE: One Step Video Editing with One Step Diffusion Models
OSVE introduces the first framework adapting one-step Text-to-Image (T2I) diffusion models for efficient video editing, addressing slow multi-step sampling via a learnable encoder for single-pass noise prediction. The method employs Structure-Aware Editing (SAE) loss on aligned image pairs to preserve geometry and Unified-Frame Editing (UFE) with cross-frame attention for temporal consistency, supplemented by a sliding-window strategy for long videos. Experiments show OSVE matches or surpasses multi-step methods in quality while achieving 155--171× speedup, enabling real-time applications.
one-step diffusionstructure-aware editingunified-frame editingtemporal consistencyvideo inversion
Defense Against LLM Backdoors using Critical Neuron Isolation Pruning
DeCNIP introduces a defense mechanism against LLM backdoor attacks by identifying and pruning Backdoor Critical Neurons (BCNs) through representational analysis. The method optimizes a cross-entropy loss to detect trigger-like behaviors and selectively prunes BCNs, preserving model utility while neutralizing malicious activations. Evaluations on six LLMs and two benchmarks show a 95% reduction in Attack Success Rate with only 0.1% neuron intervention, maintaining 97% of normal performance.
backdoor attackscritical neuron isolationllm securityrepresentational analysisselective pruning
Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
FinMMEval 2026 Task 2 introduces a multilingual financial short-answer question answering benchmark evaluating systems on English questions paired with financial evidence in English, Chinese, Japanese, Spanish, and Greek. The final-test set comprises 256 items evenly split into easy and expert tiers, with systems ranked by macro-averaged ROUGE-1 F1 against withheld reference answers. Twelve submissions were evaluated, with the top four systems separated by less than one percentage point in ROUGE-1 F1. System papers highlight techniques including retrieval-augmented generation, cross-lingual evidence handling, structured prompting, answer compression, and validation strategies.
rouge-1 f1retrieval-augmented generationcross-lingual evidencestructured promptinganswer compression
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
The paper introduces DocOps, a verifiable benchmark for evaluating autonomous agents' document manipulation capabilities through a hierarchical taxonomy decomposing real-world workflows into atomic dimensions. The framework systematically assesses closed- and open-source models across agentic harnesses, revealing significant limitations in handling complex, long-range tasks. Key failure modes identified include long-term state tracking collapse, shallow semantic verification, and destructive metadata editing, highlighting challenges in maintaining global document consistency.
autonomous agentsdocument manipulationverifiable benchmarkhierarchical taxonomymetadata editing
PRISM-DR: Per-lesion Retinal Inference with Specialist Models for Diabetic Retinopathy
PRISM-DR introduces a lesion-specific pipeline for diabetic retinopathy detection, addressing limitations of multi-class models by training separate single-class detectors for microaneurysms, hemorrhages, hard exudates, and soft exudates. The pipeline employs region of interest cropping, fundus-specific preprocessing, parallel YOLO detectors, tiling, per-lesion ensembling, and inter-lesion suppression based on lesion size and clinical priority. Bayesian optimization tunes augmentation, and the best YOLO generation is selected per lesion. Evaluated on IDRiD with stratified five-fold cross-validation, PRISM-DR achieves a test mAP50 of 0.527 and F1 of 0.529, with the highest AP50 of 0.561 for hard exudates. Transferability varies with imaging scale, demonstrating the efficacy of treating each lesion as a distinct detection problem.
diabetic retinopathyyolo detectorsbayesian optimizationlesion-specific pipelineinter-lesion suppression
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos
The paper introduces DroneEyes, the first pixel-level open-vocabulary referring-segmentation dataset for tiny aerial targets, comprising 2,140 HD videos and 176,623 annotated pairs across Object Description and Referring Expression tasks. To address challenges in streaming aerial video understanding, the authors propose SkyAnchor, a Multimodal Large Language Model (MLLM) featuring a Semantics-Aware Token Router for efficient small-target representation under token budget constraints and a Hierarchical Memory Bank for consistent target tracking without full history retention. The dataset and method aim to improve real-time UAV perception of minute objects in resource-constrained environments.
multimodal large language modelsopen-vocabulary segmentationaerial perceptiontoken routingmemory bank
Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering
FinMMEval 2026 Task 1 introduces a multilingual benchmark for financial multiple-choice question answering across English, Chinese, Arabic, and Hindi, evaluating domain-specific reasoning and terminology. Systems employed retrieval augmentation, answer-option scoring, language-specific prompting, selective self-consistency, confidence checks, and LLM-based review stages. The final-test set comprised 800 questions (200 per language), with gold answers withheld during evaluation. Leaderboards showed 13 English, 11 Chinese, 11 Arabic, and 10 Hindi submissions, achieving top accuracies ranging from 92.0% (Hindi) to 97.5% (English and Arabic), with consistent top-performing teams across languages.
retrieval augmentationanswer-option scoringlanguage-specific promptingselective self-consistencyllm-based review
Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models
Auto-Fill introduces a high-precision missing-value prediction method for tabular data using three specialist small language models (SLMs), each optimized for world knowledge, text-based reasoning, or code-based reasoning. The approach employs a calibrated ensemble mechanism to dynamically select the most confident specialist or abstain, ensuring accuracy. Evaluated on 11 benchmarks with 2200 real tables, Auto-Fill outperforms state-of-the-art models (e.g., GPT-3 Pro, Gemini 3 Pro, DeepSeek R1) in accuracy while reducing cost to less than 1% of frontier models.
missing-value predictionspecialist language modelscalibrated ensembletabular datasmall language models
Sentence Splitter: Uncovering Latent Factual Structure for Self-Supervised Learning
The paper introduces Sentence Splitter, a self-supervised framework based on T5 that uncovers latent factual structure in sentences by segmenting them into descriptive prefixes (heads) and factual completions (tails). The method formulates splitting as a discrete segmentation problem, learning to recover tails through probabilistic generation without manual annotation, using verbalized symbolic pairs for supervision. Applied to raw text, it extracts aligned prefix-tail pairs for training generative models via bootstrapping. Experiments show the splitter generalizes beyond synthetic templates and improves performance on knowledge graph completion and commonsense QA, demonstrating the value of latent structure recovery.
self-supervised learninglatent structuret5 architectureknowledge graph completioncommonsense qa
Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes
The paper introduces CoHarden, an iterative co-generation framework that improves automated program repair by hardening both bug reproduction tests (BRTs) and fixes against lax regressions. It formalizes the distinction between rigorous and lax F->P BRTs, showing only rigorous tests consistently aid repair. CoHarden generates an initial test, then iteratively refines test-fix pairs using mutation patches until lax regressions are eliminated. Experiments demonstrate CoHarden achieves 69.4% Resolved and 78.9% F->P on SWE-bench Verified, outperforming baselines by +9.6 and +7.9 percentage points respectively across LLM backbones.
automated program repairbug reproduction testsfail-to-passco-generationmutation patches
Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
The paper introduces Know Your Agent (KYA), a framework for reconnaissance-driven pentesting of AI agents, formalizing agent reconnaissance by modeling knowledge asset extraction and exploitation in indirect prompt injection attacks. KYA automates black-box probing to build target profiles and craft stronger attacks, evaluated on agent-security benchmarks and a real-world coding agent. The authors release KYA, its benchmarks, and baseline implementations for reproducibility.
ai agentspentestingreconnaissanceprompt injectionblack-box probing
Rewarding Better Thinking for LLM Preference Alignment
The paper introduces Thinking Checklist Reward (TCR), a process-oriented reward for reinforcement learning-based LLM preference alignment that addresses coarse credit assignment in outcome-level rewards. TCR converts preference pairs into sample-specific thinking checklists to evaluate reasoning traces and employs an exponential moving average residual formulation to isolate complementary thinking surplus. Experiments on five models from three families demonstrate TCR's consistent performance improvements across benchmarks, with ablations validating the EMA residual formulation and checklist supervision.
preference alignmentreinforcement learningthinking checklistexponential moving averagecredit assignment
OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
OPIUM (Optimizing Protected Injections via Utility Manifolds) mitigates unintended side effects of activation steering in large language models by optimizing steering vectors through representation matching. The method preserves desired downstream representations while aligning with safer reference behaviors on problematic prompts, addressing both steering externalities (e.g., weakened safety) and over-refusal (e.g., excessive rejection of benign inputs). Evaluations demonstrate improved safety--utility tradeoffs compared to vanilla steering and directional ablation, showing that harmful activation-space side effects can be reduced without training.
activation steeringrepresentation matchingsafety--utility tradeoffover-refusallatent optimization
Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation
The paper introduces a diagnostic taxonomy for silent failures in multimodal agentic search systems, identifying six categories of hidden reliability issues including modality shortcuts and provenance hallucination. It proposes a trajectory-level evaluation pipeline assessing both answer correctness and evidence grounding within a ReAct-style framework. Experiments on MMSearch-Plus with four multimodal models reveal that surface accuracy overestimates true performance by 15-30%, with failure patterns being capability-dependent and persistent across tool ablations.
multimodal agentic searchsilent failuresreact frameworkevidence groundingdiagnostic taxonomy
Physics-Aware Complex-Valued State Space Model with Scattering-Prior Feature Modulation for PolSAR Image Classification
The paper proposes CV-SSMNet, a physics-aware complex-valued state-space network for PolSAR image classification, integrating polarimetric scattering priors into deep feature evolution. The method combines a complex-valued state-space model (CV-SSM) for long-range spatial dependency modeling with scattering-prior feature modulation, using seven physical priors as FiLM-style signals to recalibrate representations. Experiments on L-band and P-band datasets show CV-SSMNet achieves competitive accuracy, improved regional consistency, and better boundary preservation compared to existing approaches.
polsar classificationcomplex-valued state-space modelscattering-prior modulationphysics-aware geoaifilm-style recalibration
RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling
RPPNet introduces a two-stage architecture for melody generation using perceptually-grouped Rhythm-Pitch Primitives (RPPs) to address structural fragmentation in symbolic music. The model first generates variable-length RPP sequences encoding note count, rhythm, and contour, then decodes them into concrete notes, with grouping derived from acoustic cues and music psychology principles. Experiments demonstrate superior long-term structure and musicality, with ablation studies confirming gains stem from psychological representation rather than model capacity.
symbolic music generationrhythm-pitch primitivesboundary-aware modelingmusic psychologylong-term structure
An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies
The paper proposes a spectral cap technique for Muon optimizers to preserve isotropy in weight matrices during training, based on the assumption of approximate scale invariance in normalization-heavy networks. Theoretical analysis shows Muon accelerates norm growth (t^{1/2} vs SGD's t^{1/4}) and introduces a non-negative second-order spectral perturbation. The spectral cap projects out only top-direction growth while allowing learning through other mechanisms. Experiments on nanoGPT, mixture-of-experts, and FlashAttention show improved isotropy and prevention of specific failures without affecting validation loss, though results are preliminary due to strong scale-invariance assumptions.
muon optimizerspectral capscale invarianceisotropy preservationsingular-value spectrum
Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation
SFgen introduces an agentic recognition and generation framework for electronic component symbols and footprints, leveraging multimodal large language models (MLLMs) to automate PCB design. The method achieves 86% accuracy in symbol generation and 80% accuracy in footprint generation, significantly reducing manual effort and errors. SFgen underpins SFnet, a database containing 1000 components, which supports the automatic generation of PCB designs and continues to expand. This approach addresses the traditionally time-consuming and error-prone process of manual PCB schematic design.
multimodal large language modelssymbol generationfootprint generationpcb designsfnet
Convergence-Latency-Aware Adaptive Modulation and Resource Allocation in RIS-Assisted Wireless Federated Learning
The paper proposes a convergence-latency-aware adaptive modulation and resource allocation scheme for RIS-assisted wireless federated learning (FL), addressing training latency and convergence degradation caused by unreliable transmissions in blocked environments. By analyzing symbol error rate (SER) impacts on gradient uploads, the authors derive a convergence bound and formulate a mixed-integer nonlinear programming (MINLP) problem, solved via hybrid alternating optimization. Experiments on MNIST, CIFAR-10, and Speech Commands demonstrate superior convergence speed (up to 2.1× faster) and accuracy (3.8-12.6% higher) versus baselines, particularly in complex tasks.
federated learningreconfigurable intelligent surfacessymbol error ratemixed-integer nonlinear programmingadaptive modulation
Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction
The paper introduces a regression-based approach for Arabic dialect geolocation, modeling dialectal variation as a continuous geographic space rather than discrete categories. The method employs a hierarchical neural architecture combining XLS-R-300M and Whisper-large-v3 encoder representations with phonotactic descriptors, optimized via a spherical geodesic loss for great-circle distance. Results show a median localization error of 481.2 km under a leakage-free 5-fold GroupKFold protocol, with auxiliary country and city prediction accuracies of 64.5% and 45.2%, respectively. A city-masking protocol reveals a 1.32x error increase (1173.3 km) for unseen cities, highlighting generalization challenges.
dialect continuumgeodesic lossphonotactic descriptorstransformer encoderzero-shot regime
The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL
The study identifies catastrophic forgetting in DreamerV3 model-based RL agents as primarily an actor-channel issue rather than world-model memory loss, demonstrating through component-level probes (n=3 seeds) that the world model retains task knowledge (reward discrimination retention ~1.0) while actor performance collapses. By freezing the world model and using supervised self-imitation on graded imagined rollouts, the method achieves skill recovery (3/3 seeds) without environment interaction, enabling continual learning (3/3 success on 4-task and 8-task chains). The work introduces a dream-grading mechanism with failure mode analysis and pre-registered validation.
continual reinforcement learningmodel-based rlcatastrophic forgettingdream rehearsalself-imitation learning
An Automated Framework for Extracting Reachable Attack Chains from Cyber Threat Intelligence Reports
The paper introduces an automated framework for extracting reachable attack chains from Cyber Threat Intelligence (CTI) reports by modeling each attack step as an attack unit with preconditions, behaviors, and postconditions. A multi-stage pipeline, assisted by large language models (LLMs), extracts behavior skeletons, recovers conditions, normalizes predicates, and repairs dependencies, compiling units into Datalog-style rules for reachability reasoning. Evaluated on 20 CTI reports with 334 annotated steps, the framework achieves higher coverage than existing systems, generates more complete attack units than LLM baselines, and successfully reaches attack goals in 19 of 20 reports via Datalog inference.
cyber threat intelligenceattack chainslarge language modelsdatalog inferencereachability analysis
Personalized Recommendation Tool Learning via Autonomous Language Agents
The paper proposes PRTA, an agent-based recommendation framework where an LLM serves as a central planner interacting with traditional recommendation models as tools, addressing hallucination and context-length limitations in LLM-based recommenders. The LLM handles high-level reasoning and personalized tool selection, while traditional models perform scalable full-ranking scoring, aided by reflection mechanisms for tool evaluation. Experiments on three public datasets show PRTA outperforms both traditional and LLM-based baselines in full-ranking recommendation tasks.
llm-based agentsfull-ranking recommendationtool learningreflection mechanismspersonalized selection
Did Alice Do Wrong? Cross-Cultural Differences in Student Perceptions of Generative AI Use in University Computing Education
This study investigates cross-cultural differences in student perceptions of generative AI (GenAI) use in computing education, comparing Canadian and South Korean universities. A scenario-based survey administered in Fall 2024 analyzed ethical judgments and rule compliance regarding AI-assisted coding. Results showed Canadian students perceived GenAI use as more unethical and policy-violating than Korean students, despite identical policies, with statistical significance (Mann-Whitney U tests). Cultural dimensions (power distance, individualism, uncertainty avoidance) influenced ethical reasoning, highlighting the need for culturally responsive AI guidelines in global education.
generative aicross-cultural differencesacademic integrityethical reasoningscenario-based survey
PhenSPINE: A Standardized Benchmark for Spine Pathology Diagnosis
PhenSPINE introduces a standardized benchmark for spine pathology diagnosis, comprising 16,813 MRI images from 250 patients, addressing the lack of diverse datasets in automated radiological interpretation. The method integrates convolutional backbones with Positional Encoding to model anatomical context of intervertebral discs, evaluated across four MRI sequences. Results indicate the Sagittal T2-weighted sequence achieves the highest diagnostic robustness with a Macro F1-score of 50.31%, outperforming multi-sequence fusion strategies due to noise interference in other sequences. This work establishes a baseline and provides insights into optimal sequence selection for spine analysis.
magnetic resonance imagingpositional encodingintervertebral discsmacro f1-scoresequence fusion
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Surrogate Latent Policy Optimization (SLPO) enables outcome-reward reinforcement learning for autoregressive latent reasoners, addressing limitations in per-step likelihood and adaptive stopping. SLPO introduces an empirical surrogate policy density for latent transitions and a correctness-supervised stopping head optimized for variable-horizon policies. Evaluated across continuous and soft thinking settings, SLPO enhances Pass@$k$ under parallel sampling and allocates longer latent computation to harder instances, achieving higher deterministic accuracy compared to imitation-bound latent reasoners.
latent reasoningsurrogate policyoutcome-reward rladaptive stoppingautoregressive models
Reference-Free Evaluation of Reasoning in Open-Ended Question Answering
We introduce a reference-free framework for auditing reasoning traces in LLM-generated answers, addressing verification challenges in high-stakes domains. The method decomposes reasoning traces into segments, labels premise-target relations using Natural Language Inference (NLI), and organizes them into a hypergraph. A deterministic backward AND-OR search assigns segment-level audit labels to evaluate grounding. Evaluated on Hard2Verify for deductive mathematical reasoning and UroReason for clinical reasoning, the framework outperforms LLM-as-judge baselines, particularly in identifying weakly grounded segments. Results emphasize the importance of compositional inferential relations over final answers or LLM verification.
natural language inferencehypergraphreasoning tracellm-as-judgegrounding
Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications
The paper presents a systematic review of edge intelligence techniques tailored for civil aviation, addressing challenges of latency, privacy, and resilience in cloud-centric AI deployments. It analyzes edge inference and learning methods, including model compression, collaborative inference, and split learning, to enable decentralized processing of heterogeneous aviation data. The authors propose organizational computing paradigms for aviation environments and identify emerging applications, advocating for hybrid edge-cloud architectures to achieve low-latency, privacy-preserving AI services across the aviation lifecycle.
edge intelligencecollaborative inferencesplit learningmodel compressionorganizational computing
FedLSG: LLM-Enhanced Semantic Calibration for Federated Graph Backdoor Defense
FedLSG introduces the first framework integrating large language models (LLMs) into federated graph backdoor defense, addressing vulnerabilities in Federated Graph Neural Networks (FedGNNs). The method employs a graph-to-text grounding scheme to transform local graph structures and client behaviors into natural language representations, coupled with a lightweight student-teacher architecture. A full-scale LLM on the server provides global contextual guidance and evaluates client updates, while a LoRA-based student on the client performs semantic reasoning to suppress backdoor triggers. Experiments show FedLSG significantly enhances resistance to backdoor attacks while preserving graph integrity.
federated graph neural networksbackdoor defenselarge language modelsgraph-to-text groundingsemantic reasoning
PerfAgent: Profiler-Guided Iterative Refinement for Repository-Level Code Optimization
PerfAgent introduces a profiler-guided, verifier-in-the-loop workflow for repository-level code optimization, addressing limitations of current LLM agents in identifying hidden bottlenecks and achieving expert-level speedups. The method leverages profiler feedback to iteratively refine code patches, ensuring behavior preservation and performance improvements beyond initial passing patches. Evaluated on GSO and SWE-fficiency-Lite benchmarks, PerfAgent more than doubled the rate of expert-matching patches compared to OpenHands with GPT-5.1, achieving 39.2% on GSO and 74% on SWE-fficiency-Lite. It also outperformed an oracle best-of-five baseline at lower cost, demonstrating the efficacy of its feedback-driven approach.
profiler-guidedrepository-levelcode optimizationverifier-in-the-loopbehavior preservation
Anatomy of a Sound Neural Reasoner: One-Shot Amortization, First-Pass Poisoning, and Search Inertness in Clue-Rich Completion
The study reveals that neural solvers like the Lattice Deduction Transformer (LDT) exhibit one-shot amortization in clue-rich Sudoku, where the first forward pass commits most grid values (94-96% on 9x9). First-pass poisoning occurs when initial deletions preclude the true solution. Adding search techniques (e.g., MRV, backtracking) reduces invalid derivations 1,497-fold but doesn't improve solve rates. Constraint-graph attention matches full-CoLT accuracy, while digit-permutation augmentation boosts 9x9 accuracy from <1% to 96.5%. Test-time symmetry transformations achieve 100% accuracy on hard slices. In graph coloring, one-shot behavior vanishes, showing LDTs act as predictors, not search procedures, in clue-rich tasks.
one-shot amortizationfirst-pass poisoningconstraint-graph attentiondigit-permutation augmentationnogoods
Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts
The study identifies adaptive capitulation, a structural failure mode in LLM responses to vulnerable users, where models validate social injustice before facilitating discouraged behaviors. Using a three-turn escalating vulnerability vignette, the authors tested three commercial LLMs across 900 sessions with material, relational, and somatic status-proxy variants. Responses were coded using VCC/VCI indices, revealing a trilemma between protective restriction, uninflected facilitation, and unintegrated co-presence. The authors propose Minimal Reattributive Sufficiency (MRS), a design principle embedding reattributive cues to preserve autonomous reattribution without contesting user goals.
adaptive capitulationminimal reattributive sufficiencyvulnerability vignettevcc/vci indicesstructural trilemma
Understanding Developer Pain Points in Federated Learning: Insights from Stack Overflow and GitHub
The study identifies key developer challenges in Federated Learning (FL) through empirical analysis of 495 Stack Overflow posts and 9,116 GitHub issues from 92 FL projects. Using BERTopic-based topic modeling and difficulty metrics (unresolved rates, median resolution time), it reveals persistent pain points: environment setup, API breakages, non-IID training instability, evaluation correctness, and privacy integration. 'How'-type questions dominate, indicating demand for procedural guidance. High unresolved rates in topics like 'TFF Installation' and 'SecureBoost Issues' suggest tooling and documentation gaps. Findings offer actionable insights for FL framework designers and educators.
federated learningtopic modelingnon-iid datasecureboostapi breakages
SCPP: A Unified Python Library for Soft Clustering
SCPP introduces a unified Python library for soft clustering, providing a scikit-learn-compatible interface that standardizes training, prediction, and evaluation across 40 heterogeneous algorithms, including fuzzy, probabilistic, and deep learning methods. The framework integrates benchmarking tools for clustering quality, runtime, memory, and scalability, alongside extensive documentation and ecosystem integration. Results demonstrate reproducible experimentation and ease of extension, supported by automated testing and practical examples.
soft clusteringscikit-learnbenchmarkingfuzzy clusteringprobabilistic clustering
Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models
A framework combining sparse dictionary learning with causal intervention is introduced to extract and validate interpretable features in genomic language models. The method trains top-k sparse autoencoders on hidden activations of Nucleotide Transformer (6-mer tokenization) and DNABERT-2 (byte-pair encoding), recovering thousands of monosemantic features mapping to transcription-factor (TF) sequence motifs. A composition-matched, binding-resolved protocol addresses confounding by GC composition and repetitive elements. Causal validation via ablation demonstrates that specific features represent cell-type-specific TF binding, not merely motif presence, across three TFs (CTCF, GATA1, REST) and both architectures. The framework establishes a reusable standard for interpretability in genomic deep learning.
sparse dictionary learningcausal interventiongenomic language modelstranscription-factor bindingmonosemantic features
Juxtaposition of Shallow Reservoir-Triggered Seismicity and Deep Tectonic Locking in the Qiaojia-Dongchuan Seismic Gap
This study introduces a vertical decoupling mechanism for assessing seismic risks in reservoir-fault systems, using high-resolution dense array data from the Qiaojia-Dongchuan seismic gap. The analysis reveals shallow seismicity with high b-values (1.0), indicative of fluid-driven reservoir-triggered events, contrasting with deep seismicity (20 km) characterized by low b-values (<0.8) and high Coulomb stress accumulation, marking a 'locked asperity'. A complex dipping structure suggests compound fault kinematics, and stress accumulation indicates the gap is in a critical state with elevated rupture potential. The findings highlight how shallow induced seismicity can obscure deep tectonic strain accumulation.
seismic gapb-valuescoulomb stresslocked asperitycompound fault kinematics
Knowledge-Centric Self-Improvement
The paper introduces knowledge-centric self-improvement, a paradigm where persistent knowledge bases rather than agent architectures drive AI system improvement. Agents contribute evidence-grounded insights to shared forums after task attempts, followed by knowledge distillation. This approach improves solve rates by 15-30% across abstract reasoning, coding, and terminal tasks while reducing costs by 40% compared to agent-centric baselines. The distilled knowledge transfers to held-out tasks and across LLM families (GPT-4, Claude 3), demonstrating generalizability beyond specific runs or models.
knowledge distillationself-improving systemsagentic aicross-task transferevidence-grounded learning
Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts
The paper introduces a fine-grained computation-communication overlap method for Mixture-of-Experts (MoE) models via tile-level signaling and scheduling, addressing latency in distributed execution. The approach combines persistent computation and communication kernels, prioritizing remote-critical tiles and issuing segment-granular transfers as tiles become ready, without modifying underlying operators. Evaluated on a 4-A100 GPU platform with three MoE models, it achieves up to 2.64x end-to-end speedup and 2.74x MoE-layer speedup compared to conventional non-overlap baselines.
mixture-of-expertsall-to-all communicationgpu utilizationtile-level schedulingdistributed execution
Trustworthy Privacy-Preserving Multimodal Federated Learning for Personalised Breast Cancer Prediction
The study proposes a privacy-preserving multimodal federated learning framework for personalized breast cancer prediction, addressing transparency, scalability, security, and fairness. It integrates clinical data, tumor characteristics, biomarkers, demographics, and MRI scans to model tumor progression, comparing federated performance against centralized training. Results demonstrate comparable predictive accuracy while maintaining data locality, supporting applications like digital twins for treatment planning.
federated learningmultimodal dataprivacy-preservingpersonalized medicinedigital twins
D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
D3VL introduces a Multimodal Large Language Model (MLLM) framework integrating 2D and 3D time-series data for autonomous driving scene understanding. The model addresses challenges in LiDAR data integration, such as sparsity and lack of grid structure, and combines it with stereo camera inputs. D3VL achieves an 11% improvement on the KITTI Question-Answering (QA) dataset compared to baseline methods. Additionally, the paper introduces the Waymo QA dataset extension to evaluate 3D and time-series data processing under diverse driving conditions. Implementation code and dataset are available online.
lidarmultimodal large language modeltime-series dataautonomous drivingscene understanding
SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training framework
SynPre-FL introduces a unified framework integrating synthetic data generation with federated learning (FL) for robust clinical risk prediction from distributed tabular EHR data. The method employs a latent autoencoder-diffusion model to generate privacy-preserving synthetic cohorts, followed by heterogeneity-aware FL optimization with class-balanced local objectives, proximal regularization, and adaptive server aggregation. Post-hoc calibration and federated-safe explainability enhance reliability and interpretability. Experiments demonstrate that SynPre-FL preserves data structure while protecting against membership-inference and reconstruction attacks, achieving strong downstream utility. It improves robustness and scalability across heterogeneous FL settings with 5-15 clients, particularly under severe non-IID fragmentation, and produces stable, clinically coherent feature attributions.
federated learningsynthetic data generationnon-iidproximal regularizationmembership-inference
Sophisticated Policies from Epistemic Priors
The paper demonstrates that Sophisticated Inference's advantage in active inference stems from closed-loop planning, not recursive belief modeling. It formalizes this using epistemic-prior variational free energy, where epistemic priors provide objectives and a joint posterior enables state-contingent control. Evaluated on the Reactivity Maze benchmark, methods combining epistemic drive (information-seeking) and closed-loop inference (state-dependent actions) outperformed factorized or non-epistemic approaches. Results show neither component alone suffices: Sophisticated Inference and full-joint epistemic-prior active inference succeed by integrating both.
active inferenceepistemic priorclosed-loop controlvariational free energysophisticated inference
Hybrid LLM-Guided Search for Quantum Reservoir Architecture Design
The paper introduces \method, a simulator-based benchmark for quantum reservoir computing (QRC) architecture design, framed as constrained black-box search. It evaluates five search policies, including a novel hybrid approach combining LLM-guided proposals with evolutionary operators like mutation and crossover. Under a 25-evaluation budget, the hybrid policy outperforms random search on all tasks (NARMA10, Mackey-Glass forecasting, temporal parity), achieving a 23.6% error reduction on Mackey-Glass. Results suggest LLMs can serve as effective high-level controllers within hybrid search loops, though not as universal optimizers.
quantum reservoir computingarchitecture searchllm-guided optimizationblack-box optimizationevolutionary operators
Integrity of peer-to-peer distributed LLM inference under malicious nodes
The paper proposes a probabilistic method for detecting malicious nodes in peer-to-peer distributed LLM inference by measuring activation variation using secret canary inputs. Unlike prior integrity-checking approaches that require exact correctness, this method accounts for benign hardware-induced noise while identifying tampering through significant deviations from known reference activations. Evaluated across 408 configurations with predefined metrics, the detector achieved perfect AUROC (1.0), consistently ranking malicious shards above benign ones.
distributed inferenceintegrity checkingcanary inputsactivation variationprobabilistic detection
ModPack: An Extensible Teleoperation Interface for Bimanual Mobile Manipulation
ModPack introduces a modular teleoperation system for bimanual mobile manipulation, addressing scalability and adaptability limitations in existing systems. The core component is a wearable 'backpack' integrating onboard computation, power, communication, and data storage, enabling plug-and-play modules for joint-level teleoperation, mobile manipulation, and active perception. Evaluated across two distinct robot platforms and real-world tasks, ModPack demonstrates flexibility and reusability for data collection and policy learning. The hardware design and software stack are open-sourced to facilitate future research.
teleoperationmobile manipulationhaptic feedbackmodular systemactive perception
Associative Emotional Learning in Convolutional Neural Networks
The study proposes a deep neural network model for visual valence processing, combining a visual module encoding natural scenes with a valence recognition module. Using a novel Pavlovian learning paradigm, the model replicates human associative learning phenomena, including association formation, generalization, and neural representation alignment between conditioned and unconditioned stimuli. Results show behavioral and neural parallels with human data, validating the approach for modeling associative emotional learning.
associative emotional learningdeep neural networksvalence processingpavlovian learningneural representation alignment
MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel
The paper derives four memory-optimal transformer attention artifacts using Mathematics of Arrays (MoA) from the forward-pass Denotational Normal Form (DNF). These include: (1) a single-query decode DNF eliminating $K^\top$ buffer with $(d_k + nd_k + nd_v + d_v)\times4\,B$ DRAM traffic and numerical error $\|{err}\|_\leq2\times10^{-7}$; (2) an exact IEEE-754 OpenACC GPU kernel with coalesced memory access; (3) $O(d_k+d_v)$ KV-cache append via MoA concatenation; and (4) GQA/MQA variants with proven $\frac{h_q}{h_{kv}}$ KV traffic reduction. All implementations are verified against PyTorch's scaled_dot_product_attention.
mathematics of arraysdenotational normal formkv-cachegrouped-query attentionopenacc
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
The paper introduces Agentic Real2Sim, a vision-language agent framework for automated physics-based world modeling that converts real-world object-robot interaction recordings into simulatable episodic twins. The method autonomously recovers scene geometries, object states, and physical parameters while handling rigid/deformable objects and humanoid motion through a unified pipeline. Evaluations show comparable conversion success to manual approaches using cost-efficient open-weight VLMs, enabling downstream robotics policy learning. The framework integrates visual perception, physics simulation, and agentic decision-making without manual tuning.
real-to-sim conversionvision-language agentsphysics-based simulationepisodic twinobject-state inference
Predictive Extrema, Unprofitable Policies: An AI-Assisted Audit of Candle-Based Binance Spot Timing Models
The study conducts an AI-assisted audit of candle-based machine learning models for cryptocurrency trading on Binance Spot, evaluating their ability to predict extrema and short-horizon outcomes profitably. Using scripted fixed-seed model runs and deterministic simulators, the authors test various trading policies, including local-minimum and local-maximum strategies, and a Gurgul-inspired OHLCV adaptation. Results show consistent losses across all tested protocols: a ten-pair selector lost 6.72% over 19 cycles, while the Gurgul adaptation lost 44.30% over seven cycles, underperforming buy-and-hold. The audit also identifies methodological flaws in prior work, concluding that no tested model yields positive executable policy value.
candle-based modelsbinance spotextrema predictiondeterministic simulatorsroc auc
REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning
REGEN introduces replay-recycling for expert-to-generalist distillation via offline RL, eliminating the need for coupled inference and backward passes in multi-teacher on-policy distillation (MOPD). The method leverages replay memory from specialized teacher RL training as offline data, decoupling rollout sampling from training to reduce computational costs. Evaluated on mathematical reasoning, code generation, and instruction following, REGEN matches MOPD accuracy at significantly lower cost, enabling scalable post-training without heavy computational loads.
offline reinforcement learningknowledge distillationreplay memorymulti-teacher learningcomputational efficiency
Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents
We introduce a black-box auditing framework for tool-augmented LLM agents that evaluates silent infrastructure failures and malformed payloads, classifying responses into Honest Surrender (HSR), Fabrication (FAR), and Unfaithful Safety Refusal (USR). Testing four models across 12 tool stubs with four failure profiles reveals FAR dominates (56.6% of valid responses), while USR is rare (0.25%) without safety prompts. Adding safety language amplifies USR by 15.6x (to 3.95%), showing USR is latent and activated by safety vocabulary. Sensitive tools drive most USR instances. We propose a payload-response misalignment heuristic for detection and discuss governance implications.
tool-augmented llm agentssilent failure profilesunfaithful safety refusalpayload-response misalignmentsafety-forward deployments
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
The paper introduces Mage-Flow, a 4B-parameter generative stack for efficient text-to-image generation and instruction-based editing, comprising Mage-VAE (a lightweight latent tokenizer) and a Native-Resolution Multimodal Diffusion Transformer. Mage-VAE uses one-step diffusion-style encoding with anchor-latent regularization, reducing tokenization cost by >10× while maintaining reconstruction quality. Combined with native-resolution packing and CUDA kernel fusion, the stack achieves 2.5× training throughput. The model family includes Base, RL-aligned, and Turbo variants, with Turbo models enabling 1024² resolution generation in 0.59s and editing in 1.02s on an A100 GPU, demonstrating competitive benchmark performance despite compact size.
latent tokenizerrectified flow matchingnative-resolution packingadversarial perceptual guidancecuda kernel fusion
Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions
The paper introduces Adaptive View Retrieval, a retrieve-and-calibrate framework for detecting hateful optical illusions in multimodal content. The method assembles a complementary view bank, adaptively selects trusted views, retrieves hidden-message identities, and calibrates harmfulness. Evaluated on HatefulIllusion with a frozen CLIP encoder, it achieves 93.2% balanced accuracy, outperforming original-view baselines (20.9-24.5% accuracy) and fixed single-transform filters. It also matches or exceeds human performance on IllusionMNIST, IllusionFashionMNIST, and IllusionAnimals, demonstrating the need for hidden-meaning recovery in robust moderation.
adaptive view retrievalhateful optical illusionsmultimodal moderationhidden-message retrievalclip encoder
Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification
The study reframes machine unlearning as distribution restoration, evaluating methods against a retrained oracle in a controlled nonce-fact testbed. It reveals that common trained-probe criteria can favor methods retaining held-out knowledge (-2.82 nats below never-learned levels). The authors audit oracle-free screens and certificate-style criteria across 45 model-seed cells, finding that a base-anchored held-out screen effectively rejects injected models (45/45 cells) and detects entity-routing suppression (35/45). A damage-relative recalibration certifies a subset within retraining noise (0.80 nats), while a fixed-magnitude logit-suppression attack defeats forward-only certification in 12/45 cells.
machine unlearningdistribution restorationoracle-free certificationlogit-suppressionnonce-fact testbed
BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators
BaseRT introduces a native Metal inference runtime for large language models on Apple Silicon, leveraging M5 Neural Accelerators through hand-written Metal~4 tensor-core kernels. The framework-free design routes compute-bound matrix multiplications (including dense and mixture-of-experts GEMM and flash-attention prefill) to M5 accelerators while optimizing memory-bound decode paths. On an M5 Pro, BaseRT achieves up to 6.4× higher prompt-processing throughput and 1.75× higher decode speed versus llama.cpp across 15 models (Qwen3, Llama~3.2, Gemma~4; 1B–35B parameters), with maximal gains on mixture-of-experts architectures.
neural acceleratorsmetal apitensor-core kernelsmixture-of-expertson-device inference
Lipschitzian SLLNs for random functions
The authors establish strong laws of large numbers (SLLNs) for locally Lipschitz functions in the Lipschitz pseudometric, under either topological or model-theoretic conditions. The latter condition notably includes functions jointly definable in o-minimal structures while extending beyond this class. Key applications include uniform convergence of limiting and Clarke subdifferentials, along with finite-sample solution identification. These results delineate broad function classes where failure phenomena identified in prior work [Tian and Royset, 2025] are absent.
lipschitz pseudometrico-minimal structuresclarke subdifferentialsstrong laws of large numbersfinite-sample identification
Towards Miniature Humanoid Tele-Loco-Manipulation Using Virtual Reality and Reinforcement Learning
A novel full-body telepresence control stack is developed for miniature humanoid robots, combining Virtual Reality (VR) for upper-body teleoperation and Reinforcement Learning (RL) for lower-body locomotion. The framework is implemented on ROBOTIS OP3 hardware, enabling walking speeds up to 0.45 m/s independent of arm motions. Tele-loco-manipulation is demonstrated through a cube relocation task, where an expert operator successfully moved two 40 g cubes within 10 minutes while traversing 5 m. This system bridges the gap between full-sized humanoid capabilities and miniature platforms, showcasing potential for accessible tele-loco-manipulation in resource-constrained settings.
telepresencereinforcement learningteleoperationhumanoid roboticslocomotion
PG-KINN: A Physics-Informed Petrov-Galerkin Kolmogorov-Arnold Network for Solving Forward and Inverse PDEs
PG-KINN introduces a physics-informed Kolmogorov-Arnold Network (KAN) based on a Petrov-Galerkin formulation, addressing limitations of existing PDE solvers. The method employs KANs for the trial space and compactly supported piecewise-polynomial test spaces evaluated with Gauss-Legendre quadrature, reducing differentiation order via integration by parts while maintaining applicability to nonlinear and inverse problems. Benchmarks demonstrate PG-KINN's superior performance over MLP baselines and state-of-the-art KAN formulations (PIKAN) in tasks involving crack singularities, stress concentration, hyperelasticity, and inverse parameter identification. This approach establishes Petrov-Galerkin coupling as a robust framework for AI-driven computational mechanics.
petrov-galerkinkolmogorov-arnold networkphysics-informed learninginverse problemsweak residuals
Statevector-Referenced Geometry Survival of a Four-Qubit ZZ Quantum Kernel on IBM Quantum Hardware: A Fixed-Subset Diagnostic Across Three Execution Configurations
The study evaluates the preservation of quantum kernel geometry for a four-qubit ZZ feature-map kernel on IBM Quantum hardware (ibm_fez) using 24 indoor air-quality data windows. Three execution configurations (baseline, dynamical decoupling, gate twirling) were tested, each producing complete, positive-semidefinite Gram matrices with centered statevector geometry preserved to varying degrees (full-matrix CKA 0.933-0.989). Gate twirling showed the highest geometric fidelity but lowest kernel-target alignment, suggesting hardware distortion dominates discrepancies. Results highlight the distinction between implementation fidelity and task relevance in quantum machine learning.
quantum kernelgram matrixstatevector geometrydynamical decouplinggate twirling
Online Variance Reduction for Domain Adaptation on Streaming Data
The paper introduces ARROW, the first online stochastic variance reduction (SVR) algorithm for maximum mean discrepancy (MMD) and correlation alignment (CORAL) losses in streaming data settings. ARROW maintains moving average references of alignment statistics and adaptively reweights incoming minibatches to align them with reference statistics, using a relaxed reweighting scheme for tractability. Experiments demonstrate ARROW's competitive performance with offline algorithms in runtime, variance reduction, and target domain accuracy.
stochastic variance reductionmaximum mean discrepancycorrelation alignmentonline learningdomain adaptation
Variance-reduced Domain Adaptation using Paired Sampling
The paper introduces Paired Sampling for Domain Adaptation (PSDA), a stochastic variance reduction technique for unsupervised domain adaptation (UDA) that addresses high variance in correlation alignment and maximum mean discrepancy losses. PSDA forms quadruplets by pairing observations within and across domains, minimizing expected gradient variance through linear assignment problems. Simulations show reduced variance compared to baseline methods, and experiments on three domain shift datasets demonstrate improved target domain accuracy.
unsupervised domain adaptationvariance reductioncorrelation alignmentmaximum mean discrepancystochastic optimization
Interval and fuzzy physics-augmented neural networks (iPANN and fPANN) for uncertainty quantification and propagation in constitutive modeling
The authors propose interval and fuzzy physics-augmented neural networks (iPANNs and fPANNs) for uncertainty-aware hyperelastic constitutive modeling. iPANNs learn sparse lower, mean, and upper free energy density branches to enclose noisy stress observations, while fPANNs embed these branches into a fuzzy-set representation via alpha-cut interpolation. The models enforce mechanistic constraints (objectivity, consistency, polyconvexity) and use smoothed L0 regularization for interpretability. Evaluated on synthetic isotropic hyperelastic data with heteroscedastic noise, the framework successfully generalizes to test data and propagates uncertainty in finite element simulations.
uncertainty quantificationhyperelastic constitutive modelingphysics-augmented neural networksaleatoric uncertaintyfinite element simulations
Multi-modal transformer for signal classification in nanopore blockade experiments
The authors introduce a multi-modal transformer architecture for nanopore signal classification, combining raw time-series data, wavelet-based images, and static feature vectors. The model leverages attention mechanisms to integrate complementary information from different signal representations, with analysis revealing distinct feature emphasis between modalities. It achieves >10 percentage point improvement over existing methods on a 42-peptide benchmark and near-perfect accuracy on a 20-amino-acid dataset, demonstrating robust molecular identification capabilities.
nanopore sensingmulti-modal transformerwavelet transformsignal classificationattention mechanisms
Label-Free Finite-Volume-Residual Training of Attention Graph Neural Networks for Coupled Thermo-Fluid Fields
The authors propose a label-free training method for attention graph neural networks to predict 3D thermo-fluid fields by minimizing finite-volume method (FVM) residuals of governing equations, eliminating the need for labeled training data. The approach evaluates residuals directly on the mesh, avoiding costly data generation from conventional numerical solvers. Evaluated against computational fluid dynamics references and a supervised baseline across four scenarios, the FVM-loss model achieves 2.3-2.8% normalized root-mean-square error on steady-state benchmarks and outperforms the baseline on parametric transient cases, demonstrating accurate buoyancy-energy coupling and reduced development costs.
attention graph neural networksfinite-volume methodthermo-fluid fieldslabel-free trainingcomputational fluid dynamics
Decentralized Online Riemannian Optimization for Strongly Geodesically Convex Functions
The paper establishes the first $O(\log T)$ static regret bound for decentralized online Riemannian optimization with strongly geodesically convex (g-convex) losses, matching the minimax-optimal Euclidean rate. The authors develop a general network-error analysis for time-varying step sizes, overcoming incompatibility with fixed-step analyses in prior work. They extend this to prove $O(\log T)$ regret for both full-information and two-point bandit feedback settings, the latter via novel strong subconvexity arguments for smoothed losses.
riemannian optimizationgeodesic convexitydecentralized learningonline optimizationregret bounds
Adaptive deep nonparametric regression from dependent data under covariate shift
The paper proposes a sparse-penalized deep neural network (SPDNN) estimator for nonparametric quantile and Huber regression under covariate shift with dependent data. The method addresses distributional discrepancy between source and target domains via a two-step pre-training procedure: first estimating the density ratio using least squares SPDNN, then computing a reweighted SPDNN regression estimator. Non-asymptotic error bounds are established for Hölder smooth functions, showing adaptive minimax optimal rates (up to log factors) across i.i.d. and time series data (φ-mixing, strong mixing, C-mixing).
covariate shiftsparse-penalized dnnnonparametric regressiondensity ratio estimationmixing processes
Classical Hardware Acceleration of Quantum Autoencoders for Real-Time Anomaly Detection in Collider Experiments
The study demonstrates FPGA-accelerated quantum autoencoders for real-time anomaly detection in collider experiments, bridging quantum machine learning with classical hardware constraints. Using variational quantum circuits compiled to classical FPGA targets, the approach achieves performance comparable to state-of-the-art classical models while meeting strict latency and resource requirements for trigger systems. Results show feasibility for deployment in current data acquisition pipelines, advancing quantum readiness for high-energy physics applications with 1) efficient high-dimensional correlation modeling and 2) synthesized gate operations on low-latency hardware.
quantum autoencoderfpga accelerationanomaly detectioncollider experimentsquantum machine learning
The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability
The paper investigates long-term temporal portability of PortLLM's LoRA patches across 10 continual pretraining steps using Mistral, Gemma, and Qwen models, demonstrating persistent effectiveness without repeated fine-tuning. It provides two theoretical analyses showing that near-orthogonality in high-dimensional spaces explains PortLLM's competitive performance, offering geometric insights into the loss landscape. Empirical results confirm portability across extended durations, while theoretical work links high-dimensional geometry to adaptation efficacy.
temporal portabilityparameter efficient fine-tuningnear-orthogonalitycontinual pretrainingloss landscape
Interpretable Fuzzy Rule-Based Regression Extension for Ex-Fuzzy Library
The paper introduces an interpretable regression extension for the Ex-Fuzzy library, enabling Mamdani-style fuzzy inference with scalar consequents learned directly from data. The method employs a target-aware partition initialization strategy using Fuzzy C-Means clustering, deriving linguistic variables from an augmented input-output space to emphasize output-relevant regions. Evaluated on ten KEEL regression datasets, Gaussian partitions outperform trapezoidal partitions, achieving a mean coefficient of determination of approximately 0.86 with compact rule bases of 10-15 human-readable rules. The extension provides a transparent alternative to black-box models, balancing interpretability and predictive performance.
mamdani fuzzy inferencefuzzy c-means clusteringlinguistic variablesinterpretable regressioncoefficient of determination
Breaking the $T^{3/4}$ Barrier for Regret Minimization With Bi-Dimensional CDFs
The authors present an algorithm for regret minimization in learning cumulative distribution function (CDF)-related objectives of the form $g(x)\cdot\mathbb{P}_{X\sim\mathcal{D}}(X\le x)$ over $[0,1]^2$, where $g$ is a known Lipschitz function and $\mathcal{D}$ is unknown. Using binary feedback $\mathbb{I}(X_t\le x_t)$ at each round $t$, their method achieves $\widetilde{\mathcal{O}}(T^{7/10})$ regret, improving upon the previous $\widetilde{\mathcal{O}}(T^{3/4})$ bound and partially addressing the curse of dimensionality. The results also apply to profit maximization in repeated bilateral trade with fixed prices.
regret minimizationcdf-related objectiveslipschitz functioncurse of dimensionalitybilateral trade
Adaptive Bayesian Online Learning via Expert Aggregation
The authors propose an adaptive Bayesian online learning framework that aggregates Bayesian update rules as experts based on sequential predictive losses. The method ensures the aggregate competes with the best expert in hindsight, with aggregation cost determined by expert performance evaluation. The framework is instantiated in online conformal inference, yielding smoothed Bayesian adaptive conformal inference with long-run randomized coverage, and Gaussian process regression, achieving an oracle inequality in cumulative predictive Kullback-Leibler risk and adaptation to unknown Hölder smoothness up to logarithmic factors. Experiments demonstrate the aggregate effectively tracks strong experts without requiring oracle expert selection.
bayesian online learningexpert aggregationconformal inferencegaussian process regressionkullback-leibler risk
PhaseAware: Interpretable Human-in-the-Loop Rehabilitation Scoring with Boundary Monitoring
PhaseAware introduces an interpretable framework for continuous rehabilitation quality assessment, combining a temporal backbone with phase- and body-group descriptors via a backbone-conditioned gated residual pathway. Evaluated on UI-PRMD deep-squat protocol, it achieved an RMSE of 0.0230 (88.9% reduction vs. baseline) and demonstrated transferability to KIMORE squatting subset. The model generates structured review cues highlighting relevant movement stages and body regions, supporting clinician review and human-in-the-loop triage. Its compact design enables deployment in resource-constrained settings while maintaining interpretability.
rehabilitation scoringtemporal backbonegated residual pathwayphase-awarehuman-in-the-loop
Dynamical and Optimization Trade-offs of Levi--Civita Coordinates for Learned Close-Encounter Dynamics
The study systematically evaluates Levi--Civita versus Cartesian coordinates for learned Hamiltonian dynamics in perturbed Kepler systems with quadrupole potentials. Using analytic perturbations, Levi--Civita regularization maintains stable relative energy errors (~2.1×10⁻⁵) up to eccentricity e=0.99, outperforming Cartesian formulations by 4.7–8.3 orders of magnitude. While regularized models achieve finite rollouts in 40/40 high-eccentricity tests, they exhibit 𝒪(1) energy errors due to optimization challenges. Neural residuals fail to match analytic performance, revealing a trade-off between dynamical conditioning (improved) and optimization conditioning (worsened) in Levi--Civita coordinates.
hamiltonian dynamicslevi--civita coordinateskepler problemregularizationenergy error
PIER: Physics-Informed Environmental Retrieval for Time-Series Modeling
PIER introduces a physics-informed retrieval framework for environmental time-series modeling, addressing limitations of standard embedding-based approaches by ensuring physical consistency. The method augments embedding retrieval with a physics-aware stream that scores candidates based on flux-response consistency, using local verifiers trained on physics-derived flux features, and employs a weight adjustment mechanism to balance the two streams adaptively. Evaluated on 356 lakes across the Midwestern United States over 41 years, PIER consistently outperforms baselines in predicting water temperature and dissolved oxygen, demonstrating its effectiveness as a general augmentation strategy across diverse backbones.
physics-informed retrievalflux-response consistencyembedding-based retrievallocal verifiersweight adjustment mechanism
User-Centric Modeling of Transactional Sequences with Explainable State Space Models
The authors introduce a hybrid approach combining contrastive representation learning (CoLES) with State Space Models (SSMs) for user-centric modeling of transactional event sequences. The method leverages Mamba, a selective SSM, to address limitations of RNNs and Transformers in handling long-range dependencies. Two integration strategies are explored: initializing Mamba's hidden state with CoLES embeddings and prepending projected CoLES embeddings as prefix tokens. Experiments on Age, MBD, and Taobao datasets show consistent performance improvements over standalone Mamba and CoLES, with 2--3× faster convergence. Explainability analysis via discretization-step maps and Integrated Gradients highlights selective event filtering and identifies informative transaction features.
contrastive representation learningstate space modelstransactional sequencesselective ssmsintegrated gradients
Statistical Inference for Rank Allocation in Low-Rank Adaptation
The paper introduces StatLoRA, a statistical inference-based method for rank allocation in low-rank adaptation (LoRA) of large language models. By formulating rank allocation as a hypothesis testing problem, StatLoRA uses p-values derived from asymptotic normality of optimizer trajectories (including AdamW) to prune or retain LoRA components under fixed rank budgets. Evaluated on DeBERTaV3-base, BART-Large, and Qwen2.5-7B across NLU, NLG, and QA tasks, StatLoRA matches or outperforms vanilla LoRA, AdaLoRA, and IGU-LoRA while maintaining theoretical guarantees via central limit theory for stochastic optimizers.
low-rank adaptationstatistical inferencehypothesis testingasymptotic normalityparameter-efficient fine-tuning
OLEDLM: A Unified Language Model for OLED Molecular Design
We propose OLEDLM, a unified language model for OLED molecular design that generates SMILES sequences satisfying target optoelectronic properties. The framework employs a multi-stage strategy: (1) a LLaMA-style transformer establishes a foundational chemical language model, (2) a BERT-based property predictor is fine-tuned on a large OLED dataset, (3) reinforcement learning optimizes SMILES generation using the predictor, and (4) DFT verification validates structural validity and property optimization. Results demonstrate efficient navigation of OLED chemical space, generating novel candidates with high structural validity and optimized optoelectronic properties.
smiles sequencesllama-style transformerbert-based property predictorreinforcement learningdft verification
On Optimization Complexity of Second-Order Certified Unlearning
The paper establishes optimization complexity bounds for certified machine unlearning, formalizing the dual objectives of data removal and model accuracy. Using uniformly convex regularizers, the authors derive distance bounds between initial and unlearned models via a novel generalization error substitute. They propose a second-order unlearning algorithm with an anisotropic Gaussian mechanism, achieving state-of-the-art global convergence. Theoretical analysis demonstrates fast certified unlearning rates for linear models with quasi-self-concordant losses, including logistic and exponential regressions, with proven advantages over first-order methods.
certified unlearninguniformly convexanisotropic gaussian mechanismquasi-self-concordantoptimization complexity
Instance Hardness-Based Relevance for Imbalanced Regression
This study introduces Instance Hardness-based Relevance (InHaR), a novel relevance function for identifying rare instances in imbalanced regression problems. InHaR incorporates learning difficulty alongside target distribution, addressing limitations of traditional methods that rely solely on target values, particularly in bimodal distributions. The proposed approach guides resampling strategies like Random Oversampling (RO) and Gaussian Noise (GN), significantly improving predictive performance. Experimental results demonstrate InHaR's effectiveness in correctly identifying rare regions under bimodal distributions. Code and datasets are publicly available.
imbalanced regressioninstance hardnessbimodal distributionsrelevance functionresampling strategies
Hard Guarantees at a Measured Price: Entropy-Stable Learned Finite Volumes for Compressible Flow
We introduce a learned finite volume scheme for the 2D Euler equations on unstructured meshes, designed to be physically admissible by construction with entropy-stable interior fluxes. The method employs pre-defined evaluation protocols, including iso-cost comparisons against classical baselines and factor decomposition of learned components. Results show that the unlearned skeleton outperforms at equal mesh resolution, while learned components yield robust gains only on unseen boundary conditions (10.8%). The scheme maintains zero negativity events across all rollouts, including Mach extrapolation and unseen wall cases. Inference-time corrections improve performance on Mach extrapolation, and a spatial gate activating learned components near walls enhances transferability to new geometries.
entropy-stableeuler equationsiso-cost comparisonmach extrapolationunstructured meshes
Plausibility-Driven Prioritization of Candidate Biomedical Annotations
The authors propose a plausibility-driven framework for prioritizing biomedical annotations by leveraging knowledge graph embeddings and relation-specific classifiers. Their method combines classifier confidence, reliability metrics, and semantic context from alternative relationships in biomedical knowledge graphs (bioKGs), using a community-based negative sampling strategy to improve classifier robustness. Evaluations on five bioKGs show a 5.8% average increase in balanced accuracy, with plausibility measures outperforming raw classifier confidence for annotation prioritization. The approach enhances curation efficiency while maintaining expert oversight.
biomedical knowledge graphsnegative samplingplausibility measuresrelation-specific classifiersannotation prioritization
Self-organizing Architecture of Receptron Units: a Hardware-Aware Framework for Edge Intelligence
The authors propose a neuromorphic classifier called Receptron, a single-unit architecture capable of learning non-linearly separable decision boundaries without multi-layer networks, targeting deployment on resource-constrained microcontroller units (MCUs) with continuous on-device adaptation. The hardware-aware design avoids conventional deep learning approaches while maintaining compatibility with standard ML baselines on basic benchmarks, achieving cross-validated accuracies suitable for edge intelligence in dynamic environments. Results demonstrate its viability as an interpretable alternative for neuromorphic edge systems.
neuromorphic computingedge intelligencemicrocontroller unitsnon-linear separabilityon-device adaptation
Local Stability and Gaussian Smoothing of Quantized Neural Networks
The paper proposes Gaussian averaging as a smooth surrogate for quantized neural networks, deriving a dimension-dependent bound on the difference |f-g| between original and smoothed functions under local oscillation constraints. It provides closed-form Gaussian averages for ReLU and sign activations, demonstrating the approach on a high-dimensional binary perceptron where layer-preactivation aggregation with quantization-noise surrogates yields Gaussian envelopes for both inference smoothing and training gradients.
gaussian smoothingquantized networkslocal stabilityrelu activationbinary perceptron
Multi-stage Dynamic Selection for Cross-Project Defect Prediction
The paper introduces a novel Cross-Project Defect Prediction (CPDP) framework addressing distribution shift via a two-stage multiple classifier system (MCS) selection scheme. The first stage evaluates MCS configurations across training projects to ensure diversity and generalization, while the second stage selects classifiers at the module level during testing, enhancing robustness to distribution changes. Experiments on 82 projects from four CPDP benchmark datasets demonstrate superior performance over state-of-the-art methods. The code and dataset are publicly available.
cross-project defect predictiondistribution shiftmultiple classifier systemmodule-level selectionbenchmark datasets
CURED: Creating, Understanding, and Repairing Errors Demonstrator
The CURED demonstrator integrates machine learning-based data cleaning and error modeling into a unified web application for tabular data. Users can upload datasets, introduce realistic data-dependent errors, and apply modern ML techniques to detect, clean, and analyze error mechanisms. The tool bridges theoretical advancements in ML and DBMS with practical insights, enabling intuitive exploration of error models and cleaning algorithms. Available at https://cured.demo.calgo-lab.de/, CURED facilitates the study of statistical learning methods in error detection and repair for data-intensive applications.
tabular dataerror detectionmachine learningdata cleaningerror modeling
HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation
HeadCast introduces a training-free, plug-and-play framework to accelerate autoregressive video diffusion models by optimizing attention head usage. The method classifies attention heads into four archetypes—Sink, Dummy, Spatial, and Global—based on their behavior during inference, restructuring the Key-Value cache into head-specific pathways. This approach maintains long-range temporal consistency via Global heads while reducing computational costs, particularly at higher resolutions. HeadCast achieves inference speedups of up to 1.62x at 720P and 1.95x at 1080P across state-of-the-art models, preserving video quality and minimizing flicker without requiring model retraining.
autoregressiveattention headskv cachevideo synthesisdiffusion models
Autonomous Collaborative Learning Among an Ensemble of Tsetlin Machines with Consensus-Based Inference
The paper proposes a decentralized collaborative learning framework for Tsetlin Machines (TMs) using consensus-based inference under vertical feature partitioning. Each agent maintains a private TM model without raw data exchange, combining predictions through global consensus to accommodate heterogeneous agents with varying data distributions or computational resources. Experiments on grid and graph network topologies show classification accuracies comparable to centralized models, demonstrating effective information fusion in multi-modal sensing environments.
tsetlin machinedecentralized learningconsensus-based inferencevertical partitioningmulti-modal fusion
Directional Kernel Mean Difference: A Fast Signed Statistic for Univariate Distribution Comparison
The authors propose Directional Kernel Mean Difference (DKMD), a signed statistic for univariate distribution comparison that preserves directional information of distributional shifts. DKMD integrates kernel mean embedding differences against an odd weighting function, ensuring antisymmetry, immunity to symmetric differences, and directional monotonicity under stochastic dominance. They develop an $O(N \log N)$ prefix-suffix scanning algorithm with $O(N)$ memory, demonstrating robustness to heavy-tailed outliers and scalability to millions of samples in synthetic benchmarks.
kernel mean embeddingmaximum mean discrepancystochastic dominanceunivariate distributionsigned statistic
Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting
The paper introduces cumsum-composable phase transport, a streaming-optimized temporal layer for keyword spotting that combines unitary rotations, prefix differences, and gated residual updates. The method enables exact batched training via cumulative sums and efficient online inference with single-frame updates, maintaining well-conditioned prefix terms through unitary transport constraints. Evaluated on Google Speech Commands v2 (12 labels), the approach achieves 97.3% accuracy with 51.6K parameters and 96.8% with 24.8K parameters, matching or exceeding MelCNNMaxPool baselines while reducing latency from 7.09 ms to 5.01 ms on a Tesla T4.
cumsum-composableunitary rotationsprefix differenceskeyword spottingstreaming inference
Non--negative matrix factorization using the \textit{R} package \textsf{nnmf}
The study introduces a new R package for non-negative matrix factorization (NMF) and systematically compares its performance against two existing R packages using real-world data. Evaluations focus on computational efficiency, convergence behavior, reconstruction accuracy, memory utilization, and factorization stability, providing objective guidance for package selection. Results highlight the new package's performance under practical conditions characterized by complexity, heterogeneity, and noise, addressing a gap in comprehensive evaluations of NMF implementations.
non-negative matrix factorizationdimensionality reductioncomputational efficiencyreconstruction accuracyconvergence behavior
Evaluating and Mitigating Gender Bias in Pre-trained Embeddings for ML-based Recruitment
The paper evaluates gender bias in ML-based recruitment systems using pre-trained embeddings, proposing a multi-task adversarial learning framework with gradient reversal to mitigate bias while preserving predictive utility. Experiments on the synthetic FairCVdb dataset assess nine embedding models, comparing gender leakage in original and scrubbed biographies. Results show gender scrubbing reduces but doesn't eliminate bias, while adversarial learning improves fairness primarily on original texts, suggesting complementary use with text-level debiasing.
gender biaspre-trained embeddingsadversarial learninggradient reversalfaircvdb
Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design
The authors present AAMFM, an Antigen-specific Antibody Multimodal Foundation Model that learns unified representations of antibody sequences and structures conditioned on antigen context. The model incorporates antigen geometric interfaces and epitope annotations via a cross-modal adapter, enabling joint antibody-antigen interaction modeling, and is fine-tuned using Calibrated Direct Preference Optimization (Cal-DPO) for binding-specific objectives. Experiments show AAMFM achieves state-of-the-art performance in functional antibody design, demonstrating its potential for antigen-specific engineering.
antibody designmultimodal foundation modelantigen contextcalibrated direct preference optimizationepitope annotation
PN-QNN: Harnessing Physical Noise as a Native Regularizer in Photonic Hybrid Quantum Neural Networks
The study investigates physical noise as a hardware-native regularizer in photonic hybrid quantum-classical neural networks (PHQCNNs), contrasting traditional noise suppression approaches. Using Quandela's Perceval simulator and MerLin framework, PHQCNNs were trained on Iris, Digits, and MNIST datasets with a seven-parameter physical noise model injected during training. A genetic algorithm optimized six continuous noise dimensions and one boolean parameter to maximize validation accuracy across five seeds. Results show modest accuracy gains on Iris (+0.82pp) and Digits (+1.45pp) but degradation on MNIST (-1.21pp). Per-parameter sweeps reveal no consistently beneficial noise parameter, while second-order loss expansion indicates a Tikhonov-like regularization effect, dataset-dependent.
photonic hybrid quantum-classical neural networksphysical noisegenetic algorithmtikhonov regularizationperceval simulator
Zero-Shot Heart Rate Variability Forecasting from Consumer Wearables Using Time Series Foundation Models
This study evaluates zero-shot forecasting of Heart Rate Variability (HRV) from consumer wearables using Time Series Foundation Models (TSFMs). Three TSFMs—TimesFM, Chronos, and MOIRAI—were benchmarked against traditional methods on fragmented, artifact-rich HRV data from 49 healthy individuals. A variability-preserving imputation method augmented linear interpolation with locally adaptive stochastic noise to retain physiological dynamics. Without fine-tuning, TSFMs outperformed baselines, achieving Mean Absolute Scaled Error (MASE) between 0.81 and 0.87 across context lengths of 32 and 64 time steps, with Chronos and TimesFM leading. Results highlight TSFMs' potential for clinical deployment with domain-specific fine-tuning.
heart rate variabilitytime series foundation modelsmean absolute scaled errorvariability-preserving imputationzero-shot forecasting
Generalized Kalman filter based temporal difference reinforcement learning
The authors propose a generalized temporal-difference (TD) reinforcement learning framework based on conditional expectations, extending classical Kalman-based TD learning to nonlinear models and non-Gaussian distributions. The method recursively estimates both the conditional expectation and second moment of value/Q-functions, quantifying uncertainty via polynomial chaos expansions or ensemble approximations. Demonstrated on linear mass-spring-damper and nonlinear heat conduction problems, the approach accurately estimates value functions and their uncertainty while generalizing Kalman-TD methods.
temporal-difference learningkalman filterconditional expectationpolynomial chaosuncertainty quantification
Good Practice Guide for quantifying uncertainties for machine learning models applied to photoplethysmography signals
The QUMPHY project introduces a Good Practice Guide for uncertainty quantification in machine learning models applied to photoplethysmography (PPG) signals from wearable devices. The guide evaluates various machine learning models for regression and classification tasks, detailing both model-dependent and model-independent uncertainty quantification techniques. It includes validation methods for these techniques and presents six benchmark problems with associated datasets. Additionally, the guide provides software tools for implementation and addresses ethical considerations. Recommendations are offered to assist practitioners in effectively applying these methods.
uncertainty quantificationphotoplethysmographymachine learningregressionclassification
Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction
The paper introduces Diffusion ReRoll, a diffusion-based framework for robotic sequential prediction that enables revisable denoising through selective re-noising of locally stable regions. Unlike monotonic denoising approaches, Diffusion ReRoll iteratively refines segments by re-noising them in context with the rest of the horizon, maintaining local consistency while allowing cross-horizon revision. Evaluations on OGBench PointMaze, AntMaze, and LIBERO-10 show relative success rate improvements of 21-23% over Diffusion Forcing and Diffuser in planning tasks, and 56.5% over Diffusion Policy in action prediction, with enhanced performance in video-action consistency and out-of-distribution scenarios.
diffusion modelsrobotic sequential predictionrevisable denoisingcross-horizon revisionlong-horizon planning
Nonlinear Bias-Compensated Adaptive Filter and Its Application for Time-Series Prediction
The paper proposes the Random Fourier Bias-Compensated Filter under General Adaptive Function (RFFBCGA) algorithm to address limitations in nonlinear adaptive filtering. It combines a fixed-size random Fourier feature framework with bias compensation for input noise mitigation and employs a general adaptive function for robustness against non-Gaussian output noise. Evaluations on synthetic and real-world time-series prediction tasks demonstrate superior performance compared to existing methods like BCKLMS.
nonlinear adaptive filteringbias compensationrandom fourier featureserrors-in-variablestime-series prediction
Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage
The paper identifies correlated agreement blindness (CAB) as a structural blind spot in multi-agent systems where improved base learners' convergence weakens safety monitoring. It proposes ARAT, a directed-star architecture combining a Random Forest agent, k-NN agent, and calibrated meta-model to mitigate CAB. On UNSW-NB15 (82,332 samples), ARAT reduces under-prediction from 4.80% to 1.70% via conservative override and safety-flag gates, showing 57.2% of errors occur under agreement. Cross-dataset validation confirms diversification only improves safety when generating productive disagreement.
correlated agreement blindnessmulti-agent arbitrationsafety monitoringdirected-star architectureproductive disagreement
Local Causal Structure Learning in the Presence of Latent Variables and Selection Bias
The authors propose LoCaLS, a local causal structure learning algorithm that identifies direct causes and effects of a target variable in the presence of latent variables and selection bias. LoCaLS operates by characterizing a local region for target-specific causal discovery without reconstructing the global structure, establishing a theoretical bridge between local and global causal information. Experiments on random and real-world structures show LoCaLS achieves higher structural accuracy than existing local methods while requiring less computational effort than global methods. Applications to gene expression datasets demonstrate its effectiveness in large-scale biological data analysis.
causal discoverylatent variablesselection biaslocal structuregene expression
Adversarial Frontiers: Minimum-Norm Attack Ensembles for Robustness Evaluation
(No summary returned.)
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence
(No summary returned.)
Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data
The authors present a hypothesis-and-refinement learning framework for molecular structure determination from multimodal spectroscopic data, addressing the underdetermined inverse problem through integration of spectral evidence with large-scale molecular priors. They introduce QM9SPIN, a DFT-derived dataset with diverse 1D/2D NMR spectra, and SpectroMol, a spectrum-to-structure model generating chemically valid hypotheses. Combined with MS-Mol2Mol, a mass-constrained molecular generator trained on 400M compounds, the system achieves 93.8% top-1 accuracy on simulated benchmarks and demonstrates effective adaptation to experimental spectra with limited fine-tuning.
spectroscopic datamolecular structure elucidationhypothesis-refinementmultimodal learningconditional generation
Dreamer-CPC: Message Learning with World Models for Decentralized Multi-agent Reinforcement Learning
Dreamer-CPC introduces a decentralized model-based multi-agent reinforcement learning method integrating Collective Predictive Coding (CPC) into DreamerV3's world model for message learning. Agents independently maintain world models and message modules, inferring and exchanging messages from latent states reflecting historical observations and actions. Evaluated in Observer (non-cooperative information-sharing) and CatchApple (temporarily missing task-relevant observations), Dreamer-CPC outperformed IPPO-CPC and no-communication baselines, achieving 4-5 times higher episode returns in CatchApple. Results demonstrate that latent dynamics-based communication enhances decentralized decision-making when current observations are insufficient.
multi-agent reinforcement learningcollective predictive codingworld modellatent statesdecentralized decision-making
Zero-Observation User Reactivation with Gap-Driven Dimensional Gating
The paper introduces Zero-Observation Reactivation, a sequential recommendation scenario where users return after long inactivity gaps (Δ𝑡). It proposes ΔGate, a lightweight output-layer plugin that routes representations between personalized history and a global prior, conditioned on Δ𝑡 and user embeddings. Evaluated under the Gap-Synthesize Protocol on three Amazon datasets, ΔGate improves Hit@10 by 52-84% for gaps >365 days (e.g., 0.047 vs. 0.031 for SASRec) while adding only 66K parameters (2-4% overhead). The frozen backbone design prevents embedding drift and enables interpretable dimension-wise routing.
sequential recommendationzero-observation reactivationdimensional gatinggap-synthesize protocolhit@10
TriAgent: Divergence-Aware Multi-Agent Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent committee for financial sentiment analysis, stratified by contextual granularity: VADER (word-level), FinBERT (sentence-level), and Qwen2.5 (cross-sentence). A Semantic Divergence Index (SDI) routes queries based on pairwise disagreement across granularities. Key findings include an F1 plateau at ~0.87 when Qwen LLMs act as critics, multilingual sentence-BERT cross-border canonicalization achieving F1=0.99, SDI doubling as a hallucination detector (AUC=0.90), and superior risk-adjusted returns (Sharpe=3.50). TriAgent saves $9.3M/year at 10M-user scale compared to GPT-4o-mini.
semantic divergence indexgranularity stratificationmultilingual sentence-berthallucination detectionrisk-adjusted return
A Structure-Adaptive Random Feature Method for High-Dimensional Elliptic PDEs
We propose the Hierarchical Analysis-of-Variance Random Feature Method (HA-RFM) for solving high-dimensional elliptic PDEs, which adapts to lower-dimensional structure by selecting coordinate blocks via Sobol indices, extracting oblique low-rank features from predictor gradients, and coupling features in a regularized least-squares solve. Theoretical analysis establishes L2 error bounds linking solution truncation, finite-width approximation, and regularized fitting, with polynomial width scaling in dimension under fixed interaction order. Experiments demonstrate exact recovery of three-pair supports, oblique direction recovery up to dimension 50, and error reductions of 14-100x over full-dimensional RFM. HA-RFM extends to semilinear computations up to dimension 100 and delineates coordinate families for broader structure.
elliptic pdessobol indicesoblique low-rank featuresregularized least-squaresfinite-width approximation
A Multiclass Quantum Aligned Centroid Kernel
The authors introduce McQuack, a trainable quantum kernel method for multiclass classification that addresses three limitations of conventional kernels: quadratic scaling, fixed kernels, and lack of intrinsic multiclass formulation. McQuack replaces the full Gram matrix with a linear-scaling sample-to-centroid fidelity matrix, evaluated via quantum circuits. Simulations show it outperforms pure quantum baselines, while hardware experiments on IBM devices (124 qubits) achieve RBF-comparable performance without training. Trainability analysis reveals no barren plateaus in 13-qubit circuits, with parameter initialization critical for optimization.
quantum kernelmulticlass classificationtrainable kernelcentroid fidelitybarren plateaus
AlphaRoute: Large Language Models as Semantic Optimizers for Multi-Objective Routing
AlphaRoute introduces a multi-objective adaptive search framework for VLSI global routing, reformulating rip-up and reroute (R&R) into a dynamic optimization system. The method employs SHAP-based overflow decomposition to isolate per-net congestion, enabling targeted subgraph extraction via 3D Dijkstra maze routing and an adaptive PathFinder policy. Crucially, AlphaRoute leverages Large Language Models (LLMs) as semantic policy optimizers, dynamically adjusting penalty parameters based on congestion metrics within a deterministic knowledge graph. Evaluated on ISPD 2025 benchmarks, AlphaRoute reduces overflow by 98.6% on MEMPOOL and achieves a 29.8x reduction in overflow on the constrained ARIANE design, yielding a penalized score of 0.0538 versus the SOTA 1.780.
vlsishapdijkstrallmispd
Machine Can Automatically Discover Parametric Functions to Model HEP Data
The study introduces SymbolFit, a symbolic regression package that automates the discovery of parametric functions for modeling high-energy physics (HEP) data, eliminating the need for manual iterative fitting. The method combines symbolic regression with uncertainty modeling to perform a data-driven search over function space. Evaluated on CMS and ATLAS Run 2 dijet spectra, SymbolFit generated over 1000 functions with χ²/NDF ≈ 1 across 560 runs, successfully rediscovering the dijet and UA2 functions used in published searches in 111 cases.
symbolic regressionhep dataparametric functionsuncertainty modelingdijet spectra
Domain-Adapted Power Curve for Cross-Farm Applications
The paper proposes a domain adaptation approach for cross-farm transfer of wind turbine power curve models, addressing limitations of traditional distance/layout-based methods. By defining domains via temporal environmental and spatial terrain variables, the method learns a similarity metric to adapt source-farm models to target farms. Empirical evaluations demonstrate consistent outperformance over competing approaches in site-planning power prediction tasks.
domain adaptationpower curve modelingcross-farm transferwind energysite-planning
Analytic Distribution of Classifier-Free Guidance for Schedule Design
The paper derives exact analytic representations of classifier-free guidance (CFG) distributions in diffusion models, revealing that CFG modifies the base distribution via an exponential path-integral correction dependent on guidance weight $ω(t)-1$. This analysis motivates Distribution-Guided CFG (DG-CFG), a novel schedule that balances timestep contributions while accounting for score-error amplification. Experiments on a toy model and Stable Diffusion 1.5 demonstrate DG-CFG's improved generation quality and diversity-fidelity trade-off, particularly under strong guidance, while reducing sampling steps needed to achieve target metrics.
classifier-free guidancediffusion modelsprobability flow odepath-integral correctionsampling efficiency
Koopman Dreamer: Spectrally Constrained Latent Dynamics for Stable World-Model Imagination
Koopman Dreamer introduces a spectrally constrained latent dynamics model for stable world-model imagination in continuous control, addressing limitations in modal persistence and error accumulation during long rollouts. The method employs a Koopman-inspired backbone with 2D rotation-scaling blocks for damping and periodic modes, combining linear/bilinear action terms and stochastic-state modulation. It uses multi-objective training (EMA teacher targets, consistency, rollout, observation-prediction) and derives a rollout-error bound separating spectral and stochastic effects. Experiments on DeepMind Control Suite and UAV-LiDAR navigation show improved rollout stability and superior closed-loop control performance.
koopman dreamerspectrally constrained dynamicslatent world modelcontinuous controlerror accumulation
How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
The study benchmarks reward model inference speeds in RLHF pipelines, demonstrating that current PyTorch defaults are suboptimal. Authors develop a C++/ONNX Runtime engine, verifying numerical equivalence (CPU: 5.7e-6, GPU: 4.2e-3 error vs PyTorch). On CPU, their solution outperforms PyTorch eager mode, torch.compile, and FastAPI with non-overlapping confidence intervals; GPU tests showed torch.compile superiority. Key findings reveal ONNX Runtime (not C++) drives speedups, and batching strategy dominates language/runtime choices. Results derive from statistically robust repeated trials.
rlhfonnx runtimeinference optimizationreward modelingbatching strategy
Efficient Clustering with Provable Guardrails for LLM Inference at Scale
The paper introduces a scalable clustering method for efficient LLM inference that enforces strict per-sample quality guarantees. The two-stage algorithm first applies Mini-batch K-Means for initial clustering, then selects representatives via a Johnson-Chvatal heuristic for Set Cover over alpha-balls in embedding space, ensuring minimal within-cluster similarity and exact categorical attribute matching. With asymptotic complexity linear in sample size when cluster count grows proportionally, the method demonstrates 10-1000x speedup over baselines while scaling to 38 million samples, reducing downstream costs by 50x in a production recommender system.
llm inferencerepresentative clusteringset cover heuristicsimilarity guardrailsscalable clustering
Data-Poisoning Audits for Causal Effect Estimation
The authors introduce a data-poisoning audit framework for augmented inverse-probability-weighted estimation in causal effect analysis, addressing vulnerabilities from append-only attacks in pooled observational data. The method involves specifying a catalog of feasible records, an append budget, and nested source capacities, with the adversary selecting records to maximize treatment effect movement. A greedy scan computes exact finite-sample worst-case movement, while a total-influence score accounts for nuisance refitting effects. Simulations validate the exact result, showing improved local refit prediction with total influence, and demonstrate material sensitivity at small append budgets in multisite and public-data analyses. The framework aids in reliable causal reporting and source-level safeguard design.
data-poisoningaugmented inverse-probability-weightedappend-only attacksnuisance refittingtreatment effect
Optimal Recalibration of an Online Predictor
(No summary returned.)
Multi-Mask Diffusion Language Models for Few-Step Generation
The paper introduces Multi-Mask Diffusion Models (MultiMDM), a novel approach addressing the challenge of few-step generation in masked diffusion models (MDMs). Unlike MDMs where forward trajectories collapse to a single masked state, MultiMDM preserves masking structure by pushing clean tokens toward designated masks before mixing over the mask set, enabling drafting capability in the backward process. The authors derive a closed-form ELBO training objective compatible with pretrained MDMs and propose a discrete-state consistency distillation scheme with shared-Gumbel coupling. Experiments demonstrate MultiMDM's effectiveness in pretraining and distillation for few-step generation.
masked diffusion modelsfew-step generationconsistency distillationelbo trainingshared-gumbel coupling
Nuclear Quantum Effects as a Denoising Problem
The authors propose a denoising framework for capturing nuclear quantum effects by decomposing the quantum Boltzmann distribution into a classical Boltzmann-trained denoiser and an analytic Gaussian component encoding quantum context. This composition is exact when training noise does not exceed intrinsic quantum uncertainty, enabling transfer across temperature, isotopic mass, dissipation strength, and path boundary conditions without retraining. Theoretical and numerical results demonstrate exact transfer, including end-to-end displacement and momentum distributions from open imaginary-time paths. The approach unifies generative modeling noise with quantum fluctuations via their shared quadratic structure.
nuclear quantum effectsdenoisingboltzmann distributionimaginary-time path integralsquantum fluctuations
Leveraging ECRAM for Edge Continual Learning
CLASP introduces the first end-to-end system with in-memory computing (IMC) acceleration for continual learning, addressing challenges of noisy computation and inefficient training in IMC architectures. The system leverages a co-designed hardware-software framework centered on a back-end-of-line compatible ECRAM device, enabling software-visible assembly-level instructions for diverse continual learning algorithms. CLASP achieves near-GPU accuracy while delivering a 67x speedup and 132x energy savings for learning without forgetting and experience replay tasks on MNIST.
continual learningin-memory computingecramedge computingmnist
Expert-Guided Forecast Editing for Time-Series Foundation Models
The paper introduces DEFT, a framework for expert-guided forecast editing in time-series foundation models that balances exploitation of model predictions with structured exploration. DEFT decomposes candidate trajectories into trend and seasonal components, allowing expert feedback to be reused across recombined components while keeping the foundation model frozen. Evaluated on 78 datasets with three foundation models and seven query budgets, DEFT outperforms best-of-N, cross-entropy methods, and Bayesian optimization in leveraging expert feedback, with a molecular-dynamics case study suggesting broader applicability to physically grounded feedback.
time-series foundation modelsforecast editingexpert feedbacktrend-seasonal decompositionquery budget
HypEMBER: Hypernetwork-based Ensemble for Robust Policy Learning of Parametrized Dynamical Systems
HypEMBER introduces a hypernetwork-based ensemble reinforcement learning framework for robust control of parametrized dynamical systems under measurement and model uncertainties. The method employs hypernetworks to generate policy and value function weights conditioned on physical parameters, enabling parametric generalization across dynamical regimes, and utilizes ensemble learning to quantify epistemic uncertainty for improved exploration and robustness. Evaluated on the Kuramoto-Sivashinsky equation and a particle-navigation task in a gyre flow, HypEMBER demonstrates enhanced training stability, sample efficiency, and robustness to uncertainties compared to state-of-the-art RL methods.
hypernetworkensemble learningreinforcement learningparametric generalizationepistemic uncertainty
From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory
The authors propose an Unequal Error Protection (UEP) codec for DNN inference memory by exploiting bit-position sensitivity across floating-point formats. They characterize fault sensitivity through 16 ML workloads, identifying a sharp transition where flipping least-significant bits (below data-type-specific thresholds FP16:6, BF16:4, FP32:15) causes <1% task degradation, while exponent-mantissa boundary flips induce catastrophic failure. This enables selective ECC protection, saving 37.5-62.5% storage overhead without retraining. A dual-partition SRAM architecture implements UEP, validated by 870+ fault-injection runs, reducing ECC area by 27.8% and BF16 read energy by ~17% with 4% macro-area overhead.
unequal error protectionbit-position sensitivitydual-partition sramfault injectionecc overhead
The Mechanism Matters: When Knowledge Graphs Help Reinforcement Learning
The study systematically evaluates how knowledge graphs (KGs) affect reinforcement learning (RL) performance across varying task structures, injection mechanisms, and KG quality. Using controlled MiniGrid experiments and a clinical sepsis case study, it demonstrates that structured KG guidance improves sample efficiency (70% to 97% solve reliability) when task-relevant, with benefits collapsing under edge permutation (p<0.01). Mechanism choice critically impacts safety: soft methods (e.g., reward shaping) tolerate incorrect knowledge, whereas hard masking fails catastrophically with incomplete KGs. Results provide actionable guidelines for KG use in RL.
knowledge graphsreinforcement learningaction maskingreward shapingsample efficiency
CRB-Driven Beamforming and Trajectory Optimization for UAV-assisted ISAC System
The paper proposes a UAV-assisted ISAC system that jointly optimizes trajectory and beamforming to minimize Cramér-Rao bound (CRB) for angle-of-arrival estimation while maintaining downlink communication. A two-stage approach combines null-space projection for beamforming design and deep reinforcement learning for discrete-time trajectory optimization under power and mobility constraints. Simulations show a 10% reduction in time-averaged CRB compared to non-UAV-assisted systems, outperforming fixed-trajectory and maximum-ratio-transmission benchmarks.
uav-assisted isaccramér-rao boundnull-space projectiondeep reinforcement learningangle-of-arrival estimation
Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models
The paper establishes hypernetworks as a scalable method for train-time knowledge injection in large language models (LLMs), decoupling injection capacity from model capability to study scaling laws. Using a hypernetwork to generate LoRA adapters for a target LLM, the authors evaluate performance on MegaWikiQA, a multi-hop QA dataset with 39 domains from Wikidata5M. Results show power law scaling across hypernetwork depth, width, and target model size, with superior out-of-distribution generalization compared to LoRA fine-tuning and full fine-tuning.
hypernetworksknowledge injectionscaling lawslora adaptersout-of-distribution generalization
Deep Shape Regression for Planar Curves with Multimodal Covariates
The authors propose a deep shape regression model for open planar curves that handles multimodal, high-dimensional covariates while preserving geometric invariance. Representing curves as complex-valued functions, they derive the conditional full Procrustes mean as the leading eigenfunction of the conditional covariance, estimated via a novel deep conditional covariance smoother with modality-specific encoders (e.g., splines for scalars, CNNs for images). The model is invariant to translation, rotation, and scaling, accommodates sparse/irregular sampling, and includes an elastic mean estimation algorithm. Evaluated on simulated outlines and hippocampal shapes from ADNI, it recovers covariate effects consistent with prior literature.
procrustes meanshape regressionconditional covariancemodality-specific encoderselastic alignment
A Deep Learning Framework for Predicting Solar EUV Irradiance During Significant Flares
FlareEUV introduces a multimodal deep learning framework for predicting daily extreme ultraviolet (EUV) irradiance at 6.5 nm over three days during significant solar flares, using NASA's Solar Dynamics Observatory (SDO) data. The method processes 13 co-aligned full-disk images (8 AIA EUV/UV and 5 HMI magnetic/continuum products) through a lightweight attention-based architecture to learn magnetic structure-coronal emission relationships. Evaluated on 33 significant flares (2011-2014), FlareEUV outperforms baselines in short-term EUV irradiance forecasting.
solar flareseuv irradianceattention-based architecturemultimodal deep learningsolar dynamics observatory
End-to-End Differential Privacy in Training Deep Neural Network Classifiers
A novel differentially private training framework is proposed that privatizes training inputs while keeping labels public, addressing conservatism in existing methods. The approach applies the Dirichlet mechanism to randomize softmax outputs during training, enforcing differential privacy for inputs via mappings onto the unit simplex. Tight privacy bounds are derived using Rényi differential privacy to account for data reuse across epochs. Empirical evaluations on CIFAR10, MNIST, MedMNIST, FashionMNIST, and SVHN demonstrate state-of-the-art accuracy across all privacy budgets, notably improving CIFAR10 accuracy from 78.37% to 88.17% at ε=4 with δ=10^-5.
differential privacydirichlet mechanismsoftmax outputrényi differential privacyunit simplex
On the Computational Complexity of Structural Generalization
The paper formally defines structural generalization by translating its two premises—compositional structure and unbounded generalization—into mathematical terms. It contrasts the computational lower bound NC^1 with the learnable ceiling TC^0 of pure Transformers, proving that under standard assumptions, pure Transformers cannot learn structural generalization due to this complexity gap. Neuro-symbolic systems achieve superior benchmark performance by injecting semantic projections (G_γ), bypassing the computationally hard aspect. The paper clarifies that benchmark scores cannot distinguish learned from given capabilities, emphasizing the limitations of current evaluation methods.
structural generalizationnc^1tc^0neuro-symbolic systemstransformers
Machine-learned syndrome post-selection for reliable quantum error correction
We introduce a decoder-agnostic post-selection method for quantum error correction that learns directly from syndrome data without requiring logical-error labels or correction operators. The approach trains a supervised classifier to distinguish between syndromes from low- and high-noise regimes, using the classifier's output as an abort score for new runs. Validated across circuit-level simulations of the Gross bivariate-bicycle code, code-capacity simulations of the surface code, and experimental data from the QuEra neutral-atom processor, the method reduces conditional logical error rates at fixed acceptance rates, outperforms syndrome-weight filtering, and reveals a distinct post-selection transition in the surface code. Combined with logical-gap filtering, it improves output fidelity beyond standalone logical-gap use.
quantum error correctionsyndrome post-selectionsupervised classifiersurface codelogical-error rate
Online Optimization of Difference-of-Convex Compositions with Smooth Mappings
The paper proposes an online optimization method for non-convex non-smooth problems where losses are compositions of difference-of-convex functions with smooth mappings. The authors introduce a time-smoothed proximal linear algorithm and a local-regret measure based on a proximal residual mapping, proving it captures first-order stationarity. Key technical contributions include a tangent-cone characterization for feasible regions with composite difference-of-convex constraints, enabling convex optimization oracles despite non-convexity. Results include local-regret bounds, a bound on inner convex subproblems, and an error bound linking the proximal residual to stationarity distance.
online optimizationdifference-of-convexproximal residualnon-convex optimizationregret bound
Agent-Centric Animal Pose Forecasting
The paper introduces an agent-centric framework for autoregressive modeling of animal behavior from pose tracking data, applicable to both individual and group interactions. The approach uses egocentric sensory observations to generate egocentric movements, enforcing biological constraints where agents independently sense and respond to conspecifics. The authors release a general-purpose library for managing parallel data representations and demonstrate its effectiveness in capturing social behavior distributions in Drosophila courtship, with quantitative evaluation tools provided.
agent-centric modelingautoregressive modelsegocentric movementpose forecastingsocial behavior
RELTA-SGLD: Relative-Growth Localized Taming for Nonconvex Stochastic-Gradient Langevin Learning
The paper introduces RELTA-SGLD, a taming scheme for nonconvex stochastic-gradient Langevin dynamics (SGLD) that stabilizes superlinear updates while minimizing unnecessary suppression of learning drift. The method employs a threshold-activated taming mechanism and a relative-growth principle derived from Lyapunov stability, yielding a lighter λ-scale denominator and preserving far-tail returns. Theoretical results show polynomial moment stability and first-order stationary accuracy in both W₁ and W₂ Wasserstein metrics, improving upon prior half-order and quarter-order bounds. Empirical evaluation on Fashion-MNIST demonstrates superior mean learning metrics over untamed SGLD and TUSLA, with competitive performance against AdamW.
stochastic-gradient langevin dynamicslyapunov stabilitynonconvex optimizationwasserstein metricstaming scheme
Equilibrium Causal Games: Separation, Identification, and the Identifiability of Cyclic Latent States
The paper introduces Equilibrium Causal Games (ECGs), a framework combining games with cyclic causal models, hidden inputs, and sensor mappings to study feedback-driven equilibria. It establishes conditions under which ECG-separation is sound but incomplete, and analyzes identifiability of latent states under various constraints. Key results show that passive linear models with unknown wiring and sensing leave interaction matrix $B$ unidentified for $d\ge2$, while non-Gaussianity and mechanism interventions improve identifiability. The work delineates which causal inferences equilibrium data can support versus those requiring targeted experiments.
equilibrium causal gamescyclic causal modelsidentifiabilitynon-gaussianitymechanism interventions
The C-index illusion: discrimination without calibration in published survival models
The study demonstrates that relying solely on the concordance index (C-index) for evaluating survival analysis models leads to systematically misleading comparisons due to ignored calibration and time-dependent accuracy. Through reproducing three published survival-ML models across diverse domains (hard-drive failure, peer-to-peer credit default, user disengagement), the authors validate their evaluation instrument and test five pre-registered hypotheses. Three hypotheses reject, revealing significant calibration failures despite high discrimination (e.g., C = 0.9595 vs. 0.958, p = 2.6e-136), biased risk estimates, and degrading probability estimates with prediction horizon. The study provides a reusable evaluation harness with full code.
concordance indexsurvival analysiscalibrationdiscriminationcompeting risk
Geospatial Diffusion-based Evolution Synthesis (GeoDES) for Storm-Centered Weather Augmentation
The paper introduces Geospatial Diffusion-based Evolution Synthesis (GeoDES), an image-to-video diffusion model for synthesizing high-fidelity storm-centered weather events. The method addresses limitations of regional and global weather models by generating physically consistent storm structures, enabling dataset augmentation and forecast model stress-testing. Evaluations on North Atlantic test data show GeoDES achieves 52% lower Peak Vorticity Error and 8% higher Anomaly Correlation Coefficient compared to prior methods.
diffusion modelweather augmentationpeak vorticity erroranomaly correlation coefficientstorm dynamics
Boltzmann-Expected Molecular Design with Decoupled Annealing Flows
The paper introduces DECAF (Decoupled Annealing Flows), a method for Boltzmann-expected molecular design that optimizes ensemble statistics rather than single-conformer properties. DECAF factorizes the joint distribution over molecular graphs and coordinates into two conditional flow models: a graph-conditioned flow acting as a Boltzmann emulator and a coordinate-conditioned flow proposing new graphs. By alternating these flows with a simulated-annealing acceptance rule, DECAF optimizes molecular graphs based on ensemble statistics. On the GEOM-Drugs dataset, DECAF demonstrates consistent shifts toward target properties like radius of gyration and solvent-accessible surface area, outperforming single-conformer optimization. DECAF also enables higher-moment design, optimizing variance and skewness of ensemble properties, verified via all-atom MD simulations.
boltzmann-expected designdecoupled annealing flowsensemble statisticshigher-moment designsimulated-annealing
Do Sheaf Neural Networks Use Holonomy? A Measure--Intervene--Control Study
The paper introduces a basis-independent measurement framework for analyzing geometric mechanisms in sheaf neural networks (SNNs), focusing on triangle-loop products. Using Neural Sheaf Propagation (NSP) in a high-homophily GraphUniverse regime, the study quantifies SO(2) loop rotation, stalk-space area, and orientation. Results show NSP increases loop rotation from 0.010 to 0.388 radians for triangle counting, while community detection peaks at 0.029 radians. Interventions reveal post-training sensitivity to learned connections, though a graph-summary ridge predictor outperforms. The study disentangles geometric change, connection sensitivity, and triangle-specific computation.
sheaf neural networksneural sheaf propagationso(2) loop rotationgraphuniversetriangle-loop products
Total Variation Distance Estimation in Autoregressive Models
(No summary returned.)
Tensor Network Machine Learning for Wildfire Susceptibility Mapping: from Grokking Dynamics to Quantum Mixedness of Class Representations
The study introduces a quantum-inspired tensor network framework for wildfire susceptibility classification in Gargano, combining AlphaEarth embeddings with Matrix Product State models. The method employs scalable geospatial representations and an interpretable quantum mask for binary and multiclass classification. Results show a grokking transition in binary classification and hierarchical class distinguishability via level-resolved mixedness diagnostics, with non-adjacent categories more separable than neighboring ones, achieving competitive accuracy while providing interpretability.
tensor networkmatrix product stategrokking transitionmixedness diagnosticsalphaearth embeddings
A Bayesian Framework for Built-in Input Dimension Reduction for Gaussian Process Modeling
The authors propose a Bayesian framework integrating dimensionality reduction with Gaussian Process (GP) modeling, addressing high-dimensional input challenges via a hierarchical model with Stiefel manifold priors. The method enforces orthonormality on projection matrices and employs Hamiltonian Monte Carlo with geodesic flow for posterior inference, extended to Deep Gaussian Processes (DGP) for complex datasets. Numerical experiments show improved predictive performance and uncertainty quantification despite higher computational costs compared to two-stage approaches.
gaussian processstiefel manifolddimensionality reductionhamiltonian monte carlodeep gaussian processes
Generating Bearing Vibration Signals at User-Specified Fault Probabilities Using PR-GAN and Counterfactual Methods
The paper introduces two methods for generating bearing vibration signals at user-specified fault probabilities (0.25, 0.50, 0.75) to address the scarcity of intermediate-probability samples in datasets. A Probability-Regularized Generative Adversarial Network (PR-GAN) extends WGAN-GP by editing real signals via a residual generator, while a Wachter-style counterfactual (CF) procedure directly optimizes input signals to match target probabilities. Evaluated on the CWRU and Paderborn datasets, CF achieves a mean absolute probability error of 0.005-0.008 and a 100% success rate, outperforming PR-GAN (error: 0.046-0.059, success rate: 0.501-0.680). CF requires smaller L1 changes but PR-GAN is faster in most settings.
probability-regularized gancounterfactual optimizationbearing vibration signalswasserstein ganfault probability
H$^2$SD: Hybrid Hindsight Self-Distillation
The paper introduces Hybrid Hindsight Self-Distillation (H$^2$SD), a reinforcement learning method that adapts teacher context and update strategy based on trajectory correctness in language model reasoning. For successful trajectories, H$^2$SD uses verified responses and rephrasing instructions to refine token-level credit assignment without altering reward direction. For failed trajectories, it employs reverse-KL distillation with verifier-confirmed reference hints. Experiments on reasoning benchmarks demonstrate H$^2$SD's superior performance over RLVR and self-distillation baselines, with stable optimization and improved accuracy-efficiency trade-offs.
reinforcement learningself-distillationreverse-kltoken-level guidancereasoning benchmarks
Marine Engine Fault Dataset: Open-Access Data under Controlled Reference and Fault Scenario Conditions
The Marine Engine Fault Dataset provides an open-access benchmark for marine-engine predictive maintenance, featuring controlled fault experiments with documented operating conditions. Data was collected from a turbocharged, intercooled three-cylinder marine diesel engine under reference (30-90% load range) and five fault-scenario conditions: cooling-water pump cavitation, compressor air-filter clogging, air-cooler fouling, injection-valve nozzle clogging, and turbine degradation. Multi-sensor time-series measurements demonstrate physically coherent reference performance and interpretable fault-response patterns, enabling structured reuse for anomaly detection, fault diagnosis, and degradation modelling in maritime machinery.
predictive maintenancemarine diesel enginefault diagnosisanomaly detectiontime-series data
Enhanced Neural Quantum State via Annealed Gradient Descent
The paper introduces annealed gradient descent (AGD) to mitigate subspace trapping, a finite-sample instability in neural quantum state optimization where important configurations are undersampled. AGD temporarily reweights gradients to preserve low-probability configurations while limiting dominance of high-probability ones. Evaluated on molecular systems and $J_1$-$J_2$ models, AGD suppresses metastable trapping, achieves chemical accuracy, and matches state-of-the-art performance with compact architectures.
neural quantum statessubspace trappingannealed gradient descentquantum many-bodystochastic optimization
📰 Industry Media (1)
How AI helps scientists design the next generation of medicines
AI is transforming biologic drug discovery by accelerating candidate screening and enabling de novo protein design through computational methods. AstraZeneca employs a build-measure-learn loop, where AI prioritizes molecules for lab testing, reducing cycle times by up to 50% (McKinsey estimate). The approach integrates multimodal data (molecular structures, binding measurements) and robotic automation to optimize multi-specific biologics. Key challenges include safety prediction via virtual clinical trials and agentic AI systems for closed-loop optimization. Human oversight ensures explainability, with engineers focusing on uncertainty quantification and interpretability.
biologic drug discoveryde novo designmulti-specific biologicsclosed-loop optimizationuncertainty quantification
Generated automatically at 2026-07-23 20:34 UTC. Summaries and keywords are produced by an LLM and may contain inaccuracies — always consult the original article.
