Daily Digest — 2026-08-02
320 items · 6 research labs, 314 arxiv papers
MarkTechPost: all feed URLs failed (last tried: https://www.marktechpost.com/feed/)AI News: all feed URLs failed (last tried: https://artificialintelligence-news.com/feed/)
🏛️ Research Labs (6)
Ten advances in mathematics and theoretical computer science
OpenAI's unreleased Astra model generated novel solutions to ten longstanding open problems across mathematics and theoretical computer science, including high-dimensional sphere packing, non-sofic groups, and quantum parallel repetition. The solutions were discovered using approximately $2,000 worth of API tokens, then formalized in Lean by humans with model assistance. Key results include disproofs of Connes’s rigidity conjecture and Erdős's unit-distance conjecture, superexponential bounds for multicolor Ramsey numbers, and polynomial-factor hardness for the closest vector problem in lattice cryptography.
non-sofic groupsquantum parallel repetitionarithmetic circuit complexitylattice cryptographyextremal combinatorics
Advancing responsible AI across Europe
OpenAI outlines its compliance strategy with the EU AI Act, emphasizing responsible AI governance through multi-stakeholder frameworks. The approach integrates internal safety protocols (Preparedness Framework, Frontier Governance Framework) with external collaborations (Frontier Model Forum, EU Cyber Action Plan) to address transparency (C2PA, SynthID), cybersecurity (Trusted Access for Cyber program), and risk management. Key initiatives include provenance tools for AI-generated content (extended to audio) and support for EU cybersecurity resilience. The methodology aligns with the EU AI Act's GPAI Code of Practice and emphasizes pragmatic, risk-based regulation.
eu ai actprovenancecybersecuritygpai codesynthid
Building abundant intelligence
OpenAI outlines a full-stack approach to democratizing advanced AI through cost reduction and capability improvements. The company reduced GPT-5.6 Luna's pricing by 80% and GPT-5.6 Terra's by 20%, while introducing a Fast mode for GPT-5.6 Sol with 2.5× speed at 2× cost. System optimizations, including speculative decoding (15% token-generation efficiency gain) and context management, improved ARC-AGI-3 benchmark performance from 13.3% to 38.3% with 6× fewer output tokens. Agentic workflows now constitute 99.8% of weekly API output, driven by adoption in enterprise settings. The strategy emphasizes balancing intelligence, speed, and cost to maximize productive outcomes.
speculative decodingagentic workflowscontext managementtoken-generation efficiencyarc-agi-3
Univé builds an AI-ready workforce
Univé implemented an organization-wide AI transformation using ChatGPT Enterprise, focusing on leadership alignment, governance-integrated deployment, and employee-led innovation. The approach combined secure platform adoption (97% license activation, 85% weekly active users) with 1,500 custom GPTs for workflows like pet insurance claims processing (reduced from hours to minutes). Results demonstrate sustained engagement (40 prompts/user/week) and cross-functional adoption, enabled by Workspace Agents that proactively structure tasks while maintaining human accountability.
chatgpt enterpriseworkspace agentscustom gptsorganizational transformationgovernance framework
Disrupting a Criminal Scam Operation
OpenAI disrupted a Cambodia-based criminal network leveraging ChatGPT for multi-modal scam operations, including investment fraud, romance scams, and law enforcement impersonation. The operation employed LLM-generated content for persona fabrication, multilingual message translation, and forged document synthesis (e.g., passports, financial records). Analysis revealed a three-phase modus operandi: initial outreach ('ping'), emotional manipulation ('zing'), and financial extraction ('sting'). The network exhibited indicators of forced labor, with internal records documenting debt bondage and recruitment coercion. OpenAI terminated associated accounts, shared threat intelligence with industry partners, and observed cross-over between cybercrime and human trafficking networks in Southeast Asia.
llm-generated contentmulti-modal scamsdebt bondagethreat intelligencepersona fabrication
How avatarin built a 24/7 retail agent with GPT-Realtime
Avatarin developed a 24/7 multilingual retail agent using OpenAI’s GPT-Realtime, integrating voice, text, and visual understanding into a single conversational interface. The agent, deployed on Yamada Denki’s online store, engaged 30,000 users in a two-week campaign, achieving 92% positive survey feedback. Key innovations included retrieval-augmented generation for accurate product information, proactive questioning for guided discovery, and adaptive conversation flows based on Yamada Denki’s expertise. The system demonstrated low-latency performance across modalities, enabling natural interactions and revealing customer insights previously obscured in conventional online shopping.
gpt-realtimeretrieval-augmented generationmultimodallow-latencyconversational interface
📜 arXiv Papers (314)
Learning to Trace Seiberg Dualities
This paper introduces a machine learning approach to efficiently identify Seiberg dualities in supersymmetric quiver gauge theories, addressing the computational challenge of establishing duality between systems. The method leverages transformer and multi-layer perceptron architectures, supplemented by pathfinder algorithms, to trace mutations of quivers with up to 10 nodes. Results demonstrate that these neural network architectures outperform deterministic algorithms in both efficiency and accuracy. The study highlights the potential of this class of problems as a benchmark for AI applications in theoretical physics.
seiberg dualitiesquiver gauge theoriestransformersmulti-layer perceptronspathfinder algorithms
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
ReToken introduces a learnable embedding to address visual retrieval challenges in vision-language models, where performance degrades with increasing visual context. The method trains a single retrieval-target token to sparsely select query-relevant visual tokens from a pre-filled KV cache, requiring only a small image-QA dataset. Evaluations show consistent improvements: on Visual Haystacks, Qwen3VL-8B and InternVL3.5 gain 13.4 and 12.4 points (>20% relative), respectively, while LVBench demonstrates an 8.0-point zero-shot gain for Qwen3VL-8B in long-video tasks. The lightweight design enables single-H100 training and inference.
learnable embeddingvisual retrievalkv cachezero-shot transfervision-language models
PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
PAC-MAN introduces a perception-aware CBF-RL framework for whole-body safety in humanoid dodgeball, combining control barrier functions (CBF) with onboard sensing. The policy uses segmentation-masked depth from a head-mounted camera during deployment, while training incorporates CBF guidance for body-link clearance and an adversarial motion prior for evasive reflexes. Evaluated on an any-link contact benchmark with seeded throws, the policy achieves near-oracle performance (95% success in real-world deployment on Unitree G1) using a fixed camera, demonstrating robustness to imperfect perception. Joint-CBF performance degrades under fixed-camera observations but recovers with ball-tracking gimbals or runtime filters.
control barrier functionshumanoid roboticsadversarial motion priorsemantic segmentationonboard sensing
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
AskChem introduces a claim-centered infrastructure for chemistry literature synthesis, shifting retrieval units from papers to provenance-carrying claims. Each paper is decomposed into atomic, typed claims, grounded by source DOIs and verbatim quotes or explicit evidence locators. The system employs a faceted taxonomy, evidence graph, and living taxonomy for hierarchical retrieval, claim linking, and paper categorization. Indexing 2.4M claims from 147K papers, AskChem offers web, REST, SDK, and MCP access. Evaluated on AskChem-Bench, grounding GPT-5.5 in AskChem achieves 100% resolvable DOIs and the highest citation density among tested systems.
claim-centeredprovenance-carryingfaceted taxonomyevidence graphliving taxonomy
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
The paper introduces Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for auditing system prompts in commercial AI products. AISPA evaluates prompts along eight user-relevant dimensions, classifying instructions as protective or problematic. The authors audit 3,249 instructions from 88 products, revealing significant variation in prompt design: 98.9% contain protective instructions, but only 24% cover all dimensions, and 40% include problematic ones. Findings indicate trends toward longer, more protective prompts but highlight persistent issues with transparency and conflicting instructions.
system promptsprompt auditinguser protectioncommercial aitransparency
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
The authors introduce OSReward, a benchmark for evaluating vision-language models (VLMs) as judges of computer-using agent (CUA) trajectories, addressing reliability concerns in task verification. OSReward includes diverse trajectories from verified instructions, rigorously annotated for ground-truth verdicts, with derived subsets OSReward-Hard and OSReward-Multi for challenging cases and fine-grained scoring. Evaluation reveals systematic leniency biases in state-of-the-art VLMs, with reliable models being cost-prohibitive. To bridge this gap, they release OS-Shepherd-100K, a corpus of reasoning-annotated judgments, and train OS-Shepherd models (9B and 35B) that match commercial judges at 30-60% lower cost.
vision-language modelscomputer-using agentstask verificationreward modelsbenchmark evaluation
PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks
PAIChecker introduces a multi-agent system to detect PR-Issue misalignment in SWE-bench-like benchmarks, addressing a 13.6% misalignment rate across five patterns in eleven scenarios. The method employs a three-phase design combining pattern identification, cross-agent label synthesis, and code-level validation for accurate detection. Evaluations on SWE-Gym and SWE-bench Multilingual show PAIChecker achieves 92.12% and 91.67% binary accuracy across four LLM backbones.
pr-issue misalignmentswe-benchmulti-agent systemcode-level validationllm backbones
DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation
DualG-MRAG proposes a decoupled architecture for Multimodal Retrieval-Augmented Generation (MM-RAG), addressing limitations in multi-hop reasoning by separating Macro-reasoning (global topological routing via Macro Graph) and Micro-matching (local verification via Micro Graph). The method formulates retrieval as query-driven message passing through a GNN Retriever and introduces dynamic programming decoding to extract explicit reasoning paths from GNN forward passes. Experiments show improvements in evidence recall and complex QA accuracy over baselines.
multimodal retrieval-augmented generationmacro-reasoning graphmicro-matching graphgnn retrieverdynamic programming decoding
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B
The study systematically evaluates seven self-improvement methods for language models (1.5B to 7B parameters) against repeated sampling at equal token cost on two math benchmarks (150 questions each). By precisely tracking all generated tokens—including those spent on critiques, reflections, and debate turns—the authors find no self-inspection method (e.g., Self-Refine, Reflexion) outperforms repeated sampling; 10/36 comparisons are significantly worse, with all 18 self-inspection cases underperforming. Key findings include: (1) model-picked answers lose 8.0-11.3 accuracy points vs. majority voting at 1.5B (gap shrinks to 1.3-2.0 at 7B), and (2) rewriting methods remain 3.6-10.1 points below baseline at 7B. All results include bootstrap confidence intervals and multiplicity correction.
self-inspectionrepeated samplingbootstrap intervalsmajority votingtoken cost
Algorithms for Structured Elections under Thiele Voting Rules
The paper analyzes computational complexity in approval-based committee elections under Thiele voting rules, a class parameterized by weight vectors determining voter satisfaction. By examining structural dependencies between candidates induced by voter approval sets, the authors develop FPT algorithms for Proportional Approval Voting (PAV) and other Thiele rules on the Voter Interval (VI) domain, where candidates are approved by consecutive voter intervals. Key results include FPT tractability for VI instances under a parameter where general cases remain NP-hard, polynomial-time solutions for ≤2-approval cases, and an FPT algorithm parameterized by winning committee score.
thiele voting rulesproportional approval votingfpt algorithmsvoter interval domaincommittee elections
Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs
This study systematically evaluates inference-time scaling for local computer-use agents (CUAs) across contextual, temporal, structural, and parallel dimensions. The authors test Qwen3-VL-8B/30B-A3B, UI-TARS-1.5-7B, and OpenCUA-7B on OSWorld, revealing diminishing returns and shifting failure modes under compute constraints. Contextual scaling improves trajectory stability but saturates with token cost, while temporal scaling reduces stalls without boosting task success. Structural decomposition introduces planning overhead, and parallel scaling mitigates failures at high computational cost. Findings highlight the need for selective compute allocation and failure-aware frameworks in local CUA design.
inference-time scalingcomputer-use agentscontextual scalingtemporal scalingparallel scaling
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
The paper introduces Atomic Policy Optimization (APO), an unsupervised framework for 3D atomic structure prediction that eliminates reliance on ground-truth reference structures. APO employs group-relative policy optimization with a dual-reward mechanism: (i) a spectral reward reinforcing dominant latent structural modes via eigen-decomposition of sample similarities, and (ii) a thermodynamic reward enforcing stability. Benchmarks on crystal and antibody structure prediction show APO outperforms supervised baselines, achieving state-of-the-art match rates and structural fidelity while improving inference efficiency through probability path straightening.
atomic policy optimization3d structure predictionunsupervised alignmentgroup-relative policy optimizationthermodynamic reward
ORCA-bench: How Ready Are Language Model Agents for Oncall?
ORCA-bench evaluates language model agents' capability for oncall root cause analysis (RCA) in production-like settings, combining OpenTelemetry-instrumented microservices (Prometheus, Jaeger, OpenSearch) with 1,079 RCA tasks varying report specificity and fault scenarios. The benchmark uses expert-curated ground truth and achieves high human-LLM judge agreement (Cohen's $κ_w=0.90$). Results show frontier models achieve 25.3% accuracy on Medium-difficulty tasks and 10.0% on Hard tasks, with hallucinations in 40% of cases for weaker models. Performance degrades without source-code access, and the benchmark's 50 GB/six-day scope represents a lower bound for real-world deployment challenges.
root cause analysisopentelemetryprometheusllm-as-judgemicroservices
MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems
We propose MANTA, a framework for Multi-Agent Network Topology Adaptation that enables self-evolving communication structures in multi-agent systems during inference. MANTA initializes task-conditioned topologies from prior structural experience and dynamically updates agent roles, communication links, execution order, information visibility, and validation pathways based on collaboration traces. Evaluated across five benchmarks for information seeking, tool use, planning, workflow execution, and mathematical reasoning, MANTA achieves an average score of 74.0, outperforming the strongest baseline by 5.8 percentage points and excelling on PlanCraft. This demonstrates the efficacy of inference-time architectural self-improvement in multi-agent collaboration.
multi-agent systemsnetwork topologyinference-time adaptationtask-conditioned initializationcollaboration traces
What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration
DAR-Net introduces a Dual-Ambiguity Rectification Network to address dual ambiguity in all-in-one image restoration: semantic ambiguity in channel-wise modulation and spatial ambiguity in restoration responses. The method employs a Degradation Archetype Representation (DAR) module to model degradation states, a Semantic Ambiguity Rectification (SeAR) module for degradation-aware prompts, and a Spatial Ambiguity Rectification (SpAR) module to orthogonalize feature responses. DAR-Net achieves state-of-the-art performance on standard benchmarks, improving average PSNR by 0.14 dB and 0.34 dB in three-degradation and five-degradation settings, respectively, and excels on CDD-11 and WeatherBench.
dual-ambiguity rectificationdegradation archetype representationsemantic ambiguity rectificationspatial ambiguity rectificationall-in-one image restoration
Selective Credibility-Limited Belief Update
The paper introduces selective credibility-limited belief update, a framework that extends standard belief update by allowing partial acceptance of compound epistemic inputs through world-specific transformations. The method transforms epistemic inputs into weaker proxies before performing credibility-limited transitions, with semantic and axiomatic characterizations provided. Two sub-classes are identified: consistency-preserving operators (ensuring credible transformed inputs when originals are consistent) and maximal consistency-preserving operators (additionally requiring maximal informativeness). The framework generalizes both Katsuno-Mendelzon update (no credibility limits) and credibility-limited update (indivisible inputs), demonstrating strictly greater expressiveness.
belief updatecredibility-limitedepistemic inputconsistency-preservingmaximal informativeness
Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation
The paper introduces confidence-scheduled restricted responses (CS-RNR), a novel opponent-exploitation method for imperfect-information games that certifies its own safety guarantees through runtime auditing. The approach uses anytime-valid confidence sequences to detect exploitable deviations from equilibrium, constructs conservative opponent models, and evaluates candidate counter-strategies via full-tree best response before deployment. In Leduc hold'em, CS-RNR achieves 6.2× the steady-state gain of money-verified binary gates while maintaining safety, with all 36,000 audited hands across three games satisfying the certificate tolerance.
nash-equilibriumconfidence sequencesrestricted responsesopponent exploitationsafety certificate
InfoOps Bench: A live information operations safety benchmark
The paper introduces InfoOps Bench, a dynamic benchmark evaluating frontier language models' susceptibility to state-backed information operations, using 2,100 live-tracked operations from Russian, Chinese, and Iranian sources. It tests 17 models from 8 providers across four prompt framings, measuring integrity scores (refusal rates) and fact-checking behavior. Results show refusal rates vary widely (8.8%-94.5%), uncorrelated with model size, with Chinese models exhibiting sharp compliance drops (48-70 percentage points) on China-critical claims. Output harmfulness and fact-checking rates (2.9%-72.9%) also vary significantly, highlighting trade-offs between safety and usability.
information operationsintegrity scoresprompt framingsfact-checking ratesmodel compliance
TCA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval
TCA-SIR introduces target-conditioned abstraction (TCA) for Scientific Inspiration Retrieval (SIR), explicitly representing how candidate inspirations transfer to a target problem. The method learns to generate target-conditioned abstractions and uses their representations to predict transferability, addressing limitations of topical similarity-based approaches. On ResearchBench, TCA-SIR outperforms prior methods like MOOSE-Chem by >10 percentage points in HitRate@top4% and provides interpretable rationales for scientific inspiration through learned abstractions.
scientific inspiration retrievaltarget-conditioned abstractiontransfer learninghypothesis generationinterpretable retrieval
SCOPE: Supply-Chain Operations through Coupled Policies for End-to-End Coordination
SCOPE introduces a composite policy model for end-to-end supply-chain coordination by representing entities as tokens and contextualizing decisions through a shared operational representation. The framework couples assortment selection, source assignment, replenishment frequency, and routing decisions, evaluating them via system-level utility. Evaluated on real data from Dingdong and JD.com, SCOPE outperforms decoupled optimization methods and practice-oriented baselines, demonstrating improved coordination across supply-chain echelons.
supply-chain coordinationcomposite policy modeloperational representationreplenishment planningsystem-level utility
A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer Networks
The study proposes a fuzzy rule-based neuro-symbolic framework for interpretable sewer pipe severity prediction, decoupling neural perception (Swin Transformer for 14 multilabel CODE degrees) from symbolic reasoning (19 IF--THEN rules derived from Weka's J48). Fuzzy logic combines t-norm activations and s-norms weighted by rule confidence to generate class evidence. Evaluated on 3,244 images with five imbalanced severity classes (labels derived from LLM consensus), the method improves accuracy, balanced accuracy, Macro F1, and MCC by 17.9%, 12.2%, 23.0%, and 17.3%, respectively, over image-only classification, while maintaining interpretability.
neuro-symbolicfuzzy logicswin transformermultilabel classificationseverity prediction
Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation
The paper proposes an autonomous aircraft surveillance system combining on-board inference with generative data augmentation to address downlink limitations and class imbalance in nanosatellite imagery. A 6U CubeSat employs a low-power edge tensor accelerator for real-time inference, while a LoRA-fine-tuned diffusion model generates synthetic minority-class aircraft images, pseudo-labeled by an intermediate detector. Results show a global mAP increase from 77.9% to 82.2%, with minority-class F1-score improving from 0.683 to 0.811, and the quantized detector achieving 25-30 FPS on orbit within memory constraints.
nanosatelliteslow-rank adaptationedge tensor acceleratorgenerative data augmentationpseudo-labeling
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
EndoCLIP introduces a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs extracted from 280,476 routine colonoscopy reports, addressing the weak linkage between clinical findings and corresponding images. The model employs progressive recovery of finding-to-frame correspondence to enable scalable supervision. EndoCLIP outperforms general-purpose and biomedical vision-language encoders in zero-shot and linear-probe settings across lesion-level image-text retrieval, structured report generation, and six multi-centre clinical classification tasks. Notably, its linear probe achieves performance comparable to expert readers in benign-versus-malignant classification, demonstrating the efficacy of leveraging routine documentation for clinical AI.
vision-language modelcolonoscopylesion-levelzero-shotlinear-probe
LeanCSP: A Framework for Certifying Constraint Reformulation and Solving in Lean
The paper presents LeanCSP, a framework for certifying constraint reformulation and solving within the Lean theorem prover. It enables two-level verification: proving semantics-preserving properties (equivalence, equisatisfiability, symmetry-breaking correctness) for entire problem families parametrically, and validating solver certificates for individual instances via MiniZinc/SMT-LIB/OPB backends. Combining both levels provides end-to-end correctness guarantees without trusting external solvers. Experiments demonstrate that verified symmetry breaking reduces solver search effort by up to 2×10^7×, while Lean-based certification remains practical (minutes for largest instances).
constraint programmingtheorem provingsemantics-preservingsymmetry breakingcertification
SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
The paper introduces Self-Verifying Refinement (SVR), a reinforcement learning framework for adaptive test-time compute in language models, eliminating reliance on external verifiers. SVR jointly learns to generate solutions, discrete correctness verdicts, and confidence scores, using self-verification to control refinement turns—retaining answers only when confident and correct. Trained with GRPO on fixed-horizon trajectories, it optimizes correctness, calibration, and early stopping. Evaluated on seven mathematical reasoning benchmarks with Qwen3.5-2B, SVR achieves 0.563 macro-average accuracy with only 2.99 average turns, outperforming baselines and fixed-budget approaches.
self-verificationtest-time computereinforcement learningadaptive stoppingmathematical reasoning
Machines that know they are aging: a framework for hardware-aware autonomous intelligence
The paper proposes Aging-Aware Autonomous Intelligence (AAAI), a framework enabling autonomous systems to adapt to hardware degradation through integrated health monitoring and adaptive reasoning. AAAI combines hardware self-awareness via physics-of-failure models, self-adaptive reasoning that adjusts inference complexity and planning horizons, and survival-centric intelligence for resource allocation. This closed-loop cognitive architecture addresses agnostic collapse by unifying prognostics and hardware-aware computing, improving resilience and operational lifetime in inaccessible or safety-critical environments like space missions and medical devices.
hardware degradationphysics-of-failure modelsadaptive reasoningagnostic collapseclosed-loop cognitive architecture
Metaphor Tracer: A Theory-Informed Analysis of Hidden States
Metaphor Tracer introduces a theory-informed method to analyze language model hidden states, scoring token positions on two properties: the aggregator, measuring text consolidation, and the differentiator, assessing transient token transport. The method, validated without training on three unrelated models, demonstrates that aggregator scores remain stable despite decreasing surprisal and attention, indicating a token's structural role in the text. Ground truth validation includes engineered registers (6/6 cells) and psychoanalyst-marked clinical transcripts (34/36 cells), showing graded increments above lexical controls. Transfer tests reveal that models with token structures tied to lexical types perform worse in singular discourse reading, while tuning enhances fidelity without altering type-transfer.
hidden statesaggregatordifferentiatortoken positiontransfer test
A foundation model of numerical intelligence with cross-disciplinary generalization
The paper introduces UNified In-Context Operator Networks (UNICON), a foundation model for numerical intelligence that generalizes across disciplines. UNICON infers predictive relations from graph-based numerical context and applies them to queries within the same system, achieving specialist-level performance without retraining, including in unseen disciplines. Combining UNICON with language-model agents further improves performance, surpassing state-of-the-art specialists in untrained domains. Results demonstrate that training-corpus diversity enhances generalization, positioning UNICON as a versatile component for broader AI ecosystems.
foundation modelnumerical intelligencein-context learningcross-disciplinary generalizationgraph-based prediction
When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence
The study introduces derived-feature over-trust (DFOT), a failure mode where LLMs treat derived measurements (e.g., PPG-derived rhythms) as direct facts despite instance-dependent validity. Using physiological sensing (PPG-ECG pairs), the authors propose five estimands—COTR, CIR, CRR, ESRM, UHR—to quantify DFOT and evaluate mitigation via privileged distillation (ECG-to-PPG supervision). On a 187-patient test set, their baseline improves repair/specificity endpoints by 1.82-6.69 percentage points (all CIs exclude zero), with minimal utility harm (+0.67pp, CI: -0.4 to +1.7). The framework enables standardized DFOT evaluation.
derived-feature over-trustprivileged distillationphysiological sensingestimandsepistemic status
WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
WIDE introduces a token-level dynamic width pruning framework for efficient LLM inference, addressing limitations of static and coarse-grained dynamic pruning methods. The approach enables fine-grained computation allocation by dynamically selecting attention-head groups and FFN-channel groups per token, extending pruning to neuron-block-level granularity. A two-stage training pipeline learns token-wise sparse execution patterns, while a pruning-kernel co-design framework decomposes dynamic sparsity acceleration into mask reordering and block-level skipping. At 50% sparsity, WIDE achieves 55.1% performance improvement over state-of-the-art dynamic depth pruning, with kernel-level speedups of up to 1.98x for prefill and 4.95x for decoding.
dynamic pruningtoken-levelattention-head groupsffn-channel groupssparsity acceleration
QQWorld: Quantile-Quantile Matching for World Model Regularization
QQWorld introduces quantile-quantile matching for latent world model regularization, addressing the vanishing gradient issue of the Epps-Pulley objective in LeWorldModel (LeWM). The method aligns projected latent samples with rank-matched Gaussian quantiles, ensuring effective corrective gradients even in tail regions. Cross-batch QQ further enhances ranking pool size using detached samples, optimizing bias-variance trade-off. Evaluated across four control environments, QQWorld improves LeWM's average planning success rate, achieves better Gaussian alignment, and reduces latent tail thickness.
quantile-quantile matchinglatent world modelepps-pulley objectivecross-batch qqgaussian alignment
On-Policy and Off-Policy Learning for Large Action Spaces
The thesis advances policy learning methods for contextual bandits with large action spaces, addressing both on-policy and off-policy settings. For on-policy learning, it proposes meTS (mixed-effect Thompson sampling) and dTS (diffusion-inspired Thompson sampling) to share information across actions, with regret bounds scaling with an effective action count. For off-policy learning, it introduces sDM (structured direct method), concave policy-weighted objectives, and differentiable pessimistic methods using exponential smoothing and PAC-Bayesian bounds to control bias-variance trade-offs in importance sampling. Theoretical guarantees are provided for both exploration efficiency and optimization stability.
contextual banditsthompson samplingimportance samplingpac-bayesiandiffusion prior
QuantWAMs: Calibrating at the Right Granularity for World Action Models
QuantWAMs introduces a post-training quantization framework tailored for World Action Models (WAMs), addressing inefficiencies in deployment caused by iterative denoising and closed-loop execution. The framework employs three strategies: shared-basis outlier calibration, co-training-objective saliency, and fixed-intervention rollout auditing, aligning quantization decisions with model structure, rollout distribution, and task objectives. Evaluated on Fast-WAM and LingBot-VA across RoboTwin 2.0, LIBERO, and AgiBot G2 real-robot manipulation tasks, QuantWAMs achieves simulation means within 0.2–0.7 percentage points of FP16. It reduces peak weight-and-activation memory to 29% of FP16 and provides 1.4–1.6× block-level speedups, demonstrating deployment feasibility.
post-training quantizationworld action modelsclosed-loop executiondenoising-step protectionempirical-fisher scores
GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation
We introduce GLM-RAG, a graph language model (GLM)-based retriever for retrieval-augmented generation (RAG) over knowledge graphs, comparing it to GNN-based and vector-search retrievers. GLM-RAG integrates graph reasoning and semantic capabilities, evaluated across single- and multi-hop RAG tasks with emphasis on domain transferability. Results show GLM retrievers achieve state-of-the-art (SOTA) on two multi-hop benchmarks in out-of-domain settings, while maintaining comparable in-domain performance with promising scaling trends. GNN-based retrievers offer higher graph coverage efficiently, and vector-search excels in single-hop tasks.
graph language modelretrieval-augmented generationgraph neural networkmulti-hop reasoningdomain transferability
When Specifications Conflict: A Symmetry-Based Framework for Measuring LLM Preferences
The authors propose a symmetry-based experimental framework for analyzing how large language models (LLMs) resolve conflicts between competing specifications. The method constructs explicit conflicts across representation types—pure natural language, formal language, naturalized formal language, and input-output examples—enabling direct observation of model preferences. Evaluated on a mathematical benchmark with 550 conflict instances spanning 11 function families, results reveal systematic preference patterns: formal ≈ naturalized formal > pure natural language > input-output examples. The framework is extended to Boolean algebra, code generation, and clinical domains, demonstrating its applicability across diverse tasks and specification forms.
conflicting specificationssymmetry-based frameworkrepresentation typessystematic preferencesfunction families
HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
HyperClaim introduces a temporal hypergraph framework for video misinformation detection, addressing limitations in global multimodal fusion by modeling higher-order cross-modal interactions. The method constructs sparse heterogeneous hypergraphs over query tokens, evidence tokens, and frames, employing confidence-aware filtering, source budgeting, and adaptive soft-incidence reasoning with residual calibration. Evaluated under the FactGuard protocol, it achieves 83.7%, 82.0%, and 87.3% accuracy on FakeSV, FakeTT, and FakeVV datasets, outperforming discriminative and reasoning-centric baselines while preserving fine-grained structure.
hypergraph reasoningcross-modal dependenciestemporal hypergraphmisinformation detectionadaptive soft-incidence
How Benchmarks Mis-Score Computer-Use Agents
The study identifies systematic reliability issues in benchmarking computer-use agents (CUA), proposing a framework to analyze failures across task construction, trajectory observation, scoring, and reporting. Through auditing 150 failure-scored trajectories from five benchmarks (web, enterprise-workflow, desktop-control), it finds 15.3% incorrect FAIL verdicts (10.7% evaluator false negatives, 4.7% broken tasks). Genuine failures are categorized into verification/feedback, planning, and execution/grounding errors, revealing limitations of scalar success rates. The findings inform design rules for long-horizon CUA benchmarks.
computer-use agentsbenchmark reliabilityfailure taxonomyevaluator false negativestrajectory observation
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
ShadowDancer introduces a novel approach for any-action, frame-level control in interactive video world models by learning unified dynamics representations from video-shadow pairs. The method employs two key innovations: (1) shadow pairs, constructed via the Shadow Library, which replay identical dynamics under independently resampled appearances, enabling precise control across diverse dynamics families; and (2) cross-shadow prediction, which learns actions by predicting one shadow from another, discarding appearance-specific information and preserving dynamics. This approach allows demonstrated clips to be reused as action assets in new environments without additional labels or fine-tuning. Experiments show improved action transfer and long action rollout, achieving an average blinded win rate of 86% across diverse dynamics families.
shadow pairscross-shadow predictionvideo world modelsdynamics representationaction transfer
Teffic-Audio: Tell Fact from Fiction
Teffic-Audio introduces a general speech deepfake detection system addressing heterogeneous spoofing mechanisms (synthesis, conversion, vocoder reconstruction, neural-codec resynthesis) and environmental variability. The architecture combines a Conformer-based encoder, multi-head attentive statistics pooling, and binary classifier, prioritizing training enhancements (multi-source data, balanced sampling, diverse augmentation) over architectural complexity. Evaluated on Speech-DF-Arena's 14 test sets, it achieves a pooled EER of 1.454%, surpassing all public systems, with lowest EER on 5 subsets and favorable compute trade-offs. The system serves as a robust reference for generalized detection.
speech deepfakeconformerequal error rateneural-codecstatistics pooling
Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners
The paper introduces Perception-Correction Distillation (PCD), a label-free method for multimodal reasoners that addresses credit assignment in perception distillation by identifying correctable perception failures. PCD uses downstream failure and teacher-student disagreement as complementary witnesses, combining them via a soft AND gate (product operation) motivated by Bayesian evidence combination. The approach employs separated perception-reasoning rollouts and mean-preserving weights without altering the reasoning objective. Evaluations across eight benchmarks show PCD improves macro averages from 44.50 to 47.28 (8B→2B) and 56.94 to 61.22 (32B→8B), with ablations confirming its necessity (2.22-point drop without PCD).
perception distillationcredit assignmentmultimodal reasoningbayesian evidenceon-policy distillation
Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents
We introduce CARP, a reputation-penalty mechanism designed for LLM marketplace agents that suppresses dishonest product listings without requiring ground-truth verification. CARP incorporates a deadband to forgive complaint noise and state-dependent severity to counteract reputation-driven detection erosion, ensuring robustness against strategic gaming. Paired with SPARC, a byte-clean code-gated reflection mechanism, CARP closes most of the consumer-welfare gap relative to a perfect-information oracle. Empirical results demonstrate that LLM agents self-correct when fabrication incurs sales penalties, with behavioral binding observed across models. CARP achieves superior welfare compared to alternative policies.
reputation-penalty mechanismdeadbandstate-dependent severitybyte-clean code-gated reflectionconsumer-welfare gap
PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?
The authors introduce PathView-Bench (PathVU), a vision-anchored benchmark for evaluating multimodal large language models (MLLMs) on fine-grained, multiscale understanding of pathology images. PathVU comprises 14 VQA-style tasks across 61,673 images and 308,070 samples, with annotations spanning 28 organs and 7,253,526 labels, assessing region localization, visual recognition, quantity estimation, spatial reasoning, and insufficient-context judgment at both Region FOV (high-resolution) and Slide FOV (whole-slide) levels. Evaluation of 18 MLLMs reveals significant limitations in fine-grained visual understanding, despite their diagnostic capabilities. PathVU provides a reproducible framework for developing pathology-specific MLLMs with explicit multiscale visual reasoning.
multimodal large language modelscomputational pathologymultiscale visual understandingvision-language benchmarkspathology image analysis
One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence
The paper introduces a framework for auditing $N$ LLM agents under a constrained budget $B \ll N$, addressing challenges of miscalibrated self-reported confidence and correlated errors. Using a two-level Gaussian copula model, it identifies a miscalibration threshold $\delta^*$ beyond which confidence-ranked auditing performs worse than random selection. Counterintuitively, $\delta^*$ increases as the budget decreases, and cross-family correlation is significant due to shared task difficulty. Empirical evaluation across five open-weight LLMs reveals operationally useless confidence levels, while a proprietary model remains informative. The study provides a quantitative criterion for vacuous oversight and validates findings via policy replay on recorded traces.
gaussian copulamiscalibration thresholdconfidence-ranked auditingcross-family correlationvacuous oversight
ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
ObjectStream introduces a training-free framework for streaming video understanding by organizing visual evidence around persistent latent objects as memory anchors. The method induces spatially coherent latent objects from frozen Video-LLM representations, links them across frames, and maintains their histories under a bounded memory budget without external detectors. It preserves persistent object histories, transient changes, and recent visual context, enabling Video-LLMs to reason over object identities and interactions without model modifications. Evaluations show ObjectStream improves Qwen2.5-VL-7B by 10.0 points on OVO-Bench Real-Time Visual Perception, reduces GPU memory and TTFT by ~50%, and discards 82.5% of visual tokens on offline benchmarks.
latent objectsstreaming videovideo-llmmemory anchorsvisual context
From Textual Requirements to Microservice Architectures - A Comprehensive Evaluation of LLM-Based Design Synthesis
The study evaluates OpenAI o3's capability to synthesize microservice architectures from textual requirements, comparing zero-shot (ZS) and few-shot (FS) prompting. Using precision, recall, and F1-score metrics, FS outperforms ZS in service identification (F1 = 0.97 vs. 0.79) and communication recovery (F1 = 0.82 vs. 0.61), reducing unsupported dependencies. Expert assessments confirm FS architectures as more modular, coherent, and plausible. Results, though limited to two small systems (Bookstore, PetClinic), suggest LLMs can bridge requirements engineering and architectural design when guided by exemplars.
microservice architectureslarge language modelsfew-shot promptingrequirements engineeringarchitectural design
MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians
MonoVoc introduces a training-free pipeline for monocular open-vocabulary 3D scene understanding by decoupling geometric reconstruction from semantic integration. The method processes monocular video to produce an object-level semantic Gaussian map, avoiding dense language embeddings in favor of lightweight post-processing. Evaluations on Replica show preserved rendering fidelity and segmentation accuracy while reducing memory usage by an order of magnitude compared to state-of-the-art approaches.
3d gaussiansopen-vocabularymonocular videosemantic segmentationmemory efficiency
CACHE-UK: A Stability-Aware Memory Editor for Sequentially Updated Quantized LLMs in Finance
CACHE-UK introduces a stability-aware memory editing framework for sequentially updated 4-bit quantized LLMs in finance, addressing the 'quantization stability crisis' where existing methods suffer catastrophic degradation. The method combines a rank-1 LoRA perturbation mechanism, financial domain prioritization, and a closed-loop Stability Controller tracking 'degradation debt' to prevent catastrophic forgetting. Evaluated on a 4-bit OpenLLaMA-3B with 88,021 UK financial documents, CACHE-UK reduces knowledge degradation by 11-17% versus baselines and achieves a 28% test success rate (6pp improvement), though absolute generalization remains low.
quantization stabilitymemory editinglora perturbationcatastrophic forgettingdegradation debt
Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3
The paper introduces Tycho, a coding-agent system for ARC-AGI-3 that constructs and utilizes game-specific programmatic world models during interactive play. Tycho separates actionable observations from intermediate frames, enabling agents to model, test, plan with, repair, or bypass executable hypotheses. Evaluated on 25 public games with Claude Opus 4.8, actor-requested delegation achieved 88.49 mean Relative Human Action Efficiency (RHAE), while GPT-5.6 Sol and Opus 5 reached 100.00 RHAE, completing all 183 levels with 61% fewer actions than human baselines.
programmatic world modelsrelative human action efficiencyactive abstractiondeterministic moore machinesinference budgets
MemHarness: Memory Is Reconstructed, Not Replayed
The paper introduces MemHarness, a framework enabling large language model agents to actively reconstruct retrieved memories based on current context, rather than replaying them verbatim. A unified policy model critiques and adapts past experiences conditioned on the present state, using end-to-end training with GRPO. Experiments on ALFWorld and WebShop demonstrate MemHarness's superiority over pure RL and static memory-augmented baselines, with notable robustness in out-of-distribution scenarios. The reconstruction objective also acts as latent guidance, enhancing the agent's intrinsic reasoning capabilities.
memory-augmented agentscontextual reconstructionnegative transferend-to-end trainingout-of-distribution robustness
Agentic Method for Deterministic Validation of Legacy Code Migration
The paper introduces the Locksmith Loop, an agentic test-synthesis method for deterministic validation of COBOL-to-Java legacy code migration. The method prepares two runtime environments—COBOL source and Java target—instrumented with mocks and executed off-mainframe. An iterative agentic loop performs Witness Search over input mocks to penetrate program branches, followed by parity-preserving mutations. When routing boundaries are reached, an analyzer identifies Locked Paragraphs, conditions preventing deeper exploration. Evaluated on three COBOL-Java case studies (430–4,114 source lines), Locksmith achieved nearly complete coverage on two open-source programs and 91.90% branch coverage on an internal production-like COBOL program, with Java matching COBOL under deterministic parity checks.
locksmith loopwitness searchparity-preserving mutationslocked paragraphcobol-java migration
Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation
The paper introduces Theia, a methodology for constructing and validating a large-scale multimodal dataset (100,000 images) from the vision-only Incidents1M for disaster response, addressing text-image misalignment in existing datasets like CrisisMMD. Using Qwen3.5 architectures (4B dense and 35B MoE), it generates high-fidelity captions and validates them via an image-blind Qwen3.5-9B evaluator, simulating modality gaps in Data-Free Knowledge Distillation (DFKD). Evaluation on 173,179 label pairs shows 78.65/100 semantic agreement, with conservative captioning (77.6% Precision, 46.0% Recall), minimizing false positives and exposing annotation inconsistencies.
vision-language modelsdata-free knowledge distillationmixture-of-expertsmultimodal datasetsemantic alignment
LLM-Guided Evolutionary Search for Constraint Model Reformulation to Improve Solver Efficiency
The paper introduces LLM-guided evolutionary search for constraint model reformulation to enhance solver efficiency, employing an Automatic Heuristic Design (AHD)-inspired framework where an LLM proposes candidate reformulations verified via benchmarking. A novel Profile-Diverse Retention (PDR) strategy, using Maximal Marginal Relevance (MMR) on runtime vectors, retains behaviorally diverse attempts, outperforming recency- or performance-based strategies. Systematic evaluation on eight CSPLib problems shows iterative reformulation yields significant held-out speedups, with diverse-context strategies and validation-based selection improving performance.
constraint model reformulationevolutionary searchlarge language modelsautomatic heuristic designmaximal marginal relevance
Operationally Guided Placement-Aware Learning for Industrial Online 3D Bin Packing
OPAL introduces an operationally guided placement-aware learning framework for industrial online 3D bin packing, combining an Operationally Guided Empty-Maximal-Space generator (OG-EMS), operational candidate representations, and a masked ranking policy trained with proximal policy optimization. OG-EMS prioritizes low, stable, and diverse placements, while an xLSTM-based Placement Encoder models geometric and operational dependencies. On BED-BPP, OPAL achieves 0.49 mean space utilization, with 15.1% and 6.3% improvements from candidate generation and learned ranking, respectively.
3d bin packingoperational guidanceempty-maximal-spacexlstmproximal policy optimization
EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE
EgoGenesis introduces a novel egocentric world-action simulator for synthesizing high-quality manipulation videos to augment scarce real-world training data. The method leverages a pretrained video generation prior enhanced by two geometry-aware conditioning mechanisms: Online Anchored Projective Memory (OAPM), which maintains a first-frame 3D scene anchor and periodically updates recent states during autoregressive generation, and Action-3D Rotary Position Embedding (A3D-RoPE), which encodes end-effector motion with camera-aware 3D rotary coordinates for precise control. These components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Augmenting 400 real trajectories with 400 synthesized trajectories improves out-of-distribution real-robot success rates from 77% to 84% on single-arm tasks and from 53% to 70% on dual-arm tasks.
egocentric videovideo generation prioronline anchored projective memoryaction-3d rotary position embeddingautoregressive generation
Agentic Metaverse Services: A New As-a-Service Paradigm
The paper introduces Agentic Metaverse Services (AMServ) as a novel service paradigm combining agentic AI and metaverse technologies, proposing Meta-AaaS (Agent-as-a-Service in metaverse) as its realization. It analyzes how generative AI enhances agent capabilities—autonomous learning, multi-modal interaction, content generation, and collaborative decision-making—to enable customized services in virtual ecosystems. The work outlines AMServ's forms, principles, and applications while identifying future research directions for this emerging service computing paradigm.
agentic servicesmeta-aaasgenerative aimetaverseservice computing
AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge
This study provides the first comprehensive empirical evaluation of generative AI reliability in Islamic knowledge domains, assessing six systems on fifty questions spanning Qur'anic interpretation, Hadith, Fiqh, and jurisprudential diversity. Using mixed-method analysis of responses from Australia and the UK, it measures accuracy, hallucinations, citation verifiability, and Madhhab consistency. Results indicate AI systems serve as introductory learning aids but lack reliability for authoritative rulings due to unverified references and inconsistent handling of jurisprudential uncertainty.
generative aiislamic knowledgehallucinationjurisprudential consistencysource verification
CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising
The paper introduces Contrastive Denoising Autoencoder (CDAE), a lightweight framework enhancing perturbation robustness in pretrained BERT embeddings by jointly optimizing contrastive and reconstruction objectives. CDAE learns perturbation-invariant representations through multiple perturbation strategies (synonym substitution, masking, word dropout) of varying strengths. Experiments show CDAE preserves higher embedding similarity under perturbations compared to BERT and SimCSE, with improvements scaling with perturbation intensity. The method maintains semantic fidelity while improving stability, demonstrating the viability of perturbation-invariant learning for sentence embeddings.
contrastive denoising autoencoderperturbation robustnesssentence embeddingsbertsimcse
EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
The EMBL AI Librarian introduces a life-sciences knowledge layer for AI agents, replacing traditional keyword-based search with natural language queries to Europe PMC (40M+ records). A single LLM orchestrates retrieval by planning subqueries, selecting papers, and extracting evidence. Evaluated on literature synthesis, claim verification (ScholarQABench: +16 F1), open-domain QA (LitQA2: +8 points for GPT-5.4), and biology tasks, it outperforms baselines and improves expert consensus alignment. The system is publicly released as a retrieval layer for agentic pipelines.
knowledge retrievallife-sciencesnatural language interfacellm orchestrationevidence extraction
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
The authors introduce Qwen-UI-Agent, a foundation GUI agent for real-world deployment across mobile, desktop, web, and DeepSearch environments. The system integrates GUI and CLI operations in a unified action space, employs an AutoResearch-style data flywheel for autonomous improvement, and utilizes online RL with >10,000 concurrent environments for long-horizon training (100+ turns). Evaluation shows SOTA mobile performance (82.1% on MobileWorld, 97.5% on AndroidDaily) and competitive results on desktop (79.5% OSWorld-Verified) and browser tasks (73.6% WebArena) against frontier models like GPT-5.6 Sol.
gui agentunified action spaceonline rldata flywheelautoresearch
Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
The survey systematizes security threats to world-model-based embodied AI across its lifecycle, identifying vulnerabilities in data construction, representation learning, state grounding, and trajectory evaluation. It demonstrates how traditional attack vectors (e.g., poisoning, adversarial examples, prompt injection) manifest uniquely when targeting learned dynamics, affordance estimates, or safety costs. The work proposes a taxonomy mapping attacks to world-model security properties, evaluation protocols for safety failures, and defenses spanning robust grounding, uncertainty-aware prediction, and deployment assurance, while noting the dual role of world models as both safety shields and potential sources of predictive illusions.
world modelsembodied aisecurity threatsadversarial examplessafety shields
Vibe-FDTR: An agent-oriented framework for reproducible frequency-domain thermoreflectance data analysis
Vibe-FDTR introduces an agent-oriented framework for reproducible frequency-domain thermoreflectance (FDTR) analysis using LLM agents. The method combines a configuration-driven FDTR package with procedural agent skills to translate natural language requests into verifiable analysis steps. Evaluated on synthetic and real-data tasks, the framework achieves 100% and 98.9% success rates, outperforming ablated variants (91.4%/36.7% without skills, 38.6%/0% without domain package), while reducing computational cost by 87.7% and execution time by 60%.
frequency-domain thermoreflectanceagent-oriented frameworkllm agentsthermal metrologyprocedural skills
The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials
The MADRS Pipeline introduces a LLM-based system to automate depression assessment in clinical trials by processing audio interviews. The method converts audio to transcripts, maps them to ten MADRS symptom items, estimates severity, and flags problematic clinical ratings. Evaluated on real clinical interviews, it achieves a 0.867 correlation with expert ratings, offering interpretable support for structured psychiatric assessments.
madrs scaleclinical trialsllm pipelinesymptom severitystructured interviews
Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation
This work evaluates the robustness of commercial image-moderation APIs based on large foundation models against simple, model-agnostic image transformations. The authors conduct a black-box analysis of three commercial services, testing seven transformations across datasets, harm categories, and perceptual-similarity constraints. Results show that all services can be bypassed using inexpensive transformations like color inversion and grayscale conversion, which preserve human-recognizable content while altering moderation decisions. Robustness varies significantly across datasets and harm categories, with multimodal content and self-harm being particularly vulnerable. The study concludes that foundation-model-based APIs alone do not provide reliable security boundaries and should be part of layered moderation pipelines.
image-moderationfoundation modelsblack-box evaluationperceptual-similaritymultimodal content
Persistent Gaussian Perturbations Prevent Oversmoothing in Recurrent Graph Neural Networks
The paper introduces persistent Gaussian perturbations as a novel mechanism to prevent oversmoothing in recurrent graph neural networks (GNNs). By modeling the system as a stochastic dynamical system with noise injection after each propagation step, the authors prove geometric ergodicity and derive a positive lower bound on the expected stationary Dirichlet energy, ensuring non-collapse of node representations. Theoretical and experimental results show the method's efficacy, with the Dirichlet energy scaling proportionally to noise variance and the graph's spectral gap.
oversmoothinggraph neural networksdirichlet energygeometric ergodicitystochastic perturbations
Integrating AI into Requirements Quality Learning in Software Engineering Education: A TPACK-Guided Empirical Study
This study presents a TPACK-guided integration of multi-agent AI into requirements engineering education, demonstrating how structured assignment design shapes student-AI interaction. Using mixed-methods analysis of 72 master-level submissions (N=100), researchers found selective AI use focused on analytical support rather than automation, with strongest alignment on concrete quality dimensions like value articulation (structural) and testability. Results revealed conditional trust, active refinement behaviors, and improved quality criteria awareness, alongside moderate usability challenges, validating TPACK's efficacy for aligning AI affordances with pedagogical goals in SE education.
requirements engineeringtpack frameworkmulti-agent aipedagogical scaffoldingquality criteria
AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach
The paper introduces Agentic Speech Recognition (AgenticASR), a novel audio-to-clean-text task that removes disfluencies and resolves self-corrections while preserving speaker intent. The proposed AgenticASR system employs an ASR--Refiner architecture with bounded active context to enable continual emission and revision of text streams. Evaluated on the bilingual AASR-Bench benchmark, AgenticASR outperforms existing systems in intent-preserving transcription, with human--AI agreement studies validating rubric-based judgments. Ablations analyze Refiner capacity, context length, and quality--latency trade-offs.
automatic speech recognitiondisfluency removalintent preservationonline revisionbilingual benchmark
Search Strategies for Optimal Classification and Regression Trees
The paper introduces a general algorithmic framework for optimal decision trees (ODTs) that unifies existing search strategies and enables the definition of new ones, addressing scalability challenges in ODT construction. The framework provides a systematic way to compare 18 search strategies, revealing their individual contributions to performance. Empirical results show that the best strategy achieves superior anytime performance for classification and reduces regression runtime by over 10x compared to state-of-the-art methods.
optimal decision treessearch strategiesscalabilityclassificationregression
Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models
We introduce LATCH, a training-free framework for candidate-aware early exit in diffusion language models (DLMs) that separates termination and acceleration decisions. LATCH combines Confidence-Verified Commit (CVC), which verifies confidence and argmax stability over dynamically extracted candidate spans, with Block-Wise Early Commit (BWEC), which accelerates non-final blocks using local rules. Evaluated on 11 zero-shot tasks using LLaDA and Dream, LATCH maintains within 2.0 percentage points of full-decoding accuracy while achieving 9.3-17.8x speedups on short-answer tasks and 2.0-3.3x on long-reasoning tasks, with a single frozen hyperparameter set.
diffusion language modelsearly exitconfidence-verified commitblock-wise early commitzero-shot tasks
OPLD: On-Policy Latent Distillation for Multimodal Reasoning
We propose On-Policy Latent Distillation (OPLD), a framework for transferring multimodal Chain-of-Thought reasoning capabilities into latent representations. Unlike prior visual-latent methods that align latent states with compressed visual features, OPLD supervises latent representations at the reasoning-process level, enabling internalization of abstract reasoning induced by multimodal CoT. Extensive experiments on diverse multimodal benchmarks demonstrate that OPLD consistently outperforms existing latent reasoning methods, achieving state-of-the-art performance. Results indicate that reasoning-process-level supervision provides a more effective paradigm for multimodal latent reasoning than conventional feature-level alignment.
latent distillationmultimodal reasoningchain-of-thoughtvisual-latentreasoning-process
Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
The study introduces ParliamentBench, an open-source benchmark framework based on Secret Hitler to evaluate LLM capabilities in deception, persuasion, and reasoning under information asymmetry. Using 1,600 simulated matches across 16 LLMs, the authors propose three novel metrics for social deduction, reasoning, and deceptive consistency. Results show frontier models (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, DeepSeek 3.1 Terminus) perform strongly, while weaker models fall below random (33%) and algorithmic (45%) baselines, with most LLMs failing to maintain deceptive consistency (<50% retention).
large language modelssocial deductioninformation asymmetrydeceptive consistencybenchmark framework
Asymmetric Communication: Large Language Models and Language Games
The paper introduces the concept of asymmetric communication to characterize human-Large Language Model (LLM) interactions, arguing that attributions of general intelligence, hallucination, agency, sentience, and alignment to LLMs constitute a category mistake. It posits that LLM outputs circulate communicatively without normative commitments or entitlements, with correctness, accountability, and practical standing enforced solely by human receivers. Drawing on Wittgenstein, Luhmann, Esposito, and Brandom, the framework reclassifies these properties as receiver-side phenomena, grounding guardrails as structural necessities rather than machine moral agency. The analysis concludes that AI alignment should be understood as institutional constraint engineering, with responsibility remaining with human institutions.
asymmetric communicationlanguage modelsnormative scorekeepingalgorithmic contingencyinstitutional constraint
Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
The study audits whether general-purpose helpfulness rubrics can distinguish answer-giving from pedagogical guidance in LLM tutoring. Using Claude Opus 4.8 and GPT-5.6 Sol as judges, it compares conversational and pedagogical policies across three tutor bases with a fixed weak simulated student. Results show no significant helpfulness difference between policies under Opus (Cliff's |δ|=0.10) but perfect rank-separation under pedagogy (|δ|=1.0). Helpfulness ordering is judge-contingent, reversing in two bases, while pedagogy contrasts retain direction. Answer-revealing turns reduce independent student work. The findings suggest pedagogy-targeted rubrics and deterministic process measures are needed for reliable tutor evaluation.
llm tutoringhelpfulness rubricpedagogical guidanceanswer leakagedeterministic detectors
ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs
ConMem introduces a contribution-aware memory framework for LLM-assisted long-horizon steel-equipment inspection, addressing the limitation of static corpus treatment in retrieval-augmented systems. The method segments inspection logs into functional evidence units, estimates diagnostic value via Shapley-style contribution scoring, and retains high-value evidence under memory constraints. Evaluated on real-world data, ConMem achieves 76.0% QA accuracy, reduces input tokens by 88.2%, and cuts response time by 86.6% versus 8K-context LLM baselines, while preserving early degradation signals for targeted inspection.
retrieval-augmented generationshapley valuememory budgetlong-horizon inspectionweak signal detection
Towards Practical Algorithm Selection for Unsupervised Domain Adaptation in Medical Imaging
We propose a label-free criterion for joint algorithm and hyperparameter selection in unsupervised domain adaptation (UDA) for medical imaging. Given a pool of candidate models from multiple algorithms trained with different hyperparameters, our approach scores each candidate against an agreement reference constructed without target labels. The reference combines multiple label-free selection signals to nominate models within each algorithm, then aggregates these nominations across algorithms to form reference predictions. The candidate with predictions most aligned with this reference is selected. Experiments on four brain MRI and four chest X-ray datasets across seven clinical transfer scenarios demonstrate superior selection performance compared to other methods, validating the approach's effectiveness across diverse algorithm pools.
unsupervised domain adaptationlabel-free selectionalgorithm selectionhyperparameter tuningmedical imaging
Information Bottleneck Learning for Faithful Time Series Forecasting Explanations
The paper introduces IB-Forecast, an interpretable multivariate time-series forecasting framework that guarantees faithful explanations via a budget-constrained information bottleneck. It decomposes predictions into learned periodic and residual components, with explainable masks over input tokens enabling direct control of explanation sparsity. Evaluations show IB-Forecast matches black-box forecasting error (e.g., requiring only 14-20% of observations for low-error predictions) while outperforming gradient-based, occlusion-based, and optimization-based baselines in faithfulness under matched sparsity constraints.
information bottlenecktime-series forecastingfaithful explanationsinterpretable modelssparsity control
BlueprintRepair: Typed Local Edits for Failed Lean Proof Blueprints
BlueprintRepair introduces a typed local edit interface for repairing failed proof blueprints in Lean, enforcing schema-checked operations that preserve theorem integrity while allowing lemma modifications. The method compares typed edits, exact patches, and complete rewrites using DeepSeek-V4-Flash and Qwen3.6-Flash on BlueprintTrace, a benchmark of 142 controlled failures. Results show typed repairs are most cost-efficient (1.30x cheaper than patching, 2.06x than rewriting) and achieve near-maximal coverage within 10k tokens, outperforming free-form interfaces in localized failure resolution.
proof blueprintlocal editschema-checkedlemma modificationcost-efficiency
Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
The study demonstrates that organizing synthetic textbook data into coherent book-level documents improves language model pre-training, beyond benefits from content or local rewriting. The authors introduce a scalable pipeline that retrieves, clusters, and assembles source-grounded sections into 686K textbooks (32B tokens) across 15,000+ disciplines, with hierarchical tables of contents. Controlled experiments show that this 'Full' setting outperforms content-matched 'Split' (+1.02 mean gain), length-matched 'RandomConcat', and retrieval-pool-matched 'Rephrase' (+1.17 gain) conditions, confirming the value of structured synthesis. Llama3-8B also benefits, validating book-level organization as a key factor in synthetic data design.
synthetic textbook databook-level organizationhierarchical tables of contentslanguage model pre-trainingcontrolled experiments
MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
Memory Intent-Aware Neural Denoising (MIND) introduces a lightweight defense framework against memory injection attacks in memory-augmented LLM-based agents. MIND leverages an intent-aware Information Bottleneck (IB) to extract compact intent-behavior representations, filtering task-irrelevant information while preserving cross-turn attack signals. A lightweight detector identifies malicious memories from these representations, mitigating information redundancy and avoiding LLM auditing overhead. Experiments on ReAct-StrategyQA demonstrate MIND's efficacy, reducing mean ASR-r and ASR-a by 55.4% and 55.3%, respectively, while maintaining task accuracy and inference efficiency comparable to undefended agents.
memory injection attacksinformation bottleneckintent-awarellm-based agentsneural denoising
An Instrument to Evaluate Governance Proposals: AI Policy Analysis at Scale
The paper introduces a multidimensional policy analysis framework for AI governance, focusing on surfacing tradeoffs and normative assumptions rather than prescribing outcomes. It employs a mixed-methods approach integrating qualitative expert insights with computational text analysis to design empirically grounded rubrics. These rubrics quantify policy objectives and enable comparative visualizations for interpretability. The framework benchmarks commercial large language models against a domain-trained rubric-calibrated model, emphasizing relevance and alignment across policy attributes. Contributions include transparent hybrid methodology, empirically grounded rubrics, and domain-trained model benchmarking.
policy analysisrubric calibrationcomputational text analysislarge language modelsai governance
PerturbMap: Cross-Context Transfer of Single-Cell Perturbation Responses
PerturbMap enables cross-context transfer of single-cell perturbation responses by combining a recipient-local low-rank base with source-to-recipient ridge experts, weighted by route reliability. The method transports measured source responses through expert models trained on paired perturbations, improving prediction accuracy for missing recipient-context effects. On the Perturb-CITE-seq melanoma cohort, PerturbMap reduces full-effect MSE by 4.1% versus a local base model and outperforms FedAvg, zero-response, and calibrated-copy controls, achieving 80.5% top-10 counterpart retrieval accuracy versus 74.5% for the base model.
single-cell perturbationcross-context transferridge expertslow-rank baseperturbation response
Diversifying Personalized Research Ideation against AI-Induced Homogenization
DivAlign introduces a four-stage pipeline for diversifying AI-assisted research ideation while preserving researcher alignment. The method extracts fine-grained researcher profiles, generates profile-conditioned candidates, scores them along Executability, Comprehensibility, and Growth Potential dimensions, and selects researcher-local directions to reduce community redundancy. Evaluated on 95 AI researchers across five subfields, DivAlign reduces average pairwise similarity from 0.331 to 0.294 and nearest-neighbor similarity from 0.704 to 0.608 compared to coarse single-shot ideation, while retaining 99.9% of researcher-direction fit versus independent top-choice baselines.
research ideationalignment-preservingde-homogenizationprofile-conditionedcommunity redundancy
Distilling Answer Set Programming Theories from Large Language Models
This work introduces a neurosymbolic approach for distilling complete and correct Answer Set Programming (ASP) theories from large language models (LLMs) within a fixed agent harness. The method employs a dataset-agnostic protocol, using a single prompt and an empty file as the starting point, with a 1-hour time limit. Evaluated on VQA benchmarks (CLEVR, GQA, CLEVRER) across nine LLMs, three frontier models achieved 100% accuracy on CLEVR and 92.8%-98.8% on GQA, while GPT-5 showed significant performance drops on GQA (41.8%) and CLEVRER (86.7%). Adding handwritten reference theories minimally impacted other models but reduced GPT-5's accuracy by 3-19 percentage points.
answer set programmingneurosymboliclarge language modelsvqa benchmarksagent harness
On a joint simultaneous learning of relevant feature subsets and subspaces in regression-like problems
The paper introduces Entropy-Optimal Manifold Regression (EOMR), an extension of Entropy-Optimal Manifold Clustering (EOMC), for joint identification of relevant feature subsets and subspaces in nonstationary, nonlinear regression. EOMR achieves robust learning with linear iteration and memory complexity. Evaluated on chaotic dynamics (Lorenz-96 with F=8, F=12) and tokamak plasma data (Hasegawa-Wakatani model), EOMR outperforms gradient-boosted random forests, deep neural networks, and TabPFN v.03, yielding significantly lower RMSE and model complexity. For Hasegawa-Wakatani, EOMR distills a linear, causal autoregressive process (8 parameters) describing Essential Orthogonal Function dynamics.
entropy-optimal manifold regressionnonlinear regressionchaotic dynamicsessential orthogonal functionautoregressive process
Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction
The authors introduce Chem World, a large-scale benchmark for chemical property prediction comprising 17 datasets with over 800,000 molecular samples spanning diverse properties like density and solubility. They propose Mixture-PINN, a physics-informed neural network that integrates chemical priors into data-driven learning to enhance accuracy and robustness. Experiments show Mixture-PINN outperforms existing methods, establishing Chem World as a standardized platform for evaluating trustworthy AI in computational chemistry.
chemical property predictionphysics-informed neural networkmolecular samplesbenchmarkcomputational chemistry
Group-Reflective Self-Distillation for Agentic Reinforcement Learning
The paper introduces Group-Reflective Self-Distillation (GRSD), a method for enhancing reinforcement learning with verifiable rewards (RLVR) by deriving capability-aligned guidance from an agent's own verified rollouts. GRSD contrasts reflections from successful and failed trajectories within an on-policy group, using a stop-gradient snapshot to construct privileged guidance. A self-teacher then refines turn-level credit assignment by modulating outcome-based advantages while preserving verifier-determined learning. Experiments across multiple environments and model scales show GRSD outperforms baselines and improves generalization to unseen tasks.
reinforcement learningself-distillationverifiable rewardscredit assignmenton-policy learning
Temporal Poisoning: Clean-Label Backdoors via Event Redistribution in SNNs
The paper introduces clean-label temporal poisoning, a novel backdoor attack on Spiking Neural Networks (SNNs) that redistributes event timestamps without altering labels or event counts. Unlike dirty-label attacks, it applies fixed temporal transformations exclusively to target-class training streams, preserving aggregated statistics while manipulating the event sequence processed by SNNs. Evaluated across three neuromorphic datasets and both convolutional and transformer-based architectures, the attack achieves a 1.00 Attack Success Rate (ASR) in optimal configurations. Ablation studies analyze poison budgets and trigger shapes, while adapted defenses reveal vulnerabilities: rate-collapsing methods fail by design, whereas feature-space detection succeeds selectively. A proposed model-free detector using per-step event mass identifies temporal transformations, delineating attack stealth boundaries. This is the first clean-label backdoor attack demonstrated on SNNs with event data.
spiking neural networksclean-label backdoortemporal poisoningneuromorphic datasetsattack success rate
Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
Echoverse introduces a framework for training computer-use agents through deep, evolving synthetic environments that co-evolve with the model. The system compiles specifications into stateful applications with graded tasks, leveraging a co-evolution loop that updates both the environment and the model based on rollout evaluations. Experiments demonstrate significant improvements: a 9B parameter model achieves 67.1% accuracy across fourteen splits, nearing frontier model performance. Deep environments enhance live-site accuracy (80.0 → 85.0), while interface control drills improve generalization to unseen widgets and the open web. Reinforcement learning with a combined reward function boosts held-out scores from 58.8% to 68.0%. Four environments are released as benchmarks.
computer-use agentssynthetic environmentsco-evolution loopstateful applicationsreinforcement learning
SemPIC: Learning Semantic Position-Independent KV Caches
SemPIC introduces a novel approach to position-independent KV caching by training a LoRA-enabled Writer to compile native per-layer document key-value (KV) states through behavioral distillation, while retaining the pretrained decoder as an unchanged Reader. This adaptation is confined to offline cache construction, preserving the standard KV interface and cache-hit decoding path. The method also incorporates KV Gradient Checkpointing to reduce peak training memory without severing gradients through cached KVs. Evaluated across three models and four tasks, SemPIC improves mean micro-F1 from 0.53 to 0.60, nearing Full Recompute performance at 0.62.
kv cachinglorabehavioral distillationgradient checkpointingmicro-f1
Stimulus-Evoked Network Dynamics in Human Cortical Organoids: From a Graph-Computational Framework to Repeated-Stimulation Depression
A graph-computational framework was developed to analyze stimulus-evoked dynamics in human cortical organoids, including stimulus-conditioned functional graphs, a graph-neural-network model for system identification, and graph-level metrics. Applied to longitudinal HD-MEA recordings from three organoids, the study found that evoked responses were fast, near-synchronous network bursts without measurable outward propagation, rendering propagation/integration-depth metrics inapplicable. Reframing analysis around synchrony and response-population size revealed that repeated daily stimulation progressively depressed and spatially contracted evoked responses. Using a developmentally-matched, stimulation-naive control, organoids with prior stimulation sessions engaged only 10% of the array compared to 93% in naive organoids.
graph-computational frameworkstimulus-evoked propagationhd-mea recordingsintegration-depth metricssynchrony
IndustryForge-27B: A Domain-Enhanced Multimodal Foundation Model for Industrial CAD
IndustryForge-27B introduces a domain-enhanced multimodal foundation model tailored for industrial CAD, addressing limitations of general-purpose models and single-task fine-tuning. Built atop Qwen3.5-VL-27B, it integrates six industrial-CAD sub-corpora (~52k multimodal samples) spanning CAD Visual QA, parametric CAD code, assembly-level CAD code, and COM APIs for Inventor/SolidWorks, trained via unified multi-task SFT. The model achieves a +33.65 pp average improvement over the base model on four CAD-domain benchmarks and outperforms GPT-5.4 across all benchmarks, while maintaining general capabilities (+1.56 pp mean, no catastrophic forgetting). IndustryForge-27B serves as a unified substrate for downstream industrial-agent projects, enabling full-stack CAD design and industrial-software operation.
multimodal foundation modelindustrial cadparametric cad codecom apimulti-task sft
SKILL-KD: Contrastive Skill Distillation for LLM Agents
SKILL-KD introduces contrastive skill distillation to improve large language model (LLM) agents by explicitly distilling actionable discrepancies between teacher and student trajectories into textual skill patches. The framework iteratively refines patches through student re-runs, employs Drift-Aware Skill Consolidation to manage edit histories, and prevents skill drift via trace-linked updates. Evaluated across five agent benchmarks and two student settings, SKILL-KD consistently outperforms fixed-model adaptation baselines for frozen student agents.
contrastive skill distillationllm agentsskill driftdrift-aware skill consolidationtextual skill patches
DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness
We introduce DataClawEval, the first benchmark for evaluating autonomous agents in end-to-end data engineering tasks, addressing the gap in existing benchmarks focused on simplified Text-to-SQL translation. Built on production-grade code from professional enterprise data engineers, it includes 100 rigorous tasks across PySpark, MySQL, HiveSQL, PrestoSQL/Trino, and FlinkSQL, executed in isolated sandboxes and graded by deterministic, rule-based scripts. Evaluation of 16 frontier agents reveals critical limitations, with the strongest model achieving only 74.9 overall accuracy and no single model dominating across all engines, indicating strict domain specialization. The dataset, containerized environments, and evaluation scripts are publicly released.
dataclavevalend-to-endpysparkdeterministicsandbox
MUL-T: Decoding Spatial Cellular Architecture in Multiplexed Tissue Images
MUL-T introduces a lightweight transformer framework for analyzing spatial cellular architecture in multiplexed tissue images by reframing tissue organization as a masked contextual prediction task over discrete cell tokens. The model learns contextualized [CLS] embeddings without task-specific supervision, capturing higher-order cellular interactions efficiently. Evaluated on tumor pattern classification, patient grading, PD-L1 positivity prediction, and cross-dataset treatment response, MUL-T outperforms classical feature-based methods and matches a foundation ViT model's performance with fewer parameters and lower training cost.
transformermultiplexed imagingcellular interactionscontextual predictiontissue architecture
VISA: A Structured Description Protocol for Agent-Based Simulation Models Towards Machine Reproducibility
VISA introduces a structured, symbol-based description protocol for agent-based simulation models (ABMs) to enhance machine reproducibility. The protocol specifies models in eight interconnected tables at agent and model levels, supported by nineteen executable consistency rules and three reusable LLM-executable skills for authoring, checking, and code generation. Validation on three independently authored ABMs demonstrates successful cross-language reproduction (NetLogo to Python) and capture of an industrial AnyLogic model, while transparently identifying reproduction barriers due to proprietary dependencies. VISA shifts reproduction challenges from implicit model assumptions to explicit, actionable dependencies.
agent-based modelsmachine reproducibilitystructured description protocolexecutable consistency rulesllm-executable skills
Scaling, Lock-In, and Proxy Compliance: A Political Economy of Responsible AI
The paper develops a sequential political-economy model to analyze AI accountability at scale, focusing on institutional dynamics between vendors, deployers, and regulators. The model examines how vendors choose auditability and mitigation strategies, deployers monitor systems post-adoption with switching costs, and enforcement relies on verifiable evidence. Results identify a proxy-compliance equilibrium where vendors meet observable procurement floors but mitigate below social optimum, explaining persistent post-deployment harms despite documentation and standardized evaluations. Key mechanisms include independent audit rights, portability, incident reporting, and outcome-linked liability, which restore deployer leverage and create detection-independent incentives. The model generates testable implications for monitoring, mitigation, and the gap between formal compliance and operational outcomes.
auditabilitymitigation strategiesswitching costsproxy-compliance equilibriumoutcome-linked liability
Flux-OPD: On-Policy Distillation with Evolving Contexts
Flux-OPD introduces an on-policy distillation (OPD) paradigm for open-ended language model training by leveraging evolving contexts as dynamic supervision. The method decomposes the reverse KL objective to reveal that distillation targets the geometric mean of context-conditioned teachers while quantifying conflicts among them. Flux-OPD injects contextual difference signals into a context-free teacher anchor, weighted by conflict terms, improving over existing OPD methods in open-ended tasks.
on-policy distillationreverse kl divergencecontext-conditioned teachersopen-ended domainscontextual corrections
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
RepBench introduces a benchmark-grounded data layer for capability-aligned representation probing in large language models, addressing limitations of paper-specific synthetic evaluations. The method compiles 13,427 benchmark papers into a taxonomy of 182 capability clusters across 13 families, and harvests 353 public benchmark datasets yielding 46,149 audited probe texts covering 94 capabilities. Cross-benchmark transfer evaluation across twelve models reveals that difference-in-means achieves the highest model-level mean on ten models, while logistic regression wins the most capability-model cells, highlighting the importance of readout method and aggregation criterion. The pipeline, corpus, and evaluation code are released as a reusable closed-loop workflow.
representation engineeringcapability clusterscross-benchmark transferdifference-in-meanslogistic regression
Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures
The study evaluates pathology foundation models (FMs) as backbones for mitotic figure detection, comparing their performance against a ResNet50 baseline. Using UNI, UNI2-h, Virchow, Virchow2, H-optimus-0, and H-optimus-1 FMs with RetinaNet, Faster R-CNN, and Deformable DETR detectors on MIDOG++ and TUPAC16 datasets, the authors find H-optimus-0 and Virchow models achieve competitive detection performance, suggesting their latent spaces are suitable for dense object detection tasks. Results indicate slight robustness improvements in out-of-domain testing.
pathology foundation modelsmitotic figure detectionlatent spaceobject detectionself-supervision
MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware
The paper introduces MMLDSum-LLM, a two-stage framework for multimodal long-document summarization addressing attention drift and cross-modal misalignment. The method combines supervised fine-tuning with visual-alignment and keyword-aware weighted losses, followed by GRPO optimization using multi-objective rewards (keyword coverage, image-text alignment, ROUGE, length). Evaluated on the novel MMLDSum-Bench benchmark against state-of-the-art models, the approach demonstrates significant improvements in key-information coverage (measured by atomic-claim precision/recall) and cross-modal consistency (via image-text alignment and LLM-as-a-judge scoring).
multimodal summarizationlong-document processingvisual-text alignmentkeyword-aware lossgrpo optimization
SKIMIX: Multi-Agent Harness-Time Scaling with Skill Mixture for Dynamic Harness Engineering
SKIMIX introduces a multi-agent framework for dynamic skill engineering, enabling agents with diverse skill portfolios to collaborate through iterative refinement. The method integrates embedding-based skill retrieval, submodular anti-dilution routing, and adaptive skill evolution. Evaluated across six reasoning benchmarks, SKIMIX demonstrates substantial improvements in open-ended mathematical reasoning but limited or negative gains on multiple-choice tasks. Agent-count scaling is non-monotonic, with most benefits occurring in the first refinement round. These findings highlight task-dependent efficacy of skill-level ensembles and provide insights for scalable agent design.
skill retrievalanti-dilution routingskill evolutionagent-count scalingiterative refinement
Driving up Inference Energy on SNNs: Per-Sample and Universal Sponge Attacks
This work demonstrates that Spiking Neural Networks (SNNs) are vulnerable to sponge attacks that exploit their energy-efficient spike-based computation to inflate inference energy consumption. The authors propose two attack models: a per-sample attack using gradient-based optimization to craft adversarial spike trains, increasing SynOps by 1.5-2.6x while maintaining 98% classification accuracy; and a universal attack applying a fixed binary perturbation via XOR to all inputs, increasing SynOps by 1.09-1.24x. Estimated energy overheads on Loihi-1 range from 14 μJ to 13.24 mJ per inference. These attacks pose practical threats to battery-powered edge systems using native event-based SNNs.
spiking neural networkssponge attackssynopsneuromorphic hardwareedge systems
Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation
The paper investigates whether domain specialization in LLM evaluators should be implemented in model weights or deferral rules. Using 99,952 rubric-conditioned examples, it shows that supplying correct rubrics improves locked-test accuracy by 2.11 points, while unrelated rubrics degrade performance by 2.66 points. Specializing via eight criterion-family LoRA judges reduces accuracy by 10.05 points and coverage by 18.01 percentage points, but initializing from a shared trained judge recovers 19.94 points. Learned deferral heads in a 0.6B-4B-8B cascade achieve 89.40% accuracy on RewardBench 2 (vs. 84.75% for 8B alone) at 0.415 normalized parameter compute, passing 95% risk audits.
llm evaluationdomain specializationlora judgesdeferral rulesrewardbench
TAPO: Transition-Aware Policy Optimization for LLM Agents
TAPO introduces transition-aware policy optimization for LLM agents, enhancing standard RL by incorporating dense environmental feedback through action-conditioned next-observation prediction. The method alternates between policy optimization and transition supervision, improving sensitivity to environmental dynamics without additional expert data or inference overhead. Experiments on WebShop and ALFWorld with various foundation models show consistent performance gains over pure policy optimization baselines.
reinforcement learningpolicy optimizationenvironmental feedbacktransition dynamicsllm agents
MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation
The paper introduces MARS-RA, a framework for credit assignment in cooperative multi-agent reinforcement learning by reformulating it as a rank aggregation problem using pairwise comparisons generated by large multimodal models. The method replaces absolute estimation with relative comparisons to enhance robustness against noise and dynamic agent participation, converting these comparisons into contribution scores via potential-based reward shaping. Theoretical analysis confirms convergence and robustness, with Shapley values providing interpretability. Experiments on diverse tasks demonstrate MARS-RA's effectiveness in guiding agents toward cooperative solutions.
credit assignmentrank aggregationmultimodal comparisonspotential-based reward shapingshapley values
Specification-Guided Synthesis of Deadlock-Free Communication Protocol Refinements with Large Language Models
Syntropy introduces a framework for synthesizing deadlock-free communication protocol refinements using large language models (LLMs) guided by multiparty session types (MPST). The method integrates refinement constraints directly into LLM generation to ensure behavioral correctness and compatibility. Evaluation shows Syntropy achieves 95.6%-99.5% validity across multiple LLMs while preserving syntactic correctness and producing diverse refinements.
protocol refinementmultiparty session typesdeadlock freedomlarge language modelsformal specification
$Σ$-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
The paper introduces $Σ$-Mem, an online reliability memory system for LLM-based multi-agent systems that tracks peer competence and inter-peer relationships via real symmetric states updated by correctness feedback. Leveraging Weyl's inequality, it ensures bounded spectral changes during updates, enabling stable adaptation without model retraining. The system supports multiple interfaces for residual steering, peer routing, and reliability-weighted voting. Evaluations across five Qwen-family models demonstrate adaptation to reliability shifts, generalization to unseen peers/tasks, and outperformance of majority voting and fixed peers on OOD evaluations. Performance scales with feedback volume, confirming progressive reliability accumulation.
reliability memorymulti-agent systemsonline adaptationspectral boundscorrectness feedback
SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions
SciSchema.org introduces a multidisciplinary collection of 16 expert-annotated schemas for structured scientific process descriptions across Biology & Biotechnology, Materials & Chemistry, Imaging & Measurement, Physics, and Psychology. The schemas, developed via a human-in-the-loop workflow leveraging large language models for candidate generation and domain experts for refinement, define fields for inputs, outputs, materials, instruments, parameters, and procedural steps. Available in JSON Schema and SHACL formats, the dataset includes intermediate model outputs, expert feedback, and validation scripts, supporting applications in metadata enrichment, knowledge graphs, and cross-study comparison.
schema-miningscientific knowledge graphsjson schemashaclinformation extraction
LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference
LAST introduces a training-free framework for query-dependent visual token pruning in edge-cloud collaborative multimodal large language model (MLLM) inference. It employs a compact edge-side vision-language model (VLM) as a guidance proxy, leveraging the last query token's attention to visual tokens under causal attention for query-aware pruning without cloud-model access or costly aggregation. LAST retains diverse query-relevant visual tokens under a fixed budget, achieving 95.4% of full-token accuracy while retaining only 12.5% of visual tokens across 11 multimodal benchmarks, with low edge-side overhead and reduced cloud-side computation.
visual token pruningedge-cloud collaborationmultimodal large language modelcausal attentionquery-aware pruning
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
The paper demonstrates a fundamental limitation in LLM safeguards that evaluate queries without considering downstream use, showing how attackers can exploit copyable context to bypass protections while maintaining utility. By decoupling model capability from verifiable evidence of intended use, the authors derive a theoretical safety trilemma proving the impossibility of simultaneously achieving Useful Capability, Reliable Safety, and Open Access. Through analysis of dual-use evaluations and adaptive attacks, they propose trusted credentials as a mitigation by introducing non-copyable evidence of legitimate use, establishing conditions where safety guarantees can be strengthened.
llm safeguardsdual-use tasksadaptive attackstrusted credentialssafety trilemma
Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting
The paper introduces Complementary Matrix Gating (CMG) for quantum-inspired fast-weight programmers (FWPs), enabling coordinate-wise memory control in quantum dynamics forecasting. CMG replaces scalar gating with sigmoid matrix gates that retain old states while writing new proposals, preserving bounded convex updates. Evaluated across four FWP architectures combining classical and QKAN-based modules, CMG improves performance in single-step forecasting benchmarks and maintains mean-squared errors below 0.001 in multi-step forecasting of Jaynes-Cummings and transmon-resonator dynamics, outperforming scalar-gated counterparts by ≥91.2%.
fast-weight programmersquantum-inspired kolmogorov-arnold networkscomplementary matrix gatingquantum dynamics forecastingcoordinate-wise memory control
Interpretable Representation via LLM-Driven Generative Disentanglement for Local-Life Service Recommendation
LGRID introduces a generative disentanglement paradigm for Semantic ID (SID) generation in local-life service recommendation, addressing semantic entanglement and black-box representation issues. The method employs an Encode -> Disentangle -> Align -> Quantize pipeline, leveraging joint LLM encoding to preserve cross-attribute dependencies, structured disentangled blocks for attribute-aligned slots, synergistic alignment learning for generative decoding and retrieval, and dual-stream residual quantization for compact SIDs. Experiments on Kuaishou and Foursquare demonstrate LGRID's superiority, achieving a 5.44% relative AUC gain, over 99% attribute-decoding accuracy for coarse geographic fields, and reducing the full-SID collision rate to 39.9%.
semantic idgenerative disentanglementdual-stream quantizationattribute-decodinglocal-life recommendation
From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
Proposes Outcome-Verified Comparative Self-Distillation (OVCSD), a method for LLM agent training that replaces action-level scoring with outcome-verified teacher supervision and comparative trajectory learning. OVCSD constructs a prefix tree from failed student rollouts, invokes a skill-conditioned teacher from reached states, retains successful continuations, and applies localized comparative learning at divergence points. Evaluated on ALFWorld and WebShop across three model scales, OVCSD outperforms skill-free RL and self-distillation baselines by up to 29.7 and 5.4 absolute success-rate points, respectively, with <3% privileged interaction overhead.
llm agentsself-distillationoutcome verificationcomparative learningprefix tree
Shapes from Examples: Foundations of Shape Learning in Recursive SHACL
The paper establishes computational bounds for shape learning in recursive SHACL, focusing on a core fragment equivalent to Description Logic ELI. Given positive (P) and negative (N) example nodes, the method computes a shape expression C validating all P and no N, leveraging recursive shape catalogs under well-founded, stable, and supported semantics. Results include tight exponential-time upper bounds for fitting existence and most-specific fitting computation, with polynomial bounds identified for special cases.
shacldescription logic elishape learningrecursive semanticscomputational bounds
The Geometric Nature and a Free Proxy for Flow-Matching Uncertainty
This work introduces denoising acceleration ($\mathrm{accel}$), a cost-free and generalizable uncertainty proxy for flow-matching (FM) models, addressing the lack of explicit uncertainty estimation in FM-generated actions. The method leverages a geometric interpretation of FM uncertainty as deviation from an ideal affine-isotropic contraction field in the velocity space, measuring trajectory bending via a single forward pass without additional training or resampling. Empirical and theoretical validation confirms $\mathrm{accel}$ as a faithful uncertainty proxy, demonstrating its effectiveness in online failure detection. Results show $\mathrm{accel}$ identifies failing rollouts significantly before termination, matching or outperforming computationally expensive baselines under realistic deployment constraints.
flow-matchinguncertainty estimationdenoising accelerationvelocity fieldonline failure detection
Meta-Task: Turning Terminal Task Synthesis into a Terminal Task for Scalable Agent Training
The paper introduces Meta-Task, a framework that reformulates terminal task synthesis as an executable terminal task to improve reliability and scalability in agent training. The method operates within real container environments to iteratively generate, verify, and execute tasks, ensuring internal consistency and executability. It employs multi-phase task specification design and optional external material support to enhance diversity. Experiments on Terminal-Bench 2.0 demonstrate that fine-tuning Qwen3 models with only 3,221 synthesized trajectories yields 22.5% and 31.8% Avg Pass@1 for Qwen3-14B and Qwen3-32B, outperforming existing approaches with less data.
terminal task synthesiscontainer environmentmulti-phase mechanismllm-as-judgepass@1
ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance
The paper introduces ARD-REFSM, a novel framework for reflection symmetry detection that addresses challenges from asymmetric interference and rotational variance. The method combines an Asymmetric Region Denoising (ARD) module to suppress background clutter and a Rotation Equivariant Feature Similarity Matching (REFSM) module that enforces consistency between original and rotated image features via rotation loss. The authors also present GMSYM, a benchmark dataset with diverse scenarios and interference types. Evaluations on DENDI, NYU, LDRS, SDRW, and GMSYM show state-of-the-art accuracy and robustness.
reflection symmetry detectionasymmetric denoisingrotation equivariancefeature similarity matchingbenchmark dataset
One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs
The paper proposes a neuron-level cross-dimensional safety alignment framework (MLS-Neurons) for multilingual and multimodal safety in large vision-language models (LVLMs). By identifying modality- and language-shared safety neurons through activation analysis and intersection operations, the method enables efficient parameter updates (~0.03%) while transferring English safety supervision to multilingual/multimodal scenarios. Experiments demonstrate superior performance over state-of-the-art approaches on diverse safety benchmarks without compromising general utility.
safety alignmentmultimodal learningneuron activationparameter efficiencycross-lingual transfer
IFHierBench: Hierarchical Instruction Following for Large Language Models
The paper introduces IFHierBench, a hierarchical instruction-following benchmark for large language models (LLMs) with 600 prompts stratified across four constraint-tree depths and 35 distinct constraints, each paired with a deterministic checker. The benchmark evaluates LLMs' ability to satisfy nested constraints at different scopes, unlike existing flat benchmarks. Results show that even top models achieve only ~50% prompt-level accuracy, with performance degrading as constraint depth increases, highlighting a significant gap in hierarchical instruction-following capability.
instruction-followinghierarchical constraintsllm evaluationbenchmarkconstraint adherence
A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models
This study conducts a cross-architecture audit of direction-based inference-time defenses in vision-language models, comparing five defense candidates across 15 model-layer cells from four architectural families. The defenses include mean image conditioning shift, CMRM refusal direction, ShiftDC attack-specific residual, prompt instruction to ignore the image, and a random control, evaluated under a magnitude-controlled protocol. Results show no single defense dominates in both refusal recovery and utility preservation, with the image conditioning shift excelling in LLaVA 1.5 and Pixtral 12B, and the prompt instruction leading in Qwen2.5 VL. The CMRM direction aligns positively with the image conditioning shift across all cells, indicating partially overlapping refusal geometry.
vision-language modelsinference-time defensesresidual streamrefusal recoverycosine alignment
Class-Aware Reinforcement Learning for Counterfactual Explanation Generation
This work introduces class-aware reinforcement learning (RL) for counterfactual explanation (CFE) generation, incorporating an instance's predicted class into the RL state representation alongside feature-derived predictors. The method is evaluated against class-blind RL (excluding class information) across seven diverse datasets, demonstrating superior convergence speed, reward optimization, and episode length reduction during training. Class-aware RL generates significantly more valid CFEs, with class-based features consistently ranking among the most influential predictors in action-selection (per SHAP and LIME analysis), highlighting its efficacy for interpretable model explanations.
counterfactual explanationsreinforcement learninginterpretabilityshap valueslime values
MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
We introduce MMHBench, a multimodal benchmark for mental health understanding in long-form videos, comprising 268 videos and 2,184 questions organized into third-person assessment (605 questions) and first-person perspective-taking (1,579 questions). The Multi-Agent Question Generation (MAQG) framework synthesizes questions from diverse social roles, refined through multi-role feedback and expert-guided verification. Evaluation of 22 multimodal large language models (MLLMs) demonstrates that long-form video mental health understanding remains highly challenging, highlighting the need for nuanced reasoning over observable behavior, interpersonal context, and latent psychological states.
multimodal benchmarkmental health understandingmulti-agent question generationlong-form videosperspective-taking
Dynamic Spectral Filtering for Temporal Graph Learning: Learning Evolving Propagation Operators
Dynamic Spectral Filtering (DSF) introduces evolving graph propagation mechanisms for temporal graph learning, treating Chebyshev polynomial filter coefficients as recurrent temporal states. DSF employs a recurrent branch for updates and multiplicative gates for regulation, maintaining node-count-independent temporal states. Evaluated on MOOC, Wikipedia, and Reddit temporal link-prediction benchmarks, DSF achieves AP scores of 0.7851, 0.9088, and 0.9860, respectively, with 93K to 133K parameters, 68 to 182 MB peak GPU memory, and 1.6 to 2.1 seconds per epoch. Compared to DEFT, DSF uses 8.3 to 8.6 times fewer parameters, 25 to 33 times less GPU memory, and 5 to 19 times less time per epoch, demonstrating computational efficiency.
temporal graph learningchebyshev polynomialdynamic spectral filteringrecurrent temporal stateslink-prediction benchmarks
Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning
The paper introduces Counterfactual Sensitivity Credit Reallocation (CSCR), a method to improve long-chain-of-thought (CoT) reasoning by reallocating credit across tokens based on counterfactual sensitivity analysis. CSCR modifies GRPO by downweighting highly sensitive tokens and renormalizing advantages, addressing limitations of uniform credit assignment in critic-free RL and unreliable privilege-induced directions in on-policy self-distillation. Experiments on long-CoT mathematical reasoning benchmarks show CSCR outperforms GRPO, with ablations confirming that moderate sensitivity-based downweighting optimizes performance while avoiding optimization instability.
reinforcement learningchain-of-thoughtcredit assignmentcounterfactual sensitivityself-distillation
RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents
RoboBRIDGE introduces a modular framework for transforming pretrained Vision-Language-Action (VLA) models into robust robotic agents, addressing limitations in failure recovery, long-horizon execution, and domain shift robustness. The framework orchestrates five modules—Monitor, Perceptor, Planner, Controller, and Robot Interface—to enhance VLA-based robotic manipulation. Key innovations include hierarchical failure recovery, asynchronous replanning and scene updating, and domain-invariant primitive skill fine-tuning via LoRA adapters. Evaluations across LIBERO, RoboCasa, and real-world case studies demonstrate consistent performance improvements over standalone policies and prior VLA deployments, highlighting the necessity of structured orchestration beyond scaling action predictors.
vision-language-action modelsrobotic manipulationfailure recoverydomain shiftlora adapters
ARES: Adaptive Reasoning-Effort Steering for PPA- and Cost-Aware RTL Optimization with LLM Agents
ARES introduces adaptive reasoning-effort steering for PPA-aware RTL optimization with LLM agents, addressing three limitations of prior work: (1) normalized cost accounting per LLM call, (2) demonstration that engineered long-term memory offers no consistent advantage over plain experience concatenation, and (3) dynamic effort allocation via a patience-based policy trained on 21 designs. On unseen test designs, ARES achieves 23-27% lower figure of merit (FoM) compared to 16-23% for fixed-effort baselines at equal cost, closes 83% of the gap to hand-optimized MAC units, and outperforms Dr. RTL by 25% FoM at 12% token cost.
llm agentsrtl optimizationadaptive reasoningppa analysisfigure of merit
An Empirical Study of Coordination Mode as the First-Class Citizen in From-Scratch Multi-Agent Coding
The paper introduces MSEval, a multi-agent from-scratch coding benchmark evaluating real-world software development tasks across 10 domains. It employs LegoGent, an execution engine testing 10 collaboration topologies with periodic sync intervals and CI/CD deployment, alongside TAgent, an automated grader measuring functional success, latency, and token cost. Results from 100 runs show topology significantly impacts performance: structured pipelines achieve fastest convergence and highest quality (30+ point score variation), while managerial oversight degrades outcomes. The benchmark establishes reproducible standards for multi-agent team coordination in software development.
multi-agent codingcollaboration topologiesci/cd pipelinesexecution engineautomated grading
Search as Computation Allocation
The paper formalizes terminal computation-allocation problems, where costly computations update beliefs and impact terminal decisions via loss functions. It derives Bellman equations for optimal computation allocation under fixed budgets, priced computation, and exact certification, linking value of computation (VOC) to information-theoretic quantities. Results show mutual information equals myopic VOC under log loss, while simple regret aligns with knowledge-gradient VOC, though information gain can poorly rank computations. The framework unifies bandit pulls, tree simulations, and node expansions under shared decision-theoretic principles, recovering weighted A* as a special case under heuristic-error models.
computation allocationvalue of computationbellman equationsinformation gainheuristic-error model
Orca: Neural Operators for Causal Reasoning in Continuous Time
Orca introduces a framework for causal reasoning in continuous-time systems using neural operators, addressing limitations of structural causal models that handle static variables and forbid cyclic dependencies. The approach models each node in a causal graph as a time-dependent function and each mechanism as a learned map between function spaces, incorporating feedback loops and irregular time observations. Extending existing neural operator architectures, Orca ensures mechanisms respect temporal causality and treat latent exogenous noise as inferable functions for counterfactual reasoning. The framework is validated on synthetic continuous-time examples, demonstrating effective counterfactual inference. Code is publicly available.
neural operatorscausal reasoningcontinuous-time systemscounterfactual inferencestructural causal models
Back to All-Entity Ranking: Sampler-Dependent Evaluation in Continuous-Time Dynamic Graphs
The paper demonstrates that next-destination prediction in continuous-time dynamic graphs (CTDGs) produces evaluation scores conditional on negative sampling distributions and candidate set sizes, altering Bayes-optimal rankings and model comparisons. Using factorial evaluation, history-based scoring, and representation interventions, the authors show that model rankings and module effects vary across Uniform-20 and full-catalog metrics in 3/4 datasets (LastFM, MOOC, Reddit, Wikipedia). They advocate for all-entity ranking, which evaluates all destinations in a fixed catalog, to eliminate sampling bias in CTDG benchmarks.
continuous-time dynamic graphsnegative samplingbayes-optimal rankingall-entity rankingnext-destination prediction
EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits
We introduce EEG-EditBench, a diagnostic benchmark for probing visual information preserved in EEG-to-image retrieval models through controlled image edits. Built from 200 THINGS-EEG2 test images, the benchmark evaluates eight representative models using 2,137 quality-controlled edits across object identity, attributes, background, and object presence. Results demonstrate that strong standard retrieval performance does not consistently transfer to edit-based evaluation, with fine-grained attribute changes posing the greatest challenge. EEG-EditBench reveals model behavior obscured by aggregate retrieval accuracy, providing a controlled framework for studying preserved visual information in EEG-image models. The code and dataset are publicly available.
eeg-to-image retrievalvisual decodingcontrolled editsfine-grained attributesdiagnostic benchmark
Simplifying Neural Networks During Training
We propose a Neural Collapse-inspired training framework for simplifying deep neural networks during training by monitoring representation dynamics through the Inverse Fisher Criterion. This identifies both the split point between feature extraction and classification layers and the training stage for simplification, enabling replacement of trailing layers with a lightweight classification head. Experiments on image-classification benchmarks with MLP, VGG, and ResNet architectures demonstrate substantial parameter reductions while maintaining accuracy comparable to full models.
neural collapseinverse fisher criterionfeature extractionparameter reductionclassification head
FinanceHarness: Autonomous Financial Deep Research Framework
The paper introduces FinanceHarness, an autonomous framework for financial deep research combining LLMs with specialized financial tools and workflows. The system addresses domain-specific challenges through a layered architecture for data construction, agent execution, and reward modeling, alongside FinanceGym, a verifiable benchmark preventing future information leakage. Expert validation shows an 82% pass rate, while baseline LLMs score below 40%. FinanceHarness improves rubric scores from 25.3% to 32.4% using the same open-weight backbone.
autonomous agentsfinancial deep researchllmsreward modelingpoint-in-time benchmark
AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
AutoSupervision introduces a framework for verifying whether scientific manuscript revisions address reviewer concerns through grounded evidence, leveraging peer-review records as supervision. The method analyzes reviewer comments, author responses, and revised manuscripts to characterize concerns, assess resolution, and identify supporting evidence. Evaluated on 56,000 Nature Communications articles, GPT-5.5 achieves 0.754 in concern characterization but only 0.501 in evidence-based verification, highlighting this as the key bottleneck.
large language modelspeer reviewevidence verificationscientific workflowsgrounded revision
Virtual Process Dossier: A Process-Aware Data Catalogue
The Virtual Process Dossier (VPD) introduces a Knowledge Graph-based data catalogue that captures workflow provenance for multi-stage manufacturing use-cases. VPD enables downstream AI-based optimization tasks by distinguishing datasets generated during individual workflow steps, providing them in a FAIR manner. The framework includes: (1) a VPD ontology as the semantic core, (2) a provenance framework integrating ontology instantiation into production environments, and (3) a user interface for human-centered interaction with the Knowledge Graph. The ontology and implementation are publicly available on GitHub.
knowledge graphworkflow provenanceontologyfair datamanufacturing
Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration
The paper introduces Margin Calibration (MC), a method to improve large language model (LLM) unlearning robustness against relearn attacks by addressing the 'margin cliff' phenomenon. MC combines a non-saturating margin hinge anchored at the reference's per-token margin with a KL probe on an instruction corpus, restoring forget-side optimization pressure. Evaluated on TOFU (Llama-3), MUSE-News (Llama-2-7B-hf), and Phi-3.5, MC reduces post-attack ROUGE-L from 0.41 to 0.18 and lowers membership AUC in 13/14 cases, with retain-side utility as the primary trade-off. Theoretical analysis links MC's efficacy to gradient-dominance conditions and stationary optimization geometry.
large language model unlearningmargin cliffrelearn attacksmargin calibrationgradient-dominance
SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification
The study introduces Spatial Attention Fusion (SAF) Gating, a novel module replacing conventional skip connections in CellViT-based models for nucleus segmentation and classification. SAF Gating concatenates encoder skip and upsampled decoder features, processes them through pointwise convolutions with ReLU, and applies channel-wise softmax to produce per-pixel trust heatmaps. This approach enhances minority class detection, particularly the 'Dead' class, improving multi-class panoptic quality (mPQ) on the PanNuke dataset. Evaluated against six gating alternatives, SAF Gating achieves the highest mPQ (0.471), with a 14.5-point improvement in Dead-class F1 score compared to the ungated baseline.
spatial attention fusionskip connectionspanoptic qualitypointwise convolutionschannel-wise softmax
MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory
MemTxn introduces a transaction boundary system for reliable memory updates and recovery in LLM agents, addressing persistent error propagation in writable memory. The method combines Ordered PatchTest for source-supported write validation, a Temporal Resolver for version selection, and durable snapshot journaling for state recovery. Evaluations show 100% accuracy on item-disjoint audits (60 supported originals accepted, 179 hard negatives rejected), complete state recovery under multi-key faults in LongMemEval-S/LoCoMo, and a 17.06–24.07 F1 improvement over Dense on MemoryAgentBench FactConsolidation across five settings.
transaction boundaryordered patchtesttemporal resolverdurable snapshotmemoryagentbench
Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
We introduce Sign Language Question Answering (SLQA), a novel task assessing semantic understanding of sign language videos through natural language question answering, addressing limitations of predefined SLU tasks. Two benchmarks, SignQA-PHOENIX14T and SignQA-CSL-Daily, are constructed by generating question-answer pairs from existing annotations using template-based methods, covering five reasoning categories: position, structural, visual search, gloss recognition, and translation understanding. A baseline model incorporating Question-Conditioned Modulated Temporal Downsampling and in-domain knowledge transfer outperforms existing vision-language models across all categories, establishing a strong benchmark for SLQA research.
sign language question answeringtemporal downsamplingin-domain transfervision-language modelssemantic understanding
STEREODISCO: Discovering Stereotypicality in LLMs
STEREODISCO introduces a framework for systematically discovering stereotypical semantic axes in LLM internal representations, addressing limitations in prior computational research. The method constructs ~2,000 candidate semantic axes from WordNet antonym synsets, recovers them as geometric axes via probing, and identifies stereotypical axes through statistical testing of concept projections. Applied to LLaMA-3-8B-INSTRUCT and Mistral-7B-INSTRUCT, results show LLMs exhibit higher inter-model agreement on social group stereotypes than human agreement, indicating divergence from social psychology findings. Novel stereotypical axes (e.g., humble vs. proud, narrow-minded vs. broad-minded) are discovered and validated by human annotators.
semantic axeswordnet antonym synsetsgeometric axisconcept projectionsstatistical testing
Deep Learning for Accelerated Long-Horizon Forecasting of Multicomponent Multiphase Microstructure Evolution in High-Entropy Alloys
The study introduces an AE-GCN-LSTM surrogate framework for accelerated long-horizon forecasting of microstructure evolution in multicomponent, multiphase high-entropy alloys. The method combines a multi-head autoencoder for latent representation of elemental concentration fields and phase-field order parameters with graph convolutional networks (GCN) and long short-term memory (LSTM) networks to model spatiotemporal evolution. The framework achieves accurate predictions over 3,000,000 timesteps, generalizes to unseen conditions (varying precipitate sizes, counts, and compositions), and provides speedups of 7,200–62,300× over conventional phase-field simulations while preserving phase morphology and compositional evolution.
phase-field modelinghigh-entropy alloysgraph convolutional networkslong short-term memorymicrostructure evolution
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
The paper introduces PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable benchmark for evaluating role-playing agents (RPAs) using user simulators. PALATE employs five per-user simulators to engage RPAs in free-form, multi-turn conversations across 300 character profiles, with personalized rubrics measuring user satisfaction. Results show higher agreement with human judgments for personalized rubrics compared to general ones, enabling interpretable evaluation of generic turn quality, long-horizon capability, and per-user experience in 16 candidate systems.
role-playing agentsuser simulatorsmulti-turn conversationspersonalized rubricsbenchmark evaluation
MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes
The paper introduces MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes with VIKR annotations (Visual clues, Identity links, Knowledge units, Reasoning mechanisms) to evaluate large vision-language models (LVLMs) on culture-dependent interpretation. It reveals a 22.6% Visual-Knowledge gap across 26 LVLMs, showing models cover visible content better than required cultural knowledge. The proposed KAR retrieval baseline, built on CultureBase, improves VIKR Success by 3.6-7.4% but trades off Visual coverage for Identity and Knowledge gains.
meme interpretationvision-language modelscultural knowledgediagnostic benchmarkentity-guided retrieval
Can AI Follow In Einstein's Footsteps?
The article identifies a divergence between AI's trajectory in physics discovery and historical human progress: while physics advanced from phenomenological laws to principle-based theories, AI has moved from symbolic regression (equation discovery) to high-accuracy predictors like AlphaFold and GraphCast that lack theoretical interpretability. This inverse trend risks limiting AI's ability to propose paradigm-level theories (e.g., quantum gravity) without developing skills in principled question-posing or theory construction. The authors argue that equipping AI with symmetry-guided, mathematically rigorous reasoning—akin to Einstein's approach—could bridge this gap and enable next-generation theoretical breakthroughs.
symbolic regressionphenomenological lawsquantum gravitytheory constructionparadigm-level discovery
Annotating Topical Legal Insights from Case Proceedings
Proposes LeDA, a web-based Legal Data Annotation system for constructing structured semantic representations from legal case proceedings via dynamic tag creation, addressing the absence of predefined ontologies. The system enables annotators to incrementally discover and tag legal concepts during document examination, supporting downstream tasks like prior case retrieval and judgment prediction. Demonstrates application with 3 assessors annotating Indian Supreme Court proceedings to build concept-based document representations.
legal annotationdynamic taggingsemantic representationcase proceedingsontology discovery
Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis
The paper proposes SentiLLM, a unified framework for Multimodal Sentiment Analysis (MSA) that leverages Semantic-Aligned Structural Abstraction to transform non-verbal modalities into text-like tokens for LLM processing. The method introduces a Dual-Stream Salience-Context Calibration Mechanism, disentangling non-verbal features into focus (salient sentiment shifts) and ambient (background states) streams, then calibrating them in a unified semantic space. Evaluated on MOSI, MOSEI, CH-SIMS, and CH-SIMS v2, SentiLLM achieves superior performance with minimal trainable parameters, demonstrating effective structural abstraction for MSA.
multimodal sentiment analysisstructural abstractiondual-stream calibrationsemantic alignmentlarge language models
SpecCal: Ambiguity-Aware Candidate Calibration for Infrared Spectrum-Based Molecular Structure Reconstruction
SpecCal introduces a training-free candidate calibration framework for infrared (IR) spectrum-based molecular structure reconstruction, addressing ambiguity in IR-to-molecule predictions. The method operates on candidate outputs from existing models, re-ranking them and introducing structurally plausible alternatives guided by spectral consistency. SpecCal is plug-and-play and model-agnostic, requiring no parameter updates. Experiments demonstrate consistent improvements in top-k reconstruction accuracy at both SMILES and scaffold levels across multiple benchmarks. The approach effectively mitigates spectral ambiguity, enhancing molecular reconstruction from IR spectra.
infrared spectramolecular reconstructionspectral ambiguitytraining-free calibrationmodel-agnostic
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts
LoRA Scaffolded Policy Optimization (LSPO) addresses gradient loss in reinforcement learning for mathematical reasoning on 'cliff' prompts, where all rollouts fail and group-normalized advantage vanishes. The method detects cliff prompts, fits a low-rank (LoRA) adapter via supervised learning on ground-truth solutions, re-rolls with the adapted model, and splices successful completions into the RL batch using importance sampling. Evaluated on DeepMath-103K with DeepSeek-R1-Distill-Qwen-1.5B, LSPO outperforms DAPO baselines across 16 benchmark configurations, achieving gains up to +10.7 points on AIME24/pass@4 and a mean improvement of +3.8 points.
reinforcement learninglow-rank adaptationgradient recoverymathematical reasoningimportance sampling
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
The paper introduces Reasoning Consensus, a framework for structural ensembling of LLM reasoning via weighted DAG aggregation from multiple reasoning chains. The method extracts Directed Acyclic Graphs (DAGs) from chain-of-thought outputs, weights steps by independent attestation frequency, and merges them to produce inspectable consensus reasoning. Evaluated across six benchmarks (statutory interpretation, graduate science, narrative multi-hop reasoning, first-order logic), it outperforms majority-vote baselines by up to 3.1% on MuSR-MM, matches self-consistency at equal compute, and shows 54.4-65.4% preference over alternatives. Ensemble weights correlate with LLM-judge rankings (Spearman ρ=0.30-0.51).
structural ensemblingdirected acyclic graphschain-of-thoughtmulti-hop reasoningself-consistency
RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy
RedFlow introduces an offline reinforcement learning framework that converts failure experiences into action-level corrections for flow-matching Vision-Language-Action (VLA) policies. The method employs Context-Aware Corrective Matching to identify failure-inducing actions and retrieve corrective targets from successful contexts, combined with an Adaptive Redirection Objective to reinforce successful actions and redirect recoverable failures. Evaluated on LIBERO and three real-world tasks, RedFlow improves success rates from 56.7% to 74.7%, matching on-policy methods (PPO, GRPO, DDPO) with 10× fewer samples.
offline reinforcement learningflow-matchingvision-language-action policiescorrective supervisionadaptive redirection
VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition
VocalRender introduces a score-native singing voice synthesis system that directly generates singing from lyrics, pitches, note values, and tempo, eliminating the need for predefined durations or time-aligned acoustic guidance. The method employs an interleaved lyric--note representation and an autoregressive diffusion model to produce continuous acoustic latents while predicting output length. Trained on a 2,300-hour dataset, it achieves strong intelligibility, melody control, and speaker similarity, outperforming baselines by 0.42 CMOS in naturalness.
singing voice synthesisscore-nativeautoregressive diffusionacoustic latentscmos
Train Small, Deploy Large: Zero-Shot GNN Transfer Through Geometric Renormalization
The paper introduces a zero-shot transfer protocol for graph neural networks (GNNs) that enables training on geometrically renormalized (GR) coarse-grained graphs and direct deployment on full-resolution graphs without retraining. The method leverages structural similarity across scales, demonstrating that GNNs trained on scaled-down replicas retain predictive performance on original networks while reducing training costs. Experiments on synthetic and real-world networks show preserved representation alignment and predictive trajectories across scales, suggesting scale-equivariant architectures may enhance GNN transferability.
graph neural networksgeometric renormalizationzero-shot transferscale-equivariantpredictive performance
Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation
The paper proposes Private Face Distillation, a framework for publishing privacy-preserving face recognition (FR) training datasets by decoupling identity semantics while preserving geometric structure. It introduces Orthogonal Geometry Preservation to create proxy identities from private representations and Relational Topology Alignment to maintain identity relations. Evaluated on IJB-C surveillance, the method improves TAR@FAR=1e-3 by 3.94% over baselines while reducing source-identity linkability, demonstrating effective utility-privacy tradeoffs for FR dataset publication.
face recognitionprivacy preservationidentity decouplinghyperspherical geometrydataset publication
Hierarchical Latent Reasoning for LLM-based Recommendation
We propose HiLaR, a Hierarchical Latent Reasoning framework for LLM-based recommendation, addressing limitations in existing latent reasoning methods by characterizing layer-wise preference roles and contributions. HiLaR constructs temporal-guided hierarchical user preference representations, aligns them with multiple LLM latent reasoning states, and organizes reasoning from broad preferences to fine-grained intents. It optimizes the reasoning trajectory using final recommendation feedback and layer-aware process rewards based on marginal target-likelihood gain. Experiments on four Amazon benchmark datasets demonstrate HiLaR's superiority over sequential, generative, and LLM-based baselines. Ablation studies confirm the efficacy of hierarchical representation learning, latent alignment, and process-level optimization.
hierarchical latent reasoningllm-based recommendationtemporal-guided representationlayer-aware reinforcementmarginal target-likelihood
A Structured Knowledge Infrastructure for Domain-Specific Data Asset Discovery
The paper introduces a structured knowledge infrastructure to improve domain-specific data asset discovery in enterprise analytics, addressing four root causes of retrieval failures (semantic gap, entity ambiguity, schema drift, asset-usage gap). The solution combines a three-tier dual-purpose knowledge base (179 documents, eight-section template) with a Graph-Guided Retriever (2,859-node knowledge graph) and Scene-Aware Ranker (19-class entity recognition). Evaluated on a commercial advertising warehouse (5,300+ Hive tables), the system achieves a 77.5pp Hit@10 improvement (19.1% to 96.6%) and 21pp knowledge coverage gain (56% to 77%) at 4.84--5.33s latency.
knowledge graphentity recognitionretrieval-augmented generationschema drifthit@10
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
The study evaluates Large Vision-Language Models (LVLMs) using visual illusions as a diagnostic tool, addressing the lack of joint perception-reasoning benchmarks in open-world settings. The authors introduce IllusionReasoning, a benchmark comprising real-world illusion images with annotated question-answer pairs. Testing diverse LVLMs reveals their reasoning capabilities fall short of claims, providing insights for future optimization.
large vision-language modelsvisual illusionsreasoning capabilitiesperception-reasoning benchmarkillusionreasoning
ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
The paper introduces Request-Oriented Compute Sharing (ROCS), a modeling and inference paradigm for large-scale recommendation systems that optimizes efficiency by deferring request-candidate interactions and sharing request-side computations across candidates. ROCS employs Generalized Layer Masking (GLM) for feature-interaction architectures and Deep Cross Attention (DCA) for sequence models, alongside In-Kernel Broadcast Optimization (IKBO) for GPU acceleration. Experiments demonstrate ROCS improves the quality-efficiency tradeoff, achieving up to 3x QPS gains in retrieval models without quality loss and 0.5% LogLoss improvement with 50% QPS increase in ranking models, with successful production deployment across diverse recommendation systems.
recommendation systemsinference efficiencyfeature-interactionsequence modelinggpu acceleration
VeriSkill: A Self-Evolution Framework for Program Verification Skills
VeriSkill introduces a self-evolution framework for program verification skills, addressing limitations of existing methods in identifying skill-specific failures and extracting actionable signals from verifier feedback. The framework attributes verification failures to skill deficiencies, distills diagnostic signatures into reusable lessons, and iteratively refines candidate skills while preserving program semantics. Experiments demonstrate that VeriSkill consistently outperforms baselines across multiple verification tools, agent frameworks, and LLM backends.
program verificationself-evolutionskill deficienciesdiagnostic signaturessemantic preservation
Towards joint scaling laws with optimal batch size schedules
The paper derives joint scaling laws for deep learning by analyzing training dynamics through convex optimization, providing a closed-form optimal batch size schedule for any given learning rate schedule. The method generalizes across optimizers and architectures, characterizing loss in terms of both schedules. Experiments demonstrate that dynamic batch size schedules consistently outperform static baselines, emphasizing their importance in large language model training.
scaling lawsbatch size schedulelearning rateconvex optimizationtraining dynamics
Baikal: Structured Search for Deep Research over Data Lakes
Baikal introduces a structured search framework for deep research over data lakes, addressing limitations of iterative retrieval by clustering heterogeneous evidence into semantic regions and adaptively searching them to balance exploration-exploitation. The method generates region-grounded subquestions, updates value estimates via finding quality rewards, and evaluates policies (random, LLM-guided, Bayesian $ε$-greedy, UCB) across 15 HybridQA (10,993 tables) and TAT-QA (2,757 tables) queries with 227K Wikipedia and 13K financial report passages. GPT-5-mini evaluations show Baikal’s best configuration improves report scores by 28% (HybridQA) and 36% (TAT-QA) over baselines (DeepSearcher, OpenCode variants), attributed to better groundedness, diversity, and utility via semantic region exploration.
data lakessemantic regionsexploration-exploitationbayesian $ε$-greedyucb
New Synchronous Computation Dynamics for Hopfield Networks
The authors propose a synchronous dynamics method for Hopfield networks, termed SD-DDF, to accelerate processing time by updating multiple neurons simultaneously while ensuring convergence and maximal energy reduction per step. This contrasts with traditional asynchronous updates. They introduce the Discrete Differential Filter (DDF) to solve the combinatorial optimization problem of selecting optimal synchronous updates. Theoretical justification is provided, and four computational experiments demonstrate empirical speedup compared to asynchronous dynamics.
hopfield networkssynchronous dynamicsdiscrete differential filtercombinatorial optimizationenergy minimization
MECA: A Mechanism-Centered Agent for Constructing Well-Specified and Valuable Mathematical Conjectures
MECA introduces a multi-agent framework for constructing well-specified mathematical conjectures by jointly developing candidate statements and their supporting mechanisms. Explorer agents propose and refine mechanisms, while critic agents assess mathematical validity and research value, guiding adjustments to assumptions, scope, and conclusions. Evaluated in two settings, MECA outperforms a generate-and-revise baseline in reconstructing target-paper conclusions and constructs 100 semi-open problems from literature-derived seeds. Results demonstrate that mechanism-centered refinement produces precise, research-worthy conjectures that remain challenging for automated provers.
mathematical conjecturesmechanism-centered refinementmulti-agent frameworkautomated proversresearch value
Albilich: Steerable Proof-State Orchestration for LLM-Based Mathematical Research with CAS Integration
Albilich is an open-source agentic framework for AI-assisted mathematical research, integrating long-horizon reasoning, computer algebra systems (CAS), literature retrieval, and SQLite-based context management. It achieves 10/10 solved problems on the RealMath benchmark with CAS and 9/10 without, while also solving open problems from the Kourovka Notebook, including a counterexample to Problem 21.142 and a strengthened proof for Problem 20.2. Ablation studies show a 32.0% token reduction with CAS and higher verifier-rejection rates without the advisor agent, demonstrating CAS-boosted efficiency and proof synthesis capability.
autoresearchcomputer algebra systemsproof synthesislong-horizon reasoningsqlite context management
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
SpatialCLI introduces a three-stage framework to bridge the capability gap between general vision-language models (VLMs) and specialist vision models for embodied tasks. The method first augments VLMs with spatial tools (Call), then optimizes tool use via Cold-Start SFT and agentic RL (Learn), and finally internalizes perceptual capabilities through trajectory verbalization (Internalize). Evaluated on the 516-example SpatialCLI-Bench, SpatialCLI improves Qwen3-VL-8B-Instruct's accuracy from 29.3% to 84.6% with tools and retains 73.8% without tools, outperforming GPT-5.6 Sol (72.1%).
vision-language modelsspatial reasoningtool internalizationcold-start sftagentic rl
RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
RefineSVG introduces a closed-loop visual feedback framework for high-fidelity image-to-SVG generation using multimodal large language models (MLLMs). The method employs an external rendering engine to compare initial SVG outputs against target images, generating a visual residual map (Diff-Map) that drives targeted corrections via ReAct-style feedback. Key innovations include an SVG-oriented semantic vocabulary (52% token compression) and a progressive training pipeline combining supervised fine-tuning, rejection-sampling data construction, and agentic reinforcement learning. Experiments demonstrate superior reconstruction fidelity, structural accuracy, and code efficiency over baselines.
image-to-svg generationmultimodal large language modelsvisual feedbackreact-style correctionsemantic vocabulary
Guiding Large Language Models with Genetic Programming-Evolved Heuristic Knowledge for Dynamic Multi-Mode Project Scheduling
This work proposes guiding large language models (LLMs) with genetic programming (GP)-evolved heuristic knowledge for dynamic multi-mode project scheduling, reversing the typical LLM--GP hybrid direction. The authors extract knowledge from high-quality GP rules and inject it into LLMs via Feature Selection, Feature Hint, Rule Reference, and Rule Follow mechanisms. Evaluations show GP-derived guidance improves unguided LLMs in scheduling performance, token efficiency, decision stability, and rationale feature focus. Feature Selection achieves best token efficiency, while Rule Follow yields strong performance at higher token cost. Guidance representation significantly impacts effectiveness, with simplified decision contexts and explicit logic outperforming feature highlighting.
genetic programminglarge language modelsheuristic knowledgedynamic schedulingtoken efficiency
LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
LabEvolver introduces a training-free framework for wet-lab agents, combining state-grounded inner trial loops (adaptive perception, online planning, safety validation) with outer evolution loops that distill trajectories into reusable skills, strategies, and safety experience. The method reduces pH-regulation completion time by 48.2% and safety-gate intercepts by 60.0% in robotic solution-preparation tasks. On ALFWorld, it improves cumulative success rate within 20 steps from 76.2% (ReAct) to 91.4% over 500 continual tasks, demonstrating generality beyond wet-lab settings.
training-freeepisodic memorystate-groundedonline planningsafety validation
Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
We introduce Rehearse, a method addressing the confidence cliff phenomenon in self-improving autoresearch, where LLM judges' selective accuracy declines from 82.8% to 56.9% as successful modifications accumulate. Rehearse proposes multiple ideas, compares them pre-execution using focused outcome memory of similar past attempts, and selects the most promising for training. This approach raises late selective accuracy to 83.5%. Evaluated across 4,000 budgeted training runs in three loops (nanochat, image classification, time-series forecasting), Rehearse improves endpoint performance under identical training-run budgets. Analysis of 366 modification pairs from AutoSOTA tasks shows initial helpful modification rates drop from 70% to 43% by iteration 6+.
autoresearchconfidence cliffselective accuracyoutcome memorytraining-run budget
Evaluating and Pricing Advertisements in AI-Generated Responses
The study introduces a psychologically grounded agent simulation framework to evaluate and price advertisements within LLM-generated responses, addressing gaps in click-through intent measurement. It distills supervision into a parameter-efficient evaluator predicting intent and ad quality via differentiable estimates. The evaluator outperforms zero-shot LLM judges in relevance sensitivity (79% vs 60-67%), generalizes to 103 fictional products, and aligns with human preferences (86% agreement). A pricing layer is derived, ensuring optimal truthful bidding, demonstrated on best-of-k allocation and extended to non-monotone cases. The differentiable signal also serves as an ad generation objective.
click-through intentagent simulationparameter-efficient evaluatordifferentiable estimatestruthful bidding
Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
The paper introduces the ProofAgent Index (PAI), a governance readiness index for AI agents, addressing the gap between capability and production readiness. PAI combines four dimensions: Evaluation (observed behavior), Context (operating environment), Compliance (rule alignment), and Governance (organizational control). Implemented in the open-source ProofAgent Harness, PAI demonstrates predictive validity in healthcare and finance, showing context engineering's impact on reliability and governance's necessity for auditable deployment. Results indicate capability alone insufficiently predicts readiness, emphasizing the need for visible governance evidence.
ai governanceproduction readinessagent evaluationcontext engineeringregulatory compliance
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
JigShape introduces a visual-geometric reasoning benchmark using tab-and-blank jigsaw puzzles with unambiguous ground truth, addressing ambiguity in texture-repeated regions. The benchmark comprises 95K instances across four grid densities (4×4 to 16×16), evaluating visual-language models (VLMs) on joint visual and geometric reasoning. Zero-shot evaluations reveal poor geometric reasoning in frontier VLMs, with only GPT-5.5 surpassing random baselines on 4×4 puzzles (70% accuracy), while others perform at chance. Supervised fine-tuning achieves >97% on 4×4 but fails on larger grids (5% on 12×12), exposing a scaling cliff in constraint satisfaction. The work highlights geometric reasoning as an unresolved challenge for VLMs.
visual-geometric reasoningjigsaw puzzlesconstraint satisfactionzero-shot evaluationscaling cliff
Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective
This work establishes a unified theoretical framework connecting Submodular Information Measures (SIMs) to classical representation learning concepts, characterizing their geometric and statistical properties. It demonstrates that Total Information objectives capture intra-class structure through variance and covariance measures, while Mutual Information objectives encode inter-class separation via centroid discrimination and Mahalanobis distance. Controlled synthetic experiments validate the theoretical insights across varied settings of variance, covariance, class imbalance, and multimodal overlap. The results provide principled guidance for selecting SIM-based objectives in representation learning tasks.
submodular information measuresrepresentation learningintra-class variancemahalanobis distancemultimodal overlap
Learning Color Grading, No Photo Sharing: Federated Aesthetic Preference Learning for Personalized Image Enhancement
The paper introduces FedPAIE, a federated framework for personalized aesthetic image enhancement that learns user preferences without centralized photo or rating collection. The method combines a lightweight dual-cue aesthetic scorer (0.787M parameters) with a CLUT-based enhancer (0.265M adapted parameters), using fidelity constraints and an excess-gap penalty to prevent over-optimization while preserving natural appearance. Evaluations on MIT-Adobe FiveK and Flickr-AES show effective open-world personalization with a 0.293M-parameter inference model, balancing user preference and image fidelity without paired retouches.
federated learningaesthetic scoringcolor gradingpersonalized enhancementclut
From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models
The authors introduce MiGUE-Bench, a benchmark for evaluating large language models (LLMs) in multi-granularity event analysis, addressing limitations of existing benchmarks in document granularity and task design. They develop MiGUE-Pipeline, an LLM-driven self-correcting annotation framework, to generate high-quality labeled event data at scale. Experiments on state-of-the-art LLMs and retrieval-augmented generation methods reveal capability boundaries in four core tasks—event detection, relation reasoning, structure induction, and future prediction—highlighting deficiencies in complex cross-document narrative understanding.
event analysislarge language modelsbenchmarkretrieval-augmented generationself-correcting annotation
HALO: Heterogeneous Admission through Localized Obligations for Safe Agentic Execution
HALO (Heterogeneous Admission through Localized Obligations) introduces a runtime protocol for safely admitting heterogeneous AI agent responses containing mixed components (notices, requests, handoffs, actions). The method preserves supported components whose declared prerequisites remain valid, rechecks actions before dispatch, and replaces blocked actions only with fresh candidates. Evaluations show HALO met all 96 admission expectations and passed 20 protocol tests, retaining 248/248 supported components in structured-response replay (vs. 0/248 for whole-response policies). In PX4/Gazebo tests, it blocked all stale routes, observed no stale setpoints, and completed all recoveries.
agentic systemsruntime protocolheterogeneous admissionlocalized obligationsstale mitigation
HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data
HealthCAT introduces an interpretable encoder-only Transformer framework for health indicator prediction from wearable sensor data, integrating Attentive Class Activation Token (AttentiveCAT) for class-specific, time-step-level interpretations. The method maps interpretations onto behavioral cycles relevant to health outcomes, enabling individual-level analysis. Evaluated on two real-world datasets with 306 participants, HealthCAT outperformed deep learning baselines by up to 17% in F1-score and 12% in accuracy (p<0.05). Masking experiments confirmed the predictive value of identified time steps (p<0.05), demonstrating their informativeness. HealthCAT advances wearable sensor analysis by combining predictive accuracy with temporal interpretability, supporting health monitoring and intervention design.
encoder-only transformerattentivecatwearable sensor datatime-step-level interpretationhealth indicator prediction
ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning
ReDiPPO introduces a reference-guided Proximal Policy Optimization framework for mathematical reasoning, addressing noisy advantage estimates in long-horizon tasks. It employs a reference-guided critic using privileged reference answers for accurate value estimation and quantifies token-level discrepancies between standard and reference-guided estimates. These discrepancies reweight token-level advantages during optimization. Experiments across mathematical reasoning benchmarks show ReDiPPO improves value-estimation accuracy and outperforms PPO, DAPO, and GSPO baselines in final reasoning performance.
proximal policy optimizationtoken-level credit assignmentvalue estimationmathematical reasoningadvantage estimates
SCOPE: Synthetic Conditional Objectives for Policy Evolution in Black-Box Combinatorial Optimization
The paper introduces SCOPE, a framework for black-box combinatorial optimization that learns synthetic objectives conditioned on search history to guide policy evolution. SCOPE's outer loop adaptively updates objectives based on their effectiveness in discovering promising regions, while the inner loop maintains a portfolio of policies to mitigate single-surrogate risk. Experiments across multiple benchmarks show SCOPE improves search performance under limited evaluation budgets and generalizes across diverse combinatorial structures.
black-box optimizationcombinatorial optimizationsynthetic objectivespolicy evolutionsearch diversity
Arm2Air: Cross-Embodiment Skeleton Transfer for 3D Relay Formation
Arm2Air introduces cross-embodiment skeleton transfer for efficient UAV relay placement in obstructed urban environments, leveraging pretrained robot-arm motion data. The method converts robot-arm motions into ordered skeletons, pretrains a transformer-based transfer platform, and adapts it to the UAV domain using Low-Rank Adaptation with limited target data. Evaluated on nine high-clutter 3D urban maps, Arm2Air reduced median end-to-end planning runtime by 64.9%, increased bottleneck capacity by 32.6%, and reduced relay-position RMSE by 53.6% with only three target-domain training maps. The approach demonstrates computationally and data-efficient transfer of structural priors across heterogeneous embodied tasks.
cross-embodiment transferlow-rank adaptationuav relay placementobstacle-avoidance skeletonstransformer-based platform
Hidden APIs in Language Models: Discovering Reusable Causal Interfaces from Forked Futures
The study introduces forked futures, a method to empirically quantify reusable causal interfaces in language models by comparing hidden states through future operation-induced response distributions. Shared, Local, Mixture, and Distributed interfaces are evaluated under prequential causal description length, with fidelity and capacity constraints. Results show Shared interfaces achieve the lowest held-out description length (0.216 nats on Qwen2.5-1.5B, 0.294 nats on Llama-3-8B) and maintain minimal future-signature distortion. Transplantation analysis confirms Shared interfaces exhibit superior target-correctness, locality, and copy-preservation. Blind testing recovers 14/16 architectures, supporting the existence of economical reusable causal interfaces within tested operation banks.
forked futurescausal interfacesprequential description lengthfuture-signature distortiontransplantation analysis
CORE: In-Context Reconstruction for Unified Tabular Anomaly Detection
CORE introduces an in-context reconstruction approach for unified tabular anomaly detection (TAD), addressing challenges in aligning heterogeneous data and capturing diverse anomaly patterns. The method employs a decorrelated feature alignment module to preserve semantic information in unified representations and formulates TAD as an in-context reconstruction problem, eliminating reliance on labeled or synthetic anomalies. By reconstructing samples using contextual normal samples, CORE captures dataset-specific distributions, with reconstruction errors indicating deviations from normality for unified TAD across unseen datasets.
tabular anomaly detectionin-context reconstructionfeature alignmentheterogeneous datadataset-specific distributions
DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
DualAnchor introduces a gloss-free sign language translation (SLT) framework addressing language-prior degradation and lexical fidelity gaps in LLM-based SLT. The method combines Token-level Prior Anchoring (TPA), which preserves LLM language priors by regularizing the multimodal decoder toward frozen LLM next-token distributions, and Optimal Transport Alignment (OTA), which enhances lexical fidelity via entropy-regularized partial optimal transport between visual and textual tokens. Evaluated on PHOENIX-2014T and CSL-Daily, DualAnchor improves fluency (TPA) and reduces fine-grained lexical errors (OTA), achieving strong overall performance.
sign language translationlanguage prioroptimal transportmultimodal decoderlexical fidelity
Revisiting the Adversarial Robustness of Graph-Based Traffic Forecasting
The study revisits adversarial robustness in graph-based traffic forecasting, critiquing prior evaluations for unrealistic threat models and proposing a physics-informed detection approach. The authors introduce a detector that feeds into a hardened forecaster, trained against adaptive attacks with fixed forecaster parameters. Evaluated across 15 model-dataset settings, their method outperforms adversarial training in 13 cases, reducing target-link error by severalfold with minimal impact on network-wide metrics, while maintaining near-zero clean performance cost. Results highlight the need for application-specific attack constraints in AI security assessments.
graph-based forecastingadversarial robustnessphysics-aware attackstargeted perturbationsdetection-mitigation
World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
World Action Planner introduces a robot planning system that combines Vision-Language Models (VLMs) with a multi-task pose-image conditioned world model for generalizable decision-making. The system enables agents to propose and iteratively refine action plans through optimization and search, leveraging imagined world model rollouts. It demonstrates superior performance in compositional tasks, new layouts, and zero-shot generalization scenarios, outperforming state-of-the-art end-to-end policy models like VLAs and WAMs.
vision-language modelsworld modelrobot planningzero-shot generalizationmulti-task conditioning
Wiring diagram extraction and gluing: a case study in classifying figure skating jumps using 3D dataset
The authors propose a theory of gluing wiring diagrams to address combinatorial complexity in Hasse clustering, enabling iterative applications to achieve equivalent results to single-pass clustering. The method extracts common patterns from sequential data and represents them graphically, with gluing operations facilitating scalable processing. This approach is validated on a 3D dataset for classifying figure skating jumps, demonstrating its feasibility in practical applications with complex sequential data. The results highlight the method's potential for reducing computational overhead while maintaining clustering accuracy.
hasse clusteringwiring diagramscombinatorial complexitysequential datafigure skating jumps
A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response
This work proposes a vision-language model (VLM)-enabled coordination framework for human-UAV teams in disaster response, integrating natural language interaction, mission control logic, and task allocation. The architecture employs Model-Based Systems Engineering (MBSE) to design a VLM Coordinator Agent, UAV Mission Control, and Task Allocator, validated via software-in-the-loop simulation and human-factors evaluation. Preliminary tests with seven participants demonstrated reduced perceived workload (mental demand, effort, frustration) and high ratings for AI trust and communication clarity, advancing scalable human-autonomy teaming for high-stakes scenarios.
vision-language modelshuman-uav collaborationmodel-based systems engineeringtask allocationincident command system
Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories
The paper introduces an agentic framework for fine-grained intertextuality analysis, where an LLM extracts and labels textual reuse between classical Chinese histories using exact character spans and a five-dimension typology. Validated on an expert-adjudicated benchmark of 2,533 pairs from the Analects and Book of Han, twelve LLMs achieve 56%-93% precision, with significant cost variation. Scaling to the Twenty-Four Histories reveals stable citation patterns but decreasing literal reuse over time, supporting cultural-attraction theory. The extraction protocol and benchmark are released.
intertextualityagentic extractionllm evaluationcultural-attractionexpert-adjudicated
Is Solving Better Than Evaluating GenAI Solutions?
This study investigates the pedagogical impact of evaluating GenAI-generated solutions versus traditional problem-solving in computing education. A randomized A/B crossover experiment (N=220) was conducted in a junior-level algorithms course, where student groups alternated between solving algorithmic problems and evaluating flawed GenAI-generated solutions across six assignments. Results showed no statistically significant differences in midterm scores, final exam scores, or overall course grades, though homework scores were higher for GenAI-evaluation tasks. Survey data indicated that students who adapted study strategies found GenAI-evaluation more helpful, suggesting that such activities redistribute effort toward verification and judgment without automatic conceptual transfer gains.
genaipedagogicalalgorithmsverificationcrossover
From Minds to Models: The Intersection of Psychology and LLM Behaviours
The study adapts psychological methods to assess racial sentiment bias in large language models (LLMs), testing GPT-3.5T, GPT-4, and GPT-4T across 126 prompts varying racial categories. Using a two-way ANOVA, a small main effect of racial condition was found (F(8, 351) = 2.04, p = .042, partial-eta² = .044), but no model effect or interaction. Sensitivity analysis weakened this finding (F(8, 351) = 1.53, p = .145), and post hoc comparisons showed no robust pairwise differences. Limitations include conflating evaluative bias with historical content valence, prompting calls for refined behavioral measures of model bias.
implicit biassentiment analysislarge language modelsanovapsychological methods
What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
The paper establishes a formal definition of prompt graph engineering, addressing the gap between practical implementations and theoretical vocabulary in AI systems. Through conceptual analysis of persistent sources and grey literature, the authors trace the evolution from dataflow graphs to modern prompt graphs. They propose four constitutive conditions: explicit structure, separation of structure and content, executable semantics, and first-class artifact status. The definition is operationalized via inclusion/exclusion tests, validated against six systems (LangGraph, DSPy, Prompt Flow, AutoGen, CrewAI, Claude Code subagents), and delineated from six neighboring concepts. The contribution provides a shared vocabulary for industry practices and outlines a research agenda across four design tension axes.
prompt graphexecutable semanticsgrey literaturefirst-class artifactconceptual analysis
DeepResearch Agent System
The DeepResearch Agent System introduces a 30B-parameter sparse activation LLM for autonomous research tasks, activating only 3B parameters per token to achieve 3.2× faster inference than dense equivalents. It combines hierarchical attention (128K context) with dual-mode reasoning (ReAct and IterResearch modes supporting 20-step iterations), improving accuracy by 31.2% over single-pass baselines. Multi-tool coordination achieves 92.1% tool-use accuracy, while GRPO-based RL optimization enhances training stability by 35%. Benchmarks show 87.3% on Humanity's Last Exam and 91.2% on WebWalkerQA, with full open-source release of training and inference pipelines.
sparse activationhierarchical attentionreact paradigmmulti-tool coordinationgrpo algorithm
Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings
Drawing-Recode introduces a framework for generating Parametric CAD sequences (SPCC format) from raster 2D CAD drawings by explicitly grounding dimensional annotations to geometric features. The method employs an image encoder for geometric feature extraction, a text recognition module for annotation parsing, and cross-attention with Annotation Grounding Loss (AGL) to link annotations to geometry, followed by CAD code generation via an LLM. Experiments demonstrate superior performance over baselines and robustness on scanned industrial drawings, facilitating digitization for manufacturing automation.
parametric cadannotation groundingraster-to-codecross-attentionspcc format
Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
The paper introduces the Function--Evidence--Validation (FEV) framework to evaluate the scientific accountability of agentic bioinformatics workflows, shifting focus from final outputs to inspectable workflow trajectories. FEV categorizes systems by demonstrated operations, traceable support for actions, and use-case-specific validation. The authors analyze 109 agentic systems and 28 evaluation resources across genomics, proteomics, and drug discovery, finding that planning and tool execution outpace reproducibility and validation. They advocate workflow correctness as a key metric for assessing scientific rigor in bioinformatics agents.
agentic bioinformaticsworkflow correctnessscientific accountabilityfunction-evidence-validationomics automation
Using Large Language Models for Idea Generation in Innovation
This study demonstrates the efficacy of large language models (LLMs) in generating high-quality product ideas, particularly GPT-4, which outperforms human-generated ideas in purchase intent metrics. Three idea pools were compared: human-generated ideas from university students, GPT-4 zero-shot prompts, and GPT-4 few-shot prompts. Evaluation employed market research techniques for purchase intent probability, text mining for idea similarity, and human raters for novelty assessment. Results indicate AI-generated ideas exhibit higher purchase intent (few-shot > zero-shot) but lower novelty and diversity. Notably, AI-generated ideas are seven times more likely to rank in the top 10% of ideas, suggesting substantial advantages in idea quality despite reduced novelty.
large language modelspurchase intentfew-shot promptingtext miningidea novelty
Cross-Embodiment Transfer via Behavior-Aligned Representations
The paper introduces behavior-aligned representations to enhance cross-embodiment transfer in vision-language-action (VLA) models for robot manipulation. These representations, including object bounding boxes, language motions, and end-effector traces, aim to unify large-scale cross-embodiment data by maintaining invariances across embodiments while predicting robot actions. A simulation-based benchmark evaluates transfer effectiveness, revealing that end-effector traces significantly improve transfer, especially with larger datasets and action-free data. The approach boosts sim-to-real transfer, increasing task completion progress by 28% for real robot policies pre-trained on simulation data.
cross-embodiment transferbehavior-aligned representationsvision-language-action modelsend-effector tracessim-to-real transfer
AI Literacy: An Exercise in Power-Knowledge
This paper critiques dominant AI literacy frameworks for overemphasizing technical competency and responsible use, arguing they reinforce passive consumption over epistemic agency. Drawing on Foucault's power-knowledge theory, Freirean critical pedagogy, and digital literacy scholarship, it proposes a tripartite critical AI literacy framework emphasizing contextual use, critical interrogation, and participatory governance. The analysis highlights how unequal AI access perpetuates epistemic injustices, positioning literacy as a means to cultivate agentic engagement with AI systems rather than compliant usage.
ai literacypower-knowledgeepistemic agencycritical pedagogyparticipatory governance
Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games
This work introduces a lightweight two-feature behavioral embedding for normal-form games, capturing the entropy of Nash equilibrium and sensitivity of optimal responses to opponent actions, to predict transfer of strategic capabilities in large language models (LLMs). Unlike structural embeddings that memorize game identities, this approach focuses on decision-making behavior rather than payoff geometry. Experiments demonstrate that the proposed embedding reliably predicts performance changes on held-out games following fine-tuning, showing that strategic capability transfer in LLMs depends on behavioral structure rather than game payoffs.
normal-form gamesnash equilibriumbehavioral embeddingstrategic capabilitiesfine-tuning
ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping
ThreatForest introduces a multi-agent system for automated threat modeling, generating attack trees from source code repositories with pluggable TTP framework mapping (MITRE ATT&CK, CAPEC) and mitigation synthesis. The pipeline decomposes threat modeling into stages—repository analysis, context refinement, threat generation, parallel attack-tree construction, and report generation—orchestrated as a directed graph with verification gates and human validation. Evaluation across seven domains shows panel-measured quality scores of 0.63-0.68 for threat statements and mitigations, but only 0.29 for TTP mapping, identifying the sentence-transformer embedding as the accuracy bottleneck. A single-call baseline doubles mapping defensibility, confirming the encoder's limitation.
threat modelingattack treesttp mappingmulti-agent systemsentence-transformer
Hierarchical Reranking for Scalable Financial RAG System
The paper introduces Hierarchical Reranker, a Retrieval-Augmented Generation (RAG) framework for financial document analysis, addressing challenges in processing hybrid text-table structures and large-scale datasets. The system combines Pre-Retrieval Optimization (query normalization, keyword expansion, table transformation), a two-stage Hierarchical Reranker Architecture for precision, and Long-Context Management via adaptive partitioning. Evaluated on FinQA, FinanceBench, and ConvFinQA, it achieves an NDCG@20 of 0.7918 and superior factual consistency, placing second in ACM-ICAIF '24 FinanceRAG Challenge. The framework enables scalable, accurate financial reasoning for applications like audit reporting and investment analysis.
retrieval-augmented generationhierarchical rerankingpre-retrieval optimizationlong-context managementfinancial document analysis
Expanding Data-Agnostic Pivotal Instances Selection Models with Proximity Trees and Ensemble Learning
The authors propose an interpretable-by-design pivot selection model for decision-making, inspired by hierarchical decision structures and similarity metrics. The method extends beyond single pivots by incorporating pivot pairs from proximity/oblique trees and ensemble techniques, while remaining modality-agnostic via pre-trained networks for data transformation. Evaluations across tabular, text, image, and time-series datasets show superior performance to alternative instance selection methods and competitive accuracy against state-of-the-art interpretable models, achieved with minimal pivot counts.
pivot selectioninterpretable modelsproximity treesensemble learningmodality-agnostic
Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks
The paper introduces automated AI scanners to detect four types of validity flaws in agentic benchmarks: ground truth access, tool failure, guessing vulnerability, and answer format ambiguity. The method involves developing grading rubrics for human labeling and evaluating scanner performance on a held-out test set from Inspect Evals benchmarks. Results show the scanners identified verified quality issues in five widely used benchmarks, though performance varied across criteria and models, highlighting standardization gaps in evaluation practices.
agentic benchmarksautomated transcript analysisvalidity flawsgrading rubricsevaluation standardization
KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models
KAISEN introduces a reproducible five-phase audit pipeline for clinical risk models, covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring. Evaluated on synthetic benchmarks (16 disease tasks, 15 social-determinant axes), key findings include: (i) significance counts correlate with standardized equalized-odds differences (ρ=0.78), (ii) per-group threshold optimization reduces disparities (mean Δ=-0.285), while Platt scaling shows inconsistent effects, (iii) mechanism diagnostics fail under proxy misspecification, and (iv) CUSUM monitoring thresholds lack cohort transferability.
subgroup fairnessequalized-odds differencemechanism diagnosticspost-hoc mitigationdrift monitoring
Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments
Change2Task introduces a system for generating executable coding agent tasks from repository history, addressing the need for scalable training and evaluation data. The method converts merged pull requests into verified tasks by aligning historical evidence with evolved code, reconstructing task states through Patch Reversal, Code Mapping, or Agent Reconstruction, and validating the lifecycle from a healthy base to task and restored states. Evaluated on five task families—Bug Fix, Feature Addition, Test Generation, API Migration, and Security Repair—the system achieves 79.6% verified task construction success from 1,130 source changes, recovering 29.2% more tasks than a pull request baseline. Historical and reconstructed cases show up to 98.0% outcome agreement, with a 10.8% reduction in pipeline expenditure.
repository historypatch reversalcode mappingagent reconstructiontask construction
MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers
MixFrag introduces a fragility-guided mixed-precision post-training quantization (PTQ) framework for Vision Transformers (ViTs), addressing heterogeneous sensitivity to quantization. The method estimates component-level fragility via Kullback-Leibler divergence between full-precision and quantized outputs, then formulates bit allocation as a Multiple-Choice Knapsack Problem for layer-wise precision assignment. On ImageNet-1K, MixFrag achieves competitive classification accuracy; on COCO, it surpasses prior mixed-precision PTQ methods by up to 9.6 AP in object detection/segmentation under MP3/MP3 constraints. Analyses confirm the fragility metric's correlation with learned bit allocations.
vision transformerspost-training quantizationmixed-precisionkullback-leibler divergencemultiple-choice knapsack problem
$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
The paper introduces $β$-OPSD, a generalized formulation of on-policy self-distillation (OPSD) that treats the KL penalty weight $β$ as a tunable parameter rather than fixed at 1. The method derives an optimal policy as a geometric interpolation between reference and teacher policies, then approximates this via efficient logit mixing rather than direct RL optimization. Experiments on mathematical reasoning benchmarks demonstrate improved stability and performance over vanilla OPSD, bridging self-distillation and policy optimization while retaining computational efficiency.
on-policy self-distillationkl penaltypolicy optimizationlogit mixingmathematical reasoning
Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories
The paper proposes Doubly Robust Functional Representation Learning (DR-FRL), a cross-fitted workflow for longitudinal causal inference with irregular functional data. DR-FRL maps irregular histories into estimand-targeted states via functional/temporal encoders, estimates nuisance functions (outcome, treatment, censoring), and provides EIF-targeted diagnostics. The estimator achieves asymptotic linearity under explicit rate, overlap, calibration, and stability conditions when the state preserves EIF-required nuisance information. Simulations demonstrate advantages in high-dimensional functional confounding, informative measurement, weak support, or heavy-tailed pseudo-outcomes. A VitalDB application shows DR-FRL's utility in analyzing ICU-disposition endpoints with irregular laboratory data.
doubly robust estimationfunctional datalongitudinal causal inferenceefficient influence functionnuisance functions
ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs
ScaFE introduces a data-efficient method for pathological scar classification by leveraging LLM-generated clinical feature programs instead of direct vision-language model decisions. The approach synthesizes executable programs from clinical evidence to measure scar attributes locally, ensuring data governance compliance and auditability. Evaluated on 600 photographs across three hospitals, ScaFE achieves 81.0% site-macro balanced accuracy, outperforming BiomedCLIP by 10.0 points and maintaining 72.0% accuracy with only 10% of training data. Iterative refinement improves program executability from 66.7% to 95.0%, with 91.7% of features verified by evidence.
scar classificationllm-generated featuresdata-efficientclinical feature programscross-site evaluation
Graph Neural Network Force Fields for Spin Dynamics in Metallic Magnets
The authors introduce a graph neural network (GNN) framework for learning magnetic force fields to simulate spin dynamics in metallic magnets, bypassing computationally expensive electronic calculations. The method learns an effective magnetic energy functional from electronic data, analogous to machine-learned interatomic potentials, enabling efficient evaluation of spin torques while capturing nonlinear, spatially extended interactions from itinerant electrons. Benchmarked on systems with collinear, noncollinear, and noncoplanar magnetic order, the GNN force fields accurately reproduce spin torques and nonequilibrium spin dynamics, matching direct electronic simulations. This establishes GNNs as a scalable approach for predictive simulations of magnetism across multiple length and time scales.
graph neural networkspin dynamicsmagnetic force fielditinerant electronsnonequilibrium magnetism
Same Graph Cross-Task Transfer in GNNs: Protocols and Predictors
The study formalizes same-graph cross-task transfer between node classification (NC) and link prediction (LP) in graph neural networks (GNNs), proposing a leakage-free evaluation protocol with fixed splits and shared message-passing graphs. Experiments with GCN, GraphSAGE, and GPS backbones reveal directional transfer: NC→LP consistently improves performance on homophilic graphs, while LP→NC is fragile and beneficial only in structure-dominant regimes where LP acts as structural pretraining. The authors introduce the CoTask Score (CTS) to optimize shared encoders for joint NC+LP tasks, showing that homophily and other statistics can predict transfer success and mitigate negative transfer.
graph neural networkscross-task transfernode classificationlink predictionhomophily
The Role of Causality in Algorithmic Recourse
The paper formalizes algorithmic recourse through a causal performative framework, addressing how recommended actions propagate via structural causal models to affect both predictions and true outcomes. It models recourse-induced behavioral responses as a non-convex optimization problem, proving conditions for performatively stable solutions via iterative dynamics. Experiments on credit data show causal recourse reduces gaming incentives and outperforms empirical risk minimization, mitigating distribution shifts from strategic behavior.
algorithmic recoursestructural causal modelperformative stabilitynon-convex optimizationdistribution shift
Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation
The study demonstrates that exact-token replay in transformer-based systems does not require live-state fidelity, as the key/value (K/V) cache alone suffices to carry divergent trajectories. Using a Qwen2.5-derived system, the authors conducted a 200-item experiment comparing retained live cache with one-shot prefill of identical tokens under BF16 and FP32 precision. Results show BF16 caused 166 suffix disagreements and 20 correctness label changes, while FP32 produced no decoded disagreements. Bidirectional transplantation of all 48 K/V layers confirmed that divergent continuations follow their cache donor, with 24/24 and 43/43 success rates at primary and later checkpoints, respectively.
kv cachebf16fp32transformertoken replay
Cybersecurity Detection Classification with Reasoning-enabled Language Models
We introduce a chain-of-thought (CoT) reasoning-enabled triage classifier for cybersecurity detection, addressing alert fatigue in Security Operations Centers (SOCs). The system combines automated prompt optimization, self-training, and reinforcement learning with verifiable rewards, trained on human-labeled Windows endpoint detections. A separate calibrator estimates verdict correctness by reading the full reasoning trace. The classifier achieves 82.6% test accuracy, improving benign recall by 43.0% and malicious recall by 18.3% over direct-label LLM classifiers. Results demonstrate the necessity of the calibrator and superior performance of a finetuned 30B model over general-purpose models.
chain-of-thoughttriage classifierprompt optimizationreinforcement learningcalibrator
Graph Neural Multilevel Preconditioners for Iterative Solvers
Proposes Graph Neural Multilevel Preconditioner (GMP), a learned preconditioner combining algebraic multigrid (AMG) hierarchy with graph neural networks (GNNs) for general sparse linear systems. GMP jointly learns smoothing, restriction, and interpolation operators within a unified GNN framework while preserving AMG's structural priors. Evaluated on 800+ sparse matrices, GMP outperforms classical AMG, ILUT, and single-level GNN preconditioners in certain regimes but incurs overhead in others, revealing trade-offs in multilevel learned preconditioning for scientific computing.
graph neural networksalgebraic multigridpreconditionerssparse linear systemskrylov solvers
Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory
The paper introduces short-term graph memory, a plug-in module for oracle-budgeted molecular optimization that learns from evaluated molecules to prioritize oracle queries. The module employs an online graph neural surrogate to pre-screen candidates, allocating the fixed budget to higher-utility molecules. On a fragment-based generator with a 1,000-call budget, it improves the mean top-10 score without extra oracle cost and maintains performance across all tested generators. Analysis reveals the module's benefit correlates with the backbone generator's search breadth and exploitation efficiency. The method provides a selective budget allocation strategy, with empirical validation of generator-specific gains.
molecular optimizationoracle budgetgraph neural surrogatefragment-based generatorexploration-exploitation
Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification
The paper introduces Kohn-Sham Spectral Embedding (KSSE), a physics-inspired energy-based model that replaces dense CNN classifiers with sparse-graph spectral embeddings evaluated at the Nishimori temperature of a Random-Bond Ising Model. The method maps pre-trained features onto quasi-cyclic low-density parity-check graphs, constructs a regularized Laplacian as a Kohn-Sham Hamiltonian, and solves spectral problems efficiently via FFT and low-order Rayleigh refinement. Graph topology is optimized using star-domain surgery, and multi-scale fractal analysis certifies landscape transitions. Theoretical results include generalized Ihara-Bass identity and fixed-point convergence. On ImageNet-1000 with frozen EfficientNet-B4 features, KSSE achieves 88.93% Top-1 accuracy with 21.24M parameters, outperforming Swin-L and matching ViT-H/14 while reducing model footprint by 10-30x.
kohn-sham spectral embeddingnishimori temperaturerandom-bond ising modellow-density parity-check graphsstar-domain surgery
Negative controls reveal volume-driven confounding in radiomics and imaging foundation model features
The authors introduce READII-2-ROQC, an open-source framework for assessing whether radiomic and deep imaging features capture independent spatial signals beyond tumour volume or acquisition artifacts. The method generates volume-preserving negative controls via voxel perturbation across tumour, background, and whole-image regions, comparing feature behaviour and model performance between original and control images. Applied to three public cancer imaging cohorts (3,552 tumour volumes), the framework extracted PyRadiomics and foundation-model features, revealing that multiple models retain performance after spatial structure destruction, indicating volume-driven confounding, while others show perturbation-sensitive signal. READII-2-ROQC enables scalable quality control for interpretable imaging biomarker development.
radiomicsnegative controlsvoxel perturbationfoundation-model featurestumour volume
QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction
QAdapt introduces a noise-adaptive neural pre-decoding framework for surface-code quantum error correction, addressing challenges posed by heterogeneous, nonstationary hardware noise and simulation-to-hardware distribution shifts. The framework captures local spatiotemporal correlations in syndrome data, sequentially adapts to evolving noise conditions while mitigating catastrophic forgetting, and forwards residual syndrome to a conventional global decoder. Evaluated on 110 synthetic out-of-distribution noise configurations and Google's Willow benchmark data, QAdapt reduces logical error rates by up to 5.79% and backend decoding latency by 9.32% without target-domain fine-tuning, enhancing robustness and decoding efficiency.
quantum error correctionsurface-codeneural pre-decodingcatastrophic forgettingsyndrome data
Windowed thinning and query complexity for the bouncy particle and Zigzag samplers
(No summary returned.)
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
The paper introduces Adaptive Anticipatory Policy Trees (AAPT) to address the latency of GUI agents caused by autoregressive decoding during decision-making. AAPT pre-compiles a bounded conditional policy tree using a frozen multimodal model during idle periods, enabling immediate action execution via lightweight observer matching without new text generation. Empirical evaluation shows AAPT improves success rates from 0.50 to 0.79 within contested decision windows (p=1.8×10^-3), with no incorrect actions. Key requirements include fast observer decoding, valid tree planning, and accurate branch routing, with branch routing identified as the bottleneck. AAPT outperforms baselines in scenarios with enumerable candidate actions.
autoregressive decodingmultimodal modelpolicy treebranch routinglightweight observer
Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs
We present a hierarchical Multilevel Monte Carlo (MLMC) neural critic that resolves the bias-cost trade-off in neural critic estimation for average-reward Constrained Markov Decision Processes (CMDPs). By debiasing simultaneously across trajectory sampling and critic optimization, the estimator achieves long-run critic bias with logarithmic sample cost. Building on this, we develop a primal-dual Natural Actor-Critic algorithm with $ ilde{O}(T^{-1/2})$ optimality gap and constraint violation, establishing the first order-optimal convergence guarantees for infinite-horizon average-reward CMDPs with general policy parameterization and neural critics, without requiring knowledge of the mixing time.
constrained markov decision processesmultilevel monte carloneural criticprimal-dual frameworknatural actor-critic
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
LedgerMind introduces provenance-constrained multimodal agentic reasoning via a Structured Evidence Ledger, addressing limitations in final-answer accuracy evaluation. The method employs a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. Tool outputs are normalized into the ledger, ensuring downstream reasoning cites only active entries, and repair is realized through typed state transitions. Experiments across multimodal reasoning benchmarks demonstrate improvements in both answer accuracy and trajectory-level faithfulness, mitigating failure patterns like unsupported reasoning and entity hallucination.
structured evidence ledgerprovenance-constrainedmultimodal reasoningentity hallucinationadaptive dual-path dispatcher
Reflected diffusion, no-flux continuity equations and confined Lagrangian flows in bounded domains
The paper establishes sufficient conditions for a density/flux pair solving a no-flux continuity equation to admit a confined regular Lagrangian flow in bounded domains, motivated by reflected diffusion models. Key assumptions include interior bounded-variation regularity, boundary collar control, one-sided divergence bounds, and vanishing normal velocity trace. The analysis leverages tangency to extend velocities for Ambrosio-DiPerna-Lions theory, proving boundary current mechanisms cannot be relaxed. Constructed counterexamples demonstrate failure of compressibility bounds despite unique marginal transport. Additional uniqueness results for no-flux Fokker-Planck equations justify ODE-based sampling in reflected diffusion models under minimal regularity.
reflected diffusionlagrangian flowno-flux continuityfokker-planckboundary current
Encryption-Compatible Clustered Federated Learning via Distributed Expectation-Maximization over Metadata
FLAMECHE proposes an encryption-compatible clustered federated learning (CFL) method via distributed expectation-maximization (EM) over metadata, addressing the CFL trilemma of privacy, communication, and computational efficiency. By reformulating metadata-based clustering as distributed EM with additive server updates, it maintains compatibility with secure federated learning schemes while preserving efficiency. Experiments across heterogeneous datasets demonstrate improved client model effectiveness and practical encryption compatibility, enhancing its position within the CFL trilemma.
clustered federated learningexpectation-maximizationmetadata-based clusteringencryption-compatibledistributed optimization
Measuring Distortion in the Empty Regions of Dimensionality Reduction Scatterplots with the Gap Index
The Gap Index (GI) is introduced as a novel quality metric for 2D dimensionality reduction projections, specifically addressing distortions in empty regions often overlooked by existing metrics. GI decomposes the projection space into empty triangles, compares them to their high-dimensional counterparts, and quantifies deformation either as a scalar or regional overlay. Results demonstrate GI's sensitivity to small structural deformations with high visual impact, while maintaining computational efficiency and interpretability. This metric enhances visual analysis by providing reliable distortion measurements in empty areas critical for layout interpretation.
gap indexdimensionality reductionvisual distortionempty trianglesdeformation
Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
Fairness Pruning introduces a lightweight structural intervention method for locating and mitigating demographic bias in large language models (LLMs) by identifying differentially activated neurons in GLU-MLP layers. The method employs minimally contrastive prompt pairs and inference-time activation capture to evaluate signals at the down_proj input, focusing on models up to 3 billion parameters (Llama-3.2 family and Salamandra-2B). Results show that zeroing identified neurons (≤40 in Llama-3.2-1B, <0.031% of MLP width) alters demographic responses while retaining 99.49% of reasoning capabilities, confirming dissociable circuits for bias and model functionality. The intervention causes bidirectional bias destabilization due to unsigned BiasScore, highlighting the need for directional behavior modulation.
fairness pruningglu-mlp layersdifferential activationsbias localizationdemographic bias
Fully Inductive Cardinality Estimation
FICE introduces the first fully inductive cardinality estimator for Basic Graph Pattern SPARQL queries over Knowledge Graphs, generalizing to unseen graphs and relations without retraining. The method employs a coupled GNN architecture: an encoder GNN generates entity and relation embeddings from a factor-graph view, while a decoder GNN composes these embeddings along query join topologies to predict log-cardinality. Trained via neighborhood sampling, FICE reduces median q-error from 13.54 to 5.34 across 10 KGs and achieves sub-millisecond latency by decoupling embedding generation from decoding.
cardinality estimationgraph neural networkknowledge graphsparqlfactor-graph
Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing
This work introduces coherent overlap as a framework for understanding sparse mixture-of-experts (MoE) routing, distinguishing route coherence, candidate quality, and candidate-by-context interaction. Using Expert Subspace Separation Index (ESSI), matched-route residuals, and prefix-controlled factorial experiments, the study analyzes six MoE architectures, including OLMoE, Mixtral, and DeepSeek. Results show substantial expert subspace overlap, yet actual routes outperform matched alternatives in token representation. Selected candidates consistently explain residual representations better than unselected rivals, though prefix context narrows this advantage. Functional redundancy is disproven, as adding experts improves next-token prediction in 24 of 39 cases, and Top-2 routing outperforms Top-1 across all seeds.
mixture-of-expertsroutingexpert subspacetoken representationprefix-controlled
A Distributed Acoustic Sensing Dataset for Vessel Detection and Localization in Submarine Cable Protection
The Marlinks-NS DAS dataset facilitates reproducible research on submarine cable protection by enabling vessel detection and localization via distributed acoustic sensing (DAS). The dataset comprises 74,771 labeled instances from ten days of continuous DAS recordings along a 2,554 m segment of a 28 km buried fiber-optic cable in the North Sea. Each instance includes spectral-energy features from 250 sensing channels, anonymized distance measurements, and AIS-derived metadata. The dataset supports machine-learning tasks for vessel detection and vessel-to-cable distance estimation under realistic marine conditions. Released in HDF5 format, it includes documentation, processing descriptions, and example code for method development and evaluation.
distributed acoustic sensingsubmarine cable protectionspectral-energy featuresais-derived metadatahdf5 format
Semi-Supervised Learning for Molecular Graphs via Ensemble Consensus
The paper introduces a semi-supervised learning method for molecular graphs using ensemble consensus to improve predictive accuracy without relying on label-preserving augmentations. The approach leverages unlabeled molecular data by training graph neural networks with an ensemble consensus objective, which enhances robustness and mimics knowledge distillation effects. Experiments show that a single model trained this way outperforms traditional supervised ensembles across diverse molecular datasets and tasks, while also reducing calibration error.
semi-supervised learningmolecular graphsensemble consensusgraph neural networksknowledge distillation
HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks
HARGO introduces heterogeneity-aware reward-guided optimization for RL post-training of LLMs on HPC tasks, addressing performance gaps in SFT models. The method computes per-response importance weights via confidence-modulated advantage, combining group-level reward contrast and reference model log-probabilities without task-type labels. Evaluated on four HPC tasks, HARGO achieves superior performance (WinRate 54.62%, Data Race F1 91.30%, PLP Similarity 0.8558) compared to nine baselines, demonstrating effective alignment for heterogeneous tasks.
reinforcement learninglarge language modelshigh-performance computingreward optimizationheterogeneity-aware
Filling the Pareto-Optimal Front for Affordance Segmentation on Embedded Devices Using RGB-D Cameras
The paper introduces two methods for affordance segmentation on embedded devices using RGB-D cameras: a hardware-aware neural architecture search with a redesigned search space for depth integration, and a fine-tuning approach with a preprocessing layer to merge depth and RGB data. Both methods aim to optimize for modern portable hardware accelerators, addressing limitations of existing tiny-like approaches. Extensive experiments on real-world datasets demonstrate that the proposed methods achieve Pareto-optimal solutions balancing generalization performance and hardware constraints. The prototype, utilizing a Jetson Nano board and RealSense RGB-D camera, achieves real-time performance within smartphone-compatible energy budgets.
affordance segmentationrgb-d camerasneural architecture searchpareto optimaljetson nano
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
A scalable evaluation framework for Large Language Models (LLMs) is introduced, approximating expert-level assessments of textual outputs. The method employs pairwise comparisons across multiple LLMs, utilizing an Elo rating system to generate stable rankings, with adjustable agreement thresholds for confidence control. Evaluations focus on competency profiles derived from scientific abstracts, demonstrating strong correlation with expert judgments while reducing human intervention. This domain-agnostic approach enhances reliability and efficiency in assessing LLM-generated content across diverse applications.
large language modelselo rating systempairwise comparisonscompetency profilesdomain-agnostic
MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek
We introduce MORFES (Morphological Open-class Recognition-and-Formation Evaluation Suite), a benchmark of 500 expert-verified items designed to evaluate inflectional competence in Modern Greek, focusing on lower-frequency lemmas to assess rule-based rather than memorized responses. MORFES tests recognition and production of inflected forms, addressing a gap in evaluating language models for morphologically rich languages. We evaluate several open language models, including LLaMA, Qwen3, DeepSeek-R1, Magistral, and Kimi K2, and release Sophea-Genesis-1, which outperforms others in inflectional morphology while maintaining general capability parity with similarly sized models.
morfesinflectional competencemodern greeklanguage modelslower-frequency lemmas
TopoFormer: Topology Meets Attention for Graph Learning
TopoFormer introduces a lightweight, scalable framework for graph representation learning by encoding topological structure into attention-friendly sequences. Its core module, Topo-Scan, decomposes graphs into ordered sequences of topological tokens via node or edge filtrations, capturing multi-scale structural patterns. These sequences are processed by a Transformer to generate expressive graph-level embeddings. The method avoids costly persistent homology computations, integrates with standard deep learning architectures, and provides theoretical stability guarantees. TopoFormer achieves state-of-the-art performance on graph classification and molecular property prediction benchmarks, matching or exceeding GNN and topology-based baselines with predictable, efficient computation.
topoformertopo-scangraph representation learningattention-friendly sequencestopological tokens
Uncertainty quantification for trustworthy deep learning: Methods and measures
The survey provides a structured review of uncertainty quantification (UQ) methods in deep learning, focusing on ensemble-based and approximate Bayesian approaches. It categorizes methods into five families: Bayesian neural networks, Monte Carlo Dropout, deep ensembles, efficient ensemble approximations, and last-layer/single-pass approaches, while also covering related topics like evidential networks and conformal prediction. The analysis includes theoretical motivation, implementation, empirical performance, and limitations, with emphasis on ensemble diversity theory and uncertainty measures. The work concludes with open research directions, including efficient epistemic measures for classification and hybrid architectures.
uncertainty quantificationbayesian neural networksdeep ensemblesconformal predictionensemble diversity
Weather Emulators at the Frontier of Heat Extremes Predictability
This study evaluates six deep learning weather emulators—Pangu-Weather, FuXi, ArchesWeather, AIFS, GraphCast, and Aurora—against dynamical systems and statistical baselines for predicting global near-surface temperature and extreme heat at 10-15 day lead times. The emulators demonstrate competitive deterministic temperature forecasting skill compared to physics-based models, albeit with reduced spectral fidelity due to blurring. While all models exhibit predictive capability for extreme heat, most emulators under-represent peak intensities, with IFS recall outperforming the emulators. The findings underscore AI's potential in extended-range temperature prediction while highlighting challenges in reliable early warning systems.
deep learningweather emulatorsspectral fidelitydeterministic forecastingextreme heat
Causal Discovery with Inverted Self-attention for Multivariate Time Series
We propose a novel framework for causal discovery in multivariate time series using an inverted causal self-attention mechanism (CSAM) within a transformer architecture. CSAM emphasizes latent and indirect causal relationships by inverting tokens and inducing sparsity in attention scores, reducing spurious correlations. The framework includes a global causal algorithm for identifying global causal links and a causal verification module for robustness. Experiments on linear and nonlinear datasets demonstrate superior performance over existing methods, highlighting the framework's effectiveness in capturing complex causal structures in multivariate time series.
causal discoveryself-attention mechanismmultivariate time seriestransformer architecturesparsity
Secure Aggregation for Privacy-Preserving Federated Learning on Clinical EEG Data
The paper proposes a privacy-preserving federated learning framework for clinical EEG data using masking-based secure aggregation, combining graph-based communication, threshold secret sharing, and dropout-resilient aggregation. The method supports semi-honest and malicious settings, implemented via the Flower framework, and includes optional Bloom filter-based record-linkage and auxiliary-notary verification. Evaluated on TUH EEG data, the framework hides individual updates while maintaining training compatibility, with semi-honest configurations offering lowest overhead and malicious variants providing stronger consistency at higher computational cost.
federated learningsecure aggregationeeg dataprivacy-preservingthreshold secret sharing
Multi-channel Uplift Policy Learning
The paper introduces ReAlloc, a fast-slow causal framework for multi-channel uplift policy learning in e-commerce marketing budget allocation. The method combines an Orthogonal Teacher for unbiased gradient extraction from short-term logs and an Explanation-Guided Student for structured marginal field distillation over long-term horizons, addressing observational confounding and extrapolation issues. Evaluations on Taobao demonstrate simultaneous improvements in pay order and income metrics.
uplift policy learningorthogonal teacherexplanation-guided studentsimplex-constrained optimizationcross-channel substitution
What Makes Deep Learning Work for Traditional Chinese Medicine Tongue Diagnosis? A Comprehensive Ablation Study
This work identifies six design principles for deep learning in traditional Chinese medicine tongue diagnosis through systematic ablation of 20+ model variants on TongueDx2 (5,109 images) and a merged 11,101-sample dataset. The study evaluates six architectures, four loss functions, five augmentation strategies, and six training approaches under 5-fold cross-validation. Key findings include ConvNeXt-Tiny's parameter efficiency (+2.7% over alternatives), BCE's superiority to Asymmetric Loss, restrained augmentation's necessity, weak-group ensemble benefits (+2.1%), data scaling gains (+20.6%), and catastrophic performance collapse (0.78→0.22) when expanding from 13 to 45 label dimensions. Best models achieved weighted-F1 of 0.6625 (976 samples) and 0.7761 (11,101 samples).
tongue diagnosisconvnext-tinymulti-label classificationweak-group ensembleasymmetric loss
LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning
The paper introduces LM-GRASP, a metaheuristic framework that reformulates GRASP's constructive phase as an online imitation learning task using instance-specific language models. The method employs a decoder-only Transformer trained via behavioral cloning on elite trajectories generated by local search, eliminating the need for offline pretraining or problem-specific feature engineering. On the Taillard PFSP benchmark (ta51-ta60), LM-GRASP improves makespan by 28.4 units over GPU-GRASP, demonstrating the viability of online-trained language models for combinatorial optimization.
combinatorial optimizationonline imitation learningtransformermetaheuristicbehavioral cloning
FinSMART: Financial Sentiment Analysis for Algorithmic Trading through Market-Aligned Reinforcement Learning
FinSMART introduces a market-aligned reinforcement learning framework for financial sentiment analysis, optimizing sentiment signals directly from realized market outcomes. The method combines market-aware data filtering with a discrete asymmetric trading reward to handle financial market noise and non-stationarity. Experiments show FinSMART improves cumulative trading returns by 220% over baselines, with superior profitability, risk-adjusted performance, and support for continuous market-aware retraining using newly observed financial articles.
reinforcement learningsentiment analysisalgorithmic tradingmarket-alignedfinancial llms
From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE Inference
The study demonstrates that mathematically equivalent expert-reduction orders in sparse mixture-of-experts (MoE) models can produce divergent execution paths, isolating this effect in DeepSeek-V4-Flash by freezing local MoE state and varying aggregation semantics. Four reduction schemes (P32, A, B, C) were evaluated, revealing that operand representation and accumulator precision affect model behavior: B-mode orders formed 360 structural classes and 11 continuation basins, while C preserved routes and outputs exactly. Controlled experiments showed post-mHC as an intra-token boundary and full persistent state as a cross-token continuation boundary, with divergence persisting across token boundaries. The findings establish numerical compatibility requirements for sparse-MoE runtimes.
sparse mixture-of-expertsexpert-reductionaccumulator precisionpost-mhc boundarynumerical divergence
Meteosat Third Generation imagery improves CNN-based SSI retrieval
A multi-imager, multi-resolution CNN architecture for 10-minute Surface Solar Irradiance (SSI) retrieval demonstrates that Meteosat Third Generation (MTG)/FCI imagery improves performance under cloudy conditions compared to Meteosat Second Generation (MSG)/SEVIRI. The hybrid SEVIRI-FCI model reduced RMSE by 8.2 W m$^{-2}$ (overcast) and 5.7 W m$^{-2}$ (cloudy) versus SEVIRI-only, with no significant improvement under clear/partly cloudy skies. Compared to the physics-based SARAH-3 product, the hybrid model achieved skill scores of 35% (overcast), 21% (cloudy), and 20% overall but underperformed in clear-sky conditions. Results indicate MTG/FCI's higher spatial resolution benefits CNN-based SSI retrieval primarily when clouds dominate irradiance variability.
surface solar irradianceconvolutional neural networkmeteosat third generationsatellite imageryrmse
GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios
GVR-Coder introduces a novel framework for generating structured Scalable Vector Graphics (SVG) diagrams from complex professional texts, addressing dataset scarcity, layout priors, and visual feedback challenges. The method leverages DocMeetSVG-100K, a large-scale dataset for document and meeting scenarios, and employs curriculum-driven rejection sampling fine-tuning to model complex structures with explicit layout constraints. Reinforcement learning from dual rendering feedback optimizes structural complexity and visual aesthetics, while a generate-verify-repair agent loop provides fine-grained feedback for refinement. Experiments show GVR-Coder outperforms baselines in producing logically coherent and visually appealing diagrams.
scalable vector graphicscurriculum-driven fine-tuninglayout constraintsreinforcement learninggenerate-verify-repair
A Query-Efficient Stochastic Volume Rendering Framework for Time-Varying Implicit Neural Volumes
The authors propose a query-efficient stochastic volume rendering framework for time-varying implicit neural volumes (INRs), addressing performance challenges in interactive rendering. Their four-stage pipeline combines heterogeneous parallelism (ray tracing cores for traversal, tensor cores for batched inference) with ray budgeting and query pruning to reduce INR evaluations. The system achieves ~30-40 FPS at 1024x1024 resolution on an RTX 4090, converges to high-fidelity images, and supports interactive temporal exploration with 1-2 ms timestep updates, enabling direct rendering of time-varying INRs without resampling or caching compromises.
implicit neural representationsvolume renderingdelta trackingheterogeneous parallelismquery pruning
ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
ClawTrack introduces a dual-assessment benchmark for evaluating LLM-based agents in complex workflows, measuring both task outcomes (Task Score) and reasoning processes (Process Score). The framework comprises 320 tasks across 8 domains with 25+ deterministic mock services, using a Process Grader to score reasoning turns along four dimensions: goal alignment, efficiency, information utilization, and result verification. Evaluations of 21 models over 16,000+ trials demonstrate that process scores effectively attribute success and failure, reveal complementary reasoning dimensions, and enable consistent post-training improvements. Result verification emerges as a systematic bottleneck, and the framework shows robustness across different judge LLMs.
llm-based agentsprocess scoretask scorereasoning dimensionsmock services
Learning features from Newton's algorithm: a way to accelerate nonlinear parametrized PDE solvers
The paper introduces a two-stage Newton initial guess strategy to accelerate nonlinear parametrized PDE solvers by leveraging learned features from parameter-space sampling and precomputed solutions. The method constructs two reduced spaces—solution feature space and corrective search direction feature space—using discrete Newton trajectories. For unseen parameters, a regression model predicts a surrogate solution, followed by a residual-minimizing correction via GMRES, yielding an improved initial guess for high-fidelity Newton convergence. This weakly intrusive approach reduces Newton iterations and CPU time, demonstrating significant speedups in numerical experiments on representative PDE problems.
newton's methodparameter-space samplinggmresreduced spacesresidual-minimizing correction
Enhancing Irregular Time Series Forecasting with Continuous-Time Modeling Framework
WrapFlow introduces a continuous-time modeling framework for irregular multivariate time series forecasting, addressing limitations of discretization-based preprocessing and ODE-based approaches. The method employs Continuous-Time Tokenization to encode raw observation events and model long unobserved intervals via gap-aware tokens, processed by a Transformer backbone. A simulation-free training paradigm for Residual Flow Matching learns conditional residual vector fields without numerical-solver simulation or backpropagation during training, enabling high-quality forecasting with fixed rollout steps. Experiments on real-world datasets demonstrate state-of-the-art performance.
continuous-time modelingirregular time seriestransformer backboneresidual flow matchinggap-aware tokens
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
The paper introduces Contrastive Reinforced Policy Optimization (CRPO), a method addressing exposure bias in On-Policy Self-Distillation (OPSD) for agentic LLMs. CRPO reformulates OPSD via contrastive learning, using predictive entropy to distinguish reflective exploration (positive positions) from exposure bias (negative positions), enabling group-wise contrast for stable optimization. Evaluations across 13 reasoning and deep-search benchmarks show CRPO outperforms existing reinforcement learning and self-distillation methods, improving training stability and generalization in long-horizon tasks.
contrastive learningon-policy self-distillationexposure biasreinforcement learningpredictive entropy
Building a User Foundation Model for the Open Web
We introduce a user foundation model tailored for open-web real-time bidding (RTB), addressing the challenge of fragmented and non-persistent user identities. The model employs self-supervised learning on user browsing histories, combining masked language modeling and sequence-level contrastive objectives, and is fine-tuned for click prediction. Pre-training optimizes via an LLM-in-the-loop search over code-level edits, implementing the LLM-as-optimizer paradigm. The encoder improves downstream tasks, yielding +1.197% RIG on bid win-rate and +1.354% RIG on CTR ranker. A 7-day A/B test confirms +2.13% CTR and -1.13% eCPC (80% CI excluding zero).
real-time biddingself-supervised learningmasked language modelingcontrastive objectivellm-in-the-loop
Generalization and Trade-off in Adversarial Training: An RKHS Perspective via Kernel Integral Operators
The paper analyzes adversarial training through an RKHS framework, deriving source-uniform generalization bounds via kernel integral operators. It establishes that adversarial robustness introduces a statistical accuracy loss due to noise-robustness interaction, proven via matching lower bounds on polynomial-spectrum models. A noise-debiased two-stage estimator is proposed, achieving near-minimax rates by removing noise contributions. Theoretical and numerical results demonstrate the trade-off between robustness and generalization, with the proposed method outperforming standard approaches.
adversarial trainingreproducing kernel hilbert spacegeneralization boundskernel integral operatorsminimax rates
It's All Just Vectorization: einx, a Universal Notation for Tensor Operations
The authors introduce einx, a universal notation for tensor operations that addresses limitations in existing frameworks like Numpy, einsum, and einops. By leveraging vectorization as a fundamental transformation mechanism, einx enables both lifting lower-order operations to higher-order ones and decomposing higher-order operations into lower-order components. The notation employs declarative, pointful expressions analogous to loop notation, reducing complex APIs to a small set of elementary operations while maintaining consistent rules across all operations. An implementation embedded in Python seamlessly integrates with existing tensor frameworks, offering improved readability and writability.
tensor operationsvectorizationdeclarative expressionsshape errorspython integration
Generalization Bounds on Optimal Control for Transformer Training and Wasserstein Distributional Robustness
The paper derives finite-sample generalization bounds for Transformers trained via dynamic programming recursions, using a measure-valued formulation of Transformer dynamics. By interpreting training as a finite-horizon Markovian control problem and analyzing a quantized model, the authors obtain explicit bounds via concentration inequalities and Lipschitz stability estimates. The bounds are extended to the base model with explicit approximation error, and the framework connects Transformer generalization to Wasserstein distributionally robust optimization.
generalization boundstransformer dynamicsmarkovian controlwasserstein robustnessquantized model
Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning
This paper introduces a reward decomposition framework for Reinforcement Unlearning (RUL), decoupling verifiability from sparsity to improve unlearning efficiency. Two novel reward functions are proposed: an exponential reward that applies graded penalties based on forbidden-concept occurrence counts, and a PageRank-inspired reward that weights penalties by semantic importance. Experiments on the Real World Knowledge Unlearning (RWKU) benchmark demonstrate that both rewards outperform binary rewards, achieving similar forgetting performance up to 3× faster while preserving model utility. The results highlight reward design as a critical factor in scalable and efficient machine unlearning.
reinforcement unlearningreward decompositionexponential rewardpagerank-inspired rewardmachine unlearning
What Makes Graph Unified? Principles and Generative Sliding-Window Transformer for Graph Foundation Models
The authors propose SliGFM, a Graph Foundation Model addressing heterogeneous node feature unification across domains via topology-aware sliding-window encoding and generative reconstruction. SliGFM orders feature dimensions by topological smoothness, applies a shared sliding-window encoder to transform features into fixed-dimensional tokens, and employs a smoothness-aware transformer to capture transferable relational patterns. The generative reconstruction objective ensures information preservation. Four desiderata guide the approach: formal uniformity, cross-domain transferability, information preservation, and backbone compatibility, enabling effective cross-domain knowledge transfer in graph learning.
graph foundation modelstopology-aware encodingsliding-window transformerfeature unificationgenerative reconstruction
AutoPref: Automatic Discovery of Task-Specific Preference Objectives for Neural Combinatorial Optimization
AutoPref introduces the first LLM-guided framework for automated discovery of task-specific preference objectives in neural combinatorial optimization (NCO). The framework factorizes objectives into a pairwise loss program and a set-aware weighting program, forming a unified programmatic objective space. A staged conditional search strategy with behavioral gates enables tractable exploration of this space. AutoPref outperforms hand-designed baselines across multiple combinatorial optimization problems (TSP, CVRP, FFSP, JSSP) and scales, demonstrating the efficacy of automated objective discovery in NCO.
neural combinatorial optimizationpreference objectivespairwise loss programset-aware weighting programstaged conditional search
TriShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement
TriShield introduces a three-layer defense against privacy backdoors in federated LLM fine-tuning, specifically countering the NeuroImprint attack. The method combines (1) a Parameter Artifact Detector to identify memorization neurons, (2) Stateful Virtual Iteration to entangle optimizer states across steps, and (3) Zero-Utility Orthogonal Projection to eliminate private gradient components. Theoretical analysis shows zero mutual information between gradients and individual samples. Experiments on GPT-2 (117M) and Llama-Guard-3-1B demonstrate 0% reconstruction rate with no utility loss and <5% GPU overhead.
federated learningprivacy backdoorgradient projectionoptimizer statemutual information
Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
We introduce a Bayesian domain weighting method for optimizing multi-domain pre-training data mixtures in Large Language Models (LLMs). The approach infers domain weights from a Dirichlet distribution by incorporating Gamma prior information learned from observations, addressing instability and computational overhead in direct optimization. This method outperforms function-fitting approaches that rely on structural assumptions like rank invariance or scaling laws, which often introduce estimation bias. Experiments demonstrate stable and efficient domain weight learning, identifying optimal mixtures with substantially reduced data requirements compared to search-based methods, enabling scalable optimization in large-scale applications.
bayesian domain weightingdirichlet distributiongamma priormulti-domain pre-traininglarge language models
ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
The paper introduces Physical-Time Flow (PT-Flow), a continuous-time latent world model called ODEWorld that parameterizes sequential data dynamics via an ordinary differential equation (ODE) in a structured representation space. By learning a continuous latent velocity field and enforcing ODE properties, ODEWorld avoids representation collapse and enables arbitrary temporal resolution, backward prediction, and high-quality long-horizon image reconstruction. Experiments show ODEWorld excels in video generation and robotic control while providing planning-oriented information.
continuous-time modelinglatent velocity fieldrepresentation collapseordinary differential equationworld model
Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control
The study evaluates reinforcement fine-tuning (RFT) with TD3 guidance to transfer control knowledge from a reasoning model (GPT-5) to a locally deployable open-weight model for multi-zone variable-air-volume (VAV) control. Using deterministic rollouts to audit the learned critic, the authors found unreliable within-state action ranking (5/10 states correct) despite high temporal correlation (r=0.9998). GPT-5 achieved a 6.2% HVAC electricity reduction without building-specific training but compromised ventilation margins. RFT failed to improve sampled-action returns over 200 steps, with transition errors persisting, suggesting supervised fine-tuning before value-based RFT.
reinforcement fine-tuningmulti-zone vav controltd3rollout verificationgpt-5
S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring
S-CEReBrO introduces a novel Windowed Alternating Attention mechanism for continuous EEG monitoring, addressing Transformer memory bottlenecks by factorizing attention into fixed-size spatiotemporal windows. This approach maintains constant KV cache memory, enabling processing of signals 100X longer than full self-attention and 3X longer than low-rank linear attention, while reducing memory usage by 45% and increasing inference throughput by 2.1X. Pre-trained on >25,000 hours of EEG recordings from >12,000 subjects, S-CEReBrO achieves state-of-the-art performance on 7 of 11 downstream tasks with up to 60% fewer parameters, advancing efficient and generalizable EEG monitoring.
windowed alternating attentionkv cacheeeg monitoringtransformerself-attention
Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances
The authors propose integrating contextual embeddings from self-supervised symbolic music models (Aria, CLaMP3) into the evaluation of expressive MIDI piano performances, addressing limitations of traditional attribute-scoped metrics. They adapt Kernel Audio Distance for symbolic music to measure conditional distributional similarity without requiring note alignment, leveraging embeddings' sensitivity to contextual perturbations. A listening study demonstrates that these embeddings serve as perceptual proxies, achieving human rating agreement comparable to traditional metrics. The authors release Pereval, an open-source library combining attribute-scoped and deep feature metrics for reproducible performance evaluation.
contextual embeddingsmidi pianokernel audio distanceself-supervised learningsymbolic music
Contrastive Concept Importance: Explaining Pairwise Class Decisions Through Automatically Extracted Concept Representations
The paper introduces contrastive concept importance (CCI), a method for explaining pairwise class decisions in black-box models by attributing logit margins between target and foil classes to automatically extracted visual concepts. CCI produces signed scores indicating concept support for target vs. foil, decomposable into target-logit and foil-logit effects, distinguishing globally important concepts from class-pair-specific ones. Evaluated on ImageNet using CRAFT-style concept bases, insertion/deletion curves, and semantic hierarchies, CCI reveals fine-grained model behavior not captured by non-contrastive methods, particularly in distinguishing shared, one-sided, or directly contrastive concept effects.
contrastive concept importancelogit marginclass-pair-specificautomated concept extractionsemantic hierarchy
ZAPs: A Reward Attribution Framework for DeFi Ecosystems with Adversarial-Robust Scoring via Parallel Anomaly Ensemble Detection
ZAPs introduces a reward attribution framework for decentralized finance (DeFi) that combines economic contribution scoring with adversarial robustness via a four-layer defense stack. The method employs protocol-specific percentile normalization, two-layer weighting, and a parallel anomaly ensemble (one-class reconstruction + isolation forest) for sybil detection. On 124,638 transactions from 1,073 malicious wallets, the ensemble achieves 0.923 ROC-AUC (+0.032 over the reconstruction model alone). Simulations show 30-90% adversarial reward reduction with <8% impact on legitimate users; live deployments cut sybil allocations by 56% and increased quality-wallet participation by 49%.
reward attributionadversarial robustnesspercentile normalizationanomaly ensemblesybil detection
Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings
The paper introduces a safety-gated supervisory control system combining rule-based constraints with LLM-generated setpoints for distillation column control. The method employs a forked-twin counterfactual gate with nine pinned constraints, tested on Skogestad's Column A across four configurations (PID-only, linear MPC, ungated agent, gated agent). Results show the gated agent (C3) outperforms linear MPC in target acquisition (IAE ratio 0.361) but requires gate interventions for disturbance rejection (16.03x improvement at upper CI). The gate reduces specification-abandonment attractor effects (P95 IAE from 11.5 to 0.77) and blocks 318 harmful proposals in 250 test cases. Performance remains model-dependent across DeepSeek-V4-Flash and NVIDIA Nemotron-3-Super.
supervisory controldistillation columncounterfactual gatesetpoint optimizationmodel-predictive control
Nanoparticle Networks for Neuromorphic Computing
The authors present a neuromorphic computing architecture using metallic nanoparticle networks interconnected by molecular junctions on a SiO2/Si substrate, demonstrating how static control electrodes transform passive networks into tunable nonlinear dynamical systems. Key design principles include operating near the system's cutoff frequency for optimal nonlinear charge tunneling and linear capacitive memory balance, tuning SiO2 thickness to control electrostatic screening length and memory type (persistent or fading), and introducing structural disorder via heterogeneous molecular junctions to overcome expressivity limits. Disorder breaks spatial symmetries, enabling independent manipulation of signal amplitudes and phases for enhanced dynamic neuromorphic performance.
neuromorphic computingnanoparticle networkselectrostatic screeningnonlinear dynamicsmolecular junctions
FeatFix: Reuse What You Verify through Local Exact-Feature Correction for Faster Cached Diffusion Inference
FeatFix introduces local exact-feature correction for cached diffusion inference, reusing previously computed exact features to reset draft residuals and reduce downstream error. The method operates at fixed sparse layer-timestep sites, replacing complete draft block outputs with exact outputs computed from the same incoming state, avoiding partial replacements or full-timestep recomputation. Experiments across four image and video backbones demonstrate FeatFix accelerates generation up to 6.70× over Vanilla while maintaining competitive output quality.
diffusion inferencefeature correctioncached inferencedraft residualtimestep sites
Robust Estimation of Sparse Numerical Vectors under Local Differential Privacy
The paper introduces Randomized Projection with Clipping (RPC), a robust method for sparse vector mean estimation under local differential privacy (LDP) with multi-item users. RPC projects user data onto random binary vectors and applies clipping to limit adversarial influence, with a novel bias correction technique that eliminates the need for bias-variance tradeoff. Theoretical guarantees for estimation error under attacks are provided, and experiments demonstrate RPC's comparable performance in trusted settings and superior robustness against poisoning attacks in untrusted environments.
local differential privacysparse vector estimationpoisoning attacksrandomized projectionclipping bias
Learning-Augmented and Randomized Algorithms for Line Aggregation with Delays
The paper introduces learning-augmented and randomized algorithms for online line aggregation with delays, focusing on robustness and consistency metrics. For λ ∈ (0,1], a deterministic learning-augmented Balance algorithm achieves (4/λ + 1/λ²)-robustness and (4 + λ)-consistency, while a randomized algorithm in the adversarial model attains (e + 1)-competitiveness, surpassing the deterministic 5-competitive benchmark. A lower bound of e for randomized algorithms is established, improving prior results. A hybrid randomized learning-augmented algorithm combines these ideas, yielding (e/λ + 1/λ²)-robustness and (e + λ)-consistency. Numerical experiments validate theoretical findings.
online aggregationlearning-augmented algorithmsrobustnessconsistencycompetitive analysis
Revisiting Predictive Process Monitoring in the Age of Foundation Models: A Comparative Study of Sequence, Tabular, and LLM Approaches
This paper systematically compares sequence models, tabular foundation models, and large language models (LLMs) for predictive process monitoring (PPM) across multiple datasets and tasks. The study evaluates next activity prediction, remaining time, and time-to-next-event forecasting using controlled benchmarks. Results indicate that sequence models (e.g., LSTMs) outperform others for next activity prediction, while tabular foundation models with in-context learning are competitive on temporal tasks; LLMs generally underperform despite higher computational costs.
predictive process monitoringsequence modelstabular foundation modelsin-context learningevent logs
Neural Network Approximation of Solutions to Fractional Parabolic Partial Differential Equations
The paper develops a neural network approximation theory for solutions to fractional parabolic PDEs with drift and potential terms, introducing anisotropic spectral Barron spaces to measure temporal and spatial regularity separately. Key innovations include using Vandermonde matrices for global-in-time extension of fractional heat semigroups and dimension-independent multiplication estimates. Results include $n^{-1/2}$ approximation bounds for two-layer networks with periodic activations and polynomial-decay conditions under additional regularity assumptions.
fractional parabolic equationsanisotropic barron spacesvandermonde matrixspectral regularityneural approximation
RIPPLE: Generating Multi-Channel Phase, Not Recovering It
RIPPLE introduces a method for generating multi-channel phase relationships directly rather than recovering them independently per channel, addressing a key limitation in audio and seismic waveform synthesis. The approach reinterprets Griffin-Lim as a phase prior initialized from source phase to preserve inter-channel structure, then refines it via rectified flow under an explicit inter-channel phase loss. Evaluated on ambisonics environment transfer and seismic cross-station translation, RIPPLE outperforms recovery-based methods: in seismic tasks, it reduces S-wave polarization error from 57.3° (random baseline) to 33.8°, while maintaining coherence metrics critical for downstream analysis.
phase generationmulti-channel synthesisgriffin-lim priorrectified flowinter-channel coherence
Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold
The paper introduces an expand-then-compress framework for improving reasoning in language models by constructing and distilling a complementary teacher union, rather than relying on a single reinforcement-learning-trained teacher. The expansion stage employs Residual Group Relative Policy Optimization (RGRPO) to train multiple teachers, each focusing on uncovered examples, while the compression stage uses reliability-gated Teacher-Union On-policy Distillation (TU-OPD) to distill knowledge into a student model, weighted by teacher reliability. Consensus-Residual Decomposition preserves specialist behaviors during aggregation. Experiments on mathematical reasoning, code generation, and instruction following demonstrate that the Qwen3-1.7B student outperforms individual teachers by 2.0%, 8.3%, and 6.9%, respectively, while maintaining single-model inference.
reinforcement learningpolicy distillationteacher unionon-policy distillationreasoning solution manifold
Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
Proposes Conditional Retrieval Alignment (CoRA), a gradient-free framework for task-conditioned retrieval in on-device in-context learning. CoRA transforms a frozen encoder into a retriever by selecting encoder layers, constructing an output-derived conditioning space, and aligning input representations via ridge regression, followed by low-rank factorization for efficient indexing. Evaluated on ten textual and four multimodal benchmarks with models including Llama-3.2-1B and MobileLLM-Pro, CoRA achieves effective retrieval without fine-tuning or backpropagation, demonstrated via Raspberry Pi~5 deployment.
in-context learningtask-conditioned retrievalridge regressionlow-rank factorizationon-device inference
DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis
The DS@GT team presents their ImageCLEFmedical Caption 2026 submissions, achieving top performance in Concept Detection (Task 1) via a late-fusion ensemble of ConvNeXt-V2, BiomedCLIP ViT-B/16, and DenseNet-169 with 'Honest Threshold Tuning' (primary F1=0.5790). A training-free KNN retrieval variant on frozen BiomedCLIP embeddings nearly matched this (F1=0.5780). For Caption Prediction (Task 2), they benchmarked scaled foundation models: fine-tuned Gemma-3 27B (0.3571, 3rd place), BLIP with Vizwins merging (0.3564), and zero-shot MedGemma-4B (0.3186), demonstrating cost-performance tradeoffs.
late-fusion ensemblehonest threshold tuningknn retrievalvizwins mergingzero-shot captioning
DAS-PMVC: A Framework for Partial Multi-View Clustering via Dual Alignment and Structure Enhancement
Proposes DAS-PMVC, a partial multi-view clustering framework addressing view misalignment via dual alignment and structure enhancement. The method combines anchor graph structure alignment for initial view consistency, structure-enhanced feature learning via pretraining and multi-view graph convolutional networks, and a dual alignment strategy using contrastive learning and the Hungarian algorithm. Evaluations on multiple datasets show DAS-PMVC outperforms state-of-the-art methods in clustering accuracy.
partial multi-view clusteringanchor graph alignmentgraph convolutional networkscontrastive learninghungarian algorithm
Improving the Robustness/Accuracy Tradeoff Against Adversarial Attacks Using Information Bottleneck Distillation Through Dual Teachers
This work enhances Information Bottleneck Distillation (IBD) by introducing a dual-teacher framework, combining a robust teacher trained with adversarial inputs and a clean teacher trained exclusively on clean data. Features from both teachers are transferred to the student via a cross-layer attention matrix. Evaluations on CIFAR-10 and CIFAR-100 demonstrate improved classification accuracy on clean samples compared to original IBD, while maintaining adversarial robustness. The method achieves competitive performance with state-of-the-art approaches, particularly in harmonic mean between clean and robust accuracy, and analyzes the impact of training settings on the attention module.
information bottleneck distillationadversarial robustnessdual-teacher frameworkcross-layer attentionharmonic mean
Error Analysis of Neural-Network-Based Engression
The paper presents a theoretical error analysis of engression, a method for learning conditional distributions via generative models $Y = f(X,\varepsilon)$ optimized under the energy score. By decomposing the excess risk into approximation error, stochastic error, and Monte Carlo error, the authors establish convergence rates for deep neural network implementations. These rates are derived under the assumption that the target conditional generator exhibits compositional smoothness structure, providing insights into the theoretical guarantees of engression.
engressionenergy scoreexcess riskcompositional smoothnessmonte carlo error
VESTIGE: A Knowledge-Guided Masking Strategy for Corruption-Aware Fine-Tuning of Genomic Transformers, Validated on Ancient DNA Reconstruction
VESTIGE introduces a knowledge-guided masking strategy for fine-tuning genomic transformers, aligning the masking distribution with position-specific corruption profiles to improve reconstruction of degraded sequences. The method replaces standard uniform masking with an empirically measured per-position corruption profile, demonstrated on ancient DNA where cytosine deamination causes position-dependent C-to-T/G-to-A gradients. Evaluated on DNABERT-2 with mammoth CDS data, VESTIGE outperforms standard MLM by +4.18 to +10.35 percentage points across all tested terminal-zone widths, achieving >0.95 Pearson correlation even under artificially amplified damage, while maintaining biosecurity with 98.2% cleared reconstructions.
masked-language-modelcytosine deaminationposition-dependent corruptiongenomic transformersancient dna reconstruction
NMINE: Normalized Mutual Information Neural Estimation
The paper introduces NMINE, a neural estimator for normalized mutual information (NMI) in continuous multidimensional variables, addressing limitations of existing k-nearest-neighbor approaches. The method combines a MINE-based mutual information estimator using the Donsker-Varadhan representation with neural marginal entropy estimators derived from divergence learning against uniform references. On Gaussian data (1-8 dimensions), NMINE outperforms KSG-based NMI baselines in accuracy, demonstrating neural estimation's viability for normalized dependency measurement in high-dimensional continuous settings.
normalized mutual informationneural estimationdonsker-varadhanmarginal entropycontinuous variables
LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference
LightRot introduces a lightweight rotation scheme and hardware accelerator for efficient low-bit LLM inference, combining Grouped Local Rotation (GLR) and Outlier Direction Aligning (ODA) with a hierarchical Fast Hadamard Transform (FHT)-based rotation unit. The design targets energy-efficient 4-bit inference, achieving 27.4 TOPS/W in a 28nm CMOS process, outperforming prior work. Evaluated on LLaMA2-13B, LLaMA3-8B, and MT-Bench, LightRot demonstrates robust performance in conversational scenarios, advancing scalable low-bit inference for sustainable AI.
quantizationhardware acceleratorlow-bit inferencehadamard transformenergy efficiency
GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
GyRot introduces an algorithm-hardware co-design framework for efficient 4-bit LLM inference by synergizing rotation and fine-grained group quantization. The method combines Coarse Rotation, Fine Grouping (CoRFiG) and Harmonic-Aligned Permutation (HAP) to reconcile global rotation with local group scaling, while reformulating asymmetric quantization for hardware efficiency. Evaluated on LLaMA-family models, GyRot achieves state-of-the-art 4-bit accuracy with 3.4x speedup and 3.6x energy efficiency over baseline accelerators.
low-bit quantizationalgorithm-hardware co-designgroup quantizationinteger dequantizationllm acceleration
Recall Before You Rank: Similarity-Guided Top-$K$ Reuse for Efficient Long-Context Attention
ReTopK introduces a training-free method to accelerate dynamic Top-$K$ sparse attention by reusing historical retrieval decisions, addressing the linear-cost bottleneck of selector operations in long-context decoding. The approach caches historical query--support pairs, retrieves similar queries for new inputs, unions their supports with a recent window, and reranks the compact candidate set via exact scoring, with fallback to full-history Exact Top-$K$ when similarity is low. Evaluated on 16K--128K contexts, ReTopK achieves 3.07× speedup at 128K (K=512) with only 0.50% perplexity degradation versus Exact Top-$K$, while outperforming approximate methods on PG19 perplexity, NIAH, and LongBench benchmarks.
sparse attentionkv cachetop-k selectionlong-context decodingperplexity
Event-Structured Physics-Informed Neural Networks for Differentiable Critical Clearing Boundaries
The paper introduces Event-Structured Physics-Informed Neural Networks (ES-PINN) for differentiable critical clearing time (CCT) estimation in power systems. ES-PINN aligns representations with pre-fault, fault-on, and post-clearing dynamics, enforcing exact state chaining across event interfaces and enabling differentiable CCT boundary approximation. The method supports accurate boundary extraction, sensitivity analysis, and optional direct CCT prediction via a distilled readout. Experiments on IEEE 9-, 14-, and 30-bus systems demonstrate improved trajectory and stability-boundary accuracy over baselines, with validation through DAE simulations and multi-fault scenarios confirming computational efficiency.
physics-informed neural networkscritical clearing timetransient stabilitydifferentiable approximationpower systems
Tight Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection
This work establishes tight sample complexity bounds for Low-Rank Adaptation (LoRA) fine-tuning of pretrained models, closing existing theoretical gaps. Using local Rademacher complexity and Fano-type packing arguments, the authors prove matching upper and lower bounds of Θ̃(rd/n) for excess risk in rank-r LoRA adaptation when the target has rank ≤ r. They demonstrate a rank-selection dichotomy: constrained empirical risk minimizers suffer from over-ranking, while adaptive nuclear-norm estimators are unaffected by over-parameterization. Experiments on synthetic trace regression and real LoRA fine-tuning of DistilBERT and RoBERTa on SST-2 and MRPC validate the theoretical predictions, showing statistically significant U-shaped validation loss curves (p = 0.016).
low-rank adaptationrademacher complexityfano-type packingexcess risknuclear-norm
FedOGL: Combating Catastrophic Forgetting in Federated Open-World Multimodal Graph Learning
FedOGL addresses catastrophic forgetting in federated open-world multimodal graph learning by preserving semantic-structural memory. Client-side mechanisms include replay, task-start distillation, and graph-propagation memory projection onto a shared structure basis. Server-side, it maintains compact category prototypes for cross-client knowledge transfer without raw data exposure. Experiments show FedOGL reduces performance degradation from catastrophic forgetting by 42.67% compared to baselines, while maintaining or improving downstream task performance.
catastrophic forgettingfederated learningmultimodal graphsknowledge transfermemory preservation
Looped Transformers with Source-Centered State Evolution
The paper proposes Source-Centered State Evolution (SCSE), a method for looped Transformers that reconciles input conditioning with reference-preserving shared recurrence. SCSE maintains input dependence via a learned anchor and initial deviation, enforces exact anchor invariance through zero-deviation masking, and ensures the anchor is a fixed point by construction. Theoretical analysis shows the zero-deviation forcing bias is a tunable design parameter, which SCSE sets to zero for optimal invariance. Experiments on WikiText-2, WikiText-103, web-corpus pretraining, and LAMBADA demonstrate improved recurrent quality, with ablations attributing gains to the learned anchor and deviation recurrence.
looped transformerssource-centered state evolutionanchor invariancezero-deviation maskrecurrent computation
Evaluation Protocols and Cross-Subject Generalization in EEG Emotion Recognition
This study investigates evaluation protocols and cross-subject generalization in EEG emotion recognition, demonstrating that reported accuracy depends on the complete evaluation procedure rather than just the classifier. Using a dynamical graph convolutional neural network (DGCNN) pathway on SEED and SEED-IV datasets, the authors conducted subject-dependent, subject-disjoint, and cross-session evaluations. Results showed checkpoint selection improved mean window accuracy from 0.7855 to 0.8892 on SEED, while held-out participant accuracy was 0.5348 (SEED) and 0.3954 (SEED-IV). The study highlights inconsistencies in train-to-held-out-subject gaps and emphasizes the need for distinct reporting of subject-dependent, subject-disjoint, and cross-session results.
eegemotion recognitiondgcnncross-subject generalizationcheckpoint selection
Certifying when decision-time information justifies adaptive experimentation
The authors introduce Opportunity-aware Policy Authorization for Laboratories (OPAL), a framework for certifying when decision-time information justifies adaptive experimentation. OPAL employs a precommitted contract to ensure non-trivial adaptation, controlled target risk, and positive executed value after cost. They establish an impossibility boundary showing that source outcomes and unlabelled target covariates cannot uniformly support non-trivial authorization under unrestricted conditional outcome shift, and derive a target-calibrated recovery. Applied to an 11,265-compound Cell Painting partition, OPAL selected 595 compounds, captured 384 positive opportunities, and achieved strictly positive executed value with a 5.18% false-activation upper bound. Among six methods, only OPAL combined non-zero activation with risk control, demonstrating its utility for safe adaptive science.
adaptive experimentationpolicy authorizationtarget-calibrated recoveryconditional outcome shiftfalse-activation
Robust Wavelength Selection for Partial Least Squares Sugar Content Estimation Using Combinatorial Bayesian Optimization
The authors propose a combinatorial Bayesian optimization method for robust wavelength-region selection in near-infrared spectroscopy to estimate sugar content. The approach formulates selection as a binary black-box optimization, using a sparse quadratic surrogate model with Thompson sampling and solving the acquisition via quadratic unconstrained binary optimization (simulated/quantum annealing). Experiments demonstrate improved partial least squares regression accuracy (lower RMSE on validation sets) and more consistent wavelength selection versus genetic algorithms and standalone simulated annealing. Selected regions exhibit local stability under one-bit perturbations, indicating convergence to smoother error landscapes without overfitting.
bayesian optimizationwavelength selectionpartial least squaresquantum annealingnear-infrared spectroscopy
First-order Constrained Trilevel Optimization Over Distributed Networks for Robust Coreset Selection
The paper introduces Federated First-order Constrained Trilevel Optimization (F²CTO), the first distributed optimization framework for robust coreset selection in IoT edge networks. F²CTO formulates the problem as a trilevel optimization with level-wise constraints, integrating hierarchical composite value-function reformulation and distributed alternating projected gradient algorithms. The method achieves a non-asymptotic convergence rate of O(ε^(-3/2)) for ε-stationary points. Empirical evaluations on reliable continual learning validate its effectiveness and efficiency in addressing computational overhead and storage bottlenecks in distributed settings.
trilevel optimizationcoreset selectiondistributed learningfederated optimizationrobust optimization
Real-Time Hard Peak Age-of-Information Safety with No-Regret Learning
OCO-PAoI-Hard introduces a no-regret learning framework for hard real-time peak Age-of-Information (peak AoI) safety in adversarial environments, guaranteeing zero per-slot violations under one-step viability. The method reformulates hard peak-AoI deadlines as affine half-space constraints on resource-allocation vectors, enabling time-varying constrained online convex optimization over a polyhedral safe set. A strictly causal proposal-shield-update loop enforces feasibility via Euclidean projection, preserving no-regret behavior. Theoretical analysis establishes O(sqrt(T)) static/dynamic regret bounds, a matching minimax lower bound, and deadline-induced competitive ratio. Empirical evaluation on a four-sensor adversarial fluid-model trap channel demonstrates zero deadline violations across all seeds, outperforming baselines by 1.65%-64.0% in slot misses.
peak age-of-informationonline convex optimizationadversarial environmentseuclidean projectionminimax lower bound
Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning
The paper introduces Kalman-Guided Prompt Selection (KGPS), a dynamic state estimation method for adaptive prompt selection during RL finetuning of LLMs. KGPS models prompt difficulty as a latent success rate in logit space using a linear-Gaussian state-space model, with process noise scaled to policy updates, and employs a Kalman filter to maintain calibrated posteriors. This approach selects prompts by maximizing posterior-expected utility, favoring intermediate difficulty while revisiting uncertain prompts. Experiments on math, planning, and geometry benchmarks demonstrate KGPS improves final accuracy and reduces rollouts by 83% versus baselines, achieving SOTA in online prompt selection.
reinforcement learningkalman filterprompt selectionstate-space modelnon-stationary dynamics
Back from the Future: Key-Value Cache Management by Counter-Causal Surprise
We propose a novel KV-cache eviction strategy for LLMs that removes redundant tokens by leveraging counter-causal attention. Our method scores cache entries by running the model with a future-only attention mask, identifying tokens predictable from more recent context as eviction candidates. A single-layer approximation restricts computation to the final transformer layer, reducing refresh cycle time with minimal accuracy loss. Evaluations on open-source LLMs and benchmark datasets demonstrate competitive or superior performance to state-of-the-art methods. The approach requires no additional training and operates directly on existing cache contents.
kv-cachecounter-causal attentioneviction strategytransformer layermemory footprint
Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
Compliance2LoRA introduces a hypernetwork-based framework for on-demand safety alignment in large reasoning models (LRMs) across arbitrary policy subsets. The method employs a LoRA adapter generator that produces policy-compliant weights based on customizable safety policy inputs, which are then added to the LRM for compliant response generation. This approach eliminates the need for training separate LRMs per policy subset while avoiding the computational overhead of in-context learning. Results demonstrate effective policy adjustments without task performance degradation across various LRM sizes and evaluation datasets, showcasing the framework's adaptability and practicality.
hypernetworklora adaptersafety alignmentlarge reasoning modelspolicy compliance
Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs
Prox introduces a training-free framework for sparse SwiGLU FFNs in LLMs, addressing activation sparsity without quality degradation. The method leverages the SwiGLU intermediate state's magnitude ranking to construct a channel mask, enabling sparse execution across all three projections. Prox operates in two stages: Stage 1 uses input sparsity and quantized proxy weights to generate a shared mask, while Stage 2 computes selected channels exactly. Evaluated across ten LLMs from six families, Prox achieves up to 1.99× end-to-end decoding speedup at 70% FFN sparsity, outperforming training-free baselines and maintaining compatibility with quantization and sparse attention.
swigluactivation sparsitychannel maskquantized proxytraining-free
MUGEN: A Unified Framework for Efficient Motion Understanding and Generation
MUGEN introduces a unified motion--language framework that eliminates quantization and multi-step generation costs by using adaptive-length continuous latent slots for motion representation. The method employs a single autoencoder to compress motion into these slots, with depth-routed hidden states enabling variable transformer depth per slot and a calibrated head predicting joint latent distributions. At a decoding cost of K language-model steps, one draw, and one decoder pass, MUGEN outperforms baselines on HumanML3D (FID), achieves top CIDEr and BLEU@4 scores, and surpasses discrete-token state-of-the-art on SnapMoGen retrieval and alignment metrics.
motion-language understandingcontinuous latent slotsdepth-routed hidden statesadaptive-length autoencoderjoint latent distribution
Heterogeneous Ranking in Industrial-Scale Recommender Systems: A Case Study
The paper presents HA-MoE, a heterogeneity-adaptive multi-gated mixture-of-experts architecture for industrial-scale heterogeneous feed ranking in Google Discover. The method incorporates explicit heterogeneity context into gating networks and expert representations, enabling specialization without operational overhead, and introduces LENS for interpretable diagnostics. Evaluation using Dual-Level AUC (DL-AUC) shows consistent offline improvements on a large-scale dataset, with online A/B tests confirming gains in feed activity and exploration metrics.
heterogeneous rankingmixture-of-expertsmulti-task learningrecommender systemsindustrial-scale
Policy Gradient Steering: Interventions from Behavioral Objectives
The paper introduces Policy Gradient Steering (PGS), a reinforcement learning-based method for dynamically steering model behavior at inference time. PGS accumulates gradients from a temporary behavioral objective over rollouts or demonstrations to construct removable task vectors. Experiments in gridworld, chess puzzles, and competitive football demonstrate PGS's calibration, reversibility, composability of tactical objectives, and transferability of behavioral adaptations across opponents.
policy gradient steeringactivation steeringbehavioral objectivestask vectorsreinforcement learning
Recognition and Label-Free Adaptation Across Recording Sessions in Surface-EMG Gesture Decoding
The study introduces a montage-agnostic encoder for surface-EMG gesture decoding that maintains performance across recording sessions without recalibration. Trained on NinaPro DB6 data from ten intact subjects, the encoder achieves a macro-F1 of 0.688 across sessions, outperforming a per-user LDA pipeline (0.540) and two published source-only baselines. Label-free test-time adaptation via feature-statistic alignment improves performance for all subjects, recovering accuracy comparable to a single labeled calibration repetition, while batch-normalization re-estimation fails entirely.
surface-emglabel-free adaptationcross-session transferfeature-statistic alignmentbatch-normalization
A Montage-Agnostic Encoder for Calibration-Light Cross-User Gesture Recognition from Surface Electromyography
The paper introduces a montage-agnostic encoder for cross-user gesture recognition from surface electromyography (sEMG), addressing the challenge of user-specific calibration. The encoder processes electrodes with shared weights, using physical coordinates instead of indices, enabling flexible channel counts without montage-specific parameters. Evaluated across three databases (DB1, DB2, DB5), it outperforms per-user Hudgins and linear-discriminant classifiers by 0.234 macro-F1 on DB1 and 0.108 on DB2, but underperforms on DB5. Ablation studies show each of its three key components contributes over half of its 3-shot macro-F1 performance. Cross-user training stability requires at least nine subjects, with performance gains plateauing beyond 39 subjects. Self-supervised pretraining offered no additional benefits over supervised training.
surface electromyographymontage-agnosticcross-user recognitionmacro-f1ablation study
Strategies for Milestone-driven Start-ups in Multi-activity Settings
This work presents a foundational stochastic control model for milestone-driven start-up strategies, explicitly characterizing optimal policies with multiple activity choices. The model captures start-up state evolution via a diffusion process, where activities influence drift, variance, and cost, determining success or failure upon reaching fixed boundaries. The optimal policy leverages an efficient frontier curve ordering controls by riskiness (drift-to-volatility ratio) and cost-effectiveness (drift-to-cost ratio), exhibiting qualitatively different structures based on model parameters. This is the first study analyzing stochastic control models admitting diverse efficient frontier curve types, offering start-ups intuitive activity evaluation measures and scenario-dependent strategic insights.
stochastic controldiffusion processefficient frontierdrift-to-volatility ratiodrift-to-cost ratio
Memory Efficient Tabular Foundation Models
This paper addresses memory efficiency in Tabular Foundation Models, focusing on practical deployment constraints. The authors investigate model compression techniques to reduce memory requirements while maintaining performance levels comparable to uncompressed models. Their approach achieves memory reductions of up to 7.6×, corresponding to an 87% decrease in deployment requirements. The study provides empirical evidence that compression methods can be effectively applied to TabPFN and similar models without significant performance degradation, offering practical insights for efficient deployment in real-world tabular machine learning applications.
tabular foundation modelsmodel compressionmemory efficiencytabpfndeployment requirements
Subtract or Replay? Exact Deletion from Language-Model Memory
The paper demonstrates that exact deletion in language models depends on memory representation, distinguishing between algebraic decrement for addressable records and replay for entangled writes. Experiments replace Gemma 3's global-attention layers with support-vector memory, achieving median KL divergence of $5.4\times10^{-15}$ over 31 token deletions at 1B parameters, with a 2.0% perplexity increase. At 4B and 12B, utility costs rise to 11.2% and 44.3%. A 48B Kimi Linear hybrid shows suffix-dependence under delta rules, while checkpointed replay achieves exact deletion for 18,842-token contexts, matching never-ingested logits and recurrent states.
exact deletionsupport-vector memoryalgebraic decrementrecurrent statesperplexity
HOMER: Huber-of-Means for Efficient and Robust Estimation in Hilbert Spaces
The paper introduces HOMER (Huber-of-Means for Efficient and Robust Estimation), a robust estimator that aggregates block means via a radial Huber center in Hilbert spaces. Two variants are proposed: canonical HOMER, which recovers the sample mean within its quadratic region, and pseudo-HOMER, which asymptotically approaches the sample mean. Theoretical analysis establishes a Hilbert-space majority theorem and deviation bounds under finite second moments, with mean inference at parametric rates under third moments. Empirical results show robustness against minority block contamination while maintaining near-optimal efficiency on Gaussian data, though sandwich intervals undercover for skewed functional data.
robust estimationhuber losshilbert spacesmedian-of-meanssandwich covariance
When Does Explicit View Routing Work? A Controlled Study of Multi-View Graph-Text Alignment
The study investigates explicit view routing in multi-view graph-text alignment, focusing on semantic aspect separation via multiple heads. Using a controlled MV-GTA setup with deterministic text segments, isolated encoders, and view-specific graph heads, the authors evaluate retrieval performance on BBBP and BACE datasets. Correct routing improves label and property nDCG by 0.305 to 0.685 over deranged training, with the expected graph head outperforming the best wrong head by 0.303 to 0.453. Property paraphrase augmentation enhances unseen-template nDCG by 0.140 and 0.147. Results highlight explicit, externally grounded label and property routing but do not support free-form routing or consistent three-view specialization.
multi-view graph-text alignmentexplicit view routingsemantic aspect separationndcgproperty paraphrase augmentation
Latent-Kernel Discrete Flow Maps for Few-Step Generation
The paper introduces Latent-Kernel Discrete Flow Maps (LKF), a novel flow-map kernel for few-step text generation that addresses the challenge of modeling token correlations in discrete diffusion models. LKF employs a mixture of M factorized components tied by a shared latent variable, enabling correlated updates while maintaining computational efficiency. Experiments on LM1B and WikiText-103 demonstrate that LKF improves generative perplexity by 2.1x-3.3x over baselines without sacrificing diversity, outperforming distilled and rectified few-step samplers at M=8.
discrete diffusionflow-matchinglatent-kernelfew-step generationtoken correlation
Sparsity Induced Identifiability in Matrix Tri-Factorisation
The paper establishes the first rigorous theoretical guarantees for sparsity-induced identifiability in general real-valued matrix tri-factorization, addressing a gap in prior work focused on two-factor models. By introducing a novel decomposition strategy that transforms the problem into two coupled auxiliary factorizations while preserving structural information, the authors derive recovery guarantees and structural consistency results. These characterize how coefficient sparsity affects recovery conditions, convergence, spectral approximation error, and structure preservation, with Monte Carlo experiments validating theoretical predictions.
matrix tri-factorizationsparsity constraintsidentifiabilityrecovery guaranteesspectral approximation
A Lightweight Foundation Model for Collider Physics with Multi-Domain Adaptation
The authors propose NEXUS, a lightweight foundation model (3M parameters) for collider physics using unsupervised pre-training on Large Hadron Collider charged particle track data. The fully connected autoencoder architecture demonstrates improved accuracy on downstream tasks (kinematic regression, event classification) with minimal labeled data compared to scratch-trained equivalents. Latent space analysis reveals cross-domain applicability to gravitational waves, flood forecasting, and neural activity, while computational efficiency enables edge deployment versus transformer-based alternatives.
foundation modelautoencodercollider physicsunsupervised pre-traininglatent space
Latent States in Neural Networks: Recovering the Temporal Structure of Drifting Data from Model Weights
The study demonstrates that latent temporal regimes in drifting data streams are recoverable from model weights, using a hidden Markov model (HMM) applied to chronologically ordered weight trajectories. Experiments on multimodal misinformation detection (Fakeddit dataset) and sentiment analysis (Yelp dataset) show classifiers generalize better within recovered states than across state boundaries, even after controlling for temporal proximity. The states correlate more with class distribution shifts than weight-space geometry, and within-state transfer advantage persists after residualizing class divergence and lag. Effects replicate across tasks but are attenuated on Yelp due to its more stable label distribution.
hidden markov modeltemporal regimesmodel weightsclass divergencetransfer advantage
Schreier-Coset Graph Rewiring
The authors propose Schreier-Coset Graph Rewiring (SCGR), a group-theoretic method to mitigate over-squashing in graph neural networks (GNNs) by augmenting input graphs with Schreier-Coset graphs derived from special linear groups. SCGR provides theoretical guarantees, including spectral gap and bounded effective resistance, creating low-resistance bypasses for long-range communication. Empirical results show SCGR reduces effective resistance by 5-40% across tasks while maintaining competitive accuracy, addressing structural bottlenecks without excessive edge growth or property distortion.
graph neural networksschreier-coset graphover-squashingeffective resistancespectral gap
OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval
OneShot introduces a holistic retrieval framework that aligns index learning with ranking objectives, addressing the traditional misalignment between ranking accuracy and indexing efficiency in large-scale recommendation systems. The method employs an end-to-end, in-model index learning approach, scaling interaction modeling beyond the dot-product bottleneck through neural scoring. Deployed in Instagram's short-video recommendation system, OneShot achieves a 20% recall gain at operational ranking volume and a 10x efficiency improvement at equivalent recall levels, significantly enhancing user engagement metrics.
retrieval frameworkindex learningneural scoringrecommendation systemsinteraction modeling
FADEx: Feature Attribution and Distortion-based Explanation of Dimensionality Reduction
FADEx introduces a local per-instance feature attribution method for explaining dimensionality reduction (DR) techniques, addressing limitations of existing approaches like multiple feature attributions and method specificity. The method employs local linear approximation via first-order Taylor expansion and Singular Value Decomposition, computing local linear models through weighted least squares without requiring out-of-sample data mapping. Evaluations demonstrate FADEx's robustness and versatility, outperforming existing methods in providing explanations and distortion analysis for DR behavior.
dimensionality reductionfeature attributionlocal linear approximationsingular value decompositiondistortion analysis
Neural Network-Assisted CLEAN for Channel Modeling in Low-SNR Regimes
The paper proposes Neural Network-Assisted CLEAN (NN-CLEAN), a hybrid framework combining multi-head residual networks with iterative CLEAN for multipath parameter estimation in low-SNR wireless channels. By replacing grid search with neural network predictions while maintaining physical residual subtraction, the method achieves 96% accuracy at 5 dB SNR, matching Grid-Search CLEAN but with significantly reduced computational complexity. Monte Carlo simulations show NN-CLEAN outperforms subspace methods and standalone neural networks while maintaining near-flat runtime scaling with batch size, enabling real-time MIMO channel estimation.
multipath parameter estimationclean algorithmlow-snr channel modelingmulti-head residual networkmimo systems
📰 Industry Media
No new items today.
Generated automatically at 2026-08-01 20:43 UTC. Summaries and keywords are produced by an LLM and may contain inaccuracies — always consult the original article.
