Daily Digest — 2026-09-11
259 items · 10 research labs, 238 arxiv papers, 11 industry media
🏛️ Research Labs (10)
How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules
César de la Fuente's lab employs OpenAI's Codex and ChatGPT alongside custom deep-learning models to accelerate antimicrobial discovery by analyzing genomic and proteomic datasets. The approach frames biology as an information system, using AI to identify functional peptide patterns in vast databases, reducing candidate screening from years to hours. The lab validates AI predictions experimentally, leveraging cross-disciplinary collaboration to bridge computational and wet-lab workflows, though clinical translation remains a multi-stage challenge.
antimicrobial resistancegenomic databasesin-context learningtransdisciplinary researchpeptide prediction
Now everyone can put data to work
OpenAI introduces the Data agent in ChatGPT Work, enabling business users to derive insights from organizational data via natural language queries without requiring SQL expertise. The agent integrates with enterprise data sources (e.g., Snowflake, Databricks) and semantic layers (e.g., dbt, Databricks Genie Ontology), enforcing existing row/column-level permissions. It supports interactive dashboard generation, metric diagnosis, and action recommendations, with adoption by OpenAI's product teams (67% of GTM) and alpha partners like NTT Data. Administrators control data-source plugins and access via Workspace settings.
semantic layersrow-level permissionsnatural language queriesinteractive dashboardsenterprise data integration
Introducing ChatGPT for Financial Services
OpenAI introduces ChatGPT for Financial Services, a domain-specific variant integrating GPT-6 Astra with curated financial datasets (Daloopa, PitchBook, LSEG News) to support investment banking workflows. The system addresses data retrieval challenges via optimized MCP connectors (S&P Global, FactSet) and offers granular citation tracing, artifact generation (Excel/PPT templates), and enterprise-grade security (SAML SSO, SCIM). Early partnerships with Morgan Stanley and Evercore informed prioritization of P&L normalization and equity research tasks. The platform achieves 50+ connector integrations (Datasite, Preqin) and enables role-based access controls for MNPI protection without model training on client data.
gpt-6 astramcp connectorssaml ssogranular citationsp&l normalization
Expanding AI access and cyber defense for federal, state, local, and tribal governments
OpenAI and the U.S. General Services Administration (GSA) announced a multi-year agreement providing free access (normally $15/user/month) and 50% usage discounts for federal, state, local, and tribal governments to OpenAI's tools, including GPT-6 Astra and Daybreak Blue for cyber defense. The initiative expands eligibility to 23 million public-sector workers, building on existing access for 1 million employees. Demonstrated applications include accelerated vulnerability research (Daybreak Access), 92% productivity gains in CDC literature reviews (30-minute reports vs. days/months), and 15-minute tax form digitization (Georgia DoR). The 27-month agreement (2026-2028) includes training, FinOps guidance, and secure deployment (no data retention for model training).
gpt-6 astradaybreak bluevulnerability researchfinopscyber defense
Build more natural voice experiences with GPT‑Live‑1 in the API
OpenAI introduces GPT-Live-1, a full-duplex voice model for API deployment, integrating speech recognition, reasoning, and synthesis in a single architecture. The model improves interruption handling by 80% over turn-based systems, leveraging joint audio processing to reduce latency. It supports delegation to backend models (e.g., GPT-6 Astra) and offers customizable tone, pace, and telephony integration. Evaluations show a 30-point improvement on Full Duplex Bench versus GPT-Realtime-2.1 and top performance on Tau3 for end-to-end voice-agent tasks. Priced at $0.05/min, it targets scalable voice workflows with enterprise support via OpenAI Presence.
full-duplexinterruption handlingtau3 benchmarktelephony integrationdelegation
Introducing the Agents API
OpenAI introduces the Agents API, a production-ready framework for deploying long-running AI agents with managed context, tool efficiency, and subagent coordination. The API leverages OpenAI's optimized harness infrastructure, offering versioned model capabilities, automatic context compaction, parallel tool execution, and multi-agent delegation. Developers can deploy agents in managed sandboxes or custom environments, with ecosystem partnerships providing varied compute/storage options. The system supports GPT-3.5/4-level models, reduces token overhead via tool search, and enables parallel subagent workflows without custom orchestration. No additional fees apply beyond standard token/tool usage costs during the public beta.
agents apicontext compactiontool searchmulti-agent delegationmanaged sandbox
The AI policy window is open. We need to act.
OpenAI advocates for urgent AI policy reforms, proposing mandatory national safety regulations, state-level legislation support (SB 813, AB 1405, SB 1119, AB 1864), and global standards to manage risks from advanced AI systems. The organization emphasizes monitoring, alignment, and security safeguards, including trajectory monitoring for models like Astra and mandatory alignment evaluations. OpenAI warns of AI-accelerated research and recursive self-improvement risks, urging policymakers to establish safety bars and human oversight before capabilities outpace governance. The proposal includes federal frameworks for independent assessments, cybersecurity, and incident reporting.
alignment-evaluation gaterecursive self-improvementtrajectory monitoringcapability-based regulationfrontier ai standards
GPT-6 Astra: The next generation in intelligence for work
OpenAI introduces GPT-6 Astra, a multimodal foundation model optimized for professional workflows with state-of-the-art performance in coding (Codex integration), enterprise applications (browser/desktop interaction), and cybersecurity (Critical capability threshold). The 1.2T-parameter model demonstrates 25× latency reduction in memory allocation tasks and 89% fewer unintended outcomes versus GPT-5.6 Sol in safety benchmarks. Enterprise features include token-efficient operation ($10/M input tokens), Zero Data Retention API endpoints, and admin controls for application whitelisting. Early adopters report 30% memory efficiency gains in GPU optimization and automated video production workflows.
multimodal foundation modelmemory allocationzero data retentionapplication whitelistingtoken-efficient
Rebuilding AUTOMATIC1111 with Gradio Workflow
Workflow1111 rebuilds AUTOMATIC1111's stable-diffusion-webui using Gradio's gr.Workflow, implementing 11 media pipelines via 73 nodes. The system integrates state-of-the-art models for text-to-image, image-to-image, hi-resolution fix, prompt-matrix grids, VLM interrogation, detection-to-inpaint masks, ControlNet-style annotators, background removal, PNG Info storage, and image-to-video. Nodes include Python functions, InferenceClient models, Gradio Spaces, and Hub datasets. The workflow supports parallel execution, local processing, and REST API endpoints, enabling GPU-free operation via Hugging Face's Inference Providers. The canvas-driven approach allows modular, browser-based workflows comparable to ComfyUI.
gr.workflowmedia pipelinesinferenceclientgradio spacesrest endpoints
3 ways to prep for your next big race with Search
Google Search integrates AI features to assist runners in race preparation through personalized training plans, custom playlists, and gear recommendations. AI Mode enables users to generate tailored training schedules, incorporating cross-training and route suggestions based on location and fitness level. Integration with YouTube Music allows for the creation of motivational playlists to enhance mental endurance during runs. Leveraging Google’s Shopping Graph with over 60 billion product listings, Search provides precise gear recommendations, including price and availability filters, to optimize race-day performance.
ai modetraining planshopping graphcustom playlistscross-training
📜 arXiv Papers (238)
Show-Harness: Just a VLM Agent Can Play Robots
Show-Harness introduces a semantic interface enabling vision-language models (VLMs) to control robots by mapping discrete action units to physical actions via deterministic interpreters. The method supports zero-shot control with closed-source VLMs and few-GPU-hour fine-tuning for open-source models, while GUMI extends the interface to GUI-based teleoperation. Experiments demonstrate robust generalization across tasks, embodiments, and environments, outperforming existing agentic and vision-language-action paradigms, suggesting foundation VLMs can achieve embodied capability without additional capacity or pretraining.
vision-language modelsrobot controlsemantic interfacezero-shot learningembodied ai
IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier
The paper introduces IB2, a protocol for measuring enterprise AI systems by serving route rather than model identifier, addressing measurement errors in current benchmarks. The protocol includes a gold-blind capability-binding preflight, reliability-inclusive first-pass scoring, and score-blind adjudication. Results from 11 systems show measurable capability availability, non-uniform discrimination across tasks, and significant score variations due to serving-arm choices (e.g., precision shift from 77.38 to 82.54). Reliability inclusion alters conclusions by accounting for failed responses.
serving routecapability-bindingreliability-inclusiveadjudicationmeasurement error
Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization
SG-JEPA extends LeWorldModel by action-conditioned physics parameterization and autoregressive latent rollout, improving zero-shot generalization across varying gravitational fields. The method jointly trains an encoder and predictor, with analysis showing encoder-learned features drive performance gains over predictor dynamics. Results show 2× lower open-loop prediction error on 2D tasks and 2.5× higher control success on 3D robotic tasks versus DINO-WM, attributed to latent dynamics consistency via multi-step loss backpropagation.
joint-embedding predictive architecturelatent dynamicsautoregressive rolloutzero-shot generalizationdiffusion policies
JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition
JarvisGUI introduces a dynamic benchmark for evaluating GUI agents on cross-device workflows across Android, Windows, and Ubuntu, addressing the gap in existing single-device benchmarks. The framework formulates tasks as input-output transformations under a lightweight type system, enabling automatic composition of multi-step workflows and unified evaluation. Results show state-of-the-art open-source GUI agents struggle with state-transfer awareness, cross-platform reasoning, and long-horizon dependencies, revealing capability gaps undetected by current benchmarks.
cross-device workflowsgui agentsdynamic task compositioninput-output transformationsheterogeneous environments
ConvMem: Convolutional Memory for Long-Context Reasoning
ConvMem introduces a training-free, parallelizable framework for long-context reasoning by reformulating it as hierarchical convolution, treating an LLM as a convolutional kernel with configurable strides and skip connections. The method employs multi-kernel convolution to disentangle semantic channels, reducing error accumulation and enabling parallel processing across text segments. Evaluated on RULER-HotpotQA and RULER-2WikiMultiHopQA, ConvMem outperforms training-free baselines and avoids overfitting risks associated with RL-trained models on out-of-distribution tasks.
long-context reasoningconvolutional memorymulti-kernel convolutiontraining-free frameworkhierarchical summarization
Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs
FOM-UL introduces a layer-selective unlearning framework for LLMs that targets transformer layers via a forget-to-retain significance score, optimizing updates to minimize residual memorization while preserving utility. The method identifies layers with high forget-set influence and low retain-set sensitivity, concentrating updates to improve robustness under post-training quantization (8-/4-bit) and adversarial prompts. Evaluations on TOFU, KnowUnDo, and MUSE benchmarks show FOM-UL outperforms GA, NPO, KLD, SURE, ReLearn, and LUNAR baselines in memorization suppression and utility retention, particularly under quantization where diffuse updates are prone to erasure.
machine unlearningtransformer layerspost-training quantizationforget-to-retain scoreresidual memorization
Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support
The study explores using GPT-4 and a knowledge graph algorithm (KGA) to screen emergency department (ED) revisit diagnosis pairs for quality assurance. Clinicians and GPT-4 assessed 99 diagnosis pairs, with GPT-4 flagging 94% for follow-up (4.4-13.3× more than clinicians). The KGA achieved 83-100% positive predictive value aligning with clinician judgments. Results suggest LLM-augmented screening could expand revisit review scope without increasing workload, though prompt engineering and validation are needed.
emergency department revisitquality assurancelarge language modelknowledge graph algorithmpositive predictive value
Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs
Fortunate Recall (FR) introduces an ontology-driven memory lifecycle management system for LLMs, classifying personal facts into an 11-category behavioral ontology with category-specific retention policies (differential decay, supersession, validity). FR-Bank, its implementation, achieves 76.9% accuracy on LifecycleBench (516 questions) and 75.2% on LongMemEval-S, outperforming Mem0 (61%) and MemoryOS (70.5%). The ontology halves confabulation (12.0% vs 24.2%) while generic lifecycle metadata maintains correctness. End-to-end, FR-Bank reduces Mem0's confabulation from 45.1% to 22.4% and improves correct answers (31.2% vs 18.6%), with consistent gains on BEAM (46.8% vs 32.9%).
memory lifecycle managementbehavioral ontologytemporal disambiguationconfabulation reductionretrieval routing
Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization
The study evaluates two paradigms for Vision-Language Model (VLM) guidance in content moderation: instruction-driven (policy precepts) and example-driven (prior precedents), using ModerationBench, a benchmark of 4,000 annotated Bluesky posts. Results show foundation models nearly triple Bluesky's moderation system performance (F1 0.60 vs. 0.22), with both paradigms achieving comparable effectiveness, demonstrating potential for scalable policy operationalization.
vision-language modelcontent moderationfoundation modelsinstruction-drivenexample-driven
MOONWALK: Mediating Operations with Intent-Evidence-Action Alignment Across Junior-Supervisor Review Workflows in Animation/VFX Pre-Production
The paper introduces MOONWALK, a structured workflow framework for animation/VFX pre-production review that aligns creative intent, grounded evidence, and executable actions. The system comprises a shared intent record, reference anchoring, structured comparison tools, and supervisor-authorized action planning, with AI handling administrative coordination while preserving human creative control. A studio study comparing MOONWALK to chat-only interfaces showed improved intent alignment (traceability +28%), decision accountability, and task executability, though aesthetic authority remained with practitioners. The results demonstrate the superiority of structured workflows over unstructured conversational AI in professional creative settings.
intent-evidence-action alignmentpre-production reviewstructured workflowdecision traceabilitycreative coordination
PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving
PACE introduces a perceived-latency-aware framework for retrieval-augmented dialogue serving, optimizing Perceived Time-to-First-Response (PTFR) under quality/cost constraints. It combines a load-adaptive cascading router, joint path-filler controller, and volatility-aware cache admission to jointly manage response composition and waiting-window content. Evaluated on 75k CarQA requests, PACE reduces P95 PTFR to 0.41s (2.4× faster than RAG at high load), cuts API calls by 94% via filler control, and eliminates stale answers with cache admission. A gating rule ensures no performance degradation versus baselines.
perceived latencycascading routingvolatility-aware cachingretrieval-augmented generationqoe optimization
OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis
OmniMed-FL introduces a multimodal federated learning framework for clinical diagnosis, addressing HIPAA/GDPR constraints by fusing medical imaging and synthetic patient records without centralized data aggregation. The framework evaluates eight fusion strategies, three initializations, and four missing-text imputation rules under non-IID Dirichlet partitioning across 3 to 20 hospital clients. Experiments on a proxy corpus of 3,000 chest radiographs and synthetic notes show multimodal fusion outperforms unimodal approaches, achieving macro-F1 scores of 0.956 (synthetic corpus) and 0.906 (radiograph corpus), with FedProx yielding the best federated performance (0.737±0.085). Label skew impacts performance more severely than client scaling, with bidirectional communication volume scaling linearly to 183.5 GiB at K=20.
multimodal federated learningnon-iid partitioningdirichlet skewmacro-f1synthetic notes
Cyber-Financial Contagion: Modeling the Propagation of an AI Vendor Compromise Through the Banking System
The paper introduces CFC-Prop, a stochastic epidemic-and-clearing model for analyzing cyber-financial contagion in banking systems dependent on shared AI vendors. The four-layer heterogeneous network couples AI vendors, financial institutions, interbank exposures, and customer accounts, simulating compromise propagation. On a synthetic dataset (60 vendors, 220 banks), CFC-Prop reproduces heavy-tailed loss distributions and patch-latency dependencies. The proposed early-warning model, CFC-GNN, achieves AUROC 0.82 and AUPRC 0.60 using vendor telemetry and graph structure. Results highlight cyber concentration as a financial-stability risk.
cyber-financial contagionstochastic epidemic modelheterogeneous networkinterbank exposuresearly-warning model
Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs
VIP-Router introduces sample-adaptive strategy routing for vision token pruning in multimodal large language models (MLLMs), addressing the limitation of fixed pruning strategies by dynamically selecting the optimal approach per input. The method conditions on low-cost visual and textual features to route between candidate pruning strategies or full-token inference, requiring only 0.017% additional trainable parameters. On VTC-Bench Group A, VIP-Router outperforms fixed-strategy baselines by 26.9% in average accuracy and 22.0% in utility-adjusted performance, demonstrating robustness across MLLM backbones and unseen benchmarks.
vision token pruningmultimodal large language modelssample-adaptive routinginference efficiencyadaptive computation
From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning
The paper presents a framework combining symbolic perception with logical deduction to enhance large language models' (LLMs) geometric reasoning without multimodal inputs. A Geometric Vision Parser converts diagrams into symbolic representations, while a Symbolic Solver performs formal deductions, reducing hallucinations and improving interpretability. Evaluated on a novel benchmark from 2025 Chinese Zhongkao examinations, the approach matches Gemini 2.5 Pro's performance while providing clearer, human-like solutions.
geometric reasoningsymbolic parsinglarge language modelsformal deductionmultimodal learning
TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards
TRACE introduces a reinforcement learning framework for causal diagnostic reasoning in domains with ambiguous ground truth, using synthesized rewards from controlled simulations. The method samples interventions, injects them into a simulator to generate observations, and trains agents to identify root causes and affected segments via Python/SQL queries. Evaluated on a digital-advertising environment with 12 root causes, TRACE achieves 0.757 FullAttr@1 with Qwen3.5-35B-A3B after supervised fine-tuning and RL, outperforming prompted baselines (including Claude Opus 5 at 0.686) while reducing tool calls. Results suggest simulation-based rewards can surpass model scale for diagnostic tasks.
reinforcement learningcausal reasoningsynthesized rewardsdiagnostic simulationtool-augmented lm
One Loop, Two Gains: Can Active Learning win the Lottery for Free?
The paper introduces Improve & Prune (I&P), a method that integrates iterative magnitude pruning into active learning retraining cycles to discover sparse subnetworks (winning tickets) without additional computational cost. I&P leverages the non-stationary data regime of active learning to perform pruning during each retraining phase, yielding deployable models at up to 95% sparsity while matching dense model accuracy. Experiments across acquisition functions, architectures (e.g., ResNet), and image datasets (e.g., CIFAR) show I&P addresses bottlenecks in per-round retraining and acquisition scoring, enabling practical deep active learning on large models and unlabeled pools.
lottery ticket hypothesisiterative magnitude pruningactive learningsparse subnetworksnon-stationary data
RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding
RiLM introduces a parameter-efficient language modeling framework that replaces the output matrix with geodesic decoding on Riemannian manifolds, reducing model size by eliminating the output layer. The method represents context as manifold trajectories and computes next-token probabilities via squared geodesic distance between states and vocabulary embeddings, sharing input-output embeddings. Evaluated on WikiText-2, HypRiLM (hyperbolic variant) achieves 54.2 validation perplexity, outperforming flat RiLM (87.6) and tied baselines (113-147), with consistent gains on Penn Treebank and larger vocabularies. The work also analyzes boundary collapse in hyperbolic recurrence and proposes Möbius stabilization for training stability.
riemannian manifoldgeodesic decodinghyperbolic recurrenceparameter-efficientperplexity
Learning Intrusion Response Strategies for OT Systems
The paper proposes a Partially Observable Markov Decision Process (POMDP) framework to model intrusion response in Operational Technology (OT) systems, incorporating partial observability based on traffic measurements. The approach employs Proximal Policy Optimization (PPO) to learn tractable response strategies, evaluated on an emulated OT system against multiple MITRE ATT&CK techniques. Results demonstrate effectiveness in mitigating specified attack types within the studied use case.
operational technologypomdpproximal policy optimizationmitre att&ckintrusion response
GANDR: Claim Auditing for Verifiable Legal Answer Generation
The paper introduces GANDR (Grounded ANswer DRafter), a two-agent system for verifiable legal answer generation, where a Drafter produces structured legal reasoning and a Critic audits each claim against cited sources, emitting per-claim audit traces. The method enforces a strict correctness criterion requiring citations to resolve to retrieved passages. On a 185-item legal benchmark, GANDR achieves 70.8% strict accuracy, outperforming baselines by 11.3 points (p<0.01). Ablations show the protocol-anchored commit rule contributes 22.7 points to accuracy. The Critic detects under-supported claims at F1 0.84 against human annotators.
grounded generationlegal reasoningclaim auditingverifiable aitwo-agent system
What Should an Agent Forget? Separating What Is Stored from What Is Used
RD-Forget introduces a training-free framework for language agents that separates persistent storage from query-conditioned memory access. The method employs a frozen language-model curator to extract and group relevant evidence into semantic slots, with same-slot replacement suppressing obsolete facts and intent-aware retrieval reactivating historical context. A rate-distortion objective optimizes memory views under budget constraints. Evaluations across conversational memory, knowledge updating, and long-context reasoning show that query-relevance and obsolete-fact suppression are critical, with slot grouping and relation preservation providing complementary benefits. Configurations lacking these mechanisms exhibit significant performance deficits.
query-conditioned memorysemantic slotsrate-distortionfrozen language-modelmulti-hop reasoning
DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs
DiSCo introduces a distribution-first framework to evaluate cultural preference bias in LLMs by isolating default cultural priors and testing steerability via a four-level context gradient (C0--C3). Using DiSCo-Bench (304 items across 12 cultures from BLEnD), the study evaluates six instruction-tuned LLMs, revealing heavily concentrated default priors favoring UK and US cultures (35% of selections). Prompt-based steering exacerbates selection gaps between high- and low-resource cultures, and explicit cultural facts minimally disrupt distributions, indicating prompt-based personalisation alone cannot resolve bias.
cultural preference biasdistribution-first evaluationcontext gradientinstruction-tuned llmssteerability
A-JIT: Agentic Just-In-Time Software Construction
The paper introduces Agentic Just-In-Time Software Construction (A-JIT), a paradigm shift from static software binaries to dynamic, continuously evolving systems. A-JIT integrates an AI agent within applications to observe runtime behavior and specialize software logic, workflows, and tool interfaces in real-time, akin to JIT compilation. This approach enables on-the-fly implementation synthesis, dynamic capability generation, and user-driven adaptation, demonstrated through trace-driven human-AI co-construction.
agentic softwarejust-in-time compilationruntime adaptationdynamic synthesishuman-ai co-construction
LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented Generation
LiteRAG introduces a cost-efficient graph-based retrieval method for multi-hop question answering, replacing LLM-controlled retrieval with query-conditioned algorithmic exploration and reasoning-chain context construction. The method employs query-adaptive thresholding and community-aware hub penalization to optimize token efficiency. On DistComp, LiteRAG achieves the highest quality (0.798) while reducing latency by 100× and cost by 99% compared to GraphRAG Global and DRIFT. On UltraDomain, it matches LinearRAG's quality while using 14× fewer tokens. Ablation studies confirm the efficiency gains stem from its adaptive thresholding and hub penalization mechanisms.
retrieval-augmented generationmulti-hop question answeringgraph-based retrievalquery-conditioned explorationtoken efficiency
Hierarchical and Permutation-Invariant Feature Transformation Learning via Policy-Guided Embedding Search
We propose a hierarchical, permutation-invariant feature transformation framework addressing three limitations in generative approaches: hierarchical feature-operation-abstraction relationships, order-sensitive embeddings, and gradient-based search inefficiency. Our method combines a self-attention pooling mechanism for consistent embeddings across semantically equivalent structures with a policy-guided multi-objective reinforcement learning strategy optimizing both predictive accuracy and transformation efficiency. Extensive experiments on diverse tabular benchmarks demonstrate superior performance against strong baselines. Code and data are publicly available.
feature transformationself-attention poolingmulti-objective reinforcement learningpermutation-invariant embeddingstabular benchmarks
Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection
FGPO (Full-Group Policy Optimization) introduces exact policy optimization for genomic tool selection by enumerating and scoring all tool subsets, eliminating the need for sampling-based approximations. Unlike GRPO (Group Policy Optimization), which suffers from reward signal degradation due to resampling, FGPO precomputes rewards for all question-subset pairs, removing frozen-reasoner calls from training. Evaluated across five frozen reasoners and three genomic benchmarks, FGPO outperforms GRPO by an average of 6.75 points (up to 14.20), reduces invoked tools per question from 2.36 to 1.40 on GenomeQA, and requires 2.4x fewer frozen-reasoner reward evaluations.
policy optimizationgenomic reasoningfrozen reasonertool selectionreward precomputation
Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries?
EvidenceNet introduces a runtime assurance layer to verify network-wide intent fulfillment across administrative domains by coordinating AI agents with scoped authority. The system employs a broker to collect post-change observations, an admission gate to validate evidence scope and freshness, and a verifier agent for content assessment. Experiments on live routing networks demonstrate that EvidenceNet reliably confirms successful outcomes using state checks, rejecting cases with incorrect sources, substitutions, or stale data, unlike configuration-action records alone.
network automationruntime assurancecompletion contractverifier agentadministrative domains
Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning
The paper introduces a contrastive modeling framework for improving reasoning path alignment in multimodal in-context learning (ICL) by reformulating demonstrations to explicitly contrast suboptimal and better responses with associated reasoning paths. The method incorporates a response-conditioned retrieval mechanism for relevant demonstration selection and a lightweight alignment controller to assess response quality. Evaluations across three multimodal tasks demonstrate consistent performance gains, particularly in visual question answering (VQA), highlighting the framework's effectiveness beyond surface-level imitation.
multimodal in-context learningcontrastive modelingreasoning path alignmentresponse-conditioned retrievalvisual question answering
Kernel-Managed Shared Memory for System-Wide Personalization
We introduce kernel-managed shared memory, a system-level abstraction enabling multi-agent personalization through centralized memory management. The kernel governs retrieval, privacy enforcement, and prompt injection while specialized agents write structured, tagged memories. Evaluated on AIOS across GPT-4o, Llama-3.1:8B, and Qwen-2.5:7B (1,800 trials), this approach improves personalization scores by 2.4-4.0 points versus unmanaged memory (Mem0) and matches retrieval-augmented injection performance while reducing latency by 15-61%. Results demonstrate centralized kernel management delivers near-unconstrained context benefits at lower computational cost.
kernel-managed shared memorymulti-agent systemsprompt injectionretrieval-augmented injectionpersonalization scores
Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning
The study analyzes Preventative Steering's temporal dynamics in adversarial fine-tuning, revealing an early compensatory adaptation phase followed by steady-state decay, with attention output projections as the primary residual-write route. Intervention Delta Preservation experiments show protection relies on active adaptation, not static weight offsets, prompting the proposal of Progressive Intensity Scheduling (PIS). PIS, which escalates injection strength after initial alignment, improves safety robustness in Qwen2.5 and Gemma-3 models while reducing harmful trait expression compared to static-strength steering.
preventative steeringadversarial fine-tuningintervention delta preservationprogressive intensity schedulingattention output projections
Agent-Based ML-LLM Fusion with Self-Optimizing Prompts for Plateau Weather Alerts
The authors propose SmartWeatherAgent, a three-stage architecture integrating intent recognition, hazard prediction, and reasoning-enhanced generation to improve plateau weather alerts. The system combines rule-based methods with large language models for query parsing and employs a LightGBM model enhanced with highland-specific features, achieving a 0.605 F1-Macro score with 1.60 ms latency. A 12-round prompt self-optimization loop improves composite warning quality from 4.2 to 8.9 (+112%), with key enhancements in data citation, physical mechanism explanation, and uncertainty statements. The system autonomously generates structured warnings integrating causal mechanisms, spatiotemporal evolution, and quantitative evidence.
intent recognitionhazard predictionlightgbmprompt self-optimizationspatiotemporal evolution
Context operations to architecture modelling output from large language models and evaluation criteria for their use in systems engineering design
The authors propose a framework for structuring context in LLM-based systems engineering design, comprising modular context units such as policy prompts, persistent reference units, and vectorized user queries. They introduce a formal evaluation method for assessing LLM-generated modeling-as-code outputs against design intent, enabling systematic assessment of LLM support for systems architecture modeling. The framework facilitates structured interactions with generative models while providing criteria for evaluating compliance with engineering design requirements.
context operationssystems architecture modelingmodeling-as-codepolicy promptsprompt vectoring
SA-Profile: Automated Sulcus Angle Profiling from Super-Resolution MRI
The study introduces SA-Profile, an automated framework for continuous sulcus angle (SA) profiling from super-resolved MRI volumes to assess trochlear dysplasia (TD). The method combines clinically acquired axial, coronal, and sagittal MR scans using implicit neural representations, followed by SA computation via two landmark-detection U-Net models across the trochlear region. Evaluated on the fastMRI dataset and an in-house TD cohort, the approach achieved a mean absolute error of 11.6° compared to manual single-slice measurements, revealing distinct SA profiles between cohorts. The framework reduces reliance on manual slice selection while maintaining clinical relevance.
sulcus angletrochlear dysplasiaimplicit neural representationssuper-resolution mriu-net
A Trust-Network-Based Federated Learning Framework for Multi-Center Aging Clock Prediction
The paper introduces TNFL, a trust-network-based federated learning framework for multi-center aging-clock prediction that addresses challenges of limited local data, directional trust relations, and model drift. TNFL combines an age-aware mixture-of-experts model with generative replay to propagate models along directed trust networks without centralized aggregation. Experiments on molecular datasets demonstrate TNFL's effectiveness in maintaining prediction accuracy (quantified implicitly via stable performance across interaction orders), interpretability of age-dependent patterns, and identification of higher-order protein interaction subnetworks spanning aging-related systems.
federated learningaging-clock predictionmixture-of-expertsgenerative replayprotein interaction networks
Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance
This paper introduces a feasibility taxonomy for inference-time AI governance, addressing the shift from training to deployment-stage regulation. The authors categorize twenty mechanisms across monitoring, verification, and enforcement, assessing their readiness on a four-point scale based on evidence from four vendors. The taxonomy is evaluated against a two-dimensional adversary model and mapped to four governance scenarios. Results indicate that fifteen mechanisms have commercial technical substrates, though governance-grade assurance varies. Adversary analysis reveals limitations against high-capability state-level deployers, and fine-tuning impacts enforcement efficacy. A substitution analysis links the taxonomy to hardware-stage mechanisms, and a reliability check yields a Cohen's kappa of 0.74.
inference-time governancefeasibility taxonomyadversary modelfine-tuningsubstitution analysis
RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases
We introduce Research Attention Prediction (RAP), a rolling benchmark evaluating LLMs' ability to track shifts in research attention across 278 AI/ML fields and 1,390 episodes. LLM agents search a temporally restricted arXiv corpus and predict paper shares across eight frozen research directions over six months. Results show that search generally improves performance, but all four diagnostic models underperform an exact-count exponentially weighted moving average (EWMA) baseline. Key bottlenecks include limited future-specific updating and Forecast-oriented policies retrieving less recent evidence. Fine-tuning on realized outcomes improves Qwen3-4B's forecast Spearman correlation by 0.105 on held-out fields.
research attention predictionllm agentsexponentially weighted moving averagearxiv corpusspearman correlation
A statistical approach to bias in zero-shot learning: the lens of handwriting recognition
The paper proposes a statistical method to mitigate misclassification bias in generalized zero-shot learning (GZSL) for large-vocabulary handwritten word recognition. The approach treats existing GZSL feature learners as black boxes and corrects their intrinsic bias toward seen classes using a two-stage hierarchical architecture: a classical GZSL model followed by an ensemble of lightweight Monte Carlo bias-correctors. Debiased classification is then performed using statistical methods (e.g., nearest neighbor, logistic regression). The method achieves a 20% relative accuracy improvement on unseen words and demonstrates that word recognition can be effectively represented in ~15 dimensions. The approach is theoretically grounded and applicable to diverse GZSL implementations.
generalized zero-shot learningbias correctionhandwritten word recognitionmonte carlo methodsstatistical inference
Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
We introduce a reference-based method for auditing bias in hidden-state representations of LLMs, addressing limitations of output-based approaches. The method encodes sentences via similarities to fixed anchor sentences, yielding relative representations that enable comparison across model variants. Representational Bias Shift ($ΔB$) quantifies shifts in group associations with positive/negative attributes. Evaluated across three model families and WildGuardMix, DecodingTrust, and ToxiGen benchmarks, $ΔB$ correlates with output-level bias change in 15/18 settings ($|r| = 0.84$, $p < 0.001$), detects bias increases with ROC AUC 0.65-0.99, and outperforms SEAT-based baselines. The method is computationally efficient (3-50× less compute than benchmarks) and stable across anchor/attribute/template variations.
hidden-state representationsrepresentational bias shiftfine-tuninganchor sentencesbias auditing
NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic Environments
NOPE-HYPE introduces a structured workflow for robust speech-to-text training by combining a controllable environment simulator, coverage-optimal Power Spectral Density (PSD) template reduction, and interpretable hyperparameter search. The method generates synthetic noise conditions to augment training data, achieving performance comparable to real-noise training for Whisper and SeamlessM4T models. A 27-run hyperparameter sweep identifies optimal simulator configurations, providing principled environment prototypes for improved generalization across diverse acoustic conditions.
speech-to-textpower spectral densityhyperparameter searchacoustic simulationwhisper
OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization
OntologyAligner introduces a three-stage framework for biomedical ontology normalization, combining ontology-aligned retrieval, large language model (LLM) candidate reranking, and hierarchy-guided refinement. The method addresses challenges in mapping free-text expressions to standardized concepts due to lexical variation and hierarchical concept distinctions. Evaluated on PhenoNormBench (13,390 samples across seven Human Phenotype Ontology datasets), OntologyAligner achieves state-of-the-art performance with 88.78% Macro Top-1 Accuracy and 86.75% Micro Top-1 Accuracy, outperforming baselines by 4.85 and 5.07 percentage points, respectively. The framework demonstrates portability to MONDO, MEDIC, and NCBITaxon ontologies.
biomedical ontology normalizationontology-aligned retrievallarge language model rerankinghierarchy-guided refinementphenotype ontology
Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training
The paper introduces Direct Diversity Optimization (DDO), an offline post-training method for improving successful strategy coverage in LLM agents for sequential decision tasks. DDO combines Divergence-Tree Collection (DTC), which constructs state-aligned branch sets from shared decision states, with the Reference-Relative Target-Odds Objective (RTO) to train models to match reference-relative targets over successful alternatives. Evaluated on BabyAI, BabaIsAI, and WebShop, DDO outperforms baseline methods in task success, successful strategy coverage, and recovery rate after local action replacement, surpassing successful-only imitation and decoding-time diversification controls.
direct diversity optimizationdivergence-tree collectionreference-relative target-odds objectivesuccessful strategy coverageoffline post-training
Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability
The Belief-State Engine (BSE) augments LLMs for principled planning under partial observability by maintaining a Bayesian posterior over latent states in a POMDP, exposing only this belief to the LLM instead of raw action-observation logs. The authors prove that the LLM-BSE pair forms a sound Markov policy on the induced belief MDP, inheriting Bellman optimality guarantees. Evaluations on the Tiger POMDP and a red-team attack-graph task show BSE improves task return (vs. 6 baselines like ReAct and POMCP), belief calibration, and decision consistency, with ablations confirming robustness across models.
pomdpbayesian posteriormarkov policybellman optimalitybelief calibration
Elastoformer: Enabling Dynamic Adaptivity via Elastic Model Transformation
Elastoformer introduces a framework for transforming conventional neural networks into Elastic Neural Networks capable of real-time elastic inference, addressing dynamic operational constraints in EdgeAI systems. Unlike traditional bag-of-models approaches requiring multiple independent models, Elastoformer provides a single, modular solution that dynamically switches between operation modes at runtime, adapting to varying computational budgets. Experimental results demonstrate up to 85% reduction in FLOPs, 50% reduction in latency, and 76% reduction in memory overhead, while maintaining architecture agnosticism across Vision Transformers and CNNs.
elastic neural networksedgeaireal-time inferencedynamic scalabilitymodular framework
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
MetroLLM-Bench introduces a 955-case benchmark evaluating language models as transit kiosk policy layers, covering six metro systems and eleven task categories. The benchmark includes deterministic (Tier 1) and semantic-quality (Tier 2) scoring, with 238 held-out cases for evaluation. A 4B Qwen 3.5 model, fine-tuned via PEFT, outperforms GPT-5.6 on Tier 1 (91.3 vs. 90.6) and matches GPT-5.4 at maximum reasoning effort (91.4), with larger models showing diminishing returns. Rule-based baselines score 84.6, with LM advantages in policy adaptation and temporal reasoning. Muse Glimmer 30B leads composite rankings.
benchmarkparameter-efficient fine-tuninglanguage-model judgetransit kioskstructured tools
What Makes Adversarial Examples Transfer Across Deepfake Detectors?
This study systematically evaluates adversarial transferability across deepfake detectors, identifying source-target compatibility as a key factor in attack success. The authors conduct controlled experiments on 60 detectors spanning six backbones, two pretraining regimes, and five training-data configurations, using AutoAttack (AA) and Carlini-Wagner with Expectation over Transformation (CW-EOT). Results show higher transferability when source and target share exact backbone, architecture family, pretraining regime, or training data, with compatibility structure being attack-dependent. Mean attack success rates are 7.21% (AA) and 19.52% (CW-EOT), while a multi-source oracle achieves 64.48%. The study releases 240,000 adversarial images, pairwise transfer results, and evaluation code.
adversarial transferabilitydeepfake detectorsautoatackexpectation over transformationsource-target compatibility
Fidelity-Aware Scheduling of Quantum Circuits on Multi-QPU Systems
A fidelity-aware scheduling framework for multi-QPU systems is proposed, utilizing a Graph Neural Network (GNN) to estimate circuit execution fidelity across Quantum Processing Units (QPUs) prior to compilation. The framework employs a tunable scheduler to balance fidelity and parallelism, avoiding exhaustive compilation on all devices. Experimental results demonstrate that this approach approximates optimal fidelity-based assignments while significantly reducing computational overhead compared to brute-force methods.
graph neural networkquantum processing unitfidelity-aware schedulingcompilation overheadmulti-qpu systems
Improving Cross-Lingual Token Representations by Adding a Pinch of SALT
The paper introduces SALT, a lightweight post-training method that enhances token representations in cross-lingual sentence encoders by injecting span-level supervision. SALT addresses the mismatch between sentence-level training and token-level applications like hallucination detection and sequence tagging. Evaluated on five multilingual token-level benchmarks, SALT achieves top performance on four, surpassing alternative fine-tuning strategies and competitive encoders. It also improves sentence-level tasks such as cross-lingual retrieval and classification, demonstrating span-level supervision's efficacy for both token and sentence representations.
cross-lingualtoken representationsspan-level supervisionsentence encoderspost-training
Structural Process Supervision for Latent Chain-of-Thought Reasoning
The paper introduces Prototype-Mediated Process Supervision (PMPS), a method for structural supervision of latent chain-of-thought (CoT) reasoning via learnable reasoning prototypes. PMPS projects latent and explicit CoT embeddings into a shared prototype space, enabling many-to-many soft alignment through prototype assignment, and employs Progressive Sequential Alignment (PSA) to guide training with positional priors. Experiments show PMPS reduces output token length to under 50% of explicit CoT on GSM8K-Aug, achieves 2.08% average accuracy gains over SIM-CoT, and surpasses CoT-SFT on GPT-2 while maintaining superior accuracy on larger models and harder tasks.
latent reasoningprocess supervisionprototype alignmentchain-of-thoughtprogressive alignment
Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models
The paper introduces Time-Frequency Geometric Cross-Attention (TFGCA), a module enhancing vision-language-action (VLA) models by addressing two motion structure limitations: frequency entanglement and cross-phase geometric relationships. TFGCA decomposes action chunks into time-frequency tokens using a learnable stationary wavelet transform and employs a cross-attention mechanism combining dot and wedge products to capture near-orthogonal motion phases. The module integrates via a zero-initialized residual, enabling fine-tuning of pretrained VLAs. Evaluations show improvements: +1.5 on LIBERO, +6.3 on LIBERO-Plus, +28.5 under RoboTwin domain randomization, and +11.67 success rate on AgiBot A2 tasks, with greater gains in out-of-distribution settings.
vision-language-action modelstime-frequency decompositiongeometric cross-attentiondomain randomizationwavelet transform
FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models
FlowCPO introduces a unified divergence framework for preference alignment in flow models, proposing an offline forward-KL objective that utilizes both preferred and dispreferred samples without online rollouts. The method derives a tractable contrastive flow matching loss under regularity conditions, showing nonnegativity unlike FlowDPO's unbounded signed regression loss. Experiments demonstrate superior in-domain performance (GenEval 0.84 vs. 0.81 for FlowDPO, OCR 0.87 vs. 0.74) but mixed out-of-domain results, with higher GenEval but lower reward scores than RFT.
preference alignmentflow modelsforward-kl divergencecontrastive lossoffline optimization
Strangers to Themselves: What Language Models Say About Themselves Is Generic
The study evaluates language models' self-knowledge by comparing their predictions of their own behavior against actual performance and control conditions. Across nine behavioral evaluations, models showed weak direct self-report correlation (r=+0.04), marginally improving to +0.24 with item-specific prompts. Predictions about 'capable AI agents in general' performed similarly (+0.28), suggesting models lack privileged self-knowledge. First-person framing introduced flattering bias, understating harmful behavior. Finetuning on behavioral records enabled narrow self-prediction but altered the predicted behavior without broad transfer. Findings indicate models' self-reports reflect general AI theory and favorable bias rather than accurate self-awareness.
self-knowledgebehavioral evaluationfirst-person biasin-context learningfinetuning
Grounded Evaluation and Repair for NL-to-PDDL Problem Generation
The paper introduces a grounded evaluation framework for NL-to-PDDL problem generation, addressing overestimation of faithfulness in LLM-generated PDDL instances due to parseability or planner success alone. The method combines LLM generation with multi-stage validation (PDDL parsing, planning, domain conformance, LLM critique) and iterative repair using fine-grained feedback from domain descriptions, NL inputs, and diagnostics. Evaluation on Planetarium, AutoPlanBench, and PDDL~2.1 benchmarks reveals divergence between operational success and reference reconstruction, with structured repair improving outcomes but PDDL~2.1 remaining challenging.
nl-to-pddlllm generationiterative repairdomain conformancesemantic equivalence
Decision Transformer for UAV-Mounted RIS-Assisted Dynamic D2D Communications
The paper proposes a Decision Transformer-based approach for optimizing UAV-mounted RIS-assisted D2D communications, jointly addressing UAV trajectory, attitude, and RIS phase configuration under mobility and energy constraints. The method leverages expert trajectories from multiple scenarios to train a deep reinforcement learning model, combining Decision Transformer with DRL for policy learning. Results show superior zero-shot cross-scenario generalization compared to direct DRL transfer, with online fine-tuning achieving competitive performance using 40% fewer interactions than baseline methods.
decision transformeruav-mounted risd2d communicationszero-shot transferonline fine-tuning
Albedo Estimation via Latent Bridge Matching
The paper proposes latent bridge matching (LBM) to address limitations in albedo estimation for intrinsic image decomposition (IID), including physical inconsistency, high inference cost, and poor generalization. The authors introduce an LBM-based architecture enforcing physical consistency via pixel reconstruction loss, leveraging LBM's efficiency, and enhancing generalization through shading conditioning. An extended version further improves reconstruction fidelity by conditioning the shading estimator on predicted albedo. The model outperforms state-of-the-art IID methods across five real and synthetic datasets.
albedo estimationlatent bridge matchingintrinsic image decompositionshading conditioningpixel reconstruction loss
Forward-Free LLM Depth Pruning via Weight Redundancy
Weight-Redundancy Pruning (WRP) is introduced as a forward-free depth-pruning method for reducing large language model inference costs. WRP estimates inter-layer redundancy directly from checkpoint weights by comparing attention output and MLP down-projection weights across layers, combining pairwise similarities with projection-scale information to construct an all-pairs similarity matrix for layer grouping and block selection. This approach eliminates the need for calibration data or model forward passes. Experimental results demonstrate that WRP consistently outperforms existing forward-free magnitude pruning methods across various pruning settings, model families, and downstream tasks, approaching the performance of activation-based methods.
depth pruningtransformer blocksweight redundancymlp down-projectionall-pairs similarity matrix
Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format
This study empirically compares scored versus generated readouts in behavioral language models, demonstrating that scored readouts outperform generated predictions in ranking accuracy across 13 model-domain cells. Using fixed model checkpoints and prompts, the authors evaluate four retail tasks in three markets, finding scored readouts superior in 12 of 13 cases (p ≈ 0.003), with AUC improvements ranging from 1.5 to 14.5 points. Analysis of 9,000 rationales reveals reduced reliance on dominant predictive features and convergence on stock formulations. A third readout method improves calibration but only for outcomes represented in training. The authors propose retaining generated rationales while using scored readouts for ranking.
scored readoutgenerated readoutaucrationalecalibration
AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents
AgentAudit introduces an extensible framework for comprehensive trust evaluation of AI agents across ten dimensions (instruction integrity, planner, memory, tool selection, invocation, correctness, alignment, faithfulness, security, execution integrity), addressing the limitation of existing benchmarks that assess isolated components. The method analyzes full execution traces without modifying agent internals, enabling failure attribution to specific stages. Evaluations of five LLMs (GPT-5, Claude Sonnet 5, Sarvam 105B, Llama 3.3 70B, Gemini 2.5 Flash) reveal significant trust score disparities (95.1 to 22.6), with frontier models outperforming others. Notably, non-frontier models exhibit unsafe compliance in adversarial settings, a nuance missed by pass/fail metrics.
trust evaluationexecution tracefailure attributionadversarial taskscomposite trust score
Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields
The article proposes a relational framework for affective computing, shifting from individual-state paradigms to interactional fields constituted by vocal dynamics. Leveraging affective resonance and vitality-contour theories, the authors employ continuous self-supervised speech representations to detect directional expressive coupling in multi-party conversations. Results show coupling is regime-specific, occurs at sub-second timescales, and collapses under negative controls, supporting a relational account of affective dynamics. The framework introduces Affective Resonance Dynamic Ontologies and null-calibrated directional coupling analyses for Artificial Affective Resonance Intelligence.
affective resonancevitality affectsrelational frameworkdirectional couplingself-supervised speech representations
With a Thermomix You Lose the Ability to Cook: A Kitchen Machine Analogy for Applications of Generative AI in Education
The paper proposes a conceptual framework for analyzing generative AI in education through analogy with the Thermomix kitchen appliance, situating usage patterns within the ICAP (Interactive-Constructive-Active-Passive) and SAMR (Substitution-Augmentation-Modification-Redefinition) pedagogical frameworks. By mapping Thermomix functionalities to AI-assisted learning scenarios, the authors demonstrate how tool integration can either enhance or undermine cognitive engagement depending on implementation modality. The analogy provides researchers and educators with a structured approach to critically evaluate AI's role in learning processes beyond binary adoption debates.
generative aipedagogical frameworkscognitive engagementtool integrationeducational technology
The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents
The Era by Eon Benchmark introduces a synthetic enterprise environment with exact ground truth for evaluating LLM agents in enterprise systems. The benchmark generates fictional companies with consistent entity graphs, product simulators (Salesforce, Zendesk, Slack, Gong), and question-conditioned internal databases, ensuring answer keys are computable from records. Design checks validate database consistency, while realism scores (improving from 61.8 to 97.0 across 23 companies) and adversarial detection ensure synthetic fidelity. In evaluations, nine models answered 33 questions thrice, yielding accuracy estimates of 42.4–76.8%, with three significant pairwise differences post-correction.
llm agentsenterprise benchmarkingentity graphsynthetic data validationquestion-conditioned generation
Can AI Agents Detect and Repair Artifact Drift in Network Experiments?
The study introduces NetArtifactBench, a benchmark assessing AI agents' ability to maintain artifact integrity in network systems by repairing inconsistent records while preserving supported claims. The benchmark comprises 52 instances with injected inconsistencies, evaluating 23 agent configurations across three runtimes using deterministic scoring. Results show a 65.3% average pass rate (5,980 outputs), but performance drops below 30% when repairs require recovering implicit relations and propagating changes, highlighting a gap between local correction and comprehensive record-level repair.
artifact integritynetwork systemsai agentsbenchmarkimplicit relations
Subgroup Membership Inference Audits of Differentially Private Synthetic Text
The study introduces subgroup-targeted membership inference audits to evaluate differential privacy (DP) risks in synthetic text releases, revealing that existing methods underestimate leakage for vulnerable subgroups. The audit framework tests 32 proxies across three attacker knowledge scenarios, four datasets, three generators (DP-SGD fine-tuning, API-based prompting, activation steering), and five privacy budgets. Results show DP reduces average leakage but concentrates residual risk: 10% of records account for 40% of leakage, with uneven protection favoring random over high-risk records. Record-level leakage depends on the release mechanism, not just the record itself.
differential privacymembership inference attacksynthetic datasubgroup leakageprivacy audit
UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model
UnitBoost introduces a non-generative meta-level operator for managing compound LLM systems, replacing traditional generative managers with a task-given unit map, constrained argmax, and explicit residual mechanism. This approach ensures order invariance, unit provenance, and testable failure conditions while avoiding opaque, order-sensitive model calls. Evaluated on three benchmarks, UnitBoost outperforms single-candidate selection with gold labels by 0.060-0.195 points and generative managers by 0.048-0.076. It improves six compound-system configurations by 0.013-0.182 and raises FanOutQA cell F1 from 0.4778 to 0.5524 through residual-directed rounds. The method quantifies cross-unit coupling as repair cost and identifies conditions where no gain is available.
compound llm systemsmeta-level operatorunit provenanceconstrained argmaxresidual mechanism
uFlowCSP: Crystal Structure Prediction using Mean flow generative models
We introduce uFlowCSP, a MeanFlow-based crystal structure prediction (CSP) model that learns the average probability-flow velocity, enabling generation of complete structures in 1-5 network evaluations. The model employs a chemistry- and symmetry-aware Transformer with canonical atom ordering, global composition, and per-token chemistry embeddings, while using a coarse crystal-system token only during training. uFlowCSP achieves 83.64% accuracy on MP-20 with 5 steps, outperforming CrystalFlow (78.34%) and DiffCSP (77.93%) while using 20x fewer evaluations. It generates 10,000 structures in 0.39-1.31 minutes, significantly faster than CrystalFlow (6.5 minutes) and DiffCSP (76.1 minutes). Under CSPBench's top-five criterion, uFlowCSP reaches 72%/72%/65% structure, space-group, and consensus match rates, demonstrating improved accuracy per evaluation.
meanflowcrystal structure predictiontransformerprobability-flow velocitycspbench
CS-Guard: Benchmarking LLM Guardrails for Code Generation Security
We introduce CS-Guard, the first benchmark for systematically evaluating guardrails in LLM code generation security. The benchmark assesses text-to-code generation using 1000 malware-generation prompts, 7 jailbreak attacks, and a novel fictional scenario attack (FSA), alongside code-to-code generation with 331 prompts spanning infilling, completion, and translation. Empirical evaluation of 9 guardrails across 7 LLMs reveals poor performance: text-to-code guardrails exhibit ~50% average attack success rate (ASR) post-jailbreak, while code-to-code guardrails show ASRs ranging from 14.4% to nearly 100%. The FSA achieves ~100% ASR, highlighting reliability concerns. CS-Guard employs a modular three-layer guardrail taxonomy and is released for community evaluation.
guardrailscode generationattack success ratejailbreakbenchmark
How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE
This work demonstrates the fragility of safety alignment in frontier-scale mixture-of-experts (MoE) models through a single-direction ablation attack on GLM-5.3-Flash (320B parameters, 288 experts). The attack removes refusal capabilities by projecting a 'refusal direction' out of residual stream weights, requiring only contrastive prompts without gradient-based training. Results show that 74% of refusal removal occurs only under joint intervention across attention, dense, and routed-expert writers, with conventional module-name matching accounting for just 0.066 of the total effect. The attack achieves 41-89 percentage-point reductions across seven harmful benchmarks without detectable capability loss, though category-concentrated residues persist across edited subspaces.
mixture-of-expertsresidual streamcontrastive promptsmodule-name matchingsafety alignment
LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios
LogiScope-VQA introduces a benchmark for evaluating Large Multimodal Models (LMMs) in logistics hazard identification, addressing the scarcity of industrial data. The dataset comprises 2,476 images, 2,918 videos, and 10,274 VQAs curated from real-world logistics parks, organized into 39 subtasks across three themes: industrial element perception, warehouse knowledge understanding, and potential risk reasoning. Dynamic thinking-budget configurations and dual-dimensional risk bias analyses are employed to assess LMM properties. Experiments reveal significant performance gaps in proprietary models (GPT-5.5, Gemini-3.1-Pro, Claude-Opus-4.7) compared to human experts, highlighting challenges in integrating perception, understanding, and reasoning. The dataset is publicly available under CC BY-NC-SA 4.0.
multimodal modelshazard identificationvqarisk reasoningindustrial dataset
Pairit: A Platform for Live Experiments on Human-AI Collaboration
Pairit introduces an online platform for designing, testing, and deploying live experiments on human-AI organizational collaboration. The platform enables researchers to declare executable experiment graphs via a single YAML configuration file, integrating pages, routing, randomization, chat, shared workspaces, server-hosted agents, surveys, timers, and custom HTML components. Pairit supports combining humans and AI agents in live sessions, facilitating high-resolution process tracing of communication, negotiation, and collaborative work. Validated through multiple live deployments and peer-reviewed studies, Pairit standardizes complex interactive protocols into auditable configuration files, offering reusable infrastructure for human-AI organizational experiments.
human-ai collaborationexperiment graphyaml configurationprocess tracingorganizational experiments
BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL
BRACE introduces an anchored Bellman-residual correction to address stale critic bias in asynchronous RL for language models. The method bounds the correction horizon to a policy token prefix and anchors a constant-weight Monte-Carlo tail beyond it, decoupling policy correction from reward propagation. BRACE achieves a 2.4% improvement in mean@1 on BrowseComp-Plus, runs 2.46× faster per step than synchronous training, and maintains stability up to 50 updates off-policy.
asynchronous reinforcement learningbellman-residual correctionpolicy lagmonte-carlo tailoff-policy stability
Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward
The paper introduces Proof-Carrying Cognition, addressing the verification gap in language-model reasoning through four contributions. First, it presents a theoretical model linking verifier-gold correlation to compute-capability trade-offs. Second, empirical demonstrations show unsound verifiers degrade under optimization, while reality-anchored settlement improves soundness log-linearly and reduces hacking gaps. Third, it proposes a paradigm where reasoning steps are probabilistic claims priced by a self-built world model. Fourth, it specifies Soundness-under-Pressure as a benchmark metric for reality-settled reasoning. Results include a 10x label efficiency gain with on-policy settlement and preservation of 6x executed reward under GRPO training.
verification gapsoundness-under-pressurereality-anchored settlementproof-carrying cognitiongrpo training
Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks
The study investigates procedural memory interference in language agents when stored routines become mismatched with task conditions. Using BrowserGym TimeWarp and synthetic shopping tasks, it evaluates memory reuse under four mismatch types: changed quantities, altered evidence representation, local-global optimization conflicts, and distributed promotion evidence. Experiments involved 32 formal cells with temperature-0 generation using a qwen3:8b configuration. Results show no predefined interference signatures when task evidence is explicit, indicating procedural memory mismatches need not disrupt behavior. The work identifies a non-interference region but does not establish general safety mechanisms or conditions triggering memory-caused errors.
procedural memoryinterference signaturestemperature-0 generationlocal-global optimizationevidence representation
Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?
We propose a hybrid approach combining KV cache-aware fine-tuning and selective KV cache recomputation to optimize Retrieval-Augmented Generation (RAG) systems for long-context inputs. Our method fine-tunes models to account for KV cache concatenation while selectively recomputing a subset of caches, addressing both response quality and time to first token (TTFT). Experiments on the RULER benchmark demonstrate a 9.7-point improvement in RULER score for 124k-token inputs compared to baseline KV cache recomputation, alongside an 80% reduction in TTFT relative to full attention.
retrieval-augmented generationkv cachetime to first tokenruler benchmarklong-context inputs
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
We introduce LexAgentHallu, a hierarchical benchmark for profiling hallucinations in legal agents, addressing limitations of existing legal benchmarks that lack multi-turn trajectory evaluation and legal-specific diagnostic capability. Constructed via a four-stage expert-in-the-loop pipeline, the benchmark comprises 3414 instances across 17 legal categories and 6 task types, annotated under a dual-layer hallucination taxonomy of 7 high-level categories and 27 fine-grained subclasses. Evaluation of 18 proprietary and open-source agents reveals a Right-Answer-Wrong-Reason effect and clustered hallucination subclasses, forming distinct agentic framework, legal task, and category profiles, demonstrating LexAgentHallu's diagnostic power.
lexagenthalluagentic hallucinationdual-layer taxonomyright-answer-wrong-reasonexpert-in-the-loop
HiRAD: A Flexible Large-Scale AGV Routing System
HiRAD introduces a hierarchical reinforcement learning framework for continuous-space AGV routing with real-time guarantees, addressing scalability and kinematic mismatches in existing methods. It employs a step-level spatiotemporal representation for differentiable RL, splits heading choice from velocity control to reduce action space, and utilizes an asynchronous event-driven decision pipeline to lower inference complexity from O(n^2) to O(n). Evaluated on random graphs and warehouse maps, HiRAD reduces makespan by 45-63% and decreases end-to-end runtime while cutting per-step latency by up to 71%.
hierarchical reinforcement learningcontinuous-space routingspatiotemporal representationasynchronous pipelinereal-time guarantees
Distilling Image Prototypes for Guided Test-Time Adaptation
The paper introduces Distilling Image Prototype for Guided Test-Time Adaptation (DIPTTA), a framework addressing error accumulation and catastrophic forgetting in Test-Time Adaptation (TTA). DIPTTA employs a Distill Image Prototype (DIP), a compact set of synthetic images that dynamically anchors source knowledge through continuous feature replay, ensuring alignment with the model's current state. Additionally, DIP enables source-calibrated uncertainty estimation, reducing bias in sample reliability assessment. Experiments across multiple benchmarks show DIPTTA outperforms state-of-the-art methods, particularly under severe domain shifts.
test-time adaptationdistill image prototypecatastrophic forgettingerror accumulationdomain shifts
Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety
CareGuard introduces an early-warning framework for cyberbullying detection to support mental health monitoring and online safety. The system combines zero-shot semantic labeling with fine-tuned transformer models (BERT, DistilBERT, RoBERTa) and employs emotion-aware filtering alongside cosine similarity-based semantic screening to prioritize contextually relevant and emotionally salient content. Evaluations on benchmark datasets show that CareGuard achieves a balance between detection accuracy and computational efficiency, suggesting viability for healthcare and online safety applications.
cyberbullying detectiontransformer modelsemotion-aware filteringzero-shot labelingsemantic screening
Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
We introduce Trimmed Logit-Gap SFT (TrimSFT), a token-level reweighting method for supervised fine-tuning that scales loss based on the logit gap between gold tokens and their strongest competitors. TrimSFT removes supervision from both mastered tokens (large logit gap) and uncertain tokens (small/negative logit gap), focusing learning on intermediate regions via a Gaussian weight centered at margin m with bandwidth τ. Evaluated on six base models across five mathematical reasoning benchmarks, TrimSFT outperforms standard SFT, achieving up to +26.9 points on MATH500. Analysis shows bandwidth τ is more critical than margin location, and half-trim variants yield inferior trade-offs. Token-level logit-gap distributions indicate TrimSFT reshapes model confidence more balanced than uniform SFT or monotonic reweighting.
supervised fine-tuninglogit gapmathematical reasoningtoken-level reweightinggaussian weight
Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation
The study investigates correctness-gated multi-teacher distillation, focusing on decision shifts and label functionality loss. Using a fixed experimental setup with eight arms, 4,330 sources, and a 63.9M-parameter student, the authors evaluate three seeds on 267 held-out examples. Results show correctness-weighted distillation improves accuracy (+0.1660) and macro-F1 (+0.1323) but reduces unsafe-action rates (-0.4979). However, the weighted arm exhibited zero Refuted recall and misclassified claims as NotEnoughInfo. An audit revealed inconclusive grounding effects, with no incremental benefit over hard filtering, which achieved 0.660 accuracy and 0.530 macro-F1.
multi-teacher distillationcorrectness-gatedmacro-f1grounding auditunsafe-action rate
Kernel-Complexity Edge Sanitization for Training-Free Defense against Structural Graph Attacks
Kernel-Complexity Edge Sanitization (KCES) is introduced as a training-free, model-agnostic defense against structural attacks on Graph Neural Networks (GNNs). The method leverages Graph Kernel Complexity (GKC), derived from a generalization upper bound on GNN test error, to compute edge-specific KC scores quantifying structural influence. KCES identifies and prunes high-KC edges, empirically enriched with adversarial perturbations, without requiring retraining. Experiments demonstrate KCES outperforms robust baselines across diverse attack settings and scales effectively to large graphs. Theoretical analysis and empirical validation support its efficacy in securing GNNs.
graph neural networksstructural attacksgraph kernel complexityedge sanitizationadversarial perturbations
When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination
The study evaluates Google Gemini 3.0 Pro's reliability in detecting planted document contaminants, revealing significant degradation in performance at scale. Using a corpus of 150 academic papers with 450 injected contaminants (typographical, semantic reversal, absurd insertion), the model's recovery rates dropped from 50-60% in single/small-batch evaluations to 2.8% in large batches, with failure modes including confident fabrication of non-existent contaminants (e.g., 'telepathic squirrel'). Detection varied by contaminant type, with absurd insertions recovered at 75% versus 50% for semantic/typographical errors. The findings highlight risks in LLM-based auditing and recommend bounded batch sizes and mechanical verification.
document contaminationbatch-size degradationconfident hallucinationllm auditingsemantic reversal
CT-SAFR: Safe and Interpretable Chain-of-Thought Reasoning for Autonomous Robots: A Multi-Layered Verification Framework for Trustworthy AI-Driven Robotic Decision Making
CT-SAFR introduces a multi-layered verification framework for enhancing Chain-of-Thought reasoning in autonomous robots, addressing hallucination and faithfulness issues. The method combines verification layers to detect 94.2% of hallucinations (n = 500, 95% CI: 91.8-95.9%) with sub-500ms latency. Evaluated through a warehouse robot case study, CT-SAFR reduces unsafe reasoning outputs by 87% (p < 0.001), demonstrating improved safety and interpretability in AI-driven robotic decision-making.
chain-of-thoughthallucination detectionverification frameworkautonomous robotsreasoning faithfulness
Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling
Looped GPT-BERT introduces a parameter-efficient architecture for small language models by combining GPT-BERT's masked next-token and causal language-modeling objectives with depth-wise parameter sharing. Trained on a 7.48M-word English corpus in the BabyLM 2026 Strict-small setting, the $4\times12$ model employs four physical layers for twelve recurrent traversals, totaling 12.18M parameters. It achieves comparable performance to GPT-2 and GPT-BERT baselines on BLiMP and GLUE benchmarks, with an Overall Average of 35.42 and NLP Average of 48.48 on BabyLM 2026. Ablations reveal that additional recurrent computation improves training but highlights limitations in representational space due to few physical layers.
parameter sharingrecurrent computationlanguage-modeling objectivesdepth-wiserepresentational space
Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QA
The paper introduces RMS-RSP, a method for selecting which labeled medical QA questions should receive rationale supervision under token budget constraints. RMS-RSP perturbs hidden states at rationale tokens and measures shifts in the gold-versus-best-distractor margin to prioritize samples. Evaluated across five medical QA datasets using MedGemma-4B-IT, RMS-RSP shows a modest average accuracy gain (60.61% vs 60.08% for Random) but significantly improves robust accuracy (+1.91 points) and semantic consistency (+2.85 points) after answer-option reorderings. Full supervision raises accuracy to 63.74% but requires 29-254x more tokens without uniformly improving robustness. Results suggest rationale-local boundary sensitivity identifies supervision improving invariance to formatting changes.
rationale supervisiontoken budgethidden state perturbationrobust accuracysemantic consistency
Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents
The paper introduces Cros, a risk-constrained stopping layer for sequential clinical diagnosis agents, addressing when to diagnose or defer based on state-wise error ranking and policy design. The method employs LTT-style exact tests for selective diagnostic error and minimum autonomous coverage, ensuring finite-sample guarantees by freezing candidate families and testing rules before calibration. On a MIMIC-derived abdominal-pain benchmark with 1,834 episodes, Cros achieves an exploratory state-error AUROC of 0.853, outperforming maximum class probability (0.715) and native stop scores (0.552). Evaluation on 367 episodes shows 16.9% selective error at 78.8% coverage, reducing costs and test counts compared to native stopping. However, Cros meets the joint criterion in only 6 of 20 development resplits, indicating exploratory feasibility rather than confirmatory safety.
risk-constrained stoppingsequential clinical diagnosisstate-wise error rankingfinite-sample guaranteeselective diagnostic error
Introducing Consort: A Spec-First Agent Framework for Enforced, Test-Driven Development on Live Database Branches
The paper introduces Consort, a spec-first agent framework enforcing test-driven development via immutable controls on live database branches. Consort employs a deterministic orchestrator to guide role agents through design and build lanes, ensuring code correctness through unmodifiable tests and gates. The authors contrast three enforcement modes—persuasion, front-loaded structure, and immutable controls—positioning Consort's approach as superior for maintainability and verifiability. Claims are framed as a pre-registered hypothesis, emphasizing empirical testability.
spec-firsttest-driven developmentdeterministic orchestratorlive database branchesimmutable tests
PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations
PRAGMA introduces a benchmark for evaluating personalized guidance in long-term conversational AI, addressing the limitations of current systems in integrating multi-conversation context and reasoning about evolving user preferences. The benchmark includes longitudinal dialogues, annotated evidence, and scenarios with dynamic user contexts and misconceptions. Experiments with retrieval systems, memory architectures, and long-context models reveal deficiencies in both evidence retrieval and its application for personalized guidance, underscoring the need for improved memory systems supporting conversational reasoning.
personalized guidancememory alignmentlong-term conversationsretrieval systemscontextual reasoning
Cascading Gradient Inversion via LT-Code Inspired Peeling in Federated Learning
The paper introduces a novel gradient inversion attack for federated learning that exceeds known upper bounds by leveraging connections to erasure-correcting codes. The method, inspired by LT-code peeling, enables exact batch recovery from a single FedSGD round, certifying reconstructions without ground-truth data. Evaluated on eight benchmarks, the attack recovers 94–100% of ImageNet batches (size ≤128) passively and >90% actively at batch sizes of several hundred, significantly outperforming prior single-round attacks and demonstrating underestimated privacy risks in federated learning.
gradient inversionfederated learningerasure-correcting codesbatch recoveryprivacy leakage
RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
The paper introduces RESCUE-BENCH, a benchmark for relation-aware multi-party emotional support conversation systems, addressing the gap in modeling interpersonal dynamics beyond one-on-one interactions. The dataset comprises 191 samples (7,079 turns, 1,064.8 minutes of video) from couple and family interviews, annotated for socio-emotional and support-related dynamics. Six tasks evaluate relational understanding and relation-sensitive support capabilities. Testing ten LLMs reveals strengths in local emotion/intervention tasks but weaknesses in relation-intensive tasks (pattern/viewpoint/support strategy prediction), highlighting limitations in modeling interpersonal relations.
emotional support conversationmulti-party interactionrelation-aware modelingbenchmark evaluationinterpersonal dynamics
Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery
The authors introduce a black-box framework for automated risk discovery in agentic AI systems, addressing limitations of single-turn evaluations. The method comprises a seven-domain taxonomy linking behaviors to risk categories, fully automated SAGE-RT red teaming generating 120 adversarial scenarios per domain, and human-validated evaluation using LLM judges. Empirical validation across CrewAI and AutoGen architectures with four base models reveals significant vulnerabilities: 56.25% governance risk, 65% privacy risk in multi-agent configurations, and agent behavior vulnerabilities reaching 85%. The framework identifies critical architectural vulnerabilities without privileged access, enabling scalable safety improvements for agent deployments.
agentic aiblack-box evaluationsage-rtmulti-agent configurationsllm judges
RobustSGPO: Search-Space Control for Agent Harness Evolution
RobustSGPO introduces search-space control for semantic-gradient-based prompt optimization (SGPO) in agent harness evolution, addressing unresolved edit scope and operation choices. The method specifies edits, constructs patches, and continues search from incumbent or retained snapshots, incorporating permission scheduling, cumulative controls, and task-family transfer. Evaluated on 120 tasks with 95 runs and 7,350 attempts, RobustSGPO improves task completion from 60.0% to 80.0% and test quality from 3.77 to 4.14 under a 20-million-token budget. Periodic scheduling outperforms fixed maximum permission by 0.28 test-score points, and category retention mitigates source-task degradation post-shift.
semantic-gradient-based prompt optimizationsearch-space controlpermission schedulingtask-family transfercategory retention
Seven Sources of Physical AI Capability Formation
The study proposes a taxonomy of seven capability-formation sources for Physical AI, addressing limitations in existing morphological and task-based classifications. Through reconstructive induction with theoretical saturation, the authors analyzed 49 evidence records across challenges including curriculum learning, active inference, and neuro-symbolic architectures. Results demonstrate that all cases were explainable by Recorded-Experience, Predictive-Modeling, Evaluative-Interaction, Surrogate-Environment, Mechanism-Grounded, Embodied-Coupling, and Evolution-Driven Formation sources, individually or in combination. The framework distinguishes capability similarity from formation similarity, supporting analysis of transfer, replication, and governance. Theoretical saturation was achieved within the defined scope, though not claiming logical completeness.
capability-formationtheoretical saturationneuro-symbolic architecturesactive inferencephysical ai
Hyperbolic Geometry for Open-World Object Detection in Remote Sensing Imagery
HyRS-OWOD introduces hyperbolic geometry for open-world object detection (OWOD) in remote sensing imagery, addressing limitations of Euclidean space representations. The method employs a Decoupled Objectness Learning (DOL) module to separate foreground proposals from background regions and a Hyperbolic Uncertainty Learning (HUL) component using hyperbolic embedding radii for known-unknown discrimination. Incremental learning is enhanced via Hyperbolic Metric Learning (HML), improving inter-class separability and mitigating catastrophic forgetting. Evaluations on three remote sensing benchmarks show consistent improvements in unknown recall and incremental learning over state-of-the-art OWOD methods.
hyperbolic geometryopen-world object detectiondecoupled objectness learninghyperbolic uncertainty learninghyperbolic metric learning
From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital Twins
The paper proposes a four-layer Cognitive Digital Twin (CDT) architecture comprising physical, digital-twin, cognitive, and task layers, enabling self-evolving closed-loop operation. The architecture integrates cognition through knowledge, memory, and attention mechanisms, facilitating task-specific cognitive model construction and decision-making under practical constraints. Operational feedback refines cognitive experience and updates digital representations, supporting evolving task interpretation and reasoning. Two operation modes—user-request-driven and self-driven cognition—are characterized, with discussions on semantic communication, knowledge querying, and task orchestration. A lightweight simulation demonstrates reliable closed-loop task feasibility under limited semantic information and improved efficiency through task experience accumulation.
cognitive digital twinself-evolving looptask orchestrationsemantic communicationknowledge querying
RouteBridge: Reliability-Routed Bidirectional Distillation Between Neural Radiance Fields and 3D Gaussian Splatting
RouteBridge introduces a reliability-routed bidirectional distillation framework between Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS), dynamically selecting the optimal teacher representation per ray to mitigate local reconstruction errors. The method combines photometric residuals and geometric evidence in a reliability estimator, enabling adaptive routing of supervision between NeRF and 3DGS without requiring shared features or point correspondence. Evaluations on mip-NeRF 360 show PSNR improvements of 1.56 dB for 3DGS exports (28.77 dB) over baseline 3DGS and 0.45 dB over NeRF-GS, with LPIPS reduced to 0.207. On DTU's three-view static scenes, RouteBridge achieves 21.12 dB, demonstrating gains from both adaptive routing and geometric ray targets.
neural radiance fields3d gaussian splattingbidirectional distillationphotometric residualsadaptive routing
Watermarks Without Verification: AI Text Watermarking After the EU AI Act
This work identifies unverifiability as the core governance failure in AI text watermarking post-EU AI Act implementation, rather than watermarking itself. The authors analyze contested claims about SynthID-Text, deployed in Claude and Gemini models, by evaluating an open-source implementation on two open-weight models. Results show watermarking's impact on prose generation is negligible compared to seed variation, while code correctness decreases by 3 points on one model with detection remaining near chance. The study maps remaining gaps to requirements including matched output release, configuration disclosure, accredited audits, shared evaluation protocols, and interoperable detection.
synthid-textwatermarkingunverifiabilitygenerative aidetection
Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints
This work investigates contact prediction for robotic manipulation through a compact visuotactile world model, trajectory-level uncertainty calibration, and behavior-initialized actor-critic learning. The model, tested on 160 MuJoCo Lift episodes, reduces endpoint-force prediction error from 1.058 N to 0.228 N and interval-peak error from 2.724 N to 0.523 N when incorporating touch, with tactile persistence achieving lower errors of 0.095 N and 0.498 N, respectively. Reward revision increases in-distribution lifting success from 20.0% to 93.3%, though success within an 8 N per-finger budget remains limited. Trajectory-level calibration improves coverage from 15.80% to 87.36% in GelSight recordings. Findings highlight distinctions between sensing improvements and force-constrained control.
visuotactileuncertainty calibrationactor-criticforce predictiontrajectory-level
Teacher Geometry Shapes Learnability in Teacher-Student Networks
The study investigates how teacher geometry impacts learnability in teacher-student neural networks, formalizing learnability as the success rate of converging to the global minimum. By analyzing the loss landscape of small networks, the authors identify two types of suboptimal local minima: out-of-bounds (OOB) and interior minima, showing that their regions of attraction vary with teacher structure. In larger networks, maximally dissimilar teachers induce more interior minima, while minimally dissimilar teachers induce more OOB minima. The authors propose adjusting learning rates differentially for the readout layer and inner biases to improve success rates, bridging the gap between theoretical teacher-student networks and practical structured functions.
teacher-student networksloss landscapelocal minimalearnabilityoverparameterization
Modality-Decoupled Federated Learning for Privacy-Preserving Embodied Intelligence in 6G
The article proposes FedMVLA, a modality-decoupled federated learning framework for privacy-preserving embodied intelligence in 6G networks. FedMVLA addresses challenges in vision-language-action (VLA) model training by introducing modality-aware federated aggregation, privacy allocation, and communication compression, alongside a modality-sliced transport design. Evaluated on federated robotic manipulation over a 3GPP-based wireless substrate, FedMVLA achieves an 84.8% task success rate, surpassing FedAvg by 22.2 percentage points, scales effectively to 128 clients, reduces per-client uplink payload by 95.6%, and maintains round-critical uplink completion time at 1.5s.
federated learningvision-language-actionmodality-decoupled6g networksrobotic manipulation
A Function-Space Approach to the Statistical Mechanics of Learning Dynamics
(No summary returned.)
CityPlanner: A Sandbox Agent for Executable Urban Planning
CityPlanner introduces a sandbox-agent framework for executable urban planning, addressing spatial optimization through task decomposition and iterative refinement. The method employs UrbanSandbox, a unified file-based environment where agents inspect tasks, generate plans, and revise decisions via executable feedback. Atomic-task reinforcement learning decomposes trajectories into BuildPlan for initial construction and ImprovePlan for feedback-based refinement. Experiments on a real-world benchmark demonstrate CityPlanner's superiority over heuristic, task-specific RL, and general LLM-agent baselines. Ablations confirm the contributions of UrbanSandbox, atomic-task RL, and iterative deployment. Code and dataset are publicly available.
urban planningsandbox-agentatomic-task reinforcement learningspatial optimizationexecutable feedback
Myocardial Strain Drift Correction in Deep Learning Based Ultrasound Tracking
The study proposes a deep learning framework for drift correction in myocardial motion tracking from echocardiography, addressing temporal drift that causes inaccurate strain estimates. The method extends TAS-Net with persistent memory tokens to share information across sliding windows over cardiac cycles, employing teacher-student fine-tuning on real echocardiographic data to enforce cyclic motion consistency. Experiments demonstrate reduced global and regional strain drift, improved agreement with clinical references, and enhanced test-retest reproducibility, supporting more reliable strain estimation in clinical practice.
myocardial strainechocardiographydeep learningtemporal driftteacher-student fine-tuning
Learning with Synthetic Data via SGD in High-Dimensional Linear Regression
This paper investigates the impact of synthetic data on generalization in high-dimensional linear regression trained via one-pass SGD, focusing on model shift. It establishes finite-sample risk bounds for mixed and two-stage training protocols, isolating standard bias/variance from source-mismatch effects like fluctuation and persistent drift. Results demonstrate that mixed training induces model collapse, while two-stage training avoids this by limiting synthetic data to the initial stage. Scaling laws reveal that larger models amplify degradation in mixed training, while high-quality synthetic pretraining reduces bias in two-stage training. A necessary-and-sufficient condition is derived for two-stage training to outperform real-only training under identical real-data budgets.
synthetic datahigh-dimensional linear regressionmodel collapsetwo-stage trainingscaling laws
Multi-Agent Agentic Graph Learning via Structural Signatures
The paper introduces Multi-Agent Agentic Graph Learning (MAAGL), a framework addressing limitations in existing agentic graph learning methods by enabling multi-agent collaboration. MAAGL partitions graphs into communities, assigns independent agents for region-specific specialization, and separates structural and semantic evidence. Structural evidence is summarized via dynamically updated, permutation-invariant signatures, while semantic evidence is filtered to top-k relevant nodes. Agents estimate confidence based on historical trajectories and engage in debate-style collaboration when needed. Experiments on four benchmark datasets demonstrate MAAGL's superiority over state-of-the-art AGL methods.
multi-agent collaborationagentic graph learningstructural signaturepermutation-invariantsemantic evidence
The Vibe Shift in Software Engineering: Evaluating AI-Led Conversational Programming for Performance, Cognition, and Responsible Adoption
The study evaluates Vibe Coding, an AI-led conversational programming paradigm, through a mixed-methods design comparing performance, cognition, and responsible adoption against traditional and AI-assisted coding. Thirty participants (professionals and students) completed tasks under three conditions, analyzed via ANOVA and thematic analysis. Results show 27% faster task completion than traditional coding and 12% faster than AI-assisted, but with lower maintainability (p<0.05) and higher security risks (SUS=71.4, NASA-TLX=55.5). Key themes include trust calibration and prompt-engineering strategies, prompting a proposed framework for hybrid human-AI integration and oversight.
conversational programminglarge language modelscognitive workloadsoftware maintainabilityprompt-engineering
High-probability guarantees for linear accessibility in feature superposition
The study establishes high-probability guarantees for linear accessibility in neural networks employing feature superposition, addressing cross-feature interference constraints. By modeling linear accessibility as a compressed sensing problem, the authors derive dimension bounds scaling linearly ($d=O_{\varepsilon}(k \log m)$) under subgaussian noise, improving upon prior quadratic worst-case limits. Theoretical results are validated via Gaussian-tail approximations across system parameters, quantifying geometric constraints of the linear representation hypothesis. The framework informs sparse autoencoder design, compositional generalization, and neural interpretability.
feature superpositionlinear accessibilitycompressed sensingsubgaussian noisesparse autoencoders
Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning
This paper demonstrates that arbitrary cipher attacks, a form of jailbreak bypassing large language model alignment, do not require fine-tuning for frontier models. Instead, models acquire cipher-based communication skills through prompting and in-context learning, significantly weakening or entirely bypassing alignment safeguards. The authors successfully execute these attacks against commercial models from Anthropic, Google, and OpenAI, evading harmfulness classifiers by encrypting harmful content into nonsensical text. This constitutes a novel attack vector against black-box large language models.
arbitrary cipher attacksjailbreakin-context learningmodel alignmentharmfulness classifiers
A Statistical Approach to Estimating Sample Size of Machine Learning Models
The authors propose a statistical framework for estimating sample size requirements in machine learning prediction models, addressing the challenge posed by nonlinear models' complex prediction surfaces. Their method approximates nonlinear ML models using localized linear representations and evaluates statistical power across these local regions. This approach enables sample size determination without requiring a priori specification of predictor-outcome relationships or effect structures, which are typically necessary in conventional power analysis.
sample size determinationmachine learningnonlinear modelsstatistical powerlocalized linear representations
Adaptive Distributed Physical-Layer Authentication and Attack Detection in 6G Non-Terrestrial Networks via Causal Meta-Learning
SAFA-MZ proposes a causal meta-learning framework for distributed physical-layer authentication (DPLA) in 6G non-terrestrial networks (NTNs), addressing distribution shifts from Doppler shifts and channel variations. The method combines multi-feature fingerprints (spatial, angular, combiner, subspace, Doppler-delay) with a structural causal model (SCM) and model-agnostic meta-learning (MAML) using invariant risk minimization (IRM) and causal consistency regularization. A two-stage scheme with graph attention (GAT) networks reduces backhaul overhead via conditional time-difference-of-arrival (TDOA) activation. Simulations show 92% accuracy and 96% AUC, outperforming centralized deep learning and single-feature baselines.
physical-layer authenticationcausal meta-learningnon-terrestrial networksinvariant risk minimizationgraph attention networks
From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function Calls
The paper compares Functional Tokens (FT) and Schema-in-Prompt (SIP) for small language models (SLMs) in vehicle function calling, introducing a benchmark of 9,822 examples across 79 Android Automotive functions. Under matched fine-tuning on SLMs (270M–1.7B parameters), FT matches SIP on trained functions (270M suffices) but fails on held-out functions, while SIP generalizes and improves with scale. SIP also better refuses out-of-scope requests but incurs higher memory and latency. The analysis shows schema representation, not scale alone, dictates SLM capabilities and failure modes.
small language modelsfunction callingin-context learninggeneralizationlatency
Distributed Physical Layer Authentication and Collaborative RSMA in Non-Terrestrial Networks via Graph Reinforcement Learning
We propose SAFA-MZ, a secure adaptive federated authentication scheme for multi-zone non-terrestrial networks (NTNs) that maximizes secrecy spectral efficiency (SSE) while ensuring authentication reliability and coverage constraints. The method embeds group-level authentication tags into a collaborative multi-layer rate-splitting multiple access (RSMA) transmission structure, jointly beamforms private and common signals, uses artificial noise to reduce leakage, and applies group differential privacy for tag protection. A joint SSE maximization problem is solved via repair-based cross-entropy method and graph-aware advantage actor-critic algorithm. Simulations show SAFA-MZ improves average SSE by 135% over single-connect transmission and 21% over schemes without artificial noise.
physical-layer authenticationrate-splitting multiple accesssecrecy spectral efficiencygraph reinforcement learningnon-terrestrial networks
ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance
We introduce CONTRACTEVAL, a diagnostic framework for evaluating procedural instruction conformance in LLM agents by explicitly matching query-active obligations against response or trace evidence. The method represents procedural instructions as query-active obligations and identifies distinct conformance failures such as omissions, wrong branches, ordering errors, extra actions, invariant breaches, and output-contract violations. On a controlled suite of audited procedural contracts, CONTRACTEVAL detects and localizes all structural failures, outperforming output-only and trace-aware LLM judges. While LLM-backed extraction preserves much of the signal, it remains calibration-sensitive, making CONTRACTEVAL an auditing tool rather than a compliance guarantee.
contractevalprocedural conformancequery-active obligationstrace evidencellm agents
Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations
The paper introduces two methods for calibrating agent confidence from internal representations in multi-turn interactions: Latent Trajectory Dynamics (LTD), which tracks residual-stream representation changes across trajectories, and Action Representation Probe (ARP), which predicts success from action-decision representations. Evaluated on Bash, SQL, and Python benchmarks using Qwen14B, Qwen7B, and DeepSeek6.7B models, these methods outperform surface-level generation and sequence-based calibration baselines without requiring prompt modifications or multi-sample rollouts.
agent confidenceresidual-stream representationsmulti-turn interactionscalibration methodszero-overhead monitoring
Efficient Leakage-Free Neural Architecture Search under Leave-One-Subject-Out Evaluation
Proposes a leakage-free, block-based Neural Architecture Search (NAS) method for Leave-One-Subject-Out (LOSO) evaluation, reducing computational cost from O(N^2) to shared runs across subjects. The approach employs nested NAS with parameter sharing, maintaining evaluation integrity while improving efficiency. On BioVid Heat Pain dataset, it achieves 83.39% mean accuracy (vs. 82.79% baseline) with up to 99.2% parameter reduction.
neural architecture searchleave-one-subject-outparameter sharingbiovid heat painnested evaluation
SCCM : Stream Cruise Control Method for Automated Drift Detection and Adaptation
The Stream Cruise Control Method (SCCM) introduces a framework for automated drift detection and adaptation in online regression, addressing evolving data distributions. SCCM integrates early-response drift detection, dynamic hyperparameter tuning, and model recalibration, employing KPI-window-based thresholding for local false-alarm mitigation and an in-memory design for real-time adaptability. Evaluated on 18 synthetic datasets with abrupt, incremental, and gradual drift, and eight real-world datasets, SCCM demonstrates improved predictive performance (R2, MSE) compared to eight baseline detector-adaptation methods across high-dimensional and large-scale settings.
concept driftonline regressiondynamic hyperparameter tuningkpi-window-based thresholdingin-memory design
XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?
The authors propose XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, and multidimensional evaluation of XAI explanation quality. The method assesses explanations along dimensions including simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability, benchmarking across datasets, ML models, and stakeholder personas. Human validation demonstrates a strong positive correlation (Spearman's rho=.693, p<.001) between LLM-generated and human ratings, indicating that LLM-based evaluations can systematically capture differences in XAI explanation quality.
llm-as-a-judgexai explanationstrust calibrationspearman's rhointerpretability
Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements
Edu-QuRating introduces a multi-dimensional educational data curation pipeline, extending QuRating with education-specific rubrics to assess text quality across six criteria. The method employs an LLM judge for pairwise document comparisons, distilling these into reusable Edu-QuRaters that score individual texts. Evaluations show the best Edu-QuRater achieves 0.917 mean accuracy in recovering GPT-4.1-mini judgements. Applications include pretraining mixture filtering (322.25M documents) and GRPO post-training, yielding improved model performance on pedagogical quality and instruction following compared to baselines like FineWeb-Edu and Qwen3-4B.
educational data curationpairwise judgementsllm judgeedu-quratergrpo post-training
Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration
The paper introduces Valerant, a training-free framework that converts pretrained action-conditioned world models (WAMs) into systems for generating navigable 3D game maps. The method integrates predictive visual rollouts with SLAM-based spatial reconstruction and exploration-driven action selection, enabling the transformation of a single image into a persistent 3D environment. This approach extends WAM-based interaction beyond 2D simulations, offering a novel solution for automated 3D game-map creation without manual design.
action-conditioned world models3d game mapsslam-based reconstructionpredictive visual rolloutsexploration-driven action selection
Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery
The paper introduces decision-focused active learning for optimizing critical-materials recovery processes, specifically in selective precipitation of rare-earth elements. The method employs a two-stage adaptive policy to minimize experiments needed to identify optimal recovery conditions, evaluated via conditional retrospective benchmarks using NdFeB and SmCo magnet recycling data. Results show adaptive policies achieve maximum enrichment in 16-24 experimental wells versus 48 for nonadaptive methods, with a hybrid batch selection strategy reducing downstream Bayes risk in simulations. Prospective validation requires standardized measurements and economic inputs.
active learningrare-earth recoveryselective precipitationbayes riskadaptive policy
Reliable Near-Field Multi-User Positioning Informed by Two-Stage MUSIC
Proposes MUSIC-Net, an end-to-end deep learning framework for near-field multi-user positioning that embeds two-stage MUltiple SIgnal Classification (MUSIC) to isolate line-of-sight (LoS) signal subspaces and identify surrogate distances, eliminating the need for non-LoS parameter estimation or path/source association. Introduces split conformal prediction (SCP) for statistically guaranteed confidence set estimation. Achieves lower mean positioning error (MPER) than benchmarks and tighter SCP-calibrated prediction regions, demonstrating accurate LoS localization and uncertainty quantification in coherent multi-path environments.
near-field localizationmultiple signal classificationsplit conformal predictionline-of-sightuncertainty quantification
An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI Frameworks
We introduce MMPIBench, a reproducible benchmark for evaluating multimodal prompt injection attacks on agentic AI frameworks. The benchmark tests six visual carriers (OCR text, overlays, EXIF metadata, QR codes, fake interfaces, hybrids) and extends to audio, measuring attack propagation from perception to tool calls. Across 720 runs involving six frameworks, five foundation models, and four attacker objectives, attacks complete in 1% of runs but are attempted in 12.8%, with planning steps mitigating most attempts. Model choice significantly impacts attack success, with one model recognizing injections in 59.7% of cases. Audio attacks complete in 49% of cases where supported, highlighting under-defended perceptual channels.
multimodal prompt injectionagentic ai frameworksmmpibenchvisual carriersperceptual channels
VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models
VANTAGE-Bench introduces a benchmark to evaluate the 'Infrastructure AI Gap' in Vision-Language Models (VLMs), focusing on fixed-camera applications across Logistics, Transportation, and Smart Spaces. It unifies image and video evaluation across semantic, spatial, temporal, and spatio-temporal capabilities, employing eight task formulations including dense captioning and spatio-temporal grounding. The benchmark includes 3,346 media assets with extensive annotations. Zero-shot evaluation of 17 models reveals concentrated shortfalls in event verification, referring expressions, and temporal localization (9-24 points below consumer-centric benchmarks), while video question answering remains within 5.3 points of VideoMME. Temporal capabilities are weakest, with no system exceeding 55.7 mIoU on temporal localization or 37.3 SODA_c on dense video captioning.
vision-language modelsinfrastructure aispatio-temporaldense captioningzero-shot evaluation
The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents
The paper introduces State-Path Tool Menu, a framework for constructing tool menus that serve as execution priors for online agents. The method learns a state path—a pre-execution route from the observable request state to the desired outcome—using an encoder to represent tool executability, output-input dependencies, and recurring order patterns. A retriever selects executable tools, missing-input producers, and the final action, while a reranker ensures producers precede consumers. On ToolBench, the approach improves online task success from 0.737 to 0.898, outperforms baselines, and achieves complete tool chains with 32 tools compared to 128 in official lists, with consistent gains across executor families.
tool menustate pathexecution prioronline agentstoolbench
An Autonomous GeoAI Agent for Arctic Eco-Navigation
The authors propose a human-in-the-loop, multi-agent GeoAI system for Arctic eco-navigation that integrates operational, physical, ecological, and community-related criteria into a unified routing framework. The system employs specialized agents for geospatial data acquisition, multi-objective route generation, and skyline-based decision support, explicitly accounting for ecological impacts on Essential Fish Habitat and seal critical habitat. This approach enables safer, transparent, and socially responsible Arctic navigation by balancing vessel efficiency with environmental and community considerations. The framework and code are publicly available.
geoaimulti-agent systemeco-navigationskyline-based decision supportessential fish habitat
Auditable Emergency Triage for Maternal and Newborn Care in India
The authors present a two-stage system for emergency triage in maternal and newborn care, combining a large language model (LLM) with a deterministic rule engine to improve interpretability and auditability. The LLM extracts canonical symptoms and patient context using a clinician-authored vocabulary, while the rule engine captures emergency scenarios. This decomposition increased recall from 0.565 to 0.810 and F1 score from 0.606 to 0.702, with structured rules driving most accuracy gains. The system enables clinical experts to inspect each stage for errors and add rules independently without costly evaluations. Deployed on 152,421 patient queries, it flagged 18.7% as emergencies with a 17.8% over-escalation rate and no increase in missed cases.
emergency triagelarge language modeldeterministic rule engineauditabilityrecall
Smart Adaptive Computing Across the Continuum: LLMs in IoT-Edge-Cloud Resource Management
We extend Wang et al.'s taxonomy of Continuum Orchestration Systems with two dimensions: AI Augmentation Paradigm, measuring LLM exploitation, and Feedback channel, capturing execution feedback paths to LLMs. Analyzing six architectures reveals a gap: none integrates full LLM orchestration with agent-layer feedback in Cloud Continuum settings, attributed to missing cross-tier feedback abstraction. This abstraction would bridge incommensurable per-tier signals and LLM Orchestrator, addressing resource management challenges across IoT-edge-cloud layers using DRL augmented by LLMs.
continuum orchestration systemsai augmentation paradigmfeedback channelllm orchestratorcloud continuum
Improving 5G AI-RAN MCS Selection by Predicting Retransmissions
NOSTRAdAMUS introduces predictive foresight to 5G NR Link Adaptation (LA) by forecasting retransmissions in the next radio frame, augmenting existing LA algorithms without redesign. Using Gradient Boosting on HARQ history, it achieves 82.9% accuracy (94.2% for high-confidence interventions) with 5.5 μs latency, trained on OTA data from X5G testbed with OpenAirInterface and NVIDIA Aerial. Deployed as a dApp, it boosts goodput by 71.5% and reduces retransmissions by 71.8% across 3GPP TDL/CDL channels, SISO/MIMO, and mobility scenarios without retraining.
5g nrlink adaptationgradient boostingharq feedbackmodulation and coding scheme
Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions
The paper proposes that gradients and Jacobians characterize phenomenal experience, tested in Gradland, an idealized differentiable neural network environment. It introduces two Jacobian structure measures: effective rank and cohesion, based on Kirchhoff complexity. These measures explain seven aspects of experience: duration (up to hundreds of milliseconds), vividness, texture perception, newborn confusion, distinct vs. confused ideas, learning processes, and the function of dense experience. The hypothesis is validated through worked examples in Gradland, demonstrating how first-order physical interactions shape experiential phenomena.
jacobian structurekirchhoff complexitygradlandphenomenal experiencedifferentiable neural networks
Support Discovery With Iteratively Reweighted Least Squares for Fixed-Charge Network Flow
A scalable continuous-optimization algorithm is proposed for large-scale single-commodity fixed-charge network flow problems (FCNFP), addressing computational challenges in mixed-integer linear programming formulations. The method employs an iteratively reweighted least-squares (IRLS) framework, replacing the discontinuous objective with a smooth nonconvex Lasry--Lions surrogate and solving weighted quadratic flow subproblems via a warm-started dual semismooth Newton method. An algorithmic variant incorporates objective-driven perturbation restarts and anchor-union restricted search to enhance arc support discovery. Evaluated on 410 instances, the method achieves a mean gap of 1.316% to a time-limited MILP reference and a win-or-tie rate of 90.0% among non-MILP methods, demonstrating effectiveness in producing high-quality solutions.
fixed-charge network flowiteratively reweighted least squareslasry--lions surrogatesemismooth newton methodarc support discovery
Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models
The study investigates gender bias in speech-to-speech (S2S) models by disentangling acoustic and content-based gender cues. Using a controlled experiment with male/female voices and stereotyped passages in English, Spanish, and Mandarin, the authors evaluate five models on voice rendering and gender attribution. Results show no stereotype drift in voice rendering, but all models prioritize content over voice for gender attribution, with odds of 'female' judgments increasing 1.7-24x per feminine content step. Misgendering reaches 90% in voice-content clashes versus 2% when aligned, revealing bias hidden in attribution.
speech-to-speechgender attributionstereotype driftacoustic biasmisgendering
Omni Interaction Agent Technical Report
Gander introduces an end-to-end model unifying omni perception, realtime interaction, and agentic capabilities within a single framework, enabling continuous full-duplex interaction across video, speech, and text. The architecture employs a Cerebellum-Brain collaborative framework, where the Cerebellum handles realtime interaction via a streaming Thinker-Talker design, while the Brain manages complex reasoning and agentic tasks. Evaluations show Gander matches SOTA open-source models in conversational ability while excelling in omni understanding, interactive robustness, and agentic intelligence under noisy, multi-party, or backchannel conditions.
omni perceptionfull-duplex interactioncerebellum-brain frameworkthinker-talker architectureagentic intelligence
DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable Connectivity
DiffLUT-Net introduces a novel FPGA-native neural network architecture with jointly trainable six-input lookup tables (LUTs) and sparse connectivity. The method employs differentiable relaxation of LUT functions and hardware-aware source selection to optimize both 64-entry truth tables and input port connections. Post-training, the network is discretized, pruned, and exported as synthesizable Verilog code. Evaluated across five benchmarks, DiffLUT-Net demonstrates superior accuracy-resource trade-offs compared to existing approaches. This work establishes the effectiveness of co-optimizing LUT functions and connectivity for compact FPGA inference.
fpgalookup tabledifferentiable trainingverilogsparse connectivity
Hi-FLoop: Hierarchical State-Feedback Loops for Multi-Timescale World Modeling
Hi-FLoop introduces a hierarchical state-feedback framework for multi-timescale world modeling in multi-agent traffic simulation, addressing cross-scale consistency and multimodal rollout challenges. The method employs eight scene-level Worlds for joint hypotheses, with Goal (8s), Preview (2s), and Control (1s) states adapting within a selected branch. A prefix-frozen A-to-B cascade handles generated-state recovery without latent state transfer. Evaluated on 955 H-D public-validation scenarios, Hi-FLoop achieves scene-joint ADE-at-joint-minFDE@8/joint-minFDE@8 of 2.048/6.384 m, with agent-centric oracle-minADE@8 of 0.526 m at 6s and 0.875 m at 8s.
multi-agent simulationhierarchical state-feedbackmulti-timescale modelingprefix-frozen cascadescene-joint ade
No Free Checker: A Survey of Verifiers for Robot Policies
The paper surveys approximately 150 verifiers for robot policies, categorizing them by judgment source (human, rule-based, learned, or model-intrinsic) and evaluating them along availability (cost, timing, frequency) and credibility (resistance to gaming). It finds an inverse relationship: credibility decreases as availability increases. The study identifies three validation metrics—human label agreement, trained policy performance, and reward hacking resilience—and proposes nine checkable metrics for future verifier development.
verifiersrobot policiesavailabilitycredibilityreward hacking
Likelihood-free inference with nuisance parameters through normalizing flows
The authors propose a neural-network-based normalizing flow decomposition that identifies near-pivotal statistics in the presence of nuisance parameters, requiring only a sample generator from the target distribution. The method minimizes average KL-divergence of p-values versus uniform and incorporates prior knowledge of group invariances like translation and scale. Results demonstrate that the approach discovers the one-sample t-test, outperforms the Welch test in worst-case size over constrained variance-ratio ranges, achieves good calibration on partial biserial correlations, and exhibits higher power and speed than profile likelihood-ratio techniques on small-to-moderate samples.
normalizing flownuisance parameterspivotal statistickl-divergencegroup invariances
A positive resolution of the gap-entropy conjecture
We resolve the gap-entropy conjecture for fixed-confidence best-arm identification in Gaussian bandits with independent unit-variance arms and a unique optimal arm. By analyzing the entropy of gap distributions and introducing a novel algorithm, we establish tight bounds on the optimal expected sample complexity. Specifically, the expected number of samples required is within constant factors of $H(\log(1/\delta)+\mathrm{Ent}(I))$, where $H$ represents the sum of inverse squared gaps and $\mathrm{Ent}(I)$ captures the entropy of the gap distribution. Our instance-independent algorithm achieves this bound plus an additive $g^{-2}\log\log(e^e/g)$ term, with $g$ being the minimal gap.
best-arm identificationgaussian banditsgap-entropysample complexityfixed-confidence
Characterizing Language Generation in the Limit: Finite Witnesses and a Separation-Width Hierarch
This work characterizes language generation in the limit for arbitrary families over countable universes, establishing that generation is possible precisely when each target language has a finite positive witness ensuring infinite common intersections across activated targets. The authors introduce a universal normalization procedure converting successful generators into history-independent ones and analyze witness size requirements through a separation-width hierarchy. They demonstrate that all hierarchy levels occur naturally, with countable families admitting singleton witnesses and explicit families realizing every finite width. The results, including a diagonal capture lemma and normalization simplification, are formally verified in Lean. Countable-support and finite-profile obstructions explain limitations of local combinatorial data.
language generationfinite witnessseparation-width hierarchycountable universelean verification
Optimal Low-Rank Quantum State Tomography with Bounded-Sample Joint Measurements
The study establishes the optimal sample complexity for low-rank quantum state tomography using bounded-sample joint measurements. For estimating a rank-$r$ state on $\mathbb{C}^d$ to trace norm error $\varepsilon$, the required sample complexity is $\Theta\left(\frac{dr}{\varepsilon^2} \max\left\{1,\frac{r}{\sqrt{t}}\right\}\right)$, where $t$ is the maximum number of samples per joint measurement. The lower bound permits adaptive measurement strategies, while the upper bound is achieved via a nonadaptive protocol based on Gaussian joint measurements. Joint measurements on $r^2$ samples are shown to be necessary and sufficient for achieving the collective rate. The analysis leverages Fisher information bounds and Gaussian measurement techniques.
quantum state tomographysample complexityjoint measurementstrace normfisher information
Quantum Feature Engineering for Credit Default Prediction: When and Why IQP Circuits Help Linear Classifiers
The study demonstrates that Instantaneous Quantum Polynomial-time (IQP) circuits can enhance credit default prediction by generating features that improve Logistic Regression performance. Using the UCI Default of Credit Card Clients dataset, 16 IQP features (8 qubits) appended to Logistic Regression increased F1 score from 0.462 to 0.517, outperforming Kernel PCA (F1 = 0.493) and other classifiers. The method encodes financial attributes as rotation angles, producing expectation values as features. Feature selection significantly impacts performance, with Random Forest-guided selection achieving F1 = 0.523, while uncorrelated features drop to 0.496, indicating IQP circuits amplify existing informative structure.
iqp circuitslogistic regressionkernel pcafeature engineeringcredit default prediction
Cross-Model Agreement as a Deployment-Time Reliability Signal for Automatic Polyp Segmentation
We propose Referee-Based Quality Estimation (RBQE), a reference-free framework for assessing reliability in automated polyp segmentation by measuring agreement between a primary segmentation model and an independently trained referee. Evaluated on a 1,223-image benchmark from four public datasets, RBQE achieves ROC-AUC scores of 0.923 with same-architecture referees and 0.960 with cross-architecture SegFormer-B0 referees, outperforming Test-Time Augmentation by 0.055 ROC-AUC. Excluding empty-mask cases, ROC-AUC drops to 0.876 (SegFormer-B0) and 0.783 (same-architecture), but RBQE's margin over baselines widens. RBQE also increases mean Dice scores as low-agreement predictions are rejected, requiring only one additional referee forward pass at inference.
polyp segmentationcross-model agreementreferee-based quality estimationroc-auctest-time augmentation
Learning with Covariance Matrices: Principal Component Analysis Meets Learning with Graphs
The article introduces theoretical foundations for covariance neural networks (VNNs), a class of graph neural networks (GNNs) operating on covariance matrices as graphs. It establishes conceptual equivalence between VNNs and principal component analysis (PCA), refines stability bounds for predictive outcomes under covariance matrix perturbations, and characterizes VNN transferability across multiscale datasets. These insights justify VNN adoption over PCA-based pipelines in applications leveraging covariance matrices, such as neuroimaging for neurodegenerative conditions. The analysis bridges signal processing and graph-based learning, offering principled design guidelines for domains where covariance matrices describe data structure.
covariance neural networksgraph neural networksprincipal component analysisstability boundstransferability
Nonmaximal sums of maximally monotone operators under Rockafellar's constraint qualification
(No summary returned.)
Deep Learning-Based Detection of Electrical Faults and Power Quality Disturbances in Aerospace Power Systems
A hardware-aware deep learning framework is proposed for multiclass detection of electrical faults and power quality disturbances in 400 Hz aerospace power systems. High-fidelity simulations generate voltage and current waveforms for 21 conditions, processed into time-series and short-time Fourier transform datasets. Signal augmentation, domain randomization, and generative adversarial networks enhance waveform diversity. Evaluations of 1D/2D CNNs, LSTMs, CNN-LSTM hybrids, ResNet, MobileNet, and VGG models reveal a compact ResNet achieves 96.94% accuracy with 175,685 parameters. Quantized deployment on a Xilinx Zynq UltraScale Plus MPSoC yields 95.87% accuracy and 6.90 ms latency, demonstrating feasibility for embedded edge AI in aircraft electrical health monitoring.
aerospace power systemsshort-time fourier transformgenerative adversarial networksquantizationembedded edge ai
Multi-Agent Reinforcement Learning for Autonomous UAV Exploration in Wildfire Response
The study proposes a deep reinforcement learning framework for autonomous wildfire monitoring using Unmanned Aerial Vehicle (UAV) agents. The method involves training UAVs in simulated wildfire environments, focusing on navigation and monitoring tasks. Results indicate that agents achieve stable and effective behaviors over time, evidenced by converging loss trends, improved reward signals, and consistent navigation patterns such as fire-boundary tracking. The findings underscore the influence of environmental structure and reward design on policy effectiveness in DRL-based UAV systems.
deep reinforcement learningunmanned aerial vehiclewildfire monitoringreward designpolicy effectiveness
Algorithmic stability via ensembling
We present a general framework for quantifying algorithmic stability guarantees achieved through ensembling strategies based on averaging. The method analyzes stability by deriving bounds in terms of the norm of a covariance operator characterizing the ensembling process. Theoretical results demonstrate that this framework provides sharper stability guarantees than those derived from privacy considerations, while yielding interpretable insights across various practical perturbation types. The approach establishes connections between ensembling and algorithmic stability for diverse data perturbation scenarios.
algorithmic stabilityensemblingcovariance operatordata perturbationaveraging
HybridFLow: SDN-Orchestrated Client Partitioning for Hybrid Federated Learning
HybridFLow introduces a Software-Defined Networking (SDN)-orchestrated framework for optimizing client partitioning in hybrid Federated Learning (FL), addressing communication delays in wide-area deployments. The method leverages the SDN controller's global network topology view to estimate per-client communication times, dynamically partitioning clients into synchronous and asynchronous groups to balance round latency and update staleness. Feedback from measured communication times refines future predictions. Experiments demonstrate that HybridFLow achieves 80% target accuracy 33-40% faster than SmartFLow, reduces average round duration by 30-40 seconds, and outperforms FedAsync in non-IID data scenarios.
federated learningsoftware-defined networkingclient partitioningcommunication delaysnon-iid data
Searching for New Physics with Reinforcement Learning
We introduce a reinforcement learning (RL) method for identifying Standard Model Effective Field Theory (SMEFT) operators that explain anomalies in particle physics. The approach addresses challenges posed by the vast number of SMEFT operators and their complex correlations at loop level, overcoming biases inherent in human phenomenological intuition. The method is tested on the CDF W-mass anomaly, reproducing and improving upon known results, and further validated in a complex scenario with multiple anomalies. Results demonstrate the RL method's efficacy in efficiently searching for new physics within the SMEFT framework.
reinforcement learningsmeft operatorsparticle physicsanomaliesloop level
A Later Test Set Is Not a New Domain: Pretraining Familiarity Survives a Contamination-Free Hold-Out
The study introduces a temporal hold-out protocol to evaluate time-series foundation models, addressing pretraining contamination by testing on data published after model release. Thirteen forecasters (four classical, three dataset-specific, six pretrained) were evaluated on seven domain groups, with pretrained models outperforming classical methods in five groups. Key findings reveal that intrinsic properties like seasonal strength and spectral entropy do not explain performance differences; instead, corpus familiarity drives success. TimesFM outperformed Chronos significantly on Wikipedia pageviews, its primary pretraining domain. The results emphasize the need for domain-specific benchmarks relative to disclosed corpora and highlight the importance of domain familiarity in model selection.
temporal hold-outpretraining contaminationcorpus familiarityseasonal strengthspectral entropy
TimeCues Studio: A Workspace for Music Annotation and Algorithm Prototyping
TimeCues Studio introduces an open-source workspace for music annotation and algorithm prototyping, addressing the scarcity of annotated training data for multimedia applications. The tool enables teams to annotate entire music collections, compare detection algorithms against annotations, and prototype new models on a grid-locked timeline visualizing multiple music features, including separated audio stems. It supports ambiguity-aware labeling, integrates a Python sandbox for model development, and includes an algorithm-comparison engine with bundled baselines. TimeCues Studio is MIT-licensed and deployable via Docker Compose, catering to both team-based annotation and solo music-sync projects.
music annotationalgorithm prototypingambiguity-aware labelingaudio stemsdocker compose
View-Structured Conformal Prediction for 3D Gaussian Splatting
The authors introduce View-Structured Conformal Prediction (VSCP), a method for uncertainty quantification in 3D Gaussian Splatting (3DGS) novel-view synthesis. VSCP decomposes the pre-calibration scale into a spatial shape and a transferable view-difficulty factor, enabling finite-sample validity across scenes via a held-out quantile over views. The approach achieves exact analysis through conformity scores and separates excess width into test-side and calibration-side terms. Evaluated on 13 real scenes, VSCP attains 91.7–92.0% view-event coverage at a 90% target, reducing width by 22.1% compared to constant scaling and matching ensemble performance with fewer resources. It also outperforms the 3DGS-U field by 4.7 points and maintains efficiency at 216–280 FPS on an RTX 4090.
3d gaussian splattingconformal predictionnovel-view synthesisuncertainty quantificationview-difficulty factor
A Dominant Diffuse Phase in the Sparse Autoencoder Phase Diagram
This study investigates feature recovery in sparse autoencoders (SAEs) through the MAIS-O43 open problem, focusing on the interplay between nesting fraction (γ), sparsity penalty (λ), and dictionary size (M). The authors conducted 200 independent fits across ten grid cells, followed by 3,300 additional fits using minibatch Adam across a 165-cell grid. Results reveal a dominant diffuse phase: perfect reconstruction but poor feature recovery (median best cosine 0.5-0.7 vs. 0.95 criterion) and denser codes than ground truth. This suggests trained SAEs may not reach objective minima, indicating a fundamental divergence between trained models and objective minimizers.
sparse autoencodersfeature recoverydiffuse phasenesting fractionsparsity penalty
The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
Brain2Semantics2Text introduces a semantic bottleneck for non-invasive speech decoding, leveraging high-level semantic representations to reconstruct text from MEG recordings. The method maps sentence-level neural responses into a semantic embedding space before inversion to natural language, avoiding word-level alignment challenges. Compared to prior Brain2Text approaches, it demonstrates improved sentence-level reconstruction by exploiting cortical semantic representations' distributed and slower-evolving properties.
speech decodingsemantic embeddingmegneural representationsbrain2text
Training Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers
The study demonstrates that training trajectory determines whether learned retrieval circuits remain functional after removing soft positional priors in Transformers. Using an annealable soft-prior Transformer with adjustable attention biases, the authors evaluate associative recall and Markov induction tasks. Models perform well with active priors (0.772 ± 0.020) but collapse without them (0.095 ± 0.009); smooth fade-to-zero training preserves zero-gate accuracy (0.734 ± 0.028), unlike forced-zero or hard-switching approaches. Mechanistic analysis reveals circuit consolidation occurs post-gate removal, with head variability across seeds. Results indicate trajectory-dependent circuit removability in small discrete tasks.
soft positional priorsannealable transformerassociative recallcircuit consolidationtraining trajectory
Structural Fusion of Bayesian Networks with Limited Treewidth Using Genetic Algorithms
The paper proposes a genetic algorithm for structural fusion of Bayesian Networks (BNs) under treewidth constraints, enabling consensus BN generation from multiple input networks while maintaining computational tractability. The method evolves BNs that preserve key structural features from the input networks while adhering to a specified treewidth limit, ensuring efficient inference. Experimental results demonstrate the algorithm's effectiveness in producing consensus BNs with restricted treewidth, facilitating information aggregation from diverse sources into computationally feasible models.
bayesian networkstreewidthgenetic algorithmstructural fusionconsensus
Maverick: Private and Verifiable LLM Inference Made Practical via Matrix-Vector Multiplication Delegation
Maverick introduces a novel protocol for private and verifiable LLM inference via matrix-vector multiplication delegation, addressing privacy and correctness concerns without significant server overhead. The approach combines an information-theoretically sound verification protocol with LPN-based pseudorandom masking to ensure input privacy. Evaluated on Qwen3-4B, Maverick achieves throughput gains of up to 17x for private inference, 45x with precomputed masks, and 44x for verification-only scenarios, demonstrating scalability across client configurations.
matrix-vector multiplicationlpn-based maskingverification protocolllm inferencethroughput gains
Through the Looking Glass: Directly Reading and Writing Transformers
The study analyzes transformer components' contributions to token predictions through direct parameter and activation examination without training or fitting. It quantifies that 53 components carry 90% of a prediction, with 8 sufficient for production alone across models ranging from 124M to 7B parameters. Findings reveal that predictions rely on 1-3% of the model, a proportion invariant to model size, and that 75% of a layer's update is a fixed linear map of its input state. The method enables component-level interventions, demonstrating that installing new associations in spare units incurs minimal loss (0.25%) compared to rank-one updates.
transformertoken predictionlinear mapactivationrank-one update
Robust Beam Prediction for V2X Networks with Multi-Modal Sensing
We propose BeamTransFuser, a hierarchical Transformer-based architecture for robust beam prediction in V2X networks, leveraging multi-modal sensor fusion of camera, LiDAR, radar, and GPS data. The framework includes a generative module to reconstruct missing modality features, enhancing robustness in incomplete sensing scenarios. Evaluations on a real-world multi-modal V2X dataset demonstrate consistent performance improvements over baseline methods, particularly in handling modality deficiencies.
beam predictionv2x networksmulti-modal fusiontransformer architecturesensor fusion
An Exponential Deterministic--Randomized Gap in ERM-Oracle Complexity for Thresholds on an Unknown Order
The paper establishes an exponential deterministic-randomized gap in ERM-oracle complexity for transductive online learning of thresholds on an unknown total order. Using a minimal-prefix rule oracle, deterministic learners require Ω(T) calls for O(log T) mistakes, while a randomized learner achieves O(log T) expected calls and mistakes. The separation depends on the oracle's selection rule: a feasible-median ERM rule permits O(log T) calls for deterministic learners, whereas a global-median rule imposes linear cost. The randomized order is shown optimal via a hard distribution requiring Ω(log T) expected calls for polylogarithmic mistakes. Partial tradeoff results for fixed query budgets and a weak consistency oracle interface are also provided.
erm-oracle complexitytransductive online learningminimal-prefix rulefeasible-median ermhard distribution
Are You Learning Biological Signal or Shortcuts? Auditing and Mitigating Bias in Protein-Protein Interaction Datasets
The study systematically characterizes biases in protein-protein interaction (PPI) datasets that lead machine learning models to learn shortcuts instead of biological signals. It analyzes HIPPIE, IntAct, STRING, and PDB-derived datasets, identifying topological shortcuts from random data splitting and persistent biases from self-interactions, taxonomic identity, and functional relatedness. The authors propose an open Nextflow pipeline combining similarity-aware dataset splitting and optimization-based negative sampling, formulated as integer linear programs, to mitigate these biases. This approach, applicable beyond PPI prediction, minimizes data loss while reducing dataset-specific shortcuts.
protein-protein interactionshortcut learninginteger linear programmingnegative samplingtopological bias
CoGe-GCD: Reframing Generalized Category Discovery with Compositional Generalization
CoGe-GCD introduces a novel approach to Generalized Category Discovery (GCD) by reframing it through compositional generalization, addressing the challenge of extrapolating to novel compositions. The method comprises two stages: (i) Compositional Perception, which structures patch tokens via primitive mapping and competitive token-primitive assignment, and (ii) Generalizing Induction, which calibrates spatial relations to preserve geometric structure. Implemented as an inductive-bias module, CoGe-GCD enhances all-class accuracy, unknown-class number estimation, and geometric quality on standard benchmarks with minimal computational overhead.
compositional generalizationgeneralized category discoveryinductive-bias moduletoken-primitive assignmentspatial relations
CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts
CompassOPD improves cross-family on-policy distillation (OPD) by isolating within-family likelihood shifts from inter-family capability offsets. The method anchors updates to a frozen student reference policy while transferring only the teacher's within-family log-likelihood improvements, preventing capability offsets from dominating the update direction. Experiments across three student families and multiple teacher families demonstrate consistent improvements over standard OPD, with reasoning accuracy gains of up to 5.50 points. For Mixture of Experts (MoE) teachers, CompassOPD constructs the reference directly by reducing expert activation, achieving a 3.43-point gain without requiring separate reference checkpoints.
on-policy distillationcross-family transferlog-likelihood shiftmixture of expertsreasoning accuracy
The Sample Complexity of Quantum Entanglement Allocation
This work characterizes the sample complexity of quantum entanglement allocation, demonstrating that memory size influences data requirements for qubit entanglement decisions. Using a fixed detector preserving coherence, the study analyzes independent commuting Pauli queries, deriving minimax excess error bounds proportional to $k^{-1}\min\{1,\sqrt{d\log(k+1)/m}\}$ for $d$-qubit paths and groups of $k$ qubits. Results show connected biclique regions can expand without increasing sample demand under bounded depth and connectivity. Experiments validate statistical predictions on a 15-qubit device and public purchase baskets, with frequency grouping outperforming basket search in high-capacity settings.
quantum entanglementminimax excess errorpauli queriesbiclique regionssample complexity
Storage-Scalable Progressive Semantic Communication via Knowledge-Base Reuse
The authors propose storage-scalable knowledge-base reuse quantization (SSKBQ) to address the storage overhead in semantic communication systems. SSKBQ reuses a compact set of knowledge bases (KBs) across multiple residual refinement stages, decoupling the number of transmission stages from the number of maintained KBs. A stage-aware residual supervision mechanism regularizes intermediate quantized representations to ensure progressive refinement. Experiments show that SSKBQ maintains competitive progressive reconstruction performance while effectively solving the storage scalability problem, outperforming single knowledge-base quantization (SKBQ) and multi-knowledge-base residual quantization (MKBQ) approaches.
semantic communicationknowledge-base reuseresidual quantizationprogressive refinementstorage scalability
A Systematic Evaluation of Molecule Generation Models for De Novo Drug Design: From Benchmarks to Practical Insights
This review systematically evaluates 82 molecule generation models for de novo drug design across five deep generative frameworks: RNNs, Transformers, VAEs, GANs, flow-based models, and diffusion models. It synthesizes benchmarks, molecular representations, and methodological principles for general and pocket-conditioned generation, providing a comparative analysis of reported performance metrics. The study highlights experimentally validated case studies and discusses future directions, including standardized 3D data, interaction-aware generation, receptor flexibility, and multi-objective molecular design. All resources are available in a public repository.
molecule generationde novo drug designdeep generative frameworksbenchmarksmolecular representations
Hybrid Quantum-Classical NLP Classification with Compact Semantic Representations: An Experimental Analysis of Representation Compression
The study introduces a hybrid quantum-classical pipeline for NLP classification, addressing the challenge of high-dimensional sentence embeddings in near-term quantum machine learning. The method combines pretrained sentence embeddings, dimensionality reduction (PCA, NCA, LDA), angle encoding, a variational quantum circuit, and a classical decision layer. Experiments on the TREC dataset show that supervised dimensionality reduction (LDA, NCA) achieves higher accuracy (85.3%, 83.1%) with 5 dimensions compared to unsupervised PCA (63.4% with 8 dimensions), demonstrating superior task-relevant information retention. These results suggest supervised compression is more effective for practical hybrid quantum-classical NLP models.
quantum machine learningdimensionality reductionvariational quantum circuitsentence embeddingsangle encoding
Physics-Informed Multi-Task Surrogate Model for the Martian Nightside Thermosphere
A physics-informed multi-task neural network is proposed for modeling the Martian nightside thermosphere, addressing challenges of sparse sampling and non-physical artifacts in data-driven approaches. The model predicts base-10 logarithmic densities of four neutral species (O, CO$_2$, N$_2$, Ar) using MAVEN/NGIMS observations from MY 32-38 (2014-2025). It employs a shared backbone for common thermospheric state representation, species-specific output heads, and a weak monotonicity prior enforced via automatic differentiation to penalize positive vertical gradients. Experiments demonstrate reduced non-physical inversions while maintaining predictive accuracy, as measured by RMSE, MAE, and $R^2$, yielding a computationally efficient surrogate with improved vertical consistency.
physics-informed neural networkmartian thermospheremulti-task learningmonotonicity priorautomatic differentiation
Orukeet: Multilingual ASR with Frozen Gabor Kernels
Orukeet introduces a multilingual ASR system that replaces half of Parakeet's temporal filters with 12,288 frozen Gabor kernels, retaining the original architecture. The remaining parameters are trained on multilingual and multi-accent data, with final adaptation using LibriSpeech test-other. Evaluated on 20,146 FLEURS recordings across 25 languages, Orukeet achieves a pooled WER of 9.85%, a 10.6% relative reduction compared to Parakeet's 11.01%. It outperforms Parakeet on 23 of 25 languages and 61 of 74 tested splits, including LibriSpeech test-clean (1.46% vs. 1.53%) and FLEURS English (3.82% vs. 4.28%). The Gabor kernels are stored as standard convolution weights, preserving inference operators.
gabor kernelsmultilingual asrword error ratetemporal filterslibrispeech
Zero-Shot Temporal Localisation of Audio Deepfakes in Multi-Speaker Conversations
We introduce a zero-shot pipeline for Temporal Deepfake Localisation in Multi-Speaker Conversations (TDLMC), addressing the challenge of detecting synthetic speech segments within genuine conversations. The method employs a five-stage pipeline that wraps a frozen binary deepfake detector, utilizing a two-threshold hysteresis finite-state-machine decoder to convert noisy window scores into coherent intervals without retraining. Evaluated on 180 multi-speaker conversations from ASVspoof 5, the system achieves a temporal intersection-over-union of 0.90, a temporal detection rate of 0.95, and an MS-DCF of 0.26, with a false-alarm rate below 6% on genuine speech. The pipeline establishes a reusable benchmark for TDLMC, demonstrating minimal performance loss compared to supervised approaches.
temporal deepfake localisationzero-shot pipelinefinite-state-machine decodermulti-speaker conversationsasvspoof
MedDeID enables locally governed clinical-text de-identification from real or synthetic training data
MedDeID introduces an on-premises framework for clinical-text de-identification, enabling local governance through in-house annotation, synthetic-note generation, and model training. The system employs compact transformers trained on hospital or synthetic data, achieving 98.9% identifier detection on a Dutch hospital benchmark while redacting only 0.24% non-identifier text. Synthetic-trained models demonstrated higher recall (90.3%) and robustness compared to hospital-trained counterparts (87.0%) on primary-care notes. An English instantiation achieved 99.7% and 98.9% identifier detection on synthetic benchmarks, showcasing cross-language transferability. Results validate MedDeID's efficacy for locally governed de-identification using real or synthetic data.
de-identificationsynthetic-note generationcompact transformerson-premises frameworkidentifier detection
Field-level prediction of mid-plane stress tensor fields in concrete target penetration: a cross-velocity graph neural operator surrogate
A graph neural operator surrogate was developed for field-level prediction of mid-plane stress tensor fields in concrete target penetration, addressing the gap in linking mesoscale heterogeneity to full-field stress-tensor prediction. The method utilized a full-scale aggregate-resolved LS-DYNA model, verified against published penetration experiments, to generate a dataset of six-component stress-tensor fields for 400 cases across four impact velocities. The surrogate achieved a single-step relative L2 error of 0.6977 and demonstrated autoregressive rollout speeds of 144 ms, yielding a 3.6x10^3 to 4.3x10^3 speedup over single-core LS-DYNA. Validation was bounded by configuration similarity and field-level self-consistency, establishing the framework as a simulation-trained decision-support method within the studied parameter space.
graph neural operatorstress tensorls-dynamesoscale heterogeneityautoregressive rollout
Dynamical Non-compensatory Multidimensional IRT Model Using Variational Approximation
We propose a dynamical extension of non-compensatory multidimensional item response theory (MIRT) models to accurately trace evolving latent skills under non-compensatory assumptions. The model combines a linear dynamical system with a non-compensatory MIRT framework, approximating the complex posterior of skills via Gaussian variational approximation by minimizing Kullback-Leibler divergence. Parameter learning employs Monte Carlo expectation maximization. Simulation studies demonstrate superior accuracy in reproducing latent skills compared to dynamical compensatory models, which exhibit significant underestimation errors. Empirical experiments confirm practical skill tracing capabilities and highlight differences between non-compensatory and compensatory model inferences.
multidimensional item response theorynon-compensatory modelvariational approximationlinear dynamical systemskill tracing
Beyond Contact Sensors: Deep learning with Pseudo-Labeling for remote Photoplethysmography
The study demonstrates that pseudo-labels generated by unsupervised signal-processing methods can effectively replace contact sensor labels for training deep learning models in remote photoplethysmography (rPPG). The authors systematically evaluate the approach under varying synchronization conditions between video data and ground truth signals. Results show that pseudo-labeling outperforms supervised training on datasets with imperfect synchronization, while performance is mixed for well-synchronized datasets. Cross-dataset evaluation favors supervised training, but removing a single outlier participant significantly improves pseudo-label performance. This reduces dependency on labor-intensive labeled dataset collection while maintaining competitive accuracy.
remote photoplethysmographypseudo-labelingunsupervised learningsignal-processingcross-dataset evaluation
Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS
The work introduces a deterministic prompting approach to stabilize speaker identity in low-resource Greek TTS, addressing degradation in prompt-based models like Parler-TTS (880M) when using LLM-generated style prompts. The method combines WhisperX-aligned audiobook data curation, deterministic prompts, and a speaker-specific LoRA (updating ~5% of parameters) trained on 3.5h of single-speaker data. Results show WER 10.7% (2.9 above ASR floor), MOS-I 4.00 (vs. 4.36 human), and near-human speaker consistency (MOS-C 4.24 vs. 4.30), demonstrating viable single-speaker synthesis with limited data.
text-to-speechlow-resourcedeterministic promptingloraspeaker consistency
Structure-Aware Unsupervised Anomaly Detection for Spacecraft Telemetry with Adaptive EVT Thresholding
The paper proposes an unsupervised anomaly detection framework for spacecraft telemetry that operates without labeled anomalies, fault knowledge, or mission-specific tuning. The method combines incremental monthly retraining, statistical model selection, and adaptive Extreme Value Theory (EVT) thresholding to control false alarms, achieving deployment readiness by the second month of operation. Evaluated on the ESA Anomalies Dataset (ESA-AD), it scores F₀.₅=0.700 on Mission 1 and F₀.₅=0.698 on Mission 2 under strict chronological evaluation.
unsupervised learninganomaly detectionspacecraft telemetryextreme value theoryincremental retraining
An Explainable Machine Learning Framework for Predicting Blood-Brain Barrier Permeability Using Molecular Descriptors
An explainable machine learning framework was developed to predict blood-brain barrier (BBB) permeability using molecular descriptors from the MoleculeNet BBBP dataset. Fifteen physicochemical descriptors from 2,039 compounds were used to train Logistic Regression, SVM, Random Forest, and XGBoost models, with hyperparameter optimization via GridSearchCV and interpretability analysis using SHAP. The optimized XGBoost classifier achieved 88.97% accuracy, 93.13% F1-score, and 0.9282 ROC-AUC, with stratified five-fold cross-validation yielding a mean ROC-AUC of 0.8982 ± 0.0130. SHAP analysis identified TPSA, HBD, and LogP as the most influential descriptors. The framework offers an accurate, interpretable approach for early-stage CNS drug screening.
blood-brain barriermolecular descriptorsshapxgboostgridsearchcv
A Sharp Barrier for Consistent Submodular Maximization: Any Improvement over $2-\sqrt{2}$ Entails Exponential Queries or Linear Recourse
The work establishes a sharp computational barrier for consistent submodular maximization, proving that no polynomial-time algorithm can achieve an approximation ratio better than $2-\sqrt{2} \approx 0.5858$ with constant recourse and polynomially many value queries. Using randomized techniques, the authors develop an algorithm attaining $\beta-\varepsilon$ approximation with $O(\varepsilon^{-2})$ changes per insertion, while showing that any fixed improvement requires either exponential queries or linear recourse. The results precisely characterize the tradeoff between consistency and optimality, with extensions to curvature-dependent thresholds and weighted coverage functions. The lower bounds hold even with unlimited post-critical queries, highlighting the inherent cost of maintaining consistency under element arrivals.
submodular maximizationapproximation ratioconstant recoursevalue queriescurvature-dependent threshold
Optimal Value Inference for Reinforcement Learning
The paper introduces a debiased estimator for optimal value inference in reinforcement learning, addressing two new nuisances derived as fixed points of a self-induced Bellman equation. The method approximates the maximum Bellman operator via softmax correspondence and leverages Neyman orthogonality to ensure asymptotic normality under diverging horizons, even with time-varying behavior policies. The proposed nuisance estimators, achievable by standard machine learning methods, are validated through synthetic experiments and real-world applications like bike repositioning and AI agentic tool use.
reinforcement learningoptimal value inferencebellman equationneyman orthogonalityasymptotic normality
Deep Neural Networks for Learning Intent from sEMG Signals to Support Hardware Devices for Post-Stroke Neurorehabilitation
The study introduces a deep learning framework for decoding five-finger motor intent from high-density surface electromyography (sEMG) signals in stroke patients, targeting post-stroke neurorehabilitation hardware. The method processes sEMG via Butterworth filtering, wavelet denoising, and segmented feature extraction, comparing LSTM, CNN, and GNN baselines. A CNN-Large variant achieves 0.593 subset accuracy and 0.714 macro F1, while a compact CNN-Micro (123K parameters) optimized via cross-channel knowledge distillation attains 0.5219 subset accuracy and 0.6095 macro F1 with four sensor inputs. The model, exported to ONNX, enables hardware-compatible intent prediction.
surface electromyographyneurorehabilitationknowledge distillationmotor intent decodingcnn-micro
Vague2Detect: Handling Ambiguous Prompts in Knowledge-Based Open-World Detection
Vague2Detect improves open-world object detection for ambiguous prompts by combining a fine-tuned Sentence-BERT with YOLO-World and dynamic KB expansion via GPT-3.5-turbo. The method retrieves candidates from a structured household KB, verifies them visually, and generates novel descriptions for out-of-KB queries. On a custom benchmark, Vague2Detect achieves 61% Vague Prompt Success Rate (VPSR) versus YOLO-World's 32%, rising to 85% with GPT fallback.
open-world detectionknowledge base retrievalsentence-bertvague prompt groundingdynamic expansion
Adversarial Training for Tabular Credit Scoring: A Multi-Attack Robustness Evaluation in P2P Lending
The study evaluates adversarial robustness in tabular credit scoring for P2P lending, addressing the gap in multi-attack generalization beyond single-attack defenses. Using a Lending Club dataset, it benchmarks three model families (logistic regression, feed-forward neural network, transformer) against four attacks (FGSM, PGD, S&P noise, DeepFool) and a mixed-attack regime. Results show adversarial training improves robustness against trained attacks and transfers well within gradient-based attacks but poorly to non-gradient corruption, while mixed training achieves balanced robustness across heterogeneous attacks without degrading clean-test performance.
adversarial trainingtabular datapeer-to-peer lendingmulti-attack robustnessgradient-based attacks
Multi-Pass, Multi-View Blended Learning for High-Fidelity Volumetric CT Synthesis from Chest X-Rays
The study introduces a Multi-Pass Multi-View Blended Learning framework for synthesizing high-fidelity volumetric CT from chest X-rays (CXRs), addressing the ill-posed inverse problem and synthetic-to-real domain gap. The method decomposes the task into two stages: unsupervised CXR-to-DRR domain adaptation and a three-pass stage (supervised DRR-to-CT transformation, unsupervised multi-view slice refinement, and progressive transfer learning). Evaluated on the LIDC-IDRI dataset, the approach improves PSNR by 14% and SSIM by 7.6% over prior methods, yielding structurally consistent and anatomically realistic CT volumes from real CXRs.
volumetric ct synthesisdomain adaptationmulti-view refinementprogressive transfer learningdigitally reconstructed radiographs
Development and Validation of a Physics-Guided Machine Learning Extrapolation Framework Using a Classical Transient Diffusion Benchmark
The study introduces a physics-guided machine learning framework for reliable extrapolation beyond training domains, addressing a key challenge in engineering applications where data outside operational ranges are scarce. The method integrates BiLSTM and PINN architectures with physics-guided coordinate transformations, boundary-aware learning, and stability-enhancing temporal marching, validated on a 1D transient diffusion benchmark with exact analytical solutions. Results show accurate, physically consistent predictions when extrapolating backward from intermediate training data toward singular initial conditions, using a recursive train-predict-validate-extend strategy.
extrapolation frameworkphysics-guided mlbidirectional lstmtransient diffusionboundary-aware learning
A Kernel-Based Modular Discriminant Analysis Framework for Small-Sample Learning
The paper presents a systematic analysis of Kernelized Linear Principal Component Discriminant Analysis (KLPCDA), a modular framework integrating variance preservation, inter-class separability, and intra-class compactness in kernel space for small-sample learning. Through cross-domain experiments on hyperspectral imaging, fault diagnosis, medical diagnosis, and face recognition, the study characterizes interactions among KLPCDA's three core objectives under noise, imbalance, and high dimensionality. Results reveal consistent patterns in objective interactions, enabling practical variant selection guidelines, with KLPCDA demonstrating robust performance and low computational complexity across domains.
kernelized discriminant analysissmall-sample learninginter-class separabilityintra-class compactnesshigh-dimensional data
Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits
The paper introduces Meta-LinEXP3, an online-within-online algorithm for adversarial linear contextual bandits (ALCBs) with random action sets, addressing the understudied problem of meta-learning in this setting. The method constructs a task-level prior from completed tasks to guide the inner LinEXP3 learner, employing a policy-centered estimator for known context distributions (achieving O(√n) per-task regret) and a past-only regularized moment estimator for unknown distributions (with O(n^(2/3)) leading regret). Theoretical analysis links prior accuracy to transfer regret, demonstrating sublinear transfer-dependent regret. Experiments validate the approach, including applications to hyperspectral tensor sampling.
meta-learningadversarial banditscontextual banditsregret boundsonline learning
Beyond Conventional Federated Learning via High-Order Regularization
The paper introduces HiFedProx, a federated learning method replacing FedProx's quadratic regularization with a power-type regularizer (p≥2) to better control client parameter displacement disparities. HiFedProx uses scale-matched regularization, where p>2 weakens responses below a reference displacement R and strengthens them above R, compressing disparities while avoiding excessive curvature. Experiments on a 60-writer FEMNIST subset show HiFedProx with p=6–7 reduces moderate- and severe-stress losses by 11.44% and 23.16% over p=2, though performance peaks at intermediate p values (5–7) due to increasing Armijo trial costs.
federated learningregularizationparameter displacementarmijo backtrackingfemnsit
ProMeta: Few-shot PROTAC-targeted degradation prediction across E3 ligases
ProMeta introduces a few-shot meta-learning approach for PROTAC degradation activity prediction across E3 ligases, addressing data scarcity and imbalance. The method employs a prototype-based graph neural network trained via episodic meta-learning on source-E3 tasks, enabling support-conditioned inference on target-E3 tasks without encoder updates. On the CRBN-to-VHL benchmark, ProMeta achieves AUROC values of 0.796 (K=2, Q=3) and 0.883 (K=2, Q=5), outperforming supervised GNN baselines by 19.9% and 6.8%, respectively. Reverse VHL-to-CRBN transfer yields AUROCs of 0.702 and 0.821 under the same protocols, demonstrating bidirectional applicability with direction-dependent performance.
protacmeta-learninggraph neural networke3 ligasesfew-shot learning
Exact Degeneracy Under Balanced k-Shot Sampling:Consequences for Small-Sample Discriminant Analysis on LLM Embeddings
The paper identifies an exact degeneracy in Kernelized Linear Principal Component Discriminant Analysis (KLPCDA) under balanced k-shot sampling, proving that two variants exhibit indifferent eigenvector selection and a third has a void objective due to within-class scatter becoming a scaled orthogonal projector. The authors derive closed-form consequences, propose an in-formula tie-break for repairable variants, and validate findings on frozen LLM embeddings and residual-stream activations. Repaired KLPCDA underperforms logistic regression probes on 3/4 text classification tasks (n≪d, up to d=4096), with geometric separability metrics ruling out anisotropy as the cause; performance gaps diminish with larger support sets (k=30-50 vs. k≤10).
kernelized discriminant analysisfew-shot learningllm embeddingsrank degeneracygeometric separability
A Unifying Perspective on Probabilities as Model Predictions
The paper proposes a unifying perspective on probabilities as outputs of prediction methods, emphasizing their dependence on abstraction construction and prediction transformation. It argues that all probabilities, including ostensibly objective ones, are model-dependent, and demonstrates that meeting a finite calibration criterion enables utility distribution anticipation and informed decision-making for finite event sets. The framework connects inductive arguments, probability calculus, and prediction methods to reconcile Bayesian and frequentist intuitions.
prediction methodscalibration criterionmodel-dependent probabilitiesutility distributioninductive arguments
When Does Low-Bit Quantization Preserve the Decisions of Vector Search?
The paper analyzes when low-bit quantization preserves decision fidelity in vector search by examining comparison-level effects rather than aggregate metrics. It presents a distribution-free decomposition showing comparison flip probability depends on zero-margin mass and residual tails, with covariance-aware bounds under a joint MGF proxy. Key results include deterministic coupling for Vamana neighbor selection, correlation gains from magnitude bits in bilinear models, and held-out block certificates for selective failure risk. Standardized exact margins outperform global rank correlation in predicting ranking/pruning flip rates across embeddings, with applicability to binary codes and product quantizers.
low-bit quantizationvector searchvamana pruningexact marginsresidual tails
TempTPI: Informer-Based trajectory prediction for maritime vessels
TempTPI introduces an Informer-based encoder with multi-channel temporal encoding for maritime vessel trajectory prediction, addressing computational inefficiency and long-term accuracy degradation in Transformer-based approaches. The framework combines ProbSparse self-attention to reduce quadratic complexity and Fourier-like frequency expansions to model cyclic behavioral patterns (hourly, daily, seasonal). Evaluated on Danish AIS data, TempTPI outperforms TPTrans across 1–5 hour horizons, achieving a 55% MSE reduction at 5 hours.
informertrajectory predictionprobsparse attentionais datatemporal encoding
In Medical Claims Data, Enhancing Predictive Performance for Major Adverse Cardiovascular Events Using Cross Attention
The study proposes a cross-attention mechanism to improve predictive performance for major adverse cardiovascular events (MACE) using medical claims data, which often lacks structured clinical information. By weighting relationships between diagnoses and treatments, the method generates more representative features. The cross-attention model achieved a ROC-AUC score of 0.7720, outperforming benchmark models including the conventional atherosclerotic cardiovascular disease model, LightGBM, and a self-attention baseline, demonstrating the efficacy of cross-attention in clinical prediction tasks.
cross-attentionmedical claims datamajor adverse cardiovascular eventsroc-aucclinical prediction
Online Inverse Integer Linear Optimization via Small-Gradient Skipping: Constant Regret and Finite Mistakes
The paper introduces Small-Gradient Skipping (SGS), a mechanism for online inverse integer linear optimization that skips updates when no mistake occurs and the correct action is uniformly separated. Applied to online gradient descent, online Newton step, and MetaGrad, SGS bounds the number of mistakes by a quantity independent of the total rounds T. For online Newton step and MetaGrad with SGS, the dimension-dependent regret improves to O(d²) for integer linear programs, eliminating the log T factor. In M-convex action sets, SGS achieves efficient regret bounds without center-of-gravity computations.
online inverse optimizationsmall-gradient skippingregret boundsinteger linear programmingm-convex sets
A practical DIRECT-type algorithm for medium-scale black-box global optimization
X-DTC-GL introduces dynamic partitioning and hybridization to enhance DIRECT-type algorithms for medium-scale black-box optimization. The method employs adaptive hyper-rectangle subdivision via local one-dimensional surrogate models and selectively applies hill-climbing in promising regions. Evaluations across four benchmark suites show 12% higher solvability, 27% improved solution quality, fastest convergence on 40% of instances, and best runtime on 17% of problems while maintaining competitive execution times.
direct-type algorithmblack-box optimizationdynamic partitioningsurrogate modelshill-climbing
Privacy-Preserving Split Learning for Federated LLM Fine-Tuning
The paper introduces a privacy-preserving split learning framework for federated fine-tuning of large language models (LLMs), addressing the inherent leakage of private inputs via intermediate activations in autoregressive models. The proposed method employs a learned obfuscate-and-recover scheme to protect distributed datasets while enabling server-side model training. Experiments show the approach achieves strong privacy guarantees with minimal utility degradation and system overhead, facilitating practical deployment in federated settings.
split learningfederated learningllm fine-tuningprivacy preservationautoregressive models
Evaluating Model Retraining under Drift: Paired Comparisons of Cumulative Subgroup Disparity
The study evaluates retraining policies for deployed classifiers by comparing cumulative subgroup disparities across scheduled, loss-triggered, and subgroup-gap-triggered approaches. Using paired differences in absolute subgroup gaps over deployment windows, the authors analyze true-positive and false-positive rates in simulations and real-world data (e.g., American Community Survey). Results show all three policies reduced mean cumulative disparity by 0.04–0.88 percentage points per window, with 69–92% agreement between finite-window and population comparisons. Subgroup-specific drift led to smaller true-positive-rate gaps but lower sensitivity. Policy comparisons require group-specific rates, action distributions, and explicit evaluation populations.
subgroup disparityretraining policiestrue-positive ratefalse-positive ratedrift regimes
NEXUS-MI: Communication-Aware Federated Personalization for Gateway-Coordinated Motor-Imagery Brain-Computer Interfaces
NEXUS-MI introduces a gateway-coordinated federated personalization framework for motor-imagery brain-computer interfaces (MI-BCIs), addressing subject/session variability via communication-aware synchronization. The method localizes raw EEG and classifier heads while maintaining a shared backbone at the edge coordinator, evaluated on BCICIV-2a (9 subjects, 4 classes) and OpenBMI (54 subjects, 2 classes) with session-wise partitioning. Communication-aware policies reduced backbone traffic by ~42% with minimal cohort-level accuracy impact, though subject-level losses reached ~12pp on BCICIV-2a versus ideal-link baselines, highlighting trade-offs between personalization, communication cost, and update freshness.
federated learningmotor-imagery bcieeg personalizationcommunication-aware synchronizationgateway coordination
MethaneFuse: Learning from Multi-Sensor Satellite Observations for Methane Plume Detection
MethaneFuse proposes a multi-sensor learning framework for methane plume detection that operates under partial sensor availability, addressing limitations of single-sensor approaches. The method leverages MethaneUnion, a temporal dataset combining Carbon Mapper plume reports with matched Sentinel-2, Landsat 8/9, EMIT, and Sentinel-5P observations (8,981 plume cases). At 480 m resolution, MethaneFuse achieves 84.87 F1 and 93.62 AUROC, improving baselines by 5.65 F1 and 8.30 AUROC points while reducing false positives by 8.19 points. The model transfers knowledge across sensors, maintaining performance when Sentinel-2 data is unavailable.
methane plume detectionmulti-sensor fusionsatellite imagerypartial observationstemporal dataset
EEGBind: Detecting Source-Level Interictal Epileptiform Discharges via EEG-Centric Multimodal Binding
EEGBind introduces an EEG-centric multimodal binding framework for five-class source-level classification of interictal epileptiform discharges (IEDs), addressing challenges like subtle source-region evidence and subject variability. The method treats EEG as primary, binding synchronized video-context features to an EEG-centric representation without early fusion to avoid perturbing source-sensitive signals, and employs view-consistent repair for hidden-set robustness. On the NeuroMM 2026 Grand Challenge Track 3 NMM-Source-IED benchmark, EEGBind achieves a weighted-F1 score of 0.8395, outperforming competitors, validating its practical utility for IED classification.
eeg-centric multimodal bindinginterictal epileptiform dischargessource-level classificationview-consistent repairneuromm 2026 benchmark
EFQ-Softmax: Exp-Free Quantization for Softmax
EFQ-Softmax introduces a low-bit probability-generation method for Transformer attention, eliminating the high-precision exp-then-quantize path by directly mapping shifted attention scores to block-scaled E2M1 operands via a single affine rule. It maintains FlashAttention-style updates while using consistent low-bit operands for both numerator and denominator updates. Evaluations on Qwen3-8B, Qwen3-VL-8B-Instruct, and WAN2.2-TI2V-5B show quality improvements (e.g., Qwen3-VL nine-task mean from 0.7826 to 0.8000) and a 40.33% latency reduction on A5 vector units for sequences up to 128K, demonstrating viable replacement of conventional softmax quantization.
softmaxquantizationtransformere2m1flashattention
Efficient Graph Neural Networks for Multicarrier Wideband Hybrid Beamforming Optimization
The paper proposes three Graph Neural Network (GNN) architectures for optimizing hybrid beamforming in multicarrier wideband systems, addressing beam squinting in 6G wireless networks. The GNNs model a shared analog beamformer via bipartite graphs, with digital beamformers represented at subcarrier nodes, edges, or integrated with singular-value decomposition. Evaluations demonstrate superior performance over traditional optimization and ML-based methods, with resilience to beam squinting and imperfect channel state information (CSI). The GNNs generalize to multi-user scenarios without retraining, achieving 15-20% higher spectral efficiency than existing hybrid designs.
graph neural networkshybrid beamformingbeam squintingmulticarrier wideband6g wireless
ALIGN-HOLD: Experience Alignment for Real-Time Hold Control in Large-Scale Ride-Hailing Matching at DiDi
ALIGN-HOLD introduces an experience alignment framework for real-time hold control in large-scale ride-hailing systems, addressing limitations of handcrafted reward designs in existing approaches. The method constructs complementary preference pairs from order trajectories, driver trajectories, and local matching graphs, training an experience Reward Model (RM) via balanced multi-view sampling and model-adaptive hard preference sampling. The RM provides dense, context-dependent rewards during simulator-based policy learning and filters low-identifiability interactions. Deployed on DiDi's platform in a 28-day A/B experiment (100,000 daily requests), ALIGN-HOLD significantly improved trip completion rates and driver income while reducing passenger cancellations, validated through ablations and RM diagnostics.
experience alignmentreward modelreal-time hold controlmulti-view samplinglow-identifiability interactions
Settling: Equilibrium Inference for Non-Convex Validity Sets
The paper introduces Settling, an equilibrium-based inference operator addressing conditional mean collapse in learning systems with non-convex validity sets. The method separates proposal generation, consistency evaluation, and equilibrium selection, refining mean-seeking proposals toward locally stable configurations via exact-gradient descent. Experiments on a 100-context geometric diagnostic show Settling achieves 99/100 success rates with lower trajectory roughness compared to baselines (0/100 for mean-seeking, 100/100 for stochastic denoising). A 1,200-run sensitivity study demonstrates robustness, with 97-100% success across obstacle-jitter ranges up to 0.20 and 94-100% across initialization perturbations from 0.05 to 0.50.
conditional mean collapseequilibrium-based inferencenon-convex validity setsexact-gradient descentlocal convergence
Muon-C: Operator-Aligned Muon for Convolutional Kernels
Muon-C introduces an operator-aligned optimizer for convolutional kernels by representing kernel momentum as frequency-wise channel-transfer matrices, polarizing these blocks independently, and using a critical Fourier grid to maintain finite kernel support. The method combines block partition and Fourier coordinates to derive an exact-polar direction, which serves as a linear minimization oracle under the critically sampled convolution norm. On CIFAR-10 flow matching, Muon-C achieves 9.87 FID at 40k iterations, outperforming unfolded Muon (22.26 FID) and Adam (51.31 FID), with 0.62× and 0.64× their model FLOPs, respectively. Under equal tuning budgets, it reaches 3.42 FID, with gains scaling across data and architectures.
convolutional kernelsoperator-aligned optimizerfourier gridlinear minimization oracleflow matching
PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling
PELM introduces a power-efficient on-device LLM inference method combining dynamic voltage frequency scaling (DVFS) with speculative decoding and variable verification depth, exploiting token-level computational redundancy. The approach jointly optimizes hardware frequency and inference workload depth, expanding the optimization space beyond prior DVFS-only methods. Evaluations show PELM achieves up to 23.1% speedup and 52.4% energy reduction versus SOTA power governing techniques while preserving task performance across hardware platforms and datasets.
speculative decodingdynamic voltage frequency scalingon-device inferencepower efficiencyllm optimization
Why Learning Rediscovers the Closed-Form Diagonal Regularizer
The paper identifies a diagonal saturation principle in modal inverse problems, demonstrating that the Bayes-optimal Tikhonov regularizer adopts a closed-form power law Γₖ ∝ λₖ^|s|, determined solely by the prior and independent of the domain. Leveraging Berry's random-wave conjecture and Weyl's eigenvalue counting law, the authors show that truncation noise decorrelates across modes, resulting in a flat loss landscape that limits the efficacy of learned diagonal regularizers. Empirical validation on FEM-simulated acoustic rooms confirms near-optimality of the closed-form solution, with three learned diagonal architectures achieving reconstruction errors within 1 percentage point. The framework extends to heat diffusion via an exponential Green's function correction, while learned iterative ridge methods exploit cross-mode coupling beyond diagonal saturation.
tikhonov regularizermodal inverse problemsberry's conjectureweyl's lawgreen's function
LightMedSeg-ISLES: Stroke Lesion Segmentation with 81x Fewer Parameters than nnU-Net
LightMedSeg-ISLES introduces a lightweight (1.26M parameters) pipeline for stroke lesion segmentation in T1-weighted MRI, achieving 97.5% of nnU-Net ResEnc-L's Dice score (0.618 vs. 0.634) with 81.4× fewer parameters. The method combines flip test-time augmentation (TTA) and lesion-wise F1 optimization, outperforming nnU-Net by 0.055 in lesion-wise F1 (0.599 vs. 0.544) while reducing FLOPs per patch by 4.7×. Extended training and augmentation further improve Dice by 0.0358 without parameter growth, surpassing UNETR++ and nnFormer in efficiency. Results are validated on ISLES'26's 146-case cohort.
stroke lesion segmentationlightweight architecturetest-time augmentationnnu-nett1-weighted mri
Distillation of Synthetic Data for Time Series Foundation Models
The paper introduces synthetic data distillation (SDD), a novel training objective for time series foundation models (TSFMs) that compares model outputs to conditional forecast distributions rather than realized future values. SDD performs Rao-Blackwellization of the stochastic gradient, reducing covariance under the Loewner partial order while preserving expectation. Evaluated on TSFMs ranging from 4M to 2.5B parameters, SDD achieves faster validation loss convergence, matching or surpassing baseline performance with 10%-40% fewer training iterations on Gaussian Process data.
time series foundation modelssynthetic data distillationrao-blackwellizationloewner partial ordergaussian process
Geometric organization of olfactory descriptor data in the Poincaré disk
The study demonstrates that two-dimensional hyperbolic embeddings effectively capture the structure of olfactory descriptor data, using hyperbolic metric multidimensional scaling on two datasets: Sagar (480 rating profiles) and GoodScents--Leffingwell (4983 molecules). The embeddings preserved pairwise descriptor distances, revealing radial and angular organization. In Sagar, profile entropy negatively correlated with hyperbolic radius, while in GoodScents--Leffingwell, active label entropy increased with radius. Sweet, musky, fruity, and pleasantness showed strong directional trends. The findings support hyperbolic mapping as an interpretable framework for olfactory data.
hyperbolic embeddingolfactory descriptorsmultidimensional scalingentropydescriptor profiles
Positional task conditioning for scalable defect detection across product families in large product catalogs
The paper introduces Positional Task Conditioning (PTC), a method for scalable defect detection in large product catalogs by decomposing the task into focused sub-tasks to mitigate long-context degradation in LLMs. PTC reinforces task identity at prompt boundaries, enabling distillation into a smaller model while maintaining performance. The approach improves F1 from 52% to 87% and achieves within 1.79% F1 of the frontier model at up to 98% lower cost, deploying across multiple countries to process 10+ million product families.
positional task conditioningdefect detectionlong-context degradationtask distillationproduct catalogs
Robust Industrial Cyber Physical Classification Using Neuromorphic Temporal Embeddings and Hybrid SNN XGBoost Under Machine Unlearning Attacks
Proposes a hybrid Spiking Neural Network (SNN) and XGBoost architecture for robust intrusion detection in power grids, combining fixed neuromorphic temporal embeddings with a retrainable classifier to resist machine unlearning attacks. The SNN serves as a feature extractor trained once on clean data, while XGBoost handles updates, reducing vulnerability to label-flipping. Achieves 99.9% accuracy (F1-macro 0.999) on Synchrophasor and 95.0% (F1-macro 0.943) on MSU/ORNL datasets, with only 0.9% F1-macro degradation at 10% poisoning and delayed collapse under 60-70% attacks compared to baselines.
spiking neural networkmachine unlearningcyber-physical systemsxgboostdata poisoning
TEFM: Token-Efficient Faithful Modeling for Structured Data
TEFM (Token-Efficient Faithful Modeling) introduces a framework for structured data analysis in critical domains, addressing token efficiency and faithfulness in LLMs. It compresses structured observations into Behavioral Code tokens, reducing token usage to ~1-2% with minimal information loss, and employs a dual-fidelity objective optimizing code-level reconstruction and prediction-level fidelity. Evaluations across clinical and security datasets with Qwen3, Gemma-2, and Phi-4 show competitive accuracy while producing faithful rationales.
token efficiencyfaithful modelingbehavioral code tokensdual-fidelity objectivestructured data analysis
Differentially Private Average Treatment Effect Estimation by Propensity Score Blocking
The authors propose two differentially private algorithms for average treatment effect (ATE) estimation in observational studies, addressing privacy concerns in sensitive domains. The first algorithm improves upon existing inverse probability weighting (IPW) methods, while the second introduces blocking on the propensity score (BPS). Both methods demonstrate reduced error and bias compared to prior approaches, with the BPS-based algorithm achieving error reductions of 75% or more in many cases. These advancements enable more accurate and privacy-preserving ATE estimation in fields such as social science and medicine.
average treatment effectdifferential privacypropensity scoreinverse probability weightingobservational studies
Oracle Complexity of Stochastic Fixed-Point Equations with Nonexpansive Maps
The paper presents an algorithm for computing fixed-point residuals with small error ε for nonexpansive self-maps T on compact convex sets, using an unbiased stochastic oracle with bounded variance σ². The method employs a recursive anchoring technique and applies to norms with weak Rademacher type q > 1. For type-2 spaces (e.g., ℓ_p for p ∈ [2, ∞]), the algorithm achieves stochastic oracle complexity Õ(σ²ε⁻³ + ε⁻¹). A near-matching lower bound is proven for ℓ_∞-norm instances in high dimensions, extending to sparse noise settings and precluding complexity improvements via non-matching ℓ_p norms.
fixed-point residualnonexpansive mapsstochastic oracle complexityrademacher typeℓ_p spaces
Inductive Biases in Field-Level Cosmological Inference from Galaxy Catalogs
This study investigates inductive biases in field-level cosmological inference by comparing machine learning architectures for estimating the matter density parameter Ω_m from simulated galaxy catalogs. Using CAMELS hydrodynamic simulations, the authors evaluate permutation-invariant Deep Sets (with MLP and KAN implementations) against graph neural networks (GNNs) across galaxy positions, peculiar velocities, and their combinations. Results show Deep Sets achieve ~18% mean relative error for Ω_m from velocities alone, while GNNs achieve ~10% by leveraging spatial relations. Peculiar velocities emerge as the dominant information source for set-based models, whereas spatial information is optimally utilized by architectures explicitly encoding galaxy-galaxy relations.
inductive biasescosmological inferencedeep setsgraph neural networkspeculiar velocities
Efficient Fairness Auditing Across Guidance Scales in Text-to-Image Diffusion Models via Causal Abstraction
We introduce a causal-abstraction-based fairness auditing method for text-to-image diffusion models that reduces computational costs by efficiently evaluating fairness across classifier-free guidance scales. The approach constructs a low-level structural causal model of the diffusion process and a corresponding high-level model over abstract denoising states, enabling identifiability of fairness-relevant interventional queries. A probabilistic transformer implements the high-level model to predict target-feature distributions amortized across guidance scales. Experiments demonstrate distributional fidelity, fairness-query accuracy, and computational efficiency, with evaluations conducted on Stable Diffusion 1.5 and StayFair, a fairness-enhanced variant.
causal abstractiondiffusion modelsfairness auditingclassifier-free guidanceprobabilistic transformer
Infra-Bench CLS: A Global, Open-Source Benchmark for Critical Infrastructure Classification with Earth Observation Foundation Models
Infra-Bench CLS introduces a global benchmark for evaluating Earth observation foundation models on facility-scale critical infrastructure classification, comprising 18,756 Sentinel-1 SAR and Sentinel-2 multispectral images across 13 classes (10 retained). Linear probing and fine-tuning (1.0x/0.3x training data) assessed seven models (e.g., SatlasPretrain S2, Prithvi-EO-2.0, DINOv3 ViT-L/16), with the best achieving 57.9% macro F1 (48% improvement over ResNet-18's 39.2%). Performance varied by class: airports (85.3%), train stations (82.1%), and data centers (77.6%) excelled, while power-sector classes lagged (27.5-46.2%). Results indicate potential for foundation models but highlight needs for higher-resolution evaluation in underperforming sectors.
earth observationfoundation modelssentinel-1linear probingmacro f1
Gaussian Approximation for Multivariate Martingale Sums from Uniformly Ergodic Markov Chains
The paper establishes Gaussian approximation bounds in higher-order Wasserstein distance $W_p$ ($p\geq2$) for sums of multivariate martingale differences from uniformly ergodic Markov chains. Under an $L^{(2+\eta)p}$-moment condition ($\eta>0$), the authors derive an explicit bound $O\left( p^3 \|A\|_4^2 + pd^{1/4}\|A\|_2^{1/2}\|A\|_4^2 \right)$, where $A$ captures increment sizes. In the balanced-increment regime, this yields the first optimal $O(n^{-1/2})$ rate for fixed $p$ and $d$. Key techniques include antisymmetric Stein couplings and a refresh-then-maximal coupling method combining independent resampling with maximal coupling for temporal dependence control.
wasserstein distancemartingale differencesmarkov chainsstein couplingsgaussian approximation
Unthrottling the Tanh Jacobian in SAC: A Negative Result on Bang-Bang Control and MetaDrive
The study investigates whether addressing the vanishing Jacobian in Soft Actor-Critic (SAC) improves performance in bang-bang control tasks. The authors propose a minimal intervention: an additional term in the actor loss to restore gradient signals at action bounds. Experiments on a double integrator (near-optimal return: -31.6 vs. -30.3) and MetaDrive show that bypassing the tanh Jacobian fails to improve returns, with ungated bypass collapsing performance (-195.5) and gated bypass also underperforming. Results suggest saturating action bounds does not equate to solving tasks optimally, and entropy tuning counters bypass effects.
soft actor-criticjacobianbang-bang controltanh saturationentropy coefficient
LeCor: Learning to Be Corrected by Meta-Learned Test-Time Training for Interactive 3D Lung-Tumour Segmentation
LeCor introduces meta-learned test-time training for interactive 3D lung-tumour segmentation, enabling case-specific adaptation through gradient updates driven by clinician corrections. The method employs a small set of case adapters, meta-learned to improve segmentation on unannotated slices with a single gradient step. Evaluated on 690 test cases from five public CT cohorts, LeCor achieves a Dice score of 0.827 after seven correction rounds, surpassing the fine-tuned SAM 3 model's 0.787. It reduces cases failing to reach a Dice of 0.80 from 47 to 27 and attains in three rounds the accuracy SAM 3 achieves in seven.
meta-learningtest-time traininglung-tumour segmentationcase adaptersgradient updates
Mode Coverage in Normalizing Flow Boltzmann Generators via Log-Ratio Variation
The paper introduces KLXX, a novel loss function for normalizing flow Boltzmann generators that addresses mode coverage issues by incorporating log-ratio variation metrics. KLXX combines two log-ratio variations: one weighted by the target distribution for accuracy and another by a mixture of quench, temper, and pushforward samples for mode exploration. Theoretical analysis shows KLXX's Fisher--Rao gradient flow properties and provides error bounds. Empirical results demonstrate improved mode coverage over forward KL, better diagnostic performance, and accurate recovery of observables compared to independent references.
normalizing flowboltzmann generatorslog-ratio variationmode coveragefisher-rao gradient
Building the Harness Automatically: Self-Play in Code Distills a Text Harness for Black-Box Optimization
The paper introduces a self-play framework where an agent learns numerical search strategies through executable code practice, then distills them into a concise text harness (197-word Harness A) for black-box optimization. During development, the agent iteratively writes and evaluates optimizer programs, later freezing the distilled text before evaluation. Harness A reduces regret by 48% for Gemini Flash (p<0.001), matches GP-BO performance on training tasks, and generalizes to BBOB landscapes and other executors (43-49% regret reduction for Claude Sonnet). An independent replication yields Harness B with comparable performance, and the method achieves state-of-the-art results on a YouTube reward-tuning benchmark.
black-box optimizationself-playregret minimizationlanguage model distillationbbob
Exact-Form Regret for Gradient Descent, Mirror Descent and Follow-the-Regularized-Leader
The paper provides a geometric characterization of deviation classes for which online gradient descent, mirror descent, and follow-the-regularized-leader (FTRL) achieve no regret, identifying exactness as the key principle. Exactness requires the displacement field to derive from a scalar potential, with the geometry varying by algorithm (Euclidean for gradient descent, regularizer-induced for mirror descent, cumulative dual state for FTRL). Under mild conditions, exactness ensures sublinear regret, while nonzero circulation leads to linear regret. The framework unifies deviation classes across algorithms, revealing distinct capabilities and implications for learning in games, including a new equilibrium notion called conservative correlated equilibrium.
exactnessdisplacement fieldsublinear regretconservative correlated equilibriumdual state
A Block Tensor Train Burer-Monteiro Framework for Low-Rank Quantum State Tomography
The authors propose a block tensor train (Block-TT) factorization for low-rank quantum state tomography, compressing the density matrix representation from exponential to linear in qubit count while preserving Hermiticity and positive semidefiniteness. Their method employs single-site and two-site DMRG algorithms operating directly on the compressed parameterization, supporting adaptive rank refinement and efficient tensor-network contractions for expectation-value evaluation. Numerical experiments show accurate state reconstruction from limited measurements with reduced memory and computational costs compared to conventional low-rank tomography methods.
quantum state tomographyblock tensor traindensity matrix renormalization grouplow-rank approximationtensor-network contractions
Uncertainty-Aware Sea-Ice Type Mapping with Multiple Ice Charts
The paper quantifies two uncertainty sources in sea-ice stage-of-development (SoD) mapping: multi-annotator label uncertainty from divergent expert ice-chart interpretations and model uncertainty from predictive deep-learning models. It evaluates their relationship using soft supervision and Monte Carlo dropout, finding that model uncertainty correlates with annotator disagreement (0.256 overall, 0.704 near ice edges). Monte Carlo dropout yields the best-calibrated confidence estimates (expected calibration error 0.050), demonstrating improved uncertainty correspondence when incorporating multiple annotators.
sea-ice stage of developmentmulti-annotator uncertaintymonte carlo dropoutmodel calibrationice-edge detection
Concept drift mitigation through community and spectral graph analysis for the detectionof cyberattacks in network traffic
The paper introduces t-robustness, a stability score for feature selection in network traffic analysis to mitigate concept drift in cyberattack detection. The method evaluates features independently of detection models by combining step-by-step distance and cumulative divergence metrics, focusing on abnormal network connectivity patterns via graph community and spectral metrics. Evaluated on the UGR16 dataset across three learning scenarios, t-robust feature spaces sustain detection performance (0.6025 expectancy) compared to graph community features (0.5230) and base NetFlow features (0.3831).
concept driftt-robustnessgraph community metricsspectral metricsnetwork traffic
MiNCE: Nonparametric, Strongly Consistent Confidence Envelopes for Band-Limited Functions and their Smoothed Spectra
The MiNCE framework establishes strongly uniform consistency for minimum-norm confidence envelopes in band-limited functions and their smoothed spectra, addressing a gap in prior finite-sample coverage analyses. Leveraging Reproducing Kernel Hilbert Spaces (RKHS) theory, MiNCE constructs nonparametric, nonasymptotic, simultaneous confidence regions under mild noise assumptions for both noise-free and noisy observation models. Theoretical guarantees are extended to frequency-domain smoothed spectra, ensuring strong uniform consistency. Empirical validation through nonparametric regression and spectral estimation experiments demonstrates envelope contraction toward target functions with increasing sample size.
minimum-norm confidence envelopeband-limited functionsreproducing kernel hilbert spacesnonparametric regressionsmoothed spectra
Tensor-Train Weak SINDy: Identifying High-Dimensional Nonlinear Dynamics
The authors propose TT-WSINDy, a method combining Multidimensional Approximation of Nonlinear Dynamics (MANDy) and Weak Sparse Identification of Nonlinear Dynamics (WSINDy) using tensor-train (TT) format to address computational and memory challenges in high-dimensional dynamical system identification. By leveraging TT decompositions, the method efficiently searches an exponentially-growing candidate function space while performing weak-form transformation, regression, and sparsification without succumbing to the curse of dimensionality. Results demonstrate its efficacy in high-dimensional settings where traditional techniques prove prohibitively expensive.
tensor-trainweak-form methodsnonlinear dynamicssparse identificationhigh-dimensional systems
Applying foundation model embeddings towards urban livability evaluation
The study presents a framework for identifying predictive geospatial indicators of urban livability using foundation model embeddings (AlphaEarth, AnySat, TerraMind). By analyzing spatial dependencies in high-resolution satellite data, the method prioritizes informative physical features for socioeconomic prediction in data-scarce regions. Results demonstrate improved livability prediction accuracy through selective feature extraction from foundation model embeddings, enabling policy interventions in underobserved areas.
foundation modelsgeospatial embeddingsurban livabilitysatellite imageryfeature prioritization
X-amine509: Predicting the Practical Risk Level of Enterprise X.509 Certificates
X-amine509 introduces a two-stage ML system for prioritizing enterprise X.509 certificate risk assessment, combining rapid ML-based triage with full deterministic analysis for high-risk cases. The method employs Extra Trees and Decision Tree models trained on 1,027,714 certificates, using 177 defect checks weighted by severity across four tiers. Results show Extra Trees achieves R²=0.993 (MAE=2.26) on 201,976 test certificates, with Decision Tree processing 3.7M certs/sec (R²=0.986), maintaining R²≥0.915 and critical-tier recall≥97.03% after 13 months. Feature importance highlights validity period, EKU configuration, negative serial encoding, and self-signed status as top predictors.
x.509 certificatesrisk prioritizationextra treesdeterministic analysisfeature importance
XAI-Refine: An Automated Explanation-Knowledge Loop for Brain-Age Prediction
XAI-Refine introduces an automated explanation-knowledge loop for brain-age prediction from resting-state functional connectivity, addressing limitations of post-hoc analyses by iteratively refining model explanations with verified neurobiological evidence. The method consolidates post-hoc analyses into structured explanations, converts these into neurobiological questions, retrieves and verifies literature, and translates verified evidence into differentiable constraints for model refinement. Multi-seed validation ensures target-directed explanatory movement while maintaining predictive performance and bounded non-target drift. Experiments demonstrate improved explanation reliability, literature alignment, and target-specific model revision in functional-connectivity-based brain-age prediction.
brain-age predictionpost-hoc analysisfunctional connectivityexplanation-knowledge loopdifferentiable constraint
Constraint-Aware Discrete Black-Box Optimization Using Tensor Decomposition
The paper introduces a tensor decomposition-based surrogate modeling approach for discrete black-box optimization that explicitly incorporates feasibility constraints. The method formulates surrogate model training as a constrained polynomial optimization problem, solved via a differentiable penalty term derived from T-norms to handle logical constraints. Experiments on synthetic benchmarks and a pressure vessel design task demonstrate improved sample efficiency by effectively avoiding infeasible regions during optimization.
discrete optimizationsurrogate modelingtensor decompositionfeasibility constraintst-norms
Explaining f-Divergence-Based Regularization via Local Curvature and Sharpness-Aware Minimization
The work establishes a theoretical connection between f-divergence-based regularization and Sharpness-Aware Minimization (SAM) by analyzing their local curvature properties under parameter- and input-space perturbations. Through second-order expansions, both methods are shown to induce curvature-sensitive penalties: divergence regularization yields a Fisher-weighted quadratic form, while SAM penalizes sharpness via the dominant Hessian eigenvalue. Empirical validation using the α-skew Jensen-Shannon divergence family demonstrates that maximal curvature penalization (achieved at α=0.5) correlates with flatter minima and improved accuracy/negative log-likelihood on four benchmark datasets.
f-divergencesharpness-aware minimizationcurvature penalizationjensen-shannon divergencefisher matrix
Encrypt What Matters: When Selective Homomorphic Inference Is Efficient
The paper introduces selective homomorphic inference, a method where only sensitive regions of interest (ROIs) are encrypted during inference, while computations independent of the ROI use plaintext. This approach matches full FHE output without retraining, with efficiency gains depending on encrypted dependency propagation. Locality-preserving architectures achieve order-of-magnitude speedups for small encrypted ROIs, whereas architectures with early global mixing show minimal improvement, highlighting locality as critical for selective homomorphic inference efficiency.
fully homomorphic encryptionselective homomorphic inferenceregion of interestlocality-preserving architecturesencrypted dependencies
Real-time and adaptive anomaly detection algorithm for cyclostationary models
PeriodicCALM introduces a real-time anomaly detection framework for cyclostationary data streams, addressing limitations of existing methods that misinterpret phase-dependent variability as anomalies. The method incorporates cycle-dependent variability to distinguish regular cyclic impulses from genuine anomalies, featuring continuous retraining and dynamic adaptation to signal evolution. Evaluations on simulated data show improved detection accuracy (vs. baseline CALM), training efficiency, and reduced prediction latency, with validation on real-world compressor vibration signals demonstrating practical utility.
cyclostationaryanomaly detectionreal-timeadaptive filteringimpulsive noise
Literati: Towards Anytime Optimal Shape Generalized Trees via AO*
Literati introduces the first algorithm for optimal Shape Generalized Tree (SGT) induction, addressing limitations of greedy approaches by jointly optimizing tree structure and shape function complexity via a novel AND/OR graph formulation. The method enhances AO* with a secondary OR-node heuristic and round-robin AND-node exploration, ensuring optimality while improving anytime performance. Evaluated on 24 real-world datasets, Literati outperforms state-of-the-art tree methods in both training and test accuracy.
shape generalized treesoptimal decision treesand/or graphao* algorithmanytime performance
"Transforming" LHCb: self-supervised maps of heavy-flavour decays
The authors propose a self-supervised transformer architecture to improve heavy-flavour decay mapping at LHCb, enhancing sensitivity to beyond-standard-model physics. The model learns decay environments by inferring masked particle identification and completing partial jets, without requiring flavour or exclusive-decay labels. Evaluated on five classification tasks in simulated LHCb Open Data, the self-supervised transformer matches a fully supervised counterpart and outperforms random initialization, achieving ~10% tagging power. Anomaly detection in 2017 LHCb proton-proton collision data confirms systematic score increases across eight heavy-flavour channels, validating the method's potential for extending discovery reach.
transformerself-supervisedheavy-flavouranomaly detectionjet completion
Tensor Network Moral Graph Recovery of Discrete Probability Distributions
The paper introduces a method for recovering the moral graph of a causal DAG from discrete probability distributions using fully connected tensor networks (FCTNs) with nuclear-norm-regularized bond corrections. Each bond matrix is parameterized as a baseline all-ones matrix plus a low-rank correction, and unnecessary bonds are driven to zero via a variational Frobenius norm penalty. Under faithfulness, positivity, and a no-implicit-rerouting assumption, optimal FCTNs with zero reconstruction error yield effective graphs matching the moral graph. For approximate cases, recovery bounds are provided using Fannes-Audenaert continuity, and a sufficient condition on the regularization parameter is derived.
moral graphtensor networksnuclear normdiscrete probabilitycausal dag
Accountable and uncertainty-aware evaluation of sensor-based AI under distribution shift: devices, subjects, and nearly three years underground
The study introduces a staged, accountable evaluation protocol for sensor-based AI systems under distribution shifts, emphasizing uncertainty quantification and decision-making based on 5% quantiles. The protocol involves four cumulative generalization stages that sequentially hold out devices, subjects, and time, assessing performance against chance references and out-of-present-scope rates. Applied to geomagnetic localization in underground mines using smartphone-based recurrent classifiers, the method evaluates models on data recorded 34 months post-training, with a held-out device generation and surveyor. Results show a 5% quantile precision of 0.39, 16.5 times the chance level, with variability across location classes exceeding repeated run spreads. The protocol highlights the inadequacy of mean-based evaluations, demonstrating bimodal configurations can mislead.
distribution shiftquantilegeomagnetic localizationrecurrent classifiersuncertainty quantification
CAST: Canonical Approximate Schur Tree for Approximate Cholesky on Graphs
The paper introduces CAST (Canonical Approximate Schur Tree), a method for constructing approximate Cholesky preconditioners on graphs by replacing dense Schur-complement cliques with weighted random spanning trees. CAST samples trees directly from the clique, ensuring connectivity and unbiased updates via edge reweighting. The authors also propose CAST-$\rho$, which uses $\rho$ copies of each neighbor to reduce sampling variability, with a $1/\rho$ bound on Schur error. Theoretical analysis shows CAST minimizes leverage-score marginals, while empirical results indicate CAST-1 is faster, whereas CAST-2 is preferable when edge contributions remain low-cost.
approximate choleskyschur complementrandom spanning treeleverage-score marginalsgraph preconditioners
📰 Industry Media (11)
Powering AI is an architecture problem
The article identifies power architecture as a critical bottleneck for AI data centers, demonstrating that legacy medium-voltage power stacks fail under gigawatt-scale, volatile AI workloads due to undersized UPS systems, inefficient protection logic, and reactive design. Proposed solutions involve relocating power infrastructure to medium-voltage (13.8+ kV) modular enclosures near substations, implementing inline energy storage, and redesigning protection schemes to absorb load swings and grid disturbances. A 2026 test at the National Laboratory of the Rockies validated this architecture, achieving compliance with ERCOT's voltage ride-through requirements while improving grid stability, permitting efficiency, and economic viability for AI facilities.
medium-voltage upsgrid reliabilityload volatilityinterconnection studyvoltage ride-through
OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call
OpenAI released the Agents API in public beta, providing developers with a managed service built on the Codex harness for running long-running agents. The API supports OpenAI-hosted, self-hosted, and partner sandboxes, handling context compaction, efficient tool use, and multi-agent coordination. Key features include programmatic tool calling, subagent support, and versioned access to the harness. Early customer results show improvements in evaluation scores (Ciridae: 0.71 to 0.85), cost reduction (SafetyKit: 60% lower), and reliability (Hypha: 86% fewer failures). The API operates under US-only data residency and lacks Zero Data Retention support.
codex harnessmulti-agent coordinationcontext compactionprogrammatic tool callingsandbox
DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse
DeepSeek AI introduces DeepSeek-V4.1-Flash, a 552B-parameter Mixture-of-Experts model with 1M-token context, addressing KV-cache bottlenecks via FP4 quantization (890 bytes/token, 4× smaller than V4-Flash) and cross-layer attention reuse. The causal encoder-decoder architecture halves prefill compute by processing tokens only in the 20-layer encoder, while Compressed Sparse Attention 2 (CSA2) shares KV states across layers in Full, Reindex, and Reuse modes. Training on 45T multimodal tokens with sparse attention from scratch achieves 90.6 on Terminal-Bench 2.1 and 3471 Codeforces rating, outperforming GPT-5.6 Sol and Opus-5 in select benchmarks. MIT-licensed weights support vLLM, SGLang, and Transformers deployments.
kv-cachemixture-of-expertscompressed sparse attentionfp4 quantizationcausal encoder-decoder
LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity
LandingAI introduces Agentic Document Extraction Gen2, featuring DPT-3 Pro and DPT-3 Verity models, which restructure document parsing as a hierarchical tree rather than flat chunks. DPT-3 Pro handles complex layouts (scans, handwriting, non-Latin scripts) with line-level grounding, while DPT-3 Verity processes digital documents deterministically with word-level confidence scores (0-1) and bounding boxes. Pricing shifts from flat per-page (3 credits under DPT-2) to a hybrid model: DPT-3 Pro charges 1 credit/page + 0.5 credits/1k output characters (priority tier), while Verity costs 0.3/page + 0.2/1k chars. The system outputs structured markdown with atomic grounding for PII redaction and document diffing, achieving 25-80% cost reductions for mixed workloads.
document parsingatomic groundinghierarchical treeconfidence scoringhybrid pricing
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
Google introduces Mantis, a modular skills toolkit enabling AI coding agents to autonomously execute the vulnerability lifecycle, including detection, reproduction, patching, and risk assessment. Mantis operates via slash commands, integrating with frameworks like Gemini CLI and Google ADK, and enforces strict execution boundaries for safety. It employs a hierarchical summary tree to reduce token overhead by over 85% and achieves sub-7% true-positive rates in vulnerability detection. The toolkit is open-sourced under Apache 2.0, currently suitable for local evaluation but not production deployment. Key stages include threat modeling, exploit chain assembly, and patch verification, with a focus on deterministic execution over LLM orchestration.
vulnerability lifecycleslash commandshierarchical summary treeexploit chainpatch verification
Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds
Gradium introduces Voice Design, a text-to-speech system generating synthetic voices from 1-500 character prompts in English, French, Spanish, Portuguese, or German without reference audio. The system employs a four-call API workflow, producing 1-5 voice candidates in 3-5 seconds with non-deterministic sampling. In a blind pairwise test across 7,627 comparisons, Voice Design achieved a 72.6% win rate against competitors, excelling in regional accents like Quebecois French (97%) and Rioplatense Spanish (86%). It also scored 83.4% prompt adherence on InstructTTSEval's English benchmark.
text-to-speechsynthetic voicesnon-deterministic samplingprompt adherenceapi workflow
Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer
Meta introduces Muse, a personal AI agent capable of executing tasks autonomously, including emailing, booking, and negotiation, while maintaining user approval for sensitive actions. Muse operates on Muse Spark 1.3, a model optimized for long-horizon agentic tasks, reducing tool calls by 20% and token usage by 25% compared to its predecessor. Each user’s agent runs on a dedicated Muse Secure VM, isolated via systemd-nspawn runtime cells, with security enforced by the Sentinel agent, which handles credential surrogation and network egress. Muse Spark 1.3 is available through Meta Model API and Muse Code, with an open-weights release planned.
agentic taskscredential surrogationsystemd-nspawnsentinel agentmuse secure vm
Supply chains detect fast, act slow: How AI agents fix it
Supply chain AI systems reduce detection latency but fail to accelerate decision-making, costing firms $184B annually (J.S. Held 2025). Current implementations focus on predictive analytics (demand sensing, ETA prediction) but require human intervention for execution, creating a 28% productivity drain in disruption response (Knosc 2026). Bounded-action agents—pre-authorized to execute decisions like lane retendering or mode switching within policy constraints—demonstrate 2.7x higher autonomous action adoption in leading firms (FourKites 2025). Three prerequisites emerge: policy formalization, machine-readable execution systems, and accountability frameworks shifting from human oversight to system design audits.
supply chain optimizationautonomous agentsdecision automationpolicy-based routingexception handling
JD.com expands physical AI in logistics with 3 million robots
JD.com announced a Physical AI Acceleration Plan targeting deployment of 3 million robots, 1 million autonomous vehicles, and 100,000 delivery drones within five years. The initiative integrates Meta Brain 3.0 for real-time logistics optimization (processing 100M+ parcel routes in seconds) and LangzuTech robotic arms using parallel reinforcement learning for parcel handling. Current infrastructure includes 1,800+ self-operated warehouses, 30+ automated LangzuTech warehouses, and 100+ drone routes, with new deployments featuring temperature-resistant systems (-20°C operation) and night-time autonomous delivery. JD Cloud plans a 100,000-GPU cluster with Moore Threads for embodied-AI training alongside a 10M-hour video dataset.
physical aireinforcement learningautonomous vehicleswarehouse automationgpu cluster
CloudNC aims to accelerate AI supply chain machining
CloudNC secured $20M in funding to scale its AI-driven precision machining technology, focusing on automating CNC programming and quoting workflows. The company's CAM Assist software generates machining strategies and toolpaths from CAD models, reducing transition time between design and production. Currently deployed in over 1,000 machine shops globally, CAM Assist has users including Lockheed Martin. CloudNC is developing Quote Agent, an AI-assisted estimating tool for rapid job costing, scheduled for 2026 release. Additionally, the company is pursuing FedRAMP certification for CAM Assist to enable deployment in US government and defense sectors.
cnc programmingcam assistquote agentfedramp certificationtoolpath generation
Samsung taps Mistral AI models for semiconductor manufacturing
Samsung Electronics partners with Mistral AI to deploy on-premises AI models, including Mistral Large, for semiconductor manufacturing optimization. The collaboration focuses on private enterprise installations to process sensitive engineering data internally, avoiding cloud exposure. Targeted applications include automated defect detection, fab machinery tuning, and yield stabilization across memory, logic, and foundry operations. Samsung's investment in Mistral AI's Series D funding round ensures long-term technical cooperation. The integration aims to accelerate development cycles, improve manufacturing precision, and maintain factory throughput in advanced semiconductor production.
on-premises modelsdefect detectionyield stabilizationfab machinery tuningsemiconductor manufacturing
Generated automatically at 2026-09-10 21:57 UTC. Summaries and keywords are produced by an LLM and may contain inaccuracies — always consult the original article.
