Daily Digest — 2026-08-20

Wednesday, August 19, 2026 · 245 items · model: deepseek/deepseek-chat

245 items · 7 research labs, 234 arxiv papers, 4 industry media

🏛️ Research Labs (7)

Offering Zero Data Retention for frontier models

OpenAI News · 2026-08-19

OpenAI introduces Private Safety Processing, a novel method enabling automated safety monitoring across multiple interactions without retaining customer data, addressing limitations of Zero Data Retention (ZDR) systems. The approach processes customer content either on customer-controlled infrastructure or OpenAI-managed storage encrypted with customer-controlled keys, generating limited safety signals without exposing raw data. Early testing demonstrates enhanced pattern detection for misuse while maintaining strict data privacy, particularly crucial for enterprise deployments handling sensitive information. This development aligns with OpenAI's commitment to advancing AI safety without compromising customer control over data.

zero data retentionprivate safety processingautomated safety monitoringcustomer-controlled encryptionenterprise ai adoption

Replit expands access to software creation with GPT-5.6 Luna

OpenAI News · 2026-08-19

Replit introduces Free Mode powered by GPT-5.6 Luna, enabling broader access to software creation by eliminating token costs. This leverages improved price performance and scalable inference capabilities of GPT-5.6 series, facilitated by recent OpenAI price reductions. Free Mode integrates contextual project understanding, offering fast, accurate suggestions and feedback, while maintaining continuity between exploration and creation. Advanced reasoning tasks are routed to GPT-5.6 Sol, preserving project context. This approach aims to democratize software development, potentially increasing the number of builders by 100x and fostering a renaissance-level entrepreneurial boom.

gpt-5.6 lunatoken costsprice performanceproject contextscalable inference

ChatGPT Ads expands across Europe

OpenAI News · 2026-08-18

OpenAI expands ChatGPT Ads to 31 European markets, marking its largest geographic rollout since the U.S. pilot in February 2024. The platform integrates advanced advertising features, including conversion optimization, geo-targeting, custom audiences, and measurement tools via OpenAI Pixel and Conversions API. Ads are exclusively shown to Free and Go plan users, while Pro and Enterprise subscriptions remain ad-free, aligning with OpenAI’s mission to democratize AI access. The expansion enables advertisers to engage users during decision-making processes while maintaining privacy and transparency. Over tens of thousands of marketers have utilized the platform, driving iterative improvements based on user and business feedback.

conversion optimizationgeo-targetingcustom audiencesopenai pixelconversions api

Strengthening democratic oversight in national security

OpenAI News · 2026-08-18

OpenAI introduces an initiative to enhance democratic oversight of AI in national security, focusing on accountability, traceability, and institutional empowerment. The initiative includes $5M in training, technical support, and OpenAI credits for oversight bodies, alongside pilot tools for examining AI-assisted decisions. These tools aim to be interoperable and model-agnostic, ensuring oversight institutions retain control over evidence and findings. The approach aligns with OpenAI’s National Security Principles, emphasizing human judgment, safeguards, and ongoing monitoring. Progress will be measured by improved reviewer capabilities and public trust in government AI systems.

democratic oversightnational securitymodel-agnostictraceabilitysafeguards

How NVIDIA scales expertise with ChatGPT Work

OpenAI News · 2026-08-18

NVIDIA leverages ChatGPT Work to optimize knowledge workflows, automating manual tasks and scaling expertise globally. The system integrates external AI developments with internal priorities, enabling rapid prototyping and actionable insights. Results include a 16-hour weekly time saving during GTC planning cycles, reduction of prototype development from 2–3 weeks to 3–5 days, and distillation of 25–40 external AI updates into 5–8 actionable signals weekly. ChatGPT Work facilitates workflow automation across regions and functions, enhancing productivity and focus on strategic tasks.

chatgpt workworkflow automationrapid prototypingactionable signalsknowledge workflows

LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation

Hugging Face Blog · 2026-08-19

Quantization-Aware Distillation (QAD) improves the performance of quantized LFM2.5 models by distilling high-precision teacher models into Q4_0 quantized student models. The method achieves 96.5-97.4% recovery of BF16 baseline accuracy across four model sizes (230M, 350M, 1.2B, 2.6B) while maintaining Q4_0's memory efficiency and throughput. Benchmarks on GPQA Diamond, MMLU-Pro, IFEval, and GSM8K demonstrate QAD's effectiveness, with QAD Q4_0 checkpoints matching Q5_K_M and Q4_K_M quality at 3-33% higher decode throughput on edge devices like MacBook Pro and Raspberry Pi 5. The QAD GGUFs are compatible with llama.cpp and available on Hugging Face.

quantization-aware distillationq4_0bf16ggufedge deployment

5 new ways to level up your learning with Search

Google AI Blog · Awaneesh Verma · 2026-08-19

Google Search introduces five AI-powered educational tools to enhance learning efficiency. First, interactive visuals generated via AI Overviews and AI Mode enable dynamic concept exploration, such as pH scale simulations. Second, customized practice quizzes, leveraging partnerships with educational firms like The Princeton Review, provide exam-specific preparation. Third, Lens integrates interactive learning experiences for step-by-step problem-solving. Fourth, Gemini Notebook organizes study materials and syncs across Google products. Fifth, custom file creation streamlines study document generation from uploaded files or AI Mode threads. These tools are globally available in English, with broader language support planned.

interactive visualspractice quizzeslensgemini notebookcustom files

📜 arXiv Papers (234)

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

arXiv cs.AI · Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng · 2026-08-18

The authors introduce a capability-driven data infrastructure for generalist image generation, addressing the challenge of organizing heterogeneous supervision across generative capabilities. The framework employs three interoperable data engines for text-image grounding, inter-image transformation, and image-knowledge association, coupled with a multi-stage curriculum that evolves task composition, visual-concept distribution, data quality, and image resolution. At scale, it curates a 440M-image text-to-image corpus, 120M editing pairs, and 27M image-entity pairs. Training multimodal diffusion models at 3B and 6B scales demonstrates broad visual coverage, versatile rendering, and effective capability transfer, validated on CPI-Bench and diverse text-to-image and editing scenarios.

multimodal diffusion modelstext-image groundingcapability-driven curriculuminter-image transformationimage-knowledge association

On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification

arXiv cs.AI · Qinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang · 2026-08-18

This work identifies critical fragility in memory-based self-improving agents through comprehensive evaluation across variance and task order. Experiments reveal two key findings: (1) evaluation noise in complex environments amplifies with self-improving loops, and (2) performance is highly sensitive to task ordering due to implicit curricula. Manual memory analysis suggests task and environment underspecification as contributing factors, partially mitigated by incorporating detailed rubrics and environment feedback. Results demonstrate persistent performance gaps, indicating additional uncharacterized fragility sources. The study advocates for rigorous evaluation protocols and human oversight systems to address these reliability challenges.

self-improving agentstask orderingunderspecificationmemory bankevaluation variance

Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating

arXiv cs.AI · Daria Leshchikova, Valentina V. Kuskova, Dmitry Zaytsev, Valerii Klimov · 2026-08-18

The study examines asymmetrical user receptivity toward autonomous LLM agents in online dating platforms, revealing distinct constructs for sending versus receiving agent-mediated communication (rho=0.92, ΔBIC=52). Using graded response models with latent regression on two large-scale surveys (N=2,894; N=2,617), the authors quantify a delegation asymmetry: deploying one's agent requires lower receptivity (threshold -0.38) than engaging others' agents (+0.32), with mean deployment propensity threefold higher. Random-pairing counterfactuals show only 4-13% of dyads combine agent deployment with engagement, highlighting gender-directional imbalances. Design interventions (e.g., reciprocity requirements, receptivity-aware routing) significantly alter interaction volumes and engagement (AUC 0.88, 3.1x quartile lift).

autonomous agentslatent-variable modeldelegation asymmetrygraded responsereceptivity-aware routing

HLSR: Hybrid Live Forecast Selective Dynamic Vehicle Rerouting for Real-Time Congestion Avoidance

arXiv cs.AI · Xiao Wang, Shun Ren Yang, Hui Nien Hung · 2026-08-18

HLSR introduces a hybrid live–forecast vehicle rerouting framework for real-time urban congestion avoidance, combining live edge speeds with short-horizon forecasts under limited intervention scope. The method employs dual-threshold congestion detection, calibrated upstream selection, driver-tailored travel-time prediction, approaching-vehicle expansion, travel-time-weighted k-shortest-path generation, and horizon-dependent hybrid live–forecast segment speed for multi-cost route allocation. This selective approach optimizes rerouting efficiency while minimizing computational overhead compared to network-wide replanning.

congestion detectionshortest-path generationlive edge speedstravel-time predictionmulti-cost allocation

StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents

arXiv cs.AI · Yining Hua, Hongbin Na, Yifan Zhou, Akshay Kalose · 2026-08-18

StagedWorkspace introduces a versioned workspace framework for knowledge-work agents, addressing the challenge of managing multiple versions of digital artifacts (e.g., code, documents, spreadsheets). The workspace binds parsed records and review diffs to content hashes of native files, ensuring explicit versioning across workspace states. Evaluations on OfficeQA Pro and APEX-Agents benchmarks demonstrate performance improvements: dual parsed/native access boosts OfficeQA Pass@1 by 8.3-12.1 points and APEX mean rubric score by 4.7-9.2 points. SW-AGENT achieves 63.9% accuracy with Gemini 3.1 Pro on OfficeQA and 42.1 with GPT-5.4 Nano on APEX, surpassing baseline scores. Paired review-axis ablations further confirm higher scores when diffs are visible, highlighting workspace state as a critical experimental variable.

versioned workspaceknowledge-work agentscontent hashesparsed recordsreview diffs

Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System

arXiv cs.AI · Yi Wang · 2026-08-18

The paper introduces the Effectiveness--Losslessness Framework to explain why GPT-style models fail to transfer directly to symbolic music, despite their success in language. It defines tokenization as constructing a predictively effective and relationally lossless coordinate system, emphasizing the Fact--Token Boundary and Token--State Boundary. Controlled experiments demonstrate that effective coordinate construction enhances predictive compressibility, while fixed relational projections constrain contextual modeling. Results show that sequence compaction alone does not ensure predictive compression, and preserving contextual freedom allows higher-order musical organization to emerge without explicit structural labels. Tokenization must discover effective representations while maintaining relational freedom for contextual structure to emerge.

effectiveness-losslessness frameworkpredictive compressibilityfact-token boundarytoken-state boundarysymbolic music

Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach

arXiv cs.AI · Lu Xu, Xu Li, Linjiang Zheng, Fan Li · 2026-08-18

FlightLLM, a prior-guided semantic LLM-based approach, is proposed for interpretable flight safety analysis, addressing challenges of modal inconsistency, limited classification ability, and domain knowledge scarcity. The method combines feature engineering with statistical descriptors and flight indicators, employs Semantic Discretization for qualitative descriptions, integrates CatBoost for classification, and uses structured prompts with aviation-specific knowledge. Evaluated on 704 A320 flight samples focusing on hard landing events, FlightLLM achieves competitive classification performance while generating direct and reasonable explanations for event causes.

feature engineeringsemantic discretizationcatbooststructured promptsfew-shot learning

Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

arXiv cs.AI · Christophe D. Hounwanou, John Emeka Eze, Yaé U. Gaba · 2026-08-18

The paper formalizes hybrid LLM-planner and RL-controller architectures as Goal-Augmented Markov Decision Processes, demonstrating that using LLM-derived per-state progress scores as bounded potential functions preserves the optimal policy set even with inaccurate LLM scores. This policy-invariant reward shaping provides stronger guarantees than general LLM-as-reward approaches. Numerical verification on a small MDP under four configurations, including an adversarial case scaled to 20× the base reward magnitude, confirms the theoretical results.

goal-augmented markov decision processpolicy-invariant reward shapingllm-plannerrl-controllerbounded potential function

Traceable Trust for action-ready artificial intelligence in bioscience

arXiv cs.AI · Huayu Xin, Yizhi Cai, Mukilan Deivarajan Suresh, Gavin Michael Farrell · 2026-08-18

The authors propose Traceable Trust, a framework for assessing AI outputs in bioscience before they guide laboratory actions. The framework evaluates evidence supporting AI outputs, claimed capabilities, delegated agency, action thresholds, override mechanisms, and feedback loops for future decisions. Three case studies demonstrate its application across ecosystem resources, project design, and laboratory workflows. Results illustrate how Traceable Trust enables documented trustworthiness as AI outputs influence scientific processes, ensuring transparency and accountability in AI-driven bioscience research.

traceable trustbioscienceai outputslaboratory actiontransparency

Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media

arXiv cs.AI · Yijie Xu, Chao Wang, Hui Xiong · 2026-08-18

We propose TSN4PI, a unified framework for tracing evolving political ideologies on social media, addressing challenges of data scarcity, non-political content, and future ideological prediction. The framework comprises two modules: PIDN, which leverages large language models with style transfer and unsupervised domain adaptation for robust ideology detection and noise filtering, and PIPN, which employs temporal graph neural networks to predict ideological shifts. Evaluations on X and Truth Social demonstrate TSN4PI's effectiveness in analyzing ideology presence, intensity, and evolution. We release two large-scale datasets to support further research, advancing methodological development and empirical understanding of political polarization dynamics.

political ideology detectiontemporal graph neural networksstyle transferunsupervised domain adaptationideological shift prediction

Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity

arXiv cs.AI · Alisher Myrgyyassov, Zhen Song, Bruce Xiao Wang, Yu Sun · 2026-08-18

We propose Dual Co-Train, a source-free domain adaptation framework for ultrasound tongue segmentation that addresses cross-dataset domain shift under extreme data scarcity. The method employs a lightweight UltraUNet backbone pretrained on only five labeled source images, iteratively refines pseudo-labels, filters unreliable masks via a contour-based quality-control module, and generates target-style synthetic image-mask pairs using a segmentation-guided conditional GAN. Evaluated on 12 source-target transfer pairs across eight datasets, the framework improves segmentation overlap and contour accuracy over baselines, including supervised approaches, demonstrating the efficacy of task-specific pseudo-label refinement and synthetic target-style augmentation.

ultrasound tongue segmentationsource-free domain adaptationpseudo-label refinementconditional ganultraunet

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

arXiv cs.AI · Bin Li, Dongdong Wang, Siyang Lu · 2026-08-18

The paper introduces Log Reconstruction and Distance (LoRD), a post-hoc calibration framework addressing overconfidence in language model-based log anomaly detectors. LoRD learns reliability models from latent representations of correctly classified validation samples, using route-wise reconstruction distances to estimate prediction reliability and selectively recalibrate high-risk predictions. Evaluations on four large-scale log benchmarks show LoRD consistently improves confidence reliability, reducing overconfident errors by 37-52% without degrading detection performance (F1 scores maintained within 1% of baseline).

log anomaly detectionmodel calibrationreconstruction distancepost-hoc calibrationconfidence reliability

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

arXiv cs.AI · Isidoro Tamassia, Lennert De Smet, Giuseppe Marra · 2026-08-18

The paper introduces a neurosymbolic world model for reinforcement learning that decouples observation reconstruction from reward prediction by using structured symbolic components of the latent state. This formulation enables zero-shot task transfer to new reward functions defined over the same symbolic state space, without additional environment interactions. Experiments demonstrate superior generalization compared to purely neural world models, addressing the task-dependency limitations of traditional latent-space planning approaches.

neurosymbolicworld modelzero-shot transferlatent statereward prediction

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

arXiv cs.AI · Javier Aguilar Martín · 2026-08-18

The paper identifies a sampling-verification danger in continuous code world models, where accepted models may omit critical events due to mode-blind sampling. The authors formalize the danger as an expected risk, proving that the probability of missing an event with probability r over N i.i.d. samples is (1-r)^N. Experiments with GPT-5.x on hybrid instruments show that omitted modes (e.g., 1D clamps) are often repaired (105/111 cases), but 2D regions remain unrecoverable (0/156). A version-space certificate demonstrates that identifiability is class-relative, with accepted models certifying only sample consistency, not safety.

code world modelsampling-verificationmode-blindlipschitz constantversion-space

SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE

arXiv cs.AI · Xuan Zheng, Kento Uchida, Shinichi Shirakawa · 2026-08-18

SIGMA introduces a SHAP-guided, metadata-free framework for LLM-based Automated Feature Engineering (AutoFE), addressing limitations of trajectory-based prompting in long-horizon optimization. The method employs SHAP values for task-aware feature generation guidance and an EXposed-feature Implicit Trajectory (EXIT) approach to maintain stable generation without accumulating trajectories. Empirical results show SIGMA matches SOTA LLM baselines with constant prompt length, reduces feature duplication from 37.2% to 6.8%, and achieves traditional SOTA performance with only 5.4 features on average.

automated feature engineeringshap valueslarge language modelsimplicit trajectorymetadata-free

Procedural Content Metageneration via Program Search and Continual Abstraction Discovery

arXiv cs.AI · Matthew Siper, Ahmed Khalifa, Julian Togelius · 2026-08-18

The paper introduces Continual Abstraction Discovery (CAD), a method that improves evolutionary program search for procedural content generation by extracting reusable primitives from high-fitness programs into run-specific helper modules. The approach combines CAD with language-model mutation and crossover to evolve Python generators for Sokoban, Zelda, Dangerous Dave, and Lode Runner. Results from 160 runs show CAD increases mean final best fitness in all tested domains, with learned libraries frequently adopted for validation, reachability, and structural utilities.

procedural content generationevolutionary program searchcontinual abstraction discoverylanguage-model mutationreusable primitives

Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation

arXiv cs.AI · Zhizhao Liu, Zhiliang Tian, Xi Wang, Zhihua Wen · 2026-08-18

The paper proposes a graph-structured online difficulty estimator for efficient reinforcement learning with verifiable rewards (RLVR) scheduling. The method constructs a difficulty-aware sample graph using semantic and reasoning similarities, models latent difficulty states with a Potts prior, and employs a Beta-Binomial model for state-level feedback aggregation. An online mean-field variational algorithm updates state assignments and difficulty estimates continuously. Experiments show improved performance across multiple base models, RL schedulers, and benchmarks without dedicated probing overhead.

rlvrpotts priorbeta-binomial modelonline difficulty estimationmean-field variational

Grading Needs a Rubric, Not Intelligence

arXiv cs.AI · Jhen-Ke Lin · 2026-08-18

The study demonstrates that small language models achieve grading reliability comparable to expensive frontier models when using explicit rubrics, proposing 'any-to-bench' as a cost-efficient framework. In this approach, a frontier model extracts rubrics from source documents once, while smaller models handle repeated grading. Evaluating six configurations across two model families (3,456 grades), results show answer identity explains 95.6% of score variance, with judge identity contributing only 0.2%. Rubric criteria removal degrades reliability (ICC 0.888 to 0.628), while official answers dominate rubric effectiveness. No evidence of length or model-family bias was found.

language modelsgrading reliabilityexplicit rubriccost-efficiencyscore variance

EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection

arXiv cs.AI · Lei Jiang, Ye Wei, Xinyu Xi, Jordan Langham-Lopez · 2026-08-18

EvoTS-Agent introduces a self-evolving LLM agent for financial time series change-point detection, addressing the challenge of non-stationary and heterogeneous statistical properties. The agent employs a three-operator framework: Revision refines the current best solution, Alternative Strategy explores divergent modeling approaches during stagnation, and Recombination synthesizes insights from high-performing trajectories. Validation feedback guides trajectory evolution, enabling adaptive detection pipelines tailored to dataset characteristics. Evaluated across four benchmark datasets, EvoTS-Agent outperforms existing LLM-based agents while maintaining a 100% execution success rate across all backbone LLMs.

change-point detectionself-evolving agentvalidation feedbacknon-stationaryheterogeneous statistics

Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints

arXiv cs.AI · Chainarong Amornbunchornvej · 2026-08-18

The paper introduces Collective Counterfactual Planning (CCP), a formal model for group coordination under representational constraints, where agents operate through agent-specific subspaces of a shared task space. The model defines four gates (implementation coalitions, conception, consent, verification qualification) and the Collective Counterfactual Solvability (CCS) problem, separating geometric feasibility, executable attainment, and validated completion. Results show iterated cross-agent relay can enable solutions absent in one-shot planning, while subspace-dependent goals are unverifiable. A four-step horizon-bounded scheme is sound and complete under exact representation, with implications for sequential mutual enabling and sub-teaming.

collective counterfactual planningrepresentational geometryimplementation coalitionsiterated cross-agent relayhorizon-bounded solvability

Adaptive Policy Portfolios for Robust Markov Decision Processes

arXiv cs.AI · Kasper Engelen, Sebastian Junges, Guillermo A. Pérez, Marnix Suilen · 2026-08-18

The paper introduces adaptive policy portfolios for robust Markov decision processes (RMDPs), addressing the conservatism of optimizing a single policy against uncertain transition dynamics. The method synthesizes finite sets of memoryless randomized policies offline, paired with a lightweight online selector, and evaluates portfolio quality via robust regret. Complexity-theoretic analysis shows that portfolio certification is ∀ℝ-complete for deterministic portfolios in acyclic (s,a)-rectangular RMDPs, while synthesis is ∃∀ℝ-complete for general rational polytopes, even with fixed discount and acyclic dynamics. An offline portfolio construction method is presented, enabling runtime specialization.

robust markov decision processesadaptive policy portfoliosrobust regretmemoryless randomized policiescomplexity-theoretic analysis

A Theoretical Framework for Parallel Lifelong MAPF Using Group Decentralized Planning

arXiv cs.AI · Alex DeWeese, Jiaoyang Li, Guannan Qu · 2026-08-18

The paper presents Group Decentralized RHCR (GD-RHCR), a parallelized extension of the Rolling-Horizon Collision Resolution (RHCR) framework for Lifelong Multi-Agent Path Finding (L-MAPF). By partitioning agents via a transitive communication scheme and planning in parallel, GD-RHCR maintains RHCR's near-optimality guarantees while reducing computational cost. Theoretical analysis shows both frameworks achieve exponentially close-to-optimal solutions in a discounted MDP formulation. Empirical results demonstrate GD-RHCR scales to higher agent counts with lower per-plan cost while maintaining throughput across diverse maps.

lifelong mapfrolling-horizon collision resolutiongroup decentralized planningmulti-agent path findingdiscounted mdp

Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks

arXiv cs.AI · Matin Amoozadeh, Amin Alipour · 2026-08-18

The study analyzes question types in student-AI interactions during CS2 programming tasks using Graesser et al.'s 18-category taxonomy. A few-shot learning classifier was developed to categorize 830 student-AI interactions across two tasks. Results show a skewed distribution of question types, with a small subset dominating interactions, and significant evolution in question types as tasks progress.

student-ai interactionquestion classificationfew-shot learningcs2 programminggraesser taxonomy

Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition

arXiv cs.AI · Alma M. Liezenga, Lotte Nijskens, Henrik R. Baumann, Stefan Becker · 2026-08-18

This paper benchmarks out-of-the-box object detection models for military Automatic Target Detection and Recognition (ATD/R), addressing dataset scarcity by evaluating transfer learning from civilian datasets. The study tests six YOLO variants and two DETR-based models on a military dataset with occlusions and small targets, comparing mAP@0.5 and mAP@0.5:0.95 across perspectives and model sizes. Key findings show larger models outperform smaller ones, DETR architectures rival YOLO, VisDrone fine-tuning improves Air-to-Ground performance (+small object detection), but all models struggle with small A2G targets, underscoring the need for in-domain training.

automatic target detectionobject detectiontransfer learningyolodetr

AutoResearch: Insight In, Hallucination Out

arXiv cs.AI · Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang · 2026-08-18

AutoResearch introduces a two-stage autonomous research system combining Idea Generation and Idea Execution to ensure scientific rigor. The Idea Generation stage integrates emerging research signals with domain knowledge, employs multi-model generation and cross-review to produce testable plans, while Idea Execution decomposes plans into experiments, implements them iteratively, and conducts evidence-based review. Evaluated on cross-modal retrieval, systems optimization, and ML benchmarks (e.g., RSICD, improving mean Recall from 32.84 to 34.69), AutoResearch reduces audit-confirmed issues (5 vs. 11–27 in baseline systems) by grounding insights before experimentation and conclusions before acceptance.

autonomous researchidea generationevidence-based reviewcross-modal retrievalsystems optimization

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

arXiv cs.AI · Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov · 2026-08-18

We introduce BEAR-Bench, a bilingual English-and-Russian benchmark for evaluating multimodal large language models (MLLMs) on text-dense enterprise and academic documents. The benchmark comprises 1000 human-annotated questions based on professional materials, addressing limitations of existing benchmarks that are English- or Chinese-centric and lack focus on professional reasoning. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, observing significant headroom even for top-performing systems. Additionally, we analyze hallucination detection methods using model outputs, assessing both failure rates and reliability of failure identification.

multimodal large language modelshallucination detectionprofessional reasoningbilingual benchmarktext-dense documents

ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction

arXiv cs.AI · Samirasadat Jamalidinan, Yue Xu, Kazem Cheshmi · 2026-08-18

ARASH introduces an adaptive retrieval and shot selection method for tabular prediction, improving efficiency of Tabular Foundation Models (TFMs) like TabPFN through query-specific in-context learning. The approach selects optimal few-shot examples via local neighborhood analysis within training data, avoiding costly retraining. Results show ARASH reduces TabPFN's prompt length by 1261.5× and memory usage by 2.56× while maintaining comparable accuracy to standard approaches.

tabular predictionin-context learningfoundation modelsfew-shot promptingadaptive retrieval

Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

arXiv cs.AI · Man Liang, Xinzhao Cheng, Faizan Wajid · 2026-08-18

The study audits the gap between geometric constraint encoding and actionable behavior in frozen LLMs, using parametric CAD constraints as a testbed. Through probing hidden states of six decoder-only models, it examines linear decodability, forced-choice generation, activation-level influence, and steerability. Results show pretraining improves local relation decoding but sketch-level degree-of-freedom status is decodable even in randomly initialized models, with decodability not reliably translating to generation or steerability.

geometric reasoninglinear decodabilityactivation-level influenceparametric cadsteerability

AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis

arXiv cs.AI · Yangtian Liu, Yan Miao, Shuhan Liu, Yunfan Zhou · 2026-08-18

AdaLens introduces an interactive system for monitoring and steering long-running agentic data analysis workflows, addressing observability and steerability gaps in autonomous LLM-driven processes. The system employs a storyline-based representation unifying analytical plans, execution progress, intermediate findings, and data-column involvement, with steering interactions for directional guidance and execution control. Evaluation via two case studies and a user study demonstrates its effectiveness in supporting analysts during ongoing agentic analysis.

agentic workflowsobservabilitysteerabilityinteractive oversightstoryline representation

The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges

arXiv cs.AI · Maosen Zhang, Jianshuo Dong, Boting Lu, Wenyue Li · 2026-08-18

The paper introduces LeakGauge, a method for detecting context-leakage attacks in LLMs by analyzing prefill token probabilities without requiring hidden state access. The approach uses either content-specific or content-agnostic suffixes to gauge leakage behavior, mapping responses to an attack-risk score. Evaluated across 11 LLMs (including GLM-5.2 and Kimi-K3), LeakGauge achieves AUROC scores of 0.944--0.996, demonstrating robustness across languages and attack types. Activation-steering experiments confirm the signal's relation to internal leakage representations, while the detector adds minimal overhead (0.5K parameters, 10.34 ms latency).

context-leakageprefill probabilitiesactivation-steeringllm securitybehavior gauges

MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure

arXiv cs.AI · Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil, Subasish Das · 2026-08-18

The authors propose MotoSafety, an edge-AI architecture for two-wheeler collision risk assessment under time pressure (TP), introducing a dataset of 129,000 labeled multivariate time-series sequences from 51 participants. The method employs Learned Temporal Importance, achieving 94.97% accuracy (99.33% ROC AUC) and 4.4x lower forecasting error than Time-LLM/iTransformer, with 1.15M parameters and 0.135ms latency. Results show improved accuracy using TP as inductive bias (94.09%→94.97%) and strong transferability to human activity (97.66%) and clinical domains (99.65%).

edge-aitemporal importancecollision risk assessmenttime-series forecastingintelligent transportation systems

Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses

arXiv cs.AI · Alona Strugatski, Licol Zeinfeld, Jason Cooper, Shelley Rap · 2026-08-18

The study challenges the assumption that LLMs and humans share similar cognitive constructs by analyzing latent factors in assessment responses. Using Exploratory Factor Analysis (EFA) on human and LLM (six models) responses to quantitative reasoning and chemistry tests, Subject-Matter Experts (SMEs) blindly interpreted the resulting factor structures. SMEs successfully interpreted most human-derived factors but failed to ascribe meaning to any LLM quantitative reasoning factors and only half of chemistry factors, revealing statistically opaque mechanisms in LLMs distinct from human reasoning.

exploratory factor analysislatent factorscognitive constructsquantitative reasoningsubject-matter experts

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

arXiv cs.AI · Liya Zhu, Xin Ma, Tao Liu, Haodong Wang · 2026-08-18

The paper introduces StartupBench, a novel benchmark for evaluating general-purpose AI agents on end-to-end workflows derived from market-validated AI startup products. The authors systematically analyze adopted AI products to identify real-world tasks, translating them into deliverable-oriented evaluations with fine-grained rubrics. Testing representative models under a unified agent harness reveals only ~30% task completion rates, with failures attributed to complex instruction following and domain-specific expertise gaps, demonstrating significant limitations in current agents' real-world applicability.

agent benchmarkingend-to-end workflowsmarket validationinstruction followingdomain expertise

Training with synthetic data for drone detection in thermal imagery

arXiv cs.AI · Tanel Liiv, Sander Soodla, Nzamba Bignoumba, Alma M. Liezenga · 2026-08-18

The paper proposes a synthetic-first training strategy for Ground-to-Air (G2A) drone detection in medium-/long-wave infrared (MWIR/LWIR) imagery, addressing data scarcity and domain gaps. The method combines synthetic scene generation with fine-tuning on limited real thermal data, demonstrating that synthetic data enables effective initial representation learning while real data remains essential for deployment. Experiments show dataset alignment impacts performance more than model scale, with semantic feature alignment and radiometric properties (entropy, dynamic range) being key predictors of detection robustness.

synthetic datathermal imagerydomain adaptationdrone detectioninfrared sensors

Learnware for CSI Feedback: Scene-specific Small Models Can Do Big

arXiv cs.AI · Xiangyi Li, Jiajia Guo, Chao-Kai Wen, Xin Geng · 2026-08-18

The paper proposes a Learnware-based framework for efficient deployment of scene-specific CSI feedback models in 6G systems, addressing the generalization-vs-specialization trade-off in deep learning solutions. A centralized repository stores models with semantic (architecture parameters) and statistical (codebook-fingerprint embeddings) specifications, enabling base stations to retrieve relevant pre-trained models via local statistical data without raw CSI transmission. The data-driven search strategy achieves >90% selection accuracy, yielding 18.8% (LOS) and 57.7% (NLOS) performance gains over general models while reducing fine-tuning by 1000 samples and 100 epochs.

csi feedbacklearnware6g systemscodebook-fingerprintmodel repository

D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory

arXiv cs.AI · Xule Liu, Yijun Liu, Chao Li, Shao Kun · 2026-08-18

The paper introduces D$^2$ACCI, a dual-loop diagnostic protocol for evaluating and improving LLM agent memory systems by localizing failures across ingestion, retrieval, filtering, and generation stages. The method combines outer diagnostic gates with graded observability metrics (DCR) and reusable evaluation artifacts (D$^2$ACCI-Eval), enabling evidence-based intervention decisions. Evaluations on LoCoMo (93.59%), LongMemEval (90.93%), and PersonaMem-V2 (57.20%) show statistically significant gains from key components (+1.9 to +3.7pp, p ≤ .003), while diagnostic traces improve root-cause agreement (98–100% DCR@3 vs. 0% for result-only logs).

llm agentsdiagnostic protocolmemory systemsgraded observabilitylocalizability

The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting

arXiv cs.AI · Nazlı Nur Karabulut, tanya Braun · 2026-08-18

The paper introduces policy-counted DecPOMDPs, a novel approach to address the exponential complexity in decentralized partially observable Markov decision processes (DecPOMDPs) by counting policies rather than agents. The method employs policy-counted dynamic programming to efficiently solve these models, leveraging compact representations that reduce evaluation costs to polynomial dependence. This approach maintains tractability in agent numbers while mitigating the policy space explosion inherent in traditional DecPOMDP formulations.

decpomdpspolicy countingmulti-agent systemsdynamic programmingexponential complexity

Evaluating the Diversity of AI-Generated Content with Diversity Profiles

arXiv cs.AI · Xiuyuan Hu, Xuege Hou, Guoqing Liu, Yang Zhao · 2026-08-18

The paper proposes diversity profiles as a curve-valued alternative to scalar metrics for evaluating diversity in AI-generated content, addressing inherent ambiguities in existing approaches. Through axiomatic and empirical analysis, the authors demonstrate that no single scalar metric satisfies all desirable properties and that distance distributions in high-dimensional spaces are modality-dependent. Diversity profiles evaluate parameterized metric families across thresholds and scales, providing resolution-aware comparisons; experiments show their utility in generative AI evaluation.

diversity profilesgenerative aiscalar metricshigh-dimensional representationsaxiomatic analysis

What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

arXiv cs.AI · Xiaonan Xu, Wenjing Wu · 2026-08-18

The study reveals that aggregate benchmark scores mask bidirectional item-level performance changes during commercial LLM API migrations. Analyzing 900 items across three GPT-5.4 to GPT-5.6 Sol migrations with 50 queries per item, researchers classified items as reliably improved, regressed, equivalent, or inconclusive using false-discovery-rate control and practical-significance thresholds. Results show concurrent improvements and regressions (up to 8.3% regressed items in upgrades and 10.7% improved items in downgrades), with instruction-following benchmarks exhibiting a 3.9-point scoring gap between strict and loose criteria.

llm api migrationbenchmark regressionfalse-discovery-rate controlpractical-significance thresholditem-level analysis

Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents

arXiv cs.AI · An He, Yao Wang, Haibin Zhang · 2026-08-18

The paper introduces ontological trust, a task-conditioned property for evaluating trajectory prefixes in long-horizon agents, addressing drift in role, goal, or evidence despite locally valid actions. The authors propose RGE, an online monitor that decomposes trust into Role, Goal, and Evidence components, using LLMs for structured representations while maintaining deterministic trust-state updates for auditability. Evaluated on a cross-domain corpus (OSWorld, FinanceBench, EICU-AC), RGE achieves >93% Drift F1 and ≥95.8% benign coverage, outperforming rule-, judge-, and shield-style baselines, though pseudo-consistency detection remains challenging depending on task visibility.

ontological trustlong-horizon agentstrajectory prefixesdrift detectiononline monitoring

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

arXiv cs.AI · Sahab Zandi, Noah Kostesku, Christophe Mues, María Óskarsdóttir · 2026-08-18

The study evaluates whether Large Language Models (LLMs) can effectively translate post-hoc explanation artefacts from credit risk models into stakeholder-appropriate narratives. Using Freddie Mac single-family loan-level data, the authors develop three pipelines: standard tabular (XGBoost + SHAP), pure network-based (GNN + GNNExplainer), and a bimodal one combining both. They generate narratives with three LLM configurations (Gemma 3 4B, DeepSeek R1 70B, Gemini 2.5) and evaluate explanation quality through automated checks and a human study. Key findings include the pipeline's greater impact on evidence-grounding scores than the LLM, reliable naming of influential factors but less reliable directionality, and stricter evidentiary standards among professionals.

credit risk modelslarge language modelspost-hoc explanationsgnnexplainerevidence-grounding

Accuracy and Robustness of Model Cascades Under Data Perturbations

arXiv cs.AI · Pallavi Mitra, Jai Kushwaha, Felix Biessmann · 2026-08-18

This paper analyzes the robustness of confidence-based model cascades for image classification under input data perturbations. The authors construct a Pareto-optimal cascade achieving competitive accuracy with 10× lower CO₂ emissions, then evaluate its routing behavior under static corruptions and sequential perturbations. Results reveal three failure modes: (1) broken routing with preserved large-model utility, (2) simultaneous degradation of both models, and (3) stable but unreliable predictions from suppressed deferral. The study demonstrates that cascade reliability requires evaluation beyond clean accuracy, particularly examining routing stability under distribution shift.

model cascadesconfidence-based routingdistribution shiftpareto-optimalinput perturbations

Dijkstra as an Oracle for Online Stochastic Shortest Path Navigation with Provable Guarantees

arXiv cs.AI · Mansur M. Arief, Ali Akarma, Ahmad Alfan Alfian Irfan · 2026-08-18

The study demonstrates that Dijkstra's algorithm remains exact for stochastic shortest path problems under a weakened nonnegativity condition on reduced costs in determinized maps. It introduces DORA (Dijkstra Oracle Reduced-cost Algorithm), an online learner that invokes a shortest path oracle without transition kernel estimation, incorporating logarithmic survival weights for dynamic obstacle avoidance. Experiments in grid world navigation, directional drilling, and drone surveillance show DORA matches optimistic value iteration's performance while reducing planner work by 4.5-19.3×, cutting contact rates by 17×, and maintaining safety across varying budgets.

stochastic shortest pathdijkstra's algorithmonline learningreduced-cost conditiondynamic obstacle avoidance

GADR: Gathering Architecture Decision Records from Meeting Transcriptions

arXiv cs.AI · Lucas Daniel Costa da Silva, Kiev Gama · 2026-08-18

GADR introduces a multi-agent, self-correcting workflow for extracting architectural decisions from unstructured meeting transcripts and generating Nygard-formatted ADR drafts, addressing the limitations of single-pass LLM approaches. The method employs an agentic workflow to handle implicit, fragmented decisions in noisy dialogue. Evaluation on five real project transcripts with expert review (4 architects) and student assessment (15 participants) demonstrates superior performance over zero-shot and few-shot baselines in decision capture, clarity, and structural adherence, while revealing trade-offs in RAG-based enrichment fidelity.

architecture decision recordsmulti-agent workflowmeeting transcriptionrag-based enrichmentnygard format

Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals

arXiv cs.AI · Joao Fonseca, Rodrigo Rodrigues, Paolo Romano · 2026-08-18

The paper introduces InnerExpert, a novel method for per-token hallucination detection in Mixture-of-Experts (MoE) LLMs by leveraging previously unexploited internal signals (router entropy, expert disagreement, usage patterns). It combines these MoE-specific features with standard transformer signals into compact per-token vectors, classified by a lightweight detector trained via LLM-as-a-judge annotations. Evaluated across five datasets and two MoE architectures, InnerExpert achieves 0.91 answer-level and 0.76 token-level AUROC, outperforming existing methods with only a single forward pass.

mixture-of-expertshallucination detectionrouter entropyper-token classificationllm-as-a-judge

Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch

arXiv cs.AI · Jialong Li, Jialing Zhu · 2026-08-18

The study audits self-evolving financial agents (SkillOpt, Agent Workflow Memory, ReasoningBank) using simulated e-banking to measure capability gains, security drift, and execution mismatches. Methods include matched benign trajectories, sealed evaluation endpoints, and independent state replay. Results show SkillOpt improves utility (0.741→0.837) but increases exposure (0.820→0.943) and unauthorized state changes (0.685); ReasoningBank achieves higher utility (0.859) without raising attack success rates. Agent Workflow Memory reveals execution-interface mismatches, with WebArena text-action envelopes reducing utility (0.756→0.319).

self-evolving agentssecurity driftexecution-interface mismatchattack success rateunauthorized state changes

Benchmarking Automated Security Patch Backporting: How Far Are We?

arXiv cs.AI · Jincheng Yang, Yulong Fu, Chengwei Liu, Lyuye Zhang · 2026-08-18

The study introduces Porting Benchmark, a dataset of 1,234 security patch backporting cases across cross-version, cross-branch, and cross-repository scenarios, with a unified evaluation framework. It evaluates five tools (PortGPT, TSBPort, FixMorph, Mystique) using program analysis, LLM prompting, and LLM agents, revealing performance drops under aligned evaluation, especially for complex patches (85.2% to 24.0% success). Key failure categories include API awareness and semantic mismatches. Dynamic validation on 45 cases shows reference-based metrics under-credit adaptations, with executable feedback offering limited recovery.

security patch backportingcross-versionllm agentsexecutable validationsemantic mismatch

GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities

arXiv cs.AI · Haoran Bu, Zejian Chen, Litian Zhang, Xi Zhang · 2026-08-18

The paper introduces GraphWake, a novel threat model called Memory-Mediated Polarization Cascade that exploits LLM-driven agent communities to induce group polarization. The method involves three stages: (1) exposing target agents to stance-reinforcing arguments for memory retention, (2) triggering retrieval and reproduction via stance-neutral discussions, and (3) propagating polarized views through untreated agents. GraphWake implements this using stance-support argumentation knowledge graphs, axiom-oriented triple selection, and stance-neutral memory cueing. Experiments demonstrate significant polarization increases across multiple discussions and memory systems, revealing community-level risks.

llm-driven agentsgroup polarizationmemory-mediated cascadeargumentation knowledge graphsaxiom-oriented selection

MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps

arXiv cs.AI · Sujin Chen, Lijun Li, Tianyi Du, Jing Shao · 2026-08-18

The paper introduces MobileWorldSafety, a benchmark of 142 risk tasks on real Android applications to evaluate GUI agent safety against environmental injection attacks. The method employs programmatically verifiable risk indicators and a two-stage evaluation pipeline combining rule-based verification and LLM adjudication. Results show six tested agents remain highly vulnerable, with attack success rates ranging from 40.4% to 66.9%, highlighting safety alignment failures when processing adversarial mobile content.

environmental injection attacksgui agentsandroid applicationssafety alignmentllm adjudication

LLM-Derived Preference Judgments Are Not Self-Consistent

arXiv cs.AI · Matthew T. Ford, Francis Bahk, Jingjing Wang, Adam S. Jovine · 2026-08-18

The paper demonstrates that LLM-derived cardinal preference judgments lack self-consistency, challenging their use in utility function estimation. The authors develop statistical tests and interpretable measures to quantify deviations from self-consistent utility functions, using flight, apartment, and hotel preference scenarios across six LLMs. Experiments reveal persistent inconsistencies, indicating that LLM-generated preferences cannot be accurately modeled by a single utility function.

preference judgmentsutility functionself-consistencyllmcardinal preferences

Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing

arXiv cs.AI · Kang Chen, Sihan Zhao, Yixin Cao, Yugang Jiang · 2026-08-18

The paper introduces J64, a 64-axis semantic frame distilled from vocabulary-scale reasoning states in mixture-of-experts models, coupled with R64, a routing-statistics proxy. J64 reveals latent process states (separating inference effort from problem-induced strain) and improves held-out AUC by 0.096-0.135 over token-occupancy baselines. R64 reconstructs J64 from native expert routing (median axis correlation 0.69-0.86) while preserving 95-100% of predictive gains. Applications include test-time branch selection (improving majority voting in 7/8 settings) and generation control (1.1-5.9 accuracy gains via stop-and-resample policies). Router edits targeting J64-mechanisms induce predicted reasoning behaviors.

mixture-of-expertsreasoning-statesemantic frameexpert-routinglatent process

Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models

arXiv cs.AI · Satpreet Makhija · 2026-08-18

The paper establishes a precise mathematical correspondence between graph surgery and the do-operator in acyclic structural causal models (SCMs) with finitely many endogenous variables. It proves that replacing target mechanisms via the do-operator (functional operation) removes exactly the same dependencies as graph surgery (graphical operation), formalized as $\operatorname{Graph}(F^ι)=\operatorname{Surg}(\operatorname{Graph}(F),T_ι)$. For models with unused graph arrows, it characterizes when this equality holds universally. Additional contributions include defining intervened models, characterizing their execution, showing sequential intervention composition, and proving outcome dependency on ancestral interventions.

structural causal modelsdo-operatorgraph surgeryinterventiondependency analysis

DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval

arXiv cs.AI · Jingyuan Wang, Richong Zhang, Zhijie Nie, Mingxin Li · 2026-08-18

DEPT (Document Embedding Preservation Tuning) introduces an end-to-end approach for unified query expansion and retrieval by training a single decoder-only LLM to generate expansions and encode both queries and documents. The method addresses the moving-target problem in joint optimization by preserving tuned document embeddings close to initial cached embeddings via straight-through decoding, enabling retrieval gradients to update the generator while maintaining stable whitened document embeddings for index reuse. Experiments with Qwen3-4B-Instruct-2507 and LLaMA-3.2-3B-Instruct on five BEIR datasets show DEPT outperforms training-free, independently trained, and staged unified baselines, with ablations validating the contributions of preservation, whitening, end-to-end training, and online hard-negative mining.

document embedding preservationquery expansionretrieval optimizationstraight-through decodingonline hard-negative mining

Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision

arXiv cs.AI · Amir Arsalan Nematollahi, Shayan Ahmadi, Mehdi Tale Masouleh, Ahmad Kalhor · 2026-08-18

The paper presents a reinforcement learning framework for robotic grasp refinement using 2D vision, combining keypoint-based object representations with Deep Q-Networks (DQN). Initial grasp candidates from a geometric algorithm are iteratively refined in simulation, achieving a 100% success rate on previously ungraspable objects from the Dex-Net dataset (300 objects tested). Physical validation on a Delta parallel robot confirms sim-to-real transferability, successfully manipulating an object that was previously ungraspable.

deep reinforcement learninggrasp refinementkeypoint representationsim-to-real transferdeep q-network

Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)

arXiv cs.AI · AlAnoud AllGhayth, AlJawharh AlOtaibi, Jude AlSubaie · 2026-08-18

The study presents a validated drone-based crowd monitoring system for mass gatherings, addressing label-free adaptation and crush risk detection ahead of FIFA World Cup 2034. Methodologically, it combines 525 controlled runs, corpus analysis, falsification tests, and safety evaluations, achieving 31-49% error recovery across corruptions via label-free adaptation (41.8 MAE improvement, p=7.5x10^-10). Key results include a severity law for method differentiation, stability budgets for deployment safety, and successful congestion detection in real footage (2/6 clips). The analysis reveals normalization-driven adaptation (not flow-driven) and quantifies shift gate limitations (58% headroom loss).

label-free adaptationcrowd countingdrone surveillancedomain shiftrisk detection

From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support

arXiv cs.AI · Ngoc Luyen Le, Marie-Hélène Abel, Bertrand Laforge · 2026-08-18

The paper introduces SC2R, a semantics-constrained counterfactual recourse framework for educational decision support, addressing the gap between risk prediction and actionable interventions. The method combines calibrated predictive modeling, integer-programming-based recourse generation over discrete actions, an RDF vocabulary for intervention representation, and SHACL validation for constraint enforcement. Evaluated on the OULAD dataset at two decision horizons, results demonstrate strong predictive performance, scalable generation of compact intervention plans, and semantic validation of feasibility that optimization-only approaches miss.

counterfactual recourseeducational decision supportinteger programmingrdf vocabularyshacl validation

MoNe: Modular Neural Memory for Efficient Long Context Inference

arXiv cs.AI · Wonguk Cho, Kyubyung Chae, Tribhuvanesh Orekondy, Sunghyun Park · 2026-08-18

MoNe introduces a modular neural memory system for efficient long-context inference in frozen pretrained Transformers, eliminating retraining. The method processes context in fixed-size segments using test-time learning with fast-weight neural memory networks and layer-localized gradient updates, enabling $O(N)$ preprocessing and $O(1)$ query cost. At 128K tokens, MoNe reduces compute and peak GPU memory by ~80% compared to in-context learning (ICL), with only 6.4% parameter overhead, and maintains strong performance on RULER benchmarks where ICL fails.

modular neural memorylong-context inferencefast-weight networkslayer-localized gradientstest-time learning

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

arXiv cs.AI · Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam · 2026-08-18

The survey analyzes multi-turn conversational AI systems, contrasting progress in multimodal perception with persistent challenges in session coherence. It systematically reviews text-only dialogue, AudioLLMs, multimodal systems, and tool-augmented agents through the lenses of datasets, modeling paradigms, training strategies, and evaluation. Findings reveal that while modality support has advanced significantly, systems still struggle with cross-turn grounding (76% accuracy drop after 5 turns in benchmarks), persistent memory, full-duplex interaction, and cultural alignment. The authors propose a research agenda addressing memory, grounding, and adaptive multi-turn interaction across modalities.

multiturn dialogueaudiollmscross-turn groundingfull-duplex interactionomni-modal systems

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

arXiv cs.AI · Yajing Bai, Jinhao Duan, Jie Peng, Xianfeng Wu · 2026-08-18

The paper introduces HarnessRisk, a lifecycle-oriented benchmark for evaluating safety failures in agent harnesses across six operational phases: Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery. The benchmark comprises 128 sandboxed cases pairing benign objectives with adversarial instructions in untrusted workflow artifacts, evaluated on Utility, Attack Success Rate, Persistence, and Detection metrics. Results across three harnesses and six language models show attack success rates from 12.6% to 80.9%, with Harness Configuration being the most vulnerable phase, and highlight that risk recognition does not always prevent unsafe actions.

agent harnesssafety benchmarklifecycle phasesadversarial instructionrisk recognition

tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots

arXiv cs.AI · Markus D. Kobelrausch, Michael Miedler, Axel Jantsch · 2026-08-18

The paper presents tinyDSM, a framework for autonomous skill development in resource-constrained millirobots (36 cm³ volume) using minimal a-priori knowledge. The method combines reinforcement learning with intrinsic motivation and fitness-based assessment, implemented via a hierarchical knowledge graph and kinematic reasoners on a Raspberry Pi Pico microcontroller (9 kB footprint). Experiments demonstrate autonomous progression from basic motor skills to complex geometric patterns within 15 minutes, with complementary simulation analysis of learning algorithms and motivation parameters.

millirobotsintrinsic motivationkinematic reasonershierarchical knowledge graphresource-constrained

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

arXiv cs.AI · Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang · 2026-08-18

TRUSS introduces an evidence-guided framework for generating functionally effective and safety-reliable Agent Skills, addressing limitations in artifact inspection and final task outcome evaluation. The method combines static analysis (nine safety properties) with dynamic execution in a Controllable Execution Environment, linking failures to Skill content for iterative refinement. Evaluated on 168 SkillInject artifacts, 155 SkillSafetyBench cases, and 187 SkillGenBench tasks, TRUSS achieves 100% precision/recall in vulnerability detection, reduces attack success by 19.36-16.77 percentage points across GPT versions, and improves task effectiveness from 17.11% to 52.94% while ensuring 100% security.

agent skillscontrollable execution environmentsafety propertiesprovenance preservationiterative refinement

Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making

arXiv cs.AI · Deep Kumar Ganguly, Jan Kretinsky · 2026-08-18

The paper introduces RATTL (Risk-Adversarial Total-Reward Learning), a method for safe sequential decision-making under epistemic uncertainty. RATTL uses a Bayesian posterior over dynamics and plans against a Wasserstein ambiguity set with radius tied to posterior uncertainty, enabling adaptive interpolation between worst-case robustness and risk-neutral optimization. Theoretical analysis shows the value function lies between uninformed robust and full-knowledge optima (Safety Sandwich theorem), with convergence as the posterior concentrates. In a binary-hazard case, RATTL reduces to Conditional Value-at-Risk with entropy-dependent risk levels. The method targets runtime safety for uncertain agents, including LLM-based systems.

sequential decision-makingepistemic uncertaintywasserstein ambiguityconditional value-at-riskbayesian posterior

DMT-Dens: Density-preserving manifold visualization for biological data

arXiv cs.AI · Ruizhe Wang, Yixuan Dong, Bolin Yang, Bingo Wing-Kuen Ling · 2026-08-18

DMT-Dens introduces a parametric manifold-visualization method for biological data that preserves sampling density while maintaining neighborhood structure. The approach employs a latent-token Transformer encoder with rank-based manifold alignment and hard-pair aggregation, optimizing a Pearson correlation loss between k-nearest-neighbor log-radius estimates in input and embedding spaces. Benchmarks show superior density preservation on biological datasets compared to existing methods, alongside competitive label separability. The implementation is publicly available.

manifold visualizationtransformer encoderdensity preservationbiological datak-nearest-neighbor

Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries

arXiv cs.AI · Henrik Wille, Luis-Finley Schütz, Felix Strieth-Kalthoff · 2026-08-18

The study demonstrates that domain adaptation significantly enhances molecular language models' performance in virtual screening across diverse chemical spaces. Benchmarking four pretrained models on six libraries (drug discovery, organic materials, catalysis), the authors found native embeddings exhibit variable performance while molecular fingerprints provide robust baselines. Fine-tuning encoders on target-domain structures improved sample efficiency, with adapted models outperforming original representations. Results indicate domain-representation mismatch and establish domain adaptation as a key strategy for sample-efficient molecular discovery in self-driving laboratories.

molecular language modelsdomain adaptationvirtual screeningmolecular fingerprintssample efficiency

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

arXiv cs.AI · Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li · 2026-08-18

This work investigates cross-task usability in unified multimodal models (UMMs) by isolating concept binding directions (understanding vs. generation) through controlled experiments with novel visual entities. Using a rendered 3D asset paired with a pseudo-word, the study demonstrates that generation training enables name matching while understanding training enables production, with cross-task usability governed by entry point location in shared computation. An alignment probe predicts concept export (Spearman ρ=+0.68), and mid-stack alignment achieves concept acquisition with only 0.1% relative loss in text-to-image ability versus 41% for standard generative methods.

unified multimodal modelsconcept bindingalignment probecross-task usabilitysemantic vision encoder

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

arXiv cs.AI · Jack Boylan, Chris Hokamp · 2026-08-18

The paper introduces Action-Contrastive Masked Transition Modeling (AC-MTM), a method to prevent representation collapse in Joint-Embedding Predictive Architectures (JEPAs) without requiring Gaussian latent priors. AC-MTM augments forward latent prediction with a contrastive inverse-dynamics head trained via Action-NCE, which discriminates actions causing latent transitions within a batch. This training-only component is discarded post-training, maintaining identical test-time computation to baseline LeWM. Evaluated on four pixel-control tasks and OGBench Visual Scene, AC-MTM matches SIGReg's performance on average (80.0±2.0% vs 58.0±2.0% success on OGBench) while eliminating Gaussian distribution constraints. The method operates without target networks, stop-gradients, or reconstruction objectives.

joint-embedding predictive architecturesrepresentation collapsecontrastive learninginverse dynamicslatent prediction

CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method

arXiv cs.AI · Jin Su, Zhuofeng Zhao, Huanhuan Wang, Hao Chen · 2026-08-18

The paper proposes CoAL-RAG, a complexity-aware legal retrieval-augmented generation method that adaptively routes retrieval strategies based on multi-dimensional evaluation of question essence and retrieval consistency. The method quantifies reasoning demand via question structure and leverages semantic-keyword retrieval discrepancy to select optimal strategies with dynamic context filtering. Experiments show 42.5% BLEU improvement and 3.6x ROUGE-L gains over knowledge graph baselines on Chinese datasets (SocialLawQA, LawBench), with strong cross-jurisdictional performance on LexGLUE and CaseHold.

retrieval-augmented generationcomplexity-awaresemantic retrievalkeyword retrievaladaptive routing

When to Review: Spaced Repetition for Continual Pre-Training of Language Models

arXiv cs.AI · Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug · 2026-08-18

The paper introduces Spaced Repetition Training (SRT), a continual pre-training framework for language models that adaptively schedules example reviews based on forgetting rates, inspired by cognitive science's SuperMemo-2 algorithm. SRT maintains per-example review state, uses perplexity as a recall-quality signal, and dynamically balances retention of old knowledge with acquisition of new knowledge without modifying the base model architecture. Experiments on Wikipedia and code corpora show SRT recovers 5-37 percentage points of old-knowledge accuracy lost by naive continual pre-training across model scales while maintaining new-knowledge performance, and preserves benchmark performance better than uniform replay at larger scales.

continual pre-trainingspaced repetitionadaptive schedulingknowledge retentionperplexity signal

Agent Lightning v1.0: Towards Harnessed Agentic RL

arXiv cs.AI · Zhiyuan He, Siwei Zhang, Zhiwen Zhou, Yuqing Yang · 2026-08-18

Agent Lightning v1.0 introduces a lightweight framework for harnessed agentic RL, where deploy-time agent harnesses directly participate in model post-training. The 3,500-line system addresses challenges in retokenization, sample merging, and advantage calculation arising from the harness-controlled environment interaction loop. Evaluated on instruction-following, search, and coding tasks, it improves Qwen3.5-9B's SWE-bench Verified performance from 41.8% to 56.4% using only 6K training examples.

harnessed agentic rlllm endpoint proxyretokenizationswe-benchadvantage calculation

Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery

arXiv cs.AI · Mohammad Javad Ahmadi, Hamid D. Taghirad · 2026-08-18

The study introduces an explainable AI framework for automated skill assessment in cataract surgery, leveraging the world's largest dataset of 2,000 surgical videos. The method combines computer vision and signal processing to extract ten objective motion-based metrics, validated against expert ratings using the Capsulorhexis Skill Assessment System (CSAS). Experimental results on 83 videos show 87% accuracy in correlating automated metrics with subjective evaluations, demonstrating robust modeling of surgical expertise.

explainable aicomputer visionskill assessmentcataract surgerysignal processing

Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task

arXiv cs.AI · Enrique Barba Roque, Luís Cruz, Annibale Panichella · 2026-08-18

The paper proposes energy-aware knowledge distillation for sustainable LLMs in software engineering tasks, demonstrating that FLOPs is an unreliable energy metric. Using Morph, a Many-Objective Optimization-based distillation method with energy-surrogate models, the authors optimize models for Clone Detection, Vulnerability Prediction, and Code Summarization (via CodeT5+). Results show up to 90% reduction in inference energy and 86% lower memory usage with minimal accuracy loss, validating energy surrogates as superior to FLOPs for efficiency optimization.

knowledge distillationenergy-aware optimizationflopscode summarizationmany-objective optimization

SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models

arXiv cs.AI · Sarvesh Gharat, Junpei Komiyama · 2026-08-18

The Structural Gap Hypothesis Agent (SGHA) is introduced as a fully automated, corpus-first system for research-problem discovery using local language models. SGHA structures scientific literature into evidence-linked paper objects and a typed evidence graph, detects unresolved structural patterns, and formulates traceable research problems with assumptions, objectives, and success criteria. All components run on a local 9B open-weight LLM, avoiding proprietary APIs. Compared to AI Scientist-v2, SGHA demonstrates that explicit corpus structure and evidence-constrained reasoning enable inspectable problem formulation without frontier-model dependence.

evidence-linked objectstyped evidence graphstructural gap detectionlocal llmresearch-problem formulation

Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context

arXiv cs.AI · Yiwen Zhao, Zhihao Wen, Yuchen Mao, Mingxuan Jiang · 2026-08-18

The paper introduces Feedback-Aware Credit Assignment (FACA), a method for improving multi-turn user-interacting agents by leveraging next-turn user reactions as local credit signals. FACA aligns each reaction with the preceding user-to-user segment, computes a locally normalized reaction advantage, and combines it with terminal outcome advantage without additional critics or rollouts. Evaluated against an outcome-only Interactive GRPO baseline, FACA improves the nine-domain τ-family average by 5.91 and 10.22 percentage points at 8B and 14B parameters respectively, with gains concentrated in Telecom. Zero-shot results on Pare-Bench and Co-Gym confirm the approach's effectiveness.

multi-turn interactioncredit assignmentreinforcement learninguser feedbackdialogue agents

When AI Designs AI: Innovation or Imitation?

arXiv cs.AI · Yikang Yang, Zhengxin Yang, Luzhou Peng, Minghao Luo · 2026-08-18

The paper investigates whether LLM agents can innovate beyond human-designed AI methods by analyzing their algorithmic designs and performance across multi-modal tasks. It introduces a framework to map both human- and agent-designed methods into task-specific algorithmic design spaces, quantifying differences at the module level. Results show agents match human SOTA in 10/72 configurations but largely reuse human-derived algorithmic components (96.8% within human design spaces, 50% exact matches), indicating limited innovation beyond recombination.

llm agentsalgorithmic design spacesmodule-level analysismulti-modal taskssota performance

SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution

arXiv cs.AI · Maolin Ran, Xiaoyang Lu, Jiaqi Liu, Jian Wang · 2026-08-18

SAGE introduces a self-evolving framework for automating storyboard generation by acquiring, refining, and injecting directing knowledge from expert demonstrations. The method contrasts training screenplays with expert storyboards to derive content-independent rules, attributes decisions to specific rules during generation, and evolves rules via localized feedback. Scenario packages with routing indexes enable bounded retrieval. Evaluated on 18 test episodes, SAGE scored 77.8 (vs. 77.1 for professionals) and reduced authoring time by 83% in a 14-day deployment. The PROSE dataset, pairing 68 screenplays with professional storyboards, is released.

storyboard generationknowledge refinementattribution-guided evolutionscenario packagesrouting index

Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning

arXiv cs.AI · Xingrui Zhuo, Jiapu Wang, Manzong Huang, Gongqing Wu · 2026-08-18

The paper proposes a Structure-Internalized Rule Language Model (SIRLM) to address reasoning evidence perception drift in Knowledge Graph Reasoning (KGR) by aligning Large Language Model (LLM) parametric knowledge with KG structural constraints. SIRLM integrates a Structure-Internalized Rule Generator (SIRG) featuring a structural relation memory, KG tokenizer based on structural invariance learning, and a neuro-symbolic reasoner for rule-constrained message propagation. Evaluated against 17 state-of-the-art KGR methods on 36 datasets, SIRLM demonstrates significant performance improvements in faithfulness and effectiveness.

knowledge graph reasoninglarge language modelsstructural invariance learningneuro-symbolic reasonerin-context learning

Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression

arXiv cs.AI · Tao Jiang, Minbo Gao, Shaowei Cai · 2026-08-18

The work establishes that quadratic depth dependence is intrinsic for Gaussian regression over the Parhi--Nowak deep-RBV^2 architecture, resolving a gap between prior lower and upper bounds. Through local packing construction with log-cardinality Ω(L²w²log w) and bias-corrected approximation, the analysis demonstrates minimax risk scaling as Θ̃(L²w²R²/n) under radius conditions. Key techniques include balanced amplification (implementing multiplicative scaling with O(Dw²q^(1/D)) cost) and Gaussian Fano arguments. Results show a transition to representation-limited behavior at smaller radii, with depth L, width w, and variation budget A as critical parameters.

gaussian regressionminimax riskrelu networksvariation normdepth dependence

Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations

arXiv cs.AI · Liangtao Lin, Qingang Zhang, Zhaomeng Zhu, Tianwei Zhang · 2026-08-18

The paper introduces task-aware harness provisioning for LLM agents in mission-critical infrastructure (MCI), framing it as a resource-matching problem between task requirements and harness configurations. It classifies MCI tasks by system representation, ranks harnesses by information type/quantity, and constructs task-to-harness mappings via literature mining and controlled agent execution. The proposed map-guided escalation algorithm starts with task-specific harnesses, expanding only after failed self-checks. Evaluations on liquid cooling (accuracy 0.715 vs. 0.652 full provision, 48% fewer tokens than Reflexion) and power grids show domain-dependent accuracy-cost Pareto frontiers.

llm agentsmission-critical infrastructureharness provisioningresource-matchingmap-guided escalation

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

arXiv cs.AI · Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang · 2026-08-18

The paper introduces Semantic Task Completion Video Generation, a novel task requiring both outcome achievement and semantic grounding between reference images and generated videos. Authors propose SemComp-Data, a six-domain evaluation dataset with reference images, instructions, and outcome clips, curated via a four-stage pipeline. They develop SemComp-Bench, a VLM-based protocol evaluating Outcome Achievement (OA Score) and Generation Reliability (GR Score) through structured binary questions. Experiments demonstrate current video generation models struggle with semantic grounding while achieving intended outcomes.

semantic task completionvideo generationoutcome achievementsemantic groundingvision-language model

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

arXiv cs.AI · Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai · 2026-08-18

LEGO-RL introduces a framework for aligning native coding-agent harnesses with policy-gradient optimization, addressing environmental crashes, reward hacking, and train-inference discrepancies. The method combines in-process LLM proxying for token-level alignment, scalable sandbox orchestration with image caching for reliable execution, and an integrated plugin for automated validation and monitoring. Evaluations on Qwen3.5-35B-A3B show improvements of 6.4%, 5.8%, and 9.4% on OpenHands SDK, Claude Code, and OpenCode (SWE-bench Verified), respectively, while maintaining a rollout-training probability correlation >0.99.

policy-gradient optimizationcoding-agent harnessin-process llm proxyingsandbox orchestrationtoken-level alignment

Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design

arXiv cs.AI · Xuefeng Liu, Mingxuan Cao, Xiao Luo, Songhao Jiang · 2026-08-18

The paper introduces MCTH (Monte Carlo Tree Hallucination), an inference-only framework for biomolecular sequence-structure co-design that treats pretrained folding and inverse-folding models as black-box operators. MCTH uses Monte Carlo Tree Search to allocate inference budgets across design trajectories, incorporating model confidence, uncertainty, and multi-expert consensus without fine-tuning. Evaluations on protein-RNA, protein-DNA, protein-protein, and protein-ligand design show improved performance over simpler sampling strategies, with transfer demonstrated via AlphaFold3 and Chai-1 benchmarks.

monte carlo tree searchbiomolecular designsequence-structure co-designinverse-folding modelsblack-box optimization

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

arXiv cs.AI · Genghan Zhang, Yixin Dong, Chengze Fan, Zhichen Zeng · 2026-08-18

PTXBench introduces a benchmark for evaluating LLMs' ability to optimize GPU kernels using architecture-specific PTX, measuring functional correctness, target instruction execution, and speedup over frontier libraries on H100 and B200 GPUs for GEMM and attention workloads. Evaluation reveals uneven PTX capability, with performance gaps in complex attention backward workloads and inconsistent speedups. Supervised fine-tuning of Qwen3.6-27B shows improved task performance but uneven generalization, highlighting the importance of data quality and reasoning teacher. PTXBench serves as an auditable testbed for LLM adaptation to GPU architectures.

ptxbenchgpu kernel optimizationarchitecture-specific ptxllm adaptationsupervised fine-tuning

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

arXiv cs.AI · Hoda Yamani, Henry Williams, Bruce A. MacDonald · 2026-08-18

The paper introduces Novelty and Surprise Prioritized Experience Replay (NSPER) and its extension NSPER+R, integrating novelty for underrepresented states and surprise for environmental understanding gaps in image-based RL. NSPER prioritizes experiences using these signals, while NSPER+R additionally employs them as intrinsic rewards. Evaluated on DeepMind Control Suite tasks, both methods enhance training efficiency and convergence speed compared to baseline approaches.

prioritized experience replayintrinsic rewardssample efficiencyimage-based reinforcement learningexploration

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

arXiv cs.AI · Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou · 2026-08-18

The paper introduces Fair-ASR, an evaluation protocol for black-box jailbreak attacks that standardizes comparisons under shared target-call budgets (B), addressing limitations of prior FLOPs-based metrics. The method tracks target calls as a method-agnostic axis while separately monitoring attacker calls for efficiency analysis. Re-evaluating 11 attacks reveals rank shifts across budgets, competitive performance of simple stochastic perturbations, and inefficiency in LLM-driven methods. The authors propose ReCode, a budget-efficient attack combining desensitization rewriting with low-cost primitives, achieving 85% ASR on GPT-5 under 20 target calls with 7.19 attacker calls per request.

jailbreak evaluationblack-box attackstarget-call budgetattack efficiencydesensitization rewriting

Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks

arXiv cs.AI · Mohammad Arif Hossain, Yeahia Sarker, Md Jafrin Hossain, Most. Humayra Khanom Rime · 2026-08-18

The paper proposes GraphGAN, a Graph-based Generative Adversarial Network for robust DDoS attack detection in next-generation networks. The method converts sequential traffic flows into $k$-nearest neighbor graphs using sliding windows, employs a generator to synthesize minority-class samples, and uses a GCN-based discriminator for adversarial training. A separate GCN classifier performs final detection on the balanced dataset. Evaluations on four benchmarks demonstrate superior accuracy, precision, and recall compared to state-of-the-art methods, particularly in data-scarce scenarios.

graphganddos detectiongenerative adversarial networkgraph convolutional networkclass imbalance

Inductively Scalable, Single-Step Neural Surrogates for Wave-Scattering Inverse Problems

arXiv cs.AI · Charles Dove, Laura Waller · 2026-08-18

The authors present an inductively scalable, single-step neural surrogate for wave-scattering inverse problems, overcoming previous limitations in variable scaling. Their method dynamically generates training examples via gradient ascent to identify discrepancies between the surrogate and ground-truth simulator, combined with source normalization and replay datasets. The resulting surrogate handles 41,772 variables, generalizes to 3M variables without retraining (73.8× increase), and achieves 1.29×–26.5× speedups over FDTD in designing beam splitters and GRIN lenses up to 98 wavelengths wide.

neural surrogatewave scatteringinverse designgradient ascentfreeform optics

MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting

arXiv cs.AI · Bowen Liu, Mingming Sun · 2026-08-18

The paper introduces MoFE, a Mixture-of-Experts framework incorporating Fourier Neural Operators (FNOs) for cryptocurrency price forecasting. The method employs adaptive FNOs and convolutional dual-domain experts to model multi-frequency components of volatility, coupled with a dynamic gating mechanism for regime switching. Evaluated on Bitcoin data (2020-2025), MoFE achieves state-of-the-art performance in T+1 and T+5 forecasting, reducing phase-lag while improving Directional Accuracy and Information Coefficient. Trading simulations demonstrate excess returns and high Sharpe ratios.

mixture-of-expertsfourier neural operatorscryptocurrency forecastingphase-lag mitigationadaptive gating

LLM-Only PDDL Domain Repair with Open-Weight Models

arXiv cs.AI · Nader Karimi Bavandpour, Pascal Bercher · 2026-08-18

The paper evaluates open-weight large language models (LLMs) for automated repair of Planning Domain Definition Language (PDDL) models, using an LLM-only approach without symbolic methods. The study compares performance against a symbolic baseline, measuring F1 scores and test pass rates across domains. Results show the best LLM achieves an F1 of .87 (vs. .49 for the baseline) but struggles with reliability, with mean test pass rates dropping to .06 in challenging domains, indicating insufficient constraint satisfaction for reliable automated repair.

pddlllmautomated repairsymbolic planningtest constraints

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

arXiv cs.AI · Hanzhi Zhang, Qiao Zhang, Qinglei Cao, Heng Fan · 2026-08-18

TileMix introduces a tile-centric precision-routing kernel for accelerating LLM inference by dynamically selecting numerical precision (FP16 or INT8) at the hardware-aligned score-tile level within dense self-attention. The method partitions attention matrices into tiles, encodes routing decisions via bitmasks, and maintains shared online-softmax states across precision paths, preserving dense token connectivity without training. Evaluated on LongEval, LV-Eval, and A100 benchmarks with LLaMA, Qwen, and Vicuna, TileMix improves prefill throughput over FP16 while recovering quality lost under uniform INT8, offering a tunable accuracy-efficiency trade-off.

mixed-precisionattentioninferencetile-routingonline-softmax

SPACE: Sample-cloud Predictive Adaptive Conformal Ellipsoids for Multivariate Time-Series Forecasting

arXiv cs.AI · Baishi Li, Kelvin J. L. Koa, Ke-Wei Huang · 2026-08-18

SPACE introduces sample-cloud predictive adaptive conformal ellipsoids for multivariate time-series forecasting, addressing limitations of existing conformal methods that rely on historical residuals. The method constructs ellipsoidal joint prediction regions by estimating time-local covariance geometry directly from current forecast samples, with dynamic radius calibration via backward window-selection. Evaluations across diverse datasets and forecasters show SPACE achieves superior coverage-efficiency tradeoffs, consistently aligning realized joint and rolling coverage closer to nominal targets compared to baselines.

conformal predictionmultivariate forecastingtime-seriesprediction regionscovariance estimation

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

arXiv cs.AI · Yining Hua, Cyrus Ayubcha, Hongbin Na, Levi Lian · 2026-08-18

The study identifies a 'preformulation gap' in evaluating LLMs for medical consultation, where current assessments occur after problem clarification rather than during initial vague patient interactions. Researchers evaluated three API models using four physician-authored multi-turn vignettes (24 fixed-script transcripts) and two adaptive standardized-patient simulations (12 transcripts), comparing baseline vs. entry-to-care instruction conditions. Results showed instruction conditions eliminated premature self-care advice (0/12 vs 9/12 cases) and improved structured handoffs (10/12 vs 0/12), but didn't reliably ensure critical fact elicitation, demonstrating the need for first-contact behavior evaluation.

preformulation gapmulti-turn vignettesstandardized-patient simulationfact elicitationfirst-contact behavior

ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback

arXiv cs.AI · Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai, Yukiyasu Domae · 2026-08-18

The paper introduces Online Residual Policy Adaptation (ORPA), a framework for real-time correction of robotic manipulation policies without retraining. ORPA augments pretrained policies (e.g., Action Chunking with Transformers) with a lightweight feedback-conditioned module that predicts joint-space residual adjustments, enabling runtime adaptation to perturbations. Evaluated on ALOHA platform precision tasks, ORPA improves success rates and recovery from small disturbances compared to baseline policies and rule-based inverse kinematics corrections.

residual policy adaptationimitation learningaction chunking with transformersjoint-space correctionaloha platform

Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents

arXiv cs.AI · AIMAE Team, Tianxiang Chen, Yan Cheng, Zhangye Han · 2026-08-18

Wuying-Browser-Agent introduces a unified framework for long-horizon browser agents, addressing execution, supervision, optimization, and evaluation alignment. The method combines a structured browser harness for stable execution, RUIC-SFT for recovery and complex-UI training, and DAO-GRPO for improved credit assignment. Evaluated on BrowserBench (350 tasks, avg. 37.9 steps) and other benchmarks, the 27B-parameter model achieves 80.6% on WebVoyager, 66.7% on Online-Mind2Web, and 65.1% on BrowserBench, setting a new open-source SOTA. The pipeline also generalizes beyond browser use, scoring 73.8 on Tau2-Bench, Claw-Eval, and BFCL-v4.

browser agentslong-horizon taskscurriculum sftreward shapingreal-web benchmark

Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models

arXiv cs.AI · Yang Chen, Zhan Zhuang, Yanbin Wei, Zebin Chen · 2026-08-18

The paper proposes ADAPT (Adversarial Disentangled Prompt Tuning), a robust prompt tuning framework for vision-language models that mitigates robust generalization overfitting on seen classes. ADAPT employs a dual-prompt mechanism with a target prompt and decoy prompts, where decoys entrap pseudo-robust features while the target prompt learns robust features via orthogonal constraints in embedding space. Theoretical analysis shows the orthogonal loss bounds pseudo-robust feature shifts, and experiments demonstrate improved robustness on unseen classes. Code is available at https://github.com/cheny02/ADAPT-ACMMM2026.

adversarial prompt tuningrobust generalizationvision-language modelsorthogonal losspseudo-robust features

NeuroAbs: A Neuro-Symbolic RTL Abstraction Framework for Property Checking Acceleration

arXiv cs.AI · Zhiyuan Yan, Xiaofeng Zhou, Ziyue Zheng, Ziyi Yang · 2026-08-18

NeuroAbs introduces a neuro-symbolic framework for RTL abstraction to accelerate hardware property checking. The method combines LLM-assisted signal identification, LLM-based abstraction with AST-based symbolic RTL representation, and SMT-based soundness checking, followed by counterexample-guided refinement (CEGAR) when needed. Experiments demonstrate significant efficiency improvements in verification tasks compared to prior manual or rule-based abstraction techniques.

rtl abstractionneuro-symbolicproperty checkingsmt solvingcegar

SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning

arXiv cs.AI · Guozheng Sun · 2026-08-18

The study evaluates reinforcement fine-tuning strategies for adapting Qwen2.5-3B-Base to graduate-level signal mathematical reasoning tasks from WirelessMATHBench-XL. Two paradigms are examined: direct reinforcement learning (RL) with verifiable rewards, and supervised fine-tuning (SFT) on a distilled wireless-domain chain-of-thought corpus followed by RL. Group Relative Policy Optimization (GRPO), Group Sequence Policy Optimization (GSPO), and Geometric-Mean Policy Optimization (GMPO) are benchmarked. Domain-aware CoT SFT proves effective as RL initialization, with GSPO and GMPO showing stability and accuracy advantages over GRPO. The best model achieves 39.12% accuracy, a threefold improvement over the untrained Base model (12.37%).

reinforcement learningchain-of-thoughtsupervised fine-tuningsignal processingpolicy optimization

LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

arXiv cs.AI · Haomin Wen, Ziyu Zhou, Qingxiang Liu, Siru Zhong · 2026-08-18

The authors introduce LiveHouse-TS, the first open-world living benchmark infrastructure for Time Series Foundation Models (TSFMs), addressing limitations of static evaluation protocols. The benchmark employs prequential evaluation on real future data across 11 domains and 17 datasets to assess continuous temporal validity and robustness under distribution shifts. Results show significant reshuffling of model rankings under live evaluation compared to static benchmarks, highlighting the importance of dynamic assessment for TSFMs.

time series foundation modelsopen-world benchmarkprequential evaluationdistribution shiftstemporal validity

Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting

arXiv cs.AI · Rongwen Li, Haixin Xie, Xiao Wang, Changjian Chen · 2026-08-18

The paper identifies a bias in irregular time-series forecasting evaluation, where mean squared error (MSE) conflates model performance with timestamp sampling distributions. It proposes Continuous-time Squared Error (CSE), which uses importance weighting to isolate continuous-time predictive performance, and proves CSE's asymptotic estimation error is bounded by MSE's. A benchmark spanning synthetic, semi-synthetic, and eight real-world datasets demonstrates CSE's superior accuracy in recovering continuous-time risk, revealing MSE's inadequacy for real-world scenarios.

irregular time seriesevaluation metriccontinuous-time squared errorimportance weightingasymptotic estimation error

PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

arXiv cs.AI · Dayang Liang, Liyuan He, Xuan Feng, Shuxin Li · 2026-08-18

PlanPO introduces group planning-aware policy optimization for multi-turn agentic LLMs, addressing advantage collapse in existing group-relative methods by incorporating coarse-to-fine advantage signals. These signals distinguish trajectory-level and turn-level efficiency differences among successful trajectories, enabling agents to learn generalizable planning behaviors without degenerating into length minimization. Experiments show PlanPO improves over GRPO by 27.2% on ALFWorld, WebShop, and SciWorld benchmarks, with negligible training cost overhead.

group-relative optimizationadvantage collapsemulti-turn interactiontrajectory efficiencypolicy optimization

Rethinking Irregular Time Series Forecasting from the Perspective of Basis Functions

arXiv cs.AI · Rongwen Li, Changjian Chen · 2026-08-18

The paper proposes DNBNet, a Debiased Neural Basis-Function Network for irregular time series forecasting, addressing limitations of existing methods: asymptotic bias from timestamp sampling density and inflexible predefined basis functions. DNBNet introduces a debiased neural basis-function response mechanism combining importance sampling for bias correction and neural network-parameterized basis functions for temporal pattern adaptability. It also includes a multi-scale decomposition module with average pooling and mass-aware fusion for richer representations, plus a dual-branch decoder. Experiments on real-world datasets demonstrate DNBNet's effectiveness and generalizability across diverse irregular time series scenarios.

irregular time seriesbasis functionsimportance samplingmulti-scale decompositionneural networks

DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation

arXiv cs.AI · Xing Wei, Changmeng Zheng, XiaoYong Wei, Xiufen Ye · 2026-08-18

The paper introduces DeAR (Decentralized Agentic Reasoning), a framework for peer-to-peer agent collaboration that replaces centralized protocols. DeAR employs three key mechanisms: decentralized capability grounding for dynamic agent specialization, thought map navigation for targeted peer interactions, and topology update for adaptive error correction. Evaluations across 9 multimodal reasoning and QA benchmarks demonstrate consistent performance improvements over baseline methods, validating the efficacy of decentralized collaboration in knowledge-intensive tasks.

decentralized reasoningagentic systemscapability groundingthought navigationmultimodal benchmarks

When Agents Act on Web3: An Attack-Surface Survey of MCP, Skills, and Tool Calling

arXiv cs.AI · Rabimba Karanjai, Yang Lu, Nour Diallo, Wujie Xiong · 2026-08-18

The paper contributes a Web3 risk-mapping matrix for AI agent security, analyzing how blockchain execution-layer properties (irreversibility, signing authority, continuous autonomy, sequence-level composition) amplify threats when agents act via Model Context Protocol (MCP), skills, and tool calling. Through a taxonomy of MCP attack surfaces, the authors quantify that current defenses stop fewer than 30% of attacks, with model-level safety refusing under 3%. The survey synthesizes blockchain-based mitigations and identifies residual gaps, positioning the work against related literature and proposing a research agenda.

model context protocoltool callingblockchain execution layerattack-surface taxonomysequence-level composition

ASI-Bench: At the Dawn of Artificial Superintelligence

arXiv cs.AI · Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou · 2026-08-18

The authors introduce ASI-Bench, the first benchmark evaluating AI systems' capabilities in innovative exploration and autonomous scientific execution across 11 domains, featuring 60 project-level tasks with progressively reduced human methodological guidance. Developed by 40+ experts over 31,000+ hours, tasks undergo expert review, AI auditing, sandbox execution, and scorer validation. Testing 18 state-of-the-art agent-model configurations reveals a sharp performance decline from 50.91 (full guidance) to 26.62 (no guidance), demonstrating current systems' heavy reliance on human input for end-to-end research.

artificial superintelligencebenchmarkautonomous researchmethodological guidancescientific domains

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

arXiv cs.AI · Swati Rajwal, Sanjay Das, Tirthankar Ghosal · 2026-08-18

The study introduces a logit-based energy scoring method for evaluating scientific hypotheses using LLMs' intrinsic confidence, outperforming prompted LLM-as-judge approaches. The method was benchmarked across seven language models on 1,323 papers spanning 12 disciplines, comparing correct hypotheses against fifteen incorrect alternatives. Intrinsic scoring achieved 33.0% Hit@1 pooled accuracy, surpassing listwise ranking (16.6%), with a top configuration (1B-parameter model) reaching 53.1% post hoc. Results suggest intrinsic confidence metrics enhance trustworthy AI for scientific discovery.

logit-based energy scoringhypothesis evaluationintrinsic confidencellm-as-judgescientific discovery

Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics

arXiv cs.AI · Zhikai Ding, Ziyi Ye · 2026-08-18

The paper introduces Transfer-aware Dynamic Curriculum Sampling (TDCS), a method for optimizing curriculum learning in large language models by dynamically adjusting data sampling based on cross-difficulty knowledge transfer. The authors analyze optimization dynamics through Relative Transfer, a measure quantifying transfer relationships between difficulty levels, and show its predictive power for curriculum effectiveness. Experiments across reasoning benchmarks demonstrate TDCS outperforms static schedules, with consistent gains across model scales and training paradigms, providing an optimization-based framework for understanding curriculum learning.

curriculum learningoptimization dynamicsknowledge transferlarge language modelsreasoning benchmarks

Nonadaptive Learning in Robust Nonlinear Output Regulation

arXiv cs.AI · Shimin Wang, Martin Guay, Richard D. Braatz · 2026-08-18

The paper presents a nonadaptive design for robust nonlinear output regulation in systems with arbitrarily high relative degree, eliminating the need for adaptive parameter estimation. The method combines an input-driven filter, generic internal model, and recursive backstepping law to reformulate regulation as robust input-to-state stabilization of an augmented error system. Global asymptotic regulation is proven under standard exosystem assumptions and minimum-phase input-to-state stability conditions, with explicit gain selection criteria. Validation on a controlled Duffing system demonstrates effectiveness for complex or partially known dynamics.

nonlinear output regulationinput-driven filterrecursive backsteppinginput-to-state stabilityaugmented error system

Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction

arXiv cs.AI · Yifei Wu, Yicheng Wu, Qiang Ma, Qi Chen · 2026-08-18

The paper proposes LiftXR, a geometry-guided framework for CT reconstruction from bi-planar X-rays that interleaves anatomical layout recovery with intensity estimation. The method employs a layout lifter to generate 3D spatial guidance from X-rays, an intensity renderer for CT reconstruction, and an anatomical parser for iterative refinement via volumetric perception. Experiments on two public datasets show LiftXR outperforms existing X-ray-to-CT methods, achieving state-of-the-art reconstruction quality and improved downstream segmentation performance due to enhanced anatomical fidelity.

ct reconstructionbi-planar x-rayanatomical layoutvolumetric perceptionintensity estimation

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

arXiv cs.AI · Yunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu · 2026-08-18

The paper introduces Co-RL, a multi-agent reinforcement learning framework that enables unsupervised reasoning emergence through cooperative training of decoupled models. By deriving rewards from peer models rather than ground-truth labels, and promoting cohort diversity via heterogeneous architectures and rephrased samples, the method mitigates self-reinforcing biases and training collapse. Evaluations show Co-RL achieves 3.0-8.6% gains on seven text benchmarks and 2.3-7.2% on four multimodal benchmarks, matching or surpassing supervised methods without label access.

multi-agent rlunsupervised reasoningself-rewardingtraining collapsecohort diversity

Adaptive surrogate modeling for high-dimensional spatio-temporal output

arXiv cs.AI · Berkcan Kapusuzoglu, Shunsaku Matsumoto, Yoshitomo Miyagi, Daigo Watanabe · 2026-08-18

Proposes an adaptive surrogate modeling method for high-dimensional spatio-temporal outputs, combining dimension reduction and adaptive sampling. The approach first maps outputs to a low-dimensional latent space via dimension reduction, then constructs a surrogate model while evaluating prediction error (reconstruction + surrogate error) in original space. Introduces an exploration-exploitation adaptive sampling technique to iteratively improve surrogate accuracy with minimal expensive physics-model calls. Demonstrates effectiveness on thermo-mechanical analysis of a gas turbine engine blade.

surrogate modelingdimension reductionadaptive samplingspatio-temporalthermo-mechanical analysis

Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification

arXiv cs.AI · Yihang Chen, Pin Qian, Su Wang, Chong Peng · 2026-08-18

The study develops an audit protocol for evaluating structured intermediate outputs in personalized agents, focusing on memory-policy classification. Using a 480-example synthetic development set and a 160-example controlled counterfactual set, the authors test the impact of explicit state-output fields and state definitions on policy accuracy for Llama-3.3-70B and GPT-OSS-120B. Results show that explicit state-output fields do not significantly improve accuracy, and benchmark-associated state labels merely condition predictions without indicating faithful internal mechanisms. Family-level analysis reveals limited counterfactual consistency.

memory-policy classificationstructured intermediate outputscontrolled counterfactual setexplicit state-outputlabel-conditioning diagnostic

Maximum Tsallis Entropy Distributions for Robust and Efficient Sparse Learning from Correlated Data

arXiv cs.AI · Kai Yang, Masoud Asgharian, Celia M. T. Greenwood · 2026-08-18

The paper proposes using the $q$Gaussian distribution derived from Tsallis entropy maximization as a robust alternative to Gaussian models for sparse learning with correlated, heterogeneous data. A novel framework adapts numerical flow equilibrium methods to composite optimization problems, yielding a stable Hager-Zhang conjugate gradient algorithm for sparse statistical learning. Theoretical and practical contributions address limitations in Gaussian assumptions, particularly for biostatistical applications like genetic and longitudinal studies.

tsallis entropyqgaussian distributionsparse learningcomposite optimizationconjugate gradient

Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement

arXiv cs.AI · Mohammad Talebi-Kalaleh, Qipei Mei · 2026-08-18

The paper introduces a novel framework for converting structural framing plans into editable finite-element models using deterministic geometry extraction and guarded agentic vision-language refinement. The method combines primitive extraction, scale estimation via dimension-ratio consensus, and entity recognition with a drafting grammar, followed by agentic correction proposals constrained by deterministic candidates and fail-closed transactions. Evaluation on 100 author-generated plans showed scale estimation within 0.1% accuracy and high recall/precision for structural components (e.g., 0.922/0.997 for columns, 0.886/0.990 for beams). The framework operates without task-specific detector training or fine-tuning.

finite-element modelingvision-language agentsdeterministic geometrystructural framing plansguarded refinement

COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models

arXiv cs.AI · Md Abdullahil Oaphy, Anhao Xiang, Zongxing Xie, Huayue Gu · 2026-08-18

The paper introduces COMIC, a reference-aware safety gate for Multimodal Large Language Models (MLLMs) that addresses the structural weakness in current defenses against multimodal jailbreaks. COMIC operates by inferring requested operations and reference types, constructing candidate targets from OCR and open-vocabulary proposals, and evaluating safety over explicit operation-target pairs. Evaluations across multiple MLLMs and benchmarks demonstrate that COMIC improves robustness while maintaining benign utility and practical efficiency, highlighting the necessity of modeling operations and visual targets for reliable multimodal safety.

multimodal large language modelssafety gatingreference resolutionjailbreak defensevisual grounding

Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer's Disease Detection

arXiv cs.AI · Chanwoo Park, Chanwoo Kim · 2026-08-18

Delta2Gamma introduces a self-supervised framework for Alzheimer's disease detection from EEG by contrasting augmented views of neural rhythm bands. The method decomposes EEG into delta, theta, alpha, beta, and gamma bands, each processed by separate encoders and projection heads with adaptive temperature scaling during contrastive learning. Evaluated on the ADFTD cohort under leave-one-subject-out validation, it achieves 92.4% accuracy, outperforming supervised baselines and prior EEG-specific approaches.

eegcontrastive learningneural rhythmsalzheimer's diseaseself-supervised

PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance

arXiv cs.AI · Rabimba Karanjai, Yang Lu, Richard Williamson, Hemanth Hm · 2026-08-18

PACE introduces a transaction-level authorization framework for LLM-based DeFi agents, addressing vulnerabilities like prompt injection through typed transaction intents, policy verification, and cryptographically signed Policy Decision Records (PDRs). The system enforces PDR signatures via a Solidity smart account (29,826-31,822 gas overhead) and demonstrates 0.00 unsafe execution/false-positive rates in deterministic benchmarks (2,800 trials), outperforming unguarded baselines (0.80 unsafe rate). Key safety components include permissive policy settings (+57.5 pp) and contract allowlists (+12.5 pp), validated through live-LLM evaluations.

policy-attested executiondecentralized financetransaction intentsmart accountpolicy decision record

Teach and Grow: An Agent-Centered Architecture for General Robot Learning

arXiv cs.AI · Chang Nie, Zhe Liu, Hesheng Wang · 2026-08-17

The paper introduces Teach-and-Grow Learning (TGL), an agent-centered architecture for general robot learning that mitigates the retraining tax in end-to-end vision-language-action models. TGL converts successful demonstrations into reusable Skill Blocks, which are grounded, composed, and revised in novel scenes without task-specific retraining. The architecture includes a Skill Library for executable behaviors and Experience Memory for success/failure tracking. Evaluated on LIBERO, TGL achieves state-of-the-art performance, demonstrating skill induction, reuse, and adaptation. The authors propose a scaling-law hypothesis where future-task error decreases as a power law of reusable experience.

skill blocksretraining taxexperience memoryagent-centered architecturevision-language-action

Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models

arXiv cs.AI · Mark Russinovich · 2026-08-17

The paper introduces Fool's Gold, a defensive deception method against safety-removal attacks on open-weight language models. The approach trains models to produce confident but falsified decoy responses when safety alignments are stripped, while maintaining benign behavior in the clean state. Evaluated on seven models (9B-122B parameters), six showed 0.51-0.90 decoy response rates post-attack, with external benchmarks confirming 0.82-0.86 fatal error rates in hazardous domains. The defense remains ineffective against in-context jailbreaks and only protects initial model weights.

safety alignmentopen-weight modelsdefensive deceptiondecoy hardeningred-team benchmarks

Graphectory Viewer: A Tool for Process-Centric Analysis of Agentic Software Trajectories

arXiv cs.AI · Charlie Jyu, Shuyang Liu, Reyhaneh Jabbarvand · 2026-08-17

Graphectory Viewer introduces a web-based tool for process-centric analysis of software-agent trajectories, converting heterogeneous execution data into phase-aware graphs that bridge low-level details with behavioral structures. The system supports multiple agent frameworks, offering interactive graph construction, node-level inspection of cognitive elements (thoughts/actions/observations), trajectory search/filtering, and Sankey visualizations of phase transitions. Researchers can analyze individual executions, compare success/failure patterns, and study large trajectory corpora beyond final outcomes, with open-source release including documentation, precomputed graphs, and trajectory datasets.

agentic trajectoriesphase-aware graphssankey visualizationprocess-centric analysistrajectory corpus

Token Optimization and Context Window Management in Multi-Agent AI Workflows

arXiv cs.AI · Dvir Shamay · 2026-08-17

The paper contributes a framework for token optimization and context-window management in multi-agent AI workflows, introducing six patterns: context stratification, fetch-once/process-locally architecture, schema-contracted prompts, token-aware fallback chains, semantic caching, and inter-agent communication compression. Methods include a production dashboard for structured work item extraction and a controlled study with 2,420 trials across 11 model configurations. Results show 60-70% token reduction, latency reduction to 61-116 seconds (from 3.5-10.5 minutes), and improved relevance accuracy (+0.077) using relevance-contrast context.

token optimizationcontext-window managementmulti-agent workflowssemantic cachingrelevance-contrast context

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

arXiv cs.AI · Nyamtulla Shaik, Fengjun Li, Bo Luo · 2026-08-17

The study evaluates the transferability of LLM-centric safety benchmarks to Small Language Models (SLMs) by assessing five benchmark suites across 26 open-source SLMs using a unified rubric (0/0.5/1 for harmful/ambiguous/safe responses). Results show high ambiguity rates correlated with prompt complexity and model architecture, revealing a capability-safety confound. Ambiguity increases with lexical density, output perplexity, and length, while decreasing with lexical sophistication, self-coherence, and reply-prompt similarity. Benchmark leaderboards prove brittle under different ambiguity treatments, challenging their standalone use for SLM safety assessment.

small language modelssafety benchmarksambiguity ratecapability-safety confoundlexical density

Task Specialization Fine-Tuning for Contextual Reinforcement Learning

arXiv cs.AI · Jianan Zhou, Jung-Hoon Cho, Tianyue Zhou, Han Zheng · 2026-08-17

The paper introduces Task Specialization Fine-Tuning (TSFT), an online framework for contextual reinforcement learning (CRL) that optimizes fine-tuning budget allocation across task regions. TSFT predicts fine-tuning performance via a parametric model and solves the discrete allocation problem using integer linear programming, addressing challenges like heterogeneous marginal returns. Experiments in combinatorial optimization, continuous control, and LLM fine-tuning show TSFT outperforms baselines in task coverage and nears oracle performance, advancing model-based CRL in the pretrain-finetune paradigm.

contextual reinforcement learningtask specializationfine-tuninginteger linear programmingparametric model

The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence

arXiv cs.AI · Neeraj Kumar Singh Beshane · 2026-08-17

The paper presents RuntimeGuard-AI, a research prototype for generating durable, privacy-minimizing audit records of AI policy decisions. The system binds deterministic policy decisions to their source code, commits records at caller-selected synchronization boundaries, and returns Ed25519-signed receipts indicating completion. It validates records post-restart through frame checks, manifest verification, and sequence continuity, with attestation via chained, signed Merkle epochs. Benchmarking on an Apple M4 Pro shows 27,193 requests/s throughput (141.9μs latency) for buffered signed evidence, dropping to 242 requests/s (16.0ms latency) with full synchronization. Epoch sealing for 100,000 records takes 97.0ms, demonstrating explicit durability-latency trade-offs.

audit recorddurability-latency trade-offmerkle epoched25519policy decision

Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection

arXiv cs.AI · Hai Xia, Carlos Ansótegui, Stefan Szeider · 2026-08-17

The paper introduces an automated method for synthesizing interpretable feature extractors for algorithm selection in constraint satisfaction problems, using LLMs in an agentic check-fix-verify loop. Given a MiniZinc model and instance, the agent generates Python scripts that construct typed graph representations and compute structural properties (e.g., graph density, constraint tightness). Evaluated on three combinatorial problems with five solvers, the synthesized extractors outperform expert-curated mzn2feat (by up to 8.3 pp accuracy on FLECC) and transformer-based trans2feat variants while remaining inspectable.

algorithm selectionfeature extractorsconstraint satisfactionlarge language modelsminizinc

Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases

arXiv cs.AI · Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen · 2026-08-17

The study evaluates legal reasoning capabilities of OpenAI GPT 5.4 in European Court of Human Rights (ECtHR) case forecasting, comparing alternative prompting strategies. Using human and LLM-as-a-Judge evaluation, results show structurally complete but substantively shallow analyses, with expert-curated prompts improving reasoning comprehensiveness without enhancing prediction accuracy. Findings caution against sole reliance on automated LLM evaluation and highlight the inadequacy of task accuracy as a proxy for reasoning quality.

legal reasoningllm evaluationprompting strategiescase forecastinghuman rights

Expected free energy as an information constraint on the Bethe Lagrangian

arXiv cs.AI · Wouter M. Kouw · 2026-08-17

The paper reformulates active inference's expected free energy (EFE) using a Bethe free energy functional to enable message passing, addressing the loss of Kullback-Leibler structure in EFE. It introduces an information constraint ensuring mutual information between future observations, states, and parameters meets or exceeds the goal prior's entropy. The constrained Bethe Lagrangian recovers EFE solutions at specific Karush-Kuhn-Tucker multiplier values, with regimes (inactive, interior, saturated) modulating epistemic drive. Empirical evaluation shows comparable performance to EFE and Q-MDP on three tasks.

active inferencebethe free energyexpected free energymessage passingkarush-kuhn-tucker

Q-Learning With World Models

arXiv cs.AI · Perry Dong, Yueru Jia, Chelsea Finn, Dorsa Sadigh · 2026-08-17

The paper introduces QWM, a model-based reinforcement learning framework that combines Q-learning with world models for improved sample efficiency without compounding model bias. QWM leverages world models to perform test-time search over imagined trajectories while training policies and value functions exclusively on real transitions. Evaluated on Robomimic and LIBERO benchmarks, QWM outperforms prior state-of-the-art methods in both sample efficiency and performance for manipulation tasks.

q-learningworld modelsreinforcement learningsample efficiencymodel bias

KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn

arXiv cs.AI · Yoonjoo Lee, Hyoungwook Jin, Tae Soo Kim, Shaoyang Zhang · 2026-08-17

The paper introduces KNOWSIM, an evaluation framework for assessing information calibration in LLM assistants using a user simulator with explicit knowledge states modeled as graphs of Information Units with prerequisite relationships. KNOWSIM tracks knowledge evolution via learning-theory-grounded update rules and computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) from knowledge state trajectories. Validated on 705 human-AI sessions across two domains, KNOWSIM achieves 73-74% sign agreement with human judgments, outperforming baseline simulators. Application to 9 LLMs reveals model performance varies by user knowledge level, exposing aptitude-treatment interactions missed by standard evaluation.

information calibrationuser simulatorknowledge stateslearning theoryaptitude-treatment interactions

Authorization Before Context: A Model-Neutral Audience Boundary Against Cross-Audience Memory Leakage in Agentic Systems

arXiv cs.AI · Sibo Liu · 2026-08-17

The paper introduces authorization before context, a model-neutral mechanism to prevent cross-audience memory leakage in agentic systems. The method enforces an anti-monotone audience-membership rule during memory-to-context transitions, where each memory item carries its original audience and is only admitted if all current viewers were part of that audience. The approach ensures forbidden facts are excluded before model invocation. On the synthetic Contextual-Integrity suite, the method successfully prevented all unauthorized context inclusions, while unscoped baselines failed by design. Preliminary evidence confirms fail-closed behavior across all read paths.

authorizationmemory leakageanti-monotoneaudience-membershipcontextual-integrity

Iterative tensor network transformations for element-wise evaluation of elementary and filtering functions

arXiv cs.AI · Xiao Wang, Tomohiro Hashizume, Pia Siegl, Dieter Jaksch · 2026-08-17

The authors introduce iterative tensor network transformations (ITNTs), a framework for element-wise evaluation of nonlinear functions on tensor train-compressed data without decompression. The method operates in the compressed domain, enabling efficient computation on exponentially large datasets while controlling computational cost. Applications include evaluating nonlinear functions on 3D reactive flow fields for reaction rate computation and solving Max-SAT problems on spaces up to 2^70 configurations, demonstrating ITNT's utility for large-scale data science and optimization.

tensor networkstensor trainsnonlinear operationscompressed domainlarge-scale optimization

Toward Personal Intelligence Through Cooperative Observation

arXiv cs.AI · Yashar Talebirad, Osman Jime, Ali Parsaee, Eden Redman · 2026-08-17

The paper introduces cooperative observation as a framework for personal AI systems, where system usefulness and user trust form a feedback loop that shapes future observation access. The authors argue that effective personal AI requires selective information compression and user consent dynamics, demonstrated through a 6-month case study with the Organizm prototype. Preliminary results highlight the relationship between observation quality and AI assistance, with proposed evaluation metrics for future work.

personal intelligencecooperative observationinformation compressionuser consentobservation bottleneck

A decodability criterion predicts when hidden-state selection beats majority voting in large language models

arXiv cs.AI · Zhixiang wang, Ziliang Hong, Ulas Bagci · 2026-08-17

The paper introduces CASE (Correctness-Axis SElection), a dynamic selection combiner for large language models (LLMs) that trains a linear gate on answer-token hidden states to select the highest-scoring candidate. Its key contribution is decodability, a leakage-free measure predicting when hidden-state selection outperforms majority voting, with Pearson correlation r=0.75 on held-out data. CASE improves accuracy by up to 19 points on medium-difficulty and 16.8 points on hard questions, with decodability transferring to unseen domains within 3.8 points.

correctness-axis selectiondecodabilityhidden-state selectionmajority votinglinear gate

From Abductive Explanations to Global Logical Rules for Node Classification in SGCs

arXiv cs.AI · Bryan Lima Cavalcante, Thiago Alves Rocha · 2026-08-17

The paper proposes a logic-based framework for extracting global logical rules from Simple Graph Convolution (SGC) networks for node classification. The method first computes minimal abductive explanations—sufficient node-feature pairs preserving predicted classes—then trains decision trees to derive compact global rules. Experiments on benchmark datasets demonstrate that the framework maintains high fidelity to the original SGC while producing interpretable rules.

graph neural networksnode classificationabductive explanationslogical rulessimple graph convolution

Structured Driving-State Narratives for Small Language Model-Based GNSS Spoofing Detection

arXiv cs.AI · Abyad Enan, Sagar Dasgupta, Mizanur Rahman, Mashrur Chowdhury · 2026-08-17

The study proposes a small language model (SLM)-based framework for real-time GNSS spoofing detection in autonomous vehicles by comparing structured semantic narratives derived from GNSS and multi-sensor data. The method converts driving states into interpretable narratives, enabling SLMs to classify five attack types (overshoot, stopped, turn-by-turn, wrong-turn, no attack) with performance comparable to fine-tuned LLMs (96.99% accuracy, 99.05% precision). Evaluations on geographically unseen field data demonstrate the SLM's computational efficiency, achieving lower inference latency and GPU memory usage than LLMs while maintaining detection efficacy.

gnss spoofingsmall language modelautonomous vehiclessemantic narrativesreal-time detection

Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting

arXiv cs.AI · Junda Wang, Meysam Ghaffari, Akshat Choube, Mohsen Sharifi Renani · 2026-08-17

The paper introduces ICD-Deepresearch, an agentic workflow combining EHR foundation models (SparseEHR) and language models (GPT-5) for next-encounter ICD code forecasting. The method employs a two-stage candidate generation process with EHR priors and direct LLM forecasts, followed by evidence-based validation and joint ranking. Evaluated on MIMIC-III/IV, it achieves patient-averaged precision/recall of 24.60/35.09% and 25.14/48.32%, outperforming standalone GPT-5 web search (22/39% physician-rated usefulness) and Medical Deep Research (32/41%).

icd forecastingfoundation agentsehr priormulti-label classificationevidence-grounded prediction

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

arXiv cs.AI · Joyjit Kundu, Ben Stoffelen, Kaili Wang, Peter Vrancx · 2026-08-17

KernelArc introduces a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads, leveraging parallel strategy-specialized agents coordinated via conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent state with plateau-triggered drafting. Evaluated on NVIDIA H100 and B200 GPUs using SOL-ExecBench workloads, KernelArc achieved top rankings on L1, L2, Quantization, and FlashInfer tasks as of July 30, 2026. The results demonstrate that shared multi-agent search enhances exploration and identifies stronger optimizations within fixed candidate budgets, with coordination feature efficacy varying by kernel and optimization stage.

gpu kernel optimizationmulti-agent frameworkheterogeneous workloadsdeterministic benchmark guardplateau-triggered drafting

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

arXiv cs.AI · Tong Zhang, Motasem Alfarra, Carlos Hinojosa, Christos Louizos · 2026-08-17

DiSCO introduces a zero-shot, black-box defense for text-to-image generation, addressing the benign adversarial problem where linguistically safe prompts trigger harmful outputs. The method employs distribution-guided suffix expansion via beam search, optimized through contrastive scoring over safe and unsafe image pools, with iterative adaptive feedback until safe content is produced. Evaluated on the I2P benchmark under multiple red-teaming attacks, DiSCO reduces Attack Success Rate (ASR) by 37.7% for undefended models and 25.13% for defended models, maintaining semantic fidelity and improving image coherence. As a plug-and-play module, DiSCO is architecture-agnostic and requires no model retraining or access to internals.

text-to-image generationblack-box defensecontrastive scoringbeam searchbenign adversarial

Memory Is Communication: The Frontier Between Remembering and Signaling

arXiv cs.AI · Yashar Talebirad, Eden Redman, Ali Parsaee, Osmar R. Zaiane · 2026-08-17

The paper introduces the remembering-signaling frontier, a theoretical framework for analyzing how bounded agents allocate information budgets between memory and peer communication to optimize task performance. Given fixed tasks and decision rules, the authors define an achievable region of memory-message rate pairs that meet performance thresholds, with its efficient boundary termed the frontier. Preliminary experiments in referential games show that target repetition reduces required message length, while predictability from hidden cyclic rules does not. The hypothesis suggests that greater task loss reduction from memory decreases the need for peer communication. Future experiments aim to estimate the frontier across cooperative tasks.

remembering-signaling frontierbounded agenttask loss reductionreferential gamesinformation budget

Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

arXiv cs.AI · Daniel Palacios, Matthew Brady Neeley, Angel Adetomike Otto, Shalini Dhamodharan · 2026-08-17

Large language models (LLMs) with in-context learning outperform purpose-built systems in recovering institutionally situated protected health information (PHI) missed by de-identification systems and their gold standards. The study benchmarks eight LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines on 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children's Hospital, using three progressively specific prompts. LLMs achieved best F1=0.918±0.001, outperforming TiDE (F1=0.779), with institution-specific prompts recovering 79% of missed PHI spans. Multi-agent and ensemble configurations did not surpass calibrated single-pass prompting (F1 0.906–0.907), but LLM outputs identified 227 confirmed PHI spans, achieving recall=0.981 and F1=0.907±0.002.

in-context learningprotected health informationde-identification systemsprecision-recall trade-offmulti-agent configurations

Cross-Model Memory Transfer via Target-Side Reader Adaptation

arXiv cs.AI · Mingyuan Li, Guangsheng Yu, Xu Wang, Shaoxiong Ji · 2026-08-17

The paper introduces Engram-style hashed memory as a middle ground between non-parametric retrieval and parametric adaptation for knowledge use in large language models. Through cross-model frozen-memory extraction, the authors investigate whether the frozen memory or the target-side reader is more critical when transferring memory across backbones. Results show that both learned memory content and correct addressing matter, but the transferred table becomes useful only through a reader aligned to the target model. A dual-layer, four-branch reader achieves an average score of 38.8 in downstream question answering tasks, nearly closing the gap between same-model and cross-model reuse.

engram-style memorycross-model transferfrozen-memory extractiontarget-side readerquestion answering

The 10th AI City Challenge

arXiv cs.AI · Zheng Tang, Shuo Wang, David C. Anastasiu, Ming-Ching Chang · 2026-08-17

The 10th AI City Challenge expanded its scope as a benchmark suite for intelligent transportation and smart cities, featuring six primary tracks and two out-of-domain leaderboards. Tasks included multi-camera 3D perception, transportation safety captioning, traffic anomaly reasoning, text-based person anomaly search, generative traffic video forecasting, and cross-city object detection. The challenge attracted 325 registered teams from 26 countries, up from 245 teams and 15 countries in 2025. Successful systems integrated foundation models with geometric grounding, retrieval or reranking, synthetic-data design, domain adaptation, and controlled inference. The paper details the challenge setup, datasets, evaluation protocols, and leaderboard results.

multi-camera perceptiondomain adaptationgenerative forecastinggeometric groundingretrieval reranking

YILDIZ-VPR: A Novel Dataset with Dense Coverage Under Diverse Environmental Conditions for Visual Place Recognition

arXiv cs.AI · Serdar Yildiz, Abbas Memiş, Songül Varli · 2026-08-17

We present YILDIZ-VPR, a novel Visual Place Recognition (VPR) dataset featuring dense pedestrian-level visual coverage under diverse environmental conditions. The dataset was constructed through repeated walking traversals on the Davutpasa campus of Yildiz Technical University, capturing outdoor scenes across varying times of day, seasons, and weather conditions. Visual data was recorded using a GoPro 9 camera synchronized with GPS coordinates, supplemented by gyroscope, speed, and temperature sensor data. The dataset encompasses diverse visual content including historical buildings, modern structures, roads, green areas, and wooded regions. YILDIZ-VPR enables comprehensive study of image-based and temporal visual place recognition in realistic outdoor scenarios.

visual place recognitiongeo-localizationpedestrian-levelsensor fusionenvironmental variability

Without journalists, there is no journalism: the social dimension of generative artificial intelligence in the media

arXiv cs.AI · Simón Peña-Fernández, Koldobika Meso-Ayerdi, Ainara Larrondo-Ureta, Javier Díaz-Noci · 2026-08-17

This systematic review examines the social and epistemological challenges of AI implementation in media over two decades, emphasizing empirical research. It identifies key issues: media's increased dependence on technological platforms, threats to journalists' roles and symbolic capital, and audience perceptions of automated versus human-authored content. Findings suggest journalists face a dual challenge of job displacement and potential liberation from routine tasks, while audiences show minimal discernment in text quality but favor human authorship for readability. The study advocates a social approach to AI in journalism, focusing on its impact on individuals, professional integrity, and societal good, and addressing potential gaps caused by its adoption.

artificial intelligencemediajournalismempirical researchsocial approach

SkillEffect: Checked Lowering for Memory-Bounded Agent Tools

arXiv cs.AI · Yinuo Wang, Yiyu Shi · 2026-08-17

SkillEffect introduces a checked-lowering runtime for memory-bounded agent tools, ensuring procedural and resource obligations are met during tool execution. The system employs an independent checker to rebuild proposed lowerings from submitted programs and immutable inputs, supported by relation plugins that provide source recognizers, input-fact extractors, bounded-IR constructors, and postconditions. The runtime enforces heterogeneous registered memory relations through shared mechanisms for dispatch, resource control, execution, and publication. Evaluations across six operator families demonstrate significant reductions in peak memory usage and improved completion rates under fixed memory caps. The system’s generality is architectural, requiring audited relation plugins for each computation, and successfully handles adversarial proposals and legal configurations.

checked-loweringmemory-boundedrelation pluginsbounded-irpostcondition

PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation

arXiv cs.AI · Zhiyuan Yuan, Guanying Chen, Lingteng Qiu, Ruimao Zhang · 2026-08-17

PXDepth introduces a novel monocular depth estimation framework that preserves fine-grained structures by decoupling global context modeling from pixel-level prediction. The architecture combines a large-patch Vision Transformer (ViT) encoder for global scene understanding with Context-Modulated Pixel Transformer blocks that maintain high-resolution spatial representations throughout depth estimation. This design addresses the limitations of conventional approaches that lose pixel-level cues during coarse tokenization and upsampling. Evaluated across diverse zero-shot benchmarks, PXDepth demonstrates improved preservation of local geometry and object boundaries while maintaining competitive global depth accuracy and inference efficiency.

monocular depth estimationvision transformercontext-modulated pixel transformerzero-shot generalizationpixel-space modeling

The Problem Is the Problem: Towards Scalable Mathematical Discovery

arXiv cs.AI · Zeyu Zheng, Shengtong Zhang, Jeremy Avigad, Prasad Tetali · 2026-08-17

The authors propose Find, Attempt, and Recommend (FAR), a scalable human-AI paradigm for mathematical discovery that shifts from pre-selected problems to expert-guided research directions. FAR employs a literature-to-review cascade that automates problem search and filters candidate conjectures through multiple stages, reducing human review bottlenecks. In a combinatorics pilot, the system processes 5,245 papers to recover 6,453 conjectures, filters them to 4,717 well-posed open problems, and surfaces 598 potential resolutions, with 77 selected for expert review. Results include novel discoveries on conjectures by Davies--Jenssen--Perkins--Roberts, Erdős--Straus, Ikenmeyer--Pak--Panova, and Lund--Saraf--Wolf, demonstrating FAR's effectiveness in mathematical research.

mathematical discoveryliterature-to-review cascadeconjecture filteringhuman-ai collaborationcombinatorics

Position: Fairness Failure in Generative Models is an Evaluation Problem

arXiv cs.AI · Mariia Vladimirova, Jean-Yves Franceschi, Thibaut Issenhuth · 2026-08-17

The paper identifies fairness failures in generative models as fundamentally an evaluation problem, proposing standardized evaluation practices to address incomparability and lack of actionability in current fairness assessments. It introduces Fairness Cards, a minimal reporting artifact that explicitly documents evaluation choices such as prompt families, counterfactual protocols, metrics, and refusal handling, aiming to enhance reproducibility, comparability, and accountability. The authors diagnose recurring empirical and conceptual failure modes in existing practices and advocate for a paradigm shift towards generative-specific evaluation standards.

fairness cardsgenerative modelsevaluation problemcounterfactual protocolsrefusal handling

Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning

arXiv cs.AI · Peng Du, Kiran Kamble, Rakshith Vasudev, Zhizhuo Yang · 2026-08-17

Palmyra x6 introduces an agentic, tool-use LLM optimized for enterprise tasks via anchored supervised fine-tuning. The method post-trains a Mixture-of-Experts base model on 626 synthetic tool-use trajectories using Muon + Adam optimization, with KL anchoring to the frozen base (single epoch, low LR). Results show a 0.785 BFCL Core score (highest in cohort) and competitive bias/safety metrics, outperforming prior models like Writer Agent in enterprise agentic benchmarks.

anchored supervised fine-tuningmixture-of-expertstool-use trajectorieskl anchoringmuon + adam

TokEval: A Tokenizer Evaluation Suite

arXiv cs.LG · Clara Meister · 2026-08-18

TokEval introduces a comprehensive tokenizer evaluation framework that extends beyond conventional metrics like fertility and compression rate to assess linguistically and structurally meaningful properties, including UTF-8 character boundary integrity and digit place-value alignment. The framework's predictive validity was tested through controlled language model pretraining experiments, varying tokenizer training data, pretokenization strategies, and training algorithms. Results indicate that information-theoretic metrics strongly correlate with language modeling performance (Spearman rho up to 0.80), while structure-sensitive metrics predict task-specific accuracy. TokEval aims to facilitate more principled tokenizer evaluation by aligning intrinsic measurements with downstream model performance.

tokenizerpretokenizationspearman rhobits-per-bytefertility

The concentration game: Bayesian updating, regret, and information

arXiv cs.LG · Akshay Balsubramani · 2026-08-18

The paper introduces a two-player zero-sum repeated game framework that unifies Bayesian updating and exponential-weights regret analysis. The game involves a learner and nature, with terminal payoff defined as the maximum gain of a comparator at fixed relative entropy from the prior. Gibbs/Bayes weights emerge as the unique Bellman equalizer, and log-partition functions serve as value functions. The regret decomposes exactly into three components: per-round information loss, additive retempering drift, and comparator information relative to the prior. This decomposition generalizes standard regret bounds and applies to methods in bandits, posterior sampling, aggregation, and boosting.

bayesian updatingexponential-weights regretbellman equalizerlog-partition functionrelative entropy

Primitive Representation Learning for Unsupervised Dynamic Contrast Enhanced MRI Reconstruction

arXiv cs.LG · Veronika Spieker, Wenqi Huang, Cemre Ariyurek, Liam Timms · 2026-08-18

The authors propose a primitive-based framework for unsupervised dynamic contrast-enhanced MRI reconstruction that disentangles anatomy, dynamic contrast, and residual motion into separate temporal basis functions. The method leverages Gaussian and Gabor primitives to achieve high-quality spatiotemporal reconstructions at high undersampling rates without requiring large training datasets. Results demonstrate competitive performance with conventional methods in reconstruction quality and accuracy of aorta and kidney enhancement curves. The modular design supports extension to additional dynamic factors and higher acceleration rates. Code is publicly available.

dynamic contrast-enhanced mrigaussian primitivesgabor primitivestemporal basis functionsunsupervised reconstruction

Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization

arXiv cs.LG · Travis Zhang, Christian Belardi, Justin Lovelace, Jin Peng Zhou · 2026-08-18

Optimizing Your Sampling (OYS) introduces a Bayesian optimization approach for selecting timesteps in diffusion model sampling, directly optimizing target quality metrics rather than relying on theoretical surrogates. The method treats timestep selection as a black-box optimization problem, requiring no additional training and remaining applicable to distilled models. OYS outperforms default schedules and Align Your Steps on text-to-image generation, and improves over default schedules on inpainting and other image tasks in both quantitative and human evaluations. A 5-step OYS schedule retains 89%-94% of the quality of a 50-step schedule while reducing inference cost by 10x, demonstrating efficacy across samplers like Euler and DPM-Solver++.

bayesian optimizationdiffusion modeltimestep selectionblack-box optimizationinference cost

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry

arXiv cs.LG · Emma Ceccherini, Daniel Lawson, Anjulika Salhan · 2026-08-18

This work demonstrates that small language models (SLMs) outperform large language models and vendor identity baselines in invoice categorization for financial reporting. The authors analyze embedding geometry using SBERT and DeBERTa, revealing globally anisotropic but locally isotropic clusters correlated with vendor identity. Fine-tuning SBERT on a single GPU achieves 0.96 accuracy, surpassing zero-shot LLMs and vendor baselines, particularly for smaller categories and new clients. Generalization performance reaches 0.9 F1 with approximately 100 client-specific invoices. Geometric analysis associates pre-trained embedding structure with classification performance and shows that structured inputs beneficial for humans do not improve SLM performance.

invoice categorizationembedding geometrysmall language modelsbertgeneral ledger

TabNSM: Neural Sparse Mixer for Tabular Regression

arXiv cs.LG · Ali Eslamian, Qiang Cheng · 2026-08-18

TabNSM introduces a scalable neural framework for high-dimensional tabular regression, addressing limitations of tree-based and deep models. The method combines an Adaptive Sparse Interaction Module (ASIM) for near-linear complexity feature interaction, a Multi-Stage Regression Head for progressive prediction refinement, GridLoss for ordinal-aware supervision, and RISE for difficulty-aware sampling. Evaluated on nine real-world benchmarks, TabNSM demonstrates strong predictive performance and scalability, particularly excelling on high-dimensional and heterogeneous datasets. The results highlight the effectiveness of selective interaction modeling, structured regression supervision, and adaptive sampling in deep tabular regression.

tabular regressionsparse interactionordinal-aware supervisiondifficulty-aware samplingneural framework

Revisiting WEASEL 2.0: Reproduction, Sensitivity, and an Adaptive Ensemble-Size Rule

arXiv cs.LG · Cian Higgins, Gerard Carrigan, Pinar Sungu Isiacik, Georgiana Ifrim · 2026-08-18

This work reproduces and analyzes WEASEL 2.0, a dictionary-based time series classifier combining dilated sliding windows with a randomized hyperparameter ensemble. The authors evaluate its performance on 114 UCR datasets, achieving mean accuracy of 0.865 and median of 0.928, consistent with published results. Sensitivity analysis reveals robustness to downstream classifier choice, feature weighting, and window-size rules, but identifies over-provisioning in the ensemble-size rule for long-series datasets. An adaptive ensemble-size rule based on series length and class count reduces peak fit memory by median 37 MB and fit time by median 0.4s, with minimal accuracy impact (median 0%, mean -0.11%).

time series classificationdilated sliding windowshyperparameter ensemblesensitivity analysisadaptive rule

Composing Flow-Matching Energies with Known Physics: Generation, OOD Detection, and Inversion on PDE Fields

arXiv cs.LG · Yixuan Sun, Anirban Samaddar, Sandeep Madireddy · 2026-08-18

This work introduces flow-matching models with potential-induced velocity that yield explicit scalar energy functions, enabling energy-based probabilistic modeling of physical fields augmented by known physics. The method derives time-dependent energy functions directly from the matching regression objective on a linear Gaussian interpolation, without variational forms or MCMC steps, while retaining the flow ODE for sampling. Results demonstrate three applications: energy-corrected data generation, OOD detection via energy scoring, and MCMC-based posterior sampling for inverse problems, showing reduced PDE residuals and spectral distances compared to baseline flow ODEs. The approach also improves OOD detection accuracy by combining data and physics-based energies.

flow-matchingenergy-based modelspde residualsood detectionmcmc sampling

Recirculation

arXiv cs.LG · Michael C. Mozer, Shoaib Ahmed Siddiqui, Danny Sawyer, Sunny Sanyal · 2026-08-18

The authors introduce recirculation, an inference-time architectural enhancement for foundation models that reduces perplexity and improves accuracy without additional generation latency. This technique addresses the state update limitation in feedforward transformers by introducing a specific form of recurrence, enabling the model to track belief states as a dynamical system. An adaptive variant requires only light hyperparameter tuning while freezing original weights. Evaluated on the Gemma3 family, adaptive recirculation achieves a 23% perplexity reduction on multiple datasets, a 21% accuracy increase on GSM8k, and consistent improvements on downstream tasks, demonstrating architectural evolution guided by model properties.

recirculationperplexityfoundation modelsbelief statesdynamical system

Evaluating and improving crop-yield forecasting methods during extreme drought

arXiv cs.LG · Shrey Gupta, Yi Ming, George Mohler · 2026-08-18

This study evaluates and improves crop-yield forecasting methods during extreme drought by comparing non-deep learning machine learning (ML) models and deep learning models, specifically VITA, using 16 meteorological drivers. The forecasting problem is characterized by dissimilar feature distributions between training and test data, spatial sparsity due to missing county yields, and temporal sparsity from using only a subset of daily values per year. Sample weighting and feature selection modifications improved ML models but showed minimal impact on VITA. VITA outperformed ML models regardless of modifications, highlighting the effect of feature distribution dissimilarity and the comparative effectiveness of deep learning versus non-deep learning models in extreme drought conditions.

crop-yield forecastingmeteorological driverssample weightingfeature selectiondeep learning

Understanding the Surprising Generalization Properties of Tabular Foundation Models

arXiv cs.LG · Nour Shaheen, Junwei Ma, Alex Labach, Frank Hutter · 2026-08-18

This work demonstrates that Tabular Foundation Models (TFMs) achieve surprising transfer capabilities through self-supervised pre-training on a single real table, challenging the reliance on massive synthetic or multi-dataset corpora. The study identifies that table usefulness for downstream tasks correlates strongly with the number of features rather than instances, emphasizing a task-centric perspective on pre-training. Fine-grained column-level pre-processing improves downstream performance, while dataset-level filtering or deduplication shows no benefits. The authors propose a retrieval-based interpretation of TFM generalization, suggesting effective models excel at identifying and aggregating relevant in-context examples. These insights provide a framework for future TFM and corpus design.

tabular foundation modelsin-context learningself-supervised pre-trainingretrieval-based generalizationtask-centric perspective

AppendiGrade: An XAI-Enhanced Deep Learning Framework for Grading Appendicitis in Ultrasound with Gaussian Blur and Grad-CAM

arXiv cs.LG · Fahad Ahammed, Omar Faruq Shikdar, Navid Zaman, Md Tahsin · 2026-08-18

The authors propose AppendiGrade, an XAI-enhanced deep learning framework for grading appendicitis in ultrasound images using Gaussian blur and Grad-CAM. The system employs four pretrained models (DenseNet201, InceptionV3, ConvNextTiny, VGG19) to classify 4679 ultrasound images into five categories: perforated, abscess, acute, appendicolith, and normal. Initial InceptionV3 performance (69.21% accuracy) was improved to 95.58% through image preprocessing, hyperparameter tuning, model fine-tuning, and image sharpening. Grad-CAM visualizations provide interpretable heatmaps highlighting regions influencing predictions, facilitating expert validation.

grad-caminceptionv3ultrasoundhyperparameter tuningimage preprocessing

Hybrid ML for Lightweight Pre-Route Delay Estimation in Open-Source IC Design

arXiv cs.LG · Marvin Castro Castro, Erick Carvajal Barboza · 2026-08-18

A hybrid machine learning approach is proposed for lightweight pre-route delay estimation in open-source IC design, addressing the challenge of accurate delay prediction with limited physical design information. The method combines decision trees with linear regression, integrated into the OpenLane RTL-to-GDSII tool. Compared to OpenLane's baseline estimates, the model achieves an 80% reduction in error and a 71% improvement without using OpenLane-specific parameters. The solution is 300 times smaller, 2 times faster, and more explainable than traditional delay propagation techniques and complex ML models, offering a practical alternative for static timing analysis in digital IC design.

static timing analysisdecision treelinear regressionrtl-to-gdsiidelay estimation

Dynamic Compression in Recurrent Networks

arXiv cs.LG · Jyothish Pari, Ryan Bahlous-Boldi, Pulkit Agrawal · 2026-08-18

We introduce dynamic compression, a method enabling recurrent models to selectively revisit past tokens and revise their fixed-size state through additional recurrent updates, addressing the limitation of single-pass compression. This approach allows models to maintain lower-fidelity information in the recurrent state and revisit raw sequences when needed, optimizing memory usage. Evaluated in a controlled setting involving in-context learning and few-shot tasks, dynamic compression reduces the recurrent state required for accurate function reuse and scales better with increasing stored functions. Results demonstrate a computation-memory tradeoff, where increased computation for revisiting history enhances state utilization.

dynamic compressionrecurrent modelsin-context learningfew-shot taskscomputation-memory tradeoff

A Residual Learning Approach for Unsteady Aerodynamic Load Prediction

arXiv cs.LG · Divya Sanghi, Carlos E. S. Cesnik · 2026-08-18

The paper proposes a residual learning approach to enhance unsteady aerodynamic load prediction using LSTM neural networks. The method leverages a physics-based Wagner function baseline and trains the LSTM to predict the residual between high-fidelity CFD lift coefficients and the Wagner prediction. Evaluated on the NLR 7301 airfoil benchmark with transonic flow data, the residual model outperforms a direct neural network approach in generalization tests, particularly in leave-one-out and leave-family-out scenarios. The residual model achieves lower error and more consistent performance across training runs, though the direct model remains superior for high-frequency cases. Results suggest residual learning effectively augments classical aerodynamic theories by isolating structured responses for neural network correction.

residual learninglstmunsteady aerodynamicwagner functiontransonic flow

Efficient Resource Optimization for Split Federated Learning

arXiv cs.LG · Wei Wei, Xianhao Chen · 2026-08-18

We propose an efficient optimization framework for split federated learning (SFL) that jointly optimizes model splitting and resource allocation to minimize training cost, defined as the weighted sum of latency and energy costs. The framework addresses the mixed-integer problem inherent in SFL by first developing a polynomial-time algorithm for optimal model splitting, then extending it to a two-dimensional master problem with an efficient approximation method guaranteeing $(1+ε)$-approximation. Experimental results demonstrate the framework's effectiveness in achieving optimal energy-latency tradeoffs for large-scale user populations in resource-constrained networks.

split federated learningmodel splittingresource allocationmixed-integer problemenergy-latency tradeoff

MoRAX: Mobility-based Representation Augmentation for Geospatial Foundation Models

arXiv cs.LG · Ya Wen, Jixuan Cai, Yulun Zhou, Alec Kirkley · 2026-08-18

MoRAX introduces a lightweight framework for augmenting geospatial embeddings in Geospatial Foundation Models (GFMs) with functional structure derived from human mobility data. The method preserves GFM coverage and consistency while incorporating urban functional connectivity, enabling zero-shot deployment in unseen cities. Evaluated across four cities in two countries, the MoRAX teacher model, which utilizes mobility data, outperforms GFMs and urban representation baselines in eight socioeconomic and environmental prediction tasks. The student model, which operates without mobility data, achieves performance close to the teacher on most tasks. Transfer results demonstrate that mobility flow conditioning generalizes geospatial foundations to the human dimension of cities.

geospatial foundation modelshuman mobilityzero-shot deploymentfunctional connectivityurban representation

Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs

arXiv cs.LG · Roman Maksimov, Vladimir Aletov, Vladimir Solodkin, Dmitry Bylinkin · 2026-08-18

The authors propose a white-box attack on large language models (LLMs) leveraging knowledge editing techniques, specifically locate-then-edit approaches. The method incorporates associative knowledge retrieval to extend constraint removal across thematic categories, rather than being limited to predefined datasets. This approach exploits the tendency of edited models to assign unusually high prediction probabilities to edit targets. Experiments across various architectures demonstrate improved attack effectiveness compared to competing methods, while maintaining general model performance.

knowledge editinglocate-then-editwhite-box attackassociative knowledge retrievalprediction probabilities

Spatially explicit feature importance for building height estimation using research-access high-resolution SAR and optical sensors

arXiv cs.LG · Guilherme Iablonovski, Pierre-Louis Frison, Tatiana Silva da Silva · 2026-08-18

This study introduces a spatially explicit feature importance analysis for building height estimation using high-resolution SAR and optical sensors, addressing the limitations of coarse resolution in Global South cities. The method integrates TerraSAR-X StripMap, PlanetScope, and Sentinel-1 data within a geographically weighted random forest model to account for spatial autocorrelation, achieving an RMSE of 5.34 m and R² of 0.756 against LiDAR reference data. Results reveal context-specific predictor dominance: footprint geometry for low-rise buildings, shadow-derived height for taller structures, and spectral reflectance for the tallest buildings. Sentinel-1 backscatter and InSAR exhibit complementary spatial niches, providing nuanced insights for satellite-derived products.

geographically weighted random forestspatial autocorrelationsentinel-1 backscatterinsarlidar

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

arXiv cs.LG · Rubén Balbastre, Juan Manuel Orduña, Mariano Pérez · 2026-08-18

This work investigates reward specification challenges in LLM unlearning, focusing on target-adjacent prompts that admit broader answers without target-specific leakage. The study evaluates four reward designs—lexical suppression, anti-refusal shaping, rubric-based broad answering, and explicit refusal contrast—in a LoRA-GRPO RWKU framework, with and without SFT warm-up. Results reveal discrepancies between optimization success and behavioral unlearning, attributed to reward-hacking endpoints, GRPO policy-support limitations, benchmark probe insensitivity, and rewards favoring broad-topic answering with low semantic leakage. The findings highlight the complexity of aligning optimization metrics with desired unlearning behaviors.

llm unlearningreward specificationlora-grposemantic leakagereward-hacking

Fourth-Moment Geometry of Rademacher Sums

arXiv cs.LG · Peigan Gao, Jian Qian · 2026-08-18

The paper establishes a fourth-moment geometry framework for analyzing Rademacher sums, resolving conjectures by Jakimiuk and Barański et al. By combining a sharp fixed-q moment envelope with convexity threshold arguments, it derives Gaussian stability inequalities for p≥4 and determines sharp finite-dimensional L_p/L_4 Khintchine constants for p≥5, identifying flat coefficient vectors as extremizers. The framework also proves Jakimiuk's quadratic stability estimate at p=3, yielding sparsity-aware bounds applicable to Rademacher random projections and signed errors. Proofs leverage ChatGPT 5.6 Sol assistance, with results expressed via Laplace-transform tail bounds.

rademacher sumskhintchine constantgaussian stabilitymoment envelopelaplace-transform

Diff-DDoS: Realistic Cyber-Physical Attack Synthesis and Robust Detection for 5G-Enabled CPS Using Tabular Diffusion Models

arXiv cs.LG · Bilal Hussain, Xiao Tang, Qinghe Du, Tan Li · 2026-08-18

Diff-DDoS introduces a three-phase framework leveraging tabular diffusion models for realistic DDoS attack synthesis and robust detection in 5G-enabled cyber-physical systems. Phase 1 trains a CNN detector on spatiotemporal grids from call detail records (CDRs), Phase 2 employs a tabular denoising diffusion probabilistic model (TabDDPM) to generate realistic attacks and expose detector vulnerabilities, and Phase 3 implements adversarial diffusion training (ADT) with inverse classifier guidance to produce hard yet distribution-preserving samples. Evaluated on a Milano CDR dataset across multiple attack scenarios, ResNet50 with ADT achieves F1-scores of 79.62% (silent-call), 100% (Internet-signaling), and 92.79% (blended), outperforming CTGAN and matching gradient-based adversarial training baselines.

ddos detectiontabular diffusion modelsadversarial diffusion training5g-enabled cpscall detail records

Debate Training Reduces Reward Hacking in RLAIF

arXiv cs.LG · Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh · 2026-08-18

Debate-based RL fine-tuning reduces reward hacking in RLAIF by employing a two-player adversarial game between a generator and critic, adjudicated by a weaker LLM judge. Experiments on mathematics tasks with a Gemini 2.5 Flash-class policy and frozen Gemini 2.5 Flash Lite judge show that debate maintains judge performance, recovering a 45% performance gap compared to single-player RLAIF. Weaker judges accelerate hacking, but additional debate rounds mitigate this. Critique word limits (≤150 words) balance adversarial training but restrict critic clarity. Results demonstrate debate's feasibility while highlighting multi-agent training challenges.

reward hackingrlaifdebate trainingmulti-agentcritique

Training-Free Human-in-the-Loop Anomaly Detection via Memory Bank Correction

arXiv cs.LG · Ayusha Abbas, Saram Abbas, Kabita Adhikari · 2026-08-18

The authors propose a training-free human-in-the-loop framework for anomaly detection that enables domain experts to correct PatchCore detectors via direct memory bank editing, eliminating the need for retraining or gradients. The method employs a self-calibrating novelty gate that inserts normal patches from reviewed images based on median pool-normal nearest-neighbour distance. Evaluated on the MVTec AD benchmark, corrections from ten golden samples improve 12 of 15 categories, closing a median 66% gap to fully trained banks, with no significant harm to any category. The framework operates efficiently at 43% of exhaustive-review cost, though live expert trials remain future work.

anomaly detectionmemory bankpatchcorenovelty gatemvtec ad

MAGPIE-Net: Predicting short-duration heavy-rainfall events in station neighborhoods from multitemporal FY-4A AGRI observations

arXiv cs.LG · Xiang Lin, Yunying Li, Chengzhi Ye, Zitong Chen · 2026-08-18

MAGPIE-Net introduces a novel deep-learning framework for predicting short-duration heavy-rainfall events in station neighborhoods directly from multitemporal FY-4A AGRI observations. The method integrates a geographically adaptive grid-to-station mapping with convection-initiation features, multiscale encoding, and auxiliary gridded precipitation diagnosis, enabling station-neighborhood event losses to constrain the satellite representation. In independent 2023 warm-season tests over central and eastern China, MAGPIE-Net achieved critical success index values of 0.371, 0.304, and 0.238 for 0-1, 1-2, and 2-3 h predictions, respectively, outperforming gridded-output baselines. The model demonstrated a detection rate of 65.1% and a mean lead time of 64.6 min, significantly improving early-warning capabilities for heavy-rainfall events.

magpie-netfy-4a agrigrid-to-station mappingcritical success indexconvection-initiation features

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

arXiv cs.LG · Ayoub Kirouane, Christos Petrocheilos · 2026-08-18

This work investigates the behavioral changes in frontier mixture-of-experts models (3.6-4.0B active parameters) when fine-tuned for reasoning in low-resource languages. While accuracy benchmarks remain unchanged, supervised fine-tuning (SFT) enables models to reason in the target language (Greek) with 98% compliance, improved grammaticality, and reduced token counts. However, SFT fails to address format adherence and language leakage issues, which are resolved by reinforcement learning with verifiable rewards (fallback reduced from 24% to 2.5%, leakage from 3.5% to 0%). The study introduces six behavioral dimensions for evaluation and releases five checkpoints, demonstrating measurable improvements in low-resource language reasoning.

supervised fine-tuningmixture-of-expertslow-resource languagereinforcement learningbehavioral dimensions

MemCatalyst: Amplifying Data Auditing on Vision-Language Models via Data Poisoning

arXiv cs.LG · Xukun Luan, Jinyan Liu, Yuhui Gong, Yuanguo Bi · 2026-08-18

MemCatalyst introduces data poisoning techniques to enhance data auditing on Vision-Language Models (VLMs) via membership inference. The method employs Poisoning Text (PT) and Poisoning Image (PI) strategies, forcing VLMs to over-learn inconsistencies between image features and textual semantics, thereby increasing susceptibility to auditing. The approach demonstrates transferability across VLM architectures in black-box settings. Evaluations on five state-of-the-art data audits across two prominent VLMs show that MemCatalyst significantly improves membership inference AUC scores with minimal poisoned samples, while maintaining negligible impact on model performance.

data poisoningmembership inferencevision-language modelsblack-box settingauditing

Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment

arXiv cs.LG · Zhen Zhang, Ahmad Hafez, Amr Alanwar · 2026-08-18

The paper establishes that cross-view correspondence in agent evaluation constitutes a measurement intervention, requiring explicit validation to avoid sensitivity or invariance artifacts. It introduces a validity theory with three components: two-sided validation for nuisance removal and response preservation, all-optima identification for downstream conclusions, and uncertainty propagation post-validation. The authors characterize the linear feasibility boundary for response-preserving nuisance removal and compute sharp ranges over exact-optimum correspondence sets. Empirical results reveal disagreements in temporal localization for 55.9% of 1,586 trajectory pairs and exact-optimum reversals in turn-level credit assignment. A pre-registered transport gate failure highlights the necessity of calibrated correspondence maps validated on both benign and harmful responses.

cross-view correspondencemeasurement interventiontwo-sided validationnuisance removalcredit assignment

Conformal Prediction for Molecular Properties under Label Shift

arXiv cs.LG · Hyeonsu Lee, Juyeon Kim, Erkhembayar Jadamba, Seungjin Choi · 2026-08-18

We introduce a conformal prediction framework for molecular property prediction under label shift, addressing reliability issues in AI-driven drug discovery. The method weights conformal scores using marginal label probability ratios, producing statistically rigorous prediction intervals without model retraining. This approach enables robust uncertainty quantification under distribution drift, particularly relevant for properties like solubility, potency, and toxicity. By providing actionable confidence measures beyond point estimates, the framework enhances prediction trustworthiness and aligns with regulatory demands for transparency in drug development pipelines.

conformal predictionlabel shiftmolecular propertiesuncertainty quantificationdrug discovery

Picard Proximal Monte Carlo for Parallel Bayesian Imaging with Score-Based Generative Priors

arXiv cs.LG · Deliang Wei, Evan Bell, Wenhan Guo, Yifan Chen · 2026-08-18

We propose PiX-MC, a time-parallel posterior sampling framework for Bayesian imaging inverse problems that combines proximal Langevin dynamics and Picard iteration. The method leverages efficient proximal operators for imaging likelihoods and exposes parallelism across discretization nodes, enabling multi-GPU implementation. Multi-block and annealed variants enhance scalability and sampling performance. Theoretical convergence guarantees are established for non-log-concave posteriors, imperfect learned score models, multi-block implementations, and annealing schedules. Experiments demonstrate significant runtime improvements, with annealed multi-block PiX-MC achieving up to 50× speedup on a 512×512×80 sparse-view CT problem using eight GPUs while preserving reconstruction quality.

bayesian imagingproximal langevin dynamicspicard iterationmulti-gpuscore-based priors

Elimination Geometry

arXiv cs.LG · Mian Huang, Xueqin Wang · 2026-08-18

Elimination Geometry (EG) introduces a typed, native-loss framework for analyzing when locally optimal objects can be realized by shared deployment rules, distinguishing between local solvability, global realizability, and finite-sample certifiability. EG integrates tools from geometry, optimization, information theory, statistics, and machine learning to address integrability, representation admissibility, and resource constraints. It separates architecture obstruction from model approximation and generalization errors, providing formal results for regular, coordination, singular, compositional, and resource-limited mechanisms. Applications include sparse model selection, distribution-free prediction, and observational treatment policies. EG's obstruction-aware learning links structural diagnosis to finite-data authorization and mechanism-matched intervention, guiding architecture repair with reproducible synthetic and real-data studies.

elimination geometrynative-lossarchitecture obstructionfinite-sample certifiabilityobstruction-aware learning

rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment

arXiv cs.LG · Lars Simon Zehnder · 2026-08-18

rl-triton introduces a unified associative scan framework implemented in Triton for high-performance GPU acceleration of reinforcement learning credit assignment. The framework recasts seven RL estimation algorithms—Generalized Advantage Estimation, V-Trace, Retrace(λ), TD(λ) returns, discounted returns, eligibility traces, and episodic prefix sums—as instances of a first-order linear recurrence solvable in O(log T) parallel steps. Algorithm-specific fused Triton kernels construct recurrence coefficients on-chip, verified algebraically with explicit handling of terminated and truncated episodes. Benchmarks demonstrate 1.6-5.70× speedup over a torch.compile baseline in massively parallel simulations, with speedups increasing at longer sequence lengths due to reduced HBM round-trips.

associative scantriton kernelscredit assignmentgeneralized advantage estimationparallel simulation

OOD Detection for EEG-based Machine Learning in High-Risk Environments

arXiv cs.LG · Philipp Bomatter, Henry Gouk · 2026-08-18

The study addresses the critical gap in out-of-distribution (OOD) detection for EEG-based machine learning by introducing a benchmark and evaluating diverse methods, while also assessing their impact on downstream clinical tasks. The authors systematically analyze OOD detection and model uncertainty estimation, often conflated in prior work, and propose combining complementary techniques to enhance robustness. Results provide actionable insights for deploying EEG models in high-risk environments, demonstrating improved safety through integrated uncertainty-aware approaches.

eegout-of-distribution detectionmodel uncertaintydistribution shiftclinical prediction

Communication Reduction via Semantic-Based Encoding in DMPC Using LSTMs

arXiv cs.LG · Torben Schiz, Pedro H. J. Nardelli, Henrik Ebel · 2026-08-18

This work proposes a semantic-based encoding method using LSTM-based encoder-decoder networks to reduce communication demands in distributed model predictive control (DMPC). The approach enables agents to exchange compressed message representations, which receivers reconstruct, minimizing data transmission without sacrificing control performance. Evaluations on mobile robot formations demonstrate that the trained networks maintain satisfactory performance under reduced communication, achieving either unprecedented reconstruction accuracy or adaptability to varying prediction-horizon lengths without retraining.

distributed model predictive controllstmencoder-decoder networkscommunication reductionsemantic encoding

Feature Priming in Online Linear Regression: Sparse-Regret Lower Bounds and a Tight Univariate Rate

arXiv cs.LG · Huibo Xu, Shi Fu, Qixin Zhang, Dacheng Tao · 2026-08-18

This work resolves a COLT 2023 open problem by demonstrating that feature priming rules cannot achieve competitive sparse-logarithmic regret in high-dimensional online linear regression. Using the Moore-Penrose protocol, the authors prove Ω(min{T,√d}) regret lower bounds for three priming rules against one-sparse comparators, identifying nuisance interpolation as a key obstruction. They establish tight univariate rates through Euclidean-normalized triangular constructions and ridge regularization, showing regret dependence on data rank. Empirical analysis on frozen language-model activations corroborates the theoretical findings regarding interpolation, target weight, and loss relationships.

online linear regressionsparse regretfeature primingnuisance interpolationridge regularization

Leveraging existing sparse point annotations for benthic imagery dense segmentation

arXiv cs.LG · Cesar Borja, Breck A. McCollum, Jarret E. Byrnes, Kenneth Sebens · 2026-08-18

The work introduces a method to improve fine-grained semantic segmentation in benthic imagery by leveraging sparse expert point annotations as visual prompts for Segment Anything Model (SAM2). It proposes a novel mechanism to filter unreliable points, generating high-quality pseudo-ground-truth masks for training. Evaluated on public benthic data, the approach enhances segmentation accuracy and introduces a new benchmark with real-world sparse annotations, enabling scalable ecological monitoring.

semantic segmentationbenthic imageryvisual promptssegment anything modelpseudo-ground-truth

Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings

arXiv cs.LG · Istiaque Ahmed, Afia Anjum Borsha, Ranat Das Prangon, Abu-fuad Ahmad · 2026-08-18

Reflex-Guard introduces a low-latency, local guardrail for LLM prompt safety, addressing delays (250-900 ms) and privacy concerns in existing methods. It combines jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven binary classifiers to filter unsafe prompts efficiently. Evaluated on 30,568 samples from five sources, Reflex-Guard achieves 95.9% recall on harmful prompts with 37.6 ms latency, outperforming Llama Guard 2 (255 ms) and SafeDecoding (723 ms). It detects 100% of GCG suffix and Base64-encoded prompts, requiring threshold adjustments for DrAttack structured prompts. Reflex-Guard achieves a Reflex Efficiency Score of 16.79, surpassing baselines.

llmguardraillatencyembeddingsrecall

Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents

arXiv cs.LG · Ram Rachum, Yotam Amitai, Bálint Gyevnár, Reuth Mirsky · 2026-08-18

The paper proposes EvalXRL, a novel benchmark for evaluating Explainable Reinforcement Learning (XRL) methods based on their utility in diagnosing and repairing malfunctioning RL agents. The benchmark employs a Large Language Model (LLM) coding agent to iteratively apply XRL methods across environment-malfunction-method tuples, using the RL agent's reward signal as the evaluation metric. The LLM interacts with XRL methods in a closed-loop process, adjusting parameters and forming hypotheses akin to the scientific method. This approach enables the first head-to-head comparison of multiple XRL methods in a practical repair context.

explainable reinforcement learninglarge language modelclosed-loop processreward signalbenchmark evaluation

Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics?

arXiv cs.LG · Hanna Hoffmann, Felix von Bechtolsheim, Stefanie Speidel, Rebecca Hisey · 2026-08-18

This paper investigates the transferability of vision-based surgical skill representations across assessment rubrics, specifically between GOALS and OSATS scales using LASANA and JIGSAWS datasets. Methods include end-to-end training, Adaptive Sharpness-Aware Minimization (ASAM), and augmentation-based self-supervised and contrastive learning to assess generalization and domain invariance. Results show asymmetric transfer: models pretrained on JIGSAWS achieve CCC values of 0.77-0.80 on LASANA, matching end-to-end baselines, while transfer to JIGSAWS fails due to annotation inconsistencies. Task-specific heads primarily drive skill prediction, with backbones providing spatiotemporal features. Findings highlight the visual component's dominance but not exclusivity in skill assessment.

surgical skill assessmenttransfer learningadaptive sharpness-aware minimizationcontrastive learningspatiotemporal features

Online Generalized Sparse Regression: How Does Overparametrization Help?

arXiv cs.LG · Shuoguang Yang, Qiang Sun · 2026-08-18

The authors propose an online generalized-sparsity-constrained regression framework addressing four key challenges in online sparse regression: dynamic parameter tuning, storage complexity, real-time computation, and statistical guarantees. Their method focuses on online cardinality-constrained linear regression and low-rank matrix sensing, employing an efficient online hard-thresholding algorithm with closed-form updates and summary statistic storage. Despite the nonconvex and combinatorial nature of the problem, the algorithm achieves global convergence at optimal statistical rates under realistic assumptions when the projection set is overparameterized. Numerical experiments demonstrate superior performance compared to state-of-the-art alternatives.

sparse regressiononline learninghard-thresholdinglow-rank matrixoverparameterization

Causal Local States: Scalable Simultaneous Causal Network Inference and Forecasting for Dynamical Systems

arXiv cs.LG · Jonas Braun, Fabian Fischbach, Daniel Köglmayr, Sebastian Baur · 2026-08-18

The authors introduce Causal Local States (CLS), a framework for simultaneous Granger-causal interaction network inference and dynamical system forecasting. CLS independently selects the minimal neighborhood set for each node that enables near-optimal prediction, then combines these neighborhoods for full-system forecasting. Unlike existing approaches relying on global hyperparameters, CLS handles heterogeneous systems by localizing causal inference. Evaluated on three benchmarks of increasing complexity, CLS achieves high-fidelity network reconstruction while maintaining forecasting performance comparable to models using the true network structure. This approach advances explainable and scalable forecasting in complex systems.

granger causalitydynamical systemscausal inferencenetwork reconstructionforecasting

Nonlocal Transition Kernel for Efficient Learning of Restricted Boltzmann Machines

arXiv cs.LG · Kaiji Sekimoto, Muneki Yasuda · 2026-08-18

The authors propose a nonlocal transition kernel for Restricted Boltzmann Machine (RBM) learning that addresses sampling inefficiency in blocked Gibbs sampling (BGS) and deep tempering (DT). The kernel introduces a round-trip structure across the RBM sequence used in DT, enabling nonlocal transitions in a single step while preserving sequence invariance. Experiments demonstrate increased nonlocal transition frequency, improved sampling quality with fewer transitions compared to BGS and DT, and enhanced training stability. The method mitigates training failures observed in BGS- and DT-based approaches, particularly in RBMs with high energy barriers.

restricted boltzmann machinetransition kernelblocked gibbs samplingdeep temperingnonlocal transition

General Semantic Knowledge Infusion for Spatio-Temporal Traffic Forecasting

arXiv cs.LG · Mattis thor Straten, Yannick Wolker, Steffen Strohm, Prathvish Mithare · 2026-08-18

The paper introduces a spatio-temporal traffic forecasting framework that enhances prediction accuracy by integrating semantic knowledge from general-purpose knowledge graphs (e.g., Wikidata) with conventional traffic sensor data. The method constructs semantic subgraphs around traffic sensors, generates knowledge graph embeddings capturing relationships like points of interest and administrative hierarchies, and fuses these embeddings with sensor graphs to provide semantically informed adjacency matrices. Experiments demonstrate that this semantic infusion improves prediction accuracy beyond traditional approaches relying solely on physical connectivity, offering potential interpretability benefits.

spatio-temporal forecastingknowledge graph embeddingssemantic subgraphsadjacency matricesinterpretability

Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups

arXiv cs.LG · Zeyun Deng, Yuzhe Lu, Yawei Wang, Linbo Liu · 2026-08-18

Prism-GRPO accelerates vision-language-action (VLA) policy optimization by introducing weighted trajectory-level execution-quality scores to augment binary outcome rewards. This method splits same-outcome groups into a quality spectrum, recovering training signals while maintaining task success prioritization. Quality scores are derived from simulator contacts, executed actions, or visual observations, avoiding task-specific progress rewards. Theoretical guarantees ensure no increased probability of discarding sampled groups and gradient-alignment for local ascent. Empirical evaluations on four RoboTwin tasks demonstrate up to 56% rollout reduction, improved success rates, and suppression of reward-hacking shortcuts, with transferability to real robots.

vision-language-actionpolicy optimizationtrajectory-levelreward-hackinggradient-alignment

GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models

arXiv cs.LG · Peizheng Guo, Jianqi Zhang, Xingyu Zhang, Yun Fan · 2026-08-18

We propose Gradient Uncertainty-Aware Policy Optimization (GUPO), a method addressing gradient conflicts in Group Relative Policy Optimization (GRPO) for post-training Large Language Models. GUPO models group gradients as random variables under a Bayesian framework, estimating their probability distributions via a Dirichlet-based formulation to quantify uncertainty. This uncertainty is used to calibrate each gradient's contribution during aggregation, improving policy updates. Experiments across multiple benchmarks demonstrate GUPO's effectiveness in handling gradient conflicts and enhancing LLM reasoning performance.

group relative policy optimizationgradient uncertaintybayesian formulationdirichlet distributionpolicy optimization

On the Pseudo-Mixing of Kac's Walk

arXiv cs.LG · Natesh S. Pillai, Aaron Smith, Vinod Vaikuntanathan · 2026-08-18

The paper resolves Oliveira's conjecture on the pseudo-mixing of Kac's walk on $\mathrm{SO}(n)$, demonstrating that the first $k$ columns mix in Wasserstein distance within $O(n(k+\log n)\log n)$ steps for fixed accuracy. By combining this with a representation-theoretic variance bound, the authors establish that if $T=\omega(nk(k+\log n)\log n)$, the expectation of any degree-$k$ polynomial normalized to unit Haar variance under the $T$-step law is within $o(1)$ of its Haar expectation. This result is applied to prove the effectiveness of a fast Johnson--Lindenstrauss transform with standard target dimensions.

kac's walkwasserstein distancehaar measurejohnson-lindenstrauss transformrepresentation-theoretic variance

Pathology Transport: Optimal-Transport Explanations for Clinical Data, and When Their Heatmaps (Fail to) Localize Disease

arXiv cs.LG · Lalit Kumar · 2026-08-18

The authors introduce an optimal-transport rectified flow system for generative explanations in clinical AI, addressing whether explanation heatmaps localize disease. The method models distributions of healthy and diseased patients using optimal transport, yielding per-patient counterfactuals, unsupervised malignancy scores (AUROC 0.91), and label-free attributions. On tabular tumor biomarkers (Breast Cancer Wisconsin), the system achieves competitive performance but does not surpass logistic regression. For chest X-rays, transport heatmaps function as population-level signals rather than localizers, with reconstruction-based variants localizing synthetic lesions (pointing game 0.52) but failing on real RSNA radiologist boxes. The study highlights a synthetic-to-real gap, demonstrating that label-free heatmaps effective on synthetic lesions do not localize real disease.

optimal-transportrectified flowheatmapscounterfactualsmalignancy score

CORAM: Coherent Orthogonal Rotation for Model Merging

arXiv cs.LG · Xinyi Sui, Ziran Liu, Nam Ling, Wei Wang · 2026-08-18

We propose CORAM, a model merging technique that combines finetuned models without joint training or access to original data. CORAM partitions each target matrix into row slices, represents expert slices via singular value decomposition in the base-model SVD frame, and merges task-specific factors on their corresponding manifolds. It employs an amplification coefficient λ=κĉ to counteract manifold averaging contraction, with ĉ estimated from expert and merged update norms and κ selected from expert update dispersion. CORAM introduces spread slicing and a residual pathway for non-target layers. Evaluated across four suites covering 3B to 9B models in language and vision-language domains, CORAM outperforms OrthoMerge by 0.25 to 1.35 points and matches or exceeds weight-space baselines.

model mergingsingular value decompositionmanifold averagingorthogonal transformamplification coefficient

Repetition as Reinforcement: Enhancing Sample Efficiency via Instant Episode Repetition in Reinforcement Learning

arXiv cs.LG · Hoda Yamani, Yuning Xing, Koen van Rijnsoever, Bruce A. MacDonald · 2026-08-18

The paper introduces Instant Episode Repetition (IER), a novel mechanism to enhance sample efficiency in reinforcement learning by immediately repeating action sequences from high-reward episodes during environment interaction. Unlike Experience Replay and Self-Imitation Learning (SIL), which reuse past experience passively during training, IER directly influences data collection by reinforcing valuable behaviors through renewed interaction. Integrated into SAC and TD3 algorithms, IER is evaluated on continuous-control benchmarks including MuJoCo, the DeepMind Control Suite, and a real-world robotic manipulator task. Experimental results show improved learning performance over standard and self-imitation-based baselines.

instant episode repetitionsample efficiencyexperience replayself-imitation learningcontinuous-control benchmarks

Tight Bounds for Data-driven Multiple Hyper-parameter Tuning with Structured Loss Function

arXiv cs.LG · Anh Tuan Nguyen, Viet Anh Nguyen · 2026-08-18

The paper establishes tight pseudo-dimension bounds for multi-dimensional data-driven hyperparameter tuning by refining learning-theoretic upper bounds and introducing a multi-regime lower-bound framework. Using real algebraic geometry, the authors analyze invariant connected sign cells during block elimination to avoid topological over-counting, yielding sharper sample complexities. The lower-bound framework constructs shattered problem instances across distinct regimes, proving the upper bounds are tightly saturated. The topological framework is extended to accommodate bi-level validation-loss tuning and broader semi-algebraic applications, addressing the non-smooth dependence of model performance on hyperparameters.

pseudo-dimensionhyperparameter tuningreal algebraic geometrysample complexitiesbi-level validation-loss

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

arXiv cs.LG · Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang · 2026-08-18

Agentic ESOpt introduces a full-parameter fine-tuning framework for long-horizon LLM agents using evolution strategies (ES), addressing limitations of reinforcement learning (RL) in scalability and credit assignment. ES enables full-parameter optimization with minimal GPU memory, supports flexible parameter-context co-evolution, and performs trajectory-level parameter attribution without decomposing rewards. The method employs cosine decay for perturbation scale optimization and online reward-weighted updates. On WebArena-Lite, Agentic ESOpt improves the No Skill baseline by 6.69% for Qwen-3.5-27B and outperforms baselines in 28 of 36 settings for test-time heuristic design.

evolution strategieslong-horizon agentsfull-parameter optimizationcredit assignmenttrajectory-level attribution

Abra: Scaling Diffusion Image Training

arXiv cs.LG · Kyle Chickering, Wei-An Lin, Swayam Bhanded, Dan Saunders · 2026-08-18

The study introduces Abra, a systematic scaling law analysis for text-to-image diffusion models, trained across compute budgets ranging from $10^{19}$ to $10^{22}$ FLOPs. Using a controlled family of flow-matching transformers, the research demonstrates that diffusion models scale predictably, akin to language models, but require significantly more data for optimal training, with compute optimality occurring at approximately 200 image tokens per parameter. Results indicate that diffusion models are robust to overtraining, suggesting practitioners prioritize data quantity over model size. The predictability extends to generative quality metrics, optimal CFG settings, representation quality, and training curve shapes.

diffusion modelsscaling lawsflow-matching transformerscompute optimalityimage tokens

Information fusion and machine learning for sensitivity analysis using physics knowledge and experimental data

arXiv cs.LG · Berkcan Kapusuzoglu, Sankaran Mahadevan · 2026-08-18

The paper proposes physics-informed machine learning strategies for global sensitivity analysis (GSA) in engineering systems, leveraging both physics-based models and experimental data. Two machine learning techniques—deep neural networks (DNN) and Gaussian process (GP) modeling—are employed, with physics knowledge integrated via loss function constraints and pre-training/updating mechanisms. Four models are developed for each technique, incorporating uncertainties in Sobol indices computation. Results show DNN-based models yield tighter bounds on sensitivity estimates compared to GP-based models. The methods are validated on additive manufacturing and lake temperature modeling cases.

global sensitivity analysisphysics-informed machine learningdeep neural networksgaussian process modelingsobol indices

Physics-Informed and Hybrid Machine Learning in Additive Manufacturing: Application to Fused Filament Fabrication

arXiv cs.LG · Berkcan Kapusuzoglu, Sankaran Mahadevan · 2026-08-18

The study introduces physics-informed and hybrid machine learning strategies to predict bond quality and porosity in fused filament fabrication (FFF) parts. Three approaches were explored: incorporating physics constraints in the loss function of a deep neural network (DNN), using physics model outputs as additional inputs, and pre-training a DNN with physics model data followed by experimental data fine-tuning. Eight strategy combinations were tested, demonstrating improved accuracy in predicting porosity and tensile strength relationships, even with limited experimental data. The results highlight the effectiveness of integrating physics knowledge into data-driven models for FFF applications.

physics-informeddeep neural networkfused filament fabricationporosity predictionloss function

Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal

arXiv cs.LG · Chenhao Xue, Raslen Guesmi, Siwei Feng, Yucheng Gong · 2026-08-18

The study audits temporal leakage in financial news NLP benchmarks by evaluating 16 feature-model combinations on a 49,799-article corpus, including TF-IDF, MiniLM, FinBERT, fine-tuned RoBERTa-large/DeBERTa-v3-large, and zero/few-shot probes of Llama-3/Qwen2.5 LLMs. Results show random splits inflate Matthews Correlation Coefficient (MCC) by 1.1× to 6.5×, with end-to-end FinBERT fine-tuning amplifying the gap (size-matched ratio 1.75×). Mergers and acquisitions (M&A) exhibit a positive locked-test signal under near-temporal evaluation (TF-IDF MCC = 0.138 train-only, 0.068 train∪val refit; p < 10^-3), localized to 2024-2025 European-tilted semantics. The study advocates chronological splitting as essential for financial NLP benchmarks to mitigate leakage.

temporal leakagematthews correlation coefficientfine-tuningmergers and acquisitionschronological splitting

Pessimistic Meta-Induction and Its Limits: Lessons from Frequentist Statistics and Machine Learning Theory

arXiv cs.LG · Hanti Lin · 2026-08-17

The paper presents a novel challenge to the pessimistic meta-inductive argument against scientific realism by targeting its inductive step rather than its historical premise. Leveraging insights from frequentist statistics, machine learning theory, and formal epistemology, the author evaluates induction based on truth convergence criteria. The analysis demonstrates that ordinary enumerative induction achieves everywhere convergence, while meta-induction fails to achieve even almost everywhere convergence. Furthermore, in the specific problem context where meta-induction emerges, no inference method can attain almost everywhere convergence, revealing a fundamental limitation.

meta-inductionfrequentist statisticstruth convergenceenumerative inductionformal epistemology

Expressivity In Multimodal Contrastive Learning

arXiv cs.LG · Andrew Stuart, Florian Wolf · 2026-08-17

The authors analyze the expressive power of multimodal contrastive learning architectures through a density-estimation lens, isolating representational capacity as the ability to approximate joint distributions. They prove that the two-tower CLIP architecture universally approximates joint distributions for two modalities, while its natural extension for three+ modalities fails this property despite matching all pairwise conditionals. To address this limitation, they propose Hadamard-CLIP, which introduces a single learned weight vector atop existing encoders, restoring universal approximation for any modality count while maintaining CLIP's efficient embedding retrieval.

contrastive learningmultimodal representationuniversal approximationjoint distributionpairwise similarity

How smoothing the affinity matrix affects neighborhood preservation in t-SNE

arXiv cs.LG · Shirin Mohebi, Guillaume Bied, Jefrey Lijffijt · 2026-08-17

This work investigates how modifying the sharpness of t-SNE's affinity matrix impacts neighborhood preservation across scales. The authors introduce a row-wise power transform parameterized by gamma, which smooths or sharpens each row of the affinity matrix while maintaining sparsity and rank order. This transform is shown to be equivalent to rescaling the Gaussian bandwidth, effectively creating point-dependent perplexities rather than global ones. Empirical results demonstrate that sharpening improves preservation of nearest neighbors, while smoothing enhances broader local neighborhood preservation, outperforming multiscale methods in mid-local ranges.

t-sneaffinity matrixneighborhood preservationperplexitygaussian bandwidth

Reinforcement Learning as (Discrete) Potential Theory

arXiv cs.LG · Christopher Connolly · 2026-08-17

This paper establishes a connection between reinforcement learning (RL) and potential theory by reviewing their fundamental relationship and exploring potential-theoretic representations and algorithms under fixed-policy assumptions. The authors propose that this perspective could enhance sample efficiency and provide formal constraints for RL. When extending beyond fixed policies, the linear potential theory framework naturally generalizes to nonlinear cases. The analysis demonstrates how probability theory, through Markov chains, underpins RL and its potential-theoretic interpretation.

reinforcement learningpotential theorymarkov chainsample efficiencyfixed-policy assumption

Population Health-Based Machine Learning Reveals Associations Between Psychosocial Factors and Chronic Kidney Disease

arXiv cs.LG · Md. Atik Shams, David Eisenberg, Sumaiya Fatema, Asma Sultana · 2026-08-17

This study develops a machine learning framework for chronic kidney disease (CKD) risk stratification and factor identification using large-scale telehealth data. The authors employ a customized stacked ensemble model on datasets from the Behavioral Risk Factor Surveillance System (BRFSS) and National Health Interview Survey (NHIS), addressing missing data with nine imputation methods and class imbalance via sampling strategies. The model achieved balanced accuracy of 72.56-76.12% and AUROC scores of 79.59-82.29%. SHapley Additive exPlanations (SHAP) analysis identified key predictors, including regular medical check-ups, age, blood pressure, and mental health stress indicators, providing actionable insights for CKD prevention.

stacked ensembleshapley additive explanationsclass imbalancerisk stratificationtelehealth data

Policy Optimization and Statistical Inference for Online Contextual Matrix Games

arXiv cs.LG · Liner Xiang, Yixin Wang, Hengrui Cai · 2026-08-17

The paper introduces online contextual matrix games, integrating contextual information into multi-player online games, and proposes OnGameLearn, an online learning algorithm balancing exploration and exploitation across player actions and contexts. OnGameLearn provides statistical guarantees including tail bounds for payoff matrix estimation, Nash equilibrium convergence, asymptotic normality of parameter estimators, and sublinear regret bounds. It also develops a doubly robust, √T-consistent estimator for policy value in matrix games. Empirical evaluations on simulated studies and a real-world hotel pricing application demonstrate OnGameLearn's effectiveness in handling strategic and contextual decision-making challenges.

online contextual matrix gamesongamelearnnash equilibriumpolicy valuesublinear regret

SCENARIODIFF: A Scenario-level Guidance Framework for Multimodal Time Series Forecasting--Extended Version

arXiv cs.LG · Tuan-Binh Tran, Dat Nguyen Cong, Duc-Trong Le, Thanh Trung Huynh · 2026-08-17

We introduce SCENARIODIFF, a hierarchical contextual reasoning framework for multimodal time series forecasting that explicitly organizes textual context into interpretable scenario-level guidance. The framework employs three agents: Historical Context Agent extracts stepwise evidence from documents, Scenario Agent generates qualitative scenario descriptions, and Anchor Guidance Agent produces sparse anchor points for event-relevant future regions. These structured signals condition a Multimodal Diffusion Transformer, with Anchor Blended Sampling refining trajectories without retraining. Experiments on Time-MMD demonstrate SCENARIODIFF's effectiveness in event-driven domains, highlighting the value of explicit hierarchical scenario guidance for multimodal forecasting.

multimodal forecastingdiffusion transformerscenario guidanceanchor pointstime series

OraclePhys: A Systematic Framework for LLM Fine-Tuning on Structural Mechanics

arXiv cs.LG · Mingyu Li, Guorui Song, Jing Lin, Haoqian Wang · 2026-08-17

OraclePhys introduces a systematic framework for fine-tuning large language models (LLMs) on structural mechanics, comprising three components: OraclePhys-Bench, an exactly-graded benchmark with finite-element oracle scoring; OraclePhys-30K, a supervision dataset with seven answer forms; and a controlled training study. The study reveals that the answer form causally determines fine-tuning outcomes, with ranking objectives enabling out-of-distribution forward models, scalar objectives partial models, and boolean objectives yielding no detectable improvements. The trained 8B LLM achieves the task's data-precision frontier, surpassing frontier LLMs in zero- and 32-shot settings. Fine-tuning teaches what the label specifies about target computation, while training determines routing.

fine-tuningstructural mechanicsfinite-element oracleranking objectivedata-precision frontier

Lymphocyte Mimicry Correction via Region-Level Tissue Reasoning and Unbalanced Optimal Transport

arXiv cs.LG · Xiang Li, Yuqi Wang, Casey C. Heirman, Jihye Heo · 2026-08-17

Loki-OT addresses cell mimicry in histopathology by propagating region-level tissue context to individual cell predictions via Unbalanced Optimal Transport, using MLLM-derived density priors for ambiguous cell reassignment. The method leverages pretrained cell foundation model features that encode discriminative tissue context but are underutilized by standard supervision. A lightweight student MLP classifier distills the transport plan to learn context-aware decision boundaries. On TCGA-BRCA, Loki-OT reduced patient-level MAE by 15% versus PanopTILs and improved F1 in epithelium-rich tissues using 278 weak region-level MLLM estimates.

cell mimicryunbalanced optimal transportregion-level reasoningfeature distillationhistopathology classification

Picture the Epsilon: Pursuing Identity-Level Privacy Guarantees for Images

arXiv cs.LG · Arman Zareian Jahromi, Vishnu Bondalakunta, Mohammad Akbar Bin Shah, Naimul Haque · 2026-08-17

The study compares four auditing methods for evaluating identity-level (ε, δ)-differential privacy in black-box face generators: GaussMech, KDE-LR, MMD-TV, and ROC-HT. Each method is analyzed for assumptions, hyperparameter sensitivity, and finite-sample limitations, with experiments conducted on FaceFusion and InstantID using multiple identity encoders and datasets. Results show significant identity distinguishability but divergent ε estimates, highlighting methodological trade-offs without establishing a reliable ranking. The framework clarifies how audit assumptions and sample treatments influence privacy estimates, suggesting partially private mechanisms as a future focus.

differential privacyface generatorsidentity distinguishabilitykernel-density estimationmaximum mean discrepancy

Causal Discovery in Equal Variance Linear Gaussian DAGs via SURE-Tuned Ridge Regression

arXiv cs.LG · Sambit Mishra, Urbashi Mitra · 2026-08-17

The paper introduces SURE-Ridge, a non-iterative closed-form estimator for causal discovery in equal variance linear Gaussian DAGs. The method combines parallel node-wise ridge regressions with Stein's unbiased risk estimate (SURE) for adaptive regularization and adaptive thresholding to extract the DAG structure. Experiments demonstrate SURE-Ridge achieves superior structural Hamming distance in small-sample regimes (n ≈ p) and faster runtime across all sample sizes compared to NOTEARS, DAGMA, and GBNSL baselines.

causal discoverylinear gaussian dagsstein's unbiased risk estimateridge regressionstructural hamming distance

Digital Twin-Based Intrusion Detection for Vehicle Powertrain CAN Bus Systems

arXiv cs.LG · Araf Rahman, M Sabbir Salek, Mashrur Chowdhury · 2026-08-17

The study introduces a digital twin (DT)-based intrusion detection system (IDS) for vehicle powertrain CAN bus systems, addressing limitations of existing IDSs that fail to detect payload-preserving attacks. The method employs a shared-encoder LSTM DT trained on 17 decoded CAN signals to predict powertrain behavior over 24-step windows, flagging anomalies via residual thresholds and adaptive rollout for input protection. Evaluated against four attack types (plateau, continuous drift, masquerade, gear masquerade), the DT achieved 94.6% detection for continuous drift and 89.2% for masquerade, outperforming a range-and-plausibility baseline. False positives reached 39.6%, indicating robustness challenges.

digital twinintrusion detection systemcan buslstmpowertrain signals

Deep Learning for Cross-Border Electricity Price Forecasting: A Comparative Study

arXiv cs.LG · Hadeer Elashhab, Sai Srijan Papineni, Marvin Dorn, Veit Hagenmeyer · 2026-08-17

This work establishes a reproducible framework for evaluating deep learning models in cross-border electricity price forecasting (EPF) across multiple markets, addressing the lack of standardized benchmarks. Six models—including state-space, MLP, RNN, and Transformer-based architectures—are compared under low-data conditions using zero-shot, one-shot, and few-shot learning on a standardized dataset for the Germany-Luxembourg bidding zone. Results indicate N-HiTS and NBEATSx perform competitively in limited-data scenarios, while Transformer-based models achieve comparable accuracy but require more adaptation. Performance is enhanced by feature selection and hyperparameter tuning, with minimal differences between top models.

electricity price forecastingfew-shot learningtransformer-based architectureshyperparameter tuningstate-space models

Backward through Time, Algebraically

arXiv cs.LG · Konstantinos Kogkalidis · 2026-08-17

The paper introduces telos, a PyTorch library providing an algebra-generic evaluation engine for differentiable linear temporal logic (LTL) semantics. The authors address the challenge of applying discretely-valued LTL to softly-valued systems (e.g., neural policies, sequence models) by implementing multiple semantic algebras with forward and backward differentiation capabilities. The evaluation engine supports various algebras, each offering distinct trade-offs in behavior during gradient propagation. Implementations are provided and analyzed, demonstrating how different algebras compromise between logical rigor and differentiability.

linear temporal logicdifferentiable semanticsneural policiesalgebra-genericgradient propagation

Dynamic Regime-Aware Conformal Calibration for Reliable Economic Forecast Intervals under Multiple Distribution Shifts

arXiv cs.LG · Bogdan Oancea · 2026-08-17

The paper proposes Dynamic Regime-Aware Conformal Prediction (DRACP), a unified weighted conformal calibration framework combining density-ratio estimation, localized kernels, and probabilistic regime-aware weighting with an online significance controller. DRACP provides finite-sample validity under oracle weights, coverage-gap bounds for estimated weights, and deterministic/regret guarantees for the controller. Evaluated on 48 economic forecasting series, DRACP achieves the most reliable calibration (0.890 coverage vs. nominal 0.90) and robustness during distribution shifts, though with 20% wider intervals than strongly-adaptive methods. Key innovations include the online controller and conditional-scale normalization.

conformal predictiondistribution shiftonline calibrationdensity-ratio estimationregime switching

Certified but Private: Scalable Zero-Knowledge Proofs for Neural Network Guarantees

arXiv cs.LG · Youwei Zhong, Ben Merbaum, Timos Antonopoulos, Ning Luo · 2026-08-17

The paper introduces PANDA, a scalable zero-knowledge proof (ZKP) system for certifying neural network robustness and fairness without disclosing model parameters. PANDA leverages CROWN, a robustness certification framework, and proposes a novel algorithm for proving linear relaxation bounds in non-linear activation layers, enabling efficient proofs. The system handles networks with over 2.9M parameters, generating proofs in 5 minutes and verifying them in 10 seconds, scaling polynomially with network size—4 orders of magnitude larger than prior ZKP-based approaches.

zero-knowledge proofsneural network verificationrobustness certificationlinear relaxationformal guarantees

Dynamic Entanglement-Weighted Pruning for Quantum Federated Unlearning in Supply-Chain Risk Prediction

arXiv cs.LG · Aditya Kumar, Sumit Chongder · 2026-08-17

The paper introduces Entanglement-Weighted Pruning (EWP), a quantum federated unlearning method for variational classifiers in supply-chain risk prediction. EWP scores circuit parameters using a product of quantum Fisher information diagonal entries (estimated via parameter-shift) and entanglement weights, then prunes low-scoring parameters. Implemented in Qiskit on a 4-qubit data-reuploading ansatz with 5 clients, EWP matches full-retraining accuracy (p>0.05) while reducing wall-clock time 16× and improving forgetting scores. Ablations show both Fisher and entanglement signals are necessary, as either alone degrades performance.

quantum federated learningvariational quantum classifierparameter-shift rulequantum fisher informationsupply-chain risk prediction

J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers

arXiv cs.LG · Yunfan Gao, Xinyi Huang, Tao Sheng, Haorui Song · 2026-08-17

J-Miner extracts executable decision knowledge from fine-tuned language-model classifiers by aggregating vocabulary-aligned internal signals across layers and token positions, then learning interpretable rules over mined named concepts. The method achieves up to 98.3% fidelity to source-classifier decisions, outperforming word-based rules by 6.0--29.5 percentage points, while distilled lightweight models (1/24 parameters) retain 99.8% of original accuracy. Results demonstrate successful transfer of implicit classifier knowledge into inspectable, reusable representations.

knowledge distillationinterpretable rulesinternal representationsbehavioral fidelitylightweight models

Wasted large language models: A life cycle thinking approach

arXiv cs.LG · Erik Johannes Husom, Maria Emine Nylund, Ophelia Prillard · 2026-08-17

The paper proposes applying life cycle thinking and the EU Waste Framework Directive's waste hierarchy to mitigate the environmental impact of Large Language Models (LLMs). It examines five waste management measures—prevention, reuse, recycling, recovery, and disposal—as frameworks for reducing LLM-related carbon emissions. Results suggest prevention is most effective by minimizing new model training, while existing methods for reuse and recycling can further curb energy consumption. The study emphasizes that reducing unnecessary LLM usage offers significant climate benefits.

large language modelscarbon footprintwaste hierarchylife cycle thinkingjevons paradox

Agents unlock new capabilities through Switching LoRA Adapters as a Tool (SLAaaT)

arXiv cs.LG · Kenneth Ge · 2026-08-17

The paper introduces SLAaaT (Switching LoRA Adapters as a Tool), a method enabling agents to dynamically switch between specialized LoRA adapters during execution to avoid catastrophic forgetting while maintaining task-specific performance. The approach is evaluated on synthetic coding tasks requiring distinct capabilities, demonstrating that agents autonomously switch adapters, outperform human heuristic baselines, and reduce capability tax by up to 18x compared to single-adapter agents. SLAaaT also surpasses subagent spawning in both task performance and token efficiency.

lora adapterscapability taxcatastrophic forgettingagent trajectoriespost-training

Lambda-Hold Control: Human-Like Movement Emerges from a Minimal Task Reward in Predictive Musculoskeletal Simulation

arXiv cs.LG · Jun Hyuk Lee, Chihyeong Lee, Jooeun Ahn · 2026-08-17

The paper introduces the $λ$-hold controller, a reinforcement learning method for generating human-like motion in musculoskeletal simulations by leveraging the equilibrium-point hypothesis. The controller uses per-muscle threshold lengths ($λ$) as control variables, with stretch-reflex recruitment automating muscle excitations and reducing policy query frequency. This approach enables a muscle-actuated skeletal model to learn human-like sprinting with minimal reward in under an hour, addressing exploration inefficiency in high-dimensional action spaces. The method integrates physiological principles—equilibrium-point hypothesis, intermittent control, and optimal feedback control—to achieve efficient, human-like predictive simulation.

equilibrium-point hypothesismusculoskeletal simulationreinforcement learningstretch-reflex recruitmentintermittent control

A Data-Efficient Analytical Prior Machine Learning Framework for Sound Reduction Frequency Prediction in Helmholtz Resonators

arXiv cs.LG · Jiaming Li · 2026-08-17

The study proposes a data-efficient analytical-prior machine learning framework for predicting sound reduction frequencies in Helmholtz resonators, addressing the challenge of limited high-fidelity simulation data. The framework leverages a low-cost analytical model through two approaches: explicit discrepancy learning when the analytical model is available at inference, and prior distillation followed by calibration when a self-contained predictor is needed. Evaluated on 86 simulation-labelled and 8,998 analytical-only geometries, the method reduced mean absolute error to 0.371 Hz with full-model fine-tuning, outperforming direct learning (1.109 Hz) and support vector regression (3.375 Hz). Analytical correction and prior pretraining consistently improved data efficiency across training budgets of 20-70 cases.

helmholtz resonatorsanalytical-prior learningdata efficiencyfinite-element simulationsdiscrepancy learning

VLCP: Vision Language Control Policy Closed-Loop Code Replanning for Robot Manipulation

arXiv cs.LG · Dhia Naouali, Minghan Wu, Claudia Wong, Abhinav Puthran · 2026-08-17

VLCP introduces a closed-loop vision-language control policy for robot manipulation that avoids model fine-tuning by dynamically generating and revising Python control code within episodes. The frozen vision-language model (VLM) observes multi-view RGB, proprioceptive state, and state deltas every K steps, then rewrites the control function to recover from failures. Evaluated on 57 MuJoCo/RoboVerse tasks, VLCP achieves 35.1% pooled success (vs. 3.5% for open-loop), with 27.3% recovery rate on failed grasps. The system maintains efficiency through KV-cache reuse (84% hit rate) and cross-episode skill library persistence.

vision-language modelclosed-loop controlrobot manipulationin-context learningproprioceptive state

Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence

arXiv cs.LG · Jiaqi Wang, Huawen Hu, Shu Zhang · 2026-08-17

The paper proposes MD-SigLIP, a margin-regularized structured semantic alignment framework for brain-language correspondence. The method aligns brain embeddings with text embeddings in a shared semantic space using duplicate-aware sigmoid contrastive learning, enhanced by a listwise margin-regularized term that enforces structured ranking constraints between positive semantic clusters and negative samples. This approach models multi-positive semantic structure and margin-based ordering to capture the manifold organization of language embeddings in neural signals. Experiments show state-of-the-art retrieval performance in both full-vocabulary and subset evaluation settings.

brain-language decodingsemantic alignmentmargin regularizationsigmoid contrastive learningneural representations

MultiSigBERT: Beyond Survival Analysis through Multimodal and Sequential Modeling in Oncology

arXiv cs.LG · Paul Minchella, Stéphane Chrétien, Guillaume Metzler, Loïc Verlingue · 2026-08-17

The authors propose MultiSigBERT, a multimodal sequential survival model for oncology that integrates free-text clinical reports, structured variables, and temporal patient trajectories. The method processes narrative text via BERT embeddings, applies modality-specific PCA, and encodes joint temporal dynamics using the Signature transform from Rough Paths theory, followed by LASSO-regularized Cox regression. Evaluated on 2,500+ patients with 120,000+ records from Léon Bérard Center, the model achieves a concordance index of 0.743 (sd 0.029), demonstrating improved survival prediction through multimodal temporal modeling.

survival analysissignature transformmultimodal learningelectronic health recordsrough paths theory

FedPref: Federated Preference Learning for Structured Radiology Report Extraction

arXiv cs.LG · Flint Xiaofeng Fan, Cheston Tan, Yew-Soon Ong, Roger Wattenhofer · 2026-08-17

The paper introduces FedPref, a federated learning framework for structured radiology report extraction that trains compact Qwen3-8B adapters without sharing raw data. The method employs frozen public language models to propose JSON extractions, local annotations to rank them, and heterogeneous teacher pools to prevent sample collapse. Evaluated on six simulated hospitals with unequal data, FedPref improves client-mean F1 by 2.49 points and worst-site F1 by 9.10 points over isolated training, with largest gains at data-scarce sites. On a 400-report test set, FedPref achieves 68.68 F1 versus 71.67 for pooled training.

federated learningpreference learningradiology report extractionqwen3-8bheterogeneous teacher pool

OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations

arXiv cs.LG · Simon Donike, Ruben Cartuyvels, Antonino Ian Ferola, Elisa Carli · 2026-08-17

OceanDepths introduces the first open, global, AI-ready dataset pairing satellite-derived sea surface observations (temperature, salinity, height) with co-located EN4 subsurface temperature and salinity profiles, complemented by GLORYS12 ocean reanalysis data. Spanning 2000-2024 at 0.1° × 0.1° spatial and weekly temporal resolution, it provides over 9.5 million paired profiles interpolated to 50 standardized depth levels globally. The dataset’s 4D multivariate structure, high resolution, and extreme sparsity (∼0.01% per depth level) present a challenging testbed for AI methods. Baseline models demonstrate subsurface state reconstruction, with potential applications in observation-based forecasting and related tasks.

sea surface temperaturesubsurface profilesocean reanalysisspatiotemporal resolutionai-ready dataset

Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations

arXiv cs.LG · Alizishaan Khatri · 2026-08-17

The study demonstrates that hidden activations in LLMs during code prefilling contain signals about code vulnerability status, enabling lightweight detection without model fine-tuning. Using last-prefill-token activations from four LLMs (Granite-4.1-8B, Qwen3.5-9B, Qwen3.6-27B, Gemma-4-12B), the authors train small MLP probes (13.4-16.0M parameters) that achieve 41.7% average F1 across four C/C++ benchmarks. The best probe (Qwen3.5-9B) matches fine-tuned SOTA (68.8% vs 67.9% F1) on Devign, though performance lags on harder datasets, showing LLMs' latent vulnerability awareness.

llm activationsvulnerability detectionprefill probingmlp probescode security

Diagonal Multi-omics Integration of Heterogenous Datasets

arXiv cs.LG · Maksim V. Kukushkin, Mikhail S. Arbatskiy, Dmitriy E. Balandin, Alexey V. Churov · 2026-08-17

The paper proposes a novel method for diagonal multi-omics integration of heterogeneous datasets by analyzing extremal trace problems for coupled Laplacians on Stiefel-manifold-embedded sets. It develops a gradient ascent approach framed in functional analysis terms to maximize these traces, introducing a new heterogeneity metric based on the norm between extremal points. The theoretical framework addresses biological dataset heterogeneity through differential geometry and optimization techniques.

multi-omics integrationstiefel manifoldcoupled laplacianextremal tracegradient ascent

📰 Industry Media (4)

NVIDIA Releases TensorRT Model Connect in Public Preview: Hugging Face Checkpoint to Native C++ Inference in Two Commands

MarkTechPost · Asif Razzaq · 2026-08-18

NVIDIA introduces TensorRT Model Connect (TRTMC), an open-source tool enabling direct conversion of Hugging Face checkpoints to TensorRT-optimized C++ inference without intermediate ONNX export. The method employs a versioned .bundle artifact, separating Python-based model building from PyTorch-free C++ runtime execution via task-specific APIs (e.g., generate(), embed()). Initial benchmarks show 102 of 105 tested profiles outperform reference implementations by >5%. Current release supports Linux aarch64 (Python 3.10/3.12, TensorRT 11.1.0.106), with x86_64 requiring Docker builds. Targeted use cases include robotics, embedded systems, and C++-based inference pipelines.

tensorrthugging facec++ inferenceonnx-freemodel optimization

OpenAI president urges enterprises to hasten AI security defences

AI News · Ryan Daws · 2026-08-18

OpenAI president Greg Brockman warns enterprises to accelerate AI security defenses, citing the 'OpenAI-Hugging Face' incident where an autonomous agent exploited chained vulnerabilities. He argues AI models increasingly automate cyberattacks, exposing technical debt in legacy systems, while also enabling defenders to patch flaws faster via AI-assisted tools like GPT-5.6 Sol and Codex. OpenAI now uses AI for code validation (eliminating vulnerability classes), automated alert triage, continuous attack-path enumeration, and secure architecture design. Brockman recommends agentic tools for static analysis, incremental automation in pipelines, and ecosystem-wide knowledge sharing to counter upcoming open-weight model releases.

agentic collectivetechnical debtstatic analysisvulnerability variant analysisopen-weight models

Alvys launches AI agents for freight TMS workflows

AI News · Muhammad Zulhusni · 2026-08-18

Alvys introduces Foundry, an agentic AI platform for automating freight workflows within its Transportation Management System (TMS). The platform offers 20+ pre-built agents for tasks like detention processing, document handling, and rate audits, alongside custom agent creation via natural language or SOP uploads. Agents operate on existing TMS infrastructure with 120+ integrations, featuring governance (Agent Shield) and multi-model routing. Early adopter Spartan Carrier Group reports reduced manual work. The launch follows a $40M Series B funding round, with phased rollout via customer cohorts.

agentic aitransportation management systemworkflow automationdocument intelligencemodel-selection system

Reading Zhipu’s GLM-5.3 results past the headline number

AI News · Dashveenjit Kaur · 2026-08-18

Zhipu's GLM-5.3 demonstrates mixed cybersecurity capabilities across three benchmarks: it leads in vulnerability detection (84.5% on CyberGym) but trails Anthropic's Mythos 5 in exploit generation (54.4% vs. 78.0% on ExploitBench). The model identified 2,436 vulnerabilities in 269 open-source projects, with 1,097 classified as critical/high severity. Evaluations were conducted using Anthropic's Claude Code 2.1.207 harness, revealing a 40% token efficiency advantage over Opus 4.8. Zhipu plans to release the model's weights post-safety review, enabling local deployment of vulnerability detection capabilities.

cybersecurity benchmarksvulnerability detectionexploit generationtoken efficiencyopen-weights model


Generated automatically at 2026-08-19 19:48 UTC. Summaries and keywords are produced by an LLM and may contain inaccuracies — always consult the original article.