AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

42 Papers today
8h Update frequency
7 Days of history
Learning Provable Neural Network Observer for Uncertain Dynamical Systems
Zhangyi Wang, Jiaxu Liu, Chen Song, Chao Xu, Shengze Cai
Theory Robotics Optimization
  • Introduction of a two-stage training framework for neural network observers that separates learning from stability certification.
  • Point-guided Lyapunov pre-training enhances estimation accuracy and local stability before LMI fine-tuning.
  • The framework provides theoretical guarantees for local stability and probabilistic coverage.
  • Empirical results show significant improvements in training speed and tracking accuracy compared to existing methods.
Read more
When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization
Gongyue Zhang, Honghai Liu
Optimization Theory
  • The optimal preconditioning exponent decreases linearly with log10 of the learning rate.
  • Negative exponents are not universally better; they are effective in high-step-size regimes.
  • Source-validation and cross-environment optimal exponents differ, indicating a model-selection conflict.
  • Lower preconditioning exponents reduce reliance on spurious and noise features.
Read more
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng, Weiqiang Zhu, Zhenhai Ji, Zhengning Wang
Reinforcement Learning Graph Learning
  • HySTAR introduces a stable credit assignment framework using anchored hypergraphs.
  • The method separates adaptive representation learning from value decomposition.
  • Extensive experiments show HySTAR outperforms MAPPO and other baselines in various MARL scenarios.
  • The framework effectively addresses structural target drift in cooperative multi-agent settings.
Read more
Neural Ideals and Neural Codes: An Algebraic Framework for Neural Network Classification and Feature Interpretation
Venkata Subbaiah Yerrapati, Rahul Dixit, Ajay Kumar Shukla
Theory Interpretability
  • Introduction of an allowed code space associated with target classes in neural networks.
  • Adaptation of neural ideals theory for artificial neural network classification.
  • Establishment of a classification–ideal correspondence for class membership characterization.
  • Development of algorithms for constructing neural codes and deriving neural ideals.
Read more
PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control
Yuhe Sui, Yingzhi Tang, Shufang Chen
Reinforcement Learning Robotics Optimization
  • Introduces PolicyAttention, a method that integrates softmax attention with Policy Mirror Descent for RL.
  • Addresses the closed-loop gap in reinforcement learning where model outputs affect future inputs.
  • Demonstrates that the proposed method retains performance under repeated control and environment shifts.
  • Achieves median normalized returned-policy losses significantly lower than existing methods.
Read more
ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers
Mohd Moin Khan, Naman Srivastava, Pandarasamy Arjunan
Efficient ML Computer Vision
  • ENAS is a CPU-only NAS framework that enhances accessibility for TinyML model development.
  • It features a flexible cell-based search space and a three-stage hybrid search strategy.
  • ENAS achieves significant search-time speedups of 2.41× and 1.70× on two benchmark datasets.
  • The framework maintains competitive accuracy while reducing peak activation RAM usage.
Read more
Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design
Fu Feng, Ruixiao Shi, Yucheng Xie, Jing Wang, Xin Geng
Robotics Optimization Reinforcement Learning
  • Introduction of Morphogene as a high-level latent blueprint for body-brain coordination.
  • GeCode reformulates morphology-control co-design as gene-driven exploration in a compact latent space.
  • Demonstrated significant performance improvements over existing methods, achieving 2.5× faster convergence.
  • Morphogene allows for coherent changes in morphology and control through adaptive conditioning.
Read more
Learning coarse-step dynamics and internal mechanical response with graph networks
Vinay Sharma, Olga Fink
Graph Learning Robotics Time Series
  • Introduces Newmark-β-DGN, a graph neural network framework for inferring mechanical responses from kinematic data.
  • Combines semi-implicit updates and operator-weighted virtual hubs to enhance prediction accuracy over coarse time scales.
  • Achieves long-horizon predictions in various physical systems without direct supervision of mechanical quantities.
  • Demonstrates the ability to infer forces and recover stiffness structures from observed trajectories.
Read more
Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
Alireza Abdollahpoorrostam, Ehsan Sharifian, Buse Åžen, Marco Cuturi, Daniel Kuhn
Optimization Theory Robotics
  • Introduces a penalized DRO formulation that incorporates Wasserstein penalties for adversarial distribution shifts.
  • Demonstrates that optimal transport maps for adversarial training are cyclically monotone, which is crucial for effective learning.
  • Presents Multi-start Particle Ascent (MPA) and ICNN-based methods to enforce cyclical monotonicity in adversarial training.
  • Empirical results show significant improvements in robustness and generalization compared to standard adversarial training methods.
Read more
Evaluating the accuracy of KV cache reuse techniques
Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona
NLP Large Language Models Efficient ML
  • Current evaluations of KV cache reuse techniques often inflate accuracy metrics.
  • Existing datasets lack the necessary dynamics to evaluate KV cache reuse effectively.
  • The proposed evaluation methodology isolates accuracy loss attributable to KV cache reuse.
  • Boxoffice generates datasets that ensure meaningful evaluation of KV cache strategies.
Read more
Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System
Hongsen Zhang, Lu Zhang, Mingjing Xu, Yi Zhang, Gregory Epiphaniou, Carsten Maple
Reinforcement Learning Theory Interpretability
  • Introduction of a PR-driven objective for generating UAPs against DRL-based IDS.
  • Development of PX-UAP, which utilizes XAI for feature attribution to enhance perturbation strategies.
  • Theoretical justification for PX-UAP with a focus on entropy-regularized perturbation allocation.
  • Extensive experimental validation showing superior performance of PX-UAP over state-of-the-art methods.
Read more
Online Learning via Learned Latent Bayesian Tracking
Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone
Time Series Efficient ML Optimization
  • AURA framework enables rapid online learning through a learned low-dimensional latent representation.
  • Utilizes extended Kalman filtering for efficient single-step updates in the latent space.
  • Meta-learns the latent dynamics and lifting map from offline data for improved adaptation performance.
  • Demonstrates significant improvements in adaptation speed and accuracy in non-stationary environments.
Read more
Block Sparse Attention with Log-Linear Complexity
Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu
NLP Large Language Models Efficient ML
  • PISA reduces the computational complexity of block-sparse attention from quadratic to log-linear.
  • The method employs a hierarchical Top-K selection strategy to efficiently narrow down key candidates.
  • Hardware-aware Triton kernels are developed for optimized training and inference.
  • PISA shows competitive performance on commonsense reasoning benchmarks and superior results on retrieval tasks.
Read more
Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model
Kimon Antonios Provatas, Ilias Georgakopoulos-Soares
Efficient ML NLP Theory
  • Conducted a reliability audit of a commercial System-1 model on biosecurity-relevant benchmarks.
  • Demonstrated significant sensitivity of model predictions to the order of answer options.
  • Introduced selective permutation averaging to improve accuracy on low-confidence items.
  • Established that the model is reasonably well-calibrated but accuracy varies significantly by task.
Read more
Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection
Jianan Liu, Chunguang Li
Time Series
  • Introduces a Bayesian Tensor Autoencoder that preserves the tensor structure of multi-dimensional time series.
  • Combines reconstruction-based and prediction-based approaches through a Physics-informed Predictive Prior.
  • Utilizes Bayesian fusion to enhance modeling capabilities for normal data.
  • Incorporates physical laws to prevent over-generalization in anomaly detection.
Read more
Differentiable RNA Secondary Structure Extraction for Deep Learning
Tyler Illman, Max Ward, Marcell Szikszai, Ryan K. Krueger
Theory Optimization
  • Introduces a novel SDSM normalization algorithm for direct output of base-pairing probability matrices.
  • Analyzes the effect of training-extraction congruence on RNA secondary structure prediction.
  • Demonstrates that the choice of extraction method significantly impacts prediction performance.
  • Shows that the SDSM model outperforms traditional methods and baselines in structure prediction accuracy.
Read more
Benchmarking the Connectomes of Caenorhabditis elegans within the Reservoir Computing Framework
Felix S. Reimers, Ola Huse Ramstad, Aliaksandr Hubin, Stefano Nichele
Graph Learning Theory
  • C. elegans connectomes were benchmarked using the reservoir computing framework.
  • Randomized null models often outperformed biological connectomes in computational tasks.
  • Performance varied significantly based on reservoir configuration and connectome derivation methods.
  • Connectomes from different ages produced varying results without clear trends.
Read more
Towards Understanding Momentum Acceleration in River-Valley Loss Landscape
Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen, Tianhao Wang, Zhiyuan Li
Optimization Theory
  • Momentum stabilizes large learning rates, enhancing optimization speed along the river in the loss landscape.
  • Theoretical analysis shows that momentum can increase the maximum tolerable learning rate significantly.
  • In flat river scenarios, the acceleration is mainly due to the larger learning rate rather than momentum.
  • Empirical experiments validate the theoretical predictions regarding momentum and learning rate interactions.
Read more
Guarded Gradient-Based Activation Steering of Shutdown Responses in Qwen3.5-0.8B: A Minimum-Step Policy
Farhad Davaripour
Large Language Models Theory Interpretability
  • Introduces a guarded gradient-based activation steering method for managing shutdown responses in AI models.
  • Focuses on the critical AI-safety issue of ensuring models accept shutdown commands instead of avoiding them.
  • Utilizes a classifier to selectively apply interventions based on context, enhancing the precision of the steering process.
  • Achieves 75% recall and 90% precision in detecting shutdown-related contexts, with minimal impact on non-shutdown behaviors.
Read more
More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting
Lewei Xie, Haoyu Zhang, Jiajun Zhou, Yulong Chen, Guanxing Chen, Yu-An Huang, Hau-San Wong, Yifan Zhang, Zhi-An Huang
Time Series Graph Learning
  • Introduces STFO, a novel framework for continual spatio-temporal forecasting that separates spatial dynamics from sensor layouts.
  • Utilizes a fixed latent grid for sensor-independent representation, allowing for the reuse of learned knowledge across different sensor configurations.
  • Implements a drift-adaptive mechanism to adjust to changes in spatial dynamics using spectral descriptors and attention mechanisms.
  • Demonstrates significant improvements in forecasting accuracy on multiple datasets, outperforming existing graph-based methods.
Read more
Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra
Matthieu Le Lain, Gaël Cessateur, Sébastien Lefèvre
Computer Vision Time Series Theory
  • Introduces a self-supervised learning method for analyzing auroral emission spectra.
  • Achieves significant improvements in classification accuracy over previous supervised models.
  • Demonstrates the limitations of transferring existing astronomical spectral models to auroral spectra.
  • Highlights the importance of in-domain pretraining for effective representation learning.
Read more
OpenHail: An Event-Driven Gymnasium Environment for Electric Ride-Hailing Fleet Control
Tommaso Schettini, Nicholas D. Kullman, Jorge E. Mendoza
Reinforcement Learning Optimization
  • OpenHail is an open-source environment tailored for electric ride-hailing fleet control using reinforcement learning.
  • The environment features a fixed-size observation-action interface for managing request assignment, repositioning, and charging.
  • An event-driven simulator captures the complexities of vehicle operations, including battery dynamics and charging constraints.
  • The decision-epoch mechanism allows for flexible policy interactions, supporting various control strategies.
Read more
NeuralCert: certified computational discovery of extremal mathematical constructions
Mark Patrick Roeling
Theory Optimization
  • NeuralCert separates discovery and certification processes, enhancing the rigor of mathematical proofs.
  • The framework allows for flexible neural parameterizations, improving the search for mathematical constructions.
  • NeuralCert has been successfully applied to extremal problems, demonstrating its capability in mathematical discovery.
  • The methodology emphasizes the importance of independent verification of mathematical claims.
Read more
Stable initialization without the CLT
Simon Kuang, Kyle Chickering, Xinfan Lin
Theory Optimization
  • Introduces uniform-phase initialization for sine activation networks, avoiding CLT-related errors.
  • Achieves full decoupling of layers, enhancing training stability.
  • Outperforms existing methods in neural representation tasks without requiring hyperparameter tuning.
  • Supports width scaling in neural networks, contributing to model performance.
Read more
PALM: Point-in-Time Adaptation for Financial Language Models
Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn
NLP Large Language Models Time Series
  • PALM offers a cost-effective alternative to annual pretraining for financial language models.
  • The necessity of annual pretraining runs is questioned; newer models do not consistently outperform older ones.
  • A low-rank adapter can effectively update a model's knowledge without modifying its pre-trained weights.
  • PALM outperforms traditional continued pretraining methods in various evaluations.
Read more
Reinforcement Learning of Communication in a Mesh of Small Language Models
Mehmet Kerem Turkcan
NLP Reinforcement Learning Large Language Models
  • TalkMesh enables decentralized communication among small language model agents to improve decision-making.
  • The system utilizes a confidence scoring mechanism to determine when agents should communicate hints and revisions.
  • Gossip consensus is employed to achieve a weighted voting mechanism without a central coordinator.
  • TalkMesh significantly improves accuracy over traditional self-consistency methods, demonstrating robustness against collusion.
Read more
LUCID: Learning Under Confounding for Inference and Discovery in Time Series
Mohammad Fesanghary
Time Series Graph Learning Theory
  • LUCID effectively identifies and adapts to different confounding regimes in time series data.
  • The method integrates seamlessly with existing causal discovery algorithms, enhancing their performance.
  • LUCID achieves a family-weighted directed, lag-resolved graph F1 score of 0.60, outperforming the best baseline by 0.19.
  • The approach is robust under various confounding conditions, including intermittent and heavy-tailed confounding.
Read more
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodolà
Efficient ML Optimization Theory
  • Introduction of Weight Pair Encoding (WeightPE) for neural network weight optimization.
  • Utilization of a lossy Re-Pair compressor to induce a smaller grammar in weights.
  • Demonstrated effectiveness on ViT-B/16 and ViT-L/16 models with CIFAR-10 dataset.
  • Achieved significant grammar size reduction with minimal impact on accuracy.
Read more
Gradient Surgery for Physics-Informed Neural Networks
Thomas Borsani, Giuseppe Di Fatta
Optimization Theory
  • PINNs face significant challenges due to conflicting task gradients during optimization.
  • The authors identify three distinct phases of gradient conflicts in PINN training.
  • PAM-GS is proposed as a solution to adaptively mitigate task interference.
  • Experimental results show PAM-GS outperforms existing methods on benchmark PDE problems.
Read more
Common-Mode Collapse and Recovery in Direct Feedback Alignment
Varun Reddy, Bernardo L. Sabatini, Houman Safaai
Theory Optimization
  • Common-mode collapse in DFA leads to stalled learning due to saturation of tanh units.
  • A mean-covariance decomposition helps understand the dynamics of error propagation in DFA.
  • Calibrating the readout to the class prior can suppress collapse and improve learning speed.
  • The severity of collapse varies with different network architectures and learning conditions.
Read more
NEXT: Physics-Informed Neuro-Spectral Exponential Time Differencing Architectures
Márcio Marques, Leonardo Mendonça, Leonardo M. Moreira, Christian Júnior de Oliveira, Vitor Balestro, Tiago Novello, Daniel Yukimura, Pavel Petrov, Lucas Nissenbaum
Theory Optimization
  • NEXT combines the strengths of Neuro-Spectral Architectures and exponential integrators to improve stability and accuracy for stiff PDEs.
  • The architecture is designed to inherently enforce causality and effectively represent high-frequency components.
  • NEXT demonstrates superior performance in benchmark tests compared to existing methods, particularly in stiff PDE scenarios.
  • The framework is applicable to inverse problems, allowing for parameter identification from limited data.
Read more
Peer-Grounded Counterfactual Path Planning for Chronic Health Management
Saman Khamesian, Hassan Ghasemzadeh
Graph Learning Optimization Time Series
  • Introduces POROS, a framework for incremental behavioral change in chronic health management.
  • Constructs a Behavioral Progression Graph that incorporates peer behavior to enhance motivation.
  • Demonstrates significant reductions in required behavioral changes for diabetes patients.
  • Aligns with self-efficacy and social comparison theories to improve patient engagement.
Read more
Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
Nux Li
NLP Large Language Models Efficient ML
  • Identifies two failure modes in quantizing looped transformers: feedback exposure and calibration blindness.
  • Demonstrates that per-channel INT4 quantization severely impacts performance at the loop-entry adapter.
  • Shows that accumulating the Hessian across recurrence steps outperforms traditional quantization methods.
  • Highlights the need for improved calibration techniques in looped transformer architectures.
Read more
WorldTS: World Modeling for Multimodal Covariate-aware Time Series Forecasting
Yuhan Zhu, Xiangfei Qiu, Hanyin Cheng, Wangmeng Shen, Chenjuan Guo, Bin Yang, Jilin Hu, Christian S. Jensen
Time Series Multimodal
  • WorldTS integrates multimodal covariates into time series forecasting to enhance predictive performance.
  • The framework utilizes a two-stage training approach to model latent state dynamics before decoding predictions back to the observation space.
  • Experiments on 21 datasets reveal that WorldTS outperforms traditional observation-space forecasting methods.
  • The model emphasizes the importance of latent representations in capturing the dynamics of complex systems.
Read more
EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting
Takumi Fujimoto, Hiroaki Nishi
Time Series
  • EPOC introduces a compressed residual state for online correction in multi-horizon forecasting.
  • The method retains low-order DCT coefficients and the final value of the preceding residual block.
  • EPOC achieves significant reductions in MSE and MAE while using less auxiliary state compared to existing methods.
  • The endpoint serves as a critical shared feature in the correction process, enhancing accuracy.
Read more
AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework
Aparajith Chandran, Juwon Kim, Saurav Jha, Pablo Castells, Florian Hottier
Large Language Models Optimization Efficient ML
  • Identification of five failure modes in production-scale AutoResearch: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation.
  • Development of a three-principle scaffolding design to mitigate identified failure modes.
  • Significant performance improvements achieved over hand-tuned baselines, including a 1.82× lift in Recall@6.
  • Autonomous design of a fallback system by the agent, increasing catalog coverage by 5.8×.
Read more
Does Uniform Discrete Diffusion Need Time?
Chunsan Hong, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Jong Chul Ye, Yuki Mitsufuji
NLP Generative Models Large Language Models
  • Population-optimal UDM predictors depend on time, but this dependence is often negligible in finite-data settings.
  • Trained language UDMs exhibit limited sensitivity to time over most of the diffusion trajectory.
  • Time-agnostic predictors can outperform time-conditioned models across various datasets and training objectives.
  • The results challenge the necessity of explicit time conditioning in UDMs, suggesting simpler architectures may be effective.
Read more
Parameters vs. Context: TRACE Fine-Tuning for Robust Retrieval-Augmented Generation
Zhengchen Huang, Yundong Sun, Minrui Song, Shuanglong Yao, Ye Liu, Ji Chen, Xing Wang
NLP Large Language Models Generative Models
  • TRACE introduces a debate-trace fine-tuning method for improved knowledge-source selection in RAG models.
  • The framework addresses answer incompleteness through regularization techniques that reinforce answer completeness.
  • Experiments show TRACE enhances robustness against misleading knowledge while preserving correct retrieved information.
  • The proposed methodology provides fine-grained supervision signals that improve the model's decision-making process.
Read more
Metacognitive Selective Ensemble for Mobile Systems
Sungmin Lee, Kichang Lee, Joonhee Lee, JaeYeon Park, Songkuk Kim, JeongGil Ko
Efficient ML Time Series
  • MetaSE reduces computational costs by maintaining a small active set of models for mobile sensing.
  • The framework utilizes temporal continuity in sensor data to evaluate model reliability efficiently.
  • MetaSE achieves performance comparable to full ensembles while being significantly faster and more memory-efficient.
  • The method outperforms traditional static and adaptive selection methods in various configurations.
Read more
Audio emotion recognition for atypical hearing
Ulysse Roussel
Audio & Speech Multimodal
  • Focus on audio emotion recognition for individuals with atypical auditory processing, particularly those with autism.
  • Exploration of generalizing affective responses from limited data to accommodate hypersensitive listeners.
  • Development of ecological data collection protocols tailored for assessing emotional reactions in real-world conditions.
  • Utilization of a fine-tuned model (CLAP) to validate methodologies on neurotypical datasets before applying them to atypical listeners.
Read more
QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices
Nicholas Foley, Devin Marinelli, Donny Moore, Diego Enriquez, Amanda Fernandez
Computer Vision Graph Learning Theory
  • QSV replaces the traditional attention mechanism with a single learned quaternion per token, coupling attention and message transformation.
  • Ablation studies indicate that the transport function is essential for model accuracy, while the routing weight can be simplified.
  • QSV shows competitive performance on CIFAR-10 and CIFAR-100 but does not outperform standard attention mechanisms on similar graphs.
  • The model's architecture leverages sparse kNN graphs on concentric Fibonacci spheres for efficient computation.
Read more
I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?
Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
Theory Reinforcement Learning Generative Models
  • Introduces a latent variable model for JEPA to study causal state recovery.
  • Develops a general information-theoretic objective combining likelihood and entropy maximization.
  • Establishes identifiability conditions for recovering latent causal states.
  • Implements an action-modulated JEPA (A-JEPA) based on theoretical findings.
Read more