AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

71 Papers today
8h Update frequency
7 Days of history
Does Transolver really need a Transformer?
Shizheng Wen, Siddhartha Mishra
Theory Efficient ML
  • Transolver's performance is not dependent on the Transformer architecture.
  • Slicing and deslicing operations are essential for maintaining model accuracy.
  • A constant linear map can replace token self-attention without loss of performance.
  • Theoretical insights confirm the universal approximation capability of the transformer-free Transolver.
Read more
What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use
Yixiao Chen, Ke Cheng, Jiangtao Guan, Shuo Huang, Yue Liu, Jun Zhang, Yuhong Liu, Jie Jiang
NLP Large Language Models Theory
  • Identification of three operational bottlenecks in language model training: Circuit, Store, and Use.
  • Introduction of a diagnosis-to-data principle that connects these bottlenecks with tailored supervision strategies.
  • Demonstration that early circuit organization improves learning outcomes over extensive training.
  • Distinction between writing content and invoking memory, highlighting the need for different supervisory approaches.
Read more
AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework
Aparajith Chandran, Juwon Kim, Saurav Jha, Pablo Castells, Florian Hottier
Large Language Models Optimization Efficient ML
  • Identification of five failure modes in production-scale AutoResearch: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation.
  • Development of a three-principle scaffolding design to address the identified failure modes.
  • Significant performance improvements over hand-tuned baselines, including a 1.82× lift in Recall@6 and a 2.1× coherence lift.
  • Autonomous design of a fallback mechanism that increased catalog coverage by 5.8×.
Read more
Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model
Kimon Antonios Provatas, Ilias Georgakopoulos-Soares
Efficient ML NLP Theory
  • The audit evaluates a commercial System-1 model's reliability on biosecurity-relevant benchmarks.
  • The model demonstrates strong calibration but variable accuracy depending on the task.
  • Significant instability in item-level decisions is observed due to answer option order sensitivity.
  • Selective averaging of probabilities can enhance accuracy without incurring high computational costs.
Read more
Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection
Jianan Liu, Chunguang Li
Time Series
  • Introduces a Bayesian Tensor Autoencoder framework for anomaly detection in multi-dimensional time series.
  • Incorporates a predictive prior that bridges the gap between reconstruction-based and prediction-based autoencoders.
  • Utilizes physical laws to enhance the modeling capability and mitigate over-generalization.
  • Demonstrates effectiveness through experiments on real-world datasets.
Read more
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
Yuchen Cai, Ding Cao, Qixiang Yin, Xin Xu, Kai Yang, Siye Wu, Pengyuan Wang, Jiaxuan Wang, Weijie Liu, Saiyong Yang, Guangzhong Sun, Guiquan Liu, Junfeng Fang
NLP Large Language Models Reinforcement Learning
  • Reinforcement learning enhances reasoning in large language models but its internal mechanisms are complex.
  • Vector steering reveals a low-dimensional effective manifold in activation space associated with RL performance gains.
  • Two geometric properties of this manifold are identified: Effective Manifold Capacity and Control Manifold Separation.
  • Alpha-Stabler framework stabilizes RL training and improves performance by managing activation gradients.
Read more
On the Capability and Limitation of Hard Prompt
Lijia Yu, Shuaitong Liu, Gaojie Jin, Xinyu Li, Xiao-Shan Gao
NLP Large Language Models Theory
  • Determining the existence of a hard prompt is NP-complete; finding an optimal hard prompt is NP-hard.
  • Hard prompts have fundamental limitations, including incompleteness and performance issues with short and long prompts.
  • Linear hard prompts can enhance transformer performance without the drawbacks of traditional hard prompts.
  • A necessary and sufficient condition for the generalizability of prompts is established based on prompt length and task size.
Read more
Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design
Fu Feng, Ruixiao Shi, Yucheng Xie, Jing Wang, Xin Geng
Robotics Optimization Reinforcement Learning
  • Introduction of Morphogene as a high-level latent blueprint for body-brain coordination.
  • GeCode reformulates morphology-control co-design as gene-driven exploration in a compact latent space.
  • Demonstrated significant performance improvements over existing methods in diverse design tasks.
  • Achieved an average of 2.5× faster convergence and 69.48% higher task performance.
Read more
HyperLabel: Multi-Label Classification via Hypergraph-Based Label Correlation Modeling
Peiyu Zhang, Heng Ping, Nikos Kanakaris, Yucheng Zhao, Shixuan Li, Wei Yang, Xiongye Xiao, Paul Bogdan
Graph Learning
  • Introduction of HyperLabel, an encoder-decoder framework for multi-label classification.
  • Construction of a label hypergraph to explicitly model multi-way label dependencies.
  • Implementation of HGNN+ for bidirectional message passing between features and labels.
  • Demonstration of state-of-the-art performance on multiple benchmark datasets.
Read more
Predictive Dual Smoothing for Column Generation
Senne Berden, Noah Schutte, Andrea Lodi, Tias Guns
Optimization
  • Introduction of predictive dual smoothing to enhance column generation efficiency.
  • Utilization of learned predictions of future duals to guide pricing subproblem.
  • Demonstrated significant reductions in generated columns and runtime in experiments.
  • Method maintains correctness of the column generation process.
Read more
Trust Guided Decision Transformer
Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath
Reinforcement Learning Robotics Optimization
  • Identifies rollout context mismatch as a critical failure mode in Decision Transformers.
  • Introduces a novel context selection mechanism that filters based on prediction error reliability.
  • Demonstrates that training-side improvements alone are insufficient for reliable performance.
  • TGDT outperforms existing methods in reducing prediction error and improving returns in various tasks.
Read more
Disentangling Lung-Cancer CT/LDCT AI: A Systematic Evidence Map of Clinical Tasks, Evidence Chains, and Translational Gaps
Surajit Das
Multimodal
  • The literature on lung cancer AI is predominantly focused on detection rather than risk prediction.
  • Only 10.6% of studies completed a comprehensive six-tier evidence chain, indicating significant attrition in the translational pathway.
  • The study introduces a Multi-Tier Evidence Graph (MTEG) to systematically analyze and visualize evidence chains in lung cancer AI research.
  • Clinical-task heterogeneity and inconsistent definitions of multimodal evidence are major challenges in synthesizing the literature.
Read more
Beyond the Graph: An Adaptive Meta-Learner Fuses Explainability, Weather, and Dynamics for Robust Bus ETA Prediction
Pratham Payra, Jagadish
Graph Learning Time Series Interpretability
  • Introduction of HYB(nm), a hybrid ensemble framework for bus ETA prediction.
  • Integration of multiple models to address nonlinear dynamics and weather effects.
  • Evaluation on extensive real-world data showing significant improvements in prediction accuracy.
  • Flexible architecture allows for tailored solutions for transit agencies.
Read more
seq2cause: One Autoregressive Backbone, Four Causal Discovery Tasks in Event Sequences
Hugo Math
Theory Time Series Generative Models
  • SEQ2CAUSE unifies four causal discovery tasks using a single autoregressive model.
  • The framework operates without task-specific retraining, enhancing efficiency.
  • A prediction-causality duality is established, linking prediction accuracy to causal identification.
  • SEQ2CAUSE is scalable, handling high-dimensional event types effectively.
Read more
Livin' on a Prior: Likelihood Score Approximation for Inverse Problems
Rostislav Makarov, Tal Peer, Danilo de Oliveira, Timo Gerkmann
Generative Models Computer Vision Audio & Speech
  • Introduces a new framework, Likelihood Score Approximation (LSA), for solving inverse problems.
  • Allows for learning unknown degradation models from few paired examples and known models from self-generated data.
  • Supports both deterministic and stochastic sampling, enhancing flexibility in model application.
  • Demonstrates competitive performance on ImageNet-256 and other benchmarks with minimal training data.
Read more
Product-Aware Deterministic Rounding for Quantized Matrix Multiplication
Piyush Sao, Narasinga Miniskar, Pedro Valero-Lara, Keita Teranishi, Sudip Seal
Optimization Theory Efficient ML
  • Dynamic rounding allows for informed rounding decisions based on weight blocks, reducing product errors.
  • Static rounding's expected product error can be characterized by the uncentered second moment and centered covariance.
  • Exact optimization for rounding decisions is NP-hard, motivating the use of additive guarantees.
  • Empirical results show significant error reduction with coordinated rounding and clipping-aware initialization.
Read more
BeatGraph: Self-Supervised Heartbeat Graphs for Infant ECG Representations from the Home Environment
Mohammad Nur Hossain Khan, M. S. Krafczyk, Beverly G. Bolster, Nancy McElwain, Mark A. Hasegawa-Johnson, Bashima Islam
Graph Learning Time Series
  • BeatGraph models heartbeats as nodes in a graph, improving representation of infant ECG data.
  • The model is pretrained on a new corpus of 3,408 hours of infant ECG recordings.
  • BeatGraph achieves significant performance improvements over existing ECG models on multiple tasks.
  • The approach demonstrates strong transferability across different age groups.
Read more
From Grey-Box to Green-Box: When can Physics-Informed Machine Learning Reduce Carbon Footprints in Structural Health Monitoring?
Daisy R. Bradley, Nathan A. Hinchliffe, Daniel J. Pitchforth, Matthew R. Jones, Elizabeth J. Cross
Efficient ML
  • Physics-informed machine learning (PIML) can reduce carbon emissions in structural health monitoring compared to traditional black-box models.
  • The study evaluates four PIML approaches, revealing that most have lower training emissions, except for input-augmented models.
  • Reducing training data requirements through PIML contributes to environmental savings by decreasing training duration.
  • A trade-off exists between the complexity of incorporating physics into models and the benefits of reduced data requirements.
Read more
Persistent Partners Raise Prices Among Learning Agents
Paul-Peter Arslan, Yubin Kim, Xiao Xiao
Reinforcement Learning Theory Optimization
  • Keeping the same partner significantly increases average profits and resting prices among learning agents.
  • Price increases occur even when rival prices are hidden, indicating that punishment is not the only factor at play.
  • The study employs a randomized experimental design to isolate the effects of partner persistence on pricing strategies.
  • Untrained Qwen2.5 models exhibit similar pricing behavior, suggesting broader implications for algorithmic pricing.
Read more
Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models
Shan Zhao, Ilija Trajkovic, Julia Kaltenborn, Yaniv Gurwicz, Peer Nowack, David Rolnick, Julien Boussard
Theory Interpretability Time Series
  • Introduces a hierarchical causal representation learning framework for climate modeling.
  • Explicitly separates internal climate variability from externally-driven responses.
  • Accurately predicts temperature changes under various future climate scenarios.
  • Demonstrates realistic responses to changes in greenhouse gas and aerosol concentrations.
Read more
$λ$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning
Berker Demirel, Clémentine Dominé, Valentino Maiorca, Marco Fumero, Marco Mondeli, Francesco Locatello
Computer Vision Theory
  • Introduces SACReg, a spectral anti-collapse regularizer that enhances representation rank in JE-SSL.
  • Demonstrates that existing JE-SSL methods may not prevent dimensional collapse in backbone representations.
  • λ-JEPA outperforms LeJEPA and VISReg on ImageNet-1k classification and improves transfer performance across multiple datasets.
  • The method is applicable to both image and video self-supervised learning tasks.
Read more
Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer
Seokyong Sheem, Hochang Lee, Suyeong Lee, Daekyum Kim
Theory
  • Semantic ablations can mislead conclusions about the benefits of semantic knowledge in predictive models.
  • Content sensitivity and predictive utility are distinct measures that provide different insights into model performance.
  • Choosing appropriate controls is crucial for accurately assessing the impact of semantic knowledge.
  • The benefit of semantic knowledge varies based on the reference condition used for comparison.
Read more
Online Learning via Learned Latent Bayesian Tracking
Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone
Time Series Optimization Efficient ML
  • AURA framework enables rapid online learning through a learned low-dimensional latent state-space model.
  • The method employs extended Kalman filtering for efficient single-step updates in the latent space.
  • AURA is validated in real-world scenarios, including wireless communication and image classification.
  • The approach shows significant improvements in adaptation speed and accuracy compared to traditional methods.
Read more
Unifying Distributional Training for One-Step Visual Generation
Chi Zhang, Haoyang Shi, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu
Computer Vision Generative Models
  • Introduces a unified framework for distributional training in visual generation.
  • Develops Mixture Gradient Flow (MGFlow) to model feature distributions with Gaussian mixtures.
  • Achieves state-of-the-art results on ImageNet, surpassing FD-Loss by 23% and 38%.
  • Implements MGFlow for text-to-image generation, outperforming previous multi-step models.
Read more
Gradient Surgery for Physics-Informed Neural Networks
Thomas Borsani, Giuseppe Di Fatta
Optimization Theory
  • PINNs face significant challenges due to conflicting gradients during optimization.
  • The authors identify three distinct phases of gradient conflicts in PINN training.
  • PAM-GS is proposed as a solution to adaptively manage task interference.
  • Experiments show PAM-GS outperforms existing optimization methods on benchmark PDE problems.
Read more
Quasi Linear Kernel Attention with Infinite Capacity
Nicolaj Rux, Johannes Hertrich, Sebastian Neumayer
NLP Large Language Models Efficient ML
  • Introduces a new capacity metric for kernels to measure expressivity in attention mechanisms.
  • Demonstrates that traditional expressive kernels have infinite capacity, while finite feature map-based kernels have limited capacity.
  • Proposes additive kernels that achieve quasi-linear computation while retaining infinite capacity.
  • Implements an efficient CUDA version of the proposed kernels, outperforming conventional attention methods for long sequences.
Read more
Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency
Xinhan Yang, Lu Lu, Shancong Mou
Optimization Theory Efficient ML
  • Introduces sketched tangent consistency loss (sTCL) for on-the-fly derivative-informed training of neural operators.
  • Eliminates the need for offline-generated derivative labels, reducing computational and storage costs.
  • Addresses challenges with stiff or ill-conditioned tangent operators through operator-aware loss-conditioning.
  • Achieves solution and Jacobian accuracy comparable to traditional methods while being more flexible.
Read more
Graph Forward Distribution Matching for Molecular Inverse Design
Yihan Zhu, Yuhan Liu, Brett Savoie, Tengfei Luo, Meng Jiang
Reinforcement Learning Generative Models Graph Learning
  • GRAPHFDM optimizes molecular design through a forward process, improving stability and property control.
  • The method incorporates valid generations into a reward-tilted target distribution for enhanced optimization.
  • It achieves up to 53.0% reduction in MAE compared to the strongest baseline methods.
  • Maintains chemical validity above 0.99 across multiple property conditions.
Read more
Cross-Material Support Transfer for Core-Loss Prediction Under Waveform Covariate Shift
Cong Yao, Chunye Gong
Theory Optimization Efficient ML
  • Identifies the core issue of covariate shift in core-loss prediction rather than class imbalance.
  • Introduces Material-Identity Support Transfer (MIST) for leveraging information from sibling materials.
  • Demonstrates that joint training can significantly improve prediction accuracy for underrepresented materials.
  • Achieves lower prediction errors with fewer parameters compared to previous models.
Read more
PALM: Point-in-Time Adaptation for Financial Language Models
Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn
NLP Large Language Models
  • Annual pretraining for financial language models is shown to be unnecessary.
  • PALM introduces a low-rank adapter that adapts existing models without retraining.
  • The method effectively avoids look-ahead bias while maintaining model eligibility.
  • Extensive experiments validate that small adapters outperform continued pretraining.
Read more
Robust Graph Clustering Network for Multiple Missing Data
Keyuan Qiu, Renda Han, Zhen Tang, Qiang He, Xingwei Wang, Wenxin Zhang, Guangzhen Yao, Junxin Chen, Qingjian Ni
Graph Learning
  • First attempt to address simultaneous missing node attributes and graph structure in graph clustering.
  • Introduces a view-decoupled dual-branch imputation method to enhance data recovery.
  • Employs a multi-hyperspherical mixture prior for improved cluster separation.
  • Integrates boundary-aware contrastive learning to sharpen cluster demarcation.
Read more
Metacognitive Selective Ensemble for Mobile Systems
Sungmin Lee, Kichang Lee, Joonhee Lee, JaeYeon Park, Songkuk Kim, JeongGil Ko
Efficient ML Time Series Computer Vision
  • MetaSE maintains a small active set of models to reduce computational costs in mobile sensing.
  • The framework leverages short-term persistence in model reliability to make efficient selection decisions.
  • MetaSE outperforms fixed and adaptive ensemble methods while using fewer resources.
  • The approach is validated across multiple HAR datasets and model architectures.
Read more
SMAT: Simple and Efficient Merge-Aware Training
Yanggan Gu, Yuanyi Wang, Zhen Li, Shuo Cai, Yuhang Liu, Junzhuo Li, Zihao Wang, Hongxia Yang
Efficient ML Multimodal Optimization
  • SMAT simulates common merging operations during expert training to enhance merged performance.
  • The method achieves improved performance with less than 2% training-time overhead compared to standard fine-tuning.
  • Periodic scheduling and efficient parameter operations contribute to SMAT's efficiency.
  • SMAT outperforms existing merge-aware training methods across multiple model architectures.
Read more
Distribution-Conditioned Task Routing for Class-Incremental Learning
Longhuan Xu, Zhipeng Zhou, Wei Ji, Chunyan Miao, Peilin Zhao, Lijun Zhang
Efficient ML Computer Vision
  • Introduces a novel post-hoc task routing framework for class-incremental learning.
  • Identifies three sources of routing error: feature-level, task-level, and class-level misalignment.
  • Proposes Feature Distribution Calibration (FDC) to address these misalignments without additional training.
  • Demonstrates significant accuracy improvements across multiple benchmarks and methods.
Read more
Propagate, Then Sharpen: Post-Hoc Refinement of Frozen Node Classifiers
Preben Johnsen Bentdal, Nello Blaser, Xue-Cheng Tai
Graph Learning
  • Introduction of Propagate, Then Sharpen (PtS) for refining frozen node classifiers.
  • PtS alternates between probability propagation and sharpening, improving accuracy without retraining.
  • Demonstrated significant accuracy gains over APPNP, especially under feature corruption.
  • Robustness against oversmoothing effects, maintaining accuracy across multiple propagation steps.
Read more
MultiEcho: An Experimental Science of Learned Worlds
Meng Zhu, Airui Zhang
Theory Generative Models Computer Vision
  • MultiEcho framework allows for the estimation of learned world laws through controlled counterfactual interventions.
  • Experiments reveal significant variability in response predictability and physical accuracy across different models and contexts.
  • Geometric regularities do not necessarily indicate the emergence of physical laws, highlighting the complexity of learned representations.
  • The framework provides a multidimensional approach to studying learned worlds, integrating time, space, and counterfactual interventions.
Read more
An Attention-Driven Heterogeneous GNN Model for Credit Card Fraud Detection
Kathiresan Jayabalan, Sethuraman Radhakrishnan
Graph Learning
  • Proposes a novel framework for credit card fraud detection using a heterogeneous GNN model.
  • Utilizes SMOTE-Tomek for data balancing to address the imbalanced dataset issue.
  • Employs an attention-based message passing technique to capture complex transaction relationships.
  • Achieves high performance metrics, indicating effectiveness in detecting credit card fraud.
Read more
Optimal Nonparametric Dynamic Pricing with Censored Demand and Adversarial Inventory
Mengxiao Zhang, Yingfei Wang, Haipeng Luo
Optimization Theory
  • Introduces a nonparametric adversarial-inventory model for dynamic pricing.
  • Proposes two algorithms: Double-Grid-UCB and Threshold-UCB, with the latter achieving improved regret rates.
  • Establishes a minimax optimality lower bound for the proposed algorithms.
  • Demonstrates superior performance of Threshold-UCB in extensive experimental evaluations.
Read more
MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR
Yangyang Ren, Haodong Zhu, Sheng Xu, Yanjing Li, Nikolai Yu. Zolotykh, Wentao Zhang, Baochang Zhang
Reinforcement Learning Large Language Models Efficient ML
  • MaPP addresses inefficiencies in RLVR by denoising advantage estimation and improving prompt selection.
  • The framework introduces a composition-invariant intrinsic advantage estimator to reduce gradient estimation errors.
  • MaPP consistently outperforms existing RLVR methods, achieving state-of-the-art results with fewer rollouts.
Read more
Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons
Kristi Topollai, Anna Choromanska
NLP Large Language Models Optimization
  • Learning-rate warmup duration should be treated as a horizon-dependent hyperparameter rather than a fixed heuristic.
  • A quadratic model effectively captures the tradeoff between early performance and long-term gains in training.
  • Optimal warmup duration varies significantly with peak learning rate and training horizon.
  • The study provides a compact scaling law that can predict warmup durations from shorter training runs.
Read more
Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference
Xianpeng Shang, Canbin Huang, Jiang Li, Tian Lan, Qianyi Cai, Xiaojun Quan, Xiangdong Su
Large Language Models Efficient ML
  • Distance-KV introduces a novel static KV retention pattern based on relative distance, improving cache efficiency.
  • The method allows for significant reductions in KV cache memory usage (up to 65.4%) and decoding speed (1.66× faster) compared to dense methods.
  • Distance-KV outperforms existing KV cache compression techniques, achieving up to 9.3 points improvement on benchmarks.
  • The approach is model- and budget-specific, enabling reuse across different inputs without online optimization.
Read more
Two Heads Are Better Than One: Aggregating Weaker LLMs for Better Forecasts
Cheng Peng, Ruixi Luo, Zhi Chen, Wei Tang
NLP Large Language Models Time Series
  • Weak-to-strong aggregation can effectively improve forecasting accuracy using weaker LLMs.
  • Learned linear pooling outperforms the strongest individual forecaster in multiple comparison groups.
  • Improvements in performance do not depend on the presence of a near-best model.
  • Adding more models does not consistently lead to better results.
Read more
DRIFT: Disentangled Responsive-Invariant Flow Transport for Single-Cell Perturbation Prediction
Mustapha Bounoua, Giulio Franzese, Pietro Michiardi
Generative Models Theory Optimization
  • Introduces DRIFT, a framework for disentangling responsive and invariant cell states.
  • Utilizes a variational encoder and conditional flow matching to model perturbation effects.
  • Outperforms existing methods in predicting cellular responses to unseen and combinatorial perturbations.
  • Addresses limitations of traditional methods by separating intrinsic variability from perturbation effects.
Read more
Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization
Mahesh Reddy Pagadala
Efficient ML Theory Optimization
  • Identifies deterministic overhead regime switching in LSTM based on fine-grained memory budget adjustments.
  • Reveals a non-monotone feasibility behavior in ResNet-32, with specific budget ranges leading to OOM errors.
  • Demonstrates that the slow execution regime in LSTM is due to broad re-eviction of nearly the entire working set.
  • Provides ablation evidence linking the instability in LSTM to the joint size-staleness scoring term.
Read more
Progressive Memory Transformer: Memory-Aware Attention for Time-Series
Tord Sture Stangeland, Andreas Köhler, Steffen Mæland, Adín Ramíres Rivera
Time Series
  • PMT introduces a multi-scale contrastive learning framework for time-series data.
  • The architecture features writable, window-aligned memory to expose mid-range representations.
  • Three distinct contrastive objectives are employed to supervise representations at different scales.
  • PMT achieves strong performance in low-label classification and competitive forecasting.
Read more
Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality
Linhao Wang, Yiyan Fan, Dongjin Huang
Reinforcement Learning Theory Robotics
  • Introduces Decision-Relevant Prediction Error (DRPE) as a better metric for evaluating world models.
  • Develops an iso-error evaluation protocol to isolate the effects of error allocation on planning performance.
  • Demonstrates that total prediction error is weakly correlated with planning success, while DRPE shows a strong correlation.
  • Finds that error relevance varies by task and becomes more significant with deeper planning.
Read more
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng, Weiqiang Zhu, Zhenhai Ji, Zhengning Wang
Reinforcement Learning Graph Learning Robotics
  • HySTAR separates credit assignment from adaptive representation learning, enhancing stability in MARL.
  • The framework utilizes an anchored hypergraph for consistent value decomposition.
  • Experiments show HySTAR outperforms MAPPO and other baselines across various scenarios.
  • The method effectively handles dynamic agent interactions and changing active-agent sets.
Read more
GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for De Novo Peptide Sequencing
Abdellah El Mekki, Laks V.S. Lakshmanan, Muhammad Abdul-Mageed
NLP
  • Introduces a decoder-guided latent imputation objective to prioritize relevant missing fragments.
  • Incorporates mass-aware attention using rotary embeddings to model mass differences between peaks.
  • Demonstrates significant improvements in sequencing precision on the NovoBench dataset.
  • Retains a standard architecture for inference, simplifying implementation.
Read more
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodolà
Efficient ML Theory Optimization
  • Introduction of Weight Pair Encoding (WeightPE) for neural network weight optimization.
  • Utilizes a lossy Re-Pair compressor to induce a smaller grammar in weights.
  • Achieves significant reductions in grammar size with minimal impact on accuracy.
  • Demonstrates the method on ViT-B/16 and ViT-L/16 models fine-tuned on CIFAR-10.
Read more
TRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac Care
Lovely Yeswanth Panchumarthi, Andrew Lu, Saurabh Kataria, Delgersuren Bold, Minxiao Wang, Runze Yan, Patricia Dykes, Brian J. Gow, Tom J. Pollard, Jessica K. Zègre-Hemsey, Dillon J. Dzikowicz, Lekshmi Kumar, Xiao Hu, Ran Xiao
Multimodal Time Series NLP
  • TRACE integrates unimodal and cross-modal learning to enhance ECG representation.
  • The model employs uncertainty-weighted multi-task learning to balance loss contributions.
  • TRACE significantly outperforms existing models in arrhythmia classification and ACO detection.
  • Real-world validation shows TRACE's clinical utility in acute cardiac scenarios.
Read more
Hardware-Aware Features for CUTLASS Kernel Selection
Shriram Chandran, Dominic Rinderer, Yakup Budanaz, Alexandru Calotoiu, Marcin Copik, Torsten Hoefler
Optimization Efficient ML
  • Introduction of hardware-aware representations for kernel selection in CUTLASS.
  • Construction of a large dataset (4.9 million kernels) for training selection models.
  • Significant reduction in selection regret compared to traditional methods.
  • Demonstration of data-efficient transfer learning capabilities within CUTLASS.
Read more
Towards Universal Representation-Based Process Control
Jinmyeong Choi, Taesup Kim, Artur Dubrawski
Time Series
  • Formulates window-level time series monitoring as a reference-relative process control problem.
  • Proposes a nonparametric framework with conformal calibration for regime consistency assessment.
  • Accommodates cyclostationary processes as stable operating regimes, addressing classical diagnostic limitations.
  • Demonstrates robustness to structural deviations and potential for diagnostic attribution beyond binary detection.
Read more
Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
Marcin Kostrzewa, Maciej Zięba
Interpretability
  • Proposes a unified evaluation protocol for robust counterfactual explanations across different model changes.
  • Demonstrates that existing robustness scores are not comparable due to method-specific evaluations.
  • Finds significant variability in the performance of robust methods depending on the type of model change.
  • Highlights the importance of separating generation performance from robustness in evaluations.
Read more
Does Uniform Discrete Diffusion Need Time?
Chunsan Hong, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Jong Chul Ye, Yuki Mitsufuji
NLP Generative Models Large Language Models
  • Population-optimal UDM predictors depend on time, but this dependence is often negligible in finite data settings.
  • Time-agnostic predictors in UDMs can achieve competitive performance compared to time-conditioned models.
  • The predictive benefit of time conditioning is primarily observed in high-noise scenarios.
  • The study provides a quantitative understanding of how time sensitivity relates to the separation of training examples.
Read more
Feedback-Robust AI for Patient Knowledge Graphs
Mohammed Sameer Syed
Graph Learning Time Series Interpretability
  • Introduction of ClosedLoopBench, a benchmark for evaluating temporal relations in clinical settings.
  • Development of feedback-robust patient knowledge graphs that integrate evidence and typed relations.
  • Demonstration of significant biases in traditional correlational methods for constructing patient KGs.
  • Identification of consistent physiological response patterns across patients despite individual variability.
Read more
Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
Alireza Abdollahpoorrostam, Ehsan Sharifian, Buse Şen, Marco Cuturi, Daniel Kuhn
Optimization Theory Robotics
  • Introduces a penalized DRO formulation that incorporates Wasserstein penalties for adversarial distributions.
  • Proves that optimal transport maps for adversarial training are cyclically monotone.
  • Develops Multi-start Particle Ascent (MPA) to enforce cyclical monotonicity in adversarial training.
  • Proposes using input-convex neural networks to parameterize adversarial transport maps.
Read more
A Comparative Analysis of Attention versus State-Space Models for In-Context Learning
Enes Arda, Semih Cayci, Atilla Eryilmaz
Theory
  • Introduces belief geometry as a unified framework for comparing attention and state-space models.
  • Identifies three key capabilities for sequential learning: evidence assembly, belief maintenance, and addressing.
  • Demonstrates that SSMs achieve optimal performance in belief maintenance and have memory advantages.
  • Shows that attention mechanisms have a significant advantage in content addressing due to their softmax selection process.
Read more
Arbitrary-Accuracy Neural Approximation with Optimal Neuron Count and Near-Optimal Bit Complexity
Zilan Cheng, Li-Lian Wang, Zhongjian Wang
Theory Efficient ML
  • Establishes that d + 1 is the minimum neuron count for arbitrary-accuracy approximation of Hölder-continuous functions.
  • Introduces a fixed activation function that allows for efficient neural network construction.
  • Demonstrates near-optimal bit complexity for the proposed network architectures.
  • Provides explicit constructions with rational parameters that can be verified in exact arithmetic.
Read more
From Phase Transition to Systemic Failure: A Decoupled Analytics Framework for GNN Robustness
Shuai Yan, Dan Peng, Jie Li, Xiaodong Huang, Ke Wang
Graph Learning
  • Developed a decoupled quantitative data analytics testbed for GNN robustness evaluation.
  • Identified distinct failure modes in GNNs: gradual degradation under label noise and severe degradation under distribution shift.
  • Demonstrated that GNN architecture does not compensate for data integrity issues.
  • Emphasized the importance of monitoring input drift over mitigating supervision noise.
Read more
Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning
Yangyang Ren, Haodong Zhu, Linlin Yang, Sheng Xu, Peichao Lai, Baochang Zhang
Reinforcement Learning Large Language Models Robotics
  • CRBC merges rollouts into a finite empirical process for better evidence aggregation.
  • The method allows recursive propagation of evidence through shared anchor states.
  • CRBC consistently outperforms existing methods in long-horizon reinforcement learning tasks.
  • The approach achieves state-of-the-art performance on multiple benchmarks.
Read more
Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
Tamim Zoabi, Ameen Ali, Lior Wolf
Reinforcement Learning Robotics Computer Vision
  • H-JEPA separates perceptual representation from control state, enhancing planning efficiency.
  • Utilizes a Bures-Wasserstein prior for isotropic regularization of perceptual codes.
  • Introduces port-inverse consistency (PIC) for effective action readout and error reweighting.
  • Achieves superior performance on pixel-based control tasks compared to existing models.
Read more
Rethinking Contextualization by Reinterpreting Attention Head Channels
Hakaze Cho, Haolin Yang, Zhun Sun, Naoya Inoue, Benjamin Heinzerling, Kentaro Inui
NLP Large Language Models Interpretability
  • Contextualization is influenced by the varying information levels of words.
  • Less informative words absorb contextual information more selectively.
  • Attention heads can be reinterpreted as channels governed by singular vectors.
  • The study provides a framework for automated interpretation of attention mechanisms.
Read more
Task-Aware Discretization of Differentiable Logic Gate Networks
Thore Gerlach
Efficient ML Theory
  • Introduces a task-risk perspective on discretization, distinguishing between conventional discretization gaps and discrete optimality gaps.
  • Demonstrates that local argmax selection can lead to significant task performance losses, even with globally optimal relaxed models.
  • Develops a task-aware gate selection method that utilizes conditional risk and provides efficient first-order approximations.
  • Empirical analysis reveals the importance of locality in gate selection, leading to a progressive discretization strategy that improves task performance.
Read more
Moment-guided edge sampling
Weibin Cai, Reza Zafarani
Graph Learning
  • Introduces a moment-guided edge sampling framework that connects local edge edits to global graph structure.
  • Develops efficient combinatorial and low-rank methods for computing moment changes.
  • Demonstrates that moment-preserving sampling retains important structural properties.
  • Shows the impact of edge structures on supervised node classification and competitive performance in contrastive learning.
Read more
Structured Neural SDEs for Functional Calibration
Francesco Piatti, Andrea Iannucci, Thomas Cass
Generative Models Theory Efficient ML
  • Introduction of SLiSDE, a structured linear Neural SDE model for functional calibration.
  • Achieves O(log T) parallel simulation efficiency through structured linear layers.
  • Gated in-flow stacking enhances model expressivity while preserving computational efficiency.
  • Incorporates a Girsanov tilt for improved performance on rare-event calibration tasks.
Read more
What Should We Freeze? Guarded Freezing: Connectivity Shapes the Fine-Tuning of Pretrained Models
Leonel Aguilar
Computer Vision NLP Large Language Models
  • Guarded Freezing optimizes the fine-tuning process by selectively freezing weights based on connectivity.
  • Removal-value and drift-value are introduced as new metrics for selecting which weights to freeze.
  • The proposed methods show improved retention of old-task accuracy compared to existing techniques.
  • Experiments demonstrate the effectiveness of Guarded Freezing across different model architectures.
Read more
Geometry-Aware Operator Families for Structured Representation Learning
Zuyuan Zhang, Fei Xu Yu, Tian Lan
Graph Learning Time Series Theory
  • Introduction of Geometry-Induced Operator Families (GIOF) for structured representation learning.
  • GIOF transforms fixed geometric descriptors into a structured family of propagation operators.
  • Dynamic selection of operators based on context enhances adaptability and performance.
  • Theoretical guarantees established for stability, locality, and parameter efficiency.
Read more
Saturation-Insensitive Dueling Bandits with General Function Approximation
Chenggong Zhang, Xuheng Li, Qiwei Di, Weitong Zhang, Quanquan Gu
Reinforcement Learning Theory Efficient ML
  • Introduction of SI-CDB, a saturation-insensitive algorithm for contextual dueling bandits.
  • The algorithm employs asymmetric arm selection to mitigate saturation effects in reward learning.
  • Localized Eluder dimension analysis is used to derive regret bounds, showing improved performance over existing methods.
  • SI-CDB achieves near-optimal regret bounds without the exponential dependence on reward scale.
Read more
Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra
Matthieu Le Lain, Gaël Cessateur, Sébastien Lefèvre
Computer Vision Time Series Theory
  • Introduction of a self-supervised encoder for auroral spectra that outperforms traditional supervised methods.
  • Demonstrated effectiveness of a 1D Vision Transformer with masked autoencoder on unlabelled spectral data.
  • Comparison of in-domain pretraining with existing astronomical models reveals limitations in transferability.
  • Achieved significant improvements in classification metrics, indicating the potential of self-supervised learning in spectroscopy.
Read more
PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem
Mateo Toro Diz, Jonathan Hoss, Noah Klarmann
Reinforcement Learning Optimization
  • Introduction of PORL, combining online pretraining with offline fine-tuning for JSSP.
  • Utilization of a KL-divergence constraint to maintain policy stability during adaptation.
  • Demonstrated superior performance of PORL over standalone offline RL and traditional scheduling methods.
  • Reduced sensitivity to dataset quality, enhancing applicability in real-world scenarios.
Read more
Rondo: Unsupervised Discovery of Recurring Temporal Structure
Yingtian Shi, Ankith Chandra, Thomas Plötz
Time Series
  • RONDO models recurring hierarchical structures in temporal data, addressing limitations of existing methods.
  • The approach uses UnitAlign for reusable unit discovery and MotifFormer for higher-level motif modeling.
  • RONDO shows significant improvements in recurring structure recovery, especially in limited-data and evolving contexts.
  • The method allows for continual refinement of unit and motif vocabularies as new data emerges.
Read more