AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

66 Papers today
8h Update frequency
7 Days of history
From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs
Tina Vartziotis, Rodopi Kosteli, Elli Vartziotis, George Dasoulas, Michael Keckeisen, Konstantinos Skianis, Sotirios Kotsopoulos, Francesca Dominici
Large Language Models Efficient ML
  • Introduces a novel methodology for estimating LLM inference energy on modern GPUs.
  • Separates energy estimates into components for detailed analysis of energy consumption.
  • Provides a framework for comparing models and optimizing for energy efficiency.
  • Focuses on transparency and reproducibility in energy estimation.
Read more
EvoCause: LLM-Guided Evolution of Causal Graphs for Root Cause Analysis
Lei Zan, Keli Zhang, Shifeng Xie, Jiale Zheng, Zehao Xiao, Zhiwei Dong, Ke Zhang, Ruichu Cai, Malik Tiomoko, Lujia Pan
Graph Learning Large Language Models Interpretability
  • EvoCause integrates LLM-guided graph editing with deterministic validation to refine causal graphs for RCA.
  • The framework allows for the incorporation of expert feedback without modifying the underlying causal discovery algorithms.
  • TeleRCA, a large-scale expert-annotated benchmark, is introduced to evaluate RCA methods.
  • EvoCause outperforms traditional methods in both synthetic and real-world scenarios, emphasizing the role of alarm title information.
Read more
When Does Explicit View Routing Work? A Controlled Study of Multi-View Graph-Text Alignment
Xiao Yue, Guangzhi Qu
Graph Learning Multimodal NLP
  • Introduces a controlled framework to differentiate between semantic routing and architectural channelization in graph-text retrieval.
  • Demonstrates that correct routing improves retrieval performance for labels and properties, but not consistently for topology.
  • Shows that a joint model can outperform separately trained specialists in aggregate retrieval scores.
  • Evaluates the effects of paraphrase augmentation on retrieval quality, revealing trade-offs between robustness and canonical retrieval.
Read more
Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding
Weiye Shi, Fanxu Meng, Muhan Zhang
Large Language Models Efficient ML Optimization
  • Conversion from MHA/GQA to MLA can lead to a loss of proposal agreement necessary for effective speculative decoding.
  • The proposed functional reconstruction method optimizes draft quality without requiring verifier supervision or changes to the inference graph.
  • The method is converter-agnostic and can be applied to various model architectures and configurations.
  • Functional reconstruction improves acceptance rates in 37 out of 64 task configurations evaluated.
Read more
RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning
Pushkal Kumar, Tucker Nielson, Tanish Kolhe, Shubham Zala, Vincent Li
NLP Large Language Models
  • RAGuard introduces a two-layer defense framework against data poisoning in RAG systems.
  • The first layer uses adversarial fine-tuning to improve the retriever's ability to downrank malicious passages.
  • ZKIP operates without requiring poison labels or ground-truth answers, making it adaptable to unseen attack types.
  • The combined defense effectively reduces attack success rates to 0.000 while maintaining retrieval quality.
Read more
DAS-PMVC: A Framework for Partial Multi-View Clustering via Dual Alignment and Structure Enhancement
Shubin Ma, Liang Zhao, Chuanye He, Zhenjiao Liu, Liang Zou, Lin Yuanbo Wu, Yu Shao
Graph Learning Multimodal Optimization
  • DAS-PMVC addresses the partial view alignment problem in multi-view clustering.
  • The framework employs a dual alignment strategy to enhance feature alignment and clustering accuracy.
  • Utilizes multi-view graph convolutional networks for improved feature extraction.
  • Experimental results show superior clustering performance compared to existing methods.
Read more
NMINE: Normalized Mutual Information Neural Estimation
Petra Eerikinharju, Marko Tuononen, Ville Hautamäki
Theory
  • Introduction of a fully neural estimator for normalized mutual information.
  • Combines MINE-based mutual information estimation with neural entropy estimation.
  • Derives entropy estimates from neural divergence relative to uniform distributions.
  • Demonstrates improved accuracy over traditional k-nearest-neighbor-based estimators.
Read more
Real-Time Hard Peak Age-of-Information Safety with No-Regret Learning
Wentao Zhang, Wentao Mo
Optimization Theory Reinforcement Learning
  • Introduces OCO-PAoI-Hard, a novel scheduling framework for peak AoI safety in IoT systems.
  • Achieves zero per-slot deadline violations under adversarial conditions while maintaining O(√T) regret.
  • Utilizes a causal proposal-shield-update loop for real-time scheduling and feasibility enforcement.
  • Provides comprehensive theoretical guarantees and empirical validation against multiple baselines.
Read more
From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE Inference
Tianyang Zhu
Large Language Models Theory Efficient ML
  • Different expert-reduction orders can produce observable differences in sparse-MoE executions.
  • The study identifies the significance of operand representation and accumulator precision in model compatibility.
  • Two state boundaries are validated for reproducing controlled divergent trajectories in autoregressive decoding.
  • Identical emitted tokens may not imply identical execution states, leading to potential output divergence.
Read more
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
Enjun Du, Hange Zhou, Chenxu Du, Siyi Liu, Zirong Chen, Ziyu Zheng, Yongqi Zhang
Multimodal
  • Introduces a Structured Evidence Ledger for multimodal agentic reasoning.
  • Identifies four recurring failure patterns that affect trajectory faithfulness.
  • Implements a Three-Layer Grounding Protocol to verify evidence integrity.
  • Enhances reasoning efficiency with an Adaptive Dual-Path Dispatcher.
Read more
THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model
Yixin Peng, Diego Collarana, Er Jin, Stefan Decker
Graph Learning
  • THGFM integrates cross-type transfer and relation-aware specialization in a unified framework.
  • The model employs a dual-path architecture with SSTA and RTTA branches.
  • Type-Conditioned Non-Competitive Gated Sum Fusion allows independent feature weighting.
  • Rotary Temporal Attention directly incorporates relative time into attention mechanisms.
Read more
From Unsupervised Subgroups to Hypothetical State-Intervention Policies: An Evaluation of Selected Subgrouping Methods in Observational Health Data
Vasundhara Acharya, Bulent Yener
Theory Interpretability Optimization
  • Proposes a framework for subgroup analysis that avoids reliance on outcome data.
  • Evaluates various unsupervised clustering methods for policy prioritization in health interventions.
  • Findings suggest no single subgrouping method consistently dominates across different policy scenarios.
  • Results should be interpreted as assumption-dependent evidence rather than definitive proof of intervention effectiveness.
Read more
Existence-Field Diffusion Model for Spatial Point Processes with Variable Cardinality
Xiaoyin Pan, Christian R. Shelton, Rakshith Mahishi, Chengkuan Hong
Generative Models
  • Introduces a novel EFDM that models spatial point processes with variable cardinality.
  • Overcomes limitations of existing methods by integrating cardinality and spatial structure in a unified diffusion process.
  • Utilizes existence variables to represent the presence of points, allowing for continuous modeling.
  • Demonstrates improved performance on synthetic and real-world datasets compared to previous approaches.
Read more
Contrastive Concept Importance: Explaining Pairwise Class Decisions Through Automatically Extracted Concept Representations
Roel Visser, Isaac Roberts, Barbara Hammer
Interpretability Computer Vision
  • Introduces Contrastive Concept Importance (CCI) for explaining pairwise class decisions.
  • Provides signed attributions indicating support for target vs. foil classes.
  • Evaluates method on ImageNet class pairs, revealing class-pair specific behaviors.
  • Distinguishes between globally important concepts and those affecting specific class distinctions.
Read more
From Interface to Inference: Eliciting Any-Order Inference from Any-Order Models
Seunggeun Kim, Jaeyeon Kim, Taekyun Lee, Yuyuan Chen, Yilun Du, Sham Kakade, Sitan Chen
NLP Large Language Models Generative Models
  • Identifies the limitations of autoregressive models in non-causal reasoning tasks.
  • Proposes two novel approaches to achieve genuine any-order inference: insertion-based and latent-space masked diffusion.
  • Demonstrates that positional uncertainty is a fundamental bottleneck in achieving any-order inference.
  • Empirical results show significant improvements in performance for code generation and mathematical reasoning tasks.
Read more
S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring
Glenn Anta Bucagu, Thorir Mar Ingolfsson, Yawei Li, Luca Benini
Time Series Efficient ML
  • S-CEReBrO introduces a Windowed Alternating Attention mechanism that maintains constant memory usage.
  • The architecture can process EEG signals 100 times longer than traditional full self-attention models.
  • It achieves state-of-the-art performance on multiple EEG analysis tasks while using fewer parameters.
  • The model demonstrates a significant increase in inference throughput and reduced memory requirements.
Read more
Benchmarking ConvLSTM for One-Day-Ahead IMDAA Rainfall-Field Prediction across Four Indian Cities
Tanmay Ghosh, Shaurabh Anand, Rakesh Gomaji Nannewar, Nithin Nagaraj
Time Series
  • ConvLSTM does not consistently outperform simpler forecasting models for rainfall prediction.
  • FC-LSTM achieved the lowest domain-mean rainfall error in three out of four cities.
  • Persistence model outperformed all others in high-rainfall day detection across all cities.
  • The effectiveness of models varies significantly based on input types and evaluation metrics.
Read more
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
Xingjian Wu, Junlin Liu, Xingchen Liu, Xuhang Zhu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai
Reinforcement Learning Large Language Models Optimization
  • CRPO enhances OPSD by applying contrastive learning principles to mitigate exposure bias.
  • The method uses predictive entropy to classify positions for effective distillation.
  • CRPO shows significant improvements in training stability and generalization across diverse benchmarks.
  • The theoretical foundation of CRPO aligns with logit-wise credit assignment, ensuring robust performance.
Read more
AgentGFM: A Graph Foundation Model with Node-Agent Information-Flow Control
Jingbo Cui, Jitao Zhao, Di Jin, Dongxiao He
Graph Learning
  • AgentGFM allows for node-specific information propagation decisions, enhancing adaptability across different graph structures.
  • The model employs a shared end-to-end trainable policy for all nodes, promoting transferability of learned behaviors.
  • The predict-act-observe-correct process enables nodes to refine their decisions based on feedback, improving information flow control.
  • Extensive experiments show that AgentGFM outperforms traditional GFMs in various transfer scenarios.
Read more
Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
Pankayaraj Pathmanathan, Furong Huang
Large Language Models NLP Reinforcement Learning
  • Compliance2LoRA provides a unified framework for dynamic safety policy compliance in large reasoning models.
  • The hypernetwork generates LoRA weights based on customizable policy inputs, enabling efficient model adaptation.
  • The framework allows for on-demand adjustments to policy compliance without sacrificing task performance.
  • Training involves both supervised and reinforcement learning techniques to enhance model generalization.
Read more
Cybersecurity Detection Classification with Reasoning-enabled Language Models
Amol Khanna, Manu Nandan, Cristian Viorel Popa, Joan Pujol-Roig, Diana Bolocan, Laura Vasilie, Alexandru Apostu, Chase Helwig, Mihaela Gaman, Michael Brautbar, Edward Raff, Chase Midler, Sven Krasser
NLP Large Language Models Reinforcement Learning
  • First detection-triage classifier trained to produce CoT reasoning for cybersecurity detections.
  • Introduced a four-stage training process that enhances classifier performance and confidence estimation.
  • Demonstrated that an untrained confidence calibrator leads to a collapse in high-confidence recall.
  • A specialized 30B model outperforms general-purpose models in cybersecurity detection tasks.
Read more
Search Strategies for Optimal Classification and Regression Trees
Jacobus G. M. van der Linden, Mim van den Bos, Emir Demirović
Optimization Interpretability Efficient ML
  • Introduces a unified algorithmic framework for Optimal Decision Trees (ODTs) that allows for the comparison of various search strategies.
  • Empirically evaluates 18 search strategies, marking one of the largest evaluations in the field.
  • Demonstrates that the best search strategy significantly improves anytime performance for classification and runtime for regression.
  • Clarifies the trade-offs between different search strategies, enhancing understanding of their strengths and weaknesses.
Read more
Kairos: Numerically Robust News Recommendation under Item Cold-Start via Cholesky-based LinUCB
Finn Hertsch
Optimization Efficient ML Theory
  • Addresses item cold-start challenges in news recommendation systems.
  • Introduces a Cholesky-based approach for numerical stability in LinUCB.
  • Integrates Matryoshka Representation Learning for efficient inference.
  • Demonstrates significant efficiency gains in empirical evaluations.
Read more
TAPO: Transition-Aware Policy Optimization for LLM Agents
Cong Li, Peixi Peng, Yisen Zhao, Xinyu Hu, Shudong Liu, Zhan Su, Zhuojian Li
Reinforcement Learning Large Language Models Optimization
  • TAPO leverages action-conditioned environmental feedback to enhance policy optimization in LLM agents.
  • The framework alternates between standard policy optimization and transition supervision, improving sensitivity to environmental dynamics.
  • TAPO requires no additional expert data or sampling costs, making it a lightweight enhancement for existing RL algorithms.
  • Empirical results show consistent performance improvements across different environments and model scales.
Read more
Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection
Nicolas Gutowski, Fabien Chhel, Alexandre Letard, Sylvain Lamprier
Optimization Reinforcement Learning Theory
  • Introduction of the Top-k Pareto bandit setting and α-approximate hypervolume regret.
  • Development of the THV-UCB algorithm for greedy selection based on optimistic marginal hypervolume gains.
  • Establishment of theoretical regret bounds that are gap-free and gap-dependent.
  • Empirical evaluation demonstrating THV-UCB's superior performance in hypervolume maximization.
Read more
Context-Informed Ship Trajectory Prediction via Conditional Attention
Yuan Guan, Chandler Squires, Timothy Hu, Pradeep Ravikumar
Multimodal Time Series Robotics
  • Introduces the Conditional Informer architecture for ship trajectory prediction.
  • Utilizes a Conditional Attention mechanism to model the influence of environmental factors on vessel dynamics.
  • Implements Modality Masking to address data intermittency and prevent shortcut learning.
  • Achieves a 15.4% improvement in prediction accuracy over traditional baselines.
Read more
$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
Jiawei Xu, Minghui Liu, Juzheng Zhang, Tom Goldstein, Furong Huang
NLP Large Language Models Reinforcement Learning
  • β-OPSD generalizes OPSD by introducing a controllable parameter β for better optimization.
  • The optimal policy is derived as a geometric interpolation between a reference policy and a privileged teacher.
  • The method efficiently approximates complex RL objectives using token-level logit interpolation.
  • Experiments show significant improvements in reasoning performance and optimization stability over vanilla OPSD.
Read more
Recursive transformers for semiconductor thermo-mechanical reliability
Kart-leong Lim
Efficient ML
  • Conventional transformers are over-parameterized for small engineering datasets, leading to overfitting.
  • Three recursive transformer architectures are proposed to optimize parameter efficiency and computational complexity.
  • The models are validated on thermo-mechanical reliability analysis and Laplace PDE numerical solving tasks.
  • Recursive weight-sharing transformers provide a practical solution for resource-constrained engineering applications.
Read more
Same Graph Cross-Task Transfer in GNNs: Protocols and Predictors
Neelam Akula, Surbhi Kumar, Murat Kantarcioglu, Baris Coskunuzer
Graph Learning
  • Formalizes same graph cross-task transfer between node classification and link prediction.
  • Introduces a leakage-free evaluation protocol to ensure fair comparisons.
  • Demonstrates that NC to LP transfer is generally beneficial, while LP to NC transfer is more fragile.
  • Introduces CoTask Score (CTS) for summarizing joint task performance.
Read more
Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?
Perry Dong, Ron Polonsky, Dorsa Sadigh, Chelsea Fin
Reinforcement Learning Robotics
  • Naive pretraining of Q-functions does not significantly enhance performance in online RL fine-tuning.
  • The Q-function learned during pretraining is fundamentally different from the one learned during online fine-tuning.
  • Initialization via Policy Ensemble (IPE) improves Q-function learning by leveraging diverse policy rollouts.
  • IPE results in an average performance improvement of 26% over naive pretraining methods.
Read more
Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark
Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal
Theory Efficient ML Interpretability
  • Marginal conformal prediction fails to adequately cover minority classes, with coverage dropping below 1%.
  • Mondrian conformal prediction restores minority coverage, achieving an average improvement of 61.7 percentage points.
  • Cost-controlled abstention mechanisms can lower expected decision costs by deferring ambiguous predictions to human experts.
  • The study quantifies break-even thresholds for human review costs, establishing economic boundaries for automated decision-making.
Read more
Policy Gradient Steering: Interventions from Behavioral Objectives
Yoann Poupart, Aurélie Beynier, Nicolas Maudet
Reinforcement Learning
  • Introduction of Policy Gradient Steering (PGS) as a method for behavioral adaptation in reinforcement learning.
  • Demonstration of PGS's effectiveness in a two-route gridworld environment, chess, and competitive football.
  • PGS allows for the creation of removable and composable behavioral interventions without retraining the base policy.
  • The method shows that compatible tactical objectives can constructively accumulate.
Read more
Understanding Context Sampling in TabPFN on Small Tabular Datasets
Mohammed Abdullah
Theory Efficient ML
  • Larger contexts in TabPFN yield more stable and accurate predictions.
  • Context diversity, rather than representativeness, is crucial for accuracy.
  • Expensive selection methods do not provide significant advantages over uniform random selection.
  • Controlled experiments demonstrate the importance of diversity in context sampling.
Read more
Lottery Tickets Are Not Deployment Tickets
Bum Jun Kim
Computer Vision Theory Efficient ML
  • Lottery tickets can achieve accuracy comparable to dense models but may behave differently in deployment.
  • Replacing a dense model with a lottery ticket can lead to significant changes in downstream decision-making.
  • Clean-accuracy recovery does not guarantee compatibility with existing decision logic.
  • A behavioral-compatibility distance metric is introduced to evaluate model behavior in practical deployment scenarios.
Read more
Q-Steer: Action-Value Guidance for Molecular Policy Optimization
Xinyu Wang, Jinbo Bi, Minghu Song
Reinforcement Learning Generative Models Optimization
  • Q-Steer provides action-value guidance during molecular generation, addressing the myopic nature of traditional methods.
  • The method utilizes a frozen PAVS-Q model to estimate the potential rewards of candidate tokens based on partial SMILES prefixes.
  • Extensive experiments show consistent improvements in mean valid-unique scores across various molecular language models and optimizers.
  • The findings highlight the importance of action identity in optimization performance.
Read more
Latent States in Neural Networks: Recovering the Temporal Structure of Drifting Data from Model Weights
Kevin Guan
Time Series NLP Multimodal
  • Latent states can be recovered from the weights of neural networks trained on temporally drifting data.
  • The study employs a hidden Markov model (HMM) to analyze the weight trajectories of classifiers.
  • Classifiers show better generalization within the same latent state compared to across state boundaries.
  • The identified latent states correlate more with class distribution shifts than with weight-space geometry.
Read more
Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning
Gong Gao, Xiao Lai, Ziqi Xie, Guojie Chen, Xianhui Liu, Weidong Zhao
Reinforcement Learning Robotics Optimization
  • Introduction of CWAC framework that combines collaborative weighting and pessimistic value estimation.
  • Adaptive reweighting of TD-errors based on predictive uncertainty to improve learning stability.
  • Stochastic sampling from return distributions to reduce overestimation and error propagation.
  • Demonstrated superior performance and sample efficiency in continuous control benchmarks.
Read more
Tight Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection
Arunan J
Theory Efficient ML Large Language Models
  • Establishes tight upper and lower bounds on the sample complexity of Low-Rank Adaptation.
  • Introduces a rank-selection dichotomy, highlighting the effects of under-ranking and over-ranking.
  • Validates theoretical predictions through experiments on synthetic and real datasets.
  • Demonstrates that over-parameterization can harm performance in constrained empirical risk minimization.
Read more
LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning
Mohand Mezmaz, Grégoire Danoy
Optimization
  • LM-GRASP reformulates the constructive phase of GRASP as an online imitation learning task.
  • The framework learns from high-quality solutions without requiring external data or offline pretraining.
  • A decoder-only Transformer network serves as the constructive policy, capturing non-myopic dependencies.
  • LM-GRASP outperforms traditional methods, achieving a notable reduction in makespan for PFSP.
Read more
Learning-Augmented and Randomized Algorithms for Line Aggregation with Delays
Tianhang Lu, Runtian Ren, Shengcai Liu, Ke Tang
Theory Optimization
  • Introduction of a deterministic learning-augmented algorithm for line aggregation with delays.
  • Development of a randomized algorithm with a competitive ratio better than existing deterministic benchmarks.
  • Establishment of a new lower bound for randomized algorithms in this context.
  • Combination of learning-augmented and randomized approaches to enhance algorithm performance.
Read more
Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
Yansen Zhang, Yilu Liu, Tianyu Liu, Jiamin Chen, Xiaokun Zhang, Kai Xie, Xue Liu, Chen Ma, Yiyan Qi
Large Language Models Optimization Efficient ML
  • Cost-blind credit assignment can lead to significant quality loss in LLM discovery under budget constraints.
  • CostAda introduces a cost-calibrated frontier utility that improves resource allocation during the search process.
  • The methodology allows for dynamic adjustments in exploration intensity and frontier allocation based on remaining budget.
  • CostAda achieves superior performance with reduced budget usage compared to existing methods.
Read more
First-order Constrained Trilevel Optimization Over Distributed Networks for Robust Coreset Selection
Yang Jiao, Kaixuan Jiao, Kai Yang, Nadjib Aitsaadi, Ilhem Fajjari, Renwei (Richard) Li
Optimization Federated Learning Theory
  • Introduces a novel framework for distributed robust coreset selection, addressing a significant gap in existing literature.
  • Formulates the problem as a constrained trilevel optimization, which is more complex than traditional approaches.
  • Proposes a distributed algorithm with proven non-asymptotic convergence guarantees.
  • Demonstrates the effectiveness of the proposed method through extensive empirical evaluations.
Read more
Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
Xiang Yuan, Kaiqing Lei, Zhenyu Jin, Jun Shu, Deyu Meng, Zongben Xu
NLP Large Language Models Optimization
  • Introduces a Bayesian domain weighting method to optimize data mixtures for LLMs.
  • Addresses limitations of existing methods that rely on strong structural assumptions.
  • Demonstrates stable and efficient learning of domain weights with reduced computational costs.
  • Empirical results show improved performance over traditional function-fitting approaches.
Read more
Universality and Approximation Rates of Graph Neural Networks with Random Features
Lukas Gonon, Thilo Meyer-Brandis, Niklas Weber
Graph Learning Theory
  • PENNs with random features can approximate any permutation-invariant or permutation-equivariant function on directed graphs.
  • The paper provides a concrete architecture for universality, improving upon existing results in the literature.
  • Upper bounds on approximation rates are derived, linking network complexity to approximation accuracy.
  • The study addresses potential overfitting issues associated with random features and suggests averaging techniques to mitigate these effects.
Read more
Evaluation Protocols and Cross-Subject Generalization in EEG Emotion Recognition
Hanting Suo, Yuwen Li
Time Series
  • Introduces a three-part EEG protocol record to clarify evaluation processes.
  • Demonstrates significant performance discrepancies between subject-dependent and subject-disjoint evaluations.
  • Establishes the importance of distinguishing between different evaluation settings in EEG studies.
  • Proposes an auditable checklist for EEG reporting to enhance reproducibility.
Read more
High-Order Markov Blanket Discovery via a k-Order Relaxation of the Faithfulness Assumption
Loong Kuan Lee, Ragavi Krishnamoorthy, Nico Piatkowski
Graph Learning Theory
  • Proposes a k-order relaxation of the faithfulness assumption for Markov blanket discovery.
  • Introduces the k-order Markov blanket (kOMB) algorithm for identifying higher-order dependencies.
  • Demonstrates that kOMB can recover Markov blankets under both true and empirical violations of faithfulness.
  • Outperforms existing methods on benchmark datasets, emphasizing the need for higher-order relationship exploration.
Read more
Semi-Supervised Learning for Molecular Graphs via Ensemble Consensus
Rasmus Tirsgaard, Laurits Fredsgaard, Marisa Wodrich, Mikkel Jordahn, Mikkel N. Schmidt
Graph Learning Theory Efficient ML
  • Introduces a novel SSL method based on ensemble consensus for molecular graphs.
  • Demonstrates significant improvements in predictive accuracy across diverse datasets.
  • Individual models trained with ensemble consensus outperform traditional supervised ensembles.
  • Reduces calibration error and enhances model robustness.
Read more
Information Bottleneck Learning for Faithful Time Series Forecasting Explanations
Xu Zheng, Wei Cheng, Zhuomin Chen, Mo Sha, Jingchao Ni, Dongsheng Luo
Time Series Interpretability
  • IB-Forecast provides a self-interpretable framework for multivariate time-series forecasting.
  • The model decomposes forecasts into a learned periodic profile and a gated deviation readout.
  • It employs an information bottleneck to ensure that only relevant historical data influences predictions.
  • IB-Forecast achieves competitive accuracy while requiring only 14-20% of observations for low-error predictions.
Read more
Weak-to-Strong On-Policy Distillation
Fangxu Yu, Zinan Lin, Xiaodong Liu, Weijia Xu, Michael Xu, Tianyi Zhou, Jianfeng Gao
NLP Large Language Models Reinforcement Learning
  • W2S-OPD allows strong models to improve using weaker models as teachers.
  • The framework constructs a proxy teacher from a contrast pair of weak models.
  • W2S-OPD consistently outperforms traditional on-policy distillation methods.
  • The method enables students to surpass their domain teachers in performance.
Read more
PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform's Perspective
Shengtian Yang, Yewen Li, Peng Jiang, Zhiyi Lyu, Bo An, Qingpeng Cai, Lei Feng
Reinforcement Learning Generative Models Optimization
  • Introduction of PlatformBid as a benchmark for auto-bidding from a unified platform perspective.
  • Definition of three competitive settings: homogeneous, heterogeneous, and promotional competition.
  • Evaluation of existing auto-bidding methods alongside the introduction of a novel method, BidFlow.
  • Demonstration of BidFlow's effectiveness in dynamic environments and its practical deployment.
Read more
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR
Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes
Reinforcement Learning Large Language Models Efficient ML
  • SARA optimizes rollout allocation by leveraging early verdicts on prompt effectiveness.
  • The method reduces the number of rollouts needed by 22% compared to dynamic sampling.
  • SARA can be integrated into existing GRPO-style pipelines without additional prediction rollouts.
  • The approach is validated through empirical results on mathematical reasoning and planning tasks using large language models.
Read more
The Role of Causality in Algorithmic Recourse
Srikanth Avasarala, Varun Gupta, Shahin Jabbari, Saber Salehkaleybar, Juba Ziani
Theory Optimization Interpretability
  • Existing algorithmic recourse methods often ignore the causal impact of suggested changes, leading to strategic manipulation.
  • The proposed causal performative recourse framework models the interaction between agents and the learning algorithm using structural causal models.
  • The authors identify conditions for stable and optimal recourse solutions that can be computed efficiently.
  • Empirical results show that causal recourse reduces incentives for gaming and enhances predictive accuracy.
Read more
Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs
Jinyi Liu, Wei Chen, Pengyu Chen, Xinyi Yuan, Minghe Bai, Guoquan Wu, Jun Wei
Large Language Models Efficient ML
  • Prox leverages the SwiGLU intermediate state for effective channel selection without incurring high computational costs.
  • The framework achieves up to a 1.99× speedup in end-to-end decoding at 70% FFN sparsity.
  • Prox is compatible with existing quantization techniques and sparse attention mechanisms.
  • The method shows superior performance compared to existing training-free baselines across various models and sparsity levels.
Read more
Expanding Data-Agnostic Pivotal Instances Selection Models with Proximity Trees and Ensemble Learning
Alessio Cascione, Mattia Setzu, Cristiano Landi, Paolo Maria Mancarella, Riccardo Guidotti
Interpretability
  • Introduction of PivotTree, a hierarchical model for selecting pivotal instances.
  • The model enhances interpretability by mimicking human cognitive processes in decision-making.
  • Incorporation of proximity and oblique trees, as well as ensemble methods, to improve performance.
  • Demonstrated effectiveness across diverse data modalities, outperforming existing strategies.
Read more
The Convergence Behavior of Adam under Heavy-Tailed Noise
Yijiang Pang
Optimization Theory
  • First convergence guarantees for Adam under heavy-tailed noise.
  • Generalization of the online-to-nonconvex conversion framework to accommodate heavy-tailed noise.
  • Establishment of a complete discounted regret analysis for the vector-form Adam update.
  • Demonstration of suboptimal iteration complexity that persists even in bounded-variance cases.
Read more
DHRCL: Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning
Shuhang Wang, Ziming Li, Hui Cheng
Large Language Models Reinforcement Learning Optimization
  • Introduction of a three-stage hierarchical reward curriculum that coordinates multiple feedback types.
  • Automatic stage progression based on validation trends, eliminating the need for fixed thresholds.
  • Implementation of stage-aware token credit redistribution to optimize learning efficiency.
  • Demonstrated consistent performance improvements across different model sizes.
Read more
ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
Xingjian Wu, Xuhang Zhu, Xingchen Liu, Junlin Liu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai
NLP Large Language Models Robotics
  • ClawTrack introduces a dual-assessment benchmark that evaluates both task outcomes and reasoning processes.
  • The Process Grader assesses reasoning quality across four dimensions, enabling detailed attribution of successes and failures.
  • Result verification is identified as a critical bottleneck in the reasoning process.
  • The framework is robust to evaluator choice, demonstrating consistent improvements in model performance.
Read more
Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing
Brandon Gower-Winter, Georg Krempl
Theory
  • Introduction of Outcome Performativity A/B Detection (OPAB) for detecting Outcome Performativity.
  • Derivation of sample complexity bounds for OPAB under different assumption classes.
  • Empirical validation of OPAB's effectiveness in detecting Outcome Performativity.
  • Identification of regions of indistinguishability where detection may be challenging.
Read more
Hierarchical Spatio-Temporal Transformer for Coherent Emergency Department Forecasting
Filipa Lino, Bárbara Tavares, Carlos Santiago, Cláudia Soares, Manuel Marques
Time Series
  • HierSTT is a novel hierarchical Transformer framework for multi-level ED demand forecasting.
  • The model incorporates a coherence-aware loss function to ensure consistency across hierarchical levels.
  • A new nationwide Portuguese ED dataset is introduced, reflecting real-world data heterogeneity.
  • HierSTT outperforms traditional statistical methods and non-hierarchical deep learning approaches.
Read more
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts
Ken Ding
Reinforcement Learning Large Language Models Optimization
  • LSPO recovers lost gradients on cliff prompts in RLVR by using a low-rank adapter during training.
  • The method integrates supervised learning with reinforcement learning without modifying the loss function.
  • Empirical results show LSPO outperforms DAPO on all tested benchmarks, with significant improvements in performance metrics.
  • The approach effectively converts a substantial portion of cliff prompts into usable learning signals.
Read more
Regularizing modality contribution drift in multimodal continual learning
Zhen Zhang, Jielei Chu, Bin Liu, Tianrui Li
Multimodal
  • Identification and formalization of modality contribution drift (MCD) in MMCL.
  • Introduction of Continual Modality Contribution Drift Regularization (CMCDR) to preserve modality contributions.
  • CMCDR supports both replay-based and replay-free learning settings.
  • Extensive experiments validate the effectiveness of CMCDR across various benchmarks.
Read more
Train Small, Deploy Large: Zero-Shot GNN Transfer Through Geometric Renormalization
Robert Jankowski, Pedro Almagro-Blanco, Marián Boguñá, Melanie Weber, M. Ángeles Serrano
Graph Learning
  • Introduction of a zero-shot transfer protocol for GNNs using geometric renormalization.
  • Empirical evidence that GR maintains GNN performance under significant graph compression.
  • Compatibility of node representations learned at different GR scales.
  • Preservation of functional training trajectories across scales.
Read more
Event-Structured Physics-Informed Neural Networks for Differentiable Critical Clearing Boundaries
Baoli Hao, Chenxi Hu, Ming Zhong, Ren Wang
Theory Optimization Time Series
  • Introduction of ES-PINN for transient stability assessment in power systems.
  • Differentiable approximation of critical clearing time (CCT) boundary.
  • Improved accuracy in trajectory and stability-boundary estimation over traditional methods.
  • Local sensitivity analysis and error estimation linked to physics residuals.
Read more
Single-Beat Cuffless Blood Pressure Estimation Using Ear-PPG and ECG with a Lightweight Hybrid Learning Framework
Kindeep K. Dhatt, Tengyue Wu, Hanbang Hua, Yayun Du
Multimodal Time Series Efficient ML
  • Introduction of a lightweight multi-modal wearable system for continuous BP estimation.
  • Demonstration that BP-related information can be effectively extracted from single PPG beats.
  • Development of a hybrid learning architecture combining CNN and LightGBM for efficient BP estimation.
  • Achieved significant reduction in mean absolute error compared to traditional models.
Read more
What Makes Graph Unified? Principles and Generative Sliding-Window Transformer for Graph Foundation Models
Dongxiao He, Siqi Liu, Jitao Zhao, Yawen Li, Yi Wang, Di Jin
Graph Learning
  • Introduces Graph Foundation Models (GFMs) for unified graph learning across diverse domains.
  • Identifies the limitations of existing feature transformation methods in preserving semantics.
  • Proposes four principles for effective cross-domain graph feature unification.
  • Develops SliGFM, which employs topology-aware sliding-window encoding for feature transformation.
Read more
Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search
Ayushman Singh, Siddharth Aphale
Reinforcement Learning Theory Optimization
  • Bilinear contrastive critics can rank actions accurately within a support set but may fail in broader candidate selection.
  • Norm drift and value decalibration are significant issues that lead to misordering of actions.
  • Experiments reveal that contrastive critics incur regret during maximization despite strong retrieval performance.
  • A value-calibrated scalar is essential for ensuring reliable action selection and restoring proper value ordering.
Read more