AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

24 Papers today
8h Update frequency
7 Days of history
DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization
Tao Zhang, Jianchao Tan, Pingwei Sun, Yanqi Yu, Zixu Jiang, Yuchen Xie, Xunliang Cai, Ziqian Zeng
NLP Large Language Models Efficient ML
  • DAMP is the first to explore post-training quantization of recurrent states in GDN and KDA models.
  • Uniform quantization leads to significant accuracy loss, particularly with lower bit formats.
  • DAMP utilizes quantization-error energy and decay-based persistence to optimize channel precision.
  • The method achieves a 69.1% reduction in recurrent-state storage and accelerates updates by up to 2.01Γ—.
Read more
REPLICANT: Learning Policies for Evading and Hardening Malware Detectors
Shae McFadden, Ilias Tsingenopoulos, Mario D'Onghia, Alexander Herzog, Myles Foley, Chris Hicks, Lorenzo Cavallaro, Fabio Pierazzi
Reinforcement Learning
  • REPLICANT achieves a mean attack success rate of 78.8%, outperforming existing methods.
  • The framework facilitates policy transfer, yielding an 82.0% relative increase in ASR.
  • Adversarial training with REPLICANT significantly reduces attacker success rates.
  • The study conducts the largest evaluation of Android malware evasion to date, with 379,680 attack evaluations.
Read more
Generalized Context in Cross Attention for Transfer Learning of Disjoint Tabular Data
Kazi F. Akhter, Ibna Kowsar, Manar D. Samad
Theory Efficient ML
  • Introduces CATTLE, a method for transfer learning in disjoint tabular datasets without requiring shared features.
  • Utilizes transformer projection weights to capture generalized context for cross-domain learning.
  • Demonstrates superior performance over nine state-of-the-art methods in terms of AUROC and ranking.
  • Addresses the unique challenges posed by heterogeneous feature types in tabular data.
Read more
An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models
Javier Aguilar MartΓ­n
Theory Robotics Reinforcement Learning
  • Acceptance-with-certainty by a sampling gate certifies models only up to the reachable query set, leaving errors beyond as gauge.
  • The study identifies three regimes of model behavior based on the angular channel width, impacting the model's exploitability and falsifiability.
  • Repair mechanisms are limited by the geometry of the omitted structure and cannot recover unreachable regions from outside evidence.
  • Mitigation strategies must match the dimensionality and direction of the model's errors to be effective.
Read more
A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation Constraint
Artem Safronov
Large Language Models Efficient ML Optimization
  • Introduces a selective mixed-precision quantization method based on layer sensitivity profiles.
  • Demonstrates significant performance improvements in LLM inference with minimal quality degradation.
  • Identifies optimal configurations for different model subsystems (FFN, Attention, lm_head).
  • Proposes additional optimizations to further accelerate model execution.
Read more
Initialization Is Critical: Advancing Federated Short-Term Load Forecasting under Load Heterogeneity via Model Initialization
Jianing Chen, Vajiheh Farhadi, Yan Li, Thomas La Porta
Federated Learning Time Series
  • Identifies structured load heterogeneity among clients affecting STLF performance.
  • Proposes two innovative model initialization strategies to mitigate client drift.
  • Demonstrates compatibility of proposed strategies with existing FL frameworks.
  • Experimental results show significant improvements in forecasting accuracy and convergence.
Read more
VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning
Pengcheng Li, Zhengyang Zhang, Dongxu Zhang, Sui Huang, Shaohua Ma
Reinforcement Learning Large Language Models
  • VICT introduces a new training-time interface for credit assignment in RL using verifier structures.
  • The method shifts credit assignment from rollout-side inference to verifier-side tracing.
  • VICT preserves original terminal rewards while enhancing credit assignment accuracy.
  • Experiments show substantial performance improvements on long-horizon tasks.
Read more
Temporal Memory-Aware Online Test-Time Adaptation on Dynamic Graphs
Bo Li, Xin Zheng, Ming Jin, Can Wang, Shirui Pan
Graph Learning Time Series Optimization
  • Introduces DGOTTA, a framework for online test-time adaptation of DGNNs on dynamic graphs.
  • Addresses the challenges of temporal and structural distribution shifts in dynamic graph data.
  • Incorporates three innovative modules: temporal-aware augmentation, memory-aware prediction, and consistency-guided adaptation.
  • Demonstrates significant improvements in generalization performance across diverse datasets and DGNN architectures.
Read more
There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation
Gabe Guo, Elon Litman, Thanawat Sornwanee, Jose Blanchet, Stefano Ermon
Generative Models Multimodal
  • BIT provides a bidirectional generative framework for multimodal translation, allowing for both T2I and I2T processes.
  • The model starts from text and interpolates to images, enhancing the flexibility of sampling algorithms.
  • BIT is derived from stochastic calculus, yielding a simulation-friendly SDE with tractable loss functions.
  • Empirical results show BIT's competitive performance against existing models in various tasks, including scientific applications.
Read more
Performative Privacy: When Differential Privacy Maximizes Utility
Uddalak Mukherjee, Edwige Cyffers, Yann Chevaleyre
Theory
  • Introduces the concept of performative privacy, linking privacy protection to long-term utility.
  • Demonstrates that a finite privacy budget can outperform non-private estimation in specific scenarios.
  • Establishes a feedback loop between data leakage and user participation, affecting future data contributions.
  • Extends the analysis to a d-dimensional average with realistic data leakage definitions.
Read more
More Data Cannot Break a Symmetry: Identifiability by Design
Jing Xu, Christopher Kanan
Theory Optimization
  • Identifiability in unsupervised alignment is constrained by the automorphism group of the stimulus geometry.
  • A design-time diagnostic can identify structural failures in experimental designs, particularly in color representation.
  • Selecting stimulus colors based on this diagnostic can reduce catastrophic alignment failures from 75% to 2%.
  • Model discrimination and correspondence recovery are largely uncorrelated objectives.
Read more
FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling
Jun Bai, Ruilin Wang, Yue Li
Federated Learning
  • Introduces FedEHR-Agents, a federated optimization framework for EHR modeling.
  • Shifts focus from model parameters to clinical modeling experience for improved collaboration.
  • Demonstrates superior performance over local and federated baselines in clinical prediction tasks.
  • Maintains patient data locality while facilitating knowledge transfer across hospitals.
Read more
Actionable CBFI: Integrating Structural Decomposition and Causal Counterfactual Recourse for Tabular Machine Learning
Sejong Oh
Interpretability
  • A-CBFI integrates structural causal models to improve actionable counterfactual recourse.
  • The framework significantly reduces the cognitive burden of interventions by focusing on root causes.
  • Empirical results demonstrate a 76.9% reduction in active intervention effort while maintaining cost-effectiveness.
  • A-CBFI addresses the limitations of existing methods by prioritizing causal bottlenecks over diffuse modifications.
Read more
Leveraging a Foundation Model for the EEG-Based Diagnosis of Alzheimer's Disease
Maggie Lin, Chung-Lin Hou, Tzyy-Ping Jung
Time Series
  • Introduces a framework combining LaBraM and Random Forest for AD diagnosis.
  • Achieves high diagnostic accuracy with only 8-second EEG segments.
  • Surpasses traditional spectral analysis methods in performance.
  • Identifies clinically relevant biomarkers linked to cognitive decline.
Read more
Exact Risk Ratios for Weighted Data Selection in Linear Regression
Guangjian Zhang
Theory Optimization
  • Determines exact values for the worst-case risk ratios in weighted data selection for linear regression.
  • Proves Fw(d, 2d - 1) = 1 + 1/d, Fw(3, 4) = 5/3, and Fw(4, 5) = 2.
  • Establishes a lower bound for Fw(d, d + k) using harmonic quantities.
  • Conjectures that the exact minimax value holds throughout the open regime.
Read more
Euclidean Fourier Neural Operators
Nathanael Bosch, Niklas Frederik Schmitz, Michael F. Herbst
Theory Efficient ML
  • EFNOs provide a domain-independent alternative to FNOs by parameterizing the spectral kernel as a continuous function of physical wavevectors.
  • EFNOs can learn operators that generalize across different periodic domains, addressing the limitations of FNOs in domain transfer.
  • The methodology was validated through experiments on a heat equation and materials science tasks, showing superior performance in generalization.
  • EFNOs maintain consistent performance across varying grid sizes and shapes, unlike FNOs, whose performance degrades with domain changes.
Read more
Beyond Flat Netlist: Hierarchical Graph Representation Learning for Scalable Analysis of Sequential Circuits
Jingyi Zhou, Zhengyuan Shi, Jiaying Zhu, Ziyang Zheng, Qiang Xu
Graph Learning
  • DeepSeq3 introduces a hierarchical representation for sequential netlists, enhancing scalability and representation richness.
  • The framework utilizes two types of graphs: combinational logic subgraphs and super-node graphs, processed hierarchically.
  • A state-centric pre-training task is introduced to improve the model's understanding of register-level semantics.
  • DeepSeq3 achieves an 18% reduction in BMC solving time on large-scale benchmarks, demonstrating its practical effectiveness.
Read more
A Deeper Analysis of Block-Sparse Featurizers
Alexandru-Iulius Jerpelea, Amith Ananthram
Computer Vision Theory Interpretability
  • The BSF introduces a block-based approach to feature representation, improving upon traditional sparse autoencoders.
  • Architectural modifications, such as the Tournament Top-K selection rule, effectively reduce feature splitting.
  • The BSF demonstrates superior performance in recovering features from low-dimensional manifolds compared to classical SAEs.
  • The study highlights the importance of seed stability in model training, showing that BSFs maintain consistency across different initializations.
Read more
Efficient Online Continual Foundation Model Fine-Tuning for Predictive Process Monitoring
Sjoerd van Straten, Marwan Hassani
Efficient ML
  • Introduces COMPASS, the first framework for online continual fine-tuning of foundation models in PPM.
  • Adapts loss-plateau drift detection for autonomous task boundary identification in event streams.
  • Combines residual subspace projection with pre-trained knowledge anchoring to achieve stability and plasticity.
  • Demonstrates superior performance over SOTA non-FM competitors in various drift scenarios.
Read more
Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining
Shuchen Zhu, Yuxin Fang, Mingze Wang, Kun Yuan
Large Language Models Optimization Efficient ML
  • Introduces a novel optimization method for LLM pretraining that enhances training dynamics along flat directions.
  • Combines slow-decay and fast-decay momentum components to balance noise reduction and curvature adaptation.
  • Employs sphere constraints to prevent parameter inflation and maintain stability during training.
  • Demonstrates significant performance improvements over existing optimizers like Muon across various architectures and model sizes.
Read more
Learning to Difference: Adaptive Reversible Differencing (AdaRDiff) for Time Series Forecasting
Morad Laglil, Younes Hlal, Marouane El Hadari, Emilie Devijver, Eric Gaussier
Time Series
  • AdaRDiff is a plug-and-play module that learns differencing weights adaptively, improving forecasting accuracy.
  • The method captures trends and seasonalities jointly and allows for autoregressive reconstruction of forecasts.
  • AdaRDiff shows significant speed improvements on GPU, achieving up to 33.7Γ— speedup over naive recurrence.
  • The two-phase training schedule enhances the learning dynamics by separating structure discovery from reconstruction.
Read more
Beyond Search-Imitation: Prior-Directed Exploration for Searchless Chess
Szymon MiΕ‚osz, Piotr Duch, Szymon Grabowski
Reinforcement Learning Theory
  • Introduces prior-directed exploration to enhance searchless chess networks.
  • Replaces traditional entropy bonuses with a forward KL divergence towards MCTS priors.
  • Achieves significant improvements in puzzle accuracy and tactical performance.
  • Demonstrates a dissociation between tactical accuracy and overall playing strength.
Read more
Explainable Uncertainty Estimation for Reliable Medical AI
Li Rong Wang, Jamie Duell, Xinran Xu, Thomas C. Henderson, Yu Yue Hew, Pik Wan Erica Chiang, Xiao Wei Alstar Ang, Bingwen Eugene Fan, Xiuyi Fan
Interpretability
  • Introduces egRUE, a method that combines uncertainty estimation with explainability.
  • Demonstrates that egRUE provides feature-level insights into prediction uncertainties.
  • Validates the method through theoretical analysis and experiments on real-world datasets.
  • User studies indicate that egRUE improves trust in AI predictions among medical professionals.
Read more
When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging
Shangge Liu, Yuehan Yin, Yinghuan Shi, Lei Wang, Wenbin Li
Optimization Theory
  • Unification of catastrophic forgetting in CL and weight-disentanglement error in MM as instances of task interference.
  • Derivation of an upper bound on task interference that isolates the spectral norm of weight updates as an optimizer-controllable factor.
  • Identification of the Muon optimizer as a mechanism that effectively regulates task interference through spectral norm control.
  • Empirical validation showing significant accuracy improvements when using Muon over AdamW in various benchmarks.
Read more