AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

23 Papers today
8h Update frequency
7 Days of history
Learning to Rank Tensor Network Contraction Plans for GPU-Accelerated Quantum Circuit Simulation
Alfred M. Pastor, Maribel Castillo, Jose M. Badia
Optimization Efficient ML Theory
  • Introduces a learning-to-rank framework for selecting efficient tensor-network contraction plans.
  • Demonstrates that contraction plan performance varies significantly on GPUs despite similar theoretical complexities.
  • Achieves high accuracy in identifying optimal plans, with the listwise model outperforming other strategies.
  • Shows that learned models retain useful decision quality across different GPU architectures.
Read more
QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding
Ayushman Garg, Akshita Gupta, Shaswata Bhattacharya, Abhishek Gupta, Sandeep Kumar, Manoj Kumar
NLP Large Language Models Efficient ML
  • QEvict introduces a recoverable eviction strategy for KV caches, addressing the limitations of traditional irreversible eviction methods.
  • The system categorizes token windows into three tiers, allowing for dynamic updates and recovery of important tokens.
  • QEvict improves attention retention and reduces missed information during long-context decoding tasks.
  • The methodology is supported by diagnostics that measure future attention and historical importance of cached states.
Read more
Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control
Xinwei Liu, Junyuan Liang, Jianting Zhang, Wuhui Chen
Reinforcement Learning Computer Vision Robotics
  • OG-SPR combines latent self-prediction and observation prediction to improve sample efficiency in visual RL.
  • The algorithm employs lightweight adapters to mitigate over-constraining of the shared representation.
  • Experimental results show OG-SPR achieves superior performance compared to existing self-predictive and observation-predictive methods.
  • The approach is validated across 28 tasks in the DeepMind Control Suite, highlighting its effectiveness in challenging domains.
Read more
Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction
Quinn Ledingham, Zhengsen Xu, Yimin Zhu, Zack Dewis, Mabel Heffring, Saeid Taleghanidoozdoozan, Motasem Alkayid, Megan Greenwood, Lincoln Linlin Xu
Interpretability
  • TabPFN outperforms other models in predicting post-wildfire debris flows with a threat score of 0.637.
  • Short-duration rainfall intensity and storm accumulation are identified as the most critical features for prediction.
  • Synthetic data augmentation significantly improves model performance, especially for deep learning approaches.
  • The study provides a systematic evaluation of 15 machine learning models, addressing gaps in previous research.
Read more
Accelerating nanodrug development in continuous flow systems using informed prediction models based on low-cost surrogate nanoparticles
Kai Dahms, Eilien Heinrich, Jochen Schmid, Michael Bortz, Iryna Savych, Regina Bleul
Optimization
  • Introduces a predictive modeling approach based on shape constraints for nanoparticle development.
  • Utilizes controlled microfluidic methods to prepare liposomes and lipid nanoparticles.
  • Validates the model with minimal empirical data, demonstrating its effectiveness.
  • Reduces the need for extensive experimental workflows in nanodrug development.
Read more
Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations
Guillaume Couairon, Alexis Jacq, Yu-Han Wu, Renu Singh, Yana Hasson, Quentin Berthet, Romuald Elie
Generative Models Efficient ML Time Series
  • Kastor improves the accuracy and efficiency of generative emulation for PDE simulations.
  • The two-stage inference scheme reduces error accumulation in long-term predictions.
  • Mean Prediction Regularization enhances model stability and performance.
  • Spatial gradient matching increases the physical fidelity of simulations.
Read more
IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games
Conor M. Artman, Nicholas Di, Scott Perkins
Reinforcement Learning Generative Models Theory
  • IFlowNets generalize AFlowNets to handle incomplete information games effectively.
  • The paper proves that existing constraints for generative flow networks are inadmissible in incomplete information contexts.
  • IFlowNets preserve the expected flow-matching property, crucial for achieving generalized Nash equilibria.
  • Preliminary results indicate that IFlowNets outperform or match the performance of traditional methods in standard game environments.
Read more
Potential Matching Optimal Transport: Continuous Normalizing Flows for Exact $p$-Wasserstein Dynamics
Lishuo Zhang, Ruizhi Huang, Yang Yu, Lei Li
Generative Models Optimization Theory
  • Introduction of PMOT, a potential-flow CNF framework for general p-cost optimal transport.
  • Establishment of zero-loss exactness, linking PMOT solutions to the generalized Benamou–Brenier optimality system.
  • Demonstration of p-specific geometric alignment and competitive likelihood modeling through empirical evaluations.
  • Flexible terminal matching capabilities using KL/NLL and MMD objectives.
Read more
Hybrid Probabilistic Zonotopes for Identifiable and Refinable Predictive Uncertainty
Zhen Zhang, Amr Alanwar
Theory Time Series Robotics
  • Introduction of Hybrid Probabilistic Zonotope (HProbZ) for better uncertainty representation in predictions.
  • HProbZ allows for identifiable decomposition of uncertainty into discrete, bounded, and stochastic components.
  • The model enables observation-driven refinement of predictions in a single forward pass.
  • Empirical results indicate HProbZ outperforms traditional mixture models in trajectory prediction tasks.
Read more
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?
Omar Coser, Antonio Orvieto, Paolo Soda, Loredana Zollo
Time Series Multimodal Robotics
  • SPT enhances transformer performance on medical time series tasks.
  • Improvements in classification accuracy range from 0-6 percentage points.
  • Deeper models benefit more from SPT due to enriched temporal representations.
  • SPT does not require task-specific architectural changes.
Read more
The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity
Iosif Lytras, Nikolaos Makras, Sotirios Sabanis
Optimization Large Language Models Theory
  • Introduction of SG-TULA for sampling from non-convex, non-smooth distributions.
  • Non-asymptotic convergence bounds in Wasserstein-2 distance with explicit constants.
  • Excess risk estimates for optimization problems associated with the sampling algorithm.
  • SG-TULA shows competitive performance in pretraining LLMs against traditional methods.
Read more
Rectifying Geometric Misalignment: Online Source-Free Adaptation for Class-Imbalanced EEG
Shiwen Chu, Shanglin Li, Motoaki Kawanabe, Reinmar Kobler
Time Series
  • Introduces OSPDIM, a novel online SFUDA framework for EEG data.
  • Addresses the issue of label shifts and class imbalance in BCI applications.
  • Implements real-time geometric bias correction using manifold optimization.
  • Demonstrates superior performance over traditional Riemannian alignment methods.
Read more
GAUGE: Granularity-Adaptive Counterfactual Gating of Evidence for Incomplete Multimodal Classification
Yunping Shi, En Yu, Kairui Guo, Jie Lu
Multimodal
  • GAUGE provides a fine-grained evidence control mechanism for incomplete multimodal classification.
  • The framework utilizes a frozen imputer and encodes inputs into fine-grained evidence units for better reliability.
  • Counterfactual effects are evaluated using Taylor evidence scores, allowing efficient computation in a single pass.
  • GAUGE outperforms existing methods across multiple benchmarks with incomplete modalities.
Read more
BaKron: Efficient Quantization with Kronecker-Factored Hessians
Johann Birnick, Rayan Saab
Efficient ML Optimization Theory
  • BaKron introduces a more efficient quantization algorithm using Kronecker-factored Hessians.
  • The algorithm reduces computational complexity significantly while capturing richer geometric information.
  • BaKron is modular, allowing for flexibility in the choice of quantizer and Hessian estimator.
  • Empirical evaluations show that BaKron outperforms existing quantization methods in terms of speed and efficiency.
Read more
An Optimal Agnostic PAC Algorithm
Markus Engelund Mathiasen, Jian Qian, Nikita Zhivotovskiy
Theory
  • The proposed learner achieves the optimal risk bound for agnostic PAC learning.
  • The construction uses a novel edge isoperimetric inequality to control approximation errors.
  • The results match the lower bounds established by previous works, confirming the optimality of the approach.
  • The methodology includes suffix averaging with variance control and a derandomization step.
Read more
KV-Skill: Forging Expertise in the Model's Native Language
Zhaowei Han, Xiang Zhang, Bing Han, Kai Liu, Danqi Hu, Jie Liu
NLP Large Language Models Optimization
  • KV-Skill introduces an external factorized operator design for task knowledge in language models.
  • The framework supports both registration of text skills and reward learning from task outcomes.
  • KV-Skill consistently outperforms traditional methods across multiple benchmarks.
  • The approach allows for independent loading and swapping of task knowledge without affecting model performance.
Read more
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, Kai Jia
Large Language Models Reinforcement Learning Optimization
  • CalibForge utilizes adversarial solver calibration to improve task synthesis for terminal agents.
  • Two calibration strategies (multi-solver and contrastive) are proposed to define a solver-relative learnable zone.
  • The system generated 5,431 calibrated tasks, leading to significant performance improvements on benchmark tasks.
  • Models trained on calibrated tasks achieved up to 47.57% accuracy on Terminal-Bench 2.0, outperforming baseline models.
Read more
ProDVI: Programmatic Dynamics Priors for Value Network Initialization
Xinwei Liu, Junyuan Liang, Jianting Zhang, Wuhui Chen
Reinforcement Learning Large Language Models Efficient ML
  • ProDVI utilizes large language models to generate dynamics priors for RL agent initialization.
  • The framework does not require pre-collected datasets or high-fidelity simulators.
  • Synthetic transitions generated from LLM outputs are used for pretraining the value network.
  • ProDVI improves sample efficiency in model-free RL tasks significantly.
Read more
SkillTFM: Gated Skill Evolution for Training-Free Adaptation of Tabular Foundation Models
Yi He, Zhengkang Guan, Anpeng Wu, Peng Cui, Fei Wu, Kun Kuang
Time Series Optimization Efficient ML
  • SkillTFM is the first skill-based adaptation system for training-free tabular foundation models.
  • It employs a gated skill-evolution mechanism that allows for selective repairs and safe fallback options.
  • The system demonstrates significant improvements in AUC and MAE across various boundary settings and real-world applications.
  • SkillTFM's learned skill state is transferable across different TFM backbones and optimizer configurations.
Read more
MS-MLB: An Open Machine Learning Benchmark for Blood-Based MS Classification
Adam Simson, Ankush Dutta, Quang Bui
Theory
  • MS-MLB is the first open benchmark for classifying MS versus healthy controls using whole blood RNA expression data.
  • The benchmark utilizes the GSE17048 dataset and implements a reproducible evaluation framework to minimize biases.
  • Gradient Boosting was identified as the top-performing algorithm with a high MS Research Score and AUC-ROC.
  • The framework allows for external model submissions, enhancing collaborative research efforts.
Read more
Beyond Rotations: AuroOFT for Expressive Quantized Orthogonal Fine-Tuning
Yue Han, Dianlin Wang, Xinkang Li, Jie Zhang, Ziyi Chen, Tao Wang, Yexin Cui, Weihong Han
NLP Large Language Models Efficient ML
  • AuroOFT enhances QOFT by adding a nonlinear residual branch, improving expressivity without sacrificing stability.
  • The method retains orthogonality in the QOFT branch while allowing for input-dependent corrections through the nonlinear branch.
  • AuroOFT shows significant performance improvements over both matched QOFT and QLoRA in various benchmarks.
  • The approach effectively reduces the number of trainable parameters while maintaining high performance.
Read more
Why the Third Axis Is Freedom
Michael Timothy Bennett
Theory Generative Models Optimization
  • Introduces 'freedom' as a key measure in generative modeling, surpassing traditional generative expressivity.
  • Demonstrates that models with greater freedom are more likely to generalize effectively.
  • Empirical results show that Explorative Modeling (XM) optimizes for freedom, leading to better performance.
  • Critiques existing measures of generative expressivity for their limitations in ranking model performance.
Read more
DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink Accounting
Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta
Federated Learning Efficient ML Optimization
  • DG-FedReuse allows for the reuse of cached updates in federated learning, enhancing communication efficiency.
  • The method employs a stochastic proxy-gradient discrepancy to determine when to use cached updates versus fresh updates.
  • Significant uplink savings were achieved in experiments, although accuracy differences were minimal.
  • The study provides a comprehensive audit of the method's performance and limitations.
Read more