AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

24 Papers today
8h Update frequency
7 Days of history
Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not
Rubén Darío Guerrero
Optimization Theory
  • Introduction of Stiefel Attention, optimizing transformer projection matrices on the Stiefel manifold.
  • Development of a Riemannian Adam optimizer that respects the geometry of the manifold.
  • Significant performance improvements on benchmarks, with validation accuracy rising from 61.1% to 97.0%.
  • The method's advantages grow with increasing data, indicating a strong relationship between geometry and optimization.
Read more
Enhanced Agriculture-informed Neural Network by Domain Knowledge
Ci Lin, Futong Li, Rose Chong-Wu, Tet Yeap, Iluju Kiringa
Interpretability
  • KAINN integrates domain knowledge into deep learning for improved N2O emission predictions.
  • The framework incorporates key environmental processes to enhance model interpretability.
  • KAINN outperforms traditional neural network models and the original AINN in accuracy.
  • The approach demonstrates reduced uncertainty and improved stability in parameter trajectories.
Read more
PhyRestore: Physics-Structured Latent-Factor Restoration
Ahmed Shafee, Chayan Lahiri
Time Series Theory Optimization
  • PhyRestore restores corrupted physical factors from bitemporal spatial observations.
  • The framework reconstructs temporal changes through the known RUSLE relationship.
  • Factor restoration improves recovery of rare high-magnitude changes under certain conditions.
  • The study compares multiple learning pathways for soil-loss prediction.
Read more
Machine-Learning Assessment of the Predictive Value of Inflammatory Biomarkers for Cognitive Impairment in an Older Hispanic Adult Cohort
Antony Garcia, Gabrielle Britton, Alcibiades Villarreal, Diana Oviedo, Giselle Rangel, Xinming Huang
Interpretability
  • I-309 (CCL1) was identified as a significant predictor of cognitive impairment, with a notable increase in predictive accuracy.
  • The study employed a leakage-safe machine learning methodology to ensure robust results.
  • The findings underscore the need for external validation of biomarkers in predicting cognitive impairment.
  • The research distinguishes between statistical significance and predictive utility in the context of small clinical datasets.
Read more
Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems
Mohammed El Hanjri, Anas Abouaomar, Hamidou Tembine, Abdellatif Kobbane
Federated Learning Time Series
  • Introduces a coalition formation approach for federated learning that adapts to client heterogeneity.
  • Utilizes opinion dynamics to form coalitions based on local model weights, enhancing model compatibility.
  • Demonstrates significant improvements in forecasting accuracy for water consumption compared to traditional methods.
  • Achieves coalition structures with no additional client-side computation or communication overhead.
Read more
Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning
Haoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang, Yian Ma, Lianhui Qin
Multimodal Robotics NLP
  • Uni-LaDiR introduces a unified latent space for multimodal reasoning, improving the integration of information from different modalities.
  • The framework employs a unified encoder to generate shared thought tokens, enhancing the reasoning process by reducing modality-specific representation issues.
  • Diffusion is used to predict the next reasoning steps, allowing for flexibility in generating multiple valid outputs.
  • Joint training of the encoder and diffusion model leads to better task-relevant thought token generation.
Read more
Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification
Zhilong Zheng, Letian Tao, Yang Guan, Yujie Yang, Wei Xiong, Kehua Sheng, Bo Zhang, Jingliang Duan, Keqiang Li, Shengbo Eben Li
Theory Efficient ML Large Language Models
  • Introduces Parameter Space Orthogonality as a feasible condition for preventing catastrophic forgetting.
  • Develops a post-hoc, tuning-agnostic framework (JANUS) that rectifies parameter updates after fine-tuning.
  • Employs a Multi-step Adaptive Rectification mechanism to dynamically adjust parameter updates.
  • Demonstrates significant improvements in knowledge recovery with minimal impact on task performance.
Read more
SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption
Wentao Zhang, Yifan Zhu, Yutong Zhang, Wentao Mo
Multimodal
  • Identifies a granularity mismatch in existing batch-level gradient modulation under heterogeneous multimodal corruption.
  • Proposes SAGG, which utilizes sample-level gating for unbiased gradient estimation.
  • Demonstrates that SAGG-based SGD converges at a rate of O(1/√T) without residual bias from corruption.
  • Provides a certified robustness framework for evaluating multimodal models.
Read more
CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning
Xiang Zou, Shengzhu Shi, Junqi Gao, Zhichang Guo
Reinforcement Learning Robotics Theory
  • Introduction of CARE-VI, a framework for improving target reliability in off-policy actor-critic learning.
  • Development of three components: CARS, SEVA, and DARE, which address candidate selection, value assessment, and residual regulation.
  • Theoretical analysis provides error bounds for the proposed methods, ensuring fixed-policy recovery.
  • Empirical results show CARE-VI achieves the highest mean return across multiple tasks and configurations.
Read more
Personalized Federated Hierarchical Gaussian Processes for Privacy-Preserving Modeling of Heterogeneous Distributed Systems
Xianjian Xie, Hao Yan
Federated Learning
  • Introduces a hierarchical decomposition for modeling heterogeneous federated clients.
  • Employs privacy-preserving federated variational inference to keep raw data local.
  • Supports uncertainty-aware predictions, crucial for risk-sensitive applications.
  • Demonstrates effectiveness in real-world applications like fault classification and air-quality modeling.
Read more
A Policy Profile for Croissant: Refusal as a Property of the Dataset
Alexander Chernov
Theory
  • Introduces an additive policy profile for Croissant with specified evaluation semantics.
  • Defines a closed operator language for dataset operations and conditions.
  • Demonstrates that the profile's decisions align with existing descriptor records.
  • Presents a minimal evaluation cost for implementing the profile.
Read more
COMPASS: Ordered Clustered Routing at 100K Scale
Ido Greenberg, Hugo Linsenmaier, Piotr Sielski, Shie Mannor, Alex Fender, Gal Chechik, Eli Meirom
Optimization
  • COMPASS is a globally-coordinated algorithm that optimizes the OCTSP using raw distance inputs.
  • The algorithm can scale to 100K synthetic nodes and 28.5K real-world e-commerce nodes, achieving state-of-the-art results.
  • COMPASS avoids a quality ceiling by modeling global dependencies between clusters.
  • The paper introduces a new benchmark suite for large-scale OCTSP, facilitating future research.
Read more
Learning-Induced Dynamical Transition in Recurrent Neural Networks
Varun Vaidya
Theory
  • Learning in RNNs can induce a transition from chaotic to stable dynamics.
  • A non-equilibrium dynamical mean-field theory (DMFT) is developed to analyze this transition.
  • The study identifies critical parameters that separate chaotic and stable regimes during learning.
  • The theory predicts the evolution of network outputs and shows quantitative agreement with simulations.
Read more
ZeroHAT: Behavior-Conditioned Zero-Shot Human Activity Trace Generation
Rongchao Xu, Dahai Yu, Lin Jiang, Guang Wang
Generative Models Time Series Optimization
  • ZeroHAT generates synthetic HATs without requiring target-region data, using only source-region data and contextual information.
  • The framework includes innovative components for intent extraction, behavioral cloning, and activity realization.
  • ZeroHAT significantly outperforms existing methods in terms of utility and fidelity across multiple target regions.
  • The approach demonstrates the effectiveness of transferring behavioral patterns across regions with different POI distributions.
Read more
Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization
Donney Fan, Colin Doumont, Aleksandra Kalisz, Paul Duckworth, Jacob R. Gardner, Henry Moss, Geoff Pleiss
Generative Models Optimization Efficient ML
  • Introduces a new LSBO algorithm that significantly reduces computational overhead.
  • Utilizes linear surrogates constrained to a spherical domain for efficient optimization.
  • Achieves over 100× speedup compared to existing BO methods while maintaining high sample efficiency.
  • Demonstrates effectiveness across molecular design and image generation tasks.
Read more
Online Adaptive Kernel Mixing for Gaussian Process Decision Making
Kavin Aravindan, Mani Tej Sriram, Gautam Dasarathy, Tejas Bodas
Optimization Theory
  • HACK GPs adaptively select kernels in Gaussian Processes to improve decision-making performance.
  • The method treats kernel selection as an online learning problem using AdaHedge for dynamic updates.
  • Two variants of HACK are introduced: Mixture of Gaussians and categorical sampling.
  • The approach shows improved performance over standard kernels and ensemble methods in empirical evaluations.
Read more
Sharp Reconstruction Bounds for Autoencoders Using the Same Forward Map
Patricia Medina, Hy P. G. Lam
Theory
  • Establishes sharp reconstruction bounds for autoencoders using the same forward map.
  • Derives a reconstruction-derivative error bound based on Jacobian singular values.
  • Demonstrates that affine maps can achieve the derived bounds at any depth.
  • Validates theoretical predictions with empirical results from a large LiDAR dataset.
Read more
QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization
Yujie Li, Zezhi Shao, Chengqing Yu, Yisong Fu, Weijie Zhu, Yifan Du, Jilin Hu, Bin Yang, Yongjun Xu, Fei Wang
Time Series Efficient ML Optimization
  • QUALS addresses data diversity issues in time series forecasting.
  • The framework includes pattern quantization and learnability synchronization mechanisms.
  • Models trained on QUALS achieve superior zero-shot performance with less data.
  • The approach effectively manages skewed pattern distributions and learnability discrepancies.
Read more
How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU?
Nguyen Dung Son, Dang Quang Minh, Nguyen Huu Loi, Truong Viet Vu, Nguyen Thai Anh
NLP Large Language Models
  • Rubric-decomposed prompting outperforms holistic prompting in most cases.
  • The mapping of trait scores to prompt scores is fragile and requires careful calibration.
  • Essay length impacts scoring accuracy, with smaller models showing variability in performance.
  • The best local model configuration achieved a QWK of 0.388, below human scoring standards.
Read more
When Does Retrieval Help Time-Series Forecasting?
Mert Onur Cakiroglu, Elham Buxton, Mehmet Dalkilic, Hasan Kurban
Time Series
  • Retrieval benefits in time-series forecasting are primarily determined by the ratio of lookback window length to dominant seasonal period.
  • A simple control method can outperform complex retrieval plug-ins under certain conditions, particularly when the seasonal structure is strong.
  • The study introduces a regime map and two statistics to predict the effectiveness of retrieval mechanisms before deployment.
  • Retrieval mechanisms are less effective when the training data lacks a concentrated seasonal structure.
Read more
Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation
Jing Zhang
Reinforcement Learning Robotics Theory
  • RSIQL improves offline GCRL by providing additional reward signals at intermediate states.
  • The method does not require a hierarchical policy, simplifying the learning process.
  • Experiments show significant performance improvements over existing methods.
  • The approach addresses the issue of delayed goal-completion supervision effectively.
Read more
SoftTri: Smooth Triangular Membership Functions for Adaptive Fuzzy Inference Systems
Babak Sarani, Rahman Ardakanian, Ali Mousavi
Optimization Interpretability Theory
  • Introduction of SoftTri, a differentiable triangular membership function that enhances optimization stability.
  • Closed-form analytical gradients derived for efficient backpropagation in training.
  • SoftTri maintains the interpretability and locality of classical triangular MFs while providing smoothness.
  • Experimental results show improved performance over classical triangular MFs and comparable results to Gaussian MFs.
Read more
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
Mobina Kashaniyan, Ali Jannesari
Large Language Models Efficient ML NLP
  • Candidate count alone does not adequately describe the system costs of multi-candidate inference.
  • Increasing the number of candidates improves accuracy but also increases energy consumption and latency.
  • Batched generation calls are significantly more efficient than serial execution of candidates.
  • The paper provides recommendations for reporting practices in multi-candidate inference studies.
Read more
Subdomain-aware representation compression for pretrained image embeddings
Poowanut Niamluang, Jittat Fakcharoenphol
Computer Vision Efficient ML
  • Dimensionality reduction techniques can be effectively tailored for subdomain representation compression.
  • PCA and LDA show significant improvements in both space efficiency and accuracy for pretrained image embeddings.
  • The approach allows for effective transfer learning capabilities using compressed representations.
  • Experiments demonstrate that compression can outperform traditional full-embedding methods.
Read more