AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

24 Papers today
8h Update frequency
7 Days of history
Learning-Based Reconstruction of Optical Properties in Bilayered Media from Single-distance Time-Resolved Reflectance Measurements
Caterina Amendola, Giulia Maffeis, Lorenzo Buffoni, Lorenzo Chicchi, Francesco Coghi, Duccio Fanelli, Raffaele Marino, Fabrizio Martelli, Riccardo Paoli, Lorenzo Pattelli, Lorenzo Spinelli
Theory Optimization Efficient ML
  • Introduces a machine learning framework for reconstructing optical properties in bilayered media.
  • Demonstrates significant improvements in accuracy and speed over traditional model-based methods.
  • Utilizes a robust synthetic dataset generated from Monte Carlo simulations.
  • Estimates parameter space dimensionality without requiring prior information on layer count.
Read more
Parallelism, critical windows, and separations among diffusion language models
Sitan Chen, Liye Wang
NLP Large Language Models Theory
  • Uniform and Gaussian diffusion can achieve sampling efficiency scaling with dual total correlation.
  • Masked diffusion requires significantly more forward passes than uniform and Gaussian diffusion for certain empirical measures.
  • The critical windows in masked diffusion are narrower, impacting its parallelism capabilities.
  • This work provides the first provable separations in parallelism among leading dLLM paradigms.
Read more
One Intervention per Component is Enough: Towards Identifiability in Linear Stochastic Dynamics from Steady State
Saber Salehkaleybar
Theory Time Series Optimization
  • One intervention per strongly connected component is sufficient for parameter recovery in OU processes.
  • A recursive algorithm is developed to order SCCs and isolate marginal dynamics for parameter estimation.
  • The paper introduces a regularized least-squares estimator for joint minimization of residuals from observational and interventional data.
  • Theoretical results are validated through empirical experiments, confirming the effectiveness of the proposed methods.
Read more
COMPASS: Ordered Clustered Routing at 100K Scale
Ido Greenberg, Hugo Linsenmaier, Piotr Sielski, Shie Mannor, Alex Fender, Gal Chechik, Eli Meirom
Optimization
  • COMPASS is a parallel algorithm that optimizes the Ordered Clustered Traveling Salesman Problem (OCTSP) by considering global dependencies between clusters.
  • The algorithm can handle general distance matrices, expanding its applicability beyond traditional coordinate-based methods.
  • COMPASS achieved state-of-the-art performance on benchmarks with up to 100K synthetic nodes and 28.5K real e-commerce nodes.
  • The paper introduces a new benchmark suite for large-scale OCTSP, facilitating further research in the area.
Read more
Personalized Federated Hierarchical Gaussian Processes for Privacy-Preserving Modeling of Heterogeneous Distributed Systems
Xianjian Xie, Hao Yan
Federated Learning
  • Introduces a hierarchical decomposition for modeling heterogeneous federated clients.
  • Utilizes privacy-preserving federated variational inference to maintain data locality.
  • Supports uncertainty-aware predictions across multiple correlated sensors.
  • Achieves high accuracy in practical applications with minimal labeled data.
Read more
Subdomain-aware representation compression for pretrained image embeddings
Poowanut Niamluang, Jittat Fakcharoenphol
Computer Vision Efficient ML
  • Dimensionality reduction techniques can enhance the quality of subdomain representations.
  • PCA and LDA are effective methods for compressing pretrained image embeddings.
  • The proposed approach improves both space efficiency and downstream task performance.
  • Transfer learning capabilities of compressed representations are demonstrated.
Read more
Explaining spatial information flow in short-term traffic forecasting models using a gated graph attention network
Yue Li, Shujuan Chen, Ying Jin
Graph Learning Time Series Interpretability
  • Introduces a gated mechanism to enhance explainability in traffic forecasting models.
  • Demonstrates that spatial information flow varies with traffic conditions.
  • Finds that both GAT layers in the architecture are largely redundant.
  • Regularization of the gating mechanism can improve model accuracy.
Read more
Generative Query Suggestion via Intent Coverage and Query-Level Credit Assignment
Xinpeng Liu, Lu Ma, Jiayi Qiao, Mengyu Zhou, Linglong Li, Xiaofeng Bian, Haonan Chen, Xiaoxi Jiang, Guanjun Jiang
NLP Large Language Models Reinforcement Learning
  • Introduces a dual-stage optimization framework for generative query suggestion.
  • Utilizes intent-aware diversity modeling to enhance intent coverage.
  • Implements query-level credit assignment for improved individual query quality.
  • Demonstrates significant performance improvements through extensive testing.
Read more
Layer-wise Curriculum Learning for Efficient LLM Compression
Donggeon Lee, Dooyeon Na, Seungmin Oh, Jongbin Ryu
Large Language Models Efficient ML Optimization
  • Introduces layer-wise curriculum learning for efficient LLM compression.
  • The method accelerates convergence and stabilizes knowledge transfer.
  • Achieves over 50% reduction in GPU memory usage and training hours.
  • Outperforms existing pruning methods on multiple language models.
Read more
CoRe: Coherence and Relational Alignment for Multivariate Time Series Forecasting
Xiaoyu Lin, Huiran Duan, Yining Liu, Zhixiang Wu, Chu Lin, Lin Lu
Time Series
  • CoRe addresses the gap in direct MTSF objectives by incorporating coherence and relational alignment in the output space.
  • The proposed method introduces no additional trainable parameters, making it easy to integrate into existing forecasting models.
  • CoRe shows significant improvements over traditional pointwise loss metrics and competitive forecasting objectives.
  • The methodology enhances both temporal coherence and cross-variable relational consistency in predictions.
Read more
Fast-varying Natural Frequencies and Damping Ratio Identification for Linear Time-Varying System
Melisa Bozaci, Alice Cicirello
Time Series
  • Introduces a hybrid EKF-LSTM approach for identifying fast-varying natural frequencies and damping ratios in LTV systems.
  • Validates the method using synthetic data from a 2-blade offshore wind turbine under realistic operational conditions.
  • Achieves high accuracy in identifying natural frequencies with minimal error, demonstrating robustness to incorrect damping assumptions.
  • Extends the methodology to effectively estimate damping ratios, improving upon existing identification techniques.
Read more
Robust Federated Q-Learning with Almost No Communication
Sreejeet Maity, Aritra Mitra
Reinforcement Learning Federated Learning Theory
  • Introduction of Robust Fed-Q, a federated Q-learning algorithm that is resilient to adversarial agents.
  • Guarantees convergence to the optimal value function despite the presence of adversarial agents.
  • Achieves near-optimal finite-time performance with minimal communication overhead.
  • Combines model-based and model-free approaches with robust statistical techniques.
Read more
Evaluating Financial Sentiment in the Age of AI
Arslan Bisharat, Oudom Hean
NLP Large Language Models
  • General-purpose LLMs achieve classification performance comparable to finance-specific models without fine-tuning.
  • Higher classification accuracy does not correlate with stronger economic relationships, particularly regarding stock returns.
  • Sentiment models are more effective for large earnings surprises than for moderate ones.
  • The study introduces a framework distinguishing between linguistic and economic validity in sentiment analysis.
Read more
A Learning Algorithm for Threshold Boolean Networks with Prescribed Fixed Points
Gonzalo A. Ruz
Theory Graph Learning Optimization
  • Introduction of a learning algorithm for inferring TBNs with fixed points.
  • Custom loss function designed to enforce fixed point preservation and reduce spurious attractors.
  • Achieved high accuracy in reconstructing the FOS-GRN model of Arabidopsis thaliana.
  • Demonstrated stability of the algorithm's performance across varying sparsity coefficients.
Read more
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang, Michael Liu, Thomas Creavin, Kurt Keutzer, Xiuyu Li, Zhaoyang Lv, Chenfeng Xu, Haiwen Feng
Computer Vision Generative Models Multimodal
  • Introduction of Video DeltaNet (VDN), a hybrid attention architecture for video generation.
  • Combines local Softmax attention with bidirectional linear memory to enhance efficiency.
  • Achieves a 14.5× speedup in denoising time compared to the Dense H3 baseline.
  • Maintains high video quality metrics while reducing computational costs.
Read more
Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification
Zhilong Zheng, Letian Tao, Yang Guan, Yujie Yang, Wei Xiong, Kehua Sheng, Bo Zhang, Jingliang Duan, Keqiang Li, Shengbo Eben Li
Theory Efficient ML Large Language Models
  • Introduction of Parameter Space Orthogonality as a feasible condition for preventing catastrophic forgetting.
  • Development of the JANUS framework for post-hoc weight rectification that integrates with any fine-tuning method.
  • Implementation of a Multi-step Adaptive Rectification mechanism to navigate the parameter space effectively.
  • Demonstration of temporal and spatial efficiency through ghost operations and SVD compression.
Read more
Towards a Unified Modality-Agnostic Multimodal Framework for Cognitive Workload Assessment
Stefanos Gkikas, Christian Arzate Cruz, Calvin Joseph, Giorgos Giannakakis, Raul Fernandez Rojas
Multimodal
  • Introduces a modality-agnostic framework for cognitive workload assessment using a hierarchical Transformer architecture.
  • Systematically evaluates all combinations of five biosignal modalities in a pilot study.
  • EEG is identified as the strongest single modality for cognitive workload assessment.
  • Combining all five modalities yields the highest performance, but not all combinations improve results.
Read more
Sharpness-Aware Minimization (SAM) Improves Classification Accuracy of Bacterial Raman Spectral Data Enabling Portable Diagnostics
Kaitlin Zareno, Jarett Dewbury, Siamak K. Sorooshyari, Hossein Mobahi, Loza F. Tadesse
Optimization
  • SAM significantly improves classification accuracy of Raman spectral data.
  • Achieved up to 10.5% accuracy improvement on single splits and 2.7% on average across all splits.
  • Addresses challenges of generalization and data quality in Raman spectroscopy diagnostics.
  • Demonstrates the effectiveness of deep learning without extensive pre-processing.
Read more
Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization
Donney Fan, Colin Doumont, Aleksandra Kalisz, Paul Duckworth, Jacob R. Gardner, Henry Moss, Geoff Pleiss
Generative Models Optimization Efficient ML
  • Introduces a nearly-instant LSBO algorithm that significantly reduces computational overhead.
  • Utilizes linear surrogates constrained to a spherical domain to improve efficiency.
  • Achieves over 100× speedup compared to traditional BO methods while maintaining performance.
  • Demonstrates effectiveness on molecular design and image generation tasks in high-dimensional latent spaces.
Read more
Smart Insole Human Activity Recognition for Continuous Monitoring in Elderly Care
Edwin Rios, Antony Garcia, Fengpei Yuan, Xinming Huang
Multimodal Time Series Robotics
  • Development of a smart insole platform for continuous monitoring of elderly activity states.
  • Integration of pressure sensors and IMU for comprehensive activity recognition.
  • High performance of Histogram-Based Gradient Boosting in classifying activities.
  • Demonstration of the potential for unobtrusive fall risk monitoring.
Read more
Improving Generalization and Robustness in Offline Reinforcement Learning via Boundary-Aware Data Augmentation
Gong Gao, Weidong Zhao, Xianhui Liu
Reinforcement Learning
  • Introduces Boundary-Aware Data Augmentation (BADA) to improve generalization and robustness in ORL.
  • Theoretical analysis shows that interpolation error is correlated with state distance, guiding the augmentation approach.
  • BADA generates synthetic data that maintains local smoothness and original distribution fidelity.
  • Extensive experiments demonstrate state-of-the-art performance across diverse benchmarks and improved robustness against adversarial attacks.
Read more
QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization
Yujie Li, Zezhi Shao, Chengqing Yu, Yisong Fu, Weijie Zhu, Yifan Du, Jilin Hu, Bin Yang, Yongjun Xu, Fei Wang
Time Series
  • QUALS improves data efficiency for time series forecasting by addressing data diversity and distribution issues.
  • The framework includes a pattern quantization mechanism for decoding heterogeneous patterns.
  • A learnability synchronization mechanism calibrates sampling weights to optimize training efficiency.
  • Models pre-trained on QUALS demonstrate superior zero-shot performance with less training data.
Read more
SoftTri: Smooth Triangular Membership Functions for Adaptive Fuzzy Inference Systems
Babak Sarani, Rahman Ardakanian, Ali Mousavi
Optimization Interpretability Theory
  • SoftTri is a differentiable triangular membership function that preserves the interpretability and locality of classical triangular MFs.
  • The function introduces a sharpness parameter that controls the smoothness of transitions, allowing for fully differentiable optimization.
  • Closed-form analytical gradients are derived, enabling efficient training without the need for subgradient heuristics.
  • SoftTri demonstrates improved optimization stability and approximation performance compared to classical triangular MFs.
Read more
Beyond Flattened Tokens: Structure-Preserving EEG Decoding with Reusable TriDim Blocks
Shiyue Su, Song Wang, Zekai Zhan, Junjie Zeng, Ziling Lu, Zongsheng Li, Xinyuan Ye, Zhiyuan Ma, Xinke Shen, Quanying Liu
Time Series
  • TriDim preserves EEG data structure by maintaining three explicit axes for spatial and temporal information.
  • TriDimEEG outperforms fifteen evaluated models with a 4.3% relative improvement in average accuracy.
  • Replacing Transformer blocks with TriDim blocks yields an average relative improvement of 7.4% in downstream accuracy.
  • The approach reduces parameter counts by 17.0% to 47.3%, enhancing efficiency.
Read more