AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

24 Papers today
8h Update frequency
7 Days of history
Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization
Jiayi Dan, Bo Li, Lu Deng, Yong Wang
Theory
  • Introduces a new doubly robust estimator for causal effect on CVR, addressing sample selection bias.
  • Derives theoretical properties of the estimator using semiparametric theory.
  • Implements targeted regularization to improve numerical stability and practical applicability.
  • Demonstrates superior performance through extensive experiments on various datasets.
Read more
Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples
Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia
Efficient ML
  • UltraIR is a foundation model for IR spectroscopy with over 100 million parameters.
  • It employs simulation-to-real transfer learning to enhance chemical inference from IR spectra.
  • The model is pretrained on simulated spectra and adapted to specific tasks with limited labeled data.
  • UltraIR outperforms traditional methods in various chemical sensing applications.
Read more
Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling
Takieddine Soualhi, Jacques Saraydaryan, Laetitia Matignon
Reinforcement Learning Robotics
  • Introduction of a proxemics-based reward model for socially compliant robot navigation.
  • Validation of the model across multiple DRL navigation methods, showing improved social metrics.
  • Systematic analysis of reward design factors for effective social navigation.
  • Focus on the importance of comfort-aware navigation assessment in addition to traditional navigation metrics.
Read more
MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Ning Ding, Xia Hu, Bowen Zhou, Chaochao Lu, Youbang Sun
NLP Large Language Models Efficient ML
  • MARCH introduces a scalable memory architecture for recurrent models that allows for efficient long-context retrieval.
  • The architecture utilizes content-conditioned state anchors to maintain historical information without increasing computational complexity.
  • MARCH consistently outperforms existing linear attention models in various benchmarks, demonstrating its effectiveness in recall-intensive tasks.
  • The method provides a controllable trade-off between historical resolution and memory cost.
Read more
Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection
Iyad Assaad Nekka, Hamida Seba, Khaled Walid Hidouci, Karima Amrouche
Graph Learning Interpretability Time Series
  • Introduction of X-AddGraph, the first post-hoc explainability framework for AddGraph and GCN+GRU models.
  • Development of the Dual Spatial-Temporal Attribution (DSTA) mechanism with three aligned components.
  • Preservation of detection performance (βˆ†AUC = 0) while providing explanations.
  • Long-term attribution identifies more relevant historical snapshots compared to random selection.
Read more
Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness
Haochen Zhang, Jiaheng Guo, Yu-Chao Huang, Nicholas Knoz, Tianlong Chen
Generative Models Time Series Multimodal
  • ReCoGen effectively addresses the issue of irregular missingness in multimodal physiological data.
  • The framework separates the representation of conditions from the generation of target signals, enhancing performance.
  • ReCoGen outperforms existing conditional generators across multiple datasets and tasks.
  • The use of masked autoencoders allows for robust and compact representations of time-variant conditions.
Read more
Adaptive $k$ Nearest Neighbors Classifier via Granular Ball Computing
Xiaoyu Lian, Shuyin Xia, Hongxuan He, Lifeng Shen, Guoyin Wang, Xinbo Gao
Efficient ML Theory
  • Introduces an adaptive KNN approach using granular-ball computing.
  • Utilizes the Fisher criterion for effective granular ball generation.
  • Dynamically determines the effective k value based on local neighborhood structure.
  • Demonstrates improved robustness against noise and local perturbations.
Read more
Training AI Scientists to Replicate Research
Damon Falck, Samer Sabri, Anja Surina, Thom Foster, Anya Sims, Sam Devlin, Dylan Rogers, Tantum Collins, Kaloyan Aleksiev, Louis Kirsch, Edward Hughes
Large Language Models Reinforcement Learning Theory
  • Introduction of Replica, a task space for replicating research papers.
  • Development of Faraday, an AI Scientist that outperforms existing models in replication tasks.
  • Implementation of a rubric-based judging system for evaluating replication quality.
  • Demonstration of Faraday's scientific rigor and creativity in handling underspecified research problems.
Read more
Incremental Evaluation and Training in Relational Deep Learning
Jakub PeleΕ‘ka, Gustav Ε Γ­r
Graph Learning Time Series Theory
  • Introduces an incremental evaluation paradigm for RDL that captures the temporal dynamics of relational databases.
  • Demonstrates the prevalence of temporal concept drift in predictive tasks within RDL.
  • Presents multiple effective incremental training strategies that outperform traditional training methods.
  • Proposes a new evaluation metric focused on near-future predictive accuracy.
Read more
The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use
Joyjeet Singh
Reinforcement Learning Optimization Robotics
  • The predictor is not the bottleneck; imagination remains informative beyond the planner's usage.
  • The planning objective saturates and can lead to counterproductive outcomes.
  • The pathology of planning issues is inherent to the method, not specific to a single implementation.
  • Reachability is a more effective planning objective than proximity.
Read more
Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration
Sabin Roman, Ljupco Todorovski, Saso Dzeroski
Theory Optimization Time Series
  • SORT provides a sparse spectral framework for learning from noisy and irregular data.
  • The technique allows for the discovery of ordinary differential equations using orthonormal basis expansions.
  • SORT shows improved stability and performance compared to traditional sparse regression methods under challenging conditions.
  • The method supports nonlinear approximation and high-dimensional integral estimation.
Read more
Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology
Yunsung Chung, Yingshuo Liu, Abboud F. Hassan, Han Feng, Mary M. Maleckar, Nassir Marrouche, Jihun Hamm
Multimodal
  • Introduces an intervention-aware latent clinical world model for dynamic post-operative outcome forecasting.
  • Utilizes a structured latent state that evolves based on irregular post-procedural events.
  • Achieves AUROC of 0.756 and AUPRC of 0.777 for recurrence prediction in atrial fibrillation ablation.
  • Provides retrospective risk estimates at various horizons, enhancing clinical decision-making.
Read more
A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family
Rishi Shah, Rishav Shrestha
Large Language Models Theory Optimization
  • Introduction of a twelve-gate contract-grade verifier for GPU kernels.
  • Audit of 2,638 machine-generated kernels reveals significant correctness issues.
  • Development of the first native Blackwell tcgen05 training backward for the GDN family.
  • The verifier successfully identifies failures in kernels accepted by existing benchmarks.
Read more
EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction
Danyu Li, Ling Zhou, Rubing Huang, Xian Zhong, Bin Zou, Kui Jiang
Graph Learning
  • EGRL introduces implicit meta-path learning to enhance relational semantics in GNNs.
  • A graph generator is included to predict interactions for cold-start RNA/protein nodes.
  • The framework employs a multi-relation-aware attention mechanism for better interaction modeling.
  • EGRL shows significant improvements in AUROC and AUPR metrics over existing methods.
Read more
Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data
Francesca Pia Panaccione, Sofia Mongardi, Marco Masseroli, Pietro Pinoli
Generative Models Graph Learning
  • Synthetic data generation can mitigate challenges of data scarcity and biases in genomic research.
  • MK-TGAN, a novel generative model, effectively integrates biological knowledge through graph neural networks.
  • Incorporating prior biological knowledge significantly enhances the realism and utility of synthetic transcriptomic data.
  • The study demonstrates state-of-the-art performance of MK-TGAN in generating biologically coherent samples.
Read more
The Impact of Temporal Context Length and Encoding Strategies on Self-Supervised ECG Representation Learning
Ahmed Sameh, Ramzi Al-Sharawi, Yogatheesan Varatharajah
Time Series
  • Longer temporal contexts (5-10 minutes) improve ECG representation learning and downstream classification accuracy.
  • Continuous convolutional patch embeddings outperform discretized vector-quantized tokens in capturing clinically relevant details.
  • Short temporal windows may not provide sufficient information for accurate rhythm diagnosis, particularly for conditions like atrial fibrillation.
Read more
HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models
Fangzhou Chen, Shiji Zhao, Mengyang Wang, Qihui Zhu, Ranjie Duan, Maoxun Yuan, Xingxing Wei
NLP Large Language Models Efficient ML
  • HiRoute establishes a cross-category safety boundary using a shared coarse-grained prompt.
  • It dynamically composes fine-grained prompt experts through a hierarchical router based on input risk.
  • The framework effectively reduces over-refusal while maintaining high safety and helpfulness in responses.
  • HiRoute's performance is validated across various instruction-tuned models and safety benchmarks.
Read more
Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency
Guo An, Zijing Wu, Honghua Dong, Yuhao Yan, Zixuan Gui, Haochong Chen, Shanzhao Ruan, Xiang Wang, Yurong Ling, Qi Tian
Reinforcement Learning Robotics Theory
  • Introduction of Action-Conditioned Predictive Consistency (ACPC) for diagnosing visual perturbations in JEPAs.
  • Establishment of Invariance Radius (IR) and Separation Rate (SR) as metrics for assessing model robustness.
  • Proven bounds on the impact of visual perturbations on multi-step prediction error and planning costs.
  • Empirical validation of ACPC across multiple control tasks and architectures, showing consistent diagnostic trends.
Read more
Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia
Xiang Guan, Roger D. Newman-Norlund, Yong Yang, Saeed Ahmadi, Regan Willis, Nadra Salman, Kalil Warren, Srihari Nelakuditi, Chris Rorden, Leonardo Bonilha, Julius Fridriksson
NLP Large Language Models Interpretability
  • PRISM adapts neuroimaging subtraction analysis for interpreting large language models.
  • The framework allows for spatially resolved testing of cognitive operation specialization in LLMs.
  • Error profiles from a perturbed LLM align with lesion patterns in aphasia patients, demonstrating the framework's validity.
  • Robust phonemic-favoring dissociations were identified in both LLM and patient data.
Read more
Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization
Jinhyung Bae
Optimization
  • First measurement of allocation value for NCO test-time sampling, showing no detectable gain in in-distribution workloads.
  • Significant gains (11-12%) observed under distribution shift with proper allocation strategies.
  • Quantification of in-sample selection bias, demonstrating that conventional methods can produce misleading results.
  • Introduction of a correction procedure that preserves real gains while eliminating phantom gains.
Read more
Into the ORBIT for Time Series: Training Regimes for Foundation Models
Hongjie Xia, Yiding Liu, Yifan Hu, Peiyuan Liu, Zewei Dong
Time Series
  • Introduction of ORBIT, a training paradigm for TSFMs that controls pre-training distribution.
  • Bootstrap Multi-Level Sampling allows for hierarchical control of data exposure.
  • Omni-Range Incremental Training enables the model to handle variable-length examples efficiently.
  • Falcon-2.0 demonstrates strong zero-shot forecasting capabilities across diverse datasets.
Read more
Knowledge-guided Pattern Discovery via Coupled Tensor Factorizations
Gaute Johannessen, Geert Roelof van der Ploeg, Evrim Acar
Time Series Optimization Interpretability
  • Introduces a knowledge-guided approach for pattern discovery that combines real and simulated data.
  • Utilizes coupled tensor factorizations to analyze multiway datasets effectively.
  • Demonstrates improved pattern discovery performance in metabolomics data.
  • Reveals potential discrepancies between real data and computational models.
Read more
Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws
Liu Ziyin, Yizhou Xu, Tomaso Poggio, Isaac Chuang
Theory Optimization
  • Introduces Neural Quadratic Forms (NQF) as a minimal model for understanding neural network learning dynamics.
  • Demonstrates that diverse neural architectures share common behaviors due to underlying symmetry.
  • Derives a universal quadratic form for the cost function expansion around initial weights.
  • Establishes a connection between training dynamics and Lotka–Volterra equations.
Read more
Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks
Farhang Yeganegi, Arian Eamaz, Mojtaba Soltanalian
Theory Optimization
  • Introduces an executable-certificate framework for assessing neural network training.
  • Defines a challenge-power modulus to characterize optimality gaps.
  • Demonstrates the framework's effectiveness on a ResNet-18 distillation problem.
  • Highlights the need for coverage mechanisms in validating training outcomes.
Read more