AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

67 Papers today
8h Update frequency
7 Days of history
Beyond Simulated Benchmarks: Evaluating Motion Representations for Fall Detection Under Real-World Data Scarcity
Timilehin B. Aderinola, Ilaria D'Ascanio, Luca Palmerini, Lorenzo Chiari, Jochen Klenk, Clemens Becker, Brian Caulfield, Georgiana Ifrim
Time Series
  • Real-world fall data is extremely scarce, making it challenging to train robust fall detection models.
  • Simulated datasets often lead to high performance in controlled settings but fail to generalize to real-world applications.
  • Interval-based representations achieve the best real-world performance, while symbolic representations with impact descriptors show resilience under data scarcity.
  • The study emphasizes the importance of representation choice for effective fall detection in real-world scenarios.
Read more
Active-Trace Complexity Bounds for Moreau--Yosida Unadjusted Langevin Sampling
Yuchen Xin, Zhihua Zhang
Theory Optimization Efficient ML
  • Introduces the concept of reference active trace (Bref) to analyze MYULA discretization error.
  • Establishes a complexity bound for the number of iterations needed to achieve desired accuracy.
  • Demonstrates that the active trace can be independent of the smoothing parameter λ for structured penalties.
  • Provides a Moreau-bias bound that helps in choosing the optimal smoothing parameter.
Read more
Training Under Challenge: Executable Certificates and Challenge-Closed Optimality for Neural Networks
Farhang Yeganegi, Arian Eamaz, Mojtaba Soltanalian
Theory Optimization
  • Introduces an executable-certificate framework for neural network training assurance.
  • Establishes a method to construct complete alternative models for objective reevaluation.
  • Defines a challenge-power modulus to characterize optimality gaps.
  • Demonstrates the framework's effectiveness through empirical results on ResNet-18.
Read more
RelShap: Relationally Consistent Shapley Explanations
Seungeun Lee, Joao Fonseca, Julia Stoyanovich
Interpretability
  • RelShap integrates relational database constraints into Shapley value computations.
  • The framework restricts background data and coalition evaluations to valid relational configurations.
  • RelShap outperforms existing methods like Kernel SHAP and Conditional SHAP in identifying dominant features.
  • The method is estimator-agnostic and can be combined with existing Shapley value methods.
Read more
The Impact of Temporal Context Length and Encoding Strategies on Self-Supervised ECG Representation Learning
Ahmed Sameh, Ramzi Al-Sharawi, Yogatheesan Varatharajah
Time Series
  • Longer temporal contexts (5-10 minutes) improve ECG representation learning and downstream classification accuracy.
  • Continuous convolutional patch embeddings outperform discretized vector-quantized tokens, preserving critical waveform details.
  • The study emphasizes the need for ECG models that capture slow-varying rhythm dynamics and individual-specific structures.
Read more
The Boolean Power of ReLU
Pablo Barceló, Floris Geerts, Matthias Lanzinger, Klara Pakhomenko, Jan Van den Bussche
Graph Learning Theory
  • ReLU-MPLang is strictly more expressive than TrReLU-MPLang for Boolean queries on graphs with Boolean features.
  • The study resolves an open problem regarding the comparative power of ReLU and truncated ReLU in Boolean contexts.
  • The findings emphasize the role of activation functions in the expressiveness of graph neural networks.
  • A specific Boolean ReLU query is demonstrated to be undefinable in Σ-MPLang for any collection of eventually constant activation functions.
Read more
Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity
Debanjan Dutta, Anish Chakrabarty, Swagatam Das
Theory Efficient ML
  • First CoT realizations of graph traversal algorithms (DFS and Dijkstra) using hard-attention Transformers.
  • Explicit small-depth CoT constructions for Strahler number and width of trees, extending previous work on binary trees to arbitrary n-ary trees.
  • Demonstrated that Dijkstra's algorithm can outperform comparison-based implementations through constant-time attention operations.
  • Established a bijection between trees and Dyck paths, allowing for independent CoT constructions for both representations.
Read more
The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity
Martin J. Wainwright
Theory Generative Models Efficient ML
  • Introduction of unmasking growth complexity (UGC) as a measure of data geometry in masking diffusion.
  • UGC increments control KL discretization error, enabling a unified analysis of unmasking schemes.
  • Development of certified-optimal samplers that achieve specified KL error with high probability.
  • Connection of UGC to classical multivariate dependence measures and previous complexity results.
Read more
Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure
Mingyuan Zhang
Theory
  • The Jaccard score's loss matrices are nonsingular with affine dimension 2s - 1.
  • Exact calibration of the Jaccard score requires exponentially many prediction coordinates.
  • Two polynomial-dimensional approximation guarantees are established, allowing for practical surrogate construction.
  • A new transfer from F1 to Jaccard provides a polynomial-time rule with controlled regret.
Read more
Multi-perspective Imbalance-Conscious 6G Beamforming Optimization and Performance
Chukwunonso Henry Nwokoye, Blessing Oluchi Iloka, Chikwue V. Umeugoji, Christopher Anene Egemba, Nnenna D. Duroha
Optimization
  • Network features are more predictive for 6G beamforming optimization than device, environmental, and vision features.
  • Unsupervised clustering methods reveal that deployment environment and device type significantly influence clustering outcomes.
  • The study emphasizes the need for adaptive beamforming techniques in dynamic 6G-IoT environments.
  • Future work will explore deep and reinforcement learning for optimizing performance indicators like throughput and latency.
Read more
The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use
Joyjeet Singh
Reinforcement Learning Optimization Robotics
  • The predictor's capacity is not the limiting factor in long-horizon planning.
  • The planning objective can saturate and lead to suboptimal performance.
  • Long-horizon success is inversely correlated with one-step prediction accuracy.
  • A simple change in the planning objective can significantly enhance performance.
Read more
Distillation of Foundation Models for Time-dependent PDEs
Daniel Musekamp, Boshra Ariguib, Andrei Manolache, Mathias Niepert
Efficient ML Theory Time Series
  • Introduction of TREX, a novel knowledge distillation framework for PDEs.
  • Demonstration of effective transfer of knowledge from large foundation models to compact student models.
  • Significant reduction in model parameters and inference time while maintaining or exceeding accuracy.
  • First systematic study of distilling general pretrained PDE foundation models into smaller models.
Read more
Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport
Bohan Zhang, Anqi Ni, Yixin Wang, Paramveer S. Dhillon
NLP Large Language Models Efficient ML
  • Weightless Fine-Tuning (WFT) is a training-free method that approximates the effects of supervised fine-tuning (SFT) without modifying model weights.
  • WFT uses a cross-prefix transport operator to compute supervised residuals and transport them to current prompts, effectively capturing the distributional effects of SFT.
  • Experimental results show that WFT outperforms other lightweight personalization methods and achieves competitive performance compared to SFT on various tasks.
  • WFT requires significantly less computational resources, achieving comparable results to SFT with less than 7% of the effective computation.
Read more
Vero: Can AI Agents Build Formally Verified Software Repositories?
Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song, Zhengxu Yan, Timothe Kasriel, Qingyang Zhang, Kaiyu Yang, Soonho Kong, Jingxuan He, Dawn Song
Theory
  • Vero is the first benchmark for evaluating joint implementation and proof synthesis at the repository level.
  • The benchmark includes 43 multi-module instances from diverse domains and programming languages.
  • An audit mechanism is implemented to identify and correct errors in specifications and reference implementations.
  • Current AI agents struggle with complex specifications, solving only 27 out of 43 instances in the best case.
Read more
Federated Compositional Muon Optimizer for Matrix-Wise Models
Wang Yan, Feihu Huang
Federated Learning Optimization Theory
  • Introduction of FedCoMuon and FedCoMuon-VR optimizers for distributed matrix-wise compositional optimization.
  • Theoretical convergence analysis under non-convex and non-i.i.d. settings, with FedCoMuon-VR showing lower sample complexity.
  • Empirical validation demonstrating competitive performance in robust federated learning and task-distributed meta learning.
Read more
When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation
Shuhan Wang, Yilin Luo, Nan Xu, Chi Wang Cheung
NLP Large Language Models Optimization
  • The only commuting orthogonal maps for distinct RoPE frequencies are independent pairwise rotations.
  • The derived rotation angle minimizes channel variance but does not improve quantization accuracy in practice.
  • The head-shared pairwise configuration results in higher perplexity compared to full-head mixing across multiple checkpoints.
  • Different statistics used by the analytic surrogate and the dynamic quantizer contribute to the observed discrepancies in performance.
Read more
Transferable Above-Ground Biomass (AGB) Estimation Model from Multi-Sensor Data with Sparse Field Calibration
Pann Thinzar Seint, Bryan Atwood, Subas Chhatkuli
Computer Vision Multimodal Optimization
  • Introduces a global CNN model for AGB estimation using multi-sensor data.
  • Employs a lightweight calibration process to adapt the model to local conditions.
  • Achieves competitive accuracy compared to existing biomass estimation methods.
  • Demonstrates the effectiveness of combining optical and radar data for biomass mapping.
Read more
Scaling Automatic Research Agents via World Models
Xiyuan Yang, Sheikh Sarwar, Jingru Cheng, Zhan Shi, Duanshun Li, Huiyuan Chen, Haiyang Zhang, Chenlei Guo, Jingrui He, Zhenyu Liao
Reinforcement Learning Large Language Models Efficient ML
  • Introduction of World Model RL (WMRL) to replace expensive environment execution in AutoResearch agents.
  • Development of Online Debiasing and Inverse-Variance Denoising mechanisms to improve training efficiency and performance.
  • Theoretical grounding of the framework with proven improvements in convergence guarantees.
  • Empirical validation showing 3-4x acceleration in training and superior performance compared to larger models.
Read more
Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation
Valentin Noël
NLP Large Language Models Interpretability
  • Measurement position is determined by the dictionary, leading to comparisons at different tokens.
  • Fixing the measurement position significantly reduces variance attributed to dictionary differences.
  • Additional reporting choices can alter the interpretation of results.
  • A new protocol for reporting ablation-based causal numbers is proposed for better comparability.
Read more
Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks
Wojciech Zarzecki, Jarosław Arabas
Optimization
  • Existing global optimization benchmarks are outdated and limited in scope.
  • Black-box adversarial attacks present a novel and relevant benchmark for optimization methods.
  • The study evaluates the performance of various evolutionary algorithms and metaheuristics on BBAA problems.
  • The findings suggest a need for modern benchmarks that reflect real-world challenges in machine learning.
Read more
Into the ORBIT for Time Series: Training Regimes for Foundation Models
Hongjie Xia, Yiding Liu, Yifan Hu, Peiyuan Liu, Zewei Dong
Time Series
  • Introduction of ORBIT, a training paradigm for TSFMs that controls effective pre-training distribution.
  • Bootstrap Multi-Level Sampling and Omni-Range Incremental Training are key components of ORBIT.
  • Falcon-2.0, a simple encoder-only Transformer, is trained under ORBIT to enhance forecasting capabilities.
  • Rank-Guided Cross-Depth Alignment improves representation alignment across Transformer depths.
Read more
Learning Discrete Decisions for MIPs with Constraint-Aware Diffusion
Vincenzo Di Vito, Mehdi Taghizadeh, Deepjyoti Deka, Kaarthik Sundar, Ferdinando Fioretto
Optimization Generative Models Graph Learning
  • Introduction of Constrained Graph Diffusion (CGD) for MIPs, which enforces feasibility in the discrete decision-making process.
  • The methodology separates the discrete and continuous components of MIPs, allowing for efficient resolution of the continuous problem post-discrete decision generation.
  • Demonstrated effectiveness on two distinct optimization problems, showcasing substantial improvements in solution quality and feasibility.
  • Achieved significant computational speedups compared to traditional numerical solvers, enhancing the practicality of solving complex MIPs.
Read more
ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening
Antoine de Mathelin, Christopher Tosh, Wesley Tansey
Efficient ML
  • ScreenShot is a foundation model that enables few-shot predictions for combination drug screening without the need for molecular profiling.
  • The model is pretrained on a large dataset, allowing it to generalize effectively to new patient samples.
  • An innovative active learning strategy is developed to optimize experimental design, reducing costs while maintaining high accuracy.
  • ScreenShot outperforms traditional methods and baselines in prediction accuracy and treatment identification.
Read more
Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints
Zhen Xu
Optimization Theory Efficient ML
  • Introduces a new algorithm for Contextual Bandits with Knapsack constraints.
  • Achieves a significant improvement in average regret bounds compared to existing literature.
  • Utilizes a UCB-guided linear programming re-solving heuristic.
  • Focuses on practical applications in digital platforms with known resource consumption.
Read more
Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency
Guo An, Zijing Wu, Honghua Dong, Yuhao Yan, Zixuan Gui, Haochong Chen, Shanzhao Ruan, Xiang Wang, Yurong Ling, Qi Tian
Reinforcement Learning Robotics Theory
  • Introduction of Action-Conditioned Predictive Consistency (ACPC) for diagnosing world models.
  • Establishment of Invariance Radius (IR) and Separation Rate (SR) as metrics for evaluating model robustness against visual perturbations.
  • Proven bounds on the impact of visual perturbations on multi-step prediction errors and planning costs.
  • Empirical validation across multiple control tasks and architectures, showing consistent diagnostic trends.
Read more
Sustaining Plasticity via Learnable Wavelet Activations in Continual Learning
Zeyang Zhang, Tieliang Gong, Junyan Lu, Weizhan Zhang
Theory Optimization
  • Introduces a learnable wavelet activation function to address plasticity loss in continual learning.
  • Combines low-frequency and high-frequency components to counteract spectral bias.
  • Employs dynamic wavelet injection and targeted regularization to enhance plasticity and stability.
  • Provides theoretical guarantees for the framework's effectiveness in L2 approximation.
Read more
Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization
Zhixin Ren, Yau Lyu, Congrong Li, Liping Zhang, Shengbo Eben Li
Optimization
  • Introduction of the AIM framework, which provides a unified interpretation of momentum-based optimization.
  • Separation of update geometry and acceleration mechanisms in optimizers, enhancing understanding of their interactions.
  • Development of RADAR, a new optimizer that combines multiple advanced techniques for improved performance.
  • Establishment of stochastic convergence for RADAR through rigorous analysis.
Read more
Learning the Mathematical Property for Designing Low Mutual Coherence Binary Sensing Matrices
Rekha, Santosh Singh, S. K. Neogy
Theory Optimization Efficient ML
  • Introduces a learning-based technique for constructing binary sensing matrices without large datasets.
  • Focuses on optimizing the mutual incoherence property for improved signal recovery.
  • Utilizes a neural network framework to generate sensing matrices, reducing computational costs.
  • Demonstrates superior performance compared to conventional deterministic and random matrix methods.
Read more
Comment on 'Modeling rapid language learning by distilling Bayesian priors into artificial neural networks'
Orr Well, Idan Tarshish, Nur Lan, Roni Katzir
Theory
  • M&G's methodology does not effectively distill a Bayesian prior into ANNs as claimed.
  • The use of cross-entropy loss leads to overfitting and poor generalization in M&G's model.
  • Early stopping may not serve as a valid regularization method in the context of M&G's approach.
  • The critique highlights the importance of properly defining and implementing Bayesian priors in machine learning.
Read more
Federated Learning for Distributed CNC Tool Wear Prediction
Afsana Khan, Morris Stallmann, Marcin Pietrasik, Charis Kouzinopoulos, Anna Wilbik
Federated Learning
  • Federated Learning enables collaborative model training without sharing raw data, addressing privacy concerns.
  • The study utilizes the MATWI dataset, which contains diverse sensor data relevant for tool wear prediction.
  • Federated models demonstrate performance close to centralized models and significantly better than local client models.
  • The findings support the feasibility of federated learning in industrial environments with distributed data.
Read more
DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees
Tianyi Li, Yaxin Luo, Xinyi Shang, Zhiqiang Shen
NLP Large Language Models Efficient ML
  • DARTree introduces a training-free speculative decoding method that extends AR correction from chains to trees.
  • The method constructs a fixed-width candidate tree, allowing for efficient parallel processing of token proposals.
  • DARTree achieves significant speedup and higher acceptance rates compared to existing diffusion-based speculative decoding methods.
  • The approach effectively decouples causal correction from sequential heap operations, addressing latency bottlenecks.
Read more
Defensive Boosting for Online Probabilistic Forecasting
Georgy Noarov, Aaron Roth
Theory Optimization Efficient ML
  • The Defensive Booster algorithm provides dual guarantees for online probabilistic forecasting.
  • It achieves competitive Brier scores while also reducing classification error under specific conditions.
  • The algorithm is efficient, requiring only one weak-class learner compared to previous methods.
  • A strongly adaptive variant offers local hard-core certificates, enhancing prediction reliability.
Read more
FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting
Rentao Gu, Yihang Ding, Junjie Li, Yi Ding, Weijing Sang, Xiaoli Huo, Xin Qin, Yuefeng Ji
Large Language Models Time Series
  • FM-LLM enables prompt-free adaptation of LLMs for time series forecasting.
  • The framework utilizes a Fourier Analysis Network for structured spectral token alignment.
  • An asymmetric Mixture-of-Experts architecture allows for specialized modeling of periodic and non-periodic dynamics.
  • FM-LLM achieves state-of-the-art performance on multiple forecasting benchmarks.
Read more
Doubly Robust Estimation of Causal Effect on CVR with Targeted Regularization
Jiayi Dan, Bo Li, Lu Deng, Yong Wang
Theory
  • Introduction of a new doubly robust estimator for CVR causal effects.
  • Theoretical guarantees based on semiparametric theory and von Mises expansion.
  • Development of a targeted regularization framework to improve numerical stability.
  • Extensive validation through experiments on synthetic and real-world data.
Read more
Forward and Inverse Virtual Metrology for Phototransistor Gain: A Hierarchical, Uncertainty-Aware Approach for Small Production Datasets
Mahshid Amirabgir, Lorenza Ferrario, Paolo Conci, Mahdieh Amirabgir, Giancarlo Orengo
Optimization Theory Efficient ML
  • Identified that approximately half of the variance in phototransistor gain is due to differences between process runs.
  • Developed a forward gain predictor with uncertainty quantification and an inverse search for recipe optimization.
  • Introduced a multi-level data quality assessment tailored to the hierarchical structure of fabrication data.
  • Released dataset and analysis code to promote reproducibility in research.
Read more
I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization
Yubo Zhang, Xinhong Ma, Zezhong Tan, Ziqiang Dong
Reinforcement Learning Large Language Models Optimization
  • I-SDPO introduces a capability-dependent routing mechanism for self-distillation in reinforcement learning.
  • The method effectively addresses the degenerate gradient problem in GRPO by selectively applying self-distillation.
  • I-SDPO achieves significant performance improvements on the SciKnowEval benchmark across multiple domains.
  • The approach reduces reliance on biased teacher signals as model capability increases, enhancing learning efficiency.
Read more
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL
Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu
Reinforcement Learning Large Language Models NLP
  • Rubric-as-reward RL can lead to reward hacking due to fixed proxy criteria.
  • Rubric Dropout mitigates reward hacking by randomly dropping rubric criteria during training.
  • The method improves out-of-distribution performance on benchmark tasks.
  • Rubric Dropout is simple to implement and requires minimal changes to the reward function.
Read more
A Multispectral Framework for the Detection of Calcium Carbide-Induced Ripening and Shelf-Life Estimation in Climacteric Fruits
Gurbhit Chaurakoti, Harshit Kumar, Hani Kumar, Anurag Singh, Ram Asrey
Multimodal
  • Developed a non-invasive multispectral framework for detecting CaC2-induced ripening in fruits.
  • Utilized visible-near infrared spectroscopy to analyze spectral profiles of mango and banana.
  • Achieved high classification accuracy (95% for mango, 81% for banana) using XGBoost algorithms.
  • Integrated environmental parameters and spectral features for improved detection and estimation.
Read more
Mechanism Design for Generative Engines: From Exploitation toward Win-Win Outcomes
Chen Xu, Zitian Guo, Chenyan Xiong
Theory Optimization Generative Models
  • The paper formulates the interaction between content providers and platforms as a repeated Stackelberg game, highlighting the strategic dynamics involved.
  • A new mechanism, Verifiable-Content Rewards (VCR), is proposed to align incentives between content creators and platforms, promoting high-quality content.
  • Simulation experiments show that VCR outperforms traditional defenses, achieving an average improvement of 12.1 percentage points in defense-utility scores.
  • The findings suggest that without proper mechanisms, the generative engine ecosystem risks devolving into citation wars that degrade content quality.
Read more
A Local Sinkhorn Framework for Conditional Distribution Reconstruction of Multidimensional Random Fields
Mingtao Xia, Qijing Shen
Theory Generative Models Efficient ML
  • Introduction of a local entropically regularized optimal transport framework for stochastic neural networks.
  • Derivation of generalization error bounds that characterize the learning efficiency of the proposed framework.
  • Demonstration of superior performance in reconstructing stochastic random fields compared to existing methods.
  • Establishment of a computationally efficient local distribution matching objective using Sinkhorn divergence.
Read more
When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits
Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
Reinforcement Learning Optimization Theory
  • Traditional offline evaluation methods can mislead in delayed-feedback contexts.
  • The proposed diagnostic protocol assesses alignment and learnability before trusting reported improvements.
  • A denser reward signal enhances online learning efficiency, revealing differences in rewards that appear tied in static estimates.
  • Personalization may not always be beneficial; sometimes, it serves as a robustness measure.
Read more
A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression
Wouter W. L. Nuijten, Esther G. van Pelt, Albert Podusenko, İsmail Şenöz, Wouter M. Kouw
Theory Efficient ML Time Series
  • Introduces a factor graph representation for multi-output Gaussian process regression.
  • Achieves linear scaling in the number of data points and handles missing observations efficiently.
  • Demonstrates competitive performance against traditional methods in terms of accuracy and computational cost.
  • Utilizes a nearest-neighbor chain to approximate high-dimensional input geometry.
Read more
GENADA: efficient generative time series adversarial attack framework
Michael Baronov, Denis Vorobev, Margarita Rusanova, Petr Sokerin, Alexey Zaytsev
Time Series Generative Models Efficient ML
  • GENADA provides a generative approach to adversarial attacks, eliminating the need for iterative optimization.
  • The framework allows for efficient perturbation generation through a single forward pass after training.
  • Empirical results show that GENADA achieves competitive attack quality with significantly reduced generation time compared to traditional methods.
  • The method is validated across multiple datasets and neural architectures, demonstrating its versatility in the time series domain.
Read more
Revisiting Overestimation Bias Problem of Q-learning: Settling Large Discrete Action Space via Action Intersection
Pu Li, Tao Tan, Hong Xie, Xiaoyu Shi, Mingsheng Shang
Reinforcement Learning
  • Identifies the overestimation bias problem in Q-learning exacerbated by large action spaces.
  • Proposes an action intersection strategy to enable semi-decoupling of Q-value estimation.
  • Demonstrates that the action intersection strategy allows for flexible bias control and fine granularity.
  • Shows significant performance improvements over state-of-the-art methods in both tabular and deep RL settings.
Read more
Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice
Ziqi Zhao, Jialin Lu, Junjie Shan, Junyuan Zhang, Shuya Yang, Ka-Ho Chow
Federated Learning
  • Identifies a gap between theoretical research and practical realities in VFL backdoor vulnerabilities.
  • Critiques existing methodologies for relying on unrealistic assumptions about threat models.
  • Introduces BVBench, a benchmark for fair evaluation of backdoor attacks in VFL.
  • Demonstrates that current understanding of VFL backdoor risks is fragile and often misleading.
Read more
Disentangling the Expressivity of RoPE
Selim Jerad, Anej Svete, Jiaoda Li, Ryan Cotterell
Theory NLP Large Language Models
  • RoPE can be categorized into periodic (RoPEP) and conventional variants, each with distinct expressivity properties.
  • RoPEP is shown to recognize languages definable in past temporal logic with modular predicates, enhancing its theoretical foundation.
  • Conventional RoPE's non-repeating rotations lead to a bounded locality bias, which can hinder performance on tasks needing long-distance context.
  • Controlled experiments validate the theoretical claims, showing that periodic RoPE outperforms conventional RoPE in specific language tasks.
Read more
Virtual Temperature Sensors in Power Transformers Using Neural Ordinary Differential Equations
Berk Hadzhamolla, Alexander Johannes Stasik, Signe Riemer-Sørensen
Time Series
  • Introduction of Neural Ordinary Differential Equations (Neural ODEs) for modeling transformer thermal behavior.
  • Integration of simplified heat-transfer equations into the Neural ODE framework for physics-aware predictions.
  • Evaluation of the model using data from fifteen transformers, showcasing its robustness and generalization capabilities.
  • Demonstration of improved forecasting accuracy compared to traditional numerical and data-driven methods.
Read more
Air Quality Station Simulation via LSTM and Attention-Based Modelling
Alexander Kostadinov, Petar O. Hristov, Dessislava Petrova-Antonova
Time Series
  • Introduction of SATADL model for simulating air quality station data during outages.
  • Utilization of spatial and temporal attention mechanisms for improved prediction accuracy.
  • Demonstrated superior performance over baseline models in forecasting PM10 concentrations.
  • Potential application in urban air quality management and smart city initiatives.
Read more
Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology
Yunsung Chung, Yingshuo Liu, Abboud F. Hassan, Han Feng, Mary M. Maleckar, Nassir Marrouche, Jihun Hamm
Multimodal
  • Introduces an intervention-aware latent clinical world model that evolves a 3D anatomical state based on post-procedural events.
  • Develops a horizon-token formulation for anytime recurrence forecasting from partial histories during the blanking period.
  • Achieves AUROC of 0.756 and AUPRC of 0.777 for recurrence prediction in atrial fibrillation ablation.
  • Provides retrospective risk estimates at different horizons, allowing for dynamic updates as new clinical data becomes available.
Read more
Dual-Model Sentiment Analysis of Consumer Reviews in the Retail Coffee Sector Using Machine Learning and Deep Learning Approaches
Muntasir Hasan Kanchan, Md. Alamgir Hossain, Md. Samiul Islam, Muhammad Masud Tarek
NLP
  • Introduction of a dual-model sentiment classification framework combining machine learning and deep learning techniques.
  • Comprehensive evaluation of multiple algorithms on a real-world, imbalanced dataset.
  • Implementation of a preprocessing pipeline using NLP techniques to enhance input quality.
  • Performance analysis based on various metrics, including accuracy and F1-score.
Read more
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement
Karim Aly, Alexei Sharpanskykh, Jacco Hoekstra
Generative Models
  • TailBooster effectively addresses the rarity of extreme events in historical flight data by augmenting datasets with synthetic extremes.
  • The framework combines generative modeling with anomaly detection to ensure operational validity of generated records.
  • Significant improvements in prediction accuracy for extreme air time and arrival delays were achieved compared to conventional synthetic data.
  • TailBooster is model-agnostic and can be applied across various domains where extreme-event prediction is critical.
Read more
TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven Cascading Failures
Orkun Irsoy, Leman Akoglu, Osman Yagan
Graph Learning Reinforcement Learning Optimization
  • TANGCO optimizes capacity allocation to enhance resilience against cascading failures in networked systems.
  • The approach utilizes a graph neural network to learn from cascade dynamics, overcoming limitations of traditional heuristics.
  • TANGCO demonstrates superior performance in robustness across multiple synthetic and real-world networks.
  • The methodology allows for transferability of learned policies to unseen graphs, enhancing practical applicability.
Read more
Symmetry-Breaking De Novo Crystal Generation via Markovian Jump Diffusion
Van Khoa Nguyen, Alexandros Kalousis
Generative Models
  • Introduces a novel diffusion-based framework for crystal generation that captures complete crystallographic specifications.
  • Utilizes a Markovian jump-diffusion process to model symmetry-breaking dynamics and inter-space-group transitions.
  • Demonstrates superior performance of SbCD over existing models in generating crystal structures.
  • Addresses the limitations of current generative models that rely on empirical distributions for space groups.
Read more
Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling
Takieddine Soualhi, Jacques Saraydaryan, Laetitia Matignon
Reinforcement Learning Robotics
  • Introduction of a proxemics-based reward model for DRL in social navigation.
  • The model is grounded in Hall's theory of interpersonal distance, enhancing interpretability.
  • Validation across multiple DRL navigation methods shows improved social metrics.
  • Emphasis on the need for explicit assessment of comfort in navigation tasks.
Read more
Geometric and Behavioral Stratification in Transformer Residual Streams
Nelson Guda
NLP Large Language Models Interpretability
  • The prediction direction acts as a content-defined privileged anchor in transformer residual streams.
  • Residual-stream variation is geometrically and behaviorally stratified based on proximity to the prediction direction.
  • A narrow prediction interface is universal across various transformer models, regardless of size.
  • Disrupting high variance directions near the prediction direction leads to immediate task-frame shifts.
Read more
Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates
Qiyao Zhou, Xujia Zhu, Pierre Joli, Yu Cong, Sibo Cheng
Optimization Theory Efficient ML
  • Introduces a physics-aware latent-space framework for parameter calibration.
  • Demonstrates the importance of observable supervision in training for better latent representation.
  • Shows that traditional reconstruction accuracy is inadequate for effective inverse modeling.
  • Evaluates the method on CFD benchmarks, highlighting its robustness in realistic measurement settings.
Read more
Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion
David Bechtoldt, Sidney Bender
Graph Learning Interpretability Generative Models
  • GDCE-I provides a comprehensive framework for generating graph counterfactuals that respect domain-specific constraints.
  • The method utilizes a discrete denoising diffusion model combined with a novel inversion technique to ensure data-manifold awareness.
  • GDCE-I outperforms existing counterfactual explanation methods across multiple benchmarks.
  • The paper introduces a standardized evaluation framework for assessing graph counterfactuals.
Read more
On the global feature importance for interpretable and trustworthy heat demand forecasting
Milan Zdravković
Interpretability Time Series
  • Introduces an ante-hoc XAI methodology for assessing global feature importance in heat demand forecasting.
  • Utilizes four different interpretability approaches, including Gradient Boosting and post-hoc methods.
  • Addresses challenges in model interpretability related to compliance and customer satisfaction.
  • Highlights the importance of understanding model decisions in complex systems like District Heating.
Read more
Long-Horizon Forecasting of Complete Financial Statements with Forma
Travis L. Johnson, Jiannan Jiang, Soumyabrata Chaudhuri, Yihao Chen, Lauren Falvey, Donal O'Cofaigh
Time Series
  • ProForma-20Q is introduced as a benchmark for forecasting complete financial statements over long horizons.
  • Forma, a transformer-based model, outperforms all tested competitors, including classical ML and large language models.
  • The model's performance improves with longer forecasting horizons, crucial for accurate valuation.
  • Forma's architecture allows for scenario analysis and maintains accounting coherence in forecasts.
Read more
A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings
Hei Ting (Una) Chan, Chenwei Wu, Xueshen Liu, Zesen Zhao, Boyuan Zheng, Luis Filipe Nakayama, Michael G. Morley, Liyue Shen, Jiasi Chen, Z. Morley Mao
Multimodal Efficient ML Interpretability
  • Introduction of a cloud-edge collaborative architecture for multimodal clinical screening.
  • Dynamic selection of diagnostic tools based on patient context to optimize modality coverage.
  • Evaluation framework includes 100 multimodal cases and simulates rural network conditions.
  • Achieved high accuracy and factual grounding while minimizing data transmission to the cloud.
Read more
Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction
Yi Liu
Theory Optimization Robotics
  • Introduces a decision-resource formulation that leverages terminal symmetry in directed construction tasks.
  • Develops SYMBUILD, which implements a transport-refine-certify approach to enhance decision-making.
  • Demonstrates significant improvements in verified efficiency across multiple domains.
  • Achieves the lowest mean capped verifier cost compared to existing planners in GRN OOD scenarios.
Read more
Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference
Zixuan Lan, Yanhong Li, Jiawei Zhou
Large Language Models NLP Efficient ML
  • RMM is a training-free, input-adaptive method that reduces matrix products in Transformers without modifying model weights.
  • The method allows for a predictable accuracy-efficiency trade-off through a simple retention ratio.
  • Larger models generally tolerate more aggressive reductions, but this tolerance is task and model-dependent.
  • Attention-side computations are more reducible than MLP components, indicating structural asymmetry in Transformers.
Read more
Consolidator: Learning Persistent Routed Memory Across Context Boundaries
Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung
Theory
  • Introduces a shared slot-local transform for STM consolidation without replay.
  • Demonstrates that consolidated LTM can guide later memory slot selection.
  • Shows significant improvements in recall performance with learned consolidation.
  • Isolates and evaluates multiple memory functions in a controlled task.
Read more
CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution
Zihao Ye, Yingyi Huang, Hongyi Jin, Bohan Hou, Junru Shao, Zhongming Yu, Jinqi Chen, Meghan Cowan, Shiyi Cao, Shanli Xing, Hanfeng Chen, Vinod Grover, Tianqi Chen, Luis Ceze
Optimization Efficient ML Theory
  • CAKE enables agents to author a typed intermediate representation (IR) for better control over GPU programming.
  • The framework provides localized correctness and performance diagnostics, enhancing the feedback loop for kernel evolution.
  • Agent-generated kernels show substantial performance improvements, outperforming traditional CUDA/PTX implementations.
  • CAKE evolves its IR based on real-world production kernels, ensuring it adapts to emerging workloads and capabilities.
Read more
A Compositional Theory of Curvature in Probabilistic Circuits
Hrithik Suresh, Sahil Sidheekh, Shelar Parth Vijay, Yasir Z, Sriraam Natarajan, Narayanan Chatapuram Krishnan
Generative Models Optimization Theory
  • Probabilistic Circuits (PCs) allow for exact inference and tractable curvature measures unlike deep neural networks.
  • Global sharpness regularization can lead to underfitting in PCs due to their compositional curvature characteristics.
  • The contribution of nodes to the Hessian trace can be decomposed into contextual usage and local sharpness, providing insights into effective regularization.
  • An adaptive sharpness-aware regularizer is proposed, which targets nodes based on their local curvature, improving generalization.
Read more
Balanced Adaptive Prototype Selection for Scalable TabPFN Inference on Large-Scale Tabular Data
Mahboobe Jadid, Melika Rezaye Garkani, Ali Mousavi
Efficient ML
  • Identifies bounded inference context as a primary scalability bottleneck for pretrained tabular models.
  • Introduces BAPS, a framework for constructing compact inference contexts that preserve critical information.
  • Demonstrates that BAPS can achieve approximately 1,953-fold context compression while retaining strong predictive performance.
  • Shows that BAPS is compatible with the original TabPFN architecture without requiring model retraining.
Read more
The Time Value of Evolution
Matthew Siper, Ahmed Khalifa, Julian Togelius
Reinforcement Learning Optimization Theory
  • Formalization of the 'time value of evolution' concept.
  • Introduction of Lineage-Value Policy Gradients (LVPG) for evolutionary search.
  • Demonstrated improvement in search efficiency and performance through long-horizon credit assignment.
  • Reduction in temporary regressions compared to immediate-return optimization.
Read more