AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

69 Papers today
8h Update frequency
7 Days of history
An Analysis of Residual-Stream Geometry Across Transformer Depth
Sunit Bhattacharya, Ravi Shankar Kolli
NLP Large Language Models Interpretability
  • Relative displacement of representations is structured by layer depth, with larger updates in early and late layers.
  • The magnitude of global Procrustes rotation remains nearly constant across depth, while residuals peak at the final layer.
  • Non-English targets exhibit greater final-layer displacement and residual compared to English targets.
  • The findings are statistically robust and suggest that depth curves are largely stable across conditions and models.
Read more
Counterfactual Shapley Credit Assignment
Mingxuan Li, Kai-Zhan Lee, Elias Bareinboim
Reinforcement Learning Theory Interpretability
  • Introduction of Counterfactual Shapley Credit Assignment framework for RL.
  • Demonstration of ϕ-values that preserve optimal policy while redistributing rewards.
  • Development of a consistent estimator for ϕ-values with amortized constant time complexity.
  • Introduction of ϕ-PPO with Prioritized Trajectory Replay, showing improved sample efficiency.
Read more
HantaWatch: Federated Learning for Hantavirus Genomic Surveillance
Shanika Iroshi Nanayakkara, Shiva Raj Pokhrel
Federated Learning
  • HantaWatch enables decentralized federated learning for hantavirus genomic surveillance without centralizing raw data.
  • The framework produces comprehensive outputs including predicted labels, risk scores, and expert-review priorities.
  • Adaptive optimization techniques are employed to manage client drift in non-IID environments.
  • Surveillance-specific model selection is based on multiple performance metrics beyond accuracy.
Read more
Thermodynamics-Informed Input Reparameterization for Neural Prediction of Real-Fluid Thermodynamic Properties in Supercritical Combustion
Haoze Zhang, Han Li, Ke Xiao, Yangchen Xu, Runze Mao, Zhi X. Chen
Theory Efficient ML
  • Introduction of target-aligned input reparameterization (TAIR) to improve neural network predictions of thermodynamic properties.
  • TAIR replaces raw enthalpy inputs with thermodynamically informed coordinates, enhancing model learnability.
  • Significant reductions in root-mean-squared error (RMSE) for temperature, density, and compressibility compared to baseline methods.
  • Demonstrated effectiveness on supercritical methane-oxygen counterflow flame data, showing improved accuracy and efficiency.
Read more
Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification
Filippo Cenacchi, Longbing Cao, Runze Yang
Time Series
  • SEB-Cal introduces a novel approach to reliability estimation in time-series classification that goes beyond traditional output-space calibration.
  • The method leverages spectral evidence from the entire sample to assess the trustworthiness of predictions.
  • Validation-gated policy ensures that spectral conditioning is applied only when it improves correctness ranking without violating error constraints.
  • Empirical results show significant improvements in selective reliability metrics across multiple datasets and backbone architectures.
Read more
A Better Start for Language Models: Domain-Conditional Position Offsets
Ye Qiao
NLP Large Language Models Efficient ML
  • Introduces domain-conditional position offsets to mitigate cold-start penalties in language models.
  • Offsets are efficient, requiring minimal training data and parameters, while maintaining model weights frozen.
  • Demonstrates significant reductions in in-domain perplexity across multiple language model architectures.
  • Offsets improve retrieval reranking and domain classification without affecting few-shot reasoning tasks.
Read more
Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression
Athanasios Ntovas, Alexandros Doumanoglou, Petros Drakoulis, Dimitris Zarpalas
NLP Large Language Models Efficient ML
  • Combines neuron importance estimation with data-aware low-rank approximation for improved model compression.
  • Introduces a dynamic compression rate allocation algorithm for better distribution of compression across layers.
  • Achieves state-of-the-art performance in compressing various foundational LLMs.
  • Demonstrates that grouping weight matrices by layer index is more effective than previous methods.
Read more
Spatio-Temporal Prediction of Unsteady Airfoil Aerodynamics Using Augmented Graph Neural Ordinary Differential Equations with Exogenous Controls
Henrik Lange, Reik Thormann, Philipp Bekemeyer
Graph Learning Time Series Optimization
  • Introduction of the GNODE framework for predicting unsteady aerodynamics.
  • Augmentation of GNNs with Neural Ordinary Differential Equations improves prediction stability and accuracy.
  • The model outperforms traditional autoregressive GNNs in handling complex aerodynamic phenomena.
  • Demonstrated effectiveness on a dataset simulating a pitching airfoil with transonic shocks.
Read more
A Reinforcement-Learning-Augmented Liquid-Fueled Reactor Network Model for Predicting Lean Blowout in Gas Turbine Combustors
Philip John, Eloghosa Ikponmwoba, Pinaki Pal, Opeoluwa Owoyele
Reinforcement Learning
  • Introduction of a reinforcement learning framework for optimal reactor network generation.
  • Goal-oriented clustering that improves lean blowout prediction accuracy.
  • Demonstration of improved predictive fidelity and computational efficiency over traditional methods.
  • First liquid-fueled reactor network capable of reproducing parametric LBO trends.
Read more
SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions
Hoang-Thang Ta
Efficient ML Theory Computer Vision
  • Introduction of SechKAN, a KAN architecture based on hyperbolic secant functions.
  • SechKAN achieves competitive performance while using fewer parameters than existing KAN variants.
  • The architecture is effective in function fitting, PDE problems, and image classification tasks.
  • Comprehensive analysis of design choices provides practical guidelines for model configuration.
Read more
Quantifying Ranking Uncertainty in LLM Benchmarks
Bitya Neuhof, Yuval Benjamini
NLP Large Language Models
  • Introduces rank confidence intervals (CIs) to quantify uncertainty in model rankings.
  • Analyzes sources of ranking uncertainty in the MMLU benchmark.
  • Demonstrates substantial variability in rankings across MMLU subjects.
  • Proposes modifications to hypothesis tests for better uncertainty estimation.
Read more
GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models
Alessandro Scalese, Santhanakrishnan Narayanan, Constantinos Antoniou
Graph Learning Optimization
  • Introduction of GUIDED, a network-agnostic feature initialization layer for GNNs.
  • Enhances spatial transferability of GNN models for traffic assignment tasks.
  • Maintains high predictive accuracy and robustness under data scarcity.
  • Enables efficient domain adaptation without structural modifications.
Read more
Program Synthesis for Simulation-Based Inference: Joint Model Selection and Parameter Estimation
Siddharth Mishra-Sharma
NLP Large Language Models Theory
  • Introduces a framework for joint model selection and parameter estimation using LLMs and SBI.
  • LLMs generate candidate simulator programs from natural language descriptions, enhancing model exploration.
  • The method iteratively refines models and evaluates them using neural density estimation.
  • Validated on benchmarks with known ground truth, showing effective model identification.
Read more
Automated Data Engineering and Feature Selection for the Case Study of Warpage Detection in Fused Deposition Modeling
Saleh Valizadeh Sotubadi, Nazanin Mahjourian, Vinh Nguyen
Reinforcement Learning Interpretability Efficient ML
  • Development of an Automated Data Processing framework for FDM warpage detection.
  • Utilization of reinforcement learning for dynamic model-feature selection.
  • Implementation of SHAP XAI for effective feature subset generation.
  • Significant improvements in model accuracy and stability demonstrated.
Read more
Unsupervised Multi-kernel Learning for Automated Algorithm Selection
Yihang Lu, Tome Eftimov, Carola Doerr
Optimization
  • Introduces an unsupervised multi-kernel learning framework for algorithm selection, avoiding the pitfalls of supervised models.
  • Utilizes a multi-kernel k-means approach to cluster problem instances based on heterogeneous landscape representations.
  • Demonstrates superior performance on Differential Evolution tasks and competitive results on Particle Swarm Optimization tasks.
  • Provides interpretable insights into the relevance of different landscape representations for algorithm selection.
Read more
Neural Controlled Differential Equations for EMT-Level Surrogate Modeling of Grid-Forming Inverters
Jiagang Qu, Yong Tao, Dan Wang, Enyi Li, Jingjing Qi, Ding Wang
Time Series
  • Introduction of Neural CDE framework for continuous-time surrogate modeling of grid-forming inverters.
  • Affine-control decomposition to separate system dynamics from control effects.
  • Dual-path control embedding for capturing multi-time-scale behaviors.
  • Physics-informed regularization to enhance stability and coherence in simulations.
Read more
L1 Augmented Attention as an Improved Vector Similarity Metric
Kurt Godden
NLP Large Language Models
  • L1-augmented attention improves upon traditional dot-product attention by incorporating L1 distance metrics.
  • The method captures both directional alignment and coordinate-wise deviations for better similarity assessment.
  • Evaluation on WikiText-2 shows a reduction in perplexity by up to 14.5% compared to the baseline.
  • The approach allows for computational efficiency through low-dimensional projections of queries and keys.
Read more
From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training
Zishang Jiang, Tingyun Li, Jinyi Han, Xinyi Wang, Sihang Jiang, Yizhou Ying, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao
NLP Large Language Models Reinforcement Learning
  • Introduction of Hindsight Policy Optimization (HPO) to improve long-horizon RL training.
  • HPO utilizes an intent space to reduce optimization variance by aggregating semantically similar actions.
  • The method employs Wasserstein distance to measure discrepancies between policy distributions.
  • Empirical results show HPO significantly reduces training noise and stabilizes optimization.
Read more
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference
Simla Burcu Harma, Danila Mishin, Zhengyuan Su, Ayan Chakraborty, Elizaveta Kostenok, Dongho Ha, Babak Falsafi, Martin Jaggi, Yunho Oh, Amir Yazdanbakhsh
Large Language Models Efficient ML
  • MXSens introduces a sensitivity-guided approach to mixed-precision quantization, optimizing bitwidth allocation based on layer and column sensitivity.
  • The method is training-free and utilizes the MXINT format for efficient low-bit inference.
  • MXSens achieves state-of-the-art performance on LLaMA models, significantly improving perplexity metrics.
  • The approach addresses the challenges posed by outliers in activations, which are common in low-bit quantization.
Read more
Spectral Higher-Order Neural Networks Have Sharp Expressivity Bounds
Gianluca Peri, Diego Febbe, Duccio Fanelli
Theory Efficient ML Interpretability
  • SHONNs reduce parameter scaling to O(N^2) through spectral parameterization and weight sharing.
  • They can operate without enforcing traditional non-linearities, enhancing model expressiveness.
  • The framework was benchmarked on N-bit parity tasks, demonstrating improved performance.
  • SHONNs offer a highly tunable hypothesis space, making them versatile for various applications.
Read more
Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior
Garrett Baker, Vinayak Pathak, Daniel Murfet, Susan Wei
NLP Large Language Models Optimization
  • Introduces Posterior Prefix Tuning (PPT) for eliciting transformer behavior without backpropagation.
  • Utilizes latent posterior models to optimize prompts based on expected utility.
  • Demonstrates efficiency by using prior samples for multiple utility functions without additional transformer calls.
  • Validates the method on Beta–Bernoulli and reinforced urn BFTs across different utility families.
Read more
WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning
Ryan Xu, Atlas Zhao, David Bao, Frank Du
Reinforcement Learning Optimization Efficient ML
  • WAR optimizes rollout generation in synchronous agentic RL by adapting to workload conditions.
  • Under low load, it employs SuffixDecoding for efficient speculative decoding without additional model overhead.
  • Under high load, it utilizes cache-aware scheduling to enhance resource utilization and reduce computation redundancy.
  • WAR achieves throughput improvements of 1.4× under low load and up to 1.6× under high load.
Read more
ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling
Chirag Vashist, Ke Li
Generative Models
  • Introduces a minimalist approach to generative modeling, focusing on simplicity and efficiency.
  • Utilizes Implicit Maximum Likelihood Estimation (IMLE) for a stable and effective training objective.
  • Achieves an FID of 2.56 on ImageNet 256 with a single forward pass, demonstrating high sample quality.
  • Reduces the number of parameters and function evaluations compared to state-of-the-art models.
Read more
Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative
David Gómez-Guillén, Mireia Diaz, Josep Lluis Arcos, Jesús Cerquides
Optimization Large Language Models Interpretability
  • Introduces an LLM-driven optimization method for calibrating grey-box simulation models in CEA.
  • Achieves lower median best error with fewer model evaluations compared to traditional methods.
  • Incorporates constraints through natural language prompts, eliminating the need for complex surrogate models.
  • Provides an auditable optimization trace, enhancing transparency and justification of modeling choices.
Read more
AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
Yu-Yang Qian, Hao-Cong Wu, Chen Chen, Jiacheng Sun, Zhenhua Dong, Peng Zhao, Zhi-Hua Zhou
NLP Large Language Models Efficient ML
  • ADAFLASH addresses high variance in draft quality from diffusion drafters.
  • The framework includes an on-policy distillation algorithm for stable convergence.
  • An adaptive length head reduces verification costs by dynamically adjusting sequence lengths.
  • Experimental results show significant performance improvements over previous methods.
Read more
HindsightBench: A Black-Box Behavioral Audit Protocol for Parametric Hindsight in Time-Indexed LLM Decision Tasks
Haozhe Jia
Large Language Models NLP Time Series
  • HindsightBench allows for cost-effective auditing of LLMs without requiring access to training data or backtesting.
  • The protocol integrates multiple components into a single causal design for comprehensive behavioral analysis.
  • Key findings include the influence of training generation on model behavior and the variability of effective cutoffs across models.
  • Audit results are sensitive to serving configurations, necessitating specific operational requirements.
Read more
Hybrid Latent-Structural Fusion (HLSF) for Cyber Anomaly Detection
Dorianis M. Perez, Maksim E. Eren, Bryan E. Kaiser
Theory
  • Introduction of Hybrid Latent-Structural Fusion (HLSF) framework for anomaly detection.
  • Integration of CP-APR and normalizing flows improves detection performance.
  • HLSF demonstrated superior results on real-world cyber anomaly data from LANL.
  • The methodology effectively captures both structural and distributional aspects of cyber behavior.
Read more
TurboVec: A Case Study in Cost-Efficient Private Retrieval for Enterprise RAG via Codebook-Oblivious Quantization
Navnit Shukla, Kamal Pandey, Omsankar Tiwari
NLP Large Language Models Efficient ML
  • TurboVec addresses codebook leakage and recall degradation in multi-tenant RAG systems.
  • TurboQuant outperforms trained FAISS Product Quantization in recall while using less memory.
  • The implementation achieves significantly lower query latency compared to traditional methods.
  • Kernel-level filtering maintains high recall rates for tenant-isolated retrieval.
Read more
Building a Neural Network from Scratch: Implementation, Evaluation, and Optimization
Yuanzhe Jia
Theory Optimization Efficient ML
  • The framework is implemented from scratch, enhancing understanding of neural network mechanics.
  • It includes diverse modules such as multi-layer architectures, activation functions, and optimization techniques.
  • The framework successfully validates its performance on a multi-class classification task.
  • It serves as both an educational tool and a baseline for future research.
Read more
Dynamic Loss Balancing for Joint SOH and RUL Prediction of Lithium-Ion Batteries via a Rotary SOH-Injected Prior Battery Transformer
Shuhao Chen, Tianyu Shi, Yiwen Huang, Chengyi Tu
Time Series Optimization Efficient ML
  • Introduction of RoSIP-Batt framework for joint SOH and RUL prediction.
  • Dynamic loss balancing through a Bayesian homoscedastic uncertainty weighting mechanism.
  • Utilization of Rotary Position Embedding for capturing degradation patterns.
  • Significant performance improvements over state-of-the-art methods on multiple datasets.
Read more
FALCON-Discover: Discovering Concentrated False-Confidence Regions for Calibration
Filippo Cenacchi, Longbing Cao, Runze Yang
Theory Interpretability
  • Introduces false-confidence concentration as a critical aspect of calibration analysis.
  • Presents FALCON-Discover, a post-hoc framework for discovering dangerous confident errors.
  • Demonstrates that different ranking strategies are effective in different regimes.
  • Establishes the importance of localizing high-risk error regions in predictive models.
Read more
Graph Neural Network-based Algorithm Selection for the Traveling Salesman Problem: A Systematic Study of Cost and Rank Losses under Distinct Budget Regimes
Zhaoxuan Li, Jiale Yang, Yifei Lu, Mustafa Misir
Graph Learning Optimization
  • Introduction of GNNAS-TSP, a GNN-based framework for TSP algorithm selection.
  • Elimination of manual feature engineering by learning directly from raw graph data.
  • Evaluation of various cost-based and rank-based loss functions for algorithm selection.
  • Significant performance improvements over the Single Best Solver under fixed computational budgets.
Read more
Selectivity Matters: Source Node Influence Pruning for Unsupervised Graph Domain Adaptation
Ridong Han, Yawen Shen, Zhongnian Li, Tongfeng Sun, Xinzheng Xu, Abdulmotaleb El Saddik
Graph Learning
  • SNIP shifts the focus from feature-level alignment to data-level refinement in UGDA.
  • The method quantifies structural discrepancies using centrality measures to assign influence scores to source nodes.
  • A rank-based normalization mechanism is introduced to standardize influence scores across different centrality measures.
  • SNIP effectively filters out structurally incompatible nodes, resulting in a refined sub-source graph.
Read more
Neural Kolmogorov Equations: Parallelizable Learning of Stochastic Dynamics under General Noise
Arthur Bizzi, Olga Fink
Time Series Generative Models Theory
  • NKEs provide a deterministic reformulation of Neural SDEs, focusing on probability density evolution.
  • The framework allows for noise-agnostic learning, accommodating diverse stochastic drivers.
  • NKEs enable parallel-in-time training, improving computational efficiency and scalability.
  • The method accurately models both deterministic and stochastic dynamics with competitive predictive accuracy.
Read more
Cost Accounting for Reactive Computational Graphs: Exhaustive Sweeps, Sequential Mutation, and the Backward-Locality Gap
Abdallah Khemais
Theory Efficient ML Interpretability
  • The speedup of exhaustive sweeps is not constant and depends on the depth profile of the network.
  • Sequential mutations incur an additional cost due to overcounting, while batched applications are order-independent and sub-additive.
  • The backward-locality gap restricts speedup during backpropagation in architectures lacking long skip connections.
  • Empirical results validate theoretical predictions across different cost profiles and mutation scenarios.
Read more
HyBDM: Multi-Scale Hybrid Experts for Time Series Forecasting with Bidirectional Dependency Modeling
Wenqiang Ma, Chen Cheng, Xue Cheng, Jiarui Ye
Time Series
  • HyBDM introduces a multi-scale hybrid framework with specialized experts for long- and short-range dependency modeling.
  • The Global Patterns Expert employs BiConv-Mamba with bidirectional convolutions and a forgetting mechanism for enhanced temporal modeling.
  • The Local Variations Expert uses a Local Window Transformer for efficient locality-aware attention.
  • HyBDM achieves superior accuracy and efficiency compared to state-of-the-art forecasting models.
Read more
Fully-sensorized smart-eyewear platform for on-device Machine Learning
Andrea Giudici, Christian Veronesi, Pietro Bartoli, Mario Caliò, Aurelio Teliti, Giacomo Gervasoni, Diana Trojaniello, Franco Zappa
Computer Vision Efficient ML Multimodal
  • ARGO leverages on-device machine learning for real-time processing, enhancing user privacy.
  • Introduces Head-wise Parallel Attention (HPA) for efficient model execution on the NPU.
  • Achieves a mean Average Precision (mAP50–95) of 24% with a memory footprint of 2.483 MB.
  • Integrates a multimodal sensor suite for comprehensive environmental monitoring.
Read more
Exposure-Based Reinforcement Learning to Rank
Harrie Oosterhuis, Rolf Jagerman, Zhen Qin, Xuanhui Wang
Reinforcement Learning Optimization Efficient ML
  • Introduces an exposure-based RL approach that avoids custom gradient computations.
  • Achieves higher sample efficiency and faster convergence compared to existing methods.
  • Integrates seamlessly with auto-differentiation, simplifying implementation for practitioners.
  • Demonstrates no additional computational costs when using GPUs.
Read more
Lightweight Wrappers for Adapting Time Series Foundation Models to Regional Drought Forecasting
Wentao Gao, Jiuyong Li, Lin Liu, Thuc Duy Le, Jixue Liu, Yanchang Zhao, Yun Chen
Time Series
  • Introduces a lightweight adaptation framework for TSFMs in drought forecasting.
  • Utilizes two wrappers (SMR2 and MBB) to enhance model performance without fine-tuning.
  • Achieves up to 26% reduction in MSE for SPEI predictions.
  • Addresses practical constraints of proprietary models and limited local data.
Read more
ALAS: Additive Learnable Alpha-Stable Kernels for Flexible Bayesian Optimization
Weibo Huang, Cheng Hua
Optimization Theory
  • Introduction of ALAS, a learnable α-stable kernel family for Bayesian Optimization.
  • ALAS adapts its smoothness based on data, capturing diverse objective structures.
  • ALAS-Sep variant improves robustness in high-dimensional optimization tasks.
  • Theoretical connections between spectral tails and information gain are established.
Read more
Physics-Guided Masked Multi-Task Network for Edge-Friendly Battery Health Diagnostics from Stochastically Fragmented Charging Profiles
Shuhao Chen, Tianyu Shi, Chengyi Tu
Time Series Efficient ML Theory
  • Introduction of PG-M2TN, a compact architecture for battery health diagnostics.
  • Utilizes a combination of BiLSTM–Attention, Masked Autoencoder, and dual-stream prediction for enhanced SOH estimation.
  • Achieves a global RMSE of 0.0781 across multiple datasets, indicating high accuracy.
  • Demonstrates robustness to data fragmentation and improved predictive sensitivity near late-life capacity collapse.
Read more
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs
Seunghyun Lee, Dongyoon Han, Sangdoo Yun
Large Language Models NLP
  • Token Inoculation allows LLMs to retain hazardous knowledge while controlling its expression using a special token.
  • The method achieves a significant reduction in hazardous query accuracy while preserving benign domain performance.
  • Refusal selectivity is influenced by the quality of the training signal during fine-tuning.
  • The approach demonstrates better safety-utility trade-offs compared to traditional methods like unlearning and refusal training.
Read more
After the Euclidean Highway: Hyperbolic Expert AI as the Next Innovation
Kwan Soo Shin, In Seok Kang, Munho Lee
NLP Large Language Models Theory
  • HySAT introduces hyperbolic geometry at the loss layer to enhance training stability and performance in domain-expert AI.
  • The paper identifies a critical placement law for integrating non-Euclidean structures into existing models without destabilizing training.
  • Extensive experiments with six SLMs show significant improvements in training efficiency and model performance.
  • The methodology allows for operational deployment of models with hierarchical domain knowledge, addressing limitations of traditional Euclidean approaches.
Read more
GEqTrain: A Configuration-Driven Framework for Retargeting Equivariant Graph Neural Networks Across 3D Scientific Tasks
Daniele Angioletti, Marco Nobile, Vittorio Limongelli
Graph Learning Generative Models
  • GEqTrain provides a modular framework for equivariant graph learning, enhancing reusability across different tasks.
  • The framework separates dataset semantics, model composition, and training objectives, allowing for flexible configuration.
  • GEqDiff extends the framework to generative modeling, enabling joint treatment of various geometric representations.
  • The framework demonstrates competitive accuracy in three distinct scientific applications without requiring separate software ecosystems.
Read more
An unsupervised clustering analysis of breast cancer data derived from electronic health records enhanced through UMAP dimensionality reduction
Davide Chicco, Nicoletta Benvenuto
Theory Efficient ML Interpretability
  • The study applies unsupervised clustering to breast cancer data from electronic health records.
  • DBSCAN is used for clustering, enhanced by UMAP for dimensionality reduction.
  • Three datasets are analyzed, revealing significant clustering patterns among patients.
  • Statistical indices confirm the effectiveness of the UMAP-DBSCAN combination.
Read more
Causal Discovery on Irregular Time Series
Martim Penim, Ricardo Ribeiro Pereira, Jacopo Bono, Hugo Ferreira, Mário A.T. Figueiredo, Pedro Bizarro
Time Series
  • Proposes a time-aware extension of PCMCI+ for irregular time series.
  • Aggregates causal influence over predefined temporal windows instead of fixed lags.
  • Demonstrates superior performance in recovering causal structures from synthetic irregular event streams.
  • Maintains interpretability and scalability of the original PCMCI+ framework.
Read more
TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
Changyue Li, Jiaming He, Youliang Yuan, Jialin Wu, Boxi Yu, Zhicong Huang, Pinjia He
NLP Large Language Models Optimization
  • TRACE shifts the focus from online merging to offline safety patch learning.
  • The framework simulates harmful tuning trajectories to optimize safety patches.
  • TRACE achieves nearly 100% safety while preserving task utility.
  • The method effectively disentangles safety recovery from user task updates.
Read more
ConceptCF: Concept-based Counterfactuals for the Explainability of Time Series
Annemarie Jutte, Faizan Ahmed, Jeroen Linssen, Maurice van Keulen
Time Series Interpretability Optimization
  • ConceptCF generates counterfactuals based on high-level human-interpretable concepts rather than individual data points.
  • The method utilizes time series decomposition to extract meaningful concepts like scale and periodicity.
  • A genetic algorithm is employed to optimize concept mutations for counterfactual generation.
  • ConceptCF outperforms existing counterfactual generation methods in terms of validity, confidence, proximity, sparsity, and plausibility.
Read more
Breaking Feedback-Blindness: Utility-Augmented Transformer for Sequential Decision Making
Yuyang Shen, Shan Dai, Daimin Chen
Reinforcement Learning Theory Interpretability
  • Introduces the concept of feedback-blind retrieval, highlighting a structural limitation in existing Transformer models.
  • Proposes the Utility-Augmented Transformer (UAT) that integrates feedback into the attention mechanism.
  • Theoretically proves that UAT can uniformly approximate feedback-dependent decision maps.
  • Demonstrates superior performance of UAT on various benchmarks compared to observation-only models.
Read more
Incomplete Observations Boost Evolutionary Performance in Ocean Modeling
Yangyang Kong, Yutong Jiang, Yanhai Gan, Junyu Dong, Feng Gao, Xiaopei Lin
Generative Models Optimization Time Series
  • Introduces a generative model that learns ocean dynamics from sparse observations.
  • Develops an optimization framework based on the expectation-maximization algorithm.
  • Demonstrates improved reconstruction and prediction performance using real-world data.
  • Addresses limitations of current models that rely on complete reanalysis datasets.
Read more
Federated Lightweight Fine-Tuning
Radhakrishna Achanta, Will Reed
Federated Learning Efficient ML
  • FLITE reduces communication overhead in federated learning by transmitting a small latent vector instead of full model weights.
  • The method achieves an 8718× reduction in communication size while maintaining competitive accuracy.
  • A low-rank factorization of the projection matrix significantly decreases memory requirements.
  • FLITE demonstrates robustness under non-IID data distributions and small client counts.
Read more
Intelligence from Learnable Novelty
Yanbo Zhang, Michael Levin
Theory Reinforcement Learning Optimization
  • Introduces 'learnable novelty' as a key concept linking different definitions of intelligence.
  • Develops a closed-form estimator for learnable novelty using a differentiable reservoir computer.
  • Demonstrates the estimator's effectiveness in complexity classification and unsupervised learning tasks.
  • Shows that using learnable novelty as an intrinsic reward improves exploration in reinforcement learning.
Read more
Apeliotes: A Diffusion-Based Modeling Framework for km-scale Multi-Level Atmospheric Fields
Evangelia Rafaela Frastali, Achyut Paudel, Maryam Golbazi, Frank Liu
Generative Models Time Series Efficient ML
  • Apeliotes leverages a global weather foundation model and a region-specific generative diffusion model for efficient km-scale weather forecasting.
  • The framework produces multi-level atmospheric fields, enhancing the detail and accuracy of weather predictions.
  • Apeliotes achieves competitive performance metrics, including less than 3% error in vertical wind profile predictions.
  • The model significantly reduces computational time, enabling rapid forecasting compared to traditional methods.
Read more
BearingNAS: Obtaining In-Sensor Intelligent Fault Diagnosis Systems for Bearings Using a Laptop
Andrea Mattia Garavagno, Edoardo Ragusa, Paolo Gastaldo, Antonio Frisoli, Rodolfo Zunino
Efficient ML Optimization Time Series
  • Introduction of BearingNAS, a framework for in-sensor fault diagnosis.
  • Elimination of reliance on GPUs by conducting searches on a laptop CPU.
  • Achieves 99.50% diagnostic accuracy on the CWRU benchmark.
  • Optimized for resource-constrained microcontrollers and ISPUs.
Read more
Robust Multi-View Classification under Noisy Supervision via Global Anchor Consensus
Yuliang Yang, Hongzhe Zhang, Huiru Wang
Multimodal
  • GALA is the first method to audit noisy labels in multi-view classification using global class anchors and cross-view consensus.
  • An adaptive label correction strategy combines soft sample weighting with a conservative rewriting rule.
  • GALA shows superior performance, ranking first in 33 out of 36 experimental settings.
  • The method improves label quality and representation learning iteratively during training.
Read more
Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning
Stella Ho, Joel Villalobos, Joseph West, Jingyang Liu, Weijie Qi, Haruhiko Kishima, Ryohei Fukuma, Takufumi Yanagisawa, Sam E. John, David B. Grayden
Time Series Multimodal Interpretability
  • Demonstrates the potential of end-to-end deep learning for visual semantic decoding from ECoG data.
  • Identifies the Transformer architecture and high-gamma frequency inputs as optimal for decoding performance.
  • Highlights the importance of specific cortical regions in visual processing and decoding.
  • Utilizes data augmentation techniques to enhance model performance despite limited training samples.
Read more
Decafs: Disentangled Conditional adversarial Flows
Anirudh Jain, Sakshi Varshney, Samuel Kaski, Vikas Garg
Generative Models
  • Introduction of a novel adversarial conditional flow framework that disentangles latent spaces for controlled generation.
  • Utilization of a Lie algebra-based generator to learn both disentangled and coupled components of the latent space.
  • Empirical validation of DECAFS in generating high-quality conditional outputs in both image and molecular domains.
  • Demonstration of improved interpretability and control over generative factors in tasks such as drug discovery.
Read more
Scalable Causal Imitation Learning
Eylam Tagor, Mingxuan Li, Elias Bareinboim
Reinforcement Learning Robotics Theory
  • Introduction of Causal SQIL and Causal IQ-Learn for continuous control tasks with unobserved confounders.
  • Development of an efficient approximation of the sequential π-backdoor criterion for causal adjustment.
  • Empirical results show significant performance improvements over prior CIL algorithms in long-horizon tasks.
  • Causal SQIL and Causal IQ-Learn can achieve competitive success rates, sometimes exceeding expert performance.
Read more
Effects of width-dependent model hyperparameters and ℓ2-regularization on the loss landscape of two-layer ReLU networks
Haruka Eshima, Makoto Yamada
Theory Optimization
  • Derivation of conditions for global minima collapse to zero in two-layer ReLU networks under ℓ2-regularization.
  • Demonstration that AdamW optimizer prevents parameter collapse, unlike SGD.
  • Analytical solutions for globally optimal parameters in one-dimensional input settings.
  • Width-invariant connectivity effects of ℓ2-regularization, with stronger dimensionality reduction as width increases.
Read more
Toward Auditable Fraud Detection: Combining Graph Features, Model Explanations, and Agentic Case Investigation
Rahil Sharma
Graph Learning Interpretability
  • The study highlights the importance of explainability and auditability in fraud detection systems.
  • Graph-derived features and anomaly signals improve detection in specific cases but not overall performance.
  • The investigation agent's reliance on model explanations does not consistently lead to better decision-making.
  • An escalation rule based on disagreement between the agent and classifier was ineffective in identifying correct decisions.
Read more
Probabilistic Physics-Aware Machine Learning Predictions of Electric Truck Energy Consumption with Field Data
Hannes Nilsson, Rafael Basso, Balázs Kulcsár, Morteza Haghir Chehreghani
Theory Optimization
  • Incorporation of physics principles improves energy consumption predictions for electric trucks.
  • Bayesian linear regression outperforms standard linear regression in reliability.
  • Complex models like neural networks and gradient boosted trees achieve higher accuracy with physics-aware inputs.
  • Uncertainty estimation is integrated into the prediction framework, enhancing decision-making under uncertainty.
Read more
FlashPDE: A Drop-in Fused Triton Operator Library for Neural PDE Solvers
Peiyu Zang, Bosen Xie, Ruoxiang Xu, Yongqiang Cai
Efficient ML
  • FlashPDE provides a drop-in library of 14 differentiable PDE operators for efficient SciML applications.
  • The library reduces memory usage from over 40 GB to 3.4 GB for complex simulations while maintaining performance.
  • It achieves significant speedups (up to 2.30×) over traditional PyTorch implementations by minimizing kernel launches.
  • FlashPDE supports various neural network architectures without requiring changes to the underlying models.
Read more
Attractor Geometry Determines the Identifiability Limits of System Discovery
Matteo Gallo, Fabio Anselmi, Paolo Lazzari
Theory
  • The geometry of the attractor significantly influences the identifiability limits of system discovery.
  • λmin(M), the smallest eigenvalue of the invariant-measure moment matrix, serves as a key metric for recovery potential.
  • Recovery is hardest in fixed-point regimes, intermediate in limit cycles, and easiest in chaotic regimes.
  • Soft F1 is introduced as a new metric for evaluating algorithm performance, capturing partial structural recovery.
Read more
Relative Positions Generalize, Absolute Positions Memorize: An Implicit-Bias Account of Length Generalization in Attention
Subham Singh, Ashutosh Mishra, Subha Raut
NLP Theory Optimization
  • Relative positional encodings (RoPE) enable transformers to generalize to longer sequences, while absolute positional encodings (APE) do not.
  • The authors provide an optimization explanation for this phenomenon, focusing on the implicit bias of attention heads.
  • RoPE maintains high accuracy across varying lengths, whereas APE's accuracy drops significantly when encountering unseen lengths.
  • The study characterizes the learned rotary rule as a low-rank carrier kernel aligned with the target offset.
Read more
Edge-Efficient Transformer for End-to-End RF Spectrum Monitoring
Zhifan Song, Haralampos-G. Stratigopoulos, Hassan Aboushady
Efficient ML
  • Introduction of E-SpecFormer, an edge-efficient Transformer for RF spectrum monitoring.
  • LiTAN mechanism reduces complexity and improves accuracy without Softmax and LayerNorm.
  • Four scalable model variants to accommodate diverse hardware constraints.
  • Achieved high accuracy on modulation recognition and covert channel detection tasks.
Read more
ChemFusion: A Multimodal Cross-Attention Network for Reaction Yield Prediction
Qiwei Han, Chi Zhou
Multimodal
  • ChemFusion integrates electronic features with 3D spatial data using a cross-attention mechanism.
  • The model significantly outperforms traditional unimodal frameworks in predicting reaction yields.
  • Attention matrices reveal the model's ability to identify and penalize steric clashes, enhancing interpretability.
  • Achieved state-of-the-art results on the Buchwald-Hartwig benchmark dataset.
Read more
What does a Bayes-filtered transformer believe? A predictive Monte Carlo approach
Afiq Abdillah Effiezal Aswadi, Haotong Ma, Susan Wei
NLP Large Language Models Interpretability
  • Introduces the Bayes-filtered transformer (BFT) concept for understanding latent task beliefs.
  • Utilizes predictive Monte Carlo (PMC) as a novel interpretability tool for BFTs.
  • Demonstrates that BFTs can approximate their pretraining prior and posterior distributions.
  • Applies PMC to various task families, revealing consistent latent phenomena.
Read more
From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning
Garvit Singla, Uma Maheswari Natarajan, Raghuram Bharadwaj Diddigi
Reinforcement Learning NLP Robotics
  • LA-MAML leverages natural language task descriptions for efficient task adaptation in reinforcement learning.
  • The framework replaces traditional trajectory collection and gradient updates with a learned embedding of task instructions.
  • Empirical evaluations on the BabyAI benchmark show LA-MAML's improved performance and reduced training time compared to baselines.
  • The study includes an ablation analysis to assess the role of language in task adaptation.
Read more
Predicting Activities in Aqueous Electrolyte Solutions with Hybrid Machine Learning
Zeno Romero, Maximilian Kohns, Fabian Jirasek
Theory Optimization Efficient ML
  • Introduction of a hybrid model (Bromley-MCM) combining physics-based and machine learning approaches.
  • Utilization of matrix completion to predict parameters for unstudied electrolytes.
  • Training on a large dataset (478 electrolytes) to enhance predictive accuracy.
  • Ability to predict activities for 9,296 electrolytes, significantly broadening the model's applicability.
Read more