AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

66 Papers today
8h Update frequency
7 Days of history
MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models
Ashitabh Misra, Madhav Agrawal, Arham Jain, Tarek Abdelzaher
NLP Large Language Models Efficient ML
  • MixQuant adapts mixed-precision quantization to varying deployment budgets without needing per-budget calibration.
  • The framework considers the sensitivity of layers based on the quantization levels of upstream layers.
  • MixQuant incorporates a tail regularization term to optimize bit allocation and improve model performance.
  • It outperforms existing methods in terms of accuracy and perplexity across various large language models.
Read more
A Physics-Informed Neural Operator for Thermal Ranking of Low-Cost Wall Materials in Hot-Dry Climates
Muhammad Akbar Khan, Fahim Raees, Ubaida Fatima
Theory
  • Introduces a two-stage computational framework for thermal ranking of indigenous wall materials.
  • Utilizes a Physics-Informed Neural Operator (PINO) for efficient learning of thermal performance.
  • Demonstrates that clay-straw adobe is the most effective low-cost material for thermal comfort.
  • Identifies a regime boundary that affects material ranking based on outdoor temperature conditions.
Read more
AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction
Ziheng Zhou, Huiyu Luo, Xiaohu Zhu, Nan Wang, Xuebiao Qin, Chaoyan Zhang, Jun Yan
Theory
  • AMPBench-MT integrates multiple evaluation endpoints for antimicrobial peptides, including potency, spectrum, and safety.
  • The benchmark employs a homology-controlled protocol to ensure robust evaluation across diverse peptide sequences.
  • High binary classification performance does not guarantee effective assay outcomes, highlighting the need for endpoint-aware assessments.
  • The study reveals that traditional performance metrics can be misleading, particularly in the context of scarce negative data.
Read more
Behavior-Driven Explainability
Caroline Dominik, Rolf Drechsler
Interpretability
  • Introduces Behavior-Driven Explainability (BDX) as a method for deriving explanations from system specifications.
  • Utilizes Behavior-Driven Development (BDD) to create structured scenarios that outline expected system behavior.
  • Demonstrates the application of BDX through a case study on a RISC-V processor, highlighting its effectiveness in explaining system exceptions.
  • Formalizes the steps for implementing BDX, providing actionable guidance for developers.
Read more
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning
Wenwu Fan, Qihong Lin, Zhijie Xia, Zhuo Zheng, Sihao Wang, Qiang Chen, Liangsheng Zhu
Reinforcement Learning Large Language Models Efficient ML
  • ACRL framework stabilizes RL training by controlling training-inference discrepancies.
  • The method enhances exploration and improves accuracy through adaptive token-level gradient adjustments.
  • Empirical results show ACRL matches BF16 accuracy and outperforms importance sampling fixes under FP8 quantization.
  • The approach addresses the challenges posed by architectural separations in training and inference engines.
Read more
Monitoring Post-Disaster Urban Recovery Using High-Resolution SAR Time Series and Unsupervised Learning: Evidence from the 2023 TΓΌrkiye-Syria Earthquake
Luigi Russo, Deodato Tapete, Silvia Liberata Ullo, Paolo Gamba
Time Series
  • Proposes an unsupervised framework for monitoring urban recovery using SAR time series.
  • Identifies heterogeneous reconstruction dynamics in cities affected by the 2023 TΓΌrkiye-Syria earthquakes.
  • Demonstrates the complementary nature of SAR and nighttime lights data in assessing recovery.
  • Highlights the effectiveness of SAR-based anomaly detection in capturing early-stage reconstruction processes.
Read more
Multiclass Classification without Labels via Posterior Simplex Geometry
RaphaΓ«l Bonnet-Guerrini, Johann Ioannou-Nikolaides, Troels Petersen, Vincenzo Piuri
Theory
  • Extension of CWoLa to multiclass classification without instance-level labels.
  • Establishment of a (K-1)-simplex geometry for Bayes-optimal mixture classifiers.
  • Development of prior-free recovery methods for latent class structures.
  • Demonstrated effectiveness on multiple datasets, improving performance in label-scarce domains.
Read more
Spend Experts Where You Are Unsure: Confidence-Adaptive Routing for Mixture-of-Experts LoRA
Tom Saliencro, Rohan Desai, Priya Nair, Maya Lindqvist, Daniel Whitmore
NLP Large Language Models Efficient ML
  • CARE adapts the number of active experts based on token uncertainty, improving computational efficiency.
  • The method uses existing router output distributions as signals for confidence and ambiguity.
  • CARE achieves better accuracy with fewer experts compared to traditional fixed top-k approaches.
  • The approach is parameter-free and can be easily integrated into existing MoE-LoRA architectures.
Read more
Human Preference aligned Tabular Similarity
Frederik Hoppe, Astrid Franz, Marianne Michaelis, Lars Kleinemeier, Udo GΓΆbel
Theory
  • Current tabular embedding methods focus on prediction tasks, neglecting human preference alignment for similarity search.
  • Standard evaluation metrics fail to capture the nuances of human-perceived similarity in industrial applications.
  • The proposed workflow integrates human feedback into the evaluation of tabular embeddings for improved trustworthiness.
  • Human preference alignment is crucial for ensuring that retrieved similar items are meaningful to domain experts.
Read more
Learning from the Descent Direction: Adaptive Gradient Descent under One-Sided HΓΆlder Regularity
Arzu Ahmadova, Ismail Huseynov
Optimization Theory
  • Introduction of a one-sided HΓΆlder descent model that focuses on directional curvature.
  • Development of an adaptive step-size rule based on local directional curvature and current gradient norm.
  • Establishment of a best-iterate stationarity guarantee with a rate determined by the HΓΆlder exponent.
  • Comparison of the proposed method against various baseline gradient descent strategies.
Read more
Inverse RL Helps Align AI by Imitating Humans
MichaΕ‚ WiliΕ„ski, Liu Leqi, Chirag Nagpal
NLP Large Language Models Reinforcement Learning
  • PARED introduces a method for deriving explicit, inspectable rewards from expert demonstrations without needing task-specific annotations.
  • The approach allows for both inference-time selection and on-policy reinforcement learning optimization.
  • PARED supports audience-conditioned alignment, demonstrating that improvements can be achieved for multiple target audiences simultaneously.
Read more
Reinformed Dreamer: An Asymmetric World Model Efficiently Trained through Latent Guidance
Gaspard Lambrechts, Adrien Bolland, Daniel Ebi, Damien Ernst
Reinforcement Learning
  • Introduction of the Reinformed Dreamer algorithm, which improves upon the Informed Dreamer.
  • Identification of limitations in the learning objectives of existing asymmetric reinforcement learning models.
  • Development of a new learning objective using a privileged variational encoder for better representation learning.
  • Demonstration of improved performance and convergence speed across multiple benchmarks.
Read more
Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
Shuai Wang, Daoan Zhang, Zhe Tang, Hao Cheng, Jiaheng Wei
Multimodal
  • NOPD enables self-improvement of VLMs without external supervision.
  • The method utilizes prediction discrepancies between clean and corrupted inputs for self-supervision.
  • NOPD outperforms traditional reinforcement learning and distillation methods in visual reasoning tasks.
  • Significant performance gains were observed on multiple benchmarks, demonstrating generalization capabilities.
Read more
Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension
Jinhao Zhang, Zeyu Liu, Zicheng Yan, Yunquan Zhang, Guangming Tan, Fangming Liu, Daning Cheng
Theory Optimization
  • Introduction of the effective alignment dimension as a measure of signal-noise geometry in neural networks.
  • Derivation of a finite-sample certificate for directional stability in model expansion, independent of previous assumptions.
  • Empirical validation showing that wider models exhibit better alignment and lower misalignment.
  • Direct interventions confirm that alignment statistics can predict changes in held-out loss.
Read more
DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment
Shiyu Teng, Haichen Yu, Jiaqing Liu, Hao Sun, Yu Song, Shurong Chai, Ruibo Hou, Lanfen Lin, Yen-Wei Chen
Multimodal
  • DynaBridge formulates mental health assessment as a structured ordinal cross-task problem, linking item responses to risk predictions.
  • The framework integrates multimodal data (acoustic, visual, textual) with LLM-generated summaries to enhance predictive accuracy.
  • It introduces DASS-aware item-to-risk reconstruction and confidence-aware refinement to improve consistency in predictions.
  • DynaBridge outperforms existing methods in predicting depression, anxiety, and stress risk, demonstrating its effectiveness in mental health assessment.
Read more
Joint Flow Matching for Generator-Consistent Classification
Hayden McAlister, Lech Szymanski
Generative Models Theory Interpretability
  • Introduction of Joint Flow Matching (JFM) for consistent generative and discriminative inference.
  • JFM assigns opposite roles to noise and data distributions, enabling bidirectional inference.
  • The framework guarantees consistency between generative and discriminative conditional distributions.
  • Empirical validation shows competitive performance on generation and classification tasks.
Read more
SPARC Segmentation to Prediction via Affine Regression and Counterfactuals
Shivani, Subhayan Roy
Optimization Theory Interpretability
  • Introduces a novel propensity modeling framework for B2B e-commerce that addresses unique challenges in transaction prediction.
  • Replaces SMOTE with a counterfactual data generation approach (DiCE) for better representation of minority-class samples.
  • Utilizes the PyPARC framework to provide calibrated propensity probabilities for customer segmentation.
  • Achieves significant improvements in precision over traditional methods, validating the effectiveness of the proposed approach.
Read more
Label-free Industrial Fault Detection via Adversarial Inverse Reinforcement Learning: A System for Run-to-Failure Prognostics
Dhiraj Neupane, Mohamed Reda Bouadjenek, Richard Dazeley, Sunil Aryal
Reinforcement Learning Time Series Optimization
  • Introduces AIRL as a novel approach for label-free fault detection in industrial machinery.
  • Demonstrates the ability to detect gradual degradation effectively without relying on labeled fault data.
  • Achieves superior performance in terms of fault detection consistency across multiple datasets.
  • Provides practical recommendations for automatic thresholding strategies in deployment.
Read more
Breaking the Periodicity Assumption: Robust Tensorial Multi-View Clustering via Graph-Spectral Low-Rank Learning
Jintian Ji, Xingsu Li, Songhe Feng
Graph Learning Theory Efficient ML
  • Existing t-SVD-based TMC methods are sensitive to sample ordering due to an implicit periodicity assumption.
  • The proposed graph-spectral low-rank tensor learning framework addresses this sensitivity by using a data-driven graph Fourier Transform.
  • An anchor-based variant is introduced to improve scalability for large datasets.
  • Extensive experiments show competitive or superior performance compared to existing TMC methods.
Read more
Capacity-Aware Deep Learning for Generalizable Traffic Volume Estimation Across Links and Cities
LΓ©o Hein, Giovanni De Nunzio, AurΓ©lie Pirayre, Laurent Najman
Graph Learning Time Series Theory
  • Proposes a link-level learning framework for traffic volume estimation using widely available data.
  • Introduces a capacity-aware formulation that incorporates traffic-theoretic constraints.
  • Evaluates the model under intra-network and inter-network generalization settings.
  • Demonstrates superior performance compared to state-of-the-art methods in spatial distribution shifts.
Read more
XGRVFL-MV: Residual-Coupled Graph-Embedded Multi-View Random Vector Functional Link Network with FleXi Guardian Loss
Yogesh Kumar, Mudasir Ganaie
Graph Learning Multimodal Optimization
  • Introduction of a novel multi-view RVFL model that incorporates graph embedding and residual coupling.
  • Utilization of FleXi Guardian loss to mitigate the impact of large prediction residuals.
  • Demonstrated competitive performance on multiple benchmark datasets.
  • Incorporation of geometric structure preservation through graph regularization.
Read more
An Embarrassingly Simple Rule-based Visiting Circulation Approach to Trip Destination Prediction
Eng-Shen Tu, Yong-Han Chen, En-Chao Liu, Hao-Yun Keng, Cheng-Te Li
Theory
  • Introduction of the Rule-based Visiting Circulation (RVC) model for trip destination prediction.
  • The model effectively utilizes origin information and individual trip behaviors without needing training data from the target area.
  • RVC outperforms supervised learning methods and heuristics in prediction accuracy.
  • The approach addresses significant challenges in destination prediction, including dataset variability and geographic complexity.
Read more
Quantum Speedups for Stochastic Optimization with Heavy-Tailed Noise
Bin Luo, Chengchang Liu, Jonathan Allcock, Shengyu Zhang, John C.S. Lui
Optimization Theory
  • Introduction of quantum mean estimators for multivariate heavy-tailed distributions.
  • Development of QNSGD and QPSGD methods that outperform classical algorithms in query complexity.
  • Establishment of quantum lower bounds demonstrating optimality of the proposed estimators.
  • Addressing the limitations of classical stochastic optimization methods under heavy-tailed noise.
Read more
Eliminating Propagation Delay: Attention-Based Spatial-Temporal Fusion Graph Convolution Network for Traffic Flow Prediction
Jinpeng Chen, Ziyu Yu, Tao Wang, Jun Ma, Hongbo Gao, Senzhang Wang, Zufeng Zhang, Kaimin Wei
Graph Learning Time Series Optimization
  • Introduces A-STFGCN to address propagation delays in traffic flow prediction.
  • Utilizes a spatial-temporal fusion block to enhance feature extraction.
  • Implements a multi-head self-attention mechanism for temporal characteristics.
  • Achieves superior performance on real-world datasets compared to existing methods.
Read more
Reinforcement Learning for Code Optimization
Pierre Chambon, Kunhao Zheng, Juliette Decugis, Benoit Sagot, Gabriel Synnaeve
Reinforcement Learning Optimization
  • Reinforcement learning can be effectively adapted for code optimization by addressing measurement noise and reward sparsity.
  • The DMC-Optim framework introduces a structured approach to testing, reward composition, and model adaptation.
  • Significant improvements in optimization performance were achieved, with notable gains in pass rates for both correctness and speed.
  • The methodology preserves correctness while enhancing execution speed, demonstrating the feasibility of combining these objectives in RL.
Read more
Score-Based Stabilization for Time-Dependent Problems
Eshed Gal, Eldad Haber, Uri Ascher
Theory Generative Models Optimization
  • Introduction of a two-stage stabilization framework that integrates a learned score-based correction into time-stepping schemes.
  • The score correction effectively mitigates nonphysical behaviors and instabilities in numerical simulations of PDEs.
  • Demonstration of conditional manifold stability, allowing for larger timesteps without traditional restrictions.
  • Validation of the approach across multiple PDEs, showcasing its versatility and robustness.
Read more
Detecting CSAM Text-to-Image LoRAs From Weights
David Demitri Africa, Cate Heine, Nadine Staes-Polet, Kimberly Mai
Generative Models Computer Vision Interpretability
  • Introduces u1, a representation for detecting harmful LoRAs from their weights.
  • Demonstrates that metadata for LoRAs is often incomplete and unreliable.
  • Shows that the method can generalize across different base models and datasets.
  • Proves robustness to noise and precision reduction in weight analysis.
Read more
From Deterministic to Generative Deep Learning for Urban Air Quality Reconstruction from Sparse Observations
Abhishek A.Sabnis, Mihai Mitrea, Lya Lugon, Karine Sartelet, Marc Bocquet, Xiaoyuan Cheng, Shupeng Zhu, Sibo Cheng
Generative Models
  • Introduction of a diffusion-based generative model for air quality reconstruction.
  • Joint modeling of multiple pollutants to capture inter-pollutant correlations.
  • Evaluation on real-world data demonstrates high accuracy and generalization.
  • Data augmentation methods facilitate transferability to real-world observations.
Read more
Stable FP4 Training via Transposition-Invariant Block Quantization
Mehdi Rahimifar, Amin Darabi, Mehran Taghian Jazi, Xing Huang, Yao Wang, Zhijun Tu, Yufei Cui, Yunke Peng, Hongliang Li
Large Language Models Efficient ML Optimization
  • Identified transposition-induced scale inconsistency as a major source of instability in FP4 training.
  • Proposed a transposition-invariant FP4 quantization framework using 2D block quantization.
  • Combined FP4 linear layers with MXFP8 attention for practical mixed-precision training.
  • Achieved stable training for models up to 30 billion parameters with minimal performance degradation.
Read more
Local Regularization Does Not Characterize Multiclass PAC Learnability
Eric Hou
Theory
  • Local regularization does not characterize multiclass PAC learnability.
  • A countable hypothesis class with low complexity cannot be learned by local regularizers.
  • The construction involves tournament structures that introduce complexity through cyclic relationships.
  • The findings challenge existing assumptions about the effectiveness of local regularization in multiclass settings.
Read more
Random Forest-Based Prediction of Bone Volume Fraction and Fracture Position from S-Parameters
Jianhe Li, Jinsui Meng, Yida Zhao, Zihe Wang, Liaoran Sun, Tao Shan
Theory Efficient ML
  • Introduces a Random Forest model for predicting bone volume fraction and fracture position.
  • Utilizes a nine-antenna microwave scanning system to gather S-parameter data.
  • Demonstrates the effectiveness of microwave sensing as a non-invasive alternative to traditional imaging techniques.
  • Validates the method through both synthetic simulations and physical experiments with bone-mimicking phantoms.
Read more
LLM as Forecasting Planner: Training-Free Text Conditioning for Time-Series Foundation Models
Huu Hiep Nguyen, Dung Nguyen, Minh Hoang Nguyen, Dai Do, Hung Le
Time Series Large Language Models Optimization
  • LAFP separates numerical generation and contextual reasoning, enhancing time-series forecasting accuracy.
  • The framework uses Monte Carlo Tree Search (MCTS) to explore candidate trajectories effectively.
  • LAFP shows consistent improvements over context-blind TSFMs and direct LLM forecasters across multiple datasets.
  • The approach is training-free, allowing for easy integration with existing TSFMs and LLMs.
Read more
Rethinking Likelihood distributions: Student's t Likelihood Boosts Bayesian Neural Network Performance
Pei-Hsuan Hsia, Lars H. Heyen, Arvid Weyrauch, Markus Goetz, Achim Streit, Sebastian Krumscheid, Charlotte Debus
Theory
  • Student's t likelihood distribution outperforms Gaussian likelihood in BNNs.
  • The choice of likelihood distribution significantly affects predictive accuracy.
  • Student's t distribution can lead to shorter training times while being easy to implement.
  • The study fills a gap in the literature regarding the exploration of likelihood distributions in BNNs.
Read more
GAUGE: Grading Agent-Built Financial Models Without a Golden Answer
Jiacheng Lu, Sinuo Wang, Wentao Zhao, Rui Sun, Cheng Hua, Tao Song, Hui Cai, Beidi Luan, Zhengze Wu, Lingjing Teng, Yijia He, Jing Li, Daxin Jiang, Zuo Bai, Haibing Guan
Large Language Models Optimization Theory
  • GAUGE provides a more realistic benchmark for evaluating financial models by considering multiple analyst outputs instead of a single expert reference.
  • The study reveals significant discrepancies in financial model outputs among professional analysts, questioning the validity of point-tolerance grading.
  • GAUGE incorporates a multi-layered scoring system that evaluates both mechanical correctness and judgment-based facets.
  • Results show that current AI agents perform better in model construction than in valuation judgment compared to human analysts.
Read more
K-Survival Means
Abdallah Alabdallah
Optimization Theory Time Series
  • K-SurvMeans explicitly incorporates survival outcomes into the clustering process.
  • The method uses Particle Swarm Optimization to solve the non-differentiable optimization problem.
  • It operates in a learned low-dimensional latent space to improve cluster separation and efficiency.
  • Experiments show superior performance in cluster separation compared to existing methods.
Read more
FinAbstain: Uncertainty-Calibrated Multimodal RAG for Selective Financial Forecasting
Dorothy Torres, Wei Cheng, Henan Huang
NLP Large Language Models Multimodal
  • Introduces a multimodal RAG framework for financial forecasting that incorporates uncertainty calibration.
  • Utilizes a point-in-time retrieval mechanism to ensure only relevant, publicly available data is used.
  • Develops a hybrid uncertainty score to guide selective predictions and abstentions.
  • Demonstrates that calibrated abstention can reduce errors and drawdowns in financial predictions.
Read more
Multi-Agent Privacy Game in Federated Learning: A Unified Mean-Field View
Kun Zhao, Xu Chen
Federated Learning Theory Optimization
  • Introduces a mean-field privacy game framework for privacy-preserving federated learning.
  • Achieves tractable equilibrium for arbitrarily many clients while accommodating heterogeneous privacy preferences.
  • Recovers existing privacy-preserving methods as special cases of the proposed framework.
  • Demonstrates effective privacy-utility trade-offs through empirical experiments.
Read more
Contextual Deconvolution for Variance-Stable Demand Sensing: Kernel-Modulated Operators in Promotional Retail
Mohammad Forouhesh
Time Series
  • Introduction of Contextual Deconvolution (CD) for improved demand sensing in promotional retail.
  • CD utilizes a kernel-modulated banded operator to separate transient shocks from structural demand.
  • The method is scalable and does not require per-SKU training, making it suitable for large retail catalogs.
  • Empirical results show significant reductions in safety stock and order variance compared to traditional methods.
Read more
Contrastive Representation Learning of Longitudinal Disease Trajectories on Temporal Graphs
Bastian Pfeifer
Graph Learning Time Series
  • Introduces a contrastive representation learning framework for longitudinal disease trajectories.
  • Models disease trajectories as temporal graphs with nodes and edges representing observations and relationships.
  • Utilizes structure-aware random walks for effective contrastive learning.
  • Achieves improved clustering performance on real-world biomedical datasets.
Read more
Mind the Missing Split: Resolving Feature Heterogeneity in Swarm Learning with Random Forests
Mohammad Tajabadi, Dominik Heider
Federated Learning
  • Addresses feature heterogeneity in Swarm Learning with Random Forests.
  • Proposes inference-time strategies to resolve missing splits without restricting feature training.
  • Evaluates methods on nine datasets, showing improved performance over traditional approaches.
  • Highlights the importance of decentralized learning in maintaining data privacy.
Read more
Charging Phase Health Indicators for Battery State-of-Health Estimation: A Systematic Comparison of CC, CV, and Combined Approaches under Cross-Battery Validation
Huy Hoang Le, Kim-Anh Nguyen
Time Series Interpretability Efficient ML
  • Combined CC+CV indicators yield the highest SOH estimation accuracy.
  • A significant performance gap exists between standard cross-validation and LOBO validation.
  • SHAP analysis identifies the CV-to-CC time ratio as a critical indicator for SOH estimation.
  • The study provides practical guidelines for indicator selection under data constraints.
Read more
Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning
Luc McCutcheon, Evangelos Chatzaroulas, Saber Fallah
Reinforcement Learning Optimization Robotics
  • CPR is introduced as a solution to prevent policy collapse in continual RL by selectively adjusting low-utility neurons.
  • The method maintains a balance between network plasticity and stability, avoiding the pitfalls of binary resets.
  • CPR outperforms existing methods in multiple benchmarks, demonstrating its effectiveness in long-term training scenarios.
  • Ablation studies reveal a tunable trade-off between plasticity and peak performance, emphasizing the utility-scaled approach.
Read more
Data-Dependent Regret and Polyak Corrections for Constrained Online Convex Optimization
Wentao Zhang
Optimization Theory
  • Introduces a refined regret analysis for constrained OCO using data-dependent quantities.
  • Proposes AdaOGD-PFS, an adaptive-step-size algorithm that achieves improved regret bounds.
  • Demonstrates significant performance improvements in experiments, with reductions in regret bounds by 38-43%.
  • Identifies the Polyak correction term as a crucial factor in achieving tighter regret bounds.
Read more
When Does Deep Representation Learning Help Single-Cell Clustering? A Sensitivity-Aware Diagnostic Benchmark for Biomedical AI Pipelines
Nguyen Thanh Phong, Truong Viet Vu, Nguyen Ha Thu, Tran An Ky, Tran Hoang Thong, Le Pham Thuy Hien, Nguyen Thai Anh
Generative Models Optimization Efficient ML
  • Deep representation learning can enhance clustering performance but is not universally superior to classical methods.
  • The study identifies three distinct dataset regimes where different methods excel.
  • Learning rate and latent dimensionality are critical hyperparameters influencing clustering outcomes.
  • A dataset-aware and compute-conscious framework is proposed for optimizing biomedical AI pipelines.
Read more
Explainable AI for Chronic Kidney Disease Prediction Using Simulated Federated Learning
Md Zahid Hasan Ontor, Md Al Amin, Anik Dev Nath, Bikash Kumar Paul
Federated Learning Interpretability
  • Federated Learning was utilized to ensure patient data privacy while predicting CKD.
  • A VotingClassifier combining Random Forest, AdaBoost, and XGBoost was employed for model prediction.
  • GridSearchCV was used for optimizing model performance on the client side.
  • Explainable AI techniques, particularly LIME, were integrated to enhance model interpretability.
Read more
EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings
Xiaocheng Fang, Jieyi Cai, Guangkun Nie, Haoyu Wang, Jiarui Jin, Yujie Xiao, Bo Liu, Chenyang He, Qinghao Zhao, Gaofeng Cheng, Hongyan Li, Shenda Hong
Multimodal
  • Introduction of EchoBridge for ECG-echocardiography text alignment.
  • Utilization of CSPP and APBC to address cross-modal representation interference and long-tailed distributions.
  • Significant performance improvements over baseline methods in classification metrics.
  • Effective in identifying low-prevalence cardiac findings.
Read more
Raven: High-Recall Sequence Modeling with Sparse Memory Routing
Arshia Afzal, Aviv Bick, Eric P. Xing, Volkan Cevher, Albert Gu
NLP Large Language Models Efficient ML
  • Raven introduces Routing Slot Memories (RSMs) to enhance long-context recall in sequence models.
  • The model selectively updates and decays memory slots, reducing interference and improving token recoverability.
  • Raven outperforms existing linear-time models on recall-intensive benchmarks, maintaining effectiveness even with extended context lengths.
  • The architecture simplifies prior models and does not depend on convolutional or sliding-window attention mechanisms.
Read more
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato
NLP Large Language Models Optimization
  • Introduces the first input-only method for minimizing targeted internal latents in LLMs.
  • Develops a behavioral + erasure protocol to differentiate genuine behavioral changes from mere activation drops.
  • Demonstrates successful suppression of evaluation-awareness latents across multiple constructions.
  • Finds that suppression does not equate to behavioral control, challenging assumptions about model evaluation.
Read more
Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework
Zeki Doruk Erden
Computer Vision Theory Interpretability
  • Introduces a multi-scale structural feature representation for visual recognition.
  • Demonstrates continual learning without destructive adaptation or reliance on past data.
  • Achieves higher accuracy on class-incremental MNIST compared to traditional methods.
  • Retains previously learned classes while integrating new observations.
Read more
Generator-Aligned Representation Interfaces for Diagnostic Soft Equivariance
Weitao Li, Gong Cheng
Computer Vision Multimodal Theory
  • Introduction of the Generator-Aligned Representation Interface (GARI) for soft equivariance.
  • GARI allows for the integration of transformation generators into generic sequence backbones.
  • Direct Equivariance Error (DEE) is proposed as a measure for assessing representation-level soft equivariance.
  • Experiments show improved task performance and representation consistency across various data modalities.
Read more
Perturbative-NeuSA: A Structured Spectral Framework for Time-Dependent PDEs
Xianli Zhu, Jia Yin
Theory Efficient ML Interpretability
  • Perturbative-NeuSA separates the solution into a low-fidelity background and a high-resolution perturbation, optimizing learning efficiency.
  • The framework outperforms the baseline NeuSA without requiring neural network training, particularly excelling in the Burgers equation.
  • The effectiveness of the neural closure is conditional, varying with background fidelity and residual organization.
  • The method provides a structured approach to diagnosing the contributions of different components in the solution process.
Read more
Semantic Space Search Trajectory Networks
Julian Agudelo, Alberto Tonda, Gabriela Ochoa, Vincent Guigue, Cristina Manfredotti, Evelyne Lutton
Optimization Theory
  • Introduction of a clustering-based discretization strategy for constructing STNs in high-dimensional semantic spaces.
  • Enables qualitative and quantitative comparisons of learning dynamics across different algorithm families.
  • Reveals systematic differences in learning dynamics between standard training and label randomization in neural networks.
  • Semantic Space STNs provide insights into the interaction between learning algorithms and data.
Read more
A Coulomb Particle Model for Learning Kernel Attention in Transformers
Masoud Badiei Khuzani, Sharath Honnaiah, Atiq Islam, Alex Cozzi, Abraham Bagherjeiran
NLP Theory Efficient ML
  • Introduces a particle-based method for learning feature distributions in kernel attention.
  • Utilizes a Hamiltonian framework to optimize kernel-target alignment.
  • Demonstrates improved performance in accuracy and robustness on NLP benchmarks.
  • Maintains linear inference complexity while enhancing attention mechanisms.
Read more
Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
Xinyi Hong, Pinjun Dong, Xinyang Yu, Binyan Jiang
NLP Large Language Models Graph Learning
  • HYSET treats tool retrieval as a set-level problem, allowing for joint evaluation of tool sets.
  • The method captures size-dependent compatibility among tools through cardinality-specific interactions.
  • HYSET functions as a pre-selection module, requiring no changes to downstream LLM agents.
  • Experimental results indicate significant improvements over existing retrieval methods in both accuracy and task success.
Read more
Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions
Quyen Tran, Hai Nguyen, Quan Dao, Zhuowei Li, Nam Le, Trung Le, Dimitris Metaxas
Theory Efficient ML
  • Identification of spectral collapse as a critical issue in RLS-based ACL methods under long-tailed distributions.
  • Introduction of Geometry-Spectral Rectification (GSR) as a solution that selectively inflates eigenvalues for tail classes.
  • Theoretical proof of GSR's effectiveness in improving numerical stability and stable rank of the Gram matrix.
  • Demonstration of GSR's superior performance compared to existing ACL methods on benchmark datasets.
Read more
Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models
Malena Loza, David Chushig-Muzo, Eva Milara, Luis Bote-Curiel, Luis Estrada-Petrocelli, Felipe Grijalva
Theory Efficient ML Interpretability
  • First empirical evaluation of TFMs under distribution shifts across three real-world datasets.
  • Systematic performance degradation of TFMs under OOD conditions, regardless of pre-training strategy.
  • Identification of a scalability gap in high-performing models requiring significant computational resources.
  • Extension of OOD analysis from traditional models to TFMs, highlighting the relationship between in-distribution and OOD performance.
Read more
Physics-Informed Neural Networks for Predicting Nitrous Oxide Flux
Freddy Yu, Jashanjeet Kaur Dhaliwal, Subhadeep Chakraborty
Theory
  • Introduces Physics-Informed Neural Networks (PINNs) for predicting N2O emissions.
  • Demonstrates significant performance improvement over traditional models like Cycles.
  • Finds that physics constraints enhance robustness but may degrade accuracy on familiar data.
  • Highlights the challenges of cross-site generalization in N2O prediction.
Read more
Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls
Md Ashikur Rahman, Md Arifur Rahman, Niamul Hassan Samin, Khandaker Rifah Tasnia, Sifat Rahman Ahona, Juena Ahmed Noshin
NLP Large Language Models Theory
  • Introduces role-stratified per-field conformal risk control for language-model tool calls.
  • Addresses the inadequacy of aggregate risk control methods that obscure failures in high-risk fields.
  • Demonstrates improved risk budget compliance and robustness under various conditions.
  • Provides formal guarantees for role-specific risk management.
Read more
A Statistical Difference between Single-Layer Learning and Hierarchical Learning in Wide Neural Networks
Sumio Watanabe
Theory
  • Training input-to-hidden weights reduces generalization error compared to fixing them.
  • Fixed input-to-hidden weights lead to singularities in the parameter space.
  • Both learning schemes are universal function approximators but exhibit different approximation properties.
  • Singularities are significant in wide neural networks and influence representation learning.
Read more
Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis
Yuntong Chen, Jianyu Liu, Guobin Zhao, Ziang Wang, Chao Chen, Ju Huang, Xitian Tian, Lijiang Huang
Time Series Interpretability Large Language Models
  • Introduction of the Diagnostic Evidence Network (DENet) for bearing fault diagnosis.
  • DENet provides a structured evidence record that includes classification, characteristic frequency prediction, and impulse localization.
  • The framework allows for independent validation of predictions against physical reality.
  • A QLoRA-adapted language model translates evidence into natural language, ensuring traceability and reducing hallucination rates.
Read more
A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks
Du Yin, Xiachong Lin, Yue Tan, Jinliang Deng, Estrid He, Hao Xue, Flora D. Salim
Graph Learning Time Series Optimization
  • A2TTA effectively addresses the challenges of topology expansion and temporal distribution shifts in traffic forecasting.
  • The framework separates adaptation into global corrections and context-specific specializations for better performance.
  • Extensive experiments validate the robustness and efficiency of A2TTA across multiple real-world traffic networks.
Read more
SpectONet: A Physics-Guided Spectral Deep Operator Network for Euler-Bernoulli Beam Dynamics
Shivani Saini, Ramesh Kumar Vats, Arup Kumar Sahoo
Theory Efficient ML
  • Introduction of SpectONet, a physics-guided spectral deep operator network for EBB dynamics.
  • Utilizes nonuniform spectral sensor placement for improved boundary response representation.
  • Incorporates physics-informed constraints into the training objective for better generalization.
  • Demonstrates significant performance improvements over traditional models in both synthetic and real-world datasets.
Read more
A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics
Zheshun Wu, Renjie Zheng, Jinhang Zuo, Zenglin Xu, Fang Kong
Reinforcement Learning Theory Efficient ML
  • Introduction of a unified framework for hybrid RL in tabular MDPs with shifted transition dynamics.
  • Development of two algorithms: MIN-UCB-VI for regret minimization and MAX-LCB-VI for best policy identification.
  • Provision of rigorous theoretical guarantees on regret and sub-optimality gaps.
  • Demonstration of the framework's effectiveness through extensive experimental validation.
Read more
Explainable Reinforcement Learning via Physics-Aware Policy Distillation
Shaker Al-Tamari, Waled Kadour
Reinforcement Learning Robotics Interpretability
  • Introduces a policy distillation framework for enhancing interpretability in DRL systems.
  • Achieves performance parity with a high-performance DRL agent while using a simpler decision tree model.
  • Demonstrates the trade-off between continuous and discrete control methods, affecting system dynamics.
  • Maintains BIBO stability, ensuring safety in autonomous applications.
Read more
Emergent Latent-State Computation under Stochastic Volatility
Xiaoyu Huang, Lulu Wang
Time Series Interpretability Theory
  • Demonstrates a two-stage computation in neural sequence models for forecasting latent volatility states.
  • Identifies specific architectural stages in Transformers where latent-state decodability emerges, varying with volatility periods.
  • Shows that in long-cycle regimes, the computation simplifies to a learned linear projection followed by normalization.
  • Highlights that forecast degradation under noisy training is linked to readout misalignment rather than latent state encoding failure.
Read more
LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning
Chen Wang, Boming Kang, Qinghua Cui
NLP Large Language Models Theory
  • Introduces LC-SEPLM, which integrates long-range contact supervision into protein representation learning.
  • Utilizes low-rank adaptation (LoRA) and cross-attention mechanisms to enhance model performance.
  • Demonstrates significant improvements in protein-level tasks, especially in remote-homology recognition.
  • Achieves superior performance on the ESM-S EC benchmark compared to previous models.
Read more