AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

24 Papers today
8h Update frequency
7 Days of history
Federated Attack Campaign Detection via Contrastive Encoding of Threat Indicators in Gradient Updates
Manuel Röder, Bibin Babu, Frank-Michael Schleif
Federated Learning
  • Introduction of FedIoC, a framework for federated attack detection that encodes threat indicators into gradient updates.
  • Utilization of a supervised contrastive loss to enhance the representation of attack campaigns in gradient updates.
  • Demonstrated effectiveness on two public datasets (CTU-13 and UNSW-NB15) for recovering attack campaign structures.
  • No raw indicators are transmitted, preserving privacy and complying with data protection regulations.
Read more
Locating and Steering Refusal Beyond Attention
Preethi Carmel Bosco, Gopalakrishnan Srinivasan
NLP Large Language Models Interpretability
  • Refusal is a shared representation across different model architectures.
  • The refusal direction must be read at the fresh write site for effective harm detection.
  • A detector-triggered gate can significantly reduce jailbreak attack success rates.
  • The refusal representation can be aligned across architectures using a rigid rotation.
Read more
WEECFP-SuRGE: Wide Embedded Extended Connectivity Fingerprint with Substructure Rotary Graph-distance Encoding
Robert W. Epps
Graph Learning
  • Introduction of WEECFP, a parameter-free continuous molecular fingerprint with enhanced substructure representation.
  • Development of WEECFP-SuRGE, a transformer architecture utilizing SuRGE for improved molecular property prediction.
  • Achieved top rankings on the TDC ADMET leaderboard without external pretraining, outperforming classical fingerprint methods.
  • Demonstrated near-lossless tokenization with high recovery rates of canonical SMILES.
Read more
A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models
E. Cho Smith, Samuel Ho, Dawn Laux
NLP Large Language Models
  • LLMs can annotate cultural texts but their reliability varies by social construct.
  • Self-esteem shows the strongest reliability, while seeking recognition is less stable.
  • Consensus labels from LLMs can be useful for supervised classification but do not confirm construct validity.
  • Establishing measurement reliability is crucial before using LLM outputs in cultural analytics.
Read more
FedDRAW: Federated Dual Reputation Annealing Weighting for Heterogeneous Multi-Institutional Chest Radiograph Classification
Maryam Moradpour, Anne-Christin Hauschild
Federated Learning
  • FedDRAW improves federated learning by adjusting aggregation weights based on model similarity rather than solely on data size.
  • The method employs dual annealing schedules to balance client influence during training.
  • FedDRAW ensures that smaller hospitals are not permanently outweighed by larger institutions in the model training process.
  • The approach outperforms eight state-of-the-art methods in chest radiograph classification tasks.
Read more
Amortizing Scaling Law Construction Costs
Abhash Kumar Jha, Diana Alexandra Onuţu, Neeratyoy Mallik, Swagatam Haldar, Sam Laing, Niccolò Ajroldi, Shiwei Liu, Joaquin Vanschoren, Aaron Klein
Large Language Models Optimization Efficient ML
  • Introduces a framework for efficient scaling law construction using Bayesian optimization.
  • Demonstrates significant computational savings (10–100x) by avoiding exhaustive grid evaluations.
  • Utilizes surrogate evaluations to enhance the recovery of scaling laws from sparse data.
  • Proposes metrics for assessing the efficiency of scaling law fitting methods.
Read more
Hessian-based molecular conformation augmentation for a scalable and efficient strategy of machine learning interatomic potentials
Bumju Kwak, Jeonghee Jo
Efficient ML Theory Optimization
  • Introduction of two Hessian-derived data augmentation methods: UniAug and ModeAug.
  • Both methods enhance MLIP training by incorporating curvature information without modifying training objectives.
  • UniAug utilizes isotropic Gaussian displacements, while ModeAug employs normal mode-weighted displacements.
  • The proposed methods show improved accuracy over traditional energy-force training and are competitive with direct Hessian supervision.
Read more
GLASS: Graph-Language Alignment with Spherical Scoring for Transferable Graph-Level Anomaly Detection
Xudong Wang, Chris Ding, Tongxin Li, Jicong Fan
Graph Learning
  • GLASS introduces a transferable GLAD paradigm using density estimation on an aligned graph-language hypersphere.
  • The framework integrates Local Topology Descriptors, GraphDP, and Spherical Multi-Modal Scoring for robust anomaly detection.
  • Empirical results show GLASS outperforms existing methods in single-domain performance and enables effective zero-shot and few-shot transfer.
  • The model captures anomalies at multiple levels of granularity through a structured representation approach.
Read more
Hakken: Predicting future discoveries to fill the gaps in today's knowledge
Tarek R. Besold, Uchenna Akujuobi, Pablo Sanchez, Alessandra Toniato, Kana Maruyama, Jihun Choi, Samy Badreddine, Frederick Gifford, Daniel Evans-Yamamoto, Sucheendra K. Palaniappan, Miquel Ferrer, Kae Nagano, Iris Rossell, Tom Joy, Hatem ElShazly, Chrysa Iliopoulou, Christoph Wehner, Thiviyan Thanapalasingam, Susana Nunes, Pedro G. Cotovio, Peter Wurman, Peter Stone, Hiroaki Kitano, Michael Spranger
Graph Learning Large Language Models NLP
  • Hakken predicts novel scientific relationships using a transformer-based model and knowledge graphs.
  • The system establishes a new benchmark for time-aware multi-label relation prediction in the biomedical domain.
  • Hakken's predictions were validated with biologists, leading to empirical confirmations of new gene interactions.
  • The model provides explainability for its predictions, aiding researchers in evaluating new hypotheses.
Read more
On the Abundance of Critical Points of the t-SNE Energy
Nakul Haridas, Ryan Murray
Theory Optimization
  • t-SNE's energy landscape is non-convex, leading to multiple local minima.
  • The authors construct infinite families of distinct critical points based on symmetry pairs.
  • Numerical examples illustrate the complexity of the t-SNE minimization problem.
  • The trivial embedding is shown not to be a local minimum for many symmetry pairs.
Read more
Conformity Breaks Conformal Prediction
Yibo Hu, Hanyu Su
Large Language Models Theory NLP
  • Introduction of the score-mechanism shift, highlighting how peer influence alters model scoring.
  • Demonstration of a significant drop in coverage rates under unanimous-wrong peer conditions.
  • Identification of a hidden conditional failure affecting low-confidence items.
  • Critique of existing defenses against peer influence, showing their inadequacy.
Read more
An Energy-Based Conservative-Dissipative Latent Neural Evolution Operator for Magnetization Dynamics
Sebastian Schaffer, Lukas Exl
Theory Efficient ML Optimization
  • Introduces a reduced-order model for micromagnetic dynamics using a latent neural ODE.
  • The model leverages an energy-based approach that ensures monotonic decrease of a learned scalar potential.
  • Demonstrates significant improvements in trajectory prediction accuracy compared to traditional methods.
  • Evaluates various latent energy formulations, with deep-quadratic energy yielding the best results.
Read more
A Data Fusion Framework for Grounding Aerospace Surrogate Model via Experimental Wind-Tunnel Observations
Nitin Nagesh Kulkarni, Dheeraj Vemula, Yin Yu, Peter Lyu, Juan J. Alonso
Optimization Efficient ML Theory
  • Introduces a correction framework that integrates experimental data into CFD-trained surrogate models.
  • Achieves high predictive fidelity for aerodynamic forces while addressing systematic discrepancies.
  • Demonstrates significant improvements in prediction accuracy using limited experimental datasets.
  • Maintains the computational efficiency of the surrogate model without retraining its parameters.
Read more
A Quantum Variational Approach to Prototypical Recurrent Unit
Mahyar Sadeghi Garjan, Tommaso Cesari, Michel Barbeau
Time Series
  • Introduction of a lightweight quantum recurrent architecture (QPRU) with fewer parameters than classical and quantum counterparts.
  • Utilization of Variational Quantum Circuits (VQCs) for modeling temporal dependencies in sequential data.
  • QPRU achieves competitive forecasting performance on time-series benchmarks.
  • The architecture offers enhanced scalability and practical advantages over existing models.
Read more
Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation
Siliang Liu, Mohammad Ghasemi, Sapan Patel, Amin Banitalebi-Dehkordi
NLP Large Language Models Efficient ML
  • Introduces a two-level framework for trade-up recommendation that combines reasoning distillation and test-time training.
  • Utilizes a retrieval-augmented LLM to generate structured labels and rationales for product pairs.
  • Achieves significant improvements in predictive performance without the need for LLM inference during scoring.
  • Demonstrates operational efficiency with a 5,000× speedup and 10,000× cost reduction compared to direct LLM inference.
Read more
Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One
Fred Zhangzhi Peng, Kaiwen Zheng, Anru R. Zhang
NLP Large Language Models Generative Models
  • PlaidQ is a continuous diffusion language model that allows for efficient code generation.
  • The model can generate code in as few as one denoising step while maintaining functional correctness.
  • Distillation techniques significantly improve the performance of the model with fewer steps.
  • PlaidQ outperforms traditional autoregressive models in code generation tasks.
Read more
Fractal basins trap latent reasoning
Jeffrey Lai, Anthony Bao, John Quinn, William Gilpin
Theory Optimization
  • Reasoning models exhibit transient chaos, leading to extended reasoning times on difficult tasks.
  • Fractal basins of attraction increase in complexity with task difficulty across various reasoning tasks.
  • Basin entropy serves as a new metric to quantify the complexity of convergence basins in reasoning models.
  • The study reveals that reasoning slowdowns are an inevitable consequence of problem hardness.
Read more
Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures
Amar Alem Koric, Qibang Liu, Seid Koric
Theory Efficient ML Optimization
  • Cross-attention mechanisms significantly improve the accuracy of DeepONets in solving PDEs.
  • Self-attention can be inconsistent, providing benefits in complex scenarios but potentially degrading performance in simpler cases.
  • The study systematically evaluates the impact of different attention configurations on model performance.
  • Increasing the depth of cross-attention improves accuracy but incurs higher computational costs.
Read more
REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation
Mohsen Nayebi Kerdabadi, Arya Hadizadeh Moghaddam, Dongjie Wang, Zijun Yao
Graph Learning Reinforcement Learning Large Language Models
  • REFINE constructs patient-specific temporal graphs from a global TKG for personalized medical concept encoding.
  • A sequential reinforcement learning policy is used to adaptively select the relational context for each observed medical code.
  • The framework integrates a heterogeneous GNN and a frozen LLM for capturing structural dependencies and refining semantic representations.
  • Experiments show that REFINE outperforms existing models and demonstrates robustness in diverse scenarios.
Read more
Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy
Margherita Mele, Andrea Castagna, Roberto Menichetti, Raffaello Potestio, Alessandro Ingrosso
Theory Efficient ML Interpretability
  • Introduces a novel unsupervised method for neuron selection based on mapping entropy.
  • Demonstrates that ME optimization can recover minimal representations in teacher-student networks.
  • Shows that ME-selected subnetworks outperform random subsets in predictive tasks.
  • Links configurational distinguishability to improved performance under compression.
Read more
Confounding-Valid Conformal Inference for Counterfactual KPIs in Wireless Networks
Abdessamed Qchohi, Jessica Moysen Cortes, Matteo Zecchin
Theory Optimization
  • CV-CCI effectively combines observational and randomized telemetry to improve counterfactual analysis in wireless networks.
  • The methodology retains finite-sample coverage guarantees while addressing hidden confounding issues.
  • Experimental results show that CV-CCI outperforms existing methods in terms of prediction set efficiency and validity.
Read more
Dynamic Heterogeneous Graph Representation Learning: A Survey
Huan Liu, Pengfei Jiao, Jie Yin, Hongjiang Chen, Zhidong Zhao
Graph Learning
  • Introduces a unified formal definition of Dynamic Heterogeneous Graphs (DHGs).
  • Proposes an algorithm-centric taxonomy categorizing DHG representation learning methods.
  • Highlights the intrinsic modeling biases of existing methods concerning dynamic granularity.
  • Summarizes applications, datasets, and benchmarks for DHG representation learning.
Read more
Representation Redundancy and Structural Complexity in Finite-Field Inversion
Zheng Zhang, Na Zhang
Theory
  • Exact redundancy in representation choice is characterized by Galois orbits.
  • Different Boolean formulations of inversion exhibit varying algebraic complexities.
  • Empirical results align with theoretical predictions regarding learning difficulty.
  • Galois orbit redundancy offers limited generalization benefits in multilayer perceptron experiments.
Read more
Beyond a Universal Forecasting Selector: Demand-Conditioned Model Selection across Demand Patterns and Horizons
Adolfo González, Víctor Parada
Time Series
  • Model selection in forecasting should be context-dependent rather than universal.
  • Five selection mechanisms were evaluated, revealing varying effectiveness across demand patterns.
  • CCG-AHSC and CCG-AHSCD are better for Smooth and Erratic demands, while OWA and ERA excel in Intermittent and Lumpy settings.
  • Selector performance is influenced by historical data availability and forecasting horizons.
Read more