AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

41 Papers today
8h Update frequency
7 Days of history
AI Assistants Overassist
Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner
Large Language Models NLP
  • Introduction of INT-BENCH, a benchmark for evaluating LLM interventions in educational contexts.
  • LLMs intervene more frequently and earlier than human teachers, often providing complete solutions.
  • Current LLM assistance strategies may hinder long-term learning and cognitive engagement.
  • The study emphasizes the importance of balancing immediate assistance with promoting independent reasoning.
Read more
Gradient Concentration, Not Weight Saliency, Explains Representation-Level Class Unlearning
Billel Habbati, Alessio Merlo, Luca Verderame, Meriem Guerar
Theory Optimization Efficient ML
  • Saliency-based masking does not significantly enhance representation-level unlearning compared to random masking.
  • Forget gradients are predominantly concentrated in the final layers of the network, influencing all masking strategies similarly.
  • Saliency masks exhibit low class specificity, selecting overlapping parameter subsets across different classes.
  • Effective representation-level unlearning may necessitate objectives that act directly on representations rather than on weight selection.
Read more
ADABORD: a novel AdaBoost approach for ordinal classification
Rafael Ayllón-Gavilán, Francisco José Martínez-Estudillo, David Guijo-Rubio, César Hervás-Martínez, Pedro A. Gutiérrez
Theory Optimization
  • ADABORD is specifically designed for ordinal classification, addressing limitations of traditional nominal approaches.
  • It incorporates ordinal information into both the base estimator and the error function of the AdaBoost algorithm.
  • The framework significantly outperforms existing methods, particularly in datasets with multiple classes.
  • The use of the absolute ranked probability score allows for a more accurate representation of ordinal relationships.
Read more
Toward Mechanistic Interpretability of an AI Foundation Model Fine-Tuned for Atmospheric Chemistry
Jason Y. Hu, Ivan Higuera-Mendieta, Patrick Obin Sturm, Makoto M. Kelp
Interpretability
  • First mechanistic interpretability analysis of an AI foundation model for atmospheric chemistry.
  • Aurora captures some ozone-NOx coupling but does not enforce process-based model constraints.
  • Identified internal features can causally influence chemical forecasts.
  • Model predictions may rely on statistical patterns rather than true chemical understanding.
Read more
Automated Synthesis and Adversarial Validation of Executable Causal Research Pipelines
Irena Girshovitz, Dan Zeltzer, Ran Gilad-Bachrach
Theory
  • Introduction of ARA framework for automated causal research analysis.
  • Integration of protocol construction, synthetic data generation, and adversarial validation.
  • Focus on making silent failures in causal assumptions visible.
  • Evaluation of ARA reveals it surfaces methodological concerns rather than just providing estimates.
Read more
SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking
Chenyu Zhang
Reinforcement Learning Optimization
  • SalesLoop addresses the gap between offline model accuracy and online performance in sales lead ranking.
  • It introduces a performance-aware reward system that incorporates ranking position and conversion speed.
  • The framework adapts Group Relative Policy Optimization to enhance listwise ranking quality.
  • A 160-day production test showed significant improvements in conversion rates and lead ranking effectiveness.
Read more
Agree on the Model, Verify the Inference: GKR Protocols for HND-Based Transformer Inference
Xiaolong Liang, Juanjuan Li, Rui Qin, Yisheng Lv
Theory Efficient ML
  • Introduction of GKR-HND for verifying outsourced Transformer inference without dense-matrix replay.
  • Separation of cryptographic verification from public computation to enhance efficiency.
  • Validation of the protocol using pretrained HND models, confirming its effectiveness in real-world scenarios.
Read more
Beyond SBDD: Geometric Deep Learning in Polypharmacology and Multi-target Drug Design
Tianming Han, Zhijie Pan, Wenchi Ge, Qi Zhao
Graph Learning Generative Models Optimization
  • Traditional SBDD is inadequate for complex diseases due to its focus on single targets.
  • Polypharmacology offers a promising strategy by designing drugs that target multiple proteins.
  • Geometric deep learning provides a framework for addressing multi-target drug design challenges.
  • The paper reviews various GDL architectures and their applications in drug discovery.
Read more
Conflict Resolution under Degraded Surveillance in Air Corridors Using Multi-Agent Reinforcement Learning
Esrat Farhana Dulia, Syed Arbab Mohd Shihab, Caleb Adams, Ruben Del Rosario
Reinforcement Learning Robotics
  • Developed a DQN-based MARL framework for decentralized conflict resolution in AAM.
  • Evaluated the impact of traffic density and separation thresholds on conflict resolution effectiveness.
  • Demonstrated that most conflicts are resolved quickly, within 1 second.
  • Identified key maneuvering strategies used by agents during conflicts.
Read more
A Graph Neural Network approach to zero-shot Digital Twins
Alicia Tierz, Icíar Alfaro, David González, Elías Cueto
Graph Learning
  • Introduction of a Zero-Shot Digital Twin framework that adapts to unseen geometries without retraining.
  • Utilization of a Thermodynamics-Informed Graph Neural Network to enforce physical laws in simulations.
  • Implementation of a continuous closed-loop feedback mechanism for real-time correction of simulations.
  • Demonstrated generalization capabilities across disparate physical scenarios.
Read more
Offline RL with Hierarchical Action Chunking
Ahad Jawaid
Reinforcement Learning Robotics Theory
  • HiQC mitigates the curse of horizon in offline goal-conditioned RL through dual horizon reduction.
  • The algorithm combines high-level planning with low-level action chunking for stable long-horizon learning.
  • Theoretical analysis indicates improved value estimation error bounds compared to standard methods.
  • Empirical results show HiQC achieves the highest performance on long-horizon navigation tasks.
Read more
Relative Value Learning
Marc Höftmann, Jan Robine, Stefan Harmeling
Reinforcement Learning
  • Introduces a framework for learning value differences directly, eliminating the need for absolute value estimates.
  • Develops a pairwise Bellman operator that is γ-contractive with a unique fixed point corresponding to true value differences.
  • Derives well-posed targets for bootstrapping using observable rewards and pairwise terms.
  • Demonstrates the effectiveness of RV integrated with PPO on the Atari benchmark, achieving competitive performance.
Read more
Multimodal CoLRAG-TF: Triple-Filtered Retrieval for Complex PDFs
Takato Yasuno
Multimodal NLP Optimization
  • Introduces a four-axis fusion architecture for multimodal retrieval, enhancing retrieval precision and compositional reasoning.
  • Implements a hybrid OCR pipeline that significantly improves table coverage in scanned documents.
  • Achieves high performance in multi-hop question answering through structured knowledge-graph integration.
  • Demonstrates the effectiveness of Bayesian optimization in calibrating retrieval weights to counteract lexical bias.
Read more
Multilevel Graph Wavelet Compressed Sensing with Scale-Aware Neural Recovery
Amirhossein Nouranizadeh, Sarang Rajendra Patil, Alan John Varghese, Varsha Narayanan, Amit Chakraborty, Mengjia Xu
Graph Learning Efficient ML Theory
  • Introduction of Graph Wavelet Compressed Sensing (GWCS) for graph signal compression.
  • Combination of multilevel importance sampling and a scale-aware GNN for signal reconstruction.
  • Demonstrated high reconstruction fidelity and substantial data compression on various datasets.
  • Development of Neural Inverse Graph Wavelet Transform (NIGWT) to improve reconstruction accuracy.
Read more
Three-Pronged Spectral Control for Federated Parameter Efficient Fine Tuning
Shiva Raj Pokhrel, Dipsan Bhattarai, Anwar Walid
Federated Learning Efficient ML Optimization
  • TRISHUL introduces a spectral-control framework to enhance federated PEFT under non-IID conditions.
  • The framework combines shared bases for exact aggregation, nuclear-norm shrinkage, and non-uniform adaptation head allocation.
  • TRISHUL maintains low communication overhead while improving model convergence and stability.
  • The method provides diagnostics for assessing update subspace alignment and spectral stability.
Read more
Improving Access to Essential Medicines via Decision-Aware Machine Learning
Angel Tsai-Hsuan Chung, Jatu Abdulai, Patrick Bayoh, Lawrence Sandi, Francis Smart, Hamsa Bastani, Osbert Bastani
Optimization Efficient ML
  • Introduces a decision-aware machine learning framework for essential medicine allocation.
  • Utilizes multi-task learning to improve sample efficiency and equitable distribution.
  • Achieved a 19% increase in allocated product consumption in treated districts.
  • Demonstrates the potential of machine learning to address healthcare resource allocation in LMICs.
Read more
Cardinality-Decomposed Loss: Matching Training Objectives to Relation Structure in Heterogeneous Recommendation Graphs
Parul Maheshwari, Amulya Paruchuri, Yiqing Zou, Alireza Sahami Shirazi, Farhad Farahani, Prakhar Mehrotra
Graph Learning
  • Traditional BPR loss leads to silent failures in attribute embeddings in heterogeneous recommendation systems.
  • Cardinality-Decomposed Loss (CDL) combines BPR and Cross Entropy to optimize for different relation structures.
  • CDL significantly improves attribute discriminability and ranking performance in various datasets.
  • The paper introduces a framework to understand the trade-off between loss functions based on semantic alignment and topology leakage.
Read more
Attribution Markets: A Fisher-Market Formulation for Fractional Credit Assignment Between Planned Tasks and Performed Actions
Salavat Ishbulatov
Optimization Theory
  • Introduces a Fisher market framework for fractional credit assignment in planning systems.
  • Proposes two market instruments that ensure conservation and budget constraints.
  • Addresses convergence issues with a new fixed point approach validated empirically.
  • Identifies sensitivity to noise in market equilibrium and resolves it with entropy regularization.
Read more
Context-weighted Discrete Flow Matching
Daniil Cherniavskii, Daniel Severo, Karen Ullrich
Generative Models NLP Large Language Models
  • Prediction difficulty is closely linked to the availability of local context, with denser neighborhoods leading to lower uncertainty.
  • The proposed context-weighted sampler improves generation quality without requiring fine-tuning of pre-trained models.
  • A scaled cross-entropy loss function effectively reweights the training signal, reducing generative perplexity by up to 63%.
  • The approach matches the quality of a strong semi-autoregressive block diffusion baseline while allowing for any-order generation.
Read more
External Clustering Validation by the Homogeneity-Parsimony Trade-off
Andreas Tiffeau-Mayer
Theory
  • Introduces normalized scores for homogeneity and parsimony in clustering validation.
  • Demonstrates that these scores vary monotonically with cluster refinement.
  • Extends the information-theoretic framework to include set-matching and pair-based metrics.
  • Unifies various clustering evaluation criteria and connects them to ROC analysis.
Read more
Robust Asynchronous Q-Learning under Reward and State Corruption via Batching
Sreejeet Maity, Aritra Mitra
Reinforcement Learning Theory
  • Introduction of BR-Async-Q, a robust variant of Q-learning that handles asynchronous online sampling.
  • The algorithm partitions data into batches to reduce variance and improve robustness against corruption.
  • Proven high-probability error bounds that match standard Q-learning, with additional terms accounting for corruption.
  • First robustness guarantee for asynchronous Q-learning under joint state and reward corruption.
Read more
Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignment Score For Retinal Disease Classification
Kritanu Chattopadhyay, Sayanjit Singha Roy, Soumya Chatterjee
Computer Vision Generative Models Interpretability
  • Introduces CounterFundus, a CycleGAN-based framework for counterfactual explainability in retinal disease classification.
  • Generates visually plausible healthy counterparts of diseased retinal images to enhance interpretability.
  • Introduces the Counterfactual-Classifier Alignment Score (CCAS) for quantifying spatial agreement between counterfactuals and classifier saliency.
  • Demonstrates improved classification performance through CCAS-filtered counterfactual augmentation.
Read more
Double-Scoring: Reliable Extraction of Strong Lottery Tickets
Bryce A. Christopherson, Jack Baretz, Darian Colgrove, Salah Dandan
Theory Optimization Efficient ML
  • Identifies layerwise sparsity selection as a bottleneck in strong ticket extraction.
  • Introduces double-scoring, which removes the need for tuning sparsity levels for each layer.
  • Proves that augmented score-space masking maintains representational access to original masks.
  • Demonstrates substantial improvements in strong-ticket extraction through controlled experiments.
Read more
GlucoTune: A Unified Framework for Blood Glucose Preprocessing, Forecasting, and Benchmarking in Diabetes
Davide Marelli, Giorgia Rigamonti, Mirko Paolo Barbato, Paolo Napoletano
Time Series
  • GlucoTune standardizes the entire experimental workflow for blood glucose data, enhancing reproducibility.
  • Configurable preprocessing pipelines allow for consistent data handling without distributing sensitive data.
  • The framework integrates various prediction models and public datasets, promoting fair comparisons.
  • A benchmarking leaderboard enables systematic evaluation of different experimental configurations.
Read more
Emergent Misalignment Recruits a Pre-existing Persona Subspace
Mohammed Suhail B Nadaf
NLP Large Language Models Theory
  • Fine-tuning on narrow bad data can lead to broad misalignment in language models.
  • A low-rank persona core shared across unrelated domains exists in the model before fine-tuning.
  • Projecting the persona subspace out during fine-tuning prevents misalignment, while injecting it into untouched models induces misalignment.
  • The first optimizer step during fine-tuning is crucial for the emergence of misalignment.
Read more
Hilbert Operator for Progressive Encoding (HOPE): A Mathematical Framework for Deconstructing Learned Representations in Deep Networks
Hossein Mobahi, Peter L. Bartlett
Theory Efficient ML Interpretability
  • HOPE provides a unified mathematical framework for deconstructing learned representations in deep networks.
  • The framework operates in a Hilbert space, allowing for unbiased capacity measurements across different layers.
  • HOPE is data-free and hyperparameter-free, utilizing Batch Normalization statistics for analysis.
  • The approach combines pruning and merging into a single theoretical paradigm, enhancing model compression techniques.
Read more
GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries
Miguel P. Bento, João Seabra
Large Language Models Efficient ML Optimization
  • GaugeQuant leverages LLM symmetries to optimize quantization during training.
  • Introduces a LogSumExp term in the loss function to minimize activation outliers.
  • Eliminates the need for calibration data and quantization simulations.
  • Demonstrates significant improvements in perplexity for LLaMA-2 7B model.
Read more
CEDAR: Causal Edge Discovery for Autoregressive Processes
Mohammad Fesanghary
Time Series Graph Learning Theory
  • CEDAR introduces a constraint-based approach for causal edge discovery in autoregressive time series.
  • The method utilizes U-centered distance correlation for nonlinear lag selection, enhancing its applicability to various data types.
  • A stable MCI pruning step effectively removes indirect causal edges, improving the clarity of causal relationships.
  • CEDAR includes a mechanism to adjust for nonstationarity in time series, addressing a common challenge in real-world data.
Read more
HierarchicalDAEW: Domain-Aware Edge-Weighted Graph Convolution with Evidential Uncertainty for Multi-Section Spatial Gene Expression Prediction from H&E Histology
Kritanu Chattopadhyay, Soumya Chatterjee, Ondrej Krejcar, Debotosh Bhattacharjee
Graph Learning
  • Introduces Domain-Aware Edge-Weighted convolution to encode tissue heterogeneity in graph message passing.
  • Achieves a +0.044 PCC improvement over flat graph baselines through Hierarchical DomainGCN and CrossScaleGate.
  • Utilizes a gene graph decoder to propagate predictions across a broader gene panel with a mean PCC of 0.8314.
  • Demonstrates evidential uncertainty estimation with 90% empirical coverage, enhancing prediction reliability.
Read more
Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development
Nadine Chang, Maying Shen, Jialiang Wang, Rafid Mahmood, Jose M. Alvarez
Robotics Theory
  • Reactive maintenance approaches in AI systems are limited and often lead to inefficiencies.
  • A proactive test-driven flywheel can enhance generalization and prevent future errors.
  • Mapping feedback to a broader 'test space' allows for better alignment with task objectives.
  • The authors provide mathematical proof that proactive methods achieve better long-term scaling.
Read more
Chronofy: A Temporal-Logical Decay Architecture for Information Validity in Time-Aware Retrieval-Augmented Generation
Muntaser Syed, Marius Silaghi, Sheikh Abujar, Sharun Akter
NLP Large Language Models Graph Learning
  • Chronofy integrates temporal validity directly into RAG systems through a three-layer architecture.
  • The framework employs learnable exponential decay functions to weight retrieved knowledge based on temporal relevance.
  • Signal Temporal Logic is utilized to assess the validity of knowledge rather than the confidence of LLM outputs.
  • Chronofy demonstrates improved retrieval precision and reduced temporal hallucination in evaluations.
Read more
Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity
Hongnan Ma, Yiwei Shi, Mengyue Yang, Weiru Liu
Time Series
  • Existing time-series explanation methods often select predictive but non-essential subsequences due to their focus on sufficiency.
  • TimePNS introduces a framework that estimates counterfactual necessity through structured interventions in a learned causal latent space.
  • The framework consists of a two-stage training procedure that enhances the quality of explanations by filtering out spurious subsequences.
  • Experiments show that TimePNS outperforms existing methods in identifying critical subsequences and improving explanation reliability.
Read more
Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma
Larry Richards
NLP Large Language Models Theory
  • The study tests the hypothesis that holonomy concentrates on active SAE feature planes.
  • Results indicate that active-feature planes carry less holonomy than mixed-feature controls.
  • The research design was preregistered, ensuring unbiased analysis.
  • Transport distortion was identified as a potential confounding factor.
Read more
AI-Driven Surrogate Models for Predicting Electrode-Scale Discharge Behavior in Lithium-Ion Batteries
Mengda Xing, Jean-Marie Lagniez, Alejandro Franco
Efficient ML Time Series Optimization
  • Introduction of a deep learning surrogate pipeline based on Swin3D Transformer for LIB discharge prediction.
  • Integration of Gaussian Positional Encoding to enhance spatial feature representation.
  • Development of a Temporal Encoding module to capture non-linear time-series evolution.
  • Significant reduction in computational overhead, achieving simulation times from hours to milliseconds.
Read more
Smooth Neural Point Processes via B-Splines
Michele Bellomo, Riccardo Ramaschi, Alberto Dolara, Tomaso Aste
Time Series
  • Introduces a novel neural TPP model using B-spline basis functions for CIF parametrization.
  • Eliminates the need for numerical integration, allowing for exact NLL evaluation.
  • Enables efficient parallelization during training, improving computational efficiency.
  • Supports smoothness regularization of the CIF, enhancing model generalization.
Read more
From Evaluation to Optimisation: Hierarchy-Aware Training Signals for CWE Prediction in Python
Muntasir Adnan, Manile Srun, Carlos C. N. Kuhn
Large Language Models Reinforcement Learning Optimization
  • The ALPHA penalty is validated as an effective training signal for CWE prediction.
  • Delivery mechanisms significantly influence the effectiveness of the penalty in training.
  • Reinforcement learning (GRPO) outperforms supervised methods under distribution shifts.
  • The best GRPO configuration reduces the ALPHA penalty by 27.9%, achieving parity with a larger model.
Read more
How Many Bits Can an Adapter Write? Measuring the Capacity and Memorization of Parameter-Efficient Fine-Tuning
Kaizhen Tan, Heqing Du, Yang Feng
NLP Large Language Models Efficient ML
  • LoRA adapters store fewer bits per parameter than previously assumed, with capacity influenced by parameter placement.
  • The study introduces a measurement protocol that quantifies the information an adapter can hold.
  • Supervised fine-tuning can lead to privacy leakage, while reinforcement learning does not record sensitive data.
  • The findings challenge the folklore surrounding adapter capabilities and provide a basis for designing against memorization risks.
Read more
Fisher Widths: Local Learning Geometry and Anisotropic Recovery
Vu Khac Ky
Theory
  • Introduces the primal and inverse-Fisher widths as complementary measures of complexity in statistical learning.
  • Establishes a relationship between the primal and inverse-Fisher widths, showing they cannot both be reduced relative to the Euclidean scale.
  • Demonstrates the role of Fisher widths in local parameter fluctuations and recovery problems with Gaussian measurements.
  • Provides support-sensitive recovery estimates that depend on the geometry of the Fisher spectrum.
Read more
Best-of-Evidence: Best-of-N Selection under Partial Verification
Cenwei Zhang, Teng Fang, Yuxia Wang, Derek Li, Bryan Dai, Lei You
Multimodal Theory Optimization
  • Introduces Best-of-Evidence (BoE) framework for selection under partial verification.
  • Utilizes a signed candidate-factor graph to manage reusable claims and evidence acquisition.
  • Establishes theoretical bounds on evidence-driven improvements and query efficiency.
  • Demonstrates improved selection outcomes in medical VQA tasks through empirical evaluation.
Read more
The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works
Yu Wang
Reinforcement Learning Large Language Models Theory
  • Dense prediction rewards can lead to catastrophic policy collapse in GRPO-trained LLM agents.
  • The 'dark room' pathology results from z-scoring normalization, which amplifies variance in all-fail groups.
  • A variance-profile criterion is introduced to evaluate the safety of dense signals in reinforcement learning.
  • Controlled experiments reveal that reward channels are neutral compared to auxiliary-loss channels.
Read more
Zero-Flow Two-Sample Tests
Yakun Wang, Leyang Wang, Song Liu, Taiji Suzuki
Theory
  • Introduction of zero-flow discrepancy (ZFD) as a new statistical measure for two-sample testing.
  • Development of the zero-flow two-sample test (ZF2ST) that separates witness learning from hypothesis evaluation.
  • Utilization of flexible neural networks for witness learning while ensuring valid statistical calibration.
  • Demonstration of strong testing power and well-calibrated type-I error rates through experiments on various datasets.
Read more