AI-generated summaries

Today's ML research,
without the noise.

Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.

24 Papers today
8h Update frequency
7 Days of history
Rethinking the Teacher-Student Framework for Test-Time Adaptation
Damian Sójka, Marc Masana, Bartłomiej Twardowski, Sebastian Cygert
Computer Vision Theory Optimization
  • The EMA teacher in TTA does not prevent model collapse over long test sequences.
  • Introducing a fixed-weight teacher (Intransigent Teacher) effectively mitigates error accumulation.
  • The IT approach allows students to potentially surpass their teachers in performance.
  • The proposed method enhances robustness to hyperparameter changes and is applicable across various architectures.
Read more
Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators
Ryota Ushio, Takashi Ishida, Masashi Sugiyama
Theory
  • Introduces soft-label-based estimators for optimal BER and AUC in binary classification.
  • Addresses the limitations of traditional accuracy metrics in imbalanced or noisy datasets.
  • Extends the FeeBee framework for practical evaluation of estimators without requiring knowledge of the optimum.
  • Validates the proposed methods through experiments on synthetic and real-world datasets.
Read more
On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers
Vignesh Prabhakar, Jialing Pan, Anil Babu Ankisettipalli
NLP Reinforcement Learning Efficient ML
  • Introduces a two-stage framework combining off-policy teacher optimization and on-policy student distillation.
  • Demonstrates significant performance improvements in compact instruction-following rerankers, especially under distribution shifts.
  • The proposed method outperforms traditional offline distillation approaches by a notable margin.
  • Achieves a favorable quality-efficiency tradeoff, making it suitable for deployment in production environments.
Read more
CliffRank: A Dual-Branch Framework for Activity-Cliff Ranking Prediction
Kewei Li, Rongying Zhang, Peiyu Yang, Zhongjian Wang, Qiuchen Zhao, Lan Huang, Fengfeng Zhou
Optimization Theory
  • CliffRank effectively combines absolute-activity regression with ranking-consistency learning.
  • The framework utilizes two parallel predictors to enhance prediction accuracy for activity cliffs.
  • CliffRank achieved the highest mean Spearman correlation of 0.6890 on small-molecule datasets.
  • Performance varied across datasets, highlighting the importance of dataset characteristics in model evaluation.
Read more
LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates
Dmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov
NLP Large Language Models Optimization
  • LoRA-TSD optimizes low-rank updates by considering the geometry of the weight space, improving efficiency.
  • The method provides a computationally cheaper retraction compared to traditional truncated-SVD approaches.
  • The paper establishes global convergence guarantees for both LoRA-TSD and LoRA-Pro.
  • LoRA-TSD outperforms existing LoRA optimizers across multiple benchmarks, demonstrating stability across adapter ranks.
Read more
UE5M3 FP4 Block Scaling for Stable Language Model Pretraining
Robert Hu, Carlo Luschi, Paul Balanca
NLP Large Language Models Efficient ML
  • Introduction of a UE5M3 FP4 pretraining recipe that simplifies the training process.
  • Demonstrated lower training and validation losses compared to the existing NVFP4 method.
  • Achieved a 21.2% increase in model-body token throughput by optimizing the execution process.
  • Identified a range bound linking UE5M3 scale targets to model performance.
Read more
hLLM: Single Pass Decoding for Generative Reranking
Emil Laftchiev, Prachi Agrawal, Moe Kayali, Bixing Yan, Qi Xu, Zijie Lei, Chen Qiu, Zhi Hua, Ke Li, Luke Simon
NLP Large Language Models Optimization
  • HLLM introduces a novel decoding strategy that reduces the number of forward passes required for generative ranking from O(N·T) to O(1).
  • The method utilizes a lightweight self-attention mechanism to create a score matrix from LLM hidden states, enabling efficient optimal assignment via the Hungarian algorithm.
  • Empirical evaluations show a 64× speed-up in inference time while preserving ranking quality on both proprietary and open datasets.
  • The framework connects generative ranking with combinatorial optimization, paving the way for further advancements in real-time ranking systems.
Read more
Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT
Madhusudhana Naidu
NLP Efficient ML
  • Introduces an efficient SciBERT-based approach for telescope bibliography classification.
  • Achieves a macro F1 score of 0.89, ranking first in the WASP-2025 Shared Task.
  • Analyzes the impact of truncation on classification performance.
  • Demonstrates the advantages of domain-specific pretraining over general models.
Read more
DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation
Wei Zhang, Hongji Li, Song Sun, Peng Yu, Xue Yang, Lei Zhao, Peng Jiang
Reinforcement Learning Optimization
  • DMRL connects skill document editing to downstream parameter optimization through structured interfaces.
  • Introduces DRPO for robust advantage estimation and LRP for long-term outcome prediction.
  • Addresses the challenges of credit assignment and population heterogeneity in advertising recommendations.
  • Demonstrated significant online improvements in a real-world advertising platform.
Read more
D-FROST: Decentralized Federated pRompt-tuning via Optimal tranSporT for Non-IID and Imbalanced Data
Quan Minh Nguyen, Hoang M. Ngo, Trong Nghia Hoang, My T. Thai
Federated Learning Optimization Computer Vision
  • First study of prompt tuning in decentralized federated learning (DFL).
  • Formulates decentralized prompt tuning as a Wasserstein-based optimization problem.
  • Introduces D-FROST, an optimal-transport-based algorithm for merging prompts.
  • Provides convergence analysis showing stability and consensus in prompt tuning.
Read more
CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
Huu Huy Nguyen, Chien Van Nguyen, Franck Dernoncourt, Ryan A. Rossi, Linh Ngo Van, Jieyang Chen, Thien Huu Nguyen
NLP Large Language Models Efficient ML
  • CRISP introduces a structural proxy (Cstruct) that simplifies routing decisions in attention mechanisms.
  • The method addresses the issue of noise accumulation in attention mass through sink-aware thresholding.
  • Empirical results show CRISP achieves up to 5.30× speedup in attention computation while maintaining performance.
  • The approach effectively adapts to long-context inputs, improving retrieval task performance significantly.
Read more
ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction
Quan Hao, Mengyue Fan, Zifan Dong, Youru Li, Jianduo Zhao, Lechuan Xu, Hao Zhang, Fei Xia, Jigang Wang, Chong Qiu, Liguo Zhang
Multimodal
  • ProbeMatchDTI improves DTI prediction by addressing the limitations of passive feature aggregation in existing methods.
  • The framework includes IterProbe and BindingProbe, which enhance the modeling of multi-scale biochemical correspondences.
  • Extensive experiments show that ProbeMatchDTI achieves superior performance on standard DTI benchmarks.
  • The model's predictions can be effectively integrated into drug-discovery workflows for candidate refinement.
Read more
DiDrive: A Risk-Aware Hierarchical Diffusion Framework for Safe Offline Reinforcement Learning in Autonomous Driving
Qisong Guo, Jingtang Chen, Zhilin Chen, Pei Xu, Mingjian Fu, Wenxi Liu, Yuanlong Yu
Reinforcement Learning Generative Models Robotics
  • DiDrive addresses the challenges of distribution shift and OOD actions in offline reinforcement learning for autonomous driving.
  • The framework consists of two main components: RHDif for state representation and 3DICE for action optimization.
  • Experimental results show DiDrive outperforms existing methods in terms of success rate and average reward in complex traffic scenarios.
  • The integration of risk-aware features and distribution correction enhances the safety and reliability of autonomous driving systems.
Read more
Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion
Md Abrar Jahin, Taufikur Rahman Fuad, Jay Pujara, Craig A. Knoblock
Graph Learning
  • Introduces QUEST, a method for initializing entity embeddings in UKGs using spectral techniques.
  • Implements a mini-batch Dirichlet energy regularizer to enforce structural consistency during training.
  • Demonstrates significant improvements in confidence and link prediction metrics over existing methods.
  • Addresses training instability issues commonly observed in dense graphs.
Read more
Entangled Representations Amplify Collateral Damage in Unlearning
Evžen Wybitul, Tim G. J. Rudner, Christian Schroeder de Witt
Interpretability Large Language Models
  • Entangled representations in neural networks complicate the unlearning process.
  • The study provides direct experimental evidence linking entanglement to collateral damage in unlearning.
  • More disentangled models incur lower retain costs during unlearning, confirming the intuition that entanglement negatively impacts unlearning efficiency.
  • The methodology can be adapted to test other structural claims in interpretability research.
Read more
A Unified Particle Filter LSTM for Data-Driven Process Simulation
Parvin Malekzadeh, Opher Baron, Dmitry Krass
Time Series
  • Introduces a Unified PF-LSTM that maintains multiple recurrent-state hypotheses to capture latent-state uncertainty.
  • Utilizes a particle filter mechanism to update beliefs about process states, enhancing prediction accuracy.
  • Demonstrates significant improvements in routing and sojourn time predictions across three emergency department datasets.
  • Shows that the framework is applicable to various sequential architectures beyond LSTMs.
Read more
Median-of-Means as an Extremal Convex Estimator and a Nonconvex Route to the Trimmed Oracle
Angshul Majumdar
Theory Optimization Efficient ML
  • Establishes a deterministic optimization framework for median-of-means (MoM) estimation.
  • Demonstrates that convex block M-estimators cannot achieve the trimmed-block oracle constant.
  • Introduces a nonconvex block-Lp family that interpolates between MoM and trimmed-block performance.
  • Shows that the energy landscape of nonconvex objectives is structured and favorable for optimization.
Read more
RINSE: Robust Target-Time Normality Estimation for Zero-Shot Graph Anomaly Detection
Taufikur Rahman Fuad, Md Abrar Jahin, Amir Hussain
Graph Learning
  • RINSE provides a gradient-free approach for estimating target normality in zero-shot graph anomaly detection.
  • The framework utilizes low-residual nodes to build a reliable normality model without requiring target labels.
  • It combines multiple evidence sources through reliability gating and rank-based ensembling to improve detection accuracy.
  • RINSE outperforms existing methods in terms of AUPRC across multiple unseen target graphs.
Read more
Differentiable Electricity-Market Clearing for Gradient-Based Planning
Luca Mungo, Maarten P. Scholl, Arnau Quera-Bofarull
Optimization
  • Introduces a differentiable optimization framework for electricity market clearing.
  • Enables gradient-based planning for data center load allocation.
  • Demonstrates high accuracy in approximating optimal solutions with gradient optimization.
  • Identifies challenges in handling discrete site transitions in planning.
Read more
Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents
Yanting Yang, Can Jin, Jinman Zhao, Jiahao Wu, Yang Zhou, Zhepeng Wang, Zhendong Wang, Mu Zhou, Dimitris N. Metaxas
Large Language Models Reinforcement Learning Robotics
  • Identifies variable-length action chunking as a critical capability for LLM agents.
  • Proposes SPACE, which uses programmatic skills to guide chunk learning and optimize decision-making.
  • Demonstrates significant improvements in task success rates and reductions in decision rounds compared to existing methods.
  • Achieves strong performance with fewer training steps, indicating higher training efficiency.
Read more
FlashKAN: B-Spline KANs via Truncated Power Form
Naveen Mysore
Efficient ML Theory Interpretability
  • Introduces a non-recursive method for evaluating B-spline activations in KANs.
  • Achieves significant speed improvements by fusing operations into a single GPU kernel.
  • Implements a stabilization technique to prevent numerical issues in B-spline evaluations.
  • Provides an open-source package for easy integration into existing frameworks.
Read more
Source-Free Class Relearning: Diagnosing Forgetting in Class Unlearning
Zahra Dehghani, Pablo Piantanida, Mohammadhadi Shateri
Theory Computer Vision Efficient ML
  • Introduces a source-free method for diagnosing class unlearning recoverability.
  • Establishes a theoretical alignment condition for class relearning.
  • Proposes the Relearning Score (RS) to measure recoverability and retain accuracy.
  • Demonstrates significant recoverability in state-of-the-art unlearning methods across multiple datasets.
Read more
Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts
Sanjaya Poudel, Nirajan Kunwor, Manish Dhakal, Debesh Jha, Sunil Kumar Gaire
Federated Learning Multimodal
  • First systematic benchmark of federated LoRA-based PEFT for biomedical vision-language models across multiple cohorts.
  • Federated LoRA adaptation significantly improves shared-class AUC from 0.687 to 0.802.
  • SVD-based aggregation is crucial for effective model updates, outperforming naive averaging.
  • Federation enhances performance in weaker cohorts while preserving strengths in stronger datasets.
Read more
CACTUS: Mask-Guided Semantic Clean-Label Backdoors in Decentralized Federated Learning
Chao Feng, Burkhard Stiller
Federated Learning Audio & Speech NLP
  • CACTUS utilizes semantic triggers for stealthy backdoor attacks in decentralized federated learning.
  • The method constructs label-consistent semantic pairs to create target-directed representation shifts.
  • Experiments show CACTUS achieves a mean attack success rate of 51.2% across various modalities with 30% malicious nodes.
  • The attack's effectiveness varies with network topology and increases with the ratio of malicious nodes.
Read more