AI-generated summaries
Today's ML research,
without the noise.
Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.
24
Papers today
8h
Update frequency
7
Days of history
Rethinking the Teacher-Student Framework for Test-Time Adaptation
Computer Vision
Theory
Optimization
- The EMA teacher in TTA does not prevent model collapse over long test sequences.
- Introducing a fixed-weight teacher (Intransigent Teacher) effectively mitigates error accumulation.
- The IT approach allows students to potentially surpass their teachers in performance.
- The proposed method enhances robustness to hyperparameter changes and is applicable across various architectures.
Read more
Rethinking the Teacher-Student Framework for Test-Time Adaptation
Summary
This paper addresses the limitations of the teacher-student framework in Test-Time Adaptation (TTA), particularly the common practice of using an Exponential Moving Average (EMA) for teacher weights. The authors argue that this method does not effectively prevent error accumulation, especially over longer test sequences, leading to potential model collapse. They propose a novel approach by introducing an 'Intransigent Teacher' (IT) that maintains fixed weights during adaptation. This adjustment significantly enhances performance across various datasets and scenarios, demonstrating improved robustness against hyperparameter changes. The findings challenge existing assumptions about the stability of the teacher-student paradigm and suggest that a more stable teacher can lead to better long-term adaptation outcomes. The proposed IT method can be easily integrated into existing TTA frameworks, providing a simple yet effective baseline for future research.
Methodology
The authors conducted a series of experiments to analyze the performance of TTA methods using both EMA and fixed-weight teachers. They evaluated the impact of these strategies on model accuracy over extended test sequences and assessed the stability-plasticity trade-off within the teacher-student framework. The proposed Intransigent Teacher was tested against standard benchmarks to demonstrate its effectiveness.
Results
The introduction of the Intransigent Teacher led to significant improvements in average classification accuracy, particularly in long test scenarios. The results indicated that while traditional methods exhibited performance degradation with increased sequence length, the IT approach maintained robust accuracy, effectively preventing model collapse and enhancing overall performance.
Implications
The findings suggest that TTA methods can be improved by reconsidering the teacher-student framework, particularly in scenarios involving long sequences or significant distribution shifts. The Intransigent Teacher approach could be widely adopted in various machine learning applications where real-time adaptation is crucial, such as in autonomous systems and dynamic environments.
Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators
Theory
- Introduces soft-label-based estimators for optimal BER and AUC in binary classification.
- Addresses the limitations of traditional accuracy metrics in imbalanced or noisy datasets.
- Extends the FeeBee framework for practical evaluation of estimators without requiring knowledge of the optimum.
- Validates the proposed methods through experiments on synthetic and real-world datasets.
Read more
Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators
Summary
This paper addresses the estimation of Bayes-optimal performance metrics in binary classification, specifically the balanced error rate (BER) and the area under the ROC curve (AUC). The authors propose novel soft-label-based estimators for these metrics, which are crucial in scenarios with class imbalance or noisy annotations. The first part of the contribution focuses on the estimation of optimal BER and AUC under two settings: a clean scenario where true soft labels and class priors are known, and a more realistic scenario where soft labels are corrupted by unknown transformations and noise. The authors employ isotonic regression to recover clean soft labels and derive finite-sample error bounds for the estimators. The second part extends the FeeBee framework for evaluating Bayes-error estimators to the optimal BER and AUC, allowing for practical evaluation of any estimator without needing to know the optimum. Experiments on synthetic and real-world datasets demonstrate the effectiveness of both the proposed estimators and the evaluation procedure, highlighting their robustness in challenging conditions.
Methodology
The authors propose soft-label-based estimators for optimal BER and AUC, starting with a clean setting where true soft labels are known. They then extend these estimators to a more realistic scenario involving unknown order-preserving transformations and noise. Isotonic regression is used to recover clean soft labels, and a clipped mean is employed to estimate the class prior. Finite-sample error bounds for the plug-in estimators are derived. The evaluation framework is adapted from FeeBee to accommodate the new metrics.
Results
The proposed estimators for optimal BER and AUC were validated through experiments, showing that they effectively estimate the metrics even in the presence of class imbalance and label noise. The evaluation procedure provided reliable scores for the estimators without needing to know the true optimal values.
Implications
The findings have significant implications for improving model evaluation in machine learning, particularly in applications where class imbalance and noisy labels are prevalent. The proposed methods can enhance the understanding of model performance and guide further improvements in model training.
On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers
NLP
Reinforcement Learning
Efficient ML
- Introduces a two-stage framework combining off-policy teacher optimization and on-policy student distillation.
- Demonstrates significant performance improvements in compact instruction-following rerankers, especially under distribution shifts.
- The proposed method outperforms traditional offline distillation approaches by a notable margin.
- Achieves a favorable quality-efficiency tradeoff, making it suitable for deployment in production environments.
Read more
On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers
Summary
This paper addresses the challenge of training compact instruction-following rerankers, which are essential for efficient deployment in retrieval systems. Traditional distillation methods limit student models to offline imitation of teacher outputs, potentially inheriting biases and failing to generalize under distribution shifts. The authors propose a novel two-stage framework that integrates off-policy teacher optimization with on-policy student distillation. In the first stage, a large 4B teacher reranker is enhanced using off-policy Generalized Reinforcement Policy Optimization (GRPO) based on feedback from a large language model (LLM) on a dataset of 88K instruction-following examples. The second stage involves training a compact 1B student model that samples its own rankings and receives soft rewards derived from the teacher, promoting exploration and effective knowledge transfer. The results demonstrate significant performance improvements, particularly under distribution shifts, with the proposed student model achieving superior nDCG scores compared to traditional offline distillation methods. The findings suggest that this approach not only enhances the performance of compact rerankers but also allows for efficient deployment without sacrificing quality.
Methodology
The methodology consists of a two-stage training process: Stage 1 involves optimizing a large 4B teacher reranker using off-policy GRPO with LLM-judge feedback, while Stage 2 focuses on training a compact 1B student reranker that samples its own rankings and receives soft rewards from the teacher's feedback. This approach leverages reinforcement learning principles to enhance exploration and knowledge transfer.
Results
The proposed method achieved an nDCG@6 score of 0.7670 on the MAIR-11 evaluation, surpassing offline listwise knowledge distillation by 4.6 points. On the broader MAIR-Full benchmark, it reached 0.6808 nDCG@6 and 0.7865 MRR@6, outperforming two released 7B RL-trained rerankers. The compact 1B reranker also achieved 0.7624 nDCG@6 on a 9,861-query validation benchmark, demonstrating a strong quality-efficiency tradeoff.
Implications
The findings suggest that the proposed on-policy distillation approach can effectively train compact models for instruction-following tasks, making it a viable option for deployment in real-world retrieval systems. This could lead to more efficient and responsive AI systems in enterprise and assistant-facing applications.
CliffRank: A Dual-Branch Framework for Activity-Cliff Ranking Prediction
Optimization
Theory
- CliffRank effectively combines absolute-activity regression with ranking-consistency learning.
- The framework utilizes two parallel predictors to enhance prediction accuracy for activity cliffs.
- CliffRank achieved the highest mean Spearman correlation of 0.6890 on small-molecule datasets.
- Performance varied across datasets, highlighting the importance of dataset characteristics in model evaluation.
Read more
CliffRank: A Dual-Branch Framework for Activity-Cliff Ranking Prediction
Summary
The paper presents CliffRank, a dual-branch framework designed to improve the prediction of activity cliffs (ACs) in antimicrobial peptides and small molecules. Activity cliffs are significant because small structural changes can lead to large variations in biological activity, yet high-quality data to understand these mechanisms is scarce. CliffRank combines absolute-activity regression with ranking-consistency learning, employing two parallel predictors that utilize mean squared error, a thresholded listwise loss, and Pairwise Preference Consistency (PPC) to align the relative ordering in preference-probability space. The framework was evaluated on three datasets of antimicrobial peptides and three datasets of small molecules, demonstrating superior performance in terms of Spearman correlation and Recall@50 metrics. The results indicate that while CliffRank achieved the highest mean Spearman correlation of 0.6890 on small-molecule datasets, its performance varied across datasets, suggesting the need for further exploration of adaptive PPC schedules and the incorporation of contextual information in future work.
Methodology
CliffRank employs a dual-branch framework that integrates absolute-activity regression with ranking-consistency learning. It trains two parallel predictors using mean squared error, a thresholded listwise loss, and Pairwise Preference Consistency (PPC) to align the relative ordering of predictions in the preference-probability space.
Results
On antimicrobial peptide datasets, CliffRank with ESM2-t12 achieved a mean Spearman correlation of 0.5393 and a mean Recall@50 of 21.4. For small-molecule datasets, CliffRank with PNA reached a mean Spearman correlation of 0.6890 and a mean Recall@50 of 30.4, matching the performance of ACANet-PNA. The results also indicated that asymmetric initialization improved some averages but not all targets.
Implications
The findings suggest that CliffRank can significantly enhance the prediction of activity cliffs, which is crucial for drug discovery and lead optimization. The methodology may be applicable to other domains where understanding the impact of small structural changes on activity is essential.
LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates
NLP
Large Language Models
Optimization
- LoRA-TSD optimizes low-rank updates by considering the geometry of the weight space, improving efficiency.
- The method provides a computationally cheaper retraction compared to traditional truncated-SVD approaches.
- The paper establishes global convergence guarantees for both LoRA-TSD and LoRA-Pro.
- LoRA-TSD outperforms existing LoRA optimizers across multiple benchmarks, demonstrating stability across adapter ranks.
Read more
LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates
Summary
The paper introduces LoRA-TSD, a novel optimizer designed for low-rank adaptation (LoRA) in fine-tuning large models. Traditional methods for updating LoRA factors independently often neglect the geometric implications of these updates, leading to inefficiencies. LoRA-TSD addresses this by treating each LoRA update as a tangent vector on the fixed-rank matrix manifold and applying a spectral-norm steepest-descent step within that tangent space. This method utilizes a retraction that is computationally cheaper than previous approaches, specifically the truncated-SVD retraction. The authors provide theoretical guarantees for the convergence of both LoRA-TSD and its Frobenius-norm counterpart, LoRA-Pro, establishing a new measure of stationarity that is computable from factor gradients alone. Empirical evaluations demonstrate that LoRA-TSD consistently outperforms existing LoRA optimizers across various benchmarks, showcasing robustness to adapter rank variations.
Methodology
LoRA-TSD employs tangent-space spectral descent, treating updates as tangent vectors on a fixed-rank matrix manifold. It utilizes a spectral-norm steepest-descent approach and a factor-induced retraction that avoids the need for full weight matrix operations, making it computationally efficient.
Results
LoRA-TSD was tested on six benchmarks using various model sizes (Llama-3.2-1B, Llama-3.1-8B, Qwen3-32B) and consistently outperformed all competing LoRA optimizers. The method also demonstrated robustness to variations in adapter rank, with theoretical convergence guarantees provided.
Implications
The development of LoRA-TSD has significant implications for efficient fine-tuning of large language models, particularly in scenarios where computational resources are limited. Its ability to optimize low-rank adaptations effectively could enhance the deployment of parameter-efficient models in real-world applications.
UE5M3 FP4 Block Scaling for Stable Language Model Pretraining
NLP
Large Language Models
Efficient ML
- Introduction of a UE5M3 FP4 pretraining recipe that simplifies the training process.
- Demonstrated lower training and validation losses compared to the existing NVFP4 method.
- Achieved a 21.2% increase in model-body token throughput by optimizing the execution process.
- Identified a range bound linking UE5M3 scale targets to model performance.
Read more
UE5M3 FP4 Block Scaling for Stable Language Model Pretraining
Summary
This technical report presents a novel approach to stable 4-bit floating-point (FP4) pretraining for language models, addressing the challenges posed by the narrow dynamic range of the E2M1 payload. The authors propose a simpler FP4 pretraining recipe that utilizes unsigned E5M3 (UE5M3) block scales, allowing for periodic tensor scaling and selective stochastic rounding of backward gradients. The methodology omits the randomized Hadamard transform (RHT) and employs FP4 in all eligible internal linear operations. The report details the pretraining of a Nemotron-H 8B model on nearly 190 billion tokens, demonstrating that the proposed method achieves lower final-window training loss and improved validation loss compared to the existing Transformer Engine NVFP4 approach. The findings indicate that the UE5M3 FP4 pretraining can enhance model performance while simplifying the training process, suggesting a potential shift towards native support for UE5M3 block scaling in future implementations.
Methodology
The authors developed a UE5M3 FP4 recipe that incorporates periodic sample-and-hold tensor scaling, 2D weight scaling, and selective stochastic rounding for upstream gradients in backward GEMMs. The approach avoids RHT and utilizes FP4 for all internal linear operations, focusing on stability and efficiency during pretraining.
Results
The pretraining of the Nemotron-H 8B model resulted in lower final-window training loss and lower validation loss measured as held-out negative log-likelihood compared to the NVFP4 method. The proposed block-16 recipe outperformed NVFP4 in downstream point estimates across multiple aggregates, and an ablation study showed a 21.2% increase in throughput when RHT was removed.
Implications
The findings suggest that the UE5M3 FP4 pretraining approach could lead to more efficient training of large language models, potentially influencing future research and development in model optimization and quantization techniques. The results advocate for the adoption of simpler, more effective training recipes in the field of machine learning.
hLLM: Single Pass Decoding for Generative Reranking
NLP
Large Language Models
Optimization
- HLLM introduces a novel decoding strategy that reduces the number of forward passes required for generative ranking from O(N·T) to O(1).
- The method utilizes a lightweight self-attention mechanism to create a score matrix from LLM hidden states, enabling efficient optimal assignment via the Hungarian algorithm.
- Empirical evaluations show a 64× speed-up in inference time while preserving ranking quality on both proprietary and open datasets.
- The framework connects generative ranking with combinatorial optimization, paving the way for further advancements in real-time ranking systems.
Read more
hLLM: Single Pass Decoding for Generative Reranking
Summary
The paper introduces HLLM (Hungarian LLM), a novel decoding strategy for generative ranking that significantly improves the efficiency of large language models (LLMs) in producing ranked outputs. Traditional autoregressive decoding requires a sequential forward pass for each emitted token, leading to high latency, especially when ranking multiple candidates. HLLM addresses this by recognizing that the output consists of N ordinal values that can be decoded in a single pass using optimal assignment techniques. The authors leverage the hidden states from the LLM's prefill phase to construct an N × K item-position score matrix, which is then processed using the Hungarian algorithm to derive the optimal permutation of rankings. This method allows for O(1) forward passes, resulting in a dramatic reduction in inference time. The paper presents empirical results demonstrating that HLLM achieves a 64× speed-up in end-to-end inference time while maintaining ranking quality comparable to traditional methods. Additionally, the authors conduct a comprehensive ablation study to isolate the contributions of various architectural and training components, further validating the effectiveness of their approach.
Methodology
The authors developed HLLM by reformulating the ordinal decoding problem as an optimal assignment task. They utilized the hidden states from the LLM's prefill phase to create a score matrix and applied the Hungarian algorithm to decode the rankings in a single pass. The approach was fine-tuned using LoRA-based techniques and teacher ranking distillation to optimize performance.
Results
HLLM achieved an end-to-end inference time of 28 ms, representing a 64× speed-up compared to traditional autoregressive methods, while maintaining comparable ranking performance. The results were validated on both proprietary datasets and the Amazon Beauty open dataset, confirming the robustness of the approach across different contexts.
Implications
The HLLM framework has significant implications for applications requiring real-time ranking, such as recommendation systems, search engines, and advertising platforms. By reducing latency while maintaining quality, it enables more efficient user experiences and can facilitate the deployment of LLMs in time-sensitive environments.
Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT
NLP
Efficient ML
- Introduces an efficient SciBERT-based approach for telescope bibliography classification.
- Achieves a macro F1 score of 0.89, ranking first in the WASP-2025 Shared Task.
- Analyzes the impact of truncation on classification performance.
- Demonstrates the advantages of domain-specific pretraining over general models.
Read more
Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT
Summary
This paper addresses the challenge of automating the classification of telescope bibliographies, which is essential for assessing the scientific impact of observatories and ensuring reproducibility in astronomy. The author presents a novel approach using SciBERT, a domain-specific language model, to classify scientific papers into four categories: science, instrumentation, mention, and not telescope. The study highlights the difficulties posed by strict context-length constraints (maximum 512 tokens) and limited computational resources. Despite these challenges, the proposed method achieved a macro F1 score of 0.89, placing it at the top of the WASP-2025 leaderboard. The paper also explores the effects of truncation on classification performance and discusses trade-offs between truncation, chunking, and long-context models, providing valuable insights into efficient scientific text curation.
Methodology
The methodology involved using SciBERT for multi-label classification of scientific papers. The input text was constructed by concatenating various sections of the papers, and a maximum sequence length of 512 tokens was enforced, leading to truncation of longer samples. The model was fine-tuned using a BCEWithLogitsLoss function and trained on Kaggle GPU resources. A baseline model using TF-IDF and Logistic Regression was also established for comparison.
Results
The SciBERT model outperformed the baseline, achieving a macro F1 score of 0.89 despite processing less than half of the full context due to truncation. This indicates that critical information for classification is concentrated in the title, abstract, and acknowledgments sections of the papers.
Implications
The findings suggest that automating bibliography classification can significantly enhance the efficiency of scientific data curation in astronomy. The insights gained from this study could inform future research on efficient text processing and classification in other scientific domains.
DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation
Reinforcement Learning
Optimization
- DMRL connects skill document editing to downstream parameter optimization through structured interfaces.
- Introduces DRPO for robust advantage estimation and LRP for long-term outcome prediction.
- Addresses the challenges of credit assignment and population heterogeneity in advertising recommendations.
- Demonstrated significant online improvements in a real-world advertising platform.
Read more
DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation
Summary
The paper introduces Document-Mediated Reinforcement Learning (DMRL), a novel framework aimed at optimizing skill documents for advertising recommendations. Traditional methods for advertising parameter tuning are labor-intensive and often rely on manual adjustments, leading to inefficiencies and suboptimal performance. DMRL addresses these limitations by modeling skill document optimization as a series of structured editing actions. It employs an upper-level agent to perform controlled edits on skill documents, while a lower-level task agent evaluates the impacts of these edits through A/B testing. The framework incorporates two innovative components: Dual-Relative Policy Optimization (DRPO), which enhances advantage estimation for robust and risk-aware decision-making, and Long-term Reward Predictor (LRP), which estimates long-term outcomes by accounting for population heterogeneity through disentangled representation learning. The DMRL framework was tested on a large-scale short-video advertising platform, demonstrating significant improvements over existing state-of-the-art methods across key advertising metrics, thus providing a principled approach to skill optimization in advertising systems.
Methodology
The DMRL framework consists of three main components: a structural decoupling mechanism that separates skill optimization from task execution, DRPO for robust advantage estimation, and LRP for predicting long-term outcomes. The methodology employs a two-stage training strategy to stabilize predictions and ensure accurate advantage estimation, allowing for effective skill document edits and parameter tuning.
Results
Extensive empirical evaluations on a large-scale short-video advertising platform showed that DMRL outperformed state-of-the-art baselines across key advertising metrics, indicating its effectiveness in optimizing advertising recommendations.
Implications
The DMRL framework has significant implications for the advertising industry, providing a more efficient and principled method for skill optimization. It can potentially reduce the reliance on manual tuning, enhance user experience, and improve commercial returns through better parameter adjustments.
D-FROST: Decentralized Federated pRompt-tuning via Optimal tranSporT for Non-IID and Imbalanced Data
Federated Learning
Optimization
Computer Vision
- First study of prompt tuning in decentralized federated learning (DFL).
- Formulates decentralized prompt tuning as a Wasserstein-based optimization problem.
- Introduces D-FROST, an optimal-transport-based algorithm for merging prompts.
- Provides convergence analysis showing stability and consensus in prompt tuning.
Read more
D-FROST: Decentralized Federated pRompt-tuning via Optimal tranSporT for Non-IID and Imbalanced Data
Summary
The paper introduces D-FROST, a novel approach to decentralized federated prompt tuning, addressing the challenges posed by non-IID and imbalanced data. Prompt tuning is a parameter-efficient method for adapting foundation models by optimizing a small set of learnable prompts while keeping the backbone model frozen. However, in decentralized federated learning (DFL), the lack of index-wise alignment among prompts learned from heterogeneous local data complicates the standard averaging methods. The authors formulate decentralized prompt tuning as a Wasserstein-based optimization problem, which captures the set-valued structure of prompts. D-FROST employs an optimal-transport-based algorithm to merge neighborhood prompts into compact representative sets, avoiding misalignment issues. The paper also provides a convergence analysis, demonstrating that D-FROST effectively controls the Wasserstein consensus error and converges to a neighborhood of stationarity for the shared objective. Experimental results across diverse vision datasets show that D-FROST outperforms existing decentralized federated prompt-tuning baselines, particularly in scenarios with data imbalance and heterogeneity.
Methodology
The authors formulate decentralized prompt tuning as a Wasserstein optimization problem, which allows for the merging of prompts based on their geometric relationships rather than direct index-wise averaging. D-FROST operates by having each client update local prompts and then applying an optimal transport-based merge function to create a compact representative prompt set. A convergence analysis is conducted to ensure that the algorithm stabilizes and converges towards a shared objective.
Results
The experimental evaluation of D-FROST across eight diverse vision datasets indicates that the proposed method consistently outperforms existing decentralized federated prompt-tuning techniques, particularly in scenarios characterized by data imbalance and extreme heterogeneity.
Implications
D-FROST presents a significant advancement in decentralized federated learning, particularly for applications involving large foundation models in environments with non-IID and imbalanced data. This approach could facilitate more efficient model adaptation in real-world applications where data privacy and communication costs are critical concerns.
CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
NLP
Large Language Models
Efficient ML
- CRISP introduces a structural proxy (Cstruct) that simplifies routing decisions in attention mechanisms.
- The method addresses the issue of noise accumulation in attention mass through sink-aware thresholding.
- Empirical results show CRISP achieves up to 5.30× speedup in attention computation while maintaining performance.
- The approach effectively adapts to long-context inputs, improving retrieval task performance significantly.
Read more
CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
Summary
The paper introduces CRISP (Cliff-awaRe Input-adaptive Sparse Prefilling), a novel approach to enhance the efficiency of long-context large language model (LLM) inference by addressing the computational bottleneck of self-attention mechanisms. Traditional sparse attention methods either use fixed patterns or rely on offline profiling, which limits their adaptability to input-dependent structures. Recent dynamic methods, such as FlexPrefill, attempt to improve adaptability through real-time routing of attention heads but face challenges due to indirect routing proxies and cumulative coverage thresholds that lead to noise accumulation. CRISP overcomes these limitations by proposing two key innovations: (1) a structural proxy, Cstruct, which directly measures attention concentration without the overhead of Jensen-Shannon Divergence (JSD), and (2) a sink-aware thresholding mechanism that effectively separates signal from noise in attention mass distributions. The empirical results demonstrate that CRISP outperforms existing sparse methods and matches or exceeds the performance of dense attention on retrieval-heavy benchmarks, achieving significant speedups in attention computation.
Methodology
CRISP employs a two-pronged approach: it utilizes Cstruct to directly assess the concentration of attention heads based on structural positions, eliminating the need for complex calculations associated with JSD. Additionally, it implements a sink-aware thresholding mechanism to navigate the mass cliff in attention distributions, ensuring that only relevant signals are selected while minimizing background noise.
Results
CRISP demonstrated superior performance across multiple benchmarks (InfiniteBench, RULER, LongBench) on two model families, achieving up to +28.0 percentage points improvement on retrieval tasks compared to baseline methods and a remarkable 5.30× speedup in attention computation at 512k tokens.
Implications
The advancements presented in CRISP could lead to more efficient implementations of large language models, particularly in applications requiring long-context processing and real-time adaptability, such as conversational AI, document retrieval, and other NLP tasks.
ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction
Multimodal
- ProbeMatchDTI improves DTI prediction by addressing the limitations of passive feature aggregation in existing methods.
- The framework includes IterProbe and BindingProbe, which enhance the modeling of multi-scale biochemical correspondences.
- Extensive experiments show that ProbeMatchDTI achieves superior performance on standard DTI benchmarks.
- The model's predictions can be effectively integrated into drug-discovery workflows for candidate refinement.
Read more
ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction
Summary
The paper presents ProbeMatchDTI, a novel framework for predicting drug-target interactions (DTI) that addresses limitations in existing biochemical representation learning methods. Traditional approaches often favor dominant molecular patterns, neglecting weaker yet relevant signals such as functional groups and residue-context patterns. ProbeMatchDTI introduces two key components: IterProbe and BindingProbe. IterProbe retains contextual states across different refinement depths and utilizes learnable probes to select relevant biochemical patterns before cross-entity matching. This mechanism preserves weak biochemical signals and enhances the association between functional groups and local motifs. BindingProbe further characterizes drug-protein complementarity at both local and global levels, enabling the model to jointly analyze fine-grained interactions and multi-scale correspondences. Extensive experiments demonstrate that ProbeMatchDTI outperforms existing methods, achieving higher AUC-ROC scores on BindingDB and DrugBank datasets. The framework's probe-driven behavior is analyzed through feature-level pattern analyses, and its predictions are integrated into a downstream drug-discovery workflow, showcasing its practical utility for candidate refinement and validation planning.
Methodology
ProbeMatchDTI employs a two-step mechanism consisting of IterProbe and BindingProbe. IterProbe retains and weighs contextual states at each position using learnable probes, allowing for the preservation of weak biochemical patterns. BindingProbe characterizes drug-protein interactions at both local and global levels, enabling a comprehensive analysis of DTI.
Results
ProbeMatchDTI achieved 2.0% and 0.5% higher AUC-ROC scores on the BindingDB and DrugBank datasets, respectively, compared to existing methods. The framework demonstrated effective preservation of weak binding-relevant associations and improved modeling of multi-scale biochemical correspondences.
Implications
The findings suggest that ProbeMatchDTI can significantly enhance the accuracy of DTI predictions, which is crucial for AI-driven drug discovery. Its integration into drug-discovery workflows can facilitate better candidate refinement and validation processes, potentially accelerating the development of new therapeutics.
DiDrive: A Risk-Aware Hierarchical Diffusion Framework for Safe Offline Reinforcement Learning in Autonomous Driving
Reinforcement Learning
Generative Models
Robotics
- DiDrive addresses the challenges of distribution shift and OOD actions in offline reinforcement learning for autonomous driving.
- The framework consists of two main components: RHDif for state representation and 3DICE for action optimization.
- Experimental results show DiDrive outperforms existing methods in terms of success rate and average reward in complex traffic scenarios.
- The integration of risk-aware features and distribution correction enhances the safety and reliability of autonomous driving systems.
Read more
DiDrive: A Risk-Aware Hierarchical Diffusion Framework for Safe Offline Reinforcement Learning in Autonomous Driving
Summary
The paper presents DiDrive, a novel framework designed to enhance safety in offline reinforcement learning (RL) for autonomous driving. The authors identify significant challenges in applying offline RL, particularly the issues of distribution shift and the risk of generating out-of-distribution (OOD) actions in complex traffic environments. DiDrive integrates two main components: the Risk-Aware Hierarchical Diffusion (RHDif) architecture, which manages input complexity by enhancing local danger features and aligning global semantics, and the Distribution Correction Estimation with Diffusion for Driving (3DICE) policy optimization paradigm, which mitigates OOD overestimation and gradient oscillation. The framework is tested in the CARLA simulation benchmark, demonstrating improved performance metrics such as success rate and average reward, particularly in high-density traffic scenarios. The study concludes that DiDrive effectively combines risk-aware representation with distribution correction optimization, providing a robust approach for safe decision-making in autonomous driving.
Methodology
The methodology involves a two-component framework: RHDif, which uses a low-level risk-gated encoder and a high-level contextual modulator to refine state representations, and 3DICE, which employs a four-stage policy optimization process to ensure safety and mitigate OOD risks. The diffusion model is guided by in-sample historical experience to adjust its denoising trajectory, focusing on high-quality actions within the training data support.
Results
In experiments conducted on the CARLA simulation benchmark, DiDrive achieved a success rate of 85% and an average reward of 4295.68 in high-density scenarios with 60 background vehicles, outperforming baseline methods such as IQL, CQL, and Diffusion-QL in terms of overall performance metrics.
Implications
The findings suggest that DiDrive can significantly enhance the safety and reliability of autonomous driving systems, particularly in complex traffic environments. This framework may be applicable to other domains requiring safe decision-making under uncertainty.
Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion
Graph Learning
- Introduces QUEST, a method for initializing entity embeddings in UKGs using spectral techniques.
- Implements a mini-batch Dirichlet energy regularizer to enforce structural consistency during training.
- Demonstrates significant improvements in confidence and link prediction metrics over existing methods.
- Addresses training instability issues commonly observed in dense graphs.
Read more
Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion
Summary
This paper addresses the challenge of Uncertain Knowledge Graph (UKG) completion, where each relational triple is assigned a continuous confidence score. Traditional methods often overlook the global community and hub structures inherent in the confidence-weighted graph, leading to suboptimal initialization of entity embeddings. The authors introduce QUEST, a novel approach that enhances the initialization of entity embeddings using the smallest non-trivial eigenvectors of the confidence-weighted graph Laplacian, thereby incorporating structural information before training. Additionally, QUEST employs a mini-batch Dirichlet energy regularizer to maintain structural consistency during the early stages of training. The proposed method does not add any trainable parameters to the existing confidence-distribution learning pipeline. The results demonstrate that QUEST significantly improves confidence and link prediction across multiple datasets, outperforming previous methods in six out of eight metric-dataset pairs and stabilizing training processes, particularly in dense graphs. This work highlights the importance of integrating spectral structural priors and graph smoothness regularization in enhancing the performance and reliability of UKG completion tasks.
Methodology
The methodology involves two main components: spectral initialization of entity embeddings using the smallest non-trivial eigenvectors of the confidence-weighted graph Laplacian, and the application of an unbiased mini-batch Dirichlet energy regularizer to ensure structural smoothness during the early training phases. The QUEST framework is designed to integrate seamlessly into existing confidence-distribution learning pipelines without introducing additional trainable parameters.
Results
On two UKG datasets, QUEST achieved improved performance in confidence prediction and link prediction, surpassing prior methods in six out of eight metric-dataset pairs and matching the best results on the remaining two. The method also mitigated the instability spikes typically observed in dense graphs, indicating enhanced training stability and reliability.
Implications
The findings suggest that incorporating spectral structural priors and smoothness regularization can significantly enhance the accuracy and stability of UKG completion tasks. This approach may have broader applications in various domains where knowledge graphs are utilized, particularly in scenarios involving uncertain or incomplete data.
Entangled Representations Amplify Collateral Damage in Unlearning
Interpretability
Large Language Models
- Entangled representations in neural networks complicate the unlearning process.
- The study provides direct experimental evidence linking entanglement to collateral damage in unlearning.
- More disentangled models incur lower retain costs during unlearning, confirming the intuition that entanglement negatively impacts unlearning efficiency.
- The methodology can be adapted to test other structural claims in interpretability research.
Read more
Entangled Representations Amplify Collateral Damage in Unlearning
Summary
This paper investigates the impact of representational entanglement on the process of unlearning in neural networks, specifically focusing on how the sharing of structure between different knowledge domains complicates the unlearning process. The authors conduct a controlled experiment using Selective Gradient Masking (SGTM) to train six 254M-parameter language models on English Wikipedia, varying the degree of disentanglement between biology and non-biology knowledge. The study applies three standard unlearning methods to evaluate the retain-forget trade-offs across these models. The findings reveal that models with higher levels of disentanglement achieve significantly better retain-forget trade-offs, demonstrating that entanglement is a contributing factor to collateral damage during unlearning. This work provides empirical evidence supporting long-held intuitions in interpretability research and opens avenues for further exploration of structural properties in neural networks.
Methodology
The authors employed Selective Gradient Masking (SGTM) to train a suite of language models with varying levels of disentanglement. They systematically altered the training process to create models that specialized in either retaining or forgetting specific knowledge domains. Three unlearning methods—Weighted Gradient Ascent (WGA), Weight Divergence Regularization (WDR), and RMU—were applied to assess the retain-forget trade-offs across the models.
Results
The results indicate that models with higher disentanglement achieve better retain-forget trade-offs. Specifically, at a fixed level of forgetting, the most disentangled models showed approximately 4× lower retain costs under WGA and RMU, and 1.3× lower under WDR compared to the most entangled models.
Implications
These findings have significant implications for the design of neural networks, particularly in applications where unlearning is critical, such as privacy-sensitive data management. The results suggest that optimizing for disentanglement could enhance the effectiveness of unlearning methods, thereby improving model interpretability and reliability.
A Unified Particle Filter LSTM for Data-Driven Process Simulation
Time Series
- Introduces a Unified PF-LSTM that maintains multiple recurrent-state hypotheses to capture latent-state uncertainty.
- Utilizes a particle filter mechanism to update beliefs about process states, enhancing prediction accuracy.
- Demonstrates significant improvements in routing and sojourn time predictions across three emergency department datasets.
- Shows that the framework is applicable to various sequential architectures beyond LSTMs.
Read more
A Unified Particle Filter LSTM for Data-Driven Process Simulation
Summary
This paper presents a novel approach to data-driven process simulation through the development of a Unified Particle Filter LSTM (Unified PF-LSTM). Traditional deep sequence models, such as LSTMs, compress observed process histories into a single deterministic state, which can lead to inaccuracies due to the partial observability of event logs. The Unified PF-LSTM addresses this limitation by maintaining a weighted set of recurrent-state hypotheses, allowing for the representation of multiple plausible interpretations of the observed history. This is achieved through a particle filter mechanism that updates beliefs about the underlying process states. The model incorporates learned features based on the moment-generating function to enhance the prediction of next activities and sojourn times. The framework is trained end-to-end using real-world data from emergency departments, demonstrating superior performance over existing data-driven baselines in terms of routing, duration, and overall system behavior, particularly in complex scenarios where event logs provide limited information about the underlying dynamics.
Methodology
The Unified PF-LSTM employs a particle filter approach to maintain a weighted particle approximation of recurrent states. It updates these beliefs through a differentiable computational graph and utilizes learned features based on the moment-generating function to predict categorical distributions for next activities and conditional quantiles for sojourn times.
Results
The proposed framework outperformed traditional data-driven models in reproducing routing and duration behaviors across all evaluated datasets. It showed particularly strong gains in scenarios with complex process dynamics, achieving higher accuracy despite longer computational times.
Implications
The Unified PF-LSTM has potential applications in various fields requiring data-driven process modeling, such as healthcare, logistics, and operational management, where understanding complex, partially observable systems is crucial for decision-making and resource allocation.
Median-of-Means as an Extremal Convex Estimator and a Nonconvex Route to the Trimmed Oracle
Theory
Optimization
Efficient ML
- Establishes a deterministic optimization framework for median-of-means (MoM) estimation.
- Demonstrates that convex block M-estimators cannot achieve the trimmed-block oracle constant.
- Introduces a nonconvex block-Lp family that interpolates between MoM and trimmed-block performance.
- Shows that the energy landscape of nonconvex objectives is structured and favorable for optimization.
Read more
Median-of-Means as an Extremal Convex Estimator and a Nonconvex Route to the Trimmed Oracle
Summary
This paper revisits the median-of-means (MoM) estimation from a deterministic optimization perspective, proposing a family of block-Lp estimators designed for robust learning in the presence of heavy-tailed and adversarially corrupted data. The author establishes that every convex block M-estimator has a worst-case robustness constant of at least 1/(1 - 2ε), which aligns with the classical MoM bound, and demonstrates that the trimmed-block oracle constant 1/(1 - ε) cannot be achieved within the convex framework. The paper introduces a nonconvex block-Lp family, where p ranges from 0 to 1, and derives finite-sample deterministic robustness bounds for all global minimizers. As p decreases, these bounds transition continuously from the MoM constant to the block-L0 oracle constant, with small p values yielding global minimizers that align with oracle solutions under specific conditions. The energy landscape of the block-Lp objectives is shown to be favorable, with all local minima situated near the true parameter value and no adverse basins present. The findings are combined with block-level concentration to yield sub-Gaussian deviation bounds under finite (2 + δ) moments, extending to high-dimensional scenarios for robust mean estimation and sparse regression with optimal rates. This analysis positions MoM estimators along a continuous path that approaches trimmed-block performance while maintaining computational feasibility and applicability to contemporary robust learning challenges.
Methodology
The paper employs a deterministic optimization approach to analyze the median-of-means (MoM) and introduces a family of nonconvex block-Lp estimators. It derives robustness bounds for global minimizers under a block contamination model, and examines the energy landscape of the proposed objectives to ensure favorable optimization properties.
Results
The study finds that the robustness constant for convex block M-estimators is at least 1/(1 - 2ε), while the nonconvex block-Lp estimators achieve a continuous interpolation of robustness constants from MoM to the trimmed-block oracle. The energy landscape is shown to be benign, with local minima near the truth and no bad basins, leading to sub-Gaussian deviation bounds for robust estimation.
Implications
The findings suggest that nonconvex block-Lp estimators can significantly enhance robustness in learning tasks involving heavy-tailed and adversarially corrupted data, providing a structured optimization path that retains the advantages of classical MoM while approaching the performance of trimmed estimators.
RINSE: Robust Target-Time Normality Estimation for Zero-Shot Graph Anomaly Detection
Graph Learning
- RINSE provides a gradient-free approach for estimating target normality in zero-shot graph anomaly detection.
- The framework utilizes low-residual nodes to build a reliable normality model without requiring target labels.
- It combines multiple evidence sources through reliability gating and rank-based ensembling to improve detection accuracy.
- RINSE outperforms existing methods in terms of AUPRC across multiple unseen target graphs.
Read more
RINSE: Robust Target-Time Normality Estimation for Zero-Shot Graph Anomaly Detection
Summary
The paper introduces RINSE (Robust Iterative Normality Self-Estimation), a novel framework for zero-shot graph anomaly detection that addresses the challenges posed by domain shifts between source and target graphs. RINSE operates without the need for target labels or gradients, maintaining a fixed source-trained detector while iteratively estimating target normality. The method identifies a reliable subset of low-residual nodes from the target graph to construct a trimmed normality model. It employs a combination of reliability-gated rank fusion and encoder ensembling to enhance anomaly detection performance. The authors evaluate RINSE across eight unseen target graphs, demonstrating that it achieves the highest average Area Under the Precision-Recall Curve (AUPRC) compared to other methods under various preprocessing protocols. The results indicate that RINSE effectively estimates normality in contaminated data and supports robust generalist graph anomaly detection.
Methodology
RINSE employs a gradient-free iterative process to estimate target normality by focusing on low-residual nodes. It constructs a trimmed normality dictionary, calibrates embeddings based on this dictionary, and integrates evidence from various anomaly detection methods using reliability gating. The framework also utilizes rank-based ensembling of independently initialized encoders to enhance performance.
Results
RINSE achieved the highest average AUPRC across eight unseen target graphs under two preprocessing protocols. The results from block ablations and sensitivity analyses validate the effectiveness of the combined design, demonstrating that the method can robustly estimate normality without target labels or tuning.
Implications
The findings suggest that RINSE can be effectively applied in real-world scenarios where labeled data is scarce or unavailable, such as fraud detection in financial networks or anomaly detection in social media platforms. The framework's ability to handle contaminated data makes it a valuable tool for generalist graph anomaly detection.
Differentiable Electricity-Market Clearing for Gradient-Based Planning
Optimization
- Introduces a differentiable optimization framework for electricity market clearing.
- Enables gradient-based planning for data center load allocation.
- Demonstrates high accuracy in approximating optimal solutions with gradient optimization.
- Identifies challenges in handling discrete site transitions in planning.
Read more
Differentiable Electricity-Market Clearing for Gradient-Based Planning
Summary
This paper addresses the challenges of planning large data centers in electricity markets, where the facility's load can influence electricity prices through market clearing, a constrained optimization problem. The authors propose treating market clearing as a differentiable optimization layer, allowing for gradient-based planning. By using reverse-mode automatic differentiation, they derive gradients of planning costs with respect to investment decisions, enabling planners to understand how to adjust their strategies based on market responses. The methodology is validated through a case study involving the allocation of 50 MW of data-center load across six candidate buses in two synthetic networks, evaluated over 36 operating states. The results demonstrate that the gradient optimization approach closely approximates the optimal allocations, with minimal objective gaps. However, the method encounters challenges near discrete site transitions, where it tends to delay site closures. Overall, this work illustrates the potential of differentiable market clearing to enhance market-aware planning by providing actionable insights through gradient information.
Methodology
The authors treat market clearing as a differentiable optimization layer, using reverse-mode automatic differentiation to propagate gradients from market-clearing prices back to planning variables. This allows for the optimization of load allocation across candidate buses while accounting for the impact of load changes on electricity prices.
Results
The gradient optimization method successfully recovers near-optimal load allocations, achieving worst-case objective gaps of 2.3% and 8.5% compared to exhaustive enumeration. The method performs consistently across multiple initializations, although it struggles with discrete transitions near site closures.
Implications
This research has significant implications for data center planning and investment strategies in electricity markets. By integrating differentiable market clearing into planning processes, stakeholders can make more informed decisions that account for market dynamics, potentially leading to more efficient energy use and cost savings.
Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents
Large Language Models
Reinforcement Learning
Robotics
- Identifies variable-length action chunking as a critical capability for LLM agents.
- Proposes SPACE, which uses programmatic skills to guide chunk learning and optimize decision-making.
- Demonstrates significant improvements in task success rates and reductions in decision rounds compared to existing methods.
- Achieves strong performance with fewer training steps, indicating higher training efficiency.
Read more
Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents
Summary
This paper addresses the inefficiencies of large language model (LLM) agents in long-horizon interactive tasks, which typically operate under a ReAct-style protocol that issues one primitive action per round. The authors identify that this approach leads to excessive decision-making rounds and suboptimal performance due to the inability to learn meaningful action chunk boundaries. To overcome this, they propose SPACE (Skill-guided Policy with Adaptive Chunk Execution), a novel framework that distills chunk-boundary supervision from successful trajectories using programmatic skills. SPACE induces two-level programmatic skills, where subskill boundaries provide direct supervision for chunk boundaries. The training involves hybrid on-/off-policy optimization with chunk-aware credit assignment, allowing the policy to generate variable-length action chunks directly. Experimental results on ALFWorld and ScienceWorld demonstrate that SPACE significantly improves task success rates by 7.0%–31.3% while reducing LLM decision rounds by up to 78.9%, achieving strong performance with fewer training steps. The findings highlight the importance of variable-length action chunking for enhancing the efficiency and capability of long-horizon LLM agents.
Methodology
The methodology involves the development of SPACE, which distills chunk-boundary supervision from successful trajectories by segmenting them into programmatic skills. The training alternates between primitive-chunk rollouts and skill-augmented rollouts, converting subskill boundaries into training signals. The policy is optimized using hybrid on-policy and off-policy techniques with chunk-aware credit assignment, allowing for the generation of variable-length action chunks without reliance on a skill library during deployment.
Results
SPACE outperforms strong prompting and reinforcement learning baselines in both success rates and efficiency metrics. Specifically, it improves task success rates by 7.0%–31.3% and reduces average LLM decision rounds by up to 78.9%. Additionally, SPACE achieves the performance of the strongest baseline using only 26.6% of the training steps, indicating a substantial improvement in training efficiency.
Implications
The findings suggest that incorporating variable-length action chunking can significantly enhance the performance and efficiency of LLM agents in long-horizon tasks. This approach could be applied in various interactive applications, such as robotics, autonomous systems, and complex decision-making environments, where efficient action execution is critical.
FlashKAN: B-Spline KANs via Truncated Power Form
Efficient ML
Theory
Interpretability
- Introduces a non-recursive method for evaluating B-spline activations in KANs.
- Achieves significant speed improvements by fusing operations into a single GPU kernel.
- Implements a stabilization technique to prevent numerical issues in B-spline evaluations.
- Provides an open-source package for easy integration into existing frameworks.
Read more
FlashKAN: B-Spline KANs via Truncated Power Form
Summary
This paper introduces FlashKAN, a novel approach to Kolmogorov-Arnold Networks (KANs) that utilizes B-spline activations on network edges. Traditional methods rely on the Cox-de Boor recursion for evaluating B-spline activations, which is computationally expensive and accounts for over 90% of the forward-pass time in KAN layers. FlashKAN replaces this recursion with a truncated power form, significantly improving efficiency. The paper presents three main contributions: (1) a torch.compile-fused implementation that consolidates operations into a single GPU kernel, eliminating recursion and data-dependent memory access; (2) a bounded-coordinate stabilization technique that prevents catastrophic cancellation by clamping inputs to a specified range; and (3) an open-source package that serves as a drop-in replacement for existing KAN layers. These advancements not only enhance computational efficiency but also maintain the interpretability of KANs, making them more practical for deployment in machine learning applications.
Methodology
The methodology involves replacing the traditional Cox-de Boor recursion with a truncated power form for cubic B-splines, which allows for a more efficient evaluation. The implementation leverages torch.compile to fuse operations into a single GPU kernel, thus optimizing performance. Additionally, a bounded-coordinate stabilization technique is introduced to mitigate numerical cancellation issues.
Results
The proposed FlashKAN approach demonstrates a drastic reduction in the computational cost associated with B-spline evaluations in KANs, achieving a forward-pass time that is significantly lower than traditional methods. The implementation is validated through profiling, showing that basis computation time is minimized, leading to overall faster KAN layer performance.
Implications
FlashKAN has the potential to enhance the efficiency of KANs in various machine learning applications, particularly in scenarios where interpretability and computational speed are critical. The open-source nature of the package encourages widespread adoption and further research in spline-based neural networks.
Source-Free Class Relearning: Diagnosing Forgetting in Class Unlearning
Theory
Computer Vision
Efficient ML
- Introduces a source-free method for diagnosing class unlearning recoverability.
- Establishes a theoretical alignment condition for class relearning.
- Proposes the Relearning Score (RS) to measure recoverability and retain accuracy.
- Demonstrates significant recoverability in state-of-the-art unlearning methods across multiple datasets.
Read more
Source-Free Class Relearning: Diagnosing Forgetting in Class Unlearning
Summary
This paper addresses the challenge of class unlearning, which aims to remove a model's ability to recognize specific classes while maintaining performance on others. The authors highlight that low forget accuracy does not guarantee the complete erasure of class structure, as some unlearning methods may only alter decision boundaries without fully eliminating the underlying representation. The study introduces a source-free approach to class relearning, where the authors investigate whether a forget class can be recovered using only the released unlearned model. They establish a theoretical framework that identifies a sufficient alignment condition for enhancing the expected logit margin of the forget class through a single gradient step on synthetic probe sets. The proposed Source-Free Relearning Audit (SFRA) generates candidate embeddings and employs model-guided confidence filtering to create high-confidence retain probes and low-confidence boundary-adjacent probes, which are then relabeled as the forget class. The authors introduce the Relearning Score (RS) to quantify recoverability, measuring both forget-class recovery and retain-accuracy preservation. Experiments conducted on CIFAR-10, CIFAR-100, and TinyImageNet with various architectures demonstrate that several state-of-the-art unlearning methods exhibit significant source-free recoverability, with some methods outperforming matched retrained references. The findings suggest that SFRA serves as a practical diagnostic tool for assessing post-unlearning recoverability without requiring access to original training samples.
Methodology
The authors develop a theoretical framework for class relearning based on alignment conditions and propose the Source-Free Relearning Audit (SFRA) to generate synthetic probes in the representation space. They utilize model-guided confidence filtering to create probes for testing recoverability and introduce the Relearning Score (RS) to quantify the effectiveness of the unlearning process.
Results
Experiments reveal that several unlearning methods exhibit substantial source-free recoverability, with some methods achieving higher recoverability than matched retrained references. The Relearning Score (RS) effectively captures the balance between forget-class recovery and retain-class performance.
Implications
The findings have significant implications for privacy-preserving machine learning, as they suggest that certain unlearning methods may leave recoverable structures that could be exploited. The proposed SFRA can help practitioners assess the effectiveness of unlearning methods in various applications, including data privacy and model updates.
Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts
Federated Learning
Multimodal
- First systematic benchmark of federated LoRA-based PEFT for biomedical vision-language models across multiple cohorts.
- Federated LoRA adaptation significantly improves shared-class AUC from 0.687 to 0.802.
- SVD-based aggregation is crucial for effective model updates, outperforming naive averaging.
- Federation enhances performance in weaker cohorts while preserving strengths in stronger datasets.
Read more
Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts
Summary
This paper presents a systematic benchmark of federated learning (FL) combined with Low-Rank Adaptation (LoRA) for adapting the BiomedCLIP model to chest radiograph classification across four international cohorts. The study addresses the challenges posed by privacy regulations and data heterogeneity in biomedical imaging, allowing institutions to collaboratively train a shared model without exchanging raw data. The authors demonstrate that federated LoRA adaptation significantly improves the area under the curve (AUC) for shared-class classification across diverse datasets from the USA, Vietnam, and Spain, achieving a mean AUC of 0.802 compared to 0.687 for the unadapted model. The paper emphasizes the importance of singular value decomposition (SVD)-based aggregation for effective model updates, revealing that naive averaging of low-rank updates can lead to substantial performance drops. The findings suggest that federated adaptation can enhance model performance in weaker cohorts while maintaining the strengths of more robust datasets, thus paving the way for collaborative biomedical image analysis without compromising patient privacy.
Methodology
The study employs federated parameter-efficient fine-tuning (PEFT) using Low-Rank Adaptation (LoRA) on the BiomedCLIP model. It involves three experimental setups: single-client baselines, one-shot aggregation of locally trained adapters, and multi-round federation where clients iteratively train and aggregate model updates. The aggregation of updates is performed using singular value decomposition (SVD) to ensure accurate representation of learned weights.
Results
The results indicate that federated LoRA adaptation leads to a mean shared-class AUC of 0.802 across four cohorts, significantly improving upon the unadapted BiomedCLIP model's AUC of 0.687. The study also finds that naive factor averaging results in a drop of 0.097 in mean AUC, highlighting the effectiveness of SVD-based aggregation. Additionally, the federated approach improves performance in weaker cohorts while maintaining the performance of stronger cohorts, approaching a centralized reference AUC of 0.812.
Implications
The findings suggest that federated learning combined with LoRA can facilitate the development of robust biomedical models without compromising patient privacy. This approach can be applied to various medical imaging tasks, enabling institutions to collaborate effectively while adhering to data protection regulations. It opens avenues for further research into federated adaptation techniques in diverse medical contexts.
CACTUS: Mask-Guided Semantic Clean-Label Backdoors in Decentralized Federated Learning
Federated Learning
Audio & Speech
NLP
- CACTUS utilizes semantic triggers for stealthy backdoor attacks in decentralized federated learning.
- The method constructs label-consistent semantic pairs to create target-directed representation shifts.
- Experiments show CACTUS achieves a mean attack success rate of 51.2% across various modalities with 30% malicious nodes.
- The attack's effectiveness varies with network topology and increases with the ratio of malicious nodes.
Read more
CACTUS: Mask-Guided Semantic Clean-Label Backdoors in Decentralized Federated Learning
Summary
The paper introduces CACTUS, a novel approach to implementing clean-label backdoor attacks in decentralized federated learning (DFL). Unlike traditional methods that utilize conspicuous synthetic triggers, CACTUS employs semantic triggers that are less detectable and can be strategically placed to ensure effective backdoor propagation across multiple aggregation rounds. The methodology involves constructing label-consistent semantic pairs of clean and triggered samples, which are then used to create target-directed representation shifts. These shifts are applied counterfactually to clean embeddings before local model aggregation, allowing malicious nodes to optimize their models while maintaining a degree of clean accuracy. The authors conduct extensive experiments across various modalities, including speech, text, tabular data, and images, demonstrating the effectiveness of CACTUS in achieving high attack success rates even in the presence of benign nodes. The results indicate that the attack's effectiveness is influenced by network topology and the proportion of malicious nodes, showcasing CACTUS's ability to propagate backdoors through repeated DFL aggregation.
Methodology
CACTUS constructs semantic pairs of clean and triggered samples offline, using modality-specific operators to isolate and apply target-directed shifts to clean embeddings. Malicious nodes optimize their models based on these shifts while submitting standard model states during local aggregation.
Results
With 30% of nodes being malicious, CACTUS achieves mean attack success rates of 51.2% on Speech Commands, 71.1% on Text, 25.7% on Tabular data, and 20.3% on CelebA. It surpasses the strongest evaluated baseline on Speech Commands, Tabular, and CelebA by significant margins.
Implications
The findings suggest that CACTUS could be used to enhance the stealth and effectiveness of backdoor attacks in decentralized federated learning environments, raising concerns about the security and robustness of federated learning systems.