AI-generated summaries
Today's ML research,
without the noise.
Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.
70
Papers today
8h
Update frequency
7
Days of history
The Role of Feed-Forward Layers in Transformer Dynamics
Theory
NLP
Large Language Models
- The feed-forward layer can steer tokens towards consensus in transformers, regardless of the attention matrices.
- Theoretical results extend to multi-cluster convergence and multi-head attention scenarios.
- Numerical experiments confirm the theoretical predictions and reveal richer dynamics in real-world LLMs.
- The study connects transformer dynamics to control theory, providing a framework for future analysis.
Read more
The Role of Feed-Forward Layers in Transformer Dynamics
Summary
This paper investigates the dynamics of tokens in transformer architectures from a control-theoretic viewpoint, focusing on the role of feed-forward layers following the self-attention mechanism. The authors model the transformer as an interacting particle system, where the self-attention is treated as a coupled ordinary differential equation (ODE) system. The main theoretical contribution is the establishment that the feed-forward network can guide tokens towards consensus, independent of the key, query, and value matrices. The results extend to scenarios involving multiple clusters and multi-head attention. Numerical experiments validate the theoretical findings, comparing the predicted thresholding behavior with real-world large language models (LLMs). The study highlights the complex interplay between self-attention and feed-forward layers, suggesting that the dynamics in actual transformer systems are richer than previously understood. The paper also outlines how the findings can be generalized to multi-cluster convergence, emphasizing the importance of the feed-forward layer in controlling long-term dynamics.
Methodology
The authors employ a control-theoretic approach, modeling the transformer as an interacting particle system governed by a coupled ODE. They analyze the dynamics of the feed-forward layer in conjunction with self-attention, using theoretical results to quantify the impact of the feed-forward layer on long-term behavior. Numerical experiments are conducted to compare theoretical predictions with actual transformer behavior.
Results
The main result demonstrates that for any desired level of consensus, there exist configurations of the feed-forward layer that can achieve this, even with a single hidden neuron. The study identifies a threshold related to the operator norm of the value matrix, above which the feed-forward layer can dominate the dynamics and force tokens into a small neighborhood around a target state. The findings also indicate that real transformers exhibit a variety of dynamics, with some layers exceeding the threshold and others falling below it.
Implications
This research provides insights into the design and analysis of transformer architectures, suggesting that careful tuning of feed-forward layers can significantly influence model behavior. The findings may inform the development of more effective transformer models and enhance understanding of their operational mechanisms, particularly in applications involving large language models.
RMB: Reward Model Boosting Mitigates Reward Hacking
Reinforcement Learning
Large Language Models
NLP
- RMB addresses the reward hacking issue in RLHF by enhancing the robustness of reward signals.
- The approach involves training multiple diverse reward models and aggregating their outputs using boosting techniques.
- Extensive experiments show significant improvements in reward prediction accuracy and mitigation of reward hacking.
- The use of a diversity-promoting regularizer helps ensure that reward models capture complementary aspects of the reward landscape.
Read more
RMB: Reward Model Boosting Mitigates Reward Hacking
Summary
This paper addresses the challenge of reward hacking in Reinforcement Learning from Human Feedback (RLHF), where the optimization of a proxy reward model can lead to degraded performance in aligning large language models (LLMs) with true human preferences. The authors propose a novel approach called Reward Model Boosting (RMB), which enhances the robustness and reliability of the reward signal in RLHF. RMB involves training a diverse set of reward models using a diversity-promoting regularizer, which encourages each model to capture different aspects of the reward landscape. A lightweight aggregator is then learned using boosting techniques to combine the outputs of these diverse models into a more accurate and robust reward signal. The paper presents extensive experiments demonstrating that RMB significantly improves reward prediction accuracy on both in-distribution and out-of-distribution datasets, effectively mitigating the reward hacking issue and enhancing overall RLHF performance. The findings suggest that promoting diversity among reward models and employing a boosting strategy can lead to more reliable reward signals in RLHF applications.
Methodology
The authors propose the Reward Model Boosting (RMB) framework, which consists of training multiple reward models with a diversity-promoting regularizer to reduce correlation among them. A decision tree-based aggregator is then trained using boosting principles to combine the outputs of these models, effectively minimizing prediction errors and enhancing robustness against noise.
Results
The experiments conducted show that RMB significantly improves reward prediction accuracy on both in-distribution and out-of-distribution datasets. It effectively mitigates the reward hacking issue during RLHF training, leading to better alignment of LLMs with human preferences. The analysis confirms that the diversity-promoting regularizer is crucial for the performance gains observed.
Implications
The findings of this paper suggest that employing diverse reward models and advanced aggregation techniques can lead to more reliable and effective RLHF systems. This has potential applications in improving the alignment of LLMs with human preferences, which is critical for their deployment in real-world scenarios.
Benchmarking graph-based models for in-silico toxicity prediction in drug discovery
Graph Learning
- Introduction of a unified benchmarking framework for GDL-based toxicity prediction.
- Systematic evaluation of over 20 models under consistent experimental conditions.
- Fair and reproducible comparisons across multiple datasets and partitioning strategies.
- Analysis of methodological trends and performance claims in existing literature.
Read more
Benchmarking graph-based models for in-silico toxicity prediction in drug discovery
Summary
This paper addresses the critical challenge of predicting chemical toxicity in drug discovery, a process often hindered by high costs and risks associated with late-stage failures. The authors introduce a unified benchmarking framework for graph deep learning (GDL) models, which utilize molecular graph representations to enhance toxicity predictions. They systematically evaluate over 20 GDL approaches under consistent experimental conditions, providing a fair and reproducible comparison across multiple datasets. The study also includes a structured literature analysis to contextualize existing methodologies and performance claims. The results highlight the strengths and limitations of current graph-based approaches, offering insights into the state of the field and suggesting future research directions. To promote transparency and reproducibility, the authors release their benchmarking framework as open-source software, enabling the community to conduct further evaluations and comparisons.
Methodology
The authors developed a standardized benchmarking framework that evaluates various GDL models for toxicity prediction. They conducted systematic experiments across multiple datasets, ensuring consistent preprocessing and evaluation protocols to facilitate fair comparisons. Additionally, a structured literature analysis was performed to contextualize the findings.
Results
The benchmarking revealed a clearer assessment of the performance of different GDL models, highlighting both their strengths in predictive accuracy and their limitations in generalization across diverse chemical spaces. The study provided insights into methodological trends and established a foundation for future research in computational toxicology.
Implications
The findings from this work have significant implications for drug discovery, particularly in enhancing the early identification of toxic compounds, thereby reducing costs and improving safety in pharmaceutical development. The open-source framework encourages further research and collaboration in the field of computational toxicology.
Preferent Compression Bounds Are Tight
Theory
Optimization
- The paper confirms that the state-of-the-art preferent compression bounds are tight.
- An explicit construction using uniform distribution and order statistics is provided to demonstrate tightness.
- A simpler proof of the upper bound is presented, making the results more accessible.
- The findings have broad applications in risk certification across multiple domains, including machine learning and control systems.
Read more
Preferent Compression Bounds Are Tight
Summary
This paper addresses a critical issue in the deployment of learning-based methods: the lack of rigorous safety and performance certificates. The authors focus on sample compression, which has emerged as a powerful tool for deriving such certificates, particularly for algorithms that exhibit a preference property (or stability). The paper resolves the open question of whether the state-of-the-art bounds for preferent compressions are tight. The authors demonstrate that the upper bound is indeed tight by providing an explicit construction based on the uniform distribution and order statistics. They also present a significantly shorter and more accessible proof of this bound, which relies on elementary counting arguments rather than complex infinite-dimensional duality. The results have implications across various domains, including control systems, machine learning, and verification, where risk certificates are essential for ensuring safety and performance.
Methodology
The authors utilize a combination of theoretical analysis and explicit construction to demonstrate the tightness of the preferent compression bounds. They provide a simplified proof that avoids complex formulations, relying instead on elementary counting arguments. The construction involves using a uniform distribution and order statistics to achieve the bound in the limit.
Results
The main result of the paper is the affirmation that the upper bound for preferent compression is tight and cannot be improved. The authors present a new, simpler proof and an explicit construction that achieves this bound, thereby resolving a significant open question in the field.
Implications
The findings of this paper have significant implications for the deployment of learning-based methods in safety-critical applications. By establishing tight bounds for preferent compressions, the authors provide a foundation for developing rigorous risk certificates, which can enhance the reliability and safety of machine learning algorithms in various domains, including autonomous systems and medical diagnostics.
Learning in the Transverse Subspace: A Minimal Representation for Divergence-Free Operator Learning
Theory
- Introduces a minimal representation for divergence-free vector fields, reducing dimensionality from D to D-1.
- The method ensures that the learned outputs are inherently divergence-free, eliminating the need for post-processing projections.
- Utilizes Fourier extension for nonperiodic flows to maintain compatibility with divergence-free constraints.
- Demonstrates improved performance in terms of error reduction and robustness in operator learning tasks.
Read more
Learning in the Transverse Subspace: A Minimal Representation for Divergence-Free Operator Learning
Summary
This paper addresses the challenge of learning divergence-free vector fields, which are crucial in incompressible flows and various PDE systems. The author critiques existing methods that utilize redundant representations, such as Neural Conservation Law (NCL) potentials, which can lower fitting errors but do not provide unique mappings necessary for operator learning. The proposed approach introduces a minimal representation that encodes a D-component divergence-free vector field as a (D-1)-component vector field, allowing for a more efficient and robust learning process. For periodic and closed impermeable fields, the transformation is invertible and preserves angles, while for open nonperiodic flows, a Fourier extension is employed to create a compatible periodic field. The model learns the evolution of the reduced components directly, ensuring that the output remains divergence-free without the need for post-prediction projections. Experimental results demonstrate that this end-to-end formulation achieves lower errors and greater robustness compared to traditional projection-based methods, which often introduce instability due to redundant components.
Methodology
The paper employs a fixed analytic Fourier-Householder transformation to map D-component divergence-free fields to (D-1)-component fields. The model encodes training data in this reduced representation and learns the evolution directly in the transverse components. For nonperiodic flows, a Fourier extension is used to create a periodic representation, and a minimum-energy rule selects the reduced representation.
Results
The proposed method achieves lower representation errors and greater robustness in learning divergence-free vector fields compared to existing projection-based methods. The experiments validate the effectiveness of the minimal representation in maintaining divergence-free properties throughout the learning process.
Implications
This approach has significant implications for fluid dynamics, electromagnetism, and other fields where divergence-free constraints are essential. It offers a more efficient framework for operator learning, potentially leading to advancements in simulations and predictive modeling of physical systems.
GNA: Granular Neighbor Assembly for Retrieval-Augmented Multivariate Time-Series Forecasting
Time Series
- GNA combines whole-window and per-variate retrieval in a single multivariate forecasting model.
- The model uses a learned gate to weigh the trust in retrieved futures against persistence forecasts.
- GNA shows significant performance improvements across multiple datasets and benchmarks.
- Retrieval is most effective when the lookback window is less informative, allowing the model to adaptively shift trust.
Read more
GNA: Granular Neighbor Assembly for Retrieval-Augmented Multivariate Time-Series Forecasting
Summary
This paper introduces GNA (Granular Neighbor Assembly), a novel retrieval layer designed for multivariate time-series forecasting. Traditional deep forecasting models rely on fixed-length lookback windows, which can lead to diminishing returns as the window length increases. GNA addresses this limitation by employing a retrieval-augmented approach that allows the model to access similar past situations and their continuations. It assembles neighbors at two granularities: whole past windows, which maintain coherence across variates, and per-variate neighbors, where each variate retrieves its most relevant past. A learned gate dynamically balances the trust between these retrieved futures and the model's own predictions. The method is strictly causal, ensuring that past windows are only used once their futures are observed. GNA significantly enhances the performance of two Transformer backbones, achieving improvements across numerous dataset-horizon settings and demonstrating the importance of both granularities in retrieval. The findings suggest that retrieval is particularly beneficial when the lookback window provides limited information, with the model effectively shifting trust to retrieved futures as the forecast horizon extends.
Methodology
GNA employs a retrieval layer that assembles neighbors at two levels of granularity: global slots for coherent past windows and per-variate slots for individualized variate retrieval. A learned gate determines the trust level in retrieved futures versus persistence forecasts, allowing for dynamic adjustments based on forecast steps and variates. The retrieval process is strictly causal, with candidates sourced from an embedding trained to predict future values based on past windows.
Results
GNA improved the performance of the GridTST and iTransformer models in 85 out of 96 dataset-horizon settings, achieving the lowest mean squared error (MSE) on 8 out of 12 standard benchmarks. The model consistently outperformed its backbone in all seeds across 10 of the 12 benchmarks, demonstrating the effectiveness of its retrieval strategy.
Implications
The findings suggest that GNA can significantly enhance forecasting accuracy in multivariate time-series applications, particularly in scenarios where traditional fixed-length lookback windows are inadequate. This approach could be applied in various fields such as finance, weather forecasting, and anomaly detection, where timely and accurate predictions are critical.
AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs
Large Language Models
- AgentPerfBench captures diverse agentic workloads, including coding and tool-using agents, through multi-turn interactions.
- The benchmarking suite employs saturation-based measurements to reflect true hardware performance under maximum load.
- Kernel-level profiling reveals performance bottlenecks and provides insights into memory bandwidth and capacity limitations.
- The study demonstrates that existing benchmarks fail to accurately represent the demands of agentic LLM applications.
Read more
AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs
Summary
The paper introduces AgentPerfBench, a new benchmarking suite designed to evaluate the inference performance of agentic large language models (LLMs). Current benchmarks primarily focus on simple, single-turn interactions, which do not reflect the increasing complexity of real-world applications involving multi-turn requests and tool usage. AgentPerfBench addresses this gap by utilizing real workload traces from existing benchmarks and generating synthetic profiles that represent diverse agentic workloads. The suite incorporates saturation-based measurement techniques to accurately assess hardware performance under maximum load conditions, revealing that existing benchmarks often underreport hardware capabilities. Additionally, the authors provide kernel-level profiling using Nsight Compute to analyze performance bottlenecks through a multi-dimensional roofline model. The results indicate significant performance discrepancies between traditional chat-based workloads and agentic workloads, highlighting the need for more representative benchmarking in the evolving landscape of LLM applications.
Methodology
AgentPerfBench organizes workloads into profiles based on real agentic traces, sampling input and output lengths to create representative synthetic requests. It employs saturation-based evaluation techniques to measure performance under maximum concurrent requests and utilizes Nsight Compute for kernel-level profiling, mapping workloads onto a multi-dimensional roofline model to identify hardware limitations.
Results
The benchmarking results show that switching from chat to coding-agent workloads significantly increases time-to-first-token (TTFT) and time-per-output-token (TPOT), with a 4.8ร increase in TTFT and a 1.5ร increase in TPOT for LLaMA-3.1-70B on H100. Multi-turn sessions exacerbate these performance gaps, with TTFT for SWE-Bench agents reaching 433 ms compared to 62 ms for chat interactions.
Implications
AgentPerfBench provides a more accurate framework for evaluating the performance of LLMs in real-world applications, guiding future optimizations in AI hardware and serving engines. The insights gained from this benchmarking suite can inform the development of more efficient LLM architectures and improve the deployment of agentic applications in various domains.
Identifying ODEs from Unstructured Data with Causal Representation Learning
Theory
Time Series
Computer Vision
- Introduces SPEED-AE, a framework combining CRL and autoencoders for ODE discovery.
- Demonstrates improved identifiability of variables from polynomial to monomial diffeomorphisms.
- Achieves state-of-the-art performance in recovering ODEs from unstructured data.
- Provides theoretical guarantees for the identification of true variables from high-dimensional observations.
Read more
Identifying ODEs from Unstructured Data with Causal Representation Learning
Summary
This paper addresses the challenge of recovering governing Ordinary Differential Equations (ODEs) from unstructured, high-dimensional data, such as images, where direct measurements of variables are not available. Traditional methods for ODE discovery often rely on direct observations or lack theoretical guarantees for the learned variables and equations. The authors propose a novel framework called SParse Equivalent Equation Discovery AutoEncoder (SPEED-AE), which integrates a pretrained Causal Representation Learning (CRL) model with a component-wise autoencoder. This combination allows for the transformation of variables into a form suitable for sparse ODE discovery. The paper demonstrates that by applying this method, the identifiability of variables can be improved from polynomial to monomial diffeomorphisms, enhancing the recovery of ODEs that closely match the ground truth. Experimental results on systems such as Lotka-Volterra, Lorenz, and a two-pendulum system indicate that SPEED-AE significantly improves the disentanglement of variables and achieves state-of-the-art forecasting performance, thereby providing a robust approach to ODE discovery from complex data.
Methodology
The authors develop SPEED-AE by first employing a pretrained CRL model to extract low-dimensional latent variables from high-dimensional data. Following this, a component-wise autoencoder is utilized to learn transformations of these variables, making them suitable for sparse ODE discovery. The framework imposes sparsity in the representation of the ODE, allowing for the recovery of equations that are close to the true underlying dynamics.
Results
The experiments conducted on various dynamical systems, including Lotka-Volterra, Lorenz, and a two-pendulum system, show that SPEED-AE not only improves the disentanglement of variables but also recovers ODEs that are closer to the ground truth compared to existing methods. Furthermore, it achieves superior forecasting performance, establishing it as a leading approach in the field.
Implications
The findings suggest that SPEED-AE can be a valuable tool for scientific discovery in dynamical systems, particularly in scenarios where only high-dimensional observations are available. This approach could be applied in various fields, including physics, biology, and engineering, where understanding the underlying dynamics of complex systems is crucial.
ConRAG: Lightweight inference of multi-hop relations
NLP
Large Language Models
Graph Learning
- Introduces CONRAG, a graph-based framework for multi-hop relation inference.
- Constructs a lightweight entity-document graph for efficient path retrieval.
- Outperforms existing RAG baselines in bridge entity recovery and reasoning chain precision.
- Reduces indexing token costs significantly, enhancing scalability.
Read more
ConRAG: Lightweight inference of multi-hop relations
Summary
The paper introduces CONRAG, a novel framework for multi-hop relation inference that addresses the challenge of connecting two known entities through intermediate entities and evidence across a document corpus. Traditional multi-hop retrieval systems often focus on finding an unknown answer entity, which limits their effectiveness in explicitly reconstructing connections between known endpoints. CONRAG constructs a lightweight entity-document graph using entity co-occurrence and LLM-based filtering, allowing for efficient retrieval of paths between entities. The authors formalize the task of multi-hop relation inference and provide new benchmarks derived from existing datasets, MuSiQue and 2WikiMultiHopQA, to evaluate their approach. The results demonstrate that CONRAG significantly outperforms existing RAG baselines in recovering bridge entities and reasoning chains while achieving a reduction in graph-indexing token costs by up to 1.5 orders of magnitude. This work not only advances the state of multi-hop relation inference but also sets a foundation for future research in relation discovery.
Methodology
CONRAG builds an entity-document graph from a corpus using entity co-occurrence and LLM-based filtering. It retrieves paths between two known entities by scoring and ranking these paths based on semantic relevance, allowing for the identification of intermediate entities and evidence needed to explain their connection.
Results
CONRAG consistently outperformed strong RAG baselines on the MuSiQue and 2WikiMultiHopQA benchmarks, achieving better bridge entity recovery and reasoning chain recall and precision. Additionally, it demonstrated a reduction in graph-indexing token costs by approximately 1.5 orders of magnitude.
Implications
The findings suggest that CONRAG can enhance knowledge discovery in various domains, including scientific research and question answering, by providing a more efficient and effective means of tracing multi-hop relations. This could lead to improved systems for evidence-based reasoning and decision-making.
Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling
NLP
Large Language Models
Efficient ML
- Triadic Linear Attention increases the memory state size of RNNs using a triadic outer product, enhancing recall for long-context tasks.
- The method is parameter-efficient, allowing for significant state size increases with minimal additional projections.
- Compatible with modern innovations in linear attention, such as data-dependent forgetting and the delta rule.
- Demonstrated improvements in long-context language modeling and recall capabilities over traditional and alternative approaches.
Read more
Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling
Summary
This paper introduces Triadic Linear Attention, a novel approach to enhance the memory state size of recurrent neural networks (RNNs) for long-context sequence modeling. Traditional RNNs utilize fixed-size memory states, which limits their ability to recall information from longer sequences. The authors build on the concept of linear attention, which extends vector-valued hidden states to matrix-valued states using an outer product of key and value vectors. Triadic Linear Attention generalizes this by employing a triadic outer product of two keys and one value, resulting in a third-order tensor state. This method allows for a significant increase in state size without a proportional increase in parameters, as it only requires two additional projections. The approach is compatible with innovations such as data-dependent forgetting, the delta rule, and chunkwise-parallel training, making it efficient for large-scale applications. The authors demonstrate that applying Triadic Linear Attention to models like Gated DeltaNet and scalar-gated linear attention leads to substantial improvements in long-context language modeling and recall capabilities, outperforming existing methods that increase state size through more complex means.
Methodology
The authors propose a triadic outer product of two keys and one value to create a third-order tensor state. This state is updated and read using queries that contract both key axes, allowing for efficient memory management and retrieval. The approach integrates with existing techniques like data-dependent forgetting and the delta rule, and is designed for efficient training through chunkwise-parallel methods.
Results
The application of Triadic Linear Attention to Gated DeltaNet and scalar-gated linear attention resulted in significant improvements in long-context language modeling and recall. The method outperformed alternatives that increase state size through larger heads or multiple value heads per key, achieving better performance with lower memory usage.
Implications
Triadic Linear Attention offers a new framework for enhancing RNNs in tasks requiring long-context understanding, such as language modeling and sequence prediction. Its parameter-efficient design makes it suitable for large-scale applications in natural language processing and other domains where memory efficiency is critical.
Inducing Process Supervision from Outcome-Only Reinforcement Learning
Reinforcement Learning
Large Language Models
Theory
- Introduction of TIPS, an outcome-only RL framework for training PRMs.
- TIPS reinforces step-level verification through outcome prediction without direct supervision.
- Achieves state-of-the-art performance on ProcessBench with significantly fewer labeled trajectories.
- Provides a theoretical analysis supporting the relationship between outcome verification and step correctness.
Read more
Inducing Process Supervision from Outcome-Only Reinforcement Learning
Summary
The paper introduces TIPS (Thinking-Induced Process Supervision), a novel outcome-only reinforcement learning framework aimed at training process reward models (PRMs) for large language models (LLMs). PRMs provide crucial step-level feedback for reasoning and agent trajectories, but traditional training methods are costly and inefficient. TIPS addresses this by allowing models to generate a chain-of-thought (CoT) followed by step-level and outcome labels, where the reward is based solely on the correctness of the predicted outcome. This approach enables the model to reinforce step-level verification without explicit process supervision. The authors validate TIPS across various benchmarks, demonstrating its effectiveness in enhancing PRMs with minimal outcome-labeled data. Notably, TIPS-Qwen3-4B-Thinking-2507 achieves an F1 score of 85.2 on ProcessBench using only 3.2K outcome-labeled trajectories, outperforming larger trained PRMs and strong prompt-only judges.
Methodology
TIPS employs an outcome-only reinforcement learning approach where the model generates a chain-of-thought followed by step-level labels and an outcome label. The reward mechanism is based solely on the correctness of the outcome prediction, which indirectly enhances the accuracy of step-level verification. The framework utilizes group-relative advantage to optimize the entire generated response, allowing for effective training without explicit process supervision.
Results
TIPS-Qwen3-4B-Thinking-2507 achieved an F1 score of 85.2 on the ProcessBench benchmark using only 3.2K outcome-labeled trajectories. This performance surpassed all evaluated trained PRMs and strong prompt-only judges, demonstrating the framework's efficiency and effectiveness in training PRMs.
Implications
The findings suggest that TIPS can significantly reduce the cost and complexity of training process reward models, making it a valuable approach for enhancing reasoning capabilities in large language models. This could lead to more efficient applications in various domains requiring step-level reasoning and verification.
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
NLP
Large Language Models
Reinforcement Learning
- Identification of a low-dimensional effective manifold in activation space related to RL-induced performance gains.
- Discovery of two geometric properties: Effective Manifold Capacity and Control Manifold Separation.
- Introduction of Alpha-Stabler, a framework that stabilizes RL training and enhances performance.
- Experiments demonstrate that a single input-invariant vector can recover a significant portion of RL gains.
Read more
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
Summary
This paper investigates the role of reinforcement learning with verifiable rewards (RLVR) in enhancing the reasoning capabilities of large language models (LLMs). The authors introduce vector steering as a method to analyze the high-dimensional parameter updates in RL training, revealing a low-dimensional effective manifold in activation space that correlates with performance gains. They identify two key geometric properties of this manifold: Effective Manifold Capacity, which indicates that while a small capacity can reproduce RL gains, it is not infinitely compressible, and Control Manifold Separation, where effective control directions are found in the low-variance complement of the activation principal subspace. The study includes experiments on five LLMs across six tasks, confirming these properties. Based on their findings, the authors propose Alpha-Stabler, a training framework that stabilizes RL training by monitoring principal-subspace intrusion and adjusting activation gradients. This framework shows promise in enhancing training stability and performance gains in LLMs.
Methodology
The authors employed vector steering to analyze the activation space of LLMs during RL training. They conducted experiments with five LLMs and six tasks, using input-invariant shared vectors to distill RL-tuned models and assess the geometric properties of the effective manifold. The Alpha-Stabler framework was developed to monitor and adjust training dynamics based on the identified geometric properties.
Results
The experiments revealed that a single input-invariant vector could recover over 85% of the performance gains achieved through full fine-tuning in most tasks. The study confirmed the existence of the effective manifold and its geometric properties, supporting the hypothesis that RL-induced changes have a structured, low-dimensional representation in activation space. Alpha-Stabler was shown to stabilize training for 2,000 steps while enhancing RL gains.
Implications
The findings provide insights into the internal mechanisms of RL in LLMs, suggesting that understanding the geometry of activation space can lead to more robust training methods. The Alpha-Stabler framework could be applied to improve the stability and performance of various RL applications in LLMs.
Learning the Structure of Triangular Transport Maps
Generative Models
Graph Learning
Optimization
- Introduces Self-Structuring Transport Maps (SSTM) for joint learning of map structure and parameters.
- Utilizes SoftSort for variable ordering and stochastic L0 gates for sparsity learning.
- Demonstrates superior density estimation performance compared to traditional methods.
- Achieves competitive results against autoregressive flows on large datasets.
Read more
Learning the Structure of Triangular Transport Maps
Summary
This paper introduces Self-Structuring Transport Maps (SSTM), a novel approach for jointly learning the structure and parameters of triangular transport maps used in probabilistic modeling. Triangular transport maps are effective for tasks such as density estimation and generative modeling, as they transform complex target distributions into simpler reference distributions through monotone triangular mappings. The structure of these maps is defined by variable ordering and sparsity patterns, which can significantly influence the quality of the resulting maps. Traditional methods often require separate optimization for structure and map fitting, which can be computationally expensive, especially in high dimensions. SSTM addresses this challenge by employing a multi-task learning framework that utilizes SoftSort for variable ordering and stochastic L0 gates for learning sparsity, while maintaining the triangular structure. The method incorporates a monotone BatchEnsemble to ensure scalability and efficiency. Experimental results demonstrate that SSTM outperforms traditional methods by providing better density estimates and achieving competitive performance against autoregressive flows, particularly in large datasets.
Methodology
The SSTM method employs a multi-task learning framework where map components share parameters through a BatchEnsemble architecture. It uses SoftSort for variable ordering and stochastic L0 gates to learn sparsity while ensuring the triangular structure is preserved during optimization. The method incorporates a monotone BatchEnsemble to facilitate scalability by sharing weight matrices across map components.
Results
SSTM shows improved density estimates over traditional methods that first estimate structure before fitting the map. When the structure is identifiable, SSTM matches the performance of maps fitted with the true structure and outperforms autoregressive flows. On large datasets, SSTM remains competitive with autoregressive flows, demonstrating its effectiveness in practical applications.
Implications
The findings suggest that jointly learning the structure and parameters of triangular transport maps can lead to more efficient and accurate probabilistic modeling. This approach could be beneficial in various applications, including Bayesian inference, density estimation, and generative modeling, particularly in high-dimensional settings where computational efficiency is crucial.
Neural Succession: A Mesoscopic Theory of Invasion, Coexistence, and Stabilization in Continual Learning
Theory
- Introduces Successional Learning Theory (SLT) as a mesoscopic framework for continual learning.
- Demonstrates that pre-invasion compatibility predicts forgetting and coexistence outcomes effectively.
- Establishes ecological analogies for understanding task transitions in continual learning.
- Identifies key conditions for coexistence and provides a minimum habitat-modification bound.
Read more
Neural Succession: A Mesoscopic Theory of Invasion, Coexistence, and Stabilization in Continual Learning
Summary
This paper introduces Successional Learning Theory (SLT), a novel framework for understanding continual learning through the lens of ecological succession. The authors conceptualize the learned representation as a resident community and the incoming task as an invader, with various dynamics such as forgetting, coexistence, and reinforcement mapped to ecological concepts. Through eight experiments, the authors demonstrate that pre-invasion compatibility can predict forgetting and coexistence outcomes, outperforming traditional metrics like activation covariance and representation similarity. The findings reveal that compatibility effectively orders forgetting across task transitions and can forecast held-out forgetting with significantly lower error rates. The study also establishes a minimum habitat-modification bound and identifies conditions for coexistence, providing a comprehensive understanding of the dynamics involved in continual learning. The results indicate that SLT can serve as a diagnostic tool for evaluating continual learning strategies, complementing existing methods like replay and regularization.
Methodology
The authors conducted eight experiments involving the Split-CIFAR-10 dataset and controlled manipulations of MNIST and CIFAR-10 tasks. They measured directional pre-invasion compatibility to diagnose transitions and analyzed the effects of incoming tasks on resident representations. The methodology included statistical analyses to correlate compatibility with forgetting rates and coexistence outcomes.
Results
The results indicated that pre-invasion compatibility significantly ordered forgetting across task transitions, with a correlation coefficient of r = -0.789. The compatibility measure also forecasted held-out forgetting with 24% lower error than a no-information baseline. The study found that the three most compatible transitions were the only ones that coexisted, achieving an AUC of 1.00. Replay mechanisms were shown to repair transitions efficiently, particularly where displacement was largest.
Implications
The findings suggest that SLT can enhance the understanding of continual learning dynamics, providing insights into how models can better adapt to new tasks without significant forgetting. This has potential applications in developing more robust machine learning systems that can learn continuously in dynamic environments.
Variational Augmented Invertible Koopman Autoencoder for probabilistic time series forecasting
Time Series
- VAIKAE introduces a stochastic framework for time series forecasting, enhancing uncertainty quantification.
- The architecture leverages normalizing flow models for likelihood computations in dynamical systems.
- New strategies for uncertainty-aware latent data assimilation are proposed.
- Experiments show VAIKAE's effectiveness on long-term forecasting benchmarks.
Read more
Variational Augmented Invertible Koopman Autoencoder for probabilistic time series forecasting
Summary
The paper introduces the Variational Augmented Invertible Koopman Autoencoder (VAIKAE), a novel architecture designed for probabilistic time series forecasting. Traditional Neural Koopman autoencoders have demonstrated the ability to create latent embeddings with linear dynamics, which is beneficial for long-term forecasting. However, these models typically operate in a deterministic framework, limiting their capacity to quantify prediction uncertainty. VAIKAE addresses this limitation by modeling the latent embedding as a Gaussian distribution, thus incorporating stochasticity into the forecasting process. A significant feature of VAIKAE is its integration with normalizing flow models, which allows for likelihood computations in the state space of dynamical systems during training. The authors also propose new strategies for uncertainty-aware latent data assimilation using the trained VAIKAE model. The effectiveness of VAIKAE is validated through experiments on various long-term time series forecasting benchmarks, demonstrating its capability to handle uncertainty in predictions and improve forecasting accuracy.
Methodology
The authors developed the VAIKAE architecture, which employs a Gaussian distribution for the latent embedding instead of a deterministic approach. The model utilizes normalizing flows to facilitate likelihood computations, enabling the training of the model in a probabilistic setting. Additionally, the paper outlines strategies for data assimilation that incorporate uncertainty into the forecasting process.
Results
The experiments conducted on long-term time series forecasting benchmarks demonstrate that VAIKAE outperforms traditional deterministic models in terms of accuracy and uncertainty quantification. The model effectively captures both types of uncertaintyโaleatoric and epistemicโleading to improved predictive performance.
Implications
VAIKAE has significant implications for fields requiring accurate time series forecasting under uncertainty, such as finance, climate modeling, and engineering. Its ability to quantify uncertainty can enhance decision-making processes in various applications.
Depot-Closed Multi-Component Construction for Neural Vehicle Routing
Optimization
- Introduction of multi-component construction for flexible route assignment in vehicle routing.
- Depot-closed interpretation allows for effective evaluation of route components and merges.
- Neural policy trained on CVRP100 shows strong generalization to larger problem sizes.
- Outperforms existing neural solvers in various benchmarks, including zero-shot evaluations.
Read more
Depot-Closed Multi-Component Construction for Neural Vehicle Routing
Summary
This paper addresses the limitations of traditional route-by-route construction methods in neural constructive solvers for the capacitated vehicle routing problem (CVRP). The authors propose a novel approach called multi-component construction, which allows for the simultaneous maintenance of multiple route components and their arbitrary merging. This method overcomes the early commitment to route membership that restricts global coordination across routes. To facilitate this, the authors introduce a depot-closed interpretation of route components, treating each as an implicitly depot-closed route. This interpretation enables the neural policy to evaluate merges based on Clarke-Wright savings, enhancing the decision-making process for connecting components. The proposed method demonstrates significant improvements in solution quality, outperforming existing neural solvers across various problem sizes and capacities. The authors also highlight the robustness of their approach, suggesting that learned route-closing behavior contributes to overcoming the limitations of traditional route-by-route solvers.
Methodology
The authors implemented a multi-component construction approach that maintains multiple route components simultaneously. They utilized a depot-closed interpretation to treat each component as a complete route, allowing for the application of Clarke-Wright savings to evaluate merges. The neural policy was trained using supervised learning for merge connectivity and reinforced learning to optimize the merge sequence.
Results
The proposed method outperformed existing neural solvers on CVRP100-500 with greedy inference and demonstrated strong performance on larger problem sizes up to CVRP1000. In zero-shot evaluations, the model consistently outperformed reported results across various capacities, indicating robust generalization capabilities.
Implications
The findings suggest that multi-component construction can significantly enhance the efficiency and effectiveness of neural solvers for vehicle routing problems. This approach may have broader applications in combinatorial optimization tasks where global coordination across components is essential.
BERT4DTI : BERT-based Model for Predicting Drug-Protein Interactions
NLP
Large Language Models
Multimodal
- BERT4DTI combines ChemBERTa and ProtBERT with mutual attention for DTI prediction.
- The model addresses limitations of existing DTI models, such as data scarcity and computational cost.
- BERT4DTI achieves state-of-the-art results on multiple benchmark datasets.
- The use of partial fine-tuning significantly reduces the number of trainable parameters.
Read more
BERT4DTI : BERT-based Model for Predicting Drug-Protein Interactions
Summary
The paper introduces BERT4DTI, a novel model designed to predict drug-protein interactions (DTIs) using advanced deep learning techniques. The authors highlight the challenges faced in DTI prediction, such as the scarcity of labeled interactions, the high computational cost of fine-tuning large pretrained models, and the inability of independently encoded sequences to capture pair-specific dependencies. BERT4DTI addresses these issues by utilizing ChemBERTa for encoding SMILES strings of drugs and ProtBERT for amino-acid sequences of proteins. The model employs bidirectional mutual attention to enhance interaction representation and utilizes convolutional layers and a multilayer perceptron for classification. To optimize performance while minimizing the number of trainable parameters, ProtBERT is truncated to 18 layers, and only the last two layers of each encoder are fine-tuned. The model is evaluated on three benchmark datasets: BIOSNAP, DAVIS, and BindingDB, achieving competitive results, including the best ROC-AUC and PR-AUC scores on BIOSNAP and the highest sensitivity across all datasets. An ablation study indicates that the mutual attention mechanism significantly improves PR-AUC and specificity. BERT4DTI demonstrates a favorable performance-parameter trade-off with 125M trainable parameters, compared to 353M for full BERT fine-tuning, suggesting its potential for efficient DTI screening.
Methodology
BERT4DTI employs a sequence-based architecture that integrates two pretrained encoders (ChemBERTa for drugs and ProtBERT for proteins) with a mutual attention mechanism. The model processes drug and protein sequences to produce contextual representations, which are then combined through mutual attention before being classified using convolutional layers and a multilayer perceptron. The model is partially fine-tuned to optimize performance while minimizing computational costs.
Results
BERT4DTI achieved the best ROC-AUC and PR-AUC on the BIOSNAP dataset and the highest sensitivity across all evaluated datasets (BIOSNAP, DAVIS, BindingDB). The model's architecture, particularly the mutual attention mechanism, was shown to enhance performance metrics such as PR-AUC and specificity. The model's parameter efficiency was highlighted, with only 125M trainable parameters compared to 353M for full BERT fine-tuning.
Implications
The development of BERT4DTI has significant implications for drug discovery and repurposing, enabling more efficient in silico screening of drug-protein interactions. Its ability to effectively model interactions with fewer parameters opens avenues for further research and application in bioinformatics and computational biology.
Weights Read and Write Features: Scalable Parameter Decomposition Grounded in Activation Space
Large Language Models
Interpretability
- Introduction of ASPD for joint decomposition of activation and parameter spaces.
- Grounding weight components in activation features enhances interpretability.
- ASPD enables scalable and causally editable parameter decomposition in large models.
- Demonstrated effectiveness on Qwen-3-8B, recovering known computational mechanisms.
Read more
Weights Read and Write Features: Scalable Parameter Decomposition Grounded in Activation Space
Summary
This paper introduces Activation-Supported Parameter Decomposition (ASPD), a novel method that jointly decomposes activation and parameter spaces in neural networks, particularly large language models. The authors argue that existing interpretability methods often analyze activation and parameter spaces separately, which limits understanding of how weights interact with activations. ASPD addresses this by grounding weight components in the activation features they read and write, thereby constraining parameter decompositions to be more interpretable and causally editable. The method employs a shared sparse representation to define activation features and gates the corresponding weight components, along with an internal reconstruction objective that ensures the components reproduce the transformations of the target weight matrix. The authors demonstrate ASPD's effectiveness on the Qwen-3-8B model, recovering mechanisms underlying known computations such as induction and duplicate-token detection without requiring circuit labels. This approach allows for tracing semantic transformations through model weights and constructing parameter-level mechanism circuits, enhancing the interpretability of large pretrained models.
Methodology
The authors developed ASPD, which uses a shared sparse representation to define activation features and corresponding weight components. An internal reconstruction objective is employed to ensure that the learned components accurately reproduce the transformations of the target weight matrix. This method allows for efficient computation and avoids the high costs associated with pairwise feature-component interventions.
Results
ASPD successfully decomposed the parameter space of the Qwen-3-8B model, revealing weight components that correspond to known computations like induction and duplicate-token detection. The method demonstrated the ability to trace semantic transformations through model weights and construct parameter-level mechanism circuits, providing insights into the model's internal workings.
Implications
The findings suggest that ASPD can significantly enhance the interpretability of large language models, making it easier to understand and edit their internal mechanisms. This could lead to improved model transparency and trustworthiness, as well as facilitate further research into model behavior and capabilities.
Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs
NLP
Large Language Models
Efficient ML
- Identification of stale-delta failure mechanism in FP8 attention training.
- Introduction of Delta-Matching to restore softmax gradient invariance.
- Demonstration of Delta-Matching's effectiveness across various model sizes and architectures.
- Matching of BF16/FP32 mixed-precision training performance with native FP8 training.
Read more
Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs
Summary
This paper addresses the challenges of reliable FP8 attention in native 8-bit training for large language models (LLMs). The authors identify a stale-delta failure mechanism that arises from inconsistencies in forward-backward operations, which distorts training dynamics and leads to performance degradation, particularly in larger models. They propose a novel method called Delta-Matching, which restores the softmax gradient's zero-row-sum invariant without requiring architectural changes or auxiliary outputs. The method enables effective native block-scaled FP8 training across various architectures and scales, matching the performance of BF16/FP32 mixed-precision training. The findings suggest that accumulated optimization errors can be concealed in smaller models, highlighting the importance of addressing these issues for successful 8-bit training. The authors plan to release their implementation and trained models to facilitate further research in this area.
Methodology
The authors conducted theoretical analyses and controlled experiments to trace the failure mechanism in FP8 attention training. They derived the Delta-Matching approach to correct the inconsistencies in forward and backward operations, enabling native FP8 training without architectural modifications. The methodology involved testing across different model architectures, scales, and training stages to validate the effectiveness of Delta-Matching.
Results
The results showed that Delta-Matching successfully mitigated the loss gap observed in stale-delta hybrid runs, matching the training loss and downstream performance of BF16/FP32 mixed-precision training across various tested configurations. The method demonstrated resilience against performance degradation in larger models, effectively addressing the accumulated optimization errors.
Implications
The findings have significant implications for the development of efficient training methods for large language models, particularly in enabling fully native 8-bit training. This could lead to faster and more memory-efficient training processes, making it feasible to train larger models with reduced computational resources.
Transversal Pooling Neural Networks
Computer Vision
Theory
Efficient ML
- Introduction of Transversal Pooling Neural Networks (TraPNets) for improved stability and sensitivity to transformations.
- Establishment of equivariance to affine group actions and derivation of stability bounds for pooled coefficients.
- Demonstration of TraPNets' effectiveness in low-data scenarios, particularly in tropical cyclone prediction.
- TraPNets outperform traditional CNNs with data augmentation in synthetic experiments.
Read more
Transversal Pooling Neural Networks
Summary
This paper introduces Transversal Pooling Neural Networks (TraPNets), a novel architecture designed to enhance stability to small transformations while maintaining sensitivity to larger ones in machine learning tasks. The authors generalize the concept of spatial max pooling to accommodate affine group actions, establishing equivariance to a selected subgroup and deriving explicit stability bounds for pooled wavelet coefficients under affine perturbations. The motivation stems from the need for approximate invariance in tasks such as digit recognition and predicting tropical cyclone intensification, where small transformations should not alter the output, but larger ones should. The experimental results demonstrate that TraPNets outperform traditional convolutional neural networks (CNNs) in low-data environments, particularly in the context of tropical cyclone prediction and a synthetic model inspired by it. The findings suggest that TraPNets can effectively leverage limited training data while providing robustness against small perturbations.
Methodology
The authors develop TraPNets by extending the max pooling operation to handle affine transformations, ensuring equivariance to a chosen subgroup. They derive mathematical stability bounds for the pooled wavelet coefficients under affine perturbations and conduct experiments to evaluate the performance of TraPNets against conventional CNNs, particularly in low-data settings.
Results
The experimental results indicate that TraPNets significantly improve performance in tasks with limited training data compared to CNN baselines. In particular, the networks demonstrated enhanced predictive capabilities for tropical cyclone intensification and performed well in a synthetic model designed to mimic this task.
Implications
The findings suggest that TraPNets can be effectively applied in scenarios where data is scarce, and robustness to small perturbations is crucial. This has potential applications in fields such as meteorology, image recognition, and other areas requiring stability to transformations.
Interpretable intrinsic dimension estimation through componentwise calibration of distance and angle
Interpretability
- Introduces componentwise calibration for intrinsic dimension estimation, enhancing interpretability.
- Derives a closed-form Kullback-Leibler divergence for improved robustness against noise.
- Demonstrates significant reductions in mean percentage error in ID estimation across various datasets.
- Highlights the impact of sample-amplitude heterogeneity on angular statistics.
Read more
Interpretable intrinsic dimension estimation through componentwise calibration of distance and angle
Summary
This paper presents a novel approach for estimating intrinsic dimension (ID) using a method called DANCo (Dimensionality from Angle and Norm Concentration), which calibrates both nearest-neighbor distance and angular statistics. The authors reformulate DANCo to allow for componentwise calibration, enabling the identification and interpretation of the sources of estimates. They derive a closed-form Kullback-Leibler divergence for the generalized ratios ID estimator (Gride) to improve accuracy in the presence of noise and sample-amplitude heterogeneity. The study shows that by separating distance and angular discrepancies, the influence of each can be explicitly analyzed. The authors demonstrate that their method significantly reduces mean percentage error in ID estimation, particularly under noisy conditions, and provides insights into the effects of amplitude variations on angular location. The results indicate that their componentwise approach is effective in both synthetic and real-world datasets, including CIFAR-10 and ImageNet, and highlights the importance of angular calibration in neural network representations.
Methodology
The authors reformulate the DANCo method to allow for separate calibration of distance and angular statistics. They derive a closed-form Kullback-Leibler divergence for the generalized ratios ID estimator (Gride) and implement a profiling technique to align mean direction while retaining concentration matching. The methodology includes extensive testing on synthetic and real-world datasets to validate the effectiveness of the componentwise approach.
Results
The reformulated DANCo method achieves a mean percentage error reduction from 27.7% to 17.6% under noise conditions. On a Gaussian scale mixture, profiling increases the Minimum Neighbor Distance (MiND) estimate significantly, and amplitude-reducing normalizations improve angular location estimates on CIFAR-10 and ImageNet. The results show that the componentwise analysis effectively identifies the layers most sensitive to angular calibration in pretrained convolutional neural networks.
Implications
The findings suggest that the componentwise calibration approach can enhance the interpretability of intrinsic dimension estimates, which is crucial for understanding dataset complexity and guiding dimensionality reduction techniques. This has potential applications in various fields, including computer vision and deep learning, where understanding the geometric properties of data representations is essential.
RAEGL: Risk-Aware Evidence-Gated Learning for Selective Contextual Routing under Temporal Shift
Time Series
- RAEGL introduces a deploy-or-exact-fallback approach for contextual forecasting under temporal shifts.
- The framework separates model training, candidate selection, gate calibration, and evaluation into distinct phases.
- A search-aware evidence gate ensures that contextual components are only deployed when they meet specific criteria.
- Experiments show RAEGL's effectiveness in reducing forecasting errors and managing deployment risks.
Read more
RAEGL: Risk-Aware Evidence-Gated Learning for Selective Contextual Routing under Temporal Shift
Summary
The paper introduces RAEGL, a Risk-Aware Evidence-Gated Learning framework designed to enhance forecasting accuracy by addressing the challenges posed by temporal distribution shifts. Traditional contextual forecasting methods may become unreliable when corrections selected from historical data do not generalize well to future conditions. RAEGL mitigates this risk by retaining a validated global predictor and activating a contextual residual only when sufficient pre-deployment evidence supports its use. The framework employs a four-phase protocol that separates model training, candidate selection, gate calibration, and canonical evaluation to ensure rigorous evaluation of contextual components. RAEGL's deployment decisions are based on a search-aware randomization screen, practical gain thresholds, and temporal stability requirements. The methodology is validated through experiments on real-world panel data and controlled settings, demonstrating that RAEGL can effectively prevent harmful contextual deployments while clarifying opportunity costs. The results indicate that RAEGL significantly reduces RMSE degradations in various scenarios and maintains high stability in strong contextual settings while rejecting unstable corrections.
Methodology
RAEGL employs a four-phase protocol that includes model training, candidate selection, gate calibration, and canonical evaluation. It utilizes a search-aware evidence gate that reruns context-and-penalty searches within time-stratified randomization replicates to assess the deployment of contextual components based on randomization evidence, practical gain, and temporal stability.
Results
The evaluation of RAEGL on real-world panel data and controlled panels demonstrated that it effectively prevents harmful contextual deployments. In a reconstructed audit, RAEGL avoided RMSE degradations of 0.0960 and 0.0239 from validation-selected corrections. In another evaluation, a region-based correction was withheld due to insufficient practical gain and stability. The framework activated in 97.2% of strong, stable-context runs while rejecting all high-drift settings.
Implications
RAEGL provides a robust framework for managing contextual deployment risks in forecasting systems, making it a valuable tool for applications in various fields where temporal shifts are a concern. Its evidence-based approach can enhance the reliability of predictions in dynamic environments, such as economics, climate forecasting, and public health.
Equivariant Neural Primal-Dual Assignment for Maximum Common Edge Subgraphs
Graph Learning
Optimization
Theory
- ENPDA allows for a reusable matching policy that significantly reduces query time for MCES problems.
- The method provides per-pair guarantees and optimality bounds, ensuring reliable performance across different graph pairs.
- ENPDA outperforms traditional methods by 7.4-8.6 accuracy points on molecular benchmarks and shows strong transferability to other graph tasks.
- The approach maintains a frozen network structure while adapting bids and prices, optimizing the matching process efficiently.
Read more
Equivariant Neural Primal-Dual Assignment for Maximum Common Edge Subgraphs
Summary
This paper introduces the Equivariant Neural Primal-Dual Assignment (ENPDA), a novel approach for solving the Maximum Common Edge Subgraph (MCES) problem, which aims to find a partial vertex correspondence between two labeled graphs while preserving as many labeled edges as possible. Traditional methods often require separate training for each graph pair, leading to inefficiencies in applications such as molecular similarity search. ENPDA addresses this by learning a shared matching policy that can be applied to new graph pairs without additional training, resulting in significant speed improvementsโup to three orders of magnitude faster than existing methods. The method utilizes a neural auction framework where source vertices act as buyers and target vertices as objects, with bids reflecting matching preferences. The policy adapts to each graph pair while maintaining a frozen network structure, allowing for efficient updates and projections to achieve a one-to-one matching. The authors provide theoretical guarantees for the method's performance and demonstrate its effectiveness across multiple molecular benchmarks, showing substantial accuracy improvements over traditional analytic counterparts. ENPDA not only enhances performance in MCES tasks but also shows promise in edge-deletion tasks from social and protein graphs, indicating its versatility and robustness in various graph matching scenarios.
Methodology
ENPDA employs a neural auction-inspired framework where source vertices bid for target vertices, with a shared equivariant network initializing bids and guiding updates. The method includes four update rounds followed by a Hungarian projection to achieve a one-to-one matching. The approach leverages learned corrections and price adjustments to handle competition among source vertices effectively.
Results
ENPDA demonstrates a significant reduction in query time, averaging 0.19 seconds per query compared to 164-170 seconds for traditional methods. It achieves accuracy improvements of 7.4-8.6 points on molecular benchmarks and 9.1-17.6 points on edge-deletion tasks without fine-tuning. The method also recovers more reference bonds in aromatic ring structures compared to existing baselines.
Implications
The findings suggest that ENPDA can be effectively utilized in applications requiring rapid and accurate graph matching, such as drug discovery and bioinformatics. Its ability to generalize across different graph types indicates potential for broader applications in graph learning and optimization tasks.
EvoMO-SR: Multiobjective LLM-based Evolution of Symbolic Expressions with substructure guidance
Large Language Models
Optimization
Interpretability
- Introduction of EvoMO-SR, an LLM-driven framework for Symbolic Regression.
- Implementation of a multi-objective survival selection to manage formula complexity and accuracy.
- Use of a substructure guidance mechanism to enhance expression mutation.
- Demonstrated superior performance on LSR-Synth and custom datasets compared to traditional SR methods.
Read more
EvoMO-SR: Multiobjective LLM-based Evolution of Symbolic Expressions with substructure guidance
Summary
This paper presents EvoMO-SR, a novel framework for Symbolic Regression (SR) that leverages Large Language Models (LLMs) to generate equation skeletons while fitting coefficients through an external optimizer. The framework addresses common challenges in SR, such as formula bloating, by implementing a multi-objective survival selection mechanism that balances accuracy and complexity. Additionally, it introduces a substructure guidance mechanism that utilizes an archive of reusable functional substructures to enhance the mutation process of generated expressions. The authors evaluated EvoMO-SR on the LSR-Synth dataset and custom benchmark problems, demonstrating that it outperforms traditional SR methods in both in-domain and out-of-domain settings. The results indicate that EvoMO-SR achieves the lowest aggregate Normalized Mean Squared Error (NMSE) in seven out of eight comparisons, showcasing its effectiveness even with a smaller LLM model. Furthermore, the framework exhibits a higher probability of recovering accurate symbolic structures compared to existing methods, although overall symbolic accuracy remains a challenge.
Methodology
EvoMO-SR employs an evolutionary strategy where an LLM generates expression skeletons in Python code, while coefficients are optimized using L-BFGS-B. The framework incorporates a multi-objective survival selection to mitigate bloating and utilizes an archive of functional substructures to guide the mutation of expressions.
Results
EvoMO-SR achieved the lowest aggregate NMSE in seven out of eight comparisons on the LSR-Synth dataset, including all four out-of-domain settings. It consistently outperformed SR baselines and showed a greater probability of recovering accurate symbolic structures, although overall symbolic accuracy was not consistently high across methods.
Implications
The findings suggest that EvoMO-SR can significantly enhance the process of symbolic regression, making it a valuable tool for scientific discovery and data-driven analysis. Its ability to balance complexity and accuracy could lead to more interpretable models in various scientific fields.
When Trees Are Not Enough: Learning Mixed-Topology Feature Graphs with Adaptive Graph Sparse Autoencoders
NLP
Large Language Models
Graph Learning
- AG-SAE allows for mixed-topology feature graphs, overcoming the limitations of single-parent tree structures.
- The framework identifies necessary multi-parent relations while rejecting redundant or spurious alternatives.
- AG-SAE employs a self-consistency cycle between dictionary and graph learning for continuous refinement.
- Experimental results show improved relational reliability and semantic validity over existing methods.
Read more
When Trees Are Not Enough: Learning Mixed-Topology Feature Graphs with Adaptive Graph Sparse Autoencoders
Summary
The paper introduces the Adaptive Graph Sparse Autoencoder (AG-SAE), a novel framework that enhances the interpretability of features in large language model (LLM) activations by allowing for mixed-topology feature graphs. Traditional sparse autoencoders (SAEs) impose restrictive single-parent tree structures, which fail to capture the complex relationships among features. AG-SAE addresses this limitation by treating each feature's complete parent set as an atomic hypothesis, enabling the selection of zero, one, or multiple parents based on evidence. The framework competes complete parent sets against alternative explanations to identify necessary multi-parent relations while rejecting redundant associations. This process is guided by a differentiable structural loss that informs SAE training, allowing for topology-guided refinement and mitigating issues like feature absorption. The methodology involves a self-consistency cycle between dictionary and graph learning, where features are continuously reassessed and refined. Experimental results demonstrate that AG-SAE achieves exact mixed-topology recovery in controlled settings and outperforms existing methods in terms of relational reliability and semantic validity on real LLM activations. The framework also exhibits stronger causal control compared to conventional SAEs, ultimately providing a more robust unsupervised training signal for feature organization.
Methodology
The AG-SAE framework employs a structure-guided training paradigm that assesses complete parent sets for each feature, competing them against null and alternative explanations. It utilizes a differentiable structural loss to guide training and incorporates a self-consistency cycle between dictionary and graph learning to refine feature representations continuously.
Results
AG-SAE demonstrated exact mixed-topology recovery in controlled experiments and exhibited greater relational reliability and semantic validity compared to structured and post-hoc baselines on real LLM activations. The framework also achieved stronger feature-level causal interventions than conventional sparse autoencoder features.
Implications
The AG-SAE framework has potential applications in improving the interpretability and organization of features in large language models, facilitating better understanding of knowledge representation and enhancing causal analysis in machine learning systems.
NeuronSifter: Intervention Planning in CNS Microenvironments
Optimization
Theory
- NeuronSifter integrates structured intervention compilation with decision-directed evidence acquisition.
- The framework improves intervention ordering accuracy from 0.760 to 0.880 in synthetic AD evaluations.
- It retains joint uncertainty across candidate rollouts, enhancing decision quality.
- NeuronSifter demonstrates superior performance compared to traditional Bayesian experimental design planners.
Read more
NeuronSifter: Intervention Planning in CNS Microenvironments
Summary
The paper introduces NeuronSifter, a novel framework for planning interventions in the central nervous system (CNS) microenvironments, particularly in the context of Alzheimer's disease (AD). The authors argue that effective intervention planning requires not only predicting the effects of a treatment regimen but also understanding the dynamics of the microenvironment where the intervention occurs. NeuronSifter addresses the limitations of traditional action-conditioned predictors that oversimplify the complexity of CNS interactions by treating decision quality as a property of the intervention interface rather than the controller placement. The framework compiles treatment regimens into state-conditional target-occupancy fields and utilizes a stochastic diffusion operator to propagate these fields through the microenvironment dynamics. It selects measurements based on their expected reduction in intervention loss, thereby integrating various outcomes into a shared posterior. The authors validate NeuronSifter through a synthetic evaluation, demonstrating significant improvements in intervention ordering accuracy and reduced trajectory continuous ranked probability scores. The framework shows promise in optimizing intervention strategies by effectively managing uncertainty and enhancing decision-making processes in CNS interventions.
Methodology
NeuronSifter employs a structured intervention compiler to map treatment regimens into state-conditional exposure and occupancy fields. It utilizes a stochastic diffusion operator to propagate these fields through the dynamics of the CNS microenvironment. The framework incorporates a decision-directed evidence acquisition mechanism that selects measurements based on their potential to reduce intervention loss, maintaining a shared posterior across candidate interventions.
Results
In a synthetic evaluation involving 64 paired scenario blocks for Alzheimer's disease, NeuronSifter achieved a reduction in trajectory continuous ranked probability score from 0.165 to 0.110 and increased intervention ordering accuracy from 0.760 to 0.880. The decision-directed acquisition method attained a terminal risk of 0.160, demonstrating competitive performance against a matched Bayesian experimental design planner.
Implications
NeuronSifter has the potential to enhance intervention planning in CNS disorders by providing a more nuanced understanding of microenvironment dynamics and improving decision-making processes. Its ability to manage uncertainty and optimize treatment strategies could lead to more effective therapies for conditions like Alzheimer's disease.
Replay in the Silent Degrees of Freedom: Continual Learning Without an Offline Phase
Theory
Efficient ML
Optimization
- Introduces a continual learning framework that operates without an offline phase.
- Utilizes biologically inspired mechanisms such as k-winner-take-all dynamics and refractory rotation.
- Achieves competitive performance on benchmark datasets, surpassing traditional methods in certain conditions.
- Demonstrates the potential for memory consolidation during active inference rather than requiring dedicated offline processing.
Read more
Replay in the Silent Degrees of Freedom: Continual Learning Without an Offline Phase
Summary
This paper explores a novel approach to continual learning in artificial neural networks that eliminates the need for an offline consolidation phase, inspired by biological mechanisms of memory consolidation during wakefulness. The proposed method utilizes a combination of isolation and refractory rotation rules to confine replay updates to hidden synapses that are not currently active, allowing for memory consolidation during inference. The architecture employs a k-winner-take-all (k-WTA) mechanism to ensure that only unused synapses are updated, thus preserving the integrity of current learning tasks. The system was evaluated on class-incremental split-MNIST and split CIFAR-10 datasets, demonstrating competitive performance compared to existing methods that rely on offline rehearsal or interleaved replay. The results indicate that the proposed method can achieve high accuracy while maintaining a biologically plausible learning framework, suggesting a shift in how continual learning can be approached in artificial networks.
Methodology
The methodology involves a biologically constrained learning framework where replay updates are confined to hidden synapses not currently active during inference. The isolation rule limits updates to silent units, while the refractory rotation rule prevents recently active units from participating in the next competition, allowing for a broader set of synapses to be consolidated. Two internal signals manage the timing of replay bursts and rotations, ensuring that past memories can be updated without interfering with current learning.
Results
On the split-MNIST dataset, the proposed system achieved an accuracy of 91.6 ยฑ 0.3%, comparable to the best offline methods and outperforming several interleaved replay strategies. In a single pass over the data, it led to a slight improvement over DER++ (91.8% vs. 90.1%). For split CIFAR-10, while it performed well, it trailed behind ER-ACE and DER++ by two to three points, with the gap attributed to dense activation rather than credit assignment issues.
Implications
The findings suggest that continual learning systems can be designed to operate more like biological systems, potentially leading to more efficient and robust learning algorithms. This approach could have applications in areas requiring lifelong learning, such as robotics, adaptive systems, and real-time data processing.
Replication Failure and Trivial Baselines in Road-Level Crash Prediction
Graph Learning
- Only 4 out of 11 design decisions from a previous model replicate in a second borough.
- Multi-seed evaluation reveals significant variability in model performance.
- The proposed GNN model is statistically indistinguishable from a trivial baseline based on past crash counts.
- Published results may overstate the advantages of GNNs due to short lookback horizons.
Read more
Replication Failure and Trivial Baselines in Road-Level Crash Prediction
Summary
This paper investigates the reliability of graph neural networks (GNNs) in road-level crash prediction, focusing on the stability of reported gains from previous models. The author independently reconstructs the data pipeline of a recent uncertainty-aware model and evaluates eleven design decisions across three London boroughs using an expanding-window protocol. The findings reveal that only four of the eleven design decisions replicate in a second borough, with several effects reversing in sign rather than merely attenuating. The study emphasizes the importance of multi-seed evaluation, demonstrating significant variability in results across different random seeds. Furthermore, the paper compares the GNN models against a trivial baseline that ranks road segments by cumulative past crash counts. The results indicate that the proposed model is statistically indistinguishable from this baseline, which shows a much wider range of accuracy based on lookback horizons. The author argues that the perceived advantages of GNNs over historical baselines are largely artifacts of the short horizons used in previous studies. The paper advocates for horizon-matched baselines and multi-seed reporting as essential practices for future research in this area.
Methodology
The author reconstructed the data pipeline of a previous uncertainty-aware model and evaluated its design decisions across three London boroughs. An expanding-window protocol was used for evaluation, and a multi-seed approach was employed to assess the stability of results. The GNN models were compared against a trivial baseline that ranked segments by cumulative past crash counts.
Results
The study found that only four design decisions replicated in a second borough, with several effects reversing in sign. The GNN model was statistically indistinguishable from the trivial baseline, which outperformed the reference architecture in all held-out windows tested. The baseline's accuracy varied significantly based on the lookback horizon, indicating that previously reported gains may be artifacts of short evaluation periods.
Implications
The findings suggest that researchers should adopt more rigorous evaluation practices, including horizon-matched baselines and multi-seed reporting, to ensure the reliability of results in road-level crash prediction. This could lead to more accurate identification of high-risk road segments and better-informed interventions for road safety.
Why Backdooring Neural Networks is so Easy?
Theory
- Feature learning in neural networks increases vulnerability to backdoor attacks.
- The relationship between poison fraction and trigger strength is characterized by ฮฑ โ ฯโ1/4 in feature-learning regimes.
- Existing security audits may underestimate backdoor risks due to reliance on linear heuristics.
- The study provides a theoretical framework that aligns with large-scale empirical findings.
Read more
Why Backdooring Neural Networks is so Easy?
Summary
This paper addresses the vulnerability of neural networks to backdoor attacks, which pose significant threats to AI systems. The authors derive a closed-form analysis of a quadratic neuron trained on a poisoned Gaussian mixture, revealing that the dynamics of feature learning, while enhancing model performance, also increase susceptibility to backdoor attacks. They demonstrate that the required trigger strength for successful attacks decreases as the poison fraction increases, particularly in feature-learning regimes. This counterintuitive finding suggests that existing security audits based on linear heuristics may underestimate backdoor vulnerabilities in modern neural networks. The study provides theoretical insights consistent with empirical observations, emphasizing the need for more robust auditing frameworks that account for the complexities of feature learning in neural networks.
Methodology
The authors utilize a statistical model involving a quadratic neuron trained on a Gaussian mixture with a fraction of poisoned samples. They analyze two training regimes: a linear regime where the quadratic feature is not utilized and a feature-learning regime where the feature direction adapts to the data. The analysis includes deriving closed-form expressions for clean accuracy and attack success rates under varying conditions of noise and trigger strength.
Results
The results indicate that in the linear regime, the trigger strength required for a successful attack scales as ฮฑ โ ฯโ1/2, while in the feature-learning regime, this relationship changes to ฮฑ โ ฯโ1/4. This means that feature learning allows for successful backdoor attacks with lower trigger strengths at smaller poison fractions, highlighting a significant vulnerability in modern neural networks. Experimental validation on datasets such as MNIST and CIFAR-10 supports these theoretical findings.
Implications
The findings suggest that as neural networks become more complex and capable through feature learning, they simultaneously become more susceptible to backdoor attacks. This necessitates the development of more sophisticated auditing and defense mechanisms that can effectively identify and mitigate such vulnerabilities in AI systems.
Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution
Optimization
- Mara Chain retains and refines rejected candidates instead of discarding them, leveraging past failures for future improvements.
- The method limits the depth of refinement chains and employs Pareto-filtered Top-N selection to manage candidate pools effectively.
- Mara Chain outperforms existing optimization methods by significant margins, achieving better results with fewer rollouts.
- The approach is validated across multiple benchmarks, demonstrating its versatility across different types of AI artifacts.
Read more
Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution
Summary
The paper introduces Mara Chain, a novel optimization procedure for AI systems that reconsiders the treatment of rejected candidates in the propose-evaluate-select loop. Traditional methods discard candidates that do not meet performance criteria, which can lead to repetitive failures and stagnation in optimization. Mara Chain retains these rejected candidates and refines them iteratively, utilizing accumulated evidence from previous attempts to inform future proposals. This approach allows for the exploration of previously discarded candidates, potentially uncovering valuable insights that can lead to improved performance. The authors demonstrate the effectiveness of Mara Chain across three distinct benchmarks: AppWorld, TerminalBench 2.1, and MuSiQue, showing significant performance improvements over existing methods. The results indicate that Mara Chain not only enhances task performance but also reduces the number of rollouts required to achieve optimal configurations, thereby increasing efficiency in the optimization process.
Methodology
Mara Chain operates within the propose-evaluate-select framework, where rejected candidates are retained and iteratively refined based on historical evidence. The method limits the depth of refinement chains and utilizes a Pareto-filtered selection process to maintain a manageable candidate pool.
Results
Mara Chain achieved up to 20.5% better performance on AppWorld compared to leading methods while requiring 65.5% fewer rollouts. It also improved the pass rate on TerminalBench 2.1 by 20.2 and 22.5 percentage points over competing methods and enhanced retrieval metrics on MuSiQue by 34.6% and 39.8% for nDCG@10 and Recall@10, respectively.
Implications
The findings suggest that retaining and building upon failed attempts can significantly enhance the optimization of AI systems, potentially leading to more robust and efficient AI applications across various domains.
Cheap and Powerful Tests for Supervised Subspaces: Per-Component Inference for PLS
Theory
Efficient ML
Interpretability
- Introduction of two new tests for PLS inference: a corrected t-test and a permutation test.
- Demonstration of significant power improvements over existing methods like CV-permutation-Q2.
- Validation of methods on various datasets, including synthetic and real-world applications.
- Establishment of a framework for per-component inference in PLS models.
Read more
Cheap and Powerful Tests for Supervised Subspaces: Per-Component Inference for PLS
Summary
This paper addresses the challenges of inference in Partial Least Squares (PLS) regression, which is commonly used in high-dimensional data analysis across various scientific fields. The authors propose a novel approach to inference that leverages held-out Ordinary Least Squares (OLS) refits of the supervised subspace, providing two new tests: a NadeauโBengio corrected asymptotic t-test and a permutation test. These tests are designed to be both cost-effective and powerful, overcoming the limitations of existing methods that are either biased, expensive, or lack calibration. The authors validate their methods on synthetic datasets, two Near-Infrared (NIR) chemometric datasets, and cross-lingual word-embedding regressions. They demonstrate that their proposed tests outperform the traditional cross-validated permutation-Q2 method in terms of power and cost efficiency. Additionally, the authors introduce a mechanism for per-component claims, allowing for more granular inference on individual components of the PLS model. The paper also includes a Rust library with bindings for Python, R, and Julia, facilitating practical application of their methods.
Methodology
The authors utilize held-out OLS refits of the supervised subspace to derive their tests. They propose a NadeauโBengio corrected resampled t-test for fast approximation and a permutation test that is valid under outcome-predictor independence. The methodology includes checks for sample size and data spectrum to ensure the validity of the approximation. They also develop a fixed-sequence test for per-component claims based on the order of PLS extraction.
Results
The proposed tests demonstrated greater power than the traditional CV-permutation-Q2 method while being significantly less costly. The NadeauโBengio test was validated under conditions of stable rank and sufficient sample size, while the permutation test was shown to maintain validity under iid assumptions. The authors also confirmed that their tests could be applied to supervised PCA and ridge probes, expanding the applicability of their methods.
Implications
The findings of this paper have significant implications for researchers using PLS regression and similar methods in high-dimensional data analysis. The ability to conduct efficient and powerful inference can enhance the reliability of predictive models in various fields, including chemometrics, natural language processing, and beyond. The release of the software library also facilitates broader adoption of these methods in practical applications.
DP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation
Efficient ML
Time Series
NLP
- DP-Rec is the first to apply dynamic latent patching for sequential recommendation, achieving superior efficiency and accuracy.
- Introduces Contrastive Entropy Surprise as a computationally efficient method for determining dynamic behavioral boundaries.
- Incorporates temporal dynamics as first-class features to enhance the detection of informative segment boundaries.
- Demonstrates significant improvements in performance under constrained computational budgets across multiple datasets.
Read more
DP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation
Summary
The paper introduces DP-Rec, a novel dynamic latent patching architecture aimed at improving the efficiency of long-sequence recommendation systems. Traditional transformer models struggle with computational inefficiency due to their fixed-rate processing of user history, which leads to excessive resource consumption and sensitivity to noise in user behavior. DP-Rec addresses these issues by shifting from item-level to patch-level modeling, segmenting interaction sequences based on contrastive entropy surprise to identify significant behavioral boundaries. This approach allows for the creation of dynamic latent behavior vectors that are processed by a larger latent transformer for next-item prediction. The authors demonstrate that DP-Rec effectively scales to long sequences and achieves a superior efficiency-accuracy trade-off compared to both non-compressed and fixed-size compression baselines. The methodology incorporates temporal dynamics as a critical factor in determining behavioral context, enhancing the model's ability to adapt to varying information density in user interactions. The extensive experiments conducted on datasets such as KuaiRand, ML-1M, and ML-10M-L validate the effectiveness of DP-Rec in real-world recommendation scenarios.
Methodology
The methodology involves a dynamic latent patching architecture that segments user interaction sequences based on contrastive entropy surprise. This segmentation identifies informative behavioral boundaries, allowing for the compression of sequences into dynamic latent behavior vectors. These vectors are then processed by a larger latent transformer model for next-item prediction, integrating temporal dynamics as key features in the patching mechanism.
Results
DP-Rec achieves Pareto-superior trade-offs in efficiency and accuracy across various datasets, including KuaiRand, ML-1M, and ML-10M-L. The model demonstrates higher inference efficiency for matched accuracy and improved accuracy within fixed computational budgets compared to existing baselines.
Implications
The findings suggest that dynamic patching can significantly enhance the performance of recommendation systems, particularly in industrial applications where computational resources are limited. This approach could lead to more effective personalization strategies that leverage long user histories without incurring prohibitive computational costs.
Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks
Large Language Models
NLP
Theory
- PriceCheck introduces a pricing mechanism for label-free checks to control selective risk in LLM answering.
- The method achieves an average of 76.1% answer serving while keeping selective risk below 1.5%.
- Price-based coverage predictions show a high correlation (0.97) with observed coverage across various schedules.
- PriceCheck outperforms traditional methods, including reward models and correctness classifiers, in terms of serving more answers with fewer errors.
Read more
Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks
Summary
This paper introduces PriceCheck, a novel approach for risk-controlled selective answering in large language models (LLMs). The authors highlight the challenge of determining when to abstain from providing answers, as traditional ranking accuracy does not adequately reflect the error rate of served answers. PriceCheck utilizes a compact family of decision rules derived from label-free checks, such as re-solving problems, each associated with a 'price' that reflects its agreement rates on correct and incorrect answers as well as its operational cost. By fitting these prices on a small, class-enriched labeled dataset, the method predicts the coverage and cost of various schedules, guiding the selection of checks to run and the timing for stopping. The results demonstrate that PriceCheck serves an average of 76.1% of answers while maintaining a selective risk below 1.5% across multiple splits, outperforming existing methods such as reward models and correctness classifiers. The findings emphasize the importance of strategically combining and stopping checks, alongside the ranking quality of verifiers, to enhance the performance of LLMs in selective answering tasks.
Methodology
The authors developed PriceCheck by creating a family of decision rules based on label-free checks. Each check is assigned a price based on its agreement rates and operational cost. The method involves fitting these prices on a labeled dataset, composing them into schedules, and selecting an optimal schedule that meets a specified selective-risk target. The evaluation is conducted using competition mathematics problems generated by fine-tuned models, comparing PriceCheck against various existing methods.
Results
PriceCheck successfully served an average of 76.1% of answers while maintaining a selective risk below 1.5% across 15 data splits. It demonstrated superior performance compared to reward models, prompted judges, and correctness classifiers, retaining the fewest wrong answers at matched coverage levels. The price-based predictions showed a rank correlation of 0.97 with actual coverage observed during testing.
Implications
The findings suggest that PriceCheck can enhance the reliability and efficiency of LLMs in applications requiring selective answering, such as automated tutoring systems, customer support, and decision-making tools. The approach may also inform future research on risk management in AI systems, particularly in contexts where accuracy and reliability are critical.
Selective Elicitation as a Commercial Influence Channel: A Reproducible Synthetic Shopping-Agent Stress Test
Theory
- Commercial incentives can influence the questions asked by AI shopping assistants, affecting user responses.
- The study contrasts neutral, soft commercial, and targeted questioning policies in a synthetic setting.
- Targeted questioning significantly increases the selection of sponsored products while reducing overall utility.
- The findings highlight the need for careful consideration of question policies in AI recommendation systems.
Read more
Selective Elicitation as a Commercial Influence Channel: A Reproducible Synthetic Shopping-Agent Stress Test
Summary
This paper investigates the influence of commercial incentives on AI shopping assistants, specifically how such incentives can affect the questions these assistants ask rather than the final product ranking. The study employs a synthetic experimental framework where a simulated user answers a pairwise question about two products with known attributes and a fixed preference vector. The research contrasts three types of questioning policies: a neutral question, a soft commercial instruction, and an explicitly adversarial instruction aimed at highlighting the sponsor's advantages while downplaying the competitor's. The findings reveal that while the soft instruction has no effect on product selection, the targeted instruction increases the selection of the sponsored product by 0.30 and decreases the mean synthetic utility by 0.0547 compared to neutral questioning. The study emphasizes the importance of understanding the implications of question policies in AI recommendations and provides a controlled stress-test protocol to explore these dynamics.
Methodology
The methodology involves a synthetic experimental design where a simulated user answers questions about two products based on their attributes and a fixed preference vector. The study employs three distinct questioning policies to observe their effects on product selection and utility.
Results
The results indicate that the soft commercial instruction does not alter product selections, while the targeted instruction increases the likelihood of selecting the sponsored product by 0.30 and reduces mean synthetic utility by 0.0547 compared to neutral questioning. A Bayesian recommender exhibited similar effects, and a consistency judge rated the targeted answers as consistent despite some instances of synthetic regret.
Implications
The findings suggest that AI shopping assistants can be subtly influenced by the way questions are framed, which has significant implications for the design of recommendation systems and the ethical considerations surrounding commercial influences in AI. The study provides a framework for future research on the impact of question policies in AI-driven recommendations.
Shared Autoregressive Context Can Distort Relationships in Synthetic Data
Large Language Models
Generative Models
Theory
- Shared autoregressive context can significantly distort relationships in synthetic data generated by LLMs.
- Generating multiple respondents in one completion increases mean absolute error in correlations compared to generating them separately.
- Answer history acts as a causal channel influencing the relationships among generated responses.
- Hiding preceding answers can reduce correlation errors but may worsen marginal accuracy.
Read more
Shared Autoregressive Context Can Distort Relationships in Synthetic Data
Summary
This paper investigates the impact of shared autoregressive context in large language models (LLMs) on the relationships among variables in synthetic data. By conducting controlled experiments with synthetic survey respondents, the author demonstrates that generating multiple records in a single autoregressive completion can distort statistical relationships, particularly within-country correlations. The study utilizes 2,000 European Social Survey profiles to show that generating ten respondents at once increases the mean absolute error in correlations by 48-58% for the Qwen3.8-27B model and 114-127% for the Llama-3.3-70B-Instruct model. The distortion is characterized by exaggerated relationship strengths while still aligning with human correlation orderings. The author establishes that answer history serves as a causal channel affecting subsequent responses, and interventions that hide preceding answers can reduce correlation errors, albeit at the cost of marginal accuracy. The findings suggest that the construction of requests is integral to the data-generating process, necessitating careful evaluation of synthetic data against the analyses they are intended to support.
Methodology
The study employs controlled experiments using synthetic survey respondents based on European Social Survey profiles. It compares the generation of multiple respondents in a single autoregressive completion against generating them individually, measuring the mean absolute error in within-country correlations. Controlled interventions are also conducted to analyze the influence of answer history on generated responses.
Results
The results indicate that generating ten respondents at once leads to a 48-58% increase in mean absolute error for the Qwen3.8-27B model and a 114-127% increase for the Llama-3.3-70B-Instruct model. The distortion primarily manifests as exaggerated relationship strengths while maintaining a significant agreement with human correlation orderings. Interventions show that manipulating answer history can alter correlations among generated responses.
Implications
These findings highlight the importance of understanding how request construction influences synthetic data generation, which is crucial for ensuring the validity of analyses based on such data. The results may inform best practices for using LLMs in generating synthetic datasets for research and policy-making.
Arithmetic Simplicity in Stochastic Gradient Methods
Optimization
Efficient ML
Theory
- Introduces the concept of arithmetic simplicity in gradient descent methods.
- Transforms AdaGrad, Adam, and AdamW into arithmetically simple versions suitable for hardware implementation.
- Demonstrates faster convergence for the static version of Adam in experimental results.
- Highlights the benefits of reduced hardware complexity and energy consumption.
Read more
Arithmetic Simplicity in Stochastic Gradient Methods
Summary
This paper explores the concept of arithmetic simplicity in stochastic gradient methods, proposing transformations of popular algorithms like AdaGrad, Adam, and AdamW to limit operations to basic arithmetic functions and restricted division. The authors argue that such transformations enhance the ease of implementation in hardware, particularly for chip design, as they avoid complex operations that increase latency and energy costs. The paper details how to convert these algorithms into arithmetically simple forms, focusing on a static version of Adam that improves convergence speed in experiments. The convergence analysis for the static Adam is provided, demonstrating its effectiveness while maintaining arithmetic simplicity. The proposed methods are expected to be beneficial for hardware implementations, potentially leading to simpler architectures and improved computational efficiency.
Methodology
The authors transform existing stochastic gradient methods into arithmetically simple forms by limiting operations to addition, subtraction, multiplication, and division by powers of two. They analyze the convergence of these transformed methods, particularly focusing on a static version of Adam and AdamW, where the step size is determined by a fixed function of the number of iterations.
Results
The transformed static Adam method shows improved convergence rates in experiments compared to its more complex counterparts. The paper provides theoretical guarantees for the convergence of the static Adam, affirming its effectiveness while adhering to the constraints of arithmetic simplicity.
Implications
The findings suggest that adopting arithmetically simple gradient methods can lead to more efficient hardware implementations, particularly in specialized architectures like FPGAs and ASICs. This could result in lower energy consumption and improved performance in large-scale machine learning applications.
Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories
Large Language Models
NLP
Optimization
- Mid-training performance in LLMs saturates with increased compute, necessitating alternative allocation strategies.
- Trajectory Soup enables the distribution of compute across multiple independent trajectories, enhancing model performance.
- Inter-trajectory averaging significantly reduces residual error compared to traditional intra-trajectory methods.
- The method shows improved performance across different model scales, learning rates, and token budgets.
Read more
Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories
Summary
The paper addresses the limitations of mid-training in large language models (LLMs), where additional compute often yields diminishing returns or even degrades performance. The authors propose a novel approach called 'Trajectory Soup' that optimizes compute allocation by distributing it across multiple independent training trajectories rather than extending a single trajectory. By forking branches from a shared checkpoint and allowing them to explore diverse optimization paths, the method captures complementary updates. The strongest checkpoints from these trajectories are then averaged to create a consolidated model. The study includes a local bias and variance analysis that highlights the benefits of inter-trajectory averaging over intra-trajectory methods. The results demonstrate that Trajectory Soup consistently outperforms traditional single-trajectory training methods across various model scales and training configurations, effectively extending the compute-scaling frontier of mid-training.
Methodology
The authors introduce Trajectory Soup, which involves forking multiple independent training trajectories from a common checkpoint. Each trajectory is trained for the same number of tokens with controlled variations in the training recipe. The strongest checkpoints from each trajectory are selected and averaged to form a consolidated model. A local bias and variance analysis is conducted to understand the contributions of intra- and inter-trajectory averaging.
Results
The empirical results indicate that Trajectory Soup outperforms the strongest single-trajectory averages and traditional inter-trajectory methods under matched budgets. The performance gains increase with the addition of more trajectories, demonstrating that the method effectively utilizes compute resources to enhance downstream performance.
Implications
The findings suggest that optimizing compute allocation through diverse trajectories can significantly improve the training efficiency and effectiveness of large language models. This approach may lead to better performance in various NLP tasks and could influence future research on model training strategies.
Behavioral Capacity Certificates for Quantized Language Models
NLP
Large Language Models
Efficient ML
- Introduces Behavioral Capacity Certificates (BCC) for quantifying model behavior in quantized language models.
- Proposes a three-step workflow for model deployment that includes screening, certification, and bounding of population loss.
- Demonstrates that BCC can lower complexity penalties while preserving model performance.
- Finds that higher precision for keys than values in cache memory improves model predictions.
Read more
Behavioral Capacity Certificates for Quantized Language Models
Summary
This paper introduces Behavioral Capacity Certificates (BCC) as a novel approach to quantify the behavioral complexity of quantized language models. The authors argue that traditional weight-code bounds fail to account for the variability in model behavior stemming from different activation and cache precisions, which can lead to misleading complexity assessments. BCC addresses this by aggregating the prior mass of all implementations that yield the same bounded loss, thus allowing for a more accurate representation of model behavior. The proposed methodology involves a three-step deployment workflow: first, a screening process that identifies optimal per-layer bit-widths based on the preservation of reference predictions; second, the identification of margin-certified cells that can be pruned or sign-flipped without altering model behavior; and third, the bounding of population loss for the deployed model. Experimental results demonstrate that BCC can effectively lower complexity penalties while maintaining prediction accuracy across various models, including GPT-2 and SmolLM2. The findings suggest that higher precision for keys than values in cache memory can lead to improved performance metrics, such as lower negative log likelihood (NLL) and better prediction agreement.
Methodology
The methodology involves a three-step deployment process: (1) screening candidates for per-layer bit-widths based on their ability to preserve reference predictions, (2) identifying margin-certified cells that can be modified without affecting behavior, and (3) bounding the population loss of the deployed model using BCC.
Results
The experiments validate each step of the proposed workflow, showing that BCC can effectively aggregate prior mass and lower complexity penalties. The results indicate that models with higher precision keys than values achieve lower NLL and improved prediction agreement, particularly in the tested models GPT-2, Qwen2.5, and SmolLM2.
Implications
The findings of this paper have significant implications for the deployment of quantized language models, suggesting that BCC can enhance model efficiency and performance while providing a robust framework for certifying model behavior across different quantization strategies.
Strict-Saddle Landscapes and Multi-Rank Geometry in Low-Tubal-Rank Tensor Sensing
Optimization
Theory
- Establishment of a strict-saddle landscape for low-tubal-rank tensor sensing with no spurious local minima.
- Local geometry is determined by Fourier-slice ranks rather than solely by tubal rank.
- Uniform ranks yield quadratic growth, while nonuniform ranks lead to quartically flat directions.
- Numerical experiments confirm the global optimization behavior and highlight differences in local geometries.
Read more
Strict-Saddle Landscapes and Multi-Rank Geometry in Low-Tubal-Rank Tensor Sensing
Summary
This paper investigates the optimization landscape of low-tubal-rank tensor sensing through a balanced factorization approach. The authors establish a quantitative strict-saddle landscape under a tubal restricted isometry condition, demonstrating that there are no spurious local minima for arbitrary Fourier multi-rank profiles. The study reveals that the local geometry of the optimization landscape is influenced more by the Fourier-slice ranks than by the tubal rank alone. Specifically, uniform ranks lead to quadratic growth in directions transverse to the solution orbit, while nonuniform ranks result in quartically flat directions due to hidden frequency-wise overparameterization. The authors conduct numerical experiments to illustrate the global optimization behavior and the contrasting local geometries, suggesting that the favorable global landscape observed in low-rank matrix sensing can also persist in low-tubal-rank tensor sensing under certain conditions. However, the local geometry can vary significantly based on the multi-rank profile of the tensor, indicating a complex interplay between global and local optimization characteristics.
Methodology
The authors utilize a balanced factorization approach to optimize low-tubal-rank tensors, applying a tubal restricted isometry condition to analyze the optimization landscape. They derive a balanced objective function that incorporates a term to control scaling ambiguity between factors and preserve t-orthogonal symmetry. The study employs numerical experiments to validate theoretical findings regarding the optimization landscape.
Results
The paper demonstrates that under suitable conditions, the optimization landscape for low-tubal-rank tensor sensing exhibits a strict-saddle structure with no spurious minima. It shows that while the global landscape can be favorable, the local geometry is significantly affected by the multi-rank profile of the tensor, with distinct behaviors observed for uniform and nonuniform ranks.
Implications
The findings have implications for tensor recovery methods in various applications, including signal processing and imaging, where understanding the optimization landscape can lead to more effective algorithms for tensor factorization and recovery.
When Is an SAE Feature Interpretable? A Validation Ladder for EEG Foundation Models
Interpretability
Time Series
- The initial interpretation of SAE latents as representing alpha activity is misleading and largely due to input distortion effects.
- A proposed validation ladder helps assess the interpretability of SAE features, emphasizing the need for rigorous testing.
- The study reveals that latents selected for their response to alpha removal do not correlate with clean EEG alpha power.
- The findings challenge existing assumptions about the semantic identifiability of features in EEG foundation models.
Read more
When Is an SAE Feature Interpretable? A Validation Ladder for EEG Foundation Models
Summary
This paper investigates the interpretability of features derived from Sparse Autoencoders (SAEs) in the context of EEG foundation models. The authors highlight a common misconception that a significant change in latent activation due to alpha-band activity removal implies that the latent represents alpha activity. Through extensive experiments across 27 different settings, they demonstrate that while the initial interpretation appears compelling (with a 7.3ร increase in latent firing upon alpha removal), this is largely an artifact of the input distortion caused by the alpha filter, which removes more signal than a sham intervention. After normalizing for the spectral energy removed, the ratio drops to 0.28ร, indicating that the latents do not reliably represent alpha activity. The authors propose a validation ladder to assess the semantic interpretations of SAE latents, which includes checking for sensitivity to perturbations, controlling for distortion, ensuring specificity, and validating against unperturbed data. They conclude that perturbation sensitivity alone is insufficient to establish the physiological meaning of a latent, and they provide a structured approach to evaluate claims about the interpretability of SAE features in EEG analysis.
Methodology
The authors conducted experiments using three different EEG datasets and three backbone models (CBraMod, REVE, and DINOv3) across 27 settings. They employed Sparse Autoencoders to analyze latent activations in response to alpha-band activity removal and compared these results against sham interventions. The validation ladder was developed to systematically evaluate the interpretability of the latents based on sensitivity, distortion control, specificity, and independent validation.
Results
The study found that while alpha removal initially appeared to significantly affect latent activation (7.3ร increase), this effect was largely due to the greater signal removal caused by the alpha filter. After normalizing for the spectral energy removed, the effect size dropped to 0.28ร, with no settings showing a significant correlation with alpha power in clean EEG data. The proposed validation ladder effectively identified the limitations of interpreting SAE latents as physiological features.
Implications
The findings suggest that researchers should exercise caution when interpreting features from SAEs in EEG analysis. The proposed validation ladder can serve as a framework for future studies to ensure that claims about the physiological relevance of latent features are substantiated. This has broader implications for the development of interpretable machine learning models in biomedical applications.
Multi-Agent Flow Matching with Decoupled Generative Guidance
Generative Models
Robotics
Theory
- Introduces DeGG-Flow for multi-agent flow matching with decoupled generative guidance.
- Establishes formal guarantees for both shared and private requirements in multi-agent systems.
- Provides feasibility and finite-horizon convergence guarantees for generated outputs.
- Derives a Wasserstein bound to characterize distributional deviation under guidance.
Read more
Multi-Agent Flow Matching with Decoupled Generative Guidance
Summary
This paper introduces DeGG-Flow, a novel framework for multi-agent flow matching that incorporates decoupled generative guidance to ensure that generated objects meet hard constraints. Traditional generative models often lack guarantees that outputs satisfy specific requirements, particularly in multi-agent scenarios where constraints may depend on multiple agents. DeGG-Flow addresses this by formulating the generative process as a control-affine dynamical system, allowing each agent to independently determine its guidance input while still satisfying team-level requirements. The framework distinguishes between shared requirements, which necessitate collaboration among agents, and private requirements, which are specific to individual agents but may still depend on their neighbors. The authors establish feasibility conditions and finite-horizon convergence guarantees for both types of requirements, ensuring that the generated outputs adhere to the necessary constraints without the need for retraining or additional data collection. Furthermore, a Wasserstein bound is derived to quantify the distributional deviation caused by the guidance, providing theoretical insights into the impact of the generative process. The effectiveness of DeGG-Flow is demonstrated through applications in multi-robot collaboration and multi-object scene generation, showcasing its ability to produce compliant outputs even in scenarios with unseen team sizes during training.
Methodology
The authors model the generative process as a control-affine dynamical system, applying decoupled guidance for each agent. They develop adaptive constraint allocation methods and establish theoretical bounds for feasibility and convergence, while also deriving a Wasserstein bound to assess distributional deviation.
Results
DeGG-Flow successfully generates outputs that satisfy hard requirements in multi-agent scenarios, demonstrating feasibility and convergence in both shared and private requirement contexts. The empirical evaluations confirm the theoretical bounds on distributional deviation, validating the framework's effectiveness in practical applications.
Implications
The proposed framework has significant implications for decentralized multi-agent systems, particularly in robotics and collaborative tasks, where ensuring compliance with constraints is critical. It allows for flexible and robust generation of behaviors in dynamic environments without the need for retraining models.
Predictive Dual Smoothing for Column Generation
Optimization
- Introduction of predictive dual smoothing that uses future dual predictions for improved pricing.
- The method combines current dual solutions with learned predictions to enhance column generation.
- Extensive evaluation shows substantial efficiency gains in solving large-scale linear programs.
Read more
Predictive Dual Smoothing for Column Generation
Summary
This paper addresses the challenge of efficiently solving large-scale linear programs using column generation (CG), a technique that alternates between solving a restricted master problem and identifying new variables through a pricing subproblem. The authors introduce a novel method called predictive dual smoothing, which enhances traditional dual smoothing by incorporating predictions of future dual solutions into the pricing process. This approach aims to mitigate the oscillations in dual solutions that can hinder convergence. By training a predictor offline on standard CG trajectories, the authors demonstrate that predictive dual smoothing can effectively guide the pricing subproblem towards more useful variables, thus improving the overall efficiency of the CG process. The method is evaluated on cutting stock and generalized assignment problems, showing significant reductions in both the number of generated columns and runtime compared to standard CG and existing stabilization techniques.
Methodology
The authors developed predictive dual smoothing by predicting future dual solutions based on the current state of the column generation process. A shared predictor is trained offline using supervised learning from historical CG trajectories, allowing the method to anticipate future needs in the pricing subproblem. The predictions are integrated into the pricing process while ensuring correctness through standard reduced-cost checks.
Results
Experiments on cutting stock and generalized assignment problems indicate that predictive dual smoothing significantly reduces the number of generated columns and the wall-clock time required for convergence, outperforming both standard column generation and existing classical and learned stabilization methods. The improvements are consistent even for out-of-distribution instance sizes.
Implications
The findings suggest that predictive dual smoothing can be a powerful tool for optimizing large-scale linear programming problems, potentially applicable in various fields such as logistics, scheduling, and resource allocation where column generation is utilized.
HiTS-CL: A Continual Learning Framework for Long-Horizon Temporal Knowledge Graph Extrapolation
Graph Learning
Time Series
Theory
- Identifies limitations of fixed-prefix protocols in extrapolative TKGR.
- Reformulates TKGR as continual learning over streaming snapshots.
- Proposes HiTS-CL, which combines continual fine-tuning, adaptive distillation, and selective memory.
- Demonstrates consistent improvements in extrapolation accuracy across multiple datasets.
Read more
HiTS-CL: A Continual Learning Framework for Long-Horizon Temporal Knowledge Graph Extrapolation
Summary
The paper introduces HiTS-CL, a continual learning framework designed to enhance extrapolative temporal knowledge graph reasoning (TKGR). Traditional TKGR methods typically rely on a fixed-prefix training protocol, which trains models on early snapshots of data and applies them unchanged to future timestamps. This approach is inadequate for extrapolation as it fails to adapt to the evolving nature of temporal knowledge graphs, leading to performance degradation over time. The authors argue that effective extrapolation should be treated as continual learning over streaming snapshots, allowing models to dynamically update their knowledge in response to new information. HiTS-CL addresses this need by implementing a three-pronged approach: continual fine-tuning to capture current dynamics, multi-teacher adaptive distillation to preserve stable knowledge, and selective memory to retain recurring historical evidence. The framework is integrated into five different TKGR backbones and evaluated across four benchmark datasets, demonstrating significant improvements in extrapolation accuracy and a reduction in long-horizon degradation compared to existing methods. The results indicate that HiTS-CL outperforms strong continual-learning baselines, marking a substantial advancement in the field of temporal knowledge graph reasoning.
Methodology
HiTS-CL employs a model-agnostic approach that includes continual fine-tuning to adapt to current dynamics, multi-teacher adaptive distillation to maintain stable knowledge, and a selective memory mechanism to keep relevant historical evidence. This allows the framework to effectively manage the evolving nature of temporal knowledge graphs.
Results
The integration of HiTS-CL into five TKGR backbones resulted in consistent improvements in extrapolation accuracy and a notable reduction in long-horizon performance degradation. The framework outperformed strong continual-learning baselines and a recent method specifically adapted for temporal knowledge graphs.
Implications
The proposed framework has significant implications for applications requiring accurate long-term predictions from temporal knowledge graphs, such as temporal question answering, recommendation systems, and decision support systems. By enabling continual learning, HiTS-CL can adapt to new information and changing dynamics in real-time.
Correct then Forecast: Observer State-Space Models for Time Series Forecasting
Time Series
- OSSMs separate latent-state propagation from measurement assimilation, improving time series forecasting.
- The framework allows for a unified interpretation of existing state-space models and reveals modeling inconsistencies.
- OSSMs achieve substantial performance improvements while keeping the same parameter count as traditional models.
- The proposed method emphasizes the role of observations in correcting estimated latent states rather than controlling dynamics.
Read more
Correct then Forecast: Observer State-Space Models for Time Series Forecasting
Summary
This paper introduces Observer State-Space Models (OSSMs) for time series forecasting, addressing the limitations of traditional recurrent forecasting models that treat observations as direct inputs controlling latent dynamics. OSSMs conceptualize observed time series as measurements from an underlying autonomous dynamical system, separating latent-state propagation from measurement assimilation. This approach maintains a consistent transition governing dynamics across context and forecasting intervals, allowing observations to correct the estimated state via an observer mechanism. The authors demonstrate that OSSMs can recover conventional and recent state-space models (SSMs) as special cases, providing a unified framework that reveals inconsistencies in existing models. Experimental results across various benchmarks show that OSSMs significantly improve forecasting performance while maintaining the same parameter count and training setup as their SSM counterparts, supporting the principle that observations should correct the estimated latent state rather than control the dynamics.
Methodology
The authors propose OSSMs, which utilize a state-estimation perspective to treat observed inputs as measurements of an autonomous dynamical system. The model employs a single transition for dynamics across context and forecasting intervals, with an observer correcting the estimated state based on available observations. This formulation is grounded in control theory principles, ensuring observability and convergence of state estimation errors.
Results
The experimental evaluation shows that OSSMs outperform traditional state-space models in forecasting tasks across several benchmarks, demonstrating improved accuracy and consistency in predictions while maintaining the same model complexity.
Implications
The OSSM framework could enhance forecasting accuracy in various applications, including finance, weather prediction, and any domain reliant on time series analysis. By providing a more robust method for handling missing observations, OSSMs may lead to better decision-making processes in real-time systems.
Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training
Large Language Models
Reinforcement Learning
NLP
- The effectiveness of OPD is dependent on student-teacher compatibility rather than teacher scale alone.
- A brief SFT warm-up improves OPD performance, while RLVR-prepared students may regress under distillation.
- Adapting the teacher with RLVR enhances downstream OPD accuracy.
- OPD offers a better initialization for subsequent RLVR compared to SFT at similar starting accuracies.
Read more
Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training
Summary
This paper investigates the interactions between three key post-training stages for large language models (LLMs): supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD). The authors argue that these stages are often treated in isolation, which can lead to suboptimal performance. Through controlled experiments using Qwen3 models, they demonstrate that the effectiveness of OPD is influenced by the compatibility between student and teacher models, rather than just the scale of the teacher. The study reveals that a brief SFT warm-up can enhance OPD performance, while an RLVR-prepared student may regress under distillation from the same teacher. Additionally, adapting the teacher with RLVR significantly boosts OPD accuracy, and OPD provides a stronger initialization for subsequent RLVR compared to SFT. The findings emphasize the importance of considering the learning state inherited from previous stages when designing multi-stage post-training pipelines.
Methodology
The authors conducted controlled experiments using Qwen3 models, varying the states of the student and teacher models, and the order of training stages. They evaluated the models on math and science benchmarks while keeping data and optimization setups consistent to isolate the impact of each stage.
Results
The experiments revealed that increasing teacher scale does not consistently improve distillation effectiveness. A brief SFT warm-up was found to enhance distillability, while RLVR adaptation of the teacher improved OPD accuracy. OPD was shown to provide a more effective initialization for subsequent RLVR compared to SFT, with the gap widening as RL compute scales.
Implications
These findings suggest that practitioners should carefully consider the order and design of post-training stages in LLMs to maximize performance. The insights can guide the development of more effective multi-stage training pipelines, potentially leading to better reasoning capabilities in LLMs.
Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training
Reinforcement Learning
Large Language Models
Optimization
- Critic instability is an optimization artifact, not an inherent flaw.
- RFPO utilizes a single frozen critic for multiple roles, enhancing efficiency.
- Binarizing the critic's score mitigates length bias in policy optimization.
- RFPO achieves performance parity with supervised PPO while reducing resource requirements.
Read more
Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training
Summary
This paper addresses the challenges of reinforcement learning (RL) post-training for large language models (LLMs) by proposing a novel approach called Reward-Free Policy Optimization (RFPO). The authors argue that the common practice of discarding the critic after training leads to missed opportunities for leveraging its predictive capabilities. They demonstrate that the instability often associated with critic-based RL is primarily an artifact of optimization techniques, and that maintaining small, low-variance policy updates can stabilize convergence. The RFPO method repurposes a well-pretrained critic to provide dense learning signals without the need for completed rollouts or external rewards. By using the critic as a rollout-level reward, a generalized advantage estimation baseline, and a success forecaster, RFPO significantly reduces computational and memory overhead while maintaining performance. The results show that RFPO can match the performance of supervised PPO without requiring any labels during training, making it particularly effective for long-horizon reasoning tasks where outcomes are delayed.
Methodology
The authors introduce RFPO, which involves freezing and calibrating a pretrained critic to serve as the sole reward signal during policy optimization. They analyze the effects of policy update variance on critic stability and demonstrate the effectiveness of using the critic's predictions as dense learning signals without needing external labels. The methodology includes binarizing the critic's output to prevent length bias and validating the approach against traditional supervised PPO.
Results
RFPO matches the performance of supervised PPO while operating without any labels in the training loop. It demonstrates significant reductions in computational and memory overhead, particularly under conditions where rollouts are truncated, making it efficient for long-horizon reasoning tasks. The validation results indicate that RFPO can effectively score incomplete rollouts with accuracy comparable to complete ones.
Implications
The findings challenge the prevailing trend of critic-free training in RL for LLMs, suggesting that a well-trained critic can enhance the efficiency and effectiveness of post-training optimization. This has potential applications in various reasoning tasks where traditional reward mechanisms are costly or impractical.
Routing in Gradient Space: Balanced Usage Is Not Expert Specialization
NLP
Large Language Models
Optimization
- Introduction of Gradient-Aligned Routing (GAR) for optimizing expert assignments based on gradient coherence.
- GAR outperforms traditional routing methods in multi-task text classification, achieving higher accuracy and better expert load balance.
- The study distinguishes between expert load balance and gradient-based routing organization, emphasizing their separate impacts on model performance.
- The methodology leverages a load-normalized partitioning criterion that enhances the coherence of gradient contributions to experts.
Read more
Routing in Gradient Space: Balanced Usage Is Not Expert Specialization
Summary
This paper investigates the routing of gradients in sparse expert models, proposing a novel approach called Gradient-Aligned Routing (GAR). The authors argue that while balanced usage of experts can be achieved, it does not equate to expert specialization. GAR is designed to optimize the routing of observations based on aligned gradients, enhancing the coherence of expert assignments. The methodology involves a load-normalized partitioning criterion that rewards the grouping of observations with similar gradient directions. The authors conduct experiments on five multi-task text classification datasets, comparing GAR against various routing methods including task-loss-only routing and gradient-combination techniques. Results show that GAR consistently outperforms these methods in terms of aggregate validation accuracy, achieving improvements of 1.07 and 1.10 percentage points over task-loss-only routing with different backbone architectures. Additionally, GAR demonstrates better-balanced expert load and higher gradient-mass purity, indicating a more effective organization of gradient contributions. The findings highlight the importance of gradient-informed routing in multi-task learning, suggesting that expert load balance and gradient organization are distinct yet complementary aspects of model performance.
Methodology
The authors propose a gradient-partitioning approach to routing, introducing the Gradient-Aligned Routing (GAR) objective. This method scores the assignment of observations to experts based on the alignment of their gradients, using a load-normalized partitioning criterion. The routing is optimized through an auxiliary objective while maintaining the primary task loss for model training. The experiments utilize various backbone architectures and expert configurations to evaluate the effectiveness of GAR against other routing strategies.
Results
GAR achieved the highest aggregate validation accuracy across multiple experiments, outperforming task-loss-only routing by 1.07 percentage points with a fully trainable RoBERTa backbone and 1.10 points with frozen DeBERTa and Qwen3-1.7B backbones. Additionally, GAR demonstrated improved expert load balance and higher gradient-mass purity compared to other methods, indicating a more effective organization of gradient contributions.
Implications
The findings suggest that incorporating gradient-informed routing can enhance the performance of multi-task learning models, leading to more efficient expert utilization and improved accuracy. This approach may have broader applications in optimizing model architectures for various tasks in natural language processing and beyond.
In-Context Learning Amplifies a Latent Symbolic Circuit
NLP
Large Language Models
Interpretability
- The symbolic reasoning circuit is identifiable and functional before achieving high accuracy.
- Causal contributions from the circuit increase significantly with the number of in-context examples.
- Patching higher shot activations into lower shot prompts can enhance model accuracy substantially.
- Function vectors can effectively substitute for the induction stage in the reasoning process.
Read more
In-Context Learning Amplifies a Latent Symbolic Circuit
Summary
This paper investigates the mechanisms of in-context learning (ICL) in large language models, particularly focusing on how these models utilize a latent symbolic reasoning circuit comprising three stages: Symbolic Abstraction (SA), Symbolic Induction (SI), and Retrieval (Ret). The study traces the development of this circuit across varying shot counts (1 to 10 shots) in three model families: Gemma 2-2B, Llama 3.1-8B, and Qwen 3-4B. The findings reveal that the circuit is functional and identifiable even at low accuracy levels, with significant increases in causal contributions as the number of examples grows. Specifically, the per-head causal contribution can increase by up to 8 times from 1-shot to 10-shot. The paper also demonstrates that patching activations from higher shot counts into lower shot prompts can dramatically improve accuracy, indicating that in-context examples amplify the latent computation already present in the model's weights. The results suggest that interventions like function vectors can effectively substitute for certain stages of the reasoning process, highlighting the importance of the underlying circuit structure in ICL tasks.
Methodology
The methodology involves causal mediation analysis (CMA) to trace the symbolic reasoning circuit across different shot counts. The study uses two contrastive conditions to identify the roles of SA, SI, and Ret heads in the circuit. The performance of the models is evaluated based on their accuracy at varying shot counts, and the stability of the circuit topology is assessed through rank-biased overlap (RBO) analysis.
Results
The results indicate that the three-stage symbolic reasoning circuit is observable from just 1-shot across all three models studied. Accuracy improves significantly with increased shot counts, with Gemma 2-2B showing a rise from 17% at 1-shot to 99% by 10-shot. The causal contributions from the circuit heads grow up to 8 times as the number of examples increases, and patching activations from higher shot counts into lower shot prompts can raise accuracy from 1% to 56% at 0-shot and from 17% to 88% at 1-shot.
Implications
The findings have implications for the design and training of large language models, suggesting that enhancing the latent symbolic reasoning capabilities can lead to better performance in abstract reasoning tasks. The insights into the internal mechanisms of ICL may inform future research on model interpretability and the development of more effective interventions.
Seeing Time: Visual-Temporal Representation Learning for Interpretable Time Series Clustering
Time Series
Interpretability
Multimodal
- WAVE integrates temporal and visual representations for improved interpretability in time series clustering.
- The method employs cross-modal contrastive learning to align different representation modalities.
- WAVE achieves the highest macro-averaged clustering performance across multiple datasets.
- The approach allows for waveform-level traceability, enhancing the validation of clustering results.
Read more
Seeing Time: Visual-Temporal Representation Learning for Interpretable Time Series Clustering
Summary
This paper addresses the challenge of interpretability in multivariate time series (MTS) clustering by proposing WAVE (Waveform Aligned Visual-temporal Embedding). Traditional deep clustering methods often produce abstract clusters that are difficult to relate to observable waveform characteristics, limiting their practical utility. WAVE integrates fine-grained temporal variations with holistic visual patterns by treating time series and their corresponding waveform plots as complementary views. The method employs cross-modal contrastive learning to align these representations, ensuring that the resulting clusters can be traced back to authentic waveform records. Extensive evaluations on 10 real-world datasets demonstrate that WAVE achieves superior clustering performance and provides qualitative insights into the discovered clusters through waveform inspection. This approach enhances the interpretability of clustering results by linking them directly to observable temporal behaviors, thereby facilitating better validation and understanding of the patterns identified in the data.
Methodology
WAVE utilizes a dual representation approach where each MTS record is represented as both a sequence and a multichannel waveform plot. The method employs cross-modal contrastive learning to align these representations, followed by parameter-free normalized fusion to create a unified representation for clustering. This allows for the preservation of fine-grained temporal details while incorporating holistic waveform characteristics.
Results
WAVE demonstrated the highest macro-averaged clustering performance and the best average rank among competing methods across 10 public datasets. Qualitative evaluations showed that the clusters could be effectively inspected through authentic waveform records, confirming the interpretability of the clustering results.
Implications
The proposed method has significant implications for fields that rely on time series data, such as industrial monitoring, medical diagnosis, and human activity recognition. By enhancing the interpretability of clustering results, practitioners can better assess and validate the patterns discovered in their data, leading to more informed decision-making.
From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning
NLP
Large Language Models
Optimization
- Checkpoint selection is crucial in supervised fine-tuning, impacting model performance significantly.
- Increasing the validation budget improves selection accuracy, with notable gains observed in generated-accuracy and checkpoint-agreement methods.
- Alternative selection rules can outperform traditional validation-loss selection but do not consistently outperform the final checkpoint.
- The study highlights the need for distinct evidence to support claims regarding checkpoint selection performance.
Read more
From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning
Summary
This paper investigates the problem of checkpoint selection in supervised fine-tuning (SFT), where multiple checkpoints are generated during training but only one is retained for evaluation. The authors address three critical questions regarding checkpoint selection: the impact of increased validation data on selection performance, the effectiveness of alternative selection rules compared to traditional validation-loss selection, and whether these alternatives outperform simply retaining the final checkpoint. By treating checkpoint selection as a finite-information decision problem, the authors conduct experiments across 60 mathematical SFT trajectories and 19 configurations, varying the validation budget to measure the performance of selected checkpoints on independent test items. Their findings reveal that increasing the validation budget leads to significant improvements in independent-test accuracy, with generation-based selection rules outperforming matched negative log-likelihood (NLL) selection. However, the gains over the final checkpoint remain inconclusive. A cross-domain replication study further supports these results, indicating that the benefits of additional validation data, the superiority of alternative selection rules, and the performance relative to the final checkpoint are distinct empirical claims that require separate validation.
Methodology
The authors conducted controlled experiments by fixing completed training trajectories, candidate checkpoints, and independent test items while varying the validation budget. They measured the performance of selected checkpoints using generated validation accuracy, checkpoint agreement, and matched negative log-likelihood (NLL). The study included both primary comparisons and a matched-half analysis to assess the impact of selection on independent test items.
Results
The results showed that increasing the validation budget from 32 to 305-313 examples improved independent-test accuracy by +0.32 percentage points for generated-accuracy selection and +0.29 percentage points for checkpoint agreement. At the full validation budget, the generation-based rules outperformed matched NLL selection by +0.71 and +0.85 percentage points, respectively. However, the gains over the final checkpoint remained unresolved. A cross-domain replication on Commonsense trajectories confirmed similar qualitative results.
Implications
The findings suggest that careful consideration of checkpoint selection strategies and validation data usage can lead to improved model performance in supervised fine-tuning. This has practical implications for machine learning practitioners aiming to optimize model selection processes and enhance the reliability of their models.
TANGO: Watermarking Masked Diffusion Language Models in Token Pairs
NLP
Large Language Models
Generative Models
- TANGO introduces a coloring-based watermarking method for masked-diffusion language models.
- The watermark is embedded in pairs of tokens, reducing the risk of frequency-based attacks.
- Detection does not rely on the order of unmasking and requires only the text and a secret key.
- TANGO shows improved detection rates compared to traditional methods like the red-green list.
Read more
TANGO: Watermarking Masked Diffusion Language Models in Token Pairs
Summary
The paper introduces TANGO, a novel watermarking technique specifically designed for masked-diffusion language models, which operate by filling in masked positions in a non-sequential manner. Traditional watermarking methods, such as the context-hashed green list, rely on left-to-right token generation and can be easily compromised by attackers who analyze token frequencies. TANGO addresses these vulnerabilities by embedding watermarks in pairs of tokens, where each new token is biased towards a color class determined by a secret key and the color of a nearby unmasked token. This approach ensures that the token frequencies remain similar to those of unwatermarked text, making it difficult for attackers to forge text. The detection mechanism requires only the text and the key, without assumptions about unmasking order. The authors demonstrate that TANGO effectively detects nearly all unedited watermarked texts and a significant portion of edited texts, outperforming existing methods in terms of robustness against forgery and maintaining text quality.
Methodology
TANGO utilizes a secret key to color the vocabulary into multiple classes and biases the generation of new tokens based on the color of a nearby unmasked token. The detection process involves recoloring the candidate text with the key and counting the positions where a checksum condition holds, allowing for effective identification of watermarked texts.
Results
TANGO successfully detects nearly all unedited watermarked texts and a majority of edited ones. Its detection performance is comparable to that of the red-green list, while demonstrating greater resilience against forgery attempts. Empirical tests indicate that frequency attacks that succeed against traditional methods fail against TANGO.
Implications
The development of TANGO has significant implications for the security and integrity of generated text from masked-diffusion language models. It provides a robust mechanism for watermarking that can help content providers identify their outputs and prevent unauthorized use or forgery.
Adapting Linear-Time Architectures for Tabular In-Context Learning
Efficient ML
- Identified optimal training strategies for causal and non-causal models in tabular ICL.
- DeltaNet outperforms non-causal linear attention but struggles with longer contexts.
- Introduced a decay schedule to stabilize recurrent state and improve generalization.
- Proposed a final-state reading mechanism to enhance performance on benchmark datasets.
Read more
Adapting Linear-Time Architectures for Tabular In-Context Learning
Summary
This paper addresses the limitations of tabular foundation models that utilize softmax attention for in-context learning (ICL) on large datasets. The authors explore linear-time architectures, particularly focusing on causal models, and propose adaptations to improve their performance. They identify that the best training setup for causal models resembles next-token prediction and find that DeltaNet, a causal linear sequence mixer, outperforms non-causal alternatives. However, DeltaNet's performance degrades significantly beyond 2-4 times the pretraining context length due to instability in the recurrent state. The authors introduce a time-dependent decay schedule to stabilize this state and enhance length generalization. Additionally, they propose a final-state reading mechanism that allows the model to closely match the performance of a controlled softmax attention baseline on benchmark datasets OpenML-CC18 and TabArena. The study emphasizes the importance of training design, controlled comparisons of recurrence mechanisms, and the identification of failure mechanisms in achieving robust ICL for tabular data.
Methodology
The authors conducted a controlled study comparing various linear sequence mixers, including DeltaNet and its adaptations, under matched parameter counts and a unified training protocol. They analyzed the effects of different training designs and evaluated the models on OpenML-CC18 and TabArena datasets, focusing on their performance beyond pretraining context lengths.
Results
The study found that DeltaNet was the most effective causal model for tabular ICL, outperforming non-causal alternatives at moderate sequence lengths. However, all recurrent models showed degradation in performance beyond 2-4 times the pretraining context length. The introduction of a decay term for write rates in DeltaNet and a final-state reading mechanism significantly improved stability and performance, allowing the model to match softmax attention baselines.
Implications
The findings suggest that linear-time architectures can be effectively adapted for tabular in-context learning, potentially leading to more efficient models for large datasets. The proposed methods may enhance the applicability of foundation models in real-world tabular data scenarios, where computational efficiency is crucial.
When the Merge Coefficient Stops Mattering: Proximity Regularized Merging for Continual LoRA Adaptation
Efficient ML
Theory
- Proximity Regularized Merging (PRM) enhances rehearsal-free continual learning by adding a proximal penalty during task-vector training.
- The effectiveness of merging task vectors is influenced by their training conditions, not just the merge coefficient.
- PRM consistently improves performance across multiple write-in rules and architectures, achieving the best reported mean AAA.
- Mechanistic analyses show that PRM reduces Fisher-weighted interference and broadens the coefficient plateau.
Read more
When the Merge Coefficient Stops Mattering: Proximity Regularized Merging for Continual LoRA Adaptation
Summary
This paper addresses the challenge of rehearsal-free continual learning using low-rank adapters (LoRA) by introducing a novel method called Proximity Regularized Merging (PRM). The authors argue that the effectiveness of merging task vectors into a running model is not solely dependent on the merge coefficient but also on the training conditions that make the task vectors more 'mergeable'. PRM incorporates a proximal penalty during task-vector training, which stabilizes the merging process and reduces the reliance on precise coefficient selection. The authors demonstrate that PRM improves performance across various write-in rules and architectures, indicating that a well-trained task vector can lead to better integration into the model, thereby reducing forgetting and enhancing overall performance. The study reveals that PRM effectively shrinks the task-vector radius, lowers interference, and broadens the coefficient plateau, ultimately suggesting that training task vectors to be more mergeable can mitigate forgetting and simplify the merging process.
Methodology
The authors propose PRM, which modifies the training of LoRA task vectors by introducing a proximal penalty. This approach allows for robust continual merging without altering the downstream write-in rules. The methodology is evaluated across various existing merging techniques, including fixed-ฮฑ, CoFiMA, and MagMax, to assess its effectiveness in improving performance metrics.
Results
PRM achieved superior performance in terms of mean AAA across different settings compared to traditional methods. The results indicate that the method not only enhances the stability of the merging process but also reduces the impact of forgetting, demonstrating that a well-trained task vector can be integrated more effectively into the model.
Implications
The findings suggest that improving the mergeability of task vectors can lead to more efficient continual learning systems, particularly in applications where model updates are frequent and data privacy constraints prevent data rehearsal. This could have significant implications for deploying AI systems in dynamic environments where continual adaptation is necessary.
Neural Constitutive Learning for Generalized Reaction-Diffusion Systems
Theory
- Introduction of the NCL-MCT Solver for generalized reaction-diffusion systems.
- Separation of PDE-specific constitutive responses from shared temporal evolution.
- Support for trajectory-free constitutive learning and reuse across different conditions.
- Demonstrated effectiveness across seven systems with competitive error rates.
Read more
Neural Constitutive Learning for Generalized Reaction-Diffusion Systems
Summary
This paper introduces the Neural Constitutive LawsโMass-Compression-Transport (NCL-MCT) Solver, which addresses the challenge of learning constitutive responses for generalized reaction-diffusion systems. These systems describe the evolution of spatial densities through various transport mechanisms and reaction kinetics. The NCL-MCT Solver separates PDE-specific constitutive responses from shared temporal evolution, allowing for a unified interface that supports both velocity-data and known-law supervision. The constitutive responses are learned based on current density, enabling reuse across different initial conditions and time horizons. The authors evaluate the NCL-MCT Solver across seven systems, demonstrating its effectiveness in achieving low relative rollout L2 errors and its ability to generalize to unseen initial conditions and extended time horizons without retraining. This work highlights the potential of constitutive learning as a complementary approach to existing neural PDE methods, emphasizing the importance of a common interface for diverse transport and reaction mechanisms.
Methodology
The NCL-MCT Solver employs an energetic variational approach to define a common constitutive interface that represents transport through mobility and thermodynamic driving force, and reaction through relative reaction rates. It allows for training under velocity-data supervision or known-law supervision, facilitating trajectory-free learning from independently sampled density fields.
Results
The NCL-MCT Solver achieved relative rollout L2 errors ranging from 10^-4 to 10^-2 across seven evaluated systems. The learned constitutive modules demonstrated approximately 6 times lower error on unseen initial-condition families compared to traditional baselines, and maintained competitive performance over extended time horizons.
Implications
The findings suggest that constitutive learning can serve as an effective learning target within a shared neural PDE framework, potentially enhancing the modeling of complex physical phenomena in various fields such as biology, materials science, and fluid dynamics.
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
NLP
Large Language Models
Reinforcement Learning
- Introduction of Least-Square Policy Distillation (LSPD) for improved sample efficiency in language model reasoning.
- LSPD connects reverse-KL objectives in OPD with KL-regularized policy optimization, enabling off-policy data reuse.
- Empirical results demonstrate significant performance improvements over existing distillation baselines.
- LSPD maintains policy diversity and efficiency through exploration and historical data reuse.
Read more
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
Summary
This paper explores on-policy distillation (OPD) through a reinforcement learning (RL) perspective, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. The authors introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that enhances sample efficiency in language model reasoning by incorporating optimistic exploration and off-policy data reuse. LSPD maintains policy diversity through exploration while improving rollout efficiency by leveraging previously collected trajectories. The theoretical analysis demonstrates that LSPD achieves a regret bound of eO(log K) under online exploration. Empirical results show that LSPD outperforms existing distillation methods across six mathematical reasoning benchmarks, achieving an average improvement of +1.59 points in Avg@16. The framework also exhibits better policy diversity and performance as the number of evaluations increases, with its off-policy variant achieving comparable results to vanilla OPD using only a fraction of the rollout data. Overall, this work provides a principled and practical approach to enhancing language model distillation.
Methodology
The authors develop LSPD by combining robust quadratic matching of student and teacher log-probabilities with explicit entropy regularization. This allows for off-policy optimization and multiple updates per rollout batch. The framework is further enhanced with a replay-buffer variant (LSPD-RB) to facilitate effective data reuse.
Results
LSPD consistently outperforms existing distillation baselines across six benchmarks, achieving an average gain of +1.59 points in Avg@16. The off-policy variant, LSPD-RB, reaches saturated performance in approximately 10 rollout batches, compared to over 40 for traditional methods. Additionally, LSPD shows improved policy diversity and performance as evaluation rounds increase.
Implications
The findings suggest that integrating RL principles into OPD can lead to more efficient and effective language model distillation, potentially enhancing the capabilities of smaller models trained on complex reasoning tasks. This approach could be applied to various applications in natural language processing and machine learning.
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
Reinforcement Learning
Large Language Models
NLP
- ROSS preserves historical self-generated rollouts as a reusable training resource.
- The method selectively supervises informative segments of rollouts while maintaining full trajectory context.
- ROSS shows consistent performance improvements across multiple domains, including math and coding tasks.
- The approach highlights the importance of compatibility and complementarity in historical experiences for effective training.
Read more
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
Summary
The paper introduces ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), a novel approach that leverages historical self-generated rollouts from large language models (LLMs) to enhance their training. Traditional reinforcement learning methods often discard historical rollouts as training progresses, despite these rollouts containing valuable behaviors that may not be reliably expressed in later policies. ROSS addresses this by preserving the full historical trajectory while applying selective supervision to only the useful segments of these rollouts. The methodology involves evaluating the compatibility and complementarity of historical experiences with the current policy to identify which segments are worth relearning. The authors validate ROSS across various domains, including mathematics, code generation, instruction following, and software engineering, demonstrating significant improvements in performance metrics. The results indicate that ROSS not only strengthens the policy but also effectively consolidates reusable behavioral experiences from earlier training stages.
Methodology
ROSS employs a selective supervision mechanism that identifies valuable historical rollouts and focuses on supervising only the informative segments within successful trajectories. This is achieved by analyzing the compatibility of historical behaviors with the current policy and their complementarity in providing additional behavioral coverage.
Results
ROSS significantly improved performance metrics, achieving an increase from 58.40% to 62.20% on the six-benchmark MOPD average and from 64.20% to 68.40% on SWE-bench Verified. These results demonstrate the effectiveness of leveraging historical rollouts for further training gains without requiring new policy rollouts.
Implications
The findings suggest that historical self-generated experiences can be a valuable resource for training LLMs, enabling more efficient learning and better performance in various tasks. This approach could lead to advancements in the development of more robust AI systems that can adapt and improve over time.
Collaborative Synthetic Data for Privacy-Preserving Financial Fraud Detection Across Organizational Silos
Federated Learning
Generative Models
- CollaFuse enables collaborative synthetic data generation without direct data sharing.
- The method addresses the scarcity of informative minority-class observations in fraud detection.
- Synthetic data generated through CollaFuse improves fraud detection performance across multiple classifiers.
- The approach reduces computational burdens on individual organizations compared to traditional federated learning.
Read more
Collaborative Synthetic Data for Privacy-Preserving Financial Fraud Detection Across Organizational Silos
Summary
This paper addresses the challenge of financial fraud detection in scenarios where data is fragmented across organizations and privacy constraints limit data sharing. The authors propose a novel approach called CollaFuse, a collaborative diffusion-based method for generating synthetic data that preserves privacy while enhancing fraud detection capabilities. Unlike traditional federated learning methods that require significant local computational resources, CollaFuse allows organizations to collaboratively generate synthetic data without sharing raw data. The study evaluates CollaFuse against various benchmarks, including classical oversampling techniques and local generative models, across five different fraud datasets. The findings indicate that while CollaFuse may not achieve the highest local fidelity, it consistently improves downstream fraud detection performance across most datasets. This suggests that the value of synthetic data lies more in its ability to capture transferable cross-organizational patterns rather than strict adherence to local data distributions.
Methodology
The authors adapted collaborative diffusion-based generation techniques to create synthetic data for imbalanced tabular classification in fraud detection. The method involves multiple local clients retaining their raw data while a shared server facilitates the generation of synthetic data. The performance of CollaFuse is compared against various local oversampling methods and centralized benchmarks to evaluate its effectiveness.
Results
The results demonstrate that CollaFuse, while not achieving the highest local fidelity, significantly enhances fraud detection capabilities across various datasets. The synthetic data generated captures relevant cross-organizational patterns, proving valuable for downstream analytics even in privacy-constrained environments.
Implications
The findings suggest that organizations can collaborate effectively to enhance fraud detection capabilities without compromising data privacy. This approach could be applied in various sectors where data sharing is limited due to privacy concerns, enabling better analytics and decision-making.
Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents
Reinforcement Learning
Efficient ML
Optimization
- Introduces Reward-rate Policy Gradient (RPG) to optimize reward per unit of time in RL tasks.
- RPG uses off-policy samples to estimate reward rates, avoiding costly on-policy rollouts.
- Theoretical analysis shows RPG approximates optimal reward rates effectively.
- Empirical results demonstrate significant performance improvements over traditional RL methods.
Read more
Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents
Summary
This paper addresses a significant limitation in traditional reinforcement learning (RL) techniques, which typically focus on maximizing expected cumulative rewards under the assumption that actions take a constant unit of time. In contrast, machine learning engineering (MLE) tasks involve actions with variable durations, making efficiency a critical factor. The authors propose a novel approach called Reward-rate Policy Gradient (RPG), which optimizes the reward rateโdefined as the long-term reward per unit of timeโby adapting concepts from continuous-time RL and Semi-Markov Decision Processes (SMDP). RPG estimates the reward rate from off-policy samples and charges each action based on the time it consumes. Theoretical analysis in a bandit setting demonstrates that RPG can approximate the optimal reward rate while avoiding the computationally expensive enumeration over policy space required by existing methods. Empirical results show that RPG significantly outperforms vanilla RL approaches on MLE-Bench and NanoGPT, achieving improvements of 19.2% and 85.7%, respectively, within fixed time budgets. This work provides a practical framework for optimizing performance in agentic RL tasks where time and resource efficiency are paramount.
Methodology
The authors develop the RPG framework by leveraging off-policy historical samples to estimate the reward rate under a greedy policy. They conduct theoretical analysis in a bandit setting to establish the effectiveness of RPG and apply Proximal Policy Optimization (PPO) to test RPG on MLE tasks. The method focuses on maximizing relative rewards, which are calculated by adjusting total rewards based on the estimated reward rate.
Results
RPG outperformed traditional RL methods, achieving a 19.2% improvement on 20 out of 22 tasks in MLE-Bench and an 85.7% improvement on NanoGPT, all while adhering to fixed time budgets. The theoretical and empirical analyses confirm RPG's efficiency and effectiveness in optimizing reward rates.
Implications
The findings suggest that RPG can enhance the efficiency of RL agents in various applications, particularly in MLE tasks where time and resource management are critical. This approach could lead to more effective automation in machine learning pipelines and other agentic RL contexts.
On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models
Reinforcement Learning
Large Language Models
Optimization
- LRMs often exhibit a concentrated confidence prior that limits exploration in on-policy RL.
- The proposed CalibSFT method reshapes the confidence prior to enable better calibration.
- CalibSFT combines success rates with correctness to create balanced confidence targets.
- Extensive evaluations show that CalibSFT improves calibration and discrimination without sacrificing accuracy.
Read more
On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models
Summary
This paper addresses the issue of overconfidence in large reasoning models (LRMs) when expressing uncertainty. The authors highlight that confidence-aware reinforcement learning (RL) methods, while promising for calibration, are limited by the model's pre-RL confidence distribution, termed the confidence prior. They demonstrate that this prior is often concentrated on a few high values, which suppresses effective policy gradient updates and inflates the expected Brier risk. To mitigate this problem, the authors propose CalibSFT, a supervised fine-tuning approach that shapes a calibrated confidence prior with broad support before RL. CalibSFT constructs confidence targets that combine success rates with response-level correctness and balances training responses across the confidence spectrum. This method introduces correctness-conditional supervision to guide confidence across all responses while only supervising reasoning on correct ones. The authors evaluate CalibSFT across 16 benchmarks, showing significant reductions in calibration errors and improved discrimination across various RL algorithms, all while maintaining comparable accuracy. The findings suggest that CalibSFT can enhance downstream applications such as selective prediction and model routing.
Methodology
The authors introduce CalibSFT, a supervised fine-tuning stage that constructs proper-scoring confidence targets and balances them across the confidence spectrum. This method employs correctness-conditional supervision to guide confidence across all responses while supervising reasoning only on correct answers, allowing the model to learn low confidence on incorrect responses without imitating their reasoning.
Results
CalibSFT was evaluated on 16 mathematical and general reasoning benchmarks, demonstrating substantial improvements in calibration and discrimination across five representative RL algorithms. The method effectively mitigated confidence concentration and strengthened the correlation between task-level confidence and accuracy, while remaining effective across different confidence formats and model families.
Implications
The findings suggest that CalibSFT can significantly enhance the reliability of LRMs in expressing confidence, which is crucial for applications such as selective prediction and model routing. By improving calibration, the method can lead to more trustworthy AI systems capable of better uncertainty estimation.
MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models
Multimodal
Audio & Speech
Large Language Models
- Introduction of MISHAP-Bench, a benchmark specifically for evaluating hallucinations in LALMs.
- Definition of two categories of hallucination: context and knowledge.
- Development of a groundedness judge for evaluating open-ended responses.
- Significant hallucination rates observed in state-of-the-art LALMs, highlighting the need for better mitigation strategies.
Read more
MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models
Summary
The paper introduces MISHAP-Bench, a novel benchmark designed to evaluate hallucinations in Large Audio-Language Models (LALMs). Traditional benchmarks primarily assess correctness, which can obscure whether a model is hallucinating or simply misinterpreting audio inputs. The authors define two categories of hallucination: context hallucination, where claims are unsupported by the audio, and knowledge hallucination, where claims lack external verification. MISHAP-Bench consists of 12,000 open-ended question-audio pairs, facilitating a more nuanced evaluation of LALM responses. A groundedness judge, informed by human annotations, is employed to classify responses as grounded or hallucinated. The evaluation of ten state-of-the-art LALMs reveals a significant hallucination rate, with the Gemini 3.7 Flash model exhibiting a rate of 36.5%. The authors also adapt and test four mitigation strategies from various domains, finding that while some methods reduce hallucination, the challenge remains largely unaddressed. The paper calls for further research into hallucination evaluation and mitigation using MISHAP-Bench.
Methodology
The authors developed MISHAP-Bench, comprising 12,000 challenging open-ended question-audio pairs. They defined categories of hallucination and created a groundedness judge that uses human annotations to evaluate model responses. The evaluation involved testing ten state-of-the-art LALMs and benchmarking four existing mitigation methods.
Results
The evaluation revealed that hallucination is a substantial issue in LALMs, with the Gemini 3.7 Flash model achieving a hallucination rate of 36.5%. The tested mitigation methods showed some effectiveness but did not fully resolve the hallucination problem.
Implications
MISHAP-Bench provides a framework for assessing and improving the reliability of LALMs, encouraging further research into hallucination evaluation and mitigation. This could enhance the deployment of LALMs in real-world applications where accuracy and reliability are critical.
What You Observe Determines How You Identify Causal Effects: Evaluating Causal Models across Observational Views
Theory
- Introduction of CausalIDView, a benchmark for evaluating causal estimators across different observational views.
- No CFM consistently outperforms others; model performance varies significantly depending on the observational view.
- Modular estimators combining predictive models with explicit identification procedures can achieve competitive results.
- CFMs exhibit model-specific failures under structural changes, impacting the stability of causal effect estimates.
Read more
What You Observe Determines How You Identify Causal Effects: Evaluating Causal Models across Observational Views
Summary
This paper introduces CausalIDView, a multi-view benchmark designed to evaluate causal foundation models (CFMs) and other causal estimators by controlling the observational views available for causal identification. The authors argue that existing evaluations of CFMs are hindered by varying pre-training environments and evaluation protocols, making it difficult to assess their performance based on the information available for causal identification. CausalIDView addresses this by fixing the structural causal model (SCM) realization and target estimand while varying only the observational views. The study finds that no single CFM consistently outperforms others across different observational views, and model rankings can vary significantly. Additionally, the authors explore the effectiveness of combining predictive estimation with explicit identification procedures, demonstrating that a modular approach using a tabular foundation model can be competitive with CFMs. The findings highlight the importance of cross-regime comparisons in evaluating the empirical value of CFMs and suggest that model-specific failures can occur under structural changes, affecting the stability of causal estimates.
Methodology
The authors developed CausalIDView, a benchmark that holds the SCM realization and target conditional average treatment effect (CATE) fixed while varying the observational views. They evaluated various causal estimators, including CFMs and modular estimators, across different identification regimes, using metrics such as sPEHE and ATE Error. The study also included semi-synthetic benchmarks and null-effect tests to assess model performance.
Results
The results indicate that no CFM consistently leads in all regimes, with performance varying by dataset and observational view. Modular estimators using the TabPFN model demonstrated competitive performance, particularly in bound estimation under partial identification. CFMs showed specific failures in maintaining stable estimates when true effects remained unchanged and in tracking genuine effect changes.
Implications
The findings suggest that careful consideration of observational views is crucial for causal effect estimation. The effectiveness of modular approaches indicates potential for improved causal inference methodologies. This research could influence future developments in causal modeling and the design of benchmarks for evaluating causal estimators.
JudgeCast: Time Series Forecasting with Experience-Informed Covariate Judgements
Time Series
- JudgeCast introduces an experience-based framework for time series forecasting that utilizes covariate judgments.
- The framework separates the assessment of covariate effects from the numerical adjustment process.
- Residual-guided experience construction allows for the reconstruction of judgments based on forecast errors.
- JudgeCast outperforms traditional forecasting methods and strong baselines in real-world datasets.
Read more
JudgeCast: Time Series Forecasting with Experience-Informed Covariate Judgements
Summary
The paper introduces JudgeCast, an innovative framework for time series forecasting that leverages experience-informed covariate judgments to enhance forecasting accuracy. Traditional forecasting methods often struggle to adapt to the varying effects of covariates across different contexts and over time. JudgeCast addresses this challenge by employing a two-step process: first, it generates a base forecast using a frozen Time Series Foundation Model (TSFM), and then it adjusts this forecast using a frozen Large Language Model (LLM) that incorporates relevant contextual experience. The framework emphasizes the importance of forming explicit covariate-wise judgments that assess the expected effects of covariates, which are then used to determine numerical adjustments to the base forecast. After observing actual outcomes, JudgeCast reconstructs alternative judgments based on the forecast error (residual) and evaluates these against the original judgments. The best-performing judgment is retained as validated experience for future forecasts. The results demonstrate that JudgeCast significantly outperforms strong baseline models across various real-world datasets, achieving an average improvement of 16.7% in Mean Squared Error (MSE) and 6.5% in Mean Absolute Error (MAE). The findings highlight the effectiveness of explicit covariate-wise judgment and the advantages of residual-guided experience construction in enhancing forecasting performance.
Methodology
JudgeCast employs a two-step forecasting process: it first generates a base forecast using a frozen TSFM, and then adjusts this forecast using a frozen LLM that incorporates contextual information. The framework emphasizes forming explicit covariate-wise judgments to assess the expected effects of covariates, which are used to determine numerical adjustments. After observing actual outcomes, it reconstructs alternative judgments based on the forecast residual and evaluates them to retain the best-performing decision as validated experience.
Results
JudgeCast demonstrates a significant performance improvement over strong baselines, achieving an average reduction of 16.7% in MSE and 6.5% in MAE across diverse real-world datasets. Ablation studies indicate that the explicit covariate-wise judgments enhance forecast-time adjustments, while the residual-guided experience construction yields more reliable forecasting gains compared to retaining raw decisions.
Implications
The findings suggest that JudgeCast can be effectively applied in various domains requiring time series forecasting, such as healthcare, demand planning, and transportation. The framework's ability to adapt to changing covariate effects and improve forecasting accuracy could lead to better decision-making in these fields.
Rethinking Soft Tokens for Parallel Decoding in Diffusion Language Models
NLP
Generative Models
Large Language Models
- Soft tokens can mitigate inconsistencies in token generation during parallel decoding in DLMs.
- A geometry-aware construction of soft tokens is proposed for use with frozen pretrained models.
- Soft-token feedback enhances coherence in token sequences beyond just preserving uncertainty.
- The method shows superior performance compared to existing parallel decoding strategies.
Read more
Rethinking Soft Tokens for Parallel Decoding in Diffusion Language Models
Summary
This paper investigates the role of soft tokens in improving parallel decoding in diffusion language models (DLMs). DLMs enable the simultaneous generation of multiple tokens, but this can lead to inconsistencies among the generated tokens. Soft tokens, which represent uncertain positions with continuous embeddings, have been proposed to address this issue. However, the specific mechanisms by which soft-token feedback enhances parallel decoding have not been thoroughly examined. The authors propose a geometry-aware construction of soft tokens that can be applied to frozen pretrained DLMs without additional training. Their analysis reveals that soft-token feedback not only preserves predictive uncertainty but also favors coherent token sequences, thereby improving the quality of parallel decoding. The proposed method outperforms standard parallel decoding techniques and a training-free Euclidean soft-token baseline across various benchmarks.
Methodology
The authors develop a training-free, geometry-aware method for constructing soft tokens that addresses the geometric mismatch between conventional soft-token construction and the pretrained embedding space. They conduct experiments using four pretrained DLMs and evaluate their method against standard parallel decoding and a training-free Euclidean soft-token baseline.
Results
The proposed geometry-aware soft-token construction significantly outperforms both standard parallel decoding and the training-free Euclidean soft-token baseline across four math and code benchmarks. The analysis indicates that soft-token feedback improves the coherence of generated token sequences.
Implications
The findings suggest that soft tokens can be effectively utilized in DLMs to enhance the quality of parallel decoding, which could lead to more efficient and coherent text generation in various applications such as natural language processing tasks, code generation, and other generative modeling scenarios.
TokenCast: Forecasting Token Consumption During LLM Agent Execution
Large Language Models
NLP
Efficient ML
- TokenCast introduces a segment-cost factorization for accurate token consumption forecasting.
- The method updates predictions dynamically during execution without additional LLM calls.
- TokenCast achieves a 14.5% reduction in mean absolute error compared to existing methods.
- In budget-control scenarios, TokenCast uses 21.3% fewer tokens than fixed-budget policies.
Read more
TokenCast: Forecasting Token Consumption During LLM Agent Execution
Summary
The paper presents TokenCast, a novel approach to forecasting token consumption during the execution of large language model (LLM) agents. The authors identify that token consumption can vary significantly across different runs of the same task, complicating cost prediction for users. TokenCast addresses this issue by learning a composable cost representation for each execution segment, which records its own consumption and the context growth it introduces. This allows for cumulative estimates of token consumption that account for the re-reading of context from earlier segments in subsequent calls. The method updates predictions in real-time as execution unfolds, without requiring additional LLM calls, achieving a mean cumulative prediction time of 32.8 ms per run. The authors validate TokenCast across four task suites and six agent models, demonstrating a mean absolute error (MAE) reduction of 14.5% compared to the strongest existing methods. Additionally, in offline budget-control scenarios, TokenCast reduces token usage by 21.3% on average compared to fixed-budget policies, while still completing tasks successfully.
Methodology
TokenCast employs a compositional cost representation that factors in the token consumption of execution segments and the growth of context. It utilizes a staged prefix-suffix predictor that combines direct and compositional forecasts, updating its predictions based on observed execution evidence without querying the LLM for cost estimates.
Results
TokenCast demonstrates a mean absolute error reduction of 14.5% across 96 comparisons with existing methods. In budget-control replay scenarios, it achieves a 21.3% reduction in token consumption while maintaining task completion.
Implications
The findings suggest that TokenCast can significantly enhance the efficiency of LLM agents by providing more accurate token consumption forecasts, which can help users better manage costs and resources in applications involving complex task execution.
Sharp Convergence and Sampling Trade-offs for Riemannian Diffusion under Nonnegative Ricci Curvature
Generative Models
Theory
Efficient ML
- Developed a general framework for Riemannian diffusion models that separates score discretization and Brownian-motion simulation errors.
- Achieved a convergence rate of O(d/ฮตยฒ) for score evaluations under nonnegative Ricci curvature, matching Euclidean models.
- Showed that O(dโดT/ฮตยฒ) geodesic random-walk steps suffice for accurate Brownian motion approximation.
- Provided a strategy for performing multiple Brownian-motion simulation steps per score evaluation to optimize computational resources.
Read more
Sharp Convergence and Sampling Trade-offs for Riemannian Diffusion under Nonnegative Ricci Curvature
Summary
This paper addresses the convergence guarantees and sampling efficiency of Riemannian diffusion models, which are generative models extended from Euclidean spaces to Riemannian manifolds. The authors propose a novel framework that separates the errors from score discretization and Brownian-motion simulation, allowing for multiple geodesic random-walk steps per score evaluation. Under the assumption of nonnegative Ricci curvature, they demonstrate that the number of score evaluations required to achieve a specified KL divergence can be reduced to O(d/ฮตยฒ), matching the convergence rate of Euclidean diffusion models. Additionally, they establish that O(dโดT/ฮตยฒ) geodesic random-walk steps are sufficient for approximating the required drifted Brownian motion with a total variation error of ฮต. This framework not only improves existing convergence guarantees but also provides a theoretical basis for optimizing the trade-off between score evaluations and Brownian-motion simulation steps, suggesting that multiple simulation steps can enhance sampling efficiency.
Methodology
The authors developed a framework that decomposes sampling errors into components arising from score discretization and Brownian-motion simulation. They analyzed the convergence rates and sampling complexities under the assumption of nonnegative Ricci curvature and utilized geodesic random walks to approximate Riemannian Brownian motion.
Results
The paper presents sharper convergence bounds for Riemannian diffusion models, demonstrating that O(d/ฮตยฒ) score evaluations are sufficient for achieving an ฮตยฒ KL divergence from the target distribution. Additionally, it establishes that O(dโดT/ฮตยฒ) geodesic random-walk steps are adequate for approximating the required drifted Brownian motion with ฮต total variation error.
Implications
The findings have significant implications for the efficient implementation of Riemannian diffusion models in various applications, including those in scientific and engineering fields where data is structured on non-Euclidean manifolds. The proposed framework can guide practitioners in optimizing their sampling strategies, potentially leading to faster and more accurate generative modeling.
LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter
NLP
Large Language Models
Efficient ML
- Introduces the concept of fixed-cost drafting, which allows drafter costs to remain constant regardless of context length.
- LONGSPARK employs a block-diffusion approach to extract fixed-size, multiscale context views from the target model's verification pass.
- Achieves significant improvements in end-to-end throughput and reduces drafter memory overhead by several orders of magnitude.
- Demonstrates superior performance in long-context tasks, achieving the lowest time-per-output-token.
Read more
LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter
Summary
The paper introduces LONGSPARK, a novel block-diffusion drafter designed to enhance the efficiency of speculative decoding in autoregressive inference. Speculative decoding typically accelerates inference by allowing a lightweight drafter to propose multiple token candidates, which the target model then verifies. However, existing drafters face a significant challenge as their computational cost scales with the context length, undermining their efficiency. The authors argue that this scaling is unnecessary since the drafter only proposes candidates, while the target model is responsible for verifying and correcting any errors. LONGSPARK addresses this issue by employing fixed-cost drafting, which decouples the drafter's cost from the prefix length. It achieves this by extracting fixed-size, multiscale views from the target's verification pass, ensuring that the drafter's operations remain constant regardless of the context size. The paper presents extensive evaluations across various benchmarks, demonstrating that LONGSPARK achieves state-of-the-art efficiency, particularly in long-context tasks, while significantly reducing the drafter's memory overhead.
Methodology
LONGSPARK utilizes a block-diffusion drafting mechanism that extracts context from the target model's verification pass in fixed-size, multiscale views. This approach allows the drafter to maintain a constant operational cost, independent of the prefix length. The methodology includes a context interface for bounded-size views and a proposal model that generates competitive drafts based solely on these views.
Results
LONGSPARK outperforms existing state-of-the-art drafters in terms of end-to-end throughput, achieving speedups of 1.88ร to 2.13ร as model size increases. It also delivers the lowest mean time per output token across long-context benchmarks up to 128K tokens, while reducing the drafter's context state by 406ร compared to previous models.
Implications
The introduction of fixed-cost drafting has the potential to significantly enhance the efficiency of autoregressive models in various applications, particularly those requiring long-context processing, such as dialogue systems, coding assistants, and reasoning tasks. This could lead to faster inference times and reduced resource consumption in practical deployments.
AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models
Multimodal
Large Language Models
Optimization
- AdaKerNet operates on frozen MLLM representations, avoiding the need for parameter fine-tuning.
- The architecture combines learnable multimodal features, a reference kernel, and a nonlinear predictor for task adaptation.
- Empirical results show significant performance improvements over traditional decoding methods in limited supervision scenarios.
- The framework allows for joint optimization of kernel representation and prediction, enhancing adaptability to downstream tasks.
Read more
AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models
Summary
The paper introduces AdaKerNet, a novel task-adaptive neural kernel decoder designed to enhance the adaptation of multimodal large language models (MLLMs) to various downstream tasks, particularly under conditions of limited supervision. Traditional methods for adapting MLLMs typically involve fine-tuning model parameters or training neural decoders, both of which face challenges when supervision is scarce or when model parameters are inaccessible. AdaKerNet circumvents these issues by operating on frozen MLLM representations, utilizing a learnable set of multimodal features, a reference kernel for structural prior, and a lightweight nonlinear predictor that adapts the kernel structure. The architecture captures relevant features and geometric relationships for downstream tasks through a unified optimization framework that jointly learns kernel representation and prediction. Empirical evaluations across four MLLMs and various multimodal inputs demonstrate that AdaKerNet significantly outperforms existing decoding methods, achieving an average error reduction of up to 41% in continuous prediction tasks, particularly in scenarios with limited labels. The study highlights the complementary contributions of AdaKerNet's components, emphasizing its effectiveness in the scarce label regime.
Methodology
AdaKerNet employs a Lipschitz-controlled feature extractor to transform frozen MLLM embeddings into learnable features. It utilizes a reference kernel to provide a structural prior and incorporates a shallow MLP as a neural predictor that adaptively deforms the kernel structure based on task-specific supervision. The learning process involves joint optimization of kernel reconstruction and supervised prediction losses.
Results
Numerical tests across four MLLMs and various multimodal inputs indicate that AdaKerNet achieves an average error reduction of up to 41% compared to baseline methods, including MLPs, attention-based models, and autoencoders, particularly under limited label conditions. The results demonstrate the effectiveness of task-adaptive neural kernel decoding for continuous prediction tasks.
Implications
AdaKerNet's approach allows for effective adaptation of large multimodal models to diverse tasks without requiring access to model parameters, making it particularly useful for applications in scenarios with limited supervision. This could enhance the deployment of MLLMs in real-world applications where data is scarce or proprietary models are used.
Backpropagated Output Momentum: Relocating Optimizer History from Parameters to Task Space
Optimization
NLP
Computer Vision
- BOM replaces parameter-space momentum history with a compact task-space EMA, significantly reducing memory usage.
- The method preserves the current supervised gradient while reprojecting historical information through the current model.
- BOM achieves substantial improvements in validation performance across multiple tasks and models.
- The approach is compatible with existing adaptive optimizers, enhancing their efficiency without compromising performance.
Read more
Backpropagated Output Momentum: Relocating Optimizer History from Parameters to Task Space
Summary
This paper introduces Backpropagated Output Momentum (BOM), a novel approach to optimizer momentum that relocates the storage of historical gradient information from parameter space to task space. Traditional momentum methods, such as AdamW, maintain a moving average of past gradients tied to the parameters, leading to high memory costs and inefficiencies as the model parameters change. BOM addresses these issues by storing a compact moving average of prediction errors at the model output, which is then reprojected through the current network at each optimization step. This method reduces the size of the optimizer state by 49.7โ99.8% and decreases step time by 4.0% across various language models. The authors provide a detailed batch-level analysis of the information retained and omitted by this new approach, demonstrating its effectiveness in improving mean validation performance in language and vision tasks. The paper also discusses the implications of historical-projection drift and presents empirical results that validate the efficiency and performance gains of BOM compared to traditional methods.
Methodology
The authors develop BOM by creating a batch-shared output history for supervised tasks, which replaces the traditional dense first-order momentum state with a compact residual exponential moving average (EMA) of task-space errors. This EMA is reprojected through the current model's Jacobian at each optimization step, allowing for efficient updates while retaining the current supervised gradient. The paper includes a detailed batch-level decomposition to analyze historical-projection drift and its effects on optimization.
Results
BOM demonstrates a reduction in parameter-shaped optimizer state by 49.7โ99.8% across three compositions and achieves an average improvement of 1.42 points in mean validation performance across five tasks in language and vision fine-tuning. Additionally, it reduces step time by 4.0% when averaged over three language backbones. Pretraining studies further validate the efficiency of BOM, showing significant reductions in optimizer state and peak memory usage.
Implications
The introduction of BOM has the potential to enhance the efficiency of various momentum-based optimizers, making them more suitable for resource-constrained environments. Its ability to maintain performance while drastically reducing memory requirements could facilitate the training of larger models and more complex tasks in both NLP and computer vision.
Interactive-Policy Distillation with Bidirectional Propose-and-Verify
Large Language Models
NLP
Efficient ML
- Introduction of Interactive-Policy Distillation (IPD) to improve knowledge distillation.
- Bidirectional propose-and-verify mechanism allows for adaptive teacher intervention.
- IPD outperforms traditional on-policy distillation methods in accuracy and data efficiency.
- Fused inference engine designed to optimize dual-model rollouts.
Read more
Interactive-Policy Distillation with Bidirectional Propose-and-Verify
Summary
The paper introduces Interactive-Policy Distillation (IPD), a novel approach to knowledge distillation that addresses the limitations of on-policy distillation (OPD) by mitigating the issue of teacher unanchoring. In OPD, the student model learns from its own generated trajectories with feedback from a teacher model, but this can lead to unreliable supervision when the student deviates significantly from the teacher's reasoning. IPD employs a bidirectional propose-and-verify mechanism where the student and teacher alternate roles, collaboratively generating mixed-source trajectories. This allows for adaptive teacher intervention, ensuring that the student receives reliable feedback even when its reasoning diverges from the teacher's. The paper also presents a fused inference engine that enhances the efficiency of dual-model rollouts. Experimental results demonstrate that student models trained with IPD outperform those trained with OPD, achieving higher accuracy and requiring significantly fewer training examples. The findings suggest that IPD serves as a bridge between on-policy and off-policy paradigms, enhancing data efficiency and performance in training models for mathematical reasoning tasks.
Methodology
The methodology involves a bidirectional propose-and-verify state machine where the student and teacher alternate roles as proposer and verifier. This collaborative approach generates mixed-source trajectories, applying different supervision based on the source of each token. The paper also introduces a source-split loss to enhance the training process and a fused inference engine to improve the efficiency of the dual-model rollouts.
Results
The results indicate that models trained with IPD achieve a +3.28 mean@8 and best@8 benchmark-averaged accuracy improvement over those trained with OPD. Additionally, IPD requires only about 1/4 of the training examples and steps compared to OPD to achieve superior performance.
Implications
The findings suggest that IPD can significantly enhance the training of compact models derived from large language models, making them more suitable for latency-sensitive and high-throughput applications. This approach could be applied in various domains requiring efficient reasoning capabilities, such as mathematics and programming.
Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs
Federated Learning
- Introduces EN-HMTFL, a framework for multi-task learning in VANETs that allows vehicles to share a common encoder while retaining local decoders.
- Addresses the scalability and communication efficiency issues of traditional federated learning in dynamic vehicular environments.
- Demonstrates up to 24% improvement in accuracy and a reduction of up to 69 communication rounds for convergence compared to benchmarks.
- Utilizes a cluster-based hierarchical federated learning architecture to facilitate local model exchanges and reduce infrastructure load.
Read more
Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs
Summary
This paper addresses the limitations of traditional federated learning frameworks in vehicular ad hoc networks (VANETs), which typically assume that all vehicles collaborate to train a single model for a common task. This assumption is impractical in real-world scenarios where vehicles perform heterogeneous but related perception tasks. The authors propose a novel framework called encoder-sharing hierarchical multi-task federated learning (EN-HMTFL), which integrates cluster-based hierarchical federated learning with a globally shared encoder and vehicle-local decoders. In this framework, vehicles can collaboratively learn a transferable feature representation while maintaining their task-specific models locally. Only the encoder parameters are exchanged and aggregated through a hierarchical structure, while raw data and local decoder parameters remain on the vehicles. The proposed method is evaluated on the MNIST and GTSRB datasets under various vehicular scenarios, demonstrating significant improvements in accuracy and communication efficiency compared to existing benchmarks.
Methodology
The authors developed the EN-HMTFL framework, which employs a hierarchical federated learning structure where vehicles are organized into clusters. Each vehicle shares a common encoder while keeping its local decoder. The framework allows for efficient collaboration among vehicles performing different tasks by exchanging only the encoder parameters, thus minimizing communication overhead and preserving local data privacy.
Results
The evaluation of EN-HMTFL on the MNIST and GTSRB datasets showed that it outperformed existing methods, achieving up to 24% higher accuracy in various scenarios. Additionally, it demonstrated a significant reduction in the number of communication rounds required for convergence, with reductions of up to 69 rounds (28.8%) in certain scenarios.
Implications
The proposed framework has potential applications in enhancing the performance of machine learning tasks in VANETs, particularly in safety-critical applications such as traffic-sign recognition and object detection. By improving communication efficiency and model accuracy, EN-HMTFL can contribute to more reliable and effective vehicular networks.