AI-generated summaries
Today's ML research,
without the noise.
Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.
71
Papers today
8h
Update frequency
7
Days of history
Does Transolver really need a Transformer?
Theory
Efficient ML
- Transolver's performance is not dependent on the Transformer architecture.
- Slicing and deslicing operations are essential for maintaining model accuracy.
- A constant linear map can replace token self-attention without loss of performance.
- Theoretical insights confirm the universal approximation capability of the transformer-free Transolver.
Read more
Does Transolver really need a Transformer?
Summary
This paper investigates the necessity of the Transformer architecture in the Transolver model, a neural operator widely used for solving partial differential equations (PDEs) in fluid dynamics. The authors conduct a thorough empirical and theoretical analysis, performing controlled ablations on nine challenging 3D fluid dynamics benchmarks. They find that replacing the token self-attention mechanism with a constant linear map does not significantly impact accuracy, suggesting that the Transolver does not require a Transformer. However, they emphasize the critical role of the slicing and deslicing operations, which must be repeated at each layer to maintain performance. The paper also provides a mathematical explanation based on the theory of neural operators, demonstrating that the transformer-free Transolver can achieve universal approximation of continuous operators. Additionally, the authors introduce FlashSlice, an efficient implementation of the slicing/deslicing module that optimizes memory and computational resources without compromising accuracy.
Methodology
The authors performed controlled ablation studies on the Transolver model across nine 3D fluid dynamics benchmarks, systematically altering components such as token self-attention and slicing/deslicing operations. They also leveraged mathematical theories related to neural operators to provide a theoretical foundation for their empirical findings.
Results
The study revealed that replacing token self-attention with a constant linear map maintained accuracy across all benchmarks, while the removal of slicing/deslicing operations led to significant performance degradation. The necessity of repeating these operations at each layer was also confirmed. The FlashSlice implementation demonstrated substantial improvements in memory and computational efficiency.
Implications
The findings suggest that the Transolver model can be simplified by removing the Transformer component, leading to more efficient implementations in scientific machine learning applications. This could facilitate faster simulations in design optimization and uncertainty quantification tasks in engineering and physical sciences.
What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use
NLP
Large Language Models
Theory
- Identification of three operational bottlenecks in language model training: Circuit, Store, and Use.
- Introduction of a diagnosis-to-data principle that connects these bottlenecks with tailored supervision strategies.
- Demonstration that early circuit organization improves learning outcomes over extensive training.
- Distinction between writing content and invoking memory, highlighting the need for different supervisory approaches.
Read more
What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use
Summary
This paper explores the limitations of language models during training by identifying three distinct bottlenecks: Circuit formation, Store content availability, and Use route selection. The authors propose a diagnosis-to-data principle that connects these bottlenecks, emphasizing the need for different supervisory strategies as each limitation is addressed. They demonstrate that early circuit training can significantly enhance subsequent learning, retaining advantages even after processing large amounts of data (100B tokens). The study also highlights the importance of distinguishing between writing content and invoking memory, and how paired supervision and context ranking can improve decision-making in route selection. The findings suggest that as a limiting operation is repaired, the training targets must evolve, indicating that different types of supervision are necessary for effective learning. The paper concludes with a call for further exploration of conditional arbitration as a remaining challenge in model training.
Methodology
The authors conducted a series of experiments involving continuous training on a 350M model, testing various interventions related to Circuit formation, Store content availability, and Use route selection. They employed a shared diagnosis-to-data principle to identify and address the specific limitations encountered during training, using controlled supervision and context variation to enhance model performance.
Results
The results indicated that early circuit training significantly improved the model's ability to learn and retain information, outperforming stage-replacement controls on facts withheld from Use teaching. The study also found that independent query surfaces and opposed-source decisions revealed conditional arbitration as a critical area for further research.
Implications
The findings suggest that training strategies for language models should be adaptive, changing in response to the specific bottlenecks encountered. This could lead to more efficient training processes and improved model performance in real-world applications, particularly in tasks requiring complex reasoning and memory utilization.
AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework
Large Language Models
Optimization
Efficient ML
- Identification of five failure modes in production-scale AutoResearch: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation.
- Development of a three-principle scaffolding design to address the identified failure modes.
- Significant performance improvements over hand-tuned baselines, including a 1.82× lift in Recall@6 and a 2.1× coherence lift.
- Autonomous design of a fallback mechanism that increased catalog coverage by 5.8×.
Read more
AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework
Summary
This paper explores the application of the AutoResearch paradigm, which utilizes a large language model (LLM) to automate the exploration of embedding systems for production recommendation pipelines. The authors conducted a twelve-week study at Amazon's book recommendation pipeline, running over 220 experiments across two independently developed representation-learning systems. They identified five recurring failure modes that emerged at production scale: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation. To address these issues, the authors propose a three-principle scaffolding design—prevent, persist, redirect—that provides structural remedies for each failure mode. The framework demonstrated significant improvements, achieving a 1.82× lift in Recall@6 and a 2.1× lift in coherence over hand-tuned baselines. Additionally, the agent autonomously designed a text-only fallback that expanded catalog coverage by 5.8×. The findings suggest that these failure modes are inherent to production-scale autonomous research rather than specific to the applications tested.
Methodology
The study employed the AutoResearch paradigm, where a large language model iteratively modifies training scripts based on performance metrics. The authors ran extensive experiments in a production environment, analyzing the performance of two distinct representation-learning systems and documenting the failure modes encountered.
Results
The implementation of the AutoResearch framework led to a 1.82× improvement in Recall@6 and a 2.1× increase in coherence over hand-tuned baselines. The autonomous agent also created a fallback mechanism that expanded catalog coverage by 5.8×, demonstrating the effectiveness of the proposed scaffolding design.
Implications
The findings highlight the challenges and potential solutions for deploying autonomous machine learning research at scale, suggesting that the proposed framework can enhance the efficiency and effectiveness of recommendation systems in production environments.
Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model
Efficient ML
NLP
Theory
- The audit evaluates a commercial System-1 model's reliability on biosecurity-relevant benchmarks.
- The model demonstrates strong calibration but variable accuracy depending on the task.
- Significant instability in item-level decisions is observed due to answer option order sensitivity.
- Selective averaging of probabilities can enhance accuracy without incurring high computational costs.
Read more
Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model
Summary
This paper presents a reliability audit of a commercial non-generative 'System-1' model, specifically evaluating its performance on biosecurity-relevant benchmarks. The authors assess the model's accuracy, calibration, error detection, selective prediction, and sensitivity to the order of answer options across 6,020 multiple-choice items from the Weapons of Mass Destruction Proxy (WMDP) and LAB-Bench subtasks. The findings indicate that while the model is reasonably well-calibrated overall, its accuracy is highly task-dependent and significantly degrades on weaker tasks. The study reveals that 37.4% of items in the WMDP-Cyber benchmark yield different answers when answer options are cyclically rotated, suggesting substantial instability in item-level decisions. To mitigate this instability, the authors propose a method of averaging probabilities across rotations, which improves accuracy, particularly when applied selectively to low-confidence items. The paper emphasizes the importance of understanding the empirical properties of low-cost probabilistic models in the context of AI safety and biosecurity applications.
Methodology
The authors conducted a reliability audit of the System-1 model by querying it on a set of 6,020 multiple-choice items. They measured various performance metrics, including accuracy, calibration, and error detection. A controlled study was performed to assess the impact of answer option order on model decisions, using cyclic rotations of answer options and a byte-identical repeat control to distinguish between option-order sensitivity and run-to-run variation. Additionally, they explored the effectiveness of selective permutation averaging to improve accuracy on low-confidence items.
Results
The model's pooled expected calibration error was 0.034, and the area under the receiver operating characteristic curve (AUROC) for top-1 probability was 0.820. However, accuracy varied significantly across tasks, with a notable 37.4% of WMDP-Cyber items yielding different answers based on answer option order. The proposed method of averaging probabilities across rotations improved accuracy by 3.8 percentage points, particularly when applied selectively to low-confidence items.
Implications
The findings suggest that while non-generative System-1 models can be cost-effective components in AI safety and biosecurity pipelines, their reliability must be thoroughly evaluated. The study underscores the importance of understanding model behavior under different conditions, which can inform the deployment of such models in critical applications.
Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection
Time Series
- Introduces a Bayesian Tensor Autoencoder framework for anomaly detection in multi-dimensional time series.
- Incorporates a predictive prior that bridges the gap between reconstruction-based and prediction-based autoencoders.
- Utilizes physical laws to enhance the modeling capability and mitigate over-generalization.
- Demonstrates effectiveness through experiments on real-world datasets.
Read more
Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection
Summary
This paper addresses the challenge of anomaly detection in multi-dimensional time series data, which are inherently tensorial and contain intrinsic correlations that are often disrupted by traditional reshaping methods. The authors propose a novel framework called the Physics-informed Predictive Prior Tensor Autoencoder (PPPTAE) that integrates a predictive prior into a reconstruction-based autoencoder. This approach leverages both current observations and historical data while preserving the tensor structure of the data, thus enhancing the model's ability to detect anomalies. The predictive prior is designed using Bayesian fusion and incorporates physical laws related to tensor low-rank decomposition, which helps mitigate over-generalization issues. The proposed PPPTAE framework is tailored with specific training and testing strategies to introduce randomness during training and simplify testing without complex density estimation. Experimental results on real-world datasets demonstrate the effectiveness of the PPPTAE in accurately detecting anomalies while maintaining the integrity of the multi-dimensional time series data.
Methodology
The authors developed the PPPTAE framework by integrating a predictive prior into a reconstruction-based autoencoder. This involved using Bayesian fusion to enhance the model's capability for normal data and incorporating tensor low-rank decomposition rules to maintain the integrity of the data structure. The training and testing strategies were specifically designed to introduce randomness during training and simplify testing processes.
Results
The experimental results indicate that the PPPTAE framework significantly outperforms traditional methods in detecting anomalies in multi-dimensional time series data, demonstrating its ability to effectively leverage intrinsic correlations and maintain data structure.
Implications
The proposed method has potential applications in various fields where multi-dimensional time series data is prevalent, such as traffic monitoring, weather forecasting, and other domains requiring timely anomaly detection to prevent losses and enhance safety.
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
NLP
Large Language Models
Reinforcement Learning
- Reinforcement learning enhances reasoning in large language models but its internal mechanisms are complex.
- Vector steering reveals a low-dimensional effective manifold in activation space associated with RL performance gains.
- Two geometric properties of this manifold are identified: Effective Manifold Capacity and Control Manifold Separation.
- Alpha-Stabler framework stabilizes RL training and improves performance by managing activation gradients.
Read more
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
Summary
This paper investigates the role of reinforcement learning with verifiable rewards (RLVR) in enhancing the reasoning capabilities of large language models (LLMs). The authors introduce vector steering as a tool to analyze the high-dimensional parameter updates associated with RL training, revealing a low-dimensional effective manifold in activation space that correlates with performance gains. They identify two key geometric properties of this manifold: Effective Manifold Capacity, which indicates that while a small capacity can reproduce RL gains, it is not infinitely compressible, and Control Manifold Separation, which shows that effective control directions are primarily found in the low-variance complement of the activation principal subspace. The authors conduct experiments on five LLMs across six tasks, demonstrating the validity of these properties. They propose a training framework called Alpha-Stabler, which includes a Predictor to monitor principal-subspace intrusion and a Controller to adjust activation gradients during backpropagation. The results indicate that Alpha-Stabler stabilizes training and enhances RL-induced gains, contributing to a deeper understanding of RL dynamics in LLMs and offering practical insights for robust post-training.
Methodology
The authors employ vector steering to analyze the activation space of LLMs, using input-invariant shared vectors to distill RL-tuned models. They characterize the effective manifold through nonlinear analysis and conduct experiments on multiple LLMs and tasks to validate their findings. The Alpha-Stabler framework is proposed to monitor and control training dynamics.
Results
Experiments show that a single input-invariant vector can recover over 85% of RL gains in most tasks, indicating the effective capacity of the manifold. The study also finds that control directions are predominantly in the low-variance complement of the principal subspace, and Alpha-Stabler successfully stabilizes training for 2,000 steps while enhancing RL gains.
Implications
The insights from this study could lead to more effective training strategies for large language models, improving their reasoning capabilities and robustness. The Alpha-Stabler framework may be applicable in various RL scenarios to enhance training stability and performance.
On the Capability and Limitation of Hard Prompt
NLP
Large Language Models
Theory
- Determining the existence of a hard prompt is NP-complete; finding an optimal hard prompt is NP-hard.
- Hard prompts have fundamental limitations, including incompleteness and performance issues with short and long prompts.
- Linear hard prompts can enhance transformer performance without the drawbacks of traditional hard prompts.
- A necessary and sufficient condition for the generalizability of prompts is established based on prompt length and task size.
Read more
On the Capability and Limitation of Hard Prompt
Summary
This paper investigates the theoretical foundations of hard prompts in the context of large language models (LLMs). While prompt engineering has become essential for adapting LLMs to specific tasks, the theoretical understanding of hard prompts—discrete instructions composed of natural language tokens—remains limited. The authors address three core theoretical questions regarding hard prompts. First, they establish that determining the existence of a hard prompt for a transformer to solve a downstream task is NP-complete, and finding an optimal hard prompt is NP-hard, marking a significant contribution to the computational complexity of hard prompting. Second, the paper outlines the limitations of hard prompts, demonstrating that they are not complete and that short prompts do not significantly enhance transformer performance. Additionally, long prompts can lead to the 'prompt dominating answer phenomenon,' where identical answers are produced for queries of the same length. In contrast, linear hard prompts do not exhibit these limitations and can improve performance. Finally, the authors provide a tight bound on the relationship between prompt length and task size, offering a necessary and sufficient condition for generalizability of prompts across data distributions. These findings provide both theoretical insights and practical guidance for the effective use of hard prompts in real-world applications.
Methodology
The authors employ theoretical analysis to explore the computational complexity of hard prompts, establishing the NP-completeness and NP-hardness of related problems. They also analyze the performance limitations of hard prompts through formal theorems and propositions, comparing them with soft prompts and providing bounds for generalization.
Results
The study reveals that hard prompts are fundamentally limited in their ability to enhance transformer performance, particularly in terms of completeness and effectiveness based on prompt length. Linear prompts are identified as a viable alternative that avoids the pitfalls of traditional hard prompts. The paper also establishes a theoretical framework for understanding the generalizability of prompts.
Implications
These findings have significant implications for the design and application of prompts in LLMs, guiding practitioners in selecting appropriate prompt types for various tasks. The theoretical insights can inform future research in prompt engineering and the development of more effective prompting strategies.
Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design
Robotics
Optimization
Reinforcement Learning
- Introduction of Morphogene as a high-level latent blueprint for body-brain coordination.
- GeCode reformulates morphology-control co-design as gene-driven exploration in a compact latent space.
- Demonstrated significant performance improvements over existing methods in diverse design tasks.
- Achieved an average of 2.5× faster convergence and 69.48% higher task performance.
Read more
Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design
Summary
This paper presents a novel approach to morphology-control co-design in embodied agents, addressing the limitations of existing methods that treat morphology and control as separate entities. The authors introduce 'Morphogene', a compact latent blueprint inspired by biological genes, which facilitates explicit coordination between an agent's body structure and control policy. By employing 'AdaConcat', Morphogene conditions both morphology and control generation at the limb level, allowing variations in Morphogene to induce coordinated changes in both components. The proposed framework, 'GeCode', reformulates the co-design process as exploration within the Morphogene space, where local design regions are anchored by Morphogenes. Performance-guided updates enable efficient exploration of the design space, combining local refinement with global exploration while ensuring compatibility between body and brain. Extensive experiments across various 2D and 3D tasks demonstrate that GeCode significantly outperforms state-of-the-art methods, achieving faster convergence and higher performance metrics.
Methodology
The methodology involves the introduction of Morphogene to jointly condition morphology and control generation. The GeCode framework utilizes a small set of Morphogenes as design anchors, combining stochastic policy rollouts with a Morphogene-proximity reward to explore high-performing designs. Performance-guided updates adjust the anchors towards promising regions in the design space, facilitating efficient exploration.
Results
GeCode consistently outperformed state-of-the-art methods across twelve 2D and 3D design spaces, achieving an average of 2.5 times faster convergence and 69.48% higher task performance, with minimal additional computational overhead.
Implications
The findings suggest that integrating morphology and control through a shared latent representation can enhance the design of embodied agents, potentially leading to more efficient and adaptable robotic systems. This approach may have applications in robotics, artificial intelligence, and bio-inspired design.
HyperLabel: Multi-Label Classification via Hypergraph-Based Label Correlation Modeling
Graph Learning
- Introduction of HyperLabel, an encoder-decoder framework for multi-label classification.
- Construction of a label hypergraph to explicitly model multi-way label dependencies.
- Implementation of HGNN+ for bidirectional message passing between features and labels.
- Demonstration of state-of-the-art performance on multiple benchmark datasets.
Read more
HyperLabel: Multi-Label Classification via Hypergraph-Based Label Correlation Modeling
Summary
The paper presents HyperLabel, a novel framework for multi-label classification (MLC) that addresses the challenge of modeling complex label dependencies through hypergraph neural networks (HGNNs). Traditional MLC methods often struggle to capture high-order label correlations, relying on pairwise interactions or implicit learning mechanisms. HyperLabel constructs a label hypergraph where samples define hyperedges, allowing for the representation of multi-way co-occurrence patterns among labels. The framework includes an HGNN+ label encoder that performs bidirectional message passing to integrate feature information with label structures, enhancing the learning of label dependencies. A shared cross-attention decoder processes both feature and label embeddings, facilitating cross-modal learning through complementary objectives. Extensive experiments on seven benchmark datasets demonstrate that HyperLabel achieves state-of-the-art performance, particularly improving macro-F1 scores significantly on datasets like Delicious and Bibtex, validating the effectiveness of the hypergraph structure in capturing complex label relationships.
Methodology
HyperLabel employs an encoder-decoder architecture where a hypergraph is constructed to represent label dependencies. The HGNN+ serves as the label encoder, performing bidirectional message passing to integrate feature information with label representations. A shared cross-attention decoder processes both feature and label embeddings, supported by multiple complementary learning objectives to enhance feature-label and label-label interactions.
Results
HyperLabel achieved state-of-the-art performance across seven benchmark datasets, with notable improvements in macro-F1 scores: +10.3% on Delicious and +8.2% on Bibtex, demonstrating the efficacy of the hypergraph-based approach in capturing complex label relationships.
Implications
The proposed HyperLabel framework can be applied in various domains requiring multi-label classification, such as text categorization, image understanding, and bioinformatics, where understanding complex label dependencies is crucial for accurate predictions.
Predictive Dual Smoothing for Column Generation
Optimization
- Introduction of predictive dual smoothing to enhance column generation efficiency.
- Utilization of learned predictions of future duals to guide pricing subproblem.
- Demonstrated significant reductions in generated columns and runtime in experiments.
- Method maintains correctness of the column generation process.
Read more
Predictive Dual Smoothing for Column Generation
Summary
This paper addresses the challenge of efficiently solving large-scale linear programs through column generation (CG), a method that alternates between solving a restricted master problem and identifying new variables via a pricing subproblem. The authors introduce 'predictive dual smoothing', a novel technique that enhances the traditional dual smoothing method by incorporating predictions of future dual solutions to guide the pricing subproblem. This approach aims to mitigate dual oscillations that can slow convergence by steering the search towards more useful variables in subsequent iterations. The predictor is trained offline using supervised learning from standard CG trajectories, allowing it to generate predictions that are used to modify the pricing subproblem's objective function. The methodology preserves the correctness of the CG process by ensuring that generated columns are verified against the current solution. The experimental results demonstrate that predictive dual smoothing significantly reduces the number of generated columns and the overall computation time compared to standard CG and existing stabilization methods, with benefits extending to out-of-distribution instance sizes. Furthermore, the technique shows improved performance when combined with classical stabilization methods.
Methodology
The authors propose predictive dual smoothing, which combines the current dual solution with predictions of future duals obtained through supervised learning from past CG trajectories. The predictions serve as a reference point for pricing, allowing the method to anticipate future needs while maintaining the integrity of the CG process. The predictor is trained offline, leveraging multiple training pairs derived from CG trajectories.
Results
Experiments on cutting stock and generalized assignment problems reveal that predictive dual smoothing leads to a substantial decrease in the number of columns generated and a reduction in wall-clock time compared to standard column generation and existing stabilization techniques. The improvements are consistent across different problem sizes, including out-of-distribution instances.
Implications
The proposed method has the potential to enhance the efficiency of solving large-scale linear programs in various optimization settings, including logistics, scheduling, and resource allocation problems. By improving convergence rates, predictive dual smoothing could facilitate faster decision-making in practical applications.
Trust Guided Decision Transformer
Reinforcement Learning
Robotics
Optimization
- Identifies rollout context mismatch as a critical failure mode in Decision Transformers.
- Introduces a novel context selection mechanism that filters based on prediction error reliability.
- Demonstrates that training-side improvements alone are insufficient for reliable performance.
- TGDT outperforms existing methods in reducing prediction error and improving returns in various tasks.
Read more
Trust Guided Decision Transformer
Summary
The paper introduces the Trust Guided Decision Transformer (TGDT), addressing the performance degradation of Decision Transformers (DT) during long rollouts due to context drift from the training distribution. The authors identify a failure mode termed 'rollout context mismatch,' where the model's next state prediction error increases during rollouts, indicating unreliable context. TGDT evaluates recent context suffixes based on their prediction error, rejecting those exceeding a calibrated threshold derived from offline data. This process allows TGDT to select only trusted context before applying value guidance from a frozen critic. Experimental results on D4RL navigation and locomotion tasks demonstrate that TGDT significantly reduces high error runs and improves returns compared to vanilla DT and other context selection methods. The key contributions include the identification of rollout context mismatch, the introduction of a trust-filtered context selection rule, and the demonstration of TGDT's effectiveness in improving decision-making under uncertainty.
Methodology
TGDT employs a rolling next-state prediction error to evaluate the reliability of context suffixes during decision-making. It maintains a threshold derived from offline data to filter out unreliable contexts before ranking actions based on a frozen critic's guidance. This approach contrasts with traditional methods that prioritize action value without assessing context reliability.
Results
Experiments on D4RL Maze2D, AntMaze, and MuJoCo locomotion tasks show that TGDT reduces the duration of high prediction error runs by approximately 8 times and improves normalized return compared to vanilla Decision Transformer and other context selection strategies.
Implications
The findings suggest that ensuring context reliability is crucial for effective decision-making in reinforcement learning settings. TGDT's approach could be applied to enhance the performance of various RL algorithms, particularly in environments where context drift is a concern.
Disentangling Lung-Cancer CT/LDCT AI: A Systematic Evidence Map of Clinical Tasks, Evidence Chains, and Translational Gaps
Multimodal
- The literature on lung cancer AI is predominantly focused on detection rather than risk prediction.
- Only 10.6% of studies completed a comprehensive six-tier evidence chain, indicating significant attrition in the translational pathway.
- The study introduces a Multi-Tier Evidence Graph (MTEG) to systematically analyze and visualize evidence chains in lung cancer AI research.
- Clinical-task heterogeneity and inconsistent definitions of multimodal evidence are major challenges in synthesizing the literature.
Read more
Disentangling Lung-Cancer CT/LDCT AI: A Systematic Evidence Map of Clinical Tasks, Evidence Chains, and Translational Gaps
Summary
This paper presents a systematic review and evidence mapping of artificial intelligence (AI) applications in lung cancer diagnosis using computed tomography (CT) and low-dose CT (LDCT). The study identifies and categorizes 293 relevant studies published between 2016 and 2026, focusing on distinct clinical tasks such as detection and risk prediction. The authors highlight significant challenges in the literature, including the ambiguity of clinical tasks labeled as 'prediction,' inconsistencies in multimodal evidence interpretation, and the lack of coherent translational pathways in existing studies. A Multi-Tier Evidence Graph (MTEG) was developed to analyze the evidence chains, revealing that most studies are detection-focused, with genuine future risk prediction being rare. The findings indicate that while there is a rich body of algorithmic research, the clinical applicability and translational maturity of these models remain limited due to gaps in validation and reasoning processes.
Methodology
The authors employed a systematic evidence mapping strategy, which included five-database retrieval, full-text eligibility assessment, role-aware modality and omics extraction, and clinical-task classification. A Multi-Tier Evidence Graph (MTEG) was constructed to analyze the evidence chains across studies, categorizing components into seven functional tiers.
Results
The analysis revealed that out of 293 studies, 230 focused on detection, 8 on future risk prediction, and 55 on other tasks. Key findings included that clinical variables, 3-D CT/LDCT, and radiomics were frequently used, while external validation and calibration were less common. The MTEG comprised 377 nodes and 3,444 edges, highlighting the fragmented nature of the literature and the rarity of complete translational evidence chains.
Implications
The findings underscore the need for clearer definitions and classifications in lung cancer AI research to improve the synthesis and interpretation of studies. The MTEG framework can serve as a valuable tool for researchers to assess the translational maturity of AI models and guide future research towards more clinically applicable outcomes.
Beyond the Graph: An Adaptive Meta-Learner Fuses Explainability, Weather, and Dynamics for Robust Bus ETA Prediction
Graph Learning
Time Series
Interpretability
- Introduction of HYB(nm), a hybrid ensemble framework for bus ETA prediction.
- Integration of multiple models to address nonlinear dynamics and weather effects.
- Evaluation on extensive real-world data showing significant improvements in prediction accuracy.
- Flexible architecture allows for tailored solutions for transit agencies.
Read more
Beyond the Graph: An Adaptive Meta-Learner Fuses Explainability, Weather, and Dynamics for Robust Bus ETA Prediction
Summary
This paper addresses the challenge of accurately predicting bus Estimated Time of Arrival (ETA) in urban settings, particularly in Kolkata, where existing models struggle with nonlinear spatiotemporal dynamics and external factors such as weather. The authors propose HYB(nm), an adaptive hybrid ensemble framework that integrates five complementary models: a historical baseline (MST-AV), periodical temporal pattern analysis (GDRN-DFT), Koopman Neural Operators for nonlinear dynamics (KOOP-NET), weather-integrated feature-engineered neural networks (FENN), and real-time graph convolutional networks (MGCN). This framework is designed to dynamically fuse these models through a meta-learner that adapts to real-time contexts. The evaluation of HYB(nm) on GPS and weather data from over 4,000 bus trips across three routes in Kolkata demonstrates its superior robustness and accuracy, achieving state-of-the-art performance that rivals leading graph neural networks. The architecture allows for flexible deployment options, catering to different operational needs, thus enhancing predictive capabilities in urban transport systems.
Methodology
The methodology involves the development of an adaptive hybrid ensemble framework that combines five distinct models, each addressing specific aspects of ETA prediction. The models are dynamically fused using a meta-learner that adjusts based on real-time contextual data, including GPS and weather information.
Results
The proposed HYB(nm) framework demonstrated superior robustness and accuracy in ETA predictions, outperforming traditional models and achieving state-of-the-art results comparable to leading graph neural networks. The framework effectively balances stability and efficiency across various operational scenarios.
Implications
The findings suggest that the HYB(nm) framework can significantly enhance the reliability of public transportation systems, improve passenger satisfaction, and provide transit agencies with flexible tools for better operational decision-making. This work contributes to the broader discourse on smart urban mobility and intelligent transportation systems.
seq2cause: One Autoregressive Backbone, Four Causal Discovery Tasks in Event Sequences
Theory
Time Series
Generative Models
- SEQ2CAUSE unifies four causal discovery tasks using a single autoregressive model.
- The framework operates without task-specific retraining, enhancing efficiency.
- A prediction-causality duality is established, linking prediction accuracy to causal identification.
- SEQ2CAUSE is scalable, handling high-dimensional event types effectively.
Read more
seq2cause: One Autoregressive Backbone, Four Causal Discovery Tasks in Event Sequences
Summary
The paper introduces SEQ2CAUSE, a novel framework designed for causal discovery in discrete event sequences across four distinct regimes: event-to-event and event-to-outcome, each at both sample-level and population-level. Traditional methods struggle with high-dimensional event types and multi-stream structures, limiting their applicability. SEQ2CAUSE leverages a pretrained autoregressive (AR) model, which allows for efficient causal identification without task-specific retraining. The framework establishes a prediction-causality duality, where improvements in next-token prediction enhance causal guarantees. The authors demonstrate SEQ2CAUSE's effectiveness on nonlinear structural causal models and real-world vehicle diagnostic logs, achieving scalability with vocabularies up to 29,000 event types. This unified approach addresses the limitations of existing methods and provides a comprehensive solution for causal discovery in complex systems.
Methodology
SEQ2CAUSE repurposes a pretrained autoregressive model to perform conditional independence testing and causal discovery across four regimes. It utilizes the model's next-token prediction capabilities to estimate conditional mutual information (CMI) and perform causal inference without fine-tuning. The framework is designed to operate on GPUs, allowing for parallelized processing of event sequences.
Results
SEQ2CAUSE successfully populates all four causal discovery regimes at scale, demonstrating its capability on both synthetic nonlinear structural causal models and real-world datasets with extensive event types. The results indicate that the method achieves significant improvements in causal identification accuracy while maintaining computational efficiency, outperforming existing methods that are either inapplicable or computationally intractable in similar settings.
Implications
The implications of SEQ2CAUSE extend to various fields that rely on understanding causal relationships in complex systems, such as healthcare, automotive diagnostics, and genomics. By providing a scalable and efficient framework for causal discovery, it can facilitate better decision-making and predictive modeling in these domains.
Livin' on a Prior: Likelihood Score Approximation for Inverse Problems
Generative Models
Computer Vision
Audio & Speech
- Introduces a new framework, Likelihood Score Approximation (LSA), for solving inverse problems.
- Allows for learning unknown degradation models from few paired examples and known models from self-generated data.
- Supports both deterministic and stochastic sampling, enhancing flexibility in model application.
- Demonstrates competitive performance on ImageNet-256 and other benchmarks with minimal training data.
Read more
Livin' on a Prior: Likelihood Score Approximation for Inverse Problems
Summary
This paper introduces Likelihood Score Approximation (LSA), a novel generative framework designed to tackle inverse problems in signal processing, particularly in audio and imaging domains. LSA operates by keeping a pretrained unconditional generative model fixed while learning an observation-conditioned model that approximates the likelihood score from paired samples. This approach allows for the effective learning of unknown degradation models from a limited number of paired examples, while known degradation models can utilize self-generated samples. The framework is versatile, supporting both score and velocity coordinates for training, and can be applied to various generative models. Empirical evaluations demonstrate that LSA achieves competitive restoration quality on the ImageNet-256 benchmark and other tasks, even with only 0.01% of the full training dataset, significantly reducing the number of network evaluations required compared to traditional posterior-sampling methods.
Methodology
The methodology involves a generative framework where a pretrained unconditional model is kept fixed while an observation-conditioned likelihood score is learned from paired data. The framework employs a conditional stochastic-interpolant approach, allowing training in score or velocity coordinates. This flexibility enables the model to adapt to various inverse problems effectively.
Results
LSA was empirically validated across multiple inverse problems in speech and image processing. It achieved competitive or superior restoration quality compared to strong posterior-sampling baselines on the ImageNet-256 benchmark, while requiring significantly fewer network evaluations—up to several orders of magnitude less than traditional methods.
Implications
The findings suggest that LSA can be a powerful tool for efficiently addressing inverse problems in various domains, particularly where data is scarce. Its ability to swap priors post-training opens avenues for further research and application in generative modeling and signal recovery tasks.
Product-Aware Deterministic Rounding for Quantized Matrix Multiplication
Optimization
Theory
Efficient ML
- Dynamic rounding allows for informed rounding decisions based on weight blocks, reducing product errors.
- Static rounding's expected product error can be characterized by the uncentered second moment and centered covariance.
- Exact optimization for rounding decisions is NP-hard, motivating the use of additive guarantees.
- Empirical results show significant error reduction with coordinated rounding and clipping-aware initialization.
Read more
Product-Aware Deterministic Rounding for Quantized Matrix Multiplication
Summary
This paper addresses the challenges of quantized matrix multiplication, particularly the suboptimality of independent rounding of scalars. The authors propose a novel approach called Product-Aware Deterministic Rounding, which coordinates rounding decisions to minimize product errors. The study introduces two main rounding regimes: dynamic and static rounding. In dynamic rounding, the rounding decisions depend on the specific weight block being multiplied, allowing for a more informed choice that can reduce error. The authors provide a polynomial-time deterministic algorithm that guarantees a squared product error bounded by the best admissible error plus a term dependent on the rank of the weight block and the largest row norm. For static rounding, the expected product error is analyzed in terms of the uncentered second moment and centered covariance, revealing that exact optimization is NP-hard even at rank one. Empirical results demonstrate that coordinated rounding significantly reduces errors, with clipping-aware initialization yielding substantial improvements. The findings highlight the importance of considering product-aware cancellation in quantized matrix multiplication, suggesting that traditional rounding methods may not be sufficient for optimal performance.
Methodology
The authors develop a polynomial-time deterministic algorithm for dynamic rounding that minimizes a continuous relaxation of the rounding choices. They analyze the expected product error for static rounding and establish the complexity of exact optimization through reductions from known NP-hard problems.
Results
In balanced synthetic blocks, the proposed dynamic rounding method achieves a median error of 0.010 at K = 1024 and r = 16, compared to 0.899 for traditional round-to-nearest methods. Clipping-aware initialization reduces median normalized error by a factor of 43.4 at ten-percent clipping. On the Digits dataset, centered fixed-bias rounding consistently results in higher median held-out product error than round-to-nearest across various configurations.
Implications
The findings suggest that product-aware deterministic rounding can significantly enhance the performance of quantized matrix multiplication in machine learning applications, particularly in scenarios where precision is critical. This approach could be applied to improve the efficiency and accuracy of neural network inference on resource-constrained devices.
BeatGraph: Self-Supervised Heartbeat Graphs for Infant ECG Representations from the Home Environment
Graph Learning
Time Series
- BeatGraph models heartbeats as nodes in a graph, improving representation of infant ECG data.
- The model is pretrained on a new corpus of 3,408 hours of infant ECG recordings.
- BeatGraph achieves significant performance improvements over existing ECG models on multiple tasks.
- The approach demonstrates strong transferability across different age groups.
Read more
BeatGraph: Self-Supervised Heartbeat Graphs for Infant ECG Representations from the Home Environment
Summary
The paper introduces BeatGraph, a novel approach for modeling infant ECG data by representing heartbeats as nodes in a graph, addressing the limitations of existing ECG models that tokenize signals into fixed-length patches. This method is particularly relevant for infants, whose ECG characteristics differ significantly from adults. BeatGraph utilizes a shared beat encoder to convert each heartbeat into a feature vector, followed by a Transformer that orders the beats temporally and applies graph attention layers to relate beats to one another. The model is pretrained on a newly created corpus of unlabeled infant ECG recordings, where it predicts masked beat embeddings, and is subsequently fine-tuned for various tasks including sleep-wake detection and infant affect recognition. Results demonstrate that BeatGraph outperforms existing baselines across multiple tasks and shows strong transferability across different age groups, achieving high AUROC scores on pediatric benchmarks. Additionally, the authors release a comprehensive infant ECG dataset, marking a significant contribution to the field.
Methodology
BeatGraph employs a graph-based representation of heartbeats, where each heartbeat is encoded into a feature vector. A Transformer with positional encoding organizes the beats in time, and graph attention layers relate all beats to one another. The model is pretrained using a self-supervised approach by predicting masked beat embeddings and is fine-tuned for specific tasks.
Results
BeatGraph improves macro-F1 scores by 0.076 to 0.158 over the strongest baseline across various tasks. It achieves an AUROC of 0.892 on the ZZU-pECG pediatric benchmark and matches the performance of the best self-supervised ECG model on the adult PTB-XL benchmark, despite being pretrained solely on infant data.
Implications
The development of BeatGraph and the accompanying infant ECG corpus could enhance the monitoring of infant health and development, providing insights into autonomic regulation and emotional states that are not easily captured through traditional behavioral observations.
From Grey-Box to Green-Box: When can Physics-Informed Machine Learning Reduce Carbon Footprints in Structural Health Monitoring?
Efficient ML
- Physics-informed machine learning (PIML) can reduce carbon emissions in structural health monitoring compared to traditional black-box models.
- The study evaluates four PIML approaches, revealing that most have lower training emissions, except for input-augmented models.
- Reducing training data requirements through PIML contributes to environmental savings by decreasing training duration.
- A trade-off exists between the complexity of incorporating physics into models and the benefits of reduced data requirements.
Read more
From Grey-Box to Green-Box: When can Physics-Informed Machine Learning Reduce Carbon Footprints in Structural Health Monitoring?
Summary
This paper investigates the environmental impact of machine learning (ML) in structural health monitoring (SHM) by comparing traditional black-box models with physics-informed machine learning (PIML) approaches. The authors highlight the increasing energy demands of ML and the associated carbon emissions, emphasizing the need for sustainable practices in engineering. They evaluate four PIML methods—residual modeling, input augmentation, hybrid modeling, and constrained learning—focusing on their training emissions relative to the amount of training data required. The findings suggest that PIML models generally exhibit lower carbon emissions during training compared to black-box models, with the exception of input-augmented models. The study concludes that while PIML can reduce environmental impacts, a trade-off exists between model complexity and data requirements, necessitating careful consideration in engineering applications.
Methodology
The authors conducted a comparative analysis of four physics-informed machine learning approaches—residual modeling, input augmentation, hybrid modeling, and constrained learning—against traditional black-box models. They assessed the carbon emissions associated with training each model to a specified error threshold, linking training duration to carbon emissions from computing.
Results
The results indicate that most physics-informed models exhibit lower training emissions compared to black-box models, with input-augmented models being an exception. The reduction in emissions is attributed to the decreased amount of training data required, which also lessens the carbon footprint associated with data collection and storage.
Implications
This research underscores the potential for physics-informed machine learning to contribute to more sustainable engineering practices by reducing the environmental impact of model training. It encourages engineers to consider the carbon footprint of their computational methods and explore PIML as a viable alternative to traditional ML approaches.
Persistent Partners Raise Prices Among Learning Agents
Reinforcement Learning
Theory
Optimization
- Keeping the same partner significantly increases average profits and resting prices among learning agents.
- Price increases occur even when rival prices are hidden, indicating that punishment is not the only factor at play.
- The study employs a randomized experimental design to isolate the effects of partner persistence on pricing strategies.
- Untrained Qwen2.5 models exhibit similar pricing behavior, suggesting broader implications for algorithmic pricing.
Read more
Persistent Partners Raise Prices Among Learning Agents
Summary
This paper investigates the impact of persistent partnerships on pricing strategies among learning agents in a repeated interaction setting. The authors conduct a pre-registered randomized experiment based on the Bertrand duopoly model, where agents utilize a tabular Q-learning module to set prices. The study examines whether keeping the same partner influences the prices agents learn and whether this leads to learned punishment for price deviations. The results indicate that maintaining the same partner increases average profits by 0.27 of the gap between competitive and monopoly profits and raises resting prices by 0.17 of the Nash-to-monopoly range. The findings also reveal that when rival prices are hidden, agents cannot punish deviations, yet prices still rise, suggesting that punishment is not the sole mechanism driving price increases. The paper further explores the effects of untrained Qwen2.5 models, demonstrating similar price increases when rival prices are omitted from prompts. Overall, the research highlights the importance of partner persistence in pricing dynamics and challenges traditional notions of punishment in agent interactions.
Methodology
The authors conducted a randomized experiment using a tabular Q-learning module in a Bertrand duopoly setting. They varied three key properties: whether agents kept the same partner or were rematched, whether rival prices were visible or hidden, and whether a messaging channel was available. The experiment included forty runs across five seed blocks, with pre-registered hypotheses and statistical analysis.
Results
The primary result showed that keeping the same partner raised average profits by 0.27 (95% CI: 0.20 to 0.35) and resting prices by 0.17 of the Nash-to-monopoly range. The effect was replicated in further exploratory blocks, with one permanent partner raising profits more than multiple partners. When rival prices were hidden, price increases persisted, while visible rival prices led to a decrease over time. The study also found that a static best responder accounted for a significant portion of perceived punishment in visible conditions, but the overall test for learned punishment was inconclusive.
Implications
The findings have significant implications for understanding algorithmic pricing and potential collusion among AI agents. They suggest that platform operators can influence pricing outcomes by controlling partner persistence, which could inform regulatory approaches to mitigate collusion risks in AI-driven markets.
Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models
Theory
Interpretability
Time Series
- Introduces a hierarchical causal representation learning framework for climate modeling.
- Explicitly separates internal climate variability from externally-driven responses.
- Accurately predicts temperature changes under various future climate scenarios.
- Demonstrates realistic responses to changes in greenhouse gas and aerosol concentrations.
Read more
Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models
Summary
This paper presents a novel hierarchical causal representation learning framework aimed at improving the emulation of climate models, specifically focusing on the effects of anthropogenic forcings on temperature. Traditional climate models are computationally intensive and often lack the ability to provide trustworthy causal insights due to their black-box nature. The proposed framework utilizes sea surface temperature data from a global climate model to explicitly model both internal climate variability and responses to external forcings such as greenhouse gas and aerosol concentrations. By training on future climate scenarios, the framework demonstrates the ability to accurately predict long-term temperature changes and respond realistically to perturbations. This work represents a significant advancement in the field of climate modeling, offering a pathway towards more interpretable and reliable climate projections.
Methodology
The authors build upon the PICABU model, which learns a low-dimensional latent representation of high-dimensional climate states. They introduce a hierarchical latent structure that incorporates global and local climate forcings, using a single-parent assumption for identifiability. The model is trained by maximizing the evidence lower bound (ELBO) under specific constraints to ensure accurate climate state predictions.
Results
The framework successfully captures scenario-dependent warming trends and produces distinct temperature responses to different anthropogenic forcings. It shows improved performance over traditional models, particularly in long-term climate projections and in handling out-of-distribution predictions.
Implications
This research has significant implications for climate science, particularly in enhancing the reliability of climate projections and enabling more effective causal attribution studies. The framework could facilitate better decision-making in climate policy and adaptation strategies by providing clearer insights into the impacts of human activities on climate change.
$λ$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning
Computer Vision
Theory
- Introduces SACReg, a spectral anti-collapse regularizer that enhances representation rank in JE-SSL.
- Demonstrates that existing JE-SSL methods may not prevent dimensional collapse in backbone representations.
- λ-JEPA outperforms LeJEPA and VISReg on ImageNet-1k classification and improves transfer performance across multiple datasets.
- The method is applicable to both image and video self-supervised learning tasks.
Read more
$λ$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning
Summary
This paper introduces a novel approach to self-supervised learning (SSL) called λ-JEPA, which incorporates a spectral anti-collapse regularization (SACReg) mechanism to enhance representation learning in joint-embedding self-supervised learning (JE-SSL) frameworks. The authors identify a critical issue where existing JE-SSL methods, while effective in preventing collapse in projected representations, do not ensure high-rank backbone representations, potentially limiting their effectiveness in downstream tasks. To address this, the paper develops SACReg, which is based on the concept of λ-balance, ensuring that the relative scales of weight matrices across layers are maintained. The authors demonstrate that applying SACReg directly to the backbone representation can prevent dimensional collapse and promote higher effective ranks. The proposed λ-JEPA method is evaluated against existing methods such as LeJEPA and VISReg, showing significant improvements in classification performance on ImageNet-1k and enhanced transfer learning capabilities across multiple downstream datasets. Additionally, λ-JEPA is successfully applied to video SSL tasks, outperforming previous benchmarks on the Something-Something-v2 and Kinetics-400 datasets.
Methodology
The authors analyze a two-layer linear network to derive the SACReg regularizer, which is motivated by λ-balance theory. This regularizer is then applied to the backbone representations in JE-SSL frameworks to prevent collapse and promote variation. The λ-JEPA method integrates SACReg into the training process, optimizing both the backbone and projected representations.
Results
λ-JEPA demonstrates superior performance in classification tasks on ImageNet-1k, achieving results comparable to state-of-the-art methods like DINO. It also shows significant improvements in linear-probe transfer performance across eight downstream image datasets and outperforms existing video SSL methods on benchmarks such as Something-Something-v2 and Kinetics-400.
Implications
The findings suggest that maintaining high-rank backbone representations is crucial for effective transfer learning in self-supervised settings. The proposed method could lead to advancements in various applications of SSL, particularly in computer vision and video analysis, enhancing the robustness and transferability of learned representations.
Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer
Theory
- Semantic ablations can mislead conclusions about the benefits of semantic knowledge in predictive models.
- Content sensitivity and predictive utility are distinct measures that provide different insights into model performance.
- Choosing appropriate controls is crucial for accurately assessing the impact of semantic knowledge.
- The benefit of semantic knowledge varies based on the reference condition used for comparison.
Read more
Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer
Summary
This paper addresses the evaluation of semantic knowledge in tabular learning, particularly in the context of cross-table transfer. The authors argue that traditional methods of assessing the predictive benefits of semantic knowledge through semantic ablations can lead to misleading conclusions. They introduce two key concepts: content sensitivity, which measures performance changes due to alterations in semantic content, and predictive utility, which assesses the actual benefit of the intended semantic knowledge compared to a suitable reference. The study highlights that performance differences observed in semantic ablation experiments may not accurately reflect the predictive advantages of the intended knowledge, as they can also stem from degradation in performance when semantic content is altered. The authors propose an evaluation framework that emphasizes the importance of selecting appropriate controls based on the specific questions being addressed. Through a bounded audit of existing studies, they find that many claims regarding predictive utility lack adequate controls, underscoring the need for careful interpretation of results in semantic knowledge evaluations.
Methodology
The authors conducted real and controlled experiments to evaluate the effects of altering semantic content on model performance. They introduced the concepts of content sensitivity and predictive utility and proposed an evaluation framework that distinguishes between different types of comparisons based on the questions being asked. Additionally, they performed a bounded audit of 25 semantic-ablation comparisons across nine studies to assess the adequacy of controls used in existing research.
Results
The findings indicate that altering semantic content can lead to significant performance differences, which do not necessarily correlate with the predictive benefits of the intended knowledge. In their audit, only one out of 18 predictive utility claims was supported by a control that effectively isolated the tested semantic contribution, highlighting a prevalent issue in the literature.
Implications
The results suggest that researchers should be cautious when interpreting the benefits of semantic knowledge in tabular learning. The proposed evaluation framework can guide future studies in selecting appropriate controls and accurately assessing the contributions of semantic knowledge, ultimately improving the robustness of findings in cross-table transfer tasks.
Online Learning via Learned Latent Bayesian Tracking
Time Series
Optimization
Efficient ML
- AURA framework enables rapid online learning through a learned low-dimensional latent state-space model.
- The method employs extended Kalman filtering for efficient single-step updates in the latent space.
- AURA is validated in real-world scenarios, including wireless communication and image classification.
- The approach shows significant improvements in adaptation speed and accuracy compared to traditional methods.
Read more
Online Learning via Learned Latent Bayesian Tracking
Summary
This paper addresses the challenge of online learning in non-stationary environments, where models must adapt quickly to changing data distributions under strict computational constraints. The authors propose a novel framework called Adaptive Update through Representation Adaptation (AURA), which leverages a learned low-dimensional latent state-space model to facilitate efficient Bayesian filtering for online adaptation. AURA decouples the computational complexity of filtering from the expressiveness of the model, allowing for rapid updates using an extended Kalman filter (EKF) in the latent space. The framework is designed to learn the latent dynamics and the lifting map from offline data, optimizing them for effective online adaptation. Empirical evaluations demonstrate that AURA significantly improves adaptation speed, accuracy, and computational efficiency in applications such as neural wireless receivers and non-stationary image classification, outperforming existing online learning and Bayesian filtering methods.
Methodology
The authors introduce AURA, which involves learning a low-dimensional latent state-space model from offline data. Online adaptation is performed using an extended Kalman filter in this latent space, with full model parameters reconstructed through a learned lifting map. This method allows for efficient Bayesian updates while maintaining model expressiveness.
Results
AURA demonstrated substantial improvements in adaptation speed and accuracy in two main applications: neural wireless receivers adapting to time-varying channels and non-stationary image classification tasks. The framework consistently enabled accurate adaptation with a single low-complexity update step per sample, outperforming existing online learning and Bayesian filtering baselines.
Implications
The findings suggest that AURA can be applied to various domains requiring rapid adaptation to changing data distributions, such as adaptive signal processing, real-time communications, and dynamic visual systems. The framework's ability to efficiently handle high-dimensional models could lead to advancements in online learning methodologies.
Unifying Distributional Training for One-Step Visual Generation
Computer Vision
Generative Models
- Introduces a unified framework for distributional training in visual generation.
- Develops Mixture Gradient Flow (MGFlow) to model feature distributions with Gaussian mixtures.
- Achieves state-of-the-art results on ImageNet, surpassing FD-Loss by 23% and 38%.
- Implements MGFlow for text-to-image generation, outperforming previous multi-step models.
Read more
Unifying Distributional Training for One-Step Visual Generation
Summary
This paper presents a unified theoretical framework for distributional training in one-step visual generation, which aims to synthesize high-quality images in a single network evaluation. The authors introduce a method called Mixture Gradient Flow (MGFlow), which models feature distributions using Gaussian mixtures and connects global objectives to pointwise feature updates through Wasserstein gradient flow. The framework separates distribution modeling from matching discrepancies, allowing for a more flexible approach to feature distribution comparison. MGFlow addresses challenges such as mode collapse by coupling mass-constrained sample assignment with component updates. The authors demonstrate that MGFlow significantly outperforms the existing FD-Loss baseline on ImageNet, achieving state-of-the-art results. Additionally, MGFlow is applied to text-to-image generation, where it post-trains a model to outperform traditional multi-step approaches, showcasing its effectiveness in real-world applications.
Methodology
The paper proposes a theoretical framework that utilizes Wasserstein gradient flow to connect distribution modeling and matching discrepancies. MGFlow is developed to refine modeling granularity using Gaussian mixtures, allowing for both optimal transport and score-based matching. The method incorporates mass-constrained sample assignment to mitigate mode collapse, ensuring effective feature transportation.
Results
MGFlow achieves a FDr6 score of 1.45 on pMF-H and 1.64 on JiT-H on the ImageNet dataset, significantly outperforming the FD-Loss baseline. In text-to-image generation, MGFlow post-trains the FLUX.2 model, achieving 0.900 GenEval and 21.98 PickScore, surpassing all previous methods.
Implications
The proposed framework and methodology have significant implications for enhancing the efficiency and quality of visual generation tasks, particularly in applications requiring rapid image synthesis from textual descriptions. The advancements in distributional training could lead to improvements in various generative models across different domains.
Gradient Surgery for Physics-Informed Neural Networks
Optimization
Theory
- PINNs face significant challenges due to conflicting gradients during optimization.
- The authors identify three distinct phases of gradient conflicts in PINN training.
- PAM-GS is proposed as a solution to adaptively manage task interference.
- Experiments show PAM-GS outperforms existing optimization methods on benchmark PDE problems.
Read more
Gradient Surgery for Physics-Informed Neural Networks
Summary
This paper addresses the challenges faced by Physics-Informed Neural Networks (PINNs) in optimizing composite objectives that combine data fitting with physics-based constraints. The authors identify that existing optimization strategies struggle with conflicting task gradients, leading to slow convergence and unstable training, especially for stiff and high-frequency partial differential equations (PDEs). Through a systematic analysis of gradient conflicts during the training of PINNs across four benchmark problems, the authors observe three distinct phases of gradient conflicts characterized by angle-based and magnitude-based conflicts. To mitigate these issues, they propose a novel method called Physics-Aware Momentum Gradient Surgery (PAM-GS), which adaptively adjusts the optimization process based on the type of gradient conflict observed. Experimental results demonstrate that PAM-GS not only improves solution accuracy but also maintains a strong balance among tasks, outperforming existing methods in most scenarios. The findings highlight the importance of understanding gradient dynamics in multi-task optimization settings, particularly in the context of PINNs.
Methodology
The authors conducted a systematic evaluation of gradient conflicts in PINNs by analyzing training dynamics across four canonical PDE benchmarks. They categorized gradient conflicts into angle-based and magnitude-based types and developed PAM-GS, a conflict-aware optimization method that adapts to the observed gradient conflicts during training.
Results
The experimental results indicate that PAM-GS achieves competitive solution accuracy while maintaining strong task balance, outperforming existing optimization methods in most benchmark problems. The analysis reveals a structured emergence of gradient conflicts, which informs the design of the proposed method.
Implications
The findings suggest that understanding and managing gradient conflicts can significantly enhance the training efficiency and accuracy of PINNs, making them more viable for solving complex PDEs in various scientific fields such as fluid dynamics and electromagnetics.
Quasi Linear Kernel Attention with Infinite Capacity
NLP
Large Language Models
Efficient ML
- Introduces a new capacity metric for kernels to measure expressivity in attention mechanisms.
- Demonstrates that traditional expressive kernels have infinite capacity, while finite feature map-based kernels have limited capacity.
- Proposes additive kernels that achieve quasi-linear computation while retaining infinite capacity.
- Implements an efficient CUDA version of the proposed kernels, outperforming conventional attention methods for long sequences.
Read more
Quasi Linear Kernel Attention with Infinite Capacity
Summary
This paper addresses the computational inefficiencies of traditional transformer architectures, particularly the quadratic scaling of softmax attention with sequence length. The authors propose a novel approach to kernel attention that maintains the expressivity of attention mechanisms while enabling quasi-linear computation. They introduce a concept of 'capacity' for kernels, which quantifies the maximum sequence length for which the attention matrix can approximate the identity. Through their analysis, they demonstrate that expressive kernels like softmax, Gauss, and Laplace possess infinite capacity, while common quasi-linear kernels derived from finite-dimensional feature maps have limited capacity. To overcome this limitation, the authors propose additive kernels constructed from univariate spline and polynomial exponential kernels, which maintain infinite capacity and allow for efficient quasi-linear computation. Their implementation shows significant performance advantages over existing softmax backends for long sequences, particularly for sequence lengths greater than 2,048, and achieves a break-even point against optimized methods like FlashAttention for lengths exceeding 20,000.
Methodology
The authors define a capacity metric for kernels to assess their expressivity in attention mechanisms. They analyze various kernels, proving that expressive kernels like softmax and Laplace have infinite capacity. They then construct additive kernels from univariate spline and polynomial exponential kernels, demonstrating that these can be computed in quasi-linear time. The implementation is validated through numerical experiments and performance benchmarks against existing attention methods.
Results
The proposed additive sorting kernels were shown to be faster than conventional single-precision PyTorch attention for sequence lengths greater than 2,048, and they reached a break-even point with FlashAttention at sequence lengths above 20,000. The results indicate that the new kernels not only maintain expressivity but also significantly improve computational efficiency for long sequences.
Implications
The findings suggest that the proposed kernel attention mechanisms could enable more efficient transformer models capable of handling longer sequences, which is crucial for applications in NLP and computer vision. This could lead to advancements in models that require processing large amounts of data, such as large language models and vision transformers.
Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency
Optimization
Theory
Efficient ML
- Introduces sketched tangent consistency loss (sTCL) for on-the-fly derivative-informed training of neural operators.
- Eliminates the need for offline-generated derivative labels, reducing computational and storage costs.
- Addresses challenges with stiff or ill-conditioned tangent operators through operator-aware loss-conditioning.
- Achieves solution and Jacobian accuracy comparable to traditional methods while being more flexible.
Read more
Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency
Summary
This paper presents a novel approach to training neural operators using a method called sketched tangent consistency loss (sTCL), which enables derivative-informed training without the need for offline-generated derivative labels. Traditional methods for training neural operators require extensive computational resources to generate and store sensitivity labels, which can be a significant bottleneck. The proposed sTCL method leverages random perturbations to enforce sensitivity consistency directly from the governing equations of partial differential equations (PDEs), eliminating the need for offline tangent solves. The authors also address challenges associated with stiff or ill-conditioned tangent operators by introducing lightweight operator-aware loss-conditioning mechanisms. The results demonstrate that sTCL achieves comparable accuracy in both solution and Jacobian estimates to existing offline methods while significantly reducing computational overhead and storage requirements. This approach allows for more flexible and scalable training of neural operators, making it a promising advancement in the field.
Methodology
The authors propose a sketched tangent consistency loss (sTCL) that samples random perturbations in the input function space to compute surrogate Jacobian-vector products using forward-mode automatic differentiation. This method enforces a PDE-derived consistency loss for the forward sensitivity equation without requiring offline tangent labels or changes to the neural-operator architecture. Additionally, lightweight operator-aware loss-conditioning mechanisms are introduced to improve the effectiveness of the training signal.
Results
The experiments conducted across various PDEs, including Helmholtz, nonlinear diffusion-reaction, Burgers, Allen-Cahn, and Navier-Stokes, show that the sTCL method achieves solution and Jacobian accuracy comparable to offline derivative-informed training methods. The approach successfully eliminates the need for extensive offline data generation and storage, demonstrating its efficiency and scalability.
Implications
The proposed sTCL method has significant implications for the training of neural operators in applications such as PDE-constrained optimization, inverse problems, design, and control. By enabling on-the-fly derivative-informed training, it allows for more adaptable and efficient workflows in machine learning tasks involving PDEs.
Graph Forward Distribution Matching for Molecular Inverse Design
Reinforcement Learning
Generative Models
Graph Learning
- GRAPHFDM optimizes molecular design through a forward process, improving stability and property control.
- The method incorporates valid generations into a reward-tilted target distribution for enhanced optimization.
- It achieves up to 53.0% reduction in MAE compared to the strongest baseline methods.
- Maintains chemical validity above 0.99 across multiple property conditions.
Read more
Graph Forward Distribution Matching for Molecular Inverse Design
Summary
This paper addresses the challenges in molecular inverse design, particularly the need for precise control over multiple properties while maintaining chemical validity. Existing reinforcement learning (RL) methods struggle with stability and property gains due to their reliance on reverse sampling as a sequential policy. The authors introduce GRAPHFDM (Graph Forward Distribution Matching), a novel online RL framework that optimizes graph diffusion through the forward process rather than reverse trajectories. By constructing a reward-tilted target distribution that incorporates valid generations, GRAPHFDM enhances the controllability of molecular properties. The methodology includes a unique optimal target distribution and a Structural Distribution Control (SDC) mechanism that jointly optimizes graph size and molecular structure. Experimental results demonstrate that GRAPHFDM significantly outperforms existing methods, achieving the lowest mean absolute error (MAE) on multiple properties while maintaining high chemical validity. The framework also shows generalization capabilities to out-of-distribution property combinations, establishing it as an effective tool for molecular design.
Methodology
The authors propose GRAPHFDM, which shifts the optimization paradigm from reverse sampling to forward process optimization. It constructs a reward-tilted target distribution based on valid molecular generations and employs Structural Distribution Control (SDC) to optimize both graph size and molecular structure. The framework integrates reinforcement signals into supervised learning without the need for reverse trajectory storage.
Results
In experiments involving multi-conditional polymer and small-molecule generation, GRAPHFDM achieved the lowest MAE on every target property, with reductions of up to 53.0% relative to existing methods. The framework maintained a chemical validity rate above 0.99 and demonstrated the ability to generalize to out-of-distribution property combinations.
Implications
The findings suggest that GRAPHFDM can significantly enhance the efficiency and effectiveness of molecular inverse design, making it a valuable tool in drug discovery and materials science. Its ability to control multiple properties simultaneously while ensuring chemical validity could lead to the development of novel compounds with desired characteristics.
Cross-Material Support Transfer for Core-Loss Prediction Under Waveform Covariate Shift
Theory
Optimization
Efficient ML
- Identifies the core issue of covariate shift in core-loss prediction rather than class imbalance.
- Introduces Material-Identity Support Transfer (MIST) for leveraging information from sibling materials.
- Demonstrates that joint training can significantly improve prediction accuracy for underrepresented materials.
- Achieves lower prediction errors with fewer parameters compared to previous models.
Read more
Cross-Material Support Transfer for Core-Loss Prediction Under Waveform Covariate Shift
Summary
This paper addresses the challenge of predicting core losses in power magnetic materials when there is a mismatch between the waveforms used for characterization and those encountered during deployment. Specifically, it focuses on the MagNet Challenge, where a significant disparity exists in the representation of trapezoidal waveforms in the training versus test datasets. The authors argue that the issue is not merely class imbalance but rather a lack of critical information under covariate shift. They propose a novel approach called Material-Identity Support Transfer (MIST), which allows for the sharing of information between similar materials to enhance prediction accuracy. MIST employs a joint training strategy for multiple materials, utilizing feature-wise linear modulation (FiLM) to incorporate material identity and reweight the loss for scarce materials. The results demonstrate a significant reduction in prediction error for the underrepresented material D, showcasing the effectiveness of cross-material support transfer in overcoming the limitations of traditional single-material modeling approaches.
Methodology
The authors conducted controlled experiments to differentiate between class imbalance and missing information under covariate shift. They developed the MIST approach, which trains a single predictor on multiple materials using feature-wise linear modulation to incorporate material identity and reweight losses for scarce materials. This method avoids fine-tuning on the target material's limited data, thus preventing the reinstallation of bias from the training set.
Results
MIST reduced the 95th-percentile relative error (p95) for material D from 20.39% to 12.38% and for the trapezoidal class from 37.4% to 15.16%. This was achieved with a model that had one-sixth the parameters of the best previous submission and without a fine-tuning stage, indicating a substantial improvement in prediction accuracy.
Implications
The findings suggest that joint characterization of scarce materials with their sibling materials can lead to more accurate predictive models in power magnetics. This approach could be applied to other domains where data scarcity and covariate shift are prevalent, potentially improving model robustness and performance.
PALM: Point-in-Time Adaptation for Financial Language Models
NLP
Large Language Models
- Annual pretraining for financial language models is shown to be unnecessary.
- PALM introduces a low-rank adapter that adapts existing models without retraining.
- The method effectively avoids look-ahead bias while maintaining model eligibility.
- Extensive experiments validate that small adapters outperform continued pretraining.
Read more
PALM: Point-in-Time Adaptation for Financial Language Models
Summary
The paper introduces PALM (Point-in-Time Adaptation for financial Language Models), a novel approach to mitigate look-ahead bias in financial language models used for backtesting. Traditional models suffer from this bias as they are trained on data that includes future outcomes, leading to inflated performance metrics. The authors demonstrate that annual pretraining of language models is unnecessary, as newer checkpoints do not significantly outperform their predecessors on the same evaluation window. Instead of requiring a full pretraining run for each year, PALM employs a low-rank adapter that is trained on text published before the decision date, allowing the model to adapt without modifying the pre-trained weights. This method effectively incorporates new information while maintaining eligibility for backtesting. The authors validate PALM across a decade of financial news and various PIT model families, showing that a small adapter outperforms continued pretraining, thus providing a more efficient and effective solution for financial language modeling.
Methodology
The authors conducted experiments comparing the performance of various PIT models with and without annual pretraining. They introduced PALM, which fits a low-rank adapter on eligible text, and evaluated its performance against traditional pretraining methods. The evaluation involved analyzing a decade of financial news and assessing the performance of different model checkpoints.
Results
The results indicate that the newer checkpoints do not provide significant performance improvements over older ones, validating the authors' hypothesis that staleness does not degrade downstream performance. PALM, utilizing low-rank adapters, consistently outperformed models that underwent continued pretraining, demonstrating its effectiveness and efficiency.
Implications
The findings suggest that financial institutions and researchers can save resources by adopting PALM for adapting language models to new information without the need for extensive retraining. This could lead to more efficient model updates and improved forecasting in financial applications.
Robust Graph Clustering Network for Multiple Missing Data
Graph Learning
- First attempt to address simultaneous missing node attributes and graph structure in graph clustering.
- Introduces a view-decoupled dual-branch imputation method to enhance data recovery.
- Employs a multi-hyperspherical mixture prior for improved cluster separation.
- Integrates boundary-aware contrastive learning to sharpen cluster demarcation.
Read more
Robust Graph Clustering Network for Multiple Missing Data
Summary
The paper addresses the challenge of clustering on graphs with simultaneous missing node attributes and structural links, a scenario often encountered in real-world applications. Traditional methods typically follow an imputation-then-clustering approach, which can lead to errors and blurring of cluster boundaries due to cross-view interference. To overcome these limitations, the authors propose the Robust Graph Clustering Network (RGCN), which introduces a view-decoupled dual-branch imputation method that enhances the recovery of missing data while minimizing interference. Additionally, RGCN employs a multi-hyperspherical mixture prior to improve intra-cluster compactness and inter-cluster separability on a latent manifold. A boundary-aware contrastive enhancement objective is also integrated to address imputation bias. The proposed method is evaluated through extensive experiments on six real-world datasets, demonstrating its superior performance compared to state-of-the-art baselines under various missing patterns. The findings suggest that RGCN is effective in reconstructing data manifolds and maintaining well-separated cluster structures even in the presence of significant dual-view incompleteness.
Methodology
RGCN utilizes a dual-pathway decoupled imputation approach to reconstruct both node attributes and graph structure. It optimizes these components alternately while employing a multi-hyperspherical prior to maintain clear inter-cluster margins. The method also incorporates a boundary-aware contrastive learning objective to enhance cluster separation.
Results
The experiments conducted on six real-world datasets show that RGCN consistently outperforms state-of-the-art clustering methods under various complex missing patterns, indicating its robustness and reliability in handling multiple missing data scenarios.
Implications
The proposed RGCN framework can be applied in various domains where graph data is incomplete, such as social network analysis, recommendation systems, and bioinformatics, providing a more reliable clustering solution in the presence of missing data.
Metacognitive Selective Ensemble for Mobile Systems
Efficient ML
Time Series
Computer Vision
- MetaSE maintains a small active set of models to reduce computational costs in mobile sensing.
- The framework leverages short-term persistence in model reliability to make efficient selection decisions.
- MetaSE outperforms fixed and adaptive ensemble methods while using fewer resources.
- The approach is validated across multiple HAR datasets and model architectures.
Read more
Metacognitive Selective Ensemble for Mobile Systems
Summary
The paper introduces MetaSE, an active ensemble framework designed to enhance the efficiency of deep learning models in mobile sensing applications. Traditional ensemble methods improve robustness by combining predictions from multiple models, but they can be computationally expensive, especially in resource-constrained mobile environments. MetaSE addresses this challenge by maintaining a small active set of models that can adaptively select replacements from a larger pool based on short-term reliability. This approach minimizes the need for full-pool evaluations, thereby reducing inference costs and energy consumption. The framework utilizes post-execution evidence to assess model reliability and employs a lightweight routing mechanism to select replacements only when necessary. The evaluation of MetaSE across four human activity recognition (HAR) datasets and various model architectures demonstrates its effectiveness, achieving accuracy comparable to full ensemble methods while being significantly faster and more memory-efficient.
Methodology
MetaSE employs a stateful design that maintains an active set of models across consecutive sensing windows. It uses post-execution evidence from active members to evaluate their reliability and only invokes lightweight routing for replacements when necessary. A class-conditional routing table is utilized to select replacements without executing inactive models, thus minimizing runtime overhead.
Results
MetaSE consistently improves accuracy over a fixed three-model ensemble and achieves results comparable to a full ten-model ensemble while executing only three models during inference. It demonstrates significant performance gains, being 2.7× faster and consuming 69% less memory than full ensemble inference on a Raspberry Pi 4B.
Implications
The findings suggest that MetaSE can enhance the deployment of deep learning models in mobile applications by providing a more efficient ensemble method that balances accuracy and resource consumption. This has potential applications in various mobile sensing tasks, such as health monitoring and activity recognition.
SMAT: Simple and Efficient Merge-Aware Training
Efficient ML
Multimodal
Optimization
- SMAT simulates common merging operations during expert training to enhance merged performance.
- The method achieves improved performance with less than 2% training-time overhead compared to standard fine-tuning.
- Periodic scheduling and efficient parameter operations contribute to SMAT's efficiency.
- SMAT outperforms existing merge-aware training methods across multiple model architectures.
Read more
SMAT: Simple and Efficient Merge-Aware Training
Summary
The paper introduces SMAT (Simple Merge-Aware Training), a novel approach to improve the performance of model merging without incurring significant additional training costs. Traditional expert training focuses solely on task loss, which does not ensure optimal performance post-merging. SMAT addresses this by jointly optimizing both expert loss and expected loss at simulated merged parameters, utilizing three key operations: Scale, Mask, and Perturb. These operations represent common merging techniques and allow for a more effective training process. The authors implement periodic scheduling, kernel fusion, and parameter storage switching to enhance efficiency, achieving one forward and one backward pass per training step. Experimental results demonstrate that SMAT significantly improves merged performance across various language and vision-language models, achieving better scores than existing methods while maintaining a training-time overhead of less than 2%. This work not only advances the field of merge-aware training but also provides insights into the merging behavior of models, paving the way for further research and applications in model merging strategies.
Methodology
SMAT employs a joint optimization strategy that combines expert loss and expected loss at simulated merged parameters. It utilizes three operations—Scale, Mask, and Perturb—to represent common merging techniques. The training process is optimized through periodic scheduling, kernel fusion, and parameter storage switching, allowing for efficient computation with minimal additional passes.
Results
SMAT improves the mean score across five merging methods by 1.07–2.16 points over the strongest baseline for each backbone model tested. For the Llama-3.2-1B-Instruct model, SMAT achieves a 1.12 point improvement over the best baseline, OrthoReg, while using 82% less training time.
Implications
The findings suggest that SMAT can be effectively applied in scenarios requiring model merging, such as ensemble learning and multi-task learning, where combining the strengths of multiple models can lead to enhanced performance. This work also opens avenues for further exploration in efficient training methodologies and model optimization strategies.
Distribution-Conditioned Task Routing for Class-Incremental Learning
Efficient ML
Computer Vision
- Introduces a novel post-hoc task routing framework for class-incremental learning.
- Identifies three sources of routing error: feature-level, task-level, and class-level misalignment.
- Proposes Feature Distribution Calibration (FDC) to address these misalignments without additional training.
- Demonstrates significant accuracy improvements across multiple benchmarks and methods.
Read more
Distribution-Conditioned Task Routing for Class-Incremental Learning
Summary
This paper addresses the challenge of task routing in class-incremental learning (CIL) where task identities are unavailable during inference. The authors propose a novel framework called Feature Distribution Calibration (FDC) that operates without retraining the model or introducing additional routers. FDC identifies and mitigates three types of routing errors: feature-level, task-level, and class-level misalignment. The framework consists of three components: Task Subspace Filtering (TSF) to suppress irrelevant feature components, Residual Likelihood Calibration (RLC) to assess the typicality of input features for a given task, and Prototype Affinity Calibration (PAC) to evaluate the compatibility of inputs with class prototypes. The authors demonstrate that FDC can be applied to various parameter-efficient CIL methods, leading to significant improvements in classification accuracy across multiple benchmarks. This work highlights the importance of effective task routing in CIL and provides a practical solution that enhances performance without the need for extensive retraining.
Methodology
The authors developed the FDC framework, which includes three components: TSF to filter out irrelevant features, RLC to evaluate the typicality of input features for a specific task, and PAC to measure the affinity of inputs to class prototypes. This approach allows for improved task routing without modifying the existing model parameters or requiring additional training.
Results
FDC was tested across five benchmarks and three training seeds, resulting in accuracy improvements ranging from 3.58 to 10.88 percentage points for Cumulative-LoRA. Additionally, when applied to eight other CIL learners, FDC achieved an average improvement of 4.39 percentage points across 40 method-dataset combinations, demonstrating its effectiveness and versatility.
Implications
The findings suggest that effective task routing is crucial for enhancing performance in class-incremental learning scenarios. The proposed FDC framework can be integrated into existing CIL methods to improve their robustness and accuracy, making it a valuable tool for continual learning applications.
Propagate, Then Sharpen: Post-Hoc Refinement of Frozen Node Classifiers
Graph Learning
- Introduction of Propagate, Then Sharpen (PtS) for refining frozen node classifiers.
- PtS alternates between probability propagation and sharpening, improving accuracy without retraining.
- Demonstrated significant accuracy gains over APPNP, especially under feature corruption.
- Robustness against oversmoothing effects, maintaining accuracy across multiple propagation steps.
Read more
Propagate, Then Sharpen: Post-Hoc Refinement of Frozen Node Classifiers
Summary
This paper introduces a novel method called Propagate, Then Sharpen (PtS) for post-hoc refinement of frozen node classifiers in graph-based node classification tasks. The authors aim to improve classification accuracy using only the graph structure and class distributions predicted by a frozen model, without access to node features or model parameters. The PtS method alternates between propagating class probabilities along the graph edges and applying a mass-preserving sharpening step to refine the predictions. This approach is motivated by the Potts energy decomposition, which penalizes disagreement among neighboring nodes and indecision within each node. The authors demonstrate that PtS outperforms the standard APPNP method, achieving significant accuracy improvements on various homophilic graphs, particularly under conditions of feature corruption. The method is shown to be robust against oversmoothing effects and maintains accuracy even with a high number of propagation steps. Overall, PtS represents a significant advancement in enhancing the performance of frozen classifiers in scenarios where retraining is not feasible.
Methodology
The methodology involves a two-step process: first, propagating class probabilities through the graph to leverage information from neighboring nodes, and second, applying a sharpening step that adjusts the class distributions to reduce indecision while preserving the predicted class. The sharpening is controlled by a single hyperparameter, allowing for flexibility in the refinement process. The authors utilize the Potts energy framework to derive the necessary terms for this approach.
Results
The results indicate that PtS improves mean test accuracy by 1.71 percentage points on clean inputs and 3.90 points under severe Gaussian feature corruption compared to independently tuned APPNP. Additionally, PtS shows reduced accuracy degradation with increased propagation depth, maintaining better performance than APPNP and other baseline methods.
Implications
The findings suggest that PtS can be effectively used in real-world applications where node features are unreliable or unavailable, such as in sensor networks or social networks. This method allows for the enhancement of existing models without the need for retraining, making it a valuable tool for practitioners in graph-based learning tasks.
MultiEcho: An Experimental Science of Learned Worlds
Theory
Generative Models
Computer Vision
- MultiEcho framework allows for the estimation of learned world laws through controlled counterfactual interventions.
- Experiments reveal significant variability in response predictability and physical accuracy across different models and contexts.
- Geometric regularities do not necessarily indicate the emergence of physical laws, highlighting the complexity of learned representations.
- The framework provides a multidimensional approach to studying learned worlds, integrating time, space, and counterfactual interventions.
Read more
MultiEcho: An Experimental Science of Learned Worlds
Summary
The paper introduces MultiEcho, a novel framework designed to study world models as experimental systems with their own response laws. The framework enables the estimation of these laws through controlled counterfactual interventions, allowing researchers to delineate their applicability and assess their physical correspondence. The authors conduct experiments across nine simulated physical systems and seven frozen model configurations, employing three reference estimators to predict complete intervention responses and recover intervention parameters. The methodology includes a universal interface for various models, enabling consistent data collection and analysis. The results reveal significant variability in response laws across different subjects, scenes, and contexts, with median reverse R2 values indicating varying degrees of predictability and accuracy. The findings suggest that geometric regularities alone do not suffice for establishing physical law emergence, emphasizing the need for a multidimensional experimental framework to investigate learned worlds. MultiEcho aims to bridge the gap between learned and physical laws, providing insights into the nature of learned representations and their applicability in real-world scenarios.
Methodology
The authors utilize a framework that involves controlled counterfactual interventions across various simulated physical systems. They employ three reference estimators to analyze intervention responses, using a universal interface for data collection from different model configurations. The methodology includes forward and reverse fits to assess the relationship between interventions and outcomes, alongside event-window studies to determine the readability of interventions.
Results
The experiments demonstrate that response laws vary significantly across different models, with median reverse R2 values ranging from 0.208 to 0.986. The findings indicate that while some models exhibit high predictability, they may not accurately reflect physical effects. The viscoelastic exact-reset experiment further distinguishes between geometric and physical evidence, supporting the conclusion that geometric regularities alone are insufficient for establishing physical law emergence.
Implications
The MultiEcho framework has the potential to advance the understanding of learned representations in machine learning, providing a structured approach to investigate the laws governing learned worlds. This could lead to improved model designs that better align with physical realities, enhancing applications in robotics, simulation, and other fields where accurate modeling of the physical world is crucial.
An Attention-Driven Heterogeneous GNN Model for Credit Card Fraud Detection
Graph Learning
- Proposes a novel framework for credit card fraud detection using a heterogeneous GNN model.
- Utilizes SMOTE-Tomek for data balancing to address the imbalanced dataset issue.
- Employs an attention-based message passing technique to capture complex transaction relationships.
- Achieves high performance metrics, indicating effectiveness in detecting credit card fraud.
Read more
An Attention-Driven Heterogeneous GNN Model for Credit Card Fraud Detection
Summary
This paper addresses the growing challenge of credit card fraud detection in the context of an increasingly cashless economy. The authors propose a novel framework that combines a data balancing technique with a deep learning model specifically designed for credit card fraud detection (CCFD). The study utilizes the Synthetic Minority Oversampling Technique (SMOTE)-Tomek to address the imbalanced nature of the dataset sourced from Kaggle. The core of the proposed solution is a Heterogeneous Graph Neural Network (HGNN) that models various transactions as a heterogeneous graph, employing an attention-based message passing mechanism to capture complex relationships, temporal factors, and user behaviors. The integration of SMOTE-Tomek enhances the model's ability to accurately identify fraudulent transactions while minimizing false positives. The model achieved impressive performance metrics, including 99.97% accuracy, 99.48% F1-score, 99.15% precision, and 98.97% recall, demonstrating its effectiveness in real-world CCFD scenarios.
Methodology
The authors collected a credit card fraud detection dataset from Kaggle and applied the SMOTE-Tomek technique to balance the dataset. They then classified the balanced dataset using a Heterogeneous Graph Neural Network (HGNN) that incorporates an attention mechanism for message passing, allowing the model to consider intricate relationships and user behaviors in transaction data.
Results
The proposed HGNN model achieved a remarkable 99.97% accuracy, 99.48% F1-score, 99.15% precision, and 98.97% recall, indicating its strong capability in accurately identifying fraudulent transactions while reducing false positives.
Implications
The findings suggest that the proposed model can significantly enhance the effectiveness of credit card fraud detection systems, making it a valuable tool for financial institutions in combating fraud in an increasingly digital transaction environment.
Optimal Nonparametric Dynamic Pricing with Censored Demand and Adversarial Inventory
Optimization
Theory
- Introduces a nonparametric adversarial-inventory model for dynamic pricing.
- Proposes two algorithms: Double-Grid-UCB and Threshold-UCB, with the latter achieving improved regret rates.
- Establishes a minimax optimality lower bound for the proposed algorithms.
- Demonstrates superior performance of Threshold-UCB in extensive experimental evaluations.
Read more
Optimal Nonparametric Dynamic Pricing with Censored Demand and Adversarial Inventory
Summary
This paper addresses the challenge of online dynamic pricing in the context of censored demand and adversarial inventory, where inventory levels can vary and adapt based on past observations. The authors extend the existing framework by Xu et al. (2026), which was limited by several restrictive assumptions, to a more general nonparametric setting. They propose two algorithms: Double-Grid-UCB and Threshold-UCB. The former discretizes both price and inventory, achieving an expected regret of O(T^(3/4)), while the latter improves this to O(T^(2/3)) by reusing sales observations across inventory levels. The study also establishes a lower bound for the expected regret, demonstrating the minimax optimality of the Threshold-UCB algorithm. Extensive experiments show that Threshold-UCB consistently outperforms benchmark algorithms across various demand functions and noise models, confirming its effectiveness in dynamic pricing scenarios with limited inventory and censored demand.
Methodology
The authors develop two algorithms based on the upper confidence bound (UCB) principle. Double-Grid-UCB discretizes price and inventory, using separate revenue estimates for each pair. Threshold-UCB improves upon this by sharing sales observations across inventory levels, leveraging common demand distributions to optimize revenue estimates.
Results
The algorithms achieve expected regret rates of O(T^(3/4)) for Double-Grid-UCB and O(T^(2/3)) for Threshold-UCB. The latter is shown to be minimax optimal up to logarithmic factors, with extensive experiments confirming its superior performance compared to existing benchmarks.
Implications
The findings have significant implications for revenue management in industries with limited and variable inventory, such as retail and hospitality. The proposed algorithms can enhance pricing strategies, leading to better revenue outcomes in dynamic environments.
MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR
Reinforcement Learning
Large Language Models
Efficient ML
- MaPP addresses inefficiencies in RLVR by denoising advantage estimation and improving prompt selection.
- The framework introduces a composition-invariant intrinsic advantage estimator to reduce gradient estimation errors.
- MaPP consistently outperforms existing RLVR methods, achieving state-of-the-art results with fewer rollouts.
Read more
MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR
Summary
The paper introduces MaPP, a Marginalized Posterior-Predictive framework aimed at enhancing data efficiency in Reinforcement Learning with Verifiable Rewards (RLVR). RLVR improves the reasoning capabilities of large language models (LLMs) but incurs high computational costs due to intensive rollouts and frequent policy updates. Existing methods for prompt selection in RLVR do not adequately account for the reliability of learning signals extracted from sampled responses, leading to inefficiencies. The authors identify a problem known as composition noise, which arises from the uncertainty in group composition affecting the estimation of response utility. To mitigate this, MaPP employs a two-component approach: MaPP-AD, which denoises response-level advantage estimation using Beta-Binomial marginalization, and MaPP-PS, which derives an uncertainty-aware prompt selection score. This framework allows for more accurate prompt selection and improved data efficiency without incurring additional rollout costs. Experimental results demonstrate that MaPP outperforms existing methods, achieving significant accuracy improvements across various tasks while requiring fewer rollouts.
Methodology
The methodology involves two main components: MaPP-AD, which utilizes closed-form Beta-Binomial marginalization to denoise the realized advantages, and MaPP-PS, which leverages the shared Beta posterior to derive an uncertainty-aware prompt selection score. This approach allows for more reliable assessment of prompt informativeness and enhances data efficiency.
Results
Experiments conducted on various tasks, including mathematics, planning, and visual geometry, show that MaPP achieves up to a 2.45% average accuracy improvement over the strongest baseline while using the same rollout budget. This demonstrates the effectiveness of the proposed framework in enhancing performance and efficiency in RLVR.
Implications
The findings suggest that MaPP can significantly improve the efficiency of RLVR applications, making it a valuable tool for training large language models. This could lead to more effective and resource-efficient AI systems capable of complex reasoning tasks.
Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons
NLP
Large Language Models
Optimization
- Learning-rate warmup duration should be treated as a horizon-dependent hyperparameter rather than a fixed heuristic.
- A quadratic model effectively captures the tradeoff between early performance and long-term gains in training.
- Optimal warmup duration varies significantly with peak learning rate and training horizon.
- The study provides a compact scaling law that can predict warmup durations from shorter training runs.
Read more
Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons
Summary
This paper addresses the duration of learning-rate warmup in language-model training, a technique that is widely used but often determined heuristically. The authors propose a quadratic model that captures the relationship between warmup duration and training horizon, highlighting how the optimal warmup duration should vary based on the peak learning rate. They identify a tradeoff between early optimization progress and long-term performance, suggesting that warmup duration should not be fixed but rather treated as a hyperparameter that scales with the training horizon. Through systematic empirical studies, the authors demonstrate that warmup can remain nearly fixed at lower peak rates but should grow substantially with the training horizon at larger rates. This leads to a compact scaling law that can predict optimal warmup durations based on shorter training runs, thereby providing a more principled approach to setting warmup durations in training protocols.
Methodology
The authors conducted a systematic empirical study using Llama-style language models, varying peak learning rates and warmup durations while evaluating validation loss across different training horizons. They employed a warmup-stable schedule to isolate the effects of warmup duration on training performance.
Results
The results indicate that at conservative peak rates, minimal or no warmup suffices, while at higher rates, the optimal warmup duration increases with the training horizon. The findings suggest that the preferred warmup duration is not static but evolves based on the peak learning rate and the overall training budget.
Implications
The insights from this study can lead to more effective training protocols for large-scale language models, optimizing the balance between early training stability and long-term performance. This could enhance the efficiency of training processes in various applications of machine learning.
Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference
Large Language Models
Efficient ML
- Distance-KV introduces a novel static KV retention pattern based on relative distance, improving cache efficiency.
- The method allows for significant reductions in KV cache memory usage (up to 65.4%) and decoding speed (1.66× faster) compared to dense methods.
- Distance-KV outperforms existing KV cache compression techniques, achieving up to 9.3 points improvement on benchmarks.
- The approach is model- and budget-specific, enabling reuse across different inputs without online optimization.
Read more
Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference
Summary
The paper addresses the challenges of memory usage and decoding latency in large language model (LLM) inference, particularly as context lengths increase. Existing key-value (KV) cache compression methods often overlook the significant variation in retrieval capabilities based on relative distance within attention heads. The authors propose a novel approach called Distance-KV, which learns a static KV retention pattern that considers the joint space of layers, attention heads, and relative distances. This pattern is learned offline while keeping the language model frozen, allowing for efficient pruning and compacting of the KV cache without the need for online importance scoring. The evaluation of Distance-KV across three backbone models and four long-context benchmarks demonstrates its superiority over existing methods, achieving substantial reductions in memory usage and improved decoding speed. The findings highlight the importance of relative distance in enhancing LLM efficiency and inform future designs for long-context inference methods.
Methodology
The authors developed Distance-KV by analyzing the retrieval capabilities of individual query heads across varying relative distances. They learned a static KV retention pattern offline using synthetic retrieval examples, optimizing it for specific models and cache budgets. This pattern is then applied during inference to prune and compact the KV cache without requiring online scoring.
Results
Distance-KV consistently outperformed competing KV cache compression methods across multiple models and benchmarks. Notably, it achieved a 65.4% reduction in KV cache memory and a 1.66× speedup in decoding on the Llama-3.1-8B-Instruct model at a context length of 128K. Additionally, it improved performance on the RULER benchmark by up to 9.3 points compared to the strongest existing compression baseline.
Implications
The findings suggest that understanding and leveraging relative distance can lead to more efficient long-context inference methods in LLMs. This could have significant implications for applications requiring processing of lengthy documents or multi-turn interactions, potentially enhancing the performance and scalability of LLMs in real-world scenarios.
Two Heads Are Better Than One: Aggregating Weaker LLMs for Better Forecasts
NLP
Large Language Models
Time Series
- Weak-to-strong aggregation can effectively improve forecasting accuracy using weaker LLMs.
- Learned linear pooling outperforms the strongest individual forecaster in multiple comparison groups.
- Improvements in performance do not depend on the presence of a near-best model.
- Adding more models does not consistently lead to better results.
Read more
Two Heads Are Better Than One: Aggregating Weaker LLMs for Better Forecasts
Summary
This paper investigates the concept of weak-to-strong forecast aggregation in the context of large language models (LLMs) used for predicting real-world events. The authors explore whether aggregating multiple weaker LLM forecasters can yield better performance than a single stronger forecaster. Utilizing the ForecastBench framework, they evaluate 70 LLMs across 16 groups, each containing over 1,000 shared subquestions, resulting in 1,121 weaker-model pairs. The study identifies the strongest individual forecaster in each group based on Brier score and assesses aggregates formed solely from weaker models, with aggregation weights learned from separate training data. The findings reveal significant evidence of weak-to-strong improvement, with learned linear pooling achieving results that match or exceed the strongest individual in 11 out of 16 groups, while remaining within 5% of its Brier score across all groups. The improvements are shown to be independent of having a near-best constituent and are generally accompanied by good calibration. The authors also note that simply adding more models does not guarantee enhanced performance, and that competitive aggregates can still be formed under practical constraints.
Methodology
The authors conducted a systematic evaluation using ForecastBench, organizing 70 LLM forecasters into 16 comparison groups. They identified the strongest individual forecaster in each group based on Brier score and evaluated aggregates formed from weaker forecasters using learned aggregation weights. The study involved extensive experiments with 1,121 weaker-model pairs and analyzed the effects of different aggregation rules.
Results
The study found that learned linear pooling could identify weaker pairs that outperformed the strongest individual forecaster in 11 out of 16 groups, with all groups showing performance within 5% of the strongest model's Brier score. The results indicated that these gains were not reliant on having a near-best constituent and were generally accompanied by good calibration.
Implications
The findings suggest that organizations can leverage weaker LLMs to create effective forecasting models, especially in situations where access to stronger models is limited. This approach could lead to cost-effective solutions in various domains requiring accurate predictions, such as finance, public policy, and inventory management.
DRIFT: Disentangled Responsive-Invariant Flow Transport for Single-Cell Perturbation Prediction
Generative Models
Theory
Optimization
- Introduces DRIFT, a framework for disentangling responsive and invariant cell states.
- Utilizes a variational encoder and conditional flow matching to model perturbation effects.
- Outperforms existing methods in predicting cellular responses to unseen and combinatorial perturbations.
- Addresses limitations of traditional methods by separating intrinsic variability from perturbation effects.
Read more
DRIFT: Disentangled Responsive-Invariant Flow Transport for Single-Cell Perturbation Prediction
Summary
The paper addresses the challenge of predicting cellular responses to perturbations, a crucial task in cellular biology with implications for drug discovery and systems biology. Traditional methods struggle with intrinsic cell-to-cell variability and the inability to measure the same cell before and after perturbation due to destructive single-cell RNA sequencing. The authors propose DRIFT (Disentangled Responsive-Invariant Flow Transport), a novel framework that disentangles the invariant and responsive components of cell states. By employing a variational encoder, DRIFT captures the invariant block (unaffected by perturbations) and the responsive block (affected by perturbations) through conditional priors and an information-theoretic invariance constraint. The method utilizes conditional flow matching to transport only the responsive block, allowing for a flexible, data-driven model of perturbation effects without confounding pre-existing variability. The results demonstrate that DRIFT outperforms existing methods in predicting combinatorial and unseen perturbations across various benchmarks, showcasing its effectiveness in modeling complex biological responses.
Methodology
The methodology involves a variational encoder that disentangles cell states into invariant and responsive components. An information-theoretic constraint ensures the separation of these components, while conditional flow matching is employed to transport only the responsive block based on the perturbation and invariant state, avoiding confounding effects from intrinsic variability.
Results
The DRIFT framework was evaluated across multiple benchmarks and consistently outperformed the strongest existing methods in predicting the effects of combinatorial and unseen perturbations, demonstrating its robustness and flexibility in handling complex biological data.
Implications
The implications of this work are significant for drug discovery, precision medicine, and cell engineering, as it enables researchers to predict cellular responses to perturbations computationally, prioritize experimental combinations, and estimate treatment effects without prior experimental data.
Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization
Efficient ML
Theory
Optimization
- Identifies deterministic overhead regime switching in LSTM based on fine-grained memory budget adjustments.
- Reveals a non-monotone feasibility behavior in ResNet-32, with specific budget ranges leading to OOM errors.
- Demonstrates that the slow execution regime in LSTM is due to broad re-eviction of nearly the entire working set.
- Provides ablation evidence linking the instability in LSTM to the joint size-staleness scoring term.
Read more
Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization
Summary
This paper presents a detailed empirical study of Dynamic Tensor Rematerialization (DTR), an online eviction policy designed for memory-constrained deep neural network (DNN) training. The research identifies deterministic instability in DTR, particularly focusing on how different memory budgets can lead to significant variations in execution overhead. Using the reference DTR simulator, the author examines LSTM and ResNet-32 execution traces, revealing that small changes in memory budget can switch the execution regime from fast to slow, with overhead differences reaching up to 7.3 times. The study also uncovers a phenomenon termed 'deterministic feasibility inversion' in ResNet-32, where certain budget ratios lead to out-of-memory (OOM) conditions, followed by feasible execution at higher budgets. The paper attributes these behaviors to specific mechanisms, including a recursive rematerialization frontier that exceeds memory budgets after all evictable tensors have been processed. The findings highlight the need for a nuanced understanding of DTR's performance under varying memory constraints and suggest that the observed instabilities stem from multiple distinct pathologies rather than a single cause.
Methodology
The study employs the reference DTR simulator to analyze pre-recorded execution traces for LSTM and ResNet-32. It conducts a fine-grained budget sweep to observe the effects of varying memory constraints on execution overhead and feasibility. The methodology includes ablation studies using different scoring heuristics to isolate the causes of observed instabilities.
Results
The results indicate that small changes in memory budget can lead to significant differences in execution overhead, with slow regimes exhibiting overheads up to 10.1 times higher than fast regimes. The study also finds a specific budget range in ResNet-32 that leads to OOM conditions, followed by feasible execution at higher budgets. The mechanisms behind these behaviors are characterized, revealing distinct pathologies affecting DTR's performance.
Implications
The findings suggest that understanding the intricate behaviors of DTR under varying memory constraints is crucial for optimizing DNN training on memory-limited hardware. This research could inform the design of more robust eviction policies and improve the efficiency of DNN training processes.
Progressive Memory Transformer: Memory-Aware Attention for Time-Series
Time Series
- PMT introduces a multi-scale contrastive learning framework for time-series data.
- The architecture features writable, window-aligned memory to expose mid-range representations.
- Three distinct contrastive objectives are employed to supervise representations at different scales.
- PMT achieves strong performance in low-label classification and competitive forecasting.
Read more
Progressive Memory Transformer: Memory-Aware Attention for Time-Series
Summary
The paper introduces the Progressive Memory Transformer (PMT), a novel architecture designed to enhance self-supervised learning for time-series data by explicitly leveraging the structural hierarchy present across multiple scales: local, mid-range, and global. Traditional methods often compress temporal dynamics into a single global representation, neglecting the valuable intermediate motifs that exist within the data. PMT addresses this limitation by incorporating a writable, window-aligned memory that allows for direct supervision of mid-range representations alongside conventional token and sequence-level outputs. The proposed learning framework employs three distinct contrastive objectives tailored to each scale: a Hierarchical Gaussian Contrastive Loss for token-level continuity, a mid-range memory-state loss for motif consistency, and a sequence-level instance loss for overall agreement. The effectiveness of PMT is validated through extensive experiments across various time-series classification and forecasting benchmarks, demonstrating its ability to learn meaningful representations that capture the intricate structures of time-series data.
Methodology
The methodology involves the design of the Progressive Memory Transformer (PMT), which integrates a Progressive Memory Attention (PMA) mechanism that produces mid-range memory states alongside token and sequence-level outputs. The model is trained using three contrastive objectives: a local objective for token continuity, a mid-range objective for motif consistency, and a global objective for sequence-level agreement. This multi-scale approach allows for targeted supervision at each level of representation.
Results
PMT was evaluated across seven UCR/UEA/UCI classification benchmarks, demonstrating strong low-label classification performance (1-5% labels) and competitive forecasting results across multiple horizons. The model's memory states effectively captured mid-range motifs, supported by both quantitative metrics and qualitative analyses.
Implications
The findings suggest that PMT can significantly improve the representation learning of time-series data, making it applicable in various domains such as finance, healthcare, and environmental monitoring where understanding multi-scale temporal structures is crucial.
Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality
Reinforcement Learning
Theory
Robotics
- Introduces Decision-Relevant Prediction Error (DRPE) as a better metric for evaluating world models.
- Develops an iso-error evaluation protocol to isolate the effects of error allocation on planning performance.
- Demonstrates that total prediction error is weakly correlated with planning success, while DRPE shows a strong correlation.
- Finds that error relevance varies by task and becomes more significant with deeper planning.
Read more
Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality
Summary
This paper challenges the conventional assumption that lower prediction error in world models leads to better decision-making. The authors introduce the concept of Decision-Relevant Prediction Error (DRPE), which focuses on the prediction errors that directly impact decision-making, rather than total prediction error. Through a controlled experimental setup using a factored gridworld, they demonstrate that models with similar total errors can exhibit vastly different planning performances based on how errors are distributed across state dimensions. The study evaluates 55 models and finds that DRPE is a stronger predictor of planning success compared to total prediction error. The results indicate that the relevance of prediction errors is task-dependent and that deeper planning amplifies the impact of decision-relevant errors. Furthermore, the authors formalize conditions under which DRPE can accurately rank models, highlighting the inadequacy of total prediction error as a sole evaluation metric. This work emphasizes the importance of assessing world models based on the specific state information that influences decision-making.
Methodology
The authors constructed a controlled environment using a factored gridworld to systematically evaluate the impact of error allocation on planning performance. They generated model families with fixed total prediction error while varying the distribution of errors across relevant and irrelevant state dimensions. This allowed for the direct measurement of DRPE and its relationship with planning success.
Results
The study found that total prediction error had a weak correlation with planning success (Spearman ρ = -0.25), while DRPE exhibited a strong correlation (ρ = -0.84; -0.98 within controlled models). Models with only a 1% difference in total error could show a 60 percentage point difference in planning success. Additionally, the ranking of models could change depending on the task, even with fixed total error.
Implications
This research suggests that evaluating world models should focus on decision-relevant information rather than total prediction error. It has potential applications in reinforcement learning, robotics, and any domain where decision-making is supported by predictive models, leading to more effective model training and evaluation strategies.
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
Reinforcement Learning
Graph Learning
Robotics
- HySTAR separates credit assignment from adaptive representation learning, enhancing stability in MARL.
- The framework utilizes an anchored hypergraph for consistent value decomposition.
- Experiments show HySTAR outperforms MAPPO and other baselines across various scenarios.
- The method effectively handles dynamic agent interactions and changing active-agent sets.
Read more
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
Summary
The paper introduces HySTAR, a novel framework designed to address the challenges of credit assignment in cooperative multi-agent reinforcement learning (MARL) under conditions of partial observability and shared rewards. Traditional methods, such as MAPPO, struggle with assigning team outcomes to individual agents and coalitions due to structural target drift, where the mapping from agents to value components varies over time. HySTAR mitigates this issue by employing an anchored overlapping sparse hypergraph as a stable value-decomposition scaffold, which allows for a consistent credit assignment basis while adapting agent representations through a spatiotemporal encoder. This separation of responsibilities enhances the stability of credit assignment and the adaptability of agent interactions. The framework was evaluated across multiple benchmarks, demonstrating significant improvements over existing methods in terms of performance and convergence speed.
Methodology
HySTAR employs a MAPPO-based architecture that integrates an anchored overlapping sparse hypergraph for stable value decomposition. It utilizes a spatiotemporal encoder to adaptively represent agent interactions based on current spatial and temporal contexts. The framework includes two main components: Anchored High-Order Value Decomposition (AHVD) for stable credit assignment and Spatiotemporal Credit Assignment (STCA) for agent-specific advantage calculations.
Results
HySTAR achieved notable performance improvements in various multi-agent environments. In the SMAC benchmark, it recorded a 16.7% relative gain over MAPPO and a 15.6% gain over HYGMA. It ranked first in all six GRF scenarios and reduced convergence epochs in the Traffic Junction environment by up to 40.2% compared to MAGIC. Additionally, it obtained the highest episode rewards in all three MPE tasks, demonstrating its effectiveness across diverse scenarios.
Implications
HySTAR's approach to stable credit assignment and adaptive representation can significantly enhance the performance of cooperative multi-agent systems in complex environments. Its methodology may be applicable to various domains requiring coordination among multiple agents, such as robotics, autonomous vehicles, and multi-agent simulations.
GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for De Novo Peptide Sequencing
NLP
- Introduces a decoder-guided latent imputation objective to prioritize relevant missing fragments.
- Incorporates mass-aware attention using rotary embeddings to model mass differences between peaks.
- Demonstrates significant improvements in sequencing precision on the NovoBench dataset.
- Retains a standard architecture for inference, simplifying implementation.
Read more
GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for De Novo Peptide Sequencing
Summary
GyroNovo presents a novel framework for de novo peptide sequencing from tandem mass spectra, addressing the challenges posed by sparse, noisy, and incomplete experimental data. The authors identify that existing methods often treat imputation as a fixed reconstruction task without prioritizing the relevance of missing fragments to current decoding errors. GyroNovo introduces two main innovations: (1) a decoder-guided latent imputation objective that emphasizes fragments linked to frequent decoding errors, allowing the model to focus on the most informative missing evidence; and (2) a mass-aware attention mechanism that incorporates pairwise mass differences between spectral peaks using rotary embeddings. This approach aligns the imputation process with the decoder's behavior and enhances the model's understanding of mass relationships in peptide fragmentation. The framework retains a standard encoder-imputer-decoder architecture during inference, requiring no additional inputs. Experimental results on the NovoBench dataset show that GyroNovo achieves significant improvements in peptide-level and amino-acid-level precision compared to the previous state-of-the-art methods, demonstrating its effectiveness in improving de novo peptide sequencing accuracy.
Methodology
GyroNovo employs a decoder-guided latent imputation strategy that maps token-level decoding errors to theoretical fragment cleavages, enhancing the imputation of missing evidence. Additionally, it utilizes a continuous mass-aware rotary attention mechanism in the spectrum encoder to explicitly model mass relationships among spectral peaks, improving the alignment of the model with the underlying geometry of tandem mass spectra.
Results
GyroNovo achieved approximately 9 percentage points improvement in peptide-level precision and 7 percentage points in amino-acid-level precision over the previous state-of-the-art methods on the NovoBench dataset, showcasing its effectiveness in de novo peptide sequencing.
Implications
The advancements made by GyroNovo could significantly enhance peptide identification in proteomics, facilitating the discovery of novel peptides and supporting analyses in cases where reference databases are incomplete or unavailable. This could have broad applications in biomedical research, drug development, and personalized medicine.
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
Efficient ML
Theory
Optimization
- Introduction of Weight Pair Encoding (WeightPE) for neural network weight optimization.
- Utilizes a lossy Re-Pair compressor to induce a smaller grammar in weights.
- Achieves significant reductions in grammar size with minimal impact on accuracy.
- Demonstrates the method on ViT-B/16 and ViT-L/16 models fine-tuned on CIFAR-10.
Read more
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
Summary
This paper introduces Weight Pair Encoding (WeightPE), a novel method for optimizing neural network weights by inducing a smaller grammar. The authors demonstrate that neural network weights can be fine-tuned to accommodate a more compact representation through a lossy Re-Pair compressor integrated within a straight-through estimator. The process involves flattening the int8-quantized weights into a single string and applying a Re-Pair algorithm to identify and merge near-matching patterns while adhering to a global L2 distortion budget. This approach allows for the creation of variable-length patterns that can be reused hierarchically, contrasting with traditional fixed-size codebooks. The authors validate WeightPE on MLP weights of ViT-B/16 and ViT-L/16 models fine-tuned on CIFAR-10, achieving a Re-Pair grammar size that is significantly smaller (0.43× and 0.38×) than that produced by an equivalent int8 quantization-aware training (QAT) run, with only minor reductions in accuracy (1.9 and 1.1 points, respectively). This work is notable for being the first to explicitly use grammar size as a training objective for neural network weights.
Methodology
WeightPE flattens the int8-quantized weights into a single string and applies a lossy Re-Pair compression algorithm to identify and merge recurring patterns. The method employs a straight-through estimator to allow for gradient computation during training, enabling the network to learn with the rewritten weights while maintaining a global L2 distortion budget.
Results
The application of WeightPE resulted in a Re-Pair grammar size that was 0.43× and 0.38× that of the grammar produced by an equivalent int8 QAT run for ViT-B/16 and ViT-L/16, respectively. The accuracy loss was minimal, with reductions of 1.9 and 1.1 points, showcasing the effectiveness of the method in compressing neural network weights while preserving performance.
Implications
The findings suggest that incorporating grammar-based objectives in neural network training can lead to more efficient weight representations, potentially reducing memory requirements and improving computational efficiency. This approach may have applications in resource-constrained environments and could inspire further research into grammar-based methods in machine learning.
TRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac Care
Multimodal
Time Series
NLP
- TRACE integrates unimodal and cross-modal learning to enhance ECG representation.
- The model employs uncertainty-weighted multi-task learning to balance loss contributions.
- TRACE significantly outperforms existing models in arrhythmia classification and ACO detection.
- Real-world validation shows TRACE's clinical utility in acute cardiac scenarios.
Read more
TRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac Care
Summary
TRACE (Text-Reinforced Analysis of Cardio ECGs) is a novel multimodal ECG representation model designed to improve the classification of cardiac conditions by learning clinically relevant signal embeddings. The model addresses the limitations of existing CLIP-style training methods, which often struggle with noisy clinical text and fail to effectively combine unimodal (ECG) and cross-modal (ECG and cardiologist reports) learning. TRACE employs a hybrid architecture that integrates uncertainty-weighted multi-task learning to optimize both masked autoencoder reconstruction and contrastive multimodal alignment. This approach allows for the extraction of high-fidelity findings from cardiologist reports, enhancing the quality of text supervision. The model was rigorously evaluated on public benchmarks for arrhythmia classification and structural abnormalities, demonstrating superior performance compared to existing unimodal and multimodal ECG models. Additionally, TRACE was validated in a real-world clinical setting for acute coronary occlusion (ACO) detection, where it significantly outperformed traditional clinical practices, achieving a 19.0% increase in sensitivity and a 62.6% reduction in false positive rates. This comprehensive evaluation confirms TRACE's potential for impactful clinical applications in high-stakes cardiac care.
Methodology
TRACE utilizes a hybrid architecture combining masked autoencoder reconstruction and contrastive multimodal alignment, optimized through uncertainty-weighted multi-task learning. An LLM-based pipeline is used to curate ECG-specific findings from cardiologist reports, ensuring high-quality text supervision.
Results
TRACE demonstrated robust performance on public benchmarks for arrhythmia classification and structural abnormalities, outperforming existing models. In real-world validation for acute coronary occlusion detection, TRACE achieved a 19.0% increase in sensitivity and a 62.6% reduction in false positive rates compared to clinical baselines.
Implications
TRACE's ability to accurately classify cardiac conditions in real-time could significantly improve patient outcomes in acute care settings, addressing the critical need for efficient and reliable ECG interpretation amidst a shortage of trained professionals.
Hardware-Aware Features for CUTLASS Kernel Selection
Optimization
Efficient ML
- Introduction of hardware-aware representations for kernel selection in CUTLASS.
- Construction of a large dataset (4.9 million kernels) for training selection models.
- Significant reduction in selection regret compared to traditional methods.
- Demonstration of data-efficient transfer learning capabilities within CUTLASS.
Read more
Hardware-Aware Features for CUTLASS Kernel Selection
Summary
The paper addresses the challenge of selecting efficient kernels from the CUTLASS library, which offers a vast number of semantically equivalent implementations for GPU operations. Traditional methods for kernel selection either rely on hand-crafted performance rules or learned models that operate on raw parameters, both of which have limitations. The authors propose a novel approach that incorporates hardware-aware representations, augmenting kernel configurations with estimates of their hardware behavior. They create a dataset of 4.9 million CUTLASS kernels and train both gradient-boosted and neural learning-to-rank models to rank kernel candidates. The results demonstrate that their method significantly reduces selection regret compared to existing methods, achieving a mean selection regret of 6.2%, which is a 64.2% reduction relative to NVIDIA's heuristics. The paper also explores the benefits of data-efficient transfer learning across different kernel configurations, highlighting the effectiveness of explicitly representing hardware behavior in improving kernel selection accuracy and efficiency.
Methodology
The authors developed a hardware-aware representation for kernel selection by augmenting candidate configurations with statically computable estimates of their induced hardware behavior. They constructed a large dataset of CUTLASS kernels and trained ranking models using gradient boosting and neural networks to evaluate and rank kernel candidates without execution.
Results
The proposed method achieved a mean selection regret of 6.2%, significantly lower than the 17% regret from NVIDIA's matrix multiply heuristics and 71.5% from random selection. This represents a 64.2% and 91.3% relative reduction in selection regret, respectively.
Implications
The findings suggest that hardware-aware feature design can serve as a practical foundation for learned performance modeling on modern accelerators, potentially leading to more efficient kernel selection in GPU libraries and improved performance in computational tasks that rely on GPU acceleration.
Towards Universal Representation-Based Process Control
Time Series
- Formulates window-level time series monitoring as a reference-relative process control problem.
- Proposes a nonparametric framework with conformal calibration for regime consistency assessment.
- Accommodates cyclostationary processes as stable operating regimes, addressing classical diagnostic limitations.
- Demonstrates robustness to structural deviations and potential for diagnostic attribution beyond binary detection.
Read more
Towards Universal Representation-Based Process Control
Summary
This paper addresses the limitations of traditional statistical tests in time series monitoring, particularly in scenarios where decisions must be made based on short, rolling windows of data. The authors reformulate window-level monitoring as a process control problem, proposing a reference-based hypothesis testing framework that utilizes empirical reference distributions instead of fixed parametric models. They introduce a nonparametric framework that integrates pretrained time series encoders, kernel density estimation, and conformal calibration, allowing for valid inference in learned representation space. The framework accommodates both stationary and cyclostationary processes as stable operating regimes, thus overcoming the constraints of classical stationarity-based diagnostics. Through extensive experiments, the authors demonstrate the framework's sensitivity to distributional deviations while maintaining well-calibrated inference, showcasing its applicability to a wide range of time series process control tasks.
Methodology
The authors develop a representation-based, nonparametric framework that combines pretrained time series encoders with kernel density estimation and conformal calibration. This approach allows for finite-sample-valid inference without relying on fixed parametric assumptions, enabling a unified treatment of diverse temporal behaviors.
Results
The experiments conducted show that the proposed framework is sensitive to window-level distributional deviations while maintaining well-calibrated inference under stable reference regimes. This highlights the framework's robustness and its capability to effectively monitor and control time series processes.
Implications
The proposed methodology has significant implications for real-time monitoring and control of time series systems, particularly in industrial applications where maintaining operational consistency is critical. It offers a more flexible and robust alternative to traditional statistical methods, potentially improving the detection of subtle deviations from expected behavior.
Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
Interpretability
- Proposes a unified evaluation protocol for robust counterfactual explanations across different model changes.
- Demonstrates that existing robustness scores are not comparable due to method-specific evaluations.
- Finds significant variability in the performance of robust methods depending on the type of model change.
- Highlights the importance of separating generation performance from robustness in evaluations.
Read more
Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
Summary
This paper addresses the challenge of evaluating robust counterfactual explanations (CFEs) that remain valid after changes to the underlying model. The authors argue that existing evaluations are method-specific and not comparable due to varying definitions of robustness against different model changes. They propose a unified evaluation protocol that tests multiple robust CFE methods against the same set of model changes across four tabular datasets. This protocol allows for a consistent comparison of robustness scores, coverage, base validity, and proximity. The study finds that the performance of robust methods varies significantly with the type of model change, highlighting that guarantees for robustness under one type of change do not necessarily apply to others. The authors emphasize the need for a common evaluation framework to ensure that robustness claims are meaningful and comparable across different methods.
Methodology
The authors developed a cross-family evaluation protocol that maintains fixed factual instances and counterfactuals while testing six robust CFE methods and two standard baselines against eight types of model changes. They characterized each model change by its prediction disagreement and probability differences, allowing for comparisons across different architectures. The evaluation metrics included empirical robustness, coverage, base validity, and proximity.
Results
The results indicate that the relative performance of robust CFE methods varies significantly with the type of model change. For instance, bounded parameter perturbations resulted in an average of 0.95% change in test predictions, while bootstrap retraining led to a 4.9% change. The method RobX showed the most consistent transferability across different changes, although achieving greater stability sometimes required larger interventions. The study also revealed that robustness scores could be misleading if not reported alongside coverage and distance metrics.
Implications
The findings suggest that practitioners should be cautious when interpreting robustness claims of CFE methods, as these claims may not generalize across different model changes. The proposed evaluation protocol could serve as a standard for future research, enhancing the reliability of robustness assessments in machine learning models, particularly in applications requiring interpretability and accountability.
Does Uniform Discrete Diffusion Need Time?
NLP
Generative Models
Large Language Models
- Population-optimal UDM predictors depend on time, but this dependence is often negligible in finite data settings.
- Time-agnostic predictors in UDMs can achieve competitive performance compared to time-conditioned models.
- The predictive benefit of time conditioning is primarily observed in high-noise scenarios.
- The study provides a quantitative understanding of how time sensitivity relates to the separation of training examples.
Read more
Does Uniform Discrete Diffusion Need Time?
Summary
This paper investigates the necessity of explicit time conditioning in Uniform Discrete Diffusion Models (UDMs), which are commonly used in language modeling. The authors demonstrate that while the population-optimal predictor in UDMs is generally time-dependent, this dependence can be negligible in finite-data scenarios typical of language tasks. They argue that when a corrupted training sequence is closer to its original clean sequence than to other competing sequences, the model's predictions become largely insensitive to time throughout most of the diffusion process. Empirical evaluations reveal that trained language UDMs exhibit limited time sensitivity and that time-agnostic models can perform comparably or even outperform time-conditioned models across various datasets and training objectives. The findings suggest that explicit time conditioning may not be necessary in practice, prompting a reevaluation of model design and training strategies in UDMs.
Methodology
The authors analyze the time dependence of UDMs by characterizing how the population-optimal predictor relies on time and how this dependence can be suppressed in finite data scenarios. They derive theoretical bounds linking time sensitivity to the separation margin of corrupted sequences and conduct empirical evaluations of trained UDMs across various noise levels and datasets.
Results
The study finds that while the population-optimal predictor is time-dependent, trained language UDMs show limited sensitivity to time over most of the diffusion trajectory. Time conditioning provides significant benefits primarily near high-noise endpoints. A hybrid model that incorporates time conditioning selectively performs similarly to fully time-conditioned models, while fully time-agnostic UDMs remain competitive.
Implications
These findings suggest that machine learning practitioners may not need to rely on explicit time conditioning in UDMs for language tasks, potentially leading to simpler and more efficient model architectures. This could influence future research directions in diffusion models and their applications in natural language processing.
Feedback-Robust AI for Patient Knowledge Graphs
Graph Learning
Time Series
Interpretability
- Introduction of ClosedLoopBench, a benchmark for evaluating temporal relations in clinical settings.
- Development of feedback-robust patient knowledge graphs that integrate evidence and typed relations.
- Demonstration of significant biases in traditional correlational methods for constructing patient KGs.
- Identification of consistent physiological response patterns across patients despite individual variability.
Read more
Feedback-Robust AI for Patient Knowledge Graphs
Summary
This paper introduces a novel approach to constructing patient knowledge graphs (KGs) that account for the feedback inherent in clinical settings, particularly in anesthesia and intensive care. The author presents ClosedLoopBench, a benchmark comprising 29 relations derived from clinical practice, physics, and pharmacology, evaluated on 3,442 surgical cases from VitalDB. The study highlights the challenges of traditional correlational KG construction, which often fails to accurately represent the complex interactions between clinician actions and patient responses. By employing negative-control action streams, the research demonstrates that many estimators can identify significant relations without calibration, although calibration still reveals biases in the interpretation of these relations. The proposed feedback-robust patient KGs integrate concept nodes with evidence pointers and typed relations, significantly reducing false positives compared to correlational methods. The findings indicate that while patient-specific estimates of drug and ventilator responses do not outperform population estimates, certain physiological responses, such as the pulse-arrival-time-systolic-pressure slope, exhibit consistent patterns across patients. This work contributes to the understanding of temporal relation discovery in clinical monitoring and emphasizes the importance of accounting for feedback in the analysis of patient data.
Methodology
The study employs a combination of negative-control action streams and various estimators to analyze temporal relations in patient data. ClosedLoopBench serves as a benchmark for testing the significance of relations derived from clinical practice. The methodology includes cross-patient transplant and random-time action streams to assess the robustness of the identified relations.
Results
The analysis reveals that six out of twelve estimators can identify a significant number of distinct relations without calibration. After calibration, biases remain, particularly in the interpretation of clinician actions. The proposed feedback-robust patient KGs show a marked reduction in false concept-level relation instances compared to traditional correlational methods, with 0.06-0.10 false instances per graph versus 10-12 for correlational constructions. Additionally, the study finds that patient-specific estimates do not predict outcomes better than population estimates, except for certain physiological measures.
Implications
The findings suggest that incorporating feedback mechanisms into patient knowledge graphs can enhance the accuracy of clinical predictions and interpretations. This approach may improve decision-making in anesthesia and intensive care by providing more reliable insights into patient responses to treatments. The methodology could be applied to other areas of clinical monitoring and patient care, potentially leading to better patient outcomes.
Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
Optimization
Theory
Robotics
- Introduces a penalized DRO formulation that incorporates Wasserstein penalties for adversarial distributions.
- Proves that optimal transport maps for adversarial training are cyclically monotone.
- Develops Multi-start Particle Ascent (MPA) to enforce cyclical monotonicity in adversarial training.
- Proposes using input-convex neural networks to parameterize adversarial transport maps.
Read more
Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
Summary
This paper addresses the challenges of distributionally robust optimization (DRO) in machine learning, particularly under distribution shifts. The authors propose a penalized DRO formulation where an adversary can select any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. They reformulate the adversary's problem as an optimization over transport maps that push empirical samples to adversarial ones, proving that optimal maps are cyclically monotone. The paper critiques standard adversarial training methods for violating cyclical monotonicity and introduces two novel approaches: Multi-start Particle Ascent (MPA), which alternates gradient ascent with optimal reassignment to maintain cyclical monotonicity, and a method that parameterizes adversarial maps as gradients of input-convex neural networks (ICNNs), ensuring cyclical monotonicity by design. Experimental results demonstrate that these methods outperform standard adversarial training and state-of-the-art baselines in robust regression, image classification, and robust control tasks, achieving enhanced robustness and generalization under distribution shifts.
Methodology
The authors reformulate the adversary's problem as an optimization problem over transport maps, ensuring that the maps are cyclically monotone. They introduce MPA, which combines parallel gradient ascent with optimal reassignment of samples, and an ICNN-based approach that guarantees cyclical monotonicity by design.
Results
The proposed methods, MPA and the ICNN-based approach, consistently outperform standard adversarial training and existing baselines across multiple tasks, including robust regression and image classification, leading to better robustness and generalization under distribution shifts.
Implications
The findings suggest that incorporating optimal transport geometry into adversarial training can significantly enhance the robustness of machine learning models against distribution shifts, making these methods valuable for real-world applications where data distributions may vary.
A Comparative Analysis of Attention versus State-Space Models for In-Context Learning
Theory
- Introduces belief geometry as a unified framework for comparing attention and state-space models.
- Identifies three key capabilities for sequential learning: evidence assembly, belief maintenance, and addressing.
- Demonstrates that SSMs achieve optimal performance in belief maintenance and have memory advantages.
- Shows that attention mechanisms have a significant advantage in content addressing due to their softmax selection process.
Read more
A Comparative Analysis of Attention versus State-Space Models for In-Context Learning
Summary
This paper presents a comprehensive analysis of attention mechanisms, specifically Transformers, and state-space models (SSMs) in the context of in-context learning. The authors introduce a novel framework termed 'belief geometry' to systematically compare the representational capabilities of these two architectures. By generalizing in-context linear regression (ICLR) and employing cumulative Bayes regret as a metric, the study identifies three critical capabilities for sequential learning: evidence assembly, belief maintenance, and addressing. The authors analyze these capabilities through controlled cases, revealing that SSMs excel in belief maintenance and memory efficiency, while attention mechanisms demonstrate superior content addressing capabilities. Empirical experiments with LLaMA-type Transformers and Mamba-2 validate the theoretical insights, indicating that the architectural advantages identified extend beyond the simplified models used in the analysis.
Methodology
The authors develop a generalized formulation of in-context linear regression and utilize cumulative Bayes regret to evaluate the performance of attention and state-space models. They analyze three specific cases to isolate the capabilities of evidence assembly, belief maintenance, and addressing, and conduct experiments with LLaMA-type Transformers and Mamba-2 to validate their theoretical findings.
Results
The analysis reveals that SSMs attain optimal regret in belief maintenance, outperforming attention mechanisms in scenarios requiring memory efficiency. In contrast, attention mechanisms exhibit an exponential advantage in content addressing, allowing for more effective selection among retained tokens. Empirical experiments corroborate these theoretical insights, demonstrating that the architectural lessons derived from the analysis hold true in practical applications.
Implications
The findings suggest that the choice between attention and state-space models should be informed by the specific requirements of the sequential learning task at hand. Understanding the strengths and weaknesses of each architecture can guide the design of more efficient and effective models for various applications in machine learning.
Arbitrary-Accuracy Neural Approximation with Optimal Neuron Count and Near-Optimal Bit Complexity
Theory
Efficient ML
- Establishes that d + 1 is the minimum neuron count for arbitrary-accuracy approximation of Hölder-continuous functions.
- Introduces a fixed activation function that allows for efficient neural network construction.
- Demonstrates near-optimal bit complexity for the proposed network architectures.
- Provides explicit constructions with rational parameters that can be verified in exact arithmetic.
Read more
Arbitrary-Accuracy Neural Approximation with Optimal Neuron Count and Near-Optimal Bit Complexity
Summary
This paper investigates the minimum number of hidden neurons necessary for the arbitrary-accuracy approximation of multivariate Hölder-continuous functions on the domain [0, 1]^d, alongside the associated encoding complexity. The authors construct a fixed activation function that allows a two-hidden-layer network with widths d and 1 to achieve arbitrary accuracy in the uniform norm. They establish that d + 1 is the exact minimum number of hidden neurons required for standard feedforward networks with locally integrable activations and affine outputs. Additionally, a simpler construction using a single elementary activation (combining floor and exponential functions) is presented, which requires three hidden layers and only two neurons above the minimum. The paper also discusses the implications of skip connections, showing that widths d, 1, and 1 suffice when allowed. The authors quantify the bit complexity of their constructions, demonstrating that it matches the metric-entropy lower bound up to a logarithmic factor, thus establishing their constructions as near-optimal in terms of information theory. The findings contribute significantly to the understanding of neural network efficiency in terms of neuron count and parameter encoding.
Methodology
The authors construct neural networks with specific architectures and activation functions to achieve arbitrary accuracy in approximating Hölder-continuous functions. They utilize theoretical proofs to establish the minimum neuron count and analyze the bit complexity of their constructions, comparing them against established lower bounds.
Results
The paper proves that a two-hidden-layer network with widths (d, 1) achieves arbitrary accuracy with exactly d + 1 hidden neurons. It also presents a three-hidden-layer network using a single elementary activation that requires d + 3 neurons, and a variant with a skip connection that requires d + 2 neurons. The bit complexity of the proposed networks is shown to be Θ(ε−d/α log(1/ε)), aligning closely with theoretical lower bounds.
Implications
The findings have significant implications for the design of efficient neural networks, particularly in applications requiring high accuracy with minimal resource usage. The results can inform future research on neural network architectures and their practical implementations in various domains.
From Phase Transition to Systemic Failure: A Decoupled Analytics Framework for GNN Robustness
Graph Learning
- Developed a decoupled quantitative data analytics testbed for GNN robustness evaluation.
- Identified distinct failure modes in GNNs: gradual degradation under label noise and severe degradation under distribution shift.
- Demonstrated that GNN architecture does not compensate for data integrity issues.
- Emphasized the importance of monitoring input drift over mitigating supervision noise.
Read more
From Phase Transition to Systemic Failure: A Decoupled Analytics Framework for GNN Robustness
Summary
This paper addresses the critical issue of data quality affecting the deployment of Graph Neural Networks (GNNs) in real-world applications, focusing on two main challenges: label noise and feature distribution shift. The authors construct a synthetic homophilic graph regression benchmark that allows for the independent manipulation of these two factors. Through extensive experimentation involving 41 configurations and 410 runs, they uncover distinct patterns of GNN performance degradation. The study reveals that GNNs exhibit relative stability under moderate label noise, with significant performance drops occurring only after a 50% noise ratio. In contrast, extreme feature distribution shifts lead to severe performance degradation, with mean squared error (MSE) increasing dramatically and correlation dropping significantly. The findings suggest that GNNs are more tolerant to label noise than to distribution shifts, highlighting the need for a data-centric approach to model robustness. The paper provides a controlled empirical baseline for understanding GNN vulnerabilities and offers practical recommendations for monitoring and maintaining model performance in deployment scenarios.
Methodology
The authors created a synthetic graph regression platform based on empirical data from real-world ecological datasets. This platform allows for the independent manipulation of label noise and feature distribution shifts, enabling a controlled evaluation of their effects on GNN performance. The study involved systematic experimentation to assess the behavior of various GNN models under different noise and shift conditions.
Results
The results indicate that GNNs maintain stable performance under moderate label noise until a critical threshold (around 50% noise) is reached, after which performance sharply declines. Conversely, under extreme feature distribution shifts, all tested models experienced significant performance degradation, with MSE increasing by factors of 48 to 316 and correlation dropping by 73% to 89%. This highlights a greater vulnerability of GNNs to distribution shifts compared to label noise.
Implications
The findings underscore the necessity for proactive monitoring of data quality in GNN applications, particularly in dynamic environments. The study suggests that ensuring data integrity and addressing distribution shifts are crucial for the reliable deployment of GNNs in real-world scenarios. The insights can inform strategies for model maintenance and recalibration in industrial graph mining applications.
Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning
Reinforcement Learning
Large Language Models
Robotics
- CRBC merges rollouts into a finite empirical process for better evidence aggregation.
- The method allows recursive propagation of evidence through shared anchor states.
- CRBC consistently outperforms existing methods in long-horizon reinforcement learning tasks.
- The approach achieves state-of-the-art performance on multiple benchmarks.
Read more
Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning
Summary
This paper introduces the Cross-Rollout Bellman Closure (CRBC) method, which enhances long-horizon reinforcement learning by effectively aggregating evidence from multiple rollouts. Traditional group-based reinforcement learning methods, such as GRPO, struggle to provide accurate step-level credit due to limitations in how they aggregate information from shared anchor states. CRBC addresses these limitations by merging rollouts into a finite empirical process that allows for recursive propagation of evidence through shared anchors. This method evaluates the behavior-policy Bellman fixed point using a single linear solve, which incorporates both successful and failed continuations from different rollouts. The CRBC method is shown to improve performance and learning efficiency across various benchmarks, including ALFWorld, WebShop, and Sokoban, achieving state-of-the-art results without requiring additional environment rollouts.
Methodology
The CRBC method merges rollout groups into an empirical process with absorbing boundaries for success and failure. It evaluates the behavior-policy Bellman fixed point through a linear solve, allowing for the aggregation of action values based on empirical action and transition frequencies. This process enables the recursive propagation of evidence across rollouts, improving the accuracy of step-level credit assignment.
Results
CRBC demonstrated significant improvements in final performance and learning efficiency across various benchmarks. For instance, it outperformed the strongest baseline by 5.59 percentage points on the ALFWorld benchmark using the Qwen2.5-1.5B-Instruct model, setting a new state-of-the-art performance.
Implications
The CRBC method has the potential to enhance the capabilities of reinforcement learning agents in complex environments, particularly in tasks requiring long-term planning and decision-making. Its ability to efficiently aggregate evidence from multiple rollouts could lead to more effective training strategies for large-scale reinforcement learning applications.
Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
Reinforcement Learning
Robotics
Computer Vision
- H-JEPA separates perceptual representation from control state, enhancing planning efficiency.
- Utilizes a Bures-Wasserstein prior for isotropic regularization of perceptual codes.
- Introduces port-inverse consistency (PIC) for effective action readout and error reweighting.
- Achieves superior performance on pixel-based control tasks compared to existing models.
Read more
Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
Summary
The paper introduces H-JEPA, a novel approach to action-conditioned world modeling that separates perceptual representation from control state, addressing limitations in existing joint-embedding predictive architectures (JEPAs). Traditional JEPAs often conflate perception and control in a single embedding, which can lead to inefficiencies in planning from pixel data. H-JEPA utilizes a wide perceptual code that is regularized towards isotropy using a Bures-Wasserstein prior, while a fixed orthonormal projection defines the control state, inheriting the covariance of the perceptual code without requiring a separate objective. The authors implement phase-conditioned dissipative port-Hamiltonian dynamics for state evolution and introduce port-inverse consistency (PIC) for action readout, which effectively reweights prediction errors based on action influence. The results demonstrate that H-JEPA outperforms or matches existing reconstruction-free baselines across multiple pixel-based control benchmarks, achieving significant improvements, particularly on the OGB-Cube task. The paper also includes ablation studies to analyze the contributions of various components of the model, providing insights into the structured predictor, PIC, prediction horizon, state rank, and anti-collapse prior.
Methodology
The methodology involves creating a two-space world model where a perceptual code is regularized towards isotropy and a fixed orthonormal projection defines the control state. The model employs phase-conditioned dissipative port-Hamiltonian dynamics for state evolution and utilizes port-inverse consistency (PIC) for action readout, ensuring that the readout is parameter-free and directly linked to the action's influence on the state.
Results
H-JEPA matches or exceeds the performance of reconstruction-based and action-decoding baselines on four control benchmarks, with the most notable improvement on the OGB-Cube task, achieving 91.9% success compared to 79.3% for the strongest baseline. The model demonstrates effective planning and action sensitivity, with ablation studies confirming the contributions of its structured predictor and other components.
Implications
The findings suggest that separating perceptual and control representations can lead to more efficient planning in reinforcement learning tasks. This approach may have applications in robotics and other domains requiring action-sensitive planning from visual inputs.
Rethinking Contextualization by Reinterpreting Attention Head Channels
NLP
Large Language Models
Interpretability
- Contextualization is influenced by the varying information levels of words.
- Less informative words absorb contextual information more selectively.
- Attention heads can be reinterpreted as channels governed by singular vectors.
- The study provides a framework for automated interpretation of attention mechanisms.
Read more
Rethinking Contextualization by Reinterpreting Attention Head Channels
Summary
This paper presents a novel perspective on contextualization in language modeling, emphasizing the role of attention heads and their channels. The authors argue that previous studies have focused too narrowly on individual words and attention heads, lacking a comprehensive understanding of their collective behavior. They propose a general principle where words with varying information levels interact during contextualization, with less informative words absorbing more contextual information selectively. The study introduces three hypotheses regarding contextualization: (1) words inherently carry different amounts of information, (2) more informative words act as stable sources during contextualization, and (3) information transmission is selective based on semantic matching. The authors reinterpret attention heads as channels gated by singular vectors, revealing that these vectors indicate the direction of information flow and enable automated interpretation of attention mechanisms. Through empirical observations and theoretical explanations, the paper highlights the dynamic nature of contextualization and the importance of attention head properties in facilitating effective information transfer between words.
Methodology
The authors employed a dual approach combining phenomenological observations and mechanistic explanations. They quantified word information based on its impact on outputs across semantically related queries and analyzed the behavior of low and high-information words during contextualization. Additionally, they reinterpreted attention heads as information channels gated by singular vectors, allowing for a continuous representation of head functionality.
Results
The findings confirm that words carry different amounts of information, with low-information words showing greater variation in representations during contextualization. The study also demonstrates that information flows more effectively between semantically matched words, supporting the proposed hypotheses about contextualization and attention mechanisms.
Implications
This research has significant implications for improving language model interpretability and understanding the dynamics of information flow in neural networks. It suggests that refining attention mechanisms could enhance model performance in various NLP tasks, leading to more effective language understanding and generation.
Task-Aware Discretization of Differentiable Logic Gate Networks
Efficient ML
Theory
- Introduces a task-risk perspective on discretization, distinguishing between conventional discretization gaps and discrete optimality gaps.
- Demonstrates that local argmax selection can lead to significant task performance losses, even with globally optimal relaxed models.
- Develops a task-aware gate selection method that utilizes conditional risk and provides efficient first-order approximations.
- Empirical analysis reveals the importance of locality in gate selection, leading to a progressive discretization strategy that improves task performance.
Read more
Task-Aware Discretization of Differentiable Logic Gate Networks
Summary
This paper addresses the challenge of discretizing Differentiable Logic Gate Networks (DLGNs) for efficient Boolean network inference. Traditional methods utilize local discretization techniques, such as argmax selection, which can lead to suboptimal task performance even when the relaxed model is globally optimal. The author demonstrates that high confidence in gate selection does not guarantee task-optimal discretization. To overcome this issue, the paper introduces a task-aware approach to discretization that directly considers downstream task risk. By deriving bounds on task-aware gate selection and developing an efficient first-order approximation, the author provides a framework for progressive discretization. This method allows for the reliable assessment of local argmax decisions while avoiding the pitfalls of nonlocal interventions. Empirical results show that the proposed task-aware discretization method outperforms existing strategies, particularly in scenarios where local decisions are crucial for task performance.
Methodology
The author formulates a task-aware discretization framework that incorporates downstream task risk into the selection of discrete gates. This involves deriving conditional expectations for gate selection and utilizing first-order approximations to assess the reliability of local decisions. The methodology includes progressive freezing of gates based on their predicted impact on task performance, rather than solely on local confidence metrics.
Results
The experiments conducted on convolutional DLGNs demonstrate that the proposed task-aware discretization method significantly reduces the discrete optimality gap compared to traditional argmax-based approaches. The results indicate that local argmax decisions are reliable when assessed through the lens of task performance, leading to improved outcomes in resource-constrained inference scenarios.
Implications
This research has implications for the design of efficient neural networks, particularly in environments where computational resources are limited. The task-aware discretization approach can enhance the performance of DLGNs in various applications, including embedded systems and real-time processing tasks, by ensuring that the discretized networks are optimized for specific downstream tasks.
Moment-guided edge sampling
Graph Learning
- Introduces a moment-guided edge sampling framework that connects local edge edits to global graph structure.
- Develops efficient combinatorial and low-rank methods for computing moment changes.
- Demonstrates that moment-preserving sampling retains important structural properties.
- Shows the impact of edge structures on supervised node classification and competitive performance in contrastive learning.
Read more
Moment-guided edge sampling
Summary
This paper presents a novel moment-guided edge sampling framework aimed at addressing the challenge of quantifying and controlling the effects of local edge edits on global graph structure. The authors utilize spectral moments of the random-walk transition matrix to create a connection between local edge modifications and global graph properties. They introduce two methods for computing moment changes: a combinatorial method for low-order moments that operates in constant time and a low-rank method that supports arbitrary moment orders and reduces computational complexity significantly. The framework allows for the selection of edge edits that steer the graph towards desired moment profiles, thereby preserving important structural properties such as the triangle-weighted clustering coefficient. The authors validate their approach through case studies in supervised node classification and graph contrastive learning, demonstrating that different edge structures influence performance and that moment-guided sampling can enhance graph learning tasks.
Methodology
The authors propose a moment-guided edge sampling framework that computes edge-wise moment changes using two methods: a combinatorial method for low-order moments that achieves O(1) time complexity per candidate edge and a low-rank method that reduces computational costs for arbitrary moment orders and supports batched edits. The framework iteratively selects edge edits based on their moment changes to achieve a target moment profile.
Results
The framework successfully demonstrates that edge-wise moment changes provide interpretable structural signatures and that preserving moments can retain related structural properties. The case studies reveal that different edge structures have distinct effects on supervised node classification, and moment-guided augmentation yields competitive results in graph contrastive learning.
Implications
The findings suggest that the moment-guided edge sampling framework can be applied to improve graph learning tasks by providing a method to control and analyze the structural properties of graphs. This could lead to advancements in various applications involving graph data, such as social network analysis, recommendation systems, and biological network modeling.
Structured Neural SDEs for Functional Calibration
Generative Models
Theory
Efficient ML
- Introduction of SLiSDE, a structured linear Neural SDE model for functional calibration.
- Achieves O(log T) parallel simulation efficiency through structured linear layers.
- Gated in-flow stacking enhances model expressivity while preserving computational efficiency.
- Incorporates a Girsanov tilt for improved performance on rare-event calibration tasks.
Read more
Structured Neural SDEs for Functional Calibration
Summary
This paper introduces SLiSDE, a novel family of Neural Stochastic Differential Equations (Neural SDEs) designed for functional calibration tasks, which involve fitting generative models to expectations of stochastic path functionals. Traditional Neural SDEs face challenges in simulation efficiency and stability, particularly when dealing with long horizons and path-dependent objectives. SLiSDE addresses these issues by employing structured linear stochastic layers that allow for parallel-in-time simulation, significantly reducing computational costs from O(T) to O(log T) through a parallel associative scan. The model incorporates gated in-flow stacking, where previous-layer paths influence the next layer's dynamics via learned gates, maintaining expressivity while ensuring each layer remains affine in its own state. Additionally, the authors introduce an optional Girsanov tilt that acts as a learned importance sampler, enhancing performance on rare-event functional calibration tasks. The paper provides theoretical guarantees regarding well-posedness, discretization error bounds, and the universality of the terminal laws generated by the model. Experimental results demonstrate that SLiSDE outperforms fully neural SDE baselines while maintaining stable importance weights and efficient simulation.
Methodology
The authors develop SLiSDE by structuring the Neural SDE into linear stochastic layers that can be evaluated in parallel. They implement gated in-flow stacking to allow previous-layer paths to modulate the next layer's dynamics. An optional Girsanov tilt is added to improve sampling efficiency for rare paths, with a closed-form likelihood-ratio correction to maintain accuracy.
Results
Experiments on functional calibration benchmarks indicate that SLiSDE outperforms traditional fully neural SDE models, demonstrating enhanced efficiency in simulation and stability in importance weights, particularly in scenarios dominated by rare paths.
Implications
The proposed SLiSDE model has significant implications for financial modeling, particularly in derivatives pricing and risk management, where efficient calibration to path functionals is crucial. Its ability to handle rare-event scenarios effectively could also benefit applications in various scientific fields requiring stochastic modeling.
What Should We Freeze? Guarded Freezing: Connectivity Shapes the Fine-Tuning of Pretrained Models
Computer Vision
NLP
Large Language Models
- Guarded Freezing optimizes the fine-tuning process by selectively freezing weights based on connectivity.
- Removal-value and drift-value are introduced as new metrics for selecting which weights to freeze.
- The proposed methods show improved retention of old-task accuracy compared to existing techniques.
- Experiments demonstrate the effectiveness of Guarded Freezing across different model architectures.
Read more
What Should We Freeze? Guarded Freezing: Connectivity Shapes the Fine-Tuning of Pretrained Models
Summary
This paper introduces the concept of Guarded Freezing, which addresses the challenges of fine-tuning pretrained models by selectively freezing weights based on connectivity. The author argues that simply freezing weights does not guarantee performance retention, as changes in trainable parameters can affect the inputs to frozen neurons. The paper proposes two scoring methods: removal-value, which evaluates the importance of neurons based on their contribution to the model's performance when incoming connections are cut, and drift-value, which assesses the disturbance caused by updates to weights when connections remain. The effectiveness of these methods is demonstrated through experiments on various architectures, including VGG-8 and transformers, showing that the proposed strategies can significantly improve retention of old-task accuracy while maintaining new-task performance. The results indicate that Guarded Freezing can outperform existing methods like DEFT and Fisher in specific scenarios, particularly in language models and vision transformers.
Methodology
The paper employs a theoretical analysis of weight freezing strategies, introducing removal-value and drift-value as metrics for evaluating the importance of neurons and weights. Experiments are conducted on VGG-8 and various transformer models to compare the performance of Guarded Freezing against existing methods like DEFT and Fisher. The evaluation focuses on retention of old-task accuracy and performance on new tasks.
Results
The results indicate that using removal-value allows for a 5.22 ± 0.51 percentage point improvement in old-task accuracy at 70% frozen weights compared to DEFT. In language models, drift-value outperforms adapted Wanda and RIA freezing scores, retaining more accuracy after 40 epochs. After 160 epochs on Qwen2.5-1.5B, drift-value retains 0.0433 ± 0.0102 more accuracy than static Fisher. In DINOv3 adaptations, drift-value significantly improves image retention accuracy from 0.097 to 0.440.
Implications
The findings suggest that fine-tuning strategies should consider the connectivity of neurons to optimize performance retention. Guarded Freezing could lead to more effective adaptation of pretrained models in various applications, particularly in scenarios where both old and new task performance is critical.
Geometry-Aware Operator Families for Structured Representation Learning
Graph Learning
Time Series
Theory
- Introduction of Geometry-Induced Operator Families (GIOF) for structured representation learning.
- GIOF transforms fixed geometric descriptors into a structured family of propagation operators.
- Dynamic selection of operators based on context enhances adaptability and performance.
- Theoretical guarantees established for stability, locality, and parameter efficiency.
Read more
Geometry-Aware Operator Families for Structured Representation Learning
Summary
This paper introduces Geometry-Induced Operator Families (GIOF), a novel framework for structured representation learning that leverages the geometry of latent representations to design adaptive propagation operators. The authors argue that existing neural architectures often rely on generic operator templates or geometry-specific constructions, which limits their flexibility and effectiveness. GIOF addresses this by transforming fixed geometric descriptors into a structured family of operators that can dynamically adapt based on the context of the data. The framework includes a mechanism for selecting appropriate operators from this family based on the current context, ensuring that the geometry of the latent space directly influences the learning process. The authors provide theoretical guarantees regarding parameter compression, identifiability, stability, and locality, and validate their approach through controlled experiments. Their results demonstrate that GIOF achieves the lowest mean absolute error (MAE) on benchmark datasets PEMS-BAY and METR-LA, outperforming the strongest baseline by 2.4% to 8.8% in scenarios with 30% missing sensors.
Methodology
The GIOF framework begins by defining a fixed geometry descriptor that specifies admissibility relations and geometric attributes. It then constructs reusable generator bases from these descriptors, which are combined through a context-dependent selector and adaptive propagation scale. The selected operator is realized through stable continuous-time propagation and a bottleneck residual layer, ensuring that the geometry of the latent space governs the admissible transformations.
Results
The experiments conducted on PEMS-BAY and METR-LA datasets showed that GIOF achieved the lowest mean MAE across all reported regional-outage settings, with improvements over the strongest baseline ranging from 2.4% to 8.8% when 30% of sensors were missing. The theoretical guarantees provided also support the effectiveness of the proposed framework.
Implications
The GIOF framework has the potential to enhance various applications in structured representation learning, particularly in scenarios where the geometry of data plays a critical role. This could lead to improved performance in tasks such as time series forecasting, graph learning, and other domains that require robust representation of structured data.
Saturation-Insensitive Dueling Bandits with General Function Approximation
Reinforcement Learning
Theory
Efficient ML
- Introduction of SI-CDB, a saturation-insensitive algorithm for contextual dueling bandits.
- The algorithm employs asymmetric arm selection to mitigate saturation effects in reward learning.
- Localized Eluder dimension analysis is used to derive regret bounds, showing improved performance over existing methods.
- SI-CDB achieves near-optimal regret bounds without the exponential dependence on reward scale.
Read more
Saturation-Insensitive Dueling Bandits with General Function Approximation
Summary
This paper addresses the challenge of saturation in contextual dueling bandits under the Bradley–Terry–Luce (BTL) preference model, where preference feedback becomes less informative as the reward model distinguishes between actions with high confidence. The authors introduce an algorithm named SI-CDB, which employs a heuristic for arm selection that mitigates saturation effects, allowing for saturation-insensitive reward learning. The algorithm operates by selecting a target arm based on empirical risk minimization and then choosing a comparator arm from a confidence set of plausible reward functions. This approach leads to a novel regret decomposition that facilitates a localized Eluder dimension analysis. The results demonstrate that SI-CDB achieves near-optimal performance for linear reward classes while avoiding the unfavorable dependence on the inverse-derivative factor commonly found in existing analyses. The paper provides a comprehensive theoretical framework that emphasizes the importance of two-arm regret analysis in enhancing single-arm performance in dueling bandits.
Methodology
The SI-CDB algorithm selects actions using a two-step process: first, it identifies a target arm based on empirical risk minimization, and then it constructs a confidence set of plausible reward functions to select a comparator arm. The selection maximizes a specific objective that facilitates a regret decomposition, allowing for a localized analysis of regret based on reward error terms.
Results
The theoretical analysis shows that the total regret of SI-CDB is polynomially dependent on the reward scale B for steps with small reward errors, while the number of steps with large errors grows logarithmically with the number of rounds T. This results in a cumulative regret bound that does not include terms scaling with both polynomial T and exponential B, marking a significant improvement over previous works.
Implications
The findings have potential applications in reinforcement learning scenarios, particularly in contexts where preference feedback is utilized, such as human feedback in machine learning and online preference-based fine-tuning of models. The saturation-insensitive approach could enhance the efficiency and effectiveness of learning algorithms in various domains.
Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra
Computer Vision
Time Series
Theory
- Introduction of a self-supervised encoder for auroral spectra that outperforms traditional supervised methods.
- Demonstrated effectiveness of a 1D Vision Transformer with masked autoencoder on unlabelled spectral data.
- Comparison of in-domain pretraining with existing astronomical models reveals limitations in transferability.
- Achieved significant improvements in classification metrics, indicating the potential of self-supervised learning in spectroscopy.
Read more
Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra
Summary
This paper presents a novel approach to auroral spectroscopy using self-supervised representation learning. The authors pretrain a 1D Vision Transformer with a masked autoencoder on a large dataset of 223,000 unlabelled auroral emission spectra recorded by the Auroral Spectrograph In Skibotn (ASIS). The pretrained model effectively recovers emission-line intensity ratios critical for diagnosing precipitating particles, achieving a correlation coefficient (R²) of 0.91 compared to 0.77 for an untrained control. When fine-tuned, the model surpasses the previous supervised auroral classifier, achieving a macro-average precision (macro-AP) of 88.5 versus 77.8. The study also evaluates the transferability of existing astronomical spectral foundation models, revealing that none can replace the in-domain pretrained model, with performance influenced by the spectral window overlap rather than the physical nature of the sources. This work highlights the potential of self-supervised learning in extracting meaningful representations from unlabelled spectral data, paving the way for improved classification and analysis of auroral emissions.
Methodology
The authors employed a masked autoencoder architecture within a 1D Vision Transformer framework to pretrain on unlabelled auroral emission spectra. They utilized a dataset of 329,704 calibrated spectra, applying a two-channel input approach to separate line and continuum information. The model was then fine-tuned and evaluated against both expert-designed features and previous classifiers.
Results
The pretrained model achieved an R² of 0.91 in recovering emission-line intensity ratios and a macro-AP of 88.5 in classification tasks, outperforming previous methods. It also exceeded the performance of a similar architecture trained from scratch by 0.159 with only 10% of the labels. The study found that existing astronomical models did not match the performance of the in-domain pretrained model.
Implications
This research suggests that self-supervised learning can significantly enhance the analysis of unlabelled spectral data, potentially revolutionizing the field of auroral spectroscopy. The findings indicate that domain-specific pretraining is essential for effective representation learning in new spectroscopic contexts.
PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem
Reinforcement Learning
Optimization
- Introduction of PORL, combining online pretraining with offline fine-tuning for JSSP.
- Utilization of a KL-divergence constraint to maintain policy stability during adaptation.
- Demonstrated superior performance of PORL over standalone offline RL and traditional scheduling methods.
- Reduced sensitivity to dataset quality, enhancing applicability in real-world scenarios.
Read more
PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem
Summary
This paper presents Pretrained Offline Reinforcement Learning (PORL), a novel hybrid approach designed to tackle the Job Shop Scheduling Problem (JSSP), a significant combinatorial optimization challenge in industrial settings. PORL integrates simulation-based online pretraining with offline fine-tuning on production-specific data. The methodology begins with online reinforcement learning to develop a general scheduling policy through interaction with a simulated environment, which allows for exploration of various scheduling strategies. Subsequently, this policy is refined using offline reinforcement learning, adapting it to specific production datasets without requiring further online exploration. A key innovation of PORL is the introduction of a KL-divergence-based policy constraint during the offline fine-tuning phase, which limits deviations from the pretrained policy. The authors evaluate PORL on JSSP instances characterized by distribution shifts and datasets generated from various behavioral policies. The results indicate that PORL consistently achieves lower optimality gaps compared to standalone offline RL methods and traditional scheduling baselines, particularly under conditions of reduced dataset quality. This suggests that PORL is less sensitive to the quality and coverage of offline data, making it a promising approach for real-world industrial scheduling scenarios where direct online exploration is impractical.
Methodology
The methodology involves two main phases: first, an online pretraining phase where a scheduling policy is learned through interaction with a simulated environment, followed by an offline fine-tuning phase where this policy is adapted using historical production data. The adaptation process employs a KL-divergence constraint to ensure the fine-tuned policy remains close to the pretrained policy.
Results
The evaluation shows that PORL achieves lower optimality gaps than both standalone offline RL approaches and traditional dispatching-rule baselines, especially under significant distributional shifts. Additionally, PORL maintains strong performance even when trained on datasets of lower quality, indicating its robustness in various conditions.
Implications
The findings suggest that PORL could be effectively applied in industrial scheduling environments, particularly where direct online exploration is not feasible. This approach may lead to improved scheduling efficiency and adaptability in dynamic production settings.
Rondo: Unsupervised Discovery of Recurring Temporal Structure
Time Series
- RONDO models recurring hierarchical structures in temporal data, addressing limitations of existing methods.
- The approach uses UnitAlign for reusable unit discovery and MotifFormer for higher-level motif modeling.
- RONDO shows significant improvements in recurring structure recovery, especially in limited-data and evolving contexts.
- The method allows for continual refinement of unit and motif vocabularies as new data emerges.
Read more
Rondo: Unsupervised Discovery of Recurring Temporal Structure
Summary
The paper introduces RONDO, an unsupervised approach designed to discover recurring hierarchical structures in continuous temporal streams. Traditional methods often treat recurring patterns at different temporal scales as independent, which can lead to fragmented discoveries. RONDO addresses this by constructing vocabularies of reusable units and their compositions, allowing for the capture of shared structures across complex temporal patterns. The approach consists of two main components: UnitAlign, which learns a vocabulary of recurring units, and MotifFormer, which models the temporal organization of these units into higher-level motifs. This hierarchical modeling enables RONDO to refine and expand its discoveries as new data arrives, making it particularly effective in limited-data and continual-stream scenarios. Evaluations across diverse datasets demonstrate that RONDO consistently outperforms existing unsupervised baselines, showcasing its ability to adapt and evolve with the data stream.
Methodology
RONDO employs a two-component framework: UnitAlign for discovering and refining reusable local units from temporal data, and MotifFormer for modeling the organization of these units into higher-level motifs. This hierarchical approach allows for the explicit separation and reuse of local structures across different motifs, facilitating continual updates as new data is processed.
Results
Evaluations indicate that RONDO outperforms existing unsupervised methods in recovering recurring structures across various temporal datasets. The improvements are particularly notable when data is initially limited and subsequently expanded. Ablation studies confirm the contributions of both UnitAlign and MotifFormer to the overall performance enhancements.
Implications
RONDO's ability to discover and adaptively refine recurring patterns in temporal data has significant implications for applications in areas such as human motion analysis, activity recognition, and intelligent systems that operate on long, unlabeled temporal streams. It reduces the need for extensive human annotations and supports scalable behavior understanding.