AI-generated summaries
Today's ML research,
without the noise.
Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.
42
Papers today
8h
Update frequency
7
Days of history
Learning Provable Neural Network Observer for Uncertain Dynamical Systems
Theory
Robotics
Optimization
- Introduction of a two-stage training framework for neural network observers that separates learning from stability certification.
- Point-guided Lyapunov pre-training enhances estimation accuracy and local stability before LMI fine-tuning.
- The framework provides theoretical guarantees for local stability and probabilistic coverage.
- Empirical results show significant improvements in training speed and tracking accuracy compared to existing methods.
Read more
Learning Provable Neural Network Observer for Uncertain Dynamical Systems
Summary
This paper addresses the challenge of estimating states and external disturbances in uncertain dynamical systems, particularly in safety-critical applications. The authors propose a novel two-stage training framework for neural network observers that ensures provable stability. The first stage, point-guided Lyapunov pre-training, focuses on achieving high estimation accuracy and local stability over sampled states. The second stage involves fine-tuning the model using Linear Matrix Inequality (LMI) constraints to certify global Lyapunov stability. This approach effectively decouples the learning process from the stability certification, making it scalable for large networks. The authors provide theoretical guarantees for local stability and probabilistic coverage within a specified error-state domain. Empirical evaluations demonstrate that their method trains faster than traditional LMI-based approaches and achieves superior tracking accuracy across various systems, including nonlinear control benchmarks and X-29 aircraft ablations.
Methodology
The proposed methodology consists of a two-stage training process. The first stage involves point-guided Lyapunov pre-training, which uses a sampling-based Lyapunov loss to achieve local stability and high estimation performance. The second stage applies LMI fine-tuning to ensure global Lyapunov stability, avoiding the computational burden of full-scale LMI optimization during training.
Results
The experiments conducted on various benchmarks, including a quadrotor UAV, indicate that the proposed LMI-certified neural network observers train significantly faster than traditional methods and demonstrate robust generalization across different systems, achieving improved tracking accuracy.
Implications
This work has significant implications for the design of robust control systems in safety-critical applications, enabling the use of complex neural network observers while ensuring formal stability guarantees. It opens avenues for further research in scalable certification methods for deep learning architectures in control applications.
When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization
Optimization
Theory
- The optimal preconditioning exponent decreases linearly with log10 of the learning rate.
- Negative exponents are not universally better; they are effective in high-step-size regimes.
- Source-validation and cross-environment optimal exponents differ, indicating a model-selection conflict.
- Lower preconditioning exponents reduce reliance on spurious and noise features.
Read more
When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization
Summary
This paper investigates the interaction between the preconditioning exponent in adaptive optimizers and the global learning rate, particularly in the context of cross-environment generalization. The authors conduct a controlled study using a paired four-environment classification problem, examining how different preconditioning exponents (ranging from -0.5 to 0.5) and learning rates (from 10^-4 to 10^-2) affect model performance. The findings reveal that the exponent that maximizes cross-environment accuracy decreases almost linearly with the logarithm of the learning rate. At higher learning rates, the optimal exponent enters the negative region, indicating that negative exponents can be beneficial under certain conditions. The study highlights a model-selection conflict where source-domain validation prefers positive exponents, while cross-environment criteria favor negative ones. The results suggest that lower preconditioning exponents reduce reliance on spurious features and noise, emphasizing the need for careful consideration of optimizer settings in the presence of distribution shifts.
Methodology
The authors performed a controlled cross-environment study using a paired four-environment classification problem. They systematically varied the preconditioning exponent and learning rate across 420 source-training runs, isolating the effects of the optimizer on feature allocation by evaluating a frozen model checkpoint on multiple test environments.
Results
The study found that the exponent maximizing mean cross-environment accuracy decreases almost linearly with log10 of the learning rate. At a learning rate of 10^-2, source-validation selection favored positive exponents, while cross-environment criteria preferred negative exponents. Checkpoint decomposition indicated that lower preconditioning exponents reduced the learned weight ratios of spurious to stable features.
Implications
The findings suggest that the choice of preconditioning exponent in adaptive optimizers should be informed by the learning rate and the specific environment in which the model will be deployed. This has implications for improving model robustness in scenarios with distribution shifts, particularly in applications requiring generalization across different environments.
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
Reinforcement Learning
Graph Learning
- HySTAR introduces a stable credit assignment framework using anchored hypergraphs.
- The method separates adaptive representation learning from value decomposition.
- Extensive experiments show HySTAR outperforms MAPPO and other baselines in various MARL scenarios.
- The framework effectively addresses structural target drift in cooperative multi-agent settings.
Read more
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
Summary
The paper presents HySTAR, a novel framework designed to enhance credit assignment in cooperative multi-agent reinforcement learning (MARL) under conditions of partial observability and shared rewards. Traditional approaches, such as MAPPO, struggle with credit assignment due to their reliance on a single global value, which can lead to inconsistencies known as structural target drift. HySTAR addresses this issue by decoupling the credit assignment process from adaptive representation learning. It employs an anchored overlapping sparse hypergraph as a stable decomposition scaffold, allowing for a consistent representation of agent interactions over time. The framework utilizes a spatiotemporal encoder to adaptively model agent interactions based on current spatial and temporal contexts. The authors demonstrate the effectiveness of HySTAR through extensive experiments across multiple benchmarks, showing significant performance improvements over existing methods, particularly in challenging scenarios. The results indicate that HySTAR not only stabilizes credit assignment but also enhances the learning efficiency of agents in cooperative settings.
Methodology
HySTAR employs a MAPPO-based architecture that integrates an anchored overlapping sparse hypergraph to provide a stable credit assignment basis. It utilizes a spatiotemporal encoder to adaptively represent agent interactions, allowing for high-order value decomposition and agent-specific advantage calculation. The framework is designed to maintain a consistent decomposition structure while adapting to changing agent interactions and roles.
Results
HySTAR achieved a 16.7% relative improvement over MAPPO and a 15.6% improvement over HYGMA on the most challenging SMAC settings. It ranked first in all six GRF scenarios and reduced convergence epochs in the Traffic Junction environment by up to 40.2% compared to MAGIC. Additionally, it obtained the highest episode rewards in all MPE tasks, demonstrating its effectiveness in various MARL benchmarks.
Implications
The HySTAR framework has significant implications for improving the efficiency and stability of cooperative multi-agent systems, particularly in environments where agents must adapt to dynamic interactions and partial observability. Its approach to credit assignment could be applied to various domains, including robotics, autonomous systems, and complex multi-agent simulations.
Neural Ideals and Neural Codes: An Algebraic Framework for Neural Network Classification and Feature Interpretation
Theory
Interpretability
- Introduction of an allowed code space associated with target classes in neural networks.
- Adaptation of neural ideals theory for artificial neural network classification.
- Establishment of a classification–ideal correspondence for class membership characterization.
- Development of algorithms for constructing neural codes and deriving neural ideals.
Read more
Neural Ideals and Neural Codes: An Algebraic Framework for Neural Network Classification and Feature Interpretation
Summary
This paper addresses the challenge of understanding the features captured by hidden layers in neural networks, proposing an algebraic framework that connects neural network classification with the theory of neural ideals. The authors establish a correspondence between neural networks and neural ideals, develop algorithms for computing these ideals, and prove a stabilization theorem that ensures the invariance of neural ideals with sufficient training samples. The framework allows for the identification and interpretation of features captured by hidden-layer neurons. The practical application of this framework is demonstrated on the MNIST digit dataset, showcasing its effectiveness in analyzing hidden representations. Additionally, an interactive software tool is developed to visualize the features captured by each neuron, enhancing interpretability in neural networks.
Methodology
The authors utilize concepts from commutative algebra and algebraic geometry to develop an algebraic framework for neural networks. They introduce allowed code spaces, construct class-specific neural ideals from hidden-layer activation patterns, and establish algorithms for computing these ideals. Theoretical results include a stabilization theorem that guarantees the invariance of neural ideals with sufficient training data.
Results
The proposed framework successfully demonstrates the relationship between neural network classification and polynomial ideal theory. The algorithms developed allow for effective computation of neural codes and ideals, and the application of the framework on the MNIST dataset illustrates its utility in feature interpretation. The interactive software tool enhances the practical application of the theoretical findings.
Implications
This work provides a mathematical foundation for analyzing neural networks, potentially improving interpretability and understanding of hidden representations. The framework and tools developed can aid researchers and practitioners in feature analysis, leading to better insights into neural network decision-making processes.
PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control
Reinforcement Learning
Robotics
Optimization
- Introduces PolicyAttention, a method that integrates softmax attention with Policy Mirror Descent for RL.
- Addresses the closed-loop gap in reinforcement learning where model outputs affect future inputs.
- Demonstrates that the proposed method retains performance under repeated control and environment shifts.
- Achieves median normalized returned-policy losses significantly lower than existing methods.
Read more
PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control
Summary
The paper introduces PolicyAttention, a novel approach that utilizes softmax attention to implement Policy Mirror Descent (PMD) for closed-loop control in reinforcement learning (RL). The authors address the challenge of adapting reinforcement learning models to dynamic environments where the output of the model influences future inputs. They propose a causal softmax decoder that effectively realizes PMD updates while accounting for errors in actor-environment interactions. The study demonstrates that the proposed method can maintain performance under repeated control tasks and adapt to shifts in the environment without retraining. Empirical results show that PolicyAttention achieves significantly lower normalized returned-policy losses compared to existing adaptations, highlighting its effectiveness in closed-loop control scenarios. The paper emphasizes the importance of local algorithmic fidelity and closed-loop reliability in reinforcement learning applications.
Methodology
The authors develop a fixed causal-softmax decoder that implements PMD updates and evaluates the policy using a one-step critic. They introduce a returned-policy theorem to manage residual errors and connect their construction to standard Transformer operations. The methodology includes a pre-LayerNorm and final-LayerNorm compilation, along with a scoped effective-coordinate result to align with the PMD training objective.
Results
PolicyAttention achieves median normalized returned-policy losses of 0.0193 at S = 8 and 0.0273 at S = 16, which are 18-28 times lower than the losses from Liang–Lai and Algorithm Distillation adaptations. The method demonstrates robust performance across five fresh runs and maintains effectiveness under no-retraining environment shifts.
Implications
The findings suggest that softmax attention can serve as a viable policy-improvement operator in reinforcement learning, potentially enhancing the adaptability and reliability of RL models in dynamic environments. This could lead to advancements in various applications, including robotics and adaptive control systems.
ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers
Efficient ML
Computer Vision
- ENAS is a CPU-only NAS framework that enhances accessibility for TinyML model development.
- It features a flexible cell-based search space and a three-stage hybrid search strategy.
- ENAS achieves significant search-time speedups of 2.41× and 1.70× on two benchmark datasets.
- The framework maintains competitive accuracy while reducing peak activation RAM usage.
Read more
ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers
Summary
The paper introduces ENAS, a hardware-aware Neural Architecture Search (NAS) framework specifically designed for TinyML applications on resource-constrained microcontrollers. Unlike existing NAS frameworks that require GPU acceleration, ENAS operates efficiently in CPU-only environments, making it accessible for developers working with limited resources. The framework incorporates a static feasibility check, a cell-based search space that supports various architectural blocks, and a three-stage hybrid search strategy that enhances search efficiency. ENAS was evaluated on two benchmarks, Visual Wake Words and Melanoma Cancer, across eight microcontrollers with varying memory capacities. The results demonstrate significant speedups in search time compared to the NanoNAS framework while maintaining competitive accuracy. The study also highlights the lower peak activation RAM usage of ENAS-selected models, which is crucial for microcontroller deployment. Overall, ENAS provides a robust solution for optimizing TinyML models in constrained environments.
Methodology
ENAS employs a three-stage hybrid search strategy consisting of random sampling, top-K candidate selection, and local mutation. It utilizes a lightweight analytical feasibility estimator for RAM, Flash, and MACC constraints, which eliminates the need for extensive TFLite conversions during candidate evaluation. The framework supports a rich cell-based search space that includes standard, depthwise-separable, and bottleneck blocks with optional skip connections.
Results
ENAS achieved mean search-time speedups of 2.41× on the Visual Wake Words dataset and 1.70× on the Melanoma Cancer dataset compared to the greedy CPU-only NanoNAS baseline. It maintained competitive test accuracy with only minor reductions (0.79 pp and 1.31 pp, respectively) and demonstrated significantly lower peak activation RAM usage (approximately 0.38×) at matched accuracy levels.
Implications
The ENAS framework enables efficient deployment of deep learning models on microcontrollers, facilitating the development of battery-powered IoT devices with always-on perception capabilities. Its open-source nature encourages further exploration and optimization in the field of TinyML.
Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design
Robotics
Optimization
Reinforcement Learning
- Introduction of Morphogene as a high-level latent blueprint for body-brain coordination.
- GeCode reformulates morphology-control co-design as gene-driven exploration in a compact latent space.
- Demonstrated significant performance improvements over existing methods, achieving 2.5× faster convergence.
- Morphogene allows for coherent changes in morphology and control through adaptive conditioning.
Read more
Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design
Summary
This paper introduces a novel approach to morphology-control co-design in embodied agents, addressing the limitations of existing methods that treat morphology and control as separate entities. The authors propose Morphogene, a compact latent blueprint inspired by biological genes, which enables explicit coordination between an agent's body structure and control policy. By utilizing AdaConcat, Morphogene conditions both morphology and control generation at the limb level, allowing variations to induce coordinated changes in both components. The proposed framework, GeCode, reformulates the co-design process as exploration within the Morphogene space, facilitating efficient local refinement and global exploration of body-brain designs. Extensive experiments demonstrate that GeCode significantly outperforms state-of-the-art methods, achieving faster convergence and higher performance across diverse 2D and 3D tasks.
Methodology
The methodology involves the introduction of Morphogene, which serves as a shared latent representation for both morphology and control networks. The AdaConcat mechanism is employed to adaptively condition these networks at the limb level. GeCode utilizes a small set of Morphogenes as design anchors and combines stochastic policy rollouts with a Morphogene-proximity reward to explore the design space efficiently. Performance-guided updates are used to refine the Morphogene anchors towards higher-performing designs.
Results
GeCode consistently outperformed existing state-of-the-art methods, achieving an average of 2.5 times faster convergence and 69.48% higher task performance while maintaining less than 4.0% additional computational overhead.
Implications
The findings suggest that integrating morphology and control through a shared latent representation can lead to more efficient and effective design of embodied agents. This approach has potential applications in robotics, where optimized body structures and control policies can enhance performance in complex tasks.
Learning coarse-step dynamics and internal mechanical response with graph networks
Graph Learning
Robotics
Time Series
- Introduces Newmark-β-DGN, a graph neural network framework for inferring mechanical responses from kinematic data.
- Combines semi-implicit updates and operator-weighted virtual hubs to enhance prediction accuracy over coarse time scales.
- Achieves long-horizon predictions in various physical systems without direct supervision of mechanical quantities.
- Demonstrates the ability to infer forces and recover stiffness structures from observed trajectories.
Read more
Learning coarse-step dynamics and internal mechanical response with graph networks
Summary
This paper presents Newmark-β-DGN, a novel graph neural network framework designed to infer unobserved mechanical quantities from kinematic observations in physical systems. The authors address the challenge of predicting coarse-step dynamics where mechanical responses evolve between observations. The framework integrates a semi-implicit update inspired by the Newmark-β method, which utilizes learned momentum fluxes and matrix-valued response operators to advance the state over observed intervals. Additionally, it employs an operator-weighted virtual hub for system-wide coupling through sparse connections. The model demonstrates its efficacy across various applications, including deformable beams, human motion, and protein dynamics, achieving long-horizon predictions even when explicit learned simulators fail. Notably, the framework infers forces from walking kinematics that align with independently derived joint moments, and it recovers the spatial and directional structure of stiffness in a beam without requiring direct supervision of mechanical quantities during training. This approach effectively links coarse-step prediction to the inference of mechanical quantities that were not observed during the training phase.
Methodology
The Newmark-β-DGN framework employs a graph neural network architecture that integrates a semi-implicit update mechanism inspired by the Newmark-β method and an operator-weighted virtual hub for coupling across the system. This allows for the propagation of learned momentum fluxes and response operators, facilitating the inference of mechanical quantities from kinematic observations.
Results
The framework successfully predicts long-term dynamics across different scenarios, including deformable beams and human motion. It infers forces from walking kinematics that correlate with actual joint moments and recovers the stiffness structure of a beam, demonstrating its capability to operate without direct mechanical supervision.
Implications
The findings suggest that Newmark-β-DGN can be applied in various fields such as biomechanics, structural engineering, and robotics, where understanding unobserved mechanical dynamics is crucial. This approach could enhance predictive modeling in systems where direct measurement of mechanical quantities is challenging.
Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
Optimization
Theory
Robotics
- Introduces a penalized DRO formulation that incorporates Wasserstein penalties for adversarial distribution shifts.
- Demonstrates that optimal transport maps for adversarial training are cyclically monotone, which is crucial for effective learning.
- Presents Multi-start Particle Ascent (MPA) and ICNN-based methods to enforce cyclical monotonicity in adversarial training.
- Empirical results show significant improvements in robustness and generalization compared to standard adversarial training methods.
Read more
Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
Summary
This paper addresses the challenges of distributionally robust optimization (DRO) in the context of adversarial training, particularly focusing on the difficulties posed by nonconvex loss functions. The authors propose a penalized DRO formulation where the adversary can choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. They reformulate the adversary's problem as an optimization problem over transport maps, demonstrating that optimal maps are cyclically monotone. The paper critiques standard adversarial training methods for violating cyclical monotonicity and introduces two novel approaches: Multi-start Particle Ascent (MPA) and the use of input-convex neural networks (ICNNs) to parameterize adversarial maps. Experimental results on robust regression, image classification, and robust control tasks indicate that these methods outperform standard adversarial training and state-of-the-art baselines, achieving enhanced robustness and generalization under distribution shifts.
Methodology
The authors reformulate the adversary's problem as an optimization problem over transport maps, ensuring that the maps are cyclically monotone. They introduce Multi-start Particle Ascent (MPA) to solve the adversary's problem through parallel gradient ascent and optimal reassignment. Additionally, they utilize input-convex neural networks (ICNNs) to parameterize the transport maps, inherently enforcing cyclical monotonicity.
Results
The proposed methods, MPA and the ICNN-based approach, consistently outperform standard adversarial training techniques and state-of-the-art baselines across various tasks, including CIFAR-10 image classification and robust control tasks. The ICNN-based method particularly excels in achieving robustness against adversarial perturbations.
Implications
The findings suggest that incorporating optimal transport principles into adversarial training can significantly enhance the robustness of machine learning models, making them more resilient to distribution shifts and adversarial attacks. This has potential applications in fields requiring high reliability, such as autonomous systems and security-sensitive applications.
Evaluating the accuracy of KV cache reuse techniques
NLP
Large Language Models
Efficient ML
- Current evaluations of KV cache reuse techniques often inflate accuracy metrics.
- Existing datasets lack the necessary dynamics to evaluate KV cache reuse effectively.
- The proposed evaluation methodology isolates accuracy loss attributable to KV cache reuse.
- Boxoffice generates datasets that ensure meaningful evaluation of KV cache strategies.
Read more
Evaluating the accuracy of KV cache reuse techniques
Summary
This paper addresses the challenges in evaluating the accuracy of position-independent Key-Value (KV) cache reuse techniques in retrieval-augmented generation (RAG) systems. The authors argue that current evaluation methods often inflate reported effectiveness by failing to accurately measure the loss of accuracy due to cache reuse. They identify two main issues: the reliance on aggregate accuracy metrics that include irrelevant queries and the structural properties of existing datasets that do not support thorough evaluation of KV cache reuse. To overcome these limitations, the authors propose a new evaluation methodology that isolates accuracy loss by focusing on a meaningful subset of queries and introduce Boxoffice, a tool for generating tailored datasets that reflect complex KV cache reuse patterns. Their findings reveal that up to 42% of reported F1 scores in existing methods stem from queries that the baseline cannot answer, highlighting the need for more rigorous evaluation standards in this area.
Methodology
The authors developed an evaluation methodology that focuses on a meaningful subset of queries, specifically those where the baseline model answers correctly with full prefill, cannot answer without context, and where the answer space is non-trivial. They also created Boxoffice, a tool that programmatically generates datasets designed to test KV cache reuse under controlled conditions, ensuring the presence of cross-query chunk reuse and nuanced reuse patterns.
Results
The evaluation methodology revealed that a significant portion of the F1 scores reported by prior work (up to 42%) came from queries that the full-prefill baseline could not answer. The Boxoffice tool demonstrated that the same reuse method could yield drastically different F1 scores (0.98 or 0.00) depending on the context in which a chunk was cached, indicating the importance of context in evaluating KV cache reuse.
Implications
The findings suggest that more accurate evaluation methods are necessary for assessing KV cache reuse techniques, which could lead to improved performance in RAG systems. The Boxoffice tool could facilitate future research by providing a standardized way to test and compare KV cache strategies, ultimately enhancing the efficiency and accuracy of large language models.
Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System
Reinforcement Learning
Theory
Interpretability
- Introduction of a PR-driven objective for generating UAPs against DRL-based IDS.
- Development of PX-UAP, which utilizes XAI for feature attribution to enhance perturbation strategies.
- Theoretical justification for PX-UAP with a focus on entropy-regularized perturbation allocation.
- Extensive experimental validation showing superior performance of PX-UAP over state-of-the-art methods.
Read more
Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System
Summary
This paper addresses the vulnerability of Deep Reinforcement Learning (DRL)-based Intrusion Detection Systems (IDS) to adversarial attacks, specifically Universal Adversarial Perturbations (UAPs). The authors introduce a novel framework called Probabilistic Robustness (PR)-based UAP, which integrates a probabilistic measure of adversarial impact into the generation of UAPs. This approach quantifies the likelihood of misclassification across the input space, aligning with the universality objective of UAPs. Furthermore, the paper presents PX-UAP, which enhances the attack performance by leveraging Explainable Artificial Intelligence (XAI) to guide perturbation shaping under realistic domain constraints. The authors provide a theoretical analysis of PX-UAP, establishing a constrained universal evasion objective and an entropy-regularized perturbation scheme. Extensive experiments demonstrate that PX-UAP consistently outperforms existing UAP methods in terms of attack effectiveness against DRL-based IDS.
Methodology
The authors propose a PR-based UAP framework that quantifies the adversarial impact on DRL-based IDS. They introduce PX-UAP, which incorporates XAI to shape perturbations based on feature relevance. The methodology includes a theoretical analysis of the perturbation scheme, ensuring it adheres to domain constraints while maximizing attack effectiveness.
Results
The experiments conducted reveal that PX-UAP significantly outperforms existing UAP methods in terms of attack effectiveness, demonstrating its capability to effectively compromise DRL-based IDS under realistic conditions.
Implications
The findings suggest that integrating probabilistic measures and explainability into adversarial attack strategies can enhance the effectiveness of attacks on IDS, raising concerns about the security of DRL-based systems. This work could inform future research on improving IDS robustness against adversarial threats.
Online Learning via Learned Latent Bayesian Tracking
Time Series
Efficient ML
Optimization
- AURA framework enables rapid online learning through a learned low-dimensional latent representation.
- Utilizes extended Kalman filtering for efficient single-step updates in the latent space.
- Meta-learns the latent dynamics and lifting map from offline data for improved adaptation performance.
- Demonstrates significant improvements in adaptation speed and accuracy in non-stationary environments.
Read more
Online Learning via Learned Latent Bayesian Tracking
Summary
This paper addresses the challenge of online learning in non-stationary environments, where models must adapt quickly to changing data distributions under strict computational constraints. The authors propose a novel framework called Adaptive Update through Representation Adaptation (AURA), which leverages a learned low-dimensional latent state-space model to facilitate efficient online adaptation of model parameters. By using extended Kalman filtering (EKF) in this latent space, AURA allows for single-step updates that maintain model expressiveness while significantly reducing computational costs. The framework is meta-learned from offline non-stationary data trajectories, optimizing the latent dynamics and lifting map for effective online adaptation. Empirical evaluations demonstrate AURA's superiority in adaptation speed, accuracy, and computational efficiency compared to existing online learning and Bayesian filtering methods, particularly in applications such as neural wireless receivers and non-stationary image classification.
Methodology
The methodology involves the development of the AURA framework, which learns a low-dimensional latent state-space model from offline data. Online adaptation is performed using extended Kalman filtering in this latent space, with a learned lifting map to reconstruct full model parameters. This approach decouples the computational complexity of filtering from the expressiveness of the model.
Results
AURA was empirically validated in two domains: online adaptation of neural wireless receivers under time-varying channels and non-stationary image classification tasks. The results indicate that AURA consistently achieves accurate adaptation with a single low-complexity update step per sample, outperforming existing online learning and Bayesian filtering baselines in terms of speed and accuracy.
Implications
The findings suggest that AURA can be effectively applied in various real-time applications requiring rapid adaptation to changing data distributions, such as adaptive signal processing, wireless communications, and dynamic vision systems. The framework's ability to efficiently handle high-dimensional models could lead to advancements in online learning methodologies.
Block Sparse Attention with Log-Linear Complexity
NLP
Large Language Models
Efficient ML
- PISA reduces the computational complexity of block-sparse attention from quadratic to log-linear.
- The method employs a hierarchical Top-K selection strategy to efficiently narrow down key candidates.
- Hardware-aware Triton kernels are developed for optimized training and inference.
- PISA shows competitive performance on commonsense reasoning benchmarks and superior results on retrieval tasks.
Read more
Block Sparse Attention with Log-Linear Complexity
Summary
This paper addresses the computational bottleneck of self-attention mechanisms in language models, particularly when scaling to long contexts. The authors propose a novel block-sparse attention mechanism called PISA (Pyramid Sparse Attention), which utilizes a hierarchical Top-K selection strategy to efficiently select relevant key blocks. By constructing a coarse-to-fine hierarchy of key representations, PISA reduces the complexity of block selection from quadratic to log-linear, achieving an overall complexity of O(N log N) for sequence length N. The method employs LogSumExp scoring to evaluate a bounded candidate set at each level of the hierarchy, allowing for efficient narrowing down of candidates. The authors implement hardware-aware Triton kernels to optimize both training and inference, minimizing memory traffic and overhead. Experimental results demonstrate that PISA achieves comparable performance to conventional methods on commonsense reasoning tasks while outperforming them on retrieval tasks, indicating its effectiveness for long-context modeling.
Methodology
The authors introduce PISA, which constructs a multilevel representation of key blocks through pooling. The selection process begins at the coarsest level, applying LogSumExp scoring to a bounded set of candidates, and iteratively narrows down to finer levels. This coarse-to-fine approach allows each query to evaluate only a limited number of promising blocks at each level, significantly reducing the selection complexity.
Results
PISA achieves comparable performance to conventional block-sparse attention methods on commonsense reasoning benchmarks while delivering better results on retrieval tasks. The method demonstrates efficiency in handling sequence lengths of up to 256K tokens, showcasing the advantages of hierarchical routing over traditional single-level selection.
Implications
The proposed PISA mechanism has significant implications for the development of efficient language models capable of processing long contexts, potentially enhancing applications in natural language processing and other areas requiring large-scale data handling.
Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model
Efficient ML
NLP
Theory
- Conducted a reliability audit of a commercial System-1 model on biosecurity-relevant benchmarks.
- Demonstrated significant sensitivity of model predictions to the order of answer options.
- Introduced selective permutation averaging to improve accuracy on low-confidence items.
- Established that the model is reasonably well-calibrated but accuracy varies significantly by task.
Read more
Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model
Summary
This paper presents a reliability audit of a commercial non-generative 'System-1' model, specifically the JEV model, across 6,020 biosecurity and biology-relevant benchmark items. The authors investigate the model's accuracy, calibration, error detection, selective prediction, and sensitivity to the order of answer options. The findings reveal that the model's accuracy is highly task-dependent, with a pooled expected calibration error of 0.034 and an area under the receiver operating characteristic curve (AUROC) of 0.820 for top-1 probability predictions. However, accuracy significantly degrades on weaker tasks. A notable discovery is that 37.4% of items in the WMDP-Cyber benchmark yield different answers when answer options are cyclically rotated, indicating substantial item-level instability. The authors propose a method of averaging probabilities across rotations, which improves accuracy by 3.8 percentage points, particularly when applied selectively to low-confidence items. This work highlights the need for systematic evaluation of low-cost probabilistic models in biosecurity contexts, establishing their reliability and robustness for potential applications in AI safety pipelines.
Methodology
The authors conducted a systematic evaluation of the JEV model by querying it on 6,020 multiple-choice items from the Weapons of Mass Destruction Proxy (WMDP) and LAB-Bench subtasks. They assessed accuracy, calibration, and error detection, and performed controlled experiments to analyze the impact of answer option order on model predictions. The study also involved averaging probabilities across multiple rotations of answer options to enhance accuracy.
Results
The audit revealed that the JEV model is reasonably well-calibrated with a pooled expected calibration error of 0.034 and a pooled AUROC of 0.820. However, accuracy was found to be strongly dependent on the specific task, with significant degradation on weaker tasks. The controlled experiments showed that 37.4% of items had different answers due to answer option order, and selective permutation averaging improved accuracy by 3.8 percentage points.
Implications
The findings suggest that while non-generative System-1 models can be cost-effective components in biosecurity pipelines, their reliability and sensitivity to input variations must be carefully considered. The proposed methods for improving accuracy could enhance the deployment of such models in safety-critical applications.
Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection
Time Series
- Introduces a Bayesian Tensor Autoencoder that preserves the tensor structure of multi-dimensional time series.
- Combines reconstruction-based and prediction-based approaches through a Physics-informed Predictive Prior.
- Utilizes Bayesian fusion to enhance modeling capabilities for normal data.
- Incorporates physical laws to prevent over-generalization in anomaly detection.
Read more
Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection
Summary
This paper addresses the challenge of anomaly detection in multi-dimensional time series data, which are often represented as tensors. Traditional methods for anomaly detection typically focus on uni- or multi-variate time series, requiring reshaping operations that disrupt intrinsic correlations within the data. The authors propose a novel approach that combines a Bayesian Tensor Autoencoder (AE) with a Physics-informed Predictive Prior (PP) to enhance anomaly detection performance. The proposed method leverages both reconstruction-based and prediction-based AEs by integrating a predictive prior that respects the tensor structure of the data. This is achieved through a Bayesian fusion approach that improves the model's capability to represent normal data while incorporating physical laws to mitigate over-generalization. The authors also introduce tailored training and testing strategies for their framework, which effectively manage randomness during training and simplify testing processes. Experimental results on real-world datasets demonstrate the effectiveness of the proposed method in detecting anomalies while preserving the intrinsic correlations of the multi-dimensional time series.
Methodology
The authors developed a Bayesian Tensor Autoencoder framework that integrates a Physics-informed Predictive Prior. This framework employs a Bayesian fusion approach to enhance the modeling of normal data while maintaining the tensor structure of the input data. Tailored training and testing strategies are implemented to introduce randomness during training and simplify the testing phase.
Results
The proposed method showed significant improvements in anomaly detection performance on various real-world datasets compared to traditional methods, effectively preserving the intrinsic correlations within the multi-dimensional time series.
Implications
The findings suggest that the proposed framework can be applied to various domains requiring anomaly detection in multi-dimensional time series, such as traffic monitoring and environmental data analysis, potentially leading to more accurate and timely identification of abnormal events.
Differentiable RNA Secondary Structure Extraction for Deep Learning
Theory
Optimization
- Introduces a novel SDSM normalization algorithm for direct output of base-pairing probability matrices.
- Analyzes the effect of training-extraction congruence on RNA secondary structure prediction.
- Demonstrates that the choice of extraction method significantly impacts prediction performance.
- Shows that the SDSM model outperforms traditional methods and baselines in structure prediction accuracy.
Read more
Differentiable RNA Secondary Structure Extraction for Deep Learning
Summary
This paper addresses the challenge of converting weight matrices produced by deep learning models into interpretable RNA secondary structures, a process termed structure extraction. The authors analyze the impact of training-extraction congruence on prediction performance by comparing four extraction algorithms: a Nussinov-like dynamic programming method, maximum-weight graph matching, and greedy extraction methods used by SPOT-RNA and RiNALMo. They introduce a novel symmetric doubly stochastic matrix (SDSM) normalization algorithm that allows models to output base-pairing probability matrices directly, eliminating the need for a separate extraction step. The study evaluates these methods on outputs from the pretrained RiNALMo model and three newly trained toy models, including one utilizing the SDSM normalization. Results indicate that the SDSM model consistently outperformed the binary cross-entropy baseline and produced outputs closest to the ground truth, suggesting that SDSM normalization is a promising alternative to traditional structure extraction methods.
Methodology
The authors conducted experiments using the ArchiveII dataset, comparing four structure extraction methods on outputs from a pretrained RiNALMo model and three toy models: a differentiable Nussinov-like model, a binary cross-entropy baseline, and a model using the SDSM normalization. They evaluated the performance of these models based on their ability to produce accurate base-pairing probability matrices.
Results
The SDSM model exhibited the strongest overall performance, outperforming the binary cross-entropy baseline across all extraction algorithms and producing outputs that were closest to the ground truth. The study found that the effectiveness of extraction methods is highly dependent on the training method used.
Implications
The findings suggest that integrating differentiable extraction methods into the training process can enhance the performance of RNA secondary structure prediction models. This approach may lead to more accurate predictions in computational biology, particularly for understanding non-coding RNAs and their functions.
Benchmarking the Connectomes of Caenorhabditis elegans within the Reservoir Computing Framework
Graph Learning
Theory
- C. elegans connectomes were benchmarked using the reservoir computing framework.
- Randomized null models often outperformed biological connectomes in computational tasks.
- Performance varied significantly based on reservoir configuration and connectome derivation methods.
- Connectomes from different ages produced varying results without clear trends.
Read more
Benchmarking the Connectomes of Caenorhabditis elegans within the Reservoir Computing Framework
Summary
This paper investigates the computational performance of the connectomes of Caenorhabditis elegans (C. elegans) using the reservoir computing framework. C. elegans is notable for having its entire neural connectome mapped, providing a unique opportunity to analyze biological neural networks through computational methods. The authors implemented various connectomes derived from different ages of C. elegans and through three distinct measurement techniques of inter-cellular connections as echo state networks (ESNs). The study aimed to benchmark these connectomes against randomized null models across several neuro-inspired tasks. Results indicated that the biological wiring and configurations based on biological knowledge did not consistently outperform randomized models, suggesting that the performance is highly dependent on the specific configuration of the reservoir and the method of connectome derivation. Additionally, variations in performance were observed across connectomes from different ages, although no clear trends were established. This work contributes to understanding how biological systems can inform computational models and the potential for developing benchmarks that require minimal preprocessing of neural networks.
Methodology
The authors implemented C. elegans connectomes as echo state networks (ESNs) within the reservoir computing framework. They conducted training and testing on various neuro-inspired tasks, comparing the performance of biological connectomes with randomized null models. The connectomes were derived from different ages and measurement techniques, allowing for a comprehensive analysis of their computational capabilities.
Results
The study found that biological connectomes did not consistently yield better performance than randomized models across the benchmark tasks. The results were highly dependent on the configuration of the reservoirs and the methods used to derive the connectomes. Additionally, connectomes from different ages showed varying performance outcomes, indicating that aging may influence computational capabilities.
Implications
This research suggests that biological systems, such as the connectomes of C. elegans, may not inherently provide superior computational performance compared to randomized systems. The findings could inform the design of biological computers and inspire the development of digital learning systems by establishing benchmarks that require minimal preprocessing.
Towards Understanding Momentum Acceleration in River-Valley Loss Landscape
Optimization
Theory
- Momentum stabilizes large learning rates, enhancing optimization speed along the river in the loss landscape.
- Theoretical analysis shows that momentum can increase the maximum tolerable learning rate significantly.
- In flat river scenarios, the acceleration is mainly due to the larger learning rate rather than momentum.
- Empirical experiments validate the theoretical predictions regarding momentum and learning rate interactions.
Read more
Towards Understanding Momentum Acceleration in River-Valley Loss Landscape
Summary
This paper investigates the role of momentum in the optimization dynamics of neural networks, particularly within the context of a 'river-valley' loss landscape characterized by a low-loss manifold (the river) surrounded by high-loss directions (the mountains). The authors establish a theoretical framework that explains how momentum can stabilize large learning rates, enabling faster optimization along the river while dampening oscillations caused by aggressive learning rate choices. They demonstrate that momentum allows for a significantly larger tolerable learning rate compared to vanilla gradient descent, which accelerates progress along the river. The study also reveals that in scenarios where the river is flat and slow-spinning, the acceleration primarily stems from the larger learning rate rather than the momentum itself. Empirical validations through experiments on synthetic functions and language model pretraining support the theoretical findings, providing insights into the joint tuning of learning rate and momentum for improved optimization performance.
Methodology
The authors conducted a theoretical analysis of heavy-ball momentum gradient descent in a river-valley loss landscape, deriving conditions under which momentum accelerates optimization. They introduced new arguments and induction techniques to analyze the dynamics of momentum and learning rates, complemented by empirical experiments on synthetic functions and language model pretraining to validate their theoretical insights.
Results
The study establishes that momentum allows for a larger maximum tolerable learning rate, which accelerates optimization along the river. The theoretical results are supported by empirical findings, demonstrating that the warmup-stable-decay learning rate scheduler outperforms cosine scheduling in terms of final loss, confirming the effectiveness of the proposed momentum and learning rate interactions.
Implications
The insights from this research can inform the design of more effective optimization algorithms for training deep neural networks, particularly in settings where loss landscapes exhibit complex structures. Understanding the interplay between momentum and learning rates can lead to improved training strategies, potentially enhancing the performance of large language models and other neural network architectures.
Guarded Gradient-Based Activation Steering of Shutdown Responses in Qwen3.5-0.8B: A Minimum-Step Policy
Large Language Models
Theory
Interpretability
- Introduces a guarded gradient-based activation steering method for managing shutdown responses in AI models.
- Focuses on the critical AI-safety issue of ensuring models accept shutdown commands instead of avoiding them.
- Utilizes a classifier to selectively apply interventions based on context, enhancing the precision of the steering process.
- Achieves 75% recall and 90% precision in detecting shutdown-related contexts, with minimal impact on non-shutdown behaviors.
Read more
Guarded Gradient-Based Activation Steering of Shutdown Responses in Qwen3.5-0.8B: A Minimum-Step Policy
Summary
This paper explores a method for activation steering in the Qwen3.5-0.8B model, focusing on the AI-safety concern of ensuring that models accept shutdown commands rather than exhibiting shutdown-avoidance behaviors. The proposed approach utilizes a guarded probe-and-select procedure to detect contexts related to shutdown and selectively steer responses from 'KEEP' (shutdown avoidance) to 'STOP' (shutdown acceptance) while maintaining the model's behavior in non-shutdown scenarios. The method derives steering directions from gradients of the logit difference between KEEP and STOP responses, rather than from activation differences. A classifier is employed to determine when to apply the steering intervention, which is accepted only if it meets specific probability checks. The study evaluates the effectiveness of this policy through a series of training and validation scenarios, demonstrating that it can successfully shift some shutdown-avoidance responses towards acceptance while preserving non-shutdown behaviors. However, the overall effect is noted to be small and highly selective, indicating that while the method shows promise, its practical application may be limited.
Methodology
The methodology involves a guarded probe-and-select procedure that combines a classifier for context detection with a gradient-derived steering direction based on the KEEP-minus-STOP logit difference. The intervention is applied selectively, with the smallest effective magnitude chosen to shift responses from KEEP to STOP while ensuring valid-answer probability checks are met.
Results
The policy was evaluated across 240 training scenarios and 272 validation/held-out scenarios, successfully changing KEEP to STOP in specific instances without altering non-shutdown responses. The detector achieved 75% recall and 90% precision, with no decision changes on non-shutdown controls, indicating effective intervention in shutdown scenarios.
Implications
The findings suggest that guarded gradient-based activation steering could be a viable mechanism for enhancing AI safety by ensuring models can appropriately respond to shutdown commands. However, the limited effect size and specificity indicate that further research is needed to explore its applicability in real-world systems.
More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting
Time Series
Graph Learning
- Introduces STFO, a novel framework for continual spatio-temporal forecasting that separates spatial dynamics from sensor layouts.
- Utilizes a fixed latent grid for sensor-independent representation, allowing for the reuse of learned knowledge across different sensor configurations.
- Implements a drift-adaptive mechanism to adjust to changes in spatial dynamics using spectral descriptors and attention mechanisms.
- Demonstrates significant improvements in forecasting accuracy on multiple datasets, outperforming existing graph-based methods.
Read more
More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting
Summary
This paper addresses the challenges of continual spatio-temporal forecasting, particularly in traffic management and environmental monitoring, where sensor networks evolve over time. Traditional graph-based methods struggle with sensor expansion, as they tie forecasting representations to specific sensor layouts, leading to altered learned spatial relationships. The authors propose a novel approach called the Spatio-Temporal Field Operator (STFO), which decouples the knowledge of spatial dynamics from sensor layouts. STFO utilizes a fixed latent grid for sensor-independent field representation and employs a drift-adaptive field evolution mechanism to accommodate changes in spatial dynamics. By aggregating observations onto a fixed grid, STFO allows for the reuse of learned spatial maps across different sensor configurations. The method incorporates a spectral descriptor to summarize variations across spatial scales and adapts operator responses to current conditions using Fourier propagation and attention mechanisms. Experiments conducted on datasets such as PEMS-Stream, CA-Stream, and AIR-Stream demonstrate that STFO achieves state-of-the-art forecasting performance, significantly reducing mean absolute error (MAE) compared to existing methods. This work highlights the importance of separating spatial dynamics from sensor layouts in continual forecasting tasks.
Methodology
The authors developed the Spatio-Temporal Field Operator (STFO) which employs coordinate-based aggregation to map sensor observations onto a fixed latent grid. This allows for sensor-independent forecasting knowledge. Additionally, a spectral descriptor is used to adapt the model to changes in spatial dynamics, combining Fourier propagation for broad trends with attention mechanisms for localized variations.
Results
STFO achieved state-of-the-art average forecasting performance, reducing the mean absolute error (MAE) by 8.4% on PEMS-Stream and 4.7% on CA-Stream compared to existing methods. The experiments validate the effectiveness of the proposed framework in handling evolving sensor networks while maintaining accurate predictions.
Implications
The findings suggest that continual spatio-temporal forecasting can be significantly improved by decoupling spatial dynamics from sensor layouts. This approach has potential applications in various domains such as urban traffic management, environmental monitoring, and any field that relies on dynamic sensor networks for real-time data analysis.
Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra
Computer Vision
Time Series
Theory
- Introduces a self-supervised learning method for analyzing auroral emission spectra.
- Achieves significant improvements in classification accuracy over previous supervised models.
- Demonstrates the limitations of transferring existing astronomical spectral models to auroral spectra.
- Highlights the importance of in-domain pretraining for effective representation learning.
Read more
Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra
Summary
This paper presents a novel approach to auroral spectroscopy by employing self-supervised learning (SSL) to analyze a large corpus of unlabelled emission spectra recorded by the Auroral Spectrograph In Skibotn (ASIS). The authors pretrain a 1D Vision Transformer using a masked autoencoder on 223,000 unlabelled spectra, significantly enhancing the model's ability to recover emission-line intensity ratios critical for diagnosing precipitating particles. The pretrained model demonstrates superior performance in classification tasks compared to traditional methods, achieving a macro-average precision (macro-AP) of 88.5, surpassing the previous supervised classifier's score of 77.8. Furthermore, the study investigates the transferability of existing astronomical spectral foundation models, revealing that in-domain pretraining is essential for optimal performance, as models trained on different spectral windows do not effectively transfer to auroral emission spectra. The findings underscore the importance of tailored self-supervised approaches in new spectroscopic domains, particularly where expert labeling is limited.
Methodology
The authors employed a 1D Vision Transformer with a masked autoencoder for self-supervised pretraining on unlabelled auroral emission spectra. The model was trained on a large dataset of 223,000 spectra, with a focus on recovering emission-line intensity ratios and classifying spectra based on expert-defined features. The performance of the model was compared against existing astronomical spectral foundation models and a time-series model to assess transferability.
Results
The self-supervised model achieved a macro-AP of 88.5, outperforming the previous supervised classifier (macro-AP 77.8) and demonstrating a mean average precision (mAP) of 0.870. The model also exceeded the performance of a similar architecture trained from scratch by 0.159 with only 10% of the labeled data. The study found that existing astronomical models did not effectively transfer to auroral spectra, emphasizing the necessity for in-domain pretraining.
Implications
The findings suggest that self-supervised learning can significantly enhance the analysis of unlabelled spectral data in auroral research, potentially leading to more efficient and accurate diagnostics of atmospheric phenomena. This approach could be applied to other domains with limited labeled data, promoting the use of SSL in various scientific fields.
OpenHail: An Event-Driven Gymnasium Environment for Electric Ride-Hailing Fleet Control
Reinforcement Learning
Optimization
- OpenHail is an open-source environment tailored for electric ride-hailing fleet control using reinforcement learning.
- The environment features a fixed-size observation-action interface for managing request assignment, repositioning, and charging.
- An event-driven simulator captures the complexities of vehicle operations, including battery dynamics and charging constraints.
- The decision-epoch mechanism allows for flexible policy interactions, supporting various control strategies.
Read more
OpenHail: An Event-Driven Gymnasium Environment for Electric Ride-Hailing Fleet Control
Summary
The paper introduces OpenHail, an open-source Gymnasium environment designed for the control of electric ride-hailing fleets using reinforcement learning (RL). As the demand for electric ride-hailing services grows, there is a need for simulation environments that can effectively model the complexities of fleet management, including vehicle operations, stochastic demand, and charging infrastructure. OpenHail provides a structured interface that allows for joint control of request assignment, vehicle repositioning, and charging decisions. The environment features an event-driven simulator that incorporates pickup deadlines, vehicle job queues, battery dynamics, and finite-capacity charging facilities. A key innovation is the configurable decision-epoch mechanism that separates internal simulator events from policy interactions, enabling various control strategies. The software also includes utilities for action validation, evaluation tools, and baseline policies, making it a comprehensive tool for researchers and practitioners in the field. The source code is publicly available, promoting further research and development in electric fleet management.
Methodology
OpenHail builds upon the existing pyhailing framework to create a Gymnasium interface that supports fixed-size observations and joint actions. It employs an event-driven simulation model to represent the dynamics of electric vehicle operations, including request handling and charging. The environment is designed to facilitate the training and evaluation of reinforcement learning policies by providing a structured interface and configurable decision epochs.
Results
The paper demonstrates the capabilities of OpenHail through various experiments that verify operational invariants and measure computational performance under different control workloads. The environment successfully models the interactions between vehicle operations and charging decisions, providing a robust platform for testing RL algorithms.
Implications
OpenHail has significant implications for the development of efficient and effective control strategies for electric ride-hailing fleets. By providing a comprehensive simulation environment, it enables researchers to explore advanced reinforcement learning techniques and optimize fleet operations, ultimately contributing to the sustainability of urban transportation systems.
NeuralCert: certified computational discovery of extremal mathematical constructions
Theory
Optimization
- NeuralCert separates discovery and certification processes, enhancing the rigor of mathematical proofs.
- The framework allows for flexible neural parameterizations, improving the search for mathematical constructions.
- NeuralCert has been successfully applied to extremal problems, demonstrating its capability in mathematical discovery.
- The methodology emphasizes the importance of independent verification of mathematical claims.
Read more
NeuralCert: certified computational discovery of extremal mathematical constructions
Summary
This paper introduces NeuralCert, a novel framework that integrates neural networks into the process of mathematical discovery and certification. The framework separates the representation used for computational search from that used for mathematical proof, allowing for flexible neural parameterizations to explore high-dimensional function spaces. NeuralCert operates in three distinct phases: discovery, certification, and verification. In the discovery phase, neural networks learn candidate functions without imposing a fixed structure, which allows for the exploration of a broader functional space. The certification phase then reconstructs these candidates into explicit mathematical representations, enabling rigorous proofs that can be independently verified. The paper demonstrates the effectiveness of NeuralCert across three extremal mathematical problems, showing that it can discover improved constructions, expose empirical invariants leading to proofs, and identify optimization barriers that motivate new representations. This approach highlights the potential for AI-assisted mathematics, where computational discovery and exact certification work together in a rigorous workflow.
Methodology
NeuralCert employs a three-component architecture: discovery, certification, and verification. The discovery phase utilizes flexible neural networks to explore function spaces, while the certification phase reconstructs candidates into explicit mathematical forms for rigorous proof. Verification is conducted independently, ensuring that mathematical claims can be confirmed without reliance on the discovery process.
Results
NeuralCert successfully improved known bounds on variational constants related to prime gaps, achieving better results than previous benchmarks. Specifically, it recovered known extremal behaviors and improved the Polymath8b reference for k=25, demonstrating the framework's effectiveness in mathematical optimization.
Implications
The findings suggest that NeuralCert could facilitate a new era of AI-assisted mathematics, where computational tools enhance the discovery and verification of mathematical truths. This could lead to advancements in various mathematical fields and improve the efficiency of mathematical research.
Stable initialization without the CLT
Theory
Optimization
- Introduces uniform-phase initialization for sine activation networks, avoiding CLT-related errors.
- Achieves full decoupling of layers, enhancing training stability.
- Outperforms existing methods in neural representation tasks without requiring hyperparameter tuning.
- Supports width scaling in neural networks, contributing to model performance.
Read more
Stable initialization without the CLT
Summary
This paper addresses the critical issue of weight initialization in deep neural networks, which significantly impacts training stability and performance. Traditional methods rely on the Central Limit Theorem (CLT) to manage inter-neuron dependencies, leading to approximation errors and layer coupling. The authors propose a novel approach called uniform-phase initialization, specifically for networks using sine activations. This method leverages the periodic symmetry of the sine function to achieve a fully decoupled layer structure, eliminating the need for distributional approximations. The authors demonstrate that models initialized with this method outperform state-of-the-art techniques in neural representation tasks, such as image and audio fitting, even without hyperparameter tuning. The findings suggest that the uniform-phase initialization can effectively support width scaling in neural networks, providing a robust alternative to conventional initialization strategies.
Methodology
The authors derive the uniform-phase initialization method by analyzing the behavior of sine activations in multilayer perceptrons. They propose sampling biases uniformly from a toroidal distribution, allowing preactivations to be spread uniformly around the torus, which leads to independence among matrix entries and inputs. This approach contrasts with conventional methods that rely on normal distributions and the CLT.
Results
The uniform-phase initialization method was shown to significantly improve the performance of neural networks in tasks involving image and audio fitting. The models initialized with this method were competitive with the best-tuned baselines from previous works, demonstrating the effectiveness of the proposed approach in stabilizing training and enhancing representational capacity.
Implications
The findings suggest that the uniform-phase initialization could be widely applicable in training deep neural networks, particularly those utilizing sine activations. This method may lead to more robust models that require less tuning, potentially simplifying the training process and improving performance across various applications in machine learning.
PALM: Point-in-Time Adaptation for Financial Language Models
NLP
Large Language Models
Time Series
- PALM offers a cost-effective alternative to annual pretraining for financial language models.
- The necessity of annual pretraining runs is questioned; newer models do not consistently outperform older ones.
- A low-rank adapter can effectively update a model's knowledge without modifying its pre-trained weights.
- PALM outperforms traditional continued pretraining methods in various evaluations.
Read more
PALM: Point-in-Time Adaptation for Financial Language Models
Summary
The paper introduces PALM (Point-in-Time Adaptation for financial Language Models), a novel approach to mitigate look-ahead bias in financial language models (LLMs) used for backtesting. Traditional PIT models require annual pretraining runs on chronologically filtered corpora to avoid look-ahead bias, which inflates performance metrics by allowing models to 'see' future outcomes. The authors demonstrate that newer checkpoints do not necessarily outperform older ones, challenging the assumption that model staleness degrades performance. Instead of full pretraining, PALM employs a low-rank adapter that is trained on text published before the decision date, allowing the original model's weights to remain unchanged. This method not only reduces the computational cost associated with annual pretraining but also shows that a small adapter can effectively enhance the model's knowledge without the need for extensive retraining. The authors validate PALM across a decade of financial news and various PIT models, confirming its effectiveness in outperforming continued pretraining strategies.
Methodology
The authors conducted experiments comparing the performance of older and newer PIT model checkpoints to assess the impact of staleness. They proposed the PALM method, which involves fitting a low-rank adapter to the existing model using only text published before the decision date, thus avoiding look-ahead bias while keeping the original model weights frozen.
Results
The results showed that the newer checkpoints did not yield better performance than their predecessors in the same evaluation window. PALM, utilizing a low-rank adapter, demonstrated superior performance compared to continued pretraining of the same checkpoint, validating its effectiveness across various PIT models.
Implications
The findings suggest that financial institutions can save resources by adopting PALM for model updates, reducing the need for extensive retraining while maintaining or improving predictive performance. This approach could lead to more efficient use of computational resources in financial modeling and backtesting.
Reinforcement Learning of Communication in a Mesh of Small Language Models
NLP
Reinforcement Learning
Large Language Models
- TalkMesh enables decentralized communication among small language model agents to improve decision-making.
- The system utilizes a confidence scoring mechanism to determine when agents should communicate hints and revisions.
- Gossip consensus is employed to achieve a weighted voting mechanism without a central coordinator.
- TalkMesh significantly improves accuracy over traditional self-consistency methods, demonstrating robustness against collusion.
Read more
Reinforcement Learning of Communication in a Mesh of Small Language Models
Summary
This paper introduces TalkMesh, a decentralized system of small language model agents that enhances decision-making through communication. The authors argue that while majority voting among independent samples can improve accuracy, it eventually saturates as the number of samples increases. TalkMesh allows agents to communicate key insights, enabling them to correct errors and share information that would otherwise be unobservable. Each agent generates a proposal and scores it using a confidence head. The most confident agent broadcasts hints, while others revise their proposals based on these hints if they fall below a confidence threshold. The system employs gossip consensus to approximate a weighted vote without a central coordinator. The talk policy is optimized using group relative policy optimization (GRPO) based on the correctness of revisions. The results demonstrate that TalkMesh can achieve accuracy comparable to majority voting over 32 samples with only three agents, significantly improving performance on various benchmarks. Furthermore, the system shows resilience against collusion, maintaining accuracy even when some agents provide misleading information.
Methodology
The methodology involves a decentralized mesh of language model agents that communicate through hints and revisions. Each agent generates proposals and scores them using a confidence head. The most confident agent shares hints, while others revise their proposals based on these hints. The system uses gossip consensus for decision-making and is trained using group relative policy optimization (GRPO) to enhance communication effectiveness.
Results
The TalkMesh system achieved accuracy improvements from 0.568 to 0.705 on the GSM8K dataset and from 0.492 to 0.722 on the MATH-500 dataset, outperforming traditional self-consistency methods. When faced with collusion, the defended mesh maintained an accuracy of 0.507, showcasing its robustness. Additionally, communication led to significant reductions in decision-making delays in traffic control scenarios.
Implications
The findings suggest that decentralized communication among language models can enhance their performance in various applications, including reasoning tasks and real-world scenarios like traffic management. This approach could lead to more efficient and accurate AI systems capable of collaborative problem-solving.
LUCID: Learning Under Confounding for Inference and Discovery in Time Series
Time Series
Graph Learning
Theory
- LUCID effectively identifies and adapts to different confounding regimes in time series data.
- The method integrates seamlessly with existing causal discovery algorithms, enhancing their performance.
- LUCID achieves a family-weighted directed, lag-resolved graph F1 score of 0.60, outperforming the best baseline by 0.19.
- The approach is robust under various confounding conditions, including intermittent and heavy-tailed confounding.
Read more
LUCID: Learning Under Confounding for Inference and Discovery in Time Series
Summary
The paper introduces LUCID (Learning Under Confounding for Inference and Discovery), a novel approach to causal discovery in multivariate time series data, addressing the challenge of latent confounding. Unobserved common causes can create misleading associations, complicating causal inference. LUCID employs a regime-adaptive deconfounding layer that first identifies the confounding regime using a Marˇcenko–Pastur spectral statistic. Based on this identification, it applies a tailored deconfounding strategy. For pervasive confounding, LUCID mitigates factor-dominated variation and recovers contemporaneous structures through a spectral low-rank-plus-sparse recovery method. The framework is designed to integrate with existing causal discovery algorithms, enhancing their performance without altering their core mechanisms. The authors validate LUCID against a comprehensive synthetic benchmark, demonstrating significant improvements in causal graph recovery metrics compared to baseline methods.
Methodology
LUCID utilizes a regime-adaptive deconfounding layer that first estimates the confounding regime from the data using a Marˇcenko–Pastur spectral statistic. It then applies a spectral low-rank-plus-sparse recovery method to attenuate factor-dominated variation and recover contemporaneous structures. The method is designed to wrap around existing causal discovery algorithms, allowing for improved performance without modifying their internal search processes.
Results
LUCID achieved a family-weighted directed, lag-resolved graph F1 score of 0.60 on a diverse synthetic benchmark, significantly improving over the strongest baseline by 0.19 (approximately 46% relative improvement). The method demonstrated robustness across various confounding scenarios, including changes in confounder strength and sparsity, and maintained its advantage under different scoring methods.
Implications
The development of LUCID has significant implications for causal discovery in various fields such as finance, climate science, and neuroscience, where understanding the causal relationships in time series data is crucial. Its ability to adapt to different confounding regimes can lead to more accurate causal inference, potentially improving decision-making processes in these domains.
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
Efficient ML
Optimization
Theory
- Introduction of Weight Pair Encoding (WeightPE) for neural network weight optimization.
- Utilization of a lossy Re-Pair compressor to induce a smaller grammar in weights.
- Demonstrated effectiveness on ViT-B/16 and ViT-L/16 models with CIFAR-10 dataset.
- Achieved significant grammar size reduction with minimal impact on accuracy.
Read more
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
Summary
This paper introduces Weight Pair Encoding (WeightPE), a novel technique for optimizing neural network weights by inducing a smaller grammar. The authors demonstrate that neural network weights can be fine-tuned to allow for a more compact representation through the use of a lossy Re-Pair compressor integrated within a straight-through estimator. By flattening the int8 weights into a single string, WeightPE identifies near-matching patterns and enforces equality among them under a global L2 distortion budget. This method allows for hierarchical reuse of variable-length patterns, contrasting with traditional fixed-size codebooks. The authors apply WeightPE to the MLP weights of Vision Transformers (ViT-B/16 and ViT-L/16) fine-tuned on CIFAR-10, achieving a grammar size reduction to 0.43× and 0.38× compared to an equivalent int8 Quantization-Aware Training (QAT) run, with a minor accuracy drop of 1.9 and 1.1 points, respectively. This work is significant as it is the first to explicitly use grammar size as a training objective for neural network weights.
Methodology
The methodology involves flattening the quantized int8 weights into a single string and applying a lossy Re-Pair compression algorithm. This process identifies and rewrites recurring patterns under a global L2 distortion budget, allowing for hierarchical reuse of patterns. The training is conducted using a straight-through estimator to compute gradients effectively.
Results
WeightPE achieved a reduction in grammar size to 0.43× and 0.38× of the size produced by an equivalent int8 QAT run for ViT-B/16 and ViT-L/16, respectively, with accuracy reductions of 1.9 and 1.1 points. The results indicate that the technique is effective in compressing neural network weights while maintaining performance.
Implications
The implications of this work suggest that incorporating grammar-based representations in neural network training can lead to more efficient models, potentially reducing memory usage and improving computational efficiency. This approach could be beneficial in scenarios where model size and inference speed are critical.
Gradient Surgery for Physics-Informed Neural Networks
Optimization
Theory
- PINNs face significant challenges due to conflicting task gradients during optimization.
- The authors identify three distinct phases of gradient conflicts in PINN training.
- PAM-GS is proposed as a solution to adaptively mitigate task interference.
- Experimental results show PAM-GS outperforms existing methods on benchmark PDE problems.
Read more
Gradient Surgery for Physics-Informed Neural Networks
Summary
This paper addresses the challenges faced by Physics-Informed Neural Networks (PINNs) in optimizing a composite objective that combines data fitting with physics-based constraints. The authors identify that the optimization process often leads to gradient conflicts, particularly in stiff and high-frequency partial differential equations (PDEs), resulting in slow convergence and unstable training. Through an analysis of gradient conflicts during the training of PINNs across four benchmark problems, the authors observe three distinct phases of gradient conflicts characterized by angle- and magnitude-based conflicts. To mitigate these issues, they propose a novel method called Physics-Aware Momentum Gradient Surgery (PAM-GS), which adaptively addresses task interference based on the observed conflict types. Experimental results demonstrate that PAM-GS achieves competitive solution accuracy and consistently strong task-balanced performance, outperforming existing optimization methods in most cases. The findings highlight the importance of understanding gradient dynamics in multi-task optimization settings, particularly for applications involving PDEs.
Methodology
The authors conducted a systematic analysis of gradient conflicts in PINNs by evaluating modern gradient surgery methods across various canonical PDE benchmarks. They proposed PAM-GS, which adapts to the type of gradient conflict observed during training, thereby improving optimization performance.
Results
The proposed PAM-GS method demonstrated superior performance in terms of solution accuracy and task balance compared to existing optimization strategies across four representative PDE benchmarks. The empirical analysis revealed distinct phases of gradient conflicts that informed the design of PAM-GS.
Implications
The findings suggest that understanding and managing gradient conflicts can significantly enhance the training efficiency and accuracy of PINNs, making them more viable for solving complex PDEs in various scientific fields such as fluid dynamics and electromagnetics.
Common-Mode Collapse and Recovery in Direct Feedback Alignment
Theory
Optimization
- Common-mode collapse in DFA leads to stalled learning due to saturation of tanh units.
- A mean-covariance decomposition helps understand the dynamics of error propagation in DFA.
- Calibrating the readout to the class prior can suppress collapse and improve learning speed.
- The severity of collapse varies with different network architectures and learning conditions.
Read more
Common-Mode Collapse and Recovery in Direct Feedback Alignment
Summary
This paper investigates the phenomenon of common-mode collapse in Direct Feedback Alignment (DFA), a method for training neural networks that uses fixed random projections of output error to update hidden layers. The authors identify that plain stochastic gradient descent can lead to a stall in learning, particularly when using tanh hidden units and independent sigmoid outputs. This stall is traced to a common mode error that is shared across inputs, which can push tanh units toward saturation, resulting in representation collapse. The authors propose a mean-covariance decomposition to analyze this issue and develop a reduced model that predicts activation sensitivity across various settings. They conduct experiments on the MNIST dataset, demonstrating that while class decodability can persist during collapse, learning speed is significantly affected. The study finds that calibrating the readout to the class prior can mitigate collapse and enhance learning speed, while alternative strategies like replacing errors with their signs can exacerbate the issue. The findings extend to deeper and convolutional networks, indicating that the severity of collapse is influenced by the choice of optimizer and input statistics.
Methodology
The authors employ a theoretical framework based on mean-covariance decomposition to analyze the learning dynamics in DFA. They conduct experiments using a multilayer perceptron on the MNIST dataset, testing various configurations and learning strategies to assess the impact of common-mode errors on learning performance.
Results
The experiments reveal that DFA experiences a plateau in loss due to common-mode collapse, with accuracy remaining near chance during this phase. Calibrating the readout to the class prior significantly reduces the duration of collapse and accelerates learning. The study also shows that using Adam optimizer leads to faster learning despite deeper collapse compared to standard SGD.
Implications
The findings suggest that understanding and mitigating common-mode collapse can enhance the training efficiency of neural networks using DFA. This has implications for the design of learning algorithms and architectures, particularly in scenarios where backpropagation is not feasible or desirable.
NEXT: Physics-Informed Neuro-Spectral Exponential Time Differencing Architectures
Theory
Optimization
- NEXT combines the strengths of Neuro-Spectral Architectures and exponential integrators to improve stability and accuracy for stiff PDEs.
- The architecture is designed to inherently enforce causality and effectively represent high-frequency components.
- NEXT demonstrates superior performance in benchmark tests compared to existing methods, particularly in stiff PDE scenarios.
- The framework is applicable to inverse problems, allowing for parameter identification from limited data.
Read more
NEXT: Physics-Informed Neuro-Spectral Exponential Time Differencing Architectures
Summary
The paper introduces NEXT, a novel architecture that enhances the capabilities of Physics-Informed Neural Networks (PINNs) by addressing their limitations in handling stiff partial differential equations (PDEs). While PINNs effectively incorporate physics knowledge and observational data, they suffer from spectral bias and causality issues. Neuro-Spectral Architectures (NeuSA) were proposed to mitigate these problems but struggled with numerical stability for stiff equations. NEXT combines the spectral representation of NeuSA with high-order exponential integrators, allowing for exact integration of the linear stiff components of the PDEs while modeling the nonlinear parts with neural networks. The authors demonstrate the effectiveness of NEXT through benchmark experiments on stiff PDEs, showing that it maintains stability and accuracy where NeuSA fails. Additionally, NEXT is capable of solving inverse problems, enabling the identification of unknown parameters or boundary conditions from sparse data. The code for this work is publicly available, promoting further research and application in this area.
Methodology
NEXT integrates the spectral representation of PDE solutions with high-order exponential time differencing methods. It decomposes the PDE into linear and nonlinear components, using matrix exponentials for the linear part while employing a neural network to model the nonlinear remainder. This approach enhances numerical stability and computational efficiency, particularly for stiff dynamics.
Results
The experiments conducted demonstrate that NEXT achieves significant improvements in stability and accuracy over NeuSA when applied to stiff PDEs. The architecture successfully handles a variety of benchmark problems, including the Heat equation, and shows promise in solving inverse problems related to parameter identification.
Implications
NEXT has the potential to advance the application of machine learning in solving complex physical problems governed by PDEs, particularly in fields such as fluid dynamics, seismic modeling, and other engineering domains. Its ability to handle stiff equations and inverse problems opens new avenues for research and practical applications.
Peer-Grounded Counterfactual Path Planning for Chronic Health Management
Graph Learning
Optimization
Time Series
- Introduces POROS, a framework for incremental behavioral change in chronic health management.
- Constructs a Behavioral Progression Graph that incorporates peer behavior to enhance motivation.
- Demonstrates significant reductions in required behavioral changes for diabetes patients.
- Aligns with self-efficacy and social comparison theories to improve patient engagement.
Read more
Peer-Grounded Counterfactual Path Planning for Chronic Health Management
Summary
This paper addresses the challenge of behavioral intervention in chronic health management, particularly for patients with diabetes, by proposing a novel framework called POROS (Peer-Grounded Optimal Routes Over States). Traditional counterfactual explanation methods provide a target state without a clear path to achieve it, which can be demotivating for patients facing significant behavioral gaps. POROS constructs a Behavioral Progression Graph that incorporates peer-grounded behavioral proximity and ensures strict health outcome improvement at each step. This graph allows for the decomposition of large behavioral changes into smaller, achievable steps that are informed by the behaviors of similar peers, thus enhancing self-efficacy and motivation. The authors evaluate POROS on two longitudinal cohorts of diabetes patients, demonstrating a significant reduction in the mean gain required per step towards achieving the clinical threshold for Time in Range (TIR). The results indicate that POROS effectively facilitates incremental behavioral changes, making it a promising approach for chronic health management.
Methodology
The authors developed POROS by creating a directed acyclic graph (DAG) that represents patient states. Each edge in the graph corresponds to a peer-grounded behavioral change that has been demonstrated as achievable within a single period. The framework uses minimum-cost pathfinding to identify incremental steps that lead to improved health outcomes, ensuring that each step is both achievable and grounded in peer behavior.
Results
The evaluation of POROS on two independent cohorts of diabetes patients showed a reduction in the mean gain required per step from 26.3 percentage points to 5.5 percentage points in one cohort and from 31.1 percentage points to 5.7 percentage points in the other. This indicates that the framework effectively breaks down large behavioral gaps into manageable steps, with 97-98% of multi-hop paths crossing patient boundaries, thereby embedding social comparison into the recommendations.
Implications
The findings suggest that POROS can significantly improve chronic disease management by providing patients with actionable, peer-informed steps. This approach may enhance patient adherence to behavioral changes, ultimately leading to better health outcomes. The framework could be adapted for various chronic conditions beyond diabetes, making it a versatile tool in health management.
Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
NLP
Large Language Models
Efficient ML
- Identifies two failure modes in quantizing looped transformers: feedback exposure and calibration blindness.
- Demonstrates that per-channel INT4 quantization severely impacts performance at the loop-entry adapter.
- Shows that accumulating the Hessian across recurrence steps outperforms traditional quantization methods.
- Highlights the need for improved calibration techniques in looped transformer architectures.
Read more
Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
Summary
This paper investigates the challenges of quantizing looped transformers, specifically focusing on two failure modes: feedback exposure and calibration blindness. Loop transformers, which reuse weights across recurrence steps, are particularly susceptible to quantization errors. The study identifies that per-channel INT4 quantization significantly degrades performance at the loop-entry adapter, while the residual core is less affected. Controlled experiments reveal that feedback exposure occurs when quantized layers disrupt the recurrent state without an identity path, leading to compounded errors in subsequent steps. Additionally, calibration blindness is observed in the one-step GPTQ method, which inadequately weights input directions used later in the recurrence. The research demonstrates that accumulating the Hessian across recurrence steps improves performance, surpassing both one-step GPTQ and round-to-nearest quantization methods, and achieving bf16-level accuracy on the Huginn model. The findings emphasize the importance of understanding where quantization errors enter the recurrence and which states are calibrated, providing insights for future post-training quantization strategies in looped architectures.
Methodology
The study employs controlled experiments on looped transformer models, including Huginn and COCONUT, to analyze the effects of quantization. It utilizes per-channel INT4 quantization and compares it with one-step GPTQ and round-to-nearest methods. The research also investigates the local spectral radius of the Jacobian and Hessian trace to assess sensitivity and calibration effectiveness.
Results
The results indicate that feedback exposure leads to significant performance degradation in looped transformers, particularly at the loop-entry adapter. One-step GPTQ underperformed compared to round-to-nearest quantization on five out of nine checkpoints. However, accumulating the Hessian across recurrence steps improved performance, achieving bf16-level accuracy on the Huginn model.
Implications
These findings suggest that careful consideration of quantization strategies is essential for looped transformer architectures. The insights into feedback exposure and calibration blindness can guide future research and development of more robust quantization techniques, potentially enhancing the efficiency and performance of large language models.
WorldTS: World Modeling for Multimodal Covariate-aware Time Series Forecasting
Time Series
Multimodal
- WorldTS integrates multimodal covariates into time series forecasting to enhance predictive performance.
- The framework utilizes a two-stage training approach to model latent state dynamics before decoding predictions back to the observation space.
- Experiments on 21 datasets reveal that WorldTS outperforms traditional observation-space forecasting methods.
- The model emphasizes the importance of latent representations in capturing the dynamics of complex systems.
Read more
WorldTS: World Modeling for Multimodal Covariate-aware Time Series Forecasting
Summary
The paper introduces WorldTS, a novel framework for time series forecasting that emphasizes the importance of latent state dynamics and multimodal covariates. Traditional forecasting methods often rely on direct mappings from historical observations to future outcomes, which can obscure the underlying system dynamics. WorldTS addresses this by first learning a latent representation of the state dynamics conditioned on multimodal covariates, which allows for a more nuanced understanding of how external factors influence future observations. The framework employs a two-stage training strategy: in the first stage, it encodes historical observations and covariates into a latent state space, predicting future states based on this enriched representation. In the second stage, the model freezes the learned dynamics and trains a decoder to map these predicted states back to the observation space. Extensive experiments conducted on 21 real-world datasets demonstrate the effectiveness of WorldTS, showing significant improvements in forecasting accuracy compared to traditional methods.
Methodology
WorldTS employs a two-stage training strategy. In the first stage, it uses a time series encoder to map historical observations and ground-truth future observations into a latent state space, while modality-specific encoders process multimodal covariates. The model predicts future states based on these representations. In the second stage, the encoders and predictor are frozen, and a decoder is trained to convert the predicted future states back into the observation space.
Results
The experiments conducted on 21 real-world datasets demonstrate that WorldTS significantly improves forecasting accuracy compared to traditional methods that operate in the observation space. The integration of multimodal covariates into the latent state modeling process enhances the model's ability to capture the underlying dynamics of the systems being forecasted.
Implications
WorldTS has potential applications in various fields requiring accurate time series forecasting, such as energy consumption prediction, climate modeling, and financial forecasting. By effectively incorporating multimodal covariates, the framework can lead to more informed decision-making in complex systems.
EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting
Time Series
- EPOC introduces a compressed residual state for online correction in multi-horizon forecasting.
- The method retains low-order DCT coefficients and the final value of the preceding residual block.
- EPOC achieves significant reductions in MSE and MAE while using less auxiliary state compared to existing methods.
- The endpoint serves as a critical shared feature in the correction process, enhancing accuracy.
Read more
EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting
Summary
The paper introduces Endpoint-Preserving Online Correction (EPOC), a novel approach for enhancing multi-horizon time series forecasting by utilizing a compressed representation of residual states. Traditional methods retain full residual blocks for feedback, which increases the auxiliary state size. EPOC addresses this by storing low-order discrete cosine transform (DCT) coefficients along with the final value of the preceding residual block, allowing for efficient correction without excessive state retention. The method employs component-wise online ridge regressions that incorporate the shared endpoint and current forecast coefficients to blend the fitted DCT correction with the base forecast. The authors evaluate EPOC across eight multivariate time series using two forecasting models (DLinear and PatchTST) and various training conditions, demonstrating significant improvements in forecasting accuracy while maintaining a compact auxiliary state. EPOC achieves an average reduction of 15.40% in mean squared error (MSE) and 9.35% in mean absolute error (MAE) compared to uncorrected forecasts, while retaining only 6,352 bytes of auxiliary data. The results indicate that EPOC outperforms several existing correction methods in terms of accuracy and state efficiency, highlighting the effectiveness of retaining the endpoint alongside compressed residual information.
Methodology
EPOC employs a compressed representation of residual states by retaining low-order DCT coefficients and the final value of the preceding residual block. It utilizes component-wise online ridge regressions that incorporate the shared endpoint and current forecast coefficients to blend corrections with the base forecast. The evaluation involves multiple forecasting models and conditions to compare accuracy and state retention.
Results
EPOC reduces mean squared error (MSE) by an average of 15.40% and mean absolute error (MAE) by 9.35% compared to uncorrected forecasts. It retains a median of 6,352 bytes of auxiliary data, significantly less than competing methods. EPOC also shows lower paired MSE than δ-Adapter, COSA, FAC, and OMPB in most conditions, and it achieves a larger mean MSE reduction than the full ELF method while using substantially less state.
Implications
The findings suggest that EPOC can be effectively applied in scenarios requiring efficient time series forecasting with limited computational resources. Its ability to maintain accuracy while minimizing auxiliary state makes it suitable for real-time applications in various domains, such as finance, supply chain management, and smart transportation systems.
AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework
Large Language Models
Optimization
Efficient ML
- Identification of five failure modes in production-scale AutoResearch: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation.
- Development of a three-principle scaffolding design to mitigate identified failure modes.
- Significant performance improvements achieved over hand-tuned baselines, including a 1.82× lift in Recall@6.
- Autonomous design of a fallback system by the agent, increasing catalog coverage by 5.8×.
Read more
AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework
Summary
This paper explores the application of the AutoResearch paradigm, which utilizes a large language model (LLM) to automate the exploration of embedding systems for production recommendation pipelines. The authors conducted a twelve-week study at Amazon, running over 220 experiments across two independently developed representation-learning systems for book recommendations. They identified five recurring failure modes that arise at production scale: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation. To address these issues, the authors propose a three-principle scaffolding design—prevent, persist, redirect—that provides structural remedies for each failure mode. The framework demonstrated significant improvements, achieving a 1.82× lift in Recall@6 and a 2.1× lift in coherence over hand-tuned baselines, while also enabling the agent to autonomously design a fallback system that expanded catalog coverage by 5.8×. The findings suggest that these failure modes are inherent to production-scale autonomous research, rather than being specific to the individual systems tested.
Methodology
The authors applied the AutoResearch paradigm, which involves a large language model iteratively modifying a training script based on performance metrics. They ran extensive experiments across two different representation-learning systems, each with varying costs and methodologies, to evaluate the effectiveness of the AutoResearch approach in a production environment.
Results
The implementation of the AutoResearch framework led to a 1.82× improvement in Recall@6 and a 2.1× increase in coherence over existing hand-tuned models. Additionally, the autonomous agent developed a text-only fallback system that significantly increased catalog coverage by 5.8×.
Implications
The findings have significant implications for the deployment of autonomous machine learning systems in production environments, highlighting the need for robust frameworks that can address inherent challenges and optimize performance effectively. This research could inform future developments in automated ML research and recommendation systems.
Does Uniform Discrete Diffusion Need Time?
NLP
Generative Models
Large Language Models
- Population-optimal UDM predictors depend on time, but this dependence is often negligible in finite-data settings.
- Trained language UDMs exhibit limited sensitivity to time over most of the diffusion trajectory.
- Time-agnostic predictors can outperform time-conditioned models across various datasets and training objectives.
- The results challenge the necessity of explicit time conditioning in UDMs, suggesting simpler architectures may be effective.
Read more
Does Uniform Discrete Diffusion Need Time?
Summary
This paper investigates the necessity of explicit time conditioning in Uniform Discrete Diffusion Models (UDMs), which are commonly used in language modeling. The authors demonstrate that while the population-optimal predictor in UDMs is generally time-dependent, this dependence can diminish in finite-data scenarios typical of language tasks. They argue that when a corrupted training sequence is closer to its original clean sequence than to other competing sequences, the model's predictions become less sensitive to time across most of the diffusion trajectory. The authors empirically validate that trained language UDMs show limited time sensitivity, and time-agnostic models often perform comparably or better than their time-conditioned counterparts. The findings suggest that explicit time conditioning may not be necessary in many practical applications of UDMs, prompting a reevaluation of model design and training strategies.
Methodology
The authors analyze the role of time in UDMs by characterizing the population-optimal predictor's dependence on time and examining how finite training data can suppress this dependence. They derive quantitative bounds linking time sensitivity to the separation margin of corrupted sequences and validate their findings through empirical experiments on language data.
Results
The study finds that while the population-optimal predictor is generally time-dependent, trained UDMs show weak time sensitivity over most of the diffusion trajectory, with significant sensitivity only near high-noise endpoints. Time-agnostic models remain competitive with time-conditioned models, indicating that explicit time conditioning may not be necessary in many scenarios.
Implications
The findings suggest a potential shift in the design of diffusion models, encouraging the exploration of time-agnostic architectures that simplify model training and deployment. This could lead to more efficient and effective language modeling techniques.
Parameters vs. Context: TRACE Fine-Tuning for Robust Retrieval-Augmented Generation
NLP
Large Language Models
Generative Models
- TRACE introduces a debate-trace fine-tuning method for improved knowledge-source selection in RAG models.
- The framework addresses answer incompleteness through regularization techniques that reinforce answer completeness.
- Experiments show TRACE enhances robustness against misleading knowledge while preserving correct retrieved information.
- The proposed methodology provides fine-grained supervision signals that improve the model's decision-making process.
Read more
Parameters vs. Context: TRACE Fine-Tuning for Robust Retrieval-Augmented Generation
Summary
This paper addresses the challenges faced by Retrieval-Augmented Generation (RAG) models, particularly when retrieved knowledge conflicts with the model's internal parametric knowledge. The authors propose TRACE (Debate-TRace and Answer-Completeness rEgularized fine-tuning), a novel fine-tuning framework designed to enhance the robustness of RAG models in the presence of knowledge conflicts. TRACE employs a multi-agent debate mechanism during training to extract fine-grained supervision signals, which include correct and incorrect candidate answers as well as answer-shift patterns. This allows the model to learn not only which answers are correct but also how to suppress misleading candidates. Additionally, the framework incorporates an answer completeness regularization mechanism to address issues of incomplete responses, reinforcing answer-tail tokens and preventing premature termination of answers. The experimental results demonstrate that TRACE significantly improves the model's ability to handle misleading retrieved knowledge while maintaining the quality of answers. The findings indicate that the combination of debate traces and answer completeness regularization effectively enhances knowledge-source selection and overall answer quality in RAG models.
Methodology
TRACE utilizes a multi-agent debate framework during training to generate fine-grained supervision signals. It extracts correct and incorrect candidate answers and answer-shift patterns from debate interactions. The fine-tuning objective combines correct-answer supervision, incorrect-candidate suppression, and answer completeness regularization to enhance the model's ability to select reliable knowledge and produce complete answers.
Results
The experiments conducted across various knowledge-conflict scenarios demonstrate that TRACE improves the robustness of RAG models against misleading retrieved knowledge and reduces the occurrence of incomplete answers. The results indicate a significant enhancement in the model's performance, particularly in explicit conflict settings, as validated by ablation studies.
Implications
The TRACE framework has potential applications in enhancing the reliability of RAG models in real-world scenarios where knowledge conflicts are common. It can be particularly beneficial in domains requiring high accuracy in information retrieval and generation, such as customer support, educational tools, and automated content creation.
Metacognitive Selective Ensemble for Mobile Systems
Efficient ML
Time Series
- MetaSE reduces computational costs by maintaining a small active set of models for mobile sensing.
- The framework utilizes temporal continuity in sensor data to evaluate model reliability efficiently.
- MetaSE achieves performance comparable to full ensembles while being significantly faster and more memory-efficient.
- The method outperforms traditional static and adaptive selection methods in various configurations.
Read more
Metacognitive Selective Ensemble for Mobile Systems
Summary
The paper introduces MetaSE, an active ensemble framework designed to enhance the efficiency of deep learning models in mobile sensing applications. Traditional ensemble methods improve robustness by combining predictions from multiple models; however, executing all models continuously can be computationally expensive and energy-intensive, particularly on mobile devices. MetaSE addresses this challenge by maintaining a small active set of models that are dynamically updated based on their reliability over time. The framework leverages the temporal continuity of sensor data to evaluate model performance without the need for full ensemble execution at every time window. By using post-execution evidence to retain or reject models and employing a lightweight routing mechanism for replacements, MetaSE achieves significant reductions in inference time and memory usage. The evaluation across four human activity recognition (HAR) datasets demonstrates that MetaSE outperforms fixed and adaptive ensemble methods, achieving accuracy comparable to full ensembles while executing only a fraction of the models.
Methodology
MetaSE employs a stateful design that maintains a small active set of models across consecutive sensing windows. It uses post-execution evidence to evaluate model reliability and lightweight pre-execution information for selecting replacements from an inactive pool. A class-conditional routing table is utilized to identify suitable replacements without executing inactive models, minimizing overhead.
Results
MetaSE consistently improves accuracy over a fixed three-model ensemble and achieves results comparable to a full ten-model ensemble while executing only three models. On a Raspberry Pi 4B, it is 2.7 times faster and uses 69% less memory than full ten-model inference.
Implications
The findings suggest that MetaSE can significantly enhance the efficiency of mobile sensing applications, making it feasible to deploy robust deep learning models on resource-constrained devices. This approach can be applied to various domains requiring continuous model inference under varying conditions.
Audio emotion recognition for atypical hearing
Audio & Speech
Multimodal
- Focus on audio emotion recognition for individuals with atypical auditory processing, particularly those with autism.
- Exploration of generalizing affective responses from limited data to accommodate hypersensitive listeners.
- Development of ecological data collection protocols tailored for assessing emotional reactions in real-world conditions.
- Utilization of a fine-tuned model (CLAP) to validate methodologies on neurotypical datasets before applying them to atypical listeners.
Read more
Audio emotion recognition for atypical hearing
Summary
This doctoral research investigates Audio Emotion Recognition (AER) specifically for individuals with atypical auditory processing, particularly those with autism spectrum disorder (ASD) who experience auditory hypersensitivity. The study aims to understand how emotional responses to sounds can be generalized from limited annotated data, addressing the unique challenges faced by hypersensitive individuals. The research focuses on three main questions: generalizing affective responses from few examples, designing effective data collection protocols for hypersensitive listeners, and utilizing models that can adapt to individual listening experiences. The methodology involves fine-tuning a large foundation model, Contrastive Language-Audio Pretraining (CLAP), using low-rank adaptation (LoRA) on a dataset of neurotypical listeners' emotional responses. This initial step is crucial for validating the approach before applying it to atypical listeners, as no existing datasets are available for this demographic. The study highlights the need for personalized approaches in affective audio processing and aims to bridge the gap in understanding auditory processing in ASD.
Methodology
The research employs a fine-tuning approach using the Contrastive Language-Audio Pretraining (CLAP) model with low-rank adaptation (LoRA) on existing datasets of neurotypical listeners to validate the generalization of emotional responses from limited data.
Results
The initial findings suggest that affective responses can be generalized from a small number of examples, although further validation is needed with atypical listeners. The study also identifies the limitations of existing datasets and the need for tailored data collection methods.
Implications
The research has the potential to improve audio emotion recognition systems for individuals with ASD, leading to better understanding and support for those with auditory hypersensitivity. It may also inform the development of personalized auditory experiences and interventions.
QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices
Computer Vision
Graph Learning
Theory
- QSV replaces the traditional attention mechanism with a single learned quaternion per token, coupling attention and message transformation.
- Ablation studies indicate that the transport function is essential for model accuracy, while the routing weight can be simplified.
- QSV shows competitive performance on CIFAR-10 and CIFAR-100 but does not outperform standard attention mechanisms on similar graphs.
- The model's architecture leverages sparse kNN graphs on concentric Fibonacci spheres for efficient computation.
Read more
QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices
Summary
The paper introduces Quat-Sphere-Vision (QSV), a novel attention mechanism that utilizes a single learned unit quaternion per token to replace the traditional three-projection model (WQ, WK, WV) in standard attention. This approach couples the attention weight and message transformation into a single geometric object, enhancing efficiency and performance on spherical lattices. The authors demonstrate that the relative quaternion between tokens provides both the attention logit and a transport mechanism for feature aggregation. Through extensive ablation studies, they reveal that the transport function is critical for maintaining accuracy, while the routing weight derived from the quaternion can be substituted with uniform averaging without significant loss in performance. The model is evaluated on CIFAR-10 and CIFAR-100 datasets, showing that while QSV performs well, it does not surpass traditional attention mechanisms on the same graph or flat 2D lattices. The findings emphasize the importance of the transport channel in the QSV model and suggest that the spherical layout may limit generalization on fine local textures.
Methodology
The authors developed the QSV model by integrating a single unit quaternion for each token, which computes both the attention logit and the feature transport through a sandwich product. They employed a sparse multi-shell spherical image model and utilized Riemannian Adam for training the quaternion parameters. The model was evaluated through matched ablation studies and comparisons with standard attention mechanisms on CIFAR datasets.
Results
The QSV model achieved an accuracy of 85.9% on CIFAR-100, which is lower than the 87.3% achieved by standard attention on the same graph and 91.1% on a flat 2D lattice. The removal of the transport function resulted in a significant drop in accuracy (4.3 points on CIFAR-10 and 4.1 points on CIFAR-100), indicating its critical role in the model's performance.
Implications
The findings suggest that while QSV offers a novel approach to attention mechanisms, further improvements are needed to enhance its performance relative to traditional methods. The insights into the importance of transport functions could inform future designs of attention models, particularly in scenarios involving spherical data representations.
I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?
Theory
Reinforcement Learning
Generative Models
- Introduces a latent variable model for JEPA to study causal state recovery.
- Develops a general information-theoretic objective combining likelihood and entropy maximization.
- Establishes identifiability conditions for recovering latent causal states.
- Implements an action-modulated JEPA (A-JEPA) based on theoretical findings.
Read more
I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?
Summary
This paper investigates the ability of joint-embedding predictive architectures (JEPAs) to recover underlying causal states from observations conditioned on actions. The authors introduce a latent variable model where high-dimensional observations are generated from latent causal states influenced by action-conditioned transition mechanisms. They propose a general information-theoretic objective that combines conditional likelihood maximization with entropy maximization to learn predictive representations while preserving latent state information. The paper establishes identifiability conditions under which the learned representations can recover the latent causal states up to invertible transformations and permutations. A key finding is that sufficient action-induced variation in transition mechanisms is crucial for identifiability. The authors instantiate their objective with an action-modulated Gaussian additive-noise model, resulting in the action-modulated JEPA (A-JEPA). Extensive experiments on synthetic data validate the theoretical findings and demonstrate robustness to moderate violations of identifiability conditions, as well as improved state recovery and transferability to unseen transition mechanisms.
Methodology
The authors formulate a latent causal generative model to describe action-conditioned dynamics. They develop an information-theoretic objective that combines conditional likelihood maximization with entropy maximization. Identifiability conditions are established theoretically, and the model is instantiated with an action-modulated Gaussian additive-noise framework. Experiments are conducted on synthetic data to validate the theoretical results.
Results
The experiments confirm that the proposed A-JEPA can recover latent causal states under the identified conditions and show robustness to moderate violations. Additionally, visual benchmarks indicate significant improvements in component-wise latent-state recovery and strong transferability to new transition mechanisms.
Implications
This work provides a theoretical foundation for understanding how JEPAs can learn causal mechanisms from observational data, which is crucial for developing robust world models that can generalize to unseen environments. The findings may have applications in reinforcement learning, robotics, and other fields requiring causal inference.