AI-generated summaries
Today's ML research,
without the noise.
Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.
48
Papers today
8h
Update frequency
7
Days of history
Training Crossroads for Recurrent Vision Transformers: Recurrence, Neural ODEs, and Deep Supervision
Computer Vision
Efficient ML
Theory
- Recurrent Vision Transformers (bViT) reduce parameter count while maintaining competitive accuracy.
- Standard ViTs outperform bViTs in FLOPs-constrained environments, while bViTs excel in memory-constrained scenarios.
- Higher-order ODE solvers introduce architectural biases rather than improving numerical accuracy.
- Stage-wise deep supervision aids in maintaining robustness beyond training but does not improve nominal accuracy.
Read more
Training Crossroads for Recurrent Vision Transformers: Recurrence, Neural ODEs, and Deep Supervision
Summary
This paper investigates the performance of single-block recurrent Vision Transformers (bViT) compared to standard Vision Transformers (ViTs) under various training and inference regimes. The authors focus on three main questions: when does recurrence outperform independently parameterized depth, the impact of using an ODE solver for training residual recurrent blocks, and the trade-off between robustness and nominal accuracy. The study finds that while standard ViTs are preferable under computational constraints (FLOPs), bViTs provide a better accuracy-to-parameter trade-off when memory is limited. The research also highlights that higher-order solvers in the context of Neural ODEs act more as architectural biases rather than improving numerical accuracy. Additionally, the implementation of stage-wise deep supervision does not enhance nominal accuracy but helps maintain performance beyond the training horizon, contrasting with naive recurrence methods that lead to performance collapse. Overall, the paper emphasizes the importance of understanding the trade-offs in model design and training strategies for efficient image recognition.
Methodology
The authors fixed the architecture of a single-block recurrent Vision Transformer (bViT) and compared it with standard ViTs across different training and inference paradigms using the CIFAR-100 dataset. They conducted experiments to analyze the impact of resource constraints on model performance, the behavior of Neural ODEs in training, and the effects of deep supervision.
Results
The experiments revealed that standard ViTs are more effective when computational resources are prioritized, while bViTs offer a favorable trade-off in terms of accuracy and parameter efficiency under memory constraints. The study also found that higher-order solvers do not uniformly improve performance and that deep supervision helps maintain accuracy beyond the training horizon.
Implications
The findings suggest that model design should consider the specific constraints of deployment environments, such as memory and computational resources. The insights into the behavior of recurrent architectures and the use of Neural ODEs could inform future research in efficient model training and architecture design for image recognition tasks.
Agentic Reinforcement Learning with Self-Distilled Reward Shaping
Reinforcement Learning
Large Language Models
NLP
- ADRS provides a structured approach to assign token-level credit in reinforcement learning, addressing the challenge of sparse rewards.
- The framework incorporates score calibration, reliability estimation, and credit integration to enhance the learning process.
- Experiments show significant performance improvements in long-horizon tasks across multiple settings.
- ADRS maintains skill-free rollouts during inference, ensuring practical applicability in real-world scenarios.
Read more
Agentic Reinforcement Learning with Self-Distilled Reward Shaping
Summary
This paper introduces Agentic Reinforcement Learning with Self-Distilled Reward Shaping (ADRS), a novel framework aimed at improving the performance of large language model (LLM) agents in interactive tasks. Traditional reinforcement learning methods often provide sparse trajectory-level rewards, making it challenging to assign credit to intermediate decisions. ADRS addresses this issue by utilizing privileged skills to rescore tokens from skill-free trajectories, allowing for denser supervision. The framework includes three main components: score calibration, reliability estimation, and credit integration. Score calibration ensures that teacher scores are comparable across interaction steps, while reliability estimation assesses the confidence of the teacher in relation to realized returns. Credit integration merges privileged guidance into the native reinforcement learning credit construction process. Experiments conducted across three interactive benchmarks demonstrate that ADRS consistently enhances performance on long-horizon tasks, showing robustness across various RL backbones, reduced-data settings, unseen tasks, and extended training periods.
Methodology
ADRS employs a self-distilled reward shaping approach that normalizes and calibrates privileged token scores at each interaction step. It utilizes a Teacher Value Advantage (TVA) gate to modulate these scores based on the association between teacher confidence and realized returns. The gated token signals are then integrated into the native reinforcement learning credit construction process, allowing for a more effective reward-to-advantage transformation.
Results
The implementation of ADRS led to consistent performance improvements across three interactive benchmarks, particularly in long-horizon tasks. The gains were observed across different reinforcement learning backbones and in scenarios with limited data or unseen tasks, indicating the robustness and versatility of the framework.
Implications
The ADRS framework has the potential to enhance the training of LLM agents in various interactive applications, such as evidence search, navigation, and web interaction. By providing a more effective method for credit assignment, it could lead to more efficient learning and improved agent performance in complex environments.
ConformalShift: Targeted Event Reordering Against Adaptive ECG Monitoring
Time Series
- ConformalShift is an adversarial attack that reorders events in ECG monitoring systems to suppress the detection of ventricular ectopic beats.
- The attack operates without modifying ECG waveforms, labels, or classifier outputs, focusing solely on the timing of event feedback.
- Experimental results show that ConformalShift can suppress 66.7% of eligible targets for Extra Trees and 60.0% for HistGradientBoosting, compared to much lower rates with random scheduling.
- The study highlights the vulnerability of adaptive conformal prediction systems to event order manipulation, emphasizing the need for robust defenses against such attacks.
Read more
ConformalShift: Targeted Event Reordering Against Adaptive ECG Monitoring
Summary
The paper introduces ConformalShift, a novel adversarial attack targeting adaptive ECG monitoring systems that utilize conformal prediction. The authors highlight the vulnerability of these systems to event reordering, which can suppress the detection of clinically significant ventricular ectopic beats without altering the ECG waveforms, labels, or classifier outputs. By manipulating the order of authentic events, ConformalShift aims to lower the ventricular threshold before a target event is evaluated, thereby compromising the monitor's ability to recover missed heartbeat classes. The study evaluates the effectiveness of this attack on two ECG datasets, demonstrating that it significantly outperforms random scheduling in suppressing targets. The findings underscore the importance of considering event order in the design of adaptive monitoring systems, revealing a previously overlooked attack surface that can be exploited in healthcare settings.
Methodology
The authors developed a constrained search method to identify feasible event reorderings that would lower the ventricular threshold for specific target events. They evaluated the effectiveness of ConformalShift using two ECG datasets and two frozen classifier families, comparing the attack's performance against random scheduling and analyzing the impact of the displacement budget on attack success.
Results
ConformalShift was able to suppress a significant percentage of ventricular ectopic beats in both datasets, achieving suppression rates of 66.7% and 60.0% for different classifiers, compared to random scheduling rates of 4.4% and 12.0%. The results indicate that the timing of authentic information can be exploited to compromise adaptive monitoring systems, even when the content remains unchanged.
Implications
The findings suggest that healthcare monitoring systems relying on adaptive conformal prediction need to account for potential adversarial manipulations of event order. This could lead to the development of more robust monitoring frameworks that are resilient to such attacks, ensuring the reliability of automated ECG analysis in clinical settings.
Sample Complexity of Multicalibration for Multilevel Properties
Theory
- Introduces a framework for multicalibration across multiple interrelated properties.
- Establishes matching upper and lower bounds for sample complexity in multicalibration tasks.
- Demonstrates that achieving multicalibration error requires exponentially many samples relative to the number of properties.
- Presents a randomized learner that significantly reduces sample complexity for finite group families.
Read more
Sample Complexity of Multicalibration for Multilevel Properties
Summary
This paper investigates the concept of multicalibration, which ensures that a predictor remains unbiased across multiple groups simultaneously, particularly for a sequence of k properties that are interrelated. The authors establish a framework for multicalibration that includes properties identifiable sequentially, such as mean, variance, and skewness. They derive matching upper and lower bounds for sample complexity, showing that achieving a multicalibration error Ξ΅ requires eβ¦(Ξ΅β(k+2)) samples, even with a limited number of binary groups. Conversely, they present a randomized learner that can achieve the same error with O(Ξ΅β(k+2) + Ξ΅β2 log |G|) samples, indicating that the sample complexity is eΞ(Ξ΅β(k+2)) for polynomial-size group families. The theoretical results are instantiated through three canonical examples, demonstrating the practical applicability of their findings in real-world scenarios.
Methodology
The authors develop a theoretical framework for multicalibration, focusing on the sequential conditional identifiability of properties. They derive sample complexity bounds using minimax analysis and construct a randomized learner to achieve multicalibration with reduced sample requirements. The methodology includes rigorous proofs and analysis of the proposed algorithms.
Results
The paper establishes that for k properties, the sample complexity required to achieve multicalibration error Ξ΅ is eΞ(Ξ΅β(k+2)) for polynomial-size group families. The authors also provide a randomized learner that can achieve this error with significantly fewer samples, demonstrating the feasibility of multicalibration in practical applications.
Implications
The findings have significant implications for areas where multiple related statistical properties need to be predicted accurately, such as healthcare, finance, and machine learning applications that require fairness and calibration across diverse groups. The results can enhance the reliability of predictive models in decision-making processes.
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
NLP
Large Language Models
Reinforcement Learning
- SPOT reformulates selective supervision in OPD into two decisions: where to probe and how to distill outcomes.
- The method employs a position-level score to prioritize probing based on teacher uncertainty and student mismatch.
- SPOT's closed-form target adjusts the teacher distribution based on verified student continuations, enhancing learning efficiency.
- Empirical results show SPOT outperforms existing methods in reasoning tasks, indicating improved solution coverage.
Read more
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
Summary
The paper introduces SPOT (Sparse Probing and Outcome-calibrated Targets for On-Policy Distillation), a novel approach to improve on-policy distillation (OPD) by addressing the limitations of standard reverse-KL training. Traditional OPD often leads to insufficient probability assignment to plausible continuations, which can hinder solution coverage. SPOT proposes a three-stage process: acquisition, exploration, and exploitation. During acquisition, a position-level score is calculated based on teacher entropy, the probability mass of top-k candidates, and student-teacher mismatch to prioritize probing positions. In the exploration phase, candidates are evaluated through verifier-scored student continuations. Finally, the exploitation phase uses these outcomes to create a KL-regularized target that favors candidates with better downstream results while remaining anchored to the teacher distribution. The empirical results demonstrate that SPOT significantly enhances reasoning performance across various student models and benchmarks, achieving the highest macro Pass@8 and competitive average accuracy.
Methodology
SPOT utilizes a three-stage procedure: acquisition, exploration, and exploitation. In acquisition, it calculates a position-level score that combines teacher entropy, top-k probability mass, and student-teacher mismatch to prioritize probing positions. During exploration, it evaluates candidates by rolling out student continuations and scoring them with a verifier. In exploitation, it derives a KL-regularized target that adjusts the teacher distribution based on the outcomes of the verified continuations.
Results
SPOT achieved the highest macro Pass@8 across three evaluated Qwen student scales and ranked highest or second-highest in macro Avg@8 among compared methods. These results indicate that SPOT not only improves reasoning performance but also enhances solution coverage while maintaining competitive accuracy.
Implications
The findings suggest that SPOT could be applied to enhance the training of smaller models by effectively transferring reasoning capabilities from larger models, potentially benefiting various applications in natural language processing and machine learning.
Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning
Multimodal
- Introduces a principled constraint definition to address one-to-many mapping ambiguities in cross-representation learning.
- Develops CoCoEvolve, a co-evolution framework that leverages cross-representation agreement as a self-supervision signal.
- Presents CoCoEvolve@Eval, a systematic evaluation suite for assessing performance across multiple tasks.
- Demonstrates significant performance gains in cross-representation understanding, both in training and test settings.
Read more
Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning
Summary
The paper introduces CoCoEvolve, a framework designed to enhance consistency in cross-representation learning among chart images, tabular data, and visualization code. The authors address the inherent challenges of one-to-many relationships in these modalities, where supervision is often ambiguous and costly. CoCoEvolve proposes a principled constraint definition that establishes explicit one-to-one correspondences, allowing for optimization based on agreement between representations without requiring additional annotations. The framework operates in two phases: CoCoEvolve@Train, which performs co-evolution during training, and CoCoEvolve@Test, which applies consistency objectives during inference for co-optimization. Additionally, the authors present CoCoEvolve@Eval, a comprehensive evaluation suite for six cross-representation tasks. Experimental results across four benchmarks demonstrate significant performance improvements, with gains of up to 37.91% on non-overlapping test sets and up to 46.88% in out-of-domain scenarios, showcasing the framework's effectiveness in enhancing model understanding across representations.
Methodology
CoCoEvolve employs a consistency-driven co-evolution approach that optimizes models through agreement between chart, table, and code representations. It defines explicit one-to-one correspondences to ground the inherent one-to-many relationships and utilizes a self-supervised learning paradigm that does not rely on labeled data. The framework operates in a cyclical manner, enforcing global semantic correctness during both training and testing phases.
Results
The experiments reveal that CoCoEvolve significantly enhances model performance in cross-representation tasks, achieving up to 37.91% improvement on non-overlapping test sets and up to 46.88% gains in out-of-domain evaluations, indicating its effectiveness and robustness.
Implications
The findings suggest that CoCoEvolve can be applied in various domains where cross-representation understanding is crucial, such as data visualization, scientific research, and financial reporting. The framework's ability to operate without extensive labeled data could facilitate broader adoption in real-world applications.
To Describe or Construct Statistical Learning Models Using the Category-theoretical Language
Theory
- Utilizes category-theoretical language to describe statistical learning models.
- Introduces a unified descriptive approach for constructing complex models.
- Summarizes classical statistical learning models and algorithms for accessibility.
- Constructs a Transformer model as a practical example of the proposed approach.
Read more
To Describe or Construct Statistical Learning Models Using the Category-theoretical Language
Summary
This paper explores the intersection of statistical learning and category theory, aiming to provide a new perspective on understanding and constructing statistical learning models. The author summarizes classical statistical learning models and algorithms, making the content accessible to newcomers in the field. The primary contributions include the use of categorical language to describe machine learning models and the introduction of a unified approach for constructing complex models, such as language models. The paper emphasizes that while it does not aim to establish a strict theoretical framework, it borrows notations from category theory to facilitate the description of intricate machine learning architectures. The author discusses various classical models, including Linear Discriminant Analysis, Gaussian Mixture Models, and Neural Networks, and presents a nontrivial example of constructing a Transformer model using a localization trick. The paper serves as a call to researchers from diverse fields, particularly mathematics, to engage with statistical learning research.
Methodology
The paper employs category theory as a descriptive language to analyze and construct statistical learning models. It introduces specific notations and concepts from category theory to facilitate understanding and communication of complex models in machine learning.
Results
The paper successfully illustrates how category theory can be applied to describe and construct statistical learning models, providing a new lens through which to view existing algorithms and models. It demonstrates the construction of a Transformer model using the proposed categorical approach.
Implications
The findings suggest that category theory can enhance the understanding and development of statistical learning models, potentially attracting researchers from mathematics and other fields to contribute to advancements in machine learning. This interdisciplinary approach may lead to innovative methodologies and improved model designs.
When Proxy Prediction Becomes Equation Reconstruction: Diagnostics and Residual Learning for Factor-Derived Proxy Supervision
Theory
- High predictive accuracy in proxy supervision may indicate equation reconstruction rather than robustness to degraded inputs.
- A comprehensive diagnostic framework is introduced to evaluate models under controlled factor degradation.
- The RASPL framework effectively combines formula preservation with adaptive contextual corrections.
- The proposed methods demonstrate superior performance in terms of degradation robustness and computational efficiency.
Read more
When Proxy Prediction Becomes Equation Reconstruction: Diagnostics and Residual Learning for Factor-Derived Proxy Supervision
Summary
This paper addresses the challenge of using proxy targets derived from known domain factors in scientific machine learning, particularly in the context of soil-loss modeling using the Revised Universal Soil Loss Equation (RUSLE). The authors identify a critical issue where high predictive accuracy may stem from the model's ability to reconstruct the proxy-generating equation rather than demonstrating robustness to degraded factor information. To tackle this, they introduce a diagnostic framework that evaluates model performance under various conditions of factor degradation. The framework incorporates several evaluation metrics, including degradation robustness scoring and tail-error analysis. The authors propose a novel approach called RASPL (Residual Adaptive Supervised Proxy Learning), which retains the degraded formula estimate as a prediction anchor while learning an adaptively gated contextual correction. This method significantly outperforms traditional direct prediction methods and exhibits enhanced robustness against degraded inputs. The study highlights the importance of formula preservation in designing robust learning systems that utilize factor-derived proxy targets.
Methodology
The authors developed a diagnostic framework that includes evaluations under complete, noisy, coarsened, and masked factor conditions. They introduced RASPL, a formula-preserving residual framework that uses degraded proxy estimates as anchors and learns contextual corrections through compact statistical encoders and convolutional representations. The evaluation metrics included tail-error analysis and degradation robustness scoring.
Results
The RASPL framework outperformed matched direct prediction methods, showing stronger degradation robustness and lower tail-error metrics. The compact statistical encoder achieved the highest macro-averaged R2 and lowest computational cost, while the convolutional encoder provided the best degradation robustness and lowest Tail95 mean absolute error (MAE).
Implications
The findings suggest that formula preservation is crucial for developing robust machine learning models in scientific applications, particularly in environments where direct measurements are scarce. This approach could enhance the reliability of predictions in various fields, including environmental science and resource management.
ArborEnum: Decision Tree Rashomon Sets over Continuous Features
Interpretability
Optimization
Theory
- Introduces ArborEnum, the first algorithm for enumerating decision tree Rashomon sets over continuous features.
- Demonstrates that existing methods fail to capture the full potential of continuous features due to binarization.
- Offers both exact enumeration and approximate solutions, achieving significant speedups in computation.
- Highlights the importance of continuous thresholds in identifying predictive multiplicity and important variables.
Read more
ArborEnum: Decision Tree Rashomon Sets over Continuous Features
Summary
The paper introduces ArborEnum, a novel algorithm for enumerating decision tree Rashomon sets over continuous features, addressing a significant limitation in existing methods that typically require binarization of data. The Rashomon effect highlights that multiple models can achieve similar performance on a given task, which has implications for model robustness and feature importance. Previous algorithms for enumerating Rashomon sets have focused on binary features, leading to a loss of information and increased computational complexity when applied to continuous features. ArborEnum leverages the ordered structure of continuous features to efficiently enumerate these sets, providing both exact and approximate solutions. The authors also propose an anytime algorithm that progressively refines candidate thresholds, allowing for a more detailed approximation of the Rashomon set. Experimental results demonstrate that ArborEnum achieves substantial speedups compared to existing methods while maintaining high recall rates, revealing that coarse binarization can overlook important trees and features. The code for ArborEnum is made publicly available, promoting further research in this area.
Methodology
ArborEnum employs threshold bounds to enumerate decision trees whose objectives fall within a specified tolerance of optimality. It propagates information from evaluated thresholds to nearby ones, allowing for efficient pruning of candidate splits. The algorithm can compute results either exactly or approximately, with an anytime variant that refines the binarization progressively.
Results
The experiments show that ArborEnum achieves median speedups of 270x over existing enumeration methods. Exact enumeration is feasible for many datasets, while approximations recover nearly all trees at a fraction of the computational cost. The study also reveals that coarse binarization can miss significant predictive features and trees, underscoring the effectiveness of ArborEnum's approach.
Implications
The findings suggest that ArborEnum can enhance model interpretability and robustness by providing a comprehensive view of high-performing models. This can be particularly useful in applications where understanding model behavior and feature importance is critical, such as in healthcare and finance.
SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts
Reinforcement Learning
Large Language Models
Efficient ML
- Introduction of SpecRoll, a speculative rollout engine for RL.
- Implementation of a two-timescale adaptation mechanism for efficient corrections.
- Achieves 1.26Γβ2.15Γ generation speedup and 1.21Γβ2.04Γ end-to-end speedup over vanilla GRPO.
- Outperforms FastGRPO in all tested settings.
Read more
SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts
Summary
The paper introduces SpecRoll, a novel speculative rollout engine designed to enhance the efficiency of reinforcement learning (RL) in large language models. Traditional autoregressive rollout generation is identified as a significant bottleneck in RL post-training, particularly when using Group Relative Policy Optimization (GRPO). SpecRoll addresses this issue by implementing a two-timescale adaptation mechanism that includes lightweight future-token heads for generating parallel proposals and a Reflex module that utilizes delayed verifier feedback for trajectory-local hidden-state corrections without backpropagation. This dual adaptation approach allows for rapid corrections of transient mismatches while reserving slower updates for persistent policy drift. The methodology also incorporates concurrency-aware sparse-tree verification to maintain the target rollout distribution. Experimental results demonstrate that SpecRoll achieves substantial speedups in generation and end-to-end processing times across various models and datasets, outperforming existing methods like FastGRPO. The findings highlight the effectiveness of combining fast and slow adaptation paths, providing a significant advancement in the efficiency of RL rollouts.
Methodology
SpecRoll employs a two-timescale adaptation mechanism that includes lightweight future-token heads for generating multiple proposals in parallel and a Reflex module for making trajectory-local corrections based on delayed verifier feedback. The system also utilizes concurrency-aware sparse-tree verification to ensure the target rollout distribution remains unchanged while adapting to policy updates.
Results
SpecRoll achieves generation speedups ranging from 1.26Γ to 2.15Γ and end-to-end speedups from 1.21Γ to 2.04Γ compared to vanilla GRPO. It consistently outperforms FastGRPO across all 15 matched model-dataset settings, with an average end-to-end gain of 1.18Γ.
Implications
The advancements presented in SpecRoll could significantly improve the efficiency of RL applications in large language models, making it feasible to deploy more complex models in real-time scenarios. This could enhance applications in natural language processing, automated reasoning, and other areas where RL is applied.
Design-Time Optimization of Deep Neural Networks for Intermittent Learning on Microcontrollers
Optimization
Efficient ML
- Introduces a method for optimizing DNNs for intermittent learning on MCUs.
- Combines energy prediction with multi-objective optimization for design-time efficiency.
- Achieves a 16.6% error rate in energy consumption predictions, facilitating reliable DNN architecture selection.
- Extends energy prediction to include both forward and backward passes for on-device training.
Read more
Design-Time Optimization of Deep Neural Networks for Intermittent Learning on Microcontrollers
Summary
This paper presents a novel method for optimizing deep neural networks (DNNs) specifically for intermittent, energy-autonomous learning on microcontroller units (MCUs). The authors address the challenge of executing AI applications in mobile environments where energy supply is unpredictable, particularly in energy-harvesting scenarios. The proposed approach integrates a hardware-aware energy prediction model with multi-objective optimization (MOO) to facilitate offline DNN optimization without the need for extensive deployment and testing on target MCUs. The energy predictor estimates per-layer energy consumption for both inference and training, accounting for intermittent checkpointing overhead. Validation is conducted using autoencoders for anomaly detection on a Cortex-M4 MCU, achieving a weighted absolute percentage error of 16.6%, which is deemed sufficient for reliable architecture selection under intermittent energy constraints. This work effectively bridges the gap between MOO, automated DNN design, and deployment on energy-harvesting systems, enabling autonomous AI at the edge.
Methodology
The authors developed a lightweight energy prediction model that estimates per-layer energy consumption for DNN inference and training. This model is integrated into a multi-objective optimization framework that allows for the design of DNN architectures while considering intermittent energy constraints. The approach does not require physical deployment or measurements, making it efficient for design-time optimization.
Results
The energy prediction model demonstrated a weighted absolute percentage error of 16.6% when validated on a Cortex-M4 MCU, indicating its reliability for architecture selection. The methodology allows for the design of DNNs that can effectively operate under intermittent energy conditions, enabling on-device training and inference.
Implications
This research has significant implications for the deployment of AI in remote or mobile applications where energy supply is limited. It enables the development of more efficient DNN architectures that can operate autonomously in energy-harvesting environments, paving the way for advanced edge AI applications.
Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't
Large Language Models
Efficient ML
Theory
- Strong alignment of representations across transformer sizes, but weak alignment of parameters.
- Dense weight projection is functionally destructive due to basis mixing that disrupts model structure.
- No effective zero-shot correction exists after applying the best-fit linear operator.
- A two-lever approach for model conversion significantly outperforms existing methods.
Read more
Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't
Summary
This paper investigates the transferability of knowledge between different sizes of transformer models within the Pythia family, specifically focusing on the conversion from a 1.4B parameter model to a 410M parameter model. The authors characterize the transfer process through four main steps: (1) they find that representations align strongly across sizes (ridge R2 = 0.84), while parameters do not align well (R2 = 0.25β0.39); (2) they diagnose the failure of dense weight projection, attributing it to basis mixing that disrupts the model's structural integrity; (3) they demonstrate that after applying the best-fit linear operator, weight residuals are statistically indistinguishable from noise, indicating no effective zero-shot correction exists; and (4) they propose a two-lever approach for effective conversion, comprising least-squares compensation and variance-preserving rescale, which significantly outperforms traditional methods like weight subcloning. The findings suggest that initialization plays a crucial role in the conversion process, and the proposed methods can achieve better performance with fewer training tokens compared to training from scratch.
Methodology
The authors conducted a controlled characterization of the transfer process by analyzing the Pythia model family. They employed statistical methods to measure representation alignment and parameter similarity, and they diagnosed the failure of dense weight projection through mechanical verification. The conversion process was decomposed into two independent levers: least-squares compensation and variance-preserving rescale, which were tested against traditional weight subcloning methods.
Results
The study found that representations between the larger and smaller models aligned strongly, while parameters did not. The proposed two-lever conversion method (Compensated Selection) outperformed weight subcloning, achieving better performance with fewer tokens and demonstrating a significant advantage over training from scratch, especially at lower budgets.
Implications
The findings have implications for the efficient deployment of transformer models, suggesting that knowledge transfer techniques can reduce the computational cost associated with training multiple model sizes from scratch. This could lead to more resource-efficient practices in model development and deployment in various applications.
Sphere Retraction Normalizations
Theory
Optimization
Large Language Models
- Unification of residual connections and GeoNorm as retraction-based updates.
- Introduction of Proj-SpheretNorm and Cay-SpheretNorm as alternative retraction methods.
- Development of p-SpheretNorm, a flexible family of normalization layers.
- Empirical results indicate Proj-SpheretNorm achieves the best performance on nanoGPT.
Read more
Sphere Retraction Normalizations
Summary
This paper addresses the challenge of training deep neural networks stably by proposing Sphere Retraction Normalizations (SpheretNorms). The authors unify residual connections and Geodesic Normalization (GeoNorm) within a single framework, demonstrating that GeoNorm can be interpreted as a Riemannian residual connection on a hypersphere. They introduce two specific retraction methods: Proj-SpheretNorm, based on metric projection, and Cay-SpheretNorm, based on the Cayley retraction. Furthermore, they generalize these methods into a single-parameter family called p-SpheretNorm, which allows for flexible tuning of the normalization process. The paper shows that Proj-SpheretNorm outperforms existing lightweight deep connection schemes in terms of validation loss and accuracy on the nanoGPT model, indicating that the exponential map is not the optimal choice for spherical residual streams. The findings suggest a new direction for improving the stability and performance of deep learning architectures.
Methodology
The authors employ a geometric approach to deep learning by utilizing retraction maps on a hypersphere. They first unify residual connections and GeoNorm, then derive Proj-SpheretNorm and Cay-SpheretNorm through specific retraction methods. They further generalize these methods into p-SpheretNorm, which allows for a tunable parameter to adjust the normalization process. The methods are evaluated using the nanoGPT architecture across pretraining and downstream tasks.
Results
The experiments conducted on nanoGPT demonstrate that Proj-SpheretNorm achieves the lowest pretraining and validation losses compared to other methods. It also shows high accuracy on downstream tasks, indicating its effectiveness as a normalization technique. The results suggest that the choice of retraction significantly impacts the performance of deep neural networks.
Implications
The findings of this paper could lead to improved training methodologies for deep learning models, particularly in architectures like Transformers. By providing a flexible normalization approach, it opens avenues for further research into the geometric properties of neural networks and their impact on stability and performance.
Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification
Computer Vision
- fANOVA is utilized to analyze the impact of design choices on model performance in multi-label classification.
- Datasets cluster based on their sensitivity to design choices rather than their performance levels.
- Fine-tuning strategy and architecture are critical for large-scale datasets, while initialization is key for data-limited scenarios.
- The study provides dataset-aware guidelines for model design choices, enhancing interpretability in benchmarking results.
Read more
Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification
Summary
This paper addresses the challenge of benchmarking deep learning models for multi-label classification (MLC) of remote sensing images (RSI), which often results in rankings that do not generalize across datasets. The authors employ functional analysis of variance (fANOVA) to quantify the contributions of various design choices and their interactions to performance variability. They conduct empirical analyses involving 48 and 20 deep learning models across seven MLC RSI datasets, focusing on design choices such as network architecture, fine-tuning strategy, learning strategy, and initialization. The study constructs dataset meta-representations that reveal sensitivity profiles to design choices, showing that datasets cluster based on their response to these choices rather than overall performance. Key findings indicate that for large-scale datasets, fine-tuning strategy and architecture are the dominant factors, while initialization plays a crucial role in data-limited scenarios. The paper reframes conclusions from previous benchmarking studies to be dataset-conditional, providing actionable insights for model design choices tailored to specific datasets.
Methodology
The authors apply an extended fANOVA framework to analyze the performance of deep learning models across multiple MLC datasets. This involves computing fANOVA importance scores to create dataset meta-representations and conducting post-hoc analysis to identify patterns in design-choice sensitivity.
Results
The analysis reveals that datasets naturally group based on their sensitivity to design choices, with significant shifts in dominant factors depending on dataset size and complexity. For large datasets, fine-tuning and architecture are paramount, while initialization is crucial in smaller datasets. The findings also suggest that interactions between design choices are significant in intermediate dataset regimes.
Implications
The insights from this study can guide practitioners in selecting appropriate model configurations for specific datasets in remote sensing applications, ultimately improving the effectiveness of multi-label classification tasks. The dataset-aware approach enhances the interpretability of benchmarking results and informs future research in deep learning model design.
An Explainable LLM Agent Layer for Open-World Anomaly Detection in Oil Wells
NLP
Large Language Models
Interpretability
- Introduction of an LLM agent layer to enhance explainability in oil well anomaly detection.
- The LLM provides natural-language justifications and names for detected anomalies.
- Achieved 35.1% top-1 and 63.9% top-3 classification accuracy across nine classes.
- Demonstrated 89.7% novelty detection rate with stable naming for clustered anomalies.
Read more
An Explainable LLM Agent Layer for Open-World Anomaly Detection in Oil Wells
Summary
This paper presents an innovative approach to enhance anomaly detection in oil wells by integrating a Large Language Model (LLM) agent layer into existing Open-World Learning (OWL) pipelines. While previous methods effectively detected and classified anomalies, they lacked the ability to explain the reasoning behind their conclusions or provide actionable insights for operators. The proposed LLM agent, utilizing the Qwen3.5-397B-A17B Mixture-of-Experts model, processes structured sensor metrics and upstream classifications to generate natural-language justifications, critiques, and human-readable names for detected anomalies. The study evaluates the agent's performance across three distinct studies involving 989 real well-file segments from the 3W dataset, demonstrating its capability to achieve significant classification accuracy and precision in novelty detection. The findings indicate that the LLM can effectively validate upstream decisions, provide explanations in sensor-grounded language, and name novelty clusters, thereby addressing the explainability gap in operational settings. This work represents a crucial step towards the deployment of OWL pipelines in the oil and gas industry, enhancing both the interpretability and usability of machine learning models in critical applications.
Methodology
The study employs a Large Language Model (Qwen3.5-397B-A17B) as a companion to an established OWL pipeline. The LLM processes structured sensor metrics derived from the 3W dataset and evaluates upstream classifications to generate explanations, critiques, and consolidated names for detected novelties. The evaluation includes three studies focusing on classification, validation, and novelty detection.
Results
The LLM agent achieved 35.1% top-1 and 63.9% top-3 classification accuracy across nine classes, with a top-2 validation accuracy of 71.7% and precision of 0.91 across seven probed classes. The novelty detection rate was 89.7%, with stable naming for five out of seven hidden classes.
Implications
The integration of an LLM agent layer into anomaly detection pipelines can significantly improve the interpretability and usability of machine learning models in the oil and gas industry. This advancement may facilitate better decision-making and operational efficiency by providing clear explanations and actionable insights for operators.
SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors
Theory
Robotics
Efficient ML
- Introduces SJEPA, a hybrid symbolic-neural predictive framework for learning elegant latent dynamics.
- Focuses on learning the simplest adequate dynamics while preventing representation collapse.
- Demonstrates that joint representation-equation learning yields simpler symbolic dynamics with improved predictive performance.
- Provides a modular learning framework that supports various learning objectives and applications.
Read more
SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors
Summary
The paper introduces SJEPA, a novel reconstruction-free Joint-Embedding Predictive Architecture (JEPA) that aims to learn predictive representations with elegant dynamics. Unlike traditional JEPA models that utilize opaque neural predictors, SJEPA combines a symbolic governing law with a neural correction to capture dynamics that cannot be adequately expressed by the symbolic grammar alone. The framework emphasizes learning the simplest adequate dynamics by imposing representation constraints that prevent trivial solutions while favoring compact symbolic transitions. The authors formalize this approach through the concept of induced-dynamics complexity and analyze the potential for representation collapse due to unconstrained operator compression. Empirical results from controlled pendulum experiments demonstrate that SJEPA can discover simpler symbolic dynamics with lower long-horizon rollout errors compared to post-hoc fitting methods. Additionally, the framework allows for both alternating representation-equation learning and symbolic dynamics fitted to fixed representations, making it versatile for various applications. The findings suggest that SJEPA can effectively balance predictive fidelity, representation quality, and symbolic parsimony, offering a promising direction for future research in predictive modeling.
Methodology
SJEPA employs a hybrid transition model that combines a symbolic governing law with a neural correction. The learning process involves constrained operator compression to ensure that the predictive coordinates remain informative and non-collapsed. The framework supports both alternating representation-equation learning and symbolic dynamics fitted to fixed representations, allowing for flexibility in model training.
Results
The experiments conducted on controlled pendulum tasks revealed that SJEPA achieved significantly simpler symbolic dynamics with lower long-horizon rollout errors compared to traditional post-hoc fitting methods. The framework also demonstrated resilience to grammar misspecification, maintaining the representable symbolic mechanism while allowing the neural component to focus on residual dynamics.
Implications
SJEPA has potential applications in areas requiring compact latent dynamics and explicit structural constraints, such as robotics, control systems, and general-purpose world modeling. Its ability to balance predictive fidelity and symbolic parsimony may enhance the interpretability and efficiency of predictive models in various domains.
Robustness Emerges Early in Training Dynamics, but Is Not Preserved
Computer Vision
Optimization
Theory
- Identification of robustness fading phenomenon in shallow layers during training.
- Introduction of Early-Phase Stabilization (EPS) and Asymmetric Weight Reversion (AWR) strategies.
- Demonstrated improvements in robustness across multiple benchmarks and architectures.
- Effective in various applications including object detection and semantic segmentation.
Read more
Robustness Emerges Early in Training Dynamics, but Is Not Preserved
Summary
This paper addresses the challenge of robustness in deep neural networks, particularly against natural corruptions. The authors identify a phenomenon termed 'robustness fading,' where shallow layers of neural networks develop robust representations and flatter loss landscapes during early training, but these properties diminish as training progresses. To combat this issue, they propose a framework that includes two parameter-free strategies: Early-Phase Stabilization (EPS) and Asymmetric Weight Reversion (AWR). EPS locks in robust configurations by halting updates to shallow subnetworks early in training, while AWR reverts shallow subnetworks to earlier robust states during subsequent training. The authors conduct extensive experiments across various benchmarks and architectures, demonstrating that their methods significantly enhance robustness in downstream transfer tasks, dynamic adaptation, and various computer vision applications. The findings suggest that training dynamics play a crucial role in the emergence and preservation of robustness in neural networks.
Methodology
The authors analyze training dynamics by tracing weight trajectories and examining representation stability in shallow subnetworks. They propose two strategies: EPS, which stabilizes early robust configurations by halting updates, and AWR, which reverts shallow subnetworks to earlier states to recover lost robustness. These methods are implemented without modifying the model architecture or introducing additional parameters.
Results
The proposed framework significantly outperforms competitive baselines in terms of robustness across various corruption benchmarks. The methods generalize well across different network architectures and improve performance in downstream transfer tasks, dynamic adaptation, and real-world scenarios. The analysis reveals that the gains are associated with smoother loss landscapes and more stable representations under distribution shifts.
Implications
The findings suggest that preserving early-emergent robustness can enhance the reliability of deep neural networks in safety-critical applications, particularly in environments subject to natural corruptions. The proposed methods can be applied to improve performance in diverse tasks beyond classification, including object detection and semantic segmentation.
Continual-Learning Physics-Informed Neural Networks for Parameterized Partial Differential Equations
Optimization
Theory
Efficient ML
- Introduction of CL-PINN to improve the efficiency and accuracy of solving parameterized PDEs.
- Utilization of continual learning techniques to sequentially learn related PDE tasks.
- Implementation of Bayesian optimization for active parameter selection to enhance training efficiency.
- Demonstration of improved accuracy and reduced computational costs compared to existing methods.
Read more
Continual-Learning Physics-Informed Neural Networks for Parameterized Partial Differential Equations
Summary
This paper presents a novel approach to solving parameterized partial differential equations (PDEs) using Continual-Learning Physics-Informed Neural Networks (CL-PINNs). Traditional methods for approximating PDE solutions often suffer from inefficiencies, overfitting, and imbalanced accuracy across different parameter values. CL-PINNs address these challenges by treating PDE instances at varying parameters as related tasks and learning them sequentially. The methodology integrates several advanced techniques, including Bayesian-optimization-based active parameter selection, dynamic loss weighting, and sparse physics-constrained replay, to enhance task allocation and knowledge retention. The proposed model does not require observational data and is designed to operate efficiently across broad parameter domains. Evaluations on five benchmarks demonstrate that CL-PINN significantly reduces objective-loss queries compared to traditional grid-greedy searches and maintains higher accuracy across tasks, thus offering a practical solution for learning PDEs that generalize well across physical parameters.
Methodology
The CL-PINN framework employs continual learning principles by treating different parameterized PDE instances as related tasks. It incorporates Bayesian optimization for selecting parameters actively, dynamic loss weighting to prioritize tasks, and sparse replay mechanisms to mitigate forgetting of previously learned tasks. An optional parameter subnetwork is included to optimize task allocation and knowledge retention.
Results
The multi-seed evaluations on five benchmarks, including one continuous function and four parameterized PDEs, indicate that CL-PINN achieves significantly lower objective-loss queries compared to grid-greedy search methods. Furthermore, it demonstrates higher and more balanced solution accuracy across various parameter values, outperforming fixed-sampling and grid-greedy baselines.
Implications
CL-PINN offers a promising approach for efficiently learning solutions to parameterized PDEs, which can be beneficial in various scientific and engineering fields. Its ability to generalize across physical parameters could lead to the development of reusable physics-informed models, facilitating large-scale parameter studies and simulations.
Can Training Logs Make Model Comparisons More Precise?
Computer Vision
Theory
Efficient ML
- Introduces an arm-specific adjustment framework for model comparisons using training logs.
- Demonstrates that early training statistics can reduce uncertainty in model performance estimates.
- Highlights the risks of poor covariate selection, which can introduce noise rather than reduce it.
- Provides empirical evidence that arm-specific adjustments can tighten confidence intervals for performance differences.
Read more
Can Training Logs Make Model Comparisons More Precise?
Summary
This paper investigates the potential of using training logs from stochastic model training to enhance the precision of model comparisons. Traditional comparisons rely on repeated runs to estimate performance differences, which inherently include uncertainty due to variations in initialization and data order. The author proposes an arm-specific covariate adjustment method, where each model is adjusted using only its own training-log statistics, thereby avoiding the pitfalls of pooled adjustments that could obscure true model differences. The study is conducted through a factorial experiment involving three model architectures (ResNet-18, ViT-Tiny, ConvNeXt-Tiny) across three datasets (CIFAR-10, CIFAR-100, Tiny-ImageNet), totaling 450 runs. The findings indicate that early training logs can significantly reduce uncertainty in performance comparisons, although the effectiveness of covariate selection is crucial. The paper emphasizes that while training logs can improve model comparison precision, careful selection of covariates is necessary to avoid introducing noise.
Methodology
The methodology involves an arm-specific covariate adjustment framework where each model is adjusted based on its own training-log covariates. The study employs a factorial experimental design to evaluate the effectiveness of this adjustment across different model architectures and datasets. The analysis includes cross-fitted coefficient estimation and coverage diagnostics to ensure the reliability of the adjustments.
Results
The results indicate that using early training logs for covariate adjustment can significantly narrow confidence intervals for performance differences between models, particularly when sufficient runs are conducted. However, when the number of runs is limited, the adjustments may lead to wider confidence intervals, highlighting the importance of adequate run budgets for reliable estimation.
Implications
The findings suggest that training logs can be a valuable resource for improving the precision of model comparisons in machine learning. This approach could lead to more reliable evaluations of model performance, particularly in research and practical applications where model selection is critical. The emphasis on careful covariate selection also points to the need for further exploration in this area to optimize model comparison methodologies.
Output-Aware Rotation for INT2 KV-Cache Quantization
NLP
Large Language Models
Efficient ML
- OptR optimizes rotations in output space to minimize post-WO attention-output error.
- The method incorporates attention-equivalent key reparameterization to reduce quantization errors.
- OptR consistently improves existing rotation-based INT2 KV-cache methods across various benchmarks.
- The approach retains the paged KV-cache format, ensuring compatibility with existing systems.
Read more
Output-Aware Rotation for INT2 KV-Cache Quantization
Summary
This paper addresses the inefficiencies in key-value (KV) cache quantization for long-context large language models (LLMs), specifically focusing on INT2 quantization. The authors introduce OptR, an output-aware rotation method that minimizes the post-output projection (WO) attention-output error, which is crucial for maintaining model performance during inference. Unlike existing methods that optimize cache statistics or proxy errors before the complete attention readout, OptR directly targets the error that propagates through the attention mechanism and affects the final output. The method involves decomposing the post-WO attention-output error into key- and value-induced components and learning per-head orthogonal corrections through the entire INT2 quantization and attention path. Additionally, OptR employs an attention-equivalent key reparameterization to mitigate large channel-wise offsets without altering the softmax distribution. The authors demonstrate that OptR significantly enhances the performance of existing rotation-based INT2 KV-cache methods across multiple benchmarks while maintaining compatibility with paged KV-cache formats and incurring negligible inference overhead.
Methodology
OptR centers the keys before rotation and quantization, optimizing the rotation parameters to minimize the post-WO attention-output error. The method learns per-head orthogonal corrections and applies an attention-equivalent key reparameterization to reduce the impact of outliers on quantization errors. The rotations are optimized on calibration data while keeping model weights frozen during inference.
Results
The implementation of OptR led to significant improvements in accuracy on the AIME25 benchmark for the Qwen3-8B model, increasing accuracy from 17.33% to 66.67% with QuaRot and from 54.67% to 66.00% with OSCAR, compared to a baseline accuracy of 68.00% for BF16.
Implications
The findings suggest that OptR can enhance the efficiency and performance of KV-cache quantization in large language models, potentially enabling more effective long-context inference and broader applications in natural language processing tasks.
GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits
Graph Learning
Large Language Models
Interpretability
- Introduction of GoT-CD, a causal discovery method that focuses on complete candidate edge sets.
- Demonstration of the fragility of post-hoc fairness audits when relying on discovered graphs.
- GoT-CD outperforms classical and LLM baselines in terms of DAG validity and structural fidelity.
- The method successfully recovers critical unfair pathways that are often overlooked by other approaches.
Read more
GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits
Summary
This paper presents GoT-CD, a novel causal discovery method that leverages the Graph-of-Thoughts framework to improve the accuracy of causal graphs derived from observational data. The authors highlight the limitations of existing causal discovery methods, particularly in the context of fairness audits, where the integrity of specific pathways is crucial. GoT-CD generates multiple candidate graphs in parallel, evaluates them using a deterministic validity function, and merges them while enforcing acyclicity. The method is benchmarked against classical and large language model (LLM) baselines across five datasets, demonstrating superior structural fidelity. A significant finding is that while GoT-CD achieves high structural accuracy, it also successfully recovers critical unfair pathways that other methods miss, underscoring the importance of pathway recovery in fairness assessments. The authors argue that causal discovery methods should be evaluated not just on aggregate structural metrics but also on their ability to preserve fairness-relevant sub-structures, which is essential for reliable fairness audits.
Methodology
GoT-CD employs a Graph-of-Thoughts framework to generate multiple candidate graphs in parallel, scoring them with a deterministic validity function. It merges these graphs under a hard union constraint to prevent the introduction of non-existent edges and enforces acyclicity through post-processing before final commitment.
Results
GoT-CD produced valid directed acyclic graphs (DAGs) across all five benchmarks, achieving the highest DAG-valid F1 scores among LLM methods on the Asia, Alzheimerβs, and COVID-Respiratory datasets. Notably, it recovered the true unfair pathway in the Alzheimerβs benchmark, while other methods failed to identify this critical pathway.
Implications
The findings suggest that causal discovery methods must be rigorously evaluated for their ability to maintain fairness-relevant pathways, especially in clinical and decision-making contexts. This has implications for the design of fairness audits and the integration of causal discovery in predictive modeling.
The Sample Complexity of Distributionally Robust PAC Learning under Cressie--Read Divergences
Theory
- Establishes sample complexity bounds for distributionally robust PAC learning under CressieβRead divergences.
- Demonstrates the impact of robustness on Ξ΅-dependence in both realizable and agnostic learning scenarios.
- Extends previous results from Ο2-divergence to the broader CressieβRead family, closing gaps in existing literature.
- Shows that ordinary empirical risk minimization can achieve optimal sample complexity rates.
Read more
The Sample Complexity of Distributionally Robust PAC Learning under Cressie--Read Divergences
Summary
This paper investigates the sample complexity of distributionally robust PAC learning under the 0β1-loss framework, specifically when adversarial perturbations of the data distribution are constrained by CressieβRead divergences of order k > 1 and radius Ο β₯ 0. The authors establish tight sample-complexity bounds for both realizable and agnostic cases, which are dependent on the VC dimension d, target accuracy Ξ΅, and confidence Ξ΄. They demonstrate that ordinary empirical risk minimization achieves these bounds up to logarithmic factors. The paper reveals that robustness alters the Ξ΅-dependence in the realizable case from Ξ΅β1 to Ξ΅βkβ as Ξ΅ approaches 0, and in the agnostic case, it changes from Ξ΅β2 to Ξ΅βkβ for 1 < k < 2, while remaining at Ξ΅β2 for k β₯ 2. The authors extend previous results concerning the Ο2-divergence to all CressieβRead orders, closing upper-lower gaps and recovering standard PAC learning rates as Ο approaches 0. This work highlights the intricate relationship between statistical estimation of classification error and its amplification through robustness, providing a clearer understanding of the transition in agnostic rates.
Methodology
The authors utilize a theoretical framework based on VC dimension and f-divergences to analyze the sample complexity of distributionally robust PAC learning. They derive bounds for both realizable and agnostic learning scenarios, focusing on the CressieβRead family of divergences. The analysis involves understanding the relationship between ordinary classification error and robust risk, particularly how the sensitivity of this relationship affects sample complexity.
Results
The paper presents tight sample-complexity bounds for distributionally robust PAC learning, showing that the sample complexity depends on the VC dimension, target accuracy, and confidence level. The results indicate a significant change in Ξ΅-dependence due to robustness, with different behaviors observed based on the value of k in the CressieβRead divergences. The authors also recover standard PAC learning rates in the limit as the radius of perturbation approaches zero.
Implications
The findings have significant implications for the design of machine learning models that are robust to distributional shifts, particularly in applications where data may be subject to adversarial perturbations. This work can guide practitioners in determining the necessary sample sizes for achieving reliable performance in uncertain environments.
Robust Counterfactual Policy Optimisation via Nondeterministic Causal Models
Reinforcement Learning
Optimization
Theory
- Introduces probabilistic nondeterministic causal models (PNSCMs) for counterfactual inference in MDPs.
- Proposes a sensitivity analysis framework to separate latent confounding from irreducible stochasticity.
- Develops a practical optimization problem for robust counterfactual policy identification.
- Validates the approach using a sepsis treatment simulator with diabetes as a confounder.
Read more
Robust Counterfactual Policy Optimisation via Nondeterministic Causal Models
Summary
This paper addresses the limitations of existing counterfactual inference methods in sequential decision-making, particularly in the context of Markov Decision Processes (MDPs), which are inherently stochastic. Traditional approaches assume deterministic causal models, where randomness is attributed solely to latent variables. The authors propose a novel framework using probabilistic nondeterministic causal models (PNSCMs) that differentiate between latent confounding and irreducible stochasticity. They formalize counterfactual policy optimization under this new model and introduce a practical optimization problem aimed at identifying robust counterfactual policies through a sensitivity analysis framework. The methodology is validated using a sepsis treatment simulator, where diabetes status serves as a hidden global confounder. The results demonstrate that the derived policies maintain robustness when evaluated against the true environment, showcasing the effectiveness of the proposed approach in addressing confounding in policy evaluation.
Methodology
The authors formalize counterfactual inference for MDPs under PNSCMs and propose optimization procedures that maximize the worst-case counterfactual value while accounting for unobserved global confounders. They employ a sensitivity parameter to bound the influence of latent variables on transition probabilities, allowing for a more nuanced understanding of the stochastic nature of the environment.
Results
The evaluation on the sepsis treatment simulator indicates that the proposed robust counterfactual policies effectively account for confounding factors, leading to improved decision-making outcomes when compared to traditional deterministic approaches.
Implications
This work has significant implications for safety-critical domains such as healthcare, where ethical constraints limit experimental interventions. The proposed methods enhance the reliability of offline policy evaluations and could be applied to various fields requiring robust decision-making under uncertainty.
Population-Robust Feature Selection via Generalized Welfare Optimization
Optimization
Theory
Interpretability
- Introduction of PopFS, a method for robust feature selection across heterogeneous populations.
- Utilization of a tunable welfare objective to balance predictive benefits and protection for less served populations.
- Scalable optimization strategy that directly searches over discrete feature sets.
- Demonstrated strong performance improvements in population-average and worst-case scenarios.
Read more
Population-Robust Feature Selection via Generalized Welfare Optimization
Summary
This paper presents PopFS, a novel method for feature selection that addresses the challenge of deploying a shared feature set across heterogeneous populations. Traditional feature selection methods optimize for a single population, while existing robust approaches typically learn a single model for all populations. PopFS allows for a shared feature set that is robust to population differences, enabling each population to train its own model. The method incorporates a tunable welfare objective, allowing practitioners to balance overall predictive performance with the protection of populations that benefit least. PopFS employs multitask sparse learning to reduce the candidate feature pool and utilizes a direct search over discrete feature sets, achieving high accuracy while being computationally efficient. The method was evaluated across multiple population splits from various datasets, demonstrating improved performance in both average and worst-case scenarios. A case study on COVID-19 nowcasting illustrated that adjusting the welfare objective can enhance performance for underrepresented populations without sacrificing overall accuracy.
Methodology
PopFS employs a two-step approach: first, it uses multitask sparse learning to filter candidate features, and then it conducts a direct search over discrete feature sets using a GaussβNewton-ranked refit search. This allows for efficient optimization while considering the welfare of different populations.
Results
PopFS achieved up to 22% improvement in both average and worst-case population performance compared to baseline methods, with a runtime of less than 15 minutes regardless of the feature set size. In the COVID-19 nowcasting study, tuning the welfare objective resulted in a 40% performance increase for the least-served states without compromising overall performance.
Implications
The findings suggest that PopFS can be effectively applied in various domains where feature selection must consider multiple populations, such as healthcare, finance, and social sciences. The ability to tune the welfare objective allows for tailored solutions that can prioritize underrepresented groups, enhancing equity in predictive modeling.
Rethinking Reservoir Pruning: A Dynamical Perspective for Echo State Networks
Time Series
Efficient ML
Theory
- Dynamical Mode Pruning (DMP) offers a new approach to pruning ESNs by focusing on the dynamic contributions of neurons.
- DMP ranks neurons based on their influence on dominant transition modes, rather than static metrics.
- The method retains or improves forecasting accuracy while reducing redundant components in the reservoir.
- DMP highlights the importance of considering input-driven dynamics in the design of reservoir computing systems.
Read more
Rethinking Reservoir Pruning: A Dynamical Perspective for Echo State Networks
Summary
This paper addresses the challenge of over-parameterization in Echo State Networks (ESNs), which can lead to redundancy and inefficiencies in temporal prediction tasks. Traditional pruning methods often rely on static metrics, failing to account for the dynamic contributions of neurons during input-driven state transitions. The authors propose a novel method called Dynamical Mode Pruning (DMP), which evaluates neuron importance based on their contributions to dominant transition modes derived from a trajectory-averaged Jacobian Gramian. By removing low-impact neurons and retraining only the readout layer, DMP effectively reduces the reservoir's complexity while maintaining or improving forecasting accuracy. Experiments conducted on chaotic and real-world time-series datasets demonstrate that DMP not only preserves predictive performance but also enhances the efficiency of the ESN architecture, suggesting that a dynamical perspective is crucial for effective reservoir refinement.
Methodology
The authors introduce Dynamical Mode Pruning (DMP), which constructs a trajectory-averaged Jacobian Gramian to assess the contributions of neurons to the reservoir's dynamic behavior. Neurons with limited dynamical influence are pruned, and only the linear readout layer is retrained, allowing for efficient reservoir simplification.
Results
Experiments on various chaotic and real-world time-series benchmarks indicate that DMP can improve or maintain forecasting accuracy while significantly reducing the number of neurons in the reservoir. This demonstrates the effectiveness of a dynamical approach to pruning in enhancing the performance of ESNs.
Implications
The findings suggest that incorporating dynamical perspectives into reservoir pruning can lead to more efficient and effective ESN architectures, with potential applications in various temporal prediction tasks across different domains.
Simulation-free and finite-time diffusion model
Generative Models
Theory
Efficient ML
- Introduces a framework for constructing reference processes that achieve simulation-free training and finite-time generation.
- Reveals that score matching is not fundamental but emerges from the reversal of the reference process.
- Demonstrates that conditional flow matching is a small-noise limit of the proposed framework.
- Provides practical constructions for both Gaussian and non-Gaussian priors.
Read more
Simulation-free and finite-time diffusion model
Summary
This paper presents a novel framework for designing reference processes in generative diffusion models that allows for both simulation-free training and finite-time generation. Traditional methods often struggle to achieve these two goals simultaneously, typically requiring a trade-off between computational efficiency and generation quality. The authors propose a new approach where time-dependent conditional distributions are prescribed first, and a reference process is constructed to realize these distributions as marginals. This innovative perspective reveals that score matching is not a fundamental aspect of diffusion model training but rather emerges from reversing the reference process. Additionally, the framework shows that conditional flow matching can be derived as a small-noise limit of the proposed method. The paper includes practical constructions for Gaussian and non-Gaussian priors and demonstrates the effectiveness of the proposed framework through numerical experiments.
Methodology
The authors reverse the conventional design procedure by first prescribing a family of time-dependent distributions and then constructing a reference stochastic differential equation (SDE) that realizes these distributions. This approach allows for the simultaneous achievement of simulation-free training and finite-time generation.
Results
The proposed framework successfully demonstrates that it can achieve both simulation-free training and finite-time generation, overcoming the limitations of conventional diffusion models. The numerical experiments validate the effectiveness of the new approach, showing improvements in generation efficiency and quality.
Implications
This work has significant implications for the development of generative models, particularly in applications requiring efficient training and rapid sample generation. The insights gained from this framework could lead to advancements in various domains, including image, audio, and text generation.
Amortized Interventional Forecasting for Multivariate CIR Processes
Time Series
- Introduces CIR-ACTIVA, a model for amortized distributional causal effect estimation.
- Develops a causal multivariate CIR data-generating process for benchmarking interventional forecasting.
- Demonstrates superior causal selectivity and calibration in short-horizon predictions compared to existing models.
- Enables what-if queries for coupled spread systems, enhancing stress testing capabilities.
Read more
Amortized Interventional Forecasting for Multivariate CIR Processes
Summary
This paper addresses the limitations of the Cox-Ingersoll-Ross (CIR) process in financial modeling, particularly its inability to capture causal influences between correlated time series. The authors propose a novel framework, CIR-ACTIVA, which combines an amortized model for distributional causal effect estimation with a causal multivariate CIR data-generating process. This framework allows for the prediction of multi-horizon shock responses without the need for retraining for each scenario. The model is evaluated using credit default swap (CDS) spreads, where it demonstrates superior performance in causal selectivity and horizon-resolved calibration compared to existing observational and causal inference baselines. The authors also introduce a causal simulator that generates paired observational and interventional samples, providing a benchmark for evaluating interventional forecasting. The results indicate that CIR-ACTIVA effectively shields non-affected series from phantom shock effects and excels in short-horizon predictions, which are crucial for stress testing in financial contexts.
Methodology
The authors developed CIR-ACTIVA, which treats financial time series as time-stamped observations and employs a horizon-bucketed decoder for causal effect estimation. They also created a causal multivariate CIR simulator that generates both observational and interventional samples, facilitating the training and evaluation of the model.
Results
CIR-ACTIVA outperformed baseline models in terms of causal selectivity and horizon-resolved calibration, particularly excelling at short horizons. The model effectively mitigated the impact of phantom shocks on unaffected series, providing more accurate predictions for stress testing scenarios.
Implications
The proposed framework has significant implications for financial forecasting, particularly in stress testing and risk management. It allows practitioners to better understand the causal dynamics of financial systems and make informed decisions based on potential interventions.
CRS-Triage: Confidence- and Reliability-Aware Selective Triage under Incomplete Clinical Evidence
Multimodal
- CRS-Triage employs a reliability-aware multimodal fusion mechanism to assess the reliability of structured data and clinical text for each patient encounter.
- The model introduces a confidence score to determine whether to accept predictions or defer cases for further assessment, reducing the risk of overconfident predictions.
- Larger penalties for under-triage errors encourage the model to prioritize accurate acuity assignments for high-acuity patients.
- CRS-Triage shows improved risk-coverage trade-off in triage predictions compared to existing methods.
Read more
CRS-Triage: Confidence- and Reliability-Aware Selective Triage under Incomplete Clinical Evidence
Summary
The paper presents CRS-Triage, a novel machine learning model designed to enhance emergency triage processes by addressing the challenges posed by incomplete and unreliable electronic health record (EHR) data. Traditional ML models often struggle with the inconsistencies and incompleteness of EHR data, leading to unreliable predictions of patient acuity levels. CRS-Triage introduces a confidence- and reliability-aware selective triage mechanism that evaluates the reliability of both structured data and clinical text separately, while also assessing their consistency. This dual evaluation allows the model to produce a confidence score for each prediction, which is then compared to a predefined threshold to decide whether to issue a prediction or defer the case. The model is designed to minimize the risk of under-triage, which can delay critical care for high-acuity patients, by incorporating larger penalties for such errors. Experiments conducted on the MIMIC-IV-ED dataset demonstrate that CRS-Triage not only achieves strong prediction performance but also maintains reliability even when faced with incomplete or inconsistent EHR data.
Methodology
CRS-Triage utilizes a reliability-aware multimodal fusion approach that evaluates the reliability of structured data and clinical text separately. It assesses the consistency between these modalities to inform the final prediction. A confidence score is calculated based on modality reliability and cross-modal consistency, guiding the decision to either issue a prediction or defer it for further assessment.
Results
The experiments on the MIMIC-IV-ED dataset indicate that CRS-Triage achieves strong prediction performance, effectively balancing the risk of under-triage and over-triage. The model demonstrates reliability even when EHR data is incomplete or inconsistent, outperforming traditional methods in terms of risk-coverage trade-off.
Implications
The CRS-Triage model has significant implications for emergency medicine, where timely and accurate triage decisions are critical. By improving the reliability of triage predictions in the face of incomplete data, this model can enhance patient safety and optimize resource allocation in emergency departments.
A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics
Time Series
Theory
Interpretability
- PI-HNO effectively predicts transient magnetization under complex conditions.
- The model integrates local and global branches to capture hysteresis and boundary conditions.
- Achieves high energy consistency in B-H trajectory predictions with minimal parameters.
- Ablation studies validate the importance of each model component.
Read more
A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics
Summary
This paper introduces the Physics-Informed Hybrid Neural Operator (PI-HNO), a novel neural model designed to predict transient magnetization in power magnetic components under complex operating conditions. Traditional core-loss models struggle to accurately represent transient behaviors due to fast transitions, dc bias, and temperature variations. The PI-HNO addresses these challenges by integrating a local recurrent branch for boundary-state representation and a global Preisach-inspired branch for waveform-level hysteresis context. The model is trained using a dataset from the MagNetX transient database, focusing on 14 different ferrite materials. Results indicate that PI-HNO achieves a balance between prediction accuracy and energy consistency in B-H trajectories, with mean and 95th percentile energy consistency errors of 1.92% and 7.60%, respectively, while utilizing only 4777 trainable parameters. Ablation studies confirm the distinct contributions of the model's components to its predictive performance, highlighting the effectiveness of incorporating physics-informed regularization in machine learning for transient magnetization prediction.
Methodology
The PI-HNO model employs a hybrid approach combining physics-informed principles with neural network architectures. It consists of a local recurrent branch for capturing boundary-state dynamics and a global branch inspired by Preisach theory for extracting hysteresis features. The model is trained on historical B-H data and operating conditions, with a focus on ensuring energy consistency through regularization techniques.
Results
The PI-HNO model demonstrated a mean B-H energy consistency error of 1.92% and a 95th percentile error of 7.60% when evaluated on the MagNetX transient database. The model's compact architecture, with only 4777 trainable parameters, allows for efficient training and deployment while maintaining high accuracy in transient magnetization predictions.
Implications
The development of the PI-HNO model has significant implications for the design and optimization of power magnetic components in high-frequency, high-power-density applications. Its ability to accurately predict transient behaviors can enhance the efficiency and reliability of power converters, leading to better performance in various industrial applications.
Approximate Speculative Decoding
NLP
Large Language Models
Efficient ML
- ASD replaces binary first-mismatch truncation with budgeted longest-prefix selection, allowing for more flexible acceptance of mismatches.
- The method does not require training or fine-tuning of the model, making it easy to implement.
- ASD improves throughput by 3.05% to 15.26% over matched strict verification and averages a 7.78% gain across multiple tasks.
- The approach enhances acceptance rates on complex tasks, demonstrating its effectiveness in practical applications.
Read more
Approximate Speculative Decoding
Summary
This paper introduces Approximate Speculative Decoding (ASD), a novel approach to enhance the efficiency of autoregressive generation in large language models (LLMs) by improving the speculative decoding process. Traditional speculative decoding methods utilize a lightweight proposal model to generate draft tokens, which are then verified by a target model. However, standard greedy verification halts at the first mismatch between the draft and target tokens, leading to potential loss of usable suffixes that could still be target-greedy. ASD addresses this limitation by implementing a budgeted longest-prefix selection strategy, allowing for the acceptance of certain mismatches while reusing contiguous target-greedy suffixes without requiring additional target model passes or fine-tuning. The method operates under a framework that includes a local target-logit regret gate, a block-level exception cap, and a persistent request-level regret budget. The authors demonstrate that ASD can improve throughput and acceptance rates significantly compared to traditional verification methods, achieving notable gains across various tasks without the need for training a new model.
Methodology
The authors propose a training-free verifier called Approximate Speculative Decoding (ASD) that modifies the verification process in speculative decoding. ASD utilizes a budgeted longest-prefix selection mechanism to allow for the acceptance of certain mismatches based on a local target-logit regret gate, while also maintaining a cumulative regret budget across requests. This enables the reuse of contiguous target-greedy suffixes without additional target model evaluations.
Results
The experimental results indicate that ASD improves fixed-workload throughput by 3.05% to 15.26% compared to matched strict verification. On average, ASD achieves a 7.78% gain across seven Qwen3-14B + DSpark-14B tasks. Additionally, on the DeepSeek-V4-Flash (284B) with DSpark, ASD increases verifier-side acceptance rates by approximately 10% to 16% on the GSM8K and MATH-500 tasks.
Implications
The introduction of ASD has significant implications for the efficiency of autoregressive generation in LLMs, potentially leading to faster inference times and improved performance on complex tasks. This method can be particularly beneficial in applications requiring real-time processing and high throughput, such as conversational AI and automated content generation.
Above-ground Biomass Estimation with Geospatial Foundation Models
Multimodal
- GFMs show promise for AGB estimation but require rich multi-modal training data.
- Frozen encoders of GFMs underperform compared to fully supervised models.
- Pre-computed embeddings from GFMs significantly enhance regression performance.
- The study establishes a benchmark for AGB estimation across diverse biomes.
Read more
Above-ground Biomass Estimation with Geospatial Foundation Models
Summary
This paper addresses the challenge of accurately estimating Above-Ground Biomass (AGB) from satellite imagery, which is crucial for monitoring carbon stocks globally. The authors explore the potential of Geospatial Foundation Models (GFMs) for this regression task, which has been underexplored in the context of quantitative applications. They present a comprehensive benchmark using the AGBD dataset, which encompasses various biomes and geographies. The study evaluates two approaches for utilizing GFMs: (i) as frozen encoders within the PAN-GAEA framework and (ii) as ready-to-use pre-computed embedding products, specifically AlphaEarth Foundations (AEF) and TESSERA. The authors compare 11 GFMs and the embedding products against a fully supervised state-of-the-art (SOTA) model, assessing their geographical and temporal generalization capabilities. The findings reveal that while frozen encoders underperform compared to the SOTA model, the pre-computed embedding products significantly enhance performance. An MLP trained on AEF embeddings outperforms the SOTA model trained on AGBD features, indicating that rich multi-modal data can improve biomass regression outcomes when distributed as accessible embedding layers.
Methodology
The authors benchmarked GFMs for AGB estimation using the AGBD dataset. They evaluated GFMs as frozen encoders and as pre-computed embedding products, comparing their performance against a fully supervised state-of-the-art model. The assessment included geographical and temporal generalization capabilities and agreement with independent reference data.
Results
The results indicated that frozen GFM features were less effective for biomass regression compared to the supervised SOTA model. However, MLPs trained on AEF embeddings outperformed the SOTA model trained on AGBD features, and the SOTA model trained on AEF embeddings achieved the best overall results, demonstrating better generalization across space and time.
Implications
The findings suggest that GFMs, particularly when utilized as pre-computed embeddings, can significantly improve the accuracy of AGB estimation from satellite imagery. This has implications for climate science, carbon accounting, and conservation efforts, enabling more reliable monitoring of carbon stocks globally.
NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning
Graph Learning
- NodeJEPA introduces a new architecture for node-level self-supervised learning that predicts latent representations instead of reconstructing inputs.
- The model employs structure-aware k-hop ego-subgraphs and a context encoder with an EMA target encoder to enhance representation learning.
- Variance-covariance and Laplacian spectral regularizers are utilized to stabilize the embedding geometry.
- NodeJEPA and its variant PatchJEPA achieve top performance on multiple node classification benchmarks, demonstrating the effectiveness of the proposed methods.
Read more
NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning
Summary
The paper introduces NodeJEPA, a novel joint-embedding predictive architecture designed for node-level graph self-supervised learning. Unlike traditional contrastive and generative methods that often rely on input reconstruction or specific augmentations, NodeJEPA focuses on predicting latent representations of masked nodes within structure-aware k-hop ego-subgraphs. This approach utilizes a context encoder that predicts these representations based on targets generated from an EMA-updated target encoder, employing a stop-gradient mechanism to prevent representation collapse. The predictor integrates various structural descriptors through cross-attention, and the model incorporates variance, covariance, and Laplacian spectral regularizers to stabilize the embedding geometry. Additionally, a curriculum learning strategy is employed to gradually increase the difficulty of masking during training. The paper also presents PatchJEPA, a variant that predicts latent representations from pre-computed graph partitions, allowing for a comparison of masking granularity effects. The evaluation of NodeJEPA on multiple node classification benchmarks demonstrates its effectiveness, achieving top ranks against several strong self-supervised baselines.
Methodology
NodeJEPA masks k-hop ego-subgraphs and predicts the latent representations of masked nodes using a context encoder. It employs an EMA target encoder to generate prediction targets and integrates structural descriptors through cross-attention. The model incorporates regularizers to prevent collapse and utilizes a curriculum learning approach to adjust masking difficulty over time. PatchJEPA serves as a comparative variant that predicts from pre-computed graph partitions.
Results
NodeJEPA and PatchJEPA achieved the best average self-supervised ranks across five node classification benchmarks, finishing first or second on four out of five datasets. The results were validated through paired significance tests and few-shot probes, confirming the effectiveness of the proposed methods.
Implications
The findings suggest that NodeJEPA can significantly enhance node representation learning in graphs, which is crucial for various applications in graph-based machine learning tasks, particularly in scenarios with limited labeled data. The approach may also inspire further research into self-supervised learning techniques that leverage structural information in graphs.
Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits
Theory
Optimization
Efficient ML
- First formulation of generalized low-rank matrix bandits with multi-objective feedback.
- Introduction of Lexi-LowGLM, an efficient online algorithm that reduces estimator-update complexity.
- Establishment of a regret bound that scales with effective low-rank dimensions.
- Numerical experiments confirm the algorithm's effectiveness and computational efficiency.
Read more
Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits
Summary
This paper addresses the problem of generalized low-rank matrix bandits with multiple prioritized objectives, proposing a novel algorithm called Lexi-LowGLM. In this setting, the learner selects matrix-valued arms and receives vector-valued rewards that correspond to multiple objectives, each with different priority levels. The authors introduce a lexicographic preference order for evaluating these arms, prioritizing higher-level objectives before lower-level ones. The proposed Lexi-LowGLM algorithm efficiently estimates objective-specific low-rank subspaces and performs online updates using a Newton-type proximal method, significantly reducing the computational complexity of estimator updates from O(TΒ²) to O(T). The authors establish a regret bound that scales with the effective low-rank dimension rather than the ambient dimension, demonstrating the efficiency of their approach. Through numerical experiments, the paper validates the effectiveness and computational efficiency of Lexi-LowGLM, marking a significant advancement in the field of online learning with matrix-valued actions and multiple objectives.
Methodology
The authors propose Lexi-LowGLM, which estimates low-rank subspaces specific to each objective and updates estimators online using a Newton-type proximal update. This approach avoids the computational burden of batch estimations, allowing for efficient learning in long-horizon settings.
Results
The proposed algorithm achieves a regret bound of eO(W lex_i βm (d1 + d2)r βT) for each objective, where W lex_i captures the lexicographic trade-off effect. The results indicate that the algorithm maintains competitive performance while significantly reducing computational complexity.
Implications
The findings have potential applications in areas requiring decision-making under uncertainty with multiple objectives, such as personalized recommendations, resource allocation, and multi-criteria optimization problems.
Sedentary Behavior Classification for Wearable Sensors with a CNN-BiLSTM Model
Time Series
- The CHAP model demonstrates strong performance on hip-worn accelerometer data but requires adaptation for wrist-worn data.
- Fine-tuning the hip-trained model on wrist data leads to improved classification accuracy compared to training from scratch.
- The study highlights the importance of posture in accurately measuring sedentary behavior using wearable sensors.
- Cross-device generalization is feasible, but wrist-specific adaptations are necessary due to signal variability.
Read more
Sedentary Behavior Classification for Wearable Sensors with a CNN-BiLSTM Model
Summary
This paper addresses the challenge of accurately classifying sedentary behavior using wearable sensors, particularly focusing on the transferability of a deep learning model from hip-worn accelerometers to wrist-worn accelerometers. The authors utilize the CHAP model, a CNN-BiLSTM architecture originally trained on hip data, to evaluate its performance on wrist data without retraining, as well as its adaptability through fine-tuning with varying amounts of labeled wrist data. The study employs the iWatch dataset, which includes ground-truth posture labels obtained from wearable cameras. Results indicate that while the hip-trained model performs well on hip data, its accuracy decreases significantly on wrist data due to differences in sensor placement. However, fine-tuning the CHAP model on wrist data consistently outperforms transformer models trained from scratch, suggesting that pretraining on hip data provides a beneficial starting point for wrist-based applications. The findings emphasize the necessity for wrist-specific adaptations to accommodate the higher variability in wrist sensor signals, ultimately contributing to improved sedentary behavior classification in real-world settings.
Methodology
The study employs a CNN-BiLSTM model (CHAP) initially trained on hip-worn accelerometer data and evaluates its zero-shot performance on wrist data. The model is then fine-tuned with varying amounts of labeled wrist data to assess performance improvements. Additionally, transformer-based models are trained from scratch for comparative analysis.
Results
The hip-trained CHAP model shows strong performance on hip data but experiences a drop in accuracy when applied to wrist data. Fine-tuning the model with wrist data consistently enhances performance, outperforming transformer models trained from scratch. The results indicate that hip-based pretraining is beneficial for wrist deployment, although wrist-specific adaptations are crucial due to higher signal variability.
Implications
The findings suggest that wearable technology for monitoring sedentary behavior can benefit from transfer learning approaches, allowing for more accurate and efficient classification of sedentary activities in diverse settings. This has potential applications in health monitoring and interventions aimed at reducing sedentary behavior.
Stochastic Emulation using Generalized Stratified Sampling for Performance-Based Risk Optimization of Structures
Optimization
- Integration of GSS with SPCE improves the estimation of extreme structural responses.
- The proposed framework reduces computational burden in nested reliability analyses.
- Application to a two-story steel building demonstrates practical effectiveness.
- Conditional exceedance probabilities are accurately estimated and recombined.
Read more
Stochastic Emulation using Generalized Stratified Sampling for Performance-Based Risk Optimization of Structures
Summary
This paper presents a novel framework that integrates Generalized Stratified Sampling (GSS) with Stochastic Polynomial Chaos Expansion (SPCE) to enhance Performance-Based Risk Optimization (PBRO) of structures subjected to stochastic loads. The authors highlight the limitations of SPCE in accurately representing extreme responses in structural response distributions, particularly in the tails. To overcome this, the GSS method partitions the input space into strata based on hazard intensity, allowing for improved representation of extreme responses. Within each stratum, independent SPCE emulators are trained, and the conditional exceedance probabilities are recombined using the total probability theorem to evaluate probabilistic constraints in the optimization problem. The framework is applied to optimize the cross-sectional areas of buckling-restrained braces in a two-story steel building, aiming to minimize construction costs while adhering to probabilistic performance constraints. The results demonstrate that the GSS-SPCE framework effectively estimates structural response distributions, including tail regions, while significantly reducing the number of nonlinear model evaluations required for PBRO.
Methodology
The methodology involves combining Generalized Stratified Sampling (GSS) with Stochastic Polynomial Chaos Expansion (SPCE). GSS partitions the input space into strata based on hazard intensity, and SPCE is used to train emulators within each stratum. The conditional exceedance probabilities are then recombined using the total probability theorem to evaluate the optimization problem's constraints.
Results
The GSS-SPCE framework successfully estimates structural response distributions, including their tail regions, while significantly reducing the number of required nonlinear model evaluations. The case study on a two-story steel building shows that the method can minimize construction costs while satisfying probabilistic performance constraints.
Implications
The proposed framework has significant implications for structural engineering, particularly in optimizing designs under uncertainty. It allows for more efficient risk assessments and can aid in making informed decisions regarding structural safety and economic viability.
Differentiating Through Dual Prices: End-to-End Policy Learning Under Capacity Constraints
Optimization
Reinforcement Learning
Theory
- Introduces an end-to-end approach for learning assignment policies under capacity constraints.
- Critiques the traditional decision-blind method for failing to consider capacity during model training.
- Proposes a bilevel optimization framework that jointly trains outcome models and dual prices.
- Demonstrates superior performance of the end-to-end methods in queueing simulations across multiple datasets.
Read more
Differentiating Through Dual Prices: End-to-End Policy Learning Under Capacity Constraints
Summary
This paper addresses the challenge of learning assignment policies for scarce resources in social services, where decisions must be made sequentially and capacity constraints are crucial. The authors critique the traditional decision-blind approach, which fits outcome models independently of capacity constraints and subsequently computes dual prices through auxiliary optimization. Instead, they propose an end-to-end method that integrates the learning of outcome models with the dual prices in a bilevel optimization framework. This approach allows for differentiation through the dual prices, enabling the model to learn from the deployed policy's value directly. The authors explore two formulations: a nonconvex one that matches the deployed policy and a convex relaxation that guarantees capacity constraints in expectation. The evaluation is conducted through a queueing simulation across six datasets, demonstrating that the end-to-end methods consistently outperform decision-blind baselines, particularly in scenarios where capacities are binding. The results indicate that the proposed method not only enhances policy value but also reduces queueing delays significantly, especially in large-scale settings like hospital resource allocation.
Methodology
The authors formulate a bilevel optimization problem where the inner problem computes dual prices based on current outcome models, while the outer problem maximizes an inverse-propensity-weighted estimate of the deployed policy's value. This allows for differentiation through the dual prices, linking model training directly to the policy's operational constraints.
Results
The end-to-end training methods consistently achieved higher policy values and reduced queueing delays across six datasets, including a significant improvement in a large hospital cohort dataset. In contrast, decision-blind baselines often violated capacity constraints and resulted in longer wait times.
Implications
The proposed method has significant implications for resource allocation in social services, such as healthcare and housing assistance, where capacity constraints are critical. It suggests that integrating operational constraints into machine learning models can lead to more effective and efficient decision-making processes.
Robust General Utility for Reinforcement Learning
Reinforcement Learning
Theory
Optimization
- Introduces robust general-utility RL to handle utility mis-specification in RL applications.
- Proposes a minimax framework that generalizes existing RL paradigms and provides a unified view.
- Develops two stochastic algorithms with convergence guarantees for both concave and nonconcave utility functions.
- Demonstrates the effectiveness of the proposed methods through experiments on practical RL tasks.
Read more
Robust General Utility for Reinforcement Learning
Summary
This paper introduces a novel framework for reinforcement learning (RL) termed robust general-utility RL, which addresses the challenges posed by utility mis-specification during policy deployment. Traditional RL with general utility assumes a fixed utility function, which can lead to performance gaps when the utility at deployment differs from that used during training. The proposed framework utilizes a minimax learning approach to train policies that are robust against such discrepancies by defining an uncertainty set for the utility functional. The authors provide a unified perspective on various existing RL paradigms, including reward-robust RL and constrained RL, through the lens of utility uncertainty. They develop two provably convergent stochastic algorithms: one for concave utilities using a projected stochastic gradient descent-ascent method, and another for nonconcave utilities employing a stochastic prox-extragradient algorithm. The experimental results demonstrate the effectiveness of the proposed methods in tasks related to large language model safety alignment and exploration maximization, confirming the theoretical convergence guarantees.
Methodology
The authors formulate the robust general-utility RL problem as a minimax objective, where they optimize policies against utility mis-specification within a defined uncertainty set. They develop a projected stochastic gradient descent-ascent method for concave utilities and a stochastic prox-extragradient algorithm for nonconcave utilities, ensuring convergence through rigorous analysis.
Results
The proposed algorithms exhibit provable convergence to first-order stationary points, with the concave utility algorithm converging at a rate of O(1/βK). The nonconcave utility algorithm effectively mitigates instability issues and demonstrates reliable performance in various RL tasks, corroborating the theoretical findings.
Implications
This work has significant implications for deploying RL systems in real-world scenarios where utility functions may differ from training conditions. It enhances the robustness of RL policies, making them more reliable in dynamic environments and applications such as safety alignment in large language models and exploration tasks.
Physics-informed reduced-order modelling with equivariant spectral submanifolds
Theory
Efficient ML
Optimization
- Introduction of equivariant spectral submanifold (eSSM) reduction for nonlinear dynamical systems.
- eSSM incorporates symmetries of the full-order model to enhance computational efficiency.
- Development of a new algorithm that improves model robustness and reduces computational costs.
- Demonstrated effectiveness through benchmark problems, outperforming traditional methods.
Read more
Physics-informed reduced-order modelling with equivariant spectral submanifolds
Summary
This paper introduces a novel approach to reduced-order modeling of nonlinear dynamical systems through the concept of equivariant spectral submanifolds (eSSMs). The author highlights the limitations of traditional spectral submanifold (SSM) reduction methods, particularly their computational expense in high-dimensional systems. The proposed eSSM framework incorporates the symmetries of the full-order model directly into the reduction process, enhancing both computational efficiency and model robustness. The mathematical foundations of eSSM are established, demonstrating that SSMs are naturally equivariant submanifolds and that the reduced dynamics inherit the appropriate group actions. A new algorithm for eSSM reduction is developed, which significantly accelerates computations while ensuring that the learned model remains aligned with the symmetric dynamics of the original system. The advantages of this method are validated through several benchmark problems, including tests from the Common Task Framework for Science, showcasing its effectiveness in capturing complex dynamics more accurately than existing linear techniques.
Methodology
The paper establishes the mathematical framework for equivariant spectral submanifolds, proving that SSMs are equivariant submanifolds. It develops an eSSM reduction algorithm that constrains both the manifold parametrization and the reduced dynamics to symmetry-adapted coefficient spaces, allowing for faster computations and improved model robustness.
Results
The eSSM reduction algorithm demonstrated significantly faster computation times and higher robustness in model predictions compared to traditional SSM methods. The approach was validated on benchmark problems, including a test from the Common Task Framework for Science, showing superior performance in capturing the dynamics of nonlinear systems.
Implications
The findings suggest that incorporating symmetries into reduced-order modeling can lead to more efficient and reliable simulations of complex dynamical systems. This has potential applications in various scientific fields where accurate modeling of nonlinear dynamics is crucial, such as fluid dynamics, structural analysis, and data-driven modeling.
Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection
Theory
Optimization
Robotics
- Critique of myopic experiment selection methods in scientific discovery.
- Formulation of goal-directed discovery as a stochastic shortest-path problem.
- Introduction of CG-Plan, a capability-aware planning approach.
- Demonstration of the limitations of myopic planners under capability gating.
Read more
Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection
Summary
This paper addresses the challenges faced by automated scientific discovery systems in selecting experiments, testing hypotheses, and acquiring necessary capabilities. The authors critique the prevalent myopic approach, which maximizes immediate expected information gain (EIG) without considering the long-term value of constructive actions that enable future experiments. They formalize the problem as a stochastic shortest-path (SSP) problem in belief space, where constructive experiments are represented as edges that alter the downstream action graph. The authors prove that myopic planners can be arbitrarily suboptimal when faced with capability gating, leading to situations where they may fail to reach the goal. They introduce CG-Plan, a capability-aware planner that incorporates a cost-to-go heuristic, allowing it to account for capabilities beyond the immediate lookahead horizon. Empirical tests demonstrate that CG-Plan outperforms traditional myopic planners, particularly in scenarios where capability acquisition is crucial for future actions.
Methodology
The authors formulate the problem as a stochastic shortest-path (SSP) problem in belief space, where nodes represent epistemic states and edges represent experiments or actions. They prove theoretical results regarding the limitations of myopic planners and introduce CG-Plan, which uses a cost-to-go heuristic that incorporates capability awareness. The methodology includes both theoretical proofs and empirical testing in a controlled environment.
Results
The paper establishes that myopic planners can have an unbounded approximation ratio relative to optimal planners when capability gating is present. CG-Plan demonstrates improved performance in a controlled testbed, showing that the performance gap persists across various lookahead horizons and is sensitive to the gating parameter.
Implications
The findings suggest that incorporating capability awareness into planning systems can significantly enhance the effectiveness of automated scientific discovery. This approach may be applicable in various domains where sequential decision-making and capability acquisition are critical, such as robotics, experimental design, and AI-driven research.
Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays
Multimodal
Robotics
Computer Vision
- Tactus achieves open-vocabulary object recognition from pressure data, matching or exceeding existing supervised models.
- The model employs a small-data recipe, emphasizing sensor calibration and masked-autoencoder pretraining.
- Error analysis reveals that model inaccuracies are structured by contact ambiguity rather than language similarity.
- Inference-time improvements, such as tuple voting, significantly enhance accuracy.
Read more
Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays
Summary
The paper introduces Tactus, an innovative open model designed for object recognition using low-cost resistive pressure arrays, which are commonly used tactile sensors. Unlike previous works that focused on optical sensors, Tactus enables recognition through pressure data alone, addressing a significant gap in tactile representation learning. The model achieves a top-1 accuracy of 0.771 Β± 0.062 on the STAG benchmark, surpassing the existing supervised closed-set CNN baseline of 0.76. The methodology employs a small-data approach, utilizing only 187 training recordings and leveraging masked-autoencoder pretraining on a large dataset of 144k unlabeled frames. The study emphasizes the importance of sensor calibration and presents a detailed error analysis, revealing that the model's errors are primarily due to contact ambiguities rather than language-related issues. The findings also highlight the effectiveness of tuple voting for inference-time improvements and document several negative results, including the detrimental effects of cross-sensor pooling and input normalization defects. The authors provide open access to the model's weights, code, and memory layer, promoting further research in this area.
Methodology
The methodology involves mapping pressure frames to a multimodal embedding space for recognition via cosine ranking against text queries. The model is trained with a small dataset of 187 recordings, utilizing masked-autoencoder pretraining and sensor calibration to enhance performance.
Results
Tactus achieves a top-1 accuracy of 0.771 Β± 0.062 and a top-3 accuracy of 0.935 on the STAG benchmark, outperforming the closed-set CNN baseline. The model's errors are primarily concentrated in contact-ambiguous classes, and two diverse frames can recover 89% of the accuracy achieved with eight frames.
Implications
The development of Tactus has significant implications for robotics and tactile sensing, enabling more robust object recognition in environments where visual data may be limited or occluded. It opens avenues for further research into low-cost tactile sensors and their integration into multimodal systems.
Bayesian Data Reweighting Improves Multimodal Retrieval for Knowledge-Based Visual Question Answering
Multimodal
- Identifies the false-negative issue in KB-VQA retrieval, where semantically relevant documents are misclassified as negatives.
- Introduces Bayesian Data Reweighting (BDR) to adaptively infer weights for query-document pairs, improving training robustness.
- Demonstrates significant improvements in retrieval accuracy across multiple benchmarks and LLM backbones.
- Achieves better-separated query-document embeddings and improved Recall@K metrics.
Read more
Bayesian Data Reweighting Improves Multimodal Retrieval for Knowledge-Based Visual Question Answering
Summary
This paper addresses the limitations of existing contrastive training methods in knowledge-based visual question answering (KB-VQA), where multimodal retrievers are used to retrieve external evidence for image-question pairs. The authors identify that current methods treat all unmatched query-document pairs as equally informative negatives, which can lead to performance degradation due to the presence of semantically relevant documents that are incorrectly classified as negatives. To overcome this issue, the authors propose a novel framework called Bayesian Data Reweighting (BDR), which models the importance of query-document pairs as latent variables and adaptively infers posterior weights to downweight likely false negatives. BDR formulates contrastive learning as a Bayesian reweighting problem, allowing for a more nuanced understanding of the relevance of unmatched documents. The method employs closed-form posterior updates and stochastic EM optimization to enhance retrieval accuracy. Extensive experiments demonstrate that BDR significantly improves multimodal retrieval performance across seven KB-VQA benchmarks and three different large language model (LLM) backbones, outperforming standard InfoNCE and existing reweighting methods. The results indicate that BDR not only enhances retrieval accuracy but also improves downstream answer generation tasks, validating its effectiveness in mitigating noisy supervision and enhancing the robustness of KB-VQA retrieval.
Methodology
The authors propose a Bayesian Data Reweighting (BDR) framework that models the importance of query-document pairs as latent variables. BDR formulates the contrastive learning process as a Bayesian reweighting problem, utilizing closed-form posterior updates under conjugate priors and optimizing the retriever with a stochastic approximation expectation-maximization algorithm. This allows the model to adaptively downweight likely false negatives without explicitly labeling them.
Results
BDR consistently outperforms standard InfoNCE and existing reweighting baselines, achieving stronger retrieval accuracy across seven knowledge-based VQA benchmarks and three different LLM backbones. It also enhances downstream answer generation performance on datasets like InfoSeek and EVQA. Detailed analyses reveal that BDR effectively assigns lower weights to false negatives and improves Recall@K across various retrieval budgets.
Implications
The proposed BDR framework has significant implications for improving the robustness of multimodal retrieval systems in KB-VQA, which can enhance applications in education, healthcare, and open-domain dialog systems by providing more accurate and contextually relevant answers.
Joint Affine Spectral Shaping: Coupling Weight and Bias Updates Beyond Weight-Only Muon
Optimization
Large Language Models
NLP
- Introduces Joint Regularized Inverse (JRI) for simultaneous weight and bias updates.
- Demonstrates that joint spectral allocation leads to improved model performance.
- Conducts a thorough ablation study to compare different optimization strategies.
- Finds that allowing bias to influence the joint SVD significantly enhances accuracy.
Read more
Joint Affine Spectral Shaping: Coupling Weight and Bias Updates Beyond Weight-Only Muon
Summary
This paper investigates the optimization of neural network parameters by exploring the coupling of weight and bias updates in affine layers. Traditional matrix spectral optimizers typically treat weight updates separately from bias updates, which may not be optimal for performance. The authors propose a novel approach called Joint Regularized Inverse (JRI), which combines weight and bias updates into a single spectral update mechanism. By formulating the affine layer as a joint momentum matrix, the authors apply a capped regularized-inverse spectral map to update both weights and biases simultaneously. Through a series of controlled experiments on a BERT-mini model trained on IMDb, the study compares the proposed JRI method against various baseline methods, including weight-only inverse shaping and bias as a spectral probe. The results indicate that JRI consistently improves test accuracy and reduces test loss compared to the baseline methods, demonstrating the effectiveness of joint spectral allocation in optimizing neural network performance.
Methodology
The authors conducted a strict five-seed ablation study on a four-layer BERT-mini model, comparing four different optimization strategies: exact-SVD Muon, weight-only inverse shaping, affine-probe inverse shaping, and the proposed JRI method. They analyzed the impact of these methods on validation accuracy, test loss, and the relationship between weight and bias updates.
Results
The JRI method achieved a test accuracy of 85.738 Β± 0.180% and a test loss of 0.3291, outperforming the weight-only inverse shaping method, which had a test accuracy of 85.562 Β± 0.308% and a test loss of 0.3345. The independent replication study confirmed the results with a test accuracy of 85.743 Β± 0.203%. The findings indicate that joint updates lead to a more effective allocation of the update budget between weights and biases.
Implications
The study suggests that incorporating joint updates for weights and biases can enhance the performance of neural networks, particularly in language modeling tasks. This approach may lead to more efficient training and improved model accuracy, potentially influencing future optimization strategies in deep learning.
GLOBE: Trajectory-Aligned Gradient Matching with Structured Sparse Optimization for Coreset Selection
Efficient ML
Optimization
- GLOBE utilizes multi-checkpoint gradient trajectories to capture the temporal evolution of training dynamics.
- The framework introduces a multi-order matching objective for aligning gradient statistics, enhancing coreset selection.
- Structured sparse optimization techniques are employed to manage correlated samples and induce sparsity.
- GLOBE maintains class balance in the selected coreset, ensuring diverse category representation.
Read more
GLOBE: Trajectory-Aligned Gradient Matching with Structured Sparse Optimization for Coreset Selection
Summary
The paper presents GLOBE (Gradient LocalβBalanced Extraction), a novel framework for coreset selection that addresses the limitations of existing gradient-based methods. Traditional methods often rely on gradients from a single model snapshot and use greedy selection techniques, which can fail to capture the evolving dynamics of optimization and struggle with correlated samples. GLOBE improves upon this by utilizing gradient trajectories that are constructed across multiple training checkpoints, allowing for a more comprehensive representation of each sample's influence throughout the training process. The authors introduce a multi-order matching objective that aligns both the first-order mean and the projected second-order moments of gradient trajectories, ensuring that the selected coreset maintains the training behavior of the full dataset. Additionally, GLOBE employs structured sparse optimization techniques, including Group LASSO and Elastic Net regularization, to induce sparsity at both group and sample levels while stabilizing the weights of correlated trajectories. The framework also incorporates a class-balanced Top-K selection strategy to ensure adequate category representation within the coreset. Experimental results across six benchmarks and five architectures demonstrate that GLOBE consistently outperforms existing coreset selection methods, particularly at low retention ratios, highlighting its effectiveness in enabling data-efficient learning.
Methodology
GLOBE formulates coreset selection as a globally optimized sparse weighting problem, leveraging gradient trajectories from multiple training checkpoints. It employs a multi-order matching objective to align first-order and second-order statistics of gradients, and incorporates structured sparsity through Group LASSO and Elastic Net regularization. A lightweight teacher-to-proxy alignment mechanism is also introduced to ensure fidelity in gradient behavior.
Results
The experiments conducted across six benchmarks and five evaluation architectures indicate that GLOBE consistently outperforms traditional coreset selection methods, particularly in scenarios with low retention ratios, leading to improved downstream test accuracy.
Implications
The proposed GLOBE framework has significant implications for on-device learning, particularly in resource-constrained environments such as edge devices. By effectively reducing the dataset size while preserving training utility, GLOBE can facilitate the deployment of deep learning models in applications like IoT and autonomous driving.
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation
Reinforcement Learning
Large Language Models
- Identifies a confounding effect in privileged replay scoring that affects token updates.
- Introduces Observation-Calibrated Self-Distillation (OCSD) to derive an observation residual for better token-level optimization.
- Demonstrates that OCSD outperforms strong baselines across multiple benchmarks and model scales.
- Shows that the calibrated residual aligns more closely with local environment feedback than traditional methods.
Read more
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation
Summary
This paper addresses the limitations of reinforcement learning in training large language model (LLM) agents, particularly the challenge of sparse trajectory-level rewards that provide insufficient guidance for token updates. The authors propose a novel approach called Observation-Calibrated Self-Distillation (OCSD) to enhance token-level supervision by contrasting two replay views: Full and Observation-Ablated. This method effectively isolates the influence of future observations from the replay scaffold, allowing for more accurate token updates during high-uncertainty steps. The experiments conducted on various benchmarks, including ALFWorld, WebShop, and Search-QA, demonstrate that OCSD consistently outperforms existing methods, providing better alignment with local environment feedback. The findings suggest that OCSD can significantly improve the training of LLM agents in complex interactive tasks by offering a more refined supervision mechanism.
Methodology
The authors developed OCSD by creating two structurally matched replay views (Full and Observation-Ablated) to derive an observation residual. This residual is used to modulate token-level updates in the Group Relative Policy Optimization (GRPO) framework, particularly during high-uncertainty steps, while maintaining the overall trajectory-level update direction.
Results
OCSD achieved superior performance across all tested tasks and model scales compared to existing methods. The diagnostic analyses confirmed that the observation residual derived from OCSD aligns more effectively with local environment feedback, enhancing the learning process of LLM agents.
Implications
The proposed OCSD method has the potential to improve the training efficiency and effectiveness of LLM agents in long-horizon interactive tasks, such as web navigation and embodied control, by providing more precise token-level supervision.
Schedule-Informed Temporal Fusion Forecasting of Hourly Airport Security-Checkpoint Throughput
Time Series
- Introduces a schedule-informed framework for forecasting airport security-checkpoint throughput.
- Transforms flight schedules into temporally aligned screening-load signals for improved accuracy.
- Achieves a weighted mean absolute percentage error of 9.33% for direct six-hour forecasts.
- Demonstrates lower errors in high-throughput percentiles compared to traditional forecasting models.
Read more
Schedule-Informed Temporal Fusion Forecasting of Hourly Airport Security-Checkpoint Throughput
Summary
This paper addresses the challenge of accurately forecasting hourly airport security-checkpoint throughput by developing a novel framework that integrates flight schedule information with historical throughput data. The authors recognize that checkpoint staffing decisions are influenced not only by the volume of passengers but also by the timing of their arrivals. Traditional forecasting methods often treat scheduled departures as immediate demand, failing to account for the lead time between passenger arrivals and flight departures. To overcome this, the authors propose a schedule-informed approach that transforms flight schedules into temporally aligned signals for predicting screening loads. Utilizing data from the Transportation Security Administration (TSA) and Cirium Diio flight schedules for HartsfieldβJackson Atlanta International Airport, the authors trained a Temporal Fusion Transformer model that combines historical throughput, scheduled activity, and seasonal variables. The model was validated and tested using a comprehensive dataset, demonstrating improved forecasting accuracy compared to recurrent neural networks and long short-term memory models. The findings indicate that the proposed framework significantly reduces forecasting errors, particularly for high-throughput percentiles, and provides a robust tool for airport operations management.
Methodology
The authors developed a framework that uses truncated Poisson lead-time kernels to convert flight schedules into pre-departure arrival-intensity features. These features were integrated into a Temporal Fusion Transformer model alongside historical TSA throughput data, scheduled activity, and temporal covariates. The model was trained and validated using data from 2023-2024, with a holdout period for testing.
Results
The proposed model achieved a weighted mean absolute percentage error of 9.33% for direct six-hour forecasts, outperforming recurrent neural networks (12.16%) and long short-term memory models (11.37%). The model maintained a consistent error rate between 10.60% and 11.04% for six-hour recursive updates across 24-96 hour forecasting horizons.
Implications
The forecasts generated by this model can significantly enhance airport operations by informing staffing decisions, lane openings, and multiday checkpoint planning. By aligning forecasts with the temporal dynamics of passenger arrivals, airports can better manage resources and improve passenger experience.
Exploiting Separability in Multi-Scale Grey-Box Bayesian Optimization
Optimization
- Introduces a bilevel reformulation for separable grey-box optimization problems.
- Reduces surrogate dimensionality by focusing on black-box variables only.
- Ensures exact satisfaction of white-box constraints without penalty functions.
- Demonstrates significant performance improvements across a suite of benchmark problems.
Read more
Exploiting Separability in Multi-Scale Grey-Box Bayesian Optimization
Summary
This paper addresses grey-box optimization problems where decision variables are divided into black-box variables, which are inputs to expensive functions, and white-box variables, which are governed by explicit equations. The authors propose a bilevel reformulation that optimizes the scalar objective using Bayesian optimization (BO) for black-box variables while solving the white-box subproblem through global optimization. This approach reduces the dimensionality of the surrogate model and ensures that white-box constraints are satisfied exactly when the inner optimizer converges to a feasible point. The method is tested on a benchmark suite of 13 problems, demonstrating significant improvements in optimization performance, including lower regret and faster execution times compared to traditional black-box BO methods. The findings suggest that this bilevel strategy effectively exploits the separability of grey-box problems, making it a valuable tool for multi-scale engineering applications.
Methodology
The authors employ a bilevel optimization framework where the outer loop uses Bayesian optimization to search over black-box variables, while the inner loop solves the white-box subproblem using global optimization techniques. This dual approach allows for dimensionality reduction and exact constraint satisfaction.
Results
The proposed bilevel optimization strategy achieved 11Γ to 108Γ lower regret than traditional black-box BO methods across 13 benchmark problems, with comparable or faster wall clock times. The results indicate robustness to different initialization sizes, exploration parameters, and choices of inner solvers.
Implications
The findings of this research have significant implications for multi-scale engineering design, where the integration of expensive black-box simulations with known white-box models can lead to more efficient optimization processes. This methodology can be applied in various fields such as chemical engineering, materials science, and any domain involving complex systems with both expensive and tractable components.
CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications
Time Series
Optimization
Theory
- CheMLFlow provides a modular and scalable platform for cheminformatics and materials informatics workflows.
- The platform reduces orchestration overhead and supports systematic benchmarking across various methods and datasets.
- CheMLFlow facilitates agent-assisted experimentation, allowing for more efficient and user-friendly scientific workflows.
- The system architecture includes core workflows and benchmarks that achieve literature performance in property predictions.
Read more
CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications
Summary
CheMLFlow is introduced as an open-source platform designed to streamline the development and execution of high-throughput workflows in cheminformatics and materials informatics. The platform addresses a significant bottleneck in scientific machine learning, where researchers often face challenges in assembling various stages of data handling, modeling, and reporting into a cohesive and reproducible pipeline. CheMLFlow offers modular components, ready-to-use reference pipelines, and standardized outputs that minimize orchestration overhead and facilitate benchmarking across diverse methods and datasets. It is built to be extensible and automation-friendly, featuring pluggable representations, deterministic splits, and structured outputs. The platform supports agent-assisted experimentation, allowing coding agents to assist users in constructing experiments and interpreting results. The paper details the system architecture, core workflows, and benchmarks that demonstrate CheMLFlow's performance in predicting quantum mechanical, physicochemical, and bioactivity properties, as well as applications in time series datasets beyond molecular chemistry. CheMLFlow aims to empower both non-expert and expert users by providing a comprehensive solution for molecular property prediction tasks while ensuring reproducibility and efficiency in scientific workflows.
Methodology
CheMLFlow employs a modular architecture that allows users to construct end-to-end workflows for data acquisition, curation, representation, model training, validation, and reporting. The platform includes ready-to-use workflows and supports high-throughput testing of configurations, treating all components as a single reproducible benchmark unit. It also features a design of experiments layer for optimizing model combinations and an agentic layer for facilitating user interactions.
Results
The benchmarks demonstrated that CheMLFlow achieves literature performance in predicting quantum mechanical, physicochemical, and bioactivity properties. Additionally, the platform's modular design allows for successful applications in time series forecasting, showcasing its versatility beyond traditional cheminformatics tasks.
Implications
CheMLFlow has the potential to significantly enhance the efficiency and reproducibility of scientific workflows in cheminformatics and materials informatics. By providing a comprehensive platform that integrates various stages of research, it can facilitate faster iterations and fair comparisons among different modeling approaches, ultimately advancing research in drug discovery and materials design.
FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs
Reinforcement Learning
Federated Learning
Efficient ML
- FedCritic-MIMO enables decentralized coordination among independently deployed cell-level controllers in 6G RANs.
- The framework utilizes peer-to-peer critic parameter exchange, avoiding the need for centralized training.
- Significant performance improvements in network throughput and user satisfaction metrics were observed.
- Communication overhead is reduced by about 76% compared to traditional distributed critic exchange methods.
Read more
FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs
Summary
This paper introduces FedCritic-MIMO, a novel framework designed for communication-efficient serverless federated multi-agent reinforcement learning (MARL) aimed at resource control in open and disaggregated 6G radio access networks (RANs). The framework addresses the challenge of coordinating independently deployed cell-level controllers that manage user scheduling, power allocation, beamforming, and interference under limited communication. Unlike traditional approaches that rely on centralized training or shared actors, FedCritic-MIMO allows each base station (BS) to execute its actor locally while exchanging only compatible shared critic parameters through peer-to-peer coordination. The methodology incorporates wireless-aware event triggering, adaptive layer-wise top-k sparse critic exchange, and balanced interference-aware fusion to facilitate effective collaboration among controllers. The paper establishes guarantees for finite-time stationarity and consensus in the proposed critic recursion model. Experimental results demonstrate that FedCritic-MIMO significantly outperforms various baselines, achieving the highest network throughput, improved user-rate distribution, and reduced communication overhead by approximately 76%. This framework showcases the potential for decentralized coordination in ultra-dense RANs without the need for centralized trajectory collection or parameter aggregation.
Methodology
The methodology involves a serverless federated learning framework where each base station executes its actor locally and retains personalized critic components. Collaboration is achieved through a compatible shared critic subnetwork, with critic exchanges triggered by local conditions such as queue urgency and interference intensity. Adaptive layer-wise top-k sparse critic exchange reduces communication load, while interference-aware fusion enhances coordination among neighboring controllers.
Results
FedCritic-MIMO demonstrated superior performance in simulations, achieving the highest held-out network throughput, improved user-rate distribution, and increased QoS satisfaction. It also attained the lowest interference cost per delivered bit and significantly reduced critic-communication overhead by approximately 76% compared to uncompressed distributed critic exchange.
Implications
The findings suggest that FedCritic-MIMO can effectively coordinate resource management in ultra-dense 6G RANs, paving the way for more efficient and scalable network architectures. This approach could be applied to various scenarios requiring decentralized learning and coordination in communication networks.