AI-generated summaries
Today's ML research,
without the noise.
Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.
24
Papers today
8h
Update frequency
7
Days of history
Data-Efficient Agentic Graph Domain Adaptation via Reliability-Aware Prototype Learning
Graph Learning
Efficient ML
- Identifies critical challenges in data-efficient agentic GDA: unreliable source anchoring, uncertain target association, and fragile target-marginal calibration.
- Proposes DEAG, a framework that stabilizes source guidance and performs soft target alignment.
- Implements a source-prior regularization technique to enhance target prediction reliability.
- Demonstrates significant performance improvements over competitive GDA baselines under the same source-data budget.
Read more
Data-Efficient Agentic Graph Domain Adaptation via Reliability-Aware Prototype Learning
Summary
This paper addresses the challenges of Graph Domain Adaptation (GDA) in data-efficient agentic learning systems, where limited labeled source graphs are available for adapting to new, unlabeled target graphs. The authors introduce DEAG, a reliability-aware prototype learning framework that stabilizes source guidance and enhances adaptation performance under constraints. DEAG estimates class reliability from retained source support and embedding compactness, creating stable source anchors by blending empirical prototypes with classifier directions. It performs prototype-aware soft target association, aligning confidence-weighted target centers with source semantics while introducing a source-prior regularizer to maintain target marginal consistency. The proposed method effectively tackles issues of unreliable source anchoring, uncertain target association, and fragile target-marginal calibration, demonstrating improved adaptation performance across various graph benchmarks with diverse domain shifts compared to existing GDA methods.
Methodology
DEAG employs a reliability-aware prototype learning approach that combines empirical prototypes with classifier directions to create stable source anchors. It utilizes soft target association techniques to align target predictions with source semantics and introduces a source-prior regularizer to maintain consistency in target marginal distributions.
Results
Experiments reveal that DEAG outperforms existing GDA methods, achieving higher average adaptation performance across various graph benchmarks while operating under the same constraints of limited source data.
Implications
The findings suggest that DEAG can be effectively applied in scenarios where labeled data is scarce, such as molecular property prediction and social network analysis, enhancing the adaptability of graph neural networks in real-world applications.
Stabilizing Performative Feedback Loops with Minimal Model Deployments
Theory
Optimization
Efficient ML
- Introduces a new algorithmic procedure for finding performatively stable models with minimal deployments.
- Achieves stability using exponentially fewer model deployments than prior approaches.
- Establishes a structural result for derandomizing the randomized predictor into a deterministic one under specific conditions.
- Connects performative stability to expected variational inequalities, providing a theoretical foundation for the proposed methods.
Read more
Stabilizing Performative Feedback Loops with Minimal Model Deployments
Summary
This paper addresses the challenge of stabilizing performative feedback loops in algorithmic predictions, where the predictions influence the data distribution they operate on. The authors introduce a new algorithmic procedure that efficiently finds a performatively stable model with minimal model deployments, significantly reducing the number of deployments required compared to previous methods. The paper establishes a connection between performative stability and expected variational inequalities, leading to a randomized approach that achieves stability without assumptions about how predictions shape distributions. Additionally, the authors present a structural result that allows for derandomization into a single predictor under certain conditions, such as well-conditioned loss and weak performative effects. This work contributes to the understanding of how to manage the complexities of social prediction in dynamic environments, where predictions and outcomes are interdependent.
Methodology
The authors develop an algorithmic procedure that leverages randomized predictors to achieve performative stability. They analyze the deployment efficiency of this approach and establish theoretical guarantees for its performance. The methodology includes a structural analysis of performative stability and its relationship to expected variational inequalities.
Results
The proposed algorithm successfully finds a performatively stable model using significantly fewer deployments than previous methods, demonstrating its efficiency in the high-accuracy regime. The structural results allow for the conversion of randomized predictors into deterministic ones under certain conditions, expanding the applicability of the findings.
Implications
This research has significant implications for fields where algorithmic predictions influence human behavior, such as finance, healthcare, and social media. By providing a method to stabilize feedback loops, it can help mitigate the risks associated with dynamic prediction environments and improve the reliability of algorithmic decision-making.
Specification Oracles
NLP
Large Language Models
Theory
- Specification oracles can effectively balance detail and compactness in knowledge representation.
- Weight-only oracles show significantly higher accuracy in structured environments compared to note-sheet oracles.
- The storage requirements for weight-only oracles are substantially higher than for note-sheet oracles.
- The study demonstrates the potential of neural networks to serve as dynamic specifications for various objects.
Read more
Specification Oracles
Summary
This paper explores the concept of specification oracles, which are designed to balance the tradeoff between detailed specifications and compactness. The authors investigate whether a language model can act as a dynamic specification oracle by learning facts about a target and directly answering related questions. They compare two methods for storing learned facts: external text notes and modifications to the model's weights. Through experiments involving four families of 596-fact worlds and two sizes of the Qwen2.5 model, they find that weight-only oracles significantly outperform note-sheet oracles in terms of accuracy when dealing with structured worlds, achieving an accuracy increase of 18.5 percentage points with the 7B model. However, this advantage comes at a cost, as the smallest adapter requires about 175 KiB of storage compared to a maximum of 16 KiB for note sheets. The study emphasizes the importance of structural information in improving the efficiency of specifications and highlights the potential of neural networks as versatile knowledge acquirers.
Methodology
The authors conducted experiments using two methods for learning oracles: a note-sheet oracle that allows the model to access a fixed-size text file, and a weight-only oracle that modifies the model's weights using LoRA parameters. They generated synthetic worlds with varying structures to evaluate the oracles' performance in terms of accuracy and storage efficiency.
Results
The weight-only oracles achieved an accuracy increase of 18.5 percentage points in structured worlds compared to 1.1 points for note-sheet oracles when using the 7B model. The smallest weight-only adapter required approximately 175 KiB of storage, while the note-sheet approach had a maximum budget of 16 KiB.
Implications
The findings suggest that neural networks can be utilized as efficient specification oracles, particularly in applications requiring dynamic knowledge representation. This could have implications for various fields, including software engineering, AI-driven documentation, and interactive systems where real-time knowledge retrieval is essential.
Multimodal deep learning from spectra for small-molecule structure identification: enhancing robustness with mixed-condition training
Multimodal
- Introduces mixed-condition training to improve robustness in multimodal deep learning for molecular structure identification.
- Utilizes a mixture-of-experts (MoE) architecture to enhance model performance under varying spectral conditions.
- Demonstrates significant improvements in mean reciprocal rank (MRR) and recall at rank 1 (R@1) through controlled experiments.
- Highlights the importance of domain knowledge in designing training conditions for better predictive performance.
Read more
Multimodal deep learning from spectra for small-molecule structure identification: enhancing robustness with mixed-condition training
Summary
This paper presents a novel approach to small-molecule structure identification using multimodal deep learning techniques that integrate various spectroscopic data. The authors introduce a mixed-condition training strategy that leverages domain knowledge from spectroscopy and chemistry to enhance the robustness of models against missing, degraded, or mismatched spectral inputs. The methodology involves a controlled evaluation protocol that simulates different conditions of spectral availability and quality, utilizing a mixture-of-experts (MoE) fusion architecture. The study evaluates 79,462 test samples across 30 conditions using simulated spectra from the Multimodal Spectroscopic Dataset (MSSD). Results demonstrate that mixed-condition training significantly improves model performance, with notable increases in mean reciprocal rank (MRR) and recall at rank 1 (R@1) across various conditions. The findings suggest that integrating domain knowledge into training strategies can lead to more robust and effective multimodal models for molecular structure elucidation.
Methodology
The authors developed a mixed-condition training strategy that incorporates domain-specific perturbations and spectrum replacements to simulate variations in spectral data. They conducted a factorial comparison of complete-input training versus mixed-condition training and vanilla concatenation versus MoE fusion to assess the impact on model performance. The candidate structure reranking task involved ranking a set of molecular structures based on available spectral data, using encoders for different spectral modalities and a scoring mechanism to determine candidate order.
Results
The mixed-condition training approach led to a mean reciprocal rank (MRR) increase from 0.9203 to 0.9763 (6.08% improvement) and a recall at rank 1 (R@1) increase from 89.50% to 96.36% (7.67% improvement). For specific modalities, MRR for IR-only and MS/MS-only increased significantly, reaching 2.15 and 2.31 times their baseline values, respectively. The results indicate that mixed-condition training is effective in enhancing model robustness while maintaining high performance with complete inputs.
Implications
The findings suggest that incorporating domain knowledge into training strategies can significantly enhance the robustness of multimodal models, making them more effective for real-world applications in molecular structure identification. This approach could be applied to other fields requiring multimodal data integration, such as drug discovery and materials science.
Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies
Theory
Large Language Models
Generative Models
- Introduces the concept of anomalies linked to main data points in random datasets.
- Establishes bounds for the minimum size of strongly dissimilar decompositions.
- Demonstrates a phase transition in decomposition size based on the number of anomalies.
- Provides criticality results for the similarity of undersampled datasets.
Read more
Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies
Summary
This paper investigates the impact of anomalies, specifically AI-generated data, on the decomposition and undersampling of random datasets, which is crucial for training large language models (LLMs). The author introduces the concept of anomalies as linked to main data points and explores how these affect the performance of LLMs. Using redundancy graphs and iterative techniques, the study establishes bounds for the minimum size of a strongly dissimilar (SD) decomposition. A notable phase transition phenomenon is demonstrated, where the minimum size of the decomposition is primarily influenced by main data points when the number of anomalies is low, but shifts to being dominated by anomalies beyond a certain threshold. The paper also presents a criticality result for the strong similarity of randomly undersampled datasets, illustrated through examples involving categorical datasets. The findings emphasize the importance of managing anomalies in training datasets to maintain model performance.
Methodology
The paper employs redundancy graphs and iterative stable set extraction techniques to analyze the decomposition of datasets. It uses first and second moment methods to estimate critical sizes for ensuring strong dissimilarity in undersampled datasets. Theoretical results are supported by examples involving categorical datasets.
Results
The study finds that the minimum size of a strongly dissimilar decomposition is determined by main data points when anomalies are few, but becomes influenced by anomalies when their number exceeds a certain threshold. Additionally, bounds for the critical size of a randomly undersampled dataset that maintains strong dissimilarity are established.
Implications
The findings have significant implications for the training of future LLMs, particularly in understanding how to effectively manage and incorporate AI-generated data to minimize its adverse effects on model performance. This research could inform strategies for dataset preparation and anomaly handling in machine learning.
A Machine Learning API for Earth Observation Data Cubes Based on openEO
Time Series
- Proposes a standardized ML specification for integrating ML into EO workflows using openEO.
- Structures ML workflows into three stages: model initialization, model actions, and model management.
- Demonstrates cross-backend interoperability through prototype implementations in R and Python.
- Identifies the need for deeper harmonization of serialization formats for full cross-backend portability.
Read more
A Machine Learning API for Earth Observation Data Cubes Based on openEO
Summary
This paper addresses the integration of machine learning (ML) methods into Earth Observation (EO) workflows, which are increasingly organized as spatio-temporal data cubes. The authors identify a representational mismatch between EO data cubes and the input requirements of ML models, necessitating complex, platform-specific transformations that hinder reproducibility and portability across cloud infrastructures. To bridge this gap, the authors propose a standardized process-level ML specification for the openEO framework, which structures ML workflows into three stages: model initialization, model actions (including training, tuning, inference, and validation), and model management. The specification supports various classical and deep learning algorithms, including Random Forests, Support Vector Machines, TempCNN, and Temporal Attention Encoders. The authors demonstrate the feasibility of their approach through three prototype implementations in R and Python, showcasing a crop type mapping use case that highlights cross-backend interoperability. The results indicate that while the proposed specification enhances reproducibility and accessibility of ML workflows, achieving full cross-backend portability requires further harmonization of serialization formats and execution semantics. The authors conclude that explicit backend conformance profiles are essential for improving the integration of ML into EO data cube workflows.
Methodology
The authors developed a process-level ML specification for openEO that organizes ML workflows into structured stages. They implemented prototypes in R and Python to validate the specification's feasibility and interoperability across different backends. Use cases were employed to demonstrate the application of classical and deep learning algorithms on EO data cubes.
Results
The implementation of the proposed ML specification showed that identical process graphs submitted to different openEO backends yielded comparable predictions and evaluation metrics. However, the study also revealed that achieving full cross-backend portability requires addressing serialization format and execution semantics inconsistencies.
Implications
The proposed ML specification can significantly enhance the integration of machine learning into Earth Observation data workflows, promoting reproducibility and interoperability across various cloud platforms. This advancement could facilitate broader applications of EO data in areas such as land-cover monitoring, crop yield estimation, and disaster response.
Scalable partial information decomposition for symptom networks via supervised embeddings
Graph Learning
Theory
Interpretability
- Introduces ePID, a scalable method for partial information decomposition in symptom networks.
- Demonstrates the ability to distinguish between redundant and synergistic information in mental health symptoms.
- ACIB embedding outperformed other methods in accurately recovering reference decompositions.
- Findings reveal significant differences in information structure between the PHQ-9 and Interpersonal Reactivity Index.
Read more
Scalable partial information decomposition for symptom networks via supervised embeddings
Summary
This paper addresses the limitations of traditional pairwise relationships in mental health symptom networks, which are often summarized as scalar edge weights. These weights fail to capture the complexity of interactions between symptoms, particularly regarding overlapping and synergistic information. The authors introduce embedding-based partial information decomposition (ePID), a scalable method that compresses non-focal symptoms into low-cardinality discrete embeddings and computes tractable two-source PID. This approach allows for the identification of unique, redundant, and synergistic components for each source-target pair. The authors benchmarked 13 candidate embeddings on synthetic Bayesian networks and 83 real-world datasets, finding that a supervised Agglomerative Conditional Information Bottleneck (ACIB) embedding performed best in recovering reference decompositions. The results indicated that in the PHQ-9 networks, most pairwise dependence was through redundant and remainder-unique channels, with synergy contributing only 6-9%. In contrast, the Interpersonal Reactivity Index showed a significant portion of its information as synergistic. The findings suggest that ePID can effectively distinguish between overlapping and interaction-dependent information, providing a model-agnostic complement to existing symptom network methodologies.
Methodology
The authors developed ePID, which compresses non-focal symptoms into low-cardinality embeddings and computes two-source partial information decomposition. They benchmarked various embeddings on synthetic and real-world datasets to evaluate performance across different PID measures.
Results
The supervised ACIB embedding achieved the highest accuracy in recovering reference decompositions across all PID measures, particularly excelling when four or more symptoms were compressed. In the PHQ-9 networks, redundancy dominated the pairwise dependence, while synergy was minimal. Conversely, the Interpersonal Reactivity Index exhibited a higher proportion of synergistic information.
Implications
The ePID method provides a new framework for analyzing symptom networks, potentially guiding clinical assessments and interventions by revealing how symptoms interact. This could lead to more effective treatment strategies that consider the complex relationships between symptoms rather than treating them in isolation.
Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport
Reinforcement Learning
Generative Models
Robotics
- Introduces OptiFlow, a framework for learning one-step flow policies in offline RL.
- Addresses issues of mode collapse and overestimation bias in multimodal action distributions.
- Utilizes state-wise entropic optimal transport to couple action samples from a reference policy and a one-step policy.
- Empirical results show strong performance across diverse offline RL benchmarks.
Read more
Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport
Summary
The paper addresses the challenges of offline reinforcement learning (RL), particularly in learning effective multimodal one-step flow policies from fixed datasets. The authors introduce a novel framework called One-step Flow policy via Optimal Transport (OptiFlow), which reformulates the learning process as a structured sample-allocation problem. OptiFlow combines a value-aware reference flow policy with an efficient one-step policy through state-wise entropic optimal transport. This approach prioritizes high-return actions while maintaining geometrically compatible pairings, thus avoiding common pitfalls such as mode collapse and overestimation bias. The experimental results demonstrate that OptiFlow successfully captures optimal multimodal behaviors and outperforms existing methods across various offline RL benchmarks, showcasing its potential for applications in domains where safe exploration is critical.
Methodology
The methodology involves reframing one-step flow policy learning as a structured sample-allocation problem. OptiFlow employs entropic optimal transport to couple samples from a value-aware reference flow policy and an efficient one-step policy. The critic-estimated values guide the prioritization of actions, while the transport cost ensures geometrically compatible pairings, allowing for flexible mapping of noise to actions without direct critic maximization.
Results
The experimental results indicate that OptiFlow effectively captures multimodal behaviors and achieves superior performance compared to existing offline RL methods across various benchmarks, demonstrating its robustness and efficiency in learning from fixed datasets.
Implications
The findings suggest that OptiFlow can be applied in critical domains such as robotics and control, where safe exploration is necessary and multimodal action distributions are prevalent. This framework could enhance decision-making processes in environments where traditional methods struggle with distribution shifts and multimodal behaviors.
Structured Features Overfit Where Random Features Grok
Theory
Optimization
- Structured feature maps do not exhibit the grokking phenomenon observed in random feature maps.
- Increasing the band width in structured feature maps leads to significant overfitting and a drop in held-out accuracy.
- The number of active modes in the feature map is a critical factor influencing model performance.
- Representability of targets can be determined in closed form, providing insights into feature map design.
Read more
Structured Features Overfit Where Random Features Grok
Summary
This paper investigates the phenomenon of 'grokking' in machine learning, specifically how structured feature maps behave differently from random feature maps in the context of over-parameterized ridge regression. Previous work established that over-parameterized models trained on unstructured random Gaussian features exhibit a delay between memorization and generalization, quantified by the weight decay parameter. However, this study shows that when using structured feature maps, such as a band-limited Fourier feature map, this delay does not manifest. Instead, as the band width increases, the model's held-out accuracy degrades significantly, indicating overfitting without a transition to generalization. The authors demonstrate that the number of active modes in the feature map, rather than the capacity ratio, is crucial in determining the model's performance. They provide a closed-form condition for representability of targets within the feature map, confirming that the structural properties of the feature map play a significant role in the learning dynamics. This work highlights the importance of feature geometry in understanding generalization behavior in machine learning models.
Methodology
The authors utilize a band-limited Fourier feature map to analyze the generalization behavior of over-parameterized ridge regression models. They conduct experiments by varying the band width and observing the effects on held-out accuracy. The representability of targets is assessed through a closed-form condition derived from Fourier analysis, allowing for a theoretical understanding of the feature map's structure.
Results
The main findings indicate that as the band width increases, the model's held-out accuracy decreases from 1.00 to 0.07, with no evidence of a memorization-then-generalization regime. A masking experiment reveals that fixing the nominal dimension while varying the active modes restores high accuracy, underscoring the importance of active support in feature maps.
Implications
This research suggests that the design of feature maps is critical for achieving effective generalization in machine learning models. Understanding the structural properties of feature maps can lead to better model architectures and training strategies, particularly in scenarios where overfitting is a concern.
High-Probability Nash Regret for Decentralized Learning in Markov $ฮฑ$-Potential Games: Episodic and Fully Online Asynchronous Algorithms with Applications to Markov Congestion Games
Reinforcement Learning
Theory
Optimization
- Establishes high-probability NE regret bounds for decentralized learning in Markov ฮฑ-potential games.
- Develops KL-projected natural policy gradient algorithms for both episodic and fully online settings.
- Eliminates distribution-mismatch coefficients from regret bounds, enhancing scalability.
- Addresses challenges of asynchronous updates and drifting state occupancies in fully online learning.
Read more
High-Probability Nash Regret for Decentralized Learning in Markov $ฮฑ$-Potential Games: Episodic and Fully Online Asynchronous Algorithms with Applications to Markov Congestion Games
Summary
This paper investigates decentralized learning of Nash equilibria (NE) in infinite-horizon discounted Markov games, specifically focusing on Markov ฮฑ-potential games. The author develops KL-projected natural policy gradient (NPG) algorithms in two distinct settings: an episodic setting with fixed policies during sampling and a fully online setting where players receive a single cost sample per time step and update policies asynchronously. The study establishes finite-time high-probability bounds on the time-averaged NE gap (NE regret) for both settings. Notably, the bounds do not include a distribution-mismatch coefficient, which can be problematic in large state spaces. In the episodic setting, the NE regret is bounded by eO(T โ1/4), while the fully online setting yields a bound of eO(T โ2/15), both accommodating various approximation errors. The paper also addresses challenges in the fully online setting, such as asynchronous state visitation and simultaneous policy adaptation, using advanced techniques like stopping-time and dynamic-tracking arguments. Furthermore, the framework is applied to independent-resource Markov congestion games (IMCGs), demonstrating its utility in constructing decentralized estimation oracles and deriving guarantees for both episodic and online settings. A novel application in strategic online job-scheduling for stochastic machines showcases the framework's scalability and effectiveness in learning stable dispatching policies.
Methodology
The paper employs KL-projected natural policy gradient algorithms in two settings: episodic (with fixed policies during sampling) and fully online (with asynchronous updates). It uses techniques such as stopping-time, coupling, and dynamic-tracking arguments to derive finite-time high-probability bounds on NE regret.
Results
The study presents NE regret bounds of eO(T โ1/4) for the episodic setting and eO(T โ2/15) for the fully online setting, accommodating various approximation errors. The framework successfully applies to independent-resource Markov congestion games, providing guarantees for decentralized learning algorithms.
Implications
The results have significant implications for decentralized learning in multi-agent systems, particularly in environments where players learn from local observations without sharing information. The framework's scalability and effectiveness in job scheduling can enhance operational efficiency in various applications, including transportation and resource management.
A Three-Axis Stress Test of LLM vs Classical ML for Network Intrusion Detection under Distribution Shift and Adversarial Evasion
Large Language Models
NLP
Theory
- Same-dataset evaluations are inadequate for assessing NIDS performance in real-world scenarios.
- XGBoost outperforms RoBERTa-LoRA in cross-dataset transfer, while RoBERTa-LoRA excels in adversarial evasion.
- The choice of model depends on the evaluation axis, emphasizing the need for multi-faceted assessments.
- Feature-leakage ablation studies reveal non-monotonic improvements in cross-dataset transfer.
Read more
A Three-Axis Stress Test of LLM vs Classical ML for Network Intrusion Detection under Distribution Shift and Adversarial Evasion
Summary
This paper evaluates the performance of large language models (LLMs) against classical machine learning (ML) models in the context of network intrusion detection systems (NIDS). The authors argue that traditional evaluations using the same dataset are insufficient, as they do not account for real-world scenarios involving distribution shifts and adversarial evasion. The study compares XGBoost and RoBERTa-LoRA across three axes: same-dataset performance, cross-dataset transfer, and adversarial evasion. Results show that while both models perform similarly on the same dataset, XGBoost significantly outperforms RoBERTa-LoRA under cross-dataset conditions, achieving higher F1 and balanced accuracy scores. Conversely, RoBERTa-LoRA excels in adversarial evasion scenarios. The findings highlight the importance of evaluating NIDS models across multiple robustness axes and suggest that model recommendations should depend on the specific evaluation context rather than relying solely on same-dataset accuracy.
Methodology
The authors conducted a comparative analysis of XGBoost and RoBERTa-LoRA on two independently collected NetFlow v2 datasets. They evaluated the models across three axes: same-dataset performance, cross-dataset transfer, and adversarial evasion. A staged feature-leakage ablation study was also performed to understand the impact of feature representation on model performance.
Results
The results indicated that both models were statistically tied in same-dataset evaluations. XGBoost achieved a 15-point advantage in F1 score and a 25-point advantage in balanced accuracy under cross-dataset conditions. In contrast, RoBERTa-LoRA outperformed XGBoost by approximately 17 points in F1 score during adversarial evasion tests, maintaining low false positive rates below 0.01.
Implications
The findings suggest that NIDS evaluations should incorporate multiple robustness axes to provide a comprehensive understanding of model performance. This approach can lead to better-informed decisions regarding model deployment in real-world cybersecurity scenarios.
WaterKron and FlipFlop Hessian: Information-Theoretically Grounded Quantization with Kronecker-factored Hessians
Theory
Efficient ML
Optimization
- Introduction of WaterKron for optimal precision allocation in quantization.
- Development of a mismatch factor ฮฆ to quantify distortion penalties in Kronecker approximations.
- Proposal of the FlipFlop Hessian, which improves upon existing Hessian approximations.
- Empirical results show significant improvements in quantization performance metrics.
Read more
WaterKron and FlipFlop Hessian: Information-Theoretically Grounded Quantization with Kronecker-factored Hessians
Summary
This paper addresses the challenge of selecting a Kronecker-factored Hessian approximation for post-training quantization of neural networks. The authors introduce WaterKron, which integrates two-sided Generalized Post-Training Quantization (GPTQ) with row- and column-dependent waterfilling scales and entropy coding. They derive a high-rate distortion measure that incorporates a Kronecker-Hessian mismatch factor, ฮฆ, which quantifies the distortion penalty associated with the Kronecker approximation. By minimizing ฮฆ, the authors formulate a Gaussian covariance-fitting problem that leads to the FlipFlop Hessian, a novel approximation that enhances quantization performance. Empirical evaluations demonstrate that the FlipFlop Hessian consistently outperforms traditional Hessian choices in terms of Kullback-Leibler divergence and perplexity, thereby providing a robust framework for efficient neural network quantization.
Methodology
The authors extend the existing waterfilling allocation methods to a two-sided GPTQ framework, deriving optimal row and column scales for quantization. They analyze the distortion in the full Hessian geometry and introduce the mismatch factor ฮฆ to guide the selection of Hessian factors. The FlipFlop Hessian is derived through Gaussian covariance fitting, optimizing the quantization process.
Results
The empirical evaluations indicate that the FlipFlop Hessian consistently yields lower Kullback-Leibler divergence and perplexity compared to input-only, marginal, and Frobenius-based Hessian approaches. The WaterKron method also demonstrates superior performance in precision allocation compared to previous methods.
Implications
The findings suggest that the proposed methods can significantly enhance the efficiency of neural network quantization, making it feasible to deploy high-performance models in resource-constrained environments. This work could lead to advancements in various applications requiring efficient model deployment, such as mobile and edge computing.
Follow the Geometry, Not the Model: Cold Start Semi-Supervised Learning
Computer Vision
Theory
Graph Learning
- VAST decouples pseudo-label inference from classifier training, addressing cold-start SSL challenges.
- The Veracity Matrix aggregates label evidence using a kernel-based approach, enhancing belief propagation.
- VAST outperforms traditional graph-based SSL methods in cold-start scenarios across multiple datasets.
- The method provides a deployable inductive classifier, avoiding the need for transductive re-solving.
Read more
Follow the Geometry, Not the Model: Cold Start Semi-Supervised Learning
Summary
This paper addresses the challenges of cold-start semi-supervised learning (SSL), where only a few labeled samples per class are available. Traditional SSL methods rely on the classifier's confidence to generate pseudo-labels, which becomes problematic in cold-start scenarios as the classifier cannot effectively supervise itself. The authors propose a novel approach called VAST (Veracity-Aware Semi-Supervised Training) that decouples the stages of pseudo-label generation and classifier training. VAST first infers probabilistic beliefs over the unlabeled data using the geometry of a frozen self-supervised embedding, and then distills these beliefs into an inductive classifier. The key innovation is the Veracity Matrix, which aggregates label evidence across the data manifold and allows for a belief propagation step that extends beyond the immediate neighborhood of labeled samples. The results demonstrate that VAST significantly outperforms existing graph-based SSL methods across multiple datasets, achieving statistically significant improvements while producing a deployable inductive classifier, contrasting with the transductive nature of many existing methods.
Methodology
The methodology involves constructing the Veracity Matrix to represent beliefs over class labels based on the geometry of a frozen self-supervised embedding. Label evidence is propagated through a kernel similarity structure, allowing for the selection of high-confidence pseudo-labels. The training loss combines cross-entropy on the labeled set with a distillation term on the soft pseudo-labels, enabling the classifier to learn from these beliefs after they have been established.
Results
VAST consistently outperformed the strongest graph-based SSL baselines across three datasets, achieving statistically significant gains in 7 out of 9 comparisons. For instance, in experiments with CIFAR-100, VAST achieved 39.2% accuracy with one label per class, compared to 30.7% for CoMatch and 9.8% for FreeMatch, demonstrating its effectiveness in low-label regimes.
Implications
The findings suggest that VAST could be particularly useful in domains where labeled data is scarce and expensive to obtain, such as medical imaging or rare-event detection. By decoupling the inference and training processes, VAST provides a robust framework for leveraging unlabeled data effectively, potentially leading to advancements in various applications of semi-supervised learning.
Graph Neural Networks for Influence Maximization in Social Networks: An Unsupervised Minimum Dominating Set Approach
Graph Learning
Optimization
- Introduces an unsupervised GNN framework for solving the MDS problem.
- Achieves significant speed improvements in inference time compared to existing methods.
- Demonstrates strong generalization to unseen graph distributions.
- Utilizes a novel multi-objective loss function for enhanced training signals.
Read more
Graph Neural Networks for Influence Maximization in Social Networks: An Unsupervised Minimum Dominating Set Approach
Summary
This paper addresses the Minimum Dominating Set (MDS) problem, a well-known NP-hard combinatorial optimization issue with significant applications in social network analysis, such as viral marketing and public health interventions. The authors propose a novel unsupervised framework utilizing Graph Neural Networks (GNNs) to tackle the MDS problem without requiring ground-truth solutions during training. The proposed method is trained on a dataset of 12,000 synthetic graphs with varying structural properties. The results indicate that the GNN approach achieves up to 55 times faster inference compared to metaheuristic baselines and up to 14 times faster than supervised learning methods, while still identifying optimal or near-optimal dominating sets on real-world social network benchmarks. The framework demonstrates strong generalization capabilities, effectively handling unseen graph distributions and maintaining performance across diverse graph structures, making it suitable for large-scale social network analysis.
Methodology
The authors developed a GNN architecture that employs a multi-objective probabilistic loss function, which combines three components to provide robust training signals. The model is trained on synthetic graphs and utilizes the Global Prediction paradigm for efficient node-wise probability generation in a single forward pass, enhancing computational efficiency.
Results
The proposed GNN framework outperformed traditional metaheuristic and supervised learning approaches, achieving up to 55ร faster inference and maintaining near-optimal performance on real-world benchmarks. The model also showed strong generalization capabilities, effectively adapting to larger and unseen graph structures.
Implications
The findings suggest that the unsupervised GNN approach can significantly enhance the efficiency of influence maximization strategies in social networks, with potential applications in marketing, public health, and information dissemination. The ability to operate without labeled data makes it particularly valuable for real-world scenarios where obtaining optimal solutions is computationally prohibitive.
Task-Directed Residual AddUNet: Perfect-Reconstruction Routing for Full-Rate Representations
Audio & Speech
Theory
Efficient ML
- Establishes the equivalence between constrained additive U-Net and critically sampled PR filter banks.
- Introduces a Residual Full-Rate PR architecture that routes task-irrelevant information while ensuring exact reconstruction.
- Demonstrates that the architecture does not require invertibility or learned decoders for reconstruction.
- Achieves improved phone recognition performance on the TIMIT dataset while maintaining exact reconstruction.
Read more
Task-Directed Residual AddUNet: Perfect-Reconstruction Routing for Full-Rate Representations
Summary
This paper presents a novel interpretation of the AddUNet architecture through the lens of perfect reconstruction (PR) and introduces a Residual Full-Rate PR architecture aimed at task-directed representation learning. The author establishes that the survivor-skip structure of a constrained additive U-Net is equivalent to a critically sampled multirate PR filter bank. By transitioning to a full-rate formulation, the complementary-subband restrictions are eliminated while maintaining perfect reconstruction. The proposed architecture allows for the explicit routing of task-irrelevant information away from the task-facing survivor, ensuring that the routed information is retained. This architecture does not require invertibility, matched synthesis banks, reconstruction loss, or learned decoders for exact reconstruction. The paper also connects this framework to identity-shortcut ResNets, demonstrating that the residual output can be viewed as a full-rate PR system. Experimental results on the TIMIT dataset show that the proposed model improves phone recognition performance while ensuring exact reconstruction, indicating that structural conservation does not inherently lead to task-specific invariance.
Methodology
The methodology involves establishing the theoretical foundation of the AddUNet architecture as a PR filter bank and developing a full-rate formulation that allows for explicit routing of information. The architecture is tested experimentally on the TIMIT dataset to validate its performance in task-directed representation learning.
Results
The proposed Residual Full-Rate PR architecture demonstrated an improvement in phone error rate (PER) from 28.60% to 25.76% on the TIMIT dataset, while ensuring exact reconstruction. The experiments confirmed the architecture's ability to route linearly separable factors accurately and maintain structural conservation.
Implications
The findings suggest that the proposed architecture can enhance representation learning in various signal processing tasks, particularly in audio and speech recognition, by allowing for more efficient routing of information and improved performance without compromising reconstruction quality.
Safe Meta-Reinforcement Learning via Information Space Reachability
Reinforcement Learning
Robotics
Theory
- Introduces a safe meta-RL framework that explicitly incorporates safety during task adaptation.
- Develops a safety value function that captures safety in the information space, enhancing the agent's ability to manage task uncertainty.
- Proposes a practical algorithm for high-dimensional systems that uses neural networks for value function and policy approximation.
- Demonstrates the effectiveness of the proposed method on standard meta-RL benchmarks, showing improved safety and performance.
Read more
Safe Meta-Reinforcement Learning via Information Space Reachability
Summary
This paper addresses the challenges of applying meta-reinforcement learning (meta-RL) in real-world scenarios where safety is a critical concern. The authors propose a novel framework that incorporates safety considerations during the adaptation of agents to new tasks. The key innovation is the introduction of a safety value function that operates within an information space, which encompasses both the physical state and the agent's belief about the task. This safety value function is shown to satisfy self-consistency and a Bellman equation, allowing it to be learned through meta-RL techniques. The proposed algorithm focuses on state-wise safety constraints, ensuring that safety is maintained at every state visited by the agent. The authors validate their approach through experiments on established meta-RL benchmarks, demonstrating that their method effectively balances safety and performance.
Methodology
The authors formulate the problem of safety preservation in meta-RL as a Bayes-adaptive reachability problem in the information space. They introduce a safety value function that is learnable via meta-RL, satisfying self-consistency and Bellman equations. The proposed algorithm approximates value functions and policies using neural networks and employs a neural encoder to infer latent task representations from online interactions.
Results
The experiments conducted on widely used meta-RL benchmarks indicate that the proposed safe meta-RL algorithm successfully maintains safety while achieving competitive performance in task adaptation. The results highlight the advantages of reasoning about safety in the information space compared to traditional methods.
Implications
This research has significant implications for deploying meta-RL in real-world applications, particularly in safety-critical domains such as robotics and autonomous systems. By ensuring safety during task adaptation, the proposed framework can help prevent catastrophic failures and enhance the reliability of intelligent agents.
An immune world model for multiscale forecasting and therapeutic hypothesis generation
Theory
Generative Models
Optimization
- Introduces the Immune World Model for multiscale immune forecasting.
- Utilizes a governed evolutionary AI Scientist for model construction.
- Demonstrates the model's ability to generalize to unseen interventions.
- Identifies IL-36ฮณ plus SIRPฮฑ inhibition as a therapeutic hypothesis.
Read more
An immune world model for multiscale forecasting and therapeutic hypothesis generation
Summary
The paper presents the Immune World Model, a novel framework designed to forecast immune responses and generate therapeutic hypotheses by integrating data across cellular, tissue, and individual levels. Traditional models often treat these scales separately, leading to incomplete predictions. The authors utilized a governed evolutionary AI Scientist, named Agent Genesis, to construct this model, which learns how interventions affect immune states across different biological hierarchies. The model was rigorously tested and frozen before independent validation, demonstrating its ability to generalize to unseen interventions and biological contexts. It effectively recovered intervention-specific cellular programs and improved predictions of ecosystem and patient responses. The analysis guided by the Immune World Model identified IL-36ฮณ plus SIRPฮฑ inhibition as a promising therapeutic hypothesis, showcasing its potential for generating testable hypotheses in immunotherapy. This work highlights the importance of a multiscale approach in understanding immune responses and offers a robust framework for future therapeutic exploration.
Methodology
The Immune World Model was constructed using Agent Genesis, an AI Scientist that evolved candidate workflows and model configurations. The model integrates multiscale immune data and encodes interventions to simulate transitions across cellular, tissue, and individual states. The model was frozen after evaluation and then tested for generalization on withheld interventions and contexts.
Results
The Immune World Model successfully linked cellular, tissue, and individual immune states, allowing for accurate predictions of immune responses to various interventions. It recovered specific cellular programs related to interventions and improved predictions of patient responses. The model also identified a promising therapeutic hypothesis involving IL-36ฮณ and SIRPฮฑ inhibition.
Implications
The Immune World Model provides a comprehensive framework for simulating immune responses, which could significantly enhance the development of immunotherapies. Its ability to generate testable hypotheses may lead to more effective treatment strategies in cancer and other diseases influenced by immune responses.
A Machine Learning Framework for Fault Detection, Isolation, and Severity Prediction of Autonomous VTOL Aircraft
Robotics
- Developed a CNN-based framework for fault detection and severity prediction in autonomous VTOL aircraft.
- Achieved over 99% accuracy in rotor fault classification and ~96% in severity estimation.
- Validated the framework using both simulated data and real-world experiments on a hexacopter.
- Addresses challenges of sensor noise and environmental disturbances in fault detection.
Read more
A Machine Learning Framework for Fault Detection, Isolation, and Severity Prediction of Autonomous VTOL Aircraft
Summary
This paper presents a novel machine learning framework aimed at enhancing the safety and reliability of autonomous Vertical Take-Off and Landing (VTOL) aircraft by addressing the critical issue of fault detection, isolation, and severity prediction. The authors highlight the challenges posed by sensor noise, environmental disturbances, and the complex nonlinear aerodynamics of multirotor platforms, which complicate real-flight fault detection. To tackle these issues, the study develops a convolutional neural network (CNN) architecture that learns spatio-temporal patterns from multivariate flight dynamics. This architecture enables the direct inference of both the faulty rotor and the severity of its damage. The framework is validated through simulations using a data-generative model and further tested on a hexacopter with controlled blade-tip breakage. The results demonstrate that the trained model achieves rotor-wise fault classification accuracies exceeding 99% and a severity estimation accuracy of approximately 96% within a ยฑ1% tolerance in experimental data. These findings indicate strong generalization capabilities and support the feasibility of real-time health monitoring for autonomous VTOL systems.
Methodology
The authors developed a convolutional neural network (CNN) architecture to analyze multivariate flight dynamics data, enabling the detection and isolation of faults in rotors and predicting the severity of these faults. The framework was first validated using simulated data and then experimentally tested on a hexacopter with controlled rotor damage.
Results
The CNN model demonstrated exceptional performance, achieving rotor-wise fault classification accuracies above 99% and a severity estimation accuracy of approximately 96% within a ยฑ1% tolerance during experimental validation. This indicates the model's strong generalization and effectiveness in real-time applications.
Implications
The proposed framework has significant implications for the safety and reliability of autonomous VTOL aircraft, enabling early detection of faults and facilitating timely interventions to prevent accidents. It can be applied in various UAV operations, including delivery, surveillance, and emergency response, where safety is paramount.
Entropy-Punctured Bloom Filters for Memory-Efficient Machine Learning
Efficient ML
- Introduction of entropy-punctured Bloom Filters for memory-efficient feature representation.
- Focus on regression tasks, addressing a gap in the application of Bloom Filters in this area.
- Evaluation of predictive efficiency as a new metric for assessing the trade-off between performance and representation size.
- Demonstration of substantial storage savings with minimal loss in predictive fidelity compared to classical methods.
Read more
Entropy-Punctured Bloom Filters for Memory-Efficient Machine Learning
Summary
This paper introduces entropy-punctured Bloom Filters (EPBFs), a novel encoding strategy aimed at enhancing memory efficiency in machine learning applications, particularly for regression tasks. Traditional Bloom Filters (BFs) provide compact probabilistic representations of features but have not been extensively explored for regression under memory constraints. The authors propose a method that selectively removes low-variability bit positions from fixed-length BF encodings based on empirical entropy, resulting in reduced representations that maintain predictive structure while improving predictive efficiency. The study evaluates EPBFs against various regression datasets, comparing their performance with raw features, Principal Component Analysis (PCA), Random Projection (RP), and other Bloom Filter variants. The evaluation focuses on predictive efficiency, defined as the R2 score relative to encoded representation size per sample. The results indicate that EPBFs achieve substantial storage savings while remaining competitive with classical compressed representations, demonstrating minimal loss in predictive fidelity. This work highlights the potential of EPBFs as an effective representation-level compression approach for memory-constrained machine learning scenarios.
Methodology
The authors developed entropy-punctured Bloom Filters by removing low-variability bit positions from fixed-length Bloom Filter encodings based on empirical entropy. They evaluated the approach on diverse regression datasets using ridge regression, XGBoost, and neural networks, comparing it against raw features and classical dimensionality reduction methods under leakage-free evaluation protocols.
Results
The results showed that entropy-punctured Bloom Filters maintained competitive predictive fidelity while significantly improving predictive efficiency. The approach yielded substantial storage savings compared to traditional methods, demonstrating the effectiveness of entropy-guided structural compression in memory-constrained machine learning.
Implications
The findings suggest that entropy-punctured Bloom Filters can be effectively utilized in scenarios where memory efficiency is critical, such as in edge computing or IoT applications. This approach could lead to more efficient machine learning models that require less storage and bandwidth, making them suitable for deployment in resource-constrained environments.
Are Gradient Boosting Models Suitable for Intermittent Demand Forecasting?
Time Series
- Gradient boosting models tend to underperform in isolation for intermittent demand forecasting.
- Specialized forecasting methods achieve the best performance among individual models.
- Combining machine learning models with specialized approaches can improve accuracy by up to 10%.
- The study addresses gaps in comparative evaluations of traditional and machine learning methods.
Read more
Are Gradient Boosting Models Suitable for Intermittent Demand Forecasting?
Summary
This paper investigates the effectiveness of gradient boosting models in forecasting intermittent demand, a challenging task characterized by infrequent and irregular sales patterns. The authors highlight the significance of accurate demand forecasting in various industries, where poor predictions can lead to substantial financial consequences. The study evaluates a range of forecasting methods, including classical time series approaches, specialized techniques for intermittent demand, machine learning models, and hybrid methods that combine these approaches. The experimental evaluation is conducted using multiple datasets, including the Royal Air Force dataset, to compare the performance of these models. The findings reveal that while specialized methods generally outperform gradient boosting models when used alone, combining gradient boosting with specialized techniques can enhance forecasting accuracy by up to 10%. This suggests that simple ensemble methods can effectively leverage the strengths of different modeling approaches, providing valuable insights for both academia and industry in improving demand forecasting strategies.
Methodology
The authors conducted an experimental evaluation of various forecasting methods, including classical time series (ETS), specialized techniques (SBA, TSB, ADIDA, IMAPA), machine learning models (XGBoost, CatBoost), and hybrid methods. The models were assessed using a unified experimental setup on the Royal Air Force dataset, with performance measured through train-test splits and careful validation to ensure robust training.
Results
The results indicate that specialized methods outperform gradient boosting models when used individually. However, hybrid models that combine gradient boosting with specialized techniques show a significant improvement in forecasting accuracy, achieving up to a 10% increase in performance.
Implications
The findings suggest that integrating machine learning with domain-specific forecasting techniques can enhance demand forecasting for products with intermittent demand. This has practical implications for inventory management and supply chain optimization across various industries, potentially leading to reduced costs and improved operational efficiency.
Temporal Self-Distillation: Faster Inference in Discrete Diffusion Language Models
NLP
Large Language Models
Generative Models
- Introduction of Temporal Self-Distillation (TSD) for dLLMs.
- TSD distills predictions across time, enhancing early token predictions.
- Eliminates the need for offline teacher generation, simplifying the training process.
- Demonstrated significant improvements in speed and quality across multiple benchmarks.
Read more
Temporal Self-Distillation: Faster Inference in Discrete Diffusion Language Models
Summary
This paper introduces Temporal Self-Distillation (TSD), a novel on-policy method designed to enhance the inference speed of discrete diffusion language models (dLLMs) while maintaining or improving generation quality. dLLMs have the potential for rapid inference by generating multiple tokens simultaneously, but they often experience significant performance degradation when parallel decoding is overly aggressive. TSD addresses this issue by training the model to distill its predictions across time, specifically aligning earlier predictions with the final output distribution at the time a token is committed. This self-distillation approach eliminates the need for offline teacher generation, allowing for a seamless integration into both base and post-trained policies. The authors demonstrate TSD's effectiveness across seven benchmarks in mathematics, planning, and coding, showing that it significantly shifts the speed-quality trade-off towards lower computational requirements. TSD achieves competitive speedups comparable to offline distillation methods while simplifying the training process into a single-stage pipeline.
Methodology
The methodology involves an on-policy self-distillation approach where the model's earlier predictions are trained to align with its final predictions at the time of token commitment. This is achieved without requiring a separate teacher model or offline data generation, allowing for a streamlined training process that operates across denoising timesteps.
Results
TSD was evaluated on seven benchmarks, showing substantial improvements in the speed-accuracy trade-off. For instance, on the MBPP benchmark, TSD achieved a 40% pass rate with approximately 2.3 times fewer forward passes compared to the base LLaDA model, and 1.5 times fewer than the offline dParallel distillation method, indicating its effectiveness in enabling aggressive parallel decoding.
Implications
The implications of this work suggest that TSD can significantly enhance the efficiency of dLLMs, making them more viable for real-time applications where computational resources are limited. This could lead to broader adoption of dLLMs in various fields such as natural language processing, automated planning, and code generation.
A Variational Optimal Transport Operator on Incompressible Flow
Generative Models
Optimization
Computer Vision
- Introduction of the VIOT operator for efficient incompressible density transport.
- Amortized optimization approach allows for real-time transport generation without trajectory supervision.
- Demonstrated significant speedup (10,000x) over traditional optimization methods.
- Utilizes a Fourier Neural Operator for flexibility across different grid resolutions.
Read more
A Variational Optimal Transport Operator on Incompressible Flow
Summary
The paper introduces the Variational Incompressible Optimal Transport (VIOT) operator, a generative neural operator designed for efficient and accurate incompressible density transport. Unlike traditional methods that require extensive optimization for each source-target density pair, VIOT predicts a divergence-free velocity field and generates transport trajectories through a feed-forward inference process. The operator is built on three main components: a stream-function or vector-potential representation that inherently enforces incompressibility, a regularized transport objective that balances accuracy and smoothness, and a Fourier Neural Operator backbone that allows for amortization across various density pairs and grid resolutions. The authors demonstrate that VIOT can produce high-quality transport results in real-time, achieving a significant speedup (approximately 10,000 times faster) compared to existing per-instance optimization methods. The system is validated on multiple 2D and 3D benchmarks, showcasing its versatility and effectiveness in generating incompressible transports for user-defined source-target pairs in an interactive manner.
Methodology
The VIOT operator employs a stream-function or vector-potential representation to ensure incompressibility, combined with a regularized transport objective that balances endpoint accuracy and flow smoothness. The use of a Fourier Neural Operator enables the model to generalize across various density pairs and grid resolutions, allowing for efficient computation of the transport trajectories through feed-forward inference.
Results
The VIOT operator was tested on various 2D and 3D density-transport benchmarks, demonstrating the ability to generate transport trajectories in seconds per pair, compared to the hour-long optimization required by traditional methods. The results showed high accuracy in maintaining volume-preserving properties and smooth flow, with effective performance across different applications, including shape and font evaluations as well as pose interpolation.
Implications
The development of the VIOT operator has significant implications for computer graphics and simulations involving fluid dynamics, enabling faster and more efficient transport processes. Its ability to handle user-defined inputs in real-time opens up new possibilities for interactive applications in animation, gaming, and virtual reality.
Land Art as a Big-Data Climate Sensor
Time Series
- Developed a 14-feature complexity signature for analyzing climate signals from satellite imagery of Spiral Jetty.
- Found strong correlations between image complexity features and climate variables, particularly lake elevation and CO2 levels.
- Refined the concept of art as a climate indicator, proposing a framework for understanding art's role in environmental monitoring.
- Released a public benchmark dataset of 1,744 satellite images and analysis code for further research.
Read more
Land Art as a Big-Data Climate Sensor
Summary
This paper investigates Robert Smithson's land artwork, Spiral Jetty, as a climate sensor by analyzing 1,744 satellite images from Landsat and Sentinel-2 spanning from 1984 to 2025. The authors develop a 14-feature complexity signature that includes various metrics such as Shannon entropy and deep learning features from a ResNet50 model. The study aims to determine the relationship between image complexity and climate variables, particularly lake elevation and CO2 levels. The findings reveal that certain complexity features, particularly permutation entropy and deep learning embeddings, correlate strongly with climate indicators, suggesting that the artwork can serve as a leading indicator of hydrological states. This research refines the notion of 'art-as-thermometer' into a more robust framework of 'art-as-leading-indicator-of-hydrological-state.' The dataset and analysis code are made publicly available, contributing to the field of remote sensing and climate studies.
Methodology
The authors conducted a comprehensive analysis of 1,744 co-registered Landsat and Sentinel-2 satellite images, creating a complexity signature that combines classical image analysis metrics and deep learning features. They validated their findings through year-aggregated bootstrap analysis, lag analysis, and partial correlations controlling for seasonal and sensor confounds.
Results
The study found that Shannon entropy is a weak proxy for climate signals, while permutation entropy and mean intensity strongly correlate with lake elevation. The third principal component of ResNet50 features emerged as a significant 'AI climate axis' with high correlations to CO2 levels and lake elevation. Additionally, image complexity was shown to lead lake stage changes by approximately three years.
Implications
This research has implications for using art and cultural heritage as tools for climate monitoring and understanding environmental changes. It opens avenues for further studies on the intersection of art, technology, and climate science, and provides a framework for analyzing other land art installations in similar contexts.
Beyond Numerical Time Series: A Unified Benchmark for Multimodal Forecasting with Heterogeneous Context
Time Series
Multimodal
- MUSE-Bench introduces a unified benchmark for multimodal time series forecasting with diverse contextual information.
- Numerical time series foundation models dominate the performance rankings, highlighting their effectiveness.
- External context significantly improves the performance of context-aware forecasting models.
- General-purpose LLMs perform poorly as direct forecasters, indicating a need for specialized approaches.
Read more
Beyond Numerical Time Series: A Unified Benchmark for Multimodal Forecasting with Heterogeneous Context
Summary
The paper addresses the limitations of existing time series forecasting benchmarks, which are predominantly numerical-centric and do not adequately evaluate the contextual information that influences real-world temporal dynamics. To overcome these challenges, the authors introduce MUSE-Bench, a comprehensive benchmark for multimodal time series forecasting that incorporates heterogeneous context. MUSE-Bench consists of fourteen datasets across eight domains, utilizing six types of contextual information: metadata, events, holidays, news, images, and numerical covariates. The authors evaluate various forecasting paradigms, including statistical methods, data-specific models, foundation models, multimodal models, and general-purpose large language models (LLMs) under consistent evaluation protocols. The findings reveal that numerical time series foundation models outperform other models, while the multimodal foundation model, Aurora, shows promise but lags behind numerical models. Additionally, the inclusion of external context enhances the performance of context-aware models, although misaligned context can hinder results. General-purpose LLMs are found to be ineffective as direct forecasters, and LLM-guided refinements do not consistently improve outcomes. MUSE-Bench serves as a foundational tool for systematically evaluating the role of context in forecasting models and paves the way for future research in multimodal forecasting.
Methodology
The authors constructed MUSE-Bench by compiling datasets from multiple domains and incorporating various types of context. They evaluated different forecasting paradigms under controlled conditions, ensuring consistent metrics and non-overlapping forecast windows for fair comparisons.
Results
The experiments revealed that numerical time series foundation models consistently ranked highest, while the multimodal model Aurora outperformed data-specific models but did not match the leading numerical models. Context-aware models benefited from external context, while misaligned context negatively impacted performance. General-purpose LLMs showed poor forecasting capabilities.
Implications
MUSE-Bench provides a robust framework for evaluating multimodal forecasting models, encouraging the development of more context-aware forecasting techniques. It can be utilized in various domains such as finance, healthcare, and transportation, where contextual information plays a crucial role in decision-making.