AI-generated summaries
Today's ML research,
without the noise.
Daily summaries of the latest machine learning papers from arXiv, processed every 8 hours.
24
Papers today
8h
Update frequency
7
Days of history
Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not
Optimization
Theory
- Introduction of Stiefel Attention, optimizing transformer projection matrices on the Stiefel manifold.
- Development of a Riemannian Adam optimizer that respects the geometry of the manifold.
- Significant performance improvements on benchmarks, with validation accuracy rising from 61.1% to 97.0%.
- The method's advantages grow with increasing data, indicating a strong relationship between geometry and optimization.
Read more
Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not
Summary
This paper introduces Stiefel Attention, a novel approach to optimizing the query and key projection matrices (WQ, WK) in transformer architectures by constraining them to the Stiefel manifold. Traditional methods use Euclidean optimizers without considering the geometric structure of these matrices, which can lead to suboptimal performance. The author proposes a Riemannian Adam optimizer that incorporates a scalar second moment per frame, a trust-region cap, and polar retraction to ensure that the optimization respects the manifold's geometry. The paper presents four propositions that demonstrate the effectiveness of this approach, including that it is steepest descent in the embedded metric and well-conditioned. The results show significant improvements in validation accuracy on benchmarks such as modular arithmetic grokking and CIFAR-10, with the Riemannian Adam achieving 97.0% accuracy compared to a baseline of 61.1%. The findings suggest that constraining WQ and WK to the Stiefel manifold not only enhances generalization but also maintains performance as data increases, challenging the conventional wisdom regarding optimizer choice in transformer models.
Methodology
The paper employs a Riemannian optimization approach using a modified Adam optimizer tailored for the Stiefel manifold. This includes a scalar second moment for each frame, a trust-region cap to ensure stability, and polar retraction to map updates back onto the manifold. The effectiveness of the method is validated through numerical propositions and extensive experiments on benchmark datasets.
Results
The proposed Riemannian Adam optimizer achieved a validation accuracy of 97.0% on modular arithmetic grokking tasks, significantly outperforming the baseline accuracy of 61.1%. Additionally, on CIFAR-10, the method demonstrated an improvement of +8.98 percentage points over 12 paired starts, with performance gains increasing as more data was introduced, indicating that the benefits of the method scale with dataset size.
Implications
The findings suggest that incorporating geometric constraints into the optimization of transformer models can lead to substantial performance improvements. This approach could influence future research on optimizer design and the architecture of neural networks, particularly in applications where attention mechanisms are critical.
Enhanced Agriculture-informed Neural Network by Domain Knowledge
Interpretability
- KAINN integrates domain knowledge into deep learning for improved N2O emission predictions.
- The framework incorporates key environmental processes to enhance model interpretability.
- KAINN outperforms traditional neural network models and the original AINN in accuracy.
- The approach demonstrates reduced uncertainty and improved stability in parameter trajectories.
Read more
Enhanced Agriculture-informed Neural Network by Domain Knowledge
Summary
This paper presents the Knowledge-enhanced Agriculture-informed Neural Network (KAINN), a novel framework designed to improve the prediction of nitrous oxide (N2O) emissions from agricultural practices by integrating domain knowledge into a hybrid neural-mechanistic modeling approach. The authors highlight the challenges in accurately modeling N2O emissions due to the complex interactions between soil, climate, and biochemical processes, compounded by the scarcity of high-quality data. KAINN builds upon the existing Agriculture-informed Neural Network (AINN) by explicitly incorporating critical environmental processes such as fertilizer diffusion, soil respiration rate, and water-filled porosity, which guide the learning process and constrain the solution space. The framework is evaluated using various architectures, including CNN, LSTM, and Transformer models, across multiple growing seasons and input configurations. Results indicate that KAINN consistently outperforms both traditional neural network models and the original AINN in terms of prediction accuracy, achieving lower Root Mean Square Error (RMSE) and Mean Absolute Error (MAE), as well as a higher coefficient of determination (R²). Additionally, the analysis of interface evolution reveals that KAINN produces smoother and more physically consistent parameter trajectories, enhancing interpretability and stability. This work underscores the significance of integrating domain knowledge into deep learning models for environmental applications, providing a scalable approach for reliable and interpretable predictions of N2O emissions in agricultural systems.
Methodology
The study employs a hybrid modeling framework that combines deep learning architectures (CNN, LSTM, Transformer) with mechanistic modeling principles. KAINN integrates domain knowledge related to environmental processes to guide the learning of the neural network, enhancing its predictive capabilities and interpretability.
Results
KAINN consistently achieved lower RMSE and MAE compared to traditional models and the original AINN, along with a higher R² value, indicating superior prediction accuracy and generalization across different environmental conditions. The model also produced smoother parameter trajectories, suggesting enhanced interpretability and reduced uncertainty.
Implications
The KAINN framework has significant implications for environmental modeling, particularly in predicting greenhouse gas emissions from agricultural systems. By integrating domain knowledge, it offers a more reliable and interpretable approach to modeling complex environmental processes, which can inform sustainable farming practices and climate mitigation strategies.
PhyRestore: Physics-Structured Latent-Factor Restoration
Time Series
Theory
Optimization
- PhyRestore restores corrupted physical factors from bitemporal spatial observations.
- The framework reconstructs temporal changes through the known RUSLE relationship.
- Factor restoration improves recovery of rare high-magnitude changes under certain conditions.
- The study compares multiple learning pathways for soil-loss prediction.
Read more
PhyRestore: Physics-Structured Latent-Factor Restoration
Summary
The paper addresses the challenge of estimating temporal soil-loss changes in the presence of noisy or corrupted input factors, particularly when significant changes are rare compared to numerous locations with minimal change. The authors introduce PhyRestore, a physics-structured latent-factor restoration framework that aims to restore corrupted physical factors and reconstruct temporal changes using the Revised Universal Soil Loss Equation (RUSLE). The study evaluates PhyRestore in a watershed-scale bitemporal raster setting, focusing on the isolated and simultaneous corruption of rainfall erosivity and cover management factors. The results indicate that factor restoration enhances recovery of high-magnitude changes when the corrupted factors are identifiable, but its effectiveness diminishes under joint corruption and when factor values fall outside the training support. The paper contributes to the understanding of signed temporal soil-loss prediction under controlled factor corruption and compares various learning pathways, including analytical estimates and machine learning models like Direct RF, XGBoost, MLP, and CNN.
Methodology
The authors formulated a framework that estimates clean values of corrupted RUSLE factors from bitemporal spatial context. They evaluated the performance of PhyRestore against degraded analytical estimates and various machine learning models (RF, XGBoost, MLP, CNN) under controlled conditions of factor corruption.
Results
The evaluation showed that PhyRestore significantly improves the recovery of high-magnitude soil-loss changes when the corrupted factors are identifiable. However, its advantages weaken in scenarios of joint corruption and when factor values are outside the training support, highlighting the limitations of the framework in certain conditions.
Implications
PhyRestore has potential applications in environmental monitoring and management, particularly in assessing soil erosion and land degradation. The framework can enhance the reliability of soil-loss estimates in regions where direct measurements are difficult to obtain, thereby aiding in better decision-making for land use and conservation practices.
Machine-Learning Assessment of the Predictive Value of Inflammatory Biomarkers for Cognitive Impairment in an Older Hispanic Adult Cohort
Interpretability
- I-309 (CCL1) was identified as a significant predictor of cognitive impairment, with a notable increase in predictive accuracy.
- The study employed a leakage-safe machine learning methodology to ensure robust results.
- The findings underscore the need for external validation of biomarkers in predicting cognitive impairment.
- The research distinguishes between statistical significance and predictive utility in the context of small clinical datasets.
Read more
Machine-Learning Assessment of the Predictive Value of Inflammatory Biomarkers for Cognitive Impairment in an Older Hispanic Adult Cohort
Summary
This study investigates the predictive utility of inflammatory biomarkers for cognitive impairment in older Hispanic adults using a machine learning approach. The authors utilized data from the Panama Aging Research Initiative–Health Disparities (PARI-HD) cohort, comprising 165 participants. They implemented a leakage-safe threshold-likelihood Bernoulli/Categorical Naive Bayes (BNB/CNB) classifier, which was designed to ensure that all data-dependent operations were performed within cross-validation folds to avoid information leakage. The baseline demographic model achieved a ROC-AUC of 0.630±0.017. Among the biomarkers analyzed, I-309 (CCL1) emerged as the most significant predictor, contributing an incremental AUC increase of +0.110. The study emphasizes the importance of distinguishing statistical association from predictive utility, highlighting I-309 as a promising candidate for further validation in predicting cognitive impairment. The findings suggest that while certain inflammatory biomarkers may be associated with cognitive decline, their predictive power varies and requires careful evaluation in clinical settings.
Methodology
The authors used a threshold-likelihood Bernoulli/Categorical Naive Bayes (BNB/CNB) classifier, performing all data-dependent operations within repeated stratified ten-fold cross-validation to prevent leakage. Continuous predictors were transformed into supervised chi-square states, and income was treated as a categorical variable.
Results
The demographic baseline model achieved a ROC-AUC of 0.630±0.017. I-309 (CCL1) was the dominant feature, showing a significant incremental AUC improvement of +0.110. The results were validated across multiple partitions, with a fixed-partition DeLong p-value of 0.0018, indicating the robustness of I-309 as a predictive feature.
Implications
The study suggests that I-309 could serve as an interpretable biomarker for predicting cognitive impairment in clinical settings, pending further validation. The methodology may also be applicable to other small clinical datasets where interpretability is crucial.
Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems
Federated Learning
Time Series
- Introduces a coalition formation approach for federated learning that adapts to client heterogeneity.
- Utilizes opinion dynamics to form coalitions based on local model weights, enhancing model compatibility.
- Demonstrates significant improvements in forecasting accuracy for water consumption compared to traditional methods.
- Achieves coalition structures with no additional client-side computation or communication overhead.
Read more
Opinion Dynamics-based Coalition Formation for Federated Learning in Heterogeneous IoT Systems
Summary
This paper addresses the challenges of federated learning (FL) in heterogeneous Internet-of-Things (IoT) systems, particularly in the context of short-term water consumption forecasting. Traditional FL methods, such as Federated Averaging (FedAvg), struggle with statistical heterogeneity, leading to suboptimal global models that do not capture client-specific patterns. The authors propose a novel approach that forms client coalitions based on local model weights using a Hegselmann–Krause (HK) bounded-confidence opinion dynamics process. This method allows for adaptive coalition formation, where the number and membership of coalitions emerge from the data rather than being predetermined. The proposed framework is evaluated against several benchmarks, including FedAvg and FedProx, demonstrating significant improvements in forecasting accuracy and stability. The results indicate that the HK-based coalition formation can reduce the mean absolute error (MAE) by up to 54% compared to FedAvg, while maintaining a high global accuracy of 83-85%.
Methodology
The authors model coalition formation as a bounded-confidence opinion dynamics process, specifically using the Hegselmann–Krause model. They develop three variants of the interaction based on Euclidean distance and cosine similarity criteria to determine compatibility among clients. The framework is instantiated for short-term water consumption forecasting using local Long Short-Term Memory (LSTM) models, and the performance is evaluated against various federated learning methods.
Results
The proposed HK-based coalition formation method produced stable coalition structures within ten iterations and significantly reduced the average MAE by up to 54% compared to FedAvg, 39% compared to FedProx, and 24% compared to Per-FedAvg. The global accuracy achieved was between 83% and 85%, indicating a strong performance in the context of heterogeneous IoT data.
Implications
This work has potential implications for improving federated learning in various IoT applications, particularly in scenarios with heterogeneous data distributions. The coalition formation approach can enhance model accuracy and stability, making it suitable for privacy-sensitive environments such as smart cities.
Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning
Multimodal
Robotics
NLP
- Uni-LaDiR introduces a unified latent space for multimodal reasoning, improving the integration of information from different modalities.
- The framework employs a unified encoder to generate shared thought tokens, enhancing the reasoning process by reducing modality-specific representation issues.
- Diffusion is used to predict the next reasoning steps, allowing for flexibility in generating multiple valid outputs.
- Joint training of the encoder and diffusion model leads to better task-relevant thought token generation.
Read more
Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning
Summary
The paper introduces Uni-LaDiR (Unified Latent Diffusion Reasoner), a novel framework designed to enhance multimodal reasoning by mapping information from various modalities into a shared latent space. Traditional methods often concatenate or interleave modality-specific tokens, which complicates the reasoning process due to the need for the model to bridge representational differences. Uni-LaDiR addresses this by employing a unified encoder that transforms teacher reasoning steps from different modalities into shared thought tokens, preserving essential information for subsequent reasoning and final outputs. The framework utilizes diffusion to predict the next block of thought tokens based on the input and preceding tokens, allowing for multiple valid reasoning paths. The joint training of the encoder and diffusion reasoner with shared weights ensures that the thought tokens are both useful for the task and predictable from the context. The model was evaluated across eleven vision-language model (VLM) benchmarks and two vision-language-action (VLA) suites, demonstrating significant improvements in reasoning accuracy and manipulation success compared to existing methods.
Methodology
The methodology involves a unified encoder that maps heterogeneous modality features into a shared latent reasoning space, generating fixed-width latent blocks as thought tokens. The model uses diffusion to predict the next thought tokens based on the input and previous tokens. Joint training of the encoder and diffusion reasoner is employed to ensure that the thought tokens retain task-relevant information and are predictable from the context.
Results
Uni-LaDiR achieved a relative improvement of 7.3% in average mathematical and logical reasoning accuracy across four VLM benchmarks and a 6.1% increase in manipulation success across ten RLBench tasks. Additionally, flow matching within the unified latent space improved performance by 14.9% over direct L2 prediction and 19.5% over cosine similarity loss.
Implications
The implications of this work suggest that separating modality-specific perception from shared latent reasoning can lead to more effective multimodal reasoning systems. This approach could be applied in various domains requiring integration of diverse data types, such as robotics, natural language processing, and computer vision.
Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification
Theory
Efficient ML
Large Language Models
- Introduces Parameter Space Orthogonality as a feasible condition for preventing catastrophic forgetting.
- Develops a post-hoc, tuning-agnostic framework (JANUS) that rectifies parameter updates after fine-tuning.
- Employs a Multi-step Adaptive Rectification mechanism to dynamically adjust parameter updates.
- Demonstrates significant improvements in knowledge recovery with minimal impact on task performance.
Read more
Past, Future, All at Once: Mitigating Stability-Plasticity Dilemma via Post-hoc JANUS Rectification
Summary
This paper addresses the stability-plasticity dilemma encountered when fine-tuning foundation models on new tasks, which often leads to catastrophic forgetting (CF). The authors critique existing methods that rely on the overly restrictive Subspace Orthogonality condition and propose a novel approach called JANUS (JAcobian NUll Space projection). This post-hoc, tuning-agnostic weight rectification framework introduces the Parameter Space Orthogonality condition, which is both necessary and sufficient for preserving historical performance. By projecting parameter updates into the JANUS, the method effectively recovers lost historical knowledge without disrupting the fine-tuning process. Additionally, the authors introduce a Multi-step Adaptive Rectification mechanism that dynamically adjusts step sizes based on the JANUS shift, ensuring efficient navigation of the parameter space. The JANUS framework integrates seamlessly with various fine-tuning methods and employs techniques like ghost projection and singular value decomposition for enhanced efficiency. Experimental results demonstrate that JANUS significantly mitigates the stability-plasticity dilemma, achieving near-perfect knowledge recovery while maintaining task adaptation.
Methodology
The authors propose the JANUS framework, which rectifies parameter updates post fine-tuning by projecting them into the Jacobian Null Space. This approach is complemented by a Multi-step Adaptive Rectification mechanism that adapts the rectification process based on a proxy metric for forgetting. Techniques such as ghost projection and singular value decomposition are utilized to enhance computational efficiency.
Results
Extensive experiments across various models and tasks show that JANUS effectively recovers historical knowledge while preserving the ability to adapt to new tasks. The method consistently outperforms existing approaches in mitigating catastrophic forgetting, achieving a significant advancement in the stability-plasticity trade-off.
Implications
The findings suggest that the JANUS framework can be widely applied to improve the performance of fine-tuning in large models across diverse applications, potentially enhancing the robustness and adaptability of AI systems in real-world scenarios.
SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption
Multimodal
- Identifies a granularity mismatch in existing batch-level gradient modulation under heterogeneous multimodal corruption.
- Proposes SAGG, which utilizes sample-level gating for unbiased gradient estimation.
- Demonstrates that SAGG-based SGD converges at a rate of O(1/√T) without residual bias from corruption.
- Provides a certified robustness framework for evaluating multimodal models.
Read more
SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption
Summary
The paper introduces Sample-Adaptive Gradient Gating (SAGG), a novel framework designed to enhance the robustness of multimodal learning systems under heterogeneous corruption. Traditional multimodal gradient balancing methods apply a uniform modulation to gradients across all samples in a batch, which fails to account for the varying levels of corruption present in individual samples. The authors demonstrate that this approach leads to an irreducible bias in gradient estimation. SAGG addresses this issue by implementing a sample-level gating mechanism that makes binary decisions to retain or discard gradients based on an online quality assessment of each sample. This method ensures unbiased gradient estimation and improves convergence rates to stationary points of the clean loss without incurring a corruption-dependent error floor. The paper also provides a certified robustness framework that connects per-modality Lipschitz constants to classification margins. Experimental results on datasets such as Kinetics-Sounds and UCF-101 show that SAGG outperforms ten existing methods, particularly in high-corruption scenarios where batch-level bias is most pronounced.
Methodology
The authors develop SAGG, which employs a binary retain-or-discard decision for each sample based on an online feature-norm quality test. This is coupled with a truncation mechanism to control variance. The theoretical analysis shows that unbiased estimation requires sample-level gating, and the convergence properties of SAGG-based SGD are established.
Results
SAGG consistently outperforms ten existing multimodal learning methods across various experiments involving Gaussian noise injection, partial modality missing, and natural contribution imbalance. The framework achieves state-of-the-art performance, particularly in high-corruption conditions where traditional methods struggle.
Implications
The findings suggest that adopting sample-level adaptive strategies can significantly enhance the robustness of multimodal learning systems in real-world applications where data corruption is heterogeneous. This could lead to improved performance in fields such as audio-visual processing, sensor fusion, and other multimodal tasks.
CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning
Reinforcement Learning
Robotics
Theory
- Introduction of CARE-VI, a framework for improving target reliability in off-policy actor-critic learning.
- Development of three components: CARS, SEVA, and DARE, which address candidate selection, value assessment, and residual regulation.
- Theoretical analysis provides error bounds for the proposed methods, ensuring fixed-policy recovery.
- Empirical results show CARE-VI achieves the highest mean return across multiple tasks and configurations.
Read more
CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning
Summary
The paper introduces CARE-VI, a framework designed to enhance the reliability of temporal-difference (TD) targets in off-policy actor-critic learning. The authors identify that direct value improvement can lead to unreliable target values due to issues such as noisy rankings and fixed enhancement weights. To mitigate these risks, they propose three key components: Conservative Adaptive Ranking and Screening (CARS), Selector-Evaluator Value Assessment (SEVA), and Dynamic Adaptive Risk-aware Enhancement (DARE). CARS manages the candidate selection process by retaining a leading prefix of candidates until a significant uncertainty gap is observed. SEVA combines a selector critic for ordering candidates with an evaluator critic for reviewing selected values, capping the reviewed value at the selector reference. DARE adjusts the residual corrections based on candidate reliability and the divergence between selector and evaluator signals. The integration of these components into CARE-VI allows for evidence-regulated target construction while maintaining existing interfaces for critic regression and actor updates. The authors provide theoretical bounds on the errors introduced by each component and demonstrate that CARE-VI achieves superior performance across various MuJoCo tasks, outperforming existing methods.
Methodology
The methodology involves the development of three main components: CARS for adaptive candidate screening, SEVA for value assessment using separate critics, and DARE for regulating residual corrections based on candidate reliability. These components are integrated into the CARE-VI framework, which maintains the existing structures for actor-critic learning while improving the reliability of target values through evidence regulation.
Results
CARE-VI demonstrated the highest mean return in all twelve experimental settings across four MuJoCo tasks when compared to baseline methods such as SAC, TD3, and TD7. The results were supported by grouped ablation studies that confirmed the contributions of each component to the overall performance.
Implications
The findings suggest that CARE-VI can significantly enhance the performance of off-policy actor-critic methods in reinforcement learning applications, particularly in continuous control tasks. The framework's ability to improve target reliability may lead to more stable and effective learning in complex environments.
Personalized Federated Hierarchical Gaussian Processes for Privacy-Preserving Modeling of Heterogeneous Distributed Systems
Federated Learning
- Introduces a hierarchical decomposition for modeling heterogeneous federated clients.
- Employs privacy-preserving federated variational inference to keep raw data local.
- Supports uncertainty-aware predictions, crucial for risk-sensitive applications.
- Demonstrates effectiveness in real-world applications like fault classification and air-quality modeling.
Read more
Personalized Federated Hierarchical Gaussian Processes for Privacy-Preserving Modeling of Heterogeneous Distributed Systems
Summary
The paper introduces Personalized Federated Hierarchical Gaussian Processes (pFedHGP), a novel framework designed for probabilistic regression and classification in heterogeneous distributed systems. The framework addresses the challenges posed by non-i.i.d. data across clients by decomposing each client's latent function into three components: a shared global component, a client-specific deviation that maintains the global kernel structure, and a flexible local residual. This hierarchical approach allows for effective modeling of diverse operational conditions while preserving data privacy. The authors employ sparse inducing-variable approximations and federated variational inference to ensure that raw data remains local, with only low-dimensional statistics being synchronized at the server level. The framework supports uncertainty-aware decision-making through full predictive distributions. In practical applications, pFedHGP achieves perfect fault classification in press tonnage monitoring with minimal labeled data and successfully recovers geographic zones in federated air-quality modeling without centralizing sensitive time series data. The paper also establishes a connection between the hierarchical model and multi-output Gaussian processes, enhancing the understanding of correlated sensor outputs.
Methodology
The methodology involves a three-level hierarchical decomposition of latent functions in Gaussian processes, integrating a global component, client-specific deviations, and local residuals. The training is conducted using a two-stage federated variational inference approach, where clients update local variational factors based on private data while the server synchronizes only compact global statistics.
Results
The pFedHGP framework achieved perfect fault classification in press tonnage monitoring using only 13.77% of labeled cycles and effectively recovered geographic zones in federated air-quality modeling without the need for centralizing station-level data. The results indicate that the model successfully captures both global trends and local variations in heterogeneous environments.
Implications
The implications of this work extend to various domains where data privacy and heterogeneity are critical, such as smart healthcare, predictive maintenance in industrial settings, and urban environmental monitoring. The framework can enhance decision-making processes in these areas by providing reliable uncertainty estimates and maintaining data locality.
A Policy Profile for Croissant: Refusal as a Property of the Dataset
Theory
- Introduces an additive policy profile for Croissant with specified evaluation semantics.
- Defines a closed operator language for dataset operations and conditions.
- Demonstrates that the profile's decisions align with existing descriptor records.
- Presents a minimal evaluation cost for implementing the profile.
Read more
A Policy Profile for Croissant: Refusal as a Property of the Dataset
Summary
This paper introduces an additive policy profile for Croissant, a machine-readable descriptor for ML datasets that utilizes JSON-LD over schema.org. The author identifies a significant gap in the existing Croissant 1.1 specification, which includes data use conditions but lacks a defined evaluation procedure for these conditions. The proposed profile allows datasets to declare the operations they permit and the conditions under which these operations are allowed, utilizing a closed set of five operators with a fully specified decision procedure. The evaluation of this profile is conducted using two distinct corpora, demonstrating that the decisions made by the profile align with the native descriptor records. The results indicate that the added evaluation cost is minimal, and the profile successfully captures the necessary semantics for automated compliance checking. The paper emphasizes that the contribution lies in the evaluation semantics rather than the vocabulary itself, providing a structured approach to dataset governance that separates caller-side and data-side authority.
Methodology
The methodology involves developing an additive policy profile that specifies evaluation semantics for dataset operations. The author employs two distinct corpora for evaluation: one from a real bioinformatics pipeline and another generated from the profile's grammar. The evaluation measures the decision-making process and timing for each operation, ensuring that the results are consistent across different descriptors.
Results
The evaluation shows that the decisions made by the profile match the native descriptor records for the datasets used, with an added evaluation cost of only 11.7 µs compared to a 119 µs decision time. The analysis of the generated corpus confirms that all valid cases agree across different decision records, reinforcing the robustness of the proposed evaluation semantics.
Implications
The implications of this work extend to enhancing dataset governance in machine learning by providing a clear framework for evaluating data use conditions. This could facilitate automated compliance checking and improve the transparency of dataset operations, ultimately contributing to responsible AI practices.
COMPASS: Ordered Clustered Routing at 100K Scale
Optimization
- COMPASS is a globally-coordinated algorithm that optimizes the OCTSP using raw distance inputs.
- The algorithm can scale to 100K synthetic nodes and 28.5K real-world e-commerce nodes, achieving state-of-the-art results.
- COMPASS avoids a quality ceiling by modeling global dependencies between clusters.
- The paper introduces a new benchmark suite for large-scale OCTSP, facilitating future research.
Read more
COMPASS: Ordered Clustered Routing at 100K Scale
Summary
The paper introduces COMPASS, an innovative algorithm designed to tackle the Ordered Clustered Traveling Salesman Problem (OCTSP), which is critical in large-scale routing scenarios where clusters of nodes must be visited in a specific order. Traditional methods often optimize clusters independently, neglecting the interdependencies between them, which can lead to suboptimal solutions. COMPASS addresses this by employing a combination of search techniques and learning-accelerated routing, utilizing parallel sub-solvers to improve solution quality without a predefined quality ceiling. The algorithm is capable of processing general distance matrices, making it versatile compared to existing solvers that rely on coordinate inputs. The authors demonstrate COMPASS's effectiveness by scaling it to 100,000 synthetic nodes and 28,500 real e-commerce nodes, achieving the largest reported routing solution over asymmetric distances, significantly surpassing previous benchmarks. The paper also presents a formal optimality guarantee for COMPASS and introduces a competitive approximate algorithm called Pairwise Slider, along with a comprehensive benchmark suite for large-scale OCTSP.
Methodology
COMPASS employs a directed graph representation of the global cluster chain, where the shortest path corresponds to the OCTSP solution. It integrates reinforcement learning techniques through an enhanced GNN architecture to compute intra-cluster solution costs. The algorithm orchestrates parallel sub-solvers to explore the solution space effectively, addressing both intra-cluster paths and inter-cluster bridges simultaneously.
Results
COMPASS consistently outperformed alternative methods in empirical tests, achieving optimal solutions for instances with up to 100K nodes and demonstrating the largest routing solution for asymmetric distances with 28.5K real e-commerce nodes. The algorithm's performance illustrates its capability to leverage compute power for improved solution costs without a quality ceiling.
Implications
The development of COMPASS has significant implications for large-scale routing applications, particularly in logistics and transportation, where optimizing routes can lead to substantial cost savings and reduced environmental impact. The benchmarks provided can serve as a foundation for further research into efficient routing algorithms and their applications in various industries.
Learning-Induced Dynamical Transition in Recurrent Neural Networks
Theory
- Learning in RNNs can induce a transition from chaotic to stable dynamics.
- A non-equilibrium dynamical mean-field theory (DMFT) is developed to analyze this transition.
- The study identifies critical parameters that separate chaotic and stable regimes during learning.
- The theory predicts the evolution of network outputs and shows quantitative agreement with simulations.
Read more
Learning-Induced Dynamical Transition in Recurrent Neural Networks
Summary
This paper explores how learning in recurrent neural networks (RNNs) can transform chaotic dynamics into stable, task-dependent behavior. The author develops a non-equilibrium dynamical mean-field theory (DMFT) to describe this transition during the learning process. The study reveals that a slow feedback-driven learning mechanism generates an evolving effective feedback strength, which facilitates a transition from chaotic to stable dynamics, characterized by a bifurcation in the DMFT solution. By deriving the two-time correlation function, the author identifies a critical feedback strength and a learning rate-dependent critical time that delineate the chaotic and stable regimes. The theory predicts the time evolution of network outputs during training and aligns well with numerical simulations, demonstrating that learning reorganizes the dynamical regime of RNNs. This work emphasizes the dynamic nature of learning, suggesting that stability emerges through gradual adaptation rather than being a fixed property of the learned structure.
Methodology
The author employs a non-equilibrium dynamical mean-field theory (DMFT) to analyze the learning dynamics of recurrent neural networks. The model incorporates feedback mechanisms and derives the two-time correlation function to track the evolution of network dynamics throughout the learning process.
Results
The study finds that the learning process leads to a bifurcation in the DMFT solution, indicating a transition from chaotic to stable dynamics. The predictions regarding the time evolution of network outputs during training are consistent with numerical simulations, validating the theoretical framework.
Implications
This research provides insights into the mechanisms by which recurrent neural networks can achieve stable computational behavior through learning. It has potential applications in understanding biological neural circuits and improving the design of artificial intelligence systems that rely on RNNs.
ZeroHAT: Behavior-Conditioned Zero-Shot Human Activity Trace Generation
Generative Models
Time Series
Optimization
- ZeroHAT generates synthetic HATs without requiring target-region data, using only source-region data and contextual information.
- The framework includes innovative components for intent extraction, behavioral cloning, and activity realization.
- ZeroHAT significantly outperforms existing methods in terms of utility and fidelity across multiple target regions.
- The approach demonstrates the effectiveness of transferring behavioral patterns across regions with different POI distributions.
Read more
ZeroHAT: Behavior-Conditioned Zero-Shot Human Activity Trace Generation
Summary
The paper presents ZeroHAT, a novel framework for generating synthetic human activity traces (HATs) in a zero-shot manner, addressing the challenges of high collection costs and privacy concerns associated with real HATs. Unlike existing methods that rely on data from the same region, ZeroHAT leverages behavioral patterns from source regions and contextual information about target regions to create realistic HATs. The framework consists of three main components: a multidimensional consistency-aware intent extractor that captures temporal, semantic, and spatial intents; a cross-region behavioral cloning module that learns region-invariant actions; and a behavior-conditioned activity realization module that constructs a dynamic action-POI graph to ground actions onto target POIs. The evaluation of ZeroHAT on a ten-city benchmark demonstrates its superior performance, achieving significantly higher normalized downstream utility and improved fidelity compared to existing baselines. This work highlights the potential of zero-shot learning in synthetic data generation for urban mobility applications.
Methodology
ZeroHAT employs a behavior-conditioned framework that includes a multidimensional intent extractor for capturing activity characteristics, a cross-region behavioral cloning module for learning invariant actions, and a dynamic action-POI graph for activity realization. It integrates these components using a consistency-guided product-of-experts to generate timestamped activities.
Results
In extensive evaluations, ZeroHAT achieved 4.5–6.4 times the normalized utility of the strongest baseline and improved average fidelity by 15.6%–40.8%. The framework demonstrated high computational efficiency, reaching approximately three times the throughput of the fastest neural baseline while using significantly less GPU memory.
Implications
The findings suggest that ZeroHAT can facilitate the generation of synthetic HATs for urban mobility applications, enhancing mobility prediction, urban simulation, and location-based services without the need for extensive real-world data collection. This approach could also inform future research on zero-shot learning in other domains.
Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization
Generative Models
Optimization
Efficient ML
- Introduces a new LSBO algorithm that significantly reduces computational overhead.
- Utilizes linear surrogates constrained to a spherical domain for efficient optimization.
- Achieves over 100× speedup compared to existing BO methods while maintaining high sample efficiency.
- Demonstrates effectiveness across molecular design and image generation tasks.
Read more
Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization
Summary
This paper addresses the challenge of making Bayesian optimization (BO) practical for de novo discovery pipelines, where generative models are used to generate candidate designs that are then filtered through virtual screens. Traditional BO methods are often too slow due to their reliance on complex surrogate models, especially when evaluations are relatively inexpensive. The authors propose a novel latent space Bayesian optimization (LSBO) algorithm that utilizes linear surrogates constrained to a spherical domain, significantly reducing computational overhead. By deriving nearly closed-form solutions for surrogate modeling and acquisition, the proposed method achieves at least a 100× speedup compared to state-of-the-art approaches while maintaining or improving performance across various benchmarks in molecular and image generation. This advancement allows BO to be effectively integrated into de novo discovery processes where it was previously impractical due to time constraints.
Methodology
The authors replace traditional nonlinear Gaussian processes with linear surrogates constrained to a spherical domain, which allows for closed-form solutions in surrogate modeling and acquisition optimization. This approach reduces the complexity of each iteration from cubic to linear in the number of samples, enabling rapid computations that complete in less than one second per iteration.
Results
The proposed LSBO method matches or exceeds the sample efficiency of existing methods on benchmarks involving latent spaces of up to 16,384 dimensions, while running over 100× faster in wall-clock time. The performance gap widens with larger observation budgets and latent dimensionality, demonstrating the method's scalability and efficiency.
Implications
This work has significant implications for the fields of drug discovery, materials science, and any domain where rapid generation and evaluation of candidate designs are critical. The efficiency of the proposed LSBO method allows for broader applications of Bayesian optimization in high-throughput settings, making it a valuable tool for researchers and practitioners in generative modeling and optimization.
Online Adaptive Kernel Mixing for Gaussian Process Decision Making
Optimization
Theory
- HACK GPs adaptively select kernels in Gaussian Processes to improve decision-making performance.
- The method treats kernel selection as an online learning problem using AdaHedge for dynamic updates.
- Two variants of HACK are introduced: Mixture of Gaussians and categorical sampling.
- The approach shows improved performance over standard kernels and ensemble methods in empirical evaluations.
Read more
Online Adaptive Kernel Mixing for Gaussian Process Decision Making
Summary
This paper introduces HACK GPs (Hedge Adaptive Cumulative Kernels), a novel approach to improve the performance of Gaussian Processes (GPs) in sequential decision-making tasks such as Bayesian optimization, level set estimation, and Bayesian active learning. The authors argue that the effectiveness of GPs is heavily reliant on the choice of kernels, and standard kernels can lead to suboptimal performance due to kernel misspecification. HACK GPs treat kernel selection as an online learning problem, where each candidate kernel is viewed as an expert. The method employs AdaHedge to dynamically update a distribution over these experts based on their predictive performance, allowing the model to adaptively focus on the most suitable kernels over time. Two variants of HACK are proposed: a mixture of Gaussians (MoG) predictive distribution and categorical sampling of a single kernel. The authors provide theoretical guarantees that the weight distribution will concentrate on the best kernel under certain conditions. Empirical results demonstrate that HACK GPs outperform standard kernels and simple ensemble methods across various tasks, showcasing robust performance and adaptability in decision-making scenarios.
Methodology
The authors propose HACK GPs, which utilize an online learning framework where each candidate kernel is treated as an expert. The AdaHedge algorithm is employed to update the weights of these experts based on their predictive performance, allowing for adaptive kernel selection. Two implementation variants are provided: a mixture of Gaussians predictive distribution and categorical sampling of a single kernel.
Results
Empirical evaluations indicate that HACK GPs consistently outperform standard kernels like Squared Exponential and Matérn-5/2, as well as simple ensemble baselines, across tasks such as Bayesian optimization, level set estimation, and Bayesian active learning. The method demonstrates robust adaptability and improved decision-making capabilities.
Implications
The proposed HACK GPs can significantly enhance the performance of Gaussian Processes in various applications requiring sequential decision-making, such as optimization problems in engineering and machine learning. This approach could lead to more efficient and effective use of GPs in real-world scenarios where kernel misspecification is a concern.
Sharp Reconstruction Bounds for Autoencoders Using the Same Forward Map
Theory
- Establishes sharp reconstruction bounds for autoencoders using the same forward map.
- Derives a reconstruction-derivative error bound based on Jacobian singular values.
- Demonstrates that affine maps can achieve the derived bounds at any depth.
- Validates theoretical predictions with empirical results from a large LiDAR dataset.
Read more
Sharp Reconstruction Bounds for Autoencoders Using the Same Forward Map
Summary
This paper investigates the reconstruction capabilities of autoencoders that utilize the same forward map for both encoding and decoding processes. The authors focus on scenarios where observed coordinates are set to zero, and they derive sharp bounds on reconstruction errors for autoencoders with equal odd input and hidden dimensions (d ≥ 3). The main contribution is the establishment of a reconstruction-derivative error bound, which is expressed as max{1 - M(M - m)/2, 0}, where M and m are the maximum and minimum singular values of the Jacobian of the transformation. The authors demonstrate that affine maps can achieve this bound at any specified depth. They also explore the implications of using a translated radial rotation, which can reconstruct inputs within a specified ball exactly, even with singular values close to one. The theoretical predictions are validated through experiments on a large terrestrial LiDAR forest scan dataset, showing that the mean theoretical bound closely aligns with the mean normalized training error. The results indicate that adding one hidden coordinate significantly reduces reconstruction error, highlighting the importance of dimensionality in autoencoder performance.
Methodology
The authors analyze the geometric properties of autoencoders that apply the same learned forward map for both encoding and decoding. They derive mathematical bounds on reconstruction errors using properties of orientation-preserving diffeomorphisms and their Jacobians. The analysis includes theoretical proofs and empirical validation through experiments on a large dataset.
Results
The study finds that the least uniform reconstruction-derivative error is max{1 - M(M - m)/2, 0}, with affine maps achieving this bound. In experiments with a 798,452-point LiDAR forest scan, the mean theoretical bound at input scale 0.05 was 0.155, which is about 84% of the mean normalized training error of 0.185. Adding one hidden coordinate reduced the mean reconstruction error to below 6 × 10^-6.
Implications
The findings suggest that careful design of autoencoder architectures, particularly regarding the dimensionality of hidden layers, can lead to significant improvements in reconstruction performance. This has potential applications in fields requiring accurate data representation and reconstruction, such as computer vision and remote sensing.
QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization
Time Series
Efficient ML
Optimization
- QUALS addresses data diversity issues in time series forecasting.
- The framework includes pattern quantization and learnability synchronization mechanisms.
- Models trained on QUALS achieve superior zero-shot performance with less data.
- The approach effectively manages skewed pattern distributions and learnability discrepancies.
Read more
QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization
Summary
The paper introduces QUALS, a novel framework aimed at improving the efficiency of time series forecasting by addressing the challenges posed by data diversity and distribution. Current approaches to universal forecasting often focus on model architecture while neglecting the complexities of the data itself, leading to suboptimal performance. QUALS enhances data efficiency by employing two main mechanisms: a pattern quantization framework that decodes heterogeneous patterns from mixed corpora using vector quantization and uniform binning, and a learnability synchronization framework that calibrates sampling weights for these patterns. This dual approach helps bridge the optimization gap between simple and complex motifs, allowing models to achieve superior zero-shot forecasting performance with significantly reduced training data. Extensive benchmarks demonstrate that pre-training on QUALS consistently yields better outcomes compared to traditional methods, even with a fraction of the original training data.
Methodology
QUALS employs a two-pronged approach: first, it utilizes a pattern quantization framework to systematically decode and categorize diverse temporal patterns from large datasets. Second, it implements a learnability synchronization framework that adjusts sampling weights for different patterns, ensuring that both simple and complex motifs are adequately represented during training.
Results
The results indicate that models pre-trained on the QUALS framework consistently outperform those trained on traditional large-scale datasets, achieving better zero-shot forecasting capabilities even when trained on significantly reduced data volumes. This demonstrates the effectiveness of QUALS in enhancing training efficiency and model performance.
Implications
The findings suggest that QUALS could be applied across various domains that rely on time series data, such as transportation, finance, and public safety, potentially leading to more robust forecasting models that require less data and computational resources.
How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU?
NLP
Large Language Models
- Rubric-decomposed prompting outperforms holistic prompting in most cases.
- The mapping of trait scores to prompt scores is fragile and requires careful calibration.
- Essay length impacts scoring accuracy, with smaller models showing variability in performance.
- The best local model configuration achieved a QWK of 0.388, below human scoring standards.
Read more
How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU?
Summary
This paper investigates the capabilities of sub-3B open language models in performing zero-shot essay scoring on a consumer-grade GPU, specifically focusing on the constraints of local inference without API calls. The authors conduct a controlled study using four instruction-tuned models from two families, evaluating their performance on eight prompts from the ASAP-AES dataset. The study compares two prompting strategies: holistic prompting and rubric-decomposed prompting, assessing their effectiveness in generating scores. Key findings indicate that rubric-decomposed prompting generally outperforms holistic prompting, particularly under batch min-max aggregation. The research also reveals that the mapping of trait scores to prompt scores is sensitive to grader calibration, and essay length significantly affects scoring accuracy. The best-performing local model configuration achieved a macro quadratic weighted kappa (QWK) of 0.388, which, while notable, remains below human inter-rater reliability and a length-only baseline. The authors conclude that sub-3B models are suitable for formative feedback in educational contexts but should not replace human scoring.
Methodology
The study employed a systematic evaluation of four instruction-tuned models (Qwen2.5 and SmolLM2) across eight ASAP-AES prompts. It utilized two prompting strategies (holistic and rubric-decomposed) and analyzed the results using bootstrap confidence intervals, Holm-corrected paired tests, and length-bias assessments, all conducted on a single 8 GB consumer GPU.
Results
The findings showed that rubric-decomposed prompting generally yielded better results than holistic prompting, particularly under batch min-max aggregation. The best local model configuration achieved a macro QWK of 0.388, indicating significant room for improvement compared to human inter-rater reliability (0.769) and a length-only baseline (0.523). Additionally, the study highlighted a systematic decline in signed error with increasing essay length.
Implications
The results suggest that sub-3B open language models can provide useful formative feedback in educational settings, particularly where privacy concerns prevent the use of third-party APIs. However, their limitations indicate that they should complement rather than replace human evaluators in essay scoring.
When Does Retrieval Help Time-Series Forecasting?
Time Series
- Retrieval benefits in time-series forecasting are primarily determined by the ratio of lookback window length to dominant seasonal period.
- A simple control method can outperform complex retrieval plug-ins under certain conditions, particularly when the seasonal structure is strong.
- The study introduces a regime map and two statistics to predict the effectiveness of retrieval mechanisms before deployment.
- Retrieval mechanisms are less effective when the training data lacks a concentrated seasonal structure.
Read more
When Does Retrieval Help Time-Series Forecasting?
Summary
This paper investigates the conditions under which retrieval mechanisms enhance time-series forecasting performance. The authors argue that the benefits of retrieval depend on the relationship between the lookback window length (S) and the dominant seasonal period (L). They demonstrate that a simple control method, which repeats the last observed seasonal period, can outperform complex retrieval plug-ins in certain conditions. The study reveals that at a short window length (S=12), this control method reduces mean squared error (MSE) significantly across various benchmarks, while retrieval mechanisms are less effective when the training data lacks a concentrated seasonal structure. The authors introduce a regime map that helps identify when retrieval is beneficial and propose two interpretable statistics to predict the effectiveness of retrieval before deployment. The findings suggest that the effectiveness of retrieval is closely tied to the phase of the seasonal cycle, highlighting the importance of understanding the underlying temporal structure in time-series data.
Methodology
The authors conducted a series of experiments varying the lookback window length (S), dominant seasonal period (L), and forecasting horizon (H) to evaluate the performance of retrieval mechanisms against a simple control method. They analyzed the mean squared error (MSE) across multiple benchmarks and introduced a regime map to categorize the conditions under which retrieval is beneficial. Additionally, they developed two interpretable statistics to assess the potential effectiveness of retrieval before deployment.
Results
The results indicate that at S=12, the simple control method reduces MSE by 8% to 44% on four out of seven benchmarks compared to standard backbones. It outperformed the strongest retrieval plug-in on the ETTm1 dataset and matched its performance on ECL. However, it performed worse on datasets lacking a concentrated seasonal structure, increasing MSE by up to 25%. The correlation between retrieval benefit and the dominant period was found to be +0.71, while the correlation with the forecasting horizon was -0.23.
Implications
The findings suggest that understanding the relationship between lookback windows and seasonal periods can significantly improve time-series forecasting strategies. The proposed regime map and predictive statistics can aid practitioners in selecting appropriate retrieval mechanisms based on the characteristics of their data, potentially leading to more effective forecasting models.
Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation
Reinforcement Learning
Robotics
Theory
- RSIQL improves offline GCRL by providing additional reward signals at intermediate states.
- The method does not require a hierarchical policy, simplifying the learning process.
- Experiments show significant performance improvements over existing methods.
- The approach addresses the issue of delayed goal-completion supervision effectively.
Read more
Improving Offline Goal-Conditioned Reinforcement Learning via Selective Reward Stimulation
Summary
This paper addresses the challenges of offline goal-conditioned reinforcement learning (GCRL), particularly in scenarios with sparse rewards and long-horizon dependencies. The author identifies that goal-completion information can be temporally distant from the actions that lead to success, complicating the learning process. To tackle this issue, the paper introduces Reward Stimulation Implicit Q-Learning (RSIQL), a non-hierarchical method that enhances training supervision by applying additional reward signals at intermediate states deemed to contribute to goal progress. This approach contrasts with hierarchical methods that typically require learning a separate high-level subgoal policy. RSIQL utilizes an auxiliary goal-conditioned value function to identify these intermediate states and applies reward stimulation accordingly. The experiments conducted on D4RL goal-reaching benchmarks and OGBench demonstrate that RSIQL outperforms traditional goal-conditioned IQL and achieves competitive performance compared to hierarchical methods while maintaining a simpler policy structure.
Methodology
The paper proposes RSIQL, which employs an auxiliary goal-conditioned value function to identify intermediate states that indicate progress toward the final goal. It applies reward stimulation at these states to enhance the training signal, thereby improving the learning of the policy and value functions using an IQL-style offline reinforcement learning objective.
Results
The experimental results indicate that RSIQL consistently outperforms traditional goal-conditioned IQL methods and achieves performance on par with hierarchical offline goal-conditioned methods, demonstrating its effectiveness in improving learning from offline datasets.
Implications
The findings suggest that enhancing training signals at informative intermediate states can significantly improve offline GCRL performance, which has implications for various applications in robotics, navigation, and autonomous systems where online exploration is impractical.
SoftTri: Smooth Triangular Membership Functions for Adaptive Fuzzy Inference Systems
Optimization
Interpretability
Theory
- Introduction of SoftTri, a differentiable triangular membership function that enhances optimization stability.
- Closed-form analytical gradients derived for efficient backpropagation in training.
- SoftTri maintains the interpretability and locality of classical triangular MFs while providing smoothness.
- Experimental results show improved performance over classical triangular MFs and comparable results to Gaussian MFs.
Read more
SoftTri: Smooth Triangular Membership Functions for Adaptive Fuzzy Inference Systems
Summary
This paper introduces SoftTri, a novel differentiable triangular membership function designed to enhance the performance of adaptive fuzzy inference systems (FIS). Traditional triangular membership functions (MFs) are popular due to their interpretability and simplicity; however, their nondifferentiability at knot points poses challenges for gradient-based optimization in adaptive neuro-fuzzy architectures. SoftTri addresses this limitation by employing a smooth soft-hinge mechanism inspired by Swish-type activations, which maintains the geometric structure and localized behavior of classical triangular MFs while ensuring C∞ smoothness with respect to both input variables and membership parameters. The authors derive closed-form analytical gradients for SoftTri, facilitating efficient backpropagation-based learning. The proposed membership function is integrated into a Takagi–Sugeno fuzzy neural network and evaluated across various benchmarks, including one-dimensional and two-dimensional nonlinear approximation tasks, as well as a real-world regression problem using the Airfoil Self-Noise dataset. The experimental results indicate that SoftTri significantly enhances optimization stability and approximation accuracy compared to classical triangular MFs, achieving performance levels comparable to or superior to Gaussian MFs under the same conditions. This work presents a valuable compromise between interpretability and differentiable optimization in modern neuro-fuzzy learning systems.
Methodology
The authors developed SoftTri by replacing hard hinge operations with a differentiable soft-hinge construction. They derived closed-form analytical gradients for the function, enabling efficient backpropagation. The function was integrated into a Takagi–Sugeno fuzzy neural network and tested on various benchmarks to evaluate its performance against classical triangular and Gaussian MFs.
Results
SoftTri demonstrated consistent improvements in optimization stability and approximation accuracy compared to classical triangular membership functions. In experiments, it achieved performance levels comparable to or better than Gaussian MFs, indicating its effectiveness in enhancing gradient-based optimization in neuro-fuzzy systems.
Implications
The introduction of SoftTri could lead to more effective and interpretable fuzzy inference systems, particularly in applications requiring adaptive learning and optimization. This approach may enhance the deployment of fuzzy systems in real-world scenarios where interpretability and performance are critical.
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
Large Language Models
Efficient ML
NLP
- Candidate count alone does not adequately describe the system costs of multi-candidate inference.
- Increasing the number of candidates improves accuracy but also increases energy consumption and latency.
- Batched generation calls are significantly more efficient than serial execution of candidates.
- The paper provides recommendations for reporting practices in multi-candidate inference studies.
Read more
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
Summary
This paper investigates the impact of candidate-generation strategies on the performance and energy efficiency of large language models (LLMs) during test-time scaling. The authors highlight that while the number of generated candidates (N) is often used to measure inference budgets, it does not account for how these candidates are executed. The study evaluates the effect of increasing N on reasoning accuracy using two models, Phi-3-mini and Qwen2.5-1.5B, across 500 GSM8K prompts, finding significant accuracy improvements with higher N. However, the authors emphasize that simply increasing N leads to increased system costs, which are not captured by accuracy metrics alone. They compare four different generation schedules (1×8, 2×4, 4×2, and 8×1) while keeping N fixed at 8, measuring latency, throughput, GPU-hours, and energy consumption. Results indicate that serial execution of candidates incurs significantly higher energy costs and latency compared to batched generation. The findings suggest that fewer generation calls with larger batch sizes are more efficient when candidates are independent, and they call for improved reporting practices in the field to include generation schedules and system metrics.
Methodology
The authors conducted experiments with two LLMs, Phi-3-mini and Qwen2.5-1.5B, using a fixed candidate count of N=8. They compared different generation schedules (1×8, 2×4, 4×2, and 8×1) and measured various performance metrics including latency, throughput, GPU-hours, and energy consumption on A100 GPUs.
Results
The study found that increasing the candidate count from 1 to 8 improved accuracy by 8.4 percentage points for Phi-3-mini and 18.4 points for Qwen2.5-1.5B. However, using eight serial calls resulted in 4.64–4.86 times more energy consumption and 5.77–6.12 times higher latency compared to a single batched call with eight candidates. This pattern was consistent across different GPU nodes and workloads.
Implications
The findings suggest that optimizing candidate generation strategies can lead to significant improvements in both performance and energy efficiency of LLMs. This has implications for the deployment of LLMs in high-performance computing environments, where resource efficiency is critical. The recommendations for reporting practices can enhance reproducibility and comparability of results in future research.
Subdomain-aware representation compression for pretrained image embeddings
Computer Vision
Efficient ML
- Dimensionality reduction techniques can be effectively tailored for subdomain representation compression.
- PCA and LDA show significant improvements in both space efficiency and accuracy for pretrained image embeddings.
- The approach allows for effective transfer learning capabilities using compressed representations.
- Experiments demonstrate that compression can outperform traditional full-embedding methods.
Read more
Subdomain-aware representation compression for pretrained image embeddings
Summary
This paper explores the use of dimensionality reduction techniques, specifically Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA), for compressing pretrained image embeddings in a subdomain-aware manner. The authors argue that traditional dimensionality reduction methods, typically applied uniformly across datasets, can be adapted to focus on subdomains, leading to improved space efficiency and computational performance without sacrificing accuracy. The study involves experiments in both unsupervised and supervised settings, demonstrating that the proposed compression methods can effectively extract important features relevant to specific tasks. Additionally, the paper investigates the transfer learning capabilities of the compressed representations, showing that they can be beneficial even when trained on different but related subdomains. The findings indicate that the proposed methods not only reduce resource requirements but also enhance downstream task performance, making them particularly suitable for edge-device machine learning applications.
Methodology
The authors employed standard dimensionality reduction techniques (PCA and LDA) to compress pretrained image embeddings. They conducted experiments in both unsupervised and supervised settings, using sampled embeddings from subdomains to learn the compression transformation. The effectiveness of the compression was evaluated through clustering tasks and transfer learning scenarios.
Results
The results indicated that the proposed compression methods led to improved space efficiency and enhanced performance in downstream tasks compared to traditional full-embedding approaches. The compression ratios achieved were significantly better than typical usage, and the methods demonstrated strong transfer learning capabilities.
Implications
The findings suggest that the proposed dimensionality reduction techniques can facilitate the deployment of machine learning models on edge devices, allowing for efficient local processing of data without the need for constant access to large remote models. This could enhance privacy and reduce computational costs in various applications.